All Products
Search
Document Center

Vector Retrieval Service for Milvus:Build a personal knowledge Q&A application

Last Updated:Sep 01, 2026

In this tutorial, you use the Alibaba Cloud Milvus knowledge base to build a personal knowledge Q&A application. You import documents, publish a version, and search the data through the SDK, and then connect a large language model (LLM) to generate answers on a local web page. By the end, you will have a runnable Q&A page that answers questions only from your own documents.

Solution overview

The end-to-end workflow has three stages: create a knowledge base and import data in the console, publish a version to make the content searchable, and then search through the software development kit (SDK) and pass the results to an LLM to generate answers.

Data import and search can be done through the SDK, but version publishing must be done in the console. No corresponding OpenAPI operation is available. Plan your automation accordingly before you start.

Parsing and chunking 10 to 100 typical documents usually takes 10 to 20 minutes in total before the data becomes searchable.

Prerequisites

  • An Alibaba Cloud account. The knowledge base is available in the following regions: China (Hangzhou), China (Beijing), China (Zhangjiakou), and China (Shenzhen).

  • RAM authorization for the calling account. Log on to the RAM console with your Alibaba Cloud account, create a RAM user, select Access using a permanent AccessKey, and then attach the system policy AliyunMilvusFullAccess to the user.

    Creating a RAM user with an AccessKey triggers security authentication (a multi-factor authentication (MFA) code, an SMS verification code, or a face scan). The AccessKey secret is shown only once after creation. Click Save results on the page to download and store it. For stricter least privilege, create a separate custom policy that contains only milvusknowledgebase:* and use it instead of the system policy above.
  • An LLM API that supports the OpenAI chat/completions protocol, for example, the OpenAI-compatible endpoint and API key of Alibaba Cloud Model Studio.

  • Python 3.9 or later installed on your local machine.

    Store your AccessKey pair and LLM API key only in local environment variables. Never write them into documents, chat logs, screenshots, or code repositories.

Step 1: Create a knowledge base

  1. Log on to the Alibaba Cloud Milvus console, select the target region from the region drop-down list at the top, and then click Knowledge Base Service in the left-side navigation pane.

  2. On the Knowledge Base page, click Create Knowledge Base and configure the following settings.

    The Data type and Embedding model settings cannot be changed after creation. This tutorial uses Omni-modal knowledge base and Built-in model. Confirm both selections before you continue.
    Configuration itemDescription
    Name2 to 64 characters. Chinese characters are not supported. The name must be unique within the same tenant.
    Data typeValid values: Omni-modal knowledge base and Structured knowledge base (chunks are generated per table row; only xls, xlsx, and csv are supported). This setting cannot be changed after creation. This topic uses Omni-modal knowledge base, which supports mp4, mp3, png, pdf, ppt, txt, markdown, docs, xlsx, csv, jsonl, and faq.
    Embedding modelValid values: Built-in model and Private/External model. This setting cannot be changed after creation. This topic uses Built-in model, which defaults to text-embedding-v3, and Vector dimensions automatically shows 1024. When you select Private/External model, Vector dimensions shows - and cannot be edited.
    SpecificationRequired. The minimum specification is currently 4 CU (4 vCPU / 16 GiB). You can adjust it on the details page after creation.
    Network settingsRequired. Specify a VPC and a vSwitch. You can also create them on the spot.

    The chunking strategy is not configured on the creation page. In the Processing strategy area of the knowledge base details page, you can view the Default strategy provided by the system (intelligent splitting with a maximum length of 512 characters), or click Create strategy to create a custom one. When you register documents through the SDK, you can specify the strategy with strategy_id.

  3. Click Create Knowledge Base and wait until the knowledge base status changes from Creating to Running.

  4. In the knowledge base list, note the Knowledge base ID (in a format such as kd-803ae9b10cc31). This ID is required for subsequent SDK calls.

Step 2: Prepare data

Start with a few Markdown, TXT, PDF, or Word documents to validate the results. File names should describe the document topic, and the body should contain a complete title and context.

If you do not have suitable documents, you can use the public Chinese legal case dataset LeCaRDv2. Legal cases contain specific case numbers, case facts, and judgments, which make it easier to verify whether an answer really comes from the knowledge base, compared with general knowledge. The dataset is used only to demonstrate retrieval technology and does not constitute legal advice.

Create a local directory named documents and copy your files into it. The sample code in the following steps reads files from ./documents.

Step 3: Set up the local environment

  1. Create a virtual environment and install the dependencies.

    python3 -m venv .venv
    source .venv/bin/activate
    pip install alibabacloud_milvusknowledgebase20260604 requests streamlit
  2. Configure the runtime parameters as environment variables. KB_ID is the knowledge base ID recorded in Step 1, and KB_REGION is the region of the knowledge base.

export ALIBABA_CLOUD_ACCESS_KEY_ID="YOUR_ACCESS_KEY_ID"
export ALIBABA_CLOUD_ACCESS_KEY_SECRET="YOUR_ACCESS_KEY_SECRET"
export KB_REGION="cn-hangzhou"
export KB_ID="kd-xxxxxxxxxxxxx"
export KB_VERSION="LATEST_PUBLISHED"

export LLM_BASE_URL="https://your-openai-compatible-api.example.com/v1"
export LLM_API_KEY="YOUR_LLM_TOKEN"
export LLM_MODEL="YOUR_MODEL_NAME"

Step 4: Write the sample code

Save the following content as kb_demo.py. This single module is the complete sample code for this tutorial. It contains four parts: client initialization, document upload, search, and LLM Q&A. The following steps call the functions defined here.

from __future__ import annotations

import os
from pathlib import Path

import requests
from alibabacloud_milvusknowledgebase20260604.client import Client
from alibabacloud_milvusknowledgebase20260604 import models as milvus_kb_models
from alibabacloud_tea_openapi import models as open_api_models

REGION = os.getenv("KB_REGION", "cn-hangzhou")
KB_ID = os.environ["KB_ID"]
KB_VERSION = os.getenv("KB_VERSION", "LATEST_PUBLISHED")
ENDPOINT = f"milvusknowledgebase.{REGION}.aliyuncs.com"

client = Client(open_api_models.Config(
    access_key_id=os.environ["ALIBABA_CLOUD_ACCESS_KEY_ID"],
    access_key_secret=os.environ["ALIBABA_CLOUD_ACCESS_KEY_SECRET"],
    endpoint=ENDPOINT,
    region_id=REGION,
    connect_timeout=10_000,
    read_timeout=60_000,
))

def upload_document(file_path: str) -> dict:
    path = Path(file_path).resolve()
    size = path.stat().st_size

    presigned = client.get_knowledge_base_pre_signed_url(
        KB_ID,
        milvus_kb_models.GetKnowledgeBasePreSignedUrlRequest(
            knowledge_base_id=KB_ID,
            documents=[
                milvus_kb_models.GetKnowledgeBasePreSignedUrlRequestDocuments(
                    path=path.name, name=path.name, size=size,
                )
            ],
            expires_in=3600,
        ),
    )
    upload_url = presigned.body.data.pre_signed_urls[0]

    # The presigned URL is signed for an empty Content-Type. Do not include a Content-Type header in the PUT request.
    with path.open("rb") as source:
        response = requests.put(upload_url, data=source, timeout=120)
        response.raise_for_status()

    added = client.add_documents(
        KB_ID,
        milvus_kb_models.AddDocumentsRequest(
            knowledge_base_id=KB_ID,
            import_type="LOCAL_UPLOAD",
            documents=[
                milvus_kb_models.AddDocumentsRequestDocuments(
                    path=path.name, name=path.name, size=size,
                )
            ],
            dedup=milvus_kb_models.AddDocumentsRequestDedup(
                doc_name_dedup=True, content_dedup=False,
            ),
        ),
    )
    return added.body

def search_knowledge_base(query: str, page_size: int = 6):
    resp = client.search_knowledge_base(
        KB_ID,
        milvus_kb_models.SearchKnowledgeBaseRequest(
            query=query,
            version=KB_VERSION,
            page_number=1,
            page_size=page_size,
            retrieval_config=milvus_kb_models.SearchKnowledgeBaseRequestRetrievalConfig(
                candidate_count=48,
                min_score=0,
                semantic_weight=0.5,
                enable_query_expansion=False,
            ),
        ),
    )
    return resp.body

def chat(messages: list[dict[str, str]]) -> str:
    base_url = os.environ["LLM_BASE_URL"].rstrip("/")
    response = requests.post(
        f"{base_url}/chat/completions",
        headers={"Authorization": f"Bearer {os.environ['LLM_API_KEY']}"},
        json={
            "model": os.environ["LLM_MODEL"],
            "messages": messages,
            "temperature": 0.1,
        },
        timeout=60,
    )
    response.raise_for_status()
    return response.json()["choices"][0]["message"]["content"].strip()

def ask(question: str) -> tuple[str, object]:
    search_query = chat([
        {
            "role": "system",
            "content": "Rewrite the user's question into a self-contained query suitable for knowledge base search. Output only the query.",
        },
        {"role": "user", "content": question},
    ])
    search_result = search_knowledge_base(search_query)
    results = search_result.results or []

    context = "\n\n".join(
        f"[Source {i}] {item.document_name or ''}\n{item.content}"
        for i, item in enumerate(results, start=1)
    )
    answer = chat([
        {
            "role": "system",
            "content": "Answer only based on the given materials. Cite materials as [Source N]. Clearly state so when the materials are insufficient.",
        },
        {"role": "user", "content": f"Question: {question}\n\nMaterials:\n{context}"},
    ])
    return answer, search_result

Step 5: Upload data

An upload has three steps: get a presigned URL, write the file to Object Storage Service (OSS), and register the document. upload_document() wraps the whole process.

Upload a single document:

from kb_demo import upload_document

result = upload_document("./documents/example.md")
print(result.request_id)

For batch uploads, start with groups of 10 documents:

from pathlib import Path
from kb_demo import upload_document

for path in sorted(Path("./documents").glob("*.md")):
    upload_document(str(path))
    print("Submitted:", path.name)

The presigned URL is signed for an empty Content-Type. The PUT request must not include a Content-Type request header. Otherwise, OSS returns 403 SignatureDoesNotMatch. If a PUT fails for network reasons, retry it directly.

Putting a file to OSS alone does not add it to the knowledge base. You must call AddDocuments to complete the registration. After a successful registration, run in the response is RUNNING, which means parsing and chunking are in progress. For 10 to 100 typical documents, this processing usually takes 10 to 20 minutes in total. Wait until the status becomes Processed before you publish a version.

When you register documents, both path and name take only the file name. Do not use the object path in the presigned URL. An internal path such as direct_upload/… returns 400 Path must not use an internal storage prefix. Leaving path empty returns 400 Path can't be empty. In a successful registration response, data.documentsmay be an empty array. This does not mean the registration failed. Verify the result with Documents or Chunks on the knowledge base details page, or with subsequent search results.

After the upload completes, check the data status and chunk count on the Data management page in the console. Publish a version only after the status becomes Processed. Click View chunks to preview the chunking results.

Step 6: Publish a version

Imported data stays unpublished and is not searchable until a version is published.

  1. On the knowledge base details page, click the Version management tab. Confirm that the page shows pending changes, and then click Publish version.

  2. In the Confirm changes step, review the change log and click Next.

  3. In the Enter description step, enter a publish description (up to 200 characters) and click Publish. Wait until the message Version published appears.

  4. In Version history, check the version number. The first published version is v1, and the status is Published.

    After publishing, update the local environment variable:
export KB_VERSION="v1"

Step 7: Search the data

Call search_knowledge_base() to search the published version:

from kb_demo import search_knowledge_base

result = search_knowledge_base("What conditions must be met to terminate a contract?")

for item in result.results or []:
    print(item.document_name, item.score)
    print(item.content[:300])

For version, use an explicit version number such as v1 or v2. You can also use LATEST_PUBLISHED to search the latest published version. Do not use DRAFT to serve stable Q&A externally.

In this example, min_score=0 also returns chunks with relatively low relevance. In a production environment, raise min_score or reduce page_size based on actual results to keep irrelevant content out of the LLM context.

You can also tune TopK, the semantic search weight, the reranking model, and the filter scope on the Search validation page in the console to quickly compare retrieval results.

Step 8: Connect the LLM for Q&A

ask() first has the LLM rewrite the question into a query suitable for search, then calls SearchKnowledgeBase, and finally asks the LLM to summarize an answer based only on the returned chunks.

from kb_demo import ask

answer, search_result = ask("What do these documents say about contract termination?")
print(answer)
print("Request ID:", search_result.request_id)

Step 9: Build the web application

  1. Save the following content as app.py.

    import streamlit as st
    from kb_demo import ask
    
    st.set_page_config(page_title="Knowledge Base Q&A")
    st.title("Knowledge Base Q&A")
    
    question = st.chat_input("Enter your question")
    if question:
        with st.chat_message("user"):
            st.write(question)
        with st.chat_message("assistant"):
            with st.spinner("Searching the knowledge base..."):
                answer, result = ask(question)
            st.write(answer)
            with st.expander("View sources"):
                for item in result.results or []:
                    st.markdown(f"**{item.document_name or 'Untitled document'}**")
                    st.caption(f"score: {item.score}  ·  Request ID: {result.request_id}")
                    st.write(item.content)
  2. Start the application.

streamlit run app.py

The browser automatically opens the local knowledge base Q&A page.

FAQ

SymptomCause and solution
An API call returns 403 Permission denied by RAM authenticationThe calling account does not have knowledge base permissions. Attach AliyunMilvusFullAccess to the RAM user and retry after the authorization takes effect.
A PUT upload returns 403 SignatureDoesNotMatchThe request includes a Content-Type request header, but the presigned URL is signed for an empty Content-Type. Remove the header and retry. For details, see Step 5: Upload data.
A search returns 404 Knowledge base version ... does not existThe knowledge base has no published version, or KB_VERSION does not match the actual version number. Publish a version in the console first, and then set KB_VERSION to the version number or LATEST_PUBLISHED.
A PUT upload occasionally fails to connectThe OSS connection fluctuates. Retry the file directly.
Knowledge Base Service is not shown in the left-side navigation pane of the consoleThe current region does not support the feature, or the menu has not finished loading. Switch to a supported region and check again.

Pre-launch checklist

  • Use a dedicated RAM user with least privilege that supports key rotation. Do not use the AccessKey pair of your Alibaba Cloud account for the long term.

  • Upload only data that you are authorized to process.

  • Every answer must be traceable to the expanded source chunks.

  • When you deploy the page publicly, add authentication and HTTPS, and mask sensitive information in access logs.

  • Treat the product documentation and the actual console behavior as the source of truth for API availability and permission requirements.