In this tutorial, you use the Alibaba Cloud Milvus knowledge base to build a personal knowledge Q&A application. You import documents, publish a version, and search the data through the SDK, and then connect a large language model (LLM) to generate answers on a local web page. By the end, you will have a runnable Q&A page that answers questions only from your own documents.
Solution overview
The end-to-end workflow has three stages: create a knowledge base and import data in the console, publish a version to make the content searchable, and then search through the software development kit (SDK) and pass the results to an LLM to generate answers.
Data import and search can be done through the SDK, but version publishing must be done in the console. No corresponding OpenAPI operation is available. Plan your automation accordingly before you start.
Parsing and chunking 10 to 100 typical documents usually takes 10 to 20 minutes in total before the data becomes searchable.
Prerequisites
An Alibaba Cloud account. The knowledge base is available in the following regions: China (Hangzhou), China (Beijing), China (Zhangjiakou), and China (Shenzhen).
RAM authorization for the calling account. Log on to the RAM console with your Alibaba Cloud account, create a RAM user, select Access using a permanent AccessKey, and then attach the system policy
Creating a RAM user with an AccessKey triggers security authentication (a multi-factor authentication (MFA) code, an SMS verification code, or a face scan). The AccessKey secret is shown only once after creation. Click Save results on the page to download and store it. For stricter least privilege, create a separate custom policy that contains onlyAliyunMilvusFullAccessto the user.milvusknowledgebase:*and use it instead of the system policy above.
An LLM API that supports the OpenAI
chat/completionsprotocol, for example, the OpenAI-compatible endpoint and API key of Alibaba Cloud Model Studio.Python 3.9 or later installed on your local machine.
Store your AccessKey pair and LLM API key only in local environment variables. Never write them into documents, chat logs, screenshots, or code repositories.
Step 1: Create a knowledge base
Log on to the Alibaba Cloud Milvus console, select the target region from the region drop-down list at the top, and then click Knowledge Base Service in the left-side navigation pane.
On the Knowledge Base page, click Create Knowledge Base and configure the following settings.
The Data type and Embedding model settings cannot be changed after creation. This tutorial uses Omni-modal knowledge base and Built-in model. Confirm both selections before you continue.Configuration item Description Name 2 to 64 characters. Chinese characters are not supported. The name must be unique within the same tenant. Data type Valid values: Omni-modal knowledge base and Structured knowledge base (chunks are generated per table row; only xls,xlsx, andcsvare supported). This setting cannot be changed after creation. This topic uses Omni-modal knowledge base, which supportsmp4,mp3,png,pdf,ppt,txt,markdown,docs,xlsx,csv,jsonl, andfaq.Embedding model Valid values: Built-in model and Private/External model. This setting cannot be changed after creation. This topic uses Built-in model, which defaults to text-embedding-v3, and Vector dimensions automatically shows 1024. When you select Private/External model, Vector dimensions shows-and cannot be edited.Specification Required. The minimum specification is currently 4 CU (4 vCPU / 16 GiB). You can adjust it on the details page after creation. Network settings Required. Specify a VPC and a vSwitch. You can also create them on the spot. The chunking strategy is not configured on the creation page. In the Processing strategy area of the knowledge base details page, you can view the Default strategy provided by the system (intelligent splitting with a maximum length of 512 characters), or click Create strategy to create a custom one. When you register documents through the SDK, you can specify the strategy with
strategy_id.Click Create Knowledge Base and wait until the knowledge base status changes from Creating to Running.
In the knowledge base list, note the Knowledge base ID (in a format such as
kd-803ae9b10cc31). This ID is required for subsequent SDK calls.
Step 2: Prepare data
Start with a few Markdown, TXT, PDF, or Word documents to validate the results. File names should describe the document topic, and the body should contain a complete title and context.
If you do not have suitable documents, you can use the public Chinese legal case dataset LeCaRDv2. Legal cases contain specific case numbers, case facts, and judgments, which make it easier to verify whether an answer really comes from the knowledge base, compared with general knowledge. The dataset is used only to demonstrate retrieval technology and does not constitute legal advice.
Create a local directory named documents and copy your files into it. The sample code in the following steps reads files from ./documents.
Step 3: Set up the local environment
Create a virtual environment and install the dependencies.
python3 -m venv .venv source .venv/bin/activate pip install alibabacloud_milvusknowledgebase20260604 requests streamlitConfigure the runtime parameters as environment variables.
KB_IDis the knowledge base ID recorded in Step 1, andKB_REGIONis the region of the knowledge base.
export ALIBABA_CLOUD_ACCESS_KEY_ID="YOUR_ACCESS_KEY_ID"
export ALIBABA_CLOUD_ACCESS_KEY_SECRET="YOUR_ACCESS_KEY_SECRET"
export KB_REGION="cn-hangzhou"
export KB_ID="kd-xxxxxxxxxxxxx"
export KB_VERSION="LATEST_PUBLISHED"
export LLM_BASE_URL="https://your-openai-compatible-api.example.com/v1"
export LLM_API_KEY="YOUR_LLM_TOKEN"
export LLM_MODEL="YOUR_MODEL_NAME"Step 4: Write the sample code
Save the following content as kb_demo.py. This single module is the complete sample code for this tutorial. It contains four parts: client initialization, document upload, search, and LLM Q&A. The following steps call the functions defined here.
from __future__ import annotations
import os
from pathlib import Path
import requests
from alibabacloud_milvusknowledgebase20260604.client import Client
from alibabacloud_milvusknowledgebase20260604 import models as milvus_kb_models
from alibabacloud_tea_openapi import models as open_api_models
REGION = os.getenv("KB_REGION", "cn-hangzhou")
KB_ID = os.environ["KB_ID"]
KB_VERSION = os.getenv("KB_VERSION", "LATEST_PUBLISHED")
ENDPOINT = f"milvusknowledgebase.{REGION}.aliyuncs.com"
client = Client(open_api_models.Config(
access_key_id=os.environ["ALIBABA_CLOUD_ACCESS_KEY_ID"],
access_key_secret=os.environ["ALIBABA_CLOUD_ACCESS_KEY_SECRET"],
endpoint=ENDPOINT,
region_id=REGION,
connect_timeout=10_000,
read_timeout=60_000,
))
def upload_document(file_path: str) -> dict:
path = Path(file_path).resolve()
size = path.stat().st_size
presigned = client.get_knowledge_base_pre_signed_url(
KB_ID,
milvus_kb_models.GetKnowledgeBasePreSignedUrlRequest(
knowledge_base_id=KB_ID,
documents=[
milvus_kb_models.GetKnowledgeBasePreSignedUrlRequestDocuments(
path=path.name, name=path.name, size=size,
)
],
expires_in=3600,
),
)
upload_url = presigned.body.data.pre_signed_urls[0]
# The presigned URL is signed for an empty Content-Type. Do not include a Content-Type header in the PUT request.
with path.open("rb") as source:
response = requests.put(upload_url, data=source, timeout=120)
response.raise_for_status()
added = client.add_documents(
KB_ID,
milvus_kb_models.AddDocumentsRequest(
knowledge_base_id=KB_ID,
import_type="LOCAL_UPLOAD",
documents=[
milvus_kb_models.AddDocumentsRequestDocuments(
path=path.name, name=path.name, size=size,
)
],
dedup=milvus_kb_models.AddDocumentsRequestDedup(
doc_name_dedup=True, content_dedup=False,
),
),
)
return added.body
def search_knowledge_base(query: str, page_size: int = 6):
resp = client.search_knowledge_base(
KB_ID,
milvus_kb_models.SearchKnowledgeBaseRequest(
query=query,
version=KB_VERSION,
page_number=1,
page_size=page_size,
retrieval_config=milvus_kb_models.SearchKnowledgeBaseRequestRetrievalConfig(
candidate_count=48,
min_score=0,
semantic_weight=0.5,
enable_query_expansion=False,
),
),
)
return resp.body
def chat(messages: list[dict[str, str]]) -> str:
base_url = os.environ["LLM_BASE_URL"].rstrip("/")
response = requests.post(
f"{base_url}/chat/completions",
headers={"Authorization": f"Bearer {os.environ['LLM_API_KEY']}"},
json={
"model": os.environ["LLM_MODEL"],
"messages": messages,
"temperature": 0.1,
},
timeout=60,
)
response.raise_for_status()
return response.json()["choices"][0]["message"]["content"].strip()
def ask(question: str) -> tuple[str, object]:
search_query = chat([
{
"role": "system",
"content": "Rewrite the user's question into a self-contained query suitable for knowledge base search. Output only the query.",
},
{"role": "user", "content": question},
])
search_result = search_knowledge_base(search_query)
results = search_result.results or []
context = "\n\n".join(
f"[Source {i}] {item.document_name or ''}\n{item.content}"
for i, item in enumerate(results, start=1)
)
answer = chat([
{
"role": "system",
"content": "Answer only based on the given materials. Cite materials as [Source N]. Clearly state so when the materials are insufficient.",
},
{"role": "user", "content": f"Question: {question}\n\nMaterials:\n{context}"},
])
return answer, search_resultStep 5: Upload data
An upload has three steps: get a presigned URL, write the file to Object Storage Service (OSS), and register the document. upload_document() wraps the whole process.
Upload a single document:
from kb_demo import upload_document
result = upload_document("./documents/example.md")
print(result.request_id)For batch uploads, start with groups of 10 documents:
from pathlib import Path
from kb_demo import upload_document
for path in sorted(Path("./documents").glob("*.md")):
upload_document(str(path))
print("Submitted:", path.name)The presigned URL is signed for an empty Content-Type. The PUT request must not include a Content-Type request header. Otherwise, OSS returns 403 SignatureDoesNotMatch. If a PUT fails for network reasons, retry it directly.
Putting a file to OSS alone does not add it to the knowledge base. You must call AddDocuments to complete the registration. After a successful registration, run in the response is RUNNING, which means parsing and chunking are in progress. For 10 to 100 typical documents, this processing usually takes 10 to 20 minutes in total. Wait until the status becomes Processed before you publish a version.
When you register documents, both path and name take only the file name. Do not use the object path in the presigned URL. An internal path such as direct_upload/… returns 400 Path must not use an internal storage prefix. Leaving path empty returns 400 Path can't be empty. In a successful registration response, data.documentsmay be an empty array. This does not mean the registration failed. Verify the result with Documents or Chunks on the knowledge base details page, or with subsequent search results.
After the upload completes, check the data status and chunk count on the Data management page in the console. Publish a version only after the status becomes Processed. Click View chunks to preview the chunking results.
Step 6: Publish a version
Imported data stays unpublished and is not searchable until a version is published.
On the knowledge base details page, click the Version management tab. Confirm that the page shows pending changes, and then click Publish version.
In the Confirm changes step, review the change log and click Next.
In the Enter description step, enter a publish description (up to 200 characters) and click Publish. Wait until the message Version published appears.
In Version history, check the version number. The first published version is
After publishing, update the local environment variable:v1, and the status is Published.
export KB_VERSION="v1"Step 7: Search the data
Call search_knowledge_base() to search the published version:
from kb_demo import search_knowledge_base
result = search_knowledge_base("What conditions must be met to terminate a contract?")
for item in result.results or []:
print(item.document_name, item.score)
print(item.content[:300])For version, use an explicit version number such as v1 or v2. You can also use LATEST_PUBLISHED to search the latest published version. Do not use DRAFT to serve stable Q&A externally.
In this example, min_score=0 also returns chunks with relatively low relevance. In a production environment, raise min_score or reduce page_size based on actual results to keep irrelevant content out of the LLM context.
You can also tune TopK, the semantic search weight, the reranking model, and the filter scope on the Search validation page in the console to quickly compare retrieval results.
Step 8: Connect the LLM for Q&A
ask() first has the LLM rewrite the question into a query suitable for search, then calls SearchKnowledgeBase, and finally asks the LLM to summarize an answer based only on the returned chunks.
from kb_demo import ask
answer, search_result = ask("What do these documents say about contract termination?")
print(answer)
print("Request ID:", search_result.request_id)Step 9: Build the web application
Save the following content as
app.py.import streamlit as st from kb_demo import ask st.set_page_config(page_title="Knowledge Base Q&A") st.title("Knowledge Base Q&A") question = st.chat_input("Enter your question") if question: with st.chat_message("user"): st.write(question) with st.chat_message("assistant"): with st.spinner("Searching the knowledge base..."): answer, result = ask(question) st.write(answer) with st.expander("View sources"): for item in result.results or []: st.markdown(f"**{item.document_name or 'Untitled document'}**") st.caption(f"score: {item.score} · Request ID: {result.request_id}") st.write(item.content)Start the application.
streamlit run app.pyThe browser automatically opens the local knowledge base Q&A page.
FAQ
| Symptom | Cause and solution |
An API call returns 403 Permission denied by RAM authentication | The calling account does not have knowledge base permissions. Attach AliyunMilvusFullAccess to the RAM user and retry after the authorization takes effect. |
A PUT upload returns 403 SignatureDoesNotMatch | The request includes a Content-Type request header, but the presigned URL is signed for an empty Content-Type. Remove the header and retry. For details, see Step 5: Upload data. |
A search returns 404 Knowledge base version ... does not exist | The knowledge base has no published version, or KB_VERSION does not match the actual version number. Publish a version in the console first, and then set KB_VERSION to the version number or LATEST_PUBLISHED. |
| A PUT upload occasionally fails to connect | The OSS connection fluctuates. Retry the file directly. |
| Knowledge Base Service is not shown in the left-side navigation pane of the console | The current region does not support the feature, or the menu has not finished loading. Switch to a supported region and check again. |
Pre-launch checklist
Use a dedicated RAM user with least privilege that supports key rotation. Do not use the AccessKey pair of your Alibaba Cloud account for the long term.
Upload only data that you are authorized to process.
Every answer must be traceable to the expanded source chunks.
When you deploy the page publicly, add authentication and HTTPS, and mask sensitive information in access logs.
Treat the product documentation and the actual console behavior as the source of truth for API availability and permission requirements.