An Alibaba Cloud Milvus knowledge base turns the policies, processes, and SOPs scattered across departments into a unified internal Q&A entry point. Employees ask questions in natural language and get process steps together with the original citations, while department tags restrict retrieval to a specified scope.
Solution overview
The end-to-end pipeline is identical to Build an intelligent customer service Q&A application with Alibaba Cloud Milvus knowledge base: define tags in the console → batch import documents by tag → publish a version → retrieve through the SDK with optional tag filtering → an LLM generates answers with source citations → Flask serves the Q&A page. Reuse the engineering code from that document directly (kb_client.py, upload.py, app.py, templates/index.html, and start.sh). This topic covers only the parts that must change for a policy scenario: the department tag scheme, retrieval parameters, the policy assistant prompt, and policy version management.
For the first version, choose only 5 to 20 documents from a single department, and expand after you verify retrieval quality and the prompt. The whole process takes about 20 to 30 minutes.
Prerequisites
Create a knowledge base and record the Knowledge Base ID (for example,
kd-803ae9b10cc31).Create a RAM user with your Alibaba Cloud account, select Use a permanent AccessKey pair, and attach the system policy
AliyunMilvusFullAccess.An LLM endpoint that supports the OpenAI
chat/completionsprotocol, and its API key.Python 3.9 or later installed locally, and the project directory set up by following the tutorial referenced in Solution overview.
Step 1: Define department and policy tags
On the knowledge base details page, under Basic Information, choose Tags > Manage and add the following three tags, all of type string:
| Tag name | Description | Example value |
| department | The department that owns the policy | Finance, Administration, IT |
| docType | The document type | Policy, Process |
| effectiveDate | The effective date | 2026-01-01 |
Tag values accept arbitrary text, so you can use the department and document type names directly.
The tag management dialog box is saved as a whole. After you open the dialog box, the tag list loads asynchronously. Wait until all existing tags are displayed before you add new tags and click Done. Otherwise, the dialog box may overwrite everything with an empty list plus the new tags, which deletes the existing tag definitions.
Tag values that are already written to data are not lost. After you re-add tags, verify the tag count on the details page.
Tag definitions mainly standardize the allowed values. Tag names that are not defined can still be written with data and used for filtering, but define them before import to make team collaboration and maintenance easier.
Step 2: Organize policy documents and the import manifest
Organize documents by department and place them in the
documents/directory. Include the policy name and the version or effective date in each file name, so each entry is easy to identify in the sources section of the answers.Create
documents.jsonland annotate each document with the department, type, and effective date:{"path": "documents/finance-travel.md", "metadata": {"department": "Finance", "docType": "Policy", "effectiveDate": "2026-01-01"}} {"path": "documents/finance-reimbursement.md", "metadata": {"department": "Finance", "docType": "Process", "effectiveDate": "2026-02-01"}} {"path": "documents/hr-leave.md", "metadata": {"department": "Administration", "docType": "Policy", "effectiveDate": "2026-01-01"}} {"path": "documents/it-troubleshoot.md", "metadata": {"department": "IT", "docType": "Process", "effectiveDate": "2026-03-01"}}Write each policy document with the effective date and the policy owner at the beginning of the body, and mark old versions with "Superseded by version X". By default, the LLM sees only the body of retrieved chunks. The
effectiveDatetag is not added to the prompt automatically and is included only after you apply the changes in Step 4.Adjust the chunking granularity. Policy documents are usually short entries. On the knowledge base details page, in Processing Policy (the document processing and chunking settings of the knowledge base), click Create Policy. The unit of maximum chunk length is characters (default 512). Set it to 580 to 770 characters (about 384 to 512 tokens) so that each process step stays intact.
Step 3: Configure retrieval parameters and the prompt
In config.json, adjust retrieval and scenario for the policy scenario. Modify only these two sections and keep the rest of the file unchanged.
{
"retrieval": {
"page_size": 6,
"candidate_count": 48,
"min_score": 0.35,
"semantic_weight": 0.6,
"enable_query_expansion": true,
"rerank_model_name": "",
"tag_filter": {
"relation": "and",
"conditions": []
}
},
"scenario": {
"title": "Enterprise policy and process Q&A",
"system_prompt": "You are an internal policy assistant. Answer only based on the retrieved published policies. Organize processes into steps, state the applicable conditions and required materials, and cite sources as [Source N]. When the materials conflict or are insufficient, clearly tell the user to contact the policy owner.",
"image_enabled": false,
"sample_questions": [
"What materials are required for travel expense reimbursement?",
"Who approves leave requests longer than three days?",
"What is the process for reporting an issue when my computer cannot connect to the network?"
]
}
}The other parameters in the example (page_size, candidate_count, enable_query_expansion, and rerank_model_name) keep the values from the referenced tutorial. The min_score, semantic_weight, and tag_filter parameters need scenario-specific choices:
min_score— There is no universal recommended value across corpora. Calibrate the threshold against your own corpus with real questions:Set
min_scoreto 0 and retrieve a batch of results.Manually label the relevance of the results.
Choose a threshold based on the recall and false-positive distribution.
Until you have such calibration data, start from 0.35, the value measured on the corpus in this topic. On that corpus, at 0.2, policies completely unrelated to the question (scores of 0.36 to 0.43) still entered the LLM context, which increased cost and the risk of wrong answers. At 0.35, these results were filtered out. Recalibrate whenever you change the corpus, switch the rerank model, or adjustsemantic_weight.
semantic_weight=0.6— Policy questions often mix proper nouns with colloquial expressions, so balance semantic and keyword matching.This parameter is the weighting factor of the final score, calculated as
score ≈ (1-semantic_weight) × keywordScore + semantic_weight × semanticScore(a rank feature may also be added).min_scorefilters on this final score after it is computed.Without reranking,
semanticScoreis the vector similarity; with reranking enabled, it is the rerank model score. The two have different scales, so after you change the weight or toggle reranking, recalibrate the threshold together withscoreDetails. Lowering the weight does not necessarily lower the total score; the score drops only when semanticScore is higher than keywordScore for the batch.tag_filter— Leaveconditionsblank by default to retrieve from all departments. Configure filter conditions only when you want to limit an entry point to a single department. Step 5: Scope retrieval to a single department explains the configuration and lists the operator behaviors you must know.
Step 4: Include department and effective date in answers
Policy Q&A must determine which version is effective and which department it applies to, so include the tag values in the LLM context. The tags field in retrieval results does not return the tag values you wrote. It comes from an internal chunk tag enhancement feature of the service and is a different field from the MetaFields written by AddDocuments. No request parameter can enable its return, so an empty value does not mean that the upload failed or that tags were lost.
Maintain the mapping on the client side: index the local import manifest by file name, look up each retrieval result by its documentName, and append the tag values to the LLM context.
The lookup relies on exact name matching. Verify that the documentName returned by retrieval matches the file name in documents.jsonl. Otherwise, the source entry silently carries no tag label.
In app.py, after app = Flask(__name__), add the mapping:
META_BY_NAME: dict[str, dict[str, Any]] = {}
_manifest = Path("documents.jsonl")
if _manifest.is_file():
for _line in _manifest.read_text(encoding="utf-8").splitlines():
if _line.strip():
_entry = json.loads(_line)
META_BY_NAME[Path(str(_entry["path"])).name] = _entry.get("metadata") or {}Then modify llm_answer() to include the tag values when it builds the context and to provide the current date in the question:
def llm_answer(question: str, results: list[dict[str, Any]]) -> str:
if not LLM.get("enabled", True):
return "The LLM is disabled. Review the retrieval results below."
if not results:
return "The current policy materials contain no relevant provisions. Contact the corresponding policy owner for confirmation."
blocks = []
for index, item in enumerate(results, 1):
title = field(item, "documentName", "DocumentName")
meta = META_BY_NAME.get(title) or {}
label = ", ".join(f"{key}={value}" for key, value in meta.items())
header = f"[Source {index}] {title}" + (f" ({label})" if label else "")
blocks.append(f"{header}\n{field(item, 'content', 'Content')}")
context = "\n\n".join(blocks)
today = datetime.date.today().isoformat()
url = str(LLM["base_url"]).rstrip("/") + "/chat/completions"
response = requests.post(
url,
headers={"Authorization": f"Bearer {LLM['api_key']}"},
json={
"model": LLM["model"],
"temperature": 0.1,
"messages": [
{"role": "system", "content": SCENARIO.get("system_prompt", "Answer only based on the provided materials.")},
{"role": "user", "content": f"Current date: {today}\nQuestion: {question}\n\nRetrieved materials:\n{context}"},
],
},
timeout=90,
)
response.raise_for_status()
return response.json()["choices"][0]["message"]["content"].strip()Add import datetime to the file header.
Never skip the step of providing the current date. The LLM does not know today's date. If you provide only effectiveDate, the model assumes a current date on its own and may reach the opposite conclusion. For example, it can judge an already effective new version as "not yet effective" and cite a superseded old standard. Only when both the tag values and the current date are provided can the model select the current version correctly and explain that old versions are superseded.
Step 5: Scope retrieval to a single department
To provide a dedicated entry point for one department, configure the corresponding filter condition:
"tag_filter": {
"relation": "and",
"conditions": [
{"field": "department", "op": "=", "value": "Finance"}
]
}You can also combine multiple conditions. For example, to search only the process documents of the Finance department:
"conditions": [
{"field": "department", "op": "=", "value": "Finance"},
{"field": "docType", "op": "=", "value": "Process"}
]Tag filter operator pitfalls
Only =, in, and not in actually take effect in a tag filter. Review the following behaviors before you configure a filter:
| Configuration | Actual behavior |
op is any other operator, including operators in the "Supported operators" list of the 400 Unsupported tag filter operator error | Tests show that the condition is silently ignored and the unfiltered full data is returned. |
op is contains or not contains | These are only aliases of the set operations in and not in, not string containment. Passing a string fragment returns 0 results. |
field is a tag name that is not defined | The API returns no error and only 0 results. |
op is an alias such as eq, ==, equal, or like | The API returns 400 Unsupported tag filter operator. |
Correct practices:
To filter by a numeric or date range, enumerate the values with
in.To check whether a tag is empty, use
= "".After you configure filter conditions, compare the total number of results with and without the filter to confirm that the filter takes effect.
Once a department filter is configured, this entry point can answer only questions of that department. Adjustsample_questionsaccordingly; otherwise, questions about other departments return only "no relevant provisions".
To switch an entry point to another department, change three places at the same time:
The metadata in
documents.jsonl, which is the source of the tags written at upload and of the answer labels added in Step 4.tag_filterinconfig.json, which defines the retrieval scope of the entry point.The sample questions, so that they match the new scope.
Step 6: Upload, publish, and verify
Upload the documents
MetaFields applies to the whole batch. The script groups documents by tag first and submits them in batches.
python upload.py --manifest documents.jsonlDuplicate uploads of files with the same name fail. When doc_name_dedup=True, if every file in the batch is deduplicated by name, the API returns 400 No OSS document can be registered. When you update policies, include the version number in the file name, or delete the old data in the console first.
Check the processing result
On the Data Management page in the console, confirm that the document status is Processing Completed and that the tag column shows the written tags (for example, docType=Process, department=IT +1).
Publish a version
On the Version Management page, click Publish Version. After you complete the three-step wizard, confirm that the new version status is Published.
After you update policies, you must republish the version. Applications retrieve new content only when they use LATEST_PUBLISHED.
Version quota. At most 3 published versions can exist at the same time by default. This limit is counted per tenant and does not increase with the instance CU specification. When the limit is reached, the Publish Version button is grayed out, while the page still shows that there are N pending changes to publish. Policy knowledge bases are updated frequently. Keep only the current effective version plus the most recent historical version, and clean up older versions before you publish a new one. The limit can be adjusted on the server side, but currently there is no self-service quota request entry for users. Submit a ticket for evaluation.
Version deletion is irreversible, and a deleted version becomes immediately unretrievable. The backend cleans up the data asynchronously. Before deletion, confirm that no application is locked to that version number, and switch applications that still use it to a new version.
Start the service and verify
python app.pySubmit a test question:
curl -sS http://127.0.0.1:7860/api/ask \
-H 'Content-Type: application/json' \
-d '{"question":"What materials are required for travel expense reimbursement?"}'Accept against the following criteria:
The page opens normally, and a submitted question returns both the answer and the retrieval sources.
Each
[Source N]in the answer has a matching policy chunk in the sources section below.The
[Source N]headers carry the tag labels, such asdepartment=,docType=, andeffectiveDate=. If a label is missing, verify that thedocumentNamereturned by retrieval matches the file name indocuments.jsonl.When you ask about something the policies do not cover, the answer clearly states that there is no relevant provision and tells you to contact the policy owner.
After the policies are updated and the version is republished, the page retrieves content from the new version. This requires the application to use
LATEST_PUBLISHED.
Considerations for policy scenarios
Old policy versions — If you keep historical versions for traceability, mark the body with "Superseded by version X" and distinguish versions with
effectiveDate. If traceability is not needed, delete old versions on the Data Management page and republish, so the model does not waver between versions.Manual fallback — For Q&A on critical policies, keep a manual fallback entry point and tell employees to follow the officially published policy documents.
FAQ
The following table lists common symptoms with their causes and resolutions.
| Symptom | Cause and resolution |
A question returns 500, and the log shows 400 Unsupported tag filter operator | op uses an alias such as eq, ==, or like. Use =, in, or not in instead. |
| After you add a tag filter, the number of results is exactly the same as without the filter | You used an operator that does not take effect, such as >, ≥, ≠, or empty. Only =, in, and not in take effect. |
| After you configure a department filter, most questions are answered with "no relevant provisions" | The entry point is limited to a single department. Confirm that the question scope matches tag_filter, or leave conditions blank. |
| Filtering by tag always returns 0 results | The tag name spelling does not match what was written. The API returns no error, only 0 results. Verify the tag name under Tags > Manage on the knowledge base details page. |
| The answer mixes old and new policy versions | The retrieved chunks lack version information. Follow Step 4 to include the tag values in the context, and mark old versions as superseded in the body. |
Upload returns 400 No OSS document can be registered. | Every file in the batch was deduplicated by name. Change the file names or delete the old data first. |
Retrieval returns 404 Knowledge base version ... does not exist | No version is published yet, or knowledge_base_version does not match the actual version number. |
| The Publish Version button is grayed out, but the page shows pending changes | The number of published versions has reached the limit of 3. Hover over the button to see the hint. Delete old versions in the version records, and then publish. |
You want to show tags in answers, but the tags field of retrieval results is empty | Retrieval results do not backfill tag values. Follow Step 4 to look them up from documents.jsonl. |
| After you add tags, the original tag definitions disappear | The tag management dialog box is saved as a whole. Reopen the dialog box, wait until the list finishes loading, and re-add the missing tag definitions. Tag values already written to data are not affected. |