The Alibaba Cloud Milvus knowledge base can turn public laws and regulations, court judgments, and corporate compliance policies into a searchable legal materials assistant: retrieve similar cases by case details, summarize judicial opinions, and display the original sources.
Solution
The end-to-end pipeline is identical to the one in Build an intelligent customer service Q&A application by using the Alibaba Cloud Milvus knowledge base : define tags in the console → import source documents in batches by tag → publish a version → SDK retrieval with optional tag filtering → an LLM generates answers with source citations → Flask serves the Q&A page. Reuse the engineering code from that tutorial directly: kb_client.py, upload.py, app.py, templates/index.html, and start.sh. This topic covers only the parts that must change for a legal scenario: case tags (including a numeric tag type), retrieval parameters, prompt constraints, and the empty-result handling that you must add.
For the first version, pick a single crime type or a single compliance topic and import only 10–50 source documents. Expand the scope after validation. The whole process takes about 25–40 minutes. Document parsing time depends on the number and size of the source documents.
This solution is a supporting tool for materials retrieval and does not produce legal opinions. Keep a manual review step for both retrieval results and model-generated summaries.
Prerequisites
-
A Milvus knowledge base is created, and you have recorded the knowledge base ID (for example,
kd-803ae9b10cc31). -
You have created a RAM user under your Alibaba Cloud account, selected Use permanent AccessKey, and attached the system policy
AliyunMilvusFullAccess. -
You have prepared the endpoint and API key of a large language model that supports the OpenAI
chat/completionsprotocol. -
Python 3.8 or later is installed locally, and the project directory has been set up by following.
-
The authorization check and de-identification of your source documents are complete. For details, see Notes for legal scenarios.
Step 1: Define case tags
On the knowledge base details page, below Basic Information, choose Tags > Manage and add the following four tags. Select the field type from the drop-down list to the right of the tag name input box. The available options are string, int64, list, float32, and bool.
|
Tag name |
Field type |
Description |
Example values |
|
docType |
string |
Document type |
Court judgment, Laws and regulations, Compliance policy |
|
court |
string |
Trial court |
The First People's Court of XX City |
|
caseType |
string |
Cause of action |
Contract fraud, Causing a traffic accident |
|
judgmentYear |
|
Judgment year |
2025 |
Defining the judgment year as int64 does not enable range queries. The retrieval API does not support integer range expressions. The purpose of defining the judgment year as int64 is that the server converts numeric strings into integers, so that =, in, and not in work correctly. Therefore, for a requirement such as "cases from the last three years", the client must first compute the list of years and then pass it with in, for example [2024, 2025, 2026]. Do not use ≥ or ≤ in tag_filter conditions. When the year range is large, there is no equally efficient alternative — split it into multiple queries.
The tag management dialog box is saved as a whole, and the tag list loads asynchronously after the dialog box opens. Wait until all existing tags are displayed before you add new tags and click Done. Otherwise, the save may overwrite everything with an "empty list + new tags", wiping out the existing tag definitions.
Note the following about tag types and values:
-
For a tag of type
int64, you can write a plain JSON number (such as2025) directly indocuments.jsonl. When filtering,valuematches whether you pass a number or a string. -
Tag values can be an empty string. Laws and regulations and compliance policies have no trial court, so you can write
"court": "". The console displays this ascourt=, and you can later usecourt = ""as a filter condition to match exactly these source documents.
Step 2: Prepare source documents and the import manifest
-
Place the source documents in the
documents/directory. PDF, DOCX, Markdown, TXT, and other formats are supported. Keep the case number, court, and judgment date in the body text or in the file name. In our tests, after writing the case number on the first line of the body, we could retrieve the corresponding judgment directly by case number. -
Create
documents.jsonland annotate each source document with the document type, court, cause of action, and judgment year:{"path": "documents/criminal-case-001.md", "metadata": {"docType": "Court judgment", "court": "The First People's Court of XX City", "caseType": "Contract fraud", "judgmentYear": 2025}} {"path": "documents/criminal-case-002.md", "metadata": {"docType": "Court judgment", "court": "The Second People's Court of XX City", "caseType": "Contract fraud", "judgmentYear": 2023}} {"path": "documents/law-excerpt.md", "metadata": {"docType": "Laws and regulations", "court": "", "caseType": "Criminal", "judgmentYear": 2024}} {"path": "documents/compliance-policy.md", "metadata": {"docType": "Compliance policy", "court": "", "caseType": "Compliance", "judgmentYear": 2024}} -
Court judgments must preserve the context of the case facts, the reasoning of the judgment, and the conclusion. On the knowledge base details page, click Create Policy under Processing Policy to adjust the chunk granularity. The unit of the maximum chunk length is characters (default 512). Set it to 770–1,150 characters (roughly 512–768 tokens) so that the key issues in dispute and the reasoning of the judgment fall into the same chunk whenever possible.
-
Import cases with different conclusions under the same cause of action. The value of legal retrieval lies precisely in presenting disagreements. In our tests, after importing two contract fraud judgments with opposite conclusions, the model cited both and explicitly pointed out that "one case recognized joint crime, while the other did not due to a lack of evidence of conspiracy". This is more valuable for reference than importing materials with a single conclusion. Require in the prompt that such cases be presented separately, so that the model does not present an individual case's conclusion as a general rule.
Step 3: Configure retrieval parameters and the prompt
In config.json, adjust aliyun.knowledge_base_version, retrieval, and scenario for the legal scenario. The following example shows only these three blocks. Keep the remaining blocks of your config.json from the referenced tutorial, including the knowledge base ID and the LLM endpoint and API key that you prepared in Prerequisites.
{
"aliyun": {
"knowledge_base_version": "LATEST_PUBLISHED"
},
"retrieval": {
"page_size": 8,
"candidate_count": 80,
"min_score": 0.25,
"semantic_weight": 0.4,
"enable_query_expansion": true,
"rerank_model_name": "qwen3-rerank",
"tag_filter": {
"relation": "and",
"conditions": [
{"field": "docType", "op": "=", "value": "Court judgment"}
]
}
},
"scenario": {
"title": "Legal and compliance case retrieval",
"system_prompt": "You are a legal materials retrieval assistant. Do not provide final legal opinions. Summarize facts, key issues in dispute, and judicial opinions and their bases strictly from the retrieved materials, and annotate each item with [Source N]. When different cases reach conflicting conclusions, present them separately. If the retrieved materials are empty or irrelevant to the question, reply only with 'Insufficient materials; manual review is recommended', and never cite any law, regulation, judicial interpretation, or case that does not appear in the retrieved materials.",
"image_enabled": false,
"sample_questions": [
"When someone defrauds payment for goods by fabricating the ability to perform a contract, how is joint crime in contract fraud determined?",
"How does voluntary surrender after causing a traffic accident affect sentencing?",
"Which cases discuss the distinction between a principal offender and an accessory?"
]
}
}
Note the following about the key parameters:
-
semantic_weight=0.4: legal retrieval makes heavy use of case numbers, crime types, and legal terminology, so lower the semantic weight and raise the keyword weight accordingly. In our tests, asking with a full case number hits the corresponding judgment precisely. You can tune this value by usingscoreDetailsin the retrieval results, which containskeywordScoreandsemanticScore. -
candidate_count=80withqwen3-rerankenabled: expand the candidate set and let the reranking model rank the candidates by similarity to similar cases. Note that rerank scores and vector scores are not on the same scale. Combined withmin_score, they may filter out some results. When tuning, fix one of the two first. -
Keep
knowledge_base_versionset toLATEST_PUBLISHEDunless compliance tracing requires a pinned version. For details about version pinning and deletion, see Manage published versions.
Only three operators work for tag filtering
In our tests, in the current version, only =, in, and not in in tag_filter.conditions actually take effect as the op value:
|
Operator |
Behavior |
Example |
|
|
Exact match; for an |
|
|
|
Enumerated match |
|
|
|
Exclude an enumeration |
|
|
|
The condition is ignored and all data is returned (no error is raised) |
— |
|
|
The condition is ignored and all data is returned |
— |
|
|
Aliases of |
— |
If you pass an operator that does not take effect, the API does not raise an error and returns all data without filtering. In a legal scenario, this means cases that should have been excluded still enter the LLM context, and nothing on the page looks wrong. Therefore:
-
Do not use
>or≥to filter by year range. Useinwith an enumeration of years instead, for example{"field": "judgmentYear", "op": "in", "value": [2023, 2024, 2025]}. -
Do not use
emptyto check for an empty tag. Use{"field": "court", "op": "=", "value": ""}instead. -
After you configure a filter condition, always compare the total result count with the unfiltered result to confirm that the filter actually takes effect. If the two are identical, the condition was not applied.
If
opis written as an alias such aseq,==, orlike, the API returns400 Unsupported tag filter operator, and every question fails. The "Supported operators" listed in that error message include operators from the table above that do not take effect, so they cannot be treated as a usable list.
Step 4: Intercept empty retrieval results in app.py
In the reused app.py, llm_answer() still calls the LLM when the retrieval results are empty. The context is then an empty string, and the model answers entirely from its own knowledge. In our tests, we asked a question the knowledge base does not cover, such as "How are damages calculated for intellectual property infringement?". The model produced lengthy damage-calculation rules and fabricated source annotations like [Source 1: Interpretation of ... Punitive Damages, Article 2], even though sources was empty. In a legal scenario, this kind of output is highly misleading and must be intercepted at the code level:
def llm_answer(question: str, results: list[dict[str, Any]]) -> str:
if not LLM.get("enabled", True):
return "The LLM is disabled; please see the retrieval results below."
if not results:
return "No materials related to this question were found in the knowledge base, so I cannot answer. Add relevant materials and try again, or escalate to manual review."
...
Prompt constraints alone are not reliable. Even when system_prompt explicitly says "state clearly when materials are insufficient", the model still answers. After adding the fallback above, uncovered questions consistently return the prompt message, and normal questions with retrieval results are unaffected.
Step 5: Upload, publish, and verify
-
Upload the source documents.
MetaFieldsapplies to the whole batch; the script groups documents by tag and submits them in batches.python upload.py --manifest documents.jsonl -
On the Data Management page of the console, confirm that the document status is Processed. A successful response from the upload API only means that the asynchronous parsing has been submitted.
-
On the Version Management page, click Publish Version, complete the wizard, and confirm that the new version status is Published. If the button is grayed out, delete old versions that are no longer needed first. See Manage published versions.
-
Start the service.
python app.py -
After the service starts, ask a test question:
curl -sS http://127.0.0.1:7860/api/ask \ -H 'Content-Type: application/json' \ -d '{"question":"When someone defrauds payment for goods by fabricating the ability to perform a contract, how is joint crime in contract fraud determined?"}'
-
The page opens normally, and submitting a question returns both the answer and the retrieval sources.
-
Every
[Source N]in the answer corresponds to a source chunk in the sources section below. -
Asking with a full case number hits the corresponding court judgment precisely.
-
When you ask about content that the knowledge base does not cover, the response returns the "Insufficient materials" message and contains no legal citations whatsoever. Make sure to test this in practice — it is the most critical acceptance criterion for this scenario.
-
After you add more materials and publish a new version, the page can retrieve the new content.
Manage published versions
LATEST_PUBLISHED automatically uses the latest published version. You can also specify an explicit version number (such as v2) to pin a version. Compliance scenarios usually require tracing which version of materials supported which decision, so record the version number actually used in the application's call logs.
After a version is deleted, it immediately becomes unretrievable, and the backend cleans up the data asynchronously afterwards. Before you delete a version, switch any application that is still pinned to it to a new version. If you delete a version that the application is using, retrieval immediately returns 404 Knowledge base version ... does not exist, and every question fails on the page.
At most three published versions can exist at the same time. Once the limit is reached, the Publish Version button is grayed out, while the page still shows "There are N pending changes". Delete old versions that are no longer needed from the version records first. If you pin a version, you must also establish a process of updating the configuration before deleting an old version.
Notes for legal scenarios
Before you share the application with end users, review the following considerations:
-
Authorized materials only — Use only materials that you are authorized to make public and process, and complete the authorization check before uploading. If a court judgment contains personal information, de-identify it first. Document chunks are sent to the LLM as context, which constitutes data egress.
-
Permanent disclaimer — The page should carry a permanent disclaimer. The sample page only says "Answers are generated from published knowledge base content". Change the notice in
templates/index.htmlto state clearly: "These results are a materials retrieval aid only and do not constitute legal opinions. Have them reviewed by a professional." -
Separate entry points by document type — For example, restrict a public consultation entry point with
docType = Laws and regulationsso that it retrieves only public laws and regulations, while an internal analysis entry point also allows court judgments. -
Production deployment — Do not keep using the Flask development server. Switch to a production WSGI server and add secret management, authentication, auditing, and throttling.
FAQ
|
Symptom |
Cause and solution |
|
Every question returns |
The version number hard-coded in the configuration has been deleted or was never published. Use |
|
Questions return 500, and the logs show |
|
|
The result count with tag filtering is exactly the same as without it |
You used an operator that does not take effect (such as |
|
Filtering by year range has no effect |
Numeric comparison operators do not take effect. Use |
|
Tag filtering always returns 0 results |
The tag name is spelled differently from what was written. The API does not raise an error and just returns 0 results. Verify it under Tags > Manage on the knowledge base details page. In our tests, the |
|
Asked a question the knowledge base does not cover, but got plausible-looking legal citations |
|
|
The |
The current version does not backfill tags. The |
|
Upload returns |
Every file in the batch was deduplicated because of duplicate file names. Rename the files, or delete the old data in the console first. |
|
Retrieval reports a failure with |
The |
|
After enabling reranking, there are fewer results |
Rerank scores and vector scores are on different scales. Combined with |