All Products
Search
Document Center

Vector Retrieval Service for Milvus:Build a legal case retrieval application with the Alibaba Cloud Milvus knowledge base

Last Updated:Aug 27, 2026

The Alibaba Cloud Milvus knowledge base can turn public laws and regulations, court judgments, and corporate compliance policies into a searchable legal materials assistant: retrieve similar cases by case details, summarize judicial opinions, and display the original sources.

Solution

The end-to-end pipeline is identical to the one in Build an intelligent customer service Q&A application by using the Alibaba Cloud Milvus knowledge base : define tags in the console → import source documents in batches by tag → publish a version → SDK retrieval with optional tag filtering → an LLM generates answers with source citations → Flask serves the Q&A page. Reuse the engineering code from that tutorial directly: kb_client.py, upload.py, app.py, templates/index.html, and start.sh. This topic covers only the parts that must change for a legal scenario: case tags (including a numeric tag type), retrieval parameters, prompt constraints, and the empty-result handling that you must add.

For the first version, pick a single crime type or a single compliance topic and import only 10–50 source documents. Expand the scope after validation. The whole process takes about 25–40 minutes. Document parsing time depends on the number and size of the source documents.

This solution is a supporting tool for materials retrieval and does not produce legal opinions. Keep a manual review step for both retrieval results and model-generated summaries.

Prerequisites

  • A Milvus knowledge base is created, and you have recorded the knowledge base ID (for example, kd-803ae9b10cc31).

  • You have created a RAM user under your Alibaba Cloud account, selected Use permanent AccessKey, and attached the system policy AliyunMilvusFullAccess.

  • You have prepared the endpoint and API key of a large language model that supports the OpenAI chat/completions protocol.

  • Python 3.8 or later is installed locally, and the project directory has been set up by following.

  • The authorization check and de-identification of your source documents are complete. For details, see Notes for legal scenarios.

Step 1: Define case tags

On the knowledge base details page, below Basic Information, choose Tags > Manage and add the following four tags. Select the field type from the drop-down list to the right of the tag name input box. The available options are string, int64, list, float32, and bool.

Tag name

Field type

Description

Example values

docType

string

Document type

Court judgment, Laws and regulations, Compliance policy

court

string

Trial court

The First People's Court of XX City

caseType

string

Cause of action

Contract fraud, Causing a traffic accident

judgmentYear

int64

Judgment year

2025

Defining the judgment year as int64 does not enable range queries. The retrieval API does not support integer range expressions. The purpose of defining the judgment year as int64 is that the server converts numeric strings into integers, so that =, in, and not in work correctly. Therefore, for a requirement such as "cases from the last three years", the client must first compute the list of years and then pass it with in, for example [2024, 2025, 2026]. Do not use ≥ or ≤ in tag_filter conditions. When the year range is large, there is no equally efficient alternative — split it into multiple queries.

Warning

The tag management dialog box is saved as a whole, and the tag list loads asynchronously after the dialog box opens. Wait until all existing tags are displayed before you add new tags and click Done. Otherwise, the save may overwrite everything with an "empty list + new tags", wiping out the existing tag definitions.

Note the following about tag types and values:

  • For a tag of type int64, you can write a plain JSON number (such as 2025) directly in documents.jsonl. When filtering, value matches whether you pass a number or a string.

  • Tag values can be an empty string. Laws and regulations and compliance policies have no trial court, so you can write "court": "". The console displays this as court=, and you can later use court = "" as a filter condition to match exactly these source documents.

Step 2: Prepare source documents and the import manifest

  • Place the source documents in the documents/ directory. PDF, DOCX, Markdown, TXT, and other formats are supported. Keep the case number, court, and judgment date in the body text or in the file name. In our tests, after writing the case number on the first line of the body, we could retrieve the corresponding judgment directly by case number.

  • Create documents.jsonl and annotate each source document with the document type, court, cause of action, and judgment year:

    {"path": "documents/criminal-case-001.md", "metadata": {"docType": "Court judgment", "court": "The First People's Court of XX City", "caseType": "Contract fraud", "judgmentYear": 2025}}
    {"path": "documents/criminal-case-002.md", "metadata": {"docType": "Court judgment", "court": "The Second People's Court of XX City", "caseType": "Contract fraud", "judgmentYear": 2023}}
    {"path": "documents/law-excerpt.md", "metadata": {"docType": "Laws and regulations", "court": "", "caseType": "Criminal", "judgmentYear": 2024}}
    {"path": "documents/compliance-policy.md", "metadata": {"docType": "Compliance policy", "court": "", "caseType": "Compliance", "judgmentYear": 2024}}
  • Court judgments must preserve the context of the case facts, the reasoning of the judgment, and the conclusion. On the knowledge base details page, click Create Policy under Processing Policy to adjust the chunk granularity. The unit of the maximum chunk length is characters (default 512). Set it to 770–1,150 characters (roughly 512–768 tokens) so that the key issues in dispute and the reasoning of the judgment fall into the same chunk whenever possible.

  • Import cases with different conclusions under the same cause of action. The value of legal retrieval lies precisely in presenting disagreements. In our tests, after importing two contract fraud judgments with opposite conclusions, the model cited both and explicitly pointed out that "one case recognized joint crime, while the other did not due to a lack of evidence of conspiracy". This is more valuable for reference than importing materials with a single conclusion. Require in the prompt that such cases be presented separately, so that the model does not present an individual case's conclusion as a general rule.

Step 3: Configure retrieval parameters and the prompt

In config.json, adjust aliyun.knowledge_base_version, retrieval, and scenario for the legal scenario. The following example shows only these three blocks. Keep the remaining blocks of your config.json from the referenced tutorial, including the knowledge base ID and the LLM endpoint and API key that you prepared in Prerequisites.

{
  "aliyun": {
    "knowledge_base_version": "LATEST_PUBLISHED"
  },
  "retrieval": {
    "page_size": 8,
    "candidate_count": 80,
    "min_score": 0.25,
    "semantic_weight": 0.4,
    "enable_query_expansion": true,
    "rerank_model_name": "qwen3-rerank",
    "tag_filter": {
      "relation": "and",
      "conditions": [
        {"field": "docType", "op": "=", "value": "Court judgment"}
      ]
    }
  },
  "scenario": {
    "title": "Legal and compliance case retrieval",
    "system_prompt": "You are a legal materials retrieval assistant. Do not provide final legal opinions. Summarize facts, key issues in dispute, and judicial opinions and their bases strictly from the retrieved materials, and annotate each item with [Source N]. When different cases reach conflicting conclusions, present them separately. If the retrieved materials are empty or irrelevant to the question, reply only with 'Insufficient materials; manual review is recommended', and never cite any law, regulation, judicial interpretation, or case that does not appear in the retrieved materials.",
    "image_enabled": false,
    "sample_questions": [
      "When someone defrauds payment for goods by fabricating the ability to perform a contract, how is joint crime in contract fraud determined?",
      "How does voluntary surrender after causing a traffic accident affect sentencing?",
      "Which cases discuss the distinction between a principal offender and an accessory?"
    ]
  }
}

Note the following about the key parameters:

  • semantic_weight=0.4: legal retrieval makes heavy use of case numbers, crime types, and legal terminology, so lower the semantic weight and raise the keyword weight accordingly. In our tests, asking with a full case number hits the corresponding judgment precisely. You can tune this value by using scoreDetails in the retrieval results, which contains keywordScore and semanticScore.

  • candidate_count=80 with qwen3-rerank enabled: expand the candidate set and let the reranking model rank the candidates by similarity to similar cases. Note that rerank scores and vector scores are not on the same scale. Combined with min_score, they may filter out some results. When tuning, fix one of the two first.

  • Keep knowledge_base_version set to LATEST_PUBLISHED unless compliance tracing requires a pinned version. For details about version pinning and deletion, see Manage published versions.

Only three operators work for tag filtering

In our tests, in the current version, only =, in, and not in in tag_filter.conditions actually take effect as the op value:

Operator

Behavior

Example

=

Exact match; for an int64 tag, passing a number or a string both work

{"field": "judgmentYear", "op": "=", "value": 2025}

in

Enumerated match

{"field": "judgmentYear", "op": "in", "value": [2024, 2025]}

not in

Exclude an enumeration

{"field": "docType", "op": "not in", "value": ["Compliance policy"]}

>, ≥, <, ≤

The condition is ignored and all data is returned (no error is raised)

—

≠, empty, not empty, start with, end with

The condition is ignored and all data is returned

—

contains, not contains

Aliases of in/not in, not string containment; passing a string fragment returns 0 results

—

If you pass an operator that does not take effect, the API does not raise an error and returns all data without filtering. In a legal scenario, this means cases that should have been excluded still enter the LLM context, and nothing on the page looks wrong. Therefore:

  • Do not use > or ≥ to filter by year range. Use in with an enumeration of years instead, for example {"field": "judgmentYear", "op": "in", "value": [2023, 2024, 2025]}.

  • Do not use empty to check for an empty tag. Use {"field": "court", "op": "=", "value": ""} instead.

  • After you configure a filter condition, always compare the total result count with the unfiltered result to confirm that the filter actually takes effect. If the two are identical, the condition was not applied.

    If op is written as an alias such as eq, ==, or like, the API returns 400 Unsupported tag filter operator, and every question fails. The "Supported operators" listed in that error message include operators from the table above that do not take effect, so they cannot be treated as a usable list.

Step 4: Intercept empty retrieval results in app.py

In the reused app.py, llm_answer() still calls the LLM when the retrieval results are empty. The context is then an empty string, and the model answers entirely from its own knowledge. In our tests, we asked a question the knowledge base does not cover, such as "How are damages calculated for intellectual property infringement?". The model produced lengthy damage-calculation rules and fabricated source annotations like [Source 1: Interpretation of ... Punitive Damages, Article 2], even though sources was empty. In a legal scenario, this kind of output is highly misleading and must be intercepted at the code level:

def llm_answer(question: str, results: list[dict[str, Any]]) -> str:
    if not LLM.get("enabled", True):
        return "The LLM is disabled; please see the retrieval results below."
    if not results:
        return "No materials related to this question were found in the knowledge base, so I cannot answer. Add relevant materials and try again, or escalate to manual review."
    ...

Prompt constraints alone are not reliable. Even when system_prompt explicitly says "state clearly when materials are insufficient", the model still answers. After adding the fallback above, uncovered questions consistently return the prompt message, and normal questions with retrieval results are unaffected.

Step 5: Upload, publish, and verify

  1. Upload the source documents. MetaFields applies to the whole batch; the script groups documents by tag and submits them in batches.

    python upload.py --manifest documents.jsonl
  2. On the Data Management page of the console, confirm that the document status is Processed. A successful response from the upload API only means that the asynchronous parsing has been submitted.

  3. On the Version Management page, click Publish Version, complete the wizard, and confirm that the new version status is Published. If the button is grayed out, delete old versions that are no longer needed first. See Manage published versions.

  4. Start the service.

    python app.py
  5. After the service starts, ask a test question:

    curl -sS http://127.0.0.1:7860/api/ask \
      -H 'Content-Type: application/json' \
      -d '{"question":"When someone defrauds payment for goods by fabricating the ability to perform a contract, how is joint crime in contract fraud determined?"}'
  • The page opens normally, and submitting a question returns both the answer and the retrieval sources.

  • Every [Source N] in the answer corresponds to a source chunk in the sources section below.

  • Asking with a full case number hits the corresponding court judgment precisely.

  • When you ask about content that the knowledge base does not cover, the response returns the "Insufficient materials" message and contains no legal citations whatsoever. Make sure to test this in practice — it is the most critical acceptance criterion for this scenario.

  • After you add more materials and publish a new version, the page can retrieve the new content.

Manage published versions

LATEST_PUBLISHED automatically uses the latest published version. You can also specify an explicit version number (such as v2) to pin a version. Compliance scenarios usually require tracing which version of materials supported which decision, so record the version number actually used in the application's call logs.

Important

After a version is deleted, it immediately becomes unretrievable, and the backend cleans up the data asynchronously afterwards. Before you delete a version, switch any application that is still pinned to it to a new version. If you delete a version that the application is using, retrieval immediately returns 404 Knowledge base version ... does not exist, and every question fails on the page.

At most three published versions can exist at the same time. Once the limit is reached, the Publish Version button is grayed out, while the page still shows "There are N pending changes". Delete old versions that are no longer needed from the version records first. If you pin a version, you must also establish a process of updating the configuration before deleting an old version.

Notes for legal scenarios

Before you share the application with end users, review the following considerations:

  • Authorized materials only — Use only materials that you are authorized to make public and process, and complete the authorization check before uploading. If a court judgment contains personal information, de-identify it first. Document chunks are sent to the LLM as context, which constitutes data egress.

  • Permanent disclaimer — The page should carry a permanent disclaimer. The sample page only says "Answers are generated from published knowledge base content". Change the notice in templates/index.html to state clearly: "These results are a materials retrieval aid only and do not constitute legal opinions. Have them reviewed by a professional."

  • Separate entry points by document type — For example, restrict a public consultation entry point with docType = Laws and regulations so that it retrieves only public laws and regulations, while an internal analysis entry point also allows court judgments.

  • Production deployment — Do not keep using the Flask development server. Switch to a production WSGI server and add secret management, authentication, auditing, and throttling.

FAQ

Symptom

Cause and solution

Every question returns 404 Knowledge base version ... does not exist

The version number hard-coded in the configuration has been deleted or was never published. Use LATEST_PUBLISHED, or specify a version name that actually exists on the Version Management page.

Questions return 500, and the logs show 400 Unsupported tag filter operator

op uses an alias such as eq, ==, or like. Use =, in, or not in instead.

The result count with tag filtering is exactly the same as without it

You used an operator that does not take effect (such as >, ≥, ≠, or empty). Use =, in, or not in instead. See Step 3.

Filtering by year range has no effect

Numeric comparison operators do not take effect. Use in with an enumeration of years.

Tag filtering always returns 0 results

The tag name is spelled differently from what was written. The API does not raise an error and just returns 0 results. Verify it under Tags > Manage on the knowledge base details page. In our tests, the contains operator returns 0 results as well.

Asked a question the knowledge base does not cover, but got plausible-looking legal citations

llm_answer() still called the LLM when the retrieval results were empty. Add the empty-result fallback from Step 4.

The tags field in retrieval results is empty

The current version does not backfill tags. The tags field does not return the court and judgment year that you wrote at upload time. It is a different field from the MetaFields of AddDocuments, and there is no parameter to enable returning it. An empty value does not mean the upload failed. If you need to show this information in answers, join the retrieval results with documents.jsonl by documentId and append it to the context.

Upload returns 400 No OSS document can be registered.

Every file in the batch was deduplicated because of duplicate file names. Rename the files, or delete the old data in the console first.

Retrieval reports a failure with code = None

The _check() logic in the reused code is wrong (getattr(body, "code", 0) != 0): a successful response has code = None. Change it to getattr(body, "code", None) not in (None, 0, "0").

After enabling reranking, there are fewer results

Rerank scores and vector scores are on different scales. Combined with min_score, more results get filtered out. Lower min_score first, or disable reranking to compare.