All Products
Search
Document Center

Vector Retrieval Service for Milvus:Build an intelligent customer service Q&A application by using the Alibaba Cloud Milvus knowledge base

Last Updated:Sep 01, 2026

In this tutorial, you use the Alibaba Cloud Milvus knowledge base to build a tagged knowledge index from customer service documents. The corpus covers product manuals, FAQs, and policies for returns and exchanges, warranties, and invoices. By the end, you will have a customer service Q&A page that retrieves answers through tag filtering and a rerank model, generates responses through an LLM, and displays answer sources.

Solution overview

The overall pipeline is: define tags and create a knowledge base in the console → batch import customer service documents by tag → publish a version → search through the SDK, optionally filtering by tag → pass the results to an LLM to generate answers with source citations → serve the Q&A page with Flask.

This topic focuses on three things specific to customer service scenarios: tag (metadata) filtering, the rerank model, and answer traceability. For the basic knowledge base pipeline (presigned upload, register data, publish a version) and the simplest Q&A implementation, see Build a personal knowledge Q&A application.

Before you start, prepare 5 to 20 documents in PDF, DOCX, Markdown, or TXT format. The whole process takes about 20 to 30 minutes. Document parsing time depends on the number and size of the documents.

Prerequisites

  • A knowledge base created in a region that supports knowledge bases: China (Hangzhou), China (Beijing), China (Zhangjiakou), or China (Shenzhen). Record the knowledge base ID, which is in the format kd-803ae9b10cc31.

  • A RAM user created from your Alibaba Cloud account, with Use Permanent AccessKey selected and the system policy AliyunMilvusFullAccess attached to the RAM user.

  • An LLM endpoint that supports the OpenAI chat/completions protocol, together with an API key. For example, use the OpenAI-compatible endpoint of Alibaba Cloud Model Studio.

  • Python 3.8 or later installed on your local machine.

    For credential security and production deployment guidance, see Pre-launch checklist.

Step 1: Define tags

Tags attach business dimensions (such as document type, product line, and effective date) to documents. They are written to the knowledge base during import and used for filtering during search. They are key to distinguishing policies, manuals, and FAQs in customer service scenarios.

  1. Log on to the Alibaba Cloud Milvus console and open the details page of the target knowledge base.

  2. Under Basic Information, find Tags and click Manage.

  3. In the Tag Management dialog box, enter a tag name, select a field type, and click Add. This topic uses the following three tags, all of the string type:

    Tag name Description Example value
    docType Document type policy, manual, faq
    productLine Applicable product line all, phone
    effectiveDate Effective date 2026-01-01
  4. After you add the three tags, click Done. Tags on the details page shows "3 tags".

    In this tutorial, tag values are written to the documents in bulk during import in Step 5, so you do not need to set values in the console now. To tag documents that already exist in the knowledge base, use Set Tags on the Data Management page, as described in FAQ.

Usage notes

  • Field types — The field type can be string, int64, list, float32, or bool.

  • Existing knowledge bases — Tags can be added to an existing knowledge base after the fact, without rebuilding the knowledge base.

  • Definitions are not a prerequisite for writing values — Defining tags in the console is not a prerequisite for writing tag values. Undefined fields can still be written with documents and used for filtering, and deleting a tag definition does not clear historical values that have been written. A definition is not an independent index field declaration. The practical effects of definitions are console display, option management for list-type tags, and value type conversion during filtering by string, int64, float32, bool, or list. The recommended practice is to define tags first to gain type validation and manageability, rather than treating it as a mandatory prerequisite.

  • Definition changes do not rebuild data — The dialog box warns that tag changes affect index building, but adding or removing tag definitions does not rebuild or migrate already imported data, and historical tag values are not cleared. Still, finalize tag definitions before batch import, so that console display, option management, and value type conversion stay consistent from the very beginning.

Important

The Tag Options column presets allowed values for a tag and only affects how values are entered in the console: when it is left blank, filtering by this tag in the console requires manually typing tag values; after you enter comma-separated values (for example after-sales,logistics,billing), the console switches to a drop-down list and avoids typos. It does not validate values written through the API — using AddDocuments to pass values outside the options still succeeds, and the value range must still be guaranteed by the import manifest itself.

Step 2: Prepare documents and the import manifest

  • Place the customer service documents in the local documents/ directory. For example:

kb-demo/
├── documents/
│   ├── shipping-policy.md
│   ├── return-policy.md
│   ├── warranty-policy.md
│   ├── phone-manual.md
│   └── invoice-faq.md
  • Create documents.jsonl. Each line describes one document and its tags. Tag names must match those defined in Step 1.

{"path": "documents/shipping-policy.md", "metadata": {"docType": "policy", "productLine": "all", "effectiveDate": "2026-01-01"}}
{"path": "documents/return-policy.md", "metadata": {"docType": "policy", "productLine": "all", "effectiveDate": "2026-01-01"}}
{"path": "documents/warranty-policy.md", "metadata": {"docType": "policy", "productLine": "phone", "effectiveDate": "2026-03-01"}}
{"path": "documents/phone-manual.md", "metadata": {"docType": "manual", "productLine": "phone", "effectiveDate": "2026-03-01"}}
{"path": "documents/invoice-faq.md", "metadata": {"docType": "faq", "productLine": "all", "effectiveDate": "2026-02-01"}}

File names should reflect the document topics, and the content should contain complete titles and context, so that retrieval can match real questions, including the sample questions configured in Step 3. Prefer importing officially effective policies and avoid keeping conflicting old versions at the same time.

Step 3: Configure runtime parameters

  • Create the project directory and install dependencies.

mkdir -p kb-demo/documents kb-demo/templates
cd kb-demo
python3 -m venv .venv
source .venv/bin/activate
pip install Flask==3.1.1 requests==2.32.4 alibabacloud-milvusknowledgebase20260604==1.0.0

The packages above install additional dependencies automatically, such as alibabacloud_tea_openapi (imported by kb_client.py) and Jinja2 (used by Flask for template rendering). You do not need to install them separately.

  • Create config.json and fill in your own credentials and knowledge base information.

{
  "aliyun": {
    "access_key_id": "YOUR_ACCESS_KEY_ID",
    "access_key_secret": "YOUR_ACCESS_KEY_SECRET",
    "region_id": "cn-hangzhou",
    "knowledge_base_id": "kd-xxxxxxxxxxxxx",
    "knowledge_base_version": "LATEST_PUBLISHED"
  },
  "llm": {
    "enabled": true,
    "base_url": "https://YOUR_OPENAI_COMPATIBLE_ENDPOINT/v1",
    "api_key": "YOUR_LLM_API_KEY",
    "model": "YOUR_MODEL_NAME"
  },
  "upload": {
    "default_meta_fields": {}
  },
  "retrieval": {
    "page_size": 6,
    "candidate_count": 48,
    "min_score": 0.35,
    "semantic_weight": 0.7,
    "enable_query_expansion": true,
    "rerank_model_name": "qwen3-rerank",
    "tag_filter": {
      "relation": "and",
      "conditions": []
    }
  },
  "scenario": {
    "title": "Intelligent Customer Service Knowledge Base Q&A",
    "system_prompt": "You are a corporate customer service assistant. Answer based only on the retrieved documents, in a clear and friendly customer service tone; cite [Source N] for key conclusions. If the retrieved documents do not directly cover the user's question, you must answer 'The available documents do not explicitly cover this.' Do not infer conclusions from partially relevant documents, and do not add contact methods or service channels beyond the documents.",
    "image_enabled": false,
    "sample_questions": [
      "How long after payment does it usually take to ship the product?",
      "Can the product still be repaired after the warranty expires?",
      "What conditions must be met to request a return?"
    ]
  }
}

Considerations for the retrieval parameter values in this scenario:

  • min_score=0.35: reduces low-relevance content from entering answers.

  • semantic_weight=0.7: leans toward semantic search to cover colloquial customer service questions.

  • rerank_model_name: re-ranks candidate chunks to improve the accuracy of the top result.

  • tag_filter.conditions: left blank means no filtering. For tag filtering, see Step 7.

Step 4: Write the knowledge base client

Save the following content as kb_client.py. It encapsulates upload with tags and search with filtering.

"""Milvus knowledge base OpenAPI SDK: local upload with tags and search against a published version."""

from __future__ import annotations

from dataclasses import dataclass
from pathlib import Path
from typing import Any, Mapping, Sequence

import requests
from alibabacloud_milvusknowledgebase20260604 import models as models
from alibabacloud_milvusknowledgebase20260604.client import Client
from alibabacloud_tea_openapi import models as openapi_models

@dataclass(frozen=True)
class LocalDocument:
    path: Path
    object_path: str

    @classmethod
    def from_path(cls, value: str | Path) -> "LocalDocument":
        path = Path(value).expanduser().resolve()
        if not path.is_file():
            raise FileNotFoundError(path)
        return cls(path=path, object_path=path.name)

@dataclass(frozen=True)
class TagCondition:
    field: str
    op: str
    value: Any

@dataclass(frozen=True)
class SearchOptions:
    version: str = "LATEST_PUBLISHED"
    page_size: int = 6
    candidate_count: int = 48
    min_score: float = 0.0
    semantic_weight: float = 0.5
    enable_query_expansion: bool = True
    rerank_model_name: str | None = None
    tag_relation: str = "and"
    tag_conditions: tuple[TagCondition, ...] = ()

class KnowledgeBaseClient:
    def __init__(
        self,
        access_key_id: str,
        access_key_secret: str,
        region_id: str,
    ) -> None:
        endpoint = f"milvusknowledgebase.{region_id}.aliyuncs.com"
        self.client = Client(
            openapi_models.Config(
                access_key_id=access_key_id,
                access_key_secret=access_key_secret,
                region_id=region_id,
                endpoint=endpoint,
                connect_timeout=10_000,
                read_timeout=60_000,
            )
        )
        self.http = requests.Session()

    @staticmethod
    def _check(body: Any, action: str) -> None:
        # Successful responses do not return the code field (its value is None), so treat only an explicit non-zero code as a failure.
        if body is None:
            raise RuntimeError(f"{action} returned an empty response")
        if getattr(body, "success", None) is False or (
            getattr(body, "code", None) not in (None, 0, "0")
        ):
            raise RuntimeError(
                f"{action} failed: {getattr(body, 'message', 'unknown')}; "
                f"requestId={getattr(body, 'request_id', '')}"
            )

    def upload(
        self,
        knowledge_base_id: str,
        file_paths: Sequence[str | Path],
        meta_fields: Mapping[str, Any] | None = None,
    ) -> dict[str, Any]:
        docs = [LocalDocument.from_path(path) for path in file_paths]
        presign_docs = [
            models.GetKnowledgeBasePreSignedUrlRequestDocuments(
                path=doc.object_path,
                name=doc.path.name,
                size=doc.path.stat().st_size,
            )
            for doc in docs
        ]
        response = self.client.get_knowledge_base_pre_signed_url(
            knowledge_base_id,
            models.GetKnowledgeBasePreSignedUrlRequest(
                knowledge_base_id=knowledge_base_id,
                documents=presign_docs,
                expires_in=3600,
            ),
        )
        body = response.body
        self._check(body, "GetKnowledgeBasePreSignedUrl")

        urls = list(body.data.pre_signed_urls or [])
        if len(urls) != len(docs):
            raise RuntimeError("The number of pre-signed URLs does not match the number of files")

        for doc, url in zip(docs, urls, strict=True):
            with doc.path.open("rb") as source:
                # The pre-signed URL is signed with an empty Content-Type; do not include Content-Type in the request.
                self.http.put(url, data=source, timeout=120).raise_for_status()

        add_docs = [
            models.AddDocumentsRequestDocuments(
                path=doc.object_path,
                name=doc.path.name,
                size=doc.path.stat().st_size,
            )
            for doc in docs
        ]
        response = self.client.add_documents(
            knowledge_base_id,
            models.AddDocumentsRequest(
                knowledge_base_id=knowledge_base_id,
                import_type="LOCAL_UPLOAD",
                documents=add_docs,
                meta_fields=dict(meta_fields) if meta_fields else None,
                dedup=models.AddDocumentsRequestDedup(
                    doc_name_dedup=True,
                    content_dedup=False,
                ),
            ),
        )
        body = response.body
        self._check(body, "AddDocuments")

        errors = list(getattr(body.data, "errors", None) or [])
        if errors:
            raise RuntimeError(f"Failed to register data: {errors}")
        return body.to_map()

    def search(
        self,
        knowledge_base_id: str,
        query: str,
        options: SearchOptions,
        image_url: str | None = None,
    ) -> dict[str, Any]:
        tag_filter = None
        if options.tag_conditions:
            tag_filter = models.SearchKnowledgeBaseRequestTagFilter(
                relation=options.tag_relation,
                conditions=[
                    models.SearchKnowledgeBaseRequestTagFilterConditions(
                        field=condition.field,
                        op=condition.op,
                        value=condition.value,
                    )
                    for condition in options.tag_conditions
                ],
            )
        response = self.client.search_knowledge_base(
            knowledge_base_id,
            models.SearchKnowledgeBaseRequest(
                query=query,
                version=options.version,
                page_number=1,
                page_size=options.page_size,
                rerank_model_name=options.rerank_model_name,
                tag_filter=tag_filter,
                image=(
                    models.SearchKnowledgeBaseRequestImage(url=image_url)
                    if image_url
                    else None
                ),
                retrieval_config=models.SearchKnowledgeBaseRequestRetrievalConfig(
                    candidate_count=options.candidate_count,
                    min_score=options.min_score,
                    semantic_weight=options.semantic_weight,
                    enable_query_expansion=options.enable_query_expansion,
                ),
            ),
        )
        body = response.body
        self._check(body, "SearchKnowledgeBase")
        return body.to_map()

Step 5: Write the batch upload script

Save the following content as upload.py. The MetaFields of AddDocuments apply to the whole batch, so the script first groups documents by tag and then submits them in batches of 100.

"""Batch upload local documents by tag; after upload, wait for parsing in the console and publish a version."""

from __future__ import annotations

import argparse
import json
from collections import defaultdict
from dataclasses import dataclass
from pathlib import Path
from typing import Any

from kb_client import KnowledgeBaseClient

@dataclass(frozen=True)
class ManifestEntry:
    path: Path
    metadata: dict[str, Any]

def load_manifest(path: Path) -> list[ManifestEntry]:
    entries: list[ManifestEntry] = []
    for line_number, raw_line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
        if not raw_line.strip():
            continue
        value = json.loads(raw_line)
        file_path = (path.parent / str(value["path"])).resolve()
        metadata = value.get("metadata") or {}
        if not isinstance(metadata, dict):
            raise ValueError(f"metadata on line {line_number} of the manifest must be an object")
        entries.append(ManifestEntry(file_path, metadata))
    return entries

def discover(paths: list[str], default_metadata: dict[str, Any]) -> list[ManifestEntry]:
    entries: list[ManifestEntry] = []
    for value in paths:
        path = Path(value).expanduser()
        files = sorted(item for item in path.rglob("*") if item.is_file()) if path.is_dir() else [path]
        entries.extend(ManifestEntry(item.resolve(), default_metadata) for item in files)
    return entries

def main() -> None:
    parser = argparse.ArgumentParser()
    parser.add_argument("paths", nargs="*", help="Files or directories; you can pass multiple")
    parser.add_argument("--config", default="config.json")
    parser.add_argument("--manifest", help="JSONL file; each line contains path and metadata")
    args = parser.parse_args()

    config = json.loads(Path(args.config).read_text(encoding="utf-8"))
    aliyun = config["aliyun"]
    upload_config = config.get("upload") or {}
    default_metadata = upload_config.get("default_meta_fields") or {}
    entries = (
        load_manifest(Path(args.manifest).expanduser().resolve())
        if args.manifest
        else discover(args.paths, default_metadata)
    )
    if not entries:
        raise SystemExit("No files found to upload")

    grouped: dict[str, list[ManifestEntry]] = defaultdict(list)
    for entry in entries:
        key = json.dumps(entry.metadata, ensure_ascii=False, sort_keys=True)
        grouped[key].append(entry)

    client = KnowledgeBaseClient(
        aliyun["access_key_id"], aliyun["access_key_secret"], aliyun["region_id"]
    )
    submitted = 0
    for metadata_key, group in grouped.items():
        metadata = json.loads(metadata_key)
        for start in range(0, len(group), 100):
            batch = group[start : start + 100]
            client.upload(
                aliyun["knowledge_base_id"],
                [entry.path for entry in batch],
                meta_fields=metadata,
            )
            submitted += len(batch)
            print(f"Submitted {submitted}/{len(entries)} documents; metadata={metadata}")

if __name__ == "__main__":
    main()

Run the upload:

python3 upload.py --manifest documents.jsonl

The output echoes batches grouped by tag. For example, 5 documents with 4 tag combinations produce 4 batches:

Submitted 2/5 documents; metadata={'docType': 'policy', 'effectiveDate': '2026-01-01', 'productLine': 'all'}
Submitted 3/5 documents; metadata={'docType': 'policy', 'effectiveDate': '2026-03-01', 'productLine': 'phone'}
Submitted 4/5 documents; metadata={'docType': 'manual', 'effectiveDate': '2026-03-01', 'productLine': 'phone'}
Submitted 5/5 documents; metadata={'docType': 'faq', 'effectiveDate': '2026-02-01', 'productLine': 'all'}
Important

A successful return from the upload API only means that asynchronous parsing has been submitted. Go to the Data Management page in the console and confirm that the document status becomes Processing Complete before you proceed to the next step. The tag column shows the written tags, such as productLine=all, docType=faq +1.

Step 6: Publish a version

After parsing completes, documents remain unpublished and can be searched only after you publish a version. This operation is currently supported only in the console.

  1. Open the knowledge base details page, click the Version Management tab, confirm that pending changes exist, and click Publish Version.

  2. In the Confirm Changes step, review the change log (newly added data is listed one by one) and click Next.

  3. In the Enter Description step, enter a release note (up to 200 characters) and click Publish.

  4. In Version History, confirm that the new version status is Published and record the version number. The first publish is v1, followed by v2 and v3.

    When knowledge_base_version in config.json uses LATEST_PUBLISHED, the latest published version is searched automatically. You can also specify an explicit version number. Do not use DRAFT to provide a stable Q&A service.

Version quotas and deletion

  • Default quota — By default, at most 3 published versions can exist at the same time.

  • Quota scope — This limit is calculated per tenant and does not increase with the instance CU specification.

  • Quota increase — To raise the limit, submit a ticket for evaluation. There is currently no self-service quota application entry for users.

  • Behavior at the limit — When the limit is reached, the Publish Version button is grayed out, while the page still shows "There are N pending changes to publish". Delete old versions that are no longer needed in Version History before publishing again.

  • Frequent updates — If documents are updated frequently, keep only the current effective version plus the most recent historical version.

Warning

Version deletion is irreversible. After deletion, the version immediately becomes unsearchable (the backend cleans up the data asynchronously). Before deleting, make sure that no application is pinned to that version number, and switch applications still using it to a new version.

Step 7: Search with tag filtering

Configure filter conditions in retrieval.tag_filter.conditions of config.json to restrict the search scope to specified tags. For example, to search only policy documents:

{
  "tag_filter": {
    "relation": "and",
    "conditions": [
      {"field": "docType", "op": "=", "value": "policy"}
    ]
  }
}

relation supports and (all conditions met) and or (any condition met). op supports the following operators:

Operator Observed behavior
= Effective. Exact match (for int64 tags, either a number or a string works)
in, not in Effective. Whether the value belongs to or does not belong to the given set
≠, >, <, ≥, ≤, empty, not empty, start with, end with Not effective. The condition is ignored, all data is returned, and no error is reported
contains, not contains Not recommended. Aliases for in/not in, not string containment

The only operators that actually work are =, in, and not in. The other operators appear in the "Supported operators" list of the 400 Unsupported tag filter operator error, but tests show that when they are passed, the condition is silently ignored and all unfiltered data is returned. In addition, contains and not contains are only aliases for in/not in, not string containment. Passing a string fragment returns 0 results.

To work around these limits:

  • To filter by numeric or date ranges, use in with enumerated values instead.

  • To check for an empty tag, use = "" instead.

  • After you configure filter conditions, always compare the total result count against the unfiltered result to confirm that the filter actually takes effect.

    Also, if field is set to an undefined tag name, the API likewise reports no error and simply returns 0 results. Setting op to aliases such as eq, ==, equal, or like returns 400 Unsupported tag filter operator.

Using the corpus in this topic as an example, the search scope under different filter conditions:

Filter condition Matched documents
No conditions set All documents
docType = policy Shipping, return and exchange, and warranty policies
docType = faq Invoice FAQ
docType = policy and productLine = phone Warranty policy
docType = faq or docType = manual (relation is or) Invoice FAQ, phone manual
effectiveDate = 2026-01-01 The two policies effective on that date

Step 8: Write the Q&A service

Save the following content as app.py.

"""Knowledge base retrieval + OpenAI-compatible LLM Q&A service."""

from __future__ import annotations

import json
from pathlib import Path
from typing import Any

import requests
from flask import Flask, jsonify, render_template, request

from kb_client import KnowledgeBaseClient, SearchOptions, TagCondition

CONFIG = json.loads(Path("config.json").read_text(encoding="utf-8"))
ALIYUN = CONFIG["aliyun"]
LLM = CONFIG.get("llm", {})
SCENARIO = CONFIG.get("scenario", {})
RETRIEVAL = CONFIG.get("retrieval", {})
KB = KnowledgeBaseClient(
    ALIYUN["access_key_id"], ALIYUN["access_key_secret"], ALIYUN["region_id"]
)
app = Flask(__name__)

def search_options() -> SearchOptions:
    raw_conditions = (RETRIEVAL.get("tag_filter") or {}).get("conditions") or []
    return SearchOptions(
        version=ALIYUN.get("knowledge_base_version", "LATEST_PUBLISHED"),
        page_size=int(RETRIEVAL.get("page_size", 6)),
        candidate_count=int(RETRIEVAL.get("candidate_count", 48)),
        min_score=float(RETRIEVAL.get("min_score", 0.0)),
        semantic_weight=float(RETRIEVAL.get("semantic_weight", 0.5)),
        enable_query_expansion=bool(RETRIEVAL.get("enable_query_expansion", True)),
        rerank_model_name=str(RETRIEVAL.get("rerank_model_name") or "") or None,
        tag_relation=str((RETRIEVAL.get("tag_filter") or {}).get("relation", "and")),
        tag_conditions=tuple(
            TagCondition(str(item["field"]), str(item["op"]), item.get("value"))
            for item in raw_conditions
        ),
    )

def find_results(payload: Any) -> list[dict[str, Any]]:
    """Extract the list of retrieved chunks from the response."""
    if isinstance(payload, list):
        return [item for item in payload if isinstance(item, dict)]
    if not isinstance(payload, dict):
        return []
    for key in ("results", "Results"):
        if isinstance(payload.get(key), list):
            return payload[key]
    for key in ("data", "Data"):
        found = find_results(payload.get(key))
        if found:
            return found
    return []

def field(item: dict[str, Any], *names: str) -> Any:
    for name in names:
        if item.get(name) not in (None, ""):
            return item[name]
    return ""

def llm_answer(question: str, results: list[dict[str, Any]]) -> str:
    if not LLM.get("enabled", True):
        return "The LLM is disabled. See the retrieval results below."
    if not results:
        return "The available documents do not explicitly cover this."
    context = "\n\n".join(
        f"[Source {index}] {field(item, 'documentName', 'DocumentName')}\n"
        f"{field(item, 'content', 'Content')}"
        for index, item in enumerate(results, 1)
    )
    url = str(LLM["base_url"]).rstrip("/") + "/chat/completions"
    response = requests.post(
        url,
        headers={"Authorization": f"Bearer {LLM['api_key']}"},
        json={
            "model": LLM["model"],
            "temperature": 0.1,
            "messages": [
                {"role": "system", "content": SCENARIO.get("system_prompt", "Answer based only on the documents.")},
                {"role": "user", "content": f"Question: {question}\n\nRetrieved documents:\n{context}"},
            ],
        },
        timeout=90,
    )
    response.raise_for_status()
    return response.json()["choices"][0]["message"]["content"].strip()

@app.get("/")
def index():
    return render_template(
        "index.html",
        title=SCENARIO.get("title", "Knowledge Base Q&A"),
        sample_questions=SCENARIO.get("sample_questions", []),
        image_enabled=bool(SCENARIO.get("image_enabled", False)),
    )

@app.post("/api/ask")
def ask():
    payload = request.get_json(silent=True) or {}
    question = str(payload.get("question", "")).strip()
    image_url = str(payload.get("image_url", "")).strip() or None
    if not question:
        return jsonify({"error": "The question cannot be empty"}), 400
    try:
        raw = KB.search(
            ALIYUN["knowledge_base_id"],
            question,
            options=search_options(),
            image_url=image_url,
        )
        results = find_results(raw)
        return jsonify({"answer": llm_answer(question, results), "sources": results})
    except Exception as exc:
        return jsonify({"error": str(exc)}), 500

if __name__ == "__main__":
    app.run(host="127.0.0.1", port=7860, debug=False)

Save the following content as templates/index.html.

<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width,initial-scale=1">
  <title>{{ title }}</title>
  <style>
    body{margin:0;background:#f6f7fb;color:#1f2937;font:15px system-ui,sans-serif}
    main{max-width:860px;margin:0 auto;padding:42px 18px}
    .card{background:#fff;border:1px solid #e5e7eb;border-radius:18px;padding:24px}
    textarea{box-sizing:border-box;width:100%;min-height:100px;border:1px solid #d1d5db;border-radius:12px;padding:14px;font:inherit}
    button{margin-top:12px;border:0;border-radius:10px;padding:11px 18px;background:#4f46e5;color:#fff;cursor:pointer}
    .chip{background:#eef2ff;color:#3730a3;margin:4px;padding:7px 10px}
    pre{white-space:pre-wrap;line-height:1.65}.muted{color:#6b7280}.source{border-top:1px solid #eee;padding:12px 0}
  </style>
</head>
<body><main><h1>{{ title }}</h1><p class="muted">Answers are generated from the published knowledge base content, with retrieval sources displayed.</p>
  <div>{% for q in sample_questions %}<button class="chip" onclick='setQ({{ q|tojson }})'>{{ q }}</button>{% endfor %}</div>
  <section class="card">
    <textarea id="q" placeholder="Enter your question"></textarea>
    <button id="ask" onclick="ask()">Send</button>
    <pre id="answer"></pre>
    <div id="sources"></div>
  </section>
</main><script>
const q=document.querySelector('#q'), answer=document.querySelector('#answer'), sources=document.querySelector('#sources');
function setQ(value){q.value=value;q.focus()}
async function ask(){
  const text=q.value.trim(); if(!text) return;
  answer.textContent='Retrieving and generating...'; sources.innerHTML='';
  const res=await fetch('/api/ask',{method:'POST',headers:{'Content-Type':'application/json'},body:JSON.stringify({question:text})});
  const data=await res.json();
  answer.textContent=data.answer||('Error: '+data.error);
  for(const [i,item] of (data.sources||[]).entries()){
    const div=document.createElement('div'); div.className='source';
    div.textContent=`Source ${i+1}: ${item.documentName||''}\n${item.content||''}`;
    sources.appendChild(div);
  }
}
</script></body></html>
Important

The onclick attribute of the sample question buttons must be wrapped in single quotes (onclick='setQ({{ q|tojson }})'). If you use double quotes, the JSON double quotes produced by tojson close the HTML attribute early, and the buttons stop responding to clicks.

Step 9: Start and verify

  • Start the service.

python3 app.py
  • Open http://127.0.0.1:7860 in a browser, click a sample question or type one manually, and click Send.

  • You can also call the API directly to verify.

curl -sS http://127.0.0.1:7860/api/ask \
  -H 'Content-Type: application/json' \
  -d '{"question":"How long after payment does it usually take to ship the product?"}'

Verify the following four acceptance criteria:

  • The page opens normally, and submitting a question returns both the answer and the retrieval sources.

  • Each [Source N] in the answer has a matching document chunk in the sources area below.

  • When a question covers content not in the documents, the answer is "The available documents do not explicitly cover this", rather than fabricated facts.

  • After you add documents and publish a version again in the console, the page can search the new content.

Tuning recommendations

  • Adjust chunk length based on document shape: customer service documents are mostly short FAQ entries and policy clauses. In Processing Policies on the knowledge base details page, click Create Policy, and choose Smart Splitting or Split by Length as the chunking method. Note that the maximum segment length is measured in characters (default 512 characters). For short-entry scenarios, set it to 380 to 580 characters (roughly 256 to 384 tokens). After the policy is created, specify it during import through AddDocumentsRequest.strategy_id. The upload.py script in this tutorial does not set this field; to use a custom processing policy, add the field to the AddDocuments request.

  • Tune reranking and min_score together: after you enable rerank_model_name, the rerank score and the vector similarity score are not on the same scale. Keeping the original min_score may trim too few results (in tests, three customer service questions each kept only one hit after reranking was enabled). If a question requires a combined answer from multiple documents, lower min_score appropriately when you enable reranking.

  • Watch out for wrong conclusions caused by partially relevant documents: when a question is literally related to a document but not semantically covered (for example, the document only describes shipping times while the user asks whether cash on delivery is supported), the LLM may draw a wrong conclusion from it. In system_prompt, explicitly require the model to answer that it does not know when the documents do not directly cover the question and forbid inference from partially relevant documents, and spot-check with real customer questions before launch.

  • Test with real customer questions, not only manual titles. A higher semantic_weight (such as 0.7) better covers colloquial expressions.

  • Prefer importing officially effective policies and avoid keeping conflicting old versions at the same time. After policy updates, publish a version again and distinguish versions with the effectiveDate tag.

Understand search scores

Each search result also returns scoreDetails, for example {"keywordScore": 0.368, "semanticScore": 0.819}, which helps you judge whether the hit comes mainly from keyword or semantic matching.

The three scores mean the following:

  • keywordScore is the keyword similarity.

  • semanticScore is the vector similarity when reranking is disabled, and the rerank model score when reranking is enabled.

  • score is the final ranking and filtering score obtained by weighting the two, calculated as score ≈ (1-semantic_weight) × keywordScore + semantic_weight × semanticScore (a rank feature may also be added). min_score filters on this final score.

    The score scale changes across models, with or without reranking, and with different weights. Scores cannot be compared across configurations, and there is no universally recommended threshold: set min_score to 0 first, retrieve a batch of results and manually label their relevance, then pick a threshold based on the distribution of recalls and false recalls. Recalibrate every time you adjust semantic_weight or toggle reranking.

FAQ

Symptom Cause and solution
Search returns 400 Unsupported tag filter operator op uses aliases such as eq, ==, or like. Use =, in, or not in instead.
Tag filtering is added but the result count is exactly the same as without filtering A non-effective operator (such as >, ≥, ≠, or empty) is used. Only =, in, and not in take effect.
Tag filtering always returns 0 results field is set to an undefined tag name (the API reports no error). Verify the tag name spelling on the knowledge base details page → Tags → Manage; also confirm that tags were actually written for these documents during import. Defining tags in the console is not a prerequisite for writing tag values. For details, see Usage notes in Step 1.
The Publish Version button is grayed out, but the page shows pending changes The limit of 3 published versions has been reached (hover over the button to see the hint). Delete old versions in Version History, then publish.
Search returns 404 Knowledge base version ... does not exist No version has been published, or knowledge_base_version does not match the actual version number. Publish a version in the console first.
Upload succeeds but nothing is searchable Upload and parsing are asynchronous. Wait until the Data Management page in the console shows Processing Complete, and publish a version again.
Calls return 401 or 403 The AccessKey is invalid, or the RAM user does not have AliyunMilvusFullAccess.
Occasional connection failures when PUTting to OSS during upload Network jitter. Simply retry that file.
Historical documents have no tags and cannot be found by tag filtering Tags are written during import; documents imported before the tag definitions are not matched by tag filtering. You can Set Tags for individual documents on the Data Management page, or re-import them.

Pre-launch checklist

  • Credential storage — Store the AccessKey pair and the LLM API key only in the local config.json. Do not commit them to code repositories or share them in chat groups. After the demo, rotate or delete the temporary credentials promptly.

  • Least privilege — Use a dedicated RAM user with least privilege that can be rotated. Do not use the Alibaba Cloud account AccessKey long term.

  • Authorized content — Import only documents you are authorized to handle. Every answer must be traceable to a source chunk.

  • Production deployment — For external deployment, do not keep using the Flask development server. Use a production WSGI server and add authentication, HTTPS, access log masking, rate limiting, and auditing.

  • Human fallback — For critical policy Q&A, keep a human fallback entry.