In this tutorial, you build a development troubleshooting assistant on an Alibaba Cloud Milvus knowledge base that holds your API references, architecture notes, Runbooks, and postmortems. You search by API name or error code and get ordered troubleshooting steps with source citations.
Solution overview
This tutorial builds on Build an intelligent customer service Q&A application with the Alibaba Cloud Milvus knowledge base. The end-to-end pipeline is identical: define tags in the console → batch-import source files by tag → publish a version → retrieve through the SDK (with optional tag filtering) → have the LLM generate answers with source citations → serve the Q&A page with Flask. You reuse the engineering code (kb_client.py, upload.py, app.py, templates/index.html, and start.sh) from that tutorial, with two modifications described in Step 4: Adapt the reused code. This document covers only what you must change for the development troubleshooting scenario: module tags, retrieval parameters for exact terms, handling of HTML source files, and constraints on high-risk commands and empty results.
For the first version, select only 5 to 20 source files from a single system and expand the scope after validation. The entire process takes about 20 to 30 minutes, and the document parsing time depends on the number and size of the source files.
Prerequisites
A Milvus knowledge base is created, and you have recorded the knowledge base ID (for example,
kd-803ae9b10cc31).A RAM user created with your Alibaba Cloud account, with Use Permanent AccessKey Access selected and the system policy
AliyunMilvusFullAccessattached.An LLM endpoint that supports the OpenAI
chat/completionsprotocol and its API key are ready.Python 3.8 or later is installed locally.
The project directory and the five reused files (
kb_client.py,upload.py,app.py,templates/index.html, andstart.sh) are set up by following the tutorial referenced in Solution overview.The source files contain no secrets, tokens, or internal network accounts.
Step 1: Define module tags
Under Basic Information on the knowledge base details page, choose Tags > Manage and add the following three tags. Set the field type of each tag to string. The type dropdown is on the right side of the tag name input box, and the options are string, int64, list, float32, and bool.
| Tag name | Description | Example value |
| module | The system or service that the source file belongs to | order-service, user-service |
| docType | The document type | API, Runbook, ErrorCode, Postmortem |
| version | The API or document version | v2 |
The tag name version has the same name as the version parameter that represents the published version of the knowledge base in a retrieval request, but the two do not affect each other: the tag is used in tagFilter.field, and the retrieval version sits at the top level of the request. If you are concerned about confusion, name the tag apiVersion instead.
The tag management dialog box is saved as a whole. After you open the dialog box, the tag list loads asynchronously. Wait until all existing tags are displayed before you add new tags and click OK. Otherwise, the save may overwrite everything with "empty list + new tags" and wipe out the existing tag definitions.
The module tag is the most critical design decision in this scenario. Different systems often have API paths with exactly the same name, and without module filtering they interfere with each other. In our test, user-service and order-service both own the same path GET /api/v2/orders/{orderId}. When we searched for this path:
| Retrieval condition | Result |
| No filter | The user-service source file ranked first (relevance 0.246) — the wrong module was hit |
module = order-service | The order-service source file ranked first, and the user-service source file was completely excluded |
Step 2: Prepare source files and the import manifest
In this step, you collect the source files, annotate each of them in the import manifest, and configure how the knowledge base chunks them.
Place the source files in the
documents/directory. Markdown, HTML, PDF, DOCX, and TXT are supported. Keep complete error codes, API paths, and version numbers in the source files — in our tests, both error codes and full paths could be retrieved exactly.Create
documents.jsonland annotate each source file with its module, document type, and version:{"path": "documents/order-api.md", "metadata": {"module": "order-service", "docType": "API", "version": "v2"}} {"path": "documents/order-timeout-runbook.md", "metadata": {"module": "order-service", "docType": "Runbook", "version": "v2"}} {"path": "documents/order-error-codes.html", "metadata": {"module": "order-service", "docType": "ErrorCode", "version": "v2"}} {"path": "documents/order-postmortem-2026-06.md", "metadata": {"module": "order-service", "docType": "Postmortem", "version": "v2"}}Import the API reference, Runbook, error code table, and postmortem as a set, and distinguish them with
docType. In our test, when we asked about an error code, all four types of source files were hit at the same time, and the model produced a complete answer of "meaning → troubleshooting steps → historical cases".
Handle HTML source files
HTML can be uploaded and parsed directly, and the content inside <table> does enter the index. In our test, querying ORD-42901, which exists only in an HTML error code table, hit that file exactly.
However, HTML has a coarser chunk granularity than Markdown (in this test, an HTML file containing two tables was parsed into 1 chunk). This is related to how HTML is parsed. Before you choose a source file format, understand two points:
Code blocks do not preserve their original layout.
<pre>and<code>are recognized as blocks, but the text is extracted recursively and joined with spaces, so indentation, line breaks, and code fences are all lost. For source files with large amounts of code or commands, convert them to Markdown before import; otherwise, the retrieved code snippets may not be directly usable.Tables are indexed as a whole. A
Conclusion: plain descriptive text and small tables can use HTML directly. When you need exact fidelity, retrieval by code segment, or splitting of large tables, prefer Markdown or structured source files, and after import, spot-check the chunks on the Data Management page.<table>is appended as a single independent segment in its original HTML string form and is not split by row, so large tables easily form coarse chunks.
Configure the chunking policy
Troubleshooting steps and code examples must not be cut in the middle. On the knowledge base details page, click Create Policy in Processing Policy to adjust the chunk granularity. The unit of the maximum segment length is characters (default 512). Set it to 580–770 characters (about 384–512 tokens) so that "one complete troubleshooting step" or "one code example" lands in the same chunk. The character-to-token ratio varies with the language and tokenizer of your source files, so use the character value as the baseline and spot-check the chunks after import.
Step 3: Configure retrieval parameters and the prompt
Development troubleshooting queries are mostly error codes, API paths, and commands. They belong to exact term matching, so the parameters differ greatly from other scenarios.
{
"retrieval": {
"page_size": 6,
"candidate_count": 64,
"min_score": 0.05,
"semantic_weight": 0.15,
"enable_query_expansion": false,
"rerank_model_name": "",
"tag_filter": {
"relation": "and",
"conditions": [
{"field": "module", "op": "=", "value": "order-service"}
]
}
},
"scenario": {
"title": "Development docs and troubleshooting assistant",
"system_prompt": "You are a development documentation assistant. Answer only based on the retrieved materials. Preserve API names, error codes, commands, and code as-is. List troubleshooting steps in order and cite them as [Source N]. When write operations or high-risk commands are involved, explicitly state the risk level and require manual confirmation. When the retrieved materials are empty, reply only that the materials are insufficient and warn against performing any change operation.",
"image_enabled": false,
"sample_questions": [
"What are the required parameters for the order query API?",
"In what order should I troubleshoot a connection timeout?",
"In which documents does this error code appear?"
]
}
}The following list describes the parameters tuned for this scenario. The remaining fields in the example carry over from the referenced tutorial.
semantic_weight: 0.15: The weighting factor of the final score. The score is calculated asscore ≈ (1-semantic_weight) × keywordScore + semantic_weight × semanticScore(a rank feature may also be added), andmin_scoreperforms post-filtering on this final score. This scenario uses 0.15 so that the keyword score of exact terms such as error codes and API paths dominates the ranking. Becausesemantic_weightinteracts withmin_score, tune the two together as described in Tune semantic_weight and min_score together.enable_query_expansion: false: Query expansion is disabled so that error codes and API names are not rewritten. In our test, after enabling it, thekeywordScoreof the queryORD-50021dropped from 0.283 to 0.226, and the number of recalled results for the queryORD-42901dropped from 2 to 1. Keep it disabled for exact term retrieval.tag_filterpinsmodule: This prevents identically named APIs of different systems from interfering with each other. See Step 1: Define module tags for the effect. If one entry point needs to cover multiple modules, use{"field": "module", "op": "in", "value": ["order-service", "user-service"]}.A blank
rerank_model_namemeans reranking is not enabled. Exact match for error codes relies on the keyword score, so reranking brings limited benefit.opsupports only=,in, andnot in: In our tests,≠,>,≥,<,≤,empty,not empty,start with, andend withare silently ignored and all data is returned (no error is reported).containsandnot containsare merely aliases ofin/not in, not string containment; passing string fragments returns 0 results. Aliases such aseq,==, andlikereturn400 Unsupported tag filter operator, and every question fails. After you configure the filter conditions, compare the total number of results with the unfiltered results to confirm that the filter actually takes effect.
Tune semantic_weight and min_score together
semantic_weight changes only the final weighted score; it does not change the two sub-scores (keywordScore and semanticScore) in scoreDetails of the retrieval results. The weighted score is approximately:
score ≈ semantic_weight × semanticScore + (1 - semantic_weight) × keywordScoreNote two points before you tune:
Lowering
semantic_weightdoes not necessarily lower the total score. The total score drops only when thesemanticScoreof the result batch is higher than itskeywordScore. This is why this document also lowersmin_scoreto 0.05. The two parameters must be calibrated together by checkingscoreDetails.When you set
For development materials,semantic_weightto 0, the current implementation no longer applies themin_scorethreshold. Do not use 0 to mean "pure keyword retrieval".keywordScoreis usually much lower thansemanticScore(in our tests, about 0.16–0.28 and 0.64–0.78 respectively), so under this distribution, loweringsemantic_weightpulls the total score toward the lower side. This is not a universal rule: only when thesemanticScoreof the result batch is higher than itskeywordScoredoes lowering the weight push the total score down; otherwise it pushes it up. Ifmin_scoreis not lowered at the same time, even semantically hit results get filtered out:
| Query | semantic_weight=0.15 | semantic_weight=0.7 |
/api/v2/orders/{orderId} | 8 results, top score 0.269 | 8 results, top score 0.602 |
kubectl -n order rollout undo deploy/order-api | 4 results | 8 results |
| Natural-language query "Why has the order API slowed down significantly recently?" | 3 results | 8 results |
When you use semantic_weight=0.15, set min_score to about 0.05 (the example above is configured this way). If you leave min_score at 0.15 or higher, the recall of command-based and natural-language queries drops by more than half. When tuning, fix one parameter and observe the two sub-scores through scoreDetails before deciding.
Constrain high-risk commands and empty results
The troubleshooting assistant outputs executable commands directly, so two things must be constrained.
High-risk commands. Mark the risk level and confirmation requirements in a table inside the source files, and the model passes them through faithfully. In our test, a restart command marked "high risk, two-person confirmation required" in the Runbook, when asked about, was returned by the model with the command plus the explicit statements "the risk level is high", "two-person confirmation is required", and "do not skip troubleshooting and run the restart directly".
Empty retrieval results. The code must intercept this case, as described in Step 4: Adapt the reused code.
Step 4: Adapt the reused code
The reused project requires two modifications before you verify the assistant: an empty-result fallback in app.py, and a return-code fix in kb_client.py.
Intercept empty retrieval results in app.py
In the reused app.py, llm_answer() still calls the LLM when the retrieval results are empty. At that point the context is an empty string, and the model answers entirely from its own knowledge. In our test, when we asked "How do I recover from a Redis cluster split-brain?", which the knowledge base does not cover, the model output a complete operations plan with parameters and gave no warning at all. In troubleshooting scenarios, users may copy the commands and run them directly against the production environment, so this must be intercepted at the code level:
def llm_answer(question: str, results: list[dict[str, Any]]) -> str:
if not LLM.get("enabled", True):
return "The LLM is not enabled. Review the retrieval results below."
if not results:
return "No relevant document was retrieved from the knowledge base, so no troubleshooting steps can be provided. Do not perform any change operation based on this. We recommend that you contact the module owner."
...Constraining this in the prompt alone is not reliable — even if you state "notify explicitly when the materials do not cover the question", the model still answers.
The last acceptance criterion in Step 5: Upload, publish, and verify covers this modification.
Fix the return-code check in kb_client.py
In the reused kb_client.py, _check() evaluates getattr(body, "code", 0) != 0. The code of a successful response is None, so successful retrievals are reported as "Failed: None". Change the check to getattr(body, "code", None) not in (None, 0, "0"). After the fix, successful retrievals no longer report "Failed: None".
Step 5: Upload, publish, and verify
Upload the source files.
MetaFieldsapplies to the whole batch; the script first groups by tag and then submits in batches.python upload.py --manifest documents.jsonlOn the Data Management page of the console, confirm that the source file status is Completed. A successful response from the upload API only means that asynchronous parsing has been submitted.
On the Version Management page, click Publish Version. After you complete the wizard, confirm that the status of the new version is Published.
ImportantAt most 3 published versions can exist at the same time. After the limit is reached, the Publish Version button is grayed out, but the page still shows "There are currently N pending changes to publish". Delete old versions that are no longer needed from the version records first. Development documents change frequently (every release may update the API reference and Runbooks), so keep only "the current version + the most recent historical version".
Start the service:
python app.pySend a test question:
curl -sS http://127.0.0.1:7860/api/ask \ -H 'Content-Type: application/json' \ -d '{"question":"In what order should I troubleshoot a connection timeout?"}'Verify against the following five acceptance criteria:
The page opens normally, and after you submit a question, the answer and the retrieval sources are returned together.
Each
[Source N]in the answer can be matched to the corresponding source file chunk in the sources section below.When you ask with a complete error code, all source files in which the error code appears are listed. In our test, querying
ORD-50021correctly listed four source files — the error code table, the API reference, the Runbook, and the postmortem — together with the location in each.When you ask about a risky operation, the answer includes the risk level and the manual confirmation requirement.
When you ask about something the knowledge base does not cover, the response is "materials are insufficient" with a warning not to perform any change operation, and no fabricated command appears. Test this case in practice.
Notes for development scenarios
Entry point per module — After
tag_filterpinsmodule, the entry point answers only questions about that module. Keep the sample questions within the same module as well.Complete exact terms — Keep complete error codes, API paths, commands, and version numbers in the source files. These exact terms are the main retrieval entry points in this scenario.
Risk marking — Mark risk levels and confirmation requirements in the source files. Do not rely on the model to judge which commands are dangerous.
No secrets in uploads — Do not upload secrets, tokens, internal network accounts, or production database connection strings. Chunk content is sent to the LLM as context.
Deprecated API descriptions — Keep the descriptions of deprecated API operations. In our test, after the API reference stated "the v1 API is deprecated and its parameter names are incompatible", the model did not mix in old parameters when answering questions about the parameters of the new API.
The
tagsfield in retrieval results does not echo the module and version that were written. It is a different field from the MetaFields ofAddDocuments, and no parameter can enable its return. An empty value does not mean the upload failed. When you need to annotate module and version in the answer, join the retrieval results withdocuments.jsonlbydocumentIdand append the information to the context.Production deployment — Do not keep using the Flask development server for online deployment. Switch to a production WSGI server and add secret management, authentication, auditing, and rate limiting.
FAQ
| Symptom | Cause and solution |
The question returns 500, and the log shows 400 Unsupported tag filter operator | op uses aliases such as eq, ==, or like. Use =, in, or not in instead. |
| A tag filter is added, but the number of results is exactly the same as without filtering | An ineffective operator (such as >, ≥, ≠, or empty) was used. Only =, in, and not in take effect. |
| Identically named APIs of other systems are retrieved | The module filter is not configured, or the filter value does not match the value written at upload time. |
| Far fewer retrieval results than expected | When semantic_weight is low, the overall scores are pushed down, and combined with min_score, more results are filtered out. Lower min_score as described in Tune semantic_weight and min_score together. |
| Error code retrieval is inaccurate | Confirm that enable_query_expansion is false. When enabled, error codes may be rewritten. |
| Tag filtering always returns 0 results | The tag name spelling does not match the one written at upload time (the API reports no error and simply returns 0 results). Verify it on the knowledge base details page by choosing Tags > Manage. |
| A question not covered by the knowledge base returned seemingly executable commands | llm_answer() still calls the LLM when the retrieval results are empty. Add the empty-result fallback as described in Step 4: Adapt the reused code. |
| Details in HTML source files cannot be retrieved | HTML has a coarser chunk granularity. Lower the maximum segment length. Convert HTML with large amounts of code to Markdown first. |
| The retrieval reports "Failed: None" | The check in _check() of the reused code is incorrect (getattr(body, "code", 0) != 0); the code of a successful response is None. Apply the fix in Step 4: Adapt the reused code. |
The retrieval returns 404 Knowledge base version ... does not exist | No version has been published yet, or the hard-coded version number in the configuration has been deleted. Use LATEST_PUBLISHED instead. |
Next steps
Expand coverage module by module: create an independent entry point for each
modulevalue, and keep the sample questions within the same module.Prepare for production: switch from the Flask development server to a production WSGI server, and add secret management, authentication, auditing, and rate limiting.
Manage versions as your documents change: publish new versions when the API references or Runbooks are updated, and keep only the current version plus the most recent historical version.