All Products
Search
Document Center

Vector Retrieval Service for Milvus:Build an education question bank Q&A application with Alibaba Cloud Milvus knowledge base

Last Updated:Aug 27, 2026

Build a question bank Q&A application for education scenarios on the Alibaba Cloud Milvus knowledge base. The knowledge base turns textbooks, lesson plans, questions, and ground truth into a searchable question bank assistant: students ask questions in natural language and receive explanations together with citations to the source material. Students can also upload a question image to find similar questions in the question bank, combined with textbook explanations.

Solution overview

The end-to-end pipeline is identical to Build an intelligent customer service Q&A application with Alibaba Cloud Milvus knowledge base: define tags in the console → import materials in bulk by tag → publish a version → SDK retrieval (tag filtering and image attachment supported) → the LLM generates answers with source citations → Flask serves the Q&A page. This tutorial reuses the project code directly from that tutorial. This topic describes only the parts that change for education scenarios: subject and knowledge point tags, how image materials are handled, the capability boundaries of image queries, and formula display.

For the first version, select only 5 to 20 materials from one grade and one subject, and expand after verification. The whole process takes about 25 to 40 minutes. Document parsing time depends on the number and size of the materials.

Prerequisites

  • You have completed Build an intelligent customer service Q&A application with Alibaba Cloud Milvus knowledge base. This tutorial reuses the following project code from that tutorial:

    • kb_client.py, upload.py, app.py, templates/index.html, and start.sh

    • The config.json configuration file

  • An omni-modal (ALL_MODAL) Milvus knowledge base, with the knowledge base ID recorded (such as kd-803ae9b10cc31). Create a new knowledge base for this tutorial, or reuse the one from the customer service tutorial. Note the following:

    • The knowledge base type cannot be changed after creation.

    • A structured knowledge base accepts only five formats: xlsx, xls, csv, jsonl, and faq. Uploading an image directly returns a parameter error. A question bank that contains image materials must therefore use an omni-modal knowledge base.

    • The console currently always uses the omni-modal (ALL_MODAL) type when creating knowledge bases, and no other option exists on the page, so no extra decision is needed. To check the type of an existing knowledge base, view the Data Type in Basic Information on the knowledge base details page.

  • A RAM user for your Alibaba Cloud account. Select Access with a permanent AccessKey for the RAM user and grant the system policy AliyunMilvusFullAccess.

  • An LLM endpoint that supports the OpenAI chat/completions protocol, and an API key.

  • Python 3.8 or later installed locally.

Step 1: Define grade and knowledge point tags

In the Basic Information section of the knowledge base details page, choose Tags > Manage and add the following four tags. All tags are of type string.

Tag name

Description

Example value

grade

Grade

Grade 8

subject

Subject

Mathematics, Physics

knowledgePoint

Knowledge point

quadratic function, Pythagorean theorem

questionType

Material type

textbook, practice question, ground truth

Tag values support Chinese. You can directly use the Chinese names of grades, subjects, and knowledge points.

Warning

The tag management dialog box is saved as a whole. After the dialog box opens, the tag list loads asynchronously. Wait until all existing tags are displayed before you add new tags and click OK. Otherwise, the whole list may be overwritten with an empty list plus the new tags, and existing tag definitions are deleted.

Step 2: Prepare materials and the import manifest

  • Place your materials in the documents/ directory. PDF, DOCX, Markdown, TXT, and images are supported. Upload textbooks, questions, and ground truth as a set: when generating explanations, the model cites methods from the textbook, the original question text, and the scoring points from the answer at the same time, making answers significantly more complete.

  • Create documents.jsonl and annotate each material with the grade, subject, knowledge point, and material type.

    {"path": "documents/math-grade8-textbook.pdf", "metadata": {"grade": "Grade 8", "subject": "Mathematics", "knowledgePoint": "quadratic function", "questionType": "textbook"}}
    {"path": "documents/math-quadratic-problem.png", "metadata": {"grade": "Grade 8", "subject": "Mathematics", "knowledgePoint": "quadratic function", "questionType": "practice question"}}
    {"path": "documents/math-quadratic-answers.docx", "metadata": {"grade": "Grade 8", "subject": "Mathematics", "knowledgePoint": "quadratic function", "questionType": "ground truth"}}
    {"path": "documents/math-pythagorean-exercise.docx", "metadata": {"grade": "Grade 8", "subject": "Mathematics", "knowledgePoint": "Pythagorean theorem", "questionType": "practice question"}}
  • For image materials, understand how they are imported. Images first go through optical character recognition (OCR) to be converted into text, and the text is then chunked and indexed. Therefore:

    • Whether the text in an image can be recognized determines whether the image can be retrieved. An image with no text, such as a geometric figure or function graph with no text, can hardly be hit and is not suitable for upload as standalone material. Place it in the same image or the same document as a question stem that contains text.

    • Keep question images clear, with a large enough font size, and prefer printed text. Recognition results may deviate. For example, the ideographic comma (、) may be recognized as another symbol, or the variable x as the multiplication sign ×. Mathematical symbols are especially prone to this.

  • Adjust the chunking granularity. Questions and explanations are mostly short items. In Processing Strategy on the knowledge base details page, click Create Strategy to adjust the chunking granularity. The unit of the maximum chunk length is characters (default 512). Set it to 580 to 770 characters (about 384 to 512 tokens) so that the question stem and explanation stay in the same chunk as much as possible.

Step 3: Configure retrieval parameters and the prompt

Adjust retrieval and scenario in config.json for the education scenario.

{
  "retrieval": {
    "page_size": 8,
    "candidate_count": 64,
    "min_score": 0.2,
    "semantic_weight": 0.75,
    "enable_query_expansion": true,
    "rerank_model_name": "qwen3-rerank",
    "tag_filter": {
      "relation": "and",
      "conditions": [
        {"field": "subject", "op": "=", "value": "Mathematics"}
      ]
    }
  },
  "scenario": {
    "title": "Education question bank retrieval and explanation",
    "system_prompt": "You are a teaching assistant. Explain only based on the retrieved textbooks, questions, and ground truth. Give the approach first, then the steps, and finally the answer, citing [Source N]. Write formulas in plain text; do not use LaTeX syntax. If a retrieved question does not match the user's description, explicitly point out the difference instead of applying the answer from the question bank directly. Do not guess the ground truth when the material is insufficient.",
    "image_enabled": true,
    "sample_questions": [
      "How do I find the maximum value from the vertex form of a quadratic function?",
      "Find an example problem that uses the Pythagorean theorem and explain it.",
      "Which knowledge points does this question test?"
    ]
  }
}

Parameter description:

  • semantic_weight=0.75 with qwen3-rerank enabled: students' questions are mostly natural-language descriptions of problems, and a higher semantic weight favors matching similar questions.

    Note that the reranking score and the vector score are not on the same scale. The final score is computed as score ≈ (1-semantic_weight) × keywordScore + semantic_weight × semanticScore (a rank feature may also be added). When reranking is disabled, semanticScore is the vector similarity. When reranking is enabled, it becomes the reranking model score, and min_score always filters on this final score. In testing, enabling reranking reduced the number of results for the same question from 5 to 4. This is the combined effect of the threshold and the new scale; it does not mean reranking makes recall worse.

    There is no recommended combination that generalizes across corpora. First set min_score to 0 to retrieve a batch of results and manually label their relevance, then choose a threshold based on the distribution of recall and false recall. Recalibrate every time you switch the reranking model, toggle reranking, or adjust semantic_weight.

  • Fixing subject in tag_filter effectively prevents cross-subject false recall. In testing, asking about a physics knowledge point under the subject=Mathematics filter returned 0 results, and the model answered "insufficient material" following the prompt.

  • For op, only the three operators =, in, and not in actually take effect. Aliases such as eq, ==, equal, or like return 400 Unsupported tag filter operator and cause every question to fail. Although ≠, >, <, ≥, ≤, empty, not empty, start with, and end with appear in the "Supported operators" list of that error message, in testing the conditions are silently ignored and all unfiltered data is returned.

    Note that contains and not contains are merely aliases for in / not in. They do not perform string containment, and passing a string fragment as the value returns 0 results.

    To filter by a range, enumerate the values with in. To check whether a tag is empty, use = "". After configuration, always compare the total number of results with and without the filter to confirm that the filter takes effect.

  • In system_prompt, explicitly state "formulas in plain text". Education prompts easily make the model output LaTeX, but the sample page displays with <pre> plain text, so formulas do not render and appear as raw dollar signs and backslashes. If you want to keep LaTeX, integrate KaTeX or MathJax into the page.

Step 4: Upload, publish, and verify

  1. Upload the materials. MetaFields applies to the whole batch; the script groups materials by tag first and submits them in batches.

    python upload.py --manifest documents.jsonl
  2. On the Data Management page of the console, confirm that the material status is Processed. For image materials, click View Chunks and confirm that the recognized text matches the content in the image before you publish a version.

  3. On the Version Management page, click Publish Version. After completing the three-step wizard, confirm that the status of the new version is Published.

    Important

    At most three published versions can exist at the same time. When the limit is reached, the Publish Version button is grayed out, but the page still shows "There are currently N pending changes to publish". You must first delete old versions that are no longer needed in the version records. Version deletion cannot be undone. Question banks usually get materials added each semester or each unit, so keep only "the current version + the most recent historical version".

  4. Start the service. The Q&A page is served at http://127.0.0.1:7860.

    python app.py
  5. If your question bank contains image materials, try an image query. The /api/ask endpoint of app.py accepts an optional image_url parameter and passes it through to the image field of SearchKnowledgeBase. Images help most with questions with unclear references. In testing, for the same question "Which knowledge points does this question test?": without an image, the hit was a Pythagorean theorem exercise (relevance 0.431, not the intended question); with an image of a quadratic function question attached, the hit was the quadratic function question in the question bank (0.583). The image provided the key semantics for retrieval.

    curl -sS http://127.0.0.1:7860/api/ask \
      -H 'Content-Type: application/json' \
      -d '{"question":"Which knowledge points does this question test?","image_url":"https://<publicly accessible image URL>"}'
  6. Verify the result against the following four checks:

    • The page at http://127.0.0.1:7860 opens normally. The education scenario additionally shows an input box for the question image URL. Submitting a question returns both the answer and the retrieval sources.

    • Each [Source N] in the answer has a corresponding material chunk in the sources section below.

    • When asking about content that the question bank does not cover, the answer explicitly says the material is insufficient instead of making up an answer.

    • After adding materials and republishing a version, the page can retrieve the new content.

  7. Add one more verification for semantic recall: ask with information that appears only in the material content, not in the file name (for example, a question number), and confirm that the corresponding material is hit. This check also applies to verifying whether an image was correctly recognized and imported.

Image query capabilities and boundaries

Images participate only in retrieval and are never sent to the LLM. The system behavior is "find the most similar question in the question bank based on the image, and then explain based on the retrieved material", not "recognize and solve the question in the image". If you upload a new question that does not exist in the question bank, the system answers with the most similar question in the bank, and the answer may differ from the question in the image without being noticed. For example, when uploading y=-(x-1)²+4 (maximum value is 4), if y=-2(x-3)²+5 exists in the question bank, the answer may give a maximum value of 5. Therefore:

  • Show a hint on the page that "the image is used to find similar questions in the question bank".

  • Require the model in the prompt to point out differences between retrieval results and the user's description.

    This capability should therefore be called "image-assisted retrieval" or "image-based question search", and must not be presented externally as "photo-to-answer solving". To solve a new question in an image, the application layer must pass the original image separately to an LLM that supports multi-modal input. Instruct the LLM to cross-check the retrieved material. Do not rely on knowledge base retrieval alone.

image_url must be a publicly accessible address. The server side validates it:

Input

Server response

Intranet or local address (such as 127.0.0.1)

400 URL resolves to a non-public or blocked address

Link to a non-image resource

400 image_query URL must point to an image.

Domain name that cannot be resolved

400 Could not resolve hostname

Left empty

Falls back to plain text retrieval (normal behavior)

Image recognition has the following boundaries, which you should know before preparing question bank materials:

  • Format: use only JPG, JPEG, PNG, and GIF. The underlying file recognition may also accept formats such as WebP and TIFF, but there is no unified public commitment; do not rely on them.

  • Size: the API has no separate hard limit on image bytes or pixels, but is actually constrained jointly by the upload gateway, image decoding, memory, and the image-to-text conversion model service. Oversized images may still fail.

  • Recognition accuracy: there is no committed accuracy metric for handwriting, mathematical formulas, geometric figures, or coordinate systems. An image with no text may get a description generated by image-to-text conversion, but retrieval hits are not guaranteed. In testing for this topic, x was recognized as ×, and the ideographic comma was recognized as a single-dot stroke character. Mathematical symbols require manual spot checks in particular.

  • No recognition quality threshold: as long as the image can be decoded and the pipeline reports no error, the document status is Processed, even if the recognized text is very sparse or wrong. Only when decoding fails or a required model call errors out does Processing Failed appear. Therefore you must spot-check the chunk text on the Data Management page after upload; do not rely on the status alone.

Usage notes for education scenarios

  • Upload textbooks, questions, and answers as a set, and distinguish them with questionType so you can filter as needed (for example, let students retrieve only "practice question", and let teachers retrieve "ground truth").

  • Create an independent entry per subject: after fixing subject in tag_filter, that entry can only answer questions for that subject. Keep the sample questions in the same subject; otherwise, students asking about other subjects only get "insufficient material".

  • Restrict access to answer materials separately: if students should not get the ground truth directly, filter with a condition such as questionType not in ["ground truth"] at the student entry.

  • Retrieval results contain an images field. By design, the field returns the images associated with chunks, and the server generates short-lived signed URLs for persisted images. This field is an implemented capability, not a reserved empty field. However, the field currently returns no image URL. When an image document is hit but the field is empty, the images of these materials were not persisted to the corresponding chunks during the parsing and chunking phase, or the signature association was not established — a pipeline issue to investigate per specific document. No request parameter can enable it, so do not treat "always empty" as product design. You do not need to maintain a "filename → image URL" mapping table yourself long-term: in the short term the page can fall back to displaying the recognized text; when the original question image needs to be shown, handle it temporarily by associating via documentId with the import manifest.

  • Keep manual review for key conclusions, and remind students that the official textbooks and teacher explanations prevail.

FAQ

Symptom

Cause and solution

A question returns 500, and the log shows 400 Unsupported tag filter operator

op used an alias such as eq, ==, or like. Use =, in, or not in.

With tag filtering added, the number of results is exactly the same as without it

A non-working operator (such as >, ≥, ≠, or empty) was used. Only =, in, and not in take effect.

Image materials cannot be found after upload

Check in order: whether the knowledge base data type supports images; whether the data status is Processed; whether the text recognized in View Chunks is empty or inconsistent with the image (images with no text, font size too small, and handwriting can all cause recognition failure).

You ask with a question image, but the answer is about a different question

Expected behavior. The image is used only to retrieve similar questions; the model explains the retrieved question.

A question returns 400 URL resolves to a non-public or blocked address

image_url points to an intranet or local address. Change it to a publicly accessible image URL.

Formulas appear as $y = a(x-h)^2 + k$ on the page

The model output LaTeX but the page renders it as plain text. Require plain-text formulas in the prompt, or integrate a formula rendering library into the page.

Fewer results after enabling reranking

Reranking scores and vector scores have different scales; combined with min_score, more results are filtered out. Lower min_score or disable reranking to compare the effect.

Tag filtering always returns 0 results

The tag name spelling differs from what was written (the API does not report an error and just returns 0 results). Verify on the knowledge base details page → Tags → Manage.

Upload returns 400 No OSS document can be registered.

All files in the batch were deduplicated because of identical names. Rename the files or delete the old data in the console first.

Retrieval returns 404 Knowledge base version ... does not exist

No version has been published, or the specified version number does not exist.