全部產品
Search
文件中心

Vector Retrieval Service for Milvus:重排序

更新時間:Aug 04, 2026

AI_RERANK 在向量或混合檢索後對召回結果進行相關性重排序(精排),提升 RAG 問答及多模態搜尋精度。系統根據候選類型自動匹配模型:文本使用 qwen3-rerank,圖片/視頻使用 qwen3-vl-rerank。

命令格式

REST 介面

POST /v2/vectordb/ai/rerank
Content-Type: application/json

{
  "model_name": "qwen3-rerank",
  "query": "<查詢文本>",
  "documents": ["<候選內容>"],
  "params": {"timeout_sec": 10}
}

Python

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=1024)
schema.add_function(
    Function(
        name="embed_content",
        function_type=FunctionType.TEXTEMBEDDING,
        input_field_names=["content"],
        output_field_names=["embedding"],
        params={
            "provider": "aliyun_milvus",
            "model_name": "text-embedding-v4",
            "dim": 1024,
        },
    )
)

reranker = Function(
    name="rerank_content",
    function_type=FunctionType.RERANK,
    input_field_names=["content"],
    params={
        "reranker": "model",
        "provider": "aliyun_milvus",
        "model_name": "qwen3-rerank",
        "queries": ["<查詢文本>"],
        "timeout_sec": 10,
    },
)

Collection Search 情境使用 FunctionType.RERANK,需設定 reranker="model"、provider="aliyun_milvus"、model_name、queries 和候選內容欄位。文本候選使用 qwen3-rerank;圖片或視頻候選使用 qwen3-vl-rerank。

多模態輸入組合

qwen3-vl-rerank 不僅支援以文字查詢圖片或視頻,也支援將圖片 URL 直接作為 Query,實現以圖搜圖或以圖搜視頻。支援的輸入組合如下。

Query

候選內容

模型

典型情境

文本

文本

qwen3-rerank

RAG 文檔重排

文本

圖片 URL

qwen3-vl-rerank

以文搜圖

圖片 URL

圖片 URL

qwen3-vl-rerank

以圖搜圖、相似商品推薦

文本

視頻 URL

qwen3-vl-rerank

以文搜視頻

圖片 URL

視頻 URL

qwen3-vl-rerank

以圖搜視頻、素材匹配

說明

REST 介面中,文本 Query 和媒體 URL 均通過 query 字串傳入;Collection Search 中通過 queries 數組傳入。媒體候選需以可訪問的 URL 字串形式儲存到 documents 或 Collection 的候選欄位中。

參數說明

參數

說明

model_name

REST 必填;模型重排 Function 也必填。文本候選使用 qwen3-rerank,圖片或視頻等多模態候選使用 qwen3-vl-rerank。

query

REST 必填。可以是非空查詢文本;多模態重排時也可以是可訪問的圖片 URL。

documents

REST 必填。至少一條候選內容;返回結果的 index 對應該數組下標。

provider

僅模型 Rerank Function 必填,固定為 aliyun_milvus。

reranker

僅 Function 使用。調用模型時設為 model;weighted、rrf、decay、boost 用於內建排序階段。

queries

僅 Search Function 必填。Query 數組,元素可以是查詢文本或圖片 URL;單 Query 搜尋傳一個元素。

max_client_batch_size

可選。單次發送給模型的最大候選數,預設 128。

max_concurrency

可選。並發數,範圍 1~64,預設 1。

timeout_sec

可選。單次模型調用逾時秒數,範圍 1~300,預設 30。

on_error

僅 Search Function 可選。模型異常時可選 fail、fallback_previous、fallback_original 或 skip_stage。

is_multimodal

可選。圖片或視頻作為 Query 或候選內容參與重排時設為 true。REST 介面使用 JSON 布爾值 true,Function 參數使用字串 "true"。

instruct

可選。排序指令,用於明確模型的排序關注點,多模態重排時建議使用英文指令。例如:圖片重排可強調主體一致性、外觀、構圖和細粒度視覺特徵;視頻重排可強調主體、動作、情境和外觀。

傳回值說明

REST 調用成功時,data.output.results 返回每個候選的 index 和 relevance_score。結果與輸入下標對應,業務側應按 relevance_score 降序排序;分數沒有固定閾值,不應依賴樣本中的具體數值。響應還包含 usage(如 total_tokens)與 request_id 欄位。

{"code":0,"data":{"output":{"results":[{"index":0,"relevance_score":0.97}]}}}

樣本一:RAG 問答的候選文檔重排(文本)

知識庫已召回三條候選文檔,使用者詢問“向量資料庫的典型應用情境”。使用 RERANK 重新打分後,應用取分數最高的文檔交給大模型產生回答。關鍵參數為 query、documents 和 timeout_sec。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

MODEL_NAME="qwen3-rerank"

BODY=$(cat <<JSON
{
  "model_name": "$MODEL_NAME",
  "query": "Typical use cases of vector databases",
  "documents": [
    "Milvus is an open-source vector database for RAG, recommendation, and image search.",
    "MySQL is a relational database for transactional applications.",
    "Vector retrieval can be combined with inverted indexes to improve recall."
  ],
  "params": {"max_concurrency": 2, "timeout_sec": 10}
}
JSON
)

RESPONSE_BODY="$(post_json "/v2/vectordb/ai/rerank" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

from __future__ import annotations
import os
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

MODEL_NAME = "qwen3-rerank"
EMBEDDING_MODEL = os.getenv("AIFUNC_RERANK_EMBEDDING_MODEL", "text-embedding-v4")
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

collection_name = "simple_ai_rerank_text"
if client.has_collection(collection_name):
    client.drop_collection(collection_name)

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=1024)
schema.add_function(Function(name="embed_content", function_type=FunctionType.TEXTEMBEDDING, input_field_names=["content"], output_field_names=["embedding"], params={"provider": "aliyun_milvus", "model_name": EMBEDDING_MODEL, "dim": 1024}))
index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, [{"content": "Milvus is an open-source vector database for RAG and recommendation."}, {"content": "MySQL is a relational database for transactional applications."}])
client.flush(collection_name)
client.load_collection(collection_name)

QUERY = "Typical use cases of vector databases"
reranker = Function(name="rerank_content", function_type=FunctionType.RERANK, input_field_names=["content"], params={"reranker": "model", "provider": "aliyun_milvus", "model_name": MODEL_NAME, "queries": [QUERY], "max_concurrency": 2, "timeout_sec": 10})
for hit in client.search(collection_name=collection_name, data=[QUERY], anns_field="embedding", limit=2, output_fields=["content"], ranker=reranker)[0]:
    print(f"score={hit['distance']:.4f} content={hit['entity']['content']}")

第 0 和第 2 條文檔通常比 MySQL 文檔更相關;Python 會按實際 relevance_score 從高到低列印候選,應用可取前兩條作為 RAG 上下文。

樣本二:商品圖片多模態重排(以文搜圖與以圖搜圖)

向量檢索先召回候選商品圖片,再使用 qwen3-vl-rerank 做精排。本樣本對同一批候選圖片連續執行兩種 Query:一是文本 Query,二是直接傳入參考圖片 URL 作為 Query,實現以圖搜圖。建議先從向量檢索結果中截取最多 40 張候選圖片,再交給模型重排;instruct 使用英文指令,明確要求優先比較主體身份、外觀、構圖和細粒度視覺特徵。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

MODEL_NAME="qwen3-vl-rerank"
QUERY_TEXT="a fashion product matching the query"
QUERY_IMAGE_URL="https://img.alicdn.com/imgextra/i3/O1CN01rdstgY1uiZWt8gqSL_!!6000000006071-0-tps-1970-356.jpg"
CANDIDATE_IMAGE_URL_1="$QUERY_IMAGE_URL"
CANDIDATE_IMAGE_URL_2="https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp"
CANDIDATE_IMAGE_URL_3="https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"
INSTRUCT="Rank the candidate images by relevance to the query, prioritizing subject identity, appearance, composition, and fine-grained visual details."

run_rerank() {
  local label="$1"
  local query="$2"
  local body

  body=$(cat <<JSON
{
  "model_name": "$MODEL_NAME",
  "query": "$query",
  "documents": [
    "$CANDIDATE_IMAGE_URL_1",
    "$CANDIDATE_IMAGE_URL_2",
    "$CANDIDATE_IMAGE_URL_3"
  ],
  "params": {
    "is_multimodal": true,
    "instruct": "$INSTRUCT",
    "max_client_batch_size": 40,
    "timeout_sec": 10
  }
}
JSON
  )

  echo "=== $label ==="
  post_json "/v2/vectordb/ai/rerank" "$body"
}

run_rerank "text query" "$QUERY_TEXT"
run_rerank "image query" "$QUERY_IMAGE_URL"

Python

from __future__ import annotations
import os
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

MODEL_NAME = "qwen3-vl-rerank"
EMBEDDING_MODEL = os.getenv("AIFUNC_RERANK_EMBEDDING_MODEL", "qwen3-vl-embedding")
VECTOR_DIM = int(os.getenv("AIFUNC_RERANK_EMBEDDING_DIM", "2560"))
QUERY_TEXT = "a fashion product matching the query"
QUERY_IMAGE_URL = "https://img.alicdn.com/imgextra/i3/O1CN01rdstgY1uiZWt8gqSL_!!6000000006071-0-tps-1970-356.jpg"
CANDIDATE_IMAGE_URLS = [
    QUERY_IMAGE_URL,
    "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp",
    "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg",
]
INSTRUCT = (
    "Rank the candidate images by relevance to the query, prioritizing "
    "subject identity, appearance, composition, and fine-grained visual details."
)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

collection_name = "simple_ai_rerank_image"
if client.has_collection(collection_name):
    client.drop_collection(collection_name)

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=VECTOR_DIM)
schema.add_function(Function(name="embed_image", function_type=FunctionType.TEXTEMBEDDING, input_field_names=["content"], output_field_names=["embedding"], params={"provider": "aliyun_milvus", "model_name": EMBEDDING_MODEL, "dim": VECTOR_DIM, "is_multimodal": "true"}))
index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, [{"content": url} for url in CANDIDATE_IMAGE_URLS[:40]])
client.flush(collection_name)
client.load_collection(collection_name)

for query_label, query in (("text", QUERY_TEXT), ("image", QUERY_IMAGE_URL)):
    reranker = Function(name=f"rerank_image_{query_label}_query", function_type=FunctionType.RERANK, input_field_names=["content"], params={"reranker": "model", "provider": "aliyun_milvus", "model_name": MODEL_NAME, "queries": [query], "is_multimodal": "true", "instruct": INSTRUCT, "max_client_batch_size": 40, "timeout_sec": 10})
    for hit in client.search(collection_name=collection_name, data=[query], anns_field="embedding", limit=len(CANDIDATE_IMAGE_URLS[:40]), output_fields=["content"], ranker=reranker)[0]:
        print(f"{query_label} score={hit['distance']:.4f} content={hit['entity']['content']}")

預期結果如下:

  • 文本 Query 按文字描述與候選圖片之間的語義、視覺相關性排序。

  • 圖片 Query 執行以圖搜圖,與 Query 相同的圖片通常排在最前。

  • REST 返回的 index 對應原始 documents 數組的下標,在映射回候選內容前不要改變原數組順序。

  • Collection Search 返回的 hit['distance'] 為重排後的相關性分數,結果已按分數排序。

樣本三:視頻素材多模態重排(以文搜視頻與以圖搜視頻)

視頻候選同樣使用 qwen3-vl-rerank。除文本 Query 外,還可以傳入參考圖片 URL 作為 Query,讓模型從候選視頻中找出主體、服飾、動作或情境最匹配的素材。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

MODEL_NAME="qwen3-vl-rerank"
QUERY_TEXT="person wearing a striped sweater"
QUERY_IMAGE_URL="https://img.alicdn.com/imgextra/i3/O1CN01rdstgY1uiZWt8gqSL_!!6000000006071-0-tps-1970-356.jpg"
VIDEO_URL="https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"
INSTRUCT="Rank the candidate videos by relevance to the query, prioritizing matching subjects, actions, scenes, appearance, and fine-grained visual details."

run_rerank() {
  local label="$1"
  local query="$2"
  local body

  body=$(cat <<JSON
{
  "model_name": "$MODEL_NAME",
  "query": "$query",
  "documents": ["$VIDEO_URL"],
  "params": {
    "is_multimodal": true,
    "instruct": "$INSTRUCT",
    "timeout_sec": 10
  }
}
JSON
  )

  echo "=== $label ==="
  post_json "/v2/vectordb/ai/rerank" "$body"
}

run_rerank "text query" "$QUERY_TEXT"
run_rerank "image query" "$QUERY_IMAGE_URL"

Python

from __future__ import annotations
import os
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

MODEL_NAME = "qwen3-vl-rerank"
EMBEDDING_MODEL = os.getenv("AIFUNC_RERANK_EMBEDDING_MODEL", "qwen3-vl-embedding")
VECTOR_DIM = int(os.getenv("AIFUNC_RERANK_EMBEDDING_DIM", "2560"))
QUERY_TEXT = "person wearing a striped sweater"
QUERY_IMAGE_URL = "https://img.alicdn.com/imgextra/i3/O1CN01rdstgY1uiZWt8gqSL_!!6000000006071-0-tps-1970-356.jpg"
VIDEO_URL = "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"
INSTRUCT = (
    "Rank the candidate videos by relevance to the query, prioritizing matching "
    "subjects, actions, scenes, appearance, and fine-grained visual details."
)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

collection_name = "simple_ai_rerank_video"
if client.has_collection(collection_name):
    client.drop_collection(collection_name)

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=VECTOR_DIM)
schema.add_function(Function(name="embed_video", function_type=FunctionType.TEXTEMBEDDING, input_field_names=["content"], output_field_names=["embedding"], params={"provider": "aliyun_milvus", "model_name": EMBEDDING_MODEL, "dim": VECTOR_DIM, "is_multimodal": "true"}))
index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, [{"content": VIDEO_URL}])
client.flush(collection_name)
client.load_collection(collection_name)

for query_label, query in (("text", QUERY_TEXT), ("image", QUERY_IMAGE_URL)):
    reranker = Function(name=f"rerank_video_{query_label}_query", function_type=FunctionType.RERANK, input_field_names=["content"], params={"reranker": "model", "provider": "aliyun_milvus", "model_name": MODEL_NAME, "queries": [query], "is_multimodal": "true", "instruct": INSTRUCT, "timeout_sec": 10})
    for hit in client.search(collection_name=collection_name, data=[query], anns_field="embedding", limit=1, output_fields=["content"], ranker=reranker)[0]:
        print(f"{query_label} score={hit['distance']:.4f} content={hit['entity']['content']}")

預期結果如下:

  • 文本 Query 根據文本描述對視頻候選重排。

  • 圖片 Query 根據參考圖中的主體、外觀和情境對視頻候選重排。

  • 生產環境建議先對向量召回結果做截斷再重排,並保持候選 URL 與業務主鍵之間的穩定映射。