AI_RERANK 在向量或混合檢索後對召回結果進行相關性重排序(精排),提升 RAG 問答及多模態搜尋精度。系統根據候選類型自動匹配模型:文本使用 qwen3-rerank,圖片/視頻使用 qwen3-vl-rerank。
命令格式
REST 介面
POST /v2/vectordb/ai/rerank
Content-Type: application/json
{
"model_name": "qwen3-rerank",
"query": "<查詢文本>",
"documents": ["<候選內容>"],
"params": {"timeout_sec": 10}
}
Python
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=1024)
schema.add_function(
Function(
name="embed_content",
function_type=FunctionType.TEXTEMBEDDING,
input_field_names=["content"],
output_field_names=["embedding"],
params={
"provider": "aliyun_milvus",
"model_name": "text-embedding-v4",
"dim": 1024,
},
)
)
reranker = Function(
name="rerank_content",
function_type=FunctionType.RERANK,
input_field_names=["content"],
params={
"reranker": "model",
"provider": "aliyun_milvus",
"model_name": "qwen3-rerank",
"queries": ["<查詢文本>"],
"timeout_sec": 10,
},
)
Collection Search 情境使用 FunctionType.RERANK,需設定 reranker="model"、provider="aliyun_milvus"、model_name、queries 和候選內容欄位。文本候選使用 qwen3-rerank;圖片或視頻候選使用 qwen3-vl-rerank。
多模態輸入組合
qwen3-vl-rerank 不僅支援以文字查詢圖片或視頻,也支援將圖片 URL 直接作為 Query,實現以圖搜圖或以圖搜視頻。支援的輸入組合如下。
|
Query |
候選內容 |
模型 |
典型情境 |
|
文本 |
文本 |
|
RAG 文檔重排 |
|
文本 |
圖片 URL |
|
以文搜圖 |
|
圖片 URL |
圖片 URL |
|
以圖搜圖、相似商品推薦 |
|
文本 |
視頻 URL |
|
以文搜視頻 |
|
圖片 URL |
視頻 URL |
|
以圖搜視頻、素材匹配 |
REST 介面中,文本 Query 和媒體 URL 均通過 query 字串傳入;Collection Search 中通過 queries 數組傳入。媒體候選需以可訪問的 URL 字串形式儲存到 documents 或 Collection 的候選欄位中。
參數說明
|
參數 |
說明 |
|
|
REST 必填;模型重排 Function 也必填。文本候選使用 |
|
|
REST 必填。可以是非空查詢文本;多模態重排時也可以是可訪問的圖片 URL。 |
|
|
REST 必填。至少一條候選內容;返回結果的 |
|
|
僅模型 Rerank Function 必填,固定為 |
|
|
僅 Function 使用。調用模型時設為 |
|
|
僅 Search Function 必填。Query 數組,元素可以是查詢文本或圖片 URL;單 Query 搜尋傳一個元素。 |
|
|
可選。單次發送給模型的最大候選數,預設 |
|
|
可選。並發數,範圍 |
|
|
可選。單次模型調用逾時秒數,範圍 |
|
|
僅 Search Function 可選。模型異常時可選 |
|
|
可選。圖片或視頻作為 Query 或候選內容參與重排時設為 |
|
|
可選。排序指令,用於明確模型的排序關注點,多模態重排時建議使用英文指令。例如:圖片重排可強調主體一致性、外觀、構圖和細粒度視覺特徵;視頻重排可強調主體、動作、情境和外觀。 |
傳回值說明
REST 調用成功時,data.output.results 返回每個候選的 index 和 relevance_score。結果與輸入下標對應,業務側應按 relevance_score 降序排序;分數沒有固定閾值,不應依賴樣本中的具體數值。響應還包含 usage(如 total_tokens)與 request_id 欄位。
{"code":0,"data":{"output":{"results":[{"index":0,"relevance_score":0.97}]}}}
樣本一:RAG 問答的候選文檔重排(文本)
知識庫已召回三條候選文檔,使用者詢問“向量資料庫的典型應用情境”。使用 RERANK 重新打分後,應用取分數最高的文檔交給大模型產生回答。關鍵參數為 query、documents 和 timeout_sec。
REST 介面
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST "$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
MODEL_NAME="qwen3-rerank"
BODY=$(cat <<JSON
{
"model_name": "$MODEL_NAME",
"query": "Typical use cases of vector databases",
"documents": [
"Milvus is an open-source vector database for RAG, recommendation, and image search.",
"MySQL is a relational database for transactional applications.",
"Vector retrieval can be combined with inverted indexes to improve recall."
],
"params": {"max_concurrency": 2, "timeout_sec": 10}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/rerank" "$BODY")"
if command -v jq >/dev/null 2>&1; then
echo "$RESPONSE_BODY" | jq .
[ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
echo "$RESPONSE_BODY"
fi
Python
from __future__ import annotations
import os
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"
def add_id(schema: Any) -> None:
schema.add_field("id", DataType.INT64, is_primary=True)
MODEL_NAME = "qwen3-rerank"
EMBEDDING_MODEL = os.getenv("AIFUNC_RERANK_EMBEDDING_MODEL", "text-embedding-v4")
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
collection_name = "simple_ai_rerank_text"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=1024)
schema.add_function(Function(name="embed_content", function_type=FunctionType.TEXTEMBEDDING, input_field_names=["content"], output_field_names=["embedding"], params={"provider": "aliyun_milvus", "model_name": EMBEDDING_MODEL, "dim": 1024}))
index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, [{"content": "Milvus is an open-source vector database for RAG and recommendation."}, {"content": "MySQL is a relational database for transactional applications."}])
client.flush(collection_name)
client.load_collection(collection_name)
QUERY = "Typical use cases of vector databases"
reranker = Function(name="rerank_content", function_type=FunctionType.RERANK, input_field_names=["content"], params={"reranker": "model", "provider": "aliyun_milvus", "model_name": MODEL_NAME, "queries": [QUERY], "max_concurrency": 2, "timeout_sec": 10})
for hit in client.search(collection_name=collection_name, data=[QUERY], anns_field="embedding", limit=2, output_fields=["content"], ranker=reranker)[0]:
print(f"score={hit['distance']:.4f} content={hit['entity']['content']}")
第 0 和第 2 條文檔通常比 MySQL 文檔更相關;Python 會按實際 relevance_score 從高到低列印候選,應用可取前兩條作為 RAG 上下文。
樣本二:商品圖片多模態重排(以文搜圖與以圖搜圖)
向量檢索先召回候選商品圖片,再使用 qwen3-vl-rerank 做精排。本樣本對同一批候選圖片連續執行兩種 Query:一是文本 Query,二是直接傳入參考圖片 URL 作為 Query,實現以圖搜圖。建議先從向量檢索結果中截取最多 40 張候選圖片,再交給模型重排;instruct 使用英文指令,明確要求優先比較主體身份、外觀、構圖和細粒度視覺特徵。
REST 介面
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST "$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
MODEL_NAME="qwen3-vl-rerank"
QUERY_TEXT="a fashion product matching the query"
QUERY_IMAGE_URL="https://img.alicdn.com/imgextra/i3/O1CN01rdstgY1uiZWt8gqSL_!!6000000006071-0-tps-1970-356.jpg"
CANDIDATE_IMAGE_URL_1="$QUERY_IMAGE_URL"
CANDIDATE_IMAGE_URL_2="https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp"
CANDIDATE_IMAGE_URL_3="https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"
INSTRUCT="Rank the candidate images by relevance to the query, prioritizing subject identity, appearance, composition, and fine-grained visual details."
run_rerank() {
local label="$1"
local query="$2"
local body
body=$(cat <<JSON
{
"model_name": "$MODEL_NAME",
"query": "$query",
"documents": [
"$CANDIDATE_IMAGE_URL_1",
"$CANDIDATE_IMAGE_URL_2",
"$CANDIDATE_IMAGE_URL_3"
],
"params": {
"is_multimodal": true,
"instruct": "$INSTRUCT",
"max_client_batch_size": 40,
"timeout_sec": 10
}
}
JSON
)
echo "=== $label ==="
post_json "/v2/vectordb/ai/rerank" "$body"
}
run_rerank "text query" "$QUERY_TEXT"
run_rerank "image query" "$QUERY_IMAGE_URL"
Python
from __future__ import annotations
import os
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"
def add_id(schema: Any) -> None:
schema.add_field("id", DataType.INT64, is_primary=True)
MODEL_NAME = "qwen3-vl-rerank"
EMBEDDING_MODEL = os.getenv("AIFUNC_RERANK_EMBEDDING_MODEL", "qwen3-vl-embedding")
VECTOR_DIM = int(os.getenv("AIFUNC_RERANK_EMBEDDING_DIM", "2560"))
QUERY_TEXT = "a fashion product matching the query"
QUERY_IMAGE_URL = "https://img.alicdn.com/imgextra/i3/O1CN01rdstgY1uiZWt8gqSL_!!6000000006071-0-tps-1970-356.jpg"
CANDIDATE_IMAGE_URLS = [
QUERY_IMAGE_URL,
"https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp",
"https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg",
]
INSTRUCT = (
"Rank the candidate images by relevance to the query, prioritizing "
"subject identity, appearance, composition, and fine-grained visual details."
)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
collection_name = "simple_ai_rerank_image"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=VECTOR_DIM)
schema.add_function(Function(name="embed_image", function_type=FunctionType.TEXTEMBEDDING, input_field_names=["content"], output_field_names=["embedding"], params={"provider": "aliyun_milvus", "model_name": EMBEDDING_MODEL, "dim": VECTOR_DIM, "is_multimodal": "true"}))
index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, [{"content": url} for url in CANDIDATE_IMAGE_URLS[:40]])
client.flush(collection_name)
client.load_collection(collection_name)
for query_label, query in (("text", QUERY_TEXT), ("image", QUERY_IMAGE_URL)):
reranker = Function(name=f"rerank_image_{query_label}_query", function_type=FunctionType.RERANK, input_field_names=["content"], params={"reranker": "model", "provider": "aliyun_milvus", "model_name": MODEL_NAME, "queries": [query], "is_multimodal": "true", "instruct": INSTRUCT, "max_client_batch_size": 40, "timeout_sec": 10})
for hit in client.search(collection_name=collection_name, data=[query], anns_field="embedding", limit=len(CANDIDATE_IMAGE_URLS[:40]), output_fields=["content"], ranker=reranker)[0]:
print(f"{query_label} score={hit['distance']:.4f} content={hit['entity']['content']}")
預期結果如下:
-
文本 Query 按文字描述與候選圖片之間的語義、視覺相關性排序。
-
圖片 Query 執行以圖搜圖,與 Query 相同的圖片通常排在最前。
-
REST 返回的
index對應原始documents數組的下標,在映射回候選內容前不要改變原數組順序。 -
Collection Search 返回的
hit['distance']為重排後的相關性分數,結果已按分數排序。
樣本三:視頻素材多模態重排(以文搜視頻與以圖搜視頻)
視頻候選同樣使用 qwen3-vl-rerank。除文本 Query 外,還可以傳入參考圖片 URL 作為 Query,讓模型從候選視頻中找出主體、服飾、動作或情境最匹配的素材。
REST 介面
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST "$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
MODEL_NAME="qwen3-vl-rerank"
QUERY_TEXT="person wearing a striped sweater"
QUERY_IMAGE_URL="https://img.alicdn.com/imgextra/i3/O1CN01rdstgY1uiZWt8gqSL_!!6000000006071-0-tps-1970-356.jpg"
VIDEO_URL="https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"
INSTRUCT="Rank the candidate videos by relevance to the query, prioritizing matching subjects, actions, scenes, appearance, and fine-grained visual details."
run_rerank() {
local label="$1"
local query="$2"
local body
body=$(cat <<JSON
{
"model_name": "$MODEL_NAME",
"query": "$query",
"documents": ["$VIDEO_URL"],
"params": {
"is_multimodal": true,
"instruct": "$INSTRUCT",
"timeout_sec": 10
}
}
JSON
)
echo "=== $label ==="
post_json "/v2/vectordb/ai/rerank" "$body"
}
run_rerank "text query" "$QUERY_TEXT"
run_rerank "image query" "$QUERY_IMAGE_URL"
Python
from __future__ import annotations
import os
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"
def add_id(schema: Any) -> None:
schema.add_field("id", DataType.INT64, is_primary=True)
MODEL_NAME = "qwen3-vl-rerank"
EMBEDDING_MODEL = os.getenv("AIFUNC_RERANK_EMBEDDING_MODEL", "qwen3-vl-embedding")
VECTOR_DIM = int(os.getenv("AIFUNC_RERANK_EMBEDDING_DIM", "2560"))
QUERY_TEXT = "person wearing a striped sweater"
QUERY_IMAGE_URL = "https://img.alicdn.com/imgextra/i3/O1CN01rdstgY1uiZWt8gqSL_!!6000000006071-0-tps-1970-356.jpg"
VIDEO_URL = "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"
INSTRUCT = (
"Rank the candidate videos by relevance to the query, prioritizing matching "
"subjects, actions, scenes, appearance, and fine-grained visual details."
)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
collection_name = "simple_ai_rerank_video"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=VECTOR_DIM)
schema.add_function(Function(name="embed_video", function_type=FunctionType.TEXTEMBEDDING, input_field_names=["content"], output_field_names=["embedding"], params={"provider": "aliyun_milvus", "model_name": EMBEDDING_MODEL, "dim": VECTOR_DIM, "is_multimodal": "true"}))
index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, [{"content": VIDEO_URL}])
client.flush(collection_name)
client.load_collection(collection_name)
for query_label, query in (("text", QUERY_TEXT), ("image", QUERY_IMAGE_URL)):
reranker = Function(name=f"rerank_video_{query_label}_query", function_type=FunctionType.RERANK, input_field_names=["content"], params={"reranker": "model", "provider": "aliyun_milvus", "model_name": MODEL_NAME, "queries": [query], "is_multimodal": "true", "instruct": INSTRUCT, "timeout_sec": 10})
for hit in client.search(collection_name=collection_name, data=[query], anns_field="embedding", limit=1, output_fields=["content"], ranker=reranker)[0]:
print(f"{query_label} score={hit['distance']:.4f} content={hit['entity']['content']}")
預期結果如下:
-
文本 Query 根據文本描述對視頻候選重排。
-
圖片 Query 根據參考圖中的主體、外觀和情境對視頻候選重排。
-
生產環境建議先對向量召回結果做截斷再重排,並保持候選 URL 與業務主鍵之間的穩定映射。