全部產品
Search
文件中心

Vector Retrieval Service for Milvus:具名實體識別

更新時間:Aug 04, 2026

AI_ENTITY_EXTRACT 函數可從文本、圖片或視頻中識別人名、組織、地點、時間和金額等具名實體,適用於新聞索引、媒體資產標註和事件檢索等情境。

命令格式

REST 介面

POST /v2/vectordb/ai/entity_extract
Content-Type: application/json

{
  "model_name": "<模型名>",
  "texts": ["<文本或媒體URL>"],
  "params": {"entity_types": ["PERSON", "ORGANIZATION"]}
}

Python

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("entities", DataType.JSON)
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2)
schema.add_function(
    Function(
        name="extract_entities",
        function_type=texttransform_function_type(),
        input_field_names=["content"],
        output_field_names=["entities"],
        params={
            "provider": "aliyun_milvus",
            "model_name": "<模型名>",
            "task": "ai_entity_extract",
            "entity_types": "PERSON,ORGANIZATION,LOCATION,DATE,PRODUCT",
            "temperature": "0",
        },
    )
)

參數說明

參數

說明

model_name

必填。模型名稱;圖片和視頻須選擇已配置多模態模型(如 qwen3.7-plus)。

texts

REST 必填。待識別文本,或可被模型訪問的圖片/視頻 URL。

entity_types

選填。目標實體類型,支援數組或逗號分隔字串,數量為 1~30。省略時預設識別人、組織、地點、產品、事件、日期、時間、金額、百分比、URL、郵箱、電話和 IP。

prompt

選填。補充識別規則,最大 5000 字元,不支援 ${...}。媒體情境建議明確要求“不根據外觀猜測身份、品牌或地點”。

media_type

選填。取值為 image 或 video,表示輸入為媒體 URL。

temperature / max_concurrency / timeout_sec

選填。模型穩定性、並發和逾時設定。

provider / task

僅 Collection Function 必填,固定為 aliyun_milvus 與 ai_entity_extract。

傳回值說明

每項輸出是 {"entities":[{"text":"實體文本","type":"實體類型"}]} 形式的 JSON 對象或可解析為該對象的 JSON 字串。entities 的每項應含非空 text 和請求中的 type;未識別到合格具名實體時,{"entities":[ ]} 是合法且更真實的結果。Schema 會將合法 JSON 寫入輸出欄位。

樣本一:為新聞建立實體索引(文本)

內容平台從新聞中抽取明確出現的人物、機構、地點、日期和產品,作為檢索過濾條件。樣本驗證 JSON 結構和類型,不要求模型返回完全相同的實體列表。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{
  "model_name": "qwen3.7-max",
  "texts": [
    "In March 2024, Zhang San joined Alibaba in Hangzhou to work on Milvus."
  ],
  "params": {
    "entity_types": ["PERSON", "ORGANIZATION", "LOCATION", "DATE", "PRODUCT"],
    "temperature": 0
  }
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/entity_extract" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # 阿里雲 Milvus 將 TEXTTRANSFORM 作為託管擴充(函數類型值 9)提供;
    # 部分 pymilvus 版本尚未內建該枚舉成員,而 Function(...) 通過 FunctionType(...) 校正。
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_entity_extract_text",
    input_fields=[("content", DataType.VARCHAR, 4096)],
    output_field=("entities", DataType.JSON, None),
    function_name="extract_entities",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-max", "task": "ai_entity_extract", "entity_types": "PERSON,ORGANIZATION,LOCATION,DATE,PRODUCT", "temperature": "0"},
    rows=[{"content": "In March 2024, Zhang San joined Alibaba in Hangzhou to work on Milvus.", "dummy_vector": [0.1, 0.2]}],
)

預期結果:entities 中每項含非空 text 且 type 屬於請求集合。實測返回 {"entities":[{"text":"March 2024","type":"DATE"},{"text":"Zhang San","type":"PERSON"},{"text":"Alibaba","type":"ORGANIZATION"},{"text":"Hangzhou","type":"LOCATION"},{"text":"Milvus","type":"PRODUCT"}]},實際邊界以模型和實體規則為準。

樣本二:審核服飾宣傳圖中的具名實體(圖片)

素材庫收到一張人物穿黑白條紋上衣、手持黑色包的圖片。這些只是可見對象,不等於可確認的人名、品牌、商品名或地點;因此樣本明確禁止猜測,並接受空 entities 作為真實結果。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{"model_name":"qwen3.7-plus","texts":["https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp"],"params":{"media_type":"image","entity_types":["PERSON","PRODUCT","LOCATION"],"temperature":0}}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/entity_extract" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # 阿里雲 Milvus 將 TEXTTRANSFORM 作為託管擴充(函數類型值 9)提供;
    # 部分 pymilvus 版本尚未內建該枚舉成員,而 Function(...) 通過 FunctionType(...) 校正。
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_entity_extract_image",
    input_fields=[("image_url", DataType.VARCHAR, 4096)],
    output_field=("entities", DataType.JSON, None),
    function_name="extract_image_entities",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_entity_extract", "media_type": "image", "entity_types": "PERSON,PRODUCT,LOCATION", "temperature": "0"},
    rows=[{"image_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp", "dummy_vector": [0.1, 0.2]}],
)

預期結果:輸出是合法實體物件,且每項類型屬於請求集合。如圖片沒有足夠證據確認具名實體,{"entities":[ ]} 是預期範圍內的真實結果,不應為了“有資料”而猜測。

樣本三:審核擬人馬角色視頻中的具名實體(視頻)

視頻素材近景呈現一個穿西裝的擬人馬角色。角色外觀不能證明真實人物身份、作品名、品牌或地點;因此該素材也採用“有明確證據才抽取”的規則,空實體數組屬於正常結果。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{"model_name":"qwen3.7-plus","texts":["https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"],"params":{"media_type":"video","entity_types":["PERSON","PRODUCT","LOCATION"],"temperature":0}}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/entity_extract" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # 阿里雲 Milvus 將 TEXTTRANSFORM 作為託管擴充(函數類型值 9)提供;
    # 部分 pymilvus 版本尚未內建該枚舉成員,而 Function(...) 通過 FunctionType(...) 校正。
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_entity_extract_video",
    input_fields=[("video_url", DataType.VARCHAR, 4096)],
    output_field=("entities", DataType.JSON, None),
    function_name="extract_video_entities",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_entity_extract", "media_type": "video", "entity_types": "PERSON,PRODUCT,LOCATION", "temperature": "0"},
    rows=[{"video_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4", "dummy_vector": [0.1, 0.2]}],
)

預期結果:輸出結構正確,實體類型均屬於請求集合,或返回空 entities。實測視頻返回 {"entities":[]}(無可確認具名實體,合法)。如需角色類型、服裝或運動描述,應使用內容摘要/分類能力,不應將普通視覺屬性冒充為具名實體。