All Products
Search
Document Center

Vector Retrieval Service for Milvus:Named Entity Recognition

Last Updated:Aug 03, 2026

The AI_ENTITY_EXTRACT function extracts named entities such as persons, organizations, locations, dates, and amounts from text, images, or videos. Use this function for news indexing, media asset annotation, and event retrieval.

Command format

You can call AI_ENTITY_EXTRACT through the REST API or the Python Collection Function.

REST API

POST /v2/vectordb/ai/entity_extract
Content-Type: application/json

{
  "model_name": "<model name>",
  "texts": ["<text or media URL>"],
  "params": {"entity_types": ["PERSON", "ORGANIZATION"]}
}

Python

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("entities", DataType.JSON)
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2)
schema.add_function(
    Function(
        name="extract_entities",
        function_type=texttransform_function_type(),
        input_field_names=["content"],
        output_field_names=["entities"],
        params={
            "provider": "aliyun_milvus",
            "model_name": "<model name>",
            "task": "ai_entity_extract",
            "entity_types": "PERSON,ORGANIZATION,LOCATION,DATE,PRODUCT",
            "temperature": "0",
        },
    )
)

Parameters

ParameterDescription
model_nameRequired. The model name. For text extraction, use a text model (for example, qwen3.7-max). For images and videos, use a configured multimodal model (for example, qwen3.7-plus).
textsRequired for REST API. Text to be recognized, or image and video URLs accessible to the model.
entity_typesOptional. Target entity types. Accepts an array or a comma-separated string, with 1 to 30 entries. When omitted, the function recognizes persons, organizations, locations, products, events, dates, times, amounts, percentages, URLs, emails, phone numbers, and IPs by default.
promptOptional. Additional recognition rules. Maximum 5,000 characters. ${...} is not supported. (Recommended) For media scenarios, include instructions such as "do not guess identity, brand, or location based on appearance."
media_typeOptional. Set to image or video to indicate that the input is a media URL.
temperature/max_concurrency/timeout_secOptional. Model stability, concurrency, and timeout settings.
provider/taskRequired for Collection Function only. Fixed as aliyun_milvus and ai_entity_extract.

Return value

Each output is a JSON object in the form {"entities":[{"text":"entity text","type":"entity type"}]}, or a JSON string that parses to the same structure. Each item in entities contains a non-empty text and a type from the request. When no qualifying named entities are recognized, {"entities":[]} is a valid result. The schema writes valid JSON to the output field.

Example 1: Building an entity index for news (text)

A content platform extracts clearly mentioned persons, organizations, locations, dates, and products from news articles to use as search filter conditions. This example verifies the JSON structure and types. It does not require the model to return the exact same entity list.

REST API

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{
  "model_name": "qwen3.7-max",
  "texts": [
    "In March 2024, Zhang San joined Alibaba in Hangzhou to work on Milvus."
  ],
  "params": {
    "entity_types": ["PERSON", "ORGANIZATION", "LOCATION", "DATE", "PRODUCT"],
    "temperature": 0
  }
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/entity_extract" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
    # some pymilvus versions do not have this enum member built-in, while Function(...) validates via FunctionType(...).
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_entity_extract_text",
    input_fields=[("content", DataType.VARCHAR, 4096)],
    output_field=("entities", DataType.JSON, None),
    function_name="extract_entities",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-max", "task": "ai_entity_extract", "entity_types": "PERSON,ORGANIZATION,LOCATION,DATE,PRODUCT", "temperature": "0"},
    rows=[{"content": "In March 2024, Zhang San joined Alibaba in Hangzhou to work on Milvus.", "dummy_vector": [0.1, 0.2]}],
)

Verification

Each item in entities contains a non-empty text field and a type value from the requested set. The following is a sample result from actual testing:

{"entities":[{"text":"March 2024","type":"DATE"},{"text":"Zhang San","type":"PERSON"},{"text":"Alibaba","type":"ORGANIZATION"},{"text":"Hangzhou","type":"LOCATION"},{"text":"Milvus","type":"PRODUCT"}]}

Actual entity boundaries depend on the model and entity rules.

Example 2: Reviewing named entities in a fashion promotional image (image)

A media library receives an image of a person wearing a black-and-white striped top and holding a black bag. These visible objects do not constitute confirmable person names, brands, product names, or locations. This example explicitly prohibits guessing and accepts empty entities as a valid result.

REST API

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{"model_name":"qwen3.7-plus","texts":["https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp"],"params":{"media_type":"image","entity_types":["PERSON","PRODUCT","LOCATION"],"temperature":0}}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/entity_extract" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
    # some pymilvus versions do not have this enum member built-in, while Function(...) validates via FunctionType(...).
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_entity_extract_image",
    input_fields=[("image_url", DataType.VARCHAR, 4096)],
    output_field=("entities", DataType.JSON, None),
    function_name="extract_image_entities",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_entity_extract", "media_type": "image", "entity_types": "PERSON,PRODUCT,LOCATION", "temperature": "0"},
    rows=[{"image_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp", "dummy_vector": [0.1, 0.2]}],
)

Verification

The output is a valid entity object where each item's type belongs to the requested set. If the image does not contain sufficient evidence to confirm named entities, {"entities":[]} is a valid result.

Example 3: Reviewing named entities in a humanoid horse character video (video)

A video asset shows a close-up of a humanoid horse character in a suit. The character's appearance does not prove a real person's identity, a work title, a brand, or a location. This example applies the "extract only with clear evidence" rule, and an empty entity array is a valid result.

REST API

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{"model_name":"qwen3.7-plus","texts":["https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"],"params":{"media_type":"video","entity_types":["PERSON","PRODUCT","LOCATION"],"temperature":0}}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/entity_extract" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
    # some pymilvus versions do not have this enum member built-in, while Function(...) validates via FunctionType(...).
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_entity_extract_video",
    input_fields=[("video_url", DataType.VARCHAR, 4096)],
    output_field=("entities", DataType.JSON, None),
    function_name="extract_video_entities",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_entity_extract", "media_type": "video", "entity_types": "PERSON,PRODUCT,LOCATION", "temperature": "0"},
    rows=[{"video_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4", "dummy_vector": [0.1, 0.2]}],
)

Verification

The output structure is correct and all entity types belong to the requested set. An empty entities array is valid when no confirmable named entities are present. The actual test returned {"entities":[]}.

To describe character types, clothing, or movement, use content summarization or classification capabilities instead of entity extraction.