The AI_MULTI_MODAL_GENERATE function converts text, images, videos, or audio into text output. Use it to describe product materials, summarize videos, answer visual questions, or transcribe speech.
Command format
Multimodal generation supports two usage patterns:
-
REST API — Call the endpoint directly to generate text from your input. Use this approach for ad-hoc or batch generation tasks.
-
Collection Function — Define a
TEXTTRANSFORMFunction in your collection schema to generate text automatically when data is inserted. Use this approach to integrate generation into your data ingestion pipeline.
The REST API examples require curl and jq. The Python examples require pymilvus. Replace the endpoint URL and credentials in each example with your Milvus instance values.
REST API
REST API
POST /v2/vectordb/ai/multi_modal_generate
{
"model_name": "<model_name>",
"texts": ["<text_or_media_url>"],
"params": {"media_type": "text | image | video | audio", "prompt": "<instruction>"}
}
Python (Collection Function)
Python (Collection Function)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("generated", DataType.VARCHAR, max_length=1024)
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2)
schema.add_function(
Function(
name="generate_text",
function_type=texttransform_function_type(),
input_field_names=["content"],
output_field_names=["generated"],
params={
"provider": "aliyun_milvus",
"model_name": "<model_name>",
"task": "ai_multi_modal_generate",
"media_type": "text",
"prompt": "Answer concisely: ${content}",
"temperature": "0",
},
)
)
Parameters
| Parameter | Description |
model_name |
Required. For text, image, and video, use qwen3.7-plus (recommended), qwen3.6-plus, qwen3.6-flash, qwen3.5-flash, or kimi-k2.6. For audio, use qwen3-asr-flash. |
texts |
Required for REST API. For the text media type, pass the original text. For other media types, pass a model-accessible URL. Images can also be passed as Base64. Results match the input order. |
media_type |
Required. Valid values: text, image, video, audio. |
prompt |
Required for image and video. Optional for text. Not supported for audio. In a Collection Function schema, use ${field_name} to reference field values. |
temperature |
Optional. Controls output generation behavior. Passed as a number in REST API calls or as a string in Collection Function params. Set to 0 in all examples. |
language / enable_itn |
Optional. Audio only. language specifies the recognition language (for example, zh, en). enable_itn controls whether to enable inverse text normalization. |
timeout_sec / max_concurrency |
Optional. Control the per-call timeout and batch concurrency. |
output_mapping |
Optional. Collection Function text, image, and video scenarios only. Maps simple paths in the model output JSON to multiple VARCHAR, TEXT, or JSON fields. Without this setting, only one text output field is allowed. |
provider / task |
Required for Collection Function only. Fixed values: aliyun_milvus and ai_multi_modal_generate. |
Return values
On success, the response contains the following fields:
-
data.output.outputs— A text array. Each element corresponds one-to-one with the inputtextsarray. -
data.usage— Usage metrics that may include:-
input_tokens,output_tokens,total_tokens— Token counts. -
image_tokens,video_tokens,audio_tokens— Media-specific token counts, when applicable. -
seconds— Media duration, when applicable.
-
Example 1: Generate a concise text response (text)
This example generates a concise text response from a text input using the text media type.
REST API
REST API
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
BODY=$(cat <<JSON
{
"model_name": "qwen3.7-plus",
"texts": ["Explain in one sentence how Milvus is used in a RAG application."],
"params": {"media_type": "text", "prompt": "Answer concisely.", "temperature": 0}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/multi_modal_generate" "$BODY")"
if command -v jq >/dev/null 2>&1; then
echo "$RESPONSE_BODY" | jq .
[ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
echo "$RESPONSE_BODY"
fi
Python
Python
from __future__ import annotations
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"
DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9
def texttransform_function_type() -> Any:
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
function_type = getattr(FunctionType, type_name, None)
if function_type is not None:
return function_type
# Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
# some pymilvus versions do not yet have this enum member built-in, and Function(...) validates via FunctionType(...).
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
def add_id(schema: Any) -> None:
schema.add_field("id", DataType.INT64, is_primary=True)
def add_dummy_vector(schema: Any) -> None:
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)
def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
for name, data_type, max_length in input_fields:
field_params = {"max_length": max_length} if max_length is not None else {}
schema.add_field(name, data_type, **field_params)
output_name, output_data_type, output_max_length = output_field
output_params = {"max_length": output_max_length} if output_max_length is not None else {}
schema.add_field(output_name, output_data_type, **output_params)
add_dummy_vector(schema)
schema.add_function(
Function(
name=function_name,
function_type=texttransform_function_type(),
input_field_names=[name for name, _, _ in input_fields],
output_field_names=[output_name],
params=function_params,
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, rows)
client.flush(collection_name)
fields = [name for name, _, _ in input_fields] + [output_name]
for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
print(row)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
run_texttransform_example(
client=client,
collection_name="simple_ai_multi_modal_generate_text",
input_fields=[("content", DataType.VARCHAR, 4096)],
output_field=("generated", DataType.VARCHAR, 1024),
function_name="generate_text",
function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_multi_modal_generate", "media_type": "text", "prompt": "Answer concisely: ${content}", "temperature": "0"},
rows=[{"content": "Explain how Milvus is used in a RAG application.", "dummy_vector": [0.1, 0.2]}],
)
Expected result: The generated field contains a concise response describing how Milvus is used in a RAG application. In the REST API response, this value is at data.output.outputs[0].
Example 2: Generate retrieval descriptions from product images (image)
This example generates an objective description for image-text retrieval before indexing clothing materials. The bash example requires jq.
REST API
REST API
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
BODY=$(cat <<JSON
{
"model_name": "qwen3.7-plus",
"texts": ["https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp"],
"params": {"media_type": "image", "prompt": "Describe the image in one concise sentence.", "temperature": 0}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/multi_modal_generate" "$BODY")"
if command -v jq >/dev/null 2>&1; then
echo "$RESPONSE_BODY" | jq .
[ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
echo "$RESPONSE_BODY"
fi
Python
Python
from __future__ import annotations
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"
DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9
def texttransform_function_type() -> Any:
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
function_type = getattr(FunctionType, type_name, None)
if function_type is not None:
return function_type
# Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
# some pymilvus versions do not yet have this enum member built-in, and Function(...) validates via FunctionType(...).
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
def add_id(schema: Any) -> None:
schema.add_field("id", DataType.INT64, is_primary=True)
def add_dummy_vector(schema: Any) -> None:
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)
def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
for name, data_type, max_length in input_fields:
field_params = {"max_length": max_length} if max_length is not None else {}
schema.add_field(name, data_type, **field_params)
output_name, output_data_type, output_max_length = output_field
output_params = {"max_length": output_max_length} if output_max_length is not None else {}
schema.add_field(output_name, output_data_type, **output_params)
add_dummy_vector(schema)
schema.add_function(
Function(
name=function_name,
function_type=texttransform_function_type(),
input_field_names=[name for name, _, _ in input_fields],
output_field_names=[output_name],
params=function_params,
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, rows)
client.flush(collection_name)
fields = [name for name, _, _ in input_fields] + [output_name]
for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
print(row)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
run_texttransform_example(
client=client,
collection_name="simple_ai_multi_modal_generate_image",
input_fields=[("image_url", DataType.VARCHAR, 4096)],
output_field=("description", DataType.VARCHAR, 2048),
function_name="describe_image",
function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_multi_modal_generate", "media_type": "image", "prompt": "Describe the image in one concise sentence.", "temperature": "0"},
rows=[{"image_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp", "dummy_vector": [0.1, 0.2]}],
)
Expected result: The description field contains a concise objective description of the image, which can be written to a product retrieval field.
Example 3: Generate retrieval summaries from concept videos (video)
This example generates retrieval summaries from concept videos to enable searching by character, action, and visual atmosphere. Only validates that the text is non-empty; no fixed wording is assumed.
REST API
REST API
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
BODY=$(cat <<JSON
{
"model_name": "qwen3.7-plus",
"texts": ["https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"],
"params": {"media_type": "video", "prompt": "Describe the video in one concise sentence.", "temperature": 0}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/multi_modal_generate" "$BODY")"
if command -v jq >/dev/null 2>&1; then
echo "$RESPONSE_BODY" | jq .
[ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
echo "$RESPONSE_BODY"
fi
Python
Python
from __future__ import annotations
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"
DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9
def texttransform_function_type() -> Any:
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
function_type = getattr(FunctionType, type_name, None)
if function_type is not None:
return function_type
# Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
# some pymilvus versions do not yet have this enum member built-in, and Function(...) validates via FunctionType(...).
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
def add_id(schema: Any) -> None:
schema.add_field("id", DataType.INT64, is_primary=True)
def add_dummy_vector(schema: Any) -> None:
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)
def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
for name, data_type, max_length in input_fields:
field_params = {"max_length": max_length} if max_length is not None else {}
schema.add_field(name, data_type, **field_params)
output_name, output_data_type, output_max_length = output_field
output_params = {"max_length": output_max_length} if output_max_length is not None else {}
schema.add_field(output_name, output_data_type, **output_params)
add_dummy_vector(schema)
schema.add_function(
Function(
name=function_name,
function_type=texttransform_function_type(),
input_field_names=[name for name, _, _ in input_fields],
output_field_names=[output_name],
params=function_params,
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, rows)
client.flush(collection_name)
fields = [name for name, _, _ in input_fields] + [output_name]
for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
print(row)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
run_texttransform_example(
client=client,
collection_name="simple_ai_multi_modal_generate_video",
input_fields=[("video_url", DataType.VARCHAR, 4096)],
output_field=("description", DataType.VARCHAR, 2048),
function_name="describe_video",
function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_multi_modal_generate", "media_type": "video", "prompt": "Describe the video in one concise sentence.", "temperature": "0"},
rows=[{"video_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4", "dummy_vector": [0.1, 0.2]}],
)
Expected result: The description field contains a one-sentence summary of the video content, which can be used as a creative material retrieval field.
Example 4: Transcribe and archive a hotline greeting (audio)
This example transcribes a hotline greeting into searchable text for archiving. Audio requests do not require a prompt. The bash example requires jq. The sample audio is a public greeting; the transcription language depends on the actual audio content.
REST API
REST API
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
BODY=$(cat <<JSON
{
"model_name": "qwen3-asr-flash",
"texts": ["https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3"],
"params": {"media_type": "audio", "language": "en", "enable_itn": true}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/multi_modal_generate" "$BODY")"
if command -v jq >/dev/null 2>&1; then
echo "$RESPONSE_BODY" | jq .
[ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
echo "$RESPONSE_BODY"
fi
Python
Python
from __future__ import annotations
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"
DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9
def texttransform_function_type() -> Any:
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
function_type = getattr(FunctionType, type_name, None)
if function_type is not None:
return function_type
# Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
# some pymilvus versions do not yet have this enum member built-in, and Function(...) validates via FunctionType(...).
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
def add_id(schema: Any) -> None:
schema.add_field("id", DataType.INT64, is_primary=True)
def add_dummy_vector(schema: Any) -> None:
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)
def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
for name, data_type, max_length in input_fields:
field_params = {"max_length": max_length} if max_length is not None else {}
schema.add_field(name, data_type, **field_params)
output_name, output_data_type, output_max_length = output_field
output_params = {"max_length": output_max_length} if output_max_length is not None else {}
schema.add_field(output_name, output_data_type, **output_params)
add_dummy_vector(schema)
schema.add_function(
Function(
name=function_name,
function_type=texttransform_function_type(),
input_field_names=[name for name, _, _ in input_fields],
output_field_names=[output_name],
params=function_params,
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, rows)
client.flush(collection_name)
fields = [name for name, _, _ in input_fields] + [output_name]
for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
print(row)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
run_texttransform_example(
client=client,
collection_name="simple_ai_multi_modal_generate_audio",
input_fields=[("audio_url", DataType.VARCHAR, 4096)],
output_field=("transcript", DataType.VARCHAR, 4096),
function_name="transcribe_audio",
function_params={"provider": "aliyun_milvus", "model_name": "qwen3-asr-flash", "task": "ai_multi_modal_generate", "media_type": "audio", "language": "en", "enable_itn": "true"},
rows=[{"audio_url": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "dummy_vector": [0.1, 0.2]}],
)
Expected result: The transcript field contains non-empty transcription text. The sample audio transcribes to Welcome to Alibaba Cloud., rendered in English; the API returns the transcription in the source language of the audio. The transcription can be archived, but do not use the example code to automatically draw compliance conclusions.
Media security and retention
Before using your own images, videos, or audio, confirm that you have obtained the necessary permissions for the content, portrait rights, and voice rights. Use least-privilege URLs with expiration, and delete personal information in transcriptions and derived text in accordance with your business retention policy.