すべてのプロダクト
Search
ドキュメントセンター

Vector Retrieval Service for Milvus:マルチモーダル生成

最終更新日:Aug 04, 2026

`AI_MULTI_MODAL_GENERATE` 関数は、テキスト、画像、動画、または音声をテキスト出力に変換します。この関数を使用して、製品資料の説明、動画の要約、視覚的な質問への回答、音声の文字起こしなどを行うことができます。

コマンドフォーマット

マルチモーダル生成は、2 つの使用パターンをサポートしています:

  • REST API — エンドポイントを直接呼び出して、入力からテキストを生成します。このアプローチは、アドホックまたはバッチ生成タスクに使用します。

  • コレクション関数 — コレクションスキーマで TEXTTRANSFORM 関数を定義し、データが挿入されたときに自動的にテキストを生成します。このアプローチは、データインジェストパイプラインに生成機能を統合するために使用します。

REST API の例では curl と jq が必要です。Python の例では pymilvus が必要です。各例のエンドポイント URL と認証情報を、ご利用の Milvus インスタンスの値に置き換えてください。

REST API

REST API

POST /v2/vectordb/ai/multi_modal_generate
{
  "model_name": "<model_name>",
  "texts": ["<text_or_media_url>"],
  "params": {"media_type": "text | image | video | audio", "prompt": "<instruction>"}
}

Python (コレクション関数)

Python (コレクション関数)

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("generated", DataType.VARCHAR, max_length=1024)
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2)
schema.add_function(
    Function(
        name="generate_text",
        function_type=texttransform_function_type(),
        input_field_names=["content"],
        output_field_names=["generated"],
        params={
            "provider": "aliyun_milvus",
            "model_name": "<model_name>",
            "task": "ai_multi_modal_generate",
            "media_type": "text",
            "prompt": "Answer concisely: ${content}",
            "temperature": "0",
        },
    )
)

パラメーター

パラメーター説明
model_name必須。テキスト、画像、動画には、qwen3.7-plus (推奨)、qwen3.6-plus、qwen3.6-flash、qwen3.5-flash、または kimi-k2.6 を使用します。音声には qwen3-asr-flash を使用します。
textsREST API では必須です。text メディアタイプの場合は、元のテキストを渡します。他のメディアタイプの場合は、モデルがアクセス可能な URL を渡します。画像は Base64 としても渡すことができます。結果は入力順に一致します。
media_type必須。有効な値: text、image、video、audio。
promptimage と video では必須です。text ではオプションです。audio ではサポートされていません。コレクション関数のスキーマでは、${field_name} を使用してフィールド値を参照します。
temperatureオプション。出力生成の動作をコントロールします。REST API 呼び出しでは数値として、コレクション関数のパラメーターでは文字列として渡されます。すべての例で 0 に設定されています。
language / enable_itnオプション。音声のみ。language は認識言語 (例:zh、en) を指定します。enable_itn は、逆テキスト正規化を有効にするかどうかをコントロールします。
timeout_sec / max_concurrencyオプション。呼び出しごとのタイムアウトとバッチの同時実行数をコントロールします。
output_mappingオプション。コレクション関数の text、image、video シナリオのみ。モデル出力 JSON 内の単純なパスを、複数の VARCHAR、TEXT、または JSON フィールドにマッピングします。この設定がない場合、許可されるテキスト出力フィールドは 1 つだけです。
provider / taskコレクション関数でのみ必須です。固定値: aliyun_milvus および ai_multi_modal_generate。

戻り値

成功した場合、応答には以下のフィールドが含まれます:

  • data.output.outputs — テキスト配列。各要素は、入力の texts 配列と 1 対 1 で対応します。

  • data.usage — 以下を含む可能性のある使用量メトリクス:

    • input_tokens、output_tokens、total_tokens — トークン数。

    • image_tokens、video_tokens、audio_tokens — 該当する場合のメディア固有のトークン数。

    • seconds — 該当する場合のメディアの持続時間。

例 1:簡潔なテキスト応答を生成する (テキスト)

この例では、text メディアタイプを使用して、テキスト入力から簡潔なテキスト応答を生成します。

REST API

REST API

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{
  "model_name": "qwen3.7-plus",
  "texts": ["Explain in one sentence how Milvus is used in a RAG application."],
  "params": {"media_type": "text", "prompt": "Answer concisely.", "temperature": 0}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/multi_modal_generate" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
    # some pymilvus versions do not yet have this enum member built-in, and Function(...) validates via FunctionType(...).
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_multi_modal_generate_text",
    input_fields=[("content", DataType.VARCHAR, 4096)],
    output_field=("generated", DataType.VARCHAR, 1024),
    function_name="generate_text",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_multi_modal_generate", "media_type": "text", "prompt": "Answer concisely: ${content}", "temperature": "0"},
    rows=[{"content": "Explain how Milvus is used in a RAG application.", "dummy_vector": [0.1, 0.2]}],
)

期待される結果: generated フィールドには、RAG アプリケーションで Milvus がどのように使用されるかを説明する簡潔な応答が含まれます。REST API の応答では、この値は data.output.outputs[0] にあります。

例 2:製品画像から検索用の説明を生成する (画像)

この例では、衣料品素材をインデックス化する前に、画像テキスト検索のための客観的な説明を生成します。bash の例では jq が必要です。

REST API

REST API

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<ユーザー名>:<パスワード>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{
  "model_name": "qwen3.7-plus",
  "texts": ["https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp"],
  "params": {"media_type": "image", "prompt": "画像を簡潔な一文で説明してください。", "temperature": 0}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/multi_modal_generate" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # Alibaba Cloud Milvus は、TEXTTRANSFORM をマネージド拡張機能 (関数タイプの値は 9) として提供します。
    # 一部の pymilvus バージョンでは、この enum メンバーがまだ組み込まれておらず、Function(...) は FunctionType(...) を介して検証を行います。
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_multi_modal_generate_image",
    input_fields=[("image_url", DataType.VARCHAR, 4096)],
    output_field=("description", DataType.VARCHAR, 2048),
    function_name="describe_image",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_multi_modal_generate", "media_type": "image", "prompt": "Describe the image in one concise sentence.", "temperature": "0"},
    rows=[{"image_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260415/hynnff/wan-video-edit-clothes.webp", "dummy_vector": [0.1, 0.2]}],
)

期待される結果: description フィールドには、画像の簡潔で客観的な説明が含まれており、これを製品検索フィールドに書き込むことができます。

例 3:コンセプト動画から検索用の要約を生成する (動画)

この例では、コンセプト動画から検索用の要約を生成し、キャラクター、アクション、視覚的な雰囲気による検索を可能にします。テキストが空でないことのみを検証し、特定の文言は想定していません。

REST API

REST API

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{
  "model_name": "qwen3.7-plus",
  "texts": ["https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/20260409/dozxak/Wan_Video_Edit_33_1.mp4"],
  "params": {"media_type": "video", "prompt": "Describe the video in one concise sentence.", "temperature": 0}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/multi_modal_generate" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
    # some pymilvus versions do not yet have this enum member built-in, and Function(...) validates via FunctionType(...).
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_multi_modal_generate_video",
    input_fields=[("video_url", DataType.VARCHAR, 4096)],
    output_field=("description", DataType.VARCHAR, 2048),
    function_name="describe_video",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3.7-plus", "task": "ai_multi_modal_generate", "media_type": "video", "prompt": "Describe the video in one concise sentence.", "temperature": "0"},
    rows=[{"video_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/20260409/dozxak/Wan_Video_Edit_33_1.mp4", "dummy_vector": [0.1, 0.2]}],
)

期待される結果: description フィールドには、動画コンテンツの一文要約が含まれており、これをクリエイティブ素材の検索フィールドとして使用できます。

例 4:ホットラインの挨拶を文字起こししてアーカイブする (音声)

この例では、ホットラインの挨拶をアーカイブ用に検索可能なテキストに文字起こしします。音声リクエストでは prompt は不要です。bash の例では jq が必要です。サンプル音声は公開されている挨拶です。文字起こしの言語は、実際の音声コンテンツに依存します。

REST API

REST API

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

BODY=$(cat <<JSON
{
  "model_name": "qwen3-asr-flash",
  "texts": ["https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3"],
  "params": {"media_type": "audio", "language": "en", "enable_itn": true}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/multi_modal_generate" "$BODY")"
if command -v jq >/dev/null 2>&1; then
  echo "$RESPONSE_BODY" | jq .
  [ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
  echo "$RESPONSE_BODY"
fi

Python

Python

from __future__ import annotations

from typing import Any

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"

DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    # Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
    # some pymilvus versions do not yet have this enum member built-in, and Function(...) validates via FunctionType(...).
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

def add_id(schema: Any) -> None:
    schema.add_field("id", DataType.INT64, is_primary=True)

def add_dummy_vector(schema: Any) -> None:
    schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)

def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
    if client.has_collection(collection_name):
        client.drop_collection(collection_name)
    schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
    add_id(schema)
    for name, data_type, max_length in input_fields:
        field_params = {"max_length": max_length} if max_length is not None else {}
        schema.add_field(name, data_type, **field_params)
    output_name, output_data_type, output_max_length = output_field
    output_params = {"max_length": output_max_length} if output_max_length is not None else {}
    schema.add_field(output_name, output_data_type, **output_params)
    add_dummy_vector(schema)
    schema.add_function(
        Function(
            name=function_name,
            function_type=texttransform_function_type(),
            input_field_names=[name for name, _, _ in input_fields],
            output_field_names=[output_name],
            params=function_params,
        )
    )
    index_params = client.prepare_index_params()
    index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
    client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
    client.insert(collection_name, rows)
    client.flush(collection_name)
    fields = [name for name, _, _ in input_fields] + [output_name]
    for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
        print(row)

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

run_texttransform_example(
    client=client,
    collection_name="simple_ai_multi_modal_generate_audio",
    input_fields=[("audio_url", DataType.VARCHAR, 4096)],
    output_field=("transcript", DataType.VARCHAR, 4096),
    function_name="transcribe_audio",
    function_params={"provider": "aliyun_milvus", "model_name": "qwen3-asr-flash", "task": "ai_multi_modal_generate", "media_type": "audio", "language": "en", "enable_itn": "true"},
    rows=[{"audio_url": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "dummy_vector": [0.1, 0.2]}],
)

期待される結果: transcript フィールドには、空でない文字起こしテキストが含まれます。サンプル音声は 欢迎使用阿里云。 に文字起こしされます。文字起こし結果はアーカイブできますが、このサンプルコードを使用して自動的にコンプライアンスの結論を導き出さないでください。

メディアのセキュリティと保持

ご自身の画像、動画、または音声を使用する前に、コンテンツ、肖像権、および音声権に関して必要な権限を取得していることを確認してください。有効期限付きの最小権限の URL を使用し、ビジネス上の保持ポリシーに従って、文字起こし結果や派生テキストに含まれる個人情報を削除してください。