AI_AUDIO_TRANSCRIBE 関数は、音声の文字起こしを行いテキストに変換します。この関数は、カスタマーサービスにおける通話録音のインデックス作成、会議議事録の生成、音声メッセージの処理、ポッドキャストコンテンツのアーカイブなどのシナリオで利用します。
独自の録音を文字起こしする前に、録音を通知する義務を果たし、必要な権限付与を得ていることを確認してください。最小権限の原則に基づき、有効期間の短いメディア URL を使用し、カスタマーサービスのデータ保持ポリシーに従って、元の音声と文字起こしされたテキストの両方の保存期間を制限することを推奨します。
コマンド形式
REST API
REST API
{
"model_name": "qwen3-asr-flash",
"texts": ["<audio_url>"],
"params": {
"language": "zh",
"enable_itn": true,
"max_concurrency": 1
}
}Python
Python
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("audio_url", DataType.VARCHAR, max_length=4096)
schema.add_field("transcript", DataType.VARCHAR, max_length=4096)
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2)
schema.add_function(
Function(
name="transcribe_audio",
function_type=texttransform_function_type(),
input_field_names=["audio_url"],
output_field_names=["transcript"],
params={
"provider": "aliyun_milvus",
"model_name": "qwen3-asr-flash",
"task": "ai_audio_transcribe",
"language": "zh",
"enable_itn": True,
},
)
)パラメーター
| パラメーター | 説明 |
model_name | 必須。 qwen3-asr-flash を使用します。 |
texts | REST API の場合に必須。 オーディオ URL の配列です。 空の文字列は含められません。 結果は入力順に対応します。 |
language | 任意。 ソース言語の言語コードです。 例:zh、en。 |
enable_itn | 任意。 ブール値。 有効にすると、話し言葉の数字や同様の表現が、書き言葉の形式に正規化されます。 |
timeout_sec / max_concurrency | 任意。 複数の音声ファイルに対し、呼び出しのタイムアウトと並行処理数をそれぞれ制御します。 |
media_type | 省略、または audio に設定可能です。 それ以外の値を指定するとエラーになります。 |
| その他 | prompt、temperature、およびその他のカスタムモデルパラメーターはサポートされていません。 音声 URL へのアクセス可否や、サポートされる形式と長さは、モデルサービスの制限に従います。 |
戻り値
data.output.outputs は、各音声入力に対応する文字起こしテキストを順番に返します。 usage.audio_tokens はオーディオトークンの使用量、usage.seconds は処理された音声の長さ (秒) を示します。
{
"code": 0,
"data": {
"output": {"outputs": ["<transcript>"]},
"usage": {"audio_tokens": 256, "total_tokens": 256, "seconds": 45}
}
}例:カスタマーサービス録音の文字起こし
次の例では、ホットラインのウェルカムメッセージをテキストに変換します。これは、バージョンのアーカイブや、営業時間、録音通知、その他の内容が明確に表現されているかを手動でレビューするために使用されます。文字起こしに含まれないフィールドは "unconfirmed" としてマークする必要があります。モデルでこれらのフィールドを補完することはできません。
デフォルトで公開されている welcome.mp3 は中国語のウェルカムメッセージのため、この例では language=zh を使用します。 AIFUNC_AUDIO_URL を置き換える場合は、AIFUNC_AUDIO_LANGUAGE も実際のソース言語に更新してください。 cURL の例の実行には jq が必要です。
REST API
REST API
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
MODEL_NAME="qwen3-asr-flash"
AUDIO_URL="${AIFUNC_AUDIO_URL:-https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3}"
AIFUNC_AUDIO_LANGUAGE="${AIFUNC_AUDIO_LANGUAGE:-zh}"
BODY=$(cat <<JSON
{
"model_name": "$MODEL_NAME",
"texts": ["$AUDIO_URL"],
"params": {"language": "$AIFUNC_AUDIO_LANGUAGE", "enable_itn": true, "max_concurrency": 1}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/audio_transcribe" "$BODY")"
if command -v jq >/dev/null 2>&1; then
echo "$RESPONSE_BODY" | jq .
[ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
echo "$RESPONSE_BODY"
fiPython
Python
from __future__ import annotations
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"
DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9
def texttransform_function_type() -> Any:
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
function_type = getattr(FunctionType, type_name, None)
if function_type is not None:
return function_type
# Alibaba Cloud Milvus では、TEXTTRANSFORM がホストされた拡張機能 (関数タイプ値 9) として提供されています。
# 一部の pymilvus バージョンにはこの列挙型メンバーがまだ組み込まれていませんが、Function(...) は FunctionType(...) を介して検証を行います。
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
def add_id(schema: Any) -> None:
schema.add_field("id", DataType.INT64, is_primary=True)
def add_dummy_vector(schema: Any) -> None:
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)
def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
for name, data_type, max_length in input_fields:
field_params = {"max_length": max_length} if max_length is not None else {}
schema.add_field(name, data_type, **field_params)
output_name, output_data_type, output_max_length = output_field
output_params = {"max_length": output_max_length} if output_max_length is not None else {}
schema.add_field(output_name, output_data_type, **output_params)
add_dummy_vector(schema)
schema.add_function(
Function(
name=function_name,
function_type=texttransform_function_type(),
input_field_names=[name for name, _, _ in input_fields],
output_field_names=[output_name],
params=function_params,
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, rows)
client.flush(collection_name)
fields = [name for name, _, _ in input_fields] + [output_name]
for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
print(row)
MODEL_NAME = "qwen3-asr-flash"
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
run_texttransform_example(
client=client,
collection_name="simple_ai_audio_transcribe_schema",
input_fields=[("audio_url", DataType.VARCHAR, 4096)],
output_field=("transcript", DataType.VARCHAR, 4096),
function_name="transcribe_audio",
function_params={"provider": "aliyun_milvus", "model_name": MODEL_NAME, "task": "ai_audio_transcribe", "language": "zh", "enable_itn": True},
rows=[{"audio_url": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "dummy_vector": [0.1, 0.2]}],
)期待される結果: data.output.outputs[0] には、4096文字以内の空ではないウェルカムメッセージの文字起こしが返され、ホットラインのバージョンアーカイブの transcript フィールドに書き込むことができます (実際のテストでの戻り値:「欢迎与使用阿里云。」)。
{"code": 0, "data": {"output": {"outputs": ["欢迎与使用阿里云。"]}}}Python のシナリオでは、入力フィールドを audio_url、出力フィールドを transcript と命名できます。 Function は insert 時に自動で文字起こしを実行します。