すべてのプロダクト
Search
ドキュメントセンター

Alibaba Cloud Model Studio:サーバーイベント

最終更新日:Sep 02, 2026

Qwen-Omni-Realtime API のサーバーイベントです。ツール呼び出し (関数呼び出し) イベントも含まれます。

Qwen-Omni-Realtime をご参照ください。

error

サーバーエラーです。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。値は常にerrorです。

errorobject

エラーの詳細です。

プロパティ

typestring

エラータイプです。

codestring

エラーコードです。

messagestring

エラーメッセージです。

paramstring

エラーに関連付けられたパラメーター、 session.modalities など。

{
  "event_id": "event_RoUu4T8yExPMI37GKwaOC",
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "code": "invalid_value",
    "message": "Invalid modalities: ['audio']. Supported combinations are: ['text'] and ['audio', 'text'].",
    "param": "session.modalities"
  }
}

session.created

クライアントが接続したときに返されます。デフォルトの会話構成が含まれます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプです。この値は常に session.created です。

sessionobject

セッションの構成。

プロパティ

objectstring

この値は常に realtime.session です。

modelstring

使用されるモデルです。

modalitiesarray

モデルの出力モダリティです。

voicestring

モデルが生成するオーディオの音声です。

input_audio_formatstring

入力音声フォーマットは pcm (16 kHz サンプルレート) のみがサポートされています。

output_audio_formatstring

出力オーディオフォーマット。サポートされるのは pcm (サンプルレート 24 kHz) のみです。

input_audio_transcriptionobject

文字起こし構成です。

プロパティ

modelstring

文字起こしモデル。値は常に qwen3-asr-flash-realtime で、設定変更はできません。

turn_detectionobject

音声区間検出 (VAD) の構成です。

プロパティ

typestring

VAD タイプ。有効な値は server_vad (デフォルト) または semantic_vad です。「クライアントイベント」をご参照ください。

thresholdfloat

VAD の検出のしきい値です。

silence_duration_msinteger

発話終了検出をトリガーする無音継続時間 (ミリ秒) です。

idle_timeout_msinteger

ミリ秒単位のアイドルタイムアウト。 server_vad モードで qwen3.5-omni-plus-realtime または qwen3.5-omni-flash-realtime モデルを使用している場合にのみ返されます。

enable_searchboolean

Web 検索を有効にするかどうかを指定します。Qwen3.5-Omni-Realtime シリーズモデルでのみサポートされています。

search_optionsobject

Web 検索のオプションです。

temperaturefloat

モデルの温度パラメーターです。

{
    "event_id": "event_RdvlSpbBb2ssyBjYrDHjt",
    "type": "session.created",
    "session": {
        "object": "realtime.session",
        "model": "qwen3-omni-flash-realtime",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "input_audio_format": "pcm",
        "output_audio_format": "pcm",
        "input_audio_transcription": {
            "model": "qwen3-asr-flash-realtime"
        },
        "turn_detection": {
            // 値は server_vad または semantic_vad (qwen3.5-omni-realtime シリーズモデルでのみサポート) です。
            "type": "server_vad",
            "threshold": 0.5,
            "prefix_padding_ms": 300,
            "silence_duration_ms": 800,
            "create_response": true,
            "interrupt_response": true
        },
        "enable_search": false,
        "search_options": {},
        "tools": [],
        "temperature": 0.8,
        "id": "sess_Ov7GOXoNXhNjlxXtOGKQS"
    }
}

session.updated

session.update リクエストが成功した後に返され、失敗した場合は代わりに error イベントが返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に session.updated です。

sessionobject

セッションの構成

プロパティ

temperaturefloat

モデルの温度パラメーターです。

modalitiesarray

モデルの出力モダリティです。

voicestring

モデルが生成するオーディオの音声です。

instructionsstring

モデルの目標とロールです。

input_audio_formatstring

入力オーディオフォーマット。pcm のみがサポートされています (サンプルレート 16 kHz)。

output_audio_formatstring

出力音声フォーマットは pcm (24 kHz サンプルレート) のみがサポートされています。

input_audio_transcriptionobject

文字起こし構成です。

プロパティ

modelstring

文字起こしモデル。常に qwen3-asr-flash-realtime であり、設定はできません。

turn_detectionobject

音声区間検出 (VAD) の構成です。

プロパティ

typestring

VAD のタイプ。有効な値は server_vad (デフォルト) または semantic_vad です。「クライアントイベント」をご参照ください。

thresholdfloat

VAD の検出のしきい値です。

silence_duration_msinteger

発話終了検出をトリガーする無音継続時間 (ミリ秒) です。

idle_timeout_msinteger

ミリ秒単位のアイドルタイムアウトです。server_vad モードで、qwen3.5-omni-plus-realtime または qwen3.5-omni-flash-realtime モデルを使用する場合にのみ返されます。

enable_searchboolean (任意)

Web 検索を有効にするかどうかを指定します。Qwen3.5-Omni-Realtime シリーズモデルでのみサポートされています。

search_optionsobject (任意)

Web 検索のオプションです。

toolsarray (任意)

ツール定義です。設定されている場合、モデルはユーザーの入力に基づいてツールを呼び出すかどうかを決定できます。

プロパティ

typestring (必須)

この値は常に function です。

function.namestring (必須)

get_current_weather や get_current_time などの関数名です。

function.descriptionstring (任意)

関数の目的の説明です。モデルはこれを使用して関数を呼び出すかどうかを決定します。

function.parametersobject (任意)

入力パラメーターのスキーマです。モデルはこれを使用してパラメーターを抽出します。関数がパラメーターを受け取らない場合は省略します。

プロパティ

typestring (必須)

この値は常にオブジェクトです。

propertiesobject (任意)

各キーはパラメーター名で、type と description を持つオブジェクトにマッピングされます。

requiredarray (任意)

どの入力パラメーターが必須かを指定します。

top_pfloat

核サンプリングの確率のしきい値です。

top_kinteger

生成中のサンプリングのための候補セットのサイズです。

max_tokensinteger

このリクエストに対してモデルが返すことができる最大トークン数です。

repetition_penaltyfloat

生成中の連続するシーケンスにおける繰り返しを制御します。

presence_penaltyfloat

生成されたコンテンツ内の繰り返しを制御します。

seedinteger

リクエストごとのモデル出力の一貫性の度合いです。

{
    "event_id": "event_X1HsXS4b4uptp6yo1LgKd",
    "type": "session.updated",
    "session": {
        "id": "sess_Aih6vAcY5Ddt6jwFx1tCa",
        "object": "realtime.session",
        "model": "qwen3-omni-flash-realtime",
        "modalities": [
            "text",
            "audio"
        ],
        "instructions": "You are Xiao Yun, a personal assistant. Answer user questions accurately and in a friendly manner. Always respond with a helpful attitude.",
        "voice": "Cherry",
        "input_audio_format": "pcm",
        "output_audio_format": "pcm",
        "input_audio_transcription": {
            "model": "qwen3-asr-flash-realtime"
        },
        "turn_detection": {
            // 値は server_vad または semantic_vad (qwen3.5-omni-realtime シリーズモデルでのみサポート) です。
            "type": "server_vad",
            "threshold": 0.1,
            "prefix_padding_ms": 500,
            "silence_duration_ms": 900,
            "create_response": true,
            "interrupt_response": true
        },
        "enable_search": true,
        "search_options": {
            "enable_source": true
        },
        "tools": [
            {
                "type": "function",
                "function": {
                    "name": "get_current_weather",
                    "description": "Useful for querying the weather in a specific city.",
                    "parameters": {
                        "type": "object",
                        "properties": {
                            "location": {"type": "string", "description": "The city name"}
                        },
                        "required": ["location"]
                    }
                }
            }
        ],
        "temperature": 0.8,
        "max_response_output_token": "inf",
        "max_tokens": 16384,
        "repetition_penalty": 1.05,
        "presence_penalty": 0.0,
        "top_k": 50,
        "top_p": 1.0,
        "seed":-1
    }
}

input_audio_buffer.speech_started

VAD モードでは、サーバーがオーディオバッファー内で発話開始を検出したときに返されます。

発話が検出される前にオーディオがバッファーに追加されるたびにトリガーされることがあります。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に input_audio_buffer.speech_started です。

audio_start_msinteger

オーディオバッファーへの書き込み開始から、最初の発話が検出されるまでの時間 (ミリ秒) です。

item_idstring

発話終了が検出されたときに作成されるユーザーメッセージアイテムの ID です。

ユーザーメッセージアイテムは、モデル推論のためにユーザー入力を会話履歴に追加します。

{
    "event_id": "event_Pvp8nEhsQuGCQbFJ9x58n",
    "type": "input_audio_buffer.speech_started",
    "audio_start_ms": 3647,
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba"
}

input_audio_buffer.speech_stopped

VAD モードでは、サーバーがオーディオバッファー内で発話終了を検出したときに返されます。

また、対応するユーザーメッセージアイテムを含む conversation.item.created イベントを返します。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプです。この値は常に input_audio_buffer.speech_stopped です。

audio_end_msinteger

会話の開始から発話終了が検出されるまでの時間 (ミリ秒) です。

item_idstring

作成されるユーザーメッセージアイテムの ID です。

{
    "event_id": "event_UhQiqNVRsgUiq4KUS5Xb5",
    "type": "input_audio_buffer.speech_stopped",
    "audio_end_ms": 4453,
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba"
}

input_audio_buffer.committed

入力オーディオバッファーがコミットされたときに返されます。

  • VAD モードでは、サーバーは発話終了を検出すると自動的にバッファーをコミットします。
  • 手動モードでは、クライアントが input_audio_buffer.commit イベントを送信した後に返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に input_audio_buffer.committed です。

item_idstring

作成されるユーザーメッセージアイテムの ID です。

{
    "event_id": "event_Iy6sUzL1nmdFgshFYxJEz",
    "type": "input_audio_buffer.committed",
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba"
}

input_audio_buffer.cleared

クライアントが input_audio_buffer.clear イベントを送信した後に返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に input_audio_buffer.cleared です。

{
  "event_id": "event_RoUu4T8yExPMI37GKwaOC",
  "type": "input_audio_buffer.cleared"
}

conversation.item.created

会話アイテムが作成されたときに返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に conversation.item.created です。

itemobject

追加する会話アイテムです。

プロパティ

idstring

会話アイテムの一意の ID です。

objectstring

この値は常に realtime.item です。

statusstring

会話アイテムのステータスです。

rolestring

メッセージのロールです。

contentarray

メッセージの内容です。このパラメーターは、タイプが message の場合に返されます。

typestring

会話アイテムのタイプ。有効な値は message または function_call です。

namestring

タイプが function_call の場合に呼び出される関数の名前です。

call_idstring

type が function_call の場合、関数呼び出しの一意の ID を表します。

argumentsstring

type が function_call の場合、このパラメーターには、関数呼び出しの引数が JSON 文字列として含まれます。

{
    "event_id": "event_JEfkrr9gO3Ny7Xcv9bGVd",
    "type": "conversation.item.created",
    "item": {
        "id": "item_YbAiGvK2H7YaS34o4R6Ba",
        "object": "realtime.item",
        "type": "message",
        "status": "in_progress",
        "role": "assistant",
        "content": [
            {
                "type": "input_audio"
            }
        ]
    }
}
// ツール呼び出しシナリオ
{
    "event_id": "event_S1hkaIQgcuQD8OEdOpGHQ",
    "type": "conversation.item.created",
    "item": {
        "id": "item_FEG9qJGNkPcdf4et3p7BV",
        "object": "realtime.item",
        "type": "function_call",
        "status": "in_progress",
        "call_id": "call_bc0a7fb7235840f69ecfe4",
        "name": "get_current_weather",
        "arguments": ""
    }
}

conversation.item.input_audio_transcription.delta

入力音声の文字起こしが有効な場合、ユーザーの発話中に頻繁に送信されます。リアルタイムの中間文字起こし結果を提供します。text + stash を連結すると、どの時点でも最も完全な文章のプレビューを取得できます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に conversation.item.input_audio_transcription.delta です。

item_idstring

関連する会話アイテムの ID です。

content_indexinteger

オーディオを含むコンテンツパートのインデックスです。

textstring

確定したテキストプレフィックス — モデルが確定し、変更されない部分です。

stashstring

未確定のテキストサフィックス — 確定部分に続く一時的な下書きで、修正される可能性があります。

languagestring

認識されたオーディオの検出された言語です。

emotionstring

認識された音声から検出された感情。有効な値: neutral、 happy、 sad、 angry、 surprised、 disgusted、 fearful。

{
    "event_id": "event_C7jzoeSFuiwOZS6tR14yx",
    "type": "conversation.item.input_audio_transcription.delta",
    "item_id": "item_ThVYhLHOdeXb4bBSvzSFF",
    "content_index": 0,
    "text": "",
    "stash": "How is the weather today?",
    "language": "en",
    "emotion": "neutral",
    "obfuscation": "ABEXGYmxdmc97u"
}

いつでも最も完全な文章のプレビューを取得するには、これら 2 つのフィールドを連結します: リアルタイムプレビュー = text + stash。

クリックして例を表示

ユーザーが「今日の天気は良いですね、晴れていて暖かいです。」と言っているとします。

以下に、受信する可能性のあるイベントストリームとその解釈方法を示します:

時間

ユーザーの発話の進行状況

API 応答 (text と stash)

クライアント UI 表示 (text + stash)

T1

「天気は...」

text: ""

stash: "天気"

天気は

T2

「...良いですね...」

text: ""

stash: "天気は良いですね"

天気は良いですね

T3

「...今日は、」

text: "天気は"

stash: "良いですね今日は、"

天気は良いですね今日は、

(「天気は」が確定し、text に移動しました)

T4

(短い間)

text: "天気は良いですね今日は、"

stash: ""

天気は良いですね今日は、

(最初の句が完全に確定しました)

T5

「晴れていて...」

text: "天気は良いですね今日は、"

stash: "sunny"

今日は天気が良く、晴れです。

T6

「...そして暖かいです。」

text: "天気は良いですね今日は、"

stash: "晴れていて暖かいです。"

天気は良いですね今日は、晴れていて暖かいです。

T7

(発話停止)

-

最終結果として conversation.item.input_audio_transcription.completed の文字起こしを使用します。

conversation.item.input_audio_transcription.completed

ユーザーの音声が、組み込みの音声認識モデル (qwen3-asr-flash-realtime) によって文字起こしされたことを示します。設定不可。

音声認識モデルによる文字起こしテキストは、Qwen-Omni-Realtime モデルによって生成された解釈と異なる場合があります。この文字起こしは参照用です。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプです。この値は常に conversation.item.input_audio_transcription.completed です。

item_idstring

ユーザーメッセージアイテムの ID です。

content_indexinteger

この値は常に 0 です。

transcriptstring

文字起こしされたテキストです。

{
    "event_id": "event_FrrZcxiDfTB9LD9p4pVng",
    "type": "conversation.item.input_audio_transcription.completed",
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba",
    "content_index": 0,
    "transcript": "Hello."
}

conversation.item.input_audio_transcription.failed

入力音声の文字起こしが有効な状態でこれに失敗した場合に返されます。error イベントとは独立しているため、クライアントは文字起こしの失敗を具体的に特定できます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプです。この値は常に conversation.item.input_audio_transcription.failed です。

item_idstring

ユーザーメッセージアイテムの ID です。

content_indexinteger

この値は常に 0 です。

errorobject

エラー情報です。

プロパティ

code string

エラーコードです。

message string

エラーメッセージです。

param string

エラーに関連するパラメーターです。

{
  "type": "conversation.item.input_audio_transcription.failed",
  "item_id": "<item_id>",
  "content_index": 0,
  "error": {
    "code": "<code>",
    "message": "<message>",
    "param": "<param>"
  }
}

response.created

サーバーが新しい応答の生成を開始したときに返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.created です。

responseobject

応答オブジェクトです。

プロパティ

id string

応答の一意の ID です。

conversation_id string

現在の会話の一意の ID です。

object string

オブジェクトタイプ。このイベントでは、この値は常に realtime.response です。

status string

応答ステータス。有効な値はcompleted, failed, in_progress, or incompleteです。

modalities array

応答モダリティです。

voice string

モデルが生成するオーディオの音声です。

output array

このイベントでは、このフィールドは空です。

{
    "event_id": "event_XuDavMzQN3KKepqGu3KRh",
    "type": "response.created",
    "response": {
        "id": "resp_HaVOPdbmX6vifiV5pAfJY",
        "object": "realtime.response",
        "conversation_id": "conv_FjJaccpnvwHNo9cPVuzGc",
        "status": "in_progress",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "output_audio_format": "pcm",
        "output": []
    }
}

response.done

応答が完全に生成された後に返されます。response オブジェクトには、raw 音声データを除くすべての出力項目が含まれます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.done です。

responseobject

応答オブジェクトです。

プロパティ

id string

応答の一意の ID です。

conversation_id string

現在の会話の一意の ID です。

object string

オブジェクトタイプで、このイベントでは値は常に realtime.response になります。

status string

応答ステータスです。

modalities array

応答モダリティです。

voice string

モデルが生成するオーディオの音声です。

output object

応答の出力です。

プロパティ

id string

応答出力の ID です。

type string

出力項目のタイプ。有効な値は、メッセージ または 関数呼び出し です。

object string

出力項目のオブジェクトタイプ。値は常に realtime.item です。

status string

出力アイテムのステータスです。

role string

出力アイテムのロールです。

content array

出力項目の内容です。このフィールドは、type が message の場合にのみ返されます。

プロパティ

type string

コンテンツタイプ。値は、プレーンテキスト出力の場合は text、オーディオ出力の場合は audio です。

text string

テキスト出力です。

transcript string

オーディオのテキスト文字起こしです。

name string

タイプが function_call の場合に呼び出される関数の名前です。

call_id string

type が function_call の場合、これは関数呼び出しの一意の ID です。

arguments string

type が function_call の場合、このフィールドには関数呼び出しの完全な引数が JSON 文字列として含まれます。

usage object

この応答のトークン使用量の詳細です。

プロパティ

total_tokens integer

この応答で使用された合計トークン数です。

input_tokens integer

入力トークンの数です。

output_tokens integer

出力トークンの数です。

input_tokens_details object

text_tokens や audio_tokens を含む、入力 トークン の使用量に関する詳細。

output_tokens_details object

text_tokens と audio_tokens を含む、出力トークンの使用量に関する詳細。

plugins object (任意)

プラグインの使用量メトリック。Web 検索 (enable_search) が有効な場合に返されます。

プロパティ

search object

検索メータリングデータです。

プロパティ

count integer

検索の数です。

strategy string

検索戦略です。

{
    "event_id": "event_CSaxRRYLvbrfexDXAEuDG",
    "type": "response.done",
    "response": {
        "id": "resp_HaVOPdbmX6vifiV5pAfJY",
        "object": "realtime.response",
        "conversation_id": "conv_FjJaccpnvwHNo9cPVuzGc",
        "status": "completed",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "output_audio_format": "pcm",
        "output": [
            {
                "id": "item_Ls6MtCUWO7LM4E59QziNv",
                "object": "realtime.item",
                "type": "message",
                "status": "completed",
                "role": "assistant",
                "content": [
                    {
                        "type": "audio",
                        "transcript": "Hello! How can I help you?"
                    }
                ]
            }
        ],
        "usage": {
            "total_tokens": 377,
            "input_tokens": 336,
            "output_tokens": 41,
            "input_tokens_details": {
                "text_tokens": 228,
                "audio_tokens": 108
            },
            "output_tokens_details": {
                "text_tokens": 9,
                "audio_tokens": 32
            },
            "plugins": {
                "search": {
                    "count": 1,
                    "strategy": "agent"
                }
            }
        }
    }
}
// ツール呼び出しシナリオ
{
    "event_id": "event_T1EFAJp43X2DWtDRmxTtx",
    "type": "response.done",
    "response": {
        "id": "resp_TucN5QgymL5MA8vkJvFlS",
        "object": "realtime.response",
        "conversation_id": "conv_SEDZESRlefT8WvLSmEn6E",
        "status": "completed",
        "modalities": ["text", "audio"],
        "voice": "Ethan",
        "output_audio_format": "pcm",
        "output": [
            {
                "id": "item_FEG9qJGNkPcdf4et3p7BV",
                "object": "realtime.item",
                "type": "function_call",
                "status": "completed",
                "call_id": "call_bc0a7fb7235840f69ecfe4",
                "name": "get_current_weather",
                "arguments": " {\"location\": \"Hangzhou\"}"
            }
        ],
        "usage": {
            "total_tokens": 567,
            "input_tokens": 524,
            "output_tokens": 43,
            "input_tokens_details": {
                "text_tokens": 487,
                "audio_tokens": 37
            },
            "output_tokens_details": {
                "text_tokens": 43
            }
        }
    }
}

response.text.delta

出力モダリティがテキストのみで、モデルが新しいテキストを増分的に生成するときに返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常にresponse.text.deltaです。

deltastring

モデルによって生成された増分テキストです。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。この値は常に 0 です。

content_indexinteger

出力アイテム内の内部パートのインデックスです。この値は常に 0 です。

{
    "delta": "Hello",
    "event_id": "event_TH49MauuPmRo1RGaMSlP7",
    "type": "response.text.delta",
    "response_id": "resp_PrRSvPVpnCExdUOGHHLuP",
    "item_id": "item_L8IRm9kRXFpxoOjDqDC96",
    "output_index": 0,
    "content_index": 0
}

response.text.done

出力モダリティがテキストのみで、モデルがテキストの生成を終了したときに返されます。

応答が中断、不完全、またはキャンセルされた場合にも返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.text.done です。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。

content_indexinteger

応答内の出力アイテムのインデックスです。

text string

モデルによって生成された完全なテキストです。

{
  "event_id": "event_B1lIeE2Nac33zn5V7h2mm",
  "type": "response.text.done",
  "response_id": "resp_B1lIdtjF4Noqpn5NOjznj",
  "item_id": "item_B1lIdJsAJlJiFs8ztWpJt",
  "output_index": 0,
  "content_index": 0,
  "text": "How can I assist you today?"
}

response.audio.delta

出力モダリティにオーディオが含まれ、モデルが新しいオーディオデータを増分的に生成するときに返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.audio.delta です。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。

content_indexinteger

応答内の出力アイテムのインデックスです。

delta string

増分オーディオデータ、Base64 エンコードされています。

{
  "event_id": "event_B1osWMZBtrEQbiIwW0qHQ",
  "type": "response.audio.delta",
  "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
  "item_id": "item_OFaPGtzfWCPyGzxnuEX9i",
  "output_index": 0,
  "content_index": 0,
  "delta": "{base64 audio}"
}

response.audio.done

出力モダリティにオーディオが含まれ、モデルがオーディオデータの生成を終了したときに返されます。

応答が中断、不完全、またはキャンセルされた場合にも返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.audio.done です。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。

content_indexinteger

応答内の出力アイテムのインデックスです。

{
    "event_id": "event_Le1TDl7VfyHQxl47DtGxI",
    "type": "response.audio.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0
}

response.audio_transcript.delta

出力モダリティにオーディオが含まれ、モデルが新しい response.audio_transcript.delta トランスクリプトテキストを増分的に生成する場合に返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.audio_transcript.delta です。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。

content_indexinteger

応答内の出力アイテムのインデックスです。

deltastring

増分テキストです。

{
    "event_id": "event_BksW7fOwnyavZdDxIzZYM",
    "type": "response.audio_transcript.delta",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "delta": "What"
}

response.audio_transcript.done

出力モダリティに音声が含まれ、モデルが音声文字起こしを完了すると、response.audio_transcript.done イベントとして返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.audio_transcript.done です。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。

content_indexinteger

応答内の出力アイテムのインデックスです。

transcriptstring

完全な文字起こしテキストです。

{
    "event_id": "event_X49tL2WerT4WjxcmH16lS",
    "type": "response.audio_transcript.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "transcript": "Hello! How can I help you?"
}

response.function_call_arguments.delta

モデルが関数呼び出しの引数をストリーミングする際に返されます。delta フィールドを連結して、引数文字列を構築します。完全なコンテンツは、後続の response.function_call_arguments.done イベントで提供されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプです。この値は常に response.function_call_arguments.delta です。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。

call_idstring

この関数呼び出しの一意の ID です。この ID は、同じターン内の done イベントと一致します。

deltastring

引数文字列の新しいセグメントです。セグメントを順番に連結します。

{
    "event_id": "event_SlKoJyEbPEqLq14DSM1u5",
    "type": "response.function_call_arguments.delta",
    "response_id": "resp_JnTOsWXlFhKcFohZbtfz6",
    "item_id": "item_Rhcms7CauTNsQprV5S4Hr",
    "output_index": 0,
    "call_id": "call_2be200f4cafe419b9530dd",
    "delta": " {\"location\": \"Beijing\"}"
}

response.function_call_arguments.done

関数呼び出しの引数がすべて生成されたことを示します。arguments フィールドには、完全な引数文字列が含まれます。ローカル関数を解析して呼び出すには、連結された delta の結果ではなく、このイベントの arguments を使用します。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプです。この値は常に response.function_call_arguments.done です。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。

call_idstring

この関数呼び出しの一意の ID です。

namestring

呼び出された関数の名前です。

argumentsstring

完全な関数呼び出しの引数を JSON 文字列として含みます。

{
    "event_id": "event_X6suLyuL5agdH7r6koesM",
    "type": "response.function_call_arguments.done",
    "response_id": "resp_JnTOsWXlFhKcFohZbtfz6",
    "item_id": "item_Rhcms7CauTNsQprV5S4Hr",
    "output_index": 0,
    "name": "get_current_weather",
    "call_id": "call_2be200f4cafe419b9530dd",
    "arguments": " {\"location\": \"Beijing\"}"
}

response.output_item.added

応答生成中に新しいアイテムが作成されると返され、アイテムタイプは message または function_call です。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプです。この値は常に response.output_item.addedです。

response_idstring

応答 ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。

itemobject

出力アイテムに関する情報です。

プロパティ

idstring

出力アイテムの一意の ID です。

objectstring

この値は常に realtime.item です。

statusstring

出力アイテムのステータスです。

rolestring

送信者のロールです。

contentarray

メッセージのコンテンツです。このフィールドは、type が message の場合に返されます。

typestring

出力項目のタイプを示します。有効な値は message または function_call です。

namestring

type が function_call の場合に呼び出す関数の名前。

call_idstring

タイプが function_call の場合、現在の関数呼び出しの一意の ID になります。

argumentsstring

type が function_call の場合、関数呼び出しの引数は JSON 文字列になります。added イベントでは、最初は空の文字列です。

{
    "event_id": "event_DsCO341DEVtiATtCB6BUY",
    "type": "response.output_item.added",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "output_index": 0,
    "item": {
        "id": "item_Ls6MtCUWO7LM4E59QziNv",
        "object": "realtime.item",
        "type": "message",
        "status": "in_progress",
        "role": "assistant",
        "content": []
    }
}
// ツール呼び出しシナリオ
{
    "event_id": "event_HXmKt5pGoiRtXx7Hq7zpN",
    "type": "response.output_item.added",
    "response_id": "resp_TucN5QgymL5MA8vkJvFlS",
    "output_index": 0,
    "item": {
        "id": "item_FEG9qJGNkPcdf4et3p7BV",
        "object": "realtime.item",
        "type": "function_call",
        "status": "in_progress",
        "call_id": "call_bc0a7fb7235840f69ecfe4",
        "name": "get_current_weather",
        "arguments": ""
    }
}

response.output_item.done

出力アイテムが完全に生成されたときに返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.output_item.done です。

response_idstring

応答 ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。

itemobject

出力アイテム情報です。

プロパティ

idstring

出力アイテムの一意の ID です。

objectstring

この値は常に realtime.item です。

statusstring

出力アイテムのステータスです。

rolestring

送信者のロールです。

contentarray

メッセージの内容です。このフィールドは、type が message の場合に返されます。

typestring

出力項目のタイプ。有効な値は message または function_call です。

namestring

タイプが function_call の場合に呼び出される関数の名前です。

call_idstring

type が function_call の場合、これは関数呼び出しの一意の ID です。

argumentsstring

type が function_call の場合、すべての関数呼び出しの引数を JSON 文字列として含みます。

{
    "event_id": "event_MEu5nlLw1LsOguHiehIP8",
    "type": "response.output_item.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "output_index": 0,
    "item": {
        "id": "item_Ls6MtCUWO7LM4E59QziNv",
        "object": "realtime.item",
        "type": "message",
        "status": "completed",
        "role": "assistant",
        "content": [
            {
                "type": "audio",
                "text": "Hello! How can I help you?"
            }
        ]
    }
}
// ツール呼び出しシナリオ
{
    "event_id": "event_FHspdfAnCyjuME3mmAwSY",
    "type": "response.output_item.done",
    "response_id": "resp_TucN5QgymL5MA8vkJvFlS",
    "output_index": 0,
    "item": {
        "id": "item_FEG9qJGNkPcdf4et3p7BV",
        "object": "realtime.item",
        "type": "function_call",
        "status": "completed",
        "call_id": "call_bc0a7fb7235840f69ecfe4",
        "name": "get_current_weather",
        "arguments": " {\"location\": \"Hangzhou\"}"
    }
}

response.content_part.added

応答生成中にアシスタントメッセージアイテムに新しいコンテンツパートが追加されたときに返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.content_part.added です。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。この値は常に 0 です。

content_indexinteger

出力アイテム内の内部パートのインデックスです。この値は常に 0 です。

partobject

出力アイテム情報です。

プロパティ

typestring

コンテンツパートのタイプです。

textstring

コンテンツパートのテキストです。

{
    "event_id": "event_AVBOmrgY3C8bjlRajfSUT",
    "type": "response.content_part.added",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "part": {
        "type": "audio",
        "text": ""
    }
}

response.content_part.done

アシスタントメッセージアイテム内のコンテンツパートのストリーミングが終了したときに返されます。

event_idstring

このイベントの一意の識別子です。

typestring

イベントタイプ。この値は常に response.content_part.done です。

response_idstring

応答 ID です。

item_idstring

メッセージアイテム ID です。

output_indexinteger

応答内の出力アイテムのインデックスです。この値は常に 0 です。

content_indexinteger

コンテンツ配列内のコンテンツパートのインデックスです。この値は常に 0 です。

partobject

出力アイテム情報です。

プロパティ

typestring

コンテンツパートのタイプです。

textstring

コンテンツパートのテキストです。

{
    "event_id": "event_Il8HD19v58Qr5IBkw7LtN",
    "type": "response.content_part.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "part": {
        "type": "audio",
        "text": "Hello! How can I help you?"
    }
}