全部产品
Search
文档中心

大模型服务平台百炼:服务端事件

更新时间:Sep 24, 2026

本文介绍 Qwen-Omni-Realtime API 的服务端事件,包括工具调用(Function Calling)相关事件。

error

服务端返回的错误信息。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为error。

errorobject

错误的详细信息。

属性

typestring

错误类型。

codestring

错误码。

messagestring

错误信息。

paramstring

与错误相关的参数,如session.modalities。

{
  "event_id": "event_RoUu4T8yExPMI37GKwaOC",
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "code": "invalid_value",
    "message": "Invalid modalities: ['audio']. Supported combinations are: ['text'] and ['audio', 'text'].",
    "param": "session.modalities"
  }
}

session.created

客户端连接后,服务端返回的第一个事件,包含本次连接的默认配置信息。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为session.created。

sessionobject

会话的配置信息。

属性

objectstring

固定为realtime.session。

modelstring

使用的模型。

modalitiesarray

模型输出模态设置。

voicestring

模型生成音频的音色。

input_audio_formatstring

用户输入音频的格式,当前仅支持设为pcm。输入音频要求为16 kHz采样率的PCM音频流。

output_audio_formatstring

模型输出音频的格式,当前仅支持设为pcm。输出音频为24 kHz采样率的PCM音频流。当前不支持自定义输出采样率。

input_audio_transcriptionobject

语音转录的配置。

属性

modelstring

语音转录模型,固定为qwen3-asr-flash-realtime,不支持修改。

turn_detectionobject

语音活动检测(VAD)的配置。

属性

typestring

VAD类型。取值:server_vad(默认值)或 semantic_vad。详情请参见客户端事件。

thresholdfloat

VAD检测阈值。

silence_duration_msinteger

检测语音停止的静音持续时间。

idle_timeout_msinteger

静默超时时间(毫秒)。仅在 server_vad 模式下,使用 qwen3.5-omni-plus-realtime 或 qwen3.5-omni-flash-realtime 模型时返回。

enable_searchboolean

是否启用联网搜索功能。Qwen3.8-Omni-Flash-Realtime 和 Qwen3.5-Omni-Realtime 系列模型支持。

search_optionsobject

联网搜索选项配置。

temperaturefloat

模型的温度参数。

{
    "event_id": "event_RdvlSpbBb2ssyBjYrDHjt",
    "type": "session.created",
    "session": {
        "object": "realtime.session",
        "model": "qwen3-omni-flash-realtime",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "input_audio_format": "pcm",
        "output_audio_format": "pcm",
        "input_audio_transcription": {
            "model": "qwen3-asr-flash-realtime"
        },
        "turn_detection": {
            // 取值为server_vad或semantic_vad(qwen3.8-omni-flash-realtime和qwen3.5-omni-realtime系列支持)
            "type": "server_vad",
            "threshold": 0.5,
            "prefix_padding_ms": 300,
            "silence_duration_ms": 800,
            "create_response": true,
            "interrupt_response": true
        },
        "enable_search": false,
        "search_options": {},
        "tools": [],
        "temperature": 0.8,
        "id": "sess_Ov7GOXoNXhNjlxXtOGKQS"
    }
}

session.updated

收到用户的 session.update 请求后,若处理成功,则返回此事件;若出错,则返回 error 事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为session.updated。

sessionobject

会话的配置信息。

属性

temperaturefloat

模型的温度参数。

modalitiesarray

模型输出模态设置。

voicestring

模型生成音频的音色。

instructionsstring

模型的目标与角色。

audioobject

回显的音频格式配置。若客户端传入了 session.audio.input.format / session.audio.output.format,服务端将在 session.updated 中按相同嵌套结构回显。未使用嵌套字段的客户端,服务端事件结构保持原有行为。

属性

audio.input.format.typestring

用户输入音频格式。可选值:pcm(默认值)、wav。

audio.input.format.sample_rateinteger

用户输入音频采样率,单位为 Hz。

audio.output.format.typestring

模型输出音频格式。可选值:pcm(默认值)、wav。

audio.output.format.sample_rateinteger

模型输出音频采样率,单位为 Hz。

input_audio_formatstring

历史兼容字段,回显客户端配置的输入音频格式。

output_audio_formatstring

历史兼容字段,回显客户端配置的输出音频格式。

input_audio_transcriptionobject

语音转录的配置。

属性

modelstring

语音转录模型,固定为qwen3-asr-flash-realtime,不支持修改。

turn_detectionobject

语音活动检测(VAD)的配置。

属性

typestring

VAD类型。取值:server_vad(默认值)或 semantic_vad。详情请参见客户端事件。

thresholdfloat

VAD检测阈值。

silence_duration_msinteger

检测语音停止的静音持续时间。

idle_timeout_msinteger

静默超时时间(毫秒)。仅在 server_vad 模式下,使用 qwen3.5-omni-plus-realtime 或 qwen3.5-omni-flash-realtime 模型时返回。

enable_searchboolean(可选)

是否启用联网搜索功能。Qwen3.8-Omni-Flash-Realtime 和 Qwen3.5-Omni-Realtime 系列模型支持。

search_optionsobject(可选)

联网搜索选项配置。

toolsarray(可选)

工具定义列表。配置后模型可根据用户输入自主决定是否调用工具。

属性

typestring(必选)

固定为 function。

function.namestring(必选)

自定义的工具函数名称,建议使用与函数相同的名称,如get_current_weather或get_current_time。

function.descriptionstring(可选)

对工具函数功能的描述,大模型会参考该字段来选择是否使用该工具函数。

function.parametersobject(可选)

对工具函数入参的描述,大模型会参考该字段来进行入参的提取。如果工具函数不需要输入参数,则无需指定。

属性

typestring(必选)

固定为 object。

propertiesobject(可选)

描述各入参的名称、数据类型与描述。Key 值为入参的名称,Value 值为包含数据类型(type)与描述(description)的对象。

requiredarray(可选)

指定哪些入参为必填项。

top_pfloat

核采样的概率阈值。

top_kinteger

模型生成过程中,采样候选集的大小。

max_tokensinteger

模型在本次请求返回的最大 Token 数。

repetition_penaltyfloat

控制模型生成时,连续序列中的重复度*。*

presence_penaltyfloat

控制模型在生成内容时的重复度。

seedinteger

模型在每次请求时,运行结果一致性程度。

{
    "event_id": "event_X1HsXS4b4uptp6yo1LgKd",
    "type": "session.updated",
    "session": {
        "id": "sess_Aih6vAcY5Ddt6jwFx1tCa",
        "object": "realtime.session",
        "model": "qwen3.5-omni-flash-realtime",
        "modalities": [
            "text",
            "audio"
        ],
        "audio": {
            "input": {
                "format": {
                    "type": "pcm",
                    "sample_rate": 16000
                }
            },
            "output": {
                "format": {
                    "type": "wav",
                    "sample_rate": 24000
                }
            }
        },
        "instructions": "你是个人助理小云,请你准确且友好地解答用户的问题,始终以乐于助人的态度回应。",
        "voice": "Tina",
        "input_audio_format": "pcm",
        "output_audio_format": "pcm",
        "input_audio_transcription": {
            "model": "qwen3-asr-flash-realtime"
        },
        "turn_detection": {
            // 取值为server_vad或semantic_vad(qwen3.8-omni-flash-realtime和qwen3.5-omni-realtime系列支持)
            "type": "server_vad",
            "threshold": 0.1,
            "prefix_padding_ms": 500,
            "silence_duration_ms": 900,
            "create_response": true,
            "interrupt_response": true
        },
        "enable_search": true,
        "search_options": {
            "enable_source": true
        },
        "tools": [
            {
                "type": "function",
                "function": {
                    "name": "get_current_weather",
                    "description": "当你想查询指定城市的天气时非常有用。",
                    "parameters": {
                        "type": "object",
                        "properties": {
                            "location": {"type": "string", "description": "城市名称"}
                        },
                        "required": ["location"]
                    }
                }
            }
        ],
        "temperature": 0.8,
        "max_response_output_token": "inf",
        "max_tokens": 16384,
        "repetition_penalty": 1.05,
        "presence_penalty": 0.0,
        "top_k": 50,
        "top_p": 1.0,
        "seed":-1
    }
}

input_audio_buffer.speech_started

在 VAD 模式下,当服务端在音频缓冲区中检测到语音开始时,会返回此事件。

若服务端尚未检测到语音,则每次向缓冲区添加音频时都可能触发此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为input_audio_buffer.speech_started。

audio_start_msinteger

从音频开始写入缓冲区到首次检测到语音所经过的毫秒数。

item_idstring

语音停止时将创建的用户消息项的 ID。

用户消息项用于将用户输入追加到对话历史,供模型后续推理与生成使用。

{
    "event_id": "event_Pvp8nEhsQuGCQbFJ9x58n",
    "type": "input_audio_buffer.speech_started",
    "audio_start_ms": 3647,
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba"
}

input_audio_buffer.speech_stopped

在 VAD 模式下,当音频缓冲区中检测到语音结束时,服务端会返回此事件。

同时,服务端还会返回一个 conversation.item.created 事件,以创建对应的用户消息项。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为input_audio_buffer.speech_stopped。

audio_end_msinteger

语音停止时刻距会话开始经过的毫秒数。

item_idstring

将创建的用户消息项的 ID。

{
    "event_id": "event_UhQiqNVRsgUiq4KUS5Xb5",
    "type": "input_audio_buffer.speech_stopped",
    "audio_end_ms": 4453,
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba"
}

input_audio_buffer.committed

当输入音频缓冲区被提交时返回此事件。

  • 在VAD模式下,当检测到用户说话结束时,服务端会自动提交音频缓冲区并返回此事件。
  • 在 Manual 模式下,当客户端发送input_audio_buffer.commit事件后,服务端返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为input_audio_buffer.committed。

item_idstring

将创建的用户消息项的 ID。

{
    "event_id": "event_Iy6sUzL1nmdFgshFYxJEz",
    "type": "input_audio_buffer.committed",
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba"
}

input_audio_buffer.cleared

客户端发送input_audio_buffer.clear事件后,服务端将返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为input_audio_buffer.cleared。

{
  "event_id": "event_RoUu4T8yExPMI37GKwaOC",
  "type": "input_audio_buffer.cleared"
}

conversation.item.created

当对话项创建时返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为conversation.item.created。

itemobject

要添加到对话中的项。

属性

idstring

对话项的唯一ID。

objectstring

始终为 realtime.item 。

statusstring

对话项的状态。

rolestring

消息的角色。

contentstring

消息的内容。当 type 为 message 时存在。

typestring

对话项的类型包括 message(常规消息)和 function_call(工具调用)。Qwen3.8-Omni-Flash-Realtime 还可返回 mcp_list_tools、mcp_call、mcp_approval_request,对应结构见MCP Item。

namestring

当 type 为 function_call 时,被调用的函数名称。

call_idstring

当 type 为 function_call 时,本次函数调用的唯一 ID。

argumentsstring

当 type 为 function_call 时,函数调用的参数(JSON 字符串)。

{
    "event_id": "event_JEfkrr9gO3Ny7Xcv9bGVd",
    "type": "conversation.item.created",
    "item": {
        "id": "item_YbAiGvK2H7YaS34o4R6Ba",
        "object": "realtime.item",
        "type": "message",
        "status": "in_progress",
        "role": "assistant",
        "content": [
            {
                "type": "input_audio"
            }
        ]
    }
}
// 工具调用场景
{
    "event_id": "event_S1hkaIQgcuQD8OEdOpGHQ",
    "type": "conversation.item.created",
    "item": {
        "id": "item_FEG9qJGNkPcdf4et3p7BV",
        "object": "realtime.item",
        "type": "function_call",
        "status": "in_progress",
        "call_id": "call_bc0a7fb7235840f69ecfe4",
        "name": "get_current_weather",
        "arguments": ""
    }
}

conversation.item.input_audio_transcription.delta

开启输入音频转录后,此事件会在用户说话过程中高频发送,用于展示实时识别的中间结果。您可以通过拼接 text + stash 获取当前最完整的句子预览。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为conversation.item.input_audio_transcription.delta。

item_idstring

关联的对话项 ID。

content_indexinteger

包含音频的内容部分的索引。

textstring

已确认的文本前缀。这是当前句子中,模型已确认不会再变更的部分。

stashstring

预识别的文本后缀。这是紧跟在已确认部分之后,模型仍在处理、可能会被修正的临时草稿。

languagestring

被识别音频的语种。

emotionstring

被识别音频的情感。可选值:neutral(平静)、happy(愉快)、sad(悲伤)、angry(愤怒)、surprised(惊讶)、disgusted(厌恶)、fearful(恐惧)。

{
    "event_id": "event_C7jzoeSFuiwOZS6tR14yx",
    "type": "conversation.item.input_audio_transcription.delta",
    "item_id": "item_ThVYhLHOdeXb4bBSvzSFF",
    "content_index": 0,
    "text": "",
    "stash": "今天天气怎么样?",
    "language": "zh",
    "emotion": "neutral",
    "obfuscation": "ABEXGYmxdmc97u"
}

在任何时刻,要获取当前最完整的句子预览,都需要将这两个字段拼接起来:实时预览句子 = text + stash。

点击查看示例

假设用户正在说:"今天天气不错,阳光明媚。"

以下是您可能会收到的事件流以及如何解读它们:

时间点

用户说话进度

API 响应 (text 和 stash)

客户端 UI 应显示 (text + stash)

T1

"今天……"

text: ""

stash: "今天"

今天

T2

"……天气……"

text: ""

stash: "今天天气"

今天天气

T3

"……不错"

text: "今天"

stash: "天气不错"

今天天气不错

(注意,"今天"已被确认并移入text)

T4

(短暂停顿)

text: "今天天气不错,"

stash: ""

今天天气不错,

(前半句完全确认)

T5

"……阳光……"

text: "今天天气不错,"

stash: "阳光"

今天天气不错,阳光

T6

"……明媚。"

text: "今天天气不错,"

stash: "阳光明媚。"

今天天气不错,阳光明媚。

T7

(结束说话)

-

使用 conversation.item.input_audio_transcription.completed 的 transcript 内容作为最终结果。

conversation.item.input_audio_transcription.completed

此事件表示用户音频写入缓冲区后生成的转录结果。其转录由内置的语音识别模型(固定为 qwen3-asr-flash-realtime)处理,不支持修改。

语音识别模型生成的转录文本可能与 Qwen-Omni-Realtime 模型的理解存在差异,仅供参考。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为conversation.item.input_audio_transcription.completed。

item_idstring

用户消息项的 ID。

content_indexinteger

当前固定为0。

transcriptstring

转录的文本内容。

{
    "event_id": "event_FrrZcxiDfTB9LD9p4pVng",
    "type": "conversation.item.input_audio_transcription.completed",
    "item_id": "item_YbAiGvK2H7YaS34o4R6Ba",
    "content_index": 0,
    "transcript": "喂,你好。"
}

conversation.item.input_audio_transcription.failed

启用输入音频转录后,若用户音频转录失败,服务端会返回此事件。此事件独立于 error 事件,便于客户端识别。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为conversation.item.input_audio_transcription.failed。

item_idstring

用户消息项的 ID。

content_indexinteger

当前固定为0。

errorobject

错误信息。

属性

code string

错误码。

message string

错误消息。

param string

错误相关的参数。

{
  "type": "conversation.item.input_audio_transcription.failed",
  "item_id": "<item_id>",
  "content_index": 0,
  "error": {
    "code": "<code>",
    "message": "<message>",
    "param": "<param>"
  }
}

response.created

当服务端生成新的模型响应时,会返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.created。

responseobject

响应对象。

属性

id string

响应的唯一 ID。

conversation_id string

当前会话的唯一ID。

object string

对象类型,此事件下固定为realtime.response。

status string

响应的状态。在[completed, failed, in_progress, or incomplete]范围内。

modalities array

响应的模态。

voice string

模型生成音频的音色。

output string

此事件下目前为空。

{
    "event_id": "event_XuDavMzQN3KKepqGu3KRh",
    "type": "response.created",
    "response": {
        "id": "resp_HaVOPdbmX6vifiV5pAfJY",
        "object": "realtime.response",
        "conversation_id": "conv_FjJaccpnvwHNo9cPVuzGc",
        "status": "in_progress",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "output_audio_format": "pcm",
        "output": []
    }
}

response.done

响应生成完成后,服务端会返回此事件。事件中的 response 对象包含除原始音频数据外的全部输出项。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.done。

responseobject

响应对象。

属性

id string

响应的唯一 ID。

conversation_id string

当前会话的唯一ID。

object string

对象类型,此事件下固定为realtime.response。

status string

响应的状态。

modalities array

响应的模态。

voice string

模型生成音频的音色。

output object

响应的输出。

属性

id string

响应输出对应的ID。

type string

输出项的类型,可选值包括 message(常规消息)、function_call(工具调用),Qwen3.8-Omni-Flash-Realtime 还支持 mcp_call。

object string

输出项的对象类型,当前固定为realtime.item。

status string

输出项的状态。

role string

输出项的角色。

content array

输出项的内容。当 type 为 message 时存在。

属性

type string

输出内容的类型。输出为纯文本时,为text;输出包含音频时,为audio。

text string

输出的文本内容。

transcript string

音频转录为文字后的内容。

name string

当 type 为 function_call 时,被调用的函数名称。

call_id string

当 type 为 function_call 时,函数调用的唯一 ID。

arguments string

当 type 为 function_call 时,函数调用的完整参数(JSON 字符串)。

usage object

本次响应的 Token 消耗信息。

属性

total_tokens integer

本次响应消耗的总 Token 数。

input_tokens integer

输入 Token 数。

output_tokens integer

输出 Token 数。

input_tokens_details object

输入 Token 的分项详情,包含 text_tokens(文本 Token 数)和 audio_tokens(音频 Token 数)。

output_tokens_details object

输出 Token 的分项详情,包含 text_tokens(文本 Token 数)和 audio_tokens(音频 Token 数)。

plugins object(可选)

插件使用计量信息。启用联网搜索(enable_search)时返回。

属性

search object

联网搜索计量信息。

属性

count integer

搜索次数。

strategy string

搜索策略。

{
    "event_id": "event_CSaxRRYLvbrfexDXAEuDG",
    "type": "response.done",
    "response": {
        "id": "resp_HaVOPdbmX6vifiV5pAfJY",
        "object": "realtime.response",
        "conversation_id": "conv_FjJaccpnvwHNo9cPVuzGc",
        "status": "completed",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "output_audio_format": "pcm",
        "output": [
            {
                "id": "item_Ls6MtCUWO7LM4E59QziNv",
                "object": "realtime.item",
                "type": "message",
                "status": "completed",
                "role": "assistant",
                "content": [
                    {
                        "type": "audio",
                        "transcript": "你好呀!有什么我可以帮你的吗?"
                    }
                ]
            }
        ],
        "usage": {
            "total_tokens": 377,
            "input_tokens": 336,
            "output_tokens": 41,
            "input_tokens_details": {
                "text_tokens": 228,
                "audio_tokens": 108
            },
            "output_tokens_details": {
                "text_tokens": 9,
                "audio_tokens": 32
            },
            "plugins": {
                "search": {
                    "count": 1,
                    "strategy": "agent"
                }
            }
        }
    }
}
// 工具调用场景
{
    "event_id": "event_T1EFAJp43X2DWtDRmxTtx",
    "type": "response.done",
    "response": {
        "id": "resp_TucN5QgymL5MA8vkJvFlS",
        "object": "realtime.response",
        "conversation_id": "conv_SEDZESRlefT8WvLSmEn6E",
        "status": "completed",
        "modalities": ["text", "audio"],
        "voice": "Ethan",
        "output_audio_format": "pcm",
        "output": [
            {
                "id": "item_FEG9qJGNkPcdf4et3p7BV",
                "object": "realtime.item",
                "type": "function_call",
                "status": "completed",
                "call_id": "call_bc0a7fb7235840f69ecfe4",
                "name": "get_current_weather",
                "arguments": " {\"location\": \"杭州\"}"
            }
        ],
        "usage": {
            "total_tokens": 567,
            "input_tokens": 524,
            "output_tokens": 43,
            "input_tokens_details": {
                "text_tokens": 487,
                "audio_tokens": 37
            },
            "output_tokens_details": {
                "text_tokens": 43
            }
        }
    }
}

response.text.delta

当输出模态仅包含文本,且模型增量生成新的文本时,服务端将返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.text.delta。

deltastring

返回的增量文本。

response_idstring

回复的ID。

item_idstring

消息项ID,可以关联同一个消息项。

output_indexinteger

响应中输出项的索引, 目前固定为 0。

content_indexinteger

响应中输出项中内部部分的索引, 目前固定为 0。

{
    "delta": "喂",
    "event_id": "event_TH49MauuPmRo1RGaMSlP7",
    "type": "response.text.delta",
    "response_id": "resp_PrRSvPVpnCExdUOGHHLuP",
    "item_id": "item_L8IRm9kRXFpxoOjDqDC96",
    "output_index": 0,
    "content_index": 0
}

response.text.done

当输出模态仅包含文本,且模型生成的文本结束时,服务端将返回此事件。

当响应中断、不完整或取消时,也会返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.text.done。

response_idstring

响应的ID。

item_idstring

消息项ID。

output_indexinteger

响应输出项的索引。

content_indexinteger

响应输出项的索引。

text string

模型输出的完整文本。

{
  "event_id": "event_B1lIeE2Nac33zn5V7h2mm",
  "type": "response.text.done",
  "response_id": "resp_B1lIdtjF4Noqpn5NOjznj",
  "item_id": "item_B1lIdJsAJlJiFs8ztWpJt",
  "output_index": 0,
  "content_index": 0,
  "text": "How can I assist you today?"
}

response.audio.delta

当输出模态包含音频,且模型增量生成新的音频数据时,服务端将返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.audio.delta。

response_idstring

响应的ID。

item_idstring

消息项ID。

output_indexinteger

响应输出项的索引。

content_indexinteger

响应输出项的索引。

delta string

模型增量输出的音频数据,使用Base64编码。

{
  "event_id": "event_B1osWMZBtrEQbiIwW0qHQ",
  "type": "response.audio.delta",
  "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
  "item_id": "item_OFaPGtzfWCPyGzxnuEX9i",
  "output_index": 0,
  "content_index": 0,
  "delta": "{base64 audio}"
}

response.audio.done

当输出模态包含音频,且模型完成生成音频数据时,服务端将返回此事件。

当响应中断、不完整或取消时,也会返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.audio.done。

response_idstring

响应的ID。

item_idstring

消息项ID。

output_indexinteger

响应输出项的索引。

content_indexinteger

响应输出项的索引。

{
    "event_id": "event_Le1TDl7VfyHQxl47DtGxI",
    "type": "response.audio.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0
}

response.audio_transcript.delta

当输出模态包含音频,且模型增量生成新的音频对应的文本时,服务端将返回 response.audio_transcript.delta 事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.audio_transcript.delta。

response_idstring

响应的ID。

item_idstring

消息项ID。

output_indexinteger

响应输出项的索引。

content_indexinteger

响应输出项的索引。

deltastring

增量文本。

{
    "event_id": "event_BksW7fOwnyavZdDxIzZYM",
    "type": "response.audio_transcript.delta",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "delta": "有什么"
}

response.audio_transcript.done

当输出模态包含音频,且模型完成音频转录后,服务端将返回 response.audio_transcript.done 事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.audio_transcript.done。

response_idstring

响应的ID。

item_idstring

消息项ID。

output_indexinteger

响应输出项的索引。

content_indexinteger

响应输出项的索引。

transcriptstring

完整文本。

{
    "event_id": "event_X49tL2WerT4WjxcmH16lS",
    "type": "response.audio_transcript.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "transcript": "你好呀!有什么我可以帮你的吗?"
}

response.function_call_arguments.delta

当模型以流式方式生成函数调用的参数字符串时,每产生一段新内容,服务端推送一次本事件。客户端应按接收顺序将各事件中的 delta 字段拼接,得到与当前进度一致的参数文本;完整内容以随后的 response.function_call_arguments.done 为准。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.function_call_arguments.delta。

response_idstring

响应的ID。

item_idstring

消息项ID。

output_indexinteger

该响应中输出项的索引。

call_idstring

本次函数调用的唯一 ID,与同一轮中的 done 事件保持一致。

deltastring

本段新增的参数字符串片段(增量)。需按顺序拼接。

{
    "event_id": "event_SlKoJyEbPEqLq14DSM1u5",
    "type": "response.function_call_arguments.delta",
    "response_id": "resp_JnTOsWXlFhKcFohZbtfz6",
    "item_id": "item_Rhcms7CauTNsQprV5S4Hr",
    "output_index": 0,
    "call_id": "call_2be200f4cafe419b9530dd",
    "delta": " {\"location\": \"北京\"}"
}

response.function_call_arguments.done

函数调用参数已全部生成完毕。本事件中的 arguments 为完整的参数字符串。客户端可在收到本事件后解析参数并调用本地工具函数;应以本事件中的完整 arguments 为准,而非 delta 拼接结果。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.function_call_arguments.done。

response_idstring

响应的ID。

item_idstring

消息项ID。

output_indexinteger

该响应中输出项的索引。

call_idstring

本次函数调用的唯一 ID。

namestring

被调用的函数名称。

argumentsstring

函数调用的完整参数,一般以 JSON 字符串形式表示。

{
    "event_id": "event_X6suLyuL5agdH7r6koesM",
    "type": "response.function_call_arguments.done",
    "response_id": "resp_JnTOsWXlFhKcFohZbtfz6",
    "item_id": "item_Rhcms7CauTNsQprV5S4Hr",
    "output_index": 0,
    "name": "get_current_weather",
    "call_id": "call_2be200f4cafe419b9530dd",
    "arguments": " {\"location\": \"北京\"}"
}

response.output_item.added

在响应生成过程中创建新项目时,服务端返回此事件。项目类型可以是 message(常规消息)、function_call(工具调用),Qwen3.8-Omni-Flash-Realtime 还支持 mcp_call。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.output_item.added。

response_idstring

响应的ID。

output_indexinteger

响应输出项的索引。

itemobject

输出项信息。

属性

idstring

输出项的唯一ID。

objectstring

始终为 realtime.item 。

statusstring

输出项的状态。

rolestring

发送消息的角色。

contentstring

消息的内容。当 type 为 message 时存在。

typestring

输出项的类型。可选值包括 message(常规消息)、function_call(工具调用),Qwen3.8-Omni-Flash-Realtime 还支持 mcp_call。

namestring

当 type 为 function_call 时,被调用的函数名称。

call_idstring

当 type 为 function_call 时,本次函数调用的唯一 ID。

argumentsstring

当 type 为 function_call 时,函数调用的参数(JSON 字符串)。在 added 事件中初始为空字符串。

{
    "event_id": "event_DsCO341DEVtiATtCB6BUY",
    "type": "response.output_item.added",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "output_index": 0,
    "item": {
        "id": "item_Ls6MtCUWO7LM4E59QziNv",
        "object": "realtime.item",
        "type": "message",
        "status": "in_progress",
        "role": "assistant",
        "content": []
    }
}
// 工具调用场景
{
    "event_id": "event_HXmKt5pGoiRtXx7Hq7zpN",
    "type": "response.output_item.added",
    "response_id": "resp_TucN5QgymL5MA8vkJvFlS",
    "output_index": 0,
    "item": {
        "id": "item_FEG9qJGNkPcdf4et3p7BV",
        "object": "realtime.item",
        "type": "function_call",
        "status": "in_progress",
        "call_id": "call_bc0a7fb7235840f69ecfe4",
        "name": "get_current_weather",
        "arguments": ""
    }
}

response.output_item.done

当新的项目输出完成时,服务端返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.output_item.done。

response_idstring

响应的ID。

output_indexinteger

响应输出项的索引。

itemobject

输出项信息。

属性

idstring

输出项的唯一ID。

objectstring

始终为 realtime.item 。

statusstring

输出项的状态。

rolestring

发送消息的角色。

contentstring

消息的内容。当 type 为 message 时存在。

typestring

输出项的类型。可选值包括 message(常规消息)、function_call(工具调用),Qwen3.8-Omni-Flash-Realtime 还支持 mcp_call。

namestring

当 type 为 function_call 时,被调用的函数名称。

call_idstring

当 type 为 function_call 时,本次函数调用的唯一 ID。

argumentsstring

当 type 为 function_call 时,函数调用的完整参数(JSON 字符串)。

{
    "event_id": "event_MEu5nlLw1LsOguHiehIP8",
    "type": "response.output_item.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "output_index": 0,
    "item": {
        "id": "item_Ls6MtCUWO7LM4E59QziNv",
        "object": "realtime.item",
        "type": "message",
        "status": "completed",
        "role": "assistant",
        "content": [
            {
                "type": "audio",
                "text": "你好呀!有什么我可以帮你的吗?"
            }
        ]
    }
}
// 工具调用场景
{
    "event_id": "event_FHspdfAnCyjuME3mmAwSY",
    "type": "response.output_item.done",
    "response_id": "resp_TucN5QgymL5MA8vkJvFlS",
    "output_index": 0,
    "item": {
        "id": "item_FEG9qJGNkPcdf4et3p7BV",
        "object": "realtime.item",
        "type": "function_call",
        "status": "completed",
        "call_id": "call_bc0a7fb7235840f69ecfe4",
        "name": "get_current_weather",
        "arguments": " {\"location\": \"杭州\"}"
    }
}

response.content_part.added

在响应生成过程中,向助手消息项中添加新内容部分时,服务端返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.content_part.added。

response_idstring

响应的ID。

item_idstring

消息项ID。

output_indexinteger

响应输出项的索引,目前固定为 0。

content_indexinteger

响应输出项中内部部分的索引, 目前固定为 0。

partobject

输出项信息。

属性

typestring

内容部分的类型。

textstring

内容部分的文本。

{
    "event_id": "event_AVBOmrgY3C8bjlRajfSUT",
    "type": "response.content_part.added",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "part": {
        "type": "audio",
        "text": ""
    }
}

response.content_part.done

在助手消息项中的内容部分完成流式传输时,服务端返回此事件。

event_idstring

本次事件唯一标识符。

typestring

事件类型,固定为response.content_part.done。

response_idstring

响应的ID。

item_idstring

消息项ID。

output_indexinteger

响应输出项的索引,目前固定为 0。

content_indexinteger

该项内容数组中内容部分的索引,目前固定为 0。

partobject

输出项信息。

属性

typestring

内容部分的类型。

textstring

内容部分的文本。

{
    "event_id": "event_Il8HD19v58Qr5IBkw7LtN",
    "type": "response.content_part.done",
    "response_id": "resp_HaVOPdbmX6vifiV5pAfJY",
    "item_id": "item_Ls6MtCUWO7LM4E59QziNv",
    "output_index": 0,
    "content_index": 0,
    "part": {
        "type": "audio",
        "text": "你好呀!有什么我可以帮你的吗?"
    }
}

MCP 事件

本节介绍 qwen3.8-omni-flash-realtime 的 WebSocket 事件与字段。基础事件与字段见前述事件说明。表中“必填”表示所属对象出现时必须提供。ID 均为不透明字符串,不应依赖其长度、前缀或生成规则。arguments、output 的外层类型为 string,读取内容时需要再次解析 JSON。MCP 连接、工具发现和调用受服务配额及超时限制。

mcp_list_tools.*

以下三个事件具有相同的字段结构:

事件 type触发时机
mcp_list_tools.in_progress开始发现某个 MCP Server 的工具
mcp_list_tools.completed工具发现成功并完成 allowed_tools 过滤
mcp_list_tools.failed工具发现失败、超时或结果超过服务限制
字段路径类型必填说明
event_idstring是本条服务端事件的唯一 ID
typestring是上表中的固定事件类型
item_idstring是本次工具发现对应的 mcp_list_tools item ID

mcp_list_tools.in_progress 可能早于 session.updated 到达。对于 completed 和 failed,服务端会先发送包含最终工具列表或错误的 conversation.item.created,再发送相同 item_id 的状态事件。

示例:

{
  "event_id": "opaque_event_id",
  "type": "mcp_list_tools.completed",
  "item_id": "opaque_item_id"
}

response.mcp_call_arguments.delta

字段路径类型必填说明
event_idstring是本条服务端事件的唯一 ID
typestring是固定为 response.mcp_call_arguments.delta
response_idstring是父 Response ID
item_idstring是当前 mcp_call item ID
output_indexinteger是当前 item 在父 Response output 数组中的零基索引
deltastring是工具参数 JSON 字符串的本次增量片段,按事件顺序拼接
obfuscationstring否可选混淆字符串;客户端可以忽略,不影响 delta 拼接和参数解析

示例:

{
  "event_id": "opaque_event_id",
  "type": "response.mcp_call_arguments.delta",
  "response_id": "opaque_response_id",
  "item_id": "opaque_item_id",
  "output_index": 0,
  "delta": "{\"city\":\"杭",
  "obfuscation": "opaque-value"
}

response.mcp_call_arguments.done

字段路径类型必填说明
event_idstring是本条服务端事件的唯一 ID
typestring是固定为 response.mcp_call_arguments.done
response_idstring是父 Response ID
item_idstring是当前 mcp_call item ID
output_indexinteger是当前 item 在父 Response output 数组中的零基索引
argumentsstring是完整工具参数,内容为 JSON 字符串;客户端应以本字段为准

示例:

{
  "event_id": "opaque_event_id",
  "type": "response.mcp_call_arguments.done",
  "response_id": "opaque_response_id",
  "item_id": "opaque_item_id",
  "output_index": 0,
  "arguments": "{\"city\":\"杭州\"}"
}

response.mcp_call.*

事件 type触发时机
response.mcp_call.in_progress参数生成完成,且审批通过或无需审批,开始调用 MCP Server
response.mcp_call.completedMCP 工具调用成功
response.mcp_call.failed连接、协议、工具业务错误、审批拒绝、审批超时、调用超时或取消

三个事件均使用以下字段:

字段路径类型必填说明
event_idstring是本条服务端事件的唯一 ID
typestring是上表中的固定事件类型
item_idstring是当前 mcp_call item ID
output_indexinteger是当前 item 在父 Response output 数组中的零基索引

示例:

{
  "event_id": "opaque_event_id",
  "type": "response.mcp_call.in_progress",
  "item_id": "opaque_item_id",
  "output_index": 0
}

MCP 对话项

以下对象是事件中 item 的取值类型,适用于 Qwen3.8-Omni-Flash-Realtime。

mcp_list_tools

该对象通过 conversation.item.created.item 返回。

字段路径类型必填说明
item.idstring是工具列表 item 的唯一 ID,与工具发现状态事件的 item_id 相同
item.typestring是固定为 mcp_list_tools
item.server_labelstring是对应 MCP Server 的标识
item.toolsarray[object]是校验并完成 allowed_tools 过滤后的工具定义;失败时为空数组
item.errorobject否工具发现失败时出现,结构见本节 MCP error

工具数组元素:

字段路径类型必填说明
namestring是MCP Server 声明的原始工具名;1~64 位,仅允许字母、数字、下划线、点和连字符
descriptionstring否MCP Server 提供的工具说明
input_schemaobject是映射自 MCP Server 的 inputSchema,根节点 type 必须为 object
annotationsobject否MCP Server 提供的 ToolAnnotations

input_schema 为 JSON Schema 对象,根节点 type 必须为 object。服务会保留并传递 properties、required、additionalProperties、$defs、oneOf、anyOf、allOf 等标准 JSON Schema 字段。服务端不执行完整的 JSON Schema 语义校验,工具参数的最终合法性由 MCP Server 校验。

示例:

{
  "event_id": "opaque_event_id",
  "type": "conversation.item.created",
  "item": {
    "id": "opaque_item_id",
    "type": "mcp_list_tools",
    "server_label": "amap",
    "tools": [
      {
        "name": "maps_weather",
        "description": "查询天气",
        "input_schema": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string",
              "description": "城市名"
            }
          },
          "required": ["city"]
        }
      }
    ]
  }
}

初始 mcp_call

模型选择 MCP 工具时,该对象出现在 response.output_item.added.item 中,也可能通过 conversation.item.created.item 返回。

字段路径类型必填说明
item.idstring是本次 MCP 调用 item 的唯一 ID
item.objectstring是固定为 realtime.item
item.typestring是固定为 mcp_call
item.statusstring是初始固定为 in_progress
item.call_idstring是本次工具调用的唯一 ID
item.server_labelstring是实际执行工具的 MCP Server 标识
item.namestring是被调用工具的原始名称
item.argumentsstring是初始通常为空字符串;完整参数以 arguments.done 和最终 item 为准

response.output_item.added 示例:

{
  "event_id": "opaque_event_id",
  "type": "response.output_item.added",
  "response_id": "opaque_response_id",
  "output_index": 0,
  "item": {
    "id": "opaque_item_id",
    "object": "realtime.item",
    "type": "mcp_call",
    "status": "in_progress",
    "call_id": "opaque_call_id",
    "server_label": "amap",
    "name": "maps_weather",
    "arguments": ""
  }
}

mcp_approval_request

该对象通过 conversation.item.created 返回。

事件顶层字段:

字段路径类型必填说明
event_idstring是本条服务端事件的唯一 ID
typestring是固定为 conversation.item.created
response_idstring是生成本次 MCP 调用的父 Response ID
item_idstring是审批请求 item ID,与 item.id 相同
previous_item_idstring是被审批的 mcp_call item ID
itemobject是审批请求对象

item 字段:

字段路径类型必填说明
item.idstring是审批请求 ID;回复时原样填入 approval_request_id
item.typestring是固定为 mcp_approval_request
item.server_labelstring是待执行工具所属的 MCP Server 标识
item.namestring是待执行工具的原始名称
item.argumentsstring是待审批调用的完整参数 JSON 字符串
item.call_idstring是被审批的工具调用 ID

示例:

{
  "event_id": "opaque_event_id",
  "type": "conversation.item.created",
  "response_id": "opaque_response_id",
  "item_id": "opaque_approval_id",
  "previous_item_id": "opaque_item_id",
  "item": {
    "id": "opaque_approval_id",
    "type": "mcp_approval_request",
    "server_label": "amap",
    "name": "maps_weather",
    "arguments": "{\"city\":\"杭州\"}",
    "call_id": "opaque_call_id"
  }
}

最终 mcp_call

MCP 调用进入终态后,该对象出现在 response.output_item.done.item 中,并进入父 response.done.response.output 的最终快照。

字段路径类型必填说明
item.idstring是本次 MCP 调用 item 的唯一 ID
item.typestring是固定为 mcp_call
item.statusstring是completed 或 failed
item.call_idstring是本次工具调用的唯一 ID
item.server_labelstring是实际执行工具的 MCP Server 标识
item.namestring是被调用工具的原始名称
item.argumentsstring是完整工具参数的 JSON 字符串
item.outputstring否MCP tools/call.result 对象序列化后的 JSON 字符串;部分失败场景也可能出现
item.errorobject否调用失败时出现,结构见本节 MCP error

成功示例:

{
  "event_id": "opaque_event_id",
  "type": "response.output_item.done",
  "response_id": "opaque_response_id",
  "output_index": 0,
  "item": {
    "id": "opaque_item_id",
    "type": "mcp_call",
    "status": "completed",
    "call_id": "opaque_call_id",
    "server_label": "amap",
    "name": "maps_weather",
    "arguments": "{\"city\":\"杭州\"}",
    "output": "{\"content\":[{\"type\":\"text\",\"text\":\"...\"}],\"isError\":false}"
  }
}

失败示例:

{
  "event_id": "opaque_event_id",
  "type": "response.output_item.done",
  "response_id": "opaque_response_id",
  "output_index": 0,
  "item": {
    "id": "opaque_item_id",
    "type": "mcp_call",
    "status": "failed",
    "call_id": "opaque_call_id",
    "server_label": "amap",
    "name": "maps_weather",
    "arguments": "{\"city\":\"杭州\"}",
    "error": {
      "type": "tool_execution_error",
      "message": "MCP tool call failed (call_timeout)."
    }
  }
}

MCP error

字段路径类型必填说明
error.typestring是固定为 tool_execution_error
error.messagestring是面向客户端的安全错误描述,不包含上游敏感响应体

MCP Server 返回 isError=true 时,最终 mcp_call 的 status 为 failed,并可能同时包含原始 output 和结构化 error。

相关文档:实时(Qwen-Omni-Realtime)。