全部產品
Search
文件中心

Alibaba Cloud Model Studio:服務端事件

更新時間:Aug 26, 2026

本文介紹 qwen3.5-livetranslate-flash-realtime API 的服務端事件。

相關文檔:即時語音/音視頻翻譯-千問

error

服務端返回的錯誤資訊。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為error

errorobject

錯誤的詳細資料。

屬性

typestring

錯誤類型。

codestring

錯誤碼。

messagestring

錯誤資訊。

paramstring

與錯誤相關的參數,如session.modalities

{
  "event_id": "event_RoUu4T8yExPMI37GKwaOC",
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "code": "invalid_value",
    "message": "Invalid modalities: ['audio']. Supported combinations are: ['text'] and ['audio', 'text'].",
    "param": "session.modalities"
  }
}

session.created

用戶端串連後,服務端返回的第一個事件,包含本次串連的預設配置資訊。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為session.created

sessionobject

會話的配置。

屬性

idstring

會話的唯一識別碼。

objectstring

固定為realtime.session

modelstring

使用的模型。

modalitiesarray

模型輸出模態設定。

voicestring

模型產生音訊音色。

input_audio_formatstring

輸入音訊格式,預設為pcm

sample_rateinteger

輸入音訊採樣率,單位為Hz。

output_audio_formatstring

輸出音訊格式,預設為pcm

turn_detectionobject

VAD(語音活動檢測)配置。

translationobject(可選)

翻譯配置。

屬性

languagestring(可選)

設定的翻譯目標語種。

corpusobject (可選)

熱詞配置,用於提升特定詞彙的翻譯準確性。

屬性

corpus.phrasesobject (可選)

熱詞映射表。key 為源語言詞彙,value 為目標語言對應翻譯,參見支援的語種

{
    "event_id": "event_QxBGpjBDmDDQQWDtrqBKB",
    "type": "session.created",
    "session": {
        "id": "sess_OozZ1vtbPt2muDflHODIH",
        "object": "realtime.session",
        "model": "qwen3.5-livetranslate-flash-realtime",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "input_audio_format": "pcm",
        "sample_rate": 16000,
        "output_audio_format": "pcm",
        "turn_detection": {
            "type": "server_vad",
            "threshold": 0.2,
            "silence_duration_ms": 1000
        },
        "translation": {
           "language": "en",
           "corpus": {
               "phrases": {
                   "人工智慧": "Artificial Intelligence",
                   "機器學習": "Machine Learning"
               }
           }
        }
    }
}

session.updated

收到使用者的 session.update 請求後,若處理成功,則返回此事件;若出錯,則返回 error 事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為session.updated

sessionobject

會話的配置。

屬性

idstring

會話的唯一識別碼。

objectstring

固定為realtime.session

modelstring

使用的模型。

modalitiesarray

模型輸出模態設定。

voicestring

模型產生音訊音色。

sample_rateinteger(可選)

輸入音訊採樣率。

input_audio_formatstring

輸入音訊格式,固定為pcm

output_audio_formatstring

輸出音訊格式,固定為pcm

input_audio_transcriptionobject

輸入音頻轉錄配置。僅在會話配置了input_audio_transcription.model參數時返回。

屬性

modelstring

語音辨識模型。

languagestring

設定的語音辨識語種。

turn_detectionobject

VAD(語音活動檢測)配置。Manual 模式下(用戶端在session.update中將該參數設為null)不返回此欄位。

屬性

typestring

VAD 類型,固定為server_vad

thresholdfloat

VAD 檢測靈敏度。

prefix_padding_msinteger

語音開始前保留的音頻時間長度(毫秒),避免丟失語音起始部分。

silence_duration_msinteger

語音結束後需保持靜音的最短時間長度(毫秒),超過該時間長度即判定語音結束。

create_responseboolean

VAD 檢測到語音結束後,是否自動觸發翻譯響應。

interrupt_responseboolean

VAD 檢測到新一輪語音開始時,是否打斷當前正在產生的翻譯響應。

translationobject(可選)

翻譯配置。

屬性

languagestring(可選)

設定的翻譯目標語種。

corpusobject(可選)

熱詞配置,用於提升特定詞彙的翻譯準確性。

屬性

corpus.phrasesobject(可選)

熱詞映射表。key 為源語言詞彙,value 為目標語言對應翻譯。

enable_voice_cloneboolean

是否啟用聲音複刻。

voice_clone_optionsobject

聲音複刻控制參數,僅在enable_voice_clonetrue時返回。

屬性

frequencystring

音色複刻頻率。

{
    "event_id": "event_QxBGpjBDmDDQQWDtrqBKB",
    "type": "session.updated",
    "session": {
        "id": "sess_OozZ1vtbPt2muDflHODIH",
        "object": "realtime.session",
        "model": "qwen3.5-livetranslate-flash-realtime",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Ethan",
        "sample_rate": 16000,
        "input_audio_format": "pcm",
        "output_audio_format": "pcm",
        "input_audio_transcription": {
            "model": "qwen3-asr-flash-realtime",
            "language": "zh"
        },
        "turn_detection": {
            "type": "server_vad",
            "threshold": 0.2,
            "prefix_padding_ms": 300,
            "silence_duration_ms": 1000,
            "create_response": true,
            "interrupt_response": true
        },
        "translation": {
           "language": "en",
           "corpus": {
               "phrases": {
                   "人工智慧": "Artificial Intelligence",
                   "機器學習": "Machine Learning"
               }
           }
        },
        "enable_voice_clone": true,
        "voice_clone_options": {
            "frequency": "once"
        }
    }
}

session.finished

會話結束事件,表示當前會話中,所有音頻翻譯已完成。

該事件在用戶端發送session.finish後才會發送,用戶端接收到該事件後可主動中斷連線。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為session.finished

{
    "event_id": "event_xxx",
    "type": "session.finished"
}

response.created

當服務端產生新的模型響應時,會返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.created

responseobject

響應對象。

屬性

id string

響應的唯一識別碼。

conversation_id string

當前會話的唯一識別碼。

object string

物件類型,此事件下固定為realtime.response

status string

響應狀態,取值範圍:

  • completed(已完成)
  • failed(失敗)
  • in_progress(進行中)
  • incomplete(不完整)

modalities array

響應的模態。

voice string

模型產生音訊音色。

output_audio_formatstring

輸出音訊格式。

output string

此事件下目前為空白。

{
    "event_id": "event_L8hHVI5jYis6BzAjnPWJh",
    "type": "response.created",
    "response": {
        "id": "resp_P79OOMs8LnrXVpiIHUCKR",
        "object": "realtime.response",
        "conversation_id": "conv_UFClXtYkRkFXrs48y8pmK",
        "status": "in_progress",
        "modalities": [
            "text",
            "audio"
        ],
        "voice": "Cherry",
        "output_audio_format": "pcm16",
        "output": []
    }
}

response.done

響應產生完成後,服務端會返回此事件。事件中的 response 對象包含除原始音頻資料外的全部輸出項。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.done

responseobject

響應對象。

屬性

id string

響應的唯一識別碼。

conversation_id string

當前會話的唯一識別碼。

object string

物件類型,此事件下固定為realtime.response

status string

響應的狀態。

modalities array

響應的模態。

voice string

模型產生音訊音色。

output_audio_formatstring

輸出音訊格式。

output object

響應的輸出。

屬性

id string

響應輸出的唯一識別碼。

type string

輸出項的類型,當前固定為message

object string

輸出項的物件類型,當前固定為realtime.item

status string

輸出項的狀態。

role string

輸出項的角色。

content array

輸出項的內容。

屬性

type string

輸出內容的類型。輸出為純文字時,為text;輸出包含音頻時,為audio

text string

輸出的常值內容。

transcript string

音頻轉錄為文字後的內容。

usage object

本次響應的 Token 消耗資訊。

{
  "event_id": "event_CNea8oXNipVanSg2VIzkO",
  "type": "response.done",
  "response": {
    "id": "resp_TfhYTqej692vsGA2jNEtH",
    "object": "realtime.response",
    "conversation_id": "conv_ZtyLfKVm8XqLwYRlsuDih",
    "status": "completed",
    "modalities": [
      "text",
      "audio"
    ],
    "voice": "Cherry",
    "output_audio_format": "pcm16",
    "output": [
      {
        "id": "item_MKtkMwN9RtcyE9eJShyWy",
        "object": "realtime.item",
        "type": "message",
        "status": "completed",
        "role": "assistant",
        "content": [
          {
            "type": "audio",
            "transcript": "Hello? "
          }
        ]
      }
    ],
    "usage": {
      "total_tokens": 56,
      "input_tokens": 47,
      "output_tokens": 9,
      "input_tokens_details": {
        "text_tokens": 20,
        "audio_tokens": 27
      },
      "output_tokens_details": {
        "text_tokens": 2,
        "audio_tokens": 7
      }
    }
  }
}

response.text.text

當輸出模態僅包含文本,且模型增量產生新的文本時,服務端將返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.text.text

textstring

返回的增量文本。

response_idstring

回複的ID。

item_idstring

訊息項ID,可以關聯同一個訊息項。

output_indexinteger

目前固定為 0。

content_indexinteger

目前固定為 0。

stashstring

初步產生的臨時文本,與當前 text 拼接後構成臨時產生結果;系統會通過 response.text.text 事件持續更新 text 和 stash,直至收到response.text.done事件,此時可通過 text 欄位擷取完整的最終文本。

{
    "event_id": "event_B1lIeyOXR7qJMEExbqtTG",
    "type": "response.text.text",
    "response_id": "resp_B1lIdtjF4Noqpn5NOjznj",
    "item_id": "item_B1lIdJsAJlJiFs8ztWpJt",
    "output_index": 0,
    "content_index": 0,
    "text": "How are",
    "stash": " you today?"
}

response.text.done

當輸出模態僅包含文本,且模型產生的文本結束時,服務端返回此事件。

當響應中斷、不完整或取消時,服務端也會返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.text.done

response_idstring

響應的唯一識別碼。

item_idstring

訊息項的唯一識別碼。

output_indexinteger

目前固定為 0。

content_indexinteger

目前固定為 0。

text string

模型輸出的完整文本。

{
    "event_id": "event_B1lIeE2Nac33zn5V7h2mm",
    "type": "response.text.done",
    "response_id": "resp_B1lIdtjF4Noqpn5NOjznj",
    "item_id": "item_B1lIdJsAJlJiFs8ztWpJt",
    "output_index": 0,
    "content_index": 0,
    "text": "How can I assist you today?"
}

response.audio.delta

當輸出模態包含音頻,且模型增量產生新的音頻資料時,服務端將返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.audio.delta

response_idstring

響應的唯一識別碼。

item_idstring

訊息項唯一識別碼。

output_indexinteger

目前固定為 0。

content_indexinteger

目前固定為 0。

delta string

模型增量輸出的音頻資料,使用Base64編碼。

{
    "event_id": "event_B1osWMZBtrEQbiIwW0qHQ",
    "type": "response.audio.delta",
    "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
    "item_id": "item_OFaPGtzfWCPyGzxnuEX9i",
    "output_index": 0,
    "content_index": 0,
    "delta": "UklGRnoGAABXQVZFZm10IBAAAAAB..."
}

response.audio.done

當輸出模態包含音頻,且模型產生音頻結束時,服務端返回此事件。

當響應中斷、不完整或取消時,服務端也會返回此事件。

該事件不返回完整音頻資料。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.audio.done

response_idstring

響應的唯一識別碼。

item_idstring

訊息項唯一識別碼。

output_indexinteger

目前固定為 0。

content_indexinteger

目前固定為 0。

{
    "event_id": "event_B1osWMWoDRYyITDyNYcBu",
    "type": "response.audio.done",
    "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
    "item_id": "item_OFaPGtzfWCPyGzxnuEX9i",
    "output_index": 0,
    "content_index": 0
}

input_audio_buffer.speech_started

當服務端 VAD(語音活動檢測)檢測到使用者開始說話時,返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為input_audio_buffer.speech_started

audio_start_msinteger

檢測到語音開始的時間點(毫秒),相對於音頻流開始的位移量。

item_idstring

關聯的訊息項唯一識別碼。

{
    "event_id": "event_xxx",
    "type": "input_audio_buffer.speech_started",
    "audio_start_ms": 568,
    "item_id": "item_xxx"
}

input_audio_buffer.speech_stopped

當服務端 VAD 檢測到使用者停止說話時,返回此事件,標誌本輪語音輸入結束。由於翻譯響應基於流式語音同步產生,實際的翻譯響應可能已經在語音輸入過程中開始,無需等待此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為input_audio_buffer.speech_stopped

audio_end_msinteger

檢測到語音結束的時間點(毫秒),相對於音頻流開始的位移量。

item_idstring

關聯的訊息項唯一識別碼。

{
    "event_id": "event_xxx",
    "type": "input_audio_buffer.speech_stopped",
    "audio_end_ms": 3900,
    "item_id": "item_xxx"
}

input_audio_buffer.committed

Manual 模式(turn_detectionnull)下,用戶端發送input_audio_buffer.commit事件後,服務端返回此事件進行確認,並自動開始產生翻譯響應。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為input_audio_buffer.committed

{
    "event_id": "event_xxx",
    "type": "input_audio_buffer.committed"
}

input_audio_buffer.cleared

用戶端發送input_audio_buffer.clear事件後,服務端返回此事件進行確認,表示已清空緩衝區中尚未提交的音頻資料。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為input_audio_buffer.cleared

{
    "event_id": "event_xxx",
    "type": "input_audio_buffer.cleared"
}

conversation.item.created

當對話中建立新的訊息項時,服務端返回此事件。以下情境會觸發此事件:

  • 服務端開始產生翻譯響應時,建立對應的 assistant 訊息項(此時content為空白數組,內容隨流式響應逐步填充)。
  • Manual 模式下,用戶端發送input_audio_buffer.commit事件後,服務端會額外建立一個對應使用者輸入音訊訊息項(content中包含{"type": "input_audio"})。

對於同一個 VAD 片段,服務端會分別建立語音辨識結果和翻譯結果訊息項。語音辨識結果訊息項的 item.id 與翻譯結果事件的 previous_item_id 相同。用戶端可據此關聯原文和譯文,並同時展示。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為conversation.item.created

previous_item_idstring

前一訊息項的唯一識別碼。對於翻譯結果事件,該值與同一 VAD 片段的語音辨識結果訊息項 item.id 相同。

itemobject

訊息項資訊。

屬性

idstring

訊息項的唯一識別碼。

typestring

固定為message

objectstring

固定為realtime.item

statusstring

訊息項的狀態。

rolestring

訊息的角色,取值為assistantuser

contentarray

訊息的內容。響應剛建立時為空白數組,隨流式響應逐步填充;Manual 模式下 commit 產生的使用者訊息項中包含{"type": "input_audio"}

{
    "event_id": "event_xxx",
    "type": "conversation.item.created",
    "previous_item_id": "item_asr_xxx",
    "item": {
        "id": "item_translation_xxx",
        "object": "realtime.item",
        "type": "message",
        "status": "in_progress",
        "role": "assistant",
        "content": []
    }
}

conversation.item.input_audio_transcription.text

當配置了input_audio_transcription.model參數時,服務端會流式返回輸入音訊語音辨識結果(源語言原文)。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為conversation.item.input_audio_transcription.text

item_idstring

訊息項唯一識別碼。

content_indexinteger

目前固定為 0。

textstring

已確認的識別文本。

stashstring

待確認的識別文本(可能會被後續事件修正)。

languagestring

檢測到的源語種。

emotionstring

被識別音訊情感。支援的情感如下:

  • surprised:驚訝
  • neutral:平靜
  • happy:愉快
  • sad:悲傷
  • disgusted:厭惡
  • angry:憤怒
  • fearful:恐懼
{
    "event_id": "event_xxx",
    "type": "conversation.item.input_audio_transcription.text",
    "item_id": "item_xxx",
    "content_index": 0,
    "text": "",
    "stash": "今天天氣真好",
    "language": "zh",
    "emotion": "neutral"
}

conversation.item.input_audio_transcription.completed

當配置了input_audio_transcription.model參數時,語音辨識完成後服務端會返回此事件,包含最終的完整識別結果。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為conversation.item.input_audio_transcription.completed

item_idstring

訊息項唯一識別碼。

content_indexinteger

目前固定為 0。

transcriptstring

完整的語音辨識結果(源語言原文)。

languagestring

檢測到的源語種。

emotionstring

被識別音訊情感。支援的情感如下:

  • surprised:驚訝
  • neutral:平靜
  • happy:愉快
  • sad:悲傷
  • disgusted:厭惡
  • angry:憤怒
  • fearful:恐懼
{
    "event_id": "event_xxx",
    "type": "conversation.item.input_audio_transcription.completed",
    "item_id": "item_xxx",
    "content_index": 0,
    "transcript": "今天天氣真好,我們一起去公園散步吧。",
    "language": "zh",
    "emotion": ""
}

conversation.item.input_audio_transcription.failed

當輸入了音頻但識別失敗時,服務端發送該事件。與其他error事件分開處理,便於用戶端識別相關的具體專案。

typestring

事件類型,固定為conversation.item.input_audio_transcription.failed

item_idstring

關聯的對話項 ID。

content_indexinteger

包含音訊內容部分的索引。

error.codestring

錯誤碼。

error.messagestring

錯誤訊息。

{
    "event_id": "event_xxx",
    "type": "conversation.item.input_audio_transcription.failed",
    "item_id": "item_xxx",
    "content_index": 0,
    "error": {
        "code": "xxx",
        "message": "xxx",
        "param": "xxx"
    }
}

response.audio_transcript.text

當輸出模態包含音頻時,服務端可能返回此事件,用於展示即時翻譯內容。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.audio_transcript.text

response_idstring

響應的唯一識別碼。

item_idstring

訊息項唯一識別碼。

output_indexinteger

目前固定為 0。

content_indexinteger

目前固定為 0。

textstring

已確認無誤的翻譯文本片段。

stashstring

初步翻譯的臨時文本,與當前 text 拼接後構成臨時翻譯結果;系統會通過 response.audio_transcript.text 事件持續更新 text 和 stash,直至收到response.audio_transcript.done事件,此時可通過 transcript 欄位擷取完整的最終翻譯文本。

{
  "event_id": "event_xxx",
  "type": "response.audio_transcript.text",
  "response_id": "resp_xxx",
  "item_id": "item_xxx",
  "output_index": 0,
  "content_index": 0,
  "text": "Hello,",
  "stash": " who are you?"
}

response.audio_transcript.done

當輸出模態包含音頻,且模型產生文本結束時,服務端返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.audio_transcript.done

response_idstring

響應的唯一識別碼。

item_idstring

訊息項唯一識別碼。

output_indexinteger

目前固定為 0。

content_indexinteger

目前固定為 0。

transcriptstring

完整文本。

{
    "event_id": "event_VN4Q4GJugLcc1S23viW8E",
    "type": "response.audio_transcript.done",
    "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
    "item_id": "item_JvJauNH2CTXb1D9WV6pD4",
    "output_index": 0,
    "content_index": 0,
    "transcript": "How can I assist you today?"
}

response.output_item.added

在響應產生過程中建立新輸出項時,服務端返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.output_item.added

response_idstring

響應的唯一識別碼。

output_indexinteger

目前固定為 0。

itemobject

輸出項資訊。

屬性

idstring

輸出項的唯一識別碼。

typestring

固定為 message

objectstring

始終為 realtime.item 。

statusstring

輸出項的狀態。

rolestring

訊息的角色。

contentstring

訊息的內容。

{
    "event_id": "event_B4O5yPt3Gjnjy5eYH3plG",
    "type": "response.output_item.added",
    "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
    "output_index": 0,
    "item": {
        "id": "item_OFaPGtzfWCPyGzxnuEX9i",
        "object": "realtime.item",
        "type": "message",
        "status": "in_progress",
        "role": "assistant",
        "content": []
    }
}

response.output_item.done

當新的項輸出完成時,服務端返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.output_item.done

response_idstring

響應的唯一識別碼。

output_indexinteger

目前固定為 0。

itemobject

輸出項資訊。

屬性

idstring

輸出項的唯一識別碼。

objectstring

始終為 realtime.item 。

typestring

固定為 message

statusstring

輸出項的狀態。

rolestring

發送訊息的角色。

contentstring

訊息的內容。

{
    "event_id": "event_XkiwbYTBC9Wcdwy6uYJ2G",
    "type": "response.output_item.done",
    "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
    "output_index": 0,
    "item": {
        "id": "item_JvJauNH2CTXb1D9WV6pD4",
        "object": "realtime.item",
        "type": "message",
        "status": "completed",
        "role": "assistant",
        "content": [
            {
                "type": "audio",
                "text": "你好,我是阿里雲研發的大規模語言模型,我叫千問。有什麼我可以協助你的嗎?"
            }
        ]
    }
}

response.content_part.added

當新的內容部分輸出時,服務端返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.content_part.added

response_idstring

響應的唯一識別碼。

item_idstring

訊息項唯一識別碼。

output_indexinteger

目前固定為 0。

content_indexinteger

目前固定為 0。

partobject

輸出項資訊。

屬性

typestring

內容部分的類型。

textstring

內容部分的文本。

{
    "event_id": "event_J2UixwYKZsXg7c9YXZetL",
    "type": "response.content_part.added",
    "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
    "item_id": "item_OFaPGtzfWCPyGzxnuEX9i",
    "output_index": 0,
    "content_index": 0,
    "part": {
        "type": "audio",
        "text": ""
    }
}

response.content_part.done

當新的內容部分輸出完成時,服務端返回此事件。

event_idstring

本次事件唯一識別碼。

typestring

事件類型,固定為response.content_part.done

response_idstring

響應的唯一識別碼。

item_idstring

訊息項唯一識別碼。

output_indexinteger

目前固定為 0。

content_indexinteger

目前固定為 0。

partobject

輸出項資訊。

屬性

typestring

內容部分的類型。

textstring

內容部分的文本。

{
    "event_id": "event_VN4Q4GJugLcc1S23viW8E",
    "type": "response.content_part.done",
    "response_id": "resp_P79OOMs8LnrXVpiIHUCKR",
    "item_id": "item_JvJauNH2CTXb1D9WV6pD4",
    "output_index": 0,
    "content_index": 0,
    "part": {
        "type": "audio",
        "text": "你好,我是阿里雲研發的大規模語言模型,我叫千問。有什麼我可以協助你的嗎?"
    }
}