全部產品
Search
文件中心

Alibaba Cloud Model Studio:WebSocket 接入概覽

更新時間:Sep 29, 2026

了解 WebSocket 的接入位址、驗證、Realtime 與 Inference 互動流程及各模型的事件差異。

前提條件

  • 已開通目標模型或應用程式,並確認其支援的地區。
  • 已取得與呼叫地區、業務空間匹配的 API Key,請參閱Token 驗證。
  • 使用業務空間專屬網域時,已取得 Workspace ID。

連線協定的選型和支援範圍請參閱 Realtime API 概述。

請求標頭

在 WebSocket 握手請求中設定 Authorization。其餘請求標頭請按對應模型的參數說明使用。

請求標頭

是否必需

說明

Authorization

是

使用 Bearer <API_KEY> 格式傳遞 API Key。

user-agent

否

識別呼叫用戶端。

X-DashScope-WorkSpace

按模型要求

指定業務空間 ID,具體使用方式請參閱對應模型的請求標頭說明。

X-DashScope-DataInspection

按模型要求

資料合規檢查設定,支援範圍與取值請參閱對應模型的請求標頭說明。

接入位址

使用 wss:// 協定,根據目標模型選擇 API 路徑和模型名稱的傳遞位置。下表列出各模型的接入方式,支援的模型及地區請參閱對應模型文件。

模型系列

API 路徑

模型名稱傳遞位置

Qwen-Omni-Realtime

/api-ws/v1/realtime

URL 查詢參數 model

Qwen-Audio-TTS/CosyVoice

/api-ws/v1/inference

run-task 的 payload.model

Qwen-TTS-Realtime

/api-ws/v1/realtime

URL 查詢參數 model

Qwen-Audio-ASR/Fun-ASR/Paraformer

/api-ws/v1/inference

run-task 的 payload.model

Qwen-ASR-Realtime

/api-ws/v1/realtime

URL 查詢參數 model

Qwen-Audio-Realtime

/api-ws/v1/realtime

URL 查詢參數 model

Qwen-LiveTranslate-Realtime

/api-ws/v1/realtime

URL 查詢參數 model

業務空間專屬網域:

地域

業務空間專屬網域

華北 2(北京)

{WorkspaceId}.cn-beijing.maas.aliyuncs.com

新加坡

{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com

將 {WorkspaceId} 替換為實際業務空間 ID。例如:

wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime
wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inference

Qwen-TTS-Realtime 也可使用公用網域:北京為 dashscope.aliyuncs.com,新加坡為 dashscope-intl.aliyuncs.com,API 路徑仍為 /api-ws/v1/realtime。API Key 必須與呼叫地區匹配。

通用互動流程

/api-ws/v1/realtime 使用會話制協定,/api-ws/v1/inference 使用任務制協定。

Realtime 協定

用戶端建立連線後,透過 session.update 設定會話,收到 session.updated 後傳送輸入。提交輸入、觸發回應和結束會話的方式由目標模型決定。

  1. 用戶端 → Realtime API:使用 Authorization 和 model 進行 WebSocket 握手。
  2. Realtime API → 用戶端:HTTP 101 和 session.created。
  3. 用戶端 → Realtime API:session.update。Realtime API → 用戶端:session.updated。
  4. 在輸入輸出循環中選擇輸入分支:音訊輸入為用戶端 → Realtime API:input_audio_buffer.append;Qwen-TTS 文字輸入為用戶端 → Realtime API:input_text_buffer.append。
  5. 可選的手動提交:用戶端 → Realtime API:按輸入類型選擇 input_audio_buffer.commit 或 input_text_buffer.commit。
  6. 對於需要顯式觸發回應的模型,用戶端 → Realtime API:response.create。
  7. Realtime API → 用戶端:文字事件 / response.audio.delta。
  8. 可選的產生回應完成事件,不適用於純識別模型:Realtime API → 用戶端:response.done。輸入輸出循環可繼續。
  9. 對於支援結束事件的模型,用戶端 → Realtime API:session.finish;Realtime API → 用戶端:session.finished。
  10. 用戶端 → Realtime API:關閉 WebSocket 連線。

詳細步驟

  1. 建立連線:在 URL 的 model 參數中指定模型,並在請求標頭中攜帶 API Key。握手成功後,伺服器會傳送 session.created。
  2. 設定會話:傳送 session.update,設定模型支援的音訊格式、輸出模態、音色或語音偵測參數,等待 session.updated。各模型支援的會話參數請參閱對應的用戶端事件文件。
  3. 傳送輸入:音訊透過 input_audio_buffer.append 傳送 Base64 資料;Qwen-TTS 透過 input_text_buffer.append 傳送文字。支援影像輸入的模型使用 input_image_buffer.append,並遵循對應模型對圖片和音訊時序的要求。
  4. 提交並取得結果:自動模式下由伺服器觸發處理;手動模式下按輸入類型傳送 input_audio_buffer.commit 或 input_text_buffer.commit。是否還需 response.create,以及應監聽哪些結果事件,請參閱下方差異對照。
  5. 結束會話:支援結束事件的模型使用 session.finish,等待 session.finished 和剩餘結果;其他模型在結果接收完成後關閉 WebSocket 連線。response.done 表示一次回應結束,不表示整個連線關閉;純語音辨識以轉錄完成事件標示辨識結束。

Inference 協定

用戶端傳送 run-task 並等待 task-started 後,才開始傳輸輸入。語音辨識傳送二進位音訊幀,語音合成透過 continue-task 傳送文字;輸入結束後傳送 finish-task,等待 task-finished。

  1. 用戶端 → Inference API:使用 Authorization 進行 WebSocket 握手。
  2. Inference API → 用戶端:HTTP 101。
  3. 用戶端 → Inference API:包含 payload.model 和 task_id 的 run-task。Inference API → 用戶端:task-started。
  4. 輸入輸出循環中的語音辨識分支:用戶端向 Inference API 傳送二進位音訊幀,Inference API 向用戶端返回 result-generated。
  5. 語音合成支線:用戶端 → Inference API:包含文字和 task_id 的 continue-task。Inference API → 用戶端:二進位音訊幀,以及 result-generated(如有)。
  6. 輸入輸出循環結束後,用戶端 → Inference API:使用同一個 task_id 傳送 finish-task。
  7. Inference API → 用戶端:剩餘結果和 task-finished。
  8. 用戶端 → Inference API:關閉連線或依模型要求複用。

詳細步驟

  1. 建立連線:使用 /api-ws/v1/inference,並在請求標頭中攜帶 API Key。
  2. 啟動任務:傳送 run-task,在 payload.model 中指定模型,並設定輸入格式等參數。為任務產生唯一的 header.task_id,等待 task-started。
  3. 傳輸輸入並接收結果:
    • 語音辨識:傳送二進位音訊幀,透過 result-generated 取得辨識結果。
    • 語音合成:透過 continue-task 的 payload.input.text 傳送待合成文字,接收二進位音訊幀;模型還可能透過 result-generated 返回時間戳記等資訊。
  4. 結束任務:傳送 finish-task 後繼續接收剩餘結果,直到收到 task-finished。同一任務的 run-task、continue-task 和 finish-task 必須使用相同的 header.task_id。之後關閉連線,或按對應模型的要求複用連線。

各模型差異

Realtime 模型

下表列出輸入觸發和會話結束方式。VAD 類型、音訊格式及其他參數的完整取值請參閱對應模型的事件文件。

模型

輸入與觸發方式

結束方式

Qwen3.8-Omni / Qwen3.5-Omni

VAD 模式自動觸發回應;Manual 模式先傳送 input_audio_buffer.commit,再傳送 response.create。

接收完結果後關閉連線。

Qwen-TTS-Realtime

server_commit 模式自動提交文字;commit 模式傳送 input_text_buffer.commit,無需 response.create。

session.finish → session.finished

Qwen-ASR-Realtime

VAD 模式自動處理;Manual 模式傳送 input_audio_buffer.commit,無需 response.create。

session.finish → session.finished

Qwen-Audio-Realtime

自動模式由伺服器判斷輪次;Manual 模式先傳送 input_audio_buffer.commit,再傳送 response.create。

接收完結果後關閉連線。

Qwen3.8-LiveTranslate

使用 output_modalities 設定輸出模態,透過 audio.input.turn_detection 設定輪次偵測。持續傳送音訊,輸入結束後傳送 session.finish。

session.finish → session.finished

Qwen3.5-LiveTranslate

使用 modalities 和 turn_detection 設定會話。Manual 模式傳送 input_audio_buffer.commit 後自動產生回應,無需 response.create。

session.finish → session.finished

主要輸出事件如下,具體返回的事件取決於設定的輸出模態。

模型

主要輸出事件

Qwen-Omni-Realtime

response.text.delta / response.audio_transcript.delta / response.audio.delta / response.done

Qwen-TTS-Realtime

response.audio.delta / response.done

Qwen-ASR-Realtime

conversation.item.input_audio_transcription.text / conversation.item.input_audio_transcription.completed

Qwen-Audio-Realtime

response.audio_transcript.delta / response.audio.delta / response.done

Qwen3.8-LiveTranslate

response.text.delta / response.audio_transcript.delta / response.audio.delta / response.done

Qwen3.5-LiveTranslate

response.text.text / response.audio_transcript.text / response.audio.delta / response.done

說明翻譯模型的事件名稱與版本有關:Qwen3.8-LiveTranslate 使用 .delta,Qwen3.5-LiveTranslate 使用 .text。例如,音訊對應的譯文分別透過 response.audio_transcript.delta 和 response.audio_transcript.text 返回。

Inference 模型

Qwen-Audio-TTS/CosyVoice 使用 run-task → task-started → continue-task → finish-task → task-finished 的任務流程;Qwen-Audio-ASR/Fun-ASR/Paraformer 語音辨識在 task-started 後傳送二進位音訊幀。兩類模型均透過 task-failed 報告任務錯誤。

各模型詳細接入方式

即時全模態

即時語音合成

Qwen-Audio-TTS

CosyVoice

Qwen-TTS-Realtime

Sambert

即時語音識別

Qwen-Audio-ASR-Message

Qwen-Audio-ASR-Streaming

Fun-ASR-Realtime

Qwen-ASR-Realtime

Paraformer

即時語音對話

Qwen-Audio-Realtime

即時音視訊翻譯

Qwen-Livetranslate-Realtime

錯誤處理

  • 握手失敗:根據 HTTP 狀態碼和錯誤訊息檢查連線位址、API Key、業務空間及模型權限。遇到 401/403 時,優先核對驗證設定。
  • 模型不存在或未開通:核對模型名稱、服務開通狀態及呼叫地區。
  • Realtime 錯誤:處理 error 事件,讀取錯誤類型和訊息。根據錯誤訊息調整請求中的事件類型或參數。
  • Inference 錯誤:收到 task-failed 表示任務失敗,可透過 header.error_code 和 header.error_message 取得失敗原因。

修正設定或處理網路中斷後,重新建立連線,並按對應協定重新設定會話或建立任務。錯誤碼說明請參閱錯誤碼。