本文檔介紹如何使用 DashScope Python SDK 調用即時語音辨識(Qwen-ASR-Realtime)模型。
重要阿里雲百鍊為華北2(北京)、新加坡地區推出了業務空間專屬網域名稱,能夠為推理請求提供卓越的效能和更高的穩定性,建議遷移至新網域名稱:
- 華北2(北京)地區:從
dashscope.aliyuncs.com遷移至{WorkspaceId}.cn-beijing.maas.aliyuncs.com - 新加坡地區:從
dashscope-intl.aliyuncs.com遷移至{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com
{WorkspaceId}需要替換為真實的Workspace ID。現有網域名稱仍可正常使用。
前提條件
- 安裝SDK,確保DashScope SDK版本不低於1.25.6。
- 擷取API Key。
- 瞭解WebSocket API。
完整範例
說明範例程式碼讀取 your_audio_file.pcm(PCM16、16 kHz、單聲道)。如僅有 MP3/WAV 等格式,可使用 ffmpeg 轉換:
ffmpeg -i your_audio.mp3 -ar 16000 -ac 1 -f s16le your_audio_file.pcm
import logging
import os
import base64
import signal
import sys
import time
import dashscope
from dashscope.audio.qwen_omni import *
from dashscope.audio.qwen_omni.omni_realtime import TranscriptionParams
def setup_logging():
"""配置日誌輸出"""
logger = logging.getLogger('dashscope')
logger.setLevel(logging.DEBUG)
handler = logging.StreamHandler(sys.stdout)
handler.setLevel(logging.DEBUG)
formatter = logging.Formatter('%(asctime)s - %(name)s - %(levelname)s - %(message)s')
handler.setFormatter(formatter)
logger.addHandler(handler)
logger.propagate = False
return logger
def init_api_key():
"""初始化 API Key"""
# 新加坡和北京地區的API Key不同。擷取API Key:https://www.alibabacloud.com/help/zh/model-studio/get-api-key
# 若沒有配置環境變數,請用阿里雲百鍊API Key將下行替換為:dashscope.api_key = "sk-xxx"
dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY', 'YOUR_API_KEY')
if dashscope.api_key == 'YOUR_API_KEY':
print('[Warning] Using placeholder API key, set DASHSCOPE_API_KEY environment variable.')
class MyCallback(OmniRealtimeCallback):
"""即時識別回調處理"""
def __init__(self, conversation):
self.conversation = conversation
self.handlers = {
'session.created': self._handle_session_created,
'conversation.item.input_audio_transcription.completed': self._handle_final_text,
'conversation.item.input_audio_transcription.text': self._handle_transcription_text,
'input_audio_buffer.speech_started': lambda r: print('======Speech Start======'),
'input_audio_buffer.speech_stopped': lambda r: print('======Speech Stop======')
}
def on_open(self):
print('Connection opened')
def on_close(self, code, msg):
print(f'Connection closed, code: {code}, msg: {msg}')
def on_event(self, response):
try:
handler = self.handlers.get(response['type'])
if handler:
handler(response)
except Exception as e:
print(f'[Error] {e}')
def _handle_session_created(self, response):
print(f"Start session: {response['session']['id']}")
def _handle_final_text(self, response):
print(f"Final recognized text: {response['transcript']}")
def _handle_transcription_text(self, response):
print(f"Got transcription result: {response['text'] + response['stash']}")
def read_audio_chunks(file_path, chunk_size=3200):
"""按塊讀取音頻檔案"""
with open(file_path, 'rb') as f:
while chunk := f.read(chunk_size):
yield chunk
def send_audio(conversation, file_path, delay=0.1):
"""發送音頻資料"""
if not os.path.exists(file_path):
raise FileNotFoundError(f"Audio file {file_path} does not exist.")
print("Processing audio file... Press 'Ctrl+C' to stop.")
for chunk in read_audio_chunks(file_path):
audio_b64 = base64.b64encode(chunk).decode('ascii')
conversation.append_audio(audio_b64)
time.sleep(delay)
def main():
setup_logging()
init_api_key()
audio_file_path = "./your_audio_file.pcm"
callback = MyCallback(conversation=None)
conversation = OmniRealtimeConversation(
model='qwen3-asr-flash-realtime',
# 以下為新加坡地區的配置,調用時請將"{WorkspaceId}"替換為真實的業務空間ID,各地區的配置不同。
url='wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime',
callback=callback,
)
callback.conversation = conversation # 把 conversation 注入回調,用於回調中調用其方法
def handle_exit(sig, frame):
print('Ctrl+C pressed, exiting...')
conversation.close()
sys.exit(0)
signal.signal(signal.SIGINT, handle_exit)
conversation.connect()
transcription_params = TranscriptionParams(
language='zh',
sample_rate=16000,
input_audio_format="pcm"
)
conversation.update_session(
output_modalities=[MultiModality.TEXT],
enable_input_audio_transcription=True,
transcription_params=transcription_params
)
try:
send_audio(conversation, audio_file_path)
# send session.finish and wait for finished and close
conversation.end_session()
except Exception as e:
print(f"Error occurred: {e}")
finally:
conversation.close()
print("Audio processing completed.")
if __name__ == '__main__':
main()
請求參數
-
以下參數通過
OmniRealtimeConversation的構造方法設定。點擊查看範例程式碼
class MyCallback(OmniRealtimeCallback): """即時識別回調處理""" def __init__(self, conversation): self.conversation = conversation self.handlers = { 'session.created': self._handle_session_created, 'conversation.item.input_audio_transcription.completed': self._handle_final_text, 'conversation.item.input_audio_transcription.text': self._handle_stash_text, 'input_audio_buffer.speech_started': lambda r: print('======Speech Start======'), 'input_audio_buffer.speech_stopped': lambda r: print('======Speech Stop======') } def on_open(self): print('Connection opened') def on_close(self, code, msg): print(f'Connection closed, code: {code}, msg: {msg}') def on_event(self, response): try: handler = self.handlers.get(response['type']) if handler: handler(response) except Exception as e: print(f'[Error] {e}') def _handle_session_created(self, response): print(f"Start session: {response['session']['id']}") def _handle_final_text(self, response): print(f"Final recognized text: {response['transcript']}") def _handle_stash_text(self, response): print(f"Got stash result: {response['stash']}") conversation = OmniRealtimeConversation( model='qwen3-asr-flash-realtime', # 以下為華北2(北京)地區的配置,調用時請將"{WorkspaceId}"替換為真實的業務空間ID,各地區的配置不同。 url='wss://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api-ws/v1/realtime', callback=MyCallback(conversation=None) # 暫時傳None,稍後注入 ) # 注入自身到回調 conversation.callback.conversation = conversation參數
類型
是否必須
說明
modelstr是
指定要使用的模型名稱。
callbackOmniRealtimeCallback是
用於處理服務端事件的回調對象執行個體。
urlstr是
語音辨識服務地址:
華北2(北京)地區:
wss://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api-ws/v1/realtime。調用時請將{WorkspaceId}替換為真實的Workspace ID。新加坡地區:
wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime。調用時請將{WorkspaceId}替換為真實的業務空間ID。
-
以下參數通過
OmniRealtimeConversation的update_session方法設定。點擊查看範例程式碼
transcription_params = TranscriptionParams( language='zh', sample_rate=16000, input_audio_format="pcm" ) conversation.update_session( output_modalities=[MultiModality.TEXT], enable_turn_detection=True, turn_detection_type="server_vad", turn_detection_threshold=0.0, turn_detection_silence_duration_ms=400, enable_input_audio_transcription=True, transcription_params=transcription_params )參數
類型
是否必須
說明
output_modalitiesList[MultiModality]是
模型輸出模態,固定為
[MultiModality.TEXT]。enable_turn_detectionbool否
是否開啟服務端語音活動檢測(VAD)。關閉後,需手動調用
commit()方法觸發識別。預設值:
True。取值範圍:
True:開啟False:關閉
turn_detection_typestr否
服務端VAD類型,固定為
server_vad。turn_detection_thresholdfloat否
VAD檢測閾值。推薦將該值設為
0.0。預設值:
0.2。取值範圍:
[-1, 1]。較低的閾值會提高 VAD 的靈敏度,可能將背景雜音誤判為語音。較高的閾值則降低靈敏度,有助於在嘈雜環境中減少誤觸發。
turn_detection_silence_duration_msint否
VAD斷句檢測閾值(ms)。靜音持續時間長度超過該閾值將被認為是語句結束。推薦將該值設為
400。預設值:
800。取值範圍:
[200, 6000]。較低的值(如 300ms)可使模型更快響應,但可能導致在自然停頓處發生不合理的斷句。較高的值(如 1200ms)可更好地處理長句內的停頓,但會增加整體響應延遲。
transcription_paramsTranscriptionParams否
語音辨識相關配置。
-
以下參數通過
TranscriptionParams的構造方法設定。點擊查看範例程式碼
transcription_params = TranscriptionParams( language='zh', sample_rate=16000, input_audio_format="pcm" )參數
類型
是否必須
說明
languagestr否
音頻源語言。
zh:中文(普通話、四川話、閩南語、吳語)
yue:粵語
en:英文
ja:日語
de:德語
ko:韓語
ru:俄語
fr:法語
pt:葡萄牙語
ar:阿拉伯語
it:意大利語
es:西班牙語
hi:印地語
id:印尼語
th:泰語
tr:土耳其語
uk:烏克蘭語
vi:越南語
cs:捷克語
da:丹麥語
fil:菲律賓語
fi:芬蘭語
is:冰島語
ms:馬來語
no:挪威語
pl:波蘭語
sv:瑞典語
sample_rateint否
音頻採樣率(Hz)。支援
16000和8000。預設值:
16000。設定為
8000時,服務端會先升採樣到16000Hz再進行識別,可能引入微小延遲。建議僅在源音頻為8000Hz(如電話線路)時使用。input_audio_formatstr否
音頻格式。支援
pcm和opus。預設值:
pcm。
關鍵介面
OmniRealtimeConversation類
OmniRealtimeConversation通過from dashscope.audio.qwen_omni import OmniRealtimeConversation方法引入。
| 方法簽名 | 服務端響應事件(通過回調下發) | 說明 |
|---|---|---|
|
| 和服務端建立串連。 |
|
| 用於更新會話配置,建議在串連建立後首先調用該方法進行設定。若未調用該方法,系統將使用預設配置。只需關注請求參數中的涉及到的參數。 |
| 無 | 將Base64編碼後的音頻資料片段追加到雲端輸入音頻緩衝區。 |
|
| 提交之前通過append添加到雲端緩衝區的音視頻,如果輸入的音頻緩衝區為空白將產生錯誤。 禁用情境:請求參數 |
|
| 通知服務端結束會話,服務端收到會話結束通知後將完成最後的語音辨識。 調用時機:
|
| 無 | 終止任務,並關閉串連。 |
| 無 | 擷取當前任務的session_id。 |
| 無 | 擷取最近一次response的response_id。 |
回調介面(OmniRealtimeCallback)
服務端會通過回調的方式,將服務端響應事件和資料返回給用戶端。
繼承此類並實現相應方法以處理服務端事件。
通過from dashscope.audio.qwen_omni import OmniRealtimeCallback引入。
| 方法簽名 | 參數 | 說明 |
|---|---|---|
| 無 | WebSocket串連成功建立時觸發。 |
| message:服務端事件 | 收到服務端事件時觸發。 |
| close_status_code:狀態代碼 close_msg:WebSocket串連關閉時的日誌資訊 | WebSocket串連關閉時觸發。 |