This topic describes the parameters and interfaces of the Python SDK for the Qwen-Audio-3.1-ASR-Flash-Message real-time speech recognition model.
ImportantAlibaba Cloud Model Studio has released workspace-specific domains for the China (Beijing) and Singapore regions. The new dedicated domains deliver superior performance and higher stability for inference requests. We recommend migrating to the new domains:
- China (Beijing): from
dashscope.aliyuncs.comto{WorkspaceId}.cn-beijing.maas.aliyuncs.com - Singapore: from
dashscope-intl.aliyuncs.comto{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com
Prerequisites
- The service is activated and Obtain an API key. To guard against security risks caused by code leaks, Configure API key as an environment variable, rather than hard-coding it in your code.
- Install the latest DashScope SDK.
Quick start
The Recognition class provides interfaces for both non-streaming and bidirectional streaming calls. Choose the call method that fits your needs:
- Non-streaming call: Recognizes a local file and returns the complete result in a single response. Suitable for processing pre-recorded audio.
- Bidirectional streaming call: Recognizes an audio stream directly and outputs results in real time. The audio stream can come from an external device, such as a microphone, or be read from a local file. Suitable for scenarios that require immediate feedback.
Non-streaming call
Submit a single real-time speech recognition task and get the recognition result synchronously by passing in a local file.
Instantiate a Recognition class, bind a Request parameters, and call call to run recognition or translation and get the final Recognition result (RecognitionResult).
from http import HTTPStatus
import dashscope
from dashscope.audio.asr import Recognition
import os
# The API Key differs between the Singapore and Beijing regions. Get an API Key: https://www.alibabacloud.com/help/model-studio/get-api-key
# If you have not configured the environment variable, replace the following line with your Model Studio API Key: dashscope.api_key = "sk-xxx"
dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY')
# The following is the configuration for the Singapore region. When calling, replace "{WorkspaceId}" with your actual workspace ID. The configuration differs by region.
dashscope.base_websocket_api_url='wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inference'
recognition = Recognition(model='qwen-audio-3.1-asr-flash-message',
format='wav',
sample_rate=16000,
callback=None)
result = recognition.call('{YOUR_AUDIO_FILE}')
if result.status_code == HTTPStatus.OK:
print('Recognition result:')
print(result.get_sentence())
else:
print('Error: ', result.message)
print(
'[Metric] requestId: {}, first package delay ms: {}, last package delay ms: {}'
.format(
recognition.get_last_request_id(),
recognition.get_first_package_delay(),
recognition.get_last_package_delay(),
))
Bidirectional streaming call
Submit a single real-time speech recognition task and stream the real-time recognition results by implementing the callback interface.
-
Start streaming speech recognition.
Instantiate a Recognition class, bind a Request parameters and a Callback interface (RecognitionCallback), and call the
startmethod to start streaming speech recognition. -
Stream the audio.
Call the
send_audio_framemethod of the Recognition class in a loop to send the binary audio stream to the server in segments. The stream is read from a local file or a device, such as a microphone.While the audio is being sent, the server returns recognition results to the client in real time through the
on_eventmethod of the Callback interface (RecognitionCallback).Send about 100 ms of audio per frame, and keep each frame between 1 KB and 16 KB.
-
End the task.
Call the
stopmethod of the Recognition class to end speech recognition.This method blocks the current thread until the
on_completeoron_errorcallback of the Callback interface (RecognitionCallback) is triggered.
import os
import signal # for keyboard events handling (press "Ctrl+C" to terminate recording)
import sys
import dashscope
import pyaudio
from dashscope.audio.asr import *
mic = None
stream = None
# Set recording parameters
sample_rate = 16000 # sampling rate (Hz)
channels = 1 # mono channel
dtype = 'int16' # data type
format_pcm = 'pcm' # the format of the audio data
block_size = 3200 # number of frames per buffer
# Real-time speech recognition callback
class Callback(RecognitionCallback):
def on_open(self) -> None:
global mic
global stream
print('RecognitionCallback open.')
mic = pyaudio.PyAudio()
stream = mic.open(format=pyaudio.paInt16,
channels=1,
rate=16000,
input=True)
def on_close(self) -> None:
global mic
global stream
print('RecognitionCallback close.')
stream.stop_stream()
stream.close()
mic.terminate()
stream = None
mic = None
def on_complete(self) -> None:
print('RecognitionCallback completed.') # recognition completed
def on_error(self, message) -> None:
print('RecognitionCallback task_id: ', message.request_id)
print('RecognitionCallback error: ', message.message)
# Stop and close the audio stream if it is running
if 'stream' in globals() and stream.active:
stream.stop()
stream.close()
# Forcefully exit the program
sys.exit(1)
def on_event(self, result: RecognitionResult) -> None:
sentence = result.get_sentence()
if 'text' in sentence:
print('RecognitionCallback text: ', sentence['text'])
if RecognitionResult.is_sentence_end(sentence):
print(
'RecognitionCallback sentence end, request_id:%s, usage:%s'
% (result.get_request_id(), result.get_usage(sentence)))
def signal_handler(sig, frame):
print('Ctrl+C pressed, stop recognition ...')
# Stop recognition
recognition.stop()
print('Recognition stopped.')
print(
'[Metric] requestId: {}, first package delay ms: {}, last package delay ms: {}'
.format(
recognition.get_last_request_id(),
recognition.get_first_package_delay(),
recognition.get_last_package_delay(),
))
# Forcefully exit the program
sys.exit(0)
# main function
if __name__ == '__main__':
# The API Key differs between the Singapore and Beijing regions. Get an API Key: https://www.alibabacloud.com/help/model-studio/get-api-key
# If you have not configured the environment variable, replace the following line with your Model Studio API Key: dashscope.api_key = "sk-xxx"
dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY')
# The following is the configuration for the Singapore region. When calling, replace "{WorkspaceId}" with your actual workspace ID. The configuration differs by region.
dashscope.base_websocket_api_url='wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inference'
# Create the recognition callback
callback = Callback()
# Call recognition service by async mode, you can customize the recognition parameters, like model, format,
# sample_rate
recognition = Recognition(
model='qwen-audio-3.1-asr-flash-message',
format=format_pcm,
# 'pcm'、'wav'、'opus'、'speex'、'aac'、'amr', you can check the supported formats in the document
sample_rate=sample_rate,
# only supports 16000 Hz
callback=callback)
# Start recognition
recognition.start()
signal.signal(signal.SIGINT, signal_handler)
print("Press 'Ctrl+C' to stop recording and recognition...")
# Create a keyboard listener until "Ctrl+C" is pressed
while True:
if stream:
data = stream.read(3200, exception_on_overflow=False)
recognition.send_audio_frame(data)
else:
break
recognition.stop()
import os
import time
import dashscope
from dashscope.audio.asr import *
# The API Key differs between the Singapore and Beijing regions. Get an API Key: https://www.alibabacloud.com/help/model-studio/get-api-key
# If you have not configured the environment variable, replace the following line with your Model Studio API Key: dashscope.api_key = "sk-xxx"
dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY')
# The following is the configuration for the Singapore region. When calling, replace "{WorkspaceId}" with your actual workspace ID. The configuration differs by region.
dashscope.base_websocket_api_url='wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inference'
from datetime import datetime
def get_timestamp():
now = datetime.now()
formatted_timestamp = now.strftime("[%Y-%m-%d %H:%M:%S.%f]")
return formatted_timestamp
class Callback(RecognitionCallback):
def on_complete(self) -> None:
print(get_timestamp() + ' Recognition completed') # recognition complete
def on_error(self, message: str) -> None:
print('Error occurred: ', message)
print('Error details: ', message)
exit(0)
def on_event(self, result: RecognitionResult) -> None:
sentence = result.get_sentence()
if 'text' in sentence:
print(get_timestamp() + ' RecognitionCallback text: ', sentence['text'])
if RecognitionResult.is_sentence_end(sentence):
print(get_timestamp() +
'RecognitionCallback sentence end, request_id:%s, usage:%s'
% (result.get_request_id(), result.get_usage(sentence)))
callback = Callback()
recognition = Recognition(model='qwen-audio-3.1-asr-flash-message',
format='wav',
sample_rate=16000,
callback=callback)
try:
audio_data: bytes = None
f = open("{YOUR_AUDIO_FILE}", 'rb')
if os.path.getsize("{YOUR_AUDIO_FILE}"):
# Read all the file data into the buffer at once
file_buffer = f.read()
f.close()
print("Start Recognition")
recognition.start()
# Send 3200 bytes from the buffer at a time
buffer_size = len(file_buffer)
offset = 0
chunk_size = 3200
while offset < buffer_size:
# Calculate the size of the data chunk to send this time
remaining_bytes = buffer_size - offset
current_chunk_size = min(chunk_size, remaining_bytes)
# Extract the current data chunk from the buffer
audio_data = file_buffer[offset:offset + current_chunk_size]
# Send the audio data frame
recognition.send_audio_frame(audio_data)
# Update the offset
offset += current_chunk_size
# Add a delay to simulate real-time transmission
time.sleep(0.1)
recognition.stop()
else:
raise Exception(
'The supplied file was empty (zero bytes long)')
except Exception as e:
raise e
print(
'[Metric] requestId: {}, first package delay ms: {}, last package delay ms: {}'
.format(
recognition.get_last_request_id(),
recognition.get_first_package_delay(),
recognition.get_last_package_delay(),
))
Request parameters
Set request parameters through the constructor (init) of the Recognition class.
| Parameter | Type | Required | Description |
|---|---|---|---|
model | str | Yes | The model name. Set to qwen-audio-3.1-asr-flash-message. |
sample_rate | int | Yes | The sample rate, in Hz. Only |
format | str | Yes | The audio format. Valid values:
Importantopus/speex: Must use Ogg encapsulation. wav: Must use PCM encoding. amr: Only the AMR-NB type is supported. |
disfluency_removal_enabled | bool | No | Whether to filter filler words and polish the output. Defaults to false. Set to true to enable this feature. Pass as a keyword argument with the same name. |
intermediate_result_enabled | bool | No | Whether to return intermediate streaming results. Defaults to false. Set to true to return intermediate streaming results. Pass as a keyword argument with the same name. |
keep_dialect | bool | No | Defaults to false, which transcribes dialects into standard Chinese. Set to true to preserve dialect expressions. Pass as a keyword argument with the same name. See Client events for details. |
vad_model | str | No | Set to near_meeting_16k (near-field) or far_field_meeting_16k (far-field, default). Pass as a keyword argument with the same name. See Client events for details. |
vocabulary_id | str | No | The ID of a precompiled hot word list. Generate this ID in advance by calling the create hot word list API. Pass the ID during recognition to use the hot words in the list. Suitable for scenarios where the vocabulary is known and relatively stable, and where you need to reuse the same word list across requests. For usage details, see Precompiled hotwords. |
vocabulary | dict | No | Instant hot words. Passed as key-value pairs, where the key is the hot word text ( Suitable for temporary, session-level hot word optimization. When instant and precompiled hotwords are configured together, the system merges both sets. If the merged set contains more than 2000 hotwords, the system randomly selects 2000 to use. For usage details, see Instant hotwords. Example: |
max_sentence_silence | int | No | The VAD silence threshold for segmentation, in ms. When the silence after a segment of speech exceeds this threshold, the system determines that the sentence has ended. Default value: 1300. Valid values: [200, 6000]. |
heartbeat | bool | No | Whether to enable heartbeat packets. Default: False.
Silent audio refers to content in an audio file or data stream that contains no sound signal. You can generate silent audio in several ways, such as using audio editing software like Audacity or Adobe Audition, or using a command-line tool like FFmpeg. This field requires SDK version 1.23.1 or later. |
speech_noise_threshold | float | No | The threshold for distinguishing speech from noise, used to adjust the sensitivity of Voice Activity Detection (VAD). Valid values: [-1.0, 1.0]. Value descriptions:
This is an advanced configuration parameter. Adjusting it can significantly affect recognition results. Recommendations:
|
callback | RecognitionCallback | No | Callback interface (RecognitionCallback). |
Pass the following parameters as keyword arguments to the call or start method of the Recognition instance.
| Parameter | Type | Required | Description |
|---|---|---|---|
raw_input | dict | No | The input object used to pass in the conversation context. Context enhancement improves recognition accuracy for domain-specific terms. For usage, see Quick start. The dict must include a
ImportantContext messages of the ImportantWhen you pass context, the messages in NoteThis field requires SDK version 1.25.23 or later. |
Key interfaces
Recognition class
Import Recognition with from dashscope.audio.asr import *.
| Member method | Method signature | Description |
|---|---|---|
call | | A non-streaming call based on a local file. This method blocks the current thread until all audio is read, and requires read permission on the file. The recognition result is returned as a |
start | | Starts speech recognition. A callback-based streaming real-time recognition. This method does not block the current thread. Use it together with |
send_audio_frame | | Pushes audio. Keep each pushed audio frame neither too large nor too small: about 100 ms per frame, between 1 KB and 16 KB. Recognition results are obtained through the on_event method of the Callback interface (RecognitionCallback). |
stop | | Stops speech recognition. Blocks until the server finishes recognizing all received audio, then ends the task. |
get_last_request_id | | Gets the request_id. Available after the constructor is called (the object is created). |
get_first_package_delay | | Gets the first-packet delay: the latency from sending the first audio packet to receiving the first recognition result. Use it after the task completes. |
get_last_package_delay | | Gets the last-packet delay: the time from sending the stop command to receiving the last recognition result. Use it after the task completes. |
get_response | | Gets the last message. Use it to retrieve a task-failed error. |
Update conversation context
Call update_context to update conversation context while a recognition task is running. The context is used to assist recognition of subsequent audio. This method requires DashScope Python SDK 1.27.5 or later.
def update_context(self, payload_input: dict)
- When to call: After starting streaming recognition with
startand before callingstop. - Parameter:
payload_inputis a dictionary containing thepayload.inputobject of acontinue-taskevent, including thecontextfield. Do not add an extrapayloadorinputwrapper. - Supported models and parameter constraints: See continue-task.
The following example uses an existing, started recognition instance.
payload_input = {
"context": [
{
"role": "user",
"content": [{"type": "input_text", "text": "Hello"}]
},
{
"role": "assistant",
"content": [{"type": "text", "text": "Hello, I am Qwen. How can I help you?"}]
}
]
}
recognition.update_context(payload_input=payload_input)
Callback interface (RecognitionCallback)
During a Bidirectional streaming call, the server returns key process information and data to the client through callbacks. Implement the callback methods to handle the information and data returned by the server.
class Callback(RecognitionCallback):
def on_open(self) -> None:
print('Connection established')
def on_event(self, result: RecognitionResult) -> None:
# Implement the logic to receive recognition results
pass
def on_complete(self) -> None:
print('Task completed')
def on_error(self, message: str) -> None:
print('An error occurred:', message)
def on_close(self) -> None:
print('Connection closed')
callback = Callback()
| Method | Parameter | Return value | Description |
|---|---|---|---|
| None | None | Called immediately after the connection to the server is established. |
| result: Recognition result (RecognitionResult) | None | Called when the server has a response. |
| None | None | Called after all recognition results are returned. |
| result: Recognition result (RecognitionResult) | None | Called when an error occurs. |
| None | None | Called after the server closes the connection. |
Response
Recognition result (RecognitionResult)
RecognitionResult represents the result of a single real-time recognition in a Bidirectional streaming call, or the result of a Non-streaming call.
| Member method | Method signature | Description |
|---|---|---|
get_sentence | | Gets the current recognized sentence and its timestamp information. A callback returns a single sentence, so this method returns Dict[str, Any]. For details, see Sentence (Sentence). |
get_request_id | | Gets the request_id of the request. |
is_sentence_end | | Determines whether the given sentence has ended. This method checks whether the end_time field in sentence is None: a non-None end_time indicates that the sentence has ended. Call it as RecognitionResult.is_sentence_end(sentence), where sentence is the single-sentence dict returned by get_sentence(), not a boolean field on a Sentence instance. |
Sentence information (Sentence)
The members of the Sentence class are as follows:
| Parameter | Type | Description |
|---|---|---|
begin_time | int | Sentence start time, in ms. |
end_time | int | Sentence end time, in ms. |
text | str | Recognized text. |
words | A list of Word-level timestamp information (Word) | Word-level timestamp information. |
Word-level timestamp information (Word)
The members of the Word class are as follows:
| Parameter | Type | Description |
|---|---|---|
begin_time | int | Word start time, in ms. |
end_time | int | Word end time, in ms. |
text | str | The word. |
punctuation | str | The punctuation. |
Error codes
If you encounter errors, see Error codes for troubleshooting.
If the issue persists, join the developer community listed in the speech SDK sample repository to report your issue and provide the Request ID for further investigation.
FAQ
Features
Q: How do I keep the connection alive during long periods of silence?
Set the heartbeat request parameter to true, and keep sending silent audio to the server.
Silent audio is content in an audio file or stream that contains no sound signal. You can generate silent audio in several ways, such as using audio editing software like Audacity or Adobe Audition, or command-line tools like FFmpeg.
Q: How do I convert audio to a supported format?
Use FFmpeg. For more usage, see the FFmpeg official website.
# Basic conversion command (all-purpose template)
# -i, purpose: input file path, example value: audio.wav
# -c:a, purpose: audio codec, example values: aac, libmp3lame, pcm_s16le
# -b:a, purpose: bitrate (audio quality control), example values: 192k, 320k
# -ar: sample rate; set to 16000 for this model
# -ac, purpose: number of channels, example values: 1 (mono), 2 (stereo)
# -y, purpose: overwrite an existing file (no value needed)
ffmpeg -i input_audio.ext -c:a codec_name -b:a bitrate -ar sample_rate -ac channels output.ext
# For example: WAV to MP3 (keep the original quality)
ffmpeg -i input.wav -c:a libmp3lame -q:a 0 -ar 16000 -ac 1 output.mp3
# For example: MP3 to WAV (16-bit PCM standard format)
ffmpeg -i input.mp3 -c:a pcm_s16le -ar 16000 -ac 1 output.wav
# For example: M4A to AAC (extract or convert Apple audio)
ffmpeg -i input.m4a -c:a copy output.aac # Only if the source is already 16000 Hz mono AAC
ffmpeg -i input.m4a -c:a aac -b:a 64k -ar 16000 -ac 1 output.aac # Re-encode as 16000 Hz mono audio
# For example: FLAC lossless to Opus (high compression)
ffmpeg -i input.flac -c:a libopus -b:a 128k -vbr on -ar 16000 -ac 1 output.opus
Q: How do I recognize a local file (recording)?
There are two ways to recognize a local file:
-
Pass the local file path directly: This way returns the complete result only after recognition finishes, so it isn't suitable for scenarios that need immediate feedback.
See Non-streaming call, and pass the file path to the
callmethod of the Recognition class to recognize the recording directly. -
Convert the local file to a binary stream for recognition: This way recognizes the file while streaming results, so it suits scenarios that need immediate feedback.
See Bidirectional streaming call, and send the binary stream to the server for recognition through the
send_audio_framemethod of the Recognition class.
Troubleshooting
Q: Why can't the speech be recognized (no recognition result)?
-
Check that the audio format (
format) and sample rate (sampleRate/sample_rate) in the request parameters are correct and meet the parameter constraints. Common errors include:- The audio file has a .wav extension but is actually in MP3 format, while the
formatrequest parameter is set to wav (incorrect parameter setting). - The audio sample rate is 3600 Hz, but the
sampleRate/sample_raterequest parameter is set to 48000 (incorrect parameter setting).
Use the ffprobe tool to get the container, codec, sample rate, channels, and other information of the audio:
ffprobe -v error -show_entries format=format_name -show_entries stream=codec_name,sample_rate,channels -of default=noprint_wrappers=1 input.xxx - The audio file has a .wav extension but is actually in MP3 format, while the
-
If none of the checks above reveal a problem, add custom hotwords to improve recognition of specific terms.