This topic describes the parameters and interface details of the HTTP API for non-real-time speech recognition with Qwen-Audio-3.x-ASR-Flash-Filetrans and Fun-ASR.
User guide:Non-real-time speech recognition. For input requirements such as supported audio formats, file size limits, and duration limits, see Audio specifications.
How it works
Unlike synchronous DashScope calls, which return the result immediately in a single request, asynchronous calls are designed for long audio files or time-consuming tasks. This mode uses a two-step submit-and-poll flow that avoids request timeouts caused by long waits:
-
Step 1: Submit the task.
- The client sends an asynchronous processing request.
- After validating the request, the server does not run the task immediately. Instead, it returns a unique
task_idto indicate that the task was created successfully.
-
Step 2: Retrieve the result.
- The client uses the returned
task_idto poll the query interface repeatedly. - When the task finishes, the query interface returns the final recognition result.
- The client uses the returned
Service endpoints
Singapore
Submit task interface: POST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/asr/transcription
Query task interface: GET https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/tasks/{task_id}
Replace {WorkspaceId} with your actual Workspace ID.
China (Beijing)
Submit task interface: POST https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/audio/asr/transcription
Query task interface: GET https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/tasks/{task_id}
Replace {WorkspaceId} with your actual Workspace ID.
ImportantAlibaba Cloud Model Studio has released workspace-specific domains for the China (Beijing) and Singapore regions. The new dedicated domains deliver superior performance and higher stability for inference requests. We recommend migrating to the new domains:
- China (Beijing): from
dashscope.aliyuncs.comto{WorkspaceId}.cn-beijing.maas.aliyuncs.com - Singapore: from
dashscope-intl.aliyuncs.comto{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com
Replace {WorkspaceId} with your actual Workspace ID. The existing domains remain fully functional.
ImportantWhen you submit a task with the new domain (https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com), the request body must include the parameters object. Even if you don't need to set any parameters, pass an empty object {}. Otherwise, the task is submitted successfully but recognition fails.
Request headers
Parameter | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Authentication token in the format |
Content-Type | string | Yes | The media type of the request body. Required only for the submit task interface. Fixed value: |
X-DashScope-Async | string | Yes | The asynchronous task flag. Required only for the submit task interface. Fixed value: |
Submit task interface
Submits a speech recognition task. This interface returns asynchronously, so poll the task status with the Query task interface.
Request bodymodel The model name. Supported values include the Qwen-Audio-3.x-ASR-Flash-Filetrans and Fun-ASR model families. For details, see Supported models and regions. input The input parameter object. parameters The request parameter object. ImportantWhen you call When you use the new domain ( | Basic callThe following example uses the Singapore region. Replace "{WorkspaceId}" with your actual workspace ID. The configuration differs by region. The Singapore region and the Beijing region use different API keys. Inline hotwordsUse inline hotwords in the following format: ContextUse context in the following format: |
Response bodyrequest_id The unique identifier of this call. output The data returned by the submit task interface. | |
Query task interface
Queries the execution status and result of a speech recognition task. Poll this interface until the task reaches a terminal state.
Request bodytask_id ImportantThis parameter is a URL path parameter. There is no request body. To query a task, specify its ID. This ID is the | The following example uses the Singapore region. Replace "{WorkspaceId}" with your actual workspace ID. The configuration differs by region. The Singapore region and the Beijing region use different API keys. |
Response bodyrequest_id The unique identifier of this call. output The data returned by the query task interface. | |
Other interfaces: batch-query task status / cancel a task
For details, see Manage asynchronous tasks: you can batch-query non-real-time speech recognition tasks submitted within the last 24 hours, and cancel tasks in the PENDING (queued) state.
Recognition result description
The recognition result is saved as a JSON file.
Click to view the recognition result example
{
"file_url":"{YOUR_AUDIO_URL}",
"properties":{
"audio_format":"pcm_s16le",
"channels":[
0
],
"original_sampling_rate":16000,
"original_duration_in_milliseconds":3834
},
"transcripts":[
{
"channel_id":0,
"content_duration_in_milliseconds":3720,
"text":"Hello world, this is Alibaba Speech Lab.",
"sentences":[
{
"begin_time":100,
"end_time":3820,
"text":"Hello world, this is Alibaba Speech Lab.",
"sentence_id":1,
"speaker_id":0, //This field is displayed only when automatic speaker diarization is enabled
"words":[
{
"begin_time":100,
"end_time":596,
"text":"Hello ",
"punctuation":""
},
{
"begin_time":596,
"end_time":844,
"text":"world",
"punctuation":", "
}
// Other content is omitted here
]
}
]
}
]
}
The following parameters are worth noting:
Parameter | Type | Description |
audio_format | string | The audio format of the source file. |
channels | array[integer] | The track index of the audio in the source file. For single-track audio, [0] is returned; for dual-track audio, [0, 1] is returned; and so on. |
original_sampling_rate | integer | The sampling rate (Hz) of the audio in the source file. |
original_duration_in_milliseconds | integer | The original audio duration (ms) in the source file. |
channel_id | integer | The track index of the transcription result, starting from 0. |
content_duration | integer | The duration (ms) of content in the track that is identified as speech. The speech recognition model service transcribes only the content in a track that is identified as speech, and meters and bills based on that duration. Non-speech content is not metered or billed. Typically, the speech content duration is shorter than the original audio duration. Because whether speech content exists is determined by an AI model, the result may differ slightly from the actual situation. |
transcript | string | The paragraph-level transcription result. |
sentences | array | The sentence-level transcription result. |
words | array | The word-level transcription result. |
begin_time | integer | The start timestamp (ms). |
end_time | integer | The end timestamp (ms). |
text | string | The transcription result. |
speaker_id | integer | The index of the current speaker, starting from 0, used to distinguish between different speakers. This field appears in the recognition result only when speaker diarization is enabled. |
punctuation | string | The punctuation predicted after the word, if any. |