This guide shows you how to use the Qwen-Audio-ASR-Streaming real-time speech recognition iOS SDK to transcribe speech into text.
Quick start
-
Get an API key: Obtain an API key
-
Download the SDK and run the sample code:
- Download the latest SDK package.
- Extract the ZIP package and add the included nuisdk.xcframework to your project.
- In Build Phases → Link Binary With Libraries, add nuisdk.xcframework.
- In General → Frameworks, Libraries, and Embedded Content, set nuisdk.xcframework to Embed & Sign.
- Open the sample project in Xcode. The sample code is in
DashFunAsrSpeechTranscriberViewController.m. Replace the API key with your own to try out the feature.
Call sequence
- Initialize the SDK.
- Set parameters for your use case: use the
nui_initializeAPI to set Connection and control parameters, and use thenui_set_paramsAPI to set Recognition quality parameters. - Call
nui_dialog_startto start the recognition process. - In the
onNuiAudioStateChangedcallback, open the recording device based on the audio state. - In the
onNuiNeedAudioDatacallback, continuously supply recording data, or callnui_update_audio_datato actively push recording data. - In the
onNuiEventCallbackcallback, listen for events and retrieve speech recognition results. - Call
nui_dialog_cancelto stop recognition, and confirm that recognition has finished by listening for the EVENT_TRANSCRIBER_COMPLETE event. - When you no longer need the recognition feature, call
nui_releaseto release the SDK resources.
Request parameters
Connection and control parameters
Configure these parameters by passing a JSON string in the parameters argument of the nui_initialize API.
Example parameters: The following JSON string is an example and does not list all parameters. Add the parameters you need when you write your code:
{
"url": "wss://dashscope.aliyuncs.com/api-ws/v1/inference",
"apikey": "st-****",
"device_id": "my_device_id",
"service_mode": "1"
}
Parameter descriptions
| Parameter | Type | Required | Description |
|---|---|---|---|
url | String | Yes | Service address:
{WorkspaceId} with your actual Workspace ID. |
apikey | String | Yes | API key. |
service_mode | String | Yes | Run mode. Fixed to "1" for real-time speech recognition. |
device_id | String | Yes | A unique string that identifies the end user. You can set it to an in-app user ID or a device identifier generated on the client. This ID is used mainly for log tracing and troubleshooting. |
audio_update_manually | String | No | Whether to enable active audio data pushing. Default: "false".If set to "true" and the SDK version supports on-device audio processing capabilities such as AEC or VAD, those capabilities are enabled by default. |
workspace | String | No | The path where on-device resource files are stored. This parameter is required when audio_update_manually is set to "true" and an on-device audio processing capability such as AEC or VAD is enabled. |
debug_path | String | No | Storage path for the log file.This parameter takes effect only when save_log is set to YES in the nui_initialize API. In that case, you must set the log file path, otherwise an error occurs.At most two log files are kept locally. |
save_wav | String | No | Whether to save an audio file for debugging. The audio file is saved under debug_path.Default: "false".Valid values:
save_log is set to true in the nui_initialize API. In addition, debug_path must also be set. |
max_log_file_size | int | No | Maximum size of the log file, in bytes.This parameter takes effect only when save_log is set to YES in the nui_initialize API.Default: 104857600 (100 * 1024 * 1024 bytes, that is, 100 MiB). |
log_track_level | int | No | Filter level for the log content sent through the log callback (onNuiLogTrackCallback).Default: 2.Valid values:
log_track_level and level (set through the nui_initialize API) together determine which logs are ultimately sent to the callback. A log is sent to the callback only when its level value is greater than or equal to both log_track_level and level. For example, if log_track_level is set to 2 (INFO) and level is set to 3 (WARNING), then only logs at WARNING level or higher (value >= 3) are sent to the callback. |
Recognition quality parameters
Configure these parameters by passing a JSON string in the params argument of the nui_set_params API.
Example parameters: The following JSON string is an example and does not list all parameters. Add the parameters you need when you write your code:
{
"service_type": 4,
"nls_config": {
"model": "qwen-audio-3.0-asr-flash-streaming",
"sr_format": "pcm",
"sample_rate": "16000"
}
}
Parameter descriptions
| Top-level parameter | Type | Required | Description |
|---|---|---|---|
service_type | int | Yes | Speech service type. Fixed to 4 for real-time speech recognition. |
nls_config | object | Yes | The core recognition configuration object. It contains key parameters such as model selection and recognition quality controls. |
nls_config.model | string | Yes | The model name. |
nls_config.sr_format | string | Yes | The audio format. Valid values:
ImportantIf you provide PCM audio data and set this parameter to |
nls_config.sample_rate | int | Yes | The sample rate in Hz. Any sample rate is supported. Important8000 Hz is not supported when on-device audio processing capabilities such as AEC or VAD are enabled. |
nls_config.semantic_punctuation_enabled | boolean | No | Whether to enable semantic segmentation. Default value: false.
Semantic segmentation is more accurate and is better suited to meeting transcription scenarios. VAD (Voice Activity Detection) segmentation has lower latency and is better suited to interactive scenarios. |
nls_config.max_sentence_silence | int | No | The VAD silence threshold for segmentation, in ms. When the silence after a segment of speech exceeds this threshold, the system determines that the sentence has ended. When semantic_punctuation_enabled is set to true, this parameter is not used as the criterion for returning sentence_end, but setting it too low may affect recognition performance.Default value: 1300. Valid values: [200, 6000]. |
nls_config.multi_threshold_mode_enabled | boolean | No | ImportantTakes effect only when Whether to enable multi-threshold mode. When enabled, this prevents VAD segments from becoming too long. Default value: false. |
nls_config.heartbeat | boolean | No | Whether to enable heartbeat packets. Default value: false.
Silent audio refers to content in an audio file or data stream that contains no sound signal. You can generate silent audio in several ways, such as using audio editing software like Audacity or Adobe Audition, or using a command-line tool like FFmpeg. |
nls_config.vocabulary_id | string | No | The ID of a precompiled hot word list. Generate this ID in advance by calling the create hot word list API. Pass the ID during recognition to use the hot words in the list. Suitable for scenarios where the vocabulary is known and relatively stable, and where you need to reuse the same word list across requests. For usage details, see Precompiled hotwords. |
nls_config.instant_vocabulary | object | No | Instant hot words. Passed as key-value pairs, where the key is the hot word text ( Suitable for temporary, session-level hot word optimization. When instant and precompiled hotwords are configured together, the system merges both sets. If the merged set contains more than 2000 hotwords, the system randomly selects 2000 to use. For usage, supported models, and limits, see Instant hotwords. |
nls_config.language_hints | array[string] | No | The language of the audio to recognize. There is no default value; if not set, the model detects the language automatically. You can set up to 4 values. If you set more, only the first 4 take effect. Click to view the supported language codes
|
nls_config.speech_noise_threshold | float | No | The threshold for distinguishing speech from noise, used to adjust the sensitivity of Voice Activity Detection (VAD). Valid values: [-1.0, 1.0]. Value descriptions:
This is an advanced configuration parameter. Adjusting it can significantly affect recognition results. Recommendations:
|
nls_config.special_word_filter | object | No | Specifies the sensitive words to process during speech recognition, and supports setting different processing methods for different sensitive words. For details, see Sensitive word filtering. |
nls_config.enable_connection_fast_check | BOOL | No | Whether to enable fast network checks so that disconnections can be reported as soon as possible. Default: NO. |
Key APIs
NeoNui
nui_initialize
Initializes the speech recognition SDK instance. The SDK is a singleton. Do not initialize it again before you call nui_release.
-(NuiResultCode) nui_initialize:(const char *)parameters
logLevel:(NuiSdkLogLevel)level
saveLog:(BOOL)save_log;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
parameters | char* | A JSON string that contains authentication, connection, and debugging parameters. See Connection and control parameters. |
level | NuiSdkLogLevel | Print level for the SDK's own logs. |
save_log | BOOL | Whether to save logs locally. If set to YES, specify a path with debug_path in Connection and control parameters, and optionally set the file size with max_log_file_size. |
nui_set_params
Sets the Recognition quality parameters in JSON format. Call this method before nui_dialog_start.
-(NuiResultCode) nui_set_params:(const char *)params;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
params | char* | Recognition quality parameters. |
nui_dialog_start
Starts recognition.
Method signature-(NuiResultCode) nui_dialog_start:(NuiVadMode)vad_mode
dialogParam:(const char *)dialog_params;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
vad_mode | NuiVadMode | VAD mode. Fixed to MODE_P2T. |
dialog_params | char* | If the apikey parameter in Connection and control parameters is a temporary API key, you can update it here when it expires. You can also provide context here to improve recognition accuracy through context enhancement.The content is in JSON format: |
nui_dialog_cancel
Ends recognition or cancels the current interaction immediately.
Method signature-(NuiResultCode) nui_dialog_cancel:(BOOL)force;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
force | BOOL | Whether to end forcibly and discard the final result.
|
nui_dialog_action
Sends a dialog action during an interaction to update the recognition context or other runtime behavior.
Method signature-(NuiResultCode) nui_dialog_action:(const char *)action_params;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
action_params | char* | A JSON string used to update the recognition context or other runtime behavior. |
action_params.type | String | Fixed to "action". |
action_params.command | String | Runtime command. Valid values:
|
action_params.context | String | When command is "context", provides the context enhancement to update immediately. Example: |
nui_update_audio_data
When audio_update_manually is set to "true", recording data is no longer supplied through onNuiNeedAudioData. Use this method to actively push it instead.
-(NuiResultCode) nui_update_audio_data:(const char *)data
Len:(int)length
FirstPack:(BOOL)first_pack;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
data | const char * | The audio data to push. |
length | int | The length of the audio data, in bytes. |
first_pack | BOOL | You do not need to use this parameter. |
nui_push_reference_data
When audio_update_manually is set to "true" and on-device acoustic echo cancellation (AEC) is enabled, use this method to push the audio played by the player as the reference signal.
-(NuiResultCode) nui_push_reference_data:(const char *)data
Len:(int)length
FirstPack:(BOOL)first_pack;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
data | const char * | The audio data to push. |
length | int | The length of the audio data, in bytes. |
first_pack | BOOL | You do not need to use this parameter. |
nui_release
Releases all internal SDK resources and forcibly terminates all running tasks. After you call this method, the SDK instance becomes unavailable. To use it again, you must call nui_initialize to initialize it again.
-(NuiResultCode) nui_release;
nui_get_version
Gets the current SDK version. This method returns a value only after nui_initialize is called.
-(const char*) nui_get_version;
Return value
The current SDK version.
nui_get_all_response
Gets the complete information for the current event callback.
Method signature-(const char*) nui_get_all_response;
Return value
The complete event information as a JSON string.
NeoNuiSdkDelegate: callback listeners
onNuiEventCallback: listen for events and speech recognition results
Method signature-(void) onNuiEventCallback:(NuiCallbackEvent)nuiEvent
dialog:(long)dialog
kwsResult:(const char *)wuw
asrResult:(const char *)asr_result
ifFinish:(BOOL)finish
retCode:(int)code;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
nuiEvent | NuiCallbackEvent | Callback event. |
dialog | long | Session ID. You do not need to use this parameter. |
wuw | char* | Voice wake-up. You do not need to use this parameter. |
asr_result | char* | Speech recognition result. |
finish | BOOL | Whether the current recognition round has finished. |
code | int | Valid only when the EVENT_ASR_ERROR event occurs. |
onNuiAudioStateChanged: listen for the audio state
The SDK uses this callback to notify you when to start or stop recording.
Method signature-(void) onNuiAudioStateChanged:(NuiAudioState)state;
NuiAudioState states
| Parameter | Description |
|---|---|
STATE_OPEN | The interaction has started. You can open the recording device and start recording. |
STATE_PAUSE | The interaction has stopped. You can stop recording. |
STATE_CLOSE | The SDK instance has been released. You can close the recording device completely. |
onNuiNeedAudioData: supply audio data for recognition
After recognition starts, this callback is triggered continuously. Supply the audio data to be recognized in it.
Method signature-(int) onNuiNeedAudioData:(char *)audioData length:(int)len;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
audioData | char * | The audio data to supply. |
len | int | The size of the supplied audio data, in bytes. |
onNuiAssistEventCallback: receive auxiliary data and information
Receives auxiliary events and related data from the SDK.
Method signature-(void) onNuiAssistEventCallback:(NuiCallbackEvent)nuiEvent
info:(char*)info
infoLen:(int)info_len
buffer:(char*)buffer
len:(int)len;
Parameter descriptions
| Parameter | Type | Description |
|---|---|---|
nuiEvent | NuiCallbackEvent | The callback event. |
info | char * | You do not need to use this parameter. |
info_len | int | You do not need to use this parameter. |
buffer | char * | Auxiliary data, such as audio data after AEC processing. |
len | int | The length of the auxiliary data, in bytes. |
onNuiLogTrackCallback: listen for tracking logs
This callback receives the SDK's detailed internal logs to help you locate and debug issues.
-(void) onNuiLogTrackCallback:(NuiSdkLogLevel)level
logMessage:(const char *)log;
NuiCallbackEvent: event types
| Event | Description |
|---|---|
EVENT_TRANSCRIBER_STARTED | The task started successfully. |
EVENT_VAD_START | Triggered right after the task starts. It does not mean that the start of speech has been detected. |
EVENT_VAD_END | The end of speech was detected. |
EVENT_ASR_PARTIAL_RESULT | An intermediate speech recognition result. |
EVENT_ASR_ERROR | An error occurred during speech recognition. |
EVENT_MIC_ERROR | Triggered because no audio data was received for 2 consecutive seconds. |
EVENT_SENTENCE_END | The end of a sentence was detected. A complete recognition result for the sentence is returned. |
EVENT_TRANSCRIBER_COMPLETE | Speech recognition finished. |
EVENT_AEC_DATA | Audio data after acoustic echo cancellation (AEC) processing. |