All Products
Search
Document Center

Alibaba Cloud Model Studio:Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR real-time speech recognition Android SDK guide

Last Updated:Sep 10, 2026

This guide shows you how to use the Android SDK for Qwen-Audio-3.0-ASR-Flash-Streaming and Fun-ASR-Realtime real-time speech recognition to convert speech to text.

User guide: For model introductions and selection advice, see Speech-to-text.

Quick start

  1. Obtain an API key

  2. Download the SDK and run the sample code:
    • Download the latest SDK package.
    • Extract the ZIP package. The AAR-format SDK is in the app/libs directory. Add it to your project dependencies. For Android C++ integration, use android_libs and android_include from the ZIP package to get the dynamic libraries and header files.
    • Open the project in Android Studio. The sample code is in DashFunAsrSpeechTranscriberActivity.java. Replace the API key to try out the feature.

Call procedure

  1. Initialize the SDK.
  2. Set the parameters for your use case. Use the parameters argument of the initialize method to set the Connection and control parameters, and use the setParams method to set the Speech recognition parameters.
  3. Call startDialog to start recognition.
  4. In the onNuiAudioStateChanged callback, start the recording device based on the audio state.
  5. Continuously supply recording data in the onNuiNeedAudioData callback, or actively push recording data by calling updateAudio.
  6. In the onNuiEventCallback callback, listen for events and get the speech recognition results.
  7. Call stopDialog to stop recognition, and listen for the EVENT_TRANSCRIBER_COMPLETE event to confirm that recognition has ended.
  8. When you no longer need recognition, call release to release the SDK resources.

Request parameters

Connection and control parameters

To configure these parameters, pass a JSON string in the parameters argument of the initialize method.

  • Example: The following is a JSON string example. Not all parameters are listed. Add others as needed when you write your code:
{
    "url": "wss://dashscope.aliyuncs.com/api-ws/v1/inference",
    "apikey": "st-****",
    "device_id": "my_device_id",
    "service_mode": "1"
}
  • Parameters

    Parameter

    Type

    Required

    Description

    url

    String

    Yes

    Service address:

    • wss://dashscope.aliyuncs.com/api-ws/v1/inference

    • China (Beijing): wss://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api-ws/v1/inference

    • Singapore: wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inference

    Replace {WorkspaceId} with your actual Workspace ID.

    apikey

    String

    Yes

    The API key.

    service_mode

    String

    Yes

    The run mode. Fixed to "1" for real-time speech recognition.

    device_id

    String

    Yes

    A unique string that identifies the end user. You can set it to an in-app user ID or a client-generated unique device identifier. This ID is mainly used for log tracing and troubleshooting.

    audio_update_manually

    String

    No

    Whether to enable active audio data pushing. Default: "false".

    If set to "true" and the SDK version supports on-device audio processing capabilities such as AEC and VAD, those capabilities are enabled by default.

    workspace

    String

    No

    The path where on-device resource files are stored. This parameter is required when audio_update_manually is set to "true" and an on-device audio processing capability such as AEC or VAD is enabled.

    debug_path

    String

    No

    The storage path for log files.

    This parameter takes effect only when save_log is set to true in the initialize method. In this case, you must set the log file path, or an error occurs.

    A maximum of two log files are kept locally.

    save_wav

    String

    No

    Whether to save the debug audio file. The audio file is saved under debug_path.

    Default: "false".

    Valid values:

    • "true": save the file.

    • "false": do not save the file.

    This parameter takes effect only when save_log is set to true in the initialize method. In addition, debug_path must also be set.

    max_log_file_size

    int

    No

    The maximum size of a log file, in bytes.

    This parameter takes effect only when save_log is set to true in the initialize method.

    Default: 104857600 (100 * 1024 * 1024 bytes, that is, 100 MiB).

    log_track_level

    int

    No

    The filter level for the log content sent through the log callback (onNuiLogTrackCallback).

    Default: 2.

    Valid values:

    • 0: LOG_LEVEL_VERBOSE

    • 1: LOG_LEVEL_DEBUG

    • 2: LOG_LEVEL_INFO

    • 3: LOG_LEVEL_WARNING

    • 4: LOG_LEVEL_ERROR

    • 5: LOG_LEVEL_NONE (disables this feature)

    Note: log_track_level and level (set through the initialize method) together determine which logs are ultimately sent to the callback. A log is sent to the callback only when its level value is greater than or equal to both log_track_level and level. For example, if log_track_level is set to 2 (INFO) and level is set to 3 (WARNING), then only logs of WARNING level or higher (value >= 3) are sent to the callback.

    enable_reconnection

    String

    No

    Whether to resume transmission after the network reconnects. Default: "false".

    aec_params

    object

    No

    The advanced configuration object for on-device acoustic echo cancellation (AEC). This object takes effect only when audio_update_manually is set to "true".

    aec_params.enable_aec

    boolean

    No

    Whether to enable on-device AEC. If audio_update_manually is set to "true" and the SDK version supports on-device AEC, this capability is enabled by default.

    aec_params.save_audio

    boolean

    No

    Whether to save audio processed by the on-device AEC module. If save_wav is set to "true" and debug_path is specified, this feature is enabled by default and the AEC audio data is saved to debug_path.

    aec_params.enable_aec_data_callback

    boolean

    No

    Whether to return the AEC-processed data to your application. Default: false. If enabled, receive the data from the EVENT_AEC_DATA event in onNuiAssistEventCallback.

    vad_params

    object

    No

    The advanced configuration object for on-device voice activity detection (VAD). This object takes effect only when audio_update_manually is set to "true".

    vad_params.enable_aec

    boolean

    No

    Whether to enable on-device VAD. If audio_update_manually is set to "true" and the SDK version supports on-device VAD, this capability is enabled by default.

    vad_params.save_audio

    boolean

    No

    Whether to save audio processed by the on-device VAD module. If save_wav is set to "true" and debug_path is specified, this feature is enabled by default and the VAD audio data is saved to debug_path.

Speech recognition parameters

To configure these parameters, pass a JSON string in the params argument of the setParams method.

  • Example: The following is a JSON string example. Not all parameters are listed. Add others as needed when you write your code:
{
    "service_type": 4,
    "nls_config": {
        "model": "qwen-audio-3.0-asr-flash-streaming",
        "sr_format": "pcm",
        "sample_rate": "16000"
    }
}
  • Parameters
    Top-level parameterTypeRequiredDescription

    service_type

    int

    Yes

    The speech service type. Fixed to 4 for real-time speech recognition.

    nls_config

    object

    Yes

    The core configuration object for speech recognition. It contains key parameters such as model selection and recognition-quality controls.

    nls_config.model

    string

    Yes

    The model name. The Qwen-Audio-3.0-ASR-Flash-Streaming and Fun-ASR-Realtime model series are supported. For details, see Supported models and regions.

    nls_config.sr_format

    string

    Yes

    The audio format.

    Valid values:

    • pcm
    • opus

    ImportantFor Opus audio, pass PCM audio to the SDK. The SDK encodes it as Opus internally.

    nls_config.sample_rate

    int

    Yes

    The sample rate, in Hz.

    Valid values: 8 kHz models support only 8000 Hz; other models support any sample rate.

    Important8000 Hz is not supported when an on-device audio processing capability such as AEC or VAD is enabled.

    nls_config.semantic_punctuation_enabled

    boolean

    No

    Whether to enable semantic segmentation.

    Default value: false.

    • true: Enables semantic segmentation and disables VAD segmentation.
    • false (default): Enables VAD segmentation and disables semantic segmentation.

    Semantic segmentation is more accurate and is better suited to meeting transcription scenarios. VAD (Voice Activity Detection) segmentation has lower latency and is better suited to interactive scenarios.

    nls_config.max_sentence_silence

    int

    No

    The VAD silence threshold for segmentation, in ms. When the silence after a segment of speech exceeds this threshold, the system determines that the sentence has ended. When semantic_punctuation_enabled is set to true, this parameter is not used as the criterion for returning sentence_end, but setting it too low may affect recognition performance.

    Default value: 1300.

    Valid values: [200, 6000].

    nls_config.multi_threshold_mode_enabled

    boolean

    No

    ImportantTakes effect only when semantic_punctuation_enabled is false.

    Whether to enable multi-threshold mode. When enabled, this prevents VAD segments from becoming too long.

    Default value: false.

    nls_config.heartbeat

    boolean

    No

    Whether to enable heartbeat packets.

    Default value: false.

    • true: Keeps the connection to the server alive even when silent audio is sent continuously.
    • false (default): Even when silent audio is continuously sent, the connection times out and closes after a period of time.

    Silent audio refers to content in an audio file or data stream that contains no sound signal. You can generate silent audio in several ways, such as using audio editing software like Audacity or Adobe Audition, or using a command-line tool like FFmpeg.

    nls_config.vocabulary_id

    string

    No

    The ID of a precompiled hot word list.

    Generate this ID in advance by calling the create hot word list API. Pass the ID during recognition to use the hot words in the list.

    Suitable for scenarios where the vocabulary is known and relatively stable, and where you need to reuse the same word list across requests.

    For usage details, see Precompiled hotwords.

    nls_config.instant_vocabulary

    object

    No

    Instant hot words.

    Passed as key-value pairs, where the key is the hot word text (string) and the value is the hot word weight (integer). No hot word list needs to be created in advance. The weight ranges from [1, 5] or is set to 50: a value in [1, 5] makes the model more likely to output the word as the value increases; a value of 50 designates a super hot word, which greatly improves recall, but the number of super hot words cannot exceed 50.

    Suitable for temporary, session-level hot word optimization.

    When instant and precompiled hotwords are configured together, the system merges both sets. If the merged set contains more than 2000 hotwords, the system randomly selects 2000 to use. For usage details, see Instant hotwords.

    ImportantOnly qwen-audio-3.0-asr-flash-streaming supports instant hot words.

    nls_config.language_hints

    array[string]

    No

    The language of the audio to recognize. There is no default value; if not set, the model detects the language automatically.

    For the Qwen-Audio-3.0-ASR-Flash-Streaming model series, you can set up to 4 values; if you set more than 4, only the first 4 take effect. For the Fun-ASR-Realtime model series, you can set only 1 value; if you set multiple values, only the first one takes effect.

    Click to view the supported language codes

    • qwen-audio-3.0-asr-flash-streaming, fun-asr-realtime, fun-asr-realtime-2025-11-07:

      • zh: Chinese
      • en: English
      • ja: Japanese
      • ko: Korean
      • vi: Vietnamese
      • th: Thai
      • id: Indonesian
      • ms: Malay
      • tl: Filipino
      • hi: Hindi
      • ar: Arabic
      • fr: French
      • de: German
      • es: Spanish
      • pt: Portuguese
      • ru: Russian
      • it: Italian
      • nl: Dutch
      • sv: Swedish
      • da: Danish
      • fi: Finnish
      • no: Norwegian
      • el: Greek
      • pl: Polish
      • cs: Czech
      • hu: Hungarian
      • ro: Romanian
      • bg: Bulgarian
      • hr: Croatian
      • sk: Slovak
    • fun-asr-realtime-2026-02-28:

      • zh: Chinese
      • en: English
      • ja: Japanese
    • fun-asr-realtime-2025-09-15:

      • zh: Chinese
      • en: English
    • fun-asr-flash-8k-realtime, fun-asr-flash-8k-realtime-2026-01-28:

      • zh: Chinese
    • fun-asr-mtl-realtime, fun-asr-mtl-realtime-2025-12-10:

      • zh: Chinese
      • en: English
      • ja: Japanese
      • ko: Korean
      • vi: Vietnamese
      • id: Indonesian
      • th: Thai

    nls_config.speech_noise_threshold

    float

    No

    The threshold for distinguishing speech from noise, used to adjust the sensitivity of Voice Activity Detection (VAD).

    Valid values: [-1.0, 1.0].

    Value descriptions:

    • The closer the value is to -1: The noise threshold decreases, so noise is more likely to be recognized as speech, which may cause more noise to be transcribed.
    • The closer the value is to +1: The noise threshold increases, so speech is more likely to be misjudged as noise, which may cause some speech to be filtered out.

    This is an advanced configuration parameter. Adjusting it can significantly affect recognition results. Recommendations:

    • Thoroughly test and verify the results before adjusting.
    • Adjust in small increments based on the actual audio environment (a step of 0.1 is recommended).

    nls_config.special_word_filter

    object

    No

    Specifies the sensitive words to process during speech recognition, and supports setting different processing methods for different sensitive words. For details, see Sensitive word filtering.

    nls_config.enable_connection_fast_check

    boolean

    No

    Whether to quickly detect network outages and report them as soon as possible. Default: false.

Key interfaces

NativeNui

initialize

Initializes a speech recognition SDK instance. The SDK is a singleton. Do not initialize it more than once before you call release.

This method blocks, so call it on a non-UI thread.

  • Method signature
public synchronized int initialize(final INativeNuiCallback callback,
                                   String parameters,
                                   final Constants.LogLevel level,
                                   final boolean save_log)
  • Parameters

    Parameter

    Type

    Description

    callback

    INativeNuiCallback

    The implementation of the event and data callback interface.

    parameters

    String

    A JSON string that contains the authentication, connection, and debug parameters. See Connection and control parameters.

    level

    Constants.LogLevel

    Controls the print level of the SDK's own logs.

    save_log

    boolean

    Whether to save local logs. If set to true, specify the path through debug_path in the Connection and control parameters, and optionally set the file size through max_log_file_size.

  • Return value

    An error code. See Error code reference.

setParams

Sets the Speech recognition parameters in JSON format. Call this method before startDialog.

  • Method signature
public synchronized int setParams(String params)

startDialog

Starts recognition.

  • Method signature
public synchronized int startDialog(VadMode vad_mode, String dialog_params)
  • Parameters
    ParameterTypeDescription

    vad_mode

    VadMode

    The VAD mode. Fixed to VadMode.TYPE_P2T.

    dialog_params

    String

    If the apikey parameter in Connection and control parameters uses a temporary API key, you can update it here when it expires.

    To improve recognition accuracy by using context, update the context here.

    The content is in JSON format:

    {
      "apikey": "st-****",
      "input_context": [
        {
          "role": "user",
          "content": [
            {
              "text": "xxxxx",
              "type": "input_text"
            }
          ]
        }
      ]
    }
    
  • Return value

    An error code. See Error code reference.

stopDialog

Ends recognition. After you call this method, the server returns the final recognition result and ends the task.

  • Method signature
public synchronized int stopDialog();

cancelDialog

Ends recognition immediately. After you call this method, the task ends at once without waiting for the server to return the final recognition result.

  • Method signature
public synchronized int cancelDialog();

updateAction

Sends an action command during an interaction to update runtime behavior, such as recognition context.

  • Method signature
public synchronized int updateAction(String params);
  • Parameters

    Parameter

    Type

    Description

    params

    String

    A JSON string used to update runtime behavior, such as recognition context.

    params.type

    String

    Set to "action".

    params.command

    String

    The runtime command. Valid values:

    • context: Immediately updates the context to improve recognition accuracy.

    • play_start: When on-device AEC is used, notifies the SDK that the player has started playing audio.

    • play_over: When on-device AEC is used, notifies the SDK that the player has finished playing audio.

    params.context

    String

    When command is set to context, immediately updates the context to improve recognition accuracy. The value is a JSON string, as shown in the following example.

{
  "context": [
    {
      "role": "user",
      "content": [
        {
          "text": "xxx",
          "type": "input_text"
        }
      ]
    }
  ]
}

updateAudio

When audio_update_manually is set to "true", call this method to actively push recording data instead of supplying the data through onNuiNeedAudioData.

  • Method signature
public synchronized int updateAudio(byte[] data, int len,
                                    boolean first_pack);
  • Parameters

    Parameter

    Type

    Description

    data

    byte[]

    The audio data to push.

    len

    int

    The number of bytes of audio data to push.

    first_pack

    boolean

    Ignore this parameter.

  • Return value

    An error code. See Error code reference.

updateRefAudio

When audio_update_manually is set to "true" and on-device AEC is enabled, call this method to push the audio played by the player as the reference signal.

  • Method signature
public synchronized int updateRefAudio(byte[] data, int len,
                                       boolean first_pack);
  • Parameters

    Parameter

    Type

    Description

    data

    byte[]

    The audio data to push.

    len

    int

    The number of bytes of audio data to push.

    first_pack

    boolean

    Ignore this parameter.

  • Return value

    An error code. See Error code reference.

release

Releases all internal resources of the SDK. After you call this method, the SDK instance becomes unavailable. To use it again, you must reinitialize it by calling initialize.

  • Method signature
public synchronized int release();

GetVersion

Gets the current SDK version information.

  • Method signature
public synchronized String GetVersion();
  • Return value

    The current SDK version information.

INativeNuiCallback: listener callbacks

onNuiEventCallback: listen for events and speech recognition results

  • Method signature
void onNuiEventCallback(NuiEvent event, final int resultCode, final int arg2, KwsResult kwsResult, AsrResult asrResult);
  • Parameters

    Parameter

    Type

    Description

    event

    NuiEvent

    The callback event.

    resultCode

    int

    The error code. Valid when the EVENT_ASR_ERROR event occurs.

    arg2

    int

    A reserved parameter.

    asrResult

    AsrResult

    The speech recognition result.

    kwsResult

    KwsResult

    The voice wake-up feature. You do not need to use this parameter.

onNuiAudioStateChanged: listen for the audio state

The SDK uses this callback to notify you when to start or stop recording.

  • Method signature
void onNuiAudioStateChanged(AudioState state);
  • AudioState states

    State

    Description

    STATE_OPEN

    The interaction has started. You can open the recording device and start recording.

    STATE_PAUSE

    The interaction has stopped. You can stop recording.

    STATE_CLOSE

    The SDK instance has been released. You can fully close the recording device.

onNuiNeedAudioData: supply the audio data to recognize

After recognition starts, this callback is triggered continuously. Supply the audio data to recognize in this callback.

  • Method signature
int onNuiNeedAudioData(byte[] buffer, int len);
  • Parameters

    Parameter

    Type

    Description

    buffer

    byte[]

    The audio data to fill.

    len

    int

    The number of bytes of audio data to fill.

  • Return value

    The number of bytes actually filled.

onNuiAssistEventCallback: receive auxiliary events and data

This callback receives auxiliary events and related data from the SDK.

  • Method signature
void onNuiAssistEventCallback_(int event, byte[] info, int info_len,
                               byte[] data);
  • Parameters

    Parameter

    Type

    Description

    event

    int

    A NuiEvent event.

    info

    String

    Ignore this parameter.

    info_len

    int

    Ignore this parameter.

    data

    byte[]

    Auxiliary data, such as audio data processed by AEC.

onNuiLogTrackCallback: listen for trace logs

This callback receives detailed internal logs from the SDK to help with troubleshooting and debugging.

default void onNuiLogTrackCallback(Constants.LogLevel level, String log)

NuiEvent: event types

Event

Description

EVENT_TRANSCRIBER_STARTED

The task started successfully.

EVENT_VAD_START

Triggered right after the task starts. This does not mean that the start of speech is detected.

EVENT_VAD_END

The end of speech is detected.

EVENT_ASR_PARTIAL_RESULT

An intermediate speech recognition result.

EVENT_ASR_WARN

A warning that does not interrupt speech recognition occurred, such as a network outage when reconnection is enabled.

EVENT_ASR_ERROR

An error occurred during speech recognition.

EVENT_MIC_ERROR

Triggered when no audio data is received for 2 consecutive seconds.

EVENT_SENTENCE_END

The end of a sentence is detected. A complete recognition result for the sentence is returned.

EVENT_TRANSCRIBER_COMPLETE

Speech recognition has ended.

EVENT_AEC_DATA

Audio data processed by AEC.