All Products
Search
Document Center

Intelligent Speech Interaction:Introduction to the speech synthesis timestamp feature

Last Updated:Sep 10, 2026

Speech synthesis timestamps mark the start and end of each Chinese character or English word in synthesized audio. Use them to synchronize video subtitles, text highlighting, and virtual character lip movements. Enable timestamps in the synthesis request to receive timing data alongside the audio stream.

Limitations

  • Word-level timestamps are available only for voices that support word-level phoneme boundaries.

  • The short-text speech synthesis RESTful API does not return timestamps. Use the WebSocket API or a corresponding SDK to receive them.

  • Subtitle text follows pronunciation and may not match the original text character by character. For example, in Chinese speech, 12 can be represented by two subtitle units: 十 and 二. Display the original text and use the returned timing data to determine sentence boundaries or highlighting times. Do not treat subtitle indexes as character offsets into the original string.

Code example

This Java example uses nls-sdk-tts 2.1.6. It reads the project AppKey and a valid NLS token from environment variables, prints subtitle information, and saves the audio to tts_test.wav. An existing file with that name is overwritten.

Add dependencies and configure credentials

Add the following dependency to the dependencies section of the Maven project:

<dependency>
    <groupId>com.alibaba.nls</groupId>
    <artifactId>nls-sdk-tts</artifactId>
    <version>2.1.6</version>
</dependency>

Set NLS_APP_KEY and NLS_TOKEN in the runtime environment to the project AppKey and NLS token. The token must be valid and belong to the same account as the project. Do not substitute an STS token or an API key for another product.

Synthesize speech and receive timestamps

import com.alibaba.fastjson.JSONArray;
import com.alibaba.nls.client.protocol.NlsClient;
import com.alibaba.nls.client.protocol.OutputFormatEnum;
import com.alibaba.nls.client.protocol.SampleRateEnum;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizer;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizerListener;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizerResponse;
import java.io.IOException;
import java.io.OutputStream;
import java.nio.ByteBuffer;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicReference;

public class SpeechSynthesizerDemo {
    private static String requiredEnv(String name) {
        String value = System.getenv(name);
        if (value == null || value.trim().isEmpty()) {
            throw new IllegalArgumentException("Missing environment variable: " + name);
        }
        return value;
    }

    public static void main(String[] args) throws Exception {
        String appKey = requiredEnv("NLS_APP_KEY");
        String token = requiredEnv("NLS_TOKEN");
        String endpoint = "wss://nls-gateway-ap-southeast-1.aliyuncs.com/ws/v1";
        NlsClient client = new NlsClient(endpoint, token);

        try (OutputStream audio = Files.newOutputStream(Paths.get("tts_test.wav"))) {
            AtomicReference<String> failure = new AtomicReference<>();
            AtomicBoolean completed = new AtomicBoolean(false);
            AtomicBoolean firstAudio = new AtomicBoolean(true);
            final long start = System.nanoTime();
            SpeechSynthesizerListener listener = new SpeechSynthesizerListener() {
                @Override
                public void onMessage(ByteBuffer message) {
                    if (firstAudio.compareAndSet(true, false)) {
                        System.out.println("First audio latency (ms): "
                                + (System.nanoTime() - start) / 1_000_000);
                    }
                    byte[] bytes = new byte[message.remaining()];
                    message.get(bytes);
                    try {
                        audio.write(bytes);
                    } catch (IOException e) {
                        failure.compareAndSet(null, "Failed to write audio: " + e.getMessage());
                    }
                }

                @Override
                public void onMetaInfo(SpeechSynthesizerResponse response) {
                    JSONArray subtitles = (JSONArray) response.getObject("subtitles");
                    if (subtitles != null) {
                        System.out.println("MetaInfo: " + subtitles.toJSONString());
                    }
                }

                @Override
                public void onComplete(SpeechSynthesizerResponse response) {
                    completed.set(true);
                    System.out.println("SynthesisCompleted: " + response.getTaskId());
                }

                @Override
                public void onFail(SpeechSynthesizerResponse response) {
                    failure.set("task_id=" + response.getTaskId()
                            + ", status=" + response.getStatus()
                            + ", status_text=" + response.getStatusText());
                }
            };

            SpeechSynthesizer synthesizer = new SpeechSynthesizer(client, listener);
            try {
                synthesizer.setAppKey(appKey);
                synthesizer.setFormat(OutputFormatEnum.WAV);
                synthesizer.setSampleRate(SampleRateEnum.SAMPLE_RATE_16K);
                synthesizer.setVoice("siyue");
                synthesizer.setPitchRate(100);
                synthesizer.setSpeechRate(100);
                synthesizer.setText("Hello world. I have 12 apples.");
                synthesizer.addCustomedParam("enable_subtitle", true);
                synthesizer.start();
                synthesizer.waitForComplete();
                if (failure.get() != null) {
                    throw new IllegalStateException(failure.get());
                }
                if (!completed.get()) {
                    throw new IllegalStateException("Synthesis did not complete");
                }
                System.out.println("Audio saved to tts_test.wav");
            } finally {
                synthesizer.close();
            }
        } finally {
            client.shutdown();
        }
    }
}

On success, the program prints MetaInfo subtitle arrays, SynthesisCompleted, and an audio-file confirmation. First-audio latency is measured from before synthesis starts until the first audio chunk arrives. It differs from the time needed to receive all audio.

Note

The example saves audio to a file. For latency-sensitive playback, play audio chunks as they arrive in onMessage instead of waiting for synthesis to finish.

Word-level timestamps

Request parameters

Before starting synthesis, set enable_subtitle to true. Timestamps are disabled by default. In the Java SDK, use:

synthesizer.addCustomedParam("enable_subtitle", true);

Server response

The service returns timestamps in MetaInfo events. Each event contains a subtitle array in payload.subtitles. The Java SDK delivers these events through the onMetaInfo callback.

Fields in each subtitles element

Field

Type

Description

text

String

A pronunciation-based text unit, such as a Chinese character or an English word.

begin_time

Integer

The start time of the text unit in the synthesized audio, in milliseconds.

end_time

Integer

The end time of the text unit in the synthesized audio, in milliseconds.

phoneme

String

When only word-level timestamps are enabled, this field contains the string "null", not JSON null.

begin_index

Integer

The zero-based start index of the text unit. This is not a character offset into the original string.

end_index

Integer

The exclusive end index of the text unit.

For example, the two words in Hello world! occupy index ranges [0, 1) and [1, 2). Punctuation can occupy an index position: in the Chinese input 你好,世界。, the character 世 has a start index of 3.

Note

A synthesis request can produce multiple MetaInfo events. Subtitle arrays can repeat or include previously returned entries. Update existing data by index and timestamp instead of appending each entire array.

Sample response

The following sample payload is returned for Hello world!. Timing values vary with the voice, speech rate, and text.

{
  "subtitles": [
    {
      "text": "Hello",
      "phoneme": "null",
      "begin_index": 0,
      "end_index": 1,
      "begin_time": 0,
      "end_time": 328
    },
    {
      "text": "world",
      "phoneme": "null",
      "begin_index": 1,
      "end_index": 2,
      "begin_time": 328,
      "end_time": 775
    }
  ]
}