Speech synthesis timestamps mark the start and end of each Chinese character or English word in synthesized audio. Use them to synchronize video subtitles, text highlighting, and virtual character lip movements. Enable timestamps in the synthesis request to receive timing data alongside the audio stream.
Limitations
-
Word-level timestamps are available only for voices that support word-level phoneme boundaries.
-
The short-text speech synthesis RESTful API does not return timestamps. Use the WebSocket API or a corresponding SDK to receive them.
-
Subtitle text follows pronunciation and may not match the original text character by character. For example, in Chinese speech,
12can be represented by two subtitle units:十and二. Display the original text and use the returned timing data to determine sentence boundaries or highlighting times. Do not treat subtitle indexes as character offsets into the original string.
Code example
This Java example uses nls-sdk-tts 2.1.6. It reads the project AppKey and a valid NLS token from environment variables, prints subtitle information, and saves the audio to tts_test.wav. An existing file with that name is overwritten.
Add dependencies and configure credentials
Add the following dependency to the dependencies section of the Maven project:
<dependency>
<groupId>com.alibaba.nls</groupId>
<artifactId>nls-sdk-tts</artifactId>
<version>2.1.6</version>
</dependency>
Set NLS_APP_KEY and NLS_TOKEN in the runtime environment to the project AppKey and NLS token. The token must be valid and belong to the same account as the project. Do not substitute an STS token or an API key for another product.
Synthesize speech and receive timestamps
import com.alibaba.fastjson.JSONArray;
import com.alibaba.nls.client.protocol.NlsClient;
import com.alibaba.nls.client.protocol.OutputFormatEnum;
import com.alibaba.nls.client.protocol.SampleRateEnum;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizer;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizerListener;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizerResponse;
import java.io.IOException;
import java.io.OutputStream;
import java.nio.ByteBuffer;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicReference;
public class SpeechSynthesizerDemo {
private static String requiredEnv(String name) {
String value = System.getenv(name);
if (value == null || value.trim().isEmpty()) {
throw new IllegalArgumentException("Missing environment variable: " + name);
}
return value;
}
public static void main(String[] args) throws Exception {
String appKey = requiredEnv("NLS_APP_KEY");
String token = requiredEnv("NLS_TOKEN");
String endpoint = "wss://nls-gateway-ap-southeast-1.aliyuncs.com/ws/v1";
NlsClient client = new NlsClient(endpoint, token);
try (OutputStream audio = Files.newOutputStream(Paths.get("tts_test.wav"))) {
AtomicReference<String> failure = new AtomicReference<>();
AtomicBoolean completed = new AtomicBoolean(false);
AtomicBoolean firstAudio = new AtomicBoolean(true);
final long start = System.nanoTime();
SpeechSynthesizerListener listener = new SpeechSynthesizerListener() {
@Override
public void onMessage(ByteBuffer message) {
if (firstAudio.compareAndSet(true, false)) {
System.out.println("First audio latency (ms): "
+ (System.nanoTime() - start) / 1_000_000);
}
byte[] bytes = new byte[message.remaining()];
message.get(bytes);
try {
audio.write(bytes);
} catch (IOException e) {
failure.compareAndSet(null, "Failed to write audio: " + e.getMessage());
}
}
@Override
public void onMetaInfo(SpeechSynthesizerResponse response) {
JSONArray subtitles = (JSONArray) response.getObject("subtitles");
if (subtitles != null) {
System.out.println("MetaInfo: " + subtitles.toJSONString());
}
}
@Override
public void onComplete(SpeechSynthesizerResponse response) {
completed.set(true);
System.out.println("SynthesisCompleted: " + response.getTaskId());
}
@Override
public void onFail(SpeechSynthesizerResponse response) {
failure.set("task_id=" + response.getTaskId()
+ ", status=" + response.getStatus()
+ ", status_text=" + response.getStatusText());
}
};
SpeechSynthesizer synthesizer = new SpeechSynthesizer(client, listener);
try {
synthesizer.setAppKey(appKey);
synthesizer.setFormat(OutputFormatEnum.WAV);
synthesizer.setSampleRate(SampleRateEnum.SAMPLE_RATE_16K);
synthesizer.setVoice("siyue");
synthesizer.setPitchRate(100);
synthesizer.setSpeechRate(100);
synthesizer.setText("Hello world. I have 12 apples.");
synthesizer.addCustomedParam("enable_subtitle", true);
synthesizer.start();
synthesizer.waitForComplete();
if (failure.get() != null) {
throw new IllegalStateException(failure.get());
}
if (!completed.get()) {
throw new IllegalStateException("Synthesis did not complete");
}
System.out.println("Audio saved to tts_test.wav");
} finally {
synthesizer.close();
}
} finally {
client.shutdown();
}
}
}
On success, the program prints MetaInfo subtitle arrays, SynthesisCompleted, and an audio-file confirmation. First-audio latency is measured from before synthesis starts until the first audio chunk arrives. It differs from the time needed to receive all audio.
The example saves audio to a file. For latency-sensitive playback, play audio chunks as they arrive in onMessage instead of waiting for synthesis to finish.
Word-level timestamps
Request parameters
Before starting synthesis, set enable_subtitle to true. Timestamps are disabled by default. In the Java SDK, use:
synthesizer.addCustomedParam("enable_subtitle", true);
Server response
The service returns timestamps in MetaInfo events. Each event contains a subtitle array in payload.subtitles. The Java SDK delivers these events through the onMetaInfo callback.
Fields in each subtitles element
|
Field |
Type |
Description |
|
text |
String |
A pronunciation-based text unit, such as a Chinese character or an English word. |
|
begin_time |
Integer |
The start time of the text unit in the synthesized audio, in milliseconds. |
|
end_time |
Integer |
The end time of the text unit in the synthesized audio, in milliseconds. |
|
phoneme |
String |
When only word-level timestamps are enabled, this field contains the string |
|
begin_index |
Integer |
The zero-based start index of the text unit. This is not a character offset into the original string. |
|
end_index |
Integer |
The exclusive end index of the text unit. |
For example, the two words in Hello world! occupy index ranges [0, 1) and [1, 2). Punctuation can occupy an index position: in the Chinese input 你好,世界。, the character 世 has a start index of 3.
A synthesis request can produce multiple MetaInfo events. Subtitle arrays can repeat or include previously returned entries. Update existing data by index and timestamp instead of appending each entire array.
Sample response
The following sample payload is returned for Hello world!. Timing values vary with the voice, speech rate, and text.
{
"subtitles": [
{
"text": "Hello",
"phoneme": "null",
"begin_index": 0,
"end_index": 1,
"begin_time": 0,
"end_time": 328
},
{
"text": "world",
"phoneme": "null",
"begin_index": 1,
"end_index": 2,
"begin_time": 328,
"end_time": 775
}
]
}