全部产品
Search
文档中心

智能语音交互:语音合成时间戳功能介绍

更新时间:Sep 10, 2026

语音合成时间戳标记汉字或英文单词在合成音频中的起止时间,可用于视频字幕同步、文字高亮和虚拟人口型同步。通过请求参数开启该功能后,在接收音频流的同时获取时间戳。

使用限制

  • 只有支持字级别音素边界接口的发音人才支持字级时间戳。

  • 短文本语音合成 RESTful API 不返回时间戳信息。获取时间戳需使用 WebSocket 接口或相应 SDK。

  • 字幕文本按发音生成,不一定与原文逐字对应。例如,12 可以对应“十”和“二”两个字幕单元。显示字幕时使用原始文本,并根据返回时间确定句首、句尾或高亮位置,不要直接将字幕索引作为原始字符串的字符偏移。

代码示例

以下 Java 示例使用 nls-sdk-tts 2.1.6,从环境变量读取项目 AppKey 和有效的 NLS Token,输出字幕信息,并将音频保存为 tts_test.wav。文件已存在时会被覆盖。

添加依赖和配置凭证

在 Maven 项目的 dependencies 中添加以下依赖:

<dependency>
    <groupId>com.alibaba.nls</groupId>
    <artifactId>nls-sdk-tts</artifactId>
    <version>2.1.6</version>
</dependency>

在运行环境中设置 NLS_APP_KEY 和 NLS_TOKEN,分别填写项目 AppKey 和 NLS Token。Token 必须有效且与项目所属账号匹配,不使用 STS Token 或其他产品的 API Key 代替。

合成并接收时间戳

import com.alibaba.fastjson.JSONArray;
import com.alibaba.nls.client.protocol.NlsClient;
import com.alibaba.nls.client.protocol.OutputFormatEnum;
import com.alibaba.nls.client.protocol.SampleRateEnum;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizer;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizerListener;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizerResponse;
import java.io.IOException;
import java.io.OutputStream;
import java.nio.ByteBuffer;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicReference;

public class SpeechSynthesizerDemo {
    private static String requiredEnv(String name) {
        String value = System.getenv(name);
        if (value == null || value.trim().isEmpty()) {
            throw new IllegalArgumentException("Missing environment variable: " + name);
        }
        return value;
    }

    public static void main(String[] args) throws Exception {
        String appKey = requiredEnv("NLS_APP_KEY");
        String token = requiredEnv("NLS_TOKEN");
        String endpoint = "wss://nls-gateway-ap-southeast-1.aliyuncs.com/ws/v1";
        NlsClient client = new NlsClient(endpoint, token);

        try (OutputStream audio = Files.newOutputStream(Paths.get("tts_test.wav"))) {
            AtomicReference<String> failure = new AtomicReference<>();
            AtomicBoolean completed = new AtomicBoolean(false);
            AtomicBoolean firstAudio = new AtomicBoolean(true);
            final long start = System.nanoTime();
            SpeechSynthesizerListener listener = new SpeechSynthesizerListener() {
                @Override
                public void onMessage(ByteBuffer message) {
                    if (firstAudio.compareAndSet(true, false)) {
                        System.out.println("First audio latency (ms): "
                                + (System.nanoTime() - start) / 1_000_000);
                    }
                    byte[] bytes = new byte[message.remaining()];
                    message.get(bytes);
                    try {
                        audio.write(bytes);
                    } catch (IOException e) {
                        failure.compareAndSet(null, "Failed to write audio: " + e.getMessage());
                    }
                }

                @Override
                public void onMetaInfo(SpeechSynthesizerResponse response) {
                    JSONArray subtitles = (JSONArray) response.getObject("subtitles");
                    if (subtitles != null) {
                        System.out.println("MetaInfo: " + subtitles.toJSONString());
                    }
                }

                @Override
                public void onComplete(SpeechSynthesizerResponse response) {
                    completed.set(true);
                    System.out.println("SynthesisCompleted: " + response.getTaskId());
                }

                @Override
                public void onFail(SpeechSynthesizerResponse response) {
                    failure.set("task_id=" + response.getTaskId()
                            + ", status=" + response.getStatus()
                            + ", status_text=" + response.getStatusText());
                }
            };

            SpeechSynthesizer synthesizer = new SpeechSynthesizer(client, listener);
            try {
                synthesizer.setAppKey(appKey);
                synthesizer.setFormat(OutputFormatEnum.WAV);
                synthesizer.setSampleRate(SampleRateEnum.SAMPLE_RATE_16K);
                synthesizer.setVoice("siyue");
                synthesizer.setPitchRate(100);
                synthesizer.setSpeechRate(100);
                synthesizer.setText("Hello world. I have 12 apples.");
                synthesizer.addCustomedParam("enable_subtitle", true);
                synthesizer.start();
                synthesizer.waitForComplete();
                if (failure.get() != null) {
                    throw new IllegalStateException(failure.get());
                }
                if (!completed.get()) {
                    throw new IllegalStateException("Synthesis did not complete");
                }
                System.out.println("Audio saved to tts_test.wav");
            } finally {
                synthesizer.close();
            }
        } finally {
            client.shutdown();
        }
    }
}

合成成功后,程序输出 MetaInfo 字幕数组、SynthesisCompleted 及音频文件保存提示。首包延迟从发起合成前开始计时,到收到第一包音频结束;它不同于接收完整音频所需的时间。

说明

示例将音频保存为文件。对实时性要求较高时,可在 onMessage 收到音频后边接收边播放,无需等到全部合成完成。

字级时间戳

参数设置

发起合成请求前,将 enable_subtitle 设置为 true。该功能默认关闭。Java SDK 设置方式如下:

synthesizer.addCustomedParam("enable_subtitle", true);

服务端响应

服务端通过 MetaInfo 事件返回时间戳,payload.subtitles 为字幕数组。Java SDK 通过 onMetaInfo 回调接收该事件。

subtitles 元素字段

字段

类型

说明

text

String

按发音生成的文本单元,如一个汉字或英文单词。

begin_time

Integer

该文本单元在合成音频中的开始时间,单位为毫秒。

end_time

Integer

该文本单元在合成音频中的结束时间,单位为毫秒。

phoneme

String

仅开启字级时间戳时,返回字符串 "null",而不是 JSON null。

begin_index

Integer

文本单元的起始索引,从 0 开始,不是原始字符串的字符偏移。

end_index

Integer

文本单元的结束索引,不包含该位置。

例如,Hello world! 的两个单词分别对应索引区间 [0, 1) 和 [1, 2);标点可能占用索引位置,你好,世界。 中“世”的起始索引为 3。

说明

一次合成可能返回多个 MetaInfo 事件,其中的字幕可能重复或包含此前返回的内容。处理字幕时,根据索引和时间更新已有字幕数据,不要直接追加每次返回的整个数组。

返回示例

以下为合成 Hello world! 时返回的 payload 示例,时间值随音色、语速和文本变化。

{
  "subtitles": [
    {
      "text": "Hello",
      "phoneme": "null",
      "begin_index": 0,
      "end_index": 1,
      "begin_time": 0,
      "end_time": 328
    },
    {
      "text": "world",
      "phoneme": "null",
      "begin_index": 1,
      "end_index": 2,
      "begin_time": 328,
      "end_time": 775
    }
  ]
}