语音合成时间戳标记汉字或英文单词在合成音频中的起止时间,可用于视频字幕同步、文字高亮和虚拟人口型同步。通过请求参数开启该功能后,在接收音频流的同时获取时间戳。
使用限制
-
只有支持字级别音素边界接口的发音人才支持字级时间戳。
-
短文本语音合成 RESTful API 不返回时间戳信息。获取时间戳需使用 WebSocket 接口或相应 SDK。
-
字幕文本按发音生成,不一定与原文逐字对应。例如,
12可以对应“十”和“二”两个字幕单元。显示字幕时使用原始文本,并根据返回时间确定句首、句尾或高亮位置,不要直接将字幕索引作为原始字符串的字符偏移。
代码示例
以下 Java 示例使用 nls-sdk-tts 2.1.6,从环境变量读取项目 AppKey 和有效的 NLS Token,输出字幕信息,并将音频保存为 tts_test.wav。文件已存在时会被覆盖。
添加依赖和配置凭证
在 Maven 项目的 dependencies 中添加以下依赖:
<dependency>
<groupId>com.alibaba.nls</groupId>
<artifactId>nls-sdk-tts</artifactId>
<version>2.1.6</version>
</dependency>
在运行环境中设置 NLS_APP_KEY 和 NLS_TOKEN,分别填写项目 AppKey 和 NLS Token。Token 必须有效且与项目所属账号匹配,不使用 STS Token 或其他产品的 API Key 代替。
合成并接收时间戳
import com.alibaba.fastjson.JSONArray;
import com.alibaba.nls.client.protocol.NlsClient;
import com.alibaba.nls.client.protocol.OutputFormatEnum;
import com.alibaba.nls.client.protocol.SampleRateEnum;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizer;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizerListener;
import com.alibaba.nls.client.protocol.tts.SpeechSynthesizerResponse;
import java.io.IOException;
import java.io.OutputStream;
import java.nio.ByteBuffer;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicReference;
public class SpeechSynthesizerDemo {
private static String requiredEnv(String name) {
String value = System.getenv(name);
if (value == null || value.trim().isEmpty()) {
throw new IllegalArgumentException("Missing environment variable: " + name);
}
return value;
}
public static void main(String[] args) throws Exception {
String appKey = requiredEnv("NLS_APP_KEY");
String token = requiredEnv("NLS_TOKEN");
String endpoint = "wss://nls-gateway-ap-southeast-1.aliyuncs.com/ws/v1";
NlsClient client = new NlsClient(endpoint, token);
try (OutputStream audio = Files.newOutputStream(Paths.get("tts_test.wav"))) {
AtomicReference<String> failure = new AtomicReference<>();
AtomicBoolean completed = new AtomicBoolean(false);
AtomicBoolean firstAudio = new AtomicBoolean(true);
final long start = System.nanoTime();
SpeechSynthesizerListener listener = new SpeechSynthesizerListener() {
@Override
public void onMessage(ByteBuffer message) {
if (firstAudio.compareAndSet(true, false)) {
System.out.println("First audio latency (ms): "
+ (System.nanoTime() - start) / 1_000_000);
}
byte[] bytes = new byte[message.remaining()];
message.get(bytes);
try {
audio.write(bytes);
} catch (IOException e) {
failure.compareAndSet(null, "Failed to write audio: " + e.getMessage());
}
}
@Override
public void onMetaInfo(SpeechSynthesizerResponse response) {
JSONArray subtitles = (JSONArray) response.getObject("subtitles");
if (subtitles != null) {
System.out.println("MetaInfo: " + subtitles.toJSONString());
}
}
@Override
public void onComplete(SpeechSynthesizerResponse response) {
completed.set(true);
System.out.println("SynthesisCompleted: " + response.getTaskId());
}
@Override
public void onFail(SpeechSynthesizerResponse response) {
failure.set("task_id=" + response.getTaskId()
+ ", status=" + response.getStatus()
+ ", status_text=" + response.getStatusText());
}
};
SpeechSynthesizer synthesizer = new SpeechSynthesizer(client, listener);
try {
synthesizer.setAppKey(appKey);
synthesizer.setFormat(OutputFormatEnum.WAV);
synthesizer.setSampleRate(SampleRateEnum.SAMPLE_RATE_16K);
synthesizer.setVoice("siyue");
synthesizer.setPitchRate(100);
synthesizer.setSpeechRate(100);
synthesizer.setText("Hello world. I have 12 apples.");
synthesizer.addCustomedParam("enable_subtitle", true);
synthesizer.start();
synthesizer.waitForComplete();
if (failure.get() != null) {
throw new IllegalStateException(failure.get());
}
if (!completed.get()) {
throw new IllegalStateException("Synthesis did not complete");
}
System.out.println("Audio saved to tts_test.wav");
} finally {
synthesizer.close();
}
} finally {
client.shutdown();
}
}
}
合成成功后,程序输出 MetaInfo 字幕数组、SynthesisCompleted 及音频文件保存提示。首包延迟从发起合成前开始计时,到收到第一包音频结束;它不同于接收完整音频所需的时间。
示例将音频保存为文件。对实时性要求较高时,可在 onMessage 收到音频后边接收边播放,无需等到全部合成完成。
字级时间戳
参数设置
发起合成请求前,将 enable_subtitle 设置为 true。该功能默认关闭。Java SDK 设置方式如下:
synthesizer.addCustomedParam("enable_subtitle", true);
服务端响应
服务端通过 MetaInfo 事件返回时间戳,payload.subtitles 为字幕数组。Java SDK 通过 onMetaInfo 回调接收该事件。
subtitles 元素字段
|
字段 |
类型 |
说明 |
|
text |
String |
按发音生成的文本单元,如一个汉字或英文单词。 |
|
begin_time |
Integer |
该文本单元在合成音频中的开始时间,单位为毫秒。 |
|
end_time |
Integer |
该文本单元在合成音频中的结束时间,单位为毫秒。 |
|
phoneme |
String |
仅开启字级时间戳时,返回字符串 |
|
begin_index |
Integer |
文本单元的起始索引,从 0 开始,不是原始字符串的字符偏移。 |
|
end_index |
Integer |
文本单元的结束索引,不包含该位置。 |
例如,Hello world! 的两个单词分别对应索引区间 [0, 1) 和 [1, 2);标点可能占用索引位置,你好,世界。 中“世”的起始索引为 3。
一次合成可能返回多个 MetaInfo 事件,其中的字幕可能重复或包含此前返回的内容。处理字幕时,根据索引和时间更新已有字幕数据,不要直接追加每次返回的整个数组。
返回示例
以下为合成 Hello world! 时返回的 payload 示例,时间值随音色、语速和文本变化。
{
"subtitles": [
{
"text": "Hello",
"phoneme": "null",
"begin_index": 0,
"end_index": 1,
"begin_time": 0,
"end_time": 328
},
{
"text": "world",
"phoneme": "null",
"begin_index": 1,
"end_index": 2,
"begin_time": 328,
"end_time": 775
}
]
}