×
Community Blog Qwen-Audio-3.0-TTS: More Multilingual, Easier to Direct

Qwen-Audio-3.0-TTS: More Multilingual, Easier to Direct

Our text-to-speech model, now across 16 languages.

cover

Qwen-Audio-3.0-TTS is our latest text-to-speech model release. It ships as two variants from the same lineage:

  • Flash: tuned for real-time interaction, with a first-packet latency at 300ms-level.
  • Plus: tuned for high-quality generation, where naturalness and timbre fidelity matter more than speed.

API: Model Studio

Release Blog: Blog

This release focuses on four things developers actually run into in production: broader language coverage, natural-language style control, fine-grained tag control, and robustness when the reference audio isn’t clean.


Qwen-Audio-3.0-TTS-Plus currently ranks #1 on Artificial Analysis, the independent third-party TTS leaderboard.

1

Here’s what changed and what the numbers look like.

Multilingual Coverage Across 16 Languages

Qwen-Audio-3.0-TTS was optimized across English, Chinese, Japanese, Korean, German, and 16 languages total, plus improved fidelity on several Chinese dialects.

Supported languages (16): Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, Vietnamese.

*Full language support rolling out soon.

  • WER/CER (lower is better): Qwen-Audio-3.0-TTS family demonstrates the strongest overall multilingual intelligibility, achieving the best WER/CER in 10 of 16 languages. Flash delivers the lowest average WER/CER at 3.87, while Plus remains highly competitive at 3.96, both outperforming the other systems on average.

2

  • Speaker Similarity (higher is better): Qwen-Audio-3.0-TTS shows a clear and consistent advantage in speaker similarity. Plus ranks first across all 16 languages with an average SS of 82.75, while Flash follows at 80.44, demonstrating strong and robust voice-preservation quality across diverse languages.

3

Style Control In Natural Language

You can describe the delivery you want in natural language instead of hand-tuning acoustic parameters.

Simple prompts with plain language steer emotion, role, scenario, and pace without any labeling expertise.

Fine-Grained Tags For Non-Verbal Details

When you need precise control over the non-verbal details — a breath, a laugh, a shift in tone — you can embed inline tags directly in the target text, like [gasp], [giggles], or [angry].

This makes the model useful for narration, games, and dubbing where the non-verbal cues carry as much as the words.

More Robust Voice Cloning From Imperfect Audio

Reference clips from the real world are rarely studio-clean. Qwen-Audio-3.0-TTS was trained with targeted acoustic simulation so speech enhancement is built into the cloning path. The model suppresses reverb and noise while preserving timbre.

In our high-noise and high-reverb tests, this release produced noticeably cleaner output than previous versions from the same degraded references.

Also In This Release

  • A curated preset voice library spanning 16 supported languages, so you can ship a voice without cloning one first.
  • 48 kHz audio output (coming soon).

Try It

Qwen-Audio-3.0-TTS is available now. Grab the model here:

If you build something with it, we’d genuinely like to hear what worked and what didn’t — the failure cases are where the next version comes from.

0 0 0
Share on

Alibaba Cloud Community

1,466 posts | 503 followers

You may also like

Comments

Alibaba Cloud Community

1,466 posts | 503 followers

Related Products