
Qwen-Audio-3.0-TTS is our latest text-to-speech model release. It ships as two variants from the same lineage:
API: Model Studio
Release Blog: Blog
This release focuses on four things developers actually run into in production: broader language coverage, natural-language style control, fine-grained tag control, and robustness when the reference audio isn’t clean.
Qwen-Audio-3.0-TTS-Plus currently ranks #1 on Artificial Analysis, the independent third-party TTS leaderboard.

Here’s what changed and what the numbers look like.
Qwen-Audio-3.0-TTS was optimized across English, Chinese, Japanese, Korean, German, and 16 languages total, plus improved fidelity on several Chinese dialects.
Supported languages (16): Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, Vietnamese.
*Full language support rolling out soon.


You can describe the delivery you want in natural language instead of hand-tuning acoustic parameters.
Simple prompts with plain language steer emotion, role, scenario, and pace without any labeling expertise.
When you need precise control over the non-verbal details — a breath, a laugh, a shift in tone — you can embed inline tags directly in the target text, like [gasp], [giggles], or [angry].
This makes the model useful for narration, games, and dubbing where the non-verbal cues carry as much as the words.
Reference clips from the real world are rarely studio-clean. Qwen-Audio-3.0-TTS was trained with targeted acoustic simulation so speech enhancement is built into the cloning path. The model suppresses reverb and noise while preserving timbre.
In our high-noise and high-reverb tests, this release produced noticeably cleaner output than previous versions from the same degraded references.
Qwen-Audio-3.0-TTS is available now. Grab the model here:
If you build something with it, we’d genuinely like to hear what worked and what didn’t — the failure cases are where the next version comes from.
1,466 posts | 503 followers
FollowAlibaba Cloud Community - November 24, 2025
Alibaba Cloud Community - November 20, 2024
Alex - August 14, 2018
Alibaba Cloud Big Data and AI - June 3, 2026
PM - C2C_Yuan - June 3, 2024
Farruh - June 23, 2025
1,466 posts | 503 followers
Follow
Alibaba Cloud Model Studio
A one-stop generative AI platform to build intelligent applications that understand your business, based on Qwen model series such as Qwen-Max and other popular models
Learn More
Intelligent Speech Interaction
Intelligent Speech Interaction is developed based on state-of-the-art technologies such as speech recognition, speech synthesis, and natural language understanding.
Learn More
Qwen
Full-range, open-source, multimodal, and multi-functional
Learn More
AI Acceleration Solution
Accelerate AI-driven business and AI model training and inference with Alibaba Cloud GPU technology
Learn MoreMore Posts by Alibaba Cloud Community