The following tables list model releases . For model deprecation rules and lists, see Model decommissioning policy .
Singapore
Model type | Date | Service scope | Model ID | Description |
|---|---|---|---|---|
Image translation | 2026-08-28 | International |
| An advanced image translation engine supporting 55 source languages. It delivers pixel-perfect replication of the original layout and typography. Gain full control over professional content with custom terminology and sensitive word filtering, ensuring a flexible, accurate, and high-performance image localization service. |
Text generation, Deep thinking, Visual understanding | 2026-08-26 | International |
| Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications. |
Video generation | 2026-08-20 | International |
| Wan3.0-Video-Prime is the high-speed version of the Wan3.0 video generation model, with capabilities aligned with the Wan3.0-Video standard version. It supports four-modal all-reference input and can generate videos up to 30 seconds long, delivering an immersive audio-visual experience with significantly improved end-to-end speed. |
Text generation, Reasoning, Visual understanding | 2026-08-19 | International |
| Kimi K3 is Kimi's most powerful flagship model with 2.8 trillion parameters. Built on KDA hybrid linear attention (Kimi Delta Attention) and Attention Residuals technologies, it natively supports vision understanding and features a 1 million token context window. It is the world's first open-source 3-trillion-parameter model designed for long-context programming, knowledge work, and advanced reasoning. |
Text generation, Reasoning, Visual understanding | 2026-08-17 | International |
| A 27B native vision-language Dense model of the Qwen3.8 series. Compared with 3.6-27B, it focuses on improving coding and office-scenario capabilities under text and visual modalities, able to more reliably complete complex tasks end-to-end and deliver trustworthy results. |
Text generation, Reasoning | 2026-08-17 | International |
| GLM-5.3 is Zhipu's strongest coding model to date, improving 50% over GLM-5.2 in internal perceptual evaluations. In cybersecurity, GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery, demonstrating strong potential for cybersecurity defense scenarios. |
Text generation, Reasoning | 2026-08-14 | International |
| A flagship Mixture-of-Experts (MoE) large language model with 1.6 trillion total parameters and 49 billion activated parameters. It natively supports context windows of up to 1 million tokens. Trained on extensive high-quality data, the model delivers strong performance in mathematical and logical reasoning, complex reasoning, professional code generation, and in-depth long-document analysis, and is suitable for demanding scenarios such as advanced scientific research, complex enterprise workflows, and sophisticated agentic applications. |
Text generation | 2026-08-12 | International |
| Qwen3.8-2.4T-A95B is the open-source version of Qwen's latest flagship series, released in August 2026. It adopts a sparse MoE architecture with 2.4 trillion total parameters and about 95 billion activated per step, combined with a hybrid attention mechanism, supporting a 1 million token context. Key benchmarks: GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, BabyVision 82.0, CodeArena global #4. |
Realtime chat | 2026-08-10 | International |
| Qwen 3.0 Realtime Speech Model Standard Edition - next-generation duplex speech model ranked #1 globally in Artificial Analysis Speech-to-Speech benchmark. Balances model intelligence with duplex dialogue rhythm for natural interaction and enhanced response quality. |
Realtime chat | 2026-08-10 | International |
| Qwen 3.0 Realtime Speech Dialogue Model Flash Edition - a next-generation duplex speech model ranked #1 globally in Artificial Analysis Speech-to-Speech benchmark. Combines high intelligence with optimized duplex rhythm for low-latency (parallel inference, full-streaming optimization) 'fast and smart' dialogue experience focusing on response speed. |
Text generation, Reasoning, Visual understanding | 2026-08-07 | International |
| Kimi K3 is Kimi's most capable flagship model to date, featuring 2.8 trillion parameters and built upon the Kimi Delta Attention (KDA) hybrid linear attention mechanism and Attention Residuals technology. It natively supports visual understanding and offers a context window of up to 1 million tokens. As the world's first open-source model at the 3-trillion-parameter scale, it is specifically designed for cutting-edge intelligent applications such as long-horizon programming, knowledge work, and reasoning. |
Text generation, Reasoning | 2026-08-07 | International |
| Direct from Zhipu AI. GLM-5.2 supports a truly usable 1M context window and maintains its leading position in long-context tasks. |
Video generation | 2026-08-06 | International |
| Wan3.0 is an all-in-one video generation model that uniformly supports multiple creative functions including reference, editing, replication, and driving. It supports four-modal all-reference input, up to 30 seconds of video generation, and can parse files, web pages, and complex images. With production-grade character consistency and realistic audio and visuals, it delivers an immersive audio-visual impact. |
Image generation | 2026-08-04 | International |
| Clear instruction understanding: supports up to 4.5k token input, generating complex image-text instructions in one pass. Stable text rendering: 10px small text, 12 languages, 20+ fonts clearly legible, ready for infographics and interfaces. Handy batch output: posters, web pages, interfaces and other daily tasks can be generated in bulk at lower cost, ideal for continuous creation. Qwen-Image-3.0 Standard pursues not only generation quality but also everyday creative convenience, making image generation a sustainable content productivity tool. |
Text generation, Reasoning, Visual understanding | 2026-08-02 | International |
| A 2.4-trillion-parameter MoE flagship with a comprehensive leap in coding and office capabilities, capable of autonomously coding for days to deliver complete projects. It handles hundreds of professional tasks such as legal, financial, and design work, delivering production-grade results end-to-end in a single conversation. Native visual understanding runs through the full planning, execution, and verification pipeline, supporting deep semantic parsing of ultra-long documents and long videos. It performs autonomous planning and closed-loop iteration in long-horizon tasks, continuously evolving. |
Text generation, Reasoning | 2026-08-01 | International |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Text embedding | 2026-07-31 | International |
| Qwen3-7-Text-Embedding is a multilingual text vector model based on Qwen3.7, offering significant improvements in text retrieval, clustering, and classification over text-embedding-v4. It achieves 20% better performance in MTEB multilingual, Chinese-English, and code retrieval tasks and supports customizable vector dimensions (256-2560). |
Speech recognition | 2026-07-30 | International |
| Added the Qwen-Audio-3.0-ASR-Flash-Streaming (real-time), Qwen-Audio-3.0-ASR-Flash-Filetrans (non-real-time), and Qwen-Audio-3.0-ASR-Flash (non-real-time) models: Dialect support: Supports the seven major Chinese dialect groups (Mandarin, Wu, Xiang, Gan, Hakka, Min, and Yue) and more than 20 regional accents; Classical poetry optimization: Improves recognition accuracy for classical Chinese poetry, making it suitable for education, culture, and audiobook scenarios; Text optimization: Enhances punctuation prediction and text normalization, automatically converting numbers, dates, and monetary amounts to standard formats; Multilingual expansion: Supports 30 languages, including Chinese, English, Japanese, and Korean; Hotwords and context: Supports hotwords (precompiled and on-the-fly) and context input to improve recognition accuracy for domain-specific terms. |
Text generation, Reasoning, Visual understanding | 2026-07-21 | International |
| The Qwen3.7 native vision-language series Flash models comprehensively enhance multimodal understanding and Agent execution capabilities compared to 3.6-Flash. Key improvements include strengthened foundational multimodal abilities, enhanced object recognition, improved real-world perception and spatial intelligence. Multimodal Agent scenarios such as Search Agent and CI Agent have seen significant upgrades, with more stable end-to-end task execution. Multimodal coding capabilities are optimized, delivering a smoother vibe coding experience. |
Image generation | 2026-07-20 | International |
| Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and 20+ fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool. |
Speech synthesis, Realtime speech synthesis | 2026-07-14 | International |
| Qwen-Audio-3.0-TTS-Plus is a high-performance speech synthesis model, designed for high-quality speech generation scenarios. Compared with the previous version, it supports more low-resource languages and Chinese dialects, significantly improves dialect authenticity, and enhances free-style instruction following and fine-grained tag control for more accurate control over emotion, tone, character, speaking rate, volume, and synthesis style. It is also more robust under noisy and reverberant acoustic conditions, with further improvements in audio quality, clarity, resolution, and overall expressiveness. The Plus version focuses more on synthesis quality and detailed expressiveness, making it suitable for professional scenarios with higher requirements for audio quality, naturalness, and expressiveness, such as content creation, audiobooks, film and video dubbing, brand voice design, and premium speech services. |
Realtime speech synthesis | 2026-07-14 | International |
| qwen-audio-3.0-tts-flash is a high-performance speech synthesis model, optimized for real-time interactive scenarios. Compared with the previous version, it supports more low-resource languages and Chinese dialects, improves dialect authenticity, and enhances free-style instruction following and fine-grained tag control for more flexible control over emotion, tone, character, speaking rate, volume, and expressive style. It is also more robust under noisy and reverberant acoustic conditions, with improved audio quality, clarity, and overall expressiveness. The Flash version focuses on real-time synthesis, with first-packet latency controlled within 200 ms, making it suitable for voice assistants, real-time dialogue, intelligent customer service, and other low-latency interactive applications. |
Text generation | 2026-07-10 | International |
| GLM-5.2-Fast-Preview is the high-speed variant of Zhipu AI's GLM-5.2, with 1M context and capabilities on par with the standard version. Inference-optimized to deliver 1.5–2× the output TPS, it fits latency-sensitive use cases such as real-time chat, multi-turn agents, and streaming code generation. |
Video generation | 2026-07-01 | International |
| Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of June 12, 2026. |
Video generation | 2026-07-01 | International |
| Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. This version is a snapshot as of June 12, 2026. |
Image generation | 2026-06-25 | International |
| The Qwen-Image-2.0 series full-fledged model integrates image generation and editing; it boasts more professional text rendering capabilities with 1k token command support, more delicate and realistic textures, meticulous depiction of realistic scenes, and stronger semantic adherence. The full-fledged version possesses the strongest text rendering capabilities and realistic textures in the 2.0 series. |
Text generation, Reasoning, Visual understanding | 2026-06-25 | International |
| kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
Text generation, Reasoning | 2026-06-25 | International |
| GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
Speech recognition | 2026-06-17 | International |
| The Bailing ASR version, updated in June 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. It supports contextualization capabilities and can transcribe audio up to 5 minutes in length. |
Video generation | 2026-06-16 | International |
| HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
Video generation | 2026-06-16 | International |
| HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
Video generation | 2026-06-16 | International |
| HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
Text generation, Reasoning, Visual understanding | 2026-06-10 | International |
| The Max model, the largest and most capable in the Qwen3.7 series, has added visual‑modal understanding compared to the May 20 snapshot, enabling it to perceive real‑world scenes and supporting multimodal interactive hybrid agent capabilities. This version is based on a snapshot taken on June 8, 2026. |
Text generation, Reasoning, Visual understanding | 2026-06-01 | International |
| Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
Text generation, Reasoning | 2026-05-21 | International |
| The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
Text generation, Reasoning | 2026-05-20 | International |
| The Max model, the largest and most capable variant in the Qwen3.7 series, is available as a preview and supports only thinking mode, offering a pure text‑only interface for experimentation. It is primarily optimized for general‑purpose conversational use cases, such as knowledge‑based question answering, instruction following, and creative writing. |
Realtime speech translation | 2026-05-19 | International |
| The real-time version of Qwen3.5-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3.5-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3.5-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 60 languages and speak 29 languages. |
Text generation, Reasoning | 2026-05-11 | International |
| A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
Text generation, Reasoning | 2026-05-11 | International |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Video generation | 2026-04-26 | International |
| Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of April 25, 2026. |
Video generation | 2026-04-26 | International |
| Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of April 25, 2026. |
Video generation | 2026-04-26 | International |
| HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
Video generation | 2026-04-26 | International |
| HappyHorse-1.0-R2V supports Reference-to-Video generation, offering enhanced stability in subject and scene referencing. Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance. |
Text generation, Reasoning, Visual understanding | 2026-04-23 | International |
| The Qwen3.5 native vision-language series Plus model has seen a substantial improvement in agentic coding capabilities compared to the February 15th snapshot. Inference speed has also been significantly enhanced, while its knowledge retention, reasoning ability, and long-context processing remain at a high level, making it well-suited for complex agent-based tasks. It is ideal for applications such as coding agents, production workflows, and high-throughput scenarios. This version is based on a snapshot taken on April 20, 2026. |
Image generation | 2026-04-23 | International |
| The full-featured Qwen-Image-2.0 series models integrate image generation and image editing, offering enhanced text rendering with support for 1,000-token prompts, more refined realistic textures, detailed depiction of photorealistic scenes, and stronger semantic adherence. The full-featured version delivers the strongest text rendering and most lifelike textures in the 2.0 series. |
Text generation, Reasoning, Visual understanding | 2026-04-22 | International |
| The Qwen3.6 27B native vision-language dense model builds upon the 3.5-27B version, with key improvements in agentic coding capabilities and enhanced STEM reasoning and inference skills. In the vision modality, it demonstrates significant advances in spatial intelligence, object localization, and detection, while video understanding, document OCR, and visual agent capabilities continue to improve steadily. |
Video generation | 2026-04-22 | International |
| HappyHorse-1.0-T2V supports text-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends text semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
Video generation | 2026-04-22 | International |
| HappyHorse-1.0-I2V enables image-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends both text and image semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
Text generation, Reasoning, Visual understanding | 2026-04-17 | International |
| The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
Text generation, Reasoning, Visual understanding | 2026-04-17 | International |
| The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance. |
Speech synthesis | 2026-04-15 | International |
| A large-model voice replication service used in conjunction with Cosyvoice-v3. Utilizing advanced large-model technology for feature extraction, it can replicate voices without a training process. Only a very short audio clip is required to quickly generate a highly similar and natural-sounding custom voice. |
Text generation, Reasoning | 2026-04-14 | International |
| The Max model, the largest and most capable variant in the Qwen3.6 series, is now available in a preview version. At present, only its plain-text capabilities are open for experimentation. Compared with the previously released Qwen3-Max and Qwen3.6-Plus, this model features enhanced vibe coding abilities, more efficient coding agent execution, and significantly improved front-end development skills. Additionally, its long-tail knowledge retention has been further upgraded. |
Text generation, Reasoning | 2026-04-14 | International |
| GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
Video generation | 2026-04-03 | International |
| Wan2.7 video edit, supports both localized and global editing with prompt. Seamlessly replace elements using image references and replicate complex dynamic processes, including motion, special effects, and camera movements. |
Video generation | 2026-04-03 | International |
| Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
Video generation | 2026-04-03 | International |
| Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. |
Video generation | 2026-04-03 | International |
| Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
Image generation | 2026-04-01 | International |
| Wan2.7–image-pro, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
Image generation | 2026-04-01 | International |
| Wan2.7 – image generation and editing, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
Text generation, Reasoning, Visual understanding | 2026-04-01 | International |
| The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |
Realtime omni-modal | 2026-03-26 | International |
| Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience. |
Omni-modal | 2026-03-26 | International |
| Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience. |
Realtime omni-modal | 2026-03-26 | International |
| Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience. |
Omni-modal | 2026-03-26 | International |
| Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience. |
Realtime omni-modal | 2026-03-25 | International |
| Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience.This version is a snapshot from March 15, 2026. |
Text generation, Reasoning | 2026-03-20 | International |
| DeepSeek-V3.2 is the official release of a model that incorporates DeepSeek Sparse Attention—a sparse attention mechanism. It's also the first model launched by DeepSeek that integrates reasoning into tool usage, supporting both reasoning-enabled and non-reasoning tool calls. |
Image generation | 2026-03-03 | International |
| The full-featured Qwen-Image-2.0 series models integrate image generation and image editing, offering enhanced text rendering with support for 1,000-token prompts, more refined realistic textures, detailed depiction of photorealistic scenes, and stronger semantic adherence. The full-featured version delivers the strongest text rendering and most lifelike textures in the 2.0 series.This version is a snapshot as of March 3, 2026. |
Image generation | 2026-03-03 | International |
| The Qwen-Image-2.0 series of accelerated models integrates image generation and image editing, offering enhanced text-rendering capabilities with support for 1,000-token prompts, more realistic textures, finely detailed photorealistic scenes, and improved semantic consistency. The accelerated version effectively strikes an optimal balance between model performance and quality. |
Speech recognition | 2026-03-02 | International |
| Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in multiple languages, while ensuring precise transcription even in complex audio environments.This version is a snapshot dated February 10, 2026. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | International |
| The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | International |
| The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | International |
| The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | International |
| The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B. |
Text generation | 2026-02-20 | International |
| The new-generation code generation model in the Qwen3 series delivers performance close to that of Qwen3-Coder-Plus while offering even better capabilities. The model has been optimized with a focus on repository-level understanding, supports multi-turn tool interactions, and enhances its compatibility with agentic coding tools. |
Text generation, Reasoning, Visual understanding | 2026-02-15 | International |
| The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities. |
Text generation, Reasoning, Visual understanding | 2026-02-15 | International |
| The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent. |
Speech recognition | 2026-02-13 | International |
| The real-time version of Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in multiple languages, while ensuring precise transcription even in complex audio environments.This version is a snapshot dated February 10, 2026. |
Speech synthesis | 2026-02-10 | International |
| Qwen3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 26, 2026. |
Speech synthesis | 2026-02-10 | International |
| Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated from the qwen-voice-enrollment service, and supports speech output in 11 languages using the same voice timbre. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 22, 2026. |
Speech synthesis | 2026-02-10 | International |
| Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments. |
Speech synthesis | 2026-02-10 | International |
| Cloning capability: CosyVoice-v3-plus is the latest large voice cloning model in the CosyVoice series from Tongyi Lab. It offers superior sound quality and cloning fidelity, ideal for professional scenarios. With just 5-20 seconds of reference audio, it can rapidly generate a highly similar and natural-sounding custom voice. Synthesis capability: CosyVoice-v3-plus is the latest large speech synthesis model in the CosyVoice series from Tongyi Lab. It features enhanced sound quality and expressiveness, ideal for professional scenarios. The model supports real-time, streaming text-to-speech synthesis. |
Speech synthesis | 2026-02-09 | International |
| Synthesis Capabilities: CosyVoice-v3-Flash is the latest high-performance speech synthesis model in the CosyVoice series from Tongyi Labs, offering improved naturalness, timbre, prosody, and emotional expressiveness compared to previous versions. This model supports real-time streaming text-to-speech synthesis. Cloning Capabilities: CosyVoice-v3-Flash is also the latest speech cloning model in the CosyVoice series from Tongyi Labs. Compared to previous versions, it improves pronunciation accuracy and timbre similarity, and adds support for more less commonly spoken languages (German, Spanish, French, Italian, Russian, Japanese). It can quickly generate highly similar and naturally sounding custom voices from just 5-20 seconds of reference audio. |
Text generation | 2026-01-30 | International |
| The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
Video generation | 2026-01-29 | International |
| Wan2.6 reference to video flash, faster and more cost-effective generation. Supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
Text embedding | 2026-01-27 | International |
| A text-ranking model trained on the Qwen LLM foundation performs relevance ranking for input queries and candidate documents. It supports over 100 languages and long-text inputs, and is suitable for applications such as text retrieval and RAG. Its performance is aligned with the open-source Qwen3-Rerank series models. |
Text generation, Reasoning | 2026-01-23 | International |
| Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model's overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026. |
Visual understanding | 2026-01-22 | International |
| The Qwen3 series of small-sized visual understanding models effectively integrates thinking and non-thinking modes. Compared with the snapshot taken on October 15, 2025, the overall performance of the model has improved significantly: it delivers enhanced capabilities in general visual recognition and reasoning, and shows marked improvements in recognition accuracy across various business scenarios such as security, in-store inspections, equipment monitoring, and photo-based problem solving. This version is a snapshot as of January 22, 2026. |
Speech synthesis | 2026-01-21 | International |
| qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments. This model is equivalent to the snapshot version released on January 22, 2026. |
Video generation | 2026-01-15 | International |
| Wan2.6 image to video flash, faster and more cost-effective generation. Intelligent shot scheduling enables multi‑camera storytelling, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
Image generation | 2026-01-15 | International |
| The Max series Qwen's image editing models delivers more stable and versatile editing capabilities: enhanced industrial design and geometric reasoning, improved character consistency, reduced offset issues, and integrated LoRA capabilities for a wider range of image editing functions. |
Speech synthesis | 2026-01-14 | International |
| Qwen3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 15, 2026. |
Speech synthesis | 2026-01-14 | International |
| Qwen 3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated from the qwen-voice-enrollment service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 15, 2026. |
Text generation | 2026-01-13 | International |
| The Qwen Role-Playing Model Series is specifically optimized for muti-language anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. This version significantly enhances Japanese linguistic localization (including dialects and honorifics), human-like role-playing authenticity, narrative coherence control, and scenario-based cognitive intelligence. |
Image generation | 2026-01-09 | International |
| The Qwen series of image-generation models boasts exceptional text-rendering capabilities and excels in complex text rendering as well as a wide range of generation and editing tasks. This version, a snapshot taken on January 9, 2026, is a distilled and accelerated variant of Qwen-Image-Max, enabling faster generation of high-quality images. |
Image generation | 2025-12-30 | International |
| The Max series of qwen's image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the "AI-like" feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering. |
Image generation | 2025-12-22 | International |
| Z-Image-Turbo is a highly efficient image-generation model that has topped the Artificial Analysis benchmark as the world's No. 1 open-source text-to-image model. With just 6 billion parameters and an 8-step inference process, it generates photo-realistic images comparable to those produced by large-scale commercial models, while excelling in bilingual Chinese–English text rendering, complex semantic understanding, and diverse thematic generation. |
Reasoning, Visual understanding | 2025-12-18 | International |
| The Qwen3 series of visual understanding models effectively integrates thinking and non-thinking modes. Compared to the snapshot released on September 23, this version delivers superior performance in reasoning and analysis tasks as well as style control, while also offering lower latency and faster response speeds. This version is based on a snapshot taken on December 19, 2025. |
Video generation | 2025-12-16 | International |
| Wan2.6 reference to video, supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
Image generation | 2025-12-15 | International |
| Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation. |
Image generation | 2025-12-15 | International |
| Wan2.6 Image, An all-round image generation model that supports joint text–image reasoning, multi-image creative fusion, commercial-grade consistency, aesthetic style transfer, and precise control of framing and lighting, significantly enhancing consistency, controllability, and expressiveness in image generation. |
Image generation | 2025-12-15 | International |
| The Qianwen series of Image Editing Plus models features enhanced character consistency, industrial design capabilities, and geometric reasoning abilities compared to the snapshot as of October 30. Additionally, it integrates LoRA capabilities such as lighting effects and effectively mitigates offset issues. This version is based on a snapshot taken on December 15, 2025. |
Speech synthesis | 2025-12-12 | International |
| Qwen 3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from December 16, 2025. |
Speech synthesis | 2025-12-12 | International |
| Qwen Voice-Design model is a series of voice design models from Qwen Speech Model. It only requires a simple text description to quickly design a suitable voice. When used in conjunction with the qwen3-tts-vd-realtime model, it can design and output speech in 10 languages. Furthermore, the synthesized audio can adaptively adjust its tone based on the text and has good processing capabilities for complex text synthesis. |
Realtime omni-modal | 2025-12-04 | International |
| The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
Omni-modal | 2025-12-04 | International |
| Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
Realtime speech recognition | 2025-12-04 | International |
| Qwen3-LiveTranslate-Flash is a high-precision, highly responsive, and robust multilingual real-time audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash provides both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, and also supports 8 Chinese dialects. |
Video generation | 2025-12-03 | International |
| Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
Video generation | 2025-12-03 | International |
| Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
— | 2025-12-01 | International |
| This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
Speech synthesis | 2025-11-27 | International |
| Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated by the qwen3-voice-enrollment service, and supports speech output in 11 languages with the same voice timbre. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust the tone according to the text, and it also has good processing capabilities for complex text synthesis.This model is provided as a snapshot version. |
Speech synthesis | 2025-11-27 | International |
| The Qwen3-TTS-Flash-Realtime model is Tongyi's latest real-time speech synthesis foundation model, featuring 17 expressive voices while delivering low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on extensive datasets, the system autonomously adjusts vocal tones based on text semantics and demonstrates robust capabilities for complex content synthesis. |
Speech synthesis | 2025-11-27 | International |
| The Qwen3-TTS-Flash is Tongyi's latest offline text-to-speech foundation model, featuring 17 expressive voices while enabling low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on massive datasets, the system automatically adjusts vocal tones based on text semantics and demonstrates robust capabilities for synthesizing complex content. |
Speech synthesis | 2025-11-27 | International |
| The Qwen Voice-Enrollment model is a series of voice replication models from the qwen speech model. It can quickly replicate highly similar voices using audio of only 5 seconds or more. When used in conjunction with the qwen3-tts-vc-realtime model, it can replicate a person's voice with high fidelity and output speech in 10 languages. Furthermore, the synthesized audio can adaptively adjust its tone according to the text and has good processing capabilities for complex text synthesis. |
Visual understanding | 2025-11-21 | International |
| This model is a snapshot version from November 20, 2025, and is based on the latest Qwen-VL3 architecture with a comprehensive upgrade. It features significant improvements in document parsing and text localization capabilities, as well as substantial reductions in end-to-end latency and illusions. |
Speech recognition | 2025-11-21 | International |
| The Fun ASR version, updated in April 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. This is a snapshot released on November 7, 2025. |
Visual understanding | 2025-11-20 | International |
| Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
Text generation | 2025-11-19 | International |
| Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
Speech recognition | 2025-11-18 | International |
| The large file transcription version of Qwen3-ASR-Flash. Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in multiple languages, ensuring precise transcription even in complex audio environments. |
Realtime speech recognition | 2025-11-17 | International |
| This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments.This is a snapshot released on November 7, 2025. |
Text generation | 2025-11-11 | International |
| Qwen-MT-Flash, a large language model from the Qwen series, has been fully upgraded with the Qwen 3 architecture for significantly enhanced performance and translation quality. It provides rapid, cost-effective translation across 92 languages, while supporting advanced features such as terminology intervention, format preservation, and domain-specific adaptation. It is the ideal choice for applications requiring a powerful balance of speed, quality, and cost. |
Video generation | 2025-11-10 | International |
| wan2.2-animate-move is a character animation generation model. Users simply upload a character photo and a reference performance video, and the model transfers the expressions and actions from the video onto the character in the image, producing a high-fidelity animated video. |
Video generation | 2025-11-10 | International |
| wan2.2-animate-mix is a character replacement model product. By uploading a character photo and a performance video, users can accurately replace the character in the original video with the character from the photo, while completely preserving environmental details such as the scene, lighting, and color tone of the original video. |
Multimodal embedding | 2025-10-31 | International |
| Embedding-Vision is a vision-centric multimodal embedding model powered by an LLM, featuring outstanding domain-specific performance and high cost-effectiveness in various domains (e.g., e-commerce, photo galleries, security, autonomous driving). With support for text, image, and video, it is applicable to downstream retrieval tasks, including text-to-image, image-to-image, text-to-video and video-to-video. |
Multimodal embedding | 2025-10-31 | International |
| Embedding-Vision is a vision-centric multimodal embedding model powered by an LLM, featuring outstanding domain-specific performance and high cost-effectiveness in various domains (e.g., e-commerce, photo galleries, security, autonomous driving). With support for text, image, and video, it is applicable to downstream retrieval tasks, including text-to-image, image-to-image, text-to-video and video-to-video. |
Image generation | 2025-10-31 | International |
| The qwen series of image editing Plus models further optimizes inference performance and system stability based on the initial Edit model, significantly reducing the response time for image generation and editing. It also supports returning multiple images in a single request, greatly enhancing user experience. |
Realtime speech recognition | 2025-10-29 | International |
| The real-time version of Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in 11 languages, while ensuring precise transcription even in complex audio environments. |
Reasoning, Visual understanding | 2025-10-21 | International |
| The largest dense model in the Qwen3-VL series, its reasoning version boasts multimodal reasoning capabilities second only to Qwen3-VL-235B-Thinking. It excels in STEM and math problem-solving, general image and video understanding, and achieves state-of-the-art performance in multimodal agent capabilities, making it ideal for complex multimodal reasoning tasks. |
Visual understanding | 2025-10-21 | International |
| The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios. |
Reasoning, Visual understanding | 2025-10-15 | International |
| The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
Reasoning, Visual understanding | 2025-10-03 | International |
| The Thinking version of the second-largest MoE model in the Qwen3-VL series features fast response speeds and enhanced multimodal understanding and reasoning capabilities, visual agents, and support for extremely long contexts such as lengthy videos and documents. It also boasts comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities, making it well-suited for complex real-world tasks. |
Visual understanding | 2025-10-03 | International |
| The Instruct version of the second-largest MoE model in the Qwen3-VL series offers rapid response speeds and supports extremely long contexts like lengthy videos and documents. It includes comprehensively upgraded image/video understanding, spatial awareness, and object recognition capabilities, as well as 2D/3D visual localization, enabling it to handle intricate real-world challenges. |
Reasoning, Visual understanding | 2025-09-30 | International |
| The Thinking version of the 8B Dense model in the Qwen3-VL series consumes less GPU memory and is capable of performing multimodal understanding and reasoning. It supports extremely long contexts such as lengthy videos and documents, 2D/3D visual localization, and features comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities. |
Visual understanding | 2025-09-30 | International |
| The Instruct version of the 8B Dense model in the Qwen3-VL series requires less GPU memory and provides comprehensively upgraded image/video understanding, support for extremely long contexts like lengthy videos and documents, spatial awareness, and object recognition capabilities, making it suitable for tackling complex real-world tasks. |
Speech recognition | 2025-09-25 | International |
| Fun's multilingual speech recognition model supports over 31 languages and allows for free language switching, making it the top choice for users expanding overseas, especially to Southeast Asia. Fun-asr is an upgraded version of this model; switching to Fun-asr is recommended. |
Image generation | 2025-09-24 | International |
| The upgraded Wan2.5 Preview text to image model, newly upgraded model architecture significantly enhances visual aesthetics, design sensibility, and realistic texture. It excels in precise instruction adherence, generates text proficiently in English, Chinese, and less common languages, and supports the generation of complex structured long texts, charts, and architectural diagrams. |
Image generation | 2025-09-24 | International |
| The upgraded Wan2.5 Preview image edit model, newly upgraded model architecture supports rich image editing capabilities via instruction control, with enhanced instruction adherence. It also enables multi-image reference generation with high consistency and demonstrates excellent text generation performance. |
Text generation, Reasoning | 2025-09-24 | International |
| The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
Realtime speech translation | 2025-09-24 | International |
| The real-time version of Qwen3-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, including 8 Chinese dialects. |
Speech recognition | 2025-09-24 | International |
| The Fun ASR version, updated in April 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. This version is equivalent to the snapshot released on November 7, 2025. |
Video generation | 2025-09-23 | International |
| The upgraded Wan2.5 Preview text to video model, newly upgraded model architecture supports synchronized audio generation with visuals, enables 10-second long video generation, and offers enhanced instruction adherence, improved motion capabilities, and superior image quality. |
Video generation | 2025-09-23 | International |
| The upgraded Wan2.5 Preview image to video model, newly upgraded model architecture supports synchronized audio generation with visuals, enables 10-second long video generation, and offers enhanced instruction adherence, improved motion capabilities, and superior image quality. |
Visual understanding | 2025-09-23 | International |
| The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
Reasoning, Visual understanding | 2025-09-23 | International |
| Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade. |
Visual understanding | 2025-09-23 | International |
| The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement. |
Realtime speech translation | 2025-09-23 | International |
| The real-time version of Qwen3-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, including 8 Chinese dialects.This version is a snapshot version from September 22, 2025. |
Text generation | 2025-09-23 | International |
| Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. This is a snapshot from 23 September, 2025.Compared to the previous version (snapshot from July 22), it demonstrates improved robustness in downstream task performance and tool invocation, along with enhanced code security. |
Image generation | 2025-09-23 | International |
| The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. |
Realtime speech recognition | 2025-09-23 | International |
| This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments. |
Speech recognition | 2025-09-19 | International |
| Qwen3-Omni-30b-a3b-Captioner is a powerful fine-grained audio analysis model designed to generate accurate and comprehensive content descriptions in complex and changing audio scenarios. It can automatically parse and describe various audio content, from complex speech and ambient sounds to music and film and television sound effects, and can maintain stable and reliable output even in multi-source and mixed environments. |
Realtime omni-modal | 2025-09-17 | International |
| The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
Omni-modal | 2025-09-17 | International |
| Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.This version is a snapshot version from September 15, 2025. |
Speech synthesis | 2025-09-16 | International |
| The Qwen3-TTS-Flash-Realtime-2025-09-18 model is Tongyi's latest real-time speech synthesis foundation model, featuring 17 expressive voices while delivering low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on extensive datasets, the system autonomously adjusts vocal tones based on text semantics and demonstrates robust capabilities for complex content synthesis.This model is provided as a snapshot version. |
Speech synthesis | 2025-09-16 | International |
| The Qwen 3-TTS-Flash-2025-09-18 is Tongyi's latest offline text-to-speech foundation model, featuring 17 expressive voices while enabling low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on massive datasets, the system automatically adjusts vocal tones based on text semantics and demonstrates robust capabilities for synthesizing complex content. This model is provided as a snapshot version. |
Speech recognition | 2025-09-16 | International |
| Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in 11 languages, ensuring precise transcription even in complex audio environments. |
Visual understanding | 2025-09-16 | International |
| Qwen-VL-Plus is the enhanced version of the large visual language model. It significantly improves detail recognition and text recognition capabilities, supporting images with resolutions exceeding one million pixels and any aspect ratio specifications. The model delivers exceptional performance across a wide range of visual tasks. |
Visual understanding | 2025-09-16 | International |
| Qwen-VL-Max is a large-scale visual language model of the Qwen series. Compared to the Plus version, it further enhances visual reasoning capabilities and instruction-following abilities, offering higher levels of visual perception and cognition. It delivers optimal performance on more complex tasks. |
Text generation, Reasoning | 2025-09-16 | International |
| Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Visual understanding | 2025-09-12 | International |
| The all-new Wan2.2 First and Last Frame to video model is here. We've optimized motion stability and success rates, enhanced prompt adherence, and enabled seamless transitions between two images. |
Text generation, Reasoning | 2025-09-11 | International |
| A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507). |
Text generation | 2025-09-11 | International |
| A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507). |
Text generation, Reasoning | 2025-09-11 | International |
| As of the September 11, 2025 snapshot, this release features enhanced instruction following and streamlined summary responses in thinking mode. Non-thinking mode provides superior Chinese text understanding and augmented logical reasoning capabilities. The model supports a 1M context length with tiered pricing. |
Text generation, Reasoning | 2025-09-05 | International |
| A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains. |
Text embedding | 2025-08-25 | International |
| The General Text Vector V4 version is a multi-language text vector model developed by the Tongyi Lab based on Qwen3. Compared to the V3 version, it significantly improves performance in text retrieval, clustering, and classification tasks. It achieves a 15% to 40% improvement in evaluation tasks such as MTEB multilingual, Chinese-English, and code retrieval. Additionally, it supports user-defined vector dimensions ranging from 64 to 2048. |
Image generation | 2025-08-18 | International |
| The first Qwen image editing model extends Qwen-Image's text rendering to editing tasks. It offers precise bilingual (Chinese/English) text editing, dual visual and semantic editing, and strong cross-benchmark performance. |
Video generation | 2025-08-15 | International |
| The upgraded Wan 2.2 image to video Flash model, delivers faster speed with optimized stability, more powerful prompt following, improved consistency for text, portraits, and products, and precise shot control. |
Realtime omni-modal | 2025-08-14 | International |
| The real-time version of Qwen's new large multimodal understanding and generation model, suitable for real-time audio interaction scenarios. It supports the understanding of audio accompanied by text, images, and video mixed inputs, and can simultaneously generate speech and text in stream, providing four natural tones. |
Image generation | 2025-08-14 | International |
| The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. |
Text generation, Reasoning | 2025-08-01 | International |
| The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
Text generation | 2025-07-31 | International |
| Qwen3-based code generation model that inherits the coding agent ability of Qwen3-Coder-480B-A35B-Instruct; code capability reaches SOTA at the same scale. |
Text generation, Reasoning | 2025-07-31 | International |
| The Qwen series model optimized for balanced performance, offering inference efficiency between Qwen-Max and Qwen-Turbo, is designed to handle moderately complex tasks effectively. This dynamically updated version implements changes without prior notice. |
Text generation, Reasoning | 2025-07-30 | International |
| Open-source Qwen3 thinking model; compared to the previous version (Qwen3-30B-A3B) excels in complex thinking tasks, including logic, math, science, code, and other challenging scenarios; instruction following, text understanding, and multilingual translation capabilities significantly improved. |
Text generation | 2025-07-29 | International |
| Based on Qwen3, this code generation model inherits the coding agent capabilities of Qwen3-Coder-Plus and supports multi-turn tool interaction. It features focused optimizations on repository-level understanding and enhanced tool-calling stability. |
Text generation | 2025-07-29 | International |
| Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-30B-A3B) shows major improvements in Chinese, English, and overall multilingual general capabilities. Optimized for subjective open-ended tasks, delivering responses significantly more aligned with user preferences and more helpful. |
Text generation, Reasoning | 2025-07-29 | International |
| Qwen3 series Plus model, integrates thinking and non-thinking modes and can switch modes during dialogue. Compared to the prior version, adds dedicated enhancements for Chinese & English capabilities and tool calling. This is a snapshot from 28 July, 2025; first to support 1 M context length, and uses tiered pricing. |
Video generation | 2025-07-28 | International |
| The upgraded Wan 2.2 Plus text to video model, delivers higher quality results with stable sweeping complex movements, cinematic vision control, more powerful prompt following, and realistic world recreation. |
Image generation | 2025-07-28 | International |
| The upgraded Wan 2.2 Plus text to image model, delivers richer image detail with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. |
Image generation | 2025-07-28 | International |
| The upgraded Wan 2.2 Flash text to image model, delivers faster speed with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. |
Video generation | 2025-07-28 | International |
| The upgraded Wan 2.2 Plus image to video model, delivers higher quality results with optimized stability, more powerful prompt following, improved consistency for text, portraits, and products, and precise shot control. |
Text generation, Reasoning | 2025-07-25 | International |
| Open-source Qwen3 thinking model; compared to the previous version (Qwen3-235B-A22B) shows major improvements in logical ability, general capabilities, knowledge enhancement, and creativity, suitable for high-difficulty, strong-thinking scenarios. |
Text generation | 2025-07-24 | International |
| Qwen-MT-Turbo is a large language model within the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 92 languages at a cost-effective price point. It also offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
Text generation | 2025-07-24 | International |
| Qwen-MT-Plus, the flagship translation model from our Qwen series, is now fully upgraded with the Qwen3 architecture. It supports 92 languages and delivers exceptionally accurate and natural-sounding translations. Its advanced capabilities in contextual understanding, terminology control, and format preservation make it a superior choice over traditional models, especially for specialized domains. |
Text generation | 2025-07-23 | International |
| Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. |
Text generation | 2025-07-23 | International |
| Qwen3-based code generation model with strong coding agent power; code capability reaches open-source SOTA. |
Text generation | 2025-07-23 | International |
| Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-235B-A22B) shows slight improvements in subjective creativity and model safety. |
Text generation, Reasoning | 2025-07-23 | International |
| The Plus model of the Qwen3 Series, achieving effective integration of thinking mode and non-thinking mode, allows switching modes during conversations. This is a snapshot from July 14, 2025. Compared to the previous version, there has been a significant improvement in both Chinese and English capabilities under non-thinking mode, with enhanced tool-calling abilities. |
Text generation | 2025-07-21 | International |
| The Qwen Role-Playing Model Series is specifically optimized for Japanese anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. This version significantly enhances Japanese linguistic localization (including dialects and honorifics), human-like role-playing authenticity, narrative coherence control, and scenario-based cognitive intelligence. |
Omni-modal | 2025-07-18 | International |
| The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. |
Image generation | 2025-05-22 | International |
| Wan2.1 Text-to-Image Turbo version, faster generation speed. Upgraded in image beauty, realism, and artistry. Stronger semantic understanding ability, rich style generalization ability, supports up to 2 million pixel generation, supports smart prompt rewriting. |
Image generation | 2025-05-22 | International |
| Wan2.1 Text-to-Image Plus version, Generate more image details. Upgraded in image beauty, realism, and artistry. Stronger semantic understanding ability, rich style generalization ability, supports up to 2 million pixel generation, supports smart prompt rewriting. |
Video generation | 2025-05-14 | International |
| Wan2.1-VACE-Plus All-in-One Video Creation and Editing model.It supports local editing, video repainting, background outpainting, duration extension, image reference, and other video editing and generation tasks, and supports multimodal conditional control through text, images, and videos. |
Text generation, Reasoning | 2025-05-12 | International |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-7B. |
Text generation, Reasoning | 2025-05-12 | International |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-32B-Instruct, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-05-12 | International |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability rivals QwQ-32B with a smaller parameter size, and its general capability significantly surpasses Qwen2.5-14B, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-05-12 | International |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-72B-Instruct, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-05-12 | International |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-14B. |
Text generation, Reasoning | 2025-04-29 | International |
| The Plus model of the Qwen3 series, effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities significantly surpass those of QwQ, and its general capabilities notably exceed those of Qwen2.5-Plus, reaching the SOTA level in the same scale within the industry. This model is the snapshot from April 28, 2025. |
Video generation | 2025-04-07 | International |
| Wan2.1 start and end frames to video Plus version, generate a smooth transition video for two images. Support for large and complex movements, adherence to physical laws, rich artistic styles, and visual quality at the film and television level. The ability to follow instructions is further enhanced, resulting in richer details in the generated video. |
Omni-modal | 2025-03-26 | International |
| The new multi-modal understanding and generation model trained based on Qwen2.5. It supports text, image, speech, video, and mixed input understanding and can simultaneously generate streams of text and speech, significantly improves the speed of multi-modal content understanding. It provides four natural tones. |
Reasoning, Visual understanding | 2025-03-26 | International |
| The Tongyi Qianwen QVQ visual reasoning model supports visual input and chain-of-thought output, demonstrating stronger capabilities in mathematics, programming, visual analysis, creation, and general tasks. |
Reasoning | 2025-03-05 | International |
| The enhanced version of the Qwen QwQ reasoning model, trained on the Qwen2.5 model, has significantly improved its reasoning capabilities through reinforcement learning. The model's core metrics in mathematics and coding (e.g., AIME 24/25, LiveCodeBench) as well as some general metrics (e.g., IFEval, LiveBench) have reached the level of the full version of DeepSeek-R1. |
Video generation | 2025-02-27 | International |
| Wan2.1 image to video Turbo version, make the static image generated video. Support for large and complex movements, adherence to physical laws, artistic styles, and visual quality of movies. The ability to follow instructions is further improved, and generation more fast. |
Text generation, Reasoning | 2025-01-30 | International |
| The Turbo model of the Qwen3 series. It effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities rival those of QwQ-32B with a smaller parameter size, while its general capabilities significantly surpass those of Qwen2.5-Turbo, achieving the SOTA level in the same scale within the industry. |
Text generation | 2025-01-30 | International |
| The Qwen series of models, which are well-balanced in capabilities, offer reasoning performance and speed that fall between Qwen-Max and Qwen-Turbo, making them suitable for moderately complex tasks. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Text generation | 2025-01-30 | International |
| Qwen-Max supports a parameter scale of hundreds of billions and multiple input languages such as Chinese and English. Qwen-Max is updated in a rolling manner. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Video generation | 2025-01-20 | International |
| Wan2.1 image to video Plus version, make the static image generated video. Support for large and complex movements, adherence to physical laws, artistic styles, and visual quality of movies. The ability to follow instructions is further improved, and better video quality. |
Video generation | 2025-01-09 | International |
| Wan2.1 text to video Turbo version, one sentence generated video. Support for large and complex movements, adherence to physical laws, artistic styles, and visual quality of movies. The ability to follow instructions is further improved, and the generation speed is faster. |
Video generation | 2025-01-09 | International |
| Wan2.1 text to video Plus version, one sentence generated video. Support for large and complex movements, adherence to physical laws, artistic styles, and visual quality of movies. The ability to follow instructions is further improved, and better video quality. |
Text embedding | 2024-07-12 | International |
| The general text vectorization model is a multilingual unified text vectorization model developed by Tongyi Lab based on the large language model (LLM) foundation. This model is designed for multiple mainstream languages worldwide and provides advanced vectorization services to help developers convert text data into high-quality vector data. |
US (Virginia)
Model type | Date | Service scope | Model ID | Description |
|---|---|---|---|---|
Text generation, Deep thinking, Visual understanding | 2026-08-26 | Global |
| Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications. |
Video generation | 2026-08-20 | Global |
| Wan3.0-Video-Prime is the high-speed version of the Wan3.0 video generation model, with capabilities aligned with the Wan3.0-Video standard version. It supports four-modal all-reference input and can generate videos up to 30 seconds long, delivering an immersive audio-visual experience with significantly improved end-to-end speed. |
Text generation, Reasoning | 2026-08-19 | US |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Text generation, Reasoning | 2026-08-14 | Global |
| A flagship Mixture-of-Experts (MoE) large language model with 1.6 trillion total parameters and 49 billion activated parameters. It natively supports context windows of up to 1 million tokens. Trained on extensive high-quality data, the model delivers strong performance in mathematical and logical reasoning, complex reasoning, professional code generation, and in-depth long-document analysis, and is suitable for demanding scenarios such as advanced scientific research, complex enterprise workflows, and sophisticated agentic applications. |
Video generation | 2026-08-06 | Global |
| Wan3.0 is an all-in-one video generation model that uniformly supports multiple creative functions including reference, editing, replication, and driving. It supports four-modal all-reference input, up to 30 seconds of video generation, and can parse files, web pages, and complex images. With production-grade character consistency and realistic audio and visuals, it delivers an immersive audio-visual impact. |
Text generation, Reasoning, Visual understanding | 2026-08-02 | Global |
| A 2.4-trillion-parameter MoE flagship with a comprehensive leap in coding and office capabilities, capable of autonomously coding for days to deliver complete projects. It handles hundreds of professional tasks such as legal, financial, and design work, delivering production-grade results end-to-end in a single conversation. Native visual understanding runs through the full planning, execution, and verification pipeline, supporting deep semantic parsing of ultra-long documents and long videos. It performs autonomous planning and closed-loop iteration in long-horizon tasks, continuously evolving. |
Text generation, Reasoning | 2026-08-01 | Global |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Text generation, Reasoning, Visual understanding | 2026-07-21 | Global |
| The Qwen3.7 native vision-language series Flash models comprehensively enhance multimodal understanding and Agent execution capabilities compared to 3.6-Flash. Key improvements include strengthened foundational multimodal abilities, enhanced object recognition, improved real-world perception and spatial intelligence. Multimodal Agent scenarios such as Search Agent and CI Agent have seen significant upgrades, with more stable end-to-end task execution. Multimodal coding capabilities are optimized, delivering a smoother vibe coding experience. |
Text generation, Reasoning | 2026-07-07 | US |
| GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
Text generation, Reasoning, Visual understanding | 2026-07-03 | US |
| The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
Text generation | 2026-07-02 | Global |
| The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
Text generation, Reasoning, Visual understanding | 2026-06-26 | US |
| Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
Text generation, Reasoning | 2026-06-26 | US |
| The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
Video generation | 2026-06-16 | Global |
| HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
Video generation | 2026-06-16 | Global |
| HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
Video generation | 2026-06-16 | Global |
| HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
Text generation, Reasoning | 2026-06-16 | Global |
| GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
Text generation, Reasoning, Visual understanding | 2026-06-15 | Global |
| kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
Text generation, Reasoning, Visual understanding | 2026-06-09 | Global |
| The Max model, the largest and most capable in the Qwen3.7 series, has added visual‑modal understanding compared to the May 20 snapshot, enabling it to perceive real‑world scenes and supporting multimodal interactive hybrid agent capabilities. This version is based on a snapshot taken on June 8, 2026. |
Text generation, Reasoning, Visual understanding | 2026-06-01 | Global |
| Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
Text generation, Reasoning | 2026-05-20 | Global |
| The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
Text generation, Reasoning | 2026-05-11 | US |
| A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
Text generation, Reasoning | 2026-05-11 | US |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Text generation, Reasoning, Visual understanding | 2026-04-29 | Global |
| Kimi-k2.5 is Moonshot AI's most versatile model with native multimodal architecture, supporting visual/text inputs, thinking/non-thinking modes, and both dialogue and agent tasks. |
Video generation | 2026-04-26 | Global |
| HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
Video generation | 2026-04-26 | Global |
| HappyHorse-1.0-R2V supports Reference-to-Video generation, offering enhanced stability in subject and scene referencing. Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance. |
Text generation, Reasoning | 2026-04-24 | Global |
| A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
Text generation, Reasoning | 2026-04-24 | Global |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Video generation | 2026-04-22 | Global |
| HappyHorse-1.0-T2V supports text-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends text semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
Video generation | 2026-04-22 | Global |
| HappyHorse-1.0-I2V enables image-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends both text and image semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
Text generation, Reasoning, Visual understanding | 2026-04-17 | Global |
| The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
Text generation, Reasoning, Visual understanding | 2026-04-17 | Global |
| The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance. |
Text generation, Reasoning | 2026-04-14 | Global |
| GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
Text generation, Reasoning, Visual understanding | 2026-04-01 | Global |
| The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |
Text generation | 2026-03-30 | US |
| Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
Visual understanding | 2026-03-14 | US |
| The Qwen3 series of small-sized visual understanding models effectively integrates thinking and non-thinking modes. Compared with the snapshot taken on October 15, 2025, the overall performance of the model has improved significantly: it delivers enhanced capabilities in general visual recognition and reasoning, and shows marked improvements in recognition accuracy across various business scenarios such as security, in-store inspections, equipment monitoring, and photo-based problem solving. This version is a snapshot as of January 22, 2026. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | Global |
| The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | Global |
| The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | Global |
| The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | Global |
| The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B. |
Text generation, Reasoning, Visual understanding | 2026-02-15 | Global |
| The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities. |
Text generation, Reasoning, Visual understanding | 2026-02-15 | Global |
| The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent. |
Video generation | 2025-12-16 | Global |
| Wan2.6 reference to video, supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
Image generation | 2025-12-15 | Global |
| Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation. |
Image generation | 2025-12-15 | Global |
| Wan2.6 Image, An all-round image generation model that supports joint text–image reasoning, multi-image creative fusion, commercial-grade consistency, aesthetic style transfer, and precise control of framing and lighting, significantly enhancing consistency, controllability, and expressiveness in image generation. |
Video generation | 2025-12-03 | US |
| Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
Video generation | 2025-12-03 | Global |
| Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
Video generation | 2025-12-03 | US |
| Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
Video generation | 2025-12-03 | Global |
| Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
Text generation, Reasoning | 2025-12-01 | US |
| This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
Text generation, Reasoning | 2025-12-01 | Global |
| This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
Visual understanding | 2025-11-21 | Global |
| This model is a snapshot version from November 20, 2025, and is based on the latest Qwen-VL3 architecture with a comprehensive upgrade. It features significant improvements in document parsing and text localization capabilities, as well as substantial reductions in end-to-end latency and illusions. |
Text generation | 2025-11-19 | Global |
| Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
Text generation | 2025-11-11 | Global |
| Qwen-MT-Flash, a large language model from the Qwen series, has been fully upgraded with the Qwen 3 architecture for significantly enhanced performance and translation quality. It provides rapid, cost-effective translation across 92 languages, while supporting advanced features such as terminology intervention, format preservation, and domain-specific adaptation. It is the ideal choice for applications requiring a powerful balance of speed, quality, and cost. |
Reasoning, Visual understanding | 2025-10-21 | Global |
| The largest dense model in the Qwen3-VL series, its reasoning version boasts multimodal reasoning capabilities second only to Qwen3-VL-235B-Thinking. It excels in STEM and math problem-solving, general image and video understanding, and achieves state-of-the-art performance in multimodal agent capabilities, making it ideal for complex multimodal reasoning tasks. |
Visual understanding | 2025-10-21 | Global |
| The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios. |
Visual understanding | 2025-10-15 | US |
| The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
Reasoning, Visual understanding | 2025-10-15 | Global |
| The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
Reasoning, Visual understanding | 2025-10-03 | Global |
| The Thinking version of the second-largest MoE model in the Qwen3-VL series features fast response speeds and enhanced multimodal understanding and reasoning capabilities, visual agents, and support for extremely long contexts such as lengthy videos and documents. It also boasts comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities, making it well-suited for complex real-world tasks. |
Visual understanding | 2025-10-03 | Global |
| The Instruct version of the second-largest MoE model in the Qwen3-VL series offers rapid response speeds and supports extremely long contexts like lengthy videos and documents. It includes comprehensively upgraded image/video understanding, spatial awareness, and object recognition capabilities, as well as 2D/3D visual localization, enabling it to handle intricate real-world challenges. |
Reasoning, Visual understanding | 2025-09-30 | Global |
| The Thinking version of the 8B Dense model in the Qwen3-VL series consumes less GPU memory and is capable of performing multimodal understanding and reasoning. It supports extremely long contexts such as lengthy videos and documents, 2D/3D visual localization, and features comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities. |
Visual understanding | 2025-09-30 | Global |
| The Instruct version of the 8B Dense model in the Qwen3-VL series requires less GPU memory and provides comprehensively upgraded image/video understanding, support for extremely long contexts like lengthy videos and documents, spatial awareness, and object recognition capabilities, making it suitable for tackling complex real-world tasks. |
Text generation, Reasoning | 2025-09-24 | Global |
| The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
Visual understanding | 2025-09-23 | Global |
| The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
Reasoning, Visual understanding | 2025-09-23 | Global |
| Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade. |
Visual understanding | 2025-09-23 | Global |
| The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement. |
Text generation | 2025-09-23 | Global |
| Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. This is a snapshot from 23 September, 2025.Compared to the previous version (snapshot from July 22), it demonstrates improved robustness in downstream task performance and tool invocation, along with enhanced code security. |
Speech recognition | 2025-09-16 | US |
| Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in 11 languages, ensuring precise transcription even in complex audio environments. |
Text generation, Reasoning | 2025-09-16 | US |
| Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Text generation, Reasoning | 2025-09-16 | Global |
| Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Text generation, Reasoning | 2025-09-11 | Global |
| A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507). |
Text generation | 2025-09-11 | Global |
| A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507). |
Text generation, Reasoning | 2025-09-11 | Global |
| As of the September 11, 2025 snapshot, this release features enhanced instruction following and streamlined summary responses in thinking mode. Non-thinking mode provides superior Chinese text understanding and augmented logical reasoning capabilities. The model supports a 1M context length with tiered pricing. |
Text generation, Reasoning | 2025-09-05 | Global |
| A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains. |
Text generation, Reasoning | 2025-08-01 | US |
| The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
Text generation, Reasoning | 2025-08-01 | Global |
| The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
Text generation | 2025-07-31 | Global |
| Qwen3-based code generation model that inherits the coding agent ability of Qwen3-Coder-480B-A35B-Instruct; code capability reaches SOTA at the same scale. |
Text generation, Reasoning | 2025-07-30 | Global |
| Open-source Qwen3 thinking model; compared to the previous version (Qwen3-30B-A3B) excels in complex thinking tasks, including logic, math, science, code, and other challenging scenarios; instruction following, text understanding, and multilingual translation capabilities significantly improved. |
Text generation | 2025-07-29 | Global |
| Based on Qwen3, this code generation model inherits the coding agent capabilities of Qwen3-Coder-Plus and supports multi-turn tool interaction. It features focused optimizations on repository-level understanding and enhanced tool-calling stability. |
Text generation | 2025-07-29 | Global |
| Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-30B-A3B) shows major improvements in Chinese, English, and overall multilingual general capabilities. Optimized for subjective open-ended tasks, delivering responses significantly more aligned with user preferences and more helpful. |
Text generation, Reasoning | 2025-07-29 | Global |
| Qwen3 series Plus model, integrates thinking and non-thinking modes and can switch modes during dialogue. Compared to the prior version, adds dedicated enhancements for Chinese & English capabilities and tool calling. This is a snapshot from 28 July, 2025; first to support 1 M context length, and uses tiered pricing. |
Text generation, Reasoning | 2025-07-25 | Global |
| Open-source Qwen3 thinking model; compared to the previous version (Qwen3-235B-A22B) shows major improvements in logical ability, general capabilities, knowledge enhancement, and creativity, suitable for high-difficulty, strong-thinking scenarios. |
Text generation | 2025-07-24 | Global |
| Qwen-MT-Plus, the flagship translation model from our Qwen series, is now fully upgraded with the Qwen3 architecture. It supports 92 languages and delivers exceptionally accurate and natural-sounding translations. Its advanced capabilities in contextual understanding, terminology control, and format preservation make it a superior choice over traditional models, especially for specialized domains. |
Text generation | 2025-07-23 | Global |
| Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. |
Text generation | 2025-07-23 | Global |
| Qwen3-based code generation model with strong coding agent power; code capability reaches open-source SOTA. |
Text generation | 2025-07-23 | Global |
| Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-235B-A22B) shows slight improvements in subjective creativity and model safety. |
Visual understanding | 2025-07-07 | Global |
| Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
Text generation, Reasoning | 2025-05-12 | Global |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-7B. |
Text generation, Reasoning | 2025-05-12 | Global |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-32B-Instruct, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-05-12 | Global |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability rivals QwQ-32B with a smaller parameter size, and its general capability significantly surpasses Qwen2.5-14B, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-05-12 | Global |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-72B-Instruct, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-05-12 | Global |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-14B. |
China (Beijing)
Model type | Date | Model ID | Description |
|---|---|---|---|
Image translation | 2026-08-28 |
| An advanced image translation engine supporting 55 source languages. It delivers pixel-perfect replication of the original layout and typography. Gain full control over professional content with custom terminology and sensitive word filtering, ensuring a flexible, accurate, and high-performance image localization service. |
Text generation, Deep thinking, Visual understanding | 2026-08-26 |
| Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications. |
Video generation | 2026-08-20 |
| Wan3.0-Video-Prime is the high-speed version of the Wan3.0 video generation model, with capabilities aligned with the Wan3.0-Video standard version. It supports four-modal all-reference input and can generate videos up to 30 seconds long, delivering an immersive audio-visual experience with significantly improved end-to-end speed. |
Text generation, Reasoning, Visual understanding | 2026-08-19 |
| Kimi K3 is Kimi's most powerful flagship model with 2.8 trillion parameters. Built on KDA hybrid linear attention (Kimi Delta Attention) and Attention Residuals technologies, it natively supports vision understanding and features a 1 million token context window. It is the world's first open-source 3-trillion-parameter model designed for long-context programming, knowledge work, and advanced reasoning. |
Text generation, Reasoning, Visual understanding | 2026-08-17 |
| A 27B native vision-language Dense model of the Qwen3.8 series. Compared with 3.6-27B, it focuses on improving coding and office-scenario capabilities under text and visual modalities, able to more reliably complete complex tasks end-to-end and deliver trustworthy results. |
Text generation, Reasoning | 2026-08-14 |
| A flagship Mixture-of-Experts (MoE) large language model with 1.6 trillion total parameters and 49 billion activated parameters. It natively supports context windows of up to 1 million tokens. Trained on extensive high-quality data, the model delivers strong performance in mathematical and logical reasoning, complex reasoning, professional code generation, and in-depth long-document analysis, and is suitable for demanding scenarios such as advanced scientific research, complex enterprise workflows, and sophisticated agentic applications. |
Text generation | 2026-08-12 |
| Qwen3.8-2.4T-A95B is the open-source version of Qwen's latest flagship series, released in August 2026. It adopts a sparse MoE architecture with 2.4 trillion total parameters and about 95 billion activated per step, combined with a hybrid attention mechanism, supporting a 1 million token context. Key benchmarks: GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, BabyVision 82.0, CodeArena global #4. |
Video generation | 2026-08-06 |
| Wan3.0 is an all-in-one video generation model that uniformly supports multiple creative functions including reference, editing, replication, and driving. It supports four-modal all-reference input, up to 30 seconds of video generation, and can parse files, web pages, and complex images. With production-grade character consistency and realistic audio and visuals, it delivers an immersive audio-visual impact. |
Image generation | 2026-08-04 |
| Clear instruction understanding: supports up to 4.5k token input, generating complex image-text instructions in one pass. Stable text rendering: 10px small text, 12 languages, 20+ fonts clearly legible, ready for infographics and interfaces. Handy batch output: posters, web pages, interfaces and other daily tasks can be generated in bulk at lower cost, ideal for continuous creation. Qwen-Image-3.0 Standard pursues not only generation quality but also everyday creative convenience, making image generation a sustainable content productivity tool. |
Text generation, Reasoning, Visual understanding | 2026-08-02 |
| A 2.4-trillion-parameter MoE flagship with a comprehensive leap in coding and office capabilities, capable of autonomously coding for days to deliver complete projects. It handles hundreds of professional tasks such as legal, financial, and design work, delivering production-grade results end-to-end in a single conversation. Native visual understanding runs through the full planning, execution, and verification pipeline, supporting deep semantic parsing of ultra-long documents and long videos. It performs autonomous planning and closed-loop iteration in long-horizon tasks, continuously evolving. |
Text generation, Reasoning | 2026-08-01 |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Speech recognition | 2026-07-30 |
| Added the Qwen-Audio-3.0-ASR-Flash-Streaming (real-time), Qwen-Audio-3.0-ASR-Flash-Filetrans (non-real-time), and Qwen-Audio-3.0-ASR-Flash (non-real-time) models: Dialect support: Supports the seven major Chinese dialect groups (Mandarin, Wu, Xiang, Gan, Hakka, Min, and Yue) and more than 20 regional accents; Classical poetry optimization: Improves recognition accuracy for classical Chinese poetry, making it suitable for education, culture, and audiobook scenarios; Text optimization: Enhances punctuation prediction and text normalization, automatically converting numbers, dates, and monetary amounts to standard formats; Multilingual expansion: Supports 30 languages, including Chinese, English, Japanese, and Korean; Hotwords and context: Supports hotwords (precompiled and on-the-fly) and context input to improve recognition accuracy for domain-specific terms. |
Text generation, Reasoning, Visual understanding | 2026-07-21 |
| The Qwen3.7 native vision-language series Flash models comprehensively enhance multimodal understanding and Agent execution capabilities compared to 3.6-Flash. Key improvements include strengthened foundational multimodal abilities, enhanced object recognition, improved real-world perception and spatial intelligence. Multimodal Agent scenarios such as Search Agent and CI Agent have seen significant upgrades, with more stable end-to-end task execution. Multimodal coding capabilities are optimized, delivering a smoother vibe coding experience. |
Image generation | 2026-07-20 |
| Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and 20+ fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool. |
Speech synthesis, Realtime speech synthesis | 2026-07-14 |
| Qwen-Audio-3.0-TTS-Plus is a high-performance speech synthesis model, designed for high-quality speech generation scenarios. Compared with the previous version, it supports more low-resource languages and Chinese dialects, significantly improves dialect authenticity, and enhances free-style instruction following and fine-grained tag control for more accurate control over emotion, tone, character, speaking rate, volume, and synthesis style. It is also more robust under noisy and reverberant acoustic conditions, with further improvements in audio quality, clarity, resolution, and overall expressiveness. The Plus version focuses more on synthesis quality and detailed expressiveness, making it suitable for professional scenarios with higher requirements for audio quality, naturalness, and expressiveness, such as content creation, audiobooks, film and video dubbing, brand voice design, and premium speech services. |
Realtime speech synthesis | 2026-07-14 |
| qwen-audio-3.0-tts-flash is a high-performance speech synthesis model, optimized for real-time interactive scenarios. Compared with the previous version, it supports more low-resource languages and Chinese dialects, improves dialect authenticity, and enhances free-style instruction following and fine-grained tag control for more flexible control over emotion, tone, character, speaking rate, volume, and expressive style. It is also more robust under noisy and reverberant acoustic conditions, with improved audio quality, clarity, and overall expressiveness. The Flash version focuses on real-time synthesis, with first-packet latency controlled within 200 ms, making it suitable for voice assistants, real-time dialogue, intelligent customer service, and other low-latency interactive applications. |
Realtime chat | 2026-07-14 |
| Qwen 3.0 Realtime Speech Model Standard Edition - next-generation duplex speech model ranked #1 globally in Artificial Analysis Speech-to-Speech benchmark. Balances model intelligence with duplex dialogue rhythm for natural interaction and enhanced response quality. |
Realtime chat | 2026-07-14 |
| Qwen 3.0 Realtime Speech Dialogue Model Flash Edition - a next-generation duplex speech model ranked #1 globally in Artificial Analysis Speech-to-Speech benchmark. Combines high intelligence with optimized duplex rhythm for low-latency (parallel inference, full-streaming optimization) 'fast and smart' dialogue experience focusing on response speed. |
Text generation | 2026-07-09 |
| GLM-5.2-Fast-Preview is the high-speed variant of Zhipu AI's GLM-5.2, with 1M context and capabilities on par with the standard version. Inference-optimized to deliver 1.5–2× the output TPS, it fits latency-sensitive use cases such as real-time chat, multi-turn agents, and streaming code generation. |
Video generation | 2026-07-01 |
| Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of June 12, 2026. |
Video generation | 2026-07-01 |
| Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. This version is a snapshot as of June 12, 2026. |
Text embedding | 2026-06-30 |
| A text-ranking model trained on the Qwen LLM foundation performs relevance ranking for input queries and candidate documents. It supports over 100 languages and long-text inputs, and is suitable for applications such as text retrieval and RAG. Its performance is aligned with the open-source Qwen3-Rerank series models. |
Image generation | 2026-06-25 |
| The Qwen-Image-2.0 series full-fledged model integrates image generation and editing; it boasts more professional text rendering capabilities with 1k token command support, more delicate and realistic textures, meticulous depiction of realistic scenes, and stronger semantic adherence. The full-fledged version possesses the strongest text rendering capabilities and realistic textures in the 2.0 series. |
Speech recognition | 2026-06-17 |
| The Bailing ASR version, updated in June 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. It supports contextualization capabilities and can transcribe audio up to 5 minutes in length. |
Visual understanding | 2026-06-16 |
| Qwen3.5-OCR is an upgraded OCR model with enhanced document parsing, text localization, and key information extraction. It significantly improves extraction performance for real-world documents (e.g., ID cards, driver's licenses). |
Video generation | 2026-06-16 |
| HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
Video generation | 2026-06-16 |
| HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
Video generation | 2026-06-16 |
| HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
Text generation, Reasoning | 2026-06-16 |
| GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
Text generation, Reasoning, Visual understanding | 2026-06-15 |
| kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
Text generation, Reasoning, Visual understanding | 2026-06-09 |
| The Max model, the largest and most capable in the Qwen3.7 series, has added visual‑modal understanding compared to the May 20 snapshot, enabling it to perceive real‑world scenes and supporting multimodal interactive hybrid agent capabilities. This version is based on a snapshot taken on June 8, 2026. |
Text generation, Reasoning, Visual understanding | 2026-06-01 |
| Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
Speech synthesis | 2026-06-01 |
| The Fun Music large model supports song creation based on open-ended lyrics or composition requirements, generating full Chinese/English songs with male/female vocals. The songs are accessible and emotionally progressive, representing a perfect fusion of human inspiration and LLM capabilities. |
Text generation, Reasoning | 2026-05-21 |
| The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
Realtime speech translation | 2026-05-19 |
| The real-time version of Qwen3.5-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3.5-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3.5-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 60 languages and speak 29 languages. |
Speech synthesis | 2026-05-06 |
| A music generation model creating complete Chinese/English songs (male/female vocals) from lyrics or creative requirements, blending human inspiration with AI capabilities. |
Video generation | 2026-04-26 |
| Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of April 25, 2026. |
Video generation | 2026-04-26 |
| Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power.This version is a snapshot as of April 25, 2026. |
Video generation | 2026-04-26 |
| HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
Video generation | 2026-04-26 |
| HappyHorse-1.0-R2V supports Reference-to-Video generation, offering enhanced stability in subject and scene referencing. Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance. |
Text generation, Reasoning | 2026-04-24 |
| A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
Text generation, Reasoning | 2026-04-24 |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Text generation, Reasoning, Visual understanding | 2026-04-23 |
| The Qwen3.5 native vision-language series Plus model has seen a substantial improvement in agentic coding capabilities compared to the February 15th snapshot. Inference speed has also been significantly enhanced, while its knowledge retention, reasoning ability, and long-context processing remain at a high level, making it well-suited for complex agent-based tasks. It is ideal for applications such as coding agents, production workflows, and high-throughput scenarios. This version is based on a snapshot taken on April 20, 2026. |
Image generation | 2026-04-23 |
| The full-featured Qwen-Image-2.0 series models integrate image generation and image editing, offering enhanced text rendering with support for 1,000-token prompts, more refined realistic textures, detailed depiction of photorealistic scenes, and stronger semantic adherence. The full-featured version delivers the strongest text rendering and most lifelike textures in the 2.0 series. |
Text generation, Reasoning, Visual understanding | 2026-04-22 |
| The Qwen3.6 27B native vision-language dense model builds upon the 3.5-27B version, with key improvements in agentic coding capabilities and enhanced STEM reasoning and inference skills. In the vision modality, it demonstrates significant advances in spatial intelligence, object localization, and detection, while video understanding, document OCR, and visual agent capabilities continue to improve steadily. |
Video generation | 2026-04-22 |
| HappyHorse-1.0-I2V enables image-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends both text and image semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
Text generation, Reasoning, Visual understanding | 2026-04-21 |
| Kimi-k2.6 is the latest intelligent model in Kimi series, featuring enhanced long-context code generation, improved instruction following and self-correction capabilities, supporting text/image/video inputs and multiple operational modes. |
Video generation | 2026-04-21 |
| HappyHorse-1.0-T2V supports text-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends text semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
Text generation, Reasoning | 2026-04-20 |
| The Max model, the largest and most capable variant in the Qwen3.6 series, is now available in a preview version. At present, only its plain-text capabilities are open for experimentation. Compared with the previously released Qwen3-Max and Qwen3.6-Plus, this model features enhanced vibe coding abilities, more efficient coding agent execution, and significantly improved front-end development skills. Additionally, its long-tail knowledge retention has been further upgraded. |
Text generation, Reasoning, Visual understanding | 2026-04-16 |
| The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
Text generation, Reasoning, Visual understanding | 2026-04-16 |
| The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance. |
Text generation, Reasoning | 2026-04-14 |
| GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
Video generation | 2026-04-03 |
| Wan2.7 video edit, supports both localized and global editing with prompt. Seamlessly replace elements using image references and replicate complex dynamic processes, including motion, special effects, and camera movements. |
Video generation | 2026-04-03 |
| Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
Video generation | 2026-04-03 |
| Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. |
Video generation | 2026-04-03 |
| Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
Image generation | 2026-04-01 |
| Wan2.7–image-pro, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
Image generation | 2026-04-01 |
| Wan2.7 – image generation and editing, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
Text generation, Reasoning, Visual understanding | 2026-04-01 |
| The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |
Realtime omni-modal | 2026-03-30 |
| Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience. |
Omni-modal | 2026-03-30 |
| Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience. |
Realtime omni-modal | 2026-03-30 |
| Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a fully evolved version of Qwen3-Omni, it supports audio input in 60+ languages, voice output in 30+ languages, and controllable voice dialogue, WebSearch and complex FunctionCall invocation, and has intelligent semantic interruption interaction capabilities. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interactive experience. |
Omni-modal | 2026-03-30 |
| Qwen 3.5-Omni is the latest generation of Qwen's multimodal large model, supporting text, image, audio, and audio-visual understanding and interaction. As a comprehensive evolution of Qwen 3-Omni, it supports over 10 hours of audio understanding and over 400 seconds of 720P (1 FPS) audio-visual understanding and dialogue. It further expands the language range, supporting audio input in 60+ languages and speech output in 30+ languages. It also possesses powerful structured audio-visual understanding capabilities and is widely used in text creation, voice assistants, multimedia analysis, and other scenarios, providing a natural and fluent multimodal understanding and interactive experience. |
Speech recognition | 2026-03-05 |
| This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments. |
Speech recognition | 2026-03-03 |
| Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in multiple languages, while ensuring precise transcription even in complex audio environments.This version is a snapshot dated February 10, 2026. |
Image generation | 2026-03-03 |
| The full-featured Qwen-Image-2.0 series models integrate image generation and image editing, offering enhanced text rendering with support for 1,000-token prompts, more refined realistic textures, detailed depiction of photorealistic scenes, and stronger semantic adherence. The full-featured version delivers the strongest text rendering and most lifelike textures in the 2.0 series.This version is a snapshot as of March 3, 2026. |
Image generation | 2026-03-03 |
| The Qwen-Image-2.0 series of accelerated models integrates image generation and image editing, offering enhanced text-rendering capabilities with support for 1,000-token prompts, more realistic textures, finely detailed photorealistic scenes, and improved semantic consistency. The accelerated version effectively strikes an optimal balance between model performance and quality. |
Speech synthesis | 2026-02-27 |
| A high-expressiveness speech synthesis model in the CosyVoice series with enhanced voice cloning and design capabilities. Supports free-style instruction control while maintaining speaker similarity, offering rich synthesis styles. Reduces first-syllable latency, improves pronunciation accuracy, and enhances prosody and sound quality. Enables ultra-natural multilingual (Chinese, English, German, French, Russian, Japanese, Korean, Portuguese, Thai, Indonesian, Vietnamese) real-time speech synthesis. |
Speech synthesis | 2026-02-27 |
| A high-performance speech synthesis model in the CosyVoice series with enhanced voice cloning and design capabilities. Supports free-style instruction control while maintaining speaker similarity, offering rich synthesis styles. Reduces first-syllable latency, improves pronunciation accuracy, and enhances prosody and sound quality. Enables ultra-natural multilingual (Chinese, English, German, French, Russian, Japanese, Korean, Portuguese, Thai, Indonesian, Vietnamese) real-time speech synthesis. |
Text generation, Reasoning | 2026-02-24 |
| MiniMax-M2.5 is MiniMax's flagship open-source large model, trained on hundreds of thousands of real-world complex scenarios through large-scale reinforcement learning, achieving or surpassing industry SOTA in productivity scenarios like programming, tool calls, search, and office tasks. |
Text generation, Reasoning, Visual understanding | 2026-02-23 |
| The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
Text generation, Reasoning, Visual understanding | 2026-02-23 |
| The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B. |
Text generation, Reasoning, Visual understanding | 2026-02-23 |
| The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B. |
Text generation, Reasoning, Visual understanding | 2026-02-23 |
| The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B. |
Text generation | 2026-02-19 |
| The new-generation code generation model in the Qwen3 series delivers performance close to that of Qwen3-Coder-Plus while offering even better capabilities. The model has been optimized with a focus on repository-level understanding, supports multi-turn tool interactions, and enhances its compatibility with agentic coding tools. |
Text generation, Reasoning | 2026-02-18 |
| GLM-5 targets coding and agent scenarios, achieving open-source SOTA in complex systems engineering with capabilities approaching Claude Opus, built on a 744B foundation with asynchronous reinforcement learning and sparse attention. |
Text generation, Reasoning, Visual understanding | 2026-02-15 |
| The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities. |
Text generation, Reasoning, Visual understanding | 2026-02-15 |
| The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent. |
Speech recognition | 2026-02-13 |
| The real-time version of Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in multiple languages, while ensuring precise transcription even in complex audio environments.This version is a snapshot dated February 10, 2026. |
Realtime speech recognition | 2026-02-12 |
| A lightweight real-time Mandarin speech recognition model optimized for Chinese call-center scenarios, supporting multi-dialect accents and achieving low-latency, high-accuracy transcription in low-sample-rate/low-SNR environments. |
Speech synthesis | 2026-02-10 |
| Qwen3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 26, 2026. |
Speech synthesis | 2026-02-10 |
| Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated from the qwen-voice-enrollment service, and supports speech output in 11 languages using the same voice timbre. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 22, 2026. |
Speech synthesis | 2026-02-10 |
| Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments. |
Text generation, Reasoning, Visual understanding | 2026-01-30 |
| Kimi-k2.5 is Moonshot AI's most versatile model with native multimodal architecture, supporting visual/text inputs, thinking/non-thinking modes, and both dialogue and agent tasks. |
Video generation | 2026-01-29 |
| Wan2.6 reference to video flash, faster and more cost-effective generation. Supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
Text embedding | 2026-01-29 |
| Qwen3-VL-Rerank is a reranking model that deeply understands rich multimodal information (text, images, video). After initial retrieval, it applies advanced cross-modal correlation to intelligently re-rank candidates, prioritizing the most relevant results. It improves cross-modal search accuracy, optimizes image/video retrieval precision, enhances image clustering quality, and enables efficient complex multimodal retrieval and tagging. |
Reasoning, Visual understanding | 2026-01-26 |
| The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
Text generation, Reasoning | 2026-01-23 |
| Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model's overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026. |
Visual understanding | 2026-01-22 |
| The Qwen3 series of small-sized visual understanding models effectively integrates thinking and non-thinking modes. Compared with the snapshot taken on October 15, 2025, the overall performance of the model has improved significantly: it delivers enhanced capabilities in general visual recognition and reasoning, and shows marked improvements in recognition accuracy across various business scenarios such as security, in-store inspections, equipment monitoring, and photo-based problem solving. This version is a snapshot as of January 22, 2026. |
Multimodal embedding | 2026-01-21 |
| Qwen3-VL-Embedding is a unified multimodal vector model based on Qwen3-VL, supporting text, image, video (single or mixed modality) input. It outputs unified representation vectors, suitable for cross-modal retrieval, image/video search, clustering, complex multimodal retrieval, and tagging. |
Speech synthesis | 2026-01-21 |
| qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments. This model is equivalent to the snapshot version released on January 22, 2026. |
Video generation | 2026-01-15 |
| Wan2.6 image to video flash, faster and more cost-effective generation. Intelligent shot scheduling enables multi‑camera storytelling, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
Image generation | 2026-01-15 |
| The Max series Qwen's image editing models delivers more stable and versatile editing capabilities: enhanced industrial design and geometric reasoning, improved character consistency, reduced offset issues, and integrated LoRA capabilities for a wider range of image editing functions. |
Speech synthesis | 2026-01-14 |
| Qwen3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 15, 2026. |
Speech synthesis | 2026-01-14 |
| Qwen 3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated from the qwen-voice-enrollment service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from January 15, 2026. |
Text generation | 2026-01-13 |
| The Qwen Role-Playing Model Series is specifically optimized for muti-language anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. This version significantly enhances Japanese linguistic localization (including dialects and honorifics), human-like role-playing authenticity, narrative coherence control, and scenario-based cognitive intelligence. |
Image generation | 2026-01-09 |
| The Qwen series of image-generation models boasts exceptional text-rendering capabilities and excels in complex text rendering as well as a wide range of generation and editing tasks. This version, a snapshot taken on January 9, 2026, is a distilled and accelerated variant of Qwen-Image-Max, enabling faster generation of high-quality images. |
Image generation | 2025-12-30 |
| The Max series of qwen's image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the "AI-like" feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering. |
Text generation, Reasoning | 2025-12-25 |
| Zhipu's latest flagship with enhanced coding and multi-step reasoning capabilities, supporting long-term task planning, tool collaboration, and immersive writing/role-playing. |
Image generation | 2025-12-18 |
| Z-Image-Turbo is a highly efficient image-generation model that has topped the Artificial Analysis benchmark as the world's No. 1 open-source text-to-image model. With just 6 billion parameters and an 8-step inference process, it generates photo-realistic images comparable to those produced by large-scale commercial models, while excelling in bilingual Chinese–English text rendering, complex semantic understanding, and diverse thematic generation. |
Reasoning, Visual understanding | 2025-12-18 |
| The Qwen3 series of visual understanding models effectively integrates thinking and non-thinking modes. Compared to the snapshot released on September 23, this version delivers superior performance in reasoning and analysis tasks as well as style control, while also offering lower latency and faster response speeds. This version is based on a snapshot taken on December 19, 2025. |
Video generation | 2025-12-16 |
| Wan2.6 reference to video, supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
Image generation | 2025-12-15 |
| Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation. |
Image generation | 2025-12-15 |
| Wan2.6 Image, An all-round image generation model that supports joint text–image reasoning, multi-image creative fusion, commercial-grade consistency, aesthetic style transfer, and precise control of framing and lighting, significantly enhancing consistency, controllability, and expressiveness in image generation. |
Image generation | 2025-12-15 |
| The Qianwen series of Image Editing Plus models features enhanced character consistency, industrial design capabilities, and geometric reasoning abilities compared to the snapshot as of October 30. Additionally, it integrates LoRA capabilities such as lighting effects and effectively mitigates offset issues. This version is based on a snapshot taken on December 15, 2025. |
Speech synthesis | 2025-12-12 |
| Qwen 3-TTS-VD model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on the voices designed by the qwen3-voice-design service, and supports speech output in 11 languages using the same voice. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust its tone according to the text, demonstrating good processing capabilities for complex text synthesis. This model is a snapshot version from December 16, 2025. |
Speech synthesis | 2025-12-12 |
| Qwen Voice-Design model is a series of voice design models from Qwen Speech Model. It only requires a simple text description to quickly design a suitable voice. When used in conjunction with the qwen3-tts-vd-realtime model, it can design and output speech in 10 languages. Furthermore, the synthesized audio can adaptively adjust its tone based on the text and has good processing capabilities for complex text synthesis. |
Realtime omni-modal | 2025-12-04 |
| The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
Omni-modal | 2025-12-04 |
| Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
Realtime speech recognition | 2025-12-04 |
| Qwen3-LiveTranslate-Flash is a high-precision, highly responsive, and robust multilingual real-time audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash provides both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, and also supports 8 Chinese dialects. |
Video generation | 2025-12-03 |
| Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
Video generation | 2025-12-03 |
| Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
Text generation, Reasoning | 2025-12-02 |
| DeepSeek-V3.2 is the official release of a model that incorporates DeepSeek Sparse Attention—a sparse attention mechanism. It's also the first model launched by DeepSeek that integrates reasoning into tool usage, supporting both reasoning-enabled and non-reasoning tool calls. |
Text generation, Reasoning | 2025-12-01 |
| This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
Speech synthesis | 2025-11-27 |
| Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. It can perform high-fidelity real-time speech synthesis on voices replicated by the qwen3-voice-enrollment service, and supports speech output in 11 languages with the same voice timbre. This model has been trained on massive amounts of data, and the synthesized audio can adaptively adjust the tone according to the text, and it also has good processing capabilities for complex text synthesis.This model is provided as a snapshot version. |
Speech synthesis | 2025-11-27 |
| The Qwen3-TTS-Flash-Realtime model is Tongyi's latest real-time speech synthesis foundation model, featuring 17 expressive voices while delivering low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on extensive datasets, the system autonomously adjusts vocal tones based on text semantics and demonstrates robust capabilities for complex content synthesis. |
Speech synthesis | 2025-11-27 |
| The Qwen3-TTS-Flash is Tongyi's latest offline text-to-speech foundation model, featuring 17 expressive voices while enabling low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on massive datasets, the system automatically adjusts vocal tones based on text semantics and demonstrates robust capabilities for synthesizing complex content. |
Speech synthesis | 2025-11-27 |
| The Qwen Voice-Enrollment model is a series of voice replication models from the qwen speech model. It can quickly replicate highly similar voices using audio of only 5 seconds or more. When used in conjunction with the qwen3-tts-vc-realtime model, it can replicate a person's voice with high fidelity and output speech in 10 languages. Furthermore, the synthesized audio can adaptively adjust its tone according to the text and has good processing capabilities for complex text synthesis. |
Speech recognition | 2025-11-21 |
| The Fun ASR version, updated in April 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. This version is equivalent to the snapshot released on November 7, 2025. |
Visual understanding | 2025-11-20 |
| Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
Text generation | 2025-11-19 |
| Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
Speech recognition | 2025-11-17 |
| The large file transcription version of Qwen3-ASR-Flash. Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in multiple languages, ensuring precise transcription even in complex audio environments. |
Realtime speech recognition | 2025-11-17 |
| This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments.This is a snapshot released on November 7, 2025. |
Speech recognition | 2025-11-17 |
| The Fun ASR version, updated in April 2026, fully supports the seven major dialect systems of traditional Chinese (Mandarin/Wu/Xiang/Gan/Hakka/Min/Yue) and is compatible with over 20 regional Mandarin accents. It features specific optimizations for the rhythm, meter, and literary expression characteristics of classical Chinese poetry, improving the accuracy of poetry recognition and making it suitable for cultural heritage, educational explanations, and audiobooks. Improved punctuation prediction and text normalization capabilities make the output text more consistent with written expression habits; numbers, dates, and amounts are automatically converted to standard formats, enhancing readability and professionalism. The language support has been expanded to include English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, and Slovak, totaling 30 languages. This is a snapshot released on November 7, 2025. |
Speech synthesis | 2025-11-17 |
| Synthesis Capabilities: CosyVoice-v3-Flash is the latest high-performance speech synthesis model in the CosyVoice series from Tongyi Labs, offering improved naturalness, timbre, prosody, and emotional expressiveness compared to previous versions. This model supports real-time streaming text-to-speech synthesis. Cloning Capabilities: CosyVoice-v3-Flash is also the latest speech cloning model in the CosyVoice series from Tongyi Labs. Compared to previous versions, it improves pronunciation accuracy and timbre similarity, and adds support for more less commonly spoken languages (German, Spanish, French, Italian, Russian, Japanese). It can quickly generate highly similar and naturally sounding custom voices from just 5-20 seconds of reference audio. |
Text generation, Reasoning | 2025-11-10 |
| Kimi-k2-thinking is a general agentic reasoning model developed by Moonshot AI, specialized in deep reasoning and multi-step tool integration to solve complex problems. |
Text generation | 2025-11-06 |
| Qwen-MT-Flash, a large language model from the Qwen series, has been fully upgraded with the Qwen 3 architecture for significantly enhanced performance and translation quality. It provides rapid, cost-effective translation across 92 languages, while supporting advanced features such as terminology intervention, format preservation, and domain-specific adaptation. It is the ideal choice for applications requiring a powerful balance of speed, quality, and cost. |
Image generation | 2025-10-30 |
| The qwen series of image editing Plus models further optimizes inference performance and system stability based on the initial Edit model, significantly reducing the response time for image generation and editing. It also supports returning multiple images in a single request, greatly enhancing user experience. |
Realtime speech recognition | 2025-10-27 |
| The real-time version of Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves highly accurate speech recognition, automatically determining the language and accurately identifying speech in 11 languages, while ensuring precise transcription even in complex audio environments. |
Reasoning, Visual understanding | 2025-10-21 |
| The largest dense model in the Qwen3-VL series, its reasoning version boasts multimodal reasoning capabilities second only to Qwen3-VL-235B-Thinking. It excels in STEM and math problem-solving, general image and video understanding, and achieves state-of-the-art performance in multimodal agent capabilities, making it ideal for complex multimodal reasoning tasks. |
Visual understanding | 2025-10-21 |
| The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios. |
Text generation, Reasoning | 2025-10-21 |
| GLM's new flagship model with comprehensive capability improvements over 4.5, featuring a 200K context window. |
Reasoning, Visual understanding | 2025-10-14 |
| The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
Reasoning, Visual understanding | 2025-09-30 |
| The Thinking version of the 8B Dense model in the Qwen3-VL series consumes less GPU memory and is capable of performing multimodal understanding and reasoning. It supports extremely long contexts such as lengthy videos and documents, 2D/3D visual localization, and features comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities. |
Visual understanding | 2025-09-30 |
| The Instruct version of the 8B Dense model in the Qwen3-VL series requires less GPU memory and provides comprehensively upgraded image/video understanding, support for extremely long contexts like lengthy videos and documents, spatial awareness, and object recognition capabilities, making it suitable for tackling complex real-world tasks. |
Reasoning, Visual understanding | 2025-09-30 |
| The Thinking version of the second-largest MoE model in the Qwen3-VL series features fast response speeds and enhanced multimodal understanding and reasoning capabilities, visual agents, and support for extremely long contexts such as lengthy videos and documents. It also boasts comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities, making it well-suited for complex real-world tasks. |
Visual understanding | 2025-09-30 |
| The Instruct version of the second-largest MoE model in the Qwen3-VL series offers rapid response speeds and supports extremely long contexts like lengthy videos and documents. It includes comprehensively upgraded image/video understanding, spatial awareness, and object recognition capabilities, as well as 2D/3D visual localization, enabling it to handle intricate real-world challenges. |
Text generation, Reasoning | 2025-09-30 |
| Experimental version introducing DeepSeek Sparse Attention (a sparse attention mechanism), exploring optimization and validation for training and inference efficiency on long texts. |
Speech recognition | 2025-09-25 |
| Fun's multilingual speech recognition model supports over 31 languages and allows for free language switching, making it the top choice for users expanding overseas, especially to Southeast Asia. Fun-asr is an upgraded version of this model; switching to Fun-asr is recommended. |
Image generation | 2025-09-23 |
| The upgraded Wan2.5 Preview image edit model, newly upgraded model architecture supports rich image editing capabilities via instruction control, with enhanced instruction adherence. It also enables multi-image reference generation with high consistency and demonstrates excellent text generation performance. |
Reasoning, Visual understanding | 2025-09-23 |
| The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos.This version is a snapshot as of September 23, 2025 |
Reasoning, Visual understanding | 2025-09-23 |
| Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade. |
Visual understanding | 2025-09-23 |
| The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement. |
Text generation, Reasoning | 2025-09-23 |
| The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
Realtime speech translation | 2025-09-23 |
| The real-time version of Qwen3-LiveTranslate-Flash, which is a high-precision, highly responsive, and robust multilingual simultaneous audio and video interpretation model. Leveraging Qwen3-Omni's powerful infrastructure, massive multimodal data, cross-language and cross-modal alignment, and visual enhancement technologies, Qwen3-LiveTranslate-Flash offers both offline and real-time audio and video translation capabilities. It can understand 19 languages and speak 10 languages, including 8 Chinese dialects. |
Text generation | 2025-09-23 |
| Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. This is a snapshot from 23 September, 2025.Compared to the previous version (snapshot from July 22), it demonstrates improved robustness in downstream task performance and tool invocation, along with enhanced code security. |
Visual understanding | 2025-09-23 |
| Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
Image generation | 2025-09-23 |
| The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. |
Realtime speech recognition | 2025-09-23 |
| This is the real-time version of Tongyi Lab's next-generation end-to-end speech recognition model, based on leading proprietary speech technology, and boasts exceptional contextual awareness and high-precision speech transcription capabilities. Based on an end-to-end architecture, Fun-ASR integrates innovative RAG technology, supporting multi-dimensional features such as large-scale hotword customization, automatic filtering of sensitive and modal particles, ITN normalization, and punctuation prediction, significantly improving overall recognition accuracy and contextual relevance. Furthermore, Fun-ASR supports flexible switching between Chinese and English, covers multiple regional dialects, and boasts enhanced noise robustness, adapting to diverse and complex environments. |
Image generation | 2025-09-22 |
| The first Qwen image editing model extends Qwen-Image's text rendering to editing tasks. It offers precise bilingual (Chinese/English) text editing, dual visual and semantic editing, and strong cross-benchmark performance. |
Video generation | 2025-09-19 |
| The upgraded Wan2.5 Preview text to video model, newly upgraded model architecture supports synchronized audio generation with visuals, enables 10-second long video generation, and offers enhanced instruction adherence, improved motion capabilities, and superior image quality. |
Image generation | 2025-09-19 |
| The upgraded Wan2.5 Preview text to image model, newly upgraded model architecture significantly enhances visual aesthetics, design sensibility, and realistic texture. It excels in precise instruction adherence, generates text proficiently in English, Chinese, and less common languages, and supports the generation of complex structured long texts, charts, and architectural diagrams. |
Video generation | 2025-09-19 |
| The upgraded Wan2.5 Preview image to video model, newly upgraded model architecture supports synchronized audio generation with visuals, enables 10-second long video generation, and offers enhanced instruction adherence, improved motion capabilities, and superior image quality. |
Video generation | 2025-09-19 |
| wan2.2-animate-move is a character animation generation model. Users simply upload a character photo and a reference performance video, and the model transfers the expressions and actions from the video onto the character in the image, producing a high-fidelity animated video. |
Video generation | 2025-09-19 |
| wan2.2-animate-mix is a character replacement model product. By uploading a character photo and a performance video, users can accurately replace the character in the original video with the character from the photo, while completely preserving environmental details such as the scene, lighting, and color tone of the original video. |
Realtime omni-modal | 2025-09-17 |
| The real-time version of the Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience. |
Omni-modal | 2025-09-17 |
| Qwen3-Omni-Flash multimodal large-scale model, based on the Thinker–Talker Mixed Expert (MoE) architecture, supports efficient understanding and speech generation of text, images, audio, and video. It can interact with text in 119 languages and speech in 20 languages, generating human-like speech for precise cross-lingual communication. The model boasts powerful command-following and system prompt customization capabilities, flexibly adapting to conversational styles and character settings. It is widely used in scenarios such as text creation, voice assistants, and multimedia analysis, providing a natural and smooth multimodal interaction experience.This version is a snapshot version from September 15, 2025. |
Speech recognition | 2025-09-17 |
| Qwen3-Omni-30b-a3b-Captioner is a powerful fine-grained audio analysis model designed to generate accurate and comprehensive content descriptions in complex and changing audio scenarios. It can automatically parse and describe various audio content, from complex speech and ambient sounds to music and film and television sound effects, and can maintain stable and reliable output even in multi-source and mixed environments. |
Speech synthesis | 2025-09-16 |
| The Qwen3-TTS-Flash-Realtime-2025-09-18 model is Tongyi's latest real-time speech synthesis foundation model, featuring 17 expressive voices while delivering low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on extensive datasets, the system autonomously adjusts vocal tones based on text semantics and demonstrates robust capabilities for complex content synthesis.This model is provided as a snapshot version. |
Speech synthesis | 2025-09-16 |
| The Qwen 3-TTS-Flash-2025-09-18 is Tongyi's latest offline text-to-speech foundation model, featuring 17 expressive voices while enabling low-latency, high-stability audio synthesis. It supports multilingual and dialect outputs with consistent voice characteristics across languages. Trained on massive datasets, the system automatically adjusts vocal tones based on text semantics and demonstrates robust capabilities for synthesizing complex content. This model is provided as a snapshot version. |
Text generation, Reasoning | 2025-09-11 |
| A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507). |
Text generation | 2025-09-11 |
| A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507). |
Text generation, Reasoning | 2025-09-11 |
| As of the September 11, 2025 snapshot, this release features enhanced instruction following and streamlined summary responses in thinking mode. Non-thinking mode provides superior Chinese text understanding and augmented logical reasoning capabilities. The model supports a 1M context length with tiered pricing. |
Speech recognition | 2025-09-08 |
| Qwen3-ASR-Flash is a highly accurate, intelligent, and robust multilingual speech recognition model based on a large language model. Leveraging a powerful foundational model, massive amounts of text and multimodal data, and tens of millions of hours of audio data, Qwen3-ASR-Flash achieves high-precision speech recognition. It can automatically determine the language and accurately recognize speech in 11 languages, ensuring precise transcription even in complex audio environments. |
Text generation, Reasoning | 2025-09-05 |
| A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains. |
Speech synthesis | 2025-09-03 |
| Cloning capability: CosyVoice-v3-plus is the latest large voice cloning model in the CosyVoice series from Tongyi Lab. It offers superior sound quality and cloning fidelity, ideal for professional scenarios. With just 5-20 seconds of reference audio, it can rapidly generate a highly similar and natural-sounding custom voice. Synthesis capability: CosyVoice-v3-plus is the latest large speech synthesis model in the CosyVoice series from Tongyi Lab. It features enhanced sound quality and expressiveness, ideal for professional scenarios. The model supports real-time, streaming text-to-speech synthesis. |
Video generation | 2025-08-25 |
| Auxiliary model for wan2.2-s2v to validate input portrait images against required specifications, ensuring compatibility with video generation via wan2.2-s2v. |
Video generation | 2025-08-25 |
| High-quality Video Generation model that creates dynamic character videos (speaking/singing/performing) from input character images and voice audio files. |
Text generation, Reasoning | 2025-08-23 |
| A hybrid inference architecture model supporting both thinking mode and non-thinking mode, featuring higher reasoning efficiency and stronger agent capabilities. |
Image generation | 2025-08-22 |
| Qwen-MT-Image specializes in image translation, converting images across 11 languages (Chinese, English, Japanese, etc.) to target languages. It accurately preserves layout and content, supporting custom features like terminology definitions, sensitive word filtering, and product detection for flexible, accurate image localization. |
Text generation | 2025-08-22 |
| Qwen-Deep-Research is an advanced agent system for complex research tasks, equipped with multi-step reasoning and global planning capabilities. It leverages tools like internet search to perform detailed task decomposition, reasoning, and analysis, ultimately generating traceable, logically rigorous research reports. |
Speech recognition | 2025-08-22 |
| This is a new-generation large speech recognition model that focuses on Mandarin Chinese, English, and Japanese. It supports a wide range of regional dialects, offers enhanced noise robustness, and adapts to diverse and complex environments, making it the top recommendation for users in China.This is a snapshot released on August 25, 2025. |
Image generation | 2025-08-13 |
| The first image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese. |
Text generation | 2025-08-05 |
| Based on Qwen3, this code generation model inherits the coding agent capabilities of Qwen3-Coder-Plus and supports multi-turn tool interaction. It features focused optimizations on repository-level understanding and enhanced tool-calling stability. |
Text generation, Reasoning | 2025-08-05 |
| The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
Text generation | 2025-07-31 |
| Qwen3-based code generation model that inherits the coding agent ability of Qwen3-Coder-480B-A35B-Instruct; code capability reaches SOTA at the same scale. |
Text generation, Reasoning | 2025-07-31 |
| The Qwen series model optimized for balanced performance, offering inference efficiency between Qwen-Max and Qwen-Turbo, is designed to handle moderately complex tasks effectively. This dynamically updated version implements changes without prior notice. |
Text generation, Reasoning | 2025-07-30 |
| Open-source Qwen3 thinking model; compared to the previous version (Qwen3-30B-A3B) excels in complex thinking tasks, including logic, math, science, code, and other challenging scenarios; instruction following, text understanding, and multilingual translation capabilities significantly improved. |
Text generation, Reasoning | 2025-07-30 |
| Qwen3 series Plus model, integrates thinking and non-thinking modes and can switch modes during dialogue. Compared to the prior version, adds dedicated enhancements for Chinese & English capabilities and tool calling. This is a snapshot from 28 July, 2025; first to support 1 M context length, and uses tiered pricing. |
Text generation | 2025-07-29 |
| Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-30B-A3B) shows major improvements in Chinese, English, and overall multilingual general capabilities. Optimized for subjective open-ended tasks, delivering responses significantly more aligned with user preferences and more helpful. |
Video generation | 2025-07-28 |
| The upgraded Wan 2.2 Plus text to video model, delivers higher quality results with stable sweeping complex movements, cinematic vision control, more powerful prompt following, and realistic world recreation. |
Image generation | 2025-07-28 |
| The upgraded Wan 2.2 Plus text to image model, delivers richer image detail with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. |
Image generation | 2025-07-28 |
| The upgraded Wan 2.2 Flash text to image model, delivers faster speed with enhanced creativity, stability, and realism. It also features stronger prompt following and native support for multiple styles. Up to 2 million pixel generation and prompt enhancement are supported as well. |
Video generation | 2025-07-28 |
| The upgraded Wan 2.2 Plus image to video model, delivers higher quality results with optimized stability, more powerful prompt following, improved consistency for text, portraits, and products, and precise shot control. |
Text generation, Reasoning | 2025-07-25 |
| Open-source Qwen3 thinking model; compared to the previous version (Qwen3-235B-A22B) shows major improvements in logical ability, general capabilities, knowledge enhancement, and creativity, suitable for high-difficulty, strong-thinking scenarios. |
Text generation | 2025-07-23 |
| Qwen-Doc-Turbo rapidly extracts precise information from documents, supporting tagging, classification, content moderation, and summarization. |
Text generation | 2025-07-22 |
| Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. |
Text generation | 2025-07-22 |
| Qwen3-based code generation model with strong coding agent power; code capability reaches open-source SOTA. |
Text generation | 2025-07-22 |
| Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-235B-A22B) shows slight improvements in subjective creativity and model safety. |
Text generation | 2025-07-22 |
| Qwen-MT-Turbo is a large language model within the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 92 languages at a cost-effective price point. It also offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
Text generation | 2025-07-22 |
| Qwen-MT-Plus, the flagship translation model from our Qwen series, is now fully upgraded with the Qwen3 architecture. It supports 92 languages and delivers exceptionally accurate and natural-sounding translations. Its advanced capabilities in contextual understanding, terminology control, and format preservation make it a superior choice over traditional models, especially for specialized domains. |
Realtime speech synthesis | 2025-07-16 |
| Qwen-TTS-Realtime is a speech synthesis model in the Qwen series with bidirectional context awareness. It enables low-latency, high-fidelity generation of multi-voice, dialectal, and long-text bidirectional streaming. |
Text generation, Reasoning | 2025-07-16 |
| The Plus model of the Qwen3 Series, achieving effective integration of thinking mode and non-thinking mode, allows switching modes during conversations. This is a snapshot from July 14, 2025. Compared to the previous version, there has been a significant improvement in both Chinese and English capabilities under non-thinking mode, with enhanced tool-calling abilities. |
Text generation | 2025-07-16 |
| Kimi-K2 is Moonshot's first open-source trillion-parameter MoE model in China, featuring 32B activated parameters with exceptional coding and tool-calling capabilities. |
Speech synthesis | 2025-06-26 |
| Qwen-TTS is the first speech synthesis model in the Qwen series, supporting Chinese, English, and mixed inputs. It adaptively adjusts output tone based on text, offers natural voice quality, and supports streaming output. |
Text generation, Reasoning | 2025-06-24 |
| The Turbo model of the Qwen3 series. It effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities rival those of QwQ-32B with a smaller parameter size, while its general capabilities significantly surpass those of Qwen2.5-Turbo, achieving the SOTA level in the same scale within the industry. |
Text generation, Reasoning | 2025-06-24 |
| Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Reasoning | 2025-06-18 |
| A large language model enhanced through extensive reinforcement learning during post-training, achieving strong reasoning capabilities with minimal labeled data. Performs well in mathematics, coding, and natural language reasoning tasks. |
Visual understanding | 2025-06-13 |
| Qwen-VL-Plus is the enhanced version of the large visual language model. It significantly improves detail recognition and text recognition capabilities, supporting images with resolutions exceeding one million pixels and any aspect ratio specifications. The model delivers exceptional performance across a wide range of visual tasks. |
Text embedding | 2025-06-05 |
| The General Text Vector V4 version is a multi-language text vector model developed by the Tongyi Lab based on Qwen3. Compared to the V3 version, it significantly improves performance in text retrieval, clustering, and classification tasks. It achieves a 15% to 40% improvement in evaluation tasks such as MTEB multilingual, Chinese-English, and code retrieval. Additionally, it supports user-defined vector dimensions ranging from 64 to 2048. |
Reasoning | 2025-06-04 |
| A large language model enhanced through extensive reinforcement learning during post-training, achieving strong reasoning capabilities with minimal labeled data. Performs well in mathematics, coding, and natural language reasoning tasks. |
Reasoning, Visual understanding | 2025-06-03 |
| Qwen QVQ Visual Reasoning Model Plus version supporting visual input and chain-of-thought output with enhanced capabilities in mathematics, programming, visual analysis, creation, and general tasks. |
Speech synthesis | 2025-05-27 |
| A generative speech synthesis model combining text understanding and speech generation through large-scale pretrained language models, supporting real-time streaming text-to-speech synthesis. |
Visual understanding | 2025-05-26 |
| Qwen-VL-Max is a large-scale visual language model of the Qwen series. Compared to the Plus version, it further enhances visual reasoning capabilities and instruction-following abilities, offering higher levels of visual perception and cognition. It delivers optimal performance on more complex tasks. |
Video generation | 2025-05-14 |
| Wanxiang 2.1 - Video Editing Unified Model - Plus. Supports localized editing, video redrawing, background expansion, duration extension, and reference-based generation. Enables multi-modal control through text, images, or videos. |
Realtime omni-modal | 2025-05-08 |
| The real-time version of Qwen's new large multimodal understanding and generation model, suitable for real-time audio interaction scenarios. It supports the understanding of audio accompanied by text, images, and video mixed inputs, and can simultaneously generate speech and text in stream, providing four natural tones. |
Text generation, Reasoning | 2025-04-29 |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-7B. |
Text generation, Reasoning | 2025-04-29 |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-32B-Instruct, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-04-29 |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability rivals QwQ-32B with a smaller parameter size, and its general capability significantly surpasses Qwen2.5-14B, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-04-29 |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-72B-Instruct, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-04-29 |
| The Plus model of the Qwen3 series, effectively integrates thinking mode and non-thinking mode, allowing for mode switching during conversations. Its reasoning capabilities significantly surpass those of QwQ, and its general capabilities notably exceed those of Qwen2.5-Plus, reaching the SOTA level in the same scale within the industry. This model is the snapshot from April 28, 2025. |
Text generation, Reasoning | 2025-04-28 |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-14B. |
Image generation | 2025-04-28 |
| A high-quality virtual try-on image generation model that produces try-on effects with enhanced image clarity, garment texture details, and logo reconstruction compared to aitryon. Requires longer generation time, suitable for non-time-critical scenarios. |
Visual understanding | 2025-04-23 |
| Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
Video generation | 2025-04-21 |
| Wanxiang 2.1 - Keyframe-to-Video - Plus. Generates smooth transition videos from two input images. Supports complex large-scale motion, physics adherence, diverse artistic styles, and cinematic-grade visual quality. Enhanced instruction-following capability and richer visual details. |
Speech synthesis | 2025-04-20 |
| Qwen-TTS is the first speech synthesis model in the Qwen series, supporting Chinese, English, and mixed inputs. It adaptively adjusts output tone based on text, offers natural voice quality, and supports streaming output. |
Omni-modal | 2025-03-27 |
| The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. This model is the snapshot from March 26, 2025, with significant improvements on visual capabilities over the snapshot of January 19, 2025. |
Omni-modal | 2025-03-26 |
| The new multi-modal understanding and generation model trained based on Qwen2.5. It supports text, image, speech, video, and mixed input understanding and can simultaneously generate streams of text and speech, significantly improves the speed of multi-modal content understanding. It provides four natural tones. |
Reasoning, Visual understanding | 2025-03-26 |
| The Tongyi Qianwen QVQ visual reasoning model supports visual input and chain-of-thought output, demonstrating stronger capabilities in mathematics, programming, visual analysis, creation, and general tasks. |
Image generation | 2025-03-25 |
| Image Editing model supporting preset and command-based tasks, including global/local editing (style transfer, inpainting, expansion, super-resolution) and reference-based generation. |
Text generation | 2025-03-20 |
| The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
Text embedding | 2025-03-20 |
| A multilingual text reranking model for semantic retrieval and RAG, ranking candidate documents by relevance to queries. |
Text generation | 2025-03-19 |
| Qwen-Long is a large language model designed for ultra-long context processing, supporting Chinese, English, and other languages. It handles up to 10 million tokens (approximately 15 million characters or 15,000 document pages) in dialogues. Integrated with document services, it supports parsing and dialogue for text files (TXT, DOCX, PDF, XLSX, EPUB, MOBI, MD, CSV) and image files (BMP, PNG, JPG/JPEG, GIF, PDF scans). Notes: HTTP requests support up to 1M tokens; file submission is recommended for longer content. |
Omni-modal | 2025-03-17 |
| The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. |
Reasoning | 2025-03-05 |
| The enhanced version of the Qwen QwQ reasoning model, trained on the Qwen2.5 model, has significantly improved its reasoning capabilities through reinforcement learning. The model's core metrics in mathematics and coding (e.g., AIME 24/25, LiveCodeBench) as well as some general metrics (e.g., IFEval, LiveBench) have reached the level of the full version of DeepSeek-R1. |
Realtime speech translation | 2025-03-04 |
| A multilingual speech-to-text and translation model providing high-accuracy real-time transcription for 10 languages (including Chinese/English/Japanese/Korean) with cross-lingual translation. |
Video generation | 2025-02-27 |
| Image-to-Video Generation-Turbo model that transforms images into dynamic videos with complex movements and cinematic aesthetics, featuring faster generation speed and enhanced instruction compliance. |
Omni-modal | 2025-02-14 |
| The new multimodal understanding and generation model, supporting text, image, speech, and video input as well as mixed input understanding. It can simultaneously stream text and speech, with significantly enhanced speed for multimodal content understanding. Additionally, it offers four natural dialogue tones. |
Text generation | 2025-02-05 |
| A distilled large language model based on Qwen2.5-Math-7B, trained using DeepSeek R1's outputs. |
Text generation | 2025-02-05 |
| A distilled large language model based on Qwen2.5-32B, trained using DeepSeek R1's outputs. |
Text generation | 2025-02-05 |
| A distilled large language model based on Qwen2.5-14B, trained using DeepSeek R1's outputs. |
Text generation | 2025-02-05 |
| A distilled large language model based on Qwen2.5-Math-1.5B, trained using DeepSeek R1's outputs. |
Text generation | 2025-02-05 |
| A distilled large language model based on Llama-3.1-70B, trained using DeepSeek R1's outputs. |
Text generation | 2025-02-03 |
| The Qwen series of models, which are well-balanced in capabilities, offer reasoning performance and speed that fall between Qwen-Max and Qwen-Turbo, making them suitable for moderately complex tasks. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Text generation, Reasoning | 2025-01-27 |
| A self-developed Mixture-of-Experts (MoE) model with 671B parameters (activating 37B), pre-trained on 14.8T tokens. Demonstrates excellent capabilities in long-text processing, coding, mathematics, encyclopedic knowledge, and Chinese language tasks. |
Video generation | 2025-01-20 |
| Image-to-Video Generation-Plus model that converts images into dynamic videos with complex motions, physical realism, cinematic quality, and improved instruction adherence for higher video fidelity. |
Image generation | 2025-01-20 |
| Enhanced Text-to-Image model specializing in realistic portraits and creative designs, upgraded in aesthetics, realism, and artistic quality with up to 2MP resolution and smart prompt rewriting support. |
Video generation | 2025-01-16 |
| Emoji is a facial animation video generation model that creates facial animation videos using face images and preset dynamic templates. |
Video generation | 2025-01-16 |
| Emoji-Detect is an image detection model assisting emoji generation, used to detect whether character images meet video generation requirements. |
Text generation | 2025-01-15 |
| Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Image generation | 2025-01-15 |
| An image segmentation model serving as a supporting module for AI try-on OutfitAnyone, capable of segmenting model images and garment images for pre/post-processing of try-on images. |
Video generation | 2025-01-09 |
| Wanxiang 2.1 - Text-to-Video - Turbo. High-speed video generation with support for complex motion, physics simulation, artistic styles, and cinematic quality. Improved instruction-following capability. |
Video generation | 2025-01-09 |
| Wanxiang 2.1 - Text-to-Video - Plus. Generates high-quality videos from text prompts. Supports complex motion, realistic physics simulation, diverse artistic styles, and cinematic visuals. Enhanced instruction-following performance. |
Image generation | 2025-01-09 |
| Wanxiang 2.1 - Text-to-Image - Turbo. Accelerated generation speed with enhanced aesthetics, realism, and artistry. Maintains strong semantic understanding, style diversity, and 2MP resolution support. Includes smart prompt rewriting functionality. |
Image generation | 2025-01-09 |
| Wanxiang 2.1 - Text-to-Image - Plus. Upgraded image generation with enhanced aesthetics, realism, and artistry. Improved semantic understanding, style generalization, and support for up to 2MP resolution. Features smart prompt rewriting and richer visual details. |
Realtime speech recognition | 2024-12-31 |
| Recommended Paraformer real-time speech recognition model supporting multilingual switching for live streaming/meetings. Enables language selection via language_hints parameter for enhanced accuracy in 8kHz customer service scenarios. Supports: Mandarin (including dialects), English, Japanese, Korean. |
Text generation | 2024-12-26 |
| Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Multimodal embedding | 2024-12-23 |
| Multimodal vector model developed by Tongyi Lab based on pre-trained multimodal foundation models. Generates high-dimensional continuous vectors from text, image, or video inputs for downstream tasks including search, classification, and content moderation. |
Speech recognition | 2024-12-19 |
| Latest Paraformer Mandarin speech recognition model with enhanced architecture, improved recognition accuracy, and 8kHz telephone speech support (Chinese hotword only). |
Text generation | 2024-12-12 |
| Intent Detection and Slot Filling model for dialogue systems, enabling joint prediction of API-based intents and slot parameters in a single output, returning standardized JSON results with multiple API commands and filled slots. |
Video generation | 2024-12-10 |
| Character Video Generation model that synthesizes lip-synced videos based on input character videos and voice audio, matching mouth movements to the audio content. |
Video generation | 2024-12-10 |
| A motion template generation model assisting AnimateAnyone, capable of extracting human motions from videos and creating templates. |
Video generation | 2024-12-10 |
| A video generation model that creates full-body motion videos based on human portraits and motion templates. |
Video generation | 2024-12-10 |
| An image detection model assisting AnimateAnyone, used to verify if human figures in images meet video generation requirements. |
Visual understanding | 2024-11-14 |
| Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
Text generation | 2024-11-12 |
| Qwen-Coder-Plus is a specialized language model for programming and code generation, delivering excellent performance and outstanding results. |
Video generation | 2024-11-07 |
| LivePortrait-detect is an auxiliary image detection model for assessing whether human figures in images meet video generation requirements. |
Video generation | 2024-11-07 |
| LivePortrait is a video generation model that creates lightweight dynamic portrait videos from static images. |
Video generation | 2024-11-07 |
| EMO is a video generation model that generates high-quality dynamic portrait videos based on character images. |
Video generation | 2024-11-07 |
| EMO-Detect is an image detection model assisting EMO, used to detect whether character images meet video generation requirements. |
Text generation | 2024-10-15 |
| Qwen-Max supports a parameter scale of hundreds of billions and multiple input languages such as Chinese and English. Qwen-Max is updated in a rolling manner. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Text generation | 2024-09-19 |
| Qwen-Math-Turbo is a specialized language model for math problem-solving, featuring high inference speed and low cost. |
Text generation | 2024-09-19 |
| Qwen-Math-Plus is a powerful math problem-solving model excelling in Chinese and English mathematical tasks, including equations, calculations, and proofs. |
Text generation | 2024-09-19 |
| Qwen-Coder-Turbo is a specialized language model for programming and code generation, featuring high inference speed and low cost. |
Speech synthesis | 2024-09-13 |
| A large-model voice replication service used in conjunction with Cosyvoice-v3. Utilizing advanced large-model technology for feature extraction, it can replicate voices without a training process. Only a very short audio clip is required to quickly generate a highly similar and natural-sounding custom voice. |
Video generation | 2024-09-13 |
| Video Style Repainting model that applies artistic styles (e.g., Japanese comic, 3D cartoon) to input video frames while preserving original appearances, generating stylized videos with diverse visual effects. |
Text generation | 2024-09-13 |
| Qwen-Math-Plus is a powerful math problem-solving model excelling in Chinese and English mathematical tasks, including equations, calculations, and proofs. |
Speech synthesis | 2024-09-13 |
| A voice cloning model leveraging advanced large model technology for feature extraction without training. Generates highly similar, natural-sounding custom voices from short audio samples. |
Speech recognition | 2024-09-06 |
| Recommended Paraformer speech recognition model supporting multilingual recognition with language_hints parameter. Supports all sampling rates and hotwords across: Mandarin (including dialects), English, Japanese, Korean. |
Realtime speech recognition | 2024-09-06 |
| Recommended Paraformer real-time speech recognition model supporting multilingual switching with language_hints parameter. Supports all sampling rates and hotwords across: Mandarin (including dialects), English, Japanese, Korean. |
Image generation | 2024-08-19 |
| Person instance segmentation employs detection and segmentation techniques to identify objects in images and generate pixel-level masks for precise object boundary delineation. |
Image generation | 2024-08-19 |
| Image erasure and completion tool removing specified elements (people, objects, text, watermarks) while preserving backgrounds using computer vision and AIGC inpainting techniques. |
Hong Kong (China)
Model type | Date | Service scope | Model ID | Description |
|---|---|---|---|---|
Text generation, Deep thinking, Visual understanding | 2026-08-26 | Global |
| Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications. |
Video generation | 2026-08-20 | Global |
| Wan3.0-Video-Prime is the high-speed version of the Wan3.0 video generation model, with capabilities aligned with the Wan3.0-Video standard version. It supports four-modal all-reference input and can generate videos up to 30 seconds long, delivering an immersive audio-visual experience with significantly improved end-to-end speed. |
Text generation, Reasoning, Visual understanding | 2026-08-19 | Global |
| Kimi K3 is Kimi's most powerful flagship model with 2.8 trillion parameters. Built on KDA hybrid linear attention (Kimi Delta Attention) and Attention Residuals technologies, it natively supports vision understanding and features a 1 million token context window. It is the world's first open-source 3-trillion-parameter model designed for long-context programming, knowledge work, and advanced reasoning. |
Text generation, Reasoning | 2026-08-14 | Global |
| A flagship Mixture-of-Experts (MoE) large language model with 1.6 trillion total parameters and 49 billion activated parameters. It natively supports context windows of up to 1 million tokens. Trained on extensive high-quality data, the model delivers strong performance in mathematical and logical reasoning, complex reasoning, professional code generation, and in-depth long-document analysis, and is suitable for demanding scenarios such as advanced scientific research, complex enterprise workflows, and sophisticated agentic applications. |
Image generation | 2026-08-06 | Global |
| Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and 20+ fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool. |
Image generation | 2026-08-04 | Global |
| Clear instruction understanding: supports up to 4.5k token input, generating complex image-text instructions in one pass. Stable text rendering: 10px small text, 12 languages, 20+ fonts clearly legible, ready for infographics and interfaces. Handy batch output: posters, web pages, interfaces and other daily tasks can be generated in bulk at lower cost, ideal for continuous creation. Qwen-Image-3.0 Standard pursues not only generation quality but also everyday creative convenience, making image generation a sustainable content productivity tool. |
Text generation, Reasoning, Visual understanding | 2026-08-02 | Global |
| A 2.4-trillion-parameter MoE flagship with a comprehensive leap in coding and office capabilities, capable of autonomously coding for days to deliver complete projects. It handles hundreds of professional tasks such as legal, financial, and design work, delivering production-grade results end-to-end in a single conversation. Native visual understanding runs through the full planning, execution, and verification pipeline, supporting deep semantic parsing of ultra-long documents and long videos. It performs autonomous planning and closed-loop iteration in long-horizon tasks, continuously evolving. |
Text generation, Reasoning | 2026-08-01 | Global |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Text generation, Reasoning, Visual understanding | 2026-07-21 | Global |
| The Qwen3.7 native vision-language series Flash models comprehensively enhance multimodal understanding and Agent execution capabilities compared to 3.6-Flash. Key improvements include strengthened foundational multimodal abilities, enhanced object recognition, improved real-world perception and spatial intelligence. Multimodal Agent scenarios such as Search Agent and CI Agent have seen significant upgrades, with more stable end-to-end task execution. Multimodal coding capabilities are optimized, delivering a smoother vibe coding experience. |
Text generation | 2026-07-02 | Global |
| The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
Text generation, Reasoning | 2026-06-16 | Global |
| GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
Text generation, Reasoning, Visual understanding | 2026-06-15 | Global |
| kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
Text generation, Reasoning, Visual understanding | 2026-06-09 | Global |
| The Max model, the largest and most capable in the Qwen3.7 series, has added visual‑modal understanding compared to the May 20 snapshot, enabling it to perceive real‑world scenes and supporting multimodal interactive hybrid agent capabilities. This version is based on a snapshot taken on June 8, 2026. |
Text generation, Reasoning, Visual understanding | 2026-06-01 | Global |
| Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
Text generation, Reasoning | 2026-05-20 | Global |
| The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
Text generation, Reasoning, Visual understanding | 2026-04-17 | Global |
| The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
Text generation, Reasoning, Visual understanding | 2026-04-01 | Global |
| The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | Hong Kong |
| The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
Text generation, Reasoning | 2026-01-23 | Hong Kong |
| Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model's overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026. |
Reasoning, Visual understanding | 2025-12-18 | Hong Kong |
| The Qwen3 series of visual understanding models effectively integrates thinking and non-thinking modes. Compared to the snapshot released on September 23, this version delivers superior performance in reasoning and analysis tasks as well as style control, while also offering lower latency and faster response speeds. This version is based on a snapshot taken on December 19, 2025. |
— | 2025-12-01 | Hong Kong |
| This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
Text generation, Reasoning | 2025-09-24 | Hong Kong |
| The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
Visual understanding | 2025-09-23 | Hong Kong |
| The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
Text generation, Reasoning | 2025-09-16 | Hong Kong |
| Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Text embedding | 2025-08-25 | Hong Kong |
| The General Text Vector V4 version is a multi-language text vector model developed by the Tongyi Lab based on Qwen3. Compared to the V3 version, it significantly improves performance in text retrieval, clustering, and classification tasks. It achieves a 15% to 40% improvement in evaluation tasks such as MTEB multilingual, Chinese-English, and code retrieval. Additionally, it supports user-defined vector dimensions ranging from 64 to 2048. |
Germany (Frankfurt)
Model type | Date | Service scope | Model ID | Description |
|---|---|---|---|---|
Text generation, Deep thinking, Visual understanding | 2026-08-26 | Global |
| Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications. |
Video generation | 2026-08-20 | Global |
| Wan3.0-Video-Prime is the high-speed version of the Wan3.0 video generation model, with capabilities aligned with the Wan3.0-Video standard version. It supports four-modal all-reference input and can generate videos up to 30 seconds long, delivering an immersive audio-visual experience with significantly improved end-to-end speed. |
Text generation, Reasoning, Visual understanding | 2026-08-19 | Global |
| Kimi K3 is Kimi's most powerful flagship model with 2.8 trillion parameters. Built on KDA hybrid linear attention (Kimi Delta Attention) and Attention Residuals technologies, it natively supports vision understanding and features a 1 million token context window. It is the world's first open-source 3-trillion-parameter model designed for long-context programming, knowledge work, and advanced reasoning. |
Text generation, Reasoning | 2026-08-14 | Global |
| A flagship Mixture-of-Experts (MoE) large language model with 1.6 trillion total parameters and 49 billion activated parameters. It natively supports context windows of up to 1 million tokens. Trained on extensive high-quality data, the model delivers strong performance in mathematical and logical reasoning, complex reasoning, professional code generation, and in-depth long-document analysis, and is suitable for demanding scenarios such as advanced scientific research, complex enterprise workflows, and sophisticated agentic applications. |
Video generation | 2026-08-06 | Global |
| Wan3.0 is an all-in-one video generation model that uniformly supports multiple creative functions including reference, editing, replication, and driving. It supports four-modal all-reference input, up to 30 seconds of video generation, and can parse files, web pages, and complex images. With production-grade character consistency and realistic audio and visuals, it delivers an immersive audio-visual impact. |
Image generation | 2026-08-06 | Global |
| Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and 20+ fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool. |
Image generation | 2026-08-04 | Global |
| Clear instruction understanding: supports up to 4.5k token input, generating complex image-text instructions in one pass. Stable text rendering: 10px small text, 12 languages, 20+ fonts clearly legible, ready for infographics and interfaces. Handy batch output: posters, web pages, interfaces and other daily tasks can be generated in bulk at lower cost, ideal for continuous creation. Qwen-Image-3.0 Standard pursues not only generation quality but also everyday creative convenience, making image generation a sustainable content productivity tool. |
Text generation, Reasoning, Visual understanding | 2026-08-02 | Global |
| A 2.4-trillion-parameter MoE flagship with a comprehensive leap in coding and office capabilities, capable of autonomously coding for days to deliver complete projects. It handles hundreds of professional tasks such as legal, financial, and design work, delivering production-grade results end-to-end in a single conversation. Native visual understanding runs through the full planning, execution, and verification pipeline, supporting deep semantic parsing of ultra-long documents and long videos. It performs autonomous planning and closed-loop iteration in long-horizon tasks, continuously evolving. |
Text generation, Reasoning | 2026-08-01 | Global |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Text generation, Reasoning, Visual understanding | 2026-07-21 | Global |
| The Qwen3.7 native vision-language series Flash models comprehensively enhance multimodal understanding and Agent execution capabilities compared to 3.6-Flash. Key improvements include strengthened foundational multimodal abilities, enhanced object recognition, improved real-world perception and spatial intelligence. Multimodal Agent scenarios such as Search Agent and CI Agent have seen significant upgrades, with more stable end-to-end task execution. Multimodal coding capabilities are optimized, delivering a smoother vibe coding experience. |
Text generation | 2026-07-02 | Global |
| The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
Video generation | 2026-06-16 | Global |
| HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
Video generation | 2026-06-16 | Global |
| HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
Video generation | 2026-06-16 | Global |
| HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
Text generation, Reasoning | 2026-06-16 | Global |
| GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
Text generation, Reasoning, Visual understanding | 2026-06-15 | Global |
| kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
Text generation, Reasoning, Visual understanding | 2026-06-09 | Global |
| The Max model, the largest and most capable in the Qwen3.7 series, has added visual‑modal understanding compared to the May 20 snapshot, enabling it to perceive real‑world scenes and supporting multimodal interactive hybrid agent capabilities. This version is based on a snapshot taken on June 8, 2026. |
Text generation, Reasoning, Visual understanding | 2026-06-01 | Global |
| Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
Text generation, Reasoning | 2026-05-20 | Global |
| The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
Text generation, Reasoning, Visual understanding | 2026-04-29 | Global |
| Kimi-k2.5 is Moonshot AI's most versatile model with native multimodal architecture, supporting visual/text inputs, thinking/non-thinking modes, and both dialogue and agent tasks. |
Video generation | 2026-04-26 | Global |
| HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
Video generation | 2026-04-26 | Global |
| HappyHorse-1.0-R2V supports Reference-to-Video generation, offering enhanced stability in subject and scene referencing. Capable of processing up to 9 reference images, it precisely preserves creative intent to deliver superior performance. |
Text generation, Reasoning | 2026-04-24 | Global |
| A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
Text generation, Reasoning | 2026-04-24 | Global |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Video generation | 2026-04-22 | Global |
| HappyHorse-1.0-T2V supports text-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends text semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
Video generation | 2026-04-22 | Global |
| HappyHorse-1.0-I2V enables image-to-video generation, featuring highly realistic dynamic rendering. It accurately comprehends both text and image semantics to produce high-quality videos that are fluid, natural, and rich in detail. |
Text generation, Reasoning, Visual understanding | 2026-04-17 | Global |
| The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
Text generation, Reasoning, Visual understanding | 2026-04-17 | Global |
| The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance. |
Text generation, Reasoning | 2026-04-14 | Global |
| GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
Text generation, Reasoning, Visual understanding | 2026-04-01 | Global |
| The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | EU |
| The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | Global |
| The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | Global |
| The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall performance is comparable to that of the Qwen3.5-27B. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | Global |
| The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B. |
Text generation, Reasoning, Visual understanding | 2026-02-23 | Global |
| The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of overall performance, this model is second only to Qwen3.5-397B-A17B. Its text capabilities significantly outperform those of Qwen3-235B-2507, and its visual capabilities surpass those of Qwen3-VL-235B. |
Text generation | 2026-02-20 | EU |
| The new-generation code generation model in the Qwen3 series delivers performance close to that of Qwen3-Coder-Plus while offering even better capabilities. The model has been optimized with a focus on repository-level understanding, supports multi-turn tool interactions, and enhances its compatibility with agentic coding tools. |
Text generation, Reasoning, Visual understanding | 2026-02-15 | Global |
| The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities. |
Text generation, Reasoning, Visual understanding | 2026-02-15 | Global |
| The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent. |
Text generation, Reasoning | 2026-01-23 | EU |
| Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model's overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026. |
Visual understanding | 2026-01-22 | EU |
| The Qwen3 series of small-sized visual understanding models effectively integrates thinking and non-thinking modes. Compared with the snapshot taken on October 15, 2025, the overall performance of the model has improved significantly: it delivers enhanced capabilities in general visual recognition and reasoning, and shows marked improvements in recognition accuracy across various business scenarios such as security, in-store inspections, equipment monitoring, and photo-based problem solving. This version is a snapshot as of January 22, 2026. |
Video generation | 2025-12-16 | Global |
| Wan2.6 reference to video, supports using a specified person or any object as a reference, precisely maintaining consistency of appearance and voice, and allows multi‑character reference for joint performances. |
Image generation | 2025-12-15 | Global |
| Wan2.6 text to image, Upgraded visual quality, aesthetics, and instruction-following deliver precise style control, realistic portraits, long-text understanding, and broad historical/cultural IP coverage, enabling high-quality, highly expressive visual generation. |
Image generation | 2025-12-15 | Global |
| Wan2.6 Image, An all-round image generation model that supports joint text–image reasoning, multi-image creative fusion, commercial-grade consistency, aesthetic style transfer, and precise control of framing and lighting, significantly enhancing consistency, controllability, and expressiveness in image generation. |
Video generation | 2025-12-03 | Global |
| Wan2.6 text to video, Intelligent shot scheduling supports multi-shot storytelling, generating multi-shot narrative videos with consistent subjects, scenes, and atmosphere, with a maximum duration of 15 seconds. It delivers better instruction following, and improved visual fidelity. |
Video generation | 2025-12-03 | Global |
| Wan2.6 image to video, intelligent shot scheduling enables multi‑camera storytelling, delivers higher‑quality voice generation, supports stable multi‑speaker dialogue with more natural and realistic vocal timbres, and supports generation of clips up to 15 seconds in length. |
— | 2025-12-01 | EU |
| This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
— | 2025-12-01 | Global |
| This version is a snapshot as of December 1, 2025, and features improved reasoning capabilities compared to the July 28 snapshot. Agent capabilities and multi-turn tool invocation abilities have been further enhanced, and performance on subjective creative tasks has improved significantly. It supports a context length of up to 1 million tokens, with tiered pricing based on context length. |
Visual understanding | 2025-11-21 | Global |
| This model is a snapshot version from November 20, 2025, and is based on the latest Qwen-VL3 architecture with a comprehensive upgrade. It features significant improvements in document parsing and text localization capabilities, as well as substantial reductions in end-to-end latency and illusions. |
Visual understanding | 2025-11-20 | Global |
| Qwen-VL_OCR is an OCR model trained based on Qwen-VL. It aggregates various image-text recognition, parsing, and processing tasks through a unified model approach, offering powerful image-text recognition capabilities. |
Text generation | 2025-11-19 | Global |
| Qwen-MT-Lite is a large language model of the Qwen model series that specializes in multi-lingual translation. It provides high-quality and rapid translation services across 32 languages at a cost-effective price. It offers features such as terminology intervention, format preservation, and domain-specific translation to cater to the diverse needs of various applications, ensuring both efficiency and performance. |
Text generation | 2025-11-11 | Global |
| Qwen-MT-Flash, a large language model from the Qwen series, has been fully upgraded with the Qwen 3 architecture for significantly enhanced performance and translation quality. It provides rapid, cost-effective translation across 92 languages, while supporting advanced features such as terminology intervention, format preservation, and domain-specific adaptation. It is the ideal choice for applications requiring a powerful balance of speed, quality, and cost. |
Reasoning, Visual understanding | 2025-10-21 | Global |
| The largest dense model in the Qwen3-VL series, its reasoning version boasts multimodal reasoning capabilities second only to Qwen3-VL-235B-Thinking. It excels in STEM and math problem-solving, general image and video understanding, and achieves state-of-the-art performance in multimodal agent capabilities, making it ideal for complex multimodal reasoning tasks. |
Visual understanding | 2025-10-21 | Global |
| The largest dense model in the Qwen3-VL series, in its non-inference version, delivers overall performance second only to Qwen3-VL-235B-Instruct. It excels in document recognition and comprehension, demonstrates strong spatial awareness and object identification capabilities, and achieves state-of-the-art performance in 2D visual detection and spatial reasoning. It is well-suited for complex perception tasks across a wide range of general-purpose scenarios. |
Reasoning, Visual understanding | 2025-10-15 | EU |
| The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
Reasoning, Visual understanding | 2025-10-15 | Global |
| The Qwen3 series of small-scale visual understanding models effectively integrates thinking and non-thinking modes, delivering superior performance compared to the open-source Qwen3-VL-30B-A3B while maintaining fast response speeds. It features a comprehensive upgrade in image/video understanding, supporting ultra-long contexts such as extended videos and documents, spatial awareness, and object recognition across various domains. Equipped with 2D/3D visual localization capabilities, it is well-suited for tackling complex real-world tasks. |
Reasoning, Visual understanding | 2025-10-03 | Global |
| The Thinking version of the second-largest MoE model in the Qwen3-VL series features fast response speeds and enhanced multimodal understanding and reasoning capabilities, visual agents, and support for extremely long contexts such as lengthy videos and documents. It also boasts comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities, making it well-suited for complex real-world tasks. |
Visual understanding | 2025-10-03 | Global |
| The Instruct version of the second-largest MoE model in the Qwen3-VL series offers rapid response speeds and supports extremely long contexts like lengthy videos and documents. It includes comprehensively upgraded image/video understanding, spatial awareness, and object recognition capabilities, as well as 2D/3D visual localization, enabling it to handle intricate real-world challenges. |
Reasoning, Visual understanding | 2025-09-30 | Global |
| The Thinking version of the 8B Dense model in the Qwen3-VL series consumes less GPU memory and is capable of performing multimodal understanding and reasoning. It supports extremely long contexts such as lengthy videos and documents, 2D/3D visual localization, and features comprehensively upgraded image/video understanding, spatial awareness, and object recognition abilities. |
Visual understanding | 2025-09-30 | Global |
| The Instruct version of the 8B Dense model in the Qwen3-VL series requires less GPU memory and provides comprehensively upgraded image/video understanding, support for extremely long contexts like lengthy videos and documents, spatial awareness, and object recognition capabilities, making it suitable for tackling complex real-world tasks. |
Text generation, Reasoning | 2025-09-24 | EU |
| The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
Text generation, Reasoning | 2025-09-24 | Global |
| The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios. |
Visual understanding | 2025-09-23 | EU |
| The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
Visual understanding | 2025-09-23 | Global |
| The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos. |
Reasoning, Visual understanding | 2025-09-23 | Global |
| Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade. |
Visual understanding | 2025-09-23 | Global |
| The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement. |
Text generation | 2025-09-23 | Global |
| Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. This is a snapshot from 23 September, 2025.Compared to the previous version (snapshot from July 22), it demonstrates improved robustness in downstream task performance and tool invocation, along with enhanced code security. |
Text generation, Reasoning | 2025-09-16 | EU |
| Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Text generation, Reasoning | 2025-09-16 | Global |
| Qwen-Plus is an enhanced version of the Qwen ultra-large language model that supports multiple input languages such as Chinese and English. Compared to previous versions, it shows significant improvements in both Chinese and English code generation, logical reasoning, and multilingual abilities. The response style has been greatly adjusted to align with human preferences, with noticeable enhancements in the level of detail and clarity of responses. Specialized improvements have been made in creative writing, adherence to JSON formatting, and role-playing abilities. |
Text generation, Reasoning | 2025-09-11 | Global |
| A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507). |
Text generation | 2025-09-11 | Global |
| A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507). |
Text generation, Reasoning | 2025-09-11 | Global |
| As of the September 11, 2025 snapshot, this release features enhanced instruction following and streamlined summary responses in thinking mode. Non-thinking mode provides superior Chinese text understanding and augmented logical reasoning capabilities. The model supports a 1M context length with tiered pricing. |
Text generation, Reasoning | 2025-09-05 | Global |
| A preview version of the Max model in the Qwen 3 series, achieving an effective integration of thinking and non-thinking modes. In thinking mode, there is a significant enhancement in capabilities such as intelligent agent programming, common-sense reasoning, and reasoning across mathematics, science, and general domains. |
Text generation, Reasoning | 2025-08-01 | Global |
| The Qwen3 Flash model offers a powerful fusion of thinking and non-thinking modes with dynamic in-conversation switching, excelling in complex reasoning while showing significant gains in instruction following and text comprehension. It supports a 1M context length and is billed on a tiered model corresponding to context usage. |
Text generation | 2025-07-31 | Global |
| Qwen3-based code generation model that inherits the coding agent ability of Qwen3-Coder-480B-A35B-Instruct; code capability reaches SOTA at the same scale. |
Text generation, Reasoning | 2025-07-30 | Global |
| Open-source Qwen3 thinking model; compared to the previous version (Qwen3-30B-A3B) excels in complex thinking tasks, including logic, math, science, code, and other challenging scenarios; instruction following, text understanding, and multilingual translation capabilities significantly improved. |
Text generation | 2025-07-29 | Global |
| Based on Qwen3, this code generation model inherits the coding agent capabilities of Qwen3-Coder-Plus and supports multi-turn tool interaction. It features focused optimizations on repository-level understanding and enhanced tool-calling stability. |
Text generation | 2025-07-29 | Global |
| Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-30B-A3B) shows major improvements in Chinese, English, and overall multilingual general capabilities. Optimized for subjective open-ended tasks, delivering responses significantly more aligned with user preferences and more helpful. |
Text generation, Reasoning | 2025-07-29 | Global |
| Qwen3 series Plus model, integrates thinking and non-thinking modes and can switch modes during dialogue. Compared to the prior version, adds dedicated enhancements for Chinese & English capabilities and tool calling. This is a snapshot from 28 July, 2025; first to support 1 M context length, and uses tiered pricing. |
Text generation, Reasoning | 2025-07-25 | Global |
| Open-source Qwen3 thinking model; compared to the previous version (Qwen3-235B-A22B) shows major improvements in logical ability, general capabilities, knowledge enhancement, and creativity, suitable for high-difficulty, strong-thinking scenarios. |
Text generation | 2025-07-24 | Global |
| Qwen-MT-Plus, the flagship translation model from our Qwen series, is now fully upgraded with the Qwen3 architecture. It supports 92 languages and delivers exceptionally accurate and natural-sounding translations. Its advanced capabilities in contextual understanding, terminology control, and format preservation make it a superior choice over traditional models, especially for specialized domains. |
Text generation | 2025-07-23 | Global |
| Qwen3-based code generation model with strong coding agent power, excels at tool calling and environment interaction, capable of autonomous programming with outstanding code capability while maintaining general ability. |
Text generation | 2025-07-23 | Global |
| Qwen3-based code generation model with strong coding agent power; code capability reaches open-source SOTA. |
Text generation | 2025-07-23 | Global |
| Open-source Qwen3 non-thinking model; compared to the previous version (Qwen3-235B-A22B) shows slight improvements in subjective creativity and model safety. |
Text generation, Reasoning | 2025-05-12 | Global |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability reaches the SOTA level in the same scale industry, and its general capability significantly surpasses Qwen2.5-7B. |
Text generation, Reasoning | 2025-05-12 | Global |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-32B-Instruct, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-05-12 | Global |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability rivals QwQ-32B with a smaller parameter size, and its general capability significantly surpasses Qwen2.5-14B, reaching the SOTA level in the same scale industry. |
Text generation, Reasoning | 2025-05-12 | Global |
| Achieves effective integration of thinking and non-thinking modes, allowing mode switching during conversations. Its reasoning capability significantly surpasses QwQ, and its general capability markedly exceeds Qwen2.5-72B-Instruct, reaching the SOTA level in the same scale industry. |
Japan (Tokyo)
Model type | Date | Service scope | Model ID | Description |
|---|---|---|---|---|
Text generation, Deep thinking, Visual understanding | 2026-08-26 | Global |
| Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications. |
Video generation | 2026-08-20 | Global |
| Wan3.0-Video-Prime is the high-speed version of the Wan3.0 video generation model, with capabilities aligned with the Wan3.0-Video standard version. It supports four-modal all-reference input and can generate videos up to 30 seconds long, delivering an immersive audio-visual experience with significantly improved end-to-end speed. |
Text generation, Reasoning, Visual understanding | 2026-08-19 | Global |
| Kimi K3 is Kimi's most powerful flagship model with 2.8 trillion parameters. Built on KDA hybrid linear attention (Kimi Delta Attention) and Attention Residuals technologies, it natively supports vision understanding and features a 1 million token context window. It is the world's first open-source 3-trillion-parameter model designed for long-context programming, knowledge work, and advanced reasoning. |
Text generation, Reasoning | 2026-08-14 | Global |
| A flagship Mixture-of-Experts (MoE) large language model with 1.6 trillion total parameters and 49 billion activated parameters. It natively supports context windows of up to 1 million tokens. Trained on extensive high-quality data, the model delivers strong performance in mathematical and logical reasoning, complex reasoning, professional code generation, and in-depth long-document analysis, and is suitable for demanding scenarios such as advanced scientific research, complex enterprise workflows, and sophisticated agentic applications. |
Video generation | 2026-08-06 | Global |
| Wan3.0 is an all-in-one video generation model that uniformly supports multiple creative functions including reference, editing, replication, and driving. It supports four-modal all-reference input, up to 30 seconds of video generation, and can parse files, web pages, and complex images. With production-grade character consistency and realistic audio and visuals, it delivers an immersive audio-visual impact. |
Image generation | 2026-08-06 | Global |
| Rich content: Supports input of up to 4.5k tokens and dense information layout with images-within-images, enabling complex layouts like newspapers, storyboards, menus, and exam papers to be generated in a single pass. Authentic detail: Supports precise rendering of text as small as 10px, and vividly reproduces fine details such as micro-expressions, pores, and individual strands of hair—approaching the quality of real photography. Deep knowledge: Supports native rendering of 12 languages and 20+ fonts, realistic simulation of mainstream interfaces such as web pages, games, and live streams, fully incorporating external knowledge. Qwen-Image-3.0-Pro isn't just pursuing "good looks"—it's pursuing "usefulness", making image generation a truly deployable productivity tool. |
Image generation | 2026-08-04 | Global |
| Clear instruction understanding: supports up to 4.5k token input, generating complex image-text instructions in one pass. Stable text rendering: 10px small text, 12 languages, 20+ fonts clearly legible, ready for infographics and interfaces. Handy batch output: posters, web pages, interfaces and other daily tasks can be generated in bulk at lower cost, ideal for continuous creation. Qwen-Image-3.0 Standard pursues not only generation quality but also everyday creative convenience, making image generation a sustainable content productivity tool. |
Text generation, Reasoning, Visual understanding | 2026-08-02 | Global |
| A 2.4-trillion-parameter MoE flagship with a comprehensive leap in coding and office capabilities, capable of autonomously coding for days to deliver complete projects. It handles hundreds of professional tasks such as legal, financial, and design work, delivering production-grade results end-to-end in a single conversation. Native visual understanding runs through the full planning, execution, and verification pipeline, supporting deep semantic parsing of ultra-long documents and long videos. It performs autonomous planning and closed-loop iteration in long-horizon tasks, continuously evolving. |
Text generation, Reasoning | 2026-08-01 | Global |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Text generation, Reasoning, Visual understanding | 2026-07-28 | Global |
| kimi-k2.7-code is Kimi's most intelligent coding model to date. It follows instructions more reliably over long contexts and completes programming tasks with higher success rates. It supports text, image, and video inputs, along with thinking mode, conversation, and agent tasks. |
Text generation, Reasoning | 2026-07-28 | Global |
| GLM-5.2 is the latest flagship model from Zhipu AI, designed for long-horizon tasks with support for an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance scenarios. |
Text generation, Reasoning, Visual understanding | 2026-07-21 | Global |
| The Qwen3.7 native vision-language series Flash models comprehensively enhance multimodal understanding and Agent execution capabilities compared to 3.6-Flash. Key improvements include strengthened foundational multimodal abilities, enhanced object recognition, improved real-world perception and spatial intelligence. Multimodal Agent scenarios such as Search Agent and CI Agent have seen significant upgrades, with more stable end-to-end task execution. Multimodal coding capabilities are optimized, delivering a smoother vibe coding experience. |
Text generation | 2026-07-02 | Global |
| The role-playing model of the Qwen series. This is a dynamically updated version, and notifications will be provided in advance for any model updates. It is suitable for anthropomorphic role-playing and has optimized capabilities in following predefined character instructions, advancing conversations, and demonstrating active listening and empathy. Additionally, it supports the deep restoration of personalized characters. |
Video generation | 2026-06-16 | Global |
| HappyHorse-1.1-T2V supports Text-to-Video generation with improved semantic understanding, cinematic shot control, and dynamic motion rendering. It more accurately captures creative intent, producing high-quality videos with smoother motion, richer details, stronger visual consistency, and more natural character actions, scene atmosphere, and physical dynamics. |
Video generation | 2026-06-16 | Global |
| HappyHorse-1.1-R2V supports Reference-to-Video generation with significantly improved stability in subject, scene style, and visual consistency. With support for up to 9 reference images, it can more accurately understand and preserve creative intent, delivering stronger controllability and expressiveness across characters, scenes, styles, and cinematic motion. |
Video generation | 2026-06-16 | Global |
| HappyHorse-1.1-I2V supports Image-to-Video generation with improved visual quality, dynamic performance, and cross-clip consistency. It more accurately understands the input image and preserves creative intent, delivering significant improvements in skin texture realism, character ID consistency across clips, motion smoothness, text rendering stability, and audio-visual synchronization, producing high-quality videos with greater realism, richer details, and stronger overall consistency. |
Text generation, Reasoning, Visual understanding | 2026-06-01 | Global |
| Among the Qwen3.7 series, the cost-effective Plus model builds on its robust text capabilities while delivering a comprehensive upgrade to its vision‑language abilities, all while preserving its full‑stack agent‑level intelligence for coding, tool use, and productivity workflows. Its key distinguishing feature is multi‑modal interactive hybrid agent capabilities, enabling it to perceive real‑world scenes, read screens and interact with GUIs, generate code based on visual references, and perform end‑to‑end navigation within mobile apps. |
Text generation, Reasoning | 2026-05-21 | Global |
| The Max model, the largest and most capable in the Qwen3.7 series, currently offers a pure‑text‑only interface for public experimentation. Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution. |
Text generation, Reasoning | 2026-05-11 | Global |
| A flagship MoE large model with 1.6 trillion parameters and 49 billion activated parameters, natively supporting context lengths of up to one million tokens. Trained on a vast corpus of high-quality data, it excels in advanced mathematical reasoning, complex logical inference, specialized coding, and deep analysis of long-form text, making it well-suited for demanding applications such as cutting-edge research, sophisticated office workflows, and advanced AI agents. |
Text generation, Reasoning | 2026-05-11 | Global |
| A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. |
Text generation, Reasoning, Visual understanding | 2026-04-29 | Global |
| Kimi-k2.5 is Moonshot AI's most versatile model with native multimodal architecture, supporting visual/text inputs, thinking/non-thinking modes, and both dialogue and agent tasks. |
Video generation | 2026-04-26 | Global |
| HappyHorse-1.0-V2V supports advanced video editing through natural language instructions. It allows for local or global editing of video elements using up to 5 reference images, precisely preserving original motion dynamics to achieve superior expressiveness. |
Text generation, Reasoning, Visual understanding | 2026-04-17 | Global |
| The Qwen3.6 native vision-language Flash model series delivers a significant performance boost over the 3.5-Flash version. This model particularly excels in agentic coding capabilities, substantially outperforming its predecessor on multiple code-agent benchmarks, as well as in mathematical and code reasoning. In terms of vision, it features markedly improved spatial intelligence, with especially notable enhancements in object localization and object detection. |
Text generation, Reasoning | 2026-04-14 | Global |
| GLM-5.1 is a model developed by Zhipu AI, specifically designed for long-horizon tasks. It has 744 billion parameters, supports an ultra-long context of 200k tokens, and can generate up to 128k tokens in a single response. GLM-5.1 excels in logical reasoning, long-text understanding, and code generation, while balancing performance with inference efficiency. It delivers outstanding results across multiple multi-task benchmarks and is well-suited for applications such as intelligent human-computer interaction, enterprise solutions, and developer assistance. |
Video generation | 2026-04-03 | Global |
| Wan2.7 video edit, supports both localized and global editing with prompt. Seamlessly replace elements using image references and replicate complex dynamic processes, including motion, special effects, and camera movements. |
Video generation | 2026-04-03 | Global |
| Wan2.7 text to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
Video generation | 2026-04-03 | Global |
| Wan2.7 reference to video, enhanced consistency & performance. Delivering superior stability for characters, props, and scenes. Supports hybrid referencing of up to 5 mixed image/video inputs and audio timbre cloning. Together with core engine upgrades, it achieves unprecedented cinematic expressive power. |
Video generation | 2026-04-03 | Global |
| Wan2.7 image to video, performance fully reimagined. Delivering nuanced and organic emotional depth in narrative arcs and visceral, bone-crunching impact in action sequences. Enhanced by rhythmic cinematic cuts for unparalleled storytelling power. |
Image generation | 2026-04-01 | Global |
| Wan2.7 – image generation and editing, supports text to image, text/image to sequential images, image editing, multi-image reference generation, and interactive editing. Delivers enhanced performance in text rendering, subject consistency, and complex instruction following. |
Text generation, Reasoning, Visual understanding | 2026-04-01 | Global |
| The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization. |