Lite
Getting Started with AI
- checkApprox. 10,000 credits/month
- checkRun 1–2 agents concurrently
- checkApply new user discount coupon to get extra $2 off.
Enjoy all-in-one access to Qwen and industry leading third party models, latest tools and expert support when you need it. Try the newest Qwen3.8-Max, HappyHorse1.1, DeepSeek-V4-Pro-0813 and GLM-5.2 models with the latest AI Token Plan.
Qwen
Wan
HappyHorse
DeepSeekGetting Started with AI
For daily use
Serious AI usage
Multi-Seat Access, Enterprise-Grade Security, Effortless Management
How to choose between Team and Individual version help_outlinePerfect for casual AI exploration
Best value for high-volume AI workflows
Built for maximum AI performance
Compare Individual vs Team Plans to find the best fit.
| Personal (Lite/Standard/Pro) | Team (Standard/Pro/Max Seat) | |
|---|---|---|
| Pricing | Prepay more to unlock bigger savings | Fixed credits per seat |
| Starting at | US$6/mo | US$20/seat/mo |
| Shared quota pack | — | Available |
| Text and LLMs | Included | Included |
| Image generation | Included | Included |
| Video generation | Included | Included |
| Speech | Included | Included |
| Billing | Fixed — same every month | Fixed — same every month |
| Enterprise features | — | ✓ Enterprise-grade data security. ✓ Enterprise management capabilities. ✓ Unified support for managing agents and usage. ✓ No model training of your content. |
| Best for | Individual users looking for good value for their AI creations | Team rollouts, budgeted projects |
Your fastest path to production AI — combining state-of-the-art models with
seamless enterprise-grade tools to build, deploy, and scale with confidence.
Native Support
Compatible with key third-party platforms
QWEN models are best in class AI models developed by Alibaba Cloud. Outperforms on intelligence, speed and affordability.
Try with 70M+ Free AI TokensHighly realistic text-to-video generation with fluid motion, natural detail, and accurate text comprehension.
Highly realistic image-to-video generation that accurately interprets both text and image semantics.
Any input, complete stories — up to 30 seconds of video in a single generation.
Next-gen text-to-video with cinematic emotional depth, rhythmic cuts, and action-sequence impact.
Next-gen image-to-video with cinematic emotional depth, rhythmic cuts, and action-sequence impact.
Prompt-driven video editing supporting localized/global edits, element replacement, and complex motion replication.
Full-featured image generation and editing with text rendering, subject consistency, and multi-image reference support.
Versatile image generation and editing supporting text-to-image, sequential images, and interactive editing.
Flagship image generation & editing: 4.5k-token rich layouts, text as small as 10px, and native rendering in 12 languages.
Image generation & editing with rich content, authentic details, and deep world knowledge.
World #1 open-source text-to-image model — 6B parameters, excels in bilingual prompts, complex scenes, and thematic generation.
Foundation image generation model with advanced complex text rendering and precise image editing capabilities.
Cost-effective vision-language model with full-stack agent intelligence for coding, tool use, and GUI interaction.
State-of-the-art vision-language model enhanced in agentic coding, front-end programming, OCR, and object localization.
Native vision-language Flash model excelling in agentic coding, math reasoning, and spatial intelligence.
Compact visual understanding model with thinking/non-thinking modes, ultra-long context, and 2D/3D localization.
World-leading visual agent model with upgrades in visual coding, spatial perception, and ultra-long video understanding.
2.4-trillion-parameter MoE flagship delivering a comprehensive leap in coding and professional work. Autonomously codes and delivers complete projects spanning 10+ days.
Largest and most capable Qwen3.7 model, excelling at agent programming, productivity tasks, and long-horizon autonomous execution.
Cost-effective vision-language model with full-stack agent intelligence for coding, tool use, and GUI interaction.
State-of-the-art vision-language model enhanced in agentic coding, front-end programming, OCR, and object localization.
Native vision-language Flash model excelling in agentic coding, math reasoning, and spatial intelligence.
Hybrid-architecture Plus model delivering state-of-the-art performance across pure-text and multimodal tasks.
Latest-gen multimodal model supporting text, image, audio, and video — 10h+ audio, 60+ language input, 30+ language speech output.
Cost-effective multimodal model supporting text, image, audio, and video with 60+ language input and 30+ language speech output.
MoE-based multimodal model for text, image, audio, and video — 119 text languages, 20 speech languages, human-like voice generation.
Realtime multilingual audio/video interpretation — understands 60 languages, speaks 29 languages.
Realtime multilingual audio/video interpretation — understands 19 languages, speaks 10 languages plus 8 Chinese dialects.
High-precision multilingual audio/video translation supporting 19 languages, 10 spoken languages, and 8 Chinese dialects.
LLM-based multilingual speech recognition — auto language detection across 11 languages with high accuracy in complex audio.
Realtime version of Qwen3-ASR-Flash — auto language detection across 11 languages with precise transcription in complex audio.
Speech recognition supporting 7 Chinese dialect systems, 20+ regional accents, and classical poetry recognition.