The Qwen3.8 27B native vision-language dense model builds upon the 3.6-27B version, with key improvements in coding and office productivity capabilities across both text and visual modalities. It enables more reliable end-to-end completion of complex tasks, delivering consistently trustworthy results.
Inference Service Provider
The inference service provider for qwen3.8-27b is Alibaba Cloud Model Studio.
Model Capabilities
Capability | Support | Capability | Support |
Input Modality | Image Text Video | Output Modality | Text |
Model Experience | Function Calling | ||
Structured Outputs | Web Search | ||
Prefix Completion | Context Caching | ||
Batch Inference | Fine-tuning |
Context Limits
Parameter | Value | Parameter | Value |
Max Input Length | 991808 | Max Output Length | 131072 |
Max Input Length (Thinking Mode) | 983616 | Max Output Length (Thinking Mode) | 131072 |
Context Window | 1000000 | Max Chain-of-Thought Length | 262144 |
The supported model length may vary depending on different combinations of API input parameters.
Pricing
This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.
China (Beijing)
Billing Item | Price (USD) | Unit |
Input | 0.424 | Per 1M tokens |
Output | 1.696 | Per 1M tokens |
Input(Implicit Cache) | 0.085 | Per 1M tokens |
Explicit Cache Creation | 0.53 | Per 1M tokens |
Explicit Cache Read | 0.042 | Per 1M tokens |
Singapore
Scope: International
Billing Item | Price (USD) | Unit |
Input | 0.5 | Per 1M tokens |
Output | 3 | Per 1M tokens |
Input(Implicit Cache) | 0.1 | Per 1M tokens |
Explicit Cache Creation | 0.625 | Per 1M tokens |
Explicit Cache Read | 0.05 | Per 1M tokens |
Rate Limits
China (Beijing)
Parameter | Value |
RPM (Requests Per Minute) | 5,000 |
TPM (Tokens Per Minute) | 5,000,000 |
Singapore
Scope: International
Parameter | Value |
RPM (Requests Per Minute) | 5,000 |
TPM (Tokens Per Minute) | 5,000,000 |