All Products
Search
Document Center

Alibaba Cloud Model Studio:qwen3.8-2.4t-a95b

Last Updated:Aug 28, 2026

Qwen3.8-2.4T-A95B is the open-source release of Qwen's latest flagship, launched August 2026. Its sparse MoE architecture holds 2.4T total parameters with ~95B activated per step, paired with hybrid attention and a 1M token context window. Key benchmarks: GPQA Diamond 92.6, PaperBench 93.0, OSWorld 86.1, BabyVision 82.0. Ranked 4th on CodeArena.

Inference Service Provider

The inference service provider for qwen3.8-2.4t-a95b is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Supported

Context Caching

Supported

Batch Inference

Supported

Fine-tuning

Unsupported

Context Limits

Parameter

Value

Parameter

Value

Max Input Length

991808

Max Output Length

131072

Context Length

1000000

Max Input Length (Thinking Mode)

983616

Max Output Length (Thinking Mode)

131072

Max Chain-of-Thought Length

131072

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Billing ItemPrice (USD)Unit

Input

1.65

Per 1M tokens

Output

4.951

Per 1M tokens

Input (Cache Hit)

0.206

Per 1M tokens

Explicit Cache Creation

2.063

Per 1M tokens

Explicit Cache Hit

0.137

Per 1M tokens

Singapore

Scope: International

Billing ItemPrice (USD)Unit

Input

2

Per 1M tokens

Output

6

Per 1M tokens

Input (Cache Hit)

0.25

Per 1M tokens

Explicit Cache Creation

2.5

Per 1M tokens

Explicit Cache Hit

0.17

Per 1M tokens

Rate Limits

China (Beijing)

ParameterValue

RPM (Requests Per Minute)

5000

TPM (Tokens Per Minute)

5,000,000

Singapore

Scope: International

ParameterValue

RPM (Requests Per Minute)

5000

TPM (Tokens Per Minute)

5,000,000