All Products
Search
Document Center

Alibaba Cloud Model Studio:deepseek-v4.1-flash

Last Updated:Sep 14, 2026

DeepSeek-V4.1-Flash is the lightweight flagship of DeepSeek's new architecture family, packing 552B total MoE parameters to deliver flagship-surpassing intelligence across key benchmarks, including out-performing DeepSeek-V4-Pro. It adopts a Causal Encoder-Decoder asymmetric design with 8B active parameters for input and 16B for output, and offers native multimodal visual understanding. KV Cache usage is cut to one-quarter of the previous generation's HBM and one-eighth of its SSD storage, dramatically lowering costs for long-context and agentic workloads. With a 1M-token context window and up to 384K-token maximum output, it balances high throughput, low latency, and exceptional cost efficiency.

Inference Service Provider

The inference service provider for deepseek-v4.1-flash is Alibaba Cloud Model Studio.

Model Capabilities

China (Beijing)

CapabilitySupportCapabilitySupport

Input Modality

Text, Image

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Unsupported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Singapore

Scope: International

CapabilitySupportCapabilitySupport

Input Modality

Text, Image

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Supported

Prefix Completion

Unsupported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Germany (Frankfurt)

Scope: Global

CapabilitySupportCapabilitySupport

Input Modality

Text, Image

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

US (Virginia)

Scope: Global

CapabilitySupportCapabilitySupport

Input Modality

Text, Image

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Japan (Tokyo)

Scope: Global

CapabilitySupportCapabilitySupport

Input Modality

Text, Image

Output Modality

Text

Model Experience

Supported

Function Calling

Supported

Structured Outputs

Supported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Supported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

1000000

Max Output Length

393216

Context Window

1000000

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

China (Beijing)

Billing ItemPrice (USD)Unit

Input (Idle hours)

0.141

Per 1M tokens

Input (Busy hours)

0.283

Per 1M tokens

Output (Idle hours)

0.565

Per 1M tokens

Output (Busy hours)

1.131

Per 1M tokens

Input(Implicit Cache, Idle hours)

0.014

Per 1M tokens

Input(Implicit Cache, Busy hours)

0.028

Per 1M tokens

Singapore

Scope: International

Billing ItemPrice (USD)Unit

Input (Idle hours)

0.15

Per 1M tokens

Input (Busy hours)

0.3

Per 1M tokens

Output (Idle hours)

0.6

Per 1M tokens

Output (Busy hours)

1.2

Per 1M tokens

Input(Implicit Cache, Idle hours)

0.015

Per 1M tokens

Input(Implicit Cache, Busy hours)

0.03

Per 1M tokens

Germany (Frankfurt)

Scope: Global

Billing ItemPrice (USD)Unit

Input (Idle hours)

0.141

Per 1M tokens

Input (Busy hours)

0.283

Per 1M tokens

Output (Idle hours)

0.565

Per 1M tokens

Output (Busy hours)

1.131

Per 1M tokens

Input(Implicit Cache, Idle hours)

0.014

Per 1M tokens

Input(Implicit Cache, Busy hours)

0.028

Per 1M tokens

US (Virginia)

Scope: Global

Billing ItemPrice (USD)Unit

Input (Idle hours)

0.141

Per 1M tokens

Input (Busy hours)

0.283

Per 1M tokens

Output (Idle hours)

0.565

Per 1M tokens

Output (Busy hours)

1.131

Per 1M tokens

Input(Implicit Cache, Idle hours)

0.014

Per 1M tokens

Input(Implicit Cache, Busy hours)

0.028

Per 1M tokens

Japan (Tokyo)

Scope: Global

Billing ItemPrice (USD)Unit

Input (Idle hours)

0.141

Per 1M tokens

Input (Busy hours)

0.283

Per 1M tokens

Output (Idle hours)

0.565

Per 1M tokens

Output (Busy hours)

1.131

Per 1M tokens

Input(Implicit Cache, Idle hours)

0.014

Per 1M tokens

Input(Implicit Cache, Busy hours)

0.028

Per 1M tokens

Rate Limits

China (Beijing)

ParameterValue

RPM (Requests Per Minute)

15000

TPM (Tokens Per Minute)

1,200,000

Singapore

Scope: International

ParameterValue

RPM (Requests Per Minute)

10000

TPM (Tokens Per Minute)

1,200,000

Germany (Frankfurt)

Scope: Global

ParameterValue

RPM (Requests Per Minute)

15000

TPM (Tokens Per Minute)

1,200,000

US (Virginia)

Scope: Global

ParameterValue

RPM (Requests Per Minute)

15000

TPM (Tokens Per Minute)

1,200,000

Japan (Tokyo)

Scope: Global

ParameterValue

RPM (Requests Per Minute)

15000

TPM (Tokens Per Minute)

1,200,000

Hong Kong (China)

Scope: Global

ParameterValue

RPM (Requests Per Minute)

15000

TPM (Tokens Per Minute)

1,200,000