All Products
Search
Document Center

Alibaba Cloud Model Studio:tongyi-embedding-vision-flash

Last Updated:Sep 28, 2026

Embedding-Vision is a vision-centric multimodal embedding model powered by an LLM, featuring outstanding domain-specific performance and high cost-effectiveness in various domains (e.g., e-commerce, photo galleries, security, autonomous driving). With support for text, image, and video, it is applicable to downstream retrieval tasks, including text-to-image, image-to-image, text-to-video and video-to-video.

Inference Service Provider

The inference service provider for tongyi-embedding-vision-flash is Alibaba Cloud Model Studio.

Model Capabilities

CapabilitySupportCapabilitySupport

Input Modality

Text Image Video

Output Modality

—

Model Experience

Unsupported

Function Calling

Unsupported

Structured Outputs

Unsupported

Web Search

Unsupported

Prefix Completion

Unsupported

Context Caching

Unsupported

Batch Inference

Unsupported

Fine-tuning

Unsupported

Context Limits

ParameterValueParameterValue

Max Input Length

—

Max Output Length

—

Context Window

—

Pricing

This page only shows the original pricing for model API calls, excluding any limited-time promotions. Visit Model Studio Console for promotional offers.

Singapore

Scope: International

Billing ItemPrice (USD)Unit

Image Input

0.03

Per 1M tokens

Text Input

0.09

Per 1M tokens

Rate Limits

Singapore

Scope: International

ParameterValue

RPM (Requests Per Minute)

600

TPM (Tokens Per Minute)

200,000

Snapshot Versions