Milvus AI Function wraps model calls as Collection Function and REST interfaces, enabling you to process text, images, videos, and audio during data ingestion, retrieval, or business workflows without maintaining separate model-call pipelines. AI Function covers common AI scenarios including embedding, reranking, content understanding and generation, media production, batch processing, and safety detection.
Common scenarios
Multi-modal retrieval and RAG — Generate embeddings for text, images, and videos, then rerank recalled candidates.
Content understanding and processing — Generate titles, translations, summaries, classifications, sentiment analysis results, keywords, entities, and structured tags.
Media production — Transcribe customer service recordings, edit product images, or create and query video generation tasks.
Batch and safety governance — Initiate asynchronous batch inference on historical assets and enable safety detection at the input or output stage of model calls.
Architecture
Input data and model services feed into the Milvus AI Function capability layer. Based on the function type, the capability layer returns embeddings, structured information, generated text, or media URLs for use by search, content production, and business systems.
Supported AI Functions
The following table lists all AI Functions. The model column lists model types and representative models. Available models are subject to server-side Provider configuration.
| AI Function | Description | Input modality | Supported model types/representative models |
| AI_EMBEDDING | Converts content into dense vectors for semantic search, RAG, clustering, recommendations, and image-text retrieval. | Text, image, video | Text embedding:text-embedding-v4; Multi-modal embedding:qwen3-vl-embedding,tongyi-embedding-vision-plus. |
| AI_EMBEDDING_CACHE | Reuses vectors from Redis exact-match cache for repeated text, reducing embedding latency and model usage. | Text | Text embedding:text-embedding-v4,text-embedding-v3. |
| AI_RERANK | Scores and reranks a query against recalled candidate text, images, or videos. | Text, image, video | Text reranking:qwen3-rerank; Multi-modal reranking:qwen3-vl-rerank. |
| AI_TEXT_TRANSFORM | Generates titles, copy, tags, or field completion results line by line based on a prompt. | Text | Text models:qwen3.7-max,deepseek-v4-pro,deepseek-v4-flash. |
| AI_TRANSLATE | Translates text into the specified target language. | Text | Text models:qwen3.7-max,deepseek-v4-pro,deepseek-v4-flash. |
| AI_SIMILARITY | Compares the semantic similarity of text pairs and returns a score from 0 to 1. | Text | Text models:qwen3.7-max,deepseek-v4-pro,deepseek-v4-flash. |
| AI_PII_MASK | Identifies and masks personal sensitive information in text. | Text | Text models:qwen3.7-max,deepseek-v4-pro,glm-5.2. |
| AI_SUMMARIZATION | Compresses text, images, or videos into concise summaries. | Text, image, video | Text models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6. |
| AI_EXTRACTION | Extracts structured fields from content based on given labels and returns JSON. | Text, image, video | Text models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6. |
| AI_ENTITY | Extracts named entities such as names, organizations, locations, times, and amounts, and returns JSON. | Text, image, video | Text models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6. |
| AI_KEYWORD | Extracts representative keywords from content and returns JSON. | Text, image, video | Text models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6. |
| AI_CLASSIFICATION | Selects the category from candidate labels that best matches the content. | Text, image, video | Text models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6. |
| AI_SENTIMENT | Selects the sentiment category from candidates that best matches the content. | Text, image, video | Text models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6. |
| AI_MULTIMODAL_GENERATION | Generates text results from text, images, videos, or audio. | Text, image, video, audio | Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6; Speech recognition:qwen3-asr-flash. |
| AI_TRANSCRIPTION | Transcribes an audio URL into searchable text. | Audio | Speech recognition model:qwen3-asr-flash. |
| AI_IMAGE_EDIT | Generates an edited image URL based on an image and a prompt. | Text, image | Image generation models:wan2.7-image-pro,wan2.7-image. |
| AI_VIDEO_EDIT | Submits text-to-video, image-to-video, or video editing tasks and asynchronously retrieves the result video URL. | Text, image, video | Video models:happyhorse-1.1-t2v,happyhorse-1.1-i2v,happyhorse-1.0-video-edit. |
| AI_BATCH | Submits Chat Completions requests in JSONL as asynchronous batch tasks, polls for status, and downloads output or error files. | Text, image, video, audio | Text models:qwen3.7-max; Vision models:qwen3-vl-plus; Omni-modal models:qwen3.5-omni-plus. |
| AI_DATA_INSPECTION | Checks for unsafe content at the input stage, output stage, or both stages of a model call. | Text | Uses the text model of the current TextTransform call, such asqwen3.7-max,deepseek-v4-pro. |
Usage
Collection Function — Configure functions in the Collection Schema for automatic execution during write, retrieval, or search stages.
REST interface — Call functions on demand from your business services. Each function page provides cURL and Python examples.
Asynchronous tasks —
AI_BATCHandAI_VIDEO_EDITreturn task identifiers. Poll until the task reaches a terminal state before reading output or downloading results.
Selection guide
Retrieval — Use
AI_EMBEDDINGfirst, then useAI_RERANKas needed. EnableAI_EMBEDDING_CACHEfor repeated text.
Content understanding — Select the summary, extraction, entity, keyword, classification, sentiment analysis, or multi-modal generation function based on your goal.
Media generation — Use
AI_IMAGE_EDITfor images andAI_VIDEO_EDITfor videos. UseAI_BATCHfor large-scale offline processing.
Sensitive content handling — Enable
data_inspectionin the TextTransform request or Function parameters.