All Products
Search
Document Center

Vector Retrieval Service for Milvus:AI Function

Last Updated:Aug 03, 2026

Milvus AI Function wraps model calls as Collection Function and REST interfaces, enabling you to process text, images, videos, and audio during data ingestion, retrieval, or business workflows without maintaining separate model-call pipelines. AI Function covers common AI scenarios including embedding, reranking, content understanding and generation, media production, batch processing, and safety detection.

Common scenarios

  • Multi-modal retrieval and RAG — Generate embeddings for text, images, and videos, then rerank recalled candidates.

  • Content understanding and processing — Generate titles, translations, summaries, classifications, sentiment analysis results, keywords, entities, and structured tags.

  • Media production — Transcribe customer service recordings, edit product images, or create and query video generation tasks.

  • Batch and safety governance — Initiate asynchronous batch inference on historical assets and enable safety detection at the input or output stage of model calls.

Architecture

Input data and model services feed into the Milvus AI Function capability layer. Based on the function type, the capability layer returns embeddings, structured information, generated text, or media URLs for use by search, content production, and business systems.

Supported AI Functions

The following table lists all AI Functions. The model column lists model types and representative models. Available models are subject to server-side Provider configuration.

AI FunctionDescriptionInput modalitySupported model types/representative models
AI_EMBEDDINGConverts content into dense vectors for semantic search, RAG, clustering, recommendations, and image-text retrieval.Text, image, videoText embedding:text-embedding-v4; Multi-modal embedding:qwen3-vl-embedding,tongyi-embedding-vision-plus.
AI_EMBEDDING_CACHEReuses vectors from Redis exact-match cache for repeated text, reducing embedding latency and model usage.TextText embedding:text-embedding-v4,text-embedding-v3.
AI_RERANKScores and reranks a query against recalled candidate text, images, or videos.Text, image, videoText reranking:qwen3-rerank; Multi-modal reranking:qwen3-vl-rerank.
AI_TEXT_TRANSFORMGenerates titles, copy, tags, or field completion results line by line based on a prompt.TextText models:qwen3.7-max,deepseek-v4-pro,deepseek-v4-flash.
AI_TRANSLATETranslates text into the specified target language.TextText models:qwen3.7-max,deepseek-v4-pro,deepseek-v4-flash.
AI_SIMILARITYCompares the semantic similarity of text pairs and returns a score from 0 to 1.TextText models:qwen3.7-max,deepseek-v4-pro,deepseek-v4-flash.
AI_PII_MASKIdentifies and masks personal sensitive information in text.TextText models:qwen3.7-max,deepseek-v4-pro,glm-5.2.
AI_SUMMARIZATIONCompresses text, images, or videos into concise summaries.Text, image, videoText models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6.
AI_EXTRACTIONExtracts structured fields from content based on given labels and returns JSON.Text, image, videoText models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6.
AI_ENTITYExtracts named entities such as names, organizations, locations, times, and amounts, and returns JSON.Text, image, videoText models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6.
AI_KEYWORDExtracts representative keywords from content and returns JSON.Text, image, videoText models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6.
AI_CLASSIFICATIONSelects the category from candidate labels that best matches the content.Text, image, videoText models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6.
AI_SENTIMENTSelects the sentiment category from candidates that best matches the content.Text, image, videoText models:qwen3.7-max; Multi-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6.
AI_MULTIMODAL_GENERATIONGenerates text results from text, images, videos, or audio.Text, image, video, audioMulti-modal models:qwen3.7-plus,qwen3.6-plus,kimi-k2.6; Speech recognition:qwen3-asr-flash.
AI_TRANSCRIPTIONTranscribes an audio URL into searchable text.AudioSpeech recognition model:qwen3-asr-flash.
AI_IMAGE_EDITGenerates an edited image URL based on an image and a prompt.Text, imageImage generation models:wan2.7-image-pro,wan2.7-image.
AI_VIDEO_EDITSubmits text-to-video, image-to-video, or video editing tasks and asynchronously retrieves the result video URL.Text, image, videoVideo models:happyhorse-1.1-t2v,happyhorse-1.1-i2v,happyhorse-1.0-video-edit.
AI_BATCHSubmits Chat Completions requests in JSONL as asynchronous batch tasks, polls for status, and downloads output or error files.Text, image, video, audioText models:qwen3.7-max; Vision models:qwen3-vl-plus; Omni-modal models:qwen3.5-omni-plus.
AI_DATA_INSPECTIONChecks for unsafe content at the input stage, output stage, or both stages of a model call.TextUses the text model of the current TextTransform call, such asqwen3.7-max,deepseek-v4-pro.

Usage

  • Collection Function — Configure functions in the Collection Schema for automatic execution during write, retrieval, or search stages.

  • REST interface — Call functions on demand from your business services. Each function page provides cURL and Python examples.

  • Asynchronous tasks —AI_BATCH andAI_VIDEO_EDIT return task identifiers. Poll until the task reaches a terminal state before reading output or downloading results.

Selection guide

  • Retrieval — UseAI_EMBEDDING first, then useAI_RERANK as needed. EnableAI_EMBEDDING_CACHE for repeated text.

  • Content understanding — Select the summary, extraction, entity, keyword, classification, sentiment analysis, or multi-modal generation function based on your goal.

  • Media generation — UseAI_IMAGE_EDIT for images andAI_VIDEO_EDIT for videos. UseAI_BATCH for large-scale offline processing.

  • Sensitive content handling — Enabledata_inspection in the TextTransform request or Function parameters.