In this tutorial, you chain the AI functions of Alibaba Cloud Milvus into one short drama production pipeline that generates keyframe images, turns them into motion shots, and ingests both into a single multimodal collection. By the end, you can retrieve any generated asset by text or by image from that collection.
Solution overview
AIGC short dramas, short-form vertical dramas, and brand short videos are produced at a fast pace. A team produces dozens or even hundreds of finished videos per week, and each finished video requires a large number of keyframes, storyboards, and candidate assets to be generated and screened. Without semantic search over an accumulated asset library, previous assets can be located only by file name and memory.
Alibaba Cloud Milvus chains this pipeline with four AI functions, in two phases — generation and retrieval — and five steps. The following table maps the four stages of a typical production pipeline to those steps.
Production stage | Phase | Step | What happens |
Storyboard script | — | Manual input | Write the storyboard scripts first, for example "On a rainy night, the female lead stands under an umbrella at the entrance of a convenience store, with warm light falling on her face". In this tutorial, the script is passed as |
Keyframe image generation | Generation | Step 1 | Convert a storyboard script into a keyframe image that anchors the shot visually. You can also replace the background or adjust the style of an existing asset. |
Image-to-video generation | Generation | Step 2 | Use a selected keyframe as the first frame to generate a motion shot. |
Asset accumulation and retrieval reuse | Retrieval | Steps 3 to 5 | Automatically generate captions for the images and videos that you produce each day, convert them into vectors, and ingest them. You can then run text-to-image search, image-to-video search, and text-to-video search. |
The four AI functions used across these steps are as follows:
AI_IMAGE_EDIT(Step 1) — Image generation and editing. Supports single-image editing and fusion of multiple reference images. Converts a storyboard reference image into a keyframe image based on the script.AI_VIDEO_EDIT(Step 2) — Image-to-video and text-to-video generation. Runs as an asynchronous task. Generates a motion shot with a keyframe as the first frame.AI_MULTI_MODAL_GENERATE(Step 3) — Multimodal content understanding. Generates titles, captions, and tags for images and videos. Generates the objective captions that are used for retrieval.AI_EMBEDDING(Steps 4 and 5) — Vectorization on write. Maps text, images, and videos into the same vector space, which is what supports cross-modal search.Three mechanisms carry this pipeline:
Ingest on generation — Keyframes and motion shots are ingested together with their captions and tags, so every generated item carries structured information and a vector.
One vector space — Images and videos share the same multimodal model and the same vector field, which is what makes cross-modal search such as text-to-image search and image-to-video search possible.
Vectorization on write and inference on query — Your application does not need to call the embedding model itself.
Prerequisites
A Milvus 2.6 instance is created. AI functions require the 2.6 kernel. You do not need to bind a model service separately after you create the instance.
pymilvus is installed. The examples in this tutorial are verified with pymilvus 3.0.0.
To access the instance over the Internet, enable Public Network Access on the Security Configuration tab of the instance details page, and add the egress IP address of your client to the public network access whitelist.
(Recommended for production pipelines) An Object Storage Service (OSS) bucket of your own, used to persist generated images and videos before ingestion.
Considerations
Before you run the pipeline, consider the following:
Endpoint port — The RESTful API and gRPC share port 19530. You must explicitly specify the port in the endpoint, for example
http://c-xxx.milvus.aliyuncs.com:19530. If you omit the port, the request goes to port 80 by default, and the connection times out.Lifetime of generated URLs — The URLs returned by image generation and video generation are OSS-signed and, measured in a test environment, remain valid for about 24 hours. Persist generated assets to your own OSS before ingestion and write the long-lived URL to
media_ref.Query-side media URLs — Make sure that the media URL passed in on the query side is accessible.
One vector field, one model — Images and videos must land in the same vector field and use the same model so that you can run cross-modal search in the same vector space.
Multimodal function parameters — A multimodal function must include
"is_multimodal": "true", which is a string value. Thedimof the vector field must match thedimin the function parameters.Flush before search — You must call
flush()after you write data. Otherwise, a search that runs immediately afterward may return an empty result.Measured values — The durations, similarity scores, and captions shown in this tutorial were measured in a test environment. Treat them as reference values rather than product specifications.
Prepare the shared code
This tutorial uses two access methods. Steps 1 to 3 call the generation and understanding functions over the RESTful API at /v2/vectordb/ai/* through the post_json utility. Steps 4 and 5 create the collection, ingest data, and search with the pymilvus MilvusClient.
The Python code blocks in this tutorial form one script that you run in order in a single Python session. Later steps consume variables produced by earlier steps: keyframe_url, video_url, image_caption, video_caption, client, and collection_name. To run a single step in isolation, supply the values of these upstream variables yourself.
The following code contains the connection settings and the REST call utility. Replace MILVUS_ENDPOINT and MILVUS_TOKEN with the information of your own instance.
from __future__ import annotations
import json
import time
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
from pymilvus import DataType, Function, FunctionType, MilvusClient
# ==== Global settings (REST and gRPC share port 19530) ====
MILVUS_ENDPOINT = "c-xxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "root:xxx"
MILVUS_REST_BASE_URL = f"http://{MILVUS_ENDPOINT}"
MILVUS_URI = MILVUS_REST_BASE_URL
def post_json(path: str, body: dict[str, Any], timeout: int = 200) -> tuple[int, dict[str, Any]]:
"""REST call utility shared by Steps 1 to 3."""
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
try:
with urlopen(request, timeout=timeout) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))The scope of this block differs by element. The connection settings MILVUS_ENDPOINT, MILVUS_TOKEN, and MILVUS_URI are used by all five steps, including the MilvusClient in Step 4. The post_json function is used by Steps 1 to 3 only. Its timeout=200 sets the client-side HTTP timeout in seconds for each REST call.
Step 1: Generate the keyframe image
Take a storyboard reference image, replace the background or adjust the style based on the storyboard script, and produce a keyframe image that serves as the first frame for image-to-video generation. The storyboard script is your own manual input: in the following code, it is the value of params.prompt.
Item | Content |
Function / model |
|
Input | Reference image URL + storyboard script |
Output | Keyframe image URL |
Call mode | Synchronous response |
# ========== Step 1: Keyframe image generation AI_IMAGE_EDIT (wan2.7-image) ==========
IMAGE_URL = "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"
PROMPT = (
"On a rainy night, a woman stands under an umbrella at the entrance of a convenience store, "
"with warm light from inside the store falling on her face, cinematic composition, keep the main subject."
)
status, data = post_json(
"/v2/vectordb/ai/image_edit",
{
"model_name": "wan2.7-image",
"texts": [IMAGE_URL],
"params": {"prompt": PROMPT, "n": 1, "size": "1024*1024",
"watermark": False, "timeout_sec": 180},
},
)
assert status == 200 and data.get("code") == 0, data
keyframe_url = data["data"]["output"]["outputs"][0]
print(f"Keyframe image URL: {keyframe_url}")For single-image input, pass the reference image URL in texts and write the editing instruction in params.prompt. n specifies the number of images to generate, which is 1 in this example.
The prompt in this example ends with keep the main subject, so the subjects of the reference image are preserved in the keyframe. The reference image contains a woman and a dog, which is why the captions generated in Step 3 mention a dog that the prompt itself does not describe.
Step 2: Generate the motion shot from the keyframe
Use the keyframe as the first frame to generate a motion shot. Video generation is a long-running asynchronous task: create the task to get a task_id, then poll for the task status.
AI_VIDEO_EDIT supports both image-to-video and text-to-video generation. This tutorial covers image-to-video generation only, which is why the keyframe is passed in media with type set to first_frame.
Item | Content |
Function / model |
|
Input | First frame image + camera movement prompt |
Output |
|
Call mode | Asynchronous: create the task with |
# ========== Step 2: Image-to-video generation AI_VIDEO_EDIT (happyhorse-1.1-i2v), asynchronous task ==========
VIDEO_MODEL = "happyhorse-1.1-i2v" # The creation and query requests must use the same model name
# Use the keyframe from Step 1 as the first frame (media.type=first_frame)
status, data = post_json(
"/v2/vectordb/ai/video_edit",
{
"model_name": VIDEO_MODEL,
"prompt": "The shot slowly pushes in, rain falls, warm light flickers gently on the umbrella, keep the main subject from the first frame.",
"media": [{"type": "first_frame", "url": keyframe_url}],
"params": {"resolution": "720P", "audio_setting": "none",
"watermark": False, "timeout_sec": 180},
},
)
assert status == 200 and data.get("code") == 0, data
task_id = data["data"]["output"]["task_id"]
print(f"Task created task_id = {task_id}")
# Poll for the result: tasks/describe queries one task_id at a time, with provider and the same model_name
video_url = None
for attempt in range(60): # Local polling limit. When it is reached, only record task_id; do not treat it as a server-side failure
status, data = post_json(
"/v2/vectordb/ai/tasks/describe",
{"provider": "aliyun_milvus", "model_name": VIDEO_MODEL, "task_id": task_id},
)
task_status = data.get("data", {}).get("output", {}).get("task_status")
print(f" Poll #{attempt:<2} -> {task_status}")
if task_status == "SUCCEEDED":
video_url = data["data"]["output"]["video_url"]
break
if task_status in ("FAILED", "CANCELED"):
raise RuntimeError(f"task ended in {task_status}: {json.dumps(data, ensure_ascii=False)}")
time.sleep(10)
assert video_url, f"Local polling limit reached. Query again later with task_id={task_id}"
print(f"Motion shot video_url: {video_url}")The request to tasks/describe must carry the same model_name and provider as the creation request, and it queries one task_id at a time. Its response body contains three fields: task_id, task_status, and video_url. The following table describes the values of task_status that polling returns.
| Description | Handling |
| The task is in progress and not complete. This is a normal return value while the task runs. | Continue polling. Do not treat it as an error. |
| The task succeeded. You can get the result from | Stop polling. |
| The task failed. | Terminal state. Stop polling and troubleshoot. |
| The task was canceled. | Terminal state. Stop polling. |
One 720P, 5-second image-to-video task takes about 2 to 3 minutes in a test environment, and task_status keeps returning UNKNOWN during that time. The polling loop in this example allows up to 60 attempts at 10-second intervals, which is a local limit of about 10 minutes and not a server-side timeout.
If the loop reaches this local limit, the task is still running on the server side. Do not treat the interruption as a failure: keep the task_id printed by the creation call, and call /v2/vectordb/ai/tasks/describe again later with the same provider, the same model_name, and that task_id to retrieve video_url.
params.resolution specifies a pixel-count tier, not fixed dimensions. The actual output dimensions depend on the aspect ratio of the first frame. In this example, the first frame is a 1024×1024 square, so 720P produces an actual output of 960×960 (the total pixel count is equivalent to 1280×720), with a duration of about 5 seconds. For 16:9 landscape output, use a first frame with the corresponding aspect ratio.
Step 3: Generate captions and tags for the assets
Before you ingest assets, use a multimodal large model to generate one objective retrieval caption for each image and video. Use the caption as the source of the title, summary, or tags, and ingest it together with the asset in Step 4.
Item | Content |
Function / model |
|
Input | Image URL or video URL |
Output | A one-sentence caption that you can use as a title, description, or tag source |
# ========== Step 3: Asset understanding AI_MULTI_MODAL_GENERATE (qwen3.7-plus) ==========
def describe_media(url: str, media_type: str, prompt: str) -> str:
status, data = post_json(
"/v2/vectordb/ai/multi_modal_generate",
{
"model_name": "qwen3.7-plus",
"texts": [url],
"params": {"media_type": media_type, "prompt": prompt, "temperature": 0},
},
)
assert status == 200 and data.get("code") == 0, data
return data["data"]["output"]["outputs"][0]
image_caption = describe_media(
keyframe_url, "image",
"Describe the subject, scene, and color tone of this keyframe image in one objective sentence, for text-to-image search.",
)
video_caption = describe_media(
video_url, "video",
"Describe the subject, action, and visual mood of this video in one objective sentence, for shot-level retrieval.",
)
print(f"[Image caption] {image_caption}")
print(f"[Video caption] {video_caption}")Use media_type=image for images and media_type=video for videos. prompt is required. The following captions were generated in a test environment:
[Image caption] On a rainy street at night, a woman in a plaid shirt stands under an umbrella at the
entrance of a convenience store, with a golden retriever sitting beside her. The overall
tone is cool blue, and the warm light reflects off the wet ground, creating a quiet,
slightly melancholy mood.
[Video caption] The video shows a rainy street at night, where a woman in a plaid shirt holding a
transparent umbrella stands with a yellow dog outside a lit storefront. The shot moves
from a wide view to a close-up of the woman's face, and the overall tone is cool with
wet reflections.The quality of the caption directly determines the recall of later text-to-image search. In the prompt, explicitly ask for an "objective description of the subject, scene, color tone, and action" and other dimensions that you use at retrieval time, and avoid captions that contain subjective judgments.
Step 4: Ingest the assets into a multimodal collection
Create a multimodal asset collection with vectorization on write, and write the image URL and video URL together with the captions and tags. Milvus automatically generates vectors for the assets and ingests them.
Item | Content |
Function / model |
|
Key parameters |
|
Input | URLs, captions, and tags produced in Steps 1 to 3 |
Output | Ingested 2560-dimensional vectors |
The example drops any existing collection that has the same name before it creates the collection. Rerunning the script therefore recreates the collection from scratch, and assets ingested by an earlier run are removed.
drop_collection deletes the collection named aigc_assets_multimodal and all entities and vectors stored in it. On an instance that carries other workloads, use a dedicated collection name, or confirm that no collection with this name exists, before you run the script.
# ========== Step 4: Create a multimodal collection and ingest vectors AI_EMBEDDING (qwen3-vl-embedding) ==========
EMBED_MODEL = "qwen3-vl-embedding"
VECTOR_DIM = 2560 # Multimodal dimension of qwen3-vl-embedding. The dimension of the vector field must match dim
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
collection_name = "aigc_assets_multimodal"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("media_type", DataType.VARCHAR, max_length=16) # image / video
schema.add_field("media_ref", DataType.VARCHAR, max_length=4096) # Image or video URL used for vectorization
schema.add_field("caption", DataType.VARCHAR, max_length=2048) # Caption generated in Step 3
schema.add_field("tags", DataType.VARCHAR, max_length=512)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=VECTOR_DIM)
# When media_ref is written, qwen3-vl-embedding is called automatically to generate a 2560-dimensional vector
schema.add_function(
Function(
name="embed_media",
function_type=FunctionType.TEXTEMBEDDING,
input_field_names=["media_ref"],
output_field_names=["embedding"],
params={
"provider": "aliyun_milvus",
"model_name": EMBED_MODEL,
"dim": VECTOR_DIM,
"is_multimodal": "true",
},
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)
# Vectorize and ingest the output of Steps 1 to 3 (url + caption + tags)
client.insert(
collection_name,
[
{"media_type": "image", "media_ref": keyframe_url,
"caption": image_caption, "tags": "keyframe,generated image"},
{"media_type": "video", "media_ref": video_url,
"caption": video_caption, "tags": "motion shot,generated video"},
],
)
client.flush(collection_name)
client.load_collection(collection_name)
print(f"Images and videos are vectorized and ingested. Collection: {collection_name}")The function parameters set is_multimodal to "true" and dim to 2560, which matches the dimension of the embedding field. The image and the video are written to the same embedding field through the same model, and the script calls flush() before it loads the collection. For the constraints behind these three settings, see Considerations.
Effects of an expired generated URL — The URLs produced in Steps 1 and 2 are temporary signed URLs, as described in Considerations. If you write such a URL to media_ref, expiry has two effects.
Scenario | Effect |
Running a text search after the URL in | The asset is still recalled correctly, because the vector was generated and stored at write time. However, the returned URL is no longer accessible, so your application fails to display the asset. |
Using an expired URL as the query for image-to-image search | The entire |
Therefore, in a production environment, persist generated assets to your own OSS before ingestion and write the long-lived URL to media_ref.
Expected result: The script prints Images and videos are vectorized and ingested. Collection: aigc_assets_multimodal. At this point the collection is loaded and holds two entities, one image and one video. Do not continue to Step 5 until this output appears. If Step 5 later returns an empty result, verify that flush() ran before the search.
Step 5: Search the asset library
After the assets are ingested, use the same multimodal model to map the query into the same vector space, and retrieve the most similar assets for reuse. Because the collection is already bound to the qwen3-vl-embedding function, you can pass raw text or a media URL directly in the data parameter of search, and Milvus vectorizes the query automatically.
Item | Content |
Function / model |
|
Input | Query text or media URL passed in |
Output | Top-K assets with similarity scores, filtered by |
Call mode | pymilvus |
# ========== Step 5: Content search, text-to-image / text-to-video / image-to-video ==========
def search_assets(query: str, top_k: int = 5, media_type: str | None = None,
min_score: float = 0.0):
"""query can be a Chinese or English text description, or the URL of a query image.
min_score: the similarity threshold. Results below this value are considered irrelevant and discarded."""
filter_expr = f'media_type == "{media_type}"' if media_type else ""
results = client.search(
collection_name=collection_name,
data=[query], # Text or image URL, both vectorized by qwen3-vl-embedding
anns_field="embedding",
limit=top_k,
filter=filter_expr,
output_fields=["media_type", "media_ref", "caption", "tags"],
)
for rank, hit in enumerate(results[0], 1):
if hit["distance"] < min_score:
continue
e = hit["entity"]
url = e["media_ref"].split("?")[0] # Remove the OSS signature parameters for cleaner display
print(f" {rank}. [Similarity {hit['distance']:.4f}] [{e['media_type']}] {e['caption'][:42]}")
print(f" {url}")
# Text-to-image search: restrict the search to images
print("[5.1] Text-to-image search:")
search_assets("A woman under an umbrella at the entrance of a convenience store on a rainy night", media_type="image")
# Text-to-video search: restrict the search to videos
print("[5.2] Text-to-video search:")
search_assets("A motion shot that slowly pushes in on a rainy street at night", media_type="video")
# Image-to-video search: use a keyframe image URL to recall video assets with a similar style
print("[5.3] Image-to-video search (recall videos from the same source with a keyframe image):")
search_assets(keyframe_url, media_type="video")
# Cross-modal hybrid search: remove the filter to search a mixed collection of images and videos
print("[5.4] Cross-modal hybrid search:")
search_assets("A warm-light shot of a convenience store on a rainy night")The following table shows the recall results and similarity scores of the four search methods in a test environment.
Search method | Query | Hit | Similarity |
Text-to-image search | A woman under an umbrella at the entrance of a convenience store on a rainy night | image | 0.4834 |
Text-to-video search | A motion shot that slowly pushes in on a rainy street at night | video | 0.4967 |
Image-to-video search | Keyframe image URL | video | 0.8715 |
Cross-modal hybrid (no filter) | A warm-light shot of a convenience store on a rainy night | video / image | 0.4967 / 0.3816 |
Image-to-video search reaches 0.8715, the highest of the four methods, because that video was generated from this keyframe: assets from the same source have very high similarity in the same vector space. Because images and videos share one vector space, one search call covers all four methods: replace the query text with an image URL to recall assets from a media query, and remove the filter to search a mixed collection of images and videos.
Keep the query text relevant to the asset topic, and apply a similarity threshold. Multimodal search returns the Top-K results sorted by similarity, so it returns results even when the query is completely unrelated to the assets. The following comparison was measured in a test environment.
Query text | Relationship to the assets | Similarity |
A woman under an umbrella at the entrance of a convenience store on a rainy night | Topic match | 0.4834 |
A striped sweater look under warm indoor light | Unrelated | 0.0715 |
A concept film with a close-up of the character and a slow push-in | Unrelated | 0.0694 |
The gap is nearly sevenfold. Therefore, filter results on the application side with a similarity threshold (the min_score parameter of search_assets in this example) and discard clearly irrelevant long-tail results, so that scores such as 0.07 are not shown to users as valid recall. Tune the exact threshold based on the distribution of your business assets.
Two details of the example affect what you see and what you should reuse:
min_scoredefaults to0.0— None of the four example calls passesmin_score, so no result is discarded and low-similarity hits such as 0.07 appear in the output. Pass your own threshold when you integrate the search into an application. Filtering happens after Top-K selection, and the example skips filtered rows without renumbering, so the printed rank numbers can be non-contiguous.The printed URL is truncated — The example removes the OSS signature parameters from
media_refto keep the printed output readable. Use the fullmedia_refvalue in your application, not the truncated form.
Troubleshooting
Symptom | Cause | Solution |
The connection to the instance times out. | The endpoint does not specify a port, so the request goes to port 80 by default. | Specify port 19530 in the endpoint, for example |
| The multimodal function does not include | Add |
A search that runs immediately after a write returns an empty result. | The written data is not flushed. | Call |
| The media URL passed in as the query cannot be downloaded, for example an expired signed URL. | Pass an accessible URL as the query. Persist generated assets to your own OSS and write the long-lived URL to |
Polling ends without a terminal | The local polling limit is reached while the task is still running on the server side. | Call |
Clean up
After you finish the tutorial, delete the collection that you created in Step 4.
Dropping the collection deletes all entities and vectors in it. Confirm that no other workload uses the collection before you run the following code.
client.drop_collection(collection_name)Solution benefits
The following table compares this pipeline with integrating each generative model from scratch and assembling your own retrieval system.
Dimension | Traditional in-house solution | Alibaba Cloud Milvus |
Generative model integration | Integrate the SDK, authentication, and response format of each text-to-image, image-to-image, and text-to-video generation service separately | AI functions provide a unified wrapper for image generation, video generation, and content understanding |
Asynchronous task management | Build your own task queue, polling backoff, terminal-state determination, and timeout fallback |
|
Asset understanding | Build or integrate your own title, caption, and tag service |
|
Vectorization pipeline | The application calls the embedding model first and then writes the data | Vectorization on write, generated automatically by the collection function |
Search entry point | Images, videos, text, and vectors are stored as four separate datasets and must be assembled across systems | A single |
Every generated item is stored with its caption, its tags, and a vector, so each generation becomes a searchable asset instead of a one-off output. When a new script calls for a similar shot, run a semantic search against the historical asset library first, reuse the assets that the search recalls, and generate only what it does not.