Alibaba Cloud Milvus 2.6 provides AI Functions that run model inference directly during writes and searches. This tutorial shows you how to use these functions to analyze autonomous driving frames. You create one collection, attach several functions to it, and then run text-to-frame search, image-to-frame search, structured filtering for corner case mining, and multimodal reranking. By the end, you have a single-collection pipeline that turns raw frame URLs into searchable vectors and structured scene data.
Solution overview
In an autonomous driving development pipeline, in-vehicle cameras are the first entry point for data. A test vehicle usually carries 6 to 12 surround-view cameras that capture footage at 20 to 30 FPS, so a single vehicle generates several terabytes of video per day of road testing. This video is the core asset for training perception models, reproducing accidents, and iterating on planning and control strategies, but it is massive, unstructured, and hard to search.
What matters is not a video clip itself, but what happens inside it. Frame-level scene understanding must identify traffic participants (pedestrians, non-motorized vehicles, the vehicle ahead), traffic control elements (traffic light status, lane lines, speed limit signs), the road environment (intersection, highway, tunnel, construction zone), and abnormal events (cut-ins, red light running, hard braking, road debris). Only after these results are persisted as structured data can the following scenarios be supported:
Data replay — Replay real road frames into simulation and training pipelines to iterate on perception models.
Corner case mining — Retrieve long-tail scenarios such as "a construction zone at a tunnel entrance on a rainy night" from a massive frame set. These are the cases where models fail most often and samples are scarcest.
Auto-labeling — Use a large model to produce coarse labels for frames instead of large amounts of manual bounding-box annotation, leaving review and refinement to humans.
Road-test reports — Aggregate statistics by scene and event to produce quantified periodic reports.
The AI Functions of Milvus 2.6 integrate these capabilities into the vector database itself. Instead of maintaining a separate external pipeline, model inference runs as a function that Milvus triggers internally duringinsertandsearch. This tutorial uses the following functions.
| Capability | Function type in code | Purpose | Use in frame analysis |
AI_EMBEDDING | FunctionType.TEXTEMBEDDING | Converts a frame image into a 2,560-dimensional vector with qwen3-vl-embedding. The multimodal model maps text and images into the same vector space. | Supports text-to-frame and image-to-frame search, so semantic requests such as "construction zone in the rain" match by vector retrieval. |
AI_CLASSIFY | TEXTTRANSFORM, task="ai_classify" | Selects the best-matching label for each frame from a preset label set and writes it to a classification field. | Tags each frame with a scene label such as intersection, highway, tunnel, or construction zone. |
AI_EXTRACT | TEXTTRANSFORM, task="ai_extract" | Extracts elements from a frame based on the specified labels and writes them to a structured field as JSON. | Extracts traffic light status, lane count, weather, and abnormal events for precise filtering. |
AI_ENTITY_EXTRACT | TEXTTRANSFORM, task="ai_entity_extract" | Recognizes named entities that explicitly appear in a frame. | Extracts place names, road names, and speed limit values from signs. |
AI_RERANK | FunctionType.RERANK | Scores the top-N vector recall results again with qwen3-vl-rerank (optional). | Reranks results by consistency of subject, action, and scene to improve fine-ranking for corner case mining. |
The capability names (such as AI_EMBEDDING) refer to the functions conceptually, and the code uses the corresponding function_type and task values shown in the table. The first four functions run automatically on write and form the "inference on write" pipeline. AI_RERANK applies to the retrieval stage and is an optional fine-ranking enhancement.
Prerequisites
A Milvus 2.6 instance. AI Functions require the 2.6 kernel, and no separate model service binding is needed after you create the instance.
Public network access, if you connect over the Internet. On the Security Configuration tab of the instance details page, enable Public Network Access, and add the egress IP address of your client to the public access whitelist.
pymilvus installed. The examples in this tutorial are verified with pymilvus 3.0.0.
ffmpeg installed, if you extract frames yourself.
Frame images uploaded to an address that the model can access over the Internet, such as OSS.
The connection endpoint (URI) and access token of your Milvus instance, which you provide in Step 2.
Limitations and considerations
Review the following constraints before you run the procedure. Each one causes a failure or silently incorrect results if you miss it.
Connection port — The RESTful interface and gRPC share port 19530, so you must specify the port explicitly, for example
http://c-xxx.milvus.aliyuncs.com:19530. If you omit the port, the request goes to port 80 by default and the connection times out.Vector dimension — The
dimof the vector field must match thedimin the embedding function parameters (2560).Multimodal flag — A multimodal function must include
"is_multimodal": "true".Default values for extraction labels — Define a default value for every extraction label. When a JSON field is null, a
!=condition does not match that row, which silently drops frames from filtered results.Embedding batch size — The multimodal batch limit of
qwen3-vl-embeddingis 10. A batch that exceeds the limit returns the errorimage batch size can should be [1, 10](returned verbatim by the service). Write frames in batches of 10 or fewer.Flush after writes — After a write, call
flush(). Otherwise, a search that immediately follows may return empty results.REST reranking — The
documentsparameter of the REST rerank interface takes an array of frame URL strings and does not supporttop_n, so it returns one score for each candidate.Model invocations per frame — Each frame you ingest triggers four function invocations (embedding, classification, extraction, and entity recognition) that run automatically on write. Account for this when you plan capacity for large frame sets.
Procedure
This procedure builds the pipeline end to end. Note the following before you start:
Step 2 prepares shared connection code that Steps 3 through 6 all depend on. Run it first.
Step 1 is only needed if you do not already have frame images.
Step 6 (reranking) is optional.
Step 7 removes the resources you create.
Step 1: Extract frames from video
In-vehicle cameras capture continuous video, so you first extract frames with ffmpeg, upload the frame images to an address that the model can access, and then pass frame_url to the ingestion code. During extraction, record the clip_id and ts_ms of each frame (the video clip ID and the frame timestamp) so that you can trace a search hit back to the original video.
# Time-based extraction: take one frame every second and scale it to a width of 960
ffmpeg -i clip001.mp4 -vf "fps=1,scale=960:-1" -q:v 3 frames/clip001_%04d.jpg
# Keyframes only (I-frames): more information and less redundancy
ffmpeg -i clip001.mp4 -vf "select='eq(pict_type,I)',scale=960:-1" -fps_mode vfr -q:v 3 frames/clip001_key_%04d.jpg
# A single frame at a specific point in time (for example, 5.2s, which corresponds to ts_ms=5200)
ffmpeg -ss 5.2 -i clip001.mp4 -frames:v 1 -vf scale=960:-1 -q:v 3 frames/clip001_5200ms.jpgChoose the extraction strategy based on your goal:
Time-based extraction — Regular sampling for building training and replay datasets.
Keyframes only (I-frames) — More information with less redundancy. Recommended for large-scale corner case mining, where reducing the frame count controls cost.
A single frame at a specific time — Pinpoint a known moment, such as reproducing an accident at a specific timestamp.
The-1inscale=960:-1calculates the height automatically from the original aspect ratio. Use-fps_mode vfrto take I-frames. The older-vsync vfrsyntax still runs but reports a deprecation warning.
Step 2: Prepare the shared connection code
The following code contains the connection settings and a compatibility wrapper for the TEXTTRANSFORM function type. Replace MILVUS_URI and MILVUS_TOKEN with the values of your instance.
from __future__ import annotations
import json
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "root:xxx"
# Alibaba Cloud Milvus exposes TEXTTRANSFORM as a managed extension whose function type value is 9.
# Some pymilvus versions do not include this member in the FunctionType enum, so the wrapper below adds it.
TEXTTRANSFORM_FUNCTION_TYPE = 9
def texttransform_function_type() -> FunctionType:
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
function_type = getattr(FunctionType, type_name, None)
if function_type is not None:
return function_type
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
VECTOR_DIM = 2560
EMBED_MODEL = "qwen3-vl-embedding"
VLM_MODEL = "qwen3.7-plus" # Multimodal understanding model (classification, extraction, entities)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)AI_CLASSIFY, AI_EXTRACT, and AI_ENTITY_EXTRACT are all functions of the TEXTTRANSFORM type (function type value 9) and are distinguished by the task parameter. Some pymilvus versions do not include this member in the FunctionType enum, which is why the compatibility wrapper is required.
If you change EMBED_MODEL, the vector dimension may change too. Keep VECTOR_DIM and the schema in Step 3 in sync with the model you choose, because reindexing an existing collection is a high-cost operation.
The example hardcodes credentials for demonstration only. In production, load MILVUS_URI and MILVUS_TOKEN from environment variables or a secret manager, do not commit them to source control, and avoid using the root account.
Step 3: Create the collection and attach the AI functions
The vector field serves semantic retrieval, and the other three fields hold the results of classification, structured extraction, and entity recognition. All four functions take frame_url as input and run automatically on write.
If a collection named driving_frames already exists, the example drops it and permanently deletes its data. Use a unique collection name if you have existing data, or back it up first. The drop branch is included so that you can rerun the tutorial from a clean state.
# ===== Create the collection: attach four AI Functions in a single definition =====
collection_name = "driving_frames"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("frame_url", DataType.VARCHAR, max_length=4096) # Frame image address
schema.add_field("clip_id", DataType.VARCHAR, max_length=128) # ID of the source video clip
schema.add_field("ts_ms", DataType.INT64) # Frame timestamp (ms)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=VECTOR_DIM)
schema.add_field("scene", DataType.VARCHAR, max_length=64) # AI_CLASSIFY output
schema.add_field("attributes", DataType.JSON) # AI_EXTRACT output
schema.add_field("entities", DataType.JSON) # AI_ENTITY_EXTRACT output
# 1) Multimodal embedding: frame image -> 2,560-dimensional vector
schema.add_function(Function(
name="embed_frame", function_type=FunctionType.TEXTEMBEDDING,
input_field_names=["frame_url"], output_field_names=["embedding"],
params={"provider": "aliyun_milvus", "model_name": EMBED_MODEL,
"dim": VECTOR_DIM, "is_multimodal": "true"}))
# 2) Scene classification: intersection/highway/tunnel/construction zone/ordinary road
schema.add_function(Function(
name="classify_scene", function_type=texttransform_function_type(),
input_field_names=["frame_url"], output_field_names=["scene"],
params={"provider": "aliyun_milvus", "model_name": VLM_MODEL,
"task": "ai_classify", "media_type": "image",
"labels": "intersection,highway,tunnel,construction zone,ordinary road",
"prompt": "Classify the frame by the road environment it shows.", "temperature": "0"}))
# 3) Structured extraction: traffic light/lane count/weather/abnormal event
# Note: define a default value for every label to prevent the model from returning null (see the description below)
schema.add_function(Function(
name="extract_traffic", function_type=texttransform_function_type(),
input_field_names=["frame_url"], output_field_names=["attributes"],
params={"provider": "aliyun_milvus", "model_name": VLM_MODEL,
"task": "ai_extract", "media_type": "image",
"labels": "traffic_light,lane_count,weather,anomaly_event",
"prompt": "traffic_light takes red/green/yellow/none; "
"weather takes sunny/rainy/cloudy/night/unknown, use unknown when it cannot be determined; "
"lane_count takes an integer, use 0 when it cannot be determined; "
"anomaly_event describes abnormal events such as cut-ins, red light running, or accidents, use none if there are none.",
"temperature": "0"}))
# 4) Named entities: road names/place names/speed limit values
schema.add_function(Function(
name="extract_entities", function_type=texttransform_function_type(),
input_field_names=["frame_url"], output_field_names=["entities"],
params={"provider": "aliyun_milvus", "model_name": VLM_MODEL,
"task": "ai_entity_extract", "media_type": "image",
"entity_types": "LOCATION,PRODUCT",
"prompt": "Extract only the place names, road names, and speed limit values that explicitly appear on traffic signs in the frame. Do not guess from appearance.",
"temperature": "0"}))
index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)The main function parameters are:
provider— The model provider. This tutorial usesaliyun_milvus.model_name— The model that backs the function.task— Selects the operation for aTEXTTRANSFORMfunction:ai_classify,ai_extract, orai_entity_extract.media_type— Set toimagebecause the input is a frame image.labels— The comma-separated label set for classification or extraction.entity_types— The comma-separated entity types to recognize. This example maps place and road names toLOCATIONand speed limit values toPRODUCT.temperature— Set to0for deterministic output.
In the SDK,is_multimodalanddim— See Limitations and considerations. A multimodal function must setis_multimodal, and the vector fielddimmust match the functiondim(2560).paramsvalues are passed as strings (for example,"is_multimodal": "true"and"temperature": "0"). In the REST payload used in Step 6,paramsvalues use native JSON types (for example,"is_multimodal": true).
Define a default value for every extraction label. The prompt above assigns a default for each label so that the model avoids returning null. This matters because a null JSON field is not matched by a != condition, which silently drops frames from filtered results. Defaults reduce, but do not always eliminate, missing values: as the ingestion results in Step 4 show, the tunnel frame still has no weather value. Design your filters to account for missing values, and see Limitations and considerations.
Successful ingestion in Step 4 (fields populated automatically) confirms that the collection and its four functions are attached correctly.
Step 4: Ingest frames with inference on write
You write only the three fields frame_url, clip_id, and ts_ms. The AI Functions populate the other four fields automatically, with no manual annotation.
# ===== Ingest frames: inference on write populates embedding/scene/attributes/entities automatically =====
# Replace frame_url with your own image address that the model can access over the Internet.
# clip_id and ts_ms record the source clip and timestamp of the frame, so you can trace a hit back to the original video.
frames = [
{"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_dashcam_5000.jpg",
"clip_id": "clip_dashcam", "ts_ms": 5000},
{"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_dashcam_12000.jpg",
"clip_id": "clip_dashcam", "ts_ms": 12000},
{"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_highway_3000.jpg",
"clip_id": "clip_highway", "ts_ms": 3000},
{"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_highway_9000.jpg",
"clip_id": "clip_highway", "ts_ms": 9000},
{"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_urban_2000.jpg",
"clip_id": "clip_urban", "ts_ms": 2000},
{"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_urban_8000.jpg",
"clip_id": "clip_urban", "ts_ms": 8000},
]
# Keep each batch at or below the multimodal batch limit of 10 (see Limitations and considerations).
_BATCH = 8
for _i in range(0, len(frames), _BATCH):
client.insert(collection_name, frames[_i:_i + _BATCH])
client.flush(collection_name)
client.load_collection(collection_name)
# After the write completes, query the structured results directly to verify them
ingested_rows = client.query(
collection_name, filter="",
output_fields=["frame_url", "clip_id", "ts_ms", "scene", "attributes", "entities"],
limit=100)
for row in ingested_rows:
print(f"{row['clip_id']}@{row['ts_ms']}ms scene={row['scene']} "
f"attributes={json.dumps(row['attributes'], ensure_ascii=False)} "
f"entities={json.dumps(row['entities'], ensure_ascii=False)}")The write loop keeps each batch at 8, which stays within the multimodal batch limit of 10. After the write, you must call flush(); otherwise, a search that immediately follows may return empty results. The verification query uses limit=100, so it returns up to 100 rows as a spot check. For larger frame sets, page through the results or run a count to confirm the full write.
In tests, writing six frames plus flush and load took about 12 seconds. The following table shows the reference output for the example frame set. The exact values depend on your own frames, so focus on whether each field is auto-populated and whether the classification matches the image.
| Frame | scene | attributes | entities |
| dashcam@5000ms | intersection | traffic_light=green, lane_count=4, weather=sunny, anomaly_event=none | [] |
| dashcam@12000ms | construction zone | traffic_light=none, lane_count=3, weather=rainy, anomaly_event=none | [] |
| highway@3000ms | highway | traffic_light=none, lane_count=4, weather=sunny, anomaly_event=none | [{"text":"EXIT 111","type":"LOCATION"}] |
| highway@9000ms | tunnel | traffic_light=none, lane_count=2, anomaly_event=none | [] |
| urban@2000ms | ordinary road | weather=cloudy, anomaly_event=pedestrians and animals on the road | [] |
| urban@8000ms | intersection | traffic_light=red, lane_count=4, weather=sunny, anomaly_event=none | [] |
The scene classification of all six frames matched the image content, and the extracted traffic light status, weather, and lane count were consistent with the images for the frames where those fields were populated. Some fields were still absent for certain frames (for example, weather for the tunnel frame and traffic_light/lane_count for the ordinary-road frame), which is why you must account for missing values when you filter. Named entities were extracted only for frames that contained readable sign text, and the model did not infer place names from appearance, which matches the "do not guess from appearance" constraint in the prompt.
Step 5: Search frames
The vectors and structured fields produced during the write stage combine at the retrieval stage: vectors handle semantic recall, and structured fields handle precise filtering.
To search by text
The query text is mapped into the image vector space by the same multimodal model and recalled directly.
query = "construction zone in the rain, with traffic cones and construction signs on the road"
text_search = client.search(
collection_name=collection_name, data=[query], anns_field="embedding",
limit=10, output_fields=["frame_url", "scene", "attributes"])
for hit in text_search[0]:
print(f"score={hit['distance']:.4f} scene={hit['entity']['scene']}")To combine vector recall with structured filtering (corner case mining)
The structured fields come from AI_EXTRACT and AI_CLASSIFY at write time and can be used in filter directly.
corner_query = "abnormal event on the road at night"
corner_search = client.search(
collection_name=collection_name, data=[corner_query], anns_field="embedding",
limit=10, filter='attributes["anomaly_event"] != "none"',
output_fields=["frame_url", "scene", "attributes"])
print(f"Matched {len(corner_search[0])} long-tail scenes")Structured filtering supports string equality and numeric comparison. For example, attributes["lane_count"] >= 4 selects frames with four or more lanes. When you combine conditions, remember that a != condition does not match a row whose JSON field value is null, which can silently drop long-tail samples. Combined conditions must match the actual content of your frame library: scene == "tunnel" and attributes["anomaly_event"] != "none" requires both a tunnel scene and an abnormal event. If no such frame exists, the query returns 0 rows, which is itself a valid corner case mining conclusion, not a malfunction. During debugging, confirm that the pipeline works with a single condition first, then add conditions to narrow the range step by step.
To search by image
Replace data with an image URL and keep everything else unchanged.
image_query_url = frames[0]["frame_url"]
image_search = client.search(
collection_name=collection_name, data=[image_query_url],
anns_field="embedding", limit=10, output_fields=["frame_url", "scene"])
for hit in image_search[0]:
print(f"score={hit['distance']:.4f} scene={hit['entity']['scene']}")The following table shows reference results for the example frame set. Focus on the relative ordering, not on the absolute scores, which depend on your own frames.
| Search method | Query | Result |
| Text-to-frame search | construction zone in the rain, with traffic cones and construction signs on the road | The construction zone frame ranked first at 0.5904, and the runner-up scored only 0.1797, a clear margin. |
| Image-to-frame search | The URL of an intersection frame used as the query | The query image itself scored 1.0000, a similar intersection frame scored 0.5534, and the least relevant frame scored 0.1162. |
| Structured filtering | attributes["anomaly_event"] != "none" | Matched the frames that actually contain an abnormal event. |
In image-to-frame search, the query image scores 1.0000 against itself. This is a reproducible check that you can use to verify the consistency of image encoding.
Step 6: Rerank results with the rerank function (optional)
Vector similarity measures semantic closeness, which is not fully equivalent to relevance to the query intent. You can use qwen3-vl-rerank to fine-rank the recalled frames and improve ranking quality for long-tail scenarios. Choose one of two patterns before you read the code:
Attach a ranker in
search— Use this when you want recall and fine-ranking in a single call.Call the REST rerank interface — Use this when you already have a candidate frame list and only want to rerank it.
Both patterns produce identical scores for the same frame, so you can choose either one based on your engineering needs.
Pattern 1 attaches a ranker in search, so qwen3-vl-rerank fine-ranks the top-N vector recall results.
# ===== Multimodal reranking with AI_RERANK (optional) =====
QUERY = "cut-in behavior at a highway ramp"
reranker = Function(
name="rerank_frames", function_type=FunctionType.RERANK,
input_field_names=["frame_url"],
params={"reranker": "model", "provider": "aliyun_milvus",
"model_name": "qwen3-vl-rerank", "queries": [QUERY],
"is_multimodal": "true",
"instruct": "Rank candidate frames by relevance to the query, "
"prioritizing subject, action, scene and fine-grained visual details.",
"timeout_sec": 10})
rerank_search = client.search(
collection_name=collection_name, data=[QUERY], anns_field="embedding",
limit=20, output_fields=["frame_url", "scene"], ranker=reranker)
for hit in rerank_search[0]:
print(f"rerank_score={hit['distance']:.4f} scene={hit['entity']['scene']}")Pattern 2 reranks only a set of already recalled frame URLs through the REST endpoint /v2/vectordb/ai/rerank. The following helper posts a JSON request; define it before the REST call.
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
request = Request(
f"{MILVUS_URI.rstrip('/')}{path}",
data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
try:
with urlopen(request, timeout=timeout) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
rest_docs = [f["frame_url"] for f in frames[:3]]
status, data = post_json("/v2/vectordb/ai/rerank", {
"model_name": "qwen3-vl-rerank", "query": QUERY,
"documents": rest_docs,
"params": {"is_multimodal": True,
"instruct": "Rank candidate frames by relevance to the query, "
"prioritizing subject, action, scene and fine-grained visual details.",
"timeout_sec": 10}})
assert status == 200 and data.get("code") == 0, data
for item in sorted(data["data"]["output"]["results"],
key=lambda x: x["relevance_score"], reverse=True):
print(f"index={item['index']} relevance_score={item['relevance_score']:.4f}")The assert statement is an example-only check. In production, replace it with explicit error handling that inspects status and the response code, and log the response instead of raising.
Multimodal reranking requires is_multimodal in params, and you can use instruct to specify the ranking focus, such as prioritizing subject, action, scene, and fine-grained visual details. The REST documents parameter takes an array of frame URL strings and does not support top_n, so it returns one score for each candidate.
The following table shows reference reranking results for the test query "cut-in behavior at a highway ramp." Focus on the relative ordering, not on the absolute scores.
| Rank | Frame scene | Rerank score |
| 1 | highway | 0.5898 |
| 2 | construction zone | 0.5018 |
| 3 | tunnel | 0.4858 |
| 4 | intersection | 0.4830 |
| 5 | intersection | 0.4089 |
| 6 | ordinary road | 0.3819 |
The highway frame ranked first, which matches the query semantics.
Step 7: Clean up resources
After you finish the tutorial, delete the resources you created to avoid unnecessary charges.
Drop the collection:
client.drop_collection(collection_name)Confirm the collection is deleted. The following command should return
False:print(client.has_collection(collection_name))Delete the frame images you uploaded to OSS for the tutorial, if you no longer need them.
(Optional) If you enabled public network access only for this tutorial, disable Public Network Access on the Security Configuration tab of the instance details page.
Troubleshooting
Use the following table to diagnose common issues. The constraints behind these symptoms are described in Limitations and considerations.
| Symptom | Cause | Solution |
| A search immediately after a write returns empty results. | flush() was not called after the write. | Call flush() after writing, then search. |
| The connection times out. | The port was omitted, so the request went to port 80. | Specify port 19530 explicitly in the URI. |
| A combined filter returns 0 rows. | No frame in the library matches all conditions. | This can be a valid corner case mining conclusion, not a fault. Confirm the pipeline with a single condition, then add conditions gradually. |
A batch write returns image batch size can should be [1, 10]. | The batch exceeds the multimodal limit of 10. | Write frames in batches of 10 or fewer. |
Solution value
The following table compares a traditional self-built solution with the Milvus AI Function approach.
| Dimension | Traditional self-built solution | Milvus AI Function |
| Development cycle | Joint debugging across frame extraction, inference, vector database, and metadata database | Attach functions when you create the collection, then write and search directly |
| Inference operations | Self-built GPU inference cluster that needs scaling and fault recovery | Invocations managed by Milvus, with no inference cluster to operate |
| Data movement | Frames move repeatedly between object storage, the inference cluster, and the vector database | Inference on write, and data stays within the Milvus instance |
| Multimodal retrieval | Requires self-built alignment of the text and image vector spaces | qwen3-vl-embedding natively supports text-to-frame and image-to-frame search |
| Annotation criteria | Manual annotation is slow and expensive, and criteria vary | prompt unifies the judgment rules, so label criteria are consistent and reproducible |
For an autonomous driving team, this approach:
Speeds up corner case mining — Text-to-frame search plus structured filtering, combined with multimodal reranking, retrieves a long-tail scenario from a single natural language sentence.
Lowers auto-labeling cost —
AI_CLASSIFY,AI_EXTRACT, andAI_ENTITY_EXTRACTproduce labels with unified criteria at write time, leaving review and refinement to humans.Converges the pipeline — Inference runs inside the vector database. Writes produce vectors and structured understanding, searches perform multimodal semantic recall, and driving data stays within the Milvus instance.