All Products
Search
Document Center

Vector Retrieval Service for Milvus:Frame analysis and multimodal retrieval for autonomous driving with Alibaba Cloud Milvus

Last Updated:Aug 13, 2026

Alibaba Cloud Milvus 2.6 provides AI Functions that run model inference directly during writes and searches. This tutorial shows you how to use these functions to analyze autonomous driving frames. You create one collection, attach several functions to it, and then run text-to-frame search, image-to-frame search, structured filtering for corner case mining, and multimodal reranking. By the end, you have a single-collection pipeline that turns raw frame URLs into searchable vectors and structured scene data.

Solution overview

In an autonomous driving development pipeline, in-vehicle cameras are the first entry point for data. A test vehicle usually carries 6 to 12 surround-view cameras that capture footage at 20 to 30 FPS, so a single vehicle generates several terabytes of video per day of road testing. This video is the core asset for training perception models, reproducing accidents, and iterating on planning and control strategies, but it is massive, unstructured, and hard to search.

What matters is not a video clip itself, but what happens inside it. Frame-level scene understanding must identify traffic participants (pedestrians, non-motorized vehicles, the vehicle ahead), traffic control elements (traffic light status, lane lines, speed limit signs), the road environment (intersection, highway, tunnel, construction zone), and abnormal events (cut-ins, red light running, hard braking, road debris). Only after these results are persisted as structured data can the following scenarios be supported:

  • Data replay — Replay real road frames into simulation and training pipelines to iterate on perception models.

  • Corner case mining — Retrieve long-tail scenarios such as "a construction zone at a tunnel entrance on a rainy night" from a massive frame set. These are the cases where models fail most often and samples are scarcest.

  • Auto-labeling — Use a large model to produce coarse labels for frames instead of large amounts of manual bounding-box annotation, leaving review and refinement to humans.

  • Road-test reports — Aggregate statistics by scene and event to produce quantified periodic reports.

    The AI Functions of Milvus 2.6 integrate these capabilities into the vector database itself. Instead of maintaining a separate external pipeline, model inference runs as a function that Milvus triggers internally during insert and search. This tutorial uses the following functions.
CapabilityFunction type in codePurposeUse in frame analysis
AI_EMBEDDINGFunctionType.TEXTEMBEDDINGConverts a frame image into a 2,560-dimensional vector with qwen3-vl-embedding. The multimodal model maps text and images into the same vector space.Supports text-to-frame and image-to-frame search, so semantic requests such as "construction zone in the rain" match by vector retrieval.
AI_CLASSIFYTEXTTRANSFORM, task="ai_classify"Selects the best-matching label for each frame from a preset label set and writes it to a classification field.Tags each frame with a scene label such as intersection, highway, tunnel, or construction zone.
AI_EXTRACTTEXTTRANSFORM, task="ai_extract"Extracts elements from a frame based on the specified labels and writes them to a structured field as JSON.Extracts traffic light status, lane count, weather, and abnormal events for precise filtering.
AI_ENTITY_EXTRACTTEXTTRANSFORM, task="ai_entity_extract"Recognizes named entities that explicitly appear in a frame.Extracts place names, road names, and speed limit values from signs.
AI_RERANKFunctionType.RERANKScores the top-N vector recall results again with qwen3-vl-rerank (optional).Reranks results by consistency of subject, action, and scene to improve fine-ranking for corner case mining.

The capability names (such as AI_EMBEDDING) refer to the functions conceptually, and the code uses the corresponding function_type and task values shown in the table. The first four functions run automatically on write and form the "inference on write" pipeline. AI_RERANK applies to the retrieval stage and is an optional fine-ranking enhancement.

Prerequisites

  • A Milvus 2.6 instance. AI Functions require the 2.6 kernel, and no separate model service binding is needed after you create the instance.

  • Public network access, if you connect over the Internet. On the Security Configuration tab of the instance details page, enable Public Network Access, and add the egress IP address of your client to the public access whitelist.

  • pymilvus installed. The examples in this tutorial are verified with pymilvus 3.0.0.

  • ffmpeg installed, if you extract frames yourself.

  • Frame images uploaded to an address that the model can access over the Internet, such as OSS.

  • The connection endpoint (URI) and access token of your Milvus instance, which you provide in Step 2.

Limitations and considerations

Review the following constraints before you run the procedure. Each one causes a failure or silently incorrect results if you miss it.

  • Connection port — The RESTful interface and gRPC share port 19530, so you must specify the port explicitly, for example http://c-xxx.milvus.aliyuncs.com:19530. If you omit the port, the request goes to port 80 by default and the connection times out.

  • Vector dimension — The dim of the vector field must match the dim in the embedding function parameters (2560).

  • Multimodal flag — A multimodal function must include "is_multimodal": "true".

  • Default values for extraction labels — Define a default value for every extraction label. When a JSON field is null, a != condition does not match that row, which silently drops frames from filtered results.

  • Embedding batch size — The multimodal batch limit of qwen3-vl-embedding is 10. A batch that exceeds the limit returns the error image batch size can should be [1, 10] (returned verbatim by the service). Write frames in batches of 10 or fewer.

  • Flush after writes — After a write, call flush(). Otherwise, a search that immediately follows may return empty results.

  • REST reranking — The documents parameter of the REST rerank interface takes an array of frame URL strings and does not support top_n, so it returns one score for each candidate.

  • Model invocations per frame — Each frame you ingest triggers four function invocations (embedding, classification, extraction, and entity recognition) that run automatically on write. Account for this when you plan capacity for large frame sets.

Procedure

This procedure builds the pipeline end to end. Note the following before you start:

  • Step 2 prepares shared connection code that Steps 3 through 6 all depend on. Run it first.

  • Step 1 is only needed if you do not already have frame images.

  • Step 6 (reranking) is optional.

  • Step 7 removes the resources you create.

Step 1: Extract frames from video

In-vehicle cameras capture continuous video, so you first extract frames with ffmpeg, upload the frame images to an address that the model can access, and then pass frame_url to the ingestion code. During extraction, record the clip_id and ts_ms of each frame (the video clip ID and the frame timestamp) so that you can trace a search hit back to the original video.

# Time-based extraction: take one frame every second and scale it to a width of 960
ffmpeg -i clip001.mp4 -vf "fps=1,scale=960:-1" -q:v 3 frames/clip001_%04d.jpg

# Keyframes only (I-frames): more information and less redundancy
ffmpeg -i clip001.mp4 -vf "select='eq(pict_type,I)',scale=960:-1" -fps_mode vfr -q:v 3 frames/clip001_key_%04d.jpg

# A single frame at a specific point in time (for example, 5.2s, which corresponds to ts_ms=5200)
ffmpeg -ss 5.2 -i clip001.mp4 -frames:v 1 -vf scale=960:-1 -q:v 3 frames/clip001_5200ms.jpg

Choose the extraction strategy based on your goal:

  • Time-based extraction — Regular sampling for building training and replay datasets.

  • Keyframes only (I-frames) — More information with less redundancy. Recommended for large-scale corner case mining, where reducing the frame count controls cost.

  • A single frame at a specific time — Pinpoint a known moment, such as reproducing an accident at a specific timestamp.

    The -1 in scale=960:-1 calculates the height automatically from the original aspect ratio. Use -fps_mode vfr to take I-frames. The older -vsync vfr syntax still runs but reports a deprecation warning.

Step 2: Prepare the shared connection code

The following code contains the connection settings and a compatibility wrapper for the TEXTTRANSFORM function type. Replace MILVUS_URI and MILVUS_TOKEN with the values of your instance.

from __future__ import annotations

import json

from pymilvus import DataType, Function, FunctionType, MilvusClient

MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "root:xxx"

# Alibaba Cloud Milvus exposes TEXTTRANSFORM as a managed extension whose function type value is 9.
# Some pymilvus versions do not include this member in the FunctionType enum, so the wrapper below adds it.
TEXTTRANSFORM_FUNCTION_TYPE = 9

def texttransform_function_type() -> FunctionType:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        function_type = getattr(FunctionType, type_name, None)
        if function_type is not None:
            return function_type
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension

VECTOR_DIM = 2560
EMBED_MODEL = "qwen3-vl-embedding"
VLM_MODEL = "qwen3.7-plus"          # Multimodal understanding model (classification, extraction, entities)
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

AI_CLASSIFY, AI_EXTRACT, and AI_ENTITY_EXTRACT are all functions of the TEXTTRANSFORM type (function type value 9) and are distinguished by the task parameter. Some pymilvus versions do not include this member in the FunctionType enum, which is why the compatibility wrapper is required.

If you change EMBED_MODEL, the vector dimension may change too. Keep VECTOR_DIM and the schema in Step 3 in sync with the model you choose, because reindexing an existing collection is a high-cost operation.

Warning

The example hardcodes credentials for demonstration only. In production, load MILVUS_URI and MILVUS_TOKEN from environment variables or a secret manager, do not commit them to source control, and avoid using the root account.

Step 3: Create the collection and attach the AI functions

The vector field serves semantic retrieval, and the other three fields hold the results of classification, structured extraction, and entity recognition. All four functions take frame_url as input and run automatically on write.

Warning

If a collection named driving_frames already exists, the example drops it and permanently deletes its data. Use a unique collection name if you have existing data, or back it up first. The drop branch is included so that you can rerun the tutorial from a clean state.

# ===== Create the collection: attach four AI Functions in a single definition =====
collection_name = "driving_frames"
if client.has_collection(collection_name):
    client.drop_collection(collection_name)

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("frame_url", DataType.VARCHAR, max_length=4096)   # Frame image address
schema.add_field("clip_id", DataType.VARCHAR, max_length=128)      # ID of the source video clip
schema.add_field("ts_ms", DataType.INT64)                          # Frame timestamp (ms)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=VECTOR_DIM)
schema.add_field("scene", DataType.VARCHAR, max_length=64)         # AI_CLASSIFY output
schema.add_field("attributes", DataType.JSON)                      # AI_EXTRACT output
schema.add_field("entities", DataType.JSON)                        # AI_ENTITY_EXTRACT output

# 1) Multimodal embedding: frame image -> 2,560-dimensional vector
schema.add_function(Function(
    name="embed_frame", function_type=FunctionType.TEXTEMBEDDING,
    input_field_names=["frame_url"], output_field_names=["embedding"],
    params={"provider": "aliyun_milvus", "model_name": EMBED_MODEL,
            "dim": VECTOR_DIM, "is_multimodal": "true"}))

# 2) Scene classification: intersection/highway/tunnel/construction zone/ordinary road
schema.add_function(Function(
    name="classify_scene", function_type=texttransform_function_type(),
    input_field_names=["frame_url"], output_field_names=["scene"],
    params={"provider": "aliyun_milvus", "model_name": VLM_MODEL,
            "task": "ai_classify", "media_type": "image",
            "labels": "intersection,highway,tunnel,construction zone,ordinary road",
            "prompt": "Classify the frame by the road environment it shows.", "temperature": "0"}))

# 3) Structured extraction: traffic light/lane count/weather/abnormal event
#    Note: define a default value for every label to prevent the model from returning null (see the description below)
schema.add_function(Function(
    name="extract_traffic", function_type=texttransform_function_type(),
    input_field_names=["frame_url"], output_field_names=["attributes"],
    params={"provider": "aliyun_milvus", "model_name": VLM_MODEL,
            "task": "ai_extract", "media_type": "image",
            "labels": "traffic_light,lane_count,weather,anomaly_event",
            "prompt": "traffic_light takes red/green/yellow/none; "
                      "weather takes sunny/rainy/cloudy/night/unknown, use unknown when it cannot be determined; "
                      "lane_count takes an integer, use 0 when it cannot be determined; "
                      "anomaly_event describes abnormal events such as cut-ins, red light running, or accidents, use none if there are none.",
            "temperature": "0"}))

# 4) Named entities: road names/place names/speed limit values
schema.add_function(Function(
    name="extract_entities", function_type=texttransform_function_type(),
    input_field_names=["frame_url"], output_field_names=["entities"],
    params={"provider": "aliyun_milvus", "model_name": VLM_MODEL,
            "task": "ai_entity_extract", "media_type": "image",
            "entity_types": "LOCATION,PRODUCT",
            "prompt": "Extract only the place names, road names, and speed limit values that explicitly appear on traffic signs in the frame. Do not guess from appearance.",
            "temperature": "0"}))

index_params = client.prepare_index_params()
index_params.add_index(field_name="embedding", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
                        index_params=index_params)

The main function parameters are:

  • provider — The model provider. This tutorial uses aliyun_milvus.

  • model_name — The model that backs the function.

  • task — Selects the operation for a TEXTTRANSFORM function: ai_classify, ai_extract, or ai_entity_extract.

  • media_type — Set to image because the input is a frame image.

  • labels — The comma-separated label set for classification or extraction.

  • entity_types — The comma-separated entity types to recognize. This example maps place and road names to LOCATION and speed limit values to PRODUCT.

  • temperature — Set to 0 for deterministic output.

  • is_multimodal and dim — See Limitations and considerations. A multimodal function must set is_multimodal, and the vector field dim must match the function dim (2560).

    In the SDK, params values are passed as strings (for example, "is_multimodal": "true" and "temperature": "0"). In the REST payload used in Step 6, params values use native JSON types (for example, "is_multimodal": true).

Define a default value for every extraction label. The prompt above assigns a default for each label so that the model avoids returning null. This matters because a null JSON field is not matched by a != condition, which silently drops frames from filtered results. Defaults reduce, but do not always eliminate, missing values: as the ingestion results in Step 4 show, the tunnel frame still has no weather value. Design your filters to account for missing values, and see Limitations and considerations.

Successful ingestion in Step 4 (fields populated automatically) confirms that the collection and its four functions are attached correctly.

Step 4: Ingest frames with inference on write

You write only the three fields frame_url, clip_id, and ts_ms. The AI Functions populate the other four fields automatically, with no manual annotation.

# ===== Ingest frames: inference on write populates embedding/scene/attributes/entities automatically =====
# Replace frame_url with your own image address that the model can access over the Internet.
# clip_id and ts_ms record the source clip and timestamp of the frame, so you can trace a hit back to the original video.
frames = [
    {"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_dashcam_5000.jpg",
     "clip_id": "clip_dashcam", "ts_ms": 5000},
    {"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_dashcam_12000.jpg",
     "clip_id": "clip_dashcam", "ts_ms": 12000},
    {"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_highway_3000.jpg",
     "clip_id": "clip_highway", "ts_ms": 3000},
    {"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_highway_9000.jpg",
     "clip_id": "clip_highway", "ts_ms": 9000},
    {"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_urban_2000.jpg",
     "clip_id": "clip_urban", "ts_ms": 2000},
    {"frame_url": "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/frames/clip_urban_8000.jpg",
     "clip_id": "clip_urban", "ts_ms": 8000},
]

# Keep each batch at or below the multimodal batch limit of 10 (see Limitations and considerations).
_BATCH = 8
for _i in range(0, len(frames), _BATCH):
    client.insert(collection_name, frames[_i:_i + _BATCH])
client.flush(collection_name)
client.load_collection(collection_name)

# After the write completes, query the structured results directly to verify them
ingested_rows = client.query(
    collection_name, filter="",
    output_fields=["frame_url", "clip_id", "ts_ms", "scene", "attributes", "entities"],
    limit=100)
for row in ingested_rows:
    print(f"{row['clip_id']}@{row['ts_ms']}ms  scene={row['scene']}  "
          f"attributes={json.dumps(row['attributes'], ensure_ascii=False)}  "
          f"entities={json.dumps(row['entities'], ensure_ascii=False)}")

The write loop keeps each batch at 8, which stays within the multimodal batch limit of 10. After the write, you must call flush(); otherwise, a search that immediately follows may return empty results. The verification query uses limit=100, so it returns up to 100 rows as a spot check. For larger frame sets, page through the results or run a count to confirm the full write.

In tests, writing six frames plus flush and load took about 12 seconds. The following table shows the reference output for the example frame set. The exact values depend on your own frames, so focus on whether each field is auto-populated and whether the classification matches the image.

Framesceneattributesentities
dashcam@5000msintersectiontraffic_light=green, lane_count=4, weather=sunny, anomaly_event=none[]
dashcam@12000msconstruction zonetraffic_light=none, lane_count=3, weather=rainy, anomaly_event=none[]
highway@3000mshighwaytraffic_light=none, lane_count=4, weather=sunny, anomaly_event=none[{"text":"EXIT 111","type":"LOCATION"}]
highway@9000mstunneltraffic_light=none, lane_count=2, anomaly_event=none[]
urban@2000msordinary roadweather=cloudy, anomaly_event=pedestrians and animals on the road[]
urban@8000msintersectiontraffic_light=red, lane_count=4, weather=sunny, anomaly_event=none[]

The scene classification of all six frames matched the image content, and the extracted traffic light status, weather, and lane count were consistent with the images for the frames where those fields were populated. Some fields were still absent for certain frames (for example, weather for the tunnel frame and traffic_light/lane_count for the ordinary-road frame), which is why you must account for missing values when you filter. Named entities were extracted only for frames that contained readable sign text, and the model did not infer place names from appearance, which matches the "do not guess from appearance" constraint in the prompt.

Step 5: Search frames

The vectors and structured fields produced during the write stage combine at the retrieval stage: vectors handle semantic recall, and structured fields handle precise filtering.

To search by text

The query text is mapped into the image vector space by the same multimodal model and recalled directly.

query = "construction zone in the rain, with traffic cones and construction signs on the road"
text_search = client.search(
    collection_name=collection_name, data=[query], anns_field="embedding",
    limit=10, output_fields=["frame_url", "scene", "attributes"])
for hit in text_search[0]:
    print(f"score={hit['distance']:.4f} scene={hit['entity']['scene']}")

To combine vector recall with structured filtering (corner case mining)

The structured fields come from AI_EXTRACT and AI_CLASSIFY at write time and can be used in filter directly.

corner_query = "abnormal event on the road at night"
corner_search = client.search(
    collection_name=collection_name, data=[corner_query], anns_field="embedding",
    limit=10, filter='attributes["anomaly_event"] != "none"',
    output_fields=["frame_url", "scene", "attributes"])
print(f"Matched {len(corner_search[0])} long-tail scenes")

Structured filtering supports string equality and numeric comparison. For example, attributes["lane_count"] >= 4 selects frames with four or more lanes. When you combine conditions, remember that a != condition does not match a row whose JSON field value is null, which can silently drop long-tail samples. Combined conditions must match the actual content of your frame library: scene == "tunnel" and attributes["anomaly_event"] != "none" requires both a tunnel scene and an abnormal event. If no such frame exists, the query returns 0 rows, which is itself a valid corner case mining conclusion, not a malfunction. During debugging, confirm that the pipeline works with a single condition first, then add conditions to narrow the range step by step.

To search by image

Replace data with an image URL and keep everything else unchanged.

image_query_url = frames[0]["frame_url"]
image_search = client.search(
    collection_name=collection_name, data=[image_query_url],
    anns_field="embedding", limit=10, output_fields=["frame_url", "scene"])
for hit in image_search[0]:
    print(f"score={hit['distance']:.4f} scene={hit['entity']['scene']}")

The following table shows reference results for the example frame set. Focus on the relative ordering, not on the absolute scores, which depend on your own frames.

Search methodQueryResult
Text-to-frame searchconstruction zone in the rain, with traffic cones and construction signs on the roadThe construction zone frame ranked first at 0.5904, and the runner-up scored only 0.1797, a clear margin.
Image-to-frame searchThe URL of an intersection frame used as the queryThe query image itself scored 1.0000, a similar intersection frame scored 0.5534, and the least relevant frame scored 0.1162.
Structured filteringattributes["anomaly_event"] != "none"Matched the frames that actually contain an abnormal event.

In image-to-frame search, the query image scores 1.0000 against itself. This is a reproducible check that you can use to verify the consistency of image encoding.

Step 6: Rerank results with the rerank function (optional)

Vector similarity measures semantic closeness, which is not fully equivalent to relevance to the query intent. You can use qwen3-vl-rerank to fine-rank the recalled frames and improve ranking quality for long-tail scenarios. Choose one of two patterns before you read the code:

  • Attach a ranker in search — Use this when you want recall and fine-ranking in a single call.

  • Call the REST rerank interface — Use this when you already have a candidate frame list and only want to rerank it.

    Both patterns produce identical scores for the same frame, so you can choose either one based on your engineering needs.

Pattern 1 attaches a ranker in search, so qwen3-vl-rerank fine-ranks the top-N vector recall results.

# ===== Multimodal reranking with AI_RERANK (optional) =====
QUERY = "cut-in behavior at a highway ramp"

reranker = Function(
    name="rerank_frames", function_type=FunctionType.RERANK,
    input_field_names=["frame_url"],
    params={"reranker": "model", "provider": "aliyun_milvus",
            "model_name": "qwen3-vl-rerank", "queries": [QUERY],
            "is_multimodal": "true",
            "instruct": "Rank candidate frames by relevance to the query, "
                        "prioritizing subject, action, scene and fine-grained visual details.",
            "timeout_sec": 10})
rerank_search = client.search(
    collection_name=collection_name, data=[QUERY], anns_field="embedding",
    limit=20, output_fields=["frame_url", "scene"], ranker=reranker)
for hit in rerank_search[0]:
    print(f"rerank_score={hit['distance']:.4f} scene={hit['entity']['scene']}")

Pattern 2 reranks only a set of already recalled frame URLs through the REST endpoint /v2/vectordb/ai/rerank. The following helper posts a JSON request; define it before the REST call.

from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen

def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
    request = Request(
        f"{MILVUS_URI.rstrip('/')}{path}",
        data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
        headers={"Authorization": f"Bearer {MILVUS_TOKEN}", "Content-Type": "application/json"},
        method="POST",
    )
    try:
        with urlopen(request, timeout=timeout) as response:
            return response.status, json.loads(response.read().decode("utf-8"))
    except HTTPError as exc:
        return exc.code, json.loads(exc.read().decode("utf-8"))

rest_docs = [f["frame_url"] for f in frames[:3]]
status, data = post_json("/v2/vectordb/ai/rerank", {
    "model_name": "qwen3-vl-rerank", "query": QUERY,
    "documents": rest_docs,
    "params": {"is_multimodal": True,
               "instruct": "Rank candidate frames by relevance to the query, "
                           "prioritizing subject, action, scene and fine-grained visual details.",
               "timeout_sec": 10}})
assert status == 200 and data.get("code") == 0, data
for item in sorted(data["data"]["output"]["results"],
                   key=lambda x: x["relevance_score"], reverse=True):
    print(f"index={item['index']} relevance_score={item['relevance_score']:.4f}")

The assert statement is an example-only check. In production, replace it with explicit error handling that inspects status and the response code, and log the response instead of raising.

Multimodal reranking requires is_multimodal in params, and you can use instruct to specify the ranking focus, such as prioritizing subject, action, scene, and fine-grained visual details. The REST documents parameter takes an array of frame URL strings and does not support top_n, so it returns one score for each candidate.

The following table shows reference reranking results for the test query "cut-in behavior at a highway ramp." Focus on the relative ordering, not on the absolute scores.

RankFrame sceneRerank score
1highway0.5898
2construction zone0.5018
3tunnel0.4858
4intersection0.4830
5intersection0.4089
6ordinary road0.3819

The highway frame ranked first, which matches the query semantics.

Step 7: Clean up resources

After you finish the tutorial, delete the resources you created to avoid unnecessary charges.

  1. Drop the collection:

    client.drop_collection(collection_name)
  2. Confirm the collection is deleted. The following command should return False:

    print(client.has_collection(collection_name))
  3. Delete the frame images you uploaded to OSS for the tutorial, if you no longer need them.

  4. (Optional) If you enabled public network access only for this tutorial, disable Public Network Access on the Security Configuration tab of the instance details page.

Troubleshooting

Use the following table to diagnose common issues. The constraints behind these symptoms are described in Limitations and considerations.

SymptomCauseSolution
A search immediately after a write returns empty results.flush() was not called after the write.Call flush() after writing, then search.
The connection times out.The port was omitted, so the request went to port 80.Specify port 19530 explicitly in the URI.
A combined filter returns 0 rows.No frame in the library matches all conditions.This can be a valid corner case mining conclusion, not a fault. Confirm the pipeline with a single condition, then add conditions gradually.
A batch write returns image batch size can should be [1, 10].The batch exceeds the multimodal limit of 10.Write frames in batches of 10 or fewer.

Solution value

The following table compares a traditional self-built solution with the Milvus AI Function approach.

DimensionTraditional self-built solutionMilvus AI Function
Development cycleJoint debugging across frame extraction, inference, vector database, and metadata databaseAttach functions when you create the collection, then write and search directly
Inference operationsSelf-built GPU inference cluster that needs scaling and fault recoveryInvocations managed by Milvus, with no inference cluster to operate
Data movementFrames move repeatedly between object storage, the inference cluster, and the vector databaseInference on write, and data stays within the Milvus instance
Multimodal retrievalRequires self-built alignment of the text and image vector spacesqwen3-vl-embedding natively supports text-to-frame and image-to-frame search
Annotation criteriaManual annotation is slow and expensive, and criteria varyprompt unifies the judgment rules, so label criteria are consistent and reproducible

For an autonomous driving team, this approach:

  • Speeds up corner case mining — Text-to-frame search plus structured filtering, combined with multimodal reranking, retrieves a long-tail scenario from a single natural language sentence.

  • Lowers auto-labeling cost — AI_CLASSIFY, AI_EXTRACT, and AI_ENTITY_EXTRACT produce labels with unified criteria at write time, leaving review and refinement to humans.

  • Converges the pipeline — Inference runs inside the vector database. Writes produce vectors and structured understanding, searches perform multimodal semantic recall, and driving data stays within the Milvus instance.