Batch inference submits JSONL-format text or multimodal Chat Completions requests as asynchronous batch tasks. Use it for long-running offline workloads such as overnight summarization, historical data labeling, model evaluation, and data annotation that can take minutes to hours.
How it works
Batch tasks follow the OpenAI Batch file input and result correlation pattern: each request uses a unique custom_id, and after completion, business data is written back from the output or error JSONL using that identifier.
-
Submit the task — Upload a UTF-8 JSONL file containing multiple requests, then use the returned file ID to create a batch task.
-
Asynchronous processing — The server validates and executes requests row by row in the background. The application queries statuses such as
validating,in_progress, andfinalizingusingbatch_id.
-
Download results — After the task reaches a terminal state (a state indicating the task has finished and will not change further), download
output. If there are failed rows, downloaderrorand usecustom_idto write back results or locate issues.
Batch inference is suitable for offline model evaluation, historical data annotation, content batch processing, and scheduled asset processing. It is not suitable for online requests such as chat or search completions that require immediate results.
Command format
Batch execution follows four steps: upload input file → create task → query status → download results.
The input file must be UTF-8 encoded JSONL. Each row is an independent request containing a unique custom_id, the fixed value method="POST", and url and body.model consistent with the task.
REST API
REST API
POST /v2/vectordb/ai/batch/files/upload
multipart/form-data: file, model_name, endpoint, purpose=batch
POST /v2/vectordb/ai/batch/jobs/create
{"input_file_id":"<file ID>","endpoint":"/v1/chat/completions","completion_window":"24h"}
POST /v2/vectordb/ai/batch/jobs/describe
{"batch_id":"<task ID>"}
POST /v2/vectordb/ai/batch/files/content
{"batch_id":"<task ID>","file_type":"output|error"}
Python
Python
upload_status, upload_data = post_multipart_upload("<input.jsonl>")
create_status, create_data = post_json(
"/v2/vectordb/ai/batch/jobs/create",
{
"input_file_id": upload_data["data"]["id"],
"endpoint": "<endpoint>",
"completion_window": "24h",
},
)
describe_status, describe_data = post_json(
"/v2/vectordb/ai/batch/jobs/describe",
{"batch_id": create_data["data"]["id"]},
)
content_status, content_data = post_json(
"/v2/vectordb/ai/batch/files/content",
{"batch_id": create_data["data"]["id"], "file_type": "output"},
)
The REST script depends on jq.
Parameters
| Parameter | Description |
file |
Required for upload. The JSONL input file. The server validates the first row's custom_id, method, url, and body.model. |
model_name |
Required for upload; also compatible with model. Must match the body.model in the first row of the JSONL. |
endpoint |
Required for upload and create. Use /v1/chat/completions for Chat Completions; use /v1/embeddings for text embedding tasks. All rows in the same file must use the same endpoint. |
purpose |
Optional for upload. Can only be empty or batch. |
input_file_id |
Required for creating a task. The file ID returned by the upload API. OSS URLs or external file identifiers are not accepted. |
completion_window |
Required for creating a task. The completion window; supports 24h to 336h or day units. |
metadata.ds_name/metadata.ds_description |
Optional. Task name can be up to 100 characters; description can be up to 200 characters. |
batch_id |
Required for query and download. The task ID returned by the create API. |
file_type |
Required for download. Use output to download successful rows; use error to download failed rows. |
provider |
Optional; defaults to aliyun_milvus. Upload, create, query, and download must use the same Alibaba Cloud account, region, and workspace. |
body.enable_thinking |
Optional, depending on the model. Must be at the same level as body.model in the JSONL row. For models with thinking enabled by default, explicitly setting this to false avoids unwanted thinking tokens. Do not place this parameter inside extra_body. |
Return values
After a task is successfully created, data.id is returned as the batch_id. Polling states include validating, in_progress, finalizing, and cancelling. Terminal states (states that indicate the task has finished and will not change further) include completed, failed, expired, and cancelled.
Batch tasks are processed asynchronously in a queue. The time from creation to a terminal state typically ranges from a few minutes to several hours, during which the status remains in_progress. You can call the query API POST /v2/vectordb/ai/batch/jobs/describe (with batch_id) at any time to get the latest status. If the example script's polling ends before the task completes, you can query again later using batch_id once the status changes to completed, then download the results. The download API returns the JSONL content directly; each row retains the input custom_id for writing results back to business data. Batch tasks are not suitable for online interactions that require immediate responses.
Example 1: Overnight batch knowledge base summary generation (text)
An operations team needs to batch-generate summaries for articles added that day. Each row contains one article request. Key parameters are body.model=qwen3.7-max, url=/v1/chat/completions, and completion_window=24h.
REST API
REST API
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
MODEL_NAME="qwen3.7-max"
INPUT_FILE="${AIFUNC_AI_BATCH_INPUT_FILE:-}"
POLL_INTERVAL_SEC="${AIFUNC_AI_BATCH_POLL_INTERVAL_SEC:-5}"
MAX_POLL_ATTEMPTS="${AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS:-120}"
if [ -z "$INPUT_FILE" ]; then
echo "Save the sample JSONL requests to a file first, then set AIFUNC_AI_BATCH_INPUT_FILE to the file path." >&2
exit 1
fi
download_batch_file() {
local file_type="$1"; local output_file="$2"; local content_body
content_body="$(jq -nc --arg batch_id "$BATCH_ID" --arg file_type "$file_type" '{provider:"aliyun_milvus", batch_id:$batch_id, file_type:$file_type}')"
curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/content" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" -H "Content-Type: application/json" \
-d "$content_body" --output "$output_file"
}
UPLOAD_RESPONSE="$(curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/upload" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-F "provider=aliyun_milvus" -F "model_name=$MODEL_NAME" -F "endpoint=/v1/chat/completions" \
-F "purpose=batch" -F "file=@$INPUT_FILE;type=application/jsonl")"
INPUT_FILE_ID="$(echo "$UPLOAD_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$INPUT_FILE_ID" ] || exit 1
CREATE_BODY="$(jq -nc --arg input_file_id "$INPUT_FILE_ID" '{provider:"aliyun_milvus", input_file_id:$input_file_id, endpoint:"/v1/chat/completions", completion_window:"24h", metadata:{ds_name:"milvus-ai-function-example", ds_description:"AI Batch REST example"}}')"
CREATE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/create" "$CREATE_BODY")"
BATCH_ID="$(echo "$CREATE_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$BATCH_ID" ] || exit 1
for ((attempt = 1; attempt <= MAX_POLL_ATTEMPTS; attempt++)); do
DESCRIBE_BODY="$(jq -nc --arg batch_id "$BATCH_ID" '{provider:"aliyun_milvus", batch_id:$batch_id}')"
DESCRIBE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/describe" "$DESCRIBE_BODY")"
BATCH_STATUS="$(echo "$DESCRIBE_RESPONSE" | jq -r '.data.status // empty')"
case "$BATCH_STATUS" in
completed)
OUTPUT_FILE="${AIFUNC_AI_BATCH_OUTPUT_FILE:-./ai_batch_output_${BATCH_ID}.jsonl}"
download_batch_file "output" "$OUTPUT_FILE"
echo "Batch output downloaded to: $OUTPUT_FILE"
exit 0 ;;
failed|expired|cancelled)
echo "Batch $BATCH_ID ended with status: $BATCH_STATUS" >&2; exit 1 ;;
validating|in_progress|finalizing|cancelling)
[ "$attempt" -lt "$MAX_POLL_ATTEMPTS" ] && sleep "$POLL_INTERVAL_SEC" ;;
*) echo "Unknown status: ${BATCH_STATUS:-empty}" >&2; exit 1 ;;
esac
done
echo "Batch $BATCH_ID did not finish after $MAX_POLL_ATTEMPTS checks." >&2
exit 1
Python
Python
from __future__ import annotations
import json
import os
import shutil
import sys
import tempfile
import time
import uuid
from pathlib import Path
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
MILVUS_REST_BASE_URL = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN = "<yourUsername>:<yourPassword>"
MODEL_NAME = "qwen3.7-max"
INPUT_MODE = os.getenv("AIFUNC_AI_BATCH_INPUT_MODE", "text").lower()
POLL_INTERVAL_SEC = int(os.getenv("AIFUNC_AI_BATCH_POLL_INTERVAL_SEC", "5"))
MAX_POLL_ATTEMPTS = int(os.getenv("AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS", "120"))
def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
try:
with urlopen(request, timeout=timeout) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
def batch_request(custom_id: str, content: Any, *, enable_thinking: bool | None = None) -> dict[str, Any]:
body: dict[str, Any] = {"model": MODEL_NAME, "messages": [{"role": "user", "content": content}]}
if enable_thinking is not None:
body["enable_thinking"] = enable_thinking
return {"custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": body}
def default_requests(input_mode: str) -> list[dict[str, Any]]:
if input_mode == "text":
return [
batch_request("milvus-summary-1", "For a support knowledge base, summarize in one English sentence: Milvus stores and searches vector embeddings for RAG, recommendation, and multimodal retrieval.", enable_thinking=False),
batch_request("milvus-summary-2", "Turn this incident note into an English FAQ title and one-sentence answer: after documents are updated, regenerate embeddings before users search the new content.", enable_thinking=False),
batch_request("milvus-summary-3", "Create a one-sentence English catalog description for an enterprise RAG case that retrieves policy passages before the assistant answers.", enable_thinking=False),
]
if input_mode == "image":
return [batch_request("milvus-image-1", [
{"type": "image_url", "image_url": {"url": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"}},
{"type": "text", "text": "For a pet-service media library, return a concise English accessibility caption and up to three searchable subject tags for this uploaded case photo."},
])]
if input_mode == "video":
return [batch_request("milvus-video-1", [
{"type": "video", "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"},
{"type": "text", "text": "For a marketing asset library, return a one-sentence English scene summary, three retrieval keywords, and whether manual brand-safety review is needed."},
])]
if input_mode == "audio":
return [batch_request("milvus-audio-1", [
{"type": "input_audio", "input_audio": {"data": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "format": "mp3"}},
{"type": "text", "text": "For a customer-service hotline greeting archive, transcribe the welcome message and assess whether its purpose, service availability, and recording or privacy notice are clear; mark details not present in the audio as unconfirmed."},
])]
print("AIFUNC_AI_BATCH_INPUT_MODE must be one of text, image, video, or audio", file=sys.stderr)
sys.exit(1)
def post_multipart_upload(input_file: str) -> tuple[int, dict[str, Any]]:
boundary = f"----milvus-ai-batch-{uuid.uuid4().hex}"
fields = {"provider": "aliyun_milvus", "model_name": MODEL_NAME, "endpoint": "/v1/chat/completions", "purpose": "batch"}
with tempfile.TemporaryFile(mode="w+b") as payload:
for name, value in fields.items():
payload.write(f"--{boundary}\r\n".encode())
payload.write(f'Content-Disposition: form-data; name="{name}"\r\n\r\n'.encode())
payload.write(value.encode())
payload.write(b"\r\n")
payload.write(f"--{boundary}\r\n".encode())
payload.write(b'Content-Disposition: form-data; name="file"; filename="input.jsonl"\r\n')
payload.write(b"Content-Type: application/jsonl\r\n\r\n")
with open(input_file, "rb") as input_handle:
shutil.copyfileobj(input_handle, payload)
payload.write(b"\r\n")
payload.write(f"--{boundary}--\r\n".encode())
payload_length = payload.tell()
payload.seek(0)
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/upload",
data=payload,
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": f"multipart/form-data; boundary={boundary}", "Content-Length": str(payload_length)},
method="POST",
)
try:
with urlopen(request, timeout=600) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
def download_batch_file(batch_id: str, file_type: str, output_file: Path) -> None:
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/content",
data=json.dumps({"provider": "aliyun_milvus", "batch_id": batch_id, "file_type": file_type}, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
with urlopen(request, timeout=600) as response, output_file.open("wb") as output:
shutil.copyfileobj(response, output)
temporary_input = None
INPUT_FILE = os.getenv("AIFUNC_AI_BATCH_INPUT_FILE")
if INPUT_FILE is None:
temporary_input = tempfile.NamedTemporaryFile(mode="w", suffix=".jsonl", delete=False, encoding="utf-8")
for request in default_requests(INPUT_MODE):
temporary_input.write(json.dumps(request, ensure_ascii=False, separators=(",", ":")) + "\n")
temporary_input.close()
INPUT_FILE = temporary_input.name
try:
status, data = post_multipart_upload(INPUT_FILE)
if status != 200 or data.get("code") != 0:
sys.exit(1)
input_file_id = data.get("data", {}).get("id")
status, data = post_json("/v2/vectordb/ai/batch/jobs/create", {
"provider": "aliyun_milvus", "input_file_id": input_file_id, "endpoint": "/v1/chat/completions",
"completion_window": "24h", "metadata": {"ds_name": "milvus-ai-function-example", "ds_description": "AI Batch example"},
})
if status != 200 or data.get("code") != 0:
sys.exit(1)
batch_id = data.get("data", {}).get("id")
for attempt in range(1, MAX_POLL_ATTEMPTS + 1):
status, data = post_json("/v2/vectordb/ai/batch/jobs/describe", {"provider": "aliyun_milvus", "batch_id": batch_id})
batch_data = data.get("data", {})
batch_status = batch_data.get("status")
if batch_status == "completed":
output_file = Path(os.getenv("AIFUNC_AI_BATCH_OUTPUT_FILE", f"ai_batch_output_{batch_id}.jsonl"))
download_batch_file(batch_id, "output", output_file)
print(f"Batch output downloaded to: {output_file}")
if batch_data.get("error_file_id"):
download_batch_file(batch_id, "error", Path(f"ai_batch_errors_{batch_id}.jsonl"))
break
if batch_status in {"failed", "expired", "cancelled"}:
print(f"Batch {batch_id} ended with status: {batch_status}", file=sys.stderr)
sys.exit(1)
if attempt < MAX_POLL_ATTEMPTS:
time.sleep(POLL_INTERVAL_SEC)
finally:
if temporary_input is not None:
os.unlink(temporary_input.name)
Expected result: after the task reaches completed, each row in the downloaded output.jsonl retains its custom_id. For example:
{"custom_id":"milvus-summary-1","response":{"status_code":200,"body":{"choices":[{"message":{"role":"assistant","content":"Milvus is a vector database that serves RAG, recommendation, and multimodal retrieval scenarios."}}]}}}
The application side writes back the summaries to the corresponding articles using custom_id. If an error.jsonl is produced, retry only the failed articles.
Example 2: Batch pet service case image description generation (image)
A pet service platform needs to generate accessibility descriptions and retrieval tags for historical case images uploaded by merchants. Each row's content includes an image_url and a text instruction. Use a multimodal model that supports image understanding; the model_name specified at upload must match.
REST API
REST API
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
MODEL_NAME="qwen3-vl-plus"
INPUT_FILE="${AIFUNC_AI_BATCH_INPUT_FILE:-}"
POLL_INTERVAL_SEC="${AIFUNC_AI_BATCH_POLL_INTERVAL_SEC:-5}"
MAX_POLL_ATTEMPTS="${AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS:-120}"
if [ -z "$INPUT_FILE" ]; then
echo "Save the sample JSONL requests to a file first, then set AIFUNC_AI_BATCH_INPUT_FILE to the file path." >&2
exit 1
fi
download_batch_file() {
local file_type="$1"; local output_file="$2"; local content_body
content_body="$(jq -nc --arg batch_id "$BATCH_ID" --arg file_type "$file_type" '{provider:"aliyun_milvus", batch_id:$batch_id, file_type:$file_type}')"
curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/content" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" -H "Content-Type: application/json" \
-d "$content_body" --output "$output_file"
}
UPLOAD_RESPONSE="$(curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/upload" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-F "provider=aliyun_milvus" -F "model_name=$MODEL_NAME" -F "endpoint=/v1/chat/completions" \
-F "purpose=batch" -F "file=@$INPUT_FILE;type=application/jsonl")"
INPUT_FILE_ID="$(echo "$UPLOAD_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$INPUT_FILE_ID" ] || exit 1
CREATE_BODY="$(jq -nc --arg input_file_id "$INPUT_FILE_ID" '{provider:"aliyun_milvus", input_file_id:$input_file_id, endpoint:"/v1/chat/completions", completion_window:"24h", metadata:{ds_name:"milvus-ai-function-example", ds_description:"AI Batch REST example"}}')"
CREATE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/create" "$CREATE_BODY")"
BATCH_ID="$(echo "$CREATE_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$BATCH_ID" ] || exit 1
for ((attempt = 1; attempt <= MAX_POLL_ATTEMPTS; attempt++)); do
DESCRIBE_BODY="$(jq -nc --arg batch_id "$BATCH_ID" '{provider:"aliyun_milvus", batch_id:$batch_id}')"
DESCRIBE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/describe" "$DESCRIBE_BODY")"
BATCH_STATUS="$(echo "$DESCRIBE_RESPONSE" | jq -r '.data.status // empty')"
case "$BATCH_STATUS" in
completed)
OUTPUT_FILE="${AIFUNC_AI_BATCH_OUTPUT_FILE:-./ai_batch_output_${BATCH_ID}.jsonl}"
download_batch_file "output" "$OUTPUT_FILE"
echo "Batch output downloaded to: $OUTPUT_FILE"
exit 0 ;;
failed|expired|cancelled)
echo "Batch $BATCH_ID ended with status: $BATCH_STATUS" >&2; exit 1 ;;
validating|in_progress|finalizing|cancelling)
[ "$attempt" -lt "$MAX_POLL_ATTEMPTS" ] && sleep "$POLL_INTERVAL_SEC" ;;
*) echo "Unknown status: ${BATCH_STATUS:-empty}" >&2; exit 1 ;;
esac
done
echo "Batch $BATCH_ID did not finish after $MAX_POLL_ATTEMPTS checks." >&2
exit 1
Python
Python
from __future__ import annotations
import json
import os
import shutil
import sys
import tempfile
import time
import uuid
from pathlib import Path
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
MILVUS_REST_BASE_URL = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN = "<yourUsername>:<yourPassword>"
MODEL_NAME = "qwen3-vl-plus"
INPUT_MODE = os.getenv("AIFUNC_AI_BATCH_INPUT_MODE", "image").lower()
POLL_INTERVAL_SEC = int(os.getenv("AIFUNC_AI_BATCH_POLL_INTERVAL_SEC", "5"))
MAX_POLL_ATTEMPTS = int(os.getenv("AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS", "120"))
def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
try:
with urlopen(request, timeout=timeout) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
def batch_request(custom_id: str, content: Any, *, enable_thinking: bool | None = None) -> dict[str, Any]:
body: dict[str, Any] = {"model": MODEL_NAME, "messages": [{"role": "user", "content": content}]}
if enable_thinking is not None:
body["enable_thinking"] = enable_thinking
return {"custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": body}
def default_requests(input_mode: str) -> list[dict[str, Any]]:
if input_mode == "text":
return [
batch_request("milvus-summary-1", "For a support knowledge base, summarize in one English sentence: Milvus stores and searches vector embeddings for RAG, recommendation, and multimodal retrieval.", enable_thinking=False),
batch_request("milvus-summary-2", "Turn this incident note into an English FAQ title and one-sentence answer: after documents are updated, regenerate embeddings before users search the new content.", enable_thinking=False),
batch_request("milvus-summary-3", "Create a one-sentence English catalog description for an enterprise RAG case that retrieves policy passages before the assistant answers.", enable_thinking=False),
]
if input_mode == "image":
return [batch_request("milvus-image-1", [
{"type": "image_url", "image_url": {"url": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"}},
{"type": "text", "text": "For a pet-service media library, return a concise English accessibility caption and up to three searchable subject tags for this uploaded case photo."},
])]
if input_mode == "video":
return [batch_request("milvus-video-1", [
{"type": "video", "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"},
{"type": "text", "text": "For a marketing asset library, return a one-sentence English scene summary, three retrieval keywords, and whether manual brand-safety review is needed."},
])]
if input_mode == "audio":
return [batch_request("milvus-audio-1", [
{"type": "input_audio", "input_audio": {"data": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "format": "mp3"}},
{"type": "text", "text": "For a customer-service hotline greeting archive, transcribe the welcome message and assess whether its purpose, service availability, and recording or privacy notice are clear; mark details not present in the audio as unconfirmed."},
])]
print("AIFUNC_AI_BATCH_INPUT_MODE must be one of text, image, video, or audio", file=sys.stderr)
sys.exit(1)
def post_multipart_upload(input_file: str) -> tuple[int, dict[str, Any]]:
boundary = f"----milvus-ai-batch-{uuid.uuid4().hex}"
fields = {"provider": "aliyun_milvus", "model_name": MODEL_NAME, "endpoint": "/v1/chat/completions", "purpose": "batch"}
with tempfile.TemporaryFile(mode="w+b") as payload:
for name, value in fields.items():
payload.write(f"--{boundary}\r\n".encode())
payload.write(f'Content-Disposition: form-data; name="{name}"\r\n\r\n'.encode())
payload.write(value.encode())
payload.write(b"\r\n")
payload.write(f"--{boundary}\r\n".encode())
payload.write(b'Content-Disposition: form-data; name="file"; filename="input.jsonl"\r\n')
payload.write(b"Content-Type: application/jsonl\r\n\r\n")
with open(input_file, "rb") as input_handle:
shutil.copyfileobj(input_handle, payload)
payload.write(b"\r\n")
payload.write(f"--{boundary}--\r\n".encode())
payload_length = payload.tell()
payload.seek(0)
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/upload",
data=payload,
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": f"multipart/form-data; boundary={boundary}", "Content-Length": str(payload_length)},
method="POST",
)
try:
with urlopen(request, timeout=600) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
def download_batch_file(batch_id: str, file_type: str, output_file: Path) -> None:
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/content",
data=json.dumps({"provider": "aliyun_milvus", "batch_id": batch_id, "file_type": file_type}, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
with urlopen(request, timeout=600) as response, output_file.open("wb") as output:
shutil.copyfileobj(response, output)
temporary_input = None
INPUT_FILE = os.getenv("AIFUNC_AI_BATCH_INPUT_FILE")
if INPUT_FILE is None:
temporary_input = tempfile.NamedTemporaryFile(mode="w", suffix=".jsonl", delete=False, encoding="utf-8")
for request in default_requests(INPUT_MODE):
temporary_input.write(json.dumps(request, ensure_ascii=False, separators=(",", ":")) + "\n")
temporary_input.close()
INPUT_FILE = temporary_input.name
try:
status, data = post_multipart_upload(INPUT_FILE)
if status != 200 or data.get("code") != 0:
sys.exit(1)
input_file_id = data.get("data", {}).get("id")
status, data = post_json("/v2/vectordb/ai/batch/jobs/create", {
"provider": "aliyun_milvus", "input_file_id": input_file_id, "endpoint": "/v1/chat/completions",
"completion_window": "24h", "metadata": {"ds_name": "milvus-ai-function-example", "ds_description": "AI Batch example"},
})
if status != 200 or data.get("code") != 0:
sys.exit(1)
batch_id = data.get("data", {}).get("id")
for attempt in range(1, MAX_POLL_ATTEMPTS + 1):
status, data = post_json("/v2/vectordb/ai/batch/jobs/describe", {"provider": "aliyun_milvus", "batch_id": batch_id})
batch_data = data.get("data", {})
batch_status = batch_data.get("status")
if batch_status == "completed":
output_file = Path(os.getenv("AIFUNC_AI_BATCH_OUTPUT_FILE", f"ai_batch_output_{batch_id}.jsonl"))
download_batch_file(batch_id, "output", output_file)
print(f"Batch output downloaded to: {output_file}")
if batch_data.get("error_file_id"):
download_batch_file(batch_id, "error", Path(f"ai_batch_errors_{batch_id}.jsonl"))
break
if batch_status in {"failed", "expired", "cancelled"}:
print(f"Batch {batch_id} ended with status: {batch_status}", file=sys.stderr)
sys.exit(1)
if attempt < MAX_POLL_ATTEMPTS:
time.sleep(POLL_INTERVAL_SEC)
finally:
if temporary_input is not None:
os.unlink(temporary_input.name)
Expected result: image-output.jsonl returns the image description and tags using the same custom_id=milvus-image-1. The application side can write the text into case library fields and combine it with image vectors for image-text retrieval. If the task completes but has failed rows, download the file_type=error result file.
Example 3: Batch video material summary generation (video)
An operations team needs to generate a one-sentence summary for historical short videos to support video asset retrieval. Each row passes a video URL through a video content block. Use a multimodal model that supports video understanding.
REST API
REST API
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
MODEL_NAME="qwen3-vl-plus"
INPUT_FILE="${AIFUNC_AI_BATCH_INPUT_FILE:-}"
POLL_INTERVAL_SEC="${AIFUNC_AI_BATCH_POLL_INTERVAL_SEC:-5}"
MAX_POLL_ATTEMPTS="${AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS:-120}"
if [ -z "$INPUT_FILE" ]; then
echo "Save the sample JSONL requests to a file first, then set AIFUNC_AI_BATCH_INPUT_FILE to the file path." >&2
exit 1
fi
download_batch_file() {
local file_type="$1"; local output_file="$2"; local content_body
content_body="$(jq -nc --arg batch_id "$BATCH_ID" --arg file_type "$file_type" '{provider:"aliyun_milvus", batch_id:$batch_id, file_type:$file_type}')"
curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/content" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" -H "Content-Type: application/json" \
-d "$content_body" --output "$output_file"
}
UPLOAD_RESPONSE="$(curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/upload" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-F "provider=aliyun_milvus" -F "model_name=$MODEL_NAME" -F "endpoint=/v1/chat/completions" \
-F "purpose=batch" -F "file=@$INPUT_FILE;type=application/jsonl")"
INPUT_FILE_ID="$(echo "$UPLOAD_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$INPUT_FILE_ID" ] || exit 1
CREATE_BODY="$(jq -nc --arg input_file_id "$INPUT_FILE_ID" '{provider:"aliyun_milvus", input_file_id:$input_file_id, endpoint:"/v1/chat/completions", completion_window:"24h", metadata:{ds_name:"milvus-ai-function-example", ds_description:"AI Batch REST example"}}')"
CREATE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/create" "$CREATE_BODY")"
BATCH_ID="$(echo "$CREATE_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$BATCH_ID" ] || exit 1
for ((attempt = 1; attempt <= MAX_POLL_ATTEMPTS; attempt++)); do
DESCRIBE_BODY="$(jq -nc --arg batch_id "$BATCH_ID" '{provider:"aliyun_milvus", batch_id:$batch_id}')"
DESCRIBE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/describe" "$DESCRIBE_BODY")"
BATCH_STATUS="$(echo "$DESCRIBE_RESPONSE" | jq -r '.data.status // empty')"
case "$BATCH_STATUS" in
completed)
OUTPUT_FILE="${AIFUNC_AI_BATCH_OUTPUT_FILE:-./ai_batch_output_${BATCH_ID}.jsonl}"
download_batch_file "output" "$OUTPUT_FILE"
echo "Batch output downloaded to: $OUTPUT_FILE"
exit 0 ;;
failed|expired|cancelled)
echo "Batch $BATCH_ID ended with status: $BATCH_STATUS" >&2; exit 1 ;;
validating|in_progress|finalizing|cancelling)
[ "$attempt" -lt "$MAX_POLL_ATTEMPTS" ] && sleep "$POLL_INTERVAL_SEC" ;;
*) echo "Unknown status: ${BATCH_STATUS:-empty}" >&2; exit 1 ;;
esac
done
echo "Batch $BATCH_ID did not finish after $MAX_POLL_ATTEMPTS checks." >&2
exit 1
Python
Python
from __future__ import annotations
import json
import os
import shutil
import sys
import tempfile
import time
import uuid
from pathlib import Path
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
MILVUS_REST_BASE_URL = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN = "<yourUsername>:<yourPassword>"
MODEL_NAME = "qwen3-vl-plus"
INPUT_MODE = os.getenv("AIFUNC_AI_BATCH_INPUT_MODE", "video").lower()
POLL_INTERVAL_SEC = int(os.getenv("AIFUNC_AI_BATCH_POLL_INTERVAL_SEC", "5"))
MAX_POLL_ATTEMPTS = int(os.getenv("AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS", "120"))
def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
try:
with urlopen(request, timeout=timeout) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
def batch_request(custom_id: str, content: Any, *, enable_thinking: bool | None = None) -> dict[str, Any]:
body: dict[str, Any] = {"model": MODEL_NAME, "messages": [{"role": "user", "content": content}]}
if enable_thinking is not None:
body["enable_thinking"] = enable_thinking
return {"custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": body}
def default_requests(input_mode: str) -> list[dict[str, Any]]:
if input_mode == "text":
return [
batch_request("milvus-summary-1", "For a support knowledge base, summarize in one English sentence: Milvus stores and searches vector embeddings for RAG, recommendation, and multimodal retrieval.", enable_thinking=False),
batch_request("milvus-summary-2", "Turn this incident note into an English FAQ title and one-sentence answer: after documents are updated, regenerate embeddings before users search the new content.", enable_thinking=False),
batch_request("milvus-summary-3", "Create a one-sentence English catalog description for an enterprise RAG case that retrieves policy passages before the assistant answers.", enable_thinking=False),
]
if input_mode == "image":
return [batch_request("milvus-image-1", [
{"type": "image_url", "image_url": {"url": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"}},
{"type": "text", "text": "For a pet-service media library, return a concise English accessibility caption and up to three searchable subject tags for this uploaded case photo."},
])]
if input_mode == "video":
return [batch_request("milvus-video-1", [
{"type": "video", "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"},
{"type": "text", "text": "For a marketing asset library, return a one-sentence English scene summary, three retrieval keywords, and whether manual brand-safety review is needed."},
])]
if input_mode == "audio":
return [batch_request("milvus-audio-1", [
{"type": "input_audio", "input_audio": {"data": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "format": "mp3"}},
{"type": "text", "text": "For a customer-service hotline greeting archive, transcribe the welcome message and assess whether its purpose, service availability, and recording or privacy notice are clear; mark details not present in the audio as unconfirmed."},
])]
print("AIFUNC_AI_BATCH_INPUT_MODE must be one of text, image, video, or audio", file=sys.stderr)
sys.exit(1)
def post_multipart_upload(input_file: str) -> tuple[int, dict[str, Any]]:
boundary = f"----milvus-ai-batch-{uuid.uuid4().hex}"
fields = {"provider": "aliyun_milvus", "model_name": MODEL_NAME, "endpoint": "/v1/chat/completions", "purpose": "batch"}
with tempfile.TemporaryFile(mode="w+b") as payload:
for name, value in fields.items():
payload.write(f"--{boundary}\r\n".encode())
payload.write(f'Content-Disposition: form-data; name="{name}"\r\n\r\n'.encode())
payload.write(value.encode())
payload.write(b"\r\n")
payload.write(f"--{boundary}\r\n".encode())
payload.write(b'Content-Disposition: form-data; name="file"; filename="input.jsonl"\r\n')
payload.write(b"Content-Type: application/jsonl\r\n\r\n")
with open(input_file, "rb") as input_handle:
shutil.copyfileobj(input_handle, payload)
payload.write(b"\r\n")
payload.write(f"--{boundary}--\r\n".encode())
payload_length = payload.tell()
payload.seek(0)
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/upload",
data=payload,
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": f"multipart/form-data; boundary={boundary}", "Content-Length": str(payload_length)},
method="POST",
)
try:
with urlopen(request, timeout=600) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
def download_batch_file(batch_id: str, file_type: str, output_file: Path) -> None:
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/content",
data=json.dumps({"provider": "aliyun_milvus", "batch_id": batch_id, "file_type": file_type}, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
with urlopen(request, timeout=600) as response, output_file.open("wb") as output:
shutil.copyfileobj(response, output)
temporary_input = None
INPUT_FILE = os.getenv("AIFUNC_AI_BATCH_INPUT_FILE")
if INPUT_FILE is None:
temporary_input = tempfile.NamedTemporaryFile(mode="w", suffix=".jsonl", delete=False, encoding="utf-8")
for request in default_requests(INPUT_MODE):
temporary_input.write(json.dumps(request, ensure_ascii=False, separators=(",", ":")) + "\n")
temporary_input.close()
INPUT_FILE = temporary_input.name
try:
status, data = post_multipart_upload(INPUT_FILE)
if status != 200 or data.get("code") != 0:
sys.exit(1)
input_file_id = data.get("data", {}).get("id")
status, data = post_json("/v2/vectordb/ai/batch/jobs/create", {
"provider": "aliyun_milvus", "input_file_id": input_file_id, "endpoint": "/v1/chat/completions",
"completion_window": "24h", "metadata": {"ds_name": "milvus-ai-function-example", "ds_description": "AI Batch example"},
})
if status != 200 or data.get("code") != 0:
sys.exit(1)
batch_id = data.get("data", {}).get("id")
for attempt in range(1, MAX_POLL_ATTEMPTS + 1):
status, data = post_json("/v2/vectordb/ai/batch/jobs/describe", {"provider": "aliyun_milvus", "batch_id": batch_id})
batch_data = data.get("data", {})
batch_status = batch_data.get("status")
if batch_status == "completed":
output_file = Path(os.getenv("AIFUNC_AI_BATCH_OUTPUT_FILE", f"ai_batch_output_{batch_id}.jsonl"))
download_batch_file(batch_id, "output", output_file)
print(f"Batch output downloaded to: {output_file}")
if batch_data.get("error_file_id"):
download_batch_file(batch_id, "error", Path(f"ai_batch_errors_{batch_id}.jsonl"))
break
if batch_status in {"failed", "expired", "cancelled"}:
print(f"Batch {batch_id} ended with status: {batch_status}", file=sys.stderr)
sys.exit(1)
if attempt < MAX_POLL_ATTEMPTS:
time.sleep(POLL_INTERVAL_SEC)
finally:
if temporary_input is not None:
os.unlink(temporary_input.name)
Expected result: video-output.jsonl returns the video summary with custom_id=milvus-video-1. After writing the summary into the video asset metadata, you can combine video vectors and keywords for retrieval. If the terminal state includes an error_file_id, retain the error file to locate failed tasks.
Example 4: Overnight archiving of customer service hotline greetings (audio)
A customer service operations team periodically archives the greetings of each service hotline, transcribing the content and checking whether the hotline purpose, service hours, and recording or privacy notices are clear, then writing the results back to the configuration center using the greeting version ID. Each row passes an MP3 greeting through input_audio using a model that supports audio understanding, with a unique custom_id to correlate the original version.
Audio and transcripts may contain personal information or internal service configuration. Confirm collection authorization and desensitization requirements before running. Retain output only for the period needed to complete write-back and spot-checks, then delete it securely according to your organization's data retention policy. Do not commit production recordings or results to a code repository.
REST API
REST API
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
MODEL_NAME="qwen3.5-omni-plus"
INPUT_FILE="${AIFUNC_AI_BATCH_INPUT_FILE:-}"
POLL_INTERVAL_SEC="${AIFUNC_AI_BATCH_POLL_INTERVAL_SEC:-5}"
MAX_POLL_ATTEMPTS="${AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS:-120}"
if [ -z "$INPUT_FILE" ]; then
echo "Save the sample JSONL requests to a file first, then set AIFUNC_AI_BATCH_INPUT_FILE to the file path." >&2
exit 1
fi
download_batch_file() {
local file_type="$1"; local output_file="$2"; local content_body
content_body="$(jq -nc --arg batch_id "$BATCH_ID" --arg file_type "$file_type" '{provider:"aliyun_milvus", batch_id:$batch_id, file_type:$file_type}')"
curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/content" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" -H "Content-Type: application/json" \
-d "$content_body" --output "$output_file"
}
UPLOAD_RESPONSE="$(curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/upload" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-F "provider=aliyun_milvus" -F "model_name=$MODEL_NAME" -F "endpoint=/v1/chat/completions" \
-F "purpose=batch" -F "file=@$INPUT_FILE;type=application/jsonl")"
INPUT_FILE_ID="$(echo "$UPLOAD_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$INPUT_FILE_ID" ] || exit 1
CREATE_BODY="$(jq -nc --arg input_file_id "$INPUT_FILE_ID" '{provider:"aliyun_milvus", input_file_id:$input_file_id, endpoint:"/v1/chat/completions", completion_window:"24h", metadata:{ds_name:"milvus-ai-function-example", ds_description:"AI Batch REST example"}}')"
CREATE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/create" "$CREATE_BODY")"
BATCH_ID="$(echo "$CREATE_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$BATCH_ID" ] || exit 1
for ((attempt = 1; attempt <= MAX_POLL_ATTEMPTS; attempt++)); do
DESCRIBE_BODY="$(jq -nc --arg batch_id "$BATCH_ID" '{provider:"aliyun_milvus", batch_id:$batch_id}')"
DESCRIBE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/describe" "$DESCRIBE_BODY")"
BATCH_STATUS="$(echo "$DESCRIBE_RESPONSE" | jq -r '.data.status // empty')"
case "$BATCH_STATUS" in
completed)
OUTPUT_FILE="${AIFUNC_AI_BATCH_OUTPUT_FILE:-./ai_batch_output_${BATCH_ID}.jsonl}"
download_batch_file "output" "$OUTPUT_FILE"
echo "Batch output downloaded to: $OUTPUT_FILE"
exit 0 ;;
failed|expired|cancelled)
echo "Batch $BATCH_ID ended with status: $BATCH_STATUS" >&2; exit 1 ;;
validating|in_progress|finalizing|cancelling)
[ "$attempt" -lt "$MAX_POLL_ATTEMPTS" ] && sleep "$POLL_INTERVAL_SEC" ;;
*) echo "Unknown status: ${BATCH_STATUS:-empty}" >&2; exit 1 ;;
esac
done
echo "Batch $BATCH_ID did not finish after $MAX_POLL_ATTEMPTS checks." >&2
exit 1
Python
Python
from __future__ import annotations
import json
import os
import shutil
import sys
import tempfile
import time
import uuid
from pathlib import Path
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
MILVUS_REST_BASE_URL = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN = "<yourUsername>:<yourPassword>"
MODEL_NAME = "qwen3.5-omni-plus"
INPUT_MODE = os.getenv("AIFUNC_AI_BATCH_INPUT_MODE", "audio").lower()
POLL_INTERVAL_SEC = int(os.getenv("AIFUNC_AI_BATCH_POLL_INTERVAL_SEC", "5"))
MAX_POLL_ATTEMPTS = int(os.getenv("AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS", "120"))
def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
try:
with urlopen(request, timeout=timeout) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
def batch_request(custom_id: str, content: Any, *, enable_thinking: bool | None = None) -> dict[str, Any]:
body: dict[str, Any] = {"model": MODEL_NAME, "messages": [{"role": "user", "content": content}]}
if enable_thinking is not None:
body["enable_thinking"] = enable_thinking
return {"custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": body}
def default_requests(input_mode: str) -> list[dict[str, Any]]:
if input_mode == "text":
return [
batch_request("milvus-summary-1", "For a support knowledge base, summarize in one English sentence: Milvus stores and searches vector embeddings for RAG, recommendation, and multimodal retrieval.", enable_thinking=False),
batch_request("milvus-summary-2", "Turn this incident note into an English FAQ title and one-sentence answer: after documents are updated, regenerate embeddings before users search the new content.", enable_thinking=False),
batch_request("milvus-summary-3", "Create a one-sentence English catalog description for an enterprise RAG case that retrieves policy passages before the assistant answers.", enable_thinking=False),
]
if input_mode == "image":
return [batch_request("milvus-image-1", [
{"type": "image_url", "image_url": {"url": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"}},
{"type": "text", "text": "For a pet-service media library, return a concise English accessibility caption and up to three searchable subject tags for this uploaded case photo."},
])]
if input_mode == "video":
return [batch_request("milvus-video-1", [
{"type": "video", "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"},
{"type": "text", "text": "For a marketing asset library, return a one-sentence English scene summary, three retrieval keywords, and whether manual brand-safety review is needed."},
])]
if input_mode == "audio":
return [batch_request("milvus-audio-1", [
{"type": "input_audio", "input_audio": {"data": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "format": "mp3"}},
{"type": "text", "text": "For a customer-service hotline greeting archive, transcribe the welcome message and assess whether its purpose, service availability, and recording or privacy notice are clear; mark details not present in the audio as unconfirmed."},
])]
print("AIFUNC_AI_BATCH_INPUT_MODE must be one of text, image, video, or audio", file=sys.stderr)
sys.exit(1)
def post_multipart_upload(input_file: str) -> tuple[int, dict[str, Any]]:
boundary = f"----milvus-ai-batch-{uuid.uuid4().hex}"
fields = {"provider": "aliyun_milvus", "model_name": MODEL_NAME, "endpoint": "/v1/chat/completions", "purpose": "batch"}
with tempfile.TemporaryFile(mode="w+b") as payload:
for name, value in fields.items():
payload.write(f"--{boundary}\r\n".encode())
payload.write(f'Content-Disposition: form-data; name="{name}"\r\n\r\n'.encode())
payload.write(value.encode())
payload.write(b"\r\n")
payload.write(f"--{boundary}\r\n".encode())
payload.write(b'Content-Disposition: form-data; name="file"; filename="input.jsonl"\r\n')
payload.write(b"Content-Type: application/jsonl\r\n\r\n")
with open(input_file, "rb") as input_handle:
shutil.copyfileobj(input_handle, payload)
payload.write(b"\r\n")
payload.write(f"--{boundary}--\r\n".encode())
payload_length = payload.tell()
payload.seek(0)
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/upload",
data=payload,
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": f"multipart/form-data; boundary={boundary}", "Content-Length": str(payload_length)},
method="POST",
)
try:
with urlopen(request, timeout=600) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
def download_batch_file(batch_id: str, file_type: str, output_file: Path) -> None:
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/content",
data=json.dumps({"provider": "aliyun_milvus", "batch_id": batch_id, "file_type": file_type}, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
with urlopen(request, timeout=600) as response, output_file.open("wb") as output:
shutil.copyfileobj(response, output)
temporary_input = None
INPUT_FILE = os.getenv("AIFUNC_AI_BATCH_INPUT_FILE")
if INPUT_FILE is None:
temporary_input = tempfile.NamedTemporaryFile(mode="w", suffix=".jsonl", delete=False, encoding="utf-8")
for request in default_requests(INPUT_MODE):
temporary_input.write(json.dumps(request, ensure_ascii=False, separators=(",", ":")) + "\n")
temporary_input.close()
INPUT_FILE = temporary_input.name
try:
status, data = post_multipart_upload(INPUT_FILE)
if status != 200 or data.get("code") != 0:
sys.exit(1)
input_file_id = data.get("data", {}).get("id")
status, data = post_json("/v2/vectordb/ai/batch/jobs/create", {
"provider": "aliyun_milvus", "input_file_id": input_file_id, "endpoint": "/v1/chat/completions",
"completion_window": "24h", "metadata": {"ds_name": "milvus-ai-function-example", "ds_description": "AI Batch example"},
})
if status != 200 or data.get("code") != 0:
sys.exit(1)
batch_id = data.get("data", {}).get("id")
for attempt in range(1, MAX_POLL_ATTEMPTS + 1):
status, data = post_json("/v2/vectordb/ai/batch/jobs/describe", {"provider": "aliyun_milvus", "batch_id": batch_id})
batch_data = data.get("data", {})
batch_status = batch_data.get("status")
if batch_status == "completed":
output_file = Path(os.getenv("AIFUNC_AI_BATCH_OUTPUT_FILE", f"ai_batch_output_{batch_id}.jsonl"))
download_batch_file(batch_id, "output", output_file)
print(f"Batch output downloaded to: {output_file}")
if batch_data.get("error_file_id"):
download_batch_file(batch_id, "error", Path(f"ai_batch_errors_{batch_id}.jsonl"))
break
if batch_status in {"failed", "expired", "cancelled"}:
print(f"Batch {batch_id} ended with status: {batch_status}", file=sys.stderr)
sys.exit(1)
if attempt < MAX_POLL_ATTEMPTS:
time.sleep(POLL_INTERVAL_SEC)
finally:
if temporary_input is not None:
os.unlink(temporary_input.name)
Expected result: call-greeting-output.jsonl or call-greeting-error.jsonl contains exactly one row with custom_id=milvus-audio-1. After limited-scope verification, successful results are written to the configuration center for spot-checking by customer service operations staff. Failed rows retain the original custom_id and can be retried individually after fixing audio permissions, MP3 format, or model parameters.