全部產品
Search
文件中心

Vector Retrieval Service for Milvus:批量推理

更新時間:Aug 04, 2026

AI_BATCH 支援將 JSONL格式的文本或多模態 Chat Completions請求作為非同步批量任務提交,適用於夜間摘要、歷史資料打標、模型評測及資料標註等分鐘至小時級的長耗時處理情境。

工作原理

批量任務遵循 OpenAI Batch 的檔案輸入和結果關聯方式:每行請求使用唯一 custom_id,完成後從輸出或錯誤 JSONL 中按該標識回寫業務資料。

  1. 提交任務:上傳包含多條請求的 UTF-8 JSONL 檔案,再使用返回的檔案 ID 建立 Batch 任務。

  2. 非同步處理:服務端在後台校正並逐行執行請求;業務側通過 batch_id 查詢 validating、in_progress、finalizing 等狀態。

  3. 下載結果:任務進入終態後下載 output;如有失敗行,再下載 error,並按 custom_id 回寫結果或定位問題。

適合離線模型評測、歷史資料標註、內容批處理和定時素材加工;不適合聊天、搜尋補全等需要立即展示結果的線上請求。

命令格式

Batch 按“上傳輸入檔案 → 建立任務 → 查詢狀態 → 下載結果”四步執行:

REST 介面

POST /v2/vectordb/ai/batch/files/upload
multipart/form-data: file, model_name, endpoint, purpose=batch

POST /v2/vectordb/ai/batch/jobs/create
{"input_file_id":"<檔案ID>","endpoint":"/v1/chat/completions","completion_window":"24h"}

POST /v2/vectordb/ai/batch/jobs/describe
{"batch_id":"<任務ID>"}

POST /v2/vectordb/ai/batch/files/content
{"batch_id":"<任務ID>","file_type":"output|error"}

Python

upload_status, upload_data = post_multipart_upload("<input.jsonl>")
create_status, create_data = post_json(
    "/v2/vectordb/ai/batch/jobs/create",
    {
        "input_file_id": upload_data["data"]["id"],
        "endpoint": "<endpoint>",
        "completion_window": "24h",
    },
)
describe_status, describe_data = post_json(
    "/v2/vectordb/ai/batch/jobs/describe",
    {"batch_id": create_data["data"]["id"]},
)
content_status, content_data = post_json(
    "/v2/vectordb/ai/batch/files/content",
    {"batch_id": create_data["data"]["id"], "file_type": "output"},
)

輸入檔案必須是 UTF-8 JSONL。每行是一條獨立請求,包含唯一 custom_id、固定值 method="POST"、與任務一致的 url 和 body.model。REST 指令碼依賴 jq。

參數說明

參數

說明

file

上傳時必填。JSONL 輸入檔案;服務端會快速校正第一行的 custom_id、method、url 和 body.model。

model_name

上傳時必填,也相容 model。必須與 JSONL 第一行的 body.model 一致。

endpoint

上傳和建立時必填。Chat Completions 使用 /v1/chat/completions,文本向量任務使用 /v1/embeddings;同一檔案的各行必須一致。

purpose

上傳時可選,只能為空白或 batch。

input_file_id

建立任務必填。由上傳介面返回的檔案 ID;不接受 OSS URL 或外部檔案標識。

completion_window

建立任務必填。完成視窗,支援 24h~336h 或天單位。

metadata.ds_name / metadata.ds_description

可選。任務名稱最長 100 個字元,描述最長 200 個字元。

batch_id

查詢和下載必填。建立介面返回的任務 ID。

file_type

下載必填。output 下載成功行,error 下載失敗行。

provider

可選,預設 aliyun_milvus。上傳、建立、查詢和下載必須使用同一百鍊帳號、地區和工作空間。

body.enable_thinking

按模型選填。應與 JSONL 行中的 body.model 同級;對於預設啟用思考的模型,顯式設為 false 可避免不需要的思考 Token。不要將該參數放入 extra_body。

傳回值說明

建立任務成功後返回 data.id,即 batch_id。輪詢狀態包括 validating、in_progress、finalizing、cancelling,終態包括 completed、failed、expired、cancelled。

批量任務採用排隊方式非同步處理,從建立到進入終態通常需要數分鐘到數小時,其間狀態持續為 in_progress。可隨時調用查詢介面 POST /v2/vectordb/ai/batch/jobs/describe(傳入 batch_id)擷取最新 status;若樣本指令碼的輪詢在任務完成前結束,可稍後憑 batch_id 重新查詢,待狀態變為 completed 後再下載結果。下載介面直接返回 JSONL 原文;每一行保留輸入時的 custom_id,用於將結果回寫到業務資料。批量任務不適合等待即時返回的線上互動。

樣本一:夜間批量產生知識庫摘要(文本)

營運團隊需要為當天新增文章批量產生摘要。每行一個文章請求,關鍵參數是 body.model=qwen3.7-max、url=/v1/chat/completions 和 completion_window=24h。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

MODEL_NAME="qwen3.7-max"
INPUT_FILE="${AIFUNC_AI_BATCH_INPUT_FILE:-}"
POLL_INTERVAL_SEC="${AIFUNC_AI_BATCH_POLL_INTERVAL_SEC:-5}"
MAX_POLL_ATTEMPTS="${AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS:-120}"

if [ -z "$INPUT_FILE" ]; then
  echo "請先將本樣本 JSONL 請求儲存為檔案,並通過 AIFUNC_AI_BATCH_INPUT_FILE 指定路徑。" >&2
  exit 1
fi

download_batch_file() {
  local file_type="$1"; local output_file="$2"; local content_body
  content_body="$(jq -nc --arg batch_id "$BATCH_ID" --arg file_type "$file_type" '{provider:"aliyun_milvus", batch_id:$batch_id, file_type:$file_type}')"
  curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/content" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" -H "Content-Type: application/json" \
    -d "$content_body" --output "$output_file"
}

UPLOAD_RESPONSE="$(curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/upload" \
  -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
  -F "provider=aliyun_milvus" -F "model_name=$MODEL_NAME" -F "endpoint=/v1/chat/completions" \
  -F "purpose=batch" -F "file=@$INPUT_FILE;type=application/jsonl")"
INPUT_FILE_ID="$(echo "$UPLOAD_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$INPUT_FILE_ID" ] || exit 1

CREATE_BODY="$(jq -nc --arg input_file_id "$INPUT_FILE_ID" '{provider:"aliyun_milvus", input_file_id:$input_file_id, endpoint:"/v1/chat/completions", completion_window:"24h", metadata:{ds_name:"milvus-ai-function-example", ds_description:"AI Batch REST example"}}')"
CREATE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/create" "$CREATE_BODY")"
BATCH_ID="$(echo "$CREATE_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$BATCH_ID" ] || exit 1

for ((attempt = 1; attempt <= MAX_POLL_ATTEMPTS; attempt++)); do
  DESCRIBE_BODY="$(jq -nc --arg batch_id "$BATCH_ID" '{provider:"aliyun_milvus", batch_id:$batch_id}')"
  DESCRIBE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/describe" "$DESCRIBE_BODY")"
  BATCH_STATUS="$(echo "$DESCRIBE_RESPONSE" | jq -r '.data.status // empty')"
  case "$BATCH_STATUS" in
    completed)
      OUTPUT_FILE="${AIFUNC_AI_BATCH_OUTPUT_FILE:-./ai_batch_output_${BATCH_ID}.jsonl}"
      download_batch_file "output" "$OUTPUT_FILE"
      echo "Batch output downloaded to: $OUTPUT_FILE"
      exit 0 ;;
    failed|expired|cancelled)
      echo "Batch $BATCH_ID ended with status: $BATCH_STATUS" >&2; exit 1 ;;
    validating|in_progress|finalizing|cancelling)
      [ "$attempt" -lt "$MAX_POLL_ATTEMPTS" ] && sleep "$POLL_INTERVAL_SEC" ;;
    *) echo "Unknown status: ${BATCH_STATUS:-empty}" >&2; exit 1 ;;
  esac
done
echo "Batch $BATCH_ID did not finish after $MAX_POLL_ATTEMPTS checks." >&2
exit 1

Python

from __future__ import annotations

import json
import os
import shutil
import sys
import tempfile
import time
import uuid
from pathlib import Path
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen

MILVUS_REST_BASE_URL = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN = "<yourUsername>:<yourPassword>"

MODEL_NAME = "qwen3.7-max"
INPUT_MODE = os.getenv("AIFUNC_AI_BATCH_INPUT_MODE", "text").lower()
POLL_INTERVAL_SEC = int(os.getenv("AIFUNC_AI_BATCH_POLL_INTERVAL_SEC", "5"))
MAX_POLL_ATTEMPTS = int(os.getenv("AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS", "120"))


def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
    request = Request(
        f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
        data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
        headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
        method="POST",
    )
    try:
        with urlopen(request, timeout=timeout) as response:
            return response.status, json.loads(response.read().decode("utf-8"))
    except HTTPError as exc:
        return exc.code, json.loads(exc.read().decode("utf-8"))


def batch_request(custom_id: str, content: Any, *, enable_thinking: bool | None = None) -> dict[str, Any]:
    body: dict[str, Any] = {"model": MODEL_NAME, "messages": [{"role": "user", "content": content}]}
    if enable_thinking is not None:
        body["enable_thinking"] = enable_thinking
    return {"custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": body}


def default_requests(input_mode: str) -> list[dict[str, Any]]:
    if input_mode == "text":
        return [
            batch_request("milvus-summary-1", "For a support knowledge base, summarize in one Chinese sentence: Milvus stores and searches vector embeddings for RAG, recommendation, and multimodal retrieval.", enable_thinking=False),
            batch_request("milvus-summary-2", "Turn this incident note into a Chinese FAQ title and one-sentence answer: after documents are updated, regenerate embeddings before users search the new content.", enable_thinking=False),
            batch_request("milvus-summary-3", "Create a one-sentence Chinese catalog description for an enterprise RAG case that retrieves policy passages before the assistant answers.", enable_thinking=False),
        ]
    if input_mode == "image":
        return [batch_request("milvus-image-1", [
            {"type": "image_url", "image_url": {"url": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"}},
            {"type": "text", "text": "For a pet-service media library, return a concise Chinese accessibility caption and up to three searchable subject tags for this uploaded case photo."},
        ])]
    if input_mode == "video":
        return [batch_request("milvus-video-1", [
            {"type": "video", "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"},
            {"type": "text", "text": "For a marketing asset library, return a one-sentence Chinese scene summary, three retrieval keywords, and whether manual brand-safety review is needed."},
        ])]
    if input_mode == "audio":
        return [batch_request("milvus-audio-1", [
            {"type": "input_audio", "input_audio": {"data": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "format": "mp3"}},
            {"type": "text", "text": "For a customer-service hotline greeting archive, transcribe the welcome message and assess whether its purpose, service availability, and recording or privacy notice are clear; mark details not present in the audio as unconfirmed."},
        ])]
    print("AIFUNC_AI_BATCH_INPUT_MODE must be one of text, image, video, or audio", file=sys.stderr)
    sys.exit(1)


def post_multipart_upload(input_file: str) -> tuple[int, dict[str, Any]]:
    boundary = f"----milvus-ai-batch-{uuid.uuid4().hex}"
    fields = {"provider": "aliyun_milvus", "model_name": MODEL_NAME, "endpoint": "/v1/chat/completions", "purpose": "batch"}
    with tempfile.TemporaryFile(mode="w+b") as payload:
        for name, value in fields.items():
            payload.write(f"--{boundary}\r\n".encode())
            payload.write(f'Content-Disposition: form-data; name="{name}"\r\n\r\n'.encode())
            payload.write(value.encode())
            payload.write(b"\r\n")
        payload.write(f"--{boundary}\r\n".encode())
        payload.write(b'Content-Disposition: form-data; name="file"; filename="input.jsonl"\r\n')
        payload.write(b"Content-Type: application/jsonl\r\n\r\n")
        with open(input_file, "rb") as input_handle:
            shutil.copyfileobj(input_handle, payload)
        payload.write(b"\r\n")
        payload.write(f"--{boundary}--\r\n".encode())
        payload_length = payload.tell()
        payload.seek(0)
        request = Request(
            f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/upload",
            data=payload,
            headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": f"multipart/form-data; boundary={boundary}", "Content-Length": str(payload_length)},
            method="POST",
        )
        try:
            with urlopen(request, timeout=600) as response:
                return response.status, json.loads(response.read().decode("utf-8"))
        except HTTPError as exc:
            return exc.code, json.loads(exc.read().decode("utf-8"))


def download_batch_file(batch_id: str, file_type: str, output_file: Path) -> None:
    request = Request(
        f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/content",
        data=json.dumps({"provider": "aliyun_milvus", "batch_id": batch_id, "file_type": file_type}, ensure_ascii=False).encode("utf-8"),
        headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
        method="POST",
    )
    with urlopen(request, timeout=600) as response, output_file.open("wb") as output:
        shutil.copyfileobj(response, output)


temporary_input = None
INPUT_FILE = os.getenv("AIFUNC_AI_BATCH_INPUT_FILE")
if INPUT_FILE is None:
    temporary_input = tempfile.NamedTemporaryFile(mode="w", suffix=".jsonl", delete=False, encoding="utf-8")
    for request in default_requests(INPUT_MODE):
        temporary_input.write(json.dumps(request, ensure_ascii=False, separators=(",", ":")) + "\n")
    temporary_input.close()
    INPUT_FILE = temporary_input.name

try:
    status, data = post_multipart_upload(INPUT_FILE)
    if status != 200 or data.get("code") != 0:
        sys.exit(1)
    input_file_id = data.get("data", {}).get("id")

    status, data = post_json("/v2/vectordb/ai/batch/jobs/create", {
        "provider": "aliyun_milvus", "input_file_id": input_file_id, "endpoint": "/v1/chat/completions",
        "completion_window": "24h", "metadata": {"ds_name": "milvus-ai-function-example", "ds_description": "AI Batch example"},
    })
    if status != 200 or data.get("code") != 0:
        sys.exit(1)
    batch_id = data.get("data", {}).get("id")

    for attempt in range(1, MAX_POLL_ATTEMPTS + 1):
        status, data = post_json("/v2/vectordb/ai/batch/jobs/describe", {"provider": "aliyun_milvus", "batch_id": batch_id})
        batch_data = data.get("data", {})
        batch_status = batch_data.get("status")
        if batch_status == "completed":
            output_file = Path(os.getenv("AIFUNC_AI_BATCH_OUTPUT_FILE", f"ai_batch_output_{batch_id}.jsonl"))
            download_batch_file(batch_id, "output", output_file)
            print(f"Batch output downloaded to: {output_file}")
            if batch_data.get("error_file_id"):
                download_batch_file(batch_id, "error", Path(f"ai_batch_errors_{batch_id}.jsonl"))
            break
        if batch_status in {"failed", "expired", "cancelled"}:
            print(f"Batch {batch_id} ended with status: {batch_status}", file=sys.stderr)
            sys.exit(1)
        if attempt < MAX_POLL_ATTEMPTS:
            time.sleep(POLL_INTERVAL_SEC)
finally:
    if temporary_input is not None:
        os.unlink(temporary_input.name)

預期結果:任務進入 completed 後下載的 output.jsonl 中每一行都保留 custom_id,例如

{"custom_id":"milvus-summary-1","response":{"status_code":200,"body":{"choices":[{"message":{"role":"assistant","content":"Milvus 是服務於 RAG、推薦和多模態檢索情境的向量資料庫。"}}]}}}

業務側按 custom_id 將摘要回寫到相應文章;如產生 error.jsonl,只重試其中失敗的文章。

樣本二:批量產生寵物案例圖片描述(圖片)

寵物服務平台需要為商家上傳的歷史案例圖片產生無障礙描述和檢索標籤。每行的 content 包含 image_url 與文本指令,使用支援映像理解的多模態模型,上傳時的 model_name 必須一致。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

MODEL_NAME="qwen3-vl-plus"
INPUT_FILE="${AIFUNC_AI_BATCH_INPUT_FILE:-}"
POLL_INTERVAL_SEC="${AIFUNC_AI_BATCH_POLL_INTERVAL_SEC:-5}"
MAX_POLL_ATTEMPTS="${AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS:-120}"

if [ -z "$INPUT_FILE" ]; then
  echo "請先將本樣本 JSONL 請求儲存為檔案,並通過 AIFUNC_AI_BATCH_INPUT_FILE 指定路徑。" >&2
  exit 1
fi

download_batch_file() {
  local file_type="$1"; local output_file="$2"; local content_body
  content_body="$(jq -nc --arg batch_id "$BATCH_ID" --arg file_type "$file_type" '{provider:"aliyun_milvus", batch_id:$batch_id, file_type:$file_type}')"
  curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/content" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" -H "Content-Type: application/json" \
    -d "$content_body" --output "$output_file"
}

UPLOAD_RESPONSE="$(curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/upload" \
  -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
  -F "provider=aliyun_milvus" -F "model_name=$MODEL_NAME" -F "endpoint=/v1/chat/completions" \
  -F "purpose=batch" -F "file=@$INPUT_FILE;type=application/jsonl")"
INPUT_FILE_ID="$(echo "$UPLOAD_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$INPUT_FILE_ID" ] || exit 1

CREATE_BODY="$(jq -nc --arg input_file_id "$INPUT_FILE_ID" '{provider:"aliyun_milvus", input_file_id:$input_file_id, endpoint:"/v1/chat/completions", completion_window:"24h", metadata:{ds_name:"milvus-ai-function-example", ds_description:"AI Batch REST example"}}')"
CREATE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/create" "$CREATE_BODY")"
BATCH_ID="$(echo "$CREATE_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$BATCH_ID" ] || exit 1

for ((attempt = 1; attempt <= MAX_POLL_ATTEMPTS; attempt++)); do
  DESCRIBE_BODY="$(jq -nc --arg batch_id "$BATCH_ID" '{provider:"aliyun_milvus", batch_id:$batch_id}')"
  DESCRIBE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/describe" "$DESCRIBE_BODY")"
  BATCH_STATUS="$(echo "$DESCRIBE_RESPONSE" | jq -r '.data.status // empty')"
  case "$BATCH_STATUS" in
    completed)
      OUTPUT_FILE="${AIFUNC_AI_BATCH_OUTPUT_FILE:-./ai_batch_output_${BATCH_ID}.jsonl}"
      download_batch_file "output" "$OUTPUT_FILE"
      echo "Batch output downloaded to: $OUTPUT_FILE"
      exit 0 ;;
    failed|expired|cancelled)
      echo "Batch $BATCH_ID ended with status: $BATCH_STATUS" >&2; exit 1 ;;
    validating|in_progress|finalizing|cancelling)
      [ "$attempt" -lt "$MAX_POLL_ATTEMPTS" ] && sleep "$POLL_INTERVAL_SEC" ;;
    *) echo "Unknown status: ${BATCH_STATUS:-empty}" >&2; exit 1 ;;
  esac
done
echo "Batch $BATCH_ID did not finish after $MAX_POLL_ATTEMPTS checks." >&2
exit 1

Python

from __future__ import annotations

import json
import os
import shutil
import sys
import tempfile
import time
import uuid
from pathlib import Path
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen

MILVUS_REST_BASE_URL = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN = "<yourUsername>:<yourPassword>"

MODEL_NAME = "qwen3-vl-plus"
INPUT_MODE = os.getenv("AIFUNC_AI_BATCH_INPUT_MODE", "image").lower()
POLL_INTERVAL_SEC = int(os.getenv("AIFUNC_AI_BATCH_POLL_INTERVAL_SEC", "5"))
MAX_POLL_ATTEMPTS = int(os.getenv("AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS", "120"))


def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
    request = Request(
        f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
        data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
        headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
        method="POST",
    )
    try:
        with urlopen(request, timeout=timeout) as response:
            return response.status, json.loads(response.read().decode("utf-8"))
    except HTTPError as exc:
        return exc.code, json.loads(exc.read().decode("utf-8"))


def batch_request(custom_id: str, content: Any, *, enable_thinking: bool | None = None) -> dict[str, Any]:
    body: dict[str, Any] = {"model": MODEL_NAME, "messages": [{"role": "user", "content": content}]}
    if enable_thinking is not None:
        body["enable_thinking"] = enable_thinking
    return {"custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": body}


def default_requests(input_mode: str) -> list[dict[str, Any]]:
    if input_mode == "text":
        return [
            batch_request("milvus-summary-1", "For a support knowledge base, summarize in one Chinese sentence: Milvus stores and searches vector embeddings for RAG, recommendation, and multimodal retrieval.", enable_thinking=False),
            batch_request("milvus-summary-2", "Turn this incident note into a Chinese FAQ title and one-sentence answer: after documents are updated, regenerate embeddings before users search the new content.", enable_thinking=False),
            batch_request("milvus-summary-3", "Create a one-sentence Chinese catalog description for an enterprise RAG case that retrieves policy passages before the assistant answers.", enable_thinking=False),
        ]
    if input_mode == "image":
        return [batch_request("milvus-image-1", [
            {"type": "image_url", "image_url": {"url": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"}},
            {"type": "text", "text": "For a pet-service media library, return a concise Chinese accessibility caption and up to three searchable subject tags for this uploaded case photo."},
        ])]
    if input_mode == "video":
        return [batch_request("milvus-video-1", [
            {"type": "video", "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"},
            {"type": "text", "text": "For a marketing asset library, return a one-sentence Chinese scene summary, three retrieval keywords, and whether manual brand-safety review is needed."},
        ])]
    if input_mode == "audio":
        return [batch_request("milvus-audio-1", [
            {"type": "input_audio", "input_audio": {"data": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "format": "mp3"}},
            {"type": "text", "text": "For a customer-service hotline greeting archive, transcribe the welcome message and assess whether its purpose, service availability, and recording or privacy notice are clear; mark details not present in the audio as unconfirmed."},
        ])]
    print("AIFUNC_AI_BATCH_INPUT_MODE must be one of text, image, video, or audio", file=sys.stderr)
    sys.exit(1)


def post_multipart_upload(input_file: str) -> tuple[int, dict[str, Any]]:
    boundary = f"----milvus-ai-batch-{uuid.uuid4().hex}"
    fields = {"provider": "aliyun_milvus", "model_name": MODEL_NAME, "endpoint": "/v1/chat/completions", "purpose": "batch"}
    with tempfile.TemporaryFile(mode="w+b") as payload:
        for name, value in fields.items():
            payload.write(f"--{boundary}\r\n".encode())
            payload.write(f'Content-Disposition: form-data; name="{name}"\r\n\r\n'.encode())
            payload.write(value.encode())
            payload.write(b"\r\n")
        payload.write(f"--{boundary}\r\n".encode())
        payload.write(b'Content-Disposition: form-data; name="file"; filename="input.jsonl"\r\n')
        payload.write(b"Content-Type: application/jsonl\r\n\r\n")
        with open(input_file, "rb") as input_handle:
            shutil.copyfileobj(input_handle, payload)
        payload.write(b"\r\n")
        payload.write(f"--{boundary}--\r\n".encode())
        payload_length = payload.tell()
        payload.seek(0)
        request = Request(
            f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/upload",
            data=payload,
            headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": f"multipart/form-data; boundary={boundary}", "Content-Length": str(payload_length)},
            method="POST",
        )
        try:
            with urlopen(request, timeout=600) as response:
                return response.status, json.loads(response.read().decode("utf-8"))
        except HTTPError as exc:
            return exc.code, json.loads(exc.read().decode("utf-8"))


def download_batch_file(batch_id: str, file_type: str, output_file: Path) -> None:
    request = Request(
        f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/content",
        data=json.dumps({"provider": "aliyun_milvus", "batch_id": batch_id, "file_type": file_type}, ensure_ascii=False).encode("utf-8"),
        headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
        method="POST",
    )
    with urlopen(request, timeout=600) as response, output_file.open("wb") as output:
        shutil.copyfileobj(response, output)


temporary_input = None
INPUT_FILE = os.getenv("AIFUNC_AI_BATCH_INPUT_FILE")
if INPUT_FILE is None:
    temporary_input = tempfile.NamedTemporaryFile(mode="w", suffix=".jsonl", delete=False, encoding="utf-8")
    for request in default_requests(INPUT_MODE):
        temporary_input.write(json.dumps(request, ensure_ascii=False, separators=(",", ":")) + "\n")
    temporary_input.close()
    INPUT_FILE = temporary_input.name

try:
    status, data = post_multipart_upload(INPUT_FILE)
    if status != 200 or data.get("code") != 0:
        sys.exit(1)
    input_file_id = data.get("data", {}).get("id")

    status, data = post_json("/v2/vectordb/ai/batch/jobs/create", {
        "provider": "aliyun_milvus", "input_file_id": input_file_id, "endpoint": "/v1/chat/completions",
        "completion_window": "24h", "metadata": {"ds_name": "milvus-ai-function-example", "ds_description": "AI Batch example"},
    })
    if status != 200 or data.get("code") != 0:
        sys.exit(1)
    batch_id = data.get("data", {}).get("id")

    for attempt in range(1, MAX_POLL_ATTEMPTS + 1):
        status, data = post_json("/v2/vectordb/ai/batch/jobs/describe", {"provider": "aliyun_milvus", "batch_id": batch_id})
        batch_data = data.get("data", {})
        batch_status = batch_data.get("status")
        if batch_status == "completed":
            output_file = Path(os.getenv("AIFUNC_AI_BATCH_OUTPUT_FILE", f"ai_batch_output_{batch_id}.jsonl"))
            download_batch_file(batch_id, "output", output_file)
            print(f"Batch output downloaded to: {output_file}")
            if batch_data.get("error_file_id"):
                download_batch_file(batch_id, "error", Path(f"ai_batch_errors_{batch_id}.jsonl"))
            break
        if batch_status in {"failed", "expired", "cancelled"}:
            print(f"Batch {batch_id} ended with status: {batch_status}", file=sys.stderr)
            sys.exit(1)
        if attempt < MAX_POLL_ATTEMPTS:
            time.sleep(POLL_INTERVAL_SEC)
finally:
    if temporary_input is not None:
        os.unlink(temporary_input.name)

預期結果:image-output.jsonl 使用同一 custom_id=milvus-image-1 返回圖片描述和標籤。業務側可將文本寫入案例庫欄位,再結合圖片向量實現圖文檢索;若任務完成但有失敗行,仍需下載 file_type=error 的結果檔案。

樣本三:批量產生視頻素材摘要(視頻)

營運團隊需要為歷史短視頻產生一句摘要,供視頻素材檢索。每行通過 video 內容塊傳入的視訊 URL;使用可使用視訊理解的多模態模型。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

MODEL_NAME="qwen3-vl-plus"
INPUT_FILE="${AIFUNC_AI_BATCH_INPUT_FILE:-}"
POLL_INTERVAL_SEC="${AIFUNC_AI_BATCH_POLL_INTERVAL_SEC:-5}"
MAX_POLL_ATTEMPTS="${AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS:-120}"

if [ -z "$INPUT_FILE" ]; then
  echo "請先將本樣本 JSONL 請求儲存為檔案,並通過 AIFUNC_AI_BATCH_INPUT_FILE 指定路徑。" >&2
  exit 1
fi

download_batch_file() {
  local file_type="$1"; local output_file="$2"; local content_body
  content_body="$(jq -nc --arg batch_id "$BATCH_ID" --arg file_type "$file_type" '{provider:"aliyun_milvus", batch_id:$batch_id, file_type:$file_type}')"
  curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/content" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" -H "Content-Type: application/json" \
    -d "$content_body" --output "$output_file"
}

UPLOAD_RESPONSE="$(curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/upload" \
  -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
  -F "provider=aliyun_milvus" -F "model_name=$MODEL_NAME" -F "endpoint=/v1/chat/completions" \
  -F "purpose=batch" -F "file=@$INPUT_FILE;type=application/jsonl")"
INPUT_FILE_ID="$(echo "$UPLOAD_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$INPUT_FILE_ID" ] || exit 1

CREATE_BODY="$(jq -nc --arg input_file_id "$INPUT_FILE_ID" '{provider:"aliyun_milvus", input_file_id:$input_file_id, endpoint:"/v1/chat/completions", completion_window:"24h", metadata:{ds_name:"milvus-ai-function-example", ds_description:"AI Batch REST example"}}')"
CREATE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/create" "$CREATE_BODY")"
BATCH_ID="$(echo "$CREATE_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$BATCH_ID" ] || exit 1

for ((attempt = 1; attempt <= MAX_POLL_ATTEMPTS; attempt++)); do
  DESCRIBE_BODY="$(jq -nc --arg batch_id "$BATCH_ID" '{provider:"aliyun_milvus", batch_id:$batch_id}')"
  DESCRIBE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/describe" "$DESCRIBE_BODY")"
  BATCH_STATUS="$(echo "$DESCRIBE_RESPONSE" | jq -r '.data.status // empty')"
  case "$BATCH_STATUS" in
    completed)
      OUTPUT_FILE="${AIFUNC_AI_BATCH_OUTPUT_FILE:-./ai_batch_output_${BATCH_ID}.jsonl}"
      download_batch_file "output" "$OUTPUT_FILE"
      echo "Batch output downloaded to: $OUTPUT_FILE"
      exit 0 ;;
    failed|expired|cancelled)
      echo "Batch $BATCH_ID ended with status: $BATCH_STATUS" >&2; exit 1 ;;
    validating|in_progress|finalizing|cancelling)
      [ "$attempt" -lt "$MAX_POLL_ATTEMPTS" ] && sleep "$POLL_INTERVAL_SEC" ;;
    *) echo "Unknown status: ${BATCH_STATUS:-empty}" >&2; exit 1 ;;
  esac
done
echo "Batch $BATCH_ID did not finish after $MAX_POLL_ATTEMPTS checks." >&2
exit 1

Python

from __future__ import annotations

import json
import os
import shutil
import sys
import tempfile
import time
import uuid
from pathlib import Path
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen

MILVUS_REST_BASE_URL = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN = "<yourUsername>:<yourPassword>"

MODEL_NAME = "qwen3-vl-plus"
INPUT_MODE = os.getenv("AIFUNC_AI_BATCH_INPUT_MODE", "video").lower()
POLL_INTERVAL_SEC = int(os.getenv("AIFUNC_AI_BATCH_POLL_INTERVAL_SEC", "5"))
MAX_POLL_ATTEMPTS = int(os.getenv("AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS", "120"))


def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
    request = Request(
        f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
        data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
        headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
        method="POST",
    )
    try:
        with urlopen(request, timeout=timeout) as response:
            return response.status, json.loads(response.read().decode("utf-8"))
    except HTTPError as exc:
        return exc.code, json.loads(exc.read().decode("utf-8"))


def batch_request(custom_id: str, content: Any, *, enable_thinking: bool | None = None) -> dict[str, Any]:
    body: dict[str, Any] = {"model": MODEL_NAME, "messages": [{"role": "user", "content": content}]}
    if enable_thinking is not None:
        body["enable_thinking"] = enable_thinking
    return {"custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": body}


def default_requests(input_mode: str) -> list[dict[str, Any]]:
    if input_mode == "text":
        return [
            batch_request("milvus-summary-1", "For a support knowledge base, summarize in one Chinese sentence: Milvus stores and searches vector embeddings for RAG, recommendation, and multimodal retrieval.", enable_thinking=False),
            batch_request("milvus-summary-2", "Turn this incident note into a Chinese FAQ title and one-sentence answer: after documents are updated, regenerate embeddings before users search the new content.", enable_thinking=False),
            batch_request("milvus-summary-3", "Create a one-sentence Chinese catalog description for an enterprise RAG case that retrieves policy passages before the assistant answers.", enable_thinking=False),
        ]
    if input_mode == "image":
        return [batch_request("milvus-image-1", [
            {"type": "image_url", "image_url": {"url": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"}},
            {"type": "text", "text": "For a pet-service media library, return a concise Chinese accessibility caption and up to three searchable subject tags for this uploaded case photo."},
        ])]
    if input_mode == "video":
        return [batch_request("milvus-video-1", [
            {"type": "video", "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"},
            {"type": "text", "text": "For a marketing asset library, return a one-sentence Chinese scene summary, three retrieval keywords, and whether manual brand-safety review is needed."},
        ])]
    if input_mode == "audio":
        return [batch_request("milvus-audio-1", [
            {"type": "input_audio", "input_audio": {"data": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "format": "mp3"}},
            {"type": "text", "text": "For a customer-service hotline greeting archive, transcribe the welcome message and assess whether its purpose, service availability, and recording or privacy notice are clear; mark details not present in the audio as unconfirmed."},
        ])]
    print("AIFUNC_AI_BATCH_INPUT_MODE must be one of text, image, video, or audio", file=sys.stderr)
    sys.exit(1)


def post_multipart_upload(input_file: str) -> tuple[int, dict[str, Any]]:
    boundary = f"----milvus-ai-batch-{uuid.uuid4().hex}"
    fields = {"provider": "aliyun_milvus", "model_name": MODEL_NAME, "endpoint": "/v1/chat/completions", "purpose": "batch"}
    with tempfile.TemporaryFile(mode="w+b") as payload:
        for name, value in fields.items():
            payload.write(f"--{boundary}\r\n".encode())
            payload.write(f'Content-Disposition: form-data; name="{name}"\r\n\r\n'.encode())
            payload.write(value.encode())
            payload.write(b"\r\n")
        payload.write(f"--{boundary}\r\n".encode())
        payload.write(b'Content-Disposition: form-data; name="file"; filename="input.jsonl"\r\n')
        payload.write(b"Content-Type: application/jsonl\r\n\r\n")
        with open(input_file, "rb") as input_handle:
            shutil.copyfileobj(input_handle, payload)
        payload.write(b"\r\n")
        payload.write(f"--{boundary}--\r\n".encode())
        payload_length = payload.tell()
        payload.seek(0)
        request = Request(
            f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/upload",
            data=payload,
            headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": f"multipart/form-data; boundary={boundary}", "Content-Length": str(payload_length)},
            method="POST",
        )
        try:
            with urlopen(request, timeout=600) as response:
                return response.status, json.loads(response.read().decode("utf-8"))
        except HTTPError as exc:
            return exc.code, json.loads(exc.read().decode("utf-8"))


def download_batch_file(batch_id: str, file_type: str, output_file: Path) -> None:
    request = Request(
        f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/content",
        data=json.dumps({"provider": "aliyun_milvus", "batch_id": batch_id, "file_type": file_type}, ensure_ascii=False).encode("utf-8"),
        headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
        method="POST",
    )
    with urlopen(request, timeout=600) as response, output_file.open("wb") as output:
        shutil.copyfileobj(response, output)


temporary_input = None
INPUT_FILE = os.getenv("AIFUNC_AI_BATCH_INPUT_FILE")
if INPUT_FILE is None:
    temporary_input = tempfile.NamedTemporaryFile(mode="w", suffix=".jsonl", delete=False, encoding="utf-8")
    for request in default_requests(INPUT_MODE):
        temporary_input.write(json.dumps(request, ensure_ascii=False, separators=(",", ":")) + "\n")
    temporary_input.close()
    INPUT_FILE = temporary_input.name

try:
    status, data = post_multipart_upload(INPUT_FILE)
    if status != 200 or data.get("code") != 0:
        sys.exit(1)
    input_file_id = data.get("data", {}).get("id")

    status, data = post_json("/v2/vectordb/ai/batch/jobs/create", {
        "provider": "aliyun_milvus", "input_file_id": input_file_id, "endpoint": "/v1/chat/completions",
        "completion_window": "24h", "metadata": {"ds_name": "milvus-ai-function-example", "ds_description": "AI Batch example"},
    })
    if status != 200 or data.get("code") != 0:
        sys.exit(1)
    batch_id = data.get("data", {}).get("id")

    for attempt in range(1, MAX_POLL_ATTEMPTS + 1):
        status, data = post_json("/v2/vectordb/ai/batch/jobs/describe", {"provider": "aliyun_milvus", "batch_id": batch_id})
        batch_data = data.get("data", {})
        batch_status = batch_data.get("status")
        if batch_status == "completed":
            output_file = Path(os.getenv("AIFUNC_AI_BATCH_OUTPUT_FILE", f"ai_batch_output_{batch_id}.jsonl"))
            download_batch_file(batch_id, "output", output_file)
            print(f"Batch output downloaded to: {output_file}")
            if batch_data.get("error_file_id"):
                download_batch_file(batch_id, "error", Path(f"ai_batch_errors_{batch_id}.jsonl"))
            break
        if batch_status in {"failed", "expired", "cancelled"}:
            print(f"Batch {batch_id} ended with status: {batch_status}", file=sys.stderr)
            sys.exit(1)
        if attempt < MAX_POLL_ATTEMPTS:
            time.sleep(POLL_INTERVAL_SEC)
finally:
    if temporary_input is not None:
        os.unlink(temporary_input.name)

預期結果:video-output.jsonl 以 custom_id=milvus-video-1 返回視頻摘要。將摘要寫入視頻素材中繼資料後,可結合視頻向量和關鍵詞完成檢索;終態帶有 error_file_id 時,應保留錯誤檔案以定位失敗任務。

樣本四:夜間歸檔客服熱線歡迎語(音頻)

客服營運團隊會定期歸檔各條服務熱線的歡迎語,轉寫其內容並檢查熱線用途、服務時間、錄音或隱私告知是否清晰,再按歡迎語版本 ID 回寫配置中心。每行通過 input_audio 傳入一段 MP3 歡迎語,使用支援音頻理解的模型,並以唯一 custom_id 關聯原始版本。

音頻和轉寫可能包含個人資訊或內部服務配置。運行前應確認採集授權及脫敏要求,輸出僅保留到完成回寫和抽檢所需的期限,之後按組織的資料保留原則安全刪除;不要將生產錄音或結果提交到代碼倉庫。

REST 介面

#!/usr/bin/env bash
set -euo pipefail

MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"

post_json() {
  local path="$1"
  local body="$2"
  curl -X POST \
    "$MILVUS_REST_BASE_URL$path" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$body"
}

MODEL_NAME="qwen3.5-omni-plus"
INPUT_FILE="${AIFUNC_AI_BATCH_INPUT_FILE:-}"
POLL_INTERVAL_SEC="${AIFUNC_AI_BATCH_POLL_INTERVAL_SEC:-5}"
MAX_POLL_ATTEMPTS="${AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS:-120}"

if [ -z "$INPUT_FILE" ]; then
  echo "請先將本樣本 JSONL 請求儲存為檔案,並通過 AIFUNC_AI_BATCH_INPUT_FILE 指定路徑。" >&2
  exit 1
fi

download_batch_file() {
  local file_type="$1"; local output_file="$2"; local content_body
  content_body="$(jq -nc --arg batch_id "$BATCH_ID" --arg file_type "$file_type" '{provider:"aliyun_milvus", batch_id:$batch_id, file_type:$file_type}')"
  curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/content" \
    -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" -H "Content-Type: application/json" \
    -d "$content_body" --output "$output_file"
}

UPLOAD_RESPONSE="$(curl --fail --silent --show-error -X POST "$MILVUS_REST_BASE_URL/v2/vectordb/ai/batch/files/upload" \
  -H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
  -F "provider=aliyun_milvus" -F "model_name=$MODEL_NAME" -F "endpoint=/v1/chat/completions" \
  -F "purpose=batch" -F "file=@$INPUT_FILE;type=application/jsonl")"
INPUT_FILE_ID="$(echo "$UPLOAD_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$INPUT_FILE_ID" ] || exit 1

CREATE_BODY="$(jq -nc --arg input_file_id "$INPUT_FILE_ID" '{provider:"aliyun_milvus", input_file_id:$input_file_id, endpoint:"/v1/chat/completions", completion_window:"24h", metadata:{ds_name:"milvus-ai-function-example", ds_description:"AI Batch REST example"}}')"
CREATE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/create" "$CREATE_BODY")"
BATCH_ID="$(echo "$CREATE_RESPONSE" | jq -r '.data.id // empty')"
[ -n "$BATCH_ID" ] || exit 1

for ((attempt = 1; attempt <= MAX_POLL_ATTEMPTS; attempt++)); do
  DESCRIBE_BODY="$(jq -nc --arg batch_id "$BATCH_ID" '{provider:"aliyun_milvus", batch_id:$batch_id}')"
  DESCRIBE_RESPONSE="$(post_json "/v2/vectordb/ai/batch/jobs/describe" "$DESCRIBE_BODY")"
  BATCH_STATUS="$(echo "$DESCRIBE_RESPONSE" | jq -r '.data.status // empty')"
  case "$BATCH_STATUS" in
    completed)
      OUTPUT_FILE="${AIFUNC_AI_BATCH_OUTPUT_FILE:-./ai_batch_output_${BATCH_ID}.jsonl}"
      download_batch_file "output" "$OUTPUT_FILE"
      echo "Batch output downloaded to: $OUTPUT_FILE"
      exit 0 ;;
    failed|expired|cancelled)
      echo "Batch $BATCH_ID ended with status: $BATCH_STATUS" >&2; exit 1 ;;
    validating|in_progress|finalizing|cancelling)
      [ "$attempt" -lt "$MAX_POLL_ATTEMPTS" ] && sleep "$POLL_INTERVAL_SEC" ;;
    *) echo "Unknown status: ${BATCH_STATUS:-empty}" >&2; exit 1 ;;
  esac
done
echo "Batch $BATCH_ID did not finish after $MAX_POLL_ATTEMPTS checks." >&2
exit 1

Python

from __future__ import annotations

import json
import os
import shutil
import sys
import tempfile
import time
import uuid
from pathlib import Path
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen

MILVUS_REST_BASE_URL = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN = "<yourUsername>:<yourPassword>"

MODEL_NAME = "qwen3.5-omni-plus"
INPUT_MODE = os.getenv("AIFUNC_AI_BATCH_INPUT_MODE", "audio").lower()
POLL_INTERVAL_SEC = int(os.getenv("AIFUNC_AI_BATCH_POLL_INTERVAL_SEC", "5"))
MAX_POLL_ATTEMPTS = int(os.getenv("AIFUNC_AI_BATCH_MAX_POLL_ATTEMPTS", "120"))


def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
    request = Request(
        f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
        data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
        headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
        method="POST",
    )
    try:
        with urlopen(request, timeout=timeout) as response:
            return response.status, json.loads(response.read().decode("utf-8"))
    except HTTPError as exc:
        return exc.code, json.loads(exc.read().decode("utf-8"))


def batch_request(custom_id: str, content: Any, *, enable_thinking: bool | None = None) -> dict[str, Any]:
    body: dict[str, Any] = {"model": MODEL_NAME, "messages": [{"role": "user", "content": content}]}
    if enable_thinking is not None:
        body["enable_thinking"] = enable_thinking
    return {"custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": body}


def default_requests(input_mode: str) -> list[dict[str, Any]]:
    if input_mode == "text":
        return [
            batch_request("milvus-summary-1", "For a support knowledge base, summarize in one Chinese sentence: Milvus stores and searches vector embeddings for RAG, recommendation, and multimodal retrieval.", enable_thinking=False),
            batch_request("milvus-summary-2", "Turn this incident note into a Chinese FAQ title and one-sentence answer: after documents are updated, regenerate embeddings before users search the new content.", enable_thinking=False),
            batch_request("milvus-summary-3", "Create a one-sentence Chinese catalog description for an enterprise RAG case that retrieves policy passages before the assistant answers.", enable_thinking=False),
        ]
    if input_mode == "image":
        return [batch_request("milvus-image-1", [
            {"type": "image_url", "image_url": {"url": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/dog_and_girl.jpeg"}},
            {"type": "text", "text": "For a pet-service media library, return a concise Chinese accessibility caption and up to three searchable subject tags for this uploaded case photo."},
        ])]
    if input_mode == "video":
        return [batch_request("milvus-video-1", [
            {"type": "video", "video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260409/dozxak/Wan_Video_Edit_33_1.mp4"},
            {"type": "text", "text": "For a marketing asset library, return a one-sentence Chinese scene summary, three retrieval keywords, and whether manual brand-safety review is needed."},
        ])]
    if input_mode == "audio":
        return [batch_request("milvus-audio-1", [
            {"type": "input_audio", "input_audio": {"data": "https://dashscope.oss-cn-beijing.aliyuncs.com/audios/welcome.mp3", "format": "mp3"}},
            {"type": "text", "text": "For a customer-service hotline greeting archive, transcribe the welcome message and assess whether its purpose, service availability, and recording or privacy notice are clear; mark details not present in the audio as unconfirmed."},
        ])]
    print("AIFUNC_AI_BATCH_INPUT_MODE must be one of text, image, video, or audio", file=sys.stderr)
    sys.exit(1)


def post_multipart_upload(input_file: str) -> tuple[int, dict[str, Any]]:
    boundary = f"----milvus-ai-batch-{uuid.uuid4().hex}"
    fields = {"provider": "aliyun_milvus", "model_name": MODEL_NAME, "endpoint": "/v1/chat/completions", "purpose": "batch"}
    with tempfile.TemporaryFile(mode="w+b") as payload:
        for name, value in fields.items():
            payload.write(f"--{boundary}\r\n".encode())
            payload.write(f'Content-Disposition: form-data; name="{name}"\r\n\r\n'.encode())
            payload.write(value.encode())
            payload.write(b"\r\n")
        payload.write(f"--{boundary}\r\n".encode())
        payload.write(b'Content-Disposition: form-data; name="file"; filename="input.jsonl"\r\n')
        payload.write(b"Content-Type: application/jsonl\r\n\r\n")
        with open(input_file, "rb") as input_handle:
            shutil.copyfileobj(input_handle, payload)
        payload.write(b"\r\n")
        payload.write(f"--{boundary}--\r\n".encode())
        payload_length = payload.tell()
        payload.seek(0)
        request = Request(
            f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/upload",
            data=payload,
            headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": f"multipart/form-data; boundary={boundary}", "Content-Length": str(payload_length)},
            method="POST",
        )
        try:
            with urlopen(request, timeout=600) as response:
                return response.status, json.loads(response.read().decode("utf-8"))
        except HTTPError as exc:
            return exc.code, json.loads(exc.read().decode("utf-8"))


def download_batch_file(batch_id: str, file_type: str, output_file: Path) -> None:
    request = Request(
        f"{MILVUS_REST_BASE_URL.rstrip('/')}/v2/vectordb/ai/batch/files/content",
        data=json.dumps({"provider": "aliyun_milvus", "batch_id": batch_id, "file_type": file_type}, ensure_ascii=False).encode("utf-8"),
        headers={"Authorization": f"Bearer {MILVUS_AUTH_TOKEN}", "Content-Type": "application/json"},
        method="POST",
    )
    with urlopen(request, timeout=600) as response, output_file.open("wb") as output:
        shutil.copyfileobj(response, output)


temporary_input = None
INPUT_FILE = os.getenv("AIFUNC_AI_BATCH_INPUT_FILE")
if INPUT_FILE is None:
    temporary_input = tempfile.NamedTemporaryFile(mode="w", suffix=".jsonl", delete=False, encoding="utf-8")
    for request in default_requests(INPUT_MODE):
        temporary_input.write(json.dumps(request, ensure_ascii=False, separators=(",", ":")) + "\n")
    temporary_input.close()
    INPUT_FILE = temporary_input.name

try:
    status, data = post_multipart_upload(INPUT_FILE)
    if status != 200 or data.get("code") != 0:
        sys.exit(1)
    input_file_id = data.get("data", {}).get("id")

    status, data = post_json("/v2/vectordb/ai/batch/jobs/create", {
        "provider": "aliyun_milvus", "input_file_id": input_file_id, "endpoint": "/v1/chat/completions",
        "completion_window": "24h", "metadata": {"ds_name": "milvus-ai-function-example", "ds_description": "AI Batch example"},
    })
    if status != 200 or data.get("code") != 0:
        sys.exit(1)
    batch_id = data.get("data", {}).get("id")

    for attempt in range(1, MAX_POLL_ATTEMPTS + 1):
        status, data = post_json("/v2/vectordb/ai/batch/jobs/describe", {"provider": "aliyun_milvus", "batch_id": batch_id})
        batch_data = data.get("data", {})
        batch_status = batch_data.get("status")
        if batch_status == "completed":
            output_file = Path(os.getenv("AIFUNC_AI_BATCH_OUTPUT_FILE", f"ai_batch_output_{batch_id}.jsonl"))
            download_batch_file(batch_id, "output", output_file)
            print(f"Batch output downloaded to: {output_file}")
            if batch_data.get("error_file_id"):
                download_batch_file(batch_id, "error", Path(f"ai_batch_errors_{batch_id}.jsonl"))
            break
        if batch_status in {"failed", "expired", "cancelled"}:
            print(f"Batch {batch_id} ended with status: {batch_status}", file=sys.stderr)
            sys.exit(1)
        if attempt < MAX_POLL_ATTEMPTS:
            time.sleep(POLL_INTERVAL_SEC)
finally:
    if temporary_input is not None:
        os.unlink(temporary_input.name)

預期結果:call-greeting-output.jsonl 或 call-greeting-error.jsonl 中恰有一行 custom_id=milvus-audio-1。成功結果經有限範圍校正後寫入配置中心並交由客服營運人員抽檢;失敗行保留原 custom_id,可在修複音頻許可權、MP3 格式或模型參數後單獨重試。