全部產品
Search
文件中心

Vector Retrieval Service for Milvus:通過阿里雲Milvus構建客服語音分析與FAQ智能匹配鏈路

更新時間:Aug 14, 2026

本文介紹如何用阿里雲 Milvus 的六個 AI Function 串聯一條客服語音處理鏈路:錄音寫入時自動轉寫為文本、文本入庫前自動脫敏 PII、FAQ 寫入即向量化,再經語義召回與大模型重排精準匹配標準答案,並對會話自動完成情感識別與工單分類。

方案概述

線上客服與電話客服每天沉澱的最大一筆資料資產,是海量通話錄音和語音信箱。一個中等規模的客服中心單日通話動輒數萬通、錄音時間長度以千小時計。這些語音裡有最真實的客戶訴求、坐席服務品質與潛在輿情風險,但多數團隊至今仍把它們當作合規留檔的冷資料。要把語音資產盤活,需要做到:

  • 錄音轉文字:把通話與語音信箱批量轉寫成可檢索、可分析的文本,這是後續一切工作的前提。

  • 建可檢索知識庫:把歷史優質應答、產品手冊、FAQ 沉澱成語義可檢索的知識庫,供機器人和坐席即時調用。

  • FAQ 智能匹配:使用者一句口語化的問題要能準確命中標準 FAQ,而不是靠關鍵詞碰運氣。

  • 情感與意圖識別:識別客戶在通話中的情緒波動,支撐即時預警與事後質檢。

  • 質檢與合規脫敏:轉寫文本常包含手機號、社會安全號碼、銀行卡號等個人敏感資訊,入庫、分析、共用前必須脫敏。

傳統方案要湊齊這套能力,通常需要串聯 ASR 服務、向量庫、外部 NLP 平台與重排服務四套以上系統。資料在多個系統間來回搬運,膠水代碼越寫越厚,任何一環抖動都會拖垮整條鏈路;原始轉寫文字資料流出到外部服務還存在個人敏感資訊泄露風險,而脫敏往往被放到鏈路末端甚至遺漏。

阿里雲 Milvus 的 AI Function 把這些能力收斂進同一個向量資料庫,模型推理在寫入與檢索時由 Milvus 內部自動觸發,資料全程不出執行個體。本文用到六個 Function:

Function

作用

在客服語音鏈路中的用途

AI_AUDIO_TRANSCRIBE

寫入時自動把錄音轉寫成文本,無需應用側先調 ASR。

把海量通話錄音變成可檢索、可分析的文本。

AI_PII_MASK

自動掩碼文本中的手機號、證件號、銀行卡號等敏感資訊。

讓進入知識庫與分析庫的都是已脫敏文本,合規前置。

AI_EMBEDDING

寫入時自動把 FAQ 文本轉成向量。

讓「執行個體怎麼連外網」這類口語化需求能被語義檢索命中。

AI_RERANK

對召回候選按與查詢的相關性重新排序。

把最貼合意圖的 FAQ 頂到首位,修正向量召回的雜訊。

AI_SENTIMENT

判斷會話情感傾向。

彙總情感標籤,支撐全量質檢與輿情預警。

AI_CLASSIFY

把會話歸類到工單類目。

支撐工單自動分類與路由。

前提條件

  • 已建立 Milvus 2.6 版本執行個體。AI Function 依賴 2.6 版本核心,建立後無需單獨綁定模型服務。

  • 如需從公網訪問執行個體,已在執行個體詳情頁的 安全配置 頁簽開啟 公網訪問 並將用戶端出口 IP 加入公網訪問白名單。

  • 已安裝 pymilvus,本文樣本基於 pymilvus 3.0.0 驗證。

  • 待轉寫的錄音已上傳到服務端可下載的真真實位址。

說明

RESTful 介面與 gRPC 共用 19530 連接埠,調用時必須顯式帶連接埠,例如 http://c-xxx.milvus.aliyuncs.com:19530;省略連接埠會預設訪問 80 連接埠並導致連線逾時。

警告

音頻地址必須是服務端能真實下載的地址,生產環境建議使用自有 OSS 的短時效簽名 URL。如果傳入的是預留位置或已失效的地址,會報 Failed to download multimodal content。

操作步驟

準備公用代碼

以下程式碼封裝含串連配置、REST 調用封裝以及 TEXTTRANSFORM 函數類型的相容兜底,供後續步驟共用。請將 MILVUS_URI 與 MILVUS_TOKEN 替換為實際執行個體資訊。

from __future__ import annotations

import json
import time
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen

from pymilvus import DataType, Function, FunctionType, MilvusClient

# ==================== 串連配置 ====================
MILVUS_URI = "http://c-xxx.milvus.aliyuncs.com:19530"  # 連接埠必須寫 19530
MILVUS_TOKEN = "root:xxx"
MILVUS_REST_BASE_URL = MILVUS_URI

client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)

# AI_AUDIO_TRANSCRIBE / AI_PII_MASK / AI_SENTIMENT / AI_CLASSIFY 都屬於 TEXTTRANSFORM
# 類型的 Function(函數類型值 9),通過 task 參數區分具體任務。
# 部分 pymilvus 版本的 FunctionType 枚舉中沒有該成員,下面做一次相容兜底。
TEXTTRANSFORM_FUNCTION_TYPE = 9


def texttransform_function_type() -> Any:
    for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
        ft = getattr(FunctionType, type_name, None)
        if ft is not None:
            return ft
    existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
    if existing is not None:
        return existing
    extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
    extension._name_ = "TEXTTRANSFORM"
    extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
    FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
    FunctionType._member_map_["TEXTTRANSFORM"] = extension
    return extension


def post_json(path: str, body: dict[str, Any], timeout: int = 120,
              retries: int = 3) -> tuple[int, dict[str, Any]]:
    """REST 即時介面統一封裝(走大模型,加重試更穩)。"""
    last: tuple[int, dict[str, Any]] | None = None
    for _ in range(retries):
        request = Request(
            f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
            data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
            headers={"Authorization": f"Bearer {MILVUS_TOKEN}",
                     "Content-Type": "application/json"},
            method="POST",
        )
        try:
            with urlopen(request, timeout=timeout) as response:
                status, data = response.status, json.loads(response.read().decode("utf-8"))
        except HTTPError as exc:
            status, data = exc.code, json.loads(exc.read().decode("utf-8"))
        last = (status, data)
        if status == 200 and data.get("code") == 0:
            return status, data
        time.sleep(1)
    return last
說明

AI_AUDIO_TRANSCRIBE、AI_PII_MASK、AI_SENTIMENT、AI_CLASSIFY 都屬於 TEXTTRANSFORM 類型的 Function(函數類型值 9),通過 task 參數區分具體任務;AI_EMBEDDING 與 AI_RERANK 則有獨立的 FunctionType 常量。

步驟一:錄音轉寫

把通話錄音批量轉寫成文本。將 AI_AUDIO_TRANSCRIBE 掛在 Collection 上後,寫入音頻地址時自動產出轉寫文本,無需應用側先調用 ASR 服務。

# ==================== 步驟 1:AI_AUDIO_TRANSCRIBE 錄音轉寫 ====================
collection_name = "cs_transcribe"
if client.has_collection(collection_name):
    client.drop_collection(collection_name)

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("audio_url", DataType.VARCHAR, max_length=4096)   # 音頻輸入欄位
schema.add_field("transcript", DataType.VARCHAR, max_length=4096)  # 轉寫輸出欄位
# Collection 必須至少包含一個向量欄位。本步驟不做向量檢索,用 2 維佔位欄位滿足約束;
# 聲明為 nullable=True 後,insert 時無需再傳該欄位。
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
    Function(
        name="transcribe_audio",
        function_type=texttransform_function_type(),
        input_field_names=["audio_url"],
        output_field_names=["transcript"],
        params={
            "provider": "aliyun_milvus",
            "model_name": "qwen3-asr-flash",
            "task": "ai_audio_transcribe",
            "language": "zh",
            "enable_itn": "true",     # 正常化口語數字,如「一三八」轉為 138
        },
    )
)

index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
                        index_params=index_params)

# audio_url 必須是服務端可下載的真真實位址,生產環境用自有 OSS 的短時效簽名 URL
audio_urls = [
    "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/calls/call_0001.mp3",
    "https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/calls/call_0002.wav",
]
client.insert(collection_name, [{"audio_url": u} for u in audio_urls])
client.flush(collection_name)

for row in client.query(collection_name, filter="",
                       output_fields=["audio_url", "transcript"], limit=10):
    print(f"{row['audio_url'].split('/')[-1]} -> {row['transcript']}")

enable_itn 開啟逆文本正常化,會把口語化的數字念法轉成規範寫法,便於後續檢索與結構化處理。

說明

Collection 必須至少包含一個向量欄位,否則建表報 schema does not contain vector field。本步驟不做向量檢索,因此用一個 2 維佔位欄位滿足約束。把它聲明為 nullable=True 後,insert 時就不必再傳該欄位;若不加 nullable,省略該欄位會報 Insert missed an field dummy_vector。

步驟二:PII 脫敏

轉寫文本常包含手機號、社會安全號碼、銀行卡號等個人敏感資訊。在文本進入知識庫或分析庫之前完成脫敏,把合規環節前置,而不是留到鏈路末端。提供兩種用法:REST 即時介面,以及寫入型 Collection。

# ==================== 步驟 2:AI_PII_MASK 脫敏 ====================
# 2.1 REST 即時介面:文本進庫前先脫敏
status, data = post_json(
    "/v2/vectordb/ai/pii_mask",
    {
        "model_name": "qwen3.7-max",
        "texts": ["您好,My Phone號是 13812345678,身份證 110101199003071234,麻煩幫我查下工單。"],
        "params": {
            "pii_types": ["PERSON", "PHONE", "ID_CARD"],
            "mask_char": "*",
            "preserve_length": True,   # 掩碼後保持原長度
            "temperature": 0,
        },
    },
)
assert status == 200 and data.get("code") == 0, data
for item in data["data"]["output"]["outputs"]:
    print(item)

# 2.2 寫入型 Collection:content 入庫時自動產生脫敏欄位 masked
collection_name = "cs_pii_mask"
if client.has_collection(collection_name):
    client.drop_collection(collection_name)

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("masked", DataType.VARCHAR, max_length=4096)
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
    Function(
        name="mask_pii",
        function_type=texttransform_function_type(),
        input_field_names=["content"],
        output_field_names=["masked"],
        params={
            "provider": "aliyun_milvus",
            "model_name": "qwen3.7-max",
            "task": "ai_pii_mask",
            "pii_types": "PERSON,PHONE,ID_CARD",
            "mask_char": "*",
            "preserve_length": "true",
            "temperature": "0",
        },
    )
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
                        index_params=index_params)

client.insert(collection_name,
              [{"content": "您好,My Phone號是 13812345678,麻煩幫我查下工單。"}])
client.flush(collection_name)
for row in client.query(collection_name, filter="", output_fields=["masked"], limit=1):
    print(row["masked"])

preserve_length 為 true 時掩碼後保持原字元長度,便於在不暴露原值的前提下保留格式特徵。實測輸出:

您好,My Phone號是 ***********,身份證 ******************,麻煩幫我查下工單。

11 位手機號掩碼為 11 個星號,18 位身份證掩碼為 18 個星號。兩種用法的區別是:REST 即時介面適合在資料流入前做一次性處理,寫入型 Collection 適合讓脫敏結果與原文一併留存、便於審計。

步驟三:建 FAQ 知識庫

把 FAQ 問題與標準答案寫入 Collection,AI_EMBEDDING 在寫入時自動完成向量化,應用側無需先調用 embedding 模型。

# ==================== 步驟 3:AI_EMBEDDING 建 FAQ 知識庫 ====================
collection_name = "cs_faq_kb"
if client.has_collection(collection_name):
    client.drop_collection(collection_name)

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)   # FAQ 問題文本
schema.add_field("answer", DataType.VARCHAR, max_length=4096)    # 標準答案
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=1024)
schema.add_function(
    Function(
        name="embed_content",
        function_type=FunctionType.TEXTEMBEDDING,
        input_field_names=["content"],
        output_field_names=["embedding"],
        params={
            "provider": "aliyun_milvus",
            "model_name": "text-embedding-v4",
            "dim": 1024,
            "max_client_batch_size": 10,
            "max_concurrency": 1,
        },
    )
)
index_params = client.prepare_index_params()
index_params.add_index(
    field_name="embedding",
    index_type="HNSW",
    metric_type="COSINE",
    params={"M": 16, "efConstruction": 200},
)
client.create_collection(collection_name=collection_name, schema=schema,
                        index_params=index_params)

faqs = [
    {"content": "如何開啟 Serverless Milvus 執行個體的公網訪問?",
     "answer": "在控制台執行個體詳情頁開啟公網並配置白名單。"},
    {"content": "忘記控制台登入密碼怎麼辦?", "answer": "通過帳號中心的找回密碼流程重設。"},
    {"content": "賬單為什麼比預期高?", "answer": "查看用量明細,重點關注計算與儲存用量。"},
    {"content": "如何建立 Collection 並寫入向量?",
     "answer": "使用 create_collection 定義 schema 後 insert。"},
    {"content": "執行個體擴容會影響線上業務嗎?", "answer": "擴容為線上操作,通常不中斷服務。"},
]
client.insert(collection_name, faqs)
client.flush(collection_name)
client.load_collection(collection_name)
說明

本例使用 text-embedding-v4(支援多語言與自訂維度)。AI 中心還提供 qwen3.7-text-embedding 等模型,可在控制台 AI 中心 的 模型服務 頁簽查看當前可用模型及其維度,按語種與效果需求選型。向量欄位的 dim 必須與 Function 參數中的 dim 一致。

步驟四與步驟五:語義召回與大模型重排

使用者的口語化問題先經向量檢索召回 Top-N 候選,再由重排模型按與查詢的真實相關性精排。這兩步合起來才能穩定命中標準 FAQ。

# ==================== 步驟 4:語義檢索召回 FAQ ====================
query = "我的執行個體怎麼才能讓外網連上?"
result = client.search(
    collection_name="cs_faq_kb",
    data=[query],
    anns_field="embedding",
    limit=3,
    output_fields=["content", "answer"],
)
for rank, hit in enumerate(result[0], 1):
    print(f"{rank}. [相似性 {hit['distance']:.4f}] {hit['entity']['content']}")


# ==================== 步驟 5:AI_RERANK 重排 ====================
# 5.1 REST 即時介面:對給定的一批候選獨立重排
status, data = post_json(
    "/v2/vectordb/ai/rerank",
    {
        "model_name": "qwen3-rerank",
        "query": query,
        "documents": [
            "如何開啟 Serverless Milvus 執行個體的公網訪問?",
            "執行個體擴容會影響線上業務嗎?",
            "如何建立 Collection 並寫入向量?",
        ],
        "params": {"max_concurrency": 2, "timeout_sec": 10},
    },
)
assert status == 200 and data.get("code") == 0, data
ranked = sorted(data["data"]["output"]["results"],
                key=lambda x: x["relevance_score"], reverse=True)
for rank, item in enumerate(ranked, 1):
    print(f"{rank}. [相關度 {item['relevance_score']:.4f}] 候選 index={item['index']}")

# 5.2 在 search 時掛載 ranker,召回與重排一次完成
reranker = Function(
    name="rerank_faq",
    function_type=FunctionType.RERANK,
    input_field_names=["content"],
    params={
        "reranker": "model",
        "provider": "aliyun_milvus",
        "model_name": "qwen3-rerank",
        "queries": [query],
        "max_concurrency": 2,
        "timeout_sec": 10,
    },
)
result = client.search(
    collection_name="cs_faq_kb",
    data=[query],
    anns_field="embedding",
    limit=3,
    output_fields=["content", "answer"],
    ranker=reranker,
)
for rank, hit in enumerate(result[0], 1):
    e = hit["entity"]
    print(f"{rank}. [重排分 {hit['distance']:.4f}] {e['content']} | 答:{e['answer']}")

以查詢「我的執行個體怎麼才能讓外網連上?」為例,實測資料如下。這組前後對比很能說明重排的價值:

候選 FAQ

步驟四 向量召回

步驟五 重排後

如何開啟 Serverless Milvus 執行個體的公網訪問?

① 0.6141

① 0.5883

忘記控制台登入密碼怎麼辦?

② 0.5613

③ 0.2607

執行個體擴容會影響線上業務嗎?

③ 0.4757

② 0.3356

向量召回把「忘記控制台登入密碼怎麼辦?」排到了第二位(0.5613),但它與「讓外網連上」在意圖上並不相關——這正是純向量檢索「召回但不精準」的典型表現:文本表面相似性不低,實際意圖相去甚遠。經重排後該條被壓到第三位(0.2607),語義更接近的「執行個體擴容」升到第二位。對客服機器人而言,Top-1 的準確性直接決定應答品質,因此重排環節不可省略。

說明

兩種重排用法的區別:REST 即時介面適合已經拿到候選列表、只需重排的情境;在 search 中掛載 ranker 則把召回與精排合并為一次調用,鏈路更短。REST 介面不支援 top_n,會為每個候選各返回一個得分,需要截斷時在應用側按得分排序後自行取前 N 條。

說明

純文字重排不需要 is_multimodal 參數。該參數僅在多模態重排(用 qwen3-vl-rerank 對圖片、視頻候選打分)時才需要,對純文字候選傳入不會改變打分結果。

步驟六:情感識別與工單分類

對脫敏後的會話文本同時掛載兩個 Function:一個判斷情感傾向用於質檢與輿情預警,一個歸類到工單類目用於自動分流。兩者都在寫入時自動完成。

# ==================== 步驟 6:AI_SENTIMENT 情感識別 + AI_CLASSIFY 工單分類 ====================
collection_name = "cs_qc"
if client.has_collection(collection_name):
    client.drop_collection(collection_name)

schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)     # 脫敏後的會話文本
schema.add_field("sentiment", DataType.VARCHAR, max_length=64)     # 情感輸出
schema.add_field("category", DataType.VARCHAR, max_length=64)      # 工單類目輸出
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
    Function(
        name="analyze_sentiment",
        function_type=texttransform_function_type(),
        input_field_names=["content"], output_field_names=["sentiment"],
        params={"provider": "aliyun_milvus", "model_name": "qwen3.7-max",
                "task": "ai_sentiment", "categories": "positive,negative,neutral",
                "temperature": "0"},
    )
)
schema.add_function(
    Function(
        name="classify_ticket",
        function_type=texttransform_function_type(),
        input_field_names=["content"], output_field_names=["category"],
        params={"provider": "aliyun_milvus", "model_name": "qwen3.7-max",
                "task": "ai_classify", "labels": "帳號,諮詢,故障,計費",
                "prompt": "根據客戶問題所屬主題分類。", "temperature": "0"},
    )
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
                        index_params=index_params)

client.insert(collection_name, [
    {"content": "你們這個問題我已經打了三次電話了,到現在還沒解決,太讓人失望了!"},
    {"content": "請問 Serverless Milvus 執行個體怎麼開啟公網訪問?"},
])
client.flush(collection_name)
for row in client.query(collection_name, filter="",
                        output_fields=["content", "sentiment", "category"], limit=10):
    print(f"情感={row['sentiment']:<10} 類目={row['category']:<6} | {row['content'][:28]}")

實測結果:

會話文本

情感

工單類目

你們這個問題我已經打了三次電話了,到現在還沒解決,太讓人失望了!

negative

故障

請問 Serverless Milvus 執行個體怎麼開啟公網訪問?

neutral

諮詢

一個 Collection 上可以掛載多個 Function,只要它們的輸出欄位不衝突。本例中 sentiment 與 category 由兩個 Function 各自填充,寫入一條資料即同時得到情感標籤與工單類目。

警告

情感與分類結果是模型判斷而非事實認定。在工單升級、服務下架等高影響動作前,應保留人工複核環節。

方案價值

維度

傳統方案(ASR + 向量庫 + NLP + 重排服務)

阿里雲 Milvus

系統數量

4 套以上,資料多處搬運

1 套,資料全程不出執行個體

錄音轉寫

應用側先調 ASR 再寫庫

寫入即轉寫

向量化鏈路

應用側調 embedding 再寫入

寫入即向量化

PII 脫敏

常被放在鏈路末端,容易遺漏

入庫即脫敏,合規前置

召回與重排

額外部署重排模型服務

AI_RERANK 內建,可一次 search 完成

情感與分類

外接 NLP 平台

AI_SENTIMENT / AI_CLASSIFY 內建

可延伸的情境包括:

  • 坐席即時輔助:把召回與重排接入坐席工作台,通話進行中即推送最貼切的話術與知識。

  • 批量質檢:歷史錄音可結合 AI_BATCH 做離線全量轉寫、脫敏與情感分類打標,把質檢覆蓋率從抽檢提升到全量。

  • 輿情預警:對 AI_SENTIMENT 輸出為 negative 的工單即時觸發預警與升級流程。

合規注意事項

  • 錄音處理前務必完成錄音告知與客戶授權。

  • 媒體一律使用短時效、最小許可權的簽名 URL,不要在 URL 中嵌入長期憑據。

  • 為原始音頻、轉寫文本、情感標籤與複核記錄設定明確的保留期限,到期刪除或匿名化;並保證派生資料隨來源資料的刪除或授權撤回而同步處置。

  • 情感與分類結果是模型判斷而非事實認定,在下架、升級等高影響動作前應保留人工複核環節。