本文介紹如何用阿里雲 Milvus 的六個 AI Function 串聯一條客服語音處理鏈路:錄音寫入時自動轉寫為文本、文本入庫前自動脫敏 PII、FAQ 寫入即向量化,再經語義召回與大模型重排精準匹配標準答案,並對會話自動完成情感識別與工單分類。
方案概述
線上客服與電話客服每天沉澱的最大一筆資料資產,是海量通話錄音和語音信箱。一個中等規模的客服中心單日通話動輒數萬通、錄音時間長度以千小時計。這些語音裡有最真實的客戶訴求、坐席服務品質與潛在輿情風險,但多數團隊至今仍把它們當作合規留檔的冷資料。要把語音資產盤活,需要做到:
錄音轉文字:把通話與語音信箱批量轉寫成可檢索、可分析的文本,這是後續一切工作的前提。
建可檢索知識庫:把歷史優質應答、產品手冊、FAQ 沉澱成語義可檢索的知識庫,供機器人和坐席即時調用。
FAQ 智能匹配:使用者一句口語化的問題要能準確命中標準 FAQ,而不是靠關鍵詞碰運氣。
情感與意圖識別:識別客戶在通話中的情緒波動,支撐即時預警與事後質檢。
質檢與合規脫敏:轉寫文本常包含手機號、社會安全號碼、銀行卡號等個人敏感資訊,入庫、分析、共用前必須脫敏。
傳統方案要湊齊這套能力,通常需要串聯 ASR 服務、向量庫、外部 NLP 平台與重排服務四套以上系統。資料在多個系統間來回搬運,膠水代碼越寫越厚,任何一環抖動都會拖垮整條鏈路;原始轉寫文字資料流出到外部服務還存在個人敏感資訊泄露風險,而脫敏往往被放到鏈路末端甚至遺漏。
阿里雲 Milvus 的 AI Function 把這些能力收斂進同一個向量資料庫,模型推理在寫入與檢索時由 Milvus 內部自動觸發,資料全程不出執行個體。本文用到六個 Function:
Function | 作用 | 在客服語音鏈路中的用途 |
| 寫入時自動把錄音轉寫成文本,無需應用側先調 ASR。 | 把海量通話錄音變成可檢索、可分析的文本。 |
| 自動掩碼文本中的手機號、證件號、銀行卡號等敏感資訊。 | 讓進入知識庫與分析庫的都是已脫敏文本,合規前置。 |
| 寫入時自動把 FAQ 文本轉成向量。 | 讓「執行個體怎麼連外網」這類口語化需求能被語義檢索命中。 |
| 對召回候選按與查詢的相關性重新排序。 | 把最貼合意圖的 FAQ 頂到首位,修正向量召回的雜訊。 |
| 判斷會話情感傾向。 | 彙總情感標籤,支撐全量質檢與輿情預警。 |
| 把會話歸類到工單類目。 | 支撐工單自動分類與路由。 |
前提條件
已建立 Milvus 2.6 版本執行個體。AI Function 依賴 2.6 版本核心,建立後無需單獨綁定模型服務。
如需從公網訪問執行個體,已在執行個體詳情頁的 安全配置 頁簽開啟 公網訪問 並將用戶端出口 IP 加入公網訪問白名單。
已安裝 pymilvus,本文樣本基於 pymilvus 3.0.0 驗證。
待轉寫的錄音已上傳到服務端可下載的真真實位址。
RESTful 介面與 gRPC 共用 19530 連接埠,調用時必須顯式帶連接埠,例如 http://c-xxx.milvus.aliyuncs.com:19530;省略連接埠會預設訪問 80 連接埠並導致連線逾時。
音頻地址必須是服務端能真實下載的地址,生產環境建議使用自有 OSS 的短時效簽名 URL。如果傳入的是預留位置或已失效的地址,會報 Failed to download multimodal content。
操作步驟
準備公用代碼
以下程式碼封裝含串連配置、REST 調用封裝以及 TEXTTRANSFORM 函數類型的相容兜底,供後續步驟共用。請將 MILVUS_URI 與 MILVUS_TOKEN 替換為實際執行個體資訊。
from __future__ import annotations
import json
import time
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
from pymilvus import DataType, Function, FunctionType, MilvusClient
# ==================== 串連配置 ====================
MILVUS_URI = "http://c-xxx.milvus.aliyuncs.com:19530" # 連接埠必須寫 19530
MILVUS_TOKEN = "root:xxx"
MILVUS_REST_BASE_URL = MILVUS_URI
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
# AI_AUDIO_TRANSCRIBE / AI_PII_MASK / AI_SENTIMENT / AI_CLASSIFY 都屬於 TEXTTRANSFORM
# 類型的 Function(函數類型值 9),通過 task 參數區分具體任務。
# 部分 pymilvus 版本的 FunctionType 枚舉中沒有該成員,下面做一次相容兜底。
TEXTTRANSFORM_FUNCTION_TYPE = 9
def texttransform_function_type() -> Any:
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
ft = getattr(FunctionType, type_name, None)
if ft is not None:
return ft
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
def post_json(path: str, body: dict[str, Any], timeout: int = 120,
retries: int = 3) -> tuple[int, dict[str, Any]]:
"""REST 即時介面統一封裝(走大模型,加重試更穩)。"""
last: tuple[int, dict[str, Any]] | None = None
for _ in range(retries):
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_TOKEN}",
"Content-Type": "application/json"},
method="POST",
)
try:
with urlopen(request, timeout=timeout) as response:
status, data = response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
status, data = exc.code, json.loads(exc.read().decode("utf-8"))
last = (status, data)
if status == 200 and data.get("code") == 0:
return status, data
time.sleep(1)
return lastAI_AUDIO_TRANSCRIBE、AI_PII_MASK、AI_SENTIMENT、AI_CLASSIFY 都屬於 TEXTTRANSFORM 類型的 Function(函數類型值 9),通過 task 參數區分具體任務;AI_EMBEDDING 與 AI_RERANK 則有獨立的 FunctionType 常量。
步驟一:錄音轉寫
把通話錄音批量轉寫成文本。將 AI_AUDIO_TRANSCRIBE 掛在 Collection 上後,寫入音頻地址時自動產出轉寫文本,無需應用側先調用 ASR 服務。
# ==================== 步驟 1:AI_AUDIO_TRANSCRIBE 錄音轉寫 ====================
collection_name = "cs_transcribe"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("audio_url", DataType.VARCHAR, max_length=4096) # 音頻輸入欄位
schema.add_field("transcript", DataType.VARCHAR, max_length=4096) # 轉寫輸出欄位
# Collection 必須至少包含一個向量欄位。本步驟不做向量檢索,用 2 維佔位欄位滿足約束;
# 聲明為 nullable=True 後,insert 時無需再傳該欄位。
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
Function(
name="transcribe_audio",
function_type=texttransform_function_type(),
input_field_names=["audio_url"],
output_field_names=["transcript"],
params={
"provider": "aliyun_milvus",
"model_name": "qwen3-asr-flash",
"task": "ai_audio_transcribe",
"language": "zh",
"enable_itn": "true", # 正常化口語數字,如「一三八」轉為 138
},
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)
# audio_url 必須是服務端可下載的真真實位址,生產環境用自有 OSS 的短時效簽名 URL
audio_urls = [
"https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/calls/call_0001.mp3",
"https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/calls/call_0002.wav",
]
client.insert(collection_name, [{"audio_url": u} for u in audio_urls])
client.flush(collection_name)
for row in client.query(collection_name, filter="",
output_fields=["audio_url", "transcript"], limit=10):
print(f"{row['audio_url'].split('/')[-1]} -> {row['transcript']}")enable_itn 開啟逆文本正常化,會把口語化的數字念法轉成規範寫法,便於後續檢索與結構化處理。
Collection 必須至少包含一個向量欄位,否則建表報 schema does not contain vector field。本步驟不做向量檢索,因此用一個 2 維佔位欄位滿足約束。把它聲明為 nullable=True 後,insert 時就不必再傳該欄位;若不加 nullable,省略該欄位會報 Insert missed an field dummy_vector。
步驟二:PII 脫敏
轉寫文本常包含手機號、社會安全號碼、銀行卡號等個人敏感資訊。在文本進入知識庫或分析庫之前完成脫敏,把合規環節前置,而不是留到鏈路末端。提供兩種用法:REST 即時介面,以及寫入型 Collection。
# ==================== 步驟 2:AI_PII_MASK 脫敏 ====================
# 2.1 REST 即時介面:文本進庫前先脫敏
status, data = post_json(
"/v2/vectordb/ai/pii_mask",
{
"model_name": "qwen3.7-max",
"texts": ["您好,My Phone號是 13812345678,身份證 110101199003071234,麻煩幫我查下工單。"],
"params": {
"pii_types": ["PERSON", "PHONE", "ID_CARD"],
"mask_char": "*",
"preserve_length": True, # 掩碼後保持原長度
"temperature": 0,
},
},
)
assert status == 200 and data.get("code") == 0, data
for item in data["data"]["output"]["outputs"]:
print(item)
# 2.2 寫入型 Collection:content 入庫時自動產生脫敏欄位 masked
collection_name = "cs_pii_mask"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("masked", DataType.VARCHAR, max_length=4096)
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
Function(
name="mask_pii",
function_type=texttransform_function_type(),
input_field_names=["content"],
output_field_names=["masked"],
params={
"provider": "aliyun_milvus",
"model_name": "qwen3.7-max",
"task": "ai_pii_mask",
"pii_types": "PERSON,PHONE,ID_CARD",
"mask_char": "*",
"preserve_length": "true",
"temperature": "0",
},
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)
client.insert(collection_name,
[{"content": "您好,My Phone號是 13812345678,麻煩幫我查下工單。"}])
client.flush(collection_name)
for row in client.query(collection_name, filter="", output_fields=["masked"], limit=1):
print(row["masked"])preserve_length 為 true 時掩碼後保持原字元長度,便於在不暴露原值的前提下保留格式特徵。實測輸出:
您好,My Phone號是 ***********,身份證 ******************,麻煩幫我查下工單。11 位手機號掩碼為 11 個星號,18 位身份證掩碼為 18 個星號。兩種用法的區別是:REST 即時介面適合在資料流入前做一次性處理,寫入型 Collection 適合讓脫敏結果與原文一併留存、便於審計。
步驟三:建 FAQ 知識庫
把 FAQ 問題與標準答案寫入 Collection,AI_EMBEDDING 在寫入時自動完成向量化,應用側無需先調用 embedding 模型。
# ==================== 步驟 3:AI_EMBEDDING 建 FAQ 知識庫 ====================
collection_name = "cs_faq_kb"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096) # FAQ 問題文本
schema.add_field("answer", DataType.VARCHAR, max_length=4096) # 標準答案
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=1024)
schema.add_function(
Function(
name="embed_content",
function_type=FunctionType.TEXTEMBEDDING,
input_field_names=["content"],
output_field_names=["embedding"],
params={
"provider": "aliyun_milvus",
"model_name": "text-embedding-v4",
"dim": 1024,
"max_client_batch_size": 10,
"max_concurrency": 1,
},
)
)
index_params = client.prepare_index_params()
index_params.add_index(
field_name="embedding",
index_type="HNSW",
metric_type="COSINE",
params={"M": 16, "efConstruction": 200},
)
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)
faqs = [
{"content": "如何開啟 Serverless Milvus 執行個體的公網訪問?",
"answer": "在控制台執行個體詳情頁開啟公網並配置白名單。"},
{"content": "忘記控制台登入密碼怎麼辦?", "answer": "通過帳號中心的找回密碼流程重設。"},
{"content": "賬單為什麼比預期高?", "answer": "查看用量明細,重點關注計算與儲存用量。"},
{"content": "如何建立 Collection 並寫入向量?",
"answer": "使用 create_collection 定義 schema 後 insert。"},
{"content": "執行個體擴容會影響線上業務嗎?", "answer": "擴容為線上操作,通常不中斷服務。"},
]
client.insert(collection_name, faqs)
client.flush(collection_name)
client.load_collection(collection_name)本例使用 text-embedding-v4(支援多語言與自訂維度)。AI 中心還提供 qwen3.7-text-embedding 等模型,可在控制台 AI 中心 的 模型服務 頁簽查看當前可用模型及其維度,按語種與效果需求選型。向量欄位的 dim 必須與 Function 參數中的 dim 一致。
步驟四與步驟五:語義召回與大模型重排
使用者的口語化問題先經向量檢索召回 Top-N 候選,再由重排模型按與查詢的真實相關性精排。這兩步合起來才能穩定命中標準 FAQ。
# ==================== 步驟 4:語義檢索召回 FAQ ====================
query = "我的執行個體怎麼才能讓外網連上?"
result = client.search(
collection_name="cs_faq_kb",
data=[query],
anns_field="embedding",
limit=3,
output_fields=["content", "answer"],
)
for rank, hit in enumerate(result[0], 1):
print(f"{rank}. [相似性 {hit['distance']:.4f}] {hit['entity']['content']}")
# ==================== 步驟 5:AI_RERANK 重排 ====================
# 5.1 REST 即時介面:對給定的一批候選獨立重排
status, data = post_json(
"/v2/vectordb/ai/rerank",
{
"model_name": "qwen3-rerank",
"query": query,
"documents": [
"如何開啟 Serverless Milvus 執行個體的公網訪問?",
"執行個體擴容會影響線上業務嗎?",
"如何建立 Collection 並寫入向量?",
],
"params": {"max_concurrency": 2, "timeout_sec": 10},
},
)
assert status == 200 and data.get("code") == 0, data
ranked = sorted(data["data"]["output"]["results"],
key=lambda x: x["relevance_score"], reverse=True)
for rank, item in enumerate(ranked, 1):
print(f"{rank}. [相關度 {item['relevance_score']:.4f}] 候選 index={item['index']}")
# 5.2 在 search 時掛載 ranker,召回與重排一次完成
reranker = Function(
name="rerank_faq",
function_type=FunctionType.RERANK,
input_field_names=["content"],
params={
"reranker": "model",
"provider": "aliyun_milvus",
"model_name": "qwen3-rerank",
"queries": [query],
"max_concurrency": 2,
"timeout_sec": 10,
},
)
result = client.search(
collection_name="cs_faq_kb",
data=[query],
anns_field="embedding",
limit=3,
output_fields=["content", "answer"],
ranker=reranker,
)
for rank, hit in enumerate(result[0], 1):
e = hit["entity"]
print(f"{rank}. [重排分 {hit['distance']:.4f}] {e['content']} | 答:{e['answer']}")以查詢「我的執行個體怎麼才能讓外網連上?」為例,實測資料如下。這組前後對比很能說明重排的價值:
候選 FAQ | 步驟四 向量召回 | 步驟五 重排後 |
如何開啟 Serverless Milvus 執行個體的公網訪問? | ① 0.6141 | ① 0.5883 |
忘記控制台登入密碼怎麼辦? | ② 0.5613 | ③ 0.2607 |
執行個體擴容會影響線上業務嗎? | ③ 0.4757 | ② 0.3356 |
向量召回把「忘記控制台登入密碼怎麼辦?」排到了第二位(0.5613),但它與「讓外網連上」在意圖上並不相關——這正是純向量檢索「召回但不精準」的典型表現:文本表面相似性不低,實際意圖相去甚遠。經重排後該條被壓到第三位(0.2607),語義更接近的「執行個體擴容」升到第二位。對客服機器人而言,Top-1 的準確性直接決定應答品質,因此重排環節不可省略。
兩種重排用法的區別:REST 即時介面適合已經拿到候選列表、只需重排的情境;在 search 中掛載 ranker 則把召回與精排合并為一次調用,鏈路更短。REST 介面不支援 top_n,會為每個候選各返回一個得分,需要截斷時在應用側按得分排序後自行取前 N 條。
純文字重排不需要 is_multimodal 參數。該參數僅在多模態重排(用 qwen3-vl-rerank 對圖片、視頻候選打分)時才需要,對純文字候選傳入不會改變打分結果。
步驟六:情感識別與工單分類
對脫敏後的會話文本同時掛載兩個 Function:一個判斷情感傾向用於質檢與輿情預警,一個歸類到工單類目用於自動分流。兩者都在寫入時自動完成。
# ==================== 步驟 6:AI_SENTIMENT 情感識別 + AI_CLASSIFY 工單分類 ====================
collection_name = "cs_qc"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096) # 脫敏後的會話文本
schema.add_field("sentiment", DataType.VARCHAR, max_length=64) # 情感輸出
schema.add_field("category", DataType.VARCHAR, max_length=64) # 工單類目輸出
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
Function(
name="analyze_sentiment",
function_type=texttransform_function_type(),
input_field_names=["content"], output_field_names=["sentiment"],
params={"provider": "aliyun_milvus", "model_name": "qwen3.7-max",
"task": "ai_sentiment", "categories": "positive,negative,neutral",
"temperature": "0"},
)
)
schema.add_function(
Function(
name="classify_ticket",
function_type=texttransform_function_type(),
input_field_names=["content"], output_field_names=["category"],
params={"provider": "aliyun_milvus", "model_name": "qwen3.7-max",
"task": "ai_classify", "labels": "帳號,諮詢,故障,計費",
"prompt": "根據客戶問題所屬主題分類。", "temperature": "0"},
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)
client.insert(collection_name, [
{"content": "你們這個問題我已經打了三次電話了,到現在還沒解決,太讓人失望了!"},
{"content": "請問 Serverless Milvus 執行個體怎麼開啟公網訪問?"},
])
client.flush(collection_name)
for row in client.query(collection_name, filter="",
output_fields=["content", "sentiment", "category"], limit=10):
print(f"情感={row['sentiment']:<10} 類目={row['category']:<6} | {row['content'][:28]}")實測結果:
會話文本 | 情感 | 工單類目 |
你們這個問題我已經打了三次電話了,到現在還沒解決,太讓人失望了! | negative | 故障 |
請問 Serverless Milvus 執行個體怎麼開啟公網訪問? | neutral | 諮詢 |
一個 Collection 上可以掛載多個 Function,只要它們的輸出欄位不衝突。本例中 sentiment 與 category 由兩個 Function 各自填充,寫入一條資料即同時得到情感標籤與工單類目。
情感與分類結果是模型判斷而非事實認定。在工單升級、服務下架等高影響動作前,應保留人工複核環節。
方案價值
維度 | 傳統方案(ASR + 向量庫 + NLP + 重排服務) | 阿里雲 Milvus |
系統數量 | 4 套以上,資料多處搬運 | 1 套,資料全程不出執行個體 |
錄音轉寫 | 應用側先調 ASR 再寫庫 | 寫入即轉寫 |
向量化鏈路 | 應用側調 embedding 再寫入 | 寫入即向量化 |
PII 脫敏 | 常被放在鏈路末端,容易遺漏 | 入庫即脫敏,合規前置 |
召回與重排 | 額外部署重排模型服務 |
|
情感與分類 | 外接 NLP 平台 |
|
可延伸的情境包括:
坐席即時輔助:把召回與重排接入坐席工作台,通話進行中即推送最貼切的話術與知識。
批量質檢:歷史錄音可結合
AI_BATCH做離線全量轉寫、脫敏與情感分類打標,把質檢覆蓋率從抽檢提升到全量。輿情預警:對
AI_SENTIMENT輸出為negative的工單即時觸發預警與升級流程。
合規注意事項
錄音處理前務必完成錄音告知與客戶授權。
媒體一律使用短時效、最小許可權的簽名 URL,不要在 URL 中嵌入長期憑據。
為原始音頻、轉寫文本、情感標籤與複核記錄設定明確的保留期限,到期刪除或匿名化;並保證派生資料隨來源資料的刪除或授權撤回而同步處置。
情感與分類結果是模型判斷而非事實認定,在下架、升級等高影響動作前應保留人工複核環節。