In this tutorial, you chain the AI Functions of Alibaba Cloud Milvus into a contact center voice pipeline that runs entirely inside one vector database. Recordings are transcribed on write, text is masked for personally identifiable information (PII) before it enters the database, FAQs are vectorized on write, and semantic retrieval plus reranking match a colloquial question to the correct standard answer. By the end, each conversation you write also carries a sentiment label and a ticket category.
Solution overview
Contact center teams accumulate call recordings and voice messages every day, yet most teams keep this data only for compliance archiving. Recordings contain customer requests, evidence of agent service quality, and early signals of negative feedback. Activating this voice asset requires the following capabilities:
Speech-to-text transcription — Transcribe calls and voice messages in bulk into text that can be searched and analyzed. Everything else depends on this step.
A searchable knowledge base — Turn high-quality historical replies, product manuals, and FAQs into a semantically searchable knowledge base that bots and agents can query in real time.
Intelligent FAQ matching — Match a colloquial customer question to the correct standard FAQ, instead of relying on keyword luck.
Sentiment and intent recognition — Detect emotional shifts during a call to support real-time alerting and post-call quality inspection.
Quality inspection and compliance masking — Transcribed text often contains PII such as mobile numbers, ID card numbers, and bank card numbers. Mask this data before storage, analysis, or sharing.
Assembling this capability set with a traditional approach usually means chaining four or more systems: an ASR service, a vector database, an external NLP platform, and a reranking service. Glue code keeps growing, and instability in any single hop drags down the whole pipeline. Sending raw transcripts to external services also risks exposing PII.
The AI Functions of Alibaba Cloud Milvus consolidate these capabilities into a single vector database. Milvus triggers model inference internally on write and on search, and data never leaves the instance. The following table describes the six AI Functions that this tutorial uses.
| Function | Purpose | Role in the contact center voice pipeline |
AI_AUDIO_TRANSCRIBE | Transcribes recordings into text on write, with no ASR call from the application. | Turns a mass of call recordings into text that can be searched and analyzed. |
AI_PII_MASK | Masks PII such as mobile numbers, ID card numbers, and bank card numbers in text. | Ensures that only masked text enters the knowledge base and the analytics store, so compliance is enforced up front. |
AI_EMBEDDING | Converts FAQ text into vectors on write. | Lets colloquial requests such as "How can I get my instance connected from the internet?" be matched by semantic retrieval. |
AI_RERANK | Reorders retrieved candidates by their relevance to the query. | Promotes the FAQ that best fits the intent to the first position and corrects noise from vector retrieval. |
AI_SENTIMENT | Determines the sentiment of a conversation. | Aggregates sentiment labels to support full-coverage quality inspection and negative-feedback alerting. |
AI_CLASSIFY | Assigns a conversation to a ticket category. | Supports automatic ticket classification and routing. |
How the AI Functions compare with a multi-system stack
The following table compares this pipeline with an equivalent pipeline assembled from separate services.
| Dimension | Traditional approach (ASR + vector database + NLP + reranking service) | Alibaba Cloud Milvus |
| Number of systems | Four or more, with data moved between them | One, with data never leaving the instance |
| Transcription | The application calls ASR first, then writes to the database | Transcription on write |
| Vectorization pipeline | The application calls embedding first, then writes | Vectorization on write |
| PII masking | Often placed at the end of the pipeline and easily skipped | Masking on write, with compliance enforced up front |
| Retrieval and reranking | A separate reranking model service must be deployed | AI_RERANK is built in and can run within a single search |
| Sentiment and classification | An external NLP platform is required | AI_SENTIMENT and AI_CLASSIFY are built in |
Prerequisites
Milvus instance — A Milvus 2.6 instance. AI Functions require the 2.6 kernel, and no separate model service binding is needed after creation.
Internet access — To access the instance over the internet, enable Public Endpoint on the Security Configuration tab of the instance details page, and add the client egress IP address to the public access whitelist.
Endpoint port — Specify port 19530 explicitly in the instance URI, for example
http://c-xxx.milvus.aliyuncs.com:19530. The RESTful API and gRPC share this port. If you omit the port, the connection falls back to port 80 and times out.Client library — pymilvus installed. The examples in this topic are verified with pymilvus 3.0.0.
Audio source — The recordings to transcribe uploaded to an address that the server can actually download from. In a production environment, use a short-lived signed URL from your own Object Storage Service (OSS) bucket. A placeholder or an expired address fails at transcription time.
Media credentials — Short-lived, least-privilege signed URLs for all media. Never embed long-term credentials in a URL.
Recording consent — Recording disclosure and customer consent completed before you process any recording.
Prepare the shared code
The following code contains the connection settings, a REST call wrapper, and a compatibility fallback for the TEXTTRANSFORM function type. Every step reuses it. This section is preparation only and does not consume one of the six step numbers.
Replace MILVUS_URI and MILVUS_TOKEN with the values of your instance.
from __future__ import annotations
import json
import time
from typing import Any
from urllib.error import HTTPError
from urllib.request import Request, urlopen
from pymilvus import DataType, Function, FunctionType, MilvusClient
# ==================== Connection settings ====================
MILVUS_URI = "http://c-xxx.milvus.aliyuncs.com:19530" # the port must be 19530
MILVUS_TOKEN = "root:xxx"
MILVUS_REST_BASE_URL = MILVUS_URI
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
# AI_AUDIO_TRANSCRIBE / AI_PII_MASK / AI_SENTIMENT / AI_CLASSIFY are all Functions of the
# TEXTTRANSFORM type (function type value 9) and are distinguished by the task parameter.
# The FunctionType enum of some pymilvus versions has no such member, so add a fallback.
TEXTTRANSFORM_FUNCTION_TYPE = 9
def texttransform_function_type() -> Any:
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
ft = getattr(FunctionType, type_name, None)
if ft is not None:
return ft
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
def post_json(path: str, body: dict[str, Any], timeout: int = 120,
retries: int = 3) -> tuple[int, dict[str, Any]]:
"""Shared wrapper for the synchronous REST APIs. These calls go through an LLM, so retries make them more stable."""
last: tuple[int, dict[str, Any]] | None = None
for _ in range(retries):
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_TOKEN}",
"Content-Type": "application/json"},
method="POST",
)
try:
with urlopen(request, timeout=timeout) as response:
status, data = response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
status, data = exc.code, json.loads(exc.read().decode("utf-8"))
last = (status, data)
if status == 200 and data.get("code") == 0:
return status, data
time.sleep(1)
return lastKeep the following properties of the shared code in mind before you run any step:
Function types —
AI_AUDIO_TRANSCRIBE,AI_PII_MASK,AI_SENTIMENT, andAI_CLASSIFYare all functions of theTEXTTRANSFORMtype and are distinguished by thetaskparameter, whereasAI_EMBEDDINGandAI_RERANKhave their ownFunctionTypeconstants.Retry scope —
post_jsonretries a call up to three times when the response is an HTTP error or the payload returns a non-zerocode. Read timeouts and network errors are not caught and propagate to the caller.Step independence — Each step creates its own collection and runs on its own. To chain the steps, feed the output field of one step into the input field of the next: the
transcriptfield from Step 1 becomes the masking input in Step 2, and the masked text becomes thecontentinput in Step 6.
Every step drops any existing collection with the same name before it creates a new one. If your instance already contains a collection named cs_transcribe, cs_pii_mask, cs_faq_kb, or cs_qc, that collection and its data are deleted. Use collection names that are not in use in your instance.
Step 1: Transcribe call recordings
Transcribe call recordings into text in bulk. After AI_AUDIO_TRANSCRIBE is attached to a collection, writing an audio address produces the transcript automatically, with no ASR call from the application.
# ==================== Step 1: AI_AUDIO_TRANSCRIBE transcription ====================
collection_name = "cs_transcribe"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("audio_url", DataType.VARCHAR, max_length=4096) # audio input field
schema.add_field("transcript", DataType.VARCHAR, max_length=4096) # transcript output field
# A Collection must contain at least one vector field. This step performs no vector search,
# so a 2-dimensional placeholder field satisfies the constraint. Declaring it as
# nullable=True means the field can be omitted at insert time.
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
Function(
name="transcribe_audio",
function_type=texttransform_function_type(),
input_field_names=["audio_url"],
output_field_names=["transcript"],
params={
"provider": "aliyun_milvus",
"model_name": "qwen3-asr-flash",
"task": "ai_audio_transcribe",
"language": "zh",
"enable_itn": "true", # normalize spoken numbers, such as one-three-eight to 138
},
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)
# audio_url must be a real address the server can download from. In production, use a
# short-lived signed URL from your own OSS bucket.
audio_urls = [
"https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/calls/call_0001.mp3",
"https://<your-bucket>.oss-cn-hangzhou.aliyuncs.com/calls/call_0002.wav",
]
client.insert(collection_name, [{"audio_url": u} for u in audio_urls])
client.flush(collection_name)
for row in client.query(collection_name, filter="",
output_fields=["audio_url", "transcript"], limit=10):
print(f"{row['audio_url'].split('/')[-1]} -> {row['transcript']}")Two parameters in this example control transcription behavior:
language— This example is set tozhand therefore transcribes Chinese-language recordings.
The placeholder vector field and itsenable_itn— Turns on inverse text normalization, which simplifies later retrieval and structured processing.nullable=Truedeclaration are both mandatory. Without the placeholder field, collection creation fails; withoutnullable=True, aninsertthat omits the field fails. For the exact error messages, see Troubleshooting.
The final query prints one line per recording, in the form call_0001.mp3 -> <transcript>. If a transcript is empty, check the audio address before you continue to Step 2.
Step 2: Mask PII before storage
Transcribed text often contains PII such as mobile numbers, ID card numbers, and bank card numbers. Mask the text before it enters the knowledge base or the analytics store, so that compliance moves to the front of the pipeline instead of the end. Choose one of the following patterns:
(Recommended for this pipeline) Synchronous REST API — Use it for one-time processing before data flows in. Only masked text reaches the database, and the masked transcript feeds Step 6 directly.
Write-time collection — Use it when you need to keep the masked result alongside the original text for auditing.
Call the synchronous REST API to mask text before it is written anywhere:
# ==================== Step 2: AI_PII_MASK masking ====================
# 2.1 Synchronous REST API: mask the text before it enters the database
status, data = post_json(
"/v2/vectordb/ai/pii_mask",
{
"model_name": "qwen3.7-max",
"texts": ["Hello, my phone number is 13800138000 and my ID card number is 110101199001010000. Please check my ticket."],
"params": {
"pii_types": ["PERSON", "PHONE", "ID_CARD"],
"mask_char": "*",
"preserve_length": True, # keep the original length after masking
"temperature": 0,
},
},
)
assert status == 200 and data.get("code") == 0, data
for item in data["data"]["output"]["outputs"]:
print(item)When preserve_length is true, the masked value keeps the original character length, which preserves format characteristics without exposing the original value. The following output is observed for the REST call above:
Hello, my phone number is *********** and my ID card number is ******************. Please check my ticket.The 11-digit mobile number is masked as 11 asterisks, and the 18-digit ID card number as 18 asterisks.
As an alternative, attach AI_PII_MASK to a collection so that the masked field is generated automatically on write:
# 2.2 Write-time Collection: the masked field is generated automatically when content is written
collection_name = "cs_pii_mask"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("masked", DataType.VARCHAR, max_length=4096)
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
Function(
name="mask_pii",
function_type=texttransform_function_type(),
input_field_names=["content"],
output_field_names=["masked"],
params={
"provider": "aliyun_milvus",
"model_name": "qwen3.7-max",
"task": "ai_pii_mask",
"pii_types": "PERSON,PHONE,ID_CARD",
"mask_char": "*",
"preserve_length": "true",
"temperature": "0",
},
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)
client.insert(collection_name,
[{"content": "Hello, my phone number is 13800138000. Please check my ticket."}])
client.flush(collection_name)
for row in client.query(collection_name, filter="", output_fields=["masked"], limit=1):
print(row["masked"])This alternative writes a shorter sample text that contains only a mobile number, so its masked value differs from the REST output shown earlier. Verify that the printed value keeps the sentence structure and replaces the digits of the mobile number with asterisks.
Step 3: Build the FAQ knowledge base
Write FAQ questions and their standard answers into a collection. AI_EMBEDDING vectorizes them on write, so the application does not need to call an embedding model first.
# ==================== Step 3: AI_EMBEDDING builds the FAQ knowledge base ====================
collection_name = "cs_faq_kb"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096) # FAQ question text
schema.add_field("answer", DataType.VARCHAR, max_length=4096) # standard answer
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=1024)
schema.add_function(
Function(
name="embed_content",
function_type=FunctionType.TEXTEMBEDDING,
input_field_names=["content"],
output_field_names=["embedding"],
params={
"provider": "aliyun_milvus",
"model_name": "text-embedding-v4",
"dim": 1024,
"max_client_batch_size": 10,
"max_concurrency": 1,
},
)
)
index_params = client.prepare_index_params()
index_params.add_index(
field_name="embedding",
index_type="HNSW",
metric_type="COSINE",
params={"M": 16, "efConstruction": 200},
)
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)
faqs = [
{"content": "How do I enable public network access for a Serverless Milvus instance?",
"answer": "Enable public network access on the instance details page in the console and configure the whitelist."},
{"content": "What do I do if I forget the console logon password?", "answer": "Reset it through the password recovery process in Account Center."},
{"content": "Why is my bill higher than expected?", "answer": "Check the usage details, with a focus on compute and storage usage."},
{"content": "How do I create a Collection and insert vectors?",
"answer": "Define the schema with create_collection, then call insert."},
{"content": "Does scaling up an instance affect online workloads?", "answer": "Scaling up is an online operation and usually does not interrupt service."},
]
client.insert(collection_name, faqs)
client.flush(collection_name)
client.load_collection(collection_name)This example uses text-embedding-v4, which supports multiple languages and custom dimensions. AI Center also provides models such as qwen3.7-text-embedding. Check the available models and their dimensions on the Model Service tab of AI Center in the console, and choose a model based on language coverage and quality requirements. Whichever model you choose, the dim of the vector field must match the dim in the function parameters.
Don't continue to Step 4 until all five FAQ entries are written and load_collection returns. Retrieval against a collection that is not loaded cannot confirm whether vectorization succeeded.
Step 4: Retrieve FAQ candidates by semantic search
Send the colloquial customer question to the FAQ knowledge base and take the Top-N candidates. Vector retrieval alone retrieves broadly but not precisely, so treat this output as a candidate set rather than a final answer.
# ==================== Step 4: retrieve FAQs by semantic search ====================
query = "How can I get my instance connected from the internet?"
result = client.search(
collection_name="cs_faq_kb",
data=[query],
anns_field="embedding",
limit=3,
output_fields=["content", "answer"],
)
for rank, hit in enumerate(result[0], 1):
print(f"{rank}. [similarity {hit['distance']:.4f}] {hit['entity']['content']}")The output lists three FAQ candidates with their COSINE similarity scores, ranked from highest to lowest. Keep this ranking at hand: Step 5 compares it with the reranked ranking for the same query.
Step 5: Rerank the candidates
A reranking model sorts the candidates by their real relevance to the query. Only retrieval and reranking together match the standard FAQ reliably. Choose one of the following patterns:
(Recommended for this pipeline) Ranker attached to
search— Merges retrieval and fine-grained ranking into a single call, which shortens the pipeline.Synchronous REST API — Use it when the candidate list already exists and only reranking is needed.
# ==================== Step 5: AI_RERANK reranking ====================
# 5.1 Synchronous REST API: rerank a given batch of candidates independently
status, data = post_json(
"/v2/vectordb/ai/rerank",
{
"model_name": "qwen3-rerank",
"query": query,
"documents": [
"How do I enable public network access for a Serverless Milvus instance?",
"Does scaling up an instance affect online workloads?",
"How do I create a Collection and insert vectors?",
],
"params": {"max_concurrency": 2, "timeout_sec": 10},
},
)
assert status == 200 and data.get("code") == 0, data
ranked = sorted(data["data"]["output"]["results"],
key=lambda x: x["relevance_score"], reverse=True)
for rank, item in enumerate(ranked, 1):
print(f"{rank}. [relevance {item['relevance_score']:.4f}] candidate index={item['index']}")
# 5.2 Attach a ranker to search so that retrieval and reranking finish in one call
reranker = Function(
name="rerank_faq",
function_type=FunctionType.RERANK,
input_field_names=["content"],
params={
"reranker": "model",
"provider": "aliyun_milvus",
"model_name": "qwen3-rerank",
"queries": [query],
"max_concurrency": 2,
"timeout_sec": 10,
},
)
result = client.search(
collection_name="cs_faq_kb",
data=[query],
anns_field="embedding",
limit=3,
output_fields=["content", "answer"],
ranker=reranker,
)
for rank, hit in enumerate(result[0], 1):
e = hit["entity"]
print(f"{rank}. [rerank score {hit['distance']:.4f}] {e['content']} | answer: {e['answer']}")The REST API does not support top_n and returns one score per candidate. When you need truncation, sort by score and take the top N on the application side.
The following data is observed for the query "How can I get my instance connected from the internet?".
| Candidate FAQ | Step 4: vector retrieval | Step 5: after reranking |
| How do I enable public network access for a Serverless Milvus instance? | ① 0.6141 | ① 0.5883 |
| What do I do if I forget the console logon password? | ② 0.5613 | ③ 0.2607 |
| Does scaling up an instance affect online workloads? | ③ 0.4757 | ② 0.3356 |
The two score columns are not on the same scale: the Step 4 column reports COSINE similarity between vectors, and the Step 5 column reports the relevance score of the reranking model. Compare the ranks, not the absolute values.
Vector retrieval ranked "What do I do if I forget the console logon password?" second at 0.5613, although its intent has nothing to do with connecting from the internet. This is the typical weakness of pure vector retrieval: surface text similarity stays high while the actual intent diverges. After reranking, that candidate drops to third place, and the semantically closer instance-scaling question rises to second. For a customer service bot, Top-1 accuracy directly determines answer quality, so the reranking stage cannot be skipped.
Text-only reranking does not need the is_multimodal parameter. That parameter applies only to multi-modal reranking, which uses qwen3-vl-rerank to score image and video candidates. Passing it for text-only candidates does not change the scores.
Step 6: Label sentiment and ticket category
Attach two functions to the conversation text: one determines the sentiment for quality inspection and negative-feedback alerting, and one assigns a ticket category for automatic request distribution. Both run automatically on write. In the full pipeline, the content field receives the masked text produced in Step 2; the following example writes raw text so that it runs on its own.
# ==================== Step 6: AI_SENTIMENT sentiment analysis + AI_CLASSIFY ticket classification ====================
collection_name = "cs_qc"
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096) # masked conversation text
schema.add_field("sentiment", DataType.VARCHAR, max_length=64) # sentiment output
schema.add_field("category", DataType.VARCHAR, max_length=64) # ticket category output
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
Function(
name="analyze_sentiment",
function_type=texttransform_function_type(),
input_field_names=["content"], output_field_names=["sentiment"],
params={"provider": "aliyun_milvus", "model_name": "qwen3.7-max",
"task": "ai_sentiment", "categories": "positive,negative,neutral",
"temperature": "0"},
)
)
schema.add_function(
Function(
name="classify_ticket",
function_type=texttransform_function_type(),
input_field_names=["content"], output_field_names=["category"],
params={"provider": "aliyun_milvus", "model_name": "qwen3.7-max",
"task": "ai_classify", "labels": "account,inquiry,fault,billing",
"prompt": "Classify by the topic of the customer question.", "temperature": "0"},
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema,
index_params=index_params)
client.insert(collection_name, [
{"content": "I have called three times about this problem and it is still not fixed. This is very disappointing!"},
{"content": "How do I enable public network access for a Serverless Milvus instance?"},
])
client.flush(collection_name)
for row in client.query(collection_name, filter="",
output_fields=["content", "sentiment", "category"], limit=10):
print(f"sentiment={row['sentiment']:<10} category={row['category']:<6} | {row['content'][:28]}")The following results are observed for the two sample conversations:
| Conversation text | Sentiment | Ticket category |
| I have called three times about this problem and it is still not fixed. This is very disappointing! | negative | fault |
| How do I enable public network access for a Serverless Milvus instance? | neutral | inquiry |
One collection can carry multiple functions as long as their output fields do not conflict. In this example, two functions populate sentiment and category separately, so a single write produces both the sentiment label and the ticket category.
Sentiment and classification results are model judgments, not established facts. Keep a human review stage before high-impact actions such as ticket escalation or service withdrawal.
Troubleshooting
The following table lists the errors that the examples in this tutorial most often return.
| Error message | Cause | Solution |
Connection timeout when MilvusClient connects | The instance URI omits the port, so the connection falls back to port 80. | Specify port 19530 explicitly, for example http://c-xxx.milvus.aliyuncs.com:19530. |
Failed to download multimodal content | The audio address is a placeholder or has expired, so the server cannot download the recording. | Write an address that the server can actually download from. In production, use a short-lived signed URL from your own OSS bucket. |
schema does not contain vector field | The collection defines no vector field. | Add a vector field. When the step performs no vector search, a 2-dimensional placeholder field satisfies the constraint. |
Insert missed an field dummy_vector | The placeholder vector field is not declared as nullable=True, and the insert call omits it. | Declare the placeholder field with nullable=True, or pass a value for it at insert time. |
Clean up
The examples create four collections: cs_transcribe, cs_pii_mask, cs_faq_kb, and cs_qc. These collections and their indexes stay in the instance, and cs_faq_kb stays loaded, until you drop them.
Drop the collections after you finish the tutorial:
for name in ("cs_transcribe", "cs_pii_mask", "cs_faq_kb", "cs_qc"):
if client.has_collection(name):
client.drop_collection(name)Verify that client.has_collection(name) returns False for each of the four names before you consider the cleanup complete.
Dropping a collection deletes the collection together with the data it holds, including the transcripts and masked text produced by the examples.
Data retention and compliance
Recording data stays sensitive after the pipeline finishes. Handle the stored artifacts as follows:
Retention period — Set an explicit retention period for raw audio, transcripts, sentiment labels, and review records. Delete or anonymize them when the period expires.
Derived data — Handle derived data together with the deletion of source data or the withdrawal of consent.
Model output review — Sentiment and classification results require a human review stage before high-impact actions. For details, see Step 6: Label sentiment and ticket category.
Consent, disclosure, and signed-URL requirements apply before you process any recording. For details, see Prerequisites.
Next steps
The following scenarios extend this pipeline:
Real-time agent assist — Connect retrieval and reranking to the agent workspace to push the most relevant scripts and knowledge during a live call.
Batch quality inspection — Combine historical recordings with
AI_BATCHfor offline full-coverage transcription, masking, and sentiment labeling, which raises quality inspection coverage from sampling-based inspection to full coverage.AI_BATCHis outside the scope of this tutorial.Negative-feedback alerting — Trigger alerts and escalation in real time for tickets whose
AI_SENTIMENToutput isnegative.