The AI_PII_MASK function identifies and masks personally identifiable information (PII) in text. It is designed for privacy protection in customer service logs, ticket archiving, and analytics data.
Prerequisites
Before you use AI_PII_MASK, make sure you have the following:
A Milvus cluster endpoint (for example,
http://c-xxxx.milvus.aliyuncs.com:19530).Authentication credentials in
<yourUsername>:<yourPassword>format.A configured text model (for example,
qwen3.7-max).(Python only) PyMilvus installed.
Command format
AI_PII_MASK supports two invocation methods:
REST interface — Call the API directly for on-demand masking of text.
Python Collection Function — Add a function to a collection schema so that PII is automatically masked during data writes.
REST interface
REST interface
Send a POST request to the /v2/vectordb/ai/pii_mask endpoint:
POST /v2/vectordb/ai/pii_mask
Content-Type: application/json
{
"model_name": "<model name>",
"texts": ["<text>"],
"params": {"pii_types": ["EMAIL", "PHONE"]}
}Python (Collection Function)
Python (Collection Function)
Define a schema with a Function that uses the TEXTTRANSFORM function type. The texttransform_function_type() helper used below handles compatibility across PyMilvus versions. For its full implementation, see the Python usage example.
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("content", DataType.VARCHAR, max_length=4096)
schema.add_field("masked", DataType.VARCHAR, max_length=4096)
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2)
schema.add_function(
Function(
name="mask_pii",
function_type=texttransform_function_type(),
input_field_names=["content"],
output_field_names=["masked"],
params={
"provider": "aliyun_milvus",
"model_name": "<model name>",
"task": "ai_pii_mask",
"pii_types": "PERSON,EMAIL,PHONE",
"mask_char": "*",
"preserve_length": "true",
"temperature": "0",
},
)
)Parameters
| Parameter | Description |
model_name | Required. Name of the configured text model. |
texts | Required for REST. Array of texts to be masked. |
pii_types | Optional. PII types to mask. Accepts an array or a comma-separated string; 1 to 30 types supported. Defaults to masking names, email addresses, phone numbers, ID numbers, bank card numbers, addresses, URLs, IP addresses, tokens, and passwords. |
mask_char | Optional. Single replacement character. Defaults to *. |
preserve_length | Optional. Whether to preserve the length of the original sensitive segment. Defaults to true. |
temperature/max_concurrency/timeout_sec | Optional. Model stability, concurrency, and timeout settings. |
provider/task | Required only for Collection Function. Fixed values: aliyun_milvus and ai_pii_mask. |
Response
data.output.outputs returns masked text in the same order as the input. When using a Collection Function, the result is written to the target text field.
Usage example: masking customer service logs before storage
Operations teams need to retain contact records for analytics but cannot store names, email addresses, or phone numbers in the analytics database. The following example masks a log entry using REST and automatically masks data during writes using a Collection Function.
REST interface
REST interface
Run the following script to mask PII in a sample log entry:
#!/usr/bin/env bash
set -euo pipefail
MILVUS_REST_BASE_URL="http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_AUTH_TOKEN="<yourUsername>:<yourPassword>"
post_json() {
local path="$1"
local body="$2"
curl -X POST \
"$MILVUS_REST_BASE_URL$path" \
-H "Authorization: Bearer $MILVUS_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d "$body"
}
MODEL_NAME="qwen3.7-max"
BODY=$(cat <<JSON
{
"model_name": "$MODEL_NAME",
"texts": ["Please contact Zhang San at zhangsan@example.com or +86 13812345678."],
"params": {
"pii_types": ["PERSON", "EMAIL", "PHONE"],
"mask_char": "*",
"preserve_length": true,
"temperature": 0
}
}
JSON
)
RESPONSE_BODY="$(post_json "/v2/vectordb/ai/pii_mask" "$BODY")"
if command -v jq >/dev/null 2>&1; then
echo "$RESPONSE_BODY" | jq .
[ "$(echo "$RESPONSE_BODY" | jq -r '.code // -1')" = "0" ] || exit 1
else
echo "$RESPONSE_BODY"
fi
# Expected: masked = Please contact ********* at ******************** or ***************.Python (Collection Function)
Python (Collection Function)
After installing PyMilvus, replace the MILVUS_URI and MILVUS_TOKEN placeholders with your actual cluster address and token, then run the script. Common helper functions are inlined in the example.
from __future__ import annotations
from typing import Any
from pymilvus import DataType, Function, FunctionType, MilvusClient
MILVUS_URI = "http://c-xxxx.milvus.aliyuncs.com:19530"
MILVUS_TOKEN = "<yourUsername>:<yourPassword>"
DUMMY_VECTOR_DIM = 2
TEXTTRANSFORM_FUNCTION_TYPE = 9
def texttransform_function_type() -> Any:
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
function_type = getattr(FunctionType, type_name, None)
if function_type is not None:
return function_type
# Alibaba Cloud Milvus provides TEXTTRANSFORM as a managed extension (function type value 9);
# some pymilvus versions do not include this enum member natively, yet Function(...) validates through FunctionType(...).
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
def add_id(schema: Any) -> None:
schema.add_field("id", DataType.INT64, is_primary=True)
def add_dummy_vector(schema: Any) -> None:
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=DUMMY_VECTOR_DIM)
def run_texttransform_example(*, client, collection_name, input_fields, output_field, function_name, function_params, rows) -> None:
if client.has_collection(collection_name):
client.drop_collection(collection_name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
add_id(schema)
for name, data_type, max_length in input_fields:
field_params = {"max_length": max_length} if max_length is not None else {}
schema.add_field(name, data_type, **field_params)
output_name, output_data_type, output_max_length = output_field
output_params = {"max_length": output_max_length} if output_max_length is not None else {}
schema.add_field(output_name, output_data_type, **output_params)
add_dummy_vector(schema)
schema.add_function(
Function(
name=function_name,
function_type=texttransform_function_type(),
input_field_names=[name for name, _, _ in input_fields],
output_field_names=[output_name],
params=function_params,
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(collection_name=collection_name, schema=schema, index_params=index_params)
client.insert(collection_name, rows)
client.flush(collection_name)
fields = [name for name, _, _ in input_fields] + [output_name]
for row in client.query(collection_name, filter="", output_fields=fields, limit=len(rows)):
print(row)
MODEL_NAME = "qwen3.7-max"
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
run_texttransform_example(
client=client,
collection_name="simple_ai_pii_mask",
input_fields=[("content", DataType.VARCHAR, 4096)],
output_field=("masked", DataType.VARCHAR, 4096),
function_name="mask_pii",
function_params={"provider": "aliyun_milvus", "model_name": MODEL_NAME, "task": "ai_pii_mask", "pii_types": "PERSON,EMAIL,PHONE", "mask_char": "*", "preserve_length": "true", "temperature": "0"},
rows=[{"content": "Please contact Zhang San at zhangsan@example.com or +86 13812345678.", "dummy_vector": [0.1, 0.2]}],
)
# Expected: masked = Please contact ********* at ******************** or ***************.Expected result: masked returns Please contact ********* at ******************** or ***************. — the name, email address, and phone number are masked with their original lengths preserved, and the data is safe to write to the analytics collection.