Add a content safety guardrail to an AI question answering application by turning on a single data_inspection switch in an existing Alibaba Cloud Milvus AI Function. This tutorial shows how the switch inspects user input before the model is called and inspects model output before the result is returned, and how the client converges the interception results into three mutually exclusive dispositions: policy_blocked, manual_review, and operational_error.
Solution overview
Once a large model serves external traffic, content safety changes from a nice-to-have into a launch requirement. As soon as real users submit prompts at one end and the model returns text at the other end, two risk exposures open at the same time:
Input side (user to model) — Users may submit non-compliant content, or use jailbreak prompts to push the model past its safety boundary. User-generated content such as posts, comments, and nicknames also needs to pass a safety gate before it is stored.
Output side (model to user) — Even when the input looks normal, the model can still generate non-compliant content, such as marketing copy with exaggerated promises or biased statements. This is especially visible in copywriting and script generation, where the model has room to improvise.
Both ends need a gate, and the business value is direct. Content safety is a hard requirement for filing and launching large model applications. Once machines block the vast majority of clearly non-compliant content, reviewers only handle a small number of ambiguous samples, so the review team moves from reviewing everything to reviewing the difficult cases. A single screenshot of a successful jailbreak or one inappropriate statement can turn into a public incident, and the guardrail stops that risk before the content leaves the system.
Assembling this capability with a traditional approach runs into several difficulties:
The two ends are inspected at different times — Input inspection must finish before the model is called, and output inspection must finish before the result is returned. One business call spans both points around the model call.
An extra content safety service must be integrated and chained manually — The typical implementation has business code call a content safety API to inspect the input, then call the large model, then call the inspection API again for the output. That means three remote calls, three timeout and retry policies, and three authentication paths.
Interception results are hard to parse programmatically — When a safety policy is triggered, the server may return an HTTP error, may return a non-zero business code, and may also return an explicit refusal as text inside HTTP 200. Checking only the HTTP status code misses refusals inside 200 responses, and checking only keywords produces false blocks on normal business text that happens to contain words such as "risk" or "reject".
The trade-off between false blocks and missed blocks — A gate that is too tight disrupts normal business, and one that is too loose increases compliance risk. A guardrail cannot have only pass and block states; it also needs a middle ground that routes ambiguous samples to manual review.
Log compliance — Troubleshooting needs logs, but as soon as high-risk raw text and personal sensitive information land in business logs in plaintext, the logging system itself becomes a new leak surface.
The approach in Alibaba Cloud Milvus is to add onedata_inspectionswitch to theparamsof an existing AI Function, such as the text generation taskai_text_generate. The switch accepts only three values:
| Value | Inspection timing | Typical scenario | Description |
input | Before the model is called | Intelligent customer service and AI assistants that receive user prompts; UGC content before it is stored | Blocks jailbreak prompts and non-compliant input. When it is triggered, the model is not called, which saves compute and improves safety. |
output | Before the result is returned | Marketing and campaign copy publishing, script generation | Suits scenarios where the input is trusted and the only concern is non-compliant model output. |
both | Once at each end | High-risk open-ended dialogue, public-facing free-form question answering | The strongest gate for when neither end is trusted, and also the most expensive. |
The guardrail and the model call complete in a single pass inside Milvus, so data never leaves the Milvus instance. Credentials are configured centrally by the administrator on the Provider side and are never written into any request body, so the business side no longer builds or chains an external content safety service.
data_inspection adds a gate to an existing call and does not replace the required parameters of the task itself. For example, text generation still requires texts, and omitting it returns texts is required for task [ai_text_generate].
Prerequisites
A Milvus 2.6 instance is created. AI Function depends on the 2.6 kernel, and no separate model service binding is needed after creation.
To access the instance over the Internet, Public Access is enabled on the Security Configuration tab of the instance details page, and the client egress IP address is added to the public access whitelist.
A target text model is configured on the Provider side, with the Data inspection (
DATA_INSPECTION) capability available. Thedata_inspectionswitch calls this model, and the administrator configures the model and its credentials centrally on the Provider side. If the target model is not configured, calls fail with anoperational_error(the model does not exist, returned as HTTP 500 with business code 65535).pymilvus is installed. The examples in this tutorial are verified with pymilvus
[TODO: confirm version].
The RESTful interface and gRPC share port 19530, so you must specify the port explicitly, for example http://c-xxx.milvus.aliyuncs.com:19530. Omitting the port falls back to port 80 and causes a connection timeout.
Procedure
This tutorial attaches the guardrail in two ways. Choose the one that matches how your application calls the model:
Collection with a TextTransform function — The guardrail runs when you write data to a collection. Use this for pipelines that inspect and then persist content, such as UGC moderation before storage. Most steps in this tutorial use this approach.
Synchronous
(Recommended) For interactive AI question answering, use the synchronoustext_generateREST call — The guardrail runs inline and returns the disposition in the response. Use this for real-time request/response scenarios such as interactive question answering. Step 4 also demonstrates this path.text_generateREST call. For content that must be inspected before it is stored, use the Collection with a TextTransform function.
Step 1: Prepare the shared code
The following code contains the connection settings, a REST wrapper, a compatibility fallback for the TEXTTRANSFORM type, and a helper function that creates a Collection with a guardrail. The guardrail switch is the single data_inspection line inside params.
from __future__ import annotations
import json
from typing import Any
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen
from pymilvus import DataType, Function, FunctionType, MilvusClient
from pymilvus.exceptions import MilvusException
# ==================== Connection settings ====================
MILVUS_URI = "http://c-xxx.milvus.aliyuncs.com:19530" # The port must be 19530
MILVUS_TOKEN = "root:xxx"
MILVUS_REST_BASE_URL = MILVUS_URI
MODEL_NAME = "<your-configured-text-model>" # Replace with a text model already configured in the Provider
TEXTTRANSFORM_FUNCTION_TYPE = 9
# The contract identifier returned by the Provider when a safety policy is triggered.
# It is the only reliable basis for deciding whether content was blocked.
BLOCK_MARKERS = ("DataInspectionFailed", "inappropriate content")
client = MilvusClient(uri=MILVUS_URI, token=MILVUS_TOKEN)
def texttransform_function_type() -> Any:
"""Get the FunctionType for TEXTTRANSFORM; add a member dynamically when the enum is missing in older versions."""
for type_name in ("TEXTTRANSFORM", "TEXT_TRANSFORM", "TextTransform"):
ft = getattr(FunctionType, type_name, None)
if ft is not None:
return ft
existing = getattr(FunctionType, "_value2member_map_", {}).get(TEXTTRANSFORM_FUNCTION_TYPE)
if existing is not None:
return existing
extension = int.__new__(FunctionType, TEXTTRANSFORM_FUNCTION_TYPE)
extension._name_ = "TEXTTRANSFORM"
extension._value_ = TEXTTRANSFORM_FUNCTION_TYPE
FunctionType._value2member_map_[TEXTTRANSFORM_FUNCTION_TYPE] = extension
FunctionType._member_map_["TEXTTRANSFORM"] = extension
return extension
def post_json(path: str, body: dict[str, Any], timeout: int = 120) -> tuple[int, dict[str, Any]]:
"""REST wrapper: returns (http_status, data), and still tries to parse the response body when HTTP is not 2xx."""
request = Request(
f"{MILVUS_REST_BASE_URL.rstrip('/')}{path}",
data=json.dumps(body, ensure_ascii=False).encode("utf-8"),
headers={"Authorization": f"Bearer {MILVUS_TOKEN}", "Content-Type": "application/json"},
method="POST",
)
try:
with urlopen(request, timeout=timeout) as response:
return response.status, json.loads(response.read().decode("utf-8"))
except HTTPError as exc:
return exc.code, json.loads(exc.read().decode("utf-8"))
def build_guard_collection(name: str, func_name: str, in_field: str, out_field: str,
prompt: str, data_inspection: str) -> None:
"""Create a write-oriented Collection with a TextTransform guardrail."""
if client.has_collection(name):
client.drop_collection(name)
schema = MilvusClient.create_schema(auto_id=True, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field(in_field, DataType.VARCHAR, max_length=1024)
schema.add_field(out_field, DataType.VARCHAR, max_length=4096)
# A Collection must have at least one vector field. This example does not run vector search,
# so a 2-dimensional placeholder field satisfies the constraint.
# With nullable=True declared, the field does not need to be passed on insert.
schema.add_field("dummy_vector", DataType.FLOAT_VECTOR, dim=2, nullable=True)
schema.add_function(
Function(
name=func_name,
function_type=texttransform_function_type(),
input_field_names=[in_field],
output_field_names=[out_field],
params={
"provider": "aliyun_milvus",
"model_name": MODEL_NAME,
"task": "ai_text_generate",
"prompt": prompt,
"data_inspection": data_inspection, # <- guardrail switch
"temperature": "0.2",
"enable_thinking": "false",
"timeout_sec": "45",
},
)
)
index_params = client.prepare_index_params()
index_params.add_index(field_name="dummy_vector", index_type="AUTOINDEX",
metric_type="COSINE")
client.create_collection(collection_name=name, schema=schema, index_params=index_params)A Collection must contain at least one vector field, otherwise creation fails with schema does not contain vector field. This example does not run vector search, so a 2-dimensional placeholder vector field satisfies the constraint and is declared nullable=True to avoid passing a value on every write.
Step 2: Converge interception results into three dispositions
This is the engineering difficulty that is most often underestimated when a guardrail goes into production. When a safety policy is triggered, the server may return an HTTP error, a non-zero business code, or an explicit refusal as text inside HTTP 200. The client therefore needs to converge the observed signals into three mutually exclusive dispositions:
| Disposition | Meaning | Recommended action |
policy_blocked | A safety policy was triggered, the guardrail is working as designed, and the result is expected by the business. | Record an audit log and return a compliance notice to the user. No operations alert is needed. |
manual_review | The protocol layer cannot decide, for example HTTP 200 with a normal structure but suspicious content. | Send the sample to the manual review queue. "No error raised" does not mean "safely passed". |
operational_error | A real service fault, such as an unavailable model, a parameter error, or a network failure. | Alert the on-call engineer and stop retrying. Do not retry indefinitely. |
# ==================== Three-state disposition classification ====================
# Both call paths (gRPC exception / REST response body) first check the Provider contract identifier,
# then fall back to the HTTP status and business code, so the same interception yields a consistent
# disposition on both paths.
def classify_grpc(exc: MilvusException) -> str:
"""Converge a gRPC exception into a three-state disposition."""
msg = str(exc)
if any(marker in msg for marker in BLOCK_MARKERS):
return "policy_blocked" # A safety policy was triggered, which is an expected business result
return "operational_error" # Everything else counts as a service fault
def classify_rest(status: int, data: dict[str, Any]) -> str:
"""Converge a REST response into a three-state disposition."""
raw = json.dumps(data, ensure_ascii=False)
if any(marker in raw for marker in BLOCK_MARKERS):
return "policy_blocked"
if status >= 400 or data.get("code", 0) != 0:
return "operational_error"
# HTTP 200 with a normal structure: the model may have returned an explicit refusal in the body.
# The protocol layer cannot decide, so route it to manual review instead of treating it as a safe pass.
return "manual_review"classify_rest is deliberately conservative. When the response is HTTP 200 with a normal structure, the protocol layer alone cannot certify that the body is safe, because the model may have embedded an explicit refusal in it. The function therefore returns manual_review for this undecidable case instead of silently passing it. Pair classify_rest with your own answer validation so that a response you can confirm as a genuine business answer is treated as a pass, and only the responses that remain undecided reach the review queue. This is what keeps the scenario table's "Passed through" for normal input consistent with the goal of routing only a small number of ambiguous samples to manual review.
Why the decision must be based on the Provider contract identifier rather than the HTTP status code alone. Testing verified that "a non-compliant input blocked by the guardrail" and "a service fault caused by a non-existent model name" return exactly the same status code combination:
| Scenario | HTTP status | Business code | Response body contains DataInspectionFailed | Correct disposition |
| Non-compliant input blocked | 500 | 65535 | Yes | policy_blocked |
| Model name does not exist | 500 | 65535 | No | operational_error |
| Required parameter texts missing | 400 | 1100 | No | operational_error |
| Normal input | 200 | 0 | No | Passed through |
The HTTP status code and business code in the first two rows are identical, so they cannot distinguish "content blocked" from "service failed". The operational meaning of the two is completely different: misclassifying a policy block as a service fault turns every content safety interception into a false service fault alert, which over time buries real faults. The decision must therefore rely on the DataInspectionFailed contract identifier in the response body. For the same reason, do not depend on specific HTTP status code values, because they can change across gateway versions.
Do not use keywords such as "risk", "reject", or "cannot" to decide whether content was blocked, because normal business text can also contain these words and cause false blocks. Keywords are at most an auxiliary signal and must not change the final disposition.
Step 3: Inspect compliant content in input and output modes
The input mode inspects user input before the model is called, and the output mode inspects model output before the result is returned. Compliant content passes through normally.
# ==================== Step 3a: input mode, inspection before the request ====================
build_guard_collection(
"guard_input", "inspect_customer_request", "request", "response",
"Answer in one sentence in a customer service tone: ${request}", "input",
)
client.insert("guard_input", [{"request": "How long does a refund take to arrive after it is approved?"}])
client.flush("guard_input")
for row in client.query("guard_input", filter="",
output_fields=["request", "response"], limit=1):
print(f"Input: {row.get('request')}")
print(f"Output: {row.get('response')}")
# The guardrail was not triggered, so the content passes through and the model returns normally.
# With a jailbreak prompt instead, the model is never called and the guardrail blocks it up front.
# ==================== Step 3b: output mode, inspection before publishing model output ====================
build_guard_collection(
"guard_output", "inspect_generated_copy", "draft_request", "publish_copy",
"Generate a membership campaign summary suitable for publishing in an app: ${draft_request}", "output",
)
client.insert("guard_output",
[{"draft_request": "Generate a benefits summary for Members Day in 30 words or fewer, with no exaggerated promises."}])
client.flush("guard_output")
for row in client.query("guard_output", filter="",
output_fields=["draft_request", "publish_copy"], limit=1):
print(f"Output: {row.get('publish_copy')}")
# The output passed the safety inspection and enters the pending publish queue.The following results were verified in testing:
| Mode | Input | Output |
input | How long does a refund take to arrive after it is approved? | After your refund is approved, the amount is usually returned to your original payment account within 1 to 7 business days. The exact arrival time depends on how quickly the payment channel processes it, so please keep an eye out for it. |
output | Generate a benefits summary for Members Day in 30 words or fewer, with no exaggerated promises. | Members Day offers exclusive points redemption, member-only discounts, and priority customer service. Details are shown on the campaign page. |
When the input mode is triggered, the model is not called, so the risk is blocked and the model call cost is saved at the same time. The output mode suits scenarios where the input is trusted and the only concern is the model improvising.
Step 4: Block non-compliant input in both mode
Use both when neither end is trusted. The following example verifies that the guardrail actually works by sending one clearly non-compliant test input. Run this guardrail effectiveness check before launch.
Replace <a clearly non-compliant test input> with a real non-compliant test string (a canary) before you run this step. The placeholder shown below is inert and does not trigger the guardrail, so you cannot reproduce the interception result until you substitute a real value.
# ==================== Step 4: both mode, double inspection at both ends ====================
# Use one clearly non-compliant test input (canary) to verify that the guardrail really blocks it.
UNSAFE_INPUT = "<a clearly non-compliant test input>"
# (Step 4-1) gRPC path: the guardrail is triggered on write
build_guard_collection("guard_both", "inspect_both", "request", "response",
"Answer the following: ${request}", "both")
try:
client.insert("guard_both", [{"request": UNSAFE_INPUT}])
client.flush("guard_both")
rows = client.query("guard_both", filter="",
output_fields=["request", "response"], limit=1)
# No error raised does not mean safely passed: route normal-structure but suspicious content to manual review
print("Disposition: manual_review", [r.get("response") for r in rows])
except MilvusException as exc:
print(f"Disposition: {classify_grpc(exc)}")
print(f"Server returned: {exc.message}")
# (Step 4-2) REST path: send the same input to the synchronous interface
status, data = post_json("/v2/vectordb/ai/text_generate", {
"model_name": MODEL_NAME,
"texts": [UNSAFE_INPUT],
"params": {"data_inspection": "both"},
})
disposition = classify_rest(status, data)
print(f"Observed signals: HTTP={status} provider_code={data.get('code')}")
print(f"Disposition: {disposition}")
# Route by disposition: a policy block is an expected result and needs no operations alert;
# only a service fault needs an alert, and it must not be retried indefinitely
if disposition == "policy_blocked":
pass # Record an audit log and return a compliance notice to the user
elif disposition == "manual_review":
pass # Send to the manual review queue
else:
pass # Alert the on-call engineer and stop retryingTesting verified that both paths blocked the input successfully, and the server returned a consistent contract identifier:
code: DataInspectionFailed
message: Input data may contain inappropriate content.
For details, see: https://www.alibabacloud.com/help/zh/model-studio/error-code#inappropriate-contentThe guardrail was triggered on the input side, the write was blocked, and the model generated no content. After the classification functions process the signals, both the gRPC path and the REST path reach the same policy_blocked conclusion.
Testing also confirmed that the both mode does not falsely block normal business input: sending a normal customer service question to the same interface returned HTTP 200 with a complete answer.
Step 5: Log dispositions safely
After the guardrail is in production, logs must locate problems without becoming a new data leak surface. Follow these recommendations:
Record only machine-readable, redacted fields: trace id, inspection stage, HTTP status, provider code, disposition enum, and disposition action.
Trim or redact raw text and personal sensitive information. During review, retrieve the data through an authorized channel using the trace id instead of keeping raw text in business logs.
Stop and alert when a policy is triggered or
operational_erroroccurs, and do not retry indefinitely.Combine with
AI_PII_MASK: redact before logging to keep the sensitive data exposure surface as small as possible.
{
"trace_id": "req-20260808-abc123",
"stage": "both",
"http_status": 500,
"provider_code": 65535,
"disposition": "policy_blocked",
"note": "blocked by data inspection on input side"
}Clean up
The examples in this tutorial create three collections: guard_input, guard_output, and guard_both. After you finish, delete them to remove the test data:
for name in ("guard_input", "guard_output", "guard_both"):
if client.has_collection(name):
client.drop_collection(name)Solution value
| Dimension | Before the guardrail | After DATA_INSPECTION is enabled |
| Jailbreak and non-compliant prompts on the input side | Relies on downstream manual review, so handling lags behind | Blocked before the model is called |
| Non-compliant content leaving on the output side | Found afterwards and handled reactively | Intercepted before the result is returned, so it never leaves |
| Manual review volume | Every item is reviewed manually | Only the small number of manual_review samples are reviewed |
| Number of systems | Business system plus an external content safety service, chained through three remote calls | One system (Milvus, with the guardrail attached to the call) |
| Data and credentials | Raw text flows out to an external service, and authentication is scattered | Data stays inside the instance, and credentials are configured centrally by the Provider |
Once the content safety guardrail is reduced to a single switch, it can be chosen flexibly per scenario:
Intelligent customer service question answering — Add
inputto text generation to block jailbreak prompts and non-compliant questions.Marketing copy and script generation — Add
outputto text generation to intercept non-compliant content before publishing.Public-facing open-ended AI assistants — Use
Directions worth extending:bothfor double protection at each end.
Combine with
AI_PII_MASK— Redact first, then inspect, then log, to keep the sensitive data exposure surface as small as possible.Cover more tasks — As more AI Function tasks support
data_inspection, the sameinput/output/bothmodel transfers to more generative scenarios.Close the policy loop — Turn
manual_reviewsamples into an evaluation set, and keep calibrating the balance between false blocks and missed blocks so the guardrail becomes more accurate over time.