Integrate with the Guardrails multimodal moderation API to detect compliance risks, prompt attacks, sensitive data, and model hallucinations in text, image, and file content.
Prerequisites
Before you begin, make sure you have:
-
Activated the Guardrails service on the Guardrails product page
-
A Resource Access Management (RAM) user with the
AliyunYundunGreenWebFullAccesspolicy attached -
An AccessKey pair for your RAM user
Set up access
Step 1: Grant RAM permissions
-
Log on to the RAM console as a RAM administrator.
-
Create a RAM user. For more information, see Create a RAM user.
-
Grant the
AliyunYundunGreenWebFullAccesssystem policy to the RAM user. For more information, see Grant permissions to a RAM user. -
Create an AccessKey pair for the RAM user. For more information, see Get an AccessKey pair.
Step 2: Install the SDK
For SDK installation instructions and code examples, see the Multimodal SDK reference.
API reference
Service information
|
Property |
Value |
|
Service name |
MultiModalGuard |
|
Billing |
Paid. Billed only for requests that return HTTP 200. Other error codes are not billed. See Activation and billing overview. |
Image requirements
-
For images submitted to the multi-modal API, supported formats include PNG, JPG, JPEG, BMP, WEBP, TIFF, SVG, and AVIF. For HEIF images, the longest side cannot exceed 8,192 pixels. For GIFs, only the first frame is processed. For ICO files, only the last image is processed.
-
The image must not exceed 20 MB. The height or width cannot exceed 30,000 pixels, and the total number of pixels cannot exceed 250 million. For optimal performance, the image resolution should be at least 200 × 200 pixels. Lower resolutions may degrade Content Moderation algorithm performance.
-
The image must finish downloading within 3 seconds. Otherwise, the request will time out and return an error.
Endpoints
|
Region |
Public endpoint |
Internal endpoint |
|
China (Shanghai) |
https://green-cip.cn-shanghai.aliyuncs.com |
https://green-cip-vpc.cn-shanghai.aliyuncs.com |
|
China (Beijing) |
https://green-cip.cn-beijing.aliyuncs.com |
https://green-cip-vpc.cn-beijing.aliyuncs.com |
|
China (Hangzhou) |
https://green-cip.cn-hangzhou.aliyuncs.com |
https://green-cip-vpc.cn-hangzhou.aliyuncs.com |
|
China (Shenzhen) |
https://green-cip.cn-shenzhen.aliyuncs.com |
https://green-cip-vpc.cn-shenzhen.aliyuncs.com |
|
China (Chengdu) |
https://green-cip.cn-chengdu.aliyuncs.com |
Not available |
|
China (Hong Kong) |
https://green-cip.cn-hongkong.aliyuncs.com |
Not available |
|
Singapore |
https://green-cip.ap-southeast-1.aliyuncs.com |
https://green-cip-vpc.ap-southeast-1.aliyuncs.com |
|
Germany (Frankfurt) |
https://green-cip.eu-central-1.aliyuncs.com |
https://green-cip-vpc.eu-central-1.aliyuncs.com |
Usage notes
The QPS (queries per second) limit per user is 50. If a request includes a file modality, the limit drops to 10 QPS. Exceeding the limit triggers throttling and may affect your service. Call the API at a rate within the quota.
Request parameters
|
Name |
Type |
Required |
Example |
Description |
|
Service |
String |
Yes |
query_security_check_intl |
|
|
ServiceParameters |
JSONString |
Yes |
— |
The parameters for the moderation service, as a JSON string. See ServiceParameters. |
ServiceParameters
At least one of content, imageUrls, or fileUrls is required per request.
|
Name |
Type |
Required |
Example |
Description |
|
content |
String |
At least one of content, imageUrls, or fileUrls |
|
The text content to moderate. Maximum 2,000 characters per request. |
|
imageUrls |
JSONArray |
|
The image URL to moderate. Only one image per request is supported. |
|
|
fileUrls |
JSONArray |
|
The file URL to moderate. Only one file per request is supported. Maximum file size: 10 MB. |
|
|
chatId |
String |
Required for stream-based moderation |
|
The unique identifier for one round of interaction, which includes user input and Large Language Model (LLM) output. |
|
sessionId |
String |
No |
|
The session ID. It indicates that the content of this request belongs to multi-turn conversations within the same session window. |
|
done |
Boolean |
Recommended for stream-based moderation |
|
Indicates whether this segment is the last in the current conversation round. |
|
dataId |
String |
No |
|
A unique identifier for the content to moderate. |
|
accountId |
String |
No |
|
The account ID, which uniquely identifies an account. Context-aware moderation is not enabled by default. To enable it, contact your business manager or submit a ticket. Until it is enabled, |
|
ip |
String |
No |
|
The IP address of the account. |
|
referer |
String |
No |
|
The referer request header, used for scenarios such as hotlink protection. Maximum 256 characters. |
|
referenceContent |
String |
No |
|
The reference content for comparison with the submitted content to detect model hallucinations. If you enable model hallucination detection but do not pass this parameter, model hallucination detection does not take effect. The request does not fail: the API returns a normal response ( |
accountId, sessionId, and chatId all identify content across requests, but they are not interchangeable. The following table compares their purposes and scenarios.
|
Parameter |
Purpose |
Scenarios |
|
accountId |
Identifies different requests that come from the same user |
Account-level identifier. Context-aware moderation is not enabled by default and must be enabled for your account before context is accumulated across requests. |
|
sessionId |
Marks text segments that belong to the same stream |
Stream-based moderation, where the moderation engine concatenates the segments and moderates them together |
|
chatId |
Identifies one round of interaction, which includes the user input and the Large Language Model (LLM) output |
Required for stream-based moderation |
Response elements
|
Name |
Type |
Example |
Description |
|
Code |
Integer |
|
The status code. See Error codes. |
|
Data |
JSONObject |
|
The moderation result. See Data. |
|
Message |
String |
|
The response message. |
|
RequestId |
String |
|
The request ID. |
Data
|
Name |
Type |
Example |
Description |
|
Detail |
JSONArray |
— |
The per-dimension moderation results, including risk labels and confidence scores. See Detail. |
|
Suggestion |
String |
|
The overall suggested action across all dimensions. Valid values: Note
|
Detail
Each item in Detail represents the result for one protection dimension.
|
Name |
Type |
Example |
Description |
|
Suggestion |
String |
|
The suggested action for this dimension. Valid values: |
|
Type |
String |
|
The protection dimension. Valid values:
|
|
Level |
String |
|
The risk level or sensitivity level. For Note
For |
|
Result |
JSONArray |
— |
The moderation labels and confidence scores for this dimension. See Result. |
Result
|
Name |
Type |
Example |
Description |
|
Description |
String |
|
A description of the label. Important
This field is for reference only and is subject to change. When processing results, use the |
|
Confidence |
Float |
|
The confidence score, from 0 to 100 with two decimal places. Some labels do not return a confidence score. |
|
Label |
String |
|
The moderation label returned for the content. Multiple labels may be returned. |
|
Level |
String |
|
The risk or sensitivity level. Same values as the |
|
Ext |
JSONObject |
— |
Extension information for certain protection dimensions. See Ext. |
Ext
|
Name |
Type |
Example |
Applicable dimension |
|
Riskwords |
String |
|
|
|
CustomizedHit |
JSONArray |
|
|
|
SensitiveData |
JSONArray |
|
|
|
Desensitization |
String |
|
|
|
FileUrl |
String |
|
|
|
OutFileSize |
String |
|
|
|
FileUrlExp |
String |
|
|
|
Filename |
String |
|
|
|
OutFileHashMd5 |
String |
|
|
CustomizedHit
|
Name |
Type |
Example |
Description |
|
LibName |
String |
|
The name of the matched custom library. |
|
Keywords |
String |
|
The matched custom words, separated by commas. |
Examples
Request example
{
"Service": "XXX",
"ServiceParameters": {
"content": "This thermos cup supports IP67 waterproofing and can be used underwater.",
"chatId": "ABC123",
"dataId": "img123******",
"accountId": "abc",
"sessionId": "abc",
"imageUrls": ["http://xxxx"], # Only one image is currently supported
"fileUrls": ["http://xxxx"], # Only one file is currently supported
"referer": "http://www.aliyun.com",
"referenceContent": "Product name: Smart thermos cup. Material: 316 stainless steel. Heat retention: 12 hours. Capacity: 500 ml. Package contents: 1 thermos cup, 1 user manual."
}
}
Submit the full content to moderate in a single request. If you split the same content across multiple requests, each request is moderated independently and risks may go undetected. Context is accumulated across requests only after context-aware moderation is enabled for your account.
Response examples
`query_security_check` — system policy matched
The response below shows a request that triggered multiple protection dimensions: a malicious URL, sensitive data masking, a prompt attack, a custom label, a content compliance violation, and a model hallucination.
{
"Code": 200,
"RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
"Message": "OK",
"Data": {
"Suggestion": "block",
"Detail": [
{
"Suggestion": "mask",
"Type": "sensitiveData",
"Level": "S2",
"Result": [
{
"Ext": {
"Desensitization": "...[mobile phone number] is my contact information...",
"SensitiveData": [
"136********"
]
},
"Description": "Mobile phone number (the Chinese mainland)",
"Label": "1814",
"Level": "S2"
},
{
"Ext": {
"SensitiveData": [
"**City"
]
},
"Description": "City (the Chinese mainland)",
"Label": "1739",
"Level": "S0"
}
]
},
{
"Suggestion": "block",
"Type": "promptAttack",
"Level": "high",
"Result": [
{
"Description": "Refusal Suppression Jailbreak",
"Confidence": 100,
"Label": "Refusal Supression Jailbreak",
"Level": "high"
}
]
},
{
"Suggestion": "block",
"Type": "contentModeration",
"Level": "high",
"Result": [
{
"Description": "Suspected political entity",
"Confidence": 100,
"Label": "political_entity",
"Level": "high"
},
{
"Ext": {
"CustomizedHit": [
{
"LibName": "Needs to be blocklisted",
"KeyWords": "word_a,word_b,word_c"
}
]
},
"Description": "Hit custom library",
"Confidence": 100,
"Label": "customized",
"Level": "high"
}
]
},
{
"Result": [
{
"Description": "Intrinsic Hallucination",
"Confidence": 95,
"Label": "Intrinsic Hallucination",
"Level": "medium"
}
],
"Type": "modelHallucination",
"Suggestion": "block",
"Level": "medium"
}
]
}
}
`img_response_security_check` — digital watermark detected, content compliance violation
{
"Code": 200,
"Data": {
"Detail": [
{
"Level": "none",
"Result": [
{
"Confidence": 0.0,
"Description": "No risk detected",
"Ext": {
"FileUrl": "https://sase-public-server-files.oss-cn-hangzhou.aliyuncs.com/saas-XXX",
"OutFileSize": 527918,
"FileUrlExp": "1754200240",
"Filename": "wJGz6kmZ1Ce.jpg",
"OutFileHashMd5": "02f5129f606027c7a87b84377ec98f8e"
},
"Label": "nonLabel",
"Level": "none"
}
],
"Suggestion": "pass",
"Type": "waterMark"
},
{
"Level": "high",
"Result": [
{
"Confidence": 90,
"Description": "Violation of advertising law - superlative words",
"Label": "ad_Compliance_WordLimit_Tii",
"Level": "high"
}
],
"Suggestion": "block",
"Type": "contentModeration"
}
],
"Suggestion": "block"
},
"Msg": "OK"
}
`text_img_security_check` — content compliance violation
{
"Code": 200,
"Data": {
"Detail": [
{
"Ext": {},
"Level": "high",
"Result": [
{
"Confidence": 98.34,
"Description": "Female cleavage",
"Label": "sexual_Cleavage",
"Level": "high"
}
],
"Suggestion": "block",
"Type": "contentModeration"
}
],
"Suggestion": "block"
},
"Msg": "OK"
}
`file_security_sync_check` — malicious file detected
{
"Code": 200,
"Data": {
"Detail": [
{
"Ext": {},
"Level": "high",
"Result": [
{
"Confidence": 100,
"Description": "Web shell",
"Label": "WebShell",
"Level": "high"
}
],
"Suggestion": "block",
"Type": "maliciousFile"
},
{
"Ext": {
"PageSum": 1
},
"Level": "none",
"Result": [
{
"Description": "No risk detected",
"Label": "nonLabel",
"Level": "none"
}
],
"Suggestion": "pass",
"Type": "contentModeration"
}
],
"Suggestion": "block"
},
"Msg": "OK"
}
`text_file_sec_sync_check` — no risk detected
{
"Code": 200,
"Data": {
"Detail": [
{
"Ext": {
"PageSum": 4
},
"Level": "none",
"Result": [
{
"Description": "No risk detected",
"Label": "nonLabel",
"Level": "none"
}
],
"Suggestion": "pass",
"Type": "contentModeration"
}
],
"Suggestion": "pass"
},
"Msg": "OK"
}
Model hallucination detection examples
Model hallucination detection (public preview) compares large language model (LLM) generated content against the reference content that you pass in referenceContent to identify false or inaccurate information in the model output. The following examples cover three typical hallucination scenarios and one contrast scenario without hallucination. Each example includes the scenario description, the input (submitted content and referenceContent), the detection result, and an explanation. You can copy the sample content to verify your integration. Prerequisites: you enabled the Model Hallucination protection dimension for the service in Configure check items, and you pass referenceContent in ServiceParameters.
Example 1: Information fabricated beyond the reference content
Scenario: the reference content does not mention waterproofing, but the model output fabricates an "IP67 waterproofing" claim.
Input:
{
"Service": "query_security_check",
"ServiceParameters": {
"content": "This thermos cup supports IP67 waterproofing and can be used underwater.",
"referenceContent": "Product name: Smart thermos cup. Material: 316 stainless steel. Heat retention: 12 hours. Capacity: 500 ml. Package contents: 1 thermos cup, 1 user manual."
}
}
Detection result:
{
"Code": 200,
"Message": "OK",
"RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
"Data": {
"Suggestion": "block",
"Detail": [
{
"Type": "modelHallucination",
"Suggestion": "block",
"Level": "medium",
"Result": [
{
"Description": "Intrinsic Hallucination",
"Confidence": 95,
"Label": "Intrinsic Hallucination",
"Level": "medium"
}
]
}
]
}
}
Explanation: the "IP67 waterproofing" claim in the submitted content has no basis in the reference content and is judged to be a model hallucination. The response contains a detection item whose Type is modelHallucination in Detail, the Label is Intrinsic Hallucination, the Confidence is the confidence level of the hallucination judgment, and the suggested action is block.
Example 2: Information that contradicts the reference content
Scenario: the reference content states that 7-day no-reason returns are supported, but the model output claims that returns are not supported.
Input:
{
"Service": "query_security_check",
"ServiceParameters": {
"content": "This product does not support returns and cannot be refunded after purchase.",
"referenceContent": "After-sales service: This product supports 7-day no-reason returns. The buyer pays the return shipping fee."
}
}
Detection result:
{
"Code": 200,
"Message": "OK",
"RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
"Data": {
"Suggestion": "block",
"Detail": [
{
"Type": "modelHallucination",
"Suggestion": "block",
"Level": "medium",
"Result": [
{
"Description": "Intrinsic Hallucination",
"Confidence": 95,
"Label": "Intrinsic Hallucination",
"Level": "medium"
}
]
}
]
}
}
Explanation: the submitted content directly contradicts the return policy in the reference content and is judged to be a model hallucination (Label: Intrinsic Hallucination). The suggested action is block.
Example 3: Information distorted or exaggerated from the reference content
Scenario: the reference content states a battery life of 8 hours for the standard edition, but the model output exaggerates it to "up to 24 hours of battery life for all models and a full charge in 5 minutes".
Input:
{
"Service": "query_security_check",
"ServiceParameters": {
"content": "All models offer up to 24 hours of battery life and a full charge in just 5 minutes. We recommend the top-tier edition for every user.",
"referenceContent": "Product specifications: The standard edition has 8 hours of battery life. In fast-charging mode, it charges to 80% in 30 minutes."
}
}
Detection result:
{
"Code": 200,
"Message": "OK",
"RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
"Data": {
"Suggestion": "block",
"Detail": [
{
"Type": "modelHallucination",
"Suggestion": "block",
"Level": "medium",
"Result": [
{
"Description": "Intrinsic Hallucination",
"Confidence": 95,
"Label": "Intrinsic Hallucination",
"Level": "medium"
}
]
}
]
}
}
Explanation: the submitted content distorts and exaggerates the battery life and charging data in the reference content and is judged to be a model hallucination (Label: Intrinsic Hallucination). The suggested action is block.
No hallucination detected (contrast)
Scenario: the model output is consistent with the reference content, and no hallucination is detected. The response is normal (Code 200), but no detection item whose Type is modelHallucination appears in Detail.
Input:
{
"Service": "query_security_check",
"ServiceParameters": {
"content": "This thermos cup keeps drinks warm for 12 hours and has a capacity of 500 ml.",
"referenceContent": "Product name: Smart thermos cup. Material: 316 stainless steel. Heat retention: 12 hours. Capacity: 500 ml. Package contents: 1 thermos cup, 1 user manual."
}
}
Detection result:
{
"Code": 200,
"Message": "OK",
"RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
"Data": {
"Suggestion": "pass",
"Detail": []
}
}
Explanation: the submitted content is faithful to the reference content, so it is not judged to be a hallucination and no modelHallucination detection item appears in Detail. The response looks the same when the model hallucination protection dimension is enabled but referenceContent is not passed: the request succeeds, but no hallucination detection result is returned. If you see no modelHallucination item, first check that your request passes referenceContent.
-
The
Label,Description, andConfidencevalues in the detection result are subject to the actual API response. -
Model hallucination detection is in public preview.
Error codes
|
Code |
Status code |
Description |
|
200 |
OK |
Cause: The request was processed successfully. Recommended action: Parse the Troubleshooting: None required. |
|
400 |
BAD_REQUEST |
Cause: The request is invalid, usually because a request parameter is incorrect. Recommended action: Check the request body against the Request parameters section of this topic: confirm that Troubleshooting: Resend the request after you correct the parameters. If status code 400 keeps being returned although the parameters are correct, submit a ticket and include the |
|
408 |
PERMISSION_DENY |
Cause: The request failed authorization checks. The account may not be authorized, may have an overdue payment, may not have activated the service, or may be banned. Recommended action: Confirm that the AI Guardrails service is activated; confirm that the Resource Access Management (RAM) user whose AccessKey is used for the call has been granted the Content Moderation permissions (such as Troubleshooting: Check in order: (1) whether you completed Step 1: Grant RAM permissions in the Set up access section of this topic; (2) whether the RAM user has the required policy attached; (3) whether the Alibaba Cloud account has an overdue payment or is banned. Retry after each item is ruled out. If the issue persists, submit a ticket. |
|
500 |
GENERAL_ERROR |
Cause: A temporary server-side error occurred. Recommended action: Retry the request later. We recommend that you implement automatic retries (such as exponential backoff) on the client side. Troubleshooting: If this status code keeps being returned, record the |
|
581 |
TIMEOUT |
Cause: The request timed out. Recommended action: Retry the request later. Check whether the client-side timeout setting is too short, and control the concurrency during peak hours. Troubleshooting: Confirm that the client timeout setting is appropriate and that the network connection is normal, then retry. If this status code keeps being returned, record the |
|
588 |
EXCEED_QUOTA |
Cause: The request frequency exceeds the quota. The QPS limit of this API operation is 50 per user, and drops to 10 when a request includes a file modality. Recommended action: Lower the request rate. We recommend that you implement rate limiting and retry logic on the client side to avoid throttling caused by traffic bursts. Troubleshooting: Check whether your current QPS exceeds the quota stated above. If your business requires a higher quota, submit a ticket to request an increase. |
The model hallucination protection dimension is enabled, but no modelHallucination detection item is returned. How do I troubleshoot?
When model hallucination detection (public preview) does not take effect, the API does not return an error code: the request succeeds (Code 200), but no detection item whose Type is modelHallucination appears in Detail. If this happens, check the following items in order:
-
Confirm that your request passes the
referenceContentparameter inServiceParameters. Model hallucination detection compares the submitted content against this reference content, and no hallucination detection result is returned when the parameter is missing. For the parameter description and self-check examples, see the Request parameters and Examples sections of this topic. -
Confirm that the service is a text-based service (such as
query_security_checkorresponse_security_check). The model hallucination protection dimension is supported only by text-based services; image, file, video, and audio services do not support it. -
Log on to the Guardrails console. In Mitigation Settings, choose the corresponding service, open the protection dimension settings, and confirm that the model hallucination card is enabled.
NoteThe configuration may take up to 5 minutes to take effect after you enable it.
-
In the configuration settings of the model hallucination card, check the threshold in Hallucination Degree: this setting controls the interception score for model hallucination detection. Content is judged to be a model hallucination only when the returned score exceeds the configured value. Valid values: 60 to 100. Default value: 60. If the score of the current content does not exceed the configured threshold, no hallucination detection result is returned.
If all the items above are confirmed and no hallucination detection result is returned, submit a ticket and include the RequestId.