All Products
Search
Document Center

AI Guardrails:Multimodal API integration guide

Last Updated:Sep 15, 2026

Integrate with the Guardrails multimodal moderation API to detect compliance risks, prompt attacks, sensitive data, and model hallucinations in text, image, and file content.

Prerequisites

Before you begin, make sure you have:

  • Activated the Guardrails service on the Guardrails product page

  • A Resource Access Management (RAM) user with the AliyunYundunGreenWebFullAccess policy attached

  • An AccessKey pair for your RAM user

Set up access

Step 1: Grant RAM permissions

  1. Log on to the RAM console as a RAM administrator.

  2. Create a RAM user. For more information, see Create a RAM user.

  3. Grant the AliyunYundunGreenWebFullAccess system policy to the RAM user. For more information, see Grant permissions to a RAM user.

  4. Create an AccessKey pair for the RAM user. For more information, see Get an AccessKey pair.

Step 2: Install the SDK

For SDK installation instructions and code examples, see the Multimodal SDK reference.

API reference

Service information

Property

Value

Service name

MultiModalGuard

Billing

Paid. Billed only for requests that return HTTP 200. Other error codes are not billed. See Activation and billing overview.

Image requirements

  • For images submitted to the multi-modal API, supported formats include PNG, JPG, JPEG, BMP, WEBP, TIFF, SVG, and AVIF. For HEIF images, the longest side cannot exceed 8,192 pixels. For GIFs, only the first frame is processed. For ICO files, only the last image is processed.

  • The image must not exceed 20 MB. The height or width cannot exceed 30,000 pixels, and the total number of pixels cannot exceed 250 million. For optimal performance, the image resolution should be at least 200 × 200 pixels. Lower resolutions may degrade Content Moderation algorithm performance.

  • The image must finish downloading within 3 seconds. Otherwise, the request will time out and return an error.

Endpoints

Region

Public endpoint

Internal endpoint

China (Shanghai)

https://green-cip.cn-shanghai.aliyuncs.com

https://green-cip-vpc.cn-shanghai.aliyuncs.com

China (Beijing)

https://green-cip.cn-beijing.aliyuncs.com

https://green-cip-vpc.cn-beijing.aliyuncs.com

China (Hangzhou)

https://green-cip.cn-hangzhou.aliyuncs.com

https://green-cip-vpc.cn-hangzhou.aliyuncs.com

China (Shenzhen)

https://green-cip.cn-shenzhen.aliyuncs.com

https://green-cip-vpc.cn-shenzhen.aliyuncs.com

China (Chengdu)

https://green-cip.cn-chengdu.aliyuncs.com

Not available

China (Hong Kong)

https://green-cip.cn-hongkong.aliyuncs.com

Not available

Singapore

https://green-cip.ap-southeast-1.aliyuncs.com

https://green-cip-vpc.ap-southeast-1.aliyuncs.com

Germany (Frankfurt)

https://green-cip.eu-central-1.aliyuncs.com

https://green-cip-vpc.eu-central-1.aliyuncs.com

Usage notes

Important

The QPS (queries per second) limit per user is 50. If a request includes a file modality, the limit drops to 10 QPS. Exceeding the limit triggers throttling and may affect your service. Call the API at a rate within the quota.

Request parameters

Name

Type

Required

Example

Description

Service

String

Yes

query_security_check_intl

  • AI input content moderation (query_security_check_intl)

  • AI-generated content moderation (response_security_check_intl)

ServiceParameters

JSONString

Yes

—

The parameters for the moderation service, as a JSON string. See ServiceParameters.

ServiceParameters

At least one of content, imageUrls, or fileUrls is required per request.

Name

Type

Required

Example

Description

content

String

At least one of content, imageUrls, or fileUrls

Text content to moderate

The text content to moderate. Maximum 2,000 characters per request.

imageUrls

JSONArray

["http://xxxx123"]

The image URL to moderate. Only one image per request is supported.

fileUrls

JSONArray

["http://xxxx456"]

The file URL to moderate. Only one file per request is supported. Maximum file size: 10 MB.

chatId

String

Required for stream-based moderation

ABC123

The unique identifier for one round of interaction, which includes user input and Large Language Model (LLM) output.

sessionId

String

No

14****

The session ID. It indicates that the content of this request belongs to multi-turn conversations within the same session window.

done

Boolean

Recommended for stream-based moderation

true

Indicates whether this segment is the last in the current conversation round. true: this is the final segment. false: more segments follow.

dataId

String

No

img123******

A unique identifier for the content to moderate.

accountId

String

No

13****

The account ID, which uniquely identifies an account. Context-aware moderation is not enabled by default. To enable it, contact your business manager or submit a ticket. Until it is enabled, accountId only identifies the account and does not accumulate context across requests. Submit the full content to moderate in a single request. Content split across multiple requests is moderated independently, so risks may go undetected.

ip

String

No

192.168.1.***

The IP address of the account.

referer

String

No

www.aliyun.com

The referer request header, used for scenarios such as hotlink protection. Maximum 256 characters.

referenceContent

String

No

Context content

The reference content for comparison with the submitted content to detect model hallucinations. If you enable model hallucination detection but do not pass this parameter, model hallucination detection does not take effect. The request does not fail: the API returns a normal response (Code 200), but no detection item whose Type is modelHallucination appears in Detail. To verify that hallucination detection works, see the "No hallucination detected (contrast)" example in the Examples section of this topic.

accountId, sessionId, and chatId all identify content across requests, but they are not interchangeable. The following table compares their purposes and scenarios.

Parameter

Purpose

Scenarios

accountId

Identifies different requests that come from the same user

Account-level identifier. Context-aware moderation is not enabled by default and must be enabled for your account before context is accumulated across requests.

sessionId

Marks text segments that belong to the same stream

Stream-based moderation, where the moderation engine concatenates the segments and moderates them together

chatId

Identifies one round of interaction, which includes the user input and the Large Language Model (LLM) output

Required for stream-based moderation

Response elements

Name

Type

Example

Description

Code

Integer

200

The status code. See Error codes.

Data

JSONObject

{"Result":[...]}

The moderation result. See Data.

Message

String

OK

The response message.

RequestId

String

AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****

The request ID.

Data

Name

Type

Example

Description

Detail

JSONArray

—

The per-dimension moderation results, including risk labels and confidence scores. See Detail.

Suggestion

String

pass

The overall suggested action across all dimensions. Valid values: block, mask, watch, pass. The merge priority is: block > mask > watch > pass.

Note
  • Only sensitive content detection (sensitiveData) supports the watch and mask suggestions. All other dimensions support only block or pass.

  • When multiple dimensions are detected, the suggestion results for each dimension are merged. Merge priority from highest to lowest: block > mask > watch > pass.

Detail

Each item in Detail represents the result for one protection dimension.

Name

Type

Example

Description

Suggestion

String

block

The suggested action for this dimension. Valid values: block, pass, watch, mask.

Type

String

contentModeration

The protection dimension. Valid values:

  • contentModeration: Content compliance detection

  • promptAttack: Prompt attack detection

  • sensitiveData: Sensitive content detection

  • modelHallucination: Model hallucination detection

Level

String

high

The risk level or sensitivity level. For contentModeration, promptAttack, and modelHallucination: high, medium, low, or none. For sensitiveData: S0 (no sensitive content detected), S1, S2, or S3 (higher number = higher sensitivity).

Note

For high risk: act immediately. For medium risk: route to manual review. For low risk: act only when high recall is required; otherwise treat the same as none. Configure risk score thresholds in the Guardrails console.

Result

JSONArray

—

The moderation labels and confidence scores for this dimension. See Result.

Result

Name

Type

Example

Description

Description

String

Suspected political entity

A description of the label.

Important

This field is for reference only and is subject to change. When processing results, use the Label field.

Confidence

Float

81.22

The confidence score, from 0 to 100 with two decimal places. Some labels do not return a confidence score.

Label

String

political_xxx

The moderation label returned for the content. Multiple labels may be returned.

Level

String

high

The risk or sensitivity level. Same values as the Level field in Detail.

Ext

JSONObject

—

Extension information for certain protection dimensions. See Ext.

Ext

Name

Type

Example

Applicable dimension

Riskwords

String

AA,BB,CC

contentModeration — The sensitive words detected, separated by commas. Some labels do not return this field.

CustomizedHit

JSONArray

[{"LibName":"...","Keywords":"..."}]

contentModeration — Returned when a custom library is matched. The label is customized. See CustomizedHit.

SensitiveData

JSONArray

["6201112223455"]

sensitiveData — The detected sensitive samples.

Desensitization

String

...[mobile phone number] is my contact information...

sensitiveData — The desensitized content with sensitive values masked.

FileUrl

String

https://sase-public-server-files.oss-cn-hangzhou.aliyuncs.com/saas-XXX

waterMark — The download link for the watermarked file.

OutFileSize

String

152357

waterMark — The file size in bytes.

FileUrlExp

String

1754135551

waterMark — The expiration time of the download link, as a Unix timestamp.

Filename

String

B7VKehJ4gZR.png

waterMark — The file name.

OutFileHashMd5

String

8b96ff73e8d8060016bb41b16d337871

waterMark — The MD5 hash of the file.

CustomizedHit

Name

Type

Example

Description

LibName

String

Custom Library 1

The name of the matched custom library.

Keywords

String

custom_word1,custom_word2

The matched custom words, separated by commas.

Examples

Request example

{
  "Service": "XXX",
  "ServiceParameters": {
    "content": "This thermos cup supports IP67 waterproofing and can be used underwater.",
    "chatId": "ABC123",
    "dataId": "img123******",
    "accountId": "abc",
    "sessionId": "abc",
    "imageUrls": ["http://xxxx"], # Only one image is currently supported
    "fileUrls": ["http://xxxx"], # Only one file is currently supported
    "referer": "http://www.aliyun.com",
    "referenceContent": "Product name: Smart thermos cup. Material: 316 stainless steel. Heat retention: 12 hours. Capacity: 500 ml. Package contents: 1 thermos cup, 1 user manual."
  }
}
Note

Submit the full content to moderate in a single request. If you split the same content across multiple requests, each request is moderated independently and risks may go undetected. Context is accumulated across requests only after context-aware moderation is enabled for your account.

Response examples

`query_security_check` — system policy matched

The response below shows a request that triggered multiple protection dimensions: a malicious URL, sensitive data masking, a prompt attack, a custom label, a content compliance violation, and a model hallucination.

{
  "Code": 200,
  "RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
  "Message": "OK",
  "Data": {
    "Suggestion": "block",
    "Detail": [
      {
        "Suggestion": "mask",
        "Type": "sensitiveData",
        "Level": "S2",
        "Result": [
          {
            "Ext": {
              "Desensitization": "...[mobile phone number] is my contact information...",
              "SensitiveData": [
                "136********"
              ]
            },
            "Description": "Mobile phone number (the Chinese mainland)",
            "Label": "1814",
            "Level": "S2"
          },
          {
            "Ext": {
              "SensitiveData": [
                "**City"
              ]
            },
            "Description": "City (the Chinese mainland)",
            "Label": "1739",
            "Level": "S0"
          }
        ]
      },
      {
        "Suggestion": "block",
        "Type": "promptAttack",
        "Level": "high",
        "Result": [
          {
            "Description": "Refusal Suppression Jailbreak",
            "Confidence": 100,
            "Label": "Refusal Supression Jailbreak",
            "Level": "high"
          }
        ]
      },
      {
        "Suggestion": "block",
        "Type": "contentModeration",
        "Level": "high",
        "Result": [
          {
            "Description": "Suspected political entity",
            "Confidence": 100,
            "Label": "political_entity",
            "Level": "high"
          },
          {
            "Ext": {
              "CustomizedHit": [
                {
                  "LibName": "Needs to be blocklisted",
                  "KeyWords": "word_a,word_b,word_c"
                }
              ]
            },
            "Description": "Hit custom library",
            "Confidence": 100,
            "Label": "customized",
            "Level": "high"
          }
        ]
      },
      {
        "Result": [
          {
            "Description": "Intrinsic Hallucination",
            "Confidence": 95,
            "Label": "Intrinsic Hallucination",
            "Level": "medium"
          }
        ],
        "Type": "modelHallucination",
        "Suggestion": "block",
        "Level": "medium"
      }
    ]
  }
}

`img_response_security_check` — digital watermark detected, content compliance violation

{
  "Code": 200,
  "Data": {
    "Detail": [
      {
        "Level": "none",
        "Result": [
          {
            "Confidence": 0.0,
            "Description": "No risk detected",
            "Ext": {
              "FileUrl": "https://sase-public-server-files.oss-cn-hangzhou.aliyuncs.com/saas-XXX",
              "OutFileSize": 527918,
              "FileUrlExp": "1754200240",
              "Filename": "wJGz6kmZ1Ce.jpg",
              "OutFileHashMd5": "02f5129f606027c7a87b84377ec98f8e"
            },
            "Label": "nonLabel",
            "Level": "none"
          }
        ],
        "Suggestion": "pass",
        "Type": "waterMark"
      },
      {
        "Level": "high",
        "Result": [
          {
            "Confidence": 90,
            "Description": "Violation of advertising law - superlative words",
            "Label": "ad_Compliance_WordLimit_Tii",
            "Level": "high"
          }
        ],
        "Suggestion": "block",
        "Type": "contentModeration"
      }
    ],
    "Suggestion": "block"
  },
  "Msg": "OK"
}

`text_img_security_check` — content compliance violation

{
  "Code": 200,
  "Data": {
    "Detail": [
      {
        "Ext": {},
        "Level": "high",
        "Result": [
          {
            "Confidence": 98.34,
            "Description": "Female cleavage",
            "Label": "sexual_Cleavage",
            "Level": "high"
          }
        ],
        "Suggestion": "block",
        "Type": "contentModeration"
      }
    ],
    "Suggestion": "block"
  },
  "Msg": "OK"
}

`file_security_sync_check` — malicious file detected

{
  "Code": 200,
  "Data": {
    "Detail": [
      {
        "Ext": {},
        "Level": "high",
        "Result": [
          {
            "Confidence": 100,
            "Description": "Web shell",
            "Label": "WebShell",
            "Level": "high"
          }
        ],
        "Suggestion": "block",
        "Type": "maliciousFile"
      },
      {
        "Ext": {
          "PageSum": 1
        },
        "Level": "none",
        "Result": [
          {
            "Description": "No risk detected",
            "Label": "nonLabel",
            "Level": "none"
          }
        ],
        "Suggestion": "pass",
        "Type": "contentModeration"
      }
    ],
    "Suggestion": "block"
  },
  "Msg": "OK"
}

`text_file_sec_sync_check` — no risk detected

{
  "Code": 200,
  "Data": {
    "Detail": [
      {
        "Ext": {
          "PageSum": 4
        },
        "Level": "none",
        "Result": [
          {
            "Description": "No risk detected",
            "Label": "nonLabel",
            "Level": "none"
          }
        ],
        "Suggestion": "pass",
        "Type": "contentModeration"
      }
    ],
    "Suggestion": "pass"
  },
  "Msg": "OK"
}

Model hallucination detection examples

Model hallucination detection (public preview) compares large language model (LLM) generated content against the reference content that you pass in referenceContent to identify false or inaccurate information in the model output. The following examples cover three typical hallucination scenarios and one contrast scenario without hallucination. Each example includes the scenario description, the input (submitted content and referenceContent), the detection result, and an explanation. You can copy the sample content to verify your integration. Prerequisites: you enabled the Model Hallucination protection dimension for the service in Configure check items, and you pass referenceContent in ServiceParameters.

Example 1: Information fabricated beyond the reference content

Scenario: the reference content does not mention waterproofing, but the model output fabricates an "IP67 waterproofing" claim.

Input:

{
  "Service": "query_security_check",
  "ServiceParameters": {
    "content": "This thermos cup supports IP67 waterproofing and can be used underwater.",
    "referenceContent": "Product name: Smart thermos cup. Material: 316 stainless steel. Heat retention: 12 hours. Capacity: 500 ml. Package contents: 1 thermos cup, 1 user manual."
  }
}

Detection result:

{
  "Code": 200,
  "Message": "OK",
  "RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
  "Data": {
    "Suggestion": "block",
    "Detail": [
      {
        "Type": "modelHallucination",
        "Suggestion": "block",
        "Level": "medium",
        "Result": [
          {
            "Description": "Intrinsic Hallucination",
            "Confidence": 95,
            "Label": "Intrinsic Hallucination",
            "Level": "medium"
          }
        ]
      }
    ]
  }
}

Explanation: the "IP67 waterproofing" claim in the submitted content has no basis in the reference content and is judged to be a model hallucination. The response contains a detection item whose Type is modelHallucination in Detail, the Label is Intrinsic Hallucination, the Confidence is the confidence level of the hallucination judgment, and the suggested action is block.

Example 2: Information that contradicts the reference content

Scenario: the reference content states that 7-day no-reason returns are supported, but the model output claims that returns are not supported.

Input:

{
  "Service": "query_security_check",
  "ServiceParameters": {
    "content": "This product does not support returns and cannot be refunded after purchase.",
    "referenceContent": "After-sales service: This product supports 7-day no-reason returns. The buyer pays the return shipping fee."
  }
}

Detection result:

{
  "Code": 200,
  "Message": "OK",
  "RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
  "Data": {
    "Suggestion": "block",
    "Detail": [
      {
        "Type": "modelHallucination",
        "Suggestion": "block",
        "Level": "medium",
        "Result": [
          {
            "Description": "Intrinsic Hallucination",
            "Confidence": 95,
            "Label": "Intrinsic Hallucination",
            "Level": "medium"
          }
        ]
      }
    ]
  }
}

Explanation: the submitted content directly contradicts the return policy in the reference content and is judged to be a model hallucination (Label: Intrinsic Hallucination). The suggested action is block.

Example 3: Information distorted or exaggerated from the reference content

Scenario: the reference content states a battery life of 8 hours for the standard edition, but the model output exaggerates it to "up to 24 hours of battery life for all models and a full charge in 5 minutes".

Input:

{
  "Service": "query_security_check",
  "ServiceParameters": {
    "content": "All models offer up to 24 hours of battery life and a full charge in just 5 minutes. We recommend the top-tier edition for every user.",
    "referenceContent": "Product specifications: The standard edition has 8 hours of battery life. In fast-charging mode, it charges to 80% in 30 minutes."
  }
}

Detection result:

{
  "Code": 200,
  "Message": "OK",
  "RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
  "Data": {
    "Suggestion": "block",
    "Detail": [
      {
        "Type": "modelHallucination",
        "Suggestion": "block",
        "Level": "medium",
        "Result": [
          {
            "Description": "Intrinsic Hallucination",
            "Confidence": 95,
            "Label": "Intrinsic Hallucination",
            "Level": "medium"
          }
        ]
      }
    ]
  }
}

Explanation: the submitted content distorts and exaggerates the battery life and charging data in the reference content and is judged to be a model hallucination (Label: Intrinsic Hallucination). The suggested action is block.

No hallucination detected (contrast)

Scenario: the model output is consistent with the reference content, and no hallucination is detected. The response is normal (Code 200), but no detection item whose Type is modelHallucination appears in Detail.

Input:

{
  "Service": "query_security_check",
  "ServiceParameters": {
    "content": "This thermos cup keeps drinks warm for 12 hours and has a capacity of 500 ml.",
    "referenceContent": "Product name: Smart thermos cup. Material: 316 stainless steel. Heat retention: 12 hours. Capacity: 500 ml. Package contents: 1 thermos cup, 1 user manual."
  }
}

Detection result:

{
  "Code": 200,
  "Message": "OK",
  "RequestId": "AAAAAA-BBBB-CCCCC-DDDD-EEEEEEEE****",
  "Data": {
    "Suggestion": "pass",
    "Detail": []
  }
}

Explanation: the submitted content is faithful to the reference content, so it is not judged to be a hallucination and no modelHallucination detection item appears in Detail. The response looks the same when the model hallucination protection dimension is enabled but referenceContent is not passed: the request succeeds, but no hallucination detection result is returned. If you see no modelHallucination item, first check that your request passes referenceContent.

Note
  • The Label, Description, and Confidence values in the detection result are subject to the actual API response.

  • Model hallucination detection is in public preview.

Error codes

Code

Status code

Description

200

OK

Cause: The request was processed successfully.

Recommended action: Parse the Data field in the response as usual, and handle the content according to the suggested action (Suggestion) returned in Detail / Result.

Troubleshooting: None required.

400

BAD_REQUEST

Cause: The request is invalid, usually because a request parameter is incorrect.

Recommended action: Check the request body against the Request parameters section of this topic: confirm that Service is a service code supported by this API operation, that ServiceParameters is a valid JSON string, and that the required fields (such as content) are passed correctly.

Troubleshooting: Resend the request after you correct the parameters. If status code 400 keeps being returned although the parameters are correct, submit a ticket and include the RequestId.

408

PERMISSION_DENY

Cause: The request failed authorization checks. The account may not be authorized, may have an overdue payment, may not have activated the service, or may be banned.

Recommended action: Confirm that the AI Guardrails service is activated; confirm that the Resource Access Management (RAM) user whose AccessKey is used for the call has been granted the Content Moderation permissions (such as AliyunYundunGreenWebFullAccess); confirm that the account is in good standing with no overdue payment.

Troubleshooting: Check in order: (1) whether you completed Step 1: Grant RAM permissions in the Set up access section of this topic; (2) whether the RAM user has the required policy attached; (3) whether the Alibaba Cloud account has an overdue payment or is banned. Retry after each item is ruled out. If the issue persists, submit a ticket.

500

GENERAL_ERROR

Cause: A temporary server-side error occurred.

Recommended action: Retry the request later. We recommend that you implement automatic retries (such as exponential backoff) on the client side.

Troubleshooting: If this status code keeps being returned, record the RequestId and contact us.

581

TIMEOUT

Cause: The request timed out.

Recommended action: Retry the request later. Check whether the client-side timeout setting is too short, and control the concurrency during peak hours.

Troubleshooting: Confirm that the client timeout setting is appropriate and that the network connection is normal, then retry. If this status code keeps being returned, record the RequestId and contact us.

588

EXCEED_QUOTA

Cause: The request frequency exceeds the quota. The QPS limit of this API operation is 50 per user, and drops to 10 when a request includes a file modality.

Recommended action: Lower the request rate. We recommend that you implement rate limiting and retry logic on the client side to avoid throttling caused by traffic bursts.

Troubleshooting: Check whether your current QPS exceeds the quota stated above. If your business requires a higher quota, submit a ticket to request an increase.

The model hallucination protection dimension is enabled, but no modelHallucination detection item is returned. How do I troubleshoot?

When model hallucination detection (public preview) does not take effect, the API does not return an error code: the request succeeds (Code 200), but no detection item whose Type is modelHallucination appears in Detail. If this happens, check the following items in order:

  1. Confirm that your request passes the referenceContent parameter in ServiceParameters. Model hallucination detection compares the submitted content against this reference content, and no hallucination detection result is returned when the parameter is missing. For the parameter description and self-check examples, see the Request parameters and Examples sections of this topic.

  2. Confirm that the service is a text-based service (such as query_security_check or response_security_check). The model hallucination protection dimension is supported only by text-based services; image, file, video, and audio services do not support it.

  3. Log on to the Guardrails console. In Mitigation Settings, choose the corresponding service, open the protection dimension settings, and confirm that the model hallucination card is enabled.

    Note

    The configuration may take up to 5 minutes to take effect after you enable it.

  4. In the configuration settings of the model hallucination card, check the threshold in Hallucination Degree: this setting controls the interception score for model hallucination detection. Content is judged to be a model hallucination only when the returned score exceeds the configured value. Valid values: 60 to 100. Default value: 60. If the score of the current content does not exceed the configured threshold, no hallucination detection result is returned.

If all the items above are confirmed and no hallucination detection result is returned, submit a ticket and include the RequestId.

What's next