Configure check items

更新时间:
复制 MD 格式

Manage detection services and protection dimensions for AI Guardrails. This topic describes how to view the Service list, enable protection dimensions, configure detection rules, and vocabularies in the console.

Prerequisites

Before you begin, ensure that you have:

Procedure

  1. Log on to the Guardrails console.

  2. In the left-side navigation pane, choose Protection Configuration > Configuration.

    The configuration list shows the available Services:

    • AI input content moderation (query_guard_pro_ec): detects text content submitted to the large model.

    • AI-generated content moderation (response_guard_pro_ec): detects text content generated by the large model.

    • Agent Log Content Detection (agent_log_guard_intl): detects text content in Agent runtime logs.

    • AIGC Input Image Security Detection (img_query_guard_intl): detects image content submitted by users.

    • AIGC Image Security Detection (img_response_guard_intl): detects image content generated by the large model.

  3. In the Actions column of the target Service, click Management to go to the Service management page.

  4. In the Protection Dimension area, view the detection capabilities supported by the current Service. Each protection dimension is displayed as a card with a toggle to enable or disable the feature.

    Text-based Services support the following protection dimensions:

    • Content compliance: detects pornographic, violent, political, and other undesirable content. Enabled by default.

    • Sensitive content detection: detects personal information or enterprise sensitive data that may be leaked.

    • Prompt injection detection: detects malicious prompts designed to bypass the safety limits of large models.

    • Malicious URL (public preview): scans content from large models for malicious links.

    • Model Hallucination (public preview): detects false or inaccurate information generated by large models.

    Note

    Enabling Sensitive content detection or Prompt injection detection incurs separate charges. For details, see Activation and billing overview.

  5. Vocabulary settings: In the configuration list, Set keyword library for a text-based Service to add blocklist or allowlist entries. For details, see Vocabulary management.

Rule management

On the Service management page, click Configuration Management on a Protection Dimension card to configure detection toggles and rules for each risk tag under that dimension.

  1. Take AI AI Input Content Security Detection_pro Pro Edition (query_guard_pro_ec) as an example. In the configuration list, click Actions column Management.

    1. In the Protection Dimension area, click Configuration Management on the target dimension card (for example, Content compliance).

    2. Select the detection type to adjust, click Edit to enter edit mode, and modify the detection status.

    3. Click Save. Changes take effect in 2 to 5 minutes.