Use Content Moderation 2.0 to identify text risks

更新时间:
复制 MD 格式

Content Moderation 2.0 features an upgraded engine that uses dynamic policies and advanced models to counter content variations. It provides moderation services for various business scenarios and identifies multiple types of risks. This topic explains how to use Content Moderation 2.0.

Features

Content Moderation 2.0 offers more features than the Text Moderation 1.0 service, including the ability to create custom moderation rules for more comprehensive content protection.

Use cases

Content Moderation 2.0 provides services tailored for multiple business scenarios to simplify integration, model selection, and compliance risk coverage. Select the service that best fits your business needs.

Scenario

Service

Business scenarios

Description

Large model moderation

Large model service for UGC text moderation - Pro (ugc_moderation_byllm_pro)

All types of text moderation for UGC scenarios that require more granular labels in the response.

A text moderation service powered by a large model that provides more granular risk labels. For more information, see LLM-based text moderation service.

Large model service for UGC text moderation (ugc_moderation_byllm)

All types of text moderation for UGC scenarios.

A text moderation service powered by a large model that efficiently and accurately identifies various violations in text. For more information, see LLM-based text moderation service.

Business scenarios

Nickname detection - Pro (nickname_detection_pro)

User nicknames, official account names, and live stream titles.

Provides more granular labels than the standard nickname detection service and allows you to enable or disable specific moderation labels. For more information, see Content Moderation 2.0 PLUS service.

Private chat content detection - Pro (chat_detection_pro)

Chat interactions between users.

Provides more granular labels than the standard private chat content detection service and allows you to enable or disable specific moderation labels. For more information, see Content Moderation 2.0 PLUS service.

Public comment content detection - Pro (comment_detection_pro)

Comments, bullet chats, public chats, and forwards.

Provides more granular labels than the standard public comment content detection service and allows you to enable or disable specific moderation labels. For more information, see Content Moderation 2.0 PLUS service.

Advertising law compliance detection - Pro (ad_compliance_detection_pro)

Product materials and ad copy.

Provides more granular labels than the standard advertising law compliance detection service and allows you to enable or disable specific moderation labels. For more information, see Content Moderation 2.0 PLUS service.

Global scenarios

Multilingual content detection for global business - Cross-border Edition (comment_multilingual_pro_cb)

Comments, chats, and nicknames in global business, with configurable controls.

Automatically detects the language from 38 supported languages and applies policies tailored for global business. For more information, see Content Moderation 2.0 Multilingual PLUS service.

Large model service for UGC text moderation - Cross-border Edition (ugc_moderation_byllm_cb)

All types of UGC text moderation for global scenarios.

A text moderation service for global scenarios, powered by a large model that efficiently and accurately identifies various violations in text. For more information, see LLM-based text moderation service.

AI-generated content detection

AI-generated text detector (text_aigc_detector)

Service providers of online information content dissemination services.

Use this service to detect and label AI-generated content.

Special scenarios

URL risk detection (url_detection)

URL publishing and sharing, built-in browsers, and more.

Identifies risks such as fraud, pornography, and gambling in third-party URLs. For more information, see Synchronous URL risk detection API.

AIGC risk moderation (For this scenario, we recommend using AI Safety)

Large model service for AIGC text moderation (aigc_moderation_byllm)

All types of text moderation for AIGC scenarios.

A text moderation service powered by a large model that efficiently and accurately identifies various violations in text. For more information, see LLM-based text moderation service.

Large language model input text detection (llm_query_moderation)

User input for large language models.

Detects baseline violations (such as pornography, politics, and violence) and harmful prompts. It also provides suggestions for handling some sensitive and misleading topics. For more information, see Text moderation PLUS service for large language models.

Large language model generated text detection (llm_response_moderation)

AI-generated content from large language models.

Detects baseline violations (such as pornography, politics, and violence) and harmful information. It can also partially detect abusive language, biases, and undesirable values that may be generated by AI. For more information, see Text moderation PLUS service for large language models.

Legacy services

The following are legacy services. We recommend using the latest PLUS services listed in the table above, as they offer more labels, flexible configuration, and more comprehensive risk mapping.

Scenario

Service

Business scenarios

Description

Business scenarios (Legacy)

Nickname detection (nickname_detection)

User nicknames, official account names, and live stream titles.

Focuses on identifying baseline violations (such as pornography, politics, and violence), impersonation of official accounts, and prohibited traffic-directing risks. It can help manage fake accounts.

Private chat content detection (chat_detection)

Chat interactions between users.

Balances user experience with identifying baseline violations (such as pornography, politics, and violence) and risks like abuse and cyberbullying.

Public comment content detection (comment_detection)

Comments, bullet chats, public chats, and forwards.

Typically high-risk scenarios with diverse and evolving risk types. Identifies baseline violations (such as pornography, politics, and violence), ad-based traffic direction, and prohibited content. It can be integrated into the decision engine. For more information, see Use the text moderation service in the decision engine.

PGC general content detection (pgc_detection)

General materials such as office documents, courseware, and promotional materials.

Suitable for low-risk content scenarios that require precise detection of baseline risks (such as pornography, politics, and violence).

AIGC-type text detection (ai_art_detection)

Text prompts for AI text-to-image generation.

Supports both Chinese and English text, focusing on identifying baseline violations (such as pornography, politics, and violence) and negative content.

Advertising law compliance detection (ad_compliance_detection)

Product materials and ad copy.

Identifies suspected violations of advertising laws, including superlative terms, industry restrictions, and red-line violations (such as pornography, politics, and violence).

Global scenarios (Legacy)

Multilingual content detection for global business (comment_multilingual_pro)

Comments, chats, and nicknames in global business.

Automatically detects the language from 38 supported languages and applies policies tailored for global business. For more information, see Content Moderation 2.0 Multilingual service.

Extensive moderation labels

If content contains multiple risk types, the service returns multiple labels. On the 机器审核增强版 > Text Moderation > Rules > Rules Management tab, click 查看标签 to view the supported labels and detection scopes for each service.

Billing

Content Moderation 2.0 supports two billing methods: pay-as-you-go and deduction by using resource packs.

Pay-as-you-go

After you enable Content Moderation 2.0, the default billing method is pay-as-you-go. You are charged based on your actual usage, and fees are settled daily. If you do not use the service, no fees are incurred. For more information, see Enable Content Moderation 2.0.

Moderation type

Services

Unit price

Text moderation standard (text_standard)

  • Nickname detection - Pro: nickname_detection_pro

  • Private chat content detection - Pro: chat_detection_pro

  • Public comment content detection - Pro: comment_detection_pro

  • Advertising law compliance detection - Pro: ad_compliance_detection_pro

  • Nickname detection: nickname_detection

  • Private chat content detection: chat_detection

  • Public comment content detection: comment_detection

  • AIGC-type text detection: ai_art_detection

  • Advertising law compliance detection: ad_compliance_detection

  • PGC general content detection: pgc_detection

  • URL risk detection: url_detection

CNY 7.5/10,000 calls

Text moderation advanced (text_advanced)

  • Multilingual content detection for global business - Cross-border Edition: comment_multilingual_pro_cb

  • Multilingual content detection for global business: comment_multilingual_pro

  • Large language model input text detection: llm_query_moderation

  • Large language model generated text detection: llm_response_moderation

CNY 15/10,000 calls

Text moderation large model edition standard (text_llm_standard)

  • Large model service for UGC text moderation - Pro: ugc_moderation_byllm_pro

  • Large model service for UGC text moderation: ugc_moderation_byllm

  • Large model service for UGC text moderation - Cross-border Edition: ugc_moderation_byllm_cb

  • Large model service for AIGC text moderation: aigc_moderation_byllm

CNY 20/10,000 calls

Text moderation large model edition advanced (text_llm_advanced)

  • AI-generated text detector: text_aigc_detector

  • Text translation

CNY 40/10,000 calls/1,000 characters

Note
  • Billing is based on the number of characters in each API call. Each call is billed in units of 1,000 characters. If the character count exceeds 1,000, it is rounded up to the nearest thousand. For example, a single API call with 3,500 characters is counted as 4 calls and costs CNY 0.016.

  • After enabling the text translation feature, each request is billed per 500 characters.

Text moderation large model edition basic (text_llm_basic)

If you enable large model capabilities for a small model service in the console, usage is added to this billing item. For more information about how to enable the feature, see Step 4: Enable large model capabilities for a small model service (Optional).

CNY 12.5/10,000 calls

Resource pack deduction

If you have a large volume of content to moderate or consistent moderation needs, we recommend that you purchase resource packs in advance. The larger the resource pack, the greater the discount. You can purchase and use multiple resource packs. For more information, see Purchase a resource pack for Content Moderation 2.0.

This resource pack is used to deduct usage of Content Moderation 2.0 and cannot be shared with Content Moderation traffic packs. The specific deduction factors are as follows:

Moderation type

Deduction factor

Text moderation standard (text_standard)

The deduction factor is 1. One call is deducted from your resource pack for each successful API call.

Note

For example, after one successful API call, a resource pack with 10 calls will have 9 calls remaining.

Text moderation advanced (text_advanced)

The deduction factor is 2. For each successful API call, 2 calls are deducted from your resource pack.

Note

For example, after one successful API call, a resource pack with 10 calls will have 8 calls remaining.

Text moderation large model edition standard (text_llm_standard)

The deduction factor is 2.67. For each successful API call, 2.67 calls are deducted from your resource pack.

Note

For example, after one successful API call, a resource pack with 10 calls will have 7.33 calls remaining.

Text moderation large model edition advanced (text_llm_advanced)

The deduction factor is 5.34. For each successful API call, 5.34 calls are deducted from your resource pack.

Note

For example, after one successful API call, a resource pack with 10 calls will have 4.66 calls remaining.

Text moderation large model edition basic (text_llm_basic)

The deduction factor is 1.67. For each successful API call, 1.67 calls are deducted from your resource pack.

Note

For example, after one successful API call, a resource pack with 10 calls will have 8.33 calls remaining.

Step 1: Enable the service

Before using Content Moderation 2.0, enable the service first.

  1. Visit the Content Moderation - Enhanced Edition page, carefully read and select the service agreement.

  2. Click Activate Now.

Step 2: Create a custom detection service (Optional)

Content Moderation 2.0 provides multiple built-in detection services for most business scenarios. For more information, see Use cases.

If you need a custom detection service, copy a built-in service and adjust its detection scope.

  1. Log on to the Content Moderation console.

  2. In the left-side navigation pane, choose Machine Moderation V2.0 > Text Moderation > Rules.

  3. On the Rules Management tab, find the service that you want to copy, click Operation in the Copy column, and then enter a Service Name and Service Description.

    The copied service inherits all configurations from the source service, including the billing method, configurable items, and custom library settings. You can then adjust the detection scope of the new service. For more information, see Step 3: Configure a custom library (Optional).

  4. Click Modify Rules to enable or disable detection items and configure the detection scope.

Step 3: Configure a custom library (Optional)

Content Moderation 2.0 provides a built-in set of moderation labels that meet most of your text moderation needs. For more information, see Risk labels.

If you need custom moderation rules, create a custom library containing a blocklist of violation keywords or an allowlist of keywords to ignore, then configure rules to match them.

  1. Log on to the Content Moderation console.

  2. On the Machine Moderation V2.0 > Text Moderation > Library Management page, follow these steps to configure the library.

    1. On the Keyword Library Management tab, click Create Library.

    2. In the Create Library panel, enter the required lexicon information.

      1. You can combine multiple keywords into a single logical expression. For example, for the expression "WeChat&Part-time", a match is triggered only if both keywords are present. The ampersand (&) represents a logical AND. The tilde (~) represents a logical NOT (exclusion). In an expression, the & operator must precede the ~ operator.

      2. Separate each keyword with a line break. A single keyword cannot exceed 50 characters in length.

      3. You can add up to 1,000 lines. To add more than 1,000 lines at a time, import them by uploading a file.

      4. You can add up to 100,000 keywords and create up to 20 libraries under a single account.

    3. Click Create Library.

      If library creation fails, an error message is displayed. Follow the instructions in the message to try again.

  3. Configure rules.

    1. On the Machine Moderation V2.0 > Text Moderation > Rules > Rules Management tab, select the target service, and click Set Thesaurus in the Operation column on the right.

    2. Select Blocklist, configure the libraries that you want to use for matching, and then click Save.

      If any keyword from the blocklist is matched in the text to be moderated, the labels field returns C_customized when you call the Content Moderation 2.0 API. This value indicates a match in a library that you created. This scenario is primarily used to detect whether the text to be moderated contains non-compliant risks.

      For example, your keyword library contains the keywords small loans and door-to-door service. When you moderate the text Our school's small loans: safe, fast, convenient, no collateral, flexible borrowing, same-day disbursement, and door-to-door service, the keywords small loans and door-to-door service are matched. When you call the Content Moderation 2.0 service by using an API, the value of the labels parameter in the response includes C_customized, in addition to any built-in labels that are matched.

    3. Select Allowlist, configure the word libraries to ignore, and then click Save.

      Keywords from an allowlist are ignored during moderation. This prevents the service from flagging them as violations.

      For example, you add the keywords convenient and fast to an allowlist. The text to be moderated is On-campus small loans, secure, fast, convenient, no collateral, borrow as you go, same-day disbursement, in-home service. The keywords convenient and fast are ignored, and the service only moderates the remaining text: On-campus small loans, secure, no collateral, borrow as you go, same-day disbursement, in-home service.

    The configuration takes effect in about 3 minutes.

Step 4: Enable LLM capabilities (Optional)

If you are using a small model text moderation service, you can enable large model capabilities with a single click.

  1. Log on to the Content Moderation console.

  2. In the left-side navigation pane, choose Machine Moderation V2.0 > Text Moderation > Rules.

  3. On the Rules Management tab, find the service that you want to enable and click Modify Rules in the Actions column.

  4. On the Detection Scope tab, select the Enable LLM Moderation checkbox.

  5. In the confirmation dialog box, click OK to enable the large model capabilities.

    Note

    The configuration takes effect in 3 to 5 minutes. After it is enabled, the moderation results from the large model are returned in addition to the results from the small model. Your current results are not affected. For details on the response parameters, see LlmContent.

Step 5: Integrate Content Moderation 2.0

Content Moderation 2.0 supports the following integration methods:

For content moderation in AI scenarios, you can use one of the following integration methods. We recommend directly integrating AI Safety.

Step 6: View moderation results (Optional)

View the moderation results to analyze common violation types in your text content.

  1. On the Machine Moderation V2.0 > Text ModerationDetection Results tab, view the audited text, hit labels, and request time.

    You can search by time range, request ID, text content, or label. You can query data from the last 30 days. The Result Query page can store up to 50,000 records. If you have higher storage requirements, you must save the API responses.

    When you search by label, you can use the following filter options:

    • Contains: Returns results where the label contains the specified value.

    • Does not contain: Returns results where the label does not contain the specified value.

    • Empty: Returns results that did not match any label.

    • Not Empty: Returns results that matched any label.

  2. Find the text record you want to inspect and click View in the Operation column to see detailed moderation information.

    If you disagree with a moderation result, you can click the Operation drop-down list in the Feedback column for that record and select No violation false alarm or Violation missed.

Step 7: View usage and risk statistics (Optional)

View call statistics to understand the recent usage of Content Moderation 2.0 across your Alibaba Cloud account and its associated RAM users.

View usage statistics:

  • On the Machine Moderation V2.0 > Text Moderation > Dashboardstab, view the number of text moderation calls. You can filter the data by a custom time range within the last 365 days, or by Alibaba Cloud account, RAM user, or service.

  • Click the 下载 icon to download the usage statistics.

View risk statistics:

  • On the Machine Moderation V2.0 > Text Moderation > Dashboards tab, view the risk statistics for text moderation. You can filter the data by a custom time range within the last 365 days, or by Alibaba Cloud account, RAM user, or service.