Content Moderation 2.0 features an upgraded engine that uses dynamic policies and advanced models to counter content variations. It provides moderation services for various business scenarios and identifies multiple types of risks. This topic explains how to use Content Moderation 2.0.
Features
Content Moderation 2.0 offers more features than the Text Moderation 1.0 service, including the ability to create custom moderation rules for more comprehensive content protection.
Use cases
Content Moderation 2.0 provides services tailored for multiple business scenarios to simplify integration, model selection, and compliance risk coverage. Select the service that best fits your business needs.
Scenario | Service | Business scenarios | Description |
Large model moderation | Large model service for UGC text moderation - Pro (ugc_moderation_byllm_pro) | All types of text moderation for UGC scenarios that require more granular labels in the response. | A text moderation service powered by a large model that provides more granular risk labels. For more information, see LLM-based text moderation service. |
Large model service for UGC text moderation (ugc_moderation_byllm) | All types of text moderation for UGC scenarios. | A text moderation service powered by a large model that efficiently and accurately identifies various violations in text. For more information, see LLM-based text moderation service. | |
Business scenarios | Nickname detection - Pro (nickname_detection_pro) | User nicknames, official account names, and live stream titles. | Provides more granular labels than the standard nickname detection service and allows you to enable or disable specific moderation labels. For more information, see Content Moderation 2.0 PLUS service. |
Private chat content detection - Pro (chat_detection_pro) | Chat interactions between users. | Provides more granular labels than the standard private chat content detection service and allows you to enable or disable specific moderation labels. For more information, see Content Moderation 2.0 PLUS service. | |
Public comment content detection - Pro (comment_detection_pro) | Comments, bullet chats, public chats, and forwards. | Provides more granular labels than the standard public comment content detection service and allows you to enable or disable specific moderation labels. For more information, see Content Moderation 2.0 PLUS service. | |
Advertising law compliance detection - Pro (ad_compliance_detection_pro) | Product materials and ad copy. | Provides more granular labels than the standard advertising law compliance detection service and allows you to enable or disable specific moderation labels. For more information, see Content Moderation 2.0 PLUS service. | |
Global scenarios | Multilingual content detection for global business - Cross-border Edition (comment_multilingual_pro_cb) | Comments, chats, and nicknames in global business, with configurable controls. | Automatically detects the language from 38 supported languages and applies policies tailored for global business. For more information, see Content Moderation 2.0 Multilingual PLUS service. |
Large model service for UGC text moderation - Cross-border Edition (ugc_moderation_byllm_cb) | All types of UGC text moderation for global scenarios. | A text moderation service for global scenarios, powered by a large model that efficiently and accurately identifies various violations in text. For more information, see LLM-based text moderation service. | |
AI-generated content detection | AI-generated text detector (text_aigc_detector) | Service providers of online information content dissemination services. | Use this service to detect and label AI-generated content. |
Special scenarios | URL risk detection (url_detection) | URL publishing and sharing, built-in browsers, and more. | Identifies risks such as fraud, pornography, and gambling in third-party URLs. For more information, see Synchronous URL risk detection API. |
AIGC risk moderation (For this scenario, we recommend using AI Safety) | Large model service for AIGC text moderation (aigc_moderation_byllm) | All types of text moderation for AIGC scenarios. | A text moderation service powered by a large model that efficiently and accurately identifies various violations in text. For more information, see LLM-based text moderation service. |
Large language model input text detection (llm_query_moderation) | User input for large language models. | Detects baseline violations (such as pornography, politics, and violence) and harmful prompts. It also provides suggestions for handling some sensitive and misleading topics. For more information, see Text moderation PLUS service for large language models. | |
Large language model generated text detection (llm_response_moderation) | AI-generated content from large language models. | Detects baseline violations (such as pornography, politics, and violence) and harmful information. It can also partially detect abusive language, biases, and undesirable values that may be generated by AI. For more information, see Text moderation PLUS service for large language models. |
Extensive moderation labels
If content contains multiple risk types, the service returns multiple labels. On the tab, click 查看标签 to view the supported labels and detection scopes for each service.
Billing
Content Moderation 2.0 supports two billing methods: pay-as-you-go and deduction by using resource packs.
Pay-as-you-go
After you enable Content Moderation 2.0, the default billing method is pay-as-you-go. You are charged based on your actual usage, and fees are settled daily. If you do not use the service, no fees are incurred. For more information, see Enable Content Moderation 2.0.
Moderation type | Services | Unit price |
Text moderation standard (text_standard) |
| CNY 7.5/10,000 calls |
Text moderation advanced (text_advanced) |
| CNY 15/10,000 calls |
Text moderation large model edition standard (text_llm_standard) |
| CNY 20/10,000 calls |
Text moderation large model edition advanced (text_llm_advanced) |
| CNY 40/10,000 calls/1,000 characters Note
|
Text moderation large model edition basic (text_llm_basic) | If you enable large model capabilities for a small model service in the console, usage is added to this billing item. For more information about how to enable the feature, see Step 4: Enable large model capabilities for a small model service (Optional). | CNY 12.5/10,000 calls |
Resource pack deduction
If you have a large volume of content to moderate or consistent moderation needs, we recommend that you purchase resource packs in advance. The larger the resource pack, the greater the discount. You can purchase and use multiple resource packs. For more information, see Purchase a resource pack for Content Moderation 2.0.
This resource pack is used to deduct usage of Content Moderation 2.0 and cannot be shared with Content Moderation traffic packs. The specific deduction factors are as follows:
Moderation type | Deduction factor |
Text moderation standard (text_standard) | The deduction factor is 1. One call is deducted from your resource pack for each successful API call. Note For example, after one successful API call, a resource pack with 10 calls will have 9 calls remaining. |
Text moderation advanced (text_advanced) | The deduction factor is 2. For each successful API call, 2 calls are deducted from your resource pack. Note For example, after one successful API call, a resource pack with 10 calls will have 8 calls remaining. |
Text moderation large model edition standard (text_llm_standard) | The deduction factor is 2.67. For each successful API call, 2.67 calls are deducted from your resource pack. Note For example, after one successful API call, a resource pack with 10 calls will have 7.33 calls remaining. |
Text moderation large model edition advanced (text_llm_advanced) | The deduction factor is 5.34. For each successful API call, 5.34 calls are deducted from your resource pack. Note For example, after one successful API call, a resource pack with 10 calls will have 4.66 calls remaining. |
Text moderation large model edition basic (text_llm_basic) | The deduction factor is 1.67. For each successful API call, 1.67 calls are deducted from your resource pack. Note For example, after one successful API call, a resource pack with 10 calls will have 8.33 calls remaining. |
Step 1: Enable the service
Before using Content Moderation 2.0, enable the service first.
Visit the Content Moderation - Enhanced Edition page, carefully read and select the service agreement.
Click Activate Now.
Step 2: Create a custom detection service (Optional)
Content Moderation 2.0 provides multiple built-in detection services for most business scenarios. For more information, see Use cases.
If you need a custom detection service, copy a built-in service and adjust its detection scope.
Log on to the Content Moderation console.
In the left-side navigation pane, choose .
On the Rules Management tab, find the service that you want to copy, click Operation in the Copy column, and then enter a Service Name and Service Description.
The copied service inherits all configurations from the source service, including the billing method, configurable items, and custom library settings. You can then adjust the detection scope of the new service. For more information, see Step 3: Configure a custom library (Optional).
Click Modify Rules to enable or disable detection items and configure the detection scope.
Step 3: Configure a custom library (Optional)
Content Moderation 2.0 provides a built-in set of moderation labels that meet most of your text moderation needs. For more information, see Risk labels.
If you need custom moderation rules, create a custom library containing a blocklist of violation keywords or an allowlist of keywords to ignore, then configure rules to match them.
Log on to the Content Moderation console.
On the page, follow these steps to configure the library.
On the Keyword Library Management tab, click Create Library.
In the Create Library panel, enter the required lexicon information.
1. You can combine multiple keywords into a single logical expression. For example, for the expression "WeChat&Part-time", a match is triggered only if both keywords are present. The ampersand (&) represents a logical AND. The tilde (~) represents a logical NOT (exclusion). In an expression, the & operator must precede the ~ operator.
2. Separate each keyword with a line break. A single keyword cannot exceed 50 characters in length.
3. You can add up to 1,000 lines. To add more than 1,000 lines at a time, import them by uploading a file.
4. You can add up to 100,000 keywords and create up to 20 libraries under a single account.
Click Create Library.
If library creation fails, an error message is displayed. Follow the instructions in the message to try again.
Configure rules.
On the tab, select the target service, and click Set Thesaurus in the Operation column on the right.
Select Blocklist, configure the libraries that you want to use for matching, and then click Save.
If any keyword from the blocklist is matched in the text to be moderated, the
labelsfield returnsC_customizedwhen you call the Content Moderation 2.0 API. This value indicates a match in a library that you created. This scenario is primarily used to detect whether the text to be moderated contains non-compliant risks.For example, your keyword library contains the keywords small loans and door-to-door service. When you moderate the text Our school's small loans: safe, fast, convenient, no collateral, flexible borrowing, same-day disbursement, and door-to-door service, the keywords small loans and door-to-door service are matched. When you call the Content Moderation 2.0 service by using an API, the value of the
labelsparameter in the response includesC_customized, in addition to any built-in labels that are matched.Select Allowlist, configure the word libraries to ignore, and then click Save.
Keywords from an allowlist are ignored during moderation. This prevents the service from flagging them as violations.
For example, you add the keywords convenient and fast to an allowlist. The text to be moderated is On-campus small loans, secure, fast, convenient, no collateral, borrow as you go, same-day disbursement, in-home service. The keywords convenient and fast are ignored, and the service only moderates the remaining text: On-campus small loans, secure, no collateral, borrow as you go, same-day disbursement, in-home service.
The configuration takes effect in about 3 minutes.
Step 4: Enable LLM capabilities (Optional)
If you are using a small model text moderation service, you can enable large model capabilities with a single click.
Log on to the Content Moderation console.
In the left-side navigation pane, choose .
On the Rules Management tab, find the service that you want to enable and click Modify Rules in the Actions column.
On the Detection Scope tab, select the Enable LLM Moderation checkbox.
In the confirmation dialog box, click OK to enable the large model capabilities.
NoteThe configuration takes effect in 3 to 5 minutes. After it is enabled, the moderation results from the large model are returned in addition to the results from the small model. Your current results are not affected. For details on the response parameters, see LlmContent.
Step 5: Integrate Content Moderation 2.0
Content Moderation 2.0 supports the following integration methods:
Integrate by calling an API. For more information, see Content Moderation 2.0 PLUS service (Recommended) or Content Moderation 2.0 general service API.
Integrate by using an SDK. For more information, see Content Moderation 2.0 PLUS service SDKs and integration guide (Recommended) or Content Moderation 2.0 general service SDKs and integration guide.
For content moderation in AI scenarios, you can use one of the following integration methods. We recommend directly integrating AI Safety.
Integrate by calling an API. For more information, see Text moderation PLUS service for large language models.
Integrate by using an SDK. For more information, see SDK for text moderation PLUS for large language models.
Step 6: View moderation results (Optional)
View the moderation results to analyze common violation types in your text content.
On the tab, view the audited text, hit labels, and request time.
You can search by time range, request ID, text content, or label. You can query data from the last 30 days. The Result Query page can store up to 50,000 records. If you have higher storage requirements, you must save the API responses.
When you search by label, you can use the following filter options:
Contains: Returns results where the label contains the specified value.
Does not contain: Returns results where the label does not contain the specified value.
Empty: Returns results that did not match any label.
Not Empty: Returns results that matched any label.
Find the text record you want to inspect and click View in the Operation column to see detailed moderation information.
If you disagree with a moderation result, you can click the Operation drop-down list in the Feedback column for that record and select No violation false alarm or Violation missed.
Step 7: View usage and risk statistics (Optional)
View call statistics to understand the recent usage of Content Moderation 2.0 across your Alibaba Cloud account and its associated RAM users.
View usage statistics:
On the tab, view the number of text moderation calls. You can filter the data by a custom time range within the last 365 days, or by Alibaba Cloud account, RAM user, or service.
Click the
icon to download the usage statistics.
View risk statistics:
On the tab, view the risk statistics for text moderation. You can filter the data by a custom time range within the last 365 days, or by Alibaba Cloud account, RAM user, or service.