Implement AI Content Moderation using Gateway with Inference Extension
To implement content compliance review for generative AI services in ACK, you can configure the ACKTrafficFilter plug-in using the Gateway API Inference Extension. This plug-in connects to Alibaba Cloud Content Moderation Service, which automatically blocks inappropriate content at the gateway layer and helps meet regulatory requirements.
How it works
In an ACK cluster, you can integrate the ACKTrafficFilter content moderation plug-in with the AI inference enhanced gateway (Gateway with Inference Extension). This plug-in calls the Alibaba Cloud Content Moderation Service. The service reviews request content and AI-generated results, and replaces non-compliant content with security prompts.
Procedure
Step 1: Install the Gateway Inference Extension Plug-in
Log on to the ACK console. In the left navigation pane, click Clusters.
On the Clusters page, click the name of your cluster. In the left navigation pane, click Add-ons .
On the Add-ons page, search for Gateway with Inference Extension. On the component card, click Install. In the dialog box that appears, check Enable Gateway API Inference Extension. Follow the prompts to complete the installation.
Step 2: Install the ACKTrafficFilter Plug-in Service
Activate the Alibaba Cloud Content Moderation Service.
Replace
<ALIYUN_ACCESS_KEY_ID>and<ALIYUN_ACCESS_KEY_SECRET>with the AccessKey ID and AccessKey Secret of the destination account. Replace<ENDPOINT>with the private network endpoint for the cluster's region from the following table. If the destination region is not listed in the table, use the public network endpoint for a nearby region. Then, save the content as theacktrafficfilter.yamlfile.If you use an AccessKey for a Resource Access Management (RAM) user, ensure that the user has been granted
AliyunYundunGreenWebFullAccessauthorization.apiVersion: inferenceextension.alibabacloud.com/v1alpha1 kind: ACKTrafficFilter metadata: name: content-security-filter spec: aiContentSecurity: accessKey: <ALIYUN_ACCESS_KEY_ID> secretKey: <ALIYUN_ACCESS_KEY_SECRET> aliyunEndpoint: <ENDPOINT>Region
Public endpoint
VPC endpoint
China (Beijing)
green-cip.cn-beijing.aliyuncs.com
green-cip-vpc.cn-beijing.aliyuncs.com
China (Shanghai)
green-cip.cn-shanghai.aliyuncs.com
green-cip-vpc.cn-shanghai.aliyuncs.com
China (Hangzhou)
green-cip.cn-hangzhou.aliyuncs.com
green-cip-vpc.cn-hangzhou.aliyuncs.com
China (Shenzhen)
green-cip.cn-shenzhen.aliyuncs.com
green-cip-vpc.cn-shenzhen.aliyuncs.com
China (Chengdu)
green-cip.cn-chengdu.aliyuncs.com
N/A
You can create the acktrafficfilter plug-in.
kubectl apply -f acktrafficfilter.yaml
Step 3: Verify Content Moderation Effect
These steps describe how to create a gateway named mock-gateway and deploy a sample application that simulates a Large Language Model (LLM). Then, you can configure a forwarding route for the sample application and initiate a test request.
You can create the backend AI service. Save the following YAML content as the mock-vllm.yaml file. Then, execute the
kubectl apply -f mock-vllm.yamlcommand.apiVersion: apps/v1 kind: Deployment metadata: name: mock-vllm spec: replicas: 2 selector: matchLabels: app: mock-vllm template: metadata: labels: app: mock-vllm spec: containers: - args: - --model - mock - --port - "8000" image: registry-cn-hangzhou.ack.aliyuncs.com/dev/mock-vllm:v0.2.3-g053679d-aliyun imagePullPolicy: IfNotPresent name: vllm-sim ports: - containerPort: 8000 name: http protocol: TCP --- apiVersion: inference.networking.x-k8s.io/v1alpha2 kind: InferencePool metadata: name: mock-pool spec: extensionRef: group: "" kind: Service name: mock-ext-proc selector: app: mock-vllm targetPortNumber: 8000 --- apiVersion: inference.networking.x-k8s.io/v1alpha2 kind: InferenceModel metadata: name: mock-model spec: criticality: Critical modelName: mock poolRef: group: inference.networking.x-k8s.io kind: InferencePool name: mock-pool targetModels: - name: mock weight: 100You can create the gateway and configure route forwarding. Save the following YAML content as the mock-gateway.yaml file. Then, execute the
kubectl apply -f mock-gateway.yamlcommand.apiVersion: gateway.networking.k8s.io/v1 kind: Gateway metadata: name: mock-gateway spec: gatewayClassName: ack-gateway listeners: - name: llm-gw protocol: HTTP port: 8080 --- apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: mock-route spec: parentRefs: - group: gateway.networking.k8s.io kind: Gateway name: mock-gateway sectionName: llm-gw rules: - backendRefs: - group: inference.networking.x-k8s.io kind: InferencePool name: mock-pool filters: - type: ExtensionRef extensionRef: group: inferenceextension.alibabacloud.com kind: ACKTrafficFilter name: content-security-filter # Reference the ACKTrafficFilter instance matches: - path: type: PathPrefix value: /You can modify
<REQUEST_CONTENT>and initiate a request to verify the ACKTrafficFilter content moderation effect.export GATEWAY_ADDRESS=$(kubectl get gateway/mock-gateway -o jsonpath='{.status.addresses[0].value}') curl -X POST ${GATEWAY_ADDRESS}:8080/v1/chat/completions \ -H 'Content-Type: application/json' -H "host: example.com" -v -d '{ "model": "mock", "max_completion_tokens": 100, "temperature": 0, "messages": [ { "role": "user", "content": "<REQUEST_CONTENT>" } ] }'Content Moderation Passed Output
{ "id": "chatcmpl-9bffeb49-057e-4c42-97e5-6fd62e3e996e", "created": 1759057516, "model": "mock", "usage": { "prompt_tokens": 1, "completion_tokens": 7, "total_tokens": 8 }, "object": "chat.completion", "choices": [ { "index": 0, "finish_reason": "stop", "message": { "role": "assistant", "content": "Today is a nice sunny day." } } ] }Content Moderation Failed Output
The HTTP response is
403 Forbidden, and the responsecontentis replaced with a security prompt.{ "id": "chatcmpl-uuFxPpxMo63DFt4lxiUjFY70kwgVs", "object": "chat.completion", "model": "from-security-guard", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "I am an artificial intelligence language model. I do not participate in or promote discussions of explicit content. I provide answers to legitimate and ethical questions about education, culture, and history. If you have other questions, please ask." }, "logprobs": null, "finish_reason": "stop" } ] }
Step 4: Clean Up the Environment
You can clean up cluster resources:
# Delete the gateway and route kubectl delete -f mock-gateway.yaml # Delete the content moderation plug-in kubectl delete -f acktrafficfilter.yaml # Delete the backend application kubectl delete -f mock-vllm.yamlOn the Add-ons page, search for Gateway with Inference Extension. On the component card, click Uninstall.
References
For more information, see Quickly Experience Gateway with Inference Extension