Intelligent Routing
Intelligent Routing dynamically analyzes request content and automatically matches the most suitable model from the candidate set, helping enterprises achieve optimal AI resource allocation in complex scenarios.
Overview
The model-code has a fixed auto-model prefix followed by 8 random characters (e.g., auto-model-abcd1234). The client calls this model-code → routing dynamically analyzes request content and matches the most suitable model from the candidate set → actual model infers → response (x-dashscope-resolved-model marks actual model, retains real token usage).
Scope:
| Dimension | Description |
|---|---|
| Region | Beijing (cn-beijing) + Singapore (ap-southeast-1) |
| Model type | Text input only |
| Call protocol | Only OpenAI-compatible interface (Chat Completions) is supported; DashScope and Anthropic protocols are not supported |
| Access domain | Only maas.aliyuncs.com (https://{WorkspaceId}.{region}.maas.aliyuncs.com/compatible-mode/v1, region is cn-beijing or ap-southeast-1); dashscope.aliyuncs.com (incl. compatible-mode) not supported |
| Billing | Pay-as-you-go |
| Throttling | By actual routed model, at account and workspace level |
Key advantages:
- Maximize model effectiveness: Automatically matches the optimal model for each task, ideal for users with extreme quality requirements.
- Precise cost reduction: Routes simple tasks to lightweight models and complex tasks to high-end models, avoiding over-provisioning and achieving precise cost-demand alignment.
- Eliminate model selection difficulty: No need for manual model evaluation or comparison; the platform auto-selects the optimal solution, drastically lowering the model selection barrier.
- Minimal learning curve: Consistent with regular model access; call via model-code with zero code changes.
Create intelligent routing
In the Bailian console under "Model Inference > Dedicated Deployment", click Deploy New Model to enter the creation page, and select Intelligent Routing under "Deployment Mode".
Prerequisites
- A Bailian workspace is activated.
- The target route model version is published, and the current account has invocation permission for the version and candidate models.
Steps
- Enter a Service Name (required, up to 50 characters).
- Under Deployment Mode, select "Intelligent Routing".
- In Route Model Version, select a published version (maintained by the platform).
- Under Routing Strategy, select "Effect First" or "Cost First".
- Under Candidate Model Set, select candidate models (at least 2).
- Click Confirm Deploy to complete creation.
After deployment, the system generates a globally unique model-code (in the form auto-model-XXXXXXXX). The model-code is not reused after deletion.
Routing strategy
The routing strategy determines which model in the candidate set executes each request. Both strategies dynamically match by request content (not fixed routing); the choice is made at creation.
Strategy | Selection mechanism | Suitable scenarios |
|---|---|---|
Effect First (EFFECT_FIRST) | Picks the strongest-effect model in the candidate set per request | Critical business, complex reasoning, high-quality content generation where quality is stringent |
Cost First (COST_FIRST) | Intelligently schedules lightweight models while meeting a quality floor | Cost-sensitive, high-concurrency, latency- and unit-cost-aware scenarios |
Candidate model configuration
The candidate model set is the range the routing service can choose from. When creating intelligent routing, select the models to participate in routing from the candidate list in the Candidate Model Set area—at least 2 candidate models are required, otherwise deployment is blocked.
NoteRAM users can only select candidate models they have
invocation permission for—if fewer than 2 are selectable, grant the RAM user model invocation authorization in that workspace first.Route model version is the standardized release form of intelligent routing capability; each version specifies the model range and core capabilities supported by the routing system for that stage. The candidate list is determined by the selected route model version—each version corresponds to one candidate set; cross-version combination is not allowed. The currently published route model version and candidate models:
model-router-2026-0908
Candidate model | Context | Input | Output | Cache hit input |
|---|---|---|---|---|
| 1M | 12 | 36 | 1.5 |
| 1M | 0.8 | 2.7 | 0.1 |
| 1M | 2 | 8 | 0.4 |
| 1M | 0.2 | 0.8 | 0.04 |
| 1M | 12 | 36 | 2.4 |
| 1M | 9 | 27 | 0.9 |
| 1M | 3 | 9 | 0.3 |
| 1M | 20 | 100 | 2 |
| 1M | 8 | 28 | 2 |
Price unit: CNY per million tokens. The deepseek-v4 series uses busy/idle time-based pricing; the table shows busy-time prices. All candidate models support thinking mode; thinking tokens are billed at output price.
NoteWhen selecting candidate models, note each model's quota limits to avoid being throttled when routed to a model due to insufficient quota. Choose models with sufficient quota, or raise the account's RPM/TPM quota for the relevant models before creation.
Invoke intelligent routing
Once created, calling intelligent routing is the same as calling a regular model—just replace the model parameter in the request with the router's model-code (e.g., auto-model-abcd1234); the client can integrate with zero code changes.
Example call:
curl -X POST https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/compatible-mode/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto-model-abcd1234",
"messages": [{"role": "user", "content": "Hello"}]
}'
from openai import OpenAI
client = OpenAI(
api_key="YOUR_DASHSCOPE_API_KEY",
base_url="https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/compatible-mode/v1",
)
resp = client.chat.completions.create(
model="auto-model-abcd1234",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)
print(resp.model) # the actual executing model name
# DashScope SDK (install first: pip install dashscope)
# Intelligent routing only supports maas.aliyuncs.com; configure DASHSCOPE_HTTP_BASE_URL
import os
os.environ['DASHSCOPE_HTTP_BASE_URL'] = 'https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1'
from dashscope import Generation
resp = Generation.call(
model="auto-model-abcd1234",
api_key="YOUR_DASHSCOPE_API_KEY",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.output.choices[0].message.content)
// OpenAI SDK (install first: npm install openai)
const OpenAI = require("openai");
const client = new OpenAI({
apiKey: process.env.DASHSCOPE_API_KEY,
baseURL: "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/compatible-mode/v1",
});
const resp = await client.chat.completions.create({
model: "auto-model-abcd1234",
messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.choices[0].message.content);
console.log(resp.model); // the actual executing model name
Response: the x-dashscope-resolved-model response header holds the actually-called model's identifier (use it to verify routing took effect); the body retains the actual executing model name and real token usage, and billing is by the actually-routed model.
Billing: Billed by the actual model each request is routed to; the intelligent routing service itself is free.
Auto-failover
Intelligent Routing has a built-in auto-failover, enabled by default with no configuration. When the target model is unavailable, the routing service reissues the whole request with another candidate model, not switching mid-inference.
The following do NOT trigger auto-failover; errors pass through with the actual model's original error info:
- 429 (rate limit)
- 400/403/404/500 and other deterministic errors (switching models would likely fail too)
- Timeout or network errors
- Errors after streaming has started returning content
Noteauto-failover does not double-bill: model A → auto-failover → model B succeeds, only model B is charged.
Real-time filtering rules
The routing service applies real-time filtering to the candidate set:
- Filters out models the current account has no permission to call.
- Filters out models that are offline.
- Filters out models incompatible with the current request's parameters (e.g., if the request carries
response_format, only candidate models supporting that parameter are routed to).
The actually-selectable subset for a single request changes dynamically with permissions, model state, and request parameters.
Usage limits
| Limit | Description |
|---|---|
| Call protocol | OpenAI-compatible interface only |
| Request type | Text Chat Completions input only; image, video input, image generation, Batch inference, Embedding/Rerank and other non-generation interfaces not supported |
| Context length | Upper bound = the minimum context window of the candidate-set models; over-long context errors directly without triggering auto-failover |
| Throttling | RPM/TPM throttled by actual routed model at account and workspace level; on failover, only the final successful model's consumption is counted |
Protocol support:
| Protocol | Supported |
|---|---|
| OpenAI-compatible interface (Chat Completions) | ✓ |
| DashScope protocol | ✗ |
| Anthropic protocol | ✗ |
Parameter notes
response_format: Routes only to candidate models supporting this parameter.enable_search: Routes only to candidate models supporting this parameter .reasoning_effort: Routes only to candidate models supporting this parameter by default; supported enum values:none,minimal,low,medium,high,xhigh,max.enable_thinking: Set tofalseto disable thinking, preferring non-thinking models; set totrueto automatically exclude non-thinking models and enable thinking.max_tokens: When set, it is automatically converted to themax_completion_tokensparameter and passed to candidate models; routing only selects candidate models that support themax_completion_tokensparameter.
In intelligent routing mode, the following parameters are not supported; setting them causes a runtime error:
| Parameter | Description |
|---|---|
top_logprobs | Limit on returned logprobs |
logit_bias | Token bias |
stop | Stop sequence |
tool_choice | Tool choice |
parallel_tool_calls | Parallel tool calls |
logprobs | Whether to return logprobs |
top_p | Nucleus sampling |
temperature | Sampling temperature |
presence_penalty | Presence penalty |
n | Number of generations |
thinking_budget | Thinking budget length |
Cache
Intelligent routing does not support explicit cache parameters:
- If the request carries an explicit cache identifier, the system ignores the explicit cache parameters.
- Explicit cache is not billed.
Only implicit cache provided by the actual model called is supported: cache may only be hit when a request is routed to the same model and the prompt prefix meets that model's cache conditions. Since different requests may be routed to different models, intelligent routing does not guarantee cache hit rate. For details on how caching works and supported models, see Context Cache.
Manage route service
View detail
After deployment, click the service name in the Dedicated Deployment list to enter the detail page:
- Basic info: service name,
model-code, running status, routing strategy tag. - Route configuration: route model version, routing strategy, candidate set and candidate model list.
- Billing: pay-as-you-go, billed by the actual model's unit price.
- Deploy config: each candidate model's TPM/RPM throttle quota.
Modify route config
Click "Edit" in the Route Configuration area to modify route version, strategy, and candidate models. Switching version changes the candidate set, taking effect within a few minutes without affecting live requests.
After selecting a specific version, online routing behavior remains long-term stable; please conduct a thorough evaluation before upgrading to avoid business impact. Current version: model-router-2026-0908. Route versions are not auto-updated.
Delete service
Delete the intelligent routing service from the Dedicated Deployment list. model-code is not reused after deletion.
Routing effect
This chapter covers routing effect metrics and monitoring. Effect metrics are viewed on the detail page's "Routing effect" tab; call monitoring is viewed on the Model Monitoring page.
Routing metrics
View the following metrics on the detail page's "Routing effect" tab. The displayed metrics are estimates, for reference only; not for billing settlement; actual billing is based on "Cost and Usage":
Metric | Meaning | Anomaly action |
|---|---|---|
Total routing requests | Total requests received (including failover-succeeded) | Investigate upstream throttling/anomalies on spikes |
Routing success rate | Forwarded requests ÷ total requests × 100% | On drop, check candidate model availability |
Actual model cost (estimated) | Full cost estimated at actual model list price (tiered billing: lowest tier; peak-valley: peak price; excludes free quota; deviates from actual billing) | On rise, check routing strategy and candidate set |
Baseline model cost (estimated) | Hypothetical cost if all handled by the baseline model (same estimation methodology) | — |
Savings | Baseline cost − actual cost (positive means routing saved cost) | — |
Billed request count | Successful requests that incur charges | — |
Spend share | Single-model cost ÷ total cost | On anomaly, check candidate-set distribution |
Cost notes
Cost analysis data has approximately a 1-hour delay and is not real-time. The above costs are all estimated by catalog price as full cost, excluding account discounts and free quotas, for reference only to evaluate routing savings. Tiered billing models use the lowest-tier catalog price; peak-valley pricing models use peak prices. Actual billing can be viewed in "Cost and Usage".
Model monitoring
For more granular call monitoring, go to the Model Monitoring page.
It shows model monitoring and logs at the auto-model dimension, not individual monitoring of the actually-routed models.