Intelligent Routing

Updated at:

Intelligent Routing dynamically analyzes request content and automatically matches the most suitable model from the candidate set, helping enterprises achieve optimal AI resource allocation in complex scenarios.

Overview

The model-code has a fixed auto-model prefix followed by 8 random characters (e.g., auto-model-abcd1234). The client calls this model-code → routing dynamically analyzes request content and matches the most suitable model from the candidate set → actual model infers → response (x-dashscope-resolved-model marks actual model, retains real token usage).

Scope:

DimensionDescription
RegionBeijing (cn-beijing) + Singapore (ap-southeast-1)
Model typeText input only
Call protocolOnly OpenAI-compatible interface (Chat Completions) is supported; DashScope and Anthropic protocols are not supported
Access domainOnly maas.aliyuncs.com (https://{WorkspaceId}.{region}.maas.aliyuncs.com/compatible-mode/v1, region is cn-beijing or ap-southeast-1); dashscope.aliyuncs.com (incl. compatible-mode) not supported
BillingPay-as-you-go
ThrottlingBy actual routed model, at account and workspace level

Key advantages:

  • Maximize model effectiveness: Automatically matches the optimal model for each task, ideal for users with extreme quality requirements.
  • Precise cost reduction: Routes simple tasks to lightweight models and complex tasks to high-end models, avoiding over-provisioning and achieving precise cost-demand alignment.
  • Eliminate model selection difficulty: No need for manual model evaluation or comparison; the platform auto-selects the optimal solution, drastically lowering the model selection barrier.
  • Minimal learning curve: Consistent with regular model access; call via model-code with zero code changes.

Create intelligent routing

In the Bailian console under "Model Inference > Dedicated Deployment", click Deploy New Model to enter the creation page, and select Intelligent Routing under "Deployment Mode".

Prerequisites

  • A Bailian workspace is activated.
  • The target route model version is published, and the current account has invocation permission for the version and candidate models.

Steps

  1. Enter a Service Name (required, up to 50 characters).
  2. Under Deployment Mode, select "Intelligent Routing".
  3. In Route Model Version, select a published version (maintained by the platform).
  4. Under Routing Strategy, select "Effect First" or "Cost First".
  5. Under Candidate Model Set, select candidate models (at least 2).
  6. Click Confirm Deploy to complete creation.

After deployment, the system generates a globally unique model-code (in the form auto-model-XXXXXXXX). The model-code is not reused after deletion.

Routing strategy

The routing strategy determines which model in the candidate set executes each request. Both strategies dynamically match by request content (not fixed routing); the choice is made at creation.

Strategy

Selection mechanism

Suitable scenarios

Effect First (EFFECT_FIRST)

Picks the strongest-effect model in the candidate set per request

Critical business, complex reasoning, high-quality content generation where quality is stringent

Cost First (COST_FIRST)

Intelligently schedules lightweight models while meeting a quality floor

Cost-sensitive, high-concurrency, latency- and unit-cost-aware scenarios

Candidate model configuration

The candidate model set is the range the routing service can choose from. When creating intelligent routing, select the models to participate in routing from the candidate list in the Candidate Model Set area—at least 2 candidate models are required, otherwise deployment is blocked.

NoteRAM users can only select candidate models they have

invocation permission for—if fewer than 2 are selectable, grant the RAM user model invocation authorization in that workspace first.

Route model version is the standardized release form of intelligent routing capability; each version specifies the model range and core capabilities supported by the routing system for that stage. The candidate list is determined by the selected route model version—each version corresponds to one candidate set; cross-version combination is not allowed. The currently published route model version and candidate models:

model-router-2026-0908

Candidate model

Context

Input

Output

Cache hit input

qwen3.8-max

1M

12

36

1.5

qwen3.8-flash

1M

0.8

2.7

0.1

qwen3.7-plus

1M

2

8

0.4

qwen3.7-flash

1M

0.2

0.8

0.04

qwen3.7-max

1M

12

36

2.4

deepseek-v4-pro-0813

1M

9

27

0.9

deepseek-v4-flash-0731

1M

3

9

0.3

kimi-k3

1M

20

100

2

glm-5.2

1M

8

28

2

Price unit: CNY per million tokens. The deepseek-v4 series uses busy/idle time-based pricing; the table shows busy-time prices. All candidate models support thinking mode; thinking tokens are billed at output price.

NoteWhen selecting candidate models, note each model's quota limits to avoid being throttled when routed to a model due to insufficient quota. Choose models with sufficient quota, or raise the account's RPM/TPM quota for the relevant models before creation.

Invoke intelligent routing

Once created, calling intelligent routing is the same as calling a regular model—just replace the model parameter in the request with the router's model-code (e.g., auto-model-abcd1234); the client can integrate with zero code changes.

Example call:

curl -X POST https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/compatible-mode/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto-model-abcd1234",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_DASHSCOPE_API_KEY",
    base_url="https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/compatible-mode/v1",
)
resp = client.chat.completions.create(
    model="auto-model-abcd1234",
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)
print(resp.model)  # the actual executing model name
# DashScope SDK (install first: pip install dashscope)
# Intelligent routing only supports maas.aliyuncs.com; configure DASHSCOPE_HTTP_BASE_URL
import os
os.environ['DASHSCOPE_HTTP_BASE_URL'] = 'https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1'
from dashscope import Generation

resp = Generation.call(
    model="auto-model-abcd1234",
    api_key="YOUR_DASHSCOPE_API_KEY",
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.output.choices[0].message.content)
// OpenAI SDK (install first: npm install openai)

const OpenAI = require("openai");

const client = new OpenAI({
  apiKey: process.env.DASHSCOPE_API_KEY,
  baseURL: "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/compatible-mode/v1",
});
const resp = await client.chat.completions.create({
  model: "auto-model-abcd1234",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.choices[0].message.content);
console.log(resp.model); // the actual executing model name

Response: the x-dashscope-resolved-model response header holds the actually-called model's identifier (use it to verify routing took effect); the body retains the actual executing model name and real token usage, and billing is by the actually-routed model.

Billing: Billed by the actual model each request is routed to; the intelligent routing service itself is free.

Auto-failover

Intelligent Routing has a built-in auto-failover, enabled by default with no configuration. When the target model is unavailable, the routing service reissues the whole request with another candidate model, not switching mid-inference.

The following do NOT trigger auto-failover; errors pass through with the actual model's original error info:

  • 429 (rate limit)
  • 400/403/404/500 and other deterministic errors (switching models would likely fail too)
  • Timeout or network errors
  • Errors after streaming has started returning content

Noteauto-failover does not double-bill: model A → auto-failover → model B succeeds, only model B is charged.

Real-time filtering rules

The routing service applies real-time filtering to the candidate set:

  • Filters out models the current account has no permission to call.
  • Filters out models that are offline.
  • Filters out models incompatible with the current request's parameters (e.g., if the request carries response_format, only candidate models supporting that parameter are routed to).

The actually-selectable subset for a single request changes dynamically with permissions, model state, and request parameters.

Usage limits

LimitDescription
Call protocolOpenAI-compatible interface only
Request typeText Chat Completions input only; image, video input, image generation, Batch inference, Embedding/Rerank and other non-generation interfaces not supported
Context lengthUpper bound = the minimum context window of the candidate-set models; over-long context errors directly without triggering auto-failover
ThrottlingRPM/TPM throttled by actual routed model at account and workspace level; on failover, only the final successful model's consumption is counted

Protocol support:

ProtocolSupported
OpenAI-compatible interface (Chat Completions)✓
DashScope protocol✗
Anthropic protocol✗

Parameter notes

  • response_format: Routes only to candidate models supporting this parameter.
  • enable_search: Routes only to candidate models supporting this parameter .
  • reasoning_effort: Routes only to candidate models supporting this parameter by default; supported enum values: none, minimal, low, medium, high, xhigh, max.
  • enable_thinking: Set to false to disable thinking, preferring non-thinking models; set to true to automatically exclude non-thinking models and enable thinking.
  • max_tokens: When set, it is automatically converted to the max_completion_tokens parameter and passed to candidate models; routing only selects candidate models that support the max_completion_tokens parameter.

In intelligent routing mode, the following parameters are not supported; setting them causes a runtime error:

ParameterDescription
top_logprobsLimit on returned logprobs
logit_biasToken bias
stopStop sequence
tool_choiceTool choice
parallel_tool_callsParallel tool calls
logprobsWhether to return logprobs
top_pNucleus sampling
temperatureSampling temperature
presence_penaltyPresence penalty
nNumber of generations
thinking_budgetThinking budget length

Cache

Intelligent routing does not support explicit cache parameters:

  • If the request carries an explicit cache identifier, the system ignores the explicit cache parameters.
  • Explicit cache is not billed.

Only implicit cache provided by the actual model called is supported: cache may only be hit when a request is routed to the same model and the prompt prefix meets that model's cache conditions. Since different requests may be routed to different models, intelligent routing does not guarantee cache hit rate. For details on how caching works and supported models, see Context Cache.

Manage route service

View detail

After deployment, click the service name in the Dedicated Deployment list to enter the detail page:

  • Basic info: service name, model-code, running status, routing strategy tag.
  • Route configuration: route model version, routing strategy, candidate set and candidate model list.
  • Billing: pay-as-you-go, billed by the actual model's unit price.
  • Deploy config: each candidate model's TPM/RPM throttle quota.

Modify route config

Click "Edit" in the Route Configuration area to modify route version, strategy, and candidate models. Switching version changes the candidate set, taking effect within a few minutes without affecting live requests.

After selecting a specific version, online routing behavior remains long-term stable; please conduct a thorough evaluation before upgrading to avoid business impact. Current version: model-router-2026-0908. Route versions are not auto-updated.

Delete service

Delete the intelligent routing service from the Dedicated Deployment list. model-code is not reused after deletion.

Routing effect

This chapter covers routing effect metrics and monitoring. Effect metrics are viewed on the detail page's "Routing effect" tab; call monitoring is viewed on the Model Monitoring page.

Routing metrics

View the following metrics on the detail page's "Routing effect" tab. The displayed metrics are estimates, for reference only; not for billing settlement; actual billing is based on "Cost and Usage":

Metric

Meaning

Anomaly action

Total routing requests

Total requests received (including failover-succeeded)

Investigate upstream throttling/anomalies on spikes

Routing success rate

Forwarded requests ÷ total requests × 100%

On drop, check candidate model availability

Actual model cost (estimated)

Full cost estimated at actual model list price (tiered billing: lowest tier; peak-valley: peak price; excludes free quota; deviates from actual billing)

On rise, check routing strategy and candidate set

Baseline model cost (estimated)

Hypothetical cost if all handled by the baseline model (same estimation methodology)

—

Savings

Baseline cost − actual cost (positive means routing saved cost)

—

Billed request count

Successful requests that incur charges

—

Spend share

Single-model cost ÷ total cost

On anomaly, check candidate-set distribution

Cost notes

Cost analysis data has approximately a 1-hour delay and is not real-time. The above costs are all estimated by catalog price as full cost, excluding account discounts and free quotas, for reference only to evaluate routing savings. Tiered billing models use the lowest-tier catalog price; peak-valley pricing models use peak prices. Actual billing can be viewed in "Cost and Usage".

Model monitoring

For more granular call monitoring, go to the Model Monitoring page. It shows model monitoring and logs at the auto-model dimension, not individual monitoring of the actually-routed models.