Best practices for decision model fine-tuning
Best practices for Alibaba Cloud Model Studio decision model fine-tuning — the full workflow of fine-tuning, deployment, and post-deployment invocation
This guide helps you complete decision model fine-tuning, deployment, and post-deployment invocation. The Model Studio decision model is provided under the model name decision-model-preview-2026-09-24 and supports direct invocation and fine-tuning; it does not generate text — a single forward pass outputs classification, scoring, and yes/no judgments together with their probabilities, making it suitable for high-frequency decision scenarios such as ticket routing, content moderation, agent routing, and result verification.
Full fine-tuning workflow: Environment preparation → Data preparation → Training submission → Model deployment → Post-deployment invocation.
Scenario | Choice | Typical use case |
|---|---|---|
Requires text generation (conversation, code, writing, reasoning) | General-purpose large model | Copywriting, coding, Q&A |
High-frequency decisions over a closed set (classification, scoring, yes/no judgment), no customization needed | Use the decision model directly (zero-shot invocation) | General ticket routing, content moderation |
Closed-set decisions, but with proprietary business rules or domain data | Fine-tune the decision model (this guide) | Business-specific ticket routing, proprietary scoring criteria |
Model Studio supports fine-tuning of the following models:
Model name | Description | Fine-tuning billing |
|---|---|---|
| Decision model; supports fine-tuning and invocation | Limited-time: 0 CNY per 1,000 tokens |
Before invocation, complete the following:
- API Key: Obtain an API Key from the Model Studio console API-KEY page
- Authorization: Contact your account manager to enable access
- Environment variable:
export DASHSCOPE_API_KEY="你的百炼 API Key"
Data preparation
Download the training data sample package (includes the sample data tickets.train.jsonl / tickets.development.jsonl, the question configuration workload.json, and the data scripts gen_data.py / split_data.py). The sample data is already split and ready to use; when using your own data, use the scripts to synthesize or split it (see README.md in the sample package).
gen_data.pycalls a Model Studio large model for automatic annotation, which incurs usage fees.
Each JSONL record contains the input state and the question table questions. Each question carries a type definition, a label (hard label), and an optional target (soft label):
{
"state": {"content": "发票抬头写错了,请帮忙重新开票。产品功能正常,也没有其他异常。"},
"questions": {
"department": {
"type": "choice",
"instructions": "应该由哪个团队处理?",
"criteria": {"billing": "支付、退款和账单问题", "technical": "产品故障和集成问题"},
"label": "billing",
"target": {"billing": 0.96, "technical": 0.04}
},
"escalate": {"type": "noul", "instructions": "是否需要立即通知值班人员?", "label": false},
"severity": {
"type": "score", "instructions": "这个问题有多严重?",
"criteria": ["轻微问题,不影响功能", "部分功能受影响,但存在替代方案", "核心功能不可用,没有替代方案", "造成严重业务或安全影响"],
"label": 0
}
}
}
Rules for filling in label and target:
Type | criteria | Required label | Optional target |
|---|---|---|---|
| An object mapping option names to their meanings | An option name, e.g. | Probabilities for each option, e.g. |
| Not required | A boolean | e.g. |
| An array of level meanings from low to high | A zero-based integer index | Keyed by string index, e.g. |
target should include all options, with probabilities between 0 and 1 that sum to 1. The question ID, type, instructions, and criteria must remain consistent across training, evaluation, and online requests for the same task; when you modify the options, update the labels and probability keys accordingly.
Training submission
Datasets are uploaded through the Model Studio OpenAPI (in JSONL format). After obtaining a file_id, reference it in the fine-tuning request. The China site endpoint is https://dashscope.aliyuncs.com and the Singapore site endpoint is https://dashscope-intl.aliyuncs.com; the API Key must match the site:
curl --request POST 'https://dashscope.aliyuncs.com/api/v1/files' \
--header 'Authorization: Bearer '${DASHSCOPE_API_KEY} \
--form 'files=@"/path/to/train.jsonl"' \
--form 'purpose="fine-tune"' \
--form 'descriptions="decision-model-preview training dataset"'
The response returns a file_id (e.g. 976bd01a-...); fill it into the fine-tuning request below. For details, see Training set and evaluation set and Fine-tuning data upload rules.
Dataset upload and fine-tuning requests go through the Model Studio OpenAPI, authenticated with
Authorization: Bearer <API-Key>.
Submit a fine-tuning job via the API:
curl --location --request POST 'https://dashscope.aliyuncs.com/api/v1/fine-tunes' \
--header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "decision-model-preview-2026-09-24",
"training_datasets": [
{"data_source_type": "file_id", "file_id": "your-file-id"}
],
"validation_datasets": [
{"data_source_type": "file_id", "file_id": "your-file-id"}
],
"hyper_parameters": {
"n_epochs": 2,
"learning_rate": "2e-5",
"batch_size": 1,
"save_strategy": "epoch",
"save_total_limit": 1
},
"training_type": "efficient_sft",
"finetuned_output_suffix": "mytune"
}'
Input parameters:
Field | Required | Description |
|---|---|---|
| Yes | List of training datasets |
| No | List of test datasets |
| Yes | Base model ID (supports |
| No | Hyperparameters; see the table below |
| Yes | Fine-tuning method; choose |
| No | Fine-tuning job name |
Hyperparameters:
Parameter | Default | Type | Purpose |
|---|---|---|---|
| 2 | Integer | Positive integer; the number of complete training epochs |
| 2e-5 | Float | Positive number; the learning rate |
| 1 | Integer | Positive integer; the number of samples per training batch |
Response example (excerpt):
{
"request_id": "your-request-id",
"output": {
"job_id": "ft-xxxxxxxx",
"status": "PENDING",
"finetuned_output": "decision-model-preview-2026-09-24-ft-xxxxxxxx",
"model": "decision-model-preview-2026-09-24",
"training_type": "efficient_sft"
}
}
finetuned_output is the model name produced by fine-tuning. After deployment, call it through the decision invocation API (pass this model name in the model field).
Training metrics
After submitting the job, you can view training progress and artifacts in Model Studio console model tuning. The metrics page is divided into three groups:
Group | Metric | Meaning | Direction |
|---|---|---|---|
Train |
| Training loss | — |
Train |
| Training epoch | — |
Eval |
| Accuracy | Higher is better |
Eval |
| Brier score; the mean squared error between the predicted probabilities and the ground truth | Lower is better |
Eval |
| Calibration error | Lower is better |
Eval |
| Negative log-likelihood | Lower is better |
Eval |
| Calibration temperature | — |
Eval |
| Number of evaluation questions | — |
Calibration |
| Accuracy after/before fine-tuning | Higher after is better |
Calibration |
| Brier score after/before fine-tuning | Lower after is better |
Calibration |
| Calibration error after/before fine-tuning | Lower after is better |
Calibration |
| Calibration temperature | — |
In the Calibration group, check whether _after improves relative to _before (acc increases, brier/ece decreases) — this indicates that fine-tuning improved calibration.
Model deployment
After training completes, the last checkpoint is automatically published to My models. To publish an intermediate checkpoint, go to the artifacts page of the Model tuning console and do it manually.
On the "My models" page, deploy the finetuned_output returned by training (e.g. decision-model-preview-2026-09-24-ft-xxxxxxxx) with deployment specification MU5×1. After deployment completes, you can call it through the decision invocation API. For deployment methods and MU unit prices, see Model deployment and Training and deployment pricing.
Post-deployment invocation
After deployment, call the fine-tuned artifact through the decision invocation API. Pass the finetuned_output returned by training (e.g. decision-model-preview-2026-09-24-ft-xxxxxxxx) in the model field; the other parameters are the same as when calling decision-model-preview:
curl
curl -sS -X POST https://dashscope.aliyuncs.com/compatible-mode/v1/systemone \
-H "Authorization: $DASHSCOPE_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "decision-model-preview-2026-09-24-ft-xxxxxxxx",
"state": {"content": "订单支付后超过 24 小时仍未到账,要求立即处理。"},
"questions": {
"department": {"type": "choice", "instructions": "应该由哪个团队处理?",
"criteria": {"billing": "支付、退款和账单问题", "technical": "产品故障和集成问题"}}
}
}'
Python
import os, requests
resp = requests.post(
"https://dashscope.aliyuncs.com/compatible-mode/v1/systemone",
headers={"Authorization": os.environ["DASHSCOPE_API_KEY"], "Content-Type": "application/json"},
json={"model": "decision-model-preview-2026-09-24-ft-xxxxxxxx",
"state": {"content": "订单支付后超过 24 小时仍未到账,要求立即处理。"},
"questions": {"department": {"type": "choice", "instructions": "应该由哪个团队处理?",
"criteria": {"billing": "支付、退款和账单问题", "technical": "产品故障和集成问题"}}}},
timeout=60,
)
print(resp.json()["answers"]["department"])
Best practices
Soft labels: teach the model calibration
In addition to the label (the answer), the training data's target carries a probability distribution auto-annotated by gen_data.py (from the sample package), which calls a Model Studio large model (default qwen3.8-max). This tells the model not just the answer but also how certain it is, making the output probabilities better calibrated. --target-temp adjusts the softness (default 1.5; the larger the value, the smoother and the more uncertainty is preserved). When resuming a run, you can reuse the saved annotations and only adjust the temperature — no further API calls and no additional fees.
Handling inputs that cannot be judged
Some inputs lack sufficient information to be judged (e.g. "help me check my ticket" with no details provided). Use --unknowable N to generate such samples; set their target to a uniform distribution and exclude them from accuracy evaluation. This specifically reduces the model's overconfidence when information is insufficient — without such samples, the model still tends to answer incorrectly with high confidence.
Keep question definitions consistent
Fine-tuning learns the mapping "input + question definition → answer" (the question definition is the ID, type, instructions, and criteria). Changing the question definition causes distribution drift — accuracy drops and calibration fails. When you must change it, re-prepare the data with the new definition and retrain.
Next steps
- Model deployment: Deployment methods and specifications
- Training and deployment pricing: Fine-tuning and deployment billing
- Fine-tune models in the console: View artifacts and publish in the console
- Fine-tuning data upload rules: Packaging rules, and size and quantity limits