Best practices for decision model fine-tuning

Updated at:

Best practices for Alibaba Cloud Model Studio decision model fine-tuning — the full workflow of fine-tuning, deployment, and post-deployment invocation

This guide helps you complete decision model fine-tuning, deployment, and post-deployment invocation. The Model Studio decision model is provided under the model name decision-model-preview-2026-09-24 and supports direct invocation and fine-tuning; it does not generate text — a single forward pass outputs classification, scoring, and yes/no judgments together with their probabilities, making it suitable for high-frequency decision scenarios such as ticket routing, content moderation, agent routing, and result verification.

Full fine-tuning workflow: Environment preparation → Data preparation → Training submission → Model deployment → Post-deployment invocation.

Scenario

Choice

Typical use case

Requires text generation (conversation, code, writing, reasoning)

General-purpose large model

Copywriting, coding, Q&A

High-frequency decisions over a closed set (classification, scoring, yes/no judgment), no customization needed

Use the decision model directly (zero-shot invocation)

General ticket routing, content moderation

Closed-set decisions, but with proprietary business rules or domain data

Fine-tune the decision model (this guide)

Business-specific ticket routing, proprietary scoring criteria

Model Studio supports fine-tuning of the following models:

Model name

Description

Fine-tuning billing

decision-model-preview-2026-09-24

Decision model; supports fine-tuning and invocation

Limited-time: 0 CNY per 1,000 tokens

Before invocation, complete the following:

  • API Key: Obtain an API Key from the Model Studio console API-KEY page
  • Authorization: Contact your account manager to enable access
  • Environment variable: export DASHSCOPE_API_KEY="你的百炼 API Key"

Data preparation

Download the training data sample package (includes the sample data tickets.train.jsonl / tickets.development.jsonl, the question configuration workload.json, and the data scripts gen_data.py / split_data.py). The sample data is already split and ready to use; when using your own data, use the scripts to synthesize or split it (see README.md in the sample package).

gen_data.py calls a Model Studio large model for automatic annotation, which incurs usage fees.

Each JSONL record contains the input state and the question table questions. Each question carries a type definition, a label (hard label), and an optional target (soft label):

{
  "state": {"content": "发票抬头写错了,请帮忙重新开票。产品功能正常,也没有其他异常。"},
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "应该由哪个团队处理?",
      "criteria": {"billing": "支付、退款和账单问题", "technical": "产品故障和集成问题"},
      "label": "billing",
      "target": {"billing": 0.96, "technical": 0.04}
    },
    "escalate": {"type": "noul", "instructions": "是否需要立即通知值班人员?", "label": false},
    "severity": {
      "type": "score", "instructions": "这个问题有多严重?",
      "criteria": ["轻微问题,不影响功能", "部分功能受影响,但存在替代方案", "核心功能不可用,没有替代方案", "造成严重业务或安全影响"],
      "label": 0
    }
  }
}

Rules for filling in label and target:

Type

criteria

Required label

Optional target

choice

An object mapping option names to their meanings

An option name, e.g. "billing"

Probabilities for each option, e.g. {"billing": 0.96, "technical": 0.04}

noul

Not required

A boolean true / false

e.g. {"true": 0.9, "false": 0.1}

score

An array of level meanings from low to high

A zero-based integer index

Keyed by string index, e.g. {"0": 1.0, "1": 0.0, "2": 0.0, "3": 0.0}

target should include all options, with probabilities between 0 and 1 that sum to 1. The question ID, type, instructions, and criteria must remain consistent across training, evaluation, and online requests for the same task; when you modify the options, update the labels and probability keys accordingly.

Training submission

Datasets are uploaded through the Model Studio OpenAPI (in JSONL format). After obtaining a file_id, reference it in the fine-tuning request. The China site endpoint is https://dashscope.aliyuncs.com and the Singapore site endpoint is https://dashscope-intl.aliyuncs.com; the API Key must match the site:

curl --request POST 'https://dashscope.aliyuncs.com/api/v1/files' \
  --header 'Authorization: Bearer '${DASHSCOPE_API_KEY} \
  --form 'files=@"/path/to/train.jsonl"' \
  --form 'purpose="fine-tune"' \
  --form 'descriptions="decision-model-preview training dataset"'

The response returns a file_id (e.g. 976bd01a-...); fill it into the fine-tuning request below. For details, see Training set and evaluation set and Fine-tuning data upload rules.

Dataset upload and fine-tuning requests go through the Model Studio OpenAPI, authenticated with Authorization: Bearer <API-Key>.

Submit a fine-tuning job via the API:

curl --location --request POST 'https://dashscope.aliyuncs.com/api/v1/fine-tunes' \
  --header "Authorization: Bearer ${DASHSCOPE_API_KEY}" \
  --header 'Content-Type: application/json' \
  --data-raw '{
    "model": "decision-model-preview-2026-09-24",
    "training_datasets": [
      {"data_source_type": "file_id", "file_id": "your-file-id"}
    ],
    "validation_datasets": [
      {"data_source_type": "file_id", "file_id": "your-file-id"}
    ],
    "hyper_parameters": {
      "n_epochs": 2,
      "learning_rate": "2e-5",
      "batch_size": 1,
      "save_strategy": "epoch",
      "save_total_limit": 1
    },
    "training_type": "efficient_sft",
    "finetuned_output_suffix": "mytune"
  }'

Input parameters:

Field

Required

Description

training_datasets

Yes

List of training datasets

validation_datasets

No

List of test datasets

model

Yes

Base model ID (supports decision-model-preview, or a model ID produced by another fine-tuning job)

hyper_parameters

No

Hyperparameters; see the table below

training_type

Yes

Fine-tuning method; choose efficient_sft (LoRA efficient fine-tuning)

job_name

No

Fine-tuning job name

Hyperparameters:

Parameter

Default

Type

Purpose

n_epochs

2

Integer

Positive integer; the number of complete training epochs

learning_rate

2e-5

Float

Positive number; the learning rate

batch_size

1

Integer

Positive integer; the number of samples per training batch

Response example (excerpt):

{
  "request_id": "your-request-id",
  "output": {
    "job_id": "ft-xxxxxxxx",
    "status": "PENDING",
    "finetuned_output": "decision-model-preview-2026-09-24-ft-xxxxxxxx",
    "model": "decision-model-preview-2026-09-24",
    "training_type": "efficient_sft"
  }
}

finetuned_output is the model name produced by fine-tuning. After deployment, call it through the decision invocation API (pass this model name in the model field).

Training metrics

After submitting the job, you can view training progress and artifacts in Model Studio console model tuning. The metrics page is divided into three groups:

Group

Metric

Meaning

Direction

Train

loss

Training loss

—

Train

epoch

Training epoch

—

Eval

acc

Accuracy

Higher is better

Eval

brier

Brier score; the mean squared error between the predicted probabilities and the ground truth

Lower is better

Eval

ece

Calibration error

Lower is better

Eval

nll

Negative log-likelihood

Lower is better

Eval

temperature

Calibration temperature

—

Eval

n_questions

Number of evaluation questions

—

Calibration

acc _after/_before

Accuracy after/before fine-tuning

Higher after is better

Calibration

brier _after/_before

Brier score after/before fine-tuning

Lower after is better

Calibration

ece _after/_before

Calibration error after/before fine-tuning

Lower after is better

Calibration

temperature

Calibration temperature

—

In the Calibration group, check whether _after improves relative to _before (acc increases, brier/ece decreases) — this indicates that fine-tuning improved calibration.

Model deployment

After training completes, the last checkpoint is automatically published to My models. To publish an intermediate checkpoint, go to the artifacts page of the Model tuning console and do it manually.

On the "My models" page, deploy the finetuned_output returned by training (e.g. decision-model-preview-2026-09-24-ft-xxxxxxxx) with deployment specification MU5×1. After deployment completes, you can call it through the decision invocation API. For deployment methods and MU unit prices, see Model deployment and Training and deployment pricing.

Post-deployment invocation

After deployment, call the fine-tuned artifact through the decision invocation API. Pass the finetuned_output returned by training (e.g. decision-model-preview-2026-09-24-ft-xxxxxxxx) in the model field; the other parameters are the same as when calling decision-model-preview:

curl

curl -sS -X POST https://dashscope.aliyuncs.com/compatible-mode/v1/systemone \
  -H "Authorization: $DASHSCOPE_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "decision-model-preview-2026-09-24-ft-xxxxxxxx",
    "state": {"content": "订单支付后超过 24 小时仍未到账,要求立即处理。"},
    "questions": {
      "department": {"type": "choice", "instructions": "应该由哪个团队处理?",
                     "criteria": {"billing": "支付、退款和账单问题", "technical": "产品故障和集成问题"}}
    }
  }'

Python

import os, requests

resp = requests.post(
    "https://dashscope.aliyuncs.com/compatible-mode/v1/systemone",
    headers={"Authorization": os.environ["DASHSCOPE_API_KEY"], "Content-Type": "application/json"},
    json={"model": "decision-model-preview-2026-09-24-ft-xxxxxxxx",
          "state": {"content": "订单支付后超过 24 小时仍未到账,要求立即处理。"},
          "questions": {"department": {"type": "choice", "instructions": "应该由哪个团队处理?",
                       "criteria": {"billing": "支付、退款和账单问题", "technical": "产品故障和集成问题"}}}},
    timeout=60,
)
print(resp.json()["answers"]["department"])

Best practices

Soft labels: teach the model calibration

In addition to the label (the answer), the training data's target carries a probability distribution auto-annotated by gen_data.py (from the sample package), which calls a Model Studio large model (default qwen3.8-max). This tells the model not just the answer but also how certain it is, making the output probabilities better calibrated. --target-temp adjusts the softness (default 1.5; the larger the value, the smoother and the more uncertainty is preserved). When resuming a run, you can reuse the saved annotations and only adjust the temperature — no further API calls and no additional fees.

Handling inputs that cannot be judged

Some inputs lack sufficient information to be judged (e.g. "help me check my ticket" with no details provided). Use --unknowable N to generate such samples; set their target to a uniform distribution and exclude them from accuracy evaluation. This specifically reduces the model's overconfidence when information is insufficient — without such samples, the model still tends to answer incorrectly with high confidence.

Keep question definitions consistent

Fine-tuning learns the mapping "input + question definition → answer" (the question definition is the ID, type, instructions, and criteria). Changing the question definition causes distribution drift — accuracy drops and calibration fails. When you must change it, re-prepare the data with the new definition and retrain.

Next steps