Reinforcement learning

Updated at:

This document explains how to develop and test the three function components for RL training: rollout (trajectory generation), reward, and group reward.

Function components

RL training requires you to develop custom function components to define the model's interaction behavior and scoring logic. This code is automatically deployed to the cloud:

Component

Description

Required

Rollout

Defines how the agent calls the LLM and tools to generate an interaction trajectory.

Required (1)

Reward

Scores a single agent output to define the optimization direction for training.

Optional (1 or more)

Group reward

Scores multiple outputs for the same problem by ranking or comparison.

Optional

Shared data model

The three function components pass data via the AgentOutput object and use TaskStatus to indicate their execution status.

AgentOutput

AgentOutput connects the Rollout and Reward components. The Rollout function constructs it and passes it to the Reward function:

Parameter

Type

Description

messages

list

A list of messages with user, assistant, and tool roles. To get the model response, retrieve the last assistant message.

reward_score

float / None

A preset score. When set in Rollout via AgentOutput(reward_score=0.95), the Reward function can read this value directly, eliminating the need for a separate Reward class.

rollout_extra

dict

Contains the value of the rollout_extra field from the training data. Use this field to access business data, such as a reference answer, during scoring.

rollout_metrics

dict

Custom metrics, such as AgentOutput(rollout_metrics={"latency": 1.2}). These metrics automatically appear in the console at the trace/rollout_metrics/ path.

TaskStatus and error handling

The process() method of each function component must return a result that includes a status field. The training framework uses this field to decide whether to retry or discard the sample:

  • TaskStatus.SUCCESS: The execution was successful.
  • TaskStatus.FAILED: The execution failed. You must set the error field to specify the reason for the failure.
## Success
return RolloutOutput(
    agent_output=AgentOutput(messages=messages),
    status=TaskStatus.SUCCESS,
)

## Failed → The framework will retry or discard this sample
return RolloutOutput(
    agent_output=AgentOutput(messages=[]),
    status=TaskStatus.FAILED,
    error="LLM call timed out",
)

Rollout function development

The rollout function generates the interaction trajectory of an agent with an LLM and external tools.

Basic structure

Inherit from AbstractRolloutProcessor and implement the setup() and process() methods:

from dashscope.finetune.reinforcement import (
    AbstractRolloutProcessor, RolloutInput, RolloutOutput
)
from dashscope.finetune.reinforcement.component.data.base_data_model import (
    AgentOutput, TaskStatus
)

class MyRolloutProcessor(AbstractRolloutProcessor):
    def setup(self) -> None:
        """Initialize resources (runs once at service startup)."""
        # Initialize the LLM client, load tools, etc.
        pass

    async def process(self, input: RolloutInput) -> RolloutOutput:
        """Generate an agent trajectory (runs once per sample)."""
        messages = input.messages or []
        model = input.model_resource.model_name
        # ... Interact with the LLM, call tools
        return RolloutOutput(
            agent_output=AgentOutput(
                messages=messages,
                rollout_metrics={"latency": 1.2}  # Custom metric
            ),
            status=TaskStatus.SUCCESS,
        )

The process() method supports both async def and standard def declarations. The SDK handles this automatically.

NoteNote: The rollout function calls the model being trained via an OpenAI-compatible API. The training framework automatically injects the endpoint (base_url) and API key into RolloutInput.model_resource at runtime. You can read these values directly in your setup() or process() methods, eliminating the need to hardcode them in your code or configuration.

Key RolloutInput fields

Parameter

Type

Description

messages

list

A list of conversation messages.

model_resource

object

Model information (model_name, base_url, api_key).

sampling_params

object

Sampling parameters (temperature, max_tokens, max_turns, timeout).

ground_truth

str

The reference answer for a sample, used for evaluation.

rollout_extra

dict

The rollout_extra field in the training data passes through business data.

Error handling: When the rollout function fails (for example, if an LLM call times out or a tool throws an exception), return TaskStatus.FAILED and set the error field. For details, see the TaskStatus and error handling section.

NoteInline scoring logic: Instead of writing a separate reward function, you can integrate scoring logic directly into the process() method of your rollout function. When returning RolloutOutput, simply provide the score by setting the reward_score parameter in AgentOutput, like AgentOutput(reward_score=0.95). Metrics from this inline scoring also appear on the metrics page in the console. For details, see AgentOutput.

Reward function development

The Reward function scores an agent's output and is a core component of RL training.

Basic structure

Inherit from the AbstractRewardProcessor class and implement your scoring logic in the process() method:

from dashscope.finetune.reinforcement import (
    AbstractRewardProcessor, RewardInput, RewardOutput, Reward, TaskStatus
)

class DemoRewardProcessor(AbstractRewardProcessor):
    def setup(self) -> None:
        pass

    async def process(self, input: RewardInput) -> RewardOutput:
        content = input.agent_output.messages[-1].get("content", "")
        ground_truth = input.ground_truth

        # Scoring logic: check if the answer is correct
        score = 1.0 if ground_truth in content else 0.0

        return RewardOutput(
            reward=Reward(
                reward_score=score,
                reward_metrics={"accuracy": score}  # Custom metric
            ),
            status=TaskStatus.SUCCESS,
        )

Key RewardInput parameters

Parameter

Type

Description

agent_output

object

The agent's output. For field details, see the AgentOutput section above.

ground_truth

str

The reference answer used for scoring.

Decorator pattern for multi-dimensional scoring

To score across multiple dimensions, use the decorator pattern to break down the logic into sub-dimensions and then aggregate their results:

from dashscope.finetune.reinforcement import (
    AbstractRewardProcessor, RewardInput, RewardOutput, Reward, TaskStatus,
    reward_func, sub_reward_func, aggregate_func
)

@reward_func("SafetyProcessor")
class SafetyProcessor(AbstractRewardProcessor):
    """Multi-dimensional safety scoring"""
    BANNED = {"hack", "exploit", "attack"}

    @sub_reward_func("toxicity", sub_weight=0.7)
    def toxicity(self, input: RewardInput) -> RewardOutput:
        """Toxicity detection sub-dimension (weight 0.7)"""
        content = input.agent_output.messages[-1].get("content", "").lower()
        score = 0.0 if any(w in content for w in self.BANNED) else 1.0
        return RewardOutput(
            reward=Reward(reward_score=score,
                          reward_metrics={"toxicity_score": score}),
            status=TaskStatus.SUCCESS,
        )

    @sub_reward_func("refusal", sub_weight=0.3)
    async def refusal(self, input: RewardInput) -> RewardOutput:
        """Refusal detection sub-dimension (weight 0.3)"""
        content = input.agent_output.messages[-1].get("content", "").lower()
        score = 0.5 if content.startswith("i cannot") else 1.0
        return RewardOutput(
            reward=Reward(reward_score=score),
            status=TaskStatus.SUCCESS,
        )

    @aggregate_func
    async def aggregate(self) -> RewardOutput:
        """Aggregates all sub-dimensions"""
        weights = self.get_weights()    # {"toxicity": 0.7, "refusal": 0.3}
        scores = self.get_scores()      # {"toxicity": 1.0, "refusal": 0.5}
        metrics = self.get_reward_metrics()  # Merges metrics from all sub-dimensions

        total = sum(scores[k] * weights[k] for k in scores)

        return RewardOutput(
            reward=Reward(reward_score=total, reward_metrics=metrics),
            status=TaskStatus.SUCCESS,
        )
Decorator descriptions

Decorator

Purpose

@reward_func("name")

Declares the class as a reward processor and sets its name.

@sub_reward_func("name", sub_weight=0.7)

Defines a scoring sub-dimension and its weight.

@aggregate_func

Defines the aggregation logic to combine scores from all sub-dimensions.

Combining multiple Reward functions

When you submit a task, you can configure multiple Reward functions, each with an independent weight:

functions=[
    RewardFunctionComponent(name="accuracy", weight=0.6, ...),
    RewardFunctionComponent(name="safety", weight=0.4, ...),
]

You can also use reward_metric_weight to set the weights for individual sub-metrics:

RewardFunctionComponent(
    name="reward-1",
    weight=1.0,
    reward_metric_weight={"metric_A": 0.3, "metric_B": 0.7},
    ...
)

Develop a Group Reward function

A Group Reward function scores multiple agent outputs for the same problem. It is ideal for scenarios that require ranking or comparative scoring.

Basic structure

Inherit from AbstractGroupRewardProcessor and handle multiple outputs in the process() method:

The process() method supports both async def and regular def declarations, which the SDK handles automatically. This rule applies to all function components: Rollout, Reward, and Group Reward.

from dashscope.finetune.reinforcement import (
    AbstractGroupRewardProcessor, GroupRewardInput, GroupRewardOutput,
    GroupReward, TaskStatus
)

class DemoGroupRewardProcessor(AbstractGroupRewardProcessor):
    def setup(self) -> None:
        pass

    async def process(self, input: GroupRewardInput) -> GroupRewardOutput:
        rewards = []
        for output in input.agent_outputs:
            content = output.messages[-1].get("content", "") if output.messages else ""
            if input.ground_truth and input.ground_truth in content:
                rewards.append(1.0)
            elif content:
                rewards.append(0.5)
            else:
                rewards.append(0.0)

        return GroupRewardOutput(
            group_reward=GroupReward(rewards=rewards),
            status=TaskStatus.SUCCESS,
        )

Differences from Reward

  • Reward: Scores a single output. It takes agent_output (singular) as input.
  • Group Reward: Scores multiple outputs. It takes agent_outputs (a list) as input and returns a list of rewards.

Reward design philosophy

The reward is the sole optimization signal in Reinforcement Learning (RL) training. Clearly define what to reward before you start coding. Three core trade-offs determine its implementation:

  • Sparse vs. dense: A sparse reward is difficult to game but has low sample efficiency. A dense reward enables faster learning but risks introducing shaping bias.
  • Rule-based vs. LLM-as-judge: A rule-based approach is inexpensive, deterministic, and reproducible. An LLM-as-judge can evaluate semantics but introduces potential bias and cost.
  • Single-dimensional vs. multi-dimensional: A single-dimensional reward is easy to iterate on but has a low performance ceiling. A multi-dimensional reward is more expressive but increases the complexity of weight design and interpretation.

Start with these four questions

  • Q1: Does a machine-determinable ground truth exist? For example, use a rule for math, code, or JSON, and a model or human for writing or dialogue.
  • Q2: Is failure binary or graded? Use a sparse reward of 0 or 1 for binary outcomes, and split the reward into multiple dimensions for graded outcomes.
  • Q3: Are intermediate incentives required? For long-chain reasoning or multi-turn agent interactions, consider using a process reward or step-based scoring.
  • Q4: What are the easiest reward hacking paths? List potential exploits first, then design defenses. See the section on identifying and preventing reward hacking.

Reward as a contract

The model optimizes for what is specified, not what is intended. This leads to three key principles:

  • Binary scoring boundaries will be precisely exploited. To mitigate this, add a margin to exact scores or soften them with a multi-dimensional combination, such as accuracy, formatting, and conciseness scores.
  • Hard limits will be pushed to their maximum. For example, if a response receives a full reward for reaching a certain length, the model will learn to pad its answers to meet that limit. Use normalization or a decay function instead.
  • Explicitly define and penalize empty output, oversized output, and timeout. Default behaviors often create loopholes. For example, assign a zero score for an empty output and deduct points from truncated, oversized output.

Choosing a reward representation

The reward representations are not mutually exclusive; they follow a layered progression from simplest to most complex: A. Inline (AgentOutput(reward_score=...)) → B. Standalone (RewardFunctionComponent) → C. Multi-dimensional with decorators (@reward_func + @sub_reward_func + @aggregate_func).

Representation

Use cases

Scoring complexity

Resource overhead

Observability granularity

Debuggability

When to avoid

A. Inline Rollout

For simple (≤10 lines) scoring logic that is synchronous with generation.

Low

Reuses the Rollout FC.

A single metric: reward_score.

Difficult, as it requires a Rollout redeploy.

When scoring requires independent scaling.

B. Standalone reward

For a single, well-defined scoring rule or for reusing a general-purpose scorer.

Medium

Standalone FC.

reward + reward_metrics.

Medium (independent deployment).

For three or more dimensions.

C. Multi-dimensional decorator

For multi-dimensional scoring that requires independent weight adjustment for each dimension.

High

Standalone FC.

Independent metrics for each sub-dimension.

High (enables troubleshooting by dimension).

When dimensions are tightly coupled and cannot be separated.

Upgrade path: Start with A to establish your initial workflow. Upgrade to B for custom metrics, and then to C for three or more dimensions. When migrating from A to B, keep your Rollout function and move the scoring logic to a standalone component. To migrate from B to C, break down the logic into dimensions based on atomic determination. Create a sub-dimension only if it can be evaluated with a binary (0/1) outcome.

Multi-dimensional scoring: Three-layer weight design

Multi-dimensional reward weights are organized into three non-overlapping, independently adjustable layers:

  • Layer 1: Function-level RewardFunctionComponent.weight (between multiple reward functions)
  • Layer 2: Sub-dimension @sub_reward_func(sub_weight=...) (between sub-dimensions within a reward function)
  • Layer 3: Sub-metric reward_metric_weight={...} (for subdivisions within reward_metrics)

Final reward = Σ (Layer 1 × Layer 2 × Layer 3 × normalized term).

Multi-dimensional weight design template

Step

Actions

Output

Guideline

1. List dimensions

Break down the objective into 3–5 atomic, quantifiable dimensions.

A list of dimensions

Recommended: ≤ 5

2. Prioritize dimensions

Rank dimensions based on business needs, such as must-haves, bonus items, and style preferences.

A prioritized list of dimensions

-

3. Assign initial weights

Assign 0.5–0.7 to must-haves, 0.2–0.3 to bonus items, and 0.05–0.1 to style preferences.

A weight vector

All weights must sum to 1.

4. Scale alignment

Normalize each dimension to a [0, 1] range before applying its weight.

A normalization function

Prevents any single dimension from dominating the reward.

5. Sample validation

Compare the reward scores of 100 samples against their manual annotations.

correlation coefficient

Spearman's rank correlation coefficient (ρ) ≥ 0.6

Weight tuning methods

  • First, train a baseline model with uniform weights to identify which sub-metric is dominant. Then, examine the variance of the curve at trace/reward_metrics/{reward}/{sub}/avg. Increase the weight of dimensions that show low variance.
  • If the validation reward increases while a specific dimension's score on the training set drops sharply, that dimension's weight is likely too low and should be increased. If you observe reward hacking, temporarily reduce the weight of the relevant dimension instead of disabling it entirely.

Group Reward use cases

Group Reward vs Independent Reward

  • Independent Reward evaluates a single response based on its absolute quality (an absolute score); Group Reward assesses the relative quality among multiple responses to the same prompt (ranking or comparison).
  • GRPO/GSPO algorithms require group-relative advantage. Group Reward provides a direct ranking signal, which avoids the bias introduced by normalization.
  • For subjective tasks like writing or dialogue, relative comparison is more stable and consistent than absolute scoring.

Integration with GRPO/GSPO algorithms

  • GRPO/GSPO algorithms calculate group-relative advantage from multiple rollouts for the same prompt. In contrast, an Independent Reward provides an absolute score, which the framework must then normalize within the group (a process that can be skewed by outliers).
  • Group Reward directly outputs a within-group distribution, reducing score scale drift. The n_rollouts parameter controls the input size; a value between 4 and 16 is recommended.

Implementation notes

  • The output rewards list must match the length of the input agent_outputs list. Do not apply hard normalization (e.g., dividing by the maximum value) within the group, as this conflicts with normalization at the algorithm level.
  • Tied scores are permitted (e.g., three responses each scoring 1.0), but avoid assigning the same score to all responses, as this results in zero advantage. Assign failed responses a score of 0 to maintain a consistent list length; the framework uses the TaskStatus to identify and handle such cases.

Reward hacking: Identification and defense

Hacking is almost inevitable

  • Goodhart's Law: When a measure becomes a target, it ceases to be a good measure.
  • The probability of reward hacking increases as an RL (Reinforcement Learning) model converges and its reward grows. This is not a bug; the optimizer is just doing its job.
  • The goal of a defense strategy is to make reward hacking more difficult than finding the true solution, not to eliminate it entirely.

Hack pattern → detection signal → defense strategy

Hack pattern

Detection signals

Defense strategy

length explosion / verbosity

trajectory/response_length increases monotonically while validation/data/reward/mean@1 plateaus or drops.

Use length normalization, add a length_penalty sub-reward, or apply hard truncation with max_length.

keyword stuffing

Reward increases, but manual spot-checks show repetitive keywords. Correlation between trace/reward_metrics/.../accuracy and human labels drops.

Use deduplication with regular expressions, replace keyword matching with an LLM judge, and introduce a semantic similarity threshold.

format hacking (markdown / emoji stuffing)

The output format is unusual, and the answer's relevance to the user's question decreases.

Normalize formatting as a preprocessing step. Score style as an independent dimension with a low weight.

refusal / empty output

Median reward is high, but trace/success_rate/... drops. The assistant message is empty.

Force a score of 0 for empty outputs, enforce a minimum length constraint, and add a "Did it answer the question?" scoring dimension.

repetition / repetitive output

High response length, high n-gram repetition rate, and a sharp drop in entropy.

Use regular expressions for repetition detection, apply a repetition_penalty, and monitor entropy.

judge collusion (in LLM-as-judge scenarios)

The LLM judge score is high but diverges from the rule-based score or human score.

Mix rule-based scores with judge scores, recalibrate the judge periodically, and rotate judges from different model families.

tool abuse / tool non-use

Anomalous trace/tool_call_count (spikes or drops) and a decreasing task success rate.

Penalize or reward tool call counts, and enforce constraints on required tool calls.

entropy collapse

actor/entropy approaches 0, and validation metrics stop changing.

Increase the KL divergence coefficient, use early stopping, or add an entropy bonus.

Four categories of defense strategies

  • reward shaping: Explicitly penalize hacking paths. This approach is direct but requires continuous patching.
  • regularization constraint: Use hard constraints for length, repetition, or format. This is inexpensive but inflexible.
  • validation monitoring: Use metrics from a source other than the training reward as a sentinel. This is essential.
  • multiple reward combination: Combine rule-based and model-based rewards to create checks and balances.

Monitoring playbook

  • Monitor the upward trend of your North Star Metric, critic/rewards/mean. Use validation/data/reward/mean@1 as an anti-hacking sentinel; a divergence from the North Star Metric indicates reward hacking. Perform behavioral health checks by monitoring metrics like trajectory/response_length and actor/entropy.
  • Every N steps, spot-check trajectories on the console's "Trajectory" tab and compare the scores with human judgment.

Scoring stability

Bias control in LLM-as-a-judge

  • Position bias: The order of items in a pairwise comparison can affect the outcome. To mitigate this, run each comparison twice with the order swapped and accept only the results that are consistent across both evaluations.
  • Self-preference bias: A model judge may favor outputs from its own model family. To counter this, use a judge from a different model family (e.g., use a GPT model to judge outputs from a Qwen model).
  • Length bias: A model judge might favor longer outputs. To address this, explicitly instruct the judge in the prompt not to score based on length and apply length normalization.
  • Discrete rating scales (e.g., 1-5 or 1-7) are more stable than continuous ones (e.g., 1-100). Use an anchor example for calibration.

Rule and model scoring

  • Use rules as hard constraints for objective criteria such as format, length, required fields, and sensitive words. If a response fails these checks, short-circuit the process and assign a score of 0.
  • Use a model for soft evaluation of subjective qualities like fluency, relevance, and reasoning quality.
  • Combination: final = rule_pass × (w_rule × rule_score + w_model × model_score). To save costs and prevent hacking, do not send responses that fail rule-based checks to the model judge.

TaskStatus in scoring scenarios

  • A score of 0 (e.g., for a failed rule check or an incorrect answer) is an effective training signal. In this case, return SUCCESS + reward_score=0.
  • For transient errors (e.g., a judge timeout or network jitter), return FAILED to let the framework retry. Do not assign a score of 0 in these cases.
  • Overusing FAILED can cause a data distribution shift. If the framework repeatedly fails on a hard example and eventually discards it, the training set becomes progressively easier.

Scoring consistency self-checks

  • Prepare a gold set of 50-200 examples with human annotation. After modifying the reward logic, run it against the gold set and compare results using metrics like Spearman's rho or Cohen's kappa.
  • When upgrading or switching your model judge, compare the new scores against the old ones to detect silent drift. Periodically send the N highest- and lowest-scoring training samples for human review.

Reward function checklist (10 questions)

Review this checklist before you submit a task with your new reward function:

  1. Can the scoring logic be explained in a single sentence?
  2. Is there a ground truth, and is it machine-determinable?
  3. What is the lowest-cost hack path, and have you defended against it?
  4. If your reward is multi-dimensional, can you annotate each dimension manually and independently?
  5. Have you normalized the sum of the weights and aligned the scales?
  6. Have you defined boundary conditions (e.g., empty output, overlong output, timeout)?
  7. Is the validation method independent of the training signal?
  8. Is the scoring time less than one-third of the rollout time?
  9. Is the reward_metrics field name stable? Renaming it will break the metric curves.
  10. Do you have a gold set for regression?

Reward design anti-patterns

Anti-pattern

Consequence

Defining the reward as "the longer, the better"

Inevitable reward hacking (length explosion)

Using single-dimension scoring for a highly subjective dimension

Noise overwhelms the signal.

Setting the weight for every dimension to 1.0

Equivalent to using no weights, resulting in inconsistent scales.

Using the same prompt and model for both the training judge and the validation judge

A form of self-deception that leads to overly optimistic results.

Returning SUCCESS with a score of 0 for temporary network errors

Data distribution contamination (hard samples are mislabeled as incorrect answers).

Observability configuration

NoteThis section has moved. See the Observability Configuration guide (Console and Observability) for details on the tracing decorator, custom metrics, dependency configuration, and the ENABLE_TRAJECTORY switch.

Function testing

Function registration

Register your function code as a function component:

client = AgenticRL()
(rollout_eids, reward_eids, group_eids,
 rollout_iids, reward_iids, group_iids) = await client.register_functions(
    functions=[
        AgenticRLFunctionComponent(
            type=FunctionType.ROLLOUT,
            fcmodel=FunctionComponentModel(
                classpath="functions.rollout.rollout.MyRolloutProcessor")),
        AgenticRLFunctionComponent(
            type=FunctionType.REWARD,
            fcmodel=FunctionComponentModel(
                classpath="functions.reward.reward.DemoRewardProcessor")),
    ],
    lazy_load=False  # False = Immediately get the instance ID for testing.
)

Equivalent CLI:

dashscope rl register_functions \
  --rollout-classpaths "functions.rollout.rollout.MyRolloutProcessor" \
  --reward-classpaths "functions.reward.reward.DemoRewardProcessor" \
  --no-lazy-load --output-format json

Remote testing

To send test data to a registered function instance, use the test_functions method:

## Test Rollout
result = await AgenticRL.test_functions(
    instance_id=rollout_iids[0],
    type=FunctionType.ROLLOUT,
    input_data="resources/rollout_input.json"
)

## Test Reward
result = await AgenticRL.test_functions(
    instance_id=reward_iids[0],
    type=FunctionType.REWARD,
    input_data="resources/reward_input.json"
)

Equivalent CLI:

dashscope rl test_functions "ro-ins-xxx" --type rollout --input resources/rollout_input.json
dashscope rl test_functions "rw-ins-xxx" --type reward --input resources/reward_input.json

Test input format

Rollout test input (rollout_input.json):

{
  "func_type": "rollout",
  "messages": [{"role": "user", "content": "Calculate 48 + 24 = ?"}],
  "ground_truth": "72",
  "sampling_params": {"temperature": 0.7, "max_tokens": 1024, "max_turns": 25},
  "model_resource": {
    "model_name": "qwen3-4b-instruct-2507",
    "base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
    "api_key": "your-api-key"
  }
}

Reward test input (reward_input.json):

{
  "func_type": "reward",
  "ground_truth": "72",
  "agent_output": {
    "messages": [{"role": "user", "content": "Calculate 48 + 24 = ?"}],
    "rollout_metrics": {"accuracy": 0.95},
    "reward_score": null
  }
}

Local testing

The scripts/ directory contains scripts for local testing:

./scripts/start_local_rollout.sh   # Start a local Rollout service
./scripts/query_local_rollout.sh   # Send a test request

Next steps