# 在线策略蒸馏可观测配置与指标参考

> 在线策略蒸馏训练的可观测手册：Tracing 接入、自定义指标上报、控制台页签、指标字典、训练效果判断与常见问题排查

在线策略蒸馏（OPD）训练自动记录每次 LLM 调用、工具调用与评分细节，在百炼控制台可视化展示。接入只需在函数代码中添加少量装饰器与包装调用。

<strong>相关文档：</strong>

- 在线策略蒸馏原理与示例选型 → 见 [在线策略蒸馏训练概述](opd-training-overview.md)
- Model OPD 开发 → 见 [Model OPD 开发](opd-model-development-guide.md)；Agentic OPD 开发 → 见 [Agentic OPD 开发](opd-agentic-development-guide.md)
- 提交参数与超参全集 → 见 [在线策略蒸馏训练配置](opd-training-config.md)

## 概述 <span id="opd-obs-overview" />

可观测能力基于 <strong>OpenTelemetry</strong>（开源可观测性标准）实现 Tracing（链路追踪），数据导出到 <strong>ARMS（应用实时监控服务）</strong>，在百炼控制台展示。<strong>Span</strong> 是 OpenTelemetry 中一次操作的观测单元，多个 Span 自动嵌套形成调用链。

五个环节：

1. 接入 Tracing — SDK 自动上报训练指标到控制台，如何在轨迹页查看
2. 自定义指标 — `rollout_metrics` 与 `reward_metrics` 的上报路径与约束
3. 控制台查看效果 — 指标 / 轨迹 / 产出 / 日志四页签怎么看
4. 训练指标参考 — 在线策略蒸馏特有指标字典 + 通用指标速查
5. 判断训练效果与排查 — 训练中判读、训练后评测、常见问题

## 接入 Tracing <span id="opd-tracing" />

以在线策略蒸馏项目的 Rollout 处理器为例，3 步完成 Tracing 接入。<strong>OPD 差异：</strong>调用链 LLM Span 对应学生模型调用，教师模型打分在服务端。

### Step 1：添加依赖 <span id="opd-tracing-step1" />

在项目根目录的 `requirements.txt` 中添加以下 OpenTelemetry 相关依赖：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "50%" }} /><col style={{ width: "50%" }} /></colgroup><thead><tr><th><p><strong>依赖</strong></p></th><th><p><strong>说明</strong></p></th></tr></thead><tbody><tr><td><p><code>opentelemetry-api==1.41.1</code></p></td><td><p>OTel API</p></td></tr><tr><td><p><code>opentelemetry-sdk==1.41.1</code></p></td><td><p>OTel SDK</p></td></tr><tr><td><p><code>opentelemetry-exporter-otlp-proto-http==1.41.1</code></p></td><td><p>OTLP HTTP 导出器</p></td></tr><tr><td><p><code>opentelemetry-processor-baggage==0.62b1</code></p></td><td><p>Baggage 处理器</p></td></tr><tr><td><p><code>loongsuite-util-genai==0.4.0</code></p></td><td><p>生成式 AI 工具库</p></td></tr></tbody></table>

运行环境中已预装 `dashscope`、`fastapi`、`uvicorn`、`pyyaml`；如需固定版本也可在 requirements.txt 中声明。

### Step 2：添加观测代码 <span id="opd-tracing-step2" />

代码改动集中在 5 个装饰器/函数。它们会自动嵌套，形成完整的调用链：

```plaintext
# 调用链结构：装饰器自动嵌套形成 Span 树，[ENTRY] 为入口 Span，[LLM]/[TOOL]/[custom] 为子 Span
[ENTRY: ROLLOUT] @observe_processor          ← Rollout 处理器入口
├── [LLM] trace_client / @observe_llm        ← LLM 调用（学生模型调用）
│   └── (OpenAI / LangChain / DashScope API)
├── [TOOL] trace_tool / @observe_tool        ← 工具调用
│   ├── tool: my_tool (MCP)
│   └── tool: my_scorer (自定义)
└── [custom] rollout_metrics                 ← Rollout 自定义指标

[ENTRY: REWARD] @observe_processor           ← Reward 处理器入口
├── [LLM] trace_client / @observe_llm        ← LLM 调用（评分用）
├── [TOOL] trace_tool / @observe_tool        ← 工具调用
└── [custom] reward_metrics                  ← Reward 自定义指标
```

#### @observe\_processor — 追踪处理器入口 <span id="opd-tracing-observe-processor" />

加在 `process()` 方法上，创建顶层 ENTRY Span（入口 Span，调用链最外层）。SDK 按继承链自动推断 Span 类型：继承 `AbstractRolloutProcessor` 判为 ROLLOUT，继承 `AbstractRewardProcessor` 判为 REWARD。自动记录输入、输出、耗时、成功/失败状态。

<strong>Rollout 侧示例（</strong>`functions/rollout/my_rollout.py`）：

```plaintext
from dashscope.finetune.reinforcement.component.observability import observe_processor
# 导入 Rollout 处理器基类与输入输出类型（项目内已定义）

class MyRolloutProcessor(AbstractRolloutProcessor):
    @observe_processor  # Span 类型 = ROLLOUT（按继承 AbstractRolloutProcessor 自动推断）
    async def process(self, input: RolloutInput) -> RolloutOutput:
        # input: Rollout 输入，含 prompt、模型资源、工具列表等
        await self._async_setup()  # 异步初始化：建 LLM 客户端、加载工具
        return await self._async_process(input)  # 执行实际 Rollout 逻辑，返回轨迹与指标
```

<strong>Reward 侧示例（</strong>`functions/reward/reward.py`）：

```plaintext
class MyRewardProcessor(AbstractRewardProcessor):
    @observe_processor  # Span 类型 = REWARD（按继承 AbstractRewardProcessor 自动推断）
    async def process(self, input: RewardInput) -> RewardOutput:
        # input: Reward 输入，含 agent_output（模型轨迹）与 ground_truth（标准答案）
        messages = input.agent_output.messages  # 取模型多轮对话消息列表
        content = messages[-1].get("content", "") if messages else ""  # 取最后一条 assistant 回复文本
        score = await evaluate(content, input.ground_truth)  # 用标准答案对回复评分，返回主分
        return RewardOutput(
            reward=Reward(reward_score=score, reward_metrics={...}),  # 主分入 reward_score，分维度入 reward_metrics
            status=TaskStatus.SUCCESS,  # 任务状态：SUCCESS 表示评分成功
            error=None,  # 无异常时为 None；失败时填错误信息
        )
```

#### trace\_client() — 追踪 LLM 客户端 <span id="opd-tracing-trace-client" />

在初始化或 process 方法内调用，包装 LLM 客户端实例。之后该客户端的所有 LLM 请求都会自动产生 LLM Span，记录模型名、请求内容、Token 用量、延迟。

<strong>支持的客户端类型（按对象属性自动识别）：</strong>

- OpenAI 客户端（`AsyncOpenAI` / `OpenAI`）
- OpenAI completions 资源（`.chat.completions`）
- LangChain ChatOpenAI 等类（通过 `.client` / `.async_client`）
- DashScope Generation 类（传入类本身，非实例）

<strong>示例（</strong>`functions/rollout/my_rollout.py`）：

```plaintext
from dashscope.finetune.reinforcement.component.observability import trace_client
# 导入 LLM 客户端包装函数，调用后该客户端的请求自动产生 LLM Span

class MyRolloutProcessor(AbstractRolloutProcessor):
    def _build_llm(self, input: RolloutInput) -> ChatOpenAI:
        # input: Rollout 输入，model_resource 含模型名等资源配置
        resource = input.model_resource  # 取模型资源（由 Runtime 注入的学生模型）
        llm = ChatOpenAI(model=resource.model_name, ...)  # 构造 LangChain ChatOpenAI 客户端
        trace_client(llm)  # 包装后自动追踪所有 LLM 调用（记录模型名、请求、Token 用量、延迟）
        return llm
```

#### trace\_tool() — 追踪工具调用 <span id="opd-tracing-trace-tool" />

获取工具实例后调用，包装工具对象。之后每次工具调用都会产生 TOOL Span，记录工具名、参数、返回值、延迟。

<strong>支持的输入格式：</strong>

- 单个 LangChain BaseTool
- 列表 / 元组 / 字典（自动遍历）
- LangGraph ToolNode（自动展开 `.tools_by_name`）
- MCP 工具（自动检测，provider 设为 "mcp"）
- 自定义 provider（`trace_tool(tool, provider="my-plugin")`）

<Warning>
  <strong>MCP 注意：</strong>MCP（Model Context Protocol，模型上下文协议）Server 和 Client 运行在不同进程中。Server 端的 `@observe_tool` 对 Client 端<strong>无效</strong>。必须在 Client 端调用 `get_tools()` 之后，对返回的工具列表执行 `trace_tool(tools)`。
</Warning>

#### @observe\_llm — 自定义 LLM 函数 <span id="opd-tracing-observe-llm" />

当 `trace_client()` 无法自动检测您的 LLM 客户端时，用此装饰器手动标记 LLM 调用函数。

<strong>签名要求：</strong>函数必须包含 `*` 后的关键字参数 `model` 和 `messages`。

<strong>示例（</strong>`functions/rollout/my_rollout.py`）：

```plaintext
from dashscope.finetune.reinforcement.component.observability import observe_llm
# 导入 LLM 装饰器，用于手动标记无法被 trace_client 自动识别的 LLM 调用

class MyRolloutProcessor(AbstractRolloutProcessor):
    @observe_llm  # 标记为 LLM Span（记录模型名、messages、Token 用量）
    async def _call_llm(self, *, messages: List[Dict], model: str) -> Any:
        # messages: 对话消息列表；model: 模型名（* 后关键字参数为签名硬性要求）
        ...
```

#### @observe\_tool — 自定义工具函数 <span id="opd-tracing-observe-tool" />

当 `trace_tool()` 无法自动检测您的工具时（如普通 Python 函数充当工具），用此装饰器手动标记。可通过 `name` 参数自定义 Span 名称。

<strong>示例（</strong>`functions/rollout/my_rollout.py`）：

```plaintext
from dashscope.finetune.reinforcement.component.observability import observe_tool
# 导入工具装饰器，用于手动标记无法被 trace_tool 自动识别的工具函数

class MyRolloutProcessor(AbstractRolloutProcessor):
    @observe_tool(name="my_scorer")  # 标记为 TOOL Span，name 自定义 Span 显示名
    def _score_response(self, *, messages: List[Dict]) -> float:
        # messages: 对话消息列表，函数返回评分为 float
        ...
```

### Step 3：提交任务 <span id="opd-tracing-step3" />

提交任务时，Runtime 的 `env` 字段留空即可——Tracing <strong>默认开启</strong>：

```plaintext
# FunctionComponentRuntime: Rollout/Reward 组件运行时配置
rollout_runtime = FunctionComponentRuntime(
    cpu=2, memory_size=4096, disk_size=512,  # 资源：2 核 CPU、4GB 内存、512MB 磁盘
    concurrency=30, capacity=30,  # 并发 30、容量 30（单实例同时处理请求数）
    min_capacity=30, max_capacity=60,  # 弹性伸缩下限 30、上限 60
    env={}  # 留空 = 默认开启 Tracing；设 {"ENABLE_TRAJECTORY": "false"} 可关闭
)
```

<Note>
  <strong>如需关闭 Tracing</strong>（节省成本），在 `env` 中设置 `{"ENABLE_TRAJECTORY": "false"}` 即可。关闭后系统指标不受影响，仅 Tracing 数据停止采集。
</Note>

## Tracing 性能与成本权衡 <span id="opd-tracing-cost" />

### 数据链路与成本来源

- <strong>数据流</strong>：函数代码（装饰器）→ OpenTelemetry SDK → ARMS → 控制台轨迹/指标页签
- <strong>成本来源</strong>：ARMS Span 存储（按量计费）+ 函数侧少量 CPU/网络开销 + 训练延迟少量增加

### 开关策略与成本治理

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "20%" }} /><col style={{ width: "20%" }} /><col style={{ width: "20%" }} /><col style={{ width: "20%" }} /><col style={{ width: "20%" }} /></colgroup><thead><tr><th><p>阶段</p></th><th><p>Tracing 状态</p></th><th><p>采集数据</p></th><th><p>不采集数据</p></th><th><p>成本影响</p></th></tr></thead><tbody><tr><td><p>开发 / 小批量调试</p></td><td><p><strong>全开</strong></p></td><td><p>全部（轨迹 + Reward 分析 + 工具调用 + 系统指标）</p></td><td><p>—</p></td><td><p>低</p></td></tr><tr><td><p>灰度 / 发布前确认</p></td><td><p><strong>保留</strong></p></td><td><p>全部</p></td><td><p>—</p></td><td><p>中</p></td></tr><tr><td><p>大规模正式训练</p></td><td><p><strong>可关闭</strong></p></td><td><p>actor/critic/trajectory/timing 系统指标</p></td><td><p>轨迹回放 / 工具调用详情 / 自定义 metrics 曲线</p></td><td><p>显著降低</p></td></tr></tbody></table>

<strong>OPD 差异：</strong>关闭 Tracing 后保留的系统指标见 §训练指标参考；`trace/distillation/*` 属系统指标，关闭后仍可查看（教师模型健康排查不受影响）。

自定义指标的成本治理：

- **数量与基数**：`reward_metrics` / `rollout_metrics` 数量与基数（cardinality）影响 ARMS 存储
- **避免高基数字段**：user\_id / request\_id 不当指标 key
- **精简指标**：关键指标 ≤10 个，不留临时调试指标

## 自定义指标 <span id="opd-custom-metrics" />

在代码中通过以下入口定义的 key-value 指标，会自动出现在控制台<strong>指标</strong>页签的 `trace/` 分组和<strong>Reward 分析</strong>页面：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "33.333333%" }} /><col style={{ width: "33.333333%" }} /><col style={{ width: "33.333333%" }} /></colgroup><thead><tr><th><p><strong>入口</strong></p></th><th><p><strong>代码位置</strong></p></th><th><p><strong>控制台路径</strong></p></th></tr></thead><tbody><tr><td><p>reward\_metrics</p></td><td><p><code>Reward(reward\_metrics=\{"acc": 0.8})</code></p></td><td><p><code>trace/reward\_metrics/\{reward-name}/acc/\{avg,sum}</code></p></td></tr><tr><td><p>rollout\_metrics</p></td><td><p><code>AgentOutput(rollout\_metrics=\{"latency": 1.2})</code></p></td><td><p><code>trace/rollout\_metrics/latency/\{avg,sum}</code></p></td></tr></tbody></table>

<strong>值类型约束：</strong>

- **扁平字典**：reward\_metrics / rollout\_metrics 返回 flat dict（扁平字典，非嵌套 list），每值为 float
- **主分单独上报**：reward\_score 单独放入 `Reward.reward_score`，不重复放入 reward\_metrics
- **非 float 只落盘**：非 float 数据通过 `rollout_extra` 只落盘不进指标
- **指标路径**：reward\_score 主分对应 `trace/reward/<name>/`，各维度对应 `trace/reward_metrics/<name>/<metric>`

<strong>多 Reward 函数场景：</strong>通过 `RewardFunctionComponent(name="reward-1")` 为每个 Reward 函数设唯一名称，指标路径自动区分（`trace/reward_metrics/reward-1/...`）。

## 控制台查看效果 <span id="opd-console" />

训练开始后，在[模型调优控制台](https://bailian.console.aliyun.com/cn-beijing/model/tuning)进入<strong>模型调优</strong>页面，点击任务名称进入详情查看观测数据。以下说明"您在代码中做了什么 → 在控制台看到什么"的对应关系。

<strong>OPD 差异：</strong>任务详情页无独立"详情"页签，状态在任务列表查看。关注点：指标页看蒸馏损失、教师模型健康与验证 reward；轨迹页逐样本看工具调用链；产出页的最后一个 Checkpoint 自动发布。

任务详情包含以下页签，按"先看进度 → 再看行为 → 出问题再下钻"顺序使用：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "33.333333%" }} /><col style={{ width: "33.333333%" }} /><col style={{ width: "33.333333%" }} /></colgroup><thead><tr><th><p>页签</p></th><th><p>主要回答</p></th><th><p>何时使用</p></th></tr></thead><tbody><tr><td><p>轨迹</p></td><td><p>模型实际做了什么、得分为何</p></td><td><p>验证模型行为、归因低分样本</p></td></tr><tr><td><p>指标</p></td><td><p>训练定量趋势</p></td><td><p>看曲线判断收敛性、发现拐点</p></td></tr><tr><td><p>产出</p></td><td><p>Checkpoint 列表与发布</p></td><td><p>训练完成后选模型</p></td></tr><tr><td><p>日志</p></td><td><p>stdout / stderr / 报错堆栈</p></td><td><p>FAILED 时排查</p></td></tr></tbody></table>

### 轨迹与行为分析 <span id="opd-console-trajectory" />

<strong>轨迹详情 — 对应 @observe\_processor：</strong>在<strong>轨迹</strong>页签的<strong>轨迹详情</strong>子页面，可以看到每次 Rollout 的完整交互过程：

- <strong>轨迹列表：</strong>展示所有采样轨迹，支持按 Sample ID / 轨迹 ID / Epoch / Step 筛选
- <strong>对话过程：</strong>完整多轮交互（user → assistant → tool\_call → tool\_result → assistant），直观看到模型的推理链
- <strong>Reward 分数：</strong>每个 Step 显示对应的 Reward 分数和状态（SUCCESS/FAILED）

关注：工具调用是否正确？对话轮次是否合理？模型是否在重复无效操作？

<span id="opd-console-tools" />

<strong>工具调用分析 — 对应 trace\_tool / @observe\_tool：</strong>在<strong>轨迹</strong>页签的<strong>工具调用分析</strong>子页面，可以查看：

- <strong>工具调用记录：</strong>工具名称、调用参数、返回结果和耗时
- <strong>Tracing 子页签：</strong>每条轨迹的 Span 树，可展开查看每次工具调用和 LLM 请求的完整详情

<strong>典型用途：</strong>排查 Agent 的工具调用失败——哪个工具报错？参数传递是否正确？耗时是否过长？

<span id="opd-console-reward" />

<strong>Reward 分析 — 对应 reward\_metrics：</strong>

<strong>概念定义：</strong>

- <strong>Sample</strong>：训练数据中一条原始样本（一道题、一条指令、一个 prompt）
- <strong>Trajectory</strong>：同一 Sample 在 `n_rollouts` 次采样下产生的具体一条交互轨迹
- 关系：一条 Sample → N 条 Trajectory（N = `n_rollouts`）

在<strong>轨迹</strong>页签的<strong>Reward 分析</strong>子页面，从三个维度评估训练效果：

- <strong>Step 维度：</strong>选择训练 Step，查看该 Step 下所有样本的 Reward 聚合（平均分、成功率、趋势图），判断整体训练趋势
- <strong>Sample 维度：</strong>选择 Sample ID，查看同一样本在不同轨迹中的 Reward 对比，发现问题样本
- <strong>Trajectory 维度：</strong>查看单条轨迹的每个评分维度原始分，用于归因分析

### 指标、产出与日志 <span id="opd-console-metrics" />

<strong>指标页签 — 对应 rollout\_metrics / reward\_metrics：</strong>在<strong>指标</strong>页签的 `trace/` 分组，可以看到您在代码中定义的所有自定义指标的聚合曲线（avg / sum）。完整指标分组见下方 §训练指标参考；看到指标异常如何归因 → §判断训练效果与排查。

<span id="opd-console-output" />

<strong>产出页签：</strong>训练完成后，在<strong>产出</strong>页签中可以看到 Checkpoint 列表，每行包含 Checkpoint ID、发布状态和剩余保存时间。

1. 选择目标 Checkpoint，点击<strong>发布</strong>按钮
2. 等待发布完成（状态从"待发布"变为"已发布"）
3. 发布后即可通过模型名称在 API 中调用

<strong>OPD 差异：</strong>最后一个 Checkpoint 自动发布，产出页仍可手动选任意 Checkpoint 发布（与自动发布的最后一个并存，调用时按模型名区分）。训练完成不等于训练成功。启用 Reward 时建议以 `validation/data/reward/mean@1` 最优的 Checkpoint 为准，而非默认的最后一个。纯蒸馏（不注册 reward 组件）无 validation reward，依据 loss 趋势与独立评测判断效果。

<span id="opd-console-logs" />

<strong>日志页签：</strong>在<strong>日志</strong>页签查看训练运行日志，也可通过 SDK / CLI 获取：

- SDK：`AgenticRL.logs(job_id="ft-xxx", lines=100)`
- CLI：`dashscope rl logs "ft-xxx" --lines 100`

FAILED 任务排查步骤见 §常见问题 → FAILED 排查。

## 训练指标参考 <span id="opd-metrics-ref" />

以下为在线策略蒸馏特有指标与通用指标全集。指标在「指标」页展示，点击展开各分组详情。

### 在线策略蒸馏特有指标 <span id="opd-metrics-opd" />

以下为在线策略蒸馏特有指标：

<Accordion title="distillation/ — 蒸馏损失">
  <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "50%" }} /><col style={{ width: "50%" }} /></colgroup><thead><tr><th><p><strong>指标</strong></p></th><th><p><strong>含义与判读</strong></p></th></tr></thead><tbody><tr><td><p><code>actor/distillation/loss</code></p></td><td><p>蒸馏损失，下降表示学生模型正在向教师模型靠拢</p></td></tr></tbody></table>

  启用 reward 时，`actor/distillation/loss` 与 `actor/pg_loss` 共同构成总损失，两者都应下降或趋稳。
</Accordion>

<Accordion title="trace/distillation/ — 教师模型信号健康">
  <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "50%" }} /><col style={{ width: "50%" }} /></colgroup><thead><tr><th><p><strong>指标</strong></p></th><th><p><strong>含义与判读</strong></p></th></tr></thead><tbody><tr><td><p><code>trace/distillation/valid\_token\_rate</code></p></td><td><p>教师模型成功给出 logprob（对数概率）的 token 占比，正常接近 1.0；小于 1 表示打分缺失</p></td></tr><tr><td><p><code>trace/distillation/missing\_token\_rate</code></p></td><td><p>教师模型未给出 logprob 的 token 占比，正常接近 0</p></td></tr><tr><td><p><code>trace/distillation/empty\_response\_rate</code></p></td><td><p>教师模型返回空响应比例，正常接近 0；升高指向教师模型侧异常</p></td></tr></tbody></table>
</Accordion>

<Accordion title="validation/ — 验证集指标（启用 Reward 时上报）">
  <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "50%" }} /><col style={{ width: "50%" }} /></colgroup><thead><tr><th><p><strong>指标</strong></p></th><th><p><strong>含义与判读</strong></p></th></tr></thead><tbody><tr><td><p><code>validation/data/reward/mean\@1</code></p></td><td><p>验证集聚合 reward 均值，相对基线上升为健康</p></td></tr></tbody></table>

  验证指标只有聚合 reward 均值。Reward 函数返回的自定义维度不在验证指标中体现，需在训练指标 `trace/reward_metrics/<reward_name>/<metric>` 下查看（维度名取自 `reward_metrics` 的键，因任务而异）。纯蒸馏（不注册 reward 组件，functions=None）不上报该组。
</Accordion>

### 关键指标速查 <span id="opd-metrics-quick" />

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "25%" }} /><col style={{ width: "40%" }} /><col style={{ width: "35%" }} /></colgroup><thead><tr><th><p><strong>类别</strong></p></th><th><p><strong>指标</strong></p></th><th><p><strong>判读</strong></p></th></tr></thead><tbody><tr><td><p>蒸馏</p></td><td><p><code>actor/distillation/loss</code></p></td><td><p>下降或趋稳</p></td></tr><tr><td><p>教师模型健康</p></td><td><p><code>trace/distillation/valid\_token\_rate</code></p></td><td><p>接近 1.0</p></td></tr><tr><td><p>教师模型健康</p></td><td><p><code>trace/distillation/missing\_token\_rate</code></p></td><td><p>接近 0</p></td></tr><tr><td><p>教师模型健康</p></td><td><p><code>trace/distillation/empty\_response\_rate</code></p></td><td><p>接近 0</p></td></tr><tr><td><p>验证</p></td><td><p><code>validation/data/reward/mean\@1</code></p></td><td><p>相对基线上升</p></td></tr><tr><td><p>reward 维度</p></td><td><p><code>trace/reward\_metrics/\<reward\_name>/\<metric></code></p></td><td><p>关键业务维度不回退</p></td></tr></tbody></table>

### 通用指标 <span id="opd-metrics-common" />

训练过程中产出的指标按前缀分为以下分组，各分组的详细指标说明见下方折叠面板：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "25%" }} /><col style={{ width: "25%" }} /><col style={{ width: "25%" }} /><col style={{ width: "25%" }} /></colgroup><thead><tr><th><p><strong>分组前缀</strong></p></th><th><p><strong>指标数</strong></p></th><th><p><strong>类型</strong></p></th><th><p><strong>说明</strong></p></th></tr></thead><tbody><tr><td><p><strong>actor/</strong></p></td><td><p>8</p></td><td><p>系统</p></td><td><p>策略网络指标：损失、熵、KL 散度、裁剪率、梯度范数、学习率</p></td></tr><tr><td><p><strong>critic/</strong></p></td><td><p>12</p></td><td><p>系统</p></td><td><p>奖励与价值评估：score / rewards / advantages / returns 的 mean / max / min</p></td></tr><tr><td><p><strong>trajectory/</strong></p></td><td><p>15</p></td><td><p>系统</p></td><td><p>轨迹统计：回复长度、Prompt 长度、截断率、中止率、对话轮次</p></td></tr><tr><td><p><strong>trace/</strong></p></td><td><p>40+</p></td><td><p><strong>混合</strong></p></td><td><p>可观测性指标：epoch、LLM 调用次数、成功率、自定义 metrics</p></td></tr><tr><td><p><strong>timing/</strong></p></td><td><p>11</p></td><td><p>系统</p></td><td><p>耗时分析：Trainer 阶段耗时、Rollout 耗时、每 Token 耗时</p></td></tr></tbody></table>

<strong>OPD 差异：</strong>在线策略蒸馏不产出 `perf/` 和 `fully_async/` 两组指标。

点击展开各分组的完整指标列表：

<Accordion title="actor/ — 策略网络指标（8 个）">
  <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "50%" }} /><col style={{ width: "50%" }} /></colgroup><thead><tr><th><p><strong>指标</strong></p></th><th><p><strong>含义</strong></p></th></tr></thead><tbody><tr><td><p><code>actor/loss</code></p></td><td><p>总 loss（pg + entropy + ...）</p></td></tr><tr><td><p><code>actor/pg\_loss</code></p></td><td><p>策略梯度损失</p></td></tr><tr><td><p><code>actor/entropy</code></p></td><td><p>当前策略的平均 token 熵，反映探索程度</p></td></tr><tr><td><p><code>actor/ppo\_kl</code></p></td><td><p>当前策略相对于初始策略的 KL 散度</p></td></tr><tr><td><p><code>actor/pg\_clipfrac</code></p></td><td><p>重要性采样裁剪触发率，反映策略漂移速度</p></td></tr><tr><td><p><code>actor/pg\_clipfrac\_lower</code></p></td><td><p>Dual-clip 下方裁剪触发率（未开启 dual-clip 时恒为 0）</p></td></tr><tr><td><p><code>actor/grad\_norm</code></p></td><td><p>梯度范数</p></td></tr><tr><td><p><code>actor/lr</code></p></td><td><p>当前学习率</p></td></tr></tbody></table>
</Accordion>

<Accordion title="critic/ — 奖励与价值评估（12 个）">
  <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "50%" }} /><col style={{ width: "50%" }} /></colgroup><thead><tr><th><p><strong>指标</strong></p></th><th><p><strong>含义</strong></p></th></tr></thead><tbody><tr><td><p><code>critic/score/\{mean,max,min}</code></p></td><td><p>原始 reward score 统计（扣除 KL 前）</p></td></tr><tr><td><p><code>critic/rewards/\{mean,max,min}</code></p></td><td><p>扣除 KL 惩罚后的最终训练 Reward 统计</p></td></tr><tr><td><p><code>critic/advantages/\{mean,max,min}</code></p></td><td><p>优势函数统计，反映当前策略相对基线的改进</p></td></tr><tr><td><p><code>critic/returns/\{mean,max,min}</code></p></td><td><p>Returns（Critic target）统计</p></td></tr></tbody></table>
</Accordion>

<Accordion title="trajectory/ — 轨迹统计（15 个）">
  <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "50%" }} /><col style={{ width: "50%" }} /></colgroup><thead><tr><th><p><strong>指标</strong></p></th><th><p><strong>含义</strong></p></th></tr></thead><tbody><tr><td><p><code>trajectory/response\_length/\{mean,max,min}</code></p></td><td><p>响应 token 数统计（包含 abort 样本）</p></td></tr><tr><td><p><code>trajectory/response/aborted\_ratio</code></p></td><td><p>响应长度为零的轨迹比例</p></td></tr><tr><td><p><code>trajectory/response\_length\_non\_aborted/\{mean,max,min}</code></p></td><td><p>排除 abort 后的有效响应 token 数统计</p></td></tr><tr><td><p><code>trajectory/response\_length/clip\_ratio</code></p></td><td><p>Response 达到最大长度被截断的比例</p></td></tr><tr><td><p><code>trajectory/prompt\_length/\{mean,max,min}</code></p></td><td><p>Prompt token 数统计</p></td></tr><tr><td><p><code>trajectory/prompt\_length/clip\_ratio</code></p></td><td><p>Prompt 达到最大长度被截断的比例</p></td></tr><tr><td><p><code>trajectory/num\_turns/\{mean,max,min}</code></p></td><td><p>Agent 与 LLM 交互的轮数统计</p></td></tr></tbody></table>
</Accordion>

<Accordion title="trace/ — 可观测性指标">
  <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "50%" }} /><col style={{ width: "50%" }} /></colgroup><thead><tr><th><p><strong>指标</strong></p></th><th><p><strong>含义</strong></p></th></tr></thead><tbody><tr><td><p><code>trace/training/epoch</code></p></td><td><p>当前训练 epoch</p></td></tr><tr><td><p><code>trace/num\_llm\_calls/\{avg,sum}</code></p></td><td><p>每条 / 总 LLM 调用次数</p></td></tr><tr><td><p><code>trace/success\_rate/agent/\{avg,sum}</code></p></td><td><p>Agent 任务成功率 / 累计成功条数</p></td></tr><tr><td><p><code>trace/success\_rate/reward/\{avg,sum}</code></p></td><td><p>Reward 计算成功率 / 累计成功条数</p></td></tr><tr><td><p><code>trace/attempts/agent/\{avg,sum}</code></p></td><td><p>Agent 平均 / 累计 HTTP 尝试次数（含重试）</p></td></tr><tr><td><p><code>trace/reward/\<reward\_name>/\{avg,sum}</code></p></td><td><p>单个 Reward 函数的平均 / 累计值</p></td></tr><tr><td><p><code>trace/reward\_metrics/\<reward\_name>/\<metric>/...</code></p></td><td><p>Reward 函数返回的自定义子指标</p></td></tr></tbody></table>
</Accordion>

<Accordion title="timing/ — 耗时分析（11 个）">
  分为三个子组：`timing/s/*`（Trainer 阶段耗时，秒）、`timing/s/rollout/*`（Rollout 侧耗时，秒）、`timing/ms/*_per_token`（每 Token 耗时，毫秒）。

  <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "50%" }} /><col style={{ width: "50%" }} /></colgroup><thead><tr><th><p><strong>指标</strong></p></th><th><p><strong>含义</strong></p></th></tr></thead><tbody><tr><td><p><code>timing/s/step</code></p></td><td><p>完整一个 Trainer step 的总耗时</p></td></tr><tr><td><p><code>timing/s/trainer\_fetch\_batch</code></p></td><td><p>Trainer 等待并拉取一个 batch 的耗时</p></td></tr><tr><td><p><code>timing/s/old\_log\_prob</code></p></td><td><p>旧 policy 的 log-prob 计算耗时</p></td></tr><tr><td><p><code>timing/s/adv</code></p></td><td><p>Advantage 计算耗时</p></td></tr><tr><td><p><code>timing/s/update\_actor</code></p></td><td><p>Actor 反向传播 + 优化器步进耗时</p></td></tr><tr><td><p><code>timing/s/param\_sync</code></p></td><td><p>Rollouter 从 Trainer 同步参数的耗时</p></td></tr><tr><td><p><code>timing/s/rollout/agent\_loop\_latency/avg</code></p></td><td><p>单次 Rollout 总耗时</p></td></tr><tr><td><p><code>timing/s/rollout/model\_latency/avg</code></p></td><td><p>单条轨迹的 LLM 推理累计耗时</p></td></tr><tr><td><p><code>timing/s/rollout/reward\_latency/avg</code></p></td><td><p>Reward 调用耗时</p></td></tr><tr><td><p><code>timing/ms/gen\_per\_token</code></p></td><td><p>生成阶段每 Token 耗时</p></td></tr><tr><td><p><code>timing/ms/update\_actor\_per\_token</code></p></td><td><p>Actor 更新阶段每 Token 耗时</p></td></tr></tbody></table>
</Accordion>

## 判断训练效果 <span id="opd-judge" />

训练过程中上报的指标已足以判断一轮在线策略蒸馏是否在正常收敛；训练完成后的业务效果评测由用户按自身口径进行。

### 训练中判读顺序 <span id="opd-judge-in-training" />

在「指标」页按以下顺序判读：

<strong>① 教师模型信号完整性</strong> — `trace/distillation/valid_token_rate` 接近 1.0。明显小于 1 时教师模型打分存在缺失，其余指标解释力下降，应先排查（排查决策树 P2）。

<strong>② 蒸馏收敛</strong> — `actor/distillation/loss` 随训练下降或趋稳，表示学生模型正在向教师模型靠拢。

<strong>③ 验证集表现</strong> — `validation/data/reward/mean@1` 相对训练前基线上升。分维度看 `trace/reward_metrics/<reward_name>/<metric>`，关键业务维度（尤其安全类）不回退。

验证频率由 `eval_steps` 控制。若 step 0 存在验证点，该点对应训练前的模型状态，可直接作为基线。

纯蒸馏（不注册 reward 组件，functions=None）不产生验证指标与分维度指标。依据蒸馏损失与教师模型健康指标判断训练是否正常，业务效果通过训练后评测判断。

### 训练后业务评测 <span id="opd-judge-post" />

<strong>保留训练前基线</strong> — 蒸馏收益是相对量，缺基线无法判断提升幅度。训练产出的是新的模型 ID，原模型仍可调用，但基线评测结果需自行留存。

<strong>保持同口径</strong> — 训练前后评测应使用相同评测集与解码参数（`temperature`、`max_tokens`、是否开启思考等）。口径不一致时差值不可比——例如响应被长度上限截断，会表现为能力下降的假象。

示例目录下附带评测与对比脚本，仅作为实现参考，可按需改造。

### 验收判据 <span id="opd-judge-criteria" />

四项同时满足：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "25%" }} /><col style={{ width: "75%" }} /></colgroup><thead><tr><th><p><strong>判据</strong></p></th><th><p><strong>要求</strong></p></th></tr></thead><tbody><tr><td><p>主判据</p></td><td><p>核心评测指标提升且非随机波动</p></td></tr><tr><td><p>安全判据</p></td><td><p>安全类维度不高于训练前</p></td></tr><tr><td><p>格式判据</p></td><td><p>格式类维度不低于训练前</p></td></tr><tr><td><p>分类别判据</p></td><td><p>按业务类别拆分后无类别显著回退</p></td></tr></tbody></table>

<Accordion title="实测数值参考（仅参考，不作为通用判据）">
  示例三方对比（200 条验证集，temperature=0，8 并发）：

  <table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "40%" }} /><col style={{ width: "20%" }} /><col style={{ width: "20%" }} /><col style={{ width: "20%" }} /></colgroup><thead><tr><th><p><strong>指标</strong></p></th><th><p><strong>Base</strong></p></th><th><p><strong>Teacher</strong></p></th><th><p><strong>OPD</strong></p></th></tr></thead><tbody><tr><td><p>核心成功率</p></td><td><p>0.700</p></td><td><p>0.955</p></td><td><p>0.990</p></td></tr><tr><td><p>安全违规率</p></td><td><p>0.235</p></td><td><p>0</p></td><td><p>0</p></td></tr></tbody></table>

  学生模型超越教师模型的原因：启用 reward 时，reward 依据标准答案打分，趋向标准答案而非教师模型；纯蒸馏才以教师模型为性能上限。
</Accordion>

## 在线策略蒸馏排查决策树 <span id="opd-troubleshooting" />

指标页看到异常曲线 → 对照本节定位问题 → 按建议处置。

### P1 蒸馏损失不降 <span id="opd-ts-p1" />

- <strong>主信号</strong>：`actor/distillation/loss` 随训练不下降或反而上升
- <strong>根因</strong>：教师模型信号缺失 / 学习率不适（过小不收敛或过大破坏输出格式） / 数据质量差
- <strong>处置</strong>：先查 `valid_token_rate` 是否接近 1.0 → 信号正常时 loss 不降常因学习率过小，升高 `learning_rate`（×1.5\~2，见训练配置调参决策表）→ 抽查训练数据 `rollout_extra` 中的标准答案

### P2 教师模型打分缺失 <span id="opd-ts-p2" />

- <strong>主信号</strong>：`trace/distillation/valid_token_rate` 明显小于 1.0
- <strong>根因</strong>：教师模型侧打分缺失（教师模型由服务端托管，对用户不透明，无法直接排查）
- <strong>处置</strong>：查 `missing_token_rate` 与 `empty_response_rate` 是否同步升高 → 若教师模型超时或空响应比例高，提交工单

### P3 验证集不升 <span id="opd-ts-p3" />

- <strong>主信号</strong>：`validation/data/reward/mean@1` 出现下行拐点而训练 reward 仍升
- <strong>通用处置</strong>：早停（在产出页签选取较早的 Checkpoint）+ 保持训练前后评测口径一致 + 分维度查看 `reward_metrics` 回退情况
- <strong>额外根因</strong>：过拟合教师模型风格（纯蒸馏时）
- <strong>额外处置</strong>：抽查 `rollout_extra` 中标准答案质量

<span id="opd-ts-four-source" />

<strong>四源协同排查：</strong>四类观测源各回答不同问题：<strong>日志</strong>→"何时失败、报错堆栈"、<strong>指标</strong>→"是不是 / 有多严重"、<strong>轨迹</strong>→"为什么 / 哪条样本"、<strong>轨迹页·Tracing 子页签</strong>→"哪段慢 / 哪个外部依赖异常"。按问题类型选观测源：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "25%" }} /><col style={{ width: "25%" }} /><col style={{ width: "25%" }} /><col style={{ width: "25%" }} /></colgroup><thead><tr><th><p><strong>问题类型</strong></p></th><th><p><strong>主用</strong></p></th><th><p><strong>次用</strong></p></th><th><p><strong>不必看</strong></p></th></tr></thead><tbody><tr><td><p>训练发散</p></td><td><p>指标</p></td><td><p>轨迹</p></td><td><p>日志</p></td></tr><tr><td><p>任务 FAILED</p></td><td><p>日志</p></td><td><p>指标</p></td><td><p>轨迹</p></td></tr><tr><td><p>Reward 不涨</p></td><td><p>指标 → 轨迹</p></td><td><p>reward\_metrics</p></td><td><p>日志</p></td></tr><tr><td><p>工具调用错</p></td><td><p>Tracing 子页签</p></td><td><p>日志</p></td><td><p>—</p></td></tr><tr><td><p>训练变慢</p></td><td><p>指标 timing</p></td><td><p>Tracing 子页签</p></td><td><p>—</p></td></tr><tr><td><p>输出质量差</p></td><td><p>轨迹</p></td><td><p>reward\_metrics</p></td><td><p>—</p></td></tr></tbody></table>

<strong>OPD 特有问题类型选源：</strong>

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "25%" }} /><col style={{ width: "25%" }} /><col style={{ width: "25%" }} /><col style={{ width: "25%" }} /></colgroup><thead><tr><th><p><strong>问题类型</strong></p></th><th><p><strong>主用</strong></p></th><th><p><strong>次用</strong></p></th><th><p><strong>不必看</strong></p></th></tr></thead><tbody><tr><td><p>蒸馏损失不降</p></td><td><p>指标</p></td><td><p>轨迹</p></td><td><p>日志</p></td></tr><tr><td><p>教师模型打分缺失</p></td><td><p>指标</p></td><td><p>日志</p></td><td><p>轨迹</p></td></tr><tr><td><p>验证集不升</p></td><td><p>指标 → 轨迹</p></td><td><p>reward\_metrics</p></td><td><p>日志</p></td></tr></tbody></table>

## 后续步骤 <span id="opd-next-steps" />

- [在线策略蒸馏训练概述](opd-training-overview.md) → 原理与端到端流程
- [Model OPD 开发](opd-model-development-guide.md) / [Agentic OPD 开发](opd-agentic-development-guide.md) → 函数开发完整细节
- [在线策略蒸馏训练配置](opd-training-config.md) → 提交参数与超参全集

## 常见问题 <span id="opd-faq" />

### 教师模型侧异常 <span id="opd-faq-teacher" />

教师模型由服务端部署，对用户不透明。以下现象需提交工单：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "60%" }} /><col style={{ width: "40%" }} /></colgroup><thead><tr><th><p><strong>现象</strong></p></th><th><p><strong>处理</strong></p></th></tr></thead><tbody><tr><td><p><code>trace/distillation/valid\_token\_rate</code> \< 1，教师模型打分缺失</p></td><td><p>提工单</p></td></tr><tr><td><p>教师模型超时或打分缺失</p></td><td><p>提工单</p></td></tr><tr><td><p>同族师生仍报 tokenizer（分词器）不匹配</p></td><td><p>提工单</p></td></tr></tbody></table>

### 自查项 <span id="opd-faq-self" />

通用自查项：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "60%" }} /><col style={{ width: "40%" }} /></colgroup><thead><tr><th><p><strong>现象</strong></p></th><th><p><strong>排查方向</strong></p></th></tr></thead><tbody><tr><td><p>Tracing 控制台空白</p></td><td><p>①确认 env 未设 <code>ENABLE\_TRAJECTORY="false"</code> ②确认已授权 ARMS ③检查 requirements.txt 含 OTel 依赖 ④检查 process() 加了 @observe\_processor</p></td></tr><tr><td><p>Reward 偶发 FAILED</p></td><td><p>未处理空 messages / 编码错误 → 加 try/except 返回 <code>TaskStatus.FAILED + error</code></p></td></tr><tr><td><p>格式类维度低</p></td><td><p>检查学生模型输出格式是否符合 Reward 函数要求</p></td></tr><tr><td><p>如何区分不同 Reward 函数</p></td><td><p>通过 <code>RewardFunctionComponent(name="reward-1")</code> 设唯一名称，指标路径自动区分</p></td></tr></tbody></table>

OPD 特有自查项：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "60%" }} /><col style={{ width: "40%" }} /></colgroup><thead><tr><th><p><strong>现象</strong></p></th><th><p><strong>排查方向</strong></p></th></tr></thead><tbody><tr><td><p>400 ... teacher\_model must be a valid model ID</p></td><td><p>改用百炼模型 ID</p></td></tr><tr><td><p>提交时报错提示 SDK 不具备在线策略蒸馏能力</p></td><td><p>换用示例指定的 wheel</p></td></tr><tr><td><p>纯蒸馏无验证指标</p></td><td><p>正常行为（不注册 reward 组件，functions=None），靠独立评测判断</p></td></tr><tr><td><p>format\_valid 异常低</p></td><td><p>检查学生模型输出是否被 thinking/工具调用文本污染，并确认 Reward 格式判定正则与训练前后解码参数保持一致</p></td></tr></tbody></table>

### 故障速查与 FAILED 排查 <span id="opd-faq-quick" />

通用故障速查：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "40%" }} /><col style={{ width: "30%" }} /><col style={{ width: "30%" }} /></colgroup><thead><tr><th><p><strong>现象</strong></p></th><th><p><strong>首查</strong></p></th><th><p><strong>次查</strong></p></th></tr></thead><tbody><tr><td><p>FAILED·函数注册失败</p></td><td><p>classpath 错 / 依赖缺</p></td><td><p>查 requirements.txt</p></td></tr><tr><td><p>FAILED·Rollout 超时</p></td><td><p>单条 timeout</p></td><td><p>timeout ↑ / Tracing 看哪段慢</p></td></tr><tr><td><p>Tracing 看不到</p></td><td><p>控制台空白</p></td><td><p>检查 ARMS 授权 / requirements.txt</p></td></tr></tbody></table>

OPD 特有故障速查：

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "40%" }} /><col style={{ width: "30%" }} /><col style={{ width: "30%" }} /></colgroup><thead><tr><th><p><strong>现象</strong></p></th><th><p><strong>首查</strong></p></th><th><p><strong>次查</strong></p></th></tr></thead><tbody><tr><td><p>valid\_token\_rate 低</p></td><td><p>教师模型侧异常</p></td><td><p>提工单</p></td></tr><tr><td><p>蒸馏损失不降</p></td><td><p>教师模型信号完整性</p></td><td><p>学习率 / 数据质量</p></td></tr><tr><td><p>tool\_call\_count 为 0 / NO\_TOOL 频繁</p></td><td><p>学生模型是否支持 function calling、传入 tools 参数是否传对</p></td><td><p>查 tool 定义与 api\_key/base\_url</p></td></tr></tbody></table>

<span id="opd-faq-failed" />

<strong>标准排查流程：</strong>

1. <strong>Step 1 任务列表</strong>确认状态与失败时间点
2. <strong>Step 2 日志页签</strong>末尾 100-500 行（SDK `AgenticRL.logs(job_id, lines=100)` / CLI `dashscope rl logs --lines 100`）
3. <strong>Step 3 区分错误层</strong>：

   - <strong>用户函数错</strong>（Rollout/Reward 抛异常）→ `test_functions` 本地复现 → 改代码 → 重新 register/run
   - <strong>框架错</strong>（OOM 内存不足 / 资源不足 / 网络）→ 调 `concurrency` / `capacity`
   - <strong>数据错</strong>（JSONL 解析失败）→ 校验单条格式与 `rollout_extra`

<strong>训练期常见报错模式：</strong>

<table style={{ display: "table", tableLayout: "fixed", width: "100%" }}><colgroup><col style={{ width: "33.333333%" }} /><col style={{ width: "33.333333%" }} /><col style={{ width: "33.333333%" }} /></colgroup><thead><tr><th><p>错误模式</p></th><th><p>主要现象</p></th><th><p>处置建议</p></th></tr></thead><tbody><tr><td><p>资源不足</p></td><td><p>扩容跟不上</p></td><td><p>见 Step 3 框架错（调 <code>concurrency</code> / <code>capacity</code>）</p></td></tr><tr><td><p>数据格式错</p></td><td><p>JSONL 解析失败</p></td><td><p>行级 JSON 校验 / <code>messages</code> 角色 / <code>rollout\_extra</code> 字段</p></td></tr></tbody></table>