Quickstart

Updated at:

Quickly get familiar with Evolution's capabilities

This guide helps you quickly get familiar with Evolution's observability, evaluation, and optimization capabilities.

Prerequisites

  1. Go to Evolution, then click Log In in the top-right corner.
  2. Sign in using your Alibaba Cloud account.
  3. If you don't yet have an Alibaba Cloud account, register one first and complete sign-in.

Observability

Evolution provides developers with full-chain, visualized observability of the execution process, completely recording every processing stage from input to output. It also supports monitoring statistics, rate-limiting statistics, and alert rule configuration.

  1. Report and view call-chain traces and model / tool events.
  2. Observe monitoring metrics such as request count, tokens, models, and tool invocations.
  3. View rate-limiting metrics such as application QPM and throttling counts.
  4. Configure alert rules based on alert templates and review alert history.

For more details, see Observability and Alert Management.

Evaluation

Evolution provides developers with systematic evaluation capabilities to assess application outputs. By analyzing experiment results, you can dig into edge cases and make targeted improvements.

  1. Create and manage evaluation datasets, organizing input samples and expected results.
  2. Create or select a grader, configuring rule- or model-based scoring dimensions.
  3. Create evaluation tasks and view results—run in batch to generate scores, and analyze score distribution and regression comparisons.

For more details, see Evaluation Tasks.

Optimization

Evolution provides developers with a systematic evaluate-and-optimize loop. Based on debugging results, evaluation dataset data, and online trace feedback, you can optimize application prompts intelligently or manually, compare versions side by side, and adopt the best result in a single click—continuously improving application output quality.

  1. Compare multiple prompt versions to debug and optimize your application prompt.
  2. Optimize your application prompt based on debugging results and human feedback.
  3. Optimize your application prompt based on high-quality evaluation datasets and evaluation task results.

For more details, see Optimization.