Manual evaluation

更新时间:
复制 MD 格式

Manual evaluation is a method for assessing the performance of large model applications. This process involves creating an evaluation set for specific business scenarios, and then manually analyzing and scoring the application's responses to generate an evaluation report.

Demonstration

Manual evaluation involves creating an evaluation set, analyzing and scoring the application's responses, and generating an evaluation report.

image

Step 1: Prepare an evaluation set

Download and complete the evaluation set template. A sample evaluation set file is available for download: Application Evaluation Set - EfmApplicationdata.xlsx

Note

Prompt: An instruction provided to the large model. A prompt can be a question, a text description, or a text description with parameters.

Completion: The response to the prompt. This can be an answer or a text description.

SessionId: The session ID. You can create a custom ID.

Step 2: Upload the evaluation set

  1. Go to the Application Evaluation Evaluation Set page.

  2. Click Create Evaluation Set, enter a custom name for the set, and upload the evaluation set file.

    Note

    Supported file extensions are .xls and .xlsx. The maximum file size is 20 MB. You can upload up to 10 files at a time.

    image

  3. Click Confirm. The uploaded file is displayed on the Evaluation Set page.

  4. Wait until the Import Status is Import Successful. Then, in the Operation column, click Publish to publish the evaluation set.

    You cannot use an evaluation set that is in a draft state. You must publish the set before you can use it.

    image

Step 3: Create an evaluation task

  1. Go to the Manual Evaluation page, click Create Evaluation Task, select a published agent application from the Batch Application Evaluation drop-down list, and then click Next.

    You can select only published agent applications.
    Note

    Batch Application Evaluation: With this method, you select an evaluation set to test the application. This is suitable for end-to-end performance validation before an application is published.

image

  1. Select an uploaded and published evaluation set, and then click Next.

image

  1. Select the evaluation dimensions and click Next.

    If a custom evaluation dimension template is not configured, select a built-in template.

image

  1. Customize the Task Name and review the task details. Click Calculation Details to view the estimated cost.

    After confirming the details, click Start Evaluation.

Note

Estimated Cost: Batch application evaluation runs inference on the evaluation set to obtain model results. If the model is deployed on public resources, token invocation fees may be incurred or your data transfer plan may be consumed. No fees are charged if the model is deployed on dedicated resources. Confirm the cost before you start the evaluation. The cost is calculated as follows:

Evaluation Cost = Tokens generated by evaluation × Unit price per model call

The total number of tokens includes the tokens in the evaluation set and the tokens in the inference results. The final cost is based on actual usage.

image

  1. When the evaluation status becomes Annotating, click the Annotate button in the Operation column to annotate the application's results.

    Annotating involves comparing the application's results with the standard answers in the evaluation set. You then rate the application's results (such as "Poor," "Fair," or "Good") or score them (such as on a scale of 1 to 5). This process helps you identify how the application performs in different scenarios.

    image

    Compare the result of the evaluation set with that of Application A. Provide an overall rating (Poor, Fair, or Good), and click Save and Next.

    image

  1. After you annotate each item in the evaluation set, click Complete and Submit Evaluation to complete the evaluation.

image

View evaluation results

After the evaluation is complete, the Evaluation Status is Completed. In the Operation column, click Result to view the details.

image

image