View training job details

Updated at:

After you submit a training job, you can view basic job information, events, resource views, and logs to fully understand the job run status. You can search by job name or ID and quickly switch between running and historical instances.

View job basic information and configuration

  1. Log on to the PAI console, select the region at the top of the page, select the workspace on the right, and then click Enter DLC.

  2. Click the job name to go to the job overview page.

  3. On the Overview page, you can view the basic information, environment information, and resource information of the job.

    The top of the overview page displays the job status (such as Succeeded), elapsed time, compute type (such as General Compute), job type (such as PyTorchJob), and resource group, along with a timeline covering phases from job creation through environment preparation and job running to job success. The basic information includes the job name, job ID, and tags. The resource information includes the resource type, number of workers, and instance type (such as ecs.gn7i-c8g1.2xlarge). The environment information includes the node image URL, dataset mount configuration, and execution command. In addition, the page contains configuration sections such as Fault Tolerance and Diagnostics (auto fault tolerance, health check switch), Network Information (VPC, security group, vSwitch, and more), and Roles and Permissions (instance RAM role, scope of visibility).

  4. Click the job name at the top of the page to expand the job switching list. The list supports fuzzy search by name or ID, helping you quickly switch between running and historical instances.

View job events

Event logs record the progress of job scheduling and resource allocation. You can view job events to locate and troubleshoot issues.

  • Click the Event tab to view the job event logs.

    The left side of the Events tab is an event timeline that displays phase nodes such as Start creation, Environment preparation, Start running, and Job success, along with the corresponding time ranges. The right side is an event log panel that records the details of PyTorchJob lifecycle events, including Job queued, Job dequeued, Pod created, Service created, Scheduling succeeded, Running, and Succeeded.

  • In the Overview area at the bottom of the Instance page, click Actions in the instance Log column to view the specific node event logs on the System Log tab.

    The system logs record the complete lifecycle events of the Pod, including status transitions (ResourcePurchasing → NetworkInitializing → Initializing → ImagePulling → WaitingForRun → Running → Succeeded), image pull records, and node events such as container creation and startup.

View resource views

The resource view provides metrics such as GPU usage, GPU memory usage, CPU usage, memory usage, and network I/O, helping you monitor resource usage in real time and optimize resource allocation.

Click the Monitoring tab to view the job resource view.

For training jobs created by using resource quotas, the following monitoring features are also supported:

  1. Monitoring metrics are available at the Job, Pod, and GPU levels.

    The top of the monitoring page displays a timeline of the job lifecycle, including phases such as Job creation, Queued, Environment preparation, Job running, and Job success, along with the duration of each phase. The lower section of the page lets you switch between the Job, Pod, and GPU tabs to view the corresponding monitoring metrics. The GPU tab displays line charts of GPU compute usage and GPU memory usage.

  2. You can filter by time and metric, and categorize metrics. Click More to customize the metric view: select the required metrics and drag them to adjust the display order to focus on core monitoring metrics.

    The optional GPU monitoring metrics include GPU compute usage, Total GPU memory, GPU memory usage, GPU memory usage amount, GPU memory device interface usage, GPU memory bandwidth usage, GPU SM device usage, GPU device power consumption, and GPU temperature. After you select the required metrics, drag them in the Metric Sort area on the right to adjust the display order, and then click OK to complete the configuration.

  3. DLC jobs also support monitoring alerts for real-time resource utilization tracking. For more information, see Training monitoring and alerts.

View job logs

When a job fails or you need to view the job execution history, you can view the job logs by using the following two methods:

  • In the Overview area at the bottom of the Instance page, click Actions in the instance Log column to view the output logs of a node.

  • Click the Log tab to search for relevant log events by keyword. For more information, see Query aggregated logs by keyword. The following is a brief guide to the query syntax:

    - Simple query: error, matches logs that contain "error"
    - Multi-word query: "Unexpected result", matches both words "Unexpected" and "result"
    - Fuzzy query: error*, matches words that start with "error"; special characters are not supported
    - Phrase query: #"abc$def", matches the exact phrase "abc$def"
    Delimiters: Logs are split by delimiters. Delimiters in keywords are treated as empty characters. If a keyword contains a delimiter, use a phrase query instead. Common delimiters: \n\t\r,;[]{}()&^*#@~=<>/\?:'"
    Note:
    The characters ( and ) in search terms are automatically replaced with spaces. Avoid using these characters in search terms.

    The left side of the log page displays the instance list, and the right side displays the user logs of the corresponding instance. In the example, the training log shows the validation results from Epoch 16 to Epoch 18, and the final output Model saved with accuracy: 98.96% indicates that model training completed successfully.

Troubleshooting empty user logs

User logs are generated from the stdout and stderr of user scripts. If the user logs are empty but the System Log shows normal status, troubleshoot the issue in the following order:

  1. Check the stdout output: Add an echo or Python print statement to the job startup command to output test content. If the test output appears in the user logs, the original job has redirected stdout/stderr to a custom location. Redirect the output to stdout or stderr instead.

  2. Check the security group outbound rules: The outbound rules of the security group associated with the DLC job must allow the internal CIDR block 100.64.0.0/10 of Log Service (SLS) on TCP ports 80 and 443. If the outbound rules deny this CIDR block, user logs cannot be reported to PAI, and the user logs appear empty on the page.

  3. Check the job duration: For jobs with a very short runtime (for example, less than 2 seconds), the user logs may be empty because SLS has not finished collecting logs. When you submit a test job, we recommend that you wait at least 30 seconds before viewing the logs.

View behavior event logs

PAI integrates with ActionTrail. You can view and query DLC behavior event logs from the last 90 days for your Alibaba Cloud account in ActionTrail. For more information, see ActionTrail.

View job restart records

If you enabled Auto fault tolerance or Health check (blacklist and rerun) when you created the job, you can click Restart count to go to the restart records page and view restart details, including the count, time, reason, result, and duration. Perform the following operations:

  • In the restart records list, click Error details to view restart details, including the restart count, restart time, node name, instance name, error code, error message, and error source.

  • Click View aggregated error details to expand a detailed list of all restart records.

Related documents

You can perform management operations based on the job status. For more information, see Manage training jobs.