Nsight performance analysis tool
DSW supports the NVIDIA Nsight performance analysis tools. Use them to run visual analytics on GPU applications, and to locate and optimize performance bottlenecks.
Limits
DSW instance type — The instance must have at least one NVIDIA GPU and must use a Lingjun AI Computing Service instance type.
Non-Hopper architectures — On non-Hopper architecture instance types, pause the AMPerf service before you use Nsight. NVIDIA hardware allows only one process to collect GPU performance metrics at a time.
Procedure
1. Create a DSW instance
Create a DSW instance with a GPU specification that meets the requirementsCreate a DSW instance. The instance must have at least one NVIDIA GPU and must use a Lingjun AI Computing Service instance type.
2. Pause AMPerf collection
If you use a Hopper architecture instance type, such as H20, skip this step. On non-Hopper architecture instance types, pause AMPerf metric collection first. After the pause, you can use tools such as nsys and ncu as usual.
Use the following commands to pause and resume AMPerf collection:
# Pause AMPerf collection
/run/amperf/bin/amperfd profmetric --pause -t 600
# Resume AMPerf collection
/run/amperf/bin/amperfd profmetric --resumeThe -t option specifies the pause duration in seconds.
When you pause AMPerf collection, consider the following:
Monitoring impact — While AMPerf is paused, performance metrics are missing, and the accuracy of the Cloud Monitor dashboard metrics on the instance product page is affected.
Pause duration — An AMPerf pause lasts at most 10 minutes (600 seconds). Collection resumes automatically when the pause expires, or you can resume it manually. After a resume, wait 1 minute before you pause again.
Segment long-running tasks — For long-running tasks, run the analysis in segments to avoid the AMPerf pause duration limit.
Resume promptly — Resume the AMPerf service promptly after the analysis completes, so that your monitoring data stays complete.
3. Install and use Nsight
Download, install, and use Nsight by following the official NVIDIA documentation.
Nsight Compute (
ncu) — A dedicated tool for CUDA kernel-level performance profiling. It supports fine-grained metrics such as instruction-level execution time and memory bandwidth utilization. For details, see Nsight Compute Command-Line Interface Guide.Nsight Systems (
nsys) — A system-level performance analysis suite. It captures the GPU-CPU execution trace and resource usage across the entire call stack. For details, see Nsight Systems Manual.