With the increasing use of AI and deep learning workloads on HPC systems, understanding application performance has become essential for improving runtime, scalability, and resource utilization. Performance issues may arise from computation, memory usage, I/O, communication, or inefficient use of GPUs and other resources. This course introduces practical performance-engineering methods and tools for analyzing AI and HPC workloads. Participants will learn the difference between monitoring, sampling, profiling, and tracing, and how each method supports a different level of performance analysis. We provide an overview of commonly used tools, including nvidia-smi/nvitop, NVIDIA Nsight Systems and Nsight Compute, Score-P ecosystem, and selected system-monitoring tools. Through hands-on exercises, participants will practice collecting and interpreting performance data. The focus is not only on running tools, but also on understanding which tool is appropriate for which performance questio
Online (BigBlueButton)
This event includes following dates: