NVIDIA GPU Architectures Series — Presentation 17

Profiling & Debug — Nsight Systems, Nsight Compute, NVTX, CUPTI

Theoretical FLOPS are easy; what your kernel actually achieves is what matters. Walk through every NVIDIA profiling tool — Nsight Systems for the timeline, Nsight Compute for the kernel, NVTX for annotations, CUPTI for programmatic capture, plus dcgmi/nvbandwidth/nvitop — and the workflow that turns "it's slow" into "fix line 42".

Nsight SystemsNsight ComputeNVTX CUPTIdcgminvbandwidth nvitopoccupancy rooflineSASS
NVTX → Nsight Systems → Nsight Compute → SASS → dcgmi → Fleet
00

Topics We'll Cover

A practical walkthrough of NVIDIA's profiling stack. Each tool answers a different question; the skill is knowing which one to reach for first.

01

The Profiling Hierarchy

Profiling is not one tool, it's a stack. A different question is asked at every level — and a different tool answers it. Picking the wrong tool wastes hours: nobody finds a bank conflict in nvidia-smi, and nobody finds a flaky power supply in Nsight Compute.

Profiling hierarchy — one question per level FLEET DCGM · dcgm-exporter · Prometheus · Grafana "why is the cluster slow?" SYSTEM Nsight Systems · NVTX timeline · nsys "why is this rank slow?" KERNEL Nsight Compute · ncu · roofline "why is this kernel slow?" SASS / SOURCE cuobjdump · nvdisasm · PTX "why does this PTX compile to that SASS?" narrower question, deeper view

The four levels in plain English

The interview answer

"Start broad, narrow only when needed." A timeline (Nsight Systems) catches 80% of perf bugs because most of them are systemic: dataloader stalls, missing NVTX context, wrong stream, a sync that didn't need to be there. Don't open Nsight Compute until you know which kernel matters — otherwise you're profiling the wrong thing very precisely.

02

NVTX — Annotate First, Profile Second

NVTX (NVIDIA Tools Extension) is the bridge between your code's mental model and what the profiler sees. It is just two things: range markers (push/pop) and category labels. The cost at runtime is essentially zero when no profiler is attached — ship NVTX in production code without guilt.

Why bother?

Without NVTX, Nsight Systems shows a sea of unnamed CUDA calls: cudaMemcpyAsync, cudaLaunchKernel, cudaStreamSynchronize, repeated thousands of times. You can't tell forward from backward, attention from MLP, or one transformer block from the next.

With NVTX ranges, the timeline shows your model's structure: bands labelled forward, attention.q_proj, backward, optimizer.step. Now patterns and gaps mean something.

What ships with it already

You don't always need to add ranges yourself:

  • PyTorch — torch.cuda.nvtx.range_push wraps autograd, optimisers, and AMP regions emit automatically when torch.profiler is active.
  • TensorFlow — tf.profiler.experimental hooks emit NVTX.
  • cuBLAS / cuDNN / TensorRT / NCCL — emit ranges per call when capture is active.

Custom ranges are still worth it for your training step, dataloader phases, and any host-side preprocessing.

PyTorch training step with NVTX (Python)
import torch
import torch.cuda.nvtx as nvtx

def train_step(model, batch, optimizer, criterion):
    nvtx.range_push("step")

    nvtx.range_push("forward")
    logits = model(batch["input_ids"])
    loss = criterion(logits, batch["labels"])
    nvtx.range_pop()                         # /forward

    nvtx.range_push("backward")
    loss.backward()
    nvtx.range_pop()                         # /backward

    nvtx.range_push("optim")
    optimizer.step()
    optimizer.zero_grad(set_to_none=True)
    nvtx.range_pop()                         # /optim

    nvtx.range_pop()                         # /step
    return loss.item()
C++ / CUDA — same idea, lower level
#include <nvtx3/nvToolsExt.h>

void run_block(int layer_id) {
    char name[64];
    snprintf(name, sizeof(name), "block.%d", layer_id);
    nvtxRangePushA(name);

    nvtxRangePushA("attn");
    launch_attention_kernel(...);
    nvtxRangePop();                          // /attn

    nvtxRangePushA("mlp");
    launch_mlp_kernel(...);
    nvtxRangePop();                          // /mlp

    nvtxRangePop();                          // /block.N
}
Practical advice

Annotate once, at the boundaries that matter to you (forward / backward / dataloader / optimizer / collective). Resist annotating every kernel — the profiler already shows those. Treat NVTX like a coarse log: it should narrate the program, not transcribe it.

03

Nsight Systems — System Timeline

Nsight Systems (nsys) is the system-wide profiler. One capture covers every CPU thread, every CUDA stream, every kernel launch, every NCCL collective, OS scheduling, and your NVTX ranges — correlated on a single time axis. It is the first tool to reach for when something is "slow".

Capture a training run with the trace categories that matter
# Single rank, full coverage:
nsys profile \
    -t cuda,nvtx,osrt,cudnn,cublas,nccl \
    --capture-range=cudaProfilerApi  \
    --capture-range-end=stop          \
    -o train_run                      \
    python train.py

# Then open in Nsight Systems UI:
nsys-ui train_run.nsys-rep

# Or generate a CLI summary without the UI:
nsys stats train_run.nsys-rep | head -60

The trace categories

What you actually look for

Idle gaps on the GPU

Empty rows on the GPU streams = the host is the bottleneck. Common causes: synchronous Python preprocessing, blocking .item() / .cpu() calls, dataloader worker starvation, torch.save mid-training.

Long bars = compute

Wide kernel bars on the GPU stream are good (work happening) but their content matters. Hover for kernel name and duration. If one kernel dominates, that's your Nsight Compute target.

Sync waits

Tall thin bars on the API row marked cudaStreamSynchronize / cudaDeviceSynchronize indicate stream serialisation. Often introduced by debug prints, metric scraping, or ill-placed .item().

NCCL on the wrong stream

If NCCL collectives appear on the same stream as compute, they serialise. The fix is a dedicated comm stream so all-reduce overlaps with the next forward.

Capture overhead

Nsight Systems is sampling-based and cheap — expect < 5% overhead, often closer to 1%. Safe to run on full-size workloads. Nsight Compute, in contrast, serialises every kernel for measurement and slows the program by 10×+; never confuse the two.

04

Reading a Timeline

What does a "good" timeline look like, vs a "bad" one? The difference is rarely subtle once you know what to look for.

Two captures of the same training step BAD — GPU busy ~35%, dataloader stuck on CPU NVTX step CPU dataloader (sync io) forward GPU s0 idle (host wait) idle (sync) GPU s1 no comm overlap (collectives on default stream) SUM: GPU util 35% · step 92 ms · bottleneck = host dataloader + sync GOOD — GPU busy 90%+, dataloader prefetched, NCCL overlapped NVTX step CPU submit prefetch GPU s0 forward backward optimizer GPU s1 all-reduce (overlapped) SUM: GPU util 92% · step 41 ms · comm overlapped · no host stall

The features to read

The lie that nvidia-smi tells

nvidia-smi's "GPU utilisation" is the percentage of time any kernel was running — even one warp, one SM. A 70 GB H100 doing 2% of useful work can show 100% utilisation. Only kernel-level analysis (Nsight Compute) tells you whether the silicon is actually loaded. Treat nvidia-smi -l as a "is it on?" check, not a perf metric.

05

Nsight Compute — Kernel Deep Dive

Nsight Compute (ncu) is the kernel-level microscope. Where Nsight Systems gives you the timeline, ncu gives you the inside of one bar on that timeline: which exact PerfMon counters fired, how many warps stalled and on what, what fraction of theoretical bandwidth and FLOPS you actually achieved.

Basic capture — "full" set on a small repro
# Profile every kernel with the full metric set, write report to file:
ncu --set full -o kernel_full python repro.py

# Most kernels you don't care about. Filter by name:
ncu --set full -k "matmul|attention" -o focused python repro.py

# Only the first occurrence of each kernel:
ncu --set full --launch-skip 0 --launch-count 1 -o once python repro.py

# Source-correlated SASS view (huge for finding the line that stalls):
ncu --set full --import-source yes -o with_src python repro.py

# Then open in the UI:
ncu-ui kernel_full.ncu-rep

What ncu reports per kernel

SectionWhat it tells you
GPU Speed of LightAchieved % of peak compute and memory bandwidth. The headline number.
Compute Workload AnalysisIssued vs executed instructions; pipe utilisation (FMA, ALU, Tensor, FP64, etc.).
Memory Workload AnalysisL1/L2 hit rates, bytes from HBM, sector loads, shared-memory bank conflicts.
Scheduler StatisticsWarps per scheduler, eligible warps, issue slots used — the occupancy story.
Warp State (stall reasons)Per-warp stall breakdown: memory dep, exec dep, IMC throttle, MIO throttle, sync barrier, etc.
Source / SASSPer-line metric overlay on PTX/SASS. Find which line dominates stalls.
RooflinePlot kernel on the FLOP/byte vs FLOPS chart. Are you memory- or compute-bound?
Why ncu is slow

To gather the full metric set, Nsight Compute serialises kernels and replays each one many times to harvest different counter groups. A kernel that runs in 0.5 ms can take 1–2 seconds inside ncu. Never run it on full-scale training — build a small repro: a single forward pass with batch=1, one transformer block, ten tokens. Same kernel pattern, profileable in seconds.

06

Roofline Analysis

The roofline plot is the single best mental model for "is this kernel slow because of memory or because of compute?". It pins the kernel on a chart with two ceilings: the memory bandwidth ceiling (left, sloped) and the compute ceiling (right, flat). Whichever is closer wins; you cannot go above either.

Roofline — arithmetic intensity vs achieved throughput peak (TC) 10x 1x low throughput (FLOPS, log) 0.1 1 10 100 1k arithmetic intensity (FLOP / byte, log) ridge: peak FLOPS / peak BW memory-bound compute-bound softmax (AI=2, far below) — memory-bound small matmul (AI=8, near ridge) large matmul FP8 (AI=120, on TC roof)

How to read it

Practical use — what the dot tells you

Matmul that's 50% memory-bound at AI=120

You're compute-bound territory by AI but achieving the slope-roof ceiling = your tile size is too small or you're hitting cache misses. Increase tile, use TMA, fix shared-memory layout.

Softmax that's 90% compute-bound at AI=2

Impossible: softmax is fundamentally memory-bound. If ncu says compute is the limit, the kernel is doing redundant work — bad reduction pattern, missing online-softmax fusion, recomputing exponentials.

Element-wise op far below the slope

Bandwidth left on the table. Common cause: launching at small grid size, kernel-launch overhead dominates. Fuse with neighbours, increase work per launch.

Large matmul sitting on the TC roof

The good case — you're hitting peak tensor-core throughput. Move on; spend optimisation budget on the next slowest kernel.

07

Common Stalls and Fixes

The Warp State section of ncu reports the average reason a warp was not eligible to issue. Each reason maps to a class of fix. Knowing the table by heart is what separates "the kernel is slow" from "the kernel needs cp.async with a deeper pipeline".

Stall reasonWhat it meansTypical fix
Memory dependency (LG/LD) Warp waiting on a global / shared load. Long-latency HBM read in flight. Increase tile size; use TMA (Hopper+) or cp.async (Ampere+) for async copies; deeper software pipeline.
Execution dependency Pipeline backed up — result of one instruction needed by the next, no other warp ready. Reduce registers per thread to raise occupancy; simplify ILP; let more warps fly in parallel.
IMC miss / throttle Immediate-constant cache miss — warp waiting on a load from constant memory (__constant__, kernel parameters, immediate values). Reduce constant-memory footprint hit per warp; avoid large jumps that defeat the constant cache; promote frequently-read values to registers.
Math pipe throttle Issue pipe saturated — back-to-back instructions targeting the same pipe (e.g. solid FFMA inner loops on the FMA pipe). Spread instruction mix; interleave loads with compute; raise warps-per-SM so other warps can fill issue slots.
MIO throttle Memory-IO unit overloaded — usually shared-memory bank conflicts. Pad shared arrays (+1 trick); swizzle layout; use vectorised loads (float4).
Sync / barrier Warps waiting at __syncthreads(). Reduce barriers; use __syncwarp when block-wide isn't needed; replace with async copies.
Tex throttle Texture / read-only cache pipeline saturated. Reduce __ldg reuse; rebalance between L1 and read-only path.
Branch divergence Threads in a warp take different paths — serial execution. Restructure conditionals; sort or bucket inputs; use predication.
PCIe transfer (host) Visible in Nsight Systems, not ncu — slow H2D copies dominate the timeline. Use pinned memory + cudaMemcpyAsync; overlap with compute on a separate stream; consider GPUDirect Storage if from disk.
Register spill --ptxas-options=-v shows local-memory bytes > 0; spills hit slow LMEM (cached in L1). Reduce live-range pressure; __launch_bounds__; smaller tiles; refactor to reuse registers.
Order of attack

Read the top-line "Speed of Light" first — if it says 80% of peak, stop optimising. If it's 20%, look at the dominant stall reason; that single number narrows the search to one of the rows above. Don't try to fix everything at once: re-profile after every change, or you'll improve one stall and silently regress another.

08

Distributed Profiling

One rank tells you about that rank. Distributed problems — stragglers, ring imbalance, cross-node latency — only show up when you correlate timelines across ranks. Nsight Systems has first-class support for this.

Per-rank capture (torchrun / SLURM / mpirun)
# torchrun: each rank gets its own .nsys-rep, named by RANK env var:
nsys profile -t cuda,nvtx,nccl -o "out_rank%q{RANK}" \
    torchrun --nproc_per_node=8 --nnodes=2 train.py

# SLURM srun: same idea with SLURM_PROCID:
srun --ntasks=16 nsys profile -t cuda,nvtx,nccl \
    -o "out_%q{SLURM_PROCID}" python train.py

# Diagnostic env vars to capture alongside the trace:
export NCCL_DEBUG=INFO              # protocol, ring, algo, channel choice
export NCCL_DEBUG_SUBSYS=COLL,ENV   # filter what NCCL prints
export NCCL_TOPO_DUMP_FILE=topo.xml # NCCL's view of the fabric — gold for "why slow ring"

Multi-report timeline

Nsight Systems' UI lets you load several .nsys-rep files together: File → Open Multiple. The tool aligns them on a common time axis (using NCCL collective endpoints as anchors). What you can then see at a glance:

NCCL_DEBUG=INFO — what you actually want

Per-rank logs print the chosen algorithm (Tree vs Ring), protocol (Simple / LL / LL128), and channel count. A sudden algorithm change between runs explains otherwise mysterious throughput regressions.

NCCL_TOPO_DUMP_FILE=topo.xml

NCCL's belief about the fabric — PCIe topology, NVLink graph, NIC affinity. If NCCL thinks two GPUs are SYS-connected when in fact they share NVLink, you'll see catastrophic perf. Diff this file against nvidia-smi topo -m.

Capture overhead at scale

1–5% per rank, but report files can balloon (gigabytes for a 30s capture across 64 ranks). Use --capture-range=cudaProfilerApi to bracket exactly the steps you want, not the whole run.

Sampling, not full collection

For very large clusters, profile a subset of ranks — rank 0 + one per node usually captures the structure without 1024× the data. Add --sample=cpu for backtraces on the host side.

09

CUPTI — Programmatic Profiling

CUPTI (CUDA Profiling Tools Interface) is the C API behind every NVIDIA profiler. It surfaces every CUDA event — kernel launches, memcpys, sync ops, stream activity — plus the GPU PerfMon counters used by Nsight Compute. You won't write CUPTI directly often; you almost certainly already use it through torch.profiler and friends.

Two CUPTI APIs

  • Activity API — trace events (kernel started, copy completed, stream sync). What you want for "what happened, in what order, when". Low overhead, push-based.
  • Profiling API (Range / Replay) — PerfMon counters: occupancy, achieved BW, FLOPS, stall reasons. What ncu uses; serialises kernels, much higher overhead.

Tools that use CUPTI

  • torch.profiler — emits Chrome trace JSON
  • TensorFlow profiler — same
  • Nsight Systems — CUDA category
  • Nsight Compute — counter API
  • NVIDIA's own DCGM, Triton, TRT-LLM tools internally
  • In-house metric scrapers / autotuners
Minimal CUPTI Activity subscriber — print kernel name + duration (C)
#include <cupti.h>
#include <stdio.h>

static void CUPTIAPI on_buffer_request(uint8_t **buf, size_t *sz, size_t *maxn) {
    *sz = 8 * 1024 * 1024;
    *buf = (uint8_t*)aligned_alloc(8, *sz);
    *maxn = 0;
}

static void CUPTIAPI on_buffer_complete(CUcontext ctx, uint32_t streamId,
                                          uint8_t *buf, size_t sz, size_t valid) {
    CUpti_Activity *rec = NULL;
    while (cuptiActivityGetNextRecord(buf, valid, &rec) == CUPTI_SUCCESS) {
        if (rec->kind == CUPTI_ACTIVITY_KIND_KERNEL ||
            rec->kind == CUPTI_ACTIVITY_KIND_CONCURRENT_KERNEL) {
            CUpti_ActivityKernel9 *k = (CUpti_ActivityKernel9*)rec;
            printf("%-40s  %.3f ms  grid=(%u,%u,%u) regs=%u\n",
                   k->name,
                   (k->end - k->start) / 1.0e6,
                   k->gridX, k->gridY, k->gridZ,
                   k->registersPerThread);
        }
    }
    free(buf);
}

int main(void) {
    cuptiActivityRegisterCallbacks(on_buffer_request, on_buffer_complete);
    cuptiActivityEnable(CUPTI_ACTIVITY_KIND_CONCURRENT_KERNEL);
    cuptiActivityEnable(CUPTI_ACTIVITY_KIND_MEMCPY);

    /* ... run your CUDA program here ... */

    cuptiActivityFlushAll(1);
    return 0;
}
When to write CUPTI directly

Custom in-house: an autotuner that picks tile size by re-measuring kernel time, a continuous metrics daemon that feeds Prometheus per-kernel, a CI test that asserts no kernel regressed by >5%. Most engineers should reach for torch.profiler first — it gives 90% of the value with one decorator.

10

Fleet & Operational Tools

Nsight is for development. In production you don't want to attach a profiler — you want continuous, low-overhead telemetry that alerts when something drifts. That's where DCGM, nvbandwidth, nvitop, and the stress testers live.

ToolWhat it doesWhen you use it
nvidia-smi Lightweight probe: power, temp, mem, processes, ECC counts. Reads NVML. Quick "is it on?" check; ad-hoc one-shot during incident.
nvitop TUI dashboard, multi-GPU + multi-process, sortable, kill-process bindings. Day-to-day "what's running where" on a shared workstation. Strictly better than watch -n1 nvidia-smi.
dcgmi (DCGM) NVIDIA's datacenter GPU manager. CLI for health checks, stress tests, group ops, MIG ops. Pre-flight on a node: dcgmi diag -r 3 runs a 5-min battery of tests.
dcgm-exporter Prometheus exporter for DCGM. Per-GPU metrics with labels. Always-on production observability. Pair with Grafana dashboards.
nvbandwidth Measures real H2D, D2H, D2D, P2P bandwidth across all GPU pairs. NVIDIA-supported successor to bandwidthTest. "Why is GPU 2 slower at all-reduce?" — nvbandwidth exposes the bad lane / cable / port.
nv-gpu-burn / gpu-fryer Stress testers — pure compute loops with optional matrix-error checking. Burn-in before production; reproducing thermal throttling.
cuobjdump / nvdisasm Disassemble a cubin: list kernels, dump SASS, dump PTX, show resource usage per kernel. Final ground-truth: "what did the compiler actually emit?"

What DCGM surfaces that nvidia-smi doesn't

Production observability stack
# On every GPU node: dcgm-exporter as a sidecar / DaemonSet
docker run -d --name dcgm-exporter --gpus all --cap-add SYS_ADMIN \
    -p 9400:9400 nvcr.io/nvidia/k8s/dcgm-exporter:latest

# Prometheus scrapes :9400/metrics every 15s
# Grafana dashboard 12239 (DCGM Exporter) is a sane default

# Pre-flight: full diagnostic before reintroducing a node
sudo dcgmi diag -r 3          # level 3 = ~5 min, full battery
nvbandwidth -t host_to_device_memcpy_ce
nvbandwidth -t device_to_device_memcpy_read_ce
Alert on these, not utilisation

The metrics that actually predict outages: rising ECC SBE rate (silicon ageing), any DBE (replace the card), XID 13/31/43/45/63 (page faults / MMU failures), throttling reason = HW Slowdown (PSU sag), NVLink replay rate > 0. Don't alert on GPU utilisation — remember, nvidia-smi lies.

11

Interactive: Pick Your Tool

Pick the question you actually have. The planner picks the tool, the exact command, what to look for, and the next tool to reach for if the first one doesn't surface the answer.

Capture overhead
—
Restart needed
—
Root needed
—
Multi-process safe
—
The whole deck in one sentence

Start at the level of the question (fleet, rank, kernel), reach for the right tool, capture once with NVTX already in the code, read the highest-priority signal first (top-line throughput → dominant stall reason → source line). Profilers don't fix bugs; they tell you where the bug is. Once you know that, the fix is usually three lines.