A visual deep dive into streaming multiprocessors, warps, SIMT execution, the thread hierarchy, and how GPU hardware maps to CUDA's programming model.
This tutorial introduces the hardware architecture that makes CUDA programming possible. No prior GPU knowledge required — we build from first principles.
None — this is the first tutorial in the series. Basic programming knowledge is helpful but not required.
CPUs and GPUs solve fundamentally different problems. Understanding why they're designed differently is the key to understanding when and how to use CUDA.
Most of a CPU die is dedicated to control logic and cache. Most of a GPU die is dedicated to arithmetic units.
Think of a CPU as a sports car — incredibly fast at getting one thing from A to B. A GPU is a bus fleet — slower per vehicle, but moves thousands of passengers simultaneously. When your problem involves moving thousands of passengers (pixels, matrix elements, neurons), the bus fleet wins.
An NVIDIA GPU is organised in a hierarchy: the full chip contains Graphics Processing Clusters (GPCs), each containing multiple Streaming Multiprocessors (SMs), each containing many CUDA cores.
The SM is the fundamental compute unit. Each SM contains:
General-purpose arithmetic units for integer and floating-point operations. 64–128 per SM in modern GPUs.
Specialised matrix multiply-accumulate units. Accelerate deep learning (FP16, BF16, INT8, FP4). Introduced in Volta.
Ray tracing acceleration units. Handle BVH traversal and ray-triangle intersection. Introduced in Turing.
The SM is the unit of independent execution. When you write a CUDA kernel, you're writing code that runs on SMs. Understanding the SM is the foundation of writing efficient CUDA programs.
SIMT stands for Single Instruction, Multiple Threads. It's NVIDIA's execution model, closely related to SIMD (Single Instruction, Multiple Data) but with a crucial difference: each thread has its own program counter and register state.
The GPU doesn't execute threads individually. Instead, it groups 32 consecutive threads into a warp. All 32 threads in a warp execute the same instruction at the same time.
When threads in a warp hit a conditional branch (like an if/else), some threads take one path and others take the other. The warp must execute both paths serially, disabling threads not on the active path. This is called warp divergence.
if (threadIdx.x < 16) {
// Path A — threads 0-15 execute, threads 16-31 are IDLE
doWorkA();
} else {
// Path B — threads 16-31 execute, threads 0-15 are IDLE
doWorkB();
}
Minimise divergence within a warp. If branching is unavoidable, try to align branch boundaries with warp boundaries (multiples of 32 threads).
CUDA organises parallel work in a hierarchy of threads, warps, blocks, and grids. This hierarchy maps directly to the hardware.
Suppose you launch a kernel to process 1,000,000 elements with a block size of 256 threads:
| Level | Count | Calculation |
|---|---|---|
| Threads per block | 256 | Chosen by the programmer |
| Warps per block | 8 | 256 / 32 = 8 |
| Blocks in grid | 3,907 | ceil(1,000,000 / 256) = 3,907 |
| Total threads | 1,000,192 | 3,907 × 256 (192 threads are "extra") |
| Total warps | 31,256 | 3,907 × 8 |
1,000,000 is not evenly divisible by 256, so the last block has threads that go past the data boundary. The kernel must include a bounds check: if (idx < N) to prevent these threads from accessing invalid memory.
Grids and blocks can be 1D, 2D, or 3D. Use the dim3 type in CUDA to specify multi-dimensional configurations — useful for matrices (2D) or volumes (3D).
// 1D: processing a flat array
dim3 block(256);
dim3 grid((N + 255) / 256);
kernel<<>>(...);
// 2D: processing a matrix (rows × cols)
dim3 block(16, 16); // 16×16 = 256 threads per block
dim3 grid((cols+15)/16, (rows+15)/16);
kernel<<>>(...);
// 3D: processing a volume
dim3 block(8, 8, 4); // 8×8×4 = 256 threads per block
dim3 grid((X+7)/8, (Y+7)/8, (Z+3)/4);
kernel<<>>(...);
The CUDA programming model maps directly to hardware. Understanding this mapping is the key to writing performant code.
| Software (CUDA) | Hardware | Notes |
|---|---|---|
| Thread | CUDA Core (lane) | Smallest unit of execution |
| Warp (32 threads) | Warp Scheduler unit | Executes in lockstep on one SM |
| Block | Streaming Multiprocessor (SM) | All threads in a block run on the same SM |
| Grid | Entire GPU | Blocks distributed across all SMs |
Multiple blocks can be assigned to the same SM (if resources permit)
Occupancy is the ratio of active warps to the maximum number of warps an SM can support. Higher occupancy generally means better latency hiding — when one warp stalls on a memory access, the warp scheduler can switch to another active warp at zero cost.
What limits occupancy:
Unlike CPU thread context switches (which are expensive), the GPU keeps the state of all active warps in registers simultaneously. When a warp stalls on a memory access, the warp scheduler can instantly switch to another ready warp — no save/restore overhead.
The GPU hides memory latency by keeping many warps in flight. This is why occupancy matters and why GPUs need thousands of threads to reach peak performance.
Each NVIDIA GPU has a compute capability version (e.g., 8.6) that determines which CUDA features it supports. The major number indicates the architecture generation; the minor number indicates incremental improvements.
| Architecture | Compute Cap. | Year | Key Features Introduced |
|---|---|---|---|
| Tesla | 1.x | 2006 | First CUDA architecture, unified shaders |
| Fermi | 2.x | 2010 | L1/L2 caches, ECC memory, 64-bit addressing |
| Kepler | 3.x | 2012 | Dynamic parallelism, Hyper-Q, shuffle instructions |
| Maxwell | 5.x | 2014 | Improved energy efficiency, shared memory improvements |
| Pascal | 6.x | 2016 | Unified memory, NVLink, FP16 support, HBM2 |
| Volta | 7.0 | 2017 | Tensor Cores (1st gen), independent thread scheduling |
| Turing | 7.5 | 2018 | RT Cores, INT8/INT4 Tensor ops, GDDR6 |
| Ampere | 8.x | 2020 | 3rd-gen Tensor Cores, TF32, sparsity, async copy |
| Ada Lovelace | 8.9 | 2022 | 4th-gen Tensor Cores, FP8, Shader Exec. Reorder |
| Hopper | 9.0 | 2022 | Transformer Engine, DPX instructions, TMA |
| Blackwell | 10.x | 2024 | 5th-gen Tensor Cores, FP4, 2nd-gen Transformer Engine |
nvcc, you target a specific compute capability with the -arch=sm_XX flag.# Target Volta and newer
nvcc -arch=sm_70 my_kernel.cu -o my_kernel
# Target Ampere (RTX 3090)
nvcc -arch=sm_86 my_kernel.cu -o my_kernel
# Target datacenter Blackwell (B200); RTX 5090 is sm_120, DGX Spark sm_121
nvcc -arch=sm_100 my_kernel.cu -o my_kernel
# Generate code for multiple architectures (fat binary)
nvcc -gencode arch=compute_70,code=sm_70 \
-gencode arch=compute_86,code=sm_86 \
-gencode arch=compute_100,code=sm_100 \
my_kernel.cu -o my_kernel
A quick-reference glossary of the terms introduced in this tutorial.
| Term | Definition |
|---|---|
| SM | Streaming Multiprocessor — the fundamental compute building block of an NVIDIA GPU. Contains CUDA cores, warp schedulers, register files, and shared memory. |
| CUDA Core | A single arithmetic execution unit within an SM. Handles one floating-point or integer operation per clock cycle. |
| Warp | A group of 32 threads that execute in lockstep under the SIMT model. The smallest unit of scheduling on the GPU. |
| Lane | A single thread's position within a warp (lane 0 through lane 31). |
| Thread Block | A programmer-defined group of threads (up to 1,024) that execute on the same SM and can share memory and synchronise. |
| Grid | The collection of all thread blocks launched by a single kernel invocation. |
| Kernel | A function written in CUDA C/C++ that executes on the GPU. Declared with the __global__ qualifier. |
| Occupancy | The ratio of active warps on an SM to the maximum supported. Higher occupancy helps hide memory latency. |
| Compute Capability | A version number (e.g., 8.6) indicating which hardware features and CUDA APIs a GPU supports. |
| SIMT | Single Instruction, Multiple Threads — NVIDIA's execution model where 32 threads (one warp) execute the same instruction simultaneously. |
| Warp Divergence | When threads within a warp take different execution paths at a branch, forcing serial execution of both paths. |
| GPC | Graphics Processing Cluster — a group of SMs within the GPU hierarchy. An organisational unit between the full chip and individual SMs. |
| Tensor Core | Specialised hardware unit for matrix multiply-accumulate operations. Accelerates deep learning training and inference. |
The warp (32 threads) is the true unit of execution. Block sizes should be multiples of 32. Minimise warp divergence.
Keep enough warps active to hide memory latency. Balance register use, shared memory, and block size.
Your First CUDA Kernel — set up the CUDA toolkit, write your first __global__ function, compile it with nvcc, and learn the host ↔ device workflow: allocate → copy → launch → copy → free.