CUDA Programming Series — Tutorial 01

GPU Architecture & the CUDA Execution Model

A visual deep dive into streaming multiprocessors, warps, SIMT execution, the thread hierarchy, and how GPU hardware maps to CUDA's programming model.

CUDA GPU Architecture SIMT Warps Streaming Multiprocessors Thread Hierarchy
CPU vs GPU → Inside the GPU → SIMT Model → Thread Hierarchy → HW ↔ SW Mapping → GPU Generations → Terminology
00

Topics We'll Cover

This tutorial introduces the hardware architecture that makes CUDA programming possible. No prior GPU knowledge required — we build from first principles.

Prerequisites

None — this is the first tutorial in the series. Basic programming knowledge is helpful but not required.

01

CPU vs GPU — Design Philosophy

CPUs and GPUs solve fundamentally different problems. Understanding why they're designed differently is the key to understanding when and how to use CUDA.

CPU — Latency Optimised

  • Few powerful cores (4–64 typical)
  • Large caches (L1/L2/L3)
  • Complex control logic & branch prediction
  • Optimised for serial tasks
  • Low latency per operation

GPU — Throughput Optimised

  • Thousands of simple cores (128–16,384)
  • Small caches per core
  • Minimal control logic
  • Optimised for parallel tasks
  • High throughput across all cores

Die Layout Comparison

Most of a CPU die is dedicated to control logic and cache. Most of a GPU die is dedicated to arithmetic units.

CPU Die

Core 0
Core 1
Core 2
Core 3
L3 Cache
Control Logic & Branch Prediction
Memory Controller

GPU Die

SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
SM
L2 Cache
Memory Controllers

The Analogy

Think of a CPU as a sports car — incredibly fast at getting one thing from A to B. A GPU is a bus fleet — slower per vehicle, but moves thousands of passengers simultaneously. When your problem involves moving thousands of passengers (pixels, matrix elements, neurons), the bus fleet wins.

02

Inside an NVIDIA GPU

An NVIDIA GPU is organised in a hierarchy: the full chip contains Graphics Processing Clusters (GPCs), each containing multiple Streaming Multiprocessors (SMs), each containing many CUDA cores.

GPU → GPC → SM → CUDA Cores

GPU Chip
→
GPC × 4–8
→
SM × 16–128
→
CUDA Cores × 64–128 per SM

Streaming Multiprocessor (SM) — The Building Block

The SM is the fundamental compute unit. Each SM contains:

CUDA Cores

General-purpose arithmetic units for integer and floating-point operations. 64–128 per SM in modern GPUs.

Tensor Cores

Specialised matrix multiply-accumulate units. Accelerate deep learning (FP16, BF16, INT8, FP4). Introduced in Volta.

RT Cores

Ray tracing acceleration units. Handle BVH traversal and ray-triangle intersection. Introduced in Turing.

Inside a Single SM

Streaming Multiprocessor (SM)
Warp Scheduler 0
Warp Scheduler 1
Warp Scheduler 2
Warp Scheduler 3
FP
FP
FP
FP
FP
FP
FP
FP
INT
INT
INT
INT
INT
INT
INT
INT
FP
FP
FP
FP
FP
FP
FP
FP
INT
INT
INT
INT
INT
INT
INT
INT
Tensor Core
Tensor Core
Tensor Core
Tensor Core
Register File (64K × 32-bit)
Shared Memory / L1 (128 KB)
Key Insight

The SM is the unit of independent execution. When you write a CUDA kernel, you're writing code that runs on SMs. Understanding the SM is the foundation of writing efficient CUDA programs.

03

The SIMT Execution Model

SIMT stands for Single Instruction, Multiple Threads. It's NVIDIA's execution model, closely related to SIMD (Single Instruction, Multiple Data) but with a crucial difference: each thread has its own program counter and register state.

Warps — The Fundamental Execution Unit

The GPU doesn't execute threads individually. Instead, it groups 32 consecutive threads into a warp. All 32 threads in a warp execute the same instruction at the same time.

One Warp = 32 Threads
T0
T1
T2
T3
T4
T5
T6
T7
T8
T9
T10
T11
T12
T13
T14
T15
T16
T17
T18
T19
T20
T21
T22
T23
T24
T25
T26
T27
T28
T29
T30
T31
All 32 threads execute the same instruction simultaneously

Warp Divergence — The Performance Killer

When threads in a warp hit a conditional branch (like an if/else), some threads take one path and others take the other. The warp must execute both paths serially, disabling threads not on the active path. This is called warp divergence.

divergence_example.cu
if (threadIdx.x < 16) {
    // Path A — threads 0-15 execute, threads 16-31 are IDLE
    doWorkA();
} else {
    // Path B — threads 16-31 execute, threads 0-15 are IDLE
    doWorkB();
}
Phase 1: Executing Path A
T0
T1
T2
T3
T4
T5
T6
T7
T8
T9
T10
T11
T12
T13
T14
T15
T16
T17
T18
T19
T20
T21
T22
T23
T24
T25
T26
T27
T28
T29
T30
T31
Phase 2: Executing Path B
T0
T1
T2
T3
T4
T5
T6
T7
T8
T9
T10
T11
T12
T13
T14
T15
T16
T17
T18
T19
T20
T21
T22
T23
T24
T25
T26
T27
T28
T29
T30
T31
Result: 2× the execution time — only 50% utilisation per phase
Rule of Thumb

Minimise divergence within a warp. If branching is unavoidable, try to align branch boundaries with warp boundaries (multiples of 32 threads).

04

Thread Hierarchy Overview

CUDA organises parallel work in a hierarchy of threads, warps, blocks, and grids. This hierarchy maps directly to the hardware.

Thread → Warp → Block → Grid

Grid
Grid (all blocks for one kernel launch)
↓ contains
Blocks
Block 0 Block 1 Block 2 Block 3 … Block N
↓ contains
Warps
Warp 0 Warp 1 Warp 2 Warp 3 …
↓ contains
Threads
T0 T1 T2 … T31

Concrete Example

Suppose you launch a kernel to process 1,000,000 elements with a block size of 256 threads:

Level Count Calculation
Threads per block 256 Chosen by the programmer
Warps per block 8 256 / 32 = 8
Blocks in grid 3,907 ceil(1,000,000 / 256) = 3,907
Total threads 1,000,192 3,907 × 256 (192 threads are "extra")
Total warps 31,256 3,907 × 8
Why 192 extra threads?

1,000,000 is not evenly divisible by 256, so the last block has threads that go past the data boundary. The kernel must include a bounds check: if (idx < N) to prevent these threads from accessing invalid memory.

Dimensionality

Grids and blocks can be 1D, 2D, or 3D. Use the dim3 type in CUDA to specify multi-dimensional configurations — useful for matrices (2D) or volumes (3D).

launch_configurations.cu
// 1D: processing a flat array
dim3 block(256);
dim3 grid((N + 255) / 256);
kernel<<>>(...);

// 2D: processing a matrix (rows × cols)
dim3 block(16, 16);           // 16×16 = 256 threads per block
dim3 grid((cols+15)/16, (rows+15)/16);
kernel<<>>(...);

// 3D: processing a volume
dim3 block(8, 8, 4);          // 8×8×4 = 256 threads per block
dim3 grid((X+7)/8, (Y+7)/8, (Z+3)/4);
kernel<<>>(...);
05

How Hardware Maps to Software

The CUDA programming model maps directly to hardware. Understanding this mapping is the key to writing performant code.

The Mapping

Software (CUDA) Hardware Notes
Thread CUDA Core (lane) Smallest unit of execution
Warp (32 threads) Warp Scheduler unit Executes in lockstep on one SM
Block Streaming Multiprocessor (SM) All threads in a block run on the same SM
Grid Entire GPU Blocks distributed across all SMs

Blocks Are Scheduled onto SMs

Block 0
→
SM 0
Block 1
→
SM 1
Block 2
→
SM 0

Multiple blocks can be assigned to the same SM (if resources permit)

Occupancy

Occupancy is the ratio of active warps to the maximum number of warps an SM can support. Higher occupancy generally means better latency hiding — when one warp stalls on a memory access, the warp scheduler can switch to another active warp at zero cost.

What limits occupancy:

Warp Scheduling — Zero-Cost Context Switching

Unlike CPU thread context switches (which are expensive), the GPU keeps the state of all active warps in registers simultaneously. When a warp stalls on a memory access, the warp scheduler can instantly switch to another ready warp — no save/restore overhead.

Warp 0: Executing instruction
↓ stalls on memory
Warp 1: Immediately starts executing
↓ stalls on memory
Warp 2: Immediately starts executing
↓ Warp 0's data arrives
Warp 0: Resumes execution
Key Insight

The GPU hides memory latency by keeping many warps in flight. This is why occupancy matters and why GPUs need thousands of threads to reach peak performance.

06

Compute Capability & GPU Generations

Each NVIDIA GPU has a compute capability version (e.g., 8.6) that determines which CUDA features it supports. The major number indicates the architecture generation; the minor number indicates incremental improvements.

Architecture Compute Cap. Year Key Features Introduced
Tesla 1.x 2006 First CUDA architecture, unified shaders
Fermi 2.x 2010 L1/L2 caches, ECC memory, 64-bit addressing
Kepler 3.x 2012 Dynamic parallelism, Hyper-Q, shuffle instructions
Maxwell 5.x 2014 Improved energy efficiency, shared memory improvements
Pascal 6.x 2016 Unified memory, NVLink, FP16 support, HBM2
Volta 7.0 2017 Tensor Cores (1st gen), independent thread scheduling
Turing 7.5 2018 RT Cores, INT8/INT4 Tensor ops, GDDR6
Ampere 8.x 2020 3rd-gen Tensor Cores, TF32, sparsity, async copy
Ada Lovelace 8.9 2022 4th-gen Tensor Cores, FP8, Shader Exec. Reorder
Hopper 9.0 2022 Transformer Engine, DPX instructions, TMA
Blackwell 10.x 2024 5th-gen Tensor Cores, FP4, 2nd-gen Transformer Engine

Why Compute Capability Matters

Compilation examples
# Target Volta and newer
nvcc -arch=sm_70 my_kernel.cu -o my_kernel

# Target Ampere (RTX 3090)
nvcc -arch=sm_86 my_kernel.cu -o my_kernel

# Target datacenter Blackwell (B200); RTX 5090 is sm_120, DGX Spark sm_121
nvcc -arch=sm_100 my_kernel.cu -o my_kernel

# Generate code for multiple architectures (fat binary)
nvcc -gencode arch=compute_70,code=sm_70 \
     -gencode arch=compute_86,code=sm_86 \
     -gencode arch=compute_100,code=sm_100 \
     my_kernel.cu -o my_kernel
07

Key Terminology Reference

A quick-reference glossary of the terms introduced in this tutorial.

Term Definition
SM Streaming Multiprocessor — the fundamental compute building block of an NVIDIA GPU. Contains CUDA cores, warp schedulers, register files, and shared memory.
CUDA Core A single arithmetic execution unit within an SM. Handles one floating-point or integer operation per clock cycle.
Warp A group of 32 threads that execute in lockstep under the SIMT model. The smallest unit of scheduling on the GPU.
Lane A single thread's position within a warp (lane 0 through lane 31).
Thread Block A programmer-defined group of threads (up to 1,024) that execute on the same SM and can share memory and synchronise.
Grid The collection of all thread blocks launched by a single kernel invocation.
Kernel A function written in CUDA C/C++ that executes on the GPU. Declared with the __global__ qualifier.
Occupancy The ratio of active warps on an SM to the maximum supported. Higher occupancy helps hide memory latency.
Compute Capability A version number (e.g., 8.6) indicating which hardware features and CUDA APIs a GPU supports.
SIMT Single Instruction, Multiple Threads — NVIDIA's execution model where 32 threads (one warp) execute the same instruction simultaneously.
Warp Divergence When threads within a warp take different execution paths at a branch, forcing serial execution of both paths.
GPC Graphics Processing Cluster — a group of SMs within the GPU hierarchy. An organisational unit between the full chip and individual SMs.
Tensor Core Specialised hardware unit for matrix multiply-accumulate operations. Accelerates deep learning training and inference.
08

Summary & Next Steps

What We Covered

Key Takeaways

Think in Warps

The warp (32 threads) is the true unit of execution. Block sizes should be multiples of 32. Minimise warp divergence.

Maximise Occupancy

Keep enough warps active to hide memory latency. Balance register use, shared memory, and block size.

Next Tutorial

Up Next — Tutorial 02

Your First CUDA Kernel — set up the CUDA toolkit, write your first __global__ function, compile it with nvcc, and learn the host ↔ device workflow: allocate → copy → launch → copy → free.