Presentations in This Series
- GPU Architecture & the CUDA Execution Model →SMs, warps, SIMT execution, thread hierarchy, and how GPU hardware maps to CUDA's programming model.
- Your First CUDA Kernel →From nvcc setup to vector addition — host/device workflow, memory allocation, kernel launch syntax, error handling.
- Thread Hierarchy & Indexing →Grids, blocks, threads, warps — mapping problem dimensions to launch configurations with worked index calculations.
- Memory Hierarchy →Global, shared, constant, texture, and register memory — access patterns, bank conflicts, coalescing rules.
- Matrix Multiplication — Naive to Tiled →Step-by-step optimisation from a naive O(n³) kernel to shared-memory tiled multiplication with benchmarks.
- Synchronisation & Atomics →__syncthreads(), warp-level primitives, atomic operations, race conditions, parallel reduction patterns.
- Profiling with Nsight →Nsight Systems and Nsight Compute — finding bottlenecks, occupancy analysis, memory throughput.
- Streams & Async Execution →Overlapping compute and data transfer, CUDA streams, events, concurrency patterns, the default stream trap.
- Libraries & Ecosystem →cuBLAS, cuDNN, Thrust, cuRAND, cuFFT — when to write custom kernels vs use optimised libraries.
- Hardware Platforms for CUDA Learning →Practical buying guide — DGX Spark, consumer GPUs, cloud instances, Colab — cost, capability, recommendations.