LLM Hub — CUDA Programming

CUDA Programming

From your first kernel to tiled matmul, streams, atomics and Nsight profiling — ten visual presentations on writing CUDA from scratch.

CUDAKernelsMemoryMatmulNsightStreams

Presentations in This Series

  1. GPU Architecture & the CUDA Execution Model →
    SMs, warps, SIMT execution, thread hierarchy, and how GPU hardware maps to CUDA's programming model.
    live
  2. Your First CUDA Kernel →
    From nvcc setup to vector addition — host/device workflow, memory allocation, kernel launch syntax, error handling.
    live
  3. Thread Hierarchy & Indexing →
    Grids, blocks, threads, warps — mapping problem dimensions to launch configurations with worked index calculations.
    live
  4. Memory Hierarchy →
    Global, shared, constant, texture, and register memory — access patterns, bank conflicts, coalescing rules.
    live
  5. Matrix Multiplication — Naive to Tiled →
    Step-by-step optimisation from a naive O(n³) kernel to shared-memory tiled multiplication with benchmarks.
    live
  6. Synchronisation & Atomics →
    __syncthreads(), warp-level primitives, atomic operations, race conditions, parallel reduction patterns.
    live
  7. Profiling with Nsight →
    Nsight Systems and Nsight Compute — finding bottlenecks, occupancy analysis, memory throughput.
    live
  8. Streams & Async Execution →
    Overlapping compute and data transfer, CUDA streams, events, concurrency patterns, the default stream trap.
    live
  9. Libraries & Ecosystem →
    cuBLAS, cuDNN, Thrust, cuRAND, cuFFT — when to write custom kernels vs use optimised libraries.
    live
  10. Hardware Platforms for CUDA Learning →
    Practical buying guide — DGX Spark, consumer GPUs, cloud instances, Colab — cost, capability, recommendations.
    live