LLM Hub — NVIDIA GPU Architectures

NVIDIA GPU Architectures

Deep technical tour of NVIDIA GPUs from Pascal to Blackwell — SMs, tensor cores, memory hierarchy, NVLink, packaging, power, plus per-architecture low-level deep dives and the DGX Spark workstation.

PascalVoltaTuringAmpereHopperAdaBlackwellDGX Spark

Presentations in This Series

  1. The NVIDIA GPU Family Tree — Pascal to Blackwell →
    Family timeline, process nodes, dies, transistor counts, memory types and what each generation unlocked. Interactive family explorer.
    live
  2. Inside the SM — How NVIDIA's Streaming Multiprocessor Evolved →
    SM internals across Pascal, Volta, Turing, Ampere, Hopper, Ada and Blackwell — schedulers, register file, tensor cores, TMA, clusters.
    live
  3. Tensor Cores — Five Generations →
    Every generation from Volta's 4×4×4 FP16 to Blackwell's MX-FP4 — formats, MMA shapes, sparsity, Transformer Engine.
    live
  4. Memory Hierarchy →
    Registers, shared, L1/L2, HBM2e/3/3e, GDDR6/6X/7, LPDDR5x unified, Hopper TMA. Decode-speed estimator.
    live
  5. NVLink & NVSwitch →
    Scale-up interconnect from NVLink 1 (Pascal) to NVLink 5 (NVL72), NVLink-C2C, the rack-scale superpod.
    live
  6. Ampere — A100, RTX 30, the LLM Era →
    GA100 + GA10x — 3rd-gen tensor cores (TF32, BF16, 2:4 sparsity), MIG, NVLink 3 with NVSwitch 2.
    live
  7. Hopper — H100, FP8, Transformer Engine →
    GH100, H100/H200/GH200 — 4th-gen tensor cores, native FP8, TMA, thread-block clusters and DSMEM, DPX, NVLink 4.
    live
  8. Ada Lovelace — RTX 40, L40S, Consumer-Class AI →
    AD102 + RTX 40 / L40S / L4 — 4th-gen RT cores, DLSS 3, the FP8 split, no-NVLink consequences.
    live
  9. Blackwell — Dual-Die, FP4, NVL72 →
    B100/B200/GB200 — dual-die NV-HBI, 5th-gen tensor cores with MX-FP4, 2nd-gen Transformer Engine, RAS engine, NVLink 5, NVL72 superpod.
    live
  10. Software Stack & Performance →
    CUDA stack — driver, cuBLAS, cuDNN, CUTLASS, Transformer Engine, NCCL, Triton, TensorRT-LLM. End-to-end LLM tok/s calculator.
    live
  11. DGX, HGX, MGX — Datacenter Reference Platforms →
    Datacenter platforms, OAM, BasePOD and SuperPOD blueprints, NVL72, the OEM ecosystem and DGX Cloud.
    live
  12. GeForce, RTX Pro, Tesla, A/H/L/B — Decoding the Lineup →
    Field guide to every NVIDIA product family — naming logic, EULA boundaries, driver branches, the same-die-different-card patterns.
    live
  13. Networking — InfiniBand, ConnectX, BlueField →
    Cross-node fabric — ConnectX HCAs, Quantum IB and Spectrum-X switches, BlueField DPUs, GPUDirect RDMA / Storage, NCCL, SHARP.
    live
  14. PCIe & GPUDirect →
    PCIe 3 to 6, Resizable BAR, IOMMU, ACS, NUMA pinning, GPUDirect P2P/RDMA/Storage. Topology lint.
    live
  15. Grace — NVIDIA's ARM CPU, GH200, GB200 →
    Grace 72-core Neoverse V2, NVLink-C2C 900 GB/s, GH200, GB200, Extended GPU Memory, DGX Spark.
    live
  16. Jetson — Edge AI & Robotics →
    Orin Nano (7W) through AGX Orin (60W) and Jetson Thor — JetPack, L4T, Holoscan, Isaac, DeepStream.
    live
  17. Profiling & Debug — Nsight, NVTX, CUPTI →
    Nsight Systems, Nsight Compute, NVTX, CUPTI, DCGM, nvbandwidth — workflow from 'cluster slow' to 'fix line 42'.
    live
  18. Sharing the GPU — MIG, MPS, vGPU →
    Hardware-partitioned MIG, MPS multiplexing, vGPU virtualisation, Kubernetes time-slicing — isolation, performance, licensing trade-offs.
    live
  19. TensorRT-LLM — NVIDIA's Optimised Inference Engine →
    Engine builder, in-flight batching, paged KV-cache, FP8/FP4 quantisation, speculative decoding, TP/PP/EP.
    live
  20. NeMo, NIM & AI Enterprise →
    NeMo Framework, Aligner (RLHF/DPO/PPO), Curator, Guardrails, NIM microservices, Base Command, Run.ai, AI Enterprise bundle.
    live
  21. PTX & SASS — The Real GPU ISAs →
    PTX portable IR and per-arch SASS — HMMA / WGMMA tensor-core ops, LDMATRIX, BAR.SYNC, predication, ptxas optimisations.
    live
  22. Warp Scheduling & SIMT →
    SM partitions, warp schedulers, instruction latencies, occupancy vs ILP, divergence, predication, independent thread scheduling.
    live
  23. HBM Internals →
    HBM2e/3/3e/4 internals — TSVs, channels and pseudo-channels, banks, refresh, ECC modes, the HBM PHY, CoWoS packaging.
    live
  24. Power & Thermal — 1000 W on a Card →
    12V-2x6, multi-phase VRMs, multi-rail, transient response, GPU Boost, P-states, vapour chambers, direct-to-chip liquid cooling.
    live
  25. Process & Packaging →
    TSMC process lineage 16FF to N3/A16, 830 mm² reticle limit, NV-HBI, CoWoS-S/L/R, defect density, the CoWoS supply bottleneck.
    live
  26. Inside Pascal — First HBM & NVLink →
    GP100 (P100) and GP102/104/106 — block diagram, FP-heavy SM, FP16x2 packed math, TSMC 16FF+, HBM2, GDDR5X, NVLink 1.0.
    live
  27. Inside Volta — First Tensor Cores →
    GV100 — TSMC 12FF, four-partition SM with split FP/INT, first-gen tensor cores at 125 TFLOPS, independent thread scheduling, NVLink 2.0 / NVSwitch 1.
    live
  28. Inside Turing — RT Cores, INT8/INT4 Tensor Cores →
    TU102/104/106/116/117 — 1st-gen RT cores, 2nd-gen tensor cores (INT8/INT4), GDDR6, NVLink 2.0 bridge, Tesla T4.
    live
  29. Inside Ampere — 3rd-Gen Tensor Cores, MIG →
    GA100 (TSMC 7N) and GA10x (Samsung 8N), TF32/BF16/2:4 sparsity, the 40 MB L2 leap, async copy, MIG, NVLink 3.0.
    live
  30. Inside Ada — AD102, 96 MB L2, 4th-Gen RT →
    AD102–107 on TSMC 4N — doubled-FP32 SM, 4th-gen tensor cores, 4th-gen RT with OMM/DMM, OFA, 96 MB L2, GDDR6X, the 12VHPWR saga.
    live
  31. Inside Hopper — FP8, TMA, Clusters →
    GH100 / H100 / H200 / GH200 — 4th-gen tensor cores at 1979 TFLOPS, WGMMA, TMA, thread-block clusters and DSMEM, DPX, NVLink 4.
    live
  32. Inside Blackwell — Dual-Die, MX-FP4, NVLink 5 →
    Dual-die B100/B200 with NV-HBI on CoWoS-L, 5th-gen tensor cores at 9 PFLOPS MX-FP4, 2nd-gen Transformer Engine, NVLink 5, NVL72, RTX 50.
    live
  33. Inside the DGX Spark — GB10 Hardware →
    GB10 SoC with 20 ARM cores + Blackwell GPU, NVLink-C2C, 128 GB unified LPDDR5x at 273 GB/s, ~1 PFLOP sparse FP4 at 170 W.
    live
  34. Setting Up DGX Spark →
    Unboxing, OOBE, DGX OS (Ubuntu ARM64), CUDA toolkit, Container Toolkit, SSH, networking, NGC login, first NIM/Ollama/vLLM model.
    live
  35. LLM Inference on DGX Spark — Numbers →
    Realistic tok/s 1B–671B, framework choice (Ollama/vLLM/NIM/TRT-LLM), quant choice (BF16/FP8/MX-FP4/AWQ-INT4), KV placement.
    live
  36. DGX Spark Development Workflow →
    VS Code Remote SSH, NGC containers, PyTorch on ARM64, JupyterHub, QLoRA at 70B, dataset streaming, HF→MX-FP4→NIM, custom Triton.
    live
  37. DGX Spark vs Alternatives →
    Mac Studio M3 Ultra, RTX 5090, RTX PRO 6000 Blackwell, used DGX Station A100, cloud H100/B200, two-Spark pair — capacity, bandwidth, cost-per-token.
    live
  38. BFP4 — NVFP4 & MXFP4 Block-Float Formats →
    Bit-level deep dive into Blackwell's two 4-bit block-FP formats — OCP MXFP4 (E8M0 scale × 32) and NVIDIA NVFP4 (FP8 + FP32 scales × 16) — with BFP history, FP4 E2M1 lattice, 5th-gen tensor-core MMA, throughput across Volta–Blackwell, quality vs FP8/FP6/INT4, interactive decoder.
    live