Presentations in This Series
- The NVIDIA GPU Family Tree — Pascal to Blackwell →Family timeline, process nodes, dies, transistor counts, memory types and what each generation unlocked. Interactive family explorer.
- Inside the SM — How NVIDIA's Streaming Multiprocessor Evolved →SM internals across Pascal, Volta, Turing, Ampere, Hopper, Ada and Blackwell — schedulers, register file, tensor cores, TMA, clusters.
- Tensor Cores — Five Generations →Every generation from Volta's 4×4×4 FP16 to Blackwell's MX-FP4 — formats, MMA shapes, sparsity, Transformer Engine.
- Memory Hierarchy →Registers, shared, L1/L2, HBM2e/3/3e, GDDR6/6X/7, LPDDR5x unified, Hopper TMA. Decode-speed estimator.
- NVLink & NVSwitch →Scale-up interconnect from NVLink 1 (Pascal) to NVLink 5 (NVL72), NVLink-C2C, the rack-scale superpod.
- Ampere — A100, RTX 30, the LLM Era →GA100 + GA10x — 3rd-gen tensor cores (TF32, BF16, 2:4 sparsity), MIG, NVLink 3 with NVSwitch 2.
- Hopper — H100, FP8, Transformer Engine →GH100, H100/H200/GH200 — 4th-gen tensor cores, native FP8, TMA, thread-block clusters and DSMEM, DPX, NVLink 4.
- Ada Lovelace — RTX 40, L40S, Consumer-Class AI →AD102 + RTX 40 / L40S / L4 — 4th-gen RT cores, DLSS 3, the FP8 split, no-NVLink consequences.
- Blackwell — Dual-Die, FP4, NVL72 →B100/B200/GB200 — dual-die NV-HBI, 5th-gen tensor cores with MX-FP4, 2nd-gen Transformer Engine, RAS engine, NVLink 5, NVL72 superpod.
- Software Stack & Performance →CUDA stack — driver, cuBLAS, cuDNN, CUTLASS, Transformer Engine, NCCL, Triton, TensorRT-LLM. End-to-end LLM tok/s calculator.
- DGX, HGX, MGX — Datacenter Reference Platforms →Datacenter platforms, OAM, BasePOD and SuperPOD blueprints, NVL72, the OEM ecosystem and DGX Cloud.
- GeForce, RTX Pro, Tesla, A/H/L/B — Decoding the Lineup →Field guide to every NVIDIA product family — naming logic, EULA boundaries, driver branches, the same-die-different-card patterns.
- Networking — InfiniBand, ConnectX, BlueField →Cross-node fabric — ConnectX HCAs, Quantum IB and Spectrum-X switches, BlueField DPUs, GPUDirect RDMA / Storage, NCCL, SHARP.
- PCIe & GPUDirect →PCIe 3 to 6, Resizable BAR, IOMMU, ACS, NUMA pinning, GPUDirect P2P/RDMA/Storage. Topology lint.
- Grace — NVIDIA's ARM CPU, GH200, GB200 →Grace 72-core Neoverse V2, NVLink-C2C 900 GB/s, GH200, GB200, Extended GPU Memory, DGX Spark.
- Jetson — Edge AI & Robotics →Orin Nano (7W) through AGX Orin (60W) and Jetson Thor — JetPack, L4T, Holoscan, Isaac, DeepStream.
- Profiling & Debug — Nsight, NVTX, CUPTI →Nsight Systems, Nsight Compute, NVTX, CUPTI, DCGM, nvbandwidth — workflow from 'cluster slow' to 'fix line 42'.
- Sharing the GPU — MIG, MPS, vGPU →Hardware-partitioned MIG, MPS multiplexing, vGPU virtualisation, Kubernetes time-slicing — isolation, performance, licensing trade-offs.
- TensorRT-LLM — NVIDIA's Optimised Inference Engine →Engine builder, in-flight batching, paged KV-cache, FP8/FP4 quantisation, speculative decoding, TP/PP/EP.
- NeMo, NIM & AI Enterprise →NeMo Framework, Aligner (RLHF/DPO/PPO), Curator, Guardrails, NIM microservices, Base Command, Run.ai, AI Enterprise bundle.
- PTX & SASS — The Real GPU ISAs →PTX portable IR and per-arch SASS — HMMA / WGMMA tensor-core ops, LDMATRIX, BAR.SYNC, predication, ptxas optimisations.
- Warp Scheduling & SIMT →SM partitions, warp schedulers, instruction latencies, occupancy vs ILP, divergence, predication, independent thread scheduling.
- HBM Internals →HBM2e/3/3e/4 internals — TSVs, channels and pseudo-channels, banks, refresh, ECC modes, the HBM PHY, CoWoS packaging.
- Power & Thermal — 1000 W on a Card →12V-2x6, multi-phase VRMs, multi-rail, transient response, GPU Boost, P-states, vapour chambers, direct-to-chip liquid cooling.
- Process & Packaging →TSMC process lineage 16FF to N3/A16, 830 mm² reticle limit, NV-HBI, CoWoS-S/L/R, defect density, the CoWoS supply bottleneck.
- Inside Pascal — First HBM & NVLink →GP100 (P100) and GP102/104/106 — block diagram, FP-heavy SM, FP16x2 packed math, TSMC 16FF+, HBM2, GDDR5X, NVLink 1.0.
- Inside Volta — First Tensor Cores →GV100 — TSMC 12FF, four-partition SM with split FP/INT, first-gen tensor cores at 125 TFLOPS, independent thread scheduling, NVLink 2.0 / NVSwitch 1.
- Inside Turing — RT Cores, INT8/INT4 Tensor Cores →TU102/104/106/116/117 — 1st-gen RT cores, 2nd-gen tensor cores (INT8/INT4), GDDR6, NVLink 2.0 bridge, Tesla T4.
- Inside Ampere — 3rd-Gen Tensor Cores, MIG →GA100 (TSMC 7N) and GA10x (Samsung 8N), TF32/BF16/2:4 sparsity, the 40 MB L2 leap, async copy, MIG, NVLink 3.0.
- Inside Ada — AD102, 96 MB L2, 4th-Gen RT →AD102–107 on TSMC 4N — doubled-FP32 SM, 4th-gen tensor cores, 4th-gen RT with OMM/DMM, OFA, 96 MB L2, GDDR6X, the 12VHPWR saga.
- Inside Hopper — FP8, TMA, Clusters →GH100 / H100 / H200 / GH200 — 4th-gen tensor cores at 1979 TFLOPS, WGMMA, TMA, thread-block clusters and DSMEM, DPX, NVLink 4.
- Inside Blackwell — Dual-Die, MX-FP4, NVLink 5 →Dual-die B100/B200 with NV-HBI on CoWoS-L, 5th-gen tensor cores at 9 PFLOPS MX-FP4, 2nd-gen Transformer Engine, NVLink 5, NVL72, RTX 50.
- Inside the DGX Spark — GB10 Hardware →GB10 SoC with 20 ARM cores + Blackwell GPU, NVLink-C2C, 128 GB unified LPDDR5x at 273 GB/s, ~1 PFLOP sparse FP4 at 170 W.
- Setting Up DGX Spark →Unboxing, OOBE, DGX OS (Ubuntu ARM64), CUDA toolkit, Container Toolkit, SSH, networking, NGC login, first NIM/Ollama/vLLM model.
- LLM Inference on DGX Spark — Numbers →Realistic tok/s 1B–671B, framework choice (Ollama/vLLM/NIM/TRT-LLM), quant choice (BF16/FP8/MX-FP4/AWQ-INT4), KV placement.
- DGX Spark Development Workflow →VS Code Remote SSH, NGC containers, PyTorch on ARM64, JupyterHub, QLoRA at 70B, dataset streaming, HF→MX-FP4→NIM, custom Triton.
- DGX Spark vs Alternatives →Mac Studio M3 Ultra, RTX 5090, RTX PRO 6000 Blackwell, used DGX Station A100, cloud H100/B200, two-Spark pair — capacity, bandwidth, cost-per-token.
- BFP4 — NVFP4 & MXFP4 Block-Float Formats →Bit-level deep dive into Blackwell's two 4-bit block-FP formats — OCP MXFP4 (E8M0 scale × 32) and NVIDIA NVFP4 (FP8 + FP32 scales × 16) — with BFP history, FP4 E2M1 lattice, 5th-gen tensor-core MMA, throughput across Volta–Blackwell, quality vs FP8/FP6/INT4, interactive decoder.