NVIDIA GPU Architectures Series — Presentation 16

Jetson — NVIDIA's Edge AI and Robotics Stack

Same CUDA, dramatically smaller power envelope. Walk through the Jetson family — Orin Nano, Orin NX, AGX Orin, and the upcoming Thor — plus Holoscan for medical edge, Isaac for robotics, DeepStream for video, and JetPack as the BSP that ties it together.

Jetson OrinAGX OrinOrin NX Orin NanoThorJetPack L4THoloscan IsaacDeepStream
SoC → JetPack → Holoscan/Isaac/DeepStream → Sensors → Edge LLM → Fleet
00

Topics We'll Cover

A grounded tour of NVIDIA's edge stack — from the SoC die through the BSP up to the application frameworks — and what it actually feels like to deploy a model on a 15-watt brick.

01

Edge AI in 2026 — Where the GPUs Live

Not every model runs in a cloud datacenter. Inference happens on devices: in cars, robots, factory cells, surgical theatres, drones, smart cameras, and increasingly on the lab desk. The hardware question isn't "how much DGX can we buy?", it's "what fits in this enclosure, this power budget, this thermal envelope, this BOM?"

The hard constraints

  • Power — 5–60 W typical; battery devices want <15 W.
  • Thermal — passive heatsink or a small fan; no datacenter air.
  • Latency — can't round-trip to cloud; control loops want <10 ms, perception <50 ms.
  • Connectivity — intermittent, expensive, or absent.
  • Cost — $200–$2000 per shipped unit, not per hour.

What the edge wants from a GPU

  • Real CUDA — the same kernels you trained with.
  • TensorRT — INT8 / FP16 / BF16 graph optimisation (FP8 / FP4 only on Thor-class Blackwell).
  • Hardware video codecs (NVENC/NVDEC) for camera streams.
  • Image signal processors and direct sensor capture (CSI, GMSL).
  • A real Linux distribution with proper drivers.
  • Containers: pull the same image you ran in the cloud.
Cars
·
Robots
·
Factories
·
Surgical theatres
·
Drones
·
Smart cameras
·
Lab desks
NVIDIA's edge answer

The Jetson family scales the same software stack down from 200+ TOPS at 60 W (AGX Orin) to 20 TOPS at 7 W (Orin Nano) — same CUDA, same TensorRT, same containers. That portability is the platform's real moat: a model trained on an H100 in the cloud runs on a 15 W brick in a robot wrist without re-engineering the model.

02

Tegra → Jetson — A Brief Lineage

Tegra is the original mobile / automotive SoC line that NVIDIA started shipping in 2008. It powered the Nintendo Switch, early Tesla autopilot, and Shield TV. Jetson is the same silicon repurposed for AI dev kits and embedded deployment — with the GPU pulled forward to whatever the current desktop architecture is.

2014
Jetson TK1
Kepler GPU
first AI dev kit
2015–17
TX1
Maxwell
TX2
Pascal
2018
Xavier / AGX Xavier
Volta
first DLA
2019–22
Nano (legacy)
Maxwell budget
$99 entry
2022+
Orin Nano / NX / AGX
Ampere SM
current workhorse
2025+
Jetson Thor
Blackwell SM
FP4 native

Tegra family beyond Jetson

The same Tegra silicon also lives inside the NVIDIA Drive automotive platform — Drive Orin for current ADAS / L2+ deployments, and Drive Thor announced for L4 / robotaxi-class compute. Drive shares Jetson's SoC but adds automotive-grade qualification (AEC-Q100, ASIL-D safety island, lock-step cores). If you understand Jetson, you understand Drive at the silicon level.

The naming pattern

Each Tegra generation gets a codename that's also the GPU architecture for the discrete card line two years earlier. Orin uses Ampere SMs (RTX 30xx, A100). Thor uses Blackwell SMs (B200, RTX 50xx). So a Jetson generation is roughly "datacenter GPU minus two years, in a 15–60 W envelope, with sensor IO bolted on."

03

Jetson Orin Family — The Workhorse

Six SKUs cover the range from a $250 student dev kit to a $1999 industrial brick. All run the same Linux for Tegra (L4T), the same JetPack BSP, the same CUDA. You pick the brick by power envelope and TOPS budget.

SKUCUDA coresSparse INT8 TOPSRAMPowerPrice
Orin Nano 4 GB512 (Ampere)~204 GB LPDDR57–15 W$250 dev kit
Orin Nano 8 GB1024~408 GB LPDDR57–15 W$499 dev kit
Orin NX 8 GB1024~708 GB LPDDR510–20 W$399 module
Orin NX 16 GB1024~10016 GB LPDDR510–25 W$599 module
AGX Orin 32 GB1792~20032 GB LPDDR515–40 W$1499 dev kit
AGX Orin 64 GB2048~27564 GB LPDDR515–60 W$1999 dev kit

Where each tier lands in practice

Orin Nano
7–15 W
smart camera, hobbyist robotics, edge sensors
Orin NX
10–25 W
drones, ROS robots, multi-cam analytics
AGX Orin 32
15–40 W
autonomous mobile robots, surgical aux
AGX Orin 64
15–60 W
flagship: edge LLMs, Holoscan, multi-stream DeepStream
Module vs dev kit

Orin NX ships only as a SoM (system-on-module) — you buy a carrier board (yours or a partner's) and bolt the module on. AGX Orin and Orin Nano have official NVIDIA dev kits with a carrier included; in production designs you typically use the bare SoM and a custom carrier. The connector and module pinout is shared between AGX Orin and AGX Xavier (and is largely shared with Thor) so a well-designed carrier outlasts one generation.

04

The Orin SoC Internals

Orin is more than a GPU bolted to an ARM — it is a heterogeneous compute SoC with eight named accelerators. Understanding which engine to schedule each pipeline stage on is what separates a 200 TOPS sticker from 200 TOPS in practice.

CPU + safety complex

  • 8× ARM Cortex-A78AE — "AE" is the functional safety variant of the A78; supports lockstep pairing for ASIL-rated workloads.
  • Cortex-R52 safety island — an independent realtime CPU for boot integrity, watchdogs and failover.
  • Security island — root of trust, fuses, secure boot key storage, crypto engine.
  • Up to 12 MB L2 + 4 MB L3, all coherent over the on-chip fabric.

GPU + AI accelerators

  • Ampere GPU — 1024–2048 CUDA cores, 32–64 tensor cores, FP16/BF16/INT8.
  • 2× DLA (Deep Learning Accelerator) — fixed-function INT8 NN engines; hand off conv/pool layers to free the GPU for higher-precision work.
  • PVA (Programmable Vision Accelerator) — classical CV (stereo, optical flow, feature tracking).
  • VIC — image format conversion / scaling without touching the GPU.
  • 2× ISP — raw sensor → YUV pipeline (Bayer demosaic, denoise, AE/AWB).
  • NVENC / NVDEC — H.264/H.265/AV1 encode and decode.
  • 12× CSI lanes — up to 16 cameras via virtual channels.
AGX Orin SoC — floor-plan sketch CPU complex 8× Cortex-A78AE (safety) Cortex-R52 safety island Security island / RoT 12 MB L2 + 4 MB L3 Ampere GPU 1024–2048 CUDA cores 32–64 tensor cores FP16 / BF16 / INT8 CUDA + cuDNN + TensorRT 2× DLA Fixed-fn INT8 NN engine Hand-off conv / pool Frees GPU for FP work Together: 200+ TOPS PVA Programmable Vision Accelerator Stereo, optflow, feature tracking VIC: format conv, scaling, blend 2× ISP Raw → YUV pipeline Bayer demosaic AE / AWB / denoise 12× CSI-2 lanes up to 16 cams via VC GMSL via deserializer NVENC H.264/265/AV1 encode NVDEC Same family decode Zero-copy to GPU Memory LPDDR5 unified 204 GB/s (AGX 64) Coherent for CPU+GPU PCIe Gen4 root 10/25 GbE, USB-C 3.2
Why the DLAs matter

Each DLA delivers ~50 TOPS INT8 at <5 W — more efficient per watt than the GPU for the layers it supports (basic conv/pool, no fancy ops). On AGX Orin 64 GB, the headline 275 sparse-INT8 TOPS is roughly 170 from the GPU plus ~100 from the two DLAs. TensorRT can split a graph across GPU and DLA automatically with setDeviceType; the practical recipe is "DLA for the perception backbone, GPU for the head".

05

Jetson Thor — The 2025 Successor

Thor is the first Blackwell-generation Jetson. It targets the workloads that didn't fit on Orin: humanoid robots, autonomous trucks, surgical robotics, large industrial automation cells — anywhere multi-modal models, transformer-based perception, and edge LLMs are part of the closed-loop stack.

Compute

  • Blackwell GPU with 5th-gen tensor cores including FP4 (MX-FP4) and FP8.
  • 2070 sparse FP4 TFLOPS (1035 dense).
  • 14× ARM Neoverse V3AE — server-class out-of-order cores with the AE safety variant.
  • Transformer Engine inherits everything from H100/B200.

Memory and IO

  • 128 GB LPDDR5x unified — CPU and GPU share one pool.
  • ~273 GB/s aggregate memory bandwidth.
  • Power envelope 40–130 W, configurable per design.
  • Mostly drop-in carrier compatibility with AGX Orin SoM-style designs — your existing carrier has a sporting chance of working.
  • QSFP / 100 GbE, multi-cam GMSL, PCIe Gen5.

Headline targets

Humanoid robots
·
Boston Dynamics
·
Figure
·
Apptronik
Autonomous trucks
·
Plus
·
TuSimple
·
Surgical robotics
·
Industrial cells
What Thor unlocks vs Orin

Two things, mostly. First, FP4 inference means a 70B-class model halves again on top of FP8 — a 70B at FP4 fits in roughly 35 GB, leaving headroom for KV cache. Second, 128 GB unified memory means you can keep an 8–14B vision-language model resident next to a perception stack and an action policy, all sharing the same RAM pool with zero copies. Orin had to choose; Thor doesn't.

06

JetPack — The Software Stack

JetPack is the BSP and SDK bundle that turns a Jetson SoM into something you can apt install things on. It is what makes a Jetson different from a generic ARM SBC with a Mali GPU.

L0 Kernel
L4T (Linux for Tegra)
Ubuntu base 22.04 / 24.04
NVIDIA kernel + display
L1 Compute
CUDA
cuDNN
TensorRT
cuBLAS / cuFFT
L2 Sensor IO
libargus (camera)
V4L2
GStreamer + nvenc/dec
VPI (vision primitives)
L3 Frameworks
PyTorch (ARM64 wheel)
TensorFlow
ONNX Runtime
Triton
L4 Apps
DeepStream
Holoscan
Isaac ROS
cuOpt
AI Enterprise

Release cadence

The headline reason JetPack matters: it ships ARM64 PyTorch / TensorFlow / ONNX Runtime wheels built against the right CUDA version for the SoC. pip install torch from PyPI gives you a CPU-only wheel; the JetPack-bundled wheel actually uses the GPU. Same applies to onnxruntime-gpu and the various Triton backends.

A first kernel on Jetson — same as desktop CUDA
// Identical to anything you'd run on an H100 or 4090.
// Same cudaMalloc, same kernel-launch syntax, same nvcc.
#include <cuda_runtime.h>

__global__ void add(const float* a, const float* b, float* c, int n) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i < n) c[i] = a[i] + b[i];
}

int main() {
    const int N = 1 << 20;
    float *a, *b, *c;
    // On Jetson, prefer cudaMallocManaged — CPU and GPU share LPDDR5.
    cudaMallocManaged(&a, N * sizeof(float));
    cudaMallocManaged(&b, N * sizeof(float));
    cudaMallocManaged(&c, N * sizeof(float));

    add<<<(N + 255) / 256, 256>>>(a, b, c, N);
    cudaDeviceSynchronize();
    cudaFree(a); cudaFree(b); cudaFree(c);
}
// Build:  nvcc -arch=sm_87 add.cu -o add   (sm_87 = Orin Ampere)
//         nvcc -arch=sm_110 add.cu -o add  (sm_110 = Thor Blackwell; sm_101 before CUDA 13)
Use unified memory on Jetson

On a discrete-GPU desktop, cudaMallocManaged hides PCIe transfers behind page faults — usually slower than explicit cudaMemcpyAsync. On Jetson the GPU sits on the same LPDDR5 bus as the CPU, so unified memory is actually free: no PCIe, no copy. Default to cudaMallocManaged on Jetson; reach for explicit copies only when you need pinned host memory for a streaming sensor.

07

Holoscan — Real-Time Sensor AI

Holoscan is NVIDIA's framework for ultra-low-latency sensor processing. It started in medical (endoscopy, ultrasound, MRI) where the AI overlay has to land on the surgeon's monitor inside one frame — ~33 ms at 30 fps, ~16 ms at 60 fps — with no jitter and no hidden host hops.

Why Holoscan exists

  • Generic sensor pipelines copy through host RAM at every step.
  • Each copy adds ~1–5 ms and a queue boundary.
  • OS schedulers introduce jitter; ROS message passing is best-effort.
  • Surgical / radiology workflows can't tolerate any of this.

What Holoscan does differently

  • Zero-copy CUDA buffers — sensor → GPU → NN → display, never touch host.
  • Deterministic scheduling — budgeted operator slots, not best-effort threads.
  • GXF (Graph Execution Framework) — declarative operator graphs.
  • Both x86 (RTX A6000, RTX PRO 6000) and Jetson Orin / Thor targets.

Endoscopy pipeline — 6 operators

Camera Capture (CSI / capture card)
↓
Format Conversion (VIC, zero-copy)
↓
Resize (CUDA NPP)
↓
Tool Segmentation (TensorRT)
↓
Overlay Composition (CUDA)
↓
Display Sink (HDMI / DisplayPort, vsync-locked)
Holoscan Python — minimal endoscopy graph
from holoscan.core import Application
from holoscan.operators import (
    VideoStreamReplayerOp, FormatConverterOp,
    InferenceOp, HolovizOp,
)

class Endoscopy(Application):
    def compose(self):
        src     = VideoStreamReplayerOp(self, name="src",
                                          directory="/data/endoscopy")
        cvt     = FormatConverterOp(self, name="cvt",
                                          out_dtype="float32")
        infer   = InferenceOp(self, name="seg",
                                    model_path_map={"tool_seg":
                                        "/models/tool_seg.engine"})
        viz     = HolovizOp(self, name="viz",
                                  tensors=[{"name":"mask","type":"color"}])

        self.add_flow(src, cvt)
        self.add_flow(cvt, infer, {("out", "receivers")})
        self.add_flow(cvt, viz)        # raw frame to compositor
        self.add_flow(infer, viz)      # mask overlay to compositor

# All tensors live in CUDA memory; the only host hop is at the model load.
Endoscopy().run()
The deployment shape

For a static surgical suite: an RTX 6000 Ada or RTX PRO 6000 Blackwell in a medical-grade workstation, capture card directly into the GPU. For mobile / cart-based: AGX Orin 64 GB or eventually Jetson Thor. Same Holoscan code in both cases; the operator graph doesn't change.

08

Isaac — Robotics Platform

Isaac is NVIDIA's robotics stack. Three components that work together: a simulator, a set of CUDA-accelerated ROS 2 packages, and a foundation-model layer for manipulation and navigation.

Isaac Sim

  • Built on NVIDIA Omniverse.
  • GPU-accelerated PhysX physics.
  • Photoreal rendering for synthetic data generation.
  • Domain randomisation for sim-to-real.
  • Used for policy training, regression suites, and synthetic dataset bring-up.

Isaac ROS

  • Drop-in CUDA-accelerated ROS 2 packages.
  • Visual SLAM (cuVSLAM).
  • Stereo depth (ESS DNN).
  • Semantic segmentation (PeopleSemSegNet, etc.).
  • Point-cloud processing.
  • All running on Orin / Thor at frame rate.

Isaac Manipulator / Perceptor

  • Manipulator — foundation models for robotic pick-and-place.
  • Perceptor — foundation models for navigation perception.
  • Pretrained on huge synthetic + real datasets.
  • Fine-tune for your gripper / floor plan / lighting.

Sim-to-real loop

Isaac Sim (RTX desktop)
→
Synthetic data + RL training
→
TensorRT engine
→
AGX Orin / Thor on robot
→
Real-world telemetry
→
(loop)
Launching Isaac ROS visual SLAM on Orin
# JetPack 6 + Isaac ROS humble container
docker run --runtime nvidia -it --rm \
    --network host \
    -v /dev/*:/dev/* \
    nvcr.io/nvidia/isaac-ros/isaac_ros_visual_slam:latest

# Inside the container, launch the cuVSLAM node:
ros2 launch isaac_ros_visual_slam isaac_ros_visual_slam.launch.py \
    setup_for_realsense_zed:=realsense

# Now publish your camera and TF; cuVSLAM publishes /visual_slam/tracking/odometry.
# 30 Hz at 720p on AGX Orin 32 GB, <5 W extra over baseline.
The Isaac promise

Train a manipulation policy in Isaac Sim on a workstation RTX. Export the same TensorRT engine. Deploy to a Thor-powered robot chassis. The Python you wrote against the simulated robot's ROS topics is the same Python that drives the real robot. That portability across the stack — Sim → ROS → SoC → deployed robot — is the whole pitch.

09

DeepStream — Video Analytics at the Edge

DeepStream is the pipeline framework for multi-camera video analytics. Built on GStreamer, it composes the canonical "decode → infer → track → encode" pipeline with hardware codecs (NVDEC/NVENC) and TensorRT inference, and scales to dozens of streams on a single Orin.

RTSP / CSI / file source × N
↓
NVDEC (hardware H.264 / H.265 / AV1 decode)
↓
nvstreammux (batch frames across cameras)
↓
Primary detector (TensorRT, INT8)
↓
nvtracker (NvDCF / Re-ID)
↓
Secondary classifier (per detection)
↓
OSD overlay (CUDA)
↓
NVENC + RTSP / Kafka / message broker out

What you get for free

Where DeepStream gets used

DeepStream config — 4 cameras + primary + secondary + tracker
# /opt/nvidia/deepstream/deepstream-7.0/samples/configs/4cam.txt

[application]
enable-perf-measurement=1

[source0]
type=4           # RTSP
uri=rtsp://cam01.local/stream1
[source1]
type=4; uri=rtsp://cam02.local/stream1
[source2]
type=4; uri=rtsp://cam03.local/stream1
[source3]
type=4; uri=rtsp://cam04.local/stream1

[streammux]
batch-size=4; width=1920; height=1080

[primary-gie]          # PeopleNet INT8 from NGC
model-engine-file=peoplenet_int8.engine
batch-size=4; interval=0

[tracker]
ll-lib-file=libnvds_nvmultiobjecttracker.so
ll-config-file=tracker_NvDCF_perf.yml

[secondary-gie0]       # Vehicle / accessory classifier
model-engine-file=vehicle_classifier.engine
operate-on-gie-id=1; operate-on-class-ids=2

[sink0]
type=6           # MsgBroker (Kafka)
msg-broker-conn-str=kafka.local;9092;deepstream-events
[sink1]
type=4           # RTSP out for review
rtsp-port=8554
The Orin sweet spot

An AGX Orin 32 GB will comfortably handle 30× 1080p30 streams with a PeopleNet primary and a secondary classifier, with the GPU running at <70% and the DLAs idle — you have room to add a re-ID model or a ANPR classifier on top. Below that, an Orin NX 16 GB handles ~16 streams; an Orin Nano 8 GB handles ~6.

10

LLMs at the Edge — Practical Numbers

Edge LLMs have become real in 2025–26. Not because the silicon got that much bigger — it didn't — but because quantisation got that much better and small models got that much smarter. A 1–8 B model on an Orin is now a viable on-device assistant.

Model + quantHardwaretok/s (decode)PowerUse case
Llama-3-1B BF16Orin Nano 8 GB~25~10 WVoice agent on a battery device
Phi-3-mini q4 (3.8B)Orin NX 16 GB~30~15 WDrone narrator, robot dialogue
Llama-3-8B q4_K_MAGX Orin 64 GB~12~30 WOn-board assistant in a vehicle
Llama-3-8B FP8Jetson Thor 128 GB~25 (proj.)~50 WVLM stack for humanoid robot
Whisper-medium (ASR)AGX Orin 64 GB~6× real-time~20 WLocal dictation / live captions
70B Q4 split DGX Spark + Thor2-box prototypeexperimental~400 W totalWorkstation-class on-prem RAG

Why the edge wins for some workloads

Latency wins

  • 50–100 ms TTFT for a small LLM on Orin.
  • Cloud round-trip is 300–1000 ms end-to-end (network + queueing + cold start).
  • For voice or closed-loop control the difference is the difference between "feels alive" and "feels broken".

Privacy wins

  • Patient data never leaves the surgical cart.
  • Factory IP never leaves the cell.
  • Soldier comms never leave the radio.
  • Your assistant doesn't need a working internet connection.
The two-box edge LLM pattern

An emerging recipe for prosumer / lab use: DGX Spark (128 GB unified, ~273 GB/s) for a big base model + Jetson Thor (128 GB unified, similar bandwidth) for a perception / VLM head, linked over 100 GbE. Together they hold a 70B-class model split FP4 across the two pools, with Thor as the gateway closest to the sensors. It's not production-grade today — bandwidth between the two boxes is the cliff — but it's where on-prem 70B at ~400 W total is heading.

11

Interactive: Edge Use-Case Picker

Pick a workload, a power constraint, and a latency budget. The planner suggests an SoM, the rough TOPS / RAM / price target, and which JetPack frameworks you'll lean on.

SoM
—
TOPS target
—
RAM
—
Approx price
—
The cheat-sheet shape

Battery devices live on Orin Nano. Mobile robots and drones live on Orin NX. Vehicles, multi-stream analytics, and surgical carts live on AGX Orin. Humanoids, autonomous trucks, and on-prem edge LLMs in the 8–14 B class live on Jetson Thor. The framework choice is workload-driven, not SoM-driven: video → DeepStream, robotics → Isaac, sensors-with-display → Holoscan, language → TensorRT-LLM on bare JetPack.