Same CUDA, dramatically smaller power envelope. Walk through the Jetson family — Orin Nano, Orin NX, AGX Orin, and the upcoming Thor — plus Holoscan for medical edge, Isaac for robotics, DeepStream for video, and JetPack as the BSP that ties it together.
A grounded tour of NVIDIA's edge stack — from the SoC die through the BSP up to the application frameworks — and what it actually feels like to deploy a model on a 15-watt brick.
Not every model runs in a cloud datacenter. Inference happens on devices: in cars, robots, factory cells, surgical theatres, drones, smart cameras, and increasingly on the lab desk. The hardware question isn't "how much DGX can we buy?", it's "what fits in this enclosure, this power budget, this thermal envelope, this BOM?"
The Jetson family scales the same software stack down from 200+ TOPS at 60 W (AGX Orin) to 20 TOPS at 7 W (Orin Nano) — same CUDA, same TensorRT, same containers. That portability is the platform's real moat: a model trained on an H100 in the cloud runs on a 15 W brick in a robot wrist without re-engineering the model.
Tegra is the original mobile / automotive SoC line that NVIDIA started shipping in 2008. It powered the Nintendo Switch, early Tesla autopilot, and Shield TV. Jetson is the same silicon repurposed for AI dev kits and embedded deployment — with the GPU pulled forward to whatever the current desktop architecture is.
The same Tegra silicon also lives inside the NVIDIA Drive automotive platform — Drive Orin for current ADAS / L2+ deployments, and Drive Thor announced for L4 / robotaxi-class compute. Drive shares Jetson's SoC but adds automotive-grade qualification (AEC-Q100, ASIL-D safety island, lock-step cores). If you understand Jetson, you understand Drive at the silicon level.
Each Tegra generation gets a codename that's also the GPU architecture for the discrete card line two years earlier. Orin uses Ampere SMs (RTX 30xx, A100). Thor uses Blackwell SMs (B200, RTX 50xx). So a Jetson generation is roughly "datacenter GPU minus two years, in a 15–60 W envelope, with sensor IO bolted on."
Six SKUs cover the range from a $250 student dev kit to a $1999 industrial brick. All run the same Linux for Tegra (L4T), the same JetPack BSP, the same CUDA. You pick the brick by power envelope and TOPS budget.
| SKU | CUDA cores | Sparse INT8 TOPS | RAM | Power | Price |
|---|---|---|---|---|---|
| Orin Nano 4 GB | 512 (Ampere) | ~20 | 4 GB LPDDR5 | 7–15 W | $250 dev kit |
| Orin Nano 8 GB | 1024 | ~40 | 8 GB LPDDR5 | 7–15 W | $499 dev kit |
| Orin NX 8 GB | 1024 | ~70 | 8 GB LPDDR5 | 10–20 W | $399 module |
| Orin NX 16 GB | 1024 | ~100 | 16 GB LPDDR5 | 10–25 W | $599 module |
| AGX Orin 32 GB | 1792 | ~200 | 32 GB LPDDR5 | 15–40 W | $1499 dev kit |
| AGX Orin 64 GB | 2048 | ~275 | 64 GB LPDDR5 | 15–60 W | $1999 dev kit |
Orin NX ships only as a SoM (system-on-module) — you buy a carrier board (yours or a partner's) and bolt the module on. AGX Orin and Orin Nano have official NVIDIA dev kits with a carrier included; in production designs you typically use the bare SoM and a custom carrier. The connector and module pinout is shared between AGX Orin and AGX Xavier (and is largely shared with Thor) so a well-designed carrier outlasts one generation.
Orin is more than a GPU bolted to an ARM — it is a heterogeneous compute SoC with eight named accelerators. Understanding which engine to schedule each pipeline stage on is what separates a 200 TOPS sticker from 200 TOPS in practice.
Each DLA delivers ~50 TOPS INT8 at <5 W — more efficient per watt than the GPU for the layers it supports (basic conv/pool, no fancy ops). On AGX Orin 64 GB, the headline 275 sparse-INT8 TOPS is roughly 170 from the GPU plus ~100 from the two DLAs. TensorRT can split a graph across GPU and DLA automatically with setDeviceType; the practical recipe is "DLA for the perception backbone, GPU for the head".
Thor is the first Blackwell-generation Jetson. It targets the workloads that didn't fit on Orin: humanoid robots, autonomous trucks, surgical robotics, large industrial automation cells — anywhere multi-modal models, transformer-based perception, and edge LLMs are part of the closed-loop stack.
Two things, mostly. First, FP4 inference means a 70B-class model halves again on top of FP8 — a 70B at FP4 fits in roughly 35 GB, leaving headroom for KV cache. Second, 128 GB unified memory means you can keep an 8–14B vision-language model resident next to a perception stack and an action policy, all sharing the same RAM pool with zero copies. Orin had to choose; Thor doesn't.
JetPack is the BSP and SDK bundle that turns a Jetson SoM into something you can apt install things on. It is what makes a Jetson different from a generic ARM SBC with a Mali GPU.
The headline reason JetPack matters: it ships ARM64 PyTorch / TensorFlow / ONNX Runtime wheels built against the right CUDA version for the SoC. pip install torch from PyPI gives you a CPU-only wheel; the JetPack-bundled wheel actually uses the GPU. Same applies to onnxruntime-gpu and the various Triton backends.
// Identical to anything you'd run on an H100 or 4090.
// Same cudaMalloc, same kernel-launch syntax, same nvcc.
#include <cuda_runtime.h>
__global__ void add(const float* a, const float* b, float* c, int n) {
int i = blockIdx.x * blockDim.x + threadIdx.x;
if (i < n) c[i] = a[i] + b[i];
}
int main() {
const int N = 1 << 20;
float *a, *b, *c;
// On Jetson, prefer cudaMallocManaged — CPU and GPU share LPDDR5.
cudaMallocManaged(&a, N * sizeof(float));
cudaMallocManaged(&b, N * sizeof(float));
cudaMallocManaged(&c, N * sizeof(float));
add<<<(N + 255) / 256, 256>>>(a, b, c, N);
cudaDeviceSynchronize();
cudaFree(a); cudaFree(b); cudaFree(c);
}
// Build: nvcc -arch=sm_87 add.cu -o add (sm_87 = Orin Ampere)
// nvcc -arch=sm_110 add.cu -o add (sm_110 = Thor Blackwell; sm_101 before CUDA 13)
On a discrete-GPU desktop, cudaMallocManaged hides PCIe transfers behind page faults — usually slower than explicit cudaMemcpyAsync. On Jetson the GPU sits on the same LPDDR5 bus as the CPU, so unified memory is actually free: no PCIe, no copy. Default to cudaMallocManaged on Jetson; reach for explicit copies only when you need pinned host memory for a streaming sensor.
Holoscan is NVIDIA's framework for ultra-low-latency sensor processing. It started in medical (endoscopy, ultrasound, MRI) where the AI overlay has to land on the surgeon's monitor inside one frame — ~33 ms at 30 fps, ~16 ms at 60 fps — with no jitter and no hidden host hops.
from holoscan.core import Application
from holoscan.operators import (
VideoStreamReplayerOp, FormatConverterOp,
InferenceOp, HolovizOp,
)
class Endoscopy(Application):
def compose(self):
src = VideoStreamReplayerOp(self, name="src",
directory="/data/endoscopy")
cvt = FormatConverterOp(self, name="cvt",
out_dtype="float32")
infer = InferenceOp(self, name="seg",
model_path_map={"tool_seg":
"/models/tool_seg.engine"})
viz = HolovizOp(self, name="viz",
tensors=[{"name":"mask","type":"color"}])
self.add_flow(src, cvt)
self.add_flow(cvt, infer, {("out", "receivers")})
self.add_flow(cvt, viz) # raw frame to compositor
self.add_flow(infer, viz) # mask overlay to compositor
# All tensors live in CUDA memory; the only host hop is at the model load.
Endoscopy().run()
For a static surgical suite: an RTX 6000 Ada or RTX PRO 6000 Blackwell in a medical-grade workstation, capture card directly into the GPU. For mobile / cart-based: AGX Orin 64 GB or eventually Jetson Thor. Same Holoscan code in both cases; the operator graph doesn't change.
Isaac is NVIDIA's robotics stack. Three components that work together: a simulator, a set of CUDA-accelerated ROS 2 packages, and a foundation-model layer for manipulation and navigation.
# JetPack 6 + Isaac ROS humble container
docker run --runtime nvidia -it --rm \
--network host \
-v /dev/*:/dev/* \
nvcr.io/nvidia/isaac-ros/isaac_ros_visual_slam:latest
# Inside the container, launch the cuVSLAM node:
ros2 launch isaac_ros_visual_slam isaac_ros_visual_slam.launch.py \
setup_for_realsense_zed:=realsense
# Now publish your camera and TF; cuVSLAM publishes /visual_slam/tracking/odometry.
# 30 Hz at 720p on AGX Orin 32 GB, <5 W extra over baseline.
Train a manipulation policy in Isaac Sim on a workstation RTX. Export the same TensorRT engine. Deploy to a Thor-powered robot chassis. The Python you wrote against the simulated robot's ROS topics is the same Python that drives the real robot. That portability across the stack — Sim → ROS → SoC → deployed robot — is the whole pitch.
DeepStream is the pipeline framework for multi-camera video analytics. Built on GStreamer, it composes the canonical "decode → infer → track → encode" pipeline with hardware codecs (NVDEC/NVENC) and TensorRT inference, and scales to dozens of streams on a single Orin.
# /opt/nvidia/deepstream/deepstream-7.0/samples/configs/4cam.txt
[application]
enable-perf-measurement=1
[source0]
type=4 # RTSP
uri=rtsp://cam01.local/stream1
[source1]
type=4; uri=rtsp://cam02.local/stream1
[source2]
type=4; uri=rtsp://cam03.local/stream1
[source3]
type=4; uri=rtsp://cam04.local/stream1
[streammux]
batch-size=4; width=1920; height=1080
[primary-gie] # PeopleNet INT8 from NGC
model-engine-file=peoplenet_int8.engine
batch-size=4; interval=0
[tracker]
ll-lib-file=libnvds_nvmultiobjecttracker.so
ll-config-file=tracker_NvDCF_perf.yml
[secondary-gie0] # Vehicle / accessory classifier
model-engine-file=vehicle_classifier.engine
operate-on-gie-id=1; operate-on-class-ids=2
[sink0]
type=6 # MsgBroker (Kafka)
msg-broker-conn-str=kafka.local;9092;deepstream-events
[sink1]
type=4 # RTSP out for review
rtsp-port=8554
An AGX Orin 32 GB will comfortably handle 30× 1080p30 streams with a PeopleNet primary and a secondary classifier, with the GPU running at <70% and the DLAs idle — you have room to add a re-ID model or a ANPR classifier on top. Below that, an Orin NX 16 GB handles ~16 streams; an Orin Nano 8 GB handles ~6.
Edge LLMs have become real in 2025–26. Not because the silicon got that much bigger — it didn't — but because quantisation got that much better and small models got that much smarter. A 1–8 B model on an Orin is now a viable on-device assistant.
| Model + quant | Hardware | tok/s (decode) | Power | Use case |
|---|---|---|---|---|
| Llama-3-1B BF16 | Orin Nano 8 GB | ~25 | ~10 W | Voice agent on a battery device |
| Phi-3-mini q4 (3.8B) | Orin NX 16 GB | ~30 | ~15 W | Drone narrator, robot dialogue |
| Llama-3-8B q4_K_M | AGX Orin 64 GB | ~12 | ~30 W | On-board assistant in a vehicle |
| Llama-3-8B FP8 | Jetson Thor 128 GB | ~25 (proj.) | ~50 W | VLM stack for humanoid robot |
| Whisper-medium (ASR) | AGX Orin 64 GB | ~6× real-time | ~20 W | Local dictation / live captions |
| 70B Q4 split DGX Spark + Thor | 2-box prototype | experimental | ~400 W total | Workstation-class on-prem RAG |
An emerging recipe for prosumer / lab use: DGX Spark (128 GB unified, ~273 GB/s) for a big base model + Jetson Thor (128 GB unified, similar bandwidth) for a perception / VLM head, linked over 100 GbE. Together they hold a 70B-class model split FP4 across the two pools, with Thor as the gateway closest to the sensors. It's not production-grade today — bandwidth between the two boxes is the cliff — but it's where on-prem 70B at ~400 W total is heading.
Pick a workload, a power constraint, and a latency budget. The planner suggests an SoM, the rough TOPS / RAM / price target, and which JetPack frameworks you'll lean on.
Battery devices live on Orin Nano. Mobile robots and drones live on Orin NX. Vehicles, multi-stream analytics, and surgical carts live on AGX Orin. Humanoids, autonomous trucks, and on-prem edge LLMs in the 8–14 B class live on Jetson Thor. The framework choice is workload-driven, not SoM-driven: video → DeepStream, robotics → Isaac, sensors-with-display → Holoscan, language → TensorRT-LLM on bare JetPack.