An 80-GB H100 sitting at 12% utilisation is a $30k mistake. Walk through every way NVIDIA lets multiple workloads share one GPU — hardware-partitioned MIG, software-multiplexed MPS, hypervisor-level vGPU, and Kubernetes time-slicing — with their isolation, performance, and licensing trade-offs.
Four sharing models, three isolation tiers, one decision to make per cluster. This deck walks the decision tree from "why share at all?" through to ready-to-paste YAML.
The single workload-per-GPU model dates from when GPUs had 8–16 GB of VRAM and one model filled the card. A 2026 H100 with 80–96 GB of HBM running a single 7B inference engine is closer to a yacht being used as a kayak.
80 GB HBM, ~3.35 TB/s bandwidth, 132 SMs. Llama-3 7B FP16 = ~14 GB weights + ~4 GB KV at modest concurrency. ~12% SM utilisation, ~18 GB used. The other 62 GB is paying rent and doing nothing. At a $30k street price the idle slack alone is ~$23k of capital.
Same card sliced 4×. Tenant A serves a 7B variant, B serves embeddings, C runs a small RAG re-ranker, D does fine-tune evals. ~70% SM utilisation, ~72 GB used. Each tenant is happy, capex amortises across 4. This is why MIG exists.
(a) Bursty workloads — aggregate utilisation rarely hits the per-tenant peak. (b) Small workloads — the model fits in a slice. (c) Memory is the constraint, not compute — sharing redistributes both. Fail any of those and a dedicated smaller card is usually cheaper than a slice of a bigger one.
NVIDIA offers four largely orthogonal sharing mechanisms. They live at different layers of the stack and can be combined.
Independent SM/L2/HBM slices, dedicated memory controllers. Each slice is a separate physical GPU instance to the OS. True isolation: noisy-neighbour impossible. A100, A30, H100, H200, B200, plus RTX PRO 6000 Blackwell (up to 4 instances).
One CUDA context, multiple client processes share SMs cooperatively via a daemon. No memory isolation; latency-friendly for many small kernels. User-space MPS has shipped since CUDA 7; Volta added hardware-accelerated MPS with per-client priority & limits.
Hypervisor splits the physical GPU into virtual GPUs. Each VM gets one. Backing can be SMs (time-sliced), a MIG instance, or a full passthrough. Licensed product (NVIDIA AI Enterprise / vGPU).
Multiple pods share one GPU at the device-plugin layer. The GPU Operator advertises N "replicas" of one device; the driver round-robins kernel launches. Loose isolation, dev-grade.
You can run K8s time-slicing on top of MIG (one slice = one schedulable resource), or MPS inside a MIG slice (many cooperating processes per tenant). vGPU can in turn back its VMs with MIG instances. Pick the lowest-numbered isolation that meets your SLO.
MIG is the only sharing mode with hardware enforcement. The GPU is physically partitioned at the GPC (Graphics Processing Cluster) and memory-controller boundary. There is no path through silicon for one slice's traffic to perturb another.
Because controllers are independent, the noisy-neighbour problem that plagues MPS and time-slicing is physically impossible: a runaway memcpy in slice 1 cannot steal HBM bandwidth from slice 2.
| Profile | SM fraction | HBM | Max instances | Typical use |
|---|---|---|---|---|
| 7g.80gb | 7/7 | 80 GB | 1 | Whole card; reverts to single-tenant |
| 4g.40gb | 4/7 | 40 GB | 1 (with smaller leftover) | One large tenant + smaller side slice |
| 3g.40gb | 3/7 | 40 GB | 2 | Mid-size LLM serving (13–30B FP8) |
| 2g.20gb | 2/7 | 20 GB | 3 | 7B inference replicas, fine-tune jobs |
| 1g.10gb | 1/7 | 10 GB | 7 | Many small models, embedders, dev pods |
| 1g.20gb new | 1/7 | 20 GB | 4 | Long-context 7B / KV-heavy small models |
Up to 7 instances per H100. Mix profiles freely as long as the total fits (e.g. 3g + 2g + 1g + 1g).
# Enable MIG mode (requires no compute in flight; persistent across reboot)
nvidia-smi -mig 1
# Create GPU Instances by profile ID: 9 = 3g.40gb, 14 = 2g.20gb, 19 = 1g.10gb
nvidia-smi mig -cgi 9,14,19,19,19
# Create one Compute Instance per GI (default profile = whole GI)
nvidia-smi mig -cci
# See what you got
nvidia-smi -L
# GPU 0: NVIDIA H100 80GB HBM3 (UUID: GPU-...)
# MIG 3g.40gb Device 0: (UUID: MIG-...) ← tenant A sees this as $CUDA_VISIBLE_DEVICES
# MIG 2g.20gb Device 1: (UUID: MIG-...)
# MIG 1g.10gb Device 2: ...
Each MIG device shows up as a UUID that workloads pin via CUDA_VISIBLE_DEVICES=MIG-<uuid>. Tenants see what looks like a normal, smaller, dedicated GPU.
MIG's hardware partitioning is its strength and its straitjacket. The same isolation that prevents noisy neighbours also prevents collaboration.
Slices on the same physical GPU cannot use NVLink to talk to each other — NVLink connects whole GPUs, not partitions. MIG and tensor parallelism don't combine. MIG is for many small workloads, never one big sharded one.
CUDA Unified Memory and CUDA IPC handles assume a shared address space. Each MIG slice is its own context, on its own MMU. Cross-slice handles fail with cudaErrorInvalidDevice.
Inside any single MIG instance the CUDA driver is normal: graphs, streams, async copies, MPS-on-MIG all behave. The boundary is the slice, not below it.
Resizing a layout requires destroying every CI then every GI then re-creating. Workloads must drain. Plan re-partitions at maintenance windows; don't expect Kubernetes-style elastic resize.
sudo nvidia-smi -pm 1. Without it the driver can unload between processes, blowing away MIG configuration on some kernel/driver combos. Datacenter cards default to on; double-check with nvidia-smi -q | grep Persistence.
MIG isolates contention but not failure. A driver crash, ECC double-bit error, or thermal trip affects the whole physical GPU. Plan your tenancy SLOs around "noisy neighbour solved" rather than "blast radius solved".
Think of MIG as n GPUs in one socket, not as a CPU-style cgroup over one GPU. Anything that requires GPUs to cooperate (TP, NCCL all-reduce, P2P, UVM) needs distinct physical GPUs. Anything that just needs a private slice with predictable latency is what MIG is for.
MPS is the older mechanism and is fundamentally different: rather than partition the hardware, it multiplexes CUDA contexts in software so multiple processes look like one to the GPU.
/tmp/nvidia-mpsWithout MPS, the GPU time-slices at the context level: kernel from A finishes, context switch, kernel from B finishes, switch back. Context switches on a GPU are not free. MPS folds all clients into a single context so kernels from any client can co-execute on different SMs concurrently.
CUDA_MPS_ACTIVE_THREAD_PERCENTAGE caps the fraction of SMs a client may dispatch onto, but it's a soft cap, per-launch, not enforced at SM level.# Start the daemon (typically as a systemd unit)
export CUDA_VISIBLE_DEVICES=0
export CUDA_MPS_PIPE_DIRECTORY=/tmp/nvidia-mps
export CUDA_MPS_LOG_DIRECTORY=/var/log/nvidia-mps
nvidia-cuda-mps-control -d
# Cap a client at 30% of SMs (best-effort)
export CUDA_MPS_ACTIVE_THREAD_PERCENTAGE=30
./my_inference_server
MPS shines when a single team runs many cooperating, latency-sensitive processes — e.g. a Triton instance with several model backends, a research fleet of small CUDA jobs, or many short-lived dev kernels. It is the wrong choice for unrelated tenants who don't trust each other.
The decision usually compresses to one question: do these workloads trust each other?
The strongest pattern in production is MIG for tenant isolation, MPS within each MIG slice for many processes per tenant. Tenant A gets a 2g.20gb slice. Inside the slice, Triton runs MPS so its 12 model backends co-execute. Tenant B's 1g.10gb slice is similarly arranged but cannot perturb A's. You get hardware isolation between strangers and software co-execution among friends.
| Property | MIG | MPS | Both |
|---|---|---|---|
| Memory isolation | hard | none | hard between slices |
| Bandwidth isolation | hard | cooperative | hard between slices |
| Concurrent kernels | across slices, not within | within client set | both axes |
| Cards supported | A100, A30, H100, H200, B200, RTX PRO 6000 Blackwell | virtually all CUDA GPUs (HW-accelerated since Volta) | MIG-capable only |
| License | included | included | included |
| Best for | multi-tenant SaaS | one team, many procs | multi-tenant SaaS with rich tenants |
vGPU is NVIDIA's enterprise virtualisation product, formerly branded GRID. It exposes one physical GPU to a hypervisor as N virtual GPUs, each of which a VM consumes as a normal device. Required where the isolation boundary is the VM, not the process.
The most common production target. ESXi + Horizon for VDI; ESXi + AI Enterprise for GPU-backed VMs.
Windows Server platform; vGPU exposed as a DDA (Discrete Device Assignment) plus virtualisation extensions.
Citrix's own hypervisor, the historical home of GRID for VDI.
RHV / OpenShift Virtualization for KVM-based RHEL shops.
Upstream KVM via mediated device (mdev) framework. The path most cloud providers customise.
Anything KVM-based can technically host vGPU if the host driver and licence are present.
The hypervisor's vGPU manager round-robins SM ownership between vGPUs. Multiple VMs share the same SMs over time. Memory is partitioned per vGPU but compute is time-multiplexed. Higher density, lower isolation.
Each vGPU is bound to a MIG instance. The VM gets hardware-isolated SMs and HBM. Best isolation in the vGPU world. Available on A100, A30, H100, H200, B200.
vGPU is licensed per concurrent vGPU, sold as named editions:
| Edition | Use case | List price/year |
|---|---|---|
| vApps | App publishing — XenApp-style | ~$120 |
| vPC | Knowledge-worker VDI | ~$250 |
| RTX Virtual Workstation (vWS) | Engineering / CAD VDI | ~$500 |
| NVIDIA AI Enterprise | Compute / AI VMs (the modern bundle) | ~$3,000 (often $4,500 list per GPU/yr) |
Discounts in volume are deep; single-seat list prices rarely apply to enterprise deals.
Multi-tenant cloud DaaS (Citrix Cloud, AWS WorkSpaces, Azure Virtual Desktop with GPU SKUs); regulated environments — banking, healthcare, defence — where workloads must live inside hardware-virtualised VMs and the security model is "VM = trust boundary".
The "NVIDIA AI Enterprise" (NVAIE) product is the 2026 packaging of vGPU plus a large software estate. If you want vGPU at all on compute cards, you are buying NVAIE.
List price is roughly $4,500 per GPU per year, often steeply discounted in large hyperscaler or enterprise commitments. Multi-year terms (3 / 5 yr) and CSP perpetual licences exist.
| Path | Cost | Support | Includes |
|---|---|---|---|
| Free CUDA toolkit | $0 | community / forum | CUDA, cuDNN, drivers |
| NVIDIA AI Enterprise | ~$4,500/GPU/yr | enterprise SLA | vGPU + NeMo + NIM + Triton + validated images + support |
| DGX Cloud | per-instance, hourly or reserved | NVIDIA-managed | NVAIE bundled, hosted on NVIDIA infrastructure |
vGPU at all (the host driver is licence-gated). Some hyperscaler L40S / H100 deployments require NVAIE under their contract with NVIDIA. The free CUDA toolkit is fine for bare-metal Kubernetes with MIG — but it is not a path to vGPU and carries no support if a driver bug lands in production.
Kubernetes' default device plugin model is unforgiving: one GPU = one pod. Anything finer-grained needs the NVIDIA GPU Operator, which manages drivers, plugins, MIG configuration and time-slicing as a coherent stack.
resources:
limits:
nvidia.com/gpu: 1 # one whole GPU. Nothing else can claim it.
Configured in the GPU Operator's ClusterPolicy: replicas: N tells the device plugin to advertise the same GPU as N "shadows". N pods can claim it concurrently; the driver round-robins their kernel launches.
apiVersion: nvidia.com/v1
kind: ClusterPolicy
spec:
devicePlugin:
config:
name: time-slicing-config
sharing:
timeSlicing:
resources:
- name: nvidia.com/gpu
replicas: 4 # 4 pods may share each GPU
cudaMalloc reduces the pool every other pod sees.For real isolation in K8s, have the GPU Operator configure MIG and surface each slice as its own resource type. Pods request slices like any other typed resource.
resources:
limits:
nvidia.com/gpu.product: H100-PCIE-80GB-MIG-1g.10gb
nvidia.com/gpu: 1
# Operator labels nodes with the available profiles per GPU,
# scheduler matches pod request to a free slice on a labelled node.
Time-slicing is the K8s equivalent of MPS (no isolation, friendly for cooperating workloads). MIG-aware GPU Operator is the K8s equivalent of MIG (hardware partitioning, real tenants). Pick by trust boundary, not convenience.
Three real-world configurations spanning the trust spectrum.
Hardware: 8× H100 PCIe nodes.
Partition: 7×1g.10gb per H100 (MIG).
Density: 56 inference tenants per node, 448 per cluster.
Stack: GPU Operator MIG-aware + Triton + autoscaled K8s deployments.
Workload: 56 fine-tuned 7B variants, each at ≤10 GB FP8 + KV.
Isolation: hardware between tenants, MPS optional inside the slice.
Hardware: 4× A40 nodes (no MIG on A40).
Partition: K8s time-slicing 4× per GPU.
Density: 16 dev users per node concurrently.
Stack: GPU Operator + JupyterHub + per-user notebook pods.
Workload: bursty notebook experiments, embedding jobs, small fine-tunes.
Isolation: none beyond container; works because workloads are bursty and trusted.
Hardware: vSphere host + 4× L40S 48 GB.
Partition: MIG-backed vGPU… on L40S not available, so time-sliced vGPU profiles (e.g. 1g.12gb per VM).
Density: 16 engineers, each with 16 vCPU + 12 GB vGPU VM.
Stack: ESXi + Horizon + NVAIE / vGPU licences.
Isolation: hypervisor-level VM isolation; licence counts per concurrent vGPU.
If your sharing scheme can lose tenant data when one tenant misbehaves, you do not have multi-tenant isolation; you have a friend group with a shared GPU. That is fine for (b), unacceptable for (a) and (c).
Pick a GPU class, a workload mix, and an isolation requirement. The picker recommends a sharing mode, estimates fan-out, flags licensing, and emits the right command or YAML snippet.