NVIDIA GPU Architectures Series — Presentation 18

Sharing the GPU — MIG, MPS, vGPU, and Kubernetes Time-Slicing

An 80-GB H100 sitting at 12% utilisation is a $30k mistake. Walk through every way NVIDIA lets multiple workloads share one GPU — hardware-partitioned MIG, software-multiplexed MPS, hypervisor-level vGPU, and Kubernetes time-slicing — with their isolation, performance, and licensing trade-offs.

MIGMPSvGPU GRIDK8stime-slicing GPU Operatorisolationmulti-tenant
One GPU → MIG / MPS / vGPU / K8s → Many Workloads → SLOs → Cost
00

Topics We'll Cover

Four sharing models, three isolation tiers, one decision to make per cluster. This deck walks the decision tree from "why share at all?" through to ready-to-paste YAML.

01

Why Share a GPU?

The single workload-per-GPU model dates from when GPUs had 8–16 GB of VRAM and one model filled the card. A 2026 H100 with 80–96 GB of HBM running a single 7B inference engine is closer to a yacht being used as a kayak.

One tenant, one H100

80 GB HBM, ~3.35 TB/s bandwidth, 132 SMs. Llama-3 7B FP16 = ~14 GB weights + ~4 GB KV at modest concurrency. ~12% SM utilisation, ~18 GB used. The other 62 GB is paying rent and doing nothing. At a $30k street price the idle slack alone is ~$23k of capital.

Four tenants, one H100

Same card sliced 4×. Tenant A serves a 7B variant, B serves embeddings, C runs a small RAG re-ranker, D does fine-tune evals. ~70% SM utilisation, ~72 GB used. Each tenant is happy, capex amortises across 4. This is why MIG exists.

Cost arithmetic

1× H100 $30,000 capex → one tenant → $30k/tenant
H100 ÷ 4 $30,000 / 4 = $7,500/tenant (MIG 2g.20gb each)
L4 each $7–15k per L4 24 GB (still less BW, no FP8 in some kernels)
When sharing pays

(a) Bursty workloads — aggregate utilisation rarely hits the per-tenant peak. (b) Small workloads — the model fits in a slice. (c) Memory is the constraint, not compute — sharing redistributes both. Fail any of those and a dedicated smaller card is usually cheaper than a slice of a bigger one.

02

The Four Sharing Models

NVIDIA offers four largely orthogonal sharing mechanisms. They live at different layers of the stack and can be combined.

MIG — Hardware Partitioning

Independent SM/L2/HBM slices, dedicated memory controllers. Each slice is a separate physical GPU instance to the OS. True isolation: noisy-neighbour impossible. A100, A30, H100, H200, B200, plus RTX PRO 6000 Blackwell (up to 4 instances).

MPS — Software Multiplexing

One CUDA context, multiple client processes share SMs cooperatively via a daemon. No memory isolation; latency-friendly for many small kernels. User-space MPS has shipped since CUDA 7; Volta added hardware-accelerated MPS with per-client priority & limits.

vGPU — VM-Level Virtualisation

Hypervisor splits the physical GPU into virtual GPUs. Each VM gets one. Backing can be SMs (time-sliced), a MIG instance, or a full passthrough. Licensed product (NVIDIA AI Enterprise / vGPU).

K8s Time-Slicing — Container Scheduling

Multiple pods share one GPU at the device-plugin layer. The GPU Operator advertises N "replicas" of one device; the driver round-robins kernel launches. Loose isolation, dev-grade.

Where each one lives in the stack

app
Tenant A Tenant B Tenant C Tenant D
k8s
GPU Operator device-plugin time-slicing
VM
vGPU (hypervisor) vSphere / KVM / Hyper-V
driver
MPS daemon CUDA runtime
silicon
MIG slices SMs / L2 / HBM
Stackable

You can run K8s time-slicing on top of MIG (one slice = one schedulable resource), or MPS inside a MIG slice (many cooperating processes per tenant). vGPU can in turn back its VMs with MIG instances. Pick the lowest-numbered isolation that meets your SLO.

03

MIG — Multi-Instance GPU

MIG is the only sharing mode with hardware enforcement. The GPU is physically partitioned at the GPC (Graphics Processing Cluster) and memory-controller boundary. There is no path through silicon for one slice's traffic to perturb another.

What gets partitioned

Because controllers are independent, the noisy-neighbour problem that plagues MPS and time-slicing is physically impossible: a runaway memcpy in slice 1 cannot steal HBM bandwidth from slice 2.

H100 80 GB — canonical profile catalogue

ProfileSM fractionHBMMax instancesTypical use
7g.80gb7/780 GB1Whole card; reverts to single-tenant
4g.40gb4/740 GB1 (with smaller leftover)One large tenant + smaller side slice
3g.40gb3/740 GB2Mid-size LLM serving (13–30B FP8)
2g.20gb2/720 GB37B inference replicas, fine-tune jobs
1g.10gb1/710 GB7Many small models, embedders, dev pods
1g.20gb new1/720 GB4Long-context 7B / KV-heavy small models

Up to 7 instances per H100. Mix profiles freely as long as the total fits (e.g. 3g + 2g + 1g + 1g).

Configuring MIG by hand

root @ host — nvidia-smi
# Enable MIG mode (requires no compute in flight; persistent across reboot)
nvidia-smi -mig 1

# Create GPU Instances by profile ID:  9 = 3g.40gb, 14 = 2g.20gb, 19 = 1g.10gb
nvidia-smi mig -cgi 9,14,19,19,19

# Create one Compute Instance per GI (default profile = whole GI)
nvidia-smi mig -cci

# See what you got
nvidia-smi -L
# GPU 0: NVIDIA H100 80GB HBM3 (UUID: GPU-...)
#   MIG 3g.40gb Device 0: (UUID: MIG-...)   ← tenant A sees this as $CUDA_VISIBLE_DEVICES
#   MIG 2g.20gb Device 1: (UUID: MIG-...)
#   MIG 1g.10gb Device 2: ...

Each MIG device shows up as a UUID that workloads pin via CUDA_VISIBLE_DEVICES=MIG-<uuid>. Tenants see what looks like a normal, smaller, dedicated GPU.

04

MIG Limits & Gotchas

MIG's hardware partitioning is its strength and its straitjacket. The same isolation that prevents noisy neighbours also prevents collaboration.

No NVLink between slices

Slices on the same physical GPU cannot use NVLink to talk to each other — NVLink connects whole GPUs, not partitions. MIG and tensor parallelism don't combine. MIG is for many small workloads, never one big sharded one.

UVM and IPC don't cross slices

CUDA Unified Memory and CUDA IPC handles assume a shared address space. Each MIG slice is its own context, on its own MMU. Cross-slice handles fail with cudaErrorInvalidDevice.

CUDA graphs work within a slice

Inside any single MIG instance the CUDA driver is normal: graphs, streams, async copies, MPS-on-MIG all behave. The boundary is the slice, not below it.

No live re-partitioning

Resizing a layout requires destroying every CI then every GI then re-creating. Workloads must drain. Plan re-partitions at maintenance windows; don't expect Kubernetes-style elastic resize.

Persistence mode strongly recommended

sudo nvidia-smi -pm 1. Without it the driver can unload between processes, blowing away MIG configuration on some kernel/driver combos. Datacenter cards default to on; double-check with nvidia-smi -q | grep Persistence.

Performance isolation, not fault isolation

MIG isolates contention but not failure. A driver crash, ECC double-bit error, or thermal trip affects the whole physical GPU. Plan your tenancy SLOs around "noisy neighbour solved" rather than "blast radius solved".

Mental model

Think of MIG as n GPUs in one socket, not as a CPU-style cgroup over one GPU. Anything that requires GPUs to cooperate (TP, NCCL all-reduce, P2P, UVM) needs distinct physical GPUs. Anything that just needs a private slice with predictable latency is what MIG is for.

05

MPS — Multi-Process Service

MPS is the older mechanism and is fundamentally different: rather than partition the hardware, it multiplexes CUDA contexts in software so multiple processes look like one to the GPU.

How it works

Client process A — CUDA app
Client process B — CUDA app
Client process C — CUDA app
↓
Unix socket: /tmp/nvidia-mps
↓
nvidia-cuda-mps-server (daemon)
↓
One CUDA context, one SM scheduler
↓
Physical GPU

Without MPS, the GPU time-slices at the context level: kernel from A finishes, context switch, kernel from B finishes, switch back. Context switches on a GPU are not free. MPS folds all clients into a single context so kernels from any client can co-execute on different SMs concurrently.

Wins

Loses

enabling MPS — per-host or per-MIG-slice
# Start the daemon (typically as a systemd unit)
export CUDA_VISIBLE_DEVICES=0
export CUDA_MPS_PIPE_DIRECTORY=/tmp/nvidia-mps
export CUDA_MPS_LOG_DIRECTORY=/var/log/nvidia-mps
nvidia-cuda-mps-control -d

# Cap a client at 30% of SMs (best-effort)
export CUDA_MPS_ACTIVE_THREAD_PERCENTAGE=30
./my_inference_server
Best fit

MPS shines when a single team runs many cooperating, latency-sensitive processes — e.g. a Triton instance with several model backends, a research fleet of small CUDA jobs, or many short-lived dev kernels. It is the wrong choice for unrelated tenants who don't trust each other.

06

MIG vs MPS — When to Use Which

The decision usually compresses to one question: do these workloads trust each other?

Use MIG when…

  • Production multi-tenant. Different teams, different blast radii, different bills.
  • Hard SLOs. p99 latency must not depend on someone else's batch.
  • Cost reporting per tenant. Slice = invoice line.
  • Predictable latency under load. No noisy-neighbour bandwidth steal.
  • Hardware available. A100 / H100 / H200 / B200 in the rack.

Use MPS when…

  • Many cooperating processes from one team.
  • Latency-sensitive small kernels — many models, low QPS each.
  • MIG-incapable GPU — V100, A40, RTX 4090, RTX A6000, L4, L40S.
  • Tight dev feedback loop — lots of short-lived experiments.
  • You can tolerate a single OOM taking out the whole party.

Use both, layered

The strongest pattern in production is MIG for tenant isolation, MPS within each MIG slice for many processes per tenant. Tenant A gets a 2g.20gb slice. Inside the slice, Triton runs MPS so its 12 model backends co-execute. Tenant B's 1g.10gb slice is similarly arranged but cannot perturb A's. You get hardware isolation between strangers and software co-execution among friends.

PropertyMIGMPSBoth
Memory isolationhardnonehard between slices
Bandwidth isolationhardcooperativehard between slices
Concurrent kernelsacross slices, not withinwithin client setboth axes
Cards supportedA100, A30, H100, H200, B200, RTX PRO 6000 Blackwellvirtually all CUDA GPUs (HW-accelerated since Volta)MIG-capable only
Licenseincludedincludedincluded
Best formulti-tenant SaaSone team, many procsmulti-tenant SaaS with rich tenants
07

vGPU — Virtualisation with NVIDIA AI Enterprise

vGPU is NVIDIA's enterprise virtualisation product, formerly branded GRID. It exposes one physical GPU to a hypervisor as N virtual GPUs, each of which a VM consumes as a normal device. Required where the isolation boundary is the VM, not the process.

Hypervisors supported

VMware vSphere

The most common production target. ESXi + Horizon for VDI; ESXi + AI Enterprise for GPU-backed VMs.

Microsoft Hyper-V

Windows Server platform; vGPU exposed as a DDA (Discrete Device Assignment) plus virtualisation extensions.

Citrix XenServer

Citrix's own hypervisor, the historical home of GRID for VDI.

Red Hat Virtualization

RHV / OpenShift Virtualization for KVM-based RHEL shops.

KVM (general)

Upstream KVM via mediated device (mdev) framework. The path most cloud providers customise.

Proxmox / others

Anything KVM-based can technically host vGPU if the host driver and licence are present.

Two backing modes

Time-sliced vGPU

The hypervisor's vGPU manager round-robins SM ownership between vGPUs. Multiple VMs share the same SMs over time. Memory is partitioned per vGPU but compute is time-multiplexed. Higher density, lower isolation.

MIG-backed vGPU

Each vGPU is bound to a MIG instance. The VM gets hardware-isolated SMs and HBM. Best isolation in the vGPU world. Available on A100, A30, H100, H200, B200.

Licensing & pricing

vGPU is licensed per concurrent vGPU, sold as named editions:

EditionUse caseList price/year
vAppsApp publishing — XenApp-style~$120
vPCKnowledge-worker VDI~$250
RTX Virtual Workstation (vWS)Engineering / CAD VDI~$500
NVIDIA AI EnterpriseCompute / AI VMs (the modern bundle)~$3,000 (often $4,500 list per GPU/yr)

Discounts in volume are deep; single-seat list prices rarely apply to enterprise deals.

Required for

Multi-tenant cloud DaaS (Citrix Cloud, AWS WorkSpaces, Azure Virtual Desktop with GPU SKUs); regulated environments — banking, healthcare, defence — where workloads must live inside hardware-virtualised VMs and the security model is "VM = trust boundary".

08

NVIDIA AI Enterprise & Licensing

The "NVIDIA AI Enterprise" (NVAIE) product is the 2026 packaging of vGPU plus a large software estate. If you want vGPU at all on compute cards, you are buying NVAIE.

What's bundled

Price shape

List price is roughly $4,500 per GPU per year, often steeply discounted in large hyperscaler or enterprise commitments. Multi-year terms (3 / 5 yr) and CSP perpetual licences exist.

PathCostSupportIncludes
Free CUDA toolkit$0community / forumCUDA, cuDNN, drivers
NVIDIA AI Enterprise~$4,500/GPU/yrenterprise SLAvGPU + NeMo + NIM + Triton + validated images + support
DGX Cloudper-instance, hourly or reservedNVIDIA-managedNVAIE bundled, hosted on NVIDIA infrastructure
Required for

vGPU at all (the host driver is licence-gated). Some hyperscaler L40S / H100 deployments require NVAIE under their contract with NVIDIA. The free CUDA toolkit is fine for bare-metal Kubernetes with MIG — but it is not a path to vGPU and carries no support if a driver bug lands in production.

09

Kubernetes — Time-Slicing & GPU Operator

Kubernetes' default device plugin model is unforgiving: one GPU = one pod. Anything finer-grained needs the NVIDIA GPU Operator, which manages drivers, plugins, MIG configuration and time-slicing as a coherent stack.

Default behaviour

Pod spec — the dumb path
resources:
  limits:
    nvidia.com/gpu: 1   # one whole GPU. Nothing else can claim it.

Time-slicing config

Configured in the GPU Operator's ClusterPolicy: replicas: N tells the device plugin to advertise the same GPU as N "shadows". N pods can claim it concurrently; the driver round-robins their kernel launches.

GPU Operator — ClusterPolicy snippet
apiVersion: nvidia.com/v1
kind: ClusterPolicy
spec:
  devicePlugin:
    config:
      name: time-slicing-config
  sharing:
    timeSlicing:
      resources:
        - name: nvidia.com/gpu
          replicas: 4      # 4 pods may share each GPU

Production multi-tenant — MIG-aware GPU Operator

For real isolation in K8s, have the GPU Operator configure MIG and surface each slice as its own resource type. Pods request slices like any other typed resource.

MIG-aware pod spec
resources:
  limits:
    nvidia.com/gpu.product: H100-PCIE-80GB-MIG-1g.10gb
    nvidia.com/gpu: 1

# Operator labels nodes with the available profiles per GPU,
# scheduler matches pod request to a free slice on a labelled node.
Mental model

Time-slicing is the K8s equivalent of MPS (no isolation, friendly for cooperating workloads). MIG-aware GPU Operator is the K8s equivalent of MIG (hardware partitioning, real tenants). Pick by trust boundary, not convenience.

10

Putting It Together — Worked Tenancy Patterns

Three real-world configurations spanning the trust spectrum.

(a) Inference factory

Hardware: 8× H100 PCIe nodes.
Partition: 7×1g.10gb per H100 (MIG).
Density: 56 inference tenants per node, 448 per cluster.
Stack: GPU Operator MIG-aware + Triton + autoscaled K8s deployments.
Workload: 56 fine-tuned 7B variants, each at ≤10 GB FP8 + KV.
Isolation: hardware between tenants, MPS optional inside the slice.

(b) Dev cluster

Hardware: 4× A40 nodes (no MIG on A40).
Partition: K8s time-slicing 4× per GPU.
Density: 16 dev users per node concurrently.
Stack: GPU Operator + JupyterHub + per-user notebook pods.
Workload: bursty notebook experiments, embedding jobs, small fine-tunes.
Isolation: none beyond container; works because workloads are bursty and trusted.

(c) VDI / engineering workstations

Hardware: vSphere host + 4× L40S 48 GB.
Partition: MIG-backed vGPU… on L40S not available, so time-sliced vGPU profiles (e.g. 1g.12gb per VM).
Density: 16 engineers, each with 16 vCPU + 12 GB vGPU VM.
Stack: ESXi + Horizon + NVAIE / vGPU licences.
Isolation: hypervisor-level VM isolation; licence counts per concurrent vGPU.

Heuristic for picking the right pattern

  1. What is the trust boundary? Process → MPS or time-slicing. VM → vGPU. Tenant → MIG.
  2. Is the GPU MIG-capable? If yes and trust boundary > process, default to MIG.
  3. Is the workload bursty and small? Time-slicing or MPS will pack better.
  4. Do you need a VM? Then vGPU; you're paying for NVAIE either way.
  5. Big sharded model? None of the above — that's TP across whole GPUs with NVLink.
A useful sanity check

If your sharing scheme can lose tenant data when one tenant misbehaves, you do not have multi-tenant isolation; you have a friend group with a shared GPU. That is fine for (b), unacceptable for (a) and (c).

11

Interactive: Sharing-Mode Picker

Pick a GPU class, a workload mix, and an isolation requirement. The picker recommends a sharing mode, estimates fan-out, flags licensing, and emits the right command or YAML snippet.

Mode
—
Max fan-out
—
Licence
—
Multi-GPU TP
—
snippet