Hardware · Architecture Series

RISC-V

An Open ISA · History, Variants, Capabilities & the Modern Ecosystem

29 Slides · Deep Dive · Interactive Encoder & Mini-CPU Stepper

00

Agenda

Origins & Philosophy

  • Why yet another ISA? The licensing & fragmentation problem
  • UC Berkeley, Patterson & Asanović (2010→)
  • The five pillars: open, clean, modular, scalable, stable
  • RISC-V International & the ratification process

Architecture Deep Dive

  • Base ISAs: RV32I / RV64I / RV32E / RV128I
  • Standard extensions: M · A · F · D · C · B · K · V · H
  • Profiles: RVA22 / RVA23 / RVB23
  • Privilege modes, CSRs, virtual memory, traps

Interactive

  • Instruction-format encoder — build a 32-bit word from fields
  • Mini-CPU stepper — watch registers update in real time
  • Live bit-field decomposition and calling-convention lookup

Ecosystem & Adoption

  • Vendor landscape & notable silicon
  • Software stack: GCC · LLVM · Linux · Zephyr
  • Case studies: NASA HPSC · WD · Tenstorrent · Meta MTIA
  • Custom extensions, CHERIoT, and what comes next
01

The ISA Landscape — Why Build Another One?

Openness / licensing freedom → Ecosystem maturity → x86/x64 Intel · AMD closed · patent-encumbered Arm licensable · royalties MIPS legacy → RISC-V (2021) OpenPOWER IBM · open 2019 RISC-V BSD-licensed spec zero royalties SPARC

The Problems RISC-V Solves

  • Licensing friction — Arm per-chip royalties (est. 1-2% ASP); architecture licenses cost $10M+ and years of negotiation
  • Legacy baggage — x86 carries 45 years of compatibility cruft; variable-length encoding, segment registers, real/protected/long modes
  • Closed standards — no published spec for x86; Arm spec behind NDA
  • Fragmentation — thousands of proprietary embedded ISAs (8051, PIC, AVR, Xtensa, ARC…)

What RISC-V Offers

  • Frozen base — RV32I/RV64I ratified & never changes; software built today runs forever
  • Extension mechanism — innovate in the margins without breaking the core
  • Scales from <10 kGE microcontrollers to out-of-order Linux-class superscalars
  • Royalty-free — CC-BY 4.0 spec, commercial use permitted, no membership required to build
02

A 15-Year Timeline

2010 2014 2015 2017 2019 2021 2023 2025 2026 ISA project begins UC Berkeley Par Lab / ASPIRE Frozen RV32I/64I User ISA v2.0 Foundation founded 25 founding members Priv. spec ratified M/S/U modes Move to Switzerland geopolitics-neutral Vector (V) & Hyp (H) RVV 1.0 ratified RVA23 profile App-class baseline Rocket Chip Chisel-generated first silicon tapeout SiFive founded first commercial IP HiFive1 ships Arduino-compatible Linux upstream kernel 4.15 (2018) Alibaba XuanTie C910 2.5 GHz OoO >10 B cores shipped cumulatively Ventana Veyron V2 3.6 GHz server-class Ubuntu / RHEL full RVA23 support From a 2010 summer research project to a multi-architecture platform in ~15 years
03

The Five Design Pillars

① Modular

  • Tiny frozen base (~40 ops for RV32I) + optional extensions
  • Implementers pick only what they need: an MCU may ship RV32IMC; a server, RV64GCHVSU
  • Avoids the x86 "carry everything forever" problem

② Clean Slate

  • No architectural status registers, flag registers, or delay slots
  • Fixed-length 32-bit encoding (16-bit with C); no mode switching
  • Single PC; load-store architecture; no implicit state

③ Open

  • Specification under CC-BY 4.0; no royalties, no membership required
  • Anyone can implement, sell, modify — commercial or not
  • Trademark "RISC-V" protects conformance, not the ISA itself

④ Scalable

  • Same encoding across XLEN = 32 / 64 / 128 bits
  • Scales from 10 kGE controllers to OoO superscalar Linux servers
  • Vector extension supports VLEN from 128 bits to 65,536 bits

⑤ Stable

  • Ratified extensions are frozen forever — software compiled today will always run
  • New work goes into new extensions, never breaking old ones
  • Formal ISA model (Sail) underpins ratification

Design Choices You Notice

  • No condition codes — compare-and-branch fused
  • 32 registers (16 for RV32E) — register pressure was a conscious trade
  • Little-endian only; big-endian optional via future Zbe
  • PC-relative addressing everywhere → relocatable code
04

Base ISAs — Four Flavours, Same DNA

BaseXLENRegistersAddress spaceTargetStatus
RV32I 32-bit x0-x31 · 32 bits 4 GiB MCUs, FPGA softcores, education Frozen 2014
RV64I 64-bit x0-x31 · 64 bits 16 EiB (virtual Sv39/48/57) Application processors, Linux, servers Frozen 2014
RV32E 32-bit x0-x15 (16 only) 4 GiB Ultra-area-constrained MCUs (<10 kGE) Ratified 2023
RV128I 128-bit x0-x31 · 128 bits 2128 bytes Future-proofing; cap-arch; capability pointers Draft / reserved

What "Base" Really Means

  • ~47 instructions: LOAD · STORE · OP · OP-IMM · BRANCH · JAL/JALR · LUI · AUIPC · SYSTEM · FENCE
  • Everything you need to run a compiler-emitted C program except integer mul/div, atomics, and FP
  • Deliberately missing: multiply, divide — provided by the M extension
  • Deliberately missing: atomic RMW — provided by the A extension

The RV64I Sign-Extension Trick

  • 32-bit integer ops on RV64 have .W suffix (ADDW, SLLW)
  • They sign-extend the 32-bit result to 64 bits — so negatives are naturally represented
  • No separate "32-bit mode" — both widths coexist in the same pipeline
  • Clever decision: avoids the x86 "mode switch" chaos while enabling efficient 32-bit code
05

Standard Extensions — The Alphabet

LetterNameAddsSize / CostStatus
MInteger Mul/DivMUL, MULH*, DIV, REM (8 ops)~2-5 k gatesRatified
AAtomicsLR/SC + AMOSWAP/ADD/AND/OR/XOR/MIN/MAX~1-2 k gatesRatified
FSingle-precision FP32 FP regs f0-f31, IEEE 754 binary32~15-30 k gatesRatified
DDouble-precision FPWidens f-regs to 64 bits; binary64 ops+30-60 k gatesRatified
QQuad-precision FPbinary128 ops (typically trapped & emulated)>100 k gatesRatified
CCompressed16-bit aliases for common 32-bit instructions (~25-30% size reduction)decoder onlyRatified
BBit manipulationZba+Zbb+Zbs (count, reverse, clz, bitfield)~5 k gatesRatified 2021
KScalar CryptoZk* — AES, SHA-2/3 round kernels~10-20 k gatesRatified 2022
VVector32 vector regs (VLEN), vector-length-agnostic ops100 k - 10 M gatesRatified 2021
HHypervisorVS-mode, 2-stage translation, virtual CSRsarea-dependentRatified 2021
ZicsrCSR accessCSRRW/CSRRS/CSRRC — implicit in any privileged coretinyRatified
ZifenceiInstr-fetch fenceFENCE.I — self-modifying code synchronisationtrivialRatified
ZicondConditional opsczero.eqz / czero.nez — branchless selectstrivialRatified 2023
Zicbom/ZicbozCache Blk Opscbo.clean / cbo.flush / cbo.inval / cbo.zeroCMO onlyRatified 2022

G = IMAFD + Zicsr + Zifencei — the "general-purpose" shorthand. A RV64GC core has everything Linux expects.

06

Application Profiles — Fighting Fragmentation

RV64I frozen base G I + M + A + F + D + Zicsr + Zifencei "general-purpose" RVA22U64 / RVA22S64 = G + C + B + Zihintpause + Zicbom/Zicboz + Svpbmt + Sstc… first Linux-grade baseline (2022) RVA23U64 / RVA23S64 = RVA22 + V + Vector Crypto + Hypervisor + AIA + more ratified 2024 — current target RVB23 (bespoke) Subset of RVA23 for MCUs no V, optional hypervisor

Why Profiles Exist

  • With hundreds of extensions, "RISC-V Linux" could mean anything
  • Profiles lock down a mandatory set for a given class of software — so a binary works on any conforming chip
  • U64 / S64 suffix: user-mode vs supervisor-mode baseline

RVA23 Highlights

  • Mandatory RVV 1.0 (min VLEN=128) — SIMDy baseline for Linux distros
  • Mandatory Hypervisor — KVM-class virtualisation
  • Mandatory AIA — modern interrupt controller (IMSIC/APLIC)
  • Cache-block management, PBMT page attributes, Sstc time CSR
  • Debian/Fedora/Ubuntu "riscv64" ports target RVA23 baseline from 2025
07

Register Model & Calling Convention

#ABI nameRoleSaver#ABI nameRoleSaver
x0zeroHard-wired to 0 (writes ignored)— x16-x17a6, a7Fn args / return 2+Caller
x1raReturn addressCaller x18-x27s2-s11Saved regsCallee
x2spStack pointer (16-byte aligned)Callee x28-x31t3-t6TemporariesCaller
x3gpGlobal pointer (linker-relative addr.)— f0-f7ft0-ft7FP temporariesCaller
x4tpThread pointer (TLS base)— f8-f9fs0-fs1FP savedCallee
x5-x7t0-t2TemporariesCaller f10-f11fa0-fa1FP args / returnCaller
x8s0/fpSaved / frame pointerCallee f12-f17fa2-fa7FP argsCaller
x9s1Saved registerCallee f18-f27fs2-fs11FP savedCallee
x10-x11a0, a1Fn args / return valuesCaller f28-f31ft8-ft11FP temporariesCaller
x12-x15a2-a5Fn argsCaller ABI: ilp32 (RV32), ilp32f/d (with F/D), lp64, lp64d (RV64). Tail calls use t1.

Why x0 = 0 is Genius

  • MV rd, rs → ADDI rd, rs, 0
  • NOP → ADDI x0, x0, 0
  • J label → JAL x0, label (discard link)
  • Unconditional zero-test: BEQ rs, x0, target

No Condition Codes

  • No flags register — all comparisons are explicit
  • Compare-and-branch fused: BEQ, BNE, BLT, BGE, BLTU, BGEU
  • Sets (SLT, SLTU, SLTI) write 1 or 0 to a register
  • Good for OoO — no implicit register renaming of flags
08

Instruction Formats — Six Shapes, One Decoder

31 25 20 15 12 7 0 funct7 rs2 rs1 funct3 rd opcode R-type ADD, SUB, AND, OR, SLL, SLT… imm[11:0] rs1 funct3 rd opcode I-type ADDI, LW, JALR, CSRRW… imm[11:5] rs2 rs1 funct3 imm[4:0] opcode S-type SW, SH, SB im12 imm[10:5] rs2 rs1 funct3 imm[4:1] im11 opcode B-type BEQ, BNE, BLT, BGE… imm[31:12] rd opcode U-type LUI, AUIPC im20 imm[10:1] im11 imm[19:12] rd opcode J-type JAL Key invariants across all formats: • rs1, rs2, rd always land in the same bit positions — register read can start before decode completes • Sign bit (bit 31) is the immediate's MSB in I/S/B/U/J formats — one sign-extend mux does them all • Bit scrambling in B/J formats lets branches/jumps reuse as much of the immediate datapath as possible
Interactive Demo · Try me
09

Live Instruction Encoder

Pick an instruction and set its fields — watch the 32-bit encoding and bit-field layout update live.

Controls

x10 (a0)
x11 (a1)
x12 (a2)
42

Assembly

ADDI a0, a1, 42

32-bit Encoding

0x02A58513
Binary (MSB → LSB):
0000001 01010 01011 000 01010 0010011

Bit-field decomposition

10

Privilege Architecture

M Machine · 0b11 HS Hypervisor · 0b01+H S Supervisor · 0b01 U User · 0b00 VS VU guest modes (H ext.) MCU (embedded): M only RTOS: M + U Linux: M + S + U KVM host: M + HS + VS + VU Rings of privilege — any combination is legal

M — Machine Mode

  • Highest privilege; has direct access to all CSRs and memory
  • Holds the boot code, SBI firmware, and platform-specific trap handlers
  • Every RISC-V core has M-mode — it's the only mandatory level
  • Key CSRs: mstatus, mepc, mtvec, mcause, mtval, mie, mip

S — Supervisor Mode

  • Where OS kernels live (Linux, FreeBSD, seL4, Zephyr-MMU)
  • Can manage virtual memory via satp, handle S-mode traps
  • SBI (Supervisor Binary Interface) is the ABI between S and M modes — ecall traps up

U — User Mode

  • Application code runs here
  • Trap to S-mode via ecall (syscall) or page faults
  • Limited to CSRs with URW/URO bits (time, cycle, instret counters)
11

Virtual Memory — Sv32 / Sv39 / Sv48 / Sv57

Sv39 — 3-level, 39-bit VA → 56-bit PA VPN[2] (9b) VPN[1] (9b) VPN[0] (9b) Page Offset (12b) 39-bit Virtual Address satp.PPN root page-table L2 PT PTE ... 512 PTEs L1 PT L0 PT 4 KiB page Physical Memory PA = (PTE.PPN << 12) | offset Superpages: 4 KiB · 2 MiB · 1 GiB PTE bits: V · R · W · X · U · G · A · D
ModeVA bitsLevelsPage sizesAddress space
Sv323224 KiB · 4 MiB4 GiB
Sv393934 KiB · 2 MiB · 1 GiB512 GiB
Sv48484+ 512 GiB256 TiB
Sv57575+ 256 TiB128 PiB

Key Mechanisms

  • satp CSR — selects mode (Bare/Sv32/39/48/57), ASID, and root PPN
  • SFENCE.VMA [rs1],[rs2] — surgical TLB invalidation by VA/ASID
  • Each PTE: V R W X U G A D + reserved + PPN
  • A (accessed) and D (dirty) can be HW- or SW-managed (Svadu)
  • Svpbmt — page-based memory types (NC, IO)
  • Svnapot — naturally-aligned power-of-2 contiguous regions (cheap huge pages)
12

Traps, Exceptions & Interrupts

User / S-mode code trap! Causes (mcause) Sync exceptions: 0 Instruction misaligned 2 Illegal instruction 3 Breakpoint 5 Load access fault 7 Store access fault 8 ECall from U 9 ECall from S 12 Instr page fault 13/15 Load/Store page fault + Async: SI/TI/EI per mode HW automatic • pc → mepc • cause → mcause • trap val → mtval • priv ← M pc ← mtvec (vectored or direct) SW saves caller-saved regs, demux on mcause, handle, then MRET MRET

Interrupt Controllers

  • CLINT — basic timer + software interrupts; in every core since spec v1.9
  • PLIC — Platform-Level Interrupt Controller; up to 1023 sources, 15 priority levels, per-hart enables/thresholds/claim-complete
  • CLIC — embedded: pre-emptive, vectored, stack-based priorities (<1 µs latency)
  • AIA — modern: IMSIC (per-hart MSI file) + APLIC (wired irq translator) — mandatory in RVA23

Trap Delegation

  • medeleg/mideleg let M-mode forward specific traps directly to S-mode — zero-cost syscalls
  • hedeleg/hideleg delegate VS traps to HS
  • WFI instruction — wait for interrupt; implementations may halt the clock
  • Reset vector is implementation-defined; typical: 0x80000000 or 0x1000
13

M & A — Multiply/Divide and Atomics

M — Integer Multiply / Divide

  • MUL rd ← rs1·rs2 (low XLEN)
  • MULH / MULHSU / MULHU — upper XLEN, signed/unsigned variants
  • DIV / DIVU / REM / REMU
  • RV64 adds MULW, DIVW, DIVUW, REMW, REMUW — 32-bit × sign-ext-to-64
  • Divide-by-zero returns all-ones for quotient, dividend for remainder — no trap
  • Typical implementations: 1–3 cycle Booth multiplier; multi-cycle sequential divider (SRT radix-4 common)
# sum-of-squares with M ext.
mul  t0, a0, a0
mul  t1, a1, a1
add  a0, t0, t1
ret

A — Atomic Memory Operations

  • LR.W/D + SC.W/D — load-reserved / store-conditional pair (forward-progress guarantee for restricted sequences)
  • AMOSWAP, AMOADD, AMOAND, AMOOR, AMOXOR, AMOMIN, AMOMAX, AMOMINU, AMOMAXU
  • aq/rl acquire/release bits — RCsc memory model hooks
  • Basis for every lock-free & lock-based primitive in Linux, pthreads, Rust atomics
  • Zaamo split-out (AMOs only); Zalrsc (LR/SC only)
# lock-free inc via LR/SC
1:  lr.w    t0, (a0)
    addi    t1, t0, 1
    sc.w    t2, t1, (a0)
    bnez    t2, 1b           # retry on contention

Memory Model

RVWMO (Weak Memory Ordering) is the default — similar to Arm, weaker than x86 TSO. Fences: FENCE [pred],[succ] where each field is a mask of i o r w. A Ztso option flips the core to Total Store Order (for easier x86 porting; Alibaba T-Head uses it).

14

Floating-Point — F, D, Q, Zfh, Zfinx

The Classic Extensions

  • F — 32 separate f0-f31 registers, 32-bit wide (binary32)
  • D — widens to 64-bit registers, adds binary64 ops (IEEE 754-2008)
  • Q — 128-bit registers, binary128 (rare; usually trap-and-emulate)
  • Fused multiply-add: FMADD, FMSUB, FNMADD, FNMSUB
  • fcsr — 8-bit CSR: 5 sticky exception flags + 3-bit rounding mode (RNE, RTZ, RDN, RUP, RMM, DYN)
  • Conversions: FCVT.{W/L/S/D}.{S/D/W/L}[U] — every pair-wise combo
  • Classify: FCLASS returns a 10-bit bitmask (-inf, -norm, -subn, -0, +0, +subn, +norm, +inf, sNaN, qNaN)

New & Specialised Flavours

  • Zfh — half-precision (binary16) scalar FP; useful for AI quant workloads
  • Zfhmin — half-precision conversions only (no arithmetic); minimum-cost support
  • Zfinx / Zdinx / Zhinx — FP "in x-registers": no separate f-regfile; saves ~30% area in MCUs
  • Zfbfmin / Zvfbfwma — BF16 support, mandated in RVA23 for ML
  • Zfa — additional FP ops: fli (load immediate), fminm/fmaxm, fround
# dot product fragment (D ext.)
  fld     ft0, 0(a0)
  fld     ft1, 0(a1)
  fmadd.d fa0, ft0, ft1, fa0
  addi    a0, a0, 8
  addi    a1, a1, 8
  bne     a0, a2, 1b
15

C — Compressed 16-bit Encodings

32-bit: ADDI sp, sp, -16 imm[11:0] = -16 x2 000 x2 OP-IMM 16-bit: C.ADDI16SP sp, -16 011 nzi5 x2 nzi[4,6,8:7,5] 01 Typical code-size reduction: 25-30% • Reuses the most common register subset x8-x15 with a 3-bit field • 16-bit and 32-bit mix freely on 2-byte alignment (bits [1:0] = 11 → 32-bit, else 16-bit) • Decoder expands each C-form to its 32-bit base instruction — zero semantic difference • Zca/Zcf/Zcd — sub-extensions for fine-grained inclusion • Zcb / Zcmp / Zcmt — "C extension extras" for further MCU code density

Popular Forms

  • C.LI — load 6-bit signed immediate
  • C.LW / C.SW — use 3-bit reg subset
  • C.ADDI, C.MV, C.ADD — 16-bit ALU ops
  • C.JR, C.JALR — 16-bit branches
  • C.BNEZ, C.BEQZ — compare-to-zero branches
  • C.SLLI, C.SRLI, C.SRAI — shifts

Why Not Thumb-2?

  • Thumb-2 is a separate ISA with mode switches (BX)
  • C extends the same ISA — no mode, PC simply advances by 2 or 4
  • Every 16-bit form has an exact 32-bit expansion, so macro-op fusion / decode is trivial
  • Enables ~100 bytes/KiB denser code than Thumb-2 in typical benchmarks (Embench)
16

Bit Manipulation (B) & Scalar Crypto (K)

B — Bitmanip (ratified Nov 2021)

  • Zba — address generation: sh1add, sh2add, sh3add, add.uw
  • Zbb — basics: clz, ctz, cpop, min/max, sext.b/h, rev8, rol/ror, orc.b, andn/orn/xnor
  • Zbc — carry-less multiply: clmul, clmulr, clmulh (CRC accelerator)
  • Zbs — single-bit: bclr, bset, binv, bext (+ immediate forms)
  • Zbkb/Zbkc/Zbkx — crypto-oriented subsets used by K
# count trailing zeros
ctz    a0, a0        # was: ~10 insns software loop

K — Scalar Cryptography (ratified 2022)

  • Zknd / Zkne — AES-128/192/256 encrypt / decrypt round helpers (aes64es, aes64ds, aes64ks1i, aes64ks2)
  • Zknh — SHA-2 (224/256/384/512) and SHA-3 Keccak round functions
  • Zksed / Zksh — SM4 and SM3 (Chinese crypto standards)
  • Zkr — entropy source CSR (seed) — NIST SP 800-90B conformant TRNG interface
  • Zkt — constant-time guarantee profile — spec promise that listed insns are data-independent latency
# AES round with K ext. (4 insns/round
# vs ~16 with tables — and sidechannel safe)
aes64es   a0, a0, a1    # encrypt + sbox
aes64es   a0, a0, a2

Why This Matters

OpenSSL, BoringSSL, and the Linux kernel crypto API all have RISC-V vector-and-scalar-crypto backends (2024–). Benchmark: AES-GCM on a SiFive U74 with Zkne+Zbkb runs ~8× faster than the pure-C path. Zkt ratification also unblocked FIPS 140-3 certification pathways for RISC-V silicon.

17

V — The Vector Extension (RVV 1.0)

v0–v31 · VLEN bits wide SEW=8 e0 e1 e2 … VL elements of SEW bits SEW=32 fewer, wider lanes LMUL = 4 (group 4 vregs) v8 v9 v10 v11 viewed as one large vector → more throughput, fewer addressable regs Segment load/store: vlseg2e32.v v4, (a0) → {v4,v5} strided by 2 elts Indexed: vluxei32.v — gather; vsuxei32.v — scatter

Vector-Length Agnostic (VLA)

  • Code does not hard-code width; hardware reports VLEN at runtime
  • Profile enforces VLEN ≥ 128; implementations today: 128–2048 bits; research: up to 65,536
  • Same binary runs on a 128-bit embedded SoC or a 2048-bit server chip

The vsetvl Pattern

  • vsetvli t0, a0, e32, m4 — "give me up to a0 lanes of 32-bit ints, 4-reg group"
  • HW returns how many elements fit → put in t0 and vl CSR
  • Natural "stripmining": loop handles any array length without epilogue
axpy:              # y = a*x + y
  vsetvli t0, a3, e32, m8
  vle32.v  v0, (a1)
  vle32.v  v8, (a2)
  vfmacc.vf v8, fa0, v0
  vse32.v  v8, (a2)
  sub      a3, a3, t0
  slli     t0, t0, 2
  add      a1, a1, t0
  add      a2, a2, t0
  bnez     a3, axpy

Sub-extensions: Zve32x/Zve32f/Zve64d (embedded vector tiers) · Zvkn, Zvkng (vector crypto: AES, SHA, GHash) · Zvfh (FP16) · Zvbb/Zvbc (vector bitmanip) · Zvknha/Zvksed (hash/SM crypto)

18

H — Hypervisor & AIA

Guest VM #1 VU app VS kernel Stage-1 VM page tables (vsatp) Guest VM #2 VU app VS kernel Stage-1 VM page tables (vsatp) HS-mode Hypervisor — KVM, Xvisor, bao-hypervisor, seL4-hyp Stage-2 translation via hgatp · controls vCPU scheduling, virtual timer & interrupts (AIA) M-mode firmware — OpenSBI / RustSBI · platform init, SBI calls

H Extension (ratified 2021)

  • Adds VS and VU modes — "virtualised" flavours of S and U
  • Two-stage address translation: VA → GPA → PA
  • New CSRs: hgatp, hstatus, hedeleg, hideleg, hcounteren, vsstatus, vsepc, vsscratch…
  • Hypervisor load/store: HLV.{B,H,W,D}, HSV.{…} — access guest memory with guest permissions
  • HFENCE.VVMA, HFENCE.GVMA — stage-1 / stage-2 TLB flush

AIA — Advanced Interrupt Arch

  • Replaces legacy PLIC for virtualisation
  • IMSIC — per-hart MSI-capable interrupt file (PCIe-native)
  • APLIC — translates wired interrupts into MSIs
  • Directly supports guest-external interrupts (no hyp-trap on every delivery)
  • Mandatory in RVA23 — fundamental for cloud workloads
19

Custom Extensions & CHERIoT

Custom Opcode Space

  • RISC-V reserves custom-0 (0x0B), custom-1 (0x2B), custom-2 (0x5B), custom-3 (0x7B) for vendor use
  • No approval required — vendors innovate; if useful, propose to RISC-V International for standardisation
  • Intel AVX-style "add AVX2 to the world" problem eliminated: new encoding space is reserved, never collides
  • Examples: SiFive Xsfvcp coprocessor interface, Andes ACE, T-Head vector subset, Tenstorrent DSP ops
  • GNU & LLVM toolchains support .insn directive and inline intrinsics for custom ops

CHERIoT / CHERI-RISC-V

  • Capability-based pointers — every pointer is a 128-bit (CHERI) or 64-bit (CHERIoT) sealed capability carrying bounds + permissions
  • Deterministic spatial safety (no buffer overflows) and temporal safety (no UAF with revocation)
  • CHERI-RISC-V was the template for Arm Morello research CPU
  • CHERIoT — MSR-Cambridge, lowRISC, SCI Semiconductor — open Ibex-based CHERI CPU for embedded/IoT security
  • Ongoing merger path via Zicfilp / Zicfiss (landing-pad + shadow-stack CFI), Smmpm (Memory-protection MMU)

Standardisation Pipeline

Draft → Public Review → Frozen → Ratified. Once Ratified, an extension's bits are fixed forever. The process is run by RISC-V International Technical Steering Committee; all specs live in GitHub riscv/riscv-isa-manual, riscv/riscv-*-ext. Each extension has an associated Sail model, a conformance test suite (riscv-arch-test), and a compliance signature mechanism.

Interactive Demo · Step the CPU
20

Mini RISC-V Stepper — Fibonacci

A minimal RV32I model. Click Step to execute one instruction. Changed registers highlight in amber. The program computes fib(N) in a0 where N starts in a1.

Program (starts at 0x1000)


      
6

CPU State

pc = 0x00001000   cycle = 0
Next instr:
0x00058863 BEQZ a1, done

Register File

21

RISC-V International — Governance

Organisation

  • Founded 2015 as RISC-V Foundation (US non-profit)
  • 2019: relocated to Zug, Switzerland as RISC-V International — explicitly geopolitics-neutral; insulates members from US export-control risks
  • ~4000+ members in 70+ countries (2025): includes NVIDIA, Intel, Qualcomm, Samsung, Google, Meta, Huawei, Alibaba, ARM(!)
  • Membership tiers: Premier, Strategic, Community · CC-BY spec means a membership is not required to implement

Technical Process

  • ~40+ Task Groups; each produces a draft → public review → freeze → ratification
  • Every ratified extension has: prose spec, Sail formal model, architectural tests
  • ACTs (Architectural Compatibility Tests) must pass for a chip to bear the "RISC-V" trademark
  • Profile and Platform specs coordinate vendor convergence around a common target

2025 Strategic Priorities

  • RVA23 ecosystem rollout — Debian, Fedora, Ubuntu all targeting it by default in 2025-26
  • Automotive (ISO 26262 certified cores from SiFive, Andes, Synopsys)
  • Data-centre: Ventana Veyron V2/V3, Tenstorrent Ascalon, SiFive P870-D
  • AI: RISC-V-based NPUs (Meta MTIA, Tenstorrent Wormhole, Esperanto ET-SoC-1)
  • Open-source silicon pipeline: OpenROAD, OpenMPW, Google+eFabless Sky130/Gf180 shuttle
22

Vendor Landscape

Commercial IP

  • SiFive (US) — E2/E7, U74, P550, P670, P870, P870-D; Linux-class and server-class
  • Andes (Taiwan) — N/D/A series incl. NX27V vector, AX45MP, QiLai multicore
  • Codasip (CZ) — customisable cores; L110, A730, X730
  • Synopsys ARC-V RMX/RHX; MIPS eVocore (post-2021 pivot)
  • Imagination Catapult MK/AX series
  • Nuclei (CN) N/N9/NA series

In-house / Fabless Silicon

  • Alibaba T-Head — XuanTie C906, C910, C920 (Linux SBCs LicheePi)
  • Espressif — ESP32-C/H/P/S family (RISC-V replaces Xtensa)
  • Microchip — PolarFire SoC (4×U54 + E51)
  • Tenstorrent — Ascalon cores in Blackhole, Wormhole AI chips; licences to LG, Japan
  • Ventana — Veyron V1/V2 chiplet server CPUs
  • Rivos, Esperanto, Condor Computing, InspireSemi

Open-Source Cores

  • Rocket — UCB reference; in-order, Scala/Chisel, highly parameterisable
  • BOOM — Berkeley Out-of-Order Machine; OoO superscalar
  • CVA6 / Ariane — ETH/OpenHW 6-stage 64-bit core; automotive-cert track
  • Ibex / CVE2 — lowRISC MCU core; base for CHERIoT
  • VexRiscv — SpinalHDL, plugin-modular, FPGA sweet-spot
  • PicoRV32, NEORV32, SweRV-EH/EL — tiny/medium embedded flavours
23

Notable Implementations — From µ to Server

CoreVendorµarchTargetClockExtensions
PicoRV32Yosys/Claire Wolfmulti-cycle, <1 kLEFPGA glue logic~100 MHzRV32IMC
IbexlowRISC2-stage in-orderOpenTitan security chip~500 MHz ASICRV32IMC + B + crypto
E31 / E51SiFive5-stage in-orderMCU / mgmt cores1 GHz (16 nm)RV32IMC / RV64IMAC
CVA6OpenHW Group6-stage in-order, 64-bitLinux-capable SoCs1 GHz ASICRV64GC + H
U74SiFive8-stage dual-issueHiFive Unmatched, VisionFive 21.2 GHz (28 nm)RV64GCB
C910Alibaba T-Head3-issue OoO, 12-stageLicheePi 4A, BeagleV2.5 GHz (12 nm)RV64GCBV (pre-1.0)
P670 / P870SiFive3/4-issue OoO + vectorMobile / auto SoCs2.6-3.6 GHzRV64GCBVH (RVA23)
Veyron V2Ventana8-issue OoO serverChiplet data-centre CPU3.6 GHz (5 nm)RV64GCBVH + custom
AscalonTenstorrent8-wide OoO + vectorAI chiplets; Samsung Foundry3+ GHz (3 nm)RV64GCBVH + custom
BOOM v3UCB / Chipyard4-issue OoO, open sourceResearch / silicon demos~1 GHz (28 nm)RV64GC (+ optional V)
SG2042Sophgo64× C910 clusterServer / workstation SBC2 GHz (12 nm)RV64GCB (+V 0.7)
ET-SoC-1Esperanto1088× ET-Minion + 4× ET-MaxionML inference accel.1 GHz (7 nm)RV64GCV + custom tensor
24

Software Ecosystem

Compilers & Toolchains

  • GCC — first-class support since 7.x; -march=rv64gc_zba_zbb_zbs_v style machine strings
  • LLVM / Clang — tier-1 since LLVM 15; RISC-V backend maintained by Igalia, SiFive, Rivos
  • Rust — riscv64gc-unknown-linux-gnu tier-1 from 2024; riscv32imac-* for bare-metal Rust
  • Go — GOARCH=riscv64 since Go 1.14; production-ready since 1.21
  • Binutils, GDB, QEMU, Spike (reference ISS)

Operating Systems

  • Linux — upstream RV64 port since 4.15 (2018); RV32 port 5.18; systemd, KVM, eBPF all supported
  • Distros: Debian, Ubuntu, Fedora, openSUSE, Gentoo, Arch-RISCV, Alpine, AOSP — all provide riscv64 ports
  • FreeBSD 13+, NetBSD, OpenBSD (2023)
  • Embedded RTOS: Zephyr, FreeRTOS, RT-Thread, NuttX, ThreadX, seL4, PX5
  • Firmware: OpenSBI (M-mode), RustSBI, U-Boot, EDK II (UEFI), coreboot/LinuxBoot

Simulation & Verification

  • Spike (riscv-isa-sim) — golden reference C++ ISS
  • Sail — formal executable ISA model (authoritative)
  • riscv-arch-test — official compliance suite
  • RISCOF — Python framework wiring Spike ↔ DUT for signature-compare tests
  • QEMU tcg+kvm; Renesas Renode; AndeSight, SiFive Studio
  • Formal: Axiomatic Alive-RV, ESOP memory model, riscv-semantics (K framework)

Libraries & Runtimes

  • glibc / musl / newlib — all have riscv ports
  • OpenSSL / BoringSSL / Botan — vector + scalar crypto backends
  • LLVM libc++, libstdc++ — tier-1
  • PyTorch, TensorFlow, JAX — run on RVA22+ via CPU backends; vector paths landing 2025
  • CUDA-equivalent: no direct equivalent; Tenstorrent uses TT-Metalium, Esperanto has its own stack
  • WebAssembly: V8, SpiderMonkey, Wasmtime all have RV64 JIT backends
25

Adoption Case Studies

NASA HPSC (2024→)

  • NASA's next-gen spaceflight computer — contract awarded to Microchip / SiFive
  • 8× SiFive X280 RISC-V cores (rad-hardened), targets 100× the perf of current RAD750
  • Expected flight hardware: Artemis follow-ons, Mars sample return
  • Drove adoption of riscv-abi rad-tolerant compilation flows

Western Digital (2018→)

  • Pledged to ship >2 B RISC-V cores/year — inside every storage controller
  • Open-sourced SweRV EH1/EH2/EL2 — in-order 32-bit MCUs
  • Replaced proprietary microcontrollers in NVMe SSDs & HDDs
  • Proof point: the boring, huge-volume embedded market adopts first

Tenstorrent Wormhole / Blackhole

  • Ascalon 8-wide OoO RISC-V cores drive & surround the Tensix AI tiles
  • LG, Samsung Foundry, and Hyundai partnered for automotive variant
  • Rivos and Jim Keller's team → first full RISC-V training-class silicon (2025)

Meta MTIA v2 (2024)

  • Meta's in-house AI inference accelerator — RISC-V cores orchestrate Tensor tiles
  • Replaces the control-plane Arm cores used in MTIA v1
  • Citation: Meta's 2024 OCP talk — RISC-V chosen for custom-extension freedom & cost
  • Similar story: Google Titan M2 (security chip), Apple rumoured use in coprocessors

Espressif

ESP32-C3/C6/H2/P4 — replaces Xtensa in hundreds of millions of Wi-Fi/BLE IoT parts.

NVIDIA

GPU System Processor (GSP) runs RISC-V NV-RISCV in every GeForce/Blackwell GPU since 2017.

Seagate

Custom high-perf RISC-V core drives Mozaic 3+ HAMR hard drives; open sourced as "Seagate Core".

Google

RVV crypto work upstreamed to BoringSSL; internal Pixel / data-centre exploration with Rivos.

26

Future Directions

ISA Extensions In-Flight

  • RAS — Reliability/Availability/Serviceability (Zicntr, Zihpm, RAS signalling)
  • Zicfilp / Zicfiss — CFI: landing-pad & shadow stack
  • Zvk* — vector crypto: ratified 2023, rolling into vendor silicon
  • Smmpt, Smepmp — enhanced PMP / memory-protection
  • RVP — packed SIMD (DSP) for MCUs that can't afford V
  • Zacas — double-word atomic compare-and-swap

Performance Frontier

  • 8-wide OoO with >256-entry ROB → SPEC2017 parity with Neoverse V2 expected by 2026
  • Big-vector (VLEN=2048+) for HPC / LLM inference
  • Decoupled vector / matrix co-processor interface (RVI-AME "matrix" extension, draft)
  • Chiplet + UCIe-based compositions — Ventana, Tenstorrent
  • Coherent accelerator interfaces via CHI-C2C and CXL

Safety & Security

  • CHERI-RISC-V / CHERIoT convergence paths into standard
  • Confidential Compute — AP-TEE / CoVE (Confidential VM Extension)
  • ISO 26262 ASIL-D certified cores (Codasip, Andes, SiFive Shield)
  • DO-254 / DAL-A for avionics (SiFive Intelligence X)
  • Post-quantum crypto intrinsics (ML-KEM, ML-DSA accelerators)

Open Questions

  • Fragmentation risk — with profiles strictly enforced, can RVA23 become the de-facto "RISC-V Linux" target before vendors diverge too far with custom extensions?
  • Software tail — SIMD libraries (SIMDe, highway, xsimd) backfilling RVV paths; AI framework vector lowering still immature vs CUDA/ROCm
  • Geopolitics — US export-control impacts on Premier-tier membership; RVI's Swiss domicile helps but doesn't immunise silicon
  • Matrix-multiply ISA standardisation — will vendors wait, or ship proprietary extensions and negotiate later?
27

Summary — What Makes RISC-V Different

The 30-Second Pitch

  • A royalty-free, modular, frozen-base ISA that scales from 10 kGE MCUs to 8-wide OoO server cores
  • Five design pillars: open · clean · modular · scalable · stable
  • Rich extension alphabet (M A F D C B K V H) + custom opcode spaces
  • Real hardware shipping: >10 billion cores/year across every tier of silicon

Why It Won Ground

  • Software stack (Linux, GCC, LLVM, Rust, Go) was upstream-mature before big vendors arrived
  • Profiles prevent ABI-level fragmentation across vendors
  • Formal ISA model (Sail) and official compliance suite keep implementations honest
  • Geopolitical neutrality — Swiss domicile shields members from unilateral export controls

Where To Next

  • RVA23 rollout in Linux distros 2025-26; Wikipedia-style everywhere adoption
  • First truly competitive RISC-V server CPU vs Graviton / Neoverse in 2026
  • RISC-V AI accelerators replace Arm control-plane cores (Meta MTIA pattern)
  • Safety-critical/avionics certified cores unlocks automotive and aerospace

Further Reading

  • riscv.org/specifications/ratified — canonical spec index
  • github.com/riscv/riscv-isa-manual — source of truth
  • Patterson & Waterman, The RISC-V Reader: An Open Architecture Atlas
  • Waterman PhD thesis (2016) — design rationale
  • RVA23 profile specification (2024)
  • Sail model: github.com/riscv/sail-riscv

"An Instruction Set Should Be Free" — Patterson & Asanović, 2014