Hardware · Architecture Series
RISC-V
An Open ISA · History, Variants, Capabilities & the Modern Ecosystem
29 Slides · Deep Dive · Interactive Encoder & Mini-CPU Stepper
00
Agenda
Origins & Philosophy
Why yet another ISA? The licensing & fragmentation problem
UC Berkeley, Patterson & Asanović (2010→)
The five pillars: open, clean, modular, scalable, stable
RISC-V International & the ratification process
Architecture Deep Dive
Base ISAs: RV32I / RV64I / RV32E / RV128I
Standard extensions: M · A · F · D · C · B · K · V · H
Profiles: RVA22 / RVA23 / RVB23
Privilege modes, CSRs, virtual memory, traps
Interactive
Instruction-format encoder — build a 32-bit word from fields
Mini-CPU stepper — watch registers update in real time
Live bit-field decomposition and calling-convention lookup
Ecosystem & Adoption
Vendor landscape & notable silicon
Software stack: GCC · LLVM · Linux · Zephyr
Case studies: NASA HPSC · WD · Tenstorrent · Meta MTIA
Custom extensions, CHERIoT, and what comes next
01
The ISA Landscape — Why Build Another One?
Openness / licensing freedom →
Ecosystem maturity →
x86/x64
Intel · AMD
closed · patent-encumbered
Arm
licensable · royalties
MIPS
legacy → RISC-V (2021)
OpenPOWER
IBM · open 2019
RISC-V
BSD-licensed spec
zero royalties
SPARC
The Problems RISC-V Solves
Licensing friction — Arm per-chip royalties (est. 1-2% ASP); architecture licenses cost $10M+ and years of negotiation
Legacy baggage — x86 carries 45 years of compatibility cruft; variable-length encoding, segment registers, real/protected/long modes
Closed standards — no published spec for x86; Arm spec behind NDA
Fragmentation — thousands of proprietary embedded ISAs (8051, PIC, AVR, Xtensa, ARC…)
What RISC-V Offers
Frozen base — RV32I/RV64I ratified & never changes; software built today runs forever
Extension mechanism — innovate in the margins without breaking the core
Scales from <10 kGE microcontrollers to out-of-order Linux-class superscalars
Royalty-free — CC-BY 4.0 spec, commercial use permitted, no membership required to build
02
A 15-Year Timeline
2010
2014
2015
2017
2019
2021
2023
2025
2026
ISA project begins
UC Berkeley
Par Lab / ASPIRE
Frozen RV32I/64I
User ISA v2.0
Foundation founded
25 founding members
Priv. spec ratified
M/S/U modes
Move to Switzerland
geopolitics-neutral
Vector (V) & Hyp (H)
RVV 1.0 ratified
RVA23 profile
App-class baseline
Rocket Chip
Chisel-generated
first silicon tapeout
SiFive founded
first commercial IP
HiFive1 ships
Arduino-compatible
Linux upstream
kernel 4.15 (2018)
Alibaba XuanTie
C910 2.5 GHz OoO
>10 B cores
shipped cumulatively
Ventana Veyron V2
3.6 GHz server-class
Ubuntu / RHEL
full RVA23 support
From a 2010 summer research project to a multi-architecture platform in ~15 years
03
The Five Design Pillars
① Modular
Tiny frozen base (~40 ops for RV32I) + optional extensions
Implementers pick only what they need: an MCU may ship RV32IMC ; a server, RV64GCHVSU
Avoids the x86 "carry everything forever" problem
② Clean Slate
No architectural status registers, flag registers, or delay slots
Fixed-length 32-bit encoding (16-bit with C); no mode switching
Single PC; load-store architecture; no implicit state
③ Open
Specification under CC-BY 4.0 ; no royalties, no membership required
Anyone can implement, sell, modify — commercial or not
Trademark "RISC-V" protects conformance, not the ISA itself
④ Scalable
Same encoding across XLEN = 32 / 64 / 128 bits
Scales from 10 kGE controllers to OoO superscalar Linux servers
Vector extension supports VLEN from 128 bits to 65,536 bits
⑤ Stable
Ratified extensions are frozen forever — software compiled today will always run
New work goes into new extensions, never breaking old ones
Formal ISA model (Sail) underpins ratification
Design Choices You Notice
No condition codes — compare-and-branch fused
32 registers (16 for RV32E) — register pressure was a conscious trade
Little-endian only; big-endian optional via future Zbe
PC-relative addressing everywhere → relocatable code
04
Base ISAs — Four Flavours, Same DNA
Base XLEN Registers Address space Target Status
RV32I
32-bit
x0-x31 · 32 bits
4 GiB
MCUs, FPGA softcores, education
Frozen 2014
RV64I
64-bit
x0-x31 · 64 bits
16 EiB (virtual Sv39/48/57)
Application processors, Linux, servers
Frozen 2014
RV32E
32-bit
x0-x15 (16 only)
4 GiB
Ultra-area-constrained MCUs (<10 kGE)
Ratified 2023
RV128I
128-bit
x0-x31 · 128 bits
2128 bytes
Future-proofing; cap-arch; capability pointers
Draft / reserved
What "Base" Really Means
~47 instructions: LOAD · STORE · OP · OP-IMM · BRANCH · JAL/JALR · LUI · AUIPC · SYSTEM · FENCE
Everything you need to run a compiler-emitted C program except integer mul/div, atomics, and FP
Deliberately missing: multiply, divide — provided by the M extension
Deliberately missing: atomic RMW — provided by the A extension
The RV64I Sign-Extension Trick
32-bit integer ops on RV64 have .W suffix (ADDW , SLLW )
They sign-extend the 32-bit result to 64 bits — so negatives are naturally represented
No separate "32-bit mode" — both widths coexist in the same pipeline
Clever decision: avoids the x86 "mode switch" chaos while enabling efficient 32-bit code
05
Standard Extensions — The Alphabet
Letter Name Adds Size / Cost Status
M Integer Mul/Div MUL, MULH*, DIV, REM (8 ops) ~2-5 k gates Ratified
A Atomics LR/SC + AMOSWAP/ADD/AND/OR/XOR/MIN/MAX ~1-2 k gates Ratified
F Single-precision FP 32 FP regs f0-f31, IEEE 754 binary32 ~15-30 k gates Ratified
D Double-precision FP Widens f-regs to 64 bits; binary64 ops +30-60 k gates Ratified
Q Quad-precision FP binary128 ops (typically trapped & emulated) >100 k gates Ratified
C Compressed 16-bit aliases for common 32-bit instructions (~25-30% size reduction) decoder only Ratified
B Bit manipulation Zba+Zbb+Zbs (count, reverse, clz, bitfield) ~5 k gates Ratified 2021
K Scalar Crypto Zk* — AES, SHA-2/3 round kernels ~10-20 k gates Ratified 2022
V Vector 32 vector regs (VLEN), vector-length-agnostic ops 100 k - 10 M gates Ratified 2021
H Hypervisor VS-mode, 2-stage translation, virtual CSRs area-dependent Ratified 2021
Zicsr CSR access CSRRW/CSRRS/CSRRC — implicit in any privileged core tiny Ratified
Zifencei Instr-fetch fence FENCE.I — self-modifying code synchronisation trivial Ratified
Zicond Conditional ops czero.eqz / czero.nez — branchless selects trivial Ratified 2023
Zicbom/Zicboz Cache Blk Ops cbo.clean / cbo.flush / cbo.inval / cbo.zero CMO only Ratified 2022
G = IMAFD + Zicsr + Zifencei — the "general-purpose" shorthand. A RV64GC core has everything Linux expects.
06
Application Profiles — Fighting Fragmentation
RV64I
frozen base
G
I + M + A + F + D
+ Zicsr + Zifencei
"general-purpose"
RVA22U64 / RVA22S64
= G + C + B + Zihintpause
+ Zicbom/Zicboz + Svpbmt + Sstc…
first Linux-grade baseline (2022)
RVA23U64 / RVA23S64
= RVA22 + V + Vector Crypto
+ Hypervisor + AIA + more
ratified 2024 — current target
RVB23 (bespoke)
Subset of RVA23 for MCUs
no V, optional hypervisor
Why Profiles Exist
With hundreds of extensions, "RISC-V Linux" could mean anything
Profiles lock down a mandatory set for a given class of software — so a binary works on any conforming chip
U64 / S64 suffix: user-mode vs supervisor-mode baseline
RVA23 Highlights
Mandatory RVV 1.0 (min VLEN=128) — SIMDy baseline for Linux distros
Mandatory Hypervisor — KVM-class virtualisation
Mandatory AIA — modern interrupt controller (IMSIC/APLIC)
Cache-block management, PBMT page attributes, Sstc time CSR
Debian/Fedora/Ubuntu "riscv64" ports target RVA23 baseline from 2025
07
Register Model & Calling Convention
# ABI name Role Saver # ABI name Role Saver
x0 zero Hard-wired to 0 (writes ignored) —
x16-x17 a6, a7 Fn args / return 2+ Caller
x1 ra Return address Caller
x18-x27 s2-s11 Saved regs Callee
x2 sp Stack pointer (16-byte aligned) Callee
x28-x31 t3-t6 Temporaries Caller
x3 gp Global pointer (linker-relative addr.) —
f0-f7 ft0-ft7 FP temporaries Caller
x4 tp Thread pointer (TLS base) —
f8-f9 fs0-fs1 FP saved Callee
x5-x7 t0-t2 Temporaries Caller
f10-f11 fa0-fa1 FP args / return Caller
x8 s0/fp Saved / frame pointer Callee
f12-f17 fa2-fa7 FP args Caller
x9 s1 Saved register Callee
f18-f27 fs2-fs11 FP saved Callee
x10-x11 a0, a1 Fn args / return values Caller
f28-f31 ft8-ft11 FP temporaries Caller
x12-x15 a2-a5 Fn args Caller
ABI: ilp32 (RV32), ilp32f/d (with F/D), lp64 , lp64d (RV64). Tail calls use t1 .
Why x0 = 0 is Genius
MV rd, rs → ADDI rd, rs, 0
NOP → ADDI x0, x0, 0
J label → JAL x0, label (discard link)
Unconditional zero-test: BEQ rs, x0, target
No Condition Codes
No flags register — all comparisons are explicit
Compare-and-branch fused: BEQ , BNE , BLT , BGE , BLTU , BGEU
Sets (SLT , SLTU , SLTI ) write 1 or 0 to a register
Good for OoO — no implicit register renaming of flags
08
Instruction Formats — Six Shapes, One Decoder
31
25
20
15
12
7
0
funct7
rs2
rs1
funct3
rd
opcode
R-type
ADD, SUB, AND, OR, SLL, SLT…
imm[11:0]
rs1
funct3
rd
opcode
I-type
ADDI, LW, JALR, CSRRW…
imm[11:5]
rs2
rs1
funct3
imm[4:0]
opcode
S-type
SW, SH, SB
im12
imm[10:5]
rs2
rs1
funct3
imm[4:1]
im11
opcode
B-type
BEQ, BNE, BLT, BGE…
imm[31:12]
rd
opcode
U-type
LUI, AUIPC
im20
imm[10:1]
im11
imm[19:12]
rd
opcode
J-type
JAL
Key invariants across all formats:
• rs1, rs2, rd always land in the same bit positions — register read can start before decode completes
• Sign bit (bit 31) is the immediate's MSB in I/S/B/U/J formats — one sign-extend mux does them all
• Bit scrambling in B/J formats lets branches/jumps reuse as much of the immediate datapath as possible
Interactive Demo · Try me
09
Live Instruction Encoder
Pick an instruction and set its fields — watch the 32-bit encoding and bit-field layout update live.
Controls
Instruction
ADD rd, rs1, rs2
SUB rd, rs1, rs2
XOR rd, rs1, rs2
SLL rd, rs1, rs2
ADDI rd, rs1, imm
ANDI rd, rs1, imm
LW rd, imm(rs1)
JALR rd, rs1, imm
SW rs2, imm(rs1)
BEQ rs1, rs2, imm
BNE rs1, rs2, imm
LUI rd, imm
AUIPC rd, imm
JAL rd, imm
rd
x10 (a0)
rs1
x11 (a1)
rs2
x12 (a2)
imm
42
32-bit Encoding
0x02A58513
Binary (MSB → LSB):
0000001 01010 01011 000 01010 0010011
Bit-field decomposition
10
Privilege Architecture
M
Machine · 0b11
HS
Hypervisor · 0b01+H
S
Supervisor · 0b01
U
User · 0b00
VS
VU
guest modes (H ext.)
MCU (embedded):
M only
RTOS:
M + U
Linux:
M + S + U
KVM host:
M + HS + VS + VU
Rings of privilege — any combination is legal
M — Machine Mode
Highest privilege; has direct access to all CSRs and memory
Holds the boot code, SBI firmware, and platform-specific trap handlers
Every RISC-V core has M-mode — it's the only mandatory level
Key CSRs: mstatus, mepc, mtvec, mcause, mtval, mie, mip
S — Supervisor Mode
Where OS kernels live (Linux, FreeBSD, seL4, Zephyr-MMU)
Can manage virtual memory via satp , handle S-mode traps
SBI (Supervisor Binary Interface) is the ABI between S and M modes — ecall traps up
U — User Mode
Application code runs here
Trap to S-mode via ecall (syscall) or page faults
Limited to CSRs with URW/URO bits (time, cycle, instret counters)
11
Virtual Memory — Sv32 / Sv39 / Sv48 / Sv57
Sv39 — 3-level, 39-bit VA → 56-bit PA
VPN[2] (9b)
VPN[1] (9b)
VPN[0] (9b)
Page Offset (12b)
39-bit Virtual Address
satp.PPN
root page-table
L2 PT
PTE
... 512 PTEs
L1 PT
L0 PT
4 KiB page
Physical
Memory
PA = (PTE.PPN << 12) | offset
Superpages:
4 KiB · 2 MiB · 1 GiB
PTE bits:
V · R · W · X · U · G · A · D
Mode VA bits Levels Page sizes Address space
Sv32 32 2 4 KiB · 4 MiB 4 GiB
Sv39 39 3 4 KiB · 2 MiB · 1 GiB 512 GiB
Sv48 48 4 + 512 GiB 256 TiB
Sv57 57 5 + 256 TiB 128 PiB
Key Mechanisms
satp CSR — selects mode (Bare/Sv32/39/48/57), ASID, and root PPN
SFENCE.VMA [rs1],[rs2] — surgical TLB invalidation by VA/ASID
Each PTE: V R W X U G A D + reserved + PPN
A (accessed) and D (dirty) can be HW- or SW-managed (Svadu )
Svpbmt — page-based memory types (NC, IO)
Svnapot — naturally-aligned power-of-2 contiguous regions (cheap huge pages)
12
Traps, Exceptions & Interrupts
User / S-mode code
trap!
Causes (mcause)
Sync exceptions:
0 Instruction misaligned
2 Illegal instruction
3 Breakpoint
5 Load access fault
7 Store access fault
8 ECall from U
9 ECall from S
12 Instr page fault
13/15 Load/Store page fault
+ Async: SI/TI/EI per mode
HW automatic
• pc → mepc
• cause → mcause
• trap val → mtval
• priv ← M
pc ← mtvec (vectored or direct)
SW saves caller-saved regs, demux on mcause, handle, then MRET
MRET
Interrupt Controllers
CLINT — basic timer + software interrupts; in every core since spec v1.9
PLIC — Platform-Level Interrupt Controller; up to 1023 sources, 15 priority levels, per-hart enables/thresholds/claim-complete
CLIC — embedded: pre-emptive, vectored, stack-based priorities (<1 µs latency)
AIA — modern: IMSIC (per-hart MSI file) + APLIC (wired irq translator) — mandatory in RVA23
Trap Delegation
medeleg /mideleg let M-mode forward specific traps directly to S-mode — zero-cost syscalls
hedeleg /hideleg delegate VS traps to HS
WFI instruction — wait for interrupt; implementations may halt the clock
Reset vector is implementation-defined; typical: 0x80000000 or 0x1000
13
M & A — Multiply/Divide and Atomics
M — Integer Multiply / Divide
MUL rd ← rs1·rs2 (low XLEN)
MULH / MULHSU / MULHU — upper XLEN, signed/unsigned variants
DIV / DIVU / REM / REMU
RV64 adds MULW, DIVW, DIVUW, REMW, REMUW — 32-bit × sign-ext-to-64
Divide-by-zero returns all-ones for quotient, dividend for remainder — no trap
Typical implementations: 1–3 cycle Booth multiplier; multi-cycle sequential divider (SRT radix-4 common)
# sum-of-squares with M ext.
mul t0, a0, a0
mul t1, a1, a1
add a0, t0, t1
ret
A — Atomic Memory Operations
LR.W/D + SC.W/D — load-reserved / store-conditional pair (forward-progress guarantee for restricted sequences)
AMOSWAP, AMOADD, AMOAND, AMOOR, AMOXOR, AMOMIN, AMOMAX, AMOMINU, AMOMAXU
aq /rl acquire/release bits — RCsc memory model hooks
Basis for every lock-free & lock-based primitive in Linux, pthreads, Rust atomics
Zaamo split-out (AMOs only); Zalrsc (LR/SC only)
# lock-free inc via LR/SC
1: lr.w t0, (a0)
addi t1, t0, 1
sc.w t2, t1, (a0)
bnez t2, 1b # retry on contention
Memory Model
RVWMO (Weak Memory Ordering) is the default — similar to Arm, weaker than x86 TSO. Fences: FENCE [pred],[succ] where each field is a mask of i o r w . A Ztso option flips the core to Total Store Order (for easier x86 porting; Alibaba T-Head uses it).
14
Floating-Point — F, D, Q, Zfh, Zfinx
The Classic Extensions
F — 32 separate f0-f31 registers, 32-bit wide (binary32)
D — widens to 64-bit registers, adds binary64 ops (IEEE 754-2008)
Q — 128-bit registers, binary128 (rare; usually trap-and-emulate)
Fused multiply-add: FMADD, FMSUB, FNMADD, FNMSUB
fcsr — 8-bit CSR: 5 sticky exception flags + 3-bit rounding mode (RNE, RTZ, RDN, RUP, RMM, DYN)
Conversions: FCVT.{W/L/S/D}.{S/D/W/L}[U] — every pair-wise combo
Classify: FCLASS returns a 10-bit bitmask (-inf, -norm, -subn, -0, +0, +subn, +norm, +inf, sNaN, qNaN)
New & Specialised Flavours
Zfh — half-precision (binary16) scalar FP; useful for AI quant workloads
Zfhmin — half-precision conversions only (no arithmetic); minimum-cost support
Zfinx / Zdinx / Zhinx — FP "in x-registers": no separate f-regfile; saves ~30% area in MCUs
Zfbfmin / Zvfbfwma — BF16 support, mandated in RVA23 for ML
Zfa — additional FP ops: fli (load immediate), fminm/fmaxm , fround
# dot product fragment (D ext.)
fld ft0, 0(a0)
fld ft1, 0(a1)
fmadd.d fa0, ft0, ft1, fa0
addi a0, a0, 8
addi a1, a1, 8
bne a0, a2, 1b
15
C — Compressed 16-bit Encodings
32-bit: ADDI sp, sp, -16
imm[11:0] = -16
x2
000
x2
OP-IMM
16-bit: C.ADDI16SP sp, -16
011
nzi5
x2
nzi[4,6,8:7,5]
01
Typical code-size reduction: 25-30%
• Reuses the most common register subset x8-x15 with a 3-bit field
• 16-bit and 32-bit mix freely on 2-byte alignment (bits [1:0] = 11 → 32-bit, else 16-bit)
• Decoder expands each C-form to its 32-bit base instruction — zero semantic difference
• Zca/Zcf/Zcd — sub-extensions for fine-grained inclusion
• Zcb / Zcmp / Zcmt — "C extension extras" for further MCU code density
Popular Forms
C.LI — load 6-bit signed immediate
C.LW / C.SW — use 3-bit reg subset
C.ADDI, C.MV, C.ADD — 16-bit ALU ops
C.JR, C.JALR — 16-bit branches
C.BNEZ, C.BEQZ — compare-to-zero branches
C.SLLI, C.SRLI, C.SRAI — shifts
Why Not Thumb-2?
Thumb-2 is a separate ISA with mode switches (BX )
C extends the same ISA — no mode, PC simply advances by 2 or 4
Every 16-bit form has an exact 32-bit expansion, so macro-op fusion / decode is trivial
Enables ~100 bytes/KiB denser code than Thumb-2 in typical benchmarks (Embench)
16
Bit Manipulation (B) & Scalar Crypto (K)
B — Bitmanip (ratified Nov 2021)
Zba — address generation: sh1add, sh2add, sh3add , add.uw
Zbb — basics: clz, ctz, cpop, min/max, sext.b/h, rev8, rol/ror, orc.b, andn/orn/xnor
Zbc — carry-less multiply: clmul, clmulr, clmulh (CRC accelerator)
Zbs — single-bit: bclr, bset, binv, bext (+ immediate forms)
Zbkb/Zbkc/Zbkx — crypto-oriented subsets used by K
# count trailing zeros
ctz a0, a0 # was: ~10 insns software loop
K — Scalar Cryptography (ratified 2022)
Zknd / Zkne — AES-128/192/256 encrypt / decrypt round helpers (aes64es, aes64ds, aes64ks1i, aes64ks2 )
Zknh — SHA-2 (224/256/384/512) and SHA-3 Keccak round functions
Zksed / Zksh — SM4 and SM3 (Chinese crypto standards)
Zkr — entropy source CSR (seed ) — NIST SP 800-90B conformant TRNG interface
Zkt — constant-time guarantee profile — spec promise that listed insns are data-independent latency
# AES round with K ext. (4 insns/round
# vs ~16 with tables — and sidechannel safe)
aes64es a0, a0, a1 # encrypt + sbox
aes64es a0, a0, a2
Why This Matters
OpenSSL, BoringSSL, and the Linux kernel crypto API all have RISC-V vector-and-scalar-crypto backends (2024–). Benchmark: AES-GCM on a SiFive U74 with Zkne+Zbkb runs ~8× faster than the pure-C path. Zkt ratification also unblocked FIPS 140-3 certification pathways for RISC-V silicon.
17
V — The Vector Extension (RVV 1.0)
v0–v31 · VLEN bits wide
SEW=8
e0
e1
e2
…
VL elements of SEW bits
SEW=32
fewer, wider lanes
LMUL = 4 (group 4 vregs)
v8
v9
v10
v11
viewed as one large vector → more throughput, fewer addressable regs
Segment load/store: vlseg2e32.v v4, (a0) → {v4,v5} strided by 2 elts
Indexed: vluxei32.v — gather; vsuxei32.v — scatter
Vector-Length Agnostic (VLA)
Code does not hard-code width; hardware reports VLEN at runtime
Profile enforces VLEN ≥ 128 ; implementations today: 128–2048 bits; research: up to 65,536
Same binary runs on a 128-bit embedded SoC or a 2048-bit server chip
The vsetvl Pattern
vsetvli t0, a0, e32, m4 — "give me up to a0 lanes of 32-bit ints, 4-reg group"
HW returns how many elements fit → put in t0 and vl CSR
Natural "stripmining": loop handles any array length without epilogue
axpy: # y = a*x + y
vsetvli t0, a3, e32, m8
vle32.v v0, (a1)
vle32.v v8, (a2)
vfmacc.vf v8, fa0, v0
vse32.v v8, (a2)
sub a3, a3, t0
slli t0, t0, 2
add a1, a1, t0
add a2, a2, t0
bnez a3, axpy
Sub-extensions: Zve32x/Zve32f/Zve64d (embedded vector tiers) · Zvkn, Zvkng (vector crypto: AES, SHA, GHash) · Zvfh (FP16) · Zvbb/Zvbc (vector bitmanip) · Zvknha/Zvksed (hash/SM crypto)
18
H — Hypervisor & AIA
Guest VM #1
VU app
VS kernel
Stage-1 VM page tables (vsatp)
Guest VM #2
VU app
VS kernel
Stage-1 VM page tables (vsatp)
HS-mode Hypervisor — KVM, Xvisor, bao-hypervisor, seL4-hyp
Stage-2 translation via hgatp · controls vCPU scheduling, virtual timer & interrupts (AIA)
M-mode firmware — OpenSBI / RustSBI · platform init, SBI calls
H Extension (ratified 2021)
Adds VS and VU modes — "virtualised" flavours of S and U
Two-stage address translation: VA → GPA → PA
New CSRs: hgatp, hstatus, hedeleg, hideleg, hcounteren, vsstatus, vsepc, vsscratch…
Hypervisor load/store: HLV.{B,H,W,D} , HSV.{…} — access guest memory with guest permissions
HFENCE.VVMA , HFENCE.GVMA — stage-1 / stage-2 TLB flush
AIA — Advanced Interrupt Arch
Replaces legacy PLIC for virtualisation
IMSIC — per-hart MSI-capable interrupt file (PCIe-native)
APLIC — translates wired interrupts into MSIs
Directly supports guest-external interrupts (no hyp-trap on every delivery)
Mandatory in RVA23 — fundamental for cloud workloads
19
Custom Extensions & CHERIoT
Custom Opcode Space
RISC-V reserves custom-0 (0x0B) , custom-1 (0x2B) , custom-2 (0x5B) , custom-3 (0x7B) for vendor use
No approval required — vendors innovate; if useful, propose to RISC-V International for standardisation
Intel AVX-style "add AVX2 to the world" problem eliminated: new encoding space is reserved, never collides
Examples: SiFive Xsfvcp coprocessor interface, Andes ACE , T-Head vector subset, Tenstorrent DSP ops
GNU & LLVM toolchains support .insn directive and inline intrinsics for custom ops
CHERIoT / CHERI-RISC-V
Capability-based pointers — every pointer is a 128-bit (CHERI) or 64-bit (CHERIoT) sealed capability carrying bounds + permissions
Deterministic spatial safety (no buffer overflows) and temporal safety (no UAF with revocation)
CHERI-RISC-V was the template for Arm Morello research CPU
CHERIoT — MSR-Cambridge, lowRISC, SCI Semiconductor — open Ibex-based CHERI CPU for embedded/IoT security
Ongoing merger path via Zicfilp / Zicfiss (landing-pad + shadow-stack CFI), Smmpm (Memory-protection MMU)
Standardisation Pipeline
Draft → Public Review → Frozen → Ratified. Once Ratified , an extension's bits are fixed forever. The process is run by RISC-V International Technical Steering Committee; all specs live in GitHub riscv/riscv-isa-manual , riscv/riscv-*-ext . Each extension has an associated Sail model, a conformance test suite (riscv-arch-test ), and a compliance signature mechanism.
Interactive Demo · Step the CPU
20
Mini RISC-V Stepper — Fibonacci
A minimal RV32I model. Click Step to execute one instruction. Changed registers highlight in amber. The program computes fib(N) in a0 where N starts in a1 .
CPU State
pc = 0x00001000 cycle = 0
Next instr:
0x00058863 BEQZ a1, done
21
RISC-V International — Governance
Organisation
Founded 2015 as RISC-V Foundation (US non-profit)
2019: relocated to Zug, Switzerland as RISC-V International — explicitly geopolitics-neutral; insulates members from US export-control risks
~4000+ members in 70+ countries (2025): includes NVIDIA, Intel, Qualcomm, Samsung, Google, Meta, Huawei, Alibaba, ARM(!)
Membership tiers: Premier, Strategic, Community · CC-BY spec means a membership is not required to implement
Technical Process
~40+ Task Groups; each produces a draft → public review → freeze → ratification
Every ratified extension has: prose spec, Sail formal model , architectural tests
ACTs (Architectural Compatibility Tests) must pass for a chip to bear the "RISC-V" trademark
Profile and Platform specs coordinate vendor convergence around a common target
2025 Strategic Priorities
RVA23 ecosystem rollout — Debian, Fedora, Ubuntu all targeting it by default in 2025-26
Automotive (ISO 26262 certified cores from SiFive, Andes, Synopsys)
Data-centre: Ventana Veyron V2/V3, Tenstorrent Ascalon, SiFive P870-D
AI: RISC-V-based NPUs (Meta MTIA, Tenstorrent Wormhole, Esperanto ET-SoC-1)
Open-source silicon pipeline: OpenROAD, OpenMPW, Google+eFabless Sky130/Gf180 shuttle
22
Vendor Landscape
Commercial IP
SiFive (US) — E2/E7, U74, P550, P670, P870, P870-D; Linux-class and server-class
Andes (Taiwan) — N/D/A series incl. NX27V vector, AX45MP, QiLai multicore
Codasip (CZ) — customisable cores; L110, A730, X730
Synopsys ARC-V RMX/RHX; MIPS eVocore (post-2021 pivot)
Imagination Catapult MK/AX series
Nuclei (CN) N/N9/NA series
In-house / Fabless Silicon
Alibaba T-Head — XuanTie C906, C910, C920 (Linux SBCs LicheePi)
Espressif — ESP32-C/H/P/S family (RISC-V replaces Xtensa)
Microchip — PolarFire SoC (4×U54 + E51)
Tenstorrent — Ascalon cores in Blackhole, Wormhole AI chips; licences to LG, Japan
Ventana — Veyron V1/V2 chiplet server CPUs
Rivos , Esperanto , Condor Computing , InspireSemi
Open-Source Cores
Rocket — UCB reference; in-order, Scala/Chisel, highly parameterisable
BOOM — Berkeley Out-of-Order Machine; OoO superscalar
CVA6 / Ariane — ETH/OpenHW 6-stage 64-bit core; automotive-cert track
Ibex / CVE2 — lowRISC MCU core; base for CHERIoT
VexRiscv — SpinalHDL, plugin-modular, FPGA sweet-spot
PicoRV32, NEORV32, SweRV-EH/EL — tiny/medium embedded flavours
23
Notable Implementations — From µ to Server
Core Vendor µarch Target Clock Extensions
PicoRV32 Yosys/Claire Wolf multi-cycle, <1 kLE FPGA glue logic ~100 MHz RV32IMC
Ibex lowRISC 2-stage in-order OpenTitan security chip ~500 MHz ASIC RV32IMC + B + crypto
E31 / E51 SiFive 5-stage in-order MCU / mgmt cores 1 GHz (16 nm) RV32IMC / RV64IMAC
CVA6 OpenHW Group 6-stage in-order, 64-bit Linux-capable SoCs 1 GHz ASIC RV64GC + H
U74 SiFive 8-stage dual-issue HiFive Unmatched, VisionFive 2 1.2 GHz (28 nm) RV64GCB
C910 Alibaba T-Head 3-issue OoO, 12-stage LicheePi 4A, BeagleV 2.5 GHz (12 nm) RV64GCBV (pre-1.0)
P670 / P870 SiFive 3/4-issue OoO + vector Mobile / auto SoCs 2.6-3.6 GHz RV64GCBVH (RVA23)
Veyron V2 Ventana 8-issue OoO server Chiplet data-centre CPU 3.6 GHz (5 nm) RV64GCBVH + custom
Ascalon Tenstorrent 8-wide OoO + vector AI chiplets; Samsung Foundry 3+ GHz (3 nm) RV64GCBVH + custom
BOOM v3 UCB / Chipyard 4-issue OoO, open source Research / silicon demos ~1 GHz (28 nm) RV64GC (+ optional V)
SG2042 Sophgo 64× C910 cluster Server / workstation SBC 2 GHz (12 nm) RV64GCB (+V 0.7)
ET-SoC-1 Esperanto 1088× ET-Minion + 4× ET-Maxion ML inference accel. 1 GHz (7 nm) RV64GCV + custom tensor
24
Software Ecosystem
Compilers & Toolchains
GCC — first-class support since 7.x; -march=rv64gc_zba_zbb_zbs_v style machine strings
LLVM / Clang — tier-1 since LLVM 15; RISC-V backend maintained by Igalia, SiFive, Rivos
Rust — riscv64gc-unknown-linux-gnu tier-1 from 2024; riscv32imac-* for bare-metal Rust
Go — GOARCH=riscv64 since Go 1.14; production-ready since 1.21
Binutils , GDB , QEMU , Spike (reference ISS)
Operating Systems
Linux — upstream RV64 port since 4.15 (2018); RV32 port 5.18; systemd, KVM, eBPF all supported
Distros: Debian, Ubuntu, Fedora, openSUSE, Gentoo, Arch-RISCV, Alpine, AOSP — all provide riscv64 ports
FreeBSD 13+, NetBSD , OpenBSD (2023)
Embedded RTOS : Zephyr, FreeRTOS, RT-Thread, NuttX, ThreadX, seL4, PX5
Firmware : OpenSBI (M-mode), RustSBI, U-Boot, EDK II (UEFI), coreboot/LinuxBoot
Simulation & Verification
Spike (riscv-isa-sim) — golden reference C++ ISS
Sail — formal executable ISA model (authoritative)
riscv-arch-test — official compliance suite
RISCOF — Python framework wiring Spike ↔ DUT for signature-compare tests
QEMU tcg+kvm; Renesas Renode ; AndeSight , SiFive Studio
Formal: Axiomatic Alive-RV, ESOP memory model, riscv-semantics (K framework)
Libraries & Runtimes
glibc / musl / newlib — all have riscv ports
OpenSSL / BoringSSL / Botan — vector + scalar crypto backends
LLVM libc++, libstdc++ — tier-1
PyTorch, TensorFlow, JAX — run on RVA22+ via CPU backends; vector paths landing 2025
CUDA-equivalent : no direct equivalent; Tenstorrent uses TT-Metalium, Esperanto has its own stack
WebAssembly : V8, SpiderMonkey, Wasmtime all have RV64 JIT backends
25
Adoption Case Studies
NASA HPSC (2024→)
NASA's next-gen spaceflight computer — contract awarded to Microchip / SiFive
8× SiFive X280 RISC-V cores (rad-hardened), targets 100× the perf of current RAD750
Expected flight hardware: Artemis follow-ons, Mars sample return
Drove adoption of riscv-abi rad-tolerant compilation flows
Western Digital (2018→)
Pledged to ship >2 B RISC-V cores/year — inside every storage controller
Open-sourced SweRV EH1/EH2/EL2 — in-order 32-bit MCUs
Replaced proprietary microcontrollers in NVMe SSDs & HDDs
Proof point: the boring, huge-volume embedded market adopts first
Tenstorrent Wormhole / Blackhole
Ascalon 8-wide OoO RISC-V cores drive & surround the Tensix AI tiles
LG, Samsung Foundry, and Hyundai partnered for automotive variant
Rivos and Jim Keller's team → first full RISC-V training-class silicon (2025)
Meta MTIA v2 (2024)
Meta's in-house AI inference accelerator — RISC-V cores orchestrate Tensor tiles
Replaces the control-plane Arm cores used in MTIA v1
Citation: Meta's 2024 OCP talk — RISC-V chosen for custom-extension freedom & cost
Similar story: Google Titan M2 (security chip), Apple rumoured use in coprocessors
Espressif ESP32-C3/C6/H2/P4 — replaces Xtensa in hundreds of millions of Wi-Fi/BLE IoT parts.
NVIDIA GPU System Processor (GSP) runs RISC-V NV-RISCV in every GeForce/Blackwell GPU since 2017.
Seagate Custom high-perf RISC-V core drives Mozaic 3+ HAMR hard drives; open sourced as "Seagate Core".
Google RVV crypto work upstreamed to BoringSSL; internal Pixel / data-centre exploration with Rivos.
26
Future Directions
ISA Extensions In-Flight
RAS — Reliability/Availability/Serviceability (Zicntr, Zihpm, RAS signalling)
Zicfilp / Zicfiss — CFI: landing-pad & shadow stack
Zvk* — vector crypto: ratified 2023, rolling into vendor silicon
Smmpt , Smepmp — enhanced PMP / memory-protection
RVP — packed SIMD (DSP) for MCUs that can't afford V
Zacas — double-word atomic compare-and-swap
Performance Frontier
8-wide OoO with >256-entry ROB → SPEC2017 parity with Neoverse V2 expected by 2026
Big-vector (VLEN=2048+) for HPC / LLM inference
Decoupled vector / matrix co-processor interface (RVI-AME "matrix" extension, draft)
Chiplet + UCIe-based compositions — Ventana, Tenstorrent
Coherent accelerator interfaces via CHI-C2C and CXL
Safety & Security
CHERI-RISC-V / CHERIoT convergence paths into standard
Confidential Compute — AP-TEE / CoVE (Confidential VM Extension)
ISO 26262 ASIL-D certified cores (Codasip, Andes, SiFive Shield)
DO-254 / DAL-A for avionics (SiFive Intelligence X)
Post-quantum crypto intrinsics (ML-KEM, ML-DSA accelerators)
Open Questions
Fragmentation risk — with profiles strictly enforced, can RVA23 become the de-facto "RISC-V Linux" target before vendors diverge too far with custom extensions?
Software tail — SIMD libraries (SIMDe, highway, xsimd) backfilling RVV paths; AI framework vector lowering still immature vs CUDA/ROCm
Geopolitics — US export-control impacts on Premier-tier membership; RVI's Swiss domicile helps but doesn't immunise silicon
Matrix-multiply ISA standardisation — will vendors wait, or ship proprietary extensions and negotiate later?
27
Summary — What Makes RISC-V Different
The 30-Second Pitch
A royalty-free , modular , frozen-base ISA that scales from 10 kGE MCUs to 8-wide OoO server cores
Five design pillars: open · clean · modular · scalable · stable
Rich extension alphabet (M A F D C B K V H) + custom opcode spaces
Real hardware shipping: >10 billion cores/year across every tier of silicon
Why It Won Ground
Software stack (Linux, GCC, LLVM, Rust, Go) was upstream-mature before big vendors arrived
Profiles prevent ABI-level fragmentation across vendors
Formal ISA model (Sail) and official compliance suite keep implementations honest
Geopolitical neutrality — Swiss domicile shields members from unilateral export controls
Where To Next
RVA23 rollout in Linux distros 2025-26; Wikipedia-style everywhere adoption
First truly competitive RISC-V server CPU vs Graviton / Neoverse in 2026
RISC-V AI accelerators replace Arm control-plane cores (Meta MTIA pattern)
Safety-critical/avionics certified cores unlocks automotive and aerospace
Further Reading
riscv.org/specifications/ratified — canonical spec index
github.com/riscv/riscv-isa-manual — source of truth
Patterson & Waterman, The RISC-V Reader: An Open Architecture Atlas
Waterman PhD thesis (2016) — design rationale
RVA23 profile specification (2024)
Sail model: github.com/riscv/sail-riscv
"An Instruction Set Should Be Free" — Patterson & Asanović, 2014