M³-AVM: A Complete ISA for Deterministic, Distributed AI Inference (All 256 Opcodes)
A deep dive into the instruction set architecture of M³-AVM — a virtual machine designed to run Transformers, SSMs, SNNs, classical ML, RAG, and full-duplex audio models in a single unified context.
Why does AI need its own ISA?
Today, running multiple AI models together means gluing together Python processes over gRPC. Want Moshi for voice, Llama for reasoning, Mamba for draft tokens, XGBoost for decisions, and RAG for retrieval? That's 5 processes, 5 serialization layers, 5 sets of deadlines, and 5 independent states that can't roll back together.
The result:
~800ms end-to-end latency for pipelines that should run in ~500ms
No coordinated rollback — barge-in in audio can't restore the LLM's KV cache atomically
No composite deadlines — audio at 80ms, text at 200ms, and video at 33ms have no shared scheduler
No shared memory — every tensor crossing a process boundary gets serialized
M³-AVM proposes a different approach: move the orchestration into the virtual machine itself . Instead of treating models as black-box services, treat them as engines that share a memory space, a scheduler, and a state-versioning mechanism — all coordinated by the ISA.
This article walks through the complete instruction set.
Instruction Format
M³-AVM uses dual-mode instructions :
Opcode range
Size
Use
0x00–0x7F
32 bytes
Legacy v1.3 (control flow, basic tensors, SSM, codec)
0x80–0xFF
64 bytes
v2.0+ (cluster, universal ops, system)
The dispatcher reads byte 0. If op >= 0x80, it reads 64 bytes; otherwise 32.
64-byte layout
┌──────────────────────────────────────────────────────────────────────┐
│ Byte 0 : OPCODE (u8) │
│ Byte 1 : FLAGS (u8) │
│ Byte 2 : RDEST (u8) │
│ Byte 3 : RSRC1 (u8) │
│ Byte 4 : RSRC2 (u8) │
│ Byte 5 : RSRC3 (u8) │
│ Byte 6 : RSRC4 (u8) │
│ Byte 7 : RSRC5 (u8) │
│ Bytes 8-15 : LAMPORT (u64 LE) — distributed logical clock │
│ Bytes 16-23: DEADLINE (u64 LE) — absolute EDF deadline (ns) │
│ Bytes 24-31: PAYLOAD_EXT (8 B) — HMAC, RDMA key, checksum │
│ Bytes 32-63: PAYLOAD_CORE (32 B) — opcode-specific │
└──────────────────────────────────────────────────────────────────────┘
Enter fullscreen mode
Exit fullscreen mode
Flags (byte 1)
Bits
Name
Values
0–1
PRECISION
0=FP32, 1=BF16, 2=FP16, 3=INT8/FP8
2
FUSED_ACT
0=none, 1=SILU, 2=ReLU, 3=GELU
3
INPLACE
0=new, 1=in-place
4
TRANSPOSE
0=normal, 1=Wᵀ·x
5
PRIORITY
0=GREEN, 1=BLUE, 2=RED
6
ATOMIC
0=normal, 1=atomic
7
PERSIST
0=volatile, 1=WAL-persisted
Register banks
Type
Count
Width
Range
GPR
256
64-bit
R0–R255
VEC
32
256-bit (8×f32)
V0–V31
STATE
64
opaque handle
S0–S63
SPECIAL
19
64-bit
SP0–SP18
Special registers include PC, SP, CTX_ID, NODE_ID, LAMPORT, DEADLINE, DEVICE_ID, POWER_BUDGET_MW, CLOCK_HEALTH, and others added by RFCs.
The Complete Opcode Map (0x00–0xFF)
Domain 0 — Control Flow (0x00–0x0F)
Hex
Name
Purpose
0x00
HALT
Terminate context
0x01
TENSOR
Allocate tensor
0x02
ATTN
Attention (with KV cache, causal, flash modes)
0x03
STREAM
Stream I/O to/from ring buffers
0x04
FORK
Snapshot context (CoW)
0x05
ABORT
Rollback to snapshot
0x06
SENSE
Read from sensor (audio PCM, tokens, etc.)
0x07
NORM
Layer/RMS/Batch normalization
0x08
FFN
Feed-forward network
0x09
EMBED
Embedding lookup
0x0A
ADD
Elementwise add
0x0B
SAMPLE
Sampling (with TOPK mode)
0x0C
COMPARE
Predicate comparison
0x0D
IF_EQUAL
Conditional branch
0x0E
JUMP
Unconditional jump
0x0F
IF_INTERRUPT
Branch on interrupt flag
Domain 1 — Matrix & Recurrence (0x10–0x1F)
Hex
Name
Purpose
0x10
MATVEC
Matrix-vector (with TRANSPOSE, FUSED_ACT)
0x11
MUL
Elementwise multiply
0x12
SILU
SiLU activation
0x13
SSM_SCAN
Mamba selective scan
0x14
SSM_RESET
Reset SSM state
0x15
CODEC_ENC
Audio codec encode (Mimi)
0x16
CODEC_DEC
Audio codec decode
0x17
AUDIO_ALIGN
Timestamp alignment
0x18
CTX_SWITCH
Switch engine (MAMBA/TRANSFORMER/AUDIO)
0x19
ROPE
Rotary position embedding
0x1A
REMOTE_SPAWN
Spawn context on remote node
0x1B
SIGNAL
Out-of-band signal (ABORT/FORK/PING)
0x1C
SEND_TENSOR
Transfer tensor to remote node
0x1D
BARRIER
Distributed barrier
0x1E
CONV
N-dimensional convolution
0x1F
GATHER
Gather/scatter by index (MoE, GNN)
Domain 2 — Memory & Universal (0x20–0x2F)
Hex
Name
Purpose
0x20
SPIKE_STEP
LIF neuron step (SNN)
0x21
DENOISE_STEP
Diffusion step (DDPM/DDIM)
0x22
FOREST
Tree ensemble evaluation (XGBoost)
0x23
DISTANCE
Vector distance + top-k
0x24
RANK1_UPDATE
Rank-1 update (DeltaNet, Titans)
0x25
ODE_STEP
ODE integration (Euler/RK2/RK4)
0x26
ARENA_ALLOC
Arena allocation
0x27
ARENA_RESET
Reset arena O(1)
0x28
SNAPSHOT
Explicit snapshot
0x29
RESTORE
Explicit restore
0x2A
MEMCPY
Memory copy (host↔GPU↔NIC)
0x2B
MEMSET
Memory set
0x2C
PREFETCH
Cache hint
0x2D
RESHAPE
Tensor reshape (view or copy)
0x2E
SLICE
Tensor slice
0x2F
CONCAT
Concatenate tensors
Domain 3 — Advanced Tensors (0x30–0x3F)
Hex
Name
Purpose
0x30
SORT
Sort by axis
0x31
TOPK
Top-k selection
0x32
ARGMAX
Argmax
0x33
REDUCE
Reduction (sum/mean/max/min/prod)
0x34
BROADCAST
Broadcast to shape
0x35
PAD
Padding
0x36
TILE
Tile/repeat
0x37
TRANSPOSE
Axis permutation
0x38
KV_TRUNCATE
Truncate KV cache
0x39
KV_COMPRESS
Compress KV cache
0x3A
FLASH_ATTN
IO-aware attention
0x3B
ATTN_SPARSE
Sparse attention (top-k)
0x3C
SOFTMAX
Softmax (with temperature)
0x3D
GELU
GELU activation
0x3E
SIGMOID
Sigmoid
0x3F
TANH
Tanh
Domain 4 — Activations & Audio (0x40–0x4F)
Hex
Name
Purpose
0x40
RELU
ReLU
0x41
EXP
Exponential
0x42
LOG
Logarithm
0x43
CLIP
Clip to range
0x44
DEPFORMER
Moshi Depformer (6×1024×16)
0x45
STREAM_MERGE
Merge 17 audio streams
0x46
VAD_DETECT
Voice activity detection
0x47
AUDIO_RESAMPLE
Sample rate conversion
0x48
AUDIO_FILTER
FIR/IIR filter
0x49
AUDIO_WINDOW
Windowing (Hann/Hamming)
0x4A
MULTICAST
Multicast to N nodes
0x4B
MIGRATE
Transactional state migration
0x4C
HEARTBEAT
Heartbeat with load info
0x4D
PING
RTT probe
0x4E
REDUCE_REMOTE
Allreduce across nodes
0x4F
GATHER_REMOTE
Allgather across nodes
Domain 5 — Retrieval & Classical ML (0x50–0x5F)
Hex
Name
Purpose
0x50
RAG_INDEX_ADD
Add to vector index
0x51
RAG_INDEX_DEL
Remove from index
0x52
RAG_SEARCH
Vector search (flat/IVF/HNSW/PQ)
0x53
EMBED_LOOKUP
Bag-of-embeddings lookup
0x54
HASH_BUCKET
Locality-sensitive hash
0x55
QUANTIZE_VEC
Vector quantization
0x56
PQ_ENCODE
Product quantizer encode
0x57
PQ_DECODE
Product quantizer decode
0x58
KMEANS_STEP
One Lloyd iteration
0x59
LINEAR_REG
Linear regression
0x5A
LOGISTIC_REG
Logistic regression
0x5B
NAIVE_BAYES
Naive Bayes
0x5C
SVM_PREDICT
SVM prediction
0x5D
PCA_STEP
Power iteration step
0x5E
STANDARDIZE
Z-score/min-max
0x5F
NEAREST_CENTROID
Nearest centroid
Domain 6 — Determinism & Conversion (0x60–0x6F)
Hex
Name
Purpose
0x60
RNG_SEED
Seed RNG per context
0x61
RNG_NEXT
Next u64
0x62
RNG_NORMAL
Box-Muller
0x63
RNG_UNIFORM
Uniform in [a,b)
0x64
HASH
SipHash/xxHash
0x65
CHECKSUM
CRC32/xxHash
0x66
HMAC
HMAC-SHA256
0x67
CAST
Cast between dtypes
0x68
QUANTIZE
Q4_0/Q4_K/Q6_K/Q8_0
0x69
DEQUANT
Dequantize
0x6A
CYCLES_COUNT
RDTSC-like counter
0x6B
TRACE_EVENT
Structured tracing
0x6C
SANITY_CHECK
Absorb NaN/Inf
0x6D
PREEMPT_CHECK
Read atomic flag
0x6E
ASSERT
Predicate trap
0x6F
DUMP
Debug snapshot
Domain 7 — Scheduler & Clock (0x70–0x7F)
Hex
Name
Purpose
0x70
YIELD
Yield CPU
0x71
SET_DEADLINE
Set EDF deadline
0x72
GET_DEADLINE
Read deadline
0x73
PRIORITY_SET
Set priority
0x74
PRIORITY_GET
Read priority
0x75
LOCK
Cross-context mutex
0x76
UNLOCK
Release
0x77
FENCE
Memory barrier
0x78
CLOCK_QUERY
Clock health
0x79
CLOCK_SYNC
Force resync
0x7A
CLOCK_ADJUST
Apply offset
0x7B
DEADLINE_NEGOTIATE
Renegotiate deadline
0x7C–0x7F
Reserved
—
Domain 8 — Extended Cluster (0x80–0x8F)
Hex
Name
Purpose
0x80
REMOTE_SPAWN_EXT
64B remote spawn
0x81
SIGNAL_EXT
64B signal with capability
0x82
SEND_TENSOR_EXT
64B send with compression
0x83
BARRIER_EXT
64B barrier
0x84
MIGRATE
WAL-backed migration
0x85
MULTICAST_EXT
64B multicast
0x86
REDUCE_REMOTE_EXT
64B allreduce
0x87
GATHER_REMOTE_EXT
64B allgather
0x88
SCATTER_REMOTE
Scatter across nodes
0x89
BROADCAST_REMOTE
Broadcast
0x8A
HEARTBEAT_EXT
64B heartbeat
0x8B
PING_EXT
64B ping
0x8C
NODE_JOIN
Join cluster (handshake)
0x8D
NODE_LEAVE
Graceful leave
0x8E
NODE_SUSPECT
SWIM suspect
0x8F
NODE_DEAD
SWIM dead
Domain 9 — MMIO & Accelerators (0x90–0x9F)
Hex
Name
Purpose
0x90
GPU_LAUNCH
Launch GPU kernel
0x91
GPU_WAIT
Wait for event
0x92
NIC_SEND
Zero-copy TX
0x93
NIC_RECV
Zero-copy RX
0x94
DMA_START
Start DMA
0x95
DMA_WAIT
Wait for DMA
0x96
ATOMIC_CAS
Compare-and-swap
0x97
ATOMIC_ADD
Fetch-and-add
0x98
ATOMIC_XCHG
Exchange
0x99
WFI
Wait-for-interrupt
0x9A–0x9F
Reserved
—
Domain 10 — System (0xA0–0xAF)
Hex
Name
Purpose
0xA0
LOAD_MODEL
Load GGUF/safetensors
0xA1
UNLOAD_MODEL
Unload
0xA2
SPAWN_CONTEXT
New context
0xA3
KILL_CONTEXT
Kill context
0xA4
SET_AFFINITY
CPU affinity
0xA5
PROFILE_START
Start profiling
0xA6
PROFILE_STOP
Stop profiling
0xA7
SET_MODEL
Switch model
0xA8
GET_MODEL
Query model
0xA9
MODEL_SWITCH
Multi-model switch
0xAA–0xAF
Reserved
—
Domain 11 — ESCAPE (0xB0–0xB6)
Hex
Name
Purpose
0xB0
ESCAPE
64B extension prefix
0xB1
ESCAPE_128
128B extension
0xB2
ESCAPE_256
256B extension
0xB3
ESCAPE_VAR
Variable length extension
0xB4
VERSION
Declare version
0xB5
CAPABILITY_QUERY
Query capabilities
0xB6
CAPABILITY_ASSERT
Assert capability
Domain 12 — Local Compute Topology (0xB7–0xBF)
Hex
Name
Purpose
0xB7
DEVICE_QUERY
Query device by capability
0xB8
DEVICE_SELECT
Set active device
0xB9
DEVICE_HEALTH
Telemetry
0xBA
PLACE
Placement hint
0xBB
SHARD
Tensor/pipeline/expert parallel
0xBC
REPLICATE
Data parallel
0xBD
REDUCE_LOCAL
Intra-rig collective
0xBE
MIGRATE_LOCAL
Device↔device migration
0xBF
BARRIER_LOCAL
Intra-rig barrier
Domain 13 — Storage Tiering (0xC0–0xC7)
Hex
Name
Purpose
0xC0
PIN
Pin to tier
0xC1
UNPIN
Release pin
0xC2
PREFETCH_TIER
Prefetch next tier
0xC3
STREAM_WEIGHTS
Sequential streaming
0xC4
EVICT
Evict from tier
0xC5
TIER_QUERY
Telemetry
0xC6
TIER_POLICY
LRU/LFU/pin-first
0xC7
WEIGHTS_HINT
Hint access pattern
Domain 14 — Power & Thermal (0xC8–0xCF)
Hex
Name
Purpose
0xC8
POWER_QUERY
mW, temperature, throttling
0xC9
POWER_BUDGET_SET
Set budget
0xCA
POWER_BUDGET_GET
Read budget
0xCB
THERMAL_QUERY
Temperature
0xCC
THROTTLE_POLICY
Degradation policy
0xCD
POWER_CAP
Hard cap
0xCE
ENERGY_ACCOUNT
Consumption
0xCF
THERMAL_HEADROOM
Headroom
Domain 15 — Composite Deadlines (0xD0–0xD7)
Hex
Name
Purpose
0xD0
DEADLINE_AND
All sub-deadlines
0xD1
DEADLINE_OR
At least one
0xD2
DEADLINE_N_OF_M
N of M
0xD3
DEADLINE_CHAIN
Temporal chain
0xD4
DEADLINE_QUERY
Status
0xD5
DEADLINE_RELAX
Relax sub-deadline
0xD6
DEADLINE_TIGHTEN
Tighten
0xD7
DEADLINE_RESET
Reset tree
Domain 16 — Security (0xD8–0xDF)
Hex
Name
Purpose
0xD8
KEY_AGREE
X25519 handshake
0xD9
KEY_ROTATE
Rotate keys
0xDA
KEY_DERIVE
HKDF
0xDB
REPLAY_CHECK
Replay window
0xDC
CRYPTO_SIGN
Sign
0xDD
CRYPTO_VERIFY
Verify
0xDE
CRYPTO_ENCRYPT
Encrypt
0xDF
CRYPTO_DECRYPT
Decrypt
Domain 17 — Observability (0xE0–0xE7)
Hex
Name
Purpose
0xE0
TRACE_SPAN_START
Open span
0xE1
TRACE_SPAN_END
Close span
0xE2
TRACE_ATTR
Add attribute
0xE3
TRACE_LINK
Link spans
0xE4
TRACE_EXPORT
Export OTLP
0xE5
TRACE_INJECT
Inject context
0xE6
TRACE_EXTRACT
Extract context
0xE7
TRACE_QUERY
Query traces
Domain 18 — KV Tiering (0xE8–0xEF)
Hex
Name
Purpose
0xE8
KV_TIER_MOVE
Move across tiers
0xE9
KV_EVICT
Evict (attention-aware)
0xEA
KV_PIN
Pin layer
0xEB
KV_RETRIEVE
Recall evicted
0xEC
KV_RECOMPUTE
Partial recompute
0xED
KV_COMPRESS_LOSSY
Adaptive quantization
0xEE
KV_SPARSE_MASK
Sparse mask
0xEF
KV_MIGRATE_REMOTE
Move to remote node
Domain 19 — Compression (0xF0–0xF3)
Hex
Name
Purpose
0xF0
COMPRESS
Compress tensor
0xF1
DECOMPRESS
Decompress
0xF2
COMPRESS_QUERY
Query best algorithm
0xF3
ENTROPY_ESTIMATE
Estimate entropy
Domain 20 — Training (0xF4–0xFB)
Hex
Name
Purpose
0xF4
GRAD_ACCUM
Accumulate gradient
0xF5
GRAD_ZERO
Zero gradient
0xF6
OPTIMIZER_STEP
Apply update
0xF7
OPTIMIZER_INIT
Initialize optimizer
0xF8
LORA_APPLY
Apply LoRA adapter
0xF9
LORA_MERGE
Merge into base
0xFA
CHECKPOINT_SAVE
Save
0xFB
CHECKPOINT_RESTORE
Restore
Domain 21 — Unified Model IR (0xFC–0xFE)
Hex
Name
Purpose
0xFC
MODEL_LOAD_UNIFIED
Load any format
0xFD
MODEL_COMPILE
Compile to M3 bytecode
0xFE
MODEL_QUERY
Metadata
Domain 22 — Terminal (0xFF)
Hex
Name
Purpose
0xFF
NOP
No-op
Memory Model — 24 Regions
Every address is 64 bits:
Bits 60–63 : region selector (16 regions)
Bits 0–59 : offset (1 EiB per region)
ID
Name
R/W
CoW
Purpose
0x0
TEXT
RO
—
Bytecode
0x1
GLOBAL
RO
—
Constants, device registry
0x2
WEIGHTS
RO (mprotect)
—
Model weights
0x3
ACTIVATION
RW
✅
Scratch per context
0x4
KV_CACHE
RW
✅
KV, SSM, SNN, diffusion state
0x5
ARENA
RW
—
Actor heap
0x6
SHARED
RW (fence)
✅
Zero-copy cross-context
0x7
WAL
append
—
Write-ahead log
0x8
SNAPSHOT
RO
—
CoW snapshots
0x9
STREAM_RING
ring
—
Audio/video buffers
0xA
RAG_INDEX
RW
—
Vector DB
0xB
CLUSTER_STAGING
RW
—
TX/RX
0xC
FEDERATED
RW
—
Federated learning (reserved)
0xD
CONFIDENTIAL
enc
—
TEE/enclave (reserved)
0xE
SCRATCH_DEVICE[i]
RW
—
Per-device VRAM
0xF
MMIO
RW
—
NIC, DMA, accelerators
0x10–0x17
Reserved by RFCs
—
—
Storage/KV tiers, keys, gradients
Example: Full-Duplex Voice Assistant (Moshi)
; ============================================================
; Full-duplex voice assistant with barge-in
; ============================================================
.data
FRAME_PCM_LEN .equ 1920 ; 80ms @ 24kHz
N_CODEBOOKS .equ 16
N_STREAMS .equ 17
.text
main:
; Composite deadline: audio 80ms AND text 200ms
DEADLINE_AND rTree, CHILDREN=[rDlAudio, rDlText]
SET_DEADLINE rDlAudio, 80000000
SET_DEADLINE rDlText, 200000000
audio_loop:
SENSE rPCM, AUDIO_PCM, LEN=FRAME_PCM_LEN
VAD_DETECT rVAD, rPCM THRESHOLD=0.5
COMPARE rVAD, 0.5
IF_GREATER rVAD, barge_in
CODEC_ENC rCodesIn, rPCM TENSOR
STREAM_MERGE rMixed, rAgent, rUser N_STREAMS=N_STREAMS MODE=mix
; Transformer (Temporal, 32 layers)
CTX_SWITCH TRANSFORMER, RED
FORK rSnap
CALL temporal_forward rMixed, rHidden
; Depformer (audio generation)
CTX_SWITCH AUDIO, BLUE
DEPFORMER rCodesOut, rHidden, rAudioEmb, rKVDep
DEP_LAYERS=6 DEP_DIM=1024 CODEBOOKS=N_CODEBOOKS
ACOUSTIC_DELAY=1 TEMPERATURE=0.8 TOP_K=50 STREAM_ID=0
CODEC_DEC rAudioOut, rCodesOut TENSOR
STREAM rAudioOut, CHANNEL=4 BLOCKING
IF_INTERRUPT rFlag, handle_barge_in
JUMP audio_loop
barge_in:
SIGNAL NODE=LOCAL, CTX=rCtxText, KIND=ABORT
JUMP handle_barge_in
handle_barge_in:
ABORT rSnap
SSM_RESET rH D_INNER=4096 D_STATE=16 LAYER=0
JUMP audio_loop
Enter fullscreen mode
Exit fullscreen mode
Example: Multi-Model Pipeline (Voice + RAG + LLM + Classifier)
pipeline_loop:
; 1. Audio input
SENSE rPCM, AUDIO_PCM, LEN=1920
; 2. RAG retrieval
EMBED rQueryEmb, rPCM, rBge MODEL=BGE_M3
RAG_SEARCH rDocs, rQueryEmb, rRag METRIC=COSINE TOPK=5 MODE=IVF
; 3. Fast classifier (XGBoost)
CONCAT rFeatures, rQueryEmb, rTurnMeta
FOREST rClassify, rFeatures, rXgb TREES=500 DEPTH=8 MODE=PROBABILITY
; 4. Route based on complexity
COMPARE rComplexity, 0.4
IF_GREATER rComplexity, deep_path
JUMP fast_path
fast_path:
; Mamba only (~50ms)
CALL mamba_forward rQueryEmb, rMamba, rOut
JUMP synthesize
deep_path:
; Speculative: Mamba draft + Llama verify
CALL speculative_forward rQueryEmb, rMamba, rLlama, rOut
synthesize:
; 5. Generate audio via Moshi
DEPFORMER rCodesOut, rOut, rAudioEmb, rKVDep
CODEC_DEC rAudio, rCodesOut
STREAM rAudio, CHANNEL=4 BLOCKING
JUMP pipeline_loop
Enter fullscreen mode
Exit fullscreen mode
Design Principles
M³-AVM follows five principles that guided every opcode:
1. ISA topology over framework topology
The ISA decides when to shard, when to migrate, and when to rollback. The user code just describes what to compute.
2. Every stateful opcode declares (a) where its state lives, (b) snapshot cost, (c) how ABORT restores it
If a new opcode can't answer these three questions, it doesn't belong in the ISA.
3. Encodings are never reused
Every opcode, feature bit, and region ID is permanent. If a design is deprecated, its number stays reserved.
4. Extensions enter via ESCAPE, never by breaking binary compatibility
The ESCAPE opcode (0xB0) can wrap arbitrary new opcodes with self-describing length. This is the same discipline that kept IBM z/Architecture alive for 60 years.
5. Dual-mode keeps I-cache pressure low
Control flow stays in 32-byte instructions. Rich payloads (cluster, hardware) use 64 bytes. The dispatcher decides by opcode range.
Status: What Works, What Doesn't
Working today (v1.3):
26 opcodes implemented in Rust
Attention, FFN, ROPE, SSM scan
Codec encode/decode, audio alignment
Fork/abort with CoW rollback
GGUF inference with Q4_K
166 passing tests
Specified but not implemented (v2.4):
~230 opcodes from RFCs 0001–0012
Cluster transport (SIGNAL, MIGRATE, BARRIER)
GPU path (wgpu)
Storage tiering
Security (ChaCha20-Poly1305)
Observability (OTLP)
How to Contribute
The project is open source under AGPL-3.0-or-later . Commercial licensing is available for teams that need proprietary use.
Repo: github.com/matheuscamarques/m3_avm
I'm actively looking for volunteers for:
Implementing opcodes in Rust (from spec to tests)
Writing benchmarks with criterion
Reviewing RFCs for design flaws
Testing on Apple Silicon, AMD, RISC-V
Building assembler/disassembler tooling
If any of this interests you, open an issue on GitHub or reach out directly.
Citation
@software { marques2026m3avm ,
author = {Marques, Matheus de Camargo} ,
title = {{M³-AVM: A Distributed Heterogeneous AI Virtual Machine}} ,
year = 2026 ,
version = {2.4} ,
license = {AGPL-3.0-or-later} ,
url = {https://github.com/matheuscamarques/m3_avm} ,
orcid = {0009-0003-4518-2258}
}
Enter fullscreen mode
Exit fullscreen mode
ORCID: 0009-0003-4518-2258
Closing Thoughts
I don't know if M³-AVM will succeed. It might be too ambitious, or the timing might be wrong, or a better architecture might emerge. But I've noticed that the AI industry currently rebuilds its serving infrastructure from scratch every 18 months — new frameworks, new serialization layers, new batching strategies. There's no stable substrate.
This project is an attempt at a stable substrate. If it fails, I want it to fail loudly, with the spec public and the RFCs documented.
If you've read this far, thank you. If you see something wrong, tell me. If you see something interesting, tell your friends.
— Matheus de Camargo Marques
Top comments (0)