Standard Transformer attention has a memory problem at scale. For a 1.05 million-token context with 28 layers, 2 KV heads, and a head dimension of 128, the FP16 Key-Value cache alone requires:
2 × 28 × 2 × 128 × 1,055,402 × 2 bytes ≈ 28.2 GiB
That is before loading a single model weight. On a 40 GB A100, the KV cache would consume over 70% of total available memory. On consumer GPUs, it simply crashes.
ISOM-R2 (Isometric State Space / Virtual SVD) is a new-gen recurrent memory architecture built to solve this directly. Instead of maintaining a linearly growing KV cache across every token, ISOM-R2 processes context through bounded isometric state projections, keeping GPU memory flat regardless of sequence length.
The Benchmark
I ran a full audited benchmark using Aazhi-Coder-1.5B, a 1.5B parameter coding model powered by ISOM-R2, built on top of Qwen2.5-Coder-1.5B-Instruct.
The corpus was real, unmodified Python source code:
181 production source files cloned directly from the official Hugging Face transformers repository
Zero synthetic tokens. Zero artificial padding.
Assembled into a continuous prompt of 1,055,402 tokens (516 discrete 2,048-token chunks)
Here are the results:
Metric Value
Total Tokens Ingested 1,055,402
Source Files 181 (real transformers repo)
Total Chunks 516
Total Runtime 123.78 seconds
Effective Throughput ~8,526 tokens/sec
Prefill VRAM Flat at 2.97 GB across all 1M tokens
Peak Execution VRAM 3.24 GB
Memory vs. Standard KV Cache 8.7× lower than 28.2 GiB requirement
Macro Retrieval Precision 100%
Prefill VRAM stayed at 2.97 GB throughout the entire million-token ingestion. Generation only raised peak to 3.24 GB — well within the capacity of consumer-grade 4 GB and 6 GB GPUs.
How ISOM-R2 Works
- Bounded Streaming Ingestion
Context is streamed in 2,048-token chunks through a bounded state-space update. No token ever materializes a full KV cache entry in GPU memory. The active GPU buffer stays under 400 MB during prefill, with the model state compressed into a fixed-size isometric manifold.
- Sub-Harmonic Lie Frequency Calibration
Standard positional embeddings lose coherence past 32k to 128k tokens due to rotational aliasing. ISOM-R2 sets a sub-harmonic frequency floor:
ω(min) < 2π/1,048,5762 ≈ 5.99×10−6 rad/token
This guarantees that the slowest coordinate in the state manifold completes less than one full rotation across the entire 1,048,576-token sequence, preserving stable temporal geometry throughout.
- Salient Micro-Window Retrieval
At generation time, instead of decoding from the full context, ISOM-R2 scores all 516 chunks for relevance and pages only the highest-scoring chunks into active attention. Each salient chunk is further sliced down to a 1,024-token anchor-centered micro-window.
In this benchmark run:
Chunk 218 (configuration_llama.py) — sliced at [518:1542]
Chunk 217 (modeling_llama.py) — sliced at [498:1522]
Active KV at generation time: 4,160 tokens
Out of 516 candidate chunks across 181 repository files, the engine isolated exactly the two target files with zero distractor chunks retrieved. Macro retrieval precision: 100%.
Honest Scope
ISOM-R2 is a retrieval-based architecture. It does not attend over all 1M tokens at generation time — it streams them efficiently into a bounded state, then pages the most relevant context back for decoding.
The memory and throughput results are measured and verifiable. The benchmark notebook is published with all cell outputs intact.
What this proves:
Streaming 1M real tokens on under 3.5 GB VRAM is achievable with bounded recurrent state spaces.
Retrieval precision at scale is viable — isolating 2 files from 181 without distractor noise.
The Proof
The fully executed benchmark notebook is publicly available on Hugging Face with all cell outputs, terminal logs, and VRAM measurements preserved:
Benchmark Notebook (ISOM-R2): https://huggingface.co/Prannesshkva/Aazhi-Coder-1.5B/blob/main/ISOM_R2.ipynb
Model Weights & Implementation: https://huggingface.co/Prannesshkva/Aazhi-Coder-1.5B
Open the notebook, inspect the output cells, and reproduce the measurements yourself.
Built on Qwen2.5-Coder-1.5B-Instruct by Alibaba Cloud (Apache 2.0). ISOM-R2 engine and Aazhi-Coder weights released under CC-BY-NC-ND 4.0.
Top comments (0)