DEV Community

Nenad Mićić
Nenad Mićić

Posted on AI-assisted

The same tiny GPT in SQL, PostScript and Brainfuck, byte for byte

This post continues the int-llm hobby experiments from my earlier post. The earlier repositories are int-llm (Q16.48 fixed-point training and TinyLlama inference), int-llm-precision-ladder (reduced stored weight precision checked against the Q16.48 reference), int-llm-coordinate-permutation (reversible coordinate permutations of Llama checkpoints), and int-llm-viz (visualization of the integer GPT weights and inference path).

Those projects produced a C program whose output is deterministic down to the last byte. This post describes three new repositories that reproduce that output in SQL, PostScript, and Brainfuck.

The reference and the gate

The model is Andrej Karpathy's microgpt: a character-level transformer with 1 layer, 32 embedding dimensions, 4 attention heads of 8 dimensions each, an MLP width of 128, a context length of 8, and a vocabulary of 26 lowercase letters plus one BOS token. It has 14,272 parameters and was trained on a list of names.

The int-llm C implementation performs every operation in Q16.48 fixed point: a signed 64-bit integer with 48 fractional bits. RMSNorm, the attention projections, softmax, exp, inverse square root, and the xorshift64 sampler are all integer code. Because no floating-point operation participates, the output depends only on the weights and the RNG seed. In this post, "the oracle" means that C program, and "the gate" means the comparison of a port's output against it.

The oracle loads a 115,576-byte MGW checkpoint, runs 122 forward passes, and prints 20 sampled names (kayla, daia, lee, …, karin). Each port must print the same bytes. Each port also computes two FNV-1a checksums inside its own code, one over every raw logit word and one over every sampled byte, and both must equal the pinned values. The checksums detect regressions; they are not cryptographic integrity checks.

A deliberately corrupted checkpoint serves as a negative control. The corrupted model prints the same 20 names but a different logit checksum (b10c3a08150e95c6 instead of 0610f72f01c199cb). A port that passes the gate with the canonical model and fails it with the corrupted model is comparing arithmetic, not only text.

int-llm-sql

microgpt.sql is a 1,410-line SQLite script. It runs with:

sqlite3 :memory: < microgpt.sql
Enter fullscreen mode Exit fullscreen mode

The script calls readfile() once to load the committed model, decodes the little-endian weights into tables, runs the forward pass and the sampling loop, prints 20 names, and compares its own checksums and step count against the pinned values before it prints SQL_GATE=PASS. No stored procedure, user-defined function, loadable extension, or floating-point value participates. The repository tests the script with SQLite 3.51.0; the shell must support recursive CTEs, generated columns, window functions, ordered aggregate arguments, and readfile().

SQL lacks four things the forward pass needs, and the script supplies each:

  • Sequencing. SQL has no loop. An insert trigger fires the next forward pass. The script delivers 160 ticks; 122 of them reach a forward pass before BOS tokens end the samples.
  • 128-bit integers. SQLite integers are signed 64-bit. The script splits each Q16.48 multiplication into limbs and implements division, exp, and inverse square root as recursive CTEs.
  • XOR. SQLite has no XOR operator. The script computes XOR(x, y) as (x | y) - (x & y).
  • Overflow detection. SQLite silently promotes an overflowing + or * to REAL. The limb arithmetic keeps every intermediate in range, and SUM() serves as the accumulator so that an overflow fails loudly.

Matrix multiplication needs no workaround: it is a join followed by an integer SUM(). During development, the SQL implementation was checked against the oracle at all 69,748 recorded activation, logit, and probability rows and at all 122 sampling records.

int-llm-postscript

microgpt_infer.ps is a 512-line PostScript program. Ghostscript runs it in about 2 seconds:

gs -q -dBATCH -dNODISPLAY -dNOSAFER microgpt_infer.ps
Enter fullscreen mode Exit fullscreen mode

PostScript provides floating-point exp, ln, and sqrt. The program uses none of them; every operation is integer add, mul, idiv, bitshift, xor, or and. The program parses model.mgw directly (header, config, tensor index, little-endian two's-complement payloads), runs the 122 forward passes, and prints the same 485 bytes as the oracle: the Machin-series Pi banner (Pi = 3.141592653589782, including the oracle's integer-truncation error in the last digits) and the 20 names.

microgpt_train.ps is a 1,090-line trainer. It initializes the model, runs the whole-sequence forward and backward pass, applies Adam with gradient clipping and bias correction under a CORDIC cosine learning-rate schedule, writes an MGW v1 checkpoint, and samples 20 names. All model and optimizer state is Q16.48. On Ghostscript 10.04.0, 5,000 steps took 3,712 seconds (about 62 minutes). The 5,000 loss lines, the 20 samples, and the 115,576-byte checkpoint matched a fresh C run byte for byte, and the checkpoint matched the committed model.mgw at the same SHA-256.

The PostScript-specific hazard: Ghostscript's mul returns a real when the product overflows 64 bits, instead of raising an error. The program therefore splits each wide product into 24-bit limbs with a carry chain, and splits division remainders into two 31-bit limbs, so that no intermediate exceeds 2^63 − 1. If any limb were wrong, the gate would fail on the first affected sample.

int-llm-brainfuck

This is the one I run once and never again, so you don't have to.

Brainfuck has eight commands: > < + - . , [ ]. A cell holds one byte. The language has no multiplication, no comparison, and no conditional other than [ ], which repeats its body while the current cell is nonzero.

The repository pins its own machine profile, because Brainfuck implementations disagree on cell width, EOF, and tape size: 4,194,304 zero-initialized cells, 8-bit wrapping arithmetic, binary stdin and stdout, pointer underflow and overflow as errors, and , writing zero at EOF. Each Q16.48 value occupies eight little-endian cells.

Python generators emit the program. A generator may expand macros and allocate static tape regions; it may not precompute logits, samples, checksums, losses, or checkpoints. The interpreter executes commands, enforces tape limits, and counts executed commands; it does not parse the model or recognize tensors. All model work happens in the eight commands.

The generated programs are not committed, because they are too large for an ordinary Git repository. The repository holds the generators, the tape map, the oracle corpus, the fixtures, and the verification logs. make regen and make training-full-program regenerate the programs.

Phase 1, inference:

  • microgpt.bf: 16,395,194,474 bytes (16.4 GB)
  • 46,883,492,331,121,926 source commands executed (about 47 quadrillion)
  • 70,843 seconds (about 19.7 hours) for 20 names
  • sample and logit checksums equal the pinned values; the corrupted checkpoint produces the expected distinct logit checksum

Phase 2, training, 5,000 steps:

  • microgpt_train.bf: 210,300,828,662 bytes (210 GB)
  • 484,906,588,575,174,327,064 source commands executed (about 485 quintillion)
  • 409 hours 29 minutes wall time on the Linux x86-64 runner, including cold AOT compilation
  • all 5,000 loss lines, the 20 samples, the wide checkpoint, and the F12 checkpoint equal the oracle's output byte for byte

The trainer also computes Pi inside Brainfuck with the oracle's integer Machin series and passes the value into the CORDIC cosine schedule, because the oracle does the same and the gate accepts nothing else.

A full run is not required to check the repository. make bootstrap verifies the pinned artifacts and runs 39 unit tests in about one second. make test-model-validation generates a 2 MB program that loads and validates the model image; it passes in about 13 seconds. make TRACE_STAGE=logits test-model-trace generates a 3.2 GB program that runs one position through the complete transformer and compares its 27 logits against the oracle; with the optimized interpreter it executed 166,850,840,789,467 source commands in 18 seconds and passed. One detail: the model image is 29,944 bytes, and the generated program contains exactly 29,944 , commands to read it.

The verification log keeps failures. The first exact stop-after-20 training run stopped with a pointer underflow on both runners. The cause was uncleared Adam serial state overlapping the next forward pass's row-copy controls. The fix clears that workspace and has a regression test. The log records the failed run as FAIL next to the fixed rerun; it does not relabel it.

The full Brainfuck training run took 409 hours. I ran it so that “it should work” could become “it did work.”

Human code review is welcome. The generated Brainfuck program is 210 GB, so reviewers wishing to avoid AI assistance should probably clear their calendars for the next few centuries.

Why

These are hobby projects, done for fun. The question behind them is how far a transformer can go, not whether it should.

Integer-only arithmetic is what makes the question answerable. A floating-point port of the same model would produce slightly different rounding in each host, and "close enough" has no gate. With Q16.48, every host either reproduces the oracle's bytes or fails on the first divergent sample. That strictness turns an unusual host into a precise test of what the computation needs: 64-bit integers, a logical right shift, a loop, and nothing else. SQL supplied the loop with a trigger. PostScript supplied the wide multiply with 24-bit limbs. Brainfuck supplied everything from eight commands and a tape of bytes.

The original int-llm project asked whether fixed-point integer arithmetic can replace floating point in training and inference, and int-llm-precision-ladder asked how little precision the stored weights can keep before the output changes. These three repositories ask how little the host language can offer. The answer so far is that the floor is low, and the gate is the reason the answer is exact.

Every result is reproducible from the repositories with a compiler, the pinned tool versions, and time.

These projects were developed and tested with AI assistance. The correctness claims do not rest on that; they rest on the byte-for-byte comparisons recorded in each repository.

Lesson learned: trust AI-assisted work through validation, not through the assumption that a human reviewed every generated line.

Repositories: int-llm-sql, int-llm-postscript, int-llm-brainfuck. Everything else is at github.com/nmicic.

The next one will be something different.

Top comments (0)