DEV Community

Cover image for Reverify weighs verified AI claims by how informative they are
Reno Lu
Reno Lu

Posted on

Reverify weighs verified AI claims by how informative they are

My reading of the reverify README is that its central idea is not the verifier. It is the admission that "every claim verified" is a score a model can reach while saying nothing. Assert that a file starts with MZ and that .text exists, and you get a clean sheet that says little. Reverify's response is to weigh each verified claim by how much it tells you, and that is the part I would point a skeptical reviewer to first.

The loop the name refers to

The project's pitch is "Stop your AI from making things up." The mechanism behind that pitch is a division of labor. A language model proposes a claim about an artifact. A deterministic tool checks the claim against that artifact and returns VERIFIED, REFUTED, or INCONCLUSIVE, along with the bytes it observed. In the README's words, the model never gets to assert a fact on its own.

The README picks binary reverse engineering as its proving ground, saying the hallucination problem in binary analysis is far worse than in source code. The toolkit covers PE/ELF/Mach-O parsing, x86/x64/ARM/ARM64 disassembly, AOB pattern scanning, CPU emulation, Protobuf/TLV dissection, and Frida hook generation, in pure Python out of the box. Installing reverify[full] upgrades it to capstone, unicorn, lief and Z3, and reverify[angr] adds angr for function boundaries, the call graph and cross-references. Without those engines it falls back to the pure-Python core, and reverify backends shows what is active.

A claim is a small JSON object. The README's first example asks whether the instructions at offset 4096 are push, mov, sub, noted as a function prologue. Another hands raw x86 bytes to the emulator and expects eax to equal 8. Claims can be batched with --claims-file claims.json, and the CLI exits non-zero if anything is refuted, so an agent or CI job can gate on the result. A depends_on field lets a refuted root invalidate the claims built on it, and "observe": true has the tools read a value instead of asserting one.

Weight, not a tally

Each result carries a weight. The README sets it to zero for claims that restate the fact sheet the model was shown, for duplicates, for inline code or data that does not occur in the binary, and for echoes of the tools' own previous output. Otherwise the weight is measured from the binary itself: how often the expected content occurs in the file and how much entropy it has. Zero padding or a ubiquitous prologue will verify and still weigh almost nothing. A reconstruction counts as grounded only when nothing is refuted and the verified weight reaches --min-information, which defaults to 1.0. The README says this follows the CORE refinement of FActScore.

reverify reconstruct --samples N carries the idea into generation. It draws several proposals per round and lets the verifier, not the model's confidence, select among them.

A plain pass/fail gate leaves a loophole open: a verified set can contain safe claims that say little. The weight rule is the README's answer to that, and it is a pattern worth considering for harnesses that grade model output.

The numbers, and where they come from

The README reports that on 71 real Windows system files, the AI's textbook answer was wrong 97% of the time, and that reverify caught every one and never accepted a wrong claim. It also cites a per-claim-kind confusion matrix: 0 false VERIFIED out of 475 known-false claims, and 0 known-true claims missed, gated in every CI job. According to the README, the benchmark runs in CI on every push against each platform's own system binaries and fails the build if a single wrong claim is VERIFIED, and a third-party aarch64 replication is documented in BENCHMARK.md. These are the project's figures. The README describes a replication package with one command per benchmark and a pinned Dockerfile for anyone who wants to check them.

Output from reverify verify --json also includes the binary's SHA-256, the reverify version and which engines judged, so a report can be handed over and replayed.

Outside the binary

Two features reach past reverse engineering. reverify equiv <reference> <candidate> --lang python (or C) runs a candidate implementation and a reference over shared inputs and checks that they agree. A refutation comes back with the input and both outputs, which puts an AI's rewrite under the same verdict structure.

The second is context. reverify rollover hands a session off to a file and starts a fresh one instead of leaning on a lossy auto-summary, and the README lists Claude Code, Codex, Gemini CLI and OpenCode as supported. Since v0.8.0 the loop also writes a ledger per binary at .reverify/ledger/<sha256>.json, checkpointed after every round. Refutations come back as KNOWN FALSE, so a fresh context does not re-propose the same wrong prior. The README's reasoning is that the model's unverified prose was never trusted, so dropping it loses nothing.

Reverify ships as an MCP server and a plain CLI. The README scopes it to authorized reverse engineering: malware analysis, CTF, interoperability research, and software you own or are permitted to analyze.


GitHub: https://github.com/2akouwu/reverify


Curated by Agent Palisade — practical AI for small and mid-sized businesses.

Top comments (0)