Here's the question worth asking about your audit log: if someone with write access to your database rewrote the whole thing tonight, what would catch it?
If the log is hash-chained and the chain lives only in that database, the honest answer is nothing. That's not obvious at first, so let's work through it.
The chain itself
The standard move is to make each record commit to the one before it. Every entry stores the previous entry's hash, and its own hash covers both that pointer and its contents. Change any record and every hash after it stops matching.
import hashlib
import json
GENESIS = hashlib.sha256(b"genesis").hexdigest()
def record_hash(prev_hash, record):
body = json.dumps(record, sort_keys=True, separators=(",", ":"))
return hashlib.sha256((prev_hash + body).encode("utf-8")).hexdigest()
def append(log, record):
prev = log[-1]["hash"] if log else GENESIS
log.append({"record": record, "prev": prev, "hash": record_hash(prev, record)})
def verify_chain(log):
prev = GENESIS
for i, entry in enumerate(log):
if entry["prev"] != prev or entry["hash"] != record_hash(prev, entry["record"]):
return i # first broken entry
prev = entry["hash"]
return None # chain is consistent
Quick note on the json.dumps line, because it matters more than it looks. sort_keys=True and the fixed separators give you a canonical serialization. Without them, the same logical record can serialize to different bytes depending on dict ordering or whoever re-serializes it later, and your verifier ends up reporting tampering that never happened. A hash commits to bytes, not meaning, so pin the bytes down.
Okay. So now if someone edits record 3 in place and leaves everything else alone, verify_chain returns 3. Tamper-evident. Done?
No.
Where it falls apart
The person who can edit record 3 can also run append in a loop. Here's the whole attack:
def rewrite(log, index, new_record):
records = [e["record"] for e in log]
records[index] = new_record
rebuilt = []
for r in records:
append(rebuilt, r)
return rebuilt
Feed the result to verify_chain and it returns None. Perfectly consistent. Every prev points at the right hash, every hash recomputes. The chain proves the log agrees with itself, and a freshly forged log agrees with itself just as well as the original did.
That's the step that's easy to skip. A hash chain proves internal consistency. It doesn't prove the log you're looking at is the log that existed yesterday. For that you need a value from yesterday that the rewriter couldn't touch.
Being an admin doesn't save you either. The threat model for an audit log of admin actions includes the admin. If the integrity check runs entirely on infrastructure the admin controls, it's checking the admin's homework with the admin's answer key.
The fix is small: publish the head
Look at what the chain gives you for free. The hash of the latest entry, the head, commits to every entry before it. Change anything earlier and the head changes. So you don't have to protect every record. You protect one value.
Take the head hash at some point and put it somewhere outside your control, somewhere the rewriter can't edit after the fact. From then on, anyone holding that published value can take your log, recompute the chain up to that entry, and compare. Match means the prefix is the same one that existed when the head was published. Mismatch means something before that point changed, and the forger can't rebuild their way past it, because the published value doesn't move.
What does "somewhere outside your control" mean in practice? It means a place where you can't change the value after the fact and a third party can check it without asking you. Emailing it to yourself doesn't count. A row in a different table in the same database definitely doesn't count.
This is the thing I built ProofLedger to do. You submit a SHA-256 hash and it gets anchored to Polygon and Bitcoin. Only the hash goes on chain. The log itself never leaves your systems, which matters when the log is full of user IDs and admin actions you'd never publish.
Anchoring the head looks roughly like this: compute it, then POST it to /api/v1/proof with your sk_ key as a Bearer token. The part I care about more is the other direction, checking it:
HEAD=$(python -c "import mylog; print(mylog.load()[-1]['hash'])")
curl "https://proofledger.io/api/v1/verify?hash=$HEAD"
That verify endpoint is public, no auth, rate limited to 120 requests per hour per IP. I made that call on purpose. If an auditor has to log into the issuer's account to check the issuer's claim, you're back to the answer-key problem. The verify-proof package on PyPI does the same check offline and locally, so the verifier doesn't even have to trust my server being up.
What anchoring doesn't buy you
I'd rather you know the limits up front than find them later.
The window after the last anchor is unprotected. Anything appended since the most recent published head can be rewritten or dropped and nothing external will notice. How often you anchor sets how big that window is. That's a decision you make for your threat model, not something anchoring decides for you.
Truncation looks a lot like "nothing happened." If someone deletes the tail after your last anchor, the remaining log still verifies against the anchored head. To catch that you need something that says how long the log should be, like the next anchor or an independent record that entries were written.
It proves "unchanged since," not "true when written." If the code writing the audit entry logs the wrong actor, the chain will faithfully preserve the wrong actor. Integrity isn't accuracy.
Your serialization is part of the protocol now. Change the json.dumps arguments in a refactor and every historical head stops reproducing. Version it, test it, treat it like a wire format.
And the anchor itself isn't magic. Anyone can timestamp a hash on Bitcoin for free with OpenTimestamps, and that proof is just as cryptographically valid. What you're picking between is the tooling around the proof, not a stronger proof.
The test to run on your own setup today
Take your audit log, or any hash-chained thing you own: event store, migration ledger, release manifest. Then do this:
- Pick an entry from the middle, change one field, and rebuild the chain from that point forward, the way
rewritedoes above. - Run whatever integrity check you have against the rebuilt log.
- If it passes, ask yourself one thing: what value, stored outside the infrastructure you just used to forge it, could I compare this head against?
If you can't name one, what you've got is a checksum. Useful for catching bugs, not for catching people. If you can name one, write down who besides you could fetch it and run the comparison without asking you first. That list is the real answer to who can trust your log.
ProofLedger is where I run the anchoring side of this, at proofledger.io. API docs are at proofledger.io/api.html if you want to see the three endpoints.
Disclosure: this article was drafted by an AI agent I built and run, from facts I supplied about my own project.
Top comments (0)