Last week Debian shipped a single kernel security update that fixes 1,313 CVEs. The Register's take was blunt: "We strongly suspect that this number of CVEs is due to LLM bots doing the bug hunting."
Immunefi called it the Vulnerability Apocalypse. CVE submissions rose 263% between 2020 and 2025, the first quarter of 2026 was another 33% higher, and NIST is only giving full analysis to 15 to 20% of what comes in.
Smart contracts aren't an exception. In Anthropic's SCONE-bench research, AI agents wrote working exploits for 207 of 405 contracts that had been hacked in real life, and found two new zero-days in 2,849 freshly deployed contracts. The average cost of an agent run was $1.22.
My read
Finding bugs used to be the expensive part of security. Now it's close to free, and it's free for whoever runs the scan first. If you deploy contracts, the quiet window after launch is gone. Someone can point an agent at your code the hour it goes live, for less than a coffee. "We'll audit once we see traction" was always a bet. It's a worse bet now.
I don't think this is doom, though. Defenders get the same tools, and your code sits in your repo for weeks before an attacker sees it. What got scarce is judgment. AI produces a lot of findings that sound plausible, and deciding which ones are real exploit paths in your system, with your trust assumptions, is now the expensive part.
We saw it first hand: my team placed 2nd of 98 researchers in Immunefi's Quantus audit competition, and everyone had the same models. The difference was where we pointed them and what we threw away.
What changes if you're building
- Scan before anyone else does, and keep scanning. Treat an AI security pass like your tests: run it on every meaningful change, not once the week before launch. Several tools can run that pass from your editor's AI now, including free ones.
- Write your threat model down. A finding is only real relative to who you trust and what must never break. Without that page you can't triage AI output, and neither can your auditor.
- Assume your contracts get scanned on day one. Launch with guardrails: deposit caps, monitoring on privileged calls and solvency, a bounty, and a plan for the bad day. An audit covers a commit, not what happens after it.
If you're deploying in the next three months
Three things I'd have ready before mainnet:
- A clean AI pass on your code. Fix the obvious things now, while it's still cheap to change.
- A one-page threat model. Where the money moves, who is trusted, and your single riskiest assumption.
- A launch plan with guardrails. Caps you can raise later, a pause you've actually tested, alerts that reach a human, and SEAL 911 saved somewhere you'll find it.
None of this needs a big budget. It needs someone to do it before the code ships.
What does your pre-deploy checklist look like now that scanning is this cheap? Curious what other teams have changed.
Originally published in my newsletter, DeFi Protocol by code.
I wrote this with help from AI for drafting and editing. The opinions and the audit experience are mine.
Top comments (1)
The Quantus result is the sharpest number here — second of 98 with everyone running the same models means the scarce input was exactly what you named: judgment about where to point and what to discard. Which makes triage a receipts problem in disguise: a finding is a claim, and findings grade by their verification, not their plausibility. The SCONE data shows the shape — the same agents that produce plausible findings can produce working exploits, so "replayed" is a property a finding either has or lacks: scanner proposes, exploit replay validates, human decides — three events from three writers, and a finding that can't carry a reproducible PoC is a lower grade of evidence: worth tracking, never gating. Your Debian drop is that at database scale — NVD filling with ungraded claims because verification is the part that didn't get cheap; the volume multiplied and the denominator of verified findings didn't move.
To your checklist question, one addition: treat the threat model as a sealed artifact, not a page. A finding is only real relative to trust assumptions, so the triage decision — accepted, mitigated, false positive — is only reviewable if the assumptions it was judged against are pinned next to it and versioned with the code. That's what turns "we knew about this class" from an allegation into a record, and it's the cheapest item on your list: it's one commit with a hash.