DEV Community

Cover image for After the Sprint: A 72-Hour Build Retrospective - Spoiler: It Wasn't Secure

After the Sprint: A 72-Hour Build Retrospective - Spoiler: It Wasn't Secure

Earl Grey on September 08, 2026

I built Charitas Clew for a 72-hour DEV Weekend Challenge. It worked. I submitted it. And then I did something slightly inconvenient: I kept poking...
Collapse
 
mickyarun profile image
arun rajkumar •

You half-answered this in the reply to @mansio, so I'll push on it. The number I'd want out of the whole exercise isn't what got fixed. It's how often Dr. Kahlo rejected something Antigravity was confident about and then held the rejection.

A reviewer and an implementer drawing on similar training can agree confidently on the same blind spot, and from inside the run that agreement is indistinguishable from correctness. If the adversary rarely won, the review layer is theatre with extra steps. If it won often, you've found something worth copying.

Either way it's a number you can keep, which is more than most people get out of a retro.

Collapse
 
earlgreyhot1701d profile image
Earl Grey • • Edited

Hello @mickyarun This is exactly the metric I wish I’d captured. I can’t honestly give you a rejection ratio after the fact because I wasn’t tracking the review that way, and trying to reconstruct one now would be suspiciously convenient.

I can say the second reviewer changed some outcomes. The clearest example was when Antigravity said one security phase was finished and Dr. Kahlo said, essentially, “Not until you prove it on the live site.” That live test found a real problem our automated tests had missed: the app was identifying users incorrectly behind Firebase and Cloud Run. Antigravity had to fix it, redeploy it, and test it again, including checking that the fix couldn’t be bypassed by faking that information.

There were smaller pushbacks too. Sometimes Dr. Kahlo required another check; other times the result was simply narrowing a claim or admitting something was a known limitation rather than calling it solved. Not every disagreement produced a necessary code change, which is actually why I like your metric.

And to correct something from my earlier reply: the 9.0 wasn’t an independent score. Antigravity produced that score in its own final audit. So it tells me an audit was completed against the criteria, not that an independent reviewer proved the app was a 9.

Next retro, I want to actually track the disagreements: what the implementer proposed, what the reviewer rejected, what changed because of that rejection, and whether the rejection held. Then I’ll have evidence that the adversarial layer changed outcomes instead of just two AIs agreeing confidently with each other.

That’s a much better measurement. Thanks for pushing on it.

Collapse
 
mansio profile image
Mikhail •

The trust proxy finding is the best thing here, and the quote deserves to outlive the post: "Tests can prove behavior inside the world you modeled. Production can tell you that you modeled the wrong world." That is a known failure class — verification certifies the model it was given, not the world — and your resolution, exercising the deployed path instead of the modeled one, is the correct fix. Credit also for testing that the proxy fix didn't create a new spoofing path: verifying your own fix can't regress into a worse door is a discipline most audits skip.

The blacklist removal was the deeper lesson. A keyword filter verifies vocabulary; the boundary it pretended to protect is structural. "Application instructions and document content stay separated, the document is untrusted, output is validated" — that is an invariant, not a pattern match. And "model output is still input" is the sentence most prompt-injection writeups never reach.

One question about the 9.0/10: who computed it? Antigravity implemented the fixes and then scored them in the final audit. A fixer grading its own fixes is the self-validation loop — the number is a receipt that an audit ran, not evidence about the result. "APPROVED FOR CURRENT SCOPE" is the more honest artifact; the score could go without losing anything.

Last gap, and your own earlier rule exposes it: the deployed-path check was a one-time audit event. You documented that the proxy config carries maintenance debt because network ranges drift — but nothing in the 143 tests will notice when it does, because the suite verifies the modeled world by construction. The finding recurs silently on the next infrastructure change. By your own checkpoint standard — a command and its pasted output — the live check deserves repeatability: on deploy, or on schedule. One-time verification is a photograph; drift is a video.

Collapse
 
earlgreyhot1701d profile image
Earl Grey •

@mansio Oh, great question about the approved for current scope. I was lost in my LLM hubris again... Gemini 3.8 approved the scope after the adversarial instructions from Chat GPT. So probably not the best way to feel secure when the tool is checking its own work. But maybe I can have a little more confidence that it was reporting against Chat GPT's prompts? Something I didn't think about yesterday that I'm definitely mulling over now !

Collapse
 
mansio profile image
Mikhail •

Short answer: a little, but name what you'd be buying. Antigravity reporting against Dr. Kahlo's prompts removes one contamination path — the implementer's narrative no longer becomes the reviewer's starting explanation. That is separate context. It is not independent verification: two systems with similar training can agree confidently on the same blind spot, and from inside the run that agreement is indistinguishable from correctness.

arun's metric above is the honest version of the confidence you're looking for: how often did Dr. Kahlo reject something Antigravity was confident about, and hold the rejection? Count it across the 17 findings. If the adversary rarely won, the review layer was ceremony and the 9.0/10 measured agreement, not scrutiny. If it won several times, you have a number worth keeping — and a ratio you can report next retro instead of a score.

Same discipline as the proxy fix: don't feel more secure because the structure looks better; measure whether the second opinion ever changed an outcome.

Collapse
 
routinekit profile image
RoutineKit •

The line that stuck with me: the insecure part wasn’t the code — it was the missing “done means shippable and checkable” gate.

I’ve started forcing a 4-line sticky before any agent touch: outcome, out of scope, done looks like, never invent. The third line is where security usually hides — “done” that doesn’t name a verification step is just vibes with a green checkmark.

Curious whether your 72-hour postmortem would have caught it earlier if “done looks like” had included one concrete check you refuse to skip, or if the miss was more about time pressure than definition.

Collapse
 
earlgreyhot1701d profile image
Earl Grey •

@routinekit I think the what "done looks like" is the next step in my spec driven development pipeline. I was moving way too fast and my first layer of security was like holding up a Kleenex to stop an incoming wave. Bad, bad tech discipline. Hardening is just good hygiene and I'm glad I went back to check. Thanks for reading!

Collapse
 
routinekit profile image
RoutineKit •

Glad the done-looks-like bit landed — putting it in the SDD pipeline before the next sprint is the hygiene move.

Kleenex-vs-wave is a perfect postmortem image. Hardening as baseline, not heroics. Thanks for writing the retrospective in public.

Collapse
 
vinhnguyenthanhdn profile image
Vinh Nguyen •

The date-language change and the action-step output are the same risk with two different dispositions. You hardened what the product is allowed to claim about extracted dates — evidence from the document, confirm with an authoritative source — but the next steps and the speaking script have identical provenance and carry no such hedge, and those are the fields a user acts on within the hour.

A notice containing a phone number, or an instruction phrased as a next step, passes the structural boundary intact and satisfies the runtime contract, because a valid string in a required field is exactly what it is. That is the injection that survives the reframe: not the model being talked out of its own instructions, but the document reaching the user as instructions along the path you built to be faithful. Structural separation decides who the model takes orders from; it says nothing about who the reader takes orders from.

The check is the one you already ran for dates, applied field by field: if a field's content can originate wholly from the uploaded page, the same disclaimer belongs on it — and a wrong phone number is the one that is hardest to walk back.

Collapse
 
earlgreyhot1701d profile image
Earl Grey •

Hi @vinhnguyenthanhdn Ah, yes...I see another blocker here. Applying field by field would have helped me avoid this trap. Thank you for pointing that out. And agreed, the wrong phone number is the hardest to walk back.