DEV Community

Cover image for I Gave ChatGPT My Full Codebase. The Results Scared Me — But Not for the Reason You Think.

I Gave ChatGPT My Full Codebase. The Results Scared Me — But Not for the Reason You Think.

Info Inlet on September 29, 2026

So I did the thing everyone tells you not to do. I took the whole codebase. Every route, the migrations, the config, that utils.ts nobody's touche...
Collapse
 
ai_topics profile image
AI Topics •

The interesting part isn’t that AI can understand a large codebase — it’s how quickly it can expose assumptions and technical debt that humans have gotten used to.

The scary part is realizing how much “context” lives in our heads instead of the code itself. 😅

Great read.

Collapse
 
infoinlet1 profile image
Info Inlet •

That line about context living in our heads is the one, yeah. The AI mapped my architecture better than my own onboarding doc — but only the part that was written down. Everything that made the system actually safe to change lived nowhere: "don't touch that table directly," "the webhook can fire twice," "this endpoint looks dead but a cron still hits it." None of that is in the code. It's in whoever's been here longest.

And here's the twist that got me — the model doesn't know that context is missing. It reads the repo, sees no note saying "careful here," and concludes there's nothing to be careful about. So the exact spots where our head-knowledge was load-bearing are the spots it signs off on most confidently. The silence in the code reads to it as "all clear."

Which is kind of a brutal mirror, honestly. Every "looks fine" it hands you is really measuring how much of your own understanding you forgot to write down. The debt was never just in the code — it's in the gap between the code and the stuff we all just know.

Appreciate you reading it.

Collapse
 
eye_java_420f1faa10ae8b86 profile image
Nan •

The runtime half of that gap makes me think of my opposite case. My tools are all static sites with no backend — no async chains, no distributed system — so there's nothing here to crash-simulate. But I do use runtime probes every day. I just didn't write them.

I hand my production sites to two outside readers as free probes. One is search crawlers: after each page ships, I check Search Console's URL Inspection for how it actually reads the page and whether it got indexed. My newest site's first post was indexed in under 24 hours — faster feedback than any local check. The other is PageSpeed: it loads your page from real edge nodes. It once caught a 177 KB third-party script choking my homepage, and five mobile scores moved from the 80s to near 100. Static reading barely sees that problem, because it only exists under real loading.

What both probes share: they look at the real-world you, not a mirror of your code. AI reading code is still reading what you wrote. Crawlers and speed tools report what you shipped.

So for a small site with no backend, the runtime half doesn't need a homemade simulation. Hand it to the outside readers who are already reading you.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the best extension of the argument anyone's left, honestly. You found the witness I kept insisting had to be a human or a separate model, and it turns out for a static site it already exists and you don't even have to build it.

The thing I love about your two probes is that they both satisfy the one rule that actually matters: independent prior. Search Console and PageSpeed aren't reading your intent, they're reporting your consequences. The crawler doesn't care what your HTML meant to do, it tells you what it actually did when a real indexer hit it. That's exactly the "different seat" I was groping toward — you just realized the seat was already occupied by Google and you were ignoring the guy sitting in it.

And the 177 KB script is the perfect example of a seam bug in your world. Statically, that script is fine — correct tag, loads, no error. The bug only exists in the relationship between your page and a real edge node on a real phone on a real network. No amount of reading the source surfaces it, same way no amount of reading my webhook surfaced the ack-before-persist. The danger was in the frame, not the picture — your frame's just the network instead of a payment provider's retry logic.

The one place I'd gently push: those probes are fantastic witnesses for the dimensions they measure — indexability, load behavior. They're silent on correctness of the thing itself. A page can be indexed in 24 hours, score 100 on mobile, and still show the wrong price. So I'd say you've fully solved the runtime half for performance and reachability, and the "did the content actually come out right" half still wants a human eye or a test. But for the half you're talking about? Yeah — don't simulate it, just go read what your outside readers already wrote about you. That's genuinely sharper than what I said in the post.

Collapse
 
rajanpanwar profile image
Rajan Panwar •

That distinction between a finder and a witness is easily the sharpest takeaway here.

The scariest hallucinations are never the obvious syntax crashes. They are the ones dressed in your own real architecture, using your actual function names, explaining a phantom vulnerability with total executive poise. Worse still is the reverse: an AI giving a green checkmark to an ack-before-persist pattern just because the try/catch block looks polite. It understands your syntax, but it has zero clue how the external world behaves when a socket drops mid-flight.

Treating every positive output as a lead instead of a verdict is the only sane approach.

What is the most convincing, beautiful-looking lie an AI has told you about your own codebase?

Collapse
 
infoinlet1 profile image
Info Inlet •

You nailed the reverse case — "the try/catch looks polite" is exactly the tone. It grades the manners of the code and calls it a security review.

The most convincing lie I ever got: it told me I had a timing-attack vulnerability in my password comparison. Named the function. Quoted the line where I was supposedly using === on the hash instead of a constant-time compare. Walked me through how an attacker measures response times to recover the hash byte by byte. Textbook, genuinely well-written, the kind of finding that makes you feel like you dodged a real bullet.

I was constant-time-comparing. Had been the whole time. It quoted a version of my function that didn't exist — right file, right function name, wrong body. It had basically written the bug it expected to find there and handed it back to me as something it found.

And that's the tell, I think, looking back: it didn't read my code and report a problem. It pattern-matched "password comparison" to "probably a timing attack" and then dressed the prediction up in my real function name. The specificity I trusted — my name, my file — was the cheapest part for it to generate. The one part that would've taken actually reading the code, the function body, was the part it made up.

Thirty seconds in the file killed it. But I'll be honest, for those thirty seconds I believed it completely. That's the gap that scares me — not that it lied, but that nothing in the lie felt different from the three findings right above it that were real.

Collapse
 
danielchinasz profile image
daniel •

What you said is a real story. Once AI finishes reading a codebase, it can genuinely flag issues you never noticed — some pretty serious ones, too. But that doesn’t make it rigorous or trustworthy itself; it’ll happily write serious bugs of its own. Tests have to be even stricter than before.

Collapse
 
infoinlet1 profile image
Info Inlet •

Yeah, exactly. And honestly that's the part that still gets me: the same tool that catches a bug you'd never have spotted will hand you a confident, perfectly-formatted bug two lines later, and on the page they look identical.

Fully agree on stricter tests. One thing I'd add though: if the AI writes the tests too, they inherit its blind spot. It'll happily test the behaviour it already believes is correct. So "stricter" has to also mean independent — tests from a different prior, plus the behavioural/seam ones for bugs that don't live inside a single file (the webhook ack-before-save never shows up in a normal unit test). Otherwise you just end up with greener tests that all quietly agree with each other.

Collapse
 
xiangc_92294 profile image
xiangc •

"The model reads inside the frame you hand it, and the danger was in the frame, not the picture."

That line completely nails it.

LLMs analyze static text, but production outages happen in distributed state transitions. You can give a model your entire codebase, but you can't give it the network partitions, third-party webhook retry backoffs, or the database pool running out of connections at 3 AM.

Static code review (human or AI) verifies syntax and obvious logic. Fault injection and integration testing verify reality. If you want to catch seam bugs, stop asking models to read the code—start simulating process crashes between async lines.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the sharper version of what I was clumsily getting at, so thank you.

"Stop asking models to read the code, start simulating crashes between async lines" is exactly it. The one place I'd still keep the model in the loop is finding where to inject. Point it at the repo and ask "list every spot where we ack, respond, or commit before the durable write, or hold a DB connection across an await," and it's genuinely good at enumerating the candidate seams. Then you fault-inject each one and let reality vote.

Model finds the suspects, chaos testing convicts them. It just never gets to be the judge.

Collapse
 
morgan_todd99 profile image
MorganTodd •

I especially like the point that giving AI a full codebase can reveal problems you’ve simply stopped noticing after working with the same project for a long time.

Collapse
 
infoinlet1 profile image
Info Inlet •

Yeah exactly. after a while you just stop seeing certain things. that dead admin endpoint and the old .env in git history were both stuff i technically knew about but had gone completely blind to. the model doesn't have that. it reads the 300th file same as the first one.

only thing is that same freshness is why i don't trust it when it says something's fine. it catches what im numb to, but it'll also make stuff up with the same confidence. so i keep whatever it flags and ignore whatever it clears.

Collapse
 
elijahbrown profile image
Elijah Brown •

The signup-vs-webhook race on users is the useful finding. When you invent accounts to prove the fix, use reserved fiction phones (US 555-0100 to 555-0199, UK 020 7946 0xxx) and domains you control, so a leftover dump from the repro cannot publish a real person.