DEV Community

Cover image for I Gave ChatGPT My Full Codebase. The Results Scared Me — But Not for the Reason You Think.
Info Inlet
Info Inlet

Posted on

I Gave ChatGPT My Full Codebase. The Results Scared Me — But Not for the Reason You Think.

So I did the thing everyone tells you not to do.

I took the whole codebase. Every route, the migrations, the config, that utils.ts nobody's touched since 2023. Dumped all of it into a model with a context window big enough to swallow the entire repo in one go, and asked it one question:

"What's wrong with this?"

Now, everyone's first reaction to that is the security angle. "You pasted your source into a chatbot?" Fair. I'll get to it, and yeah, it matters. But that's not the part that got to me.

What got to me was the answer. It came back as a clean numbered list, and I sat there realizing I couldn't actually tell which items were true and which weren't. They all read the same.

Let me walk through it the way it happened.

First it impressed the hell out of me

I expected file summaries. What I got was my architecture, described back to me. It knew auth ran through one middleware. It knew two of my services were quietly sharing a table they had no business sharing. It knew the payment webhook and the signup path both wrote to users from totally different places in the code.

Honestly? It mapped the system better than the onboarding doc I'd written for new hires. Faster, too.

If you've never tried this, do it once just for this. A model sitting on your entire repo can answer "where does X happen, and what breaks if I change it" better than half the people who actually work in the thing. That part is real, and it's worth your time.

Keep that in mind though, because it's exactly what set me up.

Then it found bugs. Real ones.

This is where I got excited. It started surfacing actual problems:

  • An old .env I'd committed back in 2022 and deleted the same week. Still sitting in git history, of course. Forever. I knew that in the abstract. I had completely forgotten it in practice.
  • A race between the webhook handler and the signup path. Both upsert the same row, no lock. Under a retry, one clobbers the other. Nobody had hit it yet, but it was absolutely there.
  • A dead admin endpoint. Unhooked from the UI two years ago, still mounted, still skipping the newer permission check. Reachable. Forgotten.

These weren't lint warnings. These were "how did I ship that" bugs. For a solid ten minutes I thought I was writing a very different post, the one titled "just give ChatGPT your codebase already, it's incredible."

And look, as a finder, it kind of is incredible. It's read more code than I ever will and it doesn't get tired on file number 340. I'm not going to pretend otherwise.

Then it started making things up. In the exact same voice.

Item 7 was a SQL injection in one of my functions.

That function doesn't exist.

It had everything going for it, though. Plausible name, plausible file, a nice confident description of how an attacker would exploit it, the whole "this could lead to data exfiltration" bit. Every signal your brain uses to go "ok this person knows what they're talking about" was firing. The only thing missing was the code it was describing.

Item 9 flagged a missing auth check on an endpoint that had the check. It was one file over from where it was looking.

Here's what actually rattled me. Items 7 and 9 looked identical to items 1 through 3. Same confidence. Same layout. Same tidy little "here's how to fix it" block. Nothing in the way it wrote the fake ones told me they were fake.

More context hadn't made it more honest. It just gave it more of my real function names to wrap a made-up story around. And a hallucination wearing your own variable names is a lot more convincing than a generic one, trust me.

And then the one that actually scared me

Buried in the webhook handler was the real landmine. Not the race this time. Something quieter:

// payment webhook
async function handle(event) {
  ack(event);            // tell the provider "got it" → 200
  await saveToDb(event); // ...then try to persist
}
Enter fullscreen mode Exit fullscreen mode

The ack goes out before the write is safe. If the process dies, or the DB throws, in that tiny window after the acknowledgement, the provider thinks it's done and never retries. The payment is just gone. Silently. And every log line is green. (If you've read the 2am outage post, you already know how this movie ends.)

This is the single most dangerous thing in the repo. Data loss with no error to point at.

So I asked it straight up: "anything wrong with the webhook handler?"

It told me the handler looked correct. Said the try/catch was good practice. Suggested I add a comment.

Same tone it had just used to correctly nuke my dead admin endpoint. Same tone it used to invent a SQL injection out of thin air. On the one bug that can actually lose a customer's money, it gave me a thumbs up and a note about code style.

And I get why. The bug isn't really in the handler. It's in the relationship between two lines, plus a fact that lives outside my codebase entirely: how the payment provider handles retries. That's a seam bug. The model reads inside the frame you hand it, and the danger was in the frame, not the picture.

The real problem: I couldn't trust my own read anymore

Step back and look at what I was actually holding:

  • 3 findings that were true and genuinely useful
  • 2 that were confidently wrong
  • 1 catastrophic bug it signed off on

And nothing in the output told me which was which. The fake vulnerability read exactly as credible as the real one. The "this is fine" on the webhook read exactly as credible as a "this is fine" on genuinely clean code would have.

That's the scary part. Not "AI writes bugs." Not even "AI hallucinates," we all know that by now. It's this:

A model sitting on your whole codebase gives you output where the confidence has basically nothing to do with whether it's right. And more context cranks up the confidence without doing anything for the bugs that live between your files.

I wanted a reviewer. What I got was the most convincing narrator of my own code I've ever seen. Equally convincing when it was right, when it was wrong, and when it was about to cost me money.

Why more context makes this worse, not better

You'd assume more context = better review. For finding stuff, sure. For judgement, it actually cuts the wrong way:

More surface area to sound expert about. It can now drop your real module names into a completely fabricated claim, and specificity reads as truth. The seam bugs are still invisible, because whole-repo context doesn't include the world your repo runs in, and that's where the expensive bugs hide. And it's reviewing with the same instincts it would've used to write the thing, so right where it would've made a mistake, it can't see the mistake, because catching it would mean disagreeing with itself.

None of this is "don't use it." It's "know which job you're actually asking it to do."

Finder vs. witness

There are two different jobs hiding inside the word "review," and I'd been treating them as one.

A finder throws candidate problems at you. Recall is what matters. Being wrong sometimes is completely fine, because a human triages the pile. A whole-repo model is a great finder. Use it as one, no notes.

A witness is the thing that says "yep, this is correct, ship it." And here, being confidently wrong is a disaster, because the entire point of a witness is that nobody checks its work again.

The model is an A+ finder and an F witness, and it hands you both in the same paragraph without flagging the difference. Every mistake I made that afternoon came from reading its witness statements like they carried a finder's stakes.

So here's what I do now:

  1. Treat it as a finder, never a witness. Every "this looks fine" is worth zero to me. Only the "here's a problem" items get my attention, and each one is a lead, not a verdict.
  2. Check every finding against the actual code. The two hallucinations died in about thirty seconds once I opened the file. That check costs nothing. Skipping it costs you a day "fixing" a bug that was never there, feeling productive the whole time.
  3. Whatever certifies the change has to be independent. Different prior than the thing that wrote it. A different model family, an adversarial prompt ("give me the input that loses money" beats "is this correct?"), a test that actually runs, and a human on the merge button. A second pass from the same mind just re-derives the same blind spot.
  4. Seam bugs need seam tests. No amount of reading catches ack-before-persist. Kill the process between those two lines and watch what the provider does. Behaviour, not opinion.

Right, the security part

Pasting a proprietary codebase into a consumer chat product is a decision, not a reflex. Before you do what I did:

  • Assume anything you put in a consumer tier might be retained or trained on unless a contract says otherwise. Read the terms for the exact tier you're on, not the marketing page.
  • Secrets are the immediate risk. The model found a secret in my git history, which means the secret was in what I pasted. Scrub credentials, tokens, customer data, internal hostnames before anything leaves your machine.
  • If this is real work, use an enterprise tier with a no-training guarantee, or run a model where your code already lives. The convenience of the chat box is not worth your source tree showing up in a training set later.

I did the whole thing on a throwaway clone with the secrets already rotated. Do that.

So, the actual takeaway

Give a model your whole codebase. Seriously, do it. As a finder it'll show you stuff you shipped and forgot, and it'll map your system faster than your own docs.

Just don't call that a review. The thing I really walked away with is that its "this is correct" isn't evidence of anything, and it shows up in the exact same voice as the findings that are pure gold. The second you let the thing that reads the code also be the thing that certifies the code, you've built yourself a witness that agrees with itself every single time.

A model saying "looks correct" was never proof it's correct. You still need a second seat whose only job is to not believe the first one. And a human on the merge.

I build xenition on exactly that split: one model writes, a separate skeptic with a different prior tries to tear it apart and isn't allowed to say "looks fine," and a human owns the merge. This whole mess is why.

Top comments (5)

Collapse
 
danielchinasz profile image
daniel •

What you said is a real story. Once AI finishes reading a codebase, it can genuinely flag issues you never noticed — some pretty serious ones, too. But that doesn’t make it rigorous or trustworthy itself; it’ll happily write serious bugs of its own. Tests have to be even stricter than before.

Collapse
 
infoinlet1 profile image
Info Inlet •

Yeah, exactly. And honestly that's the part that still gets me: the same tool that catches a bug you'd never have spotted will hand you a confident, perfectly-formatted bug two lines later, and on the page they look identical.

Fully agree on stricter tests. One thing I'd add though: if the AI writes the tests too, they inherit its blind spot. It'll happily test the behaviour it already believes is correct. So "stricter" has to also mean independent — tests from a different prior, plus the behavioural/seam ones for bugs that don't live inside a single file (the webhook ack-before-save never shows up in a normal unit test). Otherwise you just end up with greener tests that all quietly agree with each other.

Collapse
 
xiangc_92294 profile image
xiangc •

"The model reads inside the frame you hand it, and the danger was in the frame, not the picture."

That line completely nails it.

LLMs analyze static text, but production outages happen in distributed state transitions. You can give a model your entire codebase, but you can't give it the network partitions, third-party webhook retry backoffs, or the database pool running out of connections at 3 AM.

Static code review (human or AI) verifies syntax and obvious logic. Fault injection and integration testing verify reality. If you want to catch seam bugs, stop asking models to read the code—start simulating process crashes between async lines.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the sharper version of what I was clumsily getting at, so thank you.

"Stop asking models to read the code, start simulating crashes between async lines" is exactly it. The one place I'd still keep the model in the loop is finding where to inject. Point it at the repo and ask "list every spot where we ack, respond, or commit before the durable write, or hold a DB connection across an await," and it's genuinely good at enumerating the candidate seams. Then you fault-inject each one and let reality vote.

Model finds the suspects, chaos testing convicts them. It just never gets to be the judge.

Collapse
 
morgan_todd99 profile image
MorganTodd •

I especially like the point that giving AI a full codebase can reveal problems you’ve simply stopped noticing after working with the same project for a long time.