I spent three weeks reviewing pull requests where most or all of the code was written by an AI assistant — Copilot, Cursor, ChatGPT, Claude, you name it. Fifty PRs across twelve repos, ranging from weekend side projects to early-stage production apps.
The code worked. Almost every PR passed its test suite. But the security posture was consistently poor in the same six ways.
1. Hard-coded secrets
The most common issue by far. API keys, database passwords, and JWT signing secrets appeared as string literals in 23 of the 50 PRs. The AI was asked to "connect to Stripe" or "set up the database," and it delivered working code with the real credential on the same line.
# Appeared in 23/50 PRs
db = psycopg2.connect("postgresql://admin:s3cret@prod-db:5432/app")
The fix is always the same — environment variables — but the AI will not reach for them unless your project instructions say to.
2. SQL injection through string formatting
Eleven PRs built SQL queries by interpolating user input. The AI used f-strings or template literals instead of parameterised queries, often inside an ORM's raw-query escape hatch.
# Appeared in 11/50 PRs
cursor.execute(f"SELECT * FROM users WHERE email = '{request.form['email']}'")
A single quote in the email field breaks this open. Parameterised queries cost the same number of keystrokes, but the AI defaults to the pattern it has seen most often in training data — and that pattern is the unsafe one.
3. Missing or broken authentication checks
Nine PRs added new API endpoints without verifying the caller's session or token. The AI focused on the business logic — "build a dashboard endpoint that returns analytics" — and delivered exactly that, minus the middleware that restricts it to logged-in admins.
This one is hard to spot in a diff. The endpoint works, the tests pass (because the tests also skip auth), and the code looks clean. You only notice it when you check whether the route is wrapped in the same auth middleware as its neighbours.
4. Overly permissive CORS
Seven PRs set Access-Control-Allow-Origin: * on APIs that serve private data. The AI treats CORS as a "make it work" checkbox and picks the widest setting. In a staging environment that is fine. In production it means any website can call your API with the user's cookies.
// Appeared in 7/50 PRs
app.use(cors({ origin: '*' }));
5. Logging sensitive data
Six PRs logged full request bodies — including passwords, tokens, and payment details — to stdout or a log file. The AI added console.log(req.body) or logger.info(payload) during debugging and never removed it.
Logs end up in monitoring dashboards, third-party log aggregators, and error-tracking tools. A plaintext password in your Datadog is a compliance incident waiting to happen.
6. Outdated or vulnerable dependencies
Five PRs imported packages with known CVEs. The AI pinned the version it was trained on — sometimes two major versions behind — and the project had no npm audit or pip-audit step to flag it.
The pattern underneath
None of these are exotic. Every item on this list is in any web-security checklist. The problem is that AI assistants optimise for "does it work?" and security bugs do not break functionality. A SQL injection endpoint returns the correct rows for correct input. An unauthenticated route serves the right data to the right user — and also to everyone else.
Human reviewers catch these, but only when they are looking for them. In a 600-line AI-generated diff, your eyes glaze over at the third helper function and you start skimming.
What actually catches them
The fix is not "stop using AI to write code." The fix is automated review that checks every diff for the patterns humans skip over.
An AI code review tool built for security can flag hard-coded secrets, SQL injection, missing auth checks, and the rest of this list as inline PR comments — before a human reviewer even opens the diff. It runs in seconds and does not get tired at PR number forty.
This is what we built Diffnix to do. Its PRInspector agent reviews every pull request in under five seconds, catches the six bug classes above and more, and posts the findings as inline comments on the exact lines that need fixing. It runs on self-hosted models and stores none of your code. There is a free plan for up to three repositories.
A quick checklist if you ship AI-generated code
- Run a secret scanner on every PR (not just commits — the diff itself).
- Require parameterised queries in your linter or project rules.
- Test every new endpoint for unauthenticated access.
- Set CORS to your actual domain list, not
*. - Strip debug logging before merge.
- Add
npm audit/pip-auditto CI. - Turn on automated PR review so these checks happen without relying on a human to remember each one.
AI-generated code is not inherently less secure, but it is consistently less careful. The same six bugs repeat because the model optimises for function, not safety. An automated reviewer that checks every diff is the cheapest way to close the gap.
Want to go deeper? Read What Is AI Code Review? A Practical Guide for Engineering Teams.
Moonlight Devs builds Diffnix, an AI code reviewer for GitHub pull requests.
Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.