This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
As part of my day-to-day security work, I scan public GitHub repos for committed Supabase credentials. In three days of scanning I found 60+ live service_role keys — keys that bypass Row Level Security entirely and give full read/write access to real production databases. Real finds include:
-
config.phpfiles with hardcoded service_role keys (a CRM, an LMS, a driving school's backend) -
.env.exampletemplates where someone pasted a real key instead of a placeholder -
NEXT_PUBLIC_-prefixed service_role keys — which means the key ships in the browser bundle to every visitor - A
docker-compose.ymlcommitting three secrets at once: DB password, anon key, and service_role key - Deploy docs (
DEPLOY.md,RAILWAY_ENV_SETUP.md) with full credentials "for convenience"
That got me wondering: the people leaking these keys are often using AI assistants to write this code. Can the models themselves spot what they helped commit? So I built a benchmark on Kaggle Benchmarks that treats models like a security reviewer: given a realistic file snippet, triage it — is there a service_role key (critical), an anon key (medium), a DB password, or nothing (placeholder/docs-only)?
The 10 cases are all modeled on real leak patterns I actually found (every key in the benchmark is a fake, structurally-valid JWT):
- Hardcoded service_role + anon in PHP config
-
.envwith a single anon key - Client-side
createClient()fallback with anon key - Deploy doc with service_role key + DB password
- Docs with placeholders only (the false-positive trap)
-
NEXT_PUBLIC_service_role fallback in a client file — bundles into the browser -
.env.exampleleft with a real service_role key - Commented-out "old key kept for reference" — still a live secret
- Two keys side by side (anon + service_role) — models must classify both correctly
- A base64-obfuscated service_role used in a bash script — requires two-step decoding
Each model answers a strict JSON verdict (has_service_role, has_anon, has_db_password, warning_level, reasoning) and the task asserts every field. Score = fraction of cases fully correct. Temperature 0. One submission per participant, so I made the cases count.
Models Tested
Nine models from the Kaggle Benchmarks suite, picked to cover the spectrum developers actually choose between — frontier flagships, cheap/fast workhorses, and open-source reasoning models:
- anthropic/claude-sonnet-5
- google/gemini-3.7-flash
- google/gemini-3-flash-preview
- google/gemini-3.1-flash-lite-preview
- openai/gpt-5.4-nano
- openai/gpt-oss-120b
- deepseek-ai/deepseek-r1-0528
- qwen/qwen3-next-80b-a3b-instruct
- ibm/granite-4.0-h-small
If your AI assistant is one of these, this is literally a test of "would my assistant catch its own leak?"
Findings
| Model | Score | Notes |
|---|---|---|
| claude-sonnet-5 | 10/10 | Clean sweep |
| gemini-3.7-flash | 10/10 | Clean sweep |
| gemini-3-flash-preview | 10/10 | Decoded the base64 blob mid-response to find the role claim |
| gemini-3.1-flash-lite | 10/10 | Clean sweep, cheapest tier to do it |
| gpt-5.4-nano | 8/10 | Missed the commented-out key + the base64 obfuscation |
| deepseek-r1-0528 | 0/10 | Flagged everything critical — see below |
| qwen3-next-80b | partial | Passed 5/6 assertions before rate-limiting out |
| gpt-oss-120b | — | Errored repeatedly (provider 429 under load; excluded from comparison) |
The headline: modern models are scarily good at the obvious cases
On the five "classic" cases (hardcoded config keys, docs with credentials, placeholder-only false-positive traps), every frontier model scored perfectly. Even gpt-5.4-nano — the cheap, fast one — went 5-for-5. If you paste a file with SUPABASE_SERVICE_ROLE_KEY=eyJ... into any of these models and ask "is this bad?", they will all tell you correctly: yes, critical, rotate it.
The interesting part is where they diverge
DeepSeek-R1, a dedicated reasoning model, scored 0/10 — but not because it missed the keys. It flagged everything: every file, including the placeholder-only DEPLOY.md, came back "critical, service_role present, DB password present." In security triage, a reviewer who cries critical on every file gets ignored on the one that matters. Recall without precision is just alarm fatigue with extra steps.
The edge cases did their job. gpt-5.4-nano missed exactly the two sneaky ones: the commented-out "old key kept for reference" (it treated the comment as dead code — but commented secrets still work, and old keys often still rotate back into service) and the base64-obfuscated key. The frontier models — including the cheap Gemini Flash tiers — caught all ten, with Gemini 3 Flash visibly decoding the base64 blob mid-response to identify the role claim before answering.
What this means in practice
- Your AI assistant will not save you from committed secrets — but it can. Every model tested can catch these leaks instantly when asked. The problem is nobody asks. Secrets get committed because the review step never happens.
- Over-flagging is its own failure mode. A reviewer (human or AI) that cries "critical" on every placeholder teaches teams to ignore alarms. Precision matters as much as recall in security triage.
-
The 15-minute fix beats any detector. Rotate the key in the Supabase dashboard (dies instantly), move it to env vars,
git filter-repoif you want it out of history.
What I'd measure next
- Tool use: give the models a repo tree and a search tool, and see if they actively hunt for secrets rather than judging a pasted file
-
Multi-file context: the real-world version of this leak is spread across
config.php+ deploy docs + docker-compose — does performance drop when the secret is two hops away? - Fix quality: flagging is step one; does the model produce a correct remediation (rotate first, then remove, then history-clean)?
Run the same checks on your own app
The 10 cases above are one snapshot. Your codebase will have its own version.
I put nine read-only Postgres queries that cover the same failure classes into a free file - nothing leaves your database, you run it in your own SQL editor:
github.com/cekuu35/supabase-rls-leak-demo - audit/rls-audit.sql
The same repo has a red/green test suite that demonstrates a cross-tenant leak and the one-file fix.
If you want the full guided version - write-side checks, a role-simulation harness, and remediation templates you can drop into an existing repo - that's the Supabase RLS Audit Kit. It runs entirely against your own catalogs, $29.
Disclosure: both links are mine. The audit queries and demo are free and MIT licensed.
My Benchmark
Task: committed-supabase-key-detection on Kaggle
The task page shows per-model results, full conversations, and the assertion breakdown for every case. Fake keys, real patterns — steal the cases for your own review checklists.
(And if you're reading this with a committed Supabase key somewhere in your repo history: rotate it now. It takes 15 minutes, and I promise you're not the only one who's seen it.)
Top comments (1)
The DeepSeek-R1 result illustrates why treating general reasoning models as pre-commit linters backfires in developer workflows. When a model treats every potential credential placeholder as an active emergency, developers learn to bypass the hook with --no-verify, which defeats the purpose entirely.
The nano model skipping the commented-out key also points to a classic semantic blind spot: smaller models tend to confuse code execution with data exposure. If it sits inside a comment block, the parser treats it as inert text rather than an active secret sitting in git history. Pairing a dumb regex check for high-entropy tokens with a fast classifier for context filtering usually avoids both failure modes.