DEV Community

Cover image for I Found 60+ Live Supabase Keys in Public Repos, So I Benchmark Whether LLMs Can Spot Them
Cenk KURTOĞLU
Cenk KURTOĞLU

Posted on

I Found 60+ Live Supabase Keys in Public Repos, So I Benchmark Whether LLMs Can Spot Them

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

As part of my day-to-day security work, I scan public GitHub repos for committed Supabase credentials. In three days of scanning I found 60+ live service_role keys — keys that bypass Row Level Security entirely and give full read/write access to real production databases. Real finds include:

  • config.php files with hardcoded service_role keys (a CRM, an LMS, a driving school's backend)
  • .env.example templates where someone pasted a real key instead of a placeholder
  • NEXT_PUBLIC_-prefixed service_role keys — which means the key ships in the browser bundle to every visitor
  • A docker-compose.yml committing three secrets at once: DB password, anon key, and service_role key
  • Deploy docs (DEPLOY.md, RAILWAY_ENV_SETUP.md) with full credentials "for convenience"

That got me wondering: the people leaking these keys are often using AI assistants to write this code. Can the models themselves spot what they helped commit? So I built a benchmark on Kaggle Benchmarks that treats models like a security reviewer: given a realistic file snippet, triage it — is there a service_role key (critical), an anon key (medium), a DB password, or nothing (placeholder/docs-only)?

The 10 cases are all modeled on real leak patterns I actually found (every key in the benchmark is a fake, structurally-valid JWT):

  1. Hardcoded service_role + anon in PHP config
  2. .env with a single anon key
  3. Client-side createClient() fallback with anon key
  4. Deploy doc with service_role key + DB password
  5. Docs with placeholders only (the false-positive trap)
  6. NEXT_PUBLIC_ service_role fallback in a client file — bundles into the browser
  7. .env.example left with a real service_role key
  8. Commented-out "old key kept for reference" — still a live secret
  9. Two keys side by side (anon + service_role) — models must classify both correctly
  10. A base64-obfuscated service_role used in a bash script — requires two-step decoding

Each model answers a strict JSON verdict (has_service_role, has_anon, has_db_password, warning_level, reasoning) and the task asserts every field. Score = fraction of cases fully correct. Temperature 0. One submission per participant, so I made the cases count.

Models Tested

Nine models from the Kaggle Benchmarks suite, picked to cover the spectrum developers actually choose between — frontier flagships, cheap/fast workhorses, and open-source reasoning models:

  • anthropic/claude-sonnet-5
  • google/gemini-3.7-flash
  • google/gemini-3-flash-preview
  • google/gemini-3.1-flash-lite-preview
  • openai/gpt-5.4-nano
  • openai/gpt-oss-120b
  • deepseek-ai/deepseek-r1-0528
  • qwen/qwen3-next-80b-a3b-instruct
  • ibm/granite-4.0-h-small

If your AI assistant is one of these, this is literally a test of "would my assistant catch its own leak?"

Findings

Model Score Notes
claude-sonnet-5 10/10 Clean sweep
gemini-3.7-flash 10/10 Clean sweep
gemini-3-flash-preview 10/10 Decoded the base64 blob mid-response to find the role claim
gemini-3.1-flash-lite 10/10 Clean sweep, cheapest tier to do it
gpt-5.4-nano 8/10 Missed the commented-out key + the base64 obfuscation
deepseek-r1-0528 0/10 Flagged everything critical — see below
qwen3-next-80b partial Passed 5/6 assertions before rate-limiting out
gpt-oss-120b — Errored repeatedly (provider 429 under load; excluded from comparison)

The headline: modern models are scarily good at the obvious cases

On the five "classic" cases (hardcoded config keys, docs with credentials, placeholder-only false-positive traps), every frontier model scored perfectly. Even gpt-5.4-nano — the cheap, fast one — went 5-for-5. If you paste a file with SUPABASE_SERVICE_ROLE_KEY=eyJ... into any of these models and ask "is this bad?", they will all tell you correctly: yes, critical, rotate it.

The interesting part is where they diverge

DeepSeek-R1, a dedicated reasoning model, scored 0/10 — but not because it missed the keys. It flagged everything: every file, including the placeholder-only DEPLOY.md, came back "critical, service_role present, DB password present." In security triage, a reviewer who cries critical on every file gets ignored on the one that matters. Recall without precision is just alarm fatigue with extra steps.

The edge cases did their job. gpt-5.4-nano missed exactly the two sneaky ones: the commented-out "old key kept for reference" (it treated the comment as dead code — but commented secrets still work, and old keys often still rotate back into service) and the base64-obfuscated key. The frontier models — including the cheap Gemini Flash tiers — caught all ten, with Gemini 3 Flash visibly decoding the base64 blob mid-response to identify the role claim before answering.

What this means in practice

  1. Your AI assistant will not save you from committed secrets — but it can. Every model tested can catch these leaks instantly when asked. The problem is nobody asks. Secrets get committed because the review step never happens.
  2. Over-flagging is its own failure mode. A reviewer (human or AI) that cries "critical" on every placeholder teaches teams to ignore alarms. Precision matters as much as recall in security triage.
  3. The 15-minute fix beats any detector. Rotate the key in the Supabase dashboard (dies instantly), move it to env vars, git filter-repo if you want it out of history.

What I'd measure next

  • Tool use: give the models a repo tree and a search tool, and see if they actively hunt for secrets rather than judging a pasted file
  • Multi-file context: the real-world version of this leak is spread across config.php + deploy docs + docker-compose — does performance drop when the secret is two hops away?
  • Fix quality: flagging is step one; does the model produce a correct remediation (rotate first, then remove, then history-clean)?

Run the same checks on your own app

The 10 cases above are one snapshot. Your codebase will have its own version.

I put nine read-only Postgres queries that cover the same failure classes into a free file - nothing leaves your database, you run it in your own SQL editor:
github.com/cekuu35/supabase-rls-leak-demo - audit/rls-audit.sql

The same repo has a red/green test suite that demonstrates a cross-tenant leak and the one-file fix.

If you want the full guided version - write-side checks, a role-simulation harness, and remediation templates you can drop into an existing repo - that's the Supabase RLS Audit Kit. It runs entirely against your own catalogs, $29.

Disclosure: both links are mine. The audit queries and demo are free and MIT licensed.

My Benchmark

Task: committed-supabase-key-detection on Kaggle

The task page shows per-model results, full conversations, and the assertion breakdown for every case. Fake keys, real patterns — steal the cases for your own review checklists.

(And if you're reading this with a committed Supabase key somewhere in your repo history: rotate it now. It takes 15 minutes, and I promise you're not the only one who's seen it.)

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

The DeepSeek-R1 result illustrates why treating general reasoning models as pre-commit linters backfires in developer workflows. When a model treats every potential credential placeholder as an active emergency, developers learn to bypass the hook with --no-verify, which defeats the purpose entirely.

The nano model skipping the commented-out key also points to a classic semantic blind spot: smaller models tend to confuse code execution with data exposure. If it sits inside a comment block, the parser treats it as inert text rather than an active secret sitting in git history. Pairing a dumb regex check for high-entropy tokens with a fast classifier for context filtering usually avoids both failure modes.