TL;DR — GSC (Git Security Checker) is a self-learning AppSec platform that doesn't stop at detecting a vulnerability. It proves it with a sandboxed exploit, fixes it with a minimal patch, verifies the fix by re-running the exploit, and heals the repo with an automated PR — then learns from your TP/FP verdicts to cut false positives over time. This post is a technical tour of the whole system: the detection engine, the proof/fix/verify loop, the adversarial AI-judging layer, and the closed-source/core split that ships it as a GitHub Action and an MCP server.
The problem with most SAST
A typical static analyzer hands you a snapshot: a flat list of findings that you then triage, verify, and fix by hand. Three pains follow immediately:
- No proof. "Possible SQL injection" is a hypothesis, not evidence. You still have to write the exploit yourself to know if it's real.
- False-positive fatigue. When every run surfaces hundreds of medium/low findings, developers learn to ignore the tool entirely.
- No loop closure. The scanner never learns which of its findings were right, so its precision never improves for your codebase.
GSC is built around closing exactly this loop. The pipeline is detect → prove → fix → verify → heal → learn, and every stage has a concrete mechanism behind it.
Try it in 10 minutes
Three entry points, depending on how you want to run it.
Local CLI:
git clone https://github.com/poliakarmai/gsc.git && cd gsc
pip install -e .
gsc external-scan https://github.com/user/repo --profile audit
With exploit proof and verified fixes:
gsc external-scan ./repo --profile audit --with-poc --with-chains
No LLM key? It still runs — GSC degrades to regex-only detection (finds vulnerabilities, just without confirmed verdicts). Add DEEPSEEK_API_KEY (or OPENROUTER_API_KEY, or LLM_BASE_URL for any OpenAI-compatible endpoint including a local Ollama) to switch on the adversarial judge.
GitHub Action — drop into any repo's CI:
- uses: poliakarmai/gsc@v1
with:
deep_scan: true
with_poc: true
fail_on_critical: true
Scans every PR/push, posts findings as a PR comment, shows a score badge, blocks merge on CRITICAL.
MCP server — let your AI agent (Claude Code, Cursor) drive it:
{
"mcpServers": {
"gsc": {
"command": "python3",
"args": ["gsc_mcp_server.py"],
"cwd": "/path/to/gsc"
}
}
}
The agent then calls scan_repo, list_findings, verify_finding inside its own context — scan, verify, fix, re-scan, without you parsing CLI output.
The closed loop, stage by stage
1. Detect — 65 detectors across four engines
Detection is a real engine, not an LLM wrapper. 65 detectors break down as 61 registry rules + 4 standalone engines (Secrets, SCA, IaC, Invariants). Coverage spans Python, JS/TS, Go, Java, Rust and more, with an AST/taint layer under the regex surface.
On top of raw detection sit the "surface" capabilities:
| Layer | What it covers |
|---|---|
| SAST | 65 detectors, adversarial AI-judged |
| SCA | dependency CVEs via OSV.dev with lock-file resolution (package-lock.json, yarn.lock, go.sum) |
| Secrets | fingerprinting with cross-repo correlation and rotation detection — stores hashes only, never values |
| Honeytokens |
GSC_HONEYTOKEN_* canaries that emit a CRITICAL finding if they leak into any scanned repo |
| IaC | Terraform, Kubernetes, Dockerfile misconfigurations |
| Supply chain | SBOM (CycloneDX/SPDX), VEX, signature verification |
| Compliance | CWE / OWASP / PCI DSS mapping on every finding |
2. Prove — a working exploit, not a guess
The finding you see is not "might be exploitable" — GSC generates a working PoC for it and runs it in an isolated sandbox. A finding becomes evidence.
There's a safe-exploitation gate: PoC generation is blocked from emitting destructive commands (rm -rf, DROP TABLE, fork-bombs). To be precise — this is a denylist filter, not a sandbox guarantee: PoC execution is sandboxed at the process level (timeouts, resource limits), not full VM isolation. It reduces blast radius; it is not a substitute for running in an ephemeral container if you point GSC at code you don't trust.
3. Fix + Verify — Proof-of-Fix
Proof-of-Fix is the mechanism that separates "a patch was written" from "a patch that actually works":
PoF_verified = (exploit succeeds BEFORE) ∧ (exploit fails AFTER) ∧ (patch is minimal)
The flow is deterministic: generate the exploit → produce a minimal patch → re-run the same exploit against the patched code in the sandbox. Only a fix that defeats its own exploit counts as verified. Two honest caveats: "minimal patch" is a diff-size heuristic, not a formal machine-checkable criterion, and a sandbox-verified fix is not the same as "safe to merge" — it proves the exploit is defeated, not that the patch is regression-free. That last part is still human review.
4. Heal — Self-Healing CI
A verified fix doesn't sit in a report — GSC can open a pull request with the patch automatically. The detect → prove → fix → verify → heal chain means the tool can carry a finding all the way to a mergeable PR without a human touching it.
5. Learn — your TP/FP verdicts tune the engine
This is the part that compounds. The Ground-Truth Trainer revalidates findings deterministically (zero LLM) against a calibration network of 16 projects (9 clean + 7 vulnerable), and auto-deactivates noisy patterns that fall below a 30% true-positive threshold. On top of that, federated learning shares DP-noised TP/FP signals across tenants to tune detectors globally. The mechanism is Laplace noise (sensitivity 1, ε=1.0 default) with per-tenant privacy-budget accounting (warn at ε=5.0, hard stop at ε=10.0). Only {tenant_hash, rule_id, tp, fp} ever leaves your environment — and tenant_hash is pseudonymous, not anonymous.
A finding, end to end
Concrete, not schematic. A user_id interpolated into an f-string SQL query:
# vulnerable.py
query = f"SELECT * FROM users WHERE id = {user_id}"
cursor.execute(query)
-
Detect — the SQLi rule flags it:
SQL injection risk: f-string in query, severity HIGH. -
Prove — GSC generates a PoC (
user_id = "1 OR 1=1 --") and runs it in the sandbox → the injection is exploitable. -
Judge — the adversarial debate returns
REAL, confidence 95 (the Red advocate produces the exploit path; the Blue advocate finds no mitigation). - Fix — a minimal patch parameterizes the query:
query = "SELECT * FROM users WHERE id = %s"
cursor.execute(query, (user_id,))
-
Verify — the same PoC is re-run against the patched code and now fails →
verified. - Heal — the patch opens as a PR you can review and merge.
The point isn't the SQLi (trivial) — it's that steps 2–5 happen automatically, and step 5 is the machine-checkable part: the fix defeats its own exploit.
Adversarial AI judging — killing hallucinated CVEs
Revalidating a finding with a single LLM call is how you get confident-looking false positives and made-up CVE/CWE identifiers. GSC uses an adversarial consensus instead:
Red Advocate — argues this is a true-positive (exploit path, impact)
Blue Advocate — argues this is a false-positive (mitigation, no taint)
Judge — reads the transcript → {REAL | FALSE_POSITIVE | VERIFY_MANUALLY}
Two details make this trustworthy rather than theatrical:
- Grounded RAG. The judge receives real CVE/CWE references from a local FTS5 knowledge base instead of whatever it might hallucinate.
-
Cost-gating. A deterministic pre-filter assigns
fp_probability(LOW/MEDIUM/HIGH) — confident cases skip the debate entirely, so expensive LLM calls fire only on genuinely ambiguous findings.
There's a BYO-LLM invariant underneath all of this: GSC never depends on its own model key. You bring your own (DeepSeek, OpenRouter, OpenAI-compatible, Ollama), or you bring no key at all — in which case GSC degrades honestly to regex-only: still finds vulnerabilities, just without confirmed verdicts. No key, no crash.
Prioritization that reflects real-world risk
Raw CVSS ranks a finding by its worst theoretical severity. GSC ranks by exploitability:
- EPSS (probability of exploitation in the wild)
- CISA KEV (known exploited vulnerabilities)
- ExploitDB (public exploit availability)
- FSTEC BDU cross-reference (Russian national vulnerability database)
The blocking engine then gates on severity ≥ HIGH ∧ confidence ≥ 80% — so a merge is blocked by findings that are both severe and provable, not by a flood of noise.
The intelligence layer
Beyond the detect→heal loop, a few capabilities round GSC out as a vulnerability-management platform rather than a scanner:
-
Security Archaeology — who introduced a vulnerability, when, and how long it lived before it was found (
git blameon findings). - Predictive Forecasting — risk heatmaps of where the next vulnerability is likely to appear.
- Exploit chains — compose individual findings into real attack paths.
- MAESTRO threat model — a 7-level threat model for agentic systems (CSA-aligned), plus auto-generated architectural threat models.
- Threat intel — export findings as STIX 2.1 / TAXII 2.1 (MISP, OpenCTI) and reconcile code against IoC feeds.
- Red-team + agent-eval — 15 adversarial prompt-injection payloads stress-test the injection detectors; agent metrics (tool-selection recall/precision, retrieval accuracy@k/MRR/NDCG, catastrophic-forgetting, coherence) keep the AI-judging layer honest.
Privacy & threat model
What leaves your machine, and what doesn't:
- Secrets are never stored as values — only a fingerprint (hash), and finding values are redacted in reports.
-
LLM revalidation sends finding snippets (rule, file path, the relevant code fragment) to your BYO model. It's your key, so traffic goes to your provider. Run fully local via
LLM_BASE_URL→ Ollama; you lose judge sharpness, not detection. -
Federated learning uploads only
{tenant_hash, rule_id, tp, fp}with Laplace noise (ε=1.0) and a per-tenant privacy budget (warn 5.0, stop 10.0).tenant_hashis pseudonymous, not anonymous. -
Closed-source Docker image — the engine ships as
ghcr.io/poliakarmai/gsc-scanner. We don't yet publish a signed SBOM / reproducible build for the scanner itself (admittedly ironic for a tool that checks your supply chain — it's on the roadmap).
Under the hood: one scan pipeline
repo (local / git / PR)
├─ inventory (multi-lang discovery, dir exclusions)
├─ detect (61 registry + 4 engines) ──► raw findings
├─ filter (disabled rules, noise tier, framework-aware)
├─ dedupe (finding_key = sha256(rule+file+snippet)[:12])
├─ LLM revalidate (BYO-LLM) ──► adversarial debate + RAG
│ └─ verdict: REAL | FALSE_POSITIVE | VERIFY_MANUALLY + confidence
├─ V3 scoring (llm_confidence + tp/fp signals + file context)
│ └─ confirmed ≥.80 | likely ≥.55 | uncertain ≥.35 | else FP
├─ attack-chain composer
├─ blocking engine (severity ≥ HIGH ∧ confidence ≥ 80%)
└─ report (json / markdown / sarif / html) + PR comment + badge
Three scan modes map to three budgets: quick (CI, regex-only), standard (daily, LLM), deep (full audit with exploit chains). Four profiles: developer-review, pr-gate, audit, candidate-review.
The architecture: public facade, closed core
This is the part most people ask about. GSC is deliberately split:
| Part | What it is | License |
|---|---|---|
Public repo poliakarmai/gsc
|
action.yml (the GitHub Action recipe), documentation, MCP protocol |
Apache 2.0 |
| Core engine |
gsc_core (detectors) + gsc_cli (scanners) + gsc_cloud (SaaS API) |
BUSL-1.1 |
| Distributed image |
ghcr.io/poliakarmai/gsc-scanner — closed-source Docker image |
BUSL-1.1 |
The engine ships as a closed-source Docker image; the public repo carries only the action.yml that runs it. Licensing is BUSL-1.1 (source-available, not open source): scanning your own repositories is always allowed; operating a competing hosted SAST service requires a commercial license; and it converts to Apache 2.0 on the Change Date. Fair, and explicit about it.
Three distribution channels hang off the same core:
-
GitHub Action —
uses: poliakarmai/gsc@v1, scans every PR/push, posts findings, shows a score badge, blocks merge on CRITICAL. -
MCP server —
scan_repo,scan_diff,list_findings,verify_finding,get_finding,list_detectors, so Claude Code / Cursor / any MCP-capable agent scans and verifies code inside its own context. Read-only with respect to your filesystem (it never edits your repo), butverify_findingdoes execute a PoC in the sandbox — so "read-only" means "no write access to your code", not "no code execution". - SaaS API — multi-tenant FastAPI with SSO, workers, webhook (code ready, deployment in progress).
One more use worth spelling out: because the MCP surface is small and read-only, GSC also plugs into other tools as a SAST add-on — not just as a standalone scanner. A DAST-only setup or an autonomous pentest agent (e.g. ARTEX) does recon and dynamic testing but has a static-analysis blind spot; point it at GSC's MCP and it gains scan_repo → verify_finding → SAST findings to feed its exploit phase. GSC closes the gap most offensive/DAST tooling leaves open.
Current state (live numbers, pulled from the codebase at write time)
| Metric | Value |
|---|---|
| Version | 1.4.1 |
| Detectors | 65 (61 registry + 4 engines) |
| Modules | 258 |
| DB schema | 34 |
| Tests | 3227 |
| Calibration network | 16 projects (9 clean + 7 vulnerable) |
| Languages | Python, JS/TS, Go, Java, Rust, and more |
We don't yet publish our own precision/recall on public repositories — the calibration network is small (16 projects) and synthetic-benchmark transfer (OWASP Benchmark v1.2 is Java-only) doesn't reflect multilingual real-world code. Related work on the multi-agent judging approach exists (QASecClaw reports F1 78.4% → 90.9% on OWASP Benchmark v1.2), but we do not claim those numbers for GSC.
Limitations (stated honestly)
No SAST post should skip this section, so:
- BUSL, not OSS. If your definition of "open source" requires OSI approval, GSC is source-available, not open source. It converts to Apache 2.0 on the Change Date.
- Precision depends on your BYO-LLM. Without a key, GSC degrades to regex-only — honest, but coarser. With a weak model, the adversarial judge is only as sharp as the judge.
- Proof ≠ every language. PoC auto-generation is strongest where the exploit sandbox has solid support; some languages fall back to detect + manual verify.
- No published precision/recall yet. The 16-project calibration network tunes thresholds, but we haven't published reproducible precision/recall/F1 on public repositories — that's the strongest open question about the tool, and the first thing we owe a skeptical reader.
-
Performance & cost.
deepmode with the adversarial judge is LLM-bound — slow on large monorepos, and BYO-LLM means your tokens and rate limits.quickmode (regex-only) is fast but noisier. - SaaS is code-ready, not deployed. The multi-tenant API exists but needs a server; today the practical paths are the Action and MCP.
What's next
On the near-term roadmap: deeper JS/XSS detection, expanding the MCP toolset (patch preview, SBOM), an IAST runtime validator for PoC/PoF, and the SaaS deployment (multi-tenant + OIDC). The self-learning loop — grounding on the calibration network, federated TP/FP tuning — is the axis the whole project keeps sharpening.
Who it's for
Good fit: GitHub-first teams who want proof + auto-fix, not just a list of findings; teams willing to spend their own LLM tokens on verification; bug-bounty and security-archaeology work; multi-agent setups (MCP).
Poor fit: air-gapped networks with no LLM (you get regex-only detection); non-GitHub CI without Docker; tiny projects where setup outweighs the value; teams that require OSI-licensed tooling; languages without PoC support (you get detect + manual verify).
Compared to Semgrep / CodeQL / Snyk: those are mature, open, and excellent at breadth and ecosystem coverage. GSC's differentiator is the closed loop — prove, fix, verify, learn — not raw detection breadth. The honest answer is usually: run Semgrep/CodeQL for coverage and GSC for the findings you actually need to close.
Glossary
- SAST — static analysis (reads source, doesn't run it). SCA — software-composition analysis (dependencies). IaC — infra-as-code scanning.
- PoC — proof-of-concept exploit. PoF — proof-of-fix (re-run the exploit against the patch).
- TP / FP — true / false positive. EPSS — probability a CVE is exploited in the wild. KEV — CISA's Known Exploited Vulnerabilities.
- SARIF — standard finding format (GitHub code scanning). MCP — Model Context Protocol (how agents call tools).
Conclusion
The difference between a scanner and a vulnerability-management platform is the loop. GSC's bet is that a tool should prove what it reports, fix what it proves, verify what it fixes, and learn from whether it was right. For a security engineer tired of triaging unverified findings by hand, that loop — not any single detector — is the feature worth looking at.
If you'd rather poke at it than read about it: the Action is poliakarmai/gsc, the MCP server is documented in the repo, and the scan is gsc external-scan <repo> --profile audit --with-poc --with-chains.
Top comments (0)