🛡️ I Built an AI Security Scanner That Eliminates False Positives
Most security scanners have one big problem: too much noise.
You run a scan, get 200 alerts, and 160 of them are false positives. You spend hours triaging instead of actually fixing real bugs. Sound familiar?
That's exactly why I built Sentinel Bug Engine — an open-source, AI-powered security scanner with a unique Joint Verification Engine (JVE) that adversarially scores each finding before it ever reaches you.
👉 GitHub: github.com/pavan67-git/sentinel-bug-engine
🤔 The Problem with Existing Tools
| Tool | What it does well | The gap |
|---|---|---|
| Semgrep | Fast static patterns | High false positive rate |
| Bandit | Python-specific | No verification step |
| Snyk | Dependency scanning | Commercial, not code-level |
| CodeQL | Deep analysis | Complex setup, GitHub-only |
None of them answer the real question: "Is this alert actually exploitable in this specific codebase?"
💡 The Solution: 4-Tier Analysis Pipeline
Sentinel uses a layered approach:
┌────────────────────────────────────────────────────┐
│ Tier 1: Universal Parser │
│ → Detects language, extracts code blocks │
├────────────────────────────────────────────────────┤
│ Tier 2: Deterministic Fast Filter │
│ → Regex + pattern rules (OWASP Top 10, CWEs) │
│ → Sub-second scanning │
├────────────────────────────────────────────────────┤
│ Tier 3: LLM Cognitive Analysis (optional) │
│ → Gemini / OpenAI / Claude / Ollama │
│ → Catches business logic bugs, race conditions │
├────────────────────────────────────────────────────┤
│ Tier 4: Joint Verification Engine (JVE) ← KEY │
│ → Adversarially evaluates each finding │
│ → Detects mitigating controls in context │
│ → Assigns confidence score, discards FPs │
└────────────────────────────────────────────────────┘
🎯 The JVE — What Makes This Different
The Joint Verification Engine is the core differentiator.
For every potential finding, the JVE:
- Reads 20 lines of surrounding code context
- Looks for mitigating controls — sanitizers, parameterized queries, type guards, allowlists
- Assigns a confidence score (0–100%)
- Labels the finding:
VERIFIED_TRUE_POSITIVE,LIKELY_TRUE_POSITIVE,SUSPECTED_FALSE_POSITIVE, orDISPROVED_FALSE_POSITIVE -
Automatically discards
DISPROVED_FALSE_POSITIVEfindings
Here's the result in practice:
# ❌ VULNERABLE — JVE flags this (LIKELY_TRUE_POSITIVE, 60%)
query = f"SELECT * FROM users WHERE id = {user_id}"
cursor.execute(query)
# ✅ SAFE — JVE correctly ignores this
cursor.execute("SELECT * FROM users WHERE id = ?", (int(user_id),))
The second query uses parameterized inputs — JVE detects this and does not raise a false alert.
📊 Live Scan Results
I ran Sentinel against a deliberately vulnerable Flask app. Here's the real output:
=== SENTINEL SCAN RESULTS ===
Target: vulnerable_flask_app.py
Total findings: 6
CRITICAL: 2
HIGH: 3
MEDIUM: 1
[SEC-001] CRITICAL | CWE-798 | Hardcoded API Secret or Private Key | JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-002] HIGH | CWE-89 | SQL Query Concatenation / Potential SQLi| JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-003] HIGH | CWE-78 | Dangerous OS Command Execution | JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-004] HIGH | CWE-78 | Dangerous OS Command Execution | JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-005] CRITICAL | CWE-502 | Unsafe Dynamic Code Evaluation (eval) | JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-006] MEDIUM | CWE-400 | Resource Leak (Unclosed File) | JVE: LIKELY_TRUE_POSITIVE (60%)
6 real findings. 0 false positives.
Every finding comes with:
- CWE classification
- OWASP Top 10 mapping
- Exact file + line location
- JVE verdict + confidence score
- Suggested fix with code diff
🔍 What It Detects
| CWE | Vulnerability | OWASP |
|---|---|---|
| CWE-89 | SQL Injection | A03:2021 |
| CWE-78 | OS Command Injection | A03:2021 |
| CWE-79 | Cross-Site Scripting | A03:2021 |
| CWE-22 | Path Traversal | A01:2021 |
| CWE-502 | Unsafe Deserialization / eval | A08:2021 |
| CWE-798 | Hardcoded Credentials | A07:2021 |
| CWE-918 | Server-Side Request Forgery | A10:2021 |
| CWE-400 | Resource Exhaustion / Leaks | — |
| CWE-362 | Race Conditions | — |
| CWE-476 | Null Pointer Dereference | — |
Supported languages: Python, JavaScript, TypeScript, Java, Go, PHP, C/C++, Rust, Ruby, C#
⚡ Quick Start
# Install
git clone https://github.com/pavan67-git/sentinel-bug-engine.git
cd sentinel-bug-engine
pip install .
# Basic scan
python -m src.cli scan ./your-project
# With SARIF export (for GitHub Advanced Security / GitLab CI)
python -m src.cli scan ./your-project --format both --sarif-out results.sarif
# Only HIGH+ severity
python -m src.cli scan ./your-project --min-severity HIGH
# With LLM reasoning (Gemini)
export GEMINI_API_KEY="your-key"
python -m src.cli scan ./your-project --enable-llm
📦 CI/CD Integration
Sentinel exports SARIF 2.1.0 — the industry standard supported by:
- ✅ GitHub Advanced Security (Code Scanning)
- ✅ GitLab Security Dashboard
- ✅ SonarQube
- ✅ Azure DevOps
Add it to your GitHub Actions workflow:
- name: Run Sentinel Security Scan
run: |
pip install .
python -m src.cli scan . --min-severity HIGH --format sarif --sarif-out results.sarif
- name: Upload SARIF to GitHub
uses: github/codeql-action/upload-sarif@v3
with:
sarif_file: results.sarif
🚀 What's Next
- [ ] More LLM-powered rules (business logic, race conditions)
- [ ] VS Code extension
- [ ] Web dashboard (already in progress)
- [ ] Benchmark against CVE database
🤝 Contribute
The project is open source and I'd love contributions — especially:
- New detection rules for more CWEs
- Language-specific pattern improvements
- Benchmark datasets
👉 github.com/pavan67-git/sentinel-bug-engine
If this helped you or the JVE concept interests you, drop a ⭐ on GitHub or leave a comment below. Always open to feedback!
Top comments (0)