DEV Community

Cover image for I Built an Open-Source AI Security Scanner That Eliminates False Positives
Pavan Jadav
Pavan Jadav

Posted on

I Built an Open-Source AI Security Scanner That Eliminates False Positives

🛡️ I Built an AI Security Scanner That Eliminates False Positives

Most security scanners have one big problem: too much noise.

You run a scan, get 200 alerts, and 160 of them are false positives. You spend hours triaging instead of actually fixing real bugs. Sound familiar?

That's exactly why I built Sentinel Bug Engine — an open-source, AI-powered security scanner with a unique Joint Verification Engine (JVE) that adversarially scores each finding before it ever reaches you.

👉 GitHub: github.com/pavan67-git/sentinel-bug-engine


🤔 The Problem with Existing Tools

Tool What it does well The gap
Semgrep Fast static patterns High false positive rate
Bandit Python-specific No verification step
Snyk Dependency scanning Commercial, not code-level
CodeQL Deep analysis Complex setup, GitHub-only

None of them answer the real question: "Is this alert actually exploitable in this specific codebase?"


💡 The Solution: 4-Tier Analysis Pipeline

Sentinel uses a layered approach:

┌────────────────────────────────────────────────────┐
│  Tier 1: Universal Parser                          │
│  → Detects language, extracts code blocks          │
├────────────────────────────────────────────────────┤
│  Tier 2: Deterministic Fast Filter                 │
│  → Regex + pattern rules (OWASP Top 10, CWEs)     │
│  → Sub-second scanning                             │
├────────────────────────────────────────────────────┤
│  Tier 3: LLM Cognitive Analysis (optional)         │
│  → Gemini / OpenAI / Claude / Ollama               │
│  → Catches business logic bugs, race conditions    │
├────────────────────────────────────────────────────┤
│  Tier 4: Joint Verification Engine (JVE) ← KEY    │
│  → Adversarially evaluates each finding            │
│  → Detects mitigating controls in context          │
│  → Assigns confidence score, discards FPs          │
└────────────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

🎯 The JVE — What Makes This Different

The Joint Verification Engine is the core differentiator.

For every potential finding, the JVE:

  1. Reads 20 lines of surrounding code context
  2. Looks for mitigating controls — sanitizers, parameterized queries, type guards, allowlists
  3. Assigns a confidence score (0–100%)
  4. Labels the finding: VERIFIED_TRUE_POSITIVE, LIKELY_TRUE_POSITIVE, SUSPECTED_FALSE_POSITIVE, or DISPROVED_FALSE_POSITIVE
  5. Automatically discards DISPROVED_FALSE_POSITIVE findings

Here's the result in practice:

# ❌ VULNERABLE — JVE flags this (LIKELY_TRUE_POSITIVE, 60%)
query = f"SELECT * FROM users WHERE id = {user_id}"
cursor.execute(query)

# ✅ SAFE — JVE correctly ignores this
cursor.execute("SELECT * FROM users WHERE id = ?", (int(user_id),))
Enter fullscreen mode Exit fullscreen mode

The second query uses parameterized inputs — JVE detects this and does not raise a false alert.


📊 Live Scan Results

I ran Sentinel against a deliberately vulnerable Flask app. Here's the real output:

=== SENTINEL SCAN RESULTS ===
Target: vulnerable_flask_app.py
Total findings: 6

  CRITICAL: 2
  HIGH: 3
  MEDIUM: 1

[SEC-001] CRITICAL | CWE-798  | Hardcoded API Secret or Private Key    | JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-002] HIGH     | CWE-89   | SQL Query Concatenation / Potential SQLi| JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-003] HIGH     | CWE-78   | Dangerous OS Command Execution          | JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-004] HIGH     | CWE-78   | Dangerous OS Command Execution          | JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-005] CRITICAL | CWE-502  | Unsafe Dynamic Code Evaluation (eval)  | JVE: LIKELY_TRUE_POSITIVE (60%)
[SEC-006] MEDIUM   | CWE-400  | Resource Leak (Unclosed File)          | JVE: LIKELY_TRUE_POSITIVE (60%)
Enter fullscreen mode Exit fullscreen mode

6 real findings. 0 false positives.

Every finding comes with:

  • CWE classification
  • OWASP Top 10 mapping
  • Exact file + line location
  • JVE verdict + confidence score
  • Suggested fix with code diff

🔍 What It Detects

CWE Vulnerability OWASP
CWE-89 SQL Injection A03:2021
CWE-78 OS Command Injection A03:2021
CWE-79 Cross-Site Scripting A03:2021
CWE-22 Path Traversal A01:2021
CWE-502 Unsafe Deserialization / eval A08:2021
CWE-798 Hardcoded Credentials A07:2021
CWE-918 Server-Side Request Forgery A10:2021
CWE-400 Resource Exhaustion / Leaks —
CWE-362 Race Conditions —
CWE-476 Null Pointer Dereference —

Supported languages: Python, JavaScript, TypeScript, Java, Go, PHP, C/C++, Rust, Ruby, C#


⚡ Quick Start

# Install
git clone https://github.com/pavan67-git/sentinel-bug-engine.git
cd sentinel-bug-engine
pip install .

# Basic scan
python -m src.cli scan ./your-project

# With SARIF export (for GitHub Advanced Security / GitLab CI)
python -m src.cli scan ./your-project --format both --sarif-out results.sarif

# Only HIGH+ severity
python -m src.cli scan ./your-project --min-severity HIGH

# With LLM reasoning (Gemini)
export GEMINI_API_KEY="your-key"
python -m src.cli scan ./your-project --enable-llm
Enter fullscreen mode Exit fullscreen mode

📦 CI/CD Integration

Sentinel exports SARIF 2.1.0 — the industry standard supported by:

  • ✅ GitHub Advanced Security (Code Scanning)
  • ✅ GitLab Security Dashboard
  • ✅ SonarQube
  • ✅ Azure DevOps

Add it to your GitHub Actions workflow:

- name: Run Sentinel Security Scan
  run: |
    pip install .
    python -m src.cli scan . --min-severity HIGH --format sarif --sarif-out results.sarif

- name: Upload SARIF to GitHub
  uses: github/codeql-action/upload-sarif@v3
  with:
    sarif_file: results.sarif
Enter fullscreen mode Exit fullscreen mode

🚀 What's Next

  • [ ] More LLM-powered rules (business logic, race conditions)
  • [ ] VS Code extension
  • [ ] Web dashboard (already in progress)
  • [ ] Benchmark against CVE database

🤝 Contribute

The project is open source and I'd love contributions — especially:

  • New detection rules for more CWEs
  • Language-specific pattern improvements
  • Benchmark datasets

👉 github.com/pavan67-git/sentinel-bug-engine


If this helped you or the JVE concept interests you, drop a ⭐ on GitHub or leave a comment below. Always open to feedback!

Top comments (0)