DEV Community

Cover image for I Built an AI Agent That Fixes Sonar and Snyk Findings.
Nitesh Rawal
Nitesh Rawal

Posted on

I Built an AI Agent That Fixes Sonar and Snyk Findings.

AutoRemediate AI is a program that automatically fixes certain code-quality and security problems in Java projects.
github url: https://github.com/nitesh401/AutoRemediate-AI

Every company that writes software gets long lists of small problems from scanners:

  • "This import is never used."
  • "This library version has a known security hole. Upgrade it."

These are boring, repetitive and numerous. Developers postpone them, and they pile up. Fixing them by hand takes time, and fixing them carelessly can break the app.

Idea: let a robot do the boring ones but make it prove each fix is safe, so humans only review results instead of doing the work.

Why not just ask an AI "fix everything"? Because an AI can be confidently wrong. It could break the app, invent a library version that does not exist, or delete something important. So we wrap the AI in strict rules and checks.

Every team has the same backlog: unused imports, unused variables, and libraries with known vulnerabilities. Boring to fix, easy to postpone, risky to fix carelessly.

So I built AutoRemediate AI, an agent that fixes the safe ones for you. The twist is that it does not trust the AI to say a fix is safe.

The AI proposes. Deterministic tools verify. A human approves.

Why "just ask the LLM to fix everything" is a bad idea

An LLM can be confidently wrong. It can break behavior, invent a library version that does not exist, or follow instructions hidden in a code comment. A passing sentence like "this change is safe" from a model is an opinion, not proof.

Proof is a green build, passing tests, and a clean re-scan. So I built the agent around that idea: the model only does the thinking part, and plain code does everything that must be reliable.

What it does

Give it a Java/Maven repository. It will:

  1. Clone it into an isolated workspace and create an ai/remediation/<job> branch (your main branch is never touched).
  2. Run a baseline: mvn clean verify, then read the test reports.
  3. Scan with Sonar (code quality) and Snyk (vulnerable dependencies).
  4. Classify each finding and pick only high-confidence ones.
  5. Ask Claude for a minimal patch, validate it, and have a second call review it.
  6. Rebuild, re-run tests, re-scan, and compare before vs after.
  7. Commit if the safety gate passes. Otherwise roll back and retry (max 3 attempts).
  8. Write a Markdown/JSON report and, optionally, open a draft pull request.
clone -> baseline -> scan -> classify
      -> for each safe finding:
           LLM plan -> validated patch -> review
           -> build -> tests -> rescan -> SAFETY GATE
              PASS: commit      FAIL: reset, diagnose, retry / revert
      -> report -> optional draft PR (never auto-merged)
Enter fullscreen mode Exit fullscreen mode

The safety gate (the part I care about most)

A fix is accepted only if all of this is true, checked by normal code with no LLM involved:

  • Baseline build was green, and the build is green after the fix
  • All tests pass, and the test count did not drop (no "fixing" by deleting tests)
  • The target finding is actually gone
  • No new Sonar issue and no new Snyk vulnerability
  • Every scanner that worked before still works afterwards (a failed rescan means "cannot prove", so reject)

Sonar results are compared as fingerprints (rule | file | normalized message), so deleting a line does not make every issue below it look new. Snyk results are compared by vulnerability ID plus package, ignoring the version the fix changes.

Keeping the AI on a short leash

  • Tiny edits, not rewrites. The model returns JSON search/replace edits. The text to replace must match the file exactly once, otherwise nothing is written.
  • Strict file and size rules. Sonar fixes may touch only .java files, Snyk fixes only pom.xml, with a cap on files and changed lines.
  • Dependency upgrades are done by code. A PomVersionUpdater changes one version number. It refuses managed versions and shared properties that would upgrade unrelated libraries. The model can only confirm risk, and the version must equal Snyk's recommendation, so it cannot invent one.
  • Fail closed. If the review step errors or is unclear, the patch is rejected.
  • Only safe categories are automated. Security findings, sensitive paths such as auth/crypto, transitive dependencies and major upgrades are reported for humans.

Security, because it runs untrusted code

The agent builds and tests other people's repositories, so:

  • One CommandRunner starts every process: allow-list of mvn, mvnw, git, snyk, no shell, timeouts, bounded output.
  • Child processes get a scrubbed environment, so build and test code never sees your API keys.
  • Repository content is treated as untrusted data (prompt injection defense) and templates substitute values in a single pass.
  • Pushes to main, master, develop and release are refused, and there is no merge capability at all.
  • Secrets live in environment variables and are redacted from logs and reports.
  • Sonar scans run against a scratch project, so your real project history is not overwritten.

The job state machine

Instead of an open-ended "agent loop", every job moves through explicit states, and illegal transitions throw:

CREATED -> CLONING -> BASELINE -> SCANNING -> CLASSIFYING -> REMEDIATING
 -> BUILDING -> TESTING -> RESCANNING -> VERIFYING
      PASS -> next finding or REPORTING -> COMPLETED
      FAIL -> RETRY -> ... -> REVERTED
Enter fullscreen mode Exit fullscreen mode

If the untouched repo already fails its own tests, the job stops with BASELINE_FAILURE. The agent will not blame itself for problems that existed before it arrived.

The stack

Java 21, Spring Boot, Maven, Git, Sonar Web API, Snyk CLI, Anthropic Messages API (model name is configuration, not code), JUnit 5, Docker. The LLM sits behind a small LlmClient interface, so swapping providers is a configuration-level change.

What I tested, and what I did not

Honest status: the deterministic core (patching, pom editing, the safety gate, classification, Git isolation, command allow-list, parsers) is covered by unit tests, and I ran a behavior harness over those classes. The full run against a live Sonar server, Snyk and the Claude API is the next milestone, and I list the known limits in the README: single-module focus, full-build verification instead of targeted tests, in-memory job store, GitHub-only PRs.

What I learned

  1. Make the world measurable first. Baselines and before/after comparisons are what turn an AI suggestion into something you can trust.
  2. Constrain the output, then validate it. A strict JSON contract plus code-level checks beats a clever prompt.
  3. Use the LLM only where thinking is needed. Git, Maven, scanning and comparing stay boring and deterministic, which is also cheaper.
  4. Design for rollback. Checkpoint, reset, retry a few times, stop.
  5. Keep a human in the loop. Draft PRs, no auto-merge.

Try it / code

Repository: https://github.com/nitesh401/AutoRemediate-AI

If you are building agents that touch real code, I would love to hear how you decide what is safe to automate. Comment below.

Top comments (0)