DEV Community

Superwavvy
Superwavvy

Posted on

How I Built an AI-Powered GitHub Action That Reviews PRs for Security Vulnerabilities

GuardPR in action


Every developer has done it. You're rushing to ship a feature, you open a pull request, and you merge it. Somewhere in that code is a SQL injection or a hardcoded secret.

Code reviews catch some of it. But reviewers are human. They miss things. And security experts are expensive,most startups don't have one on staff.

So I wanted to see if I could build a cheap alternative: an AI-powered GitHub Action that scans every PR for the OWASP Top 10, and posts a comment with the findings. Completely automated & No signup required.

What GuardPR Does

When a developer opens a pull request, GuardPR:

  1. Fetches the full content of every changed code file via the GitHub API
  2. Sends the code to an LLM with a security engineering prompt
  3. Parses the structured JSON response into a clean Markdown table
  4. Posts a comment on the PR with each vulnerability, severity, and suggested fix

Tech Stack
· Runtime: Node.js
· GitHub Integration: @octokit/rest (GitHub REST API)
· AI Analysis: Groq API (free tier, fast inference)
· CI/CD: GitHub Actions

Everything runs on the free tier. No paid APIs or expensive infrastructure.

Fetching the PR Files

The first step is getting the code. The GitHub API gives you two options:

  1. Fetch the diff: only the lines that changed
  2. Fetch the full file content: the entire file at the PR's head commit

Here's how I fetch the full content of every changed code file:

const { Octokit } = require("@octokit/rest");
const octokit = new Octokit({ auth: process.env.GITHUB_TOKEN });

async function getPRFiles(owner, repo, pullNumber) {
    const filesResponse = await octokit.pulls.listFiles({
        owner, repo, pull_number: pullNumber
    });

    let combinedContent = "";

    for (const file of filesResponse.data) {
        if (file.status === "removed") continue;
        if (!file.filename.match(/\.(js|ts|py|java|go|rb|php|sol)$/)) continue;

        const contentResponse = await octokit.repos.getContent({
            owner, repo,
            path: file.filename,
            ref: `refs/pull/${pullNumber}/head`
        });

        const decoded = Buffer.from(
            contentResponse.data.content, "base64"
        ).toString("utf-8");

        combinedContent += `\n\n===== FILE: ${file.filename} =====\n${decoded}`;
    }

    return combinedContent;
}
Enter fullscreen mode Exit fullscreen mode

The LLM Analysis

This is the heart of the tool. The prompt is everything:

async function analyzeCode(code) {
    const prompt = `You are an expert application security engineer.
    Analyze the following code for OWASP Top 10 vulnerabilities.

    Focus on:
    1. SQL Injection
    2. Cross-Site Scripting (XSS)
    3. Hardcoded Credentials
    4. Insecure Direct Object References (IDOR)
    5. Server-Side Request Forgery (SSRF)
    6. Broken Authentication
    7. Command Injection
    8. Path Traversal

    Return ONLY valid JSON in this exact format:
    {
      "vulnerabilities": [
        {
          "type": "SQL Injection",
          "severity": "HIGH",
          "location": "line 16",
          "description": "...",
          "fix": "..."
        }
      ]
    }`;

    const response = await fetch(
        "https://api.groq.com/openai/v1/chat/completions",
        {
            method: "POST",
            headers: {
                "Authorization": `Bearer ${process.env.GROQ_API_KEY}`,
                "Content-Type": "application/json"
            },
            body: JSON.stringify({
                model: "openai/gpt-oss-120b",
                messages: [{ role: "user", content: prompt }],
                response_format: { type: "json_object" },
                temperature: 0.1
            })
        }
    );

    const data = await response.json();
    return JSON.parse(data.choices[0].message.content);
}
Enter fullscreen mode Exit fullscreen mode

Three things to notice:

· response_format: { type: "json_object" } — this forces the LLM to return valid JSON. No more parsing chaos.
· temperature: 0.1 — this makes the LLM deterministic. Security scanning needs consistency, not creativity.
· Role assignment starting with "You are an expert security engineer" primes the model to think like an auditor.


The Bug That Almost Broke It

Here's the part I want to be honest about, because it taught me something important.

My first version only fetched the diff of the PR (the lines that changed). I tested it on a PR that added a single comment line. The bot said:

✅ No vulnerabilities detected. Great work!
Enter fullscreen mode Exit fullscreen mode

But my test file (login.js) had a blatant SQL injection sitting right there. It had been in the file for weeks.

The LLM wasn't wrong. The diff was clean. I was scanning the wrong thing.

The fix: fetch the full content of every changed file, not just the diff.

The Fail-Safe Lesson

The second bug was worse.

When the Groq API failed (wrong model name, network timeout), my code did this:

if (!data.choices || !data.choices[0]) {
    console.error("LLM response error:", JSON.stringify(data));
    return { vulnerabilities: [] };  // ❌ BAD
}
Enter fullscreen mode Exit fullscreen mode

It returned an empty array. Which meant "no vulnerabilities found." Which got formatted as "✅ Safe."

That's called fail-open, which is the most dangerous pattern in security. If your security tool can't do its job, it should never silently say everything is fine.

The fix was one line:

if (!data.choices || !data.choices[0]) {
    throw new Error("LLM analysis failed: " + JSON.stringify(data));
}
Enter fullscreen mode Exit fullscreen mode

Now, if the LLM fails, the whole script crashes. GitHub marks the check as red and there no false sense of safety.

This is one thing I've learnt that a security tool must always fail-safe.

The GitHub Action Wrapper

To make it run automatically on every PR, I added this workflow file:

name: GuardPR Security Scan

on:
  pull_request:
    types: [opened, synchronize, reopened]

jobs:
  security-scan:
    runs-on: ubuntu-latest

    permissions:
      pull-requests: write
      contents: read

    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: '20'
      - run: npm install
      - run: node scan.js
        env:
          GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
          GROQ_API_KEY: ${{ secrets.GROQ_API_KEY }}
          GITHUB_REPOSITORY: ${{ github.repository }}
          PR_NUMBER: ${{ github.event.pull_request.number }}
Enter fullscreen mode Exit fullscreen mode

Note the permissions block. I'm explicitly limiting what the Action can do like read code and write PR comments, nothing else. This is the principle of least privilege, and it's a security fundamental.

What I Learned

A few things that surprised me:

  1. Prompt engineering is half the work. The difference between a good and bad prompt is the difference between "found nothing" and "caught the SQL injection on line 16."
  2. Structured output is non-negotiable. response_format: json_object turned fragile string parsing into clean JSON.
  3. Fail-safe > fail-open. Silent failures in security tools are worse than loud ones. If you can't scan, say so.
  4. Context matters: Scanning a diff is cheap but dumb. Scanning the full file is slower but actually useful.

You can try it yourself:

· Repo: github.com/superwavvy/guardpr
· Live demo: guardpr-test/pull/1

The workflow runs in about 30 seconds per PR. It costs nothing on the Groq free tier. And it's found real bugs in my test repo.

Top comments (0)