DEV Community

Daniel Paiva
Daniel Paiva

Posted on

I built a PR reviewer that survived its own model being deprecated

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🀝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

groq-pr-reviewer-net β€” a .NET 10 CLI that reads your git diff and returns a structured code review in your terminal, before you ever open the pull request.

dotnet run -- --staged
Enter fullscreen mode Exit fullscreen mode

That's it. It reviews bugs and correctness, security, performance, and readability.

The friend I built it for is the developer who codes alone. No teammate to tag for a second opinion, no budget for a paid review bot β€” the solo maintainer, the student, the person shipping a side project at 1am. I've been that person on every side project I've ever started.

The problem it solves is specific: the moment right before you commit, when you know you should have someone look at this, and there is nobody to ask. Paid AI review tools answer that with a seat license and a company card. This answers it with a terminal and a free API key.

Demo

A CLI has no deployed link, so here is the whole thing end to end.

1. Setup you can verify without leaking anything

Terminal running dotnet run --check, showing key source, length, gsk_ prefix ok, and the model name

--check reports where the key was loaded from, how long it is, and whether it has the right shape β€” without ever printing the key. One command tells you if you are ready, before you spend a single request.

2. The open-weight catalogue

Terminal listing the models available on Groq, with no Llama chat model present

Every model my key can actually reach. Pay attention to what is missing from that list β€” it matters in a moment.

3. The model I built the whole thing on, dead

Terminal showing a 404 error: the model llama-3.3-70b-versatile does not exist, with a hint to use --model

No Llama chat model is left in the catalogue. Notice that the failure explains itself: it names the likely cause, links the live model list, and tells you which flag fixes it. That was the one thing I most wanted to get right, because I'd just lost an hour to an error message that told me nothing.

4. Same diff, one flag different, working review

Terminal showing a full code review with findings grouped into bugs, security, performance and best practices

Identical command minus one argument. That gap between screenshot 3 and screenshot 4 is the entire argument of this post, and it cost me about forty seconds.

Dogfooding: the tool reviewing its own source

Screenshot 4 is the tool reviewing the commit that implements it. Earlier passes, against earlier versions of this code, caught three things worth showing β€” all since fixed, which is why you won't find them in the screenshot above.

It found a real bug I had introduced minutes earlier β€” I'd changed the default model constant and the README, but forgot the --help text:

Inconsistent default model description – PrintUsage() documents the default model as llama-3.3-70b-versatile, while the code constant DefaultModel is openai/gpt-oss-120b. Users may be confused about which model is actually used.

It caught an unguarded JSON access that would throw if the API schema ever shifted:

Potential JSON-parsing failure in FetchModelIds – the method assumes the response contains a top-level "data" array. If Groq changes the schema or returns an error payload without that property, GetProperty("data") will throw an unhandled KeyNotFoundException.

And it flagged something missing from my own threat model entirely:

Potential secret leakage – the diff itself is sent to Groq unchanged; if the diff contains passwords, tokens, or other secrets they will be transmitted to an external service.

That last one became a warning section in the README. The tool wrote its own security disclaimer.

And it keeps earning its keep: the run in screenshot 4 flagged that I never dispose the HttpResponseMessage in either API call β€” a socket leak I had walked straight past.

Being honest about the failure modes

It is not magic. Across two earlier review passes it insisted, both times, that net10.0 was not a valid target framework and that I was missing a using System.Linq;. Both wrong β€” the project builds clean with zero warnings. The model's knowledge cutoff predates .NET 10, and ImplicitUsings already covers LINQ.

Screenshot 4 has its own share: it claims the diff is "printed to stdout before being sent," which simply never happens. Roughly one in four findings is noise.

I put that in the README verbatim, because a tool that oversells itself is worse than no tool:

Treat the output as a fast second opinion, not as truth. Read it the way you would read a well-meaning junior reviewer.

A junior reviewer who is occasionally confidently wrong but catches things you missed is still worth having. Pretending otherwise would be the actual failure.

Code

GitHub logo danhpaiva / groq-pr-reviewer-net

Terminal code review for your git diff, powered by an open-weight model on Groq. One C# file, no SDK, no cost.

groq-pr-reviewer-net

.NET Model Weights Inference License Hacktoberfest

A C# (.NET 10) CLI that reviews your git diff using an open-weight model (openai/gpt-oss-120b, Apache 2.0) running on Groq's ultra-fast inference.

Point it at any git repository and it prints a structured code review in your terminal β€” before you open the pull request.

dotnet run -- --staged
Enter fullscreen mode Exit fullscreen mode

A terminal showing a full code review with findings grouped into bugs and correctness, security, performance, and best practices

Above: the tool reviewing its own source code.

Why open-source AI here?

  • Open weights. The default model ships under Apache 2.0: you can swap it for another one at any time, or run it elsewhere. No vendor lock-in β€” and when a model is retired, --list-models plus --model gets you moving again without touching the code.
  • Zero cost. Groq's free tier comfortably covers an individual developer's usage.
  • No dependency on paid review tooling. Anyone in the community can clone this and use it today.
  • Your diff stays in your control. You choose the provider and the model, instead…

The whole thing is under 400 lines across five small files β€” deliberately small enough to read in one sitting and fork.

The core is unremarkable on purpose, which is the point: an open-weight model behind an OpenAI-compatible endpoint means no SDK, no framework, no abstraction layer to learn.

// GroqClient.cs
private const string ChatCompletionsUrl = "https://api.groq.com/openai/v1/chat/completions";
private const string ModelsUrl = "https://api.groq.com/openai/v1/models";

// CliOptions.cs
// Open-weight (Apache 2.0) and currently the strongest chat model on Groq.
// Run --list-models if this one is ever retired, then pass --model.
public const string DefaultModel = "openai/gpt-oss-120b";
Enter fullscreen mode Exit fullscreen mode

The system prompt is plain text, not a framework construct β€” which is exactly why swapping models costs nothing:

private const string SystemPrompt = """
    You are a senior code reviewer. Analyse the pull request diff below and reply in English,
    as bullet points organised into these sections:
    - Bugs and correctness
    - Security
    - Performance
    - Best practices / readability

    Be concise and specific. If a section has nothing worth raising, write "Nothing to flag".
    Ignore trivial formatting changes.
    """;
Enter fullscreen mode Exit fullscreen mode

MIT licensed.

How I Built It

The open-source AI: openai/gpt-oss-120b, an open-weight model released under Apache 2.0, served on Groq for inference. No proprietary model touches this project.

The stack: .NET 10, five small files, and HttpClient. No agent framework, no orchestration library, no vector database. The entire dependency list is the .NET base class library.

That minimalism is a design choice tied directly to the open-weight decision. Because open models are served behind the same OpenAI-compatible /chat/completions contract, I didn't need an abstraction layer to stay portable β€” the HTTP contract is the abstraction layer. A 30-line HttpClient call is already provider-agnostic. Anything heavier would have added lock-in rather than removing it, and would have been one more thing for a forker to learn.

The flow is four steps: shell out to git diff, truncate at 60k characters to respect the context window, POST it with the system prompt, print the response.

Two things took longer than the happy path:

Reading streams without deadlocking. Draining git's stdout and stderr sequentially will hang the process if stderr fills its pipe buffer while you're still reading stdout. Both reads have to be in flight before you wait:

var outputTask = process.StandardOutput.ReadToEndAsync();
var errorTask = process.StandardError.ReadToEndAsync();
await Task.WhenAll(outputTask, errorTask);
await process.WaitForExitAsync();
Enter fullscreen mode Exit fullscreen mode

Making the setup failure legible. Which brings me to the hour I lost.

Groq is not Grok

Groq is an inference provider for open models; keys start with gsk_. xAI's Grok is an unrelated company with proprietary models; keys start with xai-. The names differ by one letter.

I generated a key on the wrong console and spent an hour staring at 401 Invalid API Key.

So I built the diagnostic I wished I'd had:

$ dotnet run -- --check
Key source  : /path/to/groq-pr-reviewer-net/.env
Length      : 56 characters
gsk_ prefix : ok
Model       : openai/gpt-oss-120b
Enter fullscreen mode Exit fullscreen mode

It reports where the key was loaded from, its length, and whether it has the right shape β€” without ever printing the key itself. The 401 message now names the gsk_ vs xai- mix-up outright, and the README carries the warning up front.

Small thing. But "build for a friend" means thinking about the person hitting your error message at midnight, and that person has no idea the two companies exist.

Why Does Open Innovation Matter?

I was going to write the usual paragraph about open weights and avoiding vendor lock-in. Then the argument proved itself while I was still building.

Halfway through, my model was deprecated.

I'd built the whole thing on Llama 3.3 70B. When I ran the first real end-to-end test against the API:

Groq API error (404 NotFound): The model `llama-3.3-70b-versatile`
does not exist or you do not have access to it.
Enter fullscreen mode Exit fullscreen mode

Not a single Llama chat model was left in the catalogue.

Here's what a closed API would have cost me: a new SDK, a new request shape, a new auth flow, a rewritten prompt, and a wait on somebody's migration guide. Possibly a dead weekend project.

What it actually cost me was a flag.

dotnet run -- --list-models      # what can I actually reach right now?
dotnet run -- --model qwen/qwen3.8-27b
Enter fullscreen mode Exit fullscreen mode

I moved to openai/gpt-oss-120b, re-ran the review, and shipped. Same endpoint, same request body, same prompt. I added --list-models during the recovery so the next person hitting a dead model can see their options in one command instead of digging through changelogs.

That is what open innovation made possible here, concretely: the difference between a flag change and a rewrite. Interchangeable models behind a common contract mean no single vendor's roadmap can end your project.

The rest of the case holds too, and matters specifically for the friend I built this for:

  • Zero cost. Groq's free tier covers an individual developer comfortably. No card, no trial clock β€” the barrier for a student or a solo maintainer is actually zero, not "zero for 14 days."
  • A real escape hatch. Apache 2.0 weights can be run elsewhere, including on your own hardware. If every hosted option vanished tomorrow, the model would still exist. You cannot say that about a closed endpoint.
  • You choose where your code goes. Point it at a different provider, or at a local inference server, and your diff never leaves your machine. With a closed tool you get whatever data policy the vendor ships that quarter.

For someone with no team and no budget, "I can swap the engine myself" isn't an abstract principle. It's whether the tool still works next year.


Built in a weekend with .NET 10 and an open-weight model that costs nothing to run.

Top comments (0)