DEV Community

Tess Ainsley
Tess Ainsley

Posted on

The self-hosted AI review decision is really a data residency decision

The most common question about self-hosting an AI code reviewer is which tool is cheapest. The test that actually answers it puts the number somewhere else. Augment Code ran ten open-source AI code review tools against a 450K-file Python, TypeScript, Java and Go monorepo over 40+ hours, published 2026-01-16 and updated 2026-08-17, and the finding that reframes the decision is that the license is the free part. Everything around it costs money.

The published cost is not the license

Augment estimates a self-hosted stack at $4,100 to $9,100 a month at any team size, combining published GPU rates with 0.25 to 0.5 FTE of maintenance at the US Bureau of Labor Statistics mean developer wage. The comparison point in the same test is $24 to $30 per developer per month for a commercial per-seat reviewer. On that math, self-hosting is not the cheap option. It is the privacy and data-residency option, and the shape of the cost changes from a subscription line item to a capital and staffing one.

That range comes with a method caveat worth stating. The hourly components are itemized and the assembly is public, but it is an estimate built from rates, not a bill from a real deployment. The bottom of the range assumes modest GPU use. A team running large models on every pull request, or needing high throughput across many repos, should plan on the high end.

Install quality filters the list before features do

Ten tools went in. Three held up. The rest lacked maintenance, broke during configuration, gated the controls that matter behind a commercial license, or reviewed files in isolation. For anyone choosing off a self-hosted query, the install path is the first real signal, because a tool that will not stand up in a 40-hour test is not a tool you will keep.

SonarQube Community Build was the strongest on detection, with near-zero false positives across 21 languages, though its analysis runs on the main branch only. Semgrep came second on custom rules, and caught framework-specific patterns the generic tools missed.

On local inference, the results separate cleanly. Tabby self-hosted as documented and runs local models through Ollama, with review staying secondary to its code-completion focus. PR-Agent also supports local inference, but a configuration issue, #2098, caused silent fallback to hosted models during testing, which defeats the purpose of a local stack if nobody catches it. Kodus and Hexmos LiveReview run local models on your own infrastructure as well.

Two smaller entries show why maintenance status matters. anc95/ChatGPT-CodeReview is a reasonable free experiment. villesau/ai-codereviewer, the more familiar name, has shipped nothing since December 2023, and Augment notes it defaults to a model snapshot that OpenAI retires in October 2026. A self-hosted tool that depends on a retired model snapshot is not self-hosted in any useful sense.

CodeQL is worth naming for a different reason. If a team already pays for GitHub Code Security at $30 per active committer per month, it is the strongest fit and it caught data-flow vulnerabilities across multiple function calls in testing. If not, it is not the cheap entry.

"Open source" and "auditor-ready" are two questions

This is where the label and the checklist come apart. Most of these tools ship an open-source core and hold the enterprise controls behind a commercial key.

SonarQube Community Build gates audit logging at Enterprise Edition. Kodus puts SSO, RBAC and audit logs behind a commercial license key. Tabby documents no audit trail at all and places single sign-on on its Enterprise plan. So the honest answer to "is it open source" is yes for the code, and not by default for the controls a security review will ask about. If the reason you are self-hosting is compliance, check the control matrix before the license, because the license being free tells you nothing about whether the deployment will pass.

The requirement open source genuinely satisfies better than any hosted option is data residency. That is the axis the choice should start from, and it is why the published cost of a self-hosted stack matters less than it looks. You are not buying a cheaper product. You are buying a different placement of the model and the data, and paying for it in infrastructure and maintenance hours.

Where all ten stop

None of the tools in the test detected cross-service breaking changes across the four languages. Every one of them operates at file level. That is a real ceiling, and it sits exactly where agent-generated code creates risk. A file-scoped reviewer can read a diff and miss the caller, the invariant, or the ordering assumption the change breaks. Teams hit the same gap on any review that spans more than one repository, and no amount of local GPU capacity changes it.

How the options compare on the criteria that decide it

Tool Self-host path Enterprise controls Notable result in the test
SonarQube Community Build Yes Audit logs gated at Enterprise Edition; main-branch analysis only Strongest overall, near-zero false positives over 21 languages
Semgrep Yes Data not published for the tested build Second on custom rules, caught framework-specific patterns
Tabby Yes, worked as documented No audit trail documented; SSO on Enterprise Local inference via Ollama; review is secondary to completion
PR-Agent Yes, local inference supported Data not published Config issue #2098 caused silent fallback to hosted models
Kodus Yes, local models on your infrastructure SSO, RBAC and audit logs behind a commercial license key Included in the tested open-source set
Hexmos LiveReview Yes, local models on your infrastructure Data not published Included in the tested set
CodeQL Yes Tied to GitHub Code Security Caught data-flow vulnerabilities across function calls; $30 per active committer per month
villesau/ai-codereviewer Yes Not documented No commits since December 2023; depends on a model snapshot retiring October 2026

What to do with this

Decide the model placement first and the tool second. If code cannot leave your network, the shortlist is SonarQube Community Build, Semgrep, Tabby, PR-Agent, Kodus, LiveReview and CodeQL, and the tiebreakers are install evidence, which controls you actually need, and whether the repo is still active. If code can leave, the hosted per-seat figure is lower and the whole exercise is a price check rather than a residency decision.

Then run your own numbers with your own GPU sizing and your own maintenance estimate, because the published per-seat figures for hosted tools are stable while the monthly cost of self-hosting is the one teams most often leave out. The license is the only free part.

Read the original test for the parts I left out, including the per-tool configuration detail and the local-inference setup notes, at Augment Code's 450K-file monorepo test.

Top comments (0)