<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Tom Seidel</title>
    <description>The latest articles on DEV Community by Tom Seidel (@tmseidel).</description>
    <link>https://dev.to/tmseidel</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2067637%2Fe2e8a4a1-f828-4841-bb86-e4a808b65216.jpg</url>
      <title>DEV Community: Tom Seidel</title>
      <link>https://dev.to/tmseidel</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tmseidel"/>
    <language>en</language>
    <item>
      <title>Low-Code Automation's 2026 Hype Isn't Fading. That's the Problem.</title>
      <dc:creator>Tom Seidel</dc:creator>
      <pubDate>Thu, 24 Sep 2026 07:18:10 +0000</pubDate>
      <link>https://dev.to/tmseidel/low-code-automations-2026-hype-isnt-fading-thats-the-problem-35p8</link>
      <guid>https://dev.to/tmseidel/low-code-automations-2026-hype-isnt-fading-thats-the-problem-35p8</guid>
      <description>&lt;p&gt;People keep asking whether the low-code automation hype — n8n, Zapier, Make, Coze, Dify, LangFlow — is finally cooling off in 2026. I researched it. The answer is worse for the skeptics than you think: the low-code automation market is not cooling. It's at its peak. The thing that's in the trough is the promise it's built on — autonomous AI agents. And that's exactly why I think code-first is the winning play.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data: low-code automation is not fading
&lt;/h2&gt;

&lt;p&gt;If you ask "is the low-code automation hype over in 2026?", the numbers say no:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;n8n more than doubled its valuation in 2026.&lt;/strong&gt; $2.5B after its $180M Series C (October 2025) [1], then $5.2B when SAP took a strategic stake in May 2026 — and signed a multi-year deal to embed n8n's workflow canvas inside Joule Studio, SAP's AI agent-building platform [2][3].&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;195k+ GitHub stars&lt;/strong&gt; [4], a growing "fair-code" community, and a flagship conference ("In The Loop 2026") that's now an industry event in its own right.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zapier shipped "Agents" with 7,000+ app integrations. Make, Coze, and Dify are all expanding their AI-node feature sets&lt;/strong&gt; [5]. The entire category is racing to add "AI" to the visual canvas.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So let's not pretend the 2010-style "it all died out" narrative applies here. It doesn't. The low-code automation market in 2026 is the strongest it's ever been.&lt;/p&gt;

&lt;p&gt;So &lt;strong&gt;the market isn't growing because the visual canvas got better. It's growing because the canvas is now the delivery vehicle for agentic AI.&lt;/strong&gt; n8n's own homepage says it: "Build visually, go deep with code" — for AI agents [6]. Zapier sells autonomous agents [5]. SAP is buying n8n &lt;em&gt;to put the visual canvas inside its agent platform&lt;/em&gt; [2][3]. The visual workflow builder has quietly become the industry's preferred way to ship "AI agents" to business users.&lt;/p&gt;

&lt;p&gt;And now we get to what &lt;em&gt;is&lt;/em&gt; in the trough.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is fading: the agentic AI promise underneath the canvas
&lt;/h2&gt;

&lt;p&gt;In April 2026, Gartner published its first dedicated &lt;strong&gt;Hype Cycle for Agentic AI&lt;/strong&gt; — separated from the broader GenAI cycle for the first time [7]. Its verdict: AI agents are at the &lt;strong&gt;peak of inflated expectations&lt;/strong&gt; and will slide into the &lt;strong&gt;trough of disillusionment throughout 2026&lt;/strong&gt;. Gartner calls 2026 the "year of disillusionment" for agentic AI [8].&lt;/p&gt;

&lt;p&gt;The supporting data:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;An 88% security incident rate for AI agents&lt;/strong&gt; (NoCode.Tech's coverage of enterprise incident data) [9] — what happens when autonomous decision-making touches live infrastructure without a governance layer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Microsoft pausing data-center construction. Anthropic delaying a major release over safety concerns&lt;/strong&gt; [10]. The broader AI sector is hitting real headwinds in 2026.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gartner's framing:&lt;/strong&gt; the trough "reflects how agentic AI was &lt;em&gt;sold and bought&lt;/em&gt;, not whether the technology works." One line, the whole 2010–2013 story.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Read those two facts together — low-code automation at an all-time peak, agentic AI heading into its trough — and the picture gets sharp: &lt;strong&gt;low-code automation has become the Trojan horse for the agentic AI hype.&lt;/strong&gt; The growth in workflow platforms is largely growth in "AI agent" workflows: LLM nodes on the canvas, natural-language workflow building, agents that "reason" between connectors. When the agentic trough lands — and Gartner says it lands in 2026 [7][8] — the platforms won't shrink, but a lot of what their users built will be graphs of LLM calls that nobody can explain, test, or trust.&lt;/p&gt;

&lt;p&gt;That's a more dangerous position than the old hype, not a less dangerous one. In 2010, a broken BPM graph was a maintenance chore. In 2026, a broken 400-node agent graph is a silent compliance and security incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the 2010 wave matters: we saw this ending once
&lt;/h2&gt;

&lt;p&gt;Around 2010 the enterprise world got a structurally identical pitch: SOA orchestration, ESB tools, Mule, TIBCO, jBPM, Camunda's rise, BPMN modelers, iPaaS. "Developers are the bottleneck — let business people wire integrations visually. No code. Citizen developers. Done."&lt;/p&gt;

&lt;p&gt;The ending is well documented: abstraction tax, version drift, connector sprawl, "who owns this when the modeler leaves." The industry went back to code. The only piece of that wave that survived is the BPMN modeler &lt;em&gt;as a companion to the IDE&lt;/em&gt; — modeling survived, implementation did not.&lt;/p&gt;

&lt;p&gt;The 2026 wave repeats the same core move (visual abstraction over real integration logic) and adds one new ingredient: nondeterminism inside the graph. That's the difference between a 2013 problem and a 2026 problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why AI makes the visual canvas worse, not better
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Two hidden layers of nondeterminism.&lt;/strong&gt; A 200-node graph already has a black box in every connector. Now 15 of those nodes are LLM calls that return something different every run. Model nondeterminism stacked on platform nondeterminism: the pipeline isn't flaky anymore, it's &lt;em&gt;unprovable&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. "AI builds your workflow from a description" just moves the drift into the prompt.&lt;/strong&gt; Hallucinated field names, mismatched connectors, silently best-effort transformations. The 2010 wave hid complexity behind boxes; this wave hides it behind a chat.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Observability and review stop working.&lt;/strong&gt; How do you trace a 400-node graph with 12 LLM calls in it? Diff it? Regression-test it? The visual format was never a good git citizen; with LLM nodes it becomes unauditable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The abstraction tax now includes the model's limits.&lt;/strong&gt; Every LLM node is a context window, a token cost, a hallucination risk, and a latency spike — wrapped in a pretty node. Your "simple" workflow is a distributed system with a stochastic dependency you can't upgrade.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. The platforms know the ceiling exists — that's why they all ship code nodes.&lt;/strong&gt; n8n's feature list literally says "fall back to code: drop into JavaScript or Python in a code node" [11]. The escape hatch is part of the product. That is the entire low-code thesis conceding defeat, shipped as a feature.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. "Low-code" was never low-maintenance; it just &lt;em&gt;looked&lt;/em&gt; that way.&lt;/strong&gt; AI makes it look even more low-maintenance. The work doesn't disappear — it becomes invisible until an LLM node returns a different JSON shape on a Tuesday at 2 a.m. and the pipeline quietly stops.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why code-first wins — especially in the AI era
&lt;/h2&gt;

&lt;p&gt;The architecture for AI automation is &lt;strong&gt;deterministic core, stochastic edges&lt;/strong&gt;: code that orchestrates, with well-defined LLM calls at the edges — schema-validated, tested, version-controlled.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;AI collapsed the cost of code.&lt;/strong&gt; Low-code's historical advantage was access for people who can't code. AI-assisted coding erased that. Code is now the &lt;em&gt;fast&lt;/em&gt; path — generated in seconds, reviewed in minutes, diffable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Determinism is a feature.&lt;/strong&gt; You can unit-test, replay, and prove a deterministic pipeline. You cannot prove a graph of LLM calls. In production, "it usually works" is a failure state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review and rollback survive.&lt;/strong&gt; A 200-line diff is reviewable. A 400-node YAML blob is not. AI writes the code; you review the diff. That combination is the whole game.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You control the whole stack.&lt;/strong&gt; Containers, CI, secrets, observability, backup, upgrades. Code-first gives you the levers; the canvas gives you a vendor roadmap — which in 2026 is literally "embedded in SAP's agent platform" or "not" [2][3].&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Complexity doesn't scale linearly.&lt;/strong&gt; A well-factored 300-line module covering 15 integrations is maintainable. 150 nodes is not.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Low-code in the AI era isn't "more automation." It's &lt;em&gt;more hidden nondeterminism with a bigger funding round.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where low-code still makes sense
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Throwaway prototypes and MVPs.&lt;/strong&gt; Weeks, not years. Validate, then rewrite in code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trivial one-offs.&lt;/strong&gt; "Email → Slack," "CSV from A to B." No state, low risk.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Business-owned, low-risk internal flows.&lt;/strong&gt; When the business genuinely owns the maintenance and the flow isn't regulated or high-volume.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When visual authoring &lt;em&gt;is&lt;/em&gt; the feature.&lt;/strong&gt; Some process changes are inherently visual — and the 2010 wave proved that modeling (BPMN) survives even when implementation doesn't.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Rule of thumb: &lt;em&gt;low complexity + low risk + short-lived = fine on a canvas.&lt;/em&gt; The moment you need state, idempotency, scale, security, testability, review, or portability — the decision flips to code, and in the AI era that flip comes &lt;em&gt;sooner&lt;/em&gt;, because AI made the code path cheaper, not more expensive.&lt;/p&gt;

&lt;h2&gt;
  
  
  The controversial part
&lt;/h2&gt;

&lt;p&gt;The same buyers who bought the 2010 SOA/iPaaS dream are buying the 2026 dream in the same clothes with a new logo: end the developer bottleneck, visually. The dynamic hasn't changed — &lt;strong&gt;abstraction doesn't remove complexity, it hides it.&lt;/strong&gt; In 2010 the hidden complexity became a maintenance liability. In 2026 it becomes &lt;em&gt;hidden nondeterminism inside an LLM&lt;/em&gt; — which is exactly where outages, runaway token costs, and compliance breaches breed.&lt;/p&gt;

&lt;p&gt;And the "is it cooling?" question deserves a straight answer: &lt;strong&gt;the low-code canvas market isn't cooling — it's at its peak, worth billions, embedded in SAP [2][3]. What's cooling is the agentic AI promise the whole category rebranded itself around, and Gartner says that trough hits in 2026 [7][8].&lt;/strong&gt; So the risk isn't that low-code automation dies. The risk is that it survives on top of a hype cycle entering its disillusionment phase, leaving every organization with hundreds of LLM-node graphs that can't be explained, tested, audited, or trusted — and a platform roadmap deciding what they can and can't do.&lt;/p&gt;

&lt;p&gt;Code-first isn't developer ego. It's the only lever that doesn't get linearly more expensive as complexity grows, and the only one that survives a trough. AI didn't change that — AI made the code path cheaper and the canvas path more dangerous, in the same year.&lt;/p&gt;

&lt;p&gt;We saw the 2010 ending in 2013. The 2026 ending won't be a shrinking market. It'll be a peak-sized market full of graphs nobody can debug.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;n8n Blog, "n8n raises $180M to get AI closer to value with orchestration" (Oct 9, 2025) — $2.5B Series C led by Accel. &lt;a href="https://blog.n8n.io/series-c/" rel="noopener noreferrer"&gt;https://blog.n8n.io/series-c/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;AyAutomate, "n8n Hits $5.2B Valuation: Enterprise-Ready in 2026" — SAP strategic stake (May 2026), canvas embedded inside Joule Studio (SAP's AI agent-building platform). &lt;a href="https://www.ayautomate.com/blog/n8n-enterprise-valuation-2026" rel="noopener noreferrer"&gt;https://www.ayautomate.com/blog/n8n-enterprise-valuation-2026&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Wikipedia, "N8n" (accessed Sep 2026) — SAP strategic investment, May 2026, roughly doubling the valuation; founded by Jan Oberhauser, launched Oct 2019. &lt;a href="https://en.wikipedia.org/wiki/N8n" rel="noopener noreferrer"&gt;https://en.wikipedia.org/wiki/N8n&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;GitHub, "n8n-io/n8n" — 195k+ stars; "fair-code platform to build and deploy AI agents and workflows, 1500+ integrations." &lt;a href="https://github.com/n8n-io/n8n" rel="noopener noreferrer"&gt;https://github.com/n8n-io/n8n&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Virtua, "n8n vs Zapier vs Make: Kosten, DSGVO &amp;amp; Datenschutz (2026)" — Zapier Agents with 7,000+ app integrations. &lt;a href="https://www.virtua.cloud/learn/de/concepts/n8n-vs-zapier-vs-make-kosten-datenschutz" rel="noopener noreferrer"&gt;https://www.virtua.cloud/learn/de/concepts/n8n-vs-zapier-vs-make-kosten-datenschutz&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;n8n homepage — "Build visually, go deep with code, connect to anything. Every step of your agents' reasoning, traceable on the canvas." &lt;a href="https://n8n.io/" rel="noopener noreferrer"&gt;https://n8n.io/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Gartner, "2026 Hype Cycle for Agentic AI" (Apr 15, 2026) — first dedicated agentic AI hype cycle; AI agents at peak of inflated expectations, trough of disillusionment predicted throughout 2026. &lt;a href="https://www.gartner.com/en/articles/hype-cycle-for-agentic-ai" rel="noopener noreferrer"&gt;https://www.gartner.com/en/articles/hype-cycle-for-agentic-ai&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Generative, Inc., "Agentic AI in 2026: How AI Evolved from Chatbots to Autonomous Agents" (Jul 10, 2026) — 2026 as the "year of disillusionment" for agentic AI. &lt;a href="https://www.generative.inc/agentic-ai-in-2026-how-ai-went-from-chatting-to-doing" rel="noopener noreferrer"&gt;https://www.generative.inc/agentic-ai-in-2026-how-ai-went-from-chatting-to-doing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;NoCode.Tech, "Gartner's 2026 Hype Cycle for Agentic AI: What the Trough of Disillusionment Actually Means for No-Code Teams" (Jul 21, 2026) — 88% AI agent security incident rate. &lt;a href="https://www.nocode.tech/article/gartners-2026-hype-cycle-agentic-ai" rel="noopener noreferrer"&gt;https://www.nocode.tech/article/gartners-2026-hype-cycle-agentic-ai&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Financial Express, "Is an AI bust coming?" (Apr 12, 2026) — Microsoft pausing data-center construction, Anthropic delaying a release over safety concerns. &lt;a href="https://www.financialexpress.com/market/global-market-pulse/is-an-ai-bust-coming/4204135/" rel="noopener noreferrer"&gt;https://www.financialexpress.com/market/global-market-pulse/is-an-ai-bust-coming/4204135/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;n8n, "Workflows App Automation Features" — "Fall back to code: drop into JavaScript or Python in a code node, add npm packages when you self-host." &lt;a href="https://n8n.io/features/" rel="noopener noreferrer"&gt;https://n8n.io/features/&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>lowcode</category>
      <category>nocode</category>
      <category>agents</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Jev Is Fast Because the Reasoning Moved Into Your Architecture</title>
      <dc:creator>Tom Seidel</dc:creator>
      <pubDate>Wed, 23 Sep 2026 06:43:57 +0000</pubDate>
      <link>https://dev.to/tmseidel/jev-is-not-a-faster-llm-it-is-a-bill-for-the-architecture-you-skipped-4k36</link>
      <guid>https://dev.to/tmseidel/jev-is-not-a-faster-llm-it-is-a-bill-for-the-architecture-you-skipped-4k36</guid>
      <description>&lt;p&gt;The hype train is packed again. Ever since TypeSafe AI released Jev, the internet has been&lt;br&gt;
full of examples that make it look like the fast, cheap alternative to a classical LLM:&lt;br&gt;
70–500 ms instead of seconds, $0.042 per million input tokens, output "too cheap to meter", 40x to 200x speedups. The conclusion is drawn implicitly and sometimes explicitly:&lt;br&gt;
&lt;em&gt;why pay for a reasoning model, when this thing answers instantly for nothing?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That conclusion is wrong, and it is wrong in a way that costs money. Jev is not a cheaper way to do what an LLM does. It is a different computation for a different problem — and it quietly moves the expensive part of the work into your codebase, where the hype slides, the demos and the LinkedIn posts never look.&lt;/p&gt;

&lt;p&gt;Here is the honest version.&lt;/p&gt;
&lt;h2&gt;
  
  
  The speed is not an optimization, it is a different problem
&lt;/h2&gt;

&lt;p&gt;Give a reasoning model this task:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Should this pull request be merged? Analyze the changes, consider security, tests,&lt;br&gt;
architecture and possible regressions.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Roughly this happens: context, understand, form hypotheses, intermediate steps, compare, decide, and finally &lt;em&gt;formulate an answer&lt;/em&gt;. That last step is where the bill is. An LLM generates autoregressively, token by token, conditioned on the token before it. Structured JSON output does not change that — it is still one token at a time, just with a schema.&lt;/p&gt;

&lt;p&gt;Jev gets a different shape of input:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;STATE:
- change touches OAuth callback handling
- security scan: no findings
- tests: 142/142 passing
- API compatibility: unchanged
- review finding: possible missing null check
- project policy: severity &amp;gt;= HIGH blocks merge

QUESTIONS:
1. finding_severity = [NONE, LOW, MEDIUM, HIGH, CRITICAL]
2. merge_blocking = [yes, no]
3. needs_human_review = [yes, no]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and returns a distribution:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;MEDIUM&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;0.74&lt;/span&gt;
  &lt;span class="na"&gt;HIGH&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;   &lt;span class="m"&gt;0.19&lt;/span&gt;
  &lt;span class="na"&gt;LOW&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;    &lt;span class="m"&gt;0.07&lt;/span&gt;

&lt;span class="na"&gt;merge_blocking&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;no&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;  &lt;span class="m"&gt;0.81&lt;/span&gt;
  &lt;span class="na"&gt;yes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;0.19&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No text. No chain of thought. All outputs in a single forward pass. TypeSafe calls the architecture a &lt;strong&gt;parallel sampler&lt;/strong&gt; and the training method &lt;strong&gt;Reinforcement Learning for Calibrated Decisions (RLCD)&lt;/strong&gt; — probabilities optimized against outcomes instead of against human raters' preference.&lt;/p&gt;

&lt;p&gt;Line those two up and the "faster LLM" story collapses. The reasoning model pays for every token of its own deliberation:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    C[Context / prompt] --&amp;gt; N[Neural pass]
    N --&amp;gt; T1["First I need to check ..."]
    T1 --&amp;gt; T2[token]
    T2 --&amp;gt; T3[token]
    T3 --&amp;gt; T4["However ..."]
    T4 --&amp;gt; T5[token]
    T5 --&amp;gt; T6["... and therefore the answer is"]
    T6 --&amp;gt; D[Decision]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Jev pays once per question and hands back a distribution:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    C["Context / state&amp;lt;br/&amp;gt;diff summary, scan results, test results,&amp;lt;br/&amp;gt;API compatibility, findings, project policy"] --&amp;gt; N[Neural pass]
    N --&amp;gt; D1["severity: MEDIUM 0.74, HIGH 0.19, LOW 0.07"]
    N --&amp;gt; D2["merge_blocking: no 0.81, yes 0.19"]
    N --&amp;gt; D3["needs_human_review: no 0.9, yes 0.1"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Comparing those as if one were the optimized version of the other is like calling a calculator a faster author. Text in either way, and a transformer under the hood according to TypeSafe — but a fundamentally different compute problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does Jev not reason at all? Wrong question
&lt;/h2&gt;

&lt;p&gt;Jev still has to see relationships in the state and weigh them. What it does &lt;em&gt;not&lt;/em&gt; do is narrate that weighing as tokens. The difference is not "reasoning versus no reasoning", it is &lt;strong&gt;where the decomposition happens&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;With a reasoning LLM, the decomposition happens at inference time, in tokens, inside the model. You pay for it on every call, and you get a rationale you cannot trust.&lt;/li&gt;
&lt;li&gt;With Jev, the decomposition has to have already happened — in your questions, in your options, in the workflow that assembled the state. If you did not do that work, no parallel sampler will do it for you at 3 a.m.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the honest sentence is not "Jev needs less reasoning". It is: &lt;em&gt;Jev relocates the reasoning, and the new location is your repository.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody advertises the pipeline that feeds Jev
&lt;/h2&gt;

&lt;p&gt;Consider the question people actually want answered:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Is this change safe?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To answer that, someone has to walk this chain:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    Q["Is this change safe?"] --&amp;gt; A[Read the diff]
    A --&amp;gt; B[Find the affected classes]
    B --&amp;gt; C[Follow the call graph]
    C --&amp;gt; D[Look up auth configuration]
    D --&amp;gt; E[Determine library versions]
    E --&amp;gt; F[Check CVE data]
    F --&amp;gt; G[Inspect the tests]
    G --&amp;gt; H[Decide]
    H --&amp;gt; I["Jev: probabilities for this one decision"]
    style Q stroke-dasharray: 5 5&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Jev does not collapse that chain. Not in 100 ms, not in 100 s. What it collapses is the last hop: instead of asking a model to &lt;em&gt;find out what is going on&lt;/em&gt; and decide in one breath, you hand it a state that is already bite-sized and ask it to &lt;em&gt;decide&lt;/em&gt;. The difference is the whole product:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Not "figure out what the situation is." But: "here is the entire situation — decide."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a legitimate, valuable, genuinely hard-to-do-well thing. It is also a thing that presupposes the chain above it exists, runs, and is correct. Every millisecond Jev saves on decision latency is a millisecond some other system spent on context assembly. That system is written by you, tested by you, and paged at 3 a.m. by you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Jev integration actually looks like
&lt;/h2&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    A["Classical software&amp;lt;br/&amp;gt;collect context, read the database,&amp;lt;br/&amp;gt;call APIs, compute, apply rules"] --&amp;gt; B["Jev&amp;lt;br/&amp;gt;fuzzy decision, 100-500 ms"]
    B --&amp;gt; C["probabilities and confidence"]
    C --&amp;gt; D["Classical software&amp;lt;br/&amp;gt;if / route / retry / stop"]
    D -.-&amp;gt;|"next iteration"| A&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;And what it does not look like:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    U[User] --&amp;gt; J[Jev] --&amp;gt; X["Autonomously solves a complex problem"]
    X --&amp;gt; Y["This is not a feature request, it is a category error"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The confusion is understandable: the name "System One" — borrowed from Kahneman's fast, intuitive thinking — plus the framing of "frontier intelligence in a function call" suggests that Jev thinks like a reasoning model, only faster. It does not. It is an extremely capable, context-dependent &lt;strong&gt;classifier with calibrated confidence&lt;/strong&gt;. An agent it is not. Calling it one is how you end up with autonomy that works in the demo and&lt;br&gt;
flakes in production, with nothing to read in the logs but floats.&lt;/p&gt;
&lt;h2&gt;
  
  
  The reasoning moves into your program graph — and that is the good news
&lt;/h2&gt;

&lt;p&gt;Here is the pattern people are actually building, for the merge decision above:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    S["Prepared state:&amp;lt;br/&amp;gt;diff summary, scan results, test results,&amp;lt;br/&amp;gt;API compat, findings, project policy"] --&amp;gt; SR["securityRisk = Jev(...)"]
    S --&amp;gt; TR["testRisk = Jev(...)"]
    S --&amp;gt; AR["apiRisk = Jev(...)"]
    S --&amp;gt; QR["qualityRisk = Jev(...)"]
    SR --&amp;gt; D1{"securityRisk above 0.8?"}
    D1 --&amp;gt;|yes| B[BLOCK]
    D1 --&amp;gt;|no| D2{"apiRisk above 0.6?"}
    D2 --&amp;gt;|yes| H[HUMAN_REVIEW]
    D2 --&amp;gt;|no| D3{"testRisk below 0.2&amp;lt;br/&amp;gt;and qualityRisk below 0.3?"}
    D3 --&amp;gt;|yes| M[MERGE]
    D3 --&amp;gt;|no| H&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The knowledge that used to be implicit in a model's chain of thought — &lt;em&gt;security outranks style, so check security first; if security is critical, the other results do not matter anymore&lt;/em&gt; — is now explicit in code. Which is genuinely better: code is diffable, reviewable, unit-testable, and it does not cost a token to branch.&lt;/p&gt;

&lt;p&gt;But notice what had to happen for this to work. Somebody had to write down that security outranks style. Somebody had to define the four questions, the option sets, the thresholds 0.8 / 0.6 / 0.2 / 0.3, and somebody has to own the answer to "why 0.8?". With an LLM agent you never wrote any of that down and it mostly worked anyway. With Jev, the design is the&lt;br&gt;
product. That is not a flaw — it is the bill arriving for the architecture you skipped.&lt;/p&gt;
&lt;h2&gt;
  
  
  The only combination worth building into an agent
&lt;/h2&gt;

&lt;p&gt;The interesting hybrid is not "Jev instead of the reasoning model". It is a division of&lt;br&gt;
labour:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TB
    L["Reasoning LLM&amp;lt;br/&amp;gt;explore, investigate, generate"] --&amp;gt; ST["state / results"]
    ST --&amp;gt; G1["Jev: is this finding relevant?"]
    ST --&amp;gt; G2["Jev: is a retry worth it?"]
    ST --&amp;gt; G3["Jev: did the review gate pass?"]
    ST --&amp;gt; G4["Jev: another tool call needed?"]
    G1 --&amp;gt; W["workflow / code"]
    G2 --&amp;gt; W
    G3 --&amp;gt; W
    G4 --&amp;gt; W
    W -.-&amp;gt;|"next step"| L&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The expensive model does only what needs exploration or generation. The dozens of small judgements per step — relevant? retry? good enough? escalate? — go to a model that answers in milliseconds with a number attached. At 40x to 200x the speed and a small fraction of the cost per call, you can afford gates that were previously too expensive to even&lt;br&gt;
consider, and that is the real win: not replacing your model, but &lt;em&gt;adding judgement where you previously made do with a hard-coded rule or nothing at all&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Now the part nobody puts on the slide: gates multiply. A chain of ten gates at 95% accuracy lands at 60% end-to-end. Twenty land at 36%. To keep twenty gates above 98% you need roughly 99.9% per gate — which is exactly why "it is just a classifier with calibrated probabilities" is not a downgrade, it is the whole requirement. And it is why your test suite now needs calibration checks on &lt;em&gt;your&lt;/em&gt; distribution, not benchmark leaderboards.&lt;/p&gt;
&lt;h2&gt;
  
  
  The myth of "it cannot hallucinate"
&lt;/h2&gt;

&lt;p&gt;This one deserves a hard stop. TypeSafe's claim is technically careful: possible outputs are defined in advance, so the model never makes type errors, and schema matching is guaranteed. Fine. That is a statement about &lt;em&gt;form&lt;/em&gt;, not about &lt;em&gt;truth&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Jev can absolutely decide wrongly. It can label a CRITICAL finding as MEDIUM at 0.74, and your code, having been told that 0.74 is the answer, will route on it. What Jev removes is the visible failure mode, not the failure:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A["LLM: a wrong sentence&amp;lt;br/&amp;gt;in fluent text"] --&amp;gt; B["Someone notices,&amp;lt;br/&amp;gt;laughs, retries"]
    C["Jev: a wrong label&amp;lt;br/&amp;gt;with a confident probability"] --&amp;gt; D["Your code branches silently"]
    D --&amp;gt; E["Weeks later:&amp;lt;br/&amp;gt;a post-mortem with&amp;lt;br/&amp;gt;no rationale to read"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Hallucination has not been abolished here. It has been converted into &lt;strong&gt;silent miscalibration&lt;/strong&gt; — and for operations, that is the worse failure mode, because a wrong sentence is loud while a wrong float is invisible, reproducible, and trusted by CI. Three consequences follow immediately:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;You cannot ask why.&lt;/strong&gt; Simon Willison's complaint is the right one: an LLM at least produces a rationale you may not trust; a decision model produces a number. If Jev flags something, which content signals tipped it off? Nobody can tell you, and bias now hides inside a float.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calibration is a claim, and it is about their distribution.&lt;/strong&gt; That a model returns honest probabilities does not mean it stays honest on your inputs. TypeSafe's own jaggedness notes admit weak spots on numbers, dates and adversarial content. Perturb a prompt slightly, watch a decision flip, and try to explain it to an auditor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The blast radius got bigger.&lt;/strong&gt; An LLM hallucinating a tool call is inconvenient. The same model buried three layers deep in a dependency chain with a latency budget is a production incident.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So: not "Jev cannot hallucinate". Rather: "Jev's wrong answers no longer look like wrong answers". If you cannot detect them, you have not automated a decision — you have automated trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you are going to build this, build it honestly
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Keep state assembly in ordinary software. Jev decides, it does not look things up.&lt;/li&gt;
&lt;li&gt;Express every decision as a question with a typed, closed option set — if you cannot, you have not thought the decision through yet.&lt;/li&gt;
&lt;li&gt;Set thresholds explicitly and write down why they are what they are. A number nobody can justify is a liability, not a configuration.&lt;/li&gt;
&lt;li&gt;Measure calibration on &lt;em&gt;your&lt;/em&gt; inputs, not on the vendor's benchmark. Then break your inputs on purpose and watch which decisions flip.&lt;/li&gt;
&lt;li&gt;Read the gate chain end to end and multiply the accuracies. Compound error is where "fast and cheap" turns into "fast, cheap and quietly wrong".&lt;/li&gt;
&lt;li&gt;Never let a gate decide something you cannot explain to a reviewer, a customer or an auditor. "The model said 0.74" is not an explanation.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Verdict: a very good tool with a very bad pitch
&lt;/h2&gt;

&lt;p&gt;Jev is a serious, interesting approach, and for the right problem shape it is the best tool currently on the market: constant, documentable, cheap judgement inside ordinary software. Used well, it is what makes agentic systems affordable — cheap gates around an expensive explorer.&lt;/p&gt;

&lt;p&gt;But it is not a replacement for LLM reasoning. It is &lt;em&gt;less&lt;/em&gt; reasoning, applied precisely where reasoning was never what you needed — only a decision. Everything that made the decision hard (knowing what matters, assembling the state, choosing the questions, setting the thresholds, owning the calibration) stays on your side of the fence.&lt;/p&gt;

&lt;p&gt;Which is the actual provocation here. The hype train is selling you speed and a price tag.&lt;br&gt;
What Jev really ships is a demand: write down what you actually decide, in code, with numbers, and test it. Teams that skipped that work are not going to be saved by a parallel sampler; they are going to discover, in production, that they never knew what they were asking.&lt;/p&gt;

&lt;p&gt;Stop calling it a faster LLM. Start treating it as a receipt — for the architecture you have been postponing all along.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>jev</category>
      <category>llm</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Before Sovereign AI, You Need Sovereign IT</title>
      <dc:creator>Tom Seidel</dc:creator>
      <pubDate>Tue, 22 Sep 2026 06:39:57 +0000</pubDate>
      <link>https://dev.to/tmseidel/sovereign-ai-is-loud-but-the-real-issue-is-sovereign-it-3cjj</link>
      <guid>https://dev.to/tmseidel/sovereign-ai-is-loud-but-the-real-issue-is-sovereign-it-3cjj</guid>
      <description>&lt;h1&gt;
  
  
  Sovereign AI Is Loud — But The Real Issue Is Sovereign IT
&lt;/h1&gt;

&lt;p&gt;Most articles here are focused on very concrete, technical problems: setups, architectures, and reproducible solutions. This one takes a step back. Still, at its core, it remains technical — because questions of “sovereignty” ultimately materialize in infrastructure, operations, and the ability to make and reverse decisions.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Current Narrative: Sovereign AI
&lt;/h2&gt;

&lt;p&gt;The debate around &lt;strong&gt;sovereign AI&lt;/strong&gt; has gained remarkable momentum. Across industry reports, policy discussions, and conference stages, there is a growing emphasis on regaining control over data, models, and infrastructure. Especially in Europe, the dependency on non-local providers has become a recurring concern, driven by both geopolitical tension and regulatory realities.&lt;/p&gt;

&lt;p&gt;Conceptually, the idea is straightforward: build, run, and control AI systems within your own sphere of influence. In practice, however, this turns out to be significantly more demanding.&lt;/p&gt;

&lt;p&gt;Operating a self-hosted LLM setup today requires assembling a stack from relatively young and still evolving components:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;inference engines such as &lt;strong&gt;vLLM&lt;/strong&gt; or &lt;strong&gt;Ollama&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;containerized deployments and orchestration, often via &lt;strong&gt;Kubernetes&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;surrounding services like vector databases, authentication, observability, and networking&lt;/li&gt;
&lt;li&gt;and, perhaps most critically, access to and operation of suitable &lt;strong&gt;GPU infrastructure&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There is no widely accepted, production-ready “standard stack” comparable to what exists for more traditional enterprise workloads. Instead, organizations are faced with a growing but fragmented ecosystem. While this provides flexibility, it also means that building such a platform still requires a considerable amount of in-house expertise and operational maturity.&lt;/p&gt;




&lt;h2&gt;
  
  
  Infrastructure Is Scaling — But Not Necessarily Becoming Sovereign
&lt;/h2&gt;

&lt;p&gt;At the same time, the physical backbone of digital infrastructure is expanding rapidly. Driven largely by AI workloads, Europe is experiencing a strong increase in data center construction and capacity.&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Data Points
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Topic&lt;/th&gt;
&lt;th&gt;Data&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Demand driver&lt;/td&gt;
&lt;td&gt;AI is now the leading driver of data center demand in Europe &lt;a href="https://www.rlbinsights.com/reports/data-centre-trends-report-2025/ai-hits-europe" rel="noopener noreferrer"&gt;1&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capacity growth&lt;/td&gt;
&lt;td&gt;Avg. deployments: 16 MW → 33 MW → 47 MW (2023–2025) &lt;a href="https://www.rlbinsights.com/reports/data-centre-trends-report-2025/ai-hits-europe" rel="noopener noreferrer"&gt;1&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Investment scale&lt;/td&gt;
&lt;td&gt;&amp;gt; €100B expected in European data centers by 2030 &lt;a href="https://www.eudca.org/state-of-european-data-centres-2025" rel="noopener noreferrer"&gt;2&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pipeline size&lt;/td&gt;
&lt;td&gt;~14 GW planned capacity in EMEA by mid-2025 &lt;a href="https://www.allianz.com/content/dam/onemarketing/azcom/Allianz_com/economic-research/publications/specials/en/2025/october/2025-10-07-construction-AZ.pdf" rel="noopener noreferrer"&gt;3&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This expansion is often interpreted as a positive signal for digital sovereignty: more infrastructure, closer to users, under European jurisdiction.&lt;/p&gt;

&lt;p&gt;Yet the picture is more nuanced. Much of this newly built capacity is designed to support hyperscale cloud and AI platforms. Even when physically located in Europe, these systems often operate within technological, contractual, and operational frameworks defined elsewhere. In that sense, infrastructure alone does not automatically translate into control.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Overlooked Context: Sovereignty Decisions Were Made Earlier
&lt;/h2&gt;

&lt;p&gt;Seen in this light, the current intensity of the AI sovereignty debate is somewhat surprising. Over the past decade, many organizations have already made fundamental decisions that reduced their control over IT systems — often quite deliberately.&lt;/p&gt;

&lt;p&gt;The shift toward cloud-based and platform-centric architectures brought undeniable advantages: faster deployment, reduced operational overhead, and access to highly sophisticated tooling. At the same time, it gradually relocated critical capabilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;infrastructure to cloud providers&lt;/li&gt;
&lt;li&gt;development workflows to hosted platforms&lt;/li&gt;
&lt;li&gt;collaboration and communication to SaaS ecosystems&lt;/li&gt;
&lt;li&gt;customer-facing systems to externally operated services&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this was irrational. In many cases, it was the most pragmatic choice available. However, it also meant that certain aspects of control — over data flows, system behavior, or long-term portability — became harder to maintain.&lt;/p&gt;




&lt;h2&gt;
  
  
  Different Industries, Different Trade-offs
&lt;/h2&gt;

&lt;p&gt;In my own experience, the way organizations approach these questions varies significantly by industry.&lt;/p&gt;

&lt;p&gt;In &lt;strong&gt;medical technology&lt;/strong&gt;, there remains a noticeable degree of caution when it comes to outsourcing critical IT systems. This is not simply a cultural preference but largely shaped by regulation. Requirements around data protection, auditability, and traceability leave relatively little room for ambiguity. As a result, questions about where data resides, who can access it, and how systems can be audited or replaced tend to be addressed early and in detail.&lt;/p&gt;

&lt;p&gt;This does not necessarily mean that cloud usage is avoided altogether. Rather, it is approached more deliberately, and often with additional safeguards or hybrid models. In that sense, sovereignty is less an abstract concept and more an operational constraint that needs to be managed.&lt;/p&gt;

&lt;p&gt;By contrast, in parts of the &lt;strong&gt;automotive industry&lt;/strong&gt;, the shift toward external platforms was often more assertive. Internal infrastructure teams were frequently seen as cost-intensive and less flexible, while cloud providers offered scalability, speed, and a rich ecosystem of services. Over time, this led to a situation where a significant portion of the IT landscape — from development pipelines to collaboration tools — relies on a relatively small number of external vendors.&lt;/p&gt;

&lt;p&gt;This approach brought clear benefits in terms of velocity and standardization. At the same time, it introduced new forms of dependency, some of which only become visible when trying to change direction. Migration paths, data portability, and operational independence tend to become more complex as systems are more deeply integrated into proprietary platforms.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why AI Changes the Tone of the Debate
&lt;/h2&gt;

&lt;p&gt;Against this backdrop, it becomes clearer why AI has triggered a more pronounced reaction.&lt;/p&gt;

&lt;p&gt;AI systems tend to sit very close to core value creation. They process sensitive data, encode domain knowledge, and increasingly influence decision-making processes. As a result, questions of control feel more immediate. Where earlier layers of abstraction — infrastructure, collaboration tools, or workflow systems — could be externalized with relatively limited visibility, AI makes dependencies more tangible.&lt;/p&gt;

&lt;p&gt;At the same time, many of the underlying challenges are not new. Vendor lock-in, limited transparency, reliance on external roadmaps, or the gradual loss of internal expertise are all dynamics that have been present for years. AI does not introduce them so much as amplify their impact.&lt;/p&gt;




&lt;h2&gt;
  
  
  A More Nuanced View on Sovereignty
&lt;/h2&gt;

&lt;p&gt;It is also worth noting that sovereignty is rarely absolute. In practice, it tends to manifest as a spectrum rather than a binary choice.&lt;/p&gt;

&lt;h3&gt;
  
  
  Levels of Sovereignty
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Full reliance on SaaS / hyperscalers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;EU hosting with contractual safeguards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Open-source, portable architectures&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maximum&lt;/td&gt;
&lt;td&gt;Fully self-hosted or air‑gapped systems&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For many organizations, a balanced approach is likely the most practical: retaining control where it matters most while leveraging external services where appropriate. What becomes increasingly important, however, is the ability to make these choices consciously — and to revisit them if needed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Early Signs of a Shift
&lt;/h2&gt;

&lt;p&gt;There are indications that the broader conversation is evolving. In the public sector, for example, initiatives are emerging that explicitly frame digital sovereignty as a strategic objective.&lt;/p&gt;

&lt;p&gt;Programs that move away from proprietary ecosystems toward open-source-based infrastructures — such as the adoption of LibreOffice, Linux, and open collaboration platforms — reflect a growing awareness of long-term dependencies and their implications. &lt;a href="https://www.schleswig-holstein.de/DE/landesregierung/ministerien-behoerden/I/Presse/PI/2024/CdS/241125_cds_open-source-strategie" rel="noopener noreferrer"&gt;4&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These efforts are not without friction. Migration, training, and organizational change can be challenging. Yet they also demonstrate that alternative approaches are possible, even at scale.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The renewed focus on sovereign AI is both understandable and valuable. At the same time, it risks narrowing the perspective if it is treated in isolation.&lt;/p&gt;

&lt;p&gt;Questions of control, dependency, and adaptability do not begin with AI — they extend across the entire IT landscape. In that sense, sovereign AI is less a starting point and more a visible manifestation of a broader theme.&lt;/p&gt;

&lt;p&gt;A more comprehensive view would therefore consider not only AI systems, but the surrounding architecture, data flows, and operational capabilities. Ultimately, sovereignty is less about complete independence and more about maintaining the flexibility to respond, adapt, and make informed choices.&lt;/p&gt;

&lt;p&gt;Or, put differently:&lt;br&gt;&lt;br&gt;
it is not necessarily about building everything yourself —&lt;br&gt;&lt;br&gt;
but about retaining the ability not to be entirely dependent on others.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>cloud</category>
      <category>security</category>
      <category>ai</category>
    </item>
    <item>
      <title>The 2026 Developer Job Market Isn't Collapsing — It's Splitting in Two</title>
      <dc:creator>Tom Seidel</dc:creator>
      <pubDate>Fri, 11 Sep 2026 06:05:46 +0000</pubDate>
      <link>https://dev.to/tmseidel/regarding-the-job-market-why-2026-feels-like-a-crisis-and-why-the-profession-isnt-dying-254n</link>
      <guid>https://dev.to/tmseidel/regarding-the-job-market-why-2026-feels-like-a-crisis-and-why-the-profession-isnt-dying-254n</guid>
      <description>&lt;h1&gt;
  
  
  Regarding the Job Market: Why 2026 Feels Like a Crisis — and Why the Profession Isn't Dying
&lt;/h1&gt;

&lt;p&gt;When ChatGPT started making its way into development workflows — back then it was all copy-paste from the chat window — and non-developers suddenly started vibe-coding prototypes that looked genuinely impressive, the question immediately surfaced: do we even need developers anymore? From day one I was a convinced believer in AI as a tool, and I still am: I'm more confident than ever that developers are needed &lt;em&gt;more&lt;/em&gt; than before, not less. And the job-market statistics from the major job boards broadly support that view. We still have a massive deficit of applications that are supposed to drive the digitalization of everyday processes, and the demand for software remains unbroken.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers: strong overall demand, a transforming job market
&lt;/h2&gt;

&lt;p&gt;The current state of the market is best described as a split, not a collapse. AI engineering roles are growing much faster than classic software engineering roles, while generalist positions are growing only modestly or sitting flat.[1] At the same time, the tooling is already everywhere: 84% of developers now use or plan to use AI tools in their development process, and 51% of professional developers use them daily.[9] Roughly 40% of software engineering job postings already mention AI coding tools — Copilot, Cursor, Claude Code — as a requirement or a plus.[23]&lt;/p&gt;

&lt;h2&gt;
  
  
  The junior paradox: the right data, the wrong conclusion
&lt;/h2&gt;

&lt;p&gt;In certain bubbles, the warning is loud: getting in as a junior developer is the hardest it has ever been. The numbers back that up. A study by Stanford's Digital Economy Lab, built on ADP payroll data covering millions of U.S. workers, found that employment for software developers aged 22–25 declined by nearly 20% from its late-2022 peak.[5][7] U.S. entry-level tech job postings dropped by roughly 67% between 2023 and 2024 in that same Stanford/ADP analysis.[3]. The reason why I've chosen statistics from the U.S. is that they are the most detailed and reliable I could find.&lt;/p&gt;

&lt;p&gt;But that is the wrong way to look at it. Getting into the profession has not become harder in an absolute sense — the &lt;em&gt;profile&lt;/em&gt; of the entry-level role has simply changed. Tasks that junior developers typically handled are now done reliably by AI. What has changed is the job description, not the demand for fresh talent. What juniors need is to adapt their skill set to the new conditions: use AI as an accelerator in their own work, focus on system design and software architecture, build more business and domain knowledge — or pick crisis-resistant industries. That work methods and tools keep changing and improving is nothing new. We have been using tools and abstractions that massively boost productivity for decades. Business as usual.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why 2026 &lt;em&gt;feels&lt;/em&gt; like a crisis year
&lt;/h2&gt;

&lt;p&gt;If demand is still there, why does 2026 subjectively feel like a crisis year for software developers? My analysis of the market psychology behind that:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. The fear of "tech debt"
&lt;/h3&gt;

&lt;p&gt;Companies are afraid to spend millions today on software architectures or large teams for projects that new AI models could render obsolete — or solve in a radically different way — in two years.&lt;/p&gt;

&lt;p&gt;The worry: anyone who commits to a specific framework or expensive in-house development today may be building tomorrow's legacy baggage.&lt;/p&gt;

&lt;p&gt;The consequence: companies wait until standards for AI-native software architectures have settled. The market's own uncertainty feeds this: Gartner predicts that 40% of agentic AI projects will be scrapped by 2027,[22] and 2026 Gartner/IDC data puts the share of enterprise AI agent pilots that stall at 89%.[13]&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The team-scaling paradox
&lt;/h3&gt;

&lt;p&gt;The core uncertainty is a simple question: &lt;em&gt;how many developers will we need in three years for the same work?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The worry: If AI agents multiply a single developer's productivity several times over, an aggressive hiring wave today becomes massive overcapacity tomorrow.&lt;/p&gt;

&lt;p&gt;The consequence: Companies currently prefer to invest in higher compute capacity and AI licenses rather than in additional full-time employees. They scale performance through technology, not through headcount. The money is visibly moving in that direction: enterprises are burning through entire AI budgets in months,[11] while the hyperscalers alone are planning roughly $700 billion in AI capex for 2026.[12] Gartner even predicts that AI coding token costs will overtake the average developer's salary by 2028.[18]&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The legacy problem
&lt;/h3&gt;

&lt;p&gt;Many companies are standing in front of gigantic migration projects — replacing old COBOL or Java monoliths, for example. Here the "wait and see" attitude is extreme&lt;/p&gt;

&lt;p&gt;The worry: Why pay consultants millions when LLMs might soon handle such migrations fully automated and error-free at the push of a button (btw: This rests on a promise that track record does not yet support)?&lt;/p&gt;

&lt;p&gt;The consequence; Vendors are aggressively selling exactly that promise,[20][21] and as long as the pain isn't acute, those mega-projects get postponed.&lt;/p&gt;

&lt;p&gt;What makes the situation feel even more strained is that the "freeze" is only selective. Companies are waiting on juniors while simultaneously hunting, almost frantically, for tech leads and software architects who can bridge existing systems and generative AI — the field is bifurcating, with implementation roles shrinking while system-level, architecture, and AI-oversight roles expand.[25]&lt;/p&gt;

&lt;h2&gt;
  
  
  The doomsday hype in the filter bubbles
&lt;/h2&gt;

&lt;p&gt;In certain online bubbles — LinkedIn, select subreddits — there is a partly hysterical fear that has little to do with reality. I see four causes:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. The "survivor bias" of the frustrated
&lt;/h3&gt;

&lt;p&gt;Extremely frustrated people who wrote 200 applications and got only rejections, then pin the blame on AI — plus attention-hungry tech influencers chasing clicks and reach with dark prophecies. It is easy to overlook the thousands of developers who say nothing, write code, and use AI to write better code.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Confusing "coding" with "software engineering"
&lt;/h3&gt;

&lt;p&gt;Most demos show an AI implementing a specific function in the shortest possible time. The thinking error is that many people believe &lt;em&gt;typing code&lt;/em&gt; is a developer's main job. Reality looks very different: research puts the time developers actually spend writing code at around 11% of the working week — about 52 minutes a day,[26] and broader industry data on how developers spend their time points in the same direction.[17] The real work consists of tasks that LLMs cannot perform on their own.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Hype distortion
&lt;/h3&gt;

&lt;p&gt;On LinkedIn or Reddit it often feels like the state of the art is already successfully deployed in enterprises — "AI agents now replace whole teams." In reality, companies are still fighting the basics: data privacy, hallucinations, intellectual property, and how to build a unified, resilient AI infrastructure inside an enterprise at all. The numbers tell the same story: a March 2026 survey found 78% of enterprises have AI agent pilots, but fewer than 15% have scaled even one agent to operational use.[14] Gartner's own forecast for "vibe coding" — 40% of new enterprise software built with these techniques — is only dated 2028, not today.[19]&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Underestimating maintenance
&lt;/h3&gt;

&lt;p&gt;"If AI writes the code, we don't need people" is completely wrong — the opposite is true. The cheaper it becomes to generate code, the more code exists in our systems that has to be understood, maintained, updated, and secured — work that will require plenty of developers for decades. You can already see this spillover in open-source projects, which are being flooded with AI-generated contributions: in February 2026, GitHub itself acknowledged the surge in low-quality AI-generated contributions and described it as an "Eternal September" for open source.[15]&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;The fear of the profession &lt;em&gt;changing&lt;/em&gt; is justified. The fear of the profession &lt;em&gt;disappearing&lt;/em&gt; is a product of filter bubbles. The profession of software developer is not dying — it is changing at high speed. Anyone who develops from a "code typist" into a "system problem-solver" will have a secure career in the future as well.&lt;/p&gt;

&lt;p&gt;The wage data points in the same direction as the conclusion: the Stanford/ADP analysis found that workers with AI skills earn roughly 62% more than their peers.[6] If AI were simply replacing developers, it should depress their wages — and it is not. The data is consistent with a market in which AI raises the value of developers who can direct it, which is precisely the bet that the "adapt, don't panic" advice is based on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;a href="https://newsletter.pragmaticengineer.com/p/state-of-the-job-market-2026" rel="noopener noreferrer"&gt;https://newsletter.pragmaticengineer.com/p/state-of-the-job-market-2026&lt;/a&gt; — The Pragmatic Engineer: State of the job market 2026&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.danilchenko.dev/posts/junior-developer-jobs-2026" rel="noopener noreferrer"&gt;https://www.danilchenko.dev/posts/junior-developer-jobs-2026&lt;/a&gt; — Danilchenko: Junior developer jobs in 2026&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://byteiota.com/junior-developer-extinction-67-hiring-collapse-explained" rel="noopener noreferrer"&gt;https://byteiota.com/junior-developer-extinction-67-hiring-collapse-explained&lt;/a&gt; — ByteIota: Junior developer extinction, 67% hiring collapse&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://hakia.com/news/junior-developer-crisis-2026" rel="noopener noreferrer"&gt;https://hakia.com/news/junior-developer-crisis-2026&lt;/a&gt; — Hakia: Junior developer crisis 2026&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://thecollegeinvestor.com/81990/stanford-study-entry-level-software-jobs-down-nearly-20-as-ai-reshapes-hiring-for-college-grads" rel="noopener noreferrer"&gt;https://thecollegeinvestor.com/81990/stanford-study-entry-level-software-jobs-down-nearly-20-as-ai-reshapes-hiring-for-college-grads&lt;/a&gt; — The College Investor: Stanford study on entry-level software jobs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.digitalecononews.com/en/articles/stanford-economist-ai-entry-level-jobs-crisis" rel="noopener noreferrer"&gt;https://www.digitalecononews.com/en/articles/stanford-economist-ai-entry-level-jobs-crisis&lt;/a&gt; — Digital Econ News: Stanford economist Brynjolfsson on AI entry-level jobs crisis&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://msftnewsnow.com/stanford-study-generative-ai-entry-level-jobs" rel="noopener noreferrer"&gt;https://msftnewsnow.com/stanford-study-generative-ai-entry-level-jobs&lt;/a&gt; — Stanford study: Generative AI's early impact on entry-level jobs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://digitaleconomy.stanford.edu/project/indicators/canaries-dashboard" rel="noopener noreferrer"&gt;https://digitaleconomy.stanford.edu/project/indicators/canaries-dashboard&lt;/a&gt; — Stanford Digital Economy Lab: Canaries Dashboard&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://survey.stackoverflow.co/2025/ai" rel="noopener noreferrer"&gt;https://survey.stackoverflow.co/2025/ai&lt;/a&gt; — Stack Overflow 2025 Developer Survey: AI&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey" rel="noopener noreferrer"&gt;https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey&lt;/a&gt; — Stack Overflow 2025 Developer Survey press release&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.forbes.com/sites/timbajarin/2026-04-29/ai-compute-surpasses-human-costs-enterprise-budgets-shift" rel="noopener noreferrer"&gt;https://www.forbes.com/sites/timbajarin/2026-04-29/ai-compute-surpasses-human-costs-enterprise-budgets-shift&lt;/a&gt; — Forbes: AI compute surpasses human costs, enterprise budgets shift&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.cnbc.com/2026-02-06/google-microsoft-meta-amazon-ai-cash.html" rel="noopener noreferrer"&gt;https://www.cnbc.com/2026-02-06/google-microsoft-meta-amazon-ai-cash.html&lt;/a&gt; — CNBC: Tech AI spending approaches 700B in 2026&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.beri.net/article/ai-agent-adoption-enterprise-2026-gartner-idc" rel="noopener noreferrer"&gt;https://www.beri.net/article/ai-agent-adoption-enterprise-2026-gartner-idc&lt;/a&gt; — Gartner/IDC 2026: 89% of enterprise AI agent pilots stall&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentmarketcap.ai/blog/2026-04-11/ai-agent-reality-check-hype-to-production-gap-2026" rel="noopener noreferrer"&gt;https://agentmarketcap.ai/blog/2026-04-11/ai-agent-reality-check-hype-to-production-gap-2026&lt;/a&gt; — AI Agent Reality Check: 78% pilots, under 15% to production&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://en.wikipedia.org/wiki/Vibe_coding" rel="noopener noreferrer"&gt;https://en.wikipedia.org/wiki/Vibe_coding&lt;/a&gt; — Wikipedia: Vibe coding&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.sonarsource.com/blog/how-much-time-do-developers-spend-actually-writing-code" rel="noopener noreferrer"&gt;https://www.sonarsource.com/blog/how-much-time-do-developers-spend-actually-writing-code&lt;/a&gt; — Sonar: How much time do developers spend actually writing code?&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.software.com/reports/code-time-report" rel="noopener noreferrer"&gt;https://www.software.com/reports/code-time-report&lt;/a&gt; — Antenna: Global Code Time Report&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.techtimes.com/articles/319333/20260629/ai-coding-costs-can-drain-budget-days-gartner-predicts-they-will-match-developer-pay.htm" rel="noopener noreferrer"&gt;https://www.techtimes.com/articles/319333/20260629/ai-coding-costs-can-drain-budget-days-gartner-predicts-they-will-match-developer-pay.htm&lt;/a&gt; — TechTimes: Gartner on AI coding costs vs developer pay&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://theoutpost.ai/news-story/vibe-coding-transforms-software-development-as-gartner-predicts-90-enterprise-adoption-by-2028-30292" rel="noopener noreferrer"&gt;https://theoutpost.ai/news-story/vibe-coding-transforms-software-development-as-gartner-predicts-90-enterprise-adoption-by-2028-30292&lt;/a&gt; — The Outpost: Gartner vibe coding 40% by 2028&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://medium.com/@hashbyt/what-happens-when-ai-meets-legacy-cobol-systems-00dae11f51d3" rel="noopener noreferrer"&gt;https://medium.com/@hashbyt/what-happens-when-ai-meets-legacy-cobol-systems-00dae11f51d3&lt;/a&gt; — Medium: AI COBOL modernization in 2026&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://ascendion.com/ai-basic/mainframe-and-cobol-modernization-with-agentic-ai-whats-actually-possible-in-2026" rel="noopener noreferrer"&gt;https://ascendion.com/ai-basic/mainframe-and-cobol-modernization-with-agentic-ai-whats-actually-possible-in-2026&lt;/a&gt; — Ascendion: Mainframe modernization with agentic AI in 2026&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.kore.ai/blog/ai-agents-in-2026-from-hype-to-enterprise-reality" rel="noopener noreferrer"&gt;https://www.kore.ai/blog/ai-agents-in-2026-from-hype-to-enterprise-reality&lt;/a&gt; — kore.ai: Gartner predicts 40% of agentic AI projects scrapped by 2027&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://gitgood.dev/blog/2026-tech-job-market-hiring-rebound-ai-roles" rel="noopener noreferrer"&gt;https://gitgood.dev/blog/2026-tech-job-market-hiring-rebound-ai-roles&lt;/a&gt; — GitGood: 40% of SE job postings mention AI coding tools&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://ucstrategies.com/news/ibm-lost-40b-because-ai-cant-actually-modernize-cobol" rel="noopener noreferrer"&gt;https://ucstrategies.com/news/ibm-lost-40b-because-ai-cant-actually-modernize-cobol&lt;/a&gt; — UC Strategies: AI COBOL modernization evidence gap&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://theboard.world/articles/technology/will-ai-replace-software-engineers-2026-reality-check" rel="noopener noreferrer"&gt;https://theboard.world/articles/technology/will-ai-replace-software-engineers-2026-reality-check&lt;/a&gt; — TheBoard: Will AI replace software engineers in 2026?&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://medium.com/@vikpoca/developers-spend-only-11-of-their-time-coding-what-3a53f65982df" rel="noopener noreferrer"&gt;https://medium.com/@vikpoca/developers-spend-only-11-of-their-time-coding-what-3a53f65982df&lt;/a&gt; — Ponamarev: Developers spend only 11% of their time coding&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>disc</category>
      <category>softwareengineering</category>
      <category>career</category>
    </item>
    <item>
      <title>Running Self-Hosted AI Code Reviews with Ollama on a Small VPS</title>
      <dc:creator>Tom Seidel</dc:creator>
      <pubDate>Thu, 10 Sep 2026 13:25:25 +0000</pubDate>
      <link>https://dev.to/tmseidel/run-ai-code-reviews-for-the-cost-of-a-5-vps-no-per-seat-saas-required-54jp</link>
      <guid>https://dev.to/tmseidel/run-ai-code-reviews-for-the-cost-of-a-5-vps-no-per-seat-saas-required-54jp</guid>
      <description>&lt;p&gt;AI-assisted code review doesn't necessarily require sending every pull request to a hosted AI service.&lt;/p&gt;

&lt;p&gt;For some teams, running the review workflow inside their own infrastructure can be an interesting alternative—particularly when they already operate Gitea, GitLab, GitHub Enterprise, or other self-hosted development infrastructure.&lt;/p&gt;

&lt;p&gt;I wanted to see how small such a setup could reasonably be.&lt;/p&gt;

&lt;p&gt;The result is a fairly simple architecture:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    git["Git platform"] --&amp;gt;|Webhook| workflow["AI workflow"]
    workflow --&amp;gt;|Repository context + prompt| ollama["Ollama"]
    ollama --&amp;gt; model["Local coder model"]
    model --&amp;gt;|Review comments| git&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;There is no per-developer component in this architecture. The main constraints are instead the compute required by the model, the size of the pull requests being reviewed, and how quickly you expect reviews to complete.&lt;/p&gt;

&lt;p&gt;For my experiments I use AI-Git-Bot as the workflow layer, Ollama as the local inference server, and an open-weight coder model.&lt;/p&gt;

&lt;p&gt;The interesting part isn't really the particular bot, though. It's what becomes possible once the Git workflow and the model runtime are separated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why run code review locally?
&lt;/h2&gt;

&lt;p&gt;Hosted AI development tools are convenient, and for many teams they're probably the simplest option.&lt;/p&gt;

&lt;p&gt;But they also come with trade-offs.&lt;/p&gt;

&lt;p&gt;Pricing commonly scales with the number of developers, while the underlying workload doesn't necessarily do so in the same way. A ten-person team might generate fewer pull requests than a three-person team working on a very active repository.&lt;/p&gt;

&lt;p&gt;There is also the question of where source code is processed.&lt;/p&gt;

&lt;p&gt;Depending on the organization, sending diffs to an external model provider can require additional security, contractual, or compliance review. In some environments it isn't allowed at all.&lt;/p&gt;

&lt;p&gt;A self-hosted setup changes those trade-offs:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    repo["Repository"] --&amp;gt;|Webhook| workflow["Workflow engine"]
    workflow --&amp;gt;|Prompt + context| model["Local model"]
    model --&amp;gt;|Review result| workflow
    workflow --&amp;gt;|Publish findings| repo&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The source code, prompts, model inference, and generated review stay within infrastructure you control.&lt;/p&gt;

&lt;p&gt;The downside is equally important: &lt;strong&gt;you now operate the infrastructure yourself.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You have to think about memory, model performance, updates, monitoring, and what happens when the model produces a poor review.&lt;/p&gt;

&lt;p&gt;So this isn't automatically a better architecture. It's just a different one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the workflow layer does
&lt;/h2&gt;

&lt;p&gt;A local LLM alone isn't enough to provide automated code reviews.&lt;/p&gt;

&lt;p&gt;Something still needs to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;react to pull-request events,&lt;/li&gt;
&lt;li&gt;retrieve the diff and repository context,&lt;/li&gt;
&lt;li&gt;construct the prompt,&lt;/li&gt;
&lt;li&gt;call the model,&lt;/li&gt;
&lt;li&gt;interpret the response,&lt;/li&gt;
&lt;li&gt;and publish findings back to the Git platform.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the role AI-Git-Bot plays in my setup.&lt;/p&gt;

&lt;p&gt;It supports GitHub/GitHub Enterprise, Gitea, GitLab, and Bitbucket Cloud and can react to normal Git events.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workflow&lt;/th&gt;
&lt;th&gt;Trigger&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PR review&lt;/td&gt;
&lt;td&gt;PR opened / review requested&lt;/td&gt;
&lt;td&gt;Summary and findings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interactive Q&amp;amp;A&lt;/td&gt;
&lt;td&gt;Bot mentioned in a PR&lt;/td&gt;
&lt;td&gt;Context-aware answer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unit-test generation&lt;/td&gt;
&lt;td&gt;PR opened&lt;/td&gt;
&lt;td&gt;Regression tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E2E / full-stack QA&lt;/td&gt;
&lt;td&gt;PR opened&lt;/td&gt;
&lt;td&gt;Test execution and results&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Issue triage&lt;/td&gt;
&lt;td&gt;Issue opened&lt;/td&gt;
&lt;td&gt;Classification or assignment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Issue implementation&lt;/td&gt;
&lt;td&gt;Issue assigned to coding agent&lt;/td&gt;
&lt;td&gt;Implementation pull request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Documentation sync&lt;/td&gt;
&lt;td&gt;PR opened&lt;/td&gt;
&lt;td&gt;Documentation changes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;i18n coverage&lt;/td&gt;
&lt;td&gt;PR opened&lt;/td&gt;
&lt;td&gt;Missing translations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Not every workflow makes sense with a small local model.&lt;/p&gt;

&lt;p&gt;That distinction turned out to be quite important.&lt;/p&gt;

&lt;h2&gt;
  
  
  A minimal deployment
&lt;/h2&gt;

&lt;p&gt;For a self-hosted experiment, the basic stack consists of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI-Git-Bot&lt;/li&gt;
&lt;li&gt;PostgreSQL&lt;/li&gt;
&lt;li&gt;Ollama&lt;/li&gt;
&lt;li&gt;an open-weight coder model&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A simplified Docker Compose setup looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;tmseidel/ai-git-bot:latest&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8080:8080"&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;SPRING_PROFILES_ACTIVE&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker&lt;/span&gt;
      &lt;span class="na"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;jdbc:postgresql://db:5432/giteabot&lt;/span&gt;
      &lt;span class="na"&gt;DATABASE_USERNAME&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;giteabot&lt;/span&gt;
      &lt;span class="na"&gt;DATABASE_PASSWORD&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;giteabot&lt;/span&gt;
      &lt;span class="na"&gt;APP_ENCRYPTION_KEY&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;change-me&lt;/span&gt;
    &lt;span class="na"&gt;depends_on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;db&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;condition&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service_healthy&lt;/span&gt;
    &lt;span class="na"&gt;restart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;unless-stopped&lt;/span&gt;

  &lt;span class="na"&gt;db&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres:17-alpine&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;POSTGRES_DB&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;giteabot&lt;/span&gt;
      &lt;span class="na"&gt;POSTGRES_USER&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;giteabot&lt;/span&gt;
      &lt;span class="na"&gt;POSTGRES_PASSWORD&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;giteabot&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;pgdata:/var/lib/postgresql/data&lt;/span&gt;
    &lt;span class="na"&gt;healthcheck&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CMD-SHELL"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pg_isready&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;-U&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;giteabot"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5s&lt;/span&gt;
      &lt;span class="na"&gt;timeout&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5s&lt;/span&gt;
      &lt;span class="na"&gt;retries&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;
    &lt;span class="na"&gt;restart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;unless-stopped&lt;/span&gt;

  &lt;span class="na"&gt;ollama&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ollama/ollama:latest&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;11434:11434"&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;ollama_data:/root/.ollama&lt;/span&gt;
    &lt;span class="na"&gt;restart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;unless-stopped&lt;/span&gt;

  &lt;span class="na"&gt;ollama-pull&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ollama/ollama:latest&lt;/span&gt;
    &lt;span class="na"&gt;entrypoint&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;sh&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;-c&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;sleep 5 &amp;amp;&amp;amp; ollama pull qwen2.5-coder:7b&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;OLLAMA_HOST&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://ollama:11434&lt;/span&gt;
    &lt;span class="na"&gt;depends_on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;ollama&lt;/span&gt;

&lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;ollama_data&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pgdata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once Ollama is running, the AI integration points at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://ollama:11434
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;with, for example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;qwen2.5-coder:7b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There are no model API credentials involved because inference happens locally.&lt;/p&gt;

&lt;h2&gt;
  
  
  How small can the model be?
&lt;/h2&gt;

&lt;p&gt;This was the more interesting question.&lt;/p&gt;

&lt;p&gt;A 7B coder model is obviously not going to compete with the largest hosted reasoning models on every task.&lt;/p&gt;

&lt;p&gt;But not every Git workflow needs the same capabilities.&lt;/p&gt;

&lt;p&gt;I've found a useful distinction between workflows that mostly generate natural language and workflows that need reliable agentic behavior.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;7B class&lt;/th&gt;
&lt;th&gt;14–32B class&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;PR review&lt;/td&gt;
&lt;td&gt;Works reasonably well&lt;/td&gt;
&lt;td&gt;Better&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Explain a change&lt;/td&gt;
&lt;td&gt;Works reasonably well&lt;/td&gt;
&lt;td&gt;Better&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Documentation updates&lt;/td&gt;
&lt;td&gt;Often sufficient&lt;/td&gt;
&lt;td&gt;Better&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Issue classification&lt;/td&gt;
&lt;td&gt;Depends on required output structure&lt;/td&gt;
&lt;td&gt;Better&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-step coding agent&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Much more suitable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strict structured output&lt;/td&gt;
&lt;td&gt;Can be unreliable&lt;/td&gt;
&lt;td&gt;More reliable&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a pull-request review, a smaller model can still inspect a diff and produce useful natural-language observations.&lt;/p&gt;

&lt;p&gt;Agent workflows are harder.&lt;/p&gt;

&lt;p&gt;Consider an issue-to-pull-request workflow:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    issue["Issue"] --&amp;gt; understand["Understand task"]
    understand --&amp;gt; inspect["Inspect repository"]
    inspect --&amp;gt; choose["Choose files"]
    choose --&amp;gt; edit["Modify code"]
    edit --&amp;gt; test["Run tests"]
    test --&amp;gt; result{"Tests pass?"}
    result --&amp;gt;|Yes| pr["Create pull request"]
    result --&amp;gt;|No| diagnose["Interpret failures"]
    diagnose --&amp;gt; edit&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Each additional decision increases the importance of reasoning quality, tool use, structured output, and instruction following.&lt;/p&gt;

&lt;p&gt;That's where the difference between a small local model and a larger model becomes much more visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  A hybrid setup may be more practical
&lt;/h2&gt;

&lt;p&gt;Self-hosting doesn't have to mean that every workflow uses the same local model.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    event["Git event"] --&amp;gt; workflow["Workflow"]
    workflow --&amp;gt;|Review task| local["Local 7B model"]
    local --&amp;gt; review["PR review"]
    workflow --&amp;gt;|Agentic task| large["Larger model"]
    large --&amp;gt; agent["Coding agent"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The inexpensive, frequent tasks can run locally.&lt;/p&gt;

&lt;p&gt;The less frequent workflows that need stronger reasoning can use a larger local model or an external provider.&lt;/p&gt;

&lt;p&gt;I find this architecture more interesting than treating the model provider as a global application setting.&lt;/p&gt;

&lt;p&gt;Different workflows have different requirements.&lt;/p&gt;

&lt;p&gt;A PR summary, security review, issue classifier, and autonomous coding agent don't necessarily benefit from the same model.&lt;/p&gt;

&lt;h2&gt;
  
  
  What about the "$5 VPS"?
&lt;/h2&gt;

&lt;p&gt;This needs a qualification.&lt;/p&gt;

&lt;p&gt;The application itself is lightweight. &lt;strong&gt;Model inference is not.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A quantized 7B model is only a few gigabytes on disk, but the machine also needs memory for the model runtime, context, operating system, database, and application.&lt;/p&gt;

&lt;p&gt;So whether an actual $5/month VPS is sufficient depends heavily on the provider, available RAM, CPU performance, context size, and acceptable latency.&lt;/p&gt;

&lt;p&gt;If you already have an 8 GB machine, a small home server, a workstation, or other unused infrastructure, the incremental software cost can be close to zero.&lt;/p&gt;

&lt;p&gt;If you're renting infrastructure specifically for inference, you should size and benchmark it rather than assuming the cheapest VPS will provide acceptable performance.&lt;/p&gt;

&lt;p&gt;For me, the useful observation isn't really "$5."&lt;/p&gt;

&lt;p&gt;It's this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cost of a self-hosted review service is primarily tied to the inference infrastructure and workload rather than directly to the number of developers using it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's a different scaling model from per-seat SaaS.&lt;/p&gt;

&lt;p&gt;Whether it's cheaper depends on your team and infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  CPU inference is possible, but latency matters
&lt;/h2&gt;

&lt;p&gt;A GPU isn't strictly required for experimenting with a 7B quantized model.&lt;/p&gt;

&lt;p&gt;CPU inference works.&lt;/p&gt;

&lt;p&gt;But "works" and "feels fast" are different things.&lt;/p&gt;

&lt;p&gt;For asynchronous workflows such as pull-request review, latency is often less critical than it would be for an interactive coding assistant.&lt;/p&gt;

&lt;p&gt;If a review takes a few minutes after a pull request is opened, that may still be perfectly usable.&lt;/p&gt;

&lt;p&gt;For interactive Q&amp;amp;A or agentic coding loops, the same latency becomes much more noticeable because every model call is part of a sequence.&lt;/p&gt;

&lt;p&gt;This is another reason I wouldn't use one deployment strategy for every AI-assisted development workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I think this architecture makes sense
&lt;/h2&gt;

&lt;p&gt;There are a few situations where I've found self-hosting particularly interesting.&lt;/p&gt;

&lt;h3&gt;
  
  
  Self-hosted Git platforms
&lt;/h3&gt;

&lt;p&gt;A lot of AI developer tooling starts with GitHub support.&lt;/p&gt;

&lt;p&gt;If your organization primarily runs Gitea or another internally hosted Git platform, connecting a generic workflow layer to a local model can be more flexible than trying to fit the environment into a GitHub-centric SaaS product.&lt;/p&gt;

&lt;h3&gt;
  
  
  Source code shouldn't leave the network
&lt;/h3&gt;

&lt;p&gt;Some teams have contractual, regulatory, or internal-security constraints around source code.&lt;/p&gt;

&lt;p&gt;Keeping the complete path&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    git["Git platform"] --&amp;gt;|Event + repository data| workflow["Workflow engine"]
    workflow --&amp;gt;|Prompt + context| model["Local model"]
    model --&amp;gt;|Generated result| workflow
    workflow --&amp;gt;|Comment or change| git&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;inside the same environment makes that boundary easier to reason about.&lt;/p&gt;

&lt;p&gt;It doesn't automatically make the system secure—you still have to secure the infrastructure—but the data flow is much simpler.&lt;/p&gt;

&lt;h3&gt;
  
  
  Existing compute is available
&lt;/h3&gt;

&lt;p&gt;If an organization already operates machines with sufficient RAM or GPU capacity, running another small inference workload may be inexpensive.&lt;/p&gt;

&lt;p&gt;The economics are quite different if hardware has to be purchased specifically for the workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this setup doesn't replace
&lt;/h2&gt;

&lt;p&gt;A self-hosted review workflow isn't a replacement for every AI developer tool.&lt;/p&gt;

&lt;p&gt;It isn't code completion.&lt;/p&gt;

&lt;p&gt;An IDE assistant helps while code is being written. A Git workflow operates after an event such as opening a pull request.&lt;/p&gt;

&lt;p&gt;Those are different places in the development lifecycle.&lt;/p&gt;

&lt;p&gt;It also doesn't remove the need for human review.&lt;/p&gt;

&lt;p&gt;I treat model-generated findings much like static-analysis findings: potentially useful signals that still need context.&lt;/p&gt;

&lt;p&gt;And a smaller local model makes that distinction especially important.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment I find interesting
&lt;/h2&gt;

&lt;p&gt;What started as an attempt to run inexpensive AI code reviews has turned into a broader architectural question for me:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much AI infrastructure actually needs the largest available model?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some development workflows clearly benefit from powerful reasoning models.&lt;/p&gt;

&lt;p&gt;Others look much more like repeated, bounded classification or transformation tasks.&lt;/p&gt;

&lt;p&gt;Once those workloads are separated, the architecture becomes more flexible:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    git["Git"] --&amp;gt; workflow["Workflow"]
    workflow --&amp;gt; route{"Choose model for task"}
    route --&amp;gt; local["Small local model"]
    route --&amp;gt; large["Large local model"]
    route --&amp;gt; hosted["Hosted reasoning model"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The workflow chooses the appropriate model for the task rather than treating "the AI" as one monolithic dependency.&lt;/p&gt;

&lt;p&gt;That's also where I'm currently experimenting with AI-Git-Bot.&lt;/p&gt;

&lt;p&gt;The project is MIT licensed, and the repository contains deployment and Ollama examples if you want to reproduce the setup.&lt;/p&gt;

&lt;p&gt;If you've tried running code review on small local models, I'd be especially interested in hearing which model and hardware combination you used—and where you found the quality threshold stopped being acceptable.&lt;/p&gt;

</description>
      <category>codereview</category>
      <category>selfhosting</category>
      <category>opensource</category>
      <category>ai</category>
    </item>
    <item>
      <title>AI-Git-Bot — an open-source gateway that connects your Git platforms with AI providers</title>
      <dc:creator>Tom Seidel</dc:creator>
      <pubDate>Sun, 12 Apr 2026 10:35:23 +0000</pubDate>
      <link>https://dev.to/tmseidel/ai-git-bot-an-open-source-gateway-that-connects-your-git-platforms-with-ai-providers-1ic6</link>
      <guid>https://dev.to/tmseidel/ai-git-bot-an-open-source-gateway-that-connects-your-git-platforms-with-ai-providers-1ic6</guid>
      <description>&lt;p&gt;🚀 Excited to share AI-Git-Bot — an open-source gateway that connects your Git platforms with AI providers for automated code reviews and autonomous issue implementation.&lt;/p&gt;

&lt;p&gt;After weeks of development, I'm proud to release v1.2.0 of AI-Git-Bot — a lightweight, self-hostable application that sits between your Git hosting and AI providers, acting as an intelligent gateway.&lt;/p&gt;

&lt;p&gt;🤖🧠 Half Bot, Half Agent&lt;/p&gt;

&lt;p&gt;As a Bot: It automatically reviews pull requests, answers questions in comments, and delivers context-aware inline feedback — like a code-review partner that never sleeps.&lt;/p&gt;

&lt;p&gt;As an Agent: Assign it to an issue and it autonomously reads the codebase, generates an implementation, validates the code with build tools, and creates a finished pull request — all on its own.&lt;/p&gt;

&lt;p&gt;🌉 The Gateway Principle&lt;/p&gt;

&lt;p&gt;What makes AI-Git-Bot unique is its gateway architecture:&lt;br&gt;
✅ 5 Git platforms — Gitea, GitHub, GitHub Enterprise, GitLab, Bitbucket Cloud&lt;br&gt;
✅ 4 AI providers — Anthropic Claude, OpenAI, Ollama, llama.cpp&lt;br&gt;
✅ Mix &amp;amp; match — Any Git platform with any AI provider&lt;br&gt;
✅ Multi-bot — Different personas (security reviewer, performance expert, mentor) on the same PR&lt;br&gt;
✅ One dashboard — Manage everything through a web UI&lt;/p&gt;

&lt;p&gt;🔒 Built for Self-Hosters&lt;/p&gt;

&lt;p&gt;Run everything on-premise with local LLMs. No code ever leaves your infrastructure — perfect for teams with regulatory or compliance requirements.&lt;/p&gt;

&lt;p&gt;Tech Stack&lt;/p&gt;

&lt;p&gt;☕ Java 21 / Spring Boot 4 &lt;br&gt;
🐘 PostgreSQL &lt;br&gt;
🐳 Single Docker image &lt;/p&gt;

&lt;p&gt;Get Started&lt;/p&gt;

&lt;p&gt;&lt;code&gt;docker compose up -d&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;That's it. Navigate to localhost:8080, create an admin account, configure your integrations, and you're reviewing code with AI.&lt;/p&gt;

&lt;p&gt;📦 GitHub: &lt;a href="https://github.com/tmseidel/ai-git-bot" rel="noopener noreferrer"&gt;https://github.com/tmseidel/ai-git-bot&lt;/a&gt;&lt;br&gt;
🐳 Docker Hub: &lt;a href="https://hub.docker.com/r/tmseidel/ai-git-bot" rel="noopener noreferrer"&gt;https://hub.docker.com/r/tmseidel/ai-git-bot&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;📄 License: MIT&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>java</category>
      <category>claude</category>
    </item>
  </channel>
</rss>
