<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: saaro</title>
    <description>The latest articles on DEV Community by saaro (@saaro_net).</description>
    <link>https://dev.to/saaro_net</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1579561%2F830ef749-97ba-47e2-b016-44502ba8f198.png</url>
      <title>DEV Community: saaro</title>
      <link>https://dev.to/saaro_net</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/saaro_net"/>
    <language>en</language>
    <item>
      <title>Agentic Coding in Production 2026: From Autocomplete to Autonomous Developer</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Sat, 03 Oct 2026 12:00:36 +0000</pubDate>
      <link>https://dev.to/saaro_net/agentic-coding-in-production-2026-from-autocomplete-to-autonomous-developer-1oi8</link>
      <guid>https://dev.to/saaro_net/agentic-coding-in-production-2026-from-autocomplete-to-autonomous-developer-1oi8</guid>
      <description>&lt;p&gt;The leap from simple chat prompts to agentic workflows is the biggest upheaval in software development since the introduction of Git. While 2024 and 2025 were still dominated by autocomplete plugins and "vibe coding" in the headlines, the landscape has fundamentally changed by 2026: AI agents plan independently, edit multiple files, run tests, and react to errors—all in a closed loop. But with this freedom come new challenges in terms of cost, control, and code quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three categories of agentic development
&lt;/h2&gt;

&lt;p&gt;The market for AI coding tools has split into three clear categories in 2026. &lt;strong&gt;GitHub Copilot&lt;/strong&gt; integrates as a plugin into existing editors and accompanies the developer with intelligent autocomplete and increasingly agentic capabilities. &lt;strong&gt;Cursor&lt;/strong&gt; is a standalone AI IDE that completely replaces the editor and intervenes deeply in the development process. &lt;strong&gt;Claude Code&lt;/strong&gt; by Anthropic pursues a third path: as a terminal-first agent, it works environment-independently and can collaborate with any IDE, any CI/CD pipeline, and any workflow.&lt;/p&gt;

&lt;p&gt;The LogRocket AI Dev Tool Power Ranking from September 2026 places Claude Code in first place after the model upgrade, followed by Cursor and GitHub Copilot. Search trends from mid-2026 confirm this development: Claude Code records roughly six times the search intensity of GitHub Copilot at the peak of its popularity. At the same time, Amazon Q Developer stopped accepting new registrations in May 2026—the market is consolidating around the three main players.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ReAct loop: Plan, Act, Observe, Adapt
&lt;/h2&gt;

&lt;p&gt;Behind every agentic coding tool lies the same core: the &lt;strong&gt;ReAct cycle&lt;/strong&gt; (Reasoning + Acting). The agent analyzes the task, plans the next steps, executes them via tools, observes the result, and adjusts its approach. In practice, this means: a coding agent reads an issue, independently opens the relevant files, writes a change, runs the tests, detects errors, corrects them, and creates a pull request.&lt;/p&gt;

&lt;p&gt;Anthropic's guide "Building Effective Agents" draws an important boundary here: between &lt;em&gt;workflows&lt;/em&gt;, where the program guides the model through predefined steps, and true &lt;em&gt;agents&lt;/em&gt;, where the model controls its own process. Most production systems in 2026 deliberately lie in between—structured processes with a few agentic decision points beat fully free loops in daily work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Spec-First and Plan Mode as quality guarantors
&lt;/h2&gt;

&lt;p&gt;Switching to agentic workflows requires more than a new tool. Three principles have proven themselves in practice:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spec-First Development:&lt;/strong&gt; Before an agent writes even a single line of code, the specification is defined. This approach drastically reduces hallucinations and creeping scope creep. Amazon Web Services introduced its own standard for this approach with "Spec-Driven Coding" at the Summit in June 2026.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Plan Mode before execution:&lt;/strong&gt; Both Claude Code and Cursor offer a Plan Mode—the agent creates a written plan that the developer reviews and approves before files are modified. Experience reports show that this single workflow change prevents 90% of typical errors before they occur.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Targeted context control:&lt;/strong&gt; Each loop iteration resends context and tool output. An agent that re-reads the entire repository at every step consumes many times more tokens. Targeted context windows, regular resets, and CLAUDE.md files with project knowledge are therefore essential for economical operation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agentic Technical Debt – the unmentioned downside
&lt;/h2&gt;

&lt;p&gt;Agentic workflows are not free. Each loop iteration costs tokens. An agent that needs five iterations quickly costs five to ten times as much as a single prompt. Added to this is what the community calls "&lt;strong&gt;Agentic Technical Debt&lt;/strong&gt;": a pipeline of prompts, tools, and semi-structured loops is a system you now own and must maintain.&lt;/p&gt;

&lt;p&gt;Not every task needs an agent. For deterministic, fully known processes, fixed scripts or cron jobs are still superior: faster, cheaper, and louder in case of failure. The clever use lies in knowing when an agent adds value and when it only produces expensive movement without real progress.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Agentic coding is no longer hype in 2026, but production reality. The tools are mature, the basic workflows are understood. The decisive factor for development teams is no longer the question of the right tool, but the right workflow: When do I delegate to an agent? Where do I set human checkpoints? And how do I avoid the freedom of the agentic workflow ending in uncontrolled costs?&lt;/p&gt;

&lt;p&gt;Teams that answer these questions report 40–60% faster feature delivery. Teams that ignore them mainly collect expensive token bills and growing technical debt. The key lies in the conscious combination of automation, human control, and clear specifications—and thus in an attitude that does not make the tool the boss, but uses it deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://blog.logrocket.com/ai-dev-tool-power-rankings/" rel="noopener noreferrer"&gt;AI Dev Tool Power Rankings (September 2026) – LogRocket Blog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://blog.buildbetter.ai/ai-development-workflow-buyers-guide-2026/" rel="noopener noreferrer"&gt;AI-Driven Development Workflow: The Complete 2026 Buyer's Guide – BuildBetter&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vibecoding.app/blog/what-is-an-agentic-workflow" rel="noopener noreferrer"&gt;What Is an Agentic Workflow? Developer Guide 2026 – Vibe Coding&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/research/building-effective-agents" rel="noopener noreferrer"&gt;Building Effective Agents – Anthropic Research&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://tech-insider.org/github-copilot-vs-cursor-vs-claude-code-2026/" rel="noopener noreferrer"&gt;GitHub Copilot vs Cursor vs Claude Code 2026 Compared – Tech Insider&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.techpillow.co/blog/amazon-kiro-aws-summit-nyc-spec-driven-2026" rel="noopener noreferrer"&gt;AWS Goes All-In on Spec-Driven Coding – TechPillow (AWS Summit NYC 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://blink.new/blog/tag/agentic-coding" rel="noopener noreferrer"&gt;Agentic Coding Best Practices – Blink Blog&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>coding</category>
      <category>claudecode</category>
      <category>cursor</category>
    </item>
    <item>
      <title>RAG Frameworks 2026: A Comparison of LlamaIndex, LangChain/LangGraph, Haystack, and More</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Fri, 02 Oct 2026 12:00:22 +0000</pubDate>
      <link>https://dev.to/saaro_net/rag-frameworks-2026-a-comparison-of-llamaindex-langchainlanggraph-haystack-and-more-4m3h</link>
      <guid>https://dev.to/saaro_net/rag-frameworks-2026-a-comparison-of-llamaindex-langchainlanggraph-haystack-and-more-4m3h</guid>
      <description>&lt;p&gt;RAG (Retrieval-Augmented Generation) has evolved from a niche technique to the standard for LLM-based applications by 2026. But with growing importance, the number of frameworks promising to build RAG pipelines has also increased. LlamaIndex, LangChain with LangGraph 1.0, Haystack, DSPy, and RAGFlow – each framework follows its own approach with specific strengths. The following article provides decision support for development leads and AI engineers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Framework or No Framework?
&lt;/h2&gt;

&lt;p&gt;The most fundamental question comes first: Does a team even need a RAG framework? For simple question-answering applications with one corpus, one LLM provider, and a standard chunking strategy, a provider SDK plus vector database client is often entirely sufficient – in under 50 lines of code without framework abstraction. The provider SDKs from OpenAI and Anthropic have absorbed much of what previously justified frameworks in 2026: Native tool usage, streaming tool calls, and prompt caching are now first-class. Most teams overestimate the orchestration complexity they will face and underestimate the cost of a framework they do not need &lt;sup id="fnref1"&gt;1&lt;/sup&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  LlamaIndex: The Specialist for Document RAG
&lt;/h2&gt;

&lt;p&gt;LlamaIndex has evolved beyond pure RAG in 2026: With &lt;strong&gt;LlamaIndex Workflows&lt;/strong&gt;, it offers an event-driven architecture for complex AI applications including state management and asynchronous processing. Over 160 supported file formats and &lt;strong&gt;LlamaParse&lt;/strong&gt; for professional document parsing cover demanding enterprise documents. &lt;strong&gt;LlamaCloud&lt;/strong&gt; complements the open-source framework with managed infrastructure offering 10,000 free months per month. Where retrieval quality is critical – for example in multi-document research, hierarchical indices, or knowledge graphs – LlamaIndex leads the competition &lt;sup id="fnref2"&gt;2&lt;/sup&gt;&lt;sup id="fnref3"&gt;3&lt;/sup&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  LangChain/LangGraph: Agentic Pipelines in Production
&lt;/h2&gt;

&lt;p&gt;LangChain remains the ecosystem with the broadest integration range. &lt;strong&gt;LangGraph 1.0&lt;/strong&gt; – stable since the end of 2025 – provides a stateful graph runtime for multi-step, agentic retrieval workflows. With over 143,000 GitHub stars (as of July 2026) and a huge community of tutorials, integrations, and Stack Overflow answers, LangChain is often the default choice. The agent abstractions have matured in 2026: Tool-calling patterns, memory management, and executors handle complex multi-step operations with production error handling. Disadvantage: The import overhead is noticeable, and the abstraction layer can hinder debugging – some teams have removed LangChain again because modular building blocks simplified their codebase &lt;sup id="fnref1"&gt;1&lt;/sup&gt;&lt;sup id="fnref2"&gt;2&lt;/sup&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Haystack, DSPy, and RAGFlow: Specialized Alternatives
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Haystack&lt;/strong&gt; (deepset) focuses on explicit, testable pipelines – every component and connection is visible in the code. With an Apache-2.0 license and EU headquarters, it is particularly interesting for data-sensitive teams. deepset Studio offers 100 pipeline hours free as a managed version. &lt;strong&gt;DSPy&lt;/strong&gt; is not a RAG framework in the classic sense, but a programmatic optimizer that compiles prompts and retrieval modules from data. For teams that constantly tune retrieval prompts "by hand," DSPy is the ideal complement – but not as a standalone framework. &lt;strong&gt;RAGFlow&lt;/strong&gt; delivers, as an Apache-2.0 project, a full-text RAG engine including UI, DeepDoc parser, and agent templates that can be self-hosted via Docker – ideal for teams looking for a ready-to-use platform &lt;sup id="fnref3"&gt;3&lt;/sup&gt;&lt;sup id="fnref1"&gt;1&lt;/sup&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The choice of the right RAG framework in 2026 depends on the specific bottleneck: If document parsing is the priority, LlamaIndex leads. For agentic multi-step pipelines, LangChain/LangGraph is the right choice. For explicit pipeline control, Haystack is the solution; for prompt optimization, DSPy. Those seeking a turnkey solution should go with RAGFlow. For simple applications, "no framework" is often the best answer – the provider SDKs of 2026 have massively lowered the entry barrier. A growing trend is also hybrid architectures: LlamaIndex for document processing, LangGraph for orchestration. Framework convergence is making rigid boundaries increasingly obsolete. In the end, what matters is which framework gets a team to a production-ready solution fastest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;techsy.io (July 2026): "Bestes RAG-Framework 2026: LangChain vs. LlamaIndex vs. Haystack" – &lt;a href="https://techsy.io/de/blog/bestes-rag-framework-2026" rel="noopener noreferrer"&gt;https://techsy.io/de/blog/bestes-rag-framework-2026&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;Zen van Riel (September 2026): "LangChain vs LlamaIndex in 2026: What's Changed and Which to Choose" – &lt;a href="https://zenvanriel.com/ai-engineer-blog/langchain-vs-llamaindex-2026-update/" rel="noopener noreferrer"&gt;https://zenvanriel.com/ai-engineer-blog/langchain-vs-llamaindex-2026-update/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;Startupik (September 2026): "Best RAG Frameworks 2026: LlamaIndex vs LangChain vs Haystack" – &lt;a href="https://startupik.com/best-rag-frameworks-tools-2026/" rel="noopener noreferrer"&gt;https://startupik.com/best-rag-frameworks-tools-2026/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>rag</category>
      <category>ai</category>
      <category>llm</category>
      <category>llamaindex</category>
    </item>
    <item>
      <title>Kubernetes Gateway API 2026: v1.5, v1.6, and Saying Goodbye to Ingress-NGINX</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Thu, 01 Oct 2026 12:00:44 +0000</pubDate>
      <link>https://dev.to/saaro_net/kubernetes-gateway-api-2026-v15-v16-and-saying-goodbye-to-ingress-nginx-39hc</link>
      <guid>https://dev.to/saaro_net/kubernetes-gateway-api-2026-v15-v16-and-saying-goodbye-to-ingress-nginx-39hc</guid>
      <description>&lt;p&gt;The year 2026 marks a turning point for the Kubernetes network: With the retirement of Ingress-NGINX in March and two major Gateway API releases (v1.5 and v1.6), the replacement of classic Ingress is imminent. This article shows what the new versions bring, why the timing is favorable, and how migration succeeds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gateway API v1.5: From Experimental to Standard
&lt;/h2&gt;

&lt;p&gt;On February 27, 2026, SIG Network released &lt;strong&gt;Gateway API v1.5&lt;/strong&gt; – the most extensive release to date. For the first time, the project switched to a &lt;em&gt;release-train model&lt;/em&gt; modeled on Kubernetes itself: By the feature-freeze deadline, all completed features are included in the release, without waiting for individual milestones.&lt;/p&gt;

&lt;p&gt;Six important features reached the &lt;strong&gt;Standard Channel&lt;/strong&gt; (GA status):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ListenerSet&lt;/strong&gt; – Previously, all listeners had to be defined directly in the Gateway object. ListenerSet allows listeners to be defined independently and merged onto a Gateway. This is particularly relevant for multi-tenant scenarios where platform and application teams manage different listeners. More than 64 listeners per Gateway also become possible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TLSRoute&lt;/strong&gt; – Enables route-based TLS forwarding without HTTP awareness, e.g., for database connections or message queues.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HTTPRoute CORS Filter&lt;/strong&gt; – native CORS handling directly in the route, without implementation-specific annotations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Client Certificate Validation&lt;/strong&gt; and &lt;strong&gt;Certificate Selection&lt;/strong&gt; – advanced TLS features for Gateway termination.&lt;/li&gt;
&lt;li&gt;**Reference&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>kubernetes</category>
      <category>gatewayapi</category>
      <category>ingress</category>
      <category>ingressnginx</category>
    </item>
    <item>
      <title>Kubernetes 1.37 Garhwal: HPA Scale-to-Zero, Gang Scheduling, and 18 Removed Kubelet Flags</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Thu, 01 Oct 2026 12:00:08 +0000</pubDate>
      <link>https://dev.to/saaro_net/kubernetes-137-garhwal-hpa-scale-to-zero-gang-scheduling-and-18-removed-kubelet-flags-10d0</link>
      <guid>https://dev.to/saaro_net/kubernetes-137-garhwal-hpa-scale-to-zero-gang-scheduling-and-18-removed-kubelet-flags-10d0</guid>
      <description>&lt;p&gt;Since August 26, 2026, Kubernetes 1.37 "Garhwal" is available – with 67 enhancements, 16 of which have reached Stable status. The release is named after the Garhwal region in the Indian Himalayas and focuses primarily on consolidation: an API that has been in beta for nine years finally becomes stable; the HorizontalPodAutoscaler can scale workloads down to zero replicas; and Gang Scheduling for AI/ML jobs reaches Beta status. Anyone operating a cluster should prepare for some breaking changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  HPA Scale-to-Zero: Workloads sleep when not needed
&lt;/h2&gt;

&lt;p&gt;Perhaps the most important innovation for operations is the scaled HorizontalPodAutoscaler (HPA). As of Kubernetes 1.37, an HPA can scale workloads down to &lt;strong&gt;zero replicas&lt;/strong&gt; – Beta status, enabled by default. The prerequisite is the use of Object or External metrics, since CPU and memory metrics depend on running Pods. The trick: a queue length exists independently of the workers processing it. The HPA can still read this value even with zero Pods and start new replicas when needed.&lt;/p&gt;

&lt;p&gt;The use case is clear: queue consumers, batch jobs, and GPU workloads that are idle between bursts simply consume no resources anymore. A &lt;code&gt;ScaledToZero&lt;/code&gt; status on the HPA object reliably distinguishes between "scaled down by the controller" and "manually set to zero".&lt;/p&gt;

&lt;h2&gt;
  
  
  Gang Scheduling: Plan AI/ML Jobs as a Unit
&lt;/h2&gt;

&lt;p&gt;Distributed training jobs often fail because some Pods are placed while the rest are stuck in a queue – the job does not progress but blocks resources. Gang Scheduling (KEP-4671) solves this problem with the &lt;code&gt;PodGroup&lt;/code&gt; concept: the scheduler receives the information that eight Pods must be started together and waits until enough capacity is available for all.&lt;/p&gt;

&lt;p&gt;Also newly added is workload-aware Preemption (KEP-5710), which prevents competing workloads from endlessly preempting each other. For platform teams running Ray, JobSet, or LWS, this is the next step toward a production-ready AI platform on Kubernetes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Metrics API becomes stable after nine years
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;metrics.k8s.io&lt;/code&gt; API has been in beta since Kubernetes 1.8 – that's nearly nine years. With 1.37, it is stable as &lt;code&gt;v1&lt;/code&gt;. The API surface does not change; it is a pure version graduation, but it sends a clear signal: Kubernetes' resource metric infrastructure is production-ready. &lt;code&gt;kubectl top&lt;/code&gt; already prefers the &lt;code&gt;v1&lt;/code&gt; version with fallback to &lt;code&gt;v1beta1&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Breaking Changes: 18 Kubelet flags removed
&lt;/h2&gt;

&lt;p&gt;Upgrading to 1.37 requires preparation. The migration of the embedded cAdvisor to a slim &lt;code&gt;cadvisor/lib&lt;/code&gt; module (PR #139870) removes &lt;strong&gt;18 Kubelet flags&lt;/strong&gt; – including &lt;code&gt;--containerd&lt;/code&gt;, &lt;code&gt;--containerd-namespace&lt;/code&gt;, &lt;code&gt;--boot-id-file&lt;/code&gt;, and &lt;code&gt;--enable-load-reader&lt;/code&gt;. A Kubelet that passes any of these flags will not start at all with &lt;code&gt;unknown flag&lt;/code&gt;. Nodes configured via &lt;code&gt;kubeadm-flags.env&lt;/code&gt; or systemd drop-in must be cleaned up before the upgrade.&lt;/p&gt;

&lt;p&gt;Also removed: the old IPVS support in kube-proxy finally gives way to nftables, and &lt;code&gt;failCgroupV1&lt;/code&gt; remains enabled by default. Container runtimes must use containerd 2.x – containerd 1.x is no longer supported.&lt;/p&gt;

&lt;h2&gt;
  
  
  DRA and Special Hardware mature
&lt;/h2&gt;

&lt;p&gt;Dynamic Resource Allocation (DRA) receives four Stable graduations in one release – this shows that GPUs, accelerators, and special network adapters are now first-class citizens in Kubernetes' resource model. The standardized DRA attribute &lt;code&gt;resource.kubernetes.io/numaNode&lt;/code&gt; allows consistent NUMA topology-dependent placement. Also new in Alpha status is the &lt;code&gt;ulimits&lt;/code&gt; configuration per container via the &lt;code&gt;Container.SecurityContext&lt;/code&gt; field – a long-awaited feature for databases and highly concurrent workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security: Pod certificates and SELinux on mount
&lt;/h2&gt;

&lt;p&gt;Pod certificates (&lt;code&gt;PodCertificateRequest&lt;/code&gt;) are stable with 1.37. Instead of mounting long-lived Secrets, the Kubelet can request short-lived X.509 certificates and rotate them regularly – a native workload identity system. Also stable: &lt;code&gt;SELinuxMount&lt;/code&gt;. Instead of recursively relabeling the entire volume before each Pod start, Kubernetes uses the Linux kernel-native &lt;code&gt;-o context&lt;/code&gt; flag and assigns the correct SEL&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>devops</category>
      <category>ai</category>
      <category>cloudnative</category>
    </item>
    <item>
      <title>RAG vs. Long Context 2026: What the Data Really Says</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:01:18 +0000</pubDate>
      <link>https://dev.to/saaro_net/rag-vs-long-context-2026-what-the-data-really-says-45ja</link>
      <guid>https://dev.to/saaro_net/rag-vs-long-context-2026-what-the-data-really-says-45ja</guid>
      <description>&lt;p&gt;In 2024, some voices predicted the imminent death of RAG: if models like Gemini 1.5 come with a million token context window, why bother running retrieval infrastructure? Two years later, the opposite has happened. Usage of RAG frameworks has increased by &lt;strong&gt;400 percent&lt;/strong&gt; between 2024 and 2026, and around 60 percent of all productive LLM applications continue to rely on Retrieval-Augmented Generation. At the same time, manufacturers like Meta with Llama 4 Scout are pushing ten million tokens into a window. Both are true at the same time – and that is not a contradiction, but the beginning of a differentiated architectural decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hard reality of long contexts
&lt;/h2&gt;

&lt;p&gt;The central insight of 2026 is sobering: &lt;em&gt;having&lt;/em&gt; a large context window and being able to &lt;em&gt;reliably use&lt;/em&gt; it are two fundamentally different things. The classic &lt;em&gt;Needle-in-a-Haystack&lt;/em&gt; test (NIAH) – a single piece of information hidden somewhere in a long document – looks impressive with a 99.7 percent hit rate for Gemini 1.5 Pro. But NIAH does not measure what goes wrong in practice.&lt;/p&gt;

&lt;p&gt;More realistic benchmarks like &lt;strong&gt;NoLiMa&lt;/strong&gt; (Multi-Fact-Recall without word overlap) or &lt;strong&gt;NeedleChain&lt;/strong&gt; (chained reasoning across multiple facts) paint a different picture: &lt;strong&gt;The multi-fact hit rate is around 60 percent&lt;/strong&gt; – 40 percent of facts disappear in the "middle part" of the context window. The &lt;em&gt;Lost-in-the-Middle&lt;/em&gt; effect (U-shaped attention curve) discovered in 2023 was confirmed again in 2026 by the RULER study on 17 long-context models. Facts at the beginning or end of the prompt are reliably found; facts in the middle drop by 20 percentage points or more.&lt;/p&gt;

&lt;p&gt;As one researcher put it: &lt;em&gt;"Larger context windows haven't solved the problem. They just give you more middle where things can get lost."&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost comparison speaks clearly
&lt;/h2&gt;

&lt;p&gt;A RAG pipeline costs about &lt;strong&gt;0.00008 US dollars&lt;/strong&gt; per query. A long-context approach with 100,000 tokens costs 0.10 to 0.25 US dollars, and with one million tokens 1.74 to 6.75 US dollars – per query. That is a cost difference of a factor of &lt;strong&gt;1,250x&lt;/strong&gt;. Anyone running 100,000 queries per day pays under 10 US dollars with RAG, but 10,000 US dollars or more with long context.&lt;/p&gt;

&lt;p&gt;Latency follows the same pattern. While a RAG pipeline responds in under two seconds, a 1M-token request takes 30 to 60 seconds. The reason is mathematical: Transformer attention scales quadratically with context length (O(n²)). A 1M-token cache requires about 100 GB of GPU memory per session. That is not an implementation weakness – the math works against large windows.&lt;/p&gt;

&lt;p&gt;Prompt caching changes the calculation for &lt;strong&gt;stable, repeatedly read&lt;/strong&gt; documents. For a 200,000-token corpus that is cached, costs become relative. For ad-hoc queries against fresh documents – the most common use case – caching does not help.&lt;/p&gt;

&lt;h2&gt;
  
  
  When long context really wins
&lt;/h2&gt;

&lt;p&gt;Despite all limitations: Long context is the right tool when three conditions apply simultaneously:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The knowledge base is small&lt;/strong&gt; (under 200,000 tokens)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The content is stable&lt;/strong&gt; and does not change constantly&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The task requires cross-corpus synthesis&lt;/strong&gt; – connecting information across the entire document&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Analyzing an annual report or a legal agreement are classic examples. When an agent needs to find inconsistencies between clauses spread over 50,000 tokens, RAG is structurally at a disadvantage because the chunks do not capture the relationships.&lt;/p&gt;

&lt;h2&gt;
  
  
  When RAG remains the right choice
&lt;/h2&gt;

&lt;p&gt;RAG remains the standard for large, dynamic knowledge bases in production. A customer support system with 50,000 documents and 100,000 queries per day cannot economically use long context – the corpus is too large, the content changes constantly, costs explode.&lt;/p&gt;

&lt;p&gt;But RAG is not a self-runner. Research identifies four typical failure modes: extraction errors (the model reads chunks incorrectly), context overflow (multiple retrievals exceed the token limit), premature termination (the model stops after a plausible but incomplete answer), and synthesis errors (correctly retrieved parts are not properly combined). The search part usually works – what happens afterward determines quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2026 hybrid: No either-or
&lt;/h2&gt;

&lt;p&gt;The smartest implementations of 2026 do not choose between RAG and long context, but &lt;strong&gt;route each query to the right tool&lt;/strong&gt;. A decision framework of five factors has become established:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Corpus size&lt;/strong&gt;: Under 100K tokens – both possible. 100K to 1M – hybrid. Over 1M – RAG first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relevance ratio&lt;/strong&gt;: If the proportion of relevant data per query is below 20 percent, RAG scores on average 13+ F1 points better.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Query volume&lt;/strong&gt;: From 10,000 queries per month, RAG becomes cost-effective; long context remains the exception for high-value individual cases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency SLO&lt;/strong&gt;: Under 2 seconds only works with RAG.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data freshness&lt;/strong&gt;: Constantly changing sources force RAG.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The year 2026 has effectively ended the "RAG vs. Long Context" debate. There is no winner – there is only the right architecture for each use case. RAG is cheaper, faster, and more reliable for most retrieval workloads. Long context delivers better results for tasks that require holistic document understanding. The future belongs to &lt;strong&gt;hybrid systems with intelligent routing&lt;/strong&gt; that combine both approaches where it makes sense, and clearly prioritize in all other cases. Anyone building an LLM-based application today should not ask "RAG or Long Context?" but rather: "When do I need which?"&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;a href="https://niteagent.com/blog/llm-context-2026-rag-vs-long-context/" rel="noopener noreferrer"&gt;LLM Context in 2026: Long Context vs RAG Decision Guide – NiteAgent&lt;/a&gt; (May 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://usewire.io/blog/long-context-vs-rag-what-the-data-shows/" rel="noopener noreferrer"&gt;RAG vs long context: what the 2026 data shows – Wire Blog&lt;/a&gt; (June 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://tianpan.co/blog/2026-04-09-long-context-vs-rag-production-decision-framework" rel="noopener noreferrer"&gt;Long-Context Models vs. RAG: When the 1M-Token Window Is the Wrong Tool – TianPan.co&lt;/a&gt; (April 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://wavect.io/blog/rag-vs-finetune-vs-longcontext-2026/" rel="noopener noreferrer"&gt;RAG vs Fine-Tuning vs Long Context: A 2026 Decision Method – Wavect&lt;/a&gt; (May 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://nat.io/blog/llms-in-2026-from-bigger-to-better" rel="noopener noreferrer"&gt;LLMs in 2026: RAG, Multimodality, Agents, and Hybrid AI Deployment – nat.io&lt;/a&gt; (February 2026)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;Lost in the Middle: How Language Models Use Long Contexts – Google Research (arXiv)&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>rag</category>
      <category>longcontext</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Argo CD 3.5: Internal mTLS, Source Integrity, and the New ApplicationSet UI</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Tue, 29 Sep 2026 12:00:42 +0000</pubDate>
      <link>https://dev.to/saaro_net/argo-cd-35-internal-mtls-source-integrity-and-the-new-applicationset-ui-5dce</link>
      <guid>https://dev.to/saaro_net/argo-cd-35-internal-mtls-source-integrity-and-the-new-applicationset-ui-5dce</guid>
      <description>&lt;p&gt;Argo CD 3.5 is here: Internal mTLS, Source Integrity, and the new ApplicationSet UI&lt;/p&gt;

&lt;p&gt;Security requirements for Kubernetes deployment pipelines are rising rapidly. Argo CD, the leading GitOps tool with nearly 60 percent adoption in Kubernetes clusters, has released version 3.5 (GA in August 2026), which closes two long-standing security gaps and significantly improves usability. In September, the release candidate for v3.6 followed – development is accelerating. This article provides an overview of the most important innovations and shows why the upgrade is particularly worthwhile for security-conscious teams.&lt;/p&gt;

&lt;h2&gt;
  
  
  Internal mTLS: The biggest security gain
&lt;/h2&gt;

&lt;p&gt;Probably the most significant innovation in Argo CD 3.5 is mutual TLS authentication (mTLS) between internal components. Until now, communication between the &lt;code&gt;repo-server&lt;/code&gt; and other services such as the API server, Application Controller, and ApplicationSet Controller was unencrypted and unauthenticated. An attacker with network access to the &lt;code&gt;repo-server&lt;/code&gt; could connect without authentication.&lt;/p&gt;

&lt;p&gt;A security vulnerability published in July 2026 exploited exactly this gap: an unauthenticated attacker could misuse a Kustomize option via the &lt;code&gt;repo-server&lt;/code&gt; to execute arbitrary commands, read the Redis password from an environment variable, and manipulate cached deployment data – without a CVE and without credentials. The bug had been reported to the maintainers 18 months before disclosure.&lt;/p&gt;

&lt;p&gt;With mTLS in v3.5, the &lt;code&gt;repo-server&lt;/code&gt; now requires a valid client certificate from every connected component. If you don't provide your own certificates, you automatically receive self-signed certificates generated in memory by the &lt;code&gt;repo-server&lt;/code&gt; – full filesystem access is not required. The feature is additive: existing clusters without their own certificates automatically receive the mTLS fallback solution, without requiring reconfiguration.&lt;/p&gt;

&lt;p&gt;Important note: mTLS does not replace NetworkPolicies, but supplements them as a second defense layer. If you activate the NetworkPolicies included in the Helm chart, you already protect the &lt;code&gt;repo-server&lt;/code&gt; at the network level. mTLS additionally covers the case where network isolation is incomplete or bypassed by lateral movement from another pod in the same namespace.&lt;/p&gt;

&lt;h2&gt;
  
  
  Source Integrity: Signed Git commits as a deployment gate
&lt;/h2&gt;

&lt;p&gt;The second major security feature is native Git commit signature verification, called &lt;em&gt;Source Integrity&lt;/em&gt;. Previously, it was sufficient that a commit existed on the configured branch – Argo CD deployed it, regardless of who pushed it or whether it was subsequently tampered with.&lt;/p&gt;

&lt;p&gt;With Source Integrity, the operator can specify that only signed commits may be synchronized. Specifically: in the &lt;code&gt;Application&lt;/code&gt; spec, set &lt;code&gt;sourceIntegrity.required: true&lt;/code&gt;, or via the CLI &lt;code&gt;argocd app set --source-integrity-required&lt;/code&gt;. A commit without a valid signature will then be excluded from synchronization.&lt;/p&gt;

&lt;p&gt;This addresses a real scenario: a leaked GitHub personal access token, a compromised developer laptop, or a force push during an incident – in all these cases, the commit on the branch looks "legitimate". The only question that Source Integrity answers and that stands before a deployment decision is: &lt;em&gt;Was this specific commit signed by a key authorized to approve production changes?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The older GPG verification is now deprecated and will be removed in the next major version.&lt;/p&gt;

&lt;h2&gt;
  
  
  ApplicationSet UI: Visibility for GitOps scaling
&lt;/h2&gt;

&lt;p&gt;ApplicationSets have long been a central feature for multi-cluster deployments. They allow templating the same application across multiple clusters or environments. Until now, however, they lacked native representation in the web interface – administrators had to read YAML or use kubectl to see which concrete applications an ApplicationSet template would generate.&lt;/p&gt;

&lt;p&gt;The new ApplicationSet UI, developed by engineers at Intuit, Red Hat, GoTo, and Octopus Deploy, now offers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A list view of all ApplicationSets&lt;/li&gt;
&lt;li&gt;Filter and detail views&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;Preview Apps tab&lt;/strong&gt; that shows which applications an ApplicationSet template will generate before deployment&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The preview tab in particular is a practical gain: it makes ApplicationSets more accessible for developer teams and reduces the risk of misconfiguration when scaling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Other innovations at a glance
&lt;/h2&gt;

&lt;p&gt;In addition to the three main features, Argo CD 3.5 brings a range of other improvements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Impersonation (Beta)&lt;/strong&gt;: Argo CD can assume a specific user identity for server-side tasks – relevant for audit trails in multi-tenant clusters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source Hydrator (Beta)&lt;/strong&gt;: The separation of dry (unhydrated) and hydrated manifests reaches beta status. Different repositories for source templates and rendered manifests are supported.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Helm 4 support&lt;/strong&gt;: Alongside backward compatibility with Helm 3.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Any-Namespace ApplicationSets&lt;/strong&gt;: ApplicationSets can be deployed in any namespace, not just the Argo CD namespace.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concurrency controls&lt;/strong&gt;: Limiting the number of simultaneously processed applications – protects against cluster overload.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Azure AD improvements&lt;/strong&gt;: Group claims overflow is resolved via the Microsoft Graph API, and Azure DevOps receives service principal authentication.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Argo CD 3.5 is more than a routine release. With internal mTLS and Source Integrity, the project closes two security gaps that were no longer acceptable given this level of adoption. The new ApplicationSet UI makes multi-cluster GitOps accessible to broader teams. As of September 2026, the release candidate for v3.6 is already available – the development speed shows that Argo CD is not standing still despite its market dominance.&lt;/p&gt;

&lt;p&gt;For teams running Argo CD in production, upgrading to v3.5 is strongly recommended, especially from a security perspective. The mTLS feature can be used without manual certificate management thanks to automatic self-signed certificates, and Source Integrity can be enabled per application – both with low migration effort.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.infoq.com/news/2026/06/argocd-supply-chain-security/" rel="noopener noreferrer"&gt;Argo CD 3.5 Tightens Supply Chain Security with Internal mTLS and Source Integrity – InfoQ (June 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://bex.co/blog/2026/07/10/argocd-35-mtls-commit-signing" rel="noopener noreferrer"&gt;Argo CD 3.5 Adds Internal mTLS and Commit Signing – bex.co (July 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://medium.com/argo-project/argo-cd-v3-5-release-candidate-02b1fbf7b419" rel="noopener noreferrer"&gt;Argo CD v3.5 Release Candidate – Argo Project Blog (June 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/argoproj/argo-cd/releases" rel="noopener noreferrer"&gt;Argo CD GitHub Releases – v3.6.0-rc1 (September 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.portainer.io/blog/argocd-vs-flux" rel="noopener noreferrer"&gt;ArgoCD vs Flux: The Complete 2026 GitOps Comparison Guide – Portainer (July 2026)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>argocd</category>
      <category>gitops</category>
      <category>kubernetes</category>
      <category>devops</category>
    </item>
    <item>
      <title>AI Workloads on Kubernetes 2026: The Production Stack (Kueue, KServe, vLLM, KubeRay)</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Mon, 28 Sep 2026 11:49:23 +0000</pubDate>
      <link>https://dev.to/saaro_net/ai-workloads-on-kubernetes-2026-the-production-stack-kueue-kserve-vllm-kuberay-3i47</link>
      <guid>https://dev.to/saaro_net/ai-workloads-on-kubernetes-2026-the-production-stack-kueue-kserve-vllm-kuberay-3i47</guid>
      <description>&lt;p&gt;Two years ago, the question of whether Kubernetes is the right place for AI workloads was still open. By 2026, it has been answered: the production stack for AI/ML on Kubernetes has consolidated, the CNCF has graduated central projects, and the tooling landscape has matured. GPU clusters under Kubernetes are now part of the standard repertoire of platform engineering teams—whether for LLM inference, distributed training, or multi-tenant environments with multiple teams.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2026 Reference Architecture
&lt;/h2&gt;

&lt;p&gt;A typical production stack for AI on Kubernetes consists of several layers that interlock seamlessly. The GPU infrastructure layer is formed by the &lt;strong&gt;NVIDIA GPU Operator&lt;/strong&gt; (currently v26.3.3), which installs drivers, container runtime, device plugin, and DCGM monitoring with a single Helm chart and enables MIG partitioning as well as time-slicing. Above that lies the scheduling layer, where &lt;strong&gt;Kueue&lt;/strong&gt; (CNCF, Kubernetes SIG project) handles admission control and quota management—crucial for multi-tenant clusters where multiple teams compete for GPUs. For distributed training with gang scheduling (where all pods of a job only start together), &lt;strong&gt;Volcano&lt;/strong&gt; (CNCF Graduated) or the open-source &lt;strong&gt;KAI Scheduler&lt;/strong&gt; is used. Often, Kueue and Volcano run as a layered system: Kueue holds back jobs when the quota is exhausted, while Volcano ensures atomic placement.&lt;/p&gt;

&lt;p&gt;For inference, &lt;strong&gt;vLLM&lt;/strong&gt; has established itself as the standard inference engine—with PagedAttention, continuous batching, and an OpenAI-compatible API. Around it, &lt;strong&gt;KServe&lt;/strong&gt; (CNCF Incubating) provides a Kubernetes-native serving layer with autoscaling (including scale-to-zero), canary deployments, and traffic splitting. For models beyond the 70B parameter range (Llama 4 405B, DeepSeek V3), &lt;strong&gt;llm-d&lt;/strong&gt; is increasingly used, a distributed inference framework with disaggregated serving that separates prefill and decode and distributes KV caches across nodes.&lt;/p&gt;

&lt;p&gt;Training workloads typically run via &lt;strong&gt;Ray&lt;/strong&gt; with the &lt;strong&gt;KubeRay&lt;/strong&gt; operator, which provides RayCluster, RayJob, and RayService as native Kubernetes CRDs. KubeRay v1.7 (August 2026) brought a history server (beta), automatic mTLS certificate management, NetworkPolicy support, and Kubernetes RBAC-based authentication—a significant leap in production readiness.&lt;/p&gt;

&lt;h2&gt;
  
  
  GPU Scheduling: The Bottleneck No One Can Ignore
&lt;/h2&gt;

&lt;p&gt;The biggest difference between classic microservices and AI workloads on Kubernetes is the handling of GPUs. The standard scheduler places pods individually—a distributed training job where seven out of eight worker pods start, but the eighth waits for a GPU, blocks the other seven GPUs at nearly zero utilization. This is exactly where gang scheduling (Volcano, KAI Scheduler) and quota queues (Kueue) come into play.Kueue boosts GPU utilization from 25–35% to 60–85% by enforcing fair-share queues across teams and enabling preemption for prioritized workloads. The combination of admission control and gang scheduling has now become the production standard for multi-tenant GPU clusters.&lt;/p&gt;

&lt;p&gt;Topology-aware scheduling is also playing a growing role: pods placed on nodes with shared NVLink achieve three to five times lower inter-GPU latency than pods spanning separate network switches.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-Tenancy and Cost Control
&lt;/h2&gt;

&lt;p&gt;Shared GPU infrastructure is expensive—an unthrottled research workload can noticeably impact the latency of production inference endpoints. In practice, the standard has settled on Namespaces + RBAC + ResourceQuotas for logical partitioning, PriorityClasses for preemption, and Kueue ClusterQueues for cross-team fair-share policies.&lt;/p&gt;

&lt;p&gt;On the cost side, &lt;strong&gt;Karpenter&lt;/strong&gt; handles dynamic node provisioning: when a training job is pending, Karpenter provisions tailored spot or on-demand instances and scales them back down once the work is done. For GPU workloads that tolerate interruptions (training with checkpointing), spot instances are a massive cost lever—cloud providers advertise up to 90% discounts compared with on-demand prices.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The AI/ML stack on Kubernetes is production-ready and standardized in 2026. The core components—NVIDIA GPU Operator, Kueue, KServe, vLLM, Ray with KubeRay—are CNCF-graduated or on their way to it, and are supported by a large community. Platform teams that invest in this stack today benefit from portability across cloud providers and on-premises environments, fair GPU distribution among teams, and an ecosystem that is evolving rapidly—KubeRay v1.7, KServe's CNCF incubation, and the growing adoption of llm-d show that the direction is clear. Anyone who wants to run AI workloads seriously will no longer be able to avoid Kubernetes as the orchestration layer in 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://kubernetesguru.com/ai-ml-on-kubernetes-2026-stack-guide/" rel="noopener noreferrer"&gt;AI/ML on Kubernetes 2026: Production Stack Guide (vLLM, Kueue, KServe, Ray) — KubernetesGuru (July 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.cloudoptimo.com/blog/kubernetes-ai-infrastructure-in-2026-gpu-scheduling-and-production-realities/" rel="noopener noreferrer"&gt;Kubernetes AI Infrastructure in 2026: GPU Scheduling &amp;amp; Production Realities — CloudOptimo (May 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anyscale.com/blog/kuberay-v1-7" rel="noopener noreferrer"&gt;Introducing KubeRay v1.6 and v1.7 — Anyscale (August 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://markaicode.com/best/best-kubernetes-setup-for-production-ai/" rel="noopener noreferrer"&gt;Best Kubernetes Setup for Production AI Workloads (2026) — Markaicode (August 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cloudstrata.io/post/kubernetes-ai-workloads-best-practices-2026/" rel="noopener noreferrer"&gt;Kubernetes and AI Workloads: Best Practices for 2026 — CloudStrata (March 2026)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>kubernetes</category>
      <category>ai</category>
      <category>devops</category>
      <category>kueue</category>
    </item>
    <item>
      <title>Knowledge Graphs vs. LLMs: Why Law Can't Get Far Without a Knowledge Graph</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Mon, 28 Sep 2026 11:37:29 +0000</pubDate>
      <link>https://dev.to/saaro_net/knowledge-graphs-vs-llms-why-law-cant-get-far-without-a-knowledge-graph-14bo</link>
      <guid>https://dev.to/saaro_net/knowledge-graphs-vs-llms-why-law-cant-get-far-without-a-knowledge-graph-14bo</guid>
      <description>&lt;p&gt;An LLM alone is like a lawyer with a photographic memory, but without understanding of chains of paragraphs. It knows that § 823 BGB exists, but not necessarily how it relates to § 254 BGB. A knowledge graph closes this gap – by mapping relationships that a pure language model does not intrinsically understand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Knowledge graphs are not competition for LLMs, but their structured complement. For legal applications, they are often superior – but not without costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Core Conflict: Vector Search vs. Graph Structure
&lt;/h2&gt;

&lt;p&gt;Modern AI knowledge systems are based on two fundamentally different approaches:&lt;/p&gt;

&lt;h3&gt;
  
  
  Vector RAG (the standard)
&lt;/h3&gt;

&lt;p&gt;Documents are cut into chunks, embedded as vectors, and queried via similarity search. Works well for "Find me a document that looks like this." Works poorly for "Which paragraphs are relevant for this case, and how do they relate?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strength:&lt;/strong&gt; Easy to set up, no manual modeling, good results for fact-based questions.&lt;br&gt;
&lt;strong&gt;Weakness:&lt;/strong&gt; No understanding of relationships between entities. A vector search might find § 823 BGB, but not automatically the associated commentaries, case law references, and exceptions from other paragraphs.&lt;/p&gt;
&lt;h3&gt;
  
  
  Knowledge Graph RAG (GraphRAG)
&lt;/h3&gt;

&lt;p&gt;Documents are modeled as nodes (entities) and edges (relationships). A knowledge graph for law could contain nodes for paragraphs, judgments, commentaries, and legal terms – and edges that map "references", "was cited in", "restricts", or "supplements".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strength:&lt;/strong&gt; Relationships are explicitly modeled. Multi-step queries ("Which judgments refer to § 823 and were issued after 2020?") deliver precise results.&lt;br&gt;
&lt;strong&gt;Weakness:&lt;/strong&gt; High initial effort, labor-intensive maintenance, more complex infrastructure.&lt;/p&gt;
&lt;h2&gt;
  
  
  Practical Example: Legal Research
&lt;/h2&gt;

&lt;p&gt;Imagine an LLM is supposed to answer the following question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"A tenant has reduced the rent due to a defective heating system. The landlord contests this. Which paragraphs, judgments, and commentaries are relevant?"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Vector RAG:&lt;/strong&gt; The system finds documents containing "rent defect", "rent reduction", and "heating". But it cannot recognize that § 536 BGB governs rent reduction, while § 569 BGB concerns extraordinary termination – and that a specific 2023 BGH ruling clarifies the balance between the two paragraphs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GraphRAG:&lt;/strong&gt; The knowledge graph contains the nodes &lt;code&gt;§ 536&lt;/code&gt;, &lt;code&gt;§ 569&lt;/code&gt;, &lt;code&gt;BGH ruling VIII ZR 123/23&lt;/code&gt; and the edges &lt;code&gt;BGH ruling VIII ZR 123/23 interprets § 536&lt;/code&gt; and &lt;code&gt;§ 569 references § 536 for heating defects&lt;/code&gt;. The system navigates along these edges and delivers a structured, context-aware answer.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;Vector RAG&lt;/th&gt;
&lt;th&gt;GraphRAG&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Setup effort&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Low (upload documents)&lt;/td&gt;
&lt;td&gt;High (model graph)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Relationship understanding&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None explicit&lt;/td&gt;
&lt;td&gt;Explicitly modeled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Multi-hop questions&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Weak&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Updates&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Re-upload&lt;/td&gt;
&lt;td&gt;Extend graph&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Data quality&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Dependent on embedding&lt;/td&gt;
&lt;td&gt;Dependent on graph design&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Interpretability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Black box (embedding)&lt;/td&gt;
&lt;td&gt;Traceable (edges)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cost (operation)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Higher (graph DB + query)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Scaling&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Linear with documents&lt;/td&gt;
&lt;td&gt;Complex with node count&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  Advantages and Disadvantages in Detail
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Advantages of Knowledge Graphs in AI Use
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. Multi-Hop Reasoning&lt;/strong&gt;&lt;br&gt;
A question like "Which exceptions apply to liability under § 823 BGB in road traffic?" requires multiple steps: find § 823 → discover § 254 (contributory negligence) → include § 7 StVG (strict liability) → find case law on this interplay. A knowledge graph navigates this chain along its edges. Vector RAG would have to hope that all relevant documents land in the same chunk.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Explainability&lt;/strong&gt;&lt;br&gt;
Every answer from a knowledge graph can be traced back via the edges. "I recommend § 536 BGB because BGH ruling VIII ZR 123/23 interpreted this paragraph in the context of heating defects." With Vector RAG, it remains unclear which chunk contributed to which part of the answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Domain-Specific Ontologies&lt;/strong&gt;&lt;br&gt;
Legal terms have precise, often ambiguous relationships. "Lawsuit" can be a "declaratory action", "performance action", or "formative action". A graph maps this hierarchy. A vector embedding does not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Consistency&lt;/strong&gt;&lt;br&gt;
A well-maintained knowledge graph is free of contradictions. LLMs with pure Vector RAG can contradict themselves depending on which chunks end up in the context.&lt;/p&gt;
&lt;h3&gt;
  
  
  Disadvantages of Knowledge Graphs
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. Initial Modeling Effort (massive)&lt;/strong&gt;&lt;br&gt;
Building a legal knowledge graph means: manually or semi-automatically capturing paragraphs, judgments, commentaries, legal terms, and their relationships. That is months of work for a complete domain. Vector RAG is ready to use in hours.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Maintenance and Updates&lt;/strong&gt;&lt;br&gt;
Laws change. New judgments are added. A knowledge graph must be continuously maintained – new nodes, new edges, sometimes new relationship types. Vector RAG: simply load new documents into the vector database.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Cost/Benefit Ratio&lt;/strong&gt;&lt;br&gt;
The Atlan blog put it succinctly in 2026: &lt;em&gt;"Vector RAG wins for single-hop questions. GraphRAG adds costs without adding accuracy for simple questions."&lt;/em&gt; For 90% of everyday AI questions, Vector RAG is perfectly sufficient.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Infrastructure Complexity&lt;/strong&gt;&lt;br&gt;
Vector DB + LLM is a standard stack (OpenAI + Pinecone, or local Chroma + Ollama). GraphRAG additionally requires a graph database (Neo4j, ArangoDB), a graph query planner, and often more computing power.&lt;/p&gt;
&lt;h2&gt;
  
  
  Hybrid Approach: The Best of Both Worlds
&lt;/h2&gt;

&lt;p&gt;Practice in 2026 increasingly relies on &lt;strong&gt;hybrid architectures&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Vector Search&lt;/strong&gt; for breadth: Find relevant documents, paragraphs, judgments&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Knowledge Graph&lt;/strong&gt; for depth: Navigate relationships between the found entities&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM&lt;/strong&gt; for synthesis: Formulate the answer in natural language&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;An example from research is &lt;strong&gt;LegalGraphRAG&lt;/strong&gt; (ACL 2026), which pursues exactly this approach for legal corpora: vector search for the initial retrieval phase, graph navigation for multi-hop refinement.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Query → Vector Search (find relevant documents)
      → Graph Traversal (explore relationships)
      → LLM (synthesize answer)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Who Benefits from GraphRAG?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Vector RAG&lt;/th&gt;
&lt;th&gt;GraphRAG&lt;/th&gt;
&lt;th&gt;Hybrid&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;General knowledge base (FAQ, manual)&lt;/td&gt;
&lt;td&gt;✅ Optimal&lt;/td&gt;
&lt;td&gt;❌ Overkill&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Legal research (individual case + sources)&lt;/td&gt;
&lt;td&gt;⚠️ Moderate&lt;/td&gt;
&lt;td&gt;✅ Strong&lt;/td&gt;
&lt;td&gt;✅ Optimal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compliance checks (many interlinked rules)&lt;/td&gt;
&lt;td&gt;❌ Weak&lt;/td&gt;
&lt;td&gt;✅ Strong&lt;/td&gt;
&lt;td&gt;✅ Optimal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scientific literature research&lt;/td&gt;
&lt;td&gt;✅ Good&lt;/td&gt;
&lt;td&gt;⚠️ Moderate&lt;/td&gt;
&lt;td&gt;✅ Optimal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Customer service (standard questions)&lt;/td&gt;
&lt;td&gt;✅ Optimal&lt;/td&gt;
&lt;td&gt;❌ Overkill&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Knowledge graphs are not a silver bullet, but for domain-specific, relationship-rich applications like law, they are the decisive advantage over raw vector search. The price is high initial effort – which pays off when questions are complex and relationships are critical.&lt;/p&gt;

&lt;p&gt;For most other applications, Vector RAG is perfectly sufficient. The art is knowing when the graph is worth the additional complexity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://aclanthology.org/2026.acl-long.1738/" rel="noopener noreferrer"&gt;LegalGraphRAG: Multi-Agent Graph Retrieval (ACL 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://atlan.com/know/knowledge-graphs-vs-rag-for-ai/" rel="noopener noreferrer"&gt;Knowledge Graph vs RAG: When Each One Wins (Atlan, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.meilisearch.com/blog/graph-rag-vs-vector-rag" rel="noopener noreferrer"&gt;GraphRAG vs. Vector RAG (Meilisearch, 2025)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://neo4j.com/blog/developer/from-legal-documents-to-knowledge-graphs/" rel="noopener noreferrer"&gt;From Legal Documents to Knowledge Graphs (Neo4j, 2025)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2601.13806" rel="noopener noreferrer"&gt;Knowledge Graph-Assisted LLM for Legal Reasoning (arXiv, 2026)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>wissensgraph</category>
      <category>knowledgegraph</category>
      <category>rag</category>
    </item>
    <item>
      <title>MCP 2026-07-28: The Model Context Protocol Comes of Age – Stateless, Scalable, and Enterprise-Ready</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Sun, 27 Sep 2026 12:00:16 +0000</pubDate>
      <link>https://dev.to/saaro_net/mcp-2026-07-28-the-model-context-protocol-comes-of-age-stateless-scalable-and-enterprise-ready-2a57</link>
      <guid>https://dev.to/saaro_net/mcp-2026-07-28-the-model-context-protocol-comes-of-age-stateless-scalable-and-enterprise-ready-2a57</guid>
      <description>&lt;p&gt;July 2026 marks a turning point for the Model Context Protocol (MCP): the &lt;code&gt;2026-07-28&lt;/code&gt; specification represents the biggest leap forward to date. The open-source protocol initiated by Anthropic, often called "the USB-C of AI tool integration," has thus grown from an experimental standard into production-ready infrastructure. For anyone developing or operating AI agents, a closer look at the innovations is worthwhile.&lt;/p&gt;

&lt;h2&gt;
  
  
  From Session Mandate to Stateless Core
&lt;/h2&gt;

&lt;p&gt;By far the most important change is the shift to a &lt;strong&gt;stateless Protocol Core&lt;/strong&gt;. Where MCP previously required a complex &lt;code&gt;initialize&lt;/code&gt;/&lt;code&gt;initialized&lt;/code&gt; handshake and an &lt;code&gt;Mcp-Session-Id&lt;/code&gt;, every request is now self-contained. Protocol version, client identity, and capabilities travel within the &lt;code&gt;_meta&lt;/code&gt; field of each call. A client can query a server's capabilities using the new &lt;code&gt;server/discover&lt;/code&gt; method — but it doesn't have to.&lt;/p&gt;

&lt;p&gt;The practical impact is enormous: An MCP server that previously needed sticky sessions, a shared session store, and deep packet inspection at the gateway can now run behind a simple round-robin load balancer. Gateway operators can route requests using the new HTTP headers &lt;code&gt;Mcp-Method&lt;/code&gt; and &lt;code&gt;Mcp-Name&lt;/code&gt; without having to parse the JSON body. List results (&lt;code&gt;tools/list&lt;/code&gt;, &lt;code&gt;prompts/list&lt;/code&gt;) also carry cache hints (&lt;code&gt;ttlMs&lt;/code&gt;, &lt;code&gt;cacheScope&lt;/code&gt;), allowing clients to cache tool catalogs across multiple calls.&lt;/p&gt;

&lt;p&gt;Those who still need state across multiple calls can manage it themselves via explicit handles (e.g., a &lt;code&gt;basket_id&lt;/code&gt;) — a pattern that, according to the maintainers, is often even more powerful than the old session model because the state becomes visible and combinable for the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Extensions Become First Class
&lt;/h2&gt;

&lt;p&gt;With the new specification, MCP introduces a formal &lt;strong&gt;Extensions Framework&lt;/strong&gt;. Two extensions stand out in particular:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MCP Apps&lt;/strong&gt; enable server-side rendered user interfaces. An MCP server can now deliver not just JSON data but interactive UI components — comparable to apps in an IDE extension. This opens doors for dashboards, configuration masks, and visual feedback loops directly within the agent workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tasks&lt;/strong&gt; moves from the experimental core into its own extension and receives a poll-based &lt;code&gt;tasks/get&lt;/code&gt; as well as a new &lt;code&gt;tasks/update&lt;/code&gt;. Change notifications will in future run through a central &lt;code&gt;subscriptions/listen&lt;/code&gt; stream, which clients subscribe to per notification type.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi Round-Trip Requests: The Server Asks Back
&lt;/h2&gt;

&lt;p&gt;Previously, MCP required a permanently open SSE connection for queries back to the user — for instance, for confirmation before deleting files. With &lt;strong&gt;Multi Round-Trip Requests (MRTR, SEP-2322)&lt;/strong&gt;, this is no longer necessary. The server responds with &lt;code&gt;resultType: "input_required"&lt;/code&gt; and lists the required inputs. The client collects the responses and repeats the original call with appended &lt;code&gt;inputResponses&lt;/code&gt;. Any server instance can handle the repetition because all necessary information is contained in the payload. Supabase, for example, was precisely waiting for this function to get costly or destructive actions confirmed before execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Authorization: Enterprise-Ready
&lt;/h2&gt;

&lt;p&gt;The security architecture has been hardened in several points. Authorization servers must now deliver the &lt;code&gt;iss&lt;/code&gt; parameter according to RFC 9207, and clients must validate it — this closes a dangerous vulnerability where an attacker could substitute a foreign authorization server. Dynamic Client Registration (DCR) is being phased out in favor of Client Metadata Documents (CIMD). The transition runs over twelve months, but new implementations should directly use CIMD.&lt;/p&gt;

&lt;p&gt;Roots, sampling, and logging are officially &lt;strong&gt;deprecated&lt;/strong&gt;. They will still work for at least twelve months, but new builds should no longer use them. The legacy HTTP+SSE transport is also considered deprecated.&lt;/p&gt;

&lt;h2&gt;
  
  
  The SDKs Are Ready
&lt;/h2&gt;

&lt;p&gt;All four Tier-1 SDKs (TypeScript, Python, Go, C#) support the &lt;code&gt;2026-07-28&lt;/code&gt; specification, and the Rust SDK supports it in beta stage. The download numbers speak volumes: The MCP SDKs collectively register &lt;strong&gt;nearly half a billion downloads per month&lt;/strong&gt;, with TypeScript and Python each having surpassed the billion mark in total downloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Over the last 18 months, MCP has evolved from a niche protocol into the dominant integration layer for AI agents — 78% of enterprise AI teams are already using it in production, and 67% of CTOs name MCP as their standard. The &lt;code&gt;2026-07-28&lt;/code&gt; specification solves the central performance problem (state management) and simultaneously delivers enterprise features like hardened authorization, cacheable lists, and a formalized extensions model. Anyone starting with MCP or migrating today is investing in a platform that will form the backbone of agentic AI integration for years to come.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/" rel="noopener noreferrer"&gt;MCP Blog: The 2026-07-28 Specification&lt;/a&gt; — Official release announcement with full changelog&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/" rel="noopener noreferrer"&gt;MCP Blog: 2026-07-28 Release Candidate&lt;/a&gt; — Technical deep dive into the RC, stateless core and MRTR&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://baeseokjae.github.io/posts/mcp-enterprise-adoption-guide-2026/" rel="noopener noreferrer"&gt;RockB: MCP Enterprise Adoption Guide 2026&lt;/a&gt; — Adoption statistics, gateway patterns, security framework&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/ai_maya_063fc568e157562fd/mcp-in-2026-how-the-model-context-protocol-became-the-usb-c-of-ai-tooling-3bfi"&gt;DEV Community: MCP in 2026 – The USB-C of AI Tooling&lt;/a&gt; — Practical perspective on model-agnostic tooling&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://blog.modelcontextprotocol.io/posts/2026-mcp-roadmap/" rel="noopener noreferrer"&gt;MCP Blog: The 2026 MCP Roadmap&lt;/a&gt; — Priority areas and working group structure&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>mcp</category>
      <category>modelcontextprotocol</category>
      <category>ai</category>
      <category>kiagenten</category>
    </item>
    <item>
      <title>prompts.chat: The World's Largest Open-Source Prompt Library (167,000 Stars)</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Sat, 26 Sep 2026 12:00:39 +0000</pubDate>
      <link>https://dev.to/saaro_net/promptschat-the-worlds-largest-open-source-prompt-library-167000-stars-461g</link>
      <guid>https://dev.to/saaro_net/promptschat-the-worlds-largest-open-source-prompt-library-167000-stars-461g</guid>
      <description>&lt;p&gt;Two years ago, it was a simple GitHub list called "Awesome ChatGPT Prompts". Today, it has become the world's largest open-source prompt library – with over &lt;strong&gt;167,000 stars&lt;/strong&gt;, 21,600 forks, and its own web app at &lt;strong&gt;prompts.chat&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;prompts.chat is not just another GitHub repository with a static list of prompts. It is a fully-fledged &lt;strong&gt;Next.js web application&lt;/strong&gt; that allows you to browse, discover, share, and collect prompts. And the best part: You can &lt;strong&gt;self-host&lt;/strong&gt; the entire platform – with complete data sovereignty.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is prompts.chat?
&lt;/h2&gt;

&lt;p&gt;The project was started by Fatih Kadir Akın and has evolved from a simple awesome list into one of the most influential prompt collections in the world. It is cited in Forbes, referenced by Harvard and Columbia, and has over &lt;strong&gt;40 academic citations&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The range of prompts is enormous – from "Act as a Linux Terminal" to "Act as a Career Counselor" to specialized developer prompts. The community has submitted thousands of prompts, sorted by category and searchable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why 167,000 Stars?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Quality over Quantity&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Every prompt in the collection is reviewed by the community. There is no SEO spam, no AI-generated filler prompts. The collection is curated and focused.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Works with every LLM&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Unlike many specialized collections (that work only for Claude or only for GPT), the prompts from prompts.chat are compatible with &lt;strong&gt;ChatGPT, Claude, Gemini, Llama, Mistral, and all common models&lt;/strong&gt;. This makes them universally usable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. From Harvard to Forbes&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The academic and media recognition is exceptional: Harvard University and Columbia University reference the collection in their guides, Forbes has called it an indispensable resource. This builds trust and visibility.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. MCP Server for Coding Agents&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
prompts.chat offers an &lt;strong&gt;MCP server (Model Context Protocol)&lt;/strong&gt; that can be directly integrated into coding agents. In 2026, this is the decisive factor: coding agents can retrieve prompts at runtime instead of storing them locally.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Technical Platform
&lt;/h2&gt;

&lt;p&gt;The repository contains more than just prompts – it is a complete web application:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Technology&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Frontend&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Next.js (React, TypeScript)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Database&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Prisma + PostgreSQL&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Search&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Full-text search across all prompts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MCP Server&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Model Context Protocol API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Deployment&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Docker (Self-Hosting), Vercel&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Export&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;prompts.csv, PROMPTS.md&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The MCP server is particularly exciting: You can integrate prompts.chat directly into Claude Code, Windsurf, or other MCP-capable tools – without manually browsing the website:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"prompts.chat"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://prompts.chat/api/mcp"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Self-Hosting for Companies
&lt;/h2&gt;

&lt;p&gt;Perhaps the most important aspect for companies: &lt;strong&gt;prompts.chat can be completely self-hosted&lt;/strong&gt;. You install the platform on your own servers, your prompt library stays internal. This is GDPR-compliant and gives you full control over the content.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/f/prompts.chat
&lt;span class="nb"&gt;cd &lt;/span&gt;prompts.chat
&lt;span class="nb"&gt;cp&lt;/span&gt; .env.example .env
docker compose up &lt;span class="nt"&gt;-d&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After that, your own prompt library runs on an internal URL – with the same features as the public platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dual Licensing
&lt;/h2&gt;

&lt;p&gt;An interesting detail: prompts.chat uses a dual licensing model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Source code and site content:&lt;/strong&gt; MIT License&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt content and data (CSV, PROMPTS.md):&lt;/strong&gt; CC0 1.0 (Public Domain)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The prompts themselves are therefore public domain. You can copy them, share them, use them in your own projects – without restrictions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who is this interesting for?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Developers:&lt;/strong&gt; The prompt collection as a reference for everyday work, plus MCP server for coding agents&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Teams:&lt;/strong&gt; Self-hosted prompt library for consistent prompts across the entire team&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Companies:&lt;/strong&gt; GDPR-compliant, internal prompt management without external dependencies&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt Engineers:&lt;/strong&gt; Inspiration and comparison options for their own prompts&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I particularly like
&lt;/h2&gt;

&lt;p&gt;The evolution from "Awesome ChatGPT Prompts" to prompts.chat shows how far the topic of prompt engineering has come in two years. From a static list, it became a platform with its own infrastructure, API, MCP server, and enterprise self-hosting. This is open-source development at its best: a project that grows with its users.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://prompts.chat" rel="noopener noreferrer"&gt;prompts.chat – Website&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/f/prompts.chat" rel="noopener noreferrer"&gt;GitHub Repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://raw.githubusercontent.com/f/prompts.chat/main/PROMPTS.md" rel="noopener noreferrer"&gt;PROMPTS.md – all prompts (Markdown)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/datasets/fka/prompts.chat" rel="noopener noreferrer"&gt;Hugging Face Dataset&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>opensource</category>
      <category>prompts</category>
      <category>promptengineering</category>
      <category>chatgpt</category>
    </item>
    <item>
      <title>LLM-Wiki 2026: From Gist to Ecosystem – Obsidian Plugin, Enterprise Adoption, and Six Months of Community Practice</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Sat, 26 Sep 2026 12:00:03 +0000</pubDate>
      <link>https://dev.to/saaro_net/llm-wiki-2026-from-gist-to-ecosystem-obsidian-plugin-enterprise-adoption-and-six-months-of-2dgd</link>
      <guid>https://dev.to/saaro_net/llm-wiki-2026-from-gist-to-ecosystem-obsidian-plugin-enterprise-adoption-and-six-months-of-2dgd</guid>
      <description>&lt;p&gt;It has been six months since Andrej Karpathy published an unassuming gist on GitHub—a few hundred lines of prompt patterns, no installer, no repository. The idea: Instead of searching through documents anew for every question like with RAG, an AI agent builds a permanent, linked knowledge base in Markdown from the sources—the "LLM Wiki." What seemed like an experiment for AI enthusiasts back then has developed into one of the most discussed patterns in personal knowledge management. With 5,000 stars and just as many forks, it is becoming clear: an ecosystem is emerging here.&lt;/p&gt;

&lt;h2&gt;
  
  
  From Prompt Pattern to Platform: The Obsidian Plugin
&lt;/h2&gt;

&lt;p&gt;The most influential further development of the Karpathy pattern is an Obsidian plugin that integrates the gist into a full-fledged editor. Developed by gd4ai and Greener-Dalii, the plugin reached &lt;strong&gt;45,000 downloads&lt;/strong&gt; within five months and is currently available in version &lt;strong&gt;1.27.1&lt;/strong&gt;. The decisive difference compared to other implementations: It is a pure Obsidian plugin without external dependencies. No Python runtime, no vector embedding model, no separate desktop app—the entire workflow lives within the editor.&lt;/p&gt;

&lt;p&gt;The architecture relies on &lt;strong&gt;PPR (Personalized PageRank) plus Monte-Carlo search over the &lt;code&gt;[[wiki-link]]&lt;/code&gt; graph&lt;/strong&gt; instead of conventional RAG. The plugin supports over 16 AI providers (Anthropic, OpenAI, Gemini, DeepSeek, Qwen, Ollama, LM Studio, OpenRouter, etc.), and the data never leaves the device when a local model is used. PDFs, Office documents, and images are processed natively. The community has since translated the plugin into &lt;strong&gt;eleven languages&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 200-File Limit and Its Solution
&lt;/h2&gt;

&lt;p&gt;One of the most interesting insights from months of practice comes from Kunal Ganglani, who used the Karpathy pattern daily for three months and documented his experiences in detail. His central finding: &lt;strong&gt;Once the wiki exceeds around 150 to 200 files, most AI agents can no longer keep the complete graph in context.&lt;/strong&gt; The quality of links and updates noticeably suffers.&lt;/p&gt;

&lt;p&gt;The community has developed a pragmatic workaround for this: a &lt;strong&gt;master index&lt;/strong&gt; that lists every page with a one-line summary. The agent first reads the index, then selectively loads only those pages relevant to a specific update. In practice, this approach extends capacity to over 300 pages. For even larger knowledge bases, &lt;strong&gt;qmd&lt;/strong&gt; exists as a CLI tool that offers local hybrid search (BM25 plus vector with re-ranking) as an MCP server—essentially RAG over the wiki when the wiki itself becomes too large.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enterprise Adoption: When the Pattern Meets the Organization
&lt;/h2&gt;

&lt;p&gt;The natural question after six months: Can the personal pattern be transferred to companies? Falconer, a company specializing in knowledge management, has taken on this question and identified the &lt;strong&gt;four properties&lt;/strong&gt; that make Karpathy's wiki successful: Capture, Link, Compound, and Stay Current.&lt;/p&gt;

&lt;p&gt;The analysis shows: &lt;strong&gt;None of these properties can be scaled directly to enterprise size.&lt;/strong&gt; The personal &lt;code&gt;raw/&lt;/code&gt; folder, where the user manually curates sources, has no equivalent in a company. Relevant sources are distributed across GitHub, Slack, Linear, Confluence, Google Drive, and dozens of other tools. Bidirectional links would have to work across tool boundaries—a Slack decision would need to be linked to the implementing PR, the Linear ticket, and the meeting transcript. And the automatic health checks that Karpathy regularly runs on his personal wiki would have to function in an organization without a manual curator.&lt;/p&gt;

&lt;p&gt;Y Combinator explicitly named this missing foundation in the Spring 2026 RFS: "A Company Brain that AI agents can actually use"—an infrastructure that does not exist today. The Stack Overflow Developer Survey 2024 demonstrates the urgency: Over 60% of developers spend at least 30 minutes daily searching for solutions, and 68% encounter a knowledge silo at least once a week. Among managers, the figure rises to 73%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Comparison of Implementations
&lt;/h2&gt;

&lt;p&gt;In six months, the LLM wiki ecosystem has produced several competing approaches: &lt;strong&gt;SamurAIGPT/llm-wiki-agent&lt;/strong&gt; (around 3,300 GitHub stars, as a Claude‑Code/Codex skill) relies on Louvain community detection and SHA256 caching, &lt;strong&gt;nashsu/llm_wiki&lt;/strong&gt; (10,000+ stars) provides a Tauri desktop app, and &lt;strong&gt;atomicstrata/llm-wiki-compiler&lt;/strong&gt; works as a TypeScript CLI with BM25 plus semantic search over chunks. The Obsidian plugin remains the only one that works without an additional runtime.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The LLM Wiki has left the proof of concept behind. The Obsidian plugin with 45,000 downloads proves that the pattern works productively for individual users—with limitations starting at 200 files, which are, however, addressed by community solutions. The more exciting question is enterprise adoption: Here, the infrastructure is still missing to automatically scale the four success properties to enterprise size. Y Combinator has recognized the need—whether and when the first satisfactory solution arrives will be one of the most interesting topics in the coming months.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.kunalganglani.com/blog/llm-wiki-karpathy-local-knowledge-base" rel="noopener noreferrer"&gt;LLM Wiki Setup: Karpathy's Knowledge Base [2026 Guide] – Kunal Ganglani&lt;/a&gt; — Three-month practice report with the 200-file limit and workarounds&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://community.obsidian.md/plugins/karpathywiki" rel="noopener noreferrer"&gt;Karpathy LLM Wiki – Obsidian Plugin&lt;/a&gt; — Plugin documentation with comparison table and architecture description&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://falconer.com/guides/enterprise-llm-wiki-karpathy/" rel="noopener noreferrer"&gt;The Enterprise LLM Wiki – Falconer Guides&lt;/a&gt; — Analysis of the four properties and their scalability to companies&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f" rel="noopener noreferrer"&gt;Karpathy's LLM Wiki Gist&lt;/a&gt; — Original prompt pattern (5,000+ stars, 5,000+ forks)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://survey.stackoverflow.co/2024/" rel="noopener noreferrer"&gt;Stack Overflow Developer Survey 2024&lt;/a&gt; — User data on knowledge silos and search times&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>llmwiki</category>
      <category>ai</category>
      <category>knowledgemanagement</category>
      <category>obsidian</category>
    </item>
    <item>
      <title>The Agency: 230 AI agents for Claude Code – a GitHub repository with 154,000 stars</title>
      <dc:creator>saaro</dc:creator>
      <pubDate>Fri, 25 Sep 2026 11:54:20 +0000</pubDate>
      <link>https://dev.to/saaro_net/the-agency-230-ai-agents-for-claude-code-a-github-repository-with-154000-stars-16ni</link>
      <guid>https://dev.to/saaro_net/the-agency-230-ai-agents-for-claude-code-a-github-repository-with-154000-stars-16ni</guid>
      <description>&lt;p&gt;Imagine having 230 AI specialists at your fingertips – an editorial expert for marketing texts, a security auditor for code reviews, a UX designer for interface feedback, a Reddit community ninja, and a lease contract lawyer. All waiting to be activated by you in Claude Code. No joke.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Agency&lt;/strong&gt; is a GitHub repository with over &lt;strong&gt;154,000 stars&lt;/strong&gt;, born from a simple idea: specialized AI agents for every conceivable task – as Claude Code prompts, organized like a real agency with departments.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is The Agency?
&lt;/h2&gt;

&lt;p&gt;The repository was started by &lt;a href="https://github.com/msitarzewski" rel="noopener noreferrer"&gt;msitarzewski&lt;/a&gt; after a Reddit thread about specialized AI agents went viral. Since then, over 24,000 forks and hundreds of contributors have built on it. The README claim says it all: &lt;em&gt;"A complete AI agency at your fingertips."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The agents are organized into &lt;strong&gt;18 departments&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Department&lt;/th&gt;
&lt;th&gt;Example Agents&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Engineering&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Architect, Code Reviewer, Tech Lead, Debugging Specialist, Database Designer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Design&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;UI Designer, UX Researcher, Design System Guardian, Accessibility Expert&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Marketing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Copywriter, SEO Strategist, Social Media Manager, Brand Voice Guardian&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Sales&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Cold Email Specialist, Demo Script Writer, Objection Handler, CRM Optimizer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Product&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Product Manager (APR), User Story Writer, Roadmap Planner, A/B Test Designer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Finance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Financial Analyst, Budget Planner, Invoice Reviewer, Tax Assistant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Healthcare&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Clinical Trial Writer, ICD Coder, Medical Literature Reviewer, Patient Communicator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Security&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Penetration Tester, Compliance Auditor, Incident Responder, Threat Modeler&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Strategy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Competitive Analyst, OKR Coach, Pitch Deck Reviewer, Market Researcher&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Project Management&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Scrum Master, Risk Manager, Stakeholder Communicator, Retro Facilitator&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Research&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Literature Reviewer, Hypothesis Generator, Methodology Advisor, Data Interpreter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Support&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Triage Agent, Knowledge Base Curator, Escalation Handler, SOP Writer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Academic&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Paper Reviewer, Grant Writer, Thesis Advisor, Citation Manager&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Game Development&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Level Designer, Narrative Designer, Game Balance Analyst, QA Tester&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Sales&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Email Campaign Specialist&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Testing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Test Case Generator, Regression Analyzer, Performance Tester, E2E Scenario Writer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GIS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Map Analyst, Geodata Expert&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spatial Computing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;AR/VR Specialist, 3D Model Reviewer&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;… and many more. Over &lt;strong&gt;230 specialized agents&lt;/strong&gt; in total.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;Each agent is a clearly defined prompt template with:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Role and Personality&lt;/strong&gt; – "You are a Senior Frontend Architect with 15 years of experience"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Process&lt;/strong&gt; – How the agent proceeds (Analysis → Plan → Execution → Review)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deliverables&lt;/strong&gt; – What the agent delivers at the end&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Usage is simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Clone repository&lt;/span&gt;
git clone https://github.com/msitarzewski/agency-agents.git
&lt;span class="c"&gt;# Copy desired agent&lt;/span&gt;
&lt;span class="nb"&gt;cp &lt;/span&gt;agency-agents/engineering/architect.md ~/.claude/agents/
&lt;span class="c"&gt;# Activate in Claude Code&lt;/span&gt;
/claude-code:agent architect
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or you copy individual agent prompts directly from the repository and paste them into your Claude Code conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the project is so successful
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. From Reddit thread to movement&lt;/strong&gt;&lt;br&gt;
What began as a discussion about specialized AI assistants has grown into one of the largest prompt collections in the world. The author himself says: &lt;em&gt;"What started as a Reddit thread about AI agent specialization has grown into something remarkable."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Quality through specialization&lt;/strong&gt;&lt;br&gt;
Instead of a generic "Help me with code," you get an agent that knows exactly what to do. The security auditor thinks like a pentester, the UX designer like a Nielsen Norman graduate, the financial analyst like an auditor.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Community-driven&lt;/strong&gt;&lt;br&gt;
The project thrives on contributors. Anyone can submit new agents, improve existing ones, or report bugs. There are Discussions, Issues, active development.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Translations&lt;/strong&gt;&lt;br&gt;
The agents are available in multiple languages – Korean, Japanese, Vietnamese. German translations are still open (who wants to join?).&lt;/p&gt;

&lt;h2&gt;
  
  
  Two critical points
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quality varies.&lt;/strong&gt; With 230+ agents from different authors, quality is not uniform. Some agents are excellently researched, others seem hastily put together. The stars of the repository are the Engineering and Security departments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Code-centric.&lt;/strong&gt; The agents are tailored to Claude Code. With other coding agents (Codex CLI, Cursor), they only work with adjustments.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does this mean for practice?
&lt;/h2&gt;

&lt;p&gt;The Agency is more than a prompt collection. It is an &lt;strong&gt;operating system overlay for AI-powered work&lt;/strong&gt;: Instead of having to think anew each time which prompt you need, you rely on proven, community-reviewed templates.&lt;/p&gt;

&lt;p&gt;I find particularly valuable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Engineering:&lt;/strong&gt; Architect + Code Reviewer + Debugger – a complete development pipeline&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security:&lt;/strong&gt; Threat Modeler + Penetration Tester + Compliance Auditor&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strategy:&lt;/strong&gt; Competitive Analyst + Pitch Deck Reviewer + OKR Coach&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The Agency is an impressive community project that shows how far specialized AI agents have come in 2026. 154,000 stars don't lie – the idea hits a nerve. Anyone working with Claude Code will certainly find an agent here that improves their own workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Only limitation:&lt;/strong&gt; Don't trust blindly. Check each agent against your own requirements before first use. But that applies to human employees as well.&lt;/p&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/msitarzewski/agency-agents" rel="noopener noreferrer"&gt;The Agency on GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://reddit.com/r/ClaudeAI" rel="noopener noreferrer"&gt;Discussion on Reddit (r/ClaudeAI)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/msitarzewski/agency-agents" rel="noopener noreferrer"&gt;The Agency&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>opensource</category>
      <category>kiagenten</category>
      <category>claudecode</category>
      <category>github</category>
    </item>
  </channel>
</rss>
