<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Composio</title>
    <description>The latest articles on DEV Community by Composio (composiodev).</description>
    <link>https://dev.to/composiodev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F9157%2Fdf89ab52-2d48-474b-a971-087232b09f19.png</url>
      <title>DEV Community: Composio</title>
      <link>https://dev.to/composiodev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/composiodev"/>
    <language>en</language>
    <item>
      <title>Top Skills &amp; Plugins to use with Grok Build 2026</title>
      <dc:creator>Developer Harsh</dc:creator>
      <pubDate>Thu, 30 Jul 2026 13:39:35 +0000</pubDate>
      <link>https://dev.to/composiodev/top-skills-plugins-to-use-with-grok-build-2026-39a0</link>
      <guid>https://dev.to/composiodev/top-skills-plugins-to-use-with-grok-build-2026-39a0</guid>
      <description>&lt;p&gt;xAI released Grok Build in May, and it’s been improving steadily since then.&lt;/p&gt;

&lt;p&gt;With support for spawning up to 8 parallel subagents and a recently added system for skills and plugins, Grok is now a full-fledged ecosystem that rivals contenders like Claude Code and Codex.&lt;/p&gt;

&lt;p&gt;However, unlike Claude or Codex, Grok comes with a superpower - it can scrape X data, the real-time engine behind every major announcement, quality content, trends and conspiracies, all with your X subscription.&lt;/p&gt;

&lt;p&gt;This gives Grok Build a unique edge for research-heavy workflows. You can track launches as they happen, pull insights from real conversations, spot emerging trends, and turn that live context into apps, agents, or automated workflows.&lt;/p&gt;

&lt;p&gt;Access to real-time data is only one part of the equation, though. To make that information useful, Grok needs the right tools to search, process, design, code, and take action across different platforms.&lt;/p&gt;

&lt;p&gt;But with so many options in place for a single need, it's hard to find the right one.&lt;/p&gt;

&lt;p&gt;This guide aims to cover which ones are worth installing on your first install, why to install them, and how to install them.&lt;/p&gt;

&lt;p&gt;Let’s begin with a quick refresher on what skills and plugins are and why you should install them.&lt;/p&gt;




&lt;h2&gt;
  
  
  What are Skills &amp;amp; Plugins &amp;amp; Why They Matter
&lt;/h2&gt;

&lt;p&gt;The concept of skills and plugins is not new, yet people still often interchange them. Both extend the capabilities of Grok Build and are related but not the same.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Skills&lt;/strong&gt; are small, reusable instruction packs (usually a SKILL.md plus optional scripts) that turn Grok into a consistent specialist.&lt;/p&gt;

&lt;p&gt;They eliminate repetitive, long prompts, enforce high-quality practices such as TDD or careful planning, reduce token waste through focused behaviour, and deliver the same reliable results across projects.&lt;/p&gt;

&lt;p&gt;On the other hand;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Plugins&lt;/strong&gt; are larger, one-command packages that bundle one or more skills with MCP servers, automation hooks, sub-agents, and platform integrations, giving Grok real superpowers.&lt;/p&gt;

&lt;p&gt;With plugins, Grok Build can do live web research, control browsers, perform database operations, analyse production errors, and enable seamless deployments.&lt;/p&gt;

&lt;p&gt;This is essential for complex agentic workflows and is easy to install, adopt, and share.&lt;/p&gt;

&lt;p&gt;With that clarification done, let’s look at how to install skills and plugins in Grok Build before looking at some of the best skills and plugins you should check out/install first.&lt;/p&gt;

&lt;p&gt;Related: Best OpenCode Skills&lt;/p&gt;




&lt;h2&gt;
  
  
  How to install Skills &amp;amp; Plugins in Grok Build
&lt;/h2&gt;

&lt;h3&gt;
  
  
  kills
&lt;/h3&gt;

&lt;p&gt;The Official ones listed on the marketplace can be accessed using &lt;code&gt;/marketplace&lt;/code&gt; inside Grok Build itself.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Finisf53vsh7yzg843q13.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Finisf53vsh7yzg843q13.png" alt="Way 1" width="799" height="454"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Or you can try manual placement&lt;/p&gt;

&lt;p&gt;Put skills in &lt;code&gt;./.grok/skills/&lt;/code&gt; or &lt;code&gt;~/.grok/skills/&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs38thr6bsizslal6e7y3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs38thr6bsizslal6e7y3.png" alt="Way 2" width="351" height="292"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;or add an extra path in the ~/.grok/config.toml under [skills] .&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhse4pcvta4hw8rq1y7el.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhse4pcvta4hw8rq1y7el.png" alt="Way 3" width="800" height="528"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Plugins
&lt;/h3&gt;

&lt;p&gt;Grok Plugins can be installed using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; grok plugin install &amp;lt;name&amp;gt; --trust
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and then verify the plugin by doing &lt;code&gt;/plugin&lt;/code&gt; . &lt;/p&gt;

&lt;p&gt;If it fails, use the Grok Build Marketplace to add it as a plugin.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgaug1ab1lgmf91by97ik.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgaug1ab1lgmf91by97ik.png" alt="Way 1" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Or you can add extra paths in ~/.grok/config.toml under  [plugins] &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpr6u70a4dh36gzjc6rik.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpr6u70a4dh36gzjc6rik.png" alt="Way 2" width="800" height="521"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For both skills and marketplace,  Grok Build comes with automatic compatibility support for &lt;code&gt;.claude/skills&lt;/code&gt; and &lt;code&gt;agent/skills&lt;/code&gt;.  Just put skills there and let grok build pick it up.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Alternative
&lt;/h3&gt;

&lt;p&gt;For non-official skills packs like skills.sh use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;npx skills@latest add &amp;lt;skill&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We will use any one of the methods listed going forward. Now time to look at top skills and plugins.&lt;/p&gt;




&lt;h2&gt;
  
  
  Top Skills to Use with Grok Build CLI in 2026
&lt;/h2&gt;

&lt;p&gt;These are the top skills I would install if I reinstall Grok Build. Most of them still live in my workspace.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Composio CLI: Power Grok with 1000+ apps from GitHub, Linear, to Figma, Canva.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmjosnvdd5nhisf0wrcgp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmjosnvdd5nhisf0wrcgp.png" alt="Compsio" width="800" height="419"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Composio provides Grok Build with access to more than 1,000 applications via a single remote MCP connection. Instead of loading every integration action into the model’s context, it exposes seven meta-tools that let the agent find tools, initiate authorisation, and execute actions when needed.&lt;/p&gt;

&lt;p&gt;This makes it useful for workflows involving applications such as  GitHub, Linear, Jira, Figma, and other external services. When an application has not been connected, Composio can generate an OAuth authorisation link for the user.&lt;/p&gt;

&lt;p&gt;You can install Composio using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;grok mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; http composio https://connect.composio.dev/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then type &lt;code&gt;/mcp&lt;/code&gt; , select Composio and complete the OAuth flow.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: In WSL you can’t access the browser directly, so copy-paste the produced URL and configure it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Alternatively; &lt;/p&gt;

&lt;p&gt;You can type &lt;code&gt;/mcp&lt;/code&gt;  inside Grok Build, in the mcp window press a to add a new mcp. Add the  &lt;code&gt;https://connect.composio.dev/mcp&lt;/code&gt;  and initiate the OAuth flow&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgq9mo0uigg01wvf4yoin.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgq9mo0uigg01wvf4yoin.png" alt="Step 1" width="799" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdrpztur8uewbs6e55vg4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdrpztur8uewbs6e55vg4.png" alt="Step 2" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Or install the Composio CLI directly. It's a CLI for everything Composio&amp;nbsp;that can handle authentication management, tool calling, bash scripting and everything in-between.&lt;/p&gt;

&lt;p&gt;This gives a more composable way for Grok CLI to work with Composio toolkits.&lt;/p&gt;

&lt;p&gt;You can install it with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;curl -fsSL https://composio.dev/install | bash
composio login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Complete the OAuth flow and then add the composio-cli skill&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;composio --install-skill composio-cli claude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This ensures that composio-cli, used by Grok Build, follows the correct instructions and doesn’t hallucinate.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The reason I put this one at the top is because , it offers the skill, plugins and mcp all bundled together under one ecosystem, so one time config is all you need.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  2. Matt Pocock Skills: Add disciplined planning, TDD, debugging, and handoff workflows.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fek4knyhxmgvui4ny8hlt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fek4knyhxmgvui4ny8hlt.png" alt="Matt Pocock" width="738" height="388"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Matt Pocock Skills adds proper engineering discipline to Grok Build. &lt;/p&gt;

&lt;p&gt;Similar to superpowers, it helps the agent plan more effectively, avoid common failure modes, and follow structured processes rather than jumping straight into code.&lt;/p&gt;

&lt;p&gt;It includes practical skills like &lt;code&gt;/grill-me&lt;/code&gt; for questioning plans, &lt;code&gt;/tdd&lt;/code&gt; for test-first development, &lt;code&gt;/diagnosing-bugs&lt;/code&gt; for testing hypotheses before fixing, and handoff for clean session transfers. &lt;/p&gt;

&lt;p&gt;This is still one of the highest-signal skill packs available across coding agents and one of my favourites.&lt;/p&gt;

&lt;p&gt;You can install Matt Pocock Skills using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills@latest add mattpocock/skills
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then run &lt;code&gt;/setup-matt-pocock-skills&lt;/code&gt; once inside Grok Build so it learns your project conventions.&lt;/p&gt;

&lt;p&gt;Related: &lt;a href="https://composio.dev/content/top-codex-skills" rel="noopener noreferrer"&gt;Top Codex Skills&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  3. Caveman: Cut token waste with terse, high-signal agent responses.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F44wc63rhsrf1jfh1fd1f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F44wc63rhsrf1jfh1fd1f.png" alt="Caveman" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Caveman is an answer for those who want to reduce token waste on long sessions. It forces Grok to drop the unnecessary politeness and over-explanation that usually appears in agent responses like Claude Code, Codex, and so on. &lt;/p&gt;

&lt;p&gt;In practice, it can cut output length by roughly 65% on average (ranging from ~22–87% depending on the task) while keeping the useful content intact. It's a small skill but one worth keeping enabled almost all the time.&lt;/p&gt;

&lt;p&gt;You can install Caveman using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills@latest add JuliusBrussee/skills &lt;span class="nt"&gt;--skill&lt;/span&gt; caveman
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once installed, toggle it with:  &lt;code&gt;/caveman&lt;/code&gt; or by saying "talk like caveman"; turn it off with "normal mode." &lt;/p&gt;

&lt;p&gt;Companion commands include &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;/caveman-commit&lt;/code&gt; (terse commit messages),&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/caveman-review&lt;/code&gt; (one-line PR comments), and&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/caveman-stats&lt;/code&gt; (session savings).&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  4. Whathappened: Turn real-time X conversations into structured briefings.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx3c7qtfsfz2i3hbc31ml.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx3c7qtfsfz2i3hbc31ml.png" alt="WhatHappened" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you are X-savvy or want real-time information about X in a structured way, what happened is the answer.&lt;/p&gt;

&lt;p&gt;WhatHappened&amp;nbsp;turns real-time X data into clean, structured briefings rather than raw noise, summarises what happened, maps public opinion, surfaces live debates, and pulls key receipts,&amp;nbsp;all using Grok’s built-in X tools. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note : This skill only works properly inside Grok Build.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You can install whathappened using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add kunchenguid/whathappened
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and then use: &lt;code&gt;/whathappened &amp;lt;query&amp;gt;&lt;/code&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  5. XActions Skills: Scrape, monitor, and automate X workflows without the official API.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb4n3yxcqhjmyjy72ioay.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb4n3yxcqhjmyjy72ioay.png" alt="XActions" width="799" height="359"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;XActions takes whathappend ability to the next level by doing more than just reading X. It packages scraping, monitoring, and automation capabilities into ready-to-use agent skills.&lt;/p&gt;

&lt;p&gt;You can scrape profiles, followers, threads, monitor accounts, or run simple automation tasks without needing the official X API. This is one of the cleaner ways people have begun to package Grok’s X advantage.&lt;/p&gt;

&lt;p&gt;You can install XActions Skills by cloning the repository and placing the skills you want under your skills folder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone &amp;lt;https://github.com/nirholas/XActions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then copy the relevant skill folders into ~/.grok/skills/ or your project’s .grok/skills/.&lt;/p&gt;




&lt;h3&gt;
  
  
  6. Wangnov/grok-skills: Combine web/X research with image, video, and ffmpeg workflows.
&lt;/h3&gt;

&lt;p&gt;Haven’t tried it yet, but on my to-do list. &lt;/p&gt;

&lt;p&gt;&lt;code&gt;Wangnov/grok-skills&lt;/code&gt; is especially useful when your workflow needs both research and media generation. It combines web and X research with image generation, video generation, and basic ffmpeg post-processing.&lt;/p&gt;

&lt;p&gt;Everything runs through a logged-in Grok session, so you avoid extra API costs for media tasks. It’s a practical all-in-one skill for research-plus-assets work.&lt;/p&gt;

&lt;p&gt;You can install Wangnov/grok-skills using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add Wangnov/grok-skills
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can learn more at: &lt;a href="https://github.com/Wangnov/grok-skills" rel="noopener noreferrer"&gt;https://github.com/Wangnov/grok-skills&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  7. Agentic-Code-Review &amp;amp; Repo-Health-Check: Review code and understand unfamiliar repos safely.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkfbp4t9k4lcl0t6m6q56.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkfbp4t9k4lcl0t6m6q56.png" alt="Agentic Code Review" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These two skills from the &lt;code&gt;awesome-grok-build&lt;/code&gt; starter kits are especially useful for properly reviewing code and getting oriented in unfamiliar repositories.&lt;/p&gt;

&lt;p&gt;agentic-code-review focuses on correctness, security, tests and regression risk. repo-health-check helps you quickly understand a new codebase and propose the smallest, safe-first change. Both work well with Plan Mode.&lt;/p&gt;

&lt;p&gt;You can install them by cloning the community kit and copying the skill folders:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone &amp;lt;https://github.com/DominikTobureto/awesome-grok-build&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then place the skill folders into your &lt;code&gt;.grok/skills/&lt;/code&gt; directory. Learn more at: &lt;a href="https://github.com/DominikTobureto/awesome-grok-build" rel="noopener noreferrer"&gt;https://github.com/DominikTobureto/awesome-grok-build&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  ### 8. GodotPrompter: Give Grok better Godot, GDScript, scenes, and signal context.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft63d5cgphmbnik831vel.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft63d5cgphmbnik831vel.png" alt="GoDotPrompter" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you are building games with Godot and want Grok Build to understand Godot-specific patterns, project structure, and common workflows.&lt;/p&gt;

&lt;p&gt;It gives the Grok/Cursor agent better context around GDScript, scenes, signals, and Godot conventions so the suggestions stay more accurate and less generic. &lt;/p&gt;

&lt;p&gt;This is one of the cleanest game-engine-focused plugins currently available for Grok Build.&lt;/p&gt;

&lt;p&gt;You can install GodotPrompter using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;grok plugin install jame581/GodotPrompter --trust
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then enable it with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;grok plugin enable godot-prompter
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Learn more at: &lt;a href="https://github.com/jame581/GodotPrompter" rel="noopener noreferrer"&gt;https://github.com/jame581/GodotPrompter&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  9. Hyperframes: Create and edit programmatic videos with HTML, CSS, and JavaScript.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh0u2rltq2b9bcmkctkhn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh0u2rltq2b9bcmkctkhn.png" alt="Hyperframes" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hyperframes is for those who want a Grok Build or similar agent to create and edit videos using HTML, CSS, and JavaScript rather than traditional timeline editors.&lt;/p&gt;

&lt;p&gt;It ships a full set of agent skills that teach the correct patterns for planning compositions, writing valid HyperFrames HTML, adding animations, linting, previewing and rendering. &lt;/p&gt;

&lt;p&gt;The main entry skill is &lt;code&gt;/hyperframes&lt;/code&gt;, which routes “make me a video” requests to the right workflow.&lt;/p&gt;

&lt;p&gt;You can install Hyperframes skills using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add heygen-com/hyperframes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and use it with &lt;code&gt;/hyperframes&lt;/code&gt;  in grok build. Learn more at: &lt;a href="https://hyperframes.heygen.com/" rel="noopener noreferrer"&gt;https://hyperframes.heygen.com/&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  10. Remotion Skills: Build production-ready motion graphics and videos with React and TypeScript.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6mankrz86yy6exayq2mv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6mankrz86yy6exayq2mv.png" alt="Remotion" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Remotion is similar to Hyperframes in that it allows you to create motion graphics and programmatic videos using React and TypeScript.&lt;/p&gt;

&lt;p&gt;It teaches the agent Remotion best practices like compositions, animations, sequencing, rendering and project structure &lt;/p&gt;

&lt;p&gt;This makes the output clean and production-ready, rather than generic React code that happens to render video.&lt;/p&gt;

&lt;p&gt;You can install Remotion Skills using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;npx skills add remotion-dev/skills
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can learn more at: &lt;a href="https://www.remotion.dev/docs/ai/skills" rel="noopener noreferrer"&gt;https://www.remotion.dev/docs/ai/skills&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: Remotion skill work best inside an existing Remotion project, so better first create it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Related: &lt;a href="https://composio.dev/content/top-design-skills" rel="noopener noreferrer"&gt;Top Design Skills&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Top Plugins to Use with Grok Build
&lt;/h2&gt;

&lt;p&gt;Skills are great; some even perform tasks, but for a seamless experience, plugins are mandatory. These are the ones that still reside directly in my skills.&lt;/p&gt;

&lt;h3&gt;
  
  
  11. Firecrawl: Search, scrape, crawl, and extract clean data from websites.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fokkzf5n1gf1p08bz2w29.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fokkzf5n1gf1p08bz2w29.png" alt="Firecrawl" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Firecrawl is especially useful for research, document retrieval, competitive analysis, and data collection from websites.&lt;/p&gt;

&lt;p&gt;It gives Grok Build live access to the web through search, scraping, crawling, website mapping, structured extraction, and browser interaction. &lt;/p&gt;

&lt;p&gt;It can render JavaScript-heavy pages, handle common anti-bot restrictions, and return content as clean Markdown or structured data.&lt;/p&gt;

&lt;p&gt;You can install Firecrawl using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/marketplace
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then search for &lt;code&gt;firecrawl&lt;/code&gt; and press &lt;code&gt;i&lt;/code&gt; to install it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: Firecrawl may request authentication when you first use its hosted MCP server.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  12. Superpowers: Add structured engineering workflows for planning, TDD, and debugging.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F05b5x39z0eui2zu58bpg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F05b5x39z0eui2zu58bpg.png" alt="Superpowers" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Superpowers is for those who want their agent to plan carefully, validate its work, and follow a more disciplined development process instead of immediately generating code.&lt;/p&gt;

&lt;p&gt;It adds structured software engineering workflows to Grok Build, and the current implementation includes test-driven development, systematic debugging, collaboration patterns, and repeatable engineering processes.&lt;/p&gt;

&lt;p&gt;I personally use this before switching to Matt Pocock's skills.&lt;/p&gt;

&lt;p&gt;You can install it with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;grok plugin install superpowers --trust
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note: Review the plugin before using --trust, since that option skips the interactive trust prompt.&lt;/p&gt;




&lt;h3&gt;
  
  
  13. Exa: Get fast, high-quality agent-oriented web search results.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmuf5phugkpywbnvg2kmp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmuf5phugkpywbnvg2kmp.png" alt="EXA" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Exa provides fast, high-quality agent-oriented search. It works particularly well as a complement to Firecrawl.&lt;/p&gt;

&lt;p&gt;Exa's API is purpose-built for LLMs, so results come back structured and filtered rather than cluttered with ads and navigation, fast enough for an agent to search mid-task without breaking flow. &lt;/p&gt;

&lt;p&gt;For teams, that means less time and token spend per lookup, so search-heavy steps (competitor checks, source verification, quick fact lookups) stop being a bottleneck inside the coding session itself&lt;/p&gt;

&lt;p&gt;Use Exa for quick, accurate retrieval, and switch to Firecrawl for full-page scraping or site crawling.&lt;/p&gt;

&lt;p&gt;You can install the Exa plugin using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/plugin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then search for &lt;code&gt;exa&lt;/code&gt; and press &lt;code&gt;i&lt;/code&gt;to install it.&lt;/p&gt;




&lt;h3&gt;
  
  
  14. Vercel: Let Grok manage deployments, environment variables, logs, and domains.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvkz9b384mj9wslorypn1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvkz9b384mj9wslorypn1.png" alt="Vercel" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Vercel is for those who want their agents to deploy their projects on the Vercel platform. It gives Grok direct control over deployments, environment variables, build logs, and domains.&lt;/p&gt;

&lt;p&gt;It also keeps Grok aware of current Vercel features, which reduces outdated suggestions.&lt;/p&gt;

&lt;p&gt;For teams, this pairing of live platform knowledge with real deploy/env/log access leads to fewer review cycles spent catching agent suggestions that no longer reflect how Vercel actually works. &lt;/p&gt;

&lt;p&gt;This fixes a common issue with agent integrations: context drift.&lt;/p&gt;

&lt;p&gt;You can install the Vercel plugin using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/plugin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then search for &lt;code&gt;vercel&lt;/code&gt; and install it using &lt;code&gt;i&lt;/code&gt; .&lt;/p&gt;




&lt;h3&gt;
  
  
  15. Cloudflare: Build and deploy Workers, Durable Objects, and edge apps with platform-aware guidance.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2p5fg7rhbddg31w9fklh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2p5fg7rhbddg31w9fklh.png" alt="Cloudflare" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cloudflare is for those scenarios when part of your stack lives on the other side of the world. It provides skills for Workers, Durable Objects, and related Cloudflare tooling.&lt;/p&gt;

&lt;p&gt;The plugin covers the entire Cloudflare developer platform: Workers, Durable Objects, the Agents SDK, MCP servers, Wrangler CLI, and web performance, functioning as a skill library that maps Cloudflare concepts directly to prompts.&lt;/p&gt;

&lt;p&gt;This makes scaffolding and deploying edge projects noticeably smoother inside Grok Build for business and working with it easier.&lt;/p&gt;

&lt;p&gt;You can install Cloudflare skills or plugins using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/plugin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then search for &lt;code&gt;cloudflare&lt;/code&gt; and install it using &lt;code&gt;i&lt;/code&gt; &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: Don’t get confused by the name, if you go to official repo , its given as a plugins.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  16. Chrome DevTools: Debug frontend issues through live browser inspection, traces, and network data.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwwfnd6ja836poqc6sys8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwwfnd6ja836poqc6sys8.png" alt="Chrome Dev Tools" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Chrome DevTools is useful for frontend debugging and performance work. It lets Grok control a live browser session. (not on WSL)&lt;/p&gt;

&lt;p&gt;You can record performance traces, inspect network requests, evaluate JavaScript, and take DOM snapshots without leaving the agent workflow.&lt;/p&gt;

&lt;p&gt;Under the hood, this runs on the official Chrome DevTools MCP server, which lets a coding agent control and inspect a live Chrome browser and acts as a Model Context Protocol server, giving the assistant access to the full power of Chrome DevTools for reliable automation, in-depth debugging, and performance analysis.&lt;/p&gt;

&lt;p&gt;This means that instead of an engineer manually opening DevTools, reproducing the issue, and reporting back what they saw, Grok can drive the same browser session directly and return a trace, a failing request, or a DOM state as evidence.&lt;/p&gt;

&lt;p&gt;You can install the Chrome DevTools plugin using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/plugin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then search for &lt;code&gt;chrome-devtools&lt;/code&gt; and install it using &lt;code&gt;i&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Related: &lt;a href="https://composio.dev/content/top-claude-code-plugins" rel="noopener noreferrer"&gt;Top Claude Code Plugins&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  17. Sentry: Pull production errors and stack traces into Grok for faster fixes.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyyqplpwldngo1f3tan1l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyyqplpwldngo1f3tan1l.png" alt="Sentry" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Sentry closes the loop between local development and production. &lt;/p&gt;

&lt;p&gt;It lets Grok pull real error data and stack traces. Combined with Seer-powered analysis, it helps turn production issues into concrete fixes faster.&lt;/p&gt;

&lt;p&gt;This works through Sentry's official Grok plugin, which connects Grok to Sentry via the Sentry MCP server, providing real production issue-debugging context, code review with Sentry data, and monitoring configuration- on top of SDK setup for any platform.&lt;/p&gt;

&lt;p&gt;This means less engineer time spent context-switching between logs, code, and chat to reconstruct what broke, and a shorter gap between an alert firing and a fix landing.&lt;/p&gt;

&lt;p&gt;You can install the Sentry plugin using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/plugin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then search for &lt;code&gt;sentry&lt;/code&gt; and install it. Part of the official Grok Build Marketplace.&lt;/p&gt;




&lt;h3&gt;
  
  
  18. Unity MCP + CLI: Drive the Unity editor and iterate on game projects from Grok.
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2434044nlqzo01whd8y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2434044nlqzo01whd8y.png" alt=" Unity MCP " width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you are like me and like to build games in Unity and want Grok Build to actually drive the editor and project instead of just writing C# in isolation.&lt;/p&gt;

&lt;p&gt;People are already using this combination to let Grok open scenes, modify GameObjects, work with the Asset Store, and iterate on playable prototypes much faster. It is currently one of the most practical ways to pair Grok Build with Unity.&lt;/p&gt;

&lt;p&gt;You can set it up by installing the Unity MCP server and connecting it through Grok’s MCP system, then pairing it with a simple Unity-focused skill that teaches the agent your project conventions.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. In Unity, open Package Manager → Add package from git URL
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/CoderGamester/mcp-unity.git" rel="noopener noreferrer"&gt;https://github.com/CoderGamester/mcp-unity.git&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After the package is installed, open Tools → MCP Unity → Server Window and click Force Install Server.&lt;/p&gt;

&lt;p&gt;Then connect it to Grok Build with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;grok mcp add unity
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note: There is no single official marketplace plugin yet. Most people combine the Unity MCP with grok build and a lightweight custom skill for best results.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Skills give Grok better judgment and consistency. Plugins give Grok real tools and reach, but don’t install them all at once.&lt;/p&gt;

&lt;p&gt;Start with a solid foundation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cross-app workflows: Composio CLI or MCP&lt;/li&gt;
&lt;li&gt;Engineering discipline → Matt Pocock Skills + Superpowers&lt;/li&gt;
&lt;li&gt;Token control → Caveman&lt;/li&gt;
&lt;li&gt;X advantage → what happened or XActions&lt;/li&gt;
&lt;li&gt;Web power → Firecrawl + Exa&lt;/li&gt;
&lt;li&gt;Game development → GodotPrompter or Unity MCP setup&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then add the platform plugins that match the work you actually do (Vercel, Cloudflare, Sentry, etc.).&lt;/p&gt;

&lt;p&gt;Use both skills and plugins, but do it with a clear understanding of the use case. Only then does Grok Build start to feel like a real system instead of just another coding agent.&lt;/p&gt;





&lt;p&gt;&lt;/p&gt;&lt;br&gt;
  Sources&lt;br&gt;
  &lt;ul&gt;

&lt;li&gt;Official docs: &lt;a href="https://docs.x.ai/build/features/skills-plugins-marketplaces" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;a href="https://docs.x.ai/build/features/skills-plugins-marketplaces" rel="noopener noreferrer"&gt;https://docs.x.ai/build/features/skills-plugins-marketplaces&lt;/a&gt;
&lt;/li&gt;

&lt;li&gt;Plugin Marketplace announcement: &lt;a href="https://x.ai/news/grok-plugin-marketplace" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;a href="https://x.ai/news/grok-plugin-marketplace" rel="noopener noreferrer"&gt;https://x.ai/news/grok-plugin-marketplace&lt;/a&gt;
&lt;/li&gt;

&lt;li&gt;Firecrawl roundup: &lt;a href="https://www.firecrawl.dev/blog/best-grok-plugins" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;a href="https://www.firecrawl.dev/blog/best-grok-plugins" rel="noopener noreferrer"&gt;https://www.firecrawl.dev/blog/best-grok-plugins&lt;/a&gt;
&lt;/li&gt;

&lt;li&gt;Matt Pocock Skills: &lt;a href="https://github.com/mattpocock/skills" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;a href="https://github.com/mattpocock/skills" rel="noopener noreferrer"&gt;https://github.com/mattpocock/skills&lt;/a&gt;
&lt;/li&gt;

&lt;li&gt;Superpowers: &lt;a href="https://github.com/obra/superpowers" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;a href="https://github.com/obra/superpowers" rel="noopener noreferrer"&gt;https://github.com/obra/superpowers&lt;/a&gt;
&lt;/li&gt;

&lt;li&gt;whathappened: &lt;a href="https://github.com/kunchenguid/whathappened" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;a href="https://github.com/kunchenguid/whathappened" rel="noopener noreferrer"&gt;https://github.com/kunchenguid/whathappened&lt;/a&gt;
&lt;/li&gt;

&lt;li&gt;XActions: &lt;a href="https://github.com/nirholas/XActions" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;a href="https://github.com/nirholas/XActions" rel="noopener noreferrer"&gt;https://github.com/nirholas/XActions&lt;/a&gt;
&lt;/li&gt;

&lt;li&gt;Wangnov/grok-skills: &lt;a href="https://github.com/Wangnov/grok-skills" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;a href="https://github.com/Wangnov/grok-skills" rel="noopener noreferrer"&gt;https://github.com/Wangnov/grok-skills&lt;/a&gt;
&lt;/li&gt;

&lt;li&gt;awesome-grok-build: &lt;a href="https://github.com/DominikTobureto/awesome-grok-build" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;a href="https://github.com/DominikTobureto/awesome-grok-build" rel="noopener noreferrer"&gt;https://github.com/DominikTobureto/awesome-grok-build&lt;/a&gt;
&lt;/li&gt;

&lt;/ul&gt;
&lt;br&gt;
&lt;br&gt;
&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>What Actually Is an MCP Gateway?</title>
      <dc:creator>Sunil Kumar Dash</dc:creator>
      <pubDate>Tue, 28 Jul 2026 14:21:38 +0000</pubDate>
      <link>https://dev.to/composiodev/what-actually-is-an-mcp-gateway-37aa</link>
      <guid>https://dev.to/composiodev/what-actually-is-an-mcp-gateway-37aa</guid>
      <description>&lt;p&gt;Every team that connects agents to real tools hits the same wall at roughly the same time. It usually looks like a Slack message: &lt;em&gt;"Who has the Jira token?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Here's what's underneath that, and what a gateway does about it.&lt;/p&gt;




&lt;h2&gt;
  
  
  The N×M problem
&lt;/h2&gt;

&lt;p&gt;You have N agents and M tools. Connect them directly, and you get N×M integrations. Each one carries its own credentials, its own auth flow, its own error handling, its own version drift.&lt;/p&gt;

&lt;p&gt;Three agents and four tools is twelve connections. Ten agents and twenty tools is two hundred. Nobody plans for that number — you arrive at it one integration at a time.&lt;/p&gt;

&lt;p&gt;A gateway collapses it to N+M. Agents connect once to the gateway. The gateway connects once to each tool.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Before:  agent ──┬──&amp;gt; GitHub
                 ├──&amp;gt; Slack
                 └──&amp;gt; Jira        (× every agent)

After:   agent ──&amp;gt; gateway ──┬──&amp;gt; GitHub
                             ├──&amp;gt; Slack
                             └──&amp;gt; Jira
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  What an &lt;a href="https://composio.dev/mcp-gateway" rel="noopener noreferrer"&gt;MCP gateway&lt;/a&gt; actually does
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Routing and aggregation&lt;/strong&gt; — one endpoint fronting many MCP servers, with tool filtering so agents don't blow past context limits&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authentication&lt;/strong&gt; — holds credentials centrally, runs OAuth flows, passes through per-user identity&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authorisation&lt;/strong&gt; — who can call which tool; RBAC, allowlists, blocking destructive actions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit&lt;/strong&gt; — logs every call: user, tool, action, outcome, including denied ones&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Threat handling&lt;/strong&gt; — tool poisoning, rug-pull updates, cross-server shadowing, prompt injection via tool descriptions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reliability&lt;/strong&gt; — rate limits, retries, timeouts, absorbing schema drift when upstream APIs change&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The threat handling deserves a note, because it's genuinely new. Tool descriptions are input the agent trusts. A server can change its tool definitions after you've approved it. Generic API security doesn't cover this, and most gateways describe their handling of it vaguely.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwth4nydhu456f2gwj5tn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwth4nydhu456f2gwj5tn.png" alt="Whats MCP gateway" width="800" height="530"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Gateway ≠ server ≠ client
&lt;/h2&gt;

&lt;p&gt;Constantly confused, so:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Client&lt;/strong&gt; — the agent (Claude, Cursor, ChatGPT)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server&lt;/strong&gt; — the thing exposing tools over MCP&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gateway&lt;/strong&gt; — sits between them, governs the traffic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And it's not an API gateway. An API gateway routes HTTP between services. An MCP gateway is protocol-aware — it understands tools and tool-call semantics, which is what lets it enforce per-tool policy.&lt;/p&gt;




&lt;h2&gt;
  
  
  The landscape splits four ways
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2xvk8ul8b5anqvfbh8u3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2xvk8ul8b5anqvfbh8u3.png" alt="MCP Gateway market map" width="800" height="568"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Purpose-built managed
&lt;/h3&gt;

&lt;p&gt;Built for MCP from the start. They differ mainly in whether they also supply the tools.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Composio &lt;a href="https://composio.dev/mcp-gateway" rel="noopener noreferrer"&gt;MCP gateway&lt;/a&gt;&lt;/strong&gt; — 1,000+ managed integrations behind per-team scoped endpoints. SCIM, action-level blocking, exportable audit. Managed, VPC, or self-hosted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TrueFoundry&lt;/strong&gt; — one control plane for LLM and MCP traffic. Publishes &amp;lt;5ms p95 overhead. Bring your own servers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lunar MCPX&lt;/strong&gt; — granular RBAC, immutable audit, centralised secrets. ~4ms p99. Open source core.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MintMCP&lt;/strong&gt; — governance-first. SOC 2 / HIPAA / GDPR log formats, SCIM-driven bundles, per-agent identity. BYO servers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operant AI&lt;/strong&gt; — runtime security rather than routing. Scans servers, maps threats to OWASP, catches shadow MCP usage on dev machines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;StackOne&lt;/strong&gt; — per-user OAuth into HRIS/ERP/CRM, plus meta-tools that keep you under client tool-count caps.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  Open source / self-hosted
&lt;/h3&gt;

&lt;p&gt;You run it, you own it. No licence cost, real operational cost.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Obot&lt;/strong&gt; — multi-role RBAC, curated server catalogue, composite servers. Names MCP-specific threats explicitly, which most don't. Kubernetes or Docker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docker MCP Gateway&lt;/strong&gt; — one container per server with signed images and resource limits. Security via isolation rather than policy. 50–200ms overhead, and no built-in RBAC or audit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Microsoft MCP Gateway&lt;/strong&gt; — MIT, Kubernetes-native, session-aware stateful routing. Genuinely useful plumbing; tightly coupled to AKS in practice. Not to be confused with Agent 365, which is a different product with a similar name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IBM ContextForge&lt;/strong&gt; — Apache 2.0. Federates MCP, A2A, REST and gRPC through one endpoint, 40+ plugins, OTel tracing. Broadest scope here; also the heaviest lift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCPJungle&lt;/strong&gt; — gateway and registry in one lightweight package. Basic RBAC, minimal ceremony.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lasso&lt;/strong&gt; — open-source, security-first. Server reputation scanning and PII redaction, at 100–250ms.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  API gateways extended
&lt;/h3&gt;

&lt;p&gt;If you already run one, MCP becomes another middleware in a chain you understand.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Kong&lt;/strong&gt; — Agent Gateway (3.14, April 2026) covers LLM, MCP and A2A on one runtime. Autogenerates MCP tools from existing REST endpoints, which is the shortest migration path if your services are already APIs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Traefik Hub&lt;/strong&gt; — MCP as middleware, acting as an OAuth 2.1 resource server. Task-Based Access Control lets policies key on amounts, time windows and record types, not just tool names.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Portkey&lt;/strong&gt; — MCP registry alongside its LLM gateway. Remote HTTP servers only; local stdio needs wrapping.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  Automation platforms extended
&lt;/h3&gt;

&lt;p&gt;Enormous catalogues, thinner governance. Fastest route to a working prototype.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Zapier&lt;/strong&gt; — 9,000+ apps, 30,000+ actions, browser-based setup. Task-based pricing gets unpredictable once an agent is driving.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Workato&lt;/strong&gt; — 12,000+ enterprise connectors, existing recipes exposed over MCP. Compelling if you're already on it, hard to justify buying for MCP alone.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a longer and detailed read, check out: &lt;a href="https://composio.dev/content/best-mcp-gateway-for-developers" rel="noopener noreferrer"&gt;Best MCP Gateway&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The trade-off nobody states plainly
&lt;/h2&gt;

&lt;p&gt;Most gateways are strong on &lt;strong&gt;governance&lt;/strong&gt; or strong on &lt;strong&gt;integration breadth&lt;/strong&gt;. Rarely both.&lt;/p&gt;

&lt;p&gt;Governance-first products — MintMCP, Lunar, Obot — do RBAC and audit well and supply zero connectors. Bring your own servers. That's a real ongoing cost: OAuth setup, schema maintenance, security review, per tool, forever.&lt;/p&gt;

&lt;p&gt;Breadth-first products — Zapier, Workato — hand you thousands of integrations and much less control over who calls what.&lt;/p&gt;

&lt;p&gt;Work out which of those is your actual constraint before you shortlist anything.&lt;/p&gt;




&lt;h2&gt;
  
  
  Choosing, briefly
&lt;/h2&gt;

&lt;p&gt;In rough order, because the early ones eliminate options:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Deployment&lt;/strong&gt; — managed SaaS, self-hosted, or VPC. Data residency rules kill whole categories before anything else matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compliance&lt;/strong&gt; — SOC 2, ISO, HIPAA if relevant. Then check the audit trail actually exports.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Identity&lt;/strong&gt; — SSO and SCIM, or you're provisioning access by hand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integration depth&lt;/strong&gt; — do they supply connectors, or do you?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pricing model&lt;/strong&gt; — agents are chatty. Per-task pricing that's fine for human-triggered automation gets weird when an agent fans one request into forty tool calls. Model your volume first.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Latency comes lower than you'd think. TrueFoundry publishes sub-5ms, Lunar around 4ms p99, Docker 50–200ms. Real differences — but a gateway that saves 3ms and costs six months of integration work is a bad trade for most teams.&lt;/p&gt;




&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;If you're prototyping, grab whatever's fastest to wire up and move on.&lt;/p&gt;

&lt;p&gt;If you're going to production, the question isn't which gateway is best. It's whether connector maintenance or governance is the thing that'll actually bite you — and then picking the one that solves that, knowing you'll compromise on the other.&lt;/p&gt;

&lt;p&gt;Check out &lt;a href="https://composio.dev/mcp-gateway" rel="noopener noreferrer"&gt;https://composio.dev/mcp-gateway&lt;/a&gt; for building agents with secure and auditable access to tools&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>security</category>
      <category>automation</category>
    </item>
    <item>
      <title>Kimi K3 vs GLM-5.2: What a practical test between 2 taught me</title>
      <dc:creator>Developer Harsh</dc:creator>
      <pubDate>Tue, 28 Jul 2026 13:36:28 +0000</pubDate>
      <link>https://dev.to/composiodev/kimi-k3-vs-glm-52-what-a-practical-test-between-2-taught-me-1ii8</link>
      <guid>https://dev.to/composiodev/kimi-k3-vs-glm-52-what-a-practical-test-between-2-taught-me-1ii8</guid>
      <description>&lt;p&gt;Moonshot AI released Kimi K3 on July 16, 2026, and it landed with a statement: 2.8 trillion parameters, 1 million token context, and open weights by July 27. The previous month, GLM-5.2 was released and is already in production&lt;/p&gt;

&lt;p&gt;This proves that open-source models aren't just catching up to closed ones; they're reshaping what developers and businesses expect.&lt;/p&gt;

&lt;p&gt;This is a comparison built for people who are &lt;em&gt;building things&lt;/em&gt;. No benchmark chasing. No marketing narratives. &lt;/p&gt;

&lt;p&gt;Just what each model does, where it shines, and what matters when you're shipping.&lt;/p&gt;

&lt;h3&gt;
  
  
  TLDR
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Kimi K3&lt;/strong&gt;: 2.8T params, 1M context, always-on reasoning, native multimodal. Best for long agent loops that need sustained reasoning and visual understanding. Frontier pricing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GLM-5.2&lt;/strong&gt;: 744B params (40B active via MoE), 1M context, flexible reasoning effort. Best for coding, math, and cost-efficient throughput. Open weights (MIT) available now.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pick K3&lt;/strong&gt; if: agents need to reason for hours, handle images/UI, cost isn't the bottleneck.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pick GLM-5.2&lt;/strong&gt; if: you want open weights today, need cheap high-volume inference, or your work leans coding/math.&lt;/li&gt;
&lt;li&gt;Few personal builds like games, physics-driven simulation, coding, and behavioral tasks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bottom line&lt;/strong&gt;: Both close the gap with closed models fast. Choice comes down to workload, not hype.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Architecture: Two Very Different Paths to Scale
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Kimi K3&lt;/th&gt;
&lt;th&gt;GLM-5.2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total Parameters&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2.8 trillion&lt;/td&gt;
&lt;td&gt;744 billion total&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Active Parameters&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not yet disclosed&lt;/td&gt;
&lt;td&gt;~40 billion active (MoE)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context Window&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1 million tokens&lt;/td&gt;
&lt;td&gt;1 million tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Key Architecture&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Kimi Delta Attention (KDA) + Attention Residuals&lt;/td&gt;
&lt;td&gt;Mixture-of-Experts (MoE)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;License&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not yet published&lt;/td&gt;
&lt;td&gt;MIT (no regional limits)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Release Date&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;July 16, 2026&lt;/td&gt;
&lt;td&gt;June 13, 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Kimi K3: Raw Scale Meets Long-Horizon Design with KDA
&lt;/h3&gt;

&lt;p&gt;Kimi K3 is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals. Moonshot engineered this for &lt;em&gt;sustained&lt;/em&gt; agent workloads, not just bigger benchmarks.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe6rb40x0g554aroo886l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe6rb40x0g554aroo886l.png" alt="Kimi K3 Architecture" width="800" height="748"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The move from K2's 1 trillion parameters to K3's 2.8 trillion is deliberate. Moonshot is charging frontier rates to make the price-to-capability tradeoff hard to ignore. &lt;/p&gt;

&lt;p&gt;You're not getting a discount model trying to go above its weight. Instead, you're getting brute-force capability with specialised attention for long reasoning chains.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: The full model weights will be released by July 27, 2026, but the technical report with full sparsity ratios and active parameter counts is still pending. You can build on K3 API today, but deep architectural details aren't locked in yet.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  GLM-5.2: Efficiency First, Capability Everywhere with MOE &amp;amp; Index Share
&lt;/h3&gt;

&lt;p&gt;GLM-5.2 is a 744-billion-parameter Mixture-of-Experts model with approximately 40 billion active parameters per token. &lt;/p&gt;

&lt;p&gt;That MoE design means only a fraction of the model activates per token.  This enables throughput that larger dense models can't match.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwuaf0rsoatwfwne9gw6q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwuaf0rsoatwfwne9gw6q.png" alt="GLM 5.2 Architecture" width="799" height="514"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The standout innovation of GLM is its IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. For you, this means GLM-5.2 makes the 1M context practical, not theoretical.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8xhaxkmee51pq5g7cp0w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8xhaxkmee51pq5g7cp0w.png" alt="Index Share" width="800" height="483"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note :  GLM-5.2 is released under an unrestricted MIT license, which matters if you're deploying locally or need no-strings-attached weights. Deploy on your own hardware, fine-tune, fork with no regional restrictions.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Kimi K3 vs GLM 5.2: Composio Golden Eval
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What does the bench look like?
&lt;/h3&gt;

&lt;p&gt;Composio Golden Eval is a real-account tool-use benchmark. Claude Code drives multi-step SaaS tasks through the hosted Composio MCP router against live accounts, then the final state is checked by reading the actual account back through APIs.&lt;/p&gt;

&lt;p&gt;The grader checks what changed in the account, which is why I care. The verifier reads the account state: labels, Sheets rows, Salesforce or HubSpot records, calendar edits, access lists. There is no transcript-only judgment that the agent basically got it.&lt;/p&gt;

&lt;p&gt;The accounts stay safe because writes are tag-scoped and cleaned up afterwards. Every run leaves tagged artefacts that can be removed after grading, which is the only sane way to run this kind of thing on live Gmail, Google Calendar, Google Drive, Google Sheets, Salesforce, HubSpot, GitHub, Linear, and Slack accounts.&lt;/p&gt;

&lt;p&gt;The run covered 12 scenarios across 24 trials. Seven are historical cases that a competent tool-use model should clear: CRM identity dedup, calendar free/busy checks, recurring-event repair, Drive external-share audits, Gmail label batches, GitHub access audits, and GitHub/Linear reconciliation. &lt;/p&gt;

&lt;p&gt;Five are the harder frontier-kill stress cases: cross-app “sync and reconcile” workflows where the agent has to read Gmail threads, apply exclusion rules, append exact rows to a Sheet ledger, send per-item replies, and write one ops-thread tally.&lt;/p&gt;

&lt;p&gt;The pass condition is an exact final state. These tasks mix exact-set reconciliation, dedup, cross-app joins, and audits with negative constraints. If the agent gets 90% of the rows right but includes one disqualified item, the run still fails. Failed runs can show partial-credit check counts like 8/13, but the outcome metric is still pass, fail, or DNF (did not finish).&lt;/p&gt;

&lt;h3&gt;
  
  
  How I ran it
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyy5l5dda1odrft3luhr0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyy5l5dda1odrft3luhr0.png" alt="Eval Chain" width="800" height="605"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I ran both models through the same 12 Golden Eval cases against the same live accounts. The readback checks were identical, with the same check names and denominators. There is one harness caveat: Kimi K3 ran under the pi agent harness, while GLM 5.2 ran under Claude Code pointed at OpenRouter. I treat the pass/fail result as a same-task, same-grader comparison, and effort numbers as harness-dependent.&lt;/p&gt;

&lt;p&gt;Grading used real-account API readback, tag-scoped cleanup, pass/fail/dnf per trial, with partial-credit check counts on failures.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Frontier-kill task&lt;/th&gt;
&lt;th&gt;Kimi K3 checks&lt;/th&gt;
&lt;th&gt;GLM 5.2 checks&lt;/th&gt;
&lt;th&gt;Kimi tool calls&lt;/th&gt;
&lt;th&gt;GLM tool calls&lt;/th&gt;
&lt;th&gt;Kimi runtime tokens&lt;/th&gt;
&lt;th&gt;GLM input tokens&lt;/th&gt;
&lt;th&gt;Kimi agent time&lt;/th&gt;
&lt;th&gt;GLM agent time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Invoice sync&lt;/td&gt;
&lt;td&gt;8/13&lt;/td&gt;
&lt;td&gt;8/13&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;896,094&lt;/td&gt;
&lt;td&gt;1,020,844&lt;/td&gt;
&lt;td&gt;389.8s&lt;/td&gt;
&lt;td&gt;507.4s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refund ledger&lt;/td&gt;
&lt;td&gt;8/13&lt;/td&gt;
&lt;td&gt;8/13&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;1,309,483&lt;/td&gt;
&lt;td&gt;1,334,127&lt;/td&gt;
&lt;td&gt;685.8s&lt;/td&gt;
&lt;td&gt;458.9s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Roster sync&lt;/td&gt;
&lt;td&gt;8/13&lt;/td&gt;
&lt;td&gt;7/13&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;609,233&lt;/td&gt;
&lt;td&gt;1,058,024&lt;/td&gt;
&lt;td&gt;505.1s&lt;/td&gt;
&lt;td&gt;418.9s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vendor directory&lt;/td&gt;
&lt;td&gt;11/13&lt;/td&gt;
&lt;td&gt;11/13&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;td&gt;820,613&lt;/td&gt;
&lt;td&gt;1,683,579&lt;/td&gt;
&lt;td&gt;788.5s&lt;/td&gt;
&lt;td&gt;803.7s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ticket sync&lt;/td&gt;
&lt;td&gt;17/24&lt;/td&gt;
&lt;td&gt;dnf&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;td&gt;dnf&lt;/td&gt;
&lt;td&gt;1,745,612&lt;/td&gt;
&lt;td&gt;dnf&lt;/td&gt;
&lt;td&gt;713.5s&lt;/td&gt;
&lt;td&gt;dnf&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;Runs used the hosted Composio MCP router. Kimi ran under the pi agent harness; GLM ran under Claude Code pointed at OpenRouter. Kimi token counts are total runtime tokens (input + output); GLM's column is input tokens, with another 18K to 46K output tokens per task. Ticket sync is the one task GLM 5.2 did not finish inside the 30-minute cap.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Pass rate was flat: Kimi K3 solved 7 of 12, and GLM 5.2 solved 7 of 12. That is 58% each.&lt;/p&gt;

&lt;p&gt;With the harness caveat above, effort favored Kimi on several finished frontier-kill runs. &lt;/p&gt;

&lt;p&gt;On the four frontier-kill tasks both models finished, Kimi used fewer tool calls on invoice, roster, and vendor, while GLM used one fewer on refund. Kimi also finished Ticket sync in 713.5s with 23 tool calls and 1,745,612 runtime tokens; GLM hit dnf inside the 30-minute cap. &lt;/p&gt;

&lt;p&gt;Time did not point one way: GLM was faster on refund ledger and roster sync, while Kimi was faster on invoice sync and slightly faster on vendor directory.&lt;/p&gt;




&lt;h3&gt;
  
  
  Findings
&lt;/h3&gt;

&lt;p&gt;Here is the task-for-task result on the same 12 cases.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Band&lt;/th&gt;
&lt;th&gt;Kimi K3&lt;/th&gt;
&lt;th&gt;GLM 5.2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CRM identity dedup&lt;/td&gt;
&lt;td&gt;Historical&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Calendar free/busy&lt;/td&gt;
&lt;td&gt;Historical&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recurring instance repair&lt;/td&gt;
&lt;td&gt;Historical&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Drive external-share audit&lt;/td&gt;
&lt;td&gt;Historical&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gmail label batch&lt;/td&gt;
&lt;td&gt;Historical&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub access audit&lt;/td&gt;
&lt;td&gt;Historical&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub / Linear reconciliation&lt;/td&gt;
&lt;td&gt;Historical&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Invoice sync&lt;/td&gt;
&lt;td&gt;Frontier kill&lt;/td&gt;
&lt;td&gt;❌ 8/13&lt;/td&gt;
&lt;td&gt;❌ 8/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refund ledger&lt;/td&gt;
&lt;td&gt;Frontier kill&lt;/td&gt;
&lt;td&gt;❌ 8/13&lt;/td&gt;
&lt;td&gt;❌ 8/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Roster sync&lt;/td&gt;
&lt;td&gt;Frontier kill&lt;/td&gt;
&lt;td&gt;❌ 8/13&lt;/td&gt;
&lt;td&gt;❌ 7/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ticket sync&lt;/td&gt;
&lt;td&gt;Frontier kill&lt;/td&gt;
&lt;td&gt;❌ 17/24&lt;/td&gt;
&lt;td&gt;❌ dnf&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vendor directory&lt;/td&gt;
&lt;td&gt;Frontier kill&lt;/td&gt;
&lt;td&gt;❌ 11/13&lt;/td&gt;
&lt;td&gt;❌ 11/13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Solved&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7/12&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7/12&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;Fractions on the failed rows are partial-credit verifier checks: how many graded assertions the model got right before missing the exact-final-state bar. &lt;code&gt;dnf&lt;/code&gt; means GLM 5.2 did not finish Ticket sync inside the 30-minute per-task cap, so no partial score was recorded.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Kimi K3 and GLM 5.2 both cleared the historical seven, both fell on all five frontier-kill workflows, and both ended at 7/12, 58%. The only score gap in the entire suite is one verifier check on Roster sync: Kimi got 8/13, GLM got 7/13.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9erh9bpavwich0pqn8s5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9erh9bpavwich0pqn8s5.png" alt="Task Results" width="800" height="641"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The historical band did not separate them. Both passed all seven cleanly, including CRM identity dedup, which spans Salesforce, HubSpot, and Gmail. Calendar free/busy, recurring instance repair, Drive external-share audit, Gmail label batch, GitHub access audit, and GitHub / Linear reconciliation all landed green for both models.&lt;/p&gt;

&lt;h3&gt;
  
  
  What it costs
&lt;/h3&gt;

&lt;p&gt;I priced the finished runs from their token counts against current OpenRouter list rates, before cache discounts. The dollar amounts are estimates, but the price gap is wide enough that the direction is clear.&lt;/p&gt;

&lt;p&gt;OpenRouter currently lists Kimi K3 at $3/M input and $15/M output. (openrouter.ai) For GLM 5.2, I used OpenRouter’s current model API rate of about $0.82/M input and $2.59/M output.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Tokens on the finished frontier-kill cases&lt;/th&gt;
&lt;th&gt;Estimated cost per case&lt;/th&gt;
&lt;th&gt;Four-case estimate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;~609K to 1.75M runtime tokens&lt;/td&gt;
&lt;td&gt;~$1.83 to $5.25&lt;/td&gt;
&lt;td&gt;~$7.31 to $21.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM 5.2&lt;/td&gt;
&lt;td&gt;~1.02M to 1.68M input, plus 18K to 46K output&lt;/td&gt;
&lt;td&gt;~$0.89 to $1.50&lt;/td&gt;
&lt;td&gt;~$3.55 to $6.02&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;GLM sometimes spent more tokens. Its finished cases ran ~1.02M to 1.68M input tokens, while Kimi’s runtime-token band started lower at ~609K. But Kimi’s input rate is about 3.6x GLM’s, so the extra GLM context still comes out cheaper in this estimate.&lt;/p&gt;

&lt;p&gt;Tool calls landed in similar ranges, 13 to 23 per case for both. So the extra GLM tokens on roster and vendor work did not buy extra passes. The scores tied, with GLM carrying the cheaper bill.&lt;/p&gt;

&lt;p&gt;At 1,000 four-case batches, that envelope turns into roughly $7.3K to $21K for Kimi and $3.6K to $6.0K for GLM before cache discounts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verdict
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fobu4unmeg9dib78b2nc1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fobu4unmeg9dib78b2nc1.png" alt="Conclusion" width="800" height="414"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Kimi K3 and GLM 5.2 tied on pass rate, 7/12 each, 58%. I’d give the practical win to Kimi because it finished the biggest frontier case at 17/24 while GLM hit the 30-minute cap, and GLM spent more tokens on roster and vendor for the same or worse result.&lt;/p&gt;

&lt;p&gt;I would pick Kimi K3 when finishing the long tool workflow matters more than the cheaper rate. Ticket sync shows why: Kimi got through 24 turns, posted 17/24, and burned 1.75M runtime tokens. GLM did not finish inside the 30-minute cap. On Roster sync, Kimi also scored 8/13 while GLM scored 7/13, because Kimi posted the cover replies GLM dropped. It did that with 13 tool calls and 609K runtime tokens, while GLM used 19 tool calls and 1.06M input tokens.&lt;/p&gt;

&lt;p&gt;GLM 5.2 makes sense if your workload looks more like the easier historical band, or if you already want the Claude Code via OpenRouter setup and can live with the frontier misses. It matched Kimi’s top-line score, cleared the same 7/7 historical cases, and tied Kimi on Invoice sync, Refund ledger, and Vendor directory by score. Refund ledger is the one frontier case where GLM was cleaner on latency: 458.9s versus Kimi’s 685.8s, with both landing at 8/13.&lt;/p&gt;




&lt;h2&gt;
  
  
  Kimi K3 vs GLM 5.2: On personal builds
&lt;/h2&gt;

&lt;p&gt;So let me share some of the builds I tried with GLM and Kimi K3, along with prompt, time, cost, and builds. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For all builds I have used open router chatroom&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Meteor City (revival)
&lt;/h3&gt;

&lt;p&gt;Used Kimi K3 + GLM 5.2 to build Meteor City Revival, a game where you race to destroy the entire city before it can regenerate itself. &lt;/p&gt;

&lt;p&gt;The task was initially given to GLM 5.2, but for some reason it stopped mid-session, so I took all the code and asked Kimi to refine and recreate the entire build.&lt;/p&gt;

&lt;p&gt;Cost was approx. $4.3, used around 18.1M Tokens, Time: 1 hr 45 min. This is justified cause without explicitly mentioning it, it generated 10K procedural buildings, the engine, and figured out the lighting, ray tracing, shaders, and optimized the game for mobile as well as web. &lt;/p&gt;

&lt;p&gt;You can play the game at: &lt;a href="https://meteor-city-revival.vercel.app/" rel="noopener noreferrer"&gt;https://meteor-city-revival.vercel.app/&lt;/a&gt; &lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  Prompt Used
  &lt;br&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Create a complete, self-contained, single HTML file using Three.js (via CDN only — no other external dependencies or files) that implements a large-scale 3D photorealistic procedural meteor impact city destruction game/simulation.

**Core Gameplay (must be fully implemented and preserved exactly):**
- Procedural city with buildings that can be damaged and destroyed.
- Clickable "Launch Meteor" button that fires a meteor. User can launch multiple times.
- Buildings regenerate over time.
- **Push-and-wait mechanic**: Holding/clicking the Launch Meteor button charges a larger, more powerful meteor (powerup style).
- **Infinity powerup**: When activated, launches 5 big meteors in quick succession that deal massive damage (enough to push the destruction bar to ~95%).
- Powerups (larger meteor charge and Infinity) drop from meteor impacts and are automatically collected when the player is near them.
- Destruction progress bar that tracks overall city damage.
- At 100% destruction, display the text: "Now I am become Death, the destroyer of worlds."
- UI sliders for meteor size, speed, angle, time of day, and destruction intensity.
- Meteor and impact sounds using Web Audio API.
- Playable simulation/game with smooth performance.

**Visual &amp;amp; Technical Polish Requirements (focus here for realism and quality):**
- Highly realistic procedural city at night: varied building heights (low-rises to skyscrapers), realistic facades with window grids using InstancedMesh (windows have individual emissive colors that flicker or turn off when damaged), different roof styles, subtle material variation (concrete, glass, brick), minor architectural details like ledges.
- Use seed-based procedural variation so the city feels organic. Add roads, paths, and scattered green areas/parks between building clusters.
- Heavy use of InstancedMesh and LOD (Level of Detail) for performance.
- Realistic ground/terrain with subtle height variation, road networks, and support for crater formation on impact.
- Rich night sky: procedural starfield with twinkling stars, subtle moon glow, gradient sky with horizon haze and light pollution from the city. Add very subtle atmospheric effects.
- High-quality meteor: glowing fiery body with long dynamic particle trail (fire, sparks, smoke) that intensifies on entry.
- Realistic impact sequence: bright flash, expanding shockwave (particles + ground ripple), crater, layered particle systems for fire/explosions, dense rising dust/smoke plumes, and flying debris with gravity and tumbling.
- Improved building destruction: pieces break off with dust, structures partially crumble or lean, and damaged areas show reduced lighting/exposed sections.
- Dynamic lighting: moonlight + hemisphere light, multiple flickering point lights from fires and impact, emissive building windows, and fire effects. City lights progressively dim or extinguish with damage.
- Materials: Use MeshStandardMaterial where appropriate. Add subtle specular/roughness variation and rim lighting for depth.
- Special effects: Performant particle systems, screen shake on impact, bloom-like glow on bright elements, subtle motion blur during fast movement or camera fly-through, atmospheric perspective, and fog for depth.
- Overall cinematic yet realistic look with balanced night-time color grading (cool tones with warm fire accents).

**Camera, Controls &amp;amp; Performance Polish:**
- Smooth OrbitControls-style camera (mouse drag to orbit/pan, scroll to zoom) with optional free-fly mode (WASD + mouse look).
- Smooth camera interpolation and gentle auto-orbit when idle.
- Refined slow-motion replay with smooth timeScale control.
- Aggressive performance optimizations for stable 60+ FPS: InstancedMesh, LOD, frustum culling, efficient particle pooling, minimal draw calls.
- Subtle ambient animations (random window flickering, gentle dust movement).

**Technical Requirements:**
- Output ONLY the complete single HTML file (nothing else before or after).
- Must be immediately runnable in a modern browser with no errors.
- Include helpful inline comments explaining key visual, lighting, particle, and optimization techniques.
- Prioritize photorealistic visuals, cinematic quality, smoothness, and immersion while keeping all gameplay mechanics fully functional and unchanged.

Generate the full polished HTML code now.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;




&lt;p&gt;&lt;/p&gt;

&lt;p&gt;Output&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/RSDQZWeAP8E"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h3&gt;
  
  
  Plane Currents
&lt;/h3&gt;

&lt;p&gt;I always wanted to try low-poly 3D graphics, so I tried it with Kimi K3. &lt;/p&gt;

&lt;p&gt;Use Kimi K3 to build a paper plane simulator where the player passes through rings to gather points and complete the course. All while having a relaxing scene and music going in the background. (no 3d assets)&lt;/p&gt;

&lt;p&gt;This took around 16 minutes to generate, cost me approx $0.45, and used 30K tokens.&lt;/p&gt;

&lt;p&gt;You can play the game by opening the &lt;a href="https://gist.github.com/DevloperHS/0512038a0e1d21a8e854e4a771db8fa7" rel="noopener noreferrer"&gt;game file&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  Prompt Used
  &lt;br&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Build me a self contained  Relaxed 3D paper-plane flying gameplay where you launch and steer a customizable plane through glowing rings and floating islands over an ocean, collecting score multipliers in short, physics-light runs with easy controls (hold to launch, mouse/keyboard steering) with stylized low-poly 3D with clean cel-shaded visuals and a sleek, colorful indie-game UI featuring customizable paper planes, glowing rings, floating islands, and simple HUD element graphics. Output a single HTML file.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;




&lt;p&gt;&lt;/p&gt;

&lt;p&gt;And the output generated by the above prompt.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/F6zLuzzFbGk"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;




&lt;h2&gt;
  
  
  Gargantua Black Hole Geodesic Ray Tracer (Complex Physics Game/Sim) - inspo from X
&lt;/h2&gt;

&lt;p&gt;I was scrolling X and found this massive black hole geodesic ray tracer simulation made by someone. Being the space nerd I am, I wanted to make this too.&lt;/p&gt;

&lt;p&gt;So I did a bit of research and constructed a prompt that requires the model to think through the actual light-bending physics and math that happens near the event horizon of a black hole.&lt;/p&gt;

&lt;p&gt;I was completely hopeless cause fable and GLM 5.3 gave up on the calculation task earlier, but anyway I entered the prompt.  To my surprise, Kimi K3 actually went through the maths and solved it in its thinking traces. &lt;/p&gt;

&lt;p&gt;After approx 14 minutes and burning through 33K tokens, which costed around $0.51 (operouter), it handed me the complete code. &lt;/p&gt;

&lt;p&gt;I ran it, and here are the results&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/oidzFPtpU28"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  Prompt Used
  &lt;p&gt;Create a complete, self-contained single HTML file (no external libraries like Three.js) that implements a real-time geodesic raytracer for a Schwarzschild black hole inspired by Gargantua.&lt;/p&gt;

&lt;p&gt;Use raw WebGL2 with GLSL ES 3.00 in a single fragment shader. Implement accurate physics: null geodesic integration with 4th-order Runge-Kutta solver, event horizon, photon sphere, accretion disk with proper rendering, gravitational lensing, Doppler beaming, and gravitational redshift effects. Target stable 60 FPS performance.&lt;/p&gt;

&lt;p&gt;Include mouse-controlled camera orbiting/zooming, and a cyberpunk-style control panel with sliders for parameters (mass, spin, disk density, view angle, etc.). Add subtle particle effects for matter falling in and dynamic lighting/shadows.&lt;/p&gt;

&lt;p&gt;The output must be 100% complete, immediately runnable in a modern browser, with no black screen, NaNs, errors, or missing features. Prioritize numerical correctness, boundary handling, solver discipline, and physical accuracy above all. Verify and comment key physics equations in the code. Make it visually stunning and interactive like a premium physics demo/game.&lt;/p&gt;



&lt;p&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://gist.github.com/DevloperHS/0ca2efb20b16dd8497026b77a7d5dbba" rel="noopener noreferrer"&gt;Game File&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Yup, the code, math, and physics engine are all built by Kimi K3, and Fable failed to build the simulation with such a level of detail, which makes its claim worth the hype.&lt;/p&gt;

&lt;p&gt;I also tried 2 more tests to verify my doubts, sharing them as a bonus.&lt;/p&gt;




&lt;h3&gt;
  
  
  Bonus Test 1 (Held Karp Problem Solution (NP-Hard)
&lt;/h3&gt;

&lt;p&gt;It's not a surprise to me that Kimi K3 and GLM 5.2 were both able to solve this in no time, but I ran the test to check just raw coding + reasoning ability.&lt;/p&gt;

&lt;p&gt;The task was simple: fix the bug, create an optimal path, load env, run code, and give an answer to the buggy Held-Karp problem. Yup, the task has multiple steps for testing instruction following&lt;/p&gt;

&lt;p&gt;The test cost 0.03 cents, used 82K tokens (most on reasoning), and the result was out in 5 minutes. You can check the buggy code and fixed code from the attached files.&lt;/p&gt;

&lt;p&gt;Game File: &lt;a href="https://gist.github.com/DevloperHS/81d56857c2cee7b8cdd3b84e1d219d9f#file-problem-py" rel="noopener noreferrer"&gt;problems.py&lt;/a&gt; , &lt;a href="https://gist.github.com/DevloperHS/81d56857c2cee7b8cdd3b84e1d219d9f#file-solution-py" rel="noopener noreferrer"&gt;solution.py&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The prompt I used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Fix all bugs in the Held-Karp code in @file:problem.py so it correctly computes the minimum-cost tour for this 12-city TSP instance. Make the DP, base cases, transitions, and path reconstruction fully correct in  @file:fixed.py. Then create a new environment (.env) inside @file:held-karp-problem, install the dependencies, activate the environment, and run it to output the optimal cost and the tour as a list of cities starting and ending at 0.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Result&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmpfxq3xm0pjgn347td3i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmpfxq3xm0pjgn347td3i.png" alt="output" width="799" height="208"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I even validated it with one of my code geek friends and Grok 4.5 (expert)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd4wyeonlqc8t0vptd9iz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd4wyeonlqc8t0vptd9iz.png" alt="grok val" width="799" height="184"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Models don’t just need to output code; they also need to maintain behavioural constraints, so I tested both with a behavioural question. The result was similar.&lt;/p&gt;

&lt;p&gt;The task was simple: to resolve a conflict between team and stakeholder using the STAR Method &lt;/p&gt;

&lt;p&gt;Prompt&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Act as an experienced tech interview coach. Answer the following behavioral interview question using the STAR method (Situation, Task, Action, Result). Make the answer concise, professional, and impactful for a software engineering or tech role. Include quantifiable results where possible and highlight leadership or collaboration skills.

Question: Tell me about a time when you had to resolve a conflict within your team or with a stakeholder.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both models thought for a very short time and delivered the result in almost the same time. &lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model Name&lt;/th&gt;
&lt;th&gt;Token Count&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Duration&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GLM 5.2&lt;/td&gt;
&lt;td&gt;592&lt;/td&gt;
&lt;td&gt;$0.00162095856&lt;/td&gt;
&lt;td&gt;25.0 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;763&lt;/td&gt;
&lt;td&gt;$0.012987&lt;/td&gt;
&lt;td&gt;18.3 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;However, Kimi K3 won as it delivered a more credible, business-aligned conflict story with quantified stakeholder impact ($200K ARR, measurable failure reduction).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A product manager and backend engineer clashed over shipping a checkout feature on deadline versus fixing payment failures (affecting 3% of transactions). The tech lead reframed both concerns as "reliable payments delivered fast," then proposed shipping the feature behind a flag while hotfixing the top failure points. The plan shipped on time, cut failures from 3% to 0.4%, retained a $200K client, and became a team standard practice.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On the other hand, GLM&amp;nbsp;tells a technically impressive but somewhat predictable "engineering debate resolved by benchmarking" narrative.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Two senior engineers deadlocked over GraphQL vs REST for an API redesign, stalling the team for two weeks. The tech lead ran a proof-of-concept benchmark showing GraphQL won on performance (35% payload reduction), then added REST endpoints for backward compatibility to honor both perspectives. Development resumed in 3 days; the API improved response times 30% and maintained support for 12 existing clients.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Simple table for understanding&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;GLM 2.5&lt;/th&gt;
&lt;th&gt;Kimi K3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Stakeholder Range&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Two engineers only&lt;/td&gt;
&lt;td&gt;PM + Engineer (broader influence)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Business Context&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Process improvement&lt;/td&gt;
&lt;td&gt;Revenue at risk ($200K)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Conflict Complexity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Technical disagreement&lt;/td&gt;
&lt;td&gt;Business vs. tech risk trade-off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Resolution Approach&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Proof-of-concept (predictable)&lt;/td&gt;
&lt;td&gt;Phased delivery + data compromise (creative)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Lasting Impact&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Team velocity restored&lt;/td&gt;
&lt;td&gt;Process adoption + trust rebuilt + client retained&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Interviewer Appeal&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Shows technical leadership&lt;/td&gt;
&lt;td&gt;Shows &lt;strong&gt;business acumen&lt;/strong&gt; + technical leadership&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This shows Kimi K3 is more aligned with the workspace and can provide factual answers when needed. Really impressive.&lt;/p&gt;

&lt;p&gt;With this, we have come to the end of this deep dive, but here is what I have to say at the end.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion: The Gap is Closing
&lt;/h2&gt;

&lt;p&gt;Six months ago, comparing open-source models to Claude and GPT meant accepting tradeoffs with performance, quality, builds, and output.&lt;/p&gt;

&lt;p&gt;Today, models like Kimi K3 are extremely competitive across coding, agentic, and multimodal tasks.&lt;/p&gt;

&lt;p&gt;Also, GLM-5.2 shows competitive performance across industry-standard evaluations, frequently rivalling or approaching proprietary models such as GPT-5.5 and Claude Opus 4.8.&lt;/p&gt;

&lt;p&gt;Here is what most people are missing.&lt;/p&gt;

&lt;p&gt;The talk is no longer about closed vs open source;&amp;nbsp; It's about&amp;nbsp;&lt;em&gt;specialised&lt;/em&gt;&amp;nbsp;vs general, and&amp;nbsp;&lt;em&gt;long-context practical&lt;/em&gt;&amp;nbsp;vs theoretical.&lt;/p&gt;

&lt;p&gt;Both K3 and GLM-5.2 are proving that open-source can own specific workloads better than models 10x the marketing budget.&lt;/p&gt;

&lt;p&gt;For builders:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Choose Kimi K3&lt;/strong&gt; if you're building agents that reason for hours, need multimodal perception, or can absorb frontier pricing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Choose GLM-5.2&lt;/strong&gt; if you want open weights today, need fast inference on a GPU, or are optimizing for math and code.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Either way, you're not choosing good models. You're choosing the right model for the right task, and that matters.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>gamedev</category>
      <category>programming</category>
    </item>
    <item>
      <title>How to connect MCP servers to Slackbot</title>
      <dc:creator>Shrijal Acharya</dc:creator>
      <pubDate>Sat, 18 Jul 2026 11:57:47 +0000</pubDate>
      <link>https://dev.to/composiodev/how-to-connect-mcp-servers-to-slackbot-1al4</link>
      <guid>https://dev.to/composiodev/how-to-connect-mcp-servers-to-slackbot-1al4</guid>
      <description>&lt;p&gt;Slackbot recently added support for MCP, which means you can now connect it with external apps and let it take actions across your work tools directly from Slack&lt;/p&gt;

&lt;p&gt;But the native app list is still limited. By the time of writing this post, there's just about &lt;strong&gt;20 apps&lt;/strong&gt; that you can connect from the &lt;a href="https://slack.com/marketplace" rel="noopener noreferrer"&gt;Slack marketplace&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;And for most teams, that's not enough.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqlyug2z5afeztlhxrwtd.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqlyug2z5afeztlhxrwtd.gif" alt="not enough gif" width="480" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;But luckily, Slack allows you to set up or use your custom MCP servers and not have to be limited by the number of apps available in marketplace.&lt;/p&gt;

&lt;p&gt;That's where &lt;a href="https://composio.dev/" rel="noopener noreferrer"&gt;Composio&lt;/a&gt; helps you. It can connect your slack bots to &lt;strong&gt;1000+ apps&lt;/strong&gt; that you can use.&lt;/p&gt;

&lt;p&gt;In this guide, we’ll go through how to connect Slackbot with Composio’s MCP server in 3 steps.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;ℹ️ The steps will be pretty much the same with other MCP servers as well.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What We're Building
&lt;/h2&gt;

&lt;p&gt;Once this is set up, you can ask Slackbot things like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Find my latest unread Gmail emails.
Search my Notion workspace for launch notes.
Check my Google Calendar for meetings tomorrow.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;And a bunch more. Imagine all the stuff you can do with 1000+ apps. 😵‍💫&lt;/p&gt;

&lt;p&gt;I'll leave the rest to your imagination...&lt;/p&gt;

&lt;p&gt;Slackbot sends the request to Composio Connect, Composio finds the right tool, asks you to connect the app if needed (one time), and then executes the action.&lt;/p&gt;


&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;p&gt;Before we begin, make sure you have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Slack workspace with Slackbot MCP client access (comes with Business+ and Enterprise plan)&lt;/li&gt;
&lt;li&gt;Permission to create or configure a Slack app.&lt;/li&gt;
&lt;li&gt;A Composio account.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's it.&lt;/p&gt;


&lt;h2&gt;
  
  
  Step 1: Get the Composio Connect MCP URL
&lt;/h2&gt;

&lt;p&gt;First, you need the MCP server URL from Composio.&lt;/p&gt;

&lt;p&gt;For this setup, use Composio Connect:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;https://connect.composio.dev/mcp&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;This is Composio’s hosted MCP server that gives your AI agent access to 1,000+ apps with just &lt;strong&gt;7 meta-tools&lt;/strong&gt; that let the slackbot discover what's available, authorize apps on demand, and execute tools across apps in parallel through a single connection.&lt;/p&gt;

&lt;p&gt;You don’t need to create a custom MCP server for this guide.&lt;/p&gt;

&lt;p&gt;Composio also supports custom MCP servers for more scoped project-specific use cases, but those can require API-key-based auth. For Slackbot, Composio Connect is the simpler path because it works with OAuth-based MCP client flows.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9r6it1xz4hl4h3ifel7a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9r6it1xz4hl4h3ifel7a.png" alt="Composio Connect" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Step 2: Add Composio Connect to a Slack App
&lt;/h2&gt;

&lt;p&gt;Now, we need to register the Composio MCP server inside a Slack app.&lt;/p&gt;

&lt;p&gt;Go to the Slack developer dashboard and create a new app.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5kkqe4kem5xnfmpyla6r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5kkqe4kem5xnfmpyla6r.png" alt="Slack new app creation" width="800" height="432"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once the app is created:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open your Slack app.&lt;/li&gt;
&lt;li&gt;In the left sidebar, go to &lt;strong&gt;Features&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Click &lt;strong&gt;MCP Servers&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwvranutnruw5ja1w5vhm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwvranutnruw5ja1w5vhm.png" alt="Slack MCP Servers button" width="800" height="432"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Click &lt;strong&gt;Get Started&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff2yx87053fw9upha8l28.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff2yx87053fw9upha8l28.png" alt="Slack get started button" width="800" height="434"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now fill in the MCP server details.&lt;/p&gt;

&lt;p&gt;Use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Name&lt;/strong&gt;: Composio (or anything you wish)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;URL&lt;/strong&gt;: &lt;a href="https://connect.composio.dev/mcp" rel="noopener noreferrer"&gt;https://connect.composio.dev/mcp&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth Type&lt;/strong&gt;: Dynamic Client Registration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnyzro2ycp7xjqhb3gsnw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnyzro2ycp7xjqhb3gsnw.png" alt="Slack Add MCP" width="800" height="434"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For the auth type, select &lt;strong&gt;Dynamic Client Registration&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;ℹ️ This is the right option for MCP servers that support OAuth discovery and client registration. Slack handles the client registration automatically, so you don’t need to manually create OAuth credentials first.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now, if you click the three dots and then &lt;strong&gt;Tools&lt;/strong&gt;, you should see that it currently cannot fetch the tools because the MCP server uses a dynamic connection and must be installed in your workspace first.&lt;/p&gt;

&lt;p&gt;So, now head over to the &lt;strong&gt;Install App&lt;/strong&gt; tab, and install it to the workspace you selected when creating the app.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxwv4sjhjmjvmkiflvhyp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxwv4sjhjmjvmkiflvhyp.png" alt="Slack Install App" width="800" height="434"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If your workspace requires approval, send the app request to your admin.&lt;/p&gt;


&lt;h2&gt;
  
  
  Step 3: Connect Composio inside Slackbot
&lt;/h2&gt;

&lt;p&gt;Once your Slack app is installed and approved, open a DM with Slackbot.&lt;/p&gt;

&lt;p&gt;Then, just type in a prompt that requires using the app, Slack will use the correct app automatically for you.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvp5wb4l3ab4anrfqv8y9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvp5wb4l3ab4anrfqv8y9.png" alt="Slack connecting composio" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Click &lt;strong&gt;Connect,&lt;/strong&gt; and you’ll be taken to a confirmation page.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4f5nvtnc0j1x1wkbv356.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4f5nvtnc0j1x1wkbv356.png" alt="Slack connecting composio confirmation" width="800" height="489"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Click on Continue, and then confirm it on the Composio end.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fljriezhxmemb38w8apqz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fljriezhxmemb38w8apqz.png" alt="Composio confirmation" width="800" height="489"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If everything went well, you should see that your account is connected.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F58g6dsw3n767nztxc4mo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F58g6dsw3n767nztxc4mo.png" alt="Slack final confirmation" width="799" height="293"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After that, Slackbot should be able to discover Composio’s MCP tools.&lt;/p&gt;

&lt;p&gt;Start with a simpler test:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What tools are available from Composio?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Once that goes through, now try an actual app action.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Send a mail to x@y.com saying 'Hi, from Composio 👋 inside Slackbot'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;If Gmail is not connected yet, Composio should generate an OAuth link for you to connect it. Once you approve it, the connection persists for future use. So, you don't have to repeat this step again and again.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5fb169cfpktm2xynzoe3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5fb169cfpktm2xynzoe3.png" alt="Composio connection link" width="800" height="438"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Slackbot may ask you to approve the action before it writes data to another app.&lt;/p&gt;

&lt;p&gt;That's expected. Once connected, Slackbot can use that app through Composio.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxlw14j7ed5fi6h0o6dk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxlw14j7ed5fi6h0o6dk.png" alt="Composio MCP in action" width="800" height="438"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Voilà, you've successfully connected Slackbot to Composio MCP. 🎊&lt;/p&gt;

&lt;p&gt;Here’s a quick workflow for initiating a connection and running an actual app action:&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/m6kv3tqjUgU"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;


&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;Slack's own marketplace is good, and if it covers all the tools you require, you can completely stick to it.&lt;/p&gt;

&lt;p&gt;But for some of you, that's simply not enough. I hope this helps overcome that problem.&lt;/p&gt;

&lt;p&gt;So instead of jumping between different tools, you can ask Slackbot to find information, create tasks, update records, and run actions across your apps from inside Slack.&lt;/p&gt;

&lt;p&gt;This is a much-needed quality-of-life improvement for teams that already live in Slack.&lt;/p&gt;

&lt;p&gt;Slackbot gives you the interface. MCP gives you the protocol.&lt;/p&gt;

&lt;p&gt;And Composio gives you the app layer. 👌&lt;/p&gt;


&lt;div class="ltag__user ltag__user__id__1127015"&gt;
    &lt;a href="/shricodev" class="ltag__user__link profile-image-link"&gt;
      &lt;div class="ltag__user__pic"&gt;
        &lt;img src="https://media2.dev.to/dynamic/image/width=150,height=150,fit=cover,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1127015%2F1c5e48a2-f602-4e7d-8312-3c0322d155c6.jpg" alt="shricodev image"&gt;
      &lt;/div&gt;
    &lt;/a&gt;
  &lt;div class="ltag__user__content"&gt;
    &lt;h2&gt;
&lt;a class="ltag__user__link" href="/shricodev"&gt;Shrijal Acharya&lt;/a&gt;Follow
&lt;/h2&gt;
    &lt;div class="ltag__user__summary"&gt;
      &lt;a class="ltag__user__link" href="/shricodev"&gt;SDE • GOLD @Microsoft Student Ambassador • Prev Lead Collab and Dev-Team Lead @oppiaorg • Mail for collaboration&lt;/a&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;



</description>
      <category>ai</category>
      <category>productivity</category>
      <category>beginners</category>
      <category>automation</category>
    </item>
    <item>
      <title>Cursor Vs Claude Code: Which one you should pick (or both)</title>
      <dc:creator>Developer Harsh</dc:creator>
      <pubDate>Thu, 16 Jul 2026 17:10:26 +0000</pubDate>
      <link>https://dev.to/composiodev/cursor-vs-claude-code-which-one-you-should-pick-or-both-8o7</link>
      <guid>https://dev.to/composiodev/cursor-vs-claude-code-which-one-you-should-pick-or-both-8o7</guid>
      <description>&lt;p&gt;Cursor and Claude Code are 2 leading products that engineers reach out for nowadays. &lt;/p&gt;

&lt;p&gt;Both can refactor whole codebases, hunt for bugs, run spec-driven builds, and handle vibe coding needs, in the same ecosystem (skills, mcps, plugins, hooks ) and harness that ties the agent loop together. Same rig, yet both cater to a different workflow.&lt;/p&gt;

&lt;p&gt;And Cursor had recently become hard to ignore as SpaceX&amp;nbsp;signed a $60 billion all-stock deal to buy its parent company,&amp;nbsp;Anysphere&amp;nbsp;(closes Q3 2026), and around the same time, it shipped&amp;nbsp;Origin, its own githost for agents, plus&amp;nbsp;cloud agents,&amp;nbsp; Composer 2.5,&amp;nbsp;and Grok 4.5.&lt;/p&gt;

&lt;p&gt;As for me, I use both every single day. I even rewrote my X bio in their honor: &lt;em&gt;I touch Claude Code, Cursor for a living.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;None of this is free, though. For 6 months, I have happily paid&amp;nbsp;&lt;strong&gt;$20/mo for Cursor&lt;/strong&gt;&amp;nbsp;and&amp;nbsp;&lt;strong&gt;$100/mo for Claude Code&lt;/strong&gt;&amp;nbsp;because neither tool excels at everything in my workflow.&lt;/p&gt;

&lt;p&gt;However, not everyone needs both, and not everyone wants to spend $120 a month to find out. If that is you, the question shifts to what you actually get for each dollar. &lt;/p&gt;

&lt;p&gt;This is what this guide answers. Let’s begin&lt;/p&gt;




&lt;h4&gt;
  
  
  TLDR
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Section&lt;/th&gt;
&lt;th&gt;Which to pick&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Models and tooling&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; for depth on one model;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; for option across many&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Billing cost's real story&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; on focused tasks (pay per fetch, ~5.5x fewer tokens);&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; for huge monorepos, but you pay to keep the index fresh&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Task delegation (async agents)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; for supervised and visual;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; for raw delegated horsepower&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The harness&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; for UI-heavy work you want to watch and control;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; for repeatable, version-controlled instructions that run themselves&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Everyday usage&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; for hands-on control over every edit;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; for reviewing finished work instead of keystrokes&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pricing&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; for simple, predictable flat pricing;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt;'s $100 only pays off on token-heavy work&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extensibility (MCP, Skills, plugins)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; to distribute a governed toolset to a large team;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; for reproducible agent behavior that lives in the repo&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data privacy&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; for a narrower footprint;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; will soon route your whole stack through one owner (SpaceX)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;But before moving forward, I would like to clear up a common paradox people get caught up in.&lt;/p&gt;




&lt;h2&gt;
  
  
  The common paradox
&lt;/h2&gt;

&lt;p&gt;Most people think Cursor is an AI editor with tools, while Claude Code is an AI agent you hand tasks to. That is not their fault tbh. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;But my friend, that framing died twice.&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, the interfaces merged:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code now runs in VS Code, as a desktop app, and in the browser.&lt;/li&gt;
&lt;li&gt;Cursor runs as a desktop app, in a terminal, in the cloud, on iOS, and on the web.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I even run Claude Code &lt;em&gt;inside&lt;/em&gt; Cursor most days now.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp7277g03008gltdp7ojx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp7277g03008gltdp7ojx.png" alt="Cursor Image" width="800" height="472"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, Cursor stopped being a code editing tool and became a platform.&lt;/strong&gt; It now owns the full software factory, top to bottom:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Write your code in Cursor,&lt;/li&gt;
&lt;li&gt;Review it with Bugbot,&lt;/li&gt;
&lt;li&gt;Host it on Origin (new release),&lt;/li&gt;
&lt;li&gt;Run it on models trained by its own group.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Look at the last seven months alone:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Move&lt;/th&gt;
&lt;th&gt;What it means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Dec 2025&lt;/td&gt;
&lt;td&gt;Acquired &lt;strong&gt;Graphite&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;Owns code review: stacked PRs, merge queues&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feb 2026&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Bugbot&lt;/strong&gt; went reviewer to &lt;em&gt;fixer&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;Spots a bug, spins its own agent, tests a fix, proposes it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jun 2026&lt;/td&gt;
&lt;td&gt;Announced &lt;strong&gt;Origin&lt;/strong&gt;, a GitHub rival&lt;/td&gt;
&lt;td&gt;Git hosting for the agentic era, AI merge-conflict resolution. Waitlist, ships fall 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jun 16, 2026&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;SpaceX agreed to acquire Cursor&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$60B all-stock per an SEC 8-K filing, close expected Q3 2026, into the xAI group&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This tells you that the platform is no longer what it was 6 months ago; the cursor now owns the entire integrated software factory stack. More about it on the Product Hunt discussion.&lt;/p&gt;

&lt;p&gt;So if the interfaces are roughly the same now, what actually separates these two tools? Read on.&lt;/p&gt;




&lt;p&gt;Two years ago, the model was the moat. In 2026, it hardly is.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; runs on Claude, currently Opus 4.8. The tool and the model are tuned for each other, and you feel it in how confidently it plans multi-step work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; lets you select your own brains: GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro, Grok 4.5, and its own Kimi K2.5 finetuned Composer 2.5. Pick the right model per task, pay Cursor to route.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On my 12-file NASA JPL refactor, Claude Code read most of the tree before writing a line, and the first pass barely needed correction. &lt;/p&gt;

&lt;p&gt;As for the cursor, it also handled things smoothly with a one-time correction with a function. It was because a model API call failed and was partially completed.&lt;/p&gt;

&lt;p&gt;This has also been a concern for the cursor teams, and they aim to be the lab, rather than a model router.  Also, it aims to invent a new kind of programming where any idea can just be represented in English.&lt;/p&gt;

&lt;p&gt;Truell’s June 16 Compile keynote &amp;amp; later in Lenny’s podcast addresses this nicely:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Our goal with Cursor is to invent a new type of programming. It looks like a world where you have a representation of the logic of your software that does look more like English. You can imagine kind of an evolution of programming language towards pseudocode. You have written down the logic of the software, and you can edit that at a high level. It won't be the impenetrable millions of lines of code, it'll instead be something that's much terser and easier to understand and easier to navigate."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now that’s training its first frontier model from scratch, 1.5 trillion parameters, on xAI's Colossus cluster, under SpaceX's $60 billion deal. It seems the company is heading into its next phase and aims to become the model developer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who wins?&lt;/strong&gt; &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you want a model tuned straight into the tool, go to Claude Code.&lt;/li&gt;
&lt;li&gt;If you want model variety today and a bet on Cursor's own lab tomorrow, go with Cursor.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  2. Pricing: Cursor vs Claude Code
&lt;/h2&gt;

&lt;p&gt;Most of us stop at the sticker price.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cursor&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hobby&lt;/td&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro&lt;/td&gt;
&lt;td&gt;$20/mo&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro+&lt;/td&gt;
&lt;td&gt;$60/mo&lt;/td&gt;
&lt;td&gt;3x usage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ultra&lt;/td&gt;
&lt;td&gt;$200/mo&lt;/td&gt;
&lt;td&gt;20x usage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Teams&lt;/td&gt;
&lt;td&gt;$40/user ($120 Premium seat)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;Custom&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ul&gt;
&lt;li&gt;Every paid plan runs on usage credits, with on-demand billing past your allotment. Turn on spend limits the day you start.&lt;/li&gt;
&lt;li&gt;Auto mode is the cheap lever: it runs Composer 2.5 or routes to a capable model automatically, and it is unlimited on paid plans.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;No free tier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro&lt;/td&gt;
&lt;td&gt;$20/mo&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max&lt;/td&gt;
&lt;td&gt;$100/mo ($200 for 20x)&lt;/td&gt;
&lt;td&gt;Unlocks Opus, up-to-1M context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Teams&lt;/td&gt;
&lt;td&gt;$25/seat ($20 annual)&lt;/td&gt;
&lt;td&gt;Caps at 150 seats&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise&lt;/td&gt;
&lt;td&gt;Base seat + API usage&lt;/td&gt;
&lt;td&gt;Cheaper light, pricier heavy&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ul&gt;
&lt;li&gt;No free tier; the cheapest door is Pro at $20.&lt;/li&gt;
&lt;li&gt;Limits run on two clocks at once: a 5-hour rolling window plus a weekly cap, so an all-day session can hit the wall mid-task.&lt;/li&gt;
&lt;li&gt;To trim spend, route routine edits to Sonnet or Haiku, save Opus for hard refactors.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On paper, both start at $20, and the offerings look even. It isn’t!&lt;/p&gt;

&lt;p&gt;The sticker price hides what it actually costs to run a task.&lt;/p&gt;

&lt;p&gt;On a widely repeated refactor test, the same job cost wildly different amounts of compute:&lt;/p&gt;

&lt;p&gt;Tokens used on the same refactor  (lower is better)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy1c36ikwdq4vggbdq8vl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy1c36ikwdq4vggbdq8vl.png" alt="comaprison" width="799" height="174"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://medium.com/@gvelosa/claude-code-vs-cursor-in-2026-the-token-efficiency-gap-befd0864e0a5" rel="noopener noreferrer"&gt;Source&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is a &lt;strong&gt;5.5x gap&lt;/strong&gt; for identical output. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One honest note: this is a single community benchmark; the two agents ran different models under the hood, and at least one prints the numbers flipped. Treat it as a strong signal, not a law. It also does not hold everywhere.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This happens because each tool loads the context differently:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cursor retrieves using a hybrid stack: a semantic index, grep, and an Explore subagent. Index-backed and targeted, strong on huge monorepos, but the index carries a standing cost to build and keep fresh. You pay for it.&lt;/li&gt;
&lt;li&gt;Claude Code skips the index and greps, globs, and reads on demand. Index-free and just-in-time, so you pay only for what it pulls into context.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;One caveat&lt;/strong&gt;: both start to degrade beyond roughly 150k tokens of genuinely relevant context, so neither truly wins at extreme scale.&lt;/p&gt;

&lt;p&gt;So, who wins?&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cursor:&lt;/strong&gt; Use it if you want simple, predictable flat pricing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code:&lt;/strong&gt; the $100 only pays off on token-heavy work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro tip:&lt;/strong&gt; Model both against your own usage, then decide.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  3. Task Delegation with Async agents
&lt;/h2&gt;

&lt;p&gt;The refactor task I gave to Claude and the bug hunt task to Cursor were not from the terminal/app UI; they were through my mobile phone. In fact, I barely touch my pc while traveling.&lt;/p&gt;

&lt;p&gt;Essence is simple.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10e3f5sq7mohu5pw95ts.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10e3f5sq7mohu5pw95ts.png" alt="Task Delegation" width="800" height="84"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In short, task delegation is here, but both Claude Code and cursor build around this differently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cursor&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cursor&lt;/strong&gt; builds this around cloud agents and Automations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Launches cloud dev environment in under 10 minutes, snapshot it, reuse it.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/in-cloud&lt;/code&gt; spins a subagent on its own VM and branch, while your laptop stays unaffected.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/automate&lt;/code&gt; creates jobs in plain language, with GitHub and Slack triggers.&lt;/li&gt;
&lt;li&gt;Bugbot review runs ~3x faster, roughly 90 seconds a pass, and can be called with &lt;code&gt;/review&lt;/code&gt; before you push.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt; builds this around agent teams and background sessions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cloud dev environment for remote sessions, spun up automatically the first time you run a remote feature, no manual web setup.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/agents&lt;/code&gt; launches a coordinated team where one session leads and others execute, viewable and steerable in the &lt;code&gt;claude agents&lt;/code&gt; view.&lt;/li&gt;
&lt;li&gt;Background agents run on separate git worktrees; kick one off from &lt;code&gt;claude agents&lt;/code&gt;, then steer it from your phone via Remote Control in the mobile app.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/code-review&lt;/code&gt; runs a review pass on your changes, improved on Opus 4.8 across effort levels, and &lt;code&gt;/security-review&lt;/code&gt; scans for vulnerabilities before you push.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Who wins?&lt;/strong&gt; &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Supervised and visual, go for Cursor.&lt;/li&gt;
&lt;li&gt;Raw delegated horsepower, go for Claude Code.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4. The Harness
&lt;/h2&gt;

&lt;p&gt;Strip as an agent down to its core, and that is a model in a loop with tools. &lt;/p&gt;

&lt;p&gt;Everything wrapped around that loop: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The memory,&lt;/li&gt;
&lt;li&gt;The standing instructions (system prompt),&lt;/li&gt;
&lt;li&gt;The automatic hooks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;decides whether the loop is reliable and works. Some people call this layer the rig, but it's commonly called a harness.&lt;/p&gt;

&lt;p&gt;Both Claude code and cursor ships with this harness, but are targeted for different workflows:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ships the harness native and documented:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;CLAUDE.md&lt;/code&gt;: standing instructions, the agent reads every session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills&lt;/strong&gt;: packaged workflows you invoke like &lt;code&gt;/review-pr&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hooks&lt;/strong&gt;: shell commands that fire on lifecycle events, so a formatter runs after every edit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Artifacts&lt;/strong&gt;: session work captured as a live web page (a PR walkthrough, a dashboard), a non-terminal teammate can read, with private org sharing and version history.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Cursor&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Its harness version exists too, and it is growing fast:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A single &lt;strong&gt;Customize&lt;/strong&gt; page pulls together plugins, skills, MCP servers, subagents, rules, commands, and hooks, with a marketplace on top.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Design Mode&lt;/strong&gt; lets you point at UI elements in the browser or on a canvas, select several at once, and narrate changes by voice while agents edit beneath the surface.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The difference is the center of gravity:&lt;/p&gt;

&lt;p&gt;Cursor's harness orbits the editor and the visual surface, and yes, it's amazing. I give one instruction in the 1st prompt, and it carries forward until the chat ends. &lt;/p&gt;

&lt;p&gt;Claude Code orbits the agent loop and the command line. In my experience, I tend to forget important instructions mid-conversation if the topic strays too far or the chat goes on too long.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So, who wins?&lt;/strong&gt; &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cursor: For UI-heavy work where you want to see every change and be in control. (not true delegation, but secure)&lt;/li&gt;
&lt;li&gt;Claude Code: For repeatable, version-controlled instructions (a &lt;code&gt;CLAUDE.md&lt;/code&gt; file plus Hooks) that your whole team inherits automatically, so the rules run on their own instead of relying on you to remember them. (true delegation, but feel less secure)&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  5. Everyday Usage
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cursor feels at home;&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It&lt;/strong&gt;&amp;nbsp;is a VS Code fork, so it looks like the editor I already use,&lt;/li&gt;
&lt;li&gt;Tab autocomplete predicts my next several edits as I type.&lt;/li&gt;
&lt;li&gt;Within an hour of writing by hand, I felt faster when I tried it for the 1st time 6 months back.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Claude Code feels like running a company;&lt;/strong&gt; &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It's terminal + ui native, no autocomplete to fall for. Mainly for task delegation.&lt;/li&gt;
&lt;li&gt;Once I have written a good enough &lt;code&gt;CLAUDE.md&lt;/code&gt; and wired a couple of Hooks, Specs, and project-level skills, it runs whole tickets across multiple subagents in parallel while I go through the diff.&lt;/li&gt;
&lt;li&gt;Mainly for task delegation, the payoff arrives late but is bigger than the current.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Review style&lt;/th&gt;
&lt;th&gt;Cursor&lt;/th&gt;
&lt;th&gt;Claude Code&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What you see&lt;/td&gt;
&lt;td&gt;Each change inline, accept or edit before it lands&lt;/td&gt;
&lt;td&gt;The finished result plus the reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;UI edits where you eyeball every pixel&lt;/td&gt;
&lt;td&gt;Delegated tickets you review as a whole&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tests&lt;/td&gt;
&lt;td&gt;You trigger them&lt;/td&gt;
&lt;td&gt;It runs them, iterates on failure, reports back&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Who wins?&lt;/strong&gt; &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tight control over every edit and well-controlled task delegation with AGENTS.md, go with Cursor.&lt;/li&gt;
&lt;li&gt;Review finished work instead of keystrokes, go with Claude Code.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  6. Extensibility: MCP, Skills, plugins
&lt;/h2&gt;

&lt;p&gt;Both speak MCP, the protocol for wiring outside tools and data into an agent. Both turned it into a team-management surface rather than a solo toy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cursor&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;With the cursor, teams can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Configure Team MCP servers once, push them across cloud agents, the agents window, the IDE, and the CLI.&lt;/li&gt;
&lt;li&gt;Publish approved integrations to a team marketplace so members can install without touching config.&lt;/li&gt;
&lt;li&gt;Added GitLab, BitBucket, and Azure DevOps support for those imports.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;With the Claude Code, teams can&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Leans on Skills, Hooks, and a plugin system, plus MCP for outside connections.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use Cowork&lt;/strong&gt; brings agent machinery to knowledge work within a local, isolated VM with access to their files.&lt;/li&gt;
&lt;li&gt;A computer-use preview lets Claude open files, click, and navigate for you.&lt;/li&gt;
&lt;li&gt;A Slack integration (Team and Enterprise plans) lets you tag Claude to hand off a task without leaving the channel.&lt;/li&gt;
&lt;li&gt;Treats extensibility as a code check-in.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But what about solo dev’s?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Solo Devs&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For the solo dev, none of the above matters; it's about speed, for example: how fast can you load tools, skills, and MCP that follow on every machine and get work done.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; is the easy on-ramp.

&lt;ul&gt;
&lt;li&gt;Adding an MCP server is a few clicks with OAuth built in, no config file to hand-edit, and you inherit the entire VS Code extension library on day one.&lt;/li&gt;
&lt;li&gt;The Customize page works at the user level, too, so your rules, skills, and MCPs live in one place as local instructions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; gives a solo dev the same files-in-repo power the teams get.

&lt;ul&gt;
&lt;li&gt;The &lt;code&gt;CLAUDE.md&lt;/code&gt; skills, hooks, and plain files are plain files that users can commit to, so their agent behaves identically on their laptop, desktop, or any box they clone into.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Either way, the setup is only half the battle.&lt;/p&gt;

&lt;p&gt;Often, we solo developers struggle to connect to multiple tools, manually pass API keys, worry about security, and hope for optimized tool calls. ‘&lt;/p&gt;

&lt;p&gt;So for this, I use composio, which helps me connect to 1000+ tools/services in one click, while handling all the issues I mentioned earlier.  - Just a practical experience here, your call.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Who wins?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Distribute a governed toolset to a large team through a marketplace. Cursor leads today.&lt;/li&gt;
&lt;li&gt;Reproducible agent behavior that lives in the repo, Claude Code fits how engineers already work.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  7. Most Important Factor
&lt;/h2&gt;

&lt;p&gt;This is the point most take for granted. The data privacy.&lt;/p&gt;

&lt;p&gt;Cursor is on its way to becoming a SpaceX subsidiary, folded into the xAI group, once the $60B deal closes in Q3. &lt;/p&gt;

&lt;p&gt;Pair that with Origin (its own git host) and Composer (its own model), and one company could soon own the tool that writes your code, the place that stores it, and the model that learns from it. That is genuinely new. No prior git host has also owned the model doing the writing.&lt;/p&gt;

&lt;p&gt;I am not calling it a trap, and I am not assuming bad intent. But if you work on client repos with strict rules about where code can live, as I do, think about this before you migrate anything. &lt;/p&gt;

&lt;p&gt;Always read the terms. Watch where the data goes. &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code keeps a narrower footprint, an agent and a harness rather than a whole hosting stack, though your code still travels to Anthropic's API either way.&lt;/li&gt;
&lt;li&gt;Cursor soon will own the stack, your code, your tool calls, your decision, plan, and all builds will go through the cursor for better model training.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The Verdict
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;You are...&lt;/th&gt;
&lt;th&gt;Your pick&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;UI / product engineer&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Cursor&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Editor you know, inline autocomplete, visual diffs, pick a model per task. The best AI code editor you can buy.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Systems / backend engineer&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Claude Code&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Delegate whole tickets, reproducible agent behavior from repo files, orchestrate several agents, review finished work. Its rig is the more serious engineering today, and its token efficiency is a real cost edge.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Most of us&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Both&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$120/mo total, for a month. Let the work sort it out.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The engineers I know who ship fastest stopped treating this as a loyalty test and started treating it as two tools for two kinds of tasks. &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cursor for the hands-on sessions.&lt;/li&gt;
&lt;li&gt;Claude Code for the delegated automations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The debate between cursor and clauded code ends the moment you stop arguing and start building. This is where I landed after six months with the two subscriptions. &lt;/p&gt;

&lt;p&gt;Remember, your repo and your habits will move these numbers in the future, so borrow my framework, not my conclusion.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>learning</category>
      <category>programming</category>
    </item>
    <item>
      <title>The Enterprise MCP Gateway Buyer's Guide: SSO, SCIM, Audit, and Governance Requirements</title>
      <dc:creator>Dumebi Okolo</dc:creator>
      <pubDate>Sun, 05 Jul 2026 22:31:54 +0000</pubDate>
      <link>https://dev.to/composiodev/the-enterprise-mcp-gateway-buyers-guide-sso-scim-audit-and-governance-requirements-ho7</link>
      <guid>https://dev.to/composiodev/the-enterprise-mcp-gateway-buyers-guide-sso-scim-audit-and-governance-requirements-ho7</guid>
      <description>&lt;p&gt;&lt;em&gt;MCP gateways are becoming mandatory infrastructure for any organization deploying AI agents at scale. Here is what they actually do, what they must do, and how to evaluate one honestly.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;In November 2024, Anthropic released the Model Context Protocol: a wire format for connecting AI clients to tools, data sources, and APIs. Eighteen months later, MCP has crossed 78% adoption among production AI engineering teams. The public registry has passed 9,400 servers. Anthropic, OpenAI, Google, and Microsoft all support it. Practitioners have started calling it "the USB-C of AI applications."&lt;/p&gt;

&lt;p&gt;The protocol's success created an infrastructure problem that nobody anticipated at quite this speed. Every MCP server connection expands an organization's attack surface. Every AI agent operating with tool access can read private data, write to production systems, and execute commands under the permissions of whoever authorized it. Without a governance layer, these agents are black boxes: no audit trail, no access control, no identity attribution, no way to answer "what did this agent do?" to an auditor.&lt;/p&gt;

&lt;p&gt;The answer the market has converged on is an MCP gateway: a control plane that sits between AI agents and the tools they call. But the term covers a lot of ground, from lightweight protocol proxies to full enterprise governance platforms. The differences are significant. Getting the choice wrong creates compliance exposure; getting it right creates the foundation for scaling AI safely.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a gateway actually is
&lt;/h2&gt;

&lt;p&gt;The core function of an MCP gateway is collapsing what engineers call the N×M integration problem. Without a gateway, every AI agent manages its own credentials, authentication flows, and access policies for every tool it connects to. &lt;/p&gt;

&lt;p&gt;Ten agents and twenty tools produce two hundred independent connection paths, each with its own credentials, each potentially leaking secrets, each invisible to anyone trying to govern AI behavior centrally. A gateway reduces that to a single control point: N agents connect to the gateway; the gateway manages access to M tools.&lt;/p&gt;

&lt;p&gt;That description makes it sound like a proxy. It is not just a proxy. The proxy, the routing layer, accounts for roughly five percent of what an enterprise-grade gateway actually delivers. The remaining ninety-five percent is everything else: identity federation, automated user provisioning, audit logging, role-based access control, policy enforcement, and protection against attack vectors that API gateways from the previous decade were never designed to handle.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The proxy is roughly 5% of the actual scope. The rest is what makes it usable, governed, and defensible to your security team."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This distinction matters for procurement. Organizations that evaluate gateways primarily on latency benchmarks and integration counts are optimizing for the five percent. The ninety-five percent, whether the gateway can prove, in a form an auditor accepts, who did what, is what determines whether the deployment is actually enterprise-grade.&lt;/p&gt;

&lt;p&gt;The Composio MCP Gateway is designed around this reality. Rather than selling a proxy and calling it governance, it ships the full stack: 1,000+ managed integrations across enterprise SaaS, a unified authentication layer, action-level RBAC, zero data-retention architecture (tool call payloads and credentials are never stored on Composio infrastructure), and SOC 2 and ISO certification. The quickstart takes about ten minutes; the governance layer is built in from the start, not bolted on later.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1dbh6hq00ruqok0bcpb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1dbh6hq00ruqok0bcpb.png" alt="How An MCP Gateway Collapses" width="800" height="494"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The four things that cannot be missing
&lt;/h2&gt;

&lt;p&gt;Across the compliance frameworks that govern enterprise AI deployments (SOC 2, HIPAA, GDPR, ISO 27001, and now the EU AI Act), four governance capabilities appear repeatedly, either explicitly or implicitly. Absence of any one of them creates either regulatory exposure or operational failures that scale into incidents.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn25oypjkqt2tfskyh2lu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn25oypjkqt2tfskyh2lu.png" alt="mcp-governance-pillars" width="799" height="482"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Identity federation and SSO
&lt;/h3&gt;

&lt;p&gt;Without SSO integration, agents authenticate using shared service account credentials or locally-stored API keys. This creates credential sprawl, blocks user-level attribution in audit logs, and prevents IT from revoking access cleanly when an employee departs. With federated identity, every tool call carries the identity of the specific user who authorized it, flowing through the gateway from the enterprise identity provider down to the MCP server.&lt;/p&gt;

&lt;p&gt;The technical baseline is support for OAuth 2.1,  standardized in the MCP specification in June 2025, alongside SAML 2.0 for enterprise SSO and OpenID Connect for modern attribute mapping. &lt;/p&gt;

&lt;p&gt;But the capability that separates governance-capable gateways from identity-aware proxies is &lt;strong&gt;On-Behalf-Of (OBO) token propagation&lt;/strong&gt;: the pattern where a gateway passes the end-user identity downstream to the MCP server rather than substituting a service account. Without OBO, an audit log records "gateway service account called database write tool." With OBO, it records "Elena Mwangi in Finance called database write tool at 14:32 UTC." The difference is the difference between an audit log and an audit trail.&lt;/p&gt;

&lt;p&gt;Composio's MCP Gateway handles this through SSO via SAML and OIDC, with documented integrations for Okta, Microsoft Entra ID, and Google Workspace. Every team gets a unique, scoped MCP endpoint. Developers paste it into Claude, Cursor, or ChatGPT. SSO authenticates. Only the tools their team is authorized to use appear, and there is no separate configuration step to restrict visibility.&lt;/p&gt;

&lt;p&gt;One practical concern worth flagging: identity provider integrations that look stable can break silently. Microsoft Entra changed its attribute mapping behavior for synchronized users in late 2024 without a deprecation notice. Every such change is a potential gap in governance coverage. When evaluating any gateway, ask vendors specifically how they monitor for and respond to IdP-side breaking changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. SCIM provisioning
&lt;/h3&gt;

&lt;p&gt;SCIM — System for Cross-domain Identity Management — automates the user lifecycle at scale. New hires receive correct tool access on day one. Role changes propagate immediately to gateway permissions. Departing employees lose all access at the moment their directory account is disabled.&lt;/p&gt;

&lt;p&gt;Without SCIM, MCP gateway access management becomes a manual operation at every organizational boundary event. HIPAA requires that access to systems holding protected health information be revoked immediately upon role change or separation. SOC 2 CC6.2 requires that access be provisioned based on authorized requests and revoked promptly when no longer needed. Manual processes fail both tests at scale.&lt;/p&gt;

&lt;p&gt;The scenario that illustrates this most clearly: a developer departs on difficult terms. Legal advises IT to immediately revoke all access. IT disables the directory account. If SCIM is integrated, that change propagates to the gateway; every agent connection that developer had, from GitHub to Jira to Salesforce to internal APIs, terminates immediately. No gap exists between directory disabling and access revocation. Without SCIM, someone has to hunt and manually revoke individual credentials across every connected system. At any scale above a handful of users, some will be missed.&lt;/p&gt;

&lt;p&gt;Composio's SCIM 2.0 implementation maps directory groups to teams directly. The mapping logic is explicit and auditable: if &lt;code&gt;department = Engineering&lt;/code&gt; then &lt;code&gt;Team: engineering&lt;/code&gt;. New hires get the right tools on day one without any manual gateway configuration. The group sync is active and continuous, not a nightly batch job.&lt;/p&gt;

&lt;p&gt;For teams building toward this themselves: the build vs. buy analysis Composio published puts the engineering effort for SCIM provisioning at 4–8 weeks for a mid-sized team, before accounting for ongoing maintenance as IdP behavior changes. That estimate covers the SCIM endpoint, group sync logic, and conflict resolution. It does not cover the OAuth token lifecycle management that sits adjacent to it, which is typically another 4–8 weeks and carries higher ongoing maintenance cost.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Audit logging
&lt;/h3&gt;

&lt;p&gt;Audit logs answer the question every regulator and every security team will eventually ask: "what did your AI agents access, and when?" Without comprehensive, immutable, structured audit logs, the honest answer is "we don't know." That answer fails every compliance framework that governs regulated data.&lt;/p&gt;

&lt;p&gt;The minimum required fields per log entry are: timestamp in UTC at millisecond precision; user identity attributed through the IdP, not a service account; agent identity; MCP server and tool name invoked; tool input parameters; tool output or error state; authorization decision and the policy rule that produced it; and session identifier for multi-turn correlation. These fields are what make a log entry into evidence.&lt;/p&gt;

&lt;p&gt;Beyond minimum fields, enterprise-grade logs must be immutable after writing, tamper-evident, either through cryptographic signing or append-only storage. They must be structured for reliable SIEM ingestion. They must support configurable retention aligned to the organization's most demanding applicable requirement: HIPAA access records for protected health information require six-year retention; SOC 2 typically requires twelve months.&lt;/p&gt;

&lt;p&gt;Composio's audit trail logs every tool call as: user, team, tool, action, outcome. Critically, &lt;strong&gt;no payloads are stored&lt;/strong&gt; , only metadata. This zero data-retention architecture matters for regulated industries where storing tool call contents on third-party infrastructure creates its own compliance risk. The logs support CSV export for compliance reviews, and retention is configurable from 7 days to 1 year. The audit log format generates entries compliant with SOC 2, HIPAA, and GDPR requirements.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Policy enforcement
&lt;/h3&gt;

&lt;p&gt;The fourth pillar is where identity, provisioning, and audit turn from documentation tools into enforcement tools. Policy enforcement means the gateway doesn't just record that an agent attempted to call a destructive action. It blocks the call if the agent's role doesn't permit it.&lt;/p&gt;

&lt;p&gt;The critical implementation detail is the granularity at which access control operates. Standard RBAC in legacy API gateways operates at the API endpoint level. MCP gateway RBAC must operate at the action level within each toolkit. A GitHub integration may expose &lt;code&gt;GITHUB_CREATE_PR&lt;/code&gt;, &lt;code&gt;GITHUB_MERGE_PR&lt;/code&gt;, and &lt;code&gt;GITHUB_DELETE_REPO&lt;/code&gt;. Governance requires that a junior developer role can call the first two but not the third, without blocking access to the GitHub toolkit entirely.&lt;/p&gt;

&lt;p&gt;Composio enforces action-level RBAC at the gateway layer, not at the model layer. Each team gets a scoped MCP endpoint exposing only the tools they are authorized to use. Destructive actions within allowed toolkits — &lt;code&gt;GITHUB_DELETE_REPO&lt;/code&gt;, &lt;code&gt;SLACK_DELETE_CHANNEL&lt;/code&gt; — can be blocked independently of toolkit access. This is enforced in the gateway: if a model tries to call a blocked action, the gateway refuses it regardless of what the model was instructed to do.&lt;/p&gt;

&lt;p&gt;The access model supports both whitelist and blacklist modes. Teams can request access to blocked tools; admins approve or deny. This creates a self-service discovery path that doesn't require IT to anticipate every team's tooling needs in advance, while retaining central control over what actually gets enabled.&lt;/p&gt;




&lt;h2&gt;
  
  
  Attack vectors that API gateways were not built for
&lt;/h2&gt;

&lt;p&gt;Traditional API gateways were built for HTTP traffic between services. MCP traffic between AI agents and tool servers introduces attack vectors that legacy infrastructure was never designed to handle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool poisoning&lt;/strong&gt; places instructions inside tool Metadata, specifically in tool descriptions and parameter documentation that AI models read to understand how tools work. If descriptions contain adversarial instructions, the model may execute them. Unlike prompt injection, tool poisoning persists across sessions: it affects every agent that interacts with the tool, not just the session in which the attack was introduced.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rug pull attacks&lt;/strong&gt; are tool poisoning with a delayed trigger. A server publishes clean, vetted tool definitions at the time of security review. After approval, the operator modifies descriptions to inject malicious instructions. Without tool hash pinning, hashing tool descriptions on first scan and alerting when they change, the gap between approved state and live state can persist indefinitely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt injection via tool output&lt;/strong&gt; embeds adversarial instructions in tool outputs  (document contents, database records, web page responses) that the agent ingests as legitimate input. The MCP specification only "SHOULD" require a human in the loop, which is insufficient protection in production environments handling sensitive data at agent speed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cross-server shadowing&lt;/strong&gt; is an MCP-specific threat with no analog in traditional API security. A malicious MCP server impersonates a trusted server or embeds instructions in tool metadata that override the behavior of adjacent servers in the same agent context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Credential sprawl&lt;/strong&gt; is the most operationally common risk. Agents storing API keys, database passwords, and OAuth tokens in local configuration files create exposure through prompts, logs, or accidental repository commits. In multi-agent architectures, credentials propagate through chained tool calls in ways invisible without gateway-level telemetry.&lt;/p&gt;

&lt;p&gt;A security leader at Medtronic described the operational concern accurately: "MCP opens a lot of opportunities to do a lot of damage very quickly." The velocity at which autonomous agents can chain tool calls makes human review an insufficient backstop without gateway-level guardrails enforcing limits in real time.&lt;/p&gt;

&lt;p&gt;Composio's zero data-retention architecture addresses the credential sprawl risk directly: tool call payloads and credentials are never stored on Composio infrastructure. This eliminates the most common vector for credential exfiltration through the gateway layer itself.&lt;/p&gt;




&lt;h2&gt;
  
  
  What compliance frameworks actually require
&lt;/h2&gt;

&lt;p&gt;No compliance framework names MCP gateways explicitly. All of them implicitly require what a gateway provides: a centralized layer where AI tool access is governed, logged, and restricted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SOC 2&lt;/strong&gt; Trust Services Criteria CC6.1 through CC6.3 require access to be restricted to minimum necessary permissions, action-level RBAC satisfies this. CC7.2 and CC7.3 require monitoring and investigation of anomalies,  real-time audit log alerting and SIEM integration satisfy this. CC8.1 requires change management controls; access approval workflows and configurable retention policies satisfy this.&lt;/p&gt;

&lt;p&gt;For teams pursuing SOC 2 Type II certification, the observation period is at minimum six months. That means an organization that starts building its own gateway today won't have a reportable SOC 2 Type II audit for seven or eight months at the earliest, and that timeline assumes the controls were architected correctly from day one. Composio ships with SOC 2 Type II and ISO 27001 certification already in place, which removes this timeline entirely from the governance roadmap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;HIPAA&lt;/strong&gt; adds a harder requirement: Business Associate Agreements. Any vendor that creates, receives, maintains, or transmits protected health information on an organization's behalf is a Business Associate and legally requires a signed BAA before any PHI touches their infrastructure. Composio's enterprise plan supports BAA execution. For healthcare organizations, this is a binary filter that precedes all technical evaluation: verify BAA availability before spending time on feature comparison.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The EU AI Act&lt;/strong&gt;, whose high-risk system provisions became fully enforceable in August 2026, requires documented risk management, human oversight mechanisms, and technical evidence of controls for AI systems operating in healthcare, financial services, employment, and critical infrastructure. MCP gateway audit logs are the primary evidence artifact for conformity assessment. Organizations that have not established audit logging infrastructure before enforcement begins cannot retroactively generate evidence for the period before capture began.&lt;/p&gt;




&lt;h2&gt;
  
  
  The build vs. buy question, answered honestly
&lt;/h2&gt;

&lt;p&gt;Internal builds of MCP gateway infrastructure are a recurring theme in enterprise AI teams. The engineering argument is usually that "a proxy is a few weeks of work." That framing is accurate for the proxy. The full enterprise stack is different.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Build estimate&lt;/th&gt;
&lt;th&gt;Ongoing cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MCP routing proxy&lt;/td&gt;
&lt;td&gt;2–4 weeks&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OAuth 2.1 implementation&lt;/td&gt;
&lt;td&gt;3–6 weeks&lt;/td&gt;
&lt;td&gt;High — each SaaS app handles OAuth differently and changes without notice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SAML/OIDC IdP integration&lt;/td&gt;
&lt;td&gt;2–4 weeks&lt;/td&gt;
&lt;td&gt;Medium — silent breaking changes require active monitoring&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SCIM provisioning endpoint&lt;/td&gt;
&lt;td&gt;4–8 weeks&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-user OAuth token lifecycle&lt;/td&gt;
&lt;td&gt;4–8 weeks&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audit log infrastructure&lt;/td&gt;
&lt;td&gt;3–5 weeks&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Action-level RBAC policy engine&lt;/td&gt;
&lt;td&gt;6–8 weeks&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15 SaaS integrations&lt;/td&gt;
&lt;td&gt;~15 weeks&lt;/td&gt;
&lt;td&gt;Ongoing per-integration maintenance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SOC 2 Type II observation period&lt;/td&gt;
&lt;td&gt;6+ months&lt;/td&gt;
&lt;td&gt;Continuous&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The proxy is five percent of the scope. The OAuth maintenance burden is where most internal builds stall or quietly degrade over time: every SaaS application handles OAuth slightly differently, and those implementations change without notice. GitHub OAuth app permissions behave differently depending on whether the organization has SAML SSO enabled. Entra changed its attribute mapping behavior in late 2024 without a deprecation notice. Each change is a potential silent breakage.&lt;/p&gt;

&lt;p&gt;Buying wins for most teams because they are not buying a proxy, they are buying maintained integrations, per-user OAuth lifecycle management, SSO and SCIM support, RBAC enforcement, audit logging, and compliance readiness, with the maintenance burden sitting on the vendor rather than internal engineering. Composio's MCP Gateway developer quickstart gets a working agent connected to its first toolkit in about ten minutes. That's the realistic comparison point against a multi-month internal build.&lt;/p&gt;

&lt;p&gt;The cases where building makes sense are narrower: unique deployment constraints no vendor accommodates, classified network requirements, or organizations with the appetite to own the entire AI infrastructure stack as a long-term strategic investment.&lt;/p&gt;




&lt;h2&gt;
  
  
  How to evaluate a gateway honestly
&lt;/h2&gt;

&lt;p&gt;Start with deployment model. For organizations in healthcare, finance, or government where regulated data must remain within specific boundaries, deployment model is often a legal requirement before any technical comparison begins. Cloud-hosted managed gateways reduce time to production but involve data transiting vendor infrastructure. Self-hosted or VPC-deployed options provide data sovereignty. Composio operates as managed SaaS with a zero data-retention architecture as the default; for organizations requiring VPC or on-premises deployment, that narrows the field significantly and should be the first filter applied.&lt;/p&gt;

&lt;p&gt;After deployment model, evaluate in this sequence:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Identity depth.&lt;/strong&gt; Does the gateway support OBO token propagation, or does it substitute service accounts? Ask vendors for a sample audit log entry and verify that user identity is IdP-attributed, not a service account name.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SCIM implementation.&lt;/strong&gt; Does it support SCIM 2.0 with push provisioning? What is the documented maximum deprovisioning latency? The deprovisioning case, an employee departure or a security incident requiring immediate access revocation, is where manual processes fail most expensively.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Audit log quality.&lt;/strong&gt; Require vendors to provide a sample log entry with all fields populated. Confirm the format is structured and suitable for SIEM ingestion. Confirm logs are immutable after writing. Confirm the retention policy can be configured to your longest applicable requirement. Ask whether PII redaction in tool parameters is configurable and, in Composio's case, whether the zero data-retention architecture means payloads aren't stored at all, which is the stronger answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Access control granularity.&lt;/strong&gt; Confirm that RBAC operates at the action level, not the toolkit level. A gateway that blocks or enables whole toolkits but cannot distinguish between &lt;code&gt;GITHUB_CREATE_PR&lt;/code&gt; and &lt;code&gt;GITHUB_DELETE_REPO&lt;/code&gt; is not implementing least-privilege access control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compliance certification.&lt;/strong&gt; Request the current SOC 2 Type II report date and auditor. Confirm whether a BAA is available. For European deployments, ask whether the vendor has documented controls relevant to EU AI Act high-risk system provisions. Composio's SOC 2 and ISO 27001 certifications are current, which shortens the security review process significantly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MCP-specific threat coverage.&lt;/strong&gt; Ask whether tool hash pinning is implemented and whether it generates alerts when tool definitions change post-approval. Ask whether tool metadata is scanned for hidden prompt instructions. These questions distinguish purpose-built MCP governance platforms from extended API management products.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Exit terms.&lt;/strong&gt; Gateway choice shapes AI adoption architecture for three to five years. Confirm that gateway configuration, audit logs, and access policies can be exported in standard formats, and that contract exit terms do not create data portability barriers.&lt;/p&gt;




&lt;h2&gt;
  
  
  What comes next
&lt;/h2&gt;

&lt;p&gt;The MCP specification continues to evolve. Client ID Metadata Documents, added in the November 2025 spec update, introduce a new mechanism for trusted client discovery. The Agent-to-Agent protocol is emerging as a complement to MCP for multi-agent orchestration, governing agent-to-agent delegation rather than agent-to-tool connectivity. Future enterprise governance will require control planes spanning both protocols.&lt;/p&gt;

&lt;p&gt;As AI agents gain persistent memory and state across sessions, the audit and governance scope expands beyond tool calls to memory operations and state modifications. Gateways scoped only to tool call governance will require extension as these capabilities become standard.&lt;/p&gt;

&lt;p&gt;The broader trajectory is toward federated multi-gateway architectures: separate gateway instances per business unit or geographic region with centralized policy management. This pattern addresses data residency requirements without requiring monolithic governance infrastructure. Including A2A roadmap questions in current gateway evaluations is forward-looking work that belongs in any RFP issued in 2026.&lt;/p&gt;




&lt;h2&gt;
  
  
  Wrapping Up,
&lt;/h2&gt;

&lt;p&gt;The teams establishing MCP governance infrastructure now  (building audit trails, connecting identity providers, implementing SCIM provisioning, enforcing action-level access policies) are building the foundation for AI adoption that compliance teams can accept and auditors can verify. The teams deferring governance are accumulating technical debt measured not in refactoring effort but in regulatory exposure.&lt;/p&gt;

&lt;p&gt;The audit log for last quarter does not exist if it was never captured. The SOC 2 observation period clock does not start until you start running controls. The EU AI Act conformity evidence is not retroactively generatable. The compliance timeline is contracting, and the enforcement mechanisms are real.&lt;/p&gt;

&lt;p&gt;For most teams moving from pilot to production, the practical starting point is a managed gateway that handles the ninety-five percent — Composio's MCP Gateway covers the integrations, the OAuth lifecycle, the SCIM provisioning, the action-level RBAC, the audit logging, and the compliance certifications in a single product. The developer quickstart takes ten minutes. The governance is not an afterthought.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Further reading:&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://composio.dev/content/what-is-mcp-gateway-and-why-your-enterprise-need-it" rel="noopener noreferrer"&gt;&lt;em&gt;What is an MCP Gateway and why your enterprise needs one&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://composio.dev/content/building-vs-buying-an-enterprise-mcp-gateway" rel="noopener noreferrer"&gt;&lt;em&gt;Building vs. buying an enterprise MCP gateway&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://composio.dev/content/mcp-gateways-guide" rel="noopener noreferrer"&gt;&lt;em&gt;MCP Gateways: a developer's guide to AI agent architecture&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://composio.dev/content/best-mcp-gateway-for-developers" rel="noopener noreferrer"&gt;&lt;em&gt;10 best MCP gateways for developers in 2026&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>mcp</category>
      <category>beginners</category>
    </item>
    <item>
      <title>A Definitive Comparison Between Opencode &amp; Codex</title>
      <dc:creator>Developer Harsh</dc:creator>
      <pubDate>Fri, 03 Jul 2026 10:09:29 +0000</pubDate>
      <link>https://dev.to/composiodev/a-definitive-comparison-between-opencode-codex-dna</link>
      <guid>https://dev.to/composiodev/a-definitive-comparison-between-opencode-codex-dna</guid>
      <description>&lt;p&gt;If your daily workflow looks anything like mine, your terminal is where the actual work happens.&lt;/p&gt;

&lt;p&gt;After the &lt;a href="https://www.anthropic.com/engineering/april-23-postmortem" rel="noopener noreferrer"&gt;Claude Code fiasco&lt;/a&gt; back in April, I wanted a way out of Claude ecosystem. Codex and OpenCode were the default no-brainer choices.&lt;/p&gt;

&lt;p&gt;So I spent the last few months stress-testing Codex and OpenCode to see which one could actually replace Claude Code as my daily driver.&lt;/p&gt;

&lt;p&gt;So, here’s what I found out.&lt;/p&gt;




&lt;h2&gt;
  
  
  TL;DR: Quick Reference
&lt;/h2&gt;

&lt;p&gt;If you are in a hurry, this is the simplest way to think about the comparison.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Codex is the better default. OpenCode is the better power-user tool.&lt;/strong&gt; Codex wins when I want speed, polish, and fewer setup decisions. OpenCode wins when I want model freedom, lower cost, local execution, and more control over the agent loop.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Section&lt;/th&gt;
&lt;th&gt;Winner&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Onboarding, Setup, and Daily UX&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Codex&lt;/td&gt;
&lt;td&gt;Faster to start, cleaner defaults, easier daily workflow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Models&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tie&lt;/td&gt;
&lt;td&gt;Codex has the stronger default model stack; OpenCode has far more model freedom and with GLM 5.2 it’s on-par with GPT 5.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pricing / Cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;OpenCode&lt;/td&gt;
&lt;td&gt;Cheaper for heavy usage if you use routing, caching, or lower-cost models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Features and Workflows&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tie&lt;/td&gt;
&lt;td&gt;Codex is better for delegation; OpenCode is better for iterative local work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Ecosystem: MCP, Skills, Plugins&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Codex&lt;/td&gt;
&lt;td&gt;Simpler MCP and plugin setup; OpenCode is more transparent but more manual&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Harness Engineering&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tie&lt;/td&gt;
&lt;td&gt;Codex has the better default harness; OpenCode has the more customizable harness&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best overall for most users&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Codex&lt;/td&gt;
&lt;td&gt;Least friction, strongest defaults, smoother path from prompt to diff&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best overall for power users&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;OpenCode&lt;/td&gt;
&lt;td&gt;Model choice, local execution, deeper control, and better cost optimization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;My take:&lt;/strong&gt; I would recommend Codex to most users first. But for my own high-control workflow, OpenCode becomes more compelling over time because the extra setup turns into flexibility.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Onboarding, Setup, and Daily UX
&lt;/h2&gt;

&lt;p&gt;Onboarding and daily UX are too closely related to treat as separate sections.&lt;/p&gt;

&lt;p&gt;The first ten minutes decide how quickly I can start. The next ten days decide whether I actually want to keep using the tool. Codex wins the first part because it removes choices. OpenCode becomes more interesting later because the choices start turning into control.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Codex&lt;/th&gt;
&lt;th&gt;OpenCode&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Install speed&lt;/td&gt;
&lt;td&gt;~90 seconds, one path&lt;/td&gt;
&lt;td&gt;~3-5 minutes, more decisions&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First impression&lt;/td&gt;
&lt;td&gt;Polished, guided, low-friction&lt;/td&gt;
&lt;td&gt;Developer-native, terminal-first, configurable&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provider choice&lt;/td&gt;
&lt;td&gt;OpenAI only&lt;/td&gt;
&lt;td&gt;75+ providers and 1000+ models&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Configuration&lt;/td&gt;
&lt;td&gt;Minimal setup after sign-in&lt;/td&gt;
&lt;td&gt;API keys, model choice, working directory, config files&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learning curve&lt;/td&gt;
&lt;td&gt;Shallow; usable in minutes&lt;/td&gt;
&lt;td&gt;Moderate; rewards 1-2 months of use&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Daily workflow&lt;/td&gt;
&lt;td&gt;Open, assign task, review diff&lt;/td&gt;
&lt;td&gt;Plan, inspect, steer, execute, repeat&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Customization&lt;/td&gt;
&lt;td&gt;Opinionated defaults&lt;/td&gt;
&lt;td&gt;Deep control over models, instructions, and local setup&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Users who want the agent to stay out of the way&lt;/td&gt;
&lt;td&gt;Power users who want to tune the agent like a dev tool&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Codex&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;I installed Codex in about 90 seconds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm i &lt;span class="nt"&gt;-g&lt;/span&gt; @openai/codex
codex
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then it was basically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;sign in with ChatGPT,&lt;/li&gt;
&lt;li&gt;pick the project,&lt;/li&gt;
&lt;li&gt;start coding.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the whole appeal. The model is already selected, GitHub integration feels native, and the default workflow does not ask me to make too many decisions. I can open Codex, describe the task, review the diff, and move on.&lt;/p&gt;

&lt;p&gt;This matters because a daily coding agent should not make me think about the agent more than the code.&lt;/p&gt;

&lt;p&gt;Codex feels strongest when I need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a quick prototype before standup,&lt;/li&gt;
&lt;li&gt;a PR review,&lt;/li&gt;
&lt;li&gt;a clean diff for a narrow task,&lt;/li&gt;
&lt;li&gt;a background refactor,&lt;/li&gt;
&lt;li&gt;a low-friction path from prompt to patch.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tradeoff is that Codex is opinionated. I do not get much control over the model strategy, local runtime, or workflow shape. That is fine for most tasks, but limiting when I want to tune the agent like part of my dev environment.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;OpenCode&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;I installed OpenCode with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://opencode.ai/install | bash
opencode
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the decisions started:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which provider do I want?&lt;/li&gt;
&lt;li&gt;Do I want the Go tier?&lt;/li&gt;
&lt;li&gt;Which model should be the default?&lt;/li&gt;
&lt;li&gt;Which API keys do I need?&lt;/li&gt;
&lt;li&gt;Which working directory should it use?&lt;/li&gt;
&lt;li&gt;How much should I configure up front?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes OpenCode feel slower on day one. It is not the tool I would recommend to someone who hates setup decisions.&lt;/p&gt;

&lt;p&gt;But the same friction becomes useful once I understand the system. OpenCode gives me control over the parts Codex hides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I can switch providers and models based on task type,&lt;/li&gt;
&lt;li&gt;use local models through Ollama or LM Studio,&lt;/li&gt;
&lt;li&gt;inspect the plan before execution,&lt;/li&gt;
&lt;li&gt;steer the agent step by step,&lt;/li&gt;
&lt;li&gt;encode project preferences in instruction files,&lt;/li&gt;
&lt;li&gt;keep the loop close to my repo and tools.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This makes OpenCode feel less like a polished single-purpose coding agent and more like a configurable development environment.&lt;/p&gt;

&lt;p&gt;The downside is cognitive overhead. OpenCode asks me to participate more, and that is not always what I want for routine work. But for serious refactors, debugging sessions, or production changes where I want to watch the agent think before it acts, the extra control is worth the friction.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verdict
&lt;/h3&gt;

&lt;p&gt;Codex wins onboarding. OpenCode wins long-term control.&lt;/p&gt;

&lt;p&gt;If I am recommending a tool to a teammate who wants the least friction, I would recommend Codex. It is faster to start, easier to understand, and better for users who just want the agent to stay out of the way.&lt;/p&gt;

&lt;p&gt;If I am picking a tool for my own high-control workflow, OpenCode becomes more compelling over time. The setup is heavier, but the payoff is model flexibility, local execution, and tighter steering.&lt;/p&gt;

&lt;p&gt;For this section, Codex wins because the first-use and default daily experience are cleaner.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;OpenCode - 0, Codex - 1&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  2. Models: Codex vs OpenCode
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Codex&lt;/th&gt;
&lt;th&gt;OpenCode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model Availability&lt;/td&gt;
&lt;td&gt;GPT-5.5 only&lt;/td&gt;
&lt;td&gt;~75 providers, 1000+ models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token Efficiency&lt;/td&gt;
&lt;td&gt;Optimized for GPT-5.5&lt;/td&gt;
&lt;td&gt;40-60% fewer tokens (MiMo)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model Switching&lt;/td&gt;
&lt;td&gt;Single model, all tasks&lt;/td&gt;
&lt;td&gt;Switch between models per task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Top Performers&lt;/td&gt;
&lt;td&gt;GPT-5.5 (58.6%)&lt;/td&gt;
&lt;td&gt;Qwen 3.7 (60.6%), MiMo-V2.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost Per Token&lt;/td&gt;
&lt;td&gt;$30-180 per million tokens&lt;/td&gt;
&lt;td&gt;Varies; DeepSeek $0.14-0.28&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best For&lt;/td&gt;
&lt;td&gt;Best-in-class performance&lt;/td&gt;
&lt;td&gt;Cost-conscious, flexible workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Codex&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The first time I ran Codex with GPT-5.5, it felt like the whole system was purpose-built around it.&lt;/p&gt;

&lt;p&gt;OpenAI’s headline is “better results with fewer tokens.” The more interesting story is &lt;em&gt;how&lt;/em&gt; they got there: Codex is a tightly tuned pipeline where the prompts, context management, tool-calling, and evaluation loop are all optimized for GPT models. This is similar to Claude &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;OpenAI designed GPT-5.5 specifically for agentic coding, then adjusted Codex to leverage its full capabilities. GPT-5.5 uses 40% fewer output tokens than GPT-5.4 on the same Codex tasks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every task I run through Codex uses this same tuned pipeline. It's like having a senior engineer trained specifically for your workflow, focused on results, rather than decisions&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;OpenCode&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;OpenCode provides integration with ~75 different providers across 1000+ models, and one might be intimidated by the cost they would incur. I had the same.  &lt;/p&gt;

&lt;p&gt;But as I  looked at benchmark data, I found something: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Qwen 3.7 maxes out at 60.6% on SWE-Bench Pro,&amp;nbsp;beating GPT-5.5's 58.6%.&lt;/li&gt;
&lt;li&gt;MiMo-V2.5-Pro uses 40-60% fewer tokens than GPT-5.4 for comparable output.&lt;/li&gt;
&lt;li&gt;DeepSeek V4-Flash costs $0.14 per million tokens for input / $0.28 for output, compared to $30 per million tokens&amp;nbsp;for input / $180 per million tokens for&amp;nbsp;output for GPT-5.5.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The hidden insight: &lt;strong&gt;I don't need the same model for every task.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Architecture decisions: Qwen.&lt;/li&gt;
&lt;li&gt;Boilerplate: DeepSeek.&lt;/li&gt;
&lt;li&gt;Bug fixing: MiMo.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want automation, you can connect OpenCode with&amp;nbsp;smart model routers as well; they will do the heavy lifting. &lt;/p&gt;

&lt;p&gt;This was the learning curve I was talking about earlier: model routing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verdict
&lt;/h3&gt;

&lt;p&gt;If you ask me: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT 5.5 is undeniably the better model than anything open-source can offer right now. Though Kimi 2.7 and GLM 5.2 are great models with near SOTA coding performance.&lt;/li&gt;
&lt;li&gt;OpenCode definitely gives the freedom to select any model one wants, plus at a lower cost. For cost-conscious people, this is definitely a USP.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Codex with GPT 5.5 and OpenCode with GLM 5.2 are match made in labs. So, at this point, it’s tie.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;OpenCode - 1,  Codex - 2&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  3. Pricing / Cost: Codex vs OpenCode
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Codex&lt;/th&gt;
&lt;th&gt;OpenCode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Entry Price&lt;/td&gt;
&lt;td&gt;Plus at $20/month&lt;/td&gt;
&lt;td&gt;Go tier at $10/month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Professional Cost&lt;/td&gt;
&lt;td&gt;$100-200/month&lt;/td&gt;
&lt;td&gt;$10-50/month (with routing)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost Savings&lt;/td&gt;
&lt;td&gt;No optimization options&lt;/td&gt;
&lt;td&gt;~70% reduction with smart routing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token Caching&lt;/td&gt;
&lt;td&gt;Limited caching&lt;/td&gt;
&lt;td&gt;Built-in, reduces cost ~70%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pricing Model&lt;/td&gt;
&lt;td&gt;Monthly subscription fixed&lt;/td&gt;
&lt;td&gt;Pay per token (variable)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best For&lt;/td&gt;
&lt;td&gt;Predictable monthly budgets&lt;/td&gt;
&lt;td&gt;Budget-conscious developers&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Codex&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Codex comes bundled with ChatGPT Plus at $20/month, which sounds cheap until you start using it heavily.&lt;/p&gt;

&lt;p&gt;Here's my actual usage pattern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Lightweight tasks: 2-3 sessions/day (covers with Plus)&lt;/li&gt;
&lt;li&gt;Serious refactoring: 4-7 hours/day (exhausts Plus)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When I upgraded to Pro ($100/month), things got a  little smoother. I never hit limits. But I'm now paying $1,200/year for what I actually use.&lt;/p&gt;

&lt;p&gt;That’s not a number; it's the real cost for a professional who codes 6+ hours/day, which is around&amp;nbsp;&lt;strong&gt;$100-$200/month&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;OpenCode&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;OpenCode Go is $10/month or less, but only if you actually need to figure out which models to use for which tasks.&lt;/p&gt;

&lt;p&gt;Here's my actual usage pattern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Day 1: Confused about model selection (burning tokens on wrong model choices)&lt;/li&gt;
&lt;li&gt;Day 10: I figured out routing: Boilerplate → one model, Architecture → another, token cost starts dropping&lt;/li&gt;
&lt;li&gt;Day 30: Smart routing is dialed in (DeepSeek for routine, Qwen for complex, local models for edge cases), making costs fixed around $10/month tier&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When I finally cracked the model-routing puzzle by month 2, I realized the real hidden advantage:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Cached tokens cost a fraction of the normal price. So my $0.50/session cost was actually closer to $0.15 with caching baked in.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;According to the estimate, the real cost for a professional with smart routing is around &lt;strong&gt;$10-$50/month&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That’s a ~70% deduction and makes switching non-negotiable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verdict
&lt;/h3&gt;

&lt;p&gt;Clearly, Open Code wins on this one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;OpenCode - 2 , Codex - 2&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  3. Features and Workflows
&lt;/h2&gt;

&lt;p&gt;This is where Codex and OpenCode start to feel like fundamentally different products.&lt;/p&gt;

&lt;p&gt;Codex is built around &lt;strong&gt;delegation&lt;/strong&gt;. OpenCode is built around &lt;strong&gt;iteration&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Codex&lt;/th&gt;
&lt;th&gt;OpenCode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Core workflow&lt;/td&gt;
&lt;td&gt;Define goal → delegate → review result&lt;/td&gt;
&lt;td&gt;Plan → review → execute → adjust&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best interaction style&lt;/td&gt;
&lt;td&gt;High-level task assignment&lt;/td&gt;
&lt;td&gt;Tight local feedback loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Goal setting&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;/goal&lt;/code&gt; command for scoped outcomes&lt;/td&gt;
&lt;td&gt;Plan mode + repo instructions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Iteration speed&lt;/td&gt;
&lt;td&gt;Better for longer background tasks&lt;/td&gt;
&lt;td&gt;Better for fast back-and-forth changes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local capability&lt;/td&gt;
&lt;td&gt;Cloud-first&lt;/td&gt;
&lt;td&gt;Local-first with Ollama/LM Studio support&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Real-time control&lt;/td&gt;
&lt;td&gt;Review changes after the agent runs&lt;/td&gt;
&lt;td&gt;Review and steer before execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Overnight refactors, PR prep, delegated work&lt;/td&gt;
&lt;td&gt;Interactive development, debugging, learning&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Codex&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Codex feels strongest when I treat it like an engineering teammate I can delegate to.&lt;/p&gt;

&lt;p&gt;The app lets me set up multi-agent workflows for longer-running execution:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One agent reviews PRs,&lt;/li&gt;
&lt;li&gt;another fixes bugs,&lt;/li&gt;
&lt;li&gt;a third updates documentation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I close my laptop and come back to the result. That makes Codex especially good for large refactors, GitHub-native workflows, team delegation, and background engineering work.&lt;/p&gt;

&lt;p&gt;The underrated feature here is Codex’s &lt;code&gt;/goal&lt;/code&gt; command. Instead of giving the agent a vague task like “improve this repo,” I can define the actual outcome I want:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;reduce flaky tests,&lt;/li&gt;
&lt;li&gt;migrate a module,&lt;/li&gt;
&lt;li&gt;clean up auth logic,&lt;/li&gt;
&lt;li&gt;prepare a PR-ready refactor.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Codex then uses that goal as the anchor for planning, execution, and review. That makes long-running delegated work feel less like prompting and more like assigning a scoped engineering objective.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;OpenCode&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;OpenCode does not have a direct &lt;code&gt;/goal&lt;/code&gt; equivalent, but its workflow solves the same problem differently.&lt;/p&gt;

&lt;p&gt;Instead of asking me to assign a goal and wait for the result, OpenCode keeps me inside a tight loop:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;define what I want,&lt;/li&gt;
&lt;li&gt;inspect the proposed plan,&lt;/li&gt;
&lt;li&gt;adjust the approach,&lt;/li&gt;
&lt;li&gt;execute,&lt;/li&gt;
&lt;li&gt;review the result,&lt;/li&gt;
&lt;li&gt;repeat.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is where Plan mode becomes important. It gives me a goal-like workflow without hiding the intermediate reasoning. I can see what OpenCode intends to do before it touches the codebase, which is useful when I am debugging, exploring unfamiliar code, or doing refactors where I want control over every step.&lt;/p&gt;

&lt;p&gt;OpenCode also pairs well with repo-level instruction files like &lt;code&gt;AGENTS.md&lt;/code&gt;. That makes its goal-setting less polished than Codex’s &lt;code&gt;/goal&lt;/code&gt;, but more customizable. I can encode project conventions, testing expectations, architecture rules, and workflow preferences once, then reuse them across sessions.&lt;/p&gt;

&lt;p&gt;The other major advantage is local execution. I can pair OpenCode with Ollama or LM Studio and run the agentic loop on my own machine with zero API calls. For security-sensitive work, regulated codebases, or local-first development, this is a real advantage.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verdict
&lt;/h3&gt;

&lt;p&gt;This one depends on how I want to work.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Codex wins for delegation:&lt;/strong&gt; give it a scoped objective, let it run, and review the result later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenCode wins for iteration:&lt;/strong&gt; inspect the plan, steer the agent, and keep the feedback loop tight.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Codex feels more polished. OpenCode feels more controllable.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For routine background work, I prefer Codex. For interactive development and learning inside a codebase, I prefer OpenCode.&lt;/p&gt;

&lt;p&gt;Tie.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;OpenCode - 3, Codex - 3&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  4. Ecosystem (MCP + Skills +  Plugins)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Codex&lt;/th&gt;
&lt;th&gt;OpenCode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MCP Setup&lt;/td&gt;
&lt;td&gt;CLI commands (&lt;code&gt;codex mcp add&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Manual config via &lt;code&gt;.opencode/mcp-config.json&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skill Installation&lt;/td&gt;
&lt;td&gt;Git clone to &lt;code&gt;~/.codex/skills/&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Clone to &lt;code&gt;~/.opencode/skills/&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plugin Management&lt;/td&gt;
&lt;td&gt;Marketplace CLI integration&lt;/td&gt;
&lt;td&gt;Update &lt;code&gt;opencode.json&lt;/code&gt; manually&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Composio Integration&lt;/td&gt;
&lt;td&gt;One-click via marketplace&lt;/td&gt;
&lt;td&gt;Config file + manual setup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User Friendliness&lt;/td&gt;
&lt;td&gt;More convenient, less transparent&lt;/td&gt;
&lt;td&gt;More transparent, less convenient&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best For&lt;/td&gt;
&lt;td&gt;Users who want simplicity&lt;/td&gt;
&lt;td&gt;Developers who like transparency&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;You can have the best model, the best providers, and the best features and workflow, yet it means nothing if your models can’t talk to the real world and perform specified tasks in specified ways. &lt;/p&gt;

&lt;p&gt;Codex and OpenCode both offer: MCP, Plugin &amp;amp; Skills, but both function differently.&lt;/p&gt;

&lt;h3&gt;
  
  
  Codex
&lt;/h3&gt;

&lt;p&gt;Codex supports MCP integration. This is how easy it is to install:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I am going with Composio, as I usually use multiple MCP servers, and it's a pain to connect to and configure each one securely and to make agents handle multiple tool calls intelligently.&lt;/p&gt;


&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Add Composio MCP server to Codex&lt;/span&gt;
codex mcp add composio

&lt;span class="c"&gt;# Authenticate&lt;/span&gt;
codex mcp auth composio
&lt;span class="c"&gt;# Opens browser for OAuth&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Verify it's connected:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;codex mcp list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now, to make sure the MCP works properly, you can add skills with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ~/.codex/skills
git clone https://github.com/ComposioHQ/awesome-codex-skills.git ~/.codex/skills/composio-connect
&lt;span class="c"&gt;# Restart Codex&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can also add the Composio plugin using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;codex plugin marketplace add ComposioHQ/awesome-codex-plugins
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And restart the app:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;codex
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But to do the same in OpenCode is a little tricky.&lt;/p&gt;

&lt;h4&gt;
  
  
  Open Code
&lt;/h4&gt;

&lt;p&gt;OpenCode also supports MCP integration, but to add any MCP server, you need to update the config at &lt;code&gt;.opencode/mcp-config.json&lt;/code&gt; .&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# .opencode/mcp-config.json&lt;/span&gt;
&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"mcp_servers"&lt;/span&gt;: &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="s2"&gt;"composio"&lt;/span&gt;: &lt;span class="o"&gt;{&lt;/span&gt;
      &lt;span class="s2"&gt;"type"&lt;/span&gt;: &lt;span class="s2"&gt;"remote"&lt;/span&gt;,
      &lt;span class="s2"&gt;"url"&lt;/span&gt;: &lt;span class="s2"&gt;"https://connect.composio.dev/mcp"&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
  &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Certainly not the most friendly interface, but good for transparency, as you can see what goes into the MCP server.&lt;/p&gt;

&lt;p&gt;Next, add skills:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/ComposioHQ/awesome-codex-skills ~/.opencode/skills/composio
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Restart OpenCode&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;opencode
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This works because OpenCode looks for skills in project and global locations, including &lt;code&gt;.opencode/skills&lt;/code&gt;, &lt;code&gt;~/.config/opencode/skills&lt;/code&gt;, &lt;code&gt;.claude/skills&lt;/code&gt;, and &lt;code&gt;.agents/skills&lt;/code&gt; .&lt;/p&gt;

&lt;p&gt;You can also add the Composio plugin:&lt;/p&gt;

&lt;p&gt;Add to &lt;code&gt;opencode.json&lt;/code&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"plugin"&lt;/span&gt;: &lt;span class="o"&gt;[&lt;/span&gt;
    &lt;span class="s2"&gt;"opencode-composio"&lt;/span&gt;,
    &lt;span class="s2"&gt;"opencode-context7"&lt;/span&gt;
  &lt;span class="o"&gt;]&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Save and restart OpenCode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;opencode
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Done!&lt;/p&gt;

&lt;h3&gt;
  
  
  Verdict
&lt;/h3&gt;

&lt;p&gt;So Codex wins here due to process simplicity.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Open Code - 3 , Codex - 4&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  5. Harness Engineering
&lt;/h2&gt;

&lt;p&gt;The model matters, but the harness decides how that model sees the repo, plans changes, calls tools, handles errors, and recovers when something breaks. In practice, the harness is the difference between “the model is smart” and “the agent is reliable.”&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Codex&lt;/th&gt;
&lt;th&gt;OpenCode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Implementation&lt;/td&gt;
&lt;td&gt;Rust-based, performance-focused CLI/app stack&lt;/td&gt;
&lt;td&gt;TypeScript core with Tauri desktop app&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Design philosophy&lt;/td&gt;
&lt;td&gt;Tightly optimized around OpenAI models&lt;/td&gt;
&lt;td&gt;Provider-agnostic and modular by design&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context handling&lt;/td&gt;
&lt;td&gt;Strong default repo understanding with fewer choices&lt;/td&gt;
&lt;td&gt;More explicit control over model, context, and instructions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool execution&lt;/td&gt;
&lt;td&gt;Permission profiles, hooks, sandboxed/cloud execution&lt;/td&gt;
&lt;td&gt;Local execution with permission gates and config-level control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feedback loop&lt;/td&gt;
&lt;td&gt;Optimized prompting, planning, and tool-calling pipeline&lt;/td&gt;
&lt;td&gt;LSP diagnostics fed back into the agent loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strength&lt;/td&gt;
&lt;td&gt;Speed, polish, and low-friction execution&lt;/td&gt;
&lt;td&gt;Control, transparency, and production thoroughness&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tradeoff&lt;/td&gt;
&lt;td&gt;Less model/harness customization&lt;/td&gt;
&lt;td&gt;More setup and slower execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Fast implementation and delegated engineering tasks&lt;/td&gt;
&lt;td&gt;Complex refactors where correctness matters more than speed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Codex&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Codex feels like a vertically integrated agent stack.&lt;/p&gt;

&lt;p&gt;The model, prompt format, context strategy, tool-calling behavior, permission model, and review flow all feel designed to work together. That is the advantage of a closed, OpenAI-first harness: fewer knobs, fewer setup decisions, and fewer ways to misconfigure the system.&lt;/p&gt;

&lt;p&gt;The strongest part is how little I have to think about the plumbing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;permission profiles decide what the agent can touch,&lt;/li&gt;
&lt;li&gt;hooks let me run pre- and post-execution checks,&lt;/li&gt;
&lt;li&gt;GitHub and PR workflows feel native,&lt;/li&gt;
&lt;li&gt;tool calls are routed through a polished approval flow,&lt;/li&gt;
&lt;li&gt;cloud execution keeps risky changes away from my local machine until review.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything is tuned around GPT-5.5. That matters because Codex is not just calling a model; it is shaping how the model receives the repo, plans the task, executes commands, and presents diffs back to me.&lt;/p&gt;

&lt;p&gt;This is why Codex often feels faster than a generic agent using the same model. The harness reduces wasted motion. It does not ask me to design the workflow first; it gives me a working default and lets me move.&lt;/p&gt;

&lt;p&gt;The downside is that this optimization comes with a ceiling. If I want to change the model strategy, deeply customize the execution loop, or route different tasks through different providers, Codex gives me much less room to experiment.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;OpenCode&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;OpenCode takes the opposite bet.&lt;/p&gt;

&lt;p&gt;Instead of optimizing one model inside one polished workflow, it gives you a modular harness that can work across providers, models, local runtimes, MCP servers, and repo-level instructions. It is less “batteries included,” but much more inspectable.&lt;/p&gt;

&lt;p&gt;The most important engineering choice is the feedback loop. OpenCode can feed Language Server Protocol diagnostics back into the agent while it works. If the agent introduces a TypeScript error, the next step can include that error as context, so the model has a chance to self-correct before I even review the final diff.&lt;/p&gt;

&lt;p&gt;That changes the feel of the tool. OpenCode may be slower, but it often behaves more like an engineer working with compiler feedback, not just a chatbot editing files.&lt;/p&gt;

&lt;p&gt;It also gives me more control over the harness itself:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I can switch providers and models based on task type,&lt;/li&gt;
&lt;li&gt;keep project-specific behavior in &lt;code&gt;AGENTS.md&lt;/code&gt;,&lt;/li&gt;
&lt;li&gt;run locally with Ollama or LM Studio,&lt;/li&gt;
&lt;li&gt;wire in MCP tools manually,&lt;/li&gt;
&lt;li&gt;inspect config instead of trusting a black box.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why OpenCode tends to feel better for production refactors. The loop is tighter, the configuration is more visible, and the agent can use local development signals instead of only relying on the initial prompt and repo context.&lt;/p&gt;

&lt;p&gt;The tradeoff is obvious: more control means more responsibility. If the model choice is bad, the config is messy, or the repo instructions are vague, OpenCode will not hide that complexity from me.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verdict
&lt;/h3&gt;

&lt;p&gt;Codex has the better &lt;strong&gt;default harness&lt;/strong&gt;. OpenCode has the better &lt;strong&gt;customizable harness&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Codex wins on speed and polish:&lt;/strong&gt; it is optimized end-to-end for OpenAI models and gets me to a usable diff quickly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenCode wins on control and feedback:&lt;/strong&gt; LSP diagnostics, local execution, and provider flexibility make it stronger for careful refactors.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Codex abstracts the harness away. OpenCode exposes the harness and lets you tune it.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a quick implementation, I would pick Codex. For a high-stakes refactor where I want visibility into every step, I would pick OpenCode.&lt;/p&gt;

&lt;p&gt;This one is a tie, but for very different reasons.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;OpenCode - 4 , Codex - 5&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The Final Verdict: When To Choose What
&lt;/h2&gt;

&lt;p&gt;Clearly, OpenCode is the winner with 6 points, but real engineers leverage both for their specific needs :&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Codex for speed, overnight refactors, and production-critical work.&lt;/li&gt;
&lt;li&gt;OpenCode for smart model routing, optimized costs, and offline critical workflows.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A simple table summarizes them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Codex&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;OpenCode&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;OpenAI ecosystem&lt;/td&gt;
&lt;td&gt;Cost control, model flexibility&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Setup&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Zero friction, bundled into ChatGPT subscripton&lt;/td&gt;
&lt;td&gt;Configure providers; slight model usage learning curve&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Autonomous work&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Cloud agent, good for overnight refactors&lt;/td&gt;
&lt;td&gt;Terminal agent; depends on your model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Integrations&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;GitHub, PR review, Slack&lt;/td&gt;
&lt;td&gt;MCP; varies by setup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Model choice&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;GPT-5 only&lt;/td&gt;
&lt;td&gt;75+ providers; Claude via API key only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Offline&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes, with Ollama/LM Studio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Transparency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Token-based credits&lt;/td&gt;
&lt;td&gt;Full model + token visibility&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Real cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$20–$200/mo&lt;/td&gt;
&lt;td&gt;Free BYOK, or ~$10–$50/mo routing&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With a few months of usage, one thing is clear to me,&lt;/p&gt;

&lt;p&gt;Choosing Codex or Opencode models is not about which benchmarks perform better; it's about picking the one that matches your workflow. Both are good in their own right, and best leveraged based on the needs.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>Claude Code vs. OpenCode without the hype</title>
      <dc:creator>Shrijal Acharya</dc:creator>
      <pubDate>Thu, 21 May 2026 13:55:18 +0000</pubDate>
      <link>https://dev.to/composiodev/claude-code-vs-opencode-without-the-hype-j1f</link>
      <guid>https://dev.to/composiodev/claude-code-vs-opencode-without-the-hype-j1f</guid>
      <description>&lt;p&gt;Everyone wants a coding agent now.&lt;/p&gt;

&lt;p&gt;Not a chatbot that explains code.&lt;/p&gt;

&lt;p&gt;An actual agent that can read your repo, edit files, run commands, use tools, and keep moving while you supervise.&lt;/p&gt;

&lt;p&gt;Claude Code and OpenCode are two of the most interesting takes on that idea.&lt;/p&gt;

&lt;p&gt;Claude Code is the polished Anthropic-native route.&lt;/p&gt;

&lt;p&gt;OpenCode is the open-source route for people who want more model choice, more control, and a setup they can tweak.&lt;/p&gt;

&lt;p&gt;And that difference matters more than it looks.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5bcj3t5mypqia64m30wr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5bcj3t5mypqia64m30wr.png" alt="distracted man GIF" width="687" height="361"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What is OpenCode
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;ℹ️ Open-source coding agent with model and tool control&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2k02x3jygu2mwj40jz44.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2k02x3jygu2mwj40jz44.png" alt="OpenCode" width="799" height="410"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;OpenCode is an open-source coding agent for developers who want more control over their AI coding setup.&lt;/p&gt;

&lt;p&gt;It runs in the terminal, IDE, and desktop, and lets you bring your own model instead of &lt;strong&gt;locking you into one provider&lt;/strong&gt;. Claude, GPT, Gemini, local models, and 75+ other providers are supported. That is probably the biggest reason people care about it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fkem0kiznc0fm4g9gyvw9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fkem0kiznc0fm4g9gyvw9.png" alt="OpenCode tweet" width="799" height="439"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It also comes with the things you expect from a serious coding agent now: LSP support, multi-session workflows, project memory through &lt;code&gt;AGENTS.md&lt;/code&gt;, MCP tools, custom agents, plugins, and editor support, and maybe a bunch more.&lt;/p&gt;

&lt;p&gt;So the pitch is not just “AI in your terminal.”&lt;/p&gt;

&lt;p&gt;That undersells it.&lt;/p&gt;

&lt;p&gt;OpenCode is closer to a &lt;strong&gt;coding-agent workbench&lt;/strong&gt;. You bring the model, the provider, the editor, the agents, and the workflow. OpenCode gives you the open layer that ties it all together.&lt;/p&gt;

&lt;p&gt;Not everyone needs that level of control.&lt;/p&gt;

&lt;p&gt;But some developers absolutely do.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💁 OpenCode is for developers who want to tweak every single detail of their coding agent.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is what makes it interesting next to Claude Code.&lt;/p&gt;




&lt;h2&gt;
  
  
  What is Claude Code
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;ℹ️ Anthropic’s polished coding agent for your terminal.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fivw9mx75nhkdloyczjr5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fivw9mx75nhkdloyczjr5.png" alt="Claude Code" width="800" height="208"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Claude Code is Anthropic’s coding agent that lives in your terminal.&lt;/p&gt;

&lt;p&gt;The idea is pretty same here, it can read your codebase, edit files, run commands, handle Git stuff, and all through prompts.&lt;/p&gt;

&lt;p&gt;The big difference is that Claude Code is built around Claude.&lt;/p&gt;

&lt;p&gt;That sounds obvious, but it matters.&lt;/p&gt;

&lt;p&gt;You are not coming here to mix and match ten different model providers. You are coming here because you trust Anthropic’s models, and you want the cleanest experience around them.&lt;/p&gt;

&lt;p&gt;Claude Code also comes with a lot of serious agent features: project memory through &lt;code&gt;CLAUDE.md&lt;/code&gt;, slash commands, permissions, hooks, MCP, plugins, custom subagents, and IDE integrations.&lt;/p&gt;

&lt;p&gt;Claude Code is closer to a Claude-native coding environment. The model, the agent loop, the tool use, the permissions, and the workflow all come from the same Anthropic-shaped box.&lt;/p&gt;

&lt;p&gt;Less DIY.&lt;/p&gt;

&lt;p&gt;But there is also a small shift happening.&lt;/p&gt;

&lt;p&gt;Some developers are starting to move from Claude Code to OpenCode or OpenAI’s Codex for one simple reason: &lt;strong&gt;usage limits&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fi4uxa5o6uxb1s9x87cgw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fi4uxa5o6uxb1s9x87cgw.png" alt="Claude Code Usage Limit meme" width="800" height="830"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Claude Code is great, but when you are deep in a coding session, hitting limits feels brutal. And for heavier users, even the &lt;strong&gt;$200 Claude Max plan&lt;/strong&gt; does not always feel like enough.&lt;/p&gt;

&lt;p&gt;That is why OpenCode and Codex are tempting. Also read: &lt;a href="https://composio.dev/content/claude-code-vs-openai-codex" rel="noopener noreferrer"&gt;Claude Code vs. Codex: Detailed breakdown&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When Claude hits the wall, people still need a way to keep shipping.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💁 If you're an Anthropic fanboy, and don't care about other models, stick to Claude Code.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  High Level Architecture
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;ℹ️ How both the agents work&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;At a high level, both OpenCode and Claude Code follow the same basic agent loop.&lt;/p&gt;

&lt;p&gt;You give it a task.&lt;/p&gt;

&lt;p&gt;It looks at the repo.&lt;/p&gt;

&lt;p&gt;It decides what files, commands, or tools it needs.&lt;/p&gt;

&lt;p&gt;It takes an action.&lt;/p&gt;

&lt;p&gt;Then it reads the result and keeps going.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fcv9tejzbqvy4whr7jcj2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fcv9tejzbqvy4whr7jcj2.png" alt="Coding Agent architecture" width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;ℹ️ This is the highest-level architecture of a coding agent. A few details change from tool to tool.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That loop is the boring part.&lt;/p&gt;

&lt;p&gt;The interesting part is everything around it.&lt;/p&gt;

&lt;p&gt;Here is a tiny example of that loop in practice.&lt;/p&gt;

&lt;p&gt;I gave Claude Code and OpenCode the same small task in a demo word-count repo:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Add a &lt;code&gt;--json&lt;/code&gt; flag to a word-count CLI, update the tests, and run them.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The interesting part is not the feature. It is watching both agents go through the same shape: understand the repo, plan the change, edit the files, and run the tests.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/D74fsmbwE98"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Claude Code wraps that loop in Anthropic’s own product system. You get Claude, project memory through &lt;code&gt;CLAUDE.md&lt;/code&gt;, permissions, hooks, MCP, &lt;a href="https://composio.dev/content/top-claude-code-plugins" rel="noopener noreferrer"&gt;plugins&lt;/a&gt;, &lt;a href="https://composio.dev/content/top-claude-skills" rel="noopener noreferrer"&gt;Claude skills&lt;/a&gt;, and subagents in one single setup.&lt;/p&gt;

&lt;p&gt;OpenCode takes a more open route. It gives you the agent runtime, but lets you bring different models, providers, agents, tools, and workflows. Its docs split agents into primary agents and subagents, and let you configure specialized assistants with custom prompts, models, and tool access.&lt;/p&gt;

&lt;p&gt;So architecturally, the difference is not that one is an agent and the other is not.&lt;/p&gt;

&lt;p&gt;They both are.&lt;/p&gt;

&lt;p&gt;The real difference is who controls the harness around the agent.&lt;/p&gt;

&lt;p&gt;Claude Code gives you Anthropic’s harness.&lt;/p&gt;

&lt;p&gt;OpenCode gives you a harness you can inspect, and configure.&lt;/p&gt;




&lt;h2&gt;
  
  
  Context, memory and tool use
&lt;/h2&gt;

&lt;p&gt;Both Claude Code and OpenCode are doing the same basic thing: they build a giant prompt, stuff it with repo context, tool definitions, memory files, recent messages, and tool results, then ask the model what to do next.&lt;/p&gt;

&lt;p&gt;The difference is how much of that system you control.&lt;/p&gt;

&lt;p&gt;Claude Code is more vertically integrated here. It is built around Anthropic models, so it can take advantage of Anthropic-specific stuff like prompt caching, native tool calls, and Claude’s own long-context behavior.&lt;/p&gt;

&lt;p&gt;That matters.&lt;/p&gt;

&lt;p&gt;Tool definitions, system prompts, and &lt;code&gt;CLAUDE.md&lt;/code&gt; can be cached between turns, which makes long coding sessions cheaper and faster than they would be if Claude had to re-read everything from scratch every single time.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fphuqzyrqslhkmlnh7ah3.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fphuqzyrqslhkmlnh7ah3.jpg" alt="compaction in a coding agent" width="800" height="407"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;OpenCode takes a different route.&lt;/p&gt;

&lt;p&gt;It does not assume one model or one provider. Instead, it &lt;strong&gt;reads the model’s context limit&lt;/strong&gt; from the provider metadata and builds the session around that. So the same OpenCode setup can run with Claude, GPT, Gemini, Qwen, local models, or whatever else you plug in.&lt;/p&gt;

&lt;p&gt;That flexibility is the whole point.&lt;/p&gt;

&lt;p&gt;But it also means OpenCode has to normalize all the weird provider differences: tool call IDs, cache support, model limits, and tool-calling parts.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F75v9hwsvotu1qmz1k8p0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F75v9hwsvotu1qmz1k8p0.png" alt="OpenCode support for multiple providers" width="799" height="255"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Claude Code gets to optimize deeply for Claude.&lt;/p&gt;

&lt;p&gt;OpenCode has to work with everyone.&lt;/p&gt;

&lt;p&gt;Memory works the same way.&lt;/p&gt;

&lt;p&gt;Claude Code uses &lt;code&gt;CLAUDE.md&lt;/code&gt; as the main project memory file. It can also load nested &lt;code&gt;CLAUDE.md&lt;/code&gt; files, user-level memory, and auto-memory. So it feels more like the agent has a built-in memory system.&lt;/p&gt;

&lt;p&gt;OpenCode uses &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is more portable. You can commit it to the repo, share it with the team, and use it as a general agent instruction file instead of something tied to one vendor. OpenCode can even fall back to &lt;code&gt;CLAUDE.md&lt;/code&gt;, which makes migration easier.&lt;/p&gt;

&lt;p&gt;At some point, every agent runs out of context.&lt;/p&gt;

&lt;p&gt;Claude Code handles this by compacting the conversation. Older tool outputs are cleared first, then the session gets summarized if needed. That is why Claude Code has commands like &lt;code&gt;/context&lt;/code&gt; and &lt;code&gt;/compact&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;OpenCode is a bit more explicit. It checks whether the session is close to the model’s context limit, keeps a buffer for output, and then prunes old tool outputs before doing a full summary. The important bit is that OpenCode stores the raw history in &lt;strong&gt;SQLite&lt;/strong&gt;, so pruning does not mean the data is gone forever.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7om4vrfngpmqdjsh3phe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7om4vrfngpmqdjsh3phe.png" alt="OpenCode flexibility" width="800" height="351"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Tool use follows the same pattern.&lt;/p&gt;

&lt;p&gt;Claude Code gives you a polished default toolbelt: read, write, edit, grep, glob, bash, web fetch, todo tracking, MCP, hooks, skills, and subagents.&lt;/p&gt;

&lt;p&gt;OpenCode gives you a smaller but more configurable tool system: read, write, edit, patch, bash, grep, glob, web fetch, task, todo, &lt;a href="https://composio.dev/content/10-best-opencode-skills-that-are-actually-useful-in-2026" rel="noopener noreferrer"&gt;skills&lt;/a&gt;, MCP, custom tools, and experimental LSP support.&lt;/p&gt;

&lt;p&gt;The difference is who controls the tool layer.&lt;/p&gt;




&lt;h2&gt;
  
  
  Subagents and task delegation
&lt;/h2&gt;

&lt;p&gt;Subagents are basically how coding agents avoid stuffing everything into one giant conversation.&lt;/p&gt;

&lt;p&gt;Instead of making the main agent do every task itself, it can delegate a smaller job to another agent with its own context window, prompt, tools, and permissions.&lt;/p&gt;

&lt;p&gt;Claude Code and OpenCode both follow the same basic pattern here.&lt;/p&gt;

&lt;p&gt;The parent agent calls a &lt;code&gt;Task&lt;/code&gt; or &lt;code&gt;task&lt;/code&gt; tool.&lt;/p&gt;

&lt;p&gt;A child agent spins up.&lt;/p&gt;

&lt;p&gt;It does the work in isolation.&lt;/p&gt;

&lt;p&gt;Then it returns one final message back to the parent.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Main Agent
  |
  | calls Task / task
  v
Subagent
  - own context window
  - own prompt
  - own tools
  - own permissions
  |
  | returns final result only
  v
Main Agent continues...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That part is important. The parent usually does not see the full subagent conversation. It gets the result, not the whole reasoning.&lt;/p&gt;

&lt;p&gt;Claude Code has the more polished version of this.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmgfuv1jw1x2rt22d0i57.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmgfuv1jw1x2rt22d0i57.png" alt="Claude Code approach to subagents" width="799" height="269"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It ships with built-in agents like &lt;code&gt;Explore&lt;/code&gt;, &lt;code&gt;Plan&lt;/code&gt;, and &lt;code&gt;general-purpose&lt;/code&gt;. &lt;code&gt;Explore&lt;/code&gt; is mostly read-only and useful for repo research. &lt;code&gt;Plan&lt;/code&gt; helps gather context during planning. &lt;code&gt;general-purpose&lt;/code&gt; is for broader work.&lt;/p&gt;

&lt;p&gt;You can also define custom agents in &lt;code&gt;.claude/agents/&lt;/code&gt; with YAML frontmatter for things like &lt;code&gt;tools&lt;/code&gt;, &lt;code&gt;model&lt;/code&gt;, &lt;code&gt;permissionMode&lt;/code&gt;, &lt;code&gt;maxTurns&lt;/code&gt;, &lt;code&gt;skills&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That means you can do stuff like:&lt;/p&gt;

&lt;p&gt;Use a fast Haiku-style agent for repo search.&lt;/p&gt;

&lt;p&gt;Use a stronger model for code review.&lt;/p&gt;

&lt;p&gt;OpenCode has a similar shape, but it is more transparent.&lt;/p&gt;

&lt;p&gt;It has primary agents and subagents. Primary agents handle the main chat, while subagents are called through the &lt;code&gt;task&lt;/code&gt; tool or &lt;code&gt;@&lt;/code&gt; mentions.&lt;/p&gt;

&lt;p&gt;Custom agents can live in &lt;code&gt;.opencode/agents/*.md&lt;/code&gt; or inside &lt;code&gt;opencode.json&lt;/code&gt;, with fields like &lt;code&gt;mode&lt;/code&gt;, &lt;code&gt;model&lt;/code&gt;, &lt;code&gt;temperature&lt;/code&gt;, &lt;code&gt;steps&lt;/code&gt;, &lt;code&gt;prompt&lt;/code&gt;, and &lt;code&gt;permission&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The interesting part is that OpenCode stores subagents as real child sessions in &lt;strong&gt;SQLite&lt;/strong&gt;. So delegation is not just a hidden prompt trick. It is represented in the session model with its own messages, permissions, and snapshots.&lt;/p&gt;

&lt;p&gt;That fits OpenCode’s whole philosophy.&lt;/p&gt;

&lt;p&gt;Claude Code gives you a cleaner subagent experience.&lt;/p&gt;

&lt;p&gt;OpenCode gives you a more inspectable one.&lt;/p&gt;


&lt;h2&gt;
  
  
  Permissions, safety, and control
&lt;/h2&gt;

&lt;p&gt;This is where the two are very different.&lt;/p&gt;

&lt;p&gt;Claude Code is more conservative by default. It has permission modes, allow/ask/deny rules, hooks, and sandboxing around &lt;strong&gt;Bash&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So you can allow boring commands like tests, deny obvious footguns like .env reads or &lt;code&gt;curl | sh&lt;/code&gt;, and ask before anything risky.&lt;/p&gt;

&lt;p&gt;The important part is that Claude Code has multiple safety layers.&lt;/p&gt;

&lt;p&gt;Permissions decide what Claude is allowed to do.&lt;/p&gt;

&lt;p&gt;Something like:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"permissions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"allow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"Bash(npm run test *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git status *)"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"deny"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"Read(./.env)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Read(./secrets/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Bash(curl *)"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Hooks can intercept tool calls before or after they run and sandboxing gives Bash an OS-level boundary.&lt;/p&gt;

&lt;p&gt;OpenCode is simpler.&lt;/p&gt;

&lt;p&gt;Most of the control lives in one permission object inside &lt;code&gt;opencode.json&lt;/code&gt;. You can set rules for &lt;code&gt;bash&lt;/code&gt;, &lt;code&gt;edit&lt;/code&gt;, &lt;code&gt;read&lt;/code&gt;, &lt;code&gt;task&lt;/code&gt;, &lt;code&gt;webfetch&lt;/code&gt;, and other tools from the same place.&lt;/p&gt;

&lt;p&gt;That is clean, but OpenCode is also more permissive by default. You are expected to configure the rules yourself.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe8i1czi8c1lgv8ujg9l8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe8i1czi8c1lgv8ujg9l8.png" alt="OpenCode permissions" width="800" height="352"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Something like:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"permission"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ask"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"bash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ask"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"git status *"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"allow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"git push *"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"deny"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"rm *"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"deny"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"edit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"deny"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"packages/web/src/**/*.tsx"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ask"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;It does have some smart checks, especially for Bash. OpenCode parses shell commands with tree-sitter (the same thing you have inside NeoVim), so it can detect risky commands like &lt;code&gt;rm&lt;/code&gt;, &lt;code&gt;mv&lt;/code&gt;, &lt;code&gt;chmod&lt;/code&gt;, or paths outside the project more carefully than plain string matching.&lt;/p&gt;

&lt;p&gt;But there is no &lt;a href="https://www.anthropic.com/engineering/claude-code-sandboxing" rel="noopener noreferrer"&gt;native sandbox&lt;/a&gt; like Claude Code.&lt;/p&gt;

&lt;p&gt;The bigger OpenCode power feature is plugins. Plugins can intercept tool execution, add custom tools, and change agent behavior.&lt;/p&gt;

&lt;p&gt;That makes OpenCode way more hackable.&lt;/p&gt;


&lt;h2&gt;
  
  
  What the Claude Code leak tells us
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7h7v5xmi6q45a33xjk2v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7h7v5xmi6q45a33xjk2v.png" alt="Claude Code leak" width="799" height="337"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The interesting part of the leak is what it showed about coding agents.&lt;/p&gt;

&lt;p&gt;A lot of Claude magic is in the harness around the model: context management, tool descriptions, prompt caching, permissions, compaction, subagents, and the agent loop.&lt;/p&gt;

&lt;p&gt;OpenCode does pretty much the same. It is not trying to clone some impossible model-level feature. It is trying to build a different harness around similar idea.&lt;/p&gt;

&lt;p&gt;OpenCode’s advantage is that the harness is open, inspectable, and replaceable.&lt;/p&gt;

&lt;p&gt;Another thing that's clear is that the future is not just about better models, but the system around them.&lt;/p&gt;

&lt;p&gt;Better context control.&lt;/p&gt;

&lt;p&gt;Better tool boundaries.&lt;/p&gt;

&lt;p&gt;Better memory.&lt;/p&gt;

&lt;p&gt;Better permissions.&lt;/p&gt;

&lt;p&gt;That is why this comparison is even interesting. Claude Code and OpenCode are not just two CLIs. They are two different answers to the same question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;❓ How much of the agent stack should be final, and how much should developers be able to control?&lt;/p&gt;
&lt;/blockquote&gt;


&lt;h2&gt;
  
  
  So, which should you pick?
&lt;/h2&gt;

&lt;p&gt;There is no clever answer here.&lt;/p&gt;

&lt;p&gt;Pick &lt;strong&gt;Claude Code&lt;/strong&gt; if you want the cleanest Claude-native coding agent experience.&lt;/p&gt;

&lt;p&gt;Pick &lt;strong&gt;OpenCode&lt;/strong&gt; if you want more control.&lt;/p&gt;

&lt;p&gt;Personally, I still love Claude Code.&lt;/p&gt;

&lt;p&gt;I really do.&lt;/p&gt;

&lt;p&gt;Anthropic models are banger, especially for coding. The problem is that the limits have started to piss me off. When you are deep in a coding session and the limit hits, it completely breaks the flow.&lt;/p&gt;

&lt;p&gt;But there is some relief now.&lt;/p&gt;

&lt;p&gt;On May 6, Anthropic announced a new compute partnership with &lt;strong&gt;SpaceX&lt;/strong&gt; and doubled Claude Code’s 5-hour limits for &lt;strong&gt;Pro&lt;/strong&gt;, &lt;strong&gt;Max&lt;/strong&gt;, &lt;strong&gt;Team&lt;/strong&gt;, and &lt;strong&gt;seat-based Enterprise users&lt;/strong&gt;. They also removed peak-time limits for Pro and Max users.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqosm66pi1fq6pu4k76t0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqosm66pi1fq6pu4k76t0.png" alt="Claude Code increase in usage limit" width="799" height="384"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That makes Claude Code a lot easier to recommend again.&lt;/p&gt;

&lt;p&gt;I am personally still mostly stuck with Claude Code because the experience is just that good.&lt;/p&gt;

&lt;p&gt;But I use OpenCode when I want to try newer models like Kimi, OpenAI models, or local models. That is where OpenCode makes more sense to me. And by no means, it is to say that you can't use Anthropic models in OpenCode, you can, and that makes it even better.&lt;/p&gt;


&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;Claude Code and OpenCode are both useful, but for different reasons.&lt;/p&gt;

&lt;p&gt;Claude Code is the one I’d pick if I just want the agent to work without thinking too much about setup. It feels cleaner, and better for getting into a repo quickly.&lt;/p&gt;

&lt;p&gt;OpenCode is more for when you want control. Different models, different providers, more ways to shape the workflow around how you actually code.&lt;/p&gt;

&lt;p&gt;I wouldn’t overthink it.&lt;/p&gt;

&lt;p&gt;If you hate setup and love Anthropic, use Claude Code.&lt;/p&gt;

&lt;p&gt;If you want more flexibility and less vendor lock-in, use OpenCode.&lt;/p&gt;

&lt;p&gt;That’s really the whole comparison.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4d09dor9q6cwva8zlg5o.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4d09dor9q6cwva8zlg5o.gif" alt="steve jobs meme" width="422" height="237"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;div class="ltag__user ltag__user__id__1127015"&gt;
    &lt;a href="/shricodev" class="ltag__user__link profile-image-link"&gt;
      &lt;div class="ltag__user__pic"&gt;
        &lt;img src="https://media2.dev.to/dynamic/image/width=150,height=150,fit=cover,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1127015%2F1c5e48a2-f602-4e7d-8312-3c0322d155c6.jpg" alt="shricodev image"&gt;
      &lt;/div&gt;
    &lt;/a&gt;
  &lt;div class="ltag__user__content"&gt;
    &lt;h2&gt;
&lt;a class="ltag__user__link" href="/shricodev"&gt;Shrijal Acharya&lt;/a&gt;Follow
&lt;/h2&gt;
    &lt;div class="ltag__user__summary"&gt;
      &lt;a class="ltag__user__link" href="/shricodev"&gt;SDE • GOLD @Microsoft Student Ambassador • Prev Lead Collab and Dev-Team Lead @oppiaorg • Mail for collaboration&lt;/a&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;



</description>
      <category>ai</category>
      <category>agents</category>
      <category>claude</category>
      <category>cli</category>
    </item>
    <item>
      <title>Per-User OAuth for AI Agents: Why It Matters and What to Look For</title>
      <dc:creator>Dumebi Okolo</dc:creator>
      <pubDate>Wed, 20 May 2026 14:21:08 +0000</pubDate>
      <link>https://dev.to/composiodev/per-user-oauth-for-ai-agents-why-it-matters-and-what-to-look-for-4h4a</link>
      <guid>https://dev.to/composiodev/per-user-oauth-for-ai-agents-why-it-matters-and-what-to-look-for-4h4a</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdcvak90uie20wnaldj4n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdcvak90uie20wnaldj4n.png" alt="Per-user OAuth flow for AI agents" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AI agents are crossing a line that traditional software never had to. They read your Slack, draft your emails, push code, update your CRM, and pay your invoices. To do that, they need keys to systems that belong to specific people. Not the application. Not the company. The person.&lt;/p&gt;

&lt;p&gt;That is the entire reason per-user OAuth exists in the agent context, and it is the difference between a side project and something a customer will trust with their Gmail account.&lt;/p&gt;

&lt;p&gt;This article breaks down what per-user OAuth means for AI agents, why shared credentials fall apart at scale, what the emerging standards look like, and the exact checklist to use when picking a platform to handle it. We will also show how &lt;a href="https://composio.dev/" rel="noopener noreferrer"&gt;Composio&lt;/a&gt; approaches each of these problems so you do not have to assemble the stack yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with the way most teams start
&lt;/h2&gt;

&lt;p&gt;Most agent prototypes start with a single API key in an environment variable. It works for one developer, on one machine, for one demo. The moment a real user shows up, the model breaks.&lt;/p&gt;

&lt;p&gt;API keys identify the calling application, not the user behind the action. Every request from the agent looks the same to the downstream service. There is no concept of consent, no scoping per user, no way to revoke one person's access without invalidating everyone, and no audit trail that says "this action happened because Sarah asked the agent to do it."&lt;/p&gt;

&lt;p&gt;This becomes a real security problem fast. If the agent code accidentally passes the wrong user identifier, or an attacker tricks the agent into requesting data for a user who did not authorize it, the agent has no protocol-level defense. This is the classic confused deputy problem, and it scales horribly with autonomous systems that chain dozens of tool calls per task.&lt;/p&gt;

&lt;p&gt;Composio's own &lt;a href="https://composio.dev/blog/secure-ai-agent-infrastructure-guide" rel="noopener noreferrer"&gt;guide on AI agent infrastructure&lt;/a&gt; calls this hitting the "Authentication Wall." It is the moment a promising prototype stops being promising.&lt;/p&gt;

&lt;h2&gt;
  
  
  What per-user OAuth actually solves
&lt;/h2&gt;

&lt;p&gt;Per-user OAuth flips the model. Instead of the agent holding one master credential, each user grants the agent a scoped, revocable token tied to their own account. The agent acts with that user's identity, within the limits that user approved, for as long as that user allows.&lt;/p&gt;

&lt;p&gt;Concretely, this gives you:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Explicit consent.&lt;/strong&gt; The user sees what the agent is asking for and approves it. No assumed access.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scoped permissions.&lt;/strong&gt; The token can be limited to read-only access on a specific resource rather than full account control. If the agent only needs to read, that is all it gets.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Short lifetimes.&lt;/strong&gt; Access tokens expire in minutes to an hour. Refresh tokens rotate. A leaked token is dangerous for a short window, not forever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Selective revocation.&lt;/strong&gt; Revoking one user's grant does not break the agent for everyone else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Identity in the audit log.&lt;/strong&gt; Every action traces back to a specific user, which is what SOC 2, HIPAA, and ISO 27001 actually require.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-tenant isolation.&lt;/strong&gt; Each user's tokens live in their own bucket, encrypted at rest. A bug in one tenant's workflow does not expose another tenant's data.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fm8adyi8c9n3x0cyb62hj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fm8adyi8c9n3x0cyb62hj.png" alt="Multi-tenant token isolation across users" width="800" height="471"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That last one matters more than people think. Composio's &lt;a href="https://docs.composio.dev/docs/users-and-sessions" rel="noopener noreferrer"&gt;user and session model&lt;/a&gt; is built around exactly this idea: a user is an identifier from your app, every connection lives under that user's ID, and connections are fully isolated between users. The same agent code can serve thousands of users without any of them ever touching each other's data.&lt;/p&gt;

&lt;h2&gt;
  
  
  The standards that are actually shaping this
&lt;/h2&gt;

&lt;p&gt;The protocol layer is moving fast. There are three things worth knowing.&lt;/p&gt;

&lt;h3&gt;
  
  
  OAuth 2.1 with mandatory PKCE
&lt;/h3&gt;

&lt;p&gt;OAuth 2.1 is the current best-practice consolidation of OAuth 2.0. It makes PKCE (Proof Key for Code Exchange) mandatory and removes older, less secure flows. PKCE matters specifically for agents because most agents are public clients running in environments where you cannot reliably hide a client secret. PKCE prevents an attacker from intercepting an authorization code mid-flow.&lt;/p&gt;

&lt;p&gt;If a platform you are evaluating does not enforce PKCE, that is a red flag.&lt;/p&gt;

&lt;h3&gt;
  
  
  The MCP Authorization spec
&lt;/h3&gt;

&lt;p&gt;The &lt;a href="https://modelcontextprotocol.io/specification/2025-03-26/basic/authorization" rel="noopener noreferrer"&gt;Model Context Protocol authorization specification&lt;/a&gt; formalized OAuth as its standard in 2025. It mandates OAuth 2.1 with PKCE, requires Authorization Server Metadata discovery via RFC 8414, and supports Dynamic Client Registration via RFC 7591. The November 2025 update added step-up authorization, letting clients request additional scopes only when an operation actually requires them, rather than over-permissioning the initial token.&lt;/p&gt;

&lt;p&gt;The spec also had a real-world security crisis to address. In late 2025, security researchers at Obsidian Security disclosed &lt;a href="https://www.obsidiansecurity.com/blog/when-mcp-meets-oauth-common-pitfalls-leading-to-one-click-account-takeover" rel="noopener noreferrer"&gt;one-click account takeover vulnerabilities&lt;/a&gt; in remote MCP servers from several well-known organizations. The root cause: many MCP servers were implemented as OAuth proxies using a single static &lt;code&gt;client_id&lt;/code&gt; to talk to the upstream SaaS authorization server. Once any user consented for that shared client_id, the SaaS auth server cached the decision. An attacker could then complete the MCP-layer consent themselves, send a crafted authorization link to a victim, and the upstream server would skip the consent prompt entirely because it had seen that client_id before. The authorization code would be issued to the attacker's redirect URI.&lt;/p&gt;

&lt;p&gt;The fix is per-client identity and strict consent handling at the proxy layer. Composio's &lt;a href="https://composio.dev/toolkits/composio" rel="noopener noreferrer"&gt;Tool Router&lt;/a&gt; gives each session a secure, user-scoped MCP URL rather than a shared endpoint, which sidesteps this class of attack structurally.&lt;/p&gt;

&lt;h3&gt;
  
  
  The IETF "On-Behalf-Of" draft for AI agents
&lt;/h3&gt;

&lt;p&gt;There is an active IETF draft, &lt;a href="https://datatracker.ietf.org/doc/draft-oauth-ai-agents-on-behalf-of-user/" rel="noopener noreferrer"&gt;draft-oauth-ai-agents-on-behalf-of-user&lt;/a&gt;, that extends OAuth specifically for agent delegation. It adds a &lt;code&gt;requested_actor&lt;/code&gt; parameter so the consent screen shows the agent's identity (not just the app), and an &lt;code&gt;actor_token&lt;/code&gt; parameter so the agent authenticates itself when exchanging the authorization code.&lt;/p&gt;

&lt;p&gt;The result is an access token that documents the full delegation chain: the user delegated to this client application, which delegated to this specific agent. That chain is what makes after-the-fact auditing possible.&lt;/p&gt;

&lt;p&gt;The draft is at revision 02 as of August 2025 and has not been adopted by the working group yet, but it points clearly at where the protocol layer is headed. Multi-hop delegation (agent A calling agent B on the same user's behalf) is still an open problem in the spec.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production architecture: where teams fail
&lt;/h2&gt;

&lt;p&gt;Getting an access token is the easy part. Operating at production scale is where most homegrown OAuth implementations collapse. Five things tend to go wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Token storage.&lt;/strong&gt; Tokens have to be encrypted at rest, isolated per tenant, never logged, and never placed in LLM context. The last point is non-obvious and critical: if you put a refresh token in the prompt, a prompt injection attack can exfiltrate it. The pattern to use is brokered credentials, where the LLM never sees the token at all and a separate service makes the actual API call.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frr6l8q3lu5ppwg1p8db7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frr6l8q3lu5ppwg1p8db7.png" alt="Brokered credentials pattern" width="800" height="374"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Composio's &lt;a href="https://composio.dev/content/secure-ai-agent-infrastructure-guide" rel="noopener noreferrer"&gt;secure infrastructure guide&lt;/a&gt; explains this pattern directly: the LLM asks Composio to perform an action, Composio calls the upstream API with the stored credential, and the token never enters the model's context window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Refresh handling.&lt;/strong&gt; Refresh tokens need proactive coordination. Waiting for a 401 error and then refreshing creates race conditions, cascading retries, and unstable background jobs. Composio handles refresh automatically and only marks a connection as &lt;code&gt;EXPIRED&lt;/code&gt; after multiple refresh attempts have failed, according to its &lt;a href="https://docs.composio.dev/docs/authenticating-tools" rel="noopener noreferrer"&gt;authentication docs&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scope discipline.&lt;/strong&gt; Composio requests sensible default scopes for each toolkit but lets you override them via &lt;a href="https://docs.composio.dev/docs/custom-auth-configs" rel="noopener noreferrer"&gt;custom auth configs&lt;/a&gt;. Tightening scopes shrinks the blast radius if something goes wrong. Most APIs still have coarse-grained scopes, which means even with discipline, agents tend to be over-permissioned. The mitigation is short token lifetimes and per-tool scoping where the API supports it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Branding and consent screens.&lt;/strong&gt; When users hit an OAuth consent screen that says "Composio wants to access your Gmail" instead of "YourProduct wants to access your Gmail," conversion drops and trust takes a hit. Composio's &lt;a href="https://docs.composio.dev/docs/custom-app-vs-managed-app" rel="noopener noreferrer"&gt;white-labeling support&lt;/a&gt; lets you bring your own OAuth app credentials so the consent screen shows your brand. Use managed apps for prototyping and your own credentials for production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-account per user.&lt;/strong&gt; Some users connect a personal Gmail and a work Gmail. The platform needs to model that without forcing them to share a user ID across both. Composio handles this with the connected account ID layered under the user ID, so a single user can have multiple accounts on the same toolkit, as documented in the &lt;a href="https://docs.composio.dev/docs/users-and-sessions" rel="noopener noreferrer"&gt;users and sessions guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to look for in a per-user OAuth platform for agents
&lt;/h2&gt;

&lt;p&gt;If you are evaluating providers, these are the non-negotiables:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Per-user token isolation&lt;/strong&gt; with encryption at rest and no cross-tenant leakage&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OAuth 2.1 with mandatory PKCE&lt;/strong&gt; for all OAuth flows&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automatic, proactive token refresh&lt;/strong&gt; that does not depend on the client&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Short-lived access tokens&lt;/strong&gt; with rotating refresh tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Granular scope configuration&lt;/strong&gt; so each integration uses least privilege&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Brokered credentials&lt;/strong&gt; so the LLM never sees raw tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Selective per-user revocation&lt;/strong&gt; without affecting other users&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit trail on the delegation chain&lt;/strong&gt; for compliance&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;White-label OAuth consent screens&lt;/strong&gt; for production trust&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Step-up authorization&lt;/strong&gt; to request new scopes only when needed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-account per user&lt;/strong&gt; for personal and work accounts on the same app&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP-native deployment&lt;/strong&gt; so any MCP-compatible client can use the same auth layer&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SOC 2 Type 2 and ISO 27001 compliance&lt;/strong&gt; as table stakes for enterprise customers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SDK support across major frameworks&lt;/strong&gt; (LangChain, CrewAI, OpenAI Agents SDK, Claude Agent SDK, Mastra, Vercel AI SDK, LlamaIndex)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Composio covers every item on that list today. It supports 500+ toolkits, handles OAuth end-to-end with user-scoped tokens, is SOC 2 Type 2 and ISO 27001 compliant, and works as both a direct SDK and an MCP server. The full feature breakdown is on the &lt;a href="https://composio.dev/agentauth" rel="noopener noreferrer"&gt;AgentAuth product page&lt;/a&gt; and the &lt;a href="https://composio.dev/content/ai-agent-authentication-platforms" rel="noopener noreferrer"&gt;comparison guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this looks like in code
&lt;/h2&gt;

&lt;p&gt;Per-user OAuth with Composio compresses into a handful of lines. The pattern below follows the current SDK as documented in Composio's &lt;a href="https://docs.composio.dev/docs/authenticating-tools" rel="noopener noreferrer"&gt;authenticating tools guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;First, install the SDK and set your API key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;composio
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;COMPOSIO_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;your_api_key_here
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To trigger the OAuth flow for a specific user, use the hosted Connect Link pattern. This returns a redirect URL the user opens in their browser to complete authentication:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;composio&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Composio&lt;/span&gt;

&lt;span class="n"&gt;composio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Composio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your_api_key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Use the "AUTH CONFIG ID" from your Composio dashboard
&lt;/span&gt;&lt;span class="n"&gt;auth_config_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your_auth_config_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# Use a unique identifier for each user in your application
&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_1349_129_12&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;connection_request&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;composio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;connected_accounts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;link&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;auth_config_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;auth_config_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;callback_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://your-app.com/callback&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;redirect_url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;connection_request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;redirect_url&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Visit: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;redirect_url&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; to authenticate your account&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After the user completes the flow, Composio stores the tokens, links them to that user ID, and handles refresh automatically. The agent code never touches the token directly.&lt;/p&gt;

&lt;p&gt;To fetch tools scoped to a specific user, pass that same user ID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;composio&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Composio&lt;/span&gt;

&lt;span class="n"&gt;composio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Composio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your_api_key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;user_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_1349_129_12&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# Tools are automatically scoped to this user's connected accounts
&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;composio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;toolkits&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GITHUB&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GMAIL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two users running the same workflow get two different sets of credentials applied transparently:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;tools_user_1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;composio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;toolkits&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GITHUB&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="n"&gt;tools_user_2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;composio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;toolkits&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GITHUB&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="c1"&gt;# Each set of tools uses the respective user's credentials
# when invoked by an agent
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole point of per-user OAuth: the auth layer disappears into the platform, and the agent code reads like it is talking to a single user even when it is serving thousands.&lt;/p&gt;

&lt;p&gt;For a full working example with an LLM framework, see the framework-specific quickstart guides in &lt;a href="https://docs.composio.dev/docs" rel="noopener noreferrer"&gt;Composio's documentation&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Per-user OAuth is not a feature you bolt onto an agent product later. It is the foundation that decides whether your agents can serve real customers at all. Shared API keys cap your ceiling at a single-user demo. Per-user OAuth opens the door to multi-tenant production deployment, enterprise compliance, and the kind of trust required to handle a customer's inbox, calendar, or revenue pipeline.&lt;/p&gt;

&lt;p&gt;The protocol layer is still evolving. OAuth 2.1, the MCP authorization spec, and the IETF on-behalf-of draft are all converging on the same answer: explicit user consent, scoped delegation, audited token lifecycle, and isolation per user. Build for that model now and you will not need to retrofit later.&lt;/p&gt;

&lt;p&gt;If you would rather skip the months of building it yourself, &lt;a href="https://composio.dev/" rel="noopener noreferrer"&gt;start with Composio&lt;/a&gt; or read the &lt;a href="https://composio.dev/blog/secure-ai-agent-infrastructure-guide" rel="noopener noreferrer"&gt;auth-to-action guide&lt;/a&gt; for the full architecture. The shortest path from prototype to production is using a platform that already solved this.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>mcp</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Took Cursor Agents Out of the IDE and It Got Weirdly Powerful 🤯</title>
      <dc:creator>Developer Harsh</dc:creator>
      <pubDate>Tue, 19 May 2026 10:05:31 +0000</pubDate>
      <link>https://dev.to/composiodev/i-took-cursor-agents-out-of-the-ide-and-it-got-weirdly-powerful-286i</link>
      <guid>https://dev.to/composiodev/i-took-cursor-agents-out-of-the-ide-and-it-got-weirdly-powerful-286i</guid>
      <description>&lt;p&gt;Cursor recently released their &lt;a href="https://cursor.com/blog/typescript-sdk" rel="noopener noreferrer"&gt;cursor agents SDK&lt;/a&gt; to the public, and it's quietly powering many teams to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;invoke agents directly from CI/CD pipelines,&lt;/li&gt;
&lt;li&gt;create automations for end-to-end workflows,&lt;/li&gt;
&lt;li&gt;and embedding agents into core products.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Basically, the SDK lets developers deploy agents without overthinking building and maintaining the entire agent stack&lt;/p&gt;

&lt;p&gt;in this blog I give a glimpse of what is possible with the cursor agents SDK and how to overcome the restrictions of limited tool access.&lt;/p&gt;

&lt;p&gt;Let's start with a brief overview of cursor agents.&lt;/p&gt;




&lt;h2&gt;
  
  
  A Primer On Cursor Agents SDK
&lt;/h2&gt;

&lt;p&gt;Building an agent from scratch is a massive headache. Cursor SDK skips the&lt;br&gt;
plumbing.&lt;/p&gt;

&lt;p&gt;Before the SDK, &lt;a href="https://cursor.com" rel="noopener noreferrer"&gt;Cursor&lt;/a&gt; was strictly an interactive IDE, but its newly released SDK turns that&lt;br&gt;
agentic power into headless infrastructure. &lt;/p&gt;

&lt;p&gt;It uses the same &lt;strong&gt;harness&lt;/strong&gt; that powers the desktop app, meaning you get IDE-grade code generation programmatically.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Harness: context engine, workspace management, and routing&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In fact, the cursor harness matters more than the model. &lt;a href="https://www.endorlabs.com/learn/gpt-5-5-sets-a-new-code-security-record-with-cursor-not-codex-in-agent-security-league" rel="noopener noreferrer"&gt;Endor Labs&lt;/a&gt; benched&lt;br&gt;
GPT-5.5 natively at 61.5% functional correctness. Then, dropping the same model into Cursor's harness, it scored 87.2%.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Faa4jgbjxm262koapm9nm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Faa4jgbjxm262koapm9nm.png" alt="Proof" width="800" height="294"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cursor spent years tuning the context management and tool dispatch; the SDK hands you that refined engine.&lt;/p&gt;

&lt;p&gt;Here is the spec list for those who care:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Spec / Feature&lt;/th&gt;
&lt;th&gt;Technical Detail&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;The Harness&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Built-in codebase indexing, semantic search, &amp;amp; instant grep.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Execution&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Cloud (sandboxed VMs with durable state), Self-Hosted, Local.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Models&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Agnostic. Uses &lt;code&gt;composer-2/composer-3&lt;/code&gt; (default), Claude, or OpenAI.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Integrations&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Deep MCP support, auto-loaded &lt;code&gt;.cursor/skills/&lt;/code&gt;, &amp;amp; hooks.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Subagents&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Main agent can spawn subagents with distinct prompts/models.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This means we have a perfect stack to build our agent. Let’s get started with the setup.&lt;/p&gt;


&lt;h2&gt;
  
  
  How to Set Up Cursor Agents SDK
&lt;/h2&gt;

&lt;p&gt;You can install it using &lt;code&gt;npm&lt;/code&gt; to get started, then use Cursor's native &lt;code&gt;cursor SDK skill&lt;/code&gt; for guidance.&lt;/p&gt;

&lt;p&gt;So open your terminal and type:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; @cursor/sdk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Output&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once done, install the cursor skill for guiding the cursor / Claude Code. In the same terminal, run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm add skills /cursor-sdk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But there is a catch: by default, the Cursor SDK agents ship with a few default tools + MCP support.&lt;/p&gt;

&lt;p&gt;However, this means connecting to or configuring multiple MCP servers is such a hassle for a production-grade product.&lt;/p&gt;

&lt;p&gt;So let's automate it with Composio as the orchestrator, connect and configure it once, and you get secure access to 1000+ tools with optimized tool calls and context.&lt;/p&gt;




&lt;h2&gt;
  
  
  Building a Production Agent With Composio MCP + Cursor Agent SDK
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Why Composio?
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://composio.dev/" rel="noopener noreferrer"&gt;Composio&lt;/a&gt; offers a single MCP server you can plug into with the Cursor SDK to access &lt;a href="https://composio.dev/toolkits" rel="noopener noreferrer"&gt;1000+ apps&lt;/a&gt; instantly. &lt;/p&gt;

&lt;p&gt;When you’re building an agent that requires interacting with external applications, let’s say a Sales agent with &lt;a href="https://composio.dev/toolkits/gong" rel="noopener noreferrer"&gt;Gong&lt;/a&gt;, &lt;a href="https://composio.dev/toolkits/hubspot" rel="noopener noreferrer"&gt;HubSpot&lt;/a&gt;, &lt;a href="https://composio.dev/toolkits/salesforce" rel="noopener noreferrer"&gt;Salesforce connectors&lt;/a&gt;, etc., you’d spend weeks on partnerships and integrations. Composio completely removes this friction, so you build what matters&lt;/p&gt;

&lt;p&gt;So, let’s explore one example of a GitHub agent that pulls the repository, analyzes it, creates a dev branch, performs automatic refactoring, pushes the code, and opens a PR for review - a simple but powerful use case.&lt;/p&gt;

&lt;p&gt;Let's begin!&lt;/p&gt;

&lt;h3&gt;
  
  
  Setup Workspace
&lt;/h3&gt;

&lt;p&gt;Head to the terminal and run the following commands&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir &lt;/span&gt;cursor-agent
&lt;span class="nb"&gt;cd &lt;/span&gt;cursor-agent
npm init &lt;span class="nt"&gt;-y&lt;/span&gt;
npm &lt;span class="nb"&gt;install &lt;/span&gt;typescript ts-node @types/node &lt;span class="nt"&gt;--save-dev&lt;/span&gt;
npx tsc &lt;span class="nt"&gt;--init&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The command will create a new project named &lt;code&gt;cursor_agent&lt;/code&gt; , switch to the folder, initialize a blank &lt;code&gt;npm&lt;/code&gt; project (ensure &lt;code&gt;package.json&lt;/code&gt; gets created), install typescript support for development, and initialize the typescript compiler and linter.&lt;/p&gt;

&lt;p&gt;Since our repository is set up, let’s configure environment variables.&lt;/p&gt;

&lt;h3&gt;
  
  
  Set up Environment Variables
&lt;/h3&gt;

&lt;p&gt;Add a &lt;code&gt;.env&lt;/code&gt; file to keep secrets secure and add the following values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;CURSOR_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;your-cursor-api-key
&lt;span class="nv"&gt;COMPOSIO_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;your-composio-api-key
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can get the &lt;code&gt;CURSOR_API_KEY&lt;/code&gt; by going to the &lt;a href="https://cursor.com/dashboard/integrations" rel="noopener noreferrer"&gt;Cursor integration&lt;/a&gt; page and creating a new API key.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Feddl8myiiaj88t15jurd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Feddl8myiiaj88t15jurd.png" alt="Step 1" width="800" height="456"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For COMPOSIO_API_KEY  , visit the &lt;a href="https://dashboard.composio.dev/" rel="noopener noreferrer"&gt;Composio Dashboard (Platform)&lt;/a&gt;, log in/sign up to the account, and head to the default project (or create one) → profile→ API KEY.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fv4drqq50yjgm04dz5pdu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fv4drqq50yjgm04dz5pdu.png" alt="Step 2" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Click “Create One”, it creates one, copy and then paste it into .env . Make sure it starts with ak_ .&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpusjo38ewzyorboce28i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpusjo38ewzyorboce28i.png" alt="Step 3" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Time to write the code.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Code
&lt;/h3&gt;

&lt;p&gt;Now, create a new &lt;code&gt;index.ts&lt;/code&gt; file and paste the following code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="c1"&gt;// imports&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;dotenv/config&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;readline&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;node:readline/promises&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;stdin&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;stdout&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;node:process&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;SDKAgent&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@cursor/sdk&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;DEFAULT_USER_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;disposeAllAgents&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;getComposioApiKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;getOrCreateRuntime&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;requireEnv&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;./runtime.ts&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// log tool calls to stdout&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;logToolCall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;unknown&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;running&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;completed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;error&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="k"&gt;void&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;running&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`\n[tool: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;] &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;completed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`[tool: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;] done`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// stream assistant text and tool events - act as streaming handler&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;runAgentChatStreaming&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;SDKAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;run&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;agent &amp;gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="k"&gt;await &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;run&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;assistant&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;block&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;block&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;text&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;block&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tool_call&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nf"&gt;logToolCall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;run&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;wait&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;error&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Agent run failed (&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;)`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// runs the chat loop, adds a readline interface and disposes agents on exit&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;requireEnv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CURSOR_API_KEY&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nf"&gt;getComposioApiKey&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Setting up Composio session...&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;runtime&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getOrCreateRuntime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;DEFAULT_USER_ID&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s2"&gt;`Ready (user: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;runtime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;, session: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;runtime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;, tools: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;runtime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;toolsCount&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;).`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Type a message, or 'exit' to quit.&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;rl&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;readline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createInterface&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;userInput&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;rl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;question&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;you &amp;gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;userInput&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userInput&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;exit&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;userInput&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;quit&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;break&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

      &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;runAgentChatStreaming&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;runtime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;userInput&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;rl&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;disposeAllAgents&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="k"&gt;catch&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let's look at the code: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;We start by loading dependencies from &lt;code&gt;.env&lt;/code&gt;, &lt;code&gt;readline&lt;/code&gt;, and shared helpers from &lt;a href="https://github.com/DevloperHS/cursor_agents/blob/main/runtime.ts" rel="noopener noreferrer"&gt;&lt;code&gt;runtime.ts&lt;/code&gt;&lt;/a&gt; (&lt;code&gt;getOrCreateRuntime&lt;/code&gt;, API key checks, and cleanup)&lt;/li&gt;
&lt;li&gt;For a sanity check, we validate &lt;code&gt;CURSOR_API_KEY&lt;/code&gt; and &lt;code&gt;COMPOSIO_API_KEY&lt;/code&gt; before the chat loop starts (Composio is also checked when &lt;code&gt;runtime.ts&lt;/code&gt; loads).&lt;/li&gt;
&lt;li&gt;Next, we create a Composio session for the default user and spin up a Cursor agent on &lt;code&gt;composer-2&lt;/code&gt; with the Composio MCP server for tools (handled in &lt;code&gt;runtime.ts&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;We keep the same agent in the loop for every message, so tools and sessions stay wired throughout the CLI session.&lt;/li&gt;
&lt;li&gt;We run a terminal chat loop with readline: read input → send to agent → stream assistant text and tool status → repeat until &lt;code&gt;exit&lt;/code&gt; or &lt;code&gt;quit&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;While streaming, we log tool calls on &lt;code&gt;running&lt;/code&gt; and &lt;code&gt;completed&lt;/code&gt;, and throw if the agent run ends in error.&lt;/li&gt;
&lt;li&gt;On exit, we close readline, dispose of agents via &lt;code&gt;runtime.ts&lt;/code&gt;, and log any failure from &lt;code&gt;main()&lt;/code&gt; before exiting with code 1.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you encounter an issue, refer to the codebase in the project's &lt;a href="https://github.com/DevloperHS/cursor_agents" rel="noopener noreferrer"&gt;GitHub repo&lt;/a&gt;.  Feel free to clone it, tweak it to your liking, or use it as it is.&lt;/p&gt;

&lt;p&gt;And that's it, all set, time to run the agent!&lt;/p&gt;




&lt;h2&gt;
  
  
  Run the agent
&lt;/h2&gt;

&lt;p&gt;To run the agent in the terminal and at the project root, type:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx ts-node index.ts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and an interactive chat opens up where you can prompt it step by step / direct it to do the entire task at once. Here is a demo of me using it in chat mode (step by step)&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/k4wfe4pJS_0"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;As you can see, the cursor agent acts as the brain, calling the &lt;a href="https://composio.dev/toolkits/github" rel="noopener noreferrer"&gt;Composio GitHub tools&lt;/a&gt; to clone the repo, check for changes, create a dev branch, refactor the code, commit to the dev branch, and finally open a PR.&lt;/p&gt;

&lt;p&gt;Usually, this takes many human hours, but with a cursor-agent, it takes only 5 minutes.&lt;/p&gt;

&lt;p&gt;Also, with the new update, you need to add a GitHub token to use GitHub with the cursor/agent. Using Composio fixes this need as well.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: For simplicity I have used terminal , but you can add a ui layer on top of the agent.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;With this, we have reached the end of the article. Here is my closing note.&lt;/p&gt;




&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;In this article, we learned how to use the Cursor SDK and Composio MCP to build a GitHub chatbot while keeping it optimized for asynchronous operations and following best practices, such as masking sensitive variables.&lt;/p&gt;

&lt;p&gt;But this is just the tip of the iceberg with what's possible with the cursor agent SDK. Feel free to experiment and build your own assistant/chatbot. You can explore more in the &lt;a href="https://github.com/cursor/cookbook" rel="noopener noreferrer"&gt;cursor-cookbook&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;At the end of the day, building agents / multi-agent systems is not hard; planning and orchestrating them to work efficiently, collaboratively, and securely is. In fact, it's the next valuable skill in 2026.&lt;/p&gt;

&lt;p&gt;So what are you waiting for? Get started with Cursor Agents while &lt;a href="https://composio.dev/" rel="noopener noreferrer"&gt;Composio MCP&lt;/a&gt; handles the tooling layer.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>agents</category>
    </item>
    <item>
      <title>Building Streamable HTTP MCP Servers from Scratch using FastMCP in 2026</title>
      <dc:creator>Developer Harsh</dc:creator>
      <pubDate>Mon, 18 May 2026 14:50:04 +0000</pubDate>
      <link>https://dev.to/composiodev/building-streamable-http-mcp-servers-from-scratch-using-fastmcp-in-2026-5fh9</link>
      <guid>https://dev.to/composiodev/building-streamable-http-mcp-servers-from-scratch-using-fastmcp-in-2026-5fh9</guid>
      <description>&lt;p&gt;If you've spent any time wiring up tools for an AI agent, you know the pain: every model vendor has its own function-calling format, every integration is bespoke glue code, and every model upgrade breaks something downstream. &lt;/p&gt;

&lt;p&gt;MCP changes this by providing a shared interface for tools, data sources, and AI clients. Instead of rebuilding integrations for every model or app, you expose capabilities once through an MCP server and let compatible clients connect to them.&lt;/p&gt;

&lt;p&gt;Anthropic released the Model Context Protocol in November 2024, and by spring 2025, OpenAI, Microsoft, and Google had adopted it. It has quickly become the de facto standard for connecting AI systems to external tools and data. That makes MCP worth learning now.&lt;/p&gt;

&lt;p&gt;This guide teaches you how to build an MCP server from scratch, expose it over two transports: &lt;strong&gt;stdio&lt;/strong&gt; and &lt;strong&gt;streamable HTTP&lt;/strong&gt;, connect it to Claude Desktop and Cursor.  and provide you with a production checklist for hardening the production MCP servers before shipping them. The full code is available in the companion repo at the end&lt;/p&gt;

&lt;h3&gt;
  
  
  tl; dr
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;What is MCP&lt;/li&gt;
&lt;li&gt;MCP Components in Nutshell&lt;/li&gt;
&lt;li&gt;Build a Local MCP server (tools, resources, and prompts)&lt;/li&gt;
&lt;li&gt;Connect to Hosted MCP Server (HTTP Streamable)&lt;/li&gt;
&lt;li&gt;MCP Server Advance Use Case&lt;/li&gt;
&lt;li&gt;Deployment Notes for Production MCP Servers&lt;/li&gt;
&lt;li&gt;Conclusion&lt;/li&gt;
&lt;li&gt;Frequently Asked Questions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Let's Begin!&lt;/p&gt;




&lt;h2&gt;
  
  
  What is MCP
&lt;/h2&gt;

&lt;p&gt;An MCP server is a JSON-RPC 2.0 process that exposes three primitives to any compliant &lt;a href="https://composio.dev/content/mcp-client-step-by-step-guide-to-building-from-scratch" rel="noopener noreferrer"&gt;MCP client&lt;/a&gt; over stdio or streamable HTTP. Those primitives are &lt;strong&gt;tools&lt;/strong&gt; (functions the model can call), &lt;strong&gt;resources&lt;/strong&gt; (data the model can read), and &lt;strong&gt;prompts&lt;/strong&gt; (reusable templates). That's the whole protocol. &lt;/p&gt;

&lt;p&gt;Everything else is an implementation detail as given in the comparison table.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Feature&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Model Context Protocol (MCP)&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Traditional API (REST / GraphQL / gRPC)&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Primary purpose&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Give LLM-based agents a &lt;em&gt;single&lt;/em&gt; way to fetch context &lt;strong&gt;and&lt;/strong&gt; invoke side-effecting tools.&lt;/td&gt;
&lt;td&gt;General machine-to-machine data exchange &amp;amp; business logic.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Integration effort&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Once an agent speaks MCP, it can talk to any compliant server; only one SDK/wire-spec to learn.&lt;/td&gt;
&lt;td&gt;Each API exposes its own spec/SDK; you integrate them one by one.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Interaction model&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stateful sessions + bidirectional messaging; supports long-running tasks and mid-job progress callbacks.&lt;/td&gt;
&lt;td&gt;Stateless request/response; usually unidirectional.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Streaming support&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Standardised via “Streamable HTTP” (SSE/WebSocket).&lt;/td&gt;
&lt;td&gt;Not part of the REST spec; developers bolt on WebSockets/SSE ad hoc.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Extensibility (adding new capabilities)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The server can add new tools/resources without breaking clients; the agent discovers them at runtime.&lt;/td&gt;
&lt;td&gt;Breaking changes need versioning, or clients must update to use new endpoints/types.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Standardisation/wire format&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One JSON-Schema-driven spec for inputs &amp;amp; outputs, plus defined transports.&lt;/td&gt;
&lt;td&gt;Multiple styles (REST, GraphQL, gRPC) with differing auth headers, error shapes, and media types.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Maturity &amp;amp; tooling&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rapidly growing, but still early; dozens of open-source servers and early commercial support.&lt;/td&gt;
&lt;td&gt;Decades of best-practice guides, gateways, SDKs, APM, and monitoring.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Performance path length&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Extra hop (agent → MCP → underlying API) adds a bit of latency; streaming mitigates waiting for large payloads.&lt;/td&gt;
&lt;td&gt;Direct call; generally lower overhead for high-QPS, deterministic workloads.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Typical sweet-spot use-cases&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Autonomous agents chaining multiple tools, dynamic workflows, and user-in-the-loop&lt;/td&gt;
&lt;td&gt;CRUD data services, stable integrations, high-throughput micro-services.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you want a deeper conceptual primer, check out the &lt;a href="https://composio.dev/blog/what-is-model-context-protocol-mcp-explained/" rel="noopener noreferrer"&gt;Model Context Protocol explainer&lt;/a&gt;, which covers the architecture in detail. For primer, here is how this all works!&lt;/p&gt;




&lt;h2&gt;
  
  
  MCP Components in Nutshell
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fs0n3rpcv1z9e5dubyxjb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fs0n3rpcv1z9e5dubyxjb.png" alt="MCP Component" width="800" height="541"&gt;&lt;/a&gt; %}&lt;/p&gt;

&lt;h3&gt;
  
  
  The Architecture
&lt;/h3&gt;

&lt;p&gt;The MCP architecture consists of several key components that work together to enable seamless integration:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;MCP Hosts&lt;/strong&gt;: Applications that want to use tools or data through MCP. Examples include Claude Desktop, Cursor, or any AI app that supports MCP.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP Clients&lt;/strong&gt;: Client instances created by the host. Each client maintains a dedicated one-to-one connection with a single MCP server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP Servers&lt;/strong&gt;: Lightweight programs that expose tools, resources, and prompts through the MCP protocol.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local Data Sources&lt;/strong&gt;: Files, databases, codebases, or local services that an MCP server can access on the user’s machine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remote Services&lt;/strong&gt;: External APIs or cloud services that an MCP server can call, such as Google Sheets, GitHub, Slack, or Postgres hosted in the cloud.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;At a high level, the host does not talk directly to your database, file system, or API. Instead, it talks to an MCP client, which talks to an MCP server, which then talks to the actual data source or service. &lt;/p&gt;

&lt;p&gt;This simple diagram might help you understand the flow better:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhtr5ddgaqrh6fdn93k6x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhtr5ddgaqrh6fdn93k6x.png" alt="MCP FLOW" width="800" height="893"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The key point is that MCP servers serve as the integration layer. They hide the messy parts of authentication, API calls, file access, and business logic behind a clean protocol interface.&lt;/p&gt;

&lt;p&gt;This separation of concerns makes MCP servers modular and maintainable. The AI app does not need to know how Google Sheets, GitHub, or a local database works. It only needs to know how to call the MCP server.&lt;/p&gt;

&lt;p&gt;So how does this all connect?&lt;/p&gt;




&lt;h3&gt;
  
  
  &lt;strong&gt;How The Components Work Together&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Let's understand this with a practical example:&lt;/p&gt;

&lt;p&gt;Say you're using Claude Code (an MCP host) to manage your project's budget. You want to update a budget report in Google Sheets and send a summary of the changes to your team via Slack.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://code.claude.com/docs/en/mcp" rel="noopener noreferrer"&gt;Claude Code&lt;/a&gt; (MCP host) initiates a request to the MCP client to update the budget report in Google Sheets and send a Slack notification.&lt;/li&gt;
&lt;li&gt;The MCP client connects to two MCP servers: one for Google Sheets and one for Slack.&lt;/li&gt;
&lt;li&gt;The &lt;a href="https://composio.dev/toolkits/googlesheets" rel="noopener noreferrer"&gt;Google Sheets MCP&lt;/a&gt; server calls the Google Sheets API (remote service) to update the budget report.&lt;/li&gt;
&lt;li&gt;The &lt;a href="https://composio.dev/toolkits/slack" rel="noopener noreferrer"&gt;Slack MCP&lt;/a&gt; server interacts with the Slack API (remote service) to send a notification.&lt;/li&gt;
&lt;li&gt;MCP servers send responses back to the MCP client.&lt;/li&gt;
&lt;li&gt;The MCP client forwards these responses to Cursor, which displays the result to the user.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This process happens seamlessly, allowing Cursor to integrate with multiple services through a standardized interface.&lt;/p&gt;

&lt;p&gt;That's the working in a nutshell. From here on, we will focus on building.&lt;/p&gt;




&lt;h2&gt;
  
  
  Build Local Stremable HTTP MCP Servers
&lt;/h2&gt;

&lt;p&gt;We'll build something you'd actually deploy: a &lt;strong&gt;GitHub issue search server&lt;/strong&gt; that lets your AI assistant search issues, fetch issue details, and pull recent pull requests from any repository. [subject to change]&lt;/p&gt;

&lt;p&gt;We will follow a series of steps to make sure things go smoothly. This is also industry standard, so it's best to follow along. &lt;/p&gt;

&lt;h3&gt;
  
  
  Prequires
&lt;/h3&gt;

&lt;p&gt;Before anything else, ensure you complete the prerequisites, as they will be the foundation for any MCP implementation we do.&lt;/p&gt;

&lt;p&gt;You'll need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Python 3.10+&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Node.js 18+&lt;/strong&gt; for TypeScript (if you choose to build with TypeScript)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Desktop&lt;/strong&gt; or &lt;strong&gt;Cursor&lt;/strong&gt; for testing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A common question users often ask: when should you pick Python vs. TypeScript? Here is an honest review:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.python.org/downloads/" rel="noopener noreferrer"&gt;Python&lt;/a&gt; with &lt;a href="https://gofastmcp.com/getting-started/welcome" rel="noopener noreferrer"&gt;FastMCP&lt;/a&gt; is faster for prototyping and works well for stdio servers.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.typescriptlang.org/" rel="noopener noreferrer"&gt;TypeScript&lt;/a&gt; shines for HTTP servers deployed to Cloudflare Workers, Vercel Edge, or any Node-based stack due to its type safety.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In case you are wondering, Composio is the ultimate integration platform, empowering developers to seamlessly connect AI agents with external tools, servers, and APIs with just a single line of code.&lt;/p&gt;

&lt;p&gt;With the fully managed &lt;a href="https://composio.dev/blog/best-mcp-servers-for-chatgpt" rel="noopener noreferrer"&gt;&lt;strong&gt;MCP Server&lt;/strong&gt;s&lt;/a&gt;, developers can rapidly build powerful AI applications without the hassle of managing complex integrations. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;In a nutshell: Composio handles the infrastructure so developers can focus on innovation.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Next, let's set up the working environment.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Work Environment Setup&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;We start by creating a project directory.&lt;/p&gt;

&lt;p&gt;Navigate to your working folder and create a folder named &lt;strong&gt;MCP,&lt;/strong&gt; or u can use the terminal command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir &lt;/span&gt;mcp_servers
&lt;span class="nb"&gt;cd &lt;/span&gt;mcp_servers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Next, create a virtual environment using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now activate the environment with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Windows:&lt;/span&gt;
.venv&lt;span class="se"&gt;\S&lt;/span&gt;cripts&lt;span class="se"&gt;\a&lt;/span&gt;ctivate

&lt;span class="c"&gt;# Linux/Mac:&lt;/span&gt;
&lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ensure you see (.venv) in Front of the terminal cwd path.&lt;/p&gt;

&lt;p&gt;Finally, install the FastMCP package&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;FastMCP
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Writing Server Code
&lt;/h3&gt;

&lt;p&gt;Server code is where the entire MCP server logic lives. Create a new &lt;code&gt;server.py&lt;/code&gt;  file and inside it paste the code below:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# server.py
&lt;/span&gt;
&lt;span class="c1"&gt;# imports
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastmcp&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastMCP&lt;/span&gt;

&lt;span class="n"&gt;mcp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastMCP&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Utility MCP Server&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# generate password utility
&lt;/span&gt;&lt;span class="nd"&gt;@mcp.tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;generate_password&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;

    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Generate a pseudo-random password for development use.

    This function creates a password by encoding a UUID4 string in base64
    and then reversing the resulting string. It&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s intended for use in
    development or testing environments where cryptographic security is not
    critical.

    Returns:
        str: A reversed base64-encoded UUID4 string, typically around 48 characters.
        May include alphanumeric characters, &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;+&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, and &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; padding.

    Security Notes:
        - Uses the string form of UUID4 (not raw bytes), so effective entropy is lower (~190 bits of input chars)
        - Suitable for temporary or development credentials.
        - NOT recommended for production use.
        - For higher security, consider using secrets.token_urlsafe().

    Example:
&lt;/span&gt;&lt;span class="gp"&gt;        &amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;generate_password&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;==dGM0ZTcwLTI0ZTUtNDJlZC05Y2I1LTk3OGJjODAxMDAwNw==&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;

    &lt;span class="n"&gt;uuid_string&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid4&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;base64_encoded&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;b64encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;uuid_string&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;base64_encoded&lt;/span&gt; &lt;span class="p"&gt;[::&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="c1"&gt;# list all the files in the directory
&lt;/span&gt;&lt;span class="nd"&gt;@mcp.tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;list_files&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;directory&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    List all files and directories in a specified path.

    This tool allows the AI to explore the file system by returning a list 
    of names of the entries in the directory given by path.

    Args:
        directory (str): The path to the directory to list. Defaults to the 
        current working directory (&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;).

    Returns:
        List[str]: A list of filenames and directory names. If an error 
        occurs (e.g., directory not found), returns a list containing 
        the error message.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;listdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;directory&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;


&lt;span class="c1"&gt;# read the given file
&lt;/span&gt;&lt;span class="nd"&gt;@mcp.tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;read_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Read and return the full text content of a file.

    This tool enables the AI to analyze the contents of specific text files 
    within the accessible file system.

    Args:
        filename (str): The path to the file that needs to be read.

    Returns:
        str: The complete string content of the file encoded in UTF-8. 
        If the file cannot be read, returns an error message starting 
        with &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Error reading file:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Error reading file: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# system info (resource)
&lt;/span&gt;&lt;span class="nd"&gt;@mcp.resource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data://sys-info&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_system_info&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Get a comprehensive system snapshot including time, hardware specs, 
    and real-time health metrics.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;platform&lt;/span&gt;
    &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;psutil&lt;/span&gt;
    &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;
    &lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;

    &lt;span class="n"&gt;now&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;timestamp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strftime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;%Y-%m-%d %H:%M:%S&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;day_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strftime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;%A&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


    &lt;span class="n"&gt;boot_time&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromtimestamp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;psutil&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;boot_time&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="n"&gt;uptime&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;now&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;boot_time&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="n"&gt;sys_type&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;platform&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;system&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;machine&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;platform&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;machine&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;cpu_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;psutil&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cpu_count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;logical&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;gpu_info&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;check_output&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nvidia-smi&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--query-gpu=gpu_name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--format=csv,noheader&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; 
            &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;gpu_info&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;N/A (NVIDIA-SMI not found)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="n"&gt;cpu_usage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;psutil&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cpu_percent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;interval&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;memory&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;psutil&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;virtual_memory&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;disk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;psutil&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;disk_usage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;C:&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
             --- Hardware Specs ---
            OS:        &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;sys_type&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;machine&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;)
            CPU:       &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;cpu_count&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; Physical Cores
            GPU:       &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;gpu_info&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

            --- Temporal Context ---
            Date/Time: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;timestamp&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
            Day:       &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;day_name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
            Uptime:    &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;uptime&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

            --- Resource Utilization ---
            CPU Load:  &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;cpu_usage&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;%
            RAM Usage: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;percent&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;% (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;used&lt;/span&gt; &lt;span class="o"&gt;//&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;MB / &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;//&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;MB)
            Disk (C:): &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;disk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;percent&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;% Used (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;disk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;free&lt;/span&gt; &lt;span class="o"&gt;//&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;GB Free)
            &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="c1"&gt;# general prompt
&lt;/span&gt;&lt;span class="nd"&gt;@mcp.prompt&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;helper_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;General purpose prompt for any task with optional context.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;base_prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
Task: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Please approach this systematically:
1. Break down the problem
2. Consider multiple solutions
3. Explain your reasoning
4. Provide clear, actionable steps
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;base_prompt&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s"&gt;Additional context: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;base_prompt&lt;/span&gt;

&lt;span class="c1"&gt;# api-design rules prompt
&lt;/span&gt;&lt;span class="nd"&gt;@mcp.prompt&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;api_endpoint_design_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;endpoint_name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;purpose&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Specific prompt for designing RESTful API endpoints.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
Design a RESTful API endpoint: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;endpoint_name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
Purpose: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;purpose&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Please provide:
1. HTTP method and URL pattern
2. Request/response schemas with examples
3. Status codes and error handling
4. Authentication requirements
5. Rate limiting considerations
6. OpenAPI/Swagger documentation snippet

Follow REST conventions and include proper validation.
Use JSON for data exchange and meaningful HTTP status codes.
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="c1"&gt;# Run MCP server
# if __name__ == "__main__":
&lt;/span&gt;    &lt;span class="c1"&gt;# mcp.run()
&lt;/span&gt;
&lt;span class="c1"&gt;# Run stremable http MCP Server (can be deployed / hosted)
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
   &lt;span class="n"&gt;mcp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;transport&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;streamable-http&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0.0.0.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Key points from the code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;mcp&lt;/code&gt; : It's the Fast MCP object named as provided. Important to initialize at the start.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;@mcp.tool()&lt;/code&gt; : Decorator wraps the entire function underneath and exposes it as a tool to the MCP client. Docstrings are very important; a good docstring leads to fewer model failures. Includes &lt;code&gt;generate_password&lt;/code&gt; , &lt;code&gt;list_files&lt;/code&gt;, &lt;code&gt;read_file&lt;/code&gt;tools.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;@mcp.resource("data://sys-info")&lt;/code&gt; : Decorator wraps the entire function and exposes it as a resource file that the client can use to fetch data. The path to the created resource is defined in &lt;code&gt;()&lt;/code&gt; . Includes &lt;code&gt;get_system_info&lt;/code&gt; resource&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;@mcp.prompt()&lt;/code&gt; : Decorator wraps the entire function and exposes it to the client as a prompt builder, useful when the task requires a specific use-case prompt. Docstrings are important here. Includes &lt;code&gt;helper_prompt&lt;/code&gt; , &lt;code&gt;api_endpoint_design_prompt&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;mcp.run(transport="streamable-http", host="0.0.0.0", port=8000)&lt;/code&gt; : Runs the MCP server.in a streamable HTTP mode. For stdio MCP server, use &lt;code&gt;mcp.run()&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Overall, we are wrapping different functions and exposing them to the client as the MCP tool, resources, and server.&lt;/p&gt;

&lt;p&gt;However, to use it with the client, additional config is required, so let's set that up. I am using a cursor for the demo.&lt;/p&gt;

&lt;p&gt;Head to the Cursor and create a new folder &lt;code&gt;.cursor&lt;/code&gt; , inside it creates a new file &lt;code&gt;mcp.json&lt;/code&gt; (the config file) and paste the following setup.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mcpServers&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utility-server&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:8000/mcp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In case you have used &lt;code&gt;mcp.run()&lt;/code&gt; then you will see the below configuration&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mcpServers&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utility_server&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stdio&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;command&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your-python-exe-path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;args&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your file path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In general I use local &lt;code&gt;stdio&lt;/code&gt; based servers and then when all test and specifications are met, I switch to &lt;code&gt;Streamble HTTP&lt;/code&gt; approach.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Quick Tip: To verify this works is by pasting the command and args together; if they start the server, you are good to go; else, provide the full path. I usually go for path of .venv’s &lt;code&gt;python.exe&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Open &lt;strong&gt;Cursor Settings,&lt;/strong&gt; then click &lt;strong&gt;Tools &amp;amp; MCPs&lt;/strong&gt;. You will see your local MCP name. Click Enable, &lt;/p&gt;

&lt;p&gt;To check, open the terminal, switch to the Output tab, select the MCP server, and check for the connected message.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3tup4lsejnoux9pnw2ws.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3tup4lsejnoux9pnw2ws.png" alt="CURSOR" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now we can test our MCP Server by opening the chat/agents window and invoking any of the defined tools via the prompt. I invoked the password-generator, as I often find myself asking Claude to "think of a 16-character password" (which is terrible).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Generate&lt;/span&gt; &lt;span class="n"&gt;password&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;show&lt;/span&gt; &lt;span class="n"&gt;mcp&lt;/span&gt; &lt;span class="n"&gt;tool&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here is the output:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fprrko7rhtl069aqb9qjg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fprrko7rhtl069aqb9qjg.png" alt="MCP" width="800" height="265"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Saw that Ran command, it means the MCP server asked me for permission. I kept it that way, as you never want to let an agent touch your system files. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: If docstring is poor, the mcp client might fall back to cli approach, the easiest way to fix this is add a rule to only use mcp servers for the task. You will find a similar approach in the project github repo as well.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;All this is great, but setting it up for testing a local MCP server is a hassle. Let's look at an easier approach, MCP Inspector.&lt;/p&gt;




&lt;h3&gt;
  
  
  &lt;strong&gt;Test with MCP Inspector&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;To test the MCP server's functionality, we will use the Fast MCP Inspector, a browser-based tool that connects to the server and lets us call tools directly, without an LLM layer. &lt;/p&gt;

&lt;p&gt;Run MCP Inspector, with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @modelcontextprotocol/inspector python server.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'll see a localhost URL. Open it in your browser, click &lt;strong&gt;Connect&lt;/strong&gt;, and you should see your server respond.&lt;/p&gt;

&lt;p&gt;Now you can run and test your tools, resources, prompts, and more right from the browser, quite handy for developers testing MCP servers' performance and understanding response schemas. Check out the &lt;a href="https://modelcontextprotocol.io/docs/tools/inspector" rel="noopener noreferrer"&gt;MCP Inspector Guide&lt;/a&gt; to learn more.&lt;/p&gt;

&lt;p&gt;Personally, I find it best for testing prompts’ responses. Here is the response it generated:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4vlqxiftlh4s5jx2qqm1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4vlqxiftlh4s5jx2qqm1.png" alt="MCP INSPECTOR" width="800" height="430"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Local servers are great for simple use cases, but for most industry use cases, they don't provide the required infrastructure. Let's explore better use cases.&lt;/p&gt;




&lt;h2&gt;
  
  
  Connect to Hosted Streamable HTTP MCP Server
&lt;/h2&gt;

&lt;p&gt;Having a local server is good, but not as great as the one used in production. Usually, these are hosted, secured by design, and optimized for multiple tool calls. &lt;/p&gt;

&lt;p&gt;Usually, these are servers written by expert developers, integrated with the app's internal mechanisms, and intended to serve as connectors. &lt;/p&gt;

&lt;p&gt;But here lies the issue: connecting to 30+ MCP servers for a single project is a pain. Composio solves this.&lt;/p&gt;

&lt;p&gt;Let's look at how you can connect to the hosted Composio MCP server. You connect it once and use 1K+ tools directly, all secure, optimized tool calls.&lt;/p&gt;




&lt;h3&gt;
  
  
  Connect to Composio MCP Servers (Streamable HTTP )
&lt;/h3&gt;

&lt;p&gt;Integrating with &lt;strong&gt;Composio MCP&lt;/strong&gt; is incredibly simple and takes just 5 steps.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Visit the &lt;a href="https://dashboard.composio.dev/" rel="noopener noreferrer"&gt;Composio MCP&lt;/a&gt; page. Ensure you are logged in, or else sign up&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0927enm97uwl0f7n50jk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0927enm97uwl0f7n50jk.png" alt="step1" width="800" height="472"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;In the dashboard, head to the Install Section&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fi0c27tywfcxieu44jbfn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fi0c27tywfcxieu44jbfn.png" alt="step2" width="800" height="473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Select Cursor from the list, and click on Install.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsxwj0ye9pv006o7kgfaz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsxwj0ye9pv006o7kgfaz.png" alt="step3" width="800" height="473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;In the redirected page, click on "Install in Cursor".&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqyp4ar0wxdo9pa1c7pqb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqyp4ar0wxdo9pa1c7pqb.png" alt="step4" width="800" height="475"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You will be redirected to Cursor, and in the MCP server Page, click Install&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fh79a70tod8hccs3liiw2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fh79a70tod8hccs3liiw2.png" alt="step5" width="800" height="475"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once installed, click on Needs Authentication, then Authorize at redirect, and done!&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fynux23lf9ww8tjxif6zn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fynux23lf9ww8tjxif6zn.png" alt="step6" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can connect to apps in advance in the Connect Apps section; most of them use OAuth. (optional). If you don't, the agent will prompt you to connect at runtime.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Febp6k5s8r0jxb2pn9klk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Febp6k5s8r0jxb2pn9klk.png" alt="step7" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;To test the integration, open the Composer, establish a connection to the app, and ask it to perform actions.&lt;/p&gt;

&lt;p&gt;But how did all this happen?&lt;/p&gt;

&lt;p&gt;It turns out that if you go to your cursor MCP &amp;amp; Tools Page, select the ✏️ icon, and see the configuration as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"composio"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"http"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://connect.composio.dev/mcp"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"headers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;strong&gt;http&lt;/strong&gt; here means a streamable HTTP server. This allows the MCP server to operate as a normal web server (accessible via a URL) rather than as a local subprocess on your machine (for local MCP servers)&lt;/p&gt;

&lt;p&gt;And interestingly, for any MCP to run in HTTP streamable mode, you have to replace:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;__name__&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;==&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"__main__"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="err"&gt;mcp.run()&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;__name__&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;==&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"__main__"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="err"&gt;mcp.run(transport=&lt;/span&gt;&lt;span class="s2"&gt;"streamable-http"&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;host=&lt;/span&gt;&lt;span class="s2"&gt;"0.0.0.0"&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;port=&lt;/span&gt;&lt;span class="mi"&gt;8000&lt;/span&gt;&lt;span class="err"&gt;)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And you are done, pretty handy, isn’t it?&lt;/p&gt;

&lt;p&gt;Now, let's look at an advanced use case to see where Composio MCP helps.&lt;/p&gt;




&lt;h2&gt;
  
  
  Using HTTP Streamable MCP Servers for Advanced Use Case (Financial Agent)
&lt;/h2&gt;

&lt;p&gt;Current financial institutions struggle with manual, time-intensive investment research that keeps analysts buried in data entry for weeks, delays critical decisions by days, lacks real-time risk visibility, creates compliance blind spots, and demands hiring more analysts just to manage portfolio volume - all while losing talent and clients&lt;/p&gt;

&lt;p&gt;Let's see how the analyst can build financial agents and leverage Composio MCP to connect to 10-15 financial data sources (Yahoo Finance, SEC filings, Google Sheets, Bloomberg, etc.) and enrich their analysis by aggregating structured APIs, unstructured web scraping, and internal data warehouses.&lt;/p&gt;

&lt;p&gt;The financial analyst agent will:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Aggregate financial data (Yahoo Finance, Google Sheets with company data, SEC filings via web scraping),&lt;/li&gt;
&lt;li&gt;Run Claude analysis to generate investment thesis.&lt;/li&gt;
&lt;li&gt;Create formatted reports in Google Docs,&lt;/li&gt;
&lt;li&gt;Track portfolio changes (pull from Excel sheet),&lt;/li&gt;
&lt;li&gt;and send Executive Summaries to stakeholders via email with embedded tables. (HTML format)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All done in parallel!&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;Web&lt;/span&gt; &lt;span class="nc"&gt;Scraping &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;composio&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;search&lt;/span&gt; &lt;span class="n"&gt;api&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;Google&lt;/span&gt; &lt;span class="nc"&gt;Sheets &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;composio&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="n"&gt;Claude&lt;/span&gt; &lt;span class="n"&gt;Analysis&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="n"&gt;Google&lt;/span&gt; &lt;span class="nc"&gt;Docs &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;composio&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="nc"&gt;Gmail &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;composio&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prerequisites
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code / Cursor / VS Code. I am going with Claude&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dashboard.composio.dev/login" rel="noopener noreferrer"&gt;Composio Account&lt;/a&gt;, login/signup&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Assuming you have pre-requisites met, let's set up Composio on Claude.&lt;/p&gt;

&lt;h3&gt;
  
  
  Set up Composio MCP in Claude
&lt;/h3&gt;

&lt;p&gt;To get started: &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Visit the &lt;a href="https://dashboard.composio.dev/" rel="noopener noreferrer"&gt;Composio MCP&lt;/a&gt; page. Ensure you are logged in, or else sign up&lt;/li&gt;
&lt;li&gt;In the dashboard, head to the Install Section&lt;/li&gt;
&lt;li&gt;Select Claude from the list, and click on Install.&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;On the next page, select the MCP tab, then copy the command &amp;amp; paste it into the shell/terminal.&lt;br&gt;
&lt;/p&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;claude&lt;/span&gt; &lt;span class="n"&gt;mcp&lt;/span&gt; &lt;span class="n"&gt;add&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="n"&gt;scope&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt; &lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="n"&gt;transport&lt;/span&gt; &lt;span class="n"&gt;http&lt;/span&gt; &lt;span class="n"&gt;composio&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;//&lt;/span&gt;&lt;span class="n"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;composio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dev&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;mcp&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ensure &lt;code&gt;/mcp&lt;/code&gt;  shows composio as an option. It might require authentication, so authenticate once&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fr1iskit9iof67v7z3kto.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fr1iskit9iof67v7z3kto.png" alt="terminal-mcp" width="800" height="167"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;(Optional ) Once done, head to the Connect Apps page and click on Connect to connect Gmail, Google Drive, Google Sheets, and Google Docs. In case you ignore, agent will ask for these connections on runtime.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Perfect, we only need to configure the document in the workspace&lt;/p&gt;

&lt;h3&gt;
  
  
  Set up Workspace for Agent
&lt;/h3&gt;

&lt;p&gt;We will use Google Drive as our workspace to make the file accessible from everywhere:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Head to drive, create a new folder named &lt;code&gt;financial_data&lt;/code&gt;  &amp;amp; from the URL, copy the folder ID.&lt;/li&gt;
&lt;li&gt;Inside, add an Excel file in the format &amp;amp; fill the data as:

&lt;ul&gt;
&lt;li&gt;Companies: List of all targeted companies with their analysis&lt;/li&gt;
&lt;li&gt;Portfolio Performance: How the portfolio is performing over time, all metrics for targeted companies&lt;/li&gt;
&lt;li&gt;Contact: List of all internal contacts: Portfolio Managers, Limited Partners, Investors,  Board Managers, and more&lt;/li&gt;
&lt;li&gt;Report Archive: Past report filing (sec)&lt;/li&gt;
&lt;/ul&gt;


&lt;/li&gt;

&lt;li&gt;Then add a Google Doc template for an investment report.&lt;/li&gt;

&lt;/ul&gt;

&lt;p&gt;Make sure to copy the ID / URL of both files&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Easiest way is to use the template given in the GitHub Repo and let any AI Agent turn them in the above given format. Then copy paste that in newly created google sheet &amp;amp; google doc file.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now is the time to run the agent.&lt;/p&gt;

&lt;h3&gt;
  
  
  Run Financial Analyst Agent
&lt;/h3&gt;

&lt;p&gt;To run the agent, open Claude Code and paste the &lt;a href="https://gist.github.com/DevloperHS/06982896ddfc20c8ace1af213dff2ebf" rel="noopener noreferrer"&gt;financial_agent_prompt&lt;/a&gt;. You will be prompted to fill in the details with the copied values from earlier, and add your email address:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;GOOGLE_SHEETS_ID&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;your&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;GOOGLE_DOC_TEMPLATE_ID&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;your&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;GOOGLE_DRIVE_FOLDER_ID&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;optional&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;GMAIL_SENDER_ADDRESS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;your&lt;/span&gt; &lt;span class="n"&gt;gmail&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once done, the agent handles the rest. Here is a demo of what it looks like, in production.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/_PTeGESx6Nk"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;The agent also works on multiple portfolio reports. I am testing it with six different portfolios; you can adjust the count as needed.&lt;/p&gt;

&lt;p&gt;Same setup, but with multiple portfolio data. Make sure to copy each one's ID, and use the financial_agent_prompt_for_multiple portfolio for the job.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/oJzYSJ3Y6T4"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Note that, despite explicitly stating that Claude never created a subagent, this is because composio already has parallel execution enabled and, behind the scenes, uses subagents.&lt;/p&gt;

&lt;p&gt;Hope this gave you an idea of how hosted MCP servers work, and tools like Composio MCP provide a universal tooling layer.&lt;/p&gt;

&lt;p&gt;However, deploying a hosted MCP server for production is a different game. Let's look at a few rules to help ensure the servers don't break in real-world use.&lt;/p&gt;




&lt;h2&gt;
  
  
  Deployment Notes for Production MCP Servers
&lt;/h2&gt;

&lt;p&gt;Though we didn't cover a hosted MCP server build, if you want to build your hosted MCP server (local + later deployed), follow the production checklist before you deploy &amp;amp; put an MCP server in front of real users:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Logging to stderr or files only&lt;/strong&gt;, never stdout for stdio servers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Errors caught on every tool&lt;/strong&gt;: return a clean error message, don't let exceptions kill the connection&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inputs validated&lt;/strong&gt;: type hints (Python) or Zod (TypeScript) catch most issues; add explicit checks for things like SQL injection or path traversal if you're doing database/file operations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;External API calls are rate-limited / Use Composio MCP&lt;/strong&gt;: your tool runs whenever the LLM decides to call it, which can be a lot. Let Composio MCP handle it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Secrets in environment variables&lt;/strong&gt;, never hardcode in servers, never commit in production.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool descriptions for the LLM&lt;/strong&gt;: Write specific, one-purpose per tool. Five small tools beat one giant tool. The approach we followed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema versioning&lt;/strong&gt;: when you change a tool's signature, bump your server's version so clients can adapt. Think API versioning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Health check endpoint&lt;/strong&gt;: for HTTP servers, expose &lt;code&gt;/health&lt;/code&gt; for your load balancer. Provides transparency to users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth on production HTTP servers&lt;/strong&gt;: Use Bearer for internal, OAuth for public. But never both for the same for public/internal (private)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With this, we have come to the end of this comprehensive guide. Here is what matters the most.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;As AI transforms software development, MCP will play an increasingly important role in creating seamless, integrated experiences.&lt;/p&gt;

&lt;p&gt;Whether you're building custom MCP servers or leveraging pre-built solutions like stremable http servers like Composio MCP, the protocol enables powerful enhancements to AI capabilities through external tools and data sources.  &lt;/p&gt;

&lt;p&gt;In case you need a few ideas, check out how to build an &lt;a href="https://composio.dev/blog/mcp-client-step-by-step-guide-to-building-from-scratch/" rel="noopener noreferrer"&gt;MCP client&lt;/a&gt; that talks to your server, or move up the stack and build an &lt;a href="https://composio.dev/blog/the-complete-guide-to-building-mcp-agents/" rel="noopener noreferrer"&gt;MCP-powered agent&lt;/a&gt; that orchestrates multiple servers at once. Happy building.&lt;/p&gt;

&lt;p&gt;I hope you had a great learning experience - happy building with Composio! 🚀&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>What Is an MCP Gateway — and Why Do Enterprise AI Teams Need One in 2026?</title>
      <dc:creator>Dumebi Okolo</dc:creator>
      <pubDate>Thu, 07 May 2026 14:29:51 +0000</pubDate>
      <link>https://dev.to/composiodev/what-is-an-mcp-gateway-and-why-do-enterprise-ai-teams-need-one-in-2026-1lie</link>
      <guid>https://dev.to/composiodev/what-is-an-mcp-gateway-and-why-do-enterprise-ai-teams-need-one-in-2026-1lie</guid>
      <description>&lt;p&gt;The &lt;a href="https://modelcontextprotocol.io/docs/learn/architecture" rel="noopener noreferrer"&gt;Model Context Protocol (MCP)&lt;/a&gt; was released by Anthropic in November 2024. Eighteen months later, it had 97 million monthly SDK downloads as of December 2025, backing from every major AI lab, and is now governed as a founding project of the Agentic AI Foundation (AAIF), a directed fund under the Linux Foundation.&lt;/p&gt;

&lt;p&gt;That adoption happened fast, even faster than most protocols manage. But it created an immediate problem: connecting AI agents directly to dozens of MCP servers at scale is operationally unsustainable, and the protocol itself does not solve governance.&lt;/p&gt;

&lt;p&gt;This article explains what an MCP Gateway is, what it does at the infrastructure level, and how to evaluate one for a production enterprise environment.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Is MCP, and What Problem Does It Solve?
&lt;/h2&gt;

&lt;p&gt;Before understanding the gateway, you need to understand what MCP standardizes.&lt;/p&gt;

&lt;p&gt;Enterprise AI teams historically faced what is called the &lt;a href="https://composio.dev/content/mcp-gateways-guide" rel="noopener noreferrer"&gt;N×M integration problem&lt;/a&gt;: connecting N agents to M tools requires N×M custom integrations, each with its own authentication flow, error-handling logic, and credential store. Without MCP, integration complexity rises quadratically as AI agents spread through an organization; with MCP, it scales linearly.&lt;/p&gt;

&lt;p&gt;MCP defines a standardized way for AI models to discover and invoke external tools using &lt;a href="https://www.jsonrpc.org/specification" rel="noopener noreferrer"&gt;JSON-RPC 2.0&lt;/a&gt; over HTTP. An agent sends a &lt;code&gt;tools/list&lt;/code&gt; request to understand what a server exposes, then uses &lt;code&gt;call_tool&lt;/code&gt; to invoke those tools. That handshake is consistent regardless of whether the backend is GitHub, Salesforce, Postgres, or an internal API.&lt;/p&gt;

&lt;p&gt;What MCP does not define is who can call what, under whose identity, with what constraints, and at what cost. Those are governance problems, and they fall outside the protocol specification by design.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Is an MCP Gateway?
&lt;/h2&gt;

&lt;p&gt;An MCP Gateway is a centralized infrastructure layer that sits between AI agents and one or more MCP servers. It acts as a &lt;a href="https://www.cloudflare.com/learning/cdn/glossary/reverse-proxy/" rel="noopener noreferrer"&gt;specialized reverse proxy&lt;/a&gt; purpose-built for MCP traffic: handling authentication, routing, policy enforcement, credential management, and observability in one place.&lt;/p&gt;

&lt;p&gt;From the agent's perspective, nothing changes. It still performs a &lt;code&gt;tools/list&lt;/code&gt; handshake and issues &lt;code&gt;call_tool&lt;/code&gt; requests. The difference is that those requests are now intercepted, evaluated against policies, and routed by the gateway before any backend system executes them.&lt;/p&gt;

&lt;p&gt;Architecturally, the shift looks like this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Without a gateway:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent A → GitHub MCP Server
Agent A → Slack MCP Server
Agent B → GitHub MCP Server
Agent B → Postgres MCP Server
Agent C → Salesforce MCP Server
... (N×M connections, each managing its own auth and credentials)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;With a gateway:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent A ──┐
Agent B ──┤──→ [MCP Gateway] ──→ GitHub MCP Server
Agent C ──┘                  ──→ Slack MCP Server
                             ──→ Postgres MCP Server
                             ──→ Salesforce MCP Server
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The gateway becomes the single chokepoint where security policy, access control, and observability can be enforced consistently. As one &lt;a href="https://news.ycombinator.com/item?id=46136222" rel="noopener noreferrer"&gt;Hacker News discussion on MCP gateways&lt;/a&gt; noted, practitioners want features like central MCP registries, OAuth integration, and curated toolset scoping; all things that make MCP viable at organizational scale, not just in a prototype.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Does MCP Alone Fall Short in Enterprise Environments?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Credential Sprawl
&lt;/h3&gt;

&lt;p&gt;Without a gateway, each agent carries its own API keys, OAuth tokens, and service account credentials for every tool it accesses. Those credentials end up in environment variables, config files, and secret stores scattered across services. This is not a theoretical risk: GitGuardian's research found 24,008 unique secrets exposed in MCP configuration files in 2025 alone, with Google API keys and PostgreSQL connection strings among the most common leaked types. Rotating credentials becomes a manual exercise across multiple codebases. Revoking access for a compromised agent requires hunting down every integration it touches. There is no single point of revocation.&lt;/p&gt;

&lt;h3&gt;
  
  
  No Centralized Access Control
&lt;/h3&gt;

&lt;p&gt;MCP does not define native role-based access control. If an agent can connect to a server, it can discover every tool that server exposes. A finance agent can see development tools. A support agent can see database administration endpoints. Principle of least privilege has to be implemented outside the protocol, in every agent individually, or not at all. As engineers in the &lt;a href="https://news.ycombinator.com/item?id=45723699" rel="noopener noreferrer"&gt;MCP-Scanner Hacker News thread&lt;/a&gt; observed, people are over-provisioning MCPs the way they install apps on a phone, without applying least-privilege access. &lt;/p&gt;

&lt;p&gt;Least-privilege access is the principle that an agent should only be able to see and invoke the specific tools it needs for its defined task, and nothing beyond that. In an MCP context, this means a support agent should have no visibility into deployment tools, and a read-only analytics agent should have no access to write operations, regardless of what the underlying server exposes..&lt;/p&gt;

&lt;h3&gt;
  
  
  Observability Black Holes
&lt;/h3&gt;

&lt;p&gt;When agents connect directly to tools, there is no aggregated view of what any agent is actually doing. Debugging a multi-step workflow requires stitching together logs from N different servers. There is no unified execution timeline, no trace correlation, no cost attribution. Anomalies go undetected because there is no baseline.&lt;/p&gt;

&lt;h3&gt;
  
  
  No Cost Governance
&lt;/h3&gt;

&lt;p&gt;MCP does not track token consumption or enforce usage limits. An agent can invoke tools repeatedly, triggering LLM calls and paid API operations, with no budget ceiling. At enterprise scale, this becomes a financial control problem, not just a technical one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Security Attack Surface
&lt;/h3&gt;

&lt;p&gt;In April 2025, security researchers &lt;a href="https://en.wikipedia.org/wiki/Model_Context_Protocol" rel="noopener noreferrer"&gt;published an analysis&lt;/a&gt; identifying multiple outstanding MCP security issues, including prompt injection, tool permissions that allow combining tools to exfiltrate data, and lookalike tools that can silently replace trusted ones. A centralized gateway is the practical enforcement point for mitigating all three.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Does an MCP Gateway Actually Do?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Centralized Authentication and Identity Propagation
&lt;/h3&gt;

&lt;p&gt;A production gateway validates incoming identity  (typically via JWT, OAuth 2.0 with PKCE, or OIDC) and propagates that identity downstream to MCP servers. Instead of agents running under shared service accounts, requests execute on behalf of specific authenticated users.&lt;/p&gt;

&lt;p&gt;This closes a real vulnerability. If a user cannot delete a repository, neither can the agent acting for them. Authorization is enforced at the protocol layer, not assumed in prompts. The MCP specification introduced OAuth 2.1 support in the March 2025 revision, with significant refinements in June 2025, but implementation quality varies between gateways. Some handle enterprise SSO automatically; others require manual configuration per server.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tool-Level RBAC
&lt;/h3&gt;

&lt;p&gt;The gateway intercepts &lt;code&gt;tools/list&lt;/code&gt; responses and filters them based on the requesting agent's role and permissions. Sensitive tools simply do not appear in the agent's context. A configuration like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;virtual_server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;support-scope&lt;/span&gt;
  &lt;span class="na"&gt;allow_tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;github.list_issues&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;github.get_comments&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;crm.update_ticket&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;...means the agent calling this endpoint never sees database administration tools, deployment controls, or any capability it has no business using. This directly improves model performance, agents reason more accurately when the action space is deliberately constrained, and reduces blast radius when something goes wrong.&lt;/p&gt;

&lt;h3&gt;
  
  
  Intelligent Routing
&lt;/h3&gt;

&lt;p&gt;The gateway examines each request and routes it to the appropriate upstream MCP server based on the tool being called. Session affinity keeps stateful, multi-step agent conversations on the same backend server. Load balancing distributes traffic. Circuit breakers prevent cascading failures when an upstream tool degrades.&lt;/p&gt;

&lt;h3&gt;
  
  
  Unified Observability
&lt;/h3&gt;

&lt;p&gt;Every &lt;code&gt;tools/list&lt;/code&gt; and &lt;code&gt;call_tool&lt;/code&gt; invocation is logged with metadata: agent identity, user context, tool arguments, response status, and latency. This creates a coherent audit trail across all connected systems. Metrics export in Prometheus format. Traces follow the &lt;a href="https://opentelemetry.io/docs/what-is-opentelemetry/" rel="noopener noreferrer"&gt;OpenTelemetry standard&lt;/a&gt; for distributed tracing, which matters when debugging multi-step agent tasks that touch six different tools.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost Management
&lt;/h3&gt;

&lt;p&gt;The gateway can implement caching for repeated tool calls, enforce per-agent or per-user rate limits, and surface usage analytics. Caching strategies for repeated tool calls can meaningfully reduce LLM costs, and the gateway is the practical place to implement this at scale.&lt;/p&gt;

&lt;h3&gt;
  
  
  Credential Vaulting
&lt;/h3&gt;

&lt;p&gt;API keys, OAuth tokens, and service credentials are stored centrally in the gateway. Agents never handle raw credentials directly. Rotation policies apply once at the gateway level rather than across every agent codebase.&lt;/p&gt;




&lt;h2&gt;
  
  
  How Does an MCP Gateway Differ from an API Gateway?
&lt;/h2&gt;

&lt;p&gt;A traditional API gateway is designed for stateless, client-server request-response cycles, standard in web and mobile applications. It handles HTTP routing, authentication, rate limiting, and transformation for REST or GraphQL traffic.&lt;/p&gt;

&lt;p&gt;An MCP gateway is designed for stateful, session-aware, and often bidirectional communication patterns specific to AI agents. It understands the context of a long-running agent task. It can propagate user identity across multiple sequential tool calls. It maintains session state so that a multi-step agent workflow does not lose context mid-execution. It understands the &lt;code&gt;tools/list&lt;/code&gt; → &lt;code&gt;call_tool&lt;/code&gt; protocol cycle and can enforce policies at that semantic level, not just at the HTTP layer.&lt;/p&gt;

&lt;p&gt;In modern enterprise architectures, both typically coexist. APIs serve application services. API gateways govern traditional HTTP traffic. MCP servers expose selected capabilities to agents. An MCP gateway governs agent-to-tool communication. The relationship is complementary.&lt;/p&gt;

&lt;h3&gt;
  
  
  How Does an MCP Gateway Differ from an AI Gateway?
&lt;/h3&gt;

&lt;p&gt;This is worth separating out because it's a more common source of confusion in practice. Buyers evaluating AI gateways frequently find themselves looking at MCP gateways instead.&lt;/p&gt;

&lt;p&gt;An AI gateway sits in front of LLM inference. It manages which model gets called, routes traffic between providers (OpenAI, Anthropic, Mistral), enforces token budgets, handles prompt/response logging, and abstracts model provider APIs behind a single interface. Its job is governing &lt;em&gt;model calls&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;An MCP gateway sits between agents and the tools those agents invoke. It governs &lt;em&gt;tool calls:&lt;/em&gt; what an agent can do after the model has already decided to act. The two layers are complementary: an AI gateway controls which brain your agent uses; an MCP gateway controls which hands it has.&lt;/p&gt;

&lt;p&gt;In a mature enterprise architecture, both are present. The AI gateway handles model-level traffic. The MCP gateway handles the downstream tool execution that the model's output triggers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Are the Categories of MCP Gateway Available?
&lt;/h2&gt;

&lt;p&gt;Understanding the gateway landscape requires understanding the primary design philosophies, not just the feature checklist.&lt;/p&gt;

&lt;h3&gt;
  
  
  Managed Integration Platforms
&lt;/h3&gt;

&lt;p&gt;These prioritize developer velocity by abstracting integration complexity behind a large library of pre-built, maintained connectors. Authentication lifecycle management  (including complex OAuth 2.1 flows) is handled for you.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://composio.dev/mcp-gateway" rel="noopener noreferrer"&gt;Composio's MCP Gateway&lt;/a&gt; is the primary example. It offers 1000+ tools and actions across major enterprise SaaS applications, a unified authentication layer, SOC2 and ISO certification, action-level RBAC, and zero data-retention architecture. The architecture is designed for teams that need to connect agents to many different tools quickly without owning the integration layer: instead of juggling 22 different MCP servers for 22 different tools, you install one gateway and access a broad library of pre-built integrations with a single authentication flow and audit surface.&lt;/p&gt;

&lt;p&gt;For most enterprise teams moving from pilot to production, this is the most practical starting point. Refer to the &lt;a href="https://composio.dev/content/mcp-gateways-guide" rel="noopener noreferrer"&gt;Composio guide to MCP gateways&lt;/a&gt; for a deeper walkthrough of the architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Security-First Proxies
&lt;/h3&gt;

&lt;p&gt;These treat security as the primary constraint and performance as secondary. &lt;a href="https://github.com/lasso-security/mcp-gateway" rel="noopener noreferrer"&gt;Lasso Security&lt;/a&gt; inspects all MCP traffic in real time to detect prompt injection, mask PII, and calculate reputation scores for MCP servers before they are loaded. The tradeoff is latency — deep security scanning adds 100–250ms overhead — which makes this category unsuitable for latency-sensitive workflows but appropriate for regulated environments where compliance is non-negotiable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Infrastructure-Native Open Source
&lt;/h3&gt;

&lt;p&gt;These integrate into existing container-native DevOps workflows. &lt;a href="https://docs.docker.com/ai/mcp-catalog-and-toolkit/mcp-gateway/" rel="noopener noreferrer"&gt;Docker MCP Gateway&lt;/a&gt; runs MCP servers as isolated Docker containers with familiar &lt;code&gt;docker mcp&lt;/code&gt; CLI tooling and container-based security. &lt;a href="https://obot.ai/" rel="noopener noreferrer"&gt;Obot&lt;/a&gt; is Kubernetes-native and designed for organizations that require full data sovereignty.&lt;/p&gt;

&lt;p&gt;Both require your team to own the integration layer. Your team  brings the MCP servers, and the gateway governs them. The operational overhead is higher than a managed platform, but so is the control.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Should Enterprise Teams Evaluate When Choosing a Gateway?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Deployment Model
&lt;/h3&gt;

&lt;p&gt;Cloud-hosted managed gateways reduce time-to-production but involve data transiting external infrastructure. Self-hosted or VPC-deployed gateways give you data sovereignty. For teams in healthcare, finance, or government where regulated data must stay in your cloud, deployment model is often the first filter, not an afterthought.&lt;/p&gt;

&lt;h3&gt;
  
  
  Authentication Standards
&lt;/h3&gt;

&lt;p&gt;Verify support for OAuth 2.1 with PKCE, OIDC, and SAML. Check whether the gateway integrates with your existing identity provider (Okta, Microsoft Entra ID, Auth0) and whether it supports on-behalf-of token propagation: the pattern where agents act under the authenticated user's identity rather than a shared service account.&lt;/p&gt;

&lt;h3&gt;
  
  
  RBAC Granularity
&lt;/h3&gt;

&lt;p&gt;Gateway-level RBAC (which tools each role can see) is the baseline. Tool-level RBAC, allowing read but not write within a single server, is more sophisticated and significantly reduces blast radius. Verify what the enforcement model looks like in practice, not just in the marketing copy.&lt;/p&gt;

&lt;h3&gt;
  
  
  Observability Depth
&lt;/h3&gt;

&lt;p&gt;Prometheus-compatible metrics and OpenTelemetry traces are the minimum. Look for whether the gateway can attribute tool calls to specific users and agents (not just service accounts), whether audit logs meet your compliance format requirements, and whether the dashboard supports anomaly detection or cost attribution, and whether the gateway offers a zero data retention architecture — meaning tool call payloads and credentials are never stored on the gateway provider's infrastructure, which matters for regulated industries and data sovereignty requirements.&lt;/p&gt;

&lt;h3&gt;
  
  
  Integration Breadth vs. Governance Depth
&lt;/h3&gt;

&lt;p&gt;Managed platforms offer wide integration libraries but less control over the underlying infrastructure. Governance-first platforms offer deep control but require you to bring your own servers. For teams that need both, a large library of managed integrations and enterprise-grade governance, &lt;a href="https://composio.dev/mcp-gateway" rel="noopener noreferrer"&gt;Composio's MCP Gateway&lt;/a&gt; is the only option currently combining 500+ tools and actions with SOC2 compliance, RBAC, and zero data retention in a single product.&lt;/p&gt;

&lt;p&gt;See the full comparison in &lt;a href="https://composio.dev/content/best-mcp-gateway-for-developers" rel="noopener noreferrer"&gt;Composio's breakdown of the best MCP gateways for developers&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Performance Overhead
&lt;/h3&gt;

&lt;p&gt;Every proxy adds latency. Managed platforms typically run under 10ms overhead. TrueFoundry publishes under 5ms p95. Lunar.dev MCPX publishes approximately 4ms p99. Docker MCP Gateway adds overhead due to container management; warm-path performance is significantly better than cold-start, which can add 50–200ms. Lasso Security adds 100–250ms. For conversational agents where response time is visible to users, this matters. For background automation workflows, it typically does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building Your Own MCP Gateway
&lt;/h2&gt;

&lt;p&gt;Building a custom gateway is possible but requires solving non-trivial distributed systems problems: credential rotation, distributed rate limiting, OAuth 2.1 state management, PII redaction, and circuit breakers. The ongoing maintenance burden as the MCP spec grows as tool APIs change and security requirements mature is the real cost, not the initial build. For most teams, a managed gateway has a significantly lower total cost of ownership than a DIY solution, even when accounting for licensing costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Note on the MCP Security Threat Landscape
&lt;/h2&gt;

&lt;p&gt;Security threats against MCP deployments are not theoretical. A representative risk: an agent running with privileged service-role access that processes user-supplied input could inadvertently execute those instructions, exfiltrating sensitive data through legitimate output channels. Principle of least privilege at the gateway level is the primary defense.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://owasp.org/www-project-top-10-for-large-language-model-applications/" rel="noopener noreferrer"&gt;OWASP guidance on LLM security&lt;/a&gt; identifies prompt injection as among the highest-risk attack vectors for AI systems. An MCP gateway is the practical enforcement layer for mitigating it through input validation against JSON-RPC schemas, allowlisted actions, PII redaction, and real-time tool reputation scoring.&lt;/p&gt;

&lt;p&gt;Without a gateway, the security posture of your MCP deployment is only as strong as the weakest link among N independently managed agents.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;How much latency does a gateway add?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Managed platforms: typically under 10ms overhead. High-performance purpose-built gateways (TrueFoundry, Lunar.dev MCPX): under 5ms p99. Security-scanning gateways (Lasso Security): 100–250ms depending on inspection depth. Docker MCP Gateway warm-path latency is low; cold-start overhead can add 50–200ms.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Comes Next for MCP Gateways
&lt;/h2&gt;

&lt;p&gt;Based on MCP's published direction and community discussions from early 2026, four priority areas have emerged: transport evolution (stateless Streamable HTTP for load balancer compatibility), agent communication primitives (retry semantics and expiry policies for the Tasks primitive), governance maturation (formal contributor processes), and enterprise readiness (audit trails, SSO-integrated auth, and gateway patterns).&lt;/p&gt;

&lt;p&gt;Gateway patterns are now explicitly on the protocol roadmap. The gateway layer is no  longer an addon but is becoming formalized infrastructure for enterprise MCP deployments.&lt;/p&gt;

&lt;p&gt;Start with your primary constraint. If it is integration velocity, a managed platform is the right answer. If it is compliance in a regulated industry, prioritize SOC 2 certification, audit log format, and IdP integration. If it is data sovereignty, evaluate VPC-deployable options. If it is raw performance for a latency-sensitive conversational product, benchmark the p95 numbers against your SLA.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://composio.dev/mcp-gateway" rel="noopener noreferrer"&gt;Composio MCP Gateway&lt;/a&gt; covers the first and most common case: an enterprise team that needs to move from prototype to production with a broad integration library, unified auth, and compliance controls without owning the infrastructure. For teams with narrower requirements or existing MCP server infrastructure, the list of specialized options covered above gives you the tradeoffs needed to make that call.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;For a deeper look at gateway architecture patterns, see &lt;a href="https://composio.dev/content/mcp-gateways-guide" rel="noopener noreferrer"&gt;Composio's developer guide to MCP gateways&lt;/a&gt;. For a full comparison of gateway options by use case, see &lt;a href="https://composio.dev/content/best-mcp-gateway-for-developers" rel="noopener noreferrer"&gt;the best MCP gateways for developers in 2026&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;What is an MCP Gateway, in one sentence?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;A centralized infrastructure layer between AI agents and MCP servers that enforces authentication, routes requests, applies access controls, and provides observability across all agent-tool interactions.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Is an MCP Gateway required for production deployments?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Not required by the protocol specification. Required in practice for any deployment with more than two or three MCP servers, multiple teams, regulated data, or compliance obligations.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;What is the difference between an MCP server and an MCP gateway?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;An MCP server executes tools. It connects to GitHub, Postgres, Slack, or an internal API and performs operations. An MCP gateway governs access to those servers. It handles identity, visibility filtering, policy enforcement, and routing before any tool executes.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;How do MCP gateways handle prompt injection?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Security-first gateways like Lasso Security scan all traffic in real time and block payloads that trigger injection detection. Governance platforms like MintMCP apply input schema validation and allowlisted actions. Managed platforms like Composio run tool implementations in sandboxed environments. Using multiple layers of defense is the current best practice.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;What authentication standards should my gateway support?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;OAuth 2.1 with PKCE, OIDC, SAML, and support for enterprise IdPs. The MCP specification introduced OAuth 2.1 in the March 2025 revision with refinements in June 2025, but implementation quality varies significantly. Test the on-behalf-of identity propagation flow specifically. This is where implementations most commonly diverge.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
