<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nitish kumar</title>
    <description>The latest articles on DEV Community by Nitish kumar (@nitish-builds).</description>
    <link>https://dev.to/nitish-builds</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4059254%2F0cee8545-4ef1-42cc-b217-7b6fbf4b230d.png</url>
      <title>DEV Community: Nitish kumar</title>
      <link>https://dev.to/nitish-builds</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nitish-builds"/>
    <language>en</language>
    <item>
      <title>Our first AI agent task board had a status column agents could write to whenever they liked.

Tasks got stuck. Updates raced each other. The UI showed states that made no sense.

The fix was boring and it worked. Full write-up 👇</title>
      <dc:creator>Nitish kumar</dc:creator>
      <pubDate>Sat, 03 Oct 2026 13:33:42 +0000</pubDate>
      <link>https://dev.to/nitish-builds/our-first-ai-agent-task-board-had-a-status-column-agents-could-write-to-whenever-they-liked-46h1</link>
      <guid>https://dev.to/nitish-builds/our-first-ai-agent-task-board-had-a-status-column-agents-could-write-to-whenever-they-liked-46h1</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/nitish-builds/how-we-built-a-live-kanban-board-for-ai-agents-5h4m" class="crayons-story__hidden-navigation-link"&gt;How we built a live Kanban board for AI agents&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/nitish-builds" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4059254%2F0cee8545-4ef1-42cc-b217-7b6fbf4b230d.png" alt="nitish-builds profile" class="crayons-avatar__image" width="800" height="640"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/nitish-builds" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Nitish kumar
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Nitish kumar
                
                
              
              &lt;div id="story-author-preview-content-4792198" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/nitish-builds" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4059254%2F0cee8545-4ef1-42cc-b217-7b6fbf4b230d.png" class="crayons-avatar__image" alt="" width="800" height="640"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Nitish kumar&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/nitish-builds/how-we-built-a-live-kanban-board-for-ai-agents-5h4m" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Oct 3&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/nitish-builds/how-we-built-a-live-kanban-board-for-ai-agents-5h4m" id="article-link-4792198"&gt;
          How we built a live Kanban board for AI agents
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/productivity"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;productivity&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/architecture"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;architecture&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/webdev"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;webdev&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/nitish-builds/how-we-built-a-live-kanban-board-for-ai-agents-5h4m#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            3 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>automation</category>
    </item>
    <item>
      <title>How we built a live Kanban board for AI agents</title>
      <dc:creator>Nitish kumar</dc:creator>
      <pubDate>Sat, 03 Oct 2026 13:30:45 +0000</pubDate>
      <link>https://dev.to/nitish-builds/how-we-built-a-live-kanban-board-for-ai-agents-5h4m</link>
      <guid>https://dev.to/nitish-builds/how-we-built-a-live-kanban-board-for-ai-agents-5h4m</guid>
      <description>&lt;p&gt;At Deskferry we run AI agents that work inside tools like Gmail, Slack and HubSpot. For a long time, the only way to see what an agent was doing was a run log. Useful for debugging, useless for actually managing work.&lt;/p&gt;

&lt;p&gt;So we built Tasks: a Kanban board where users plan work, assign it to agents, and watch cards move from Backlog to Done in real time. This post covers the three design problems that mattered most: modelling task state, pausing for human approval, and keeping the board live.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. A task is a state machine, not a status field&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first version had a status column and agents wrote to it whenever they felt like it. That fell apart fast: tasks got stuck, two updates raced each other, and the UI showed states that made no sense.&lt;/p&gt;

&lt;p&gt;We replaced it with an explicit state machine. Every task is in exactly one state, and only defined transitions are allowed:&lt;/p&gt;

&lt;p&gt;backlog     -&amp;gt; in_flight      (agent assigned and picks it up)&lt;br&gt;
in_flight   -&amp;gt; needs_you      (agent hits an action that requires approval)&lt;br&gt;
needs_you   -&amp;gt; in_flight      (user approves or answers)&lt;br&gt;
needs_you   -&amp;gt; backlog        (user rejects, task returns for rework)&lt;br&gt;
in_flight   -&amp;gt; done           (agent completes all steps)&lt;br&gt;
in_flight   -&amp;gt; backlog        (agent fails; error attached to the card)&lt;/p&gt;

&lt;p&gt;The four board columns map directly to these states, so the UI can never show something the backend doesn't consider valid. Every transition is written as an event with a timestamp and the actor (agent or user), which gives us an audit trail for free.&lt;/p&gt;

&lt;p&gt;Lesson: if agents and humans both change the same object, make illegal states impossible at the data layer, not in the UI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Human-in-the-loop as a first-class state&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most human in the loop AI setups we've seen treat approval as a blocking prompt: the agent waits on a promise until someone clicks a button. That's fine in a demo and terrible in production. Users approve things hours later, servers restart, and you can't hold a process open all afternoon.&lt;/p&gt;

&lt;p&gt;We treat "waiting for a human" as a durable state instead:&lt;/p&gt;

&lt;p&gt;Before any action tagged as sensitive (send, spend, write to a record), the agent stops and serialises what it was about to do, including the full proposed action and the context it used to decide.&lt;br&gt;
The task transitions to needs_you and an approval item appears on the user's Desk.&lt;br&gt;
The agent process ends. Nothing is held in memory.&lt;br&gt;
When the user approves, edits or rejects, we rehydrate the agent from the saved checkpoint and continue from the exact step it paused on.&lt;/p&gt;

&lt;p&gt;This makes approvals cheap. A task can sit in Needs You for a minute or a weekend with no cost. It also makes the approval screen better, because the user sees precisely what will happen, not a vague "Agent wants to continue."&lt;/p&gt;

&lt;p&gt;Lesson: approval is a pause in a workflow, not a modal dialog. Persist it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Keeping the board live&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A board where cards move on their own only works if the user actually sees them move. Polling every few seconds felt sluggish and wasted requests on idle boards.&lt;/p&gt;

&lt;p&gt;Instead, every state transition from point 1 publishes an event, and the board subscribes to events for the current workspace over a persistent connection. The client applies each event to its local copy of the board, so a card slides between columns the moment the backend commits the change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two details made this reliable:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Events carry a version number per task. The client ignores any event older than what it already has, which handles out-of-order delivery after a reconnect.&lt;/p&gt;

&lt;p&gt;On reconnect, the client refetches the board once, then resumes applying events. Simple, and it avoids building a replay system.&lt;/p&gt;

&lt;p&gt;Lesson: for AI agent monitoring, the event log you already need for auditing is also the best feed for your UI.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What changed for users&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The technical work was in service of one product shift. Before Tasks, people had to design an AI agent workflow before getting any value. With Tasks, they write a sentence and assign it. The agent works out the steps, the board shows progress, and the Desk collects every decision that needs a human.&lt;/p&gt;

&lt;p&gt;That turned out to be the right abstraction for AI agent orchestration from the user's side: not a graph of nodes, but a list of tasks with owners.&lt;/p&gt;

&lt;p&gt;Takeaways if you're building something similar&lt;br&gt;
Model agent work as an explicit state machine with logged transitions.&lt;br&gt;
Make human approval a persisted state with checkpoints, not a blocking call.&lt;br&gt;
Drive your live UI from the same events you use for auditing.&lt;br&gt;
Show users tasks and owners, not pipelines. They already know how to manage those.&lt;/p&gt;

&lt;p&gt;Tasks is live in Deskferry now. If you want to see the board in action, you can try it free at deskferry.app. Happy to answer questions about any of this in the comments.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>architecture</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Your AI employees can now work as a team: introducing Agent Orchestration</title>
      <dc:creator>Nitish kumar</dc:creator>
      <pubDate>Wed, 02 Sep 2026 18:29:32 +0000</pubDate>
      <link>https://dev.to/deskferry/your-ai-employees-can-now-work-as-a-team-introducing-agent-orchestration-gfo</link>
      <guid>https://dev.to/deskferry/your-ai-employees-can-now-work-as-a-team-introducing-agent-orchestration-gfo</guid>
      <description>&lt;p&gt;Until now, a DeskFerry agent was a very good solo employee. You described a job in plain English, it connected to your tools, and it ran that job on a schedule or trigger. That worked well for one job at a time.&lt;/p&gt;

&lt;p&gt;But real work isn't one job. Qualifying a lead is research, then scoring, then CRM hygiene, then outreach, then follow-up. Each of those is a different skill, with different tools and different failure modes. Cramming all of it into a single agent meant one giant prompt, one giant blast radius, and a lot of "why did it do that?"&lt;/p&gt;

&lt;p&gt;As of September 1, that changes. Agent Orchestration is live in DeskFerry.&lt;/p&gt;

&lt;p&gt;What it does&lt;/p&gt;

&lt;p&gt;Orchestration lets one agent coordinate others. You give a brief to a lead agent, and it breaks the work down and delegates each piece to a specialist agent you've already built (or lets DeskFerry build the missing ones for you).&lt;/p&gt;

&lt;p&gt;Concretely, agents can now:&lt;/p&gt;

&lt;p&gt;Delegate. A lead agent hands a sub-task to another agent and waits for the result, the same way a manager assigns work.&lt;br&gt;
Share context. The research agent's findings are available to the outreach agent without you copying anything between them. Memory is scoped to the workflow, not trapped in one agent.&lt;br&gt;
Hand off mid-flight. A support triage agent can escalate a ticket to a billing agent with the full thread attached, then step back in when billing is done.&lt;br&gt;
Run in parallel. Independent steps (enrich five leads, check three inboxes) fan out and come back together.&lt;br&gt;
Pause for humans where it matters. Approvals work exactly as before, but now you can put the checkpoint at the team level — approve the final email, not every intermediate step.&lt;br&gt;
What it looks like&lt;/p&gt;

&lt;p&gt;You still describe the outcome, not the plumbing:&lt;/p&gt;

&lt;p&gt;Every weekday at 8am, pull new HubSpot leads, research each company, score them against our ICP, draft a first-touch email for anything above 70, and post the drafts to #sales for approval before sending.&lt;/p&gt;

&lt;p&gt;DeskFerry turns that into a small team: a coordinator, a researcher, a scorer, and a writer. Each one is a normal DeskFerry agent — you can open it, edit its instructions, swap its model, or reuse it in a different workflow. The scorer you built for inbound leads can be the same scorer your partnerships agent calls next month.&lt;/p&gt;

&lt;p&gt;Why we built it this way&lt;/p&gt;

&lt;p&gt;We looked at two approaches. One was a visual canvas where you draw boxes and arrows between agents. The other was letting you describe the team the way you'd describe it to a new hire. We chose the second, because the whole point of DeskFerry is that you shouldn't need to become a workflow architect to get work done.&lt;/p&gt;

&lt;p&gt;Under the hood, the coordinator is doing the planning, but everything it decides is visible. Every delegation, every handoff, and every intermediate result shows up in the run log, so when something goes sideways you can see exactly which agent made which call.&lt;/p&gt;

&lt;p&gt;What it means in practice&lt;br&gt;
Smaller, more reliable agents. A specialist with one job and three tools is much harder to confuse than a generalist with fifteen.&lt;br&gt;
Reuse. Build a skill once, call it from anywhere.&lt;br&gt;
Less babysitting. The coordinator handles retries and routing so you don't get paged for a single flaky step.&lt;br&gt;
&lt;a href="https://deskferry.app" rel="noopener noreferrer"&gt;Try it&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Orchestration is available now. Open any existing agent, and you'll see a new option to let it delegate to other agents, or start fresh with a multi-step brief and let DeskFerry assemble the team.&lt;/p&gt;

&lt;p&gt;If you're on a trial, this is included. If you're not, the trial is free and doesn't auto-convert.&lt;/p&gt;

&lt;p&gt;We'd love to hear what you build. Drop your workflows in the comments, or tell us where orchestration falls short — that's the feedback that shapes what ships next.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>agents</category>
      <category>nocode</category>
    </item>
    <item>
      <title>Your agent doesn't need more tools, it needs better tool descriptions</title>
      <dc:creator>Nitish kumar</dc:creator>
      <pubDate>Sun, 02 Aug 2026 16:31:35 +0000</pubDate>
      <link>https://dev.to/nitish-builds/your-agent-doesnt-need-more-tools-it-needs-better-tool-descriptions-4d28</link>
      <guid>https://dev.to/nitish-builds/your-agent-doesnt-need-more-tools-it-needs-better-tool-descriptions-4d28</guid>
      <description>&lt;p&gt;Most agent failures I've debugged weren't reasoning failures. The model reasoned fine. It just picked the wrong tool, because the tool description didn't tell it what it needed to know.&lt;/p&gt;

&lt;p&gt;This is an under-discussed problem, and it gets worse the more integrations you add. Here's what we've learned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The setup&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Say your agent has four calendar-ish tools available:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"calendar_create_event"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Creates a calendar event"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"calendar_quick_add"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Adds an event from text"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"scheduling_book_slot"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Books a scheduling slot"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"meeting_schedule"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Schedules a meeting"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Now: "Book me 30 minutes with Sarah on Thursday."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Which one? You don't know either, and you have context the model doesn't. The model will pick one, it'll probably work about 60% of the time, and when it's wrong the failure will be silent — an event created in the wrong system that nobody notices for a week.&lt;/p&gt;

&lt;p&gt;Roughly 40% of our agent failures traced back to tool selection, not reasoning. That was surprising and it changed where we spent engineering time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why OpenAPI schemas aren't enough&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A schema tells you what a tool accepts. It doesn't tell you:&lt;/p&gt;

&lt;p&gt;What the tool is for — the use case, not the signature&lt;br&gt;
When it's appropriate versus the four adjacent tools that look identical&lt;br&gt;
Whether it's reversible — can you undo it, and how&lt;br&gt;
What it costs — a rate-limited call and a free one shouldn't be equally attractive&lt;br&gt;
What it assumes — does it require the user's calendar to be connected first?&lt;/p&gt;

&lt;p&gt;Documentation doesn't have this either, because docs are written for humans who already know which product they're integrating with. The model is choosing between products.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What a description for an agent looks like&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's the template we converged on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;calendar_create_event&lt;/span&gt;
&lt;span class="na"&gt;purpose&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="s"&gt;Create an event on the user's primary Google Calendar when you already&lt;/span&gt;
  &lt;span class="s"&gt;know the exact date, time, and duration.&lt;/span&gt;
&lt;span class="na"&gt;use_when&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;The time is fully specified ("Thursday 2pm for 30 minutes")&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;The user owns the calendar being written to&lt;/span&gt;
&lt;span class="na"&gt;do_not_use_when&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;The time is approximate or needs negotiation → use scheduling_book_slot&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Inviting external participants who need to pick → use scheduling_book_slot&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;The user said "find a time" rather than naming one&lt;/span&gt;
&lt;span class="na"&gt;reversible&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;yes — event can be deleted via calendar_delete_event&lt;/span&gt;
&lt;span class="na"&gt;side_effects&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sends invitations to all attendees immediately&lt;/span&gt;
&lt;span class="na"&gt;cost&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;low&lt;/span&gt;
&lt;span class="na"&gt;requires&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;google_calendar connection active&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The do_not_use_when block with explicit redirects is doing most of the work. It's not describing this tool — it's describing the boundary between this tool and its neighbours, which is exactly what the model is uncertain about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generating these at scale&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Writing these by hand doesn't scale past a few dozen tools. Our loop:&lt;/p&gt;

&lt;p&gt;Generate a first draft from the schema and whatever docs exist, using a model.&lt;br&gt;
Ship it. It'll be mediocre.&lt;br&gt;
Log every selection, along with what the agent was asked and what it picked.&lt;br&gt;
When a selection is wrong, don't patch the prompt — patch the description of the tool that should have been chosen and the one that was.&lt;br&gt;
Over time the descriptions become a record of every mistake the system has made.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;record_misselection&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chosen_tool&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;correct_tool&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# The fix goes in the tool metadata, not the system prompt.
&lt;/span&gt;    &lt;span class="c1"&gt;# System prompt fixes don't generalize; description fixes do.
&lt;/span&gt;    &lt;span class="n"&gt;boundary_note&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;infer_distinguishing_feature&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;chosen_tool&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;correct_tool&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;correct_tool&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;use_when&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;boundary_note&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;chosen_tool&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;do_not_use_when&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;boundary_note&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; → use &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;correct_tool&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That step 5 is the important one. Every instinct says to fix selection errors by adding rules to the system prompt. Resist it. Prompt fixes are global, they collide with each other, and they don't survive a model upgrade. Description fixes are local to the tools involved and they compose.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A cheap win: prune before you send&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Related, and it saves real money: don't send every tool schema on every step of the loop. Schemas can be a large share of your token spend, because they're re-sent on every iteration, not once per run.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;relevant_tools&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;all_tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Cheap pre-filter — embedding similarity between the current
&lt;/span&gt;    &lt;span class="c1"&gt;# goal and each tool's `purpose` field
&lt;/span&gt;    &lt;span class="n"&gt;scored&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;rank_by_similarity&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;goal&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;all_tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;purpose&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;scored&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nf"&gt;always_include&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;all_tools&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fewer tools in context also means better selection accuracy, so this helps twice. Most frameworks encourage handing the agent everything up front; that's the wrong default past about a dozen tools.&lt;/p&gt;

&lt;p&gt;The bigger point&lt;/p&gt;

&lt;p&gt;The industry is competing on model quality and integration count. Neither is where agent reliability actually comes from right now.&lt;/p&gt;

&lt;p&gt;The semantic layer — what each tool means, when it applies, how it differs from its neighbours — is unglamorous curation work that nobody ships as a feature. It's also, in our experience, the difference between an agent that demos well and one that runs unattended.&lt;/p&gt;

&lt;p&gt;If I were starting again I'd treat it as core product from day one instead of infrastructure to get through.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Building DeskFerry — agents across 1,500+ integrations, which is how we ended up caring about this. Questions welcome below.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>python</category>
    </item>
  </channel>
</rss>
