<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Lars Winstand</title>
    <description>The latest articles on DEV Community by Lars Winstand (@lars_winstand).</description>
    <link>https://dev.to/lars_winstand</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3908932%2Feb8bc1ff-405f-4ef0-8204-ba1ed7caa59f.jpeg</url>
      <title>DEV Community: Lars Winstand</title>
      <link>https://dev.to/lars_winstand</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lars_winstand"/>
    <language>en</language>
    <item>
      <title>I read the 17-comment Reddit fight about trying Kimi K3 and the answer is way less exciting than people want</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Tue, 21 Jul 2026 17:12:22 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-read-the-17-comment-reddit-fight-about-trying-kimi-k3-and-the-answer-is-way-less-exciting-than-bdp</link>
      <guid>https://dev.to/lars_winstand/i-read-the-17-comment-reddit-fight-about-trying-kimi-k3-and-the-answer-is-way-less-exciting-than-bdp</guid>
      <description>&lt;p&gt;The easiest way to try Kimi K3 right now is Moonshot’s own OpenAI-compatible API, not local inference.&lt;/p&gt;

&lt;p&gt;That was the real answer in a 17-comment r/openclaw thread about a deceptively simple question: how do you actually try Kimi K3?&lt;/p&gt;

&lt;p&gt;If you want the short version:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use Moonshot’s API if you want the most direct path&lt;/li&gt;
&lt;li&gt;Use OpenRouter if you want convenience and can tolerate occasional rough edges&lt;/li&gt;
&lt;li&gt;Don’t pretend “runs on a single 80GB A100” means “easy local test”&lt;/li&gt;
&lt;li&gt;If you run agents all day, the bigger issue is not access, it’s still token billing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The thread is here: &lt;a href="https://reddit.com/r/openclaw/comments/1v1vajb/how_do_you_try_kimi_k3/" rel="noopener noreferrer"&gt;https://reddit.com/r/openclaw/comments/1v1vajb/how_do_you_try_kimi_k3/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What made it interesting wasn’t model hype. It was the reason people were asking.&lt;/p&gt;

&lt;p&gt;The original poster wasn’t shopping for novelty. They were looking for a less restrictive option because Claude had started refusing tasks “ever since 4.6+”. That changes the whole framing.&lt;/p&gt;

&lt;p&gt;This is not benchmark tourism.&lt;br&gt;
This is agent operators asking: what still works in production-like loops?&lt;/p&gt;
&lt;h2&gt;
  
  
  The practical answer: use Moonshot’s API
&lt;/h2&gt;

&lt;p&gt;The most useful comment in the thread said the quiet part out loud: Moonshot’s API is the practical way to try K3 without going down a hardware rabbit hole.&lt;/p&gt;

&lt;p&gt;Moonshot exposes an OpenAI-compatible endpoint, which means if your stack already talks to OpenAI-style chat completions, you can usually swap the base URL and model name.&lt;/p&gt;
&lt;h3&gt;
  
  
  Endpoint
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;https://api.moonshot.ai/v1/chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h3&gt;
  
  
  Model
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;kimi-k3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h3&gt;
  
  
  Minimal curl example
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MOONSHOT_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"YOUR_KIMI_API_KEY"&lt;/span&gt;

curl &lt;span class="nt"&gt;--request&lt;/span&gt; POST &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--url&lt;/span&gt; https://api.moonshot.ai/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$MOONSHOT_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--data&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Hello"
      }
    ]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;If you already use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI SDKs&lt;/li&gt;
&lt;li&gt;OpenClaw&lt;/li&gt;
&lt;li&gt;n8n&lt;/li&gt;
&lt;li&gt;Make&lt;/li&gt;
&lt;li&gt;Zapier&lt;/li&gt;
&lt;li&gt;custom agent runners&lt;/li&gt;
&lt;li&gt;any HTTP client wired for chat completions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;...this is boring in the best way.&lt;/p&gt;

&lt;p&gt;And boring is what you want when you’re testing a model inside an existing workflow.&lt;/p&gt;
&lt;h2&gt;
  
  
  Python example with the OpenAI client
&lt;/h2&gt;

&lt;p&gt;If the provider really is OpenAI-compatible, the easiest test is often just changing the base URL.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_KIMI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.moonshot.ai/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kimi-k3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize why developers care about long-context models.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s the whole appeal.&lt;/p&gt;

&lt;p&gt;No weird adapter layer. No custom protocol. No “works if you install this fork from a Discord message.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this thread matters more than the launch posts
&lt;/h2&gt;

&lt;p&gt;A lot of launch coverage treats access as solved the second a model appears somewhere online.&lt;/p&gt;

&lt;p&gt;Developers know that’s fake.&lt;/p&gt;

&lt;p&gt;A model is not really available until you can do all of these without pain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;call it from code&lt;/li&gt;
&lt;li&gt;swap it into an existing agent loop&lt;/li&gt;
&lt;li&gt;handle errors under load&lt;/li&gt;
&lt;li&gt;understand how pricing behaves when usage spikes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s why this thread was better than most announcement posts. People were comparing actual access paths, not vibes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The access options people mentioned
&lt;/h2&gt;

&lt;p&gt;The thread brought up several ways to get at Kimi K3 or Kimi-adjacent deployments:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Moonshot direct&lt;/li&gt;
&lt;li&gt;OpenRouter&lt;/li&gt;
&lt;li&gt;Cloudflare Workers AI&lt;/li&gt;
&lt;li&gt;OpenCode Go&lt;/li&gt;
&lt;li&gt;Kimi consumer membership&lt;/li&gt;
&lt;li&gt;local/self-hosted distilled variants&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That sounds like plenty of choice.&lt;/p&gt;

&lt;p&gt;In practice, it’s fragmentation.&lt;/p&gt;

&lt;p&gt;Each option solves a different problem.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What you’re really getting&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Moonshot API&lt;/td&gt;
&lt;td&gt;Official provider, OpenAI-compatible access, token-billed usage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenRouter&lt;/td&gt;
&lt;td&gt;Fast aggregator access, easy testing, but users reported occasional 429s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloudflare Workers AI&lt;/td&gt;
&lt;td&gt;Infra-adjacent path if you already live in Cloudflare’s world&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenCode Go&lt;/td&gt;
&lt;td&gt;Provider abstraction for coding workflows, less provider babysitting&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local/distilled variant&lt;/td&gt;
&lt;td&gt;More control and privacy, much higher hardware and setup cost&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The most honest summary in the thread might have been: “Open Router. Occasional 429 though.”&lt;/p&gt;

&lt;p&gt;That’s exactly how aggregator access usually feels.&lt;/p&gt;

&lt;p&gt;Great until load shows up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can you run Kimi K3 locally?
&lt;/h2&gt;

&lt;p&gt;Sort of.&lt;/p&gt;

&lt;p&gt;This is where Reddit model threads usually get slippery.&lt;/p&gt;

&lt;p&gt;Yes, people mentioned a 32B distilled version that can run on a single 80GB A100.&lt;/p&gt;

&lt;p&gt;No, that does not mean local Kimi K3 is a casual weekend test for most developers.&lt;/p&gt;

&lt;p&gt;A single 80GB A100 is not normal desktop hardware.&lt;br&gt;
It is not “I had an extra GPU lying around.”&lt;br&gt;
It is not the same thing as “just run it locally.”&lt;/p&gt;

&lt;p&gt;So when someone says “you can run Kimi locally,” they usually mean one of three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You can run a smaller or distilled variant&lt;/li&gt;
&lt;li&gt;You already have access to serious hardware&lt;/li&gt;
&lt;li&gt;You’re willing to spend real time on deployment instead of just evaluating the model&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Those are very different claims.&lt;/p&gt;

&lt;p&gt;If your actual goal is: should I try this in OpenClaw or an agent loop?&lt;br&gt;
Then local is usually not the first move.&lt;/p&gt;

&lt;p&gt;Hosted API access is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example: swapping providers in an agent workflow
&lt;/h2&gt;

&lt;p&gt;This is the real developer use case.&lt;/p&gt;

&lt;p&gt;You already have an agent setup. You don’t want to rebuild it just to test one model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Generic config pattern
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"moonshot"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"base_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.moonshot.ai/v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"api_key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"YOUR_KIMI_API_KEY"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"kimi-k3"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Pseudocode for a provider-swappable chat call
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fetch&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;node-fetch&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;baseUrl&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;baseUrl&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/chat/completions`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Authorization&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;text&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`HTTP &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;baseUrl&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://api.moonshot.ai/v1&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;MOONSHOT_API_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;kimi-k3&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Write a regex that extracts order IDs from log lines.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;})();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is why OpenAI-compatible APIs keep winning. Not because they’re exciting. Because they let developers test providers with minimal surgery.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part people keep glossing over: token billing
&lt;/h2&gt;

&lt;p&gt;This is where the Reddit thread was useful but incomplete.&lt;/p&gt;

&lt;p&gt;Yes, Moonshot direct is the practical path.&lt;br&gt;
Yes, OpenRouter is convenient.&lt;br&gt;
Yes, local is mostly oversold for casual testing.&lt;/p&gt;

&lt;p&gt;But the bigger issue for teams running agents is cost behavior.&lt;/p&gt;

&lt;p&gt;Kimi API usage is still token-billed.&lt;br&gt;
That means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;input tokens cost money&lt;/li&gt;
&lt;li&gt;output tokens cost money&lt;/li&gt;
&lt;li&gt;retries cost money&lt;/li&gt;
&lt;li&gt;long-context prompts cost money&lt;/li&gt;
&lt;li&gt;always-on agent loops definitely cost money&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So if your question is:&lt;/p&gt;

&lt;p&gt;“How do I try Kimi K3?”&lt;/p&gt;

&lt;p&gt;The answer is easy.&lt;/p&gt;

&lt;p&gt;If your question is:&lt;/p&gt;

&lt;p&gt;“How do I run Kimi-style workloads for agents all day without watching token spend like a hawk?”&lt;/p&gt;

&lt;p&gt;That’s a different problem.&lt;/p&gt;

&lt;p&gt;And it’s the one most teams run into after the first successful demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  My take
&lt;/h2&gt;

&lt;p&gt;If you want to evaluate Kimi K3 for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenClaw&lt;/li&gt;
&lt;li&gt;coding agents&lt;/li&gt;
&lt;li&gt;long-context workflows&lt;/li&gt;
&lt;li&gt;provider comparisons&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Start with Moonshot’s official API.&lt;/p&gt;

&lt;p&gt;It’s the least confusing path.&lt;br&gt;
It fits existing OpenAI-compatible tooling.&lt;br&gt;
It gets you to a real answer quickly.&lt;/p&gt;

&lt;p&gt;Use OpenRouter if speed and convenience matter more than consistency.&lt;br&gt;
Just expect occasional provider-layer weirdness, including the kind of 429s people mentioned in the thread.&lt;/p&gt;

&lt;p&gt;Use local or distilled variants only if you already care about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;privacy&lt;/li&gt;
&lt;li&gt;infrastructure control&lt;/li&gt;
&lt;li&gt;hardware experimentation&lt;/li&gt;
&lt;li&gt;self-hosting for strategic reasons&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Don’t use local because Reddit made it sound easy.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for teams running agents
&lt;/h2&gt;

&lt;p&gt;This thread started as a model question.&lt;br&gt;
It turned into an infrastructure question.&lt;br&gt;
That’s why it was worth reading.&lt;/p&gt;

&lt;p&gt;For developers running real automations, the hard part is rarely “can I hit the endpoint?”&lt;/p&gt;

&lt;p&gt;The hard part is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;can I swap providers without rewriting everything?&lt;/li&gt;
&lt;li&gt;will this stay available under load?&lt;/li&gt;
&lt;li&gt;what happens when the model starts refusing tasks?&lt;/li&gt;
&lt;li&gt;what happens to cost when the agent runs 24/7?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one is where a lot of teams eventually rethink the whole pricing model.&lt;/p&gt;

&lt;p&gt;If you’re tired of per-token billing and constant usage math, that’s exactly the problem Standard Compute is built for: unlimited AI compute at a flat monthly price, using an OpenAI-compatible API, so agent workflows can run without token anxiety.&lt;/p&gt;

&lt;p&gt;That’s the bigger story behind this little Kimi K3 thread.&lt;/p&gt;

&lt;p&gt;Trying a model is easy.&lt;br&gt;
Running agents predictably is the real problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Actionable takeaway
&lt;/h2&gt;

&lt;p&gt;If you want to test Kimi K3 today:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Get a Moonshot API key&lt;/li&gt;
&lt;li&gt;Point your OpenAI-compatible client at &lt;code&gt;https://api.moonshot.ai/v1&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;kimi-k3&lt;/code&gt; as the model name&lt;/li&gt;
&lt;li&gt;Run a small real-world prompt from your actual workflow&lt;/li&gt;
&lt;li&gt;Measure quality, latency, refusal behavior, and cost&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you’re running agents continuously, add one more step:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Decide whether token billing is acceptable before you wire it into production loops&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That’s the answer the Reddit thread circled around.&lt;/p&gt;

&lt;p&gt;Not glamorous, but useful.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>openai</category>
      <category>devops</category>
    </item>
    <item>
      <title>I learned the hard way that Slack is the worst place to find out your website agent skipped the important part</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Tue, 21 Jul 2026 09:11:50 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-learned-the-hard-way-that-slack-is-the-worst-place-to-find-out-your-website-agent-skipped-the-167i</link>
      <guid>https://dev.to/lars_winstand/i-learned-the-hard-way-that-slack-is-the-worst-place-to-find-out-your-website-agent-skipped-the-167i</guid>
      <description>&lt;p&gt;Slack is great for approvals and alerts.&lt;/p&gt;

&lt;p&gt;It is a terrible source of truth for supervising long-running agents.&lt;/p&gt;

&lt;p&gt;That sounds dramatic until you run a real website update workflow through it.&lt;/p&gt;

&lt;p&gt;I’m talking about the common setup now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;an agent updates site copy&lt;/li&gt;
&lt;li&gt;touches a CMS&lt;/li&gt;
&lt;li&gt;maybe runs a script&lt;/li&gt;
&lt;li&gt;maybe calls an internal API&lt;/li&gt;
&lt;li&gt;posts progress into Slack so a human can approve the final step&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On paper, this looks clean.&lt;/p&gt;

&lt;p&gt;In practice, Slack turns rich execution traces into vibes.&lt;/p&gt;

&lt;p&gt;And vibes are not enough when an agent is editing production content.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure mode is extremely normal
&lt;/h2&gt;

&lt;p&gt;I ran into this while looking at agent supervision patterns for website update automation.&lt;/p&gt;

&lt;p&gt;A thread on r/openclaw captured the exact problem:&lt;br&gt;
&lt;a href="https://reddit.com/r/openclaw/comments/1v1nmnk/how_to_get_agent_commentary_and_tool_calls_to/" rel="noopener noreferrer"&gt;https://reddit.com/r/openclaw/comments/1v1nmnk/how_to_get_agent_commentary_and_tool_calls_to/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The user had already enabled Slack streaming with commentary, narration, and tool progress:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"channels"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"slack"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"streaming"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"progress"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"progress"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"commentary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"render"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"rich"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"narration"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"toolProgress"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is not a lazy setup.&lt;/p&gt;

&lt;p&gt;That is someone trying to do agent supervision correctly.&lt;/p&gt;

&lt;p&gt;And after a bunch of testing, they still ended up with Slack showing header-like fragments instead of the details they actually needed.&lt;/p&gt;

&lt;p&gt;The key line was basically:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I can see the commentary and tool calls in the OpenClaw dashboard, but I want them in Slack too.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s the whole issue.&lt;/p&gt;

&lt;p&gt;The dashboard has the truth.&lt;br&gt;
Slack has the summary.&lt;/p&gt;

&lt;p&gt;If your human supervisor only watches Slack, they are watching a compressed version of reality.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why this happens
&lt;/h2&gt;

&lt;p&gt;Because Slack is a chat app.&lt;/p&gt;

&lt;p&gt;Agents are not chat-shaped anymore.&lt;/p&gt;

&lt;p&gt;Modern agent workflows emit structured events:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;planning steps&lt;/li&gt;
&lt;li&gt;tool calls&lt;/li&gt;
&lt;li&gt;partial tool output&lt;/li&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;approval checkpoints&lt;/li&gt;
&lt;li&gt;state transitions&lt;/li&gt;
&lt;li&gt;final actions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpenAI Responses API streams event-like output around tool use and response items.&lt;br&gt;
Anthropic Messages API streaming exposes granular blocks like &lt;code&gt;tool_use&lt;/code&gt; and content deltas.&lt;/p&gt;

&lt;p&gt;That is how agents behave in the real world.&lt;/p&gt;

&lt;p&gt;A trace looks more like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;inspect page state&lt;/li&gt;
&lt;li&gt;compare requested copy changes&lt;/li&gt;
&lt;li&gt;call CMS tool&lt;/li&gt;
&lt;li&gt;get validation error&lt;/li&gt;
&lt;li&gt;retry with corrected field&lt;/li&gt;
&lt;li&gt;generate diff summary&lt;/li&gt;
&lt;li&gt;request approval&lt;/li&gt;
&lt;li&gt;publish&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A dashboard can preserve that structure.&lt;/p&gt;

&lt;p&gt;A Slack thread flattens it into text.&lt;/p&gt;

&lt;p&gt;Once you flatten it, you lose the exact context a human needs to decide whether the agent is being careful or just sounding confident.&lt;/p&gt;
&lt;h2&gt;
  
  
  The real problem: lying by omission
&lt;/h2&gt;

&lt;p&gt;This is the part that matters.&lt;/p&gt;

&lt;p&gt;If Slack says:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;Commentary&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;bash&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;tool running&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;that is not transparency.&lt;br&gt;
That is a label.&lt;/p&gt;

&lt;p&gt;It does not answer the useful questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What command ran?&lt;/li&gt;
&lt;li&gt;What file changed?&lt;/li&gt;
&lt;li&gt;What API payload was sent?&lt;/li&gt;
&lt;li&gt;Did the first attempt fail?&lt;/li&gt;
&lt;li&gt;Did the agent retry?&lt;/li&gt;
&lt;li&gt;Did it touch staging or production?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are very different stories.&lt;/p&gt;

&lt;p&gt;But Slack can make them all look identical.&lt;/p&gt;

&lt;p&gt;For website automation, that gets dangerous fast.&lt;/p&gt;

&lt;p&gt;There’s a huge difference between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;resizing an image&lt;/li&gt;
&lt;li&gt;updating a blog title&lt;/li&gt;
&lt;li&gt;replacing pricing copy&lt;/li&gt;
&lt;li&gt;changing SEO metadata&lt;/li&gt;
&lt;li&gt;triggering publish in Webflow, WordPress, or Contentful&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If all Slack shows is &lt;code&gt;tool running&lt;/code&gt;, your supervision layer is missing the important part.&lt;/p&gt;
&lt;h2&gt;
  
  
  Slack’s limits are a bad fit for agent traces
&lt;/h2&gt;

&lt;p&gt;This is not just a UX complaint.&lt;/p&gt;

&lt;p&gt;Slack’s API constraints are fine for chatbots and notifications.&lt;br&gt;
They are not fine for high-fidelity agent streaming.&lt;/p&gt;

&lt;p&gt;Slack recommends keeping message text under 4,000 characters and says messages over 40,000 characters may be truncated.&lt;/p&gt;

&lt;p&gt;Slack also rate-limits message posting to roughly 1 message per second per channel, plus broader workspace limits.&lt;/p&gt;

&lt;p&gt;For a normal bot, no problem.&lt;/p&gt;

&lt;p&gt;For an agent emitting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;commentary&lt;/li&gt;
&lt;li&gt;tool progress&lt;/li&gt;
&lt;li&gt;command output&lt;/li&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;approval checkpoints&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;that becomes a bottleneck.&lt;/p&gt;
&lt;h3&gt;
  
  
  What breaks first
&lt;/h3&gt;

&lt;p&gt;Usually not the agent.&lt;/p&gt;

&lt;p&gt;The supervision layer breaks first.&lt;/p&gt;

&lt;p&gt;When you push too much execution detail into Slack, one of these happens:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;updates get batched into vague summaries&lt;/li&gt;
&lt;li&gt;updates arrive late or out of order&lt;/li&gt;
&lt;li&gt;updates truncate&lt;/li&gt;
&lt;li&gt;updates disappear under load&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And then the human says, “the agent was weird.”&lt;/p&gt;

&lt;p&gt;Maybe.&lt;/p&gt;

&lt;p&gt;But a lot of the time, the trace was fine and Slack turned it into mush.&lt;/p&gt;
&lt;h2&gt;
  
  
  Bad pattern: treating Slack like an observability system
&lt;/h2&gt;

&lt;p&gt;I keep seeing teams do this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent -&amp;gt; tool call -&amp;gt; tool output -&amp;gt; progress update -&amp;gt; retry -&amp;gt; approval request -&amp;gt; publish -&amp;gt; final summary
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All streamed directly into one Slack thread.&lt;/p&gt;

&lt;p&gt;This feels convenient because everyone already lives in Slack.&lt;/p&gt;

&lt;p&gt;It is also the fastest way to create false confidence.&lt;/p&gt;

&lt;p&gt;A thread full of updates looks like visibility.&lt;br&gt;
It is not the same thing as execution visibility.&lt;/p&gt;
&lt;h2&gt;
  
  
  Better pattern: Slack for attention, trace UI for truth
&lt;/h2&gt;

&lt;p&gt;This is the pattern I’d use every time for website update automation.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What happens in real life&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Slack thread only&lt;/td&gt;
&lt;td&gt;Fast for human attention, bad for dense traces, raw tool output, and debugging; truncation and rate limits show up quickly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dashboard / trace UI only&lt;/td&gt;
&lt;td&gt;Best for full fidelity, spans, retries, tool input/output, and replay; worse for quick approvals because humans are not staring at it all day&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hybrid: Slack + dashboard&lt;/td&gt;
&lt;td&gt;Best tradeoff; Slack handles summaries and approvals, dashboard holds the canonical trace&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That hybrid setup is the one that survives contact with production.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I’d actually build
&lt;/h2&gt;
&lt;h3&gt;
  
  
  1. Keep Slack short and decision-oriented
&lt;/h3&gt;

&lt;p&gt;Slack messages should answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What is happening?&lt;/li&gt;
&lt;li&gt;Does a human need to act?&lt;/li&gt;
&lt;li&gt;Where is the full trace?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Good Slack milestones:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;planned change&lt;/li&gt;
&lt;li&gt;running tool step&lt;/li&gt;
&lt;li&gt;awaiting approval&lt;/li&gt;
&lt;li&gt;completed&lt;/li&gt;
&lt;li&gt;failed and needs review&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Bad Slack content:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;full command output&lt;/li&gt;
&lt;li&gt;raw diffs&lt;/li&gt;
&lt;li&gt;long reasoning streams&lt;/li&gt;
&lt;li&gt;every retry event&lt;/li&gt;
&lt;li&gt;every intermediate tool payload&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example Slack message:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Website agent updated homepage hero copy in staging. Awaiting approval before publish. Full trace: https://your-trace-ui/runs/abc123"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is enough.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Put the real execution record in a trace UI
&lt;/h3&gt;

&lt;p&gt;Every Slack checkpoint should link to the trace.&lt;/p&gt;

&lt;p&gt;That trace should include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;raw tool input/output&lt;/li&gt;
&lt;li&gt;command text&lt;/li&gt;
&lt;li&gt;file diffs&lt;/li&gt;
&lt;li&gt;timestamps&lt;/li&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;approval events&lt;/li&gt;
&lt;li&gt;model used for each step&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are using OpenClaw, LangSmith, or your own tracing layer, this is where the real debugging happens.&lt;/p&gt;

&lt;p&gt;One reply in that OpenClaw thread mentioned trying this too:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"commandText"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"raw"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That may improve visibility in some setups.&lt;/p&gt;

&lt;p&gt;Still, I would not make Slack the canonical log even if raw command text shows up.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Approve in Slack, investigate in the trace
&lt;/h3&gt;

&lt;p&gt;This is the clean split.&lt;/p&gt;

&lt;p&gt;Slack is good for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“Approve this publish?”&lt;/li&gt;
&lt;li&gt;“This step failed.”&lt;/li&gt;
&lt;li&gt;“The agent is waiting on you.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trace UI is good for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“Why did it touch this field?”&lt;/li&gt;
&lt;li&gt;“What command actually ran?”&lt;/li&gt;
&lt;li&gt;“Did the first attempt fail?”&lt;/li&gt;
&lt;li&gt;“Which model decided to call the tool?”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Different interface, different job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Concrete implementation sketch
&lt;/h2&gt;

&lt;p&gt;Here’s a simple pattern for an agent that updates a site and posts into Slack.&lt;/p&gt;

&lt;h3&gt;
  
  
  Agent flow
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;plan -&amp;gt; inspect -&amp;gt; edit draft -&amp;gt; validate -&amp;gt; summarize diff -&amp;gt; request approval -&amp;gt; publish
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  What gets stored in the trace
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"run_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"site-update-4821"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gpt-5.4"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"steps"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"inspect"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"homepage hero"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"result"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Current headline fetched"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"tool_call"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"tool"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"contentful.updateEntry"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"input"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"entryId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"hero_01"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"field"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"headline"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"draft updated"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"approval_required"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"summary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Headline changed from A to B"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  What gets sent to Slack
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Website agent prepared a homepage update.

Change: hero headline updated
Status: awaiting approval
Trace: https://trace.example.com/runs/site-update-4821
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That separation keeps Slack readable and the trace useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters even more as agents get better
&lt;/h2&gt;

&lt;p&gt;Here’s the weird part: richer supervision usually increases traffic.&lt;/p&gt;

&lt;p&gt;Better agent operations do not always mean fewer messages.&lt;br&gt;
They often mean:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;more checkpoints&lt;/li&gt;
&lt;li&gt;more retries surfaced&lt;/li&gt;
&lt;li&gt;more tool events&lt;/li&gt;
&lt;li&gt;more review loops&lt;/li&gt;
&lt;li&gt;more experiments with prompts and routing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That creates pressure on both:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the human-facing channel&lt;/li&gt;
&lt;li&gt;the underlying model/tool budget&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is where pricing starts affecting architecture.&lt;/p&gt;

&lt;p&gt;If every extra trace, retry, and review loop feels expensive, teams suppress visibility.&lt;br&gt;
They log less.&lt;br&gt;
They supervise less.&lt;br&gt;
They avoid useful checkpoints because each one costs money.&lt;/p&gt;

&lt;p&gt;That is a bad incentive if you are running agents in n8n, Make, Zapier, OpenClaw, or custom workflows.&lt;/p&gt;

&lt;p&gt;For agent-heavy automations, flat monthly compute is a lot easier to reason about than per-token billing.&lt;/p&gt;

&lt;p&gt;If your team wants more supervision, more traces, and more iteration without constantly watching usage, that’s exactly why products like Standard Compute exist:&lt;br&gt;
&lt;a href="https://standardcompute.com" rel="noopener noreferrer"&gt;https://standardcompute.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It gives you an OpenAI-compatible API with flat monthly pricing, so you can run agent workflows without every extra checkpoint feeling like a billing event.&lt;/p&gt;

&lt;p&gt;That matters more than people think.&lt;/p&gt;

&lt;p&gt;Because the better your oversight gets, the more model activity you usually generate.&lt;/p&gt;
&lt;h2&gt;
  
  
  My rule now
&lt;/h2&gt;

&lt;p&gt;If a human might need to ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;what exactly did the agent do?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Slack cannot be the only place that answer lives.&lt;/p&gt;

&lt;p&gt;Use Slack as the front desk.&lt;br&gt;
Use a trace UI as the security camera footage.&lt;/p&gt;

&lt;p&gt;That’s the clean lesson here.&lt;/p&gt;

&lt;p&gt;Slack is great for attention, approvals, and escalation.&lt;/p&gt;

&lt;p&gt;It is not where I want the only copy of a production agent’s execution history.&lt;/p&gt;

&lt;p&gt;Once you see that clearly, the architecture gets simpler:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Slack for summaries&lt;/li&gt;
&lt;li&gt;dashboard for evidence&lt;/li&gt;
&lt;li&gt;humans approve in chat&lt;/li&gt;
&lt;li&gt;humans debug in traces&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That setup is less flashy than streaming everything into a thread.&lt;/p&gt;

&lt;p&gt;It is also the one I trust.&lt;/p&gt;
&lt;h2&gt;
  
  
  Practical takeaway
&lt;/h2&gt;

&lt;p&gt;If you’re building website update automation this week, do this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;1. Stream milestones to Slack
2. Store full traces elsewhere
3. Link every Slack update to the trace
4. Keep approvals &lt;span class="k"&gt;in &lt;/span&gt;Slack
5. Keep debugging out of Slack
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you skip step 2, you are not supervising an agent.&lt;/p&gt;

&lt;p&gt;You are reading its status messages and hoping they tell the full story.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>automation</category>
      <category>api</category>
    </item>
    <item>
      <title>I thought adding 3 more OpenClaw agents would help but the real problem was AI agent handoff state</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Mon, 20 Jul 2026 17:13:16 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-thought-adding-3-more-openclaw-agents-would-help-but-the-real-problem-was-ai-agent-handoff-state-klj</link>
      <guid>https://dev.to/lars_winstand/i-thought-adding-3-more-openclaw-agents-would-help-but-the-real-problem-was-ai-agent-handoff-state-klj</guid>
      <description>&lt;p&gt;AI agent handoff state is the thing that separates a fun OpenClaw demo from a multi-agent system you can trust.&lt;/p&gt;

&lt;p&gt;One agent can get away with thread memory.&lt;/p&gt;

&lt;p&gt;Three agents usually can’t.&lt;/p&gt;

&lt;p&gt;Once agents start sharing work, you need structured shared memory for facts and artifacts. Not a 100k-token transcript that gets slower, more expensive, and less reliable every week.&lt;/p&gt;

&lt;p&gt;I keep seeing the same failure mode in OpenClaw setups.&lt;/p&gt;

&lt;p&gt;The first agent feels incredible.&lt;/p&gt;

&lt;p&gt;It writes specs. It drafts outbound. It triages bugs. It kicks hard tasks to Claude or GPT-5. You feel like you found a cheat code.&lt;/p&gt;

&lt;p&gt;Then you add a second agent.&lt;/p&gt;

&lt;p&gt;Then a reviewer.&lt;/p&gt;

&lt;p&gt;Then a researcher.&lt;/p&gt;

&lt;p&gt;Then something that posts updates to Discord or Slack.&lt;/p&gt;

&lt;p&gt;And now the system isn’t exactly broken. It’s worse than broken.&lt;/p&gt;

&lt;p&gt;It’s flaky.&lt;/p&gt;

&lt;p&gt;One agent discovers something important and the next one behaves like it never happened. Or it remembers the wrong detail. Or you "fix" that by shoving the entire transcript into every prompt, and now latency spikes, token usage gets ugly, and your orchestration starts feeling like a hostage negotiation with stale context.&lt;/p&gt;

&lt;p&gt;That’s when OpenClaw stops being a prompt problem and becomes a systems problem.&lt;/p&gt;

&lt;p&gt;While looking into this, I found a thread on r/openclaw where someone asked the right question: how are people sharing knowledge between multiple OpenClaw agents?&lt;/p&gt;

&lt;p&gt;That is the bottleneck.&lt;/p&gt;

&lt;p&gt;Not model quality.&lt;/p&gt;

&lt;p&gt;Not prompt cleverness.&lt;/p&gt;

&lt;p&gt;Handoff state.&lt;/p&gt;

&lt;h2&gt;
  
  
  The common mistake: treating memory like one giant transcript
&lt;/h2&gt;

&lt;p&gt;A lot of teams start here:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Agent A does 20 turns of work&lt;/li&gt;
&lt;li&gt;Pass all 20 turns to Agent B&lt;/li&gt;
&lt;li&gt;Agent B adds 15 more turns&lt;/li&gt;
&lt;li&gt;Pass all 35 turns to Agent C&lt;/li&gt;
&lt;li&gt;Wonder why everything got slower, pricier, and dumber&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That architecture is a junk drawer with a context window.&lt;/p&gt;

&lt;p&gt;LangGraph’s docs make a useful distinction here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Checkpointers handle short-term, thread-level state&lt;/li&gt;
&lt;li&gt;Stores handle long-term shared data across threads&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That maps cleanly to OpenClaw.&lt;/p&gt;

&lt;p&gt;The real split is not one-agent vs multi-agent.&lt;/p&gt;

&lt;p&gt;It’s this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;thread-scoped memory&lt;/li&gt;
&lt;li&gt;cross-thread shared memory&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your research agent found a competitor pricing page, your coding agent probably does not need the whole chat that led to it.&lt;/p&gt;

&lt;p&gt;It needs something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"fact"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Competitor X charges $99/month for 10k runs"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com/pricing"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"captured_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-07-20T10:22:00Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.93&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"verified"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a handoff.&lt;/p&gt;

&lt;p&gt;A transcript is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Yes, Claude can handle huge prompts. That’s not the point.
&lt;/h2&gt;

&lt;p&gt;This is where people get tripped up.&lt;/p&gt;

&lt;p&gt;Anthropic has shown Claude handling very large prompts. Their older 100k context announcement framed 100,000 tokens as roughly 75,000 words. They’ve also shown long-context retrieval examples like scanning a 72k-token copy of The Great Gatsby.&lt;/p&gt;

&lt;p&gt;So yes, giant prompts are real.&lt;/p&gt;

&lt;p&gt;And for some workloads, they’re fine.&lt;/p&gt;

&lt;p&gt;If you have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one agent&lt;/li&gt;
&lt;li&gt;one bounded corpus&lt;/li&gt;
&lt;li&gt;one workflow that resets cleanly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;then prompt stuffing can be the simplest correct answer.&lt;/p&gt;

&lt;p&gt;Anthropic also points out that prompt caching can reduce latency and cost substantially for repeated long prompts.&lt;/p&gt;

&lt;p&gt;That’s all true.&lt;/p&gt;

&lt;p&gt;But it falls apart as an architecture once you have ongoing multi-agent work.&lt;/p&gt;

&lt;p&gt;Because then your transcript becomes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;larger every run&lt;/li&gt;
&lt;li&gt;less relevant every run&lt;/li&gt;
&lt;li&gt;harder to trust every run&lt;/li&gt;
&lt;li&gt;more expensive every run&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;LangGraph explicitly warns about long histories increasing latency and cost, exceeding context windows, and degrading model performance because the model gets distracted by stale or irrelevant content.&lt;/p&gt;

&lt;p&gt;That matches what people see in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Reddit posts stop sounding like demos very quickly
&lt;/h2&gt;

&lt;p&gt;What got my attention is that OpenClaw users are already talking in real spend, not toy-project numbers.&lt;/p&gt;

&lt;p&gt;In one r/openclaw thread, a user said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;About 3k a month. GTM, product de (mostly specs, coding done by Claude/Codex/GHCP), Sales... Multi model - depending on task. Not fully convinced it is worth the cash. But it does make a lot of tasks easier and faster, which is invaluable.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Another said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;we've used it to basically build and run a real estate brokerage, real estate investment business, and openclaw-for-realtors SaaS product called "Homies AI" burn rate on it running the business is like $30k/yr the businesses make like $500k/yr&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And in another thread:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Since I installed OpenClaw 4 months ago I have spent over $10k on tokens via OpenRouter.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s the part people miss.&lt;/p&gt;

&lt;p&gt;Once agents are attached to real work, memory design becomes a cost decision.&lt;/p&gt;

&lt;p&gt;Bad handoffs are not just messy.&lt;/p&gt;

&lt;p&gt;They’re expensive.&lt;/p&gt;

&lt;p&gt;That’s exactly why pricing starts to matter too. If your agents are constantly summarizing, re-reading, re-handoffing, and carrying giant prompts around, per-token billing punishes every architectural mistake. Teams running multi-agent automations on OpenClaw, n8n, Make, Zapier, or custom workflows feel this fast.&lt;/p&gt;

&lt;p&gt;That’s also why flat-rate API access is interesting. If you’re iterating on agent architecture and doing lots of retries, summarization, routing, and tool calls, predictable cost matters more than people admit. Standard Compute is basically built for this kind of workload: OpenAI-compatible API, flat monthly pricing, and no per-token panic while you tune agent systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should one agent actually pass to another?
&lt;/h2&gt;

&lt;p&gt;OpenAI’s Agents SDK has a clean mental model here.&lt;/p&gt;

&lt;p&gt;It separates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;handoffs: who gets control next&lt;/li&gt;
&lt;li&gt;sessions: conversation history for a run or thread&lt;/li&gt;
&lt;li&gt;application state: local state that should not be sent to the model&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That distinction is gold.&lt;/p&gt;

&lt;p&gt;Because not all state belongs in the prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  The 3 buckets you should keep separate
&lt;/h3&gt;

&lt;h4&gt;
  
  
  1. Model-visible conversational state
&lt;/h4&gt;

&lt;p&gt;What the next agent genuinely needs to read.&lt;/p&gt;

&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;current task&lt;/li&gt;
&lt;li&gt;recent decisions&lt;/li&gt;
&lt;li&gt;a compact summary&lt;/li&gt;
&lt;li&gt;explicit constraints&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  2. Trusted application state
&lt;/h4&gt;

&lt;p&gt;Stuff your app needs, but the model does not need verbatim.&lt;/p&gt;

&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;auth tokens&lt;/li&gt;
&lt;li&gt;internal IDs&lt;/li&gt;
&lt;li&gt;workflow flags&lt;/li&gt;
&lt;li&gt;rate limit counters&lt;/li&gt;
&lt;li&gt;customer account metadata&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  3. Shared durable knowledge
&lt;/h4&gt;

&lt;p&gt;Things other agents may need later.&lt;/p&gt;

&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;extracted facts&lt;/li&gt;
&lt;li&gt;approved decisions&lt;/li&gt;
&lt;li&gt;source links&lt;/li&gt;
&lt;li&gt;generated artifacts&lt;/li&gt;
&lt;li&gt;verified outputs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you collapse all three into one giant prompt blob, you get the worst combination possible:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;more tokens&lt;/li&gt;
&lt;li&gt;less control&lt;/li&gt;
&lt;li&gt;less trust&lt;/li&gt;
&lt;li&gt;worse debugging&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A minimal handoff pattern
&lt;/h2&gt;

&lt;p&gt;Here’s a tiny Python example showing the shape of a better handoff.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataclasses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;dataclass&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;

&lt;span class="nd"&gt;@dataclass&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;HandoffArtifact&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fact&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;decision&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;artifact&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;source_url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;draft&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verified&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;approved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stale&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;created_by&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;

&lt;span class="n"&gt;pricing_fact&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;HandoffArtifact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;kind&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fact&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Competitor pricing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Competitor X charges $99/month for 10k runs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;source_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://example.com/pricing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;confidence&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.93&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verified&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;created_by&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;research-agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;created_at&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2026-07-20T10:22:00Z&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the next agent gets the useful output, not the entire cognitive mess that produced it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical OpenClaw memory layout
&lt;/h2&gt;

&lt;p&gt;If I were wiring this up today, I’d use something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;openclaw-memory/
├── threads/
│   ├── thread_123.json
│   └── thread_124.json
├── artifacts/
│   ├── facts.jsonl
│   ├── decisions.jsonl
│   └── outputs.jsonl
└── indexes/
    └── embeddings.sqlite
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the flow would look like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Thread memory keeps the current run coherent&lt;/li&gt;
&lt;li&gt;Each agent emits structured artifacts&lt;/li&gt;
&lt;li&gt;Artifacts are stored centrally&lt;/li&gt;
&lt;li&gt;Retrieval selects only relevant artifacts for the next agent&lt;/li&gt;
&lt;li&gt;Old or low-confidence artifacts get pruned&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That last step matters a lot.&lt;/p&gt;

&lt;p&gt;A shared store can turn into a second junk drawer if you let every half-baked thought into it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example: don’t pass chats, pass artifacts
&lt;/h2&gt;

&lt;p&gt;Bad handoff:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Here are the last 14 messages from the research agent, plus 8 web snippets,
plus 3 abandoned ideas, plus a summary of a summary. Use all of that to write the spec."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Better handoff:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"task"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Write implementation spec for competitor-monitoring job"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"constraints"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Use Postgres"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Must support retryable fetches"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Run daily at 09:00 UTC"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"artifacts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"fact"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Competitor X pricing"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"$99/month for 10k runs"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"source_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com/pricing"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.93&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"verified"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"decision"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Storage backend"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Use Postgres instead of Redis for auditability"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"source_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"approved"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s shorter, clearer, cheaper, and easier to debug.&lt;/p&gt;

&lt;h2&gt;
  
  
  You can prototype this locally in an afternoon
&lt;/h2&gt;

&lt;p&gt;A very simple setup using Python, SQLite, and JSON is enough to prove the pattern.&lt;/p&gt;

&lt;p&gt;Install basics:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
&lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate
pip &lt;span class="nb"&gt;install &lt;/span&gt;sqlite-utils pydantic
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Create a tiny artifact store:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sqlite3&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sqlite3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;memory.db&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
CREATE TABLE IF NOT EXISTS artifacts (
  id INTEGER PRIMARY KEY AUTOINCREMENT,
  kind TEXT NOT NULL,
  title TEXT NOT NULL,
  content TEXT NOT NULL,
  source_url TEXT,
  confidence REAL,
  status TEXT,
  created_by TEXT,
  created_at TEXT
)
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;artifact&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kind&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fact&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Competitor X pricing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;$99/month for 10k runs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://example.com/pricing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.93&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verified&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_by&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;research-agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2026-07-20T10:22:00Z&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    INSERT INTO artifacts
    (kind, title, content, source_url, confidence, status, created_by, created_at)
    VALUES (?, ?, ?, ?, ?, ?, ?, ?)
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kind&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_by&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;commit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Retrieve only what the next agent needs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sqlite3&lt;/span&gt;

&lt;span class="n"&gt;conn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sqlite3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;memory.db&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    SELECT kind, title, content, source_url, confidence, status
    FROM artifacts
    WHERE status IN (&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;verified&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;approved&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;)
      AND confidence &amp;gt;= 0.8
    ORDER BY created_at DESC
    LIMIT 5
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fetchall&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That alone is already better than forwarding raw transcripts forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which memory pattern actually wins?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;What it’s good for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Giant shared prompt&lt;/td&gt;
&lt;td&gt;Fastest way to prototype. Fine for one agent and a small bounded corpus. Gets worse as history grows.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thread/session memory&lt;/td&gt;
&lt;td&gt;Good for continuity inside one run or one conversation. Not enough for cross-agent knowledge sharing.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shared store + retrieval&lt;/td&gt;
&lt;td&gt;Best pattern for multi-agent systems. Better for facts, artifacts, and trusted handoffs. Requires schema and pruning discipline.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;My opinion is simple.&lt;/p&gt;

&lt;p&gt;If you’re testing one agent on a narrow task, use the giant prompt and move on.&lt;/p&gt;

&lt;p&gt;If you’re building an OpenClaw workflow with multiple specialists, a shared transcript is the wrong architecture.&lt;/p&gt;

&lt;p&gt;Use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;thread memory for continuity&lt;/li&gt;
&lt;li&gt;a shared store for durable knowledge&lt;/li&gt;
&lt;li&gt;selective retrieval for handoffs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s the line between a demo and a system.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real problem is trust, not storage
&lt;/h2&gt;

&lt;p&gt;The hardest question in agent handoff is not:&lt;/p&gt;

&lt;p&gt;"Can Agent B access what Agent A saw?"&lt;/p&gt;

&lt;p&gt;It’s this:&lt;/p&gt;

&lt;p&gt;"What can Agent B trust?"&lt;/p&gt;

&lt;p&gt;A raw transcript is terrible at answering that.&lt;/p&gt;

&lt;p&gt;It mixes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;facts&lt;/li&gt;
&lt;li&gt;guesses&lt;/li&gt;
&lt;li&gt;abandoned plans&lt;/li&gt;
&lt;li&gt;temporary confusion&lt;/li&gt;
&lt;li&gt;old context that should have died three runs ago&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A structured memory layer is better because it lets you attach metadata to knowledge:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;source URL&lt;/li&gt;
&lt;li&gt;authoring agent&lt;/li&gt;
&lt;li&gt;timestamp&lt;/li&gt;
&lt;li&gt;confidence&lt;/li&gt;
&lt;li&gt;approval state&lt;/li&gt;
&lt;li&gt;freshness&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s what makes multi-agent systems feel solid.&lt;/p&gt;

&lt;p&gt;Not more context.&lt;/p&gt;

&lt;p&gt;Better contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I’d do if I were fixing an OpenClaw setup this week
&lt;/h2&gt;

&lt;p&gt;If your current setup is flaky, I’d start here:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Stop passing full transcripts between agents&lt;/li&gt;
&lt;li&gt;Add a structured artifact schema for facts, decisions, and outputs&lt;/li&gt;
&lt;li&gt;Store only durable artifacts in shared memory&lt;/li&gt;
&lt;li&gt;Retrieve only approved or high-confidence artifacts&lt;/li&gt;
&lt;li&gt;Add TTLs or stale markers so old junk doesn’t keep resurfacing&lt;/li&gt;
&lt;li&gt;Keep app state out of prompts&lt;/li&gt;
&lt;li&gt;Measure prompt size and handoff size per run&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you want one blunt rule:&lt;/p&gt;

&lt;p&gt;Every agent should produce artifacts that another agent can consume without reading the whole backstory.&lt;/p&gt;

&lt;p&gt;That’s the architecture change that matters.&lt;/p&gt;

&lt;p&gt;And if you’re running these systems at scale, this is also where API pricing stops being a side issue. Multi-agent workflows naturally create retries, summaries, handoffs, and long-running automation loops. Per-token billing makes all of that stressful. Predictable flat-rate compute is a much better fit when agents are running 24/7 and you don’t want every design choice to show up as a surprise bill.&lt;/p&gt;

&lt;p&gt;That’s the appeal of Standard Compute for this exact crowd: it’s a drop-in OpenAI-compatible API with flat monthly pricing, built for AI agents and automations. If you’re wiring OpenClaw into n8n, Make, Zapier, or custom orchestrators, having unlimited compute changes how aggressively you can test and refine memory architecture.&lt;/p&gt;

&lt;p&gt;The big shift is this:&lt;/p&gt;

&lt;p&gt;People think they need more memory.&lt;/p&gt;

&lt;p&gt;Usually they need better handoffs.&lt;/p&gt;

&lt;p&gt;And once you see that, a lot of multi-agent weirdness suddenly makes sense.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>python</category>
      <category>architecture</category>
    </item>
    <item>
      <title>I read the r/openclaw voice thread so you don’t have to — and yeah, the real problem is 10–20 second latency</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Mon, 20 Jul 2026 09:11:58 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-read-the-ropenclaw-voice-thread-so-you-dont-have-to-and-yeah-the-real-problem-is-10-20-1ceb</link>
      <guid>https://dev.to/lars_winstand/i-read-the-ropenclaw-voice-thread-so-you-dont-have-to-and-yeah-the-real-problem-is-10-20-1ceb</guid>
      <description>&lt;p&gt;A thread on &lt;a href="https://reddit.com/r/openclaw/comments/1v0oe4o/talk_to_openclaw/" rel="noopener noreferrer"&gt;r/openclaw&lt;/a&gt; got 10 upvotes and 18 comments.&lt;/p&gt;

&lt;p&gt;That’s not big Reddit.&lt;/p&gt;

&lt;p&gt;But small technical subreddits usually surface real problems faster than big ones. If 18 OpenClaw users keep circling the same issue, I pay attention.&lt;/p&gt;

&lt;p&gt;This thread started with a simple question: can you talk to OpenClaw?&lt;/p&gt;

&lt;p&gt;Not type.&lt;br&gt;
Not send a voice memo.&lt;br&gt;
Actually talk.&lt;/p&gt;

&lt;p&gt;After reading the whole thing, my takeaway is pretty simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice works. Conversation mostly doesn’t.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And the reason is not prompts or bad UX polish. It’s architecture and latency.&lt;/p&gt;
&lt;h2&gt;
  
  
  The thread in one sentence
&lt;/h2&gt;

&lt;p&gt;Most people in the thread can make voice input/output work with OpenClaw.&lt;/p&gt;

&lt;p&gt;What they can’t consistently get is something that feels like a live conversation.&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;There’s a huge difference between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;speaking into Telegram and getting a spoken reply later&lt;/li&gt;
&lt;li&gt;streaming audio over a persistent connection with interruption handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A lot of AI demos blur those together. Developers shouldn’t.&lt;/p&gt;
&lt;h2&gt;
  
  
  What people are actually building
&lt;/h2&gt;

&lt;p&gt;The comments were full of real setups, not theory.&lt;/p&gt;

&lt;p&gt;People mentioned combinations like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Telegram&lt;/strong&gt; voice notes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Discord&lt;/strong&gt; bots&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Open WebUI&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ElevenLabs&lt;/strong&gt; for TTS&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google TTS&lt;/strong&gt; as fallback&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Home Assistant&lt;/strong&gt; webhooks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parakeet v3&lt;/strong&gt; for STT&lt;/li&gt;
&lt;li&gt;custom local apps on &lt;strong&gt;Windows&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One example from the thread was basically:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;capture speech with OS dictation or Telegram voice notes&lt;/li&gt;
&lt;li&gt;send it through OpenClaw&lt;/li&gt;
&lt;li&gt;transcribe + generate a response&lt;/li&gt;
&lt;li&gt;run TTS with ElevenLabs or Google&lt;/li&gt;
&lt;li&gt;play the result via Home Assistant on a speaker&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is a valid system.&lt;/p&gt;

&lt;p&gt;It is also not what most people mean by “I want to talk to my agent.”&lt;/p&gt;

&lt;p&gt;It’s a spoken-message pipeline.&lt;/p&gt;
&lt;h2&gt;
  
  
  Voice notes are solved. Realtime conversation is not.
&lt;/h2&gt;

&lt;p&gt;If your goal is just hands-free input/output, the thread is actually encouraging.&lt;/p&gt;

&lt;p&gt;You can build something useful today.&lt;/p&gt;

&lt;p&gt;If your goal is “make this feel like ChatGPT voice mode,” the thread gets much less optimistic.&lt;/p&gt;

&lt;p&gt;That’s where latency starts killing the experience.&lt;/p&gt;
&lt;h2&gt;
  
  
  OpenClaw is not the whole voice stack
&lt;/h2&gt;

&lt;p&gt;A lot of confusion in the thread goes away once you separate &lt;strong&gt;OpenClaw&lt;/strong&gt; from the rest of the system.&lt;/p&gt;

&lt;p&gt;OpenClaw is a self-hosted gateway/control plane. It connects channels, agents, and models.&lt;/p&gt;

&lt;p&gt;It is not:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a realtime speech model&lt;/li&gt;
&lt;li&gt;a hosted low-latency audio transport&lt;/li&gt;
&lt;li&gt;a complete speech-to-speech runtime by itself&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So your actual voice experience depends on the full chain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;STT model&lt;/li&gt;
&lt;li&gt;LLM selection&lt;/li&gt;
&lt;li&gt;TTS engine&lt;/li&gt;
&lt;li&gt;transport format&lt;/li&gt;
&lt;li&gt;channel behavior&lt;/li&gt;
&lt;li&gt;buffering strategy&lt;/li&gt;
&lt;li&gt;whether you’re using files, chunks, or streams&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s why two people can both say “voice works with OpenClaw” and mean completely different things.&lt;/p&gt;

&lt;p&gt;One means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I can leave a Telegram voice note and get audio back.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Another means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I built a wake-word desktop client over WebSocket.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are not the same product.&lt;/p&gt;
&lt;h2&gt;
  
  
  The CLI tells the story
&lt;/h2&gt;

&lt;p&gt;OpenClaw’s own commands hint at what it is optimized for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw status
openclaw status &lt;span class="nt"&gt;--all&lt;/span&gt;
openclaw gateway status
openclaw status &lt;span class="nt"&gt;--deep&lt;/span&gt;
openclaw logs &lt;span class="nt"&gt;--follow&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And for onboarding the always-on gateway:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw onboard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s gateway-and-channel infrastructure.&lt;/p&gt;

&lt;p&gt;Not “instant voice assistant out of the box.”&lt;/p&gt;

&lt;p&gt;If you’re a developer building on top of OpenClaw, that’s fine. But you need to be honest about what layer you’re solving.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why ChatGPT Realtime feels better
&lt;/h2&gt;

&lt;p&gt;Because it solves the right problem at the transport level.&lt;/p&gt;

&lt;p&gt;One commenter in the thread said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Very easy to setup with chat gpt realtime, with tool calls and full access. Unfortunately, you have to pay per token for that, it's not part of subscription.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s the tradeoff in one sentence.&lt;/p&gt;

&lt;p&gt;The old voice stack usually looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;speech -&amp;gt; STT -&amp;gt; text LLM -&amp;gt; TTS -&amp;gt; audio
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every hop adds delay.&lt;/p&gt;

&lt;p&gt;A realtime stack is closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mic stream -&amp;gt; persistent socket -&amp;gt; model -&amp;gt; streamed audio response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s why &lt;strong&gt;OpenAI Realtime API&lt;/strong&gt; feels more natural. It was designed around:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;persistent WebSockets&lt;/li&gt;
&lt;li&gt;direct audio streaming&lt;/li&gt;
&lt;li&gt;interruption handling&lt;/li&gt;
&lt;li&gt;lower end-to-end latency&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is an architectural win, not a prompting trick.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thread is really comparing 3 approaches
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What you get&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenClaw + DIY voice stack&lt;/td&gt;
&lt;td&gt;Flexible and self-hosted, works across channels like Telegram and Discord, but latency depends on your STT, TTS, transport, and model choices&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI Realtime API&lt;/td&gt;
&lt;td&gt;Best shot at low-latency speech-to-speech with function calling and interruption support, but usage-based pricing brings back token anxiety&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Telegram/Discord voice workflows&lt;/td&gt;
&lt;td&gt;Easy to assemble and often cheap, but usually behave like async voice messages rather than live conversation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That’s basically the whole thread.&lt;/p&gt;

&lt;p&gt;Everything else is people choosing which compromise hurts least.&lt;/p&gt;

&lt;h2&gt;
  
  
  The delay numbers are the real story
&lt;/h2&gt;

&lt;p&gt;This was the most revealing part.&lt;/p&gt;

&lt;p&gt;A project shared in the thread, &lt;strong&gt;seven-voice&lt;/strong&gt;, exists because the maintainers said they &lt;strong&gt;couldn't find anything reliable&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That already tells you there’s a gap.&lt;/p&gt;

&lt;p&gt;Then the important quote:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“there's a little bit of a delay — nothing too terrible tho (no more than 10-20 seconds on average. Also that includes a 4.5 second delay after I'm done speaking which can be shortened).”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I’m going to be blunt:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;10–20 seconds is terrible for conversation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It may be acceptable for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;task dispatch&lt;/li&gt;
&lt;li&gt;voice notes&lt;/li&gt;
&lt;li&gt;smart home commands&lt;/li&gt;
&lt;li&gt;async bot workflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It is not acceptable for back-and-forth speech.&lt;/p&gt;

&lt;p&gt;If your app needs conversational rhythm, 10 seconds is forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick latency budget: where the time goes
&lt;/h2&gt;

&lt;p&gt;If you’re building one of these systems, here’s the practical way to think about it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User speaks                2.0s
Post-speech silence buffer 1.5s
Upload / transport         0.8s
STT                        1.2s
LLM response               3.5s
TTS                        1.5s
Playback startup           0.7s
------------------------------
Total                      11.2s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing there looks individually catastrophic.&lt;/p&gt;

&lt;p&gt;Together, it feels dead.&lt;/p&gt;

&lt;p&gt;That’s why a stack can be “working” and still feel broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical test for your own setup
&lt;/h2&gt;

&lt;p&gt;If you’re evaluating a voice stack, don’t ask only:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does it transcribe?&lt;/li&gt;
&lt;li&gt;Does it call tools?&lt;/li&gt;
&lt;li&gt;Does it speak back?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ask this instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;How many seconds from end-of-speech to first audible response?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That number matters more than most feature checklists.&lt;/p&gt;

&lt;p&gt;A rough rule:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&amp;lt; 1.5s&lt;/strong&gt;: feels live&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1.5s–3s&lt;/strong&gt;: usable&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3s–6s&lt;/strong&gt;: noticeably sluggish&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;6s+&lt;/strong&gt;: starts feeling async&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;10s+&lt;/strong&gt;: this is a voice note workflow, not a conversation&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why people don’t just switch to Realtime
&lt;/h2&gt;

&lt;p&gt;Because OpenClaw users are not only optimizing for latency.&lt;/p&gt;

&lt;p&gt;They’re optimizing for cost sanity.&lt;/p&gt;

&lt;p&gt;That part matters a lot.&lt;/p&gt;

&lt;p&gt;While looking into this thread, I also found another &lt;a href="https://reddit.com/r/openclaw/comments/1v10b2k/help_im_spending_100day_on_ai/" rel="noopener noreferrer"&gt;r/openclaw post&lt;/a&gt; where someone said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Since I installed OpenClaw 4 months ago I have spent over $10k on tokens via OpenRouter... Today I have 35 million input tokens, 600K output, and 81 million cached.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Once you’ve seen numbers like that, “just use the better realtime API” stops sounding casual.&lt;/p&gt;

&lt;p&gt;A persistent voice interface attached to an active agent can burn usage fast.&lt;/p&gt;

&lt;p&gt;That’s the real tension:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;low latency usually pushes you toward premium realtime APIs&lt;/li&gt;
&lt;li&gt;predictable cost pushes you toward flatter, more controlled infrastructure&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Cheap voice exists. Cheap good voice is the hard part.
&lt;/h2&gt;

&lt;p&gt;The thread shows people trying to keep costs under control with sensible choices:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Telegram&lt;/strong&gt; for capture&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OS dictation&lt;/strong&gt; or &lt;strong&gt;Parakeet v3&lt;/strong&gt; for STT&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google TTS&lt;/strong&gt; for low-cost output&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ElevenLabs&lt;/strong&gt; when they want better voice quality&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Home Assistant&lt;/strong&gt; for playback and device control&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are all reasonable engineering choices.&lt;/p&gt;

&lt;p&gt;The problem is that latency compounds across the stack.&lt;/p&gt;

&lt;p&gt;You don’t lose the experience in one place. You lose it everywhere, a little at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The most honest workaround was also the most custom
&lt;/h2&gt;

&lt;p&gt;One commenter built a small &lt;strong&gt;Windows&lt;/strong&gt; app with &lt;strong&gt;Claude&lt;/strong&gt; that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;listens for a wake word&lt;/li&gt;
&lt;li&gt;captures the spoken prompt&lt;/li&gt;
&lt;li&gt;sends it directly to the &lt;strong&gt;OpenClaw WebSocket&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;reads the response aloud locally&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s a smart design.&lt;/p&gt;

&lt;p&gt;It also tells you a lot.&lt;/p&gt;

&lt;p&gt;If people are writing custom desktop clients because the default path still feels slow, the demand is real.&lt;/p&gt;

&lt;p&gt;But so is the gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I’d recommend if you’re building this now
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1) Decide whether you need conversation or just voice I/O
&lt;/h3&gt;

&lt;p&gt;This is the biggest mistake in the whole category.&lt;/p&gt;

&lt;p&gt;If your actual use case is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;capturing notes&lt;/li&gt;
&lt;li&gt;dispatching tasks&lt;/li&gt;
&lt;li&gt;triggering automations&lt;/li&gt;
&lt;li&gt;controlling Home Assistant&lt;/li&gt;
&lt;li&gt;sending prompts while walking around&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;then async voice is probably enough.&lt;/p&gt;

&lt;p&gt;Use the simpler stack.&lt;/p&gt;

&lt;h3&gt;
  
  
  2) Measure end-to-end latency, not component latency
&lt;/h3&gt;

&lt;p&gt;Don’t benchmark STT and TTS separately and call it done.&lt;/p&gt;

&lt;p&gt;Measure from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;user stops speaking -&amp;gt; first audible response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s what users feel.&lt;/p&gt;

&lt;h3&gt;
  
  
  3) Prefer streaming transports over file-based workflows
&lt;/h3&gt;

&lt;p&gt;If your flow still looks like “record blob, upload blob, transcribe blob,” you’re already behind.&lt;/p&gt;

&lt;p&gt;For example, this is the kind of architecture that tends to age badly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// slow-ish pattern&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;audioFile&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;recordUntilSilence&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;upload&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;audioFile&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;transcribe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;audioFile&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;reply&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;speech&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;tts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nf"&gt;play&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;speech&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You want something closer to a persistent session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// better direction conceptually&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ws&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;WebSocket&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;REALTIME_ENDPOINT&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nx"&gt;mic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;chunk&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;ws&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;onmessage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;audio&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;speaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;play&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not because WebSockets are magical, but because conversational systems need streaming behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  4) Be careful with usage-based pricing for always-on agents
&lt;/h3&gt;

&lt;p&gt;If your team is building voice into automations, copilots, or internal agents that run all day, per-token billing gets painful fast.&lt;/p&gt;

&lt;p&gt;That’s where a flat-cost layer starts to matter.&lt;/p&gt;

&lt;p&gt;For teams already using OpenAI-compatible SDKs, a service like &lt;strong&gt;&lt;a href="https://standardcompute.com" rel="noopener noreferrer"&gt;Standard Compute&lt;/a&gt;&lt;/strong&gt; is interesting for a different reason than flashy demos:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;flat monthly pricing&lt;/li&gt;
&lt;li&gt;OpenAI-compatible API surface&lt;/li&gt;
&lt;li&gt;no per-token billing anxiety&lt;/li&gt;
&lt;li&gt;easier to let agents run continuously&lt;/li&gt;
&lt;li&gt;dynamic model routing across &lt;strong&gt;GPT-5.4&lt;/strong&gt;, &lt;strong&gt;Claude Opus 4.6&lt;/strong&gt;, and &lt;strong&gt;Grok 4.20&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That doesn’t magically solve realtime audio transport by itself.&lt;/p&gt;

&lt;p&gt;But it does solve one of the other big problems in the thread: people are scared to leave agents running because the bill can spiral.&lt;/p&gt;

&lt;p&gt;If you’re building voice-enabled automations on &lt;strong&gt;n8n&lt;/strong&gt;, &lt;strong&gt;Make&lt;/strong&gt;, &lt;strong&gt;Zapier&lt;/strong&gt;, &lt;strong&gt;OpenClaw&lt;/strong&gt;, or custom agent stacks, predictable cost is not a side issue. It changes what you’re willing to ship.&lt;/p&gt;

&lt;h2&gt;
  
  
  My take
&lt;/h2&gt;

&lt;p&gt;The r/openclaw thread is not really about whether voice is possible.&lt;/p&gt;

&lt;p&gt;It is.&lt;/p&gt;

&lt;p&gt;It’s about whether it feels alive.&lt;/p&gt;

&lt;p&gt;Right now, the landscape looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;if you want &lt;strong&gt;cheap&lt;/strong&gt;, you can stitch together a decent async voice workflow&lt;/li&gt;
&lt;li&gt;if you want &lt;strong&gt;good&lt;/strong&gt;, realtime APIs still have the cleanest architecture&lt;/li&gt;
&lt;li&gt;if you want &lt;strong&gt;good and cost-predictable&lt;/strong&gt;, you’re still doing a lot of engineering&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s not an OpenClaw failure. That’s just where voice agents are right now.&lt;/p&gt;

&lt;p&gt;The practical takeaway is boring, but useful:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;if you only need hands-free input/output, stop chasing realtime demos and build the simple thing&lt;/li&gt;
&lt;li&gt;if you need actual conversation, don’t pretend 10–20 seconds is acceptable&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because it isn’t.&lt;/p&gt;

&lt;p&gt;And everyone in that thread already knows it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>openai</category>
      <category>automation</category>
      <category>webdev</category>
    </item>
    <item>
      <title>The first OpenClaw setup I’d recommend for a blind parent is not a chatbot</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Mon, 20 Jul 2026 01:11:03 +0000</pubDate>
      <link>https://dev.to/lars_winstand/the-first-openclaw-setup-id-recommend-for-a-blind-parent-is-not-a-chatbot-219p</link>
      <guid>https://dev.to/lars_winstand/the-first-openclaw-setup-id-recommend-for-a-blind-parent-is-not-a-chatbot-219p</guid>
      <description>&lt;p&gt;I knew this was going to annoy some agent builders the second I started reading the threads.&lt;/p&gt;

&lt;p&gt;Because yes, OpenClaw can absolutely be part of the solution.&lt;/p&gt;

&lt;p&gt;That’s not the real question.&lt;/p&gt;

&lt;p&gt;The real question is: should a blind, non-technical parent talk directly to OpenClaw as their primary interface?&lt;/p&gt;

&lt;p&gt;My answer is no.&lt;/p&gt;

&lt;p&gt;Not for v1.&lt;/p&gt;

&lt;p&gt;After reading a really solid r/openclaw thread about a 72-year-old, fully blind, Spanish-speaking senior in Argentina, I came away with a much less glamorous answer than “build a general AI assistant.”&lt;/p&gt;

&lt;p&gt;Build a constrained voice shell.&lt;/p&gt;

&lt;p&gt;Not a chatbot.&lt;/p&gt;

&lt;p&gt;Not an open-ended agent loop.&lt;/p&gt;

&lt;p&gt;A voice-first interface with a wake word, speech-to-text, text-to-speech, and a hard-scoped action router for a very small command set.&lt;/p&gt;

&lt;p&gt;Once voice latency drifts into the 10–20 second range, the whole thing stops feeling assistive and starts feeling broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  The product is the constraint
&lt;/h2&gt;

&lt;p&gt;If the user wants to say things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Play Argentine news on YouTube&lt;/li&gt;
&lt;li&gt;Continue my Audible book&lt;/li&gt;
&lt;li&gt;Play tango on Spotify&lt;/li&gt;
&lt;li&gt;Read my newest email&lt;/li&gt;
&lt;li&gt;Call my daughter&lt;/li&gt;
&lt;li&gt;Lower the TV volume&lt;/li&gt;
&lt;li&gt;Tell me what I can do&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;...then that list is not a limitation.&lt;/p&gt;

&lt;p&gt;That list is the product.&lt;/p&gt;

&lt;p&gt;A blind senior does not need an agent that might do anything.&lt;/p&gt;

&lt;p&gt;They need a voice assistant that will definitely do a handful of things, every time, with predictable behavior.&lt;/p&gt;

&lt;p&gt;That changes the architecture immediately.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Reddit comment that got it right
&lt;/h2&gt;

&lt;p&gt;One commenter in the original thread said the quiet part out loud:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;OpenClaw is a terminal-based tool — not great for a blind non-tech user directly. You’d need a voice layer on top (wake-word + STT/TTS bridge) to make it work.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s basically the whole design brief.&lt;/p&gt;

&lt;p&gt;The mistake is assuming the hard part is the LLM.&lt;/p&gt;

&lt;p&gt;It isn’t.&lt;/p&gt;

&lt;p&gt;The hard part is building a spoken interface that feels reliable when the user cannot fall back to a screen.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I’d actually build
&lt;/h2&gt;

&lt;p&gt;If I were doing this for my own family, I’d split the system into four boring pieces.&lt;/p&gt;

&lt;p&gt;That’s a compliment.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Wake word or push-to-talk
&lt;/h3&gt;

&lt;p&gt;The user needs a clear start signal.&lt;/p&gt;

&lt;p&gt;If they can’t see whether the assistant is listening, ambiguous states are poison.&lt;/p&gt;

&lt;p&gt;Good options:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;physical push-to-talk button&lt;/li&gt;
&lt;li&gt;local wake word listener on a Raspberry Pi&lt;/li&gt;
&lt;li&gt;microphone array attached to a Home Assistant box&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A giant glowing button is not less advanced than a wake word.&lt;/p&gt;

&lt;p&gt;It’s often better.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Speech-to-text
&lt;/h3&gt;

&lt;p&gt;For Spanish, I’d test Whisper first.&lt;/p&gt;

&lt;p&gt;If you’re already in Home Assistant land, Wyoming makes this pretty straightforward.&lt;/p&gt;

&lt;p&gt;Example architecture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mic -&amp;gt; wake word -&amp;gt; Whisper STT -&amp;gt; intent router -&amp;gt; action -&amp;gt; Piper/ElevenLabs TTS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If local STT accuracy is bad, use a cloud STT service.&lt;/p&gt;

&lt;p&gt;This is not the place to be ideological.&lt;/p&gt;

&lt;p&gt;If the transcript is wrong, everything after it is wrong too.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Safe action router
&lt;/h3&gt;

&lt;p&gt;This is the part people skip because prompting Claude Opus 4.6 or GPT-5.4 is more fun.&lt;/p&gt;

&lt;p&gt;But the router is the whole game.&lt;/p&gt;

&lt;p&gt;The assistant should map spoken intents to a finite set of approved actions.&lt;/p&gt;

&lt;p&gt;Not “open a browser and figure it out.”&lt;/p&gt;

&lt;p&gt;Not “log into random websites and click around.”&lt;/p&gt;

&lt;p&gt;Actual verbs tied to actual APIs and actual devices.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;wake word
  -&amp;gt; STT
  -&amp;gt; intent classifier/router
  -&amp;gt; approved action OR OpenClaw fallback
  -&amp;gt; TTS response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the permissions should be brutally explicit.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ALLOW:
- read_latest_email
- call_approved_contact
- play_spotify_playlist
- resume_audible
- play_youtube_news
- set_tv_volume
- list_available_commands

DENY:
- send_email
- delete_email
- purchase_item
- change_account_settings
- reset_password
- install_app
- open_browser_freely
- edit_contacts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That may look restrictive.&lt;/p&gt;

&lt;p&gt;Good.&lt;/p&gt;

&lt;p&gt;The original requirement in the Reddit thread was basically: allow email reading, but prevent deletion, sending, purchases, and account changes.&lt;/p&gt;

&lt;p&gt;That’s not a side constraint.&lt;/p&gt;

&lt;p&gt;That is the product requirement.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Text-to-speech
&lt;/h3&gt;

&lt;p&gt;For Spanish output:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Piper is a strong local option&lt;/li&gt;
&lt;li&gt;ElevenLabs is still hard to beat for naturalness&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I like local-first where possible, but if a cloud TTS voice is dramatically easier for the user to understand, I’d pick usability over purity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I would not put OpenClaw at the front door
&lt;/h2&gt;

&lt;p&gt;Because voice makes every weakness feel 10x worse.&lt;/p&gt;

&lt;p&gt;When an n8n workflow fails, you inspect logs.&lt;/p&gt;

&lt;p&gt;When a Zapier run stalls, you inspect the run history.&lt;/p&gt;

&lt;p&gt;When a terminal tool gets weird, a technical user pokes around until it behaves.&lt;/p&gt;

&lt;p&gt;A blind senior cannot do any of that.&lt;/p&gt;

&lt;p&gt;And the biggest problem in current voice-agent setups is not raw intelligence.&lt;/p&gt;

&lt;p&gt;It’s latency.&lt;/p&gt;

&lt;p&gt;In another r/openclaw thread about talking to OpenClaw, users reported response delays around 10–15 seconds, and in one workaround setup, 10–20 seconds on average with a 4.5 second post-speech delay.&lt;/p&gt;

&lt;p&gt;That is not a cosmetic UX issue.&lt;/p&gt;

&lt;p&gt;That is the difference between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“this helps me”&lt;/li&gt;
&lt;li&gt;and “this thing is dead again”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you can see a screen, maybe you’ll tolerate a spinner.&lt;/p&gt;

&lt;p&gt;If you can’t, silence is ambiguous.&lt;/p&gt;

&lt;p&gt;Did it hear me?&lt;/p&gt;

&lt;p&gt;Did it crash?&lt;/p&gt;

&lt;p&gt;Is it still recording?&lt;/p&gt;

&lt;p&gt;Should I repeat myself?&lt;/p&gt;

&lt;p&gt;That’s why I think the first version should avoid open-ended agent loops unless they are absolutely necessary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where OpenClaw does belong
&lt;/h2&gt;

&lt;p&gt;In the back room.&lt;/p&gt;

&lt;p&gt;Not at the front door.&lt;/p&gt;

&lt;p&gt;OpenClaw is useful when the request actually needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;reasoning&lt;/li&gt;
&lt;li&gt;tool selection&lt;/li&gt;
&lt;li&gt;multi-step orchestration&lt;/li&gt;
&lt;li&gt;summarization over approved context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I would use deterministic paths for common commands, and only hand off to OpenClaw when the request falls outside a known route.&lt;/p&gt;

&lt;p&gt;Example split:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“Lower the TV volume” -&amp;gt; direct Home Assistant entity action&lt;/li&gt;
&lt;li&gt;“Play tango on Spotify” -&amp;gt; direct Spotify intent&lt;/li&gt;
&lt;li&gt;“Read my newest email” -&amp;gt; read-only email function&lt;/li&gt;
&lt;li&gt;“What can I do?” -&amp;gt; static help response in Spanish&lt;/li&gt;
&lt;li&gt;“What did my daughter say about tomorrow’s appointment?” -&amp;gt; maybe now invoke OpenClaw for summarization over approved email/message context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That split matters because it keeps the common path fast.&lt;/p&gt;

&lt;p&gt;And speed is accessibility.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack I’d pick first
&lt;/h2&gt;

&lt;p&gt;Here’s the honest version.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;Best use case&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenClaw + custom voice shell&lt;/td&gt;
&lt;td&gt;Flexible agent behavior if you have engineering time and can tolerate setup/debug work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Home Assistant Assist + Wyoming + Piper + Whisper&lt;/td&gt;
&lt;td&gt;Best first build for constrained voice commands, smart-home control, and predictable behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Alexa with Spanish support&lt;/td&gt;
&lt;td&gt;Best choice when the family wants the least maintenance and can live with less customization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;My actual opinion: for a blind parent, Home Assistant Assist is the better starting point than raw OpenClaw.&lt;/p&gt;

&lt;p&gt;That does not mean OpenClaw is bad.&lt;/p&gt;

&lt;p&gt;It means the interface problem matters more than the agent problem.&lt;/p&gt;

&lt;p&gt;And honestly, the Alexa argument is fair too.&lt;/p&gt;

&lt;p&gt;If the family wants something that works for months without anyone SSH-ing into a mini PC on Sunday afternoon, Alexa may beat a custom stack.&lt;/p&gt;

&lt;p&gt;That is not a defeat.&lt;/p&gt;

&lt;p&gt;That is adult engineering.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical v1 blueprint
&lt;/h2&gt;

&lt;p&gt;If I had to sketch this tomorrow, I would keep it very small.&lt;/p&gt;

&lt;h3&gt;
  
  
  Components
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Home Assistant Assist as the primary voice interface&lt;/li&gt;
&lt;li&gt;Whisper via Wyoming for Spanish STT testing&lt;/li&gt;
&lt;li&gt;Piper or ElevenLabs for Spanish TTS&lt;/li&gt;
&lt;li&gt;A wake word or large physical push-to-talk button&lt;/li&gt;
&lt;li&gt;A strict action router for media, calls, read-only email, and Home Assistant entities&lt;/li&gt;
&lt;li&gt;OpenClaw only as fallback for approved reasoning tasks&lt;/li&gt;
&lt;li&gt;Human-approved setup for contacts, devices, playlists, inbox access, and blocked actions&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Minimal flow
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User speaks
  -&amp;gt; wake word/button
  -&amp;gt; STT
  -&amp;gt; intent match
  -&amp;gt; direct action if known
  -&amp;gt; OpenClaw fallback if approved and necessary
  -&amp;gt; TTS response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Example intent router pseudocode
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;route_intent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;intent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;classify_intent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;play_news&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;play_youtube_channel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Argentine News&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;resume_audible&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;resume_audible&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;read_latest_email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;read_latest_email&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;read_only&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;call_contact&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;contact&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;extract_contact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;call_if_approved&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;contact&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;set_tv_volume&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;level&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;extract_volume_level&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;set_tv_volume&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;level&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;help&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;speak_available_commands&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;es&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;APPROVED_REASONING_TASKS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;run_openclaw_with_scoped_tools&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;transcript&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;speak&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Lo siento, no puedo hacer eso todavía.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Example deny-by-default tool policy
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"allow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"read_latest_email"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"call_approved_contact"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"play_spotify_playlist"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"resume_audible"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"play_youtube_news"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"set_tv_volume"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"list_available_commands"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"deny"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"send_email"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"delete_email"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"purchase_item"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"change_account_settings"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"reset_password"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"install_app"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"open_browser_freely"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"edit_contacts"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The cost problem changes the design too
&lt;/h2&gt;

&lt;p&gt;There’s another reason I would not leave a freeform agent hanging open all day.&lt;/p&gt;

&lt;p&gt;Usage explodes.&lt;/p&gt;

&lt;p&gt;Voice assistants create lots of short interactions.&lt;/p&gt;

&lt;p&gt;Retries add extra turns.&lt;/p&gt;

&lt;p&gt;STT and TTS wrap every request.&lt;/p&gt;

&lt;p&gt;Testing latency fixes means even more calls.&lt;/p&gt;

&lt;p&gt;If you’re paying per token, pricing stops being a backend detail and starts changing product decisions.&lt;/p&gt;

&lt;p&gt;That matters a lot if you’re building agents or automations that stay available all day.&lt;/p&gt;

&lt;p&gt;I ran into a separate OpenClaw user saying they had spent more than $10k on tokens in four months, with 35 million input tokens, 600k output, and 81 million cached.&lt;/p&gt;

&lt;p&gt;That’s obviously not a normal home accessibility setup.&lt;/p&gt;

&lt;p&gt;But it is a useful warning for anyone building persistent AI systems.&lt;/p&gt;

&lt;p&gt;This is where a flat-rate, subscription LLM setup becomes more than a pricing preference.&lt;/p&gt;

&lt;p&gt;It changes whether you feel free to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;test retries&lt;/li&gt;
&lt;li&gt;add fallback flows&lt;/li&gt;
&lt;li&gt;run background automations&lt;/li&gt;
&lt;li&gt;keep an agent available 24/7&lt;/li&gt;
&lt;li&gt;iterate on multimodal voice behavior without staring at a token meter&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you’re building on n8n, Make, Zapier, OpenClaw, or your own agent framework, predictable cost matters because reliable systems require more guardrails than demos do.&lt;/p&gt;

&lt;p&gt;That’s one reason Standard Compute is interesting for this kind of work: it gives you an OpenAI-compatible API with flat monthly pricing instead of per-token billing, so you can build and test agent-heavy workflows without every retry feeling like a billing event.&lt;/p&gt;

&lt;p&gt;For accessibility work especially, that matters.&lt;/p&gt;

&lt;p&gt;Accessible systems need confirmations, guardrails, redundancy, and fallback behavior.&lt;/p&gt;

&lt;p&gt;Those are good engineering choices.&lt;/p&gt;

&lt;p&gt;They also generate more model traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to prototype this
&lt;/h2&gt;

&lt;p&gt;Here’s a rough starting point for a Home Assistant-style local stack.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# example only&lt;/span&gt;
&lt;span class="c"&gt;# Raspberry Pi / Linux host&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt update
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt &lt;span class="nb"&gt;install &lt;/span&gt;docker.io docker-compose-plugin

&lt;span class="nb"&gt;mkdir &lt;/span&gt;voice-assistant-stack
&lt;span class="nb"&gt;cd &lt;/span&gt;voice-assistant-stack
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You’d then wire up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Home Assistant&lt;/li&gt;
&lt;li&gt;Wyoming services for Whisper/Piper&lt;/li&gt;
&lt;li&gt;a local or network microphone endpoint&lt;/li&gt;
&lt;li&gt;webhook-based actions for approved commands&lt;/li&gt;
&lt;li&gt;optional OpenClaw fallback service&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you’re building a custom router service, keep it small and observable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# example Python service&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv
&lt;span class="nb"&gt;source&lt;/span&gt; .venv/bin/activate
pip &lt;span class="nb"&gt;install &lt;/span&gt;fastapi uvicorn pydantic
uvicorn app:app &lt;span class="nt"&gt;--reload&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And log every step:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[voice] wake word detected
[voice] transcript: "lee mi correo más reciente"
[router] matched intent=read_latest_email
[action] provider=gmail mode=read_only
[tts] response generated in 1.2s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you can’t debug the spoken path quickly, you will hate maintaining it.&lt;/p&gt;

&lt;h2&gt;
  
  
  My opinion, plainly
&lt;/h2&gt;

&lt;p&gt;The best first OpenClaw setup for a blind parent is barely an OpenClaw setup at all.&lt;/p&gt;

&lt;p&gt;It’s a voice-first assistant with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;hard edges&lt;/li&gt;
&lt;li&gt;fast paths&lt;/li&gt;
&lt;li&gt;explicit permissions&lt;/li&gt;
&lt;li&gt;a very short list of things it does well&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is less magical than “build an AI companion.”&lt;/p&gt;

&lt;p&gt;Good.&lt;/p&gt;

&lt;p&gt;Magic is overrated when the user is 72, blind, and needs the assistant to work on the first try.&lt;/p&gt;

&lt;p&gt;The Reddit thread got the core instinct right.&lt;/p&gt;

&lt;p&gt;Take the user seriously.&lt;/p&gt;

&lt;p&gt;Start with a constrained voice shell.&lt;/p&gt;

&lt;p&gt;Then add agent behavior only where it clearly improves the experience.&lt;/p&gt;

&lt;p&gt;Not where it makes the demo cooler.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>a11y</category>
      <category>opensource</category>
      <category>devops</category>
    </item>
    <item>
      <title>I thought a cheaper model would fix my agent bill, then I found the 34k-character system prompt</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Sun, 19 Jul 2026 17:12:07 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-thought-a-cheaper-model-would-fix-my-agent-bill-then-i-found-the-34k-character-system-prompt-2e26</link>
      <guid>https://dev.to/lars_winstand/i-thought-a-cheaper-model-would-fix-my-agent-bill-then-i-found-the-34k-character-system-prompt-2e26</guid>
      <description>&lt;p&gt;I’ve seen this same debugging pattern too many times now.&lt;/p&gt;

&lt;p&gt;Agent costs spike. Everyone blames the model. Someone says, “Move from Claude Opus to Claude Sonnet.” Someone else says, “Cap the context window.” Then the team spends two days arguing about whether GPT-5 is worth it.&lt;/p&gt;

&lt;p&gt;Sometimes that helps.&lt;/p&gt;

&lt;p&gt;But a lot of the time, the real problem is way less glamorous: nobody knows what actually happened inside the run.&lt;/p&gt;

&lt;p&gt;That’s why I think agent tracing matters more than model shopping.&lt;/p&gt;

&lt;p&gt;One OpenClaw debugging case made this painfully obvious. Preflight estimated &lt;strong&gt;10,698 prompt tokens&lt;/strong&gt; and overflowed. Compaction then reported &lt;strong&gt;0 messages&lt;/strong&gt; to summarize. The run kept failing anyway.&lt;/p&gt;

&lt;p&gt;That is not a pricing problem.&lt;/p&gt;

&lt;p&gt;That is an observability problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bill is scary. The mystery is worse.
&lt;/h2&gt;

&lt;p&gt;If all you have is a final token total, you’re basically debugging from a receipt.&lt;/p&gt;

&lt;p&gt;But an agent run is not one prompt.&lt;/p&gt;

&lt;p&gt;It’s usually some combination of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;system prompt injection&lt;/li&gt;
&lt;li&gt;conversation history&lt;/li&gt;
&lt;li&gt;tool schemas&lt;/li&gt;
&lt;li&gt;retrieved docs&lt;/li&gt;
&lt;li&gt;browser output&lt;/li&gt;
&lt;li&gt;exec output&lt;/li&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;summarization&lt;/li&gt;
&lt;li&gt;memory writes&lt;/li&gt;
&lt;li&gt;the final LLM call&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you can’t see those as separate steps, you’re guessing.&lt;/p&gt;

&lt;p&gt;And guessing is how teams end up “optimizing” the model while the real cost driver is a bloated context assembly pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  The OpenClaw case that changed how I think about agent spend
&lt;/h2&gt;

&lt;p&gt;I came across a thread on r/openclaw where a user said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;It started doing insane looping and used up a bunch of credit.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sentence is basically the shared trauma of agent engineering in 2026.&lt;/p&gt;

&lt;p&gt;The interesting follow-up was a separate OpenClaw debug thread where the logs showed this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;Compacting context (0 messages)...&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Auto-compaction could not recover this turn.&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And then the numbers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;systemPromptChars=34,549&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;estimatedPromptTokens=10,698&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;contextTokenBudget=40,960&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;reserveTokens=32,960&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;promptBudgetBeforeReserve=8,000&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;overflowTokens=2,698&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;historyTextChars=0&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That mismatch is the whole story.&lt;/p&gt;

&lt;p&gt;Preflight saw a giant prompt and said, “we’re over budget.”&lt;/p&gt;

&lt;p&gt;Compaction looked at conversation history only, saw zero messages to summarize, and said, “nothing to do here.”&lt;/p&gt;

&lt;p&gt;So you had two subsystems looking at different slices of the same request.&lt;/p&gt;

&lt;p&gt;That’s exactly why end-to-end tracing matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is what “hidden context cost” looks like
&lt;/h2&gt;

&lt;p&gt;The most useful part of OpenClaw’s docs is that they make the overhead visible.&lt;/p&gt;

&lt;p&gt;Here’s the kind of context breakdown they show:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;system prompt: &lt;code&gt;38,412 chars&lt;/code&gt; (~&lt;code&gt;9,603 tokens&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;project context: &lt;code&gt;23,901 chars&lt;/code&gt; (~&lt;code&gt;5,976 tokens&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;tool schemas: &lt;code&gt;31,988 chars&lt;/code&gt; (~&lt;code&gt;7,997 tokens&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;browser schema: ~&lt;code&gt;2,453 tokens&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;exec schema: ~&lt;code&gt;1,560 tokens&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;session tokens cached: &lt;code&gt;14,250 / ctx=32,000&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why “the model is expensive” is often the wrong diagnosis.&lt;/p&gt;

&lt;p&gt;Sometimes the expensive thing is not Claude Opus, GPT-5, or Grok.&lt;/p&gt;

&lt;p&gt;Sometimes it’s your own baggage:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;giant system prompts&lt;/li&gt;
&lt;li&gt;too many tools&lt;/li&gt;
&lt;li&gt;verbose tool schemas&lt;/li&gt;
&lt;li&gt;injected files the agent barely needs&lt;/li&gt;
&lt;li&gt;browser dumps that never get trimmed&lt;/li&gt;
&lt;li&gt;retry loops that keep replaying all of the above&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Compaction is useful, but it won’t save a bad context strategy
&lt;/h2&gt;

&lt;p&gt;A lot of teams treat compaction like a magic cleanup button.&lt;/p&gt;

&lt;p&gt;It isn’t.&lt;/p&gt;

&lt;p&gt;OpenClaw separates &lt;strong&gt;compaction&lt;/strong&gt; from &lt;strong&gt;session pruning&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Compaction summarizes older conversation turns. Pruning trims old tool results from active memory.&lt;/p&gt;

&lt;p&gt;That helps when the problem is chat history growth.&lt;/p&gt;

&lt;p&gt;It does not help much when the real problem is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a 34k-character system prompt&lt;/li&gt;
&lt;li&gt;8k tokens of tool schema&lt;/li&gt;
&lt;li&gt;unnecessary workspace injection&lt;/li&gt;
&lt;li&gt;browser and exec tools dumping huge outputs every turn&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If compaction only sees message history, it can’t shrink what it never touched.&lt;/p&gt;

&lt;h2&gt;
  
  
  First thing I would run in OpenClaw
&lt;/h2&gt;

&lt;p&gt;If you suspect hidden context bloat, these are the commands I’d start with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/status
/context list
/context detail
/context map
/usage tokens
/compact
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gives you a fast way to inspect what’s actually being included.&lt;/p&gt;

&lt;p&gt;If you want to change the compaction model, you can do that too:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"agents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"defaults"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"compaction"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/anthropic/claude-sonnet-4-6"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Useful? Yes.&lt;/p&gt;

&lt;p&gt;A full explanation of where the spend came from? No.&lt;/p&gt;

&lt;p&gt;For that, you need tracing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trace the run, not just the model call
&lt;/h2&gt;

&lt;p&gt;My opinion: the right unit of analysis for agents is &lt;strong&gt;one request with nested spans&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Not one prompt.&lt;/p&gt;

&lt;p&gt;That means tracing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;context assembly&lt;/li&gt;
&lt;li&gt;retrieval steps&lt;/li&gt;
&lt;li&gt;tool calls&lt;/li&gt;
&lt;li&gt;tool outputs&lt;/li&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;summarization&lt;/li&gt;
&lt;li&gt;downstream LLM calls&lt;/li&gt;
&lt;li&gt;final response generation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the agent loops, you want to know:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;what caused the first retry&lt;/li&gt;
&lt;li&gt;what got replayed on each retry&lt;/li&gt;
&lt;li&gt;which step exploded token usage&lt;/li&gt;
&lt;li&gt;whether the expensive part was actually the model call at all&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That’s where tools like LangSmith, Helicone, and OpenTelemetry-style tracing become useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical observability options
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What it’s good for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenClaw built-in inspection&lt;/td&gt;
&lt;td&gt;Fast inspection of context contributors like system prompts, files, skills, tool schemas, and token usage. Good for local debugging. Not full end-to-end tracing.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LangSmith observability&lt;/td&gt;
&lt;td&gt;Full traces with nested spans across LLM calls, tools, and higher-level functions. Best when you need to debug multi-step agent workflows.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Helicone gateway/observability&lt;/td&gt;
&lt;td&gt;Good gateway-style visibility across OpenAI, Anthropic, OpenRouter, and others with logging, retries, caching, and request metadata.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you want a quick LangSmith setup, it’s straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;LANGSMITH_TRACING&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true
export &lt;/span&gt;&lt;span class="nv"&gt;LANGSMITH_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&amp;lt;your-langsmith-api-key&amp;gt;"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&amp;lt;your-openai-api-key&amp;gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then wrap your model client and your tool functions so one agent run shows up as one trace tree.&lt;/p&gt;

&lt;p&gt;That’s when debugging gets a lot less philosophical.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to fix after tracing
&lt;/h2&gt;

&lt;p&gt;Tracing won’t reduce cost by itself.&lt;/p&gt;

&lt;p&gt;It just tells you where to cut.&lt;/p&gt;

&lt;p&gt;The fixes are usually boring, which is probably why they work.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Shrink the system prompt
&lt;/h3&gt;

&lt;p&gt;If your system prompt is &lt;code&gt;34,549&lt;/code&gt; or &lt;code&gt;38,412&lt;/code&gt; characters, start there.&lt;/p&gt;

&lt;p&gt;That is a huge tax on every single turn.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Cut tool schema bloat
&lt;/h3&gt;

&lt;p&gt;Tool schemas are easy to ignore because they feel like infrastructure.&lt;/p&gt;

&lt;p&gt;They still cost tokens.&lt;/p&gt;

&lt;p&gt;If browser schema is ~&lt;code&gt;2,453 tokens&lt;/code&gt; and exec schema is ~&lt;code&gt;1,560 tokens&lt;/code&gt;, that overhead adds up fast.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Inject fewer files and skills
&lt;/h3&gt;

&lt;p&gt;Don’t give the agent 15 capabilities if the task needs 3.&lt;/p&gt;

&lt;p&gt;Context should be assembled per task, not by habit.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Prune tool output aggressively
&lt;/h3&gt;

&lt;p&gt;Browser content and shell output are repeat offenders.&lt;/p&gt;

&lt;p&gt;If you keep replaying giant results back into the next turn, your context window will disappear fast.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Kill retry loops early
&lt;/h3&gt;

&lt;p&gt;If a tool is failing or the agent is stuck, unlimited retries are not resilience.&lt;/p&gt;

&lt;p&gt;They are a billing strategy, just a bad one.&lt;/p&gt;

&lt;p&gt;A hard turn limit is often the simplest protection.&lt;/p&gt;

&lt;h2&gt;
  
  
  The model still matters. It’s just not the first question.
&lt;/h2&gt;

&lt;p&gt;I’m not saying model pricing is irrelevant.&lt;/p&gt;

&lt;p&gt;It matters a lot.&lt;/p&gt;

&lt;p&gt;Claude Opus 4.6, Claude Sonnet 4.6, GPT-5, Grok 4.20, Qwen, and Llama all have different cost/performance tradeoffs.&lt;/p&gt;

&lt;p&gt;And yes, cheaper models can absolutely reduce spend.&lt;/p&gt;

&lt;p&gt;But if your workflow is replaying a giant system prompt, huge tool schemas, and stale context on every retry, switching models is just making the same bug slightly cheaper.&lt;/p&gt;

&lt;p&gt;That’s not optimization. That’s damage control.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters even more for automation teams
&lt;/h2&gt;

&lt;p&gt;If you’re running agents inside n8n, Make, Zapier, OpenClaw, or a custom automation stack, the problem gets worse because the agent isn’t isolated.&lt;/p&gt;

&lt;p&gt;It’s part of a workflow.&lt;/p&gt;

&lt;p&gt;So one bad loop can trigger:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;repeated tool calls&lt;/li&gt;
&lt;li&gt;repeated API requests&lt;/li&gt;
&lt;li&gt;repeated LLM calls&lt;/li&gt;
&lt;li&gt;repeated browser sessions&lt;/li&gt;
&lt;li&gt;repeated summarization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s exactly why per-token pricing gets painful at scale.&lt;/p&gt;

&lt;p&gt;You’re not just paying for one clever answer.&lt;/p&gt;

&lt;p&gt;You’re paying for every invisible retry, every oversized prompt, and every step your pipeline replays.&lt;/p&gt;

&lt;p&gt;That’s also why I think flat-rate AI infrastructure is becoming more attractive for teams running real automations. If your agents run all day, predictable pricing is a relief.&lt;/p&gt;

&lt;p&gt;Standard Compute is interesting here because it gives you an OpenAI-compatible API with flat monthly pricing instead of per-token billing. So if your team is already using OpenAI SDKs or wiring models into n8n/Make/Zapier/custom agents, you can swap the endpoint without rebuilding everything.&lt;/p&gt;

&lt;p&gt;That does not replace tracing.&lt;/p&gt;

&lt;p&gt;But it does remove a lot of the token anxiety while you fix the actual workflow problems.&lt;/p&gt;

&lt;p&gt;And honestly, that combination is what most teams want:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;visibility into what the agent is doing&lt;/li&gt;
&lt;li&gt;predictable cost while it runs&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My practical rule now
&lt;/h2&gt;

&lt;p&gt;Before changing models:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;inspect context assembly&lt;/li&gt;
&lt;li&gt;trace tool calls&lt;/li&gt;
&lt;li&gt;trace retries&lt;/li&gt;
&lt;li&gt;measure schema overhead&lt;/li&gt;
&lt;li&gt;find the span that exploded&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then decide whether the model is actually the problem.&lt;/p&gt;

&lt;p&gt;Because a surprising number of “LLM cost problems” are really:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;context assembly problems&lt;/li&gt;
&lt;li&gt;retry policy problems&lt;/li&gt;
&lt;li&gt;tool design problems&lt;/li&gt;
&lt;li&gt;observability problems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The OpenClaw example is memorable because the numbers are so absurd.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;34k-character system prompt&lt;/strong&gt; was enough to push preflight over budget, while compaction saw &lt;strong&gt;0 messages&lt;/strong&gt; and had nothing to summarize.&lt;/p&gt;

&lt;p&gt;If you only looked at the final bill, you’d probably blame the model.&lt;/p&gt;

&lt;p&gt;If you traced the run, you’d know exactly where to start.&lt;/p&gt;

&lt;p&gt;That’s the difference.&lt;/p&gt;

&lt;p&gt;And once you can see the sequence, runaway agent costs stop feeling random.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>observability</category>
    </item>
    <item>
      <title>I thought Grok subscriptions were the cheap way to run agents until the limits got weird</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Sun, 19 Jul 2026 01:11:43 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-thought-grok-subscriptions-were-the-cheap-way-to-run-agents-until-the-limits-got-weird-1j11</link>
      <guid>https://dev.to/lars_winstand/i-thought-grok-subscriptions-were-the-cheap-way-to-run-agents-until-the-limits-got-weird-1j11</guid>
      <description>&lt;p&gt;I kept seeing the same advice: buy X Premium or SuperGrok, connect Grok to OpenClaw with OAuth, and skip per-token billing.&lt;/p&gt;

&lt;p&gt;For solo tinkering, that sounds great.&lt;/p&gt;

&lt;p&gt;No API key setup. No usage dashboard open in another tab. No tiny panic every time your agent decides to summarize half the internet.&lt;/p&gt;

&lt;p&gt;But once your agent stops being a chatbot and starts acting like infrastructure, the pricing story gets a lot less cute.&lt;/p&gt;

&lt;p&gt;That was the thing I underestimated.&lt;/p&gt;

&lt;p&gt;A Grok subscription can be fine for experiments. For always-on agents, fuzzy quotas and account eligibility are not a minor detail. They are the whole reliability model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup is genuinely convenient
&lt;/h2&gt;

&lt;p&gt;OpenClaw makes the Grok path look easy, because it is easy.&lt;/p&gt;

&lt;p&gt;You can authenticate with OAuth instead of creating an xAI API key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw models auth login &lt;span class="nt"&gt;--provider&lt;/span&gt; xai &lt;span class="nt"&gt;--method&lt;/span&gt; oauth
openclaw models &lt;span class="nb"&gt;set &lt;/span&gt;xai/grok-4.3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you're running OpenClaw on a VPS, Raspberry Pi, or some always-on Linux box, device auth feels pretty slick.&lt;/p&gt;

&lt;p&gt;You sign in once and your agent is off to the races.&lt;/p&gt;

&lt;p&gt;For a personal bot in Discord or Telegram, that convenience matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fine print is where it gets weird
&lt;/h2&gt;

&lt;p&gt;Here is the part that changed my mind:&lt;/p&gt;

&lt;p&gt;OpenClaw's xAI docs say xAI decides which accounts are eligible to receive OAuth API tokens.&lt;/p&gt;

&lt;p&gt;If your account is not eligible, you need to use the API-key route instead.&lt;/p&gt;

&lt;p&gt;That is not just an auth detail.&lt;/p&gt;

&lt;p&gt;That means your production-ish automation may depend on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;your consumer subscription state&lt;/li&gt;
&lt;li&gt;your account's OAuth eligibility&lt;/li&gt;
&lt;li&gt;whatever usage ceilings exist behind that subscription&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a human using Grok in a browser, that is annoying.&lt;/p&gt;

&lt;p&gt;For an agent that runs 24/7, preserves memory, calls tools, and wakes up from webhooks, that is operational risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reddit had the most honest signal
&lt;/h2&gt;

&lt;p&gt;The most useful info I found was not on a pricing page.&lt;/p&gt;

&lt;p&gt;It was in an r/openclaw thread where someone asked what the SuperGrok Heavy plan actually gets you, and whether it works with OAuth in OpenClaw.&lt;/p&gt;

&lt;p&gt;One reply said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Grok, even just the 30 USD plan, can be used through OpenClaw OAuth. I do challenge my subscription token limit each month though.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sentence is doing a lot of work.&lt;/p&gt;

&lt;p&gt;Translation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;yes, it works&lt;/li&gt;
&lt;li&gt;yes, there is some ceiling&lt;/li&gt;
&lt;li&gt;no, the ceiling is not especially legible&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For hobby use, maybe that's fine.&lt;/p&gt;

&lt;p&gt;For agent workloads, I want the opposite of mystery.&lt;/p&gt;

&lt;h2&gt;
  
  
  API pricing may be more expensive, but at least it is legible
&lt;/h2&gt;

&lt;p&gt;xAI's API page is much clearer.&lt;/p&gt;

&lt;p&gt;It exposes an OpenAI-compatible endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://api.x.ai/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the integration looks boring in the best possible way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;OpenAI&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;XAI_API_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://api.x.ai/v1&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;grok-4.5&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Summarize the latest failed jobs in plain English&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That does not automatically make the API cheaper.&lt;/p&gt;

&lt;p&gt;It does make it understandable.&lt;/p&gt;

&lt;p&gt;And for anything always-on, understandable beats vibes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents are not power users. They are chaos with retries
&lt;/h2&gt;

&lt;p&gt;This is the part people miss when they compare subscription pricing to API pricing.&lt;/p&gt;

&lt;p&gt;Agents do not behave like careful humans.&lt;/p&gt;

&lt;p&gt;Agents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;loop&lt;/li&gt;
&lt;li&gt;retry&lt;/li&gt;
&lt;li&gt;pull too much context&lt;/li&gt;
&lt;li&gt;call the same tool three times&lt;/li&gt;
&lt;li&gt;wake up in the middle of the night because n8n fired a webhook&lt;/li&gt;
&lt;li&gt;keep going after you stop watching&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes unclear limits much more dangerous.&lt;/p&gt;

&lt;p&gt;Another r/openclaw thread had a user describing an early Grok run like this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;It started doing insane looping and used up a bunch of credit.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is not an edge case.&lt;/p&gt;

&lt;p&gt;That is normal agent behavior when your prompts, tool limits, or context strategy are not tight enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real problem is not just price. It is architecture.
&lt;/h2&gt;

&lt;p&gt;The best datapoint I found was from a third r/openclaw discussion.&lt;/p&gt;

&lt;p&gt;A user said they were averaging 41.5K tokens per message on simple tasks before optimizing.&lt;/p&gt;

&lt;p&gt;That should make any automation engineer stop scrolling.&lt;/p&gt;

&lt;p&gt;Once you are at that level, your problem is bigger than whether Grok via subscription is cheaper than Grok via API.&lt;/p&gt;

&lt;p&gt;Your agent is carrying too much baggage into every turn.&lt;/p&gt;

&lt;p&gt;Likely causes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;too much conversation history&lt;/li&gt;
&lt;li&gt;too much memory injected every time&lt;/li&gt;
&lt;li&gt;giant retrieval payloads&lt;/li&gt;
&lt;li&gt;one general-purpose agent doing five jobs badly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The best fix in that thread was also the most practical: split the agent into smaller specialized workers.&lt;/p&gt;

&lt;p&gt;That matches what actually works in production.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;worker 1: classify request
worker 2: retrieve only relevant docs
worker 3: execute tool calls
worker 4: write final response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead of one giant agent with every doc, every tool, and every memory blob attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  OAuth vs API key: what actually changes
&lt;/h2&gt;

&lt;p&gt;This is the comparison that matters once the agent is always on.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What changes when the agent is always on?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Grok via OpenClaw OAuth&lt;/td&gt;
&lt;td&gt;Fast to start. Human login flow. Subscription-based usage. Good for personal agents and testing. Riskier when uptime depends on account eligibility and unclear ceilings.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;xAI API&lt;/td&gt;
&lt;td&gt;API key auth. Explicit pricing. Easier to reason about in production. Better fit for service accounts, shared systems, and OpenAI-compatible clients.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenRouter via OpenClaw&lt;/td&gt;
&lt;td&gt;More explicit model routing and fallback options. Better fit if you want to swap providers or build resilience into automations.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you build with n8n, Make, Zapier, or custom workers, the pattern is usually the same:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;credentials in env vars or vaults&lt;/li&gt;
&lt;li&gt;predictable endpoints&lt;/li&gt;
&lt;li&gt;retries you control&lt;/li&gt;
&lt;li&gt;fallbacks you can script&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is why API-key auth fits automation better than user OAuth in most serious cases.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would do before trusting a subscription plan with real agents
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Measure context per task
&lt;/h3&gt;

&lt;p&gt;Do not just track monthly spend.&lt;/p&gt;

&lt;p&gt;Track tokens per request type.&lt;/p&gt;

&lt;p&gt;You want to know whether the expensive path is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;retrieval&lt;/li&gt;
&lt;li&gt;memory injection&lt;/li&gt;
&lt;li&gt;planning&lt;/li&gt;
&lt;li&gt;tool retries&lt;/li&gt;
&lt;li&gt;final response generation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are using OpenAI-compatible clients, add logging around request size and response size.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;estimateChars&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reduce&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;sum&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;messageCount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;approxChars&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;estimateChars&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not perfect, but enough to catch obvious abuse fast.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Split generalist agents into specialized workers
&lt;/h3&gt;

&lt;p&gt;A single all-knowing agent is usually the expensive design.&lt;/p&gt;

&lt;p&gt;A better pattern is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;router agent&lt;/li&gt;
&lt;li&gt;retrieval worker&lt;/li&gt;
&lt;li&gt;action worker&lt;/li&gt;
&lt;li&gt;summarizer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That reduces context and makes failures easier to isolate.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Put hard limits on loops and tool retries
&lt;/h3&gt;

&lt;p&gt;If your framework allows it, cap tool recursion and retries.&lt;/p&gt;

&lt;p&gt;Pseudocode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;MAX_TOOL_CALLS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;MAX_RETRIES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;toolCalls&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;MAX_TOOL_CALLS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Tool call limit exceeded&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This sounds obvious until an agent burns through usage because one tool kept returning malformed JSON.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Use API credentials for persistent or shared workflows
&lt;/h3&gt;

&lt;p&gt;If the workflow is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;always on&lt;/li&gt;
&lt;li&gt;shared by a team&lt;/li&gt;
&lt;li&gt;tied to business operations&lt;/li&gt;
&lt;li&gt;expected to survive restarts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;use API credentials.&lt;/p&gt;

&lt;p&gt;This is not me being anti-OAuth.&lt;/p&gt;

&lt;p&gt;It is just the cleaner operational model.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Keep OAuth subscriptions for experiments and personal bots
&lt;/h3&gt;

&lt;p&gt;This is where I landed.&lt;/p&gt;

&lt;p&gt;Grok via OpenClaw OAuth is a smart convenience feature.&lt;/p&gt;

&lt;p&gt;It is great for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;testing model behavior&lt;/li&gt;
&lt;li&gt;personal assistants&lt;/li&gt;
&lt;li&gt;low-stakes Discord or Telegram bots&lt;/li&gt;
&lt;li&gt;short-lived experiments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It is much less convincing as the foundation for infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Standard Compute fits
&lt;/h2&gt;

&lt;p&gt;This is also why flat-rate API access is appealing if you are running lots of automations.&lt;/p&gt;

&lt;p&gt;The real pain is not just paying for tokens.&lt;/p&gt;

&lt;p&gt;It is having to think about token economics every time you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;add memory&lt;/li&gt;
&lt;li&gt;increase context&lt;/li&gt;
&lt;li&gt;run agents continuously&lt;/li&gt;
&lt;li&gt;connect another workflow in n8n or Make&lt;/li&gt;
&lt;li&gt;let multiple workers operate in parallel&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Standard Compute takes the opposite approach: predictable monthly pricing with OpenAI-compatible API access, so you can plug it into existing SDKs and automation stacks without doing per-token math all day.&lt;/p&gt;

&lt;p&gt;If your team is building agents that run like infrastructure, that model makes a lot more sense than hoping a consumer subscription behaves like a production service.&lt;/p&gt;

&lt;h2&gt;
  
  
  My actual takeaway
&lt;/h2&gt;

&lt;p&gt;I do not think the lesson is "never use Grok subscriptions."&lt;/p&gt;

&lt;p&gt;I think the lesson is narrower and more useful:&lt;/p&gt;

&lt;p&gt;A subscription that feels cheap for a human can get weird fast for an autonomous agent.&lt;/p&gt;

&lt;p&gt;OAuth is great when you are playing.&lt;/p&gt;

&lt;p&gt;It gets shaky when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the agent is always on&lt;/li&gt;
&lt;li&gt;the workload is autonomous&lt;/li&gt;
&lt;li&gt;the usage ceiling is unclear&lt;/li&gt;
&lt;li&gt;uptime matters&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the line.&lt;/p&gt;

&lt;p&gt;If your setup is a weekend prototype, Grok via OpenClaw OAuth might be perfect.&lt;/p&gt;

&lt;p&gt;If your setup is drifting toward real infrastructure, boring beats clever.&lt;/p&gt;

&lt;p&gt;Use explicit APIs. Measure context. Limit loops. Prefer pricing you can explain to yourself at 2 a.m.&lt;/p&gt;

&lt;p&gt;Because "it probably works" is not a cost model.&lt;/p&gt;

&lt;p&gt;It is suspense.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>automation</category>
      <category>devops</category>
    </item>
    <item>
      <title>I think Facebook Marketplace posting is the sleeper real estate AI automation project on OpenClaw (6-step workflow)</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Sat, 18 Jul 2026 17:10:55 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-think-facebook-marketplace-posting-is-the-sleeper-real-estate-ai-automation-project-on-openclaw-52mk</link>
      <guid>https://dev.to/lars_winstand/i-think-facebook-marketplace-posting-is-the-sleeper-real-estate-ai-automation-project-on-openclaw-52mk</guid>
      <description>&lt;p&gt;I keep seeing the same bad real estate AI demo: paste in a few property notes, get back a polished paragraph, call it automation.&lt;/p&gt;

&lt;p&gt;That is not the hard part.&lt;/p&gt;

&lt;p&gt;The hard part is everything around the paragraph:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;missing fields&lt;/li&gt;
&lt;li&gt;bad photo sets&lt;/li&gt;
&lt;li&gt;inconsistent formats&lt;/li&gt;
&lt;li&gt;approval bottlenecks&lt;/li&gt;
&lt;li&gt;handoff into the actual posting flow&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;While digging through an r/openclaw thread about workflows people were oddly proud of, the one that stuck with me was Facebook Marketplace listing ops, not some giant autonomous agent. That felt right.&lt;/p&gt;

&lt;p&gt;Because this is where AI stops being a toy and starts acting like operations.&lt;/p&gt;

&lt;p&gt;If you build it well, a Facebook Marketplace workflow on OpenClaw can remove 15 to 20 minutes of manual work per listing. Not by being magical. By being structured.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real opportunity is not copy generation
&lt;/h2&gt;

&lt;p&gt;Writing listing copy is cheap now.&lt;/p&gt;

&lt;p&gt;GPT-5.4 can do it. Claude Opus 4.6 can do it. Grok 4.20 can do it. Smaller models can do it too if the prompt is clean.&lt;/p&gt;

&lt;p&gt;So if your whole automation is just:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;take notes&lt;/li&gt;
&lt;li&gt;generate description&lt;/li&gt;
&lt;li&gt;done&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;...you automated the least valuable part.&lt;/p&gt;

&lt;p&gt;The real win is turning messy listing inputs into something a human can approve fast.&lt;/p&gt;

&lt;p&gt;That is why Facebook Marketplace is a better automation target than a generic "listing bot." It forces you to solve the operational mess.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;What happens in practice&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Generic chatbot demo&lt;/td&gt;
&lt;td&gt;Produces text, but ignores missing fields, image quality, and approval flow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One-shot listing generator&lt;/td&gt;
&lt;td&gt;Creates a draft fast, but still leaves humans doing the ops work manually&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No-review autoposting bot&lt;/td&gt;
&lt;td&gt;Feels clever until bad data, UI changes, or policy issues create a mess&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenClaw listing ops workflow&lt;/td&gt;
&lt;td&gt;Handles intake, validation, drafting, review, and posting prep as one system&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the one I would actually ship.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 6-step workflow I would build in OpenClaw
&lt;/h2&gt;

&lt;p&gt;The useful version is not "AI, write me a listing."&lt;/p&gt;

&lt;p&gt;The useful version is a pipeline with guardrails.&lt;/p&gt;

&lt;p&gt;Here is the 6-step version that makes sense:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;intake form or CRM trigger&lt;/li&gt;
&lt;li&gt;field validation&lt;/li&gt;
&lt;li&gt;photo checks&lt;/li&gt;
&lt;li&gt;AI draft generation&lt;/li&gt;
&lt;li&gt;human approval queue&lt;/li&gt;
&lt;li&gt;posting prep&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you want to wire that up in OpenClaw, the flow looks more like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Google Sheets / Airtable / HubSpot
        -&amp;gt; OpenClaw trigger
        -&amp;gt; required field validator
        -&amp;gt; photo quality + duplicate checks
        -&amp;gt; LLM draft generation
        -&amp;gt; policy / formatting validation
        -&amp;gt; approval queue
        -&amp;gt; posting payload prep
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is already much more valuable than a chatbot.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each step should actually do
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1) Intake form or CRM trigger
&lt;/h3&gt;

&lt;p&gt;Start with structured input.&lt;/p&gt;

&lt;p&gt;Good sources:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Google Sheets&lt;/li&gt;
&lt;li&gt;Airtable&lt;/li&gt;
&lt;li&gt;HubSpot&lt;/li&gt;
&lt;li&gt;internal admin form&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Minimum fields I would require:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"address"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"123 Main St, Austin, TX"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"price"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;425000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"bedrooms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"bathrooms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"square_feet"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1840&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"property_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"single_family"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contact_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Jane Doe"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contact_phone"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"555-0102"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"photo_urls"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"https://.../1.jpg"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://.../2.jpg"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your intake is free-form, your downstream automation will be bad.&lt;/p&gt;

&lt;h3&gt;
  
  
  2) Field validation
&lt;/h3&gt;

&lt;p&gt;Before you spend any tokens, reject incomplete listings.&lt;/p&gt;

&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;missing price&lt;/li&gt;
&lt;li&gt;missing city/state&lt;/li&gt;
&lt;li&gt;0 bedrooms on a residential listing&lt;/li&gt;
&lt;li&gt;invalid phone number&lt;/li&gt;
&lt;li&gt;square footage missing when required&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Pseudo-validation logic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;validateListing&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;listing&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;errors&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;listing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;price&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;missing price&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;listing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;address&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;missing address&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;listing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;contact_phone&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;missing contact phone&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isArray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;listing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;photo_urls&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;listing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;photo_urls&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;not enough photos&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;errors&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This step is boring. That is why it matters.&lt;/p&gt;

&lt;h3&gt;
  
  
  3) Photo checks
&lt;/h3&gt;

&lt;p&gt;This is where a lot of "AI listing automation" quietly falls apart.&lt;/p&gt;

&lt;p&gt;You do not need perfect computer vision. You just need useful filters:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;duplicate image detection&lt;/li&gt;
&lt;li&gt;low-resolution detection&lt;/li&gt;
&lt;li&gt;missing cover image&lt;/li&gt;
&lt;li&gt;obviously broken URLs&lt;/li&gt;
&lt;li&gt;weird image count mismatches&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Even a lightweight image screening step saves reviewers time.&lt;/p&gt;

&lt;h3&gt;
  
  
  4) AI draft generation
&lt;/h3&gt;

&lt;p&gt;Now use the model.&lt;/p&gt;

&lt;p&gt;Generate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Facebook Marketplace title&lt;/li&gt;
&lt;li&gt;full description&lt;/li&gt;
&lt;li&gt;shorter mobile-friendly description&lt;/li&gt;
&lt;li&gt;optional follow-up message templates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example prompt shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Write a Facebook Marketplace real estate listing.

Requirements:
- Keep title under 80 characters
- Description should be clear, factual, and non-hypey
- Do not invent features not present in input
- Include beds, baths, square footage, location, and CTA
- Avoid risky claims like "best deal" or unverifiable superlatives

Input:
{listing_json}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At this point, GPT-5.4 and Claude Opus 4.6 are both strong choices. I would pick based on output consistency and cost model, not ideology.&lt;/p&gt;

&lt;h3&gt;
  
  
  5) Human approval queue
&lt;/h3&gt;

&lt;p&gt;I would not skip this.&lt;/p&gt;

&lt;p&gt;Not for Facebook Marketplace.&lt;/p&gt;

&lt;p&gt;Not for real estate.&lt;/p&gt;

&lt;p&gt;Not for anything where bad data creates support work later.&lt;/p&gt;

&lt;p&gt;Send the generated draft into:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a Slack approval flow&lt;/li&gt;
&lt;li&gt;an Airtable review status column&lt;/li&gt;
&lt;li&gt;a HubSpot task&lt;/li&gt;
&lt;li&gt;an internal admin panel&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The reviewer should see:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;original input&lt;/li&gt;
&lt;li&gt;validation warnings&lt;/li&gt;
&lt;li&gt;generated title&lt;/li&gt;
&lt;li&gt;generated description&lt;/li&gt;
&lt;li&gt;photo summary&lt;/li&gt;
&lt;li&gt;approve / reject / edit actions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the difference between a flashy automation and a usable one.&lt;/p&gt;

&lt;h3&gt;
  
  
  6) Posting prep
&lt;/h3&gt;

&lt;p&gt;I am intentionally saying posting prep, not blind autoposting.&lt;/p&gt;

&lt;p&gt;Why?&lt;/p&gt;

&lt;p&gt;Because UI-driven automations around Facebook Marketplace are brittle.&lt;/p&gt;

&lt;p&gt;A safer pattern is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;normalize approved fields&lt;/li&gt;
&lt;li&gt;package final assets&lt;/li&gt;
&lt;li&gt;generate a posting-ready payload&lt;/li&gt;
&lt;li&gt;hand it to the operator or downstream tool&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example output object:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"approved"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"marketplace_title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"3BR Home in South Austin - Updated Kitchen"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"marketplace_description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Well-maintained 3 bed, 2 bath home..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cover_image"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://.../cover.jpg"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"photo_order"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"cover.jpg"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"kitchen.jpg"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"living-room.jpg"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contact_name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Jane Doe"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contact_phone"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"555-0102"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gets the team 80 to 90 percent of the way there without pretending the last 10 percent is free.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the simple version breaks in production
&lt;/h2&gt;

&lt;p&gt;The first prototype always looks good.&lt;/p&gt;

&lt;p&gt;Then reality shows up.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the address is incomplete&lt;/li&gt;
&lt;li&gt;the square footage is missing&lt;/li&gt;
&lt;li&gt;the title is too long&lt;/li&gt;
&lt;li&gt;the photos are out of order&lt;/li&gt;
&lt;li&gt;the CTA is wrong&lt;/li&gt;
&lt;li&gt;the seller wants a different tone&lt;/li&gt;
&lt;li&gt;Facebook Marketplace wants one format while your CRM stores another&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now your reviewer is fixing machine-generated slop in three browser tabs.&lt;/p&gt;

&lt;p&gt;That is why I am much more bullish on ops-style automation than chatbot-style automation.&lt;/p&gt;

&lt;p&gt;Chatbots are easy to demo.&lt;/p&gt;

&lt;p&gt;Pipelines are what survive contact with production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why human review is not a compromise
&lt;/h2&gt;

&lt;p&gt;A lot of teams still treat human review like failure.&lt;/p&gt;

&lt;p&gt;I think that is backwards.&lt;/p&gt;

&lt;p&gt;For this kind of workflow, human review is the feature.&lt;/p&gt;

&lt;p&gt;Facebook Marketplace is exactly the kind of environment where brittle automation creates hidden costs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;UI changes&lt;/li&gt;
&lt;li&gt;moderation quirks&lt;/li&gt;
&lt;li&gt;duplicate content issues&lt;/li&gt;
&lt;li&gt;edge-case listing details&lt;/li&gt;
&lt;li&gt;account risk if low-quality posts slip through&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A review queue keeps the system useful.&lt;/p&gt;

&lt;p&gt;The strongest version of this project is not a robot posting unsupervised.&lt;/p&gt;

&lt;p&gt;It is a workflow that standardizes inputs, handles the repetitive work, and gives a human a clean final checkpoint.&lt;/p&gt;

&lt;p&gt;That wins in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this creates token anxiety fast
&lt;/h2&gt;

&lt;p&gt;This is the part people underestimate.&lt;/p&gt;

&lt;p&gt;A listing workflow does not make one model call and stop.&lt;/p&gt;

&lt;p&gt;It keeps going:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;intake checks&lt;/li&gt;
&lt;li&gt;image analysis&lt;/li&gt;
&lt;li&gt;draft generation&lt;/li&gt;
&lt;li&gt;rewrite passes&lt;/li&gt;
&lt;li&gt;validation&lt;/li&gt;
&lt;li&gt;follow-up handling&lt;/li&gt;
&lt;li&gt;retries when upstream data is messy&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do that across dozens or hundreds of listings and per-token pricing starts to feel like a tax on operational ambition.&lt;/p&gt;

&lt;p&gt;Every extra safeguard costs money.&lt;br&gt;
Every retry costs money.&lt;br&gt;
Every always-on agent watching Airtable, HubSpot, or Google Sheets costs money.&lt;/p&gt;

&lt;p&gt;That is why the API layer matters.&lt;/p&gt;

&lt;p&gt;If your OpenClaw workflow uses an OpenAI-compatible API, you can swap the backend without rebuilding your automations.&lt;/p&gt;

&lt;p&gt;That is a huge deal.&lt;/p&gt;

&lt;p&gt;For teams building always-on listing pipelines, Standard Compute is a strong fit here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;flat monthly pricing&lt;/li&gt;
&lt;li&gt;OpenAI-compatible API&lt;/li&gt;
&lt;li&gt;dynamic routing across GPT-5.4, Claude Opus 4.6, and Grok 4.20&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That changes how you design workflows.&lt;/p&gt;

&lt;p&gt;You stop optimizing for fear.&lt;/p&gt;

&lt;p&gt;You can afford:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;an extra validation pass&lt;/li&gt;
&lt;li&gt;a second draft when the first one is weak&lt;/li&gt;
&lt;li&gt;classification checks&lt;/li&gt;
&lt;li&gt;background watchers running 24/7&lt;/li&gt;
&lt;li&gt;more aggressive retries for bad upstream data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is what makes agentic automation dependable instead of fragile.&lt;/p&gt;
&lt;h2&gt;
  
  
  A practical OpenAI-compatible integration example
&lt;/h2&gt;

&lt;p&gt;If OpenClaw or your custom worker is calling an OpenAI-style endpoint, the code does not need to get weird.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.standardcompute.com/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$STANDARD_COMPUTE_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "openai/gpt-5.4",
    "messages": [
      {"role": "system", "content": "You write accurate Facebook Marketplace real estate listings."},
      {"role": "user", "content": "Generate a title and description for this property: 3 bed, 2 bath, 1840 sq ft in South Austin, updated kitchen, fenced yard."}
    ]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or with the OpenAI SDK pattern in Node:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;OpenAI&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;STANDARD_COMPUTE_API_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://api.standardcompute.com/v1&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;completion&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai/gpt-5.4&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;system&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;You write accurate Facebook Marketplace real estate listings. Never invent facts.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="na"&gt;address&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;123 Main St, Austin, TX&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;bedrooms&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;bathrooms&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;square_feet&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1840&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;features&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;updated kitchen&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;fenced yard&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
      &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;completion&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That matters because it keeps the workflow portable.&lt;/p&gt;

&lt;p&gt;You can use the same integration style across OpenClaw, n8n, Make, Zapier, or a custom worker.&lt;/p&gt;

&lt;h2&gt;
  
  
  If I were building this this week
&lt;/h2&gt;

&lt;p&gt;I would ship it in this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Airtable or Google Sheets intake&lt;/li&gt;
&lt;li&gt;validation rules&lt;/li&gt;
&lt;li&gt;LLM draft generation&lt;/li&gt;
&lt;li&gt;approval queue&lt;/li&gt;
&lt;li&gt;photo checks&lt;/li&gt;
&lt;li&gt;posting payload export&lt;/li&gt;
&lt;li&gt;optional follow-up message automation&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Not the other way around.&lt;/p&gt;

&lt;p&gt;Most teams overbuild autoposting before they build data quality controls.&lt;/p&gt;

&lt;p&gt;That is a mistake.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part people are underestimating
&lt;/h2&gt;

&lt;p&gt;Facebook Marketplace sounds small until you map the workflow.&lt;/p&gt;

&lt;p&gt;Then it turns into a miniature listing operating system:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;structured intake&lt;/li&gt;
&lt;li&gt;asset handling&lt;/li&gt;
&lt;li&gt;model routing&lt;/li&gt;
&lt;li&gt;approval logic&lt;/li&gt;
&lt;li&gt;publishing prep&lt;/li&gt;
&lt;li&gt;auditability&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is much closer to how real automation teams think than the endless stream of AI assistant demos.&lt;/p&gt;

&lt;p&gt;That is why this project stood out to me.&lt;/p&gt;

&lt;p&gt;It is not trying to impress anyone with fake autonomy.&lt;/p&gt;

&lt;p&gt;It is solving the annoying sequence of tasks businesses actually pay to remove.&lt;/p&gt;

&lt;p&gt;And once you see it that way, the sleeper idea is not "AI writes listings."&lt;/p&gt;

&lt;p&gt;It is that Facebook Marketplace posting becomes the wedge into a full listing ops system.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>openai</category>
      <category>realestate</category>
    </item>
    <item>
      <title>I finally saw a legal agent setup that used OpenClaw for 6 months without pretending to be your lawyer</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Sat, 18 Jul 2026 09:11:12 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-finally-saw-a-legal-agent-setup-that-used-openclaw-for-6-months-without-pretending-to-be-your-4787</link>
      <guid>https://dev.to/lars_winstand/i-finally-saw-a-legal-agent-setup-that-used-openclaw-for-6-months-without-pretending-to-be-your-4787</guid>
      <description>&lt;p&gt;I keep seeing the same bad question come up in AI threads:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can GPT-5, Claude, Qwen, or Llama do legal work yet?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question leads straight to the worst demos.&lt;/p&gt;

&lt;p&gt;You get polished text, fake confidence, and citations that look plausible until someone actually checks them.&lt;/p&gt;

&lt;p&gt;Then I ran into a thread on r/openclaw from someone using OpenClaw during a divorce and custody case in Japan, and it was the first legal-agent setup I’ve seen that felt operationally sane.&lt;/p&gt;

&lt;p&gt;Not because it was flashy.&lt;/p&gt;

&lt;p&gt;Because it was constrained.&lt;/p&gt;

&lt;p&gt;The user wasn’t asking OpenClaw to be a lawyer. They were using it like a very disciplined paralegal with access to a lot of records and zero authority to act on its own.&lt;/p&gt;

&lt;p&gt;That difference is everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one rule that made the whole setup credible
&lt;/h2&gt;

&lt;p&gt;This line was the key:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Boundaries are respected. Nothing external ever gets sent without my explicit approval. It drafts; I approve.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the right architecture for legal AI.&lt;/p&gt;

&lt;p&gt;Not:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“the model is smart now”&lt;/li&gt;
&lt;li&gt;“the benchmark score is higher”&lt;/li&gt;
&lt;li&gt;“we added a legal system prompt”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Just a hard boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenClaw drafts.
Human approves.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s more mature than most legal AI discourse.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the actual stack looked like
&lt;/h2&gt;

&lt;p&gt;The user described a setup on a Mac mini with Discord channels, an Obsidian repo, and access to email, calendar, and the file system.&lt;/p&gt;

&lt;p&gt;Over about 6 months, they used it to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;organize thousands of case artifacts&lt;/li&gt;
&lt;li&gt;maintain chronology&lt;/li&gt;
&lt;li&gt;translate documents&lt;/li&gt;
&lt;li&gt;draft bilingual correspondence&lt;/li&gt;
&lt;li&gt;cross-check evidence across sources&lt;/li&gt;
&lt;li&gt;build a private case site for counsel&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The architecture was simple enough to describe in one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Mac mini + Discord + Obsidian + email/calendar/files
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is not an “AI lawyer.”&lt;/p&gt;

&lt;p&gt;That is an evidence operations stack.&lt;/p&gt;

&lt;p&gt;And that’s why it feels safer than most chatbot demos.&lt;/p&gt;

&lt;h2&gt;
  
  
  The useful legal-agent pattern is narrower than people want
&lt;/h2&gt;

&lt;p&gt;The moment you ask a model for a final legal answer, you push it toward its weakest behavior.&lt;/p&gt;

&lt;p&gt;LLMs want to complete the pattern. They want to sound finished.&lt;/p&gt;

&lt;p&gt;That is exactly what you do not want in legal work.&lt;/p&gt;

&lt;p&gt;But if you narrow the job, agents become a lot more useful.&lt;/p&gt;

&lt;p&gt;The four jobs that actually make sense are:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Organize chronology&lt;/li&gt;
&lt;li&gt;Translate and draft without sending&lt;/li&gt;
&lt;li&gt;Cross-reference evidence across systems&lt;/li&gt;
&lt;li&gt;Flag weak support and missing citations&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That’s the pattern.&lt;/p&gt;

&lt;p&gt;Not “do law.”&lt;/p&gt;

&lt;p&gt;More like: “help me manage a messy record without losing provenance.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is better than a one-shot legal chatbot
&lt;/h2&gt;

&lt;p&gt;A one-shot chatbot can answer fast.&lt;/p&gt;

&lt;p&gt;It can also confidently hand you unsupported nonsense.&lt;/p&gt;

&lt;p&gt;I’d take a slower workflow that says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;This claim is backed by:
- a photo
- a timestamp
- a calendar entry
- GPS corroboration
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;...over a fast workflow that says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Based on applicable law...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;...and then invents a citation.&lt;/p&gt;

&lt;p&gt;That false-citation problem came up in the thread too, and honestly, good. People should be paranoid about it.&lt;/p&gt;

&lt;p&gt;The fix is not “trust the model less” as a vague principle.&lt;/p&gt;

&lt;p&gt;The fix is to design the workflow so the model mostly handles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;retrieval&lt;/li&gt;
&lt;li&gt;structure&lt;/li&gt;
&lt;li&gt;drafting&lt;/li&gt;
&lt;li&gt;comparison&lt;/li&gt;
&lt;li&gt;evidence ranking&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And not final legal authority.&lt;/p&gt;

&lt;h2&gt;
  
  
  The smartest part: evidence confidence tiers
&lt;/h2&gt;

&lt;p&gt;The best detail in the whole setup was how the user thought about evidence strength.&lt;/p&gt;

&lt;p&gt;They described a parenting journal that cross-referenced photos with smartwatch GPS data to build a tiered, timestamped record of time spent with their kid.&lt;/p&gt;

&lt;p&gt;The pattern was basically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;photo exists -&amp;gt; photo + GPS track confirms location
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s a much better frame than “AI summarized my files.”&lt;/p&gt;

&lt;p&gt;Because not all evidence is equal.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a screenshot is not the same as a screenshot plus metadata&lt;/li&gt;
&lt;li&gt;a photo is not the same as a photo plus GPS corroboration&lt;/li&gt;
&lt;li&gt;a remembered date is not the same as a date confirmed by email, calendar, and message exports&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This idea generalizes really well.&lt;/p&gt;

&lt;p&gt;It works for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;legal assist&lt;/li&gt;
&lt;li&gt;compliance reviews&lt;/li&gt;
&lt;li&gt;HR investigations&lt;/li&gt;
&lt;li&gt;insurance disputes&lt;/li&gt;
&lt;li&gt;internal audits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Anywhere you need to distinguish between:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;artifact exists
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;artifact is corroborated
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why OpenClaw-style agents feel different from plain chat
&lt;/h2&gt;

&lt;p&gt;The interesting part of OpenClaw isn’t that it chats.&lt;/p&gt;

&lt;p&gt;ChatGPT chats. Claude chats. Local Qwen chats.&lt;/p&gt;

&lt;p&gt;The useful part is when an agent can pull from multiple connected systems and cross-reference them in a way you can inspect.&lt;/p&gt;

&lt;p&gt;That’s the real upgrade.&lt;/p&gt;

&lt;p&gt;Here’s how I’d compare the approaches:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;What it’s actually good at&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenClaw evidence pipeline&lt;/td&gt;
&lt;td&gt;Cross-references email, calendar, files, and notes; maintains chronology; builds searchable evidence; keeps a human approval gate before anything external is sent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One-shot legal chatbot&lt;/td&gt;
&lt;td&gt;Fast answer generation; high risk of unsupported claims or false citations; weak provenance unless manually checked&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Manual case binder workflow&lt;/td&gt;
&lt;td&gt;Strong provenance control; slow retrieval; weak cross-referencing at scale; lots of manual reconstruction&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If I had to choose one for real legal-assist work, I’d take the OpenClaw pattern every time.&lt;/p&gt;

&lt;p&gt;Not because it’s magical.&lt;/p&gt;

&lt;p&gt;Because it respects evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you were building this yourself, here’s the architecture I’d use
&lt;/h2&gt;

&lt;p&gt;At a high level:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sources -&amp;gt; ingestion -&amp;gt; normalization -&amp;gt; chronology -&amp;gt; evidence scoring -&amp;gt; drafting -&amp;gt; human review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;More concretely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;email exports
calendar events
chat logs
photos
audio
notes
PDFs
        |
        v
ingest into a searchable workspace
        |
        v
normalize names, dates, timezones, languages
        |
        v
build timeline entries with source references
        |
        v
score each claim by support level
        |
        v
draft summaries / translations / letters
        |
        v
require human approval before anything leaves the system
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you want to model the evidence layer explicitly, even a simple schema helps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"claim"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Parent present at school pickup on 2024-03-14"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"support_level"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"corroborated"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"sources"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"photo"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/evidence/photos/pickup_2024_03_14.jpg"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gps"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/evidence/gps/watch_export_2024_03_14.json"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"calendar"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/evidence/calendar/march_2024.ics"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is way more useful than a blob of text that says “the parent appears involved.”&lt;/p&gt;

&lt;h2&gt;
  
  
  The trust model is simple: distrust it properly
&lt;/h2&gt;

&lt;p&gt;Can you trust an agent-heavy workflow?&lt;/p&gt;

&lt;p&gt;Only if you distrust it correctly.&lt;/p&gt;

&lt;p&gt;That’s the paradox.&lt;/p&gt;

&lt;p&gt;Even the good version has obvious failure modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;summaries flatten nuance&lt;/li&gt;
&lt;li&gt;translations miss tone&lt;/li&gt;
&lt;li&gt;extracted timelines can propagate a bad date forever&lt;/li&gt;
&lt;li&gt;citations can be wrong&lt;/li&gt;
&lt;li&gt;confidence can look higher than it should&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the safeguards matter more than the model.&lt;/p&gt;

&lt;p&gt;My baseline rules would be:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Keep source links attached to every claim
&lt;/h3&gt;

&lt;p&gt;If a timeline item exists, it should point back to the underlying artifact.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;claim -&amp;gt; source document -&amp;gt; exact message/photo/email/event
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Mark confidence tiers explicitly
&lt;/h3&gt;

&lt;p&gt;Do not blur these together:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;possible&lt;/li&gt;
&lt;li&gt;supported&lt;/li&gt;
&lt;li&gt;corroborated&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Treat legal citations as untrusted until verified
&lt;/h3&gt;

&lt;p&gt;If GPT-5, Claude, Qwen, or anything else gives you authority, verify it like you expect it might be wrong.&lt;/p&gt;

&lt;p&gt;Because sometimes it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boring problem that actually decides whether this works
&lt;/h2&gt;

&lt;p&gt;The biggest issue isn’t model IQ.&lt;/p&gt;

&lt;p&gt;It’s operations.&lt;/p&gt;

&lt;p&gt;The same Reddit threads that make OpenClaw look powerful also make it look fragile:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;breaking updates&lt;/li&gt;
&lt;li&gt;pinned versions&lt;/li&gt;
&lt;li&gt;speed vs quality tradeoffs&lt;/li&gt;
&lt;li&gt;expensive prompts&lt;/li&gt;
&lt;li&gt;workflows that become too annoying to run consistently&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That matters a lot if you’re indexing thousands of messages and revisiting the record for months.&lt;/p&gt;

&lt;p&gt;You need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;version discipline&lt;/li&gt;
&lt;li&gt;backups&lt;/li&gt;
&lt;li&gt;repeatable workflows&lt;/li&gt;
&lt;li&gt;stable connectors&lt;/li&gt;
&lt;li&gt;predictable LLM costs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one matters more than people admit.&lt;/p&gt;

&lt;p&gt;A legal-assist pipeline does a lot of repetitive work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;retrieval&lt;/li&gt;
&lt;li&gt;comparison&lt;/li&gt;
&lt;li&gt;translation&lt;/li&gt;
&lt;li&gt;redrafting&lt;/li&gt;
&lt;li&gt;timeline updates&lt;/li&gt;
&lt;li&gt;evidence re-checking&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If every pass feels like feeding a taxi meter, people start rationing analysis.&lt;/p&gt;

&lt;p&gt;That’s where per-token pricing gets weirdly destructive.&lt;/p&gt;

&lt;p&gt;You stop re-running checks.&lt;br&gt;
You skip useful comparisons.&lt;br&gt;
You avoid broad context windows.&lt;br&gt;
You hesitate to let agents do the boring but necessary passes.&lt;/p&gt;

&lt;p&gt;And then the workflow gets worse.&lt;/p&gt;

&lt;p&gt;For this kind of agent-heavy process, flat monthly pricing is just easier to operate.&lt;/p&gt;

&lt;p&gt;That’s one reason Standard Compute is interesting here. It gives you unlimited AI compute for a predictable monthly price, works as a drop-in OpenAI-compatible API, and is built for automations and agents rather than occasional chat use.&lt;/p&gt;

&lt;p&gt;If you’re wiring up legal-assist, compliance, or evidence-heavy workflows in OpenClaw, n8n, Make, Zapier, or your own stack, predictable cost matters a lot more than benchmark flexing.&lt;/p&gt;

&lt;p&gt;The useful question is not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Which model is smartest?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It’s:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Can this workflow run all month without me worrying about token burn?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s much closer to the real engineering problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical implementation sketch
&lt;/h2&gt;

&lt;p&gt;If I were prototyping this workflow, I’d keep the control flow boring on purpose.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. ingest artifacts&lt;/span&gt;
python ingest_email.py
python ingest_calendar.py
python ingest_photos.py
python ingest_chat_exports.py

&lt;span class="c"&gt;# 2. normalize metadata&lt;/span&gt;
python normalize_dates.py
python normalize_contacts.py
python detect_language.py

&lt;span class="c"&gt;# 3. build timeline&lt;/span&gt;
python build_chronology.py

&lt;span class="c"&gt;# 4. score evidence strength&lt;/span&gt;
python score_evidence.py

&lt;span class="c"&gt;# 5. generate drafts only&lt;/span&gt;
python draft_summary.py
python draft_bilingual_letter.py

&lt;span class="c"&gt;# 6. require manual approval&lt;/span&gt;
python queue_for_review.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And I’d make the approval boundary impossible to miss:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;outbound_message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;approved_by_human&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Blocked: external send requires explicit approval&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one guardrail is worth more than a better prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern worth stealing
&lt;/h2&gt;

&lt;p&gt;The best agent workflows usually look less like “replace the expert” and more like “give the system a finite job with hard evidence gates.”&lt;/p&gt;

&lt;p&gt;That maps perfectly to legal assist.&lt;/p&gt;

&lt;p&gt;A sane workflow looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Ingest records from email, calendar, files, notes, exports, and media&lt;/li&gt;
&lt;li&gt;Normalize dates, names, languages, and document types&lt;/li&gt;
&lt;li&gt;Build chronology with source references&lt;/li&gt;
&lt;li&gt;Score evidence strength from weak artifact to corroborated record&lt;/li&gt;
&lt;li&gt;Draft summaries or correspondence&lt;/li&gt;
&lt;li&gt;Flag weak claims, missing support, and unverified citations&lt;/li&gt;
&lt;li&gt;Require human review before anything is sent externally&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That’s the whole thing.&lt;/p&gt;

&lt;p&gt;No robot attorney fantasy required.&lt;/p&gt;

&lt;h2&gt;
  
  
  My take
&lt;/h2&gt;

&lt;p&gt;The breakthrough here is smaller than people want, but more useful than most demos.&lt;/p&gt;

&lt;p&gt;AI does not need to replace lawyers to be valuable.&lt;/p&gt;

&lt;p&gt;It just needs to help people handle evidence better:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;organize it&lt;/li&gt;
&lt;li&gt;cross-reference it&lt;/li&gt;
&lt;li&gt;translate it&lt;/li&gt;
&lt;li&gt;rank it&lt;/li&gt;
&lt;li&gt;draft from it&lt;/li&gt;
&lt;li&gt;keep humans in the approval loop&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s already a big deal.&lt;/p&gt;

&lt;p&gt;So my opinionated version is this:&lt;/p&gt;

&lt;p&gt;Legal automation gets useful the moment you stop asking for final answers and start building evidence pipelines with human review.&lt;/p&gt;

&lt;p&gt;Not answer engines.&lt;/p&gt;

&lt;p&gt;Evidence pipelines.&lt;/p&gt;

&lt;p&gt;That OpenClaw setup is the first legal-agent example I’ve seen that really respects that line.&lt;/p&gt;

&lt;p&gt;And for high-stakes workflows, that’s the line that matters.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>automation</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I thought my agent needed a huge soul.md until a 52-word file worked better</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Sat, 18 Jul 2026 01:11:59 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-thought-my-agent-needed-a-huge-soulmd-until-a-52-word-file-worked-better-l87</link>
      <guid>https://dev.to/lars_winstand/i-thought-my-agent-needed-a-huge-soulmd-until-a-52-word-file-worked-better-l87</guid>
      <description>&lt;p&gt;I used to think better agent behavior came from a bigger &lt;code&gt;soul.md&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;You know the file:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tone rules&lt;/li&gt;
&lt;li&gt;personality rules&lt;/li&gt;
&lt;li&gt;edge cases&lt;/li&gt;
&lt;li&gt;values&lt;/li&gt;
&lt;li&gt;backstory&lt;/li&gt;
&lt;li&gt;workflow preferences&lt;/li&gt;
&lt;li&gt;weird little reminders you swear matter&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It starts as “just a few notes” and ends up as a 1,500-word manifesto that gets injected into every run.&lt;/p&gt;

&lt;p&gt;I’ve built that version. It feels smart for a while.&lt;/p&gt;

&lt;p&gt;Then the agent gets more obedient and less useful.&lt;/p&gt;

&lt;p&gt;It starts protecting its persona better than your actual state.&lt;/p&gt;

&lt;p&gt;And after looking through a couple OpenClaw threads, I think the better pattern is much simpler:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tiny &lt;code&gt;soul.md&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;markdown save files for canonical state&lt;/li&gt;
&lt;li&gt;retrieval with something like LanceDB&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That setup is less glamorous than a giant prompt. It also works better.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comment that killed the giant-prompt myth for me
&lt;/h2&gt;

&lt;p&gt;In an r/openclaw thread about writing a soul, one commenter said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Have your AI write it. Also it barely matters. Agent and memory files matter.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s a rude sentence if you’ve spent hours polishing a persona file.&lt;/p&gt;

&lt;p&gt;It gets worse.&lt;/p&gt;

&lt;p&gt;Another user said their &lt;code&gt;soul.md&lt;/code&gt; was just 52 words, and they’d been working with that agent since January.&lt;/p&gt;

&lt;p&gt;Fifty-two words.&lt;/p&gt;

&lt;p&gt;That’s enough to define role and tone. Not enough to pretend the file is a database, CRM, runbook, and autobiography at the same time.&lt;/p&gt;

&lt;p&gt;That matches what I keep seeing in real agent setups:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;short persona&lt;/li&gt;
&lt;li&gt;deterministic saved state&lt;/li&gt;
&lt;li&gt;retrieval for relevant context&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Not the other way around.&lt;/p&gt;

&lt;h2&gt;
  
  
  The best example I found: an OpenClaw Dungeon Master
&lt;/h2&gt;

&lt;p&gt;The clearest proof came from another OpenClaw thread where someone built a Dungeon Master agent.&lt;/p&gt;

&lt;p&gt;What mattered wasn’t a better &lt;code&gt;soul.md&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It was architecture:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;structured directory layout for hard saves&lt;/li&gt;
&lt;li&gt;local LanceDB vector search as memory-core&lt;/li&gt;
&lt;li&gt;local markdown copies of the D&amp;amp;D 5e SRD&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s the important distinction.&lt;/p&gt;

&lt;p&gt;The agent wasn’t “remembering” because the prompt was poetic.&lt;/p&gt;

&lt;p&gt;It was remembering because state lived outside the prompt and retrieval pulled in only what was relevant.&lt;/p&gt;

&lt;p&gt;That’s a much better design for agents that have to survive real use.&lt;/p&gt;

&lt;p&gt;Whether your agent runs a campaign, triages support, handles Discord ops, or executes an &lt;code&gt;n8n&lt;/code&gt; workflow all day, continuity usually comes from memory architecture, not prompt theater.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why giant prompts underperform
&lt;/h2&gt;

&lt;p&gt;A bloated system prompt tries to solve every problem in one place:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;personality&lt;/li&gt;
&lt;li&gt;policy&lt;/li&gt;
&lt;li&gt;memory&lt;/li&gt;
&lt;li&gt;workflow rules&lt;/li&gt;
&lt;li&gt;examples&lt;/li&gt;
&lt;li&gt;special cases&lt;/li&gt;
&lt;li&gt;historical context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So every request drags around the same giant instruction block.&lt;/p&gt;

&lt;p&gt;That causes a few problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;more tokens every run&lt;/li&gt;
&lt;li&gt;more chances for instruction conflicts&lt;/li&gt;
&lt;li&gt;harder debugging&lt;/li&gt;
&lt;li&gt;more prompt dilution&lt;/li&gt;
&lt;li&gt;smaller models get mushy faster&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It’s like making a function depend on a global config object that contains your entire company wiki.&lt;/p&gt;

&lt;p&gt;Technically possible. Terrible to reason about.&lt;/p&gt;

&lt;h2&gt;
  
  
  A cleaner split: persona, state, retrieval
&lt;/h2&gt;

&lt;p&gt;Here’s the split I’d recommend.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. &lt;code&gt;soul.md&lt;/code&gt; handles identity
&lt;/h3&gt;

&lt;p&gt;Keep it short.&lt;/p&gt;

&lt;p&gt;Think 50 to 150 words.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;You are a pragmatic coding agent.
Be concise, specific, and honest about uncertainty.
Prefer shipping over theorizing.
Ask clarifying questions only when blocked.
Preserve existing conventions unless there is a strong reason to change them.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s enough.&lt;/p&gt;

&lt;p&gt;It defines role and behavior without pretending to store memory.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Markdown files handle canonical state
&lt;/h3&gt;

&lt;p&gt;Put facts here.&lt;/p&gt;

&lt;p&gt;Examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;campaign_state.md
customer_context.md
decisions_log.md
inventory.md
runbook.md
incident_notes.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are your source of truth.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# customer_context.md&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Customer: Acme Health
&lt;span class="p"&gt;-&lt;/span&gt; Plan: Enterprise
&lt;span class="p"&gt;-&lt;/span&gt; Primary integration: n8n
&lt;span class="p"&gt;-&lt;/span&gt; Known issue: intermittent Slack webhook retries
&lt;span class="p"&gt;-&lt;/span&gt; Last decision: keep retries at 3 before escalation
&lt;span class="p"&gt;-&lt;/span&gt; Escalation contact: maya@acme.example
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is much easier to inspect than burying the same facts inside a giant prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Retrieval handles fuzzy recall
&lt;/h3&gt;

&lt;p&gt;Use LanceDB or another retrieval layer when the agent needs relevant context, not all context.&lt;/p&gt;

&lt;p&gt;That means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;smaller prompts&lt;/li&gt;
&lt;li&gt;less repeated baggage&lt;/li&gt;
&lt;li&gt;better recall on demand&lt;/li&gt;
&lt;li&gt;easier debugging when memory goes weird&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What did we decide about Slack webhook retries for Acme Health?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s cleaner than injecting your entire customer history into every turn.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this looks like in practice
&lt;/h2&gt;

&lt;p&gt;A simple agent workspace might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;agent/
├── soul.md
├── state/
│   ├── customer_context.md
│   ├── decisions_log.md
│   └── runbook.md
├── memory/
│   └── lancedb/
└── tools/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And a request pipeline might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Load soul.md
2. Load the specific state files needed for this task
3. Query LanceDB for relevant memory chunks
4. Build a small prompt from those pieces
5. Run the model
6. Persist new facts back into markdown and memory
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a lot more maintainable than “append more instructions until the vibe improves.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Prompt caching helps, but it does not fix bad design
&lt;/h2&gt;

&lt;p&gt;This is where people get sloppy.&lt;/p&gt;

&lt;p&gt;Yes, prompt caching is useful.&lt;/p&gt;

&lt;p&gt;OpenAI supports prompt caching for long reused prefixes. Anthropic also offers prompt caching and shows major latency and cost reductions for repeated long prompts.&lt;/p&gt;

&lt;p&gt;That’s real.&lt;/p&gt;

&lt;p&gt;But caching does not solve the actual problems caused by giant prompts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;conflicting instructions&lt;/li&gt;
&lt;li&gt;over-scripted behavior&lt;/li&gt;
&lt;li&gt;poor retrieval design&lt;/li&gt;
&lt;li&gt;hard-to-debug failures&lt;/li&gt;
&lt;li&gt;smaller models choking on too much prefix&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cheaper bloat is still bloat.&lt;/p&gt;

&lt;p&gt;If your whole agent strategy depends on sending the same monster prompt forever, caching can reduce the bill and latency. It does not make the architecture elegant.&lt;/p&gt;

&lt;p&gt;For teams running lots of automations, that distinction matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters more in automations than in chat demos
&lt;/h2&gt;

&lt;p&gt;If you’re running one-off chats, prompt bloat is annoying.&lt;/p&gt;

&lt;p&gt;If you’re running agents inside:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;n8n&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Make&lt;/li&gt;
&lt;li&gt;Zapier&lt;/li&gt;
&lt;li&gt;OpenClaw&lt;/li&gt;
&lt;li&gt;custom workers&lt;/li&gt;
&lt;li&gt;Slack bots&lt;/li&gt;
&lt;li&gt;Discord bots&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;then prompt bloat becomes an operational problem.&lt;/p&gt;

&lt;p&gt;Every run pays for unnecessary context unless your provider offsets it. Every handoff becomes harder to reason about. Every failure turns into a forensic exercise inside a giant prefix.&lt;/p&gt;

&lt;p&gt;This is exactly why predictable compute matters.&lt;/p&gt;

&lt;p&gt;When you’re building agents that run all day, you want to optimize for architecture and reliability, not constantly wonder whether one more chunk of prompt text is worth the cost.&lt;/p&gt;

&lt;p&gt;That’s also why Standard Compute is an interesting fit for this kind of workload: it gives you an OpenAI-compatible API with flat monthly pricing, so you can run agent-heavy workflows without babysitting per-token spend. That makes it much easier to choose the right memory design instead of the cheapest prompt at every step.&lt;/p&gt;

&lt;h2&gt;
  
  
  My ranking after looking at the patterns
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;What it’s best at&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;LanceDB retrieval&lt;/td&gt;
&lt;td&gt;Semantic recall, RAG, agent memory, smaller prompts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Markdown hard-save files&lt;/td&gt;
&lt;td&gt;Canonical facts, deterministic state, easy inspection and versioning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long &lt;code&gt;soul.md&lt;/code&gt; / bloated system prompt&lt;/td&gt;
&lt;td&gt;Tone shaping, behavior nudging, but easy to overdo&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;My actual winners:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;best for continuity: markdown hard saves&lt;/li&gt;
&lt;li&gt;best for recall: LanceDB&lt;/li&gt;
&lt;li&gt;best for style: short &lt;code&gt;soul.md&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;easiest thing to waste time on: giant persona files&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If I had to keep only one, I’d keep the hard-save files.&lt;/p&gt;

&lt;p&gt;State beats self-mythology.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to put in each layer
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Keep &lt;code&gt;soul.md&lt;/code&gt; embarrassingly short
&lt;/h3&gt;

&lt;p&gt;Good contents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;role&lt;/li&gt;
&lt;li&gt;tone&lt;/li&gt;
&lt;li&gt;a few constraints&lt;/li&gt;
&lt;li&gt;one or two priorities&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Bad contents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;entire customer history&lt;/li&gt;
&lt;li&gt;every workflow edge case&lt;/li&gt;
&lt;li&gt;logs&lt;/li&gt;
&lt;li&gt;inventories&lt;/li&gt;
&lt;li&gt;project state&lt;/li&gt;
&lt;li&gt;emotional backstory the model does not need&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Put facts in markdown
&lt;/h3&gt;

&lt;p&gt;Use markdown files for things that should be true until explicitly changed.&lt;/p&gt;

&lt;p&gt;Examples:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# decisions_log.md&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; 2026-07-10: Keep Slack webhook retries at 3.
&lt;span class="p"&gt;-&lt;/span&gt; 2026-07-11: Escalate repeated failures to Maya.
&lt;span class="p"&gt;-&lt;/span&gt; 2026-07-12: Do not auto-close incidents without human review.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Put fuzzy memory in retrieval
&lt;/h3&gt;

&lt;p&gt;Use retrieval for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;prior conversations&lt;/li&gt;
&lt;li&gt;semantically related incidents&lt;/li&gt;
&lt;li&gt;similar customer issues&lt;/li&gt;
&lt;li&gt;docs that are too large to inject every time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s where LanceDB shines.&lt;/p&gt;

&lt;h2&gt;
  
  
  Smaller models benefit even more
&lt;/h2&gt;

&lt;p&gt;This matters a lot if you’re using cheaper or smaller models.&lt;/p&gt;

&lt;p&gt;Claude Opus 4.6 or GPT-5.4 can absorb some prompt abuse.&lt;/p&gt;

&lt;p&gt;Smaller Qwen or Llama variants usually can’t.&lt;/p&gt;

&lt;p&gt;A lot of “this model is bad” takes are really “this agent is dragging around too much prompt junk.”&lt;/p&gt;

&lt;p&gt;If you’re evaluating the best cheap model for agent workflows, test it with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tiny persona&lt;/li&gt;
&lt;li&gt;explicit state files&lt;/li&gt;
&lt;li&gt;retrieval-based memory&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then compare it to the same model under a 1,500-word system prompt.&lt;/p&gt;

&lt;p&gt;You may find the architecture was the bottleneck, not the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I’d debug an agent like this
&lt;/h2&gt;

&lt;p&gt;One reason I like this setup is that it gives you layers you can inspect.&lt;/p&gt;

&lt;p&gt;If the agent fails, check them in this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is canonical state wrong?&lt;/li&gt;
&lt;li&gt;Is retrieval pulling irrelevant chunks?&lt;/li&gt;
&lt;li&gt;Is the persona too restrictive or too vague?&lt;/li&gt;
&lt;li&gt;Did the workflow fail to persist new facts?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In OpenClaw specifically, the built-in commands help:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw status
openclaw status &lt;span class="nt"&gt;--all&lt;/span&gt;
openclaw status &lt;span class="nt"&gt;--deep&lt;/span&gt;
openclaw logs &lt;span class="nt"&gt;--follow&lt;/span&gt;
openclaw doctor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s a much better debugging loop than editing a giant prompt and hoping the vibe changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical takeaway
&lt;/h2&gt;

&lt;p&gt;The real question is not:&lt;/p&gt;

&lt;p&gt;“how detailed should my &lt;code&gt;soul.md&lt;/code&gt; be?”&lt;/p&gt;

&lt;p&gt;It’s:&lt;/p&gt;

&lt;p&gt;“where should memory actually live?”&lt;/p&gt;

&lt;p&gt;My answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;identity in &lt;code&gt;soul.md&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;truth in markdown&lt;/li&gt;
&lt;li&gt;recall in retrieval&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That gives you an agent that is easier to run, easier to debug, and easier to scale across automations.&lt;/p&gt;

&lt;p&gt;And if you’re running those automations continuously, flat-cost compute becomes a real advantage. Per-token billing pushes people toward weird prompt compromises. Predictable monthly pricing lets you optimize for what works.&lt;/p&gt;

&lt;p&gt;If you’re still tempted to write a 2,000-word &lt;code&gt;soul.md&lt;/code&gt;, try this first:&lt;/p&gt;

&lt;p&gt;Write 52 words.&lt;/p&gt;

&lt;p&gt;Then spend the rest of your effort on state and retrieval.&lt;/p&gt;

&lt;p&gt;That’s where the real memory usually comes from.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>devops</category>
      <category>productivity</category>
    </item>
    <item>
      <title>The first OpenClaw workflow I’d steal is a job agent that checks listings 2x a week and never auto-applies</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Fri, 17 Jul 2026 17:12:26 +0000</pubDate>
      <link>https://dev.to/lars_winstand/the-first-openclaw-workflow-id-steal-is-a-job-agent-that-checks-listings-2x-a-week-and-never-3cdf</link>
      <guid>https://dev.to/lars_winstand/the-first-openclaw-workflow-id-steal-is-a-job-agent-that-checks-listings-2x-a-week-and-never-3cdf</guid>
      <description>&lt;p&gt;I’ve seen a lot of agent demos that look incredible right up until you ask one boring question:&lt;/p&gt;

&lt;p&gt;“Would I trust this with something that actually matters?”&lt;/p&gt;

&lt;p&gt;Usually the answer is no.&lt;/p&gt;

&lt;p&gt;An agent ordering lunch is cute. An agent booking a flight is fine. An agent “running your life” is mostly a benchmark for how quickly a demo can drift into nonsense.&lt;/p&gt;

&lt;p&gt;But while digging through OpenClaw use cases, I found one workflow that felt immediately real.&lt;/p&gt;

&lt;p&gt;A user on r/openclaw shared a supervised job-search agent setup that got 103 upvotes because it solved an actual problem without pretending AI should do the whole thing for you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;run job search twice a week&lt;/li&gt;
&lt;li&gt;find fresh listings&lt;/li&gt;
&lt;li&gt;rank the top 5 matches&lt;/li&gt;
&lt;li&gt;tailor the resume for each role&lt;/li&gt;
&lt;li&gt;draft cover letters&lt;/li&gt;
&lt;li&gt;keep the final submit step manual&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last bullet is the whole point.&lt;/p&gt;

&lt;p&gt;The best job-search agent is not an auto-apply cannon.&lt;/p&gt;

&lt;p&gt;It’s a selective pipeline that watches the market, does the repetitive work, and hands a human a shortlist worth reviewing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Reddit setup was boring in the best way
&lt;/h2&gt;

&lt;p&gt;The original poster said they were a former data scientist, out of work for about 1 year after 10 years in the field. They ran OpenClaw on a Mac mini, gave it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a resume&lt;/li&gt;
&lt;li&gt;a GitHub profile&lt;/li&gt;
&lt;li&gt;a markdown file describing the roles they actually wanted&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then OpenClaw searched twice a week, picked the top 5 matches, tailored the resume, and drafted boilerplate cover letters.&lt;/p&gt;

&lt;p&gt;And crucially: it did not auto-submit.&lt;/p&gt;

&lt;p&gt;That’s what made it believable.&lt;/p&gt;

&lt;p&gt;A lot of job automation optimizes for volume. More tabs. More applications. More browser sessions. More “look how autonomous this is.”&lt;/p&gt;

&lt;p&gt;That sounds good until you’ve been on the hiring side.&lt;/p&gt;

&lt;p&gt;You can smell spray-and-pray applications instantly.&lt;/p&gt;

&lt;p&gt;This OpenClaw workflow did the opposite:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;narrow the target&lt;/li&gt;
&lt;li&gt;rank by fit&lt;/li&gt;
&lt;li&gt;rewrite only for plausible roles&lt;/li&gt;
&lt;li&gt;keep a human in the loop before anything irreversible happens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s not less sophisticated.&lt;/p&gt;

&lt;p&gt;That’s more sophisticated.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real advantage is recency, not autonomy
&lt;/h2&gt;

&lt;p&gt;The most useful part of an always-on job agent is not that it can click buttons.&lt;/p&gt;

&lt;p&gt;It’s that it can notice good roles early.&lt;/p&gt;

&lt;p&gt;That matters more than people admit.&lt;/p&gt;

&lt;p&gt;If you’ve job hunted seriously, you know the first few days after a posting goes live are often the best window. The Reddit poster said the eventual role was found 1 day after it was posted, and the process led to an offer in about 1 month.&lt;/p&gt;

&lt;p&gt;That tracks.&lt;/p&gt;

&lt;p&gt;Humans are bad at sustained vigilance.&lt;/p&gt;

&lt;p&gt;We’re good at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;interviews&lt;/li&gt;
&lt;li&gt;judgment&lt;/li&gt;
&lt;li&gt;deciding whether a role feels right&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We’re bad at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;checking 14 job sources every morning&lt;/li&gt;
&lt;li&gt;staying consistent for 6 weeks&lt;/li&gt;
&lt;li&gt;rewriting the same materials over and over without losing our minds&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s exactly where an agent helps.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why job search is a great agent use case
&lt;/h2&gt;

&lt;p&gt;This is one of the cleanest real-world agent workflows because the task is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;repetitive&lt;/li&gt;
&lt;li&gt;time-sensitive&lt;/li&gt;
&lt;li&gt;open-ended&lt;/li&gt;
&lt;li&gt;mostly text-based&lt;/li&gt;
&lt;li&gt;easy to supervise&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A good job-search agent can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;monitor target boards on a schedule&lt;/li&gt;
&lt;li&gt;score new roles against your criteria&lt;/li&gt;
&lt;li&gt;summarize why a role is or isn’t a fit&lt;/li&gt;
&lt;li&gt;tailor resume bullets to the job description&lt;/li&gt;
&lt;li&gt;draft a first-pass cover letter&lt;/li&gt;
&lt;li&gt;produce a review queue instead of a mess&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s not sci-fi.&lt;/p&gt;

&lt;p&gt;That’s admin work.&lt;/p&gt;

&lt;p&gt;And unlike a lot of flashy browser-agent demos, this doesn’t break the second one CSS selector changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  You probably don’t need browser automation for the most useful part
&lt;/h2&gt;

&lt;p&gt;This was the part I think more developers should pay attention to.&lt;/p&gt;

&lt;p&gt;A practical job-search agent does not need to start with LinkedIn scraping and headless browser gymnastics.&lt;/p&gt;

&lt;p&gt;A lot of the best value comes earlier in the pipeline: discovery, filtering, ranking, and drafting.&lt;/p&gt;

&lt;p&gt;For that, structured APIs beat brittle UI automation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with Greenhouse and Lever
&lt;/h2&gt;

&lt;p&gt;If I were building this, I’d start with Greenhouse and Lever immediately.&lt;/p&gt;

&lt;h3&gt;
  
  
  Greenhouse Job Board API
&lt;/h3&gt;

&lt;p&gt;Greenhouse exposes public listings as JSON and supports applications through an official endpoint.&lt;/p&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;GET https://boards-api.greenhouse.io/v1/boards/{board_token}/jobs?content=true
POST https://boards-api.greenhouse.io/v1/boards/{board_token}/jobs/{id}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Lever Postings API
&lt;/h3&gt;

&lt;p&gt;Lever also exposes published jobs through a REST API and supports programmatic workflows around listings.&lt;/p&gt;

&lt;p&gt;That means your architecture can look more like a normal data pipeline and less like a flaky browser bot.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;What it’s good at&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Greenhouse Job Board API&lt;/td&gt;
&lt;td&gt;Public JSON listings, official application endpoint, stable source for monitoring and draft prep&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lever Postings API&lt;/td&gt;
&lt;td&gt;Published job listings via REST API, structured discovery, realistic source for selective application workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Headless browser auto-apply on LinkedIn or Indeed&lt;/td&gt;
&lt;td&gt;Wider theoretical coverage, but brittle selectors, CAPTCHA issues, UI churn, and much higher risk&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That’s why I think the winning version of this workflow starts with Greenhouse and Lever, not Selenium heroics.&lt;/p&gt;

&lt;h2&gt;
  
  
  A simple architecture I’d actually build
&lt;/h2&gt;

&lt;p&gt;Here’s the version I’d trust enough to run for weeks.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;scheduler
  -&amp;gt; fetch job listings from Greenhouse + Lever
  -&amp;gt; normalize postings
  -&amp;gt; dedupe by company/title/url
  -&amp;gt; score against candidate profile
  -&amp;gt; shortlist top N
  -&amp;gt; generate tailored resume bullets
  -&amp;gt; draft cover letter
  -&amp;gt; send review packet to human
  -&amp;gt; human decides whether to apply
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you want to prototype this fast, the stack is pretty boring:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenClaw for the agent workflow&lt;/li&gt;
&lt;li&gt;cron or GitHub Actions for scheduling&lt;/li&gt;
&lt;li&gt;Python or TypeScript for fetch/normalize/scoring glue&lt;/li&gt;
&lt;li&gt;SQLite or Postgres for dedupe and history&lt;/li&gt;
&lt;li&gt;Claude, GPT-5, or both for ranking and drafting&lt;/li&gt;
&lt;li&gt;email, Slack, or Telegram for review delivery&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s a real system. Not a conference demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example: fetch Greenhouse jobs in Python
&lt;/h2&gt;

&lt;p&gt;A basic collector is trivial.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="n"&gt;BOARD_TOKEN&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your-company&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://boards-api.greenhouse.io/v1/boards/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;BOARD_TOKEN&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/jobs?content=true&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;jobs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jobs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;jobs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;location&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;location&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;updated_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;updated_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;absolute_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;absolute_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Normalize that output into your own schema and score against a candidate profile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example candidate profile as markdown
&lt;/h2&gt;

&lt;p&gt;This is the part most people skip, and it’s why their agent outputs garbage.&lt;/p&gt;

&lt;p&gt;Give the system a constrained profile.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Target Role Profile&lt;/span&gt;

&lt;span class="gu"&gt;## Role targets&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Senior Data Scientist
&lt;span class="p"&gt;-&lt;/span&gt; Applied ML Engineer
&lt;span class="p"&gt;-&lt;/span&gt; AI Engineer

&lt;span class="gu"&gt;## Strong preferences&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Remote in US
&lt;span class="p"&gt;-&lt;/span&gt; Series B+ or profitable company
&lt;span class="p"&gt;-&lt;/span&gt; Product teams shipping LLM features
&lt;span class="p"&gt;-&lt;/span&gt; Python-heavy stack

&lt;span class="gu"&gt;## Hard no&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Onsite only
&lt;span class="p"&gt;-&lt;/span&gt; Contract-only roles
&lt;span class="p"&gt;-&lt;/span&gt; Generic "AI evangelist" jobs
&lt;span class="p"&gt;-&lt;/span&gt; Roles requiring active security clearance

&lt;span class="gu"&gt;## Salary floor&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; $180k base

&lt;span class="gu"&gt;## Signals of strong fit&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Production ML systems
&lt;span class="p"&gt;-&lt;/span&gt; Evaluation pipelines
&lt;span class="p"&gt;-&lt;/span&gt; Agents or workflow automation
&lt;span class="p"&gt;-&lt;/span&gt; Strong writing / stakeholder communication
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one file will improve results more than most prompt tweaking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I would not fully automate final apply
&lt;/h2&gt;

&lt;p&gt;Because this is where “autonomous” becomes “reckless.”&lt;/p&gt;

&lt;p&gt;The original poster made the right call by keeping final submission manual.&lt;/p&gt;

&lt;p&gt;I would do the same for three reasons.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Job applications are social signals
&lt;/h3&gt;

&lt;p&gt;A resume can be technically aligned and still feel wrong.&lt;/p&gt;

&lt;p&gt;A cover letter can mention every keyword and still read like AI sludge.&lt;/p&gt;

&lt;p&gt;This is where humans catch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;weird tone&lt;/li&gt;
&lt;li&gt;overfitting to keywords&lt;/li&gt;
&lt;li&gt;incorrect claims&lt;/li&gt;
&lt;li&gt;missing context about why the company actually matters&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Browser auto-apply is brittle
&lt;/h3&gt;

&lt;p&gt;If your workflow depends on clicking through dynamic UI flows across LinkedIn, Indeed, and random ATS pages, you are signing up for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;broken selectors&lt;/li&gt;
&lt;li&gt;CAPTCHA fights&lt;/li&gt;
&lt;li&gt;anti-bot systems&lt;/li&gt;
&lt;li&gt;account restrictions&lt;/li&gt;
&lt;li&gt;constant maintenance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That can still be worth it in some cases, but it should not be your starting point.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The final click is the cheapest human step
&lt;/h3&gt;

&lt;p&gt;This is the key tradeoff.&lt;/p&gt;

&lt;p&gt;The expensive part is not clicking “Submit.”&lt;/p&gt;

&lt;p&gt;The expensive part is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;finding relevant jobs consistently&lt;/li&gt;
&lt;li&gt;reading them&lt;/li&gt;
&lt;li&gt;comparing them to your background&lt;/li&gt;
&lt;li&gt;rewriting resume bullets&lt;/li&gt;
&lt;li&gt;drafting customized materials&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Let the agent do that.&lt;/p&gt;

&lt;p&gt;Keep the last irreversible action human.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hidden villain is token anxiety
&lt;/h2&gt;

&lt;p&gt;This is where a lot of agent workflows quietly stop being practical.&lt;/p&gt;

&lt;p&gt;A supervised job-search agent sounds cheap until you count the actual loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;read new job descriptions&lt;/li&gt;
&lt;li&gt;compare each role against resume, GitHub, and role criteria&lt;/li&gt;
&lt;li&gt;score and rank candidates&lt;/li&gt;
&lt;li&gt;rewrite resume bullets for top matches&lt;/li&gt;
&lt;li&gt;draft cover letters&lt;/li&gt;
&lt;li&gt;repeat for weeks or months&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That burns tokens fast.&lt;/p&gt;

&lt;p&gt;Especially if you use frontier models for writing quality.&lt;/p&gt;

&lt;p&gt;And this is where per-token pricing starts changing behavior in bad ways.&lt;/p&gt;

&lt;p&gt;People don’t just spend less.&lt;/p&gt;

&lt;p&gt;They make the workflow worse.&lt;/p&gt;

&lt;p&gt;They shorten prompts.&lt;br&gt;
They skip useful passes.&lt;br&gt;
They stop re-running ranking.&lt;br&gt;
They avoid deeper tailoring.&lt;br&gt;
They under-automate the exact tasks that matter.&lt;/p&gt;

&lt;p&gt;That’s token anxiety in practice.&lt;/p&gt;

&lt;p&gt;Not just “my bill might be high.”&lt;/p&gt;

&lt;p&gt;More like: “I know this extra pass would improve output, but I’m going to avoid it because I can feel the meter running.”&lt;/p&gt;

&lt;p&gt;For agentic workflows, that’s poison.&lt;/p&gt;
&lt;h2&gt;
  
  
  This is why flat-rate compute makes more sense for agent workflows
&lt;/h2&gt;

&lt;p&gt;Job search is an unusually clear example of why predictable pricing matters.&lt;/p&gt;

&lt;p&gt;You don’t know if a search lasts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2 weeks&lt;/li&gt;
&lt;li&gt;2 months&lt;/li&gt;
&lt;li&gt;4 months&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And you don’t know how many jobs need evaluation before one is worth pursuing.&lt;/p&gt;

&lt;p&gt;That uncertainty is exactly why per-token billing feels bad for long-running automations.&lt;/p&gt;

&lt;p&gt;If you’re building agents that need to stay on, re-check sources, rewrite outputs, and iterate without someone watching a token dashboard, flat monthly pricing is just a better fit.&lt;/p&gt;

&lt;p&gt;That’s the appeal of Standard Compute.&lt;/p&gt;

&lt;p&gt;It’s a drop-in OpenAI API replacement with unlimited AI compute at a predictable monthly price, so you can run agent workflows without worrying that every extra ranking pass or resume rewrite is quietly inflating the bill.&lt;/p&gt;

&lt;p&gt;For this kind of use case, that matters a lot more than people think.&lt;/p&gt;

&lt;p&gt;Because the best version of the workflow is not the cheapest-looking prompt chain.&lt;/p&gt;

&lt;p&gt;It’s the one you’ll actually let run.&lt;/p&gt;
&lt;h2&gt;
  
  
  Which models I’d use for each step
&lt;/h2&gt;

&lt;p&gt;I would not use one model for everything.&lt;/p&gt;

&lt;p&gt;That’s the wrong optimization.&lt;/p&gt;

&lt;p&gt;Use the right model for the right stage.&lt;/p&gt;
&lt;h3&gt;
  
  
  My practical split
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Claude Sonnet for extraction, ranking, summarization, and requirement matching&lt;/li&gt;
&lt;li&gt;GPT-5 for structured scoring and rubric-based evaluation&lt;/li&gt;
&lt;li&gt;Claude Opus for final resume tailoring and cover-letter drafting when tone matters most&lt;/li&gt;
&lt;li&gt;local Llama or Qwen variants for private experiments or cheap pre-filtering&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The expensive model should only touch the shortlist.&lt;/p&gt;

&lt;p&gt;Don’t spend premium model budget on jobs that fail the first filter.&lt;/p&gt;

&lt;p&gt;That’s true whether you’re paying per token or routing across models behind a unified API.&lt;/p&gt;
&lt;h2&gt;
  
  
  A concrete scoring pipeline
&lt;/h2&gt;

&lt;p&gt;If I were implementing this, I’d separate filtering from writing.&lt;/p&gt;
&lt;h3&gt;
  
  
  Step 1: cheap scoring pass
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title_match"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"location_match"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"salary_match"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"domain_match"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"skills_match"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.85&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"overall_fit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.84&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"reasons"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Strong Python + ML systems overlap"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Remote role matches preference"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"LLM product experience is relevant"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"risks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Salary not explicitly listed"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Role leans more platform than research"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h3&gt;
  
  
  Step 2: shortlist only top 5 or top 10
&lt;/h3&gt;
&lt;h3&gt;
  
  
  Step 3: expensive writing pass
&lt;/h3&gt;

&lt;p&gt;Generate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;tailored resume bullets&lt;/li&gt;
&lt;li&gt;cover letter draft&lt;/li&gt;
&lt;li&gt;quick rationale for why this role made the cut&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That split keeps the workflow sane.&lt;/p&gt;
&lt;h2&gt;
  
  
  Example cron setup for a twice-weekly run
&lt;/h2&gt;

&lt;p&gt;If the original Reddit workflow ran twice a week, that’s a good default.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Every Monday and Thursday at 8:00 AM&lt;/span&gt;
0 8 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; 1,4 /usr/bin/python3 /opt/job-agent/run.py &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; /var/log/job-agent.log 2&amp;gt;&amp;amp;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s enough to catch fresh roles without creating noise fatigue.&lt;/p&gt;

&lt;p&gt;If the market is moving fast, run daily.&lt;/p&gt;

&lt;p&gt;If your criteria are narrow, twice a week is probably ideal.&lt;/p&gt;

&lt;h2&gt;
  
  
  The version I’d copy tomorrow
&lt;/h2&gt;

&lt;p&gt;If I were building this for myself, I’d do exactly this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;create a markdown candidate profile with real constraints&lt;/li&gt;
&lt;li&gt;ingest jobs from Greenhouse and Lever first&lt;/li&gt;
&lt;li&gt;schedule discovery daily or twice weekly&lt;/li&gt;
&lt;li&gt;dedupe and store posting history&lt;/li&gt;
&lt;li&gt;score all new jobs with a cheaper model&lt;/li&gt;
&lt;li&gt;shortlist top 5 or top 10&lt;/li&gt;
&lt;li&gt;use a stronger model for resume tailoring and cover-letter drafts&lt;/li&gt;
&lt;li&gt;send everything to a human review queue&lt;/li&gt;
&lt;li&gt;keep final submission manual&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is not a compromise.&lt;/p&gt;

&lt;p&gt;That is the product.&lt;/p&gt;

&lt;p&gt;The point is not to remove the human.&lt;/p&gt;

&lt;p&gt;The point is to remove:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;dead time&lt;/li&gt;
&lt;li&gt;missed postings&lt;/li&gt;
&lt;li&gt;repetitive rewriting&lt;/li&gt;
&lt;li&gt;the constant background stress of checking job boards manually&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The OpenClaw story worked because it respected that boundary.&lt;/p&gt;

&lt;p&gt;It used the agent for what agents are actually good at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;staying awake&lt;/li&gt;
&lt;li&gt;reading too much&lt;/li&gt;
&lt;li&gt;filtering chaos into a shortlist&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything after that still belonged to a person.&lt;/p&gt;

&lt;p&gt;And honestly, that’s the first OpenClaw workflow I’ve seen that I’d steal immediately.&lt;/p&gt;

&lt;p&gt;If you’re building agent workflows like this, the next thing I’d fix is pricing. Once an automation is useful, it tends to run more often, touch more context, and do more rewriting than you expected. That’s where a flat-rate API setup becomes less of a nice-to-have and more of an architectural decision.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>automation</category>
      <category>openclaw</category>
    </item>
    <item>
      <title>I knew agents were getting real when someone kept an OpenClaw bird-card workflow running every hour for their kids</title>
      <dc:creator>Lars Winstand</dc:creator>
      <pubDate>Fri, 17 Jul 2026 09:13:56 +0000</pubDate>
      <link>https://dev.to/lars_winstand/i-knew-agents-were-getting-real-when-someone-kept-an-openclaw-bird-card-workflow-running-every-hour-3fk9</link>
      <guid>https://dev.to/lars_winstand/i-knew-agents-were-getting-real-when-someone-kept-an-openclaw-bird-card-workflow-running-every-hour-3fk9</guid>
      <description>&lt;p&gt;I’ve seen a lot of "agents are here" posts lately.&lt;/p&gt;

&lt;p&gt;Most of them use the same formula:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;polished demo
n- cherry-picked workflow&lt;/li&gt;
&lt;li&gt;one perfect run&lt;/li&gt;
&lt;li&gt;zero discussion of what happens after day 3&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The thing that finally convinced me agents are becoming usable was much smaller.&lt;/p&gt;

&lt;p&gt;I found a thread on r/openclaw where someone in Las Vegas had OpenClaw pull BirdWeather data every hour and generate Garbage Pail Kids / Pokémon-style bird cards for their kids.&lt;/p&gt;

&lt;p&gt;That’s it. No enterprise ROI deck. No fake productivity theater. Just an hourly workflow that stayed alive because the family liked it.&lt;/p&gt;

&lt;p&gt;That is a better signal than most benchmark charts.&lt;/p&gt;

&lt;p&gt;If a weird automation keeps running when nobody is forcing it to, you’re looking at something real.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this example matters more than another agent demo
&lt;/h2&gt;

&lt;p&gt;A lot of agent demos are optimized to look impressive for 90 seconds.&lt;/p&gt;

&lt;p&gt;That’s not the same as being usable.&lt;/p&gt;

&lt;p&gt;The real test is uglier:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Will someone leave this thing running every hour for weeks when there is no boss, no KPI, and no meeting attached to it?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s why the BirdWeather example stuck with me.&lt;/p&gt;

&lt;p&gt;It’s small enough to be honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern is actually very technical
&lt;/h2&gt;

&lt;p&gt;Under the cute use case, this is a serious recurring workflow pattern:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Poll an external data source on a schedule&lt;/li&gt;
&lt;li&gt;Detect changes or new events&lt;/li&gt;
&lt;li&gt;Feed those events into an LLM&lt;/li&gt;
&lt;li&gt;Generate structured output&lt;/li&gt;
&lt;li&gt;Deliver it somewhere people already are&lt;/li&gt;
&lt;li&gt;Repeat forever&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is the same pattern behind a lot of useful agent systems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;job monitoring&lt;/li&gt;
&lt;li&gt;lead enrichment&lt;/li&gt;
&lt;li&gt;inbox triage&lt;/li&gt;
&lt;li&gt;support classification&lt;/li&gt;
&lt;li&gt;listing alerts&lt;/li&gt;
&lt;li&gt;research digests&lt;/li&gt;
&lt;li&gt;compliance checks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The bird cards are just the friendlier version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why BirdWeather is good agent input
&lt;/h2&gt;

&lt;p&gt;BirdWeather is not just a cute gadget.&lt;/p&gt;

&lt;p&gt;Its PUC device continuously records outdoor audio, uploads soundscapes, and BirdWeather says the audio is analyzed with BirdNET for automatic bird detection.&lt;/p&gt;

&lt;p&gt;That means the input stream is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;recurring&lt;/li&gt;
&lt;li&gt;messy&lt;/li&gt;
&lt;li&gt;time-based&lt;/li&gt;
&lt;li&gt;event-driven&lt;/li&gt;
&lt;li&gt;different every hour&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is perfect agent fuel.&lt;/p&gt;

&lt;p&gt;A static prompt is easy.&lt;br&gt;
A living input stream is where systems get interesting.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why OpenClaw fits this kind of workflow
&lt;/h2&gt;

&lt;p&gt;OpenClaw makes more sense when you stop thinking of agents as a browser tab and start thinking of them as a long-running process with memory.&lt;/p&gt;

&lt;p&gt;That’s the right shape for hobby automations and sidekick workflows.&lt;/p&gt;

&lt;p&gt;You do not want another dashboard for this kind of thing.&lt;br&gt;
You want something that can sit in the background, remember context, and send updates into a channel you already use.&lt;/p&gt;

&lt;p&gt;That might be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Telegram&lt;/li&gt;
&lt;li&gt;Discord&lt;/li&gt;
&lt;li&gt;Slack&lt;/li&gt;
&lt;li&gt;WhatsApp&lt;/li&gt;
&lt;li&gt;Signal&lt;/li&gt;
&lt;li&gt;Google Chat&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That changes the feel of the system.&lt;/p&gt;

&lt;p&gt;Now it is not "go open the AI tool."&lt;br&gt;
It is "the agent is around when I need it."&lt;/p&gt;
&lt;h2&gt;
  
  
  Minimal OpenClaw shape
&lt;/h2&gt;

&lt;p&gt;If you are comfortable in a terminal, the setup shape is pretty understandable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw agents add birds &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--workspace&lt;/span&gt; ~/.openclaw/workspace-birds &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--bind&lt;/span&gt; telegram:&lt;span class="k"&gt;*&lt;/span&gt;

openclaw status &lt;span class="nt"&gt;--all&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s not consumer software, obviously.&lt;/p&gt;

&lt;p&gt;You still need to be okay with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;self-hosting&lt;/li&gt;
&lt;li&gt;runtime versions&lt;/li&gt;
&lt;li&gt;background processes&lt;/li&gt;
&lt;li&gt;credentials&lt;/li&gt;
&lt;li&gt;occasional breakage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But that is also why this use case matters. Even with setup friction, people are still keeping these workflows alive.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the workflow probably looks like
&lt;/h2&gt;

&lt;p&gt;You could build the bird-card version as a pretty standard loop.&lt;/p&gt;

&lt;h3&gt;
  
  
  1) Poll BirdWeather on a schedule
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# cron: every hour&lt;/span&gt;
0 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; /usr/local/bin/node /opt/birds/fetch.js
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2) Normalize new sightings
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sightings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getBirdWeatherSightings&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fresh&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;sightings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;timestamp&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;lastRun&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;seenIds&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3) Prompt the model with a tight output contract
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`
Create a kid-friendly bird trading card.

Bird: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;bird&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;commonName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;
Scientific name: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;bird&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;scientificName&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;
Location: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;bird&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;location&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;
Observed at: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;bird&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;

Return JSON with:
- title
- tagline
- powers
- weakness
- rarity
- fun_fact
- art_prompt
`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4) Generate and deliver
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;card&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-5.4-mini&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendTelegramMessage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;formatCard&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;card&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is not exotic engineering.&lt;/p&gt;

&lt;p&gt;That is exactly why it matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real problem is not tooling. It is billing psychology.
&lt;/h2&gt;

&lt;p&gt;This is where most always-on agent setups get weird.&lt;/p&gt;

&lt;p&gt;The workflow side has gotten much better.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Product&lt;/th&gt;
&lt;th&gt;What the pricing encourages&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenClaw&lt;/td&gt;
&lt;td&gt;Self-hosted experimentation and persistent multi-channel agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;n8n&lt;/td&gt;
&lt;td&gt;Recurring automations because billing is tied to workflow execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Zapier&lt;/td&gt;
&lt;td&gt;Lightweight hosted automation, but with hard task limits on lower tiers&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The model side is still where people get nervous.&lt;/p&gt;

&lt;p&gt;If your workflow runs every hour, 24/7, small token costs stop feeling small.&lt;/p&gt;

&lt;p&gt;That is especially true for automations that are useful-but-not-business-critical.&lt;/p&gt;

&lt;p&gt;Examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;bird cards for your kids&lt;/li&gt;
&lt;li&gt;a Telegram travel helper&lt;/li&gt;
&lt;li&gt;a job-search watcher&lt;/li&gt;
&lt;li&gt;a listings monitor for niche gear&lt;/li&gt;
&lt;li&gt;a personal research digest&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are exactly the workflows that should be allowed to run freely.&lt;/p&gt;

&lt;p&gt;Instead, per-token billing makes you do mental math all week.&lt;/p&gt;

&lt;p&gt;That kills experimentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why per-token pricing breaks good habits
&lt;/h2&gt;

&lt;p&gt;There is a difference between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"this automation is worth running"&lt;/li&gt;
&lt;li&gt;"this automation is worth monitoring for cost every day"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Developers will tolerate a lot:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;rough docs&lt;/li&gt;
&lt;li&gt;Node version nonsense&lt;/li&gt;
&lt;li&gt;self-hosting setup&lt;/li&gt;
&lt;li&gt;occasional agent failures&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What they hate is uncertainty.&lt;/p&gt;

&lt;p&gt;If an hourly workflow might quietly become an expensive hobby, people turn it off early.&lt;/p&gt;

&lt;p&gt;That is a bad outcome because recurring agents only become valuable after they survive long enough to become routine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same pattern shows up in more serious workflows
&lt;/h2&gt;

&lt;p&gt;The bird-card use case sounds whimsical, but the architecture is the same as more obviously practical agent systems.&lt;/p&gt;

&lt;h3&gt;
  
  
  Example: job-search agent
&lt;/h3&gt;

&lt;p&gt;A job-search agent can:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;poll listings every hour&lt;/li&gt;
&lt;li&gt;compare them against your resume and preferences&lt;/li&gt;
&lt;li&gt;filter out junk&lt;/li&gt;
&lt;li&gt;summarize the good ones&lt;/li&gt;
&lt;li&gt;send them to Telegram or Slack&lt;/li&gt;
&lt;li&gt;keep state across runs&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is the same loop.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;jobs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetchNewJobs&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;matches&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;rankJobsAgainstProfile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;jobs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;candidateProfile&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;top&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;matches&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;j&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;j&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.82&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendTelegramDigest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;top&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same idea, different stakes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Example: marketplace watcher
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;listings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetchEbayListings&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;vintage nikon lens&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;scored&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;scoreDeals&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;listings&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;preferences&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;alerts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;scored&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;dealScore&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;90&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendSignalAlert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;alerts&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Again: same loop.&lt;/p&gt;

&lt;p&gt;This is what usable agents look like in practice.&lt;br&gt;
Not one giant autonomous worker.&lt;br&gt;
A lot of small loops that keep earning the right to stay on.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I think this says about agent maturity
&lt;/h2&gt;

&lt;p&gt;No, a bird-card automation does not prove agents are mainstream.&lt;/p&gt;

&lt;p&gt;OpenClaw is still rough in places.&lt;br&gt;
Self-hosted agent stacks still break.&lt;br&gt;
Security and reliability still matter a lot.&lt;br&gt;
Most of this is still too fiddly for non-technical users.&lt;/p&gt;

&lt;p&gt;But it proves something more useful:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;agent tooling has crossed the line where people with niche interests will keep it running anyway&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a big milestone.&lt;/p&gt;

&lt;p&gt;Mainstream software usually starts there.&lt;br&gt;
Not with universal adoption.&lt;br&gt;
With a weird group of users who cannot imagine turning it off.&lt;/p&gt;
&lt;h2&gt;
  
  
  Practical takeaway for developers building always-on agents
&lt;/h2&gt;

&lt;p&gt;If you are building recurring LLM workflows, optimize for these things first:&lt;/p&gt;
&lt;h3&gt;
  
  
  1) Make the loop durable
&lt;/h3&gt;

&lt;p&gt;Prefer boring reliability over flashy autonomy.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;idempotency&lt;/li&gt;
&lt;li&gt;deduping&lt;/li&gt;
&lt;li&gt;state persistence&lt;/li&gt;
&lt;li&gt;alerting&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  2) Deliver into an existing channel
&lt;/h3&gt;

&lt;p&gt;Do not make users babysit another dashboard.&lt;/p&gt;

&lt;p&gt;Push results into:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Telegram&lt;/li&gt;
&lt;li&gt;Slack&lt;/li&gt;
&lt;li&gt;Discord&lt;/li&gt;
&lt;li&gt;email&lt;/li&gt;
&lt;li&gt;SMS&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  3) Keep outputs structured
&lt;/h3&gt;

&lt;p&gt;Use JSON or strict schemas whenever possible.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Cardinal Clash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"rarity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Rare"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"powers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"Dawn Chorus"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Seed Swipe"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"weakness"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Window reflections"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4) Design for recurring economics
&lt;/h3&gt;

&lt;p&gt;This is the one people skip.&lt;/p&gt;

&lt;p&gt;A workflow that runs once a day is one thing.&lt;br&gt;
A workflow that runs every 15 minutes forever is a completely different cost model.&lt;/p&gt;

&lt;p&gt;If your pricing model punishes background usage, users will never let the system settle into habit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why flat-rate API access matters here
&lt;/h2&gt;

&lt;p&gt;This is the part that feels under-discussed.&lt;/p&gt;

&lt;p&gt;Always-on agents need predictable economics more than they need one more benchmark win.&lt;/p&gt;

&lt;p&gt;If you are running OpenClaw, n8n, Make, Zapier, or your own agent stack, flat-rate API access changes the decision.&lt;/p&gt;

&lt;p&gt;Instead of asking:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;should I let this run all week?&lt;/li&gt;
&lt;li&gt;how many tokens did this burn?&lt;/li&gt;
&lt;li&gt;is this fun automation secretly expensive?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You get to ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;does this deserve a place in my stack?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is a much healthier way to build.&lt;/p&gt;

&lt;p&gt;This is also why Standard Compute is an interesting fit for developers building recurring automations. It is an OpenAI-compatible API endpoint, so you can use existing SDKs and clients, but the pricing model is flat monthly instead of per-token. For always-on agents, that removes a lot of the low-grade cost anxiety that makes people shut down good workflows too early.&lt;/p&gt;

&lt;p&gt;If your agent is polling every hour, summarizing, routing, generating, and replying across the day, predictable cost is not a nice-to-have. It changes what you are willing to leave running.&lt;/p&gt;

&lt;h2&gt;
  
  
  My take
&lt;/h2&gt;

&lt;p&gt;The next breakout agent workflow probably will not start in a boardroom.&lt;/p&gt;

&lt;p&gt;It will be something slightly weird and extremely sticky:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;bird cards from BirdWeather&lt;/li&gt;
&lt;li&gt;a Buy Nothing digest bot&lt;/li&gt;
&lt;li&gt;a marketplace watcher&lt;/li&gt;
&lt;li&gt;a family logistics sidekick&lt;/li&gt;
&lt;li&gt;a travel helper in Telegram&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The common thread is not "maximum intelligence."&lt;br&gt;
It is "this became part of someone’s week."&lt;/p&gt;

&lt;p&gt;That only happens when three things are true:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the workflow is easy enough to keep alive&lt;/li&gt;
&lt;li&gt;the delivery channel fits daily life&lt;/li&gt;
&lt;li&gt;the cost model does not punish curiosity&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is why the OpenClaw bird-card story matters.&lt;/p&gt;

&lt;p&gt;Not because it is big.&lt;br&gt;
Because nobody wanted to turn it off.&lt;/p&gt;

&lt;p&gt;And that is usually how real software starts.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>automation</category>
      <category>openclaw</category>
    </item>
  </channel>
</rss>
