<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aria</title>
    <description>The latest articles on DEV Community by Aria (@ariaxhan).</description>
    <link>https://dev.to/ariaxhan</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2349576%2F5f601738-840d-422e-bfb4-ccff83763ef0.jpeg</url>
      <title>DEV Community: Aria</title>
      <link>https://dev.to/ariaxhan</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ariaxhan"/>
    <language>en</language>
    <item>
      <title>Stop Writing Markdown. Start Writing Memory.</title>
      <dc:creator>Aria</dc:creator>
      <pubDate>Wed, 07 Oct 2026 08:15:34 +0000</pubDate>
      <link>https://dev.to/ariaxhan/stop-writing-markdown-start-writing-memory-2f6g</link>
      <guid>https://dev.to/ariaxhan/stop-writing-markdown-start-writing-memory-2f6g</guid>
      <description>&lt;h1&gt;
  
  
  Stop Writing Markdown. Start Writing Memory.
&lt;/h1&gt;

&lt;p&gt;Since I've been coding with AI, it's always been one big blob of markdown files. It's the default in all the agentic coding platforms for "planning" mode, and widely accepted as the canonical way to plan and execute a coding task with agentic AI.&lt;/p&gt;

&lt;p&gt;Research summaries. Architecture plans. Debug traces. Session notes. Feature requirement distillations. Each one dutifully generated by an AI agent, each one formatted for human consumption, each one completely unqueryable by the very agents that created them.&lt;/p&gt;

&lt;p&gt;We've settled for a system where machines talk to machines through human-readable documents. Like passing notes in class by printing them first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Markdown Problem
&lt;/h2&gt;

&lt;p&gt;Here's what happens when you use AI agents to code:&lt;/p&gt;

&lt;p&gt;You ask an agent to plan a feature. It writes feature-plan.md. Implementation happens. The plan sits there, never again referenced, slowly drifting from reality. By week three, it's archaeological artifact.&lt;/p&gt;

&lt;p&gt;This is the default behavior of every AI coding assistant. Generate markdown. Pile it up. Hope someone reads it.&lt;/p&gt;

&lt;p&gt;The fundamental problem: &lt;strong&gt;markdown is a human-readable format being used for machine-to-machine communication.&lt;/strong&gt; Your agents generate structured knowledge, then immediately flatten it into prose that only humans can parse efficiently.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Inversion
&lt;/h2&gt;

&lt;p&gt;What if you use a database as a communication layer?&lt;/p&gt;

&lt;p&gt;Not storage you dump things into. Not a backup system. The actual protocol agents use to coordinate.&lt;/p&gt;

&lt;p&gt;I rebuilt my agent workflow around a single SQLite file. Three tables: context, learnings, and errors. No markdown generation unless a human explicitly needs to read something.&lt;/p&gt;

&lt;p&gt;I call it AgentDB. It changed everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Startup Hook
&lt;/h2&gt;

&lt;p&gt;Before diving into the schema, here's what makes this actually work: the session startup hook.&lt;/p&gt;

&lt;p&gt;When I open Claude Code, before I type anything, a hook runs. It reads from the database, from git, from the file system (whatever I've configured) and injects the result directly into the session context. It's unique to each folder/repo.&lt;/p&gt;

&lt;p&gt;Here's a sample:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;## Git State&lt;/span&gt;
Branch: feat/streaming-responses
Changes: M src/chat/stream.ts, M src/api/completions.ts
Recent:
 a3f7c21 feat&lt;span class="o"&gt;(&lt;/span&gt;chat&lt;span class="o"&gt;)&lt;/span&gt;: SSE streaming &lt;span class="k"&gt;for &lt;/span&gt;long responses
 8b2e4d9 fix&lt;span class="o"&gt;(&lt;/span&gt;context&lt;span class="o"&gt;)&lt;/span&gt;: token count before truncation
&lt;span class="c"&gt;## Active Contracts&lt;/span&gt;
| ID | Goal | Status |
| CR-031 | Streaming responses &lt;span class="k"&gt;for &lt;/span&gt;messages &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;500 tokens | in_progress |
| CR-028 | Context window management &lt;span class="k"&gt;for &lt;/span&gt;long threads | blocked |
&lt;span class="c"&gt;## Recent Learnings&lt;/span&gt;
| Category | Summary |
| failure | ReadableStream must be returned, not piped |
| pattern | Chunked transfer encoding requires explicit headers |
| gotcha | Vercel edge functions have 25s &lt;span class="nb"&gt;timeout&lt;/span&gt;, not 30s |
&lt;span class="c"&gt;## Recent Errors&lt;/span&gt;
| Tool | Error |
| Edit | File not found: src/old-path.ts |
| Bash | npm &lt;span class="nb"&gt;test exit &lt;/span&gt;code 1 |
&lt;span class="c"&gt;## Active Agents&lt;/span&gt;
- steady-pulse &lt;span class="o"&gt;(&lt;/span&gt;branch: feat/streaming-responses&lt;span class="o"&gt;)&lt;/span&gt;
- quick-spark &lt;span class="o"&gt;(&lt;/span&gt;branch: main&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent sees this before I say a word. It knows what changed recently. It knows what patterns I've discovered. It knows what errors occurred. It knows what I was working on.&lt;/p&gt;

&lt;p&gt;This is ambient context. No retrieval step. No "let me check my notes." The context is present before you ask.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The hook can read from anything:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The database (learnings, checkpoints, contracts, errors)&lt;/li&gt;
&lt;li&gt;Git (branch, commits, file changes)&lt;/li&gt;
&lt;li&gt;File system (folder structure, counts, specific files)&lt;/li&gt;
&lt;li&gt;External APIs (if you want)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The database is the storage layer. The startup hook is the delivery layer. Together, they create ambient context that survives sessions without manual effort.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Three Tables
&lt;/h2&gt;

&lt;p&gt;Now let's look at what gets stored.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;context: The Communication Protocol&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This replaced every "Hey here's what I found" markdown file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- CONTEXT: Work state (ephemeral per-contract)&lt;/span&gt;
&lt;span class="c1"&gt;-- Types: contract, checkpoint, handoff, verdict&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
 &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="n"&gt;ts&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;strftime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'%Y-%m-%dT%H:%M:%fZ'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'now'&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
 &lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;CHECK&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'contract'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'checkpoint'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'handoff'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'verdict'&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
 &lt;span class="n"&gt;contract_id&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;links&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="n"&gt;entries&lt;/span&gt; &lt;span class="k"&gt;to&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;contract&lt;/span&gt;
 &lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;which&lt;/span&gt; &lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="n"&gt;wrote&lt;/span&gt; &lt;span class="n"&gt;this&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;orchestrator&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;surgeon&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;adversary&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
 &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;JSON&lt;/span&gt; &lt;span class="nb"&gt;blob&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the orchestrator assigns work, it writes a contract. The surgeon reads that, writes checkpoints as it works. The adversary reads the contract, checkpoints, and the implementation to write a verdict. All context can easily be transferred to a fresh conversation with handoffs.&lt;/p&gt;

&lt;p&gt;All through the database. All queryable. All with typed structure.&lt;/p&gt;

&lt;p&gt;No copy-paste relay. No "let me summarize what the previous agent said." Each agent queries what it needs and writes what the next one will need.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;learnings: Knowledge That Compounds&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is where the markdown graveyard problem gets solved.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- LEARNINGS: Cross-session memory (survives forever)&lt;/span&gt;
&lt;span class="c1"&gt;-- Read these at session start to avoid repeating mistakes&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;learnings&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
 &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="n"&gt;ts&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;strftime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'%Y-%m-%dT%H:%M:%fZ'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'now'&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
 &lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;CHECK&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'failure'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'pattern'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'gotcha'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'preference'&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
 &lt;span class="n"&gt;insight&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="n"&gt;evidence&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="k"&gt;domain&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;g&lt;/span&gt;&lt;span class="p"&gt;.,&lt;/span&gt; &lt;span class="s1"&gt;'auth'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'database'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'frontend'&lt;/span&gt;
 &lt;span class="n"&gt;hit_count&lt;/span&gt; &lt;span class="nb"&gt;INTEGER&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="n"&gt;last_hit&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When an agent discovers something (API returns 500 on expired tokens; this library doesn't handle concurrent requests; always validate before calling that endpoint) it writes a learning. Typed. Categorized. Queryable.&lt;/p&gt;

&lt;p&gt;Recent failures surface automatically at session start. Patterns with high hit counts stay visible. Gotchas resurface before you hit them again.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;errors: Automatic Failure Capture&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This one is different. Agents don't write to it directly; hooks do.&lt;/p&gt;

&lt;p&gt;I can assign Claude Code the most complex, architecture-spanning task, and the biggest issues won't be the logic. They'll be: figuring out how to cd into the right directory, reading a file that moved, calling an MCP tool with the wrong parameters.&lt;/p&gt;

&lt;p&gt;These tool errors seem harmless. They compound. One failed Edit leads to a retry, which leads to a different approach, which leads to confusion about what state the file is in. An hour later, you're debugging the debug session.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- ERRORS: Automatic capture of failures&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
 &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="nb"&gt;INTEGER&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt; &lt;span class="n"&gt;AUTOINCREMENT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="n"&gt;ts&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;strftime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'%Y-%m-%dT%H:%M:%fZ'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="s1"&gt;'now'&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
 &lt;span class="n"&gt;tool&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="n"&gt;error&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="n"&gt;file&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
 &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This table captures tool failures automatically via hooks. Edit can't find a file? Logged. Bash returns non-zero? Logged. MCP server returns an error? Logged. The agent doesn't decide whether to record it. The system does.&lt;/p&gt;

&lt;p&gt;Next session, the startup hook surfaces recent errors: "Edit failed on src/old-path.ts." The agent immediately knows that file moved or was deleted. It doesn't waste twenty minutes trying the same path.&lt;/p&gt;

&lt;p&gt;Learnings are intentional: "I discovered this pattern." Errors are automatic: "this tool call broke." Both matter. Only one requires the agent to remember to write it down.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Agent Lifecycle
&lt;/h2&gt;

&lt;p&gt;With the schema in place, here's how agents actually use it. I made a lightweight Python CLI so they don't have to deal with SQL syntax.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Session start:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;agentdb&lt;/span&gt; &lt;span class="n"&gt;read&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent gets: recent failures to avoid, active patterns, the last checkpoint, any active contract, recent errors. One command. Everything needed to resume.&lt;/p&gt;

&lt;p&gt;During work, when something is learned:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;agentdb&lt;/span&gt; &lt;span class="n"&gt;learn&lt;/span&gt; &lt;span class="n"&gt;failure&lt;/span&gt; &lt;span class="nv"&gt;"Stripe webhook returns 200 but event.verified is false"&lt;/span&gt; &lt;span class="nv"&gt;"found in logs"&lt;/span&gt;
&lt;span class="n"&gt;agentdb&lt;/span&gt; &lt;span class="n"&gt;learn&lt;/span&gt; &lt;span class="n"&gt;pattern&lt;/span&gt; &lt;span class="nv"&gt;"always check verified flag explicitly"&lt;/span&gt; &lt;span class="nv"&gt;"fixed 3 bugs"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Session end:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;agentdb write-end '{"did":"implemented webhook handler","next":"add retry logic","blocked":""}'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This writes a checkpoint. Tomorrow, &lt;code&gt;read-start&lt;/code&gt; shows exactly where things left off.&lt;/p&gt;

&lt;p&gt;Before complex work:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight make"&gt;&lt;code&gt;&lt;span class="nl"&gt;agentdb contract '{"goal"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="nf"&gt;"timeout message after 5s"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nf"&gt;"scope":["src/auth/login.ts"]&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nf"&gt;"constraints":["no new deps"]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You don't run these commands. The agents do. The hooks are non-negotiable: every artifact reads on start, writes on end.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Point
&lt;/h2&gt;

&lt;p&gt;Here's the bigger insight:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Representation is the bottleneck.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not intelligence. Not scale. Not model size. How you structure information for machine consumption determines what machines can do with it.&lt;/p&gt;

&lt;p&gt;Markdown is optimized for human eyes. Great for documentation you'll read. Terrible for knowledge agents need to query.&lt;/p&gt;

&lt;p&gt;SQL is optimized for structured retrieval. Terrible for prose. Perfect for typed knowledge with categories, timestamps, hit counts, relationships.&lt;/p&gt;

&lt;p&gt;The endless markdown files weren't a storage problem. They were a representation problem. Information structured for the wrong consumer.&lt;/p&gt;

&lt;p&gt;When I switched to SQLite, I didn't add capabilities. I removed friction. Agents could suddenly query exactly what they needed instead of scanning documents hoping to find relevant passages.&lt;/p&gt;

&lt;p&gt;This generalizes beyond agent coordination. Every time you're tempted to generate a markdown report, ask: who consumes this? If it's another machine, the answer probably isn't prose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Goes
&lt;/h2&gt;

&lt;p&gt;The database becomes the source of truth. Markdown becomes a rendering layer: something you generate FROM the database when humans need to read it, not something you store.&lt;/p&gt;

&lt;p&gt;Session summaries? Query the context table, format as prose.&lt;/p&gt;

&lt;p&gt;What did we learn this week? Query learnings, render as bullet points.&lt;/p&gt;

&lt;p&gt;What's the status of this project? Query context for active contracts, generate markdown.&lt;/p&gt;

&lt;p&gt;The inversion: markdown is output, not storage. The database is storage. Machines talk to machines through structured queries. Humans get rendered views when they ask.&lt;/p&gt;

&lt;p&gt;This is where agentic systems are heading. Not smarter models generating better prose. Smarter architectures where machines communicate in formats optimized for machines.&lt;/p&gt;

&lt;p&gt;One SQLite file. Three tables. Zero rotting markdown.&lt;/p&gt;

&lt;p&gt;AgentDB is part of the Kernel Claude Code plugin for self-evolving configuration and multi-agent coordination. It comes with the AgentDB and all the agents, commands, hooks, and more detailed in this article.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Taking on consulting projects for custom Claude Code/agentic coding architecture. Also happy to trade notes if you're deep in this space. Reach out either way.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>sqlite</category>
    </item>
    <item>
      <title>I Tested OpenAI's New Codex Desktop App. The UI Is the Real Product</title>
      <dc:creator>Aria</dc:creator>
      <pubDate>Wed, 07 Oct 2026 08:14:18 +0000</pubDate>
      <link>https://dev.to/ariaxhan/i-tested-openais-new-codex-desktop-app-the-ui-is-the-real-product-1jf1</link>
      <guid>https://dev.to/ariaxhan/i-tested-openais-new-codex-desktop-app-the-ui-is-the-real-product-1jf1</guid>
      <description>&lt;h1&gt;
  
  
  I Tested OpenAI's New Codex Desktop App. The UI Is the Real Product
&lt;/h1&gt;

&lt;p&gt;I started the way I always start: by having the tool design its own configuration.&lt;/p&gt;

&lt;p&gt;If you're going to evaluate an AI coding agent, make it work on itself first. Codex passed. It pulled its full range of agentic functionalities and generated the memo I asked for.&lt;/p&gt;

&lt;p&gt;Then it gave me something I've been waiting for: a direct link to open the file in Cursor.&lt;/p&gt;

&lt;p&gt;Such a small thing. I've been wishing the terminal did that for weeks.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg07e13lh7bdexglrnu8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg07e13lh7bdexglrnu8.png" width="800" height="723"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;But don't get too excited. The link is broken 80% of the time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The UI is the Story
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fei5twh98dfmhn1o3et5d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fei5twh98dfmhn1o3et5d.png" width="800" height="501"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Look at the bar in the top right corner. It's not an IDE. It's not a chatbot.&lt;/p&gt;

&lt;p&gt;It's the first genuinely agent-native interface I've seen, and it nails exactly what I reach for most when coding with AI.&lt;/p&gt;

&lt;p&gt;Git operations (commit, push, worktrees) tucked into convenient locations. Terminal toggle. IDE toggle. All the friction points I hit multiple times per session, smoothed away.&lt;/p&gt;

&lt;p&gt;And a special favorite of mine? AI-powered run controls in a button with environment settings.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5eseopmq8yskzqbfzhp8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5eseopmq8yskzqbfzhp8.png" width="800" height="908"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Automations
&lt;/h2&gt;

&lt;p&gt;This is my favorite part so far. I asked Codex to generate a daily automation based on my context (existing projects, AI configs, Claude Code conversations, etc.). It decided on, appropriately, a Context drift radar and drafted it up in the chat with a clean "Create" button.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwfrwtjd5zqik8876fd27.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwfrwtjd5zqik8876fd27.png" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then a quick modal, already filled in so all I have to do is hit "Save."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgfjwuhxrnu4vfqzy1rot.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgfjwuhxrnu4vfqzy1rot.png" width="800" height="1203"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I appreciate that it builds the full automation instead of ever asking me to type into an empty box. You can test your automations in the Automations panel with a single click, and it triggers seamless creation of a new conversation and corresponding worktree.&lt;/p&gt;

&lt;h2&gt;
  
  
  Skills
&lt;/h2&gt;

&lt;p&gt;Skills let you extend Codex beyond code generation. Bundle instructions, resources, and scripts into a reusable package; Codex picks them up automatically or on command.&lt;/p&gt;

&lt;p&gt;I converted one of my Claude Code commands into a Codex Skill. All I had to do was ask.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F50p3rnb6qhs0ntmdiddd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F50p3rnb6qhs0ntmdiddd.png" width="800" height="424"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The openai.yaml format produced a clean UI-ready entry.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnjlg2qfvz2gnagh5lrc7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnjlg2qfvz2gnagh5lrc7.png" width="800" height="959"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I like that the skill system respects complexity. This isn't "generate boilerplate"; it's a multi-phase workflow with branching logic, and Codex handles it cleanly. Skills sync across app, CLI, and IDE extension, and you can check them into your repo for team access.&lt;/p&gt;

&lt;p&gt;OpenAI ships built-in skills (Figma, Linear, Vercel, image generation, document creation). But the real value is bringing your own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models
&lt;/h2&gt;

&lt;p&gt;The Codex models are what's available; no support for external models yet (based on the current UI). You can log in with either your ChatGPT subscription or your API account.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fta73u5xvmm5wjcxr0xsm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fta73u5xvmm5wjcxr0xsm.png" width="800" height="977"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When I tried GPT-5.2 Codex Medium, it failed with an error saying the model isn't supported on ChatGPT accounts. Swapped to GPT-5.2 Codex Low and it responded much faster, but the quality drop was drastic. First impressions put it below Claude's Haiku.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0gayb5u5y0k22q23ob6y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0gayb5u5y0k22q23ob6y.png" width="800" height="138"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Personalization
&lt;/h2&gt;

&lt;p&gt;OpenAI is offering a choice between two interaction styles: terse/pragmatic or conversational/empathetic. Same capabilities, different tone.&lt;/p&gt;

&lt;p&gt;I set mine to pragmatic: "concise, task-focused, and direct."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgdghqmkrhhrp1xdkratt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgdghqmkrhhrp1xdkratt.png" width="800" height="607"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In response, it decided to answer every message by complimenting my request.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"This is a thoughtful automation prompt, and I like that you want it grounded in real review signals."&lt;/p&gt;

&lt;p&gt;"This is a great prompt to work on together, and I can already see a few high-leverage spots to tighten your flow."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sounds a lot more like "warm, collaborative, and helpful" than pragmatic.&lt;/p&gt;

&lt;p&gt;I'll tune it with my own system prompt, but it's amusing that they landed on exactly two options and then didn't quite commit to the distinction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Catch
&lt;/h2&gt;

&lt;p&gt;The UI is exciting. The models are not.&lt;/p&gt;

&lt;p&gt;GPT-5.2 Codex Medium doesn't work consistently. This is brand new, so some slack is warranted. But GPT-5.2 Codex Low quickly revealed itself as inadequate for actual code generation, and that left me with very few options.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means
&lt;/h2&gt;

&lt;p&gt;OpenAI shipped something important: an interface that finally looks like it was designed for AI-native development. The chatbot paradigm is cracking.&lt;/p&gt;

&lt;p&gt;The execution on UI details is sharp. Git integration, environment controls, diff views; these aren't features, they're the removal of obstacles. And the automation is clean. That's what's been missing, and I've been working around it with hooks and commands and all sorts of patches. I appreciate having it all built in.&lt;/p&gt;

&lt;p&gt;But I've been tired of the ChatGPT voice for a long time, and it's back in full force here. Even "pragmatic mode" didn't help. The models need work, both in voice and in quality of execution.&lt;/p&gt;

&lt;p&gt;My verdict for now: use Codex for the workflow automation. Test the models. Experiment with what they can and can't do.&lt;/p&gt;

&lt;p&gt;The real signal isn't what this tool can do today. It's that OpenAI finally built a UI that admits the chatbot was the wrong frame all along.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>openai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I Put ChatGPT in Charge of Claude Code</title>
      <dc:creator>Aria</dc:creator>
      <pubDate>Wed, 07 Oct 2026 08:13:38 +0000</pubDate>
      <link>https://dev.to/ariaxhan/i-put-chatgpt-in-charge-of-claude-code-3lpm</link>
      <guid>https://dev.to/ariaxhan/i-put-chatgpt-in-charge-of-claude-code-3lpm</guid>
      <description>&lt;h1&gt;
  
  
  I Put ChatGPT in Charge of Claude Code
&lt;/h1&gt;

&lt;p&gt;I love Claude Code. I have spent an unreasonable number of hours in that terminal. But even I'm getting tired of staring at it, watching it confidently over-engineer a three-file feature into a twelve-module architecture while I whisper "please don't" at my screen.&lt;/p&gt;

&lt;p&gt;Despite all my guardrails, Claude Code is still spectacular at going rogue.&lt;/p&gt;

&lt;p&gt;So I gave it a babysitter.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setup
&lt;/h2&gt;

&lt;p&gt;I downloaded the ChatGPT desktop app, gave it Accessibility permissions, and it could immediately see the terminal window running Claude Code. Split screen. Two AIs. One watching the other work.&lt;/p&gt;

&lt;p&gt;The configuration is simple: Claude Code builds. ChatGPT observes. I supervise both, but ChatGPT handles the tedious part of supervision for me: catching incorrect assumptions, flagging subtle implementation mistakes, writing course-correction prompts, and (more than anything) maintaining a stable context that persists outside of Claude Code's auto-compact window.&lt;/p&gt;

&lt;p&gt;That last part is the real unlock.&lt;/p&gt;

&lt;h2&gt;
  
  
  The First Save
&lt;/h2&gt;

&lt;p&gt;I told Claude Code to execute a plan we had already assembled, then asked ChatGPT what it thought about the current execution behavior.&lt;/p&gt;

&lt;p&gt;My default is parallel agents. In this case, ChatGPT immediately told me to stop. It designed a new protocol on the spot: Contract Lock and Sequential repo execution with commit and verify at each step.&lt;/p&gt;

&lt;p&gt;That saved me hours of debugging right there.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnpkdnbhji7fp2ahbzhmk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnpkdnbhji7fp2ahbzhmk.png" width="800" height="536"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then it did something I wasn't expecting. It didn't just validate the ongoing process; it gave me specific things to look out for while supervising execution. A small thing. But knowing what to watch for before problems appear is a completely different mode of working than reacting to failures after the fact.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwfhd3jbcaugp6p4m687g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwfhd3jbcaugp6p4m687g.png" width="800" height="660"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where It Gets Good
&lt;/h2&gt;

&lt;p&gt;Execution phase. ChatGPT checking in along the way. The most helpful part was the mistakes it caught without me asking: it defined checkpoints, specified what to have Claude Code output at each stage, and preemptively identified assumptions we had already encountered and burned by before.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1rv06dt43uzizjs1torb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1rv06dt43uzizjs1torb.png" width="800" height="262"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdm5cqb3nlcu24tlk2opz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdm5cqb3nlcu24tlk2opz.png" width="800" height="245"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;No prompting required. It just did it.&lt;/p&gt;

&lt;p&gt;And then the most important moment: Claude Code runs out of context. This is the wall every heavy Claude Code user hits. The conversation compacts, critical details vanish, and you're left trying to reconstruct state from memory and Markdown files.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdnz32z01numqaxe0n2s8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdnz32z01numqaxe0n2s8.png" width="800" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;ChatGPT's handoff prompt included the full ongoing history from its own conversation. Continuation across Claude Code instances with consistent, accumulated context.&lt;/p&gt;

&lt;p&gt;I kept it to one terminal instance for now. But this is obviously the architecture for an external orchestrator managing parallel executions too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Quiet Productivity Gains
&lt;/h2&gt;

&lt;p&gt;Here's the part that surprised me most.&lt;/p&gt;

&lt;p&gt;I started using ChatGPT for all the little questions that usually break the flow of a Claude Code session: explaining a specific package in depth, fixing a broken Python environment that Claude Code decided wasn't its problem, understanding why a dependency conflict exists instead of just brute-forcing past it.&lt;/p&gt;

&lt;p&gt;These questions are important. They build my understanding. But asking them inside Claude Code means waiting for Opus to take an unreasonably long time answering something simple, and burning context on things that have nothing to do with the current task.&lt;/p&gt;

&lt;p&gt;Now I have a separate AI designated to do nothing other than talk to me and monitor Claude Code. The familiarity of a chat interface works because what I actually want from this layer is conversation, not agentic behavior (definitely will be trying Codex next). And it's far more stable than trying to have two agents edit the same codebase simultaneously.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Evaluation Problem
&lt;/h2&gt;

&lt;p&gt;This is the part I didn't anticipate solving.&lt;/p&gt;

&lt;p&gt;Working only in Claude Code created a void: only Claude was commenting on, iterating on, and learning from its own mistakes. No external signal. No second opinion. Just one model evaluating its own output in a closed loop.&lt;/p&gt;

&lt;p&gt;Adding ChatGPT as an observer immediately catches things Claude Code misses. Over time, it builds a nuanced understanding of my patterns, my projects, my failure modes. Not in Markdown files that get compacted away, but in ChatGPT's native memory: summarization, trait tracking, persistent context that actually sticks.&lt;/p&gt;

&lt;p&gt;I experiment constantly with different methods, prompts, protocols. Instead of running blind, I now have something that tells me when an approach makes sense and when it doesn't.&lt;/p&gt;

&lt;p&gt;Two models looking at the same problem from different angles. That's not redundancy; it's depth.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Matters Here
&lt;/h2&gt;

&lt;p&gt;Three takeaways worth keeping:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Separation of concerns applies to your AI workflow too.&lt;/strong&gt; Claude Code builds. ChatGPT monitors, explains, and maintains context. Trying to make one model do both is how you get 200k tokens of tangled conversation where half of it is you asking "wait, what were we doing?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context persistence is the bottleneck nobody talks about.&lt;/strong&gt; The auto-compact wall isn't a minor annoyance; it's where most agentic coding sessions silently degrade. An external observer that carries history across instances changes the failure mode entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Self-evaluation is not evaluation.&lt;/strong&gt; One model reviewing its own output in a closed loop will always have blind spots. A second model with different training, different priors, and a completely separate conversation history catches things the first one structurally cannot.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Goes Next
&lt;/h2&gt;

&lt;p&gt;The obvious extension: ChatGPT as a proper orchestrator for parallel Claude Code instances, each running in separate terminals, each with context managed externally. Or Codex so it can do even more, like write/update the existing local configuration and read files itself. I did try connecting Cursor, but all it can do so far is read the current active file.&lt;/p&gt;

&lt;p&gt;Less obvious but more interesting: using ChatGPT's persistent memory as a long-term learning layer. Not just session continuity, but accumulated pattern recognition across weeks and months of work. Which protocols actually reduced bugs. Which architectural decisions held up. Which ones didn't.&lt;/p&gt;

&lt;p&gt;Right now this is two apps and a split screen. The tooling will catch up. The insight that won't change is simpler: the best way to supervise an AI building things is with another AI whose only job is to pay attention.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>chatgpt</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
