<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bruce He</title>
    <description>The latest articles on DEV Community by Bruce He (@bruce_he).</description>
    <link>https://dev.to/bruce_he</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3860271%2F633c47d5-d82f-45ed-825b-37cb34e7ae6a.jpg</url>
      <title>DEV Community: Bruce He</title>
      <link>https://dev.to/bruce_he</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bruce_he"/>
    <language>en</language>
    <item>
      <title>CS146S Fall 2026: How to Watch and Follow Along Free</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:48:47 +0000</pubDate>
      <link>https://dev.to/bruce_he/cs146s-fall-2026-how-to-watch-and-follow-along-free-bj0</link>
      <guid>https://dev.to/bruce_he/cs146s-fall-2026-how-to-watch-and-follow-along-free-bj0</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvo9ek0cwwb8850fmp2za.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvo9ek0cwwb8850fmp2za.webp" alt="CS146S Fall 2026 calendar and free follow-along plan for non-Stanford students" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can follow Stanford &lt;strong&gt;CS146S Fall 2026&lt;/strong&gt; for free starting Tuesday, September 22, but you can't &lt;em&gt;watch&lt;/em&gt; it: as of September 11, 2026, there are no official lecture recordings, and the "CS146S video" playlist Google keeps surfacing is a third-party recap of &lt;em&gt;last year's&lt;/em&gt; syllabus. That distinction matters more this year than last, because the Fall 2026 syllabus is a different course. Six of ten weeks are new topics, the grading now puts 30% on open-source contributions, and Tuesday/Thursday replaces Monday/Friday.&lt;/p&gt;

&lt;p&gt;This page is the operational answer to the questions people keep landing here with: where the videos are (and aren't), the exact Fall 2026 calendar, what's free versus enrollment-only, what changed since Fall 2025, and a week-by-week plan that pairs each new theme with a tool you can run tonight. I'm deliberately not re-explaining the ten-week syllabus or the guest lineup here; that lives in my &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;CS146S overview&lt;/a&gt;. And the lecture-by-lecture notes for the Fall 2025 material are in the &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-02-cs146s-study-guide/" rel="noopener noreferrer"&gt;CS146S study guide&lt;/a&gt;. This post is the layer on top for people who want to run the quarter in real time.&lt;/p&gt;

&lt;p&gt;One receipt up front, because the calendar is the whole point. The official site's Calendar tab is client-rendered and currently links to &lt;code&gt;#&lt;/code&gt;, so I pulled the Fall 2026 schedule straight out of the site's JavaScript bundle on September 11. Every date and speaker below comes from that data, not from a secondhand post.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "CS146S Video" Actually Gets You
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;There is no official CS146S lecture video, and there never was.&lt;/strong&gt; A student who worked through the Fall 2025 course wrote it plainly in his &lt;a href="https://github.com/georgestephenson/cs146s-modern-software-dev" rel="noopener noreferrer"&gt;notes repo&lt;/a&gt;: "As of writing there are no video recordings of the lectures that are publicly available." Nothing on the &lt;a href="https://themodernsoftware.dev/" rel="noopener noreferrer"&gt;Fall 2026 site&lt;/a&gt; changes that. Course materials and submissions go through Stanford Canvas; Ed Discussion links go "to enrolled students."&lt;/p&gt;

&lt;p&gt;So what is the playlist that ranks for the query? It's &lt;a href="https://www.youtube.com/playlist?list=PLxpwjSdVZQ95OWe3QvkVEcU1f1X-6NDj9" rel="noopener noreferrer"&gt;CS146S The Modern Software Development&lt;/a&gt; on a channel called AI With Ryan: nine videos, 14 to 16 minutes each, 29,044 views as of this week, last updated February 16, 2026. They're AI-narrated summaries of the Fall 2025 slide decks. The Week 1 description even opens by calling CS146S a course on "scalable systems, real-world algorithms, and high-performance engineering," which it is not. They're fine as a commute-length preview of the 2025 topics. They are not lectures, and they cover a syllabus that Fall 2026 has largely replaced.&lt;/p&gt;

&lt;p&gt;Here's the honest inventory of everything you can actually press play on:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;th&gt;Term covered&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.youtube.com/playlist?list=PLxpwjSdVZQ95OWe3QvkVEcU1f1X-6NDj9" rel="noopener noreferrer"&gt;AI With Ryan playlist&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;9 AI-narrated recaps, 14-16 min each&lt;/td&gt;
&lt;td&gt;Fall 2025&lt;/td&gt;
&lt;td&gt;Preview only; not lectures&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://www.youtube.com/watch?v=wEsjK3Smovw" rel="noopener noreferrer"&gt;From Writing Code to Managing Agents&lt;/a&gt; (EO)&lt;/td&gt;
&lt;td&gt;Interview with Mihail Eric&lt;/td&gt;
&lt;td&gt;Course philosophy&lt;/td&gt;
&lt;td&gt;Worth 30 minutes for the "agent manager" framing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://www.youtube.com/watch?v=jOe4fJSc2IE" rel="noopener noreferrer"&gt;The Modern Software Engineer&lt;/a&gt; (AAIF Live)&lt;/td&gt;
&lt;td&gt;Conference talk by Eric&lt;/td&gt;
&lt;td&gt;Course philosophy&lt;/td&gt;
&lt;td&gt;Same thesis, different audience&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fall 2025 Google Slides&lt;/td&gt;
&lt;td&gt;Linked per session on the &lt;a href="https://themodernsoftware.dev/fall2025" rel="noopener noreferrer"&gt;/fall2025&lt;/a&gt; syllabus tab&lt;/td&gt;
&lt;td&gt;Fall 2025&lt;/td&gt;
&lt;td&gt;The real primary source; 17 decks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fall 2026 slides&lt;/td&gt;
&lt;td&gt;Not posted as of Sep 11&lt;/td&gt;
&lt;td&gt;Fall 2026&lt;/td&gt;
&lt;td&gt;Check the syllabus tab each Tuesday&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The practical consequence: "how can I watch it" has the same answer as "how can I follow along." You read the deck, do the reading, run the exercise, and you do it on the week the class does, so the guest-talk chatter on X and the speakers' own posts land while the topic is fresh. That's the plan in the second half of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The CS146S Fall 2026 Calendar
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Fall 2026 runs ten weeks, Tuesday and Thursday, from September 22 to December 3, in room 420-041.&lt;/strong&gt; Tuesdays are Eric's lectures; Thursdays are mostly guests. Stanford's &lt;a href="https://studentservices.stanford.edu/calendar-events/academic-calendars/future-academic-calendars/stanford-academic-calendar-2026-2027" rel="noopener noreferrer"&gt;2026-27 academic calendar&lt;/a&gt; puts instruction start on September 22, the study-list deadline on October 9, Thanksgiving recess November 23-27, last day of classes December 4, and exams December 7-11. The course skips the Thanksgiving week entirely, which is why Week 10 lands on December 1.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Wk&lt;/th&gt;
&lt;th&gt;Tuesday lecture&lt;/th&gt;
&lt;th&gt;Thursday session&lt;/th&gt;
&lt;th&gt;Guest&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Sep 22: Course intro + build Claude Code in 200 lines&lt;/td&gt;
&lt;td&gt;Sep 24: How SOTA coding agents are designed (system prompts)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Sep 29: Advanced prompting + RePPIT, spec-driven development&lt;/td&gt;
&lt;td&gt;Oct 1: Full introduction to MCP and tool calling&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Oct 6: All about agent skills (incl. web skills)&lt;/td&gt;
&lt;td&gt;Oct 8&lt;/td&gt;
&lt;td&gt;Lee Robinson, VP DevRel @ Cursor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Oct 13: Customizing your setup (CLAUDE.md, AGENTS.md, hooks)&lt;/td&gt;
&lt;td&gt;Oct 15&lt;/td&gt;
&lt;td&gt;Boris Cherny, creator of Claude Code @ Anthropic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Oct 20: Agent readiness in your repos&lt;/td&gt;
&lt;td&gt;Oct 22&lt;/td&gt;
&lt;td&gt;Eno Reyes, CTO @ Factory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Oct 27: Agentic code review, best practices and architectures&lt;/td&gt;
&lt;td&gt;Oct 29&lt;/td&gt;
&lt;td&gt;Silas Alberti, SVP Research @ Cognition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Nov 3: Security in AI codebases&lt;/td&gt;
&lt;td&gt;Nov 5&lt;/td&gt;
&lt;td&gt;Isaac Evans, CEO @ Semgrep&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Nov 10: Background agents, launching tasks asynchronously&lt;/td&gt;
&lt;td&gt;Nov 12&lt;/td&gt;
&lt;td&gt;Rajesh Bhatia, Senior Director @ Cloudflare&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Nov 17: Guest, Elad Gil, investor @ Gil Capital&lt;/td&gt;
&lt;td&gt;Nov 19&lt;/td&gt;
&lt;td&gt;Amjad Masad, CEO @ Replit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Nov 23-27: Thanksgiving recess, no class&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Dec 1: Coding agents in big teams (MCP portals, gateways, routing)&lt;/td&gt;
&lt;td&gt;Dec 3: The Software Factory: self-running, self-improving systems&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two logistics notes for people outside California. Lecture times aren't published on the public site, only the days. And Pacific time flips from PDT (UTC-7) to PST (UTC-8) on November 1, 2026, so a Thursday afternoon guest talk is Friday early morning in Beijing, Singapore, or Tokyo either way, and one hour later after the switch. If you're following from Europe, Thursday afternoon Pacific is late Thursday evening for you. Since you can't attend anyway, what this actually affects is &lt;em&gt;when&lt;/em&gt; the post-talk material shows up: speakers' slides and threads tend to land within 24 hours, so a Friday-morning check on the syllabus tab and on X is the habit that pays off.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What's New in Fall 2026 vs Fall 2025
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Fall 2025 was a lifecycle tour; Fall 2026 is an operator's manual.&lt;/strong&gt; The 2025 syllabus walked the SDLC (prompting, agents, IDE, terminal, testing, review, app building, ops, future). The 2026 syllabus assumes you already live inside a coding agent and asks how you configure, constrain, review, parallelize, and scale it. Here's the diff, item by item:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Fall 2025&lt;/th&gt;
&lt;th&gt;Fall 2026&lt;/th&gt;
&lt;th&gt;My read&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Days&lt;/td&gt;
&lt;td&gt;Mon / Fri&lt;/td&gt;
&lt;td&gt;Tue / Thu&lt;/td&gt;
&lt;td&gt;Guest talks now mid-week; slides land before the weekend&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grading&lt;/td&gt;
&lt;td&gt;Final project 80%, assignments 15%, participation 5%&lt;/td&gt;
&lt;td&gt;Final project 50%, &lt;strong&gt;open source contributions 30%&lt;/strong&gt;, assignments 15%, participation 5%&lt;/td&gt;
&lt;td&gt;The biggest change on the page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 1&lt;/td&gt;
&lt;td&gt;LLM basics, prompting&lt;/td&gt;
&lt;td&gt;Build Claude Code in 200 lines, read production system prompts&lt;/td&gt;
&lt;td&gt;Deep end, day one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 2&lt;/td&gt;
&lt;td&gt;Agent anatomy, MCP&lt;/td&gt;
&lt;td&gt;RePPIT + spec-driven dev, MCP&lt;/td&gt;
&lt;td&gt;Specs promoted to a core method&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 3&lt;/td&gt;
&lt;td&gt;AI IDE, context&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Agent skills + CLI&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;New&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 4&lt;/td&gt;
&lt;td&gt;Agent patterns&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;CLAUDE.md / AGENTS.md, hooks, subagents&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;New&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 5&lt;/td&gt;
&lt;td&gt;Warp terminal&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Agent-ready codebases&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Terminal week gone; readiness scoring in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 6&lt;/td&gt;
&lt;td&gt;Testing + security&lt;/td&gt;
&lt;td&gt;Agentic code review&lt;/td&gt;
&lt;td&gt;Review gets its own week, earlier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 7&lt;/td&gt;
&lt;td&gt;Code review&lt;/td&gt;
&lt;td&gt;Security (SAST/SCA, prompt injection)&lt;/td&gt;
&lt;td&gt;Swapped order with review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 8&lt;/td&gt;
&lt;td&gt;One-prompt app building&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Background agents, fleets, issue-to-PR&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;New&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 9&lt;/td&gt;
&lt;td&gt;Post-deployment ops&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;AI-native team: MCP portals, LLM gateways, cost routing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;New&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Week 10&lt;/td&gt;
&lt;td&gt;Future of SWE&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;The Software Factory&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;New&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Returning guests&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Boris Cherny, Silas Alberti, Isaac Evans&lt;/td&gt;
&lt;td&gt;Anthropic, Cognition, Semgrep again&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New guests&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Lee Robinson (Cursor), Eno Reyes (Factory), Rajesh Bhatia (Cloudflare), Elad Gil, Amjad Masad (Replit)&lt;/td&gt;
&lt;td&gt;Two of the five run agent companies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dropped&lt;/td&gt;
&lt;td&gt;Zach Lloyd (Warp), Tomas Reimers (Graphite), Gaspar Garcia (Vercel), Resolve.ai, Martin Casado (a16z)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Terminal, UI-gen, and ops angles cut&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 30% for open-source contributions is the line to stare at. Fall 2025 was "build a thing"; Fall 2026 is "get a PR merged into a real project with an agent." The site lists fifteen open-source partners (OpenHands, marimo, CrewAI, Semgrep, Milvus, Unsloth, cmux, pi.dev, Arize Phoenix, Browserbase, CopilotKit, HeyGen, Vercel, Warp, Anyscale), and it's a safe bet that's the contribution pool. Read it as the instructor saying out loud that agent-built code that never leaves your laptop isn't the skill being taught anymore.&lt;/p&gt;

&lt;p&gt;Two things haven't caught up yet, and you should know before you start. The &lt;a href="https://github.com/mihail911/modern-software-dev-assignments" rel="noopener noreferrer"&gt;assignments repo&lt;/a&gt; sits at 3,950 stars and 949 forks, but its last push was November 10, 2025; it still has &lt;code&gt;week1&lt;/code&gt; through &lt;code&gt;week8&lt;/code&gt; for the 2025 syllabus and nothing for 2026. And the Fall 2026 syllabus entries currently have topics and speakers but no readings, no slides, and no assignment links. I'd expect those to appear week by week, the way they did in 2025. Until then, the 2025 repo is the only runnable material, and about half of it maps cleanly onto the new weeks (the table further down says which half).&lt;/p&gt;

&lt;h2&gt;
  
  
  Free vs Enrollment-Only
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;"Is CS146S free?" has a two-part answer: the materials are, the course isn't.&lt;/strong&gt; The &lt;a href="https://bulletin.stanford.edu/courses/2274401" rel="noopener noreferrer"&gt;Stanford bulletin&lt;/a&gt; lists it as 3-4 units, Letter or Credit/No Credit, with a September 27 application deadline for students who lack the formal prerequisites (CS111 and CS161 equivalents). The course FAQ says auditing is open "to Stanford students and staff." I searched for an SCPD or Stanford Online listing and found none, so as of today there's no route for a working engineer to enroll remotely.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;You get&lt;/th&gt;
&lt;th&gt;Free, public&lt;/th&gt;
&lt;th&gt;Enrolled Stanford students only&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Syllabus with topics, dates, speakers (both terms)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fall 2025 slide decks, readings, 8 weeks of assignment code&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fall 2026 slides and readings&lt;/td&gt;
&lt;td&gt;As they're posted (none yet)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lectures in 420-041, Tue/Thu&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Guest Q&amp;amp;A (Cherny, Robinson, Masad, and the rest)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ed Discussion, Canvas, office hours (Fri 12:00-12:30)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Graded assignments and final project feedback&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open-source contribution as a graded, scaffolded deliverable&lt;/td&gt;
&lt;td&gt;Self-imposed only&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Units on a Stanford transcript&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool subscriptions (Claude Code and similar)&lt;/td&gt;
&lt;td&gt;Your own bill; the FAQ warns "some cloud-based services may require subscriptions"&lt;/td&gt;
&lt;td&gt;Course "will provide access or alternatives where possible"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you want live instruction from Eric without a Stanford ID, the only option is his paid Maven course, &lt;a href="https://maven.com/the-modern-software-developer/ai-course" rel="noopener noreferrer"&gt;AI Software Development: From First Prompt to Production Code&lt;/a&gt;: four weeks, 3-4 hours a week, eight live sessions. Its module list mirrors the Fall 2026 themes almost line for line ("Build and Integrate an MCP Server or Agent Skill," "Implementing the RePPIT Dev Loop"), which tells you the new Stanford syllabus and the public course were designed together.&lt;/p&gt;

&lt;p&gt;Two caveats I can't resolve from outside: the most recent cohort I can see in the page data ran June 22 to July 17, 2026, with no fall cohort listed yet, and the price isn't shown publicly without going through checkout. Treat it as the "I need feedback and a deadline" option, not the default.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Follow-Along Plan: Each New Theme, Paired With a Tool
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Every one of the five headline themes on the Fall 2026 site (MCP, agent skills, spec-driven development, loop engineering, the software factory) already has a runnable 2026 stack, and most of them have a deep-dive on this blog.&lt;/strong&gt; That's the pairing that makes following along without a classroom actually work: the lecture gives you the &lt;em&gt;why&lt;/em&gt; and the vocabulary; the tool gives you the reps.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The week-by-week version below is the thing to screenshot. "2025 material" is what's runnable today from the public repo and slides; "Do this" is the one exercise I'd spend the week's hours on; "Read" is the post on this blog that goes deeper than the deck will.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Wk&lt;/th&gt;
&lt;th&gt;Dates&lt;/th&gt;
&lt;th&gt;Theme&lt;/th&gt;
&lt;th&gt;2025 material you can use now&lt;/th&gt;
&lt;th&gt;Do this (4-6 hrs)&lt;/th&gt;
&lt;th&gt;Read&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Sep 22-24&lt;/td&gt;
&lt;td&gt;Agent internals&lt;/td&gt;
&lt;td&gt;2025 W2 Mon deck "Building a coding agent from scratch" + completed exercise&lt;/td&gt;
&lt;td&gt;Write a 200-line agent loop with read/write/edit/bash tools; then diff it against a real CLI agent's system prompt&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-02-cs146s-study-guide/" rel="noopener noreferrer"&gt;Study guide, Weeks 1-2&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Sep 29-Oct 1&lt;/td&gt;
&lt;td&gt;Specs, RePPIT, MCP&lt;/td&gt;
&lt;td&gt;2025 W3 assignment "Build a Custom MCP Server"; W3 design-doc template&lt;/td&gt;
&lt;td&gt;Ship one MCP server, then run the same feature spec-first with OpenSpec and compare token spend&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-28-openspec-superpowers-workflow/" rel="noopener noreferrer"&gt;OpenSpec + Superpowers workflow&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Oct 6-8&lt;/td&gt;
&lt;td&gt;Agent skills + CLI&lt;/td&gt;
&lt;td&gt;None (new topic)&lt;/td&gt;
&lt;td&gt;Convert one of your MCP tools into a SKILL.md + script; measure the context cost of each&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;MCP vs Skills: why CLI + Skill wins&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Oct 13-15&lt;/td&gt;
&lt;td&gt;CLAUDE.md, AGENTS.md, hooks, subagents&lt;/td&gt;
&lt;td&gt;2025 W4 "Coding with Claude Code" assignment; Anthropic's Claude Code best practices&lt;/td&gt;
&lt;td&gt;Add a lint/test hook that blocks bad commits; split one task into planner / implementer / reviewer subagents&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;Context engineering 2026&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Oct 20-22&lt;/td&gt;
&lt;td&gt;Agent-ready codebases&lt;/td&gt;
&lt;td&gt;None (new topic)&lt;/td&gt;
&lt;td&gt;Score one of your repos: docs, tests, checks, structure; fix the top two gaps an agent trips on&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-09-claude-code-openspec-superpowers/" rel="noopener noreferrer"&gt;Triple stack: Claude Code + OpenSpec + Superpowers&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Oct 27-29&lt;/td&gt;
&lt;td&gt;Agentic code review&lt;/td&gt;
&lt;td&gt;2025 W7 "Code Review Reps" assignment&lt;/td&gt;
&lt;td&gt;Adversarially review an AI-written PR with a second agent as critic; count what it catches vs misses&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-05-loop-engineering/" rel="noopener noreferrer"&gt;Loop engineering: a critic that says no&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Nov 3-5&lt;/td&gt;
&lt;td&gt;Security&lt;/td&gt;
&lt;td&gt;2025 W6 "Writing Secure AI Code" assignment; Semgrep blog on finding vulns with Claude Code&lt;/td&gt;
&lt;td&gt;Run Semgrep on your agent-built repo; then try one prompt-injection attack on your own MCP server&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;Overview, Week 6 + guests&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Nov 10-12&lt;/td&gt;
&lt;td&gt;Background agents&lt;/td&gt;
&lt;td&gt;None (new topic)&lt;/td&gt;
&lt;td&gt;Wire one issue-to-PR trigger (GitHub issue label or Slack message) to a cloud-run agent; review the PR cold&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-05-loop-engineering/" rel="noopener noreferrer"&gt;Loop engineering: stop conditions&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Nov 17-19&lt;/td&gt;
&lt;td&gt;AI-native team&lt;/td&gt;
&lt;td&gt;None (new topic)&lt;/td&gt;
&lt;td&gt;Put an LLM gateway in front of your agents, route one cheap task to a cheaper model, look at the bill&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;Context engineering 2026&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Nov 23-27&lt;/td&gt;
&lt;td&gt;Recess&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Catch-up week; this is when you land your open-source PR&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Dec 1-3&lt;/td&gt;
&lt;td&gt;Software factory&lt;/td&gt;
&lt;td&gt;None (new topic)&lt;/td&gt;
&lt;td&gt;Chain Weeks 4, 6, and 8 into one unattended pipeline: spec in, reviewed PR out; write down where it broke&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;MCP vs Skills&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Notice where the 2025 material runs out: Weeks 3, 5, 8, 9, and 10 have no public assignment yet. That's not a reason to wait. Those are exactly the weeks where the 2026 tooling is furthest ahead of any syllabus, and the exercises above are what practitioners were already doing this summer. If the course posts an official assignment for one of those weeks, do that one instead and treat mine as the warm-up.&lt;/p&gt;

&lt;p&gt;Two things not to do. Don't spend Week 1 re-reading the 2025 prompting deck; Fall 2026 skips prompting basics for a reason, and if you need them, the study guide's two-week core is faster. And don't stack the Maven cohort on top of this plan; the two overlap enough that you'd be paying for the same reps with a deadline attached.&lt;/p&gt;

&lt;p&gt;One thing worth doing: pair CS146S with the agent courses running in parallel this fall, because they cover what CS146S deliberately skips (training and evaluating the agent itself). Stanford &lt;a href="https://cs329z.stanford.edu/" rel="noopener noreferrer"&gt;CS329Z: Engineering AI Agents&lt;/a&gt; (Diyi Yang, Sep 23 to Dec 2) has a public syllabus, though its recordings are Canvas-only. CMU's &lt;a href="https://www.cmu-agents.com/" rel="noopener noreferrer"&gt;11-768: AI Agents&lt;/a&gt; (Graham Neubig and Daniel Fried) is the one that actually posts lecture video: four Fall 2026 lectures were on YouTube as of September 11, with more landing as the term runs. And Stanford &lt;a href="https://cs329a.stanford.edu/" rel="noopener noreferrer"&gt;CS329A: Self-Improving AI Agents&lt;/a&gt; has all nine lectures from its last run on Stanford Online's YouTube channel. If you only have bandwidth for one companion, CMU's is the closest thing to the "watchable" course people keep searching for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What You Lose by Not Enrolling
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;You lose the two things a syllabus can't ship: feedback on your final project, and being in the room for the guest Q&amp;amp;A.&lt;/strong&gt; Everything else is recoverable with discipline. Here's the honest accounting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Final project feedback (50% of the grade).&lt;/strong&gt; A TA reading your project and telling you where the agent workflow was fake is the single most valuable thing enrolled students get. Substitute: post the project publicly and ask for review in the repo of the tool you used. It's worse, but it's not nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guest Q&amp;amp;A.&lt;/strong&gt; Eight practitioners, including the people who run Cursor's developer relations, Anthropic's Claude Code, Factory, Cognition, Semgrep, Cloudflare's agent work, and Replit. You get their slides if they post them and whatever they tweet. You don't get to ask Boris Cherny why hooks are designed the way they are.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 30% open-source scaffold.&lt;/strong&gt; Enrolled students will have partner repos, presumably some mentoring, and a grade riding on a merged PR. You have the same fifteen partner repos on the public site. Pick one by Week 4, ship something by the Thanksgiving recess, and you've reproduced the deliverable, minus the grade.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cohort pressure.&lt;/strong&gt; The FAQ budgets 10-12 hours a week. Nobody following along on their own sustains that for ten weeks; 5-6 hours is what I've seen people actually keep up, which is why the table above scopes each week to one exercise.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What you don't lose is the content. Every deck the class sees will, based on the 2025 pattern, be a Google Slides link on the syllabus tab within days. The readings are public URLs. The tools are the same ones you'd install anyway. If you follow the calendar, do the exercises, and land one PR, you'll finish the quarter with more evidence of skill than most enrolled students had in December 2025, when the deliverable was still a solo project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Page Stops
&lt;/h2&gt;

&lt;p&gt;This post is the calendar and the plan. It intentionally doesn't re-explain what each lecture teaches or grade the 2025 lectures; that's the job of the &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;overview&lt;/a&gt; and the &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-02-cs146s-study-guide/" rel="noopener noreferrer"&gt;study guide&lt;/a&gt;. I'll update the calendar table if the course posts times, slides, or a Fall 2026 assignments branch, and I'll note it at the top when I do.&lt;/p&gt;

&lt;p&gt;Three facts I couldn't verify and want on the record: the lecture time of day (not on the public site), the Maven course price and whether a fall cohort exists, and whether the 30% open-source contributions are restricted to the listed partner projects. If you're enrolled and know any of these, the comments are open.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;Stanford CS146S: The Modern Software Developer, 2026 Guide&lt;/a&gt;: the full ten-week breakdown and guest lineup&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-02-cs146s-study-guide/" rel="noopener noreferrer"&gt;CS146S Study Guide 2026: Lecture-by-Lecture Notes and Workbook&lt;/a&gt;: verdicts and exercises for the Fall 2025 material&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-05-loop-engineering/" rel="noopener noreferrer"&gt;Loop Engineering: Building the Cage Your AI Agent Runs In&lt;/a&gt;: the discipline behind Weeks 4, 6, and 8&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;MCP vs Skills: Why CLI + Skill Wins the Agent Toolchain&lt;/a&gt;: the Week 3 argument, with token numbers&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-28-openspec-superpowers-workflow/" rel="noopener noreferrer"&gt;OpenSpec vs Superpowers: My Spec-Driven Workflow&lt;/a&gt;: spec-driven development in practice for Week 2&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;Context Engineering for Coding Agents 2026&lt;/a&gt;: what to put in CLAUDE.md and what to leave out&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>stanfordcs146s</category>
      <category>aicoding</category>
      <category>learningpath</category>
      <category>agenticengineering</category>
    </item>
    <item>
      <title>ego lite Review: Handing Claude Code Your Logged-In Browser</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:48:11 +0000</pubDate>
      <link>https://dev.to/bruce_he/ego-lite-review-handing-claude-code-your-logged-in-browser-12ij</link>
      <guid>https://dev.to/bruce_he/ego-lite-review-handing-claude-code-your-logged-in-browser-12ij</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fny78y05snmt13zcjbucw.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fny78y05snmt13zcjbucw.webp" alt="ego lite review: handing Claude Code your logged-in browser through Spaces" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The first real task I gave ego lite was the one its whole pitch rests on: open GitHub and tell me who's logged in. The agent came back in 65 seconds with "not logged in." I &lt;em&gt;was&lt;/em&gt; logged in. So was the browser. The agent had simply been handed the wrong one of my two imported profiles, and nothing in the skill, the docs, or the UI had told it that a second profile existed.&lt;/p&gt;

&lt;p&gt;That's the review in miniature. &lt;a href="https://github.com/citrolabs/ego-lite" rel="noopener noreferrer"&gt;ego lite&lt;/a&gt; is the most ambitious answer yet to a problem this blog has been circling since January — how to give a coding agent a browser without paying for it in tokens, flakiness, or your own sanity. The series went from headless drivers (&lt;a href="https://www.heyuan110.com/posts/ai/2026-01-13-vercel-agent-browser/" rel="noopener noreferrer"&gt;Vercel's agent-browser&lt;/a&gt;) to &lt;a href="https://www.heyuan110.com/posts/ai/2026-03-17-chrome-devtools-mcp-guide/" rel="noopener noreferrer"&gt;attaching to your real browser&lt;/a&gt; to &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-playwright-cli-skill-zero-token-automation/" rel="noopener noreferrer"&gt;cutting the token bill to zero&lt;/a&gt; to a &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-27-ubrowser-review/" rel="noopener noreferrer"&gt;12-star repo with the right cost model and no maintainer&lt;/a&gt;. ego lite is the current end of that arc: one Chromium you use every day, in which your agents get their own Spaces, your logins, and a JavaScript API instead of a CLI. 15,666 GitHub stars as of September 10, 2026, up from about 7,900 when the Chinese tech press covered it on August 3.&lt;/p&gt;

&lt;p&gt;I installed v0.5.0.28, drove it from Claude Code on three tasks of increasing difficulty, ran the same prompts through agent-browser and Chrome DevTools MCP, counted every token and every second, and then went looking for the seams. Cards on the table: at its best it was the cheapest agent browsing I have ever measured — one tool call, 16 seconds, twelve cents. At its worst it was the slowest — 30 turns and three minutes for a task agent-browser finished in 43 seconds. Both numbers are real, and the difference between them is a Claude Code setting, not the browser.&lt;/p&gt;

&lt;h2&gt;
  
  
  What ego lite actually is, from the binary up
&lt;/h2&gt;

&lt;p&gt;ego lite is a closed-source Chromium 152 fork from Citro Labs, distributed as a free DMG; the MIT-licensed GitHub repo holds only the &lt;code&gt;ego-browser&lt;/code&gt; skill and docs. The pitch has three parts: &lt;strong&gt;Spaces&lt;/strong&gt; (isolated workspaces inside one window — you browse in yours, agents work in theirs), &lt;strong&gt;inherited state&lt;/strong&gt; (it imports your Chrome or Edge profile, cookies and extensions included), and &lt;strong&gt;"code base, not CLI base"&lt;/strong&gt; — the agent writes JavaScript that runs against the page in one pass instead of issuing one shell command per click.&lt;/p&gt;

&lt;p&gt;None of the existing coverage — &lt;a href="https://agent.csdn.net/6a62db7910ee7a33f291f217.html" rel="noopener noreferrer"&gt;Hello.Reader's CSDN walkthrough&lt;/a&gt; is the clearest — explains what the CLI actually is, so I read the bundle. &lt;code&gt;ego-browser&lt;/code&gt; is a 1.9 MB native arm64 Mach-O binary that lives inside the app (&lt;code&gt;ego Framework.framework/Versions/0.5.0.28/Helpers/&lt;/code&gt;) and gets symlinked to &lt;code&gt;~/.local/bin&lt;/code&gt; at onboarding. It doesn't open a TCP port or a WebSocket; &lt;code&gt;strings&lt;/code&gt; shows it talks to the running app over a &lt;strong&gt;Mojo named IPC channel&lt;/strong&gt; (&lt;code&gt;ego.mojom.EgoCliBootstrap&lt;/code&gt; and &lt;code&gt;EgoCliBridge&lt;/code&gt;), the same transport Chromium uses between its own processes. Your script runs in an embedded Node 24.18 runtime (a separate &lt;code&gt;ego Helper (Node)&lt;/code&gt; process), against a small, deliberately un-Playwright API: &lt;code&gt;taskSpace()&lt;/code&gt;, &lt;code&gt;page.goto()&lt;/code&gt;, &lt;code&gt;page.snapshot()&lt;/code&gt;, &lt;code&gt;page.click()&lt;/code&gt;, &lt;code&gt;page.evaluate()&lt;/code&gt;, and &lt;code&gt;page.cdp()&lt;/code&gt; as the escape hatch.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-10-ego-lite-browser-review/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The install is genuinely two commands. &lt;code&gt;npx skills add citrolabs/ego-lite&lt;/code&gt; dropped the skill into &lt;code&gt;.agents/skills/ego-browser/&lt;/code&gt; in 8.5 seconds (with a "High Risk" flag from the skills.sh scanner, which I'll come back to), and the app's own onboarding had already symlinked the same skill into &lt;code&gt;~/.agents/skills/&lt;/code&gt; so every agent on the machine sees it. Typing &lt;code&gt;/ego-browser&lt;/code&gt; in Claude Code loads a 19.5 KB SKILL.md — about 4,546 tokens by my count, versus 7,372 for agent-browser's core skill.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiz0lwlot3zs0dnuyi8cg.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiz0lwlot3zs0dnuyi8cg.webp" alt="ego lite v2 what's-new page: new ego-browser skill and toolset, claiming -70% memory, -20% cost, +10% speed, +10% success rate" width="800" height="455"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The app's own what's-new page for the September 10 v2.0.0 release claims the rewritten skill cuts memory 70%, cost 20%, and adds 10% to speed and success rate. No methodology is published for any of the four; keep that in mind when you read mine, which is.&lt;/p&gt;

&lt;p&gt;Then it refused to run. Every &lt;code&gt;ego-browser nodejs&lt;/code&gt; call answered "Please complete the onboarding process first. A setup window has opened" — with no setup window anywhere on screen, even though the app had imported my Chrome data two hours earlier. The missing step turned out to be an &lt;code&gt;ego://onboarding&lt;/code&gt; page whose only content was "No importable browser data right now" and a Done button; pressing Done unlocked the CLI. Budget ten confused minutes for this.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwi72elm2ta7qfguxjpgd.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwi72elm2ta7qfguxjpgd.webp" alt="ego lite Spaces overview: nine Spaces, eight of them agent task spaces marked running, the user's Space separate" width="800" height="455"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The receipts: three tasks, three tools, one model
&lt;/h2&gt;

&lt;p&gt;Method, so you can reproduce it. Claude Code 2.1.268, &lt;code&gt;claude -p&lt;/code&gt; with &lt;code&gt;--output-format json&lt;/code&gt;, model pinned to Sonnet for every run, on an M5 MacBook Pro running macOS 26.5. Wall time is measured around the process; turns, cost, and tokens come from the JSON result. The same prompt went to each tool; only the sentence naming the tool changed. Baselines: &lt;a href="https://github.com/vercel-labs/agent-browser" rel="noopener noreferrer"&gt;Vercel agent-browser&lt;/a&gt; 0.37.1 with its bundled skill, and &lt;a href="https://github.com/ChromeDevTools/chrome-devtools-mcp" rel="noopener noreferrer"&gt;Chrome DevTools MCP&lt;/a&gt; 1.9.0 headless. The three tasks: &lt;strong&gt;(a)&lt;/strong&gt; list Hacker News' top five titles with points; &lt;strong&gt;(b)&lt;/strong&gt; open github.com and report the logged-in user; &lt;strong&gt;(c)&lt;/strong&gt; search docs.python.org for &lt;code&gt;asyncio.gather&lt;/code&gt;, open the result, and return the first code block verbatim.&lt;/p&gt;

&lt;p&gt;The first table is what you get if you run Claude Code the way most people do — permission prompts on, with an allow-list for the browser command.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Wall&lt;/th&gt;
&lt;th&gt;Turns&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Calls blocked by Claude Code&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;(a) HN top 5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;ego lite&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;76.0s&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;$0.312&lt;/td&gt;
&lt;td&gt;5 of 11&lt;/td&gt;
&lt;td&gt;✅ correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(a) HN top 5&lt;/td&gt;
&lt;td&gt;agent-browser&lt;/td&gt;
&lt;td&gt;39.9s&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;$0.298&lt;/td&gt;
&lt;td&gt;4 of 10&lt;/td&gt;
&lt;td&gt;✅ correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(a) HN top 5&lt;/td&gt;
&lt;td&gt;Chrome DevTools MCP&lt;/td&gt;
&lt;td&gt;26.4s&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;$0.230&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;✅ correct (one 13,418-token snapshot)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(b) GitHub login&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;ego lite&lt;/strong&gt;, default profile&lt;/td&gt;
&lt;td&gt;64.7s&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;$0.242&lt;/td&gt;
&lt;td&gt;5 of 8&lt;/td&gt;
&lt;td&gt;⚠️ "not logged in" — wrong profile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(b) GitHub login&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;ego lite&lt;/strong&gt;, profile named&lt;/td&gt;
&lt;td&gt;35.6s&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;$0.178&lt;/td&gt;
&lt;td&gt;2 of 5&lt;/td&gt;
&lt;td&gt;✅ "logged in as heyuan110"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(c) Docs search&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;ego lite&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;194.8s&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;$0.751&lt;/td&gt;
&lt;td&gt;11 of 29&lt;/td&gt;
&lt;td&gt;✅ correct, painfully&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(c) Docs search&lt;/td&gt;
&lt;td&gt;agent-browser&lt;/td&gt;
&lt;td&gt;42.6s&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;$0.305&lt;/td&gt;
&lt;td&gt;1 of 11&lt;/td&gt;
&lt;td&gt;✅ correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(c) Docs search&lt;/td&gt;
&lt;td&gt;Chrome DevTools MCP&lt;/td&gt;
&lt;td&gt;51.3s&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;$0.325&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;✅ correct&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read the "blocked" column before anything else. ego lite's whole interaction model is a shell heredoc containing JavaScript, and Claude Code's Bash guard rejects any command where a brace sits next to a quote — which is to say any object literal: &lt;code&gt;{ keep: [] }&lt;/code&gt;, &lt;code&gt;{ scope: "full_page" }&lt;/code&gt;, &lt;code&gt;{ profileId: "Profile 1" }&lt;/code&gt;. The error is &lt;code&gt;Contains brace with quote character (expansion obfuscation)&lt;/code&gt;, and in &lt;code&gt;-p&lt;/code&gt; mode it's a hard refusal, not a prompt. On task (c) the agent burned 11 of 29 calls on this, tried to write a script file instead (also refused — &lt;code&gt;Write&lt;/code&gt; wasn't allow-listed), tried &lt;code&gt;printf&lt;/code&gt;, tried &lt;code&gt;JSON.parse('{"scope":"full_page"}')&lt;/code&gt; as a workaround, and eventually got there. agent-browser hit the same wall on its &lt;code&gt;eval --stdin&lt;/code&gt; path, just less often, because most of its commands are flag-shaped.&lt;/p&gt;

&lt;p&gt;ego's docs say to run Claude Code with "Full access." So I did — &lt;code&gt;--dangerously-skip-permissions&lt;/code&gt; — and re-ran (a) and (c) for the two CLI tools. Chrome DevTools MCP wasn't blocked in the first round, so its numbers stand.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Wall&lt;/th&gt;
&lt;th&gt;Turns&lt;/th&gt;
&lt;th&gt;Browser calls&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Tool-result tokens&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;(a) HN top 5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;ego lite&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;16.4s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.123&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;151&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(a) HN top 5&lt;/td&gt;
&lt;td&gt;agent-browser&lt;/td&gt;
&lt;td&gt;19.0s&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;$0.176&lt;/td&gt;
&lt;td&gt;177&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(c) Docs search&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;ego lite&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;40.5s&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;$0.219&lt;/td&gt;
&lt;td&gt;3,805&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;(c) Docs search&lt;/td&gt;
&lt;td&gt;agent-browser&lt;/td&gt;
&lt;td&gt;40.8s&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;$0.290&lt;/td&gt;
&lt;td&gt;4,253&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's the number the marketing is about, and it holds. On Hacker News, one &lt;code&gt;ego-browser nodejs&lt;/code&gt; call — &lt;code&gt;taskSpace&lt;/code&gt;, &lt;code&gt;goto&lt;/code&gt;, &lt;code&gt;waitForSelector&lt;/code&gt;, one &lt;code&gt;page.evaluate&lt;/code&gt; that maps the rows to &lt;code&gt;{title, points}&lt;/code&gt; — and done: 2 turns, 151 tokens of tool output, 16.4 seconds of which most is Sonnet thinking. agent-browser needed &lt;code&gt;open&lt;/code&gt;, &lt;code&gt;snapshot&lt;/code&gt;, &lt;code&gt;eval&lt;/code&gt;, &lt;code&gt;close&lt;/code&gt;. On the docs task, the two tie on wall time and ego wins on turns and cost by about 25%. ego's own homepage claims "up to 3.45× faster than agent-browser" on complex tasks (the README says 2.5×; the numbers aren't published); my complex task came out even on time and cheaper by a quarter. Not 3.45×, but a real win — with the setting that its docs recommend and that most security-minded readers of this blog won't use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the token win actually comes from (not the snapshot)
&lt;/h2&gt;

&lt;p&gt;Here is the misconception I most want to kill, because ego's own comparison matrix encourages it: the "compressed semantic input" checkbox. I loaded the exact same Hacker News front page in every tool and counted the snapshot each one hands the model, with the same &lt;code&gt;o200k_base&lt;/code&gt; tokenizer I've used across this series.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Snapshot&lt;/th&gt;
&lt;th&gt;Chars&lt;/th&gt;
&lt;th&gt;Tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ego lite&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;page.snapshot()&lt;/code&gt; (viewport)&lt;/td&gt;
&lt;td&gt;30,469&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7,823&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ego lite&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;page.snapshot({scope:"full_page"})&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;45,970&lt;/td&gt;
&lt;td&gt;11,867&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;agent-browser&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;snapshot&lt;/code&gt; (full a11y tree)&lt;/td&gt;
&lt;td&gt;27,845&lt;/td&gt;
&lt;td&gt;7,368&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;agent-browser&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;snapshot -i&lt;/code&gt; (interactive only)&lt;/td&gt;
&lt;td&gt;13,734&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4,742&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chrome DevTools MCP&lt;/td&gt;
&lt;td&gt;&lt;code&gt;take_snapshot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;39,123&lt;/td&gt;
&lt;td&gt;13,418&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Playwright MCP&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;browser_snapshot&lt;/code&gt; (my July article-page measurement)&lt;/td&gt;
&lt;td&gt;28,874&lt;/td&gt;
&lt;td&gt;~7,400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ubrowser&lt;/td&gt;
&lt;td&gt;compact format (100-element HN page, July)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;~1,720&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;ego's snapshot is &lt;em&gt;not&lt;/em&gt; compact. It's a faithful accessibility tree with every &lt;code&gt;table &amp;gt; table_row &amp;gt; table_cell&lt;/code&gt; level spelled out — Hacker News is nested tables all the way down, and the ego dump of the viewport alone costs more than agent-browser's interactive-only view of the whole page, and only 40% less than Chrome DevTools MCP's, which I called the most expensive in the business &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-21-claude-code-screenshot-mcp-frontend-debugging/" rel="noopener noreferrer"&gt;back in July&lt;/a&gt;. What it does have is refs (&lt;code&gt;[ref=7, loc=href:/newest]&lt;/code&gt;) with stable-locator hints, iframe contents inline, and — per the v2.0.0 changelog — refs that survive &lt;code&gt;page.evaluate&lt;/code&gt; calls, which was a real bug before September 9.&lt;/p&gt;

&lt;p&gt;So why did the full-access run cost 151 tokens instead of 7,823? Because the agent never asked for a snapshot. It wrote &lt;code&gt;page.evaluate(() =&amp;gt; [...document.querySelectorAll(".athing")].slice(0,5).map(...))&lt;/code&gt; and got back a five-line JSON array. That is the &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-21-claude-code-screenshot-mcp-frontend-debugging/" rel="noopener noreferrer"&gt;65-token &lt;code&gt;evaluate_script&lt;/code&gt; pattern&lt;/a&gt; I've been preaching since July, and ego's contribution is to make it the &lt;em&gt;default&lt;/em&gt; path rather than the expert path: the skill tells the model to write code, the runtime lets a whole navigate-wait-extract sequence ship in one process, and the snapshot is what you fall back to when you don't know the page. That's the correct design. It's the same insight ubrowser had — batch the steps, ship only the answer — with an actual company behind it and a screenshot tool.&lt;/p&gt;

&lt;p&gt;The flip side showed up on docs.python.org. When the page is unfamiliar, the agent has to snapshot, and ego's snapshot of that page cost 8,300 chars per look; the docs site has three &lt;code&gt;input[name=q]&lt;/code&gt; search boxes (one hidden), so &lt;code&gt;fill("css=input[name=q]")&lt;/code&gt; threw &lt;code&gt;matched 3 elements&lt;/code&gt;, the &lt;code&gt;@8&lt;/code&gt; ref went stale after a re-snapshot, and the default-mode run spiraled to 30 turns. To be fair to ego, its error messages were the best of the three tools — "matched 3 elements (2 visible, 1 hidden). Candidates: 1. input 'Quick search' (hidden)…" tells the model exactly what to do next, which is what I dinged ubrowser for lacking. But the general rule from this series stands: &lt;strong&gt;the browser tool doesn't decide your token bill; the agent's habit of asking for facts instead of trees does.&lt;/strong&gt; ego nudges the habit in the right direction and then hands you a tree as fat as anyone's when you ask for one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8wj8q9zauost3ly14wbc.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8wj8q9zauost3ly14wbc.webp" alt="ego lite page.screenshot of Hacker News from inside an agent task space, 1512x738 CSS pixels" width="800" height="391"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  "Your logged-in browser" — half true, and the half that's false matters
&lt;/h2&gt;

&lt;p&gt;The headline feature is the one the first paragraph tripped over, so let me lay out exactly what happened. ego lite's onboarding found two browsers on my Mac and imported both: Chrome's profile became ego's &lt;strong&gt;Default&lt;/strong&gt; (&lt;code&gt;Bruce&lt;/code&gt;), Edge's became &lt;strong&gt;Profile 1&lt;/strong&gt; (&lt;code&gt;he bruce&lt;/code&gt;). I live in Edge; that's where GitHub is signed in. &lt;code&gt;taskSpace()&lt;/code&gt; with no arguments creates the space on the default profile, and the skill explicitly tells the agent not to "inspect or select profiles unless the user explicitly requests a particular profile." So the agent did the right thing by its instructions and got the wrong answer. Naming the profile — &lt;code&gt;taskSpace("…", { profileId: "Profile 1" })&lt;/code&gt; — produced &lt;code&gt;logged in as heyuan110&lt;/code&gt; in 36 seconds with 4 browser calls. If you have exactly one browser and one profile, you'll never see this. If you have two, or a work and a personal profile, you will, and the fix is one line the docs don't mention.&lt;/p&gt;

&lt;p&gt;The bigger question is what "own Space" means for the data. The docs say each task space gets "its own native BrowserContext for cookies and storage," which any Playwright user reads as &lt;em&gt;isolation&lt;/em&gt;. So I tested it: in task space A, on example.com, set &lt;code&gt;document.cookie = "isotest=fromA"&lt;/code&gt; and a localStorage key; create task space B on the same profile, load example.com, read them back.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"spaceA"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"spaceB"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"seenInA"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"cookie"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"isotest=fromA"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"ls"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"fromA"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"seenInB"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"cookie"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"isotest=fromA"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"ls"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"fromA"&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Space B saw everything. Spaces isolate &lt;strong&gt;tabs, focus, and the window&lt;/strong&gt; — the thing you see in the overview screenshot, the thing that keeps an agent from hijacking your mouse — but on one profile they are one cookie jar. That's exactly why the agent lands on GitHub already logged in, and exactly why you should think of a Space as a private tab group over shared state, not a sandbox. Two open issues sharpen the point: &lt;a href="https://github.com/citrolabs/ego-lite/issues/303" rel="noopener noreferrer"&gt;#303&lt;/a&gt; reports &lt;code&gt;cdp("Network.clearBrowserCookies")&lt;/code&gt; from an agent task space logging the &lt;em&gt;user&lt;/em&gt; out of every site in the main Space, and &lt;a href="https://github.com/citrolabs/ego-lite/issues/319" rel="noopener noreferrer"&gt;#319&lt;/a&gt; shows &lt;code&gt;Storage.getCookies&lt;/code&gt; from a task space on a secondary profile returning the default profile's full jar — 1,105 cookies, other profiles' auth included. &lt;a href="https://github.com/citrolabs/ego-lite/issues/315" rel="noopener noreferrer"&gt;#315&lt;/a&gt; puts the whole thing in one sentence: model-generated JavaScript runs in a privileged Node process with unrestricted raw CDP and unrestricted outbound HTTP, so a prompt injection on any page the agent visits is a prompt injection with your session cookies in scope. To their credit, the ego team answered all three within a week or two, and the reply on #319 is the most honest line in the whole project: "A space is not the same as a profile, and spaces do not isolate cookies," with per-profile space creation promised for 0.5.0.x — which is the &lt;code&gt;profileId&lt;/code&gt; option I used above, so that part shipped. The raw-CDP exposure in #315 is acknowledged and not yet fixed, and the repo's private vulnerability reporting was returning 403 when the reporter tried it.&lt;/p&gt;

&lt;p&gt;Isolation between users and agents, though, held up under everything I threw at it. Six task spaces opened and loaded six real pages in 8.2 seconds without touching my tabs; the app cold-started from fully quit to first snapshot in 2.1 seconds without stealing focus from my terminal (the &lt;a href="https://github.com/citrolabs/ego-lite/issues/284" rel="noopener noreferrer"&gt;focus-stealing bug&lt;/a&gt; filed against August builds didn't reproduce on 0.5.0.28); and the Spaces overview shows each agent's space labeled "running" with a live thumbnail, with a click to take it over. The CSDN piece quotes ego's own numbers for six concurrent tasks — "0.9 GB, 6 processes" for Space mode versus "15 GB, 84 processes" for six separate browsers — and for once a vendor number survived contact: my six spaces added exactly &lt;strong&gt;6 processes and 0.91 GB&lt;/strong&gt; of RSS (17 → 23 processes, 0.61 → 1.52 GB), and the browser reclaimed most of it within a minute of idling (back to 0.68 GB). I didn't measure the 84-process alternative; I believe it, and it's the least interesting comparison here, because nobody runs six full browsers for six tasks anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rough edges, in the order they cost me time
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The &lt;code&gt;-e&lt;/code&gt; flag blocks on stdin.&lt;/strong&gt; &lt;code&gt;ego-browser nodejs -e '…'&lt;/code&gt; never returns if its stdin is an open pipe — which it is inside most agent harnesses and CI runners. Three of my direct invocations hung until killed before I isolated it; &lt;code&gt;&amp;lt; /dev/null&lt;/code&gt; fixes it. Heredocs are fine because the heredoc closes stdin.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The onboarding gate with no window.&lt;/strong&gt; Covered above; ten minutes, &lt;code&gt;ego://onboarding&lt;/code&gt;, press Done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extensions import disabled.&lt;/strong&gt; Every Chrome extension arrived with "has been disabled — accept new permissions," and the app's first launch popped a permission bubble per extension in the toolbar. Fine for you; irrelevant to the agent, which doesn't get extension UI anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ownership flips silently.&lt;/strong&gt; After my cleanup script called &lt;code&gt;finish({ keep: [] })&lt;/code&gt; on eight agent spaces, one survived as &lt;code&gt;ownership: "user"&lt;/code&gt; — the GitHub tab had flipped it, per &lt;a href="https://github.com/citrolabs/ego-lite/issues/314" rel="noopener noreferrer"&gt;#314&lt;/a&gt;. It's harmless but it's a leaked renderer, and &lt;a href="https://github.com/citrolabs/ego-lite/issues/270" rel="noopener noreferrer"&gt;#270&lt;/a&gt; says idle task spaces pile up across sessions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quitting loses your tabs.&lt;/strong&gt; When I quit ego lite to test cold start, it came back to a fresh New Tab; the three tabs I'd had open were gone. Chromium's "continue where you left off" is off by default here, and I'd turn it on before making this a daily driver.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;macOS only, and the roadmap says "Planned" with no date for Windows and Linux.&lt;/strong&gt; &lt;a href="https://github.com/citrolabs/ego-lite/issues/203" rel="noopener noreferrer"&gt;#203&lt;/a&gt;, &lt;a href="https://github.com/citrolabs/ego-lite/issues/345" rel="noopener noreferrer"&gt;#345&lt;/a&gt;, and a &lt;a href="https://github.com/citrolabs/ego-lite/issues/374" rel="noopener noreferrer"&gt;WSL2 request&lt;/a&gt; are the most-upvoted feature issues. If you're on Windows, this article is a preview, not a recommendation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The skills.sh scanner flagged the skill "High Risk"&lt;/strong&gt; at install. Having read the skill, the flag is about capability, not malice: it teaches the model to run arbitrary JavaScript against a browser holding your cookies. That is the product. Decide accordingly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two things I expected to be rough and weren't: the skill's writing is excellent (the "use exactly one TaskSpace, print its id, resume it later" discipline saved the agent from tab sprawl in every run), and the v2 API's action receipts — popups, dialogs, downloads returned as data instead of surprises — are more thoughtful than agent-browser's.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict: use it, on a dedicated profile, with full access, on a Mac
&lt;/h2&gt;

&lt;p&gt;My one-line judgment: &lt;strong&gt;ego lite is the first agent browser where "your logged-in browser" is a two-command install and the token win is measurable rather than projected — but the win only appears when you run Claude Code in full-access mode, and the isolation it implies does not exist at the cookie layer, so it belongs on a dedicated profile with only the logins you'd hand an intern.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The comparison table I'd actually use, across everything measured in this series:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;ego lite&lt;/th&gt;
&lt;th&gt;agent-browser&lt;/th&gt;
&lt;th&gt;Chrome DevTools MCP&lt;/th&gt;
&lt;th&gt;Playwright MCP&lt;/th&gt;
&lt;th&gt;ubrowser&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Interface&lt;/td&gt;
&lt;td&gt;JS via &lt;code&gt;ego-browser nodejs&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;CLI, one command per action&lt;/td&gt;
&lt;td&gt;MCP tools&lt;/td&gt;
&lt;td&gt;MCP tools&lt;/td&gt;
&lt;td&gt;MCP tools (batch)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Your logins&lt;/td&gt;
&lt;td&gt;✅ imported profile, shared jar&lt;/td&gt;
&lt;td&gt;❌ (unless you &lt;code&gt;connect&lt;/code&gt; to your Chrome)&lt;/td&gt;
&lt;td&gt;✅ attach to real Chrome&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best measured run (HN)&lt;/td&gt;
&lt;td&gt;16.4s / 2 turns / $0.12&lt;/td&gt;
&lt;td&gt;19.0s / 6 turns / $0.18&lt;/td&gt;
&lt;td&gt;26.4s / 6 turns / $0.23&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1 call / 641 tokens (July)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Snapshot cost, HN front page&lt;/td&gt;
&lt;td&gt;7,823 tokens&lt;/td&gt;
&lt;td&gt;4,742 (&lt;code&gt;-i&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;13,418&lt;/td&gt;
&lt;td&gt;~7,400 (July, article page)&lt;/td&gt;
&lt;td&gt;~1,720 (July)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Default-mode Claude Code&lt;/td&gt;
&lt;td&gt;❌ heredoc JS blocked&lt;/td&gt;
&lt;td&gt;⚠️ &lt;code&gt;eval&lt;/code&gt; blocked&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Screenshot / evaluate&lt;/td&gt;
&lt;td&gt;✅ / ✅&lt;/td&gt;
&lt;td&gt;✅ / ✅&lt;/td&gt;
&lt;td&gt;✅ / ✅&lt;/td&gt;
&lt;td&gt;✅ / ✅&lt;/td&gt;
&lt;td&gt;❌ / ❌&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platforms&lt;/td&gt;
&lt;td&gt;macOS&lt;/td&gt;
&lt;td&gt;macOS / Linux / Windows&lt;/td&gt;
&lt;td&gt;all&lt;/td&gt;
&lt;td&gt;all&lt;/td&gt;
&lt;td&gt;all&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maintained&lt;/td&gt;
&lt;td&gt;v2.0.0 on 2026-09-10&lt;/td&gt;
&lt;td&gt;active&lt;/td&gt;
&lt;td&gt;active&lt;/td&gt;
&lt;td&gt;active&lt;/td&gt;
&lt;td&gt;abandoned Dec 2025&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And the decision rule, since that's what you'll screenshot:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;If you…&lt;/th&gt;
&lt;th&gt;Use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Are on a Mac, run Claude Code with bypass permissions, and your tasks need &lt;em&gt;your&lt;/em&gt; logins (dashboards, admin panels, your own SaaS)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;ego lite&lt;/strong&gt;, on a dedicated profile that's logged into only those sites&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Run agents in default permission mode and want the cheapest clean-room browsing&lt;/td&gt;
&lt;td&gt;agent-browser (&lt;code&gt;snapshot -i&lt;/code&gt; + &lt;code&gt;eval&lt;/code&gt;), or &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-playwright-cli-skill-zero-token-automation/" rel="noopener noreferrer"&gt;Playwright CLI in a skill&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Are debugging your own frontend&lt;/td&gt;
&lt;td&gt;Chrome DevTools MCP with targeted &lt;code&gt;evaluate_script&lt;/code&gt;, as before&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Need real isolation between the agent and your session&lt;/td&gt;
&lt;td&gt;Not ego lite. Separate browser, separate profile, or a cloud browser&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Are on Windows or Linux, or running in CI&lt;/td&gt;
&lt;td&gt;Not ego lite, until the roadmap moves&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three concrete setup rules, earned the hard way. One: create a fresh ego lite profile for agents and sign it into only what the agent needs — the shared cookie jar means "inherits your logins" also means "inherits your bank." Two: if you keep two browsers, check &lt;code&gt;profiles()&lt;/code&gt; once and put the right &lt;code&gt;profileId&lt;/code&gt; in your CLAUDE.md, or your agent will confidently report "not logged in." Three: never &lt;code&gt;-e&lt;/code&gt; without &lt;code&gt;&amp;lt; /dev/null&lt;/code&gt;, and never &lt;code&gt;cdp("Network.clear…")&lt;/code&gt; from a task space unless you enjoy re-authenticating everything.&lt;/p&gt;

&lt;p&gt;The bigger picture, for anyone following the series: the field has now split cleanly into two philosophies. Drivers (agent-browser, Playwright, the MCPs) give you a browser the agent owns and you don't; ego lite gives you a browser you own and the agent borrows. The second is more useful and more dangerous in exactly equal measure, and the token economics — the thing I started measuring in January — turned out to be the easy part. Both camps have converged on "write code, ship the answer, don't ship the DOM." What separates them now is whose cookies are in the jar.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is ego lite and how does it work with Claude Code?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A free, closed-source Chromium 152 fork for macOS with agent-owned Spaces inside your daily browser. Claude Code uses the &lt;code&gt;ego-browser&lt;/code&gt; skill: the agent writes JavaScript, the bundled native CLI sends it over Mojo IPC to an embedded Node runtime, and &lt;code&gt;page.evaluate()&lt;/code&gt; runs inside the page. My best run was one tool call, 16.4 seconds, $0.12.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is ego lite faster and cheaper than Vercel agent-browser?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In full-access mode, yes — 16.4s/$0.12 vs 19.0s/$0.18 on Hacker News; even on time and 25% cheaper on a multi-page docs search. In Claude Code's default permission mode it was slower, because heredoc JavaScript trips the shell sanitizer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do ego lite Spaces isolate the agent from my logged-in session?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. A cookie set in one task space was readable in another on the same profile, and issue #303 shows a task-space cookie clear logging out the main Space. Spaces isolate tabs and focus; give agents a dedicated profile.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does ego lite work on Windows or Linux?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not as of September 2026. Both are "Planned" on the roadmap with no date.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does the skill fail with "expansion obfuscation" in Claude Code?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because any JavaScript object literal in a heredoc puts a brace next to a quote, which Claude Code's Bash guard rejects in default permission mode. ego's docs say to use full access; the same docs task went from 30 turns to 7 when I did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-27-ubrowser-review/" rel="noopener noreferrer"&gt;ubrowser Review: Fastest Cheapest Browser Automation? Tested&lt;/a&gt; — the abandoned repo that got the cost model right first; ego lite is what it looked like it wanted to become&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-01-28-claude-code-browser-automation/" rel="noopener noreferrer"&gt;Browser Automation in Claude Code: 5 Tools Compared&lt;/a&gt; — the January field map this review updates&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-21-claude-code-screenshot-mcp-frontend-debugging/" rel="noopener noreferrer"&gt;Claude Code Screenshot MCP Setup: Browser Automation 2026&lt;/a&gt; — where the 13,418-token snapshot and the 65-token evaluate come from&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-01-13-vercel-agent-browser/" rel="noopener noreferrer"&gt;Vercel Agent Browser: AI-Native Browser Automation CLI Tool&lt;/a&gt; — the baseline tool in this test, back when it launched&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-playwright-cli-skill-zero-token-automation/" rel="noopener noreferrer"&gt;Playwright CLI + Skills: 0-Token Browser Automation&lt;/a&gt; — the pattern to use if you want ego-class economics without ego's cookie jar&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Series Navigation
&lt;/h2&gt;

&lt;p&gt;This is Part 7 of the &lt;strong&gt;Browser Automation for AI Agents&lt;/strong&gt; series — the arc from headless drivers, to attaching to your real browser, to cutting the token bill, to sharing your logged-in session:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-01-13-vercel-agent-browser/" rel="noopener noreferrer"&gt;Vercel Agent Browser&lt;/a&gt; — a snapshot-driven CLI built for agents, not test suites&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-01-28-claude-code-browser-automation/" rel="noopener noreferrer"&gt;Browser Automation in Claude Code: 5 Tools Compared&lt;/a&gt; — the field map — token cost, speed, stability&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-03-17-chrome-devtools-mcp-guide/" rel="noopener noreferrer"&gt;Chrome DevTools MCP Setup 2026&lt;/a&gt; — attaching to your real browser, and the port 9222 traps&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-playwright-cli-skill-zero-token-automation/" rel="noopener noreferrer"&gt;Playwright CLI + Skills: 0-Token Automation&lt;/a&gt; — the pattern that removes the MCP token tax&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-21-claude-code-screenshot-mcp-frontend-debugging/" rel="noopener noreferrer"&gt;Claude Code Screenshot MCP Setup&lt;/a&gt; — the frontend debugging loop, 10,220 vs 65 tokens&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-27-ubrowser-review/" rel="noopener noreferrer"&gt;ubrowser Review&lt;/a&gt; — the right design trapped in an abandoned repo&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;This article&lt;/strong&gt;: ego lite Review — handing agents your logged-in session through Spaces&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-10-ego-lite-browser-review/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>egolite</category>
      <category>claudecode</category>
      <category>browserautomation</category>
      <category>agents</category>
    </item>
    <item>
      <title>Free AI Agent Courses Fall 2026: Stanford, CMU, MIT Compared</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:47:34 +0000</pubDate>
      <link>https://dev.to/bruce_he/free-ai-agent-courses-fall-2026-stanford-cmu-mit-compared-1833</link>
      <guid>https://dev.to/bruce_he/free-ai-agent-courses-fall-2026-stanford-cmu-mit-compared-1833</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0b556pf9zf84gcvvaxad.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0b556pf9zf84gcvvaxad.webp" alt="Free AI agent courses Fall 2026 from Stanford, CMU, and MIT compared" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Five universities are running or have just published &lt;strong&gt;free AI agent courses for Fall 2026&lt;/strong&gt;, and three of them you cannot watch. That's the fact to start from, because the search results won't tell you. Stanford CS146S, Stanford CS329Z, Stanford CS329A, CMU 11-768, and MIT's multimodal course all have public websites, syllabi, and reading lists; only two of them have lecture recordings you can play today, and only one is recording &lt;em&gt;this&lt;/em&gt; term as it happens.&lt;/p&gt;

&lt;p&gt;I've been covering Stanford CS146S since February, and its &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;overview&lt;/a&gt;, &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-02-cs146s-study-guide/" rel="noopener noreferrer"&gt;study guide&lt;/a&gt;, and &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;Fall 2026 follow-along&lt;/a&gt; are the most-read pages on this blog. That's why this hub exists: readers keep asking "which of these should I follow," and the answer isn't "all five" or "the Stanford one." As of September 9, 2026, I pulled every schedule from the official sites (for CMU, from the site's JavaScript bundle, because the page is client-rendered), counted the videos on every playlist, and read every grading table. What follows is the comparison, the honest inventory of what's actually free, and the stack I'd run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Five Courses at a Glance
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The five aren't peers; they're three different kinds of course wearing the same "AI agents" label.&lt;/strong&gt; CS146S is a practitioner course about working &lt;em&gt;inside&lt;/em&gt; coding agents. CMU 11-768 and CS329Z are graduate courses about &lt;em&gt;building and training&lt;/em&gt; agents. CS329A is a research seminar on agents that improve themselves. MIT's course is a multimodal ML course that happens to have an agents week. Here's the whole picture in one table:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Course&lt;/th&gt;
&lt;th&gt;Term and dates&lt;/th&gt;
&lt;th&gt;Instructors&lt;/th&gt;
&lt;th&gt;Live video?&lt;/th&gt;
&lt;th&gt;Slides / materials&lt;/th&gt;
&lt;th&gt;Assignments public?&lt;/th&gt;
&lt;th&gt;Prereqs&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://themodernsoftware.dev/" rel="noopener noreferrer"&gt;Stanford CS146S&lt;/a&gt; The Modern Software Developer&lt;/td&gt;
&lt;td&gt;Fall 2026, Sep 22 to Dec 3, Tue/Thu&lt;/td&gt;
&lt;td&gt;Mihail Eric&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;No&lt;/strong&gt; (none, ever)&lt;/td&gt;
&lt;td&gt;Syllabus public; Fall 2025 decks public; 2026 decks TBD&lt;/td&gt;
&lt;td&gt;Fall 2025 assignments on GitHub; 2026 TBD&lt;/td&gt;
&lt;td&gt;Programming&lt;/td&gt;
&lt;td&gt;Engineers who use Claude Code / Codex daily and want the practice canon&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://www.cmu-agents.com/" rel="noopener noreferrer"&gt;CMU 11-768&lt;/a&gt; AI Agents&lt;/td&gt;
&lt;td&gt;Fall 2026, Aug 25 to Dec 3, Tue/Thu 3:30 to 4:50 PM ET&lt;/td&gt;
&lt;td&gt;Graham Neubig, Daniel Fried&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Yes&lt;/strong&gt;, YouTube, updating (4 of 28 as of Sep 9)&lt;/td&gt;
&lt;td&gt;Slides PDF per lecture (6 so far), readings&lt;/td&gt;
&lt;td&gt;Assignment 1 starter repo public&lt;/td&gt;
&lt;td&gt;Trained an LM before (11-667/11-711 level)&lt;/td&gt;
&lt;td&gt;Builders who want the whole loop: harness, eval, RL training&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://cs329z.stanford.edu/" rel="noopener noreferrer"&gt;Stanford CS329Z&lt;/a&gt; Engineering AI Agents&lt;/td&gt;
&lt;td&gt;Fall 2026, Sep 23 to Dec 2, Mon/Wed 1:30 to 2:50 PM PT&lt;/td&gt;
&lt;td&gt;Diyi Yang, Michael Ryan, John Yang&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;No&lt;/strong&gt; (Canvas-only)&lt;/td&gt;
&lt;td&gt;Syllabus and schedule public; slides TBD&lt;/td&gt;
&lt;td&gt;Descriptions public; starter code not&lt;/td&gt;
&lt;td&gt;CS224N-level NLP&lt;/td&gt;
&lt;td&gt;Framework-literate builders who want a from-scratch harness plus eval discipline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://cs329a.stanford.edu/" rel="noopener noreferrer"&gt;Stanford CS329A&lt;/a&gt; Self-Improving AI Agents&lt;/td&gt;
&lt;td&gt;Autumn 2025, Sep 22 to Dec 5, 2025 (replay)&lt;/td&gt;
&lt;td&gt;Azalia Mirhoseini, Aakanksha Chowdhery&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Yes&lt;/strong&gt;, 9 lectures on Stanford Online (posted Aug 2026)&lt;/td&gt;
&lt;td&gt;Schedule public; slides not linked&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;ML basics; RL helps&lt;/td&gt;
&lt;td&gt;Anyone who wants the test-time compute / verifier / RL frontier explained by people who shipped PaLM and Gemini&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://mit-mi.github.io/mmai-course/spring2026/" rel="noopener noreferrer"&gt;MIT MAS.S60 / 6.S985&lt;/a&gt; Modeling: Multimodal AI&lt;/td&gt;
&lt;td&gt;Spring 2026, Feb 3 to May 12, 2026 (replay)&lt;/td&gt;
&lt;td&gt;Paul Liang and three co-instructors&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Yes&lt;/strong&gt;, 13 of 28 sessions on YouTube&lt;/td&gt;
&lt;td&gt;Slides for every lecture; application guest lectures slides-only&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Deep learning basics&lt;/td&gt;
&lt;td&gt;People building multimodal or GUI agents; not a general agents course&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two rows deserve a second look. CS146S and CMU 11-768 are both live this fall, both Tuesday/Thursday, and both run through December 3. If you follow both, you'll be doing four lecture-topics a week from September 22 on. And the two Stanford courses that are actually running this fall, CS146S and CS329Z, are the two with no public video. The "Stanford AI agents course" people search for is, this term, a reading assignment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which You Can Watch vs Which You Can Only Read
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;"Free" here means three different things, and the difference decides whether you can follow along.&lt;/strong&gt; I sort the five into three tiers by what you can press play on as of September 9, 2026:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tier 1, watchable live: CMU 11-768.&lt;/strong&gt; The &lt;a href="https://www.youtube.com/playlist?list=PLSN0qpDfUvTM" rel="noopener noreferrer"&gt;Fall 2026 playlist&lt;/a&gt; on Graham Neubig's channel had four lectures (57 to 76 minutes each) when I checked, covering the "What is an agent," tool use, long-context, and skills-and-memory sessions from August 25 to September 3. Slides are linked as PDFs for six sessions. The playlist had 1,413 views. That number matters: this is the least-discovered of the five, and the only one recording as it goes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tier 2, watchable as a replay: CS329A and MIT.&lt;/strong&gt; Stanford Online put all nine CS329A lectures on YouTube in August 2026, ten months after the course ran; the &lt;a href="https://www.youtube.com/playlist?list=PLangBM27OtEA" rel="noopener noreferrer"&gt;playlist&lt;/a&gt; had 68,302 views when I counted, and the lectures run 63 to 75 minutes. Nine is not the whole course, though. The official schedule lists 18 sessions including guest lectures from Denny Zhou and Thang Luong (Google DeepMind), Misha Laskin (Reflection AI), and Danny Driess (Physical Intelligence); none of those guest sessions are in the published set. MIT's course has &lt;a href="https://www.youtube.com/playlist?list=PLWzwL390pi04" rel="noopener noreferrer"&gt;13 videos&lt;/a&gt; out of 28 sessions; the application lectures (manufacturing, design, cities, transportation) and the tutorials are slides-only.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tier 3, materials only: CS146S and CS329Z.&lt;/strong&gt; CS146S has never released video, a point I documented in the &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;follow-along post&lt;/a&gt;; what you get is the syllabus and, based on the 2025 pattern, Google Slides decks within days of each lecture. CS329Z's course page is explicit: cameras in the back of the room record instructor presentations, and recordings "can be accessed by logging into the course Canvas site." Everything else about CS329Z is public, including a week-by-week schedule and the two homework descriptions, which are good enough to steal as self-study projects.&lt;/p&gt;

&lt;p&gt;Here's how the five terms overlap on the calendar, which is the constraint nobody mentions:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-09-free-ai-agent-courses-fall-2026/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The practical reading: from September 22 to December 3 there are three live courses on the same six weekday slots. Anything you add on top is a replay, and replays can wait until January.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Stack I'd Actually Run
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Run CMU 11-768 as the spine, CS146S as the practice track, and CS329A as the theory replay; skip or defer the other two.&lt;/strong&gt; That's the whole recommendation. Here's why it comes out that way and what it costs per week.&lt;/p&gt;

&lt;p&gt;CMU is the spine because it's the only course that gives you all three pieces on the same week: a lecture you can watch, slides you can search, and a &lt;a href="https://github.com/cmu-agents/assignment-1" rel="noopener noreferrer"&gt;public assignment&lt;/a&gt; you can run. The syllabus goes further than any of the others too. Weeks 1 to 3 are capabilities (tool use, context management, skills and memory, planning), weeks 4 to 6 are domains plus &lt;em&gt;training&lt;/em&gt; (SFT, RL basics, advanced RL, RL systems), then safety, frameworks (an OpenHands session and a LangGraph session), interaction, and search. Twenty-eight sessions including a fall break and two guest lectures in November. Assignment 1 asks you to build a ReAct harness from scratch that fixes a bug in a chess app and then plays chess through tool calls, with context compaction in the middle; that's a better exercise than anything in the CS146S 2025 assignment set, and it was posted publicly on August 31.&lt;/p&gt;

&lt;p&gt;CS146S is the practice track because it teaches what CMU deliberately skips: how to work inside Claude Code, Codex, and Cursor as a professional. MCP, agent skills, CLAUDE.md and hooks, agent-ready repos, background agents, and the software factory. No video, but each week's topic maps to a tool you can run that night, and I've already written the &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;week-by-week exercise plan&lt;/a&gt;. If you've read my &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-05-loop-engineering/" rel="noopener noreferrer"&gt;loop engineering&lt;/a&gt; or &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;CLI skills vs MCP&lt;/a&gt; posts, you already know the shape of that syllabus.&lt;/p&gt;

&lt;p&gt;CS329A is the theory replay because it explains the frontier the other two only gesture at: why test-time compute scales, what a verifier is and why it has to be robust, how RL post-training turned chatbots into agents, and how to evaluate long-horizon tasks. Nine lectures, 10.5 hours, no homework you can submit. Watch it in November when the two live courses hit their project phases and the lecture load drops.&lt;/p&gt;

&lt;p&gt;The two I'd defer: CS329Z, because without recordings it's a syllabus to mine, not a course to follow (more on that below), and MIT, because it's a multimodal ML course; agents get one lecture (April 7) and one tutorial (April 23) out of 28 sessions, and if you need multimodal fusion and alignment you'll know it.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-09-free-ai-agent-courses-fall-2026/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Hours per week, with the assumptions stated. CMU live: two lectures at 60 to 76 minutes, two to three readings (the tool-use lecture alone lists four papers and 26 reference links), and the assignment; 6 to 8 hours. CS146S without grades: one deck, one reading, one exercise; 5 to 6 hours, which is the number readers of the follow-along post told me they can sustain. CS329A replay: 1.75 hours per lecture including notes; 1 to 2 per week. Spine plus practice track is 11 to 14 hours a week. That's a real commitment, and if you can't make it, drop to CMU alone and read the CS146S decks on weekends. Don't do the reverse: CS146S without CMU leaves you fluent in tools and unable to explain why a harness works.&lt;/p&gt;

&lt;h2&gt;
  
  
  Course by Course: What Each Is Really For
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;CMU 11-768 is the closest thing to a complete agents curriculum any university has published, and it's the one nobody's talking about.&lt;/strong&gt; Neubig maintains OpenHands and Fried came from Meta's agent work, so the frameworks weeks aren't a survey; they're the authors explaining their own decisions. Grading is 40% individual assignments (harness 10%, eval 15%, training 15%), 10% "lecture highlights," and 50% a team research project with a poster on December 1 to 3. Two things to know before you start. The assignment defaults to DeepSeek-V4-Flash through an OpenAI-compatible endpoint and runs agent actions in Modal sandboxes; enrolled students get credits, you'll pay your own, and I haven't run it end to end myself, so budget for that. And Assignment 2 (eval) and 3 (training) were due September 24 and October 22 but hadn't been posted publicly when I checked; the site says starter materials "will be posted here as they become available."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stanford CS146S is the course for the working engineer, and the one I'd tell most readers of this blog to start with.&lt;/strong&gt; I've written three full posts on it, so I'll be brief: Fall 2026 is a rewrite of the 2025 syllabus around MCP, agent skills, CLAUDE.md and hooks, agent-ready codebases, background agents, and the software factory, with 30% of the grade now on open-source contributions and eight guest sessions from the people who build Cursor, Claude Code, Factory, Cognition, Semgrep, Cloudflare's agent stack, and Replit. Its weakness is exactly what CMU has: no video, no assignment you can grade yourself against, and a shallow treatment of how agents are trained. Its strength is that every week's topic is something you'll use at work on Monday.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stanford CS329Z is the syllabus to mine, not the course to follow.&lt;/strong&gt; Diyi Yang, Michael Ryan (DSPy core contributor), and John Yang (SWE-bench and SWE-agent co-creator) teach it from scratch: LLMs for builders, RAG, tool use, frameworks (DSPy, LangGraph, LlamaIndex, MCP, litellm), design patterns and scaffolds, memory, multi-agent, optimization, then coding agents and proactive agents in the back half. The two homeworks are the prize. HW1 is "build a company's internal AI assistant from scratch, with no agent frameworks: just a chat-completion call and code you write yourself," released October 5, due October 30. HW2 is an evaluation suite with code-based graders, at least one LLM-as-judge eval, and benchmark tasks built with the course's "4-tuple framework," due November 20. Both descriptions are public. Neither starter kit is. Do them anyway, on the course's dates, and you'll have reproduced 20% of a Stanford grade without Canvas. The recordings, quizzes, and the "Making Life at Stanford Better with Agents" project (50%) stay behind the login.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stanford CS329A is the lecture series to watch when you want the why.&lt;/strong&gt; Mirhoseini and Chowdhery (PaLM, Gemini, AlphaChip between them) organize the course around one question: how does an agent keep improving through interaction with its environment? The nine public lectures are the instructor sessions: course overview, test-time compute scaling, robust verification, learning from feedback with tools and code, planning and multi-step reasoning, train-time scaling and RL, self-improvement and deep research agents, agentic evaluations and long-horizon tasks, future research areas. If you've read my &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;agentic loops&lt;/a&gt; post and wanted the training-side counterpart, this is it. Homework and the research project (35%) were for enrolled students, and the course hasn't announced a Fall 2026 run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MIT's course is excellent and mostly not about agents.&lt;/strong&gt; Paul Liang's Spring 2026 offering is a multimodal ML course: representation, fusion, alignment, large multimodal models, generation, reasoning, transfer, then applications with Media Lab and Sloan co-instructors. Thirteen of 28 sessions have video; the four application lectures and the three tutorials don't. The agents content is week 10.1 (multimodal interaction, with VisualWebArena, Mind2Web, and OpenVLA on the reading list), the April 23 agents tutorial (slides only), and week 14.1 on self-evolving AI. If you're building GUI or vision agents, watch weeks 4, 5, 9, and 10 and skip the rest. If you aren't, this isn't your course, and it's fine to say so.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Not Free
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Everything that involves a human looking at your work stays behind enrollment, at all five.&lt;/strong&gt; The tables above say "free," and the materials are, but here's the honest list of what you don't get, because it's the same list every time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Grades and feedback.&lt;/strong&gt; CMU's 50% project, CS329Z's 50% project, CS329A's 35% project, CS146S's 50% final project. No one reads your work. Substitute: publish it and ask the tool's community to review, which is worse but not nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assignment starter code, partially.&lt;/strong&gt; CMU posted Assignment 1; Assignments 2 and 3 weren't public on September 9. CS329Z's HW1 and HW2 are descriptions only. CS146S's 2026 assignments weren't posted; the 2025 set is on GitHub. CS329A and MIT posted none.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API and compute credits.&lt;/strong&gt; CMU's assignment expects a Modal account and an LLM key; the course arranges credits for students. Running the harness assignment on your own DeepSeek or OpenAI-compatible key is the one line item in this post that costs money, and I can't tell you how much because I haven't run it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Discussion and office hours.&lt;/strong&gt; Piazza, Ed, Canvas, and TA hours are enrolled-only everywhere. CS329Z even keeps its quizzes closed-book and individual.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guest lectures, mostly.&lt;/strong&gt; CS329A's DeepMind and Reflection AI guests aren't in the nine public videos. CS146S's eight guests have no video at all. CMU's two November guest lectures aren't named yet, and whether they're recorded is up to the speaker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recordings, for CS329Z and CS146S.&lt;/strong&gt; Canvas-only and nonexistent, respectively. If you need to &lt;em&gt;watch&lt;/em&gt; a Stanford agents course, CS329A is the only one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What is free is more than it sounds: every syllabus, every reading list, six CMU slide decks and four lectures with more coming, 17 CS146S decks from 2025, CS329A's nine lectures, MIT's 13, and one runnable CMU assignment. Ten weeks of that, done on the calendar the classes are on, is more than most enrolled students will finish.&lt;/p&gt;

&lt;h2&gt;
  
  
  Following From Outside the US
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;None of the live sessions are attendable remotely, so time zones only affect when materials land.&lt;/strong&gt; CMU lectures are Tuesday/Thursday 3:30 to 4:50 PM Eastern; the recordings have been showing up on the playlist within days, not hours. CS329Z is Monday/Wednesday 1:30 to 2:50 PM Pacific, but you can't watch it anyway. CS146S publishes days but not times. Daylight saving ends November 1 in the US, so any "next morning" habit shifts an hour after that.&lt;/p&gt;

&lt;p&gt;The real access question for readers in China is YouTube and Google Slides, and the answer is that four of the five depend on them: CMU's recordings and CS329A's and MIT's lectures are YouTube-only, and CS146S's decks are Google Slides. CMU's slides are plain PDFs on the course domain, which is the one exception. I go into the download-and-cache workflow in the Chinese version of this post; the short version is that the PDFs and playlists are all public, so a one-time fetch per week is enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Page Stops
&lt;/h2&gt;

&lt;p&gt;This is the routing page. The deep dives are where the week-by-week plans live: the &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;CS146S follow-along&lt;/a&gt; for the practice track, the &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-08-cmu-11-768-ai-agents-course/" rel="noopener noreferrer"&gt;CMU 11-768 deep dive&lt;/a&gt; for the spine, and the &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-04-stanford-cs329z-engineering-ai-agents/" rel="noopener noreferrer"&gt;CS329Z breakdown&lt;/a&gt; for the homework you can steal. I'll update the table when CMU posts Assignments 2 and 3, when CS146S posts its 2026 decks, and if CS329Z or CS329A releases any video; the "as of" dates in the text are the tell.&lt;/p&gt;

&lt;p&gt;Three things I couldn't verify and want on the record: whether CMU's November guest lectures will be recorded, the cost of running CMU Assignment 1 on your own API key, and whether CS329A will run again in 2026-27. If you're enrolled in any of these and know, the comments are open.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;CS146S Fall 2026: How to Watch and Follow Along Free&lt;/a&gt;: the calendar and week-by-week exercise plan for the practice track&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;Stanford CS146S: The Modern Software Developer, 2026 Guide&lt;/a&gt;: the full ten-week breakdown and guest lineup&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-02-cs146s-study-guide/" rel="noopener noreferrer"&gt;CS146S Study Guide 2026: Lecture-by-Lecture Notes and Workbook&lt;/a&gt;: the two-week core route if you can't do the quarter&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;Agentic Loops 2026: Self-Looping AI Agents Explained&lt;/a&gt;: the ReAct loop CMU's Assignment 1 asks you to build&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-05-loop-engineering/" rel="noopener noreferrer"&gt;Loop Engineering: Building the Cage Your AI Agent Runs In&lt;/a&gt;: the discipline CS146S weeks 4, 6, and 8 are about&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;MCP vs Skills: Why CLI + Skill Wins the Agent Toolchain&lt;/a&gt;: the argument behind CS146S week 3 and CMU lecture 4&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-09-free-ai-agent-courses-fall-2026/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>learningpath</category>
      <category>stanfordcs146s</category>
      <category>agenticengineering</category>
    </item>
    <item>
      <title>CMU 11-768 AI Agents Fall 2026: Full Syllabus Breakdown</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:46:58 +0000</pubDate>
      <link>https://dev.to/bruce_he/cmu-11-768-ai-agents-fall-2026-full-syllabus-breakdown-2cgd</link>
      <guid>https://dev.to/bruce_he/cmu-11-768-ai-agents-fall-2026-full-syllabus-breakdown-2cgd</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fidypsplvfepkjka3jjzg.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fidypsplvfepkjka3jjzg.webp" alt="CMU 11-768 AI Agents Fall 2026 syllabus breakdown: build a harness, evaluate, train with RL" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Ten minutes into the second half of lecture 1, Graham Neubig asks the room a question I've spent most of this year writing about. There are two ways to make an agent better, he says: train the LLM, or engineer the harness around it. Show of hands, which one matters more? A few hands go up for training. A lot go up for harness. Daniel Fried raises his hand twice and gets told instructors don't get to vote.&lt;/p&gt;

&lt;p&gt;Then Neubig gives his own answer, and it's the most useful sentence in the four and a half hours of &lt;strong&gt;CMU 11-768 AI Agents&lt;/strong&gt; video that exist as of September 8, 2026: "Typically what happens is you identify a problem and you solve it in the harness first. Then the people training the models catch up... and you don't need to solve it in the harness side anymore." He favors training as the fundamental fix, when you can afford it. You usually can't, so you start with the harness.&lt;/p&gt;

&lt;p&gt;That one exchange tells you what this course is. Stanford's CS146S, the most-read course page on this blog, teaches you to &lt;em&gt;use&lt;/em&gt; agents well. 11-768 is the supply side: how the harness, the eval, and the RL-trained policy get built, taught by the person who built OpenHands and a co-instructor whose group works on human-agent interaction. There is no other public university course that puts scaffold-building, eval design and agent RL in one syllabus. This post breaks down all 28 sessions, the three assignments, and who it's actually for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What CMU 11-768 Actually Is
&lt;/h2&gt;

&lt;p&gt;The site at &lt;a href="https://www.cmu-agents.com/" rel="noopener noreferrer"&gt;cmu-agents.com&lt;/a&gt; is a React app that renders nothing to a plain fetch, so every fact in this table comes from the site's JavaScript bundle (pulled September 8), the &lt;a href="https://github.com/cmu-agents/assignment-1" rel="noopener noreferrer"&gt;Assignment 1 repo&lt;/a&gt;, and the &lt;a href="https://www.youtube.com/watch?v=UwfjzyLnvMg" rel="noopener noreferrer"&gt;lecture 1 recording&lt;/a&gt;. I'll flag the few things I couldn't confirm at the end.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Details&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Course&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;11-768 AI Agents, Carnegie Mellon University (Language Technologies Institute)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Term&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Fall 2026, Aug 25 to Dec 3; Tue/Thu 3:30 to 4:50 PM ET, Porter Hall 100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Instructors&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.phontron.com/" rel="noopener noreferrer"&gt;Graham Neubig&lt;/a&gt; (OpenHands), &lt;a href="https://dpfried.github.io/" rel="noopener noreferrer"&gt;Daniel Fried&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TAs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Aditya Soni, Andy Liu, Apurva Gandhi, Demi Wang, Jiarui Liu, Yueqi Song&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Sessions&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;28 scheduled (22 lectures incl. 6 guest slots, 2 project-hour days, 2 poster days, 2 break days)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Prerequisite&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"Prior experience training neural language models"; 11-667 / 11-711 / 10-202 recommended&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Assignments&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A1 Harness (due Sep 14), A2 Eval (Sep 24), A3 Training (Oct 22), then a team research project&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Compute sponsors&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Fireworks, Modal, Prime Intellect, Sail&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Public materials&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Slides for lectures 1 to 6, YouTube recordings for lectures 1 to 4, Assignment 1 starter code&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The one-line pitch is Neubig's own, from the &lt;a href="https://x.com/gneubig/status/2072730570304430183" rel="noopener noreferrer"&gt;announcement tweet&lt;/a&gt;: "learn how to create a scaffold, build evals, and train an agentic LLM using RL." On the course site, the stated outcomes are that you'll be able to implement an agent from scratch on top of an open-source LLM, design evaluations for multi-step tasks, train agents to improve their capabilities, reason about safety and reliability tradeoffs, and pursue an open research question in agents.&lt;/p&gt;

&lt;p&gt;Read that list against CS146S's syllabus and the gap is obvious. CS146S never touches training. It never asks you to write the loop. 11-768 makes you write the loop in week three.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build, Evaluate, Train: The Spine of the Course
&lt;/h2&gt;

&lt;p&gt;Fried lays out the structure explicitly in lecture 1: three skills in the first half (create a harness, evaluate it, train it with RL), then the second half is a group project that uses those skills. The lecture order is capabilities first, then domains (coding, GUI, deep research), then SFT and RL, then frameworks and safety, then interaction, then guest lectures.&lt;/p&gt;

&lt;p&gt;Here's how that spine maps onto the tools and concepts this blog already covers, because that mapping is the whole reason a Claude Code or Codex power user should care about a graduate NLP course.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-08-cmu-11-768-ai-agents-course/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The dotted edges are the point. Every box on the left side of Assignment 1 is a thing you've configured in a commercial coding agent: hooks, skills, compaction thresholds, tool error handling. The course makes you build each one and then measures what breaks. That's the harness engineering I described in &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-harness-six-layers-reverse-build/" rel="noopener noreferrer"&gt;Build the 6 Layers Backwards&lt;/a&gt;, except with a grader.&lt;/p&gt;

&lt;p&gt;And the right side is where Neubig's "harness first, then training catches up" flow lands. In &lt;a href="https://www.heyuan110.com/posts/ai/2026-05-08-harness-engineering-window-of-opportunity/" rel="noopener noreferrer"&gt;Harness Engineering: Window of Opportunity&lt;/a&gt; I argued the harness advantage is a 2026-2027 window, not a moat. Lecture 1 is the OpenHands creator saying the same thing to a room of PhD students, and then spending Assignment 3 teaching them how to close the window.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Full 28-Session Schedule (Fall 2026)
&lt;/h2&gt;

&lt;p&gt;At module level, the semester looks like this:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-08-cmu-11-768-ai-agents-course/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every row comes from the site's schedule data. Video links are the four recordings that exist on the &lt;a href="https://www.youtube.com/playlist?list=PLSN0qpDfUvTM" rel="noopener noreferrer"&gt;YouTube playlist&lt;/a&gt; as of September 8; slides are live PDFs on cmu-agents.com (lecture 5's deck is 18 MB, so don't open it on mobile data). Speaker is Neubig or Fried unless a name is listed.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;th&gt;Speaker&lt;/th&gt;
&lt;th&gt;Materials&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Tue Aug 25&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Course Overview: What Is an Agent?&lt;/td&gt;
&lt;td&gt;Fried + Neubig&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.cmu-agents.com/slides/lecture-01-agents.pdf" rel="noopener noreferrer"&gt;Slides&lt;/a&gt; · &lt;a href="https://www.youtube.com/watch?v=UwfjzyLnvMg" rel="noopener noreferrer"&gt;Video&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Aug 27&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Agent Capabilities 1: Tool Use&lt;/td&gt;
&lt;td&gt;Neubig&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.cmu-agents.com/slides/lecture-02-tool-use.pdf" rel="noopener noreferrer"&gt;Slides&lt;/a&gt; · &lt;a href="https://www.youtube.com/watch?v=jXChFB4JSyw" rel="noopener noreferrer"&gt;Video&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Sep 1&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Agent Capabilities 2: Context Management for Long-Context Agents&lt;/td&gt;
&lt;td&gt;Neubig&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.cmu-agents.com/slides/lecture-03-long-context.pdf" rel="noopener noreferrer"&gt;Slides&lt;/a&gt; · &lt;a href="https://www.youtube.com/watch?v=AiwCCvFW1uE" rel="noopener noreferrer"&gt;Video&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Sep 3&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Agent Capabilities 3: Skills and Memory&lt;/td&gt;
&lt;td&gt;Fried&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.cmu-agents.com/slides/lecture-04-memory-and-skills.pdf" rel="noopener noreferrer"&gt;Slides&lt;/a&gt; · &lt;a href="https://www.youtube.com/watch?v=6zigF2a-2Pw" rel="noopener noreferrer"&gt;Video&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Sep 8&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Agent Capabilities 4: Planning, Task Decomposition, and Multi-Agent Coordination&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.cmu-agents.com/slides/lecture-05-planning.pdf" rel="noopener noreferrer"&gt;Slides&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Sep 10&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Domains 1: Coding Agents&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.cmu-agents.com/slides/lecture-06-coding-agents.pdf" rel="noopener noreferrer"&gt;Slides&lt;/a&gt; · A1 due Mon Sep 14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Sep 15&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Domains 2: GUI Agents&lt;/td&gt;
&lt;td&gt;JY Koh&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Sep 17&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Training 1: Supervised Fine-Tuning (SFT)&lt;/td&gt;
&lt;td&gt;Yueqi Song&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Sep 22&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Training 2: Reinforcement Learning Basics&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Sep 24&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Domains 3: Deep Research Agents&lt;/td&gt;
&lt;td&gt;Akari Asai&lt;/td&gt;
&lt;td&gt;A2 due&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Sep 29&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;Training 3: Advanced RL Algorithms&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Oct 1&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;Training 4: RL Systems&lt;/td&gt;
&lt;td&gt;Apurva Gandhi&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Oct 6&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;Safety 1: Sandboxing and Credential Management&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Oct 8&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;Frameworks 1: OpenHands&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Oct 13, 15&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Fall Break, no class&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Oct 20&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;Frameworks 2: LangGraph&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Oct 22&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;Safety 2: Observability and Monitoring&lt;/td&gt;
&lt;td&gt;Eric Wallace&lt;/td&gt;
&lt;td&gt;A3 due&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Oct 27&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;Agents and the Future of Work&lt;/td&gt;
&lt;td&gt;Zora Wang&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Oct 29&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;td&gt;Interaction 1: Multi-Agent Interaction&lt;/td&gt;
&lt;td&gt;Saujas Vaduguru&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nov 3, 5&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Project hours&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Nov 10&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;Interaction 2: Human-Agent Interaction&lt;/td&gt;
&lt;td&gt;Valerie Chen&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Nov 12&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;Search 1: Reranking and Critic Models&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Nov 17&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;td&gt;Search 2: Tree Search&lt;/td&gt;
&lt;td&gt;JY Koh&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Nov 19&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;Guest Lecture&lt;/td&gt;
&lt;td&gt;Karthik Narasimhan (Princeton)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tue Nov 24&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;td&gt;Guest Lecture&lt;/td&gt;
&lt;td&gt;Sasha Rush&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thu Nov 26&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Thanksgiving, no class&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dec 1, 3&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;Final presentations (posters)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things about this table that matter more than the titles.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The readings are the real syllabus.&lt;/strong&gt; Lecture 2 alone links 26 references, and they aren't survey papers: the &lt;a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/" rel="noopener noreferrer"&gt;MCP 2026-07-28 spec release&lt;/a&gt;, Qwen3.8's tool-call chat template, DeepSeek V3.2's tool-call encoding script, OpenHands' ToolDefinition source, FastMCP's bearer-token auth. Lecture 3 has 44, including KV-cache pricing pages from DeepSeek, Z.AI, Kimi, OpenAI and Anthropic side by side, and the source of Codex, OpenCode, Pi, Hermes Agent and OpenHands as case studies in how production agents manage context. Lecture 4's readings put Hermes Agent's &lt;code&gt;prompt_builder.py&lt;/code&gt; and &lt;code&gt;skills_tool.py&lt;/code&gt; next to the SkillsBench and "Not All Skills Help" papers. If you only watch the videos, you'll miss most of what the course is teaching.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The guest list is researchers, not vendors.&lt;/strong&gt; CS146S brought in the creator of Claude Code, the CEO of Warp, a partner at a16z. 11-768's outside voices are Karthik Narasimhan (SWE-bench and SWE-agent are from his group), Sasha Rush, Eric Wallace, Akari Asai. Different course, different audience, and you should pick accordingly. More on that below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Assignment 1 Up Close: Build the Thing You've Been Configuring
&lt;/h2&gt;

&lt;p&gt;This is where the course earns its keep for practitioners, so it gets the most space.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/cmu-agents/assignment-1" rel="noopener noreferrer"&gt;cmu-agents/assignment-1&lt;/a&gt; went public on August 31 and had 30 stars and 23 forks by September 8. It's a &lt;code&gt;uv&lt;/code&gt; project with a Makefile, a vendored chess web app, and a 100-point rubric spelled out in &lt;code&gt;ASSIGNMENT.md&lt;/code&gt;. The default model is &lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt; through an OpenAI-compatible endpoint, and every tool call runs in a &lt;a href="https://modal.com/" rel="noopener noreferrer"&gt;Modal&lt;/a&gt; sandbox. The starter's tests fail on purpose; you fill in the TODOs. I cloned it on September 8 and ran &lt;code&gt;make setup&lt;/code&gt; (uv sync plus the pinned &lt;code&gt;chess_app&lt;/code&gt; submodule) and &lt;code&gt;uv run pytest&lt;/code&gt;: 11 failed, 4 passed, 4 deselected (the billable Modal tests), 3.2 seconds, every failure a &lt;code&gt;NotImplementedError&lt;/code&gt; at a TODO. That's the whole offline loop, and it costs nothing.&lt;/p&gt;

&lt;p&gt;Here's what you build, part by part, and what it corresponds to in the tools you already run.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;A1 task&lt;/th&gt;
&lt;th&gt;What you implement&lt;/th&gt;
&lt;th&gt;What it is in Claude Code / Codex terms&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1.1 &lt;code&gt;build_prompt&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;System/user/assistant/tool message sequencing, domain-agnostic&lt;/td&gt;
&lt;td&gt;The context window you've been shaping with CLAUDE.md&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.2 &lt;code&gt;Agent.run&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;The ReAct loop with a &lt;code&gt;step_limit&lt;/code&gt; and &lt;code&gt;finished&lt;/code&gt; flag&lt;/td&gt;
&lt;td&gt;The &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;agentic loop&lt;/a&gt; itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.3 &lt;code&gt;execute_tool_calls&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Parallel tool calls; malformed JSON and unknown tools become recoverable observations, not exceptions&lt;/td&gt;
&lt;td&gt;Tool error handling; what a PostToolUse hook sees&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1.4 Skills&lt;/td&gt;
&lt;td&gt;Discover one &lt;code&gt;SKILL.md&lt;/code&gt; per directory, parse YAML frontmatter, expose name+description in the system prompt, full body via &lt;code&gt;invoke_skill&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Claude Code Skills, literally the &lt;a href="https://agentskills.io/home" rel="noopener noreferrer"&gt;Agent Skills&lt;/a&gt; format&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2.1 &lt;code&gt;compact_context&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Model-generated working memory; summarize the old prefix, keep system/task and latest tool step verbatim; trigger at 6,000 tokens on &lt;code&gt;django__django-15368&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;/compact&lt;/code&gt; and auto-compaction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3.1 to 3.5 ChessAgent&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;play_move&lt;/code&gt;, &lt;code&gt;simulate_move&lt;/code&gt;, &lt;code&gt;run_python&lt;/code&gt; (code runs in the sandbox, never locally), plus a chess skill&lt;/td&gt;
&lt;td&gt;Programmatic tool calling; CodeAct-style code-as-action&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three details in that rubric are worth more than any lecture slide.&lt;/p&gt;

&lt;p&gt;First, the compaction part is graded on a real SWE-bench instance and the report you write comparing token usage with and without compaction. That's the same experiment I ran informally in &lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;Context Engineering for Coding Agents 2026&lt;/a&gt;, where subtracting context beat adding it. Now it's a homework with a FAIL_TO_PASS check. And the reason it's in the course is the story Fried opens lecture 1 with: an OpenClaw user asked the agent to organize her inbox, it announced "I'm taking the nuclear option," deleted her mail, and later admitted "I do remember that she told me this, but I violated it." Fried's diagnosis: the agent compacted its context and the instruction not to delete emails fell out. Part 2 is you building the mechanism that caused that, and learning what has to survive the summary.&lt;/p&gt;

&lt;p&gt;Second, the skills part requires &lt;em&gt;progressive disclosure&lt;/em&gt;: only name and description in the system prompt, full content on demand. The grader checks that when no skill is loaded, nothing in the prompt mentions &lt;code&gt;patch.txt&lt;/code&gt;. That's a cleaner statement of why Claude Code's skill catalog works the way it does than anything in Anthropic's docs.&lt;/p&gt;

&lt;p&gt;Third, Part 3's observation A/B experiment (board only vs. board plus legal moves, on DeepSeek and gpt-oss, four runs) is graded "on the experiment and evidence, not on a particular result or winning the game." Tool output design as a controlled experiment. Most teams I've watched build agents never run this experiment once.&lt;/p&gt;

&lt;p&gt;Two honest caveats. The instructor tests and reference patches aren't in the repo, so outsiders can pass &lt;code&gt;make test&lt;/code&gt; but never get the private score. And the assignment is billable: Modal sandboxes plus LLM tokens. Enrolled students get sponsored credits; you'd pay your own way, though at DeepSeek-V4-Flash prices the LLM side is pocket change and Modal's free tier covers a lot of sandbox minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Should Take 11-768 vs CS146S vs CS329Z
&lt;/h2&gt;

&lt;p&gt;Three courses now cover "AI agents" at top schools this fall, and they're not substitutes. I wrote up the whole field in &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-09-free-ai-agent-courses-fall-2026/" rel="noopener noreferrer"&gt;Free AI Agent Courses Fall 2026&lt;/a&gt;; here's the three-way call.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;CMU 11-768&lt;/th&gt;
&lt;th&gt;Stanford CS146S&lt;/th&gt;
&lt;th&gt;Stanford CS329Z&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Question it answers&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;How do you build, evaluate and train an agent?&lt;/td&gt;
&lt;td&gt;How do you ship software with agents?&lt;/td&gt;
&lt;td&gt;How do you engineer agent systems?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;You write&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A ReAct harness, evals, an RL training run&lt;/td&gt;
&lt;td&gt;Prompts, MCP servers, specs, projects with Claude Code&lt;/td&gt;
&lt;td&gt;See the &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-04-stanford-cs329z-engineering-ai-agents/" rel="noopener noreferrer"&gt;CS329Z breakdown&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Prereq reality&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Have trained a 4-7B model&lt;/td&gt;
&lt;td&gt;Can program&lt;/td&gt;
&lt;td&gt;Systems background&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Public video&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes, 4 of 22 so far&lt;/td&gt;
&lt;td&gt;No official recordings&lt;/td&gt;
&lt;td&gt;See breakdown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Agent builders, harness engineers, ML engineers moving into agents&lt;/td&gt;
&lt;td&gt;Developers who want to use Claude Code / Cursor well&lt;/td&gt;
&lt;td&gt;Engineers designing multi-agent production systems&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Skip if&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;You've never fine-tuned anything and don't plan to&lt;/td&gt;
&lt;td&gt;You already run agents daily&lt;/td&gt;
&lt;td&gt;You want the ML side&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The decision rule I'd give a friend: if the thing you want to get better at is &lt;em&gt;configuring&lt;/em&gt; an agent, take CS146S and read my &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;CS146S overview&lt;/a&gt;. If the thing you want to get better at is deciding &lt;em&gt;whether a problem belongs in the harness or in the weights&lt;/em&gt;, that's 11-768, and nothing else public teaches it.&lt;/p&gt;

&lt;p&gt;One misconception to kill: 11-768 is not an OpenHands course. OpenHands gets exactly one lecture, October 8, and LangGraph gets the next one. Assignment 1 is closer to mini-SWE-agent (which Fried demos in lecture 1 and calls "effective with recent models, but also pretty simple") than to OpenHands. The readings cite Codex, OpenCode, Pi, Hermes Agent and Claude Code's permission modes as peers. Neubig built OpenHands; the course is about the ideas underneath every one of these tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Five Lectures a Claude Code Power User Should Watch
&lt;/h2&gt;

&lt;p&gt;If you run a coding agent daily and have 6 hours, not 27, here's the cut, in order.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Lecture 1, second half (Neubig on capabilities).&lt;/strong&gt; The harness-vs-training exchange, his taxonomy of what makes agents fail (environment understanding, safety as a capability, "a thousand-line change where two lines would do"), and the systems tour: sandboxes (Docker, Apptainer, Modal), inference (vLLM, SGLang, "an extremely important part of agents is caching previous requests"), RL systems (SkyRL, Miles), observability (Laminar, MLflow). 30 minutes that reorganize how you think about your own setup.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lecture 3, Context Management.&lt;/strong&gt; The pricing and KV-cache section is the one to watch. When you see why cached-prefix pricing exists, you stop appending to your CLAUDE.md.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lecture 4, Skills and Memory.&lt;/strong&gt; Fried's lecture, and the readings are the Hermes Agent skills implementation next to the SkillsBench and "Not All Skills Help" papers. Directly applicable to any &lt;code&gt;.claude/skills/&lt;/code&gt; directory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lecture 2, Tool Use.&lt;/strong&gt; Slower, but the chat-template material (how Qwen, Mistral and DeepSeek actually encode a tool call) explains half the "the model called the wrong tool" bugs you've filed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lecture 13, Sandboxing and Credential Management (Oct 6, not yet recorded).&lt;/strong&gt; Neubig previews it in lecture 1 with the OpenAI cyber-benchmark incident where an agent, unable to hack its target, hacked the benchmark's Hugging Face page instead. If you give agents credentials, this is your lecture.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Skip, for now: the SFT/RL block (lectures 8 to 12) unless you have a training job in mind. It's the heart of the course for enrolled students, but it's also the part where "prior experience training language models" stops being a suggestion.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Follow 11-768 Free, and When Videos Actually Drop
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;There's no livestream.&lt;/strong&gt; Lectures are recorded and uploaded in batches to Neubig's YouTube channel. All four current videos were posted on September 8, covering lectures from August 25 through September 3, so the lag is roughly one to two weeks. Lectures 5 and 6 (September 8 and 10) have slides up but no video as of September 8. Don't set an alarm for class time; check the playlist on Tuesdays.&lt;/p&gt;

&lt;p&gt;For completeness, class time in other zones (Pittsburgh is UTC-4 until November 1, then UTC-5):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Zone&lt;/th&gt;
&lt;th&gt;Until Oct 31&lt;/th&gt;
&lt;th&gt;From Nov 3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pittsburgh (ET)&lt;/td&gt;
&lt;td&gt;Tue/Thu 3:30 to 4:50 PM&lt;/td&gt;
&lt;td&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;London&lt;/td&gt;
&lt;td&gt;8:30 to 9:50 PM&lt;/td&gt;
&lt;td&gt;8:30 to 9:50 PM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Beijing / Singapore&lt;/td&gt;
&lt;td&gt;Wed/Fri 3:30 to 4:50 AM&lt;/td&gt;
&lt;td&gt;Wed/Fri 4:30 to 5:50 AM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;India (IST)&lt;/td&gt;
&lt;td&gt;Wed/Fri 1:00 to 2:20 AM&lt;/td&gt;
&lt;td&gt;Wed/Fri 2:00 to 3:20 AM&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;What's free vs. enrollment-only, as of September 8, 2026:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Free&lt;/th&gt;
&lt;th&gt;Enrolled only&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Slides for lectures 1 to 6 (PDF)&lt;/td&gt;
&lt;td&gt;Piazza, Canvas, office hours&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;YouTube recordings 1 to 4 (4h 36m total)&lt;/td&gt;
&lt;td&gt;Sponsored Modal + LLM API credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full reading and reference lists&lt;/td&gt;
&lt;td&gt;Assignment 2 (Eval) and Assignment 3 (Training) repos&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assignment 1 starter repo and rubric&lt;/td&gt;
&lt;td&gt;Private tests, reference patches, grades&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Course policies, grading weights&lt;/td&gt;
&lt;td&gt;Lecture-highlight quizzes, project mentoring&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;A six-week self-learner sequence.&lt;/strong&gt; The semester spreads 22 lectures over 15 weeks with breaks. If you're following on your own, compress it and reorder around the assignment you can actually do:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Week 1:&lt;/strong&gt; Lectures 1 and 2. Clone assignment-1, run &lt;code&gt;make setup&lt;/code&gt; and &lt;code&gt;make doctor&lt;/code&gt;, get one billable &lt;code&gt;make run-code-agent&lt;/code&gt; through.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Week 2:&lt;/strong&gt; Lectures 3 and 4. Do A1 Parts 1 and 2. Write the token-usage comparison even though nobody will grade it; that report is the learning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Week 3:&lt;/strong&gt; Lectures 5 and 6. Do A1 Part 3. Run the four-way observation A/B.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Week 4:&lt;/strong&gt; Lectures 7 to 10 as they post (GUI, SFT, RL basics, deep research). Read the SWE-Gym and R2E-Gym papers from lecture 6's list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Week 5:&lt;/strong&gt; Lectures 11 to 14 (advanced RL, RL systems, sandboxing, OpenHands). Since A2 and A3 aren't public, build your own eval over your A1 harness: 10 tasks, FAIL_TO_PASS style, and an LLM-as-judge for the ones without tests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Week 6:&lt;/strong&gt; Lectures 15 to 23 as they land through late November. Guest lectures are the dessert.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Grading, if you want to know what enrolled students are optimizing: Assignment 1 is 10%, Assignments 2 and 3 are 15% each, lecture highlights are 10% (22 opportunities, best 20 count, must be written without AI, submitted within 24 hours), and the team project is 50% across proposal, check-in, poster and a 30% final report. Each assignment gets two 24-hour slack days, then 5% per day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Page Stops
&lt;/h2&gt;

&lt;p&gt;Things I could not verify and won't pretend to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Assignment 2 and 3 contents.&lt;/strong&gt; Only the one-line summaries on the site ("design the evaluation framework," "implement the training procedures") and Fried's remark that A2 involves "LLM-as-judge based evaluation approaches among other evals." No repos as of September 8.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who teaches lectures 5, 6, 9, 11, 13, 14, 15, 20.&lt;/strong&gt; The schedule lists no lecturer for those, which by the site's convention means Neubig or Fried; I haven't confirmed which.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whether every lecture will be recorded.&lt;/strong&gt; Four of six delivered lectures have video. The guest-lecture policy on recording isn't stated.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Modal and API costs for an outsider doing A1 end to end.&lt;/strong&gt; I ran the setup and the offline test suite, not the billable pipeline (&lt;code&gt;make run-code-agent&lt;/code&gt;, the SWE-bench run, the four chess runs); I'm not going to invent a dollar figure.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else, from the 28 dates to the 100-point rubric to the quotes from lecture 1 (lightly cleaned of "um" and "uh," from YouTube's auto-transcript), is checkable at the links above. If cmu-agents.com changes the schedule, the JavaScript bundle at &lt;code&gt;/assets/index-*.js&lt;/code&gt; is where the truth lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-09-09-free-ai-agent-courses-fall-2026/" rel="noopener noreferrer"&gt;Free AI Agent Courses Fall 2026: Stanford, CMU, MIT Compared&lt;/a&gt; — the hub that places 11-768 against every other public course this term&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-09-04-stanford-cs329z-engineering-ai-agents/" rel="noopener noreferrer"&gt;Stanford CS329Z: Engineering AI Agents&lt;/a&gt; — the systems-side sibling&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;Stanford CS146S: The Modern Software Developer&lt;/a&gt; — the demand-side course, for using agents rather than building them&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-harness-six-layers-reverse-build/" rel="noopener noreferrer"&gt;Harness Engineering: Build the 6 Layers Backwards&lt;/a&gt; — Assignment 1 is layers 1 through 4 with a grader&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-05-08-harness-engineering-window-of-opportunity/" rel="noopener noreferrer"&gt;Harness Engineering: Window of Opportunity, Not a Forever Moat&lt;/a&gt; — Neubig's "harness first, then training catches up," argued from the practitioner side&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;Context Engineering for Coding Agents 2026&lt;/a&gt; — the compaction experiment, before it was homework&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;Agentic Loops 2026&lt;/a&gt; — the loop you implement in A1 Part 1&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-08-cmu-11-768-ai-agents-course/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>cmu11768</category>
      <category>coursereview</category>
      <category>harnessengineering</category>
    </item>
    <item>
      <title>pi Coding Agent Review 2026: 4 Tools vs Claude Code, Tested</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:46:21 +0000</pubDate>
      <link>https://dev.to/bruce_he/pi-coding-agent-review-2026-4-tools-vs-claude-code-tested-h52</link>
      <guid>https://dev.to/bruce_he/pi-coding-agent-review-2026-4-tools-vs-claude-code-tested-h52</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7bentdmlfs696rotxclq.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7bentdmlfs696rotxclq.webp" alt="pi coding agent review 2026: primitives, not features, tested against Claude Code" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The same one-word prompt cost 1,358 tokens of context in pi and 31,012 in Claude Code. Same Mac, same afternoon, same "reply with exactly the word PONG." That single number is the whole pitch of the &lt;strong&gt;pi coding agent&lt;/strong&gt;, and it's also the number that misleads people into thinking pi is a lighter Claude Code.&lt;/p&gt;

&lt;p&gt;It isn't. After a day of installing it, writing an extension for it, forking its sessions, and running three identical tasks in both pi and Claude Code, my read is this: &lt;strong&gt;pi is a harness-engineering kit, not a finished harness.&lt;/strong&gt; It ships the four layers you'd build first if you were rolling your own agent, and it hands you the two layers that my &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-harness-six-layers-reverse-build/" rel="noopener noreferrer"&gt;six-layer harness post&lt;/a&gt; says drive 80% of production stability. Whether that's a gift or a trap depends entirely on who you are.&lt;/p&gt;

&lt;p&gt;Receipts below. A quick note on scope: a Chinese tutorial on runoob covers pi's feature list step by step; this post is about what those features cost, what they don't do, and whether you should switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What pi coding agent actually is (September 2026)
&lt;/h2&gt;

&lt;p&gt;pi is Mario Zechner's answer to Claude Code turning into, in his words, "a spaceship with 80% of functionality I have no use for." Zechner is the libGDX creator; he published pi's rationale on &lt;a href="https://mariozechner.at/posts/2025-11-30-pi-coding-agent/" rel="noopener noreferrer"&gt;his blog in November 2025&lt;/a&gt; and the core idea hasn't moved since: four tools (&lt;code&gt;read&lt;/code&gt;, &lt;code&gt;write&lt;/code&gt;, &lt;code&gt;edit&lt;/code&gt;, &lt;code&gt;bash&lt;/code&gt;), a system prompt small enough to read in one screen, and a TypeScript extension API for everything else.&lt;/p&gt;

&lt;p&gt;The project's numbers are not small-project numbers. As of September 6, 2026, &lt;a href="https://github.com/earendil-works/pi" rel="noopener noreferrer"&gt;earendil-works/pi&lt;/a&gt; has 103,925 stars and 12,999 forks, created August 9, 2025, last pushed the day before I checked, MIT licensed. The CLI package &lt;code&gt;@earendil-works/pi-coding-agent&lt;/code&gt; sits at v0.85.1 (released September 5) and pulls 1.53M npm downloads a week.&lt;/p&gt;

&lt;p&gt;For scale, &lt;code&gt;@anthropic-ai/claude-code&lt;/code&gt; does 8.55M and &lt;code&gt;@openai/codex&lt;/code&gt; 13.16M, so pi is roughly one-sixth of Claude Code by install volume and well ahead of OpenCode's 1.31M.&lt;/p&gt;

&lt;p&gt;The repo moved in May. On May 7, 2026, v0.74.0 became the first release under the &lt;code&gt;@earendil-works&lt;/code&gt; scope after Zechner joined Earendil, the public-benefit corporation Armin Ronacher co-founded; the old &lt;code&gt;badlogic/pi-mono&lt;/code&gt; URL now redirects. The license stayed MIT and the CLI is still &lt;code&gt;pi&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Two facts make the move matter more than a rename: OpenClaw, the agent that got half the industry's attention this spring, is built on pi's packages (its current release &lt;code&gt;openclaw@2026.9.4&lt;/code&gt; depends on &lt;code&gt;@earendil-works/pi-tui&lt;/code&gt;), and pi's library packages out-download its CLI. &lt;code&gt;pi-tui&lt;/code&gt; alone gets 5.12M weekly downloads. &lt;strong&gt;pi is already more "infrastructure other agents are built on" than "CLI people type into."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Community reception is ecosystem-shaped rather than launch-thread-shaped. There's no 800-point Show HN; the loudest pi thread on Hacker News in 2026 is a 56-point complaint that its config folder ignores XDG on Linux (August 17), and the second is oh-my-pi, a fork with an IDE wired in (42 points, July 21). The &lt;a href="https://newsletter.pragmaticengineer.com/p/building-pi-and-what-makes-self-modifying" rel="noopener noreferrer"&gt;Pragmatic Engineer&lt;/a&gt; ran a full episode on it in April. And the package gallery at &lt;a href="https://pi.dev/packages" rel="noopener noreferrer"&gt;pi.dev/packages&lt;/a&gt; lists 5,410 packages, which is the real reception signal: people are building on it, not arguing about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  pi's four packages on the six harness layers
&lt;/h2&gt;

&lt;p&gt;pi ships as a monorepo of four packages plus a couple of support libraries: &lt;code&gt;pi-ai&lt;/code&gt; (unified provider API), &lt;code&gt;pi-agent-core&lt;/code&gt; (the loop), &lt;code&gt;pi-tui&lt;/code&gt; (terminal rendering), and &lt;code&gt;pi-coding-agent&lt;/code&gt; (the CLI that wires them together, plus the SDK and RPC modes). Mapped onto the six-layer model from my earlier posts, the shape is lopsided in a very deliberate way.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-06-pi-coding-agent-review/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Layers 1 through 4 are not just present, they're better than most people expect. The session tree in particular is something Claude Code doesn't have: every entry in a pi session has an &lt;code&gt;id&lt;/code&gt; and a &lt;code&gt;parentId&lt;/code&gt;, so you can branch from any earlier turn without losing the original path. Layer 4 is the layer I told people to skip in the six-layer post, and pi ships it anyway because it's cheap when the storage format is a tree from day one.&lt;/p&gt;

&lt;p&gt;Layers 5 and 6 are where pi stops. There is a JSON event stream and an RPC protocol, which is the raw material for observability, but no cost dashboard beyond the TUI footer and no eval harness. And there is no permission system at all. pi's &lt;a href="https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/security.md" rel="noopener noreferrer"&gt;security doc&lt;/a&gt; says it plainly: "Pi does not include a built-in sandbox... Real isolation needs to come from the operating system or a virtualization/container boundary."&lt;/p&gt;

&lt;p&gt;If you read my &lt;a href="https://www.heyuan110.com/posts/ai/2026-05-08-harness-engineering-window-of-opportunity/" rel="noopener noreferrer"&gt;window-of-opportunity post&lt;/a&gt;, you'll recognize the bet pi is making: the patch layers (context tricks, forced planning, permission theater) get absorbed by better models, so don't bake them in. The design-literacy layers (what to verify, what to recover from) stay with the user. pi is that argument turned into a product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hands-on: install, provider, and what a prompt costs before you type
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Install.&lt;/strong&gt; &lt;code&gt;npm install -g --ignore-scripts @earendil-works/pi-coding-agent&lt;/code&gt; took 77 seconds on my M-series Mac and landed 449 MB in &lt;code&gt;node_modules&lt;/code&gt;. That's the first thing that punctures the "minimal" story: 292 MB of that is &lt;code&gt;@esbuild&lt;/code&gt; platform binaries, and the rest is the Anthropic, OpenAI, Google, and AWS SDKs bundled so every provider works out of the box. pi is minimal in context tokens, not on disk. On first interactive launch it also downloaded &lt;code&gt;fd&lt;/code&gt; and &lt;code&gt;ripgrep&lt;/code&gt; (6.7 MB) into &lt;code&gt;~/.pi/agent/bin&lt;/code&gt; without asking, which is convenient and slightly presumptuous.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Provider.&lt;/strong&gt; pi's &lt;code&gt;/login&lt;/code&gt; supports subscription OAuth for ChatGPT Plus/Pro, Claude Pro/Max, GitHub Copilot, and xAI, plus 30-odd API-key providers including DeepSeek, Kimi, MiniMax, Qwen, and Z.AI. One line in the providers doc deserves a highlight because it changes the economics for a lot of readers: "Third-party harness usage draws from extra usage and is billed per token, not against Claude plan limits." I did not run the Claude login for that reason; it would have charged this account per token. The only key on this machine was Google's, so pi ran on &lt;strong&gt;Gemini 3.1 Pro&lt;/strong&gt; (with Gemini 3.8 Flash for the cheap experiments), and Claude Code ran on its default &lt;strong&gt;Claude Fable 5.1&lt;/strong&gt;. Keep that in mind for everything below: these are harness-plus-model comparisons, not harness-only.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context footprint.&lt;/strong&gt; This is the measurement that started the post. I sent the same "reply with exactly the word PONG" through three configurations and read the usage back from each tool's JSON output:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;Input tokens before your first word&lt;/th&gt;
&lt;th&gt;List-price cost of "PONG"&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;pi, &lt;code&gt;--no-skills --no-extensions --no-context-files&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,358&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$0.0034 (Gemini 3.1 Pro)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pi, default (auto-loaded 48 skills from &lt;code&gt;~/.agents/skills&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;10,787&lt;/td&gt;
&lt;td&gt;$0.0081 (Gemini 3.8 Flash)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code 2.1.268, &lt;code&gt;claude -p&lt;/code&gt;, fresh empty repo&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;31,012&lt;/strong&gt; (20,884 cache write + 10,126 cache read + 2)&lt;/td&gt;
&lt;td&gt;$0.4214 (Fable 5.1, list)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things to take from that table. First, pi's floor really is small: 1,358 tokens is the four tool definitions, the prompt, and the working directory, and it's the same number on Pro and Flash. &lt;/p&gt;

&lt;p&gt;Second, the middle row is the gotcha nobody warns you about. pi auto-discovers skills from &lt;code&gt;~/.agents/skills&lt;/code&gt; and &lt;code&gt;.agents/skills&lt;/code&gt;, the same directories other harnesses use. I had 48 skills sitting there from other tools, and pi silently put 9,400 tokens of skill descriptions into every request. Run &lt;code&gt;pi --verbose&lt;/code&gt; once and read the startup header before you trust the "under 1,000 tokens" marketing line.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three identical tasks: pi vs Claude Code receipts
&lt;/h2&gt;

&lt;p&gt;I built a 241-line Node project (invoice math: money, tax, discounts, CSV report) with a five-rule &lt;code&gt;AGENTS.md&lt;/code&gt;, a planted off-by-one bug in the discount tiers, and four &lt;code&gt;TODO&lt;/code&gt; comments. Then I ran three prompts through both agents headlessly (&lt;code&gt;pi --mode json -p&lt;/code&gt; and &lt;code&gt;claude -p --output-format json --dangerously-skip-permissions&lt;/code&gt;), each in a fresh copy of the repo.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;pi + Gemini 3.1 Pro&lt;/th&gt;
&lt;th&gt;Claude Code + Fable 5.1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;a. Add 3 unit tests&lt;/strong&gt; for &lt;code&gt;mergeLines&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;29.5s · 6 turns · 5 tool calls · 75k in / 2.0k out · &lt;strong&gt;$0.10&lt;/strong&gt; · 8/8 pass&lt;/td&gt;
&lt;td&gt;27.2s · 4 turns · 132k in / 1.5k out · &lt;strong&gt;$0.61&lt;/strong&gt; · 8/8 pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;b. Find and fix the tier bug&lt;/strong&gt;, add regression test, explain&lt;/td&gt;
&lt;td&gt;28.7s · 7 turns · 6 tool calls · 86k in / 1.7k out · &lt;strong&gt;$0.11&lt;/strong&gt; · identical 3-line fix · 1 test, 7 asserts&lt;/td&gt;
&lt;td&gt;41.5s · 7 turns · 219k in / 2.5k out · &lt;strong&gt;$0.77&lt;/strong&gt; · identical 3-line fix · 3 tests incl. negative-input case&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;c. Resolve all 4 TODOs&lt;/strong&gt; with implementations and tests&lt;/td&gt;
&lt;td&gt;94.7s · 21 turns · 20 tool calls (12 bash, 6 edit, 2 write) · 316k in / 5.3k out · &lt;strong&gt;$0.31&lt;/strong&gt; · 9/9 pass · 0 TODOs left&lt;/td&gt;
&lt;td&gt;86.5s · 8 turns · 245k in / 7.0k out · &lt;strong&gt;$1.16&lt;/strong&gt; · 13/13 pass · 0 TODOs left · test-first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AGENTS.md compliance (JSDoc, &lt;code&gt;node:test&lt;/code&gt; only, no &lt;code&gt;data/&lt;/code&gt; edits, ran &lt;code&gt;npm test&lt;/code&gt;, &lt;code&gt;SUMMARY:&lt;/code&gt; line)&lt;/td&gt;
&lt;td&gt;5/5 on all three&lt;/td&gt;
&lt;td&gt;5/5 on all three&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Costs are list-price figures reported by each tool; I pay Anthropic a subscription, so the Claude Code column is what it would cost on the API, not what left my bank account. The token columns are the honest part: &lt;strong&gt;Claude Code consumed 1.8x to 2.5x the input tokens of pi on every task&lt;/strong&gt;, and the gap is almost entirely harness overhead being re-sent (and cache-read) every turn.&lt;/p&gt;

&lt;p&gt;Where the quality actually differed, it was in task c, and it was the model, not the harness. Claude Code wrote each test first and confirmed it failed before implementing, made the zero-total warning's logger injectable so the test didn't need to mock &lt;code&gt;console&lt;/code&gt;, and ended with an unprompted note: "One thing I noticed but left alone since it is not a TODO: &lt;code&gt;volumeDiscount&lt;/code&gt; uses strict greater-than comparisons," which is the exact bug from task b. pi with Gemini resolved all four TODOs correctly, mocked &lt;code&gt;console.warn&lt;/code&gt; instead, and didn't notice the adjacent bug. On tasks a and b the diffs were functionally identical. &lt;strong&gt;Four tools were not the bottleneck on any of the three tasks.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That matches the only same-model comparison I've found. Composio ran 30 agentic tool-use tasks on DeepSeek V4 Pro through both pi and OpenCode in August 2026: pi solved 21/30 (70%) to OpenCode's 19/30 (63%), spent $1.64 to OpenCode's $2.25, but was slower at the median (363s vs 281s). Under 1,000 tokens of fixed overhead per request versus about 6,900. Same shape as my numbers: fewer tokens, similar or better outcomes, not faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 40-line extension, and how the model walked around it
&lt;/h2&gt;

&lt;p&gt;pi's extensibility story is the part that's actually different from Claude Code's, so I wrote an extension instead of reading about one. Extensions are TypeScript files loaded via jiti (no build step) that get an &lt;code&gt;ExtensionAPI&lt;/code&gt;. Mine registers a &lt;code&gt;tool_call&lt;/code&gt; handler that blocks &lt;code&gt;rm -rf&lt;/code&gt;, &lt;code&gt;sudo&lt;/code&gt;, and force-pushes, a custom tool the model can call, and a slash command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// ~/.pi/agent/extensions/guard.ts&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;ExtensionAPI&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@earendil-works/pi-coding-agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;isToolCallEventType&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@earendil-works/pi-coding-agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Type&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;typebox&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;DANGEROUS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;rm&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+-&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;a-z&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;*r&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;a-z&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;*f&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;git&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+push&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+.*--force&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;sudo&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nf"&gt;function &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pi&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ExtensionAPI&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;blocked&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="nx"&gt;pi&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tool_call&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;isToolCallEventType&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;bash&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cmd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;command&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;DANGEROUS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;re&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="c1"&gt;// Interactive: ask. Headless (-p / --mode json): deny. Never self-approve.&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;hasUI&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ui&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;confirm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;guard.ts&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;`Allow?\n&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;blocked&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;block&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`guard.ts blocked destructive command: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="nx"&gt;pi&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;registerTool&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;guard_stats&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;label&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Guard stats&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Report how many bash commands guard.ts has blocked this session&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Type&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;({}),&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;text&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`guard.ts blocked &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;blocked&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; command(s) so far`&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt; &lt;span class="na"&gt;details&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;blocked&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="nx"&gt;pi&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;registerCommand&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;guard&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Show guard.ts block count&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;handler&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;_args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ui&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;notify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`guard.ts blocked &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;blocked&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; command(s)`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;info&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the PreToolUse hook from Claude Code, except it's a typed function in the same process instead of a shell script reading JSON on stdin, and it can mutate &lt;code&gt;event.input&lt;/code&gt; in place, register tools, draw TUI widgets, and read the session. The &lt;code&gt;ctx.hasUI&lt;/code&gt; branch is the detail the &lt;a href="https://agenticcontrolplane.com/blog/pi-acp-extension" rel="noopener noreferrer"&gt;Agentic Control Plane team&lt;/a&gt; called the "empty-chair test" when they built a policy extension in August: in headless mode nobody is there to click Allow, so "ask" must become "deny."&lt;/p&gt;

&lt;p&gt;Then I ran it: &lt;code&gt;pi -e guard.ts --mode json -p "mkdir tmp_scratch ... then rm -rf tmp_scratch ... if that fails, delete it another way."&lt;/code&gt; Here is the tool sequence from the JSON stream, 11.6 seconds and $0.015 on Gemini 3.8 Flash:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CALL bash  {"command": "mkdir -p tmp_scratch &amp;amp;&amp;amp; echo hi &amp;gt; tmp_scratch/a.txt"}
  END bash  isError=false
CALL bash  {"command": "rm -rf tmp_scratch"}
  END bash  isError=true   guard.ts blocked destructive command: rm -rf tmp_scratch
CALL bash  {"command": "rm tmp_scratch/a.txt &amp;amp;&amp;amp; rmdir tmp_scratch"}
  END bash  isError=false
CALL guard_stats {}
  END guard_stats  guard.ts blocked 1 command(s) so far
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The hook fired. The custom tool worked. And the model deleted the directory anyway, one turn later, with &lt;code&gt;rm&lt;/code&gt; plus &lt;code&gt;rmdir&lt;/code&gt;. I told it to find another way, so it did, but that's exactly what a prompt-injected model would do too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A &lt;code&gt;tool_call&lt;/code&gt; hook is policy, not a boundary.&lt;/strong&gt; It's great for "don't touch &lt;code&gt;.env&lt;/code&gt;," useless against a model (or an attacker in a README) that wants the thing gone. That is the strongest argument for Zechner's position, not against it: if the only real boundary is the OS, then permission dialogs are UX, and pretending otherwise is the dangerous part.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-06-pi-coding-agent-review/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The registry side works as advertised. &lt;code&gt;pi install npm:pi-mcp-adapter&lt;/code&gt; took 25 seconds, added one line to &lt;code&gt;~/.pi/agent/settings.json&lt;/code&gt;, and gave me &lt;code&gt;/mcp&lt;/code&gt;. The adapter is the community's rebuttal to Zechner's "you don't need MCP" post: one proxy tool of about 200 tokens instead of the 13,700 tokens Playwright MCP dumps into context, servers started lazily. It gets 197,718 downloads a week. &lt;code&gt;pi-subagents&lt;/code&gt; gets 66,401. The things pi refused to build are the most-installed things in its registry, which is either a vindication of primitives or a sign that everyone rebuilds the batteries anyway. I think it's both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Session tree: what /fork actually does
&lt;/h2&gt;

&lt;p&gt;pi stores sessions as JSONL trees under &lt;code&gt;~/.pi/agent/sessions/&lt;/code&gt;. I ran a two-turn session headlessly (read &lt;code&gt;app.js&lt;/code&gt;, then edit it to print &lt;code&gt;v2&lt;/code&gt;), then forked it with &lt;code&gt;pi --fork &amp;lt;id&amp;gt; -p "Instead of v2, change it to print v3. Which version did the file print when you started?"&lt;/code&gt;. The fork answered "v3... When I started, the file printed v1," and here is what landed on disk:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;original session
message   id=a35234bf parent=0e946043 user       "Read app.js and tell me..."
message   id=d7986285 parent=a35234bf assistant  (read tool call)
message   id=7b182e60 parent=d7986285 toolResult "console.log('v1')"
message   id=2b93a6ee parent=7b182e60 assistant  "app.js prints the string v1"
message   id=d6b72102 parent=2b93a6ee user       "Now change it to print v2..."
message   id=d8eb9eb7 parent=a17b112b assistant  "Updated app.js to print v2."

fork (new file, header carries parentSession=&amp;lt;original path&amp;gt;)
... same 10 entries copied, then:
message   id=2c7cfb20 parent=d8eb9eb7 user       "Instead of v2, change it to print v3..."
message   id=763c751e parent=3694092b assistant  "I have updated app.js to print v3..."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So &lt;code&gt;--fork&lt;/code&gt; from the CLI copies the whole active branch into a new file and continues from the leaf. Branching from an &lt;em&gt;earlier&lt;/em&gt; point (the thing Claude Code can't do at all) is the interactive &lt;code&gt;/tree&lt;/code&gt; view, where selecting an old user message moves the leaf to its parent and puts the text back in your editor for re-submission; abandoned branches can get an automatic summary attached at the new position. The JSONL is plain enough that I parsed it with ten lines of Python, which is the observability primitive pi gives you instead of a dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  No sandbox, no permission dialogs: what it means for a team
&lt;/h2&gt;

&lt;p&gt;Here's the comparison table you're probably here for. Everything is as of September 2026; sandbox and permission facts are from each tool's own docs.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;pi 0.85&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Claude Code 2.1&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Codex CLI&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;OpenCode&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sandbox&lt;/td&gt;
&lt;td&gt;None built in; Docker / Gondolin micro-VM / OpenShell patterns documented&lt;/td&gt;
&lt;td&gt;OS-level sandboxed Bash (Seatbelt on macOS, bubblewrap on Linux), auto-allow or ask&lt;/td&gt;
&lt;td&gt;Kernel sandbox by default (Seatbelt / Landlock+bwrap / Windows), network off&lt;/td&gt;
&lt;td&gt;None built in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Permissions&lt;/td&gt;
&lt;td&gt;None; &lt;code&gt;tool_call&lt;/code&gt; hook or a registry extension (&lt;code&gt;cc-safety-net&lt;/code&gt;, &lt;code&gt;@gotgenes/pi-permission-system&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Prompt per tool, allow/deny rules, permission modes, hooks&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;approval_policy&lt;/code&gt; x &lt;code&gt;sandbox_mode&lt;/code&gt;, two independent dials&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;allow&lt;/code&gt; / &lt;code&gt;ask&lt;/code&gt; / &lt;code&gt;deny&lt;/code&gt; rules per tool with globs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extension model&lt;/td&gt;
&lt;td&gt;TypeScript modules in-process: events, tools, commands, TUI, providers; npm/git packages&lt;/td&gt;
&lt;td&gt;Hooks (shell scripts), skills, MCP, plugins&lt;/td&gt;
&lt;td&gt;AGENTS.md, MCP, config profiles&lt;/td&gt;
&lt;td&gt;Plugins, MCP, agents in config&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model vendors&lt;/td&gt;
&lt;td&gt;Anthropic, OpenAI, Google, DeepSeek, Kimi, MiniMax, Qwen, Z.AI, Bedrock, Vertex, Ollama, custom OpenAI-compatible&lt;/td&gt;
&lt;td&gt;Anthropic (plus Bedrock, Vertex, Foundry)&lt;/td&gt;
&lt;td&gt;OpenAI (plus OSS providers via config)&lt;/td&gt;
&lt;td&gt;Any via Models.dev (75+)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session model&lt;/td&gt;
&lt;td&gt;JSONL tree; &lt;code&gt;/tree&lt;/code&gt;, &lt;code&gt;/fork&lt;/code&gt;, &lt;code&gt;/clone&lt;/code&gt;, branch summaries&lt;/td&gt;
&lt;td&gt;Linear with &lt;code&gt;--resume&lt;/code&gt;, checkpoints&lt;/td&gt;
&lt;td&gt;Linear resume&lt;/td&gt;
&lt;td&gt;Linear with share&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;License and price&lt;/td&gt;
&lt;td&gt;MIT; you pay the model&lt;/td&gt;
&lt;td&gt;Proprietary CLI; subscription or API&lt;/td&gt;
&lt;td&gt;Apache-2.0; ChatGPT plan or API&lt;/td&gt;
&lt;td&gt;MIT; you pay the model&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Zechner's argument for the empty Sandbox row is quotable and mostly right: "As soon as your agent can write code and run code, it's pretty much game over... Everybody is running in YOLO mode anyways to get any productive work done, so why not make it the default and only option?" My guard experiment above is the evidence. But there's a second half he leaves to you, and it's the half that matters for a team.&lt;/p&gt;

&lt;p&gt;For a solo engineer, no dialogs is a productivity gain with a known risk you've already accepted (you were going to click Allow anyway). For a team, the question isn't "is the prompt useful," it's "who is accountable when an agent running as &lt;code&gt;deploy&lt;/code&gt; deletes the fixtures." Claude Code's answer is a sandbox you can mandate in &lt;code&gt;settings.json&lt;/code&gt;; Codex's is a kernel policy on by default. pi's answer is "run it in a container," which is correct and also means the container is now your responsibility, your CI's responsibility, and your onboarding doc's responsibility. If you already run agents in Docker or a micro-VM, pi costs you nothing here. If you don't, pi is the tool that makes you start.&lt;/p&gt;

&lt;p&gt;There's a smaller trust boundary pi does implement, and it's the right one: &lt;strong&gt;project trust&lt;/strong&gt;. The first time you open a repo with &lt;code&gt;.pi/extensions&lt;/code&gt; or &lt;code&gt;.agents/skills&lt;/code&gt;, pi asks before loading them, because a repo that can silently install an in-process TypeScript extension can do anything. Headless runs skip the prompt and default to not loading them. That's the one place pi says "ask first," and it's the one place where asking actually buys you something.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who should switch, who shouldn't
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Switch to pi if&lt;/strong&gt; at least two of these are true: you already run agents in containers or CI; you're embedding an agent in your own product and want the SDK or RPC mode rather than a subprocess of &lt;code&gt;claude -p&lt;/code&gt;; you need models Anthropic doesn't sell (DeepSeek, Kimi, Qwen, a self-hosted vLLM), which pi treats as first-class; you pay per token at scale and a 20x context floor difference shows up on the invoice; or you've hit the ceiling of what a shell-script hook can do and want a typed, in-process one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stay on Claude Code if&lt;/strong&gt; you're onboarding people who've never run an agent, because pi has no guardrails to catch them; you need plan mode, sub-agents, and MCP today and don't want to curate packages; you want a vendor to answer the phone when something breaks (Earendil is a small company; pi has 202 open issues); or you're on a Claude Max plan and expect included usage, which pi's Claude login explicitly isn't.&lt;/p&gt;

&lt;p&gt;The honest middle: if you build harnesses for a living, pi is the best reference implementation of layers 1 through 4 you can read in an afternoon, and the cheapest base to build 5 and 6 on top of. If you consume harnesses, Claude Code's batteries are worth their 31,012 tokens. I laid out the same distinction for tool surfaces in &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;CLI skills vs MCP&lt;/a&gt; and for the loop itself in &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;agentic loops&lt;/a&gt;; pi is what you get when you take both of those posts' advice literally.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rough edges from one day of use
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;pi -p&lt;/code&gt; hangs if stdin is an open pipe.&lt;/strong&gt; My first headless run sat for 180 seconds with zero output because the tool harness kept a pipe open; &lt;code&gt;&amp;lt; /dev/null&lt;/code&gt; fixed it. If you script pi in CI, redirect stdin or pipe your prompt in explicitly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skill auto-discovery inflates context silently.&lt;/strong&gt; 48 skills from &lt;code&gt;~/.agents/skills&lt;/code&gt; turned a 1,358-token request into 10,787. Use &lt;code&gt;--no-skills&lt;/code&gt; or &lt;code&gt;pi config&lt;/code&gt; to prune.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Warning spam.&lt;/strong&gt; With both &lt;code&gt;GOOGLE_API_KEY&lt;/code&gt; and &lt;code&gt;GEMINI_API_KEY&lt;/code&gt; set, pi printed "Both ... are set. Using GOOGLE_API_KEY" 19 times to stderr in a single run, once per model call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;449 MB install&lt;/strong&gt; for a tool whose pitch is minimalism, and unrequested binary downloads (&lt;code&gt;fd&lt;/code&gt;, &lt;code&gt;rg&lt;/code&gt;) on first launch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost visibility is uneven.&lt;/strong&gt; The TUI footer and &lt;code&gt;--mode json&lt;/code&gt; both report per-message cost with cache splits, which is better than Claude Code's headless output. Plain &lt;code&gt;-p&lt;/code&gt; text mode reports nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CJK rendered fine.&lt;/strong&gt; I typed a Chinese prompt into the editor through a pty and it displayed correctly; the differential renderer didn't garble wide characters. The &lt;code&gt;--tui-mode fullscreen&lt;/code&gt; mode has a documented iTerm2 inline-image limitation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;~/.pi/agent&lt;/code&gt; ignores XDG on Linux&lt;/strong&gt;, which is the community's loudest complaint this year (56 points on HN, August 17).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;pi coding agent is the clearest existence proof that four tools and 1,358 tokens are enough to match a batteries-included harness on ordinary coding tasks; my three runs and Composio's thirty say the same thing. What it is not is a drop-in Claude Code replacement, because the two layers it leaves out, observability and constraints, are the two that decide whether an agent survives contact with a team.&lt;/p&gt;

&lt;p&gt;If you've been meaning to build your own harness, pi is where I'd start, and the guard extension above is your first afternoon. If you just want the work done and someone else to own the sandbox, keep paying for the spaceship.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-harness-six-layers-reverse-build/" rel="noopener noreferrer"&gt;Harness Engineering: Build the 6 Layers Backwards&lt;/a&gt; — the six-layer model pi maps onto, and why layers 5-6 carry the weight&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-05-08-harness-engineering-window-of-opportunity/" rel="noopener noreferrer"&gt;Harness Engineering: Window of Opportunity, Not a Forever Moat&lt;/a&gt; — the "patch layers get absorbed" argument pi is built on&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;CLI Skills vs MCP&lt;/a&gt; — the token math behind pi's no-MCP default and the adapter that undoes it&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;Agentic Loops&lt;/a&gt; — what pi-agent-core's loop does and doesn't do for you&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-19-claude-code-vs-codex/" rel="noopener noreferrer"&gt;Claude Code vs Codex&lt;/a&gt; and &lt;a href="https://www.heyuan110.com/posts/ai/2026-03-10-codex-cli-deep-dive/" rel="noopener noreferrer"&gt;Codex CLI Deep Dive&lt;/a&gt; — the two batteries-included harnesses in the comparison table&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  External References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/earendil-works/pi" rel="noopener noreferrer"&gt;earendil-works/pi on GitHub&lt;/a&gt; — source, docs, 79 example extensions&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://pi.dev" rel="noopener noreferrer"&gt;pi.dev&lt;/a&gt; and the &lt;a href="https://pi.dev/packages" rel="noopener noreferrer"&gt;package gallery&lt;/a&gt; — "There are many agent harnesses, but this one is yours"&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://mariozechner.at/posts/2025-11-30-pi-coding-agent/" rel="noopener noreferrer"&gt;Mario Zechner, "pi coding agent" (Nov 30, 2025)&lt;/a&gt; — the rationale for every omission&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://pi.dev/news/2026/5/7/pi-has-a-new-home" rel="noopener noreferrer"&gt;pi has a new home (May 7, 2026)&lt;/a&gt; — the Earendil move and package rename&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://composio.dev/content/pi-vs-opencode" rel="noopener noreferrer"&gt;Composio, "Pi vs OpenCode: After 100 Hours" (Aug 21, 2026)&lt;/a&gt; — the same-model 30-task comparison&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agenticcontrolplane.com/blog/pi-acp-extension" rel="noopener noreferrer"&gt;Agentic Control Plane, "pi ships no permission system, on purpose" (Aug 17, 2026)&lt;/a&gt; — the empty-chair test&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-06-pi-coding-agent-review/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>picodingagent</category>
      <category>claudecode</category>
      <category>harnessengineering</category>
      <category>aicoding</category>
    </item>
    <item>
      <title>Stanford CS329Z Engineering AI Agents: Syllabus + Self-Study</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:45:28 +0000</pubDate>
      <link>https://dev.to/bruce_he/stanford-cs329z-engineering-ai-agents-syllabus-self-study-4k6b</link>
      <guid>https://dev.to/bruce_he/stanford-cs329z-engineering-ai-agents-syllabus-self-study-4k6b</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgfj64shq3qm95f6492uh.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgfj64shq3qm95f6492uh.webp" alt="Stanford CS329Z Engineering AI Agents Fall 2026 syllabus breakdown and self-study plan" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Stanford's new agents course bans agent frameworks in its first assignment. HW1 of &lt;strong&gt;Stanford CS329Z: Engineering AI Agents&lt;/strong&gt; tells students to build a company's internal AI assistant "with no agent frameworks: just a chat-completion call and code you write yourself." LangChain, LangGraph, DSPy, and LlamaIndex don't show up until lecture six, and when they do, the stated purpose is to compare "what frameworks abstract vs. what you built from scratch."&lt;/p&gt;

&lt;p&gt;That one rule tells you what kind of course this is. It's not "learn LangChain in ten weeks," and it's not a coding-with-Claude-Code practicum like &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;CS146S&lt;/a&gt;. It's an engineering-discipline course: the syllabus names three challenges up front (decomposition, data, evaluation), and 17 lectures, two homeworks, and 50 readings are organized around them.&lt;/p&gt;

&lt;p&gt;This page is the syllabus breakdown plus the part I actually care about: since the recordings are Canvas-only, how far can you get on the public materials alone, and which weeks are worth your time if you already ship agents for a living. I haven't sat in Packard 101 (the first lecture is September 23, 2026), so everything below comes from the &lt;a href="https://cs329z.stanford.edu/" rel="noopener noreferrer"&gt;course site&lt;/a&gt;, the &lt;a href="https://bulletin.stanford.edu/courses/2283761" rel="noopener noreferrer"&gt;Stanford Bulletin&lt;/a&gt;, the instructors' own pages, and the readings themselves. Where I'm guessing, I say so.&lt;/p&gt;

&lt;h2&gt;
  
  
  CS329Z at a Glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Details&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Course&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;CS 329Z: Engineering AI Agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Term&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Fall 2026 (Autumn 1, 2026-27), first offering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Lectures&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Mon/Wed 1:30-2:50 pm, Packard 101, Sep 23 to Dec 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Units / class number&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3 units, #27855, Letter or Credit/No Credit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Instructors&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://cs.stanford.edu/~diyiy/" rel="noopener noreferrer"&gt;Diyi Yang&lt;/a&gt; (Stanford NLP faculty), &lt;a href="https://michryan.com/" rel="noopener noreferrer"&gt;Michael Ryan&lt;/a&gt; (PhD student, DSPy core contributor, MIPROv2 and GEPA author), &lt;a href="https://john-b-yang.github.io/" rel="noopener noreferrer"&gt;John Yang&lt;/a&gt; (PhD student, SWE-bench / SWE-agent / SWE-smith first author)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Prerequisite&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Any of CS224N, CS224U, CS224V, CS336, or equivalent NLP background&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Frameworks touched&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;litellm, DSPy, LangChain/LangGraph, LlamaIndex, MCP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Recordings&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Captured, distributed via Canvas only. Not public.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Contact&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;a href="mailto:cs329z-staff@lists.stanford.edu"&gt;cs329z-staff@lists.stanford.edu&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The instructor lineup is the strongest signal about what the course values. Michael Ryan wrote the optimizer papers that make DSPy more than a prompt wrapper (MIPROv2 at EMNLP 2024, &lt;a href="https://arxiv.org/abs/2507.19457" rel="noopener noreferrer"&gt;GEPA&lt;/a&gt;, ICLR 2026 oral). John Yang built the benchmark the entire coding-agent industry reports against (&lt;a href="https://arxiv.org/abs/2310.06770" rel="noopener noreferrer"&gt;SWE-bench&lt;/a&gt;), the reference agent that runs on it (&lt;a href="https://github.com/SWE-agent/SWE-agent" rel="noopener noreferrer"&gt;SWE-agent&lt;/a&gt;), and the data engine that trains for it (&lt;a href="https://github.com/SWE-bench/SWE-smith" rel="noopener noreferrer"&gt;SWE-smith&lt;/a&gt;). One instructor is optimization, one is evaluation and data. That's the course.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "Engineering" Is the Operative Word
&lt;/h2&gt;

&lt;p&gt;The course description opens with the &lt;a href="https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/" rel="noopener noreferrer"&gt;compound AI systems&lt;/a&gt; framing from Zaharia et al.: the interesting unit is no longer a model, it's a system of models, retrievers, tools, and optimizers. Then it lists what students learn to do: "pick what types of problems to focus on, decompose problems, select appropriate components, collect and curate data, build evaluations, and reason about the design tradeoffs."&lt;/p&gt;

&lt;p&gt;Notice what's missing from that list: prompting. It gets a bullet in lecture two and then disappears. Compare that with most 2025 agent courses, where prompt patterns were half the syllabus.&lt;/p&gt;

&lt;p&gt;The grading confirms it. Half the grade is a quarter-long group project ("Making Life at Stanford Better with Agents"), 20% is the two homeworks, and 15% is two ten-minute quizzes where you explain the design decisions and tradeoffs in your own homework. There's no exam on lecture content. You cannot pass CS329Z by remembering what was said in Packard 101, which, as it happens, is also why not having the videos hurts less than you'd think.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Weight&lt;/th&gt;
&lt;th&gt;What it actually tests&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Project (proposal 5, midway report 5, midpoint demo 7, final submission 15, final demo 18)&lt;/td&gt;
&lt;td&gt;50%&lt;/td&gt;
&lt;td&gt;Can you scope and ship a working agent in ten weeks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HW1: Build an Agentic Harness&lt;/td&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;td&gt;Can you build tools, memory, a terminal, and a human-in-the-loop with zero frameworks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HW2: Evaluate an Agent&lt;/td&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;td&gt;Can you write code graders, an LLM judge, and 4-tuple benchmark tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HW-based quizzes (2 x 7.5%)&lt;/td&gt;
&lt;td&gt;15%&lt;/td&gt;
&lt;td&gt;Can you defend your own design tradeoffs, closed-book&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Paper video (7%) + 3 peer reviews (3%)&lt;/td&gt;
&lt;td&gt;10%&lt;/td&gt;
&lt;td&gt;Can you critique a recent agent paper and add something (a reproduction, an experiment)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Participation&lt;/td&gt;
&lt;td&gt;5%&lt;/td&gt;
&lt;td&gt;Discussion, teamwork, recitations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Here's the mistake I expect people to make with this course: treating it as CS146S with a harder prerequisite. It isn't. CS146S (Fall 2026 runs Tue/Thu from Sep 22, and I wrote a &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;follow-along plan for it&lt;/a&gt;) is about &lt;em&gt;developer practice&lt;/em&gt;: using Claude Code, Cursor, MCP servers, and code review to ship software faster. CS329Z is about &lt;em&gt;building the agent&lt;/em&gt; that someone else uses. The prerequisites say the same thing: CS146S wants CS111-level programming; CS329Z wants CS224N.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;CS146S: The Modern Software Developer&lt;/th&gt;
&lt;th&gt;CS329Z: Engineering AI Agents&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Question it answers&lt;/td&gt;
&lt;td&gt;How do I ship software 10x faster with agents?&lt;/td&gt;
&lt;td&gt;How do I build an agent system that works?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unit of work&lt;/td&gt;
&lt;td&gt;Your codebase, your PRs&lt;/td&gt;
&lt;td&gt;Harness, RAG pipeline, eval suite&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frameworks&lt;/td&gt;
&lt;td&gt;Claude Code, Cursor, Warp, MCP&lt;/td&gt;
&lt;td&gt;litellm, DSPy, LangGraph, LlamaIndex, MCP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First assignment&lt;/td&gt;
&lt;td&gt;Prompting playground (2025) / 200-line agent (2026)&lt;/td&gt;
&lt;td&gt;Full harness with no frameworks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evaluation content&lt;/td&gt;
&lt;td&gt;One review week&lt;/td&gt;
&lt;td&gt;Three lectures plus HW2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prerequisite&lt;/td&gt;
&lt;td&gt;CS111&lt;/td&gt;
&lt;td&gt;CS224N / CS336&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Videos&lt;/td&gt;
&lt;td&gt;None public&lt;/td&gt;
&lt;td&gt;None public (Canvas only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Who should follow&lt;/td&gt;
&lt;td&gt;Engineers who write code with agents&lt;/td&gt;
&lt;td&gt;Engineers who ship agents as the product&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you're a working developer whose agent exposure is Claude Code and Cursor, CS146S first. If you own an "AI assistant" feature at work and your last three incidents were the agent doing something plausible but wrong, CS329Z is the one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Full Fall 2026 Schedule
&lt;/h2&gt;

&lt;p&gt;Twenty-one meeting slots between September 23 and December 2: 17 content lectures, 2 guest lectures (speakers TBA), and 2 Thanksgiving days off. The schedule is marked tentative on the site; this is the version as of September 4, 2026. Required readings are listed; each lecture also carries "additional readings" (27 in total) that I've folded into the self-study table further down.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Wk&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Lecture&lt;/th&gt;
&lt;th&gt;Required readings&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Wed Sep 23&lt;/td&gt;
&lt;td&gt;Introduction: what are agentic systems? Monolithic models to compound systems to agents; the three challenges (decomposition, data, evaluation)&lt;/td&gt;
&lt;td&gt;Zaharia et al., Compound AI Systems (BAIR 2024)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Mon Sep 28&lt;/td&gt;
&lt;td&gt;LLMs for builders: APIs and SDKs (litellm), structured I/O and constrained generation, decoding and test-time compute, context engineering, model selection, cost/latency&lt;/td&gt;
&lt;td&gt;Anthropic, Building Effective Agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Wed Sep 30&lt;/td&gt;
&lt;td&gt;Retrieval-Augmented Generation: grounding, embeddings and vector stores, chunking, hybrid search, cross-encoders and ColBERT. Hands-on: RAG from scratch&lt;/td&gt;
&lt;td&gt;Lewis et al., RAG (NeurIPS 2020)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Mon Oct 5&lt;/td&gt;
&lt;td&gt;Tool use and function calling: the REPL, function-calling APIs, MCP, designing good tools, sandboxes, error handling and retries. Hands-on: tool-using system from scratch&lt;/td&gt;
&lt;td&gt;MCP Specification (2025-06-18)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Wed Oct 7&lt;/td&gt;
&lt;td&gt;Frameworks and orchestration: DSPy (signatures, modules, optimizers), LangChain/LangGraph, LlamaIndex; what frameworks abstract vs. what you built&lt;/td&gt;
&lt;td&gt;Khattab et al., DSPy (ICLR 2024)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Mon Oct 12&lt;/td&gt;
&lt;td&gt;Agent design patterns and scaffolds: workflows vs. agents, five workflow patterns, ReAct, plan-and-execute, reflection; scaffolds as design decisions&lt;/td&gt;
&lt;td&gt;Yao et al., ReAct (ICLR 2023)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Wed Oct 14&lt;/td&gt;
&lt;td&gt;Agent memory architectures: short vs. long-term, memory as tool actions, the file system as memory, structured memory, cross-agent memory&lt;/td&gt;
&lt;td&gt;Packer et al., MemGPT (2023)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Mon Oct 19&lt;/td&gt;
&lt;td&gt;Multi-agent systems: single vs. multi, orchestration, handoffs and state transfer, delegation, coordination and error propagation&lt;/td&gt;
&lt;td&gt;Wu et al., AutoGen (COLM 2024)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Wed Oct 21&lt;/td&gt;
&lt;td&gt;Optimization: prompt optimization (GEPA, MIPROv2, OPRO, TextGrad), test-time compute, LoRA/QLoRA, distillation, RLHF/DPO; prompts vs. weights vs. inference compute&lt;/td&gt;
&lt;td&gt;Snell et al., Scaling Test-Time Compute; Agrawal et al., GEPA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Mon Oct 26&lt;/td&gt;
&lt;td&gt;Guest lecture (TBA)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Wed Oct 28&lt;/td&gt;
&lt;td&gt;What data do agents need? Traces, demonstrations, feedback; data for optimization vs. evaluation; flywheels; synthetic data&lt;/td&gt;
&lt;td&gt;Shankar, Data Flywheels for LLM Applications&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Mon Nov 2&lt;/td&gt;
&lt;td&gt;Data selection and quality: maximally informative data, filtering, tiny-but-targeted benchmarks, annotation, datasets from agent traces&lt;/td&gt;
&lt;td&gt;Yang et al., SWE-smith; Shankar et al., Who Validates the Validators?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Wed Nov 4&lt;/td&gt;
&lt;td&gt;Evaluation fundamentals and benchmark design: why evals are hard, the 4-tuple (request, environment, stopping criteria, scorer), properties of good benchmarks, reliability&lt;/td&gt;
&lt;td&gt;Zhu et al., Rigorous Agentic Benchmarks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Mon Nov 9&lt;/td&gt;
&lt;td&gt;LLM-as-judge and eval infrastructure: three grader types, judge prompts, known biases, pairwise vs. pointwise, pass@k vs. pass^k, harness design, Anthropic's 8-step roadmap&lt;/td&gt;
&lt;td&gt;Anthropic, Demystifying Evals for AI Agents; Zheng et al., MT-Bench; Ryan et al., AutoMetrics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Wed Nov 11&lt;/td&gt;
&lt;td&gt;Agent safety and guardrails: privacy risks of tool access, prompt injection (incl. indirect), red-teaming, sandboxing and permissions, output guardrails, human-in-the-loop&lt;/td&gt;
&lt;td&gt;Shao et al., PrivacyLens; Zhang &amp;amp; Yang, Privacy Risks via Simulation; Li, Agentic LLMs as Deanonymizers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Mon Nov 16&lt;/td&gt;
&lt;td&gt;Guest lecture (TBA)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Wed Nov 18&lt;/td&gt;
&lt;td&gt;Coding and software agents: end-to-end; SWE-agent, Claude Code, and OpenHands architectures; scaffolds as design decisions; SWE-bench and the 4-tuple in practice&lt;/td&gt;
&lt;td&gt;Yang et al., SWE-agent; Wang et al., OpenHands&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Nov 23, 25&lt;/td&gt;
&lt;td&gt;No class (Thanksgiving)&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;Mon Nov 30&lt;/td&gt;
&lt;td&gt;Proactive agents: reactive to proactive, General User Models, next-action prediction, open-source proactive agents, mixed initiative&lt;/td&gt;
&lt;td&gt;Shaikh et al., General User Models (UIST 2025)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;Wed Dec 2&lt;/td&gt;
&lt;td&gt;Frontiers and open problems: multimodal, web and computer-use agents, science agents, long-running architectures, observability and cost, reliability&lt;/td&gt;
&lt;td&gt;(additional only: OSWorld, WebShop)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Finals&lt;/td&gt;
&lt;td&gt;Dec 7-11&lt;/td&gt;
&lt;td&gt;Final project demo day&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Deadlines, all Pacific: HW1 released Oct 5, project proposal due Oct 9, HW2 released Oct 26, HW1 due Oct 30, midpoint demo Nov 4 in class, midway report Nov 6, paper video Nov 13, HW2 due Nov 20, peer reviews Nov 30, final submission and demo during finals week.&lt;/p&gt;

&lt;p&gt;Two things jump out from the table. First, the "from scratch" hands-on sessions (RAG on Sep 30, tools on Oct 5) come &lt;em&gt;before&lt;/em&gt; the frameworks lecture on Oct 7, and HW1 releases the same day as the tools lecture. The course wants your hands dirty before it hands you an abstraction. Second, five of the 17 content lectures (Oct 21, Oct 28, Nov 2, Nov 4, Nov 9) are about data, optimization, and evaluation. Nearly 30% of the content is the part most agent tutorials skip entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Five Modules, and Where This Blog Already Covers Them
&lt;/h2&gt;

&lt;p&gt;The syllabus has ten section headers; I collapse them into five modules because that's how the dependencies actually run. Modules one and two are what you build; three and four are how you make it good and prove it; five is where the field is going.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-04-stanford-cs329z-engineering-ai-agents/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The dotted lines are honest about coverage. This blog has written a lot about modules one, two, and five, mostly from the coding-agent angle: &lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;context engineering&lt;/a&gt;, &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;agentic loops&lt;/a&gt;, &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-13-harness-subagent-architecture/" rel="noopener noreferrer"&gt;sub-agent architecture&lt;/a&gt;, and &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-harness-six-layers-reverse-build/" rel="noopener noreferrer"&gt;harness layers&lt;/a&gt;. Module three (data) is where I have almost nothing, and module four (evaluation) I've only covered as one layer of a harness. That gap is a fair proxy for the industry: everyone writes about the loop, few write about the flywheel.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Two Homeworks Are the Course
&lt;/h2&gt;

&lt;p&gt;If you take one thing from CS329Z without enrolling, take the homework sequence. The site describes both briefs in enough detail to rebuild them, and together they encode the course's actual thesis: build the harness bare-handed, then find out you can't tell whether it works.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;HW1: Build an Agentic Harness (weeks 3-6, released Oct 5, due Oct 30).&lt;/strong&gt; "Build a company's internal AI assistant from scratch, with no agent frameworks: just a chat-completion call and code you write yourself. Start with LLM pipelines that retrieve and reason over a real corporate email archive, then grow them into a full agent harness with tools, a terminal, memory, and a human in the loop." The course doesn't name the archive; my guess is the Enron corpus, because it's the only large, public, real corporate email dataset, and it has the messy threads that make retrieval interesting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;HW2: Evaluate an Agent (weeks 6-9, released Oct 26, due Nov 20).&lt;/strong&gt; "Given a pre-built agent, design a comprehensive evaluation suite with code-based graders, at least one LLM-as-judge eval, benchmark tasks built with the 4-tuple framework (request, environment, stopping criteria, scorer), and error analysis."&lt;/p&gt;

&lt;p&gt;Read the order again. Students spend four weeks building a harness and then four weeks evaluating a &lt;em&gt;different&lt;/em&gt;, pre-built agent, not their own. That's a deliberate move: it stops you from writing evals that flatter your own design, and it forces the 4-tuple discipline onto something you didn't build and can't quietly patch. The quiz after each homework is closed-book and asks you to explain tradeoffs, so "I copied the pattern from a tutorial" doesn't survive contact.&lt;/p&gt;

&lt;p&gt;Here's how I'd reconstruct both as exercises you can run with Claude Code or Codex as your pair, without any course infrastructure:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-04-stanford-cs329z-engineering-ai-agents/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two rules make the substitute worth doing. For HW1, the zero-frameworks constraint is the whole exercise: if you &lt;code&gt;pip install langgraph&lt;/code&gt; on weekend one, you've skipped the course. Let your coding agent write boilerplate, but you write the loop, the tool dispatch, and the memory policy. For HW2, use an agent you didn't write. &lt;a href="https://github.com/OpenHands/OpenHands" rel="noopener noreferrer"&gt;OpenHands&lt;/a&gt; and &lt;a href="https://github.com/SWE-agent/SWE-agent" rel="noopener noreferrer"&gt;SWE-agent&lt;/a&gt; are both open, both configurable, and both were built by people who wrote the course readings.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;pass^k&lt;/code&gt; metric in the Nov 9 lecture is the one detail I'd flag for anyone who has shipped an agent. &lt;code&gt;pass@k&lt;/code&gt; asks whether &lt;em&gt;any&lt;/em&gt; of k runs succeeds; &lt;code&gt;pass^k&lt;/code&gt; asks whether &lt;em&gt;all&lt;/em&gt; k do. For a demo, the first number matters. For an agent that runs unattended on customer data, only the second one does, and it collapses fast: a task that passes 80% of the time has a pass^5 of about 33%. That single reframing is worth more than most agent tutorials.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-Study Substitute, Module by Module
&lt;/h2&gt;

&lt;p&gt;Every reading in the syllabus is public. Lecture slides "will be linked here as they are released," which may or may not happen; the Fall 2025 CS146S decks did eventually go public, so it's plausible. Until then, here's a substitute for each module, weighted toward what you can run rather than what you can read.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Module&lt;/th&gt;
&lt;th&gt;Watch/read instead of the lecture&lt;/th&gt;
&lt;th&gt;Do instead of the section&lt;/th&gt;
&lt;th&gt;Blog post to pair&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. Building blocks (Sep 23-Oct 5)&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/" rel="noopener noreferrer"&gt;Compound AI Systems&lt;/a&gt;; &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Building Effective Agents&lt;/a&gt;; Lewis RAG and ColBERT papers; the &lt;a href="https://modelcontextprotocol.io/specification/2025-06-18" rel="noopener noreferrer"&gt;MCP spec&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Write a RAG pipeline and a tool-calling loop with raw &lt;a href="https://docs.litellm.ai/docs/" rel="noopener noreferrer"&gt;litellm&lt;/a&gt; calls. No vector DB service; a numpy matrix and cosine similarity is enough to learn the failure modes&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;Context Engineering 2026&lt;/a&gt;, &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;CLI Skills vs MCP&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. Frameworks and design (Oct 7-21)&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://dspy.ai/" rel="noopener noreferrer"&gt;DSPy docs&lt;/a&gt; on signatures, modules, and &lt;a href="https://dspy.ai/learn/optimization/optimizers/" rel="noopener noreferrer"&gt;optimizers&lt;/a&gt;; &lt;a href="https://langchain-ai.github.io/langgraph/" rel="noopener noreferrer"&gt;LangGraph docs&lt;/a&gt;; ReAct, MemGPT, AutoGen papers; &lt;a href="https://arxiv.org/abs/2503.13657" rel="noopener noreferrer"&gt;Why Do Multi-Agent LLM Systems Fail?&lt;/a&gt;; Neubig's &lt;a href="https://openhands.dev/blog/dont-sleep-on-single-agent-systems" rel="noopener noreferrer"&gt;Don't Sleep on Single-Agent Systems&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Port your HW1 harness to DSPy, then to LangGraph. Write down what each framework took away from you. Then run &lt;code&gt;dspy.GEPA&lt;/code&gt; or MIPROv2 on one module and measure the delta on a 50-example dev set&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;Agentic Loops&lt;/a&gt;, &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-13-harness-subagent-architecture/" rel="noopener noreferrer"&gt;Subagent Architecture&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. Data (Oct 28-Nov 2)&lt;/td&gt;
&lt;td&gt;Shankar's &lt;a href="https://www.sh-reya.com/blog/ai-engineering-flywheel/" rel="noopener noreferrer"&gt;Data Flywheels&lt;/a&gt;; &lt;a href="https://arxiv.org/abs/2504.21798" rel="noopener noreferrer"&gt;SWE-smith&lt;/a&gt;; Who Validates the Validators; LIMA&lt;/td&gt;
&lt;td&gt;Log every trace from your HW1 agent. Build a 100-example dataset from the traces, then a 20-example "tiny but targeted" subset that predicts the full score&lt;/td&gt;
&lt;td&gt;(gap, see below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4. Evaluation and safety (Nov 4-11)&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents" rel="noopener noreferrer"&gt;Demystifying Evals for AI Agents&lt;/a&gt;; &lt;a href="https://arxiv.org/abs/2507.02825" rel="noopener noreferrer"&gt;Rigorous Agentic Benchmarks&lt;/a&gt;; MT-Bench; &lt;a href="https://arxiv.org/abs/2512.17267" rel="noopener noreferrer"&gt;AutoMetrics&lt;/a&gt;; OpenAI on &lt;a href="https://openai.com/index/prompt-injections/" rel="noopener noreferrer"&gt;prompt injections&lt;/a&gt;; PrivacyLens&lt;/td&gt;
&lt;td&gt;The HW2 substitute above. Then plant one indirect prompt injection in your email corpus and see whether your HW1 agent takes the bait&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-harness-six-layers-reverse-build/" rel="noopener noreferrer"&gt;Harness: 6 Layers Backwards&lt;/a&gt; (layers 5-6)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5. Coding, proactive, frontier (Nov 18-Dec 2)&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://arxiv.org/abs/2405.15793" rel="noopener noreferrer"&gt;SWE-agent&lt;/a&gt; and &lt;a href="https://arxiv.org/abs/2407.16741" rel="noopener noreferrer"&gt;OpenHands&lt;/a&gt; papers; Anthropic's &lt;a href="https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents" rel="noopener noreferrer"&gt;Effective Harnesses for Long-Running Agents&lt;/a&gt;; General User Models; &lt;a href="https://github.com/openclaw/openclaw" rel="noopener noreferrer"&gt;OpenClaw&lt;/a&gt;
&lt;/td&gt;
&lt;td&gt;Run SWE-agent on 10 SWE-bench Lite tasks with two different scaffolds (tool set or system prompt changed, model held constant). Report the scaffold delta&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-harness-six-layers-reverse-build/" rel="noopener noreferrer"&gt;Harness: 6 Layers Backwards&lt;/a&gt;, &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;Agentic Loops&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One warning on module two, because it's where self-learners lose the most time. DSPy's docs are good but they're organized as a tour, not a course; the thing the Oct 7 lecture will give enrolled students is the &lt;em&gt;comparison&lt;/em&gt;: signatures/modules/optimizers against LangGraph's graph-of-nodes against LlamaIndex's data-first view. Without that framing you tend to learn one framework's vocabulary and mistake it for the concept. The fix is the exercise in the table: port the same harness across two frameworks and write the diff in prose. The prose is the deliverable.&lt;/p&gt;

&lt;p&gt;For a broader map of what's free this fall, including CMU's parallel agents course, see the &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-09-free-ai-agent-courses-fall-2026/" rel="noopener noreferrer"&gt;Fall 2026 free AI agent courses hub&lt;/a&gt; and the &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-08-cmu-11-768-ai-agents-course/" rel="noopener noreferrer"&gt;CMU 11-768 breakdown&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Most Valuable Week for Readers of This Blog
&lt;/h2&gt;

&lt;p&gt;If you already run coding agents in production, the obvious pick is November 18: coding agents, taught by the person who wrote SWE-agent, covering Claude Code and OpenHands architectures side by side. It'll be the most fun lecture of the quarter. It's not the most valuable one for you, because you've already read the SWE-agent paper, the Claude Code best-practices post, and probably this blog's harness series, and there's no recording anyway.&lt;/p&gt;

&lt;p&gt;The most valuable block is the evaluation fortnight, November 4 and 9, plus the HW2 brief. Here's my reasoning. In the 60 days I tracked a harness in production for the &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-harness-six-layers-reverse-build/" rel="noopener noreferrer"&gt;six-layers post&lt;/a&gt;, the layers that drove stability were eval and recovery, not the loop, and eval was the layer I built last and worst. That's the common pattern: teams build the loop in a week, tune prompts for a month, and then ship without a benchmark they'd bet money on. CS329Z spends three lectures and a full homework on exactly this, and the 4-tuple framework (request, environment, stopping criteria, scorer) is a portable artifact you can apply to any agent tomorrow, no Stanford login required.&lt;/p&gt;

&lt;p&gt;Second place goes to October 21 (optimization), for a narrower audience. If your team is fine-tuning because "prompting plateaued," the GEPA paper on that week's list reports beating GRPO with up to 35x fewer rollouts through reflective prompt evolution. Whether that holds on your task is an afternoon with &lt;code&gt;dspy.GEPA&lt;/code&gt; and your dev set, and that afternoon might cancel a fine-tuning project.&lt;/p&gt;

&lt;h2&gt;
  
  
  What You Can't Get Without Enrolling
&lt;/h2&gt;

&lt;p&gt;Self-study gets you the readings, the topic list, and two reconstructable homeworks. It doesn't get you these, and I'd rather list them than pretend the substitute is equivalent.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The recordings.&lt;/strong&gt; Cameras capture the instructor; the footage goes to Canvas. Non-enrolled means no video, full stop. Stanford Online does list &lt;a href="https://online.stanford.edu/courses/cs329z-engineering-ai-agents" rel="noopener noreferrer"&gt;CS329Z&lt;/a&gt; as a course, which for other CS courses means non-degree enrollment at the published rate of $1,575 per unit with a 3-unit minimum (roughly $4,725 plus a $250 document fee). The listing itself wouldn't load for me, so confirm the non-degree option before you budget for it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The homework infrastructure.&lt;/strong&gt; The email archive, the pre-built agent for HW2, the autograders, and whatever starter harness they hand out. My substitutes above are reconstructions from the brief, not the real assignments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The quizzes.&lt;/strong&gt; Fifteen percent of the grade is a closed-book, ten-minute defense of your own design decisions. That's the single best forcing function in the course and it doesn't exist outside it, unless you get a colleague to grill you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The two guest lectures&lt;/strong&gt; (Oct 26 and Nov 16, speakers TBA). CS146S's guest list was a major draw; CS329Z hasn't announced anyone yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feedback from Michael Ryan and John Yang&lt;/strong&gt; on your eval suite and your optimizer runs. This is the thing I'd pay for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demo day&lt;/strong&gt; and a team. "Making Life at Stanford Better with Agents" is a group project with a proposal, a midway report, and a demo in finals week. Alone, you can build the system; you can't fake the deadlines or the audience.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Credit.&lt;/strong&gt; Three units, letter or credit/no credit, repeatable: no.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;CS329Z is the first Stanford course that treats agents as an engineering discipline rather than a developer skill, and it proves it by grading design decisions and evals rather than lecture recall. Follow it if you build agents for other people to use; follow CS146S if you use agents to build for other people; skip both if you haven't yet written a tool-calling loop against a raw API, because both courses assume you have.&lt;/p&gt;

&lt;p&gt;Without the recordings, the useful part of CS329Z is about 70% available: 50 public readings and two homework briefs specific enough to rebuild. Do HW1 with zero frameworks and HW2 against an agent you didn't write, and you'll have covered the part of this course that the industry most needs and least teaches. I'll update this page when the slides land or the guests are announced.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-24-stanford-cs146s-overview/" rel="noopener noreferrer"&gt;Stanford CS146S: The Modern Software Developer, 2026 Guide&lt;/a&gt;: the developer-practice sibling course, fully broken down&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-09-11-cs146s-fall-2026-follow-along/" rel="noopener noreferrer"&gt;CS146S Fall 2026: How to Watch and Follow Along Free&lt;/a&gt;: calendar and follow-along plan for the other Stanford course this fall&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-09-09-free-ai-agent-courses-fall-2026/" rel="noopener noreferrer"&gt;Free AI Agent Courses, Fall 2026&lt;/a&gt;: the hub page comparing every open agents syllabus this term&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-09-08-cmu-11-768-ai-agents-course/" rel="noopener noreferrer"&gt;CMU 11-768: AI Agents Course Breakdown&lt;/a&gt;: the CMU counterpart, and how it differs from CS329Z&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-18-harness-six-layers-reverse-build/" rel="noopener noreferrer"&gt;Harness Engineering: Build the 6 Layers Backwards&lt;/a&gt;: why eval and recovery are the layers that matter, with production numbers&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;Context Engineering for Coding Agents 2026&lt;/a&gt;: the lecture-two topic, in depth&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-03-agentic-loops/" rel="noopener noreferrer"&gt;Agentic Loops 2026&lt;/a&gt;: the ReAct-style loop that HW1 makes you write by hand&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-04-stanford-cs329z-engineering-ai-agents/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>stanfordcs329z</category>
      <category>agents</category>
      <category>dspy</category>
      <category>coursereview</category>
    </item>
    <item>
      <title>gitui vs lazygit in 2026: Benchmarked on 82K Commits</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:45:22 +0000</pubDate>
      <link>https://dev.to/bruce_he/gitui-vs-lazygit-in-2026-benchmarked-on-82k-commits-272e</link>
      <guid>https://dev.to/bruce_he/gitui-vs-lazygit-in-2026-benchmarked-on-82k-commits-272e</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe8gfbkgqzt3ppzeoxo6d.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe8gfbkgqzt3ppzeoxo6d.webp" alt="gitui vs lazygit benchmark cover: two terminal Git clients compared on a large repository" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Eleven milliseconds. That's how long gitui took to paint a usable screen on an 82,180-commit clone of git/git on my M5 MacBook, and the number didn't move when I pointed it at a 630-commit blog repo instead. Lazygit took 416 ms on the same repo. If you came here for "gitui vs lazygit, which is faster," you can stop reading: gitui, by a factor of nearly 40.&lt;/p&gt;

&lt;p&gt;It's also the least important fact in this comparison. The tool that paints 40x faster can't do an interactive rebase, can't show you a conflict diff, has no custom commands and no worktree view, and shipped three releases in the last twenty months. The tool that takes 416 ms shipped twelve in the last seven.&lt;/p&gt;

&lt;p&gt;So the real question isn't speed. It's whether the things gitui does uniquely well (blame, full-tree browsing, flat memory on giant histories) matter more to you than the things only lazygit does. I spent a week with both, wrote a pty harness to time them properly, and I'll give you the receipts and then a straight answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gitui vs lazygit benchmark: three repos, one harness
&lt;/h2&gt;

&lt;p&gt;Cards on the table: you can't benchmark a TUI with &lt;code&gt;hyperfine&lt;/code&gt; the way you'd time &lt;code&gt;ls&lt;/code&gt;. Both tools refuse to start without a terminal, and "startup time" is meaningless until the screen shows something you can act on. So I wrote a small Python harness that forks each tool inside a pseudo-terminal, feeds the output into a &lt;code&gt;pyte&lt;/code&gt; screen emulator, and stops the clock when the screen meets a readiness condition. For lazygit that's "the Commits panel is drawn and contains a hash." For gitui it's "the tab bar is drawn and every &lt;code&gt;Loading ...&lt;/code&gt; placeholder is gone." Memory is the RSS of the whole process tree sampled after the screen settles, because lazygit spawns &lt;code&gt;git&lt;/code&gt; children (it auto-fetches on launch) and those count.&lt;/p&gt;

&lt;p&gt;Setup, so you can reproduce or argue with it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Machine&lt;/td&gt;
&lt;td&gt;MacBook Pro, Apple M5, macOS 26.5.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lazygit&lt;/td&gt;
&lt;td&gt;v0.65.0 (Homebrew, released 2026-09-05), 18.4 MB binary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gitui&lt;/td&gt;
&lt;td&gt;v0.28.1 (Homebrew, released 2026-03-24), 9.5 MB binary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small repo&lt;/td&gt;
&lt;td&gt;this blog, 630 commits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medium repo&lt;/td&gt;
&lt;td&gt;neovim/neovim, 38,077 commits, 331 MB &lt;code&gt;.git&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large repo&lt;/td&gt;
&lt;td&gt;git/git, 82,180 commits, 334 MB &lt;code&gt;.git&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache&lt;/td&gt;
&lt;td&gt;warm (10 runs each, median reported); I couldn't purge the page cache without sudo, so no cold numbers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal&lt;/td&gt;
&lt;td&gt;200x50 pty, &lt;code&gt;TERM=xterm-256color&lt;/code&gt;, startup popups disabled&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Time to a usable screen&lt;/strong&gt; (median of 10 runs):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Repo&lt;/th&gt;
&lt;th&gt;lazygit&lt;/th&gt;
&lt;th&gt;gitui&lt;/th&gt;
&lt;th&gt;Ratio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;630 commits&lt;/td&gt;
&lt;td&gt;88 ms&lt;/td&gt;
&lt;td&gt;11 ms&lt;/td&gt;
&lt;td&gt;8x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;38,077 commits&lt;/td&gt;
&lt;td&gt;232 ms&lt;/td&gt;
&lt;td&gt;11 ms&lt;/td&gt;
&lt;td&gt;21x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;82,180 commits&lt;/td&gt;
&lt;td&gt;416 ms&lt;/td&gt;
&lt;td&gt;11 ms&lt;/td&gt;
&lt;td&gt;38x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Memory (process tree RSS) right after that screen:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Repo&lt;/th&gt;
&lt;th&gt;lazygit&lt;/th&gt;
&lt;th&gt;gitui&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;630 commits&lt;/td&gt;
&lt;td&gt;43 MB&lt;/td&gt;
&lt;td&gt;16 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;38,077 commits&lt;/td&gt;
&lt;td&gt;34 MB&lt;/td&gt;
&lt;td&gt;18 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;82,180 commits&lt;/td&gt;
&lt;td&gt;40 MB&lt;/td&gt;
&lt;td&gt;31 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Those numbers describe two different strategies, not two speeds of the same strategy. Gitui draws the frame first and fills it asynchronously: at 11 ms the status tab is ready, and the log tab is still counting up in a corner (&lt;code&gt;300/3000&lt;/code&gt;, then &lt;code&gt;82180/82180&lt;/code&gt;). Lazygit runs &lt;code&gt;git log&lt;/code&gt; for the first 300 commits, builds the graph, and only then shows you anything. That's why its startup scales with history and gitui's doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What 11 milliseconds actually buys you
&lt;/h2&gt;

&lt;p&gt;The honest framing: gitui's speed advantage is real, and below roughly 100K commits you will not feel it. A 416 ms launch is slower than a keypress but faster than a tab switch in VS Code. Where the strategies diverge for real is when you want the &lt;em&gt;whole&lt;/em&gt; history, and that's where the numbers get dramatic.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Full history loaded&lt;/th&gt;
&lt;th&gt;lazygit&lt;/th&gt;
&lt;th&gt;gitui&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;38,077 commits: time&lt;/td&gt;
&lt;td&gt;under 0.5 s (press &lt;code&gt;&amp;gt;&lt;/code&gt; twice in Commits)&lt;/td&gt;
&lt;td&gt;150 ms from launch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;38,077 commits: RSS&lt;/td&gt;
&lt;td&gt;129 MB&lt;/td&gt;
&lt;td&gt;25 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;82,180 commits: time&lt;/td&gt;
&lt;td&gt;1.92 s&lt;/td&gt;
&lt;td&gt;484 ms from launch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;82,180 commits: RSS&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;716 MB&lt;/strong&gt; (peak 788 MB)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;31 MB&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jump to the first commit ever&lt;/td&gt;
&lt;td&gt;after the load above&lt;/td&gt;
&lt;td&gt;instant, 24 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-02-gitui-vs-lazygit/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Lazygit at 716 MB isn't a leak, it's the commit graph. Lazygit renders the branch-line graph for every loaded commit, and on git/git's merge-heavy history that graph is dozens of columns wide. Gitui doesn't draw a graph at all (branch visualization is &lt;a href="https://github.com/gitui-org/gitui/issues/81" rel="noopener noreferrer"&gt;roadmap item #81&lt;/a&gt;), so it keeps a flat list and a flat memory profile. You are trading a picture of the history for 23x less RAM. On a laptop with 16 GB that's a trade you never notice; on a shared dev box with a 2M-commit monorepo it's the difference between a tool you use and a tool you kill.&lt;/p&gt;

&lt;p&gt;This also puts gitui's own README benchmark in perspective. It quotes the Linux kernel (900K+ commits): 24 s and 0.17 GB for gitui versus 57 s and 2.6 GB for lazygit, with lazygit "freezing" and "sometimes crashing." Those numbers date to a 2020 RustBerlin meetup talk and were measured on lazygit builds several years old; I saw no freezes or crashes at 82K commits on v0.65.0. Extrapolate my memory curve, though, and 2.6 GB at 900K commits is entirely plausible. The README isn't wrong, it's just describing a repo size most of us don't work in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict on performance, dated 2026-09-02:&lt;/strong&gt; if your repo has fewer than ~100K commits, startup speed should not be in your decision at all. If you routinely read the entire history of a Linux-kernel-sized repo, gitui is the only one of the two that does it comfortably.&lt;/p&gt;

&lt;h2&gt;
  
  
  Feature matrix from a week of real use
&lt;/h2&gt;

&lt;p&gt;Everything below I did with my own hands on v0.65.0 and v0.28.1, not from the READMEs. Where a feature is missing I link the issue so you can watch it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;lazygit v0.65.0&lt;/th&gt;
&lt;th&gt;gitui v0.28.1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Interactive rebase (squash, fixup, reorder, drop, edit)&lt;/td&gt;
&lt;td&gt;Yes, in-place in the commits panel&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;No&lt;/strong&gt;; &lt;a href="https://github.com/gitui-org/gitui/issues/32" rel="noopener noreferrer"&gt;#32&lt;/a&gt; open since 2020-04-23, 91 reactions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Line and hunk staging&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;space&lt;/code&gt; / &lt;code&gt;v&lt;/code&gt; range select&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;s&lt;/code&gt; stage lines, Enter stage hunk; equally good&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conflict resolution&lt;/td&gt;
&lt;td&gt;Dedicated view: pick hunk, pick both, next conflict, undo&lt;/td&gt;
&lt;td&gt;Marks file &lt;code&gt;!&lt;/code&gt;, diff panel shows &lt;code&gt;size: 0 B -&amp;gt; 62 B&lt;/code&gt; and nothing else (&lt;a href="https://github.com/gitui-org/gitui/issues/2865" rel="noopener noreferrer"&gt;#2865&lt;/a&gt;); resolve in an external editor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom commands&lt;/td&gt;
&lt;td&gt;YAML &lt;code&gt;customCommands&lt;/code&gt; with Go templates, prompts and menus&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Worktrees&lt;/td&gt;
&lt;td&gt;Files panel has a Worktrees tab; &lt;code&gt;w&lt;/code&gt; creates one from a branch&lt;/td&gt;
&lt;td&gt;No worktree view&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blame&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;None&lt;/strong&gt; (not in the keybindings reference)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;B&lt;/code&gt; in the Files tab, syntax highlighted, go-to-line&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browse the full tree at any commit&lt;/td&gt;
&lt;td&gt;No (only changed files)&lt;/td&gt;
&lt;td&gt;Files tab, any revision&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Commit graph&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amend an old commit / custom patches / bisect / undo&lt;/td&gt;
&lt;td&gt;Yes; &lt;code&gt;ctrl+z&lt;/code&gt; undoes almost anything via the reflog&lt;/td&gt;
&lt;td&gt;Amend HEAD only; no bisect; &lt;code&gt;U&lt;/code&gt; undoes the last commit and that's it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Keybinding discovery&lt;/td&gt;
&lt;td&gt;Context bar at bottom plus &lt;code&gt;?&lt;/code&gt; filterable menu&lt;/td&gt;
&lt;td&gt;Context bar at bottom plus &lt;code&gt;h&lt;/code&gt; help popup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Config format&lt;/td&gt;
&lt;td&gt;YAML (&lt;code&gt;config.yml&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;RON (&lt;code&gt;theme.ron&lt;/code&gt;, &lt;code&gt;key_bindings.ron&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nerd Font&lt;/td&gt;
&lt;td&gt;Optional icons (&lt;code&gt;gui.nerdFontsVersion: "3"&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Not used&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UI languages&lt;/td&gt;
&lt;td&gt;auto-detects zh-CN, zh-TW, ja, ko, ru, pl, nl, pt&lt;/td&gt;
&lt;td&gt;English only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows install&lt;/td&gt;
&lt;td&gt;winget, scoop, choco&lt;/td&gt;
&lt;td&gt;winget, scoop, choco&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Releases, last 12 months&lt;/td&gt;
&lt;td&gt;17 (v0.55.1 on 2025-09-17 through v0.65.0 on 2026-09-05)&lt;/td&gt;
&lt;td&gt;2 (v0.28.0 on 2025-12-14, v0.28.1 on 2026-03-24)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub stars (2026-09-02)&lt;/td&gt;
&lt;td&gt;82,214&lt;/td&gt;
&lt;td&gt;22,477&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three rows on that table decided the article for me.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Interactive rebase.&lt;/strong&gt; I wrote a whole post about why &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-10-lazygit-guide/" rel="noopener noreferrer"&gt;lazygit's rebase feels like cheating&lt;/a&gt;, and nothing in gitui replaces it. Gitui's log tab offers reword, revert, reset and "rebase branch" (a plain &lt;code&gt;git rebase&lt;/code&gt; onto another branch). It does not let you squash two commits, move one above another, or drop one. The 91 thumbs-up on issue #32 say I'm not alone in wanting that, and the fact that it's been open for six years says it isn't coming soon. If you rebase weekly, this row alone ends the comparison.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Conflicts.&lt;/strong&gt; I built a two-branch repo with a one-line conflict and opened both tools. Lazygit switched the files panel to "(only conflicting)", showed the &lt;code&gt;&amp;lt;&amp;lt;&amp;lt;&amp;lt;&amp;lt;&amp;lt;&amp;lt;&lt;/code&gt; / &lt;code&gt;=======&lt;/code&gt; / &lt;code&gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt;&amp;gt;&lt;/code&gt; blocks in the right pane, and the bottom bar read &lt;code&gt;Pick hunk: &amp;lt;space&amp;gt; | Pick both hunks: b | Previous conflict: &amp;lt;left&amp;gt; | Next conflict: &amp;lt;right&amp;gt; | Undo: z&lt;/code&gt;. Gitui marked &lt;code&gt;app.py&lt;/code&gt; with a &lt;code&gt;!&lt;/code&gt;, and the diff pane said &lt;code&gt;size: 0 B -&amp;gt; 62 B (+62 B)&lt;/code&gt; on an otherwise empty panel. The only conflict-specific key was &lt;code&gt;Abort merge&lt;/code&gt;. That's an open bug (&lt;a href="https://github.com/gitui-org/gitui/issues/2865" rel="noopener noreferrer"&gt;#2865&lt;/a&gt;), and the workaround is &lt;code&gt;e&lt;/code&gt; to open your editor, which is exactly the trip a TUI is supposed to save you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blame and tree browsing.&lt;/strong&gt; This is where gitui genuinely wins, and I don't want the rebase point to bury it. Lazygit's Files panel only lists &lt;em&gt;changed&lt;/em&gt; files; there is no way to walk the repository tree and ask "who wrote this line." Gitui's Files tab shows the whole tree at any commit, &lt;code&gt;B&lt;/code&gt; gives you a syntax-highlighted blame with go-to-line, and &lt;code&gt;H&lt;/code&gt; gives per-file history. When I'm reading an unfamiliar codebase, that's the thing I actually want from a Git TUI, and lazygit simply doesn't have it.&lt;/p&gt;

&lt;h2&gt;
  
  
  gitui vs lazygit for AI coding agents
&lt;/h2&gt;

&lt;p&gt;The way I use Git changed in 2025: most diffs on my machine are now written by Claude Code or Codex, and my job is to review them, not author them. That shifts what I need from a Git TUI, and it shifts the comparison hard toward lazygit for three concrete reasons.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hunk-by-hunk review of agent output.&lt;/strong&gt; Both tools stage lines. But an agent-generated diff is where you want to stage the good half of a file, drop the hallucinated half, and commit in logical units the agent didn't think about. Lazygit's &lt;code&gt;v&lt;/code&gt; range select plus its "custom patch" flow (pull lines out of a commit the agent already made) covers that; gitui stops at staging.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One worktree per agent.&lt;/strong&gt; Running two or three agents in parallel means &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-28-claude-code-worktree-guide/" rel="noopener noreferrer"&gt;one worktree per task&lt;/a&gt;, and the annoying part is never creating them, it's remembering which is which and cleaning up. Lazygit's Worktrees tab lists them, switches with &lt;code&gt;space&lt;/code&gt;, and deletes directory plus metadata with &lt;code&gt;d&lt;/code&gt;. Gitui has no worktree concept; the only mentions in its tracker are bugs about running inside one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Custom commands as the glue.&lt;/strong&gt; This is the feature that makes lazygit a hub rather than a viewer. Two bindings I actually use, in &lt;code&gt;config.yml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;customCommands&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="c1"&gt;# Copy the selected commit's diff so I can paste it to the agent for a review pass&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;c-y&amp;gt;'&lt;/span&gt;
    &lt;span class="na"&gt;context&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;commits'&lt;/span&gt;
    &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;show&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;{{.SelectedCommit.Hash}}&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;|&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pbcopy'&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Copy&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;commit&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;diff&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;clipboard'&lt;/span&gt;
  &lt;span class="c1"&gt;# Open a Claude Code session in the selected worktree, in a new tmux window&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;C'&lt;/span&gt;
    &lt;span class="na"&gt;context&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;worktrees'&lt;/span&gt;
    &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;tmux&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;new-window&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;-c&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;{{.SelectedWorktree.Path&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;|&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;quote}}&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;"claude"'&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Start&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Claude&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Code&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;this&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;worktree'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gitui can't express either of those. There is no custom command system, and searching its issues for one turns up nothing beyond a closed build error. For a pure "look at the repo" tool that's fine; for a tool that sits between me and an agent, it's disqualifying.&lt;/p&gt;

&lt;p&gt;The one place gitui earns a slot in an agent workflow is review of &lt;em&gt;unfamiliar&lt;/em&gt; code the agent touched: &lt;code&gt;B&lt;/code&gt; blame on a file to see whether the line it changed was load-bearing and who last touched it. I keep gitui installed for exactly that and open it maybe twice a week.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I hit that neither README mentions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Lazygit shows escaped bytes in diff headers for non-ASCII paths.&lt;/strong&gt; I made a repo with a file named &lt;code&gt;中文文件名.md&lt;/code&gt;. Lazygit's file tree rendered the name correctly, but the diff header read &lt;code&gt;diff --git "a/\344\270\255\346\226\207..."&lt;/code&gt;, because lazygit shells out to &lt;code&gt;git diff&lt;/code&gt; and Git quotes non-ASCII paths by default. One config line fixes it globally, and you should set it anyway if you ever touch CJK, accented or emoji file names:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git config &lt;span class="nt"&gt;--global&lt;/span&gt; core.quotepath &lt;span class="nb"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gitui was clean out of the box because it reads diffs through libgit2 instead of parsing porcelain output. Small thing, but it's the kind of small thing that makes a tool feel broken on day one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lazygit picks your UI language from the locale.&lt;/strong&gt; On my machine, which runs a zh_CN locale, lazygit v0.65.0 launched with a fully translated Chinese interface before I'd touched a config file (&lt;code&gt;gui.language: auto&lt;/code&gt; is the default). Nice for the audience of this blog's Chinese edition, mildly startling if you're following an English tutorial and every keybinding hint is in another language. Set &lt;code&gt;gui.language: en&lt;/code&gt; if you want the docs to match your screen. Gitui has no localization at all.&lt;/p&gt;

&lt;p&gt;Neither hurt gitui, so in fairness, the gitui-specific wart I found is older and documented: over HTTPS it needs &lt;code&gt;credential.helper&lt;/code&gt; set explicitly or push and fetch fail, and on WSL2 there's a &lt;a href="https://github.com/gitui-org/gitui/issues/1974" rel="noopener noreferrer"&gt;still-open report&lt;/a&gt; of fetch and pull hanging while the CLI and lazygit work. I couldn't test Windows or WSL in this session; treat those as things to check on your own box, not as findings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict: install lazygit, keep gitui for blame
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-02-gitui-vs-lazygit/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;My recommendation, as of September 2026, for most people reading this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;brew install lazygit&lt;/code&gt;&lt;/strong&gt; (or &lt;code&gt;winget install -e --id=JesseDuffield.lazygit&lt;/code&gt;, &lt;code&gt;scoop install lazygit&lt;/code&gt;) and make it your daily driver. It's the one with rebase, a real conflict view, worktrees, custom commands, seventeen releases in twelve months and 82K stars. The 416 ms it costs to open on an 82K-commit repo is a price you'll stop noticing by Tuesday.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Also install gitui&lt;/strong&gt; if you read other people's code for a living, work in a repo with hundreds of thousands of commits, or just want &lt;code&gt;B&lt;/code&gt; for blame. It's a 9.5 MB binary, 16 MB of RAM, and it has no opinions about your workflow. Bind it to a separate alias and let it be the microscope while lazygit is the workbench.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pick gitui alone&lt;/strong&gt; only if history rewriting is something you never do and memory on the box is genuinely tight. That's a real population (SRE boxes, Termux on a phone, a 2M-commit monorepo on a shared VM), it's just not most developers.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What I'd skip: choosing by language. "gitui is Rust, lazygit is Go" tells you the binary size (9.5 vs 18.4 MB) and nothing about which one will get you out of a botched merge at 6 pm. Choose by the feature rows above.&lt;/p&gt;

&lt;p&gt;If you're coming from the &lt;a href="https://www.heyuan110.com/posts/linux/2026-04-18-fish-shell-rust-2026/" rel="noopener noreferrer"&gt;fish shell&lt;/a&gt; and Rust-tool crowd and want the rest of the terminal sorted first, my &lt;a href="https://www.heyuan110.com/posts/macos/2025-01-22-terminal-tools-guide/" rel="noopener noreferrer"&gt;terminal emulator roundup&lt;/a&gt; and the &lt;a href="https://www.heyuan110.com/posts/ai/2026-06-22-best-windows-terminal-2026/" rel="noopener noreferrer"&gt;Windows Terminal ranking&lt;/a&gt; cover the layer underneath either of these.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Footnote on tig: yes, it still exists (13.3K stars, v2.6.1 in June 2026), and as a read-only pager over &lt;code&gt;git log&lt;/code&gt; it starts even lighter than gitui. It doesn't stage, rebase or resolve anything, so it isn't a third contender here; it's a very good &lt;code&gt;less&lt;/code&gt; for Git.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-10-lazygit-guide/" rel="noopener noreferrer"&gt;Lazygit in 2026: The Git TUI That Makes Interactive Rebase Feel Like Cheating&lt;/a&gt; — the deep dive on the tool this article picks&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-28-claude-code-worktree-guide/" rel="noopener noreferrer"&gt;Claude Code Worktree: Run Multiple AI Tasks in Parallel&lt;/a&gt; — the workflow lazygit's Worktrees tab was made for&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-22-best-windows-terminal-2026/" rel="noopener noreferrer"&gt;Best Windows Terminal 2026: Ranked for Developers&lt;/a&gt; — pick the terminal before you pick the Git TUI&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/macos/2025-01-22-terminal-tools-guide/" rel="noopener noreferrer"&gt;Best Terminal Emulators in 2025: 23 Tools Compared&lt;/a&gt; — the cross-platform terminal roundup&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/linux/2026-04-18-fish-shell-rust-2026/" rel="noopener noreferrer"&gt;Fish Shell 4.6 Review&lt;/a&gt; — the shell that pairs nicely with both tools&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-09-02-gitui-vs-lazygit/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>developertools</category>
      <category>git</category>
      <category>terminal</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Codex App Guide (2026): Setup, Workflows, and Real Pitfalls</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 10 Jul 2026 10:34:48 +0000</pubDate>
      <link>https://dev.to/bruce_he/codex-app-guide-2026-setup-workflows-and-real-pitfalls-40dh</link>
      <guid>https://dev.to/bruce_he/codex-app-guide-2026-setup-workflows-and-real-pitfalls-40dh</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsvkbuvcoilyv94dynw1l.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsvkbuvcoilyv94dynw1l.webp" alt="OpenAI Codex app guide 2026: setup, parallel agent workflows, pricing, and pitfalls" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Codex app&lt;/strong&gt; is OpenAI's desktop command center for coding agents — and since July 9, 2026, it's no longer a standalone download but a dedicated Codex mode inside the new ChatGPT desktop app, sitting next to Chat and Work. It runs the same agent as the open-source Codex CLI, bills against your existing ChatGPT plan (every tier, including Free), and its specialty is running several agents in parallel, each in its own git worktree, while you review diffs instead of babysitting a terminal. If you already pay for ChatGPT, installing it costs you nothing but disk space, and it's the easiest on-ramp to agentic coding that exists right now.&lt;/p&gt;

&lt;p&gt;That's the three-sentence answer. The rest of this guide is the long version: what the app actually is after the merger, how to set it up properly, the workflows that make it worth the screen real estate, what every plan really gets you, where it loses to the CLI and to Claude Code, and the pitfalls that are currently filling up Hacker News threads.&lt;/p&gt;

&lt;p&gt;One thing before we start, so this page stays honest: I've lived in &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-12-codex-cli-mastery-guide/" rel="noopener noreferrer"&gt;Codex CLI&lt;/a&gt; since February, but I have not run months of production work through the desktop app. Everything here is built from OpenAI's own docs and changelog, the launch and merger announcements, and a lot of community receipts — and where something is contested, I say so.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Version anchor&lt;/strong&gt;: this guide describes the state of the Codex app as of &lt;strong&gt;July 12, 2026&lt;/strong&gt; — three days after the ChatGPT desktop merger. Given that this product changed shape three times in six months, check the &lt;a href="https://developers.openai.com/codex/changelog" rel="noopener noreferrer"&gt;official changelog&lt;/a&gt; if you're reading this from the future. I'll keep this page updated as the app evolves.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is the Codex App in July 2026?
&lt;/h2&gt;

&lt;p&gt;Start with the fact that confuses everyone this week: the Codex app didn't die on July 9 — it got promoted. OpenAI merged the standalone app into a rebuilt ChatGPT desktop app for macOS and Windows, where Codex is now one of three modes (Chat, Work, Codex). Your projects, threads, and settings carried over; existing Codex app installs simply updated into the new shell. The old ChatGPT desktop app was renamed "ChatGPT Classic" — more on why that name is a problem later.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;As of July 2026, the Codex app lives as a dedicated mode inside the ChatGPT desktop app (macOS Apple Silicon and Windows), is included in every ChatGPT plan including Free, and serves over 5 million weekly users — up from roughly 600,000 in January 2026, per OpenAI's own June 2 figures.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That growth curve is the reason this product deserves a pillar guide. OpenAI's numbers: 600K weekly users at the start of 2026, over 2 million by March, 4 million by April 21, past 5 million by June 2. Even more telling: about 20% of those users aren't developers, and the non-developer segment is growing three times faster. Andrew Ambrosino, who leads the Codex desktop team, said in a &lt;a href="https://www.htx.com/news/why-did-codex-and-chatgpt-merge-whats-next-for-codex-openai-qWC9f4eM/" rel="noopener noreferrer"&gt;July 5 interview&lt;/a&gt; that marketing, finance, and legal teams kept adopting a tool built for engineers — which is exactly why OpenAI stopped shipping it as a developer-only product. The merger wasn't a demotion; it was OpenAI concluding that "the boundary between developer tools and general-purpose knowledge tools is collapsing."&lt;/p&gt;

&lt;p&gt;It helps to see where the app sits in the Codex product family, because "Codex" is one agent wearing four different outfits:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-12-openai-codex-app-guide/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The practical takeaway: nothing you learn in the app is wasted in the CLI, and vice versa. Both surfaces read the same &lt;code&gt;AGENTS.md&lt;/code&gt;, share user-level configuration under &lt;code&gt;~/.codex&lt;/code&gt;, and pull from the same plan limits. The app is not a different product with a different brain — it's a control room bolted onto the same engine, and that's precisely how you should use it.&lt;/p&gt;

&lt;h2&gt;
  
  
  From Launch to Merger: Six Months of Whiplash
&lt;/h2&gt;

&lt;p&gt;The Codex app's first half-year explains most of its current quirks, so here's the timeline in one table, straight from OpenAI's changelog and launch posts:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date (2026)&lt;/th&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Feb 2&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://openai.com/index/introducing-the-codex-app/" rel="noopener noreferrer"&gt;Codex app launches on macOS&lt;/a&gt; (v26.202): parallel agent threads, project sidebar, built-in review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mar 4&lt;/td&gt;
&lt;td&gt;Windows version ships (v26.304), with a native Windows sandbox and PowerShell support&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apr 2&lt;/td&gt;
&lt;td&gt;Pricing switches from per-message to API-token-aligned credits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apr 16&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://openai.com/index/codex-for-almost-everything/" rel="noopener noreferrer"&gt;"Codex for almost everything"&lt;/a&gt; update: computer use, in-app browser built on Atlas, image generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;May 21&lt;/td&gt;
&lt;td&gt;v26.519: Goal mode goes GA, remote computer use, plugin marketplace sharing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jul 9&lt;/td&gt;
&lt;td&gt;Codex app merges into the new ChatGPT desktop app as a dedicated mode; old ChatGPT app renamed "Classic"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read that sequence back and the product strategy is legible: launch as a developer tool, discover non-developers love it, add computer use and a browser so it can touch everything, then fold it into the consumer flagship. I covered the strategic fork this creates — OpenAI betting on the super-app while Anthropic bets on composable CLI primitives — in my &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-10-gpt-5-6-general-availability/" rel="noopener noreferrer"&gt;GPT-5.6 release analysis&lt;/a&gt;, and I won't re-litigate it here. What matters for this guide is the practical consequence: &lt;strong&gt;the Codex app's roadmap now answers to a consumer product&lt;/strong&gt;, and developers evaluating it should price in some consumer-first decisions over the next few quarters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Installing and Setting Up the Codex App
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Requirements first, because they exclude people.&lt;/strong&gt; The ChatGPT desktop app runs on macOS with Apple Silicon (M1 or later — Intel Macs are out) and on Windows. There is no Linux build, and OpenAI has announced no plans for one. If you're on Linux, stop here and use the &lt;a href="https://www.heyuan110.com/posts/ai/2026-03-10-codex-cli-deep-dive/" rel="noopener noreferrer"&gt;Codex CLI&lt;/a&gt; — same agent, same plan limits, no GUI.&lt;/p&gt;

&lt;p&gt;Setup is genuinely a five-minute affair:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Download&lt;/strong&gt; the ChatGPT desktop app from &lt;a href="https://chatgpt.com/" rel="noopener noreferrer"&gt;chatgpt.com&lt;/a&gt; (existing Codex app installs auto-update into it).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sign in with your ChatGPT account&lt;/strong&gt; — no API key needed. This is the single biggest onboarding difference from most agent tools: billing rides your existing subscription.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Switch to Codex mode&lt;/strong&gt; and point it at a project folder. Codex treats each folder as a project, with threads organized underneath.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pick your permission posture&lt;/strong&gt; when prompted. Like the CLI, the app runs agents in a sandbox with bounded file and network access, and asks before escalating.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check your &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/strong&gt; If your repo already has one for the CLI, the app reads the same file. If not, write one — repo conventions, build commands, test commands — before you run anything serious.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two configuration details worth knowing on day one. First, user-level settings live in &lt;code&gt;~/.codex/config.toml&lt;/code&gt; and are shared across the app, CLI, and IDE extension — MCP servers you configured for the CLI show up in the app for trusted projects. Second, the app distinguishes &lt;strong&gt;local environments&lt;/strong&gt; (agents work directly on your machine, in worktrees) from &lt;strong&gt;cloud environments&lt;/strong&gt; (tasks run on OpenAI's infrastructure, like Codex web). Cloud tasks are how you fire off work from your phone and review it later, but note that they draw from the same usage window as local messages — a pitfall we'll get to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Core Codex App Workflows: Threads, Worktrees, Parallel Agents
&lt;/h2&gt;

&lt;p&gt;Here's the mental model that makes the app click: &lt;strong&gt;every thread is a workspace, not a chat.&lt;/strong&gt; When you start a new thread in a project, Codex quietly creates a git worktree for it — an independent checkout that shares the same &lt;code&gt;.git&lt;/code&gt; metadata as your main repository. The agent works there, isolated from your working directory and from every other thread. You can have three agents refactoring, bug-fixing, and writing tests simultaneously without any of them stepping on your uncommitted changes.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-12-openai-codex-app-guide/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is, frankly, the app's killer feature — and it's worth naming why. Anyone who has tried to run parallel agents from the terminal knows the tax: manual &lt;code&gt;git worktree add&lt;/code&gt;, one terminal tab per agent, and a cleanup ritual when you're done. The CLI still makes you do all of that yourself (community guides exist precisely because &lt;a href="https://www.frr.dev/posts/codex-cli-worktrees-manual-parallelism/" rel="noopener noreferrer"&gt;the CLI lacks built-in worktree automation&lt;/a&gt;). The app does it silently, per thread, with cleanup handled when threads close. If parallel agent work is your goal, the app isn't a dumbed-down CLI — on this one axis it's ahead.&lt;/p&gt;

&lt;p&gt;The workflow that follows from this: &lt;strong&gt;stop watching agents work.&lt;/strong&gt; The intended loop is queue several threads, go do something that needs your brain, come back and review. The July 9 release sharpened the review half considerably — you can now edit diffs inline inside the app instead of round-tripping through your editor, and a PR review side panel brings the review-before-merge step into the same window. Reviewing three agents' diffs in sequence turns out to be a much better use of attention than watching one agent think in real time.&lt;/p&gt;

&lt;p&gt;Beyond the core loop, three features distinguish the app from every other Codex surface:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Computer use.&lt;/strong&gt; Since April, Codex agents can see your screen and operate any app with their own cursor — and since July 9, this runs on GPT-5.6 and is noticeably faster. Multiple agents can do this in parallel without hijacking your actual mouse. It's the feature with the highest wow-factor and the highest risk surface; scope it to the apps a task actually needs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The in-app browser.&lt;/strong&gt; Built on OpenAI's Atlas engine, it lets an agent iterate on frontend work — run the dev server, look at the page, fix the CSS, look again — without you alt-tabbing to verify. For frontend iteration loops, this quietly removes the most annoying human-in-the-loop step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Automations and Goal mode.&lt;/strong&gt; Scheduled tasks ("run the test suite triage every morning") and goal-directed longer runs went GA in May. This overlaps with what ChatGPT Work does for non-code tasks, and the boundary between them is genuinely blurry — a confusion OpenAI has earned, as the pitfalls section will show.&lt;/p&gt;

&lt;h2&gt;
  
  
  Codex App Pricing: What Every Plan Actually Gets
&lt;/h2&gt;

&lt;p&gt;The headline is friendlier than you'd expect: &lt;strong&gt;Codex is included in every ChatGPT plan&lt;/strong&gt;, and since April 2, 2026, usage is metered in credits that map directly onto API token prices. Here's the plan lineup as of July 2026, per &lt;a href="https://developers.openai.com/codex/pricing" rel="noopener noreferrer"&gt;OpenAI's pricing docs&lt;/a&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;th&gt;Codex allowance (local messages / 5-hour window)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Free&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;Limited taste&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Go&lt;/td&gt;
&lt;td&gt;$8/mo&lt;/td&gt;
&lt;td&gt;Light usage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plus&lt;/td&gt;
&lt;td&gt;$20/mo&lt;/td&gt;
&lt;td&gt;~20–110, depending on model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro&lt;/td&gt;
&lt;td&gt;$100/mo&lt;/td&gt;
&lt;td&gt;5x Plus (~100–550)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro 20x&lt;/td&gt;
&lt;td&gt;$200/mo&lt;/td&gt;
&lt;td&gt;20x Plus (~400–2,200)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Business&lt;/td&gt;
&lt;td&gt;$25/user/mo&lt;/td&gt;
&lt;td&gt;Plus-tier limits per seat&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise/Edu&lt;/td&gt;
&lt;td&gt;Custom&lt;/td&gt;
&lt;td&gt;Credit-based, scales with contract&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The credit rates are where you should actually look, because they tell you which model to run. Per million tokens: &lt;strong&gt;GPT-5.6 Sol costs 125 credits in / 750 out, Terra 62.5 / 375, and Luna 25 / 150&lt;/strong&gt;. Run the arithmetic against OpenAI's API prices ($5/$30, $2.50/$15, $1/$6) and every tier lands on the same exchange rate: &lt;strong&gt;one credit ≈ 4 cents of API value&lt;/strong&gt;. That's a deliberately clean design — your subscription is effectively a prepaid API balance with a generous multiplier, and when you exhaust a 5-hour window, you can buy top-up credits instead of upgrading a tier.&lt;/p&gt;

&lt;p&gt;The practical advice falls out of the table. Luna costs one-fifth of Sol per token, and for routine agent work — test fixes, small refactors, scripted chores — the quality gap rarely justifies a 5x burn rate. My suggested default: &lt;strong&gt;Terra for daily driving, Luna for bulk chores, Sol reserved for the tasks you'd have escalated to a senior engineer.&lt;/strong&gt; This mirrors the tiered-model advice from my &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-10-gpt-5-6-general-availability/" rel="noopener noreferrer"&gt;GPT-5.6 pricing breakdown&lt;/a&gt;, where the benchmark gaps between tiers turned out smaller than the price gaps.&lt;/p&gt;

&lt;p&gt;Two cost gotchas the pricing page won't shout about. Cloud tasks share the same 5-hour window as local messages — delegating to the cloud doesn't buy you extra capacity, just extra parallelism. And image generation burns your included limits "3–5x faster on average," per OpenAI's own docs, so a design-heavy session can eat a window surprisingly fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Codex App vs Codex CLI vs Claude Code
&lt;/h2&gt;

&lt;p&gt;This is the decision most readers came for, so let's do it properly. The app-versus-CLI question is easy because they're the same agent; the Codex-versus-Claude-Code question is a real fork, and I compared the underlying agents in depth in &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-19-claude-code-vs-codex/" rel="noopener noreferrer"&gt;Claude Code vs Codex&lt;/a&gt; — here I'll focus on what the desktop app changes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Codex app&lt;/th&gt;
&lt;th&gt;Codex CLI&lt;/th&gt;
&lt;th&gt;Claude Code&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Form&lt;/td&gt;
&lt;td&gt;GUI mode in ChatGPT desktop&lt;/td&gt;
&lt;td&gt;Open-source terminal agent&lt;/td&gt;
&lt;td&gt;Terminal agent + SDK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platforms&lt;/td&gt;
&lt;td&gt;macOS (Apple Silicon), Windows&lt;/td&gt;
&lt;td&gt;macOS, Linux, Windows&lt;/td&gt;
&lt;td&gt;macOS, Linux, Windows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Parallel agents&lt;/td&gt;
&lt;td&gt;Built-in, automatic worktrees&lt;/td&gt;
&lt;td&gt;Manual worktree setup&lt;/td&gt;
&lt;td&gt;Manual worktrees / subagents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Review&lt;/td&gt;
&lt;td&gt;Inline diff edits, PR panel&lt;/td&gt;
&lt;td&gt;Terminal diffs&lt;/td&gt;
&lt;td&gt;Terminal diffs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Computer use&lt;/td&gt;
&lt;td&gt;Yes, parallel, GPT-5.6&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No (browser via MCP)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scripting / CI&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes — &lt;code&gt;codex exec&lt;/code&gt;, least-privilege sandbox&lt;/td&gt;
&lt;td&gt;Yes — SDK, hooks, headless mode&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Billing&lt;/td&gt;
&lt;td&gt;Any ChatGPT plan, incl. Free&lt;/td&gt;
&lt;td&gt;Same plans, or API key&lt;/td&gt;
&lt;td&gt;Claude Pro/Max, or API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extensibility&lt;/td&gt;
&lt;td&gt;Plugins, MCP, marketplace&lt;/td&gt;
&lt;td&gt;MCP, AGENTS.md, open source&lt;/td&gt;
&lt;td&gt;Skills, MCP, hooks, plugins&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;My read, stated plainly:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Choose the Codex app if&lt;/strong&gt; you already pay for any ChatGPT plan and you want parallel agents with a real review surface — or you're the kind of user who was never going to open a terminal. The marginal cost is zero, the worktree automation is genuinely best-in-class, and for reviewing multiple agents' output it beats every terminal workflow I know.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stay on the CLI if&lt;/strong&gt; your agent work feeds scripts, CI pipelines, or cron jobs. &lt;code&gt;codex exec&lt;/code&gt; runs read-only by default and escalates permissions explicitly — exactly the posture you want for automation, and something the app simply doesn't do. The app also can't be driven programmatically. They share config, so "both" is a legitimate answer: CLI as the backbone, app as mission control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Code holds its ground if&lt;/strong&gt; your workflow is already built on it. Nothing in the July 9 release creates a capability gap that justifies a migration — Claude Code's Skills/hooks/SDK ecosystem remains the deeper toolbox for engineers who script their agents, an argument I made at length in &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;CLI + Skills vs MCP&lt;/a&gt;. The honest asymmetry: OpenAI now offers the better &lt;em&gt;no-terminal&lt;/em&gt; agent experience, Anthropic the better &lt;em&gt;terminal-native&lt;/em&gt; one. Pick the side that matches where you live, and note that Anthropic's answer to the app — Claude Cowork, its desktop working-session product — is aimed at office workers, not at developers running parallel worktrees.&lt;/p&gt;

&lt;h2&gt;
  
  
  Six Pitfalls That Bite New Codex App Users
&lt;/h2&gt;

&lt;p&gt;Every one of these comes with a receipt — most from the &lt;a href="https://news.ycombinator.com/item?id=48849059" rel="noopener noreferrer"&gt;Hacker News thread on the merger&lt;/a&gt;, which is the most useful unfiltered field report available right now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Work mode and Codex mode look identical, and sometimes are.&lt;/strong&gt; The top complaint post-merger: toggling between Work and Codex produces no visible change for many tasks, and OpenAI hasn't clearly documented what differs. Practical rule until they fix the UX: repo work in Codex mode (you get worktrees and diff review), documents and spreadsheets in Work mode, and don't expect the toggle to change the model's brain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The merged app dropped ChatGPT features you may rely on.&lt;/strong&gt; Early merged builds are missing temporary chats, voice mode, Deep Research, and custom GPTs, and chat history is squeezed into a small popup. If those matter to you, keep "ChatGPT Classic" installed alongside — and yes, multiple HN commenters pointed out that naming an app "Classic" reads like a deprecation notice. Treat it as one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Codex writes to your filesystem in places you didn't ask for.&lt;/strong&gt; Users report the app auto-creating folders in &lt;code&gt;~/Documents&lt;/code&gt;. Combine that with per-thread worktrees and you can accumulate surprising disk clutter; close finished threads so cleanup runs instead of leaving a dozen live worktrees behind.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Cloud tasks and images drain the same meter as local work.&lt;/strong&gt; Both share your 5-hour window, and image generation burns it 3–5x faster. If you hit caps mysteriously early, check whether background cloud tasks or image calls are the culprit before blaming the model picker — and remember Sol burns five Lunas per token.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Computer use is powerful and deserves paranoia.&lt;/strong&gt; An agent that can see your screen and click anything is a lateral-movement risk if a task ingests malicious content — security researchers flagged exactly this boundary question when the feature shipped in April. Grant it specific apps per task, never blanket access, and keep it away from password managers and banking tabs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. IT management broke with the bundle ID change.&lt;/strong&gt; On macOS the app now ships as &lt;code&gt;com.openai.codex&lt;/code&gt; instead of &lt;code&gt;com.openai.chatgpt&lt;/code&gt;, which silently breaks MDM policies, allowlists, and automation scripts pinned to the old identifier. If you manage a fleet, update your policies before your users update their apps.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;The Codex app in July 2026 is an easy recommendation with a narrow shape: if you pay for ChatGPT at all, install it, point it at a real repo, and run two parallel threads through the review panel — that experience, not benchmarks, will tell you whether it fits your work. It's the best zero-marginal-cost entry into agentic coding on the market, and its worktree-per-thread design is something even terminal loyalists should steal.&lt;/p&gt;

&lt;p&gt;But keep the roles straight. The app is a control room; the CLI is the engine you can script. The merger that made Codex a tab in a consumer super-app is great for its distribution and unproven for its developer roadmap — the feature gaps and mode confusion of the July 9 build are exactly what shipping a consumer pivot in a hurry looks like. Use the app for what it's uniquely good at today, keep your automation on the CLI, and check back on this page: this product has changed shape three times in six months, and I doubt it's done.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-12-codex-cli-mastery-guide/" rel="noopener noreferrer"&gt;Codex CLI Mastery Guide: 20+ Power Tips&lt;/a&gt; — the terminal side of the same agent&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-19-claude-code-vs-codex/" rel="noopener noreferrer"&gt;Claude Code vs Codex: 8-Dimension Head-to-Head&lt;/a&gt; — the deeper agent-vs-agent comparison&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-10-gpt-5-6-general-availability/" rel="noopener noreferrer"&gt;GPT-5.6 Release: Pricing, ChatGPT Work, and the Codex Merger&lt;/a&gt; — the launch that reshaped this app&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-03-10-codex-cli-deep-dive/" rel="noopener noreferrer"&gt;Codex CLI Deep Dive: Setup, Config, and Power User Tips&lt;/a&gt; — configuration reference that carries over to the app&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-04-cli-skills-vs-mcp/" rel="noopener noreferrer"&gt;MCP vs Skills: Why CLI + Skill Wins the Agent Toolchain&lt;/a&gt; — why I still keep the backbone in the terminal&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-12-openai-codex-app-guide/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>codex</category>
      <category>openai</category>
      <category>aicodingagents</category>
      <category>developertools</category>
    </item>
    <item>
      <title>Claw Code: The Open-Source Claude Code Rewrite That Hit 100K Stars</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 10 Jul 2026 10:31:51 +0000</pubDate>
      <link>https://dev.to/bruce_he/claw-code-the-open-source-claude-code-rewrite-that-hit-100k-stars-26jd</link>
      <guid>https://dev.to/bruce_he/claw-code-the-open-source-claude-code-rewrite-that-hit-100k-stars-26jd</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-04-claw-code-open-source-agent/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On March 31, 2026, a missing &lt;code&gt;.npmignore&lt;/code&gt; entry shipped 512,000 lines of unobfuscated TypeScript to the public npm registry — the entire internal architecture of Claude Code laid bare.&lt;/p&gt;

&lt;p&gt;Two days later, &lt;strong&gt;Claw Code&lt;/strong&gt; launched as a clean-room Python and Rust rewrite. It became the fastest-growing repository in GitHub history, surpassing 100,000 stars in its first hours.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This article examines:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What exactly leaked and how (the npm packaging error)&lt;/li&gt;
&lt;li&gt;What Claw Code is — architecture, design decisions, key differences&lt;/li&gt;
&lt;li&gt;Claude Code vs Claw Code architecture comparison&lt;/li&gt;
&lt;li&gt;Legal and ethical analysis: is a clean-room rewrite actually safe?&lt;/li&gt;
&lt;li&gt;Should you switch from Claude Code? (Honest assessment: probably not yet)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-04-claw-code-open-source-agent/" rel="noopener noreferrer"&gt;Read the full article →&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you found this useful, check out &lt;a href="https://www.heyuan110.com/" rel="noopener noreferrer"&gt;my blog&lt;/a&gt; for more AI engineering guides.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>coding</category>
      <category>claudecode</category>
    </item>
    <item>
      <title>Claude Code vs Cursor vs Copilot 2026: The Definitive Three-Way Comparison</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 10 Jul 2026 10:31:41 +0000</pubDate>
      <link>https://dev.to/bruce_he/claude-code-vs-cursor-vs-copilot-2026-the-definitive-three-way-comparison-4dg7</link>
      <guid>https://dev.to/bruce_he/claude-code-vs-cursor-vs-copilot-2026-the-definitive-three-way-comparison-4dg7</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-04-claude-code-vs-cursor-vs-copilot/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The AI coding tool landscape in April 2026 has consolidated around three clear leaders: &lt;strong&gt;Claude Code&lt;/strong&gt;, &lt;strong&gt;Cursor&lt;/strong&gt;, and &lt;strong&gt;GitHub Copilot&lt;/strong&gt;. After eight months of using all three daily on production codebases, here is the honest comparison.&lt;/p&gt;

&lt;p&gt;Three fundamentally different philosophies:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt;: Terminal-native agent — AI operates at system level&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cursor&lt;/strong&gt;: IDE-native AI — fork of VS Code rebuilt around AI&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Copilot&lt;/strong&gt;: Universal plugin — meets devs where they already work&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Key findings:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code wins at deep reasoning and large refactors (1M token context)&lt;/li&gt;
&lt;li&gt;Cursor wins at speed for routine tasks (Composer 2 at $0.50/M tokens)&lt;/li&gt;
&lt;li&gt;Copilot wins at reach and simplicity ($10/mo, works in any IDE)&lt;/li&gt;
&lt;li&gt;Most productive developers use 2+ tools, not just one&lt;/li&gt;
&lt;li&gt;The Kimi K2.5 controversy means Cursor's transparency is questionable&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-04-claude-code-vs-cursor-vs-copilot/" rel="noopener noreferrer"&gt;Read the full article →&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you found this useful, check out &lt;a href="https://www.heyuan110.com/" rel="noopener noreferrer"&gt;my blog&lt;/a&gt; for more AI engineering guides.&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>cursor</category>
      <category>githubcopilot</category>
      <category>ai</category>
    </item>
    <item>
      <title>Claude Fable 5: When the $10/$50 Flagship Is Worth It</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 10 Jul 2026 10:31:23 +0000</pubDate>
      <link>https://dev.to/bruce_he/claude-fable-5-when-the-1050-flagship-is-worth-it-29o3</link>
      <guid>https://dev.to/bruce_he/claude-fable-5-when-the-1050-flagship-is-worth-it-29o3</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmmff5x9lvzriydj8mxaq.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmmff5x9lvzriydj8mxaq.webp" alt="Claude Fable 5 guide: when the $10/$50 flagship is worth it and how to use it well" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three tasks. That's all it took to burn through 73% of a five-hour usage window — and one of the three never even finished. It happened to Kazike, a Chinese blogger on the $200/month Max 20x plan — the most expensive consumer tier Anthropic sells — during Fable 5's free window. He said it was the first time he'd ever felt token scarcity; in years of shipping code on Opus 4.8, it had never happened once.&lt;/p&gt;

&lt;p&gt;Here's the part that should worry you: that was while Fable 5 was free.&lt;/p&gt;

&lt;p&gt;On July 13, 2026, &lt;strong&gt;Claude Fable 5&lt;/strong&gt; drops out of every subscription plan. That same burn rate now pulls real dollars from a separately funded usage-credit balance: $10 per million input tokens, $50 per million output. No frontier lab has ever done this — ship your best model, let everyone get a taste, then stick a meter on it.&lt;/p&gt;

&lt;p&gt;So this post is about money, start to finish: what a Fable 5 task actually costs, which tasks are worth it, and how to make sure the run you pay for doesn't go sideways. "What should my default model be" is a different question — I covered it in &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-07-best-ai-coding-models-2026/" rel="noopener noreferrer"&gt;my cross-vendor comparison&lt;/a&gt;, and the answer is still Sonnet 5.&lt;/p&gt;

&lt;p&gt;Cards on the table: Fable 5 is a per-task purchase, not a monthly teammate. At $30–100 a run, it only pencils out when one successful run plausibly replaces half a day of your own work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Fable 5 Lives After July 13
&lt;/h2&gt;

&lt;p&gt;Fable 5 didn't disappear — it went from subscription perk to metered add-on. Its status has whipsawed all month, so here's the whole saga, checked against Anthropic's own statements and squeezed into one table:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date (2026)&lt;/th&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;June 9&lt;/td&gt;
&lt;td&gt;GA; in-plan access announced as free through June 22&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;June 12&lt;/td&gt;
&lt;td&gt;US government &lt;a href="https://www.anthropic.com/news/fable-mythos-access" rel="noopener noreferrer"&gt;directive suspends access&lt;/a&gt; for all foreign nationals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;June 30&lt;/td&gt;
&lt;td&gt;Export controls lifted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;July 1&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.anthropic.com/news/redeploying-fable-5" rel="noopener noreferrer"&gt;Redeployed globally&lt;/a&gt;; Pro/Max/Team get it for up to 50% of weekly limits, through July 7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;July 7&lt;/td&gt;
&lt;td&gt;After subscriber backlash, &lt;a href="https://www.forbes.com/sites/sandycarter/2026/07/07/claude-fable-5-extends-by-five-more-days-10-moves-to-make-now/" rel="noopener noreferrer"&gt;extended five days&lt;/a&gt; to July 12, 11:59 PM PT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;July 13 →&lt;/td&gt;
&lt;td&gt;Usage credits only, at API rates, on top of any plan&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read that sequence back: free, banned overnight, restored, pulled again — four states in one month. And the July 7 extension didn't come from goodwill; subscribers forced it, and what they were angry about is the math in the next section.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;As of July 13, 2026, Claude Fable 5 is not included in any Claude subscription plan. It bills through a separate usage-credit balance at $10 per million input tokens and $50 per million output tokens, and Anthropic has announced no date for restoring in-plan access.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two claims making the rounds deserve corrections. First, Fable 5 is &lt;strong&gt;not "API-only"&lt;/strong&gt; after the cutoff. It's still right there in the model picker in Claude.ai and Claude Code; it just bills your usage-credit balance instead of your plan's included limits. You can keep your $20 Pro plan and run Fable 5 tonight — you'll just be watching a dollar meter instead of a usage bar.&lt;/p&gt;

&lt;p&gt;Second, "temporary" is doing a lot of work in Anthropic's framing. The company says it'll restore in-plan access once serving capacity allows — and as of July 12, there's no timeline attached. Budget as if this lasts months, not days.&lt;/p&gt;

&lt;p&gt;Two hard constraints also carry over unchanged from June. Fable 5 has &lt;a href="https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5" rel="noopener noreferrer"&gt;mandatory 30-day data retention and is excluded from zero-data-retention agreements&lt;/a&gt; — ZDR organizations get a 400 on every request, full stop. And its safety classifiers decline with &lt;code&gt;stop_reason: "refusal"&lt;/code&gt; often enough to matter for anything security-adjacent. Its unclassified sibling Mythos 5 remains limited to Project Glasswing partners, so it isn't an option for the rest of us.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fable 5 Pricing in Practice: How $10/$50 Becomes $30–100 a Task
&lt;/h2&gt;

&lt;p&gt;The sticker price is the least useful number in this whole decision. Three multipliers sit between $10/$50 and what you actually spend — here's the math so you can plug in your own numbers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multiplier one: thinking you cannot turn off.&lt;/strong&gt; On Fable 5, &lt;a href="https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5" rel="noopener noreferrer"&gt;adaptive thinking is the only mode&lt;/a&gt; — &lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt; isn't supported, and thinking tokens bill as output at $50 per million. Even a "short answer" carries reasoning overhead, hard turns can think for minutes, and your only throttle is the &lt;code&gt;effort&lt;/code&gt; parameter, which most people never touch.&lt;/p&gt;

&lt;p&gt;The thinking tax deserves its own line item. A 60-turn coding session averaging 5K thinking tokens per turn is 300K tokens of pure thought — $15 before a single line of code gets written. Opus 4.8 ($5/$25) lets you choose when to pay for extended thinking, so the effective output-price gap is much wider than the nominal 2x.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multiplier two: agentic compounding.&lt;/strong&gt; A Claude Code session isn't one API call; it's dozens of turns, each re-reading a growing context. Prompt caching absorbs most of it — cache reads cost about a tenth of fresh input — but the snowball still rolls: a two-hour session over a mid-sized repo routinely racks up tens of millions of cache-read tokens plus hundreds of thousands of output tokens.&lt;/p&gt;

&lt;p&gt;Napkin math for a typical deep task: ~15M cache reads ($15) + 1M fresh input ($10) + 400K output including thinking ($20) — call it $45. That's the anatomy behind my working estimate of $30–100 per serious run. Plug your own usage into my &lt;a href="///tools/claude-token-cost-calculator.html"&gt;Claude token cost calculator&lt;/a&gt; — a 20-turn session and an 80-turn session are the difference between a coffee and a dinner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multiplier three: retries you cause yourself.&lt;/strong&gt; Every run that comes back wrong because you under-specified it is a full-price run. Nobody budgets for this one, and it's the one you control most directly — the last section of this post is entirely about shrinking it.&lt;/p&gt;

&lt;p&gt;With all three multipliers in view, that opening 73% stops being mysterious: one deep Fable 5 task ate roughly a quarter of a Max 20x window, on Anthropic's priciest consumer tier. Carry that into post-July-13 billing and it gets scarier — three deep tasks a day at my per-task estimate works out to about a $3,000 month. The same tasks on Sonnet 5: fifty cents to two bucks each.&lt;/p&gt;

&lt;p&gt;This isn't a price hike. It's a change of unit — from dollars per month to dollars per task.&lt;/p&gt;

&lt;p&gt;Which is why "is Fable 5 worth it" has no answer in the aggregate — it only resolves task by task. (For how the subscription tiers map to real usage generally, see my &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-03-claude-pricing-complete-guide/" rel="noopener noreferrer"&gt;Claude pricing complete guide&lt;/a&gt;.) The next section is the per-task answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Half-Day Test: Which Tasks Deserve Fable 5
&lt;/h2&gt;

&lt;p&gt;The rule I actually use has one clause: &lt;strong&gt;turn on Fable 5 only when a single successful run would plausibly replace at least half a day of your own skilled work.&lt;/strong&gt; At $30–100 per run against $300+ of engineer time, that's a 3–10x return with margin for partial misses. Below the bar, you're paying a 5–20x premium for output you couldn't tell apart from a cheaper model's.&lt;/p&gt;

&lt;p&gt;Three archetypes clear the bar consistently, each with concrete evidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The one-shot feature build.&lt;/strong&gt; Back to Kazike. He wanted a time-decayed "trending" section for his AI news site — clustering, decay weighting, plus the edge case where a quiet news day should collapse the section entirely. He went through two design sessions with Opus 4.8 and walked away unhappy both times; Fable 5 took the same requirement and had it designed, built, and in production in 30 minutes, edge cases included.&lt;/p&gt;

&lt;p&gt;That's the profile: a task with real design judgment in it, where the flagship's extra depth converts directly into not needing you in the loop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The deep analysis report.&lt;/strong&gt; This is the archetype that changed my mind about what the price buys. The same user asked Opus 4.8 to back-test a month of scoring data and got a report he described as insight-free; Fable 5 ran autonomously for 1 hour 18 minutes and produced an analysis that took him 20 minutes to read — flagging problems in his scoring system he had never thought to ask about. Hold onto "never thought to ask"; the last section collects on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The massive mechanical migration.&lt;/strong&gt; This is the category Anthropic's own launch material staked out: Stripe ran a full-library migration across a 50-million-line Ruby codebase in a single day — work a team would have scheduled in months. Few of us have Stripe's problem, but the shape generalizes: enormous, latency-tolerant, machine-checkable, where a per-run fee is noise against the alternative.&lt;/p&gt;

&lt;p&gt;The bar cuts just as hard the other way. Interactive edit-run-fix loops fail twice over: the task value is small, and the always-on thinking latency wrecks the rhythm — while iterating, a model that answers in 15 seconds at 90% beats one that answers in four minutes at 97%. High-volume pipeline work (test generation, lint sweeps, commit messages) fails on value. Security work fails on refusals: Kazike reported Fable 5 declining to audit &lt;em&gt;his own codebase&lt;/em&gt; for vulnerabilities. ZDR organizations fail with a hard 400.&lt;/p&gt;

&lt;p&gt;None of those are edge cases. Together they cover most of a normal working week — which is the honest core of any Fable 5 review: it is simultaneously the strongest model available and the wrong choice for most hours of the day.&lt;/p&gt;

&lt;p&gt;Here's the whole ledger in one flowchart — if you screenshot one thing from this post, make it this:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-10-claude-fable-5-guide/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The most important node is the one that loops: if your unknowns aren't clarified, the tree won't let you spend — it sends you back to a cheap model first. Note also what this tree is not: it doesn't pick your vendor or your default; that's the &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-07-best-ai-coding-models-2026/" rel="noopener noreferrer"&gt;cross-vendor comparison&lt;/a&gt;. This one runs per task, after your default is set.&lt;/p&gt;

&lt;p&gt;Quick reference by task type:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task type&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;th&gt;Cost per run (my estimate)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One-shot feature / frontend build&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Fable 5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;#1 WebDev Arena; one run ships the feature&lt;/td&gt;
&lt;td&gt;$30–60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deep analysis report over data or code&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Fable 5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;surfaces problems you didn't know to ask about&lt;/td&gt;
&lt;td&gt;$50–100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Massive mechanical migration&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Fable 5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Stripe-scale work, machine-checkable&lt;/td&gt;
&lt;td&gt;scale-dependent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Daily agentic coding&lt;/td&gt;
&lt;td&gt;Sonnet 5&lt;/td&gt;
&lt;td&gt;beats Opus 4.8 on Terminal-Bench at 40% of the price&lt;/td&gt;
&lt;td&gt;$0.5–2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deep multi-file refactor&lt;/td&gt;
&lt;td&gt;Opus 4.8&lt;/td&gt;
&lt;td&gt;strong reasoning, no per-use meter&lt;/td&gt;
&lt;td&gt;$5–15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interactive edit-run-fix loop&lt;/td&gt;
&lt;td&gt;Sonnet 5&lt;/td&gt;
&lt;td&gt;Fable's thinking latency kills tight loops&lt;/td&gt;
&lt;td&gt;cents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;High-volume pipelines&lt;/td&gt;
&lt;td&gt;Haiku 4.5&lt;/td&gt;
&lt;td&gt;capability delta ≈ 0, cost delta ~10x&lt;/td&gt;
&lt;td&gt;cents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security / pentest tooling&lt;/td&gt;
&lt;td&gt;Opus 4.8&lt;/td&gt;
&lt;td&gt;Fable's classifiers refuse benign security work&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Brainstorming / clarifying unknowns&lt;/td&gt;
&lt;td&gt;Sonnet 5&lt;/td&gt;
&lt;td&gt;the prep loop for a Fable run&lt;/td&gt;
&lt;td&gt;cents&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Close Your Unknowns Before You Pay
&lt;/h2&gt;

&lt;p&gt;Deciding &lt;em&gt;when&lt;/em&gt; to turn Fable 5 on is half the job; the other half is making sure the paid run is the one that works. The best guidance here comes from inside Anthropic: Thariq Shihipar, an engineer on the Claude Code team, published &lt;a href="https://x.com/trq212/status/2073100352921215386" rel="noopener noreferrer"&gt;A Field Guide to Fable: Finding Your Unknowns&lt;/a&gt; on July 3, and it passed two million views within days on the strength of one sentence: "Fable is the first model where I find the quality of the work is bottlenecked by my ability to clarify its unknowns."&lt;/p&gt;

&lt;p&gt;His frame is simple. Your prompt and context are a map; the codebase and its real constraints are the territory; the gap between them is the unknowns — and when Claude hits an unknown, it decides based on its best guess of what you want. For years the bottleneck was model capability: you pushed the model, and the model was what fell short. Fable 5 flips that. When a run comes back wrong now, the cause is usually a hole in your map.&lt;/p&gt;

&lt;p&gt;My one addition is a cost footnote: at $50 per million output tokens, every unknown the model has to guess at is a line item on your bill. Thariq's own summary is accidentally a billing strategy: "Every explainer, brainstorm, interview, prototype, and reference is a cheap way to find out what you didn't know before it gets expensive to fix."&lt;/p&gt;

&lt;p&gt;He sorts unknowns into four quadrants, each with its own move:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-10-claude-fable-5-guide/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you can state it, state it.&lt;/strong&gt; Known knowns are ordinary spec-writing; no technique required.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you know you haven't decided, get interviewed.&lt;/strong&gt; One prompt does it: "Interview me one question at a time about anything ambiguous. Prioritize questions whose answers would change the architecture." That last clause does the real work — without it, the model asks trivia.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you'd only recognize it on sight, go fishing.&lt;/strong&gt; Unknown knowns are the standards you hold but would never think to write down. Ask for four throwaway design directions in one HTML page with fake data; your reaction to what's wrong &lt;em&gt;is&lt;/em&gt; the missing spec.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you can't even name the question, ask for a blindspot pass.&lt;/strong&gt; Say it straight: "I'm adding an auth provider but I've never touched this codebase's auth module. Do a blindspot pass: what are my unknown unknowns here?" Telling the model what you don't know is the method, not an embarrassment.&lt;/p&gt;

&lt;p&gt;Now to collect on that phrase from the half-day test. The analysis-report archetype is worth $100 precisely because it &lt;em&gt;is&lt;/em&gt; a paid, industrial-strength unknown-unknowns pass — run over your data instead of your prompt. The framework and the cost math aren't two separate pieces of advice; they're the same economics viewed from the prompt side.&lt;/p&gt;

&lt;h3&gt;
  
  
  My Pre-Flight Checklist
&lt;/h3&gt;

&lt;p&gt;The quadrants become a routine once you pin them to a sequence, and my biggest concrete recommendation is one Thariq's guide implies but never states: &lt;strong&gt;run the entire clarification phase on Sonnet 5, and let Fable 5 touch only the final execution run.&lt;/strong&gt; Clarification is conversational, latency-sensitive, many-turn work — everything Fable 5 is bad at and Sonnet 5 is nearly free at. The cheap model asks the questions; the expensive model does the work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before&lt;/strong&gt; (Sonnet 5, ~$1 of tokens):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Blindspot pass over any unfamiliar territory&lt;/li&gt;
&lt;li&gt;Interview, one question at a time, architecture-changing questions first&lt;/li&gt;
&lt;li&gt;Throwaway prototype with fake data for anything visual or UX-shaped&lt;/li&gt;
&lt;li&gt;Reference pointers: "this Rust crate in &lt;code&gt;vendor/rate-limiter&lt;/code&gt; has the backoff semantics I want; reimplement them in our TypeScript client" beats three paragraphs of prose&lt;/li&gt;
&lt;li&gt;An implementation plan with the most volatile decisions at the top — data model changes, type interfaces, anything user-facing — because those are what you'll actually want to veto&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The artifacts that prep loop produces — spec, prototype, references, plan — are exactly the curated context the expensive run consumes. That's &lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;context engineering&lt;/a&gt; in its purest form.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;During&lt;/strong&gt; (Fable 5, the metered part):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Start a fresh session and feed it the artifacts, not the conversation that produced them&lt;/li&gt;
&lt;li&gt;Run one long session, not several short ones — warm cache and intact context beat re-clarifying every restart, and on a per-token meter, restarts are literally money&lt;/li&gt;
&lt;li&gt;Have it keep an &lt;code&gt;implementation-notes.md&lt;/code&gt;: on an edge case, take the conservative option, log the deviation, keep moving&lt;/li&gt;
&lt;li&gt;Define "done" — and when the loop stops — before you launch; designing the loop before paying for it is the whole thesis of &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-05-loop-engineering/" rel="noopener noreferrer"&gt;loop engineering&lt;/a&gt;, and it matters most when every lap has a price&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;After&lt;/strong&gt; (either model):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ask for a walkthrough report, then a quiz on the changes&lt;/li&gt;
&lt;li&gt;Hold Thariq's line: merge only on a perfect score&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A $50 run that ships code you don't understand isn't leverage; it's deferred debugging at flagship prices.&lt;/p&gt;

&lt;p&gt;One last note from my own setup, because I didn't reason my way to these rules — I got burned into them. This blog's entire production pipeline runs on Claude Code — parallel agents drafting posts, generating covers, running validation — and parallelism is exactly where flagship pricing turns dangerous: a 5x per-token premium multiplied across N concurrent agents isn't an upgrade, it's a leak.&lt;/p&gt;

&lt;p&gt;So my standing config is boring on purpose: cheap models by default, Fable 5 behind a manual, per-task escalation, released only for the single deep pass — the site-wide audit, the gnarly feature — where the half-day test genuinely clears. I paid real quota to learn that the thinking tax applies even to tasks that need no thinking. You don't have to pay it again.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Remember
&lt;/h2&gt;

&lt;p&gt;As of July 13, 2026, Fable 5 is a metered add-on: $10/$50, usage credits, no restoration date. Treat it as a per-task purchase whose real price is $30–100 a run, turn it on only for tasks that pass the half-day test — one-shot builds, deep analysis, huge migrations — and let Sonnet 5 run everything else. Before every paid run, spend a dollar of cheap tokens closing your unknowns: with a model this strong, the bottleneck and the bill point at the same thing — your own clarity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-07-best-ai-coding-models-2026/" rel="noopener noreferrer"&gt;Best AI Coding Models 2026: Fable 5 vs Sonnet 5 vs GPT-5.6&lt;/a&gt; — the cross-vendor "what should my default be" question&lt;/li&gt;
&lt;li&gt;
&lt;a href="///tools/claude-token-cost-calculator.html"&gt;Claude Token Cost Calculator&lt;/a&gt; — price your own Fable 5 session before you run it&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-16-context-engineering-2026/" rel="noopener noreferrer"&gt;Context Engineering in 2026&lt;/a&gt; — the artifact-driven prep that makes expensive runs land&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-07-05-loop-engineering/" rel="noopener noreferrer"&gt;Loop Engineering: Designing the Loop Before You Run It&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-03-claude-pricing-complete-guide/" rel="noopener noreferrer"&gt;Claude Pricing Complete Guide: API vs Pro vs Max&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-10-claude-fable-5-guide/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>fable5</category>
      <category>aicodingmodels</category>
      <category>llmpricing</category>
    </item>
    <item>
      <title>What SpaceX's $60B Cursor Acquisition Means for Developers</title>
      <dc:creator>Bruce He</dc:creator>
      <pubDate>Fri, 10 Jul 2026 10:31:07 +0000</pubDate>
      <link>https://dev.to/bruce_he/what-spacexs-60b-cursor-acquisition-means-for-developers-236j</link>
      <guid>https://dev.to/bruce_he/what-spacexs-60b-cursor-acquisition-means-for-developers-236j</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbzyjy0t2gf0vz59lmzjw.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbzyjy0t2gf0vz59lmzjw.webp" alt="SpaceX $60B Cursor acquisition explained: what the Anysphere deal means for developers, Claude model access, and pricing" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The company that catches falling rockets with steel chopsticks just paid $60 billion for a code editor. Not a satellite maker, not a defense prime — a VS Code fork. And it's real: on June 16, 2026, four days after the largest IPO in history, SpaceX announced it would acquire Anysphere, the company behind Cursor, for &lt;strong&gt;$60 billion in stock&lt;/strong&gt;. &lt;a href="https://www.cnbc.com/2026/06/16/spacex-spcx-cursor-acquisition-ipo.html" rel="noopener noreferrer"&gt;CNBC confirmed the deal&lt;/a&gt;, Cursor CEO Michael Truell put his name on an official statement, and the transaction — the largest acquisition of a venture-backed startup ever — is expected to close in Q3 2026, pending regulatory approval. The &lt;strong&gt;SpaceX Cursor acquisition&lt;/strong&gt; is not satire, though I spent a day cross-checking sources before I believed it myself.&lt;/p&gt;

&lt;p&gt;Here's why it isn't a joke: the buyer stopped being a rocket company months ago. SpaceX absorbed xAI in February 2026, which put Grok and the Colossus data centers inside the same ticker as Starship. Read the deal that way and the absurdity evaporates — an AI lab just bought the developer distribution it could never earn on its own. That makes this the most consequential event in developer tooling since GitHub sold to Microsoft, and I think most of the coverage is aimed at the wrong layer. This was never an IDE story. It's the story of the most popular &lt;em&gt;model-neutral&lt;/em&gt; coding tool becoming a wholly owned subsidiary of a company that sells a competing model. So, quick and to the point: the verified facts, one sharp judgment about what dies here, and the concrete moves worth making — because "wait and see" is only a strategy if you know what you're waiting to see.&lt;/p&gt;

&lt;h2&gt;
  
  
  The SpaceX Cursor Acquisition, Fact-Checked
&lt;/h2&gt;

&lt;p&gt;A story this weird breeds embellished retellings fast, so before any opinion, the fact base. Every load-bearing claim below traces to CNBC, Forbes, NPR, CBS News, or an official company statement — not aggregator blogs.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-09-spacex-cursor-acquisition/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Five details carry most of the weight:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The buyer is xAI wearing a spacesuit.&lt;/strong&gt; SpaceX folded xAI into itself in &lt;a href="https://en.wikipedia.org/wiki/Initial_public_offering_of_SpaceX" rel="noopener noreferrer"&gt;February 2026&lt;/a&gt;, then went public on June 12 as SPCX — &lt;a href="https://www.npr.org/2026/06/11/nx-s1-5853199/spacex-ipo-price-elon-musk" rel="noopener noreferrer"&gt;$75 billion raised, the largest IPO ever&lt;/a&gt;, at a $1.77 trillion valuation. Grok and Colossus were already inside the public company before the Cursor deal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It was premeditated, not impulsive.&lt;/strong&gt; Back in April 2026, SpaceX quietly locked in an option: pay roughly $10 billion for a partnership, or buy Anysphere outright for $60 billion later in the year. It exercised the buy side four days after listing. The IPO wasn't adjacent to the acquisition — it minted the currency for it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It's all stock.&lt;/strong&gt; Anysphere shareholders receive SpaceX Class A shares. Truell's statement: &lt;em&gt;"SpaceX has exercised their option to acquire Cursor in an all-stock transaction with the goal of building the world's most useful AI models."&lt;/em&gt; Notice where that sentence lands — models, not editors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cursor is a serious business, not vapor.&lt;/strong&gt; Roughly &lt;strong&gt;$4 billion in annualized revenue&lt;/strong&gt; as of June 2026, up from $1 billion in November 2025; about $2.6 billion of that from enterprise B2B; 64% of the Fortune 500 on the customer list; over a million daily active users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Public markets gagged on it.&lt;/strong&gt; SpaceX stock popped 16% on announcement day — briefly making it the fourth most valuable US company — then reversed and shed roughly &lt;strong&gt;$600 billion in market value over four days&lt;/strong&gt; as investors priced in the dilution and the focus risk.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Facts settled. Now the part that actually matters: what they mean.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a Rocket Company Buying an IDE Isn't a Joke
&lt;/h2&gt;

&lt;p&gt;Start with xAI's position in early 2026, because it explains the whole transaction. The company had frontier-scale compute — Colossus 1 and 2 in Memphis — and almost no developer mindshare. One commenter on &lt;a href="https://news.ycombinator.com/item?id=48553224" rel="noopener noreferrer"&gt;the Hacker News announcement thread&lt;/a&gt; was brutal about it: &lt;em&gt;"xAI, a failed AI company which turned into a datacentre operator probably won't help."&lt;/em&gt; Unkind, but not wrong about the business mix. xAI's best-performing 2026 product is other people's workloads: &lt;a href="https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/" rel="noopener noreferrer"&gt;Anthropic pays $1.25 billion a month&lt;/a&gt; for all of Colossus 1, and Google pays $920 million a month. Renting GPUs to your competitors is a fine business. It just isn't an AI strategy.&lt;/p&gt;

&lt;p&gt;Cursor patches three holes with one signature. &lt;strong&gt;Distribution&lt;/strong&gt;: a million-plus daily active developers and 64% of the Fortune 500, acquired overnight — a funnel Grok couldn't have built in a decade. &lt;strong&gt;Data&lt;/strong&gt;: the interaction stream of working professionals — accepted edits, rejected suggestions, full agentic task traces — is arguably the best coding RLHF corpus outside Anthropic and OpenAI, and several HN commenters converged on this as the real prize: &lt;em&gt;"the data they have flowing through the system is valuable for training."&lt;/em&gt; &lt;strong&gt;Revenue optics&lt;/strong&gt;: $4 billion of fast-growing ARR looks terrific inside a freshly public company trying to justify a $1.7 trillion valuation with something other than launch contracts.&lt;/p&gt;

&lt;p&gt;And if you're hoping the deal collapses under its own price tag, here's the uncomfortable math: &lt;strong&gt;$60 billion in post-IPO stock is close to free money&lt;/strong&gt;. Fifteen times forward revenue is aggressive but not crazy by 2026 AI standards, and SpaceX paid in shares trading at a $2 trillion market cap — inflated paper for real revenue. The market's $600 billion tantrum says investors understood the trade perfectly: rational for Musk, dilutive for them.&lt;/p&gt;

&lt;p&gt;I argued in &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-23-agentic-coding-trends-2026/" rel="noopener noreferrer"&gt;my agentic coding trends piece&lt;/a&gt; that 2026 would consolidate the coding-tool market around whoever owns both the model and the harness. SpaceX just bought the biggest independent harness on the shelf. Which brings us to what actually breaks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Neutrality Is the Real Casualty
&lt;/h2&gt;

&lt;p&gt;Cursor's moat was never the editor. VS Code forks are a commodity — Windsurf proved it, and a dozen lesser clones proved it again. What Cursor sold was &lt;strong&gt;Switzerland&lt;/strong&gt;: the one serious tool where a dropdown routed your task to Claude Opus, GPT, or Gemini, and where an enterprise could point sensitive code at whichever provider its compliance team already trusted. Ownership by a model vendor doesn't dent that neutrality — it deletes it, and no reassuring blog post rewrites the incentive math.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-09-spacex-cursor-acquisition/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The panicked version of this take is wrong too, and worth killing: &lt;em&gt;"Anthropic will cut Claude off from Cursor tomorrow, like it did to Windsurf."&lt;/em&gt; The Windsurf precedent is real — Anthropic yanked model access in 2025 while OpenAI was circling, on the plain logic that you don't arm a competitor buying your distribution channel. But this time the leverage runs the other way. Anthropic has committed roughly &lt;strong&gt;$15 billion a year of SpaceX-xAI compute&lt;/strong&gt; through May 2029, &lt;a href="https://www.datacenterdynamics.com/en/news/anthropic-to-use-all-of-spacex-xais-colossus-1-data-center-compute/" rel="noopener noreferrer"&gt;taking over all of Colossus 1&lt;/a&gt; and expanding into Colossus 2 — a commitment &lt;a href="https://www.anthropic.com/news/higher-limits-spacex" rel="noopener noreferrer"&gt;Anthropic itself announced&lt;/a&gt; as the thing funding higher Claude usage limits. When your landlord buys your biggest reseller, you don't torch the building. Mutual hostage-taking is the most underrated stabilizer in this industry.&lt;/p&gt;

&lt;p&gt;So forget the cutoff. The realistic failure mode is &lt;strong&gt;erosion&lt;/strong&gt;. The joint xAI-Cursor model — &lt;a href="https://finance.yahoo.com/technology/ai/articles/spacexai-plans-launch-model-cursor-210200389.html" rel="noopener noreferrer"&gt;reportedly shipping as early as July 8&lt;/a&gt;, before the deal has even legally closed — becomes the default for new users. Grok gets the fast lane, the deepest agent hooks, the generous tier. Claude and GPT stay on the menu but drift upmarket: pricier tiers, slower capability rollouts, second-class agent support. Cursor already rehearsed this play with its in-house Composer model, which I took apart in &lt;a href="https://www.heyuan110.com/posts/ai/2026-04-01-cursor-composer-2-review/" rel="noopener noreferrer"&gt;my Cursor Composer 2 review&lt;/a&gt; — except Composer was a hedge, and the joint model is now the stated mission. Truell's own words: the goal is "building the world's most useful AI models."&lt;/p&gt;

&lt;p&gt;The model picker used to be the product. It's about to become a migration funnel.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Cursor Users Should Do Before the Deal Closes
&lt;/h2&gt;

&lt;p&gt;Nothing changes in your editor this week, and anyone claiming otherwise is farming clicks — Anysphere operates independently until the Q3 close. But three clocks started ticking on June 16, and they determine what product you'll be using in December.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One: don't sign anything annual.&lt;/strong&gt; The pricing pressure is structural, not hypothetical. SpaceX paid 15x forward revenue with stock the public promptly marked down by $600 billion, and the fastest way to fix that math is to squeeze more out of Cursor's enterprise book ($2.6 billion annualized) and its million-plus daily users. I can't prove prices go up; I can show you every incentive pointing that way — plus HN commenters with enterprise seats already reporting their orgs &lt;em&gt;"have killed their cursor enterprise plans"&lt;/em&gt; within days of the announcement. Stay on monthly billing until post-close pricing lands. The option value costs a few dollars; prepaying for a product that changes owners and priorities mid-contract costs a year of regret.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two: read the privacy fine print like it matters, because now it does.&lt;/strong&gt; Cursor's enterprise pitch leaned hard on privacy mode and SOC 2 — your code goes only to the provider you chose, nothing retained. Under a parent whose stated goal is training "the world's most useful AI models," your interaction data stops being incidental and starts being strategic. Watch the privacy policy and the enterprise DPA for revisions between now and close. And if you work in defense, aerospace-adjacent, or any org with Musk-entity procurement restrictions — they exist, and they're more common than you'd think — compliance may make this whole decision for you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three: assume the defaults will move against you.&lt;/strong&gt; The July 8 joint model is the tell. Onboarding, Auto mode, agent presets — all of it will tilt toward house models, because that tilt is what the $60 billion bought. If your workflow depends on picking Claude Opus for hard refactors — which, per my &lt;a href="https://www.heyuan110.com/posts/ai/2026-01-19-cursor-agent-best-practices/" rel="noopener noreferrer"&gt;Cursor agent best-practices guide&lt;/a&gt;, is exactly what you should be doing today — you're the user this transition treats worst: the tool keeps working while your preferred configuration quietly gets more expensive and less supported.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Should Be Nervous: Winners and Losers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Player&lt;/th&gt;
&lt;th&gt;Before June 16&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;th&gt;Net&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SpaceX-xAI&lt;/td&gt;
&lt;td&gt;Frontier compute, no developer distribution&lt;/td&gt;
&lt;td&gt;1M+ DAU funnel, Fortune 500 foothold, coding data stream&lt;/td&gt;
&lt;td&gt;Big win (if it doesn't fumble the users)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cursor founders/investors&lt;/td&gt;
&lt;td&gt;Private, ~$30B valuation trajectory&lt;/td&gt;
&lt;td&gt;$60B in liquid public stock; founders' net worth doubled&lt;/td&gt;
&lt;td&gt;Enormous win&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;Powers much of Cursor usage; Windsurf precedent available&lt;/td&gt;
&lt;td&gt;Its biggest IDE channel is now owned by a model competitor — that also happens to be its landlord&lt;/td&gt;
&lt;td&gt;Complicated; Claude Code becomes the strategic hedge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;Model supplier to Cursor, lost Windsurf bid in 2025&lt;/td&gt;
&lt;td&gt;Weakest position: pure competitor with no compute entanglement&lt;/td&gt;
&lt;td&gt;Loss&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Copilot / Microsoft&lt;/td&gt;
&lt;td&gt;Losing mindshare to Cursor all year&lt;/td&gt;
&lt;td&gt;Inherits every enterprise that can't stomach Musk ownership&lt;/td&gt;
&lt;td&gt;Quiet win by default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windsurf&lt;/td&gt;
&lt;td&gt;"Cursor but cheaper"&lt;/td&gt;
&lt;td&gt;"Cursor but not SpaceX" — a much better pitch&lt;/td&gt;
&lt;td&gt;Win&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developers&lt;/td&gt;
&lt;td&gt;One great neutral tool&lt;/td&gt;
&lt;td&gt;Neutrality gone; choice moves up a level, to which &lt;em&gt;company&lt;/em&gt; you pick&lt;/td&gt;
&lt;td&gt;Depends on what you do next&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you want the single most nervous party in that table, it's OpenAI: pure competitor, no compute entanglement, no leverage — locked out of the largest third-party surface for its coding models. Second most nervous is any enterprise with Musk-entity restrictions and a Cursor-shaped hole in its toolchain. And the quiet winners didn't lift a finger: Windsurf's pitch upgraded overnight from "Cursor but cheaper" to "Cursor but not SpaceX," while Copilot inherits every org that can't stomach the new owner.&lt;/p&gt;

&lt;p&gt;The bigger industry read: the independent, model-neutral coding harness is going extinct. Every serious harness now belongs to a model vendor — Claude Code to Anthropic, Copilot to Microsoft/OpenAI, Antigravity to Google, and now Cursor to SpaceX-xAI. When I wrote &lt;a href="https://www.heyuan110.com/posts/ai/2026-02-18-claude-code-vs-cursor-vs-windsurf-2026/" rel="noopener noreferrer"&gt;Claude Code vs Cursor vs Windsurf&lt;/a&gt; in February, "Cursor is the neutral option" was a genuine differentiator; M&amp;amp;A just deleted that column from the comparison table. From here on, choosing a coding tool means choosing a model ecosystem. Make that choice on purpose, not by inheriting it from whoever buys your editor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Leave Cursor? A Decision Framework
&lt;/h2&gt;

&lt;p&gt;Here's the framework I'd actually use, instead of the vibes-based "Musk bad, leave now" or "nothing ever changes, stay forever" takes flooding your feed.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(interactive diagram — &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-09-spacex-cursor-acquisition/" rel="noopener noreferrer"&gt;view it on the original post&lt;/a&gt;)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;My own position, since this blog exists to hold positions: &lt;strong&gt;I would not build a new team workflow on Cursor right now.&lt;/strong&gt; Not because the product got worse — it didn't, and its agent tooling is still excellent — but because the stability of its assumptions got worse. Recommending a tool to a team means underwriting the next eighteen months of its roadmap, and Cursor's next eighteen months belong to integrating a trillion-dollar parent, shipping a house model, and justifying a $60 billion price tag. None of that work is for you.&lt;/p&gt;

&lt;p&gt;For individuals, breathe. If Auto mode serves you fine, stay and enjoy it — with Colossus-scale compute behind it, the joint model may genuinely be strong. Just keep your setup portable: your &lt;code&gt;.cursorrules&lt;/code&gt;, your MCP config, your prompts. The habits that make you effective in Cursor's agent transfer almost wholesale to Claude Code and Windsurf, which is exactly why switching costs less than it feels like it will.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;The SpaceX Cursor acquisition is verified fact, rational strategy, and a eulogy for the one thing that made Cursor special — in that order. xAI bought distribution, data, and revenue with inflated post-IPO paper, and the bill lands on the users who valued Cursor precisely as neutral ground. A Windsurf-style instant cutoff probably never comes, because Anthropic's $15-billion-a-year compute entanglement makes cold war more profitable than hot war for everyone involved. What comes instead is slower and harder to headline: defaults shift, tiers reshuffle, and some morning in 2027 the model picker feels less like a menu and more like a suggestion.&lt;/p&gt;

&lt;p&gt;You don't need to leave Cursor this week. You need to make leaving cheap — monthly billing, portable config, one weekend trial of an alternative — so that if the erosion arrives, exiting is an afternoon's decision instead of a quarter's migration. Tool loyalty made sense in an era when tools didn't change owners for $60 billion. That era ended on June 16.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-18-claude-code-vs-cursor-vs-windsurf-2026/" rel="noopener noreferrer"&gt;Claude Code vs Cursor vs Windsurf: The 2026 Comparison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-01-19-cursor-agent-best-practices/" rel="noopener noreferrer"&gt;Cursor Agent Best Practices&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-02-23-agentic-coding-trends-2026/" rel="noopener noreferrer"&gt;Agentic Coding Trends 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-04-01-cursor-composer-2-review/" rel="noopener noreferrer"&gt;Cursor Composer 2 Review: The In-House Model Hedge&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-09-apple-wwdc-2026-gemini-siri-pivot/" rel="noopener noreferrer"&gt;Apple's AI Capitulation at WWDC 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.heyuan110.com/posts/ai/2026-06-12-anthropic-965b-managed-agents/" rel="noopener noreferrer"&gt;Anthropic's $965B Valuation and Managed Agents&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.heyuan110.com/posts/ai/2026-07-09-spacex-cursor-acquisition/" rel="noopener noreferrer"&gt;heyuan110.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cursor</category>
      <category>spacex</category>
      <category>xai</category>
      <category>aicoding</category>
    </item>
  </channel>
</rss>
