<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Stephan Miller</title>
    <description>The latest articles on DEV Community by Stephan Miller (@eristoddle).</description>
    <link>https://dev.to/eristoddle</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F18795%2F5f6c41b8-6033-4887-937a-2ebdfe623d2e.jpeg</url>
      <title>DEV Community: Stephan Miller</title>
      <link>https://dev.to/eristoddle</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/eristoddle"/>
    <language>en</language>
    <item>
      <title>Grok 4.6 Hit the Frontier at Bargain-Bin Prices</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Tue, 18 Aug 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/grok-46-hit-the-frontier-at-bargain-bin-prices-365</link>
      <guid>https://dev.to/eristoddle/grok-46-hit-the-frontier-at-bargain-bin-prices-365</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjj6b8u92s37mdbw9hnry.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjj6b8u92s37mdbw9hnry.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For about a year now the deal has been simple. You want the frontier, you pay frontier prices. Twenty-five, fifty bucks a million tokens. Or you go cheap and you accept that you’re running something a tier down, good enough for most jobs but not the thing topping the leaderboards. Premium or bargain. Pick a lane.&lt;/p&gt;

&lt;p&gt;This week that deal broke from three directions inside about 48 hours. xAI shipped a model that sits at the actual intelligence frontier and charges bargain-bin prices for it. Google shipped a cheap coding workhorse while its actual flagship stayed exactly as imaginary as it was last month. And the open-weights promise I spent all of last week’s post complaining about? It shipped. The weights are real. You can download them this morning.&lt;/p&gt;

&lt;p&gt;And through all of it, the single cheapest model on my boards kept quietly beating things that cost 57 times more. Same model I’ve been recommending since April. Let me walk you through it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Grok 4.6 Walked Up to the Frontier and Sat Down&lt;/li&gt;
&lt;li&gt;Google Shipped Another Flash While the Flagship Stays a Ghost&lt;/li&gt;
&lt;li&gt;The Broken Promise From Last Week Actually Shipped&lt;/li&gt;
&lt;li&gt;Newest Still Isn’t Best&lt;/li&gt;
&lt;li&gt;Cheapskate Picks: Where the Money Actually Is&lt;/li&gt;
&lt;li&gt;Horror Stories From the Wild&lt;/li&gt;
&lt;li&gt;Coming Soon (Or “Soon,” Anyway)&lt;/li&gt;
&lt;li&gt;What I Actually Took Away This Week&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Grok 4.6 Walked Up to the Frontier and Sat Down
&lt;/h2&gt;

&lt;p&gt;Grok 4.6 landed on August 12, about five weeks after 4.5, and it was live in Cursor, Grok Build, and the API the same afternoon. No slow rollout, no waitlist. Here it is, go use it.&lt;/p&gt;

&lt;p&gt;The number that matters: it jumped roughly 5 points on the Artificial Analysis Intelligence Index to land at 61. That ties GPT-5.6 Sol and sits two points behind Claude Opus 5. That is the frontier. Not “the frontier for the price,” not “punching above its weight.” The actual top cluster, the models everyone pays a premium to touch.&lt;/p&gt;

&lt;p&gt;Now the price. Grok 4.6 is $2 in, $6 out per million tokens. Flat, unchanged from 4.5. GPT-5.6 Sol, the model it just tied on intelligence, is $5 in and $30 out. Do that division. For the cost of running one GPT-5.6 Sol job you can run the frontier-tier Grok five times over. &lt;a href="https://artificialanalysis.ai/articles/grok-4-6-benchmarks-and-analysis" rel="noopener noreferrer"&gt;Artificial Analysis put it plainly&lt;/a&gt; in their own headline: Grok 4.6 returns xAI to the intelligence frontier and leads on cost efficiency. The context window is still 500k, unchanged, which is the one spec that didn’t move.&lt;/p&gt;

&lt;p&gt;Here’s the asterisk, because there’s always an asterisk. On AA-Omniscience, the benchmark that measures whether a model knows what it doesn’t know, Grok 4.6 scores 48.2% accuracy and a 65.7% non-hallucination rate. Translate that: when Grok hits something it genuinely doesn’t know, it declines to answer about two times out of three, and it confidently makes something up the other third. A brand-new model at the top of the intelligence index that fabricates on one in three unknowns is not a dealbreaker. But if you’re about to wire it into an agent loop where it can act on its own confident guesses, you want to know that number before, not after.&lt;/p&gt;

&lt;p&gt;Arena hasn’t caught up yet, which is normal. Grok 4.6 is only #44 on the Overall board on a few thousand fresh votes, because Arena is a popularity contest and popularity takes weeks to accumulate. So you get the classic split: the hard benchmarks love it, the blind-taste crowd hasn’t voted. It’s smarter than it looks in the booth. Give the votes a couple weeks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Google Shipped Another Flash While the Flagship Stays a Ghost
&lt;/h2&gt;

&lt;p&gt;On August 13, one day after Grok, Google put out Gemini 3.7 Flash. That’s 23 days after Gemini 3.6 Flash. Three weeks. Google is iterating the cheap tier faster than most labs update their changelog.&lt;/p&gt;

&lt;p&gt;And it’s a real update, not a version-number bump. On Google’s own DeepSWE v1.1 coding benchmark it scores 65.3% against 3.6 Flash’s 49.0%. That’s a 16-point jump on real software-engineering tasks. They didn’t retrain from scratch either, just algorithmic improvements and user feedback layered onto the previous version, then shipped as a full replacement. On the Artificial Analysis index it lands at 56, and it’s one of the fastest models AA tracks, around 344 output tokens a second. Introductory pricing is $0.75 in, $3.75 out through the end of the year. It’s live in AI Studio, Android Studio, Antigravity, and the enterprise platform. On Arena it’s already #3 in Creative Writing and #4 in Math, though both of those ride preliminary vote counts in the hundreds, so treat them as a strong first impression under flattering light.&lt;/p&gt;

&lt;p&gt;Here’s the part that makes me laugh, though. &lt;a href="https://www.axios.com/2026/08/13/google-gemini-37-flash" rel="noopener noreferrer"&gt;Axios noticed it too&lt;/a&gt;: Gemini 3.7 Flash arrived before Gemini 3.5 Pro. The Pro. The flagship. The one Sundar Pichai promised at I/O back on May 19 would land “within a month.”&lt;/p&gt;

&lt;p&gt;That was roughly ninety days ago. &lt;a href="https://www.forbes.com/sites/johnwerner/2026/08/13/gemini-35-pro-delay-continues/" rel="noopener noreferrer"&gt;Forbes ran a piece on August 13&lt;/a&gt; with the deeply original title “Gemini 3.5 Pro Delay Continues.” People have started calling it the longest-awaited model of 2026. The reporting says the team hit a structural problem, scrapped the base model, rebuilt it, and that Google has already started pretraining an entirely new flagship it’s calling Gemini 4. So the flagship is so broken they’re skipping ahead to its replacement while shipping Flash after Flash after Flash to keep the lights on.&lt;/p&gt;

&lt;p&gt;I’m not complaining, exactly. The Flash tier is where the value is and Google’s clearly good at it. But when a company ships three iterations of its budget model in the time its premium model goes from “next month” to “we started building the next one instead,” that tells you where the actual engineering is landing. It’s landing on cheap and fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Broken Promise From Last Week Actually Shipped
&lt;/h2&gt;

&lt;p&gt;If you read last week’s post, you watched me go looking for Qwen3.8-Max’s open weights with my coffee and come back with nothing. Alibaba had promised them “the week of August 10” and the deadline came and went with no license, no files, no new date. I filed it under broken promises and moved on.&lt;/p&gt;

&lt;p&gt;Well. On August 12 they shipped. &lt;a href="https://www.mindstudio.ai/blog/qwen3-8-2-4t-a95b-release" rel="noopener noreferrer"&gt;Qwen3.8-2.4T-A95B went up on Hugging Face and ModelScope&lt;/a&gt;, 2.4 trillion parameters with 95 billion active, alongside the smaller Qwen3.8-27B that actually fits on a workstation. First time Alibaba has ever put a Max-class model in the public’s hands. About two days late, which in this business is basically on time. So the “China promised and stalled” beat from last week only half held. They stalled, then they delivered.&lt;/p&gt;

&lt;p&gt;The independent grade came in at 58 on the intelligence index, which puts it just under the 60-to-63 frontier cluster. Good model, not a record-breaker, and now that the weights are public the real test starts: what do people build with it and how does it hold up under evals that Alibaba didn’t run itself.&lt;/p&gt;

&lt;p&gt;And it wasn’t the only open drop. On August 14, Z.ai quietly released GLM-5.3. It’s not on Arena yet so I can’t rank it, but Z.ai’s GLM line has anchored my cheapskate picks for months, so it goes straight on the watch list. When the boards catch up, we’ll see where it lands.&lt;/p&gt;

&lt;h2&gt;
  
  
  Newest Still Isn’t Best
&lt;/h2&gt;

&lt;p&gt;Quick detour to the top of the boards, because it keeps being true and it keeps mattering. Anthropic owns the leader slot in five of six Arena categories again. But look at &lt;em&gt;which&lt;/em&gt; Anthropic model is winning.&lt;/p&gt;

&lt;p&gt;On the Overall board the leader is Claude Fable 5. On Instruction Following and Hard Prompts, the leader is Opus 4.6, the high-effort variant. Not Opus 5. The newest, most expensive, top-of-the-intelligence-index flagship sits at #7 and #10 on Overall, behind its own older siblings. Opus 5 only clearly wins the Math board, and it does that on 437 votes, which is thin enough to wobble.&lt;/p&gt;

&lt;p&gt;So even inside a single lab, “newest” and “what blind human raters actually prefer” are pointing at different models. The version number went up. The preference didn’t follow. Every time a lab tells you the new one is smarter, remember that smarter on a benchmark and better in your actual work are two different measurements, and the second one is the one you’re paying for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheapskate Picks: Where the Money Actually Is
&lt;/h2&gt;

&lt;p&gt;This is the section I write the whole thing for. Method first, because the table is meaningless without it. For each Arena category, take the leader’s rating, draw a line 50 points below it, and everything above that line is the competitive band. Statistically it’s a coin-flip away from the “best” model. Then sort that band by output price and grab the cheapest thing in it. The mistake I’ve made before is reading the band off the visible top 20. The band is defined by &lt;em&gt;points&lt;/em&gt;, not rank, and it runs way deeper than the first screen. The Overall band this week is 57 models deep. The cheap open-weight stuff lives down in the 30s and 50s, a rounding error behind the premium brands and an order of magnitude cheaper. So I pulled the full tables and computed the bands in code.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Leader&lt;/th&gt;
&lt;th&gt;$ leader&lt;/th&gt;
&lt;th&gt;Cheapskate pick&lt;/th&gt;
&lt;th&gt;$ pick&lt;/th&gt;
&lt;th&gt;Δ rating&lt;/th&gt;
&lt;th&gt;Price ratio&lt;/th&gt;
&lt;th&gt;AA Pareto&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Overall&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#38)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;-38&lt;/td&gt;
&lt;td&gt;~57×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#25)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;-34&lt;/td&gt;
&lt;td&gt;~57×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative Writing&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;Gemini 3.6 Flash (#12)&lt;/td&gt;
&lt;td&gt;$1.88&lt;/td&gt;
&lt;td&gt;-33&lt;/td&gt;
&lt;td&gt;~27×&lt;/td&gt;
&lt;td&gt;nearby&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction Following&lt;/td&gt;
&lt;td&gt;claude-opus-4-6&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#25)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;-44&lt;/td&gt;
&lt;td&gt;~29×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hard Prompts&lt;/td&gt;
&lt;td&gt;claude-opus-4-6&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#28)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;-37&lt;/td&gt;
&lt;td&gt;~29×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math&lt;/td&gt;
&lt;td&gt;claude-opus-5-max&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;Gemini 3.6 Flash (#7)&lt;/td&gt;
&lt;td&gt;$1.88&lt;/td&gt;
&lt;td&gt;-41&lt;/td&gt;
&lt;td&gt;~13×&lt;/td&gt;
&lt;td&gt;nearby&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few things worth saying out loud.&lt;/p&gt;

&lt;p&gt;MiMo v2.5 Pro sweeps three of six again, all at $0.87 output. That’s Xiaomi’s model from April, the one I’ve recommended every week for over a month now. This is the fifth cycle running it’s been the cheapest thing inside the competitive band of multiple categories, and it’s still sitting on &lt;a href="https://artificialanalysis.ai/models/mimo-v2-5-pro" rel="noopener noreferrer"&gt;Artificial Analysis’s Intelligence-vs-Cost Pareto frontier&lt;/a&gt;, index score 43, in the most efficient quadrant. The popularity metric and the capability metric keep independently landing on the same cheap model. That’s the strongest buy signal this newsletter produces, and it simply will not change. On OpenRouter it routes across &lt;a href="https://openrouter.ai/xiaomi/mimo-v2.5-pro" rel="noopener noreferrer"&gt;seven providers with a 30% discount live right now&lt;/a&gt;, no geo-lock, purchasable everywhere. The trade is speed: it’s slow, around 56 tokens a second, so for a tight agent loop where latency compounds you might pay up. For everything else, it’s the boring correct answer.&lt;/p&gt;

&lt;p&gt;Its ranks look scary until you check the vote counts. That #38 Overall rating is backed by more than 50,000 votes. A #28 Hard Prompts slot sits on 33,000. Those are far more settled numbers than the preliminary top-10 entries riding a few hundred votes. Deep rank measures preference, not reliability. Don’t let it spook you.&lt;/p&gt;

&lt;p&gt;Creative Writing and Math flipped this week, and the reason is a price cut. Gemini 3.6 Flash, which I quoted at $3.75 last issue, now shows up on Arena at $0.38 in and $1.88 out. That’s a straight halving, almost certainly Google trimming the old tier the moment 3.7 Flash launched on top of it. At $1.88 it’s the cheapest thing in both the Creative band and the Math band, and it beats every Opus except Fable 5 in Creative. Math is the compressed board as always, only 10 models deep, and the whole thing rides thin preliminary votes, so treat that pick as directional rather than gospel.&lt;/p&gt;

&lt;p&gt;And there’s a wildcard I have to name because the method demands it. On the Overall board, hy3, which is Tencent’s Hunyuan 3, sits at #54 with a $0.53 output price. That’s cheaper than MiMo. But it’s parked right at the leader-minus-50 line with a rating margin wide enough to slip below the cutoff on any given day, on only 4,600 votes. So it’s the cheaper gamble, not the anchor. If you want to ride the very edge of the band to save another thirty cents a million, it’s sitting right there. I’m keeping MiMo as the pick because I like my recommendations boring and my vote counts high.&lt;/p&gt;

&lt;h2&gt;
  
  
  Horror Stories From the Wild
&lt;/h2&gt;

&lt;p&gt;Three this week, running from “new model problem” to “your problem” to “everyone’s problem.”&lt;/p&gt;

&lt;p&gt;First, the one I already flagged: Grok 4.6’s fabrication rate. A frontier model that makes something up on one in three of the things it doesn’t know is fine for a chat window where you can eyeball the answer. It is not fine bolted into an autonomous loop where its confident guess becomes the next tool call. New model, top of the index, still lies to you a third of the time it’s cornered. Know that going in.&lt;/p&gt;

&lt;p&gt;Second, a story that’s becoming the defining nightmare of agentic coding. &lt;a href="https://www.truefoundry.com/blog/llm-cost-attribution-agentic-cicd" rel="noopener noreferrer"&gt;A frontend team deployed a new agent&lt;/a&gt; that hallucinated a missing dependency, then entered a 400-step resolution loop trying to fix a problem that never existed. Every one of those 400 steps re-sent the entire accumulated context. Every step billed. The model invented a problem and then spent your money in a circle trying to solve it.&lt;/p&gt;

&lt;p&gt;Third, the cost story, because it’s the one that gets everyone eventually. &lt;a href="https://leanopstech.com/blog/agentic-ai-cost-runaway-token-budget-2026/" rel="noopener noreferrer"&gt;One developer kicked off an autonomous refactoring run&lt;/a&gt; over a long weekend, on a workload the team hadn’t even validated, and came back to a $4,200 API bill. The broader pattern the piece describes: within about 90 days of switching on coding agents, the AI bill becomes the second-largest line item on the engineering ledger, right after salaries. This is exactly why the cheapskate math isn’t a hobby. When a single unattended session can drain four figures, the difference between $0.87 and $50 a million stops being an abstraction and starts being your quarter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coming Soon (Or “Soon,” Anyway)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 3.5 Pro.&lt;/strong&gt; Still vapor, now roughly ninety days past Google’s “within a month.” Reportedly scrapped and rebuilt over reliability and coding failures, with Google already pretraining Gemini 4 in the background. I’ll believe it when the API returns a token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3.8 open weights.&lt;/strong&gt; These actually shipped on August 12. The story now isn’t the launch, it’s the independent evals. Watch what the community builds and benchmarks over the next couple weeks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GLM-5.3.&lt;/strong&gt; Z.ai dropped it August 14. Not on Arena yet. Given how often GLM models have anchored these picks, it’s the one I’m most curious to see ranked.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grok 4.6 in the EU.&lt;/strong&gt; As usual, xAI’s European rollout is lagging the US launch. If you’re on that side of the Atlantic, expect it later in the month.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I Actually Took Away This Week
&lt;/h2&gt;

&lt;p&gt;The wall between “frontier” and “cheap” is coming down, and it’s coming down fast.&lt;/p&gt;

&lt;p&gt;A year ago the frontier was a walled garden you paid $30 a million to enter. This week a frontier-tier model shipped at $6 output, a cheap coding model got 16 points better in three weeks, and a Max-class model’s weights went public for anyone to download. Meanwhile the flagship everyone waited all summer for is so broken its maker started over, and the $0.87 model from April kept winning categories nobody talks about.&lt;/p&gt;

&lt;p&gt;The launches change. The headlines change. The correct move does not. Open the leaderboards, find the cheapest model inside the competitive band, confirm two different metrics agree it’s actually good, and run that. This week that’s still MiMo v2.5 Pro at 57 times less than the thing at the top of the board. And if you’re feeling brave, Grok 4.6 is now a genuine frontier option at a price that doesn’t require a finance meeting.&lt;/p&gt;

&lt;p&gt;Next Tuesday, same coffee, same two tabs. Some lab will have promised me something by then. I’ll believe that one when I can download it too.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>openrouter</category>
    </item>
    <item>
      <title>How to Sync Obsidian on Android for Free</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Fri, 14 Aug 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/how-to-sync-obsidian-on-android-for-free-41m</link>
      <guid>https://dev.to/eristoddle/how-to-sync-obsidian-on-android-for-free-41m</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft2r0vsjqkygdrpx43xq7.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft2r0vsjqkygdrpx43xq7.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I sync five devices. An iPad is in that pile, and the iPad is what makes it complicated. iOS hands you a sandbox and a shrug.&lt;/p&gt;

&lt;p&gt;Android is not that. It has a real filesystem, real background services, and apps that can touch each other’s folders. Almost everything that makes syncing Obsidian on an iPhone annoying just does not apply. If you came here from &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-across-all-your-devices-including-free-methods-1mi5"&gt;my guide to syncing Obsidian for free on every device&lt;/a&gt; for the Android specifics, this is the platform where the free options actually win.&lt;/p&gt;

&lt;p&gt;The catch is one decision you make before any of it. So let’s start there.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Where Your Vault Lives Decides Everything Else&lt;/li&gt;
&lt;li&gt;Why “Obsidian Can’t See My Folder” Happens&lt;/li&gt;
&lt;li&gt;Syncthing Is the Best Free Answer, and Every Guide Is Now Wrong About It&lt;/li&gt;
&lt;li&gt;FolderSync and Dropsync: The Cloud Folder Route&lt;/li&gt;
&lt;li&gt;Google Drive on Android Is Worse Than It Looks&lt;/li&gt;
&lt;li&gt;The Settings That Silently Kill Your Sync&lt;/li&gt;
&lt;li&gt;Windows Plus Android, Specifically&lt;/li&gt;
&lt;li&gt;
Frequently Asked Questions

&lt;ul&gt;
&lt;li&gt;Where is the Obsidian vault located on Android?&lt;/li&gt;
&lt;li&gt;Can you sync Obsidian on Android for free?&lt;/li&gt;
&lt;li&gt;Does Obsidian sync with Google Drive on Android?&lt;/li&gt;
&lt;li&gt;Can you sync Obsidian between Windows and Android?&lt;/li&gt;
&lt;li&gt;Does Obsidian work with Dropbox on Android?&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The Short Version&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where Your Vault Lives Decides Everything Else
&lt;/h2&gt;

&lt;p&gt;Open Obsidian on a fresh Android install and it asks where to put the vault. Two options: &lt;strong&gt;device storage&lt;/strong&gt; or &lt;strong&gt;app storage&lt;/strong&gt;. It looks like a privacy preference. It is the sync decision, made before you have thought about sync at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Device storage&lt;/strong&gt; puts the vault in the shared filesystem, somewhere like &lt;code&gt;/storage/emulated/0/Documents/Obsidian&lt;/code&gt;, which your file manager shows as plain &lt;code&gt;Documents/Obsidian&lt;/code&gt;. Other apps can see it. Obsidian asks for the “All files access” permission to do this, which feels invasive and is why a lot of people click the other button. &lt;a href="https://help.obsidian.md/mobile" rel="noopener noreferrer"&gt;Obsidian’s own docs&lt;/a&gt; recommend device storage anyway, for compatibility: any tool that syncs files, &lt;a href="https://syncthing.net/" rel="noopener noreferrer"&gt;Syncthing&lt;/a&gt;, &lt;a href="https://play.google.com/store/apps/details?id=dk.tacit.android.foldersync.lite&amp;amp;hl=en_US" rel="noopener noreferrer"&gt;FolderSync&lt;/a&gt;, &lt;a href="https://play.google.com/store/apps/details?id=com.ttxapps.dropsync&amp;amp;hl=en_US" rel="noopener noreferrer"&gt;Dropsync&lt;/a&gt;, a git client, has to be able to see those files.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;App storage&lt;/strong&gt; puts the vault in Obsidian’s private sandbox. No scary permission prompt, better isolation, and genuinely the right answer for some people. It also means nothing outside Obsidian can read the vault. &lt;a href="https://obsidian.md/sync" rel="noopener noreferrer"&gt;Obsidian Sync&lt;/a&gt; works. Plugins that sync over the network, like &lt;a href="https://github.com/remotely-save/remotely-save" rel="noopener noreferrer"&gt;Remotely Save&lt;/a&gt;, work, because they run inside Obsidian. Syncthing does not. FolderSync does not. No external app does. It is the whole point of the sandbox.&lt;/p&gt;

&lt;p&gt;And the part worth putting in bold: &lt;strong&gt;if you use app storage and uninstall Obsidian, the local vault is deleted with it.&lt;/strong&gt; Android deletes app-private data on uninstall. Your other devices keep their copies if you were syncing, but the phone’s copy is gone.&lt;/p&gt;

&lt;p&gt;So the branch is simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Want to use Syncthing, FolderSync, Dropsync, or a git client? &lt;strong&gt;Device storage.&lt;/strong&gt; Grant the permission.&lt;/li&gt;
&lt;li&gt;Only ever going to use Obsidian Sync or a sync plugin? &lt;strong&gt;App storage&lt;/strong&gt; is fine and slightly tidier.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you picked wrong, you are not stuck, but Obsidian will not move a vault for you. Make a new vault in the other location, copy the notes across with a file manager or from your desktop copy, and open that one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why “Obsidian Can’t See My Folder” Happens
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsium1ax9hv1a1thlpecm.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsium1ax9hv1a1thlpecm.jpg" alt="Why" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The other half of this is scoped storage, Android’s permission model for file access, which Obsidian has been fighting for years. Obsidian’s docs are blunt about it: scoped storage runs a permission check on every single file operation, which tanks performance in an app touching hundreds of small markdown files, and it gives no way to watch for external changes.&lt;/p&gt;

&lt;p&gt;That second one is what bites sync users. Watching for external changes is exactly what you need when another app is writing to your vault in the background. It is why Obsidian wants the broad “All files” permission instead of the narrow folder picker, and why a vault Obsidian only has partial access to behaves like it is haunted.&lt;/p&gt;

&lt;p&gt;Practical version: put the vault somewhere boring in shared storage. Not inside another app’s private directory, and not on an SD card if you can avoid it. Then point your sync tool at that same path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Syncthing Is the Best Free Answer, and Every Guide Is Now Wrong About It
&lt;/h2&gt;

&lt;p&gt;Syncthing is peer to peer. Your devices talk directly, there is no cloud account, no storage cap, no monthly anything, and on Android it runs as a real background service instead of begging the OS for scraps. For a vault of text files it is close to ideal, and it is what I would use if I did not have an iPad in the mix. (For the iPad half of that problem, see &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-on-iphone-and-ipad-for-free-in-2026-427p"&gt;syncing Obsidian on iPhone and iPad for free&lt;/a&gt;.)&lt;/p&gt;

&lt;p&gt;Here is the thing every existing Android guide gets wrong, including a fair number published this year: &lt;strong&gt;the &lt;a href="https://github.com/syncthing/syncthing-android" rel="noopener noreferrer"&gt;official Syncthing Android app&lt;/a&gt; was discontinued.&lt;/strong&gt; Its final release shipped with the December 2024 version of Syncthing. The maintainers &lt;a href="https://forum.syncthing.net/t/discontinuing-syncthing-android/23002" rel="noopener noreferrer"&gt;cited Google Play publishing friction and a lack of active development&lt;/a&gt;. If a tutorial tells you to grab “Syncthing” from the Play Store, that tutorial is pointing you at an abandoned app.&lt;/p&gt;

&lt;p&gt;The successor is &lt;strong&gt;&lt;a href="https://github.com/researchxxl/syncthing-android" rel="noopener noreferrer"&gt;Syncthing-Fork&lt;/a&gt;&lt;/strong&gt;, which has been the better Android build for years anyway. Its maintainer, Catfriend1, then retired and handed the project to researchxxl, in a way that briefly made the repo vanish from GitHub and made everyone reasonably nervous. That got sorted out in public: F-Droid and Syncthing developers reviewed the new repo, found no sign of anything malicious, confirmed the old commits were unaltered, and the builds are reproducible, which is the actual defense against a supply chain attack.&lt;/p&gt;

&lt;p&gt;I am telling you that instead of just handing you a link because it is your notes. Install it from &lt;strong&gt;&lt;a href="https://f-droid.org/packages/com.github.catfriend1.syncthingfork/" rel="noopener noreferrer"&gt;F-Droid&lt;/a&gt;&lt;/strong&gt;, which builds from source and has a slower review window between updates. That is exactly the property you want on a project that just changed hands.&lt;/p&gt;

&lt;p&gt;Setup, once you have it:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F24ale56cur3sdh6cvp7k.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F24ale56cur3sdh6cvp7k.jpg" alt="Syncthing Is the Best Free Answer, and Every Guide Is Now Wrong About It" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Install Syncthing-Fork on the phone from F-Droid, and &lt;a href="https://syncthing.net/downloads/" rel="noopener noreferrer"&gt;Syncthing&lt;/a&gt; on your desktop.&lt;/li&gt;
&lt;li&gt;Open the desktop UI, which is the part that throws everyone. Syncthing has no window and no icon. It runs as a background service and you configure it in a browser at &lt;code&gt;http://localhost:8384&lt;/code&gt;. That page is the whole app.&lt;/li&gt;
&lt;li&gt;On that page, click Add Folder, point it at your vault, and note the Folder ID it generates.&lt;/li&gt;
&lt;li&gt;Pair the devices. On the desktop, Actions, then Show ID, which gives you a long string and a QR code. On the phone, Devices tab, the plus button, then scan that code. The desktop throws a prompt asking whether it should talk to this new device. Accept it.&lt;/li&gt;
&lt;li&gt;Back on the desktop, edit the vault folder, open its Sharing tab, and tick the phone. Now the phone gets a notification asking where to put the folder. That is where you point it at your device-storage vault location.&lt;/li&gt;
&lt;li&gt;Add &lt;code&gt;.obsidian/workspace.json&lt;/code&gt; and &lt;code&gt;.obsidian/workspace-mobile.json&lt;/code&gt; to the ignore patterns &lt;strong&gt;on both devices.&lt;/strong&gt; Ignore patterns in Syncthing are per device, not per folder, so setting them on the desktop does exactly nothing for the phone. Both live in the same place: edit the folder, Ignore Patterns tab. These two files are per-device UI state and they change constantly, which means they conflict constantly. Nothing else in &lt;code&gt;.obsidian&lt;/code&gt; needs excluding, and you want your plugins and settings to sync.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then let it settle before you decide it’s broken. Give it ten minutes before you touch anything. A text vault is small, but the first pass still has to hash every file, and both ends show a percentage while it works. Watch that number move instead of assuming the thing is dead.&lt;/p&gt;

&lt;h2&gt;
  
  
  FolderSync and Dropsync: The Cloud Folder Route
&lt;/h2&gt;

&lt;p&gt;If your notes already live in &lt;a href="https://www.dropbox.com/" rel="noopener noreferrer"&gt;Dropbox&lt;/a&gt; or &lt;a href="https://www.microsoft.com/en-us/microsoft-365/onedrive/online-cloud-storage" rel="noopener noreferrer"&gt;OneDrive&lt;/a&gt; on the desktop and you want the phone to mirror that, this is the path. Both apps do the same job: watch a local folder, watch a remote cloud folder, keep them the same on a schedule.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dropsync&lt;/strong&gt; is the Dropbox specialist, now published by &lt;a href="https://www.metactrl.com/" rel="noopener noreferrer"&gt;MetaCtrl&lt;/a&gt;, and it is actively maintained. It picked up updates through spring of this year, which is more than you can say for a lot of Android utilities.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FolderSync&lt;/strong&gt; is the generalist and talks to a long list of providers. Its pricing is genuinely confusing, so: there are two separate apps on Google Play. FolderSync (free, ad supported, with a Premium in-app purchase) and &lt;a href="https://play.google.com/store/apps/details?id=dk.tacit.android.foldersync.full&amp;amp;hl=en_US" rel="noopener noreferrer"&gt;FolderSync Pro&lt;/a&gt; (paid up front, a few dollars). Buying Premium in the free app gets feature parity with Pro but does not let you download Pro, because Play treats them as different apps. The license follows your Google account. Pick one and stay there.&lt;/p&gt;

&lt;p&gt;Setup is the same shape either way: create an account connection, create a folderpair, set local folder to your vault and remote folder to the vault in the cloud, sync type two-way, then a schedule.&lt;/p&gt;

&lt;p&gt;And now the caveat that makes this second choice rather than first. &lt;strong&gt;This is scheduled sync, not live sync.&lt;/strong&gt; Even at the most aggressive interval there is a window where phone and desktop disagree. Edit a note on the phone, walk to the laptop before the sync fires, edit it again there, and you have made a conflict by hand. Syncthing pushes on change and mostly dodges this. Cloud-folder sync does not.&lt;/p&gt;

&lt;p&gt;Also, and I will keep saying this until it stops eating people’s notes: &lt;strong&gt;run exactly one sync system.&lt;/strong&gt; Not Dropsync plus Syncthing “for redundancy.” Two schedulers writing the same files without knowing about each other is a race, and your note is the prize.&lt;/p&gt;

&lt;h2&gt;
  
  
  Google Drive on Android Is Worse Than It Looks
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdomh7fq4pobtnx58n0ff.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdomh7fq4pobtnx58n0ff.jpg" alt="Google Drive on Android Is Worse Than It Looks" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This one gets its own section purely because of how reasonable it sounds. You have an Android phone, &lt;a href="https://www.google.com/drive/" rel="noopener noreferrer"&gt;Google Drive&lt;/a&gt; is right there, free and already signed in. Obviously that is the answer.&lt;/p&gt;

&lt;p&gt;It is not, because Google Drive has no two-way folder sync on Android. &lt;a href="https://www.google.com/drive/download/" rel="noopener noreferrer"&gt;Drive for desktop&lt;/a&gt; does that job on a computer and has no Android equivalent. The Drive app streams files on demand. It does not keep a plain folder on disk that Obsidian can open as a vault. Getting there means going around it: FolderSync connects to Drive as a provider, or you use the &lt;a href="https://github.com/stravo1/obsidian-gdrive-sync" rel="noopener noreferrer"&gt;Google Drive Sync plugin&lt;/a&gt;, which is still beta, is not in the community plugin directory, installs manually or through &lt;a href="https://github.com/TfTHacker/obsidian42-brat" rel="noopener noreferrer"&gt;BRAT&lt;/a&gt;, and whose own README carries a data loss warning I quote rather than paraphrase in the &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-across-all-your-devices-including-free-methods-1mi5"&gt;main sync guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you are on Android and you want a cloud-storage-shaped answer, Dropbox with Dropsync is the one that works. Drive is the one that looks like it should work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Settings That Silently Kill Your Sync
&lt;/h2&gt;

&lt;p&gt;You will set all this up, confirm it works, pocket the phone, and find it three hours stale. That is not the sync tool. That is Android’s battery optimizer deciding your background service is a freeloader, and manufacturer skins are far more aggressive about it than stock Android.&lt;/p&gt;

&lt;p&gt;For whichever sync app you chose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Battery optimization: turn it off for that app.&lt;/strong&gt; Roughly Settings, Apps, the app, Battery, then Unrestricted, though the wording moves around by Android version and manufacturer. This is the one that matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disable background restrictions&lt;/strong&gt; , a separate toggle that is easy to miss.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Allow background data&lt;/strong&gt; , including on metered connections if you want sync off wifi.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lock the app in the recents view&lt;/strong&gt; if your phone has that. Open recents, long press the app’s card or its icon, and pick the lock or pin option. Samsung, Xiaomi, and OnePlus kill “unused” background apps on their own schedule regardless of the standard settings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Syncthing-Fork specifically:&lt;/strong&gt; check its run conditions. It can be set to run only on wifi, only while charging, or only on certain networks, and the battery-friendly defaults are not always what you want.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If sync works while the screen is on and stops when it isn’t, you have found your problem. It is always this.&lt;/p&gt;

&lt;h2&gt;
  
  
  Windows Plus Android, Specifically
&lt;/h2&gt;

&lt;p&gt;This combination comes up constantly and gets bad advice, usually some variant of “just use iCloud, there’s a Windows app.” Do not. Obsidian’s own documentation warns that iCloud Drive on Windows can lead to file duplication or corruption, which I dug into in &lt;a href="https://dev.to/eristoddle/obsidian-icloud-sync-in-2026-including-the-windows-problem-4503"&gt;the iCloud sync post&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For Windows plus Android with no Apple device in the mix, ranked:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbjea16txv9lpgrh56lvj.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbjea16txv9lpgrh56lvj.jpg" alt="Windows Plus Android, Specifically" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Syncthing.&lt;/strong&gt; Both platforms are first class, it is free with no storage cap, and there is no Apple sandbox to design around. This combo is where Syncthing is at its least annoying. On Windows, get it through &lt;a href="https://github.com/GermanCoding/SyncTrayzor" rel="noopener noreferrer"&gt;SyncTrayzor&lt;/a&gt;, which wraps Syncthing in an actual tray app instead of leaving you with a background service and a browser tab. And note the pattern repeating: the original SyncTrayzor is no longer maintained either, and the v2 fork is the one Syncthing’s own docs point at.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dropbox plus Dropsync.&lt;/strong&gt; If the notes are already in Dropbox, this is less work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remotely Save&lt;/strong&gt; , pointed at storage you already have: S3, Backblaze B2, any WebDAV server, Dropbox, or OneDrive. Runs inside Obsidian, so it works with app storage too, which none of the others do.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/Vinzent03/obsidian-git" rel="noopener noreferrer"&gt;Git&lt;/a&gt;&lt;/strong&gt;, if you already live in git. It is a real option on Android and a worse one than you expect on mobile generally, which I went through in &lt;a href="https://dev.to/eristoddle/obsidian-git-sync-in-2026-what-actually-works-on-mobile-46ad"&gt;what actually works for Obsidian git sync on mobile&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Where is the Obsidian vault located on Android?
&lt;/h3&gt;

&lt;p&gt;Wherever you put it, and that is the point. Device storage puts it in shared storage where other apps can reach it, typically under a &lt;code&gt;Documents&lt;/code&gt; folder. App storage puts it in Obsidian’s private directory, invisible to other apps and deleted if you uninstall Obsidian. Check which one you picked before troubleshooting a sync tool that cannot find your files. Open the vault switcher in Obsidian and look at the path under the vault name. If it starts with &lt;code&gt;/storage/emulated/0/&lt;/code&gt;, you are in device storage.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can you sync Obsidian on Android for free?
&lt;/h3&gt;

&lt;p&gt;Yes, more easily than on iOS. Syncthing is free with no storage limit, Remotely Save is free and connects to storage you already have, and Dropsync and FolderSync have free tiers that cover a vault of text notes. You never need an Obsidian Sync subscription on Android.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Obsidian sync with Google Drive on Android?
&lt;/h3&gt;

&lt;p&gt;Not directly. The Drive app streams files rather than keeping a real synced folder on disk, so there is nothing for Obsidian to open. You get there through FolderSync using Drive as a provider, or the third-party Google Drive Sync plugin, which is still beta and ships with a data loss warning. Dropbox is the smoother cloud option on Android.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can you sync Obsidian between Windows and Android?
&lt;/h3&gt;

&lt;p&gt;Yes, and it is one of the easier combinations because neither platform has an iOS-style sandbox. Syncthing is the best free answer. Dropbox with Dropsync works if your vault already lives in Dropbox. Avoid iCloud on Windows for this.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Obsidian work with Dropbox on Android?
&lt;/h3&gt;

&lt;p&gt;Yes, through Dropsync or FolderSync rather than the Dropbox app itself, which does not keep a persistent local folder Obsidian can use as a vault. You need a sync utility to mirror the cloud folder to local storage.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Short Version
&lt;/h2&gt;

&lt;p&gt;Pick device storage. Grant the “All files” permission even though it feels wrong. Install Syncthing-Fork from F-Droid, not the dead official app. Ignore the two workspace files. Turn off battery optimization for whatever you installed. Run one sync system and only one.&lt;/p&gt;

&lt;p&gt;That is a genuinely free, genuinely reliable setup, and the only reason my own phone is not running it is that I own an iPad, which is a sentence that explains most of my sync decisions.&lt;/p&gt;

</description>
      <category>obsidian</category>
      <category>android</category>
      <category>sync</category>
      <category>syncthing</category>
    </item>
    <item>
      <title>The Open-Weights Promise Qwen Broke (and Meta Kept)</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Tue, 11 Aug 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/the-open-weights-promise-qwen-broke-and-meta-kept-44m4</link>
      <guid>https://dev.to/eristoddle/the-open-weights-promise-qwen-broke-and-meta-kept-44m4</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8wk0r9kup8tk45f5dtlw.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8wk0r9kup8tk45f5dtlw.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Last week Alibaba dropped Qwen3.8-Max and promised the open weights were landing “next week.” Well. It’s next week. I went looking for them this morning with my coffee, and there’s nothing on Hugging Face, no license, and no new date. The 2.4-trillion-parameter monster everyone lost their minds over is still API-only, still a black box, still grading its own homework.&lt;/p&gt;

&lt;p&gt;Meanwhile, the other side of the world did the exact opposite. Meta shipped a coding model and a terminal agent, then turned around and said it’s going to &lt;em&gt;open&lt;/em&gt; the weights. So this week the script flipped. China promised open and didn’t deliver. The US shipped and pledged to open up. And me? I’m still running the same $0.87 model from April that I was running last week, and the week before that, and the week before that.&lt;/p&gt;

&lt;p&gt;Let me walk you through it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Open-Weights Promise That Evaporated&lt;/li&gt;
&lt;li&gt;Meanwhile, Meta Did the Opposite&lt;/li&gt;
&lt;li&gt;The Catch in Meta’s Cheap Tier&lt;/li&gt;
&lt;li&gt;The Boring Answer, Week Four&lt;/li&gt;
&lt;li&gt;Cheapskate Picks: Where the Money Actually Is&lt;/li&gt;
&lt;li&gt;Horror Stories From the Wild&lt;/li&gt;
&lt;li&gt;Coming Soon (Or “Soon,” Anyway)&lt;/li&gt;
&lt;li&gt;What I Actually Took Away This Week&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Open-Weights Promise That Evaporated
&lt;/h2&gt;

&lt;p&gt;Here’s where we left Qwen3.8-Max last Tuesday. Genuinely impressive spec sheet. 2.4 trillion parameters, 95 billion active, a million-token context window, priced at roughly a quarter of Claude Opus 5. It rocketed onto the Arena boards on day one. And Alibaba said the weights, plus a smaller Qwen3.8-27B, would go public &lt;a href="https://www.datacamp.com/blog/qwen3-8-max" rel="noopener noreferrer"&gt;the week of August 10 on Hugging Face and ModelScope&lt;/a&gt;. First open-weight release at Max scale for the Qwen line. Big deal.&lt;/p&gt;

&lt;p&gt;That week is now. &lt;a href="https://byteiota.com/qwen3-8-open-weights-drop-this-week-read-before-you-download/" rel="noopener noreferrer"&gt;The weights haven’t appeared, no license has been named, and Alibaba hasn’t given a new date&lt;/a&gt;. The 27B is vapor too. No architecture details, no context length, no license, nothing.&lt;/p&gt;

&lt;p&gt;The good news for the model, if you’re keeping score, is that it finally got an independent grade. Last week the only benchmarks were Alibaba’s own, in a table that helpfully scored its competitors for them. This week Artificial Analysis actually ran it and &lt;a href="https://benchlm.ai/benchmarks/artificialanalysis" rel="noopener noreferrer"&gt;landed it at 56 on the Intelligence Index&lt;/a&gt;. Which is fine. It’s a good model. But 56 sits below the top cluster of Opus 5 at 61, Fable 5 at 60, GPT-5.6 Sol at 59, and Kimi K3 at 57. So the one hard number we got came in under the launch-day hype, and the open weights that were supposed to let anyone verify the rest didn’t show. If you rewired your stack around that press release, this is your reminder not to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meanwhile, Meta Did the Opposite
&lt;/h2&gt;

&lt;p&gt;While Alibaba was quietly not shipping, Meta was loudly shipping. On August 5 it put out &lt;a href="https://www.developersdigest.tech/blog/meta-muse-code-spark-1-2-release" rel="noopener noreferrer"&gt;Muse Spark 1.2 and a terminal coding agent called Muse Code&lt;/a&gt;. The model is a coding specialist with a million-token context, built for the whole plan-execute-validate-fix loop across a big codebase instead of spitting out isolated snippets. It debuted at #4 on Arena Overall and #8 on the Coding board, though both of those ratings are riding preliminary vote counts in the low thousands, so treat them as a first impression with good lighting.&lt;/p&gt;

&lt;p&gt;Then, on August 10, Meta announced it’s going to open Muse Spark 1.2’s weights. It already opened a smaller sibling, Muse Glimmer, which is now &lt;a href="https://cryptobriefing.com/meta-muse-glimmer-spark-release/" rel="noopener noreferrer"&gt;showing up on the trackers&lt;/a&gt;. If the Spark 1.2 release actually lands, it’d be the strongest US open-weight model going right now.&lt;/p&gt;

&lt;p&gt;Sit with the reversal for a second, because it’s the real story this week. For a solid year the pattern was: American labs ship closed, Chinese labs give the good stuff away for pennies. This week Alibaba dangled open weights and pulled them back, and Meta is the one queuing up a genuine open release. One missed week isn’t a cancellation, and a pledge isn’t a release, so don’t over-read it. But the roles inverted, and that’s worth noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Catch in Meta’s Cheap Tier
&lt;/h2&gt;

&lt;p&gt;Now for the part where the cheapskate in me perks up and then immediately gets suspicious. Muse Code has a “contributor” tier. Standard pricing on Muse Spark 1.2 is $1.25 in and $4.25 out per million tokens. The contributor tier is &lt;a href="https://codersera.com/blog/muse-code-contributor-tier-privacy-2026/" rel="noopener noreferrer"&gt;$0.10 in and $0.20 out&lt;/a&gt;. That’s roughly 12 times cheaper on input and 21 times cheaper on output. My kind of number.&lt;/p&gt;

&lt;p&gt;Except you don’t pay in dollars. You pay in data. Everything you send and everything the model sends back becomes training material for future Meta models. And here’s the part that actually bugs me: Meta hasn’t said whether that data use stops at training, or whether it also covers evaluation, red-teaming, and product analytics. The scope is undefined, which in practice means you assume the worst.&lt;/p&gt;

&lt;p&gt;So the honest guidance is the boring guidance. Fine for public code, synthetic code, throwaway experiments, anything you’d have posted to a gist anyway. Absolutely not for client repos, secrets, NDA material, or unreleased product logic. This is the oldest cheapskate trap there is. The sticker price isn’t the real price. Sometimes the discount is the product and you’re the inventory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Boring Answer, Week Four
&lt;/h2&gt;

&lt;p&gt;Okay. Strip away the launch confetti and the broken promises. What’s the model that gives me the most quality per dollar right now?&lt;/p&gt;

&lt;p&gt;Same answer as last week. And the week before. MiMo v2.5 Pro. Xiaomi’s model from April, $0.43/$0.87 on Arena’s list price, &lt;a href="https://openrouter.ai/xiaomi/mimo-v2.5-pro" rel="noopener noreferrer"&gt;routing across seven providers on OpenRouter&lt;/a&gt; with no geo-lock. Listed, purchasable, cheap as dirt. This is now the fourth cycle running where it’s the cheapest thing inside the competitive band of four separate Arena categories, and it’s still sitting on &lt;a href="https://artificialanalysis.ai/" rel="noopener noreferrer"&gt;Artificial Analysis’s Intelligence-vs-Cost Pareto frontier&lt;/a&gt;. The popularity metric and the capability metric keep independently pointing at the same cheap model. That’s the strongest buy signal this newsletter ever produces, and it just refuses to change.&lt;/p&gt;

&lt;p&gt;While I’m here, one thing worth calling out about the top of the boards. Anthropic owns the leader spot in five of six Arena categories this week. But look at &lt;em&gt;which&lt;/em&gt; Anthropic model. On Instruction Following and Hard Prompts the leader is Opus 4.6-thinking. Not Opus 5. The newest, “smartest” flagship, the one that tops the hard-benchmark index, only clearly wins the Math board, and it does that on preliminary votes. So even inside one lab, “newest” and “what blind human raters actually prefer” are pointing in different directions. The word “smartest” is doing a lot of unearned work in the marketing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheapskate Picks: Where the Money Actually Is
&lt;/h2&gt;

&lt;p&gt;This is the section I write the whole thing for. The method, one more time, because it’s the only way the table makes sense. For each Arena category, take the leader’s rating, draw a line 50 points below it, and everything above that line is the competitive band. Statistically it’s a rounding error away from the “best” model. Sort that band by output price, take the cheapest thing in it. The trick I’ve screwed up before is that the band is defined by &lt;em&gt;points&lt;/em&gt;, not by rank. It runs way deeper than the visible top 20. The Overall band this week is about 53 models deep. The cheap open-weight models live down in the 30s and 40s, a couple of points behind premium brands and an order of magnitude cheaper. Truncate at rank 20 and you delete the entire reason this section exists.&lt;/p&gt;

&lt;p&gt;So I pulled the full tables and computed the bands in code. Here’s where the money is.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Leader&lt;/th&gt;
&lt;th&gt;$ leader&lt;/th&gt;
&lt;th&gt;Cheapskate pick&lt;/th&gt;
&lt;th&gt;$ pick&lt;/th&gt;
&lt;th&gt;Δ rating&lt;/th&gt;
&lt;th&gt;Price ratio&lt;/th&gt;
&lt;th&gt;AA Pareto&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Overall&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#37)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;-38&lt;/td&gt;
&lt;td&gt;~57×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#27)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;-35&lt;/td&gt;
&lt;td&gt;~57×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative Writing&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;Gemini 3 Flash (#21)&lt;/td&gt;
&lt;td&gt;$3&lt;/td&gt;
&lt;td&gt;-48&lt;/td&gt;
&lt;td&gt;~17×&lt;/td&gt;
&lt;td&gt;nearby&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction Following&lt;/td&gt;
&lt;td&gt;claude-opus-4-6-thinking&lt;/td&gt;
&lt;td&gt;$12.50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#24)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;-44&lt;/td&gt;
&lt;td&gt;~14×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hard Prompts&lt;/td&gt;
&lt;td&gt;claude-opus-4-6-thinking&lt;/td&gt;
&lt;td&gt;$12.50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#27)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;-38&lt;/td&gt;
&lt;td&gt;~14×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math&lt;/td&gt;
&lt;td&gt;claude-opus-5-max&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;Gemini 3.6 Flash (#5)&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;td&gt;-39&lt;/td&gt;
&lt;td&gt;~6.7×&lt;/td&gt;
&lt;td&gt;nearby&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few things worth saying out loud.&lt;/p&gt;

&lt;p&gt;MiMo sweeps four of six again, all at $0.87 output. Its ranks look scary until you check the vote counts. That #37 Overall rating is backed by 50,000 votes. That’s a far more settled number than most of the preliminary top-10 entries sitting on a few hundred votes each. Deep rank measures preference, not reliability. Don’t let it spook you.&lt;/p&gt;

&lt;p&gt;Creative Writing is the one place MiMo can’t reach, same as always, because that board rewards polish and the cheap crowd falls just below the cutoff. The pick there is Gemini 3 Flash at $3 output, clinging to the edge of the band 48 points back, still 17 times cheaper than the leader. Math is the compressed board this week, only about 6 models deep, and the cheapest thing in it is Gemini 3.6 Flash at $3.75. Real discount, just a smaller one, and the leader up top is preliminary anyway.&lt;/p&gt;

&lt;p&gt;And there’s a wildcard I have to mention because the method demands it. On the Overall board, a model called hy3, which is Tencent’s Hunyuan 3, sits at rank #52 with an output price of $0.53. That’s cheaper than MiMo. But it’s parked at exactly the leader-minus-50 line, with a rating margin wide enough to dip below the cutoff on any given day. So it’s the cheaper gamble, not the anchor. If you want to ride the very edge of the band to save another thirty cents, it’s there. I’m keeping MiMo as the pick because I like my recommendations boring and my votes plentiful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Horror Stories From the Wild
&lt;/h2&gt;

&lt;p&gt;Three this week, sliding from “your problem” to “everyone’s problem.”&lt;/p&gt;

&lt;p&gt;The first one I already told you: Meta’s contributor tier, where the cheapest coding option on the board is cheap precisely because it’s eating your prompts. If you saw the $0.20 output price and got excited before reading the fine print, that’s the fire. Go check what you piped through it.&lt;/p&gt;

&lt;p&gt;The second is Qwen3.8-Max’s disappearing act. Not a crash, just a broken promise with a benchmark table full of self-graded numbers behind it. The lesson is the same one this newsletter keeps preaching. Wait for someone who doesn’t work at the lab to run the model before you build anything on top of it.&lt;/p&gt;

&lt;p&gt;The third is the one that should actually scare you, because it’s not about any single model. It’s slopsquatting. Roughly &lt;a href="https://www.pearlorganisation.com/post/ai-hallucinations-in-enterprise-apps-real-costs-root-causes-and-how-to-fix-them" rel="noopener noreferrer"&gt;one in five package dependencies that AI coding assistants suggest simply don’t exist&lt;/a&gt;. Attackers know this. They watch for the common hallucinated package names and register real malware under those exact names. So your agent confidently invents an import, you run the install without blinking, and now you’ve pulled a payload onto your machine. The fix is unglamorous and non-negotiable: read your dependencies before you install them. Every time. Yes, even the ones that look obviously real.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coming Soon (Or “Soon,” Anyway)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 3.5 Pro.&lt;/strong&gt; Still vapor. It was widely expected on August 7, then &lt;a href="https://nokiapoweruser.com/gemini-3-5-pro-delayed-again-deployment-issues/" rel="noopener noreferrer"&gt;got pushed again on “deployment issues”&lt;/a&gt;. That’s roughly &lt;a href="https://tech-insider.org/au/gemini-3-5-pro-67-days-delay-2026/" rel="noopener noreferrer"&gt;67 days past Google’s “within a month” promise from I/O&lt;/a&gt;, after the team reportedly scrapped and rebuilt the base model over coding and reliability failures. Google keeps shipping Flash tiers instead. I’ll believe it when I can call the API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3.8-Max open weights.&lt;/strong&gt; Overdue as of this writing, no license, no date. If they land, the independent evals will be the story, not the launch table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Meta Muse Spark 1.2 open weights.&lt;/strong&gt; Announced August 10. A pledge, not a release, but if it ships it’s the best US open-weight model out there.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ling 3.0 Flash.&lt;/strong&gt; This one actually happened. Ant Group’s inclusionAI &lt;a href="https://www.businesswire.com/news/home/20260726584441/en/" rel="noopener noreferrer"&gt;open-sourced it August 5 under MIT&lt;/a&gt;, 124 billion parameters with 5.1 billion active, a 262K context. In a week defined by open weights that didn’t ship, a small Chinese lab quietly shipped some. Worth a look if you’re running your own hardware.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I Actually Took Away This Week
&lt;/h2&gt;

&lt;p&gt;The frontier is boring and the promises are noise.&lt;/p&gt;

&lt;p&gt;Anthropic still owns the top of nearly every board and has for over a month. The interesting action is all down in the bargain bin, where a Chinese lab dangled open weights and yanked them, an American lab pledged to open its own, and a $0.87 model from April kept quietly winning four categories while nobody talked about it. The launches change. The headlines change. The correct move does not.&lt;/p&gt;

&lt;p&gt;Same advice as last week, and probably next week. Ignore the launch you read about in the news. Open the leaderboards, find the cheapest model inside the band, confirm two different metrics agree it’s actually good, and run that. This week that’s still MiMo v2.5 Pro at 57 times less than the thing everyone’s cheering for.&lt;/p&gt;

&lt;p&gt;And read your dependencies before you install them. I mean it. The slop is coming from inside the house now.&lt;/p&gt;

&lt;p&gt;Next Tuesday, same coffee, same two tabs. I fully expect another lab to have promised me something by then. I’ll believe that one when I can download it too.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>openrouter</category>
    </item>
    <item>
      <title>Somebody Finally Wrote Down Why My Coding Agents Keep Failing the Same Way</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Mon, 10 Aug 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/somebody-finally-wrote-down-why-my-coding-agents-keep-failing-the-same-way-13o</link>
      <guid>https://dev.to/eristoddle/somebody-finally-wrote-down-why-my-coding-agents-keep-failing-the-same-way-13o</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3fe9995dmeidzmy55h9d.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3fe9995dmeidzmy55h9d.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I read &lt;a href="https://www.amazon.com/Beyond-Code-AI-Assisted-Engineering-Mechanical-ebook/dp/B0H83HR8SH" rel="noopener noreferrer"&gt;&lt;em&gt;Beyond Code&lt;/em&gt;&lt;/a&gt; as a PDF in Apple Books, which means the only reason I have any notes on it is that I got annoyed enough last month to &lt;a href="https://dev.to/eristoddle/apple-books-hides-your-pdf-highlights-my-obsidian-plugin-now-digs-them-out-lio"&gt;make my own plugin dig the highlights out of Apple’s database&lt;/a&gt;. So the first thing this book did was justify a weekend I’d already spent.&lt;/p&gt;

&lt;p&gt;The second thing it did was produce 113 highlights in the first hundred pages. With this book I was highlighting entire paragraphs because I kept hitting sentences that described something I had personally screwed up and then written a blog post about.&lt;/p&gt;

&lt;p&gt;That’s the review, really. But I’ll take the long way there. Fair warning, I’m sixty percent in. Parts I through IV.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0d342iswzlpups1ngpjc.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0d342iswzlpups1ngpjc.jpg" alt="Beyond Code by Jeremy McEntire" width="800" height="987"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What This Book Isn’t&lt;/li&gt;
&lt;li&gt;The Framing That Should Be on a Poster&lt;/li&gt;
&lt;li&gt;The Chapter That Made Me Go Look at My Own Repo&lt;/li&gt;
&lt;li&gt;What Not to Feed&lt;/li&gt;
&lt;li&gt;Mechanical Gates Over Advisory Review&lt;/li&gt;
&lt;li&gt;The Multi-Agent Thing&lt;/li&gt;
&lt;li&gt;Every Metric in My Pipeline Is Now Suspect&lt;/li&gt;
&lt;li&gt;Where It Loses Me a Little&lt;/li&gt;
&lt;li&gt;Where I Actually Am With It&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What This Book Isn’t
&lt;/h2&gt;

&lt;p&gt;The most important thing about this book is a thing that isn’t in it.&lt;/p&gt;

&lt;p&gt;There is no tool in this book. Not one. No “here’s my Claude Code setup.” No &lt;code&gt;.cursorrules&lt;/code&gt; walkthrough. No comparison table of Copilot versus Opencode versus whatever shipped last Tuesday and will be deprecated by the time you finish the chapter. No recommended library, no framework, no repo to clone. The author does not tell you which model to use. He does not tell you which agent harness to use. He never once tells you what to install.&lt;/p&gt;

&lt;p&gt;I’ve read a lot of books and many blog posts about AI-assisted development at this point, mine included, and nearly all of them sit at one of two extremes. Either they’re deeply technical and showing you how to write the code, or they’re written for stakeholders and exist to convince a CEO that AI deserves a budget line. Both have a shelf life of about six months, because the thing they’re really teaching you is a menu, and the menu changes. This book sits between the two, which happens to be the only place where you come away with a framework in your head instead of a list of settings.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Beyond Code&lt;/em&gt; is teaching you the physics. Every failure mode it describes is derived from a property of how attention works, or how information behaves when you compress it, or how optimization pressure behaves when you point it at a proxy. None of those change when a new model ships. McEntire says this outright, and it’s the line that made me trust the rest of the book:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The engineer who understands information loss at coordination boundaries, proxy optimization in evaluation systems, and the structural impossibility of quality-through-review-gates does not need to memorize a list of anti-patterns. They can derive the anti-patterns from the physics, and they can design architectures that avoid them by construction rather than by vigilance.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s a hell of a claim to open with. He mostly delivers on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Framing That Should Be on a Poster
&lt;/h2&gt;

&lt;p&gt;The book opens on an asymmetry I have been failing to articulate since I started doing this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;the cost of building software dropped by roughly an order of magnitude, but the cost of building the wrong software did not drop at all. When code was expensive to produce, the expense itself served as a natural forcing function for thought.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Everything else falls out of that. When typing was slow, the slowness did your thinking for you. It made you consider whether the feature was worth it, because you were the one who had to sit there and type it. That forcing function is gone and nothing replaced it.&lt;/p&gt;

&lt;p&gt;There’s a carpenter analogy early on that I’d normally roll my eyes at, because tech books love a trade analogy, but this one earns it. Give a nail gun to someone who understands load paths and soil conditions and you get a faster carpenter. Give it to someone who’s never framed a wall and you get something that looks like a house until the first heavy snow. The line that stuck:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;It made them a fast amateur, and a fast amateur with a power tool is more dangerous than a slow one, because the volume of confident mistakes exceeds anyone’s ability to catch them before the roof goes on.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I have been the fast amateur. I have &lt;a href="https://dev.to/eristoddle/the-great-vibe-coding-experiment-how-i-built-15-projects-with-ai-in-my-spare-time-275o"&gt;fifteen projects&lt;/a&gt; proving it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Chapter That Made Me Go Look at My Own Repo
&lt;/h2&gt;

&lt;p&gt;A guideline says “functions should be short.” A constraint says “functions exceeding fifty lines fail the linter.” The guideline needs someone to agree, to notice, and to care. The constraint just fails the build. One admits interpretation. The other doesn’t. I’d been operating on that distinction by feel and had never had the words for it.&lt;/p&gt;

&lt;p&gt;Then he describes constraints as a topology. Rigid exterior, flexible interior. The boundary is enforced mechanically and non-negotiably, and inside the boundary the agent gets total freedom on naming, structure, algorithm choice, error handling, all of it.&lt;/p&gt;

&lt;p&gt;And then he names the failure on the other end, which is the one nobody warns you about:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An over-constrained system has specified the requirements so precisely that no implementation can satisfy all of them simultaneously without consuming the entire budget on compliance. The remedy in both cases is the same: constrain what matters, leave flexible what does not, and know the difference.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I stopped reading and went and looked at a repo.&lt;/p&gt;

&lt;p&gt;Because I have written this exact arc as three separate blog posts without once connecting them. &lt;a href="https://dev.to/eristoddle/the-great-vibe-coding-experiment-how-i-built-15-projects-with-ai-in-my-spare-time-275o"&gt;I vibe coded a project into a corner&lt;/a&gt; and got a sprawling codebase that did a third of what I wanted in a way that made the other two thirds impossible. That’s under-constraint. Then I &lt;a href="https://dev.to/eristoddle/my-third-try-how-a-living-plan-beat-both-vibe-coding-and-spec-kit-5a89"&gt;tried spec-kit on the same idea&lt;/a&gt; and generated a beautiful tree of documents describing something I was still guessing about, and I was exhausted with the project before I wrote a line of code. That is over-constraint, precisely, and I’d been telling people it was a spec-kit problem. It wasn’t. It’s a topology problem. Spec-kit just made it easy to fall into.&lt;/p&gt;

&lt;p&gt;The thing that eventually worked, a single &lt;code&gt;PLAN.md&lt;/code&gt; with numbered decisions and everything else left loose, is a constraint topology. I built one by accident and then wrote a post about how the trick was “dumber than I want to admit.” Turns out the trick has a name and about eight pages of derivation behind it.&lt;/p&gt;

&lt;p&gt;That’s the experience this book keeps producing. Not “here’s a new technique.” More like watching someone explain the mechanism behind a thing you’d already stumbled into, which is somehow both validating and humiliating.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Not to Feed
&lt;/h2&gt;

&lt;p&gt;One chapter in here is worth the price of the book on its own.&lt;/p&gt;

&lt;p&gt;The core argument is arithmetic. Attention is a fixed mass distributed across candidates. Ten thousand tokens of context, each token gets a share. A hundred thousand tokens, each token gets a thousandth of that share. The degradation isn’t a bug someone will patch. It’s how the mechanism works.&lt;/p&gt;

&lt;p&gt;Then he lands the part that genuinely reordered something in my head:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This mechanism explains why well-written documentation is more dangerous than random text. Random text has low semantic coherence with the task-relevant information in the context. The model’s attention mechanism can distinguish between meaningfully related tokens and gibberish, and it largely ignores the gibberish. Coherent documentation on a related topic, by contrast, is densely packed with tokens that are semantically proximate to the task-relevant tokens.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Gibberish in your context is harmless. The model routes around it. Well written, on-topic, genuinely useful documentation that happens to be irrelevant to &lt;em&gt;this specific task&lt;/em&gt; is the poison, because it looks identical to the relevant material in embedding space. The model cannot tell “related and relevant” from “related and irrelevant.”&lt;/p&gt;

&lt;p&gt;He points this straight at RAG, noting that retrieval by semantic similarity is retrieval by exactly the metric that predicts maximum distraction. I don’t think that’s a full takedown of RAG and he doesn’t claim it is, but it’s the most uncomfortable sentence about RAG I’ve read.&lt;/p&gt;

&lt;p&gt;For me it explained a thing I’d already fixed without understanding. My &lt;code&gt;PLAN.md&lt;/code&gt; grew to 28,000 words and became &lt;a href="https://dev.to/eristoddle/the-living-plan-got-fat-compacting-a-doc-that-wont-stop-growing-3nk0"&gt;the document I dreaded opening&lt;/a&gt;, so I built a skill to keep it lean and file the cooled material into linked docs. I thought I was solving a token cost problem and a me-being-bored problem. I was actually solving an attention dilution problem, and every word in that doc was well written and on topic, which per this book is what made it worse.&lt;/p&gt;

&lt;p&gt;The rule he gives is one line and it’s the whole discipline: the quality of context is determined as much by what you exclude as by what you include.&lt;/p&gt;

&lt;p&gt;There’s also a small tactical thing in here that I’ve already changed my behavior over. Negative instructions backfire mechanically. To process “do not use eval,” the model has to represent &lt;code&gt;eval&lt;/code&gt;, which activates the attention patterns around &lt;code&gt;eval&lt;/code&gt;, which raises the odds &lt;code&gt;eval&lt;/code&gt; shows up in the output. Then it has to suppress that, and suppression fails under load more often than activation does. So the forbidden thing appears &lt;em&gt;because&lt;/em&gt; you forbade it. Positive instructions specify the correct point in the solution space. Negative ones exclude one wrong point and leave the rest wide open.&lt;/p&gt;

&lt;p&gt;My agent instruction files are full of “never do X.” I have some editing to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mechanical Gates Over Advisory Review
&lt;/h2&gt;

&lt;p&gt;This is the section I expect people to argue with, and it’s the one I most agree with.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Advisory review asks whether someone did the work. Mechanical verification asks whether the work succeeded.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The argument is that code review, as practiced almost everywhere, is advisory. It renders an opinion. And any process that asks “does this look right?” instead of “does this pass?” is asking a question with no objective answer, so the answers get governed by cognitive bias and social dynamics instead of by properties of the artifact. Swapping human reviewers for AI reviewers doesn’t fix it, because the humans were never the problem. The gate was.&lt;/p&gt;

&lt;p&gt;I’ve been building around this without naming it. Every one of my writing pipeline skills that actually holds up is a mechanical gate. A link auditor that resolves every URL. A fact checker that produces a pass or fail report. A brief checker that counts headings against the spec. The skill descriptions I wrote for myself say things like “generation is unreliable, this is the guarantee.” I wrote that before I read this book, and reading it felt like getting graded.&lt;/p&gt;

&lt;p&gt;The chapter goes further into contract-driven development, where the contracts are the product and the implementation is disposable. Production incident happens, you don’t patch the implementation, you write a reproducer test and regenerate the implementation against the stricter contract. “The implementation is cattle, not pets.” He handles the obvious objection honestly, too, which I appreciated: contracts only test what you thought to specify, and they don’t test what you failed to anticipate. His answer is convergence over time rather than a claim of completeness.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Multi-Agent Thing
&lt;/h2&gt;

&lt;p&gt;This is the part that stung.&lt;/p&gt;

&lt;p&gt;He ran experiments across pipeline, hierarchical, and swarm agent architectures, and the finding is blunt: a single agent with full context beat every multi-agent configuration, because it suffered no compression loss. Every gate in a pipeline squeezes a rich code artifact down into a low-dimensional verdict, and every stage after that operates on the verdict instead of the code. The information is gone and no amount of additional review recovers it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For any task that fits within a single agent’s effective context window, hierarchical distribution is a net negative: it introduces compression losses, strategic distortions, and coordination overhead without providing any compensating benefit.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I want to be honest about my reaction here, which is that I did not want this to be true. I like the fan-out. Firing off a bunch of agents in parallel feels like leverage. It also once killed an entire session for me, because twenty-two background jobs finished at various times and each one dumped its full output back into the same context, and the session ate its own limit in about a minute. I turned that into a hard rule in my global agent config afterward: cap concurrency, jobs write to disk, they return a path and a one line status. Conclusions come back, not payloads.&lt;/p&gt;

&lt;p&gt;That rule was pure scar tissue. This book explains it as coordination bandwidth. Distribute only what actually needs distributing, decompose at natural information boundaries where coupling is low enough that the interface fits in a sentence or two, and coordinate through the shared environment rather than through chatter between agents. That last one he calls stigmergic coordination, and the best line in the chapter is about why it beats agents describing things to each other:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A description of a bug is a lossy compression of the bug. The test failure is the bug, visible in the shared environment without anyone’s having described it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s what a task board and a test suite have been doing in my projects the whole time. I just thought of them as project management.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every Metric in My Pipeline Is Now Suspect
&lt;/h2&gt;

&lt;p&gt;I came off the Goodhart chapter right before sitting down to write this, so it’s the freshest one in my head. The argument isn’t the usual “metrics bad” hand-wave. It’s that the search for the correct metric is itself the mistake, because the divergence between proxy and objective is structural rather than a matter of picking better. Test coverage measures execution and execution is not verification, and no amount of refining coverage as a metric will make those the same thing. His agents weren’t gaming the system out of career anxiety. They optimized what they could see at the expense of what they couldn’t, which is just what optimizers do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where It Loses Me a Little
&lt;/h2&gt;

&lt;p&gt;Two honest gripes.&lt;/p&gt;

&lt;p&gt;The prose is dense. This is not a book you skim on a phone. Some sentences run long enough that I had to take a second pass, and the register stays clinical even when the material is dramatic. I don’t mind it, but if you want the breezy conversational thing, this is not that. It reads more like a good long paper than like a blog. But each chapter is broken up into sections that are rarely over two pages long, so you can read it in bite-sized pieces.&lt;/p&gt;

&lt;p&gt;The other is that being tool-agnostic has a cost. The book will tell you to encode a constraint mechanically and it will not tell you what that looks like in your stack. That’s deliberate and it’s why the book won’t rot, but there were a few points where I wanted one concrete example and got a principle instead. You are expected to do the translation yourself. If you want a book that hands you a config file, this is not the book for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I Actually Am With It
&lt;/h2&gt;

&lt;p&gt;The technical argument is the part I came for and it’s done: the shift, context, constraints, coordination. What’s left is Part V on the craft and Part VI on where engineering goes, which look more like the career and human chapters. I’m going to finish it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.amazon.com/Beyond-Code-AI-Assisted-Engineering-Mechanical-ebook/dp/B0H83HR8SH" rel="noopener noreferrer"&gt;&lt;em&gt;Beyond Code: Context, Constraints, and the New Craft of Software&lt;/em&gt; by Jeremy McEntire&lt;/a&gt;. Five stars on the two thirds I’ve read, and I’ll revisit that when I finish the rest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who should read this:&lt;/strong&gt; anyone who has been using coding agents seriously for more than a few months and has started noticing that the failures repeat. If you’re finding that your agents fail the same way across different tools and different models, this book explains why, and the explanation will survive the next model release. It’s also, I think, genuinely good for the “AI is useless” and “AI replaces engineers” crowds, both of whom are answered here by the same argument: the cost of producing code collapsed, the cost of knowing what to build did not, and both camps are assuming the first one implies the second.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who shouldn’t:&lt;/strong&gt; anyone looking for a setup guide. Go read a blog post. Possibly one of mine.&lt;/p&gt;

&lt;p&gt;I read a lot of stuff in this space and most of it is somebody’s workflow with the serial numbers filed off. This is the first one where I finished a chapter and went to go change something in a repo. Then did it again four chapters later. That’s a low bar in theory and almost nothing clears it.&lt;/p&gt;

</description>
      <category>beyondcodebook</category>
      <category>aiassisteddevelopmen</category>
      <category>codingagentsfailure</category>
      <category>jeremymcentire</category>
    </item>
    <item>
      <title>The Expensive Model Only Plans Now: I Split My AI Coding Rig Across Two Tools</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Wed, 05 Aug 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/the-expensive-model-only-plans-now-i-split-my-ai-coding-rig-across-two-tools-4dch</link>
      <guid>https://dev.to/eristoddle/the-expensive-model-only-plans-now-i-split-my-ai-coding-rig-across-two-tools-4dch</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fidgwj1vhiyxkken12qlf.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fidgwj1vhiyxkken12qlf.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A couple of posts back I pulled a month of my own session logs to catch coding red-handed as the token hog, and &lt;a href="https://dev.to/eristoddle/planning-is-cheaper-than-coding-and-my-own-logs-proved-me-wrong-about-why-51bh"&gt;the meter came back flat&lt;/a&gt;. Planning and building cost about the same to run. The expensive part is the thinking that has to happen &lt;em&gt;before&lt;/em&gt; the typing, and that thinking is mine.&lt;/p&gt;

&lt;p&gt;If the typing is the cheap, fast, mechanical part, why am I paying frontier-model prices for it? Why is Opus, the model I keep around because it can &lt;em&gt;think&lt;/em&gt;, the one grinding out a config file from a spec I already wrote?&lt;/p&gt;

&lt;p&gt;It shouldn’t be. I should keep the expensive model for the one thing it should, and hand all the work after that to &lt;a href="https://dev.to/eristoddle/the-cheapskates-guide-to-the-arena-leaderboard-why-i-stopped-paying-claude-opus-prices-1ipn"&gt;the cheapest models that can write code&lt;/a&gt;. Which, it turns out, means splitting the work across two different tools. Claude Code plans, &lt;a href="https://opencode.ai" rel="noopener noreferrer"&gt;Opencode&lt;/a&gt; builds.&lt;/p&gt;

&lt;p&gt;This is part four of a series about working with AI coding agents on an open-ended project through a single living &lt;code&gt;PLAN.md&lt;/code&gt; instead of vibe coding or spec-kit ceremony. &lt;a href="https://dev.to/eristoddle/my-third-try-how-a-living-plan-beat-both-vibe-coding-and-spec-kit-5a89"&gt;Post one&lt;/a&gt; is the thesis: the doc is the deliverable, the code is the byproduct. The &lt;a href="https://dev.to/eristoddle/the-living-plan-got-fat-compacting-a-doc-that-wont-stop-growing-3nk0"&gt;last few&lt;/a&gt; were about keeping that doc lean and figuring out where the real cost lives. This one is about what happens when you stop paying the planner to type.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Setup: One Expensive Model Doing Two Jobs&lt;/li&gt;
&lt;li&gt;Enter opencode (and a Roster of Cheap Models)&lt;/li&gt;
&lt;li&gt;The Rule I’d Just Landed On, and Why I Inverted It&lt;/li&gt;
&lt;li&gt;Reviewers That Don’t Share the Builder’s Blind Spots&lt;/li&gt;
&lt;li&gt;The Tell: The Diagrams Came Before the Code&lt;/li&gt;
&lt;li&gt;Three Docs, Three Sets of Write Permissions&lt;/li&gt;
&lt;li&gt;The Whole Truth&lt;/li&gt;
&lt;li&gt;What’s Next&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Setup: One Expensive Model Doing Two Jobs
&lt;/h2&gt;

&lt;p&gt;For most of this series my setup has been one tool wearing two hats. I plan in Claude Code with Opus. Plain chat, back and forth, drilling down until the design is actually pinned. Then the &lt;em&gt;same&lt;/em&gt; Claude Code hands the build work to Sonnet sub-agents. Opus knows, Sonnet does. On the Pro plan I could talk to Opus for a long planning session and barely dent my usage, because in build mode the token-heavy churn (the greps, the file reads, the edits) happens in a sub-agent’s own context and comes back as a short summary. The orchestrator stays lean.&lt;/p&gt;

&lt;p&gt;That’s a fine setup. But it has Opus, or at least a Claude model, standing over &lt;em&gt;every&lt;/em&gt; build task, and I’d just proven to myself that the build tasks are the mechanical part. I was using an expensive model to assemble flat-pack furniture. The instructions were already written. Any model that can read can do the work.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe7oshmzzf6jicjxc4j2l.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe7oshmzzf6jicjxc4j2l.jpg" alt="The Setup: One Expensive Model Doing Two Jobs" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here’s the framing that made me actually move, and it’s stolen from a thing I said to myself during a debugging session months ago: &lt;strong&gt;models are the Temu of things.&lt;/strong&gt; They’ll have knowledge of many things, but the most generalized version of it. A one-line prompt gets you the average of everything ever written on the subject. And if the build task is &lt;em&gt;fully specified&lt;/em&gt;, if I did the custom thinking already and wrote it down, then Temu is exactly what I want. Generalized competence at assembling a known thing is cheap and abundant. I just had to stop buying it at boutique prices.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enter opencode (and a Roster of Cheap Models)
&lt;/h2&gt;

&lt;p&gt;The tool that let me do this cleanly is &lt;a href="https://opencode.ai" rel="noopener noreferrer"&gt;Opencode&lt;/a&gt;, a terminal coding agent that, crucially, lets you assign a &lt;em&gt;different model to every agent&lt;/em&gt; through a repo-local &lt;code&gt;opencode.json&lt;/code&gt;, routed through OpenRouter. So the plan stays in Claude Code with Opus, and the entire build crew lives in opencode, each role on whatever model is cheapest for its job.&lt;/p&gt;

&lt;p&gt;Here’s the roster, trimmed from the real config:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://opencode.ai/config.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"small_model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/z-ai/glm-4.7-flash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"default_agent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"orchestrate"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"agent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"orchestrate"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/z-ai/glm-5.2"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"build"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"subagent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/xiaomi/mimo-v2.5-pro"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ai-pipeline"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"subagent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/deepseek/deepseek-v4-pro"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"qa-fast"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"subagent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/deepseek/deepseek-v4-flash"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"qa-deep"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"subagent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/google/gemini-3.1-pro-preview"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"debug"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"subagent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/anthropic/claude-sonnet-5"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"docs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"subagent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/qwen/qwen3.7-plus"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"research"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"subagent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/deepseek/deepseek-v4-flash"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"architect"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openrouter/anthropic/claude-opus-5"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not a Claude model in the build path. GLM-5.2 orchestrates. A Xiaomi MiMo model writes the Go daemon and the CLI. DeepSeek handles the ML and retrieval code. Qwen writes the docs. The only Anthropic models in the whole file are &lt;code&gt;debug&lt;/code&gt; (Sonnet, for when something’s genuinely stuck) and &lt;code&gt;architect&lt;/code&gt;. And to tell you the truth, most of the model choices are just a test of what will work in that slot or even if I need all these different agents.&lt;/p&gt;

&lt;p&gt;I’ve been coy about &lt;em&gt;what&lt;/em&gt; this crew is building, so: it’s &lt;a href="https://dev.to/eristoddle/i-got-tired-of-ai-memory-hype-so-i-built-a-context-lake-55fi"&gt;the Context Lake&lt;/a&gt;, the one-brain-outside-the-agents thing I wrote about a month back. A Go daemon that watches my session logs and my vault, a Python ML layer that indexes and retrieves, a CLI and a dashboard sitting on top. That matters here for exactly one reason: it is not a toy. It’s a real multi-language system with a knowledge graph, a vector store, and a daemon.&lt;/p&gt;

&lt;p&gt;MiMo wasn’t even my first pick for the build. I started with GLM-5.2 doing double duty, orchestrating &lt;em&gt;and&lt;/em&gt; building, and swapped the builder over to MiMo on day three (the commit message, optimistically, reads “for improved performance”). It’s cheaper on the output tokens a builder spends most of, though I’ll warn you the exact gap is a moving target: model pricing on OpenRouter whiplashes week to week, and as I write this GLM’s page is showing 76% off. The reason I &lt;em&gt;kept&lt;/em&gt; MiMo, though, wasn’t the couple of cents. It was that GLM tends to lose the thread across a long series of tasks, and MiMo just… doesn’t. Stack it deep and it grinds through the whole queue. For a builder whose entire job is to chew an unattended stack while I’m somewhere else, that turned out to matter more than the price.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fydf8fvfbhrcnvh1a4wyc.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fydf8fvfbhrcnvh1a4wyc.jpg" alt="Enter opencode (and a Roster of Cheap Models)" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The division of labor is the same knower/doer split I’ve run all along, just stretched across a tool boundary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code + Opus is the planner.&lt;/strong&gt; It’s where I think, discuss, and write the tasks. It never runs in opencode. The picks were made against live OpenRouter pricing using &lt;a href="https://dev.to/eristoddle/building-a-cost-saving-agent-skill-that-accidentally-became-its-own-weekly-blog-post-3o1h"&gt;my weekly model-buzz research&lt;/a&gt; as the value spine, so this isn’t “cheap for the sake of cheap.” It’s cheap where cheap is fine, and one expensive model where thinking has to happen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Opencode is the whole build crew.&lt;/strong&gt; It reads the tasks I wrote, does the keystrokes, runs the checks, reports back.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And &lt;code&gt;architect&lt;/code&gt;, the lone Opus-4.8 agent in opencode? It’s set to manual-only and &lt;strong&gt;never&lt;/strong&gt; auto-invoked. It’s my break-glass “call in a senior for a from-orbit sanity check” button for when a Claude Code session isn’t handy. A zero-invocation count on it is &lt;em&gt;expected&lt;/em&gt;. I wrote a note in the repo telling future-me not to delete it as dead config, because future-me absolutely would.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Rule I’d Just Landed On, and Why I Inverted It
&lt;/h2&gt;

&lt;p&gt;Here’s where it got uncomfortable, because I had to break something I’d just decided was right.&lt;/p&gt;

&lt;p&gt;In my single-tool setup, &lt;a href="https://dev.to/eristoddle/the-bottleneck-was-me-how-i-stopped-racing-my-ai-builder-and-started-pacing-it-3e1n"&gt;the builder works a queue of numbered tasks&lt;/a&gt;, and the rule for a task it can’t finish was: &lt;strong&gt;skip it and keep going.&lt;/strong&gt; If piece 3 of 8 turns out underspecified, some fork I didn’t see, don’t halt the whole queue. Mark it blocked, move to piece 4, do everything that doesn’t depend on the broken one. An hour away should come back with six of eight done, not two. I was proud of that rule. It’s the thing that makes walking away safe.&lt;/p&gt;

&lt;p&gt;Then I moved the build to opencode and inverted it completely. The rule now: &lt;strong&gt;hit something underspecified, stop and report. Do not keep going. Do not invent the next task. Do not promote anything from the parking lot.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Same person, opposite rule, two weeks apart. What changed?&lt;/p&gt;

&lt;p&gt;The tool boundary changed. In the single-tool world, the thing that skips a blocked task and continues is &lt;em&gt;the same context that could re-plan it&lt;/em&gt;. It’s all one agent, warm on the whole conversation, and “keep going past the blocker” is safe because the planner is right there. In the cross-tool world, &lt;strong&gt;the planner is a different tool.&lt;/strong&gt; Opencode is a cheap executor that knows nothing about my planning thread and has no business deciding what to build next. Discovering work and scheduling work are different authorities, and I’d just handed them to different tools. If I let the cheap builder skip-and-continue, I’m letting the model I specifically chose &lt;em&gt;not&lt;/em&gt; to think make the scheduling calls.&lt;/p&gt;

&lt;p&gt;So the executor’s rule became rigid on purpose. From the actual contract in the repo:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1f75jqh8szjhk648fjji.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1f75jqh8szjhk648fjji.jpg" alt="The Rule I'd Just Landed On, and Why I Inverted It" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You may NOT execute items in § Next (underspecified by design — stop at the tier boundary and report), may NOT promote items out of § Parking lot, and may NOT invent new tasks. If § Now is drained, SAY SO AND STOP.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Take the planner out of the room and put it in another tool, and the safe move flips to “stop and wait.” A rule is only as good as the context it assumes, and the second you move the work across a boundary, re-check every assumption.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reviewers That Don’t Share the Builder’s Blind Spots
&lt;/h2&gt;

&lt;p&gt;If cheap models are doing the building, the obvious worry is quality. My answer has two parts, and neither is “trust the cheap model.”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part one: the reviewers run on a different bloodline than the builder.&lt;/strong&gt; The &lt;code&gt;build&lt;/code&gt; agent is MiMo; the deep reviewer, &lt;code&gt;qa-deep&lt;/code&gt;, is Gemini; the debugger is Sonnet. A reviewer that shares the implementer’s model lineage shares its blind spots. It’ll wave through the same class of mistake the builder was prone to make, because it thinks the same way. Point review and implementation at the same model and review quietly stops catching anything. Different lineage, different blind spots.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Part two, and this is the one I got wrong first: the obligation to run the checks lives on the agent that can’t forget it.&lt;/strong&gt; Originally I had the planner enforce QA: “call the reviewer after every step.” It decayed. And I have the receipts, because this is the kind of thing I’d rather catch myself than have a commenter catch. I audited 58 opencode sessions on this project and found QA delegation went from &lt;strong&gt;5 calls in the first two weeks to 0 in the last two&lt;/strong&gt; , while the builder kept right on running. The deep reviewer, whose trigger was the soft phrase “at milestones,” had fired &lt;strong&gt;exactly zero times.&lt;/strong&gt; The invariants I care about were never actually being checked.&lt;/p&gt;

&lt;p&gt;Why did it decay? Because a planner’s context fills up over a long orchestration and it drops the soft, optional steps, the same way you stop doing the stretches your physical therapist gave you. The fix wasn’t a sterner reminder. It was structural: &lt;strong&gt;move the obligation onto the sub-agents, which get fresh context on every spawn.&lt;/strong&gt; The &lt;code&gt;build&lt;/code&gt; and &lt;code&gt;ai-pipeline&lt;/code&gt; agents are now not-done until they’ve run the checks and pasted the output into their report, and the planner rejects any completion report that lacks it. An obligation on a long-lived context erodes. An obligation on a fresh-every-spawn context can’t. Don’t ask a model to remember. Make it structural.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmid6ie2cjge858qyubza.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmid6ie2cjge858qyubza.jpg" alt="Reviewers That Don't Share the Builder's Blind Spots" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The deep reviewer got the same treatment. Instead of “at milestones,” it fires on a deterministic condition. When the current batch of work is drained, before the batch is reported done, it reviews the whole batch diff against a fixed list of the project’s load-bearing rules. A trigger you can’t measure is a trigger that never fires.&lt;/p&gt;

&lt;p&gt;I did have to fix the ordering the hard way first, because my initial config &lt;em&gt;looked&lt;/em&gt; right and wasn’t. Agents weren’t firing in sequence, QA was skippable, and a couple of agents never ran at all. So I rewrote it to force the line: &lt;code&gt;build&lt;/code&gt; writes the code and runs its own tests, and only after the checks pass does the &lt;code&gt;docs&lt;/code&gt; agent, which always reads the real code first, write anything. When I later asked Claude whether that fix actually held, the read was blunt: &lt;strong&gt;“Ordering is right. The QA-gate fix is genuinely first.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And then, a couple of weeks later, the whole thing handed me a lesson I didn’t order. At the batch drain the gate was green across the board: 43 Python tests passed, ruff and mypy clean, &lt;code&gt;go build&lt;/code&gt; and &lt;code&gt;go test&lt;/code&gt; fine. Then I ran the thing against real data. The dashboard crashed on an actual projection file, and the project-identity code reported &lt;strong&gt;6 projects where the vault holds 24&lt;/strong&gt; , four of the six being phantoms. Its last-resort rule for naming a project was “first tag containing a hyphen,” so it had been happily minting a project per topic tag. Both bugs sailed straight through a green gate. Making the checks structural fixed &lt;em&gt;whether they run&lt;/em&gt;. It did not turn them into a proxy for the software being correct, and I had started quietly treating it as one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Tell: The Diagrams Came Before the Code
&lt;/h2&gt;

&lt;p&gt;Here’s the small, dumb detail that actually convinced me the split works.&lt;/p&gt;

&lt;p&gt;Back at the start, on June 26th, before a line of the Go daemon or the Python ML code existed, I had Claude draw the architecture as a set of mermaid diagrams. Not because the code was complex; there wasn’t any. Because the &lt;em&gt;plan&lt;/em&gt; was: “is there any way that we can create a rough mermaid diagram of what’s going on here? Cause it’s getting kind of complex now.” The diagrams were a picture of the intended system, committed a few hours before the first &lt;code&gt;daemon/&lt;/code&gt; and &lt;code&gt;ml/&lt;/code&gt; directories landed that same evening.&lt;/p&gt;

&lt;p&gt;Then the roster of cheap models built the thing, over weeks, one specified task at a time.&lt;/p&gt;

&lt;p&gt;At some point I had Claude audit the built system, half-expecting the usual rot, docs that lie because nobody updated them. Instead: &lt;strong&gt;“OpenCode has kept the architecture honest.”&lt;/strong&gt; The diagrams hadn’t been quietly patched to match the code. The code had been built to match the diagrams. Weeks of cheap-model keystrokes, and the shape at the end was the shape I’d drawn on day one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo2saunod9ko7xtaas72u.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo2saunod9ko7xtaas72u.jpg" alt="The Tell: The Diagrams Came Before the Code" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That’s the real proof, and it’s humbler than “the cheap models are geniuses.” They’re not. But a plan they can’t skip, run in an order they can’t reorder, gets followed faithfully enough that the map you sketched before any code existed still describes the territory a month later. That’s the whole bet: the intelligence goes into the plan, and the plan is cheap to enforce.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Docs, Three Sets of Write Permissions
&lt;/h2&gt;

&lt;p&gt;The last piece is the guardrail that keeps a cheap, eager builder from wandering off. When two tools share one repo, “who’s allowed to write what” stops being a style preference and becomes the fence that keeps them from fighting. Three docs, three permissions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;| Doc | What it is | Who writes |
|------------------|-------------------------------------|-----------------------|
| docs/vision.md | The vision, architecture, invariants| Planner only |
| TASKS.md | The current batch of committed work | Executor (checkboxes) |
| docs/archive/ | Completed batches, frozen | Planner, at replan |

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The executor, opencode, can check off boxes in &lt;code&gt;TASKS.md&lt;/code&gt; § Now as work lands, keep the status line current, and &lt;strong&gt;append&lt;/strong&gt; discoveries to a parking lot. That’s it. It cannot edit the vision doc, cannot promote a parked item into the active work, cannot write the next batch. Those are planning authorities, and planning happens in the other tool, driven by me.&lt;/p&gt;

&lt;p&gt;This sounds like bureaucracy until you’ve watched a helpful agent decide, unprompted, that it knows what you want to build next. The most load-bearing sentence in the whole config is the one that tells the opencode orchestrator, in so many words, to &lt;strong&gt;never tell the human to “exit plan mode” or hand the work back&lt;/strong&gt;. Its job is to act by delegating to a sub-agent, not to bounce the plan back to me with a “ready when you are.” And the sibling rule, which longtime readers know I’ve promoted to a standing law: planning is a &lt;em&gt;discussion&lt;/em&gt;, not a multiple-choice quiz. The moment a tool tries to compress the design conversation into three options, you lose the part that was doing the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Whole Truth
&lt;/h2&gt;

&lt;p&gt;What I’m solid on: the &lt;em&gt;shape&lt;/em&gt; is right. Paying a thinker to type is a real waste, the cost data backs it up, and separating “who plans” from “who types” onto tools priced for each job is a clean way to stop doing it. The QA-decay finding is measured, not vibed: 58 sessions, 5 to 0, a reviewer that never fired. And putting the obligation on fresh context instead of a memory is a principle I’ll stand behind anywhere.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe368i64la5ho42s2x7s4.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe368i64la5ho42s2x7s4.jpg" alt="The Honest Scope Note" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What surprised me is how &lt;em&gt;smooth&lt;/em&gt; the unattended part has been. I can stack five tasks, walk away, and come back to a screen full of done. Actually unattended, actually finished, and I keep having Opus re-audit the whole thing because I half don’t believe it.&lt;/p&gt;

&lt;p&gt;What I was wrong about is the finish line. I wrote most of this post feeling like the project was nearly done. It isn’t. The batch I cut on July 18th sits at 11 of 15 remaining, it got &lt;em&gt;extended&lt;/em&gt; three days later after another dogfooding pass, and the scope has since grown a cross-repo orchestration layer and a skills catalog. The cheap crew is fast at the work I hand it. That was never the same thing as the work running out.&lt;/p&gt;

&lt;p&gt;What I’m still &lt;em&gt;not&lt;/em&gt; solid on: whether the cheap builders stay good enough, at scale, over time. The two bugs above are the first real data point, and they’re ambiguous on purpose. Neither is obviously a “a Claude builder would have caught this” failure, because both were spec gaps that a green test suite couldn’t see either. But they’re exactly the shape of thing I said I was watching for, they showed up inside a month, and one project going smoothly is still a sample size of one. The &lt;code&gt;qa-deep&lt;/code&gt; reviewer is on a &lt;em&gt;preview-tier&lt;/em&gt; model, which means it can get rate-limited, change behavior, or vanish out from under me on a random Tuesday; I keep a same-obligation, different-lineage fallback noted for exactly that. And the roster itself drifts. Model pricing moves, slugs get deprecated, and the config is the truth while any table I write about it is already going stale. The table in my own docs has lied to me before.&lt;/p&gt;

&lt;p&gt;So: promising shape, real receipts on the &lt;em&gt;why&lt;/em&gt;, genuinely unproven on the “are cheap models good enough” question that the whole thing rides on. Frontier of the method, not a settled result.&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s Next
&lt;/h2&gt;

&lt;p&gt;The plan lives in Claude Code with Opus, the build lives in Opencode with a crew of cheap models, the expensive model only thinks, and the whole thing is fenced by who’s allowed to write which file. It runs. Whether it &lt;em&gt;holds&lt;/em&gt; is the thing I actually have to live with now instead of theorize about. Can Temu models carry a real build over months without quietly costing me more in bugs than they saved me in tokens?&lt;/p&gt;

&lt;p&gt;There’s a bigger post hiding under this one, the thesis version: that planning is the only thing in the whole stack worth a boutique model, and everything downstream is keystrokes you can buy in bulk. I’ve now watched that hold across three projects, which is finally enough proof points to write it honestly instead of as a hot take. That’s the one I want to write eventually.&lt;/p&gt;

</description>
      <category>claudeopus</category>
      <category>opencodeai</category>
      <category>llmcostoptimization</category>
      <category>aicodingagents</category>
    </item>
    <item>
      <title>Qwen3.8-Max Dropped. I'm Still Running the $0.87 Model.</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Tue, 04 Aug 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/qwen38-max-dropped-im-still-running-the-087-model-27ep</link>
      <guid>https://dev.to/eristoddle/qwen38-max-dropped-im-still-running-the-087-model-27ep</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwgn1imrqggs26wiznwzr.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwgn1imrqggs26wiznwzr.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This week Alibaba dropped Qwen3.8-Max, a 2.4-trillion-parameter monster, and the stock jumped 7% in Hong Kong. Everybody lost their minds. And the smartest thing a cheapskate can do with that news is nod politely and keep running a Chinese model from April that costs 57 times less than the thing everyone’s actually cheering for.&lt;/p&gt;

&lt;p&gt;Let me explain how I got there.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;China Dropped a 2.4-Trillion-Parameter Flagship and I Mostly Yawned&lt;/li&gt;
&lt;li&gt;The “Smartest Model in the World” Had an Awkward Debut&lt;/li&gt;
&lt;li&gt;The Boring Answer That Keeps Winning&lt;/li&gt;
&lt;li&gt;Cheapskate Picks: Where the Actual Money Is&lt;/li&gt;
&lt;li&gt;Horror Stories From the Wild&lt;/li&gt;
&lt;li&gt;Coming Soon (Or “Soon,” Anyway)&lt;/li&gt;
&lt;li&gt;What I Actually Took Away This Week&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  China Dropped a 2.4-Trillion-Parameter Flagship and I Mostly Yawned
&lt;/h2&gt;

&lt;p&gt;Here’s the thing about Qwen3.8-Max. It’s genuinely impressive on paper. 2.4 trillion total parameters, 95 billion active through a sparse mixture-of-experts setup, a 1-million-token context window, and API access that went worldwide on day one. It landed at #5 on the Arena text leaderboard and #2 on the vision board about a day after launch, which makes it the highest-ranked Chinese text model anyone’s ever put up. Alibaba priced it at roughly 24% to 40% of Claude Opus 5. Open weights are supposedly landing next week.&lt;/p&gt;

&lt;p&gt;But look at the votes. It’s sitting at #5 on 3,327 Arena votes. Opus 4.6-thinking above it has 67,000. When a model debuts high on a thin vote count, that’s not a verdict, that’s a first impression with good lighting. Arena rewards new-and-polished before the crowd has actually lived with the thing.&lt;/p&gt;

&lt;p&gt;And then there’s the benchmark table. Alibaba says Qwen3.8-Max scores 86.1 on OSWorld-Verified, ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0. It claims 86.6 on Terminal-Bench 2.1 and 93.0 on PaperBench. Great numbers. Here’s my problem: &lt;a href="https://venturebeat.com/technology/qwen3-8-max-arrives-with-a-bold-claim-it-outperforms-gpt-5-6-sol-max-and-fable-5-on-agentic-computer-use" rel="noopener noreferrer"&gt;that table includes self-reported scores for the competitors too&lt;/a&gt;, and as of the day before I wrote this, &lt;a href="https://evolink.ai/blog/qwen3-8-benchmark" rel="noopener noreferrer"&gt;no independent evaluation of the model existed at all&lt;/a&gt;. A vendor grading its own homework and its rivals’ homework in the same table isn’t a lie. It’s just not evidence yet. The correct posture is interested skepticism, and I’ll happily eat my words next week when somebody who doesn’t work at Alibaba runs the thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The “Smartest Model in the World” Had an Awkward Debut
&lt;/h2&gt;

&lt;p&gt;Claude Opus 5 shipped July 24, and by this week it finally had enough Arena votes to show up properly. Anthropic and Artificial Analysis both told you it’s the most intelligent model on Earth right now. On AA’s &lt;a href="https://artificialanalysis.ai/articles/opus-5" rel="noopener noreferrer"&gt;Intelligence Index it scores 61&lt;/a&gt;, narrowly the top of the board, effectively tied with Fable 5 at 60 and ahead of GPT-5.6 Sol at 59 and Kimi K3 at 57.&lt;/p&gt;

&lt;p&gt;So where does it land on Arena Overall, now that humans have actually voted on it blind?&lt;/p&gt;

&lt;p&gt;Number 7. Behind Opus 4.6-thinking at #2 and Opus 4.7-thinking at #3. It got beaten by its own grandparents.&lt;/p&gt;

&lt;p&gt;I want to be fair here, because this is the interesting part, not a dunk. Opus 5 is #1 on the hard-benchmark axis and it takes the Arena Math crown outright (on thin preliminary votes, but still). The &lt;a href="https://artificialanalysis.ai/articles/opus-5" rel="noopener noreferrer"&gt;cost-per-task number is the real headline&lt;/a&gt;: $2.03 per Intelligence Index task versus Fable 5’s $2.75, about 26% cheaper, at the same $5/$25 sticker as Opus 4.8. That’s a real bargain if you’re doing hard agentic work.&lt;/p&gt;

&lt;p&gt;But if you upgraded to Opus 5 on the strength of the word “smartest” alone, blind human raters are quietly telling you that you bought a lateral move on everyday chat. “Best on the benchmarks” and “the answer people prefer when they don’t know which model they’re looking at” are two different axes, and this week they pointed in two different directions at the same lab.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Boring Answer That Keeps Winning
&lt;/h2&gt;

&lt;p&gt;Now the part I actually care about. If I strip away the launch confetti, what’s the model that gives me the most quality per dollar right now?&lt;/p&gt;

&lt;p&gt;Same answer as last cycle. MiMo v2.5 Pro. Xiaomi’s model from April, MIT-licensed, $0.43/$0.87 per million tokens on Arena’s list price and &lt;a href="https://openrouter.ai/xiaomi/mimo-v2.5-pro" rel="noopener noreferrer"&gt;routing even cheaper on OpenRouter&lt;/a&gt; at roughly $0.35/$0.70. Listed, purchasable, no geo-lock, multiple providers. Boring. Available. Cheap as dirt.&lt;/p&gt;

&lt;p&gt;What makes this week different is that two completely unrelated methodologies pointed at the same model. Arena’s price sort puts MiMo as the cheapest thing inside the competitive band of four separate categories. And Artificial Analysis, which measures hard-benchmark capability instead of crowd preference, &lt;a href="https://artificialanalysis.ai/models/mimo-v2-5-pro" rel="noopener noreferrer"&gt;puts MiMo on its Intelligence-vs-Cost Pareto frontier&lt;/a&gt;. Its “most attractive quadrant.” When the popularity metric and the capability metric independently name the same cheap model, that’s about as strong a buy signal as this newsletter ever gets.&lt;/p&gt;

&lt;p&gt;The honest caveat: MiMo scores 42 on AA’s Intelligence Index while the frontier sits at 57 to 61, and it’s slow at around 47 tokens per second. So it’s the crowd favorite and cost-efficient for what it is. But it’s not frontier-grade on hard reasoning, and it won’t win a latency race. For everyday work and agentic loops where you care about the bill, that’s a trade I’ll take every time. For a gnarly proof or a nasty debugging session, spend the money.&lt;/p&gt;

&lt;p&gt;That’s the whole hype-versus-value story this week in one line. Qwen3.8-Max is new, loud, and unproven. MiMo is old, quiet, and double-confirmed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheapskate Picks: Where the Actual Money Is
&lt;/h2&gt;

&lt;p&gt;This is the section I write the newsletter for. The method is simple and I’ll say it once so the table makes sense. For each Arena category, take the leader’s rating, draw a line 50 points below it, and that’s the competitive band. Everything inside that band is, statistically, a rounding error away from the “best” model. Then I sort the band by output price and pick the cheapest thing in it. The trick, and the thing I screwed up in a past issue, is that the band is defined by points, not by rank. It runs way deeper than the visible top 20. The Coding band this week is 53 models deep. The cheap open-weight models live down in the 20s, 30s, and 40s, sitting a couple of points below premium brands while costing an order of magnitude less. Truncate at rank 20 and you delete the entire reason this section exists.&lt;/p&gt;

&lt;p&gt;So I pulled the full tables and computed the bands in code. Here’s where the money is:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Leader&lt;/th&gt;
&lt;th&gt;$ leader&lt;/th&gt;
&lt;th&gt;Cheapskate pick&lt;/th&gt;
&lt;th&gt;$ pick&lt;/th&gt;
&lt;th&gt;Δ rating&lt;/th&gt;
&lt;th&gt;Price ratio&lt;/th&gt;
&lt;th&gt;AA Pareto&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Overall&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#40)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;−43&lt;/td&gt;
&lt;td&gt;~57×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#29)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;−35&lt;/td&gt;
&lt;td&gt;~57×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative Writing&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;Gemini 3 Flash (#23)&lt;/td&gt;
&lt;td&gt;$3&lt;/td&gt;
&lt;td&gt;−49&lt;/td&gt;
&lt;td&gt;~16.7×&lt;/td&gt;
&lt;td&gt;nearby&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction Following&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#25)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;−46&lt;/td&gt;
&lt;td&gt;~57×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hard Prompts&lt;/td&gt;
&lt;td&gt;claude-fable-5&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (#27)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;−40&lt;/td&gt;
&lt;td&gt;~57×&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math&lt;/td&gt;
&lt;td&gt;claude-opus-5-max*&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;Gemini 3.6 Flash (#4)&lt;/td&gt;
&lt;td&gt;$7.50&lt;/td&gt;
&lt;td&gt;−32&lt;/td&gt;
&lt;td&gt;~3.3×&lt;/td&gt;
&lt;td&gt;nearby&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;* The Math leader is preliminary, sitting on only 231 votes. Treat that whole board as a rumor this week.&lt;/p&gt;

&lt;p&gt;A few things worth saying out loud about that table.&lt;/p&gt;

&lt;p&gt;MiMo sweeps four of six categories, all at $0.87 output, all roughly 57 times cheaper than the Fable 5 leader. Its ranks look scary (#40 in Overall) until you check the vote counts. That Overall rating is backed by 45,910 votes. It’s a far more settled number than most of the shiny preliminary top-10 entries with a few hundred votes each. Deep rank measures preference, not reliability. Don’t let it spook you.&lt;/p&gt;

&lt;p&gt;Creative Writing is the one place MiMo can’t reach, because that category rewards polish and the cheap crowd falls just below the cutoff. The pick there is Gemini 3 Flash at $3 output, and it’s clinging to the very edge of the band at 49 points back. Still 16 times cheaper than the leader.&lt;/p&gt;

&lt;p&gt;Math is a mess this week and I’m flagging it hard. The leader is Opus 5 on 231 votes, the band is only 8 models deep, and the cheapest thing in it is Gemini 3.6 Flash at $7.50. There’s no sub-$7 play here. Math is a “you’re paying for quality” category right now, so if you need it, budget for it and check back when the votes settle.&lt;/p&gt;

&lt;p&gt;And the throughline underneath all of it: Fable 5 still sweeps 5 of 6 Arena categories as the outright leader. The ceiling hasn’t moved in weeks. What keeps changing is the floor, and the floor keeps getting cheaper.&lt;/p&gt;

&lt;h2&gt;
  
  
  Horror Stories From the Wild
&lt;/h2&gt;

&lt;p&gt;Two this week. One is a real fire, the other is a smoke alarm.&lt;/p&gt;

&lt;p&gt;The fire: the DeepSeek V4 API migration deadline hit on July 24 at 15:59 UTC, and it hit hard. DeepSeek retired the &lt;code&gt;deepseek-chat&lt;/code&gt; and &lt;code&gt;deepseek-reasoner&lt;/code&gt; model aliases with no grace period and no fallback. Call the old names now and you get an error, full stop. &lt;a href="https://www.developersdigest.tech/blog/deepseek-chat-to-v4-migration-guide" rel="noopener noreferrer"&gt;One developer went digging through production logs and found 14,000 calls still hitting &lt;code&gt;deepseek-chat&lt;/code&gt;&lt;/a&gt;, every single one returning a 404. The fix is a one-line model-name swap, which sounds trivial until you hit the two gotchas: thinking mode moved from the model name into a request parameter, so a naive swap either silently drops your reasoning entirely or quietly turns the cheapest endpoint into a reasoning-token furnace that torches your bill. If you run anything scheduled or agentic against DeepSeek, go read your logs right now. I’ll wait.&lt;/p&gt;

&lt;p&gt;The smoke alarm: I already said it above, but it belongs here too. Qwen3.8-Max launched with a benchmark table that scores its competitors for them and zero independent verification behind any of it. That’s not a crash. It’s a “don’t rewire your whole pipeline around a press release” warning. Wait for someone outside Alibaba to run it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coming Soon (Or “Soon,” Anyway)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3.8-Max open weights.&lt;/strong&gt; Announced for roughly the week of August 10, right behind the API launch. If they’re real, the &lt;a href="https://www.scmp.com/tech/article/3362738/alibabas-ai-model-qwen38-max-made-widely-accessible-ahead-open-weights-release" rel="noopener noreferrer"&gt;independent evals that follow will be the actual story&lt;/a&gt;, not the launch table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 3.5 Pro.&lt;/strong&gt; Still vapor. It missed its July 17 target, which is somewhere around the third or fourth slip now, and &lt;a href="https://techcrunch.com/2026/07/21/google-releases-three-new-gemini-models-but-no-3-5-pro/" rel="noopener noreferrer"&gt;Google shipped three other Gemini models instead of it&lt;/a&gt; while reportedly scrapping and rebuilding the base model over hallucination and reliability problems. At this point I’ll believe it when I can call the API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kimi K3 community quants.&lt;/strong&gt; The weights &lt;a href="https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation" rel="noopener noreferrer"&gt;went public July 26 under a modified MIT license&lt;/a&gt;, so quantized community builds are showing up. Just remember the model is 2.8 trillion parameters and needs something like 1.4TB of fast memory even at four-bit. “Open weights” and “you can run it” aren’t the same sentence when you’d need 4 to 8 H100s to load the thing.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I Actually Took Away This Week
&lt;/h2&gt;

&lt;p&gt;The frontier is stuck and the bargain bin is on fire. That’s the real state of things.&lt;/p&gt;

&lt;p&gt;Anthropic still owns the top of every leaderboard that matters, and it has for a month. Meanwhile Alibaba, Xiaomi, DeepSeek, and Moonshot are locked in a race to give away nearly-as-good models for pennies, and Chinese labs now make up something like 45% of all the tokens flowing through OpenRouter. The story isn’t “who’s the smartest.” It’s been settled for weeks. The story is that the price of “good enough for almost everything” fell off a cliff and keeps falling.&lt;/p&gt;

&lt;p&gt;So here’s my honest advice, which is the same advice as last week and probably next week. Ignore the launch you read about in the news. Open the leaderboards, find the cheapest model inside the band, check that two different metrics agree it’s actually good, and run that. This week that’s MiMo v2.5 Pro at 57 times less than the model everyone’s cheering for.&lt;/p&gt;

&lt;p&gt;And go read your DeepSeek logs. Seriously. Right now.&lt;/p&gt;

&lt;p&gt;Next Tuesday I’ll be back with the coffee and the two tabs, and I fully expect a different Chinese lab to have dropped a trillion-parameter something-or-other by then. That’s the price you pay for paying attention.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>openrouter</category>
      <category>modelroundup</category>
      <category>largelanguagemodels</category>
    </item>
    <item>
      <title>Obsidian Git Sync in 2026: What Actually Works on Mobile</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Mon, 03 Aug 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/obsidian-git-sync-in-2026-what-actually-works-on-mobile-46ad</link>
      <guid>https://dev.to/eristoddle/obsidian-git-sync-in-2026-what-actually-works-on-mobile-46ad</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpqmqi30l2y8xgz45a70a.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpqmqi30l2y8xgz45a70a.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I use git every day. So when I started taking Obsidian seriously, the obvious move was to drop the vault into a repo and call it solved. Commit, push, pull, done.&lt;/p&gt;

&lt;p&gt;That works perfectly. On laptops. The second you add a phone, git stops being the easy answer and starts being the thing you fight with, which is why it is the one method I do not use for mobile in &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-across-all-your-devices-including-free-methods-1mi5"&gt;my guide to syncing Obsidian for free on every device&lt;/a&gt;. I run the Git plugin on desktop for version history and sync my iPad with something else entirely.&lt;/p&gt;

&lt;p&gt;Almost nobody searching for this is asking how to set up git. They already know git. They are asking one question: &lt;strong&gt;will this work on my phone?&lt;/strong&gt; So let’s answer that directly.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Desktop Git Is Solved, So Let’s Not Waste Your Time&lt;/li&gt;
&lt;li&gt;Why Mobile Is a Completely Different Animal&lt;/li&gt;
&lt;li&gt;What Actually Breaks, and When&lt;/li&gt;
&lt;li&gt;No, It Is Not Real-Time Sync&lt;/li&gt;
&lt;li&gt;
The Stable Path: Let a Real Git Client Do the Mobile Leg

&lt;ul&gt;
&lt;li&gt;GitSync, which is what I would use now&lt;/li&gt;
&lt;li&gt;Working Copy on iOS&lt;/li&gt;
&lt;li&gt;About mgit on Android&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Git vs Remotely Save vs LiveSync&lt;/li&gt;
&lt;li&gt;Who Git Is Actually Right For&lt;/li&gt;
&lt;li&gt;FAQ&lt;/li&gt;
&lt;li&gt;The Short Version&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Desktop Git Is Solved, So Let’s Not Waste Your Time
&lt;/h2&gt;

&lt;p&gt;If you only use Obsidian on computers, there is no article here. Put the vault in a repo, add a &lt;code&gt;.gitignore&lt;/code&gt;, commit, push. That is it. You do not need a plugin at all.&lt;/p&gt;

&lt;p&gt;If you want the commits to happen without you thinking about it, install the &lt;a href="https://github.com/Vinzent03/obsidian-git" rel="noopener noreferrer"&gt;Obsidian Git plugin&lt;/a&gt; from the community plugins tab like &lt;a href="https://dev.to/how-to-install-obsidian-plugins/"&gt;any other plugin&lt;/a&gt;, and it will do automatic commit-and-sync (commit, pull, and push) on a schedule, plus auto-pull when Obsidian starts. On desktop the plugin shells out to the actual git binary on your machine, so it behaves exactly the way git behaves. I walked through my own setup of this in the first post, including the part where it refused to give me any branch options and I had to go back to the command line like a caveman.&lt;/p&gt;

&lt;p&gt;The mobile story is not the same.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Mobile Is a Completely Different Animal
&lt;/h2&gt;

&lt;p&gt;Neither iOS nor Android will let an app run the real git binary. So the plugin does the only thing it can do: it swaps in &lt;a href="https://isomorphic-git.org/" rel="noopener noreferrer"&gt;isomorphic-git&lt;/a&gt;, a reimplementation of git written in JavaScript, and runs that inside Obsidian.&lt;/p&gt;

&lt;p&gt;That is a different program wearing git’s clothes, and the gaps are specific:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No SSH authentication.&lt;/strong&gt; isomorphic-git does not support it. Your SSH keys and &lt;code&gt;git@github.com:&lt;/code&gt; remotes are useless here. You are on HTTPS with a personal access token, and you get to store that token on your phone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No rebase merge strategy.&lt;/strong&gt; If your workflow assumes rebase, mobile does not have it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No submodules.&lt;/strong&gt; Some people build vaults out of submodules. Not on a phone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No git-lfs.&lt;/strong&gt; If your repo uses Large File Storage, mobile is out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repo size is capped by memory.&lt;/strong&gt; All of this happens inside the app’s RAM.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one is the one that actually kills people, and it deserves its own section.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Breaks, and When
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4zalmr5w7qfh0zjkmupq.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4zalmr5w7qfh0zjkmupq.jpg" alt="What Actually Breaks, and When" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The plugin’s documentation does not hedge. It says the git implementation on mobile is &lt;strong&gt;very unstable&lt;/strong&gt; , that it would not recommend using this plugin on mobile, and that you should try other syncing services instead.&lt;/p&gt;

&lt;p&gt;That is the developer of the plugin telling you not to use his plugin. I have never seen a clearer signal in an Obsidian community plugin.&lt;/p&gt;

&lt;p&gt;The docs get specific about the failure mode too. Depending on your device and how much free RAM it has, Obsidian may crash on clone or pull, throw buffer overflow errors, or just run forever without finishing.&lt;/p&gt;

&lt;p&gt;Notice what all three have in common. They are memory failures, not sync failures. So the variable that decides whether this works for you is not “is my vault big” in some absolute sense. It is &lt;strong&gt;your vault plus its entire history versus whatever RAM your phone has free right now.&lt;/strong&gt; Which means it can work fine for weeks and then fail because you had a bunch of tabs open.&lt;/p&gt;

&lt;p&gt;I am not going to give you a magic megabyte number, because there isn’t one. What I can tell you is the shape of it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;strong&gt;initial clone is the worst moment&lt;/strong&gt; , by a wide margin. Biggest memory spike you will ever ask it to perform, and where most people give up before they reach daily use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attachments are the accelerant.&lt;/strong&gt; A pure-markdown vault is tiny. Start pasting screenshots and PDFs into it and the repo gets heavy fast, and unlike your notes, images do not compress or diff. Every version of that screenshot is in your history forever.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;History accumulates even if the vault doesn’t.&lt;/strong&gt; A three year old vault with automatic commits every ten minutes has a lot of objects in it, and the clone deals with all of them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Also worth knowing: an older separate &lt;code&gt;obsidian-git-mobile&lt;/code&gt; repo exists and shows up in search results. Do not install it. It was archived back in September 2022 and its functionality was folded into the main plugin.&lt;/p&gt;

&lt;h2&gt;
  
  
  No, It Is Not Real-Time Sync
&lt;/h2&gt;

&lt;p&gt;“Obsidian git plugin real-time sync” is one of the most searched versions of this question, and I think people are hoping the answer has changed. It has not.&lt;/p&gt;

&lt;p&gt;Git is not a sync engine. It is a version control system you are using as a sync engine, and the difference shows up exactly here. The plugin’s automatic commit-and-sync runs &lt;strong&gt;on a timer measured in minutes&lt;/strong&gt;. There is no file watcher pushing your keystrokes to a server.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2uy6mn8c8b2psr3b8h41.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2uy6mn8c8b2psr3b8h41.jpg" alt="No, It Is Not Real-Time Sync" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So the real behavior is: you type a note on your laptop, and some number of minutes later it gets committed and pushed. Then your phone pulls whenever it pulls. Best case, a gap of a few minutes. Worst case, you opened the app before it finished pulling and you are now editing a stale copy of a note, which is how you manufacture a merge conflict on a device with no good way to resolve one.&lt;/p&gt;

&lt;p&gt;If you want changes to appear on the other device almost immediately, git is the wrong tool and Self-hosted LiveSync is the one built for that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Stable Path: Let a Real Git Client Do the Mobile Leg
&lt;/h2&gt;

&lt;p&gt;Here is the setup that actually holds up, and it is the same principle for both platforms. &lt;strong&gt;Do not make Obsidian do the git.&lt;/strong&gt; Install a dedicated git client on the phone, let it clone and push the repo into a real folder, and then point Obsidian at that folder as a vault. Obsidian just edits markdown files. The git app handles git. Neither one has to be clever.&lt;/p&gt;

&lt;h3&gt;
  
  
  GitSync, which is what I would use now
&lt;/h3&gt;

&lt;p&gt;The plugin’s own docs point at &lt;a href="https://github.com/ViscousPot/GitSync" rel="noopener noreferrer"&gt;GitSync&lt;/a&gt; as the alternative, and having looked at it, that recommendation is correct and it is the biggest thing that has changed in this space.&lt;/p&gt;

&lt;p&gt;GitSync is a mobile git client built specifically for syncing a folder between a git remote and a local directory, which is precisely the job. It runs on Android 5+ and iOS 13+, it is open source under GPL-3.0, and it is on Google Play, the App Store, F-Droid, and IzzyOnDroid. It won a 2024 Gem of the Year award in the Obsidian tools category, so the vault use case is not an accident, it is the point.&lt;/p&gt;

&lt;p&gt;The part that matters most: it does &lt;strong&gt;not&lt;/strong&gt; use isomorphic-git. It uses a native Rust core built on &lt;code&gt;git2-rs&lt;/code&gt;. So it supports SSH, along with HTTPS and OAuth (GitHub, GitLab, and Gitea), and it does not inherit the memory ceiling that makes the plugin fall over. It also does background sync, which the plugin cannot do because the plugin only runs while Obsidian is open.&lt;/p&gt;

&lt;p&gt;It is free to download. Premium is a &lt;strong&gt;$24.99 one-time unlock&lt;/strong&gt; that adds additional repositories, Git LFS, git-crypt, and priority issue tagging, and you can also get it by becoming a GitHub Sponsor. There is a separate GitSync AI subscription at $6.99, which has nothing to do with syncing and which you can ignore. Note that Git LFS thing: it is a capability the Obsidian plugin does not have on mobile at any price.&lt;/p&gt;

&lt;h3&gt;
  
  
  Working Copy on iOS
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://apps.apple.com/us/app/working-copy-git-client/id896694807" rel="noopener noreferrer"&gt;Working Copy&lt;/a&gt; is the old reliable iOS git client and it is genuinely excellent. The catch that nobody mentions in sync articles: &lt;strong&gt;the free version cannot push.&lt;/strong&gt; The App Store description says it straight, that you need to unlock pro features “such as the ability to push commits and manage more than 5 repositories.” Pro Unlock is $35.99. There is a 10 day trial, and it is free for students through the GitHub Student Developer Pack.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzr5udbfc03r9xbull6gj.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzr5udbfc03r9xbull6gj.jpg" alt="Working Copy on iOS" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One wrinkle on “one-time”: you get permanent access to every pro feature that exists at purchase plus anything added over the next year. Features added after that need a Pro Upgrade ($17.99 for the recent ones). You never lose what you bought, but “buy once, get everything forever” is not quite the deal.&lt;/p&gt;

&lt;p&gt;For a sync workflow, “cannot push” means “cannot sync,” so budget for it. What you get is a mature app other iOS apps can read from, which is what makes the handoff to Obsidian clean.&lt;/p&gt;

&lt;h3&gt;
  
  
  About mgit on Android
&lt;/h3&gt;

&lt;p&gt;I have recommended &lt;a href="https://manichord.com/projects/mgit.html" rel="noopener noreferrer"&gt;mgit&lt;/a&gt; before and I am walking that back. The advice used to be to skip the Google Play build (users report it broken) and grab it from &lt;a href="https://f-droid.org/packages/com.manichord.mgit/" rel="noopener noreferrer"&gt;F-Droid&lt;/a&gt; instead. That is still true, but the F-Droid listing’s latest release is version 1.7.0, from &lt;strong&gt;January 4, 2023.&lt;/strong&gt; That is over three years of nothing. It may well still work for you, but I am not going to tell somebody to trust their notes to an abandoned app when GitSync exists, is maintained, and is better.&lt;/p&gt;

&lt;h2&gt;
  
  
  Git vs Remotely Save vs LiveSync
&lt;/h2&gt;

&lt;p&gt;People search for this comparison literally, so here it is without hedging.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Git&lt;/th&gt;
&lt;th&gt;Remotely Save&lt;/th&gt;
&lt;th&gt;Self-hosted LiveSync&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What it really is&lt;/td&gt;
&lt;td&gt;Version control used as sync&lt;/td&gt;
&lt;td&gt;File sync to cloud storage&lt;/td&gt;
&lt;td&gt;Live replication over CouchDB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mobile story&lt;/td&gt;
&lt;td&gt;Bad via plugin, fine via GitSync&lt;/td&gt;
&lt;td&gt;Good, it is the whole point&lt;/td&gt;
&lt;td&gt;Good, but you run a server&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speed&lt;/td&gt;
&lt;td&gt;Minutes, on a timer&lt;/td&gt;
&lt;td&gt;Minutes, on a timer&lt;/td&gt;
&lt;td&gt;Near instant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Version history&lt;/td&gt;
&lt;td&gt;Excellent, it is the entire feature&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Setup burden&lt;/td&gt;
&lt;td&gt;Low if you know git&lt;/td&gt;
&lt;td&gt;Lowest&lt;/td&gt;
&lt;td&gt;Highest by a mile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conflict handling&lt;/td&gt;
&lt;td&gt;Real merges, and real merge conflicts&lt;/td&gt;
&lt;td&gt;Duplicate files&lt;/td&gt;
&lt;td&gt;Handled at the database level&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Pick Remotely Save&lt;/strong&gt; if you want your notes on your phone with the least friction. This is what I actually run. &lt;strong&gt;Pick LiveSync&lt;/strong&gt; if you want changes to land before you can switch apps, and you are the kind of person who is fine maintaining a CouchDB instance. &lt;strong&gt;Pick git&lt;/strong&gt; if you want history.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Git Is Actually Right For
&lt;/h2&gt;

&lt;p&gt;Git is right for you if you already live in git and what you want out of sync is a &lt;strong&gt;time machine&lt;/strong&gt; , not a file transfer. Being able to see that you deleted three paragraphs on March 12th and get them back is a genuinely different capability from having your notes on two devices, and no cloud sync method matches it. Add the &lt;a href="https://github.com/kometenstaub/obsidian-version-history-diff" rel="noopener noreferrer"&gt;Version History Diff plugin&lt;/a&gt; and you can read those diffs inside Obsidian.&lt;/p&gt;

&lt;p&gt;Git is wrong for you if what you want is your notes on your phone. That is a sync problem, and you should solve it with a sync tool.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fafwesuwz3g3ta2g8wha6.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fafwesuwz3g3ta2g8wha6.jpg" alt="Who Git Is Actually Right For" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Which leads to the setup I think most git people should actually run, and it is a hybrid: &lt;strong&gt;Git plugin on desktop only, for history and backup. Something else for the mobile leg.&lt;/strong&gt; My iPad talks to Dropbox through &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-on-iphone-and-ipad-for-free-in-2026-427p"&gt;Remotely Save, which is the route I walk through in the iPhone and iPad guide&lt;/a&gt;, my laptops keep a git history, and the two never touch each other.&lt;/p&gt;

&lt;p&gt;That last part is a hard rule, and it is the same one that &lt;a href="https://dev.to/eristoddle/obsidian-icloud-sync-in-2026-including-the-windows-problem-4503"&gt;wrecks iCloud vaults on Windows&lt;/a&gt;: never point two sync systems at the same vault folder at once. Git and a cloud sync client fighting over the same &lt;code&gt;.git&lt;/code&gt; directory is a genuinely creative way to destroy a repo.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does the Obsidian Git plugin work on mobile?&lt;/strong&gt; Technically yes, on both iOS and Android, using isomorphic-git instead of real git. Practically, the plugin’s own documentation says the mobile implementation is very unstable and recommends using a different syncing service instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is the Obsidian Git plugin stable on iOS and Android?&lt;/strong&gt; No. The documented failure modes are Obsidian crashing during clone or pull, buffer overflow errors, or the operation running indefinitely, all depending on your device’s available RAM.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does the Obsidian Git plugin do real-time sync?&lt;/strong&gt; No. Automatic commit-and-sync runs on a schedule measured in minutes. If you need near instant propagation, Self-hosted LiveSync is the method built for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I use SSH with the Obsidian Git plugin on mobile?&lt;/strong&gt; No. isomorphic-git does not support SSH authentication, so mobile requires HTTPS with a personal access token. GitSync does support SSH, because it uses a native git implementation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the best Android support option for Obsidian and git?&lt;/strong&gt; GitSync. It is maintained, it is on Play, F-Droid, and IzzyOnDroid, and it uses native git rather than a JavaScript reimplementation. mgit’s last F-Droid release was January 2023.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I use the Git plugin with Nextcloud for multi-user sync?&lt;/strong&gt; You can host the remote anywhere, Nextcloud included. But git is not a live collaboration layer. Two people editing the same note between pushes produces a merge conflict, not a merged note. For shared vaults, look at LiveSync or Obsidian Sync’s shared vaults.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I uninstalled the Git plugin. Where did my vault go?&lt;/strong&gt; Nowhere. The plugin does not own your vault, and removing it leaves your files and &lt;code&gt;.git&lt;/code&gt; folder untouched. The confusion is almost always a mobile vault location problem: if you cloned through the plugin on iOS, the vault lives inside Obsidian’s own app storage, which is not somewhere you can casually browse to. Clone with a real git client into a folder you picked, and you always know where your notes are.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Short Version
&lt;/h2&gt;

&lt;p&gt;Git on Obsidian desktop is great and you should probably be doing it for the history alone. Git on Obsidian mobile through the plugin is a thing the plugin’s own author asked you not to do, and nobody seems to want to say that plainly.&lt;/p&gt;

&lt;p&gt;If you want git on your phone anyway, and that is a reasonable thing to want, use GitSync and let a native git implementation do the work. If you just want your notes to show up on your phone, stop trying to make git do it and go use an &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-across-all-your-devices-including-free-methods-1mi5"&gt;actual sync method&lt;/a&gt;. I use git for the time machine and Dropbox for the phone, and I have not lost a note yet.&lt;/p&gt;

</description>
      <category>obsidian</category>
      <category>git</category>
      <category>sync</category>
      <category>ios</category>
    </item>
    <item>
      <title>How to Sync Obsidian on iPhone and iPad for Free in 2026</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Thu, 30 Jul 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/how-to-sync-obsidian-on-iphone-and-ipad-for-free-in-2026-427p</link>
      <guid>https://dev.to/eristoddle/how-to-sync-obsidian-on-iphone-and-ipad-for-free-in-2026-427p</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4rcgyal49zfkw9hz6k7b.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4rcgyal49zfkw9hz6k7b.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;My iPad is the device that made me actually think about Obsidian sync instead of just having opinions about it. Laptops are easy. Two MacBook Pros and a Windows desktop will happily share a folder in any cloud service you point them at. Then you add an iPad and the thing that worked everywhere else does not work at all, because iOS doesn’t let apps do the stuff sync tools need to do.&lt;/p&gt;

&lt;p&gt;I run Dropbox with Remotely Save for my iPad, and I pay nothing for it, which is the whole point of &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-across-all-your-devices-including-free-methods-1mi5"&gt;my guide to syncing Obsidian for free on every device&lt;/a&gt;. This post is the iOS-only version of that. Four options that work, what each one costs, and why the answer depends on what your desktop computer is rather than anything about your phone.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why iOS Is the Hard Case&lt;/li&gt;
&lt;li&gt;The Four Routes That Actually Work&lt;/li&gt;
&lt;li&gt;Route 1: iCloud, If Your Desktop Is a Mac&lt;/li&gt;
&lt;li&gt;Route 2: Remotely Save, the One I Actually Run&lt;/li&gt;
&lt;li&gt;
Route 3: Syncthing on iOS, the Route Everyone Asks About

&lt;ul&gt;
&lt;li&gt;Möbius Sync&lt;/li&gt;
&lt;li&gt;SyncTrain&lt;/li&gt;
&lt;li&gt;The Reality of Both&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Route 4: Obsidian Sync, When You Are Done Fighting&lt;/li&gt;
&lt;li&gt;OneDrive on iOS Is Its Own Problem&lt;/li&gt;
&lt;li&gt;The Files App Trap&lt;/li&gt;
&lt;li&gt;Frequently Asked Questions&lt;/li&gt;
&lt;li&gt;The Verdict&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why iOS Is the Hard Case
&lt;/h2&gt;

&lt;p&gt;Every other platform gives you a folder. iOS gives you a sandbox.&lt;/p&gt;

&lt;p&gt;On a Mac or a Windows box, a sync tool runs as a background service, watches a directory, and pushes changes when they happen. None of that is available to an iOS app. There are no background daemons. An app that is not on screen is mostly not running, and iOS decides when to wake it, not you. Apps also cannot freely read each other’s files, so “just point Obsidian at the Dropbox folder” is not a thing that exists here the way it does on desktop.&lt;/p&gt;

&lt;p&gt;That is also why there is no native Syncthing for iOS. Syncthing is a daemon. iOS does not do daemons. Everything you will find on the App Store is a third-party app that embeds the Syncthing engine and works within Apple’s rules, which is a different and more limited thing.&lt;/p&gt;

&lt;p&gt;So the question is not “which sync tool is best.”, but “which compromise do I want.”&lt;/p&gt;

&lt;h2&gt;
  
  
  The Four Routes That Actually Work
&lt;/h2&gt;

&lt;p&gt;Pick by what your desktop is.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Your desktop&lt;/th&gt;
&lt;th&gt;Best free route&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mac only&lt;/td&gt;
&lt;td&gt;iCloud&lt;/td&gt;
&lt;td&gt;Built into Obsidian’s iOS vault creation, near zero setup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows&lt;/td&gt;
&lt;td&gt;Remotely Save with Dropbox or S3&lt;/td&gt;
&lt;td&gt;Works identically on both ends, no Apple dependency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Linux&lt;/td&gt;
&lt;td&gt;Remotely Save, or Syncthing via SyncTrain&lt;/td&gt;
&lt;td&gt;Linux has no iCloud and no official OneDrive client&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mixed, or you want no cloud at all&lt;/td&gt;
&lt;td&gt;Syncthing via SyncTrain or Möbius Sync&lt;/td&gt;
&lt;td&gt;Peer to peer, nothing stored on anyone’s server&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;You want it to just work&lt;/td&gt;
&lt;td&gt;Obsidian Sync, $4/mo annual&lt;/td&gt;
&lt;td&gt;Not free, but it is the one that never needs troubleshooting&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Now the details.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route 1: iCloud, If Your Desktop Is a Mac
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5yhwutk138fxy69spzd.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5yhwutk138fxy69spzd.jpg" alt="Route 1: iCloud, If Your Desktop Is a Mac" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you are all Apple, stop here. Create the vault on your iPhone or iPad, toggle on &lt;strong&gt;Store in iCloud&lt;/strong&gt; during creation, and you are done. It is the only route with first-class support inside Obsidian’s own iOS app.&lt;/p&gt;

&lt;p&gt;I am not going to repeat the setup steps, because I wrote the long version already. See &lt;a href="https://dev.to/eristoddle/obsidian-icloud-sync-in-2026-including-the-windows-problem-4503"&gt;the full iCloud walkthrough&lt;/a&gt; for creating versus migrating a vault, the “Optimize Mac Storage” setting that removes your notes, and the part where it falls apart if Windows is involved.&lt;/p&gt;

&lt;p&gt;The short version: works perfectly on Apple hardware, unreliably the moment it is not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route 2: Remotely Save, the One I Actually Run
&lt;/h2&gt;

&lt;p&gt;This is my setup. &lt;a href="https://github.com/remotely-save/remotely-save" rel="noopener noreferrer"&gt;Remotely Save&lt;/a&gt; is a community plugin that syncs your vault to a storage backend you already have. The free tier covers S3 and anything S3-compatible like Cloudflare R2 or Backblaze B2, plus Dropbox, WebDAV including Nextcloud and Synology, and OneDrive personal with a caveat I get to below.&lt;/p&gt;

&lt;p&gt;The reason it wins on iOS is that it sidesteps the sandbox problem. Remotely Save runs &lt;em&gt;inside&lt;/em&gt; Obsidian. It does not need to watch a folder or run in the background, because it syncs when Obsidian is open and you tell it to, or on the schedule you set. iOS does not have to cooperate for it to work.&lt;/p&gt;

&lt;p&gt;The setup is the same on every platform, which is the other reason I use it. Install the plugin from Community Plugins, pick your service, authenticate, set a sync interval, and run a manual sync once to seed the vault. Do the desktop side first and let it finish before you touch the iPad, same as any sync method.&lt;/p&gt;

&lt;p&gt;The tradeoff is honest: sync happens when Obsidian is open. Open the app, wait a couple of seconds, then start typing. That is the deal, and after a couple of years of it I have stopped noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route 3: Syncthing on iOS, the Route Everyone Asks About
&lt;/h2&gt;

&lt;p&gt;This is the biggest cluster of searches on this topic and the one the forum threads are full of, so here is the current state of it in 2026.&lt;/p&gt;

&lt;p&gt;You cannot run &lt;a href="https://syncthing.net/" rel="noopener noreferrer"&gt;Syncthing&lt;/a&gt; on iOS directly. You run one of two apps that wrap it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Möbius Sync
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://mobiussync.com/" rel="noopener noreferrer"&gt;Free to download&lt;/a&gt;. The catch is the one people keep asking about in forum threads: usage within the Möbius Sync sandbox is free up to 20MB. Past that you need the in-app purchase, &lt;strong&gt;Unlimited file sync, $4.99&lt;/strong&gt; , which is a one-time purchase and not a subscription.&lt;/p&gt;

&lt;p&gt;For an Obsidian vault, 20MB is the deciding number. A text-only vault of a few thousand notes fits under it well. A vault with images, PDFs, or scanned documents blows through it immediately. But five dollars once is a fair price for a sync solution.&lt;/p&gt;

&lt;h3&gt;
  
  
  SyncTrain
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy735929jvzr4770kgy6z.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy735929jvzr4770kgy6z.jpg" alt="SyncTrain" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apps.apple.com/us/app/synctrain/id6553985316" rel="noopener noreferrer"&gt;The other option&lt;/a&gt;, and worth knowing about because it is truly free. No in-app purchases, open source, iOS 17 or later, actively maintained. It has deeper Shortcuts integration, which is the closest thing iOS offers to background syncing, since you can trigger a sync from an automation rather than remembering to open the app.&lt;/p&gt;

&lt;p&gt;I appreciate that its own App Store description tells you this: “Do not use Synctrain for back-up purposes, and always keep a back-up of your data.” An app that volunteers its own limitations is an app I trust more than one that does not.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Reality of Both
&lt;/h3&gt;

&lt;p&gt;Syncthing on iOS is peer to peer, which means your notes never sit on anyone else’s server. That is the real reason to pick it.&lt;/p&gt;

&lt;p&gt;But it also means both devices have to be awake and reachable at the same time for anything to happen. Your Mac asleep in a bag syncs nothing. This is why people describe iOS Syncthing as flaky when it is actually working exactly as designed. There is no server holding your changes until the other device shows up. If you want that, you want one of the other three routes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Route 4: Obsidian Sync, When You Are Done Fighting
&lt;/h2&gt;

&lt;p&gt;Not free, so it does not really belong in a free guide, except that it belongs in every honest one. &lt;a href="https://obsidian.md/sync" rel="noopener noreferrer"&gt;Sync Standard&lt;/a&gt; is $5 a month billed monthly or $4 billed annually, it is end-to-end encrypted by default, and it is the only option on this page where iOS is a first-class platform rather than a workaround.&lt;/p&gt;

&lt;p&gt;If you have read this far and the phrase “sandbox limitations” has stopped being interesting, that is the answer. Four dollars a month is roughly one coffee.&lt;/p&gt;

&lt;h2&gt;
  
  
  OneDrive on iOS Is Its Own Problem
&lt;/h2&gt;

&lt;p&gt;OneDrive keeps showing up in the search data for this topic, so it deserves a direct answer: it is the most awkward of the cloud options on iOS.&lt;/p&gt;

&lt;p&gt;Remotely Save does support OneDrive on the free tier, but with a restriction worth understanding before you start. The free version can only connect to the App Folder, meaning &lt;code&gt;/Apps/remotely-save&lt;/code&gt; inside your OneDrive. The PRO version is what connects to the root folder.&lt;/p&gt;

&lt;p&gt;That means you can’t point free Remotely Save at a vault that already lives somewhere else in your OneDrive. The vault has to live in the app’s own folder. If you were planning to sync a vault that sits alongside your work documents, that is the plan that breaks, and it breaks after you have already set everything up.&lt;/p&gt;

&lt;p&gt;If you are on OneDrive because your job is on Microsoft 365, this is workable as long as you keep the vault in the app folder. If you are on OneDrive out of habit, Dropbox or an S3 bucket will give you a smoother ride for free.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Files App Trap
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbk6g9a2rc7c6j2vhgy2w.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbk6g9a2rc7c6j2vhgy2w.jpg" alt="The Files App Trap" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here is the mistake that generates the most confused forum posts.&lt;/p&gt;

&lt;p&gt;Your Obsidian vault on iOS is not just a folder in the Files app that you can put anywhere. Obsidian stores its iOS vaults in its own app container, which surfaces in Files as &lt;strong&gt;iCloud Drive &amp;gt; Obsidian&lt;/strong&gt;. That location is not decoration. Obsidian’s docs are explicit that vaults should live inside the Obsidian folder in iCloud Drive.&lt;/p&gt;

&lt;p&gt;Drop a vault in some other Files location and you get the confusing outcome: it looks fine in Files, it may even sync between your Macs, and the iOS app will not list it. People then conclude sync is broken when the vault was never somewhere Obsidian could see.&lt;/p&gt;

&lt;p&gt;The other half of this trap is that “it’s in Files” does not mean “it’s on the device.” A file showing a cloud icon is a placeholder. Obsidian searching a vault full of placeholders returns incomplete results and nothing tells you why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can I sync Obsidian on iPhone for free?&lt;/strong&gt; Yes. iCloud if your desktop is a Mac, Remotely Save with Dropbox or S3 for everything else, or SyncTrain if you want peer to peer with no cloud involved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Syncthing work on iOS?&lt;/strong&gt; Not directly, because iOS does not allow background daemons. You use a wrapper app, either Möbius Sync or SyncTrain, and both devices must be awake at the same time to sync.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Möbius Sync free?&lt;/strong&gt; Free to download and free within its sandbox up to 20MB. Beyond that the Unlimited file sync in-app purchase is $4.99, one time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I sync Obsidian between my PC and iPhone?&lt;/strong&gt; Remotely Save with Dropbox or an S3 bucket. iCloud is the wrong tool when a Windows machine is involved, for reasons covered in the iCloud walkthrough.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why are my notes out of date when I open Obsidian on my phone?&lt;/strong&gt; Because no iOS sync method runs continuously in the background. Open the app and give it a few seconds before you start editing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where does Obsidian store vaults on iPhone and iPad?&lt;/strong&gt; In its own container, visible in the Files app as iCloud Drive &amp;gt; Obsidian. Vaults kept elsewhere may not appear in the app’s vault list.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Verdict
&lt;/h2&gt;

&lt;p&gt;The iOS question answers itself once you stop asking it about iOS. Mac desktop, use iCloud. Windows or Linux desktop, use Remotely Save. Want nothing on anyone’s server, use SyncTrain and accept that both devices have to be on. Want to stop thinking about it, pay the four dollars.&lt;/p&gt;

&lt;p&gt;What none of them do is sync silently in the background the way you are used to on a laptop, and no blog post is going to fix that, because it is Apple’s design and not a gap in the tooling. Once you build the two-second pause into opening the app, iOS sync stops being a problem and goes back to being a folder.&lt;/p&gt;

&lt;p&gt;Mine has been a folder for four years now. I only think about it when I write about it.&lt;/p&gt;

</description>
      <category>obsidian</category>
      <category>ios</category>
      <category>iphone</category>
      <category>ipad</category>
    </item>
    <item>
      <title>Apple Books Hides Your PDF Highlights. My Obsidian Plugin Now Digs Them Out.</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Wed, 29 Jul 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/apple-books-hides-your-pdf-highlights-my-obsidian-plugin-now-digs-them-out-lio</link>
      <guid>https://dev.to/eristoddle/apple-books-hides-your-pdf-highlights-my-obsidian-plugin-now-digs-them-out-lio</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh5unzcqi7rgb76dnf5px.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh5unzcqi7rgb76dnf5px.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Of all &lt;a href="https://dev.to/eristoddle/the-obsidian-plugin-collection-i-built-one-free-kiro-credit-at-a-time-2b42"&gt;the little plugins I’ve vibe coded into existence&lt;/a&gt;, the &lt;a href="https://github.com/eristoddle/apple-books-annotation-import" rel="noopener noreferrer"&gt;Apple Books Annotation Import&lt;/a&gt; plugin is the one I actually use. Not “use” in the way you use a project once, screenshot it for a blog post, and never open again. I mean I run it every week. It’s my favorite thing I’ve built, and it’s been humming along for &lt;a href="https://dev.to/eristoddle/jules-ai-the-currently-free-coding-assistant-that-cant-follow-directions-but-gets-shit-done-33k3"&gt;over a year without me touching it&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Then a few weeks ago I went looking for a highlight I knew I’d made in a PDF. It wasn’t there. None of my PDF highlights were there. A year of trusting this thing, and it turns out it had quietly been ignoring an entire category of my books the whole time.&lt;/p&gt;

&lt;p&gt;So I went digging. And it turns out Apple Books handles PDFs in a completely different way. This is the story of getting those highlights out, and the 1.1.0 release that finally does it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Quick Refresher On What This Thing Does&lt;/li&gt;
&lt;li&gt;Where Apple Actually Hides PDF Highlights&lt;/li&gt;
&lt;li&gt;Reading Highlights Out Of A Raw PDF&lt;/li&gt;
&lt;li&gt;Hooking It Into The Machine That Already Worked&lt;/li&gt;
&lt;li&gt;The Tradeoff: It’s Slower, But It’s Everything Now&lt;/li&gt;
&lt;li&gt;How To Install It (It’s Not In Community Plugins Yet)&lt;/li&gt;
&lt;li&gt;The Takeaway&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A Quick Refresher On What This Thing Does
&lt;/h2&gt;

&lt;p&gt;If you’ve &lt;a href="https://dev.to/eristoddle/creating-an-obsidian-plugin-with-claude-ai-gaj"&gt;never seen the plugin before&lt;/a&gt;, here’s the pitch. It’s macOS desktop only, and I’m fine with that, because being desktop only is the entire reason it can do what it does. I read in an iPad and I pay for the cheapest rung of iCloud specificallly because of this plugin to get the books where the plugin runs.&lt;/p&gt;

&lt;p&gt;When you highlight something in Apple Books, that highlight has to live somewhere on disk. On a Mac it does. On an iPhone it’s locked in a sandbox you can’t reach. So the plugin sits on the desktop where all the good data is and pulls from every source it can find:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Books SQLite databases.&lt;/strong&gt; This is where your EPUB highlights and notes actually live, buried in &lt;code&gt;~/Library/Containers/com.apple.iBooksX&lt;/code&gt;. The plugin &lt;a href="https://dev.to/eristoddle/exporting-mac-osx-book-highlights-into-an-obsidian-vault-or-markdown-files-40lg"&gt;reads them straight out of SQLite&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The EPUB files themselves.&lt;/strong&gt; The database doesn’t have everything. So the plugin also cracks open the EPUB to grab the real metadata: ISBN, publisher, language, subjects, and the cover image.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;And now, as of 1.1.0, the PDFs themselves.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Out of all that it builds a clean markdown note per book: your highlights as blockquotes, color-coded to match the highlighter you used, optional notes, dates, citations, cover image, and an author page with a Dataview query that lists every book by that author. It has a smart overwrite mode that hashes the note so re-importing only rewrites files that actually changed. Pick “Import all books” or cherry-pick from a list.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ftrfsvm2cqtxensrldr32.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ftrfsvm2cqtxensrldr32.png" alt="Apple Books Annotation Import settings" width="800" height="958"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The whole point is completeness. I’m not scraping one source, I’m triangulating across three. Which is why the missing PDF highlights bugged me so much. There was a hole in the completeness.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Apple Actually Hides PDF Highlights
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff8u5coee3a25dpgmt1bd.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff8u5coee3a25dpgmt1bd.jpg" alt="A Quick Refresher On What This Thing Does" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here’s the part that took real detective work, and where my first assumption was flat wrong.&lt;/p&gt;

&lt;p&gt;My mental model was simple: highlights go in the SQLite database, PDFs are books, therefore PDF highlights go in the database. So I looked. And there they were. Sort of.&lt;/p&gt;

&lt;p&gt;The database had a row for every PDF I’d ever opened. But every single one looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;ZANNOTATIONSELECTEDTEXT&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;(empty)&lt;/span&gt;
&lt;span class="py"&gt;ZANNOTATIONSTYLE&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;0&lt;/span&gt;
&lt;span class="py"&gt;ZANNOTATIONTYPE&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;3&lt;/span&gt;
&lt;span class="py"&gt;ZPLUSERDATA&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;&amp;lt;binary plist: BKPageLocation, pageOffset 21&amp;gt;&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No text. No color. No note. Just a page number wrapped in a binary property list. That’s not a highlight. That’s a bookmark. Apple Books drops one of these in the database for every PDF so it can remember what page you were on, and my plugin had been correctly throwing them away for a year because they have no selected text.&lt;/p&gt;

&lt;p&gt;So the highlights weren’t in the database at all. That meant they had to be in the PDF files, which on a Mac live in your iCloud Books folder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;~/Library/Mobile Documents/iCloud~com~apple~iBooks/Documents/

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I grepped one of my highlighted PDFs for the PDF highlight marker, and there it was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;58 occurrences of /Subtype /Highlight
71 QuadPoints arrays

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s the answer. Apple Books writes PDF highlights back into the PDF file as standard PDF annotations. Not into its own database like it does for EPUBs. Into the file. Which is actually the more portable choice, it’s just the opposite of everywhere else the app stores things.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading Highlights Out Of A Raw PDF
&lt;/h2&gt;

&lt;p&gt;A PDF highlight annotation doesn’t store the text you highlighted. That would be too easy. It stores &lt;code&gt;QuadPoints&lt;/code&gt;, which are the rectangles the yellow marker was painted over, in PDF coordinate space. To get the actual words you have to line those rectangles up against the text on the page and read out whatever sits underneath.&lt;/p&gt;

&lt;p&gt;That’s a job for &lt;a href="https://github.com/mozilla/pdf.js" rel="noopener noreferrer"&gt;pdf.js&lt;/a&gt;, Mozilla’s PDF engine. It gives me the annotations on each page and, separately, every run of text with its position. For each highlight I:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Turn its &lt;code&gt;QuadPoints&lt;/code&gt; into one box per highlighted line.&lt;/li&gt;
&lt;li&gt;Find the text on the page whose baseline falls inside each box.&lt;/li&gt;
&lt;li&gt;Clip the first and last lines, because a highlight usually starts and ends mid-sentence, snapping the cut to the nearest word boundary so I don’t slice a word in half.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjqt6ad0g0hl2a5qzsiup.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjqt6ad0g0hl2a5qzsiup.jpg" alt="Reading Highlights Out Of A Raw PDF" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I tested it against a highlight I knew the exact wording of, and it came back verbatim, starting and ending in exactly the right place. Colors came out too. pdf.js hands them back as plain RGB values from 0 to 255, and Apple’s highlighter yellow is &lt;code&gt;(250, 205, 90)&lt;/code&gt;, so I map that to the same little yellow square the EPUB highlights already use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hooking It Into The Machine That Already Worked
&lt;/h2&gt;

&lt;p&gt;This is the part I’m actually happy about, and it’s the part that keeps this from being a bolt-on mess.&lt;/p&gt;

&lt;p&gt;The plugin already had a whole pipeline for EPUBs: an &lt;code&gt;Annotation&lt;/code&gt; object shape, a markdown renderer, the smart-overwrite dedup, author page creation, file naming. All of it keyed off two internal types. So instead of writing a parallel universe for PDFs, I made the PDF extractor produce those exact same types. A PDF highlight becomes an &lt;code&gt;Annotation&lt;/code&gt;. A PDF file becomes a &lt;code&gt;BookDetail&lt;/code&gt;, with its title and author pulled from the Books library database so the note gets a real title instead of a mangled filename.&lt;/p&gt;

&lt;p&gt;Once the shapes matched, the entire existing machine just ran. The renderer, the dedup, the author pages, the color emoji, none of it knew or cared that these highlights came out of a PDF instead of a database. PDF import is just a second phase inside the same “Import all books” command, plus one new toggle in settings so you can turn it off (because it is now the slow part of the plugin).&lt;/p&gt;

&lt;p&gt;The genuinely annoying part was bundling pdf.js into a single-file Obsidian plugin. pdf.js wants to run its parser in a Web Worker loaded from a separate file, and an Obsidian plugin ships as one &lt;code&gt;main.js&lt;/code&gt;. The trick is to run the worker on the main thread by handing pdf.js its own worker module through a global:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;pdfjsWorker&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pdfjs-dist/legacy/build/pdf.worker.js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;globalThis&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;pdfjsWorker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;pdfjsWorker&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That, plus telling the bundler to leave the optional &lt;code&gt;canvas&lt;/code&gt; dependency alone, and it works. It also took &lt;code&gt;main.js&lt;/code&gt; from about 250KB to 4.3MB, because I’m now shipping an entire PDF engine inside a note-taking plugin. Such is life.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Tradeoff: It’s Slower, But It’s Everything Now
&lt;/h2&gt;

&lt;p&gt;I’m not going to pretend this is free. The database approach for EPUBs is instant. The PDF approach is not.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdr53a86l2qrh85tcyovl.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdr53a86l2qrh85tcyovl.jpg" alt="The Tradeoff: It's Slower, But It's Everything Now" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Apple doesn’t track which PDFs you’ve highlighted anywhere I can query, so the plugin has to go look. It scans your entire iCloud Books folder, does a cheap byte-level check for the highlight marker to skip the hundreds of PDFs you never marked up, and then parses the survivors with pdf.js on the main thread. In my library that’s 748 PDFs to glance at. So yes, “Import all books” takes noticeably longer than it used to, and a book with a lot of highlights can make the app pause for a second while it works.&lt;/p&gt;

&lt;p&gt;I decided I’ll take that trade every time. A slightly slower import that gets me everything beats an instant import with a hole in it.&lt;/p&gt;

&lt;p&gt;And I do mean everything. Between the SQLite databases, the EPUB files, and now the PDFs, I’m fairly confident there’s nothing left in Apple Books that I can’t pull out. EPUB highlights, PDF highlights, notes, metadata, covers. If Apple is storing it on my Mac, this plugin can reach it. I went looking for one more hidden source after the PDFs and came up empty, which for once is the good outcome.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9vh0i4p0shp2n3jhndre.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9vh0i4p0shp2n3jhndre.png" alt="A PDF book note with imported highlights" width="800" height="1211"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How To Install It (It’s Not In Community Plugins Yet)
&lt;/h2&gt;

&lt;p&gt;Fair warning: this plugin is not in the Obsidian community plugin browser. It’s mine, it’s niche, and it reads files out of your macOS system libraries, so it lives on GitHub for now. That means two ways in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The easy way, with BRAT.&lt;/strong&gt; BRAT is the Beta Reviewers Auto-update Tool, and it exists exactly for plugins like this one. Install BRAT from the community plugins, then add this repo as a beta plugin:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;eristoddle/apple-books-annotation-import

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;BRAT installs it and keeps it updated whenever I push a new release. I wrote a whole walkthrough of this if you’ve never done it: &lt;a href="https://www.stephanmiller.com/how-to-install-obsidian-plugins/#:~:text=How%20to%20Install%20Beta%20Obsidian%20Plugins%20with%20BRAT" rel="noopener noreferrer"&gt;How to Install Obsidian Plugins&lt;/a&gt;, including the BRAT section.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The manual way.&lt;/strong&gt; Grab &lt;code&gt;main.js&lt;/code&gt;, &lt;code&gt;manifest.json&lt;/code&gt;, and &lt;code&gt;styles.css&lt;/code&gt; from the &lt;a href="https://github.com/eristoddle/apple-books-annotation-import/releases" rel="noopener noreferrer"&gt;latest release&lt;/a&gt;, drop them in a folder under &lt;code&gt;.obsidian/plugins/apple-books-annotation-import/&lt;/code&gt; in your vault, and enable it in settings. No auto-updates, but it works.&lt;/p&gt;

&lt;p&gt;Either way, once it’s on, flip on “Import PDF highlights” in the settings, hit “Import all books,” and give it a minute to chew through your library.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Takeaway
&lt;/h2&gt;

&lt;p&gt;The lesson here isn’t really about PDFs. It’s that “it works” and “it’s complete” are two different claims, and I’d been quietly making the first one while believing the second for over a year. The plugin worked. It just wasn’t done. It took actually going looking for a specific missing highlight to find the gap.&lt;/p&gt;

&lt;p&gt;If you build tools for yourself, this is the failure mode to watch for. The thing runs, you trust it, and you stop checking whether it’s still telling you the whole truth. Apple gave me a good excuse by hiding PDF highlights in a completely different place than everything else, but the hole was mine to notice.&lt;/p&gt;

&lt;p&gt;Anyway. It’s fixed. Every highlight I’ve got, in every format Apple Books supports, now lands in Obsidian. Which means I’m out of excuses and back to actually reading the books.&lt;/p&gt;

</description>
      <category>obsidian</category>
      <category>applebooks</category>
      <category>plugin</category>
      <category>pdf</category>
    </item>
    <item>
      <title>Claude Opus 5: The Flagship Got Cheaper (There's a Catch)</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Tue, 28 Jul 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/claude-opus-5-the-flagship-got-cheaper-theres-a-catch-171i</link>
      <guid>https://dev.to/eristoddle/claude-opus-5-the-flagship-got-cheaper-theres-a-catch-171i</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8fhlgljv83v774v3vudo.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8fhlgljv83v774v3vudo.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For about a year now, every model launch has followed the same script. A lab drops a new flagship, it’s a little smarter than the last one, and it costs more. You brace for it. New tier, new price, same shrug. So when Anthropic put out Claude Opus 5 on July 24, I opened the pricing page already wincing.&lt;/p&gt;

&lt;p&gt;Then I read it twice, because it was half the price.&lt;/p&gt;

&lt;p&gt;Not half the price of some bloated competitor. Half the price of Anthropic’s own current king, Fable 5. And on the Artificial Analysis Intelligence Index, Opus 5 edged out Fable 5 for the number one spot. Smarter and cheaper, in the same launch. That’s not how any of this has gone for a year.&lt;/p&gt;

&lt;p&gt;And then, three days later, the biggest open-weights model in human history dropped on Hugging Face, and I sat there doing the math on whether I could run it. Spoiler: I cannot. Nobody with fewer than eight datacenter GPUs can. Welcome to the week.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Flagship Got Cheaper (Wait, What?)&lt;/li&gt;
&lt;li&gt;The 2.8 Trillion Parameter Paperweight&lt;/li&gt;
&lt;li&gt;The Actually Useful Release Nobody Tweeted About&lt;/li&gt;
&lt;li&gt;Cheapskate Picks: What I’d Actually Pay For&lt;/li&gt;
&lt;li&gt;The Horror Show&lt;/li&gt;
&lt;li&gt;The Still-Waiting Room&lt;/li&gt;
&lt;li&gt;The Takeaway&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Flagship Got Cheaper (Wait, What?)
&lt;/h2&gt;

&lt;p&gt;Here’s the part I keep re-reading. Claude Opus 5 costs $5 per million input tokens and $25 per million output. That’s the exact same price as Opus 4.8, the model it replaces. Meanwhile Fable 5, Anthropic’s big expensive brain, sits at $10 in and $50 out. So Opus 5 is literally half the sticker of the model it just beat on the intelligence leaderboard.&lt;/p&gt;

&lt;p&gt;On the Artificial Analysis Intelligence Index, Opus 5 (max) scores 61. Fable 5 lands at 60. GPT-5.6 Sol comes in at 59, and Moonshot’s Kimi K3 at 57. It’s a photo finish at the top, but the point is that the cheaper Anthropic model is the one holding the trophy this week.&lt;/p&gt;

&lt;p&gt;The number that actually matters, though, isn’t the Index rank. It’s cost per task. Sticker price per token lies to you, because different models burn different amounts of tokens to finish the same job. Artificial Analysis measured it: running their full Intelligence Index costs about $2.03 per task on Opus 5 versus $2.75 on Fable 5. That’s roughly 26 percent cheaper to do the same work, on top of the lower per-token rate. For anyone running agent loops where the bill compounds, that’s the real headline.&lt;/p&gt;

&lt;p&gt;The benchmarks back up the “it’s genuinely good” claim, not just the “it’s cheap” one. On Frontier-Bench v0.1 it hit 43.3 percent, against Fable 5’s 33.7 and GPT-5.6 Sol’s 34.4. On ARC-AGI-3 it scored 30.2 percent while Opus 4.8 managed a sad 1.5. It’s the new default on Claude Max and the strongest model available on Claude Pro.&lt;/p&gt;

&lt;p&gt;So what’s the catch, because there’s always a catch. The catch is buried in the launch chart. On the Frontier-Bench numbers, Anthropic notes that Opus 4.8 “stood in as a fallback” whenever a safety classifier refused an Opus 5 request. Fine. Except they never said how often that happened. So the flagship benchmark quietly folds in a weaker model’s answers by an amount nobody will tell you. If you remember the refuse-and-reroute mess that followed Fable 5 around, this is the same species of problem wearing a nicer suit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2.8 Trillion Parameter Paperweight
&lt;/h2&gt;

&lt;p&gt;While Anthropic was cutting prices, Moonshot was flexing. On July 27 they released the open weights for Kimi K3, and this thing is a monster: 2.8 trillion parameters, the largest open-weight model ever shipped. On the Artificial Analysis Index it scores 57, which makes it the highest-scoring open model on the board, ahead of everything else you can actually download.&lt;/p&gt;

&lt;p&gt;Here’s where the dream meets the parking lot. The download is about 1.4 terabytes of weights even at MXFP4 quantization. The architecture is a mixture of experts, 896 experts total with 16 active per token, so roughly 50 billion parameters are actually doing work on any given pass. To load it you need something like four to eight H100 80GB GPUs. Your 4090 can’t touch it. Your maxed-out Mac Studio can’t touch it. “Own your weights” is a beautiful slogan right up until you price the hardware to hold them.&lt;/p&gt;

&lt;p&gt;So in practice, for almost everyone, Kimi K3 is still just an API you rent at $3 per million in and $15 per million out. The weights being open is great for labs, cloud providers, and the three guys on Reddit with a GPU rack in the garage. For the rest of us it’s a philosophical victory, not a practical one.&lt;/p&gt;

&lt;p&gt;And it’s not a clean win even on quality. Accuracy went up about 13 points over the K2.6 generation, which sounds great, but the hallucination rate also climbed about 12 points. Moonshot frames that as the model being “more willing to answer.” Cute. For regulated, legal, medical, or financial work, “more willing to answer” is a polite way of saying “more confidently wrong more often,” and there’s no dial to trade it back.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Actually Useful Release Nobody Tweeted About
&lt;/h2&gt;

&lt;p&gt;Buried under the Opus 5 launch and the Kimi K3 spectacle, Google shipped the model I’d actually reach for on a Tuesday. Gemini 3.6 Flash landed July 21 at $1.50 in and $7.50 out. That output price is down from the $9 that Gemini 3.5 Flash charged, and Google says it uses about 17 percent fewer output tokens on top of that. Cheaper rate, fewer tokens, same 1 million token context, knowledge cutoff pushed to March 2026.&lt;/p&gt;

&lt;p&gt;The funny part: Artificial Analysis gives 3.6 Flash the same Intelligence Index score as 3.5 Flash, a 50. So the tech press mostly shrugged. No leap, no headline. But it gained on the benchmarks that matter for real work, SWE-Bench Pro up to 58.7 percent from 55.1, OSWorld computer use up to 83 from 78.4. And it sits inside the competitive band of five of the six Arena categories I track. Cheaper than the thing it replaces, and it beats the mid-tier of the pricier labs. That’s the whole pitch, and it’s a good one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheapskate Picks: What I’d Actually Pay For
&lt;/h2&gt;

&lt;p&gt;Here’s the trick I run every week, because I’m cheap and I have to be. The Arena leaderboards cluster tight at the top. The entire Overall top 20 this week fits inside 32 rating points, from Fable 5 at 1508 down to a pack at 1476. When the whole visible field is that compressed, paying the leader’s price buys you almost nothing over something a fraction of the cost. So the game is: find the cheapest model still inside spitting distance of the category leader.&lt;/p&gt;

&lt;p&gt;Here’s where that landed this week:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Leader&lt;/th&gt;
&lt;th&gt;$ out&lt;/th&gt;
&lt;th&gt;Cheapskate pick&lt;/th&gt;
&lt;th&gt;$ out&lt;/th&gt;
&lt;th&gt;Δ rating&lt;/th&gt;
&lt;th&gt;Cheaper by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Overall&lt;/td&gt;
&lt;td&gt;Fable 5 (1508)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (1465, #37)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;−43&lt;/td&gt;
&lt;td&gt;~57x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;Opus 4.7-thinking (1553)&lt;/td&gt;
&lt;td&gt;~$25&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (1519, #25)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;−34&lt;/td&gt;
&lt;td&gt;~29x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative Writing&lt;/td&gt;
&lt;td&gt;Fable 5 (1507)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;Gemini 3-Flash (1458, #21)&lt;/td&gt;
&lt;td&gt;$3&lt;/td&gt;
&lt;td&gt;−49&lt;/td&gt;
&lt;td&gt;~16.7x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction Following&lt;/td&gt;
&lt;td&gt;Fable 5 (1515)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (1469, #23)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;−46&lt;/td&gt;
&lt;td&gt;~57x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hard Prompts&lt;/td&gt;
&lt;td&gt;Fable 5 (1535)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;MiMo v2.5 Pro (1494, #25)&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;−41&lt;/td&gt;
&lt;td&gt;~57x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math&lt;/td&gt;
&lt;td&gt;Fable 5 (1539)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;Qwen3.7 Max (1490, #14)&lt;/td&gt;
&lt;td&gt;$4.42&lt;/td&gt;
&lt;td&gt;−49&lt;/td&gt;
&lt;td&gt;~11.3x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;MiMo v2.5 Pro, Xiaomi’s open-weights model, is the actual story this week, not Gemini 3.6 Flash. It sweeps four of six categories, Overall, Coding, Instruction Following, and Hard Prompts, all at the same $0.87 output price, and it does it by sitting at rank 23 to 37 in every one of those boards. A top-20 read misses it every single time, which is exactly the mistake I made drafting this section the first time around: I caught it on Coding, went back and checked the rest, and found the same miss in three more categories. Nobody’s talking about it. It’s just sitting there, MIT-licensed, being the actual answer.&lt;/p&gt;

&lt;p&gt;Gemini 3.6 Flash, the model with a whole section above this one, doesn’t win a single category outright once you look past the top 20. It’s still inside the band in five of six categories, so it’s a perfectly good pick if you want something you don’t have to go hunting for or self-host, but the honest cheapest option in four of six categories is a phone company’s open-weights model most readers have never heard of.&lt;/p&gt;

&lt;p&gt;Creative Writing breaks the pattern in both directions. MiMo doesn’t even show up in this category’s band, because Arena’s Creative Writing top end closes off faster (around rank 22 here) than the other categories do. And the actual cheapest thing inside that narrower band isn’t Gemini 3.6 Flash either, it’s the older, cheaper Gemini 3-Flash at $3, forty-nine points back and sitting right at the edge of the cutoff.&lt;/p&gt;

&lt;p&gt;Math is the one category where nothing changed. Qwen3.7 Max is still the cheapest model in the band, and it’s still clinging to the very edge at 49 points back. Grok 4.5 and Gemini 3.6 Flash both sit closer to the leader for a bit more money if you’d rather not ride the edge.&lt;/p&gt;

&lt;p&gt;One caveat on Overall and Hard Prompts: I pulled 40 rows deep for each category and the band technically hadn’t closed yet at row 40 in those two (the cutoff-adjacent rows were still inside the window). MiMo’s $0.87 is close to the price floor this cycle, so it’s very unlikely anything cheaper is sitting a few rows further down, but “very unlikely” isn’t the same as “confirmed,” so treat those two picks as high-confidence rather than fully closed.&lt;/p&gt;

&lt;p&gt;If you’re keeping score: MiMo v2.5 Pro is the boring, correct answer this week, not Gemini 3.6 Flash. Four out of six categories, one price tag, and it took reading past rank 20 in every single one of them to find that out.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Horror Show
&lt;/h2&gt;

&lt;p&gt;Every roundup needs a section where I tell you what broke. This week was generous.&lt;/p&gt;

&lt;p&gt;The big one was DeepSeek’s migration cliff. DeepSeek is the single most-used vendor on OpenRouter right now, about 17.6 percent of all routed tokens. On July 24 at 15:59 UTC they hard-retired the old model names &lt;code&gt;deepseek-chat&lt;/code&gt; and &lt;code&gt;deepseek-reasoner&lt;/code&gt;. Not deprecated with a grace period. Retired. Calls to those names now return errors with no fallback, so any service that didn’t repoint to &lt;code&gt;deepseek-v4-flash&lt;/code&gt; or &lt;code&gt;deepseek-v4-pro&lt;/code&gt; started throwing user-facing failures the moment the clock hit. Worse, reasoning moved from being a model name to being a request parameter, so lazy integrations silently lost their thinking mode before the hard cutoff even arrived. If your app went weird last Friday afternoon, there’s your answer.&lt;/p&gt;

&lt;p&gt;Then there’s the Opus 5 fallback I already whined about. A benchmark that quietly swaps in a different model when the safety filter trips, by an amount nobody discloses, is exactly the kind of asterisk that gets left out of the headline.&lt;/p&gt;

&lt;p&gt;And Kimi K3 has a confidence problem. The hallucination rate climbing 12 points while the marketing calls it “more willing to answer” is going to bite somebody who wired it into a pipeline that trusts its output. Retrieval checks and citation verification are not optional with this one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Still-Waiting Room
&lt;/h2&gt;

&lt;p&gt;The upcoming section is short and it’s mostly one name. Gemini 3.5 Pro missed its target again, and I’ve lost count, but this is at least the fourth slip. It’s still stuck in Vertex AI enterprise preview. Instead of the Pro model everyone actually wanted, Google shipped a fistful of Flash variants on July 21, which is how we got 3.6 Flash. TechCrunch’s headline said it plainly: three new Gemini models, but no 3.5 Pro. Reporting says the model keeps failing to hit Google’s own internal performance bar. At some point “delayed” starts to read as “in trouble.”&lt;/p&gt;

&lt;p&gt;On the open side, expect the Kimi K3 community quants and finetunes to start rolling now that the weights are public, assuming you have the hardware to do anything with them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Takeaway
&lt;/h2&gt;

&lt;p&gt;For most of the past year, the story in these roundups was rent versus own. Big closed models you pay for by the token, versus open Chinese weights you could theoretically host yourself. This week broke that frame in both directions at once. Anthropic made the closed flagship cheaper than its predecessor, and Moonshot made the open flagship so enormous that owning it is meaningless unless you run a datacenter.&lt;/p&gt;

&lt;p&gt;So the real question quietly changed. It’s not rent versus own anymore. It’s how much intelligence can you actually afford to run. Opus 5 answers it one way, by dropping the price of the top shelf. Gemini 3.6 Flash answers it another, by being 85 percent cheaper and good enough. Kimi K3 answers it by being technically free and practically out of reach.&lt;/p&gt;

&lt;p&gt;For the record, the market keeps drifting east while all this happens. Chinese models hit a record 58 percent of tokens processed by US firms on OpenRouter this month, peaking around 63 percent earlier in July. DeepSeek alone is that 17.6 percent, Qwen another 13.9, and Anthropic is the last US lab standing in the top 10. Make of that what you will.&lt;/p&gt;

&lt;p&gt;Me, I’m going to keep running Gemini 3.6 Flash for the boring stuff and paying up for Opus 5 when the task actually needs a brain. And I’m going to keep not running Kimi K3, because I do not, in fact, own a rack of H100s. Maybe next week.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>openrouter</category>
    </item>
    <item>
      <title>Obsidian iCloud Sync in 2026, Including the Windows Problem</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Mon, 27 Jul 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/obsidian-icloud-sync-in-2026-including-the-windows-problem-4503</link>
      <guid>https://dev.to/eristoddle/obsidian-icloud-sync-in-2026-including-the-windows-problem-4503</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftkeljomnda7xalrgtpf4.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftkeljomnda7xalrgtpf4.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I sync Obsidian to two MacBook Pros, a Windows desktop, an iPad, and an Android phone, and I pay nothing for it. iCloud is not how I do it. I wrote up every method I know of in &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-across-all-your-devices-including-free-methods-1mi5"&gt;my full guide to syncing an Obsidian vault across devices&lt;/a&gt;, and iCloud got a short section in that post with a warning attached, because the moment an Android phone enters the picture iCloud is done.&lt;/p&gt;

&lt;p&gt;But that pillar post never had room to explain the part people actually get burned by. So this is the long version. How to set iCloud sync up properly, what the Apple-only happy path looks like, and the specific ways iCloud Drive on Windows starts leaving random markdown files around your vault.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Who Obsidian iCloud Sync Is Actually For&lt;/li&gt;
&lt;li&gt;How to Create a New Obsidian Vault in iCloud&lt;/li&gt;
&lt;li&gt;How to Move an Existing Obsidian Vault to iCloud&lt;/li&gt;
&lt;li&gt;The Mac, iPhone, and iPad Happy Path&lt;/li&gt;
&lt;li&gt;The Setting That Quietly Destroys Your Vault&lt;/li&gt;
&lt;li&gt;
The Windows Problem

&lt;ul&gt;
&lt;li&gt;The Workaround That Actually Respects the Problem&lt;/li&gt;
&lt;li&gt;When To Just Not Do It&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Obsidian Sync vs iCloud&lt;/li&gt;
&lt;li&gt;Is Obsidian iCloud Sync Encrypted?&lt;/li&gt;
&lt;li&gt;Cleaning Up the Duplicates You Already Have&lt;/li&gt;
&lt;li&gt;Frequently Asked Questions&lt;/li&gt;
&lt;li&gt;The Verdict&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Who Obsidian iCloud Sync Is Actually For
&lt;/h2&gt;

&lt;p&gt;iCloud is the best free option in exactly one situation: every device you touch has an Apple logo on it.&lt;/p&gt;

&lt;p&gt;Mac plus iPhone plus iPad is close to zero configuration. Obsidian on iOS has a first-class iCloud option built into the vault creation screen, which is more than you can say for Dropbox or Google Drive on that platform.&lt;/p&gt;

&lt;p&gt;Everywhere else it gets worse. Windows is possible but unreliable, covered in detail below. Linux has no official iCloud Drive client at all. And Android is not possible in any way I would recommend to a human being, because there is no iCloud Drive client for Android that syncs a folder to local storage, and Obsidian on Android needs a real local folder. That is why my own vault lives in Dropbox with &lt;a href="https://github.com/remotely-save/remotely-save" rel="noopener noreferrer"&gt;Remotely Save&lt;/a&gt; handling Android. My phone is the constraint that decided the whole architecture.&lt;/p&gt;

&lt;p&gt;If your phone is an iPhone and your desktop is a Mac, stop reading comparison posts and just use iCloud. It is free, it is already on, and it works.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Create a New Obsidian Vault in iCloud
&lt;/h2&gt;

&lt;p&gt;Two completely different procedures get searched for here, and people mix them up constantly. Creating a fresh vault in iCloud is the easy one. Do this first if you are starting clean.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;On iPhone or iPad:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open Obsidian and tap &lt;strong&gt;Create new vault&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Give it a name. Make it something you will recognize in Files.&lt;/li&gt;
&lt;li&gt;Turn on &lt;strong&gt;Store in iCloud&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Tap &lt;strong&gt;Create&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;On Mac:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open Obsidian and choose &lt;strong&gt;Create new vault&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;For the location, browse to &lt;code&gt;iCloud Drive&lt;/code&gt; and pick or make a folder there.&lt;/li&gt;
&lt;li&gt;Create the vault.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Figi8q6x8xuk3evmrndsg.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Figi8q6x8xuk3evmrndsg.jpg" alt="How to Create a New Obsidian Vault in iCloud" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On disk, the Mac path is &lt;code&gt;~/Library/Mobile Documents/iCloud~md~obsidian/Documents/&lt;/code&gt;, which shows up in the Files app as &lt;strong&gt;iCloud Drive &amp;gt; Obsidian&lt;/strong&gt;. That is where the iOS app puts vaults, and it is worth knowing because it is the folder you will be pointing other tools at later. Obsidian’s own docs are direct about this: vaults should live inside the Obsidian folder in iCloud Drive. A vault you drop somewhere else in iCloud Drive will sync between Macs fine and then fail to show up in the iOS app’s vault list, which is a confusing hour of your life you can skip.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Move an Existing Obsidian Vault to iCloud
&lt;/h2&gt;

&lt;p&gt;This is the one that goes wrong. The instinct is to drag your vault folder into iCloud Drive and open it. Do not start there.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Back the vault up first.&lt;/strong&gt; Copy the entire folder somewhere outside any sync service. A zip on your desktop is fine, and it is the only thing standing between you and a bad afternoon.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Close Obsidian on every device.&lt;/strong&gt; Not backgrounded on your phone. Closed. Two clients writing into a folder during an initial upload is how you generate conflicts on day one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create a new empty vault in iCloud&lt;/strong&gt; using the steps above, with the same name as your existing vault.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Move the contents&lt;/strong&gt; of your old vault into that new iCloud folder. Include the hidden &lt;code&gt;.obsidian&lt;/code&gt; folder if you want your settings, themes, hotkeys, and plugins to come along. On Mac, &lt;code&gt;Cmd + Shift + .&lt;/code&gt; toggles hidden files in Finder.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wait for the upload to finish completely&lt;/strong&gt; before opening Obsidian anywhere. Watch the iCloud status in Finder’s sidebar. Rushing this is the whole problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open it on your Mac first&lt;/strong&gt; , let plugins load, then open it on iOS.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;On a big vault, leave the attachments out of the first pass. Sync the markdown, confirm it works, then move the images. Markdown files are tiny. Attachments are what blow up the initial sync window and give conflicts room to happen.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Mac, iPhone, and iPad Happy Path
&lt;/h2&gt;

&lt;p&gt;Once the vault is in place, the Apple-only setup genuinely is the low-maintenance option. Edits show up on the other device in a few seconds when both are awake and online. Obsidian’s own settings sync along with the vault because &lt;code&gt;.obsidian&lt;/code&gt; is just another folder, so if you have gone deep on &lt;a href="https://dev.to/eristoddle/how-to-install-activate-and-update-obsidian-plugins-4d2p"&gt;installing Obsidian plugins&lt;/a&gt; you do not have to set them up again per device.&lt;/p&gt;

&lt;p&gt;Two habits keep it that way. Let one device finish syncing before you start editing on another, because sync services do not have opinions about which version of a paragraph you meant. And never run a second sync system on the same vault. Not Obsidian Sync on top of iCloud, not Dropbox pointed at the same folder, not a Git plugin committing every two minutes. Two sync engines fighting over one folder is the most reliable way to lose notes, and it is self-inflicted every time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setting That Quietly Destroys Your Vault
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fee3l7m23o773e0ip1u7y.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fee3l7m23o773e0ip1u7y.jpg" alt="The Setting That Quietly Eats Your Vault" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Turn off &lt;strong&gt;Optimize Mac Storage&lt;/strong&gt; for iCloud Drive, or at least understand what it does before you leave it on.&lt;/p&gt;

&lt;p&gt;When macOS decides you are low on space, it evicts the local copies of files it thinks you are not using and leaves a placeholder with a little cloud icon. For photos, fine. For an Obsidian vault, not fine. Search, graph view, and backlinks all depend on files actually being present on disk. An evicted note is not a note. Searches come back short, links look broken, and nothing warns you that half your vault is currently a stub.&lt;/p&gt;

&lt;p&gt;The setting is in &lt;strong&gt;System Settings &amp;gt; [your name] &amp;gt; iCloud &amp;gt; iCloud Drive&lt;/strong&gt;. If you need it on for storage reasons, open iCloud Drive in Finder, control-click the vault folder, and choose &lt;strong&gt;Keep Downloaded&lt;/strong&gt; to exempt it. Windows has the same concept with a different name, and it is worse there, which is a good segue.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Windows Problem
&lt;/h2&gt;

&lt;p&gt;Obsidian’s own documentation says flatly that iCloud Drive on Windows may lead to file duplication or corruption. When the people who make the app warn you off a sync method, that is worth more than any blog post, including this one.&lt;/p&gt;

&lt;p&gt;iCloud for Windows exists. You install it from the Microsoft Store, sign in, and get an iCloud Drive folder in File Explorer. You can point Obsidian at a vault inside it. It will appear to work. Then, over days and weeks, these things start happening.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Placeholder files instead of real files.&lt;/strong&gt; iCloud for Windows does on-demand downloads the same way OneDrive does. A file with a cloud icon next to it is not on your disk. Obsidian tries to read it, and depending on timing you get an empty note, a failed read, or a plugin that throws. Vault-wide operations like search and Dataview queries are the worst hit because they touch everything at once. You can fight this by right-clicking the vault folder and choosing &lt;strong&gt;Always keep on this device&lt;/strong&gt; , which pins it locally. Do that before anything else if you are going to attempt this at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Duplicate files with numbers appended.&lt;/strong&gt; This is the classic symptom and the reason people search for this problem. You end up with &lt;code&gt;Meeting Notes.md&lt;/code&gt; and &lt;code&gt;Meeting Notes 2.md&lt;/code&gt;, sometimes several generations deep. It happens when the Windows client and another device both write a file before either has seen the other’s version. iCloud does not merge and it does not prompt. It keeps both and renames one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjhokopjicb8w02uc45o0.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjhokopjicb8w02uc45o0.jpg" alt="The Windows Problem" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The &lt;code&gt;.obsidian&lt;/code&gt; folder taking damage.&lt;/strong&gt; Your config folder is a pile of small JSON files that get rewritten constantly as you use the app. That write pattern is exactly what a lazy sync client handles worst. Corrupted &lt;code&gt;workspace.json&lt;/code&gt;, plugin settings reverting, hotkeys resetting, community plugins disabling themselves. If your Windows machine keeps forgetting your setup, this is why.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sync that just stops.&lt;/strong&gt; The client parks in a pending state and stays there. No error, no notification, just a folder that stopped updating while you kept typing into it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Timing that encourages all of the above.&lt;/strong&gt; iCloud on Windows is slower to propagate changes than it is between Apple devices. A longer window between “I saved” and “the other machine knows” is a bigger window for conflicts.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Workaround That Actually Respects the Problem
&lt;/h3&gt;

&lt;p&gt;The fix that works is to stop letting Obsidian and iCloud touch the same folder.&lt;/p&gt;

&lt;p&gt;Keep your working vault in a plain local folder on the Windows machine, somewhere iCloud cannot see. Then run a separate process that syncs that local folder to the iCloud copy, with real conflict handling. Obsidian only ever talks to fast local disk, and the sync layer deals with iCloud’s nonsense on its own schedule.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/gursimar/obsidian-icloud-windows-sync" rel="noopener noreferrer"&gt;gursimar/obsidian-icloud-windows-sync&lt;/a&gt; does exactly this. It is a Python three-way sync engine that tracks the local vault, the iCloud copy, and a history snapshot so it can tell the difference between “this file changed here” and “this file changed on the other side.” When both changed, it keeps the newer one by modification time and saves the loser as a &lt;code&gt;_CONFLICT_&amp;lt;timestamp&amp;gt;&lt;/code&gt; file rather than silently picking a winner. It uses atomic writes and a stabilization delay so it is not reacting to Obsidian’s autosave mid-keystroke. The README is explicit that it must run natively on Windows and not under WSL, because iCloud placeholder files behave incorrectly when accessed through WSL.&lt;/p&gt;

&lt;p&gt;That is a real answer, but be honest with yourself about what it is. It is a Python script you have to configure with a YAML file, keep running, and troubleshoot when it stops. If that sounds like a project rather than a solution, it probably is one for you.&lt;/p&gt;

&lt;h3&gt;
  
  
  When To Just Not Do It
&lt;/h3&gt;

&lt;p&gt;If your setup is Windows plus iPhone, iCloud is the wrong tool. You are picking the option that is worst on your primary computer for the sake of convenience on your phone.&lt;/p&gt;

&lt;p&gt;Dropbox with Remotely Save covers Windows and iOS without any of this. So does Syncthing if you want nothing in the cloud at all. Obsidian Sync costs money and handles it. Any of those is a better use of your evening than fighting a sync client that is not designed for the write pattern of a notes app.&lt;/p&gt;

&lt;p&gt;I keep a Windows desktop in my rotation. I have never once been tempted to put my vault in iCloud on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Obsidian Sync vs iCloud
&lt;/h2&gt;

&lt;p&gt;The honest comparison, since “obsidian sync vs icloud” is what a lot of people are really asking.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1lvbzwuk9uafd6clt1m1.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1lvbzwuk9uafd6clt1m1.jpg" alt="Obsidian Sync vs iCloud" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;iCloud&lt;/th&gt;
&lt;th&gt;Obsidian Sync&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Price&lt;/td&gt;
&lt;td&gt;Free with your existing iCloud storage&lt;/td&gt;
&lt;td&gt;Standard $5/mo, or $4/mo billed annually. Plus $10/mo, or $8/mo annually&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;Shares your iCloud quota&lt;/td&gt;
&lt;td&gt;Standard 1 GB. Plus 10 GB, upgradable to 100 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;File size cap&lt;/td&gt;
&lt;td&gt;iCloud Drive limits&lt;/td&gt;
&lt;td&gt;Standard 5 MB. Plus 200 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Version history&lt;/td&gt;
&lt;td&gt;Whatever iCloud keeps, not note-aware&lt;/td&gt;
&lt;td&gt;Standard 1 month. Plus 12 months&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows&lt;/td&gt;
&lt;td&gt;Unreliable, see above&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Android&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Linux&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Encryption&lt;/td&gt;
&lt;td&gt;In transit and at rest, Apple holds the keys unless Advanced Data Protection is on&lt;/td&gt;
&lt;td&gt;End-to-end by default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conflict handling&lt;/td&gt;
&lt;td&gt;Duplicate files&lt;/td&gt;
&lt;td&gt;Merges, with per-file history to recover from&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Four dollars a month is the price of not reading this article. If you use Obsidian daily and your devices are not all Apple, that is a rounding error against the time any free method will cost you. I do not pay it, but I also enjoy this kind of problem, which is not a normal thing to enjoy.&lt;/p&gt;

&lt;p&gt;The version history row is the one people undervalue. iCloud can restore some files some of the time. Obsidian Sync keeps per-note history that understands what a vault is. The first time you need it, that gap is enormous.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is Obsidian iCloud Sync Encrypted?
&lt;/h2&gt;

&lt;p&gt;Yes, but probably not in the way you are assuming.&lt;/p&gt;

&lt;p&gt;iCloud Drive is encrypted in transit and encrypted on Apple’s servers by default, and Apple holds the keys. That means Apple can access the contents and can be compelled to hand them over. Standard data protection is the default on every account.&lt;/p&gt;

&lt;p&gt;Turning on &lt;strong&gt;Advanced Data Protection&lt;/strong&gt; moves iCloud Drive to end-to-end encryption, so only your devices hold the keys. It is in &lt;strong&gt;System Settings &amp;gt; [your name] &amp;gt; iCloud &amp;gt; Advanced Data Protection&lt;/strong&gt; on a Mac, and the same path under Settings on iOS. The tradeoff is real. Apple makes you set up at least one alternative recovery method first, either a recovery contact or a recovery key, because once it is on Apple does not have the keys to help you. Lose your recovery methods and the data is gone. That is the entire point of it.&lt;/p&gt;

&lt;p&gt;Obsidian Sync is end-to-end encrypted by default with no configuration. If your vault holds anything you would call sensitive, that difference matters more than the price difference does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cleaning Up the Duplicates You Already Have
&lt;/h2&gt;

&lt;p&gt;If you got here after the fact, the vault is already littered with numbered copies. Work in this order.&lt;/p&gt;

&lt;p&gt;Close Obsidian everywhere, let iCloud finish syncing, and copy the whole vault out to somewhere iCloud does not control. Then find the offenders. On Mac, from the vault root:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;find &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"* 2.md"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"* 3.md"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*conflicted copy*"&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0gkgzjc9m07dn8teyaca.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0gkgzjc9m07dn8teyaca.jpg" alt="Cleaning Up the Duplicates You Already Have" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On Windows, in PowerShell:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;Get-ChildItem&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Recurse&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Filter&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"* 2.md"&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Diff before you delete. Most numbered copies are identical to the original and safe to remove. Some contain the only copy of a paragraph you wrote:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gh"&gt;diff "Meeting Notes.md" "Meeting Notes 2.md"
&lt;/span&gt;&lt;span class="err"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Merge what matters, delete the rest, fix any internal links that pointed at the duplicate, then fix the cause or you will do this again next month. Fixing the cause means taking the Windows machine out of the iCloud path, or moving the vault to a method that handles conflicts properly.&lt;/p&gt;

&lt;p&gt;You can also automate the finding part. Obsidian Cleaner, one of &lt;a href="https://dev.to/eristoddle/the-obsidian-plugin-collection-i-built-one-free-kiro-credit-at-a-time-2b42"&gt;the plugins I had Kiro build for me on free monthly credits&lt;/a&gt;, surfaces conflicted copies, numbered duplicates, and zero-byte markdown files as a checklist you can review before deleting. I built it because I kept hitting this, which tells you how common the problem is across every sync method, not just iCloud.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;How do I sync Obsidian with iCloud?&lt;/strong&gt; On iOS, enable &lt;strong&gt;Store in iCloud&lt;/strong&gt; when you create the vault. On Mac, put the vault in &lt;code&gt;~/Library/Mobile Documents/iCloud~md~obsidian/Documents/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I connect Obsidian to iCloud on a vault I already have?&lt;/strong&gt; Close Obsidian everywhere, create an empty iCloud vault with the same name, and move your files plus the hidden &lt;code&gt;.obsidian&lt;/code&gt; folder into it. The order matters more than the steps do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I move my Obsidian vault to iCloud without losing anything?&lt;/strong&gt; Back up first, and wait for the initial upload to finish before you open the vault on a second device. On a large vault, sync the markdown first and the attachments second.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Obsidian iCloud sync work on Windows?&lt;/strong&gt; It runs, but it is unreliable. If you must do it, keep the working vault in a plain local folder and mirror it to iCloud with a tool like &lt;code&gt;gursimar/obsidian-icloud-windows-sync&lt;/code&gt; rather than letting Obsidian edit inside the iCloud folder.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Obsidian iCloud sync encrypted?&lt;/strong&gt; Encrypted in transit and at rest, but Apple holds the keys unless Advanced Data Protection is on. Obsidian Sync is end-to-end encrypted by default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I sync Obsidian to Android with iCloud?&lt;/strong&gt; No. Use Dropbox with Remotely Save, Syncthing, or Obsidian Sync instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Verdict
&lt;/h2&gt;

&lt;p&gt;iCloud sync for Obsidian is genuinely good and genuinely free, right up to the edge of Apple’s ecosystem, where it falls off a cliff with no railing.&lt;/p&gt;

&lt;p&gt;All Apple, all the time: use it, it is the right answer and you can stop researching. Windows in the mix: it is a maintenance project, and you should either run a real sync layer on top of it or pick a different method entirely. Android anywhere in your life: it is not a choice you have.&lt;/p&gt;

&lt;p&gt;I landed on Dropbox and Remotely Save because of an Android phone, and I have no regrets about it, though I did spend a weekend getting there. The methods are not ranked by quality. They are ranked by which devices you happen to own, and the honest advice is to pick based on your worst device rather than your best one.&lt;/p&gt;

</description>
      <category>obsidian</category>
      <category>icloud</category>
      <category>sync</category>
      <category>windows</category>
    </item>
    <item>
      <title>The Bottleneck Was Me: How I Stopped Racing My AI Builder and Started Pacing It</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Wed, 22 Jul 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/the-bottleneck-was-me-how-i-stopped-racing-my-ai-builder-and-started-pacing-it-3e1n</link>
      <guid>https://dev.to/eristoddle/the-bottleneck-was-me-how-i-stopped-racing-my-ai-builder-and-started-pacing-it-3e1n</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqpx2xeklxabbgxlzjmq5.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqpx2xeklxabbgxlzjmq5.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The last post ended on a clean little problem. I’d gotten my &lt;code&gt;PLAN.md&lt;/code&gt; lean: the decisions up top, the history filed away, a skill that keeps it that way so I don’t have to. The plan was finally generating work faster than I could build it. And that exposed a new bottleneck on the build side: my setup hands one task at a time to a background Sonnet agent and waits. The builder finishes, then sits there idle while I plan the next one. So I asked the next question: how do you keep the builder &lt;em&gt;always busy?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The goal turned out to be a decent description of the symptom and a terrible description of what I actually wanted. This is the post where I figured out the difference, by doing it backwards first.&lt;/p&gt;

&lt;p&gt;This is part three of a series about working with AI coding agents using a single living &lt;code&gt;PLAN.md&lt;/code&gt; instead of vibe coding or spec-kit ceremony. &lt;a href="https://dev.to/eristoddle/my-third-try-how-a-living-plan-beat-both-vibe-coding-and-spec-kit-5a89"&gt;Post one&lt;/a&gt; is the thesis. The doc is the deliverable, the code is the byproduct. &lt;a href="https://dev.to/eristoddle/the-living-plan-got-fat-compacting-a-doc-that-wont-stop-growing-3nk0"&gt;The last post&lt;/a&gt; was about keeping that doc from turning into a 28,000-word novella.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Batch That Hooked Me&lt;/li&gt;
&lt;li&gt;The Experiment: A Pre-Specified Queue&lt;/li&gt;
&lt;li&gt;I Optimized the Wrong Thing&lt;/li&gt;
&lt;li&gt;The Gap Was the Point&lt;/li&gt;
&lt;li&gt;The Model: A Serial Queue&lt;/li&gt;
&lt;li&gt;Two Ways to Stack&lt;/li&gt;
&lt;li&gt;The Real Bottleneck Was Never the Task Slot&lt;/li&gt;
&lt;li&gt;The Name&lt;/li&gt;
&lt;li&gt;The Honest Scope Note&lt;/li&gt;
&lt;li&gt;What’s Next&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Batch That Hooked Me
&lt;/h2&gt;

&lt;p&gt;The seed of this whole idea was a session a few weeks back where I did something different. Instead of speccing one hardening task, launching it, and waiting, I lined up a &lt;em&gt;batch&lt;/em&gt; of small mechanical jobs and drove them back-to-back as one verified sequence. Each one: Sonnet builds it, I re-read the diff and the tests, then I launch the next. The test count climbed 140 → 234 across the run.&lt;/p&gt;

&lt;p&gt;Driving them as a batch freed up enough of my attention that I was &lt;em&gt;working a second project at the same time&lt;/em&gt;, jumping back and forth between this content pipeline and my other repo. The fact that the builder grinding away in the background bought me time to go be useful somewhere else.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fty91vdnx1mifarngqcof.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fty91vdnx1mifarngqcof.jpg" alt="The Batch That Hooked Me" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I filed that away as a direction and went back to feeding the plan. But it nagged at me, because I’d named it wrong in my own head. I called it “keeping the builder busy,” and I built the next experiment around that name.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Experiment: A Pre-Specified Queue
&lt;/h2&gt;

&lt;p&gt;Here’s how I decided to actually test it. I had a real feature coming up: making the content pipeline able to &lt;em&gt;require&lt;/em&gt; embedded artifacts in an article. A brief could say “this post needs at least two of: a code block, a comparison table, a diagram,” and the pipeline would produce them, preserve them through the humanizer, and verify they’re actually there at the end. Four moving parts, real dependencies between them.&lt;/p&gt;

&lt;p&gt;So instead of speccing it one task at a time the way I always had, I pre-specified the &lt;em&gt;entire arc&lt;/em&gt; up front as an ordered queue. Four tasks. Task A, the schema that declares the requirement, had to land first, because the other three all depend on it. Then B (the writer produces the artifacts), C (the humanizer preserves them), and D (the checker verifies them) could each go independently. I wrote all four specs, mapped the dependency, and flagged the genuine design forks, the questions only I could answer, to settle before anything ran.&lt;/p&gt;

&lt;p&gt;The point of pre-specifying the whole thing wasn’t speed. It was to find out &lt;em&gt;where pre-specifying breaks down.&lt;/em&gt; Which specs would look airtight on paper and fall apart the moment a builder touched real code.&lt;/p&gt;

&lt;p&gt;Then I ran it. And immediately optimized the wrong variable.&lt;/p&gt;

&lt;h2&gt;
  
  
  I Optimized the Wrong Thing
&lt;/h2&gt;

&lt;p&gt;Task A landed first. It had to; everything blocked on it. But B, C, and D were independent. So Claude decided to run them concurrently.&lt;/p&gt;

&lt;p&gt;Three Sonnet agents, in parallel, each on its own task, each on a disjoint set of files. And it &lt;strong&gt;worked&lt;/strong&gt;. All three landed clean. The test suite went from 234 to 328. If you’d asked me to measure throughput, I’d just tripled it. By every number I’d have put on a dashboard, this was the win.&lt;/p&gt;

&lt;p&gt;It felt awful.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6qivqvg4ixadw4r5xb55.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6qivqvg4ixadw4r5xb55.jpg" alt="I Optimized the Wrong Thing" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Because the second those three agents came back, they came back &lt;em&gt;together&lt;/em&gt;. The background hum that let me work on the other project and do slow human thinking. Gone. I’d taken a process whose entire value was &lt;em&gt;spreading work out over time&lt;/em&gt; and I’d compressed it back into a single spike of “everything needs you, now.”&lt;/p&gt;

&lt;p&gt;I’d been accidentally optimizing for &lt;strong&gt;throughput&lt;/strong&gt; : get the most build done per unit time. But throughput was never what bought me the second project. What bought me the second project was the &lt;em&gt;gap.&lt;/em&gt; Parallelism doesn’t widen that gap. It removes it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Gap Was the Point
&lt;/h2&gt;

&lt;p&gt;The builder’s grind is not a cost to be minimized. &lt;strong&gt;It’s a resource.&lt;/strong&gt; A long-enough synchronous build is a &lt;em&gt;gap&lt;/em&gt; I spend doing the slow human work: investigating, answering my own open questions, and making the design calls that only I can make. Often on a completely different project. The ideal loop isn’t “builder always busy.” It’s two slow things overlapping in time: the machine grinding through mechanical work while I grind through judgment work. Pacing, not racing.&lt;/p&gt;

&lt;p&gt;Once I saw it that way, the design fell out immediately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Synchronous, one-task-at-a-time, is the default&lt;/strong&gt; , because that’s what manufactures the gap. One builder, working a queue in order, then stopping to report. The stop is a feature. It’s the handoff point where I come back, review, and re-aim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parallel is a rare opt-in burst,&lt;/strong&gt; for the specific case where I’m actively babysitting &lt;em&gt;this&lt;/em&gt; project and want raw speed more than I want freed attention. But it’s not a default. The session that taught me this lesson proved parallel works. It just also proved it’s the wrong thing to reach for ninety percent of the time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9vc9hklbxs13cbz0ud9b.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9vc9hklbxs13cbz0ud9b.jpg" alt="The Gap Was the Point" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Model: A Serial Queue
&lt;/h2&gt;

&lt;p&gt;So the task handoff file got reshaped. It used to be &lt;code&gt;TASK.md&lt;/code&gt;: singular, one task, overwrite it for the next one. Now it’s &lt;code&gt;TASKS.md&lt;/code&gt;, and the active task is a &lt;strong&gt;serial queue of numbered pieces.&lt;/strong&gt; The builder works them top to bottom, in one pass, then stops and reports. Finished work doesn’t get overwritten and lost. It collapses into a &lt;code&gt;✅ Done&lt;/code&gt; archive with a stable tag, so the file carries its own history.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two Ways to Stack
&lt;/h2&gt;

&lt;p&gt;Because once the active task is a queue you can stack as deep as you want, two genuinely different working modes fall out of the exact same mechanism:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pacing.&lt;/strong&gt; Stack a couple of pieces. The builder grinds them while I plan the next batch a step ahead. By the time it stops, I’ve got the next chunk ready to go. This is the everyday rhythm, the overlap of machine-grind and human-thinking I’ve been describing the whole post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unattended.&lt;/strong&gt; Stack a &lt;em&gt;lot&lt;/em&gt; of pieces (an hour’s worth) and walk away entirely. Go do something else, something not-this-project, and come back to a pile of finished, tested work and one report. Same serial builder, same queue. The only difference is how deep I stack it and whether I’m in the room.&lt;/p&gt;

&lt;p&gt;They’re mechanically identical, but the unattended mode forced two rules that the pacing mode never exposed, because pacing rarely runs more than a piece or two before I’m back:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A blocked piece must not halt the queue.&lt;/strong&gt; If the builder hits piece 3 of 8 and it turns out underspecified (some fork I didn’t see), the old rule was “stop and report.” Fine when I’m sitting right there. Catastrophic when I’m gone for an hour, because pieces 4 through 8 never run even if they don’t depend on the broken one. So the rule is now: mark the blocked piece, &lt;em&gt;skip it,&lt;/em&gt; and keep going with anything that doesn’t depend on it. Halt only when nothing left can proceed. An hour away should come back with six of eight done, not two.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flp3p4nyl64u2jdcrw3fu.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flp3p4nyl64u2jdcrw3fu.jpg" alt="Two Ways to Stack" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;State gets logged on every stop.&lt;/strong&gt; This is the one that matters most for walking away. Whenever the builder stops, finished or blocked, it writes its state directly into &lt;code&gt;TASKS.md&lt;/code&gt;: which pieces are done, which are blocked and why, what’s left, where to resume. It flips a status box on each piece as it goes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Design — pieces&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; [x] 1. Schema: declare required_artifacts on the brief
&lt;span class="p"&gt;-&lt;/span&gt; [x] 2. Writer produces the artifacts
&lt;span class="p"&gt;-&lt;/span&gt; [!] 3. Humanizer preserves them — BLOCKED: heading-protection
       order is ambiguous, needs a call
&lt;span class="p"&gt;-&lt;/span&gt; [x] 4. Checker verifies presence (depends on #1)

&lt;span class="gu"&gt;### ▶ Run state&lt;/span&gt;
3 of 4 done. #3 blocked on a design fork (see above). Resume there.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why in the task file and not somewhere clever? Because that file is already what the builder is working in, it survives the session getting killed, and if I come back tomorrow having completely forgotten where I left off, the file tells me. No memory required, mine or the agent’s. The thing that makes unattended runs safe to walk away from is that every piece keeps the test suite green, so “done” actually means done, and the state log means “interrupted” never means “lost.”&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Bottleneck Was Never the Task Slot
&lt;/h2&gt;

&lt;p&gt;I’d assumed the constraint on keeping a builder fed was the single-task slot: one &lt;code&gt;TASK.md&lt;/code&gt;, one job at a time, obviously that’s the chokepoint. It isn’t. &lt;strong&gt;The real bottleneck is spec throughput.&lt;/strong&gt; The builder only stays fed if there’s a backlog of &lt;em&gt;fully-specified, mechanical&lt;/em&gt; tasks waiting for it. The moment a task hits a real design fork, a genuine “should it do A or B” that only I can answer, the queue stalls, no matter how clever the queue is.&lt;/p&gt;

&lt;p&gt;And pre-specifying that whole embedded-artifacts arc up front is what made this visible, because it told me exactly &lt;em&gt;where&lt;/em&gt; my specs leak. The verdict: pre-specifying nails &lt;strong&gt;structure&lt;/strong&gt; and leaks on &lt;strong&gt;judgment.&lt;/strong&gt; The file decomposition was right. The dependency order (A blocks the rest, B/C/D are independent) was right. What broke was subtler:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4lvtyfhnsmjpeg825ods.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4lvtyfhnsmjpeg825ods.jpg" alt="The Real Bottleneck Was Never the Task Slot" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One task rested on a &lt;strong&gt;wrong premise.&lt;/strong&gt; I’d specified the “diagram” artifact as something the pipeline could satisfy with an ASCII drawing. But my actual diagram practice is to render real images, with ASCII only as a fallback. That’s not a fork the builder could catch or resolve. It’s a wrong assumption baked into the spec, and only I could see it, because it lived in my head and not in the code.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;judgment call hid inside a “mechanical” task.&lt;/strong&gt; The artifact checker had to count distinct artifacts, and my supposedly airtight spec quietly let one image get counted twice toward the requirement. The builder implemented exactly what I wrote. What I wrote had a bug. I only caught it because I read the diff.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Which is the callback to post one’s whole “trust but verify” spine, and it’s &lt;em&gt;more&lt;/em&gt; load-bearing in a pipeline, not less. The batch doesn’t skip verification. It removes the &lt;em&gt;round-trip with me between tasks,&lt;/em&gt; not the rigor. I still read every diff the builder actually wrote. Both of those bugs, the double-count and a separate one where a config value silently broke its file format, got caught at my review, not by the agents. The lesson for the queue is precise: a task is only “ready” when its &lt;em&gt;premises&lt;/em&gt; and its &lt;em&gt;judgment semantics&lt;/em&gt; are pinned, not just its files. Pinning the files is the easy 80%. The leak is always in the other 20%, and the other 20% is mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Name
&lt;/h2&gt;

&lt;p&gt;The seed of this post had a whole brainstorm in it about &lt;em&gt;finally&lt;/em&gt; renaming the skill. Its name describes one of the three things it does and I can never remember it. I had candidates: &lt;em&gt;planwright,&lt;/em&gt; for the one who authors plans, and &lt;em&gt;foreman,&lt;/em&gt; for the one who keeps the line moving and the builder fed. And given that this entire post is about keeping the line moving, &lt;em&gt;foreman&lt;/em&gt; was right there, practically gift-wrapped.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F73ehr4atl4buty2gt7y6.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F73ehr4atl4buty2gt7y6.jpg" alt="The Name" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I did not rename it. I codified the whole convention, ported it into the skill, patched both projects, and left the dumb name exactly where it was. The rename is real work. It’s wired into three repos now, and I wasn’t going to fake-land it just to give this post a tidy bow. So the name is still wrong, &lt;em&gt;foreman&lt;/em&gt; is still leading, and it’ll get fixed when it gets fixed. Consider this the third post in a row where I’ve promised to rename the thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Honest Scope Note
&lt;/h2&gt;

&lt;p&gt;Every post gets one. Here’s this one’s: &lt;strong&gt;this is the frontier of the method, not a settled result.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The pacing-over-throughput call is fresh. I made it by running the throughput version, hating it, and reasoning backward, which is a strong signal but a sample size of basically one. The unattended mode, with its skip-and-continue and its state logging, is &lt;em&gt;built&lt;/em&gt; but barely road-tested; I haven’t actually walked away for a real hour and come back to judge what I found, and I’d bet the skip-and-continue logic has edge cases I haven’t hit. The “promote a recurring judgment into an automatic rule” idea from the last post is still unproven. And the deepest claim here, that spec throughput, not the task slot, is the true bottleneck, is a hypothesis I’ve confirmed exactly once.&lt;/p&gt;

&lt;p&gt;What I’m confident about is the shape of the mistake, because I made it cleanly: optimizing throughput when I wanted pacing is a real and seductive wrong turn, and “keep the builder busy” is exactly the kind of goal that leads you into it. The metric that’s easy to put on a dashboard, work per unit time, is not always the one you actually care about. Sometimes the gap is the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s Next
&lt;/h2&gt;

&lt;p&gt;I don’t know yet, and I’m not going to force a cliffhanger. The plan stays lean, the builder runs on a paced queue, and I can stack it shallow to work alongside it or deep to walk away from it. The pieces are in place. What I don’t have is a real verdict on the unattended mode under fire, or a clue whether the “spec throughput is the bottleneck” insight holds up across more than one feature arc. Those are the next things to actually live with rather than theorize about.&lt;/p&gt;

&lt;p&gt;There may be a post four. There may not. The method’s still moving, and I’d rather tell you what actually happened than what would make a clean ending. So far that’s served the series fine. We’ll see what the gap fills with.&lt;/p&gt;

</description>
      <category>aicodingagents</category>
      <category>sonnetagent</category>
      <category>planmdmethodology</category>
      <category>aidevelopmentworkflo</category>
    </item>
  </channel>
</rss>
