<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kirill Lukyanov</title>
    <description>The latest articles on DEV Community by Kirill Lukyanov (@klukyanov).</description>
    <link>https://dev.to/klukyanov</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4092798%2F719db81c-244e-4965-a433-0662ed094a53.png</url>
      <title>DEV Community: Kirill Lukyanov</title>
      <link>https://dev.to/klukyanov</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/klukyanov"/>
    <language>en</language>
    <item>
      <title>Lunar Fasting: an intermittent fasting app that puts your Apple Health activity on the lunar calendar</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Sun, 04 Oct 2026 16:34:20 +0000</pubDate>
      <link>https://dev.to/klukyanov/lunar-fasting-an-intermittent-fasting-app-that-puts-your-apple-health-activity-on-the-lunar-4e2n</link>
      <guid>https://dev.to/klukyanov/lunar-fasting-an-intermittent-fasting-app-that-puts-your-apple-health-activity-on-the-lunar-4e2n</guid>
      <description>&lt;p&gt;Intermittent fasting has one simple rule: eat inside a window, don't eat outside it. The rule is easy, keeping it is not. Some days the window holds effortlessly, other days you break by lunch and have no idea why. I wanted an app that doesn't just tick a timer but shows me, from my own data, which days are easier and which are harder.&lt;/p&gt;

&lt;p&gt;That's how &lt;strong&gt;Lunar Fasting&lt;/strong&gt; came about. It went live on the App Store on October 2, 2026: 157 days from the first commit, 28 builds, 656 automated tests and one App Review rejection. Below: how to pick a fasting interval, what the app does, and the two features it was really built for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fasting intervals: which one to pick
&lt;/h2&gt;

&lt;p&gt;"16:8" means 16 hours without food and an 8-hour eating window. The longer the fast, the stricter the regime. The app supports six protocols:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Protocol&lt;/th&gt;
&lt;th&gt;Fast&lt;/th&gt;
&lt;th&gt;Eating window&lt;/th&gt;
&lt;th&gt;In real life&lt;/th&gt;
&lt;th&gt;Good for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;12:12&lt;/td&gt;
&lt;td&gt;12 h&lt;/td&gt;
&lt;td&gt;12 h&lt;/td&gt;
&lt;td&gt;Dinner at 8 pm, breakfast at 8 am — just no night snacks&lt;/td&gt;
&lt;td&gt;Beginners&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;14:10&lt;/td&gt;
&lt;td&gt;14 h&lt;/td&gt;
&lt;td&gt;10 h&lt;/td&gt;
&lt;td&gt;Dinner at 7 pm, breakfast at 9 am&lt;/td&gt;
&lt;td&gt;A gentle next step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16:8&lt;/td&gt;
&lt;td&gt;16 h&lt;/td&gt;
&lt;td&gt;8 h&lt;/td&gt;
&lt;td&gt;Eat from noon to 8 pm&lt;/td&gt;
&lt;td&gt;The most popular one, the default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;18:6&lt;/td&gt;
&lt;td&gt;18 h&lt;/td&gt;
&lt;td&gt;6 h&lt;/td&gt;
&lt;td&gt;Eat from 1 pm to 7 pm, two meals&lt;/td&gt;
&lt;td&gt;When 16:8 feels easy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;20:4&lt;/td&gt;
&lt;td&gt;20 h&lt;/td&gt;
&lt;td&gt;4 h&lt;/td&gt;
&lt;td&gt;Late lunch and early dinner in one short window&lt;/td&gt;
&lt;td&gt;Experienced, for a while&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OMAD&lt;/td&gt;
&lt;td&gt;23 h&lt;/td&gt;
&lt;td&gt;1 h&lt;/td&gt;
&lt;td&gt;One meal a day&lt;/td&gt;
&lt;td&gt;The strictest, only deliberately&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The app is not a doctor: with diabetes, pregnancy, eating disorders or any chronic condition, talk to a doctor before you start.&lt;/p&gt;

&lt;p&gt;My advice: start with 12:12 or 14:10 and move up when the current step stops taking effort. You can switch protocols any time, and past days won't be recolored — each day is judged against the goal that was active when it ended.&lt;/p&gt;

&lt;h3&gt;
  
  
  Fasting without a "start" button
&lt;/h3&gt;

&lt;p&gt;Most fasting trackers have a timer you start and stop by hand. Forget to tap it and the day is lost. Lunar Fasting doesn't store fasting windows at all: they are &lt;strong&gt;derived from logged meals&lt;/strong&gt;. Log breakfast and dinner, and the app knows how long you went without food.&lt;/p&gt;

&lt;p&gt;A few rules came out of this, each pinned by tests. A 2 am snack honestly breaks the fast instead of stretching it. If you forgot to log yesterday's dinner, you don't get a 40-hour "record" — you get "no data". Log a meal retroactively and the whole chain recalculates.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why lunar days, not calendar dates
&lt;/h2&gt;

&lt;p&gt;A lunar day runs from one moonrise to the next, about 24 h 50 min. There are roughly thirty of them in a lunar month. Most lunar calendars compute them with a mean formula and are off by up to half a day. Lunar Fasting computes moonrise astronomically for your location and uses true new moon instants (Meeus), not averages.&lt;/p&gt;

&lt;p&gt;To be clear: I don't claim the Moon affects appetite or how much you move. Science hasn't shown that. The app doesn't promise — it &lt;strong&gt;checks, on your own data&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calories from a photo
&lt;/h2&gt;

&lt;p&gt;Snap your plate and a vision model splits it into items — side dish, cutlet, salad — each with its weight in grams. Add a text hint for ambiguous dishes ("buckwheat with mushrooms, no oil") and accuracy goes up noticeably.&lt;/p&gt;

&lt;p&gt;One deliberate decision: &lt;strong&gt;the model never computes the total&lt;/strong&gt;. It only returns item names, grams and kcal per 100 g; the app does the arithmetic. Small models are surprisingly bad at mental math, and this way every number is visible and editable.&lt;/p&gt;

&lt;p&gt;Recognition runs either &lt;strong&gt;on-device via Apple Intelligence&lt;/strong&gt; (no network, no key, the photo never leaves the phone) or through &lt;strong&gt;your own key&lt;/strong&gt; for OpenAI, Gemini, Qwen or DeepSeek, stored in the Keychain and sent straight to the provider. "Auto" mode tries on-device first. Meals are written to Apple Health too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main feature: your activity on the lunar calendar
&lt;/h2&gt;

&lt;p&gt;On first launch the app pulls a year of activity from Apple Health — steps, active energy, workouts and stand hours — and keeps it updated, re-syncing the last week each time because watch data arrives late.&lt;/p&gt;

&lt;p&gt;Four numbers become one &lt;strong&gt;activity index from 0 to 100&lt;/strong&gt;. Each metric is normalized to a daily target and capped, so overachieving doesn't inflate it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Daily target&lt;/th&gt;
&lt;th&gt;Weight&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Steps&lt;/td&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;35%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Active energy&lt;/td&gt;
&lt;td&gt;500 kcal&lt;/td&gt;
&lt;td&gt;30%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workouts&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stand hours&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;15%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Without an Apple Watch, workouts and stand hours are almost always zero and the index would cap at about 65. So in statistics the app drops components you have no data for and renormalizes the weights.&lt;/p&gt;

&lt;p&gt;Calendar days and lunar days don't line up, so each lunar day gets a weighted value from the two calendar days it overlaps, proportional to time. The charts show your average index for each of the thirty lunar days and each moon phase — your own year of movement, not someone's theory.&lt;/p&gt;

&lt;h2&gt;
  
  
  A forecast for the lunar month
&lt;/h2&gt;

&lt;p&gt;The forecast is an additive model, and the app shows each day's breakdown:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Baseline&lt;/strong&gt; — your mean index with a 45-day half-life, so recent weeks weigh more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weekday effect&lt;/strong&gt; — 7 and 29.5 are incommensurable; without this axis, the weekly rhythm would leak into a fake lunar signal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recent trend&lt;/strong&gt; — the last two weeks vs. baseline, decaying over the horizon.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Moon phase and lunar day number&lt;/strong&gt; — hierarchical, and only if your history actually shows it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The last part matters most: lunar effects go through &lt;strong&gt;empirical-Bayes shrinkage&lt;/strong&gt;. With few observations they shrink to zero, so with no lunar pattern in your data the forecast honestly falls back to baseline, weekday and trend. There's an uncertainty band that widens with the horizon and a confidence level: the forecast appears after 21 days of history, medium confidence from 60 days, high from 180. With 60+ days the app also backtests itself: it hides the last two weeks, predicts them and compares the error with a naive "same as average" forecast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coming in 1.1: Google Health
&lt;/h2&gt;

&lt;p&gt;If you wear a Google or Fitbit band with an iPhone, its data reaches Apple Health poorly: workouts go missing, calories and heart rate disagree, steps double. Lunar Fasting 1.1 adds a direct bridge: sign in with Google, and the app pulls workouts, steps, distance, active energy and heart rate from the Google Health API and writes them to Apple Health — deduplicated via &lt;code&gt;HKMetadataKeySyncIdentifier&lt;/code&gt;, with Google's echo of Apple Health data filtered out. Band activity then feeds the index and forecast just like Apple Watch data.&lt;/p&gt;

&lt;p&gt;One caveat: Google requires OAuth verification and a security audit for health scopes. Until then the feature runs in testing mode, capped at 100 users, so 1.1 ships it to a limited group first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Free vs. Pro
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Free:&lt;/strong&gt; six fasting protocols, meals and manual calories, water, weight, daily ratings, lunar calendar, stats for steps, water, weight and calories, CSV export. No account, no server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro:&lt;/strong&gt; activity index, monthly forecast, calories from a photo. Weekly or yearly subscription with a 3-day free trial.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;iPhone and iPad, iOS 17+, English and Russian. &lt;a href="https://apps.apple.com/app/id6772260495" rel="noopener noreferrer"&gt;Get it on the App Store&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Would you trust a forecast that openly tells you when it has nothing to say?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/lunar-fasting/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ios</category>
      <category>swift</category>
      <category>healthkit</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Kimi K3, DeepSeek and GLM free from NVIDIA: I tested the claim and here is the catch</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Thu, 01 Oct 2026 09:42:31 +0000</pubDate>
      <link>https://dev.to/klukyanov/kimi-k3-deepseek-and-glm-free-from-nvidia-i-tested-the-claim-and-here-is-the-catch-2m9b</link>
      <guid>https://dev.to/klukyanov/kimi-k3-deepseek-and-glm-free-from-nvidia-i-tested-the-claim-and-here-is-the-catch-2m9b</guid>
      <description>&lt;p&gt;A claim is going around: NVIDIA hands out a free API key for Kimi K3, DeepSeek V4.1 Flash and two flavours of GLM 5.3 — no credit card, fifteen minutes. I went through it myself. The models are real, the free tier is more generous than the reposts say, and there is one catch the headline leaves out: a phone number check that does not cover every country.&lt;/p&gt;

&lt;h2&gt;
  
  
  What NVIDIA gives away
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://build.nvidia.com" rel="noopener noreferrer"&gt;NVIDIA Build&lt;/a&gt; is NVIDIA's catalog of hosted open models. The API is OpenAI-compatible: same &lt;code&gt;/v1/chat/completions&lt;/code&gt; request, you only swap the base URL and the key. As of October 1, 2026, &lt;code&gt;integrate.api.nvidia.com/v1/models&lt;/code&gt; lists 81 models, including the four most interesting open-weight releases of this autumn.&lt;/p&gt;

&lt;p&gt;The part most reposts get wrong is the money. New accounts used to get roughly 1,000 credits, and people still quote that number. Credits are gone. NVIDIA's own account verification dialog now says &lt;em&gt;"Unlimited API requests without daily limits"&lt;/em&gt;. What is limited is the rate: threads on the NVIDIA developer forum converge on about 40 requests per minute per key, and a forum moderator stated in July that the limit depends on the model, use case and overall traffic and cannot be officially raised on the free tier.&lt;/p&gt;

&lt;p&gt;40 RPM is 57,600 requests a day if you hammer it nonstop. For prototyping, a personal agent or evaluating a model on your own tasks, that is effectively unlimited. For production, NVIDIA points you to paid options: partner endpoints or self-hosted deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four models
&lt;/h2&gt;

&lt;p&gt;All four are mixture-of-experts models, and all four ship with a 1,048,576-token context window.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;API ID&lt;/th&gt;
&lt;th&gt;Total params&lt;/th&gt;
&lt;th&gt;Active per token&lt;/th&gt;
&lt;th&gt;Highlights&lt;/th&gt;
&lt;th&gt;License&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;moonshotai/kimi-k3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;~2.8T&lt;/td&gt;
&lt;td&gt;not stated&lt;/td&gt;
&lt;td&gt;Long-horizon agentic coding, tool use, image input; thinking always on&lt;/td&gt;
&lt;td&gt;Modified MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM 5.3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;z-ai/glm-5.3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;753B&lt;/td&gt;
&lt;td&gt;~40B&lt;/td&gt;
&lt;td&gt;Text, reasoning, tool calling&lt;/td&gt;
&lt;td&gt;MIT with a clause for &amp;gt;$10B-revenue resellers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM 5.3 Flash&lt;/td&gt;
&lt;td&gt;&lt;code&gt;z-ai/glm-5.3-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;320B&lt;/td&gt;
&lt;td&gt;18B&lt;/td&gt;
&lt;td&gt;Text + images, tools, structured output, thinking budget&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4.1 Flash&lt;/td&gt;
&lt;td&gt;&lt;code&gt;deepseek-ai/deepseek-v4.1-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;552B&lt;/td&gt;
&lt;td&gt;8B prefill / 16B decode&lt;/td&gt;
&lt;td&gt;Multimodal, reasoning effort adjustable 1–100&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Source: model cards on build.nvidia.com.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For coding I would start with Kimi K3 — it is the reason for the hype, and running it yourself takes a rack you do not have. Free access on someone else's GPUs is the only sensible way for most of us to try it. GLM 5.3 Flash and DeepSeek V4.1 Flash are for fast, cheap calls where you do not need the smartest answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The catch: phone verification
&lt;/h2&gt;

&lt;p&gt;Sign-up itself is smooth: email, password, NVIDIA Developer Program account. No card, as promised. But on the API keys page you get a modal — &lt;em&gt;"We'll need to verify your phone number"&lt;/em&gt; — and no key until you enter a one-time SMS code. NVIDIA frames it as fraud and abuse protection.&lt;/p&gt;

&lt;p&gt;The country list is not global. Russia is not on it at all; under the list NVIDIA says it is "rapidly expanding worldwide availability". Kazakhstan is listed, but in my attempt with a Kazakh number the "Send Code to Phone" button never became active, and the page showed no error. I could not confirm whether Kazakh numbers go through.&lt;/p&gt;

&lt;p&gt;I did not try to get around the check and would not recommend it: virtual numbers and borrowed SIMs violate the terms, and a key obtained that way can be revoked along with the account. So "fifteen minutes, no card" is true only if your phone number is from a supported country.&lt;/p&gt;

&lt;h2&gt;
  
  
  First request
&lt;/h2&gt;

&lt;p&gt;If your number is supported, the first call takes a minute. Any client that lets you change the base URL works, including coding agents.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;NVIDIA_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"nvapi-..."&lt;/span&gt;

curl https://integrate.api.nvidia.com/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$NVIDIA_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "moonshotai/kimi-k3",
    "messages": [{"role": "user", "content": "Explain Swift async/await in three sentences"}],
    "max_tokens": 1024
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Swap &lt;code&gt;model&lt;/code&gt; for any ID from the table. One Kimi K3 detail: thinking is always on, and in multi-turn chats or tool calls you must send back the full previous assistant message, including &lt;code&gt;reasoning_content&lt;/code&gt; and &lt;code&gt;tool_calls&lt;/code&gt;, or it loses the thread.&lt;/p&gt;

&lt;h2&gt;
  
  
  My take
&lt;/h2&gt;

&lt;p&gt;The offer itself is great: four of the strongest open models of the season, no payment, no daily cap, standard API. If your phone number qualifies, it is the easiest way to try Kimi K3 without renting hardware. Just know that "no card, fifteen minutes" quietly skips the phone step — and that step is where some of us stop.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/nvidia-build-free-models/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>api</category>
    </item>
    <item>
      <title>Listen to your coding agent instead of reading it: how I save my eyes</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:42:49 +0000</pubDate>
      <link>https://dev.to/klukyanov/listen-to-your-coding-agent-instead-of-reading-it-how-i-save-my-eyes-4ci0</link>
      <guid>https://dev.to/klukyanov/listen-to-your-coding-agent-instead-of-reading-it-how-i-save-my-eyes-4ci0</guid>
      <description>&lt;p&gt;My working day at the computer used to be code and documentation. Now there's a third stream, and it turned out to be the biggest one: the agent's replies. Claude Code reports every step, explains what it changed, asks clarifying questions and sends summaries. All of it is text, and all of it is read with your eyes, in small terminal font.&lt;/p&gt;

&lt;p&gt;By evening you feel it physically. So I put together a pair of my own macOS apps. The first one, SafeYourEye, makes sure my eyes actually get rest. The second one, PromtsMaker, takes part of the load off them: the agent talks to me by voice, and I answer by voice too. Here's how it works for me and how to set it up in ten minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The agent writes more than I can read
&lt;/h2&gt;

&lt;p&gt;With an agent your eyes don't get a break, they work harder. It writes the code, but I read its reports, check diffs and answer questions. Instead of typing text I read it, constantly, in small portions, switching between windows.&lt;/p&gt;

&lt;p&gt;Yet most of what the agent writes doesn't need a screen. "Starting", "tests are green", "found two places where this is used, fix both?" can simply be heard. You need the screen for diffs and code; for status and questions, voice is enough.&lt;/p&gt;

&lt;p&gt;Hence two fixes: don't let your eyes work without breaks, and move to audio everything that doesn't have to be read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Half one: SafeYourEye keeps the breaks
&lt;/h2&gt;

&lt;p&gt;Everyone knows the 20-20-20 rule: every 20 minutes, look at something about 20 feet away for 20 seconds. Almost nobody follows it, because 20 minutes fly by at work. Here's what the app does:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Timer.&lt;/strong&gt; Counts a work session (5 to 60 minutes) and shows a break window with a 20-second countdown. A window is harder to dismiss than a banner.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Idle detection.&lt;/strong&gt; Walk away from the computer and the session doesn't burn: time without keyboard and mouse isn't counted as work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distance and light.&lt;/strong&gt; Using the camera, it hints when you sit too close to the screen or the room gets dark. Frames are processed on the Mac itself and never leave it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vision checks.&lt;/strong&gt; Four self-checks on one screen, with history and a PDF export for your doctor. Not a diagnosis, a way to notice a trend.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With an agent the timer turned out to be especially useful. While the agent runs a long task I don't sit staring at the terminal, I take the break. The SafeYourEye window and the agent's summary often arrive at about the same time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Half two: PromtsMaker gives the agent a voice
&lt;/h2&gt;

&lt;p&gt;PromtsMaker started as a prompt editor: dictate a thought, the app cleans up the text and adds project context. You don't have to use prompts at all, though. Since version 1.4 the app has a bridge to the agent, and that alone is worth installing it for.&lt;/p&gt;

&lt;p&gt;Inside PromtsMaker runs a local MCP server. Claude Code connects to it like to any other tool, and the agent gets a mouth and ears:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;say&lt;/strong&gt; — speak a short line during the work: "starting", "build passed";&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ask_user&lt;/strong&gt; — ask by voice and wait for a spoken answer: the app starts dictation itself and returns the transcript to the agent;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;report_result&lt;/strong&gt; — read the summary aloud, while details (diff, command output) stay as text in the panel;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;wait_for_prompt&lt;/strong&gt; — pick up a new task I dictated in PromtsMaker, without copying it into the terminal.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent has one rule: don't read walls of text. One or two sentences aloud; a long result only after "read it all?". Without that rule the first version of the bridge honestly read everything.&lt;/p&gt;

&lt;p&gt;I run the bridge live with Claude Code and OpenCode. The protocol is standard, MCP over HTTP, so any agent that can connect MCP servers over HTTP should work. Codex CLI supports such servers, but I haven't tested it with the bridge myself yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it looks like on an exercise bike
&lt;/h2&gt;

&lt;p&gt;My favourite scenario. A long task: rewrite a module, run the tests, fix whatever breaks. No reason to watch it in the terminal.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Dictate the task&lt;/strong&gt; into PromtsMaker as a normal sentence. The app strips the "umm"s and fixes terms using the project glossary. Sending it to the agent is ⌘⏎, no window switching.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Get on the bike.&lt;/strong&gt; The screen is far away; I don't need to read it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The agent works and talks.&lt;/strong&gt; Short updates are spoken, and I know where it is without looking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The agent asks, I answer by voice.&lt;/strong&gt; At a fork it asks out loud, the app starts dictation, I say the answer, the agent continues.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The summary is spoken.&lt;/strong&gt; If it's clear, I dictate the next task right away. I look at the diff later, in one pass, instead of in bits for an hour.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Double benefit: my eyes rest on the far distance instead of the terminal, and my body moves while the agent does work that needs only a couple of spoken answers from me.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eyes or ears: what moved where
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What's happening&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;Now&lt;/th&gt;
&lt;th&gt;Who helps&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Give a task&lt;/td&gt;
&lt;td&gt;Type or paste into the terminal&lt;/td&gt;
&lt;td&gt;Dictate, ⌘⏎&lt;/td&gt;
&lt;td&gt;PromtsMaker&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Know which step the agent is on&lt;/td&gt;
&lt;td&gt;Watch terminal output&lt;/td&gt;
&lt;td&gt;A short spoken line&lt;/td&gt;
&lt;td&gt;PromtsMaker, &lt;code&gt;say&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answer a clarifying question&lt;/td&gt;
&lt;td&gt;Read it, type an answer&lt;/td&gt;
&lt;td&gt;Hear it, answer by voice&lt;/td&gt;
&lt;td&gt;PromtsMaker, &lt;code&gt;ask_user&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get the result&lt;/td&gt;
&lt;td&gt;Read the whole report&lt;/td&gt;
&lt;td&gt;Spoken summary, details as text&lt;/td&gt;
&lt;td&gt;PromtsMaker, &lt;code&gt;report_result&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Review diff and code&lt;/td&gt;
&lt;td&gt;Eyes&lt;/td&gt;
&lt;td&gt;Eyes, in one pass at the end&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;While the agent works long&lt;/td&gt;
&lt;td&gt;Sit at the screen and wait&lt;/td&gt;
&lt;td&gt;Break or exercise bike&lt;/td&gt;
&lt;td&gt;SafeYourEye&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Don't overstay at the screen&lt;/td&gt;
&lt;td&gt;Hope you remember&lt;/td&gt;
&lt;td&gt;Break window every 20 minutes&lt;/td&gt;
&lt;td&gt;SafeYourEye&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sat too close, room got dark&lt;/td&gt;
&lt;td&gt;Don't notice&lt;/td&gt;
&lt;td&gt;Camera hint&lt;/td&gt;
&lt;td&gt;SafeYourEye&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The code I always review with my eyes, and that's fine. Tool names are for the agent; users never need to know them.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to set up the same pair
&lt;/h2&gt;

&lt;p&gt;Both apps are for macOS and are in the Mac App Store. &lt;a href="https://apps.apple.com/us/app/promtsmaker/id6786639355?mt=12" rel="noopener noreferrer"&gt;PromtsMaker&lt;/a&gt; is free (macOS 14+). &lt;a href="https://apps.apple.com/us/app/safeyoureye-eye-care-timer/id6770624870?mt=12" rel="noopener noreferrer"&gt;SafeYourEye&lt;/a&gt; runs on macOS 13+, it's free to download, with full features in a Pro subscription ($2.99/month or $19.99/year).&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Install SafeYourEye&lt;/strong&gt; and set the session length. You can skip camera access: distance and blink hints turn off, the timer keeps working.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Install PromtsMaker&lt;/strong&gt; and open your project folder in it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enable the agent bridge.&lt;/strong&gt; The app starts the server on 127.0.0.1 only, generates an access token and shows a ready-made connect command.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Connect Claude Code&lt;/strong&gt; with that command. It looks like this (the app picks the port, the token is added as an Authorization header):
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; http promtsmaker http://127.0.0.1:8765/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Tell the agent "wait for tasks from PromtsMaker".&lt;/strong&gt; On first connect it stores the bridge rules in its own memory; you don't need to add instructions to project files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Allow speech recognition&lt;/strong&gt; when macOS asks. Better do it right away: the agent's first voice question otherwise waits on the system prompt, and you can't answer until it's dismissed.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What's next: read any reply with one key
&lt;/h2&gt;

&lt;p&gt;The bridge works when the agent decides to say something. But sometimes the session runs in a plain terminal, without the bridge, and you'd rather listen to the reply than read it. For that, the next PromtsMaker update adds a global hotkey: it reads the latest Claude Code reply for the project in the open tab, and pressing it again stops. The app strips Markdown first, so you don't hear asterisks and hashes.&lt;/p&gt;

&lt;p&gt;It already works in my dev build and will ship in the next App Store version.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/listen-to-agent/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>macos</category>
      <category>productivity</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Fitbit Air and iPhone: why your band's data gets lost on the way to Apple Health</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Mon, 28 Sep 2026 09:30:09 +0000</pubDate>
      <link>https://dev.to/klukyanov/fitbit-air-and-iphone-why-your-bands-data-gets-lost-on-the-way-to-apple-health-43g8</link>
      <guid>https://dev.to/klukyanov/fitbit-air-and-iphone-why-your-bands-data-gets-lost-on-the-way-to-apple-health-43g8</guid>
      <description>&lt;p&gt;I wear a Google Fitbit Air on my wrist and carry an iPhone. Sounds like a normal combo. In practice it's two health worlds that barely understand each other: the band counts steps, heart rate, sleep and workouts, the Google Health app shows all of it, but Apple Health — where every other iPhone app gets its data — receives only part of it, and not always in the right shape.&lt;/p&gt;

&lt;p&gt;I ran into this as a developer: one of my apps builds activity analytics on top of Apple Health, and with the band's data the analytics started to lie. Here's what's actually broken, why it's hard to fix, what the paid bridges on the App Store cost, and how I built my own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two "Health" apps that refuse to be friends
&lt;/h2&gt;

&lt;p&gt;Apple and Google store health data in fundamentally different ways. Apple Health (HealthKit) lives only on the phone: an encrypted on-device database with no cloud API. Google Health is the opposite — a cloud service: the band sends data to the phone, the phone sends it to your Google account, and that's where it lives.&lt;/p&gt;

&lt;p&gt;To move data from one world to the other you need a program that runs on the iPhone, downloads data from Google's cloud and writes it into Apple's local store. Nobody else can do it: Google's cloud can't reach the database on your iPhone, and Apple only lets apps on the device itself write to it.&lt;/p&gt;

&lt;p&gt;For years Google simply didn't ship such a program. When HealthKit launched in 2014, Fitbit said outright it had no plans to support it, and Fitbit owners with iPhones lived on third-party utilities for a decade. Only in August 2026, in version 5.05, did the Google Health app learn to write to Apple Health itself. Problem solved? Not quite.&lt;/p&gt;

&lt;h2&gt;
  
  
  What gets lost on the way
&lt;/h2&gt;

&lt;p&gt;The official sync works, but with losses. On my own data:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Workouts go missing.&lt;/strong&gt; Some workouts recorded by the band never show up in Apple Health, although they're in Google Health.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calories and heart rate don't match&lt;/strong&gt; between the two apps for the same day.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duplicates.&lt;/strong&gt; The same steps and calories land in Apple Health twice and inflate the daily total.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Some metrics don't transfer at all&lt;/strong&gt;, e.g. heart rate variability (HRV) — even the coverage of the update admits it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you only glance at steps, that's cosmetic. But fitness apps, food diaries and training plans read Apple Health and draw conclusions from an incomplete, partly doubled picture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it's hard to do right
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Different data models.&lt;/strong&gt; In HealthKit a workout is a single object with calories and distance inside. In Google's model it's a set of separate points: an exercise segment, per-minute calories, per-minute distance. Copy everything as-is and calories inside a workout are counted twice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The echo.&lt;/strong&gt; Google Health also &lt;em&gt;reads&lt;/em&gt; Apple Health (iPhone steps, Apple Watch workouts), uploads them to Google's cloud, and they come back as "Google data". A bridge that can't tell its own records from the echo grows duplicates on every sync.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The locked phone.&lt;/strong&gt; Apple Health's store is encrypted while the iPhone is locked, so nothing can be written. Background sync works in fits and starts, and long history imports die when the screen turns off.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Old paths are closing.&lt;/strong&gt; Many cheap bands have been writing to Google Fit for years. Since May 1, 2024 new developers can't sign up for its APIs, and they shut down at the end of 2026. The replacement is the Google Health API, which has its own strings attached (see below).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Paid bridges on the App Store and what they cost
&lt;/h2&gt;

&lt;p&gt;All of them are free to download and charge via in-app purchases. US App Store prices as of September 28, 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;App&lt;/th&gt;
&lt;th&gt;Subscription&lt;/th&gt;
&lt;th&gt;Lifetime&lt;/th&gt;
&lt;th&gt;Rating&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fitbit Sync To Apple Health (Syncfit)&lt;/td&gt;
&lt;td&gt;$4.99–9.99/mo, $19.99–39.99/yr&lt;/td&gt;
&lt;td&gt;$59.99&lt;/td&gt;
&lt;td&gt;4.6 (19,688)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fitbit to Apple Health Sync (myFitnessSync)&lt;/td&gt;
&lt;td&gt;$5.99–9.99/mo, $9.99/wk, $39.99/yr&lt;/td&gt;
&lt;td&gt;$39.99–59.99&lt;/td&gt;
&lt;td&gt;4.3 (29,189)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Power Sync for Fitbit&lt;/td&gt;
&lt;td&gt;auto-sync $1.99/mo, $7.99/yr (daily totals only)&lt;/td&gt;
&lt;td&gt;$14.99&lt;/td&gt;
&lt;td&gt;4.2 (15,642)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Power Sync: Fitness to Health&lt;/td&gt;
&lt;td&gt;$5.99/wk, $10.99/mo, $19.99–24.99/yr&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;4.5 (12,888)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fitbit to Health Sync Solver&lt;/td&gt;
&lt;td&gt;$9.99/wk, $39.99/yr&lt;/td&gt;
&lt;td&gt;$59.99&lt;/td&gt;
&lt;td&gt;4.5 (3,250)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fitbit Sync to Health App&lt;/td&gt;
&lt;td&gt;$9.99/wk, $19.99/mo, $39.99/yr&lt;/td&gt;
&lt;td&gt;$59.99&lt;/td&gt;
&lt;td&gt;4.1 (1,628)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fitbit to Apple Health Sync · (StepsApp)&lt;/td&gt;
&lt;td&gt;Pro: $4.99 or $19.99&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;4.3 (374)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The typical price is $39.99/year or $59.99 lifetime. The Fitbit Air costs about $100, so you're asked to pay almost half the band's price again just to get its data onto your iPhone properly. Weekly $9.99 plans add up to roughly $520 a year — five times the band.&lt;/p&gt;

&lt;p&gt;One caveat matters more than price: nearly all of these are built around the Fitbit / Google Health account. If your band writes to the old Google Fit, a bridge may not see its data. Test on the free trial before paying for a year.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I built my own bridge
&lt;/h2&gt;

&lt;p&gt;I didn't want to pay for a bridge, but my app's analytics had to be fixed, so I wrote the sync myself: the app pulls data straight from the Google Health API and writes it into Apple Health. (Turn off Google's own sync in that case, so you don't have two sources of the same thing.) The core work took one day; the interesting part is the pitfalls.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;API.&lt;/strong&gt; Google Fit REST is closed to new developers, so it's the Google Health API v4 (&lt;code&gt;dataTypes/{type}/dataPoints&lt;/code&gt;). All its scopes are classified as Restricted — more on that below.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sign-in without a secret.&lt;/strong&gt; Google officially expects a web-server client with a client secret, and you can't ship a secret inside an iOS app. An iOS client with PKCE via &lt;code&gt;ASWebAuthenticationSession&lt;/code&gt; worked — no backend, no Google SDK. Tokens live in the Keychain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dedup the Apple way.&lt;/strong&gt; Every HealthKit sample gets &lt;code&gt;HKMetadataKeySyncIdentifier&lt;/code&gt; like &lt;code&gt;gh:&amp;lt;type&amp;gt;:&amp;lt;point id&amp;gt;&lt;/code&gt; plus &lt;code&gt;HKMetadataKeySyncVersion&lt;/code&gt;. A repeated sync replaces the record instead of adding a new one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Echo filtering.&lt;/strong&gt; Points that came into Google from Apple Health (platform &lt;code&gt;HEALTH_KIT&lt;/code&gt;) are skipped, otherwise you get a loop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whole workouts.&lt;/strong&gt; Workouts are built with &lt;code&gt;HKWorkoutBuilder&lt;/code&gt; with calories and distance inside, and per-minute calories/distance that fall into the workout window are dropped so the day doesn't count them twice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;History in chunks.&lt;/strong&gt; Fresh days first, then history from newest to oldest: workouts and sleep up to a year, steps and calories 90 days, heart rate 30. Progress is stored per metric, so an interrupted sync resumes where it stopped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the screen on.&lt;/strong&gt; During a manual sync the idle timer is disabled, and if the phone gets locked anyway the user sees a clear hint instead of a database error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Format surprises.&lt;/strong&gt; 64-bit numbers arrive as strings, distance comes in millimeters and sits somewhere other than the docs suggested — the first parser version crashed on a real response. The mapper is covered by tests on trimmed real data.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The main limitation isn't technical. Until the app passes Google's OAuth verification and a third-party security audit, the feature runs in testing mode: at most 100 users, and the Google sign-in expires every 7 days. Fine for personal use, not for everyone. The paid bridges above prove the path is passable — it just takes weeks and money.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do right now if you own a Google band
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Turn on the official sync&lt;/strong&gt; in Google Health: add Apple Health in the connections settings. It's free, and for steps and sleep it's often enough.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Set source priority&lt;/strong&gt; in Apple Health: open a metric → Data Sources &amp;amp; Access → Edit, and put the band above the iPhone so phone and wrist steps don't add up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare a week of workouts&lt;/strong&gt; in Google Health and Apple Health. No differences — you need nothing else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If workouts go missing and they matter&lt;/strong&gt;, pick a bridge from the table — but start with the trial and check &lt;em&gt;your&lt;/em&gt; data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep one bridge.&lt;/strong&gt; Official sync plus a paid app at the same time is the shortest path to duplicates.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;The sync already works in my test build, and I'm going to ship it as a separate app — just the Google Health → Apple Health bridge, nothing else. No name yet: Google verification comes first. I'll write about it when it's out.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/fitbit-air-apple-health/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;br&gt;
&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ios</category>
      <category>healthkit</category>
      <category>swift</category>
      <category>android</category>
    </item>
    <item>
      <title>Which AI Subscription Should You Pick for a Task: a Benchmark Table (September 2026)</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Tue, 22 Sep 2026 11:23:10 +0000</pubDate>
      <link>https://dev.to/klukyanov/which-ai-subscription-should-you-pick-for-a-task-a-benchmark-table-september-2026-5hd6</link>
      <guid>https://dev.to/klukyanov/which-ai-subscription-should-you-pick-for-a-task-a-benchmark-table-september-2026-5hd6</guid>
      <description>&lt;p&gt;Yesterday SpaceXAI shipped Grok 4.7 with a benchmark table in which Grok 4.7 wins. Three weeks ago OpenAI shipped GPT-6 Astra with a table in which Astra wins. Anthropic did the same with Claude Fable 5.1. None of these tables is fake — but every one of them was put together by the team that ships the model.&lt;/p&gt;

&lt;p&gt;So before renewing my own subscriptions I did the boring thing: pulled numbers from &lt;strong&gt;eight sources&lt;/strong&gt; into one matrix. Three independent leaderboards — &lt;a href="https://artificialanalysis.ai/leaderboards/models" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt;, &lt;a href="https://www.vals.ai/home" rel="noopener noreferrer"&gt;Vals AI&lt;/a&gt; and &lt;a href="https://arena.ai/" rel="noopener noreferrer"&gt;Arena&lt;/a&gt; (blind human voting) — plus the launch posts from OpenAI, Anthropic, Google and SpaceXAI, which, put side by side, check each other surprisingly well.&lt;/p&gt;

&lt;p&gt;Snapshot date: &lt;strong&gt;September 22, 2026&lt;/strong&gt;. Every number below is labeled as independent or vendor-reported.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lineup: six models, four subscriptions
&lt;/h2&gt;

&lt;p&gt;The only rule: the model must be available in a regular paid plan an individual can buy.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Released&lt;/th&gt;
&lt;th&gt;Subscription&lt;/th&gt;
&lt;th&gt;Price / month&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5.1&lt;/td&gt;
&lt;td&gt;Sep 1&lt;/td&gt;
&lt;td&gt;Claude Pro / Max&lt;/td&gt;
&lt;td&gt;$20 / $100–200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;before Sep&lt;/td&gt;
&lt;td&gt;Claude Pro / Max&lt;/td&gt;
&lt;td&gt;$20 / $100–200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;Sep 3&lt;/td&gt;
&lt;td&gt;ChatGPT Plus / Pro&lt;/td&gt;
&lt;td&gt;$20 / $200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;before Sep&lt;/td&gt;
&lt;td&gt;ChatGPT Plus / Pro&lt;/td&gt;
&lt;td&gt;$20 / $200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;Sep 2&lt;/td&gt;
&lt;td&gt;Google AI Pro / Ultra&lt;/td&gt;
&lt;td&gt;$19.99 / $250&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4.7*&lt;/td&gt;
&lt;td&gt;Sep 21&lt;/td&gt;
&lt;td&gt;SuperGrok&lt;/td&gt;
&lt;td&gt;$30&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;* At launch Grok 4.7 is available in Cursor, Grok Build and the API; SpaceXAI hasn't yet confirmed it in the SuperGrok consumer docs. Check before paying.&lt;/p&gt;

&lt;p&gt;Left out on purpose: &lt;strong&gt;Meta Muse Spark 1.3&lt;/strong&gt; (free in the Meta AI app — not a subscription), &lt;strong&gt;Claude Mythos 5.1&lt;/strong&gt; (gated to verified security and life-science programs), and open-weight models like DeepSeek, Qwen, Kimi and GLM — they get their own article.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why you shouldn't take a launch table at face value
&lt;/h2&gt;

&lt;p&gt;Not because vendors lie. Because the same benchmark can be run several ways, and every vendor picks the way that flatters them. Four examples from September alone:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ARC-AGI-3.&lt;/strong&gt; OpenAI reports 99.9% for GPT-6 Astra — in a harness that preserves reasoning state between steps. In the neutral harness everyone uses, Astra scores 62.7%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HealthBench Professional.&lt;/strong&gt; SpaceXAI's table gives Claude Fable 5.1 62.1%. The length-adjusted independent leaderboard gives the same model 56.6% — a gap larger than first-to-third place.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OSWorld 2.0.&lt;/strong&gt; OpenAI lists Claude Opus 5 at 70.2%. Anthropic lists the same model at 39.6% — in "strict" mode. Both are correct; they're just not comparable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CursorBench.&lt;/strong&gt; SpaceXAI uses v4.0, Anthropic uses v3.2. Fable 5.1 scores 51.8% and 73.4% respectively.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three rules I now apply: independent rankings beat vendor tables even when they have fewer rows; always check &lt;strong&gt;who is missing&lt;/strong&gt; from a table (Grok 4.7's chart has no GPT-6 Astra); and a number without its mode (max, xhigh, with tools, strict) is not a number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Independent rankings: a tie at the top
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Artificial Analysis Intelligence Index:&lt;/strong&gt; Fable 5.1 and GPT-6 Astra — 53 each, Opus 5 — 51, GPT-5.6 Sol — 47, Grok 4.7 — 46.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vals Index&lt;/strong&gt; (finance, legal research, code migration, medical documentation): Fable 5.1 68.8%, Opus 5 67.2%, Astra 66.6%, Gemini 3.8 Flash 62.3%, Grok 4.7 60.2%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Arena Text&lt;/strong&gt; (blind voting, Sep 13): Fable 5.1 1498, Opus 5 and Gemini 3.8 Flash 1493 each — within the confidence interval. &lt;strong&gt;Arena Code&lt;/strong&gt; is a different story: Astra 1800, Fable 1758, Opus 1687.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're picking one subscription "for everything", Claude and ChatGPT are effectively equal right now. The decision comes down to the tasks where they diverge — and there are plenty.&lt;/p&gt;

&lt;h2&gt;
  
  
  The full matrix: 6 models × 19 tests
&lt;/h2&gt;

&lt;p&gt;Leader of each row in &lt;strong&gt;bold&lt;/strong&gt;. "—" means the model wasn't tested or the result isn't published. &lt;strong&gt;I&lt;/strong&gt; = independent, &lt;strong&gt;V&lt;/strong&gt; = vendor-reported.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task · test&lt;/th&gt;
&lt;th&gt;Fable 5.1&lt;/th&gt;
&lt;th&gt;Opus 5&lt;/th&gt;
&lt;th&gt;GPT-6 Astra&lt;/th&gt;
&lt;th&gt;GPT-5.6 Sol&lt;/th&gt;
&lt;th&gt;Gemini 3.8 Flash&lt;/th&gt;
&lt;th&gt;Grok 4.7&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Overall · AA Index&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;53&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;51&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;53&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;47&lt;/td&gt;
&lt;td&gt;&amp;lt;46&lt;/td&gt;
&lt;td&gt;46&lt;/td&gt;
&lt;td&gt;I · Artificial Analysis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Professional tasks · Vals Index&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.8&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;67.2&lt;/td&gt;
&lt;td&gt;66.6&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;62.3&lt;/td&gt;
&lt;td&gt;60.2&lt;/td&gt;
&lt;td&gt;I · Vals AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blind chat · Arena Text&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1498&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1493&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1493&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;I · Arena&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code, blind · Arena Code&lt;/td&gt;
&lt;td&gt;1758&lt;/td&gt;
&lt;td&gt;1687&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1800&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;I · Arena&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal agent · Terminal-Bench 4.0&lt;/td&gt;
&lt;td&gt;55.8&lt;/td&gt;
&lt;td&gt;52.3&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;57.7&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;37.3&lt;/td&gt;
&lt;td&gt;19.1&lt;/td&gt;
&lt;td&gt;38.0&lt;/td&gt;
&lt;td&gt;V · OpenAI, Anthropic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-horizon SWE · DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;70.0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;74.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;72.7&lt;/td&gt;
&lt;td&gt;73.7&lt;/td&gt;
&lt;td&gt;71.0&lt;/td&gt;
&lt;td&gt;V · Google, SpaceXAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;In-editor coding · CursorBench 4.0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;51.8&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;41.7&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;46.3&lt;/td&gt;
&lt;td&gt;V · SpaceXAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clinical reasoning · HealthBench Pro&lt;/td&gt;
&lt;td&gt;56.6&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;63.4&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;60.5&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;56.7&lt;/td&gt;
&lt;td&gt;I · aggregated leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medical coding · Vals MedCode&lt;/td&gt;
&lt;td&gt;#7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;63.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;#23&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;I · Vals AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Legal agent · Harvey LAB&lt;/td&gt;
&lt;td&gt;6.7&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;2.5&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;19.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;V · SpaceXAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Legal research · Vals Legal Research&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;55.3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;#21&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;I · Vals AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Electrical engineering · EEBench&lt;/td&gt;
&lt;td&gt;56.4&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;39.4&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;64.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;V · SpaceXAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Documents · GDPval-AA (Elo)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1853&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1824&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1711&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;V · Anthropic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-hour office work · AA Briefcase&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1678&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1487&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1657&lt;/td&gt;
&lt;td&gt;V · SpaceXAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Computer use · OSWorld 2.0&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;70.2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;72.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;65.7&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;V · OpenAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Research math · FrontierMath T4&lt;/td&gt;
&lt;td&gt;87.8&lt;/td&gt;
&lt;td&gt;73.2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;97.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;83.0&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;V · OpenAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expert knowledge · HLE with tools&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;65.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;63.6&lt;/td&gt;
&lt;td&gt;57.2&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;V · Anthropic, OpenAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grad-level Q&amp;amp;A · GPQA Diamond&lt;/td&gt;
&lt;td&gt;93.7&lt;/td&gt;
&lt;td&gt;93.7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;96.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;94.6&lt;/td&gt;
&lt;td&gt;95.3&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;V · OpenAI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Images &amp;amp; documents · MMMU Pro&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;90.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;89.9&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;I · Vals AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API price, $ / 1M tokens (in / out)&lt;/td&gt;
&lt;td&gt;10 / 50&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;10 / 50&lt;/td&gt;
&lt;td&gt;4 / 20&lt;/td&gt;
&lt;td&gt;0.75 / 3.75&lt;/td&gt;
&lt;td&gt;2 / 6&lt;/td&gt;
&lt;td&gt;V · price lists&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Percentages unless noted; Elo and indices in points. For Terminal-Bench, Fable 5.1 uses Anthropic's figure (55.8); SpaceXAI's max-effort run shows 57.9. For HealthBench, the length-adjusted version is used; SpaceXAI's table shows 62.1 for Fable 5.1.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it means, task by task
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Coding depends on what you call coding.&lt;/strong&gt; If the model works alone in a terminal, it's a tie between Astra and Fable. If you sit next to it in the editor, Fable leads on real Cursor sessions. If you need a lot of code cheaply, Gemini 3.8 Flash is within four points of the leaders on DeepSWE at roughly 1/13 of the token price.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Medicine: ChatGPT for clinical reasoning, Claude for documentation.&lt;/strong&gt; Astra tops HealthBench Professional (an OpenAI-built benchmark — worth keeping in mind). On medical coding and scribing, independent Vals AI puts Opus 5 and Fable 5.1 first. None of this replaces a doctor: even the leader meets physician criteria in fewer than two thirds of hard cases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Law and electrical engineering: the only rows Grok wins.&lt;/strong&gt; 19.6% on Harvey's Legal Agent Benchmark vs 6.7% for Fable; 64.0% on EEBench vs 56.4%. Both numbers are SpaceXAI's own. On independent legal research, Opus 5 leads.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Office work: Claude for documents, ChatGPT for clicking around.&lt;/strong&gt; Fable leads GDPval-AA and AA Briefcase; Astra leads OSWorld 2.0 and Vals CUA-bench.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Math: Astra by a mile&lt;/strong&gt; — 97.6% on FrontierMath Tier 4, ten points ahead. But on Humanity's Last Exam with tools, Fable (65.0%) beats Astra (57.2%).&lt;/p&gt;

&lt;h2&gt;
  
  
  Task → subscription cheat sheet
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pair-coding in the editor&lt;/td&gt;
&lt;td&gt;Fable 5.1&lt;/td&gt;
&lt;td&gt;Claude Pro / Max&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Autonomous coding agent&lt;/td&gt;
&lt;td&gt;Astra or Fable 5.1&lt;/td&gt;
&lt;td&gt;ChatGPT or Claude&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lots of code, low budget&lt;/td&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;Google AI Pro&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clinical questions&lt;/td&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;ChatGPT Plus&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medical documentation&lt;/td&gt;
&lt;td&gt;Opus 5 / Fable 5.1&lt;/td&gt;
&lt;td&gt;Claude Pro&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Legal research&lt;/td&gt;
&lt;td&gt;Opus 5&lt;/td&gt;
&lt;td&gt;Claude Pro&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Legal agent work, circuits&lt;/td&gt;
&lt;td&gt;Grok 4.7&lt;/td&gt;
&lt;td&gt;SuperGrok (once confirmed)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reports, spreadsheets, decks&lt;/td&gt;
&lt;td&gt;Fable 5.1&lt;/td&gt;
&lt;td&gt;Claude Pro / Max&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;UI automation&lt;/td&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;ChatGPT Plus&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math&lt;/td&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;ChatGPT Plus / Pro&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What I actually use
&lt;/h2&gt;

&lt;p&gt;Two subscriptions, and this table didn't change them. Claude for code — I build iOS and macOS apps and spend most of the day in the editor with the model, which is exactly where Fable 5.1 leads. ChatGPT Pro for everything else — images for my site, calculations, tasks where the model has to finish the job on its own.&lt;/p&gt;

&lt;p&gt;The main takeaway: in 2026, "which AI is the best" is the wrong question. The top three differ less than the same model does across modes. You don't pick the best model — you pick the right one for the task.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Next up:&lt;/strong&gt; the same task-by-task table for open-weight models you can run locally — where they've already caught up with the $20 subscriptions, and where they're still a year behind.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/subscription-models-by-task/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>chatgpt</category>
      <category>claude</category>
    </item>
    <item>
      <title>Yandex open-sourced an 80B model trained from scratch: what's inside and where it wins</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:45:41 +0000</pubDate>
      <link>https://dev.to/klukyanov/yandex-open-sourced-an-80b-model-trained-from-scratch-whats-inside-and-where-it-wins-3c3g</link>
      <guid>https://dev.to/klukyanov/yandex-open-sourced-an-80b-model-trained-from-scratch-whats-inside-and-where-it-wins-3c3g</guid>
      <description>&lt;p&gt;Yandex just open-sourced a language model it trained entirely from scratch — no borrowed weights, no initialization from Qwen or Llama. It's called &lt;strong&gt;AliceAI-Foundation-80B-A3B-Base&lt;/strong&gt;, it's on Hugging Face under Apache 2.0, and it's a surprisingly interesting release if you care about MoE architecture or non-English models.&lt;/p&gt;

&lt;p&gt;Here's what's inside, where it actually wins, and where the benchmark table deserves a second look.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Parameters&lt;/td&gt;
&lt;td&gt;80B total, 3B active per token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Architecture&lt;/td&gt;
&lt;td&gt;MoE: 512 experts, top-10 routed + 1 shared&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Layers&lt;/td&gt;
&lt;td&gt;48 — hybrid Kimi Delta Attention + gated attention (3:1)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context&lt;/td&gt;
&lt;td&gt;262,144 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training&lt;/td&gt;
&lt;td&gt;~18T tokens, from scratch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Languages&lt;/td&gt;
&lt;td&gt;Russian, English&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;License&lt;/td&gt;
&lt;td&gt;Apache 2.0 (commercial use OK)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It's a &lt;strong&gt;base model&lt;/strong&gt; — not instruction-tuned, not a chatbot. Yandex calls it experimental: a testbed for the architecture of its upcoming unified reasoning model, which will power agentic features in its Alice AI assistant.&lt;/p&gt;

&lt;p&gt;Compared to Yandex's previous flagship (Alice AI LLM, 235B, October 2025), it's almost 3x smaller overall and 7x smaller in active parameters — and, per the technical report, beats it on facts, math, code and long context.&lt;/p&gt;

&lt;h2&gt;
  
  
  80 billion parameters, three at work
&lt;/h2&gt;

&lt;p&gt;Every token is routed through 10 of 512 small experts plus one shared expert that's always on. You get the knowledge capacity of an 80B model with roughly the per-token compute of a 3B one.&lt;/p&gt;

&lt;p&gt;The attention stack is the other notable bit. Three out of every four layers use &lt;strong&gt;Kimi Delta Attention&lt;/strong&gt; — a linear attention variant that folds history into a fixed-size state instead of keeping a KV entry for every past token. Only every fourth layer is classic (gated) attention. In practice that means 262K context without the usual KV-cache blowup — exactly the bottleneck you hit with "read this whole contract" or "understand this repo" workloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks: where it wins and where it doesn't
&lt;/h2&gt;

&lt;p&gt;Yandex compared base versions against open models in the same weight class and above. A selection from the official model card, including rows where it loses:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;AliceAI 80B-A3B&lt;/th&gt;
&lt;th&gt;Qwen3.5 35B-A3B&lt;/th&gt;
&lt;th&gt;Nemotron-3 Super 120B-A12B&lt;/th&gt;
&lt;th&gt;DeepSeek-V4 Flash 284B-A13B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;WikiWebFacts (Yandex's own)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;86.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;62.4&lt;/td&gt;
&lt;td&gt;72.8&lt;/td&gt;
&lt;td&gt;83.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HardMultiQA (Yandex's own)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;67.9&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;47.2&lt;/td&gt;
&lt;td&gt;54.5&lt;/td&gt;
&lt;td&gt;65.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EduBench Russian&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;74.2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;42.9&lt;/td&gt;
&lt;td&gt;44.0&lt;/td&gt;
&lt;td&gt;67.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ExpertFactsQA Law&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;49.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;27.9&lt;/td&gt;
&lt;td&gt;24.3&lt;/td&gt;
&lt;td&gt;40.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MATH-500&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;91.1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;81.9&lt;/td&gt;
&lt;td&gt;84.8&lt;/td&gt;
&lt;td&gt;80.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LiveCodeBench v5-6&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;50.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;50.4&lt;/td&gt;
&lt;td&gt;50.4&lt;/td&gt;
&lt;td&gt;38.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BigCodeBench&lt;/td&gt;
&lt;td&gt;48.3&lt;/td&gt;
&lt;td&gt;43.5&lt;/td&gt;
&lt;td&gt;48.8&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;49.1&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TriviaQA&lt;/td&gt;
&lt;td&gt;79.0&lt;/td&gt;
&lt;td&gt;71.4&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;89.8&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;89.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMLU-Pro&lt;/td&gt;
&lt;td&gt;66.8&lt;/td&gt;
&lt;td&gt;63.2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;69.9&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;66.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LongMemEval 128k&lt;/td&gt;
&lt;td&gt;64.6&lt;/td&gt;
&lt;td&gt;55.6&lt;/td&gt;
&lt;td&gt;64.8&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern is clear. Anything that needs knowledge of the Russian-speaking world — language, law, school curriculum, local facts — it beats models 4x its active size. That's the training corpus talking. Math is strong too: HMMT Feb 2026 at 96.9 (pass@32) vs 87.9 for Qwen3.5.&lt;/p&gt;

&lt;p&gt;On English trivia and general knowledge (TriviaQA, MMLU-Pro), the bigger Nemotron and DeepSeek lead. No magic: 80B parameters can't hold as much about the world as 284B.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can you run it locally?
&lt;/h2&gt;

&lt;p&gt;3B active sounds laptop-friendly, but all 80B have to sit in memory: ~160 GB in bf16, roughly 45 GB at 4-bit by my estimate. So in theory it's 64 GB Mac territory — and Apple Silicon is where MoE models shine, since you only read a small slice of weights per token.&lt;/p&gt;

&lt;p&gt;In practice, not yet. The official recipes are transformers and vLLM on NVIDIA GPUs (the example uses four), and the KDA layers need the &lt;code&gt;flash-linear-attention&lt;/code&gt; kernels. I couldn't find GGUF or MLX quants on release day, and non-standard attention is exactly what tends to delay llama.cpp support by weeks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It's a base model.&lt;/strong&gt; It continues text; it doesn't chat. Instruction tuning is on you (or on Yandex's own future releases).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The two most impressive benchmarks are Yandex's own.&lt;/strong&gt; To their credit, WikiWebFacts and HardMultiQA were published alongside the weights with evaluation protocols — anyone can re-run them. Until someone does, I'd weight third-party benchmarks more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It's not a frontier model.&lt;/strong&gt; This is a compact model winning on data and architecture, not a GPT or Claude competitor.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Still, a fully in-house model with a permissive license, open benchmarks and a detailed tech report is a genuinely useful foundation for anyone building products in Russian — especially in legal, education and reference use cases where its lead is largest.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/alice-ai-foundation-llm/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Best Ollama Models for Coding, Writing and Medicine (September 2026)</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Sun, 20 Sep 2026 07:27:42 +0000</pubDate>
      <link>https://dev.to/klukyanov/best-ollama-models-for-coding-writing-and-medicine-september-2026-4a89</link>
      <guid>https://dev.to/klukyanov/best-ollama-models-for-coding-writing-and-medicine-september-2026-4a89</guid>
      <description>&lt;p&gt;A 0.9-billion-parameter model beats a 27-billion one — as long as you hand it a scanned document instead of a conversation. That single fact is why "top 10 local models" lists are useless: they rank by size, and size is not what you pick a model by.&lt;/p&gt;

&lt;p&gt;I spent a weekend going through the Ollama library sorted by &lt;em&gt;task&lt;/em&gt; rather than by parameter count. Here is what the landscape looks like at the end of September 2026, and — more importantly — the three rules that will still hold when every name below has been replaced.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule 1: read the tag, not the name
&lt;/h2&gt;

&lt;p&gt;Sort the Ollama library by newest and the top rows are &lt;code&gt;deepseek-v4.1-flash&lt;/code&gt;, &lt;code&gt;glm-5.3&lt;/code&gt;, &lt;code&gt;kimi-k3&lt;/code&gt;, &lt;code&gt;minimax-m3&lt;/code&gt;. All four carry a &lt;code&gt;cloud&lt;/code&gt; tag.&lt;/p&gt;

&lt;p&gt;That means &lt;code&gt;ollama run&lt;/code&gt; ships your prompt to someone else's servers. It is a great way to try a model that would never fit on your machine. It is not local inference, and if you came here for privacy or offline work, those rows get crossed out first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule 2: specialisation beats size on its own turf
&lt;/h2&gt;

&lt;p&gt;The clearest example of the month is &lt;code&gt;glm-ocr&lt;/code&gt;: 0.9B parameters, 2.2 GB on disk, and the top spot on OmniDocBench V1.5 at 94.62. Feed it a messy invoice with nested tables, footnotes and formulas and it holds. A general-purpose 27B vision model does not — nobody trained it for document structure.&lt;/p&gt;

&lt;p&gt;The right pipeline is boring and effective: &lt;code&gt;glm-ocr&lt;/code&gt; turns the document into structured text, then a general model reasons over that text.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rule 3: the benchmark is not your task
&lt;/h2&gt;

&lt;p&gt;SWE-bench, the number everyone quotes, measures fixing a bug in an unfamiliar repository from an issue description. That is very different from "write me this function" and nothing at all like "explain what this code does". A high SWE-bench score on a model you keep around for autocomplete tells you close to nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coding
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://ollama.com/library/qwen3.8" rel="noopener noreferrer"&gt;Qwen3.8-27B&lt;/a&gt; landed on 14 August: 27.78B dense parameters, Apache 2.0, 17.7 GB in the Ollama build, a 256K context window, vision included. Qwen reports 73.0 on Terminal-Bench 2.1 (up from 63.4), 61.7 on SWE-bench Pro, 90.3 on LiveCodeBench v6 and 84.3 on OSWorld-Verified, up from 63.9. Its Artificial Analysis index went up 14 points over the previous generation &lt;em&gt;on an identical architecture&lt;/em&gt; — all of it from post-training.&lt;/p&gt;

&lt;p&gt;If 18 GB of weights is out of reach, the choice is between two smaller ones:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;th&gt;Speed&lt;/th&gt;
&lt;th&gt;Strength&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;qwen3.8:27b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;dense 27B, vision, 256K&lt;/td&gt;
&lt;td&gt;~7 tok/s on a Mac mini M4&lt;/td&gt;
&lt;td&gt;terminal agent, long tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;devstral-small-2:24b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;dense 24B from Mistral, vision&lt;/td&gt;
&lt;td&gt;139.2 tok/s&lt;/td&gt;
&lt;td&gt;responsiveness, multi-file edits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;qwen3-coder:30b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;MoE, 3B active of 30&lt;/td&gt;
&lt;td&gt;90.6 tok/s&lt;/td&gt;
&lt;td&gt;quality per gigabyte&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Those throughput numbers come from Artificial Analysis on provider hardware — at home they will be lower, the ratio is what matters. One detail most reviews skip: Devstral's context costs roughly twice as much memory, about 5 GB per 32K tokens versus 3 GB for the MoE Qwen. On a machine where RAM is tight, that decides it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents: Meta is back, and that is the surprise of the month
&lt;/h2&gt;

&lt;p&gt;Three weeks ago Meta was the cautionary tale of open weights — the company that invented the strategy, down to under one percent of tokens on OpenRouter. Today &lt;a href="https://ollama.com/library/muse-glimmer" rel="noopener noreferrer"&gt;muse-glimmer&lt;/a&gt; sits in the Ollama library: 30B parameters, Apache 2.0, 18.2 GB, 128K context, and over two hundred thousand pulls in three weeks.&lt;/p&gt;

&lt;p&gt;The model card is refreshingly unambitious: tuned for "tool use, long tasks, and failure recovery". Not &lt;em&gt;smartest&lt;/em&gt; — the one that finishes. Meta reports 75.5 on MCP Atlas, 51.2 on SWE-Bench Pro and 94.7 on AIME 2026.&lt;/p&gt;

&lt;p&gt;That framing is the actual lesson. An agentic task is not one answer, it is a hundred in a row. A model that is right 90% of the time but never notices its own mistake loses to one at 85% that catches and redoes it.&lt;/p&gt;

&lt;p&gt;Next to it sits NVIDIA's &lt;code&gt;nemotron-3.5-lightning&lt;/code&gt; (11 August): 30B total, 3B active, interleaved Mamba-2 and MoE layers, 51.56 on SWE-bench Verified, 75.44 on GPQA Diamond, 81.94 on MMLU Pro. The headline 1,200 tokens per second is a provider-hardware measurement, not a laptop one — but the architecture is genuinely built to emit tokens fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing: the honest answer is still "not locally"
&lt;/h2&gt;

&lt;p&gt;On &lt;a href="https://eqbench.com/creative_writing.html" rel="noopener noreferrer"&gt;EQ-Bench Creative Writing v3&lt;/a&gt; the top of the board is closed: Claude Opus 5 at 2121 Elo, GPT-5.6 Sol at 1963. The best open-weights score belongs to Kimi K3 at 2071 — a 2.8-trillion-parameter model that needs a multi-node cluster. "Open" in the same sense an operating system's source code is open to someone without a chip fab.&lt;/p&gt;

&lt;p&gt;What &lt;em&gt;does&lt;/em&gt; work locally is working text: email, summaries, tightening a paragraph, turning a call transcript into structure. The pick there is &lt;code&gt;gemma4&lt;/code&gt;, with 25.3M pulls the most downloaded model in the library.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://habr.com/ru/articles/1033808/" rel="noopener noreferrer"&gt;Russian-language test from May&lt;/a&gt; on an RTX 5070 Ti ran Gemma 4, Qwen 3.6 and Qwen3-Coder through twelve practical tasks. Gemma 4 in fast mode scored 12/12, Qwen 3.6 got 9/12. And a finding worth more than any index: turning thinking mode &lt;em&gt;on&lt;/em&gt; made instruction-following worse — 11/12 instead of 12 for Gemma, 7 instead of 9 for Qwen.&lt;/p&gt;

&lt;h2&gt;
  
  
  Medicine: the one area where a specialised model is mandatory
&lt;/h2&gt;

&lt;p&gt;Start with the boundary, not the models. Open medical models ship under the &lt;a href="https://developers.google.com/health-ai-developer-foundations/medgemma/model-card" rel="noopener noreferrer"&gt;Health AI Developer Foundations&lt;/a&gt; terms, and Google says plainly that they are &lt;strong&gt;not clinical grade&lt;/strong&gt; and require task-specific validation. This is a building block for a developer, not a patient-facing chatbot and not a second opinion.&lt;/p&gt;

&lt;p&gt;Within that boundary they are genuinely useful. &lt;code&gt;medgemma1.5:4b&lt;/code&gt; is 3.3 GB with a 128K context and vision; version 1.5 added whole-slide histopathology, longitudinal imaging, anatomical localisation and document understanding. It scores 89.6% on EHRQA and turns a lab report into structured JSON at 91.0 macro-F1. The text-only &lt;code&gt;medgemma:27b&lt;/code&gt; reaches 87.7% on MedQA. For broad health conversations graded by physicians, &lt;code&gt;gpt-oss-120b&lt;/code&gt; leads open models on HealthBench at 0.576 — but that is 65 GB of weights.&lt;/p&gt;

&lt;p&gt;The realistic use for a private person: put your own records into readable shape, merge several documents into one table, prepare questions for an appointment — all on your own machine, without shipping medical data into someone's cloud.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieval: embeddings matter more than the generator
&lt;/h2&gt;

&lt;p&gt;The most common mistake in a homegrown RAG setup is spending all the memory on a big generator and grabbing whatever embedding model came first. It works the other way round: if retrieval returns the wrong chunks, the model on top will confidently summarise the wrong thing.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;qwen3-embedding:0.6b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.6 GB&lt;/td&gt;
&lt;td&gt;best quality per gigabyte, tunable dimensions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;qwen3-embedding:8b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4.7 GB&lt;/td&gt;
&lt;td&gt;maximum quality, 70.58 MTEB multilingual&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;bge-m3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.2 GB&lt;/td&gt;
&lt;td&gt;multilingual corpora, hybrid dense+sparse, 8K context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;embeddinggemma:300m&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.6 GB&lt;/td&gt;
&lt;td&gt;high-volume indexing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;nomic-embed-text&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.3 GB&lt;/td&gt;
&lt;td&gt;simplest start, 86M pulls&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Qwen3-Embedding family lets you set the vector dimension anywhere from 32 to 1024, which is a direct saving on your vector store — you choose the trade-off instead of inheriting it.&lt;/p&gt;

&lt;p&gt;For the generator on top, something modest does the job. IBM's &lt;code&gt;granite4.2&lt;/code&gt; (3B/8B/30B, Apache 2.0, 128K) was tuned for exactly this: RAG, tool calls and structured JSON output.&lt;/p&gt;

&lt;h2&gt;
  
  
  The table
&lt;/h2&gt;

&lt;p&gt;Sizes below are not approximations — I pulled them from the Ollama registry manifests, so this is the sum of layers you will actually download, as of 20 September 2026. In memory it will be more, by the size of your context.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Command&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal agent, long coding tasks&lt;/td&gt;
&lt;td&gt;Qwen3.8-27B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull qwen3.8:27b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;17.7 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fast coding assistant&lt;/td&gt;
&lt;td&gt;Devstral Small 2&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull devstral-small-2:24b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;15.2 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding on tight memory&lt;/td&gt;
&lt;td&gt;Qwen3-Coder 30B-A3B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull qwen3-coder:30b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;18.6 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool-using agent&lt;/td&gt;
&lt;td&gt;Muse Glimmer (Meta)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull muse-glimmer:30b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;18.2 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Many small steps, speed matters&lt;/td&gt;
&lt;td&gt;Nemotron 3.5 Lightning&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull nemotron-3.5-lightning:30b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;25.4 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Writing, editing, email&lt;/td&gt;
&lt;td&gt;Gemma 4&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull gemma4:12b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;7.6 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same on weak hardware&lt;/td&gt;
&lt;td&gt;Gemma 4 E2B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull gemma4:e2b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;7.2 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medical documents and imaging&lt;/td&gt;
&lt;td&gt;MedGemma 1.5&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull medgemma1.5:4b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3.3 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medical text only&lt;/td&gt;
&lt;td&gt;MedGemma 27B&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull medgemma:27b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;17.4 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scans, PDFs, tables, formulas&lt;/td&gt;
&lt;td&gt;GLM-OCR&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull glm-ocr&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2.2 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Images on weak hardware&lt;/td&gt;
&lt;td&gt;MiniCPM-V 4.6&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull minicpm-v4.6:1b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.6 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Search over your own documents&lt;/td&gt;
&lt;td&gt;Qwen3-Embedding&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull qwen3-embedding:0.6b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.6 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same, multilingual corpus&lt;/td&gt;
&lt;td&gt;BGE-M3&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull bge-m3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.2 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answers over documents, strict JSON&lt;/td&gt;
&lt;td&gt;Granite 4.2&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ollama pull granite4.2:8b&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5.3 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What is deliberately missing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Translation.&lt;/strong&gt; I meant to write that section, but the honest answer is that the best model in the niche is closed. Qwen3.8-LiveTranslate, released 19 September, does simultaneous interpretation at 2.3 seconds of lag across 60 input languages — API only, no weights. Locally there is nothing but general models that translate "fine".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ornith-1.5.&lt;/strong&gt; Sizes of 9B, 35B and 397B, 256K context, a self-improvement loop that generates its own training tasks, three hundred thousand pulls in a month. The most interesting row in the library, and exactly why it is not in my list: there are no comparable independent numbers yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually run
&lt;/h2&gt;

&lt;p&gt;On a 16 GB MacBook: &lt;code&gt;gemma4:12b&lt;/code&gt; as the generalist, plus &lt;code&gt;glm-ocr&lt;/code&gt; and &lt;code&gt;qwen3-embedding:0.6b&lt;/code&gt; for documents and search over my own archive. Everything heavier is a different machine or a rented endpoint.&lt;/p&gt;

&lt;p&gt;If I had to keep exactly one model forever, it would be Gemma 4 at whatever size fits. Not because it wins benchmarks — it does not — but because it does what you asked instead of what it found more interesting. Twenty-five million pulls is a lot of people arriving at the same conclusion independently.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/ollama-models-by-task/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>ollama</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Running Flux locally on a Mac: install, commands and two models compared</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Sun, 20 Sep 2026 05:55:52 +0000</pubDate>
      <link>https://dev.to/klukyanov/running-flux-locally-on-a-mac-install-commands-and-two-models-compared-4h9l</link>
      <guid>https://dev.to/klukyanov/running-flux-locally-on-a-mac-install-commands-and-two-models-compared-4h9l</guid>
      <description>&lt;p&gt;People write about local image generation in one of two ways. Either "it installs in three clicks, why pay for anything," or "it doesn't work on a laptop." Both are equally useless when you're sitting in front of an empty terminal and just want a picture out of it.&lt;/p&gt;

&lt;p&gt;So here's the manual instead. Every command below was actually run on this machine, every number comes from the log of a specific run: Apple M5, 16 GB of unified memory, macOS 26.5, &lt;code&gt;stable-diffusion.cpp&lt;/code&gt; build &lt;code&gt;master-650&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;We'll install two models from the same family, and the difference between them is the whole point. &lt;strong&gt;Flux.1-schnell&lt;/strong&gt; draws a frame from scratch out of text. &lt;strong&gt;Flux.1-Kontext-dev&lt;/strong&gt; takes a finished picture and changes exactly what you asked for, leaving the rest alone. Different jobs — one doesn't replace the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why keep a model on your own machine
&lt;/h2&gt;

&lt;p&gt;An honest list of exactly three reasons, because there is no fourth one.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Free.&lt;/strong&gt; There is no spend counter at all. Trying thirty variations of a scene costs you time and nothing else. That changes behaviour: with a cloud API you think before you press enter, locally you just run batches in the background.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offline.&lt;/strong&gt; No internet needed at any point once the weights are downloaded. Planes, cabins, dead VPNs, corporate networks that block everything — the model doesn't care.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Private.&lt;/strong&gt; The picture never leaves your machine. This is the one argument the cloud cannot beat: NDA material, someone else's mockups, internal screenshots.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And now the part these articles usually skip. &lt;strong&gt;Local will not be faster than the cloud.&lt;/strong&gt; In an earlier benchmark on identical prompts, local Flux produced a frame in 172 seconds and GPT Image in 167 — and the cloud model was visibly better at composite scenes. Local wins on cost, autonomy and privacy. Not on speed, and not on quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1. Build the engine
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;xcode-select &lt;span class="nt"&gt;--install&lt;/span&gt;
brew &lt;span class="nb"&gt;install &lt;/span&gt;cmake

git clone &lt;span class="nt"&gt;--recursive&lt;/span&gt; https://github.com/leejet/stable-diffusion.cpp ~/Tools/stable-diffusion.cpp
&lt;span class="nb"&gt;cd&lt;/span&gt; ~/Tools/stable-diffusion.cpp
&lt;span class="nb"&gt;mkdir &lt;/span&gt;build &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;build
cmake .. &lt;span class="nt"&gt;-DSD_METAL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;ON
cmake &lt;span class="nt"&gt;--build&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;--config&lt;/span&gt; Release &lt;span class="nt"&gt;-j&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two places where people trip.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;--recursive&lt;/code&gt; is mandatory.&lt;/strong&gt; The &lt;code&gt;ggml&lt;/code&gt; compute core is a submodule; without it the build dies on missing headers. Already cloned without it? &lt;code&gt;git submodule update --init --recursive&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;-DSD_METAL=ON&lt;/code&gt; is what turns the GPU on.&lt;/strong&gt; The &lt;code&gt;.cpp&lt;/code&gt; suffix in the project name is misleading — it reads like "the CPU version, therefore slow." It isn't. The first line of the log says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ggml_metal_device_init: GPU name: MTL0 (Apple M5)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The GPU does the work. The CPU gets a separate and fairly humiliating job: cleaning up after what's broken in Metal. More on that below, and it's the most important section here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2. Download the weights
&lt;/h2&gt;

&lt;p&gt;Flux is not one file. It's four, and all four are required.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;flux1-schnell-Q4_K_S.gguf&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;6.3 GB&lt;/td&gt;
&lt;td&gt;the diffusion model itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;t5xxl-Q4_K_M.gguf&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2.7 GB&lt;/td&gt;
&lt;td&gt;T5 text encoder — parses long phrasing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;clip_l.safetensors&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;235 MB&lt;/td&gt;
&lt;td&gt;second text encoder, CLIP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ae.safetensors&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;320 MB&lt;/td&gt;
&lt;td&gt;VAE — turns the result into pixels&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two encoders is a Flux architecture thing: CLIP gives the general mood of the prompt, T5 reads the wording word by word. That's also where its strength comes from — you write long descriptive English sentences, not comma-separated tags.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ~/Tools/sd-models/flux &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; ~/Tools/sd-models/flux

curl &lt;span class="nt"&gt;-L&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; - &lt;span class="nt"&gt;-o&lt;/span&gt; flux1-schnell-Q4_K_S.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  https://huggingface.co/city96/FLUX.1-schnell-gguf/resolve/main/flux1-schnell-Q4_K_S.gguf

curl &lt;span class="nt"&gt;-L&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; - &lt;span class="nt"&gt;-o&lt;/span&gt; t5xxl-Q4_K_M.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  https://huggingface.co/city96/t5-v1_1-xxl-encoder-gguf/resolve/main/t5-v1_1-xxl-encoder-Q4_K_M.gguf

curl &lt;span class="nt"&gt;-L&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; - &lt;span class="nt"&gt;-o&lt;/span&gt; clip_l.safetensors &lt;span class="se"&gt;\&lt;/span&gt;
  https://huggingface.co/comfyanonymous/flux_text_encoders/resolve/main/clip_l.safetensors

curl &lt;span class="nt"&gt;-L&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; - &lt;span class="nt"&gt;-o&lt;/span&gt; ae.safetensors &lt;span class="se"&gt;\&lt;/span&gt;
  https://huggingface.co/second-state/FLUX.1-schnell-GGUF/resolve/main/ae.safetensors
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;About the VAE: the obvious place to get it would be the official Black Forest Labs repo, but it's license-gated and returns 401 without a token and an accepted agreement. Hence the &lt;code&gt;second-state&lt;/code&gt; mirror — same file, byte for byte.&lt;/p&gt;

&lt;h3&gt;
  
  
  The gotcha that costs you an evening
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;curl&lt;/code&gt; can break in the middle of a large file and &lt;strong&gt;exit with code 0&lt;/strong&gt;. The file looks complete, the &lt;code&gt;GGUF&lt;/code&gt; header is there, and the model simply won't load. The error you get is unhelpful:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read tensor data failed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Check the size against the header the server sends:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;https://huggingface.co/city96/FLUX.1-schnell-gguf/resolve/main/flux1-schnell-Q4_K_S.gguf
&lt;span class="nv"&gt;EXPECTED&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;curl &lt;span class="nt"&gt;-sIL&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$URL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'BEGIN{IGNORECASE=1} /^content-length:/{l=$2} END{print l+0}'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;ACTUAL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;stat&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt;%z flux1-schnell-Q4_K_S.gguf&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ACTUAL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-eq&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$EXPECTED&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"OK"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"TRUNCATED: &lt;/span&gt;&lt;span class="nv"&gt;$ACTUAL&lt;/span&gt;&lt;span class="s2"&gt; of &lt;/span&gt;&lt;span class="nv"&gt;$EXPECTED&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;-C -&lt;/code&gt; resumes from where it broke, so re-running loses nothing. If your connection drops regularly, add &lt;code&gt;--speed-limit 51200 --speed-time 30&lt;/code&gt;: that pair kills a dead connection after half a minute instead of hanging on a 30-minute timeout.&lt;/p&gt;

&lt;p&gt;One more thing that isn't about the tooling: &lt;strong&gt;don't generate while the weights are downloading.&lt;/strong&gt; The model pushes memory into swap, and under swap the network on this machine dies — HuggingFace starts serving kilobytes per second, which then gets blamed on "bad internet."&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3. The first picture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;~/Tools/stable-diffusion.cpp/build/bin/sd-cli &lt;span class="nt"&gt;-M&lt;/span&gt; img_gen &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--diffusion-model&lt;/span&gt; ~/Tools/sd-models/flux/flux1-schnell-Q4_K_S.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--vae&lt;/span&gt;        ~/Tools/sd-models/flux/ae.safetensors &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--clip_l&lt;/span&gt;     ~/Tools/sd-models/flux/clip_l.safetensors &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--t5xxl&lt;/span&gt;      ~/Tools/sd-models/flux/t5xxl-Q4_K_M.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"A vintage green enamel mug standing on a weathered wooden windowsill, warm morning light, a rainy blurred street outside the window, photorealistic, 50mm lens, shallow depth of field, no legible text anywhere"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cfg-scale&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--sampling-method&lt;/span&gt; euler &lt;span class="nt"&gt;--steps&lt;/span&gt; 4 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-W&lt;/span&gt; 1024 &lt;span class="nt"&gt;-H&lt;/span&gt; 1024 &lt;span class="nt"&gt;--seed&lt;/span&gt; 42 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--vae-on-cpu&lt;/span&gt; &lt;span class="nt"&gt;--diffusion-fa&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-o&lt;/span&gt; ~/Desktop/first.png
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two flags you must not touch.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;--cfg-scale&lt;/code&gt; stays at 1.0
&lt;/h3&gt;

&lt;p&gt;The 7.0 that floats around every Stable Diffusion tutorial turns the picture into coloured noise here. Schnell is a distilled model: it was trained to hit the target in a handful of steps without a guidance mechanism. At 1.0 that mechanism is effectively off, which is the correct mode.&lt;/p&gt;

&lt;p&gt;Same root cause for the second surprise: &lt;strong&gt;negative prompts do nothing.&lt;/strong&gt; Not "work poorly" — they physically don't participate in the computation. "No people in frame" has to become "an empty street."&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;code&gt;--vae-on-cpu&lt;/code&gt; is the most expensive thing I know about this stack
&lt;/h3&gt;

&lt;p&gt;Without it you get a &lt;strong&gt;blank white frame&lt;/strong&gt;. Not an error, not a crash — a white rectangle, and the log cheerfully says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;save result image (success)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Flux's VAE decoder returns NaN on Metal and the result silently collapses. You can spot it instantly by file size: &lt;strong&gt;a broken PNG is about 17 KB, a live one is 450 KB and up&lt;/strong&gt; (mine came out at 1.7–1.9 MB). If you see seventeen kilobytes, don't rewrite the prompt and don't re-download the weights — just add the flag.&lt;/p&gt;

&lt;p&gt;The flag moves decoding to the CPU. That isn't free: on a 1024×1024 frame the CPU VAE costs about 58 seconds, and those 58 seconds are in every single picture no matter how many steps you run.&lt;/p&gt;

&lt;h2&gt;
  
  
  How many steps you actually need
&lt;/h2&gt;

&lt;p&gt;The common "put it on twenty steps for more detail" advice is not just useless for schnell, it's actively wrong. Three runs, same prompt, same seed, only the step count differs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Steps&lt;/th&gt;
&lt;th&gt;Sampling&lt;/th&gt;
&lt;th&gt;VAE on CPU&lt;/th&gt;
&lt;th&gt;Total&lt;/th&gt;
&lt;th&gt;Incl. loading weights&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;26.4 s&lt;/td&gt;
&lt;td&gt;57.3 s&lt;/td&gt;
&lt;td&gt;84.2 s&lt;/td&gt;
&lt;td&gt;99 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;105.2 s&lt;/td&gt;
&lt;td&gt;58.9 s&lt;/td&gt;
&lt;td&gt;164.5 s&lt;/td&gt;
&lt;td&gt;185 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;210.6 s&lt;/td&gt;
&lt;td&gt;59.0 s&lt;/td&gt;
&lt;td&gt;270.0 s&lt;/td&gt;
&lt;td&gt;284 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three things fall out of this table, none of them obvious up front.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A step costs exactly 26.3 seconds, linearly.&lt;/strong&gt; No warm-up: the first step costs the same as the eighth. So you can do the arithmetic in your head — 58 seconds plus 26 per step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;At one step, two thirds of the time isn't drawing.&lt;/strong&gt; 57 seconds out of 84 is the CPU VAE, i.e. working around the Metal bug. On fast modes you're mostly paying to patch Metal, not to run the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Eight steps don't give you "the same, but better."&lt;/strong&gt; They give you a &lt;em&gt;different frame&lt;/em&gt;: different angle, different light, different street outside the window. For a distilled model the step count is another dimension of the scene, like the seed. So "let me add steps to fix that weird handle" doesn't work — you'll get a different mug.&lt;/p&gt;

&lt;p&gt;Practical default: &lt;strong&gt;4 steps for real work, 1 step for browsing ideas.&lt;/strong&gt; One step in a minute and a half already gives a usable frame; run a dozen concepts that way, then re-run the good one at 4 steps with the same seed.&lt;/p&gt;

&lt;p&gt;One more thing from the 8-step frame: the model decorated the buildings with shop signs whose letters are not letters. &lt;strong&gt;Flux renders plausible gibberish instead of text&lt;/strong&gt;, and the more steps, the more eagerly. Put &lt;code&gt;no legible text anywhere&lt;/code&gt; in the prompt — and if you genuinely need readable text in the image, a local model is the wrong tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second model: editing a finished frame
&lt;/h2&gt;

&lt;p&gt;Here's the reason to keep a second model on disk at all. A generator has a built-in limitation you can't prompt your way around.&lt;/p&gt;

&lt;p&gt;Picture this: the frame came out well, you like everything, but you'd like the mug replaced with a cactus. Asking the generator for that is pointless — it will draw a &lt;em&gt;new picture&lt;/em&gt;: different light, different window, different street. You can burn an hour on retries and never get the same scene back, because "the same scene" doesn't exist in its world.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Flux.1-Kontext-dev&lt;/strong&gt; works differently: it takes a finished image plus an instruction about what to change. Same command as before with three differences — a different model file, a &lt;code&gt;-r&lt;/code&gt; flag pointing at the source image, and a prompt in the imperative:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;~/Tools/stable-diffusion.cpp/build/bin/sd-cli &lt;span class="nt"&gt;-M&lt;/span&gt; img_gen &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--diffusion-model&lt;/span&gt; ~/Tools/sd-models/flux-kontext/flux1-kontext-dev-Q4_K_S.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--vae&lt;/span&gt;        ~/Tools/sd-models/flux/ae.safetensors &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--clip_l&lt;/span&gt;     ~/Tools/sd-models/flux/clip_l.safetensors &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--t5xxl&lt;/span&gt;      ~/Tools/sd-models/flux/t5xxl-Q4_K_M.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-r&lt;/span&gt; first.png &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"Replace the green enamel mug with a small potted cactus in a terracotta pot, keep everything else exactly the same"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cfg-scale&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--guidance&lt;/span&gt; 2.5 &lt;span class="nt"&gt;--sampling-method&lt;/span&gt; euler &lt;span class="nt"&gt;--steps&lt;/span&gt; 8 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-W&lt;/span&gt; 1024 &lt;span class="nt"&gt;-H&lt;/span&gt; 1024 &lt;span class="nt"&gt;--seed&lt;/span&gt; 42 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--vae-on-cpu&lt;/span&gt; &lt;span class="nt"&gt;--clip-on-cpu&lt;/span&gt; &lt;span class="nt"&gt;--diffusion-fa&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-o&lt;/span&gt; cactus.png
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Kontext weights live separately, the encoders and VAE are reused — it's exactly one more file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ~/Tools/sd-models/flux-kontext &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; ~/Tools/sd-models/flux-kontext

curl &lt;span class="nt"&gt;-L&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; - &lt;span class="nt"&gt;-o&lt;/span&gt; flux1-kontext-dev-Q4_K_S.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  https://huggingface.co/QuantStack/FLUX.1-Kontext-dev-GGUF/resolve/main/flux1-kontext-dev-Q4_K_S.gguf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same gated-repo trap here: the &lt;code&gt;city96&lt;/code&gt; Kontext build needs an accepted license and returns 401 without a token. The &lt;code&gt;QuantStack&lt;/code&gt; mirror is open.&lt;/p&gt;

&lt;p&gt;The mug became a cactus, and the street behind the glass, the frame, the raindrops, the windowsill planks and the light all stayed exactly as they were. None of those were mentioned in the prompt — the model understood it wasn't asked to touch them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Parameter&lt;/th&gt;
&lt;th&gt;schnell&lt;/th&gt;
&lt;th&gt;Kontext&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Source image&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;&lt;code&gt;-r file.png&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--guidance&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;unused&lt;/td&gt;
&lt;td&gt;2.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--steps&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1–4&lt;/td&gt;
&lt;td&gt;8–20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--clip-on-cpu&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;not needed&lt;/td&gt;
&lt;td&gt;recommended&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt&lt;/td&gt;
&lt;td&gt;describes the frame&lt;/td&gt;
&lt;td&gt;says what to change&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Kontext is not distilled&lt;/strong&gt; — unlike schnell it's an ordinary model, so it wants both &lt;code&gt;--guidance&lt;/code&gt; and more steps. Eight is the working minimum; complex edits are worth twenty.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;--clip-on-cpu&lt;/code&gt; pushes the text encoders into regular RAM.&lt;/strong&gt; Kontext's diffusion part is heavier, and on 16 GB that flag decides whether everything fits at once: it splits the load into 6.5 GB on the GPU and 3.5 GB in RAM instead of nearly ten in one place.&lt;/p&gt;

&lt;h3&gt;
  
  
  It preserves the scene, not the pixels
&lt;/h3&gt;

&lt;p&gt;It looks like the model neatly cut out the mug and pasted a cactus. In reality the whole frame is re-rendered, just with the original in view. I measured the difference pixel by pixel:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Difference&lt;/th&gt;
&lt;th&gt;Share of the frame&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&amp;gt; 8 levels&lt;/td&gt;
&lt;td&gt;81.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&amp;gt; 16 levels&lt;/td&gt;
&lt;td&gt;33.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&amp;gt; 32 levels&lt;/td&gt;
&lt;td&gt;11.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The "untouched" background is in fact touched — you just can't see it. Practical consequence: &lt;strong&gt;you cannot layer a Kontext result over the original&lt;/strong&gt;, and you cannot use it to edit a photo where pixel authenticity matters. For an article cover it's irrelevant. For a document or someone else's photo it isn't.&lt;/p&gt;

&lt;h3&gt;
  
  
  What it costs in time
&lt;/h3&gt;

&lt;p&gt;Editing is more expensive than generating, for an obvious reason: the model has to read the source image first, not just the text. One run, 1024×1024, eight steps:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;th&gt;Where&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Loading weights&lt;/td&gt;
&lt;td&gt;14.9 s&lt;/td&gt;
&lt;td&gt;disk → memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reading the source image&lt;/td&gt;
&lt;td&gt;30.9 s&lt;/td&gt;
&lt;td&gt;CPU, same VAE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Parsing the prompt&lt;/td&gt;
&lt;td&gt;9.2 s&lt;/td&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sampling, 8 steps&lt;/td&gt;
&lt;td&gt;469.2 s&lt;/td&gt;
&lt;td&gt;GPU, 58.7 s/step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decoding the result&lt;/td&gt;
&lt;td&gt;51.0 s&lt;/td&gt;
&lt;td&gt;CPU, forced&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;561 s ≈ 9.5 min&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nine and a half minutes per edit is not "playing around." It's a go-make-coffee workflow, and you should plan around that.&lt;/p&gt;

&lt;p&gt;The result is reproducible, though. I ran the same request twice: the first time the machine was thrashing and each step took two minutes, the second time memory was free. &lt;strong&gt;Both runs produced a byte-identical PNG.&lt;/strong&gt; Machine load affects the time, not the picture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Q4 vs Q6: are three gigabytes worth it
&lt;/h2&gt;

&lt;p&gt;The file name tells you how hard the model was squeezed. &lt;strong&gt;Quantization&lt;/strong&gt; stores weights at reduced precision: Q4 gives each number roughly four bits instead of sixteen, Q6 roughly six. Higher number, closer to the original, fatter file.&lt;/p&gt;

&lt;p&gt;For Kontext the gap is three gigabytes, and on a 16 GB machine that isn't an abstraction. I ran both on the same picture with the same prompt and seed.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Quant&lt;/th&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Memory for the model&lt;/th&gt;
&lt;th&gt;Per step&lt;/th&gt;
&lt;th&gt;Total, 8 steps&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Q4_K_S&lt;/td&gt;
&lt;td&gt;6.3 GB&lt;/td&gt;
&lt;td&gt;9.8 GB&lt;/td&gt;
&lt;td&gt;58.7 s&lt;/td&gt;
&lt;td&gt;561 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q6_K&lt;/td&gt;
&lt;td&gt;9.2 GB&lt;/td&gt;
&lt;td&gt;12.6 GB&lt;/td&gt;
&lt;td&gt;57.0 s&lt;/td&gt;
&lt;td&gt;605 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;You can't tell the outputs apart by eye, and that's measurable, not a feeling:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SSIM — &lt;strong&gt;0.983&lt;/strong&gt; out of 1.0&lt;/li&gt;
&lt;li&gt;PSNR — &lt;strong&gt;33.9 dB&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1% of pixels&lt;/strong&gt; differ by more than 16 levels; mean difference across the frame is 1.7 levels out of 255&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For comparison, the edit itself changed a third of the frame. &lt;strong&gt;The difference between quants is thirty times smaller than the difference you asked for.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Speed, interestingly, is a wash too: 58.7 s/step for Q4 against 57.0 for Q6 is run-to-run noise, not an advantage. The extra three gigabytes don't slow the math down. They eat memory — and that's the real price.&lt;/p&gt;

&lt;p&gt;So on 16 GB the answer is simple: &lt;strong&gt;take Q4.&lt;/strong&gt; Three gigabytes of disk is the lesser problem; the worse one is that those same three gigabytes move the machine closer to swap. And swap on these models costs you multiples, not percentages: I left a second generation process running by accident, and a 58-second step turned into 142.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheat sheet
&lt;/h2&gt;

&lt;p&gt;Two shell functions save you from retyping the long command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;FLUX&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;~/Tools/sd-models/flux
&lt;span class="nv"&gt;SD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;~/Tools/stable-diffusion.cpp/build/bin/sd-cli

&lt;span class="c"&gt;# generate from scratch: img "prompt" [steps]&lt;/span&gt;
img&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="nv"&gt;$SD&lt;/span&gt; &lt;span class="nt"&gt;-M&lt;/span&gt; img_gen &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--diffusion-model&lt;/span&gt; &lt;span class="nv"&gt;$FLUX&lt;/span&gt;/flux1-schnell-Q4_K_S.gguf &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--vae&lt;/span&gt; &lt;span class="nv"&gt;$FLUX&lt;/span&gt;/ae.safetensors &lt;span class="nt"&gt;--clip_l&lt;/span&gt; &lt;span class="nv"&gt;$FLUX&lt;/span&gt;/clip_l.safetensors &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--t5xxl&lt;/span&gt; &lt;span class="nv"&gt;$FLUX&lt;/span&gt;/t5xxl-Q4_K_M.gguf &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--cfg-scale&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--sampling-method&lt;/span&gt; euler &lt;span class="nt"&gt;--steps&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;2&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;4&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-W&lt;/span&gt; 1024 &lt;span class="nt"&gt;-H&lt;/span&gt; 1024 &lt;span class="nt"&gt;--vae-on-cpu&lt;/span&gt; &lt;span class="nt"&gt;--diffusion-fa&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-o&lt;/span&gt; ~/Desktop/img_&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%H%M%S&lt;span class="si"&gt;)&lt;/span&gt;.png
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;# edit a finished frame: edit source.png "what to change"&lt;/span&gt;
edit&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="nv"&gt;$SD&lt;/span&gt; &lt;span class="nt"&gt;-M&lt;/span&gt; img_gen &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--diffusion-model&lt;/span&gt; ~/Tools/sd-models/flux-kontext/flux1-kontext-dev-Q4_K_S.gguf &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--vae&lt;/span&gt; &lt;span class="nv"&gt;$FLUX&lt;/span&gt;/ae.safetensors &lt;span class="nt"&gt;--clip_l&lt;/span&gt; &lt;span class="nv"&gt;$FLUX&lt;/span&gt;/clip_l.safetensors &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--t5xxl&lt;/span&gt; &lt;span class="nv"&gt;$FLUX&lt;/span&gt;/t5xxl-Q4_K_M.gguf &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$2&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--cfg-scale&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--guidance&lt;/span&gt; 2.5 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--sampling-method&lt;/span&gt; euler &lt;span class="nt"&gt;--steps&lt;/span&gt; 8 &lt;span class="nt"&gt;-W&lt;/span&gt; 1024 &lt;span class="nt"&gt;-H&lt;/span&gt; 1024 &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--vae-on-cpu&lt;/span&gt; &lt;span class="nt"&gt;--clip-on-cpu&lt;/span&gt; &lt;span class="nt"&gt;--diffusion-fa&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-o&lt;/span&gt; ~/Desktop/edit_&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%H%M%S&lt;span class="si"&gt;)&lt;/span&gt;.png
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And one habit that saves hours — check swap before starting a heavy model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;sysctl &lt;span class="nt"&gt;-n&lt;/span&gt; vm.swapusage
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it's full, don't hunt for speed by tuning steps and resolution. Let memory free up first. On 16 GB of unified memory the rule is blunt but accurate: &lt;strong&gt;either the model, or everything else.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually use it for
&lt;/h2&gt;

&lt;p&gt;Local Flux isn't a replacement for the cloud, it's a separate tool with its own niche. It's not faster and not better — it's free, autonomous, and it shows your pictures to nobody.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Drafts.&lt;/strong&gt; Checking whether a cover concept reads at all — ten variants at one step each while the tea brews. The one that works goes to a cloud model for the final.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edits.&lt;/strong&gt; "Keep everything, just change this" requests go to Kontext, because a generator answers those with a new picture.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anything that can't leave the machine.&lt;/strong&gt; No alternative here.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fifteen minutes to build, an evening to download, 17 GB on disk — and after that it just works, with no counter and no internet.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/flux-mac-setup/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>macos</category>
      <category>tutorial</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Local generation on a Mac: where it is actually free, and where it costs two hours per second</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Fri, 18 Sep 2026 09:04:44 +0000</pubDate>
      <link>https://dev.to/klukyanov/local-generation-on-a-mac-where-it-is-actually-free-and-where-it-costs-two-hours-per-second-3aol</link>
      <guid>https://dev.to/klukyanov/local-generation-on-a-mac-where-it-is-actually-free-and-where-it-costs-two-hours-per-second-3aol</guid>
      <description>&lt;p&gt;Two hours and fourteen minutes of compute for one second of video. That is not a joke about running things on the CPU — it is a number from a log.&lt;/p&gt;

&lt;p&gt;I spent a week putting the whole local generation stack through its paces on a single machine: Apple M5, 16 GB of unified memory, macOS 26.5, a locally built &lt;code&gt;stable-diffusion.cpp&lt;/code&gt;. Stills, editing existing frames, video. The goal was to find out where "just run it locally, it's free" is honest advice and where it stops being an option at all.&lt;/p&gt;

&lt;p&gt;Every number below comes from the log of a specific run.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, kill the CPU myth
&lt;/h2&gt;

&lt;p&gt;Because of the &lt;code&gt;.cpp&lt;/code&gt; in the name, &lt;code&gt;stable-diffusion.cpp&lt;/code&gt; gets filed under "the slow CPU version". The first line of the log says otherwise:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ggml_metal_device_init: GPU name: MTL0 (Apple M5)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the ggml Metal backend. The work runs on the GPU. The CPU gets one job here, and it is a humiliating one: finishing what is broken on Metal. That happens twice in this article.&lt;/p&gt;

&lt;p&gt;The second thing that shapes everything: on Apple Silicon, memory is shared. A model has no private VRAM — it takes system memory, the same pool your browser and editor live in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;total params memory size = 9959.28MB (VRAM 9864.71MB, RAM 94.57MB)
  text_encoders     3395.09MB (VRAM)
  diffusion_model   6469.62MB (VRAM)
  vae                 94.57MB (RAM)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ten gigabytes out of sixteen for one model. Everything else on the machine splits the rest. Most of what follows grows out of that single fact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stills: works, and genuinely free
&lt;/h2&gt;

&lt;p&gt;Flux.1-schnell, 4-bit quant, 1024×512, four steps:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;th&gt;Where&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Loading weights from disk&lt;/td&gt;
&lt;td&gt;~15 s&lt;/td&gt;
&lt;td&gt;disk → memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sampling, 4 steps&lt;/td&gt;
&lt;td&gt;46.83 s&lt;/td&gt;
&lt;td&gt;GPU, 11.70 s/step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;VAE decode&lt;/td&gt;
&lt;td&gt;24.36 s&lt;/td&gt;
&lt;td&gt;CPU, not by choice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total, launch to PNG&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;87 s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Ninety seconds per frame is a working pace. Not instant, but fast enough to iterate on composition without watching a billing counter. schnell is distilled — four steps is what it was trained for, cranking it to twenty buys nothing.&lt;/p&gt;

&lt;p&gt;Now the trap that makes people conclude local generation is broken. One flag:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--vae-on-cpu&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without it, the Flux VAE decoder produces NaN on Metal and you get an empty white frame. The nasty part: there is no error in the log. It cheerfully prints &lt;code&gt;save result image (success)&lt;/code&gt;. You identify it by file size — a broken PNG is about 17 KB, a real one starts at 450 KB. The flag moves decoding to the CPU, and that is exactly the 24 seconds in the table above. A third of total generation time goes into working around a bug, and that is still the good news in this article.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 29x memory cliff
&lt;/h2&gt;

&lt;p&gt;Flux generates from scratch. Kontext, from the same family, edits: it takes an existing image and changes what you asked while keeping the rest. In practice it covers the one request a generator cannot serve — "keep everything, move this one thing." Ask a generator that and you get a new picture: different light, different furniture, different angle.&lt;/p&gt;

&lt;p&gt;Then came the most instructive measurement of the week. Identical run — same model, same prompt, 768×432, eight steps, same flags — on the same machine:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Machine state&lt;/th&gt;
&lt;th&gt;Per step&lt;/th&gt;
&lt;th&gt;Eight steps&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Swap at 8.5 GB of 9.2 (after a series of Flux runs)&lt;/td&gt;
&lt;td&gt;445 s&lt;/td&gt;
&lt;td&gt;~1 hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Swap collapsed, 6+ GB free&lt;/td&gt;
&lt;td&gt;15.36 s&lt;/td&gt;
&lt;td&gt;~2 minutes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Twenty-nine times, with nothing changed but the state of memory. Once the model runs out of physical RAM, every step starts going to disk and stops keeping up.&lt;/p&gt;

&lt;p&gt;The practical rule is dull but saves hours: check &lt;code&gt;sysctl -n vm.swapusage&lt;/code&gt; before starting a heavy model, and if swap is full, let memory settle instead of tuning steps and resolution. I spent two days convinced Kontext was hopelessly slow. It was queuing behind its own swap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Video: three walls in a row
&lt;/h2&gt;

&lt;p&gt;Then I tried to animate finished frames locally — even rough drafts would do. I took the smallest sane video model available, deliberately one that fits in memory. Three walls followed, each behind the previous one.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Wall&lt;/th&gt;
&lt;th&gt;What the log says&lt;/th&gt;
&lt;th&gt;Way around&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Quantized weights won't load&lt;/td&gt;
&lt;td&gt;&lt;code&gt;invalid number of dimensions: 5 &amp;gt; 4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;official safetensors, 4.1 GB instead of 1.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;VAE fails on Metal&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;unsupported op&lt;/code&gt; in &lt;code&gt;WanVAERunner::_compute&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;--vae-on-cpu&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Denoiser fails on Metal&lt;/td&gt;
&lt;td&gt;&lt;code&gt;unsupported op 'IM2COL_3D'&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Wall one: 5 &amp;gt; 4.&lt;/strong&gt; A video model convolves over time as well as width and height, so its first convolution kernel is a five-dimensional tensor. In ggml, the maximum tensor rank is four — a constant in the library core, not a build option:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;gguf_init_from_file_ptr: tensor 'patch_embedding.weight'
  has invalid number of dimensions: 5 &amp;gt; 4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I assumed a bad conversion and downloaded a quant from a different repository by a different author. Identical failure — which is the signal that the file is not the problem. Official safetensors load fine through a different code path, at the cost of size: 4.1 GB instead of 1.7, which defeats the point of quantizing.&lt;/p&gt;

&lt;p&gt;Useful corollary: updating the engine for this error is pointless. The limit is architectural.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Walls two and three: two unsupported ops.&lt;/strong&gt; The model loaded, then the VAE encoder died with &lt;code&gt;unsupported op&lt;/code&gt; — fixed by the same &lt;code&gt;--vae-on-cpu&lt;/code&gt; flag as Flux, for an entirely different reason. Then the denoiser died for good:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ggml_metal_op_encode_impl: error: unsupported op 'IM2COL_3D'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before blaming an old build, go read the source — which I did. &lt;code&gt;IM2COL_3D&lt;/code&gt; does not exist in the Metal backend of either the fork &lt;code&gt;stable-diffusion.cpp&lt;/code&gt; builds on or upstream ggml inside llama.cpp: zero mentions in &lt;code&gt;ggml-metal-device.cpp&lt;/code&gt; and &lt;code&gt;ggml-metal-ops.cpp&lt;/code&gt; in both repositories. The operation the model cannot run without was simply never written for Metal. Not "misconfigured", not "needs an update" — absent.&lt;/p&gt;

&lt;h2&gt;
  
  
  The CPU fallback, priced
&lt;/h2&gt;

&lt;p&gt;That leaves &lt;code&gt;--backend cpu&lt;/code&gt;: compute everything on the processor, around Metal and its holes. I ran it to close the question with a number instead of an opinion. Minimal task: one second of video, 512×288, 17 frames, 8 steps.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reference encode (VAE)&lt;/td&gt;
&lt;td&gt;77 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sampling, 8 steps&lt;/td&gt;
&lt;td&gt;7728 s — 256 to 1079 s per step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decode (VAE)&lt;/td&gt;
&lt;td&gt;115 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total, one second of video&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8027 s = 2h 14m&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Note the spread: four to eighteen minutes per step, slower toward the end as the machine sinks deeper into swap.&lt;/p&gt;

&lt;p&gt;The output was worse than the time. Composition drifted from the reference, objects disappeared and new ones appeared, there was almost no motion between first and last frame, and anatomy broke on limbs. Also worth knowing: in this mode the input image is a &lt;em&gt;reference&lt;/em&gt;, not a locked first frame — the model treats it as a hint and owes you nothing. Locking first and last frames is a different mode that lives in 14B models, which is not a conversation you have with 16 GB of RAM.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generation strangles your network
&lt;/h2&gt;

&lt;p&gt;A side finding that first looked like a bad ISP. Downloading weights crawled at 2.5 KB/s, while the same link gave 1.6–3.3 MB/s on an idle machine. The culprit was local: a generation was running, swap was full, and the network died along with everything else.&lt;/p&gt;

&lt;p&gt;Hence a rule that is not obvious until it bites: never run generation and downloads at the same time — chain them, weights first, compute second. And on downloading itself: &lt;code&gt;hf download&lt;/code&gt; creates a &lt;em&gt;new&lt;/em&gt; temp file on every broken attempt instead of resuming the old one, so on a flaky link it never finishes. After a series of attempts I had three stumps of one file: 0, 52 and 755 MB. What works is a resumable curl in a retry loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sL&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; - &lt;span class="nt"&gt;--speed-limit&lt;/span&gt; 51200 &lt;span class="nt"&gt;--speed-time&lt;/span&gt; 30 &lt;span class="nt"&gt;--max-time&lt;/span&gt; 600 &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$F&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$U&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;-C -&lt;/code&gt; resumes at the break, and &lt;code&gt;--speed-limit&lt;/code&gt; with &lt;code&gt;--speed-time&lt;/code&gt; kill a dead socket in 30 seconds instead of hanging in a half-hour timeout.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "free" actually costs
&lt;/h2&gt;

&lt;p&gt;No money, but it bills you in disk:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Family&lt;/th&gt;
&lt;th&gt;On disk&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Flux (schnell, encoders, VAE)&lt;/td&gt;
&lt;td&gt;9.6 GB&lt;/td&gt;
&lt;td&gt;works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flux Kontext (two quants)&lt;/td&gt;
&lt;td&gt;16 GB&lt;/td&gt;
&lt;td&gt;works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video model (weights, encoder, VAE)&lt;/td&gt;
&lt;td&gt;9.8 GB&lt;/td&gt;
&lt;td&gt;computes nothing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;35 GB&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nearly ten of those gigabytes are a paid lesson: weights for a model that will not render a single frame on this machine. Plus the time to download them, plus two hours of fan noise for one second of garbage. For reference, that second costs roughly $0.10 (480p) to $0.23 (720p) in the cloud at public Seedance 2.5 pricing — and comes back usable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I landed
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Local:&lt;/strong&gt; drafts and idea checks — composition, angle, "does the joke read". Ninety seconds a shot, no billing counter. The cover image of the original article was made exactly this way: local draft to approve the concept, paid generation for the final.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local:&lt;/strong&gt; editing an existing frame with Kontext. Two minutes for something a generator will not do at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local:&lt;/strong&gt; anything that should not leave the machine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Paid:&lt;/strong&gt; final quality. A 4-bit quant on 16 GB loses detail, and a cover is what people see before the text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Paid:&lt;/strong&gt; video, entirely. Not because local is more expensive, but because on this hardware local does not exist.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One direction I have not closed: MLX, Apple's tensor framework for unified memory. It runs outside ggml and its limits, the Flux port is mature, video ports exist but are raw. If I come back to local video on Apple Silicon, it will be through MLX, not through the path in this article.&lt;/p&gt;

&lt;p&gt;If there is one line to take away: local generation is not a free replacement for the cloud, it is a different tool with its own domain. It covers drafts, iteration and privacy very well, and final quality and video not at all. Both slogans — "why pay for anything" and "it doesn't work on a laptop" — get in the way equally.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/local-generation-mac/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>macos</category>
      <category>machinelearning</category>
      <category>performance</category>
    </item>
    <item>
      <title>Russia Just Started Selling Quantum Computers as Hardware, Not Cloud Time</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Wed, 16 Sep 2026 18:32:42 +0000</pubDate>
      <link>https://dev.to/klukyanov/russia-just-started-selling-quantum-computers-as-hardware-not-cloud-time-5bma</link>
      <guid>https://dev.to/klukyanov/russia-just-started-selling-quantum-computers-as-hardware-not-cloud-time-5bma</guid>
      <description>&lt;p&gt;Yesterday's headlines all said the same thing: "Russia starts selling quantum computers." Technically true — for the first time, a Russian quantum computing effort isn't offering "cloud access," it's issuing an invoice and shipping hardware. But behind that headline is a much narrower and much more interesting story than "Russia catches up to Google."&lt;/p&gt;

&lt;p&gt;What's actually being sold isn't a personal computer, or even a data-center server — it's a research installation the size of a walk-in freezer, priced like an apartment, built for "try real quantum hardware," not "solve a business problem faster than a classical machine." The details of what's in the box matter a lot more than the word "quantum" in the headline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually for sale
&lt;/h2&gt;

&lt;p&gt;The seller is Quantum Park — a joint center of Bauman Moscow State Technical University and VNIIA (a Rosatom research institute), led by the same team (Ilya Rodionov's group) that shipped Russia's first batch of photonic chips back in July. The product line is called SnowDrop and comes in four configurations of superconducting quantum coprocessors:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Qubits&lt;/th&gt;
&lt;th&gt;Topology&lt;/th&gt;
&lt;th&gt;Single-qubit fidelity&lt;/th&gt;
&lt;th&gt;Two-qubit fidelity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SnowDrop 4QTCS&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;T-shaped&lt;/td&gt;
&lt;td&gt;up to 99.96%&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SnowDrop 8Q2TCS&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;T-shaped&lt;/td&gt;
&lt;td&gt;99.927%&lt;/td&gt;
&lt;td&gt;99.21%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SnowDrop 8QGCL&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Grid&lt;/td&gt;
&lt;td&gt;99.9%+&lt;/td&gt;
&lt;td&gt;up to 99.55%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SnowDrop 10QGCL&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Grid&lt;/td&gt;
&lt;td&gt;99.9%+&lt;/td&gt;
&lt;td&gt;up to 99.55%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The platform is marketed as scalable to 48 qubits — that's a roadmap, not what ships today. Every model comes as a complete kit: the quantum chip itself, a cryogenic readout system, a Russian-made ultra-low-temperature cryostat called YARANGA with control signal switching, qubit control electronics, and a software stack for running algorithms. They're not selling a chip — they're selling a turnkey lab.&lt;/p&gt;

&lt;h2&gt;
  
  
  The price: 170 million rubles, and the questions that go unanswered
&lt;/h2&gt;

&lt;p&gt;Only one number is public: the entry-level platform with the 4-qubit processor starts at 170 million rubles (roughly $1.9M) — comparable, as several outlets pointed out, to a four-bedroom apartment in central Moscow. Pricing for the 8- and 10-qubit versions hasn't been disclosed. Given that more qubits mean tighter cooling and connectivity requirements, a higher price is a safe bet, but I won't put a number on it — there isn't one in any public source.&lt;/p&gt;

&lt;p&gt;This isn't really a complaint about secrecy: per-unit price lists are rare for this class of hardware globally. IBM and Google don't publish quantum system pricing either — they negotiate contracts individually. But it does mean there's currently no way to compare "cost per Russian qubit" to "cost per IBM qubit" in any meaningful way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fidelity: the numbers are real, the bar is different
&lt;/h2&gt;

&lt;p&gt;The most concrete thing in this release is the fidelity data, and it's genuinely good for a superconducting architecture: single-qubit gate errors under 0.2%, two-qubit gate errors under 1% on the best qubit pairs. For context, the widely cited threshold below which error correction can actually suppress accumulated noise rather than add to it sits around 99% two-qubit fidelity. So the claimed numbers are formally above that threshold — that's a real engineering result, not marketing rounding.&lt;/p&gt;

&lt;p&gt;The catch: "operation fidelity" and "useful quantum computation" are different metrics. SnowDrop tops out at 4-10 computational qubits, and any algorithm with real error correction needs orders of magnitude more physical qubits per logical qubit. So "fidelity on par with world-class systems" is an honest claim about individual operations — not a claim that this machine is solving problems classical computers can't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this sits globally
&lt;/h2&gt;

&lt;p&gt;Compared head-to-head on qubit count, it's not close: IBM already has chips with over a thousand physical qubits in its fleet, Google demonstrated the 105-qubit Willow chip below the error-correction threshold back in late 2024, and IonQ has dozens of algorithmic qubits on trapped-ion hardware with a completely different physics and different strengths. Formally, the Russian lineup is two orders of magnitude behind at launch.&lt;/p&gt;

&lt;p&gt;But the qubit-count race isn't the only metric, and it's often not the most honest one — D-Wave has thousands of qubits, but that's an annealing architecture for a narrow class of optimization problems, not universal gate-based computing like SnowDrop. Realistically, the Russian team isn't competing with 2026-vintage IBM and Google flagships — it's roughly where those companies were five to seven years ago: "the first reliably working, non-one-off superconducting gate-model computer in the country." For the domestic market, that's a real milestone. For the global race, it's an entry, not a claim to leadership.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this is actually for
&lt;/h2&gt;

&lt;p&gt;The buyer list isn't hidden: universities, research centers, and companies that need infrastructure to develop and test quantum algorithms — not consumer or even typical enterprise users. Before this announcement, the only way to access a Russian quantum coprocessor legally was through the Bauman Octillion cloud platform: since July 2025, it's run over 103,000 hybrid quantum-classical algorithms, with more than 9,000 cumulative hours of uptime. Selling hardware is the next step for anyone who's outgrown queue time on a shared cluster and needs their own installation for continuous experimentation.&lt;/p&gt;

&lt;p&gt;Bauman's rector, Mikhail Gordin, framed the project's logic this way: an engineering university's job isn't just to create technology, but to bring it to practical use. After further equipping Quantum Park, the team claims a production capacity of 10-20 multi-qubit installations per year — not mass manufacturing in the usual sense, but hand-built units for specific customers, which is normal for hardware at this level.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway: an infrastructure story, not a consumer one
&lt;/h2&gt;

&lt;p&gt;Strip away the "Russia vs. the world" headline effect, and what's left is a modest but genuinely interesting story: the country now has a reproducible product, tested across 103,000 runs, with an actual price tag on at least the base version — not a one-off prototype in a single lab. For a university or R&amp;amp;D team with 170+ million rubles and a real need to run quantum algorithms on dedicated hardware instead of queuing for cloud time, this is a reasonable purchase. For the market broadly, it's a signal that after years of cloud-only access, Russia's quantum program has moved into a phase where it's willing to sell a physical product, not just demo prototypes.&lt;/p&gt;

&lt;p&gt;I wouldn't read this as "a quantum laptop in five years," though. The path from 10 research qubits to something that actually beats a classical computer on a real business problem is measured in decades worldwide, and Russia isn't an exception here.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/russian-quantum-computer-snowdrop/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shorter weekly write-ups (in Russian) — on &lt;a href="https://t.me/+KSAy8mfFOsgwYWMy" rel="noopener noreferrer"&gt;Telegram&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>quantumcomputing</category>
      <category>hardware</category>
      <category>science</category>
      <category>russia</category>
    </item>
    <item>
      <title>iPhone Duo: What Apple's 'Boring' Keynote Actually Means for iOS Developers</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Mon, 14 Sep 2026 14:28:34 +0000</pubDate>
      <link>https://dev.to/klukyanov/iphone-duo-what-apples-boring-keynote-actually-means-for-ios-developers-20dp</link>
      <guid>https://dev.to/klukyanov/iphone-duo-what-apples-boring-keynote-actually-means-for-ios-developers-20dp</guid>
      <description>&lt;p&gt;Apple's September keynote felt like "nothing interesting": a slightly faster Pro, a slightly pricier Pro, new watches, new earbuds, and a foldable that looked almost exactly like the leaks. For buyers, that's a fair summary. For iOS developers, it's the opposite — the evening quietly delivered the biggest layout change since the iPhone X notch, and the clock on it runs out on October 23.&lt;/p&gt;

&lt;p&gt;The day before the event I published 13 predictions with probabilities. Here's how they did, what was actually shown, and — the part I really care about — what the iPhone Duo's inner display means for your code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scorecard: 11 out of 13
&lt;/h2&gt;

&lt;p&gt;Probabilities were set on the morning of September 9, before the stream, and haven't been touched since. Anything below 50% means I was betting &lt;em&gt;against&lt;/em&gt; the event.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Prediction&lt;/th&gt;
&lt;th&gt;My odds&lt;/th&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A foldable iPhone is shown&lt;/td&gt;
&lt;td&gt;95%&lt;/td&gt;
&lt;td&gt;Shown — iPhone Duo&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;It's called iPhone Ultra&lt;/td&gt;
&lt;td&gt;55%&lt;/td&gt;
&lt;td&gt;It's called iPhone Duo&lt;/td&gt;
&lt;td&gt;✗&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Foldable starts at $2,000&lt;/td&gt;
&lt;td&gt;85%&lt;/td&gt;
&lt;td&gt;$1,999 for 256 GB&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Touch ID instead of Face ID on the foldable&lt;/td&gt;
&lt;td&gt;75%&lt;/td&gt;
&lt;td&gt;Touch ID in the side button&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No telephoto on the foldable&lt;/td&gt;
&lt;td&gt;70%&lt;/td&gt;
&lt;td&gt;Main 48 MP + ultra wide; 2x is a crop of the main sensor&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No base iPhone 18&lt;/td&gt;
&lt;td&gt;90%&lt;/td&gt;
&lt;td&gt;Moved to spring 2027&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A20 Pro on 2 nm in all three iPhones&lt;/td&gt;
&lt;td&gt;90%&lt;/td&gt;
&lt;td&gt;All three&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;iPhone 18 Pro gets at least $100 pricier&lt;/td&gt;
&lt;td&gt;70%&lt;/td&gt;
&lt;td&gt;$1,199 / $1,299 — exactly +$100&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Variable aperture on the Pro main camera&lt;/td&gt;
&lt;td&gt;65%&lt;/td&gt;
&lt;td&gt;Yes, plus manual aperture and shutter&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ceramic Apple Watch Series 12&lt;/td&gt;
&lt;td&gt;70%&lt;/td&gt;
&lt;td&gt;Ceramic finish in the lineup&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AirPods 5 with a new H3 chip&lt;/td&gt;
&lt;td&gt;65%&lt;/td&gt;
&lt;td&gt;H3, $129 / $149 — ANC on both models&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Home hub is shown (bet against)&lt;/td&gt;
&lt;td&gt;35%&lt;/td&gt;
&lt;td&gt;Not shown&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Something not in any leak (bet against)&lt;/td&gt;
&lt;td&gt;30%&lt;/td&gt;
&lt;td&gt;Front camera under the foldable's inner display&lt;/td&gt;
&lt;td&gt;✗&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Brier score: &lt;strong&gt;0.118&lt;/strong&gt; (answering 50/50 on everything scores 0.25). I missed exactly where I warned I would: specs leak for months because thousands of factory workers know them, while the name is known by about twenty people who stay quiet. And I got cocky betting against a surprise.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was announced
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Device&lt;/th&gt;
&lt;th&gt;US price&lt;/th&gt;
&lt;th&gt;Pre-order&lt;/th&gt;
&lt;th&gt;Available&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;iPhone 18 Pro&lt;/td&gt;
&lt;td&gt;from $1,199&lt;/td&gt;
&lt;td&gt;Sep 12&lt;/td&gt;
&lt;td&gt;Sep 18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;iPhone 18 Pro Max&lt;/td&gt;
&lt;td&gt;from $1,299&lt;/td&gt;
&lt;td&gt;Sep 12&lt;/td&gt;
&lt;td&gt;Sep 18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;iPhone Duo&lt;/td&gt;
&lt;td&gt;from $1,999&lt;/td&gt;
&lt;td&gt;Oct 16&lt;/td&gt;
&lt;td&gt;Oct 23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apple Watch Series 12&lt;/td&gt;
&lt;td&gt;from $399&lt;/td&gt;
&lt;td&gt;now&lt;/td&gt;
&lt;td&gt;Sep 18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apple Watch Ultra 4&lt;/td&gt;
&lt;td&gt;from $799&lt;/td&gt;
&lt;td&gt;now&lt;/td&gt;
&lt;td&gt;Sep 18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AirPods 5&lt;/td&gt;
&lt;td&gt;$129 / $149&lt;/td&gt;
&lt;td&gt;now&lt;/td&gt;
&lt;td&gt;Sep 18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;iOS 27, macOS Golden Gate &amp;amp; co&lt;/td&gt;
&lt;td&gt;free&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Sep 14&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The good:&lt;/strong&gt; a foldable whose two displays share (almost) the same proportions, so content scales instead of reflowing; Split View on an iPhone for the first time; a real variable aperture; 45 hours of video on the Pro Max; ANC in the $129 AirPods; and Apple Reference Image, which signs sensor data so a photo can later be proven to be a photo — arguably the most underrated announcement of the night.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The not-so-good:&lt;/strong&gt; +$100 on the Pro and +$300 on the 1 TB tier, price bumps on older models; Siri AI ships in beta, English-only, with daily usage limits and "expanded access for a fee in the future" (and not at all on iPhone in the EU or mainland China); the iPhone 18 Pro display is unchanged apart from a narrower Dynamic Island; the Duo is eSIM-only and doesn't support Apple Pencil Pro.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not shown:&lt;/strong&gt; the base iPhone 18, the home hub, HomePod mini, Apple TV, any Mac — and, for developers, the iOS 27.1 SDK.&lt;/p&gt;

&lt;h2&gt;
  
  
  The new screen in numbers
&lt;/h2&gt;

&lt;p&gt;For 19 years the iPhone has been a tall, narrow strip: 1.5 on the original, 2.17 on today's Pro Max. Every mobile layout we write grew up inside that strip. The Duo's inner display breaks it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;iPhone 18 Pro Max&lt;/th&gt;
&lt;th&gt;Duo, outer&lt;/th&gt;
&lt;th&gt;Duo, inner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Diagonal&lt;/td&gt;
&lt;td&gt;6.9″&lt;/td&gt;
&lt;td&gt;5.4″&lt;/td&gt;
&lt;td&gt;7.6″&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pixels&lt;/td&gt;
&lt;td&gt;1320 × 2868&lt;/td&gt;
&lt;td&gt;1398 × 2034&lt;/td&gt;
&lt;td&gt;2670 × 1878&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Points (w × h)&lt;/td&gt;
&lt;td&gt;440 × 956&lt;/td&gt;
&lt;td&gt;466 × 678&lt;/td&gt;
&lt;td&gt;890 × 626&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Density&lt;/td&gt;
&lt;td&gt;460 ppi&lt;/td&gt;
&lt;td&gt;460 ppi&lt;/td&gt;
&lt;td&gt;430 ppi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aspect ratio&lt;/td&gt;
&lt;td&gt;2.17&lt;/td&gt;
&lt;td&gt;1.45&lt;/td&gt;
&lt;td&gt;1.42&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Size classes (portrait)&lt;/td&gt;
&lt;td&gt;compact × regular&lt;/td&gt;
&lt;td&gt;compact × regular&lt;/td&gt;
&lt;td&gt;regular × regular&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Supported orientations&lt;/td&gt;
&lt;td&gt;honored&lt;/td&gt;
&lt;td&gt;honored&lt;/td&gt;
&lt;td&gt;ignored&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three takeaways:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;1.42 is roughly √2 — the A4 paper ratio.&lt;/strong&gt; Fold such a sheet in half and each half keeps the same shape. In Split View the inner display splits into two panes of roughly 445 × 626 points, each almost the shape of the outer display. Two apps side by side each get an "outer screen". That's hardware designed with software in mind.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The open Duo is wider than tall.&lt;/strong&gt; It's not a phone turned sideways — it's the new default posture. The Dynamic Island sits vertically on the side, so safe-area insets are asymmetric, and with the 27.1 SDK vertical bars move to the left edge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;16:9 video loses.&lt;/strong&gt; At 890 points wide a video is 890 × 501, leaving 125 points of black — a fifth of the screen.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  How the Duo shows your app if you do nothing
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Built without the iOS 27 SDK:&lt;/strong&gt; a familiar phone-sized window in the middle of the inner display, surrounded by black. A postcard in a giant frame.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Built with the iOS 27 SDK:&lt;/strong&gt; wider, but part of the screen is still empty.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Built with the iOS 27.1 SDK:&lt;/strong&gt; edge to edge, toolbars and tab bars move to the left. This is the native look — and the SDK for it isn't even in beta yet.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What to grep for this week
&lt;/h2&gt;

&lt;p&gt;The good news: almost all of this is stuff Apple has asked us to do for years. The bad news: iPhones were always narrow, and plenty of codebases quietly relied on it. The Duo is still an iPhone, but inside it's regular × regular, like an iPad.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;In your code&lt;/th&gt;
&lt;th&gt;Why it breaks on Duo&lt;/th&gt;
&lt;th&gt;Replace with&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;UIScreen.main.bounds&lt;/code&gt;, hardcoded widths&lt;/td&gt;
&lt;td&gt;Two differently shaped screens, the app moves between them live&lt;/td&gt;
&lt;td&gt;Your view's geometry (&lt;code&gt;onGeometryChange&lt;/code&gt;), screen via &lt;code&gt;UIWindowScene&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interface orientation checks&lt;/td&gt;
&lt;td&gt;The inner display ignores supported orientations&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;horizontalSizeClass&lt;/code&gt; / &lt;code&gt;verticalSizeClass&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;userInterfaceIdiom == .phone&lt;/code&gt; as "narrow screen"&lt;/td&gt;
&lt;td&gt;Duo is an iPhone with regular × regular and Split View&lt;/td&gt;
&lt;td&gt;Layout by size class, not by device type&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Symmetric safe-area insets&lt;/td&gt;
&lt;td&gt;The Dynamic Island is on the side&lt;/td&gt;
&lt;td&gt;Handle each inset and margin independently&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;UIRequiresFullScreen&lt;/code&gt; in Info.plist&lt;/td&gt;
&lt;td&gt;No longer prevents resizing from iOS 27&lt;/td&gt;
&lt;td&gt;Support live resizing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A hardcoded "Sign in with Face ID" label&lt;/td&gt;
&lt;td&gt;Duo has Touch ID in the side button&lt;/td&gt;
&lt;td&gt;Branch on &lt;code&gt;LAContext.biometryType&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One window per app&lt;/td&gt;
&lt;td&gt;Split View opens multiple scenes, new windows only on the inner display&lt;/td&gt;
&lt;td&gt;Support multiple scenes and handle refused requests&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Biometrics is the easy one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;import&lt;/span&gt; &lt;span class="kt"&gt;LocalAuthentication&lt;/span&gt;

&lt;span class="kd"&gt;func&lt;/span&gt; &lt;span class="nf"&gt;biometryLabel&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="kt"&gt;String&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="nv"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;LAContext&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;canEvaluatePolicy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deviceOwnerAuthentication&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;switch&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;biometryType&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;faceID&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="s"&gt;"Sign in with Face ID"&lt;/span&gt;
    &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;touchID&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Sign in with Touch ID"&lt;/span&gt;
    &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nv"&gt;opticID&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"Sign in with Optic ID"&lt;/span&gt;
    &lt;span class="k"&gt;default&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;       &lt;span class="s"&gt;"Sign in with passcode"&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Layout isn't rocket science either if you're on SwiftUI — &lt;code&gt;NavigationSplitView&lt;/code&gt; collapses to one column in compact width and expands on the inner display:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight swift"&gt;&lt;code&gt;&lt;span class="kd"&gt;struct&lt;/span&gt; &lt;span class="kt"&gt;RootView&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;View&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;@Environment&lt;/span&gt;&lt;span class="p"&gt;(\&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;horizontalSizeClass&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;hSize&lt;/span&gt;

    &lt;span class="k"&gt;var&lt;/span&gt; &lt;span class="nv"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kd"&gt;some&lt;/span&gt; &lt;span class="kt"&gt;View&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kt"&gt;NavigationSplitView&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;SidebarList&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="nv"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="kt"&gt;DetailView&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
                &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;padding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hSize&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;regular&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Apple's "Preparing your app for iPhone Duo" material also names foldable-specific APIs: &lt;code&gt;onHingeChange&lt;/code&gt; / &lt;code&gt;UIHingeInteraction&lt;/code&gt; for hinge states and angle, &lt;code&gt;ArrangementView&lt;/code&gt; / &lt;code&gt;UIArrangementViewController&lt;/code&gt; for two-pane layouts around the fold, &lt;code&gt;reservedRegion&lt;/code&gt; on &lt;code&gt;GeometryProxy&lt;/code&gt; / &lt;code&gt;UIView&lt;/code&gt; to keep controls off the fold, &lt;code&gt;CameraCaptureAccessory&lt;/code&gt; for a preview or teleprompter on the outer display, and &lt;code&gt;AVCaptureDeviceDirectionCoordinator&lt;/code&gt; — because camera position is no longer camera direction.&lt;/p&gt;

&lt;p&gt;Two easy-to-miss extras: StandBy on the Duo runs on either display even when not charging, so apps without widgets or Live Activities simply don't exist there. And health apps should expect a lot more data — Series 12 and Ultra 4 sample heart rate every five seconds all day.&lt;/p&gt;

&lt;h3&gt;
  
  
  The next six weeks
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sep 14&lt;/strong&gt; — iOS 27 ships. Its SDK (already in RC) decides whether your app resizes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sep 16–17&lt;/strong&gt; — Apple group labs on iPhone Duo; Q&amp;amp;As on Sep 23.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Xcode 27.1 beta&lt;/strong&gt; — iOS 27.1 SDK and a Duo simulator you can open, close, rotate and partially fold.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Oct 23&lt;/strong&gt; — iPhone Duo ships on iOS 27.1, and people who paid $1,999 start opening your app.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;If you own an iPhone 16 Pro or 17 Pro, you could have skipped this keynote. If you write iOS apps, this is the busiest autumn since 2017, when the iPhone X taught everyone the words "safe area". This time the lesson is that an iPhone can be wider than it is tall, that size classes on a phone aren't a formality, and that one window per app is no longer a given. The keynote was boring precisely because everything important in it was for Xcode, not for the store window.&lt;/p&gt;

&lt;p&gt;How many &lt;code&gt;UIScreen.main&lt;/code&gt; references are in your project?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/apple-event-september-2026-itogi/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ios</category>
      <category>swiftui</category>
      <category>apple</category>
      <category>mobile</category>
    </item>
    <item>
      <title>The Keynote Without Tim Cook: My Bets on What Apple Announces Today</title>
      <dc:creator>Kirill Lukyanov</dc:creator>
      <pubDate>Wed, 09 Sep 2026 10:37:49 +0000</pubDate>
      <link>https://dev.to/klukyanov/the-keynote-without-tim-cook-my-bets-on-what-apple-announces-today-50jm</link>
      <guid>https://dev.to/klukyanov/the-keynote-without-tim-cook-my-bets-on-what-apple-announces-today-50jm</guid>
      <description>&lt;p&gt;Today at 8pm Moscow time, Apple is holding its first fall keynote in fifteen years without Tim Cook on stage — and, if the leaks hold up, its first-ever foldable iPhone reveal. I'm writing this before the event starts. Not another leak roundup — there will be a hundred of those today — but my own predictions, with probabilities attached, so I can come back after the dust settles and check where I was wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  A new person on stage — and it matters
&lt;/h2&gt;

&lt;p&gt;On September 1, 2026, Tim Cook moved into the role of Executive Chairman, and John Ternus became Apple's CEO. Cook spent fifteen years in the role and leaves with the company valued around $4.6 trillion. Today's keynote is Ternus's first major public appearance in the new job.&lt;/p&gt;

&lt;p&gt;His background matters here. Ternus has spent twenty-five years at Apple, trained as a mechanical engineer, and most recently ran hardware engineering — Apple Watch, AirPods, Vision Pro, and the Mac's move to Apple silicon all passed through his hands. He's not a finance guy or a marketer; his whole career has been about how the physical object works.&lt;/p&gt;

&lt;p&gt;And here's the coincidence that's hard to ignore: his first keynote is the one where Apple, for the first time in its history, shows a foldable phone. A hinge, a 4.5mm folded body, a panel rated for 200,000 folds — that's exactly the kind of engineering problem Ternus grew up solving. If Apple timed the leadership handoff to land on the most "hardware" launch of the decade, the timing landed well.&lt;/p&gt;

&lt;p&gt;The practical takeaway: expect hardware today, not services. I wouldn't bet on a big subscription or ad-platform announcement closing out the evening.&lt;/p&gt;

&lt;h2&gt;
  
  
  My bets: thirteen predictions with probabilities
&lt;/h2&gt;

&lt;p&gt;This is the part the whole piece was written for. The probabilities are mine — not a betting line, not analyst consensus — based on how many leaks agree, how many independent supply-chain sources confirm a given spec, and Apple's long habit of surprising everyone specifically on names and prices.&lt;/p&gt;

&lt;p&gt;Read it like this: 90%+ is essentially locked in — only a real surprise changes it. 60-80% is likely, but I wouldn't be shocked by the opposite. Below 50% and I'm leaning against it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Prediction&lt;/th&gt;
&lt;th&gt;Probability&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Foldable iPhone is shown today&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;95%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Confirmed across independent supply chains, and the "Surprise and shine" slogan is clearly about this. You don't build hype this size for a camera bump.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;It's called iPhone Ultra&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;55%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The name that shows up most often in leaks, but naming is the one thing Apple keeps secret until the end. "Fold" was on the list too.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Foldable starts at $2000+&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;85%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Two OLED panels, a titanium hinge, a new frame. Leaks consistently land on $2000 for the base and $2500+ for higher storage.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Touch ID instead of Face ID on the foldable&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;75%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4.5mm folded simply doesn't leave room for a TrueDepth module. A side button is the only way out.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No telephoto on the foldable&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;70%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Thickness beats periscope optics. Leaks consistently show two cameras: wide and ultra-wide.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No standard iPhone 18 today&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;90%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Every major source confirms the affordable line is pushed to spring 2027. Apple is splitting the lineup across two seasons.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A20 Pro on a 2nm process across all three iPhones&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;90%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The 2nm jump is the year's biggest hardware story — no reason to split it by model.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;iPhone 18 Pro prices go up at least $100&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;70%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rumored at $1199/$1299. The new process is expensive, but Apple has held Pro pricing before, so a freeze isn't off the table.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Variable aperture on the Pro's main camera&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;65%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Long-rumored and fits the "pro" narrative, but a moving mechanical part in a thin body is easy to slip a year.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Apple Watch Series 12 gets a ceramic case option&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;70%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;It's been in the lineup before and shows up in leaks, but Apple has delayed such variants to spring before.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AirPods 5 with a new H3 chip&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;65%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Follows the usual refresh cycle, and two versions (with/without ANC) is plausible. Easy to get overshadowed by the foldable though.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A smart home hub with a 7-inch display&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;35%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Delayed twice already, tied to a new Siri. I'm betting against a proper reveal sharing the stage with the foldable.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Something not in any leak at all&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;30%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Recent years' leaks cover almost everything. But a new CEO's first keynote is exactly the moment to keep one card up your sleeve.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If I had to compress it into one line: I'm nearly certain about the hardware and nearly unsure about names and prices. Specs leak months in advance because thousands of people on factory floors know them. Names and prices are known by maybe twenty people, and they don't talk.&lt;/p&gt;

&lt;h2&gt;
  
  
  The foldable iPhone: the night's biggest bet
&lt;/h2&gt;

&lt;p&gt;For eight years the foldable phone market existed without Apple. Samsung shipped seven generations, Chinese makers got the crease nearly invisible, and Apple stayed silent. Today, apparently, the silence ends.&lt;/p&gt;

&lt;p&gt;What the leaks say: a book-style fold like the Galaxy Z Fold, a ~5.5-inch outer display and ~7.8-inch inner one. 4.5mm unfolded, about 9mm folded. A 5400-5800 mAh battery, A20 Pro chip, 12GB of RAM. Two rear cameras, no telephoto. Two colors: silver-white and deep indigo. MagSafe included. Starting around $2000.&lt;/p&gt;

&lt;p&gt;What interests me most isn't the thinness record — it's the &lt;strong&gt;Face ID compromise&lt;/strong&gt;. If Apple really is going back to a fingerprint sensor in a button, this is the first flagship in eight years where biometrics take a step backward. A company that built half its interface — unlocking, Apple Pay, password autofill — around Face ID is accepting a different method because physics leaves no other option. For a company that usually delays a product rather than back off its own standard, that's a real concession.&lt;/p&gt;

&lt;p&gt;And that's the real question behind the price. Selling a $2000 device with no telephoto and no Face ID is harder than it sounds. The pitch has to be the screen itself — email, a spreadsheet, and a messaging app side by side, real multitasking, unfolding in one motion. If Apple spends today's demo on real use cases instead of millimeters of thickness, the product could take off. If the whole story is a thinness record and hinge materials, it's a very expensive toy for enthusiasts.&lt;/p&gt;

&lt;h2&gt;
  
  
  iPhone 18 Pro: an evolution you'll pay more for
&lt;/h2&gt;

&lt;p&gt;Next to the foldable, the regular Pro models look boring — and that seems deliberate. Design barely changes from the 17 Pro. What changes is what's inside: the A20 Pro on a 2nm process (roughly 18% faster, 30% more efficient), a variable aperture main camera with a Samsung three-layer stacked sensor, LTPO+ display, a bigger battery in the Pro Max (~5567 mAh), new colors (dark cherry, light blue, silver — no black, if the leaks hold), and a $100+ price bump landing around $1199/$1299.&lt;/p&gt;

&lt;p&gt;The variable aperture is the most interesting bit. Every computational trick of recent years — portrait mode, night mode, background blur — has been software compensating for optics too small to fit in a phone. A mechanically variable aperture is a step in the other direction: not estimating a shot, but actually capturing it. For anyone who shoots seriously on a phone, that's a bigger deal than another megapixel bump.&lt;/p&gt;

&lt;h2&gt;
  
  
  What won't be shown today — and why that matters more
&lt;/h2&gt;

&lt;p&gt;The most underrated story of this fall isn't what gets announced, it's what doesn't. There's likely no standard iPhone 18 today — the affordable models are moving to spring 2027.&lt;/p&gt;

&lt;p&gt;That breaks a decade-plus ritual: September used to mean the entire lineup, base to Pro Max, all at once. Now the year splits in two. Expensive models in fall, mass-market ones in spring.&lt;/p&gt;

&lt;p&gt;The logic makes sense. A fall keynote stops being six devices competing for attention — the expensive models get the whole stage and head into holiday sales without internal competition. The mass-market iPhone gets a quiet spring window all to itself with no news cycle competing for headlines.&lt;/p&gt;

&lt;p&gt;For buyers, the downside is real: if you just want a good new iPhone without the "Pro" label, tonight doesn't concern you — wait until spring. And the price gap widens: the cheapest new iPhone now arrives six months after the most expensive one, and September's lineup starts at $1200-plus.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watches, earbuds, and the home
&lt;/h2&gt;

&lt;p&gt;Everything else tonight is background noise, and that's telling too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Apple Watch Series 12&lt;/strong&gt; — no design change, new processor, a returning ceramic case option, and reportedly continuous heart-rate monitoring throughout the day. That last part is the genuinely interesting one: continuous measurement instead of spot checks is a meaningfully different quality of health data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Apple Watch Ultra 4&lt;/strong&gt; — same processor, incremental improvements. An upgrade for first-gen Ultra owners.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AirPods 5&lt;/strong&gt; — H3 chip, better sound, lower latency, in ANC and non-ANC versions. Splitting into two versions is a way to occupy a lower price tier without touching the Pro line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A smart home hub&lt;/strong&gt; with a 7-inch square display, camera, and video calling is the shakiest bet of the night. It's tied to a new Siri, and Siri has been rocky for Apple the last couple of years. I'm betting it doesn't get a proper showing today.&lt;/p&gt;

&lt;p&gt;Separately: iOS 27, iPadOS 27, macOS Golden Gate, watchOS 27, tvOS 27, and visionOS 27 are expected to release September 14. Preorders likely shift from the usual Friday to Saturday, September 12 — September 11 isn't a day Apple opens sales on. Retail availability is expected September 18.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means if you write iOS code
&lt;/h2&gt;

&lt;p&gt;If you build for iOS, tonight leaves you work for the fall. A foldable screen isn't "just another size" — it's several new states your app has to handle honestly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two genuinely different screens in one device.&lt;/strong&gt; 5.5 inches outside and 7.8 inside are different size classes entirely — compact phone outside, near-tablet inside. Anything laid out for a single fixed-width column will look stretched on the inner display. Adaptive size classes — the thing many teams have skipped for years because "iPhone is always compact" — suddenly matter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Live transitions between states.&lt;/strong&gt; Users will unfold the device mid-interaction. That's a runtime geometry change with state preservation: unfinished text input, scroll position, an open navigation screen. This is where anything that stores state in the view hierarchy instead of the model breaks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Biometrics.&lt;/strong&gt; If Touch ID really is back, any code or copy that hardcodes Face ID assumptions will show users something untrue. Check the biometry type on-device — never assume it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;App Store screenshots.&lt;/strong&gt; A new form factor means a new set of storefront assets. Anyone shipping an app will need another screenshot pass by spring.&lt;/p&gt;

&lt;p&gt;The good news: almost all of this is fixed by doing what was already considered best practice — not nailing your layout to fixed sizes, and keeping state separate from the view. The bad news: far fewer apps do this cleanly than people like to think.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd bet on myself
&lt;/h2&gt;

&lt;p&gt;If I had exactly one bet for tonight, I wouldn't put it on the foldable — that's nearly a lock already. I'd bet that &lt;strong&gt;the real drama turns out to be pricing, not hardware&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Almost everything about tonight's devices is already known technically. What's not known is how much Apple will actually charge for the first foldable, and whether it will simultaneously raise Pro prices. $2000 for the foldable plus $100 on top of Pro would make this the most expensive September in the company's history — and it'll be announced by someone who's run Apple for nine days.&lt;/p&gt;

&lt;p&gt;That's what to watch for at 8pm: not the millimeters of the hinge, but the pricing slide, and how confidently it gets delivered.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;All figures above are rumors and leaks as of the morning of September 9, 2026, and the probabilities are my own personal estimate, not analyst consensus. I'll come back to this piece after the keynote and count how many bets landed.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://klukyanov.ru/notes/apple-event-september-2026/" rel="noopener noreferrer"&gt;klukyanov.ru&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>apple</category>
      <category>ios</category>
      <category>tech</category>
      <category>swift</category>
    </item>
  </channel>
</rss>
