<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Super Lewis</title>
    <description>The latest articles on DEV Community by Super Lewis (@super_lewis).</description>
    <link>https://dev.to/super_lewis</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4064426%2F6d43e7fa-ca92-4bee-82ad-7abc88cee021.jpg</url>
      <title>DEV Community: Super Lewis</title>
      <link>https://dev.to/super_lewis</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/super_lewis"/>
    <language>en</language>
    <item>
      <title>Nano Banana 2.1 vs Nano Banana Pro: Same Prompts, Real Latency Data, and When Pro Is Worth It</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Thu, 08 Oct 2026 02:59:03 +0000</pubDate>
      <link>https://dev.to/super_lewis/nano-banana-21-vs-nano-banana-pro-same-prompts-real-latency-data-and-when-pro-is-worth-it-je0</link>
      <guid>https://dev.to/super_lewis/nano-banana-21-vs-nano-banana-pro-same-prompts-real-latency-data-and-when-pro-is-worth-it-je0</guid>
      <description>&lt;p&gt;&lt;strong&gt;Short answer:&lt;/strong&gt; for most product images, banners and text-heavy graphics, Nano Banana 2.1 does the job at roughly half to a third of Nano Banana Pro's price and about twice the speed. Pick Pro when the scene needs more world knowledge, richer detail or more character references (5 versus 4).&lt;/p&gt;

&lt;p&gt;Nano Banana 2.1 is Google's Flash-tier image generation and editing model (Gemini API id &lt;code&gt;gemini-nano-banana-2.1&lt;/code&gt;, released on 6 October 2026, built on Gemini 3.6 Flash). Nano Banana Pro is the Pro tier (&lt;code&gt;gemini-3-pro-image&lt;/code&gt;). Google's own docs call 2.1 "the more efficient counterpart to Gemini 3 Pro Image".&lt;/p&gt;

&lt;p&gt;Disclosure up front: the production numbers below come from apimodels.app, the API platform I work on, and the test images were generated through it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I compared them
&lt;/h2&gt;

&lt;p&gt;I fixed the criteria before looking at any output:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Price per image&lt;/strong&gt; at 1K, 2K and 4K, from Google's pricing page (updated 7 October 2026, standard tier).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Documented capabilities&lt;/strong&gt;: reference-image limits, thinking controls, grounding, aspect ratios.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same-prompt output&lt;/strong&gt; at 2K, 16:9: one text-heavy infographic, one edit from a reference image.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-world latency and success rate&lt;/strong&gt; from one week of production traffic (1–8 October 2026).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  At a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Nano Banana 2.1&lt;/th&gt;
&lt;th&gt;Nano Banana Pro&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini API id&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemini-nano-banana-2.1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemini-3-pro-image&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google price, 1K / 2K / 4K&lt;/td&gt;
&lt;td&gt;$0.0336 / $0.0504 / $0.113&lt;/td&gt;
&lt;td&gt;$0.134 / $0.134 / $0.24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Object references&lt;/td&gt;
&lt;td&gt;up to 10&lt;/td&gt;
&lt;td&gt;up to 6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Character references&lt;/td&gt;
&lt;td&gt;up to 4&lt;/td&gt;
&lt;td&gt;up to 5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Style references&lt;/td&gt;
&lt;td&gt;up to 3&lt;/td&gt;
&lt;td&gt;not documented&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking&lt;/td&gt;
&lt;td&gt;minimal / medium / high (default medium)&lt;/td&gt;
&lt;td&gt;always on&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Search grounding&lt;/td&gt;
&lt;td&gt;Web and Image Search&lt;/td&gt;
&lt;td&gt;Web Search&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extra wide ratios&lt;/td&gt;
&lt;td&gt;1:4, 4:1, 1:8, 8:1&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Our median time, 2K&lt;/td&gt;
&lt;td&gt;about 25 s&lt;/td&gt;
&lt;td&gt;about 40 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;volume, banners, infographics, many object refs&lt;/td&gt;
&lt;td&gt;complex scenes, brand-exact work, more character refs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 2K, Pro costs 2.7 times as much as 2.1 on Google's list; at 1K the gap is 4 times.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 1: a text-heavy infographic
&lt;/h2&gt;

&lt;p&gt;Prompt (same for both, 2K, 16:9):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A clean flat-design infographic poster, 16:9, titled "How an image API request works". Four numbered steps left to right, each with an icon and a short label: "1. Send prompt", "2. Queue task", "3. Render image", "4. Download result". A small footer line reads "Average time: 25 seconds". White background, navy and coral accents, crisp legible typography.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa033ckflyftw3f9fs2hy.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa033ckflyftw3f9fs2hy.jpg" alt="Side by side: Nano Banana 2.1 (25 s) on the left and Nano Banana Pro (57 s) on the right, both rendering the same four-step infographic with every label spelled correctly. AI-generated images." width="800" height="249"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Both models spelled every string correctly: the title, the four labels and the footer. Pro drew richer icons inside numbered circles; 2.1 went flatter and turned the footer into a navy bar. 2.1 finished in 25 seconds, Pro in 57. For text-heavy graphics at this size, I could not tell you which one is "better" without a style preference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 2: editing with a reference image
&lt;/h2&gt;

&lt;p&gt;The reference was a 1280×720 title card from our own site (that is why the brand name appears on the screen). Prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Use the reference image as the screen content. Show it playing on a large modern TV mounted on a living-room wall at night, warm lamp light, a plant on the left, a soft reflection on the floor. Keep the on-screen poster recognisable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffy97kjbgwe6kcrecw8fs.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffy97kjbgwe6kcrecw8fs.jpg" alt="Side by side: Nano Banana 2.1 (28 s) shows a wall-mounted TV with the reflection on a foreground table; Nano Banana Pro (49 s) puts the TV on a cabinet with a purple reflection on the floor. AI-generated images." width="800" height="249"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Scored against the prompt:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Wall-mounted TV&lt;/strong&gt;: 2.1 yes; Pro placed it on a cabinet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plant on the left, warm lamp&lt;/strong&gt;: both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reflection on the floor&lt;/strong&gt;: Pro yes; 2.1 put it on a foreground table (with the letters correctly mirrored).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On-screen text kept exact&lt;/strong&gt;: both.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each model missed one instruction. 2.1 took 28 seconds, Pro 49.&lt;/p&gt;

&lt;h2&gt;
  
  
  A week of production timings
&lt;/h2&gt;

&lt;p&gt;Two prompts prove nothing about reliability, so here is what real traffic looked like on apimodels.app from 1 to 8 October 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model and job&lt;/th&gt;
&lt;th&gt;Requests&lt;/th&gt;
&lt;th&gt;Median&lt;/th&gt;
&lt;th&gt;90th percentile&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pro, 4K with references&lt;/td&gt;
&lt;td&gt;617&lt;/td&gt;
&lt;td&gt;61 s&lt;/td&gt;
&lt;td&gt;252 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro, 2K with references&lt;/td&gt;
&lt;td&gt;203&lt;/td&gt;
&lt;td&gt;40 s&lt;/td&gt;
&lt;td&gt;76 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro, 2K text only&lt;/td&gt;
&lt;td&gt;83&lt;/td&gt;
&lt;td&gt;38 s&lt;/td&gt;
&lt;td&gt;143 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2.1, 4K with references&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;45 s&lt;/td&gt;
&lt;td&gt;50 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2.1, 2K (all)&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;about 25 s&lt;/td&gt;
&lt;td&gt;35 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2.1, 1K text only&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;13 s&lt;/td&gt;
&lt;td&gt;15 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Success rates over the same week: Pro 98.9% (971 requests; all 11 failures were content-moderation refusals), 2.1 97.6% (41 requests since it went live on 7 October; one upstream failure). The 2.1 sample is small, so treat its percentiles as a first look. The striking number is Pro's 4K tail: one request in ten took over four minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;Node 18+ (uses the built-in &lt;code&gt;fetch&lt;/code&gt;). The same script runs either model; swap the id.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// compare.mjs — APIMODELS_API_KEY in your environment&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;BASE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://api.apimodels.app/v1&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;H&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;Authorization&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;APIMODELS_API_KEY&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;A clean flat-design infographic poster, 16:9, titled "How an image API request works" ...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;t0&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;created&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;BASE&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/images/generations`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;H&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;aspect_ratio&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;16:9&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;resolution&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2K&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;then&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;created&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;taskId&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(;;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;4000&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;BASE&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/images/generations?task_id=&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;H&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;then&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pending&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;processing&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;seconds&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;t0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="na"&gt;urls&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;resultUrls&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;nano-banana-2-1&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gemini-3-pro-image&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)]))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To edit instead of generating, add &lt;code&gt;image_urls: ['https://…/your-reference.jpg']&lt;/code&gt; to the body (2.1 accepts up to 10 references here).&lt;/p&gt;

&lt;h2&gt;
  
  
  Which one should you pick
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;High-volume banners, thumbnails, product shots&lt;/strong&gt; → 2.1. About half the time per image and a third to a half of the price.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infographics and labelled diagrams at 2K or 4K&lt;/strong&gt; → 2.1 first. It matched Pro on spelling in my test, and Google's model card reports infographic factuality of 0.521 for 2.1 versus 0.179 for Nano Banana 2.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Very wide banners (4:1, 8:1)&lt;/strong&gt; → 2.1. Pro does not offer those ratios.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scenes that lean on world knowledge, exact brand assets or five recurring characters&lt;/strong&gt; → Pro. That is where Google positions it, and its extra detail showed in the icons.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Small text at 1K or long paragraphs&lt;/strong&gt; → neither. Google lists blurry small text at 1K as a known 2.1 limitation; render text at 2K or add it in post.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where the numbers came from, and when not to use apimodels.app
&lt;/h2&gt;

&lt;p&gt;apimodels.app is a multi-model API gateway: one API key and OpenAI- and Anthropic-compatible endpoints for about 150 image, video, audio and language models. Over the week above, Pro and 2.1 ran there at 98.9% and 97.6% success, failed calls are not charged, and the per-image prices are $0.024 / $0.04 / $0.064 for 2.1 and $0.08 / $0.08 / $0.13 for Pro, below Google's standard list.&lt;/p&gt;

&lt;p&gt;Do not route through it if you need 2.1's thinking-level control, Google Search grounding or all 14 reference images: our endpoint does not expose those yet, and Google's API does. The same applies if you only use Gemini and can wait for Google's batch tier, which is half the standard price. Model pages: &lt;a href="https://apimodels.app/models/nano-banana-2-1" rel="noopener noreferrer"&gt;Nano Banana 2.1&lt;/a&gt; and &lt;a href="https://apimodels.app/models/gemini-3-pro-image" rel="noopener noreferrer"&gt;Nano Banana Pro&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits of this comparison
&lt;/h2&gt;

&lt;p&gt;Two prompts, one run each, at 2K only. Image models vary run to run, so a single miss (Pro's cabinet, 2.1's table reflection) is an anecdote, not a rate. The latency data is real traffic but skewed toward 4K jobs with references for Pro.&lt;/p&gt;

&lt;p&gt;Which job would you trust to 2.1 and which would you still send to Pro? I'm curious where others draw the line.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This post was written with AI assistance from the test data above and reviewed before publishing. The comparison images are AI-generated by the two models being compared.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>gemini</category>
      <category>google</category>
    </item>
    <item>
      <title>How to Render a Motion Graphics Video with Claude Opus 5.5: Prompt, Frame Renderer, ffmpeg and Real Cost</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Wed, 07 Oct 2026 05:39:15 +0000</pubDate>
      <link>https://dev.to/super_lewis/how-to-render-a-motion-graphics-video-with-claude-opus-55-prompt-frame-renderer-ffmpeg-and-real-3e1g</link>
      <guid>https://dev.to/super_lewis/how-to-render-a-motion-graphics-video-with-claude-opus-55-prompt-frame-renderer-ffmpeg-and-real-3e1g</guid>
      <description>&lt;p&gt;Claude Opus 5.5 cannot output a single pixel, yet people keep posting motion-graphics videos "made by Opus 5.5". The trick is that the model writes a program that draws every frame, and you render that program into an MP4. This post is the full, reproducible pipeline we used for a 12-second 1280x720 promo: the exact prompt, the API call, a 25-line frame renderer, the ffmpeg command, and what it actually cost.&lt;/p&gt;

&lt;p&gt;Disclosure up front: we ran this through apimodels.app, the API gateway I work on. apimodels.app is a multi-model API gateway: one API key and an OpenAI-compatible endpoint for about 150 image, video, audio and language models, including &lt;code&gt;claude-opus-5-5&lt;/code&gt;. Every step below works the same against any OpenAI-compatible endpoint that serves Opus 5.5; only the base URL changes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4cjxido1vr72urymsu4i.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4cjxido1vr72urymsu4i.jpg" alt="First frame of the 12-second promo that Claude Opus 5.5 wrote as an HTML canvas program" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The idea in one sentence
&lt;/h2&gt;

&lt;p&gt;Ask Opus 5.5 for &lt;strong&gt;an HTML page with a deterministic &lt;code&gt;renderFrame(t)&lt;/code&gt; function&lt;/strong&gt;, not for "a video", then let a headless browser call that function once per frame.&lt;/p&gt;

&lt;p&gt;Opus 5.5 outputs text. A page that animates itself with &lt;code&gt;requestAnimationFrame&lt;/code&gt; looks fine in a browser, but you cannot export it frame-accurately, because what you see depends on timing. A page whose every frame is a pure function of &lt;code&gt;t&lt;/code&gt; can be rendered frame 0, frame 1 … frame 359 in any order, and each call returns the same picture. That is exactly what a video encoder needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you need
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;An API key for an endpoint that serves &lt;code&gt;claude-opus-5-5&lt;/code&gt; (OpenAI chat-completions format).&lt;/li&gt;
&lt;li&gt;Node.js 20+ and &lt;code&gt;puppeteer-core&lt;/code&gt; (&lt;code&gt;npm i puppeteer-core&lt;/code&gt; works; we use the Chrome already installed on the machine).&lt;/li&gt;
&lt;li&gt;ffmpeg 6+.&lt;/li&gt;
&lt;li&gt;About seven minutes of patience for the model call. More on that below.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 1: write the prompt as a frame contract
&lt;/h2&gt;

&lt;p&gt;The rules in the first half of this prompt are what make the output renderable. The story comes second, the style last. This is the prompt we sent, unedited:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are writing a program that renders a video. Output one complete, self-contained HTML file and nothing else.

Rules for the program (these make it renderable frame by frame):
1. A single &amp;lt;canvas&amp;gt; of exactly 1280x720, no CSS scaling, black page background.
2. Expose a global function window.renderFrame(t) that draws the frame at time t (seconds, 0 &amp;lt;= t &amp;lt; 12) from scratch. Everything on screen must be a pure function of t: no Date.now, no performance.now, no Math.random (use a seeded PRNG if you need noise), no accumulated state between calls, no requestAnimationFrame loop.
3. Also expose window.DURATION = 12 and window.FPS = 30. When the page is opened normally, play it once in a loop with requestAnimationFrame by calling renderFrame, so a human can preview it.
4. No external assets, fonts, images or libraries. Use system-ui for text.
5. The rhythm is 120 BPM: a beat every 0.5 s. Put every major cut, text entrance and impact on a beat time (a multiple of 0.5). List the beat plan as a comment at the top of the script.

The story (tell this, then decide the visuals):
A 12-second promo for "APIMODELS", an API gateway. Problem, then turn, then payoff:
- 0-3 s: a developer's screen is cluttered with many different API keys and SDK logos flying in from every side, overlapping, getting chaotic (draw them as simple rounded labels such as "image", "video", "LLM", "audio", "key_1", "key_2", "sdk", not real company logos).
- 3-4.5 s: everything snaps together on a beat and collapses into one glowing key.
- 4.5-9 s: from that one key, lines fan out to a grid of model cards that light up one per beat (label them "145 models", "image", "video", "LLM", "audio", "one endpoint").
- 9-12 s: clean end card: the word APIMODELS, the line "One key. Every model.", and "apimodels.app" small underneath; hold still for the last second.

Style: dark background, one accent colour (electric violet #7c5cff) plus white, motion with easing (easeOutCubic / easeInOutQuad written by you), subtle depth (scale and blur fall-off), no clutter in the final 3 seconds.

Before the code, do not explain. Return only the HTML.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three details mattered more than the wording:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ban every source of non-determinism by name.&lt;/strong&gt; &lt;code&gt;Date.now&lt;/code&gt;, &lt;code&gt;performance.now&lt;/code&gt;, &lt;code&gt;Math.random&lt;/code&gt; and accumulated state are the usual culprits. The model followed all four.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix the rhythm in numbers.&lt;/strong&gt; "120 BPM, cuts on multiples of 0.5 s, write the beat plan as a comment" produced a 24-beat plan, and every cut landed on one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Say what must not appear.&lt;/strong&gt; We asked for plain rounded labels instead of real company logos, which keeps the result usable.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Step 2: call Opus 5.5 with streaming and no small max_tokens
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.apimodels.app/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$APIMODELS_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;jq &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="nt"&gt;--rawfile&lt;/span&gt; p prompt.md &lt;span class="se"&gt;\&lt;/span&gt;
        &lt;span class="s1"&gt;'{model:"claude-opus-5-5", stream:true, messages:[{role:"user", content:$p}]}'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--no-buffer&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; stream.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things will bite you here if you skip them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Opus 5.5 thinks before it writes, and thinking counts as output.&lt;/strong&gt; On this prompt the first character of the answer arrived after &lt;strong&gt;110 seconds&lt;/strong&gt;, and the full reply took &lt;strong&gt;6 minutes 37 seconds&lt;/strong&gt;. Stream the request so the connection stays alive. A client with a 60-second timeout will give up long before the first token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not send a small &lt;code&gt;max_tokens&lt;/code&gt;.&lt;/strong&gt; The reply used &lt;strong&gt;31,786 output tokens&lt;/strong&gt; for about 16,000 characters of HTML, most of it thinking. A &lt;code&gt;max_tokens: 4096&lt;/code&gt; cap cuts the file off mid-script. If you omit the field, apimodels.app uses the model's own ceiling instead of inventing a small default.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Join the &lt;code&gt;delta.content&lt;/code&gt; pieces from &lt;code&gt;stream.txt&lt;/code&gt; and save the result as &lt;code&gt;promo.html&lt;/code&gt;. Open it in a browser first: it should play its own preview loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: render every frame with headless Chrome
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// render.mjs: pnpm add puppeteer-core (or npm i puppeteer-core)&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;puppeteer&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;puppeteer-core&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;mkdirSync&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;writeFileSync&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node:fs&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="nf"&gt;mkdirSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;frames&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;recursive&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;browser&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;puppeteer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;launch&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;executablePath&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/Applications/Google Chrome.app/Contents/MacOS/Google Chrome&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;browser&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;newPage&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setViewport&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;width&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1280&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;height&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;720&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;goto&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;file://&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cwd&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/promo.html&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;evaluate&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;requestAnimationFrame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="c1"&gt;// stop the preview loop&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;DURATION&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;FPS&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;evaluate&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;DURATION&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;FPS&lt;/span&gt; &lt;span class="p"&gt;}))&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;DURATION&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;FPS&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;png&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;evaluate&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;renderFrame&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;querySelector&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;canvas&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toDataURL&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;image/png&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;FPS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="nf"&gt;writeFileSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`frames/f&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;padStart&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;0&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;.png`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;png&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;base64&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;browser&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Linux, point &lt;code&gt;executablePath&lt;/code&gt; at your Chromium binary. 360 frames took 14 seconds on a laptop, and the page threw no errors.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: encode with ffmpeg
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-framerate&lt;/span&gt; 30 &lt;span class="nt"&gt;-i&lt;/span&gt; frames/f%04d.png &lt;span class="nt"&gt;-c&lt;/span&gt;:v libx264 &lt;span class="nt"&gt;-pix_fmt&lt;/span&gt; yuv420p &lt;span class="nt"&gt;-crf&lt;/span&gt; 20 &lt;span class="nt"&gt;-movflags&lt;/span&gt; +faststart promo.mp4
&lt;span class="c"&gt;# optional soundtrack, keep the BPM you gave the model:&lt;/span&gt;
&lt;span class="c"&gt;# ffmpeg -i promo.mp4 -i beat.mp3 -c:v copy -shortest promo-with-audio.mp4&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;yuv420p&lt;/code&gt; and &lt;code&gt;faststart&lt;/code&gt; are what make the file play everywhere, including in browsers and on phones. The result was a 1.8 MB, 12-second, 30 fps MP4.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: check frames at the story beats
&lt;/h2&gt;

&lt;p&gt;Before judging the video, pull frames at the moments the brief describes and put them side by side: 2.3 s (the clutter), 3.5 s (one key), 7.2 s (cards lighting up) and 11.3 s (end card).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwfc0jzg63v08a9sigujo.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwfc0jzg63v08a9sigujo.jpg" alt="Four frames at 2.3 s, 3.5 s, 7.2 s and 11.3 s, matching each section of the brief" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;All four matched on the first attempt. To change something, ask by timecode and beat ("at 7.5 s the sixth card should light up with a pulse") and ask for the whole file back, so the frame contract stays intact.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it cost
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-opus-5-5&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First token / total time&lt;/td&gt;
&lt;td&gt;110 s / 6 min 37 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output tokens&lt;/td&gt;
&lt;td&gt;31,786 (mostly thinking)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost at Anthropic list price ($4 / $20 per 1M tokens)&lt;/td&gt;
&lt;td&gt;about $0.64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost at apimodels.app list price ($2.40 / $12 per 1M tokens)&lt;/td&gt;
&lt;td&gt;about $0.38&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rendering and encoding&lt;/td&gt;
&lt;td&gt;local, free&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Output tokens dominate. A clip that needs 30,000 to 80,000 tokens of thinking and code costs roughly $0.36 to $0.96 per attempt at the lower rate, so budget for two or three attempts per finished clip.&lt;/p&gt;

&lt;h2&gt;
  
  
  When this approach is the wrong tool
&lt;/h2&gt;

&lt;p&gt;Code-rendered video is good at kinetic typography, UI walkthroughs, animated charts, logo reveals, explainers and social templates you can re-render with new text in seconds. It is the wrong tool for anything that has to look filmed: photoreal people, natural camera motion, skin, fabric, physics. For those shots use a text-to-video model and let Opus 5.5 write the shot list and prompts instead.&lt;/p&gt;

&lt;p&gt;It also has no audio (add it with ffmpeg), and the model never runs its own page. Treat the HTML as untested code: render it and look at the frames before you trust it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to go from here
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The model page for &lt;a href="https://apimodels.app/models/claude-opus-5-5" rel="noopener noreferrer"&gt;Claude Opus 5.5 on apimodels.app&lt;/a&gt; has the API parameters and this run's files.&lt;/li&gt;
&lt;li&gt;We collected &lt;a href="https://apimodels.app/claude-opus-5-5-video-prompts" rel="noopener noreferrer"&gt;60+ Opus 5.5 video prompts with the original posts and methods&lt;/a&gt;, sorted into motion design, product promos, explainers, music videos and 3D.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What would you render first with a &lt;code&gt;renderFrame(t)&lt;/code&gt; contract: a chart, a logo reveal, or a UI walkthrough? I'm curious which kind of clip breaks the approach.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This post was written from our own run's code and numbers; AI helped with structure and editing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Photoshop's Pen Tool Is Done: How AI Splits a Flat Image into Layers</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Tue, 06 Oct 2026 13:27:39 +0000</pubDate>
      <link>https://dev.to/super_lewis/photoshops-pen-tool-is-done-how-ai-splits-a-flat-image-into-layers-1pec</link>
      <guid>https://dev.to/super_lewis/photoshops-pen-tool-is-done-how-ai-splits-a-flat-image-into-layers-1pec</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://layergrab.com/blog/photoshop-is-dead" rel="noopener noreferrer"&gt;LayerGrab blog&lt;/a&gt;. Disclosure: I build LayerGrab, an AI tool that splits images into layers; this post was drafted with AI assistance and edited by me.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Separating a flat image into layers used to be the one job you could not hand off: Pen tool, Layer via Copy, Content-Aware Fill, repeat for every element. &lt;strong&gt;Image layer decomposition&lt;/strong&gt; is the AI task that does it in one pass: a model takes a flat RGB image and returns a stack of RGBA layers, one per element, with the regions behind each element filled in so the layers recompose into the original.&lt;/p&gt;

&lt;p&gt;Here is how it works, what we measured on a real poster today, and where it breaks.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the model has to do
&lt;/h2&gt;

&lt;p&gt;Three things at once, which is why this was hard until recently:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Decide what the elements are.&lt;/strong&gt; Not pixels by colour, but "headline", "badge", "cup".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cut each one out as RGBA&lt;/strong&gt; with a usable alpha edge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Invent what was behind each element&lt;/strong&gt;, because the camera or designer never stored it. Without this, moving a layer leaves a hole.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Segmentation models solve step 2 and inpainting models solve step 3. A decomposition model does all three, which is what makes the output a real layered file instead of cutouts on a broken background.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running the open model: what we measured
&lt;/h2&gt;

&lt;p&gt;Qwen-Image-Layered (Apache 2.0, weights released December 2025) is the open model for this. We ran it through fal's hosted API on an 880 × 1184 coffee shop poster on 2026-10-06. Example request, from fal's public docs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Example (fal queue API). FAL_KEY is your own key.&lt;/span&gt;
curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://queue.fal.run/fal-ai/qwen-image-layered &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Key &lt;/span&gt;&lt;span class="nv"&gt;$FAL_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"image_url": "https://example.com/poster.png", "num_layers": 4, "output_format": "png"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What came back with &lt;code&gt;num_layers: 4&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;4 PNGs, each &lt;strong&gt;544 × 736&lt;/strong&gt;, the full canvas size: the model works at around 640 px on the long side, and its README recommends that resolution.&lt;/li&gt;
&lt;li&gt;Layer 1 was the fully opaque background; layers 2–4 were transparent everywhere except their element.&lt;/li&gt;
&lt;li&gt;Text, the discount badge and three small leaves shared one layer; the cup got its own; one large leaf got its own.&lt;/li&gt;
&lt;li&gt;Stacked back together, the layers matched the downscaled original with a 1.5% mean pixel difference.&lt;/li&gt;
&lt;li&gt;18.6 s of inference, 65 s end to end including the queue.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With &lt;code&gt;num_layers: 2&lt;/code&gt;, the cup and headline stayed in the background layer, and the second layer held only the small decorations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pitfalls that decide whether the layers are usable
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You choose the layer count, and it groups by depth, not by object.&lt;/strong&gt; Too few layers and the main subject stays in the background. Start at 4–6.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resolution drops.&lt;/strong&gt; Expect ~640 px output. Upscale afterwards if you need print size.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VRAM.&lt;/strong&gt; Users in the model's GitHub issues report failures on 16 GB cards and trouble on 24 GB; ComfyUI's guide measured a 45 GB peak at 1024 px.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Text is pixels.&lt;/strong&gt; You get the headline on its own layer, not as editable type. Hide it and retype.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hosted APIs bill per output image.&lt;/strong&gt; fal charges per image returned, so 4 layers cost 4×.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When you don't want to run it yourself
&lt;/h2&gt;

&lt;p&gt;LayerGrab is an AI tool that splits one image into up to 16 separate transparent layers, one per element, names each layer from what it shows, fills the background behind them, and exports PNG, ZIP or a layered PSD. On the same poster it returned 10 named layers in 100 seconds at the original resolution. It also runs Qwen-Image-Layered itself, with a free 2-layer try, if you want to compare: &lt;a href="https://layergrab.com/qwen-image-layered-free" rel="noopener noreferrer"&gt;layergrab.com/qwen-image-layered-free&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Don't use it if your images must stay on your own machine: run the open model locally instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is Photoshop dead?
&lt;/h2&gt;

&lt;p&gt;As a finishing tool, no: we even ship a Photoshop plugin. As the place you go to cut a flat image apart by hand, yes. That job is now a model call.&lt;/p&gt;

&lt;p&gt;Have you tried any of these models on real client work? I'm curious where the output broke for you.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Claude Opus 5.5 vs GPT-6.1 Sol: which one to call, and how to measure it on your own prompts</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:51:37 +0000</pubDate>
      <link>https://dev.to/super_lewis/claude-opus-55-vs-gpt-61-sol-which-one-to-call-and-how-to-measure-it-on-your-own-prompts-2g8e</link>
      <guid>https://dev.to/super_lewis/claude-opus-55-vs-gpt-61-sol-which-one-to-call-and-how-to-measure-it-on-your-own-prompts-2g8e</guid>
      <description>&lt;p&gt;Claude Opus 5.5 and GPT-6.1 Sol are the two models most teams are choosing between this week. Opus 5.5 is Anthropic's everyday flagship, released on 22 September 2026. GPT-6.1 Sol is OpenAI's upgraded mid-tier model, released at DevDay on 29 September 2026. The short answer: &lt;strong&gt;pick GPT-6.1 Sol for high-volume and cost-sensitive work, and Claude Opus 5.5 for long agentic coding and tasks where a wrong answer costs more than the tokens.&lt;/strong&gt; The longer answer depends on reasoning effort, and you can check it on your own prompts in about ten minutes.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure up front: I work on apimodels.app, a multi-model API gateway that serves both models. Prices for it below are its public prices on 1 October 2026. This post was drafted with AI assistance and checked against the vendors' pages.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How I compared them
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;List price per 1M tokens&lt;/strong&gt;, from the vendors' pages on 1 October 2026.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Specs that change your code&lt;/strong&gt;: context window, max output, reasoning controls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Benchmarks both models were run on in the same table.&lt;/strong&gt; The labs mostly publish different suites, so I only use rows where both appear, and I name who ran them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost per call&lt;/strong&gt; for a realistic agent turn, computed from the prices above.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  At a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Claude Opus 5.5&lt;/th&gt;
&lt;th&gt;GPT-6.1 Sol&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Released&lt;/td&gt;
&lt;td&gt;22 Sep 2026 (Anthropic)&lt;/td&gt;
&lt;td&gt;29 Sep 2026 (OpenAI DevDay)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;List price, input / output&lt;/td&gt;
&lt;td&gt;$4 / $20&lt;/td&gt;
&lt;td&gt;$2 / $10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cached input (list)&lt;/td&gt;
&lt;td&gt;$0.20 read, $5 write&lt;/td&gt;
&lt;td&gt;$0.10 read, $2.50 write&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context / max output&lt;/td&gt;
&lt;td&gt;1M / 128K&lt;/td&gt;
&lt;td&gt;1.05M / 128K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge cutoff&lt;/td&gt;
&lt;td&gt;June 2026&lt;/td&gt;
&lt;td&gt;30 April 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning control&lt;/td&gt;
&lt;td&gt;effort low to max, default medium, always thinks&lt;/td&gt;
&lt;td&gt;reasoning_effort low to max, default medium, no &lt;code&gt;none&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long context surcharge&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;above 272K input: input and cache 2×, output 1.5×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;long agent sessions, review, careful writing&lt;/td&gt;
&lt;td&gt;volume, extraction, cost-capped agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Not for&lt;/td&gt;
&lt;td&gt;high-volume bulk classification&lt;/td&gt;
&lt;td&gt;code that relies on &lt;code&gt;reasoning_effort: none&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The benchmarks they share
&lt;/h2&gt;

&lt;p&gt;OpenAI's GPT-6.1 Sol launch included Opus 5.5 in several rows. These are OpenAI's runs and OpenAI's choice of benchmarks, so read them as a vendor's claim:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark (OpenAI launch)&lt;/th&gt;
&lt;th&gt;Effort&lt;/th&gt;
&lt;th&gt;GPT-6.1 Sol&lt;/th&gt;
&lt;th&gt;Claude Opus 5.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AutomationBench&lt;/td&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;31.7%&lt;/td&gt;
&lt;td&gt;29.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AutomationBench&lt;/td&gt;
&lt;td&gt;max&lt;/td&gt;
&lt;td&gt;36.1% ($0.30/task)&lt;/td&gt;
&lt;td&gt;42.5% ($1.44/task)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GDP.pdf&lt;/td&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;30.0% ($0.34/task)&lt;/td&gt;
&lt;td&gt;25.6% ($0.80/task)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench Science 0.1&lt;/td&gt;
&lt;td&gt;max&lt;/td&gt;
&lt;td&gt;57.0% ($5.47/task)&lt;/td&gt;
&lt;td&gt;63.3% ($23.21/task)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern is consistent: &lt;strong&gt;at medium effort GPT-6.1 Sol edges ahead, at max effort Opus 5.5 pulls ahead, and Sol costs less per task in every row.&lt;/strong&gt; Independent aggregates point the same way on capability. Reported max-effort figures put Opus 5.5 at 58 and GPT-6.1 Sol at 52 on the Artificial Analysis Intelligence Index, with cost per index task of $5.98 and $0.72.&lt;/p&gt;

&lt;p&gt;Anthropic's own launch numbers for Opus 5.5 (66.4% on Terminal-Bench 4.0, 57.8% on CursorBench 4.0) have no GPT-6.1 Sol row, so they don't settle the head-to-head.&lt;/p&gt;

&lt;h2&gt;
  
  
  What one agent turn costs
&lt;/h2&gt;

&lt;p&gt;Take a coding-agent turn that re-sends 50,000 input tokens, 40,000 of them cached, and returns 2,000 output tokens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Claude Opus 5.5&lt;/th&gt;
&lt;th&gt;GPT-6.1 Sol&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;At list price&lt;/td&gt;
&lt;td&gt;$0.088&lt;/td&gt;
&lt;td&gt;$0.044&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;On apimodels.app&lt;/td&gt;
&lt;td&gt;$0.053&lt;/td&gt;
&lt;td&gt;$0.022&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things move these numbers in practice. Opus 5.5 always thinks, so its reasoning tokens land in the output bill even on short answers. GPT-6.1 Sol at &lt;code&gt;low&lt;/code&gt; effort spends far fewer reasoning tokens, which is where most of its cost-per-task advantage comes from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure it on your own prompts
&lt;/h2&gt;

&lt;p&gt;Benchmarks are someone else's tasks. The fastest way to decide is to send the same twenty prompts to both models and compare answers and cost. Both are OpenAI-compatible on apimodels.app, so one client covers both, and every response carries the billed amount in an &lt;code&gt;x-apimodels-cost&lt;/code&gt; header.&lt;/p&gt;

&lt;p&gt;Tested with Python 3.11 and &lt;code&gt;openai&lt;/code&gt; 1.x:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;APIMODELS_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.apimodels.app/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;MODELS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6.1-sol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;t0&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;with_raw_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;t0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;out&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completion_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;billed_usd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x-apimodels-cost&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;  &lt;span class="c1"&gt;# what you were actually charged
&lt;/span&gt;        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;prompts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompts.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;prompts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;MODELS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;ask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;s  in=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; out=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;out&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  $&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;res&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;billed_usd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Score the answers yourself, or with a third model as judge, then divide the total billed cost by the number of answers you would ship. That number is the one that matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gotchas when switching
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;reasoning_effort: "none"&lt;/code&gt; is not on GPT-6.1 Sol's list.&lt;/strong&gt; OpenAI's model page lists low, medium, high, xhigh and max. Code written for GPT-6 Sol that sends &lt;code&gt;none&lt;/code&gt; should send &lt;code&gt;low&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Small &lt;code&gt;max_tokens&lt;/code&gt; can return an empty answer.&lt;/strong&gt; Reasoning counts against the budget on both models. If &lt;code&gt;finish_reason&lt;/code&gt; is &lt;code&gt;length&lt;/code&gt; and the content is empty, raise &lt;code&gt;max_tokens&lt;/code&gt; or lower the effort.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Long prompts change Sol's price.&lt;/strong&gt; Above 272K input tokens the whole request is billed at 2× input and 1.5× output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Opus 5.5 can't be made to cost less by turning thinking off.&lt;/strong&gt; If a task is simple enough that thinking is waste, it is probably a Sol task.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Which one should you pick
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Classification, extraction, routing, bulk summarisation → &lt;strong&gt;GPT-6.1 Sol&lt;/strong&gt; at low effort.&lt;/li&gt;
&lt;li&gt;Agents with a hard cost cap per task → &lt;strong&gt;GPT-6.1 Sol&lt;/strong&gt; at medium effort.&lt;/li&gt;
&lt;li&gt;Long agentic coding sessions, code review, careful long-form writing → &lt;strong&gt;Claude Opus 5.5&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The hardest reasoning tasks → run both at max effort on twenty real examples before you commit.&lt;/li&gt;
&lt;li&gt;You need a direct contract with OpenAI or Anthropic, such as zero data retention or a regulated-industry agreement → &lt;strong&gt;call the vendor directly&lt;/strong&gt;, not a gateway.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  About apimodels.app
&lt;/h2&gt;

&lt;p&gt;apimodels.app is a multi-model API gateway: one key and one OpenAI-compatible endpoint for 140+ image, video, audio and language models. When an upstream channel fails, requests switch to another channel automatically, failed calls are not charged, and long-running jobs can report back through async callbacks. On price, on 1 October 2026 it lists &lt;a href="https://apimodels.app/models/claude-opus-5-5?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=opus55-vs-gpt61sol" rel="noopener noreferrer"&gt;Claude Opus 5.5 at $2.40 / $12 per 1M tokens&lt;/a&gt;, 40% below Anthropic's list, and &lt;a href="https://apimodels.app/models/gpt-6.1-sol?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=opus55-vs-gpt61sol" rel="noopener noreferrer"&gt;GPT-6.1 Sol at $1 / $5&lt;/a&gt;, half of OpenAI's list, billed per token. It is not the right choice if you only ever call one model and already have the vendor's SDK and contract in place.&lt;/p&gt;

&lt;p&gt;Which prompts in your own workload flip between the two models? I'd like to hear where your results disagree with the launch tables.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
      <category>openai</category>
    </item>
    <item>
      <title>We Replaced a 2,300-Character Image Prompt With One Sentence</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Sun, 27 Sep 2026 02:02:56 +0000</pubDate>
      <link>https://dev.to/super_lewis/we-replaced-a-2300-character-image-prompt-with-one-sentence-1lh0</link>
      <guid>https://dev.to/super_lewis/we-replaced-a-2300-character-image-prompt-with-one-sentence-1lh0</guid>
      <description>&lt;p&gt;Our prompt for making AI images look real used to be 2,300 characters long. It listed every AI tell to remove, every camera detail to add, and rules for colour, shadows and lens optics. It produced faces covered in freckles and skin so drained it looked like a corpse. The version in production today is one sentence.&lt;/p&gt;

&lt;p&gt;The lesson generalises to any image-editing prompt: a modern image model already knows what a phone photo looks like, so naming a familiar camera does the work of a page of instructions, and every extra clause is something the model can over-apply.&lt;/p&gt;

&lt;p&gt;sparkpix.ai's AI Image Humanizer is a web tool that re-renders an AI-generated image so it reads like a real phone-camera photo, keeping pose and framing. Disclosure: I build sparkpix.ai. This post was drafted with AI assistance for structure and wording; the prompt text, comments and code are from the sparkpix codebase (Next.js App Router on Vercel, calling GPT Image 2 through apimodels.app, September 2026).&lt;/p&gt;

&lt;h2&gt;
  
  
  Long prompts fail by arguing with themselves
&lt;/h2&gt;

&lt;p&gt;Every clause in the long prompt was added to fix a problem the previous clause created. "Add natural skin texture" produced freckles; "don't add freckles" produced flat skin; "keep healthy colour" fought "remove the saturated AI grade". The model obeyed every clause a little too hard, and the clauses pulled against each other.&lt;/p&gt;

&lt;p&gt;What replaced it, tested directly against GPT Image 2 and preferred by eye on 2026-07-30:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;DE_AI_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Remove the AI-generated look from this image so it looks natural — for a &lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;photographic image, like it was taken with an iPhone 12 camera. Keep the &lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;original art style: an illustration or anime image must stay an &lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;illustration, not become a photograph.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;"Like it was taken with an iPhone 12 camera" carries the sensor noise, the colour science, the skin rendering and the slightly imperfect light, all at once, because the model has seen millions of those photos. There is no clause left to over-apply.&lt;/p&gt;

&lt;h2&gt;
  
  
  One word can change the subject instead of the camera
&lt;/h2&gt;

&lt;p&gt;We tried "selfie" in that sentence, and it re-posed people. "Selfie" does not only describe a camera; it describes a photo genre: arm raised, head tilted, shot close and from above. The model read it as an instruction about the person and changed the pose the user had already chosen.&lt;/p&gt;

&lt;p&gt;The rule we took from it: &lt;strong&gt;name the device, never the genre.&lt;/strong&gt; "iPhone 12 camera" is about the lens. "Selfie", "portrait shot", "street photo" are about the picture, and the model will happily give you a new picture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scope a style instruction, don't negate it
&lt;/h2&gt;

&lt;p&gt;The unscoped sentence, "make it look like an iPhone photo", is also an instruction to convert the medium, and anime uploads started coming back as photographs. The model was doing exactly what it was told.&lt;/p&gt;

&lt;p&gt;The tempting fix is to append "do not change the art style". That recreates the original failure: two clauses arguing. Instead, the camera reference is scoped to photographic input ("for a photographic image, like it was taken with…"), and the second sentence adds a constraint rather than a contradiction.&lt;/p&gt;

&lt;p&gt;Honest status: the scoped version has not been through the same side-by-side test as the one-sentence version. It is in production because it fixed the anime-to-photo reports, not because it was benchmarked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the prompt in exactly one place
&lt;/h2&gt;

&lt;p&gt;The first time we shortened the prompt, nothing changed for real users. The wording lived in three places: two tool defaults on the server, and a pre-filled textarea on each landing page. The landing-page copy is what actually reached the model, because the editor sends its textarea as &lt;code&gt;prompt&lt;/code&gt; and the server only falls back to its default when &lt;code&gt;prompt&lt;/code&gt; is empty. We had edited the fallback.&lt;/p&gt;

&lt;p&gt;Now there is one exported constant, imported by both the tool config and the pages, and the old long prompt is not kept as a "fallback", because an unused second copy is how the drift started. It is in git history if anyone needs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failures must refund exactly once
&lt;/h2&gt;

&lt;p&gt;A humanizer call can fail after credits are taken: the model's safety filter refuses the image, or the upstream API errors. The refund has to happen once, even though a failure can be observed by several code paths (the background job, a client poll, a cleanup cron). The pattern is to let the database decide who gets to refund:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;failAndRefund&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;errorMessage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s2"&gt;`UPDATE generations SET status = 'failed', error_message = $1
     WHERE id = $2 AND status = 'processing'`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;errorMessage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;substring&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nx"&gt;generationId&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rowCount&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rowCount&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;addCredits&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;creditsCharged&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only the caller whose &lt;code&gt;UPDATE&lt;/code&gt; actually flips the row from &lt;code&gt;processing&lt;/code&gt; to &lt;code&gt;failed&lt;/code&gt; sees &lt;code&gt;rowCount &amp;gt; 0&lt;/code&gt;, so only that caller refunds. A &lt;code&gt;SELECT&lt;/code&gt; followed by an &lt;code&gt;UPDATE&lt;/code&gt; looks equivalent and is not: two pollers can both read &lt;code&gt;processing&lt;/code&gt; before either writes. We learned that the expensive way, with a polling endpoint that refunded the same failure on every poll.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this approach falls short
&lt;/h2&gt;

&lt;p&gt;A one-sentence prompt gives the user no dial. If you need to control grain strength or skin detail per image, a slider-based tool is the better fit. A generative re-render can also shift very fine details, and GPT Image 2's safety filter refused about 1 in 20 of these requests in early September 2026, which is why the refund path matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why an iPhone 12 and not the latest model?&lt;/strong&gt; Nothing magic; it is a very common, well-represented camera. Users do edit it: we see iPhone 13 to 17 in their prompts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does naming a camera work with other models?&lt;/strong&gt; We only tested GPT Image 2. The idea, naming a familiar device instead of describing its output, should transfer, but test it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not just add grain in post?&lt;/strong&gt; Grain fixes the surface only. Plastic lighting, colour grades and over-smooth skin are structural, and a texture overlay cannot touch them.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The tool is at &lt;a href="https://sparkpix.ai/ai-image-humanizer" rel="noopener noreferrer"&gt;sparkpix.ai/ai-image-humanizer&lt;/a&gt;, 5 credits per image with the first image free. For how it compares with nine other tools, see &lt;a href="https://sparkpix.ai/blog/best-ai-image-humanizers" rel="noopener noreferrer"&gt;Best AI Image Humanizers in 2026&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>promptengineering</category>
      <category>software</category>
    </item>
    <item>
      <title>How to Read C2PA and IPTC AI Labels From Image Bytes in Node.js</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Sat, 26 Sep 2026 14:42:19 +0000</pubDate>
      <link>https://dev.to/super_lewis/how-to-read-c2pa-and-iptc-ai-labels-from-image-bytes-in-nodejs-5bcl</link>
      <guid>https://dev.to/super_lewis/how-to-read-c2pa-and-iptc-ai-labels-from-image-bytes-in-nodejs-5bcl</guid>
      <description>&lt;p&gt;Most AI images carry a note from the tool that made them, and you can read it with about 60 lines of TypeScript. The hard part is not parsing it; it is knowing what its absence means, which is nothing at all.&lt;/p&gt;

&lt;p&gt;sparkpix.ai's AI Image Detector is a web tool that estimates whether an image was AI-generated by reading provenance labels in the file first and then scoring the pixels. This post walks through how the first layer works in Node.js, how the two layers combine into a verdict, and the one upload bug that silently deletes the evidence.&lt;/p&gt;

&lt;p&gt;Disclosure: I build sparkpix.ai and apimodels.app, and both appear below. This post was drafted with AI assistance for structure and wording; the code is from the sparkpix codebase (Next.js App Router, Node 24, September 2026).&lt;/p&gt;

&lt;h2&gt;
  
  
  Two kinds of label live inside AI images
&lt;/h2&gt;

&lt;p&gt;AI tools write provenance into the file in two standard ways, and a detector should look for both.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;C2PA Content Credentials.&lt;/strong&gt; A signed manifest, defined by the Coalition for Content Provenance and Authenticity, that records which tool produced the file and what was done to it. OpenAI, Adobe Firefly and Microsoft embed one. It sits in a JUMBF box: JPEG APP11 segments, a &lt;code&gt;caBX&lt;/code&gt; chunk in PNG, a &lt;code&gt;C2PA&lt;/code&gt; chunk in WebP.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IPTC Digital Source Type.&lt;/strong&gt; An XMP field whose value &lt;code&gt;trainedAlgorithmicMedia&lt;/code&gt; means "created by a generative model" and &lt;code&gt;compositeWithTrainedAlgorithmicMedia&lt;/code&gt; means "edited with one". Google writes it into Gemini and Imagen images, and C2PA action assertions reuse the same vocabulary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When either is present, the question is settled: the maker declared it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A byte scan is enough to detect presence
&lt;/h2&gt;

&lt;p&gt;You do not need a full C2PA validator to answer "does this file claim to be AI?". We deliberately scan bytes and report presence plus the claimed generator, not signature validity:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;readMetadataSignals&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;MetadataSignals&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;latin&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;latin1&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;c2pa&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;latin&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;jumb&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;latin&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;c2pa&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
    &lt;span class="nx"&gt;latin&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;caBX&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
    &lt;span class="sr"&gt;/C2PA&lt;/span&gt;&lt;span class="se"&gt;[\s\S]{0,8}&lt;/span&gt;&lt;span class="sr"&gt;jumb/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;latin&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;composite&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;latin&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;compositeWithTrainedAlgorithmicMedia&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;generated&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;composite&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;latin&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;trainedAlgorithmicMedia&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;generator&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;c2pa&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nf"&gt;claimGenerator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;generator&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="sr"&gt;/Made with Google AI/i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;latin&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="nx"&gt;generator&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Google AI&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;c2pa&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;generator&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;aiSource&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;composite&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;composite&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;generated&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;generated&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note the order of the two IPTC checks: &lt;code&gt;compositeWithTrainedAlgorithmicMedia&lt;/code&gt; contains &lt;code&gt;trainedAlgorithmicMedia&lt;/code&gt; as a substring, so the composite test has to run first or every AI-edited photo reads as fully generated.&lt;/p&gt;

&lt;p&gt;The generator name comes from the C2PA claim, which is CBOR. Newer (v2) claims store it as &lt;code&gt;claim_generator_info: [{ name: "ChatGPT" }]&lt;/code&gt;; older (v1) claims use a single &lt;code&gt;claim_generator&lt;/code&gt; string such as &lt;code&gt;Adobe_Photoshop/25.0 adobe_c2pa/0.7.6&lt;/code&gt;. Reading a CBOR text string by hand is a few lines, because the first byte tells you the length:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;readCborText&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;at&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;at&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mh"&gt;0x60&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mh"&gt;0x77&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mh"&gt;0x60&lt;/span&gt;      &lt;span class="c1"&gt;// length in the type byte&lt;/span&gt;
  &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mh"&gt;0x78&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;at&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;at&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;            &lt;span class="c1"&gt;// 1-byte length&lt;/span&gt;
  &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mh"&gt;0x79&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;readUInt16BE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;at&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;at&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="c1"&gt;// 2-byte length&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;buf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;utf8&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;len&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What this does not do:&lt;/strong&gt; it does not verify the signature, so a forged manifest would pass. For "is this probably AI?" that is an acceptable trade; for legal provenance, use the official c2pa-rs or c2pa-node libraries.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your upload pipeline may be deleting the evidence
&lt;/h2&gt;

&lt;p&gt;The most common reason a detector finds no label is that your own code stripped it. Our site compresses every upload in the browser and re-encodes it to WebP before sending it to storage, which is right for an image editor and fatal for a detector: re-encoding drops the metadata segments where both labels live. The detector's upload path skips that optimizer for any file up to 10 MB and stores it byte-for-byte; larger files still get compressed, and the user is told the metadata check may miss.&lt;/p&gt;

&lt;p&gt;Social platforms and chat apps do the same thing to everyone, and a screenshot creates a brand-new file. That is why a missing label proves nothing, and why a second layer is needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second layer scores the pixels
&lt;/h2&gt;

&lt;p&gt;When there is no label, a classifier looks at texture, noise and structure. We call an external model for this, apimodels.app's &lt;code&gt;ai-image-detector&lt;/code&gt; (POST &lt;code&gt;/v1/images/detections&lt;/code&gt;, synchronous, about $0.015 per image), which also reads C2PA itself and returns a 0-1 score:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Upload is stored unmodified, then two checks run.&lt;/li&gt;
&lt;li&gt;Byte scan for C2PA and IPTC labels. If IPTC says generated, the verdict is &lt;strong&gt;AI Generated (99%+)&lt;/strong&gt;; if it says composite, &lt;strong&gt;Digitally Edited&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Otherwise the pixel classifier's score picks the band: 85-100 &lt;strong&gt;AI Generated&lt;/strong&gt;, 60-84 &lt;strong&gt;Likely AI&lt;/strong&gt;, 40-59 &lt;strong&gt;Uncertain&lt;/strong&gt;, 16-39 &lt;strong&gt;Likely Real&lt;/strong&gt;, 0-15 &lt;strong&gt;Real Photo&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A metadata declaration always wins over the pixel score, because it is the maker's own statement. Without either signal, the endpoint refuses to guess: if the classifier is down and the file has no label, the user gets an error and their credits back, not a made-up "Uncertain".&lt;/p&gt;

&lt;h2&gt;
  
  
  How accurate is it?
&lt;/h2&gt;

&lt;p&gt;On our own set of 28 images, 14 AI images from GPT Image 2 and Gemini plus the 14 real photos they were modelled on, the combined detector got 27 right: no real photo was flagged and one AI image passed as real. Calls took 1.6 to 3.1 seconds. The classifier returned near-binary scores (0.99 or 0.001), and it named the generator on none of the 28, so attribution is best-effort.&lt;/p&gt;

&lt;p&gt;That test is small and covers two generators. Don't read it as a benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  When not to build this
&lt;/h2&gt;

&lt;p&gt;If you need provenance that stands up in court or in a newsroom's publishing workflow, a byte scan is the wrong tool: use a real C2PA validator that checks signatures and the trust list. And if your images come from social media, skip the metadata layer's expectations entirely; almost everything will arrive stripped.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can I detect AI images without calling any API?&lt;/strong&gt; Partly. The metadata layer is free and local, and it is decisive when a label exists. It answers nothing for stripped files, which is most images shared online.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does SynthID show up in this scan?&lt;/strong&gt; No. Google's SynthID is an invisible watermark in the pixels, not a metadata field, and there is no public API to read it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is 40-59% labelled "Uncertain" instead of a lean?&lt;/strong&gt; Because a classifier score near the middle carries almost no information, and presenting it as "slightly AI" invites people to act on noise.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Try the detector at &lt;a href="https://sparkpix.ai/ai-image-detector" rel="noopener noreferrer"&gt;sparkpix.ai/ai-image-detector&lt;/a&gt;. For the non-code version, including reverse image search and visual checks, see &lt;a href="https://sparkpix.ai/blog/how-to-tell-if-an-image-is-ai-generated" rel="noopener noreferrer"&gt;How to Tell If an Image Is AI-Generated&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>node</category>
      <category>programming</category>
      <category>typescript</category>
    </item>
    <item>
      <title>We Replaced a 2,300-Character Image Prompt With One Sentence</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Sat, 26 Sep 2026 14:41:17 +0000</pubDate>
      <link>https://dev.to/super_lewis/we-replaced-a-2300-character-image-prompt-with-one-sentence-3hl6</link>
      <guid>https://dev.to/super_lewis/we-replaced-a-2300-character-image-prompt-with-one-sentence-3hl6</guid>
      <description>&lt;p&gt;Our prompt for making AI images look real used to be 2,300 characters long. It listed every AI tell to remove, every camera detail to add, and rules for colour, shadows and lens optics. It produced faces covered in freckles and skin so drained it looked like a corpse. The version in production today is one sentence.&lt;/p&gt;

&lt;p&gt;The lesson generalises to any image-editing prompt: a modern image model already knows what a phone photo looks like, so naming a familiar camera does the work of a page of instructions, and every extra clause is something the model can over-apply.&lt;/p&gt;

&lt;p&gt;sparkpix.ai's AI Image Humanizer is a web tool that re-renders an AI-generated image so it reads like a real phone-camera photo, keeping pose and framing. Disclosure: I build sparkpix.ai. This post was drafted with AI assistance for structure and wording; the prompt text, comments and code are from the sparkpix codebase (Next.js App Router on Vercel, calling GPT Image 2 through apimodels.app, September 2026).&lt;/p&gt;

&lt;h2&gt;
  
  
  Long prompts fail by arguing with themselves
&lt;/h2&gt;

&lt;p&gt;Every clause in the long prompt was added to fix a problem the previous clause created. "Add natural skin texture" produced freckles; "don't add freckles" produced flat skin; "keep healthy colour" fought "remove the saturated AI grade". The model obeyed every clause a little too hard, and the clauses pulled against each other.&lt;/p&gt;

&lt;p&gt;What replaced it, tested directly against GPT Image 2 and preferred by eye on 2026-07-30:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;DE_AI_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Remove the AI-generated look from this image so it looks natural — for a &lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;photographic image, like it was taken with an iPhone 12 camera. Keep the &lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;original art style: an illustration or anime image must stay an &lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;illustration, not become a photograph.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;"Like it was taken with an iPhone 12 camera" carries the sensor noise, the colour science, the skin rendering and the slightly imperfect light, all at once, because the model has seen millions of those photos. There is no clause left to over-apply.&lt;/p&gt;

&lt;h2&gt;
  
  
  One word can change the subject instead of the camera
&lt;/h2&gt;

&lt;p&gt;We tried "selfie" in that sentence, and it re-posed people. "Selfie" does not only describe a camera; it describes a photo genre: arm raised, head tilted, shot close and from above. The model read it as an instruction about the person and changed the pose the user had already chosen.&lt;/p&gt;

&lt;p&gt;The rule we took from it: &lt;strong&gt;name the device, never the genre.&lt;/strong&gt; "iPhone 12 camera" is about the lens. "Selfie", "portrait shot", "street photo" are about the picture, and the model will happily give you a new picture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scope a style instruction, don't negate it
&lt;/h2&gt;

&lt;p&gt;The unscoped sentence, "make it look like an iPhone photo", is also an instruction to convert the medium, and anime uploads started coming back as photographs. The model was doing exactly what it was told.&lt;/p&gt;

&lt;p&gt;The tempting fix is to append "do not change the art style". That recreates the original failure: two clauses arguing. Instead, the camera reference is scoped to photographic input ("for a photographic image, like it was taken with…"), and the second sentence adds a constraint rather than a contradiction.&lt;/p&gt;

&lt;p&gt;Honest status: the scoped version has not been through the same side-by-side test as the one-sentence version. It is in production because it fixed the anime-to-photo reports, not because it was benchmarked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the prompt in exactly one place
&lt;/h2&gt;

&lt;p&gt;The first time we shortened the prompt, nothing changed for real users. The wording lived in three places: two tool defaults on the server, and a pre-filled textarea on each landing page. The landing-page copy is what actually reached the model, because the editor sends its textarea as &lt;code&gt;prompt&lt;/code&gt; and the server only falls back to its default when &lt;code&gt;prompt&lt;/code&gt; is empty. We had edited the fallback.&lt;/p&gt;

&lt;p&gt;Now there is one exported constant, imported by both the tool config and the pages, and the old long prompt is not kept as a "fallback", because an unused second copy is how the drift started. It is in git history if anyone needs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failures must refund exactly once
&lt;/h2&gt;

&lt;p&gt;A humanizer call can fail after credits are taken: the model's safety filter refuses the image, or the upstream API errors. The refund has to happen once, even though a failure can be observed by several code paths (the background job, a client poll, a cleanup cron). The pattern is to let the database decide who gets to refund:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;failAndRefund&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;errorMessage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s2"&gt;`UPDATE generations SET status = 'failed', error_message = $1
     WHERE id = $2 AND status = 'processing'`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;errorMessage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;substring&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nx"&gt;generationId&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rowCount&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rowCount&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;addCredits&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;creditsCharged&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only the caller whose &lt;code&gt;UPDATE&lt;/code&gt; actually flips the row from &lt;code&gt;processing&lt;/code&gt; to &lt;code&gt;failed&lt;/code&gt; sees &lt;code&gt;rowCount &amp;gt; 0&lt;/code&gt;, so only that caller refunds. A &lt;code&gt;SELECT&lt;/code&gt; followed by an &lt;code&gt;UPDATE&lt;/code&gt; looks equivalent and is not: two pollers can both read &lt;code&gt;processing&lt;/code&gt; before either writes. We learned that the expensive way, with a polling endpoint that refunded the same failure on every poll.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this approach falls short
&lt;/h2&gt;

&lt;p&gt;A one-sentence prompt gives the user no dial. If you need to control grain strength or skin detail per image, a slider-based tool is the better fit. A generative re-render can also shift very fine details, and GPT Image 2's safety filter refused about 1 in 20 of these requests in early September 2026, which is why the refund path matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why an iPhone 12 and not the latest model?&lt;/strong&gt; Nothing magic; it is a very common, well-represented camera. Users do edit it: we see iPhone 13 to 17 in their prompts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does naming a camera work with other models?&lt;/strong&gt; We only tested GPT Image 2. The idea, naming a familiar device instead of describing its output, should transfer, but test it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not just add grain in post?&lt;/strong&gt; Grain fixes the surface only. Plastic lighting, colour grades and over-smooth skin are structural, and a texture overlay cannot touch them.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The tool is at &lt;a href="https://sparkpix.ai/ai-image-humanizer" rel="noopener noreferrer"&gt;sparkpix.ai/ai-image-humanizer&lt;/a&gt;, 5 credits per image with the first image free. For how it compares with nine other tools, see &lt;a href="https://sparkpix.ai/blog/best-ai-image-humanizers" rel="noopener noreferrer"&gt;Best AI Image Humanizers in 2026&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>showdev</category>
      <category>web</category>
    </item>
    <item>
      <title>14,000 production calls later: image and video APIs fail in two opposite ways</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:19:54 +0000</pubDate>
      <link>https://dev.to/super_lewis/14000-production-calls-later-image-and-video-apis-fail-in-two-opposite-ways-51d4</link>
      <guid>https://dev.to/super_lewis/14000-production-calls-later-image-and-video-apis-fail-in-two-opposite-ways-51d4</guid>
      <description>&lt;p&gt;Media generation endpoints fail in two ways that look identical in your logs and need opposite handling. One is a content filter refusing the prompt, which no amount of retrying will fix. The other is a saturated GPU pool, which almost always clears within seconds. If you lump them into one &lt;code&gt;catch&lt;/code&gt; block, you will both burn money retrying rejected prompts and give up on requests that would have succeeded.&lt;/p&gt;

&lt;p&gt;I pulled the numbers to see how the split actually falls, across 14,069 image and video generations over 30 days. The ratio is not close, and it inverts completely between the two media types. Disclosure up front: the data is from apimodels.app, which is the API platform I work on, so treat it as one operator's dataset rather than an industry benchmark. Everything below is written so you can run the same measurement against whatever provider you use.&lt;/p&gt;

&lt;h2&gt;
  
  
  The split
&lt;/h2&gt;

&lt;p&gt;The two tiers and the video model below are the same three models the code at the end calls. For image generation, content rejections were 90% of all failures. For video generation on the same platform in the same window, they were 4%.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                        content  capacity  timeout  upstream  bad input
image model, fast tier      823        17       29        20         24
image model, detail tier    232        69       44        16         14
video model                   2        31        3          9         6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Turned into rates against total calls, the picture is sharper:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model                calls   success   infra failures   p50      p90
image, fast tier    10,149     91.0%            0.65%   40.1s    71.2s
image, detail tier   3,630     89.7%            3.55%   52.0s   135.9s
video                  290     82.4%           14.83%   36.5s   104.1s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fast image tier looks unreliable at 91% success and is not. Strip out the prompts a filter refused and actual infrastructure failure is 66 calls in 10,149. Meanwhile the video model, which looks worse at 82.4%, has essentially no content problem at all: it failed 43 times because capacity was not there.&lt;/p&gt;

&lt;p&gt;That inversion has a boring physical cause. Image inference finishes in tens of seconds and the filters sit on a text prompt, which users write freely and sometimes carelessly. Video inference holds a GPU for a minute or more, so pools saturate under load and "come back later" is a normal operating state rather than an incident. Any provider running the same hardware economics will show some version of this shape.&lt;/p&gt;

&lt;p&gt;One honest caveat on my own numbers: these models are served through more than one upstream provider, so the capacity column partly reflects routing between them rather than a single vendor's raw availability. Your numbers against a single-vendor API will look different in magnitude. The two-class split is the part that transfers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure it on your own provider
&lt;/h2&gt;

&lt;p&gt;You need failures bucketed into "the user can fix this" and "time can fix this", and most APIs will not hand you that distinction. They return a 4xx or 5xx and a message string.&lt;/p&gt;

&lt;p&gt;The mapping that has held up for us is three buckets, not two, because the third one is the one that wakes people up at night:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="c1"&gt;// The user can fix this. Never retry. Show them the provider's own wording.&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/safety|content policy|moderation|nsfw|prohibited|violat/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CONTENT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;400&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="sr"&gt;/invalid|unsupported|too large|must be/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;BAD_INPUT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// Time can fix this. Retry with backoff.&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;429&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="sr"&gt;/busy|capacity|overload|rate limit|try again/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CAPACITY&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="sr"&gt;/timeout|timed out|econnreset|gateway/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;TRANSIENT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// Nobody knows yet. This is the bucket you alert on.&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;UNKNOWN&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Alert on the size of &lt;code&gt;UNKNOWN&lt;/code&gt;, not on your overall error rate. A rising &lt;code&gt;UNKNOWN&lt;/code&gt; means the provider changed their error wording and your retry logic has quietly stopped working. A rising &lt;code&gt;CONTENT&lt;/code&gt; means your users changed, or someone shipped a prompt template that trips a filter, and no amount of infrastructure work will help.&lt;/p&gt;

&lt;p&gt;I am not going to pretend this regex table is elegant. It is hand-maintained, it drifts, and every provider that collapses both classes into a generic &lt;code&gt;500 Internal Error&lt;/code&gt; makes it worse. It is still better than one retry policy for everything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The retry loop the split implies
&lt;/h2&gt;

&lt;p&gt;Retry capacity and transient failures with exponential backoff, and fail content rejections immediately with the message attached.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;RETRYABLE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CAPACITY&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;TRANSIENT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;tries&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;tries&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;submitAndPoll&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;bucket&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;RETRYABLE&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;bucket&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;UserFacingError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;bucket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// 2s, 4s, 8s&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;still failing after retries&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three attempts at 2/4/8 seconds covers the capacity failures we see, because a saturated pool usually has a slot within ten seconds. Do not push this to ten attempts; if a pool is still full after fifteen seconds it is having a real incident and you want to fail visibly rather than pile on.&lt;/p&gt;

&lt;p&gt;Surface content rejections immediately, with the provider's own wording. It is the one failure class a human can act on, and hiding it behind a spinner that retries three times means the user waits thirty seconds to find out their prompt was the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline these numbers came from
&lt;/h2&gt;

&lt;p&gt;The workload was a two-call pipeline: generate a still frame, then hand that frame to a video model as its first frame. Almost every hosted media API uses the same asynchronous shape, so the structure ports even though the field names will not.&lt;/p&gt;

&lt;p&gt;Runtime for everything below: Node 24.19 with the built-in fetch, no SDK. Measurements taken 2026-09-20 over the preceding 30 days.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;BASE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;MEDIA_API_BASE&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;// swap for your provider&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;HEAD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;Authorization&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;API_KEY&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="c1"&gt;// 1. still frame&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;img&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;BASE&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/images/generations`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-image-2.5-flare&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;a minimalist oak desk with a laptop showing a blue analytics dashboard, soft morning window light&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;aspect_ratio&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;16:9&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;resolution&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;1K&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;quality&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;low&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;frame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;poll&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;BASE&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/images/generations`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;img&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;taskId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// 2. animate it&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;vid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;BASE&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/video/generations`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;grok-imagine-video-1.5&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;images&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;          &lt;span class="c1"&gt;// array, even for a single reference&lt;/span&gt;
  &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;slow push-in on the laptop screen, dust motes in the light&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;resolution&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;720p&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You POST, get a task id, and poll until the state is terminal. A synchronous-looking wrapper around a 40-second median will time out somewhere in your stack, usually at a proxy you forgot about.&lt;/p&gt;

&lt;p&gt;That still frame cost $0.008 and took 37 seconds.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvf7qe1xy17xl8syzru2d.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvf7qe1xy17xl8syzru2d.jpg" alt="A minimalist oak desk with a laptop showing a blue analytics dashboard, a white ceramic mug and soft morning window light, generated at 1K low quality in a 16:9 frame" width="799" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The delivered size was 1672x941, not the 1920x1080 a 16:9 request implies. Image models snap dimensions to whatever grid their tiler uses; if a downstream step needs exact pixels, resize after generation rather than trusting the request.&lt;/p&gt;

&lt;p&gt;The economics push one way: iterate on the still image, not on the clip. A second of 720p video cost more than six times the entire still frame, so refining motion when the real problem is the composition in frame one is the expensive mistake.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things that cost me time
&lt;/h2&gt;

&lt;p&gt;The poll parameter was &lt;code&gt;task_id&lt;/code&gt; while the create response returned &lt;code&gt;taskId&lt;/code&gt;. A poll loop that returns an empty object for eight minutes rather than erroring is a genuinely miserable debugging session, and snake case versus camel case across create and read is common enough to check first.&lt;/p&gt;

&lt;p&gt;Result files expire. Seven days on this platform, and every hosted media API has some version of it. Download to your own storage in the job that created the file, not in a nightly batch that will eventually run against dead URLs.&lt;/p&gt;

&lt;p&gt;Billing lands on success. That is good for the bill and confusing for reconciliation, because a task that burned real GPU time upstream and then failed shows as free to you. Do not reconcile spend against request counts.&lt;/p&gt;

&lt;h2&gt;
  
  
  When not to do this at all
&lt;/h2&gt;

&lt;p&gt;Do not generate when you need the same output twice. Seeds help and do not survive model version changes, so if your product promises a user that what they saved last month still looks the same, generate once and store the file. Treat the model as a source of assets, not a renderer you can call again.&lt;/p&gt;

&lt;p&gt;Skip generation entirely for exact text inside an image. Every model in this class still garbles multi-word text at small sizes, and composing type in SVG over a generated background is more reliable than rolling dice on a headline.&lt;/p&gt;

&lt;p&gt;And if your failure volume is low enough that you would never notice a 3% difference, skip the classifier too. It earns its keep somewhere north of a few thousand calls a month; below that, a single retry and a clear error message is the correct amount of engineering.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would like to know
&lt;/h2&gt;

&lt;p&gt;If you run generation in production, how are you separating these two classes? I am specifically curious about providers that return a bare &lt;code&gt;500&lt;/code&gt; for both, because the regex table above is the part of our stack I would most like to delete.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>performance</category>
    </item>
    <item>
      <title>$30 to $8,880/mo in Six Months: How a PM Who Can't Code Runs an API Gateway Doing 300K Calls a Month</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Sat, 29 Aug 2026 04:10:49 +0000</pubDate>
      <link>https://dev.to/super_lewis/30-8880-in-six-months-how-a-pm-who-cant-code-runs-an-api-gateway-doing-300k-calls-a-month-adi</link>
      <guid>https://dev.to/super_lewis/30-8880-in-six-months-how-a-pm-who-cant-code-runs-an-api-gateway-doing-300k-calls-a-month-adi</guid>
      <description>&lt;p&gt;I'm a product manager. I can't write code — not "I'm rusty," not "I know some Python." I cannot write the code that runs my business.&lt;/p&gt;

&lt;p&gt;Six months ago my product made $30.32. This month it made $8,880. I have never paid for a click.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers first
&lt;/h2&gt;

&lt;p&gt;Revenue, straight out of the production database, admin accounts excluded:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Month&lt;/th&gt;
&lt;th&gt;Revenue&lt;/th&gt;
&lt;th&gt;Paying customers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-03&lt;/td&gt;
&lt;td&gt;$30.32&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-04&lt;/td&gt;
&lt;td&gt;$230&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-05&lt;/td&gt;
&lt;td&gt;$460&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-06&lt;/td&gt;
&lt;td&gt;$2,410&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-07&lt;/td&gt;
&lt;td&gt;$6,730&lt;/td&gt;
&lt;td&gt;63&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-08&lt;/td&gt;
&lt;td&gt;$8,880&lt;/td&gt;
&lt;td&gt;127&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;293x in six months. Trailing 30 days is $9,180 — so no, I haven't crossed $10k/month yet. September probably. I'd rather publish the real number than a rounder one.&lt;/p&gt;

&lt;p&gt;On the system side:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;August: 299,668 API calls at a 95.5% success rate&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Excluding content-moderation rejections — a customer's prompt refused upstream, not my system failing — it's &lt;strong&gt;98.9%&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;From July to August, &lt;strong&gt;volume grew 2.6x (116k → 300k calls) while the success rate went from 93.9% to 98.9%&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Volume nearly tripled and stability went &lt;em&gt;up&lt;/em&gt;. That last line says more than the success rate on its own.&lt;/p&gt;

&lt;p&gt;But a success rate can lie to you — in a minute I'll show you exactly how it lied to me. &lt;strong&gt;The number I actually trust: only 20 customers have ever topped up in more than one month, and they account for 70.8% of all revenue I've made.&lt;/strong&gt; Heavy users renewing month after month is the one stability proof that can't be faked. Nobody keeps paying for an API that isn't good.&lt;/p&gt;

&lt;p&gt;The product is &lt;a href="https://apimodels.app" rel="noopener noreferrer"&gt;apimodels.app&lt;/a&gt;: &lt;strong&gt;one API key for 182 AI models&lt;/strong&gt; — image, video, audio, and LLMs — behind a single OpenAI-compatible interface. Every line of it was written by an AI coding agent working under my direction.&lt;/p&gt;




&lt;h1&gt;
  
  
  How does someone who can't write code run this?
&lt;/h1&gt;

&lt;p&gt;It's the question I get most. The answer has four layers, and the first one matters more than the rest combined.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. This business has exactly two product cores
&lt;/h2&gt;

&lt;p&gt;Nobody picks a model aggregator because the UI is nice, the feature list is long, or the model catalog is big. They pick you for two reasons: &lt;strong&gt;the same model is cheaper here, and you don't fall over.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the whole list. Everything else is decoration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Get the core right and the product follows.&lt;/strong&gt; That sounds like a truism. The operational meaning is sharper: every hour you have should go into those two things, and you need to know, at all times, whether you're cheating on one of them.&lt;/p&gt;

&lt;p&gt;Because they will lie to each other. Here's the time I nearly got away with cheating on "cheaper."&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Official model channels only — a rule I paid for in customers
&lt;/h2&gt;

&lt;p&gt;Early on I wired up a &lt;strong&gt;very&lt;/strong&gt; cheap upstream for a popular video model. Cheap enough that the margin looked absurd.&lt;/p&gt;

&lt;p&gt;Then I noticed something no dashboard flagged: &lt;strong&gt;several users tried it once and never came back.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No errors. No failed calls. Every request returned 200.&lt;/p&gt;

&lt;p&gt;So I ran the same prompt through the official channel and put the two clips side by side. &lt;strong&gt;It was watered down.&lt;/strong&gt; Almost certainly a lower resolution upscaled and delivered as the higher one — or the cheap tier being passed off as the premium tier.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That upstream wasn't failing. It was substituting.&lt;/strong&gt; Every request "succeeded," my error rate was immaculate, and my retention was dying.&lt;/p&gt;

&lt;p&gt;That's when the rule got written, and it hasn't been broken since: &lt;strong&gt;official model channels only. Never a reseller priced far below the market.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If something is cheap in a way that seems impossible, it is guaranteed to be both unstable and quality-compromised.&lt;/strong&gt; There's no third explanation — somebody has to pay for the GPU, and if it isn't you, it's coming out of what you receive.&lt;/p&gt;

&lt;p&gt;So "cheap" as a product core needs restating. It isn't &lt;em&gt;cheapest possible&lt;/em&gt;. It's &lt;strong&gt;the lowest price at a quality level you have actually verified.&lt;/strong&gt; Undercut that line and you haven't won a customer — you've rented one for a single call.&lt;/p&gt;

&lt;p&gt;And notice which instrument caught it: &lt;strong&gt;not the error rate. Retention.&lt;/strong&gt; Silent quality degradation is invisible to every monitor you have. The only thing that reveals it is whether people come back.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Everything built in-house — no open-source gateway
&lt;/h2&gt;

&lt;p&gt;There are several good open-source API gateways. You can be running on one in a day. I used none of them.&lt;/p&gt;

&lt;p&gt;The reason is simple: &lt;strong&gt;price and stability both live in the same layer of code&lt;/strong&gt; — the forwarding logic between the customer's request and the upstream model. That layer is exactly what an off-the-shelf gateway occupies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When a failure lives inside a dependency, your move is to file an issue and wait. When it lives in your own code, you fix it that afternoon. At 300,000 calls a month, the gap between those two response times is the product.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Much slower to start. And the only way to actually control the two things customers are paying for.&lt;/p&gt;

&lt;p&gt;As for how the success rate got from 93.9% to 98.9%: &lt;strong&gt;every mechanism in this system exists because of a specific production failure.&lt;/strong&gt; Not one of them was designed up front. "Polishing," in plain language, means something breaks, you understand it, you turn it into a rule that stops it recurring, and then you do that three hundred times. There's no shortcut version — and it happens to be work a non-coder can do.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. What I own isn't the code. It's defining what counts as evidence.
&lt;/h2&gt;

&lt;p&gt;The most misunderstood part: this is &lt;strong&gt;not&lt;/strong&gt; "the AI writes it and I click deploy."&lt;/p&gt;

&lt;p&gt;If that's your loop, you'll ship something that appears to work and quietly does the wrong thing, and you'll hear about it from a customer.&lt;/p&gt;

&lt;p&gt;The agent generates. &lt;strong&gt;I define what counts as evidence that it works.&lt;/strong&gt; Four rules, each one written down because breaking it cost me money:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Never deploy automatically.&lt;/strong&gt; Test locally, show me the result, I say OK, &lt;em&gt;then&lt;/em&gt; commit. No exceptions, no "this one's trivial."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Test it" means test it.&lt;/strong&gt; Not test-and-then-fix-what-you-found. When I ask for a diagnosis I want the diagnosis — a fix applied before I understand the problem is a fix I can't evaluate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A deploy succeeded only if the code running on the server equals the version I pushed.&lt;/strong&gt; Not if the command printed no errors. I lost a full day to a deploy that reported success at every step while the server ran week-old code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verify against production data, never against a summary of production data.&lt;/strong&gt; "I checked and there were no errors" and "here is the query and here are the zero rows" are completely different claims.&lt;/p&gt;

&lt;p&gt;None of these require reading code. They require refusing to accept confidence as evidence — which turns out to be a product skill, not an engineering one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The bottleneck in AI-assisted development was never code generation. It's verification.&lt;/strong&gt; And verification is exactly the part a non-coding PM can own completely.&lt;/p&gt;




&lt;h1&gt;
  
  
  Then three things I got wrong
&lt;/h1&gt;

&lt;h2&gt;
  
  
  My free tier couldn't afford a single call
&lt;/h2&gt;

&lt;p&gt;New users got $0.10 in signup credit. The video models my content pushed hardest cost between $0.89 and $2.66 per successful call.&lt;/p&gt;

&lt;p&gt;A new user signs up, lands on the page my SEO worked hardest to rank, tries the thing that page is about, and gets rejected for insufficient balance. &lt;strong&gt;Their first experience of my product is a billing error.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The obvious fix is raising the credit. I almost did. Then I sorted every model by actual per-call spend instead of by how much I liked it, and two things stopped me cold:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;veo-3.1-fast-fhd&lt;/code&gt;: $0.0684 per clip&lt;/strong&gt; — 1080p, native audio. $0.10 covers it with change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;gpt-image-2-lite&lt;/code&gt;: $0.0078 per image.&lt;/strong&gt; The same $0.10 buys &lt;strong&gt;twelve pictures&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the free tier was never too small. It was perfectly adequate for a path no new user was ever shown. I'd built a funnel that took strangers by the hand and walked them into the most expensive thing I sell.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A free tier isn't a generosity number. It's the answer to one question: can a brand-new user complete one full happy path with it?&lt;/strong&gt; If not, you have two levers, and raising the credit is the expensive one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Never build a product page for a model you can't actually call
&lt;/h2&gt;

&lt;p&gt;To go after a new model's search terms, the instinct is to put the page up first.&lt;/p&gt;

&lt;p&gt;But that reads as a claim that you offer it. The API will 400 every time, and the specs and prices can only be invented. Worse — &lt;strong&gt;you've fed a false fact to the LLMs, and it gets cemented.&lt;/strong&gt; By the time you really launch, the version in their heads is still the one you made up.&lt;/p&gt;

&lt;p&gt;The right move is an honest placeholder: state plainly it isn't open yet, show how you verified that, promise it the day it launches, and put a &lt;strong&gt;working&lt;/strong&gt; neighboring model in the code sample. &lt;strong&gt;Honest pages rank fine, and they leave no debt.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In an era where AI remembers what you said, this only gets more important.&lt;/p&gt;

&lt;h2&gt;
  
  
  Healthy SEO metrics and a working product are independent
&lt;/h2&gt;

&lt;p&gt;Two of my flagship model pages failed on every single call because of a config error — &lt;strong&gt;for ten days, and nobody noticed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because the page status was active, it was indexed normally, and Search Console kept right on reporting impressions and clicks. Those metrics measure "the page exists and got crawled." None of them measures "the person who clicked could use it." &lt;strong&gt;If GSC is your only dashboard, everything looks great.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This and the watered-down upstream are the same disease: &lt;strong&gt;your dashboards will show everything is fine while the product rots.&lt;/strong&gt; The only cure is finding one metric that can't lie. For me that's retention.&lt;/p&gt;




&lt;h1&gt;
  
  
  One thing I still haven't gotten right
&lt;/h1&gt;

&lt;p&gt;I've never bought a click. That I can prove.&lt;/p&gt;

&lt;p&gt;The only acquisition work I've done is content: &lt;strong&gt;1,204 indexable pages across 7 languages, a page for every model.&lt;/strong&gt; I believe the traffic comes from search and LLM citations, because it's the only thing I've ever done.&lt;/p&gt;

&lt;p&gt;But I have to say this out loud: &lt;strong&gt;I can't prove it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is &lt;strong&gt;no analytics on my site&lt;/strong&gt;. No Google Analytics, no Plausible, no PostHog. My user table has a registration IP and a timestamp and &lt;strong&gt;no source field of any kind.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My largest customer paid $3,000 in August. &lt;strong&gt;I do not know how they found me.&lt;/strong&gt; Not which page, not which query, not which language. I know their invoice total and nothing else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Install analytics on day one, and add a first-touch source column to your user table on day one.&lt;/strong&gt; Not because you'll read the dashboard — you won't, for months. Because on the day you finally have revenue worth explaining, the data either exists or it doesn't. &lt;strong&gt;It cannot be backfilled. Ever.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It's the biggest mistake I've made, and it's the one that looks least urgent while you're making it.&lt;/p&gt;




&lt;h3&gt;
  
  
  About apimodels.app
&lt;/h3&gt;

&lt;p&gt;One API key for &lt;strong&gt;182 AI models&lt;/strong&gt; — image, video, audio, and LLMs — through an OpenAI-compatible interface, with no separate signups or quota applications per vendor.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Image generation&lt;/strong&gt; from &lt;strong&gt;$0.0078 per image&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1080p video with native audio&lt;/strong&gt; from &lt;strong&gt;$0.0684 per clip&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;299,668 calls served in August 2026 at a 95.5% success rate&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Official model channels only&lt;/strong&gt; — no resellers priced far below market&lt;/li&gt;
&lt;li&gt;Interface and docs in English, Chinese, Japanese, Korean, Spanish, Portuguese, and Russian&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every figure here is queried straight from production. Happy to answer anything in the comments.&lt;/p&gt;

</description>
      <category>startup</category>
      <category>webdev</category>
      <category>ai</category>
      <category>apigateway</category>
    </item>
    <item>
      <title>The Seedance 2.5 Prompting Guide, in English</title>
      <dc:creator>Super Lewis</dc:creator>
      <pubDate>Wed, 05 Aug 2026 15:16:24 +0000</pubDate>
      <link>https://dev.to/super_lewis/the-seedance-25-prompting-guide-in-english-4hen</link>
      <guid>https://dev.to/super_lewis/the-seedance-25-prompting-guide-in-english-4hen</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Cross-posted&lt;/strong&gt; from &lt;a href="https://apimodels.app/blog/seedance-2-5-prompting-guide" rel="noopener noreferrer"&gt;apimodels.app&lt;/a&gt; — the canonical version lives there and is kept up to date.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you have only ever prompted image models, video models feel like a different instrument. You are no longer describing a frame; you are directing time, camera and sound at once, and the failure modes are unfamiliar — subjects that duplicate halfway through, props that change hands, edits that quietly re-roll the whole shot.&lt;/p&gt;

&lt;p&gt;ByteDance released &lt;strong&gt;Seedance 2.5&lt;/strong&gt; on 31 July 2026 and, unusually, shipped a real prompting manual with it, written by the people who trained the model. It is specific in the way vendor docs rarely are: how many reference assets it accepts, how to address each one, how to stage a 30-second clip so it does not collapse, how to change one object in existing footage without disturbing anything else, and a blunt list of nine things it will not do.&lt;/p&gt;

&lt;p&gt;This is that manual walked through in English with &lt;strong&gt;every template kept&lt;/strong&gt;. The limits, syntax and stated boundaries are ByteDance's; the worked examples are mine, written to be copy-pasted rather than translated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contents&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1. Foundations&lt;/li&gt;
&lt;li&gt;2. The techniques that matter&lt;/li&gt;
&lt;li&gt;3. Advanced usage&lt;/li&gt;
&lt;li&gt;4. Checklist before you submit&lt;/li&gt;
&lt;li&gt;5. What it will not do&lt;/li&gt;
&lt;li&gt;Running these prompts today&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  1. Foundations
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1.1 The base formula
&lt;/h3&gt;

&lt;p&gt;A prompt is assembled from six parts. Only the first two are required; drop any of the rest that you do not need.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FORMULA

subject + action or event
  + scene and environment    (optional)
  + visual style             (optional)
  + camera work or cutting   (optional)
  + sound                    (optional)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What each part is responsible for:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Subject + action or event&lt;/strong&gt; — who or what is doing what. This is the floor of the whole prompt. Summarise the main process first; add concrete detail only for the key action; never describe the same action twice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scene and environment&lt;/strong&gt; — place, time, weather, spatial relationships, state of the background.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Visual style&lt;/strong&gt; — light, colour, material, image texture, overall mood.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Camera work or cutting&lt;/strong&gt; — shot size, camera position, movement, focus target, how shots join.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sound&lt;/strong&gt; — dialogue, timbre, ambience, effects, music.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE

&amp;lt;SUBJECT&amp;gt; &amp;lt;main action or event&amp;gt; in &amp;lt;scene and environment&amp;gt;.
The image is &amp;lt;visual style&amp;gt;.
The camera uses &amp;lt;shot size, position, movement or cutting&amp;gt;.
Sound includes &amp;lt;dialogue, ambience, effects or music&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

A luthier finishes a violin in a narrow workshop, lifting the body off
the bench and setting it upright in the drying rack.

Late afternoon light rakes through fine sawdust; fresh varnish shows a
deep amber, the bench is orderly, tools laid out in a worn leather roll.

The camera opens on a medium shot of the hands releasing the clamp, then
pushes slowly in on the grain of the top plate, then cuts to the rack
seen straight on.

Keep the scrape of the plane, the click of the clamp, and quiet room tone.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Generation parameters do not belong in the prompt. Resolution, duration and aspect ratio are set on the generation page or through the API.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.2 With reference assets: prepare them, then assign roles
&lt;/h3&gt;

&lt;p&gt;You can combine up to &lt;strong&gt;50 reference assets&lt;/strong&gt; in one job. Each type has a hard input range and a narrower recommended range; the recommended range exists to raise stability and is not a statement about capability.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LIMITS

Images    input range   up to 30, each no larger than 4K
          recommended   1-8 subjects

Video     input range   up to 10 clips, 30s total across all clips
          recommended   1-5 subjects, 5-10s per clip

Audio     input range   up to 10 clips, 30s total across all clips
          recommended   keep only dialogue, timbre, ambience or music
                        that the task actually needs

Editing   input range   source video plus reference images together
          recommended   source under 20s, 1-5 reference images
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You may go beyond the recommended range: image subjects can extend to 9-12, audio/video subjects to 6-10, editing reference images to 6-8. The guide is blunt about the trade-off — the more assets, the more easily stability drops.&lt;/p&gt;

&lt;p&gt;When a subject needs more than five image references and still needs multiple viewpoints, &lt;strong&gt;split the viewpoints across separate images&lt;/strong&gt;. Several independent view images are usually more stable than several views packed into one grid.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assign a role to every asset.&lt;/strong&gt; After uploading, state what each asset supplies. When a person, background or composition inside an asset is likely to bleed into the result, also state what not to take from it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The mapping between assets and subjects must be written into the prompt.&lt;/strong&gt; Do not rely on text labels burned into the image, and do not leave the model to work out which asset corresponds to which person, prop or scene.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — asset roles

@image1 supplies &amp;lt;subject&amp;gt;'s &amp;lt;appearance, clothing, structure or material&amp;gt;.
@video1 supplies &amp;lt;action, camera movement or rhythm&amp;gt;.
@audio1 supplies &amp;lt;character or sound type&amp;gt;'s &amp;lt;timbre, dialogue,
        ambience or music&amp;gt;.

&amp;lt;SUBJECT&amp;gt; completes &amp;lt;main action or event&amp;gt; in &amp;lt;scene&amp;gt;.
The image is &amp;lt;visual style&amp;gt;; the camera uses &amp;lt;camera expression&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@image1 supplies &amp;lt;BAKER&amp;gt;'s features, hair and flour-dusted grey apron.
        Do not take the background.
@image2 supplies &amp;lt;BAKERY&amp;gt;'s counter layout, tiled wall and window
        position. Do not take the people in the frame.
@video1 supplies the rhythm of shaping the dough, lifting it and sliding
        the tray in. Do not take the person's identity, clothing or room.

&amp;lt;BAKER&amp;gt; shapes a sourdough loaf in &amp;lt;BAKERY&amp;gt; and slides it into the deck
oven.

The camera records the shaping in a medium shot, then pushes slowly in on
the crust; keep the scrape of the peel, the oven door, and room tone.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When several images are different viewpoints of the same person or product, say so explicitly, and state how many of the object should exist in the result.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE — multiple viewpoints

@image1 defines the front of the same folding lamp.
@image2 defines its left-side structure.
@image3 defines its right-side structure.
@image4 defines its back structure.

The four images together define one single folding lamp; there is only
ever one folding lamp in the finished video.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When a reference video already carries the action, camera work and ordering accurately, the prompt only needs to say &lt;strong&gt;what it inherits&lt;/strong&gt; — do not restate the choreography move by move. Restating it invites a conflict with the asset itself.&lt;/p&gt;

&lt;p&gt;A white-box video mainly supplies motion and a spatial skeleton. You still have to write out the subject, scene, action and style you want generated.&lt;/p&gt;

&lt;h3&gt;
  
  
  1.3 Special characters for sound and text
&lt;/h3&gt;

&lt;p&gt;Natural language works on its own. When you need to separate music, effects, dialogue and on-screen text, four bracket types disambiguate them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SYNTAX

Music       ( )     (calm piano plays underneath)
Sound fx    &amp;lt; &amp;gt;     &amp;lt;a bell rings in the distance&amp;gt;
Dialogue    { }     {Hello, welcome back}
Subtitle    【 】    【Chapter One: Departure】
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To control subtitles or audio, simply write which sound categories to keep and which not to use. Picture content should still be written as positive description — only asset roles, edit scope, and things likely to bleed in need explicit prohibition.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

No background music; keep only dialogue, ambience and action effects.
No subtitles.
No audio at all.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Reinforcing the dialogue language.&lt;/strong&gt; When a line is not in Chinese, name the language before it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The girl says softly in Japanese: {もう大丈夫です}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the line text is English but the model speaks it in Chinese, or when you need to control a regional variety, be more specific. The recommended formula:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FORMULA

dialogue language + regional variant or accent + delivery
  + speaker + {line}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

Dialogue language: American English. The girl says it naturally and
colloquially: {I thought you weren't coming.}

Dialogue language: native Los Angeles American English. A young man says
it in natural LA speech: {No way, you actually made it.}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  2. The techniques that matter
&lt;/h2&gt;

&lt;h3&gt;
  
  
  2.1 Multi-asset work: tell the model which asset each scene uses
&lt;/h3&gt;

&lt;p&gt;Images, video and audio can be combined. Once there are many assets, the job of the prompt is &lt;strong&gt;not&lt;/strong&gt; to cram them all into one sentence — it is to make the correspondence between people, props, scenes, actions and sound unambiguous. Organise it in this order:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ORDER

per-asset roles → subject mapping → grouping by type
  → subject specification → per-scene invocation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 01 — name each subject individually.&lt;/strong&gt; Different people, products and props each get their own binding.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&amp;lt;PERSON A&amp;gt; corresponds to @image1; take appearance, hair and clothing only.
&amp;lt;PERSON B&amp;gt; corresponds to @image2; take appearance, hair and clothing only.
&amp;lt;PROP A&amp;gt;   corresponds to @image3; take structure, material and colour only.
&amp;lt;SCENE A&amp;gt;  references @image4; take spatial layout, architecture and light
           only, not the people in the frame.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do &lt;strong&gt;not&lt;/strong&gt; write "@image1 through @image4 define four characters". That never says which image is which character.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 02 — group the assets by type.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[PEOPLE]
&amp;lt;CONSERVATOR&amp;gt; corresponds to @image1; appearance, hair, clothing only.
&amp;lt;REGISTRAR&amp;gt;   corresponds to @image2; appearance, hair, clothing only.
&amp;lt;INSTALLER&amp;gt;   corresponds to @image3; appearance, hair, clothing only.
&amp;lt;DOCENT&amp;gt;      corresponds to @image4; appearance, hair, clothing only.
The four do not exchange appearance, clothing, actions, positions or lines.

[PROPS]
&amp;lt;SAMPLE BOX&amp;gt;   corresponds to @image5; belongs to &amp;lt;CONSERVATOR&amp;gt; only.
&amp;lt;RECORD BOARD&amp;gt; corresponds to @image6; belongs to &amp;lt;REGISTRAR&amp;gt; only.

[SCENES]
&amp;lt;LAB&amp;gt;     references @image7; space, materials and light only.
&amp;lt;GALLERY&amp;gt; references @image8; space, materials and light only.

[ACTION AND SOUND]
@video1 supplies &amp;lt;CONSERVATOR&amp;gt; opening the sample box; do not take the
        person or the room from it.
@audio1 supplies &amp;lt;DOCENT&amp;gt;'s timbre and the specified line.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 03 — write a specification for an important subject.&lt;/strong&gt; When one person is used across scenes with several assets, collect their definition in one block.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[SUBJECT SPEC: CONSERVATOR]
Appearance and clothing: @image1.
Fixed prop:              &amp;lt;SAMPLE BOX&amp;gt; from @image5.
Appears in:              &amp;lt;LAB&amp;gt; and &amp;lt;GALLERY&amp;gt;.
Action references:       the box-opening in @video1, the sample-placing
                         in @video2.
Do not:                  wear another character's clothing; hold the
                         &amp;lt;RECORD BOARD&amp;gt; or any presentation equipment.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 04 — invoke assets per scene.&lt;/strong&gt; State what the scene uses, what happens, and the state at the end.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SCENE ONE | inspection in the lab
Uses:   &amp;lt;CONSERVATOR&amp;gt;, &amp;lt;SAMPLE BOX&amp;gt;, &amp;lt;LAB&amp;gt;, and the box-opening from @video1.
Event:  &amp;lt;CONSERVATOR&amp;gt; opens the sample box at the bench and inspects the
        sample inside.
Ends:   &amp;lt;CONSERVATOR&amp;gt; stands at the inner side of the bench; the sample box
        stays beside their own right hand, i.e. on the left of frame.

SCENE TWO | registration in the gallery
Uses:   &amp;lt;REGISTRAR&amp;gt;, &amp;lt;RECORD BOARD&amp;gt;, &amp;lt;GALLERY&amp;gt;.
Event:  &amp;lt;REGISTRAR&amp;gt; checks the numbers on the record board beside a case.
Ends:   the record board is still held in both hands by &amp;lt;REGISTRAR&amp;gt;; no
        other character enters the case area.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal of multi-asset work is for the model to pick the &lt;strong&gt;right&lt;/strong&gt; asset in the current scene. It is not a requirement that every asset appears on screen at once.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.2 Thirty-second video: organise events by stage and end state
&lt;/h3&gt;

&lt;p&gt;When there are many events, break the story into consecutive stages. Give each stage exactly one main state change, and write the state that is &lt;strong&gt;directly visible on screen&lt;/strong&gt; when that stage ends.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — long video

[GOAL]
Generate a &amp;lt;video type&amp;gt;. The core subject is &amp;lt;subject&amp;gt;; the main event is
&amp;lt;story summary&amp;gt;.

[STAGE ONE]
Opens with:   &amp;lt;initial state of people, props and scene&amp;gt;.
Main event:   &amp;lt;one main action or event&amp;gt;.
Ends with:    &amp;lt;person position, prop ownership or picture state&amp;gt;.

[STAGE TWO]
Carried over: &amp;lt;state that must be preserved&amp;gt;.
Main event:   &amp;lt;one main action or event&amp;gt;.
Ends with:    &amp;lt;observable state&amp;gt;.

[STAGE THREE]
Main event:   &amp;lt;closing event&amp;gt;.
Ends with:    &amp;lt;final picture state&amp;gt;.

[KEEP CONSISTENT]
Keep &amp;lt;identity, headcount, clothing, prop ownership, spatial orientation
and sound relationships&amp;gt; stable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE — florist order-packing walkthrough

[GOAL]
A walkthrough of a florist packing an order. &amp;lt;FLORIST&amp;gt; and &amp;lt;ASSISTANT&amp;gt;
together trim, wrap and hand off a bouquet.

[STAGE ONE]
Opens with:   &amp;lt;FLORIST&amp;gt; behind the bench; loose stems, shears and wrapping
              paper on the surface.
Main event:   &amp;lt;FLORIST&amp;gt; arranges the stems and trims them to length.
Ends with:    the bouquet is held in &amp;lt;FLORIST&amp;gt;'s left hand; the shears are
              back on the right side of the bench.

[STAGE TWO]
Carried over: both keep the same identity and clothing; the bouquet is
              still held by &amp;lt;FLORIST&amp;gt;.
Main event:   &amp;lt;ASSISTANT&amp;gt; unrolls the paper; &amp;lt;FLORIST&amp;gt; lays the bouquet in
              and ties a green ribbon.
Ends with:    the wrapped bouquet lies flat in the centre of the bench, the
              ribbon knot facing the camera.

[STAGE THREE]
Main event:   &amp;lt;ASSISTANT&amp;gt; lifts the bouquet onto the pickup shelf.
Ends with:    the bouquet is alone in the centre of the pickup shelf; both
              stand behind the bench looking at the finished order.

[KEEP CONSISTENT]
Keep &amp;lt;FLORIST&amp;gt; and &amp;lt;ASSISTANT&amp;gt;'s identities, clothing, the orientation of
the bench, the position of the shears and ownership of the bouquet stable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Timestamps and pacing.&lt;/strong&gt; For ordinary narrative, prefer stages. Reach for time expressions — in whole seconds — only when a handover, an entrance or exit, a transition or a specific beat needs control. Time &lt;strong&gt;ranges&lt;/strong&gt; allocate story pacing; a time &lt;strong&gt;point&lt;/strong&gt; pins one key event; &lt;strong&gt;relative&lt;/strong&gt; time describes waiting between events.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

0-5s:   an empty wooden display stand; a hand sets down a white ceramic
        plate; at the end the hand has left and only the white plate is
        on the stand.
5-10s:  the white plate is removed and a clear glass is set down; at the
        end only the clear glass is on the stand.
10-15s: the clear glass is removed and a green ceramic bottle is set down;
        at the end only the green bottle is on the stand.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Time range        0-3s … 3-7s … 7-12s …
Explicit point    At 5s the camera whips left and completes the transition.
Relative time     Three seconds after the button is pressed, the room
                  lights fade out.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Segments must be continuous and must not overlap. A segment is a &lt;strong&gt;time budget for an event, not a precise cut point&lt;/strong&gt; — an action may land slightly before or after a boundary. Too little content in a segment widens the model's room to invent; too much causes over-cutting or dropped story beats. Do not use timestamps to demand a frequency such as three actions in one second.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.3 Editing, first/last frame and extension: parameters that lock themselves
&lt;/h3&gt;

&lt;p&gt;Video editing, first-frame or first-and-last-frame generation, and video extension automatically lock some generation parameters based on the input.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AUTO LOCK

Video editing
  aspect ratio   automatically keeps the input video's ratio; not settable
  duration       automatically stays essentially the same as the input;
                 not settable. Frame handling can produce a difference of
                 up to about 0.3 seconds.

First / first+last frame
  aspect ratio   automatically uses the first-frame image's ratio. First
                 and last frame should use the same ratio, otherwise the
                 last frame can be stretched.
  duration       settable

Video extension
  aspect ratio   automatically keeps the input video's ratio; not settable
  duration       settable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Locked parameters cannot be overridden on the generation page or through the API. Everything else follows the options currently available.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.4 Video editing: declare the master, the scope, and what stays
&lt;/h3&gt;

&lt;p&gt;When editing existing footage, first define the source as the &lt;strong&gt;sole master&lt;/strong&gt;, then state the edit target, the scope it applies to, the target asset, and what must be preserved.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — general edit

[EDIT GOAL]
Edit @video1; across &amp;lt;the whole clip or a specific time range&amp;gt;, perform
&amp;lt;add / remove / replace / adjust&amp;gt; on &amp;lt;picture object, region or sound
category&amp;gt;.

[SOURCE ROLE]
@video1 is the sole editing master, responsible for &amp;lt;people, scene,
action, composition, camera, occlusion relationships, sound and event
order&amp;gt;.

[TARGET ASSET ROLE]
@image1 or @audio1 supplies &amp;lt;the specified attributes of the target object
or sound&amp;gt;.

[EDIT SCOPE]
Process only &amp;lt;object, region, time range or sound category&amp;gt;.

[PRESERVE]
Keep &amp;lt;the picture, action, sound and timing relationships that must not
change&amp;gt; as in @video1.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

[EDIT GOAL]
Edit @video1; between 4s and 7s only, change the cold blue light on the
right-hand wall to warm amber.

[SOURCE ROLE]
@video1 is the sole editing master, responsible for the people, the room
layout, the action, composition, camera movement, sound and event order.

[EDIT SCOPE]
Adjust only the colour of the light on the right-hand wall and the area it
illuminates; skin tone may change naturally with the ambient light.

[PRESERVE]
Identity, clothing, expression, position, movement, room structure, camera
movement, dialogue and room tone stay as in @video1.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Case A — replacing a subject.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE

[EDIT GOAL]
Edit @video1; change &amp;lt;original object&amp;gt; to &amp;lt;target object&amp;gt; only.

[SOURCE ROLE]
@video1 is the sole editing master, responsible for the original scene,
camera position, camera movement, motion paths, occlusion relationships
and event order.

[TARGET ASSET ROLE]
@image1 supplies &amp;lt;target object&amp;gt;'s &amp;lt;appearance, structure or material&amp;gt;;
do not take &amp;lt;unrelated background, people or composition&amp;gt;.

[EDIT OBJECT AND SCOPE]
Modify only &amp;lt;the explicit object and region&amp;gt;. The count of the target
object across the whole clip is &amp;lt;number&amp;gt;. Do not modify &amp;lt;what must stay&amp;gt;.

[TIMELINE INHERITANCE]
&amp;lt;Target object&amp;gt; inherits every appearance, movement, occlusion and exit of
&amp;lt;original object&amp;gt;: the same timings, durations, paths and changes of speed.
Apart from the object or region explicitly modified above, all other
people, props, scene content, camera movement, cuts and event order in
@video1 stay as they are.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

[EDIT GOAL]
Edit @video1; replace the yellow folding lamp with the white folding lamp
from @image1 only.

[SOURCE ROLE]
@video1 is the sole editing master, responsible for the desk, the books,
the hand movements, camera position, camera movement, occlusion
relationships and event order.

[TARGET ASSET ROLE]
@image1 is used only for the white folding lamp's appearance, structure
and material; do not take the background, composition or other objects in
the image.

[EDIT OBJECT AND SCOPE]
There is exactly one white folding lamp throughout. Replace only the
original yellow folding lamp; do not modify the books, desk, hands or
background.

[TIMELINE INHERITANCE]
The white folding lamp inherits every appearance, swing of the arm,
occlusion by the hand and exit from frame of the original yellow lamp:
the same timings, paths and changes of speed. Apart from the object
explicitly modified above, all other people, props, scene content, camera
movement, cuts and event order in @video1 stay as they are.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Case B — replacing a background.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE

[EDIT GOAL]
Edit @video1; replace &amp;lt;the original background region&amp;gt; with &amp;lt;the target
environment&amp;gt; from @image1 only.

[SOURCE ROLE]
@video1 is the sole editing master, responsible for the people, foreground
objects, action, composition, camera movement and event order.

[TARGET ASSET ROLE]
@image1 is used only for &amp;lt;target environment&amp;gt;'s spatial layout, materials,
depth of field, ambient colour and light direction; do not take the people
or foreground objects in the image.

[EDIT OBJECT AND SCOPE]
Modify only &amp;lt;the background region outside the subject's outline&amp;gt;. Do not
modify &amp;lt;the subject's identity, features, hair, clothing, expression,
position, size or movement&amp;gt;.

[TIMELINE INHERITANCE]
The subject's movement and occlusion relationships stay as in @video1.
Apart from the object or region explicitly modified above, all other
people, props, scene content, camera movement, cuts and event order in
@video1 stay as they are.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@video1 is the sole editing master, responsible for the person, the action,
composition, camera and event order.
@image1 supplies only the spatial layout, depth, ambient colour and light
direction of a daytime glass greenhouse; do not take the people in it.

Replace only the pale grey background outside the person's outline in
@video1 with the daytime glass greenhouse from @image1.

The person's identity, features, hair, clothing, expression, position,
size and raised-hand movement stay as in @video1. Apart from the region
explicitly modified above, all other people, props, scene content, camera
movement, cuts and event order in @video1 stay as they are.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Editing sound.&lt;/strong&gt; Dialogue, language, timbre, background music and effects can each be handled separately. State the speaker or sound category, the intended change, and whether the other sounds are preserved.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

Edit @video1; remove the original background music only, keeping the
dialogue, lip sync, ambience and action effects. Picture content, camera
and cutting rhythm stay as in @video1.

Edit @video1; change &amp;lt;DOCENT&amp;gt;'s dialogue language to natural American
English, keeping the line content and the timing of speech unchanged.
Other characters' voices, background music, ambience and picture stay as
in @video1.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2.5 Extension: align the boundary frame first, then describe what comes before or after
&lt;/h3&gt;

&lt;p&gt;Extension generates new material outside the boundary of an existing clip. The output ratio automatically follows the input; the added duration is set by you.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BOUNDARY

          extend backward  ←  [ original video ]  →  extend forward
                              first frame  last frame
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The boundary frame must be continuous. Extending &lt;strong&gt;forward&lt;/strong&gt;, the first frame of the new segment has to continue from the source's last frame. Extending &lt;strong&gt;backward&lt;/strong&gt;, the last frame of the new segment has to join the source's first frame. Beyond the boundary frame itself, people, props, background, motion trend and sound must also stay continuous.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Extending forward.&lt;/strong&gt; Describe the continuous state of the source's last frame first, then what happens after it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — forward, basic

@video1 is the source video to extend forward.

Extend @video1 forward. The first frame of the extension continues
directly from the last frame of @video1: keep &amp;lt;subject pose and facing&amp;gt;,
&amp;lt;prop positions&amp;gt;, &amp;lt;background and spatial relationships&amp;gt;, &amp;lt;camera position
and composition&amp;gt;, &amp;lt;light&amp;gt;, &amp;lt;sound state&amp;gt; and &amp;lt;motion trend&amp;gt; continuous.

Then, &amp;lt;describe the new action, event, camera move or sound&amp;gt;.

Throughout the extension keep &amp;lt;identity and clothing&amp;gt;, &amp;lt;key props&amp;gt;,
&amp;lt;background layout&amp;gt;, &amp;lt;the camera axis&amp;gt; and &amp;lt;the original sound
environment&amp;gt; continuous. The same subject remains one continuous object —
no duplication, no splitting; the subject's form and the number of parts
stay stable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@video1 is the source video to extend forward.

Extend @video1 forward. The first frame of the extension continues
directly from the last frame of @video1: keep the same locked-off medium
shot, the position and heading of the orange paper plane, the classroom
window background, the afternoon light and the drift toward the right of
frame continuous.

Then let the orange paper plane glide further right and leave frame, while
the white curtain at the window sways slightly. Camera and classroom
background stay as at the source's last frame.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — forward, with other reference assets

@image1 defines &amp;lt;PERSON A&amp;gt;'s facial features.
@image2 defines &amp;lt;PERSON A&amp;gt;'s clothing.
@image3 defines &amp;lt;key prop&amp;gt;'s structure and material.
@video1 is the source video to extend forward.

Extend @video1 forward. The first frame of the extension continues
directly from the last frame of @video1: keep &amp;lt;the boundary picture and
sound state&amp;gt; continuous.

Then, &amp;lt;the new action or event PERSON A performs with the key prop&amp;gt;.

Throughout the extension keep &amp;lt;identity and clothing&amp;gt;, &amp;lt;key props&amp;gt;,
&amp;lt;background layout&amp;gt;, &amp;lt;the camera axis&amp;gt; and &amp;lt;the original sound
environment&amp;gt; continuous. The same subject remains one continuous object —
no duplication, no splitting; the subject's form and the number of parts
stay stable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@image1 defines &amp;lt;GARDENER&amp;gt;'s facial features.
@image2 defines &amp;lt;GARDENER&amp;gt;'s pale green work apron.
@image3 defines &amp;lt;WICKER BASKET&amp;gt;'s structure and material.
@video1 is the source video to extend forward.

Extend @video1 forward. The first frame of the extension continues
directly from the last frame of @video1: keep the greenhouse bench,
&amp;lt;GARDENER&amp;gt;'s standing position and the position of &amp;lt;WICKER BASKET&amp;gt;
continuous.

Then &amp;lt;GARDENER&amp;gt; lifts &amp;lt;WICKER BASKET&amp;gt; with both hands and sets it on the
middle shelf of the wooden rack behind them.

Throughout, keep &amp;lt;GARDENER&amp;gt;'s face, apron, the greenhouse layout and the
camera direction continuous.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Other assets may supply people, props or sound, but they do &lt;strong&gt;not&lt;/strong&gt; replace the source's last frame as the control on the opening picture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Extending backward.&lt;/strong&gt; Describe what happens before the source starts first, then write the source's first frame as an explicit end state. Writing only "then it joins the original video" tends to make the model introduce later characters or effects early, or keep changing the picture after it has already reached the target state.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — backward, basic

@video1 is the source video to extend backward.

Extend @video1 backward. Before the source begins, &amp;lt;describe the preceding
action, event, camera move or sound&amp;gt;.

The last frame of the extension joins the first frame of @video1 naturally:
&amp;lt;subject pose and facing&amp;gt;, &amp;lt;prop positions&amp;gt;, &amp;lt;background and spatial
relationships&amp;gt;; keep &amp;lt;camera position and composition&amp;gt;, &amp;lt;light&amp;gt;, &amp;lt;sound
state&amp;gt; and &amp;lt;motion trend&amp;gt; consistent with @video1's first frame.

Throughout the extension keep &amp;lt;identity and clothing&amp;gt;, &amp;lt;key props&amp;gt;,
&amp;lt;background layout&amp;gt;, &amp;lt;the camera axis&amp;gt; and &amp;lt;the original sound
environment&amp;gt; continuous. The same subject remains one continuous object —
no duplication, no splitting; the subject's form and the number of parts
stay stable.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@video1 is the source video to extend backward.

Extend @video1 backward. Before the source begins, show the same glass
greenhouse empty: morning mist lying low and slowly dispersing, the shade
blind at the top gradually rising, no people in frame.

The last frame of the extension joins the first frame of @video1 naturally:
keep the central aisle, the planting benches on both sides, the glass
frame, the soft morning light and the locked-off wide composition
consistent with @video1's first frame; at the end the blind is fully
raised, no one is in the aisle, and the leaves are still moving faintly.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — backward, with other reference assets

@image1 defines &amp;lt;PERSON A&amp;gt;'s facial features.
@image2 defines &amp;lt;PERSON A&amp;gt;'s clothing.
@image3 defines &amp;lt;key prop&amp;gt;'s structure and material.
@video1 is the source video to extend backward.

Extend @video1 backward. Before the source begins, &amp;lt;PERSON A completes the
preceding action or event&amp;gt;.

The last frame of the extension joins the first frame of @video1 naturally:
&amp;lt;PERSON A's pose and facing&amp;gt;, &amp;lt;the position and state of the key prop&amp;gt;,
&amp;lt;where the other characters stand&amp;gt;; keep &amp;lt;background and spatial
relationships&amp;gt;, &amp;lt;camera position and composition&amp;gt;, &amp;lt;light&amp;gt;, &amp;lt;sound state&amp;gt;
and &amp;lt;motion trend&amp;gt; consistent with @video1's first frame.

Throughout the extension keep &amp;lt;identity and clothing&amp;gt;, &amp;lt;key props&amp;gt;,
&amp;lt;background layout&amp;gt;, &amp;lt;the camera axis&amp;gt; and &amp;lt;the original sound
environment&amp;gt; continuous. The same subject remains one continuous object —
no duplication, no splitting; the subject's form and the number of parts
stay stable.
&amp;lt;Assets that appear only after the source begins&amp;gt; must not appear early in
the backward extension.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@image1 defines &amp;lt;CURATOR&amp;gt;'s facial features.
@image2 defines &amp;lt;CURATOR&amp;gt;'s dark blue work coat.
@image3 defines &amp;lt;WOODEN DISPLAY CASE&amp;gt;'s structure and material.
@image4 defines the grey work clothes of the two &amp;lt;INSTALLERS&amp;gt;.
@image5 defines the space and light of &amp;lt;PREP ROOM&amp;gt;.
@video1 is the source video to extend backward.

Extend @video1 backward. Before the source begins, &amp;lt;CURATOR&amp;gt; walks to the
bench, picks up the closed &amp;lt;WOODEN DISPLAY CASE&amp;gt; and opens the lid.

The last frame of the extension joins the first frame of @video1 naturally:
&amp;lt;CURATOR&amp;gt; stands in the centre of frame holding the opened &amp;lt;WOODEN DISPLAY
CASE&amp;gt; in both hands; the two &amp;lt;INSTALLERS&amp;gt; stand behind them, one to each
side. Keep the vertical front-on medium shot, the bench position, the prep
room background and the morning light from the left consistent with
@video1's first frame.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The boundary should look naturally joined; it should not be read as pixel-identical. The loudness of the extension may vary slightly from the source. Extending a video generated by the same model generation usually joins more smoothly in both picture and sound. When reviewing, check the picture, the sound &lt;strong&gt;and&lt;/strong&gt; the whole extended segment on both sides of the boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Advanced usage
&lt;/h2&gt;

&lt;h3&gt;
  
  
  3.1 Keyframes, storyboards and white-box references
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;First and last frame together with reference assets.&lt;/strong&gt; In multimodal reference mode you can declare @image1 as the first frame and &lt;a class="mentioned-user" href="https://dev.to/image2"&gt;@image2&lt;/a&gt; as the last frame in the opening line of the prompt — there is no need to switch to a separate first/last-frame mode. The system locks the output ratio to the first frame's ratio; duration is set on the generation page or through the API. First and last frame should use the same ratio, or the last frame may be stretched. Remaining reference images can still define people, props, scenes and materials separately.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE

@image1 is the first frame, defining the composition, subject position,
pose, prop state, scene and camera direction at the start.
@image2 is the last frame, defining the composition, subject position,
pose, prop state, scene and camera direction at the end.
@image3 supplies &amp;lt;SUBJECT A&amp;gt;'s &amp;lt;appearance, clothing, structure or
material&amp;gt;; it does not change the first-frame composition defined by
@image1, nor the last-frame composition defined by @image2.
@image4 supplies &amp;lt;SUBJECT B, prop or scene&amp;gt;'s &amp;lt;specified attributes&amp;gt;; it
does not change the first-frame composition defined by @image1, nor the
last-frame composition defined by @image2.

&amp;lt;Describe one continuous action or event.&amp;gt;
The picture begins naturally from the first frame defined by @image1 and,
through continuous action, arrives at the last frame defined by @image2.
Between the two, keep &amp;lt;identity, prop structure and ownership, scene
layout and camera direction&amp;gt; continuous.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@image1 is the first frame, defining the composition of the perfume
workshop at the start: the person's position, pose, the state of the props
on the bench and the camera direction.
@image2 is the last frame, defining the same at the end.
@image3 supplies &amp;lt;PERFUMER&amp;gt;'s face, hair and dark green apron; it does not
change the compositions defined by @image1 and @image2.
@image4 supplies &amp;lt;GLASS BOTTLE&amp;gt;'s shape, material and label position; it
does not change the compositions defined by @image1 and @image2.

Starting from the first-frame pose, &amp;lt;PERFUMER&amp;gt; picks up a pipette and
&amp;lt;GLASS BOTTLE&amp;gt;, drips amber concentrate into it, swirls it gently, seats
the stopper, and places the finished bottle in the centre of the bench,
arriving naturally at the last frame defined by @image2.

Between the two, keep &amp;lt;PERFUMER&amp;gt;'s identity and clothing, the number and
structure of the bottles, the layout of the wooden bench, the warm side
light and the camera direction continuous.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each anchor image needs its own sentence — do not merge them into "images 1 and 2 as the first and last frames". First and last frame should use the same ratio. Other reference images supply specified attributes only; they do not override the composition of the anchors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multiple keyframes in order.&lt;/strong&gt; When several independent images define different stages of a process, open with the ordering statement, then describe the key state each image represents.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE

Use @image1 through @imageN as keyframes, in that order.

@image1 is the first frame, defining &amp;lt;the composition, subject position,
pose, prop state and camera direction at the start&amp;gt;.
@image2 defines the second keyframe: &amp;lt;the visible state at the end of
stage one&amp;gt;.
@image3 defines the third keyframe: &amp;lt;the visible state at the end of
stage two&amp;gt;.
@imageN is the last frame, defining &amp;lt;the composition, subject position,
pose, prop state and camera direction at the end&amp;gt;.

The picture passes in turn through the states defined by @image1, @image2,
@image3 … @imageN, moving between stages with continuous action.
Throughout, keep &amp;lt;subject identity, prop structure and ownership, scene
layout, light and the camera axis&amp;gt; continuous.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

Use @image1 through @image4 as keyframes, in that order.

@image1 is the first frame: an orange paper plane at rest on the left of a
classroom desk, nose pointing to the right of frame, locked-off medium shot.
@image2 defines the second keyframe: the same orange paper plane lifted off
the desk by a hand, nose direction unchanged.
@image3 defines the third keyframe: the same plane passing the window, the
curtain swaying slightly to the right.
@image4 is the last frame: the same plane landing on the middle shelf of
the bookcase on the right, nose still pointing right.

The picture passes in turn through the states defined by @image1 to
@image4, keeping flight direction and speed continuous between stages.
Throughout, keep the plane's orange stock, size and fold lines, the
classroom layout, the afternoon side light and the camera axis continuous.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Independent keyframes usually align more easily than several panels packed into one grid image. Understand what they control: &lt;strong&gt;stage order and key states — not frame-by-frame reproduction&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Storyboard grids.&lt;/strong&gt; A grid gives the overall story, the shot order and rough composition. It is not suited to demanding that each cell be reproduced in detail. Keep it under about 15 cells, use clean line art or a tidy schematic, and minimise text labels. State the reading order in the prompt, then write each shot's subject action, shot size or camera move, plus the final look and sound.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE

@image1 supplies the shot order and rough composition of an &amp;lt;N-cell
storyboard&amp;gt;, read &amp;lt;left to right, top to bottom&amp;gt;; do not take &amp;lt;the line-art
style, text labels or placeholder figures&amp;gt; from it.
@image2 defines &amp;lt;SUBJECT A&amp;gt;'s &amp;lt;appearance and clothing&amp;gt;.
@image3 defines &amp;lt;key prop or scene&amp;gt;'s &amp;lt;structure, material or light&amp;gt;.

Shot 1: &amp;lt;shot size, subject action, scene state&amp;gt;.
Shot 2: &amp;lt;shot size, subject action, camera move or cut&amp;gt;.
…
Shot N: &amp;lt;closing action and final picture state&amp;gt;.

The final image uses &amp;lt;visual style&amp;gt;; sound includes &amp;lt;dialogue, ambience,
action effects or music&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@image1 supplies the shot order and rough composition of a four-cell
pottery storyboard, read left to right, top to bottom; do not take the
line-art style or the text labels.
@image2 defines &amp;lt;CERAMICIST&amp;gt;'s face, short hair and dark grey apron.
@image3 defines &amp;lt;BLUE GLAZED CUP&amp;gt;'s proportions, glaze colour and curved
handle.

Shot 1: a wide shot of the quiet pottery studio, &amp;lt;CERAMICIST&amp;gt; seated at
        the wheel.
Shot 2: a medium side shot of &amp;lt;CERAMICIST&amp;gt;'s hands steadying the spinning
        wet clay as the wall rises.
Shot 3: a close-up of fingers refining the join between rim and handle,
        slip running slowly off the fingertips.
Shot 4: a medium close-up of the fired &amp;lt;BLUE GLAZED CUP&amp;gt; being set on the
        wooden rack as &amp;lt;CERAMICIST&amp;gt; withdraws both hands.

The final image is realistic documentary; keep the hum of the wheel, the
friction of wet clay and studio room tone.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;White-box references.&lt;/strong&gt; These split into coarse and fine, and picking the wrong one wastes the asset.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;COMPARE

Coarse white-box
  Suits      previsualising action, paths, blocking, camera moves or cuts
             with simple geometry
  Asset      clear relationships between solids, a complete action
             sequence; can be layered with people, prop and scene images
  Prompt     map every white-box solid one by one, and state which
             temporal and spatial information is inherited

Fine white-box
  Suits      a finished model that needs different people, materials,
             colour, scene or style
  Asset      complete structure and a clean plate — no trajectory lines,
             axes or camera frustums
  Prompt     keep structure, action and camera; state which attributes
             are to be re-rendered
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A coarse white-box is good for locking motion paths, direction of travel, where people stand, entrances and exits, camera paths, cut points, light changes and sound rhythm. Every solid should map to a final subject or prop; appearance comes from other reference images.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MAPPING — what the white-box supplies → what you must state

Blocking   motion paths, direction of travel, where people stand,
           entrance and exit order
Camera     position, camera path, direction and changes of speed
Light      light direction, changes in brightness, and when they happen
Cuts       cut points, and the subject and composition either side
Sound      whether dialogue, music, ambience or action effects are
           inherited
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Prefer simple solids with clear relationships. Appendages such as arms or wings are only suitable as white-box information when the action sequence is complete; otherwise they tend to produce stiff movement or misread structure.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — coarse white-box

@video1 is a coarse white-box reference. It supplies only &amp;lt;motion paths,
blocking, camera position, camera movement, cuts, light changes, sound
rhythm or spatial relationships&amp;gt;; do not take the white-box appearance,
materials or scene from it.
The &amp;lt;white-box subject A&amp;gt; in @video1 corresponds to &amp;lt;SUBJECT A&amp;gt;.
The &amp;lt;white-box subject B or geometric prop&amp;gt; in @video1 corresponds to
&amp;lt;SUBJECT B or key prop&amp;gt;.
@image1 defines &amp;lt;SUBJECT A&amp;gt;'s &amp;lt;appearance, clothing or structure&amp;gt;.
@image2 defines &amp;lt;SUBJECT B, key prop or scene&amp;gt;'s &amp;lt;specified attributes&amp;gt;.

&amp;lt;SUBJECT&amp;gt; completes &amp;lt;the main action or event&amp;gt; in &amp;lt;scene&amp;gt;.
Keep &amp;lt;the motion paths, blocking, camera movement, cuts, light or sound
rhythm&amp;gt; from @video1.
The final image uses &amp;lt;people, scene, materials and visual style&amp;gt;; sound
includes &amp;lt;dialogue, ambience or action effects&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@video1 is a coarse white-box reference. It supplies only the walking path,
the direction the cart travels, the locked-off camera position, one push-in
and two cuts; do not take the grey solids' appearance or the empty scene.
The tall cylinder in @video1 corresponds to &amp;lt;DOCENT&amp;gt;.
The long box in @video1 corresponds to &amp;lt;MOBILE CART&amp;gt;.
@image1 defines &amp;lt;DOCENT&amp;gt;'s face, blue uniform and name badge.
@image2 defines &amp;lt;MOBILE CART&amp;gt;'s white metal frame and clear cover.
@image3 defines the curved wall, grey floor and linear ceiling lights of a
science gallery.

&amp;lt;DOCENT&amp;gt; pushes &amp;lt;MOBILE CART&amp;gt; along the curved wall, stops in front of the
central plinth and opens the clear cover.
Keep the walking path, blocking, push-in direction and cut points from
@video1.
The image is bright, realistic museum-documentary; keep footsteps, the
sound of the wheels and gallery room tone.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A fine white-box already has complete people, props or scene structure, and suits changing materials, colour, cast, scene and overall style. The plate should be clean — no trajectory lines, axes, controllers or camera frustums.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — fine white-box

@video1 is a fine white-box reference. Keep &amp;lt;subject structure, action,
spatial layout, camera position, camera movement and cuts&amp;gt;; do not take
the original grey-model materials or the empty background.
@image1 defines &amp;lt;subject&amp;gt;'s &amp;lt;cast, material, colour or surface detail&amp;gt;.
@image2 defines &amp;lt;scene&amp;gt;'s &amp;lt;space, materials, light or visual style&amp;gt;.

Re-render the &amp;lt;subject&amp;gt; in @video1 as &amp;lt;final subject&amp;gt;, and the scene as
&amp;lt;final scene&amp;gt;.
Keep &amp;lt;the structure, action, camera and spatial relationships&amp;gt; from
@video1; the image shows &amp;lt;materials, colour and style&amp;gt;; sound includes
&amp;lt;ambience, effects or music&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@video1 is a fine white-box reference. Keep the complete structure of the
ring assembly, the rotational relationship between the three rings, the
plinth position, the orbiting camera move and the cuts; do not take the
original grey-model materials or the empty background.
@image1 defines the brushed brass material of the outer ring.
@image2 defines the translucent blue glass material of the inner blades.
@image3 defines the white curved wall, dark grey floor and soft ceiling
light of a contemporary art gallery.

Re-render the ring assembly in @video1 as a kinetic sculpture of brass and
blue glass, and the scene as the contemporary art gallery.
Keep the structure, rotation rhythm, orbiting camera move and cuts from
@video1; keep the low hum of the mechanism and quiet interior room tone.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3.2 One-click film
&lt;/h3&gt;

&lt;p&gt;One-click film assembles several images — or images plus a style reference video — into a finished piece with consistent pacing and packaging. The prompt has to state each asset's role, the image order, how much the picture moves, the cutting rhythm, the visual packaging and the sound. Do not just write "turn these into a video".&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ORDER

asset roles → image order → amount of motion → cutting style
  → visual packaging → sound
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE

[ASSET ROLES]
@image1 supplies &amp;lt;person, product, scene or opening shot&amp;gt;.
@image2 supplies &amp;lt;person, product, scene or process shot&amp;gt;.
@image3 supplies &amp;lt;person, product, scene or closing shot&amp;gt;.
@video1 supplies &amp;lt;cutting rhythm, transitions, caption packaging or music
        style&amp;gt; only; do not take the identities or locations in it.
        (optional)

[ARRANGEMENT]
The images appear in &amp;lt;upload order / a specified order / freely arranged
by theme&amp;gt;.
&amp;lt;State the person, product, place and event relationships to preserve.&amp;gt;

[PICTURE MOTION]
Each image uses &amp;lt;a slight live effect, parallax, push/pull, lateral move
or localised motion&amp;gt;.
Keep &amp;lt;subject appearance, product structure, text or background
relationships&amp;gt; stable.

[FINISHED STYLE]
Use &amp;lt;cutting rhythm, transition style, caption or graphic packaging,
colour style&amp;gt;.

[SOUND]
Include &amp;lt;dialogue, ambience, effects or music&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE — night-market travel short

[ASSET ROLES]
@image1 supplies the market entrance and the opening environment.
@image2 supplies &amp;lt;TRAVELLER&amp;gt; walking along the street.
@image3 supplies the lantern stall and craft details.
@image4 supplies three friends around a table eating.
@image5 supplies the riverside at night and its reflections.
@image6 supplies the closing shot of the three on the bridge.
@video1 supplies the brisk cutting rhythm, hand-drawn stickers and
        transition style only; do not take the identities or locations.

[ARRANGEMENT]
The images appear in the order @image1 to @image6, forming the sequence
arrive → wander → eat → walk → group photo.
Keep the three friends' appearance and clothing stable; do not blend them
into each other.

[PICTURE MOTION]
Environment images use a slow push-in and slight parallax; images with
people add only natural blinking, head turns, raised cups and clothing
moving in the wind.
Keep the stall structure, table positions and the bridge railing stable.

[FINISHED STYLE]
Use a bright travel-short rhythm; join scenes with natural wipes and
related colours; hand-drawn stickers appear only at the edge of frame.

[SOUND]
Keep market crowd noise, the clink of tableware and river wind, with light
instrumental music that does not dominate.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When image order matters, write the order image by image. When you are happy for the model to arrange them, say so explicitly — "may be arranged freely by theme". With several people or several products, name and bind each one as usual.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.3 Seamless transitions between two videos
&lt;/h3&gt;

&lt;p&gt;A seamless transition generates continuous material between two clips. State which clip comes before and which after, then the trigger action, camera movement, how the picture changes, the state it arrives at, and how the sound joins.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ORDER

before-clip → after-clip → trigger action → camera movement
  → picture morph → arrival state → sound
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FIVE WAYS — what each one needs stated

Dive or whip-back    camera direction, changes of speed, and when the
                     next scene is entered
Subject rotation     the pose, the direction of rotation, and how clothing
                     or background changes continuously
Foreground wipe      when the occluder fills the frame, and the
                     composition revealed after it
Object morph         the corresponding shapes and materials before and
                     after, and the morphing process
Push/pull or focus   camera movement, the focus target, and the
                     continuous spatial relationship between the two
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE

@video1 is the before-clip; take its &amp;lt;closing subject, action,
composition, camera direction and sound&amp;gt;.
@video2 is the after-clip; take its &amp;lt;opening subject, composition, camera
direction and sound&amp;gt;.
Keep the &amp;lt;identities, product structure, scenes and main actions&amp;gt; in the
original @video1 and @video2 stable.

At the end of @video1, &amp;lt;subject or foreground object&amp;gt; triggers the
transition through &amp;lt;action&amp;gt;.
The camera &amp;lt;direction and change of speed&amp;gt;, and the &amp;lt;shape, material,
light or space&amp;gt; in frame gradually becomes &amp;lt;the corresponding element&amp;gt; at
the start of @video2.
The transition ends arriving naturally at @video2's opening composition,
keeping &amp;lt;subject position, camera direction and motion trend&amp;gt; continuous.
Sound moves smoothly from &amp;lt;the first clip's sound&amp;gt; to &amp;lt;the second clip's
sound&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

@video1 is the before-clip; take the rainy night street, the red umbrella,
the slow push-in and the rain.
@video2 is the after-clip; take the circular skylight of the gallery, the
rising camera and the quiet interior reverb.
Keep the people, the street, the gallery structure and the main actions in
both original clips stable.

At the end of @video1, the red umbrella moves toward the lens and
gradually covers the whole frame, triggering the transition.
The camera continues pushing forward; the circular edge of the umbrella
gradually becomes the metal ring of the gallery skylight, and the red
canopy gradually becomes the white daylight coming through it.
The transition ends arriving naturally at @video2's opening low-angle
composition, the camera turning smoothly from a forward push to a rise.
The rain fades and passes smoothly into footsteps echoing inside the
gallery.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The aim is continuity of picture and sound. You can ask for the main content of both source clips to be preserved, but a generative transition is not a pixel-identical edit splice.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.4 Emotional direction and observable performance
&lt;/h3&gt;

&lt;p&gt;Words like "tense", "warm" or "oppressive" give the model an overall direction, and leave the performance wide open. To control acting reliably, name what can be directly seen or heard: eyes, brow, mouth corner, breathing, gaze direction, hand movement. You do not need to list every facial detail. For a single emotional turn, two to four of the clearest signals is usually enough; stage it across multiple beats only when the emotion genuinely turns more than once.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — single turn, the default

Overall the emotion moves from &amp;lt;starting emotion&amp;gt; to &amp;lt;ending emotion&amp;gt;.
After &amp;lt;trigger event&amp;gt;, &amp;lt;subject&amp;gt; first shows &amp;lt;the immediate observable
reaction&amp;gt;.
Then &amp;lt;eyes, brow, mouth corner, breathing, gaze or hand movement&amp;gt;
gradually &amp;lt;changes&amp;gt;.
Finally &amp;lt;subject&amp;gt; expresses &amp;lt;the target emotion&amp;gt; through &amp;lt;a restrained or
explicit outward sign&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TEMPLATE — multiple stages, keyed to triggers

On hearing or seeing &amp;lt;the first trigger&amp;gt;, &amp;lt;the subject's first observable
reaction&amp;gt;.
After &amp;lt;the second trigger&amp;gt; appears, &amp;lt;the change in the subject's
expression, gaze or breathing&amp;gt;.
Once &amp;lt;the key information&amp;gt; is confirmed, &amp;lt;the emotion the subject is
trying to contain or hide&amp;gt; gradually surfaces through &amp;lt;observable signs&amp;gt;.
Finally, &amp;lt;the subject's closing action, expression or manner of speaking&amp;gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EXAMPLE

Applause for the end of the performance carries from behind the stage. The
young actor's fingers stop on the programme; her gaze turns slowly toward
the curtain while her shoulders stay tight.

Once the curtain call is confirmed, she lets out a small breath, her
shoulders gradually loosen, a restrained smile reaches the corner of her
mouth and her eyes slowly well up — but she never turns to leave.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3.5 Professional camera terms
&lt;/h3&gt;

&lt;p&gt;Basic camera language and popular moves can go straight into the prompt. When a term is uncommon, open to several readings, or when you need precise control of how the picture changes, also state what it acts on, how the picture changes, and the result you want.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BASIC

Shot size    extreme wide, wide, medium, close, extreme close-up
Movement     push, pull, pan, track, follow, orbit, dive, pull-back,
             tilt-up, handheld shake
Position     low angle, overhead, first person
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POPULAR — seven moves and what to add

Oner              the subjects, spaces and event order the camera passes
                  through continuously
Dolly zoom        the size the subject holds, and whether the background
                  is drawn in or pushed away
Aerial            the height, direction of travel, and how much
                  environment to reveal
FPV               the first-person flight or run path, speed and turns
Bullet time       which action is frozen or slowed, and the direction of
                  the orbit
Handheld          who is being followed and how much shake — avoid a
                  "handheld feel" with no subject
Speed ramp        where the action accelerates, decelerates or snaps
                  back, and the final resting state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These can be used as-is; but when several subjects are in frame, still state who the move is built around, where it starts and where it ends.&lt;/p&gt;

&lt;p&gt;For terms that are genuinely niche, whose meaning varies across the industry, or that the model may not know, keep the term and translate it into observable change.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FORMULA

term + what it acts on + how the picture changes
  + foreground/background relationship + direction or speed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Rack focus: focus moves smoothly from the foreground leaves to the figure
behind. The leaves gradually soften; the figure's face resolves from
blurred to sharp.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a precise transition, also write the trigger moment, the occluder, the camera direction, how the cut happens, and the composition or motion trend to preserve afterwards.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FIVE EXAMPLES

1 · Shallow-depth portrait
  &amp;lt;PASTRY CHEF&amp;gt;'s eyes and face stay sharp; the glass jars and lights
  behind fall away into soft round bokeh.

2 · Tracking shot
  The camera moves horizontally at the same speed as &amp;lt;SKATER&amp;gt;; the figure
  stays sharp while the street wall smears horizontally from right to left.

3 · Golden hour
  Warm low-angle sun enters from behind &amp;lt;CLIMBER&amp;gt;'s left; the ridge
  ground carries long shadows.

4 · Natural vignette
  The four corners darken gradually while the central &amp;lt;PIANIST&amp;gt; keeps
  normal brightness and skin tone; no black border appears.

5 · Whip-pan transition
  At 5s the camera whips left; when the foreground bookcase fills the
  frame it cuts to the next scene, and after the cut the camera keeps
  moving left at a similar speed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Aperture, focal length and shutter values can go into the prompt, but the visible result is usually clearer than a number on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Checklist before you submit
&lt;/h2&gt;

&lt;p&gt;The guide closes with fourteen questions. Run them before you spend a generation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; 1  Are the subject and the main action or event stated clearly?
 2  Does every reference asset state what to take AND what not to take?
 3  Is each person, product and prop named and bound to its asset?
 4  Are assets selected per scene rather than all forced on screen at once?
 5  Does each stage of a long video have one main change and a written
    end state?
 6  Are headcount, clothing, prop ownership and spatial relationships
    stable?
 7  Does the edit declare a sole master, the scope, the target count and
    what stays?
 8  Are abstract emotions and camera terms cashed out as something
    directly seen or heard?
 9  Do first/last frames and multi-keyframes each state their own role,
    and do first and last frame share one aspect ratio?
10  Does the storyboard state which structure is inherited?
11  For a white-box, have you decided coarse or fine, and written which
    timing, structure, material and style are inherited?
12  Do editing, first/last-frame generation and extension respect the
    auto-locked ratio and duration rules?
13  For an extension, have you checked the boundary frame, the motion
    trend AND sound continuity?
14  Does one-click film state asset roles, image order, amount of motion,
    cutting style and sound? Does a seamless transition state both clips'
    roles, the trigger, the transition process and the arrival state?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  5. What it will not do
&lt;/h2&gt;

&lt;p&gt;Nine stated boundaries. These are the requests that end in disappointment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1  Timestamps allocate event pacing only; they are not frame-level cut
   points.
2  An edit prompt raises the probability that key events align with the
   source, but cannot guarantee frame-exact overlap.
3  Multi-asset work is about selecting and combining the right assets, not
   about showing them all at once.
4  Subtitles, formulas, signage, product specs and frame-precise timing
   that must be exactly right should be handled with pre-composited assets
   plus generation plus post-production.
5  Video editing locks the input's ratio and essential duration; neither is
   settable, and output duration may differ from the input by up to about
   0.3 seconds.
6  First-frame or first/last-frame generation locks the ratio to the first
   frame; duration is settable. Mismatched first/last ratios can stretch
   the last frame.
7  Video extension locks the input's ratio; the added duration is settable.
   The extension's loudness may differ slightly from the source.
8  In one-click film, if image order or character correspondence matters,
   it must be specified in the prompt.
9  A seamless transition aims at visual and audio continuity; it does not
   mean the two source clips stay pixel-identical.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Running these prompts today
&lt;/h2&gt;

&lt;p&gt;Seedance 2.5 is not open through us yet — verified against our own Volcengine Ark account on 4 August 2026, and we do not list a model whose generate button would fail. &lt;strong&gt;Seedance 2.0 is live from $0.092/s&lt;/strong&gt; and shares most of the grammar above: named subjects, per-asset roles, staged events with end states, native synced audio, up to 9 reference images plus reference video and audio. The formula, the naming discipline, the stage-with-end-state structure and the inheritance clause all transfer. When 2.5 opens, the same key and the same endpoint reach it — you change one model string.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://apimodels.app/blog/seedance-2-5-prompting-guide" rel="noopener noreferrer"&gt;apimodels.app&lt;/a&gt;, where the same guide is also available &lt;a href="https://apimodels.app/zh/blog/seedance-2-5-prompting-guide" rel="noopener noreferrer"&gt;in Chinese&lt;/a&gt; with Chinese-language templates.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Source: ByteDance's official Seedance 2.5 prompting guide — the limits, syntax, auto-locked parameters, checklist and stated boundaries are theirs; the worked examples and the English phrasing are mine. Status figures were measured on 4 August 2026.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If you want to try this grammar today, &lt;a href="https://apimodels.app/models/seedance-2.0" rel="noopener noreferrer"&gt;Seedance 2.0&lt;/a&gt; shares most of it and runs from $0.092/s.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>videogeneration</category>
      <category>prompt</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
