<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: WesLin</title>
    <description>The latest articles on DEV Community by WesLin (@codesugar_lin_037a57b06a4).</description>
    <link>https://dev.to/codesugar_lin_037a57b06a4</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3688057%2F0e4fd75d-bad1-40db-870d-1b00c924cd55.jpg</url>
      <title>DEV Community: WesLin</title>
      <link>https://dev.to/codesugar_lin_037a57b06a4</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/codesugar_lin_037a57b06a4"/>
    <language>en</language>
    <item>
      <title>Build Clean Animated GIFs From Sprite Sheets With ffmpeg</title>
      <dc:creator>WesLin</dc:creator>
      <pubDate>Sun, 13 Sep 2026 11:28:07 +0000</pubDate>
      <link>https://dev.to/codesugar_lin_037a57b06a4/build-clean-animated-gifs-from-sprite-sheets-with-ffmpeg-5eeh</link>
      <guid>https://dev.to/codesugar_lin_037a57b06a4/build-clean-animated-gifs-from-sprite-sheets-with-ffmpeg-5eeh</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;If you point ffmpeg straight at your raw sprite frames, you get a bloated 611KB GIF that shimmers on every frame. You can fix both problems with a two-pass ffmpeg command that builds one shared palette:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Step 1: Generate a single 64-colour palette across all frames&lt;/span&gt;
ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; frame-%02d.png &lt;span class="nt"&gt;-vf&lt;/span&gt; &lt;span class="s2"&gt;"palettegen=max_colors=64:stats_mode=diff"&lt;/span&gt; &lt;span class="nt"&gt;-update&lt;/span&gt; 1 palette.png

&lt;span class="c"&gt;# Step 2: Render the GIF using that palette with dithering disabled&lt;/span&gt;
ffmpeg &lt;span class="nt"&gt;-framerate&lt;/span&gt; 12 &lt;span class="nt"&gt;-i&lt;/span&gt; frame-%02d.png &lt;span class="nt"&gt;-i&lt;/span&gt; palette.png &lt;span class="nt"&gt;-lavfi&lt;/span&gt; &lt;span class="s2"&gt;"paletteuse=dither=none"&lt;/span&gt; run-cycle.gif
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These two steps shrink your file size by 78%. But palette work only fixes half your problem. If your animation still jitters, your cut is the reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  The size ladder: why palettegen matters
&lt;/h2&gt;

&lt;p&gt;A standard GIF supports at most 256 colours. A raw image from a model holds about 48,000 colours across subtle shifts in your background. &lt;/p&gt;

&lt;p&gt;When you export without a custom palette, ffmpeg picks a separate 256-colour map for every single frame. That makes your background flicker, and it blows up your file size.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Export method&lt;/th&gt;
&lt;th&gt;File size&lt;/th&gt;
&lt;th&gt;What changed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Even grid cut, default palette&lt;/td&gt;
&lt;td&gt;611,294 bytes&lt;/td&gt;
&lt;td&gt;The bloated baseline.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sprite-aligned, default palette&lt;/td&gt;
&lt;td&gt;241,987 bytes&lt;/td&gt;
&lt;td&gt;Alignment alone cuts 60% of the byte weight.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sprite-aligned, shared 64-colour palette&lt;/td&gt;
&lt;td&gt;165,345 bytes&lt;/td&gt;
&lt;td&gt;The shared palette removes another 31%.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4-frame sheet, aligned, shared palette&lt;/td&gt;
&lt;td&gt;130,635 bytes&lt;/td&gt;
&lt;td&gt;The final production build.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Setting &lt;code&gt;stats_mode=diff&lt;/code&gt; tells ffmpeg to build the palette around the pixels that move between frames. Setting &lt;code&gt;dither=none&lt;/code&gt; stops speckling on your flat pixel-art surfaces.&lt;/p&gt;

&lt;h2&gt;
  
  
  The math: why even grid slicing fails
&lt;/h2&gt;

&lt;p&gt;Most developers slice sprite sheets by dividing the canvas into equal boxes. If your image is 1672 pixels wide with 4 columns, you expect each box to be 418 pixels wide.&lt;/p&gt;

&lt;p&gt;We tested sheets from the &lt;a href="https://seadanse.com/models/gpt-image-2-5" rel="noopener noreferrer"&gt;GPT Image 2.5 generator&lt;/a&gt; across Flare and Sunburst checkpoints. We asked for a 1280x720 sheet, but every one came back at 1672x941 pixels. &lt;/p&gt;

&lt;p&gt;Here's what your actual pixel offsets look like on the raw sheet:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Canvas width: 1672 px
Target columns: 4
Calculated cell width: 1672 / 4 = 418 px

Actual frame start positions (Flare 8-frame):
Frame 1: 145 px
Frame 2: 554 px (gap: 409 px)
Frame 3: 928 px (gap: 374 px)
Frame 4: 1295 px (gap: 367 px)

Drift on frame 4: 418 - 367 = 51 px off-centre
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The characters don't sit on an even grid. The gap between frame starts shrinks from 409 pixels down to 367 pixels across the row. &lt;/p&gt;

&lt;p&gt;If you slice by fixed 418-pixel steps, your character drifts 104 pixels sideways and bounces 48 pixels up and down. That's half the character's body width in drift.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0zdk9v273d19t9ww7j68.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0zdk9v273d19t9ww7j68.png" alt="Eight frames averaged into one image twice: the grid cut on the left is a blur of eight scattered characters, the aligned cut on the right is one sharp character with only the limbs blurred" width="799" height="456"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Align on the baseline, not the cell
&lt;/h2&gt;

&lt;p&gt;To get rid of the bounce, you need to ignore the canvas grid and measure the pixels directly:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Find each sprite's actual box inside the raw image.&lt;/li&gt;
&lt;li&gt;Crop tightly to your sprite's edges.&lt;/li&gt;
&lt;li&gt;Place each cropped sprite onto one shared canvas.&lt;/li&gt;
&lt;li&gt;Centre the crop from left to right, and line up the bottom edge where the feet land.&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Slicing approach&lt;/th&gt;
&lt;th&gt;Horizontal drift&lt;/th&gt;
&lt;th&gt;Vertical bounce&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Even 4x2 grid&lt;/td&gt;
&lt;td&gt;104 px&lt;/td&gt;
&lt;td&gt;48 px&lt;/td&gt;
&lt;td&gt;Character slides and bounces constantly.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bounding box, bottom-aligned (8 frames)&lt;/td&gt;
&lt;td&gt;19 px&lt;/td&gt;
&lt;td&gt;0 px&lt;/td&gt;
&lt;td&gt;Vertical bounce drops to zero.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bounding box, bottom-aligned (4 frames)&lt;/td&gt;
&lt;td&gt;1 px&lt;/td&gt;
&lt;td&gt;0 px&lt;/td&gt;
&lt;td&gt;Character stays locked in place.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Lining up the feet eliminates the 48-pixel vertical hop entirely. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fycx1okvlglwgkrzo8z2m.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fycx1okvlglwgkrzo8z2m.gif" alt="A four-frame animated GIF of the pixel-art courier running in place with no bounce" width="245" height="353"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why you should make 4 frames instead of 8
&lt;/h2&gt;

&lt;p&gt;When you prompt for an 8-frame run cycle across two rows, the model duplicates poses. &lt;/p&gt;

&lt;p&gt;We measured frame similarity using mean absolute difference (MAD), which scores the average per-pixel difference between two images. A low score means two frames are near copies:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Top frame 3 vs bottom frame 7: 4.95 MAD&lt;/li&gt;
&lt;li&gt;Top frame 2 vs bottom frame 6: 5.62 MAD&lt;/li&gt;
&lt;li&gt;Top frame 4 vs bottom frame 8: 5.73 MAD&lt;/li&gt;
&lt;li&gt;Top frame 1 vs bottom frame 5: 7.21 MAD&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Genuinely distinct animation frames score 12.36 MAD or higher. The second row of an 8-frame grid merely repeats the first row with minor errors.&lt;/p&gt;

&lt;p&gt;On Sunburst sheets, the gaps between frame starts measured 400, 366, and 417 pixels. On Flare, the gaps measured 409, 374, and 367 pixels. &lt;/p&gt;

&lt;p&gt;A single-row 4-frame sheet keeps your frame starts within 7 pixels of each other. Sprite widths on the 4-frame sheet varied by only 2 pixels (227 to 229 pixels), compared to a 39-pixel spread on the 8-frame layout.&lt;/p&gt;

&lt;h2&gt;
  
  
  The video model route vs direct cutting
&lt;/h2&gt;

&lt;p&gt;You can also pass your raw sprite sheet into an image-to-video model like Seedance 2.0 Mini. &lt;/p&gt;

&lt;p&gt;We tested this with our 4-frame sheet at 480p. The model produced an 864x496 video with 121 frames in 93 seconds.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Direct sheet crop&lt;/th&gt;
&lt;th&gt;Video generation route&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Generation compute cost&lt;/td&gt;
&lt;td&gt;1x&lt;/td&gt;
&lt;td&gt;about 19x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frame control&lt;/td&gt;
&lt;td&gt;4 exact approved frames&lt;/td&gt;
&lt;td&gt;121 generated frames&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Palette size per frame&lt;/td&gt;
&lt;td&gt;64 colours&lt;/td&gt;
&lt;td&gt;7,433 to 8,194 colours&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Final GIF size&lt;/td&gt;
&lt;td&gt;130,635 bytes&lt;/td&gt;
&lt;td&gt;530,249 bytes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Baseline drift&lt;/td&gt;
&lt;td&gt;0 px&lt;/td&gt;
&lt;td&gt;3 px&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The video route keeps your character stable, with just 3 pixels of drift across five seconds. &lt;/p&gt;

&lt;p&gt;But the video model redraws your character on every single frame. You lose your clean pixel outlines, and the colour count multiplies by more than a hundred times.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to make the raw sheet
&lt;/h2&gt;

&lt;p&gt;To try these steps yourself, open the &lt;a href="https://seadanse.com/models/gpt-image-2-5" rel="noopener noreferrer"&gt;GPT Image 2.5 tool&lt;/a&gt; and pick your settings. &lt;/p&gt;

&lt;p&gt;OpenAI positions Flare for fast runs and Sunburst for tasks where you need higher editing precision. The composer lets you set resolution tiers (1K, 2K, 4K) and aspect ratios from one menu.&lt;/p&gt;

&lt;p&gt;Use this prompt to make your source frames:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A 2D pixel-art sprite sheet of one character: a green-hooded courier with a brown satchel, side view facing right, running cycle, 4 frames in a single horizontal row, evenly spaced with equal margins on a flat mid-grey background, identical character height, proportions and colour palette in every frame, crisp 1px outlines, limited 16-colour palette, no text, no frame numbers, no grid lines, no drop shadows, no background scenery.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three limits will stay on the raw output:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Background colour values shift slightly between runs (such as RGB 122,121,122 vs 127,126,125), so don't use hardcoded hex keys.&lt;/li&gt;
&lt;li&gt;Files download as JPGs without an alpha layer.&lt;/li&gt;
&lt;li&gt;Two-row requests duplicate poses across rows.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can run your prompts directly on &lt;a href="https://seadanse.com/models/gpt-image-2-5" rel="noopener noreferrer"&gt;GPT Image 2.5 on Seadanse&lt;/a&gt; with starter credits upon sign-up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Workflow checklist
&lt;/h2&gt;

&lt;p&gt;Apply this checklist when you build your next sprite animation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Ask for 4 frames in a single horizontal row, not 8 frames in a grid.&lt;/li&gt;
&lt;li&gt;Read your image dimensions from the file header instead of assuming your requested canvas size.&lt;/li&gt;
&lt;li&gt;Crop around individual sprite boxes rather than using equal grid cuts.&lt;/li&gt;
&lt;li&gt;Align each cropped frame to a common bottom baseline where the feet land.&lt;/li&gt;
&lt;li&gt;Make a shared 64-colour palette with &lt;code&gt;palettegen=stats_mode=diff&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Build your final GIF using &lt;code&gt;paletteuse=dither=none&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We work on Seadanse, which runs GPT Image 2.5, and the tests in this post were run there.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Wan 3.0 API spec and the three reference arrays</title>
      <dc:creator>WesLin</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:43:57 +0000</pubDate>
      <link>https://dev.to/codesugar_lin_037a57b06a4/wan-30-api-spec-and-the-three-reference-arrays-4kj1</link>
      <guid>https://dev.to/codesugar_lin_037a57b06a4/wan-30-api-spec-and-the-three-reference-arrays-4kj1</guid>
      <description>&lt;p&gt;The landing page for Wan 3.0 promises "up to 20 reference assets", but the underlying API does not accept a single list of twenty items. It splits those inputs across three separate arrays, each with its own count, file type, and duration cap.&lt;/p&gt;

&lt;h2&gt;
  
  
  What one wan3.0-video request actually takes
&lt;/h2&gt;

&lt;p&gt;The API uses a single model identifier, &lt;code&gt;wan3.0-video&lt;/code&gt;, across every generation mode. The fields you include in the payload determine which mode runs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;prompt&lt;/code&gt;: Text string up to 20,000 characters describing the shot.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;resolution&lt;/code&gt;: Output size, accepting &lt;code&gt;480P&lt;/code&gt;, &lt;code&gt;720P&lt;/code&gt;, or &lt;code&gt;1080P&lt;/code&gt;. The provider default is &lt;code&gt;1080P&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;duration&lt;/code&gt;: Video length from 2 to 30 seconds, or &lt;code&gt;-1&lt;/code&gt; to let the model select the duration. Default is 5.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;audio&lt;/code&gt;: Boolean controlling sound generation, defaulting to &lt;code&gt;true&lt;/code&gt;. Sound generates in the same pass as the video frames rather than as a post-process.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;generation_type&lt;/code&gt;: Processing mode flag, set to &lt;code&gt;frame&lt;/code&gt; for frame-driven generation or &lt;code&gt;reference&lt;/code&gt; for asset-conditioned generation.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;image_urls&lt;/code&gt;: Array of image URLs for single-frame generation or reference images.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;image_with_roles&lt;/code&gt;: Array of image objects containing URLs and explicit roles (&lt;code&gt;first_frame&lt;/code&gt;, &lt;code&gt;last_frame&lt;/code&gt;, or &lt;code&gt;reference_image&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;video_urls&lt;/code&gt;: Array of reference video URLs.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;audio_urls&lt;/code&gt;: Array of reference audio URLs.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;file_url&lt;/code&gt;: URL for one reference document, up to 100 MB and 50 pages.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;link_url&lt;/code&gt;: URL for one public web page reference.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;size&lt;/code&gt;: Output dimensions and aspect ratio.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;seed&lt;/code&gt;: Integer seed for output repeatability.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;nsfw_check&lt;/code&gt;: Content filtering toggle.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;watermark&lt;/code&gt;: Watermark inclusion toggle.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Payload construction determines the generation mode. Passing two frames into &lt;code&gt;image_with_roles&lt;/code&gt; triggers first-and-last frame interpolation. Passing a single image into &lt;code&gt;image_urls&lt;/code&gt; with &lt;code&gt;generation_type: frame&lt;/code&gt; runs standard image-to-video. Passing image URLs with &lt;code&gt;generation_type: reference&lt;/code&gt; runs reference-conditioned generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The twenty is three arrays, not one
&lt;/h2&gt;

&lt;p&gt;Alibaba groups reference handling under the name Omni-Creation. The advertised twenty-asset capacity is divided into three fixed buckets:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Group&lt;/th&gt;
&lt;th&gt;How many&lt;/th&gt;
&lt;th&gt;Other limits&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reference images&lt;/td&gt;
&lt;td&gt;up to 10&lt;/td&gt;
&lt;td&gt;20 MB each; JPG, JPEG, PNG, BMP or WebP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reference video&lt;/td&gt;
&lt;td&gt;up to 5 clips&lt;/td&gt;
&lt;td&gt;1-15 seconds each, 15 seconds in total; MP4 or MOV&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reference audio&lt;/td&gt;
&lt;td&gt;up to 5 clips&lt;/td&gt;
&lt;td&gt;1-15 seconds each, 15 seconds in total; WAV or MP3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The total of 10 images, 5 video clips, and 5 audio clips makes up the advertised twenty. Because each category is isolated, you cannot submit 20 reference images or a single 20-second reference video.&lt;/p&gt;

&lt;p&gt;References bind to prompt text using explicit tags: &lt;code&gt;@Image1&lt;/code&gt;, &lt;code&gt;@Video1&lt;/code&gt;, and &lt;code&gt;@Audio1&lt;/code&gt;. The numeric index maps to the position of the asset inside its respective request array. The API also accepts standalone reference audio in &lt;code&gt;audio_urls&lt;/code&gt; without any attached visual reference.&lt;/p&gt;

&lt;h2&gt;
  
  
  One code block that does the arithmetic
&lt;/h2&gt;

&lt;p&gt;Pricing scales along resolution and duration tiers without offering audio discounts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Resolution cost scaling (fixed duration):
  480p  = 1x  (baseline)
  720p  = 2x
  1080p = 4x

Duration cost scaling (fixed resolution):
  5s    = 1x  (baseline)
  30s   = 6x

Audio setting:
  audio: true  = 1x
  audio: false = 1x (disabling audio does not reduce cost)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scaling a batch from 480p at 5 seconds up to 1080p at 30 seconds increases resource consumption by a factor of 24. You can calculate specific requirements using the &lt;a href="https://seadanse.com/tool/ai-video-cost-calculator-wan-3-0" rel="noopener noreferrer"&gt;Wan 3.0 cost calculator&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we shipped through it, and what the clips showed
&lt;/h2&gt;

&lt;p&gt;We tested the pipeline by sending four clips through the API at 480p, 16:9 aspect ratio, 5 seconds per clip, with audio enabled. Prompts were transmitted exactly as written.&lt;/p&gt;

&lt;p&gt;Clip 1 structured the prompt using a five-element order: shot type, subject, action, lighting, and sound. The render delivered the ceramicist lifting the bowl from the wheel, low window light entering from the left, background studio elements in shadow, and a visible camera push-in.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pub-d877f0c06b75414f98b2a43f60b265e5.r2.dev/seed-dance-2-5/blog/wan-3-0-prompt-guide/ours/M1-five-part-order.mp4" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F82p5ow5awyf098p5ks25.jpg" alt="A ceramicist lifts a wet bowl off the wheel and turns it to check the rim in low window light, generated with Wan 3.0" width="800" height="462"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Generated with Wan 3.0, 480p, 5 seconds, audio on, from the prompt above, unedited — 2026-09-05.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Clip 2 formatted scene ambience as active clauses rather than descriptor lists: rain striking the glass, sodium streetlights crossing the subject's face sequentially, interior lighting restricted to an overhead strip, and a locked-off camera position.&lt;/p&gt;

&lt;p&gt;Clip 3 supplied two reference stills from Nano Banana 2 tagged as &lt;code&gt;@Image1&lt;/code&gt; and &lt;code&gt;@Image2&lt;/code&gt;. The yellow jacket, dark hair, and bag strap mapped from the first image, while the magenta neon, fire escape, and wet ground mapped from the second. Because text-to-image calls retain no state between separate generations, the open-shutter reference still had to be created by editing the closed-shutter still directly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pub-d877f0c06b75414f98b2a43f60b265e5.r2.dev/seed-dance-2-5/blog/wan-3-0-prompt-guide/ours/M4-reference-images.mp4" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8hdd1s8danllo2jqrces.jpg" alt="A courier in a yellow jacket walks into a neon-lit alley and stops to check a package label, generated with Wan 3.0 from two reference images" width="800" height="462"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Generated with Wan 3.0, 480p, 5 seconds, audio on, with two stills attached as @Image1 and &lt;a class="mentioned-user" href="https://dev.to/image2"&gt;@image2&lt;/a&gt; — 2026-09-05.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A single reference test confirms that label binding functioned in this instance, but it does not isolate whether the model resolved the assets via the &lt;code&gt;@Image&lt;/code&gt; tokens or via matching text descriptions like "the courier" and "the alley".&lt;/p&gt;

&lt;p&gt;Clip 4 tested frame-to-video with defined start and end points. The initial frame showed a closed shutter, which raised across the runtime before settling on the final framing. An additional run on 2026-08-16 evaluated a full 30-second continuous shot at 1080p, executing an unbroken crane pull-back from a rooftop herb garden to a dawn skyline. Full prompt text and media outputs are documented in the &lt;a href="https://seadanse.com/blog/wan-3-0-prompt-guide" rel="noopener noreferrer"&gt;Wan 3.0 prompt guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What our own front end does not expose
&lt;/h2&gt;

&lt;p&gt;The interface on the &lt;a href="https://seadanse.com/models/wan-3-0" rel="noopener noreferrer"&gt;Seadanse model page&lt;/a&gt; provides access to text-to-video and frame-to-video generation, durations from 4 to 30 seconds, three resolution profiles, six aspect ratio presets plus Auto, a 10,000-character prompt field, and native audio generation.&lt;/p&gt;

&lt;p&gt;Certain API capabilities are omitted from the front end:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Duration minimum:&lt;/strong&gt; The composer enforces a 4-second minimum rather than the API's 2-second limit due to a shared schema constraint across our interface layer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Document parsing:&lt;/strong&gt; The API accepts PDF and text files via &lt;code&gt;file_url&lt;/code&gt;, which is not exposed in the web composer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Web page parsing:&lt;/strong&gt; The &lt;code&gt;link_url&lt;/code&gt; input is not wired to the UI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;In-place video editing:&lt;/strong&gt; Direct video transformation via API flags is omitted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Upload constraints:&lt;/strong&gt; While the upload panel lists accepted file extensions (JPG, PNG, WebP, MP4, MOV), it does not display the API's per-group quantity and duration limits.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A checklist for the next model API page you read
&lt;/h2&gt;

&lt;p&gt;When evaluating a video model API from its documentation, check these five constraints before writing integration code:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Array segmentation:&lt;/strong&gt; Verify whether multi-asset reference claims allow arbitrary file combinations or mandate fixed allocations across image, video, and audio arrays.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-array duration caps:&lt;/strong&gt; Check if video and audio inputs have aggregate duration ceilings in addition to individual file limits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Default resolutions:&lt;/strong&gt; Check the default API resolution value, as provider defaults like 1080p will run at higher cost multipliers than baseline 480p settings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Duration floors:&lt;/strong&gt; Compare the documented API minimum length against front-end limits to identify where interface validation schemas diverge from underlying endpoints.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weight availability:&lt;/strong&gt; Verify if model weights are downloadable. For Wan 3.0, no public weights exist on Hugging Face; open checkpoints under Alibaba's organisation stop at Wan 2.2 releases such as &lt;code&gt;Wan2.2-TI2V-5B-Diffusers&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The author works on Seadanse, which runs Wan 3.0.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Automating Google Sites: What Worked, What Failed, and What Cost a Rebuild</title>
      <dc:creator>WesLin</dc:creator>
      <pubDate>Wed, 09 Sep 2026 17:44:00 +0000</pubDate>
      <link>https://dev.to/codesugar_lin_037a57b06a4/automating-google-sites-what-worked-what-failed-and-what-cost-a-rebuild-39ji</link>
      <guid>https://dev.to/codesugar_lin_037a57b06a4/automating-google-sites-what-worked-what-failed-and-what-cost-a-rebuild-39ji</guid>
      <description>&lt;p&gt;We needed to publish a 2,921-word article into new Google Sites from a script as part of shipping a model comparison. There isn't an API for the current version of the editor, so we had to test every programmatic input path to find what actually persists.&lt;/p&gt;

&lt;p&gt;Most injection methods failed immediately, and one failed silently after reporting success. Here's every route we tried, the exact way each one broke, and the final &lt;a href="https://sites.google.com/view/seadanse-review/hailuo-3-vs-seedance-2-5" rel="noopener noreferrer"&gt;model comparison on Google Sites&lt;/a&gt; that came out the other end.&lt;/p&gt;

&lt;h2&gt;
  
  
  There is no API, and that is official
&lt;/h2&gt;

&lt;p&gt;If you look for a REST endpoint or an official client library, you won't find one. As &lt;a href="https://developers.google.com/workspace/sites/changelog" rel="noopener noreferrer"&gt;Google's own deprecation notice&lt;/a&gt; explains:&lt;/p&gt;

&lt;p&gt;"Sites API is deprecated and might stop working at any time. Sites API can only access classic Sites. Sites API can't access the rebuilt version of Sites that was launched on November 22, 2016."&lt;/p&gt;

&lt;p&gt;Both the changelog and the developer guide carry that exact banner. That means classic Sites had an API, but the rebuilt version from 2016 never got one. That leaves you with the live editor in a browser tab, and whatever you can make that browser do through automation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we tried, and what came back
&lt;/h2&gt;

&lt;p&gt;We tested seven different routes against the live editor on 2026-09-09 and 2026-09-10.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Route&lt;/th&gt;
&lt;th&gt;What came back&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;document.execCommand('insertHTML', …)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"This document requires 'TrustedHTML' assignment."&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;new DOMParser().parseFromString(html, 'text/html')&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The same Trusted Types error.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Build the DOM under &lt;code&gt;trustedTypes.createPolicy(...)&lt;/code&gt;, then &lt;code&gt;Range.insertNode&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Renders correctly. Autosave says saved to Drive. Gone on reload.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;navigator.clipboard.write([new ClipboardItem(...)])&lt;/code&gt; from an eval&lt;/td&gt;
&lt;td&gt;Never settles. No transient activation.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Synthetic &lt;code&gt;cmd+v&lt;/code&gt; through CDP or an extension&lt;/td&gt;
&lt;td&gt;Nothing happens.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;osascript&lt;/code&gt; System Events &lt;code&gt;keystroke "v" using command down&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;"osascript is not allowed to send keystrokes."&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typing, character by character, through the extension&lt;/td&gt;
&lt;td&gt;Works, and persists.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The third row is the one that's genuinely dangerous. When you build DOM nodes under a trusted types policy and insert them with &lt;code&gt;Range.insertNode&lt;/code&gt;, the page renders your content cleanly.&lt;/p&gt;

&lt;p&gt;The editor shows 18,482 characters, 11 h2 elements, 10 h3 elements, and 11 links. The autosave pill at the top turns into a checkmark and reads "all changes saved to Drive."&lt;/p&gt;

&lt;p&gt;Then you reload the tab. The canvas clears, and you're staring at 0 characters.&lt;/p&gt;

&lt;p&gt;Nothing in that sequence tells you it failed. The reason it disappears is that Google Sites doesn't treat the DOM as its source of truth. It keeps an internal document model in JavaScript.&lt;/p&gt;

&lt;p&gt;When you mutate the DOM directly, that mutation doesn't update the internal model. On the next reload, the editor renders the state of its internal model, which is completely empty.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one route that works
&lt;/h2&gt;

&lt;p&gt;The only approach that updates the document model and survives a reload is sending synthetic keystrokes through a browser automation extension. It simulates a user typing characters one by one. That works, but three specific behaviors will bite you along the way.&lt;/p&gt;

&lt;p&gt;First, synthetic typing silently drops non-ASCII characters. We had an em dash inside the string "TL;DR —", and it showed up in the editor canvas as two plain spaces.&lt;/p&gt;

&lt;p&gt;The console threw no errors, and the extension reported a clean run. You've got to normalise every string to ASCII before you start typing.&lt;/p&gt;

&lt;p&gt;Second, typing "- " at the start of any line triggers a retroactive list format. The editor detects the markdown shortcut and turns the entire text box into a bulleted list, including every paragraph you typed before that line.&lt;/p&gt;

&lt;p&gt;You can't just backspace out of it. The fix is to run a select-all command and toggle the list button twice to strip the formatting.&lt;/p&gt;

&lt;p&gt;Third, adding hyperlinks needs careful handling. You apply a link by selecting the anchor text in the live document and pressing cmd+k. That opens a link dialog with two traps.&lt;/p&gt;

&lt;p&gt;If you press Return inside the URL input, the editor doesn't submit the form; it types the URL directly into the document and replaces your selected anchor with a bare URL. You've got to click the input field and click the apply button with synthetic click events.&lt;/p&gt;

&lt;p&gt;On top of that, the text selector matches the first occurrence of an anchor phrase. If that phrase appears twice on the page, the script links the wrong sentence while the document looks completely normal.&lt;/p&gt;

&lt;p&gt;Headings worked through keyboard shortcuts, but they needed explicit pacing. Pressing &lt;code&gt;cmd+alt+2&lt;/code&gt; and &lt;code&gt;cmd+alt+3&lt;/code&gt; applied real &lt;code&gt;&amp;lt;h2&amp;gt;&lt;/code&gt; and &lt;code&gt;&amp;lt;h3&amp;gt;&lt;/code&gt; tags to the selected line.&lt;/p&gt;

&lt;p&gt;But the editor needed about three seconds between events. When we fired heading shortcuts back to back in a single batch, only the first one took.&lt;/p&gt;

&lt;h2&gt;
  
  
  A code block
&lt;/h2&gt;

&lt;p&gt;Here's the terminal log showing what each programmatic input returned during our tests.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ node -e "…"                      # no API to call
&amp;gt; execCommand('insertHTML', …)     Error: This document requires 'TrustedHTML' assignment.
&amp;gt; new DOMParser().parseFromString  Error: This document requires 'TrustedHTML' assignment.
&amp;gt; trustedTypes.createPolicy(...)   ok
  ... build nodes, Range.insertNode
  editor shows 18,482 chars, 11 h2, 10 h3, 11 links
  autosave: "all changes saved to Drive"
  reload                           0 chars
&amp;gt; navigator.clipboard.write(...)   (never settles)
&amp;gt; key: cmd+v                       (no change)
&amp;gt; osascript keystroke "v"          Error: osascript is not allowed to send keystrokes.
&amp;gt; type("Disclosure: we own ...")   persists
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Two things about the output that cost us a rebuild
&lt;/h2&gt;

&lt;p&gt;Two structural quirks forced us to tear down and rebuild our published layout.&lt;/p&gt;

&lt;p&gt;The first quirk is how Google Sites generates &lt;code&gt;&amp;lt;title&amp;gt;&lt;/code&gt; tags. On a subpage, the &lt;code&gt;&amp;lt;title&amp;gt;&lt;/code&gt; tag matches the page name you assign. On the home page, the &lt;code&gt;&amp;lt;title&amp;gt;&lt;/code&gt; tag is always the site name, no matter what you name the root page.&lt;/p&gt;

&lt;p&gt;We published our article at the site root first, and Google Sites threw the article's title tag away. The fix was opening the page menu (the ⋮ icon) and clicking "Duplicate page."&lt;/p&gt;

&lt;p&gt;That moved the content to a subpage and let us set the subpage name and custom path in a single dialog, so we didn't have to retype anything.&lt;/p&gt;

&lt;p&gt;The second quirk is how the editor handles images. A text box holds text only. Images live in separate layout blocks that you position by dragging, which means an image can't sit inside the paragraph it illustrates.&lt;/p&gt;

&lt;p&gt;Even worse, uploading an image through the editor's file input crashed the client interface with a runtime error. That crash happened on a 1.7 MB PNG and an 89 KB JPEG, so it wasn't a file size limit. That part of the pipeline still needs a human operator.&lt;/p&gt;

&lt;h2&gt;
  
  
  The checklist
&lt;/h2&gt;

&lt;p&gt;If you're asked to automate publishing into a closed WYSIWYG editor you don't control, run these checks before writing your script:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Check whether the editor uses an internal document model instead of the live DOM.&lt;/li&gt;
&lt;li&gt;Reload the page right after synthetic DOM mutations to confirm that autosave actually wrote data to the backend.&lt;/li&gt;
&lt;li&gt;Test whether synthetic keystrokes drop non-ASCII characters like em dashes without logging warnings.&lt;/li&gt;
&lt;li&gt;Check whether markdown triggers like leading dashes reformat earlier text blocks retroactively.&lt;/li&gt;
&lt;li&gt;Test whether modal dialogs accept keyboard Enter or need explicit mouse clicks on action buttons.&lt;/li&gt;
&lt;li&gt;Test heading shortcut pacing to make sure rapid shortcut events don't get dropped.&lt;/li&gt;
&lt;li&gt;Test file inputs directly to see if scripted file uploads crash the client app.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Disclosure: we build &lt;a href="https://seadanse.com/" rel="noopener noreferrer"&gt;Seadanse&lt;/a&gt;, an AI video tool, and the page we were publishing is a model comparison on a site of ours.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>automation</category>
      <category>javascript</category>
      <category>programming</category>
    </item>
    <item>
      <title>Seedance 2.5 API: The Official Endpoint and 6 Gotchas</title>
      <dc:creator>WesLin</dc:creator>
      <pubDate>Mon, 24 Aug 2026 08:43:04 +0000</pubDate>
      <link>https://dev.to/codesugar_lin_037a57b06a4/seedance-25-api-the-official-endpoint-and-6-gotchas-a22</link>
      <guid>https://dev.to/codesugar_lin_037a57b06a4/seedance-25-api-the-official-endpoint-and-6-gotchas-a22</guid>
      <description>&lt;p&gt;Three sites sell access to "the Seedance 2.5 API". Only one of them belongs to ByteDance. Here's the official endpoint, the model ID, and the six things that cost us time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The official endpoint
&lt;/h2&gt;

&lt;p&gt;ByteDance hosts the official API on BytePlus ModelArk. If you want to know &lt;a href="https://seadanse.com/blog/seedance-2-5-official-website" rel="noopener noreferrer"&gt;which sites are actually ByteDance&lt;/a&gt;, check the domain on the endpoint before you write any code. Everything else selling a "Seedance API" is a reseller sitting in front of this.&lt;/p&gt;

&lt;p&gt;You need an authorization bearer token and a JSON body to start a task:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://ark.ap-southeast.bytepluses.com/api/v3/contents/generations/tasks &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$ARK_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "dreamina-seedance-2-5-260628",
    "content": [
      { "type": "text", "text": "A slow push-in on a red door at the end of a corridor, matching the light in @Image1." },
      { "type": "image_url", "image_url": { "url": "https://example.com/door.png" }, "role": "reference_image" }
    ],
    "generate_audio": true,
    "ratio": "16:9",
    "duration": 15
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model ID for Seedance 2.5 is &lt;code&gt;dreamina-seedance-2-5-260628&lt;/code&gt;. If you pass the older Seedance 2.0 ID instead, you'll need &lt;code&gt;dreamina-seedance-2-0-260128&lt;/code&gt;. You can confirm both IDs and the route directly in ByteDance's ModelArk docs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The request body
&lt;/h2&gt;

&lt;p&gt;The API takes a single JSON payload. Use these exact field names:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;model&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;required&lt;/td&gt;
&lt;td&gt;the model id&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;content&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;object[]&lt;/td&gt;
&lt;td&gt;required&lt;/td&gt;
&lt;td&gt;the input list: text, images, videos&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;omni_reference_task_type&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;&lt;code&gt;auto&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;task-type hint; marked new in the docs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;resolution&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;output resolution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ratio&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;aspect ratio&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;duration&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;integer&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;frames&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;integer&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;frame count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;generate_audio&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;td&gt;&lt;code&gt;true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;makes the sound in the same pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;watermark&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;td&gt;&lt;code&gt;false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;output_format&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;&lt;code&gt;mp4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;marked new in the docs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;seed&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;integer&lt;/td&gt;
&lt;td&gt;&lt;code&gt;-1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;camera_fixed&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;td&gt;&lt;code&gt;false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;return_last_frame&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;td&gt;&lt;code&gt;false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;draft&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;td&gt;&lt;code&gt;false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;service_tier&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;&lt;code&gt;default&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;callback_url&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;execution_expires_after&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;integer&lt;/td&gt;
&lt;td&gt;&lt;code&gt;172800&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;seconds, so 48 hours&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;priority&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;integer&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;safety_identifier&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;string&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;end-user identifier&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three defaults in this schema will surprise you if you don't check them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;generate_audio&lt;/code&gt; defaults to &lt;code&gt;true&lt;/code&gt; — you'll generate sound tracks by default even if you only wanted silent video clips.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;execution_expires_after&lt;/code&gt; defaults to &lt;code&gt;172800&lt;/code&gt; — the docs call it the task expiration threshold, so that's 48 hours.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;omni_reference_task_type&lt;/code&gt; defaults to &lt;code&gt;auto&lt;/code&gt; — the system guesses your reference mode unless you tell it what task you're running.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Gotcha 1: your references are bound inside the prompt string
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;content&lt;/code&gt; array holds text, reference images, and reference videos. The model takes up to 30 images, 10 video clips, and 10 audio clips in one call.&lt;/p&gt;

&lt;p&gt;Instead of passing separate fields for character shots or motion guides, you point to your files directly inside the prompt string using &lt;code&gt;@Image1&lt;/code&gt;, &lt;code&gt;@Video1&lt;/code&gt;, and &lt;code&gt;@Video2&lt;/code&gt;. ByteDance's own example prompt contains: "The strawberry flavor refers to @Image1" and "referring to the composition of @Video1" and "referring to the impact of @Video3".&lt;/p&gt;

&lt;p&gt;The number after &lt;code&gt;@Video&lt;/code&gt; or &lt;code&gt;@Image&lt;/code&gt; counts against the order of items inside &lt;code&gt;content&lt;/code&gt;. Nothing in the parameter table says this; the doc's example is the only place it appears. If you reorder items in &lt;code&gt;content&lt;/code&gt;, &lt;code&gt;@Video1&lt;/code&gt; silently binds to a completely different clip.&lt;/p&gt;

&lt;p&gt;Here's the minimal shape for a text prompt paired with an image reference:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"A slow push-in on a red door at the end of a corridor, matching the light in @Image1."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"image_url"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"image_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com/door.png"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"reference_image"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Gotcha 2: the result URL dies in 24 hours
&lt;/h2&gt;

&lt;p&gt;When a job finishes, the API returns a presigned URL (a link that carries its own expiry and signature in the query string) in &lt;code&gt;content.video_url&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Look at the query parameters on that URL and you'll find &lt;code&gt;X-Tos-Expires=86400&lt;/code&gt;. That's 86,400 seconds — so the link is dead in 24 hours, and anything you didn't copy is gone. If you save the URL in your database instead of copying the file, every video in your product's library is a dead link the next day, and nothing in the API tells you it happened. We copy every finished file to our own storage for exactly that reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gotcha 3: audio is on unless you turn it off
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;generate_audio&lt;/code&gt; setting defaults to &lt;code&gt;true&lt;/code&gt;. That means the model makes synchronized sound in the same pass as the video frames.&lt;/p&gt;

&lt;p&gt;So you'll get generated audio tracks even when your application only asked for background video. If your product adds its own soundtrack or needs silent video, you'll need to pass &lt;code&gt;"generate_audio": false&lt;/code&gt; in your request body. If you leave the field unset, it won't be silent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gotcha 4: there is no 4K tier on 2.5
&lt;/h2&gt;

&lt;p&gt;Don't build a 4K option into your UI for Seedance 2.5. The official ModelArk rate table lists only a 480p/720p tier and a 1080p tier for &lt;code&gt;dreamina-seedance-2-5-260628&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;4K exists on Seedance 2.0 (&lt;code&gt;dreamina-seedance-2-0-260128&lt;/code&gt;), but it isn't available on 2.5. Several platform pages advertise 4K for 2.5 anyway. For example, a page titled "Seedance 2.5 API Now Available - 30s 4K AI Video on Kie.ai" mentions 4K in its title, but its own rate card lists only 480P, 720P, and 1080P.&lt;/p&gt;

&lt;p&gt;Set your resolution to 480p, 720p, or 1080p. We haven't tested what a 4K request does on 2.5 — there's no tier to bill it against, so don't find out in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gotcha 5: it is async, and the only thing you get back is an id
&lt;/h2&gt;

&lt;p&gt;Task creation is strictly asynchronous. The API won't keep an HTTP connection open while the video finishes.&lt;/p&gt;

&lt;p&gt;When you send a creation request, the only thing you get back is a task ID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"cgt-2026******-****"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You have to poll the task endpoint to check the status:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; GET &lt;span class="s2"&gt;"https://ark.ap-southeast.bytepluses.com/api/v3/contents/generations/tasks?page_size=3&amp;amp;filter.status=succeeded"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$ARK_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When a task completes, the response gives you the full state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"cgt-2026******-****"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"dreamina-seedance-2-5-260628"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"succeeded"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"video_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://ark-content-generation-ap-southeast-1.tos-ap-southeast-1.volces.com/...?X-Tos-Expires=86400&amp;amp;X-Tos-Signature=***"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"usage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"completion_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;108900&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"total_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;108900&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"resolution"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"720p"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"ratio"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"16:9"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"duration"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"framespersecond"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;24&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A finished task object carries several fields, but these are the ones that matter most for your app:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;status&lt;/code&gt; — tells you if the job &lt;code&gt;succeeded&lt;/code&gt;, failed, or is still running.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;content.video_url&lt;/code&gt; — the temporary storage link to the generated MP4 file.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;usage.completion_tokens&lt;/code&gt; — the exact token count you're billed on.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Billing runs per million tokens, and &lt;code&gt;usage.completion_tokens&lt;/code&gt; is what ByteDance charges you for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gotcha 6: first-and-last-frame is a documented mode
&lt;/h2&gt;

&lt;p&gt;The create-task documentation includes tabs for seven distinct modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;multimodal reference (takes text, images, video and audio together)&lt;/li&gt;
&lt;li&gt;edit video&lt;/li&gt;
&lt;li&gt;extend video&lt;/li&gt;
&lt;li&gt;audio video first frame&lt;/li&gt;
&lt;li&gt;audio video first and last frames&lt;/li&gt;
&lt;li&gt;image to video from base64&lt;/li&gt;
&lt;li&gt;text to video&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Worth knowing, because a wrapper only exposes the modes its authors wired up. If you need a clean transition between two states, first-and-last-frame is documented right there in ByteDance's own endpoint tabs — check whether the API you picked passes it through.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you would rather not write the integration
&lt;/h2&gt;

&lt;p&gt;If you want to &lt;a href="https://seadanse.com/" rel="noopener noreferrer"&gt;run Seedance 2.5 in the browser&lt;/a&gt; without managing task queues, storage copies, and polling workers, you can use Seadanse. It takes reference images and reference clips for 5, 10, 15, 20, 25, or 30-second clips at 480p, 720p, or 1080p. It doesn't support reference audio, it has no video extension mode, and it caps at 1080p and 30 seconds. You can check &lt;a href="https://seadanse.com/pricing" rel="noopener noreferrer"&gt;what a clip costs&lt;/a&gt; directly on our pricing page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where these facts came from
&lt;/h2&gt;

&lt;p&gt;Every parameter, endpoint path, and model name here comes from ByteDance's official documentation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://docs.byteplus.com/en/docs/ModelArk/1520757" rel="noopener noreferrer"&gt;BytePlus ModelArk Task Creation&lt;/a&gt;, read 2026-08-21, page last updated August 18, 2026.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.byteplus.com/en/docs/ModelArk/1521675" rel="noopener noreferrer"&gt;BytePlus ModelArk Task Polling and Retrieval&lt;/a&gt;, read 2026-08-21.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.byteplus.com/en/docs/ModelArk/1544106" rel="noopener noreferrer"&gt;BytePlus ModelArk Pricing and Model IDs&lt;/a&gt;, read 2026-08-21.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5" rel="noopener noreferrer"&gt;ByteDance Seedance 2.5 Announcement&lt;/a&gt;, read 2026-08-21.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>video</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Is Seedance 2.5 free? I read every pricing page so you don't have to</title>
      <dc:creator>WesLin</dc:creator>
      <pubDate>Fri, 21 Aug 2026 14:05:52 +0000</pubDate>
      <link>https://dev.to/codesugar_lin_037a57b06a4/is-seedance-25-free-i-read-every-pricing-page-so-you-dont-have-to-4i1f</link>
      <guid>https://dev.to/codesugar_lin_037a57b06a4/is-seedance-25-free-i-read-every-pricing-page-so-you-dont-have-to-4i1f</guid>
      <description>&lt;p&gt;Landing pages for video models don't always match their buttons. When you look for Seedance 2.5, you'll find plenty of search results that claim free access. Most pricing pages treat "free" as a headline rather than a clear spec.&lt;/p&gt;

&lt;p&gt;If you build things, you read pricing pages like API docs. You look for inputs, outputs, and hard limits. You don't want marketing copy when you're planning a build.&lt;/p&gt;

&lt;p&gt;To see what's actually real, I checked every platform running Seedance on August 21, 2026 without an account. For each site, I treated the free tier as a spec with three fields: what the page claims, what it publishes as a number, and what you can make before giving up a card. When a number is missing, that missing number is the finding.&lt;/p&gt;

&lt;h2&gt;
  
  
  The state of free Seedance access
&lt;/h2&gt;

&lt;p&gt;ByteDance built the Seedance family of models. Its own consumer tool is Dreamina (&lt;a href="https://dreamina.capcut.com/seedance/seedance-2-5" rel="noopener noreferrer"&gt;Dreamina&lt;/a&gt;). Dreamina's page is headed "Free Seedance 2.5 AI Video Generator with Audio". Its body says "Create cinematic 4K videos with Seedance 2.5 unlimited in Dreamina".&lt;/p&gt;

&lt;p&gt;That same Dreamina page publishes no free-credit number anywhere. The tool behind it opens on AI image without an account. The video model list and any credit balance sit behind a sign-in box. You can't verify the grant without making an account first.&lt;/p&gt;

&lt;p&gt;Other sites show similar gaps when you open them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Seedance Studio publishes clear terms on its page (&lt;a href="https://tryseedance.ai/seedance-2-0-free" rel="noopener noreferrer"&gt;Seedance Studio&lt;/a&gt;). It lists a free plan of "160 credits/mo" and "Up to 3 videos or 20 images to try it out" at 480p with no card needed. In its own words, "Free videos use the same Seedance 2.5 model as paid plans, same quality, native sound and references. Not a stripped-down demo."&lt;/li&gt;
&lt;li&gt;seedance.tv lists Seedance 2.5 in its model menu (&lt;a href="https://www.seedance.tv/" rel="noopener noreferrer"&gt;seedance.tv&lt;/a&gt;). Its footer says "This platform is an independent product and is not affiliated with Bytedance." We checked its flow on August 13, 2026 and found its sign-up grant lands short of any Seedance 2.5 clip. You won't see its pack prices until you log in.&lt;/li&gt;
&lt;li&gt;EaseMate AI titles its page "Create 30-Second HD Videos with Seedance 2.5 Free Online" (&lt;a href="https://www.easemate.ai/seedance-2-5-ai-video-generator" rel="noopener noreferrer"&gt;EaseMate AI&lt;/a&gt;). The on-page tool lists "Seedance 2.5 — Live Now". But its pricing page publishes no video allowance for the free tier (&lt;a href="https://www.easemate.ai/pricing" rel="noopener noreferrer"&gt;EaseMate AI Pricing&lt;/a&gt;). It says "Limited credits for images/videos" and gives 30 free credits on sign-up.&lt;/li&gt;
&lt;li&gt;aiimagetovideo.pro titles its page "Unlimited Free Seedance 2.5 AI Video Generator by ByteDance (No Sign Up)" (&lt;a href="https://aiimagetovideo.pro/seedance-2-5-free-ai-video-generator" rel="noopener noreferrer"&gt;aiimagetovideo.pro&lt;/a&gt;). A banner reads "Free Users Get 2 Daily Videos for AI Video Generator. Subscribers Get Priority Access." The generator on the page offers only one model, "Video Fast 1.0", at 3 or 5 seconds and 480p or 720p. It doesn't offer Seedance at all.&lt;/li&gt;
&lt;li&gt;Luma runs Seedance 2.5 on its platform (&lt;a href="https://lumalabs.ai/pricing" rel="noopener noreferrer"&gt;Luma Labs&lt;/a&gt;). Its Plans &amp;amp; Pricing page lists only paid individual plans. It publishes no free plan at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's how each platform stacks up when you look at it without an account:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Platform&lt;/th&gt;
&lt;th&gt;What a visitor with no account gets&lt;/th&gt;
&lt;th&gt;Free Seedance 2.5?&lt;/th&gt;
&lt;th&gt;What it means for you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Dreamina&lt;/td&gt;
&lt;td&gt;Page says free and unlimited; the tool opens on images and the numbers sit behind sign-in&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;You have to sign in before you can find out what you were promised&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seedance Studio&lt;/td&gt;
&lt;td&gt;160 credits a month, up to 3 videos, 480p, no card&lt;/td&gt;
&lt;td&gt;Yes, capped&lt;/td&gt;
&lt;td&gt;The clearest free route, if three 480p clips a month is enough&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;seedance.tv&lt;/td&gt;
&lt;td&gt;Sign-up grant that lands short of one 2.5 render; pack prices after login&lt;/td&gt;
&lt;td&gt;Not on the grant alone&lt;/td&gt;
&lt;td&gt;You can look at 2.5 but you cannot finish a clip for nothing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EaseMate AI&lt;/td&gt;
&lt;td&gt;30 credits on registration; the free video allowance is not published&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;The model is offered, the allowance is not written down&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;aiimagetovideo.pro&lt;/td&gt;
&lt;td&gt;A generator with one model, Video Fast 1.0, and 2 videos a day&lt;/td&gt;
&lt;td&gt;No — the page does not run Seedance&lt;/td&gt;
&lt;td&gt;The page ranks for the phrase; the tool answers a different one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seadanse&lt;/td&gt;
&lt;td&gt;10 stills a day before sign-in, then 60 credits on sign-up with no card&lt;/td&gt;
&lt;td&gt;Not on the grant alone&lt;/td&gt;
&lt;td&gt;Free covers the rehearsal; you pay once for the take you keep&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sources: &lt;a href="https://dreamina.capcut.com/seedance/seedance-2-5" rel="noopener noreferrer"&gt;Dreamina&lt;/a&gt;, &lt;a href="https://tryseedance.ai/seedance-2-0-free" rel="noopener noreferrer"&gt;Seedance Studio&lt;/a&gt;, &lt;a href="https://www.seedance.tv/" rel="noopener noreferrer"&gt;seedance.tv&lt;/a&gt;, &lt;a href="https://www.easemate.ai/seedance-2-5-ai-video-generator" rel="noopener noreferrer"&gt;EaseMate AI&lt;/a&gt;, &lt;a href="https://aiimagetovideo.pro/seedance-2-5-free-ai-video-generator" rel="noopener noreferrer"&gt;aiimagetovideo.pro&lt;/a&gt;, &lt;a href="https://seadanse.com/" rel="noopener noreferrer"&gt;Seadanse&lt;/a&gt;. Filled on August 21, 2026 by opening each page without an account.&lt;/p&gt;

&lt;h2&gt;
  
  
  The credit arithmetic on Seadanse
&lt;/h2&gt;

&lt;p&gt;On &lt;a href="https://seadanse.com/" rel="noopener noreferrer"&gt;Seadanse&lt;/a&gt;, you can test the setup before you spend anything. You don't have to guess what things cost.&lt;/p&gt;

&lt;p&gt;Before signing in, you can make reference images. The composer's first step has a button labelled "Create visual". It works with no sign-up and lets you make up to 10 stills a day. Signed-in visitors also get 10 a day.&lt;/p&gt;

&lt;p&gt;When you sign up with an email, you get 60 credits once. You don't need a credit card. &lt;/p&gt;

&lt;p&gt;Credits are the site's own unit. Every video costs a set number of them. Here's the credit arithmetic for a new $0 account:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Sign-up grant (one-time, no card):           60 credits
Cost of one Mini clip (480p, 5 seconds):     60 credits
Balance after one Mini clip:                  0 credits

Cost of one 2.5 clip (480p, 5 seconds):     125 credits
Shortfall for Seedance 2.5 clip:             65 credits (60 - 125)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The composer opens on a 480p, 5-second Mini clip by default. That means you can press Generate once without changing any setting.&lt;/p&gt;

&lt;p&gt;You'll always see the exact cost on the Generate button before you click it. If a run fails, the platform returns your credits automatically. You can check the ratios on &lt;a href="https://seadanse.com/pricing" rel="noopener noreferrer"&gt;the credit calculator&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why debugging on Mini saves credits
&lt;/h2&gt;

&lt;p&gt;You shouldn't debug prompts on an expensive checkpoint. &lt;a href="https://seadanse.com/models/seedance-2-0-mini" rel="noopener noreferrer"&gt;Seedance 2.0 Mini&lt;/a&gt; and Seedance 2.5 come from the same model family. They take the same kind of written prompt.&lt;/p&gt;

&lt;p&gt;Because both checkpoints share syntax, a prompt that reads correctly on Mini reads correctly on 2.5. You can read &lt;a href="https://seadanse.com/blog/seedance-2-5-prompt-guide" rel="noopener noreferrer"&gt;how to write a Seedance 2.5 prompt&lt;/a&gt; and run your tests on Mini first. &lt;/p&gt;

&lt;p&gt;Mini runs at 480p and 720p. A 480p clip on Mini for 5 seconds costs less than half of what 2.5 costs at the same size and length (60 credits against 125).&lt;/p&gt;

&lt;p&gt;That buys you a cheap rehearsal loop. You debug camera movement, framing, and action on the cheap checkpoint. Once the shot works, you switch the dropdown to 2.5 and pay once for the final take.&lt;/p&gt;

&lt;p&gt;This ratio isn't unique to one site. Luma's published cost table charges Seedance 2.0 Mini at well under half the Seedance 2.5 rate at the same resolution (&lt;a href="https://lumalabs.ai/pricing" rel="noopener noreferrer"&gt;Luma Labs&lt;/a&gt;). The same ratio shows up across hosts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Seedance 2.5 runs in production
&lt;/h2&gt;

&lt;p&gt;When you're ready for final output, the &lt;a href="https://seadanse.com/" rel="noopener noreferrer"&gt;Seedance 2.5 AI video maker&lt;/a&gt; exposes the full model spec. You can check &lt;a href="https://seadanse.com/blog/what-is-seedance-2-5" rel="noopener noreferrer"&gt;what Seedance 2.5 changed&lt;/a&gt; before setting up your run.&lt;/p&gt;

&lt;p&gt;When you run &lt;a href="https://seadanse.com/" rel="noopener noreferrer"&gt;Seedance 2.5 on Seadanse&lt;/a&gt;, you get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Resolutions of 480p, 720p, and 1080p (there is no 4K on 2.5; 4K belongs to the Seedance 2.0 preset).&lt;/li&gt;
&lt;li&gt;Lengths of 5, 10, 15, 20, 25, or 30 seconds.&lt;/li&gt;
&lt;li&gt;Six aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, and 21:9.&lt;/li&gt;
&lt;li&gt;Sound made with the picture, on by default.&lt;/li&gt;
&lt;li&gt;A prompt box that takes up to 10,000 characters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cost scales with duration. A 30-second clip costs six times a 5-second one at the same settings. If you need more credits, you can buy a $9.9 pack that does not need a subscription. Credits bought in a pack stay valid for 12 months.&lt;/p&gt;

&lt;p&gt;To test your prompt on the default settings, &lt;a href="https://seadanse.com/#generate" rel="noopener noreferrer"&gt;open the composer&lt;/a&gt; and run your first rehearsal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist for evaluating video generation pricing pages
&lt;/h2&gt;

&lt;p&gt;When you land on any page that promises free model access, run through these four checks:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Find the exact credit cost&lt;/strong&gt;: Look for published numbers that show the credit price per clip length and resolution. If the number isn't there, you can't price your pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check the actual model in the composer&lt;/strong&gt;: Check whether the tool loads the model from the headline or swaps in an unrelated fast tier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare the free grant to clip costs&lt;/strong&gt;: Subtract the cost of one baseline clip from the sign-up grant. If the balance is negative, you can't finish a clip for free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check for standalone packs&lt;/strong&gt;: See if the site sells simple packs like a $9.9 pack, or if it locks all usage behind a monthly subscription.&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I work on Seadanse, one of the platforms in this comparison. Everything said about the other platforms comes from their own public pages, linked above.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>I Built an Open-Source macOS Screen-Cleaning &amp; Pet-Lock Utility in 3 Hours with AI</title>
      <dc:creator>WesLin</dc:creator>
      <pubDate>Tue, 11 Aug 2026 05:16:09 +0000</pubDate>
      <link>https://dev.to/codesugar_lin_037a57b06a4/i-built-an-open-source-macos-screen-cleaning-pet-lock-utility-in-3-hours-with-ai-k42</link>
      <guid>https://dev.to/codesugar_lin_037a57b06a4/i-built-an-open-source-macos-screen-cleaning-pet-lock-utility-in-3-hours-with-ai-k42</guid>
      <description>&lt;h2&gt;
  
  
  The Problem Nobody Talks About
&lt;/h2&gt;

&lt;p&gt;Here's a dumb problem that bugged me for years: &lt;strong&gt;I couldn't clean my MacBook screen properly.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every time I grabbed a cloth to wipe the display, I'd accidentally brush the keyboard or trackpad. The screen would wake up, show my wallpaper, and suddenly I couldn't tell where the smudges were anymore. I'd end up "blind-wiping" and hoping for the best. Shutting down the whole machine just to wipe a screen felt absurd.&lt;/p&gt;

&lt;p&gt;Then there was the cat problem. I saw complaints in App Store reviews: people watching videos on their Macs, and their cats would walk across the keyboard — pausing the video, switching tabs, opening Spotlight. Classic cat behavior, zero good solutions.&lt;/p&gt;

&lt;p&gt;So I built &lt;a href="https://github.com/opensource-works/CleanMyMac" rel="noopener noreferrer"&gt;&lt;strong&gt;CleanMyScreen&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What It Does
&lt;/h2&gt;

&lt;p&gt;Three modes, one tiny app:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;What Stays Visible&lt;/th&gt;
&lt;th&gt;What Gets Locked&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cleaning&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Pure-black overlay on all displays&lt;/td&gt;
&lt;td&gt;Keyboard + pointer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pet / Kid&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Your current video or app&lt;/td&gt;
&lt;td&gt;Keyboard + trackpad&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Selective&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Your desktop&lt;/td&gt;
&lt;td&gt;Any combination you choose&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every session has a 3-second countdown before locking, an optional auto-timeout, and a 3-second Escape hold as an emergency unlock.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Stack
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SwiftUI&lt;/strong&gt; for the UI&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AppKit&lt;/strong&gt; for window management and overlays&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Core Graphics&lt;/strong&gt; for display brightness control&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IOKit&lt;/strong&gt; for low-level HID device seizing (this is how you actually lock a specific trackpad vs. keyboard at the hardware level)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No Electron. No web views. No accounts. No telemetry. No network calls at all. The app literally cannot phone home.&lt;/p&gt;

&lt;h2&gt;
  
  
  Built in 3 Hours with Codex
&lt;/h2&gt;

&lt;p&gt;Here's the part that still blows my mind.&lt;/p&gt;

&lt;p&gt;I'd had this idea sitting in my backlog for over a year. It wasn't complicated conceptually, but between work and other projects, I never prioritized it — partly because the monetization path was unclear for something this simple.&lt;/p&gt;

&lt;p&gt;Then I had some Codex credits about to reset, so I thought: why not just ship it?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The entire functional app was built in under 3 hours.&lt;/strong&gt; Codex handled the SwiftUI views, the IOKit device enumeration, the overlay window management, and the brightness control logic. I spent the remaining time fine-tuning UX details — countdown animations, the emergency unlock flow, making sure device seizing was properly cleaned up on session end.&lt;/p&gt;

&lt;p&gt;This is exactly the kind of project where AI coding tools shine: the requirements are clear, the scope is tight, and the individual pieces (IOKit APIs, CGDisplay functions) are well-documented but tedious to wire together manually.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try It / Contribute
&lt;/h2&gt;

&lt;p&gt;CleanMyScreen is MIT-licensed and works on both Apple Silicon and Intel Macs.&lt;/p&gt;

&lt;p&gt;🔗 &lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/opensource-works/CleanMyMac" rel="noopener noreferrer"&gt;github.com/opensource-works/CleanMyMac&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can build from source with just the Swift command-line tools — no Xcode project needed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;./Scripts/run-tests.sh
./Scripts/build-app.sh
open dist/CleanMyScreen.app
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you have ideas, bugs, or feature requests, open an issue. I'd especially love input on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;External display brightness control (it's best-effort right now due to macOS API fragmentation)&lt;/li&gt;
&lt;li&gt;Additional use cases for the Selective mode&lt;/li&gt;
&lt;li&gt;Localization — the app is English-only for now&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>swift</category>
    </item>
    <item>
      <title>htdemucs vs BS-RoFormer vs Spleeter: A 2026 Audio Source Separation Benchmark</title>
      <dc:creator>WesLin</dc:creator>
      <pubDate>Wed, 29 Apr 2026 16:25:01 +0000</pubDate>
      <link>https://dev.to/codesugar_lin_037a57b06a4/htdemucs-vs-bs-roformer-vs-spleeter-a-2026-audio-source-separation-benchmark-2ll8</link>
      <guid>https://dev.to/codesugar_lin_037a57b06a4/htdemucs-vs-bs-roformer-vs-spleeter-a-2026-audio-source-separation-benchmark-2ll8</guid>
      <description>&lt;p&gt;If you've spent any time looking at AI music separation in the last twelve months, you've probably run into the same three names: &lt;strong&gt;Spleeter&lt;/strong&gt;, &lt;strong&gt;htdemucs&lt;/strong&gt; (Hybrid Transformer Demucs), and &lt;strong&gt;BS-RoFormer&lt;/strong&gt;. They show up in every comparison post, every research paper, and every "how to extract vocals" tutorial — but the way they're compared is usually wrong. Most posts cite a single SDR number from a 2019 paper and call it a day.&lt;/p&gt;

&lt;p&gt;That's not useful if you're trying to ship a product, build a pipeline, or pick a model for real audio.&lt;/p&gt;

&lt;p&gt;This post compares the three on the dimensions that actually matter when you're deploying audio separation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Quality&lt;/strong&gt; — SDR scores from peer-reviewed sources, not vibes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inference speed&lt;/strong&gt; — what you'll actually wait for in production&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost per song&lt;/strong&gt; — running on commodity GPUs at 2026 prices&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output flexibility&lt;/strong&gt; — 2 stems vs 4 stems vs 6 stems&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When each one is the right choice&lt;/strong&gt; — and when it isn't&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Everything below is based on published benchmarks plus our own production deployment of htdemucs at scale. Where we cite numbers, we cite the source.&lt;/p&gt;




&lt;h2&gt;
  
  
  TL;DR (for people who want the answer now)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;th&gt;Output stems&lt;/th&gt;
&lt;th&gt;Quality (avg SDR)&lt;/th&gt;
&lt;th&gt;Speed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spleeter&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Real-time, low-resource, batch processing&lt;/td&gt;
&lt;td&gt;2, 4, or 5&lt;/td&gt;
&lt;td&gt;~5.9 dB (vocals)&lt;/td&gt;
&lt;td&gt;~100× real-time on GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;htdemucs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Production C2C apps, balance of quality and speed&lt;/td&gt;
&lt;td&gt;4 or 6&lt;/td&gt;
&lt;td&gt;~9.0 dB (avg)&lt;/td&gt;
&lt;td&gt;~5–8× real-time on A40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;BS-RoFormer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Highest-fidelity offline work, mastering, archival&lt;/td&gt;
&lt;td&gt;4 (typically)&lt;/td&gt;
&lt;td&gt;~9.80 dB (avg)&lt;/td&gt;
&lt;td&gt;~2–3× real-time on A40&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you take only one thing from this post: &lt;strong&gt;htdemucs is the right default for almost any product, and you should probably be running &lt;code&gt;htdemucs_ft&lt;/code&gt; rather than the default checkpoint.&lt;/strong&gt; On Replicate's serverless pricing, all three Demucs variants (default, 6s, ft) cost essentially the same per call — but ft delivers meaningfully better separation. We didn't expect this when we started; it only became clear after looking at our actual billing.&lt;/p&gt;

&lt;p&gt;BS-RoFormer is meaningfully better only on bass and only when latency doesn't matter. Spleeter is a 2019 model running on 2026 hardware — fast, but the quality gap is now audible.&lt;/p&gt;

&lt;p&gt;The rest of this post explains why.&lt;/p&gt;




&lt;h2&gt;
  
  
  What we mean by "quality" — SDR explained briefly
&lt;/h2&gt;

&lt;p&gt;Music source separation quality is usually measured in &lt;strong&gt;Signal-to-Distortion Ratio (SDR)&lt;/strong&gt;, in decibels. Higher is better. The reference dataset is &lt;strong&gt;MUSDB18&lt;/strong&gt; (or MUSDB18-HQ for high-quality audio), which contains 150 full-length tracks with isolated stems for vocals, drums, bass, and "other."&lt;/p&gt;

&lt;p&gt;A few practical anchors:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&amp;lt;6 dB SDR&lt;/strong&gt;: noticeable artifacts, "phasey" vocals, audible bleed between stems&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;6–8 dB SDR&lt;/strong&gt;: usable for casual purposes (karaoke, learning songs, sketching ideas)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;8–10 dB SDR&lt;/strong&gt;: clean enough for content creation and most DJ applications&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&amp;gt;10 dB SDR&lt;/strong&gt;: approaching transparent for the average listener; suitable for release-quality work after light cleanup&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Anything above ~9 dB on vocals is generally past the point where most listeners can tell the difference in a blind test. The gains from there are about edge cases — heavy reverb, doubled vocals, complex mixes.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A note on SI-SDR&lt;/strong&gt;: Some recent papers report SI-SDR (scale-invariant SDR), which corrects for simple gain differences and is more robust. When numbers in this post differ from other sources, the metric definition is usually the reason.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The three models, briefly
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Spleeter (Deezer, 2019)
&lt;/h3&gt;

&lt;p&gt;Released by the Deezer research team in 2019, Spleeter is a U-Net architecture operating in the spectrogram domain. It comes in 2-stem (vocals/accompaniment), 4-stem (vocals/drums/bass/other), and 5-stem (adds piano) configurations.&lt;/p&gt;

&lt;p&gt;It was a landmark release at the time — the first time anyone could run good-enough source separation on a laptop CPU without licensing fees. Six years later, it's been overtaken on quality by every modern model, but it remains the fastest and lightest option by a wide margin.&lt;/p&gt;

&lt;h3&gt;
  
  
  htdemucs (Meta AI, 2022)
&lt;/h3&gt;

&lt;p&gt;The fourth-generation Demucs model from Meta AI's research team. Unlike Spleeter, htdemucs is a &lt;strong&gt;hybrid&lt;/strong&gt; model — it operates in both the time domain (waveform) and frequency domain (spectrogram), with a Transformer backbone connecting them. The original paper &lt;a href="https://arxiv.org/pdf/2111.03600" rel="noopener noreferrer"&gt;reports a 1.4 dB SDR improvement&lt;/a&gt; over the previous Demucs generation on MUSDB-HQ.&lt;/p&gt;

&lt;p&gt;Two variants matter in practice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;htdemucs&lt;/code&gt;&lt;/strong&gt; — the standard 4-stem model&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;htdemucs_6s&lt;/code&gt;&lt;/strong&gt; — a 6-stem variant that adds isolated guitar and piano stems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There's also &lt;code&gt;htdemucs_ft&lt;/code&gt;, a fine-tuned version that's slower but slightly more accurate on individual stems.&lt;/p&gt;

&lt;p&gt;htdemucs placed competitively in the 2021 Sony Music Demixing Challenge and remains the default for most production pipelines that aren't chasing the absolute SOTA.&lt;/p&gt;

&lt;h3&gt;
  
  
  BS-RoFormer (2023)
&lt;/h3&gt;

&lt;p&gt;The current state of the art on MUSDB18-HQ, BS-RoFormer (Band-Split RoPE Transformer) is a pure-Transformer architecture that replaces RNN modules with a hierarchical RoPE Transformer. It splits the input spectrogram into multiple non-overlapping frequency sub-bands, exploiting the fact that different instruments occupy characteristic frequency ranges (bass low, cymbals high, etc.).&lt;/p&gt;

&lt;p&gt;BS-RoFormer trained on MUSDB18-HQ plus 500 extra songs &lt;strong&gt;won first place in the Music Source Separation track of the Sound Demixing Challenge 2023 (SDX23)&lt;/strong&gt;. Even the smaller version trained without extra data &lt;a href="https://www.researchgate.net/publication/370763841_Benchmarks_and_leaderboards_for_sound_demixing_tasks" rel="noopener noreferrer"&gt;reports 9.80 dB average SDR on MUSDB18-HQ&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The downside: it's slower and more memory-intensive than htdemucs, and the production-ready open weights are still scattered across community implementations rather than a single canonical release.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Quality benchmark (published SDR scores)
&lt;/h2&gt;

&lt;p&gt;This is where most comparison posts fall apart — they cherry-pick a single number. Here are the per-stem SDR scores from the published literature, on MUSDB18-HQ (no extra training data unless noted):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Vocals&lt;/th&gt;
&lt;th&gt;Drums&lt;/th&gt;
&lt;th&gt;Bass&lt;/th&gt;
&lt;th&gt;Other&lt;/th&gt;
&lt;th&gt;Average&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Spleeter (4-stem)&lt;/td&gt;
&lt;td&gt;~5.9 dB&lt;/td&gt;
&lt;td&gt;~5.9 dB&lt;/td&gt;
&lt;td&gt;~5.5 dB&lt;/td&gt;
&lt;td&gt;~4.5 dB&lt;/td&gt;
&lt;td&gt;~5.4 dB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;htdemucs (default)&lt;/td&gt;
&lt;td&gt;~8.1 dB&lt;/td&gt;
&lt;td&gt;~8.4 dB&lt;/td&gt;
&lt;td&gt;~8.6 dB&lt;/td&gt;
&lt;td&gt;~5.9 dB&lt;/td&gt;
&lt;td&gt;~7.7 dB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;htdemucs_ft (fine-tuned)&lt;/td&gt;
&lt;td&gt;~8.9 dB&lt;/td&gt;
&lt;td&gt;~9.5 dB&lt;/td&gt;
&lt;td&gt;~9.4 dB&lt;/td&gt;
&lt;td&gt;~6.4 dB&lt;/td&gt;
&lt;td&gt;~8.5 dB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BS-RoFormer (no extra data)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;~11.28 dB&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;~9.80 dB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BS-RoFormer (with 500 extra songs)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;~9.76 dB+&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;: Spleeter scores from the &lt;a href="https://joss.theoj.org/papers/10.21105/joss.02154.pdf" rel="noopener noreferrer"&gt;Spleeter JOSS paper&lt;/a&gt; and the &lt;a href="https://beatstorapon.com/ai-stem-splitter" rel="noopener noreferrer"&gt;BeatsToRapOn separation benchmark&lt;/a&gt;. htdemucs scores from &lt;a href="https://arxiv.org/pdf/2111.03600" rel="noopener noreferrer"&gt;Hybrid Spectrogram and Waveform Source Separation&lt;/a&gt; and &lt;a href="https://www.researchgate.net/publication/370763841_Benchmarks_and_leaderboards_for_sound_demixing_tasks" rel="noopener noreferrer"&gt;Benchmarks and leaderboards for sound demixing tasks&lt;/a&gt;. BS-RoFormer scores from the SDX23 results documented in &lt;a href="https://www.researchgate.net/publication/370763841_Benchmarks_and_leaderboards_for_sound_demixing_tasks" rel="noopener noreferrer"&gt;the same paper&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A few observations from the table:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Spleeter → htdemucs gap is bigger than the htdemucs → BS-RoFormer gap.&lt;/strong&gt; Going from Spleeter to htdemucs gets you roughly +2.3 dB on average. Going from htdemucs to BS-RoFormer gets you roughly +1.3 dB. This is why htdemucs is the practical sweet spot for most use cases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;BS-RoFormer's biggest win is on bass.&lt;/strong&gt; Bass separation jumps from ~8.6 dB (htdemucs) to ~11.28 dB (BS-RoFormer) — a difference you can hear in a blind test. The vocal and drum gains are smaller. If you're building something that specifically needs clean bass (DJ tools, transcription, music education for bass players), BS-RoFormer is worth the extra compute. For everything else, the gain is on the edge of perceptible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;htdemucs_ft is underrated.&lt;/strong&gt; Many comparison posts only test the default &lt;code&gt;htdemucs&lt;/code&gt; checkpoint. The fine-tuned version (&lt;code&gt;htdemucs_ft&lt;/code&gt;) closes most of the gap to BS-RoFormer at the cost of roughly 4× the inference time — still faster than BS-RoFormer in practice.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Inference speed (real-world, not theoretical)
&lt;/h2&gt;

&lt;p&gt;Approximate end-to-end time for a 3-minute song on a single A40 GPU, measured from API call to download-ready output:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;End-to-end time&lt;/th&gt;
&lt;th&gt;Real-time multiplier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Spleeter (4-stem, GPU)&lt;/td&gt;
&lt;td&gt;~2–5 seconds&lt;/td&gt;
&lt;td&gt;~40–90× real-time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;htdemucs (default, 4-stem)&lt;/td&gt;
&lt;td&gt;~30–45 seconds&lt;/td&gt;
&lt;td&gt;~4–6× real-time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;htdemucs_6s (6-stem)&lt;/td&gt;
&lt;td&gt;~40–60 seconds&lt;/td&gt;
&lt;td&gt;~3–5× real-time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;htdemucs_ft (fine-tuned)&lt;/td&gt;
&lt;td&gt;~90–150 seconds&lt;/td&gt;
&lt;td&gt;~1.2–2× real-time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BS-RoFormer&lt;/td&gt;
&lt;td&gt;~60–120 seconds&lt;/td&gt;
&lt;td&gt;~1.5–3× real-time&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;End-to-end time ≠ pure GPU inference time.&lt;/strong&gt; Public benchmarks usually report just the model forward pass on clean inputs. Real production time includes container cold start (5–30s on serverless), audio I/O (file download, ffmpeg pre-processing), and result upload. Our numbers above are end-to-end on Replicate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spleeter is in a different league for speed.&lt;/strong&gt; It's the only one that runs comfortably faster than real-time on CPU alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;htdemucs's &lt;code&gt;overlap&lt;/code&gt; parameter is a big speed lever.&lt;/strong&gt; The default &lt;code&gt;overlap=0.25&lt;/code&gt; is a reasonable trade-off; setting &lt;code&gt;overlap=0.5&lt;/code&gt; improves quality slightly at ~2× the cost; setting &lt;code&gt;overlap=0&lt;/code&gt; makes it noticeably faster but introduces audible chunking artifacts at segment boundaries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BS-RoFormer's reference implementations vary wildly in speed&lt;/strong&gt; depending on whose checkpoint and inference code you use. Numbers above are for the &lt;a href="https://mvsep.com/en/algorithms" rel="noopener noreferrer"&gt;community-popular MVSep BS-RoFormer SW&lt;/a&gt; build.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're shipping a consumer product where users wait for results, anything slower than ~60 seconds for a 3-minute song starts to hurt conversion in our experience. That keeps htdemucs (default and 6s) inside acceptable territory and pushes htdemucs_ft and BS-RoFormer toward async/queued flows where the user can come back later.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Cost per song (production deployment economics)
&lt;/h2&gt;

&lt;p&gt;This is the section where most online comparisons are completely wrong. Public pricing on Replicate looks straightforward — A40 at $0.000725/second, multiply by inference time, done. In practice, that calculation is off by roughly 2× from your actual bill, and there's a more interesting wrinkle that almost no comparison post mentions.&lt;/p&gt;

&lt;h3&gt;
  
  
  The headline finding from our production deployment
&lt;/h3&gt;

&lt;p&gt;We've been running htdemucs in production at &lt;a href="https://aistemsplitter.org" rel="noopener noreferrer"&gt;aistemsplitter.org&lt;/a&gt; for several months across all three Demucs variants — &lt;code&gt;htdemucs&lt;/code&gt; (default 4-stem), &lt;code&gt;htdemucs_6s&lt;/code&gt; (6-stem), and &lt;code&gt;htdemucs_ft&lt;/code&gt; (fine-tuned). On Replicate's A40 GPU instances, &lt;strong&gt;all three variants cost approximately the same per call in our actual billing&lt;/strong&gt;: roughly &lt;strong&gt;22 calls per $1&lt;/strong&gt;, or about &lt;strong&gt;$0.045 per song&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That's worth pausing on, because it contradicts what you'd expect from the published inference times.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Naive cost (public pricing × inference time)&lt;/th&gt;
&lt;th&gt;Our actual measured cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Spleeter (GPU)&lt;/td&gt;
&lt;td&gt;&amp;lt;$0.002&lt;/td&gt;
&lt;td&gt;&amp;lt;$0.005&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;htdemucs (default)&lt;/td&gt;
&lt;td&gt;~$0.022&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$0.045&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;htdemucs_6s (6-stem)&lt;/td&gt;
&lt;td&gt;~$0.029&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$0.045&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;htdemucs_ft (fine-tuned)&lt;/td&gt;
&lt;td&gt;~$0.11&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$0.045&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BS-RoFormer&lt;/td&gt;
&lt;td&gt;~$0.065&lt;/td&gt;
&lt;td&gt;~$0.06–0.10 (varies)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Why all three Demucs variants converge to the same cost
&lt;/h3&gt;

&lt;p&gt;The naive pricing model assumes you pay only for pure GPU inference time. In reality, every Replicate call also includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Container cold-start time (5–30 seconds when scaling from zero)&lt;/li&gt;
&lt;li&gt;Model weight loading into GPU memory&lt;/li&gt;
&lt;li&gt;Audio file download and ffmpeg pre-processing&lt;/li&gt;
&lt;li&gt;Result encoding and upload back to storage&lt;/li&gt;
&lt;li&gt;A minimum billable duration per call&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These overheads are roughly &lt;strong&gt;fixed costs&lt;/strong&gt; per invocation — they don't scale with how complex your model is. When the GPU forward pass goes from 30 seconds (htdemucs default) to 90 seconds (htdemucs_ft), the &lt;em&gt;additional&lt;/em&gt; compute matters less to the bill than you'd expect, because the per-call overhead is already eating most of the budget.&lt;/p&gt;

&lt;p&gt;The practical implication: &lt;strong&gt;if you're already on the htdemucs platform, there's almost no economic reason not to use the highest-quality variant your latency budget allows.&lt;/strong&gt; If your users will wait 60 seconds, use &lt;code&gt;htdemucs_6s&lt;/code&gt; (6 stems, default speed). If they'll wait 2 minutes, use &lt;code&gt;htdemucs_ft&lt;/code&gt; (fine-tuned, near-BS-RoFormer quality on most stems). The bill is the same.&lt;/p&gt;

&lt;p&gt;This is the opposite of the conclusion you'd reach by reading academic papers and Replicate's posted GPU pricing. It only shows up when you actually look at your bill at the end of the month.&lt;/p&gt;

&lt;h3&gt;
  
  
  Implications for unit economics
&lt;/h3&gt;

&lt;p&gt;If you're modeling unit economics for a stem separation product, plan for &lt;strong&gt;$0.04–$0.05 per song&lt;/strong&gt; as your floor, regardless of which Demucs variant you choose. That sets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Free tier ceiling&lt;/strong&gt; — at 10 free minutes per user (≈3 free songs), you're absorbing roughly $0.13 per signup before any conversion&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Minimum viable credit pack pricing&lt;/strong&gt; — anything below ~$0.10/song retail leaves no margin for Stripe fees, support, and infrastructure overhead&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bulk processing cost&lt;/strong&gt; — at 10,000 songs/month you're looking at ~$450 in pure inference, before storage, bandwidth, and any other infrastructure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two important caveats:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Cold starts dominate at low traffic.&lt;/strong&gt; If your service is processing fewer than a few hundred songs per day, the cold-start overhead becomes proportionally larger. At very low traffic, the actual cost can drift up toward $0.06–$0.07 per song.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Self-hosting only beats this above ~$2k/mo in inference spend.&lt;/strong&gt; Until you have enough sustained traffic to keep a dedicated GPU &amp;gt;40% utilized, serverless GPU is cheaper than RunPod, Vast.ai, or your own colo. We've measured this directly — Replicate stayed cheaper than dedicated infrastructure throughout our launch period.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  4. Output flexibility (stem count and format)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Available stem configurations&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Spleeter&lt;/td&gt;
&lt;td&gt;2, 4, or 5 stems&lt;/td&gt;
&lt;td&gt;5-stem adds piano (separate model)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;htdemucs&lt;/td&gt;
&lt;td&gt;4 or 6 stems&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;htdemucs_6s&lt;/code&gt; adds guitar + piano&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BS-RoFormer&lt;/td&gt;
&lt;td&gt;4 stems (mostly); some 6-stem community builds&lt;/td&gt;
&lt;td&gt;Quality drops on the rarer guitar/piano stems&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;This is where htdemucs_6s genuinely stands alone.&lt;/strong&gt; If your use case requires isolated guitar or piano stems (music education, multi-track remixing, transcription), htdemucs_6s is the only widely-deployed model that delivers them at production quality. BS-RoFormer 6-stem variants exist in the community but are less mature; the canonical BS-RoFormer is a 4-stem system.&lt;/p&gt;

&lt;p&gt;For "vocals only" or "instrumental only" use cases (the karaoke crowd), all three models work fine, and you should pick on speed, not quality. Spleeter at 90× real-time will give you a usable instrumental in milliseconds.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. When to pick which one
&lt;/h2&gt;

&lt;p&gt;After running these in production for several months, here's the simple decision tree we'd give someone starting from scratch:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pick Spleeter when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You need to process audio in real-time or near-real-time&lt;/li&gt;
&lt;li&gt;You're running on CPU or constrained hardware&lt;/li&gt;
&lt;li&gt;You need batch-processing throughput (e.g., feature extraction over a music catalog)&lt;/li&gt;
&lt;li&gt;The quality bar is "usable" not "good"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Pick htdemucs when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're building a consumer-facing product where users wait &amp;lt;60 seconds&lt;/li&gt;
&lt;li&gt;You need 6 stems (use &lt;code&gt;htdemucs_6s&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;You want the best quality-per-dollar ratio in production&lt;/li&gt;
&lt;li&gt;You don't want to maintain custom inference code (it's well-supported on every major model-serving platform)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Pick BS-RoFormer when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're running offline or batch jobs where 1–2 minutes per song is fine&lt;/li&gt;
&lt;li&gt;Bass quality specifically matters (DJ tools, transcription, audio analysis)&lt;/li&gt;
&lt;li&gt;You're producing release-quality work and the marginal SDR matters&lt;/li&gt;
&lt;li&gt;You're willing to invest engineering time in keeping up with community model releases&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Don't pick any of these when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You only need vocal removal for karaoke. Use Spleeter 2-stem; the quality difference doesn't matter for sing-along audio that's going to be played over a microphone.&lt;/li&gt;
&lt;li&gt;You need real-time stem separation in a DJ application. None of these are real-time on consumer hardware. Use a DAW with built-in real-time separation (Ableton 12, etc.) or pre-process tracks offline.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What this looks like in practice
&lt;/h2&gt;

&lt;p&gt;We run htdemucs_6s in production at &lt;a href="https://aistemsplitter.org" rel="noopener noreferrer"&gt;aistemsplitter.org&lt;/a&gt; — a hosted version of 6-stem separation aimed at people who don't want to set up the local toolchain (which, between PyTorch versions, CUDA versions, and audio dependency hell, takes most people a full afternoon).&lt;/p&gt;

&lt;p&gt;A few things we learned that aren't in the papers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Real production cost is roughly 2× what naive calculations suggest, and roughly &lt;em&gt;flat&lt;/em&gt; across Demucs variants.&lt;/strong&gt; Public GPU pricing × inference time gives you a number that ignores platform overhead. Our actual Replicate bill works out to about $0.045 per song — and it's the same number whether we run &lt;code&gt;htdemucs&lt;/code&gt;, &lt;code&gt;htdemucs_6s&lt;/code&gt;, or &lt;code&gt;htdemucs_ft&lt;/code&gt;. The fixed overhead per call swamps the marginal compute difference between models. This single fact changed how we think about model selection: pick on quality, not on theoretical compute cost, because the cost difference doesn't actually show up in your bill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Format conversion matters more than the model.&lt;/strong&gt; htdemucs only accepts WAV input. Users upload MP3, FLAC, M4A, OGG, and increasingly weird WebM containers. The pre-processing ffmpeg layer is non-trivial to get right at scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;YouTube/SoundCloud URL ingestion is half the UX win.&lt;/strong&gt; Asking users to download a file and upload it loses ~40% of them. Direct URL ingestion via yt-dlp is fiddly to maintain (age-restricted videos, region locks, livestreams) but worth it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 6-stem case is where users see the magic.&lt;/strong&gt; When someone hears guitar isolated from piano on their favorite song for the first time, they tell their friends. The 4-stem case is "neat"; the 6-stem case is "wait, how is this possible".&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want to hear what 6-stem htdemucs sounds like on real audio without setting up the toolchain, our site has free credits to try a few songs.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's next in this space
&lt;/h2&gt;

&lt;p&gt;A few open questions worth watching in 2026:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Will 8-stem (vocals/backing-vocals/drums/bass/guitar/piano/synth/other) become standard?&lt;/strong&gt; Community fine-tunes are moving in this direction, but training data for individual synth and backing-vocal stems is the bottleneck.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time on consumer hardware?&lt;/strong&gt; No current open model runs at real-time speed on a CPU at acceptable quality. This will change with model distillation, but probably not in 2026.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multilingual / non-Western vocal separation.&lt;/strong&gt; Most published benchmarks are dominated by English pop and rock. We see noticeably lower performance on languages with different vocal techniques (Mandarin, Cantopop with heavy auto-tune, Bollywood vocal stacks). This is a genuine gap in the field, not a model deployment issue.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you're working in this space and have data we'd find interesting — or you've hit something on these models we haven't — &lt;a href="https://aistemsplitter.org" rel="noopener noreferrer"&gt;drop us a line&lt;/a&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;htdemucs&lt;/strong&gt; — Rouard, S., Massa, F., Défossez, A. &lt;em&gt;Hybrid Transformers for Music Source Separation&lt;/em&gt;. &lt;a href="https://arxiv.org/abs/2211.08553" rel="noopener noreferrer"&gt;arXiv:2211.08553&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demucs v4 (hybrid)&lt;/strong&gt; — Défossez, A. &lt;em&gt;Hybrid Spectrogram and Waveform Source Separation&lt;/em&gt;. &lt;a href="https://arxiv.org/pdf/2111.03600" rel="noopener noreferrer"&gt;arXiv:2111.03600&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BS-RoFormer&lt;/strong&gt; — Lu, W.-T., Wang, J.-C., et al. &lt;em&gt;Music Source Separation with Band-Split RoPE Transformer&lt;/em&gt;. &lt;a href="https://www.researchgate.net/publication/370763841_Benchmarks_and_leaderboards_for_sound_demixing_tasks" rel="noopener noreferrer"&gt;SDX23 Challenge results&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spleeter&lt;/strong&gt; — Hennequin, R., Khlif, A., Voituret, F., Moussallam, M. &lt;em&gt;Spleeter: a fast and efficient music source separation tool with pre-trained models&lt;/em&gt;. &lt;a href="https://joss.theoj.org/papers/10.21105/joss.02154.pdf" rel="noopener noreferrer"&gt;JOSS 2020&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MUSDB18 dataset&lt;/strong&gt; — Rafii, Z., Liutkus, A., Stöter, F.-R., Mimilakis, S. I., Bittner, R. &lt;em&gt;The MUSDB18 corpus for music separation&lt;/em&gt;. &lt;a href="https://doi.org/10.5281/zenodo.1117372" rel="noopener noreferrer"&gt;Zenodo&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sound Demixing Challenge 2023&lt;/strong&gt; — &lt;a href="https://www.researchgate.net/publication/370763841_Benchmarks_and_leaderboards_for_sound_demixing_tasks" rel="noopener noreferrer"&gt;Mitsufuji et al., SDX23 results&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MVSep model leaderboard&lt;/strong&gt; — &lt;a href="https://mvsep.com/en/algorithms" rel="noopener noreferrer"&gt;mvsep.com/en/algorithms&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;Last updated: April 2026. If you find an error in the data, the SDR numbers, or any of the practical claims, &lt;a href="https://aistemsplitter.org" rel="noopener noreferrer"&gt;send us a correction&lt;/a&gt; and we'll update the post with attribution.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>deeplearning</category>
      <category>machinelearning</category>
      <category>performance</category>
    </item>
    <item>
      <title>Building a Kirkify AI Generator with Next.js and Google's Nano Banana Pro API</title>
      <dc:creator>WesLin</dc:creator>
      <pubDate>Thu, 15 Jan 2026 08:05:14 +0000</pubDate>
      <link>https://dev.to/codesugar_lin_037a57b06a4/building-a-viral-meme-generator-with-nextjs-and-googles-nano-banana-pro-api-536f</link>
      <guid>https://dev.to/codesugar_lin_037a57b06a4/building-a-viral-meme-generator-with-nextjs-and-googles-nano-banana-pro-api-536f</guid>
      <description>&lt;p&gt;&lt;strong&gt;How I turned a trending internet meme into a full-stack AI-powered web app (and why Nano Banana Pro is a game-changer for face manipulation)&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fr2f982lyl9ffc3jw3mex.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fr2f982lyl9ffc3jw3mex.png" alt="Charlie Kirk face swapped onto a lion - demonstrating Kirkify AI's capabilities" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you've spent any time on Twitter or Reddit lately, you've seen them: the "kirkification" memes. Pictures of politicians, celebrities, or random people with Charlie Kirk's distinctively proportioned face seamlessly swapped onto them. They're hilarious, they're viral, and they're surprisingly hard to make well.&lt;/p&gt;

&lt;p&gt;I wanted to create these memes. But every time I tried, I hit the same wall: manual face-swapping in Photoshop took 20+ minutes per image, online meme generators produced low-quality results, and I never quite nailed the subtle facial proportions that make a kirkification truly cursed.&lt;/p&gt;

&lt;p&gt;So I did what any developer would do: I built Kirkify AI, a web app that generates photorealistic Charlie Kirk face swaps in seconds using Google's Nano Banana Pro API. Here's how I went from frustration to a working product, and why Nano Banana Pro turned out to be the perfect tool for this absurdly specific problem.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem: When Memes Take Longer Than They Should
&lt;/h2&gt;

&lt;p&gt;I wanted to create quality kirkification memes, but every approach had major drawbacks:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Option 1: Photoshop Manual Labor&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;15-20 minutes per meme masking, aligning, and blending&lt;/li&gt;
&lt;li&gt;Required significant design skills&lt;/li&gt;
&lt;li&gt;Results often looked fake despite the effort&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Option 2: Existing Online Tools&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Generic face-swap apps that didn't understand kirkification&lt;/li&gt;
&lt;li&gt;Low-quality results with obvious cutout edges&lt;/li&gt;
&lt;li&gt;Often watermarked or paywalled&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Option 3: Manual AI Prompting (Midjourney, DALL-E)&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Wildly inconsistent results&lt;/li&gt;
&lt;li&gt;No reference image support for consistent likeness&lt;/li&gt;
&lt;li&gt;AI didn't understand Charlie Kirk's specific proportions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fundamental issue? A speed versus quality trade-off that felt impossible to resolve. I wanted a tool that could generate photorealistic kirkifications in under 30 seconds without requiring Photoshop expertise or AI prompting knowledge.&lt;/p&gt;

&lt;p&gt;That's when I started looking into AI image APIs specifically designed for face manipulation, and that's how I discovered Nano Banana Pro.&lt;/p&gt;




&lt;h2&gt;
  
  
  Discovering Nano Banana Pro: The API I Didn't Know I Needed
&lt;/h2&gt;

&lt;p&gt;My requirements were clear: handle face manipulation with high fidelity, support reference images (not just text prompts), process requests asynchronously, and deliver photorealistic results consistently.&lt;/p&gt;

&lt;p&gt;I explored OpenAI's DALL-E, Replicate's Flux models, and Stability AI—all excellent for general image generation, but face-swapping is a specialized task requiring precise facial feature mapping and seamless blending.&lt;/p&gt;

&lt;p&gt;Then I discovered Google's Nano Banana Pro. Here's what made it stand out:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Nano Banana Pro Won:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Face manipulation as core functionality&lt;/strong&gt; - Not an afterthought, but engineered specifically for this&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Native reference image support&lt;/strong&gt; - Feed it Charlie Kirk's actual face for consistent results&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Async callback architecture&lt;/strong&gt; - Perfect for time-intensive operations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Photorealistic quality&lt;/strong&gt; - Outputs indistinguishable from real photographs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Intelligent context preservation&lt;/strong&gt; - Maintains hair, clothing, lighting, and backgrounds automatically&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Granular facial proportion understanding&lt;/strong&gt; - Actually understands facial geometry, not just overlaying faces&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point was crucial. Charlie Kirk's face has very specific proportions that make the meme work. Nano Banana Pro understands facial geometry—the spatial relationships between eyes, nose, mouth—and maps them accurately while adapting to different angles and lighting.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3d3v1jm7xa4kt2eo306q.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3d3v1jm7xa4kt2eo306q.webp" alt="Deadpool kirkification example" width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Even complex costumes and comic book characters work flawlessly—Deadpool gets the full kirkification treatment&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Building Kirkify AI: Tech Stack &amp;amp; Architecture Decisions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Tech Stack:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Next.js 16&lt;/strong&gt; (App Router) - Server-side rendering, API routes, React 19&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TypeScript&lt;/strong&gt; - Type safety for async operations and API contracts&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PostgreSQL + Drizzle ORM&lt;/strong&gt; - Track generation tasks and status&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TailwindCSS + Radix UI&lt;/strong&gt; - Fast, accessible UI components&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vercel AI SDK&lt;/strong&gt; - Multi-provider support foundation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The Async Architecture Challenge:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Nano Banana Pro takes 10-30 seconds per face manipulation. I needed a system that could handle this without blocking the UI or timing out. Here's the flow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;User submits image&lt;/strong&gt; → Frontend calls &lt;code&gt;/api/generate-images&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend creates task&lt;/strong&gt; → Saves to database (status: "queued"), returns run ID&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend calls Nano Banana Pro&lt;/strong&gt; → Via APImart with callback URL, then moves on&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nano Banana Pro processes&lt;/strong&gt; → Async on their infrastructure&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Webhook receives result&lt;/strong&gt; → &lt;code&gt;/api/ai-callback/apimart&lt;/code&gt; updates database (status: "success")&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frontend polls status&lt;/strong&gt; → &lt;code&gt;/api/ai-task-status&lt;/code&gt; with exponential backoff (1s → 2s → 4s → 10s max)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User sees result&lt;/strong&gt; → When polling detects "success" status&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This webhook + polling hybrid gives real-time feedback without hammering the API. The exponential backoff keeps it responsive early on but backs off for longer operations.&lt;/p&gt;

&lt;p&gt;I also built multi-provider support (OpenAI, Replicate, FAL) for flexibility and redundancy, though Nano Banana Pro became the star performer for face manipulation.&lt;/p&gt;


&lt;h2&gt;
  
  
  The Prompt Engineering Journey: Less Work Than Expected
&lt;/h2&gt;

&lt;p&gt;I expected this to be the hardest part—getting AI to consistently understand "kirkification" while preserving context seemed like weeks of work.&lt;/p&gt;

&lt;p&gt;Here's what actually happened: Nano Banana Pro just worked. After a few test runs, I landed on this prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A photorealistic face swap based on this reference image.
Replace the face of the person in the image completely with
the face of Charlie Kirk. Capture Charlie Kirk's distinct
facial features, proportions, and likeness. Crucially,
maintain the exact original context: keep the same hair,
clothing, body pose, background setting, lighting conditions,
and film grain. Seamless blending, highly detailed,
realistic photograph.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. No complex parameter tuning, no multi-step pipelines, no special case handling. This single prompt works reliably across all scenarios.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why This Worked:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Natively understands facial geometry and proportions&lt;/li&gt;
&lt;li&gt;Handles context preservation holistically without explicit per-element instructions&lt;/li&gt;
&lt;li&gt;Performs seamless blending across different lighting conditions automatically&lt;/li&gt;
&lt;li&gt;Renders photorealistic results without over-processing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nano Banana Pro reduces complex face manipulation to a simple, descriptive prompt. The API handles the hard parts—facial geometry mapping, lighting adaptation, edge blending, context preservation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwt0k4d11at6qz4ef3m70.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwt0k4d11at6qz4ef3m70.webp" alt="Spider-Man kirkification" width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The same prompt works across wildly different scenarios—Spider-Man in action&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzlv3603f5cjcncbk9bql.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzlv3603f5cjcncbk9bql.webp" alt="Iron Man kirkification" width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;From superheroes to mech suits, consistent quality across all inputs&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Result: From Idea to Production
&lt;/h2&gt;

&lt;p&gt;After integrating everything, I had a working product: upload any image, wait 20-30 seconds, and receive a photorealistic Charlie Kirk face swap.&lt;/p&gt;

&lt;p&gt;The user experience is deliberately simple. Upload a photo, and within half a minute you get back a seamless face swap with no cutout lines, no weird blending artifacts, and the original context perfectly intact. Everything except the face—hair, clothing, background, lighting—remains untouched.&lt;/p&gt;

&lt;p&gt;Here are some real examples:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0pksuk5orkrblkmksult.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0pksuk5orkrblkmksult.png" alt="Lion kirkification - before and after comparison" width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;A majestic lion gets the kirkification treatment—notice how Charlie Kirk's facial proportions are preserved while maintaining the original lighting and mane&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9k1np3co91bgutcalsxh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9k1np3co91bgutcalsxh.png" alt="Trump kirkification - before and after comparison" width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Even complex facial expressions and professional photography lighting are handled seamlessly&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhdcn5ptx8k1zv6np9564.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhdcn5ptx8k1zv6np9564.webp" alt="Fox kirkification" width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Animals work particularly well—this fox demonstrates perfect facial proportion mapping&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The results speak for themselves. These look like actual photographs that could plausibly exist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Learnings:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Choosing the right API saves months&lt;/strong&gt; - Nano Banana Pro's specialized capabilities eliminated the need to build face-swapping logic from scratch&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Async architecture is essential&lt;/strong&gt; - Quality face manipulation takes time; trying to force sync would cause timeouts and frustrated users&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Invest in prompt templates upfront&lt;/strong&gt; - One well-crafted prompt has powered thousands of generations with ~95% success rate&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-provider redundancy works&lt;/strong&gt; - Fallback providers keep the app reliable during peak loads&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Performance-wise: most generations complete in 15-25 seconds, and based on the increasingly cursed memes being created, user satisfaction is high.&lt;/p&gt;




&lt;h2&gt;
  
  
  Try Kirkify AI (and Let the Cursed Memes Begin)
&lt;/h2&gt;

&lt;p&gt;Kirkify AI is live at &lt;a href="https://ai-kirkify.com" rel="noopener noreferrer"&gt;ai-kirkify.com&lt;/a&gt;. Upload any photo—friends, celebrities, even pets—and watch the face-swapping magic happen.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For Developers:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The architecture patterns here—async callbacks, webhook handling, polling with exponential backoff—apply to any long-running API operations. If you're building with AI image generation or video processing, these patterns are broadly applicable.&lt;/p&gt;

&lt;p&gt;Nano Banana Pro is accessible through &lt;a href="https://apimart.ai/" rel="noopener noreferrer"&gt;APImart&lt;/a&gt;, an AI API marketplace. The documentation is clear, the API is straightforward, and the results deliver on the promise of photorealistic face manipulation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's Next:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'm exploring batch processing for multiple images, style transfer options, and expanding to other viral meme formats. The multi-provider architecture makes adding new capabilities relatively painless.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Takeaway:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Sometimes the best way to learn a technology is to build something absurd with it. Kirkify AI started as "I want faster Charlie Kirk memes" and evolved into a robust AI-powered web app with proper architecture and production-grade error handling. If you're curious about AI image generation, pick a ridiculous idea and start building. You'll learn more from shipping than from tutorials alone.&lt;/p&gt;

&lt;p&gt;Now go forth and kirkify responsibly. Or irresponsibly. I'm not your boss.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Want to see the code or contribute?&lt;/strong&gt; Check out the &lt;a href="https://github.com/codeugar/ai-kirkify.com" rel="noopener noreferrer"&gt;GitHub repository&lt;/a&gt; or try the tool at &lt;a href="https://ai-kirkify.com" rel="noopener noreferrer"&gt;ai-kirkify.com&lt;/a&gt;.&lt;/p&gt;




</description>
      <category>nextjs</category>
      <category>ai</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
