<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Humza Tareen</title>
    <description>The latest articles on DEV Community by Humza Tareen (@humzakt).</description>
    <link>https://dev.to/humzakt</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1039696%2F3f502619-3a9f-4833-b6d1-40ca618daef0.jpg</url>
      <title>DEV Community: Humza Tareen</title>
      <link>https://dev.to/humzakt</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/humzakt"/>
    <language>en</language>
    <item>
      <title>One Brief Instead of Four: Unifying Create Flows Into a VSL Generator and an AI Presenter</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:13:30 +0000</pubDate>
      <link>https://dev.to/humzakt/one-brief-instead-of-four-unifying-create-flows-into-a-vsl-generator-and-an-ai-presenter-276l</link>
      <guid>https://dev.to/humzakt/one-brief-instead-of-four-unifying-create-flows-into-a-vsl-generator-and-an-ai-presenter-276l</guid>
      <description>&lt;p&gt;The main video-generation service had accumulated four separate ways to start an ad: clone an ad, original from script, video sales letter, and — briefly — a fourth. Each had its own page, its own header, its own place in the sidebar. They looked like four products. Underneath, they were four thin page files sitting over one shared wizard component, and the words that told them apart never made it past the page layer into the code that actually planned and rendered anything. Collapsing that into one brief-driven flow, and then building two genuinely new product surfaces on top of the result, took about three weeks and touched navigation, script splitting, sample pricing, and — closing the loop — the first real finished output the new flow ever produced.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A chooser that asks a question whose answers converge is not a choice. It is a detour.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Four routes, two runs
&lt;/h2&gt;

&lt;p&gt;PR #993 found the tell before it fixed anything: &lt;code&gt;videoTypeForCreateMode&lt;/code&gt; mapped two of the four entry points — "AI Presenter" and "video sales letter" — to the exact same &lt;code&gt;short_vsl&lt;/code&gt; job tag. Different headers, different marketing copy, byte-identical jobs underneath. The fix collapsed the chooser to one card, with the four original types demoted to a picker inside a single brief: storyboard, UGC video, VSL, or reference recreation. &lt;code&gt;/create/vsl&lt;/code&gt; became the only page that actually mounts the form; the other three routes now just redirect into it. The pipeline itself — &lt;code&gt;/api/wizard/*&lt;/code&gt; — wasn't touched at all, which is the property that made the change checkable: nothing about how a job renders was supposed to change, only how many doors led to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A second card, justified by a different job
&lt;/h2&gt;

&lt;p&gt;PR #994 is the interesting counter-move to #993, and it's worth sitting with the order these two shipped in: the very next day. A stakeholder asked for a dedicated AI Presenter wizard — one person to camera, no reference video, none of the fancier modes — in the same conversation that asked to delete the "no reference" button from reference recreation. The PR's own comment names the standard #993 had just set: a chooser that converges to the same job is a detour, not a choice. So the new card only earns its place by producing something the merged flow didn't already produce — a distinct &lt;code&gt;videoType: "ai_presenter"&lt;/code&gt; rather than a second front door to &lt;code&gt;short_vsl&lt;/code&gt;. Two products that looked identical a day earlier now diverge by design, not by accident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eight back buttons, none of them the same
&lt;/h2&gt;

&lt;p&gt;The AI Presenter card shipped with an unglamorous bug: no way back to the chooser screen. PR #995 could have patched that one path. Instead, asked to audit the whole tool's navigation rather than add a ninth arrow, it found eight hand-rolled back treatments across five different glyphs — &lt;code&gt;←&lt;/code&gt;, &lt;code&gt;‹&lt;/code&gt;, &lt;code&gt;↺&lt;/code&gt;, a chevron, a bare dropdown caret — and &lt;code&gt;router.back()&lt;/code&gt; appearing exactly zero times anywhere in the codebase. Two of the eight were the same bordered &lt;code&gt;←&lt;/code&gt; chip routed to two different destinations depending on which screen rendered it. The fix is one shared back component with one behavior, replacing all eight, closing not just the reported bug but the entire class it came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Splitting a script where a viewer would actually cut
&lt;/h2&gt;

&lt;p&gt;PR #996 fixed a scene-splitting bug that had been hiding behind a workaround written into the product's own documentation. A line like "When she told me about this, I smiled and said I'd look into it" is two beats of performance carried by one beat of picture — a natural cut point sits at the comma. The splitter had no way to see it: both the sentence-partition logic and the scene-count cap treated a sentence, not a clause, as the smallest unit a cut could land on. The operator-facing guidance for this exact problem was to punctuate a script with extra full stops to manufacture more cut points — grammatically wrong on purpose, as a workaround for a limit that shouldn't have existed. The fix moves the atom from sentence to clause boundary, following the same rule every subtitling authority independently converges on for exactly this reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  A sample that finally shows what the money buys
&lt;/h2&gt;

&lt;p&gt;The richest thread in this whole arc starts with a plain product complaint: the pre-review stage — the one screen positioned specifically to catch a bad plan before the pipeline renders and bills for the full fleet — showed nothing but still frames. Pacing and line delivery are both fundamentally time-domain; a still frame can't show either. PR #997 changes pre-review to actually render two real scenes before the operator commits to the rest, with a feedback box underneath so a correction can be applied to the sample and re-checked before the expensive part starts.&lt;/p&gt;

&lt;p&gt;Three fast-follow PRs hardened it, and the middle one is worth reading as a genuine self-correction rather than a straight line of fixes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;#998&lt;/strong&gt; closed a race: applying a feedback note re-renders and re-stitches the sample, which is safe only at the exact gate it was written for — mid-sample it collides with a render already in flight, and after approval it silently re-bills work the operator already paid for. The route now refuses anything outside that one window by name, because a safety property that depends on the screen never sending the request late is one stale tab away from being violated anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;#999&lt;/strong&gt; found the quote was reading half its own bill: the cost estimator priced a voice-bound clip at the bare per-second rate, while the actual billing code multiplied that same rate whenever the request carried image-set elements — which, on the UGC path, is every clip, single presenter included. A mirror test had been passing the whole time because it only checked that the base rate matched; nothing checked the multiplier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;#1000&lt;/strong&gt; then made a call that turned out to be wrong, and the PR that follows it says so directly rather than quietly overwriting it. #1000 reasoned that the sample quote and the full-run quote should share one number — the plan's average beat length — since they're "an inch apart" and answer what looked like the same question. #1001, written against a real paid run, showed they don't: the sample quote and the full quote answer two different questions ("what does this click spend" vs. "what does the whole ad cost"), and forcing them onto the same average produced a 12% under-quote on the exact number an operator is about to commit money against. The fix reverted to pricing each figure off the beats it actually renders — the sample from the sampled beats, the full estimate from the whole plan — and the PR is explicit that this reverses a call made one PR earlier, with the measured run that proved it wrong attached as evidence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That reversal is the part worth remembering longer than any individual bug: the fix that shipped first wasn't obviously wrong when it shipped. It took a real paid run, not more code review, to show the two numbers needed to diverge on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  A vendor's zero exit code is a claim, not proof
&lt;/h2&gt;

&lt;p&gt;Three defects landed on the same day, and the RCA that followed them is a small case study in not trusting your own fix. PR #1002 traced a transcription failure back to its real cause: ffmpeg had already refused to demux the uploaded file and logged exactly why, and the code downgraded that free, authoritative, local signal to a warning and handed the same unreadable file to a paid remote transcription vendor instead — which then failed with a much less specific error. The fix makes a failed local probe stop the pipeline immediately rather than pass the problem downstream.&lt;/p&gt;

&lt;p&gt;PR #1005 is the sequel, and its own opening line states the discipline plainly: it exists because the author tested the #1002 fix instead of trusting it, by uploading a deliberately truncated file to the live deployment. The new guard didn't catch it. Every individual check the previous PR had added — video probe, ffmpeg exit code, non-empty output — passed cleanly on a corrupted result: ffmpeg reported success while writing a 799-byte MP3 with no decodable audio frames inside it, because a zero exit code only claims the tool didn't crash, not that its output means anything. The real fix probes the extracted audio file itself, not the process that produced it.&lt;/p&gt;

&lt;p&gt;In between, PR #1003 traced a "scene cuts not applied" report back to the clause-splitter from #996: one sentence in a real script carried five commas and a dash, every individual clause candidate fell under the minimum word floor, and the merge-short-atoms step folded all of them into one indivisible 33-word atom — 30% of the whole script riding on one shot. The fix caps how large a single clause atom is allowed to grow before it's forced to split anyway, so a floor meant to prevent scenes that are too short can no longer produce one scene that eats a third of the ad. And PR #1004 traced a "the character doesn't look like Pixar, it looks like our normal avatar" report to a filename: the animated character had been generated and paid for correctly, uploaded to Drive under a name the serving route's lookup table didn't recognize, so every request for it silently 400'd and the pipeline fell back to the photographic original nobody had asked for.&lt;/p&gt;

&lt;h2&gt;
  
  
  A progress bar that says how long "generating" actually takes
&lt;/h2&gt;

&lt;p&gt;PR #1006 is a smaller, purely UX fix worth including because of the number behind it: the render stage — the one that takes roughly five minutes per scene, the better part of an hour on a nine-scene ad — showed a bar frozen at 0% with the word "starting…" for the entire duration. It was one of roughly forty hand-rolled waiting states across the app, and only two of them used the shared progress component that already existed. The fix consolidates onto one component that names the current scene, elapsed time, and a realistic estimate range, mounted everywhere a screen can run for minutes rather than seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  A seventh stage, because "done" was a dead end
&lt;/h2&gt;

&lt;p&gt;Before PR #1007, a delivered ad was terminal. Every edit allowlist in the repository named only the review steps that come before delivery, so the instant an ad shipped, the only way to change anything about it was to pay for an entirely new render — a standing do-not-retry. #1007 adds an eighth-hour-of-the-day feature that sounds simple and touches everything: a seventh pipeline stage, Edit, with per-scene regenerate priced on the button itself, saved takes, master version history, and a genuinely free editorial cut. It's registered once, in the shared &lt;code&gt;STAGES&lt;/code&gt; array every flow already reads from — there's no separate per-flow wiring to forget.&lt;/p&gt;

&lt;p&gt;The three PRs that followed it are the now-familiar shape of shipping something and then actually using it in production before calling it finished. #1008 found two of the new stage's refusal messages contradicting the condition they were checking — a delivered ad refused with "can only be edited once the ad is finished," describing a state the ad was already in. #1009 found the new edit route was the only mutating wizard route doing a bare in-memory job lookup instead of falling back to the Drive-backed restore every other route already used, which meant reopening an edit on an ad that had survived a container restart — the exact case the feature exists for — 404'd. #1010 found a fourth refusal site with the same contradictory wording that #1008's by-hand fix had missed, and closed it by centralizing the message next to the predicate it describes rather than leaving four hand-written copies of one rule for a future bug to hide in.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first finished VSL, and what it actually surfaced
&lt;/h2&gt;

&lt;p&gt;PR #1019 is where this whole arc reports back. Taking a real product script through the new VSL Generator, in the Pixar visual style, with an AI-generated presenter, produced the first ad this exact configuration had ever finished rendering — and finishing it surfaced nine real defects, four fixed same-day and five more detailed in this PR. The most interesting one is a retraction: an earlier report had blamed a density control for doing nothing. It wasn't inert. A second, unrelated cap on how many generated lifestyle spans could ship was silently deleting the density control's own output after placement — the density setting worked exactly as configured, and something downstream threw its answer away. Measured against two real finished runs, the arithmetic lines up exactly: four dropped spans totaling 10.42 seconds, and 20.56 minus 10.42 is 10.14 — the shipped total to the decimal. The bug wasn't where the first report pointed; the first report was simply reading the wrong layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: a flow isn't unified until someone finishes an ad through it
&lt;/h2&gt;

&lt;p&gt;Every stage of this arc reads clean in isolation — collapse four routes to one, add a card that's actually different, fix the navigation it broke, split scripts at the right boundary, make the sample honest, make the quote honest, give delivered ads a way back in. What ties them together is where almost every fix came from: not code review, but a real run, watched end to end, with the actual numbers pulled from logs and job documents rather than assumed from reading the code that produced them. The reversal in the sample-pricing arc and the retraction in the closing PR are the clearest examples, but the shape repeats throughout — a fix that looks complete when it merges and turns out to be answering the wrong layer of the question, caught only once someone pushed a real script all the way through to a finished ad.&lt;/p&gt;

</description>
      <category>aivideo</category>
      <category>productengineering</category>
      <category>costengineering</category>
      <category>ux</category>
    </item>
    <item>
      <title>Stop Losing Paid Renders: Hardening a Portrait-to-Widescreen Reframing Tool</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:12:56 +0000</pubDate>
      <link>https://dev.to/humzakt/stop-losing-paid-renders-hardening-a-portrait-to-widescreen-reframing-tool-53da</link>
      <guid>https://dev.to/humzakt/stop-losing-paid-renders-hardening-a-portrait-to-widescreen-reframing-tool-53da</guid>
      <description>&lt;p&gt;The reframing tool does one job: take a 9:16 clip and outpaint it into 16:9, so a vertical ad can run on a widescreen placement without cropping the subject out. It's a young tool — its vendor integration dates back one month at the time of this audit — and it had reached the point where editors were running real, paid renders through it every day. That transition is where it stopped being a demo and started needing to survive redeploys, vendor flakiness, and its own money math being wrong.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Clean HTTP metrics don't mean nothing is failing. They mean nothing is failing at the layer you're measuring.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A render abandoned by a coincidence of timing
&lt;/h2&gt;

&lt;p&gt;PR #1 traces a specific, maddening failure: a clip errors with "Lost connection to server... the render may have finished — check before re-running," which is a client-side give-up, not a real server error. The browser polls a status endpoint through a hub proxy in front of the tool, and while that hub redeploys — routinely 60 to 180 seconds of 502s on Railway — eight consecutive poll failures trip a client-side timeout and the UI marks the item as errored. The server was never told anything was wrong. The worker thread kept running. The job row survived. The editor just had no way to know that, so a perfectly good render looked abandoned because of unlucky timing against an unrelated deploy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Clean metrics, invisible losses
&lt;/h2&gt;

&lt;p&gt;PR #2 generalizes the fix, and its opening framing is the most useful sentence in the whole cluster: Railway's HTTP metrics were clean — 10.4k requests, zero 5xx, a flat 0.0% error rate. Nothing was failing at the web layer. &lt;strong&gt;Jobs were being lost, not erroring&lt;/strong&gt;, and the only record of a failure was &lt;code&gt;str(exc)&lt;/code&gt; in a dict that dies with the container the moment it restarts. Four distinct loss paths came out of the metrics and a production log — including the vendor rejecting a submission with a 429 when a shared account was busy — and the fix was to retry every vendor call and make failures diagnosable by writing them somewhere that outlives the process, instead of trusting a clean-looking dashboard to mean the system was healthy.&lt;/p&gt;

&lt;h2&gt;
  
  
  720p, billed as 1080p
&lt;/h2&gt;

&lt;p&gt;PR #4 starts from an editor's finished render answering its own download link with Google's 401 page, an hour after the run completed with the file intact in Drive — and chasing that turned up something worse in the same batch of runs. Every run was supposed to use a specific native-1080p pipeline. Four independent measurements — wall-clock render time, bitrate, blur at 1fps sampled over 259 frames, and PSNR after a 1080→720→1080 round trip — separated the runs into two groups with &lt;strong&gt;no overlap&lt;/strong&gt; between them:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;wall-clock&lt;/th&gt;
&lt;th&gt;bitrate&lt;/th&gt;
&lt;th&gt;blur (n=259)&lt;/th&gt;
&lt;th&gt;PSNR round-trip&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;native 1080p&lt;/td&gt;
&lt;td&gt;6.9 min median&lt;/td&gt;
&lt;td&gt;6.40–7.74 Mbps&lt;/td&gt;
&lt;td&gt;7.50 (sd 1.46)&lt;/td&gt;
&lt;td&gt;48.2–51.0 dB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;downgraded to 720p&lt;/td&gt;
&lt;td&gt;lower&lt;/td&gt;
&lt;td&gt;lower&lt;/td&gt;
&lt;td&gt;higher (softer)&lt;/td&gt;
&lt;td&gt;lower (out of range)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Seven of twenty-eight delivered segments had shipped as upscaled 720p while billed at native-1080p price — ten of thirty-seven segment rows were missing the field that records which upscale pass actually ran, while still reading &lt;code&gt;status='complete'&lt;/code&gt;. The fix wasn't just correcting the price. It was serving the actual durable copy of a finished render instead of a transient one that could expire and start answering with a login page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two UI bugs nobody had reported, found by auditing the fix
&lt;/h2&gt;

&lt;p&gt;PR #5 is the follow-up audit on #4, and it catches the kind of bug that survives specifically because it fails quietly in a direction nobody's watching. The segment progress strip branched its color logic on &lt;code&gt;"done"&lt;/code&gt;, &lt;code&gt;"rendering"&lt;/code&gt;, &lt;code&gt;"processing"&lt;/code&gt;, and &lt;code&gt;"abandoned"&lt;/code&gt; — none of which is an actual value of the status field it was reading. The real "finished" value is &lt;code&gt;complete&lt;/code&gt;; &lt;code&gt;abandoned&lt;/code&gt; belongs to an entirely different enum. Only the &lt;code&gt;failed&lt;/code&gt; branch ever matched anything real, so every complete, reused, or submitted segment rendered grey, as if it hadn't started — for the entire lifetime of the feature, with CSS that would have shown the correct color sitting right there, unused, the whole time. The fix pins the mapping end to end: a SQL CHECK constraint, an enum, and now two runtime guards, one in each direction, because the audit's own conclusion was that the second direction — catching an enum value with no matching UI branch — was the one that actually mattered.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same missing-import bug, twice
&lt;/h2&gt;

&lt;p&gt;PR #6 and PR #8 are the same failure mode a month apart, and the second time it happened is what made it worth fixing structurally rather than patching again. In #6, &lt;code&gt;formatMin&lt;/code&gt; was called in two modules that never imported it — one inside a &lt;code&gt;confirm()&lt;/code&gt; dialog string for the main "render N clips" button, meaning that button silently started zero renders, with no visible error, because the thrown &lt;code&gt;ReferenceError&lt;/code&gt; pre-empted the dialog entirely. The per-row "run one clip" button worked fine, which is exactly why nobody noticed the batch button was dead. In #8, the same shape of bug hit &lt;code&gt;APP_PREFIX&lt;/code&gt;: read in one module, imported in a sibling module it had been split out of, unreachable until a render actually succeeded and then thrown on every subsequent paint — so a row would sit on "Submitting..." for two real hours while the backend had already finished and failed nineteen minutes in. PR #9 is what actually closes the class of bug: a lint guard extended to flag any bare reference to an unimported binding, not just function calls — and the guard itself had a gap, since testing a name inside an &lt;code&gt;if (APP_PREFIX)&lt;/code&gt; condition had been silently counted as a valid "use" that proved the name was bound, hiding 91 such cases across the codebase from the very check meant to catch them.&lt;/p&gt;

&lt;h2&gt;
  
  
  An adopted row that couldn't actually retry
&lt;/h2&gt;

&lt;p&gt;PR #10 is a sharp one: a row restored from a prior partial run — an "adopted" row — carries a placeholder file object with &lt;code&gt;size: 0&lt;/code&gt;, because the real bytes live on the server and the browser tab looking at it never had them. That fact had been recorded on the row since PR #3. No code had ever read it. So clicking "Retry anyway" on an adopted row ran the exact same fallback path as a fresh upload — &lt;code&gt;FormData.append("video", placeholder)&lt;/code&gt; — which JavaScript happily stringifies into the literal text &lt;code&gt;"[object Object]"&lt;/code&gt; as a form field, sent to a server that (correctly) rejected it as garbage, discarding a render that should have simply been re-run server-side from data it already had.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two deprecated vendor models and a fallback that made things worse
&lt;/h2&gt;

&lt;p&gt;PR #11 through #15 form the last arc, and it starts with a question someone asked directly: is the model outdated? PR #12 confirms it in the vendor's own words — both models the tool was using are explicitly deprecated, and the integration dated to the tool's very first commit, a month before the replacement model even shipped, never revisited since. But before reaching for the obvious "just upgrade," PR #11 measured what was actually happening: across 13 days and 64 submissions, the fast tier of the old model had a 0% failure-to-start rate; the tier the system fell back to when something went wrong had a 25% failure-to-start rate and was reached exclusively through the exact failure path meant to recover from trouble. &lt;strong&gt;The recovery path was less reliable than what it was recovering from&lt;/strong&gt; — a fallback that can make an outage worse than doing nothing.&lt;/p&gt;

&lt;p&gt;Moving to the current model (PR #12) opened a second, subtler problem: the two vendor integration paths — "brokers," in the tool's own language — disagreed with each other's documentation about a basic parameter. One broker's docs said an omitted duration field defaults to matching the source; the other's said it defaults to a flat 5 seconds. PR #13 fixed the code to match what each vendor's docs claimed. PR #14 then paid for three real generations to check the docs against reality, and found two of the three documented behaviors were simply wrong — the field made no observable difference to either broker's output. PR #15 finally exercised the previously-never-successful integration path end to end with a real purchase, confirming the correct behavior directly rather than trusting anyone's documentation, including the vendor's own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: measure the money, not just the uptime
&lt;/h2&gt;

&lt;p&gt;Nothing in this cluster failed loudly. A render abandoned by a proxy redeploy still looked like a normal error. A dashboard reporting 0% error rate was telling the truth about the layer it measured and nothing about the layer that mattered. Wrong-resolution output still played back fine in a preview. A fallback path can pass code review and still be worse than not falling back at all. The throughline across all fifteen PRs is the same: don't trust an instrument that's clean by construction — measure the actual bytes delivered, the actual price charged, and the actual vendor behavior under a real paid call, because that's the only place these bugs were ever visible.&lt;/p&gt;

</description>
      <category>aivideo</category>
      <category>reliability</category>
      <category>costengineering</category>
      <category>videorendering</category>
    </item>
    <item>
      <title>Building an Internal Ops Platform: Employee Management, Evaluations, and a Proxy That Was Leaking Secrets</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:12:52 +0000</pubDate>
      <link>https://dev.to/humzakt/building-an-internal-ops-platform-employee-management-evaluations-and-a-proxy-that-was-leaking-16k6</link>
      <guid>https://dev.to/humzakt/building-an-internal-ops-platform-employee-management-evaluations-and-a-proxy-that-was-leaking-16k6</guid>
      <description>&lt;p&gt;An internal ops platform for a video-production team started as a project tracker and grew, over about three weeks, into something closer to an HR system: employee scorecards, evaluation dashboards, a product catalog, and a shared proxy layer that every other internal tool routes through. That last part is what makes this cluster of PRs worth writing up together — a proxy sitting in front of several tools accumulates security bugs differently than any single tool does, because a mistake there isn't scoped to one screen. It's scoped to everything behind it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A permission check that lives only in the UI is a suggestion. A permission check that lives at the route is a rule. Ship the rule.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Turning hardcoded lists into admin tables
&lt;/h2&gt;

&lt;p&gt;The project tracker's list of valid projects lived in two places that had to agree with each other: a TypeScript union in the board component, and a database &lt;code&gt;CHECK&lt;/code&gt; constraint. Adding a project meant a pull request and a deploy for what is, functionally, one row of data. PR #16 replaced both with a real &lt;code&gt;projects&lt;/code&gt; table and an admin UI, and PR #18 shipped it to production with a detail worth keeping: a project's board section only renders once it actually has a task, so the board doesn't fill with empty tables as the project list grows ahead of real work. PR #19 then gave Projects its own sidebar entry under Admin instead of leaving it stacked under Users &amp;amp; Access — a small move, but it's the difference between an admin feature existing and an admin feature being findable.&lt;/p&gt;

&lt;p&gt;The same pattern repeated twice more, each time for a reason specific enough to be worth its own PR. PR #47 moved the export tool's product dropdown out of hardcoded TypeScript in a sibling repo into an &lt;code&gt;export_products&lt;/code&gt; table with its own admin page — and the PR body is explicit about why this didn't just reuse the existing &lt;code&gt;projects&lt;/code&gt; table, walking through the field-level differences that made a shared table the wrong shortcut. PR #50 did the same for evaluation categories, and it's the richest of the three: &lt;code&gt;detectCategory&lt;/code&gt; is an ordered ladder of string-matching rules where the order itself carries meaning — "LEAN UV Cleaner Upsell C1" is a real production row, and it scores correctly as a Lead only because the L-prefix rule is checked before anything that matches "Upsell." Turning that ladder into editable data meant preserving the ordering as a first-class &lt;code&gt;priority&lt;/code&gt; column, not just a list — getting the migration wrong would have silently reclassified real historical rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Applying a rule retroactively means measuring before writing
&lt;/h2&gt;

&lt;p&gt;PR #49 added Upsell as its own scoring category — a substring match, the only one in the whole ladder, deliberately placed last so it can only claim what no earlier, more specific rule already claimed. PR #52 then applied that rule to all of history, and the PR description is a small model of how to do a backfill safely: before writing anything, it measured read-only what was actually in the database — confirming nothing had ever been scored in this tool before a specific date, so there was no prior scorecard the backfill could silently contradict. PR #53 closes the loop with something that isn't a code change at all: a written record of a deliberate decision — asked directly whether any id containing "upsell" should always score as Upsell, the category owner said no, categories should strictly follow the first-letter code, because some upsells are really just leads or bodies underneath. That answer is exactly why the substring rule sits last in the ladder rather than first, and writing the reasoning down next to the rule is what stops a future well-intentioned refactor from "fixing" the ordering into something that quietly breaks four real rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Employee Management: access enforced in the API, not the component
&lt;/h2&gt;

&lt;p&gt;PR #26 shipped the Employee Management tool, and its own description states plainly why it isn't simply the feature as originally built: the manager's board component was a 3,822-line client component making 34 direct Postgrest calls straight from the browser. Every one of those calls was only as safe as the row-level security policy behind it — a UI-only gate, dressed up as a feature. The shipped version moves access enforcement into the API layer and normalizes the editor reference so editor data is looked up once, not repeated across dozens of call sites.&lt;/p&gt;

&lt;p&gt;PR #27 is the verification that followed the launch, and it's a good model for what "verify a permission boundary" actually looks like in practice: a manager asked directly whether the Evaluation tab was private to each editor, and rather than answer from memory, the check walked all three layers — route, API, and rendered payload — and found the tab genuinely was private, but also found a real bug hiding next to the correct behavior. An editor's page was unconditionally requesting a manager-only dashboard resource, which meant a guaranteed 403 firing silently on every single editor page load. Confirming a security property is correct doesn't mean the code path exercising it is clean, and this PR is the difference between checking the outcome and checking the mechanism.&lt;/p&gt;

&lt;p&gt;PR #28 shipped a "you vs. the team" comparison card on an editor's own scorecard — four KPI averages plus their own numbers, side by side — and the type design does real work here: the payload type is a closed shape with exactly four numbers and a rating band, structurally unable to carry a per-editor name or id. Widening it into per-editor data later can't happen by accident, because there's nowhere in the type for that data to go. That's access control enforced by the shape of the data, not by a check someone has to remember to write.&lt;/p&gt;

&lt;h2&gt;
  
  
  The proxy: four bugs, each found while fixing a different one
&lt;/h2&gt;

&lt;p&gt;The most consequential work landed on the shared proxy every tool in the platform sits behind, and it's a clean example of how security bugs cluster: fixing one often surfaces the next, sitting right next to it.&lt;/p&gt;

&lt;p&gt;PR #29 found that GET and HEAD requests were proxied with &lt;code&gt;redirect: "follow"&lt;/code&gt;, so the underlying HTTP client resolved any 3xx server-side without checking whose origin it was following to. One tool redirects a finished render to Google Drive; the proxy followed that redirect itself, with no Google session attached, got a 401 back from Google, and returned Google's raw HTML error page to the browser — under this app's own origin. A user who successfully finished a render an hour earlier saw a third-party error page served as if it came from the platform they were using. The fix scopes &lt;code&gt;follow&lt;/code&gt; to same-origin redirects only, which is also the fix for the credential-leak shape this bug could have taken with a maliciously-controlled redirect target.&lt;/p&gt;

&lt;p&gt;PR #31 closed a related but separate issue found while reviewing #29: the proxy stripped &lt;code&gt;content-length&lt;/code&gt; from every upstream response, which kills the browser's download progress bar. On a 221MB deliverable, a progress bar with no length to report against looks exactly like a stuck download. The fix is narrower than the original strip — content length only needs to be dropped when the body is being transcoded, not when it passes through byte-for-byte unchanged.&lt;/p&gt;

&lt;p&gt;PR #33 is the sharpest bug in this cluster, and it was found purely by re-reading code that PR #31 had just touched for an unrelated reason — a genuinely different defect sitting one function away from the one being fixed. Response headers were rebuilt using &lt;code&gt;Headers.forEach&lt;/code&gt; plus &lt;code&gt;.set()&lt;/code&gt;, and &lt;code&gt;set-cookie&lt;/code&gt; is the one HTTP header that doesn't comma-fold into a single value — the Fetch API yields each cookie as a separate entry specifically so they don't get merged. Looping with &lt;code&gt;.set()&lt;/code&gt; silently overwrote each prior cookie with the next one, so only the &lt;em&gt;last&lt;/em&gt; upstream cookie ever survived the proxy. Measured directly against a real response: two Set-Cookie headers in, one out.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// before: only the last Set-Cookie survives&lt;/span&gt;
&lt;span class="nx"&gt;upstream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forEach&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;responseHeaders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// .set() overwrites on repeat keys&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// after: append preserves every cookie in the response&lt;/span&gt;
&lt;span class="nx"&gt;upstream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forEach&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;set-cookie&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;responseHeaders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;responseHeaders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PR #34 is the one I'd flag as the most important, because it's not a proxy bug — it's an authentication check that only looked like one. A "fast path" for serving tool sub-resources authenticated requests by testing the &lt;em&gt;name&lt;/em&gt; of a Supabase auth cookie — does it start with &lt;code&gt;sb-&lt;/code&gt; and contain &lt;code&gt;-auth-token&lt;/code&gt; — without ever parsing or verifying the value inside it. Confirmed directly against production: a cookie literally named &lt;code&gt;sb-probe-auth-token&lt;/code&gt; holding a single arbitrary character passed the check, the asset was served, and the request was forwarded upstream with the platform's own internal secret header injected on top. The upstream tools trusted that secret unconditionally, which meant the proxy's cookie-name check was the &lt;em&gt;only&lt;/em&gt; gate in the entire chain — and it was satisfied by a cookie any client could set itself. The fix verifies the token's actual claims offline, without a network round-trip, closing a gap where "looks like an auth cookie" and "is a valid session" had quietly become the same check.&lt;/p&gt;

&lt;h2&gt;
  
  
  Small operational fixes that mattered in practice
&lt;/h2&gt;

&lt;p&gt;Not everything here is a security finding. PR #38 fixed a UX bug reported directly by an editor — the tracker board polled every fifteen seconds by bumping a counter that was part of its data-fetching cache key, so every poll looked like a brand-new request and reset the scroll position, several times a minute, on a page someone was actively working in. PR #36 added a visual distinction the queue was missing: a launch-queue row could land on the board for either of two different reasons — actually approved, or simply because its target date arrived — and both rendered identically, so a ready-to-ship row was visually indistinguishable from one that still needed sign-off. PR #37 turned a dead lint step back on after a framework upgrade had silently stopped enforcing it, and fixed everything it immediately found. PR #41 and PR #54 both fix the same category of trust problem from different ends: a scorecard was visible before its reporting period had actually posted, and a manager's written review and score corrections were silently never reaching the editor they were about — two separate defects sitting behind one hardcoded line of placeholder text.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: a shared proxy multiplies the blast radius of every shortcut
&lt;/h2&gt;

&lt;p&gt;Four real defects — an off-origin redirect follow, a stripped content-length, a cookie-folding bug, and an authentication check that verified a name instead of a value — all sat in the same few hundred lines of proxy code, each found while looking for a different one. That's not a coincidence particular to this codebase; it's what happens when one piece of infrastructure sits in front of everything else. A shortcut in a single tool's own code affects that tool. A shortcut in the layer every tool routes through affects all of them at once, silently, until someone reads that code specifically looking for what else might be wrong nearby. The right response to finding one bug in shared infrastructure isn't just fixing it — it's treating the surrounding code as suspect and reading it again.&lt;/p&gt;

</description>
      <category>security</category>
      <category>accesscontrol</category>
      <category>productengineering</category>
      <category>typescript</category>
    </item>
    <item>
      <title>The Storyboard Layer: Verifying an AI Ad Against Its Reference, Beat by Beat</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:12:18 +0000</pubDate>
      <link>https://dev.to/humzakt/the-storyboard-layer-verifying-an-ai-ad-against-its-reference-beat-by-beat-5hbl</link>
      <guid>https://dev.to/humzakt/the-storyboard-layer-verifying-an-ai-ad-against-its-reference-beat-by-beat-5hbl</guid>
      <description>&lt;p&gt;"Approve" is the button that turns a plan into paid clips. On the main video-generation service, that button sits behind a storyboard gate: every establishing frame gets checked against the reference ad it's swiping — same instruments in frame, same hands occupied or empty, same product visible where the reference shows it. The idea is simple. The implementation spent a month finding out that a check which is itself unverified is just a second place for the bug to hide.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A gate that can be wrong is not a gate. It's an opinion with a veto.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A gate that compared the reference to itself
&lt;/h2&gt;

&lt;p&gt;PR #773 is the cleanest example of the whole month's theme. A check called &lt;code&gt;cutaway-ad-as-talking-head&lt;/code&gt; blocked auto-advance with the message "our plan covers 38% of it with cutaways" — a number that sounded like a real measurement of the plan. It wasn't. &lt;code&gt;gate.ts&lt;/code&gt; supplied &lt;code&gt;referencePersonlessShare(referenceShots)&lt;/code&gt; as the stand-in for "our share," so both sides of the comparison were measuring the &lt;em&gt;reference&lt;/em&gt;, with two different instruments, and the check fired on their disagreement with each other rather than on any actual gap between plan and reference. The plan was never consulted.&lt;/p&gt;

&lt;p&gt;PR #774 is the same failure mode from the opposite direction: the board gate refused to approve a plan it had already paid fourteen frames for ($5.87), because a critic looked at a cutaway render of the pipeline's own device and read it as a different device. The plan's motion description for that panel explicitly called for a translucent variant — the critic was correct that panel 4 looked different from panel 8, and wrong that "different" meant "wrong." A board that "does not agree with itself" is a real signal worth having. It just needs to be checked against what was actually asked for, not against an average of its own panels.&lt;/p&gt;

&lt;h2&gt;
  
  
  A reroll loop that could never close
&lt;/h2&gt;

&lt;p&gt;PR #780 found a genuinely nasty state bug: the board gate's designed remedy for a bad panel is a cheap reroll (~$0.10 via &lt;code&gt;regenerate&lt;/code&gt; with &lt;code&gt;ugc-dyn-frame&lt;/code&gt;) instead of a full re-render (~$1.15). The reroll invalidated the old checkpoint and rendered a new frame — and then never re-graded it. &lt;code&gt;approve&lt;/code&gt; kept reading the stale &lt;code&gt;failed[]&lt;/code&gt; array from before the reroll, so once a panel triggered a failure, the loop had no path back to a clean state. The fix isn't clever; it's just closing the loop the first version left open, which is exactly the kind of bug that only shows up when someone actually exercises the recovery path instead of only testing the happy one.&lt;/p&gt;

&lt;p&gt;PR #785 caught what shipped right behind that fix: after a reroll re-graded correctly for the first time, the board wrote &lt;code&gt;storyboardReport: checked: true, panels: 2, failed: []&lt;/code&gt; on a run with fourteen scenes and fourteen rendered frames. A clean pass over 2 of 14 panels read as a pass on the whole board — on the exact gate guarding roughly $16 of clips per run. Probing all fourteen frames directly showed every one of them was still fetchable and gradable; nothing was actually lost, the report was just wrong about how much of the board it had covered. A short sheet reading as a full pass is worse than a gate that's too strict, because it looks identical to success from the outside.&lt;/p&gt;

&lt;h2&gt;
  
  
  Messages that reported the question, not the answer
&lt;/h2&gt;

&lt;p&gt;PR #782 is a smaller fix with an outsized effect on trust in the system: seven gate messages led with the fixed, code-authored criterion instead of the model's actual finding. The real answer was there — thirty to eighty words later, in a trailing parenthetical the operator had to dig past boilerplate to reach, on the exact 409 response a human reads before deciding whether to override a block. When the thing you show first is the question rather than the answer, every gate reads as more opaque than it is, independent of whether the underlying check is right.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three PRs to find the real ceiling
&lt;/h2&gt;

&lt;p&gt;The token-limit saga across #796, #803, and #805 is worth walking through in full because it's a clean example of iterating toward a root cause instead of stopping at the first plausible fix. A two-hander re-plan (two people in one scene, doubling the planning complexity) truncated at a 16,000-token ceiling. PR #796 raised it to 32,000 — reasonable on its face. It broke every plan call in production within minutes: the Anthropic SDK refuses a non-streaming request whose &lt;code&gt;max_tokens&lt;/code&gt; implies a response that could run over ten minutes, and the run function called &lt;code&gt;messages.create&lt;/code&gt;, not &lt;code&gt;messages.stream&lt;/code&gt;. PR #803 reverted the raise as an emergency fix, with an honest note that its own original reasoning had been wrong. PR #805 found the actual root cause at the SDK source rather than inferring it from docs, and fixed the transport instead of the symptom: stream every call, so the real ceiling becomes whatever the model itself supports, not an arbitrary number picked to dodge a client-side timeout heuristic. The two-hander plan that #796 was trying to unblock finally worked — through the fix that took three PRs to actually name the constraint correctly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Facts that died before the vendor call
&lt;/h2&gt;

&lt;p&gt;PR #781 traces a contradiction all the way to its source: a composite closing shot — presenter, brand mascot, and several product units grouped on a table, mirroring the reference's own end card — was fully wired in the data model and still rendered wrong. The scene's plan data was completely correct: &lt;code&gt;mascotEndCard: true&lt;/code&gt;, the reference board correctly typed as an end-card at that position, the mascot correctly anchored to bookend the reference's own opening shot. The frame that actually rendered was a macro of one closed case, alone. The facts were right at every stage of the pipeline except the last one — they simply never survived into the prompt actually sent to the image model.&lt;/p&gt;

&lt;p&gt;PR #798 found the same family of bug from a rendering angle: a library b-roll span painted over 41.9 to 45.5 seconds of a delivered master's closing composite shot, so the composite showed for about 1.2 of its intended 5 seconds. The root cause: &lt;code&gt;end-card&lt;/code&gt; is a member of &lt;code&gt;CUTAWAY_SHOT_TYPES&lt;/code&gt;, so whenever the reference's own last shot happens to be typed &lt;code&gt;end-card&lt;/code&gt;, the ordinary b-roll span selector treats it as fair game for a cutaway candidate — nothing in that selector knew a composite closing shot needed protection from the exact category it was classified under.&lt;/p&gt;

&lt;h2&gt;
  
  
  A rule that kept overruling its own plan
&lt;/h2&gt;

&lt;p&gt;PR #792 documents a recurring bug in the staging logic: an "empty hands" default rule kept overriding scenes where the plan explicitly described hands as occupied — this time with the presenter's motion reading "slides the closed case into his cardigan pocket and pats it flat," and the rendered frame prompt asserting "BOTH HANDS ARE COMPLETELY EMPTY." The PR notes this is the third recurrence with new verbs, after twelve prior beats across six runs — which is exactly the kind of pattern that signals a detector matching against a fixed vocabulary rather than the actual semantic claim, and each new verb choice in the script slips past it again.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "measure, don't guess" restraint looks like
&lt;/h2&gt;

&lt;p&gt;PR #800 is worth including for what it deliberately doesn't do. Comparing the reference and the pipeline's own output, the reference consistently builds toward the product touching the body — hands massaging cream into skin, an overlay glowing through contact — while the pipeline's beats never show the device making contact at all. The natural fix is to auto-generate a contact beat. The PR ships measurement only: a contact table across all seven reference/output pairs, agreeing exactly on the gap, with an explicit note that auto-generating the missing beat is a separate, riskier decision than measuring that it's missing. Not every finding needs to become a feature in the same PR that discovers it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: verify the verifier
&lt;/h2&gt;

&lt;p&gt;Every fix in this cluster is one layer removed from the bug you'd expect. The storyboard gate isn't "buggy" in the sense of producing wrong grades on correct inputs most of the time — it's buggy in the sense that the machinery grading the ad needed exactly as much scrutiny as the machinery generating it, and for a while got less. A gate that compares the reference to itself, a critic that can't distinguish its own product's variants, a reroll that never re-grades, a message that shows the question instead of the answer — none of these are algorithmically hard. They're the direct consequence of a verification layer being trusted the moment it started reporting failures, instead of being measured with the same discipline applied to the thing it verifies.&lt;/p&gt;

</description>
      <category>aivideo</category>
      <category>videogeneration</category>
      <category>promptengineering</category>
      <category>computervision</category>
    </item>
    <item>
      <title>Pin One Hook Into Every Variant: Fixing a Hook-Management Builder</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:12:15 +0000</pubDate>
      <link>https://dev.to/humzakt/pin-one-hook-into-every-variant-fixing-a-hook-management-builder-4pj</link>
      <guid>https://dev.to/humzakt/pin-one-hook-into-every-variant-fixing-a-hook-management-builder-4pj</guid>
      <description>&lt;p&gt;The ad-variant export tool exists to reassemble sped-up ad variants with a swappable opening "hook" — an editor picks a fast-cut visual, pairs it against a library of alternate hooks, and the tool builds every combination as its own export. Eight PRs over three weeks fixed it from "editors keep reporting the same class of bug in different clothes" to something closer to what they'd actually asked for — and the most interesting finding wasn't a UI bug at all, it was a test suite that had been silently asserting nothing for months.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A green test suite is a claim about what ran, not about what passed. Check the first thing before trusting the second.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A naming bug that erased the library
&lt;/h2&gt;

&lt;p&gt;PR #1 opens with the kind of bug that looks impossible until you see the exact ordering: &lt;code&gt;_clip_from_file&lt;/code&gt; ran &lt;code&gt;parse_sop()&lt;/code&gt; &lt;em&gt;before&lt;/em&gt; it read the file's own stamped slot-type metadata. Any filename shaped like &lt;code&gt;&amp;lt;adname&amp;gt;_&amp;lt;product&amp;gt;_&amp;lt;format&amp;gt;&lt;/code&gt; — which is exactly how the tool names its own short leads, e.g. &lt;code&gt;SL872_PRO99_916.mp4&lt;/code&gt; — got reclassified as a finished ad purely from its filename shape, before the code ever checked the field that actually said what it was. Every short lead an editor exported vanished from the builder's own library, including leads the tool had generated itself moments earlier. The same PR also ported hook reordering and per-ad prelead handling from the tool's sibling "Export" repo, and fixed archiving that had been silently unreliable — three real fixes bundled together, with a fourth request (adding a product tracker) explicitly scoped out to a different repo rather than absorbed here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pin one hook, not five clicks
&lt;/h2&gt;

&lt;p&gt;PR #2 tackled the feature editors actually wanted: apply one hook across an entire batch of variants in a single action, instead of dragging it into each one by hand. The PR also fixed something unrelated that had been quietly broken for a while — the preview panel stopped updating after the first build of a session, because &lt;code&gt;recompute()&lt;/code&gt; short-circuited into the finished-job list whenever a prior batch's job list was non-empty, and nothing ever cleared that list once a job finished. Every keystroke, every product change, every rearrangement after that first build just repainted the old result. The only way to see a new arrangement was to commit to another full render — verifying your edit by paying for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Built the wrong shape, twice — on purpose, not by accident
&lt;/h2&gt;

&lt;p&gt;PR #4 is a clean example of shipping something reasonable that still isn't what was asked for, and treating that as a spec problem rather than a regression. What landed in #2 was "pin to all" — one hook joins every other variant. What the editor actually needed was a &lt;em&gt;subset&lt;/em&gt;: one specific hook applied to three of five variants, not all five. The fix — tick the hooks that should get it, then &lt;em&gt;Add a hook to each of 3…&lt;/em&gt; — is asserted twice: once against the exact expected filename list in the unit suite, and again by driving the real browser controls end to end, so the assertion can't quietly stop matching what a user actually clicks.&lt;/p&gt;

&lt;p&gt;PR #5 found the same feature was still subtly wrong one layer down. The picker that chose which clip to apply listed individual &lt;em&gt;clips&lt;/em&gt;, not &lt;em&gt;lanes&lt;/em&gt; — so if the source hook was a two-part opener, picking "one of its clips" split the pair in half: one half joined the target variants, the other was left behind as an orphaned alternate nobody asked for. Measured on a real five-hook batch, the fix took the lane count from six back down to the correct five. The picker now offers lanes, not clips, which makes "half an opener" structurally impossible to select.&lt;/p&gt;

&lt;h2&gt;
  
  
  A test suite that reported success having asserted nothing
&lt;/h2&gt;

&lt;p&gt;PR #3 is the one worth sitting with longest. Two follow-up audits after #2 shipped — an accessibility/UX pass and a speed-mode pass — surfaced a batch of an editor's work disappearing entirely on a mid-task mode switch, and table rows wrapping awkwardly throughout the interface. But the finding that undermines trust in everything shipped before it is this: &lt;code&gt;test_builder_e2e.py&lt;/code&gt; could report success having verified nothing at all. Flask auto-loads a local &lt;code&gt;.env&lt;/code&gt; file, and on any machine where &lt;code&gt;HUB_SHARED_SECRET&lt;/code&gt; happened to be set, every request the test suite made came back &lt;code&gt;403&lt;/code&gt;. That 403 arrived as an &lt;code&gt;HTTPError&lt;/code&gt; — a subclass of &lt;code&gt;URLError&lt;/code&gt; — so the suite's own retry loop caught it, printed a friendly &lt;code&gt;SKIP: the Flask dev server did not come up&lt;/code&gt;, and exited zero. Green. The suite that had been specifically added because a DOM stub is blind to real layout was itself blind, on exactly the machines most likely to have a shared secret configured — which is to say, on anything resembling a real environment. The fix forces the secret empty for the suite's own server and distinguishes a closed connection from an actual HTTP rejection, so a real failure can no longer disguise itself as an environment quirk.&lt;/p&gt;

&lt;h2&gt;
  
  
  Small frictions, fixed for the reason an editor gave
&lt;/h2&gt;

&lt;p&gt;PR #6 closed a naming inconsistency that had been open since an earlier task: two different ways of including a hook in a batch — "include in every hook" and "add a hook to each" — defaulted to opposite ordering, producing two different filenames for what editors meant as the same operation. The fix wasn't a preference call; it was resolved from the editors' own worked example, and the reasoning is concrete: &lt;code&gt;buildSopName&lt;/code&gt; emits the timeline in reverse, so a code placed &lt;em&gt;last&lt;/em&gt; in the name plays &lt;em&gt;first&lt;/em&gt; in the render — meaning "plays-first" was the only default that made both routes agree.&lt;/p&gt;

&lt;p&gt;PR #7 investigated a report that the library picker "reloads" every time a filter is touched, and the finding was that it never actually reloads — the listing is fetched once per mode and cached, and a filter click issues zero network requests. What was real: reopening a cached library flashed a loading state it didn't need, and one specific filter path was needlessly rebuilding 480 row nodes in three seconds for a change that touched none of them. Three plausible-sounding causes, only one of which was doing real work — the kind of bug report that's only solvable by actually instrumenting the claim instead of guessing at it.&lt;/p&gt;

&lt;p&gt;PR #8 closed the loop on a claim that turned out to be about a missing feature, not a broken one. A report read "apply hook to all — this is not applied," but the feature had been live and tested since #4. What was actually missing was the specific &lt;em&gt;route&lt;/em&gt; to it that the sibling Export tool already had: a &lt;code&gt;⧉ Duplicate&lt;/code&gt; gesture, stacking each copy with its visual. Rather than re-invent that interaction from scratch, this PR ported the row controls that already existed one repo over — the same workflow, available here for the first time, built the way it had already been proven to work elsewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: check what the test actually exercised
&lt;/h2&gt;

&lt;p&gt;Most of these bugs are unremarkable individually — a metadata read in the wrong order, a stale short-circuit, a picker granularity mismatch. What ties them together is that nearly every fix in this batch came from checking a claim against reality rather than trusting that a passing test, a shipped feature, or a plausible-sounding bug report was accurately describing what was actually happening. The test suite that reported green while testing nothing is the sharpest version of that lesson, but PR #2's stale preview, PR #7's "it reloads" report, and PR #8's "it's not applied" report are all the same shape: the thing that looked true on the surface wasn't, and the fix in every case started with measuring, not assuming.&lt;/p&gt;

</description>
      <category>frontendengineering</category>
      <category>videoproduction</category>
      <category>typescript</category>
      <category>ux</category>
    </item>
    <item>
      <title>Retiring a Drive Folder Scan: A Queryable Run Index and a Ledger That Doesn't Lose the Delta</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:11:41 +0000</pubDate>
      <link>https://dev.to/humzakt/retiring-a-drive-folder-scan-a-queryable-run-index-and-a-ledger-that-doesnt-lose-the-delta-4ae3</link>
      <guid>https://dev.to/humzakt/retiring-a-drive-folder-scan-a-queryable-run-index-and-a-ledger-that-doesnt-lose-the-delta-4ae3</guid>
      <description>&lt;p&gt;&lt;code&gt;/runs&lt;/code&gt; had no queryable store. The only durable copy of a run's state was one &lt;code&gt;job-state.json&lt;/code&gt; file per run, sitting in that run's own Google Drive folder — so rendering a table of thirteen fields meant enumerating every Drive folder and downloading every run's entire job document to read them. Measured: a mean of 65 KB and a max of 445 KB per run, for a table that showed thirteen columns. The in-memory cache that was supposed to make this bearable was module-level, so the first load after every deploy returned an empty list. A sibling tool's own schema file had already named the smell: "read live state from a 5-minute-stale Drive folder scan."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Four of this round's defects were mine, and every one was found by measuring production rather than by reading the diff.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's not a stray line — it's the author's own summary of the nine-PR round that closed this migration out. It's worth sitting with, because it sets the tone for everything below: replacing a Drive scan with a real index sounds like a clean infrastructure swap, and it produced instead a small, honest cluster of self-found production bugs, several of them shipped by the same person who then caught them.&lt;/p&gt;

&lt;h2&gt;
  
  
  A queryable index, kept deliberately fail-open
&lt;/h2&gt;

&lt;p&gt;PR #890 built a Supabase-backed run index as a projection, and made one important choice explicit: the old Drive-scan path didn't get deleted yet. It stayed as a fallback if Supabase failed, and — just as useful — as a free oracle to diff the new projection against. Two systems both answering "what runs exist" isn't waste when one of them exists specifically to catch the other being wrong.&lt;/p&gt;

&lt;p&gt;That oracle earned its keep. Run twice over the same production data, the comparison came back &lt;code&gt;$0&lt;/code&gt; both times: 293 of 293 records identical, then 290 of 290, with job-ID sets matching exactly and 14 of 16 fields byte-identical on every row. The one legitimate difference was a projector-written row that was a few hours fresher than its frozen Drive snapshot — which is the index being &lt;em&gt;right&lt;/em&gt;, not the index disagreeing. Only once that verification held up did PR #928 retire the Drive scan and make the index the single source of truth, closing the issue that had opened the whole migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  A cache that became a write's merge base by accident
&lt;/h2&gt;

&lt;p&gt;The migration's first real data-loss bug arrived immediately, in PR #892, and it's a sharp lesson in how a read path and a write path can silently start sharing state that was never designed to be shared. Every catalog writer in the pipeline is a read-modify-write: read the whole index, merge one record into it, upload the merged result back over the Drive copy. PR #890 had made the public &lt;code&gt;listProfiles()&lt;/code&gt;/&lt;code&gt;listRips()&lt;/code&gt; functions prefer the new Supabase mirror — and those were exactly the functions the writers used as their merge base. The mirror write was deliberately fail-open, so it could silently fail while the Drive upload succeeded, leaving the mirror stale. The next writer would then read that stale mirror, merge its own change in, and upload the result over Drive — permanently erasing whatever had changed in between. A cache that's allowed to be wrong is fine for reads. It is never allowed to be the base of a write.&lt;/p&gt;

&lt;p&gt;The same shape reappeared for a second catalog a few days later. PR #910 found that &lt;code&gt;setLabel&lt;/code&gt; — the writer behind display-name renames — re-read the current state before writing, which is the correct instinct against a stale cache. But its re-read path, &lt;code&gt;ensureLabelsFresh&lt;/code&gt; plus &lt;code&gt;currentLabels()&lt;/code&gt;, could not report failure: it failed open on a Drive error, and on a cold container with no existing cache, its fallback was an empty object. A cold container plus one failed Drive request was enough to upload an effectively-empty map over every rename anyone had ever made. PR #924 closed it by mirroring the display-name index into the same Supabase system as the others — the fifth and last of five app-written Drive-JSON indexes, and the only one still reading Drive on a cold start.&lt;/p&gt;

&lt;h2&gt;
  
  
  A run's real spend, not just its final-video cost
&lt;/h2&gt;

&lt;p&gt;PR #893 is the kind of bug that's invisible until someone reads a real number. &lt;code&gt;total_cost_usd&lt;/code&gt; summed &lt;code&gt;models[].costUsd&lt;/code&gt; — the per-model cost of a &lt;em&gt;finished&lt;/em&gt; video. A run that paused at a review gate, or failed before producing a final video, reported &lt;code&gt;$0&lt;/code&gt; spent no matter how much detection and planning it had actually paid for. This was found by watching the very first row the new projector wrote in production: a run parked mid-pipeline, genuinely billed, reading as free. The "spend by week" query shipped in the same schema summed exactly that broken column.&lt;/p&gt;

&lt;p&gt;The deeper fix landed in PR #930, and its own framing is the cleanest summary: "the delta was in hand and thrown away." The run's live ledger — &lt;code&gt;job.costSoFarUsd&lt;/code&gt; plus a list of cost notes — lived only in-memory, persisted solely inside the Drive snapshot. Giving the run &lt;em&gt;total&lt;/em&gt; a durable home (PR #893) wasn't enough, because a total can't answer "which specific charge was this for" — and that's exactly the question both of this round's documented money bugs turned on. The function that recorded every charge, &lt;code&gt;addCost&lt;/code&gt;, already computed the exact dollar delta for each one. It just wrote that number to a log line and discarded it instead of persisting it anywhere durable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// before: the delta existed for one log line and nowhere else&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;addCost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Job&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;deltaUsd&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;note&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;costSoFarUsd&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;deltaUsd&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`[cost] +$&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;deltaUsd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toFixed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt; — &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;note&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="c1"&gt;// deltaUsd is gone the moment this function returns&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// after: the same number, persisted per-charge&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;addCost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Job&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;deltaUsd&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;note&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;costSoFarUsd&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;deltaUsd&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;ledger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;record&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;runId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;deltaUsd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;note&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PR #907 used that same underlying accounting to answer a question that had never had a real number attached to it: when an operator cancels a wedged run, how much money are they walking away from? Before this, "it will stop and be marked failed" was the entire message — no figure, no way to distinguish killing a forty-cent run from killing a fourteen-dollar one. A dedicated module now derives three figures — spent, committed, and not-yet-started — specifically so the number that actually matters for a cancel decision has somewhere to be shown.&lt;/p&gt;

&lt;h2&gt;
  
  
  The age-out bug that emptied the dashboard's most important group
&lt;/h2&gt;

&lt;p&gt;This is the round's most instructive regression, because it's a bug introduced by a refactor that was itself a fix. PR #925 moved the stalled-run age-out logic from a client-side module into the server-side run-index reader — a legitimate move, made its own PR specifically because the team's own rule says a deletion must never share a PR with a behavior change. But the move carried the threshold across and quietly dropped the &lt;em&gt;ordering&lt;/em&gt; of the checks — and the ordering turned out to be the whole rule.&lt;/p&gt;

&lt;p&gt;The display status of a run is decided in three sequential steps: terminal state, then review-gate state, then freshness/staleness. Ageing a row out on the server rewrote its status to a terminal value directly, which meant the review-gate check downstream could never fire for it. A run sitting at a review gate isn't stalled — it's waiting for a human, has no running process that could have died, and its idle time carries no information at all. PR #927 caught the consequence in production: the age-out was emptying the dashboard's "needs you" group, silently hiding exactly the runs an operator most needed to see. The fix wasn't reverting the move — it was restoring the order the original code had gotten right by accident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deleting the code that used to be true
&lt;/h2&gt;

&lt;p&gt;Alongside the index migration, a separate but related discipline ran through the same weeks: finding and removing code nobody calls anymore, using tools rather than instinct. PR #851 used the repo's own call-graph script to find two genuinely dead functions among thirty-four flagged candidates — most of the rest were framework callbacks the tool correctly knows to warn about rather than delete. PR #856 found three scratch files sitting at the repo root, invisible to every guard because the guard's glob only matched &lt;code&gt;.ts&lt;/code&gt;/&lt;code&gt;.tsx&lt;/code&gt; — the fourth time in this effort that a coverage gap, not a threshold, was the actual defect.&lt;/p&gt;

&lt;p&gt;PR #906 is the capstone: two commits, the first clearing 1,747 dead import bindings left behind by an earlier file-splitting migration (every extracted file had kept importing what its pre-split self used, and nothing in the existing lint config could see the leftovers), the second turning on &lt;code&gt;no-unused-vars&lt;/code&gt; so the count can't silently grow back. Fixing the tooling itself surfaced real defects: a file-wide bracket cleanup regex had turned &lt;code&gt;insertShape: shape,\n}&lt;/code&gt; into a syntactically valid but semantically broken line across 162 files, and three separate source-scanning tools had been quietly missing files because of it. The sweep found four real bugs this way — not by someone going looking for bugs, but by making 1,897 pieces of dead weight visible enough that the four live wires tangled up in them couldn't hide anymore.&lt;/p&gt;

&lt;p&gt;Not every "dead code" candidate survives contact with measurement, and that's worth stating plainly rather than glossing over. PR #973 audited eight items flagged as unread and found five were load-bearing — one route had four live callers serving labels to three different surfaces, one field appeared in eight separate schema assertions governing what a forked run inherits. The count of what's actually safe to delete keeps shrinking under real inspection, which is the entire point of measuring before cutting rather than after.&lt;/p&gt;

&lt;h2&gt;
  
  
  A security deletion, not a hardening
&lt;/h2&gt;

&lt;p&gt;PR #921 is a smaller finding with an outsized blast radius if it had gone unnoticed longer: a legacy image proxy sat on a route excluded from the hub's auth gate, ostensibly so an image-editing model could pull frames by URL. In practice, it fetched an arbitrary URL out of job state and streamed the response body back — an unauthenticated request relay running from the service's own egress IP, returning someone else's content under this origin. Measured before deleting rather than assumed: of 86 archived job-state documents, exactly two carried the field this route needed, and every URL in both pointed at the service's own host. The fix was deletion, not hardening, because by the time anyone looked, the branch was already dead — which didn't make it any less of a live vulnerability while it existed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: infrastructure migrations create their own bug class
&lt;/h2&gt;

&lt;p&gt;Nothing in this cluster is a hard algorithm. A cache became a write's merge base because two independently-reasonable pieces of code started sharing state neither was designed to share. An age-out check regressed because a refactor preserved the threshold and silently dropped the order it ran in. A cost total was wrong because "spent" and "billed for a finished video" look identical until a run stops partway through. Every one of these is the specific bug class that infrastructure migrations create: not the new system being wrong, but the seam between the new system and everything that still assumes the old one — caches, readers, orderings, assumptions about what a field means — being wrong in ways that only show up once real production data runs through it. The fix, consistently, was the same discipline repeated four times: measure the live behavior, not the diff.&lt;/p&gt;

</description>
      <category>costengineering</category>
      <category>architecture</category>
      <category>codequality</category>
      <category>database</category>
    </item>
    <item>
      <title>Building a Reference-Less Mode: When There's No Video to Mirror</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:11:38 +0000</pubDate>
      <link>https://dev.to/humzakt/building-a-reference-less-mode-when-theres-no-video-to-mirror-1khc</link>
      <guid>https://dev.to/humzakt/building-a-reference-less-mode-when-theres-no-video-to-mirror-1khc</guid>
      <description>&lt;p&gt;The main video-generation service had always been built around one assumption: somewhere in every run sits a reference video — a competitor's ad, or a past winner — that the plan gate measures the output against. Scene count is checked against the reference's cut count. Runtime against its runtime. Narration pace against its pace. Every one of those checks has a right-hand side because a reference always supplied one. Then an editor asked for a mode with no reference at all — just a script, because "your script is the whole brief." That single missing input turned out to be load-bearing under nearly every quality gate in the pipeline, and finding each place it was load-bearing took running real, paid jobs rather than reading the code.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A check that has quietly depended on an input for months doesn't fail loudly when that input disappears. It reports a pass — because zero minus zero still reads as "nothing lost."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Blocked from the day it shipped
&lt;/h2&gt;

&lt;p&gt;PR #915 is the mode's origin story and its first bug in one PR. The route's own header read "No reference needed — your script is the whole brief," and the screenshot showed every step ticked, including a green "No reference" — directly above an error reading "missing referenceVideo file or link." The reject had been added to the request parser six weeks before the reference-less mode existed, and nobody had gone back to teach the parser about the new case when the mode shipped. The mode had never worked, not once, from the moment it went live.&lt;/p&gt;

&lt;h2&gt;
  
  
  A word budget with no speaker to measure
&lt;/h2&gt;

&lt;p&gt;Once runs could actually start, PR #949 found every single reference-less plan logging the same warning on every run: four scenes over their word budget, every time, no exceptions. A warning that always fires teaches people to ignore the one time it matters. The cause was a clamp — &lt;code&gt;budgetWps = min(refWps, ours)&lt;/code&gt; — where &lt;code&gt;refWps&lt;/code&gt; silently fell back to a default whenever the reference's speaker hadn't been measured. On a swipe, that fallback is load-bearing: a real reference sits behind it most of the time. On a reference-less run, there is no reference speaker to measure, ever, so the fallback wasn't a fallback anymore — it was the only value the budget could ever see, quietly capping every script-only run at the same generic delivery rate regardless of what the actual voice model could speak.&lt;/p&gt;

&lt;h2&gt;
  
  
  A prompt that argued with itself
&lt;/h2&gt;

&lt;p&gt;PR #948 is the fix I'd point to as the cleanest example of the whole arc's failure shape. The reference-less planner prompt opened with an explicit, correct instruction: "THE SCRIPT IS THE ONLY SOURCE OF STRUCTURE — there is no reference ad, and nothing about one is being withheld from you." Measured by actually building the prompt sent to the model — not by reading the template that assembles it, which the team's own engineering notes insist is the only way to know what a model actually receives — twenty-some lines later the same prompt instructed the planner to mirror a reference anyway. A leftover block from before reference-less mode existed had never been removed, and it directly contradicted the instruction two paragraphs above it. The planner wasn't ignoring a correct instruction. It was being given two instructions and picking one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The green tick that meant nothing
&lt;/h2&gt;

&lt;p&gt;PR #955 found the sharpest version of the pattern running through this whole arc. On a real job with no reference, the plan gate displayed: "✅ Script fills the reference — No issues found." The underlying check compared the plan's scene count against the reference's cut count. There was no reference, so the right-hand side was zero. &lt;code&gt;lostCuts = max(0, 0 - planScenes)&lt;/code&gt; evaluates to zero no matter what the plan actually contains, and zero reads as "lost nothing" — a structurally guaranteed pass dressed up as a measurement. The team's own documentation makes the same point about a different guard elsewhere in the codebase: answering "no file is over the line-count cap" while several files actually are is worse than answering nothing at all, because the number reads as evidence when it's actually an artifact of nothing being measured. The fix doesn't make the check smarter — it makes an unmeasurable check say &lt;code&gt;unmeasured&lt;/code&gt; instead of quietly manufacturing a pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four near-identical shots in one room
&lt;/h2&gt;

&lt;p&gt;PR #954 is where the fixes so far collided with a genuinely new problem rather than a leftover assumption. A verification render, once the prompt no longer contradicted itself, came back with a presenter addressing the camera, two cutaway spans, and passing identity checks — and also four near-identical seated medium close-ups in one room, with internal-cut detection reporting no cut inside any clip. Forty-six percent of a 28-second ad was one unbroken thirteen-second take. Removing the contradictions had fixed &lt;em&gt;who&lt;/em&gt; the ad addressed. Nothing had ever told the system &lt;em&gt;how&lt;/em&gt; to shoot it, because on a swipe that answer always came from the reference's own cuts — there was no equivalent instruction for a mode with nothing to cut against. The fix is a beat grammar: an explicit shot-framing vocabulary keyed off the ad's declared format, so the system has somewhere to look for pacing and composition that isn't a reference video.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cast from speaker labels, not a vision pass
&lt;/h2&gt;

&lt;p&gt;PR #956 closes the same gap one layer up, on casting rather than shot framing. A swipe proposes its cast by running a vision pass over the reference and counting distinct speakers. A reference-less run has no video to run a vision pass over, so &lt;code&gt;castProposal&lt;/code&gt; was simply never written, and the cast gate silently offered one solo slot on every run — even though the mode's own chooser told editors, in as many words, that the system plans the scenes, the cast, &lt;em&gt;and the shots&lt;/em&gt; from the words. Two of those three were true. The fix proposes cast directly from the script's own speaker labels, which was information the input already contained and nothing had been reading.&lt;/p&gt;

&lt;h2&gt;
  
  
  Proving a negative for $0
&lt;/h2&gt;

&lt;p&gt;PR #947 is worth calling out on its own, because of what it cost to find the problem it prevents recurring. Every existing diagnostic instrument in the pipeline — the plan gate, the fork replay, the free fork-scope quote, the render fleet — starts from either a reference video or an existing run. A script-only plan is neither, so the only way to see what one would produce was to start a real job and pay for it: $14.68, in this case, to discover the planner-prompt contradiction that PR #948 later fixed. The response was a standalone script that replays the real sentence-splitting and scene-count logic against a script file directly, for $0, plus a small corpus of known scripts that pin down what a correct beat count should look like — so the next time this specific question comes up, it doesn't cost fourteen dollars and a full render cycle to answer it again.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same defect, six times
&lt;/h2&gt;

&lt;p&gt;By the time the round's handoff was rewritten, what had started as "the planner still writes near-identical shots" had grown from four confirmed instances of the same root defect to six. Every one of them was the identical shape: a check, a budget, or a proposal step that had quietly depended on a reference being present, silently degrading to a default, a fallback, or a false pass the moment that reference stopped existing. None of them threw an error. All of them looked, from a passing test suite, like working code. The fix in each case wasn't cleverness — it was tracing one more input all the way from "the reference supplies this" to "what happens when nothing does," and refusing to let a missing measurement read as a good one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern
&lt;/h2&gt;

&lt;p&gt;Building a genuinely new mode into a pipeline that has one deeply-ingrained assumption baked into it — here, "a reference video exists" — is less about writing the new mode's own logic and more about auditing every place downstream that assumption quietly leaked into. The parser assumed it. The word budget assumed it. The planner's own prompt assumed it, twice, contradicting an instruction two lines away. The plan gate assumed it hard enough to manufacture a false pass out of a zero-versus-zero comparison. None of those were bugs in the new mode. They were bugs in the old assumption, invisible for as long as the assumption held, and each one only surfaced once a real, paid run went looking for the answer instead of trusting that a green check meant what it said.&lt;/p&gt;

</description>
      <category>aivideo</category>
      <category>promptengineering</category>
      <category>productengineering</category>
      <category>aiplanning</category>
    </item>
    <item>
      <title>Letting Operators Request Animated People Without Letting a Reference Do It By Accident</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:11:04 +0000</pubDate>
      <link>https://dev.to/humzakt/letting-operators-request-animated-people-without-letting-a-reference-do-it-by-accident-4a47</link>
      <guid>https://dev.to/humzakt/letting-operators-request-animated-people-without-letting-a-reference-do-it-by-accident-4a47</guid>
      <description>&lt;p&gt;The main video-generation service had never had a real concept of "animated." It had a side effect: if a reference ad happened to be heavily CGI, the pipeline's presenter sometimes came back looking like a cartoon too, because "animated" was derived silently from how much of the reference was computer-generated and used to drive rendered inserts and rendered people together, as one undifferentiated signal. When an operator actually wanted a deliberately animated, Pixar-style ad — not an accident of a CGI-heavy reference, a real creative choice — that capability didn't exist. It had, in fact, just been removed, one dismantled channel at a time, by earlier fixes aimed at stopping presenters from accidentally turning into cartoons. Building it back as an intentional feature, rather than restoring the accident, is what this entire arc is about.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Two questions that look like one question — "should the medium be photoreal or animated" and "who is allowed to decide that" — need two different answers, and conflating them is how a feature disappears while every individual fix along the way looks correct.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The design principle: an operator may ask, a reference may not
&lt;/h2&gt;

&lt;p&gt;PR #844 states the governing rule directly in its own title: an operator may ask for animated people; a reference may not. Before this PR, "animated" was never a separate, nameable choice — it was derived purely from the reference's measured CGI share and used to drive both non-human inserts and on-camera people from the same signal. On a reference that measured 43% CGI, the presenter kept coming back rendered as a cartoon, and the fix each time had been to dismantle another one of the four channels that could trigger it — across three separate earlier PRs, ending with the character-label channel closed in #834. Each of those fixes was individually correct: it stopped an unwanted cartoon presenter on that specific reference. Collectively, they had also deleted the only path by which an operator who genuinely wanted an animated ad could get one, and nobody had made that trade-off on purpose.&lt;/p&gt;

&lt;p&gt;The reasoning behind PR #834 itself is worth sitting with, because it's a clean example of verifying against the actual artifact rather than the code's intent. Every one of seven person-beats in a delivered video had asked the image model, in the persisted prompt, for the subject to be "rendered as a high-fidelity, photoreal 3D-animated character that preserves his likeness precisely" — while, in the very same request, a separate rule promised "this shot is the PHOTOGRAPHED half ... the photograph underneath stays a photograph." The request was contradicting itself inside a single call, in exactly the clause that establishes a person's identity, and the only way to find that was reading the actual prompt the model received, not the code that assembled it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A negative prompt asking the wrong question
&lt;/h2&gt;

&lt;p&gt;PR #849 found a downstream consequence of that same medium/people conflation. A negative prompt meant to suppress the visual failure modes of a &lt;em&gt;photoreal human&lt;/em&gt; render — "smooth plastic skin, exaggerated gestures" — was being stripped whenever the run's medium was animated, based on a check that read only &lt;code&gt;medium.medium&lt;/code&gt; and ignored &lt;code&gt;medium.source&lt;/code&gt;. Before the operator-choice split landed, medium and "are there animated people" really were the same fact, so that shortcut was harmless. After #844, they can disagree — a mixed ad can have an animated medium with photoreal people, or vice versa — and the negative prompt needs to answer the people question specifically, not the medium question by proxy. A term that used to be a safe stand-in for another term stopped being safe the moment the feature it was standing in for became independently choosable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verifying with a controlled pair, not one lucky reference
&lt;/h2&gt;

&lt;p&gt;PR #857 is the handoff that reports the animated lane verified in both directions, and the verification method is worth naming: the same reference, avatar, voice, and lines launched twice on the identical build, changing only whether the medium was pinned as operator-requested-animated or reference-derived-animated — with both pins confirmed &lt;em&gt;before&lt;/em&gt; any money was spent, not inferred afterward from the output. That's a controlled A/B on a single variable, which is a stronger claim than "we tried it and it looked right," because it rules out the reference's own content as the explanation for whatever came back.&lt;/p&gt;

&lt;h2&gt;
  
  
  A face-check built for a face that was never photographed
&lt;/h2&gt;

&lt;p&gt;PR #863 found that the identity-drift critic — the one that verifies a rendered face actually matches the reference face — was asking the wrong question on a genuinely operator-animated run. Its predicate for "does this beat show live action" consulted only what the &lt;em&gt;reference&lt;/em&gt; had filmed, never the newer &lt;code&gt;peopleMayBeAnimated&lt;/code&gt; flag. On an operator-driven Pixar run, the pipeline itself is rendering the presenter as a character in every beat — there is no photographed face anywhere in that run for a face-similarity model to compare against. The check still ran, scored a drawn face against a real photograph, and reported the mismatch as a defect. A check firing on an input it has no meaningful answer for is a distinct failure mode from a check firing correctly on the wrong input — the first one needs a "this doesn't apply" branch, not a bug fix to its comparison logic.&lt;/p&gt;

&lt;p&gt;PR #882 is the companion fix on the generation side of the same seam: a Pixar run's identity anchor — the reference image the model is told to match a face against — was still a real photograph, even though every text channel in the prompt correctly described an animated character. Every sentence was right and the frames still came back as photoreal CG humans, inconsistently across scenes, because the one non-text signal in the request — the anchor image itself — was still telling the model "match this specific real face," which pulls toward photorealism regardless of what the words around it say. The fix swaps the anchor to an actual generated Pixar-style character reference once one exists, so the image evidence and the text instruction finally agree instead of quietly fighting each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  A clause that never reached the manifest it was shipped to prove
&lt;/h2&gt;

&lt;p&gt;PR #876 caught the first real operator-animated render coming back photoreal-leaning rather than stylised, and the useful part of the fix isn't the bug itself — it's the diagnostic discipline in how it was isolated. The finished scene manifests listed camera, shot size, posture, wardrobe, stature, and setting, and said nothing at all about medium, on the one run whose entire purpose was proving the medium reaches the picture. The PR explicitly separates two different claims that a symptom like this can support — "the clause never arrived" versus "the clause arrived and the model disobeyed it" — because the fix for each is completely different, and conflating them wastes a diagnostic cycle chasing the wrong one. PR #878 found the same class of gap one layer over: the clip critic that judges a finished ad had never once referenced &lt;code&gt;peopleMayBeAnimated&lt;/code&gt; across any of its four modules, meaning it was asking "are these people drawn" using the wrong signal on a mixed-medium ad where some people are animated on purpose and some aren't.&lt;/p&gt;

&lt;h2&gt;
  
  
  A body-effect that authored zero times
&lt;/h2&gt;

&lt;p&gt;PR #932 shipped an "anatomical glow" overlay — a stylised body-part highlight the reference used on six of fourteen shots and a human editor's own manual recreation used on thirteen of twenty — behind a flag, defaulted off, specifically because it's a genuinely re-cast creative effect rather than a literal copy of the reference. PR #936 caught, at the plan-gate verification step and before a single frame was rendered, that the effect had authored on &lt;em&gt;zero&lt;/em&gt; of the beats that should have carried it. The body-site lookup asked the storyboard graph's entity list which body part to target, and that entity list — by construction, elsewhere in the pipeline — only ever contained the literal string &lt;code&gt;"mascot"&lt;/code&gt;, never an actual body-part name the glow logic could match against. The feature was fully built, flagged, and structurally incapable of authoring on any input, caught before it could waste a single paid render finding that out. PR #938 found a second, more subtle version of the same class of bug one step later: the glow's face-box coordinates were being converted using the render target read live at &lt;em&gt;paint&lt;/em&gt; time, while a run's actual aspect ratio is pinned once at first dispatch specifically so later code doesn't re-resolve it — meaning the math was only correct by coincidence, whenever the live setting happened to still match the pin.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verified, retested, and still not quite right
&lt;/h2&gt;

&lt;p&gt;PR #1012 is a real editor's second-round retest of an earlier fix round, and it's a good example of distinguishing "the fix didn't ship" from "the fix shipped and is incomplete." The Pixar character-generation machinery itself worked correctly — billed five cents, mirrored, byte-for-byte the animated file rather than the photograph — but a "look" clause still needed to reach the actual frame-generation call, and the character's face needed rounding to read as more stylised. PR #1013, found by literally running the new cutting-pace control rather than trusting that it was wired, discovered the control could report success on a pass that changed nothing — the same beat count came back paced and unpaced, correctly, because the math genuinely produced an identical result on that particular script length, and the fix is making the tool say "this had no effect" explicitly rather than let a no-op read as an applied change. PR #1015 found the visual-style clause reaching every render except the manifest that's supposed to record it, on the one render whose entire purpose was proving the look reaches the picture — the same "arrived but unrecorded" shape as #876, recurring on a different clause. And PR #1018, the most recent in this arc, found that an animated ad's &lt;em&gt;generated&lt;/em&gt; cutaways carried the run's style correctly, while cutaways pulled from the house b-roll bank of real filmed footage did not — a real human hand holding a real product, cut directly between cartoon beats, because the guard governing which cutaway source is allowed had never been taught that an animated run needs to exclude live-action b-roll specifically, not just avoid animating people who were never in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: a capability and a decision-right are two different features
&lt;/h2&gt;

&lt;p&gt;Nearly every bug in this arc is the same shape wearing different clothes: something that looks like one flag — "is this animated" — is actually two separable facts, "what medium should this render in" and "who gets to decide that," and treating them as one value is what let the capability disappear in the first place, then let it come back in pieces that individually worked and collectively didn't reach the screen. The identity anchor, the negative prompt, the face-check predicate, the body-site lookup, the b-roll source guard — each one had its own copy of "does animated apply here," derived a different way, and each had to be taught the same operator-vs-reference distinction separately before the feature was actually whole. Naming the decision-right explicitly, in #844, is what turned a scattered set of ad hoc "is this a cartoon" checks into one governing rule the rest of the pipeline could be checked against.&lt;/p&gt;

</description>
      <category>aivideo</category>
      <category>computervision</category>
      <category>productengineering</category>
      <category>promptengineering</category>
    </item>
    <item>
      <title>Roll-Forward Versioning and Concurrent Golden-Data Forks in an Enterprise Review Pipeline</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:10:59 +0000</pubDate>
      <link>https://dev.to/humzakt/roll-forward-versioning-and-concurrent-golden-data-forks-in-an-enterprise-review-pipeline-274k</link>
      <guid>https://dev.to/humzakt/roll-forward-versioning-and-concurrent-golden-data-forks-in-an-enterprise-review-pipeline-274k</guid>
      <description>&lt;p&gt;On the enterprise workflow platform, a task moving through review can be claimed, released, reworked, and re-claimed by a different reviewer — sometimes several times before it's done. The original design cleared a reviewer's draft state the moment a claim was released, on the theory that a fresh claimant should start clean. In an auditable training-data pipeline, that theory is wrong: clearing state destroys exactly the record you need when something needs to be traced back later. AGT-1431 replaced claim-time draft clearing with roll-forward versioning — release moves the task forward to a new version rather than erasing what was there. That one architectural change is small to describe and took the rest of the month to get right.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Versioning forward instead of clearing in place is the correct call for auditability — and it moves every assumption your navigation, lineage, and concurrency code was quietly making.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The core change: release moves forward, it doesn't erase
&lt;/h2&gt;

&lt;p&gt;PR #831 (AGT-1431) is the foundation: manual release and timeout release both now roll the task forward to a new version instead of clearing the current draft. The PR shipped with a shared version-copy policy, release async guards, assignment-release metadata events, and lineage exposed through the version API and UI — plus, notably, a documented roll-forward contract and validation checklist committed alongside the code. That documentation mattered more than it might look: almost every bug that followed was a case of some other part of the system not yet knowing the new contract existed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Navigation bugs: correct data, wrong screen
&lt;/h2&gt;

&lt;p&gt;PR #869 (AGT-1593, AGT-1594) found two bugs introduced by the roll-forward change, and the first is a good reminder that "the data is right" and "the feature works" are different claims. A same-user re-claim was redirecting to the &lt;em&gt;previous&lt;/em&gt; version instead of the current head — a different user claiming the same task redirected correctly, so the underlying versioning was sound; only the same-user navigation path had the bug. The second, AGT-1594, was a lineage-tracking bug: a source version was supposed to stay pinned at its base (v4) but was incrementing on every subsequent action, so the version dropdown showed "v6 (Source: v5)" when it should have shown "v6 (Source: v4)" — the lineage pointer was chasing the current version instead of recording where the chain actually started.&lt;/p&gt;

&lt;h2&gt;
  
  
  A redirect that fired for every version, not just the rework source
&lt;/h2&gt;

&lt;p&gt;PR #988 (AGT-1760) is the sharpest of the navigation bugs. Golden-data rework forks a task to a new version, and an effect was supposed to redirect a viewer straight to that new version — but the condition triggering the redirect was simply "the manifest's version is greater than whatever version you're looking at." That's true for &lt;em&gt;every&lt;/em&gt; older version, not just the one the fork actually copied from. The practical effect: an admin trying to inspect version history on a task with a pending rework got redirected away from v1, v2, and v3 alike, whenever the manifest sat at v4 — making it impossible to look at anything but the newest version while a rework was in flight. The fix extracts the redirect condition into its own named function, scoped specifically to the rework's actual source version, so inspecting history stops fighting the same redirect that's supposed to help active reviewers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cross-task lineage and a 500 that only fired on clones
&lt;/h2&gt;

&lt;p&gt;PR #893 (AGT-1604) traced a 500 error on the version-history endpoint back to lineage resolution rejecting a legitimate shape: a test-task clone whose version chain contained a cross-task terminal ancestor &lt;em&gt;in the middle&lt;/em&gt; of the chain, not just as the starting point. The resolver had only ever been exercised with a terminal ancestor as an origin; a chain that passed through one and continued was outside what it had been built to expect, and it failed closed with a 500 instead of resolving. The fix widens acceptance to cover cross-task terminals wherever they legitimately appear in a chain, verified against two real QA tasks pulled directly from Cloud Run logs — the kind of bug that specifically needs production traces to notice, since it only manifests on a lineage shape that generated tests are unlikely to construct by accident.&lt;/p&gt;

&lt;h2&gt;
  
  
  A GCS marker race that flashed a false failure
&lt;/h2&gt;

&lt;p&gt;PR #894, following directly from #893, chased down a brief "task version resolution failed" flash that appeared every time a user opened a task from the dashboard. Two root causes stacked: server-side, an unpinned GCS read had no 404 handling around the case where a version's visibility marker exists as a prefix but is still mid-write — a genuine TOCTOU race between roll-forward, a calibration fork, or initial version creation writing the marker, and a read arriving in that same narrow window. Client-side, the auto-seed list page was separately navigating to a stale version. The fix normalizes the 404 race so it's treated as "still loading," not "failed," and keeps the loading state visible through legitimate retries instead of surfacing a transient race as a hard error to the user.&lt;/p&gt;

&lt;h2&gt;
  
  
  The race that actually wedged a task
&lt;/h2&gt;

&lt;p&gt;PR #987 (AGT-1759) is the one real concurrency bug in this cluster, found live on QA: a specific task was stuck with v4's golden-data step showing a QC error while v5 sat in the trainer review gate — a genuinely confusing state for whoever opened it next. The root cause was concurrent calls to the version-resolution function racing to fork the same rework version through a copy-then-write-then-commit sequence against GCS. Two forks landing on the same source version at the same time collided on the compare-and-swap for the result file, and the copy failed outright with a parallel-copy error. The fix serializes concurrent golden-data QC-rework forks, closing the exact window where two reviewers (or a reviewer and a retry) could both try to fork the same version at once.&lt;/p&gt;

&lt;p&gt;PR #981 (AGT-1354) is the PR that originally built the fork-on-rework mechanism this race lived inside — rebased onto main after the shared QC engine adapter (from a separate PR stack, #841–#843) had landed independently and implemented an overlapping piece of the same file. Rather than pick one implementation and discard the other, the rebase combined both: the shared adapter's load/writeStatus/finalizeOnPass machinery stayed, and roughly 640 lines of the original PR's custom copy/rollback/resolve infrastructure were replaced by calls into that shared adapter instead of a parallel implementation living next to it. Reconciling two independently-built solutions to the same problem into one, rather than shipping whichever landed second, is the less obvious and more valuable version of "resolving a merge conflict."&lt;/p&gt;

&lt;h2&gt;
  
  
  A conflict filter with no memory of "released"
&lt;/h2&gt;

&lt;p&gt;PR #990 (AGT-1761) is a smaller bug with an outsized effect on reviewer experience: the task pool hides tasks from a user holding a conflicting prior assignment, but all four of the exclusion predicates matched on status alone (&lt;code&gt;CLAIMED&lt;/code&gt; or &lt;code&gt;CLOSED&lt;/code&gt;) with no check on &lt;em&gt;why&lt;/em&gt; the assignment closed. A user who claimed a task and then released it back to the pool was hidden from it identically to the user who actually submitted it — meaning releasing a task could permanently lock the releaser out of ever claiming it again, the opposite of what "release" is supposed to mean. The fix uses the platform's existing release-reason constants to exclude only genuinely-closed (submitted) assignments from the conflict filter, letting a released claim go back into the open pool the way it should.&lt;/p&gt;

&lt;h2&gt;
  
  
  Roll-forward's ripple into a separate storage tier
&lt;/h2&gt;

&lt;p&gt;PR #896 (AGT-1606) found that roll-forward versioning, as originally shipped, only copied a version tree into the primary storage bucket — nothing touched the long-horizon storage bucket used for larger, hour-scale tasks. Both the golden and model regrade dispatch paths assumed a specific package existed in that second bucket regardless, and failed with an opaque rsync error when it didn't. The fix self-heals the missing package and makes sure a required pipeline step runs before a regrade dispatch fires, rather than assuming a prerequisite that roll-forward's original scope had silently left out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two honest error-message fixes
&lt;/h2&gt;

&lt;p&gt;Two smaller PRs are worth including because they're a different kind of fix — not wrong behavior, just an error message that hid the real problem. PR #895 (AGT-1604) replaced a generic "Failed to resolve task versions" fallback with a descriptive error carrying the task ID, the failure stage, and the underlying cause, because the generic version made it impossible to diagnose anything on beta or QA without cross-referencing server logs by hand. PR #1006 fixed the same class of problem on the admin side: an uploaded QC feedback file in the wrong shape produced a raw Zod validation message — "Invalid input: expected string, received undefined" — with no field name and no hint at the expected shape, leaving an admin to guess which of several possible fields was missing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shipping a whole second review path alongside the first
&lt;/h2&gt;

&lt;p&gt;Separately from the versioning work, PRs #941 and #943 shipped a complete manual-QC review path as a new first-class alternative to the existing automated one. #941 laid reusable infrastructure first — auth guards and services, a Cloud Tasks integration with idempotency and retry, structured operation logging, and configuration for runtime roles — deliberately with no manual-QC-specific code in the same PR. #943 then built the actual Next.js UI on top of that infrastructure: a three-pane review console, a typed API client mirroring the backend's DTOs, and React Query hooks for review, submit, and transfer actions. Landing the reusable plumbing and the feature-specific UI as two separate, sequential PRs — rather than one large PR mixing both — is the same discipline as the roll-forward contract doc: each piece is reviewable and revertible on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: an architectural fix moves the bugs, it doesn't remove them
&lt;/h2&gt;

&lt;p&gt;Every bug in this cluster traces back to one change: release now versions forward instead of clearing in place. That change was correct — clearing draft state on release genuinely was the wrong design for an auditable pipeline. But "correct architecture" and "correct in every caller" are different claims, and this month is the gap between them: navigation code that assumed release meant "start over" redirected to the wrong place, lineage tracking that assumed a source version was static drifted, a second storage tier the original change never touched broke silently, and concurrent access to the new fork-on-rework path raced in a way the old clear-in-place design never could have, because it never had two versions to race between. None of that is a reason not to make the architectural fix. It's the reason an architectural fix needs a month of follow-through, not a day.&lt;/p&gt;

</description>
      <category>enterpriseplatform</category>
      <category>dataversioning</category>
      <category>typescript</category>
      <category>distributedsystems</category>
    </item>
    <item>
      <title>Shipping a New Tool With the Sibling Repos' Engineering Standards From Day One</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:10:26 +0000</pubDate>
      <link>https://dev.to/humzakt/shipping-a-new-tool-with-the-sibling-repos-engineering-standards-from-day-one-1h4p</link>
      <guid>https://dev.to/humzakt/shipping-a-new-tool-with-the-sibling-repos-engineering-standards-from-day-one-1h4p</guid>
      <description>&lt;p&gt;Most of the tools in this ad-production suite learned their engineering standards the hard way — a 400-line file cap earned by splitting god files after the fact, deploy verification earned by chasing a stale build for a week, dead-code audits earned by tripping over five PRs of code nothing called. The endcard conversion tool — a small service that turns AppLovin MRAID HTML endcards into MP4 and back — is the first one to start with those lessons already applied, in its very first PR, rather than paying to relearn them.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The cheapest way to fix a class of bug is to never let a new codebase repeat it. That only works if day one remembers what day two hundred learned.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Measuring the endcard instead of guessing at its length
&lt;/h2&gt;

&lt;p&gt;PR #1 is the tool's actual first feature, and it's a good one to lead with because it replaces a guess with a measurement. Converting an endcard to video used to mean typing a duration and hoping it lined up with the page's own animation loop. The new panel offers four real options instead: &lt;strong&gt;One play&lt;/strong&gt; — one cycle of the biggest repeating element, usually the background video, with anything smaller that doesn't divide evenly into it cut where it lands and the UI naming exactly what got cut and where; &lt;strong&gt;Seamless loop&lt;/strong&gt; — the shortest duration at which every animated element is simultaneously back at its starting position, so the loop point is invisible; &lt;strong&gt;Until it settles&lt;/strong&gt; — render until nothing on the page is still moving; or a fixed number of seconds, for when an editor genuinely wants a specific length regardless of what the page is doing.&lt;/p&gt;

&lt;p&gt;On the two real endcards used to validate this, the measured lengths were 5.000s and 7.000s — not the flat 10s a typed guess would have produced for either. That's the difference between a video that ends where the page's own motion actually resolves and one that just stops mid-loop because a human picked a round number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Standards adopted, not rediscovered
&lt;/h2&gt;

&lt;p&gt;The same PR's title says the quiet part out loud: it adopts the sibling repos' engineering standards as part of the initial build, not as a retrofit. That means the file-size discipline earned the hard way on the main video-generation service — the one that eventually needed an explicit 400-line cap enforced at lint time after files had grown past 15,000 lines — shows up here from commit one, when the codebase is still small enough that following it costs nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  A changelog script that flagged its own example as real work
&lt;/h2&gt;

&lt;p&gt;PR #2 is a small bug with a genuinely funny cause: the automated changelog-stamping tool, ported from a sibling repo, refused to run on this repo's own changelog — one hour after being ported in. The reason was that &lt;code&gt;CHANGELOG.md&lt;/code&gt;'s own documentation includes a fenced code example showing what an &lt;code&gt;## Unreleased&lt;/code&gt; heading looks like, and the stamper's parser couldn't tell a worked example living inside a code fence from a real pending section waiting to be stamped. It saw two candidates where there was exactly one. The fix shipped with a new test file that pins down the fence case specifically — watched failing against the old logic first, so the test actually proves something — plus cases for the grouping rules that must not read as accumulation, genuine accumulation, and the empty case. Worth noting: every refusal in this script exits 0 on purpose, so a broken changelog stamper can never fail a merge — which is exactly why a bug in it can sit invisible until someone happens to look.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing the gaps flagged at merge time
&lt;/h2&gt;

&lt;p&gt;PR #3 closes three things explicitly flagged as follow-ups when PR #1 merged, rather than left to rot as someday-maybe items. A &lt;code&gt;GET /api/version&lt;/code&gt; endpoint reports which commit is actually answering — and, more importantly, &lt;em&gt;where that answer came from&lt;/em&gt;, because on this host the deploy platform's variable list carries no git metadata by default, so the Dockerfile also bakes the commit sha in as a build argument and the endpoint reports which source won. &lt;code&gt;unknown&lt;/code&gt; is a distinct, honest answer for "a deployment that can't identify itself," which is a different failure than simply reporting the wrong build with false confidence. A &lt;code&gt;GET /api/queue&lt;/code&gt; endpoint answers whether a redeploy right now would lose in-flight work — the job table lives in the container's own memory, with nothing durable behind it, so a merge landing mid-conversion silently drops it. And the endcard's call-to-action hook, previously empty, got filled in.&lt;/p&gt;

&lt;h2&gt;
  
  
  A documentation bug caught by checking the platform, not assuming it
&lt;/h2&gt;

&lt;p&gt;PRs #4 and #5 are a small, honest pair worth reading together. PR #4 discovered — by actually querying the deploy platform's API rather than trusting what the README claimed — that this service had received exactly one deployment ever, covering every merge made since. Merging to &lt;code&gt;main&lt;/code&gt; was doing nothing; only a manual deploy command actually shipped anything. The root cause was a permissions gap: the deploy platform's GitHub App had never been granted access to this specific repository under the org, so the auto-deploy connection that every sibling tool takes for granted silently didn't exist here. The doc got corrected to say so, with the exact command and the exact place to grant access.&lt;/p&gt;

&lt;p&gt;PR #5, same day, corrects the correction: once the GitHub App was actually granted access, deploys started working again, verified against the platform's own reported commit sha rather than assumed from the README changing. The stopgap &lt;code&gt;ENDCARD_BUILD_SHA&lt;/code&gt; variable that had been covering for the missing auto-deploy got deleted once the real mechanism was confirmed live. Two documentation PRs in one day, each one checking a claim against the platform's actual state before writing it down, is a small thing — but it's the same discipline that, applied at scale on the main pipeline, is what turned "the docs said X" from a source of truth into something that has to earn trust every time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meanwhile, on the render-batch builder tool
&lt;/h2&gt;

&lt;p&gt;A short update on the sibling tool this one shares engineering lineage with, already covered in an earlier post on fixing its instance-repeat semantics. A finished ad naming two clips instead of one exposed a classification hazard: an automated "migrate misclassified files" job moved anything matching a simple SOP-naming pattern out of the shared source library and into Full Ads — and a short lead clip named &lt;code&gt;SL872&lt;/code&gt; matched that pattern purely by containing the substring &lt;code&gt;L8&lt;/code&gt;. Measured against the live library, 15 of 196 visible clips were real short leads that had been silently reclassified and moved out of the library the builders actually read from — a namespace collision doing real data-model damage, caught by checking the actual library contents rather than trusting the classifier's logic in isolation.&lt;/p&gt;

&lt;p&gt;Two more fixes rounded out the update: editors were seeing a bare &lt;code&gt;unauthorized&lt;/code&gt; error — a string this repo doesn't even produce internally, traced to the shared hub proxy rejecting the session on its fast path — which was also killing in-progress exports outright rather than degrading gracefully, alongside a separately-discovered ffmpeg flag that had been removed and needed restoring. And the product dropdown moved from a hardcoded list to reading &lt;code&gt;export_products&lt;/code&gt; live from the shared Supabase project the whole tool family now shares — the render-batch tool is a read-only consumer of a catalog owned and edited once, centrally, so a new product ships to every tool that reads it without a single per-tool PR.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: cheap lessons the first time, free the second
&lt;/h2&gt;

&lt;p&gt;None of the standards this new tool adopted were invented for it. The file-size cap, the deploy-verification discipline, the habit of checking a platform's actual state before documenting a claim about it — all of them were paid for, expensively, by other tools in this family over the preceding months. The only thing that made them free here was writing them into the very first PR instead of waiting for the codebase to grow large enough to need them the hard way. A lesson that isn't carried forward gets paid for again, in full, by the next tool.&lt;/p&gt;

</description>
      <category>videoproduction</category>
      <category>softwareengineering</category>
      <category>typescript</category>
      <category>tooling</category>
    </item>
    <item>
      <title>Four Stops Instead of Thirty: Rebuilding the Dashboard and Retiring Hand-Rolled UI</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:10:22 +0000</pubDate>
      <link>https://dev.to/humzakt/four-stops-instead-of-thirty-rebuilding-the-dashboard-and-retiring-hand-rolled-ui-1a35</link>
      <guid>https://dev.to/humzakt/four-stops-instead-of-thirty-rebuilding-the-dashboard-and-retiring-hand-rolled-ui-1a35</guid>
      <description>&lt;p&gt;An editor's real complaint about the main video-generation service's internal tool wasn't any single bug. It was the shape of using it: thirty confirmation stops on an attended run, a dashboard that buried the one thing you actually owed a click against three finished runs above it, and six different hand-rolled versions of a toggle scattered across screens that had each been built under deadline by whoever needed one that week. Over roughly thirty-five PRs across a month, the fix wasn't a redesign in the usual sense — it was systematically finding every place the tool disagreed with itself and picking one answer.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A tool that has quietly drifted its own visual language, its own vocabulary, and its own confirmation pattern isn't many small inconsistencies. It's one large tax, paid by whoever has to learn all of them at once.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Thirty stops, or zero — nothing in between
&lt;/h2&gt;

&lt;p&gt;PR #869 names the actual defect precisely: oversight was bimodal. An attended run parks at &lt;code&gt;2N + 2&lt;/code&gt; gates — script, board, then a frame gate and a clip gate for every scene after the first, then final — and the fleet's real runs run 14–15 scenes, so a fully attended run confirms roughly thirty times. The only alternative was &lt;code&gt;runAll&lt;/code&gt;, which skips straight to autopilot. Thirty stops or zero was the entire range an editor could choose from. The fix collapses that to four real checkpoints while deliberately keeping the ones that actually gate money — the framing is explicit about which stops earned their place and which were just accumulated ceremony.&lt;/p&gt;

&lt;h2&gt;
  
  
  Front doors that didn't open
&lt;/h2&gt;

&lt;p&gt;PR #865 is a small, sharp audit: inventory every screen and every way a user is supposed to reach it, then actually click through. Three paths dead-ended. The failed-run recovery link — the one button offered after a run has already failed — built a URL against a route deleted three weeks earlier, so the recovery path from a failure was itself a 404. PR #874 found the same shape in the diagnostics page: its only inbound link gated on a benchmark-model field that the create form had stopped writing entirely, so the gate could never open, on any run the app could currently start — a page that existed, worked, and was structurally unreachable. PR #880 found the whole library section — Products, B-roll, Avatars, Profiles — rendered only while its own section was already active in the URL, meaning from anywhere else in the app, four real destinations were simply invisible in the navigation that was supposed to expose them.&lt;/p&gt;

&lt;h2&gt;
  
  
  A control that lied about what it did
&lt;/h2&gt;

&lt;p&gt;PR #871 catches something worse than a broken link: a control that renders as functional and isn't. The create form's flag system has a dependency graph — a flag whose parent requirement is unmet resolves to &lt;code&gt;false&lt;/code&gt; no matter how the operator sets it. The UI didn't know that. All 56 flags rendered as equal toggles, so six of them could be clicked, visibly read "On," and change nothing about the run — on flags that decide what a paid generation actually does. The fix mirrors the same transitive resolution client-side rather than re-deciding it independently, so the control can't say something the system won't honor. PR #879 closed the readability half of the same problem: 32 of 32 settings had a group label: zero of 56 flags did, so the ones that spend real money sat unmarked in one long ungrouped scroll — a gap the component's own code comment had already conceded before anyone fixed it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A dashboard that answers the wrong question first
&lt;/h2&gt;

&lt;p&gt;PR #887 found the landing page sorting runs newest-first and showing five — so a run genuinely parked at a gate, waiting on a person, could sit below three finished runs and fall off the bottom of the strip entirely, while a paid render waited on a click nobody knew was owed. The fix cost zero new derivation: &lt;code&gt;groupRunRows&lt;/code&gt; already computed which runs needed a person, oldest-wait-first, and already existed elsewhere in the app — the landing page had just never called it. PR #964 pushed the same idea one step further: rather than one sorted list with a badge as the only signal, the dashboard now renders two headed blocks, "Waiting for you (N)" and "Everything else," because an editor scanning the page is answering exactly one question, and a heading answers it faster than a sort order can.&lt;/p&gt;

&lt;h2&gt;
  
  
  Clips you already paid for, hidden behind a spinner
&lt;/h2&gt;

&lt;p&gt;PR #883 is the kind of bug that costs real wall-clock time rather than just looking bad: the clip review gate keyed its &lt;em&gt;entire view&lt;/em&gt; on whether the currently-focused scene had a finished take. While scene 3 rendered, scenes 0 through 2 — rendered, billed, fully playable — simply disappeared behind a loading panel. On a fourteen-scene run at four to six minutes a clip, that's most of an hour spent staring at one spinner with nothing to watch, for content that already existed and had already been paid for.&lt;/p&gt;

&lt;h2&gt;
  
  
  The run page learns to say where the money went
&lt;/h2&gt;

&lt;p&gt;PR #881 traced a real cost surprise — one operator-animated run's total read $9.05 against a roughly $1.40 estimate, with nothing on screen explaining the gap. The answer was already being recorded in &lt;code&gt;costNotes&lt;/code&gt; and simply never surfaced: 49 image calls for 14 frames, broken down by vendor and purpose (42 Anthropic clip-critic calls, 36 OpenAI-image frame generations, 27 Anthropic continuity checks, and so on). The fix isn't a new measurement — it's making an existing, accurate ledger finally visible on the one screen someone would actually look at after a surprising bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eight ways of saying the same two things
&lt;/h2&gt;

&lt;p&gt;The most concentrated part of this cluster is a ten-PR "editor-UI rebuild" run, each one stacked on the last. It starts with groundwork: PR #961 found two of the guardrails meant to detect hand-rolled UI duplication were themselves reading through re-export barrels rather than the actual leaf files, silently blind to exactly the drift they existed to catch. Once the detector could actually see, the list it produced was the whole story: &lt;strong&gt;eight different phrasings&lt;/strong&gt; — "accordion header row," "inline expand toggles," "inline disclosure link," "dashed full-width add affordance," and five more — describing what were structurally only two shapes: a disclosure and an insert row. PR #968 replaced all eight with two shared controls. PR #962 did the same for booleans and one-of-N choices, retiring four hand-rolled toggle/segmented-control copies that shared no code with each other. PR #969 unified six separately-drawn media-thumbnail components behind one primitive specifically built to not re-fetch on every dashboard poll — a real performance bug hiding inside what looked like a pure styling duplication. PR #967 found the app's only focus-trapping modal had no test at all and sat on the exception list for its own close button; it became the shared overlay primitive, joined immediately by a drawer the library screens needed, so two hand-written Tab traps couldn't quietly disagree with each other about keyboard behavior.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// before: eight components, eight names, zero shared code&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;InlineDisclosureLink&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="cm"&gt;/* ... */&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;DashedInsertAffordance&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="cm"&gt;/* ... */&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;StagedEditToggle&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="cm"&gt;/* ... */&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c1"&gt;// ...five more, each solving the same two problems independently&lt;/span&gt;

&lt;span class="c1"&gt;// after: two primitives, everything else composes them&lt;/span&gt;
&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Disclosure&lt;/span&gt; &lt;span class="nx"&gt;summary&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;...&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;children&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="sr"&gt;/Disclosure&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;
&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;InsertRow&lt;/span&gt; &lt;span class="nx"&gt;onInsert&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{...}&lt;/span&gt; &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  One verb, one price formatter, one name for money
&lt;/h2&gt;

&lt;p&gt;PR #972 found five different names spread across nine UI labels for the single action "make this again" — "re-roll," "Regenerate frame," "Regenerate as new version," "Re-do this scene + the ones after it," among others — plus two different progress words for the identical underlying work on adjacent gates. An editor who learned one gate's vocabulary learned nothing transferable about the next one. The fix collapses all of it to one verb, "Regenerate," with scope carried as a separate, explicit parameter rather than folded into the verb's name. PR #974 found a subtler duplication in the same spirit: a currency formatter defined twice, with two different signatures, because two real contracts existed — a table cell that must never collapse to nothing gets an em-dash, while a chart label or sentence sometimes correctly shouldn't render at all. The fix wasn't to force one signature; it named both contracts explicitly and gave each one exactly one implementation instead of four call sites independently deciding how to paper over the difference.&lt;/p&gt;

&lt;h2&gt;
  
  
  One wizard, not secretly two
&lt;/h2&gt;

&lt;p&gt;PR #988 found the create flow and the wizard flow were, structurally, two different pipeline models rendering on top of each other: the create form drew a six-stage progress bar, then submitted with a full page navigation into a wizard drawing an entirely separate five-phase stepper — so at the exact same script gate, one chart said "3 of 6" and the other said "2 of 5" for the identical run. PR #950 gave each pipeline phase its own URL specifically so a gate could be linked directly and the browser's Back button would do something meaningful between phases — reported by an editor, verbatim, as difficulty navigating the pipeline at all. PR #991 closed out four more pieces of the same feedback thread: a completed run no longer auto-navigates into History the instant it finishes, the two competing pipeline visualizations became one, and the interface dropped its warnings and blockers in favor of a state an editor described wanting as simply "clean." PR #992 finished the thread by retiring the app's last two native &lt;code&gt;window.confirm()&lt;/code&gt; dialogs — the ones for restart and cancel — into the same custom confirmation pattern already used for the two gates that release money, so a spend confirmation and a destructive-action confirmation finally look and read the same way everywhere they appear.&lt;/p&gt;

&lt;h2&gt;
  
  
  Speaking the editor's language, not the pipeline's
&lt;/h2&gt;

&lt;p&gt;PR #963 is a no-behavior-change PR that's really about respect for the audience: the library screens displayed pipeline internals directly to the people using them — a column literally labeled &lt;code&gt;ASL&lt;/code&gt; (film-school shorthand for average shot length) where "Cuts every 1.83s" would do, and an empty state that handed an editor a shell script to run. PR #976 found the underlying library itself had silently shrunk to a fraction of its real size — 121 clips existed, but the app could browse only 29, because nothing in the UI could write the tag index that made the rest searchable; a project doc had already named re-tagging as "free and the prerequisite" for retrieval working at all, and it hadn't been free, because the write path simply didn't exist yet. PR #977 closed a related gap: reference videos in the library had already been measured once, at real cost, and every new run silently re-measured the same reference from scratch anyway — sometimes inconsistently, with the same file returning 13 shots on one run and 15 on the next. The fix keys a stored measurement by the file's own content hash, so picking a known reference reuses what was already paid for instead of re-deriving it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: consistency is a feature, not a cleanup task
&lt;/h2&gt;

&lt;p&gt;Nothing in this cluster added a capability the tool didn't already have in some form. Every PR either exposed a destination that already existed, reused a value that was already being computed, or replaced several independently-built versions of the same idea with one. The through-line worth naming is that a tool which drifts its own vocabulary, its own confirmation pattern, and its own component shapes doesn't fail all at once — it fails one small relearning cost at a time, paid by every person who has to move between its screens. Collapsing thirty gates to four, or eight component phrasings to two, isn't polish layered on top of working software. It's removing the tax the software had been quietly charging its own users to operate it.&lt;/p&gt;

</description>
      <category>frontendengineering</category>
      <category>ux</category>
      <category>react</category>
      <category>designsystems</category>
    </item>
    <item>
      <title>The Connective Tissue of an AI Platform: Workflow, Taxonomy, Auth, and Memory</title>
      <dc:creator>Humza Tareen</dc:creator>
      <pubDate>Tue, 25 Aug 2026 21:27:58 +0000</pubDate>
      <link>https://dev.to/humzakt/the-connective-tissue-of-an-ai-platform-workflow-taxonomy-auth-and-memory-15b9</link>
      <guid>https://dev.to/humzakt/the-connective-tissue-of-an-ai-platform-workflow-taxonomy-auth-and-memory-15b9</guid>
      <description>&lt;p&gt;When you're building an AI evaluation platform with multiple microservices, the "core" services get all the attention — the evaluation engine, the scoring system, the RAG pipeline. But a platform doesn't work without the connective tissue: the workflow orchestration that keeps humans in the loop, the taxonomy engine that classifies tasks intelligently, the platform service that ties authentication together, and the evaluation suites that ensure models actually remember context.&lt;/p&gt;

&lt;p&gt;These four services don't make headlines, but they're what turned a collection of microservices into an actual platform. Here's what went into each one and why the engineering decisions mattered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Workflow Orchestration: The Human-in-the-Loop Engine
&lt;/h2&gt;

&lt;p&gt;AI evaluation is not fully automated — and it shouldn't be. Certain decisions require human judgment: Is this model response harmful? Does this evaluation rubric make sense for this domain? Is this edge case a genuine failure or acceptable behavior?&lt;/p&gt;

&lt;p&gt;The workflow orchestrator manages these decision points. It coordinates multi-step evaluation workflows where some steps are automated (LLM scoring, data validation) and others require human approval before the pipeline continues.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Architecture
&lt;/h3&gt;

&lt;p&gt;The core is a &lt;strong&gt;state machine&lt;/strong&gt; built on FastAPI and PostgreSQL. Each workflow is a DAG (directed acyclic graph) of tasks, where each node can be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Automated:&lt;/strong&gt; Runs immediately, calls another service (scoring, data enrichment), stores the result&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human gate:&lt;/strong&gt; Pauses the workflow, notifies the assigned reviewer via the notification service, waits for approval/rejection&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conditional:&lt;/strong&gt; Routes to different branches based on previous step outcomes (e.g., if confidence score &amp;lt; threshold, escalate to senior reviewer)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;State transitions are persisted in PostgreSQL with Alembic-managed migrations. Every transition is logged — who approved what, when, and with what context. This audit trail turned out to be critical for client reporting.&lt;/p&gt;

&lt;h3&gt;
  
  
  Real-Time Updates with WebSocket
&lt;/h3&gt;

&lt;p&gt;The original system polled the API every 5 seconds to check workflow status. With dozens of reviewers working concurrently, this created unnecessary load and a poor user experience — you'd approve a task and see nothing happen for up to 5 seconds.&lt;/p&gt;

&lt;p&gt;I replaced this with &lt;strong&gt;WebSocket connections&lt;/strong&gt; that push state changes in real-time. When a reviewer approves a step, every connected client watching that workflow sees the update instantly. The implementation uses FastAPI's WebSocket support with Redis Pub/Sub as the message broker, so it works across multiple Cloud Run instances.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Simplified WebSocket broadcast pattern
&lt;/span&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;broadcast_workflow_update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;workflow_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;channel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;workflow:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;workflow_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;publish&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;state_change&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;workflow_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;workflow_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;step&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;step&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;new_status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;actor&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;actor_email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timestamp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;utcnow&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;isoformat&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;}))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Production Logging Overhaul
&lt;/h3&gt;

&lt;p&gt;The existing codebase used &lt;code&gt;print()&lt;/code&gt; statements everywhere. In production on Cloud Run, these were effectively invisible — they'd show up as unstructured text in Cloud Logging with no way to filter, search, or correlate them.&lt;/p&gt;

&lt;p&gt;I replaced the entire logging infrastructure with &lt;strong&gt;structured JSON logging&lt;/strong&gt;. Every log entry includes a correlation ID that traces a request across the workflow orchestrator, the notification service, and whatever downstream service is involved. When a workflow fails at step 4 of 7, you can now trace exactly what happened at each step, in each service, with a single query.&lt;/p&gt;

&lt;h2&gt;
  
  
  Taxonomy Workflow Engine: Intelligent Task Classification
&lt;/h2&gt;

&lt;p&gt;Not all evaluation tasks are the same. A code generation task requires different rubrics, different evaluators, and different tooling than a conversational AI task. The taxonomy engine is the routing layer that classifies incoming tasks and determines which evaluation workflow to apply.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Problem It Solves
&lt;/h3&gt;

&lt;p&gt;Before this service existed, task classification was manual. A project manager would look at incoming evaluation requests, decide which team should handle them, and assign the appropriate rubric. This worked at 50 tasks per day. It didn't work at thousands.&lt;/p&gt;

&lt;h3&gt;
  
  
  How It Works
&lt;/h3&gt;

&lt;p&gt;The engine uses a combination of &lt;strong&gt;keyword matching&lt;/strong&gt;, &lt;strong&gt;metadata analysis&lt;/strong&gt;, and &lt;strong&gt;configurable rule sets&lt;/strong&gt; to classify tasks. Each classification determines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which evaluation rubric to apply&lt;/li&gt;
&lt;li&gt;Which reviewer pool to draw from (by expertise)&lt;/li&gt;
&lt;li&gt;Whether the task requires single or multi-reviewer consensus&lt;/li&gt;
&lt;li&gt;SLA targets for completion time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The file upload system allows clients to submit evaluation tasks in bulk via CSV/JSON uploads to GCS. The engine parses, validates, classifies each row, and enqueues them into the appropriate workflow — all asynchronously via Cloud Tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Infrastructure:&lt;/strong&gt; Cloud SQL for taxonomy rules and classification history, GCS for bulk file uploads, Cloud Run for the API layer, Cloud Tasks for async processing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Core Platform Service: The Authentication Backbone
&lt;/h2&gt;

&lt;p&gt;Every microservice in the platform needs to answer two questions: "Who is making this request?" and "Are they allowed to do this?" The core platform service provides those answers.&lt;/p&gt;

&lt;h3&gt;
  
  
  JWT Authentication Fixes
&lt;/h3&gt;

&lt;p&gt;The existing JWT implementation had a subtle but critical bug: token validation was checking expiration time against the &lt;em&gt;server's local time&lt;/em&gt; rather than UTC. Cloud Run instances can have slight clock drift, and this meant tokens would occasionally be rejected as "expired" when they were still valid, or accepted when they should have been rejected.&lt;/p&gt;

&lt;p&gt;The fix was straightforward — normalize all time comparisons to UTC — but finding it required tracing sporadic 401 errors across multiple services to realize the pattern correlated with specific Cloud Run instances, not specific users.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Before: clock-sensitive comparison
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;token_exp&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;  &lt;span class="c1"&gt;# Local time — unreliable on Cloud Run
&lt;/span&gt;    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;HTTPException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Token expired&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# After: UTC-normalized comparison
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;token_exp&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timezone&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;utc&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;  &lt;span class="c1"&gt;# Always consistent
&lt;/span&gt;    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;HTTPException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Token expired&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  User Management and GDPR Compliance
&lt;/h3&gt;

&lt;p&gt;Built the user deletion endpoint — which sounds simple until you realize that "deleting a user" in a system with audit trails, evaluation history, and cross-service references means carefully cascading the deletion while preserving anonymized audit records. The implementation soft-deletes the user profile, anonymizes their evaluation history (replacing PII with hashed identifiers), and propagates the deletion event to downstream services via Pub/Sub.&lt;/p&gt;

&lt;h3&gt;
  
  
  Developer Experience
&lt;/h3&gt;

&lt;p&gt;Improved the local development workflow by rewriting &lt;code&gt;start_dev.sh&lt;/code&gt; to properly handle Docker container lifecycle. The previous script would silently fail if the PostgreSQL container was already running from a previous session, leading to "connection refused" errors that wasted 10-15 minutes of debugging time per developer, multiple times per week. The new script checks for existing containers, handles cleanup, and provides clear status messages.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory Evaluation Suite: Does the Model Remember?
&lt;/h2&gt;

&lt;p&gt;One of the harder problems in LLM evaluation is measuring &lt;strong&gt;context retention&lt;/strong&gt;. When you give a model a long conversation or a complex document, does it actually use information from the beginning when answering questions at the end? Or does it "forget" earlier context?&lt;/p&gt;

&lt;p&gt;The memory evaluation suite provides structured tests for this. It generates conversations with deliberate information planted at various positions (beginning, middle, end), then asks questions that require recalling that information. The scoring tracks not just accuracy, but &lt;strong&gt;where&lt;/strong&gt; in the context window the model starts losing information — which is critical data for the teams training these models.&lt;/p&gt;

&lt;h3&gt;
  
  
  Code Quality as a Feature
&lt;/h3&gt;

&lt;p&gt;This was also where I implemented the team's first &lt;strong&gt;pre-commit workflow&lt;/strong&gt; using Ruff for linting and formatting, enforced via GitHub Actions. The motivation wasn't just code aesthetics — inconsistent formatting was causing unnecessary merge conflicts across the team. Two developers would change the same file, both would reformat it differently, and the merge conflict had nothing to do with the actual logic.&lt;/p&gt;

&lt;p&gt;After rolling out the pre-commit pipeline on this service and proving it reduced merge conflicts, we adopted it across every service in the platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Engineering Pattern
&lt;/h2&gt;

&lt;p&gt;What ties these four services together isn't the domain logic — it's the &lt;strong&gt;systematic engineering discipline&lt;/strong&gt; I applied to each one. Every service I touched got the same treatment:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;Why It Matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Structured JSON logging&lt;/strong&gt; with correlation IDs&lt;/td&gt;
&lt;td&gt;One query to trace a request across all services. Reduced mean time to diagnosis from hours to minutes.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Pre-commit hooks&lt;/strong&gt; (Ruff, type checking)&lt;/td&gt;
&lt;td&gt;Eliminated formatting merge conflicts. Caught type errors before they hit production.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Custom exception hierarchies&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Consistent error responses across services. Clients can programmatically handle errors instead of parsing strings.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Alembic migrations&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Version-controlled schema changes. Zero-downtime deployments with reversible migrations.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Security audit per service&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Found hardcoded credentials, missing auth checks, and SQL injection vectors before they became incidents.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Why This Work Matters
&lt;/h2&gt;

&lt;p&gt;It's easy to dismiss "I worked on four more services" as a breadth play. But the reality is that &lt;strong&gt;platform engineering requires breadth&lt;/strong&gt;. The workflow orchestrator doesn't exist without the platform service providing authentication. The taxonomy engine doesn't work without the workflow orchestrator to route tasks into. The memory evaluation suite's code quality pipeline became the template for every other service.&lt;/p&gt;

&lt;p&gt;These aren't four independent projects. They're four layers of a system that only works because someone cared enough to apply the same engineering rigor to the "boring" services that they applied to the "interesting" ones.&lt;/p&gt;

</description>
      <category>microservices</category>
      <category>hitl</category>
      <category>websocket</category>
      <category>python</category>
    </item>
  </channel>
</rss>
