<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Michael Truong</title>
    <description>The latest articles on DEV Community by Michael Truong (@michaeltruong).</description>
    <link>https://dev.to/michaeltruong</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3965775%2F868d43f8-59c8-45ca-93f1-3f2428fb222d.jpg</url>
      <title>DEV Community: Michael Truong</title>
      <link>https://dev.to/michaeltruong</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/michaeltruong"/>
    <language>en</language>
    <item>
      <title>I visualised my job search. The data model started breaking.</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:43:21 +0000</pubDate>
      <link>https://dev.to/michaeltruong/records-first-chart-second-the-sankey-disagreed-3841</link>
      <guid>https://dev.to/michaeltruong/records-first-chart-second-the-sankey-disagreed-3841</guid>
      <description>&lt;p&gt;I wanted to see how my job applications progressed: where they started, which stages they went through, and where they ended.&lt;/p&gt;

&lt;p&gt;I had assumed those histories would look like a conventional funnel: application, recruiter, hiring manager, technical, outcome. They did not. Some skipped stages. Some repeated them. Inbound opportunities did not start with an application at all.&lt;/p&gt;

&lt;p&gt;So I modelled each opportunity's path in structured YAML first. Then I wired those histories into a Sankey diagram, a &lt;a href="https://sankeymatic.com/" rel="noopener noreferrer"&gt;flow chart&lt;/a&gt; showing how applications moved between stages. Records first, chart second. The validators passed. The chart still kept surfacing problems the records alone hid.&lt;/p&gt;

&lt;h2&gt;
  
  
  The operational board was not funnel history
&lt;/h2&gt;

&lt;p&gt;I had already &lt;a href="https://dev.to/michaeltruong/build-looked-absurd-under-a-recruiter-deadline-1145"&gt;built a job search system&lt;/a&gt; that kept each opportunity in Notion alongside its research and generated resume in a repo. As the search ramped up, overlapping hiring processes replaced a handful of quiet opportunities. With enough of them accumulating, I wanted the cross-search view the board could not provide.&lt;/p&gt;

&lt;p&gt;I use the Notion board to keep track of where each opportunity is, whether I am waiting on a company, and what I need to do next. What it does not preserve is the path each opportunity took to get there.&lt;/p&gt;

&lt;p&gt;When a role ends in rejection, &lt;strong&gt;Phase&lt;/strong&gt; often moves to &lt;strong&gt;Completed&lt;/strong&gt;. That is useful on a kanban, but it loses the history I needed for analysis. The board no longer answers "how far did this application get before it ended?" I could see that a process closed, not the path it took.&lt;/p&gt;

&lt;p&gt;Funnel analytics needs observed transitions between stages, grounded in substantiated events. Company A might run recruiter, then hiring manager, then technical. Company B might skip recruiter entirely. No global ordering of stage names can represent both without lying.&lt;/p&gt;

&lt;p&gt;I needed to normalise those histories without forcing them back into a global funnel: how an opportunity entered, the ordered process events that followed, and any terminal outcome.&lt;/p&gt;

&lt;p&gt;That is why &lt;code&gt;lifecycle.yml&lt;/code&gt; exists as a separate canonical store: ordered &lt;code&gt;events&lt;/code&gt;, one file per role posting, validated in CI. Notion stays operational. The repo holds history.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the chart kept surfacing
&lt;/h2&gt;

&lt;p&gt;The first problem was provenance: mixing how an opportunity entered the funnel with what happened after. Inbound outreach and cold applications became difficult to read on the same chart. A LinkedIn InMail that never went through a portal looked like it had already "reached" an application stage before any recruiter conversation. Entry provenance is how the opportunity entered the funnel. Recruiter, hiring manager, and technical rounds are what happened after. Those are different layers.&lt;/p&gt;

&lt;p&gt;Order was the next fight. Repeating a stage is normal: two technical rounds are two technical rounds, not one node with a count badge. The chart keys nodes by sequence position so it preserves order without inventing a global stage ladder. Direct exits are valid too. One application went directly from cold application to rejected with no process events in between. The YAML allows &lt;code&gt;events: []&lt;/code&gt; while a search is still open. The chart had to render &lt;code&gt;Cold application → Rejected&lt;/code&gt; without padding imaginary recruiter steps.&lt;/p&gt;

&lt;p&gt;Not every wrong branch was a modelling bug. The backfill had created a lifecycle file for another posting at the same company: researched, never submitted. Once the Sankey existed, that record looked wrong. There was no application process to chart, only notes where a hiring path should be. I deleted the lifecycle file rather than inventing another state for something that had never entered the funnel. The next day I applied the same boundary to the operational board and removed its Notion row too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Separating data, projection, and presentation
&lt;/h2&gt;

&lt;p&gt;Failures came from different parts of the system, so I kept three concerns separate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Canonical YAML&lt;/strong&gt; stores how an opportunity entered and the ordered process events that followed. A terminal outcome such as rejected, stalled, withdrawn, or accepted ends that history. If the research only substantiates entry so far, the file can stop there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Projection code&lt;/strong&gt; turns each history into Sankey edges. Histories without a terminal outcome append an &lt;strong&gt;Active&lt;/strong&gt; branch on the chart only. That state never gets written back to &lt;code&gt;lifecycle.yml&lt;/code&gt;. It answers "where does this open path end on the chart right now?" not "what is the canonical outcome?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Presentation&lt;/strong&gt; is layout: labels, column alignment, tooltips. I avoided a layout mode that forces every terminal into a shared rightmost column. Rejections after one recruiter screen and rejections after three rounds looked equivalent when all sinks lined up. Outcome nodes now terminate at natural depth, so a quick rejection after application does not share a column with a rejection after several rounds.&lt;/p&gt;

&lt;p&gt;That separation made triage faster. When a branch looked wrong, I could ask whether the canonical data was wrong, the projection code misread it, or the layout was misleading. Sometimes the problem was that an opportunity should never have been included in the lifecycle history at all.&lt;/p&gt;

&lt;p&gt;Spreadsheets and YAML validators catch schema mistakes. They do not show you that two rejection depths collapsed into one visual column, or that a role you researched but never submitted should not be in the dataset at all. The Sankey stress tested my representation of these hiring histories in specific, fixable ways.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; If you are modelling a multi-step human process with irregular ordering, build structured records first, then visualise early to see whether your categories mean what you think. A valid schema does not necessarily mean you have modelled the process correctly.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>analytics</category>
    </item>
    <item>
      <title>The agent host didn't have the lifecycle boundaries I assumed</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Fri, 25 Sep 2026 02:00:15 +0000</pubDate>
      <link>https://dev.to/michaeltruong/the-agent-host-didnt-have-the-lifecycle-boundaries-i-assumed-5akh</link>
      <guid>https://dev.to/michaeltruong/the-agent-host-didnt-have-the-lifecycle-boundaries-i-assumed-5akh</guid>
      <description>&lt;p&gt;As I started handing agents longer pieces of work, I stopped being present at every lifecycle boundary.&lt;/p&gt;

&lt;p&gt;In shorter copilot-style sessions I had been part of the control loop without really thinking about it. Longer agent runs broke that assumption: I was no longer present to know when work finished, what mattered, or what context should survive into the next task.&lt;/p&gt;

&lt;p&gt;One place this surfaced was memory. I wanted an agent to review completed work and decide whether anything was worth carrying into future sessions. I call the system I built for that Savepoints.&lt;/p&gt;

&lt;p&gt;Savepoints already had the semantic path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent work
→ Observe
→ emit-savepoint (worth keeping?)
   → emit-learning / no_capture
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Observe what happened, decide whether anything was worth keeping, then emit a learning or explicitly capture nothing.&lt;/p&gt;

&lt;p&gt;The weak point was how that review started. It depended on the agent following instructions, but I had no reliable way to tell whether it had.&lt;/p&gt;

&lt;p&gt;Across four substantial agent sessions, observe hooks confirmed real activity, but the &lt;code&gt;emit-savepoint&lt;/code&gt; skill was consulted inconsistently and &lt;code&gt;emit-learning&lt;/code&gt; never ran. I could not distinguish "reviewed the work and found nothing" from "the review never happened."&lt;/p&gt;

&lt;p&gt;I added two boundaries around the semantic review:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent work
→ Observe
→ Review opportunity       ← new boundary
→ emit-savepoint (worth keeping?)
   → emit-learning / no_capture
→ mark reviewed            ← new boundary
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open a review opportunity when observed work needed review, then mark that evidence as reviewed once the opportunity had been handled.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Observation and semantic judgment stay separate.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I assumed the host could establish
&lt;/h2&gt;

&lt;p&gt;I already had a suspected shape for the larger system, and checking half a dozen agent-memory systems reinforced it: opportunity IDs, idempotent review closure, fail-open hooks, semantic judgment left to the agent.&lt;/p&gt;

&lt;p&gt;What remained uncertain was whether those invariants had reliable lifecycle boundaries to attach to in the host.&lt;/p&gt;

&lt;p&gt;For this implementation, the host was Cursor, which exposes lifecycle hooks such as &lt;code&gt;beforeSubmitPrompt&lt;/code&gt;, &lt;code&gt;afterFileEdit&lt;/code&gt;, and &lt;code&gt;stop&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The host seemed to expose enough information to establish three boundaries: whether a turn was review scaffolding or source work, whether lifecycle events belonged to the same piece of work, and whether that work had actually affected this repository.&lt;/p&gt;

&lt;p&gt;If those answers were reliable, the adapter was straightforward: register the work, watch its evidence, then open review when it completed.&lt;/p&gt;

&lt;p&gt;I expected the next pass to be production wiring.&lt;/p&gt;

&lt;p&gt;Live probes said otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Probe 1: Host provenance is not scaffolding identity
&lt;/h2&gt;

&lt;p&gt;The straightforward implementation used Cursor's &lt;code&gt;stop&lt;/code&gt; hook as the completion signal. When ordinary work hit &lt;code&gt;stop&lt;/code&gt;, Savepoints would open a review opportunity and send a follow-up asking the agent to review what had just happened.&lt;/p&gt;

&lt;p&gt;That immediately created a recursion problem. &lt;strong&gt;The review was itself another agent turn.&lt;/strong&gt; If that turn also ended in &lt;code&gt;stop&lt;/code&gt;, Savepoints could mistake its own review for more completed work and open another review.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ordinary work
→ stop
→ open review
→ review follow-up
→ stop? (could trigger another review)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The natural first question was whether the host could tell me that this new turn was the review follow-up I had created.&lt;/p&gt;

&lt;p&gt;I tried making that distinction when &lt;code&gt;beforeSubmitPrompt&lt;/code&gt; fired. If the host could tell me whether a prompt came from the user or from the follow-up I had generated, I could classify the turn before any work happened.&lt;/p&gt;

&lt;p&gt;Live probes showed that distinction was not reliable enough to build on.&lt;/p&gt;

&lt;p&gt;So I stopped asking the host to infer an identity I controlled. When Savepoints creates a review follow-up, it now marks it explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SAVEPOINTS_CAPTURE_REVIEW_V1 opportunity_id=&amp;lt;uuid&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A prompt with that marker is review scaffolding. It does not register as new source work, so the review cannot recursively open another review.&lt;/p&gt;

&lt;p&gt;Instead of inferring scaffolding identity from host metadata, I made it part of the protocol I owned.&lt;/p&gt;

&lt;h2&gt;
  
  
  Probe 2: &lt;code&gt;stop&lt;/code&gt; did not establish which work had finished
&lt;/h2&gt;

&lt;p&gt;Explicitly marking the review turn solved one problem: Savepoints no longer had to infer whether a prompt was its own scaffolding.&lt;/p&gt;

&lt;p&gt;But an earlier attempt to contain that recursion had exposed a different assumption about &lt;code&gt;stop&lt;/code&gt;. Before I added the marker, I had tried a simpler guard: after opening review, suppress the next &lt;code&gt;stop&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That assumed the next &lt;code&gt;stop&lt;/code&gt; belonged to the review I had just opened. It didn't.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;open review
→ expect review stop
→ no stop arrives
→ guard remains armed

~100 seconds later
→ unrelated work stops
→ stale guard eats it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I tried several variations on the same idea. None gave me a reliable way to distinguish "the review I just started has finished" from "some unrelated work has finished."&lt;/p&gt;

&lt;p&gt;The stale guard had assumed an ordering relationship the host did not guarantee. A &lt;code&gt;stop&lt;/code&gt; told me that something had ended. It did not prove that the thing ending was the review I had just opened.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;generation_id&lt;/code&gt; gave me a stronger relationship: events carrying the same ID could be connected to the same agent generation. I no longer had to assume that the next &lt;code&gt;stop&lt;/code&gt; belonged to the work I was tracking.&lt;/p&gt;

&lt;p&gt;But correlation only told me which events belonged together. It did not tell me whether that work had affected this repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  Probe 3: Hook scope is not repository evidence
&lt;/h2&gt;

&lt;p&gt;Multi-root workspaces exposed why that distinction mattered.&lt;/p&gt;

&lt;p&gt;A single agent generation can touch several repositories in one session. Session-wide hooks still run in each Savepoints-enabled repository, even when the edit happened somewhere else.&lt;/p&gt;

&lt;p&gt;That meant &lt;code&gt;generation_id&lt;/code&gt; could tell me that events belonged to the same agent generation, but not which repository that work had affected. A generation did not need a single repository owner. It could span several repositories.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one agent generation
│
├─ edits repo A/
│  └─ afterFileEdit under repo A's root → evidence for repo A
│
└─ edits repo B/
   └─ afterFileEdit under repo B's root → evidence for repo B

Each repo evaluates its own evidence independently.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seeing the hook fire was therefore not evidence that this repository had been affected. Each repository asked a narrower question: &lt;strong&gt;did this generation produce evidence here?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An &lt;code&gt;afterFileEdit&lt;/code&gt; event counted only when its &lt;code&gt;file_path&lt;/code&gt; resolved under that repository's &lt;code&gt;repoRoot&lt;/code&gt;. If the session touched another repository but none of those edits belong here, this repository does not open review.&lt;/p&gt;

&lt;p&gt;Shell and MCP activity does not count toward this gate because I cannot reliably tie it to a repository. Savepoints can still learn from what happens around a tool call; it just does not try to observe what happens inside a tool boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three boundaries that must stay separate
&lt;/h2&gt;

&lt;p&gt;By this point, each failed assumption had removed an inference from the adapter. What remained were three boundaries that needed separate evidence:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Boundary I needed&lt;/th&gt;
&lt;th&gt;Reliable signal&lt;/th&gt;
&lt;th&gt;What I could not infer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Is this Savepoints' own review turn?&lt;/td&gt;
&lt;td&gt;&lt;code&gt;SAVEPOINTS_CAPTURE_REVIEW_V1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Scaffolding identity from host provenance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Do these lifecycle events belong to the same agent generation?&lt;/td&gt;
&lt;td&gt;&lt;code&gt;generation_id&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Which work had finished from event order alone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Did this work affect this repository?&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;afterFileEdit&lt;/code&gt; under its &lt;code&gt;repoRoot&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Repository evidence from hook scope&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Sometimes the correct adapter is unsupported
&lt;/h2&gt;

&lt;p&gt;On Desktop, I could now follow one piece of work all the way through: it started, this repository produced evidence, the same work stopped, and Savepoints opened review.&lt;/p&gt;

&lt;p&gt;Then I ran the same design in Cloud.&lt;/p&gt;

&lt;p&gt;Every piece seemed to be there. I could see work start, observe repository-local file edits, and see a final &lt;code&gt;stop&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The individual lifecycle events existed. The boundary I needed between them did not.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;this work started
→ this repository produced evidence
→ this same work stopped
→ open review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Desktop, &lt;code&gt;generation_id&lt;/code&gt; connected that chain. In Cloud, the ID I saw when work started and while evidence was collected did not reliably match the one I saw at &lt;code&gt;stop&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I could have tried to guess which events belonged together using &lt;code&gt;conversation_id&lt;/code&gt;, timing, or rules for normalizing the IDs.&lt;/p&gt;

&lt;p&gt;I didn't. Probe 2 had already shown the problem with guessing at correlation: seeing events in the expected sequence was not proof that they belonged to the same work.&lt;/p&gt;

&lt;p&gt;If Savepoints could not reliably establish that the work it observed was the work that just stopped, it could not safely open review for it. So in Cloud, the adapter treats capture review that depends on lifecycle hooks alone as &lt;code&gt;unsupported&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That does not mean Savepoints cannot run in Cloud, or that an agent cannot save a learning. It means this adapter cannot establish the evidence needed to guarantee this particular review path.&lt;/p&gt;

&lt;p&gt;I was also reluctant to build my own correlation scheme on top of a lifecycle surface that was still evolving. &lt;a href="https://forum.cursor.com/t/generation-id-has-extra-suffix-for-afteragentthought/166275" rel="noopener noreferrer"&gt;A later public Cursor report&lt;/a&gt; documented different &lt;code&gt;generation_id&lt;/code&gt; shapes across hook types. I would rather wait for a stable primitive I can verify than own an approximation that the host may eventually make unnecessary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; An agent host can expose lifecycle events without exposing the lifecycle boundaries your system needs. I had to establish scaffolding identity, lifecycle correlation, and repository-local evidence separately instead of inferring them from host events.&lt;/p&gt;

&lt;p&gt;I could establish review identity myself, but correlation and repository evidence still needed independent proof. When the host could not provide that proof in Cloud, &lt;code&gt;unsupported&lt;/code&gt; was more accurate than another heuristic.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>workflow</category>
    </item>
    <item>
      <title>AI made translation cheap. The hard part was knowing what architecture to keep.</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Thu, 17 Sep 2026 06:28:58 +0000</pubDate>
      <link>https://dev.to/michaeltruong/ai-made-translation-cheap-the-hard-part-was-knowing-what-architecture-to-keep-4hp3</link>
      <guid>https://dev.to/michaeltruong/ai-made-translation-cheap-the-hard-part-was-knowing-what-architecture-to-keep-4hp3</guid>
      <description>&lt;p&gt;I ship &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=ai-made-translation-cheap-localization-still-needed-human-judgment&amp;amp;utm_content=intro" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;, a web game where an LLM plays Codenames with you. One of my longtime readers, &lt;a class="mentioned-user" href="https://dev.to/xulingfeng"&gt;@xulingfeng&lt;/a&gt;, asked whether the game could support multi-language selection, and that was enough to pull forward work that had been sitting on the backlog: new word sets and languages, including Simplified Chinese.&lt;/p&gt;

&lt;p&gt;I am not a native Chinese speaker. AI could generate candidate translations far faster than I could judge whether they were natural, ambiguous in the right way, or simply plausible-looking mistakes. So I needed machinery around the generation: deterministic checks for what I could verify mechanically, and explicit review points for what required human judgment.&lt;/p&gt;

&lt;p&gt;I deliberately did not turn this into full application localization. The interface could stay in English; what mattered was the language of the board and the AI playing against it. The board language is carried through to the model, so it can reason about the right words and answer in that language.&lt;/p&gt;

&lt;p&gt;Generating the translations was fast. The harder work was discovering which parts of the system were real product invariants, and which were temporary scaffolding created by the first localization.&lt;/p&gt;

&lt;h2&gt;
  
  
  A useful signal hardened into a rule
&lt;/h2&gt;

&lt;p&gt;I assumed a strict substring-collision validator would protect card-word quality. If no playable value could appear inside another, ambiguous substring interactions would shrink and the pack would feel cleaner.&lt;/p&gt;

&lt;p&gt;That assumption was half right. Collisions matter in Codenames. Treating collision purity as a hard invariant made the Chinese worse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The first build treated overlap as a hard error.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If one playable value appeared inside another (水 (water) inside 水星 (Mercury), 手 (hand) inside 手表 (watch)), the build failed until I shortened or rewrote the translation. On paper that sounds responsible. In practice it pushed translations toward single-character fragments and awkward abbreviations just to satisfy the checker. The validator was green. The card words were becoming less natural.&lt;/p&gt;

&lt;p&gt;I relaxed the rule so overlap became something to review. Exact duplicates and invalid values still stop the build. Substring overlap in Chinese is often compositional (水, 水星 (Mercury), 水槽 (sink)) and sometimes a deliberate trade. Once collisions were reviewable instead of fatal, compact natural cards became viable again: &lt;strong&gt;WATER&lt;/strong&gt; → 水, &lt;strong&gt;WIND&lt;/strong&gt; → 风, &lt;strong&gt;HAND&lt;/strong&gt; → 手, &lt;strong&gt;SNOW&lt;/strong&gt; → 雪. The failure was elevating a useful authoring signal to the same severity as duplicates and empty values.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;But passing validation still wasn't enough.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When I tried to preserve English double meanings too aggressively, I got unnatural Chinese. Codenames lives on ambiguity in English. Chinese often forces you to choose a meaning more explicitly. A validator that only asks "is this collision-free?" cannot tell you which meaning belongs on the card.&lt;/p&gt;

&lt;p&gt;Some English concepts do not survive as a single clean Chinese word without choosing a meaning:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ROCK&lt;/strong&gt; → 岩石 (stone), not 摇滚 (music)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SPRING&lt;/strong&gt; → 春 (season), not 弹簧 (coiled metal)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SEAL&lt;/strong&gt; → 海豹 (animal), not 封 (to close)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those choices are product calls, not lint results. A validator can flag overlap. It cannot tell you that 春 (spring) is a better card word than 春天 (springtime) for this grid, or that a natural standalone word beats an awkward shortening invented only to satisfy the checker.&lt;/p&gt;

&lt;p&gt;This is a different question from clue-time substring rules in my English fairness validator (&lt;a href="https://dev.to/michaeltruong/schema-first-prompt-second-valid-json-wasnt-enough-3nhm"&gt;schema first, prompt second&lt;/a&gt;). Clue validation asks whether a spymaster hint is fair on today's board. Content validation asks whether a codename is a good standalone word on a tile. Same word "substring," different layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first implementation looked like the architecture
&lt;/h2&gt;

&lt;p&gt;I initially modeled Simplified Chinese as a localization of the English Classic pack: English concepts with Chinese mappings, validation metadata, and a build step that produced the word list the game loads at runtime. Codenames AI does not translate card words when a game starts.&lt;/p&gt;

&lt;p&gt;That structure was useful while I was doing the localization work. It gave me somewhere to record ambiguous translations and collision judgments.&lt;/p&gt;

&lt;p&gt;I made a mistake, though: I started treating that authoring workflow as the product model.&lt;/p&gt;

&lt;p&gt;I briefly compared the localized Classic pack against an existing Chinese community list. That was useful as a spot check, but it also exposed another category error: a native Chinese Codenames list is not automatically a translation of English Classic. I was treating cross-language concept identity as a runtime requirement.&lt;/p&gt;

&lt;p&gt;Later I added an Extended word set with independently sourced English and Simplified Chinese lists. There was no meaningful one-to-one mapping between them. They were simply two playable pools under the same product variant.&lt;/p&gt;

&lt;p&gt;And the game did not care.&lt;/p&gt;

&lt;p&gt;The runtime requirement was much smaller:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;wordSet × language → playable word pool&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Classic English, Classic Chinese, Extended English, and Extended Chinese could all use the same path.&lt;/p&gt;

&lt;p&gt;Once that was obvious, the original localization machinery became scaffolding rather than architecture. I removed the mapping layer, generator scripts, and old comparison artifacts: roughly 4,100 lines deleted in a follow-up cleanup, while keeping the shipped words intact.&lt;/p&gt;

&lt;p&gt;The useful distinction was not "translation versus native word list." It was &lt;strong&gt;authoring process versus runtime requirement&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I had made the same mistake inside the validator: a useful signal during authoring had hardened into a rule the product did not actually require.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where AI helps and where it stops
&lt;/h2&gt;

&lt;p&gt;AI made producing the translations cheap. It did not tell me what deserved to become permanent structure.&lt;/p&gt;

&lt;p&gt;A substring collision was useful evidence, not a hard invariant. A translation mapping was useful authoring scaffolding, not a runtime requirement.&lt;/p&gt;

&lt;p&gt;Both mistakes came from the same place: &lt;strong&gt;confusing the machinery that helped build a feature with the model the product actually needs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That smaller runtime requirement was what survived. The rest could stay where it belonged: in the implementation history, not the architecture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; Useful machinery from building a feature is not automatically part of the product model. Before a validator rule or authoring workflow hardens into architecture, ask what the runtime still requires without it.&lt;/p&gt;




&lt;p&gt;Play &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=ai-made-translation-cheap-localization-still-needed-human-judgment&amp;amp;utm_content=footer" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt; in Classic or Extended mode, in English or Simplified Chinese.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>Throwaway experiments are easy to start. Retiring one safely is not</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Thu, 10 Sep 2026 06:04:45 +0000</pubDate>
      <link>https://dev.to/michaeltruong/throwaway-experiments-are-easy-to-start-retiring-one-safely-is-not-2afe</link>
      <guid>https://dev.to/michaeltruong/throwaway-experiments-are-easy-to-start-retiring-one-safely-is-not-2afe</guid>
      <description>&lt;p&gt;I was closing out a throwaway repo from an agent-workflow experiment. I had treated experiment repos as cheap to delete once the hypothesis felt answered. The prototype had to go because leaving both checkouts live gave later agents two competing sources of precedent. Deleting it meant deciding what had been validated, writing it down somewhere durable, and removing experiment surfaces only after that record was complete.&lt;/p&gt;

&lt;p&gt;Standing up the narrow prototype had been genuinely fast. &lt;a href="https://dev.to/michaeltruong/i-was-solving-agent-portability-at-the-wrong-boundary-1406"&gt;Portable agent policy, skills, and repo bootstrap&lt;/a&gt; had already made that part easy. Safe retirement was not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was validated before deletion
&lt;/h2&gt;

&lt;p&gt;The experiment had tested one narrow claim: a working agent could turn captured learnings into retrospective summary cards in a single pass, without a dedicated second inference service. The prototype wired learnings into Notion, ran that synthesis step, and left card formats, retrieval defaults, and cap rules as unproven demo choices. Most of the wiring was unhardened prototype.&lt;/p&gt;

&lt;p&gt;By retirement time, that conclusion was already in architecture notes. Everything else in the prototype wiring was discardable.&lt;/p&gt;

&lt;p&gt;The audit question was not "did I copy every file?" It was "did I record every experimentally supported conclusion before I delete the experiment?"&lt;/p&gt;

&lt;p&gt;The answer was yes. That cleared the way for the part where things actually got scary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three ways experiment retirement goes wrong
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The first failure mode is preserving provisional choices because they feel unique.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When I asked an agent to plan the retirement, its first draft tried to absorb all the demo logic into architecture notes before deletion: card formats, retrieval defaults, cap rules, idempotency guesses. That treats retirement like archival absorption. The audit question above already ruled that out.&lt;/p&gt;

&lt;p&gt;Agents completing a "retire safely" goal default toward documenting every visible artifact, because nothing marks validated conclusions versus provisional hacks unless you write an explicit discard list. An earlier architecture pass had already elevated some of that demo protocol as if it were product design; retirement needed de-specification, not more absorption.&lt;/p&gt;

&lt;p&gt;Two different leftovers would have read as precedent. One run had produced five retrospective cards from eight captured learnings; that still looked like evidence even though the pipeline was unproven. A forty-card demo cap would have landed the same way: the next agent writes "last time we capped retros at forty cards" into architecture even when the experiment never validated that limit.&lt;/p&gt;

&lt;p&gt;My retirement plan named what not to absorb so the real product could be designed from validated direction plus product constraints, not inherited defaults.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The second failure mode is deleting before you unlink.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Live GitHub and Notion links in architecture notes look harmless until the repo is gone and the URLs 404. Agents (and future-you) follow those links, infer missing context, or treat dead references as signals that something was lost mid-migration.&lt;/p&gt;

&lt;p&gt;A docs-only absorption step had to land first: reframe the experiment in past tense, strip live experiment URLs, drop protocol details the experiment never validated, and point readers at the written record.&lt;/p&gt;

&lt;p&gt;Only after architecture notes were self-contained did I delete the GitHub repo, the Notion databases, and the local checkout.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The third failure mode is treating search results as deletion targets.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When you have three related Notion databases (raw learnings, retrospectives, improvement ideas) linked by relations, you need a verified deletion set before you touch anything. My retirement plan listed exact database IDs and data source IDs, verified independently against the live workspace, and deleted them in reverse dependency order: improvement ideas, then retrospectives, then raw learnings.&lt;/p&gt;

&lt;p&gt;Stray-page search was allowed only as inspection: find standalone pages that might belong to the experiment, report them, and do not delete anything that search returns unless you can prove it is experiment-only.&lt;/p&gt;

&lt;p&gt;That rule exists because search is retrieval, not enumeration. A query that mentions "learning" or the experiment project name will surface unrelated pages. An agent under pressure to "clean up" can easily delete the wrong surface if search hits become the allowlist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why retirement needs semantics
&lt;/h2&gt;

&lt;p&gt;Agentic workflows lower the cost of spinning up experiment repos, skills, Notion schemas, and multi-root checkouts to test a hypothesis in isolation. Nothing in the default toolchain lowers the cost of retiring them safely, so teams accumulate ambiguous surfaces that agents treat as canonical.&lt;/p&gt;

&lt;p&gt;Those three failure modes are what retirement looks like without explicit decisions. The sequence that held up looked like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;What it did&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Plan&lt;/td&gt;
&lt;td&gt;Write down what will be deleted, in what order, and what must be captured first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Doc absorption&lt;/td&gt;
&lt;td&gt;Make architecture notes self-contained: past tense, no live experiment links&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delete external&lt;/td&gt;
&lt;td&gt;Delete GitHub repo, Notion DBs, and local checkout with exact ID allowlists, under human verification&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Whether that retirement contract should live in a central registry, per-repo retirement plans, or something else is still an open design question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; Answering the hypothesis is not retirement. Capture validated conclusions in durable docs, make those docs self-contained, then delete external experiment surfaces from a verified allowlist under human verification.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>workflow</category>
    </item>
    <item>
      <title>The board came back. The highlights lied.</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Fri, 04 Sep 2026 02:26:58 +0000</pubDate>
      <link>https://dev.to/michaeltruong/the-board-came-back-the-highlights-lied-18bo</link>
      <guid>https://dev.to/michaeltruong/the-board-came-back-the-highlights-lied-18bo</guid>
      <description>&lt;p&gt;I ship &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=persist-game-state-not-ephemeral-ui-intent&amp;amp;utm_content=intro" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;, a web game. Each game mode keeps its own save in &lt;code&gt;localStorage&lt;/code&gt;. Reload the tab, switch to a different mode, come back later: the board, turn, clue history, and in-progress results all come back. That felt like a win until I added spymaster clue targeting.&lt;/p&gt;

&lt;p&gt;While drafting a clue, I can click cards on my team. Those clicks sync the clue count and show which words I had in mind. They are visual intent only. They are not part of the clue submission payload.&lt;/p&gt;

&lt;p&gt;I assumed that if I reloaded the same game, the UI could restore those highlights too. Same session, same cards, same mental model. On a single-mode refresh that was mostly harmless.&lt;/p&gt;

&lt;p&gt;Switching game modes was not. Each mode loads its own 25-card board. Leftover targeting clicks stayed in memory. If a word on the new board matched a card I had highlighted in the previous mode, that new card lit up. I had never clicked it. The board was truthful. The UI was lying about what I was drafting.&lt;/p&gt;

&lt;p&gt;Clearing those highlights when the board changed would have stopped the lie. That cheaper fix was not enough. I had already assumed a same-game reload should bring the clicks back. Keeping that assumption and stopping the leak meant persisting the clicks per mode, the same way we persist the board. That would have treated a thinking aid like a move. The collision forced the real question: should coming back restore those clicks at all?&lt;/p&gt;

&lt;h2&gt;
  
  
  What coming back restores
&lt;/h2&gt;

&lt;p&gt;People now expect drafts to survive a reload. Google Docs made that the default: leave, come back, the paragraph is still there. Autosave is already on in this game. A lot of players never press Save. They just come back.&lt;/p&gt;

&lt;p&gt;A saved snapshot carries the board, the turn, typed clue fields, history, and any in-flight result. It omits targeting clicks on purpose. The snapshot type documents that omission in &lt;code&gt;gamePersistence.ts&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A clue you typed already works the Docs way. An unsubmitted AI-generated clue does not. Refresh asks the model again. The clue word and the highlighted targets can change (&lt;code&gt;INSECT&lt;/code&gt; and one card can come back as &lt;code&gt;BODYPART&lt;/code&gt; and two different cards). That generation has not crossed into accepted game state. A useful consequence is that a model experiment, upgrade, swap, or config change can take effect on the next reload instead of replaying the last output.&lt;/p&gt;

&lt;p&gt;Human targeting clicks are the same shape: a thinking aid around a later submit, not a move on the board.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Restored from snapshot&lt;/th&gt;
&lt;th&gt;Cleared or regenerated on restore&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Board, revealed cards, team, outcome, and pending guess result&lt;/td&gt;
&lt;td&gt;Human-intended target highlights&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Human-typed clue word and count&lt;/td&gt;
&lt;td&gt;Unsubmitted AI clue and targets&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Submitted clue history&lt;/td&gt;
&lt;td&gt;Manual clue-count override flag&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is why "persist domain state, discard UI state" is the wrong summary. Some UI state should persist (the typed clue). Some AI-generated state should not, even though it occupies the same fields. React state versus a domain object does not tell you enough. The snapshot has to encode what the product has accepted as true, not whatever happened to exist in the UI, and not because anyone pressed Save.&lt;/p&gt;

&lt;h2&gt;
  
  
  What restore actually does
&lt;/h2&gt;

&lt;p&gt;Switching modes and reloading the tab both restore the saved game, then clear the drafting pose. Targets, the manual count override, and any cached AI overlay go with it. The same clear runs on new game and after a successful submit, so a leftover thinking aid cannot leak into the next turn.&lt;/p&gt;

&lt;p&gt;So restore is not "rehydrate everything the component used to know." It is "restore the game, then clear the drafting pose."&lt;/p&gt;

&lt;p&gt;The restored count is durable user work: the number you typed comes back. What does not come back is the live session flag that blocked auto-sync from targets. After reload, the first target selection re-derives count from the grid. That is the same interaction as changing targets in the same session without reloading (&lt;code&gt;type 5 → change targets → derived count&lt;/code&gt; matches &lt;code&gt;type 5 → reload → select a target → derived count&lt;/code&gt;). The override flag is transient drafting pose, not durable provenance like human-typed clue word vs unsubmitted AI clue.&lt;/p&gt;

&lt;p&gt;The regression tests I care about are behavioral:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Grid reference:&lt;/strong&gt; highlight on board A, switch to board B that shares the word, assert nothing is lit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provenance:&lt;/strong&gt; after reload, a human-typed clue survives while an unsubmitted model-generated clue is regenerated. The test controls the model's second answer so replaying the saved output fails deterministically.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count after reload:&lt;/strong&gt; manually typed count survives reload; the first target selection afterward re-derives count from the grid (same as in-session target toggles).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Highlights only reflect choices made in the current drafting session on the current grid.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same problem, different surfaces
&lt;/h2&gt;

&lt;p&gt;The leftover highlight was not a styling bug. Dropping the clicks from the snapshot fixed the category: a thinking aid is not current truth.&lt;/p&gt;

&lt;p&gt;If you are building agent UIs with drafts, wizards, or "thinking aloud" interactions, coming back poses the same question. Autosave does not settle what on the screen is current truth. Keep proposals and thinking-aloud hints ephemeral, even when they sit in the same inputs as accepted work. Document that acceptance boundary in the snapshot type so the next contributor does not "helpfully" persist whatever the component last held.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; Coming back should restore what the product treats as current truth, not whatever happened to be on screen. Accepted work stays. Proposals and thinking-aloud hints do not, even when they share the same fields. The persistence boundary is semantic, not architectural. Test what coming back looks like, not whether the save still loads.&lt;/p&gt;




&lt;p&gt;If you'd like to see the project that inspired these lessons, you can try &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=persist-game-state-not-ephemeral-ui-intent&amp;amp;utm_content=footer" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>typescript</category>
    </item>
    <item>
      <title>I was solving agent portability at the wrong boundary</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Fri, 28 Aug 2026 01:12:13 +0000</pubDate>
      <link>https://dev.to/michaeltruong/i-was-solving-agent-portability-at-the-wrong-boundary-1406</link>
      <guid>https://dev.to/michaeltruong/i-was-solving-agent-portability-at-the-wrong-boundary-1406</guid>
      <description>&lt;p&gt;Copying the last repo's agent setup into a new one worked at first. It also copied product-specific assumptions. Once I had several active projects, there was no longer a single canonical repo I could copy from. Every improvement now had several places it could drift.&lt;/p&gt;

&lt;p&gt;I was doing that across &lt;a href="https://codenames-ai.com/" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;, a portfolio site, and a resume generator. Shared workflows for planning, editorial work, dependency upgrades, review, and repo bootstrap had accumulated around them. Some still lived inside product repos simply because that was where they had evolved.&lt;/p&gt;

&lt;p&gt;I wanted the next repo to start with the methodology already available, without cloning the implementation details of the last product.&lt;/p&gt;

&lt;p&gt;My first instinct was to solve that with an MCP-shaped architecture. Extract the shared behavior into a separate repository. Expose it through a remote tool boundary. Every repo could call the same capability when it needed merge-safe planning rules or editorial workflow guidance.&lt;/p&gt;

&lt;p&gt;That felt rigorous. One service. One contract. One place to version policy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the MCP-shaped instinct looked right
&lt;/h2&gt;

&lt;p&gt;The idea lived in notes and conversation: treat portable agent policy the way you would treat a remote tool.&lt;/p&gt;

&lt;p&gt;Building it would have bought a clean contract. It would also have attached the baggage that belongs to real tools:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a runtime or service boundary&lt;/li&gt;
&lt;li&gt;an MCP contract and deployment story&lt;/li&gt;
&lt;li&gt;versioning and invocation decisions&lt;/li&gt;
&lt;li&gt;"when do we call this?" routing inside every agent session&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That overhead makes sense when the capability is genuinely external: query Notion, pull PostHog metrics, deploy through Vercel. It does not make sense when the capability is mostly &lt;strong&gt;operating methodology&lt;/strong&gt;: how to slice plans, when to stop after opening a PR, how to keep merge-safe invariants explicit.&lt;/p&gt;

&lt;p&gt;Planning standards and merge-safe workflows are agent policy and procedure. They are not remote resources waiting behind a tool boundary.&lt;/p&gt;

&lt;p&gt;I never built that service. I did not need to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What shipped instead: four durable scopes
&lt;/h2&gt;

&lt;p&gt;The decomposition was the work: what should follow me into every repo, what should stay behind a tool boundary, what should stay with me as procedure, and what must live in the repo itself. I happened to implement that split in Cursor (user-level rules, MCP config, installable skills, repo files). The architecture is the scopes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;What belongs here&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Always-on policy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Invariants in every repo&lt;/td&gt;
&lt;td&gt;Short rules: execution authority, repository topology, stop-after-open, tool preferences, pointer to planning methodology.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Shared tools&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;External tool boundaries&lt;/td&gt;
&lt;td&gt;Notion, PostHog, Vercel after you authenticate the service.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Reusable procedures&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Workflow installation&lt;/td&gt;
&lt;td&gt;Full portable skills such as staged planning and new-repo bootstrap.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Repo files&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;What must live in the repo&lt;/td&gt;
&lt;td&gt;Stable mechanics (local hooks, remote environment lifecycle, CI) and product knowledge (&lt;code&gt;AGENTS.md&lt;/code&gt;, domain rules, product skills, review guides)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Always-on policy&lt;/strong&gt; kept a pointer to the planning methodology instead of a second copy of it. Copying one product repo's harness into another just to match would have recreated the drift.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shared tools&lt;/strong&gt; stay behind that authenticated service boundary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reusable procedures&lt;/strong&gt; stay at user level. The new repo does not store those skills. I run bootstrap once to write &lt;strong&gt;stable mechanics&lt;/strong&gt; into repo files. That step does not copy the skill into the repo, and it does not write &lt;strong&gt;product knowledge&lt;/strong&gt; because it varies by product.&lt;/p&gt;

&lt;p&gt;The scopes also have different update semantics: policy changes flow across existing repos, while bootstrap changes become the baseline for new ones unless I explicitly migrate older repos.&lt;/p&gt;

&lt;h2&gt;
  
  
  The interview showed what still had to be reconstructed
&lt;/h2&gt;

&lt;p&gt;The first serious cold-start test was a timed AI-native product-build interview. I used the same scopes in a genuinely new repo under time pressure.&lt;/p&gt;

&lt;p&gt;The interview proved the methodology did not depend on my existing repos. It also showed that too much generic setup still had to be reconstructed in an empty one. The agent put instructions in the README instead of &lt;code&gt;AGENTS.md&lt;/code&gt;. Hooks that should have wired the remote agent environment were not reliably set up. Setup that was obvious in my established repos was not obvious when an agent had to invent it under time pressure.&lt;/p&gt;

&lt;p&gt;Policy, procedures, and shared tools were already available. What failed was leaving stable repo mechanics to be rediscovered. An empty repo still has to run those hooks and that remote environment. The following week I converted more of that baseline into deterministic bootstrap: known-good scripts and templates for local hooks, remote environment lifecycle, CI, and other baseline infrastructure, not another round of agent redesign.&lt;/p&gt;

&lt;p&gt;The goal is not zero bootstrap. It is to stop spending agent reasoning on decisions I have already made.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; Splitting the problem by ownership and lifecycle showed I did not need an MCP-shaped architecture. Policy could stay always-on, procedures could stay at user level, stable mechanics could be materialized by running bootstrap, and the agent could spend its reasoning on the product.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Editor's note (September 2026)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Since publishing this, I extracted the implementation behind this approach into an open-source workflow and shipped it as a Cursor plugin. It includes the portable agent policy, reusable skills, and deterministic repo bootstrap described above.&lt;/p&gt;

&lt;p&gt;The implementation is available in the &lt;a href="https://github.com/multipliers-dev/cursor-team-marketplace" rel="noopener noreferrer"&gt;multipliers-dev/cursor-team-marketplace&lt;/a&gt; repository on GitHub.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>workflow</category>
      <category>cursor</category>
    </item>
    <item>
      <title>The pipeline was green. The product was underspecified</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Mon, 17 Aug 2026 08:23:00 +0000</pubDate>
      <link>https://dev.to/michaeltruong/the-pipeline-was-green-the-product-was-underspecified-1fnj</link>
      <guid>https://dev.to/michaeltruong/the-pipeline-was-green-the-product-was-underspecified-1fnj</guid>
      <description>&lt;p&gt;The checks all passed. I still would not have sent the resume.&lt;/p&gt;

&lt;p&gt;I had been using a Cursor agent to implement a private &lt;strong&gt;facts → prose&lt;/strong&gt; resume generator. Structured career claims in, recruiter-facing PDFs out. Generation, rendering, and ATS checks all stayed green. The PDFs looked plausible.&lt;/p&gt;

&lt;p&gt;I treated that as enough.&lt;/p&gt;

&lt;p&gt;Several product requirements were still wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What implementation QA was checking
&lt;/h2&gt;

&lt;p&gt;The useful split is structured facts on one side and disposable rendered artifacts on the other:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Holds&lt;/th&gt;
&lt;th&gt;Does not hold&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Structured facts&lt;/td&gt;
&lt;td&gt;Stable claims (actions, outcomes, metrics, scope)&lt;/td&gt;
&lt;td&gt;Resume bullet phrasing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Application config&lt;/td&gt;
&lt;td&gt;Which facts to include, tone, theme, page length&lt;/td&gt;
&lt;td&gt;New career claims&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generated output&lt;/td&gt;
&lt;td&gt;Markdown and PDF resumes&lt;/td&gt;
&lt;td&gt;Source of truth&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Implementation QA in this workflow means the pipeline runs end to end and the automated checks pass. Generation succeeds. PDFs render. ATS scripts assert page counts, required sections, and a few structural rules about separators and headings.&lt;/p&gt;

&lt;p&gt;Those checks are real. They caught broken builds and regressions I did not want to ship.&lt;/p&gt;

&lt;p&gt;They did not answer a different question: did the specification describe the resume I actually wanted?&lt;/p&gt;

&lt;p&gt;That gap showed up in three places. None of them failed the scripts at first.&lt;/p&gt;

&lt;h3&gt;
  
  
  When page count is not product fit
&lt;/h3&gt;

&lt;p&gt;The workflow produced one-page and two-page variants for the same application. Both PDFs passed &lt;code&gt;check:ats&lt;/code&gt;. Both stayed inside their page limits.&lt;/p&gt;

&lt;p&gt;The two-page resume still read like a stretched one-pager.&lt;/p&gt;

&lt;p&gt;Experience on the shorter version used concise evidence entries selected for a tight one-page fit. The longer version reused that same condensed slice, then filled the remaining space with additional facts. Page count was correct, but the shape was wrong: the original concise bullets never expanded into fuller evidence. A two-page resume should deepen the roles that already earned a place, not keep the one-page wording and pad with more items.&lt;/p&gt;

&lt;p&gt;Mismatched typography could fake the same green result by enlarging text on the longer PDF. Unifying the shared typographic scale and margins removed that shortcut: a two-page count had to come from the evidence itself.&lt;/p&gt;

&lt;p&gt;Implementation QA had no opinion about which facts belonged on which page length, or whether two pages meant more evidence or just larger fonts. It only knew the PDF had two pages.&lt;/p&gt;

&lt;h3&gt;
  
  
  When valid data reads wrong to a human
&lt;/h3&gt;

&lt;p&gt;Not every miss was about page length.&lt;/p&gt;

&lt;p&gt;One Program Lead role was technically valid in the data: correct dates, correct employer, correct title. In the Experience section it rendered with a &lt;code&gt;Full-time&lt;/code&gt; employment label. On paper that is accurate enough for a schema. On a resume it reads like a sequential primary job when the role was actually concurrent with other work.&lt;/p&gt;

&lt;p&gt;The requirement was not "store valid employment metadata." It was "make concurrent work legible to a recruiter scanning the ladder." Renaming the label to &lt;code&gt;Concurrent program&lt;/code&gt; was a product fix, not a pipeline fix. No ATS script flagged the old wording.&lt;/p&gt;

&lt;h3&gt;
  
  
  When structural correctness stood in for finish
&lt;/h3&gt;

&lt;p&gt;The last category looked optical. It was still underspecification.&lt;/p&gt;

&lt;p&gt;I only noticed after opening the PDF: the contact block and Skills sidebar shared a column edge on paper, but the Skills heading sat a few points lower than Experience, so the two-column header row looked crooked even though every section and separator rule still passed. ATS-safe contact separators are a constraint, not a design.&lt;/p&gt;

&lt;p&gt;The workflow had encoded structural correctness: page counts, required sections, separator rules, heading shape. It had never named a human visual-acceptance criterion. The scripts were never asked to stand in for a finish requirement the spec had never named: would I send this?&lt;/p&gt;

&lt;h2&gt;
  
  
  What manual implementation used to smuggle in
&lt;/h2&gt;

&lt;p&gt;When I wrote this kind of tooling by hand, implementation and requirements review were harder to separate. Every intermediate decision was visible. Choosing a font size, picking which bullet to cut, or rewriting a concurrent-role label forced a product judgment in the same session as the code change.&lt;/p&gt;

&lt;p&gt;Manual implementation was accidentally doing requirements QA. Every ambiguous decision eventually became my problem because I had to turn it into code myself. An agent can absorb that ambiguity instead.&lt;/p&gt;

&lt;p&gt;The two-page PDF was that pattern in miniature. It looked finished. The missing product decision (whether two pages meant deeper evidence) never came back through me.&lt;/p&gt;

&lt;p&gt;That speed is the danger now. Incomplete specification can arrive dressed as a finished product.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/michaeltruong/build-looked-absurd-under-a-recruiter-deadline-1145"&gt;An earlier piece&lt;/a&gt; was about why building this system suddenly made economic sense. This is the other side of that shift: once implementation got cheaper, I needed to make requirements review more explicit.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a requirements QA stage could look like
&lt;/h2&gt;

&lt;p&gt;The portable version is a few jobs, not a particular toolchain:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Force the vague request into concrete behavior before implementation: what changes, what stays invariant, and what counts as done&lt;/li&gt;
&lt;li&gt;Review against that written intent, not only the local diff&lt;/li&gt;
&lt;li&gt;Accept the finished artifact as a product ("would I send this?"), not only as a green pipeline&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first two are where requirements QA is cheapest. If those questions stay implicit, an agent can execute the wrong thing extremely efficiently.&lt;/p&gt;

&lt;p&gt;In my workflow that looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;plan                        (overall intent)
→ slice                     (concrete behavior)
→ review slices             (gaps and boundaries before any code)
→ for each slice:
    → implementation
    → review after code     (still matches plan?)
→ artifact acceptance       (ready to send?)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I commit the plan alongside the work so the intent lives in the repo, not only in the chat. A reviewer is not limited to the current diff. It can see where this slice is supposed to lead, and flag something that is locally correct but inconsistent with a later slice. Reviewing the slices can still challenge assumptions, missing requirements, and slice boundaries before any code exists.&lt;/p&gt;

&lt;p&gt;The ladder above is how I recreate the interrogation that manual implementation used to provide implicitly. Planning and review still miss gaps that never made it into the written requirements. Artifact acceptance is the last gate for those: it catches product misses the earlier rungs never named.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; Agents are good at satisfying the requirements they are given, quickly enough that a miss looks finished. Implementation QA proves the system did what you asked. Somebody still has to QA whether those requirements describe the product you actually wanted.&lt;/p&gt;




&lt;p&gt;If you'd like to see the same gap between green checks and product acceptance on a product I keep revising in public, try &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=ai-workflows-need-a-requirements-qa-stage&amp;amp;utm_content=footer" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>workflow</category>
    </item>
    <item>
      <title>AI changed the build-vs-buy threshold</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Thu, 13 Aug 2026 06:39:39 +0000</pubDate>
      <link>https://dev.to/michaeltruong/build-looked-absurd-under-a-recruiter-deadline-1145</link>
      <guid>https://dev.to/michaeltruong/build-looked-absurd-under-a-recruiter-deadline-1145</guid>
      <description>&lt;p&gt;Building custom software to solve a two-afternoon problem sounded absurd.&lt;/p&gt;

&lt;p&gt;A Riot Games recruiter reached out while I was still preparing to return to the job market. Suddenly I needed a current resume to send back, and I had roughly two afternoons to produce one.&lt;/p&gt;

&lt;p&gt;Normally that is an obvious &lt;strong&gt;buy&lt;/strong&gt; decision. Under a short deadline you are not optimizing for reuse. You are optimizing for a PDF in someone's inbox. A resume builder gives you templates, export, and enough polish to look professional without inventing infrastructure.&lt;/p&gt;

&lt;p&gt;Historically, I built when the reuse justified the setup cost. I bought or assembled manually when I only needed the artifact.&lt;/p&gt;

&lt;p&gt;The same rule still applied. What had changed was the cost.&lt;/p&gt;

&lt;p&gt;AI had lowered not only how much it cost to build the first version, but how much it cost to keep revising the architecture underneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built instead
&lt;/h2&gt;

&lt;p&gt;I built a private &lt;strong&gt;facts → prose&lt;/strong&gt; resume repository with Cursor.&lt;/p&gt;

&lt;p&gt;The idea is to separate career evidence from application wording:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Holds&lt;/th&gt;
&lt;th&gt;Does not hold&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Structured facts&lt;/td&gt;
&lt;td&gt;Stable claims (actions, outcomes, metrics, scope)&lt;/td&gt;
&lt;td&gt;Resume bullet phrasing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Application config&lt;/td&gt;
&lt;td&gt;Which facts to include, tone, theme&lt;/td&gt;
&lt;td&gt;New career claims&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generated output&lt;/td&gt;
&lt;td&gt;Markdown and PDF resumes&lt;/td&gt;
&lt;td&gt;Source of truth&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Career claims live once in structured YAML. Each application selects, reorders, and rephrases them. &lt;code&gt;npm run generate&lt;/code&gt; renders recruiter-facing prose. &lt;code&gt;npm run pdf&lt;/code&gt; prints it. &lt;code&gt;npm run check:ats&lt;/code&gt; runs structural ATS checks on the output.&lt;/p&gt;

&lt;p&gt;You do not need my private repo to apply the pattern. The useful split is structured facts on one side and disposable rendered artifacts on the other.&lt;/p&gt;

&lt;p&gt;Before generating a resume, the workflow researched the company and role, then used that context to decide which evidence from my career inventory belonged in the application. The system knew about far more career evidence than any one resume should contain. The inventory stayed put. What changed was which slice mattered.&lt;/p&gt;

&lt;p&gt;For a gaming company like Riot Games, the research made a university game-design award relevant enough to surface on a software engineering resume where it normally would not belong. The same run also produced a one-page and two-page resume plus a recruiter reply.&lt;/p&gt;

&lt;p&gt;A week later, I used the same career inventory for an AI product engineering company. The underlying evidence had not changed, but what mattered had. The workflow selected a different slice of it. The reuse I had been optimizing for was already real.&lt;/p&gt;

&lt;p&gt;Later it expanded past resumes into interview prep materials, a use I had not planned when the recruiter first wrote. Different company, different evidence selection. Different stage, different artifact altogether. The system that looked overbuilt for one reply kept finding uses.&lt;/p&gt;

&lt;h2&gt;
  
  
  The system did not start here
&lt;/h2&gt;

&lt;p&gt;The repository did not begin with that model.&lt;/p&gt;

&lt;p&gt;It evolved through increasingly useful abstractions across the same short build window:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A single resume document&lt;/li&gt;
&lt;li&gt;A reusable resume template&lt;/li&gt;
&lt;li&gt;Career facts separated from prose&lt;/li&gt;
&lt;li&gt;Capabilities grouping related achievements&lt;/li&gt;
&lt;li&gt;Evidence entries inside each capability&lt;/li&gt;
&lt;li&gt;Application-specific selection over that inventory&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first workable version was not the last one. Because implementation and revision were cheap, I could keep moving instead of freezing at "good enough for tonight."&lt;/p&gt;

&lt;p&gt;Without that cost shift, the rational stopping point would probably have been a reusable template or a lightly parameterized document. Fine for one application. Weak as career infrastructure.&lt;/p&gt;

&lt;p&gt;That was the advantage of the build path here: room to discover the right abstraction after the first one works.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two costs moved, not one
&lt;/h2&gt;

&lt;p&gt;AI did not magically make &lt;strong&gt;build&lt;/strong&gt; correct for every problem.&lt;/p&gt;

&lt;p&gt;It lowered two costs at once:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Implementation cost:&lt;/strong&gt; scaffolding the generator, themes, and checks stopped being a multi-week side project.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Architectural iteration cost:&lt;/strong&gt; revising the underlying model (facts vs capabilities vs application selection) stayed cheap enough to do in the same session.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Buying still wins when the problem is narrow, the tool fits, and you will not reuse the result. Building still loses when maintenance will crush you.&lt;/p&gt;

&lt;p&gt;What changed is the boundary. Problems that used to land firmly on the buy side can cross over when reuse matters and you can afford to iterate past the first design.&lt;/p&gt;

&lt;h2&gt;
  
  
  Buy was still faster for one send
&lt;/h2&gt;

&lt;p&gt;This is not an argument that custom software always beats SaaS. I chose to build because the deadline still left room for something reusable. If the goal had been a single resume, buying an off-the-shelf builder would have been the faster path.&lt;/p&gt;

&lt;p&gt;The build path made sense here because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The career inventory is the stable core.&lt;/li&gt;
&lt;li&gt;Research and role context drive application-specific selection.&lt;/li&gt;
&lt;li&gt;The same system produced multiple resume variants and a recruiter reply.&lt;/li&gt;
&lt;li&gt;It later expanded into interview prep materials.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cheap iteration on both implementation and revision is what made exploring these abstractions affordable.&lt;/p&gt;

&lt;p&gt;The same economics show up outside resumes: internal tooling, dashboards, documentation pipelines, code generators, personal workflow infrastructure. Anywhere the old math was "custom software is too expensive to build &lt;strong&gt;and revise&lt;/strong&gt;", the revise term got smaller.&lt;/p&gt;

&lt;p&gt;I suspect that generalizes beyond software engineering into knowledge work where bespoke systems used to lose to packaged tools on turnaround alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; AI moves the build-vs-buy threshold by cutting both implementation cost and the cost of changing your mind about architecture. When reuse matters and you can iterate in the same sprint, building a small internal system can become rational where buying would previously have won on cost and turnaround.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Editor's note (October 2026)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This piece is about why building the job-search system made economic sense under a recruiter deadline. A follow-up, &lt;a href="https://dev.to/michaeltruong/records-first-chart-second-the-sankey-disagreed-3841"&gt;I visualised my job search. The data model started breaking.&lt;/a&gt;, picks up once enough hiring histories had accumulated. By then, extending the system with lifecycle analytics was cheap enough to be another incremental feature, and putting those histories on a Sankey diagram exposed semantic problems the YAML validators alone could not see.&lt;/p&gt;




&lt;p&gt;If you'd like to see the same economics on a product I keep revising in public, try &lt;a href="https://codenames-ai.com/" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>workflow</category>
    </item>
    <item>
      <title>One skill per action looked like the safe boundary</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Fri, 07 Aug 2026 05:24:54 +0000</pubDate>
      <link>https://dev.to/michaeltruong/one-skill-per-action-looked-like-the-safe-boundary-13pj</link>
      <guid>https://dev.to/michaeltruong/one-skill-per-action-looked-like-the-safe-boundary-13pj</guid>
      <description>&lt;p&gt;I started with a rule that felt like good engineering: &lt;strong&gt;one skill per action&lt;/strong&gt;. Create a card here. Enrich it there. Reclassify it somewhere else. Each prompt got a clean boundary. Each file stayed small.&lt;/p&gt;

&lt;p&gt;Then that decomposition started to fight the domain.&lt;/p&gt;

&lt;p&gt;I ran into this while building an AI-assisted editorial workflow in Cursor, but the problem was not really about Cursor or publishing. It was about where an agent capability should begin and end.&lt;/p&gt;

&lt;p&gt;In that setup, a &lt;strong&gt;skill&lt;/strong&gt; is a markdown file the agent loads for a workflow. My editorial workflow uses those skills to create and manage Notion cards before drafting posts.&lt;/p&gt;

&lt;h2&gt;
  
  
  When every action gets its own skill
&lt;/h2&gt;

&lt;p&gt;The first version of my inbox skill only created Inbox cards. That matched early usage: capture an observation, normalize it into a canonical shape, attach a small set of grounded references, and stop. The skill was create-only and treated each new card as immutable once it left Inbox.&lt;/p&gt;

&lt;p&gt;As capture matured, the same card kept needing more work &lt;strong&gt;while it was still Inbox&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Create&lt;/strong&gt; when a new observation arrived&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enrich&lt;/strong&gt; when new evidence or framing changed the normalized shape&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reclassify&lt;/strong&gt; when routing rules decided the card should be sparse vs rich, or a quick note vs a planned blog post&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those felt like three different jobs. They had different verbs. They had different retrieval triggers. Splitting them into separate skills seemed obvious.&lt;/p&gt;

&lt;p&gt;The split failed in ordinary corrections. I would enrich an Inbox card with new evidence, then realize it should change from a quick note into a planned blog post. That meant a second skill invocation for reclassify, with a second copy of the same lifecycle rules. For a moment it was unclear which skill was still responsible for keeping the page as one current write-up instead of an accumulating edit history. Get the order wrong and you did the work twice: an enrich that left the card type stale, or a reclassify that ignored the evidence rewrite you still needed.&lt;/p&gt;

&lt;p&gt;The deeper problem was ownership. All three operations touched the &lt;strong&gt;same owned object&lt;/strong&gt;: a single Inbox card in a Notion database. They shared the same lifecycle gate (the card must stay in Inbox), the same mutation boundaries (never change lifecycle status, source path, or published URLs on existing pages), and the same routing rules for how sparse or rich the card should be and whether it was a quick note or a planned post. They also shared the same rule: after every update, the page must hold exactly one current normalized write-up and exactly one captured observation. Notion history is the revision log; the operational card is not an append-only audit log.&lt;/p&gt;

&lt;p&gt;Treating Create, Enrich, and Reclassify as three skills meant three prompts trying to enforce one coherent capability. The boundaries were at the wrong layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Consolidation made the skill clearer
&lt;/h2&gt;

&lt;p&gt;The fix was to stop pretending those were separate capabilities. One inbox skill now owns the full Inbox lifecycle: &lt;strong&gt;Create&lt;/strong&gt;, &lt;strong&gt;Enrich&lt;/strong&gt;, and &lt;strong&gt;Reclassify&lt;/strong&gt; while the card remains in Inbox.&lt;/p&gt;

&lt;p&gt;The operator still invokes one skill. The skill routes internally:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;User-facing command&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;create inbox card&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Creates an Inbox card with a canonical write-up and grounded references&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;enrich inbox card&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Folds new evidence into the existing Inbox page as one current write-up&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same-thread continuation&lt;/td&gt;
&lt;td&gt;Treats further capture in the same chat as enrich on the page just created&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Reclassify&lt;/strong&gt; is not a separate user command. It is detected inside enrich when the routing rules decide the card should be richer or thinner than before, or should shift from a quick note to a planned post. Sparse captures can become richer field reports. The skill rebuilds the current Inbox representation when those derived choices change; it does not append enrichment history sections.&lt;/p&gt;

&lt;p&gt;That consolidation expanded the inbox skill from create-only immutability to Inbox-lifecycle ownership. The skill file grew, but the &lt;strong&gt;system&lt;/strong&gt; got simpler: one place owns Inbox normalization, one shared rule set, one place with authority over the rules.&lt;/p&gt;

&lt;p&gt;That was the opposite of what I expected. I thought a bigger skill file would feel heavier. Instead routing got easier. I stopped wondering which inbox skill to invoke for a correction vs a note-to-post change. I invoked the inbox skill, and the routing rules decided whether enrich included reclassify.&lt;/p&gt;

&lt;h2&gt;
  
  
  One capability, many internal operations
&lt;/h2&gt;

&lt;p&gt;The incident suggests a short consolidate-vs-split test:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Same owned object under the same lifecycle gate.&lt;/strong&gt; If lifecycle stage or artifact type diverges, stop consolidating.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compatible mutation rules and safety boundaries.&lt;/strong&gt; If the operations need conflicting write permissions or rejection rules that cannot share one authority, keep them separate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Same type-and-shape rules across operations.&lt;/strong&gt; Given the same input, if the operations would disagree about what kind of thing it should become, keep them separate.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Those are internal operations, not separate capabilities.&lt;/p&gt;

&lt;p&gt;A parallel already existed elsewhere in the pipeline. A triage skill scores Inbox cards, recommends promotions out of Inbox, and archives the weakest rows. Those are different mutations, but they live inside one triage skill because they share the same queue-review ownership. I did not split score, promote, and archive into three skills. The operations differ; the owned workflow does not.&lt;/p&gt;

&lt;p&gt;That comparison has limits. Triage promotion requires an explicit &lt;strong&gt;apply&lt;/strong&gt; command after a dry-run report. Inbox enrich rejects cards that have already left Inbox. The internal gates differ without splitting ownership. The pattern is still recognizable: &lt;strong&gt;one skill per coherent capability&lt;/strong&gt;, with internal routing between operations.&lt;/p&gt;

&lt;p&gt;That does not mean every related action belongs in one skill. Scheduling, drafting, critique, and publishing own different artifacts and stop lines, so they remain separate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The resource mattered more than the verb
&lt;/h2&gt;

&lt;p&gt;You do not need Cursor to recognize the shape. A REST API does not usually become a separate service for every operation on the same resource.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;POST&lt;/code&gt; creates. &lt;code&gt;PATCH&lt;/code&gt; updates. &lt;code&gt;DELETE&lt;/code&gt; removes. Different operations, same resource contract. Reclassification in my system is just another mutation of that same Inbox card resource.&lt;/p&gt;

&lt;p&gt;Splitting those operations into separate agent skills was like building one service for create, another for update, and a third for delete. The endpoints looked clean in isolation. Ownership of the resource was fragmented.&lt;/p&gt;

&lt;p&gt;Agentic workflows drift toward that decomposition because &lt;strong&gt;actions are easier to name than ownership&lt;/strong&gt;. "Create card" and "enrich card" are vivid verbs. "Own Inbox normalization throughout the Inbox lifecycle" is accurate but abstract. The verbs made the skills easy to name. The resource revealed where the boundary actually belonged.&lt;/p&gt;

&lt;p&gt;Your domains may decompose differently. The useful question is not "how many skills do I have?" but "what object does this capability own, and are these verbs operations on that object or different capabilities entirely?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; Start with one skill per action if that helps you ship. When multiple operations share an owned object, lifecycle, and compatible safety boundaries, consolidate them into one capability module and route internally. The risk was no longer a skill becoming too broad. It was one capability having multiple competing owners.&lt;/p&gt;




&lt;p&gt;If you'd like to see the project behind these workflow experiments, try &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=skills-should-own-capabilities-not-individual-actions&amp;amp;utm_content=footer" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>workflow</category>
    </item>
    <item>
      <title>I expected pair programming with a Cloud Agent. I got a new hire.</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Fri, 31 Jul 2026 05:53:25 +0000</pubDate>
      <link>https://dev.to/michaeltruong/the-first-cloud-agent-felt-less-like-pair-programming-and-more-like-hiring-an-engineer-18j4</link>
      <guid>https://dev.to/michaeltruong/the-first-cloud-agent-felt-less-like-pair-programming-and-more-like-hiring-an-engineer-18j4</guid>
      <description>&lt;p&gt;I thought a cloud coding agent was still pair programming: the same local conversation, just running somewhere else.&lt;/p&gt;

&lt;p&gt;The first thing that surprised me was not the code. It was how little the run needed me.&lt;/p&gt;

&lt;p&gt;My evidence is a first Cursor Cloud Agent run against a real monorepo. What initially looked like a Cursor feature turned out to be a different execution model. The run started over in a fresh environment, did the onboarding work, and left proof. Useful. Autonomous. And missing almost everything I already knew in the local session.&lt;/p&gt;

&lt;h2&gt;
  
  
  The wrong picture
&lt;/h2&gt;

&lt;p&gt;In local Cursor I already had a planning thread: investigation, tradeoffs, and tools I had already authenticated in that session (including Notion MCP). When I pointed a Cloud Agent at &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=the-first-cloud-agent-felt-less-like-pair-programming-and-more-like-hiring-an-engineer&amp;amp;utm_content=intro" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;, an npm-workspaces monorepo with an Express API, a Vite frontend, Playwright E2E, and CI, I expected continuity. Same decisions. Same auth. Same half-finished reasoning, just remote.&lt;/p&gt;

&lt;p&gt;That assumption failed in the first hour.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the first run actually did
&lt;/h2&gt;

&lt;p&gt;The task itself was small: make the repository ready for future Cloud Agent runs and document the non-obvious setup in &lt;code&gt;AGENTS.md&lt;/code&gt;. Completing it required establishing and proving the whole execution environment.&lt;/p&gt;

&lt;p&gt;After that prompt, the agent worked without me sitting in the loop. It cloned the repo, installed dependencies, ran lint, typecheck, build, and tests, installed Playwright browsers, ran the E2E suite, played a Solo turn in the UI (a single-player practice game), captured screenshots and a walkthrough video, and opened a pull request.&lt;/p&gt;

&lt;p&gt;That pull request added the &lt;code&gt;AGENTS.md&lt;/code&gt; notes and recorded the verification trail from the clean VM, including a hello-world Solo turn in the UI. The change set was small. The &lt;em&gt;behavior&lt;/em&gt; was large: autonomous environment setup plus artifacts local agent workflows rarely leave behind.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two failures I did not expect
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Local reasoning stayed on the laptop.&lt;/strong&gt; Choices I had already made in Cursor (what mattered, what to skip, how I was framing the job) were not present in that new cloud task. Anything that depended on that judgment had to be re-established. The agent could reach the repo and CI. It could not inherit the argument I had already had with myself. Other Cursor flows can move a conversation into the cloud. This fresh task did not arrive with the local framing behind it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool auth did not ride along.&lt;/strong&gt; Notion MCP worked only after separate authentication for the cloud run. Local Cursor access was not session continuity. "The cloud can use MCP" and "the cloud already has my MCP sessions" turned out to be different claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it felt like hiring
&lt;/h2&gt;

&lt;p&gt;Async handoff fit. Open-ended design debate in the cloud did not. The run wanted a bounded job and a definition of done, not a remote pair for figuring out the product.&lt;/p&gt;

&lt;p&gt;You do not onboard someone by forwarding a half-finished Slack thread. You give them a machine, a checklist, and a brief. Missing context hurts more as autonomy increases, because a new clean-environment job starts without your local reasoning.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;AGENTS.md&lt;/code&gt; mattered once the agent was in an unfamiliar checkout: non-obvious Node version floors, optional &lt;code&gt;.env&lt;/code&gt; for local runs, E2E setup, font-sensitive visual snapshots. You do not need the file itself. The point is that this is briefing material for an execution worker, not a substitute for the planning conversation that happened on my laptop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not another local worktree?
&lt;/h2&gt;

&lt;p&gt;A second local git worktree can also provide parallel execution, and it has one obvious advantage: it can reuse the tools, credentials, caches, and machine state I already have. For a single developer, that can make the handoff cheaper.&lt;/p&gt;

&lt;p&gt;But it leaves me managing another local workspace, and it keeps the result coupled to my machine: extra checkouts, local processes, port conflicts, and the risk that the result only works in my environment. A teammate cannot reproduce that run exactly without inheriting my laptop state.&lt;/p&gt;

&lt;p&gt;A cloud agent starts with less inheritance, so the handoff matters more. In return, I get an isolated task that is easy to parallelize, with less "works on my machine" risk, and I can start or monitor work remotely without babysitting another local workspace.&lt;/p&gt;

&lt;p&gt;Cursor was where I encountered the boundary clearly. Clean environments, onboarding, and missing laptop state are not new lessons if you have been shipping to remote servers for years. What was new was seeing that old systems idea reappear inside an AI coding workflow. The useful response is an explicit handoff, not a remote continuation of your session. Wherever coding agents become independently executable, I expect the same trade-off to appear. The boundary is architectural.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where each side fits
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Local workspace
(planning, judgment, brief)
        ↓
Cloud agent
(bounded execution in a clean environment)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That first run left open questions: whether later runs would reuse the onboarding work, and how the pattern would hold beyond one carefully bounded task.&lt;/p&gt;

&lt;p&gt;Looking back, I kept choosing Cloud Agents for bounded execution: maintenance loops (including &lt;a href="https://dev.to/michaeltruong/upgrades-dont-have-to-be-a-blind-trust-exercise-13mj"&gt;dependency upgrades&lt;/a&gt;), targeted bug fixes, and focused investigations such as tracing a review finding or validating a specific question. Not because the first-run costs disappeared, but because isolation and the shared task model reduced coordination overhead. The run still leaves a reviewable trail (screenshots, video, a PR), the same kind of demo evidence you would expect from another engineer handing work back.&lt;/p&gt;

&lt;p&gt;The lasting surprise was that the clean boundary is not only a limitation. It is also the feature that makes the model scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; A cloud agent gives up some inherited local context in exchange for isolation and scalable delegation. Brief it like a new hire, not like a remote continuation of your existing session. The context it needs must cross that boundary deliberately.&lt;/p&gt;




&lt;p&gt;If you'd like to see the project behind these workflow experiments, try &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=the-first-cloud-agent-felt-less-like-pair-programming-and-more-like-hiring-an-engineer&amp;amp;utm_content=footer" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>workflow</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Active players looked real until we asked which sessions counted</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Thu, 23 Jul 2026 04:28:12 +0000</pubDate>
      <link>https://dev.to/michaeltruong/active-players-looked-real-until-we-asked-which-sessions-counted-11em</link>
      <guid>https://dev.to/michaeltruong/active-players-looked-real-until-we-asked-which-sessions-counted-11em</guid>
      <description>&lt;p&gt;I've been building &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=active-players-looked-real-until-we-asked-which-sessions-counted&amp;amp;utm_content=intro" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;, a small web game where an LLM plays Codenames with you. Like most solo products, I glance at a Product Health dashboard when I want a quick read on whether anyone is actually playing.&lt;/p&gt;

&lt;p&gt;One morning in June, three weeks after launching the site, the Active players tile said &lt;strong&gt;64&lt;/strong&gt;. Next to it sat &lt;strong&gt;122&lt;/strong&gt; starts and restores. The number looked like traction. My first instinct was to treat it as confirmation and keep shipping.&lt;/p&gt;

&lt;p&gt;That instinct did not survive the next question: which sessions were actually in that count?&lt;/p&gt;

&lt;h2&gt;
  
  
  The dashboard answered a wider question than I asked
&lt;/h2&gt;

&lt;p&gt;I was reading Product Health as if every event in the project came from real players on the production site. The tile did not lie about its math. It counted distinct people who started or restored a game. What it could not tell me, from the chart alone, was which runtime those people were in.&lt;/p&gt;

&lt;p&gt;I had reasons to trust the number:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;PostHog init only ran when &lt;code&gt;VITE_POSTHOG_KEY&lt;/code&gt; was set. Local Vite and Playwright runs did not ship that key, so I treated laptop and E2E traffic as silent by configuration. There was no &lt;code&gt;analytics_environment&lt;/code&gt; property yet, and no environment-conditional init path. "Do not put the key in this build" was one guardrail.&lt;/li&gt;
&lt;li&gt;Returning users looked safe too. On production, game state restores from origin-scoped &lt;code&gt;localStorage&lt;/code&gt;, and PostHog keeps an anonymous ID on that same origin. Come back later and you still count as one Active player via &lt;code&gt;game_restored&lt;/code&gt;. We do not call &lt;code&gt;identify&lt;/code&gt;; continuity is browser storage on that host. I assumed testing on review URLs worked the same way: me again, already counted.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Outside PostHog, the acquisition picture did not match. Real arrivals were mostly organic Google Search. In Search Console, we had not yet hit the first “30 clicks from Google Search in the past 28 days” milestone. We had only just started posting on &lt;a href="https://dev.to/"&gt;dev.to&lt;/a&gt;, so that channel was not a material source either.&lt;/p&gt;

&lt;p&gt;Sixty-four unique players on a site that young, against a search funnel that had not cleared thirty clicks in a month, and early publishing that barely existed, was already a little suspicious. The starts/restores volume next to it made it worse. My working note was blunt: investigate further; something was minting unique players that real arrivals could not explain.&lt;/p&gt;

&lt;h3&gt;
  
  
  Review deploys were the hole
&lt;/h3&gt;

&lt;p&gt;Review deploys (for us, Vercel preview URLs) look like the real app, often share the same analytics project key, and show up whenever you click a pull-request review link. They were not "local without a key," and they were not the same origin as &lt;code&gt;codenames-ai.com&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A review hostname gets its own empty save store and its own anonymous PostHog identity, so a click-through during review can land as a new unique player (&lt;code&gt;game_started&lt;/code&gt;) instead of folding into the production self I already knew. Without a way to separate those runtimes, that 64 was still a hypothesis about whether review-deploy traffic, and new identities on those hosts, were in the count.&lt;/p&gt;

&lt;p&gt;That investigation became a concrete plan: stop treating every capture in the project as if it were production traffic.&lt;/p&gt;

&lt;p&gt;An early cut disabled PostHog for E2E. Silencing one runtime would still leave review deploys sharing the key; we needed an explicit boundary instead of relying on some environments staying silent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should count as production?
&lt;/h2&gt;

&lt;p&gt;Two fixes landed together.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Client tagging.&lt;/strong&gt; On PostHog init, the frontend resolves an &lt;code&gt;analytics_environment&lt;/code&gt; of &lt;code&gt;production&lt;/code&gt;, &lt;code&gt;preview&lt;/code&gt;, &lt;code&gt;local&lt;/code&gt;, or &lt;code&gt;e2e&lt;/code&gt;, then attaches it to every event and to the user profile. Hostname and the host’s build-time environment distinguish the runtimes.&lt;/p&gt;

&lt;p&gt;Non-production traffic is excluded by dashboard filters, not by skipping PostHog init. Tagging every runtime, including ones we used to silence by omitting the key, is what makes the filter meaningful.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dashboard filters.&lt;/strong&gt; Product Health keeps events where &lt;code&gt;analytics_environment = production OR not set&lt;/code&gt;, so older production events from before tagging remain visible. Newer views can use an exact &lt;code&gt;production&lt;/code&gt; filter once tagging coverage is trusted.&lt;/p&gt;

&lt;p&gt;The missing dimension wasn't another metric. It was the production boundary. Once that existed, Product Health could filter on it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Healthy numbers resist questions
&lt;/h3&gt;

&lt;p&gt;The harder lesson wasn't that the dashboard was wrong. It was that healthy-looking numbers are the least likely ones to get questioned.&lt;/p&gt;

&lt;p&gt;While working on &lt;a href="https://dev.to/michaeltruong/model-experiments-became-an-architectural-stress-test-3gc0"&gt;model experiments&lt;/a&gt;, failure exposed hidden assumptions. Here nothing looked broken, so curiosity had to do the same job: notice that the system was faithfully answering a different question than the one I thought I was asking.&lt;/p&gt;

&lt;p&gt;How we ask the dashboard questions is a separate story. This post stays on the quieter failure mode: one project key, a review runtime that looked like production, and a number that looked clean until we asked.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd check on the next dashboard
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Start by asking what question the metric actually answers, not the one you hope it answers.&lt;/li&gt;
&lt;li&gt;Compare it against an independent signal. If the numbers do not fit together, investigate before celebrating.&lt;/li&gt;
&lt;li&gt;Look for missing dimensions that collapse different kinds of traffic into one KPI: environment, internal users, bots, staging, or another hidden segment.&lt;/li&gt;
&lt;li&gt;Only then decide whether the fix is better tagging, better filtering, or a different metric altogether.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You do not need our dashboards or our app code to apply the pattern. Review deploys were the incident that exposed the gap here.&lt;/p&gt;

&lt;p&gt;I cannot put a clean contamination percentage, from today’s data alone, on the period before we added tagging; the point is the missing question, not a guessed share of noise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; Production engineering isn't just responding to broken signals. It's occasionally distrusting reassuring ones. Metrics answer exactly the question you instrumented, not necessarily the one you think you asked.&lt;/p&gt;




&lt;p&gt;If you'd like to see the project that inspired these lessons, you can try &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=active-players-looked-real-until-we-asked-which-sessions-counted&amp;amp;utm_content=footer" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>analytics</category>
    </item>
    <item>
      <title>Model experiments became an architectural stress test</title>
      <dc:creator>Michael Truong</dc:creator>
      <pubDate>Fri, 17 Jul 2026 15:41:21 +0000</pubDate>
      <link>https://dev.to/michaeltruong/model-experiments-became-an-architectural-stress-test-3gc0</link>
      <guid>https://dev.to/michaeltruong/model-experiments-became-an-architectural-stress-test-3gc0</guid>
      <description>&lt;p&gt;I've been tuning &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=model-experiments-became-an-architectural-stress-test&amp;amp;utm_content=intro" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;, a small web game where an LLM plays Codenames with you. Clue generation is tightly constrained: one word, a count, optional intended targets, JSON on the wire, then deterministic validation before anything reaches the board.&lt;/p&gt;

&lt;p&gt;As the project started attracting regular players, I wanted to improve the gameplay experience without blowing out costs. Moving one model generation from &lt;code&gt;gpt-4o-mini&lt;/code&gt; to &lt;code&gt;gpt-5-mini&lt;/code&gt; was my first instinct.&lt;/p&gt;

&lt;p&gt;The default reasoning setting made responses an order of magnitude slower for this workload. Minimal reasoning looked like the obvious compromise: newer model, responsive gameplay.&lt;/p&gt;

&lt;p&gt;I expected to compare clue quality, latency, and cost while the surrounding prompt, validator, and consumer contracts stayed put.&lt;/p&gt;

&lt;p&gt;That last part was wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment stopped behaving like an A/B test
&lt;/h2&gt;

&lt;p&gt;What showed up was structural, and it showed up in places that had been stable for months.&lt;/p&gt;

&lt;p&gt;Validation failures started rising. Retries started rising. Entire candidate batches started failing before the game ever saw a clue. The sharpest signal came from a clue-selection path that had run untouched for months, and it hard-failed for the first time. They weren't latency regressions so much as architectural ones.&lt;/p&gt;

&lt;p&gt;It is easy to read that as "minimal reasoning made the model worse." More often, the failures were exposing gaps in contracts that had looked fine under the previous model.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each failure actually invalidated
&lt;/h2&gt;

&lt;p&gt;Eventually every failure traced back to one of three layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Prompt contracts&lt;/strong&gt; ask for exactly &lt;code&gt;count&lt;/code&gt; targets and, in batch mode, several distinct candidates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic validators&lt;/strong&gt; reject target/count mismatches and filter invalid candidates before anything downstream runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Downstream consumers&lt;/strong&gt; only see survivors. Empty batches retry with rejection feedback, then fall back if needed.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Those layers share one job: enforce the same invariants. The failures below cut across all three rather than mapping one to one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Side commentary could kill an otherwise usable turn.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;To pick a clue, one strategy (Strange mode) simulates how the AI guesser would respond to each candidate clue, then scores those simulated turns and keeps the best one. I thought those simulations would fail only when the guesses themselves were bad. After the swap, they could also fail because the model attached commentary about other words it had considered, including words that were not even on the board. Because the payload schema included that commentary, the validator had to treat it as part of the same all-or-nothing contract. A payload with usable guesses still got rejected, and when every candidate died that way, the turn came back as a controlled API failure instead of a clue.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Target cardinality had to match the clue count.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I thought my validator was protecting the game. Instead I discovered the previous model had been consistently producing outputs that satisfied those contracts.&lt;/p&gt;

&lt;p&gt;Say the prompt asks for &lt;code&gt;count: 2&lt;/code&gt; and a &lt;code&gt;targets&lt;/code&gt; array with exactly two unrevealed friendly codenames. Under the old model, a clue like &lt;code&gt;{"word": "BUILDING", "count": 2, "targets": ["TOWER", "CASTLE"]}&lt;/code&gt; usually meant two real board words. After the swap, I started seeing the same shape with one valid target and one word that is not on the grid at all, or only a single target when &lt;code&gt;count&lt;/code&gt; was 2. Valid JSON. Perfect keys. Intent status: invalid.&lt;/p&gt;

&lt;p&gt;The validator rejects clues whose validated targets don't match &lt;code&gt;count&lt;/code&gt;. Valid JSON wasn't enough.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retries assumed the contracts were already specific enough.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I thought retries were simply robustness. Instead they became diagnostic tooling because they finally told me which invariant had actually failed. When a batch fails validation, the retry path can attach rejection feedback (failed clue words plus reason strings) so the next attempt is not a blind redo. That only helps if the contracts are specific enough to name the failure. Vague "try again" prompts hide whether you have a model problem or an underspecified invariant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The failures showed up in the product, not just the logs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every rejected clue meant another retry before the player saw a move. On an AI spymaster turn, the game shows a clue, a count, and highlights the board words that clue is meant to cover. When the validated targets came back shorter than &lt;code&gt;count&lt;/code&gt;, the UI looked broken: &lt;code&gt;count: 2&lt;/code&gt; with only one word highlighted. The AI guesser still trusted the clue count and started reasoning from a board state that never actually existed.&lt;/p&gt;

&lt;p&gt;None of this required a different product thesis from &lt;a href="https://dev.to/michaeltruong/schema-first-prompt-second-valid-json-wasnt-enough-3nhm"&gt;schema-first validation&lt;/a&gt;. Valid JSON was never enough. The migration stress-tested whether prompt text, deterministic checks, and consumer assumptions still agreed after the model changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable part
&lt;/h2&gt;

&lt;p&gt;On paper, the clue path already looked responsible. Prompt, validator, consumer. Clean separation.&lt;/p&gt;

&lt;p&gt;The migration revealed a hidden layer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Prompt
  ↓
Model capability
  (compensating for weak contracts)
  ↓
Validator
  ↓
Consumer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I expected to compare models. Instead I ended up comparing how much of my architecture each model had been compensating for.&lt;/p&gt;

&lt;p&gt;While a more capable model kept quietly covering those weak contracts, the dashboards looked fine. Drop reasoning effort, and the same prompts start producing outputs that are honest about what you actually specified. Once that stopped happening, I was no longer measuring model quality. I was measuring how much of the gameplay experience had been resting on those hidden assumptions.&lt;/p&gt;

&lt;p&gt;That is uncomfortable and useful. Apparent regressions (count mismatches, partial batches, more retries, collapsed guess simulations) are a signal to ask which layer was doing the work: the model, or the application.&lt;/p&gt;

&lt;p&gt;Subjective "does this clue feel clever?" still matters for gameplay. It should not be the only scoreboard when the pipeline can reject an entire batch before the server ever picks a clue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat migrations as compatibility tests
&lt;/h2&gt;

&lt;p&gt;What I want out of a model swap now:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Align invariants across prompt, validator, and consumer.&lt;/strong&gt; If the prompt says "exactly &lt;code&gt;count&lt;/code&gt; targets," the validator must reject mismatches, and the API response shape must not pretend invalid intent is OK.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep structural correctness in deterministic code.&lt;/strong&gt; Use the model for association quality. Use pure functions for board membership, cardinality, illegal clue shapes, and survivor lists.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instrument validation failures by category.&lt;/strong&gt; First-pass success rate, retry rate, and failure reasons tell you whether you tightened a contract or uncovered a real model gap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evaluate end-to-end workflow metrics&lt;/strong&gt;, not only single-call latency or token price. Retries and fallbacks change the bill and the player experience; measuring only the happy path lies.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Takeaway:&lt;/strong&gt; A model migration tests the model and the architecture around it. If prompt, validator, and consumer contracts do not enforce the same invariants, stronger models can mask weaknesses in those contracts until a cheaper or more literal model exposes them. The lesson is not really about which LLM you pick. It is about architectural coupling: the model itself had become part of the contract without me noticing.&lt;/p&gt;




&lt;p&gt;If you'd like to see the project that inspired these lessons, you can try &lt;a href="https://codenames-ai.com/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=model-experiments-became-an-architectural-stress-test&amp;amp;utm_content=footer" rel="noopener noreferrer"&gt;Codenames AI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
