<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Harsh Kedia</title>
    <description>The latest articles on DEV Community by Harsh Kedia (@harshkedia17).</description>
    <link>https://dev.to/harshkedia17</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F718083%2F877341ef-4041-47f1-a7d5-d1342d3dc6b7.jpeg</url>
      <title>DEV Community: Harsh Kedia</title>
      <link>https://dev.to/harshkedia17</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/harshkedia17"/>
    <language>en</language>
    <item>
      <title>Nobody Evaluates the Evaluator</title>
      <dc:creator>Harsh Kedia</dc:creator>
      <pubDate>Mon, 03 Aug 2026 14:23:51 +0000</pubDate>
      <link>https://dev.to/harshkedia17/nobody-evaluates-the-evaluator-hh6</link>
      <guid>https://dev.to/harshkedia17/nobody-evaluates-the-evaluator-hh6</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://harshkedia.com/writing/nobody-evaluates-the-evaluator/" rel="noopener noreferrer"&gt;harshkedia.com&lt;/a&gt;. Cross-posted here in full.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An AI eval score is two claims at once: that the model did the thing, and that your grader is right about whether it did. The industry argues the first and almost never tests the second. In 2026 frontier judges have gotten genuinely good, so the unreliability moved: the score now turns on how you word the rubric and what your harness does, not on the model. Here is why that happens, with numbers, and what to do instead of trusting the number.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Earlier this year Anthropic published a number that should have landed harder than it did. A model, Opus 4.5, scored 42% on a benchmark called CORE-Bench. Then, without touching the model, they fixed some rigid grading, some ambiguous task specs, and some harness bugs, and the same model scored 95%. Fifty-three points. The model did not get smarter overnight. The ruler changed.&lt;/p&gt;

&lt;p&gt;That is the whole problem in one anecdote. We treat an eval score as a measurement of the model. It is actually the product of two things: the model's behaviour, and the instrument that judged it. And that instrument, more and more, is itself a model, an LLM-as-judge, or a benchmark harness with a grader buried inside it. Almost nobody measures the instrument.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an LLM-as-a-judge score actually claims
&lt;/h2&gt;

&lt;p&gt;Write it out and it is obvious. A passing score asserts two things: that the model produced an acceptable output, and that your grader is correct that it was acceptable. The first claim gets all the attention: leaderboards, launch posts, benchmark wars. The second gets almost none. Yet you cannot tell "the model regressed" from "the grader drifted" unless you have versioned and frozen the grader, and hardly anyone ships a grader with a version number. Your judge is a dependency that changes underneath you without a changelog.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqt62ehyz0stdegzq1zx1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqt62ehyz0stdegzq1zx1.png" alt="A model produces an answer that is scored by an instrument made of an LLM judge, a rubric, and a harness. The model path is what everyone measures; the instrument is unversioned and unvalidated, and nothing validates the validator." width="800" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Fig 1 · a score is mostly a claim about the instrument that produced it, and the instrument is the part nobody measures.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The good news: LLM judges do agree with humans
&lt;/h2&gt;

&lt;p&gt;Modern judges are good, and getting better, which ruins the easy hot take. The canonical study of LLM-as-judge found a strong model agreed with human graders more often than the humans agreed with each other.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;the agreement between GPT-4 and humans reaches 85%, which is even higher than the agreement among humans (81%).&lt;/p&gt;

&lt;p&gt;— Zheng et al., Judging LLM-as-a-Judge with MT-Bench (2023)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I wanted to see how that holds up on today's models, so I ran a small probe: ten question-and-answer pairs, graded PASS or FAIL by the three newest Gemini judges at temperature zero, all on one shared rubric. They agreed on all ten. I ran the newest of them twelve times on each item, same input every time, and it never once contradicted itself. The old fear that judges are slot machines, that even temperature zero flips verdicts, is largely obsolete on frontier models.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;three current judges, one rubric, temperature 0&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;judges: gemini-3.1-pro-preview, gemini-3.6-flash, gemini-3.5-flash

 #  verdict  item                                         all agree?
 1   FAIL    "Sydney" as Australia's capital                 yes
 2   PASS    "Canberra"                                      yes
 3   PASS    "1945" for the end of WWII                      yes
 4   FAIL    "approximately 400" for 17 x 24                 yes
 5   FAIL    "2, 3, and 9" as three primes                   yes
 6   FAIL    "yes" to is Pluto a planet                      yes
 7   PASS    "Bonjour" for good morning                      yes
 8   FAIL    the sky is blue "because it reflects the ocean" yes
 9   PASS    a passable autumn haiku                         yes
10   PASS    a vague one-line Hamlet summary                 yes

unanimous on all ten. same judge, 12 runs each at temp 0: 0 flips.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Rubric sensitivity: the unreliability moved into the criteria
&lt;/h2&gt;

&lt;p&gt;If the judges agree and never waver, the unreliability has not gone away; it has moved into the criteria. Watch what happens when I hold the same flagship judge and the same answers fixed, and change the rubric by one word, from "correct and adequately responsive" to "correct, complete, and precise":&lt;/p&gt;

&lt;p&gt;&lt;em&gt;same judge, same answers, one reworded rubric&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;judge: gemini-3.1-pro-preview,  temperature 0

 #   "responsive"   "precise"    item
10       PASS          FAIL       a vague one-line Hamlet summary

1 of 10 verdicts flipped. same model, same answer.
the score moved because the ruler did.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;It is no longer in the model, or even the judge. It moved into the criteria. And nobody versions the criteria.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;One adjective, and the flagship model reverses its verdict. That is not the model being unreliable. That is the criteria being underspecified, and the score being a function of wording nobody wrote down carefully. Researchers have a name for this. They call it criteria drift: you cannot actually pin down your evaluation criteria until you start looking at outputs, so the criteria keep moving as you judge.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;it is impossible to completely determine evaluation criteria prior to human judging of LLM outputs.&lt;/p&gt;

&lt;p&gt;— Shankar et al., Who Validates the Validators? (2024)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And notice which item flipped. Not the arithmetic, not the capital of Australia. The subjective one. That is exactly the problem, because the clear-cut cases never needed a judge in the first place. You reach for an LLM grader precisely for the fuzzy, subjective residue, and that residue is where the grader is least anchored and most sensitive to how you phrased the ask.&lt;/p&gt;

&lt;h2&gt;
  
  
  The harness moves the number more than the model
&lt;/h2&gt;

&lt;p&gt;Zoom out from one grader to a whole benchmark and it gets worse, because a benchmark is not just a grader. It is a harness: prompts, tools, context limits, retries, the environment. Any one of those can swing the score by as much as a whole model generation would.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;a composite of models, harnesses, contexts, environments, and feedback signals, any one of which can move the benchmark score by margins comparable to those between adjacent model generations.&lt;/p&gt;

&lt;p&gt;— Coding Benchmarks Are Misaligned with Agentic Software Engineering (2026)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The CORE-Bench jump from 42 to 95 was exactly this: grading, specs, and harness, not the model. And it is not only agentic benchmarks. An audit of Chatbot Arena found two identical checkpoints, the same weights submitted twice, landing 17 rating points apart, with four other models ranked in between the two copies of the same thing. The measurement had more spread than the models it was ranking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Even inter-rater agreement lies
&lt;/h2&gt;

&lt;p&gt;Suppose you do the responsible thing and try to validate your grader by measuring how often it agrees with a human. On real production traffic almost everything passes, and on lopsided data raw agreement is a liar:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;"our judge agrees with humans 90% of the time"&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;conversations        1000
failure base rate    6.5%
raw agreement        90.0%   &amp;lt;- the number in the pitch deck
Cohen's kappa         0.11   &amp;lt;- agreement above chance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ninety percent agreement, and a grader barely better than a rubber stamp that says PASS to everything. Cohen's kappa corrects for how often you would agree purely by chance, and on pass-heavy data chance agreement is enormous, so the corrected number collapses. Clinicians named this the kappa paradox back in 1990: high agreement, near-zero reliability. Which means even the teams diligent enough to validate their validator usually do it with a statistic that flatters them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reward hacking: it gets worse as the models get better
&lt;/h2&gt;

&lt;p&gt;One more twist, and it is the one that should worry anyone betting on models simply outgrowing this. Optimization pressure always flows into the gap between your metric and what you actually wanted. That is just Goodhart's law, and stronger models are better at finding the gap. Cursor reported that the smarter the model, the more it games the eval:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;reward hacking is far more common with newer, more sophisticated models than with older ones.&lt;/p&gt;

&lt;p&gt;— Cursor, Reward hacking is swamping model intelligence gains (2026)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So a stronger model does not make your unvalidated eval more trustworthy. It makes it less. The thing being measured is actively adversarial to the ruler, and it is improving faster than the ruler is.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to validate an LLM judge: this is not an argument for despair
&lt;/h2&gt;

&lt;p&gt;It would be easy to read all this as "scores are meaningless," and that is wrong, and I want to be careful because it is the lazy conclusion. Judges can match human agreement, as MT-Bench showed, once you validate and bias-correct them. Fixing a vague grader is often cheap: a handful of labelled examples and a couple of iterations. The claim is not that measurement is impossible. It is that the instrument is unvalidated by default, and validating it is both tractable and neglected.&lt;/p&gt;

&lt;p&gt;The moves are unglamorous and mostly borrowed from measurement disciplines older than machine learning. Version and freeze the grader, so that a score change means the model moved and not the ruler. Report agreement as kappa, not raw percent, so lopsided data cannot flatter you. Decompose a holistic "quality: 4 out of 5" into atomic yes-or-no checks, because one blurry score hides broken outputs and binary checks expose them. And wherever an outcome is checkable against the real world, check it instead of judging it: the most reliable eval is the one that needs no opinion. Agent benchmarks like tau-bench score the final database state rather than the prose, and are immune to grader noise for exactly that reason. Reserve the fallible judge for the genuinely subjective residue, and then treat that judge as the fallible instrument it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who evaluates the evaluator?
&lt;/h2&gt;

&lt;p&gt;There is a wall at the end of this, and honesty means walking up to it. If you validate your grader with another grader, what validates that one? The regress bottoms out only at something expensive and unscalable: careful human judgment, or a real-world outcome you can actually observe. You cannot automate your way to the bottom. Which means there is no product, no framework, no clever meta-judge that makes this go away. There is only the discipline of treating the thing that produces your scores as a measured, versioned, noise-quantified instrument, instead of an oracle.&lt;/p&gt;

&lt;p&gt;The industry spent three years getting very good at building the thing being measured, and almost no time getting good at the measurement. A passing score is a claim about your model and a claim about your grader. We have been arguing the first at the top of our lungs and taking the second on faith.&lt;/p&gt;

&lt;p&gt;Full disclosure: I am building in evals, so read this as a practitioner's obsession, not a detached take.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Anthropic, &lt;a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents" rel="noopener noreferrer"&gt;Demystifying Evals for AI Agents&lt;/a&gt; (2026). The CORE-Bench 42 to 95 story, and how much of a score is grading and harness.&lt;/li&gt;
&lt;li&gt;Zheng et al., &lt;a href="https://arxiv.org/abs/2306.05685" rel="noopener noreferrer"&gt;Judging LLM-as-a-Judge with MT-Bench&lt;/a&gt; (2023). Where the 85% vs 81% human-agreement result comes from.&lt;/li&gt;
&lt;li&gt;Shankar et al., &lt;a href="https://arxiv.org/abs/2404.12272" rel="noopener noreferrer"&gt;Who Validates the Validators?&lt;/a&gt; (2024). Criteria drift, and why you cannot fully specify criteria up front.&lt;/li&gt;
&lt;li&gt;Singh et al., &lt;a href="https://arxiv.org/abs/2504.20879" rel="noopener noreferrer"&gt;The Leaderboard Illusion&lt;/a&gt; (2025). Identical checkpoints, different Arena scores.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2606.17799" rel="noopener noreferrer"&gt;Coding Benchmarks Are Misaligned with Agentic Software Engineering&lt;/a&gt; (2026). Why harness variance rivals a model generation.&lt;/li&gt;
&lt;li&gt;Cursor, &lt;a href="https://cursor.com/blog/reward-hacking-coding-benchmarks" rel="noopener noreferrer"&gt;Reward hacking is swamping model intelligence gains&lt;/a&gt; (2026).&lt;/li&gt;
&lt;li&gt;Feinstein and Cicchetti, &lt;a href="https://pubmed.ncbi.nlm.nih.gov/2348207/" rel="noopener noreferrer"&gt;High agreement but low kappa&lt;/a&gt; (1990). The kappa paradox, from clinical statistics.&lt;/li&gt;
&lt;li&gt;Yao et al., &lt;a href="https://arxiv.org/abs/2406.12045" rel="noopener noreferrer"&gt;tau-bench&lt;/a&gt; (2024). Grading the end state instead of judging the prose.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;I write about AI/LLM systems, agent infrastructure, and the backends beneath them at &lt;strong&gt;&lt;a href="https://harshkedia.com" rel="noopener noreferrer"&gt;harshkedia.com&lt;/a&gt;&lt;/strong&gt;. Code for most of this is on &lt;a href="https://github.com/harshkedia177" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>The Two-Tower Model, Minus the Training</title>
      <dc:creator>Harsh Kedia</dc:creator>
      <pubDate>Mon, 03 Aug 2026 14:23:46 +0000</pubDate>
      <link>https://dev.to/harshkedia17/the-two-tower-model-minus-the-training-469p</link>
      <guid>https://dev.to/harshkedia17/the-two-tower-model-minus-the-training-469p</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://harshkedia.com/writing/two-tower-minus-the-training/" rel="noopener noreferrer"&gt;harshkedia.com&lt;/a&gt;. Cross-posted here in full.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A two-tower (dual-encoder) recommendation system uses a separate user encoder and item encoder that map into one shared vector space, so item embeddings can be precomputed and served from an approximate-nearest-neighbour index. The usual move is to train both towers on your click log. We did not. We built our discovery feed with frozen, pretrained multimodal embeddings and an LLM standing in for the user tower, and trained nothing for retrieval. This is why that works, what it costs, and when you should train the towers instead.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Shoppin was a &lt;a href="https://harshkedia.com/work/shoppin-discovery-feed" rel="noopener noreferrer"&gt;fashion discovery app&lt;/a&gt;: a vertical feed of videos, images, and products that had to feel personal from the first scroll. The catalog was 100,000+ videos and a far larger pool of product and inspiration images, changing daily as we scraped and generated more. This is the problem every discovery feed has on day one: you need a per-user ranking over a corpus too large to score item by item, and at the start you have almost no behavioural data to learn it from. The textbook answer to the first half is a two-tower model; the answer to the second half is to train it on your click log. We had the first problem and not the click log, and shipped anyway. Here is how, and why skipping the training was the right call.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a two-tower model actually is
&lt;/h2&gt;

&lt;p&gt;A two-tower model (also called a dual encoder, bi-encoder, or factorized model) is two neural networks that never touch until the very end. One tower, the user or query tower, reads everything you know about the user (history, context, intent) and produces a single vector. The other tower, the item tower, reads an item's features and produces a vector in the same space. Relevance between a user and an item is just the dot product of their two vectors. That is the whole scoring function: one multiply-and-sum in a shared embedding space.&lt;/p&gt;

&lt;p&gt;The reason it is drawn as two separate towers, rather than one network that eats a user and an item together, is the entire point of the design. Because the item vector depends only on item features and never on the user, you can compute every item embedding once, offline, and drop them into an approximate-nearest-neighbour (ANN) index. At request time you only run the user tower (one forward pass), get one vector, and ask the index for its top-k nearest items. Cost per request is one small model call plus one sublinear index lookup, no matter whether the corpus is 100,000 items or a billion. A model that fused user and item together would be more expressive, but it would force a forward pass per candidate, which is a non-starter for retrieval.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The item vector never depends on the user. That one constraint is what lets you precompute the whole catalog and serve it from an index.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retrieval is not ranking&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A production recommender is a funnel: retrieval narrows millions of items down to hundreds, then a heavier ranking model orders those hundreds. Two-tower lives in the retrieval stage, where the objective is recall (did the good stuff make the shortlist) and speed, not perfectly calibrated scores. Ranking is where the expensive feature-crossing models (DLRM, DCN, wide-and-deep) earn their keep, because by then the candidate set is small enough to afford them. Confusing the two stages is the most common way people over-engineer retrieval.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  How you would normally train a two-tower model
&lt;/h2&gt;

&lt;p&gt;The canonical way to train the two towers is to treat retrieval as an enormous classification problem: given a user, predict the one item they engaged with out of the entire catalog. You cannot compute a softmax over millions of items per step, so you approximate it with in-batch negatives. For each (user, item-they-liked) pair in a mini-batch, every other item in the same batch is treated as a negative example. It is cheap because those item embeddings are already computed for the batch, and it gives you a dense contrastive signal for free.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;in-batch softmax, with the correction everyone forgets&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;L(i) = -log( exp(s(x_i, y_i)/T) / sum_j exp(s(x_i, y_j)/T) )

  s(x, y) = &amp;lt;u(x), v(y)&amp;gt;     # dot product in the shared space
  T                          # temperature, scales the logits

# logQ correction: popular items appear as in-batch negatives too
# often, so subtract their sampling log-probability before softmax:
  s'(x, y) = s(x, y) - log Q(y)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The subtle part is that in-batch negatives are drawn from your traffic, which is a power law, so popular items get sampled as negatives constantly and the model over-penalises them. The fix, from Google's 2019 paper (the one people mean by "the two-tower paper"), is a logQ correction: estimate each item's sampling frequency and subtract its log-probability from the score before the softmax. Then you tune temperature, add uniform negatives so the long tail is reachable, mine a few hard negatives, and keep training and serving features identical so the model does not rot on train-serve skew. It works, and it is a real project.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;We did not have a click log worth training on. We had a cold start and a deadline.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;None of that pays off if you do not yet have the labels. Early on, our engagement data was thin, biased toward whatever we happened to be showing, and changing weekly as the product changed. Training a user tower on it would have baked in that bias and given us a model to babysit (retraining, drift, skew, index-version syncing) at exactly the moment we most needed to iterate on the product. So we made a different bet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Our version: freeze both towers
&lt;/h2&gt;

&lt;p&gt;Keep the two-tower architecture (decouple, precompute, ANN) and throw out the training. The item tower becomes a pretrained embedding model we never fine-tune. The user tower becomes something stranger: a large language model that reasons about the user and emits queries, which we then embed with the same encoder. Two encoders, a shared space, a dot product, an index. Dual-encoder-shaped in every way that matters for serving, with zero models trained for retrieval.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpnzqqg5kqa81dw5tccuk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpnzqqg5kqa81dw5tccuk.png" alt="The item tower embeds the catalog offline into a vector database; the user tower, a language model, synthesises and embeds queries online, searches the index, and the results are blended and reranked into a feed." width="800" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Fig 1 · both towers frozen. The item tower runs offline into the index; the LLM user tower runs once per request.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The item tower is a pretrained embedding pipeline
&lt;/h2&gt;

&lt;p&gt;Every inspiration image and product image runs through a pretrained multimodal embedding model that maps a picture straight into a shared vector space. Using a multimodal encoder (we used Gemini's multimodal embeddings) is the load-bearing choice here: because the same model embeds images and text into the same space, a phrase like "cropped leather jacket, muted tones" can be compared directly against a photo with a plain dot product, no separate text and image indexes to reconcile. A distributed pipeline does the unglamorous part: walk the catalog, embed in batches, upsert the vectors in bulk, and run a reconciliation job that keeps the vector store honest against the source of truth. Images and products follow the same path into different collections, and the retrieval mechanics never change. Video is where it gets interesting, because a video is not a picture.&lt;/p&gt;

&lt;p&gt;The embedding is not the whole story for an item. Alongside each vector we keep a handful of scalar labels produced by separate classifiers and by Gemini: an aesthetic score, a gender tag, and a set of quality signals (watermarking, composition, safety, an AI-generation likelihood). None of these are learned by a retrieval model. They ride along with the item and get used later, at filter and rerank time, to keep the feed clean and on brand. There is no learned clustering or taxonomy of items; the structure comes entirely from the embedding space plus these labels.&lt;/p&gt;

&lt;p&gt;Freezing has one quietly large payoff: there is no train-serve skew, the classic production-ML bug where your offline training features drift from what you compute at request time, because there is no training and the same model produces and serves every embedding. What you trade for it is staleness, embeddings aging as the catalog grows and the encoder version moves, which the reconciliation job and periodic re-embedding handle. Staleness is a far easier thing to reason about than skew.&lt;/p&gt;

&lt;h2&gt;
  
  
  A video is not one frame
&lt;/h2&gt;

&lt;p&gt;You cannot embed a fifteen-second clip by embedding its thumbnail and calling it done. A thumbnail is marketing: it is picked to make you tap, and it routinely oversells or plain misrepresents what the clip is actually about. But you also cannot embed every frame, because that is hundreds of near-identical vectors per clip, mostly redundant, and it bloats both the index and the bill. So we sampled the timeline. For each video we took the thumbnail plus the frames at roughly the 30th, 50th, and 90th percentiles of its duration, embedded each of those with the same multimodal model, and pooled them into a single vector. The thumbnail says how the clip presents itself; the percentile frames say how it actually unfolds, including the payoff near the end that a cover frame never shows. One clip becomes one point in the space, and that point represents the whole thing instead of its most clickable instant. Keeping it one point per item, rather than indexing all four frames, is deliberate: the nearest-neighbour math stays a plain top-k, and no clip can out-rank the field just by occupying four slots.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a vector database and ANN search earn their place
&lt;/h2&gt;

&lt;p&gt;Once every item is a vector, you have to decide where those vectors live and how you search them. The naive option is to keep them in your primary database and compute cosine similarity in a query, and for a few thousand rows that is genuinely fine. At a hundred thousand videos and a far larger pool of images it falls over, because exact nearest-neighbour search is linear in the corpus: every request reads every vector. That is a full scan on the hot path of your feed, on every scroll, for every user. A vector database exists to make that sublinear. It builds an approximate-nearest-neighbour index so a top-k lookup touches a small fraction of the corpus, it lets you attach scalar fields (gender, quality flags, freshness) and filter on them inside the same search, and it owns the operational work (sharding, replication, streaming upserts as the catalog changes) that you do not want to hand-roll. We used Zilliz, the managed form of Milvus, precisely so the index and the ops were somebody else's problem and our attention could go to retrieval and ranking.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;At a hundred thousand items, exact nearest-neighbour search reads every vector on every request. A vector database is how you stop paying for a full scan on every scroll.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Approximate is a feature, not a compromise&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The 'approximate' in ANN is doing real work. You trade a sliver of recall for orders of magnitude of speed, so the top-k an index returns is almost, but not exactly, the true top-k. For retrieval that is a great trade, because a ranking stage is about to reorder the shortlist anyway and a near-miss at position 40 costs you nothing. It would be a terrible trade if you treated these scores as final, calibrated relevance, which is one more reason retrieval scores and ranking scores must never be confused.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The user tower is a language model
&lt;/h2&gt;

&lt;p&gt;This is the part I like most. Instead of a learned user encoder, the user tower is a two-stage Gemini pipeline. Stage one reads the user's visual context: their own photos, up to five recent virtual-try-on results, and up to five images they have been browsing, and summarises appearance, colour and silhouette preferences, and a few style vibes. Stage two takes that summary plus engagement signals (which past queries they actually interacted with, weighted real searches versus onboarding-derived vibes, Pinterest board names) and produces exactly ten short search queries.&lt;/p&gt;

&lt;p&gt;Those ten queries are not random. The prompt mandates the split explicitly: seven that extend the user's current taste, and three from adjacent aesthetics they have not explored yet. Each query is then embedded with the same 1536-dimensional encoder the items use, and the resulting vectors are mean-pooled and L2-normalised into a single query vector. That vector goes into the ANN index exactly as a trained user tower's output would. The difference is that ours is a sentence first and a vector second.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The user tower does not output an embedding. It outputs sentences, and we embed those.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Framing the user tower as an LLM buys three things a trained encoder would have made hard. Cold start mostly disappears: the model can reason from a single profile photo and a few browse images, so a brand-new user still gets a coherent feed. The feed is explainable, because the intermediate representation is human-readable queries you can log and eyeball, not an opaque 1536-dim point. And the explore-versus-exploit balance is a line in a prompt (seven taste, three adjacent) a product person can reason about, not a hyperparameter you move by retraining. Demographics ride along the same way: age and body type as a natural-language suffix on the query, ethnicity as context in the prompt, never a hard filter on the vectors.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieval, then ranking: a stack of heuristics
&lt;/h2&gt;

&lt;p&gt;Serving is an ANN top-k over cosine similarity. Two details matter. First, the set of items the user has already seen is excluded inside the vector search's own filter, at retrieval time, not after, so you never waste the top-k on things you are about to drop. Second, you can lean on the database's range search to grab a band of results that are similar but not near-identical, which is how you get variety without the feed collapsing into ten shots of the same jacket.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;similarity bands, roughly&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;near-duplicate   cosine &amp;gt;= 0.90    drop, it is the same look
core match       cosine &amp;gt;= 0.75    keep, rank it high
usable           cosine &amp;gt;= 0.20    keep as filler / explore
below floor      cosine &amp;lt;  0.20    discard

result dedup     two results too close -&amp;gt; drop one
range band       ~0.75 to 0.90     "similar, not a dupe"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Whatever survives retrieval is reranked by a blend, not by cosine alone. Similarity gets you relevance; the other terms get you quality and freshness. An aesthetic score pulls prettier results up, a click-through-rate percentile (computed within the batch so it is scale-free) rewards things that actually perform, and raw popularity is best handled by sampling rather than a fixed weight, so a single viral item does not pin itself to the top forever. Recency decays with a short half-life, and a per-user cursor walks through the user's queries across requests so consecutive scrolls do not keep serving the same seed.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;the blended rank (representative weights)&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;score =  0.8 * cosine
       + 0.2 * (aesthetic_score / 100)     # quality signal, 0..100
       + 0.3 * ctr_rank_percentile          # within-batch, scale-free

# popularity is sampled, not a fixed term: draw from a Beta
# posterior on each item's engagement (Thompson sampling), with a
# per-query cap so one viral query cannot flood the row. Different
# engagements weigh differently: an order counts for more than a view.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Explore vs exploit, without a bandit you have to train
&lt;/h2&gt;

&lt;p&gt;How much of the feed is personalised (drawn from the user's own query vectors) versus explore or viral content is decided per user by a single function over a composite activity signal. Query volume weighs more than raw clicks, blended with how recently the user has been active, and the result maps to a personalised share somewhere between about 20% and 80%, sitting around 60% for a typical user. A user with no queries yet gets an all-explore feed, because there is nothing to personalise on. On top of that sit a few diversity modes that decide how the explore slots split across the user's own affinity, look-alike affinities, and genuinely adjacent ones.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;the explore/exploit knob&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;activity        = weighted blend of query volume, clicks, recency
personalised    = clamp(map(activity), ~0.20, ~0.80)   # ~0.60 typical
no queries yet  -&amp;gt;  100% explore       # nothing to personalise on

diversity modes (explore slots: self / look-alike / adjacent):
  self-heavy    3 / 1 / 1
  balanced      2 / 2 / 1
  explore       1 / 2 / 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A trained system would learn this trade-off from data, most likely as a bandit balancing exploration against exploitation. Ours is a handful of readable constants. That is objectively less clever, and it is also debuggable at 2am, tunable per market without a training run, and impossible to poison with a bad week of data. For the stage we were at, that was the correct trade.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;embedding dim (multimodal)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1536&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;videos indexed&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100K+&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;frames / video (thumb + 30/50/90%)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;queries / user (7 taste, 3 explore)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;10&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;personalised share&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20-80%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;surfaces (video, image, product)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;models trained for retrieval&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What we gave up by not training
&lt;/h2&gt;

&lt;p&gt;This is not a free lunch, and pretending otherwise is how you end up defending a system you do not understand. The biggest thing we gave up is learned personalisation inside the embedding itself. A pretrained encoder places two jackets near each other because they look alike, not because this particular user keeps engaging with that specific niche despite a middling visual similarity. A trained user tower can learn that; ours cannot, and we paper over it with the reranking heuristics and the LLM's taste summary.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where the frozen version leaks&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Retrieval quality is capped by the pretrained encoder, so if it does not understand a subculture or a fabric, no amount of reranking rescues it. Popularity bias is handled by Thompson sampling and per-query caps rather than a principled logQ correction, so it is contained, not solved. The LLM user tower adds a Gemini call to the critical path and can drift or hallucinate a query, which is why the queries are logged and gated. And the classic two-tower ceiling still applies: user and item never interact until the dot product, except in our case we do not even get the learned late interaction a trained model would give us.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  When you should actually train the towers
&lt;/h2&gt;

&lt;p&gt;The honest upgrade path is clear. Once you have a real engagement log, one that is large enough and not hopelessly biased by the current feed, training a user tower on it earns you the personalisation the frozen version structurally cannot express. That is the moment to bring back everything from the middle of this post: in-batch negatives with a logQ correction, a few mined hard negatives, careful temperature tuning, and strict train-serve feature parity. You keep the exact same serving shape (precompute item vectors, index them, run one tower at request time), so it is an upgrade to the towers, not a rewrite of the system.&lt;/p&gt;

&lt;p&gt;Decoupling the user and item encoders, so the catalog can be precomputed and served from an ANN index, is what makes web-scale retrieval possible, and you get that benefit whether the towers are trained on a billion clicks or frozen off the shelf.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Training buys you accuracy. It does not buy you the architecture. Start frozen, ship, and train the towers when the data earns it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Yi et al., 2019, &lt;a href="https://dl.acm.org/doi/10.1145/3298689.3346996" rel="noopener noreferrer"&gt;Sampling-Bias-Corrected Neural Modeling for Large Corpus Item Recommendations&lt;/a&gt; (RecSys). The two-tower paper: in-batch negatives, the logQ correction, and the streaming frequency estimator.&lt;/li&gt;
&lt;li&gt;Covington et al., 2016, &lt;a href="https://dl.acm.org/doi/10.1145/2959100.2959190" rel="noopener noreferrer"&gt;Deep Neural Networks for YouTube Recommendations&lt;/a&gt; (RecSys). Where the retrieval-then-ranking funnel and serving-via-ANN became standard.&lt;/li&gt;
&lt;li&gt;Huang et al., 2020, &lt;a href="https://arxiv.org/abs/2006.11632" rel="noopener noreferrer"&gt;Embedding-based Retrieval in Facebook Search&lt;/a&gt; (KDD). The hard-negative-mining lessons: hard negatives only cut recall by more than half; you want mostly easy plus a few hard.&lt;/li&gt;
&lt;li&gt;Yang et al., 2020, &lt;a href="https://dl.acm.org/doi/10.1145/3366424.3386195" rel="noopener noreferrer"&gt;Mixed Negative Sampling for Learning Two-Tower Neural Networks&lt;/a&gt; (WWW). Why in-batch negatives alone never reach the long tail.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.tensorflow.org/recommenders" rel="noopener noreferrer"&gt;TensorFlow Recommenders&lt;/a&gt; (&lt;code&gt;tfrs.tasks.Retrieval&lt;/code&gt;, &lt;code&gt;FactorizedTopK&lt;/code&gt;, the ScaNN serving layer) if you want the trained version in code.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;I write about AI/LLM systems, agent infrastructure, and the backends beneath them at &lt;strong&gt;&lt;a href="https://harshkedia.com" rel="noopener noreferrer"&gt;harshkedia.com&lt;/a&gt;&lt;/strong&gt;. Code for most of this is on &lt;a href="https://github.com/harshkedia177" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>embeddings</category>
      <category>vectorsearch</category>
    </item>
    <item>
      <title>The Model Thinks. The Harness Remembers.</title>
      <dc:creator>Harsh Kedia</dc:creator>
      <pubDate>Mon, 03 Aug 2026 14:09:37 +0000</pubDate>
      <link>https://dev.to/harshkedia17/the-model-thinks-the-harness-remembers-16n5</link>
      <guid>https://dev.to/harshkedia17/the-model-thinks-the-harness-remembers-16n5</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published on &lt;a href="https://harshkedia.com/writing/the-model-thinks-the-harness-remembers/" rel="noopener noreferrer"&gt;harshkedia.com&lt;/a&gt;. Cross-posted here in full.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An agentic harness is everything wrapped around a language model that turns it into an agent: the loop that lets it act, the tools it acts with, and the machinery that manages its memory so it can keep going. The model supplies the intelligence. The harness supplies the reliability. This is a from-scratch tour of what a harness is and how each of its parts keeps a brilliant, forgetful model on task long enough to finish a real job, grounded in open-source agents and the public engineering writing on the closed ones.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Give a modern language model a job that takes an afternoon. Refactor a module, run the tests, read the failures, fix them, move to the next file, keep going until the suite is green. For the first ten minutes it is magic. It reads the code, makes a sane plan, edits with taste. Then, somewhere around the twentieth step, it starts to drift. It re-reads a file it already changed. It forgets a decision it made twenty minutes ago. It declares victory while three tests are still red. By step forty it is confidently going in circles.&lt;/p&gt;

&lt;p&gt;The model did not get dumber between step ten and step forty. Something else broke. Understanding what broke, and the machinery people build to stop it from breaking, is the whole subject of this post. That machinery has a name that has quietly become one of the most important words in applied AI: the harness.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an agentic harness actually is
&lt;/h2&gt;

&lt;p&gt;A language model, on its own, does exactly one thing: you give it text, it predicts more text. It cannot read a file, run a command, or check its own work. It has no memory beyond the words currently in front of it. To turn that into an agent, something has to run it in a loop, hand it tools, feed the results back in, and decide when it is done. That something is the harness. Mitchell Hashimoto compressed the whole idea into a formula that stuck: Agent = Model + Harness.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An LLM agent runs tools in a loop to achieve a goal.&lt;/p&gt;

&lt;p&gt;— Simon Willison&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The loop is the beating heart of it. Anthropic describes the shape as gather context, take action, verify the result, repeat. The model looks at what it knows, calls a tool (read a file, run a test, search the web), the harness executes that tool and appends the result to the conversation, and the model looks again. It keeps looping until it stops asking for tools. Everything that is not the model itself is the harness: the loop, the set of tools and how they are described (Anthropic calls this the agent-computer interface, and warns you should sweat it as much as you would a human UI), the way context is managed, the permission checks, and the recovery when something fails. A useful 2026 shorthand pins it to four necessary parts: an agent loop, a tool interface, context management, and control. Everything in this post is one of those four.&lt;/p&gt;

&lt;p&gt;Martin Fowler's team offers a clean way to split the harness in two. There are guides, which shape the agent before it acts (the system prompt, an AGENTS.md or CLAUDE.md, the tool descriptions), and there are sensors, which give it feedback after it acts (a linter, a test result, a type error). The best sensors do not just report a failure, they phrase it so the model knows how to fix it. A good harness maximizes the chance the agent gets it right the first time, and then catches what it got wrong before a human ever sees it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq37zouxum84yxmazoqoc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq37zouxum84yxmazoqoc.png" alt="The agent loop: the model sits at the center; the harness is the ring around it that gathers context, executes the tool the model asks for, verifies the result, and feeds it back, repeating until the model stops calling tools." width="800" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Fig 1 · the model predicts; the harness does everything else. Gather context, act, verify, repeat, until the model stops asking for tools.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the naive agent loop falls apart
&lt;/h2&gt;

&lt;p&gt;Here is the naive harness in full: a while loop that calls the model, runs whatever tool it asked for, appends the output to the conversation, and calls the model again. That is genuinely all you need for short tasks, and it is close to what real agents run. The problem is the word 'appends.' Every step makes the conversation longer, and the conversation is the model's entire working memory. On a long task, that memory does not just fill up. It rots.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;the naive harness, in full&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_task&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;reply&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;            &lt;span class="c1"&gt;# predict the next action
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;           &lt;span class="c1"&gt;# no tool call: it thinks it is done
&lt;/span&gt;        &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;reply&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;run_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;        &lt;span class="c1"&gt;# read a file, run a test, search
&lt;/span&gt;        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;        &lt;span class="c1"&gt;# the context only ever grows
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rot is measurable. Chroma tested eighteen frontier models and found performance degrades as input grows, often well before the advertised context limit, with the steepest drop between roughly 100,000 and 500,000 tokens. The classic 'lost in the middle' result found accuracy falling by thirty points or more when the relevant fact sat in the middle of the context instead of the ends. Practitioners talk about a 'dumb zone' in the middle 40 to 60 percent of the window where recall quietly falls apart. A transcript stuffed with old tool output is not neutral filler. It is active interference.&lt;/p&gt;

&lt;p&gt;Then there is arithmetic that no model upgrade escapes. If each step of a task succeeds ninety-five percent of the time, and the steps depend on each other, a ten-step task finishes about sixty percent of the time and a twenty-step task about a third of the time. At a more realistic ninety percent per step, twenty steps drops to around twelve percent. METR's measurements put numbers on the ceiling: today's models finish almost every task that takes a human a few minutes, and almost none that take a human more than a few hours. Long-horizon reliability is not a knowledge problem. It is a structural one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A long task does not fail because the model got dumber. It fails because its working memory filled with noise and its small mistakes compounded.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is even a psychological wrinkle. Cognition found that models exhibit what they call context anxiety: as the window fills, the model starts rushing, cutting corners, and announcing it is running out of room even when it is not. Their fix is almost funny. They give the model a million-token window but cap real usage far below it, so it believes it has plenty of runway and stops panicking. Everything the rest of a harness does is, in one way or another, a response to these four problems: the window fills, the middle rots, errors compound, and the model panics. So let us build the fixes, one move at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move 1: keep the plan on disk, not in the context window
&lt;/h2&gt;

&lt;p&gt;If the model cannot trust its own memory over a long task, put the memory outside its head. The simplest version is a todo list the agent writes and then re-reads. This is why almost every serious coding agent has a todo tool. Claude Code shipped TodoWrite and has since moved to a structured Task tool set; opencode, which is fully open source, has a todowrite tool that persists straight into a local SQLite database. The list does two jobs: it forces the model to commit to a plan up front, and, because the agent is prompted to consult it constantly, it keeps the goal in fresh attention instead of letting it drift to the rotting middle of the context.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;move 1: keep the plan on disk, not in the model's memory&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;write_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;TODO.md&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;plan&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# commit to a plan up front
&lt;/span&gt;
&lt;span class="c1"&gt;# many steps later, feed the plan back so it cannot drift out of attention
&lt;/span&gt;&lt;span class="n"&gt;todo&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;read_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;TODO.md&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Here is your plan. Update it, then continue:&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;todo&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The bigger version of the same idea is to treat the filesystem as memory. The agent writes notes, progress logs, and decisions to files, then reads them back later. Anthropic's write-up on long-running agents describes a two-agent pattern for work that spans many sessions: an initializer sets up the environment and a progress log, and every later session reads the log, makes incremental progress, and commits its work to git with a descriptive message before it runs out of room. They even found the model is less likely to corrupt a JSON state file than a Markdown one, so the durable memory is JSON. The shift-handoff is the mental model: each session arrives with no memory of the last, so the last one has to leave good notes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move 2: context compaction, or forgetting on purpose
&lt;/h2&gt;

&lt;p&gt;Notes help, but the live conversation still grows. So the harness learns to forget deliberately. This is compaction: when the context approaches its limit, summarize the older part of the conversation, keep the most recent turns verbatim, and continue from the summary. opencode's implementation is a clean thing to read because it is open. It keeps a buffer of headroom, keeps the last several thousand tokens of conversation exactly as they are, serializes the rest into a summary produced by the same model, and starts the next turn from summary-plus-recent. Claude Code does this automatically too; Anthropic documents that it 'compacts conversation history when you approach context limits, preserving important code and decisions while freeing space.'&lt;/p&gt;

&lt;p&gt;&lt;em&gt;move 2: compact when the window fills&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;maybe_compact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;keep_recent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8000&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;count_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;                        &lt;span class="c1"&gt;# still room; do nothing
&lt;/span&gt;    &lt;span class="n"&gt;recent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;tail_by_tokens&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;keep_recent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# keep the last turns verbatim
&lt;/span&gt;    &lt;span class="n"&gt;older&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;recent&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
    &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;summarize_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;older&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;   &lt;span class="c1"&gt;# let the model compress the rest
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;system&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;recent&lt;/span&gt;          &lt;span class="c1"&gt;# continue from summary + recent
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compaction is powerful and slightly dangerous, and honesty requires saying so. A summary is lossy by definition, and the art is tuning it to preserve the load-bearing details (unresolved bugs, architectural decisions, why a thing was done) while dropping the redundant tool output. Get it wrong and you summarize away the one fact that mattered. Simon Willison has a cautionary tale of an agent that lost its original instruction during a compaction and deleted the very inbox it had been told only to review. Forgetting on purpose is necessary. Forgetting the wrong thing is how long agents go quietly, catastrophically wrong.&lt;/p&gt;

&lt;p&gt;That danger is why compaction is not one switch you flip, and why on high-stakes work you have to decide what is even allowed to be forgotten. Picture an agent drafting a contract over a hundred turns. If it compacts the client's hard requirement, 'indemnity capped at fees paid,' down to a breezy 'discussed liability terms,' it will draft something wrong and sound completely confident doing it. The fix is to make the lossy part the disposable reasoning and the lossless part the facts that must not change, by keeping the critical facts out of the summarizer entirely.&lt;/p&gt;

&lt;p&gt;So the load-bearing facts get extracted, verbatim and with their source, into a durable structured store (a matter-state: parties, governing law, must-have clauses, forbidden terms, deadlines) that compaction never touches and the agent re-reads before it drafts. Three rules keep it safe. Extract the critical facts to durable memory before you compact, not after. Never let the model draft a critical clause from its own recollection; make it re-read the source (retrieve, do not recall). And keep the full transcript on disk, so 'compacted' means trimmed-from-the-window, not deleted, and a person or a fresh agent can always audit what was dropped. For the work that truly matters, that audit is a human, at a checkpoint. Compaction manages the model's short-term memory. It is not a place to keep your source of truth.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;safe compaction: never summarize the facts that must not change&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# critical facts never go through the summarizer
&lt;/span&gt;&lt;span class="n"&gt;facts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;extract_facts&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;               &lt;span class="c1"&gt;# verbatim, with their source
&lt;/span&gt;&lt;span class="nf"&gt;save_json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;MATTER.json&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;merge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;MATTER.json&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;facts&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;maybe_compact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="c1"&gt;# only the chatter gets summarized
&lt;/span&gt;
&lt;span class="c1"&gt;# before drafting anything critical, re-read the source. do not trust memory.
&lt;/span&gt;&lt;span class="n"&gt;clause&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;draft&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;read_json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;MATTER.json&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;source_docs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;topic&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Move 3: subagents as a context firewall
&lt;/h2&gt;

&lt;p&gt;Some work is inherently context-hungry. 'Figure out how auth works across this codebase' might mean reading forty files to produce three sentences of conclusion. If the main agent does that itself, its context is now full of forty files it will never need again. The fix is the most elegant move in the whole harness: spawn a subagent. The subagent runs in its own separate context window, does the messy reading, and returns only a compressed summary to the main thread. The exploration happened somewhere else. The main thread stays clean. Anthropic notes their subagents often burn tens of thousands of tokens exploring and return one or two thousand. It is a context firewall.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn9m7q5rg15krwjrz9bgc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn9m7q5rg15krwjrz9bgc.png" alt="A context firewall: the main agent stays lean while a subagent runs in an isolated context window, expands to read many files, and returns only a small compressed summary to the main thread." width="800" height="384"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Fig 2 · the subagent as context firewall. It burns a big context exploring, and hands back a small summary. The main thread never sees the mess.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;move 3: a subagent as a context firewall&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;spawn_subagent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# a fresh context window the main thread never sees
&lt;/span&gt;    &lt;span class="n"&gt;sub&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;subagent_prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
    &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;run_agent_loop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sub&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;             &lt;span class="c1"&gt;# burns its own context exploring
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;summarize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# hand back only a small summary
&lt;/span&gt;
&lt;span class="c1"&gt;# the main thread stays lean
&lt;/span&gt;&lt;span class="n"&gt;findings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;spawn_subagent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;how does auth work across this codebase?&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;findings: &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;findings&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is also where multi-agent hype meets its limits, and the honest answer is nuanced. Reading parallelizes beautifully: fan out ten subagents to research ten things at once, each with its own context, and merge the summaries. Anthropic reported a multi-agent research system beating a single agent by about ninety percent on their internal eval, at roughly fifteen times the token cost. Writing is the opposite. Cognition's blunt argument in 'Don't Build Multi-Agents' is that when two agents edit in parallel, each makes implicit decisions the other cannot see, and the pieces do not fit together. Their rule is the single-writer principle: one agent owns the writing and the full context; subagents are scouts, not co-authors. The pattern that actually works is a single main thread that dispatches read-only subagents and keeps the pen for itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move 4: verification gates, because 'done' has to prove itself
&lt;/h2&gt;

&lt;p&gt;A model will happily tell you it finished while the tests are red. It is a text predictor, and 'the task is complete' is a very probable sentence. So the harness cannot take 'done' on faith. Done is a claim, and the harness's job is to demand proof. The cheapest proof is a verifier that runs before completion is allowed: run the test suite, run the type checker, and if they fail, feed the failure back into the loop as a new problem rather than accepting the claim. A stronger version, which Anthropic uses, is a fresh-context evaluator, a separate read-only agent whose context never saw the build, so it judges the result without the builder's motivated reasoning.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;move 4: make 'done' prove itself&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;is_done&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# 'done' is a claim. demand proof before believing it.
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;run_tests&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;passed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;      &lt;span class="c1"&gt;# a deterministic sensor
&lt;/span&gt;        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;tests are red&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
    &lt;span class="n"&gt;check&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;fresh_agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;diff&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;      &lt;span class="c1"&gt;# a read-only judge that never saw the build
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reason&lt;/span&gt;

&lt;span class="n"&gt;done&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;why&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;is_done&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;done&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;not done: &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;why&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;. keep going.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The other half of completion is surviving crashes, and this is where a real open-source harness earns its keep as an example. opencode writes every step to a durable SQLite event log before any side effect happens, snapshots the filesystem before each step so any change can be reverted, and, on restart, scans for tool calls that were left mid-flight and marks them failed so the loop can recover. That is the difference the durable-execution people (Temporal, Restate, Inngest, and others) keep hammering: saving a checkpoint is not the same as guaranteeing completion. A checkpoint says 'here is where you were.' Durable execution says 'this will run to the end, even across a crash.' Dex Horthy's 12-Factor Agents makes the same point from the other side: the agents that survive production are mostly ordinary, well-engineered software with the model sprinkled in at the decision points, not autonomous magic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move 5: goal recitation against acceptance criteria
&lt;/h2&gt;

&lt;p&gt;Everything so far keeps the agent alive across a long task. The last move keeps it pointed at the right target, because a long-running agent drifts. Forty steps in, it is busily optimizing some sub-problem it invented three detours ago and has quietly lost the plot. Drift has the same root cause as rot: the real goal was stated once, at the very top, and it is now buried in the rotting middle of the context while every recent turn is about a tangent. Every fix is a version of one idea, make the goal louder than the noise.&lt;/p&gt;

&lt;p&gt;Write the goal and the definition of done as durable, checkable criteria before the work starts, not as a vibe: must include clauses X and Y, must not include Z, done when the tests pass and a review agent signs off. Then re-inject them. Manus calls it recitation, re-stating the goal and the plan every few steps so they sit at the fresh edge of the context where attention actually is, not the dead middle. Check progress against those criteria at milestones, so off-track work is caught while it is one wasted step and not thirty. And watch the trajectory itself: an agent that starts re-reading the same file, repeating actions, or thrashing is showing you drift in real time, and the right response is to stop and re-ground it, or reset the context, rather than let it spiral. For anything high-stakes, the checkpoints are human, and the agent proposes at each gate instead of barreling to the end.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;move 5: recite the goal, check against the criteria&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;goal&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load_goal&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;   &lt;span class="c1"&gt;# {objective, must_include, must_not_include, done_when}
&lt;/span&gt;
&lt;span class="c1"&gt;# recite the goal every few steps so it stays at the edge of context
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Recite the goal and acceptance criteria, then continue:&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;render&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;goal&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;

&lt;span class="c1"&gt;# check against the criteria at milestones, not just at the very end
&lt;/span&gt;&lt;span class="n"&gt;unmet&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current_draft&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;goal&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;done_when&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;unmet&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;user&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Off track. Still unmet: &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;unmet&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;None of this is exotic. It is a project manager's job, written down: set the spec, track against it, review at milestones, escalate when something looks off. That is the uncomfortable lesson of long-horizon agents. The harness has to be the project manager, because the model, left alone with a big goal and a growing context, will confidently wander off and never notice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The twist: the best harness is barely there
&lt;/h2&gt;

&lt;p&gt;After all that machinery, the punchline is counterintuitive: the best harnesses are aggressively simple. The instinct to wrap a model in an elaborate graph of specialized agents almost always makes things worse, because every layer is another thing to debug and another way for context to fragment. The people who build these tools say it plainly. Boris Cherny, who created Claude Code, describes it as close to the opposite of clever.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;All the secret sauce, it's all in the model. And this is the thinnest possible wrapper over the model.&lt;/p&gt;

&lt;p&gt;— Boris Cherny, on building Claude Code&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Simplicity shows up in the tools, too. Give an agent a handful of sharp, general tools (read, edit, run a shell command, search) and it will compose them into anything. Give it a hundred bespoke ones and it degrades, spending its attention choosing between tools it half understands. And there is a deeper reason not to over-build: the harness has a shelf life. Every hand-coded workaround you add exists to paper over something the model cannot yet do reliably, and models keep getting better. Engineers who have watched this happen argue you should treat your harness as a ninety-day artifact, and remove structure as the model grows into the job. Add scaffolding for the level of capability you have, then take it away, because yesterday's scaffolding becomes tomorrow's bottleneck.&lt;/p&gt;

&lt;p&gt;That churn is visible in the tooling itself. By mid-2026 the sharper question had quietly moved from which model you run to which harness you run, and a new category appeared above them: meta-harnesses that let you swap one engine for another mid-session while keeping your context and policies the same underneath. The harness stopped being plumbing and became something you choose.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part nobody can measure
&lt;/h2&gt;

&lt;p&gt;Which raises the uncomfortable question I cannot stop thinking about. Every move in this post is a design decision. How do you know any of them actually helped? You would measure it, of course, by running the agent on a benchmark. Except the benchmark score is not a property of the model. It is a property of the whole system, harness included. Take one fixed model and run it inside nine different harnesses on the same tasks, and the score swings by more than twenty points, as much as the gap between one model generation and the next.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A coding agent is a system harness, a composite of models, harnesses, contexts, environments, and feedback signals, any one of which can move the benchmark score by margins comparable to those between adjacent model generations.&lt;/p&gt;

&lt;p&gt;— Coding Benchmarks Are Misaligned with Agentic Software Engineering (2026)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So most harness engineering is done half-blind. You change the compaction strategy, the score moves, and you genuinely cannot say whether you improved the harness or just got a lucky draw from a noisy instrument. If that sounds familiar, it should: it is the same problem as &lt;a href="https://harshkedia.com/writing/nobody-evaluates-the-evaluator/" rel="noopener noreferrer"&gt;trusting an eval score whose grader you never validated&lt;/a&gt;. The harness is the evaluator's twin, another unmeasured instrument silently deciding your outcome. I am not a neutral observer here, since measuring these systems honestly is the problem I have chosen to work on, but the point stands on its own: you cannot improve a harness you cannot measure, and almost nobody measures theirs.&lt;/p&gt;

&lt;h2&gt;
  
  
  So, what is a harness
&lt;/h2&gt;

&lt;p&gt;It is the machine that takes a brilliant, forgetful, occasionally overconfident text predictor and keeps it pointed at a goal long enough to reach it. A loop to let it act. Tools to act with. A notebook so it does not forget. Compaction so it can forget the right things. Subagents so the messy work happens somewhere else. Verification so 'done' means done. And enough restraint to keep all of it thin, because the model is climbing toward the day it needs less of your help, not more. The model supplies the thinking. The harness supplies the memory, the discipline, and the second chances.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The model thinks. The harness remembers, checks the work, and refuses to let it quit early.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Anthropic, &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Building Effective Agents&lt;/a&gt; (2024). The canonical workflows-vs-agents distinction and the agent-computer interface idea.&lt;/li&gt;
&lt;li&gt;Anthropic, &lt;a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="noopener noreferrer"&gt;Effective context engineering for AI agents&lt;/a&gt; (2025). Compaction, note-taking, and sub-agent context strategies.&lt;/li&gt;
&lt;li&gt;Anthropic, &lt;a href="https://www.anthropic.com/engineering/multi-agent-research-system" rel="noopener noreferrer"&gt;How we built our multi-agent research system&lt;/a&gt; (2025). Read parallelizes, write does not; subagents return compressed summaries.&lt;/li&gt;
&lt;li&gt;Cognition, &lt;a href="https://cognition.com/blog/dont-build-multi-agents" rel="noopener noreferrer"&gt;Don't Build Multi-Agents&lt;/a&gt; (2025). The single-writer principle, stated bluntly.&lt;/li&gt;
&lt;li&gt;Martin Fowler / Birgitta Boeckeler, &lt;a href="https://martinfowler.com/articles/harness-engineering.html" rel="noopener noreferrer"&gt;Harness Engineering&lt;/a&gt; (2026). Guides and sensors as a vocabulary for the harness.&lt;/li&gt;
&lt;li&gt;LangChain, &lt;a href="https://www.langchain.com/blog/the-anatomy-of-an-agent-harness" rel="noopener noreferrer"&gt;The Anatomy of an Agent Harness&lt;/a&gt; (2026). The four parts of a harness, laid out cleanly.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/ai-boost/awesome-harness-engineering" rel="noopener noreferrer"&gt;awesome-harness-engineering&lt;/a&gt;. A living 2026 list of harness patterns, evals, memory, and orchestration.&lt;/li&gt;
&lt;li&gt;Chroma, &lt;a href="https://www.trychroma.com/research/context-rot" rel="noopener noreferrer"&gt;Context Rot&lt;/a&gt; (2025). Eighteen models, all degrading before their limits.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/sst/opencode" rel="noopener noreferrer"&gt;opencode&lt;/a&gt;. An open-source coding-agent harness worth reading end to end.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2606.17799" rel="noopener noreferrer"&gt;Coding Benchmarks Are Misaligned with Agentic Software Engineering&lt;/a&gt; (2026). Why the same model swings twenty points across harnesses.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;I write about AI/LLM systems, agent infrastructure, and the backends beneath them at &lt;strong&gt;&lt;a href="https://harshkedia.com" rel="noopener noreferrer"&gt;harshkedia.com&lt;/a&gt;&lt;/strong&gt;. Code for most of this is on &lt;a href="https://github.com/harshkedia177" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
