<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Daniel Romitelli</title>
    <description>The latest articles on DEV Community by Daniel Romitelli (@daniel_romitelli_44e77dc6).</description>
    <link>https://dev.to/daniel_romitelli_44e77dc6</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2564609%2F45e9921e-df6d-47a9-a7b5-344290cb30a0.jpg</url>
      <title>DEV Community: Daniel Romitelli</title>
      <link>https://dev.to/daniel_romitelli_44e77dc6</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/daniel_romitelli_44e77dc6"/>
    <language>en</language>
    <item>
      <title>Messages Need a Protocol Before They Need a Chat UI</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Fri, 31 Jul 2026 20:07:15 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/messages-need-a-protocol-before-they-need-a-chat-ui-37oj</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/messages-need-a-protocol-before-they-need-a-chat-ui-37oj</guid>
      <description>&lt;p&gt;A patient taps a kiosk and asks for help. A staff member replies from another screen. The same message may appear in the app, trigger a push notification, fall back to Short Message Service (SMS), and later show as read.&lt;/p&gt;

&lt;p&gt;If each screen decides what those events mean, the thread turns into a rumor mill. One client marks the message as sent. Another treats a push attempt as delivery. A third retries after a network pause. In a clinic, that ambiguity creates audit gaps, extra staff work, and a patient expectation problem: sent must mean something actionable.&lt;/p&gt;

&lt;p&gt;The rule I wanted was simple: every message has a durable state, every transition has an owner, and urgent communication can leave the app channel when policy allows it.&lt;/p&gt;

&lt;p&gt;That rule is why I built &lt;code&gt;BidirectionalMessagingService&lt;/code&gt; as the coordination layer for kiosk messaging, rather than treating chat as a screen-level feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  The invariant comes before the interface
&lt;/h2&gt;

&lt;p&gt;The kiosk already had services for push, SMS, email, and templates. Those services deliver through channels. They do not decide the truth of the thread.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;BidirectionalMessagingService&lt;/code&gt; owns the message lifecycle: creation, delivery policy, acknowledgement handling, read receipts, fallback decisions, and teardown. Push, SMS, and email remain transports.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
patient[Patient Kiosk] --&amp;gt; service[BidirectionalMessagingService]
staff[Staff Client] --&amp;gt; service
service --&amp;gt; thread[Message Record]
service --&amp;gt; push[Push Notification]
service --&amp;gt; sms[SMS Fallback]
service --&amp;gt; email[Email Delivery]
service --&amp;gt; receipt[Read Receipt]```



The cost is ceremony. A basic widget can append text and look finished. This version asks callers to provide role, priority, permitted channels, and acknowledgement behavior. That extra shape buys one authority for changing the record.

## The row carries the user-facing truth

The message model stores identity, direction, status, timestamps, and attachments. The durable status set is intentionally small: `sent`, `delivered`, `read`, and `failed`.

Queue membership stays in memory while work is active. It never becomes a database value. A queued item is a worker condition, not a patient-visible fact. Persisting it would force every client to explain whether the item is safe, blocked, or retrying.

Delivery options are policy, rather than a single flag:

| Field | What it decides | Cost |
|---|---|---|
| `channels` | Which transports may carry the message | More combinations to test |
| `urgency` | How aggressively the system alerts users | Greater risk of alert fatigue |
| `requireDeliveryConfirmation` | Whether delivery needs acknowledgement | More bookkeeping |
| `fallbackToSMS` | Whether the message may leave the app channel | Higher external dependency surface |
| `translationLanguage` | Whether content needs language adaptation | More transformation risk |

Timeouts live in the same policy layer. In my implementation, the clock starts with the active delivery attempt. If that attempt expires and no permitted route can prove delivery, the message resolves to `failed`. Presence can influence routing, but it does not create another stored status.

## Receipts and teardown are side effects

A read action moves the message to `read`. Notifying the sender is a side effect of that transition. If the sender notification fails, the patient still read the message. Mixing those facts would make the history less reliable.

The same ownership applies to runtime resources. Realtime subscriptions and pending work need an explicit end, so the service owns cleanup instead of scattering it across screens. In the implementation, cleanup unsubscribes active channels, clears channel tracking, and clears the in-memory message queue. That prevents stale listeners from interpreting later kiosk activity.

That matters in a shared device flow. One patient can walk away, another can begin check-in, and an old subscription should have no vote in the new session.

## The review table

Once the lifecycle is explicit, each operation has one test: it changes the durable record, triggers a side effect, or both.

| Event | Owner | Durable status outcome | Side effects |
|---|---|---|---|
| Staff sends message | Messaging service | `sent` | Channel delivery attempts |
| Transport confirms delivery | Channel adapter | `delivered` | Optional confirmation notice |
| Patient opens message | Messaging service | `read` | Read receipt notification |
| Active attempt expires | Messaging service | `failed` | Optional fallback path before failure |
| Sender receives receipt notice | Receipt handler | No new message status | User interface acknowledgement |

Consumer chat can tolerate fuzzy indicators. A clinic kiosk has less room for soft meaning because staff coordinate care through the thread and patients expect a reply to reach someone. With explicit states and owners, one status record survives channel retries, fallback routing, and client differences. The interface can vary; the thread still tells one story.

---

🎧 **Listen to the audiobook** — [Spotify](https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D) · [Google Play](https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en) · [All platforms](https://www.craftedbydaniel.com/audiobook)
🎬 [Watch the visual overviews on YouTube](https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6)
📖 [Read the full 13-part series](https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>messaging</category>
      <category>healthcare</category>
      <category>typescript</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>A Cache Key Is an Equivalence Relation</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Tue, 28 Jul 2026 08:54:10 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/a-cache-key-is-an-equivalence-relation-51g5</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/a-cache-key-is-an-equivalence-relation-51g5</guid>
      <description>&lt;p&gt;A retry should not become a new creative decision.&lt;/p&gt;

&lt;p&gt;When a video job fails halfway through and runs again, the user still asked for the same scene. If a timestamp or retry count changes the lookup, the system pays for another clip. If the lookup ignores a model route or seed, it can return the wrong artifact with a perfectly valid URL.&lt;/p&gt;

&lt;p&gt;I built the video generation pipeline’s cache around that split. The hash is the system’s definition of sameness.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The request definition gets trimmed before hashing
&lt;/h2&gt;

&lt;p&gt;The compiler produces a &lt;code&gt;GenerationContract&lt;/code&gt;: prompt inputs, selected model route, generation mode, constraints, seed, and execution metadata. Only the artifact-defining fields enter the digest.&lt;/p&gt;

&lt;p&gt;The implementation lives in &lt;code&gt;lib/scene-compiler/ast-cache.ts&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="cm"&gt;/**
 * Hash the compiled AST to a 64-char hex string (SHA-256).
 * Identical contracts produce identical hashes regardless of
 * volatile per-request fields like timestamps or scene IDs.
 */&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;hashCompiledAST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;GenerationContract&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;extractHashableContent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;createHash&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;sha256&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;hex&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="cm"&gt;/**
 * Build a cache key with the `ast:` prefix for Supabase lookup.
 */&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;buildCacheKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;GenerationContract&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;`ast:&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;hashCompiledAST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;contract&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Abstract Syntax Tree (AST) hashing is the mechanism; classification is the design. The hashable content includes &lt;code&gt;sourceInputs.prompt&lt;/code&gt;, &lt;code&gt;sourceInputs.imageUrl&lt;/code&gt;, sorted &lt;code&gt;referenceImageUrls&lt;/code&gt;, &lt;code&gt;chosenModel&lt;/code&gt;, &lt;code&gt;chosenEndpoint&lt;/code&gt;, &lt;code&gt;generationMode&lt;/code&gt;, sorted &lt;code&gt;constraints&lt;/code&gt;, and &lt;code&gt;seed ?? null&lt;/code&gt;. Volatile fields such as &lt;code&gt;compiledAt&lt;/code&gt;, &lt;code&gt;sceneId&lt;/code&gt;, &lt;code&gt;idempotencyKey&lt;/code&gt;, and &lt;code&gt;retries&lt;/code&gt; stay out.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  contract[GenerationContract] --&amp;gt; split{Classify each field}
  split --&amp;gt;|artifact identity| keep["prompt, imageUrl, referenceImageUrls,&amp;lt;br/&amp;gt;chosenModel, chosenEndpoint,&amp;lt;br/&amp;gt;generationMode, constraints, seed"]
  split --&amp;gt;|run history| drop["compiledAt, sceneId,&amp;lt;br/&amp;gt;idempotencyKey, retries"]
  keep --&amp;gt; norm["Sort constraints by type, then target"]
  norm --&amp;gt; sha["SHA-256 over the serialized fields"]
  sha --&amp;gt; key["Cache key, ast prefix plus 64 hex"]
  drop --&amp;gt; excluded["Never reaches the digest"]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The cost is ongoing ownership. Every new field in the request shape needs a decision: artifact identity or run history. Ambiguity becomes either wasted generation or false reuse.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Normalization removes accidental difference
&lt;/h2&gt;

&lt;p&gt;Constraints are structured values, and array order can reflect construction path rather than meaning. I normalize them before serialization:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;sortConstraints&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;constraints&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;SceneConstraint&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;unknown&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;constraints&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt; &lt;span class="p"&gt;}))&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sort&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;typeCompare&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;localeCompare&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;typeCompare&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;typeCompare&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;localeCompare&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That rule makes two request definitions with the same constraint set hash together even if they were assembled in a different order. The tradeoff is explicit: ordering cannot carry semantic weight here. If priority later depends on position, this function has to change before the data model does.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The two bugs sit on opposite sides
&lt;/h2&gt;

&lt;p&gt;Hash too much, and every retry misses. &lt;code&gt;compiledAt&lt;/code&gt; lets the clock break reuse. &lt;code&gt;retries&lt;/code&gt; turns attempt two into a separate artifact. &lt;code&gt;idempotencyKey&lt;/code&gt; belongs to transport safety, so it should not alter creative identity.&lt;/p&gt;

&lt;p&gt;Hash too little, and stale output looks correct. Prompt-only reuse ignores &lt;code&gt;chosenModel&lt;/code&gt;, &lt;code&gt;chosenEndpoint&lt;/code&gt;, &lt;code&gt;generationMode&lt;/code&gt;, and &lt;code&gt;seed&lt;/code&gt;. That failure is worse than a miss: the pipeline receives a real video URL and continues with the wrong clip.&lt;/p&gt;

&lt;p&gt;Tests pin the promises: metadata changes such as &lt;code&gt;compiledAt&lt;/code&gt;, &lt;code&gt;sceneId&lt;/code&gt;, and &lt;code&gt;idempotencyKey&lt;/code&gt; preserve &lt;code&gt;hashCompiledAST&lt;/code&gt;, while &lt;code&gt;buildCacheKey&lt;/code&gt; must match &lt;code&gt;^ast:[a-f0-9]{64}$&lt;/code&gt;. The same file also defines provenance-aware steering hashes with a separate &lt;code&gt;steer:&lt;/code&gt; prefix, so namespacing is part of the storage contract.&lt;/p&gt;

&lt;p&gt;Retries keep a stable artifact identity, and different routes and seeds stay isolated. A content-addressable lookup is only as correct as the equality rule behind it; write that rule in the language of the generated thing, then let the hash enforce it.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>caching</category>
      <category>typescript</category>
      <category>videogeneration</category>
      <category>systemsdesign</category>
    </item>
    <item>
      <title>Workflow JSON Is Generated Code</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 22:51:11 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/workflow-json-is-generated-code-1nfk</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/workflow-json-is-generated-code-1nfk</guid>
      <description>&lt;p&gt;A screen recording can show a whole job without explaining it. Someone opens an inbox, checks a sender, copies a value into a customer record, compares it with a spreadsheet, sends a summary, and moves on.&lt;/p&gt;

&lt;p&gt;That work is visible, but automation still has to survive a harder test: can the system rebuild the job without dropping a step, wiring the wrong action, or importing something that looks right and fails later?&lt;/p&gt;

&lt;p&gt;I built the n8n side of the screen-analysis project around that problem. The generator does not treat the final file as a bag of text. It turns discovered automations into &lt;code&gt;N8NWorkflow&lt;/code&gt;, &lt;code&gt;N8NNode&lt;/code&gt;, and connection objects, then emits n8n-compatible JavaScript Object Notation (JSON). That adds code ceremony, but it catches a class of mistakes that string assembly invites.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Keep the platform vocabulary narrow
&lt;/h2&gt;

&lt;p&gt;Every emitted node type comes from &lt;code&gt;NodeType&lt;/code&gt; in &lt;code&gt;n8n_workflow_generator.py&lt;/code&gt;. The enumeration covers triggers, language model nodes, and application integrations the generator knows how to create. When the project needs another n8n node, I add it there before generation can use it.&lt;/p&gt;

&lt;p&gt;That costs editing speed. A quick one-off node cannot slip through by spelling a new identifier in a prompt. The gain is sharper failure: unsupported platform names fail in Python instead of hiding inside an importable file.&lt;/p&gt;

&lt;p&gt;The same idea applies to agent configurations in &lt;code&gt;n8n_agent_templates.py&lt;/code&gt;. &lt;code&gt;AgentTemplate&lt;/code&gt; names the available patterns; &lt;code&gt;AgentConfig&lt;/code&gt; carries the prompt, tools, integrations, trigger preferences, model choice, temperature, and iteration limit. A prompt is one field, not the container for everything else.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;__future__&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;annotations&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataclasses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;dataclass&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;field&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;enum&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Enum&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AgentTemplate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;EMAIL_TRIAGE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email_triage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;CRM_DATA_SYNC&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;crm_data_sync&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;CALENDAR_ASSISTANT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;calendar_assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;DOCUMENT_PROCESSOR&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;document_processor&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;COMMUNICATION_ROUTER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;communication_router&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;REPORT_GENERATOR&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;report_generator&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;LEAD_QUALIFIER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lead_qualifier&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;TASK_MANAGER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_manager&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;VOICE_ASSISTANT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;voice_assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;MULTI_AGENT_ORCHESTRATOR&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;multi_agent_orchestrator&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="nd"&gt;@dataclass&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AgentConfig&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;template&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;AgentTemplate&lt;/span&gt;
    &lt;span class="n"&gt;system_prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;default_factory&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;integrations&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;default_factory&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;triggers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;default_factory&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;llm_model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-2.5-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.7&lt;/span&gt;
    &lt;span class="n"&gt;max_iterations&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The schema is also a constraint. If a new automation needs a concept &lt;code&gt;AgentConfig&lt;/code&gt; cannot express, I extend the model first. That slows experiments, and it keeps the export path honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Build objects before JSON
&lt;/h2&gt;

&lt;p&gt;The generator’s structured form is the n8n graph itself: nodes plus named connections. It does not maintain a second private graph format. &lt;code&gt;N8NWorkflow&lt;/code&gt;, &lt;code&gt;N8NNode&lt;/code&gt;, and connection records are the representation between analysis results and the saved JSON file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  analysis[Discovered automation] --&amp;gt; workflow[N8NWorkflow]
  workflow --&amp;gt; nodes[N8NNode objects]
  workflow --&amp;gt; connections[Connection records]
  nodes --&amp;gt; json[n8n JSON export]
  connections --&amp;gt; json
  json --&amp;gt; importer[REST importer]
  importer --&amp;gt; status[Import status]```



That distinction matters. A detected step such as “classify this email” is mapped to concrete n8n nodes only when the generator has enough context to choose a trigger, model, integration, and connection order. Positioning is computed separately from identity, so the canvas stays readable without tying layout to node IDs.

The tradeoff is flexibility. Deterministic placement cannot match a hand-arranged canvas, and typed construction is heavier than editing a JSON file directly. For generated automations, I prefer predictable inspection over perfect visual layout.

The core object shape is simple:



```python
from __future__ import annotations

from dataclasses import dataclass, field
from typing import Any, Dict, List


@dataclass
class N8NNode:
    id: str
    name: str
    type: str
    position: List[int]
    parameters: Dict[str, Any] = field(default_factory=dict)
    credentials: Dict[str, Any] = field(default_factory=dict)
    type_version: float = 1.0

    def to_dict(self) -&amp;gt; Dict[str, Any]:
        node_dict = {
            "id": self.id,
            "name": self.name,
            "type": self.type,
            "position": self.position,
            "parameters": self.parameters,
            "typeVersion": self.type_version,
        }
        if self.credentials:
            node_dict["credentials"] = self.credentials
        return node_dict
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the part that makes the JSON feel like generated code. The object owns identity, type, parameters, credentials, version, and position before serialization happens. By the time the file exists, the important decisions have already passed through inspectable Python structures.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Treat import as deployment state
&lt;/h2&gt;

&lt;p&gt;Generation ends at a file; operation begins when that file reaches n8n through the Representational State Transfer (REST) API. In &lt;code&gt;n8n_importer.py&lt;/code&gt;, importer failures have a named exception, and import progress has explicit states.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;enum&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Enum&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;N8NError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Custom exception for n8n API errors.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ImportStatus&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;PENDING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;IMPORTING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;importing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;SUCCESS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;success&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;FAILED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;REQUIRES_CREDENTIALS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;requires_credentials&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A credential problem and a failed import need different recovery paths, so they get different labels. The deploy flow can generate automations, create supporting agent files, import them, and leave activation off while credentials are handled.&lt;/p&gt;

&lt;p&gt;That separation removes convenience. A single button that generates, imports, credentials, and activates would be faster for a demo. In production, splitting those actions makes partial failure recoverable.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The table I test against
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Question it answers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AgentTemplate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Which automation pattern is being built?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;AgentConfig&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Which prompt, tools, integrations, trigger, and model settings describe it?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;NodeType&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Which n8n identifiers may be emitted?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;N8NWorkflow&lt;/code&gt; / &lt;code&gt;N8NNode&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Which graph becomes JSON?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;ImportStatus&lt;/code&gt; / &lt;code&gt;N8NError&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;What happened when the artifact reached n8n?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This costs more than string interpolation: extra enums, dataclasses, object construction, save steps, and importer reports. It also makes schema changes explicit. I pay that cost because automations inferred from screen recordings already start with uncertainty; the export path should reduce it.&lt;/p&gt;

&lt;p&gt;Workflow JSON is code the moment it can move data, call models, and route work. Treating it as generated code is how I keep a discovered process from becoming an imported accident.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>n8n</category>
      <category>workflowautomation</category>
      <category>codegeneration</category>
      <category>screenanalysis</category>
    </item>
    <item>
      <title>Why I Kept Search Scope Inside a Single Supabase RPC</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:03:04 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/why-i-kept-search-scope-inside-a-single-supabase-rpc-212a</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/why-i-kept-search-scope-inside-a-single-supabase-rpc-212a</guid>
      <description>&lt;p&gt;A retrieval system returned the right embedding, the right similarity score, and a completely wrong answer. The chunk it surfaced shared vocabulary, naming conventions, even architectural patterns with the query — but it came from the wrong repository. Nothing crashed. Nothing timed out. The system just lied with perfect confidence, and I did not catch it for two days.&lt;/p&gt;

&lt;p&gt;That is the most dangerous failure mode in vector search: plausible neighbors from the wrong scope. The embedding math was fine. The rows were valid. The search logic was simply assembling too many decisions in too many places, and by the time the request reached PostgreSQL, the database had to guess which parts belonged together.&lt;/p&gt;

&lt;p&gt;I fixed the problem by making the RPC the single source of truth. The caller sends the query embedding, the candidate count, and the JSONB filter together. The SQL function receives those same values together. The index is built for the same embedding column the function reads. Once those pieces move as one unit, the path stops drifting.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  query[Query embedding] --&amp;gt; payload[RPC payload]
  count[match_count] --&amp;gt; payload
  filter[JSONB filter] --&amp;gt; payload
  payload --&amp;gt; rpc[Supabase RPC]
  rpc --&amp;gt; sql[search_embeddings]
  sql --&amp;gt; rows[Ranked rows]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The bug was scope drift, not similarity math
&lt;/h2&gt;

&lt;p&gt;The first mistake was treating search as a handful of knobs instead of a single request object. If the embedding is built in one place, the metadata filter is built in another, and the candidate depth is decided a layer above that, the call boundary becomes mushy. The function still runs, but nobody can point to one object and say: this is the exact search intent.&lt;/p&gt;

&lt;p&gt;That matters most when retrieval is scoped. In this codebase, the filter is not an afterthought. It can describe a repo, a file path, a language, a type, or any combination of metadata fields that belong in the same search slice. If I am looking for TypeScript files in a specific repository, I want that to be expressed as part of the same request that carries the vector. I do not want the application to infer scope from session state, hidden defaults, or a previous call.&lt;/p&gt;

&lt;p&gt;The bug showed up as plausible-but-wrong neighbors because vector similarity is happy to rank related text from the wrong place. That is what makes retrieval bugs so slippery. The results do not look random. They look close enough to distract you. A chunk from the wrong repo can still share concepts, terminology, or naming conventions with the query. If the filter is applied too late, the wrong row can look like a good answer until you inspect the metadata closely.&lt;/p&gt;

&lt;p&gt;The fix was to make the request shape boring and explicit. The caller decides the search scope. The database enforces that scope. The index serves that same scope. There is no second pass that tries to patch over a weaker request after the fact.&lt;/p&gt;

&lt;h2&gt;
  
  
  The caller sends one object
&lt;/h2&gt;

&lt;p&gt;The wrapper I use is intentionally small. It does not hide the request shape, and it does not smuggle in extra search behavior. It passes the vector, the count, and the filter straight into &lt;code&gt;search_embeddings&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;SupabaseClient&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@supabase/supabase-js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;SearchFilter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Record&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;unknown&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;SearchResult&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;document_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;chunk_text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;similarity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Record&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;unknown&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;runSearch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;SupabaseClient&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;queryEmbedding&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;[],&lt;/span&gt;
  &lt;span class="nx"&gt;matchCount&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;SearchFilter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;SearchResult&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rpc&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;SearchResult&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;search_embeddings&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;query_embedding&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;queryEmbedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;match_count&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;matchCount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I like this shape because it is hard to misunderstand. The only inputs that matter at call time are the embedding, the number of rows to return, and the structured scope filter. If I need to search within one repo, I pass a repo filter. If I need to narrow by language, I add a language key. If I need to restrict by file type or path, I add that too. The caller is not building a query plan; it is declaring intent.&lt;/p&gt;

&lt;p&gt;The same function can be called with a narrow filter or an empty one. An empty filter means search the whole corpus. A populated filter means search the subset that matches the metadata predicate. That is a very different outcome, and I want that difference visible right where the request is created.&lt;/p&gt;

&lt;p&gt;A typical call site stays just as clear:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;runSearch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;supabase&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;repo&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;the author/portfolio&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;language&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;typescript&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the point. The search boundary should be obvious at a glance. If I am debugging a bad answer later, I should be able to inspect one payload and know exactly what the database was asked to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  The database function matches the same interface
&lt;/h2&gt;

&lt;p&gt;On the PostgreSQL side, the &lt;code&gt;search_embeddings&lt;/code&gt; function accepts the same three inputs the caller sends. The metadata filter stays inside SQL, where it belongs. The rows are filtered first, then ranked by vector distance.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="n"&gt;EXTENSION&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;OR&lt;/span&gt; &lt;span class="k"&gt;REPLACE&lt;/span&gt; &lt;span class="k"&gt;FUNCTION&lt;/span&gt; &lt;span class="n"&gt;search_embeddings&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;query_embedding&lt;/span&gt; &lt;span class="n"&gt;vector&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;match_count&lt;/span&gt; &lt;span class="nb"&gt;integer&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;filter&lt;/span&gt; &lt;span class="n"&gt;jsonb&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="s1"&gt;'{}'&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;jsonb&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;RETURNS&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;document_id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;chunk_text&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;similarity&lt;/span&gt; &lt;span class="nb"&gt;double&lt;/span&gt; &lt;span class="nb"&gt;precision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;metadata&lt;/span&gt; &lt;span class="n"&gt;jsonb&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;LANGUAGE&lt;/span&gt; &lt;span class="k"&gt;sql&lt;/span&gt;
&lt;span class="k"&gt;STABLE&lt;/span&gt;
&lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="err"&gt;$$&lt;/span&gt;
  &lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;document_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chunk_text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;query_embedding&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;similarity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;metadata&lt;/span&gt;
  &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;embeddings&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;
  &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;filter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'{}'&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="n"&gt;jsonb&lt;/span&gt; &lt;span class="k"&gt;OR&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;metadata&lt;/span&gt; &lt;span class="o"&gt;@&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;filter&lt;/span&gt;
  &lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;query_embedding&lt;/span&gt;
  &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="n"&gt;match_count&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="err"&gt;$$&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="n"&gt;idx_embeddings_embedding&lt;/span&gt;
  &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;embeddings&lt;/span&gt;
  &lt;span class="k"&gt;USING&lt;/span&gt; &lt;span class="n"&gt;hnsw&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;embedding&lt;/span&gt; &lt;span class="n"&gt;vector_cosine_ops&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important line is the metadata predicate: &lt;code&gt;e.metadata @&amp;gt; filter&lt;/code&gt;. That is not a clean-up step after ranking. It is part of the search itself. Rows that do not belong to the requested scope never enter the ranked candidate set.&lt;/p&gt;

&lt;p&gt;That design matters because the database is the only place that can apply the filter consistently at the same moment it applies similarity. If the application filters after ranking, the query can still surface neighbors from the wrong scope first. If the filter is inside SQL, then ranking only happens across rows that already belong to the same metadata neighborhood as the request.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;similarity&lt;/code&gt; field is there for the caller's benefit. Internally, cosine distance still drives the ordering. Externally, I want a score where larger reads as better. Returning both the score and the raw chunk text gives the next stage enough context to render, inspect, or rerank without another round trip.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;STABLE&lt;/code&gt; marker also fits the way I use the function. For a fixed snapshot and a fixed input payload, this is a deterministic retrieval step. It is not a side-effect machine. It is a search function.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why JSONB belongs in the request, not outside it
&lt;/h2&gt;

&lt;p&gt;The filter is JSONB because the scope is structured, not free-form. A metadata filter can say more than one thing at once, and it needs to do that without collapsing into a pile of ad hoc parameters. The same object can describe a repository slice, a file path constraint, a language constraint, or a file-type constraint.&lt;/p&gt;

&lt;p&gt;That gives me a single place to express the search boundary. If I need to retrieve only TypeScript chunks from a particular repository, I do not want to assemble a special query for that case. I want to build a JSONB object like &lt;code&gt;{ repo: 'the author/portfolio', language: 'typescript' }&lt;/code&gt; and pass it straight through. The SQL predicate then enforces exactly that constraint.&lt;/p&gt;

&lt;p&gt;This is also the reason I keep the filter visible at the RPC boundary instead of burying it in a helper that silently rewrites inputs. Hidden rewrite logic is how search calls become hard to reason about. The JSONB object is simple enough to inspect, easy to log, and unambiguous in SQL.&lt;/p&gt;

&lt;p&gt;There is a second benefit that shows up during debugging. When a query returns too much, I can loosen the filter and see the effect immediately. When it returns too little, I can inspect the metadata keys that are actually present in the table. Because the filter is part of the request, there is no mystery about which layer decided the corpus was too broad or too narrow.&lt;/p&gt;

&lt;p&gt;The same applies when I expand the metadata model. If I start attaching more structure to a chunk, such as file type or path segments, I do not need to change the RPC signature. I update the JSONB shape, then let the same &lt;code&gt;search_embeddings&lt;/code&gt; function enforce the new predicate. The request boundary stays stable while the metadata vocabulary evolves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the index is part of the same story
&lt;/h2&gt;

&lt;p&gt;I do not think about the index as a separate optimization pass. It is part of the same guarantee that the RPC makes. If the function says it will search the &lt;code&gt;embeddings&lt;/code&gt; table with cosine distance, the index should be built for that exact access path.&lt;/p&gt;

&lt;p&gt;That is why the HNSW index sits beside the function in my mental model. The function defines the agreement. The index makes that agreement fast enough to use all the time. &lt;code&gt;vector_cosine_ops&lt;/code&gt; matches the ranking strategy, so the storage layer is not fighting the retrieval layer.&lt;/p&gt;

&lt;p&gt;The nice thing about HNSW here is that it matches the shape of the workload I care about: lots of dense vector searches, with a metadata filter that keeps the working set scoped before ranking. I am not trying to make the index do the job of the filter. I am making both pieces do the job they are good at. The metadata predicate narrows the rows. The vector index ranks the rows that remain.&lt;/p&gt;

&lt;p&gt;That separation is what keeps the system predictable. If I ever need to inspect performance, I know where to look. If the wrong neighbors are showing up, I inspect the filter. If the right neighbors are slow, I inspect the index and the shape of the vector column. The responsibilities do not blur together.&lt;/p&gt;

&lt;h2&gt;
  
  
  What went wrong the first time
&lt;/h2&gt;

&lt;p&gt;The first broken version of this path made the request feel more flexible than it really was. The caller knew one thing, the search function inferred another, and the database had to reconcile them later. That is exactly the sort of accidental complexity that makes retrieval bugs hard to pin down.&lt;/p&gt;

&lt;p&gt;The visible symptom was a result set that looked sane at a glance. The hidden problem was that the request did not fully describe the scope of the search. That is why the bug survived long enough to matter. Nothing crashed. Nothing timed out. The system just answered the wrong question with confidence.&lt;/p&gt;

&lt;p&gt;Once I stopped spreading that decision across layers, the failure mode disappeared. The request object became the single place where I could answer three questions at once:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What text or embedding is this search about?&lt;/li&gt;
&lt;li&gt;How deep should the candidate set be?&lt;/li&gt;
&lt;li&gt;Which metadata fields are allowed to participate?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is a much better debugging surface than a trail of local variables and implicit defaults. If I have to reason about a bad answer, I want to reason about one request object and one SQL function. That is enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I trust the boundary now
&lt;/h2&gt;

&lt;p&gt;The version I trust is the one where the search request is explicit enough that the database never has to guess. The caller passes the vector, the candidate count, and the JSONB filter in one place. The SQL function applies the filter inside the ranking query. The HNSW index is built on the same embedding column the function reads. Every step agrees on the same shape.&lt;/p&gt;

&lt;p&gt;That is the whole reason I keep search scope inside a single Supabase RPC. Not because it is fashionable, and not because it makes the code shorter, but because it keeps the search intent attached to the request that asked for it. The RPC boundary becomes the line where scope is declared and enforced.&lt;/p&gt;

&lt;p&gt;Once I made that change, retrieval stopped feeling like a chain of guesses and started feeling like a reliable interface again. That matters in a system where a correct answer is only useful if it comes from the right slice of data. In the next pass, I am going to push that same discipline further into the ingestion side, because retrieval only stays honest when the embeddings and metadata that feed it are just as deliberate.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>supabase</category>
      <category>postgres</category>
      <category>pgvector</category>
      <category>rag</category>
    </item>
    <item>
      <title>The AgentGroupChat Pattern That Keeps the Mapper from Drifting</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:02:52 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/the-agentgroupchat-pattern-that-keeps-the-mapper-from-drifting-101e</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/the-agentgroupchat-pattern-that-keeps-the-mapper-from-drifting-101e</guid>
      <description>&lt;p&gt;The first version of this orchestration failed in a very specific way: the mapper kept producing templates that looked plausible to a human and were wrong for the platform. The failure was not that the model was incapable of reasoning. The failure was that I had asked too much of a single, broad conversation, so the wrong structure got accepted early and then propagated all the way to the end. By the time the validator complained, the chain had already lost the distinction between analysis, mapping, generation, and approval.&lt;/p&gt;

&lt;p&gt;That is the point where I stopped treating the system like a prompt stack and started treating it like a state machine. In the workflow analyzer SaaS, the code in &lt;code&gt;runner/azure_foundry/src/orchestrator.py&lt;/code&gt; does the important work: it creates the kernel, registers the plugins, assembles the agents, defines who speaks next, and decides when the run is done. The state manager in &lt;code&gt;runner/azure_foundry/src/state_manager.py&lt;/code&gt; sits around that orchestration so a run can resume from the last durable message history instead of inventing a fresh conversation every time. That single design choice changed the behavior of the whole pipeline.&lt;/p&gt;

&lt;p&gt;The architecture I wanted was simple to describe and annoying to get right: each agent gets one contract, each contract produces one kind of output, and the validator is the gate that decides whether the output can move forward. The analyzer looks at the workflow analysis and identifies automation opportunity. The mapper turns that opportunity into abstract integration steps. The generator turns the mapped steps into a platform-shaped template. The validator checks the structure and either approves the run or sends it back to the generator.&lt;/p&gt;

&lt;p&gt;The part that made this work was not just splitting the roles. It was making the next speaker explicit. That is what &lt;code&gt;AgentGroupChat&lt;/code&gt; gives me here: selection strategy, termination strategy, and a history buffer that can be rehydrated before the chat starts. Once those pieces are in place, the system stops behaving like a free-form exchange and starts behaving like a controlled pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What broke first
&lt;/h2&gt;

&lt;p&gt;The first architecture looked tidy on paper and failed under repetition. I had an analysis step, a mapping step, a generation step, and a validation step, but the conversation itself was too loose. The mapper started inventing platform syntax because nothing in the orchestration made that impossible. The generator then built on top of that invented syntax, and the validator only saw the problem after the wrong shape had already been repeated several times.&lt;/p&gt;

&lt;p&gt;The deeper problem showed up on retries. If the run timed out or the validator rejected the output, the next attempt often lost the conversation history that explained why the output had failed. That meant the system retried from a weak starting point and repeated the same mistake. The loop did not need more creativity. It needed a strict memory of what had already happened and a strict rule about which agent was allowed to correct which class of mistake.&lt;/p&gt;

&lt;p&gt;That is why I moved to a narrow contract chain with validator-led convergence. The validator does not try to repair everything. It does one job: decide whether the generated template is acceptable. If the answer is no, the generator gets another pass. If the answer is yes, the run terminates. That keeps the correction local instead of letting every agent participate in every mistake.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core loop in &lt;code&gt;AgentGroupChat&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;The orchestration in &lt;code&gt;runner/azure_foundry/src/orchestrator.py&lt;/code&gt; is built around &lt;code&gt;AgentGroupChat&lt;/code&gt; with two strategies attached: a selection strategy that decides the next speaker, and a termination strategy that decides when to stop. That is the heart of the pattern.&lt;/p&gt;

&lt;p&gt;The selection strategy reads the conversation history and follows a fixed pipeline:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Analyzer speaks first.&lt;/li&gt;
&lt;li&gt;Mapper speaks second.&lt;/li&gt;
&lt;li&gt;Generator speaks third.&lt;/li&gt;
&lt;li&gt;Validator speaks fourth.&lt;/li&gt;
&lt;li&gt;If the validator returns INVALID, the next speaker is the Generator again.&lt;/li&gt;
&lt;li&gt;If the validator returns WORKFLOW_APPROVED, the chat ends.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is a very different shape from a single model answering everything inside one giant instruction block. The selection strategy makes the pipeline visible to the runtime. It also gives me a place to encode the retry rule explicitly instead of hiding it inside a paragraph of instructions.&lt;/p&gt;

&lt;p&gt;The termination strategy is just as important. The validator is the only agent whose result can end the run, and the run still has a maximum iteration cap of 10 so a bad loop cannot spin forever. That cap matters because agent loops do not usually fail in dramatic ways. They fail by wobbling. A small wobble repeated enough times turns into wasted tokens, delayed jobs, and outputs that never settle.&lt;/p&gt;

&lt;p&gt;Here is the shape of that control flow in a standalone Python example that mirrors the same idea. This version does not depend on Azure or Semantic Kernel, but it shows the exact contract I want the orchestrator to enforce:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataclasses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;dataclass&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;enum&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Enum&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Dict&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;ANALYZER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Analyzer&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
    &lt;span class="n"&gt;MAPPER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Mapper&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
    &lt;span class="n"&gt;GENERATOR&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Generator&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
    &lt;span class="n"&gt;VALIDATOR&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Validator&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;


&lt;span class="nd"&gt;@dataclass&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;rehydrate_history&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history_state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Message&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="nc"&gt;Message&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;''&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;history_state&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;next_speaker&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Message&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Optional&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ANALYZER&lt;/span&gt;

    &lt;span class="n"&gt;last&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;last&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;role&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;VALIDATOR&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;WORKFLOW_APPROVED&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;last&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;INVALID&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;last&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GENERATOR&lt;/span&gt;

    &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ANALYZER&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MAPPER&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GENERATOR&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;VALIDATOR&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;seen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;role&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;role&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ANALYZER&lt;/span&gt;

    &lt;span class="n"&gt;last_agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;AgentName&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;idx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;index&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;last_agent&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;idx&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;


&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;restored&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;rehydrate_history&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Analyzer&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Detected CRM handoff&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Analyzer&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Mapper&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Trigger + action pair identified&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Mapper&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Generator&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;triggers&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;: [], &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;actions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;: []}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Generator&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Validator&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;INVALID: missing triggers and actions&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Validator&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;next_speaker&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;restored&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That small example captures the most important behavior: the validator does not restart the chain, and the chain does not forget where it left off. If the validator rejects the output, the next speaker is the Generator, not the Analyzer, because the analysis step did its job already.&lt;/p&gt;

&lt;h2&gt;
  
  
  How history gets rehydrated before the run
&lt;/h2&gt;

&lt;p&gt;The resumability story matters because the orchestration runs inside Prompt Flow, not inside a single in-memory toy conversation. In the real flow, &lt;code&gt;state_manager.py&lt;/code&gt; loads the prior session history before the orchestrator node runs. That history is passed in as &lt;code&gt;history_state&lt;/code&gt;, and the orchestrator reconstructs &lt;code&gt;ChatMessageContent&lt;/code&gt; objects from each saved message so &lt;code&gt;AgentGroupChat&lt;/code&gt; can continue the conversation instead of starting over.&lt;/p&gt;

&lt;p&gt;That detail is easy to miss and expensive to ignore. Without rehydration, every timeout becomes a reset. With rehydration, the system can pick up from the last durable state and keep going. If the validator has already rejected a malformed template, that rejection stays in history. If the mapper already established the target platform and the apps involved, that context remains available. If the previous run ended halfway through generation, the next run does not need to relearn the same facts.&lt;/p&gt;

&lt;p&gt;The lifecycle is straightforward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;load session history from &lt;code&gt;state_manager.py&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;reconstitute the saved messages into the chat history&lt;/li&gt;
&lt;li&gt;run the orchestrator with the restored history&lt;/li&gt;
&lt;li&gt;persist the resulting execution state after the run&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is what makes the pipeline durable. The system is not pretending the previous attempt never happened. It is continuing the same conversation with the same constraints.&lt;/p&gt;

&lt;h2&gt;
  
  
  The template library is not decoration
&lt;/h2&gt;

&lt;p&gt;The template library in &lt;code&gt;runner/template_library&lt;/code&gt; gives the generator a starting point that already matches the platform family. That matters because the generator should not be inventing a whole import shape from memory. It should be filling a known skeleton with detected apps, triggers, and actions.&lt;/p&gt;

&lt;p&gt;When the generator starts from a template base, it can spend its effort on the parts that actually need reasoning: mapping the workflow into the right platform structure, inserting the right application names, and filling the right fields. The template library keeps it from drifting into malformed top-level structure.&lt;/p&gt;

&lt;p&gt;That is also why the validator can stay narrow. If the generator is working from a known shape, the validator does not need to be a general-purpose critic. It only needs to check the contract that matters for the chosen platform: do the expected keys exist, is the structure complete, and does the result satisfy the platform rules well enough to be accepted?&lt;/p&gt;

&lt;p&gt;The combination is what matters. The template library constrains the starting point. The generator fills it. The validator checks it. The selection strategy decides whether the system should move forward or send the work back to the generator.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real runtime wiring in the orchestrator
&lt;/h2&gt;

&lt;p&gt;The orchestrator itself is where the system becomes code-first instead of prompt-first. The kernel is initialized with Azure OpenAI configuration from the environment, the validation plugin is registered, and the template library plugin is registered before the conversation begins. Then the agents are assembled and the selection and termination strategies are attached.&lt;/p&gt;

&lt;p&gt;That ordering is important. The model cannot speak before the kernel has its service. The generator cannot rely on the template library unless the plugin is present. The validator cannot enforce the schema unless the validation plugin is loaded. The orchestration code makes those dependencies explicit instead of leaving them implicit in a prose prompt.&lt;/p&gt;

&lt;p&gt;The runtime wiring also keeps the Azure configuration outside the orchestration logic itself. The environment supplies the deployment name, endpoint, and API key, and the kernel uses those values when it constructs the Azure chat completion service. That keeps the orchestration focused on behavior rather than connection plumbing.&lt;/p&gt;

&lt;p&gt;In practice, the structure looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  load[Load state by session_id] --&amp;gt; rehydrate[Rebuild chat history]
  rehydrate --&amp;gt; analyzer[Analyzer]
  analyzer --&amp;gt; mapper[Mapper]
  mapper --&amp;gt; generator[Generator]
  generator --&amp;gt; validator[Validator]
  validator -- WORKFLOW_APPROVED --&amp;gt; stop[Terminate]
  validator -- INVALID --&amp;gt; generator
  validator -- other --&amp;gt; generator
  stop --&amp;gt; persist[Persist state]```



The important arrow is the one that returns from Validator to Generator. That is the local correction path. Invalid output does not reset the analysis, erase the mapping, or jump back to the start. It goes back to the one agent responsible for producing the template.

## A runnable validator that actually checks structure

The earlier draft included a validator that always returned success, which would have defeated the entire point of the system. The validator must reject malformed output or it is just another polite participant in the conversation. Here is a minimal, runnable version that checks a template shape and returns a real verdict:



```python
import json
from typing import Any, Dict, List


SUPPORTED_PLATFORMS = {'zapier', 'make', 'n8n'}


def select_template(platform: str) -&amp;gt; Dict[str, Any]:
    platform_key = platform.lower()
    if platform_key not in SUPPORTED_PLATFORMS:
        raise ValueError(f'Unsupported platform: {platform}')

    base = {
        'name': f'{platform_key}_automation',
        'triggers': [],
        'actions': [],
    }
    return json.loads(json.dumps(base))


def validate_template(template_json: str, platform: str) -&amp;gt; Dict[str, Any]:
    errors: List[str] = []

    try:
        template = json.loads(template_json)
    except json.JSONDecodeError as exc:
        return {'status': 'invalid', 'errors': [f'Invalid JSON: {exc}']}

    for key in ('name', 'triggers', 'actions'):
        if key not in template:
            errors.append(f'Missing key: {key}')

    if platform.lower() not in SUPPORTED_PLATFORMS:
        errors.append(f'Unsupported platform: {platform}')

    if not isinstance(template.get('triggers'), list) or not template.get('triggers'):
        errors.append('At least one trigger is required')

    if not isinstance(template.get('actions'), list) or not template.get('actions'):
        errors.append('At least one action is required')

    return {
        'status': 'valid' if not errors else 'invalid',
        'errors': errors,
    }


if __name__ == '__main__':
    bad = json.dumps({'name': 'demo', 'triggers': [], 'actions': []})
    good = json.dumps({'name': 'demo', 'triggers': ['new_event'], 'actions': ['create_record']})
    print(validate_template(bad, 'Zapier'))
    print(validate_template(good, 'Zapier'))
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the behavior I want the real validator path to approximate: reject missing structure, reject empty structure, and return a verdict that the selection strategy can use to decide the next speaker. The important part is not the exact shape of this standalone example. The important part is the discipline it expresses. Validation is not a suggestion. It is a gate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this pattern is easier to trust
&lt;/h2&gt;

&lt;p&gt;The reason this design works is that every failure has one owner. If analysis is weak, the Analyzer is the problem. If the mapping invents structure, the Mapper is the problem. If the generated template does not fit the platform, the Generator is the problem. If the template violates the contract, the Validator says so. That clarity makes debugging much easier because I am not trying to infer which layer went wrong from a cloud of blended instructions.&lt;/p&gt;

&lt;p&gt;It also makes the retry logic sane. When the validator rejects the output, I do not want the entire conversation to restart. I want the generator to take another pass with the same context still in memory. That is what the selection strategy enforces. The orchestrator reads the last message, sees the invalid verdict, and routes execution back to the generator. No one else needs to renegotiate the earlier steps.&lt;/p&gt;

&lt;p&gt;State rehydration matters here because the retry is not the same as a fresh run. A fresh run throws away the exact information I need most: what failed, what was already agreed, and which platform shape was already chosen. Rehydration preserves that state so the next attempt can repair the actual fault instead of repeating the whole conversation.&lt;/p&gt;

&lt;p&gt;The result is a system that can fail locally without collapsing globally. That sounds like a small thing until you watch a long run survive a rejection, recover from persisted history, and land on a valid template without re-deriving the entire workflow from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing the loop
&lt;/h2&gt;

&lt;p&gt;The strongest part of this orchestration is not that it uses multiple agents. It is that the agents are constrained by code, the next speaker is chosen by strategy, the validator owns the final gate, and the session history survives retries. Once those pieces were in place, the mapper stopped drifting and the generator stopped improvising against the wrong shape.&lt;/p&gt;

&lt;p&gt;That gave me exactly what I wanted from the workflow analyzer SaaS: a chain that can analyze, map, generate, validate, and resume without pretending a failed attempt never happened. The next thing I am interested in is pushing more of the template shape into the library itself so the generator begins with an even tighter platform skeleton and the validator has less to reject in the first place.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>semantickernel</category>
      <category>azurefoundry</category>
      <category>promptflow</category>
      <category>agentorchestration</category>
    </item>
    <item>
      <title>Validation Geometry Is Part of the Model</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:02:41 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/validation-geometry-is-part-of-the-model-nhn</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/validation-geometry-is-part-of-the-model-nhn</guid>
      <description>&lt;p&gt;The dangerous part was not the classifier. It was the label file sitting there looking convenient.&lt;/p&gt;

&lt;p&gt;A saved &lt;code&gt;train_y.npy&lt;/code&gt; artifact existed, but for this baseline it was a trap: it contained magnitude-filtered positive targets only, which made it unusable for a directional classifier. If I had trained on it anyway, the result would have looked like a model comparison while quietly being a dataset-artifact comparison.&lt;/p&gt;

&lt;p&gt;That is the reason I made the LightGBM minute baseline read Pramaana's per-asset feature parquet files directly. The point was not to tune trees until they confessed. The point was to close the model-class capacity objection in the minute-level ceiling experiment without letting label construction, overlapping windows, or temporal bleed sneak into the room wearing a lab coat.&lt;/p&gt;

&lt;p&gt;The research question behind the ICAIF 2026 paper is deliberately narrow: minute-scale cryptocurrency direction from OHLCV candles appears reproducibly capped near 52% across a broad set of model and feature configurations, while the same research stack recovers a materially larger directional signal at the hourly horizon. Across seven minute configurations and approximately 36 million rows of minute-scale OHLCV history, the observed range is 51.4% to 52.3%. The LightGBM baseline is the seventh configuration: 46 microstructure-proxy features, a 15-minute forward return target in basis points, a stride-15 de-overlap, a no-trade filter at &lt;code&gt;|target| &amp;gt; 10 BPS&lt;/code&gt;, and a per-asset 85/15 temporal split with a 15-row purge before validation.&lt;/p&gt;

&lt;p&gt;The transferable idea is simple and annoyingly easy to violate: in time-series ML, validation geometry is not bookkeeping. It is part of the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The baseline was a geometry test, not a tuning contest
&lt;/h2&gt;

&lt;p&gt;A naive capacity comparison asks, “Does a different model class beat the neural setup?” That sounds reasonable until the labels are temporal, overlapping, filtered, and asset-scoped. Then the better question is, “Can I compare model classes without changing the target semantics or leaking nearby time into validation?”&lt;/p&gt;

&lt;p&gt;That is why the script documents the protocol before it imports anything. The docstring is doing more than explaining a file; it is pinning down the shape of the experiment.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;#!/usr/bin/env python3
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Run a LightGBM directional baseline on the frozen M6 sniper matrix.

This closes the model-class capacity objection for the paper&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s minute-level
ceiling section. The script reads Pramaana&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s per-asset
``data/tmp_sniper_feat_*.parquet`` files directly because the saved
``train_y.npy`` artifact currently contains magnitude-filtered positive targets
only and is therefore not usable for a directional classifier.

Protocol matched to ``scripts/preprocess_sniper.py``:
  - 46 engineered microstructure-proxy features
  - 15-minute forward return target in BPS
  - stride-15 de-overlap
  - |target| &amp;gt; 10 BPS no-trade filter
  - per-asset 85/15 temporal split with a 15-row purge before validation
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;__future__&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;annotations&lt;/span&gt;

&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;lightgbm&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;lgb&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;polars&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pl&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;scipy&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;stats&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sklearn.metrics&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;balanced_accuracy_score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;roc_auc_score&lt;/span&gt;


&lt;span class="n"&gt;TRAIN_RATIO&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.85&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I like this kind of comment because it is falsifiable. Every important experimental choice is named: feature count, target horizon, de-overlap, no-trade filtering, split ratio, and purge width. If the number later changes, the protocol has to change with it. (The real script also exposes these knobs as &lt;code&gt;argparse&lt;/code&gt; CLI flags; I trimmed the flag-parsing boilerplate from this excerpt to keep the protocol in focus.)&lt;/p&gt;

&lt;p&gt;The saved label artifact failed the most basic requirement for this comparison: it did not represent both directions for a directional classifier. Directional classification needs the sign of the target. A target file containing magnitude-filtered positive targets only is not a neutral shortcut; it changes the task. So the baseline reconstructs the binary label from the 15-minute forward return in basis points inside the feature-parquet path, rather than inheriting a label artifact whose semantics were already wrong for this purpose.&lt;/p&gt;

&lt;p&gt;There is a useful mental model here: a time-series split is less like cutting a deck of cards and more like cutting wet paint. The boundary smears unless you leave room for it to dry. In this baseline, the 15-row purge is that dry strip between training and validation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The timeline is the experiment
&lt;/h2&gt;

&lt;p&gt;The pipeline has only a few stages, but the order matters. The baseline starts from per-asset parquet files, reconstructs the directional target from the forward return, filters out the no-trade zone, de-overlaps with stride 15, applies a per-asset temporal split, inserts a purge gap, and only then fits and validates the model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  parquet[Per asset feature parquet] --&amp;gt; labels[Reconstruct sign label]
  labels --&amp;gt; filter[No trade filter]
  filter --&amp;gt; stride[Stride 15 de overlap]
  stride --&amp;gt; split[Per asset temporal split]
  split --&amp;gt; train[Training interval]
  split --&amp;gt; purge[15 row purge gap]
  purge --&amp;gt; validate[Validation interval]
  validate --&amp;gt; metrics[Accuracy and AUC]```



The diagram looks almost too ordinary, which is exactly the trap. The ordinary-looking arrows are where most of the statistical damage would happen if they were skipped or reordered.

For each asset, the geometry is anchored by time rather than random assignment. The training interval comes first. The validation interval comes last. The purge gap sits between them. The overlapping-window hazard is caused by the target construction itself: a 15-minute forward return means nearby rows can share future information unless the split respects the horizon. The stride-15 de-overlap reduces that hazard, and the purge gap protects the split boundary.

| Geometry choice | What it protects against | Concrete value in this baseline |
|---|---|---|
| Per-asset temporal split | Cross-time contamination within each asset | 85/15 train/validation |
| Purge gap | Boundary bleed from nearby rows | 15 rows before validation |
| De-overlap | Repeated labels from overlapping horizons | stride-15 |
| No-trade filter | Tiny targets treated as tradable direction | `|target| &amp;gt; 10 BPS` |
| Label reconstruction | Wrong task inherited from saved labels | sign of 15-minute forward return in BPS |
| Raw tree features | Scaling mismatch for tree baseline | tree model does not require RobustScaler |

The naive version would be shorter. Load `X`, load `y`, fit classifier, report accuracy. It would also be wrong in exactly the way that makes a result hard to debug: the code would run, the metrics would print, and the comparison would look scientific. The error would live in the meaning of `y` and the geometry of the split, not in a stack trace.

That distinction matters because most bad financial ML baselines do not fail loudly. They often fail by making the wrong thing convenient. A cached label array, a random split helper, a global shuffle, a validation set reused for early stopping, a forward-return target created before de-overlap: none of these choices necessarily creates an obvious programming error. They create an evidentiary error. The model may be implemented correctly while the experiment answers a question I did not intend to ask.

## Why the saved label file had to be rejected

A directional classifier is only as honest as its labels. In this baseline, the target is “15-minute forward return in BPS; binary label is sign(target).” That gives the model a two-sided problem: up versus down after filtering out the no-trade zone.

The existing saved `train_y.npy` artifact did not satisfy that contract. It contained magnitude-filtered positive targets only. That makes it unsuitable for a directional classifier because it no longer represents the binary sign task the baseline is supposed to measure. There is no clever model-side fix for that. Once the target artifact encodes the wrong task, using a different classifier just gives the wrong task a new costume.

The baseline therefore goes back to `data/tmp_sniper_feat_*.parquet`. That choice matters because the feature parquet is upstream of the bad shortcut. It lets the script reconstruct labels under the same protocol used by the sniper preprocessing path: 46 engineered microstructure-proxy features, 15-minute forward return in basis points, stride-15 de-overlap, the no-trade filter, and the per-asset temporal split with purge.

This is the part of baseline design that feels unglamorous but decides whether the result means anything. A model-class objection says, “Maybe the neural family is the reason the minute result clusters near 52%.” A contaminated label artifact would make the answer meaningless. Reconstructing labels from feature parquet keeps the comparison focused on capacity instead of accidentally comparing two different tasks.

The same principle applies outside this specific paper. If an intermediate artifact was built for a different objective, it is not a neutral cache. It is an encoded research decision. A saved target array can carry filtering, horizon, class definition, censoring, asset selection, and split assumptions. If those assumptions no longer match the experiment, downstream code should not pretend the file is just bytes on disk. It is a contract, and in this case the contract was wrong for the classifier I needed to run.

## The capacity objection needed a clean target

The minute-ceiling section reports seven configurations. The first three use tens of millions of labels and increasingly rich feature sets. The fourth and fifth reduce sample size but de-overlap and alter the loss. The sixth uses a different 15-minute basis-point target and compact microstructure-proxy features. The seventh replaces the neural CQR family with a classical tree model on the frozen microstructure-proxy matrix.

That seventh row is the key capacity-control move. If the ceiling were merely an artifact of the neural setup, a different model class on the same frozen feature/target construction should have had a chance to break away. Instead, the LightGBM baseline reports 52.231% accuracy with CI `[52.042, 52.420]`, AUC `0.53046`, balanced accuracy `52.21%`, and `n=268342` validation samples. In the minute configurations table, that appears as 52.23% ± 0.189 for row 7.

&amp;gt; **The ceiling holds — 52.231% accuracy, 95% CI `[52.042, 52.420]`.**
&amp;gt; Swap the neural CQR family for a classical tree on the same frozen feature and target matrix, and the result lands in the same 51.4–52.3% band. The minute ceiling is not an artifact of the neural setup.

For metric provenance, this was a CPU-only LightGBM run from `scripts/run_lightgbm_sniper_baseline.py` using the Python scientific stack in the paper environment: Polars for parquet loading, NumPy for arrays, SciPy for the binomial interval and significance calculation, scikit-learn metrics for balanced accuracy and AUC, and the LightGBM scikit-learn API for the classifier. The run writes `data/lightgbm_sniper_baseline.json` and `data/lightgbm_sniper_baseline_model.txt`; the reported metric is taken from that JSON artifact, not copied by hand into the manuscript.

The important part is not that LightGBM has a particular personality. It is that the baseline swapped model class while holding the validation geometry and target construction in place. Without that, “LightGBM versus neural” would be a noisy argument about everything except the model.

The training block reflects that narrow purpose. It uses a tree classifier with raw engineered features, reserves the final 15% of the training side for early stopping, and computes class weighting from the fit subset. The held-out validation interval remains untouched by early stopping.



```python
X_train, y_train, X_val, y_val, feature_names, asset_stats = _load_sniper_arrays(pramaana)

fit_end = int(len(X_train) * 0.85)
X_fit, X_es = X_train[:fit_end], X_train[fit_end:]
y_fit, y_es = y_train[:fit_end], y_train[fit_end:]

n_pos = int(y_fit.sum())
n_neg = int(len(y_fit) - n_pos)
scale_pos_weight = n_neg / max(n_pos, 1)

clf = lgb.LGBMClassifier(
    n_estimators=2000,
    learning_rate=0.03,
    num_leaves=31,
    max_depth=6,
    min_child_samples=200,
    subsample=0.85,
    colsample_bytree=0.85,
    reg_alpha=0.5,
    reg_lambda=1.0,
    scale_pos_weight=scale_pos_weight,
    objective="binary",
    metric="binary_logloss",
    n_jobs=-1,
    random_state=42,
    force_col_wise=True,
    verbosity=-1,
)

clf.fit(
    X_fit,
    y_fit,
    eval_set=[(X_es, y_es)],
    eval_metric="binary_logloss",
    callbacks=[
        lgb.early_stopping(stopping_rounds=100),
        lgb.log_evaluation(period=100),
    ],
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The hyperparameters are not the story I care about here. The non-obvious detail is the split inside the training side: the final 15% of train is used only for early stopping, which keeps validation as the final held-out interval rather than letting it become a tuning surface.&lt;/p&gt;

&lt;p&gt;That one choice prevents a common baseline failure. If I had passed the paper validation interval as the LightGBM &lt;code&gt;eval_set&lt;/code&gt;, early stopping would have made the validation set part of the training procedure. The model would not directly fit labels from validation, but the selected number of boosting rounds would be chosen by validation performance. That is enough to contaminate the final metric. The point of the baseline is not to squeeze the last basis point out of LightGBM; it is to answer whether a classical tree classifier breaks the minute ceiling under the same target and split discipline.&lt;/p&gt;

&lt;p&gt;The results JSON captures the protocol in a compact form. This is the artifact I want beside the paper because it records not only the metric, but also the target, preprocessing, split, and scaling assumptions that make the metric interpretable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"LightGBM M7 sniper baseline"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"created_utc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-05-04T02:58:24Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"source"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/home/the author/Development/Python/crypto-fl-v2/data/tmp_sniper_feat_*.parquet"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"protocol"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"feature_set"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"46 microstructure proxy features"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"15-minute forward return in BPS; binary label is sign(target)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"preprocessing"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stride-15 de-overlap; |target| &amp;gt; 10 BPS no-trade filter"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"split"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Per-asset 85/15 temporal split with 15-row purge before validation; final 15% of train used only for early stopping"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"scaling"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Raw engineered features; tree model does not require RobustScaler"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"train_samples"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1520254&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"early_stop_samples"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;228039&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"validation_samples"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;268342&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"n_features"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;46&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I care more about the &lt;code&gt;protocol&lt;/code&gt; object than the timestamp. A metric without this surrounding geometry is just a number looking for a story; the protocol is what prevents the wrong story from attaching itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Temporal bleed is quiet
&lt;/h2&gt;

&lt;p&gt;The hard part about temporal bleed is that it rarely announces itself. You do not get an exception that says, “Your validation rows are too close to your training rows.” You get a better-looking number, and that number is seductive because it can be explained as model quality.&lt;/p&gt;

&lt;p&gt;The minute experiments are especially exposed to this because the target horizon is short and overlapping labels are easy to create accidentally. A 15-minute forward return target means row &lt;code&gt;t&lt;/code&gt; and row &lt;code&gt;t+1&lt;/code&gt; can be describing heavily shared future intervals. If the split boundary cuts through those neighborhoods without a purge, the validation side can remain too close to what the training side has already seen.&lt;/p&gt;

&lt;p&gt;That is why the geometry is engineered at multiple levels instead of relying on one protective trick.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Hazard&lt;/th&gt;
&lt;th&gt;Naive baseline failure&lt;/th&gt;
&lt;th&gt;Geometry response&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Saved label artifact has wrong semantics&lt;/td&gt;
&lt;td&gt;Classifier trains on a target that is not the directional task&lt;/td&gt;
&lt;td&gt;Reconstruct labels from feature parquet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Adjacent windows share future information&lt;/td&gt;
&lt;td&gt;Validation resembles training near the boundary&lt;/td&gt;
&lt;td&gt;Insert 15-row purge before validation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Overlapping targets inflate sample familiarity&lt;/td&gt;
&lt;td&gt;Many rows encode nearly the same horizon&lt;/td&gt;
&lt;td&gt;Apply stride-15 de-overlap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Asset timelines differ&lt;/td&gt;
&lt;td&gt;Global shuffle mixes unrelated time positions&lt;/td&gt;
&lt;td&gt;Split per asset temporally&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Early stopping consumes validation signal&lt;/td&gt;
&lt;td&gt;Held-out set becomes part of model selection&lt;/td&gt;
&lt;td&gt;Use final 15% of train only for early stopping&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is also why I resist treating train/test split as an afterthought in financial ML. For IID tabular data, the split is often a convenience. For time series, the split is a claim about causality. It says what information was available before the prediction time and what was not. If that claim is false, the model can be perfectly implemented and still be scientifically useless.&lt;/p&gt;

&lt;p&gt;The per-asset split is important for a second reason: cryptocurrency pairs do not all have identical listing histories, liquidity regimes, or missing-data structure. A global split by row count would mix assets at unrelated calendar positions. A global random split would be worse. The baseline needs each asset’s validation samples to come from the end of that asset’s own history, with the purge applied at that asset boundary. That preserves the meaning of “future held out” even when the panel is irregular.&lt;/p&gt;

&lt;p&gt;The no-trade filter is a different kind of guardrail. It does not prevent leakage; it prevents the classifier from being evaluated on economically tiny moves as if every infinitesimal return were a meaningful directional event. The target is still a research abstraction, not a full trading simulator with costs and execution, but the &lt;code&gt;|target| &amp;gt; 10 BPS&lt;/code&gt; filter keeps the sign label from being dominated by noise around zero. That matters when the entire empirical question is about a narrow 51.4% to 52.3% band. At that scale, sloppy target construction can easily masquerade as signal.&lt;/p&gt;

&lt;h2&gt;
  
  
  The minute ceiling becomes more credible when the baseline cannot cheat
&lt;/h2&gt;

&lt;p&gt;The paper's data-and-methods section frames the smaller de-overlapped and microstructure-proxy runs as controls against cleaner labels, different short-horizon targets, and a classical tree classifier. The LightGBM row belongs to that logic. It is not a new headline result; it is a stress test against a specific objection.&lt;/p&gt;

&lt;p&gt;The minute ceiling is not “models cannot beat chance.” The reported range is above chance in a statistical sense. The issue is effect size. Across the seven configurations, directional accuracy remains inside a 0.9 percentage-point band, from 51.4% to 52.3%. With sample sizes in the millions for the larger configurations, the question is not whether the models detect something. The question is why substantial changes in features, scaling, loss functions, target construction, model class, and sample count do not move the result into a stronger range.&lt;/p&gt;

&lt;p&gt;The hourly positive control is what keeps this from becoming nihilism. At the hourly horizon, the FT-Transformer/CQR walk-forward run reached 54.69% directional accuracy on 386,056 future-held-out hourly samples, with a 95% interval of 54.53–54.84 and 80.53% conformal coverage. That result comes from the five-fold expanding walk-forward CQR run recorded in &lt;code&gt;backtest_results/walkforward_cqr_hourly.json&lt;/code&gt; and the corresponding resume log, executed with the PyTorch FT-Transformer/CQR training stack on the project’s CUDA workstation environment. The important runtime distinction is that the hourly result is not a LightGBM CPU run: it is a neural FT-Transformer/CQR evaluation under expanding walk-forward temporal validation, with the metric emitted by the walk-forward script and then pulled into &lt;code&gt;data/results_manifest.json&lt;/code&gt; for paper table generation.&lt;/p&gt;

&lt;p&gt;That contrast matters because it shows the research stack does not mechanically inflate every task. Minute OHLCV direction clusters near the ceiling; hourly direction recovers a larger signal under stricter temporal protocol.&lt;/p&gt;

&lt;p&gt;But that contrast only means something if the minute baseline is clean. If the LightGBM run had used a broken label artifact, then the row would not close the capacity objection. It would open a new hole. Reconstructing labels from parquet and enforcing split geometry is what lets the baseline answer the narrow question it was built to answer.&lt;/p&gt;

&lt;p&gt;The evidence-bundle design reinforces this. The paper repository is separate from the trading system because the paper is the manuscript and reproducibility layer, while Pramaana is the experiment apparatus. The paper artifacts point back to the scripts and result files that generated each table row: &lt;code&gt;scripts/run_lightgbm_sniper_baseline.py&lt;/code&gt;, &lt;code&gt;data/lightgbm_sniper_baseline.json&lt;/code&gt;, &lt;code&gt;data/lightgbm_sniper_baseline_model.txt&lt;/code&gt;, and the per-asset parquet inputs for the LightGBM capacity control; &lt;code&gt;backtest_results/walkforward_cqr_hourly.json&lt;/code&gt;, the walk-forward logs, and &lt;code&gt;scripts/walkforward_cqr_hourly.py&lt;/code&gt; for the hourly positive control. That is the level at which a result becomes inspectable. A table row should not be a manually typed number; it should be the visible tip of an artifact chain.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the feature importances can and cannot say
&lt;/h2&gt;

&lt;p&gt;The LightGBM run also writes a model file and feature importances. The top features include &lt;code&gt;return_60m&lt;/code&gt;, &lt;code&gt;z_effort_60m&lt;/code&gt;, &lt;code&gt;price_vs_vwap_5m&lt;/code&gt;, and &lt;code&gt;z_effort_15m&lt;/code&gt;. I do not treat those importances as a market theory. Tree feature importance is useful for checking that the model is not obviously broken, but it is not a causal explanation of minute returns.&lt;/p&gt;

&lt;p&gt;What it can say is more modest: the classifier found most of its splits in recent return, effort, and VWAP-relative features, which is consistent with the feature family the baseline was meant to test. That helps catch the opposite failure mode: a model that reports a plausible metric while accidentally training on an ID column, timestamp encoding, or mislabeled target. Feature inspection is not proof, but it is a useful diagnostic after the validation geometry is already correct.&lt;/p&gt;

&lt;p&gt;This is also why I report AUC alongside accuracy. Accuracy answers the sign-decision question at the default threshold. AUC checks whether the probability ranking carries directional information independent of that threshold. In the LightGBM run, AUC is &lt;code&gt;0.53046&lt;/code&gt;, which aligns with the story told by 52.231% accuracy: the model is detecting a small signal, not discovering a dramatically separable classification boundary.&lt;/p&gt;

&lt;p&gt;The high-confidence slice is similarly restrained. The JSON records a threshold of &lt;code&gt;p &amp;gt;= 0.60 or p &amp;lt;= 0.40&lt;/code&gt;, with 2,305 samples, 0.859% coverage, and 57.007% accuracy. That is interesting as a calibration and selectivity diagnostic, but it does not rescue the minute problem. A tiny high-confidence region with better accuracy is not the same as a broad tradable edge, especially before costs, slippage, and execution constraints. For the paper’s central claim, the full validation metric is the right number to emphasize.&lt;/p&gt;

&lt;h2&gt;
  
  
  The discipline is to make shortcuts impossible
&lt;/h2&gt;

&lt;p&gt;The small engineering choice I would generalize is this: when a saved artifact has ambiguous or wrong semantics, do not patch around it downstream. Rebuild from the nearest artifact whose meaning is still compatible with the experiment.&lt;/p&gt;

&lt;p&gt;In this case, the nearest compatible artifact was the per-asset feature parquet. That forced the script to carry the target definition, filtering rule, de-overlap policy, split geometry, purge width, and scaling assumption in the same place as the model run. The result is less convenient than loading &lt;code&gt;train_y.npy&lt;/code&gt;, but it is more honest.&lt;/p&gt;

&lt;p&gt;There is a tradeoff. Reconstructing labels from parquet couples the baseline to the upstream feature files and requires the paper repo to reference Pramaana's data path. The evidence bundle handles that by recording local evidence anchors rather than pretending the manuscript repository contains every large matrix. Compact artifacts get tracked directly with metadata and hashes; larger upstream matrices are referenced by path and size unless full hashing is requested. That is an engineering compromise, but it keeps the paper repository focused on reproducibility rather than becoming a warehouse for heavyweight experiment outputs.&lt;/p&gt;

&lt;p&gt;The payoff is that the baseline result has a shape I can defend. It says: on the frozen M6 sniper matrix, with 46 microstructure-proxy features, 15-minute forward-return sign labels, stride-15 de-overlap, a no-trade filter, per-asset temporal splitting, a 15-row purge, and a held-out validation interval, a classical tree model lands at 52.231% accuracy. That is a capacity-control statement, not a vague leaderboard entry.&lt;/p&gt;

&lt;p&gt;It also changes how I evaluate future baselines. I want the experiment script to make the invalid path difficult. If the wrong label file is easy to load, someone eventually will. If early stopping on validation is a one-line convenience, it will creep in. If a random split helper is available in the same file as time-series code, it will eventually be called in the wrong place. Good research code does not merely implement the correct protocol; it removes temptations that produce attractive but meaningless numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed in how I think about baselines
&lt;/h2&gt;

&lt;p&gt;I used to think of baselines as simpler models. I now think of them as simpler claims. A good baseline should remove one objection at a time. This LightGBM run removes the objection that the minute ceiling is merely a neural architecture artifact. It does not claim to solve execution, transaction costs, order-book dynamics, or richer market state. It only says that replacing the model family on this engineered short-horizon matrix does not break the ceiling.&lt;/p&gt;

&lt;p&gt;That restraint is the virtue. If a baseline tries to answer every objection simultaneously, it becomes another opaque system. If it answers one objection under carefully preserved geometry, it becomes useful evidence.&lt;/p&gt;

&lt;p&gt;For time-series ML, the geometry is the evidence. The split, purge, stride, and label source are not clerical details after the “real” modeling work. They are the rails that keep the model from learning yesterday's shadow of tomorrow.&lt;/p&gt;

&lt;p&gt;A classifier trained on the wrong artifact can look competent; a classifier trained inside the wrong timeline can look brilliant. I would rather have the modest number I can trust than the impressive one produced by a boundary I forgot to draw. The next credible test is not another minute-candle model with a larger parameter count; it is adding the state variables that the minute horizon is missing — order-book pressure, queue dynamics, spread formation, and execution-aware microstructure — while keeping the same discipline about labels, time, and evidence.&lt;/p&gt;

&lt;p&gt;In time-series work, validation geometry is not an implementation detail — it is the contract that keeps the model honest. The same discipline that protects a live trading gate or a self-indexing memory write protects the scientific claim.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>timeseries</category>
      <category>financialml</category>
      <category>lightgbm</category>
    </item>
    <item>
      <title>Vector Split by Chunk: Why My Retrieval Stops at the Boundary I Drew</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:02:31 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/vector-split-by-chunk-why-my-retrieval-stops-at-the-boundary-i-drew-3llf</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/vector-split-by-chunk-why-my-retrieval-stops-at-the-boundary-i-drew-3llf</guid>
      <description>&lt;p&gt;I watched a draft miss the exact file span I needed, and the failure was embarrassingly clean: the vector was "close," but the chunk I wanted was buried inside a larger blob. That is the kind of miss that looks acceptable in a demo and annoying in production, because the answer is almost right in the way a blurry photo is almost a portrait.&lt;/p&gt;

&lt;p&gt;The mechanic underneath that failure is what dense retrieval people call &lt;strong&gt;representational collapse&lt;/strong&gt;: when you ask one embedding to summarize an entire document, the model picks the dominant mode and discards the long tail. That tail is where the precise sentence lives. So the retrieval system returns the right neighborhood and the wrong house — coherent topically, useless operationally.&lt;/p&gt;

&lt;p&gt;The fix is to stop asking one vector to do that work. Split the text before embedding, embed each slice independently, and let retrieval search the slices instead of the document. The rest of this post is what falls out of that decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boundary I chose on purpose
&lt;/h2&gt;

&lt;p&gt;I split embeddings by chunk because the retrieval layer needed a smaller unit than a document. A whole-file vector is too coarse when the system has to answer from a specific path, a specific code block, or a specific rewrite. The chunk is the right compromise: small enough to isolate one semantic mode, large enough to preserve the local context around it.&lt;/p&gt;

&lt;p&gt;The implementation starts with a simple windowing function in &lt;code&gt;lib/supabaseVectorStorage.ts&lt;/code&gt;. It does one thing — slice text into overlapping fixed-size chunks. That is the dumbest viable splitter, and it is on purpose. Smarter chunkers exist (sentence-aware splitters, recursive structural splitters, semantic-similarity splitters like LangChain's &lt;code&gt;SemanticChunker&lt;/code&gt;), but they all bring their own failure modes and tuning surfaces. A naïve windowed splitter has exactly two knobs and zero hidden behavior, which is the right baseline before you let cleverness in.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="cm"&gt;/**
 * Split text into chunks for embedding
 */&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;splitTextIntoChunks&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;chunkSize&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;overlap&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;chunkSize&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;end&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;end&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;overlap&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The overlap is doing more work than it looks like. It is encoding a locality assumption — that an important sentence often straddles a chunk boundary, and that the embedding for either neighbor will still anchor near that sentence in vector space. Drop the overlap and you turn coherent thoughts into disjoint halves with weaker mutual similarity. Common starting values in production are 10–20% overlap; this one runs at 20% (200 of 1000), which sits at the upper end and trades a bit of storage for resilience at boundaries. The Pinecone team's &lt;a href="https://www.pinecone.io/learn/chunking-strategies/" rel="noopener noreferrer"&gt;chunking strategies guide&lt;/a&gt; is the canonical writeup if you want the longer comparison.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why the naive version fails
&lt;/h3&gt;

&lt;p&gt;Embedding the whole document and calling it done works until retrieval needs a narrow answer. Then the vector starts behaving like a résumé summary: it remembers the broad shape, but not the sentence where the detail actually lives. That is the precision side of the precision/recall tradeoff failing — you keep finding &lt;em&gt;related&lt;/em&gt; documents and missing the &lt;em&gt;exact&lt;/em&gt; one.&lt;/p&gt;

&lt;p&gt;Chunk-level embedding changes the unit of truth. Instead of asking one embedding to represent everything, I ask many embeddings to represent adjacent slices, then let cosine similarity decide which slice the query is closest to. More embeddings means more candidate matches per document, which means the retrieval layer has a much better chance of finding the specific span that answers the question instead of the document that contains it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the chunk becomes a vector
&lt;/h2&gt;

&lt;p&gt;Once the text is split, each chunk is embedded independently with OpenAI. The embedding call is tied to the chunk text, not the parent document.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="cm"&gt;/**
 * Generate embedding for a given text using OpenAI
 */&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;generateEmbedding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;embeddings&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;text-embedding-3-large&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;dimensions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Error generating embedding:&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things in that call are worth naming. First, &lt;code&gt;input: text&lt;/code&gt; is fed the chunk, not the document — the whole purpose of the split. Second, &lt;code&gt;dimensions: 2000&lt;/code&gt; exploits OpenAI's &lt;strong&gt;Matryoshka representation learning&lt;/strong&gt;: &lt;code&gt;text-embedding-3-large&lt;/code&gt; is natively 3072-dimensional, but the model is trained so that any leading prefix of the vector is itself a usable embedding. Truncating to 2000 dims keeps roughly 99% of the retrieval quality at two-thirds of the storage cost and roughly two-thirds of the cosine computation per query. For a system that stores tens of thousands of chunks behind an HNSW index, that compounds fast.&lt;/p&gt;

&lt;p&gt;The storage path follows the same logic when chunks are inserted. Each chunk gets its own record with its own index and metadata.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunkPromises&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunkText&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;index&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;embedding&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generateEmbedding&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunkText&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;document_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;documentId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;chunk_index&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;index&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;chunk_text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;chunkText&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// Pass as array, Supabase will handle vector conversion&lt;/span&gt;
    &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;chunk_index&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;index&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;total_chunks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunksWithEmbeddings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;all&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunkPromises&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// Insert chunks with embeddings&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;embError&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;supabase&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;embeddings&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;insert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunksWithEmbeddings&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The non-obvious part is that the metadata carries &lt;code&gt;chunk_index&lt;/code&gt; and &lt;code&gt;total_chunks&lt;/code&gt; alongside the vector. That makes each chunk usable on its own &lt;em&gt;and&lt;/em&gt; still locatable in the original document — which is what enables a downstream pattern called &lt;strong&gt;context expansion&lt;/strong&gt;: when a chunk matches, you can also fetch its neighbors by index to widen the window the LLM sees without polluting the similarity score. Retrieval stays narrow; presentation gets to be generous.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why this pattern beats one vector per file
&lt;/h3&gt;

&lt;p&gt;A single file vector is cheap to reason about and expensive to trust. It compresses too much. Splitting first and embedding second preserves the local neighborhood around each idea, which is what lets the search layer return the exact span instead of the general topic. The honest cost is more rows, more embeddings, and more storage — but in retrieval, precision is usually the thing that keeps the rest of the system from sounding vaguely confident and slightly wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The retrieval path that makes the split worth it
&lt;/h2&gt;

&lt;p&gt;The strongest evidence that the split matters is the retrieval helper that bypasses vector similarity entirely when I already know the file path. In &lt;code&gt;supabase/functions/_shared/blog-utils.ts&lt;/code&gt;, the helper queries embeddings by &lt;code&gt;metadata-&amp;gt;&amp;gt;'file_path' LIKE '%&amp;lt;path&amp;gt;%'&lt;/code&gt; so I can guarantee the accuracy agent sees chunks from the file the draft explicitly mentions.&lt;/p&gt;

&lt;p&gt;This is &lt;strong&gt;hybrid retrieval&lt;/strong&gt; in its simplest form: lexical exact-match (the &lt;code&gt;LIKE&lt;/code&gt; on file path) and dense vector search (the cosine query) live side by side and compose by union. Vector search is great when I need semantic recall — "find me the chunk that talks about this idea, even if it uses different words." Path-based retrieval is what I reach for when I need the system to stop being poetic and start being literal — "find me chunks from this exact file." Production retrieval stacks at scale usually go further and combine BM25 lexical scores with dense vectors via reciprocal rank fusion; this is the minimal version of that pattern, sized to the problem.&lt;/p&gt;

&lt;p&gt;The chunk-level split is what makes the path-based path useful, because the file can now return multiple precise spans instead of one oversized blob. If I am asking about a function, I want the function's neighborhood, not the whole neighborhood's autobiography.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the storage model
&lt;/h2&gt;

&lt;p&gt;The storage row carries the &lt;code&gt;document_id&lt;/code&gt;, &lt;code&gt;chunk_index&lt;/code&gt;, &lt;code&gt;chunk_text&lt;/code&gt;, &lt;code&gt;embedding&lt;/code&gt;, and metadata. That shape is what makes downstream retrieval and debugging sane: when I inspect a result, I can tell where it came from, where it sits in the source, and how many sibling chunks exist around it. The schema is simple enough to survive maintenance and specific enough to survive scrutiny — a combination that is rarer than it should be.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;DocumentChunk&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;id&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;document_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;chunk_index&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;chunk_text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;embedding&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;[];&lt;/span&gt;
  &lt;span class="nl"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="nb"&gt;Record&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;unknown&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The practical tradeoff: granularity buys control, not magic
&lt;/h2&gt;

&lt;p&gt;Chunking does not make retrieval smart by itself. It gives the search layer a better surface to work with — and that surface has a goldilocks zone. Too large and chunks blur together into mid-density topical clouds; too small and they lose enough context to become syntactic fragments whose embeddings collapse into the average. The 1000-character window here is roughly two paragraphs of prose or one tight function — large enough to carry meaning, small enough to admit only one dominant idea per row.&lt;/p&gt;

&lt;p&gt;The overlap is the design's safety net. It encodes the assumption that meaning is &lt;em&gt;local but not aligned&lt;/em&gt; — important sentences refuse to respect window boundaries — and pays a storage tax to keep adjacent chunks neighbors in vector space. This is the same tradeoff that shows up in sliding-window attention (preserve locality at the cost of duplicated work) and in n-gram lexical indexes (overlap at the boundary to avoid losing cross-token matches).&lt;/p&gt;

&lt;p&gt;The other cost is operational. More chunks mean more embeddings, and more embeddings mean more insert work and more bytes in the HNSW graph. I accept that because the alternative is a retrieval system that keeps returning the right topic and the wrong answer, which is a very expensive way to be unhelpful.&lt;/p&gt;

&lt;h2&gt;
  
  
  The retrieval story, drawn plainly
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  A["Source text"] --&amp;gt; B["Chunk split"]
  B --&amp;gt; C["OpenAI embed"]
  C --&amp;gt; D["Embeddings table"]
  D --&amp;gt; E["File-path query"]
  D --&amp;gt; F["Vector search"]
  E --&amp;gt; G["Chunk result"]
  F --&amp;gt; G
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why I kept the split at the chunk boundary
&lt;/h2&gt;

&lt;p&gt;I could have pushed more logic into the embedding step — late-interaction models like ColBERT preserve per-token embeddings and aggregate at query time, which gives you precision without choosing a chunk size up front. I could have tried to recover precision later with cross-encoder reranking. But both move the problem instead of solving it; they paper over a representation that is wrong at the unit of storage.&lt;/p&gt;

&lt;p&gt;So I split by chunk, embed by chunk, store by chunk, and retrieve by chunk. The system is easier to inspect because every stage speaks the same language. That is the real payoff: not just better recall, but a retrieval stack whose granularity matches the way I actually ask questions.&lt;/p&gt;

&lt;p&gt;When the answer lives inside a file, I want the search layer to arrive with a scalpel, not a shovel.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>embeddings</category>
      <category>rag</category>
      <category>vectorsearch</category>
      <category>supabase</category>
    </item>
    <item>
      <title>Four Vectors, One Record: How I Split Embeddings Before They Hit Search</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:02:20 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/four-vectors-one-record-how-i-split-embeddings-before-they-hit-search-211i</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/four-vectors-one-record-how-i-split-embeddings-before-they-hit-search-211i</guid>
      <description>&lt;p&gt;The failure that pushed me into this design was not subtle. I had a blended embedding that kept returning candidate matches that looked reasonable at a glance and wrong on inspection. A query about a very specific certification would drag in people with the right industry background but no credential. A query about deep experience in a narrow domain would surface candidates whose skills summary happened to mention the right keywords but whose work history did not support the match. The single vector was doing exactly what I asked it to do: compressing the entire record into one semantic point. The problem was that a candidate record is not one thing.&lt;/p&gt;

&lt;p&gt;This is the same failure dense-retrieval people call &lt;strong&gt;representational collapse&lt;/strong&gt;: when you ask one embedding to summarize a heterogeneous object, the model gravitates to the dominant mode and discards the long tail. For a document, that tail is a paragraph. For a candidate record, the tail is an entire field. Credentials get smoothed by industry text. Career trajectory gets smoothed by a polished summary. A single vector turns a structured record into a compromise the search layer has to live with.&lt;/p&gt;

&lt;p&gt;A profile has several different centers of meaning — multiple &lt;strong&gt;semantic modes&lt;/strong&gt; that do not collapse cleanly. Work history answers one class of search intent. Skills and designations answer another. A broad profile summary is useful for discovery, but it is a poor substitute for the details that actually separate one candidate from another. The fix is to stop forcing one embedding to carry all of them.&lt;/p&gt;

&lt;p&gt;That is what &lt;code&gt;app/agents/embedding_agent.py&lt;/code&gt; is for in this system. It generates four parallel embeddings for the same logical record: &lt;code&gt;profile_vector&lt;/code&gt;, &lt;code&gt;experience_vector&lt;/code&gt;, &lt;code&gt;skills_vector&lt;/code&gt;, and &lt;code&gt;general_vector&lt;/code&gt;. This is &lt;strong&gt;late fusion&lt;/strong&gt; in the retrieval sense — embed each view independently, defer the "which view matched" decision until query time, and let the search layer pick the best-matching vector per query. Then &lt;code&gt;app/jobs/embedding_generator.py&lt;/code&gt; takes over the production concerns: it validates the input, applies retry logic, and routes terminal failures to the dead-letter queue. The separation is deliberate. The agent handles semantic decomposition and caching. The job handles durability and failure control. That two-layer split is the &lt;strong&gt;bulkhead pattern&lt;/strong&gt; — keep the failure modes of one concern from contaminating the other, the way watertight compartments work on a ship.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the pipeline
&lt;/h2&gt;

&lt;p&gt;I like systems that are honest about their steps.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  sourceRecord[Source record] --&amp;gt; fieldSplit[Split into 4 field views]
  fieldSplit --&amp;gt; embeddingAgent[Embedding agent generates 4 vectors]
  embeddingAgent --&amp;gt; redisCache[Redis cache with 24h TTL]
  embeddingAgent --&amp;gt; embeddingJob[Embedding job]
  embeddingJob --&amp;gt; validate[Validate length and dimension]
  validate --&amp;gt; retryFailures[Retry transient failures]
  validate --&amp;gt; deadLetter[Route terminal failures to DLQ]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The agent, its constants, and the cache key
&lt;/h2&gt;

&lt;p&gt;The embedding agent is specialized on purpose. Its job is not "make vectors." Its job is "make four specific vectors, keep them cached, and do it fast enough that repeated calls do not punish the API."&lt;/p&gt;

&lt;p&gt;The contract is declared as constants at the top of the file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;EMBEDDING_MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text-embedding-3-large&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;EMBEDDING_DIMENSIONS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3072&lt;/span&gt;  &lt;span class="c1"&gt;# text-embedding-3-large native dimensions
&lt;/span&gt;&lt;span class="n"&gt;EMBEDDING_TTL_SECONDS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;86400&lt;/span&gt;  &lt;span class="c1"&gt;# 24 hours
&lt;/span&gt;&lt;span class="n"&gt;CACHE_KEY_PREFIX&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;emb:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those four lines are the boundary conditions for the whole embedding pipeline — what a working systems engineer would call a &lt;strong&gt;contract module&lt;/strong&gt;. The model choice defines the embedding family. The dimension count defines what downstream systems must expect, and what the HNSW or IVF index was sized for — if it ever drifts, the mismatch is immediate. The TTL is the literal floor on how stale the cache can get, which makes it a direct knob on the freshness/cost tradeoff. The prefix scopes the embedding namespace inside Redis, because a multi-tenant cache without prefixes becomes archaeology the moment anything goes wrong.&lt;/p&gt;

&lt;p&gt;The cache key itself is where multi-vector output stays safe:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_generate_cache_key&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vector_type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Generate a deterministic cache key from query text.

    Returns: emb:{md5_hash}:{vector_type}
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;normalized&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;query_hash&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;md5&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;normalized&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;CACHE_KEY_PREFIX&lt;/span&gt;&lt;span class="si"&gt;}{&lt;/span&gt;&lt;span class="n"&gt;query_hash&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;vector_type&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;vector_type&lt;/code&gt; token is the part that makes this cache safe for a four-vector bundle. This is &lt;strong&gt;key-space scoping&lt;/strong&gt; in the small — each vector view gets its own deterministic slot, so a lookup for &lt;code&gt;profile_vector&lt;/code&gt; cannot accidentally return the value cached for &lt;code&gt;skills_vector&lt;/code&gt;. Without the token, the keys would collide and a lookup would return whichever vector wrote last; with it, the cache stays correct under concurrent writes from all four views. The &lt;code&gt;lower().strip()&lt;/code&gt; normalization is what gives the key its &lt;strong&gt;deterministic property&lt;/strong&gt; — semantically identical queries that differ only in whitespace or case hit the same slot. Repeated content returns quickly. New content pays the model cost once. The cache stays legible because nothing else in Redis uses the &lt;code&gt;emb:&lt;/code&gt; prefix.&lt;/p&gt;

&lt;p&gt;The agent then runs the four model calls in parallel on a cache miss. The vectors are independent semantic views, so there is no reason to serialize them — &lt;code&gt;asyncio.gather&lt;/code&gt; runs the four embedding requests concurrently and the worst-case latency is &lt;strong&gt;fan-out latency&lt;/strong&gt;, not four-times-sequential. In practice the four calls complete in roughly the wall-clock time of a single call. The agent also tracks cache hits, cache misses, embeddings generated, errors, tokens used, and total latency. Those metrics are what separate a slow cache miss from a model issue from a bad keying strategy when the system gets noisy — observability has to be richer than the failure mode you are trying to debug, or you end up guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The job that makes the system durable
&lt;/h2&gt;

&lt;p&gt;The agent solves the semantic problem. The job in &lt;code&gt;app/jobs/embedding_generator.py&lt;/code&gt; solves the production problem.&lt;/p&gt;

&lt;p&gt;Its responsibility is not to invent vectors. Its responsibility is to make sure the vectors that enter the index are valid, and to make sure failures are classified correctly when something goes wrong. The job handles three things: content-length validation, embedding-dimension validation, and retry/DLQ routing.&lt;/p&gt;

&lt;p&gt;The retry policy is &lt;strong&gt;exponential backoff with capped intervals&lt;/strong&gt; — 30 seconds, 2 minutes, 10 minutes for transient failures. That growth pattern (~4× between steps) gives temporary upstream issues time to clear without flooding the pipeline with immediate retries, and the cap keeps the worst-case retry budget bounded. The pattern is older than queue infrastructure itself — TCP's RTO calculation uses the same shape — and it works because most transient failures recover on a timescale that does not require sub-second polling.&lt;/p&gt;

&lt;p&gt;But not every failure deserves another attempt. Some failures are &lt;strong&gt;terminal&lt;/strong&gt; by definition. If the content is too long, the input needs to change. If the dimension count does not match the expected shape, the embedding is not fit for indexing. Those cases go straight to the DLQ as &lt;strong&gt;poison messages&lt;/strong&gt; — input that no amount of retrying will fix because the retry policy operates on time, not on the cause of the failure.&lt;/p&gt;

&lt;p&gt;That is the right line to draw. A retryable failure says "try again later." A terminal failure says "this record needs intervention or upstream correction." Mixing those two up is how queues become &lt;strong&gt;retry-storm amplifiers&lt;/strong&gt; — the same bad input cycling through the worker pool, burning capacity that healthy traffic needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why terminal failures go straight to the DLQ
&lt;/h2&gt;

&lt;p&gt;I treat the dead-letter queue as an operational control surface, not as a waste bin. If a record lands there because the content is too long or the dimensions are wrong, that is useful signal. It tells me exactly what kind of correction is needed.&lt;/p&gt;

&lt;p&gt;That approach keeps the normal path clean. Retryable failures get another chance. Invalid records get isolated. The DLQ becomes the place where malformed input, schema drift, or upstream data quality issues are surfaced — instead of being hidden under repeated attempts that look like noise on a dashboard. That is also why the terminal failure list is explicit: &lt;code&gt;content_too_long&lt;/code&gt; and &lt;code&gt;dimension_mismatch&lt;/code&gt;. Those are not transient conditions. Retrying them would only burn time and make the failure harder to interpret.&lt;/p&gt;

&lt;p&gt;The job's logic protects downstream search from poison data. That is a boring sentence, and it is the whole point of the job.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real benefit of four vectors
&lt;/h2&gt;

&lt;p&gt;The payoff is not an abstract claim that multi-vector is better. It is that a candidate record can answer multiple search intents without being flattened into a single compromise representation. This is exactly the problem classical information retrieval solved with &lt;strong&gt;BM25F&lt;/strong&gt; — field-weighted BM25 that lets a search over "title, body, anchor text" weight each field separately rather than concatenating them into a single bag of words. The four-view embedding is the dense-retrieval analog: each field gets its own representation, and the search layer composes them at query time instead of relying on a pre-flattened average.&lt;/p&gt;

&lt;p&gt;The work-history view preserves sequence and career motion. The skills view preserves credentials, designations, and specific capabilities. The profile view keeps the higher-level picture intact. The general view gives broad coverage when the query is not narrowly about one field — the safety net for queries that do not map cleanly to any one mode.&lt;/p&gt;

&lt;p&gt;The record keeps its internal structure while still becoming searchable. If I had stayed with one blended embedding, queries about certification would have kept drifting toward domain-adjacent candidates, and queries about career history would have kept overvaluing polished summary language. Splitting the record before embedding makes those tradeoffs explicit instead of accidental — the search layer can now choose which view drives the score, and which one only gets to break ties.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this changes the way I debug search
&lt;/h2&gt;

&lt;p&gt;One of the best side effects of the split is that debugging got a lot cleaner. When the embedding layer is monolithic, every retrieval complaint feels like one problem. With four vectors and a separate job layer, &lt;strong&gt;each failure mode has a coordinate&lt;/strong&gt; — I can point at which vector, which stage, and which agent owns the problem instead of triangulating.&lt;/p&gt;

&lt;p&gt;If latency is high, I look at cache misses and generation time. If the system is burning API calls, I look at cache hit rate. If records are failing to persist, I check validation and DLQ counts. If the embedding shapes are wrong, I know the issue is in the contract, not in the ranking layer. That mapping from symptom to layer is what makes the system &lt;strong&gt;observable&lt;/strong&gt; in the precise sense — not "we have dashboards," but "every reasonable question has a corresponding metric to read."&lt;/p&gt;

&lt;p&gt;That split in responsibilities is what makes the system tractable. The embedding agent is about generating the right semantic inputs. The job is about making sure those inputs survive production constraints. Each layer is narrow enough to inspect without guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I kept the implementation narrow
&lt;/h2&gt;

&lt;p&gt;I did not want the embedding path to become a catch-all orchestration layer. A lot of systems become hard to reason about because one component starts doing semantic prep, retry control, persistence, error handling, and cache management all at once — what designers call &lt;strong&gt;god-object accretion&lt;/strong&gt;, where convenience eats coherence one feature at a time.&lt;/p&gt;

&lt;p&gt;This one stays focused. The agent prepares and caches the vectors. The job validates and routes outcomes. The search layer consumes the resulting embeddings through the normal indexing path. That separation keeps each piece easier to test, easier to replace, and free of accidental coupling — changing retry behavior never forces a rewrite of semantic preparation, and adjusting the vector split never touches dead-letter routing. It is the same logic that drives &lt;strong&gt;single-responsibility design&lt;/strong&gt; at the class level, applied at the service-boundary level.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the split preserved
&lt;/h2&gt;

&lt;p&gt;The thing I wanted to preserve was not just accuracy — it was the shape of the candidate record itself. A blended embedding makes a record searchable but smooths away the distinctions that separate one candidate from the next: experience is not the same as skills, and a polished title can hide a work history that tells a different story. Four semantic views keep those distinctions where search can still use them.&lt;/p&gt;

&lt;p&gt;That is the real win. The record remains one record, but it no longer has to pretend it only means one thing.&lt;/p&gt;

&lt;p&gt;The next step is to connect this multi-vector representation to the retrieval side in a way that keeps the same discipline: explicit field intent, explicit score handling, and no hidden magic between the index and the ranking layer. &lt;strong&gt;Late-interaction models&lt;/strong&gt; like ColBERT push this idea even further — preserving per-token vectors and aggregating at query time, which gives precision without committing to a field schema up front. But that is a much bigger lift, and the four-view split is the right step from where I started. Once the record is no longer flattened, the search layer has to earn its keep.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>embeddings</category>
      <category>search</category>
      <category>redis</category>
      <category>python</category>
    </item>
    <item>
      <title>Coverage Before Creativity: The RAG Gate That Keeps My Blog Pipeline Honest</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:02:09 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/coverage-before-creativity-the-rag-gate-that-keeps-my-blog-pipeline-honest-5gcj</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/coverage-before-creativity-the-rag-gate-that-keeps-my-blog-pipeline-honest-5gcj</guid>
      <description>&lt;p&gt;The first failure I had to eliminate in the blog pipeline was not a bad paragraph. It was a bad evidence set. The system was finding a few nearby chunks, mistaking density for coverage, and then drafting as if that narrow slice represented the whole repository. That produces text that sounds confident right up until you compare it with the code. The fix was to stop treating topic selection like a writing problem and start treating it like a retrieval coverage problem.&lt;/p&gt;

&lt;p&gt;That distinction matters. If the upstream evidence is thin, no amount of prompt polish saves the result. The draft will still overfit the first cluster of files that happened to match the query. I wanted the pipeline to be disciplined about breadth before it was creative about prose. So the gate moved earlier: query fan-out, file-path-aware dedupe, breadth validation, pinned excerpts, and only then the writing pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gate lives before the writing step
&lt;/h2&gt;

&lt;p&gt;In my pipeline, the important work happens before generation starts. The dispatcher is where the topic search fans out through multiple lanes: curated highlights, a fixed RAG query pool, and recent commit-derived queries. That is deliberate. A single semantic search tends to collapse into the same dense corners of the codebase, which is exactly where a system becomes persuasive and shallow at the same time.&lt;/p&gt;

&lt;p&gt;The dispatcher does not need to know how the post will read yet. Its job is to prove that the candidate topic has enough distinct evidence behind it to deserve a draft. That means the retrieval layer has to do more than collect relevant chunks. It has to show spread. It has to show that the match did not come from one file, one subsystem, or one repetitive cluster of adjacent chunks.&lt;/p&gt;

&lt;p&gt;That is why the shared blog utilities matter. The generator path imports &lt;code&gt;checkRagSufficiency&lt;/code&gt; and &lt;code&gt;fetchFailureEvidence&lt;/code&gt; from &lt;code&gt;supabase/functions/_shared/blog-utils.ts&lt;/code&gt;, and that is the right place for this logic: the part of the system that decides whether retrieval is good enough should sit close to the code that evaluates it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  A[Dispatcher: curated highlights] --&amp;gt; B[Query fan-out]
  C[Dispatcher: fixed RAG query pool] --&amp;gt; B
  D[Dispatcher: recent commit queries] --&amp;gt; B
  B --&amp;gt; E[RAG retrieval]
  E --&amp;gt; F[Dedupe by repo and file path]
  F --&amp;gt; G[Breadth check]
  G --&amp;gt;|passes| H[Pin excerpts]
  G --&amp;gt;|fails| I[Reject candidate]
  H --&amp;gt; J[Stage-compose merge]
  J --&amp;gt; K[Generator / drafting]```



That flow is the real control surface. The generator is downstream of a decision that already happened: is the evidence wide enough to trust?

## Three query lanes, three different jobs

The query fan-out is not random, and it is not a single blended prompt pretending to be a strategy. I built it as three separate signals because each one catches a different failure mode.

The curated highlight queries keep the system anchored in the kinds of features and systems I already know are worth revisiting. Those are the posts that usually come from places I have touched repeatedly: workflow orchestration, retrieval, caching, parsing, state management, security boundaries, or data transformation. They help the pipeline remember what is already interesting in the repository family.

The fixed query pool is the broadest lane. It exists to force coverage across architectural themes and implementation patterns rather than letting one topic family dominate. This is the part that looks for the general system shape: event-driven flows, retrieval logic, prompt construction, orchestration, caching, retry paths, auth boundaries, model inference, ETL, and state machines. If I let the selector live only inside the curated highlights, it would become too self-referential. If I let it live only inside the fixed pool, it would become too generic. The combination is what keeps the output grounded and varied.

The recent commit queries add the temporal dimension. They bias the selector toward what actually changed recently instead of letting the system drift into evergreen topics that no longer reflect the repository’s current shape. That matters because the most obvious topic is often the wrong one when recent work has shifted the architecture. A topic can be semantically relevant and still be stale in practice.

The point of splitting those lanes is simple: no single lane is trusted to decide the topic alone. They feed the same retrieval pass, but they do so for different reasons. One lane preserves editorial continuity. One lane broadens architectural search. One lane keeps the system current. When those three are merged, I get a candidate set that is much harder to fool with local similarity alone.

## Why I dedupe by repo and file path, not just text similarity

Once retrieval returns a pile of chunks, the next problem is repetition. Similarity search loves repetition. A single file can dominate a result set by surfacing multiple overlapping excerpts, especially when the file is dense or when several queries land in the same section of code. If I let that happen, the draft starts building itself around one artifact instead of one system.

That is why file-path-aware dedupe matters so much. I want repeated hits from the same repo and file path to collapse early. I do not care whether five adjacent chunks all sound relevant if they are all pointing at the same paragraph of the same file. What I care about is whether the sample spans distinct parts of the codebase.

This is also why repo identity belongs in the dedupe key. In a multi-repo setup, two chunks can look similar for completely different reasons. They may both describe retrieval logic, orchestration, or prompt shaping, but they live in different systems and should not be treated as interchangeable evidence. Repo plus file path tells me whether the system is sampling breadth or simply rediscovering the same neighborhood under different search terms.

The practical effect is that the retrieval layer becomes less greedy. It stops rewarding the first obvious cluster with extra representation. It stops counting repeated evidence as coverage. That makes the candidate set smaller, but it also makes it much more trustworthy.

There is a second-order benefit here too: dedupe reduces the risk that a single implementation detail becomes the skeleton of the whole post. Without it, the draft can end up over-explaining one helper, one file, or one branch of logic simply because retrieval happened to hit it several times. Dedupe breaks that bias before the generator ever sees the prompt.

## Breadth is not a vibe; it is a threshold

After dedupe, I do not ask whether the chunks feel diverse. I measure whether the sample is wide enough to support a post. That is the entire point of the breadth gate. It exists to reject candidate sets that are semantically plausible but structurally weak.

This is where `checkRagSufficiency` fits into the pipeline. The name is exactly what the behavior needs to be: a sufficiency check. If the retrieved set cannot prove enough spread across the repository, it should not advance. The system should fail closed, not guess.

The important thing about the breadth threshold is that it changes the meaning of retrieval. Retrieval is no longer a convenience layer that collects whatever looks closest. It becomes a gate that must establish evidence quality before writing begins. That changes the failure mode from "draft written from a narrow slice" to "candidate rejected because the sample is too narrow." I will take the second failure every time.

That rejection path matters in practice. It catches cases where the query pool lands too hard in one subsystem, cases where recent changes dominate the semantic neighborhood, and cases where one file generates too many overlapping hits. A lot of retrieval bugs look like success until the prompt is assembled. The breadth check is the thing that stops those bugs from turning into published text.

I also like that the failure can produce concrete evidence. `fetchFailureEvidence` belongs in the same shared utility layer because it gives me a way to inspect why a candidate was rejected. That is useful during tuning. If a topic keeps failing breadth, I can see whether the issue is query bias, inadequate file-path diversity, or a retrieval window that is too small for the amount of material I want to cover.

## The stage-compose step adds a second guardrail

File-path retrieval and semantic retrieval are not competing systems. File-path retrieval gives me boundary-aware evidence. Semantic retrieval gives me breadth across related concepts. In `blog-stage-compose`, I merge both and dedupe a second time, because the two strategies can land on the same excerpt from different angles — and the evidence set should not regress back into repetition right before drafting.

## Pinned excerpts keep the evidence from drifting

Once a candidate survives the coverage gate, I pin the excerpts that explain why the topic is worth writing about. That is not a cosmetic step. It is a stability step. Without pinned evidence, the drafting stage has too much freedom to wander away from the exact chunks that earned the topic in the first place.

Pinned excerpts act like an anchor for the generation pass. They preserve the evidence trail. They keep the prompt honest about the source material the draft was built from. That matters because the strongest failure mode in a retrieval-driven blog system is not complete hallucination. It is drift: the draft starts from real evidence, then gradually generalizes beyond what was actually retrieved.

Pinning the surviving excerpts makes that harder. It forces the generation stage to stay connected to the specific implementation details that passed the gate. It also makes review easier, because I can inspect exactly which chunks were considered important enough to carry forward.

The combination of pinning and breadth checking is what gives the pipeline its shape. Breadth says, "This sample is wide enough." Pinning says, "These are the exact pieces that justify the topic." Together, they stop the generator from inventing confidence where the retrieval layer did not earn it.

## Why I prefer rejection over a weak draft

I am completely comfortable with a pipeline that says no. In fact, I want it to say no when the evidence is bad. A weak candidate should not be rescued by a polished prompt. If the retrieval set is narrow, the honest response is rejection.

That discipline keeps the blog output specific. It keeps the writing tied to actual systems instead of generic patterns. It also saves me from having to edit around a draft that was born from a bad evidence shape. A rejected topic costs less time than a published post that looks right but misses the structure of the thing it claims to describe.

This is especially important in a multi-repo environment. Once you have a handful of systems with overlapping concepts, retrieval can become too eager to collapse them into one theme. A good gate has to resist that collapse. It has to respect repository boundaries, file boundaries, and evidence density boundaries. If those boundaries are not visible in the sample, the draft should not happen yet.

That is the philosophical change I stopped treating as optional. The system does not owe me a draft. It owes me a trustworthy sample. The sample has to prove that the topic is broad enough, current enough, and distinct enough to justify writing.

## The actual win is not creativity; it is control

What changed here was not my ability to generate prose. What changed was the quality of the evidence stack that the prose starts from. The dispatcher fans out through curated highlights, a fixed query pool, and recent commit queries. The retrieval layer dedupes by stable keys. The sufficiency gate checks for breadth. Stage-compose merges file-path and semantic evidence and dedupes again. Pinned excerpts keep the final prompt from drifting.

That sequence turns retrieval into a control system instead of a suggestion engine. It enforces a minimum standard before the draft is allowed to exist. That is the kind of discipline a blog pipeline needs if it is going to write about real systems with real precision.

The result is not just fewer bad drafts. It is a pipeline that knows what it does not know early enough to stop itself. When the evidence is wide, the writing stage does the part it is good at. When the evidence is narrow, the system does the part I am more grateful for: it refuses to pretend it knows.

---

🎧 **Listen to the audiobook** — [Spotify](https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D) · [Google Play](https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en) · [All platforms](https://www.craftedbydaniel.com/audiobook)
🎬 [Watch the visual overviews on YouTube](https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6)
📖 [Read the full 13-part series](https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>rag</category>
      <category>supabase</category>
      <category>nextjs</category>
      <category>typescript</category>
    </item>
    <item>
      <title>The Reward Calibrator That Learns the Shape of Its Own Judgment</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:01:59 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/the-reward-calibrator-that-learns-the-shape-of-its-own-judgment-2782</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/the-reward-calibrator-that-learns-the-shape-of-its-own-judgment-2782</guid>
      <description>&lt;p&gt;I hit the wall when a single noisy signal kept dragging candidate selection around like a shopping cart with one bad wheel. The composite score looked stable on paper, but the moment I watched different weight sets compete on real samples, it became obvious that I was tuning a judgment function by instinct and hoping the surface was kind.&lt;/p&gt;

&lt;h2&gt;
  
  
  The idea behind the calibrator
&lt;/h2&gt;

&lt;p&gt;The part that changed my approach was treating the weights as a search space, not a set of sacred constants. The reward mixer already exposed five signals — visualDrift, colorHarmony, motionContinuity, compositionStability, and narrativeCoherence — and the calibrator sits above that layer to ask a different question: which combination of weights best matches the benchmark data?&lt;/p&gt;

&lt;p&gt;That is a very different job from adjusting thresholds. A threshold says, “accept or reject at this line.” A calibrator says, “given these signals, what should the system believe more or less strongly?” That distinction matters because the calibrator is searching the parameter surface of the judgment function itself. It is not just moving a gate; it is reshaping the scorer that the gate depends on.&lt;/p&gt;

&lt;p&gt;The reward mixer starts with explicit signal types, and that structure is what makes calibration possible in the first place.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;RewardSignals&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="cm"&gt;/** CLIP embedding cosine similarity to source frame (0-1) */&lt;/span&gt;
  &lt;span class="nl"&gt;visualDrift&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;
  &lt;span class="cm"&gt;/** Color palette consistency score (0-1) */&lt;/span&gt;
  &lt;span class="nx"&gt;colorHarmony&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;
  &lt;span class="cm"&gt;/** Motion direction alignment score (0-1) */&lt;/span&gt;
  &lt;span class="nx"&gt;motionContinuity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;
  &lt;span class="cm"&gt;/** Spatial composition stability score (0-1) */&lt;/span&gt;
  &lt;span class="nx"&gt;compositionStability&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;
  &lt;span class="cm"&gt;/** LLM-derived narrative coherence score (0-1) */&lt;/span&gt;
  &lt;span class="nx"&gt;narrativeCoherence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What I like about this shape is that every signal is visible, named, and separable. That makes the failure mode obvious too: if one signal is noisy, it does not deserve to dominate the entire score just because it arrived with confidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I combine signals before calibration
&lt;/h2&gt;

&lt;p&gt;The non-obvious part is that the composite score cannot assume every signal is present. In the calibrator, null values are skipped and the weights are renormalized over the remaining signals. That guardrail keeps a missing or unavailable signal from pretending it should still count.&lt;/p&gt;

&lt;p&gt;Here is the core scoring helper from the calibrator.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;computeComposite&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;weights&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;RewardWeights&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;signals&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;RewardSignals&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;totalWeight&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;weightedSum&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;SIGNAL_KEYS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;signal&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;signals&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;signal&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;signal&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;w&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;weights&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
      &lt;span class="nx"&gt;weightedSum&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;w&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;signal&lt;/span&gt;
      &lt;span class="nx"&gt;totalWeight&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;w&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;totalWeight&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;weightedSum&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nx"&gt;totalWeight&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That renormalization is the quiet guardrail in the whole system. If I had simply multiplied every signal by its weight and summed the result, a missing signal would have behaved like a silent penalty. Instead, the score is computed from the signals that actually exist, and the denominator follows the evidence.&lt;/p&gt;

&lt;p&gt;The calibrator's own config spells out the bounds the search runs under.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;DEFAULT_CONFIG&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Required&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;CalibratorConfig&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;gridSteps&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;refinementIterations&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;refinementStep&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.02&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;minWeight&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.05&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;maxWeight&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.50&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;normalizeWeights&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;holdoutFraction&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those bounds matter because they keep one signal from swallowing the rest during the search. The calibrator is allowed to explore, but not to wander into a regime where a single weight becomes the whole story.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the search loop actually does
&lt;/h2&gt;

&lt;p&gt;The calibrator uses two optimization strategies: a grid search followed by coordinate descent. The first is broad and exhaustive over a discretized weight space. The second starts from the best grid point and nudges one dimension at a time until the score stops improving.&lt;/p&gt;

&lt;p&gt;That structure is easier to trust than a single clever optimizer because it gives me a baseline and a refinement pass. I can see whether the coarse search found a decent region before I ask the local search to polish it.&lt;/p&gt;

&lt;p&gt;I wanted the calibrator to feel more like a lab bench than a black box: try a weight set, score it against the samples, keep what wins, move on. That makes the search legible when I inspect it later, which is exactly when these systems usually try to become mysterious.&lt;/p&gt;

&lt;p&gt;The calibrator's result type shows that I care about more than just the final weights.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;CalibrationResult&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="cm"&gt;/** Optimized weights */&lt;/span&gt;
  &lt;span class="nl"&gt;weights&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;RewardWeights&lt;/span&gt;
  &lt;span class="cm"&gt;/** Accuracy on holdout set (if k-fold used) */&lt;/span&gt;
  &lt;span class="nx"&gt;accuracy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
  &lt;span class="cm"&gt;/** Correlation between predicted composite and actual quality */&lt;/span&gt;
  &lt;span class="nx"&gt;correlation&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
  &lt;span class="cm"&gt;/** Number of samples used */&lt;/span&gt;
  &lt;span class="nx"&gt;sampleCount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
  &lt;span class="cm"&gt;/** Number of weight configurations evaluated */&lt;/span&gt;
  &lt;span class="nx"&gt;configurationsEvaluated&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
  &lt;span class="cm"&gt;/** Phase 1 vs Phase 2 improvement */&lt;/span&gt;
  &lt;span class="nx"&gt;gridBestAccuracy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
  &lt;span class="nx"&gt;refinedAccuracy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;gridBestAccuracy&lt;/code&gt; and &lt;code&gt;refinedAccuracy&lt;/code&gt; are the before-and-after story — both hold the combined objective the calibrator actually optimizes, which weights accuracy at 70% and Pearson &lt;code&gt;correlation&lt;/code&gt; with the quality score at 30%. If the two numbers are close, refinement did not earn its keep; if they diverge, coordinate descent found a ridge the grid stepped past. &lt;code&gt;accuracy&lt;/code&gt; and &lt;code&gt;correlation&lt;/code&gt; on their own are the holdout-set read of the final weight set, computed after the train/validation split. &lt;code&gt;configurationsEvaluated&lt;/code&gt; is less glamorous but load-bearing for debugging: at &lt;code&gt;gridSteps: 5&lt;/code&gt; the Phase 1 alone walks 3,125 weight vectors, so the count tells me whether Phase 2 contributed fifty more candidates or five hundred — the difference between a run that bounced off its bounds immediately and one that actually polished the grid point. &lt;code&gt;sampleCount&lt;/code&gt; is the denominator everything else is calibrated against, and the calibrator short-circuits to defaults when it drops below ten.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is not hand-tuning
&lt;/h2&gt;

&lt;p&gt;Hand-tuning starts with a belief and tries to make the numbers agree with it. Calibration starts with data and asks which belief survives contact with the benchmark set.&lt;/p&gt;

&lt;p&gt;The calibrator operates on logged samples from real pipeline runs. Each sample carries the raw reward signals, a quality score, and whether the candidate was accepted. That means the search is not guessing in a vacuum; it is optimizing against examples where the pipeline already revealed something about what "good" looked like.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;CalibrationSample&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="cm"&gt;/** Raw reward signals for this candidate */&lt;/span&gt;
  &lt;span class="nl"&gt;signals&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;RewardSignals&lt;/span&gt;
  &lt;span class="cm"&gt;/** Human or automated quality assessment (0-1) */&lt;/span&gt;
  &lt;span class="nx"&gt;qualityScore&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
  &lt;span class="cm"&gt;/** Was this candidate selected/accepted? */&lt;/span&gt;
  &lt;span class="nx"&gt;accepted&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt;
  &lt;span class="cm"&gt;/** Optional metadata for filtering */&lt;/span&gt;
  &lt;span class="nx"&gt;sceneType&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;
  &lt;span class="nx"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important limitation here is that calibration is only as honest as the samples feeding it. If the benchmark data is narrow, the search can still find a polished lie. So the calibrator is not a replacement for judgment; it is a way to make the judgment function answer to evidence instead of habit.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Phase 1 already told me
&lt;/h2&gt;

&lt;p&gt;The weights in production today — &lt;code&gt;visualDrift: 0.30&lt;/code&gt;, &lt;code&gt;colorHarmony: 0.25&lt;/code&gt;, &lt;code&gt;motionContinuity: 0.15&lt;/code&gt;, &lt;code&gt;compositionStability: 0.15&lt;/code&gt;, &lt;code&gt;narrativeCoherence: 0.15&lt;/code&gt; — came from the Phase 1 baseline run on 2026-02-21 across 100 contracts, and formal hyperparameter search was deferred in that phase while the HITL gate was calibrated first. But Phase 1 surfaced exactly the kind of signal evidence the calibrator is built to reconcile. On the 69 GPU-rendered contracts, &lt;code&gt;visualDrift&lt;/code&gt; averaged 0.529 and drove zero HITL flags, while &lt;code&gt;narrativeCoherence&lt;/code&gt; averaged 0.317 and drove 65.2% of them. &lt;code&gt;motionContinuity&lt;/code&gt; came in at 0.387 with another 17.4%. The two signals sitting at weight 0.15 together accounted for over 80% of the failure signal, while the signal carrying the largest weight at 0.30 silently passed everything.&lt;/p&gt;

&lt;p&gt;That is the distribution the calibrator exists to rebalance. A weight vector that puts its largest share on the signal doing no discrimination is not a tuned instrument — it is a ritual. The &lt;code&gt;correlation&lt;/code&gt; field is exactly the statistic that would catch it: Pearson correlation between the composite score and the logged quality score, which punishes weight vectors that let a flat signal dominate.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the baseline comparison keeps me honest
&lt;/h2&gt;

&lt;p&gt;The sharpest part of the design is the before-and-after comparison. The calibrator does not just report a weight vector and call it done. It compares the refined result against the grid-search baseline, which makes the improvement legible instead of assumed.&lt;/p&gt;

&lt;p&gt;That matters because search algorithms can easily produce movement without progress. A local refinement step can feel productive while actually circling the same hill. By keeping the baseline in view, I can tell whether the search surface rewarded the new weights or merely entertained the optimizer.&lt;/p&gt;

&lt;p&gt;The reward calibrator is also explicit about its two phases in the code comments: coarse grid search first, then coordinate descent around the best grid point. That sequence is not an accident. The grid gives me coverage; the refinement gives me precision. If I skipped the grid, I would be polishing a guess. If I skipped the refinement, I would be leaving useful accuracy on the table.&lt;/p&gt;

&lt;p&gt;The guardrails are doing real work here. The minimum and maximum weight bounds keep the search from overcommitting to one signal. The normalization option keeps the total scale stable. And the null-skipping composite score means missing evidence does not masquerade as a negative vote.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architecture Overview
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  subgraph signalLayer[Signal Layer]
    visualDrift[Visual Drift Signal]
    colorHarmony[Color Harmony Signal]
    motionContinuity[Motion Continuity Signal]
    compositionStability[Composition Stability Signal]
    narrativeCoherence[Narrative Coherence Signal]
  end

  subgraph scoringLayer[Reward Scoring Layer]
    signalNormalizer[Signal Normalizer]
    rewardMixer[Reward Mixer]
    compositeScore[Composite Score]
  end

  subgraph calibrationLayer[Calibration Layer]
    candidateWeights[Candidate Weight Sets]
    rewardCalibrator[Reward Calibrator]
    benchmarkData[Benchmark Data]
    baselineModel[Baseline Model]
    scoreComparator[Score Comparator]
    bestWeights[Best Weight Set]
  end

  visualDrift --&amp;gt; signalNormalizer
  colorHarmony --&amp;gt; signalNormalizer
  motionContinuity --&amp;gt; signalNormalizer
  compositionStability --&amp;gt; signalNormalizer
  narrativeCoherence --&amp;gt; signalNormalizer

  signalNormalizer --&amp;gt; rewardMixer
  candidateWeights --&amp;gt; rewardMixer
  rewardMixer --&amp;gt; compositeScore

  benchmarkData --&amp;gt; rewardCalibrator
  baselineModel --&amp;gt; scoreComparator
  compositeScore --&amp;gt; scoreComparator
  candidateWeights -.-&amp;gt; rewardCalibrator
  rewardCalibrator --&amp;gt; scoreComparator
  scoreComparator --&amp;gt; bestWeights
  bestWeights ==&amp;gt; candidateWeights
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Where the calibrator fits in the larger scoring stack
&lt;/h2&gt;

&lt;p&gt;The reward calibrator extends the evaluation layer that already handles quality-predictor weights, but it targets the reward weights used by the multi-signal reward mixer. That separation is what makes the architecture clean: the mixer defines how scores are combined, and the calibrator searches for better parameters for that combination.&lt;/p&gt;

&lt;p&gt;In practice, that means I can improve the judgment function without rewriting the signals themselves. If the system learns that composition stability should matter a little more than motion continuity on the benchmark corpus, the calibrator can discover that. If the opposite is true, it can discover that too. I do not need to hard-code a philosophical stance into the scorer.&lt;/p&gt;

&lt;p&gt;That is the real trick. The system is not merely deciding what threshold to cross. It is learning how to weigh its own senses.&lt;/p&gt;

&lt;p&gt;The result is a scoring layer that feels less like a fixed rulebook and more like a tuned instrument: bounded, testable, and willing to admit when one note has been playing too loudly for too long.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>typescript</category>
      <category>mlsystems</category>
      <category>scoring</category>
      <category>optimization</category>
    </item>
    <item>
      <title>Small in Code. Large in Behavior.</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:01:48 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/small-in-code-large-in-behavior-5hk7</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/small-in-code-large-in-behavior-5hk7</guid>
      <description>&lt;p&gt;Neuroloq's attachment pipeline turns on one boundary that comes before transcription: whether to treat an upload as document-like at all, and only then which OCR mode should read it. The shape of that decision is the product.&lt;/p&gt;

&lt;p&gt;The thing Neuroloq is optimizing for is cheap, deterministic, searchable text recovery — not maximum accuracy on every messy upload. If the goal were maximum accuracy on hostile inputs, the right move would be a vision-language model handling routing and transcription in one pass. Neuroloq does not do that, because predictable transcripts that survive search a week later matter more than the marginal accuracy bought with unpredictable latency and per-attachment bills. That tradeoff shapes every decision below.&lt;/p&gt;

&lt;p&gt;The interesting move is also smaller than it sounds. Selecting an OCR mode ahead of transcription is standard in any OCR pipeline. The earlier branch — deciding whether an attachment deserves OCR at all — is where the wins and the failures both live.&lt;/p&gt;

&lt;h2&gt;
  
  
  What broke first
&lt;/h2&gt;

&lt;p&gt;The earliest version trusted the filename, sent every image to a single vision pass, and stored the description back as if it were a transcript. Generic image analysis is good at describing what a frame contains. It is not good at preserving the text surface a learner expects to search later.&lt;/p&gt;

&lt;p&gt;Two failures came out of this. Screenshots of code and terminal output were interpreted as visual scenes instead of text-bearing material. Handwritten material was fed through the same assumptions as typed text, producing transcripts that looked plausible enough to ship and noisy enough to be useless later.&lt;/p&gt;

&lt;p&gt;The fix was to classify the attachment first, then pick an OCR mode only when the upload behaves like a document.&lt;/p&gt;

&lt;h2&gt;
  
  
  The split that mattered
&lt;/h2&gt;

&lt;p&gt;The split lives in &lt;code&gt;classifyAttachment&lt;/code&gt; and &lt;code&gt;detectOcrMode&lt;/code&gt;. The classifier answers whether the attachment deserves OCR. The mode detector answers how to OCR it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;AttachmentKind&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;document_like&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;freeform&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;OcrMode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;printed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;handwritten&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;AttachmentInput&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;userText&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;DOCUMENT_KEYWORDS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;screenshot&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;code&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;doc&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;scan&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pdf&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;note&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;text&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;terminal&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;log&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;error&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;config&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;output&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;whiteboard&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;handwritten&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;notebook&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;marker&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;written&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;board&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;OCR_MODE_HINTS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;handwritten&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;whiteboard&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;board&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;marker&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;handwritten&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;my writing&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;my notes&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;printed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;screenshot&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;code&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pdf&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;scan&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;terminal&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;log&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;error&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;config&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;output&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;classifyAttachment&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;AttachmentInput&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;AttachmentKind&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;haystack&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;userText&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DOCUMENT_KEYWORDS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;keyword&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;haystack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;keyword&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;hit&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;document_like&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;freeform&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;detectOcrMode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;AttachmentInput&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;OcrMode&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;haystack&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;userText&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;handwrittenHit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;OCR_MODE_HINTS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;handwritten&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;keyword&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;haystack&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;keyword&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;handwrittenHit&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;handwritten&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;printed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a v1 heuristic. The classifier only sees the filename and the user's caption, which means a learner who uploads &lt;code&gt;IMG_2847.png&lt;/code&gt; of a terminal pane and types "help with this" falls straight through to freeform vision. The next iteration needs at least one image-side signal — aspect ratio, EXIF, a cheap zero-shot pass on the pixels — to close that hole. Until then, the classifier reads "did the user or the OS hint that this is text-bearing?", not "is this actually text-bearing?"&lt;/p&gt;

&lt;p&gt;The two functions are also not really two questions. Every keyword in &lt;code&gt;OCR_MODE_HINTS.printed&lt;/code&gt; already appears in &lt;code&gt;DOCUMENT_KEYWORDS&lt;/code&gt;. By the time &lt;code&gt;detectOcrMode&lt;/code&gt; runs, the input is known to be document-like, so the only question being asked is whether any handwritten keyword fired. Printed is the residual. Handwritten wins ties on purpose: misrouting a handwritten note through printed TrOCR produces confident nonsense, while the reverse produces messy but recoverable text.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  U[Upload attachment] --&amp;gt; C[classifyAttachment]
  C --&amp;gt;|document_like| M[detectOcrMode]
  C --&amp;gt;|freeform| V[GPT-5.2 vision analysis]
  M --&amp;gt;|printed| P[TrOCR printed]
  M --&amp;gt;|handwritten| H[TrOCR handwritten]
  P --&amp;gt; S[GPT-5.2 text-only synthesis]
  H --&amp;gt; S
  S --&amp;gt; PERSIST[Persist searchable text]
  V --&amp;gt; PERSIST
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The decision tree is the product. The rest of the implementation is just code around that tree.&lt;/p&gt;

&lt;h2&gt;
  
  
  The OCR client is thin on purpose
&lt;/h2&gt;

&lt;p&gt;The OCR client does not invent policy. It receives an image URL and an OCR mode, fetches the image, runs the matching Hugging Face TrOCR model, and returns the extracted text with metadata.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;HfInference&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@huggingface/inference&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;OcrMode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;printed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;handwritten&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;OCR_MODELS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;Record&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;OcrMode&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;printed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;microsoft/trocr-base-printed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;handwritten&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;microsoft/trocr-large-handwritten&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;OCR_PROVIDER&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;huggingface&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;OcrResult&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nl"&gt;extractedText&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;ocrProvider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;ocrModel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;ocrMode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;OcrMode&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;runOcr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;imageUrl&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;OcrMode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;printed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;OcrResult&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;apiKey&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;HUGGING_FACE_API_KEY&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;HUGGING_FACE_API_KEY is not set — cannot run OCR&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;imageUrl&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Failed to fetch image for OCR (HTTP &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;)`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;imageBlob&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;blob&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hf&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;HfInference&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;OCR_MODELS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;hf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;imageToText&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;imageBlob&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;generated_text&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;score&lt;/span&gt;&lt;span class="p"&gt;?:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;extractedText&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;generated_text&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="na"&gt;ocrProvider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;OCR_PROVIDER&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;ocrModel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;ocrMode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result shape carries provider, model, and mode together so transcripts stay traceable. If a transcript looks odd, the stored mode tells me immediately whether the wrong model ran or the source image was the problem.&lt;/p&gt;

&lt;p&gt;The model choice deserves a caveat. &lt;code&gt;microsoft/trocr-base-printed&lt;/code&gt; is a single-line transformer; it struggles with multi-column code, deep indentation, and any spatial structure that needs preserving. Tesseract with layout analysis, PaddleOCR, or a vision-language model with a structured prompt would probably do better on that class. I picked TrOCR for its latency and cost profile, not on a benchmark — that benchmark is on the list, and the model is the part of this pipeline with the least evidence behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why GPT-5.2 sits after OCR
&lt;/h2&gt;

&lt;p&gt;When OCR succeeds, Neuroloq passes the raw transcript to GPT-5.2 in text-only mode. The pass cleans line breaks, trims noise, and reshapes the transcript into something stable enough to store and search. The image is not passed in again.&lt;/p&gt;

&lt;p&gt;This is a tradeoff, not a virtue. Text-only synthesis is cheaper, more deterministic, and easier to debug than a multimodal cleanup pass with the original image attached — but it also means TrOCR errors are permanent. If printed-mode TrOCR confuses a &lt;code&gt;0&lt;/code&gt; for an &lt;code&gt;O&lt;/code&gt; or drops the indentation on a Python block, the synthesis pass has no pixels left to recover from. A multimodal cleanup is strictly more informative; it just costs more and produces less predictable output. Reproducibility and cost are the reasons to keep this text-only, not discipline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the route has to stay explicit
&lt;/h2&gt;

&lt;p&gt;The hybrid case is the one the heuristic handles worst. Picture a notebook page with a printed derivation on the top half and a learner's handwritten step-by-step on the bottom. Handwritten-first routes the whole upload to handwritten TrOCR. That model reads the handwriting cleanly and turns the printed formula into characters that look almost right — &lt;code&gt;x²&lt;/code&gt; becomes &lt;code&gt;x2&lt;/code&gt;, the radical sign becomes a slash, line breaks land wherever the model felt strongest. The reverse default would mirror the failure on the bottom half: clean printed transcription, garbled handwriting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Neither default is correct.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Until segmentation happens before OCR — split the page, OCR each region with the right model — the system commits to one failure mode per attachment. Handwritten-first picks the mode that fails less catastrophically when wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The persistence layer follows the decision
&lt;/h2&gt;

&lt;p&gt;Storage stays downstream of the analysis decision. Extracted text, provider, model, and OCR mode travel with the message record so the tutor thread stays searchable and explainable. If GPT-5.2 normalized the OCR output, that version gets indexed. If vision handled the attachment, the descriptive text from that path gets indexed. The interesting work happens earlier in the pipeline; persistence is straightforward once the route has decided.&lt;/p&gt;

&lt;h2&gt;
  
  
  What stays right—and what is still soft
&lt;/h2&gt;

&lt;p&gt;The improvement was not a broader claim about OCR. It was a tighter boundary around the route that handles document-like attachments — and an honest accounting of where that boundary is still soft. The classifier is a substring check that misses pixel-only signals. The synthesis pass trades accuracy for determinism. TrOCR is a constraint, not a verdict. The hybrid case still picks its failure mode rather than solving it.&lt;/p&gt;

&lt;p&gt;What stays right, even with all of that, is the decision the route enforces: not whether an image can be analyzed, but whether it can be read correctly enough to survive search, review, and a return visit days later.&lt;/p&gt;

&lt;p&gt;Small in code. Large in behavior.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ocr</category>
      <category>typescript</category>
      <category>huggingface</category>
      <category>neuroloq</category>
    </item>
    <item>
      <title>Model Failure Is a Time Series</title>
      <dc:creator>Daniel Romitelli</dc:creator>
      <pubDate>Mon, 27 Jul 2026 16:01:24 +0000</pubDate>
      <link>https://dev.to/daniel_romitelli_44e77dc6/model-failure-is-a-time-series-275e</link>
      <guid>https://dev.to/daniel_romitelli_44e77dc6/model-failure-is-a-time-series-275e</guid>
      <description>&lt;p&gt;A quiet model is easy to trust for the wrong reason. The failure mode that forced my hand was not a dramatic crash or an exception in the inference server. It was worse: the primary predictor could keep emitting ordinary-looking forecasts while the property I actually needed for execution — whether the next prediction would be correct enough to act on — was changing underneath it.&lt;/p&gt;

&lt;p&gt;That is why I built CTF as a separate reliability forecaster in the Pramaana crypto research stack. Not another directional signal. Not a committee. Not a narrative layer around the model. CTF is a lightweight inference helper in &lt;code&gt;inference/ctf_predictor.py&lt;/code&gt; that turns recent uncertainty behavior into a fixed reliability feature vector and predicts the probability that the next primary-model call deserves trust.&lt;/p&gt;

&lt;p&gt;The transferable idea is the title taken literally: model failure is a time series. Reliability is not just a scalar attached to one row. If the main predictor has an uncertainty trail — entropy, confidence, calibration width, agreement, drift, recent hit rate — then trust is itself a time series. CTF exists because I wanted the execution layer to stop treating confidence as a one-frame photograph and start treating it as telemetry.&lt;/p&gt;

&lt;p&gt;This post is not about adding yet another confidence threshold. I had already built several forms of gating, calibration, and meta-evaluation into the system: conformal intervals, per-asset calibration, cost-baked EV checks, and post-deployment filter audits. CTF is a narrower contribution. It models the evolution of uncertainty behavior as a supervised reliability target. The primary model answers, “What happens next?” CTF answers, “Is this model currently in a state where its next answer is likely to be right?”&lt;/p&gt;

&lt;h2&gt;
  
  
  Predict correctness, not price
&lt;/h2&gt;

&lt;p&gt;The naive version of uncertainty handling is to read one confidence value and make a decision from it. High confidence means act. Low confidence means stay flat. That is clean, but it is too shallow for a regime-shifting market.&lt;/p&gt;

&lt;p&gt;A single uncertainty snapshot tells me how the model feels about the current input. It does not tell me whether confidence has been deteriorating for the last hour, whether entropy has become unstable across assets, whether conformal width is expanding, whether the model has been right recently, or whether the current confidence value is abnormal relative to its own local history.&lt;/p&gt;

&lt;p&gt;Those are not properties of one prediction. They are properties of a sequence.&lt;/p&gt;

&lt;p&gt;CTF is built around that distinction. The primary model emits a forecast and uncertainty telemetry. CTF consumes a bounded recent-history window and predicts correctness of the next forecast. That keeps the tasks separate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the primary predictor models market path behavior;&lt;/li&gt;
&lt;li&gt;CTF models the primary predictor’s current reliability state;&lt;/li&gt;
&lt;li&gt;the execution layer consumes both through a cost-aware gate.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A price predictor tries to model the market. A correctness predictor models the predictor.&lt;/p&gt;

&lt;p&gt;The analogy I keep coming back to is a racing engine. The primary model is the engine producing torque. CTF is not a second engine bolted to the hood; it is the telemetry system watching temperature, vibration, and pressure over the last few laps. You do not add telemetry because the engine never works. You add it because engines often sound fine right before they stop being fine.&lt;/p&gt;

&lt;p&gt;Here is the shape of the layer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flowchart TD
  marketInput[Market features] --&amp;gt; primaryModel[Primary predictor]
  primaryModel --&amp;gt; primaryOutput[Prediction plus uncertainty telemetry]
  primaryOutput --&amp;gt; historyBuffer[Recent history buffer]
  historyBuffer --&amp;gt; featureVector[Windowed reliability features]
  featureVector --&amp;gt; ctfModel[CTF correctness predictor]
  ctfModel --&amp;gt; trustProbability[Probability next prediction is correct]
  trustProbability --&amp;gt; gate[Cost-baked execution gate]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important transformation happens in the middle. The system stops asking only, “What did the model just say?” and starts asking, “How has the model been behaving?”&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is different from a confidence gate
&lt;/h2&gt;

&lt;p&gt;A confidence gate is memoryless unless you explicitly give it memory. It takes the current output, compares it to a threshold, and either passes or blocks the trade. That can be useful, but it collapses the reliability problem into a single coordinate.&lt;/p&gt;

&lt;p&gt;CTF is not that. CTF is a supervised model trained on windowed behavior. It can learn that the same current confidence value means different things depending on the preceding trajectory.&lt;/p&gt;

&lt;p&gt;For example, suppose the primary predictor emits a confidence value of 0.58. A simple gate sees 0.58 and applies one rule. CTF sees the surrounding state:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;confidence was 0.74, 0.69, 0.63, then 0.58 over the last few observations;&lt;/li&gt;
&lt;li&gt;attention entropy has been rising;&lt;/li&gt;
&lt;li&gt;cross-asset dispersion has widened;&lt;/li&gt;
&lt;li&gt;recent correctness has fallen below the local baseline;&lt;/li&gt;
&lt;li&gt;conformal width is expanding for the same assets that now produce trade candidates.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is a different situation from a stable 0.58 in a low-dispersion regime where the model has been consistently right. The current scalar is identical. The reliability state is not.&lt;/p&gt;

&lt;p&gt;This is the novelty that matters. The model is not “more cautious” because a human wrote a more conservative if-statement. It is trained to map temporal uncertainty telemetry to a correctness probability.&lt;/p&gt;

&lt;p&gt;That distinction also keeps CTF separate from the older XGBoost meta-filter work I ran before. The post-deployment audit of that filter was painful but useful: the apparent win-rate lift was not statistically solid, and the filter selected a fatter-tailed loss distribution. A meta-filter that passes trades based on static signal descriptors can look good in aggregate while quietly changing the tail profile of the trades it allows through.&lt;/p&gt;

&lt;p&gt;CTF is designed around the lesson from that failure. It does not try to be a second trading strategy hidden behind a pass/fail switch. It predicts whether the primary model’s next answer is likely to be correct, using a time-indexed reliability state. That is a smaller job and a cleaner supervised target.&lt;/p&gt;

&lt;h2&gt;
  
  
  The supervised target: correctness at the next prediction
&lt;/h2&gt;

&lt;p&gt;The target for CTF is not future return. It is not realized P&amp;amp;L. It is not “would this trade have made money after fees?” Those are execution outcomes, and they mix model quality with sizing, spread, slippage, funding, stops, and path-dependent order handling.&lt;/p&gt;

&lt;p&gt;The CTF label is correctness of the primary prediction at the prediction horizon. For a directional head, that means the predicted direction matches the realized direction over the same horizon. For a path-passage formulation, it means the predicted passage event matches the realized barrier event the execution layer evaluates. The important rule is alignment: the correctness label must be defined against the same event the primary model claimed to predict.&lt;/p&gt;

&lt;p&gt;That sounds obvious, but it is where a lot of meta-models go bad. If the primary model predicts a 60-minute direction and the reliability model is labeled against post-cost trade profitability, the meta-model is no longer measuring predictor correctness. It is measuring a mixture of predictor correctness and execution mechanics. Sometimes that is useful, but it is not CTF.&lt;/p&gt;

&lt;p&gt;The CTF label is deliberately close to the primary model’s semantic contract:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;at time &lt;code&gt;t&lt;/code&gt;, the primary model emits a prediction for horizon &lt;code&gt;h&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;the CTF feature vector is built only from telemetry available at or before &lt;code&gt;t&lt;/code&gt;;&lt;/li&gt;
&lt;li&gt;when the future outcome at &lt;code&gt;t + h&lt;/code&gt; is known, the row receives a binary correctness label;&lt;/li&gt;
&lt;li&gt;that label becomes the supervised target for the reliability model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The no-leakage rule is strict. The CTF row at time &lt;code&gt;t&lt;/code&gt; cannot include the correctness of the prediction made at &lt;code&gt;t&lt;/code&gt;, because that outcome is not known yet. It can include rolling correctness for earlier predictions whose horizons have already resolved. If the horizon is one hour, the most recent eligible correctness observation is from a prediction at or before &lt;code&gt;t - h&lt;/code&gt;, not from the prediction being scored now.&lt;/p&gt;

&lt;p&gt;That one detail is the difference between a reliability model and a disguised lookahead bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  The history buffer is the boundary between models
&lt;/h2&gt;

&lt;p&gt;A recent-history window sounds mundane until it becomes the contract between the primary predictor and the reliability predictor.&lt;/p&gt;

&lt;p&gt;The primary model produces a stream of values: prediction, confidence, entropy, conformal width, asset, timestamp, and eventually realized correctness. CTF does not consume the raw stream as an unbounded sequence. The helper maintains a bounded buffer and compresses recent behavior into a fixed feature vector.&lt;/p&gt;

&lt;p&gt;That compression is intentional. I did not want the execution gate to depend on variable-length sequence handling or call-time improvisation. A fixed schema makes the reliability layer testable, serializable, and comparable between training and inference.&lt;/p&gt;

&lt;p&gt;The buffer holds only information that would have existed when a live decision was made:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;current primary-model telemetry for the candidate prediction;&lt;/li&gt;
&lt;li&gt;prior telemetry for the same asset or asset group;&lt;/li&gt;
&lt;li&gt;resolved correctness from older predictions;&lt;/li&gt;
&lt;li&gt;contemporaneous cross-asset uncertainty measurements;&lt;/li&gt;
&lt;li&gt;derived rolling statistics over the recent window.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It excludes anything that requires future knowledge:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;realized correctness of the current prediction;&lt;/li&gt;
&lt;li&gt;realized return over the current prediction horizon;&lt;/li&gt;
&lt;li&gt;post-trade P&amp;amp;L from a decision that has not completed;&lt;/li&gt;
&lt;li&gt;any feature recomputed with future rows accidentally included in a rolling window.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last class of bug is easy to miss. Pandas rolling operations, joins between prediction logs and realized outcomes, and cross-asset aggregates can all leak if the index semantics are sloppy. The implementation discipline is to treat CTF rows exactly like live inference snapshots. If the value was not known at the timestamp being scored, it does not belong in the feature vector.&lt;/p&gt;

&lt;p&gt;The fixed feature vector then captures several kinds of reliability evidence, each answering a separate question:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Telemetry signal&lt;/th&gt;
&lt;th&gt;Reliability question it answers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Level&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Where has uncertainty been sitting recently?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Change&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;How abruptly is uncertainty moving right now?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Trend&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Is that movement persistent, or just noise?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Dispersion&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Is the uncertainty isolated to one asset, or spreading across many?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Recent accuracy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Have resolved prior predictions actually been right?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Calibration state&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Is interval width or calibrated probability shifting?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Drift&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Is the current telemetry distribution departing from its recent baseline?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Means, deltas, trends, dispersion, and rolling hit rates are not decorative statistics. Each answers a separate reliability question. The level tells me whether the model is operating in a generally uncertain state. The delta tells me whether the state just changed. The trend tells me whether the change is noise or a persistent move. Dispersion tells me whether the issue is asset-local or system-wide. Rolling correctness tells me whether the model’s recent self-assessment has been earning trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attention entropy is telemetry, not a verdict
&lt;/h2&gt;

&lt;p&gt;In the FT-Transformer path, attention entropy is one of the most useful telemetry signals because it is closer to the model’s internal allocation of attention than a final probability alone. A final softmax collapses the whole forward pass into one number over the outputs; attention entropy instead measures how evenly the model spread its focus across input features on the way there, so it can expose the model hedging across many weak cues or fixating on a single brittle one even when the output probability looks unchanged. But I do not treat entropy as a magic score.&lt;/p&gt;

&lt;p&gt;A single entropy value can be misleading for the same reason a single confidence value can be misleading. High entropy might mean the model is genuinely uncertain. It might also be normal for a specific asset, horizon, or volatility regime. Low entropy might mean the model has found a clean structure. It might also mean the model has collapsed onto a brittle shortcut.&lt;/p&gt;

&lt;p&gt;The trajectory is the point. CTF watches how entropy behaves over time and relative to the local population of assets. Rising entropy across the universe is different from rising entropy in one thin asset. A sudden entropy jump after a long stable period is different from an asset whose baseline is noisy. A high-entropy state with improving resolved correctness is different from a high-entropy state with deteriorating correctness.&lt;/p&gt;

&lt;p&gt;That is why I prefer to call these inputs telemetry rather than explanations. They are measurements from the model while it works. CTF learns which telemetry patterns have historically preceded correct or incorrect primary-model calls.&lt;/p&gt;

&lt;p&gt;The same logic applies to conformal width. Wider intervals often mean less precise forecasts, but width by itself is not the decision. What matters is whether width is expanding, whether it is expanding faster than usual, whether it is expanding across correlated assets, and whether the model has recently remained correct under similar width dynamics.&lt;/p&gt;

&lt;p&gt;The feature vector gives the reliability model those comparisons without turning the primary predictor into a monolith that must simultaneously forecast price, calibrate itself, diagnose drift, and decide execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Train/serve parity is a first-class constraint
&lt;/h2&gt;

&lt;p&gt;A reliability layer can fail even when the concept is right if the training features and inference features drift apart. I had already seen this class of problem in the earlier meta-filter work, where train/serve skew around &lt;code&gt;kelly_fraction&lt;/code&gt; was a likely contributor to bad live behavior. CTF was built with that scar tissue in mind.&lt;/p&gt;

&lt;p&gt;The feature schema is treated as a contract. Training and inference use the same ordered feature list, the same definitions, the same window semantics, and the same missing-value policy. The CTF helper emits a fixed vector rather than a loose dictionary that downstream code can accidentally reorder.&lt;/p&gt;

&lt;p&gt;The important pieces of the contract are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;feature names are explicit and versioned with the model artifact;&lt;/li&gt;
&lt;li&gt;column order is preserved at serialization time;&lt;/li&gt;
&lt;li&gt;training rows are generated by replaying historical predictions as if they were live;&lt;/li&gt;
&lt;li&gt;inference rows are generated by the same feature builder, not a hand-written approximation;&lt;/li&gt;
&lt;li&gt;missing values from insufficient warm-up history are handled the same way in both paths;&lt;/li&gt;
&lt;li&gt;assets that lack enough resolved history are either scored conservatively or withheld until the buffer is warm.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Warm-up behavior deserves special attention. If the model needs a 32-observation window, the first few rows for an asset cannot pretend to have a full history. There are only three honest choices: do not score until the buffer is warm, use features that explicitly encode the short history length, or route the candidate through a conservative fallback. What I avoid is silently filling the window with values that make the first live rows look cleaner than they are.&lt;/p&gt;

&lt;p&gt;Cross-asset features have their own parity trap. In training, it is tempting to compute dispersion across all rows in a timestamp bucket after the full dataset has been assembled. In live inference, only the assets scored at that time are available, and some may be missing due to exchange, ingest, or latency issues. The CTF feature builder has to make that mismatch explicit. The dispersion feature cannot depend on a perfect historical panel if live inference will see an imperfect one.&lt;/p&gt;

&lt;p&gt;This is why I think of the history buffer as a boundary, not just a container. It enforces what the CTF model is allowed to know.&lt;/p&gt;

&lt;h2&gt;
  
  
  Validation before the trust probability reaches execution
&lt;/h2&gt;

&lt;p&gt;A probability is useful only if it behaves like a probability. Before CTF can influence a gate, I need to know whether its scores are ordered correctly and whether their magnitudes are calibrated well enough to consume.&lt;/p&gt;

&lt;p&gt;The first validation pass is ranking: when CTF assigns higher trust probabilities, does realized correctness actually rise? I check this by binning predictions into score buckets and comparing empirical correctness across buckets. If the model cannot rank reliability states, it has no business gating trades.&lt;/p&gt;

&lt;p&gt;The second pass is calibration: when CTF emits 0.70, does that bucket land near 70% correctness, after accounting for sample size and regime splits? Perfect calibration is not realistic in a drifting market, but gross miscalibration is dangerous because the execution gate may treat the score as composable with expected value, risk, or size.&lt;/p&gt;

&lt;p&gt;The third pass is regime stratification. A reliability model that only works during one volatility regime is not a reliability model; it is a regime artifact. I care about performance across realized-volatility buckets, asset tiers, time folds, and market phases. This is especially important in crypto, where a one-year window can contain calm bearish drift, violent bullish expansion, and corrective chop. A model that passes on aggregate can still fail exactly when the system most needs it.&lt;/p&gt;

&lt;p&gt;The fourth pass is decision impact. CTF is not validated only by AUC or log loss. I also inspect what the gate would have done differently:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;which candidates would have been blocked;&lt;/li&gt;
&lt;li&gt;whether blocked candidates were actually lower quality;&lt;/li&gt;
&lt;li&gt;whether the model changes the tail of accepted losses;&lt;/li&gt;
&lt;li&gt;whether it reduces participation in a way that destroys opportunity;&lt;/li&gt;
&lt;li&gt;whether it creates asset concentration by vetoing some names more than others.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point came directly from the meta-filter postmortem. A filter can improve the headline win rate and still pass worse losses. For CTF, I inspect the distribution of allowed and vetoed outcomes, not only the average.&lt;/p&gt;

&lt;p&gt;Only after those checks does the trust probability become eligible for the execution path.&lt;/p&gt;

&lt;p&gt;The bar I hold it to in shadow mode is narrow and specific. The vetoes should cluster in exactly the windows where the primary model later proves unreliable — rising cross-asset entropy, widening conformal intervals, rolling correctness slipping below its local baseline — and the gate should stay out of the way when the model is stable and recently right. The goal was never a headline win-rate lift; it was making the veto fire for a reason I can name, in the degradation regimes the primary model is already known to fall into, while leaving the high-quality candidates untouched.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the gate consumes CTF
&lt;/h2&gt;

&lt;p&gt;The live trading philosophy in this stack is intentionally narrow. The path-passage strategy is governed by a cost-baked expected-value gate. The older multi-agent advisory path was removed from live decisioning; I kept the system pointed at one operational question: does this candidate clear the execution rule after costs and risk controls?&lt;/p&gt;

&lt;p&gt;CTF fits that philosophy because it emits a number the gate can consume. It does not create a debate. It does not explain the market. It does not vote with other agents. It estimates the probability that the primary model’s next prediction is correct.&lt;/p&gt;

&lt;p&gt;There are two clean ways for the gate to consume that score.&lt;/p&gt;

&lt;p&gt;The first is a hard reliability floor: if &lt;code&gt;ctf_confidence&lt;/code&gt; is below the configured threshold, the candidate is vetoed regardless of the primary model’s directional confidence. This is the most conservative integration and the easiest to reason about operationally.&lt;/p&gt;

&lt;p&gt;The second is EV adjustment: the trust probability modifies the expected value calculation or position eligibility without replacing the rest of the cost model. That route is more expressive, but it requires stronger calibration because the probability is being treated as a numerical ingredient rather than a pass/fail guard.&lt;/p&gt;

&lt;p&gt;In both cases, CTF remains subordinate to the execution rule. It does not say “buy” or “sell.” It says, “The predictor is currently in a reliability state where its next answer is or is not worth using.”&lt;/p&gt;

&lt;p&gt;That separation matters. A reliability model should not quietly become a shadow strategy. If it starts making directional decisions, the label and validation setup must change. CTF stays focused on correctness probability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a gate, not a council
&lt;/h2&gt;

&lt;p&gt;I used to route decisions through a retired multi-agent advisory console. The lesson from retiring it was simple: a narrow probabilistic gate beats broad advisory debate when the operational question is just &lt;em&gt;“does this candidate clear the EV rule after costs?”&lt;/em&gt; — so CTF lives next to inference and gating, emitting one calibrated reliability estimate instead of arguments in a dashboard panel.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure modes I watch for
&lt;/h2&gt;

&lt;p&gt;CTF is not an escape hatch from model risk. It is another model, trained on historical relationships between uncertainty telemetry and correctness. If that relationship breaks, CTF can be wrong too.&lt;/p&gt;

&lt;p&gt;The difference is that its failures are more diagnosable than a monolithic predictor’s failures. If the primary model degrades and CTF remains overconfident, I know the reliability model is missing a failure signature. If CTF becomes too conservative while the primary model remains useful, I know the reliability model is overreacting to telemetry patterns that no longer imply failure. If CTF works on BTC and ETH but fails on smaller assets, I know the cross-asset or asset-tier behavior needs separate treatment.&lt;/p&gt;

&lt;p&gt;The main failure modes are predictable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;leakage in the correctness label or rolling features;&lt;/li&gt;
&lt;li&gt;train/serve skew in feature construction;&lt;/li&gt;
&lt;li&gt;overfitting to a narrow volatility regime;&lt;/li&gt;
&lt;li&gt;miscalibrated probability magnitudes;&lt;/li&gt;
&lt;li&gt;vetoes that improve accuracy but worsen tail exposure;&lt;/li&gt;
&lt;li&gt;excessive conservatism that blocks the few candidates with real edge;&lt;/li&gt;
&lt;li&gt;asset concentration caused by uneven telemetry quality.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I do not treat any of those as theoretical. They are exactly the kinds of problems that show up when a research result becomes an execution component.&lt;/p&gt;

&lt;p&gt;The defense is not faith in the model. The defense is a strict feature contract, walk-forward validation, calibration checks, regime stratification, and shadow evaluation before the score affects live decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The broader pattern
&lt;/h2&gt;

&lt;p&gt;The CTF idea generalizes beyond this crypto stack. Any system with a primary model that emits uncertainty over time can treat reliability as its own supervised problem.&lt;/p&gt;

&lt;p&gt;The ingredients are modest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a primary predictor with a defined forecast target;&lt;/li&gt;
&lt;li&gt;uncertainty telemetry emitted at inference time;&lt;/li&gt;
&lt;li&gt;a recent-history buffer with no lookahead;&lt;/li&gt;
&lt;li&gt;a fixed feature transform shared by training and serving;&lt;/li&gt;
&lt;li&gt;a correctness label aligned to the primary model’s target;&lt;/li&gt;
&lt;li&gt;calibration and regime validation for the reliability score;&lt;/li&gt;
&lt;li&gt;a downstream gate that knows how to consume trust probability.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The mistake is waiting for failures to become obvious in the business metric. By then, the evidence has already been paid for. If the model’s own uncertainty behavior contains early warning signs, the right move is to model those signs directly.&lt;/p&gt;

&lt;p&gt;That is why I think of CTF as a missing sense organ. The primary model sees the market. CTF watches the model seeing the market.&lt;/p&gt;

&lt;p&gt;A forecast is an answer. Reliability is a condition of using that answer. Treating those as separate inference problems made the system cleaner because the model no longer had to be trusted or distrusted all at once. It could be measured while it worked.&lt;/p&gt;

&lt;p&gt;The final design lesson is the one I keep applying across Pramaana: do not ask one model to carry every responsibility. Let the forecaster forecast. Let calibration quantify uncertainty. Let the execution gate account for costs. Let CTF estimate whether the forecaster is currently in a state where its answer should be admitted into that gate. The moment those responsibilities are separated, debugging becomes less mystical, validation becomes sharper, and failure becomes something I can model before it becomes a line item in P&amp;amp;L.&lt;/p&gt;




&lt;p&gt;🎧 &lt;strong&gt;Listen to the audiobook&lt;/strong&gt; — &lt;a href="https://open.spotify.com/show/4ABVd5yDVfbX9HlV5JjT7D" rel="noopener noreferrer"&gt;Spotify&lt;/a&gt; · &lt;a href="https://play.google.com/store/audiobooks/details/How_to_Architect_an_Enterprise_AI_System_And_Why_t?id=AQAAAECafz8_tM&amp;amp;hl=en" rel="noopener noreferrer"&gt;Google Play&lt;/a&gt; · &lt;a href="https://www.craftedbydaniel.com/audiobook" rel="noopener noreferrer"&gt;All platforms&lt;/a&gt;&lt;br&gt;
🎬 &lt;a href="https://youtube.com/playlist?list=PLRteDbGJPYDb9XNjecvHplGlgW7tIv_q6" rel="noopener noreferrer"&gt;Watch the visual overviews on YouTube&lt;/a&gt;&lt;br&gt;
📖 &lt;a href="https://www.craftedbydaniel.com/blog/series/how-to-architect-an-enterprise-ai-system-and-why-the-engineer-still-matters" rel="noopener noreferrer"&gt;Read the full 13-part series&lt;/a&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>timeseries</category>
      <category>modelreliability</category>
      <category>tradingsystems</category>
    </item>
  </channel>
</rss>
