DEV Community

Cover image for Bridging Written Policy and Past Decisions Through the Lens of Hindsight
Srimukhi Padamatinti
Srimukhi Padamatinti

Posted on

Bridging Written Policy and Past Decisions Through the Lens of Hindsight

Bridging Written Policy and Past Decisions Through the Lens of Hindsight

A cashless pre-authorisation packet can be “complete” according to a policy and still come back with a query because one date was on the wrong letterhead. I built PayerLine around that gap: the rules people learn from actual insurer replies are often more useful than the rules printed in the policy.

A desk that remembers outcomes

PayerLine helps hospital revenue-cycle staff decide whether to hold or send a cashless packet before it reaches an insurer. The desk captures a patient, procedure, package name, and document details, then returns a recommendation with missing items, documents to leave out, confidence, and the case history behind the advice. Staff remain responsible for sending the packet; the system does not choose treatment or submit anything to an insurer portal.

The patient file: today's admission details feed both the first page and the live check.

The case I use to explain the system is a laparoscopic cholecystectomy. The procedure stays the same, but the payer changes what matters. Meridian’s outcomes make the ultrasound date on hospital letterhead and the exact package name significant. Northline’s history emphasizes a current fitness certificate, labs, and a matching culture report. Harbour cares about signed consent and an itemised estimate. Sable asks for photo identity, a policy e-card, and a CBC.

That is not one checklist with four names on it. It is four bodies of operational history, and confusing them is a correctness failure. PayerLine gives each insurer its own Hindsight bank, so the review starts inside the selected insurer’s memory boundary. Hindsight’s agent memory system on GitHub and memory documentation describe the underlying platform; in this application, its value is that outcomes can be retained, recalled, and reflected on as an evolving case history rather than kept as a pile of static examples. The broader idea of agent memory for systems that need to use past experience maps directly to this workflow.

One insurer bank is selected; the other three stay closed.

The important boundary is the payer bank

The Express API receives the case and reviewAdmission resolves the payer before it calls memory. The review runs recall, reflection, and recent outcome loading concurrently, but each memory operation receives the same payer ID:

const [recalled, reflection, outcomes] = await Promise.all([
  recallCase(admission.payerId, query, admission.procedure),
  reflectCase(admission.payerId, query, admission.procedure),
  loadRecentOutcomes(admission.payerId, admission.procedure, admission.packageName),
]);
Enter fullscreen mode Exit fullscreen mode

This small piece of orchestration carries an architectural decision: payer selection is not merely a prompt hint. It chooses the bank passed to the Hindsight client. recallCase and reflectCase both resolve the payer’s bank ID before making their calls, and recording an outcome follows the same route. Procedure tags narrow the memory search further.

I prefer that explicit routing to a single shared memory with an instruction saying “ignore other insurers.” Natural-language instructions help a model reason, but they are not a tenant boundary. The application knows which payer is being reviewed, so it should choose the corresponding bank in code. A prompt can say to use only that history; it should not be the only thing preventing cross-payer contamination.

Memory should reason over evidence, not invent policy

A memory system can be confidently wrong if every anecdote becomes a rule. PayerLine provisions each bank with directives about the limits of its evidence. For example, one directive says not to invent a payer rule from a single case and to report thin confidence until at least two outcomes agree. Another asks the system to show both cases when outcomes disagree. Hindsight’s retain mission extracts document defects, package names, and outcome reasons, while its observations mission asks for durable insurer-level patterns and revision when newer evidence contradicts them.

Reflection also has a constrained response shape. The application asks for a hold/send decision, confidence, policy conflict, summary, fixes, documents not to add, and evidence. That makes the response usable in the desk UI and easier to validate than free-form prose. The system asks Hindsight to cite dated approvals, queries, or denials, so a billing specialist can inspect why a recommendation appeared.

This matters most when policy and practice disagree. Meridian’s written policy says any clinically reasonable package name is acceptable; the modeled outcomes say the exact procedure name has mattered. The review can surface that conflict and show the relevant history. Memory is not silently substituted for policy: the disagreement itself becomes part of what the user sees.

I did not leave the final checklist to reflection

Recall and reflection are useful for matching a new packet to prior cases, but several known requirements are straightforward predicates. Those belong in ordinary code. shared/packet.ts checks the current fields and constructs a stable set of missing documents and items to remove. For Meridian, for example:

case "meridian":
  if (!packageMatchesProcedure(admission)) {
    missingRequired.push(`Package name: ${procedure}`);
  }
  if (!admission.ultrasoundDateOnLetterhead) {
    missingRequired.push("Ultrasound on letterhead");
  }
  if (admission.cultureReportAttached && admission.cultureContradictsDiagnosis) {
    removeFromPacket.push("Conflicting culture report");
  }
  break;
Enter fullscreen mode Exit fullscreen mode

The final response is stabilized against that checklist:

return {
  ...decision,
  decision: checklist.mustHold ? "hold" : "send",
  policyConflict: checklist.policyConflict,
  fixes: checklist.missingRequired,
  doNotAdd: checklist.removeFromPacket,
  summary: checklist.message,
};
Enter fullscreen mode Exit fullscreen mode

The consequence is deliberate. Hindsight contributes historical context, cited memories, and a confidence estimate. The deterministic assessment controls the final hold/send status and document lists for requirements encoded in the application. That division makes the behavior easier to explain: the model can identify a relevant past denial, while a missing letterhead date is not left to the model to notice consistently.

There is a tradeoff. If the hard-coded checklist and memory evidence become inconsistent, the application may hold a case because a rule remains in code even when newer outcomes suggest it has changed. The memory layer cannot automatically repair an outdated predicate. Production operation therefore needs an explicit process for reviewing and updating these rules, as well as monitoring for conflicts between coded checks and fresh outcomes. Making a rule deterministic improves repeatability; it does not make the rule permanently correct.

The feedback loop is an engineering feature

PayerLine does not stop at advice. When the insurer replies, staff record whether the case was approved, queried, or denied, along with the reason. recordOutcome stores a snapshot of the submitted packet: package name, letterhead status, culture, fitness note, consent, estimate, identity documents, and CBC. It retains that record with a procedure tag and an outcome tag in the selected payer’s bank.

That snapshot is important. A bare memory that says “denied” has little reuse value. The next review needs to connect the insurer’s result to what was actually in the packet. The reason preserves the insurer’s explanation; the structured flags make the documents and their state searchable alongside it. Stable document IDs are used for seeded history so the bank can be provisioned repeatedly without duplicating the same records.

The desk also exposes a ledger, making stored memories visible to the staff using them. A memory system that only influences a hidden answer is difficult to audit. Showing recent observations and outcomes gives a human a way to see what the system has retained and whether its remembered history looks relevant.

What happens in one review

Suppose the current packet is for Meridian. The package field says “management,” the ultrasound is attached but its date is not on hospital letterhead, and a culture report contradicts the diagnosis. The UI sends those fields to the API. The API queries the Meridian bank, scoped to the procedure, while the checklist marks the package name and letterhead issue as fixes and flags the culture report to remove.


The packet checklist for today's file, ready to review against memory.

The review can then show a hold recommendation, the specific changes to make, and dated prior outcomes that explain why the packet is risky despite the broad written policy. If the insurer later queries or denies it for another reason, the desk records that result into Meridian’s bank. A future review can retrieve that outcome. Switching the selected payer changes the bank as well as the applicable checklist; Harbour’s consent requirement should not leak into this Meridian case.

Same file, Northline selected: a different bank, a different checklist, and a send recommendation.


Meridian selected: hold, with the written-policy conflict and dated outcomes shown.

The note-entry flow can optionally use Groq function calling to extract structured fields from a clinical note. I kept that concern separate from payer memory: extraction turns text into case fields, packet assessment checks those fields, and Hindsight supplies history. That separation makes it clearer which part of the system produced a value and where to investigate when a review is wrong.

What I learned

  1. Put isolation in routing, not just in instructions. When the application already knows the tenant or payer, use that identity to choose the memory store before reasoning starts.
  2. Make evidence thresholds explicit. A system should have a way to say that a remembered pattern is based on too few outcomes, and a way to preserve disagreement instead of smoothing it away.
  3. Use deterministic code for known invariants. Repeated checks such as a required document or a date limit are easier to reason about as predicates. Keep their ownership and update path clear because code can also become stale.
  4. Treat citations and visible memory as part of the interface. A recommendation is more useful when staff can see the dated cases behind it and inspect the ledger that future decisions will use.
  5. Capture the outcome, not only the recommendation. The insurer’s eventual response and the packet state create the feedback needed for future reviews to improve.

The hardest design question was not how to ask a model “hold or send?” It was deciding which knowledge belongs in a payer’s evolving history, which rules must remain deterministic, and how a person can see the evidence behind the result. Keeping those boundaries visible is what makes memory useful in a workflow where a plausible answer is not enough.

Top comments (0)