DEV Community

Amrrutha K
Amrrutha K

Posted on

When Correct Memories Become Wrong Decisions: Building Context-Aware Manufacturing Memory with Hindsight

By Team ThinkMates — Amrutha K, Kammar Akshay, Rashmika K

1. The scenario

Sealer-02 is a heat-sealing machine on a packaging line. For months it ran Film-A from supplier PackCo under recipe R10, and the fix for any Weak Seal defect was reliable: Increase Temperature +5°C — four successes, zero failures.

Then production changed to Film-B from FlexPack under recipe R11. The Weak Seal defect returned — and the remembered fix, four successes and zero failures behind it, failed twice.

2. The problem with ordinary AI memory

Most AI memory answers one question well: what worked before?

They store experiences, retrieve them by similarity, and surface the past fix.

That works until the world moves.

"Temperature +5°C worked 4 out of 4 times" is technically correct — and operationally dangerous — when those times happened under a context that no longer exists.

3. Why remembering what worked is insufficient

A fix is never valid in the abstract.

It is valid for a defect, under a context:

  • Machine
  • Material
  • Supplier
  • Recipe
  • Firmware

Change the film and the seal physics change with it.

The old memory is still true — the fix really did work under Film-A/R10.

But truth and current validity differ, and a system that cannot tell them apart may recommend a fix that the new reality has already contradicted.

The fix did not become false. Its validity boundary changed.

4. Enter Validrift

Validrift is change-aware manufacturing memory: persistent Hindsight memory plus a deterministic context-validity engine deciding whether learned knowledge should still be trusted.

Hindsight remembers the manufacturing history.

The Validrift engine decides whether that knowledge remains valid under the current production context.

Its main ideas are:

  • Persistent Hindsight memory
  • Context-scoped fix validity
  • Change-triggered memory audits
  • Outcome-based revalidation
  • Preservation of historical truth

Nothing is deleted or globally overwritten just because the production context changed.

5. The Film-A/R10 history

Under:

  • Machine: Sealer-02
  • Material: Film-A
  • Supplier: PackCo
  • Recipe: R10
  • Defect: Weak Seal

the corrective action Increase Temperature +5°C had:

4 successes / 0 failures

Status:

VALIDATED

Validrift preserves this as historical evidence.

It happened, it worked, and that historical fact remains true.

6. The process change

Production later changed:

Film-A → Film-B

PackCo → FlexPack

R10 → R11

Validrift records process changes as first-class events in its structured ledger and in Hindsight.

A process change tells the system that previously learned knowledge may need to be re-evaluated.

7. The drift of Temperature +5°C

Under Film-B/R11, the same fix failed twice:

0 successes / 2 failures

Status:

DRIFTED

The old record is not erased.

The Fix Passport can show both facts at the same time:

Film-A/R10

Temperature +5°C

4 successes / 0 failures

VALIDATED

Film-B/R11

Temperature +5°C

0 successes / 2 failures

DRIFTED

The memory wasn't wrong.

Its validity drifted.

8. Context-scoped validity

Validrift's central idea is that validity is a function of:

(fix, defect, context)

—not of the fix alone.

Every intervention record carries the production context.

For a new incident, the deterministic engine separates evidence into:

  • Current-context evidence
  • Historical evidence

It then evaluates each fix using one of these statuses:

  • VALIDATED
  • SUPPORTED
  • UNVERIFIED
  • REVALIDATION REQUIRED
  • DRIFTED
  • INSUFFICIENT EVIDENCE

This lets the same corrective action have different validity states under different production conditions.

9. The Fix Passport

A Fix Passport is the biography of one corrective action.

It contains:

  • Previous attempts
  • Outcomes
  • Production contexts
  • Current-context evidence
  • Historical evidence
  • Validity state

If a quality engineer asks:

"Why not Temperature +5°C?"

Validrift can answer with both halves of the truth:

Film-A/R10:

4/0 — VALIDATED

Film-B/R11:

0/2 — DRIFTED

There is no global overwrite and no silent forgetting.

10. The Memory Validity Audit

A process change can trigger a Memory Validity Audit.

Known fixes are re-evaluated against the current context.

For example:

  • A fix with only historical success may become REVALIDATION REQUIRED
  • A historically successful fix that repeatedly fails in the current context may become DRIFTED
  • A fix with repeated current-context success can become VALIDATED

Instead of treating every remembered fix as permanently trustworthy, Validrift continuously asks:

Does this knowledge still apply here?

11. Hindsight RETAIN

Manufacturing events are retained into the Hindsight memory bank.

These can include:

  • Incidents
  • Interventions
  • Outcomes
  • Process changes
  • Recommendations
  • Recommendation outcomes

HindsightService.retain uses stable document IDs and records a MemoryTrace row in SQLite.

This makes memory operations auditable with:

  • Operation type
  • Status
  • Latency
  • Returned memory IDs
  • Request information

12. Hindsight RECALL

When a new Weak Seal incident occurs on Sealer-02 under Film-B/R11, Validrift recalls relevant manufacturing experience from Hindsight.

This can include:

  • Similar incidents
  • Previous corrective actions
  • Outcomes
  • Relevant context

During the verified live run, the recommendation flow received 49 real recalled memory IDs.

The important distinction is:

Recall supplies evidence. It does not supply the final validity verdict.

That decision belongs to Validrift's deterministic validity engine.

13. Hindsight REFLECT

Hindsight REFLECT is used to generate higher-level understanding from accumulated experience.

For example, Validrift can ask:

How did the effectiveness of Temperature +5°C change between Film-A/R10 and Film-B/R11?

Reflection helps explain how knowledge evolved across contexts.

However, REFLECT does not determine the validity status.

The deterministic engine remains responsible for:

VALIDATED, SUPPORTED, DRIFTED, and the other validity states.

14. The deterministic validity engine

The heart of Validrift is deliberately not an LLM.

A simplified part of the real logic is:

if hs > 0 and cf >= settings.drift_min_failures and cs == 0:
    status = ValidityStatus.DRIFTED
    reason = (
        f"Historically successful evidence exists, "
        f"but {cf} repeated failure(s) occurred in the current context."
    )

elif cs >= settings.validated_min_successes and ratio >= 0.75:
    status = ValidityStatus.VALIDATED
    reason = (
        f"{cs} successful current-context outcomes "
        f"support repeated validation."
    )
Enter fullscreen mode Exit fullscreen mode

There are no language-model probabilities deciding whether a fix is valid.

The thresholds are deterministic configuration, and every status is backed by evidence and a reason.

15. Outcome learning: the learning loop

A recommendation is only a hypothesis until its outcome is recorded.

The recommendation flow writes its reasoning back into memory:

retain_text = (
    f"Validrift recommended '{best.fix_name}' for incident "
    f"{incident.id} ({incident.defect}) under "
    f"machine {incident.machine}, "
    f"material {incident.material}, "
    f"supplier {incident.supplier}, "
    f"recipe {incident.recipe}, "
    f"firmware {incident.firmware}. "
    f"Validity status: {best.status}. "
    f"Current-context evidence: "
    f"{best.current_successes} success(es), "
    f"{best.current_failures} failure(s)."
)

await memory_service.retain(
    db,
    request_id=request_id,
    content=retain_text,
    context="Validrift recommendation",
    document_id=f"recommendation:{rec_id}",
    timestamp=rec.created_at,
    metadata={
        "recommendation_id": rec_id,
        "incident_id": incident.id,
        "fix": best.fix_name
    },
    tags=[
        "validrift",
        "recommendation",
        f"defect:{incident.defect.lower().replace(' ', '-')}"
    ]
)
Enter fullscreen mode Exit fullscreen mode

In the verified end-to-end flow:

Pressure +8%

started with:

2 successes / 0 failures

Status:

SUPPORTED

One additional successful recorded outcome changed the evidence to:

3 successes / 0 failures

Status:

VALIDATED

The loop becomes:

Recommend → Record Outcome → Retain → Re-evaluate → Improve Future Recommendations

16. Architecture

Quality Engineer
        │
        ▼
Validrift Frontend
        │
        ▼
FastAPI Backend
        │
        ├────────► SQLite
        │          Incidents
        │          Interventions
        │          Process Changes
        │          Recommendations
        │          Outcomes
        │          Memory Traces
        │
        ├────────► Hindsight Cloud
        │          RETAIN
        │          RECALL
        │          REFLECT
        │
        └────────► Deterministic Validity Engine
                         │
                         ▼
                  Context-Scoped Validity
                         │
                         ▼
                    Recommendation
                         │
                         ▼
                   Recorded Outcome
                         │
                         ▼
                   Hindsight RETAIN
                         │
                         ▼
                Better Future Recommendation
Enter fullscreen mode Exit fullscreen mode

Nothing in this architecture automatically controls manufacturing equipment.

Validrift is a decision-support system, not autonomous machine control.

17. From the actual codebase

The system defines the production context using five core fields:

CORE_CONTEXT_FIELDS = (
    "machine",
    "material",
    "supplier",
    "recipe",
    "firmware"
)
Enter fullscreen mode Exit fullscreen mode

Evidence is represented explicitly:

@dataclass
class EvidenceItem:
    incident_id: str
    intervention_id: str
    date: datetime
    defect: str
    fix_name: str
    outcome: str
    context: dict
    notes: str

    def to_dict(self):
        d = asdict(self)
        d["date"] = self.date.isoformat()
        return d
Enter fullscreen mode Exit fullscreen mode

Each evaluated fix keeps current and historical evidence separately:

@dataclass
class FixEvaluation:
    fix_name: str
    defect: str
    status: str
    current_successes: int
    current_failures: int
    current_partial: int
    historical_successes: int
    historical_failures: int
    current_evidence: list[EvidenceItem]
    historical_evidence: list[EvidenceItem]
    relevant_process_change: dict | None
    reason: str
    score: float
Enter fullscreen mode Exit fullscreen mode

And context equality is deterministic:

def context_of(incident: Incident) -> dict:
    return {
        k: getattr(incident, k)
        for k in CORE_CONTEXT_FIELDS
    }


def contexts_match(a: dict, b: dict) -> bool:
    return all(
        (a.get(k) or "") == (b.get(k) or "")
        for k in CORE_CONTEXT_FIELDS
    )
Enter fullscreen mode Exit fullscreen mode

Every Hindsight operation also leaves an auditable memory trace:

def _trace(
    self,
    db: Session,
    request_id: str,
    operation: str,
    summary: str,
    started: float,
    memory_ids: list[str] | None = None,
    status: str = "SUCCESS",
    details: dict | None = None
):
    db.add(
        MemoryTrace(
            id=new_id("MT"),
            request_id=request_id,
            operation=operation,
            query_summary=summary,
            memory_ids=memory_ids or [],
            latency_ms=round(
                (time.perf_counter() - started) * 1000,
                2
            ),
            status=status,
            details=details or {},
        )
    )

    db.commit()
Enter fullscreen mode Exit fullscreen mode

18. Verified results

Check Result
Backend test suite 13/13 passed
Frontend pages 5/5 hydrated, zero JS errors
Hindsight RETAIN / RECALL / REFLECT PASS / PASS / PASS
/api/health hindsight: {mode: live, ok: true}
Temperature +5°C, Film-A/R10 4/0, VALIDATED
Temperature +5°C, Film-B/R11 0/2, DRIFTED — history preserved
Pressure +8%, Film-B/R11 2/0 SUPPORTED → 3/0 VALIDATED after recorded outcome

19. Degraded-mode behavior

Validrift is designed to fail honestly.

If Hindsight becomes unavailable, the deterministic engine can still compute validity statuses from the structured SQLite ledger.

The UI reports:

Memory service temporarily unavailable.

If natural-language explanation is unavailable, the system reports:

Evidence is available, but natural-language explanation is temporarily unavailable.

It does not fabricate recalled memories, successful memory events, or reflection output.

20. Limitations

Validrift still has important limitations.

  • Validity is only as good as the recorded outcomes. Unrecorded interventions cannot become reliable evidence.
  • Context matching currently uses exact matching across five production fields.
  • Near-matching contexts are not fuzzy-matched.
  • The LLM explains evidence; it does not determine validity.
  • Validrift is decision support and does not automatically control manufacturing equipment.

21. Conclusion

The instinct — in manufacturing and in AI memory design — is to treat the past as a promise.

Validrift treats the past as evidence:

kept forever, trusted conditionally, and re-examined whenever the context changes.

A corrective action can deserve:

VALIDATED under Film-A/R10

and:

DRIFTED under Film-B/R11

at the same time.

That distinction is the core of change-aware memory.

The memory wasn't wrong. Its validity drifted.

Top comments (0)