DEV Community

saisharanya thammali
saisharanya thammali

Posted on

Memory filtering Keeping the Live Incident Out of the Past: A Memory Filtering Problem in Incident Response

Keeping the Live Incident Out of the Past: A Memory Filtering Problem in Incident Response

The most dangerous result in an incident memory search may be the one that looks perfect.

Imagine investigating HTTP 500 errors on the Payment API while database connections are unusually high after a traffic spike. Hindsight returns a record with the same service and symptoms. It appears to confirm the diagnosis—until you discover that the record is today’s incident, already retained during an earlier attempt to investigate it. The search result is highly similar, but it is not historical evidence at all.

That subtle distinction became an important engineering problem in IncidentMind. Retrieval is only part of a useful memory system. Before the model or engineer can rely on returned incidents, the application must decide whether they are actually previous incidents, whether results repeat the same event, and how many to show.

Recall is a starting point

IncidentMind builds a natural-language query from the current service, error, and symptoms. It asks Hindsight to find prior incidents involving similar services, errors, symptoms, root causes, triggers, and resolutions. The app then takes up to ten returned memory units for further filtering:

result = hindsight.recall(
    bank_id=HINDSIGHT_BANK_ID,
    query=query
)

memories = getattr(
    result,
    "results",
    []
)

# Get extra memories so filtering does not leave us
# with an unnecessarily small historical set.
memories = memories[:10]

memories = filter_historical_memories(
    memories,
    service,
    error,
    symptoms
)
Enter fullscreen mode Exit fullscreen mode

The ten-result cap is a buffer. Some returned items may describe the live incident or duplicate another memory, so filtering only the first five could leave the interface with fewer than five usable historical entries even if later results were available. After cleaning, the application returns at most five unique memories for display and reasoning.

This is not a formal retrieval-accuracy scheme. Hindsight does the recall; the application applies practical checks to the retrieved text. The code does not claim to compute a calibrated similarity score or prove that each remaining result is relevant.

The live incident can leak into history

Why would today’s incident be in the memory bank at all? IncidentMind has a button that retains an investigation for future use. An engineer may run an investigation, store it, and then investigate the same still-active event again. The query naturally resembles the record that was just stored. Similarity alone cannot distinguish “a previous incident with the same symptoms” from “this incident, stored a few minutes ago.”

The app’s first defense looks for current-date and text clues. It checks whether a memory contains today’s date; if so, it treats a matching current service as a sign that the memory may be the live incident. If the service is not enough, it counts matching words from the current error or symptoms and filters the memory when at least two qualifying words appear. It also checks explicit text markers such as live-001, current incident, and current production incident.

Here is the core shape of the date-based logic:

current_date = date.today().isoformat()

if current_date in text:
    if service_text in text:
        return True

    # Error and symptom keyword checks follow...
Enter fullscreen mode Exit fullscreen mode

This approach tries to avoid a broad rule like “discard every result that mentions the service.” Such a rule would suppress useful older Payment API incidents whenever the current event also concerns the Payment API. Date plus service or symptom overlap is a narrower clue that a result may refer to today’s event.

It remains a heuristic. A legitimate historical record updated today could be removed; a current incident without the expected date or marker could slip through. Text matching also depends on how incident details are phrased. The checks reduce a known class of mistakes but cannot establish identity with certainty.

Similar memories can be duplicates

Hindsight may return multiple memory units that refer to the same incident. Showing each as a separate past outage would distort the engineer’s impression of how many independent precedents exist. IncidentMind deduplicates the filtered results before displaying them.

When the text contains an ID matching INC- followed by digits, the first matching ID is normalized to uppercase and used as the identity key:

ids = re.findall(
    r"\bINC-\d+\b",
    text,
    flags=re.IGNORECASE
)

if ids:
    incident_id = ids[0].upper()

    if incident_id in seen_incident_ids:
        continue

    seen_incident_ids.add(incident_id)
    unique_memories.append(memory)
Enter fullscreen mode Exit fullscreen mode

If no such ID is present, the app lowercases the text, collapses runs of whitespace into a single space, and drops exact matches of that normalized string. This catches duplicated memory units whose only textual difference is capitalization or spacing.

There is a tradeoff here too. Two different records that share the same first incident ID will collapse into one. Conversely, two paraphrases of the same event without an ID will remain separate because their normalized text differs. The implementation does not perform semantic clustering or fuzzy duplicate detection. Its simple rules are inspectable and predictable, but they do not cover every way duplicate memories can be represented.

Why the engineer should see what was retrieved

After filtering, IncidentMind exposes the supporting memory text in expandable “Past Experience” entries. It also extracts dates from retrieved text for a small timeline. This gives the engineer a chance to catch mistakes the rules miss: a memory might be labeled as historical by the filter yet describe the wrong service, an unrelated incident, or the current event in different words.

That visibility matters because the retrieved context influences the LLM analysis. If the UI showed only a root-cause hypothesis, an accidental current-incident match could look like independent corroboration. Showing the source makes it possible to ask: Is this actually a prior event? What matches? What differs? Is the system counting one outage twice?

The application also reports recurring patterns, but those labels come from straightforward text checks. For example, detect_patterns searches the combined historical memory text for phrases such as http 500, database connection, connection pool, traffic spike, workers, and timeout. It returns matching labels, up to five. This can help summarize visible history, but it is not semantic analysis and does not prove that a pattern recurs across independent incidents.

A separate “Memory Signal” label is similarly explicit in the code: it counts predefined signal strings found in both the current incident text and the combined historical text, then maps the total to weak, moderate, or strong recurring-pattern labels. The comments call it a transparent heuristic and say it is not an ML similarity score. Because it counts signal presence in the combined text, repeated occurrences do not necessarily represent independent incidents. It is a UI explanation aid, not a probability of a correct match.

A concrete failure the filters are trying to prevent

Suppose the current Payment API event has HTTP 500s, high database connections, and a traffic spike. A record retained for that event includes today’s date and the same service and symptoms. Hindsight may rank it highly because it is almost identical to the query. If presented as “a past incident,” it could falsely suggest that the same root cause has already been confirmed before.

The date-and-text check is intended to remove that result. Older sample record INC-009, by contrast, describes a similar Payment API failure after a traffic spike and connection-pool exhaustion. Because its date differs, it can remain useful historical context. The distinction is not “similar versus dissimilar”; it is “same current event versus a genuinely earlier event.” Similarity is necessary for useful recall, but it is not sufficient for historical relevance.

Limits and lessons

The current filter relies on dates, a small set of text markers, keyword overlap, IDs in a particular format, and normalized exact-text equality. It can produce false positives and false negatives. The UI’s recurring-signal labels are also keyword heuristics. The repository provides no benchmark for retrieval quality, duplicate detection, or filtering effectiveness, so there is no basis for claiming formal accuracy.

A more robust design could attach a stable incident identifier and explicit event lifecycle metadata to every retained record, then filter by identity rather than infer identity from prose. It could also separate confirmed historical outcomes from an AI-generated investigation and preserve provenance for each memory. These are potential improvements, not capabilities in the current code.

The lesson from building IncidentMind is that memory retrieval needs a careful boundary between resemblance and history. A near-perfect match may be a useful prior—or it may be the very incident on the screen. Filtering, deduplication, and showing the underlying records give engineers ways to manage that ambiguity. They do not eliminate it, and they should not pretend to. In incident response, a memory earns trust through provenance and inspection, not similarity alone.

Top comments (0)