DEV Community

Yelem Venkata Rachika
Yelem Venkata Rachika

Posted on

A restart fixed the incident. Eight minutes later, it was back.

A restart fixed the incident. Eight minutes later, it was back. Remembering Operational Anti-Patterns and Failed Fixes using Hindsight

If you have spent time on call in a complex microservice environment, you have lived this exact sequence. An alert fires, error rates surge, and someone attempts a quick service restart. The metrics momentarily dip back into the green, everyone relaxes on the incident call, and ten minutes later the error rate spikes right back up to 18%.

Most AI incident tools and runbook automation systems are obsessed with recording successful solutions. When an incident closes, they store the final patch or configuration fix. But in real-world operations, remembering what failed during troubleshooting is just as valuable as remembering what worked.

When building DeployLens for our NovaCart production environment, we designed the memory architecture to explicitly capture operational anti-patterns using the [Hindsight agent memory SDK]@code.in (https://github.com/vectorize-io/hindsight).

Here is how we store, format, and recall failed fix attempts so engineering teams stop repeating the same troubleshooting mistakes.

The Problem with Success-Only Memory

When an AI investigation agent only remembers successful outcomes, it suffers from a critical blind spot: it cannot warn engineers away from plausible-sounding workarounds that have repeatedly failed in the past.

Consider a common scenario: checkout-api latency spikes due to downstream Redis connection pool exhaustion after a release of payment-service. To a generic troubleshooting assistant, restarting checkout-api sounds perfectly reasonable. Restarting worker processes clears transient memory and flushes stale socket connections.

However, if three previous SREs already tried restarting checkout-api during identical past outages and found it provided only 8 minutes of temporary relief before connection queues filled up again, that knowledge is vital. If that experience isn't captured, every new engineer on call will repeat the exact same ineffective restart step.

Using Vectorize persistent agent memory, DeployLens treats failed troubleshooting attempts as first-class memory units.

Formatting Failure into Anti-Pattern Memories

To turn failed engineering actions into usable memories, we created dedicated formatting rules inside MemoryFormatter (backend/app/memory/memory_formatter.py).

Here is how we convert failed troubleshooting actions into semantic memories:

@staticmethod
def format_investigation(
    incident_id: str,
    service_name: str,
    engineer: str,
    action: str,
    result: str,
    outcome_type: str
) -> str:
    if outcome_type in ("FAILED_TEMPORARY", "FAILED_NO_EFFECT"):
        return (
            f"During [{incident_id}] on {service_name}, engineer {engineer} attempted action: '{action}'. "
            f"Result: {result}. Outcome: This action WAS NOT a permanent fix ({outcome_type})."
        )
    else:
        return (
            f"During [{incident_id}] on {service_name}, engineer {engineer} performed action: '{action}'. "
            f"Result: {result}. Outcome: {outcome_type}."
        )
Enter fullscreen mode Exit fullscreen mode

Why This Matters

Notice the structure of the generated text string. We explicitly include phrases like "WAS NOT a permanent fix" and append outcome types such as FAILED_TEMPORARY.

When Hindsight indexes these strings into its vector space, queries searching for historical fixes on checkout-api retrieve these failure memories alongside resolutions. The language choices ensure the LLM reasoning agent understands that the action was attempted but failed to solve the root problem.

Writing Outcome Memories to Hindsight

When an engineer resolves an incident in DeployLens, they submit a postmortem form containing both the confirmed root cause and any failed fix attempts.

Inside backend/app/memory/memory_service.py, store_incident_outcome processes these failed attempts and writes them to the Hindsight memory bank:

def store_incident_outcome(
    self,
    incident_id: str,
    service_name: str,
    root_cause: str,
    successful_fix: str,
    failed_attempts: List[str],
    timestamp: str = ""
) -> str:
    ts = timestamp or datetime.utcnow().strftime("%Y-%m-%d %H:%M")
    content = MemoryFormatter.format_resolution(
        incident_id, service_name, root_cause, successful_fix, failed_attempts
    )
    meta = {
        "service": service_name,
        "incident_id": incident_id,
        "source_type": "resolution",
        "outcome": "SUCCESS"
    }
    mem_id = self.remember_event(content, "resolution", meta, tags=["resolution", service_name, incident_id])

    # Store failed fix memories explicitly if any
    for failed in failed_attempts:
        failed_content = MemoryFormatter.format_investigation(
            incident_id, service_name, "Incident Team", failed,
            "Temporary recovery or no permanent effect", "FAILED_TEMPORARY"
        )
        self.remember_event(failed_content, "failure", {
            "service": service_name,
            "incident_id": incident_id,
            "source_type": "failure"
        }, tags=["failure", service_name])

    return mem_id
Enter fullscreen mode Exit fullscreen mode

Why This Matters

This function creates two distinct types of records in Hindsight:

  1. A resolution memory capturing the successful fix.
  2. Individual failure memories for every failed attempt submitted by the SRE, tagged explicitly with source_type: "failure".

When future incidents occur, Hindsight recall returns both the verified fix and the anti-pattern warnings.

Surfacing Anti-Patterns in the User Interface

Storing failed fixes in memory is only half the battle; the UI must present them as explicit warnings so engineers don't overlook them.

In DeployLens, when Hindsight recalls a failed action pattern, the frontend renders a dedicated FailedFixCard component displaying:

  • Previously Attempted Action: e.g., "Restarted checkout-api service instance 3 times".
  • Historical Outcome: "0/3 Permanent Resolutions".
  • DeployLens Memory Warning: "This action historically provided temporary relief (5-8 mins) but did NOT permanently resolve the root cause."

Before vs After: Success-Only Memory vs Anti-Pattern Memory

Here is how incident investigation changes when your memory system retains failed troubleshooting attempts.

WITHOUT FAILED-FIX MEMORY

During active incident INC-2051 (Checkout API Error Rate Spike):

  1. Telemetry shows HTTP 500 errors at 18.4% and Redis timeout warnings.
  2. The agent searches past records and finds that checkout-api had high latency two weeks ago.
  3. Because failed troubleshooting steps weren't recorded, the agent suggests: "Try restarting the checkout-api pods to restore connection pools."
  4. The engineer restarts the service. Errors disappear for 7 minutes, then return to 18.4%. Time wasted: 15 minutes.

WITH HINDSIGHT FAILED-FIX MEMORY

With anti-pattern indexing enabled via Hindsight agent memory documentation:

  1. Telemetry shows HTTP 500 errors at 18.4% and Redis timeout warnings.
  2. Hindsight recalls INC-1042: "Restarting checkout-api was attempted 3 times and provided only 8 minutes of temporary relief."
  3. DeployLens displays an explicit warning: "DO NOT RESTART CHECKOUT-API."
  4. DeployLens directs the engineer to the verified fix: "Increase PAYMENT_REDIS_POOL_SIZE from 20 to 50 in payment-service configuration."
  5. The engineer applies the config change. Errors drop to <0.1% permanently. Time to resolution: 4 minutes.

Tradeoffs and Limitations

Capturing failed fix attempts introduces specific operational tradeoffs:

  1. Postmortem Form Discipline: Failure memories rely on engineers filling out the "Failed Fix Attempts" field when closing an incident. If SREs skip this field during resolution, anti-pattern memory cannot accumulate.
  2. False Equivalency Risk: An action that failed under one root cause might be the correct resolution for a different root cause. We mitigate this by attaching service tags and symptom metadata to every failure memory unit.
  3. Over-reliance on Historical Warnings: Engineers must balance anti-pattern warnings with active telemetry inspection to ensure novel failures are not misdiagnosed based solely on past occurrences.

Practical Lessons for Engineers

  1. Treat Failures as Assets: Record what didn't work during an outage. Ineffective actions are valuable organizational anti-patterns.
  2. Tag Failure Outcomes Explicitly: Use outcome markers like FAILED_TEMPORARY when writing memory units to Hindsight.
  3. Warn Prominently in the UI: Display anti-patterns with clear visual warnings so engineers don't repeat ineffective workarounds.
  4. Link Failures to Root Causes: Ensure failure memories retain symptom and service context to prevent false positive warnings.
  5. Differentiate Temporary Relief from Root Cause Fixes: Explicitly distinguish workarounds that lower error rates briefly from fixes that restore long-term stability.


Top comments (0)