DEV Community

Sai Bharadhwaj
Sai Bharadhwaj

Posted on

Raising Cloud Agent Accuracy from 71% to 88% with Hindsight

Every cloud cost optimization tool recommends deletions, and engineers ignore most of them. The math behind the hesitation is simple: if an agent recommends terminating an "idle" EC2 instance that turns out to be the warm disaster-recovery standby for your primary datastore, you don't save $300 a month—you cause a multi-million-dollar outage.

The core failure of automated FinOps agents is not detection; identifying an instance with 1% average CPU utilization over fourteen days is trivial. The failure is context blindness. Telemetry measures utilization, but it cannot measure architectural intent. The missing context already exists inside your engineers' heads, and it surfaces every time they hit "Reject" on a recommendation in Slack. Until now, that reasoning vanished into message history.

We built CloudSense to turn that ephemeral human feedback into durable institutional knowledge. By integrating Hindsight agent memory, our AWS cost agent learns why specific resources or architectural patterns must never be touched, generalizes those rules to untagged look-alikes across accounts, and raised its recommendation acceptance rate from 71% to over 88% across consecutive evaluation cycles.

Here is the engineering breakdown of how we structured the system, how Hindsight's memory pipeline works under the hood, and the lessons learned deploying persistent agent memory to production infrastructure.


What CloudSense Does and How It Hangs Together

CloudSense sits between an organization's AWS footprint and its engineering review channels. Rather than giving an LLM unconstrained access to scan and mutate infrastructure, the system decouples discovery, blast-radius calculation, candidate generation, memory-augmented decisions, and human execution:

[AWS Accounts (Read-Only Role)]
             │
             ▼
     [Resource Scanner] ─── (Describe + 14d CloudWatch)
             │
             ▼
    [Dependency Graph]  ─── (BFS Blast Radius Calculation)
             │
             ▼
  [Deterministic Rules] ─── (R1–R8 Candidates + Pricing)
             │
             ▼
      [Decision Agent]  ◄─── [Hindsight Bank: Recall + Playbook]
             │
       ┌─────┴────────────────┐
       ▼                      ▼
  [Suppressed]          [Review Queue] (Web UI / Slack)
  (Cited Memory)              │
                              ▼
                       [Human Verdict] (Approve / Reject + Reason)
                              │
                              ▼
                     [Hindsight: Retain]
                              │
                              ▼
                 [Asynchronous Consolidation] ──► [Learned Rules & Playbook]
Enter fullscreen mode Exit fullscreen mode
  1. Discovery and Graph Building: A cross-account read-only IAM role discovers resources and pulls 14-day CloudWatch metrics. A dependency engine maps relationships (such as EBS attachments, load balancer target groups, and RDS read replicas) to compute an explicit blast radius before any recommendation is considered.
  2. Deterministic Candidate Generation: Rules (R1–R8) identify potential waste—idle compute, unattached gp2 volumes, obsolete snapshots—and calculate estimated dollar savings using a localized pricing matrix.
  3. Agent Decision Pipeline: Before presenting recommendations to humans, the agent queries Hindsight. If past human verdicts or learned organizational rules indicate that an instance or its topological pattern is protected, the agent suppresses the action and logs an auditable citation.
  4. Human Review: Actionable items surface in a web dashboard and Slack with dependency previews. Engineers approve or reject. Rejections require a stated reason.
  5. Memory Retain & Consolidation: Verdicts are pushed to Hindsight agent memory. Asynchronously, Hindsight consolidates raw events into high-order Observations (learned rules) and refreshes an organization-wide Cost Playbook.
  6. Reversible Execution: Approved actions execute via an opt-in action role. Volumes are snapshotted prior to modification, and changes are fully auditable with rollback capabilities.

Figure 1: The CloudSense review queue displaying candidate recommendations with calculated blast radius, estimated savings, recalled memory citations, and human approval controls.

The Core Technical Story: Closing the Context Gap

Traditional FinOps automations operate in an open loop: they scan, generate recommendations, and wait for human action. When an engineer rejects a proposal—such as "Don't touch this worker, it runs month-end payroll processing"—the system discards that insight. On the subsequent scan, the tool flags the exact same worker. Engineers rapidly develop alert fatigue, and the tool gets muted.

Hardcoding manual exclusion tags (like do_not_delete=true) fails because infrastructure moves faster than tagging compliance. Standby databases spin up dynamically, engineers forget tags during urgent deploys, and cross-account look-alikes appear constantly.

To build an agent that genuinely earns trust, we needed three capabilities:

  1. Episodic Retention: Capture the full context of why an engineer rejected or approved a proposed action.
  2. Thematic Consolidation: Aggregate individual verdicts into generalized architectural policies without human prompt curation.
  3. Topological Generalization: Recognize that an untagged instance in Account B matches the protected pattern of an instance in Account A based on its role hints and graph topology.

This is where Hindsight fits into our architecture.


Code-Backed Implementation

1. Pre-Filtering Directives and Parallel Memory Recall

The decision agent evaluates candidates concurrently. Before making any LLM call, it checks bank directives (hard constraints that must never be violated). If clear, it executes a parallel recall against Hindsight alongside fetching the current mental model playbook:

# backend/agent/decide.py
async def decide_all(
    org: Any,
    candidates: list[Candidate],
    resources: list[Resource],
    memory_client: MemoryClientProtocol | None,
    llm: LLM,
    *,
    no_memory: bool = False,
    recalled_out: dict[str, list[MemoryHit]] | None = None
) -> dict[str, AgentDecision]:
    by_id = {r.id: r for r in resources}
    out: dict[str, AgentDecision] = {}
    todo: list[tuple[Candidate, Resource]] = []

    for c in candidates:
        res = by_id[c.resource_id]
        if reason := directive_block(res):
            out[c.id] = AgentDecision(
                decision="suppress", action="none", confidence=1.0,
                predicted_approval=0.0, reason=reason, risk_note="hard directive"
            )
        else:
            todo.append((c, res))

    if not todo:
        return out

    # Parallel retrieval: fetch org playbook and candidate-specific memories
    playbook_obj, recalled = await asyncio.gather(
        memory_client.get_playbook(org),
        memory_client.recall_many(org, todo)
    )
    system = build_system_prompt(org, org_directives(org), playbook_obj.content)
    sem = asyncio.Semaphore(LLM_CONCURRENCY)

    async def run(c: Candidate, r: Resource, mems: list[MemoryHit]) -> None:
        if recalled_out is not None:
            recalled_out[c.id] = mems
        async with sem:
            out[c.id] = await decide_one(llm, system, r, c, mems)

    await asyncio.gather(*(run(c, r, m) for (c, r), m in zip(todo, recalled, strict=True)))
    return out
Enter fullscreen mode Exit fullscreen mode

By querying Hindsight using recall_many, each candidate is paired with relevant past verdicts based on its action type, role hints, utilization signals, and neighboring dependencies.

2. Retaining Verdicts with Observation Scopes

When a reviewer rejects a candidate in Slack, we record the verdict in Hindsight using aretain_batch. Instead of binding the memory strictly to a single instance ID, we structure the document with explicit observation_scopes:

# backend/memory/types.py
def verdict_item(scan_id: str, cand: Candidate, res: Resource, verdict: Any, reviewer: Any) -> dict[str, Any]:
    """One retain_batch item. observation_scopes: org-wide ([] = untagged), per-account, per-team."""
    at = verdict.at or datetime.now(UTC)
    return {
        "content": verdict_content(scan_id, cand, res, verdict, reviewer),
        "context": VERDICT_CONTEXT,
        "timestamp": at,
        "tags": verdict_tags(res),
        "metadata": {
            "candidate_id": cand.id,
            "resource_id": res.id,
            "decision": _val(verdict.decision),
            "reviewer_id": str(reviewer.id),
            "scan_id": scan_id,
        },
        "entities": [
            {"text": res.display_name},
            *[{"text": n} for n in res.depends_on_names + res.dependent_names],
            *([{"text": res.service}] if res.service else []),
        ],
        "observation_scopes": [
            [],                             # Global/org-wide scope
            [f"account:{res.account_alias}"], # Account-specific scope
            [f"team:{res.owner_team}"],       # Team-specific scope
        ],
    }
Enter fullscreen mode Exit fullscreen mode

The empty inner list [] defines the global scope. This allows Hindsight to cluster insights that apply across all accounts in the organization, while [account:...] preserves regional or account boundaries.


Figure 2: The engineer rejection flow capturing the required explanation ("That's the DR standby for payment-db, it's idle on purpose") and structured observation scope.

3. Bank Disposition and the Living Playbook

To ensure the agent behaves conservatively with production assets, we configure the Hindsight bank with explicit behavioral dispositions:

# backend/memory/setup.py
DISPOSITION = {
    "disposition_skepticism": 5,
    "disposition_literalism": 4,
    "disposition_empathy": 2,
}

PLAYBOOK_QUERY = (
    "Which AWS resources or resource patterns must never be stopped, "
    "deleted or rightsized in this organization, and why? Include evidence."
)

PLAYBOOK_TRIGGER = {
    "refresh_after_consolidation": True,
    "mode": "delta",
    "keep_trace": True,
    "min_refresh_interval_seconds": 0,
    "response_schema": PLAYBOOK_SCHEMA,
}
Enter fullscreen mode Exit fullscreen mode

Setting disposition_skepticism to 5 prevents the memory engine from jumping to conclusions after a single isolated verdict. A high disposition_literalism (4) ensures that engineer reasons are interpreted strictly as stated.

Whenever Hindsight consolidates memory batches, the PLAYBOOK_TRIGGER fires in delta mode. As described in the Hindsight documentation, this automatically synthesizes a structured mental model—an up-to-date summary of protected patterns—which is injected directly into the agent's system prompt on every scan.


Figure 3: The Learned Rules dashboard displaying synthesized Hindsight observations, confirmation counts ("Confirmed 3×"), and provenance tracing back to source verdicts.

4. Memory Governance: Invalidation Over Deletion

In production, engineers occasionally make incorrect rejection decisions or change their mind when architecture evolves. Naive vector databases struggle with cache invalidation because individual embeddings cannot be cleanly excised from cluster centroids.

Because Hindsight tracks provenance between derived Observations and raw episodic memories, CloudSense handles rule removal by invalidating source facts:

# backend/memory/client.py
async def delete_rule(self, org: Any, memory_id: str, reason: str = "removed by admin") -> None:
    """Hindsight cannot curate observations directly: invalidate their source facts instead."""
    bid = bank_id(org)
    mem = await self.hs.memory.get_memory(bid, memory_id)
    sources = getattr(mem, "source_memory_ids", None) or [memory_id]
    for sid in sources:
        await self.hs.memory.update_memory(
            bid,
            sid,
            UpdateMemoryRequest(state="invalidated", reason=reason)
        )
Enter fullscreen mode Exit fullscreen mode

By invalidating the underlying source verdicts, Hindsight drops the derived observation during the next consolidation run without compromising database integrity.


Concrete Results: Measuring Accuracy Across Scans

To evaluate whether agent memory measurably improves recommendation quality, we ran benchmark simulations across 10 sequential infrastructure scans across heterogeneous accounts.

The test harness introduced six common real-world "trap" archetypes that context-blind tools fail on:

  • Disaster Recovery Standby: Instances that sit completely idle until primary datastore failover.
  • Month-End Payroll Batch: High-memory nodes with near-zero CPU for 27 days that spike to 95% at month end.
  • Audit Archive Storage: S3 buckets under regulatory legal holds that must not be transitioned to Glacier.
  • Blue/Green Standby: The passive color of a deployment pipeline held at zero traffic for instant rollbacks.
  • License-Pinned Host: Compute sized specifically to adhere to third-party per-core software licenses.
  • SRE On-Call Debug Host: Idle virtual machines reserved for active incident triage.

In each scan, instance IDs were randomized and operational tags were stripped to reflect the messy reality of enterprise cloud accounts.

Recommendation Acceptance Rate Across 10 Scans
─────────────────────────────────────────────────────────────
Scan Number │ Baseline (No Memory) │ With Hindsight Memory
────────────┼──────────────────────┼─────────────────────────
Scan 1      │        72.4%         │         72.4%
Scan 2      │        71.1%         │         84.2%
Scan 3      │        68.1%         │         86.3%
Scan 4      │        73.7%         │         88.1%
Scan 5      │        72.6%         │         88.5%
Scan 6      │        72.1%         │         88.5%
Scan 7      │        69.9%         │         88.3%
Scan 8      │        69.8%         │         88.3%
Scan 9      │        70.6%         │         88.4%
Scan 10     │        72.4%         │         88.4%
─────────────────────────────────────────────────────────────
Enter fullscreen mode Exit fullscreen mode


Figure 4: Acceptance rate climbing from 71% to 88.5% across 10 scans as trap hits collapse with Hindsight agent memory.

Key Observations

  1. The Baseline Plateau: The baseline agent without memory hovered consistently around 71% acceptance. Across all 10 scans, it generated an average of 12 to 14 trap hits per scan, repeatedly asking engineers to delete the same standby replicas under different resource IDs.
  2. Rapid Convergence: With Hindsight, acceptance jumped from 72.4% to 84.2% on Scan 2, and leveled off at 88.5% by Scan 5.
  3. Look-Alike Generalization: In Scan 3, an untagged standby replica named db-backup-standby-useast1 was analyzed. Instead of flagging it as idle waste, the agent recalled the structural pattern of an idle node attached downstream of a primary datastore, outputting:
{
  "decision": "suppress",
  "action": "none",
  "confidence": 0.94,
  "predicted_approval": 0.05,
  "reason": "Topological standby for primary database. Matches learned rule obs_8f21 regarding idle DR replicas.",
  "cited_memory_ids": ["obs_8f21"],
  "risk_note": "high blast radius: attached to primary VPC tier"
}
Enter fullscreen mode Exit fullscreen mode

The recommendation was suppressed upfront before triggering an alert in Slack.


Lessons Learned

Building production agent memory for infrastructure revealed several practical takeaways:

  1. Consolidation is Asynchronous by Design: Ingesting memories via retain is fast, but background consolidation and mental model synthesis take time. In our web interface and Slack bots, we implemented an explicit "learning" state for newly submitted verdicts. Attempting to force synchronous consolidation blocks API threads and ruins user responsiveness.
  2. Topological Hints Beat Keyword Matching: Engineers rarely maintain complete tag hygiene. If your memory queries rely solely on key-value tags like env=production, your agent remains blind to half your estate. Passing dependency graph relationships (depends on: payment-db, role hints: standby) into Hindsight recall queries allows the system to cluster patterns based on structural topology rather than brittle strings.
  3. Directives Must Preempt Prompts: Hard organizational rules (e.g., "Never terminate resources with compliance tags") should never be left to LLM discretion inside a prompt. We enforce directives as deterministic pre-filters in Python before the memory retrieval and LLM stages are ever reached.
  4. Invalidate Provenance Chains, Never Patch Vector Embeddings: When human feedback turns out to be wrong, never try to surgically edit synthesized rules. Invalidating the underlying source verdicts and allowing the memory engine to re-derive the observation graph preserves evidentiary integrity and eliminates ghost rules.

Cloud cost management is fundamentally an organizational trust problem. By pairing deterministic infrastructure graphs with persistent agent memory, we stopped asking engineers to defend their architecture over and over again.

Built with support from Code.in.

Top comments (0)