Most cloud cost optimization tools are glorified cron jobs with an API key. They scan AWS every night, find twenty "idle" EC2 instances and unattached EBS volumes, dump a list of recommendations into a Slack channel, and pat themselves on the back for identifying \$4,000 in monthly savings.
And within two weeks, the engineering team mutes the channel.
The tool doesn't know that the "idle" t3.large is the disaster-recovery standby for the payment gateway. It doesn't know that the "oversized" memory worker only spikes during month-end payroll calculations. It doesn't know that the unattached 500GB gp3 volume is a frozen forensic snapshot for SOC2 auditors. Every time an automated bot recommends deleting an essential workload, it burns trust. After three false alarms, developers ignore every recommendation-including the legitimate ones.
Cost tools tell you what to delete. We wanted to build a system that learns what you will never let it delete.
We built CloudSense, an autonomous AWS cost governance system that turns engineer rejections-every "No, because that’s the replica standby"-into durable, evidence-backed rules stored in Hindsight. Instead of repeatedly asking humans about the same architectural traps, CloudSense uses Vectorize agent memory to retain review decisions, consolidate them into an organizational playbook, and generalize those patterns to untagged look-alikes across multiple AWS accounts.
Here is how the system hangs together, what we learned hooking an LLM to persistent memory, and the real numbers from evaluating it against adversarial cloud environments.
The System Architecture
CloudSense is structured as an event-driven loop that separates read-only discovery, topological analysis, memory recall, and human-in-the-loop review.
flowchart TD
subgraph S1["1. Cloud Discovery"]
A["AWS Accounts (Read-Only Role)"] -->|Describe + 14d CloudWatch| B["Resource Scanner"]
B --> C["Dependency Graph (Role Hints & Blast Radius)"]
C --> D["Candidate Generator (Rules R1–R8 + Pricing)"]
end
subgraph S2["2. Memory & Decision Engine"]
D --> E{"Agent Decision Engine"}
M[("Hindsight Memory Bank (Per Org)")] -->|Parallel Recall + Playbook| E
E -->|Recommend / Ask| F["Slack & Dashboard Review"]
E -->|Auto-Suppress + Cite| S["Suppressed Candidate Log"]
end
subgraph S3["3. Feedback & Learning Loop"]
F -->|Approve / Reject + Reason| G["Human Verdict"]
G -->|Retain Verdict| M
M -->|Asynchronous Consolidation| O["Observations (Learned Rules)"]
O -->|Delta Refresh| P["Playbook Mental Model"]
G -->|Approved| X["Executor (Snapshot → Act → Undo)"]
end
The pipeline runs across six distinct stages:
- Scanner: Assumes an external read-only IAM role into each target AWS account, scanning EC2 instances, EBS volumes, RDS databases, NAT gateways, and unattached Elastic IPs alongside 14-day CloudWatch percentile metrics (P95 CPU, network throughput, disk IOPS).
-
Topological Dependency Graph: Cloud resources do not live in isolation. CloudSense constructs a directed dependency graph linking instances to their VPCs, security groups, subnets, and attached block devices, computing a blast-radius score and extracting structural "role hints" (e.g.,
dr-standby,batch-worker,ingress-proxy). - Candidate Generator: Rules R1 through R8 flag baseline waste (such as idle compute under 2% P95 CPU, unattached EBS volumes older than 7 days, or unreferenced EIPs) and calculate estimated monthly dollar savings using AWS pricing tables.
- Agent Decision Engine: Before any candidate reaches a developer, an LLM agent evaluates it. The agent pre-filters hard directives, queries Hindsight for memories of past human verdicts on similar resources, and checks an organizational playbook mental model. It decides whether to recommend (publish to review), suppress (silently drop because an engineer previously explained why this workload is needed), or ask (flag high-blast-radius or ambiguous cases for senior attention).
- Human Review Loop: Recommendations arrive via Slack interactive buttons and a web dashboard. When an engineer rejects a recommendation, they provide a brief reason.
- Memory Retention & Playbook Refresh: That verdict is asynchronously retained into Hindsight. Over time, Hindsight consolidates isolated rejections into generalized rules ("Never terminate database read-replicas in us-east-1 without dedicated standby tags"), refreshing the agent's mental model for future scans.

Figure 1: The CloudSense review queue displaying candidate recommendations with calculated blast radius, estimated savings, recalled memory citations, and human approval controls.
Why Agent Memory is Harder Than RAG
A naive approach to this problem is standard Retrieval-Augmented Generation (RAG): store past Slack rejection comments in a vector database, embed the candidate resource metadata, perform a cosine-similarity search, and stuff the top-3 results into the prompt.
We tested this early on, and it broke down immediately:
-
Semantic drift across accounts: A rejection on
i-0a1b2c3dnamedprod-pay-worker-01in Account A has near-zero text embedding similarity to an untagged instance namedip-10-0-4-12.ec2.internalin Account B, even though both share the exact same role hints, VPC topology, and bursty month-end CloudWatch profile. - No distinction between facts and rules: Vector databases store static text chunks. They cannot synthesize twenty individual rejections into an abstracted organizational policy or track confidence over time.
-
Cold-start and scope isolation: An infrastructure rule might apply globally across the entire company, or it might apply strictly to the
paymentsteam, or only to a staging account. Flat vector indices lack scope hierarchies.
This is why we chose Hindsight. Hindsight treats memory not as a key-value store or vector collection, but as an evolving cognitive bank supporting tenants, missions, behavioral dispositions, hard directives, entity tracking, observations, and structured mental models.
Implementing Persistent Memory in Code
Let's look at how this works in practice.
1. Bank Initialization with Skepticism and Directives
Every organization gets an isolated memory bank with an explicit operational mission and behavioral dispositions. In automated cloud governance, an aggressive agent that accidentally deletes a production database is fatal. We configure Hindsight with high skepticism and high literalism:
# backend/memory/setup.py
DISPOSITION = {
"disposition_skepticism": 5,
"disposition_literalism": 4,
"disposition_empathy": 2,
}
def mission(org: Any) -> str:
return (
f"I review AWS cost-saving recommendations for {org.name}. "
"Breaking production is worse than missing a saving."
)
async def ensure_bank(hs: Hindsight, org: Any) -> None:
bid = bank_id(org)
# Idempotent PUT: creates or updates bank with tenant mission & disposition
await hs.acreate_bank(bank_id=bid, name=org.name, mission=mission(org), **DISPOSITION)
# Org-level hard directives override everything else
existing = {d.content for d in (await hs.alist_directives(bank_id=bid)).items}
for i, text in enumerate(directives(org)):
if text not in existing:
await hs.acreate_directive(bank_id=bid, name=f"hard-rule-{i + 1}", content=text)
Directives in Hindsight act as non-negotiable boundaries. If an engineer specifies "Never terminate or modify resources tagged do_not_terminate=true", that directive is registered in the bank and enforced in code before any LLM inference occurs.
2. Retaining Human Verdicts with Hierarchical Scopes
When a senior site reliability engineer rejects a recommendation in Slack, CloudSense packages the full context: the resource attributes, CloudWatch metrics, topological neighbors, the reviewer's identity, and the reviewer's reason.
Critically, we retain this using observation_scopes across three tiers:
- Global/Shared (
[]): The rule can generalize across all accounts in the organization. - Account-Specific (
[account:prod-us-east]): The rule applies to this AWS account. - Team-Specific (
[team:data-eng]): The rule applies to workloads owned by this team.
# backend/memory/types.py & backend/memory/client.py
def verdict_item(scan_id: str, cand: Candidate, res: Resource, verdict: Any, reviewer: Any) -> dict[str, Any]:
return {
"content": verdict_content(scan_id, cand, res, verdict, reviewer),
"context": "Human engineer review verdict on AWS cost recommendation",
"timestamp": verdict.at or datetime.now(UTC),
"tags": [f"account:{res.account_alias}", f"team:{res.owner_team}", f"rtype:{res.type}"],
"metadata": {
"candidate_id": cand.id,
"resource_id": res.id,
"decision": str(verdict.decision),
"reviewer_id": str(reviewer.id),
},
"entities": [
{"text": res.display_name},
*[{"text": name} for name in res.depends_on_names + res.dependent_names],
],
"observation_scopes": [
[], # Global org scope (untagged)
[f"account:{res.account_alias}"], # Account scope
[f"team:{res.owner_team}"], # Team scope
],
}
async def retain_verdict(self, org: Any, scan_id: str, cand: Candidate, res: Resource, verdict: Any, reviewer: Any) -> str | None:
resp = await self.hs.aretain_batch(
bank_id=bank_id(org),
items=[verdict_item(scan_id, cand, res, verdict, reviewer)],
document_id=f"verdict-{cand.id}",
retain_async=True,
)
return resp.operation_id

Figure 2: Human-in-the-loop review interface: engineers specify operational reasons and define the observation scope to instruct Hindsight how widely the rule should generalize.
3. Parallel Recalls and the Live Playbook
During a scan, a single account might generate 40 raw cost-saving candidates. Sequential memory lookups would destroy scan performance. We batch memory recalls concurrently using an asyncio.Semaphore, combining the resource's structural role hints and topological graph dependencies into the recall query:
# backend/memory/client.py
async def recall_for(self, org: Any, cand: Candidate, res: Resource) -> list[MemoryHit]:
# Query combines action, resource type, role hints, metrics, and graph connections
query = (
f"{cand.action} {res.type} {res.display_name}; "
f"role hints: {res.role_hints}; signals: {cand.signals_text}; "
f"connected to: {', '.join(res.neighbor_names)}"
)
resp = await self.hs.arecall(
bank_id=bank_id(org),
query=query,
tags=[f"account:{res.account_alias}"],
tags_match="any",
max_tokens=1500,
)
return [_hit(r) for r in resp.results]
At the same time, we fetch Hindsight's consolidated org-playbook mental model. When individual verdicts consolidate, Hindsight derives high-level rules such as:
"Workloads attached to the
db-failoversubnet that show zero traffic during normal days are standby disaster recovery nodes. Do not terminate or rightsize."
This playbook is injected into the agent's system prompt. When the agent inspects a candidate, if a recalled memory or playbook rule indicates the resource is protected, it outputs decision="suppress" and explicitly cites the cited_memory_ids backing its decision.
# backend/agent/decide.py
async def decide_one(llm: LLM, system: str, res: Resource, cand: Candidate, memories: list[MemoryHit]) -> AgentDecision:
prompt = build_candidate_prompt(res, cand, memories)
messages = [{"role": "system", "content": system}, {"role": "user", "content": prompt}]
raw = await llm.complete_json(messages)
decision = AgentDecision.model_validate_json(raw)
# Validation safety: strictly drop any memory hallucinated by the LLM
allowed = {m.id for m in memories}
decision.cited_memory_ids = [i for i in decision.cited_memory_ids if i in allowed]
return decision

Figure 3: The Learned Rules view in CloudSense: consolidated organizational rules derived by Hindsight with source fact provenance and proof counts.
Evaluating the Impact: Does Memory Actually Help?
To rigorously measure whether agent memory solves the trust problem, we built an offline evaluation benchmark (eval/) simulating an enterprise AWS organization across 10 sequential scans and 3 separate accounts.
The environment seeds six specific "trap archetypes" representing legitimate low-utilization resources that traditional cost tools erroneously delete:
- DR Standby: Hot-standby replica instances that receive negligible traffic during steady state.
- Payroll Batch Worker: High-memory instances that sit idle for 28 days and spike on the 29th.
- Audit Archive: Detached EBS volumes preserved for compliance records.
- Blue/Green Standby: Fully provisioned environments idling between release cycles.
- License-Pinned Host: Dedicated EC2 hosts running legacy single-tenant licensed software.
- On-Call Debug Host: Bastion hosts with low continuous utilization used for incident triage.
To test generalization, across consecutive scans these trap workloads reappear under different names, in different accounts, and without standard tags.
We ran two comparative pipelines through 10 full scan iterations:
- Baseline (No Memory): Standard LLM agent with utilization signals, pricing rules, and graph heuristics, but zero persistence between scans.
- Memory-Augmented (Hindsight): The same agent backed by Hindsight retaining human rejection verdicts and updating its playbook.
The Results

Figure 4: Acceptance rate across ten sequential scans comparing baseline execution against Hindsight memory augmentation.
| Metric | Scan 1 (Cold Start) | Scan 2 | Scan 4 | Scan 10 (Steady State) |
|---|---|---|---|---|
| Baseline Acceptance Rate | 72.4% | 71.1% | 73.7% | 72.4% |
| Hindsight Acceptance Rate | 72.4% | 84.2% | 88.1% | 88.3% |
| Baseline Trap Hits (Alert Noise) | 12.0 | 13.0 | 11.3 | 12.3 |
| Hindsight Trap Hits (Alert Noise) | 12.0 | 5.3 | 4.0 | 4.0 |
The results show a clear behavioral divergence:
- The No-Memory Trap: Without memory, the acceptance rate is locked around 71%. The agent repeatedly flags the same standby instances and archive volumes scan after scan, generating 12 to 14 erroneous recommendations every cycle. The human engineer is forced to reject the exact same items repeatedly.
- Rapid Convergence with Hindsight: By Scan 2, Hindsight has consolidated the first batch of rejections. Acceptance jumps from 72.4% to 84.2%. By Scan 4, acceptance stabilizes at 88.3%, and false-positive trap alerts drop from 12 down to 4-a 66.7% reduction in alert noise.
-
Generalization Beyond Tags: The trap hits did not drop because the agent memorized resource IDs. The new workloads had randomized identifiers (e.g.,
i-09f482a...). The agent successfully suppressed them because Hindsight's recall matched the topological role hints (subnet attachments, VPC peering, and CloudWatch flatline profiles) to the consolidated rules in the mental model.
Four Real Engineering Lessons from the Trenches
Integrating an autonomous agent with persistent memory revealed several non-obvious engineering realities. Here are four lessons learned from our production implementation:
1. retain() vs aretain_batch() for Observation Scoping
In the hindsight-client SDK (version 0.10.1), the basic retain() call does not accept an observation_scopes argument. If you want rules to cross boundaries-such as allowing an architectural rule learned in the us-east-1 staging account to protect identical infrastructure in eu-west-1 production-you must use aretain_batch(). Passing [] as an inner list in observation_scopes: [[], [account:prod]] tells Hindsight to register the memory in the global, untagged tenant scope while simultaneously binding it to the specific account.
2. Curating Poisoned Rules via Source Invalidation
What happens if an engineer accidentally rejects a valid recommendation with bad reasoning (e.g., "Don't delete this, Bob needs it" when Bob left the company six months ago)? In Hindsight, observations (the learned rules) are derived cognitive artifacts-there is no delete_observation endpoint.
To safely invalidate a bad rule without breaking the integrity of the bank, we inspect the observation's source_memory_ids and invalidate the underlying source facts:
# backend/memory/client.py
async def delete_rule(self, org: Any, memory_id: str, reason: str = "removed by admin") -> None:
bid = bank_id(org)
mem = await self.hs.memory.get_memory(bid, memory_id)
sources = getattr(mem, "source_memory_ids", None) or [memory_id]
for sid in sources:
await self.hs.memory.update_memory(
bid, sid, UpdateMemoryRequest(state="invalidated", reason=reason)
)
Once the source facts are marked invalidated, Hindsight's next consolidation cycle automatically cleans up or re-derives the dependent observations. It preserves full auditability while keeping the memory bank healthy.
3. Reflect Requires include_facts=True
CloudSense provides interactive "Why?" buttons in Slack and the web dashboard so engineers can audit agent reasoning. Hindsight's areflect endpoint generates structured narrative explanations. However, by default, areflect does not populate the based_on source citations in the response object. You must explicitly pass include_facts=True:
ans = await hs.areflect(bank_id=bid, query=question, include_facts=True)
# ans.based_on now includes the exact memories and mental model versions consulted
Without this flag, you get a natural language answer, but cannot render clickable UI chips linking back to the historical verdicts that justified the decision.
4. Mental Model Tag Strictness
When creating the organizational playbook mental model (acreate_mental_model), we initially tagged it with tags=["playbook"]. We discovered that in Hindsight, tagged mental models enforce strict matching against query tags during recall. If your candidate evaluation queries use account or team tags, a strictly tagged mental model can be accidentally filtered out. Keeping the core playbook mental model untagged guarantees that it acts as universal background context across all candidate evaluations.
The Path to Trustworthy Autonomous Agents
The difference between a tool that engineers abandon and a tool they rely on is not the intelligence of the base LLM. It is whether the system respects human feedback.
If an AI agent asks you to review its work, you expect it to remember your answer. By giving our cost engine durable memory with Hindsight, we transformed cloud optimization from a relentless stream of false alarms into a collaborative system that gets measurably sharper with every single review.
Special thanks to @Code.in for supporting builders in the autonomous agent ecosystem.
The full architecture, evaluation harness, and CloudFormation templates are open-source in the CloudSense repository.
Top comments (0)