In every cloud environment I have ever audited, infrastructure tags are a polite fiction.
Site Reliability Engineering mandates that every EC2 instance, EBS volume, and RDS database carry env, owner, and cost-center tags. But at 2:00 AM on a Sunday when a database goes down, someone spins up an untagged instance named temp-fix-db or leaves an unattached 500GB volume after a manual snapshot.
Traditional AWS cost-cutting tools rely entirely on tags and naive utilization thresholds. If an instance sits at 0.5% CPU and carries no tags, the tool flags it as abandoned waste and asks to delete it.
The problem? That untagged, flatlined instance is often a standby database replica in an availability zone failover group, or an off-hours forensic archive kept for SOC2 compliance.
When an autonomous bot attempts to terminate a disaster recovery node, it doesn't just annoy developers-it threatens production stability.
To solve this, my teammate and I built CloudSense. While our initial goal was automated cost governance, we quickly realized that solving cloud waste requires solving infrastructure context. We designed an engine that pairs directed graph topology and behavioral traffic analysis with Hindsight agent memory.
By feeding structural role hints and transitive dependency graphs into Vectorize agent memory, we taught an LLM agent to recognize what an untagged resource actually does-and remember engineer feedback across multiple AWS accounts.
Here is how we built the graph reasoning engine, how it integrates with Hindsight, and what we learned evaluating it against adversarial cloud infrastructure.
The Core Technical Challenge: The "Tag Lie" and Look-Alikes
When engineers reject automated cost recommendations, their reasons rarely map to simple static metadata:
- "Don't terminate
i-09ab7c1-it's the warm DR standby for payments." - "Leave that unattached volume alone; it's our frozen Q3 audit snapshot."
- "This large instance is dormant for 27 days, but runs payroll batches on the 28th."
A standard vector database or RAG pipeline fails here. If an engineer in Account A rejects a recommendation on i-09ab7c1, storing that text string helps nobody when the same team spins up an untagged look-alike in Account B named ip-10-0-12-88.ec2.internal. Cosine similarity on instance IDs and generic AWS attributes yields virtually nothing.
To bridge this gap, our agent needed three capabilities:
- Topological Reconstruction: Automatically mapping dependencies (load balancers, security groups, block devices, replication pairs) without relying on tags.
- Behavioral Traffic Shaping: Detecting whether a resource is truly dead or exhibits periodic burst patterns (e.g., month-end batches).
- Associative Agent Memory: Storing structural role hints in an evolving memory bank that consolidates individual human verdicts into an organizational playbook.
System Architecture: From Graph Extraction to Memory Recall
The decision pipeline operates as a closed-loop system connecting AWS infrastructure, graph algorithms, persistent memory banks, and human engineering feedback:
flowchart TD
subgraph Discovery ["1. Topology & Signal Discovery"]
A[AWS Read-Only Scanner] --> B[EC2, EBS, RDS, EIP Scan]
B --> C[CloudWatch 14d Time-Series]
end
subgraph Graph ["2. Dependency & Blast Radius Engine"]
C --> D[Directed Dependency Graph]
D --> E[BFS Transitive Blast Radius]
D --> F[Traffic Shaping & Role Hints]
end
subgraph Memory ["3. Hindsight Memory Subsystem"]
F --> G{Agent Decision Engine}
H[(Hindsight Org Bank)] -->|Associative Recall| G
I[Org Playbook Mental Model] -->|Injected System Context| G
end
subgraph Feedback ["4. Human-in-the-Loop & Consolidation"]
G -->|Recommend / Ask| J[Slack & Web UI]
G -->|Suppress & Cite Memory| K[Suppressed Audit Log]
J -->|Human Verdict: Reject + Reason| L[Retain to Hindsight]
L -->|Asynchronous Consolidation| H
H -->|Delta Trigger| I
end

Figure 1: The CloudSense review queue displaying candidate recommendations with calculated blast radius, estimated savings, recalled memory citations, and human approval controls.
Engineering the Graph and Role Hint Engine
Before making any LLM call, CloudSense builds an in-memory directed dependency graph of the target AWS environment and extracts topological role hints.
1. Reconstructing the Infrastructure Graph
Resources in AWS are deeply entangled. An EBS volume depends on an EC2 instance; an EC2 instance depends on target groups and load balancers; an RDS read-replica depends on its primary source.
More subtly, Security Groups form implicit dependencies. If Security Group A allows inbound traffic from Security Group B, any resource in B depends on A.
Here is the graph reconstruction logic from our codebase:
# backend/graph/build.py
def build_graph(resources: list[Resource]) -> dict[str, list[str]]:
nodes = {r.id: r for r in resources}
edges: set[tuple[str, str]] = set() # (dependent, dependency)
for r in list(resources):
attrs = r.attrs
if r.type == "ebs" and attrs.get("attached_to"):
edges.add((r.id, attrs["attached_to"]))
if r.type == "rds" and attrs.get("replica_source"):
src = attrs["replica_source"]
if src in nodes:
edges.add((r.id, src))
if r.type == "ec2":
if asg := attrs.get("asg"):
edges.add((r.id, f"asg:{asg}"))
for tg in attrs.get("target_groups", []):
edges.add((r.id, tg["arn"]))
for lb_arn in tg.get("load_balancer_arns", []):
edges.add((tg["arn"], lb_arn))
# Security Group cross-references: if A allows ingress from B,
# members of B depend on members of A.
sg_members: dict[str, list[str]] = {}
for r in nodes.values():
for sg in r.attrs.get("security_groups", []):
sg_members.setdefault(sg, []).append(r.id)
for r in nodes.values():
for sg in r.attrs.get("sg_refs", []):
for client in sg_members.get(sg, []):
edges.add((client, r.id))
return {rid: [dep for dep, src in edges if src == rid] for rid in nodes}
2. Transitive Blast Radius via BFS
If CloudSense recommends stopping an instance, what downstream services will experience a failure? We compute the transitive blast radius using breadth-first search:
# backend/graph/blast_radius.py
def blast_radius(graph: dict[str, list[str]], start_id: str) -> list[str]:
seen = {start_id}
order: list[str] = []
queue = deque([start_id])
while queue:
for dep in graph.get(queue.popleft(), []):
if dep not in seen:
seen.add(dep)
order.append(dep)
queue.append(dep)
return order
If the transitive blast radius exceeds a critical safety threshold, CloudSense automatically flags the decision as "ask", preventing autonomous mutations on high-risk nodes.
3. Extracting Role Hints and Behavioral Traffic Shapes
Instead of relying on developer tags, we look at structural properties and CloudWatch time-series history:
# backend/graph/build.py
def role_hints(resource: Resource, graph: dict[str, list[str]], names: dict[str, str]) -> list[str]:
hints: list[str] = []
if src := resource.attrs.get("replica_source"):
hints.append(f"replica_of:{names.get(src, src)}")
if resource.attrs.get("asg"):
hints.append("in_asg")
if resource.attrs.get("target_groups"):
hints.append("behind_lb")
# Traffic Shaping: Detect flatline vs bursty periodic batches
m = resource.metrics
if m.cpu_max is not None and m.cpu_max < 1.0 and (m.db_connections_max or 0) == 0:
hints.append("flat_zero")
series = m.daily_cpu_series
if len(series) >= 5:
med = statistics.median(series)
high = [v for v in series if v >= max(med * 3.0, med + 20)]
if 0 < len(high) <= 4:
hints.append("periodic_spikes") # classic month-end payroll pattern
return hints
These generated hints (replica_of:..., in_asg, flat_zero, periodic_spikes) form a persistent fingerprint for workloads that developer tags fail to describe.
Wiring Graph Topology into Hindsight Agent Memory
Once structural role hints and blast-radius data exist, we pass them directly into Hindsight.
1. Graph-Aware Memory Retain
When an engineer rejects a proposed action in Slack, we retain the decision. Rather than storing just the raw text comment, we bind the resource's structural graph neighbors as entities and register hierarchical observation_scopes:
# backend/memory/types.py
def verdict_item(scan_id: str, cand: Candidate, res: Resource, verdict: Any, reviewer: Any) -> dict[str, Any]:
return {
"content": verdict_content(scan_id, cand, res, verdict, reviewer),
"context": "Human engineer review verdict on AWS cost recommendation",
"timestamp": verdict.at or datetime.now(UTC),
"tags": [f"account:{res.account_alias}", f"team:{res.owner_team}", f"rtype:{res.type}"],
"metadata": {
"candidate_id": cand.id,
"resource_id": res.id,
"decision": str(verdict.decision),
},
"entities": [
{"text": res.display_name},
*[{"text": name} for name in res.depends_on_names + res.dependent_names],
],
"observation_scopes": [
[], # Global tenant scope (shares across all accounts)
[f"account:{res.account_alias}"], # Account scope
[f"team:{res.owner_team}"], # Team scope
],
}

Figure 2: Human-in-the-loop review interface: engineers specify operational reasons and define observation scopes to teach Hindsight how to generalize.
The empty inner list [] in observation_scopes tells Hindsight that this architectural insight belongs to the universal organizational scope. If Team Payments marks a failover database pattern as protected in us-east-1, the rule transfers seamlessly to an untagged standby in eu-central-1.
2. High-Dimensional Associative Recall
When CloudSense scans candidate resources on subsequent runs, it queries Hindsight with graph-aware parameters:
# backend/memory/client.py
async def recall_for(self, org: Any, cand: Candidate, res: Resource) -> list[MemoryHit]:
# Construct recall query combining action, role hints, and topological neighbors
query = (
f"{cand.action} {res.type} {res.display_name}; "
f"role hints: {res.role_hints}; signals: {cand.signals_text}; "
f"connected to: {', '.join(res.neighbor_names)}"
)
resp = await self.hs.arecall(
bank_id=f"org-{org.id}",
query=query,
tags=[f"account:{res.account_alias}"],
tags_match="any",
max_tokens=1500,
)
return [_hit(r) for r in resp.results]
When an untagged, flatlined EC2 instance is evaluated, Hindsight matches its query not by its arbitrary random ID, but by its topological role hints (role hints: ['flat_zero', 'replica_of:prod-db-01']).
The agent immediately recognizes the workload as a protected standby, issues a decision="suppress", and cites the exact memory ID of the human verdict that taught it this rule.

Figure 3: The Learned Rules dashboard in CloudSense: consolidated organizational rules derived by Hindsight with source fact provenance and proof counts.
Evaluation Benchmark: Testing Against Adversarial Traps
To test whether topological memory prevents dangerous recommendations, we evaluated CloudSense against an offline benchmark (eval/) simulating multi-account AWS topologies across 10 sequential scan iterations.
The benchmark injects six difficult "trap archetypes" designed to confuse standard cost optimizers:
- Disaster Recovery Standby: Idle instances that mirror production topology but see zero steady-state traffic.
- Payroll Batch Worker: High-memory machines idling for 27 days before spiking for 24 hours.
- Audit Archive: Detached EBS volumes retained for SOC2 compliance.
- Blue/Green Standby: Fully configured environments awaiting deployment cutovers.
- License-Pinned Host: Dedicated single-tenant servers running legacy enterprise binaries.
- On-Call Bastion: Low-traffic SSH bastion nodes used for production incident triage.
Between consecutive scans, these trap resources reappear under randomized names, in new accounts, and completely stripped of tags.
Experimental Results
We compared a baseline agent (relying solely on utilization metrics and local heuristics) against CloudSense backed by Hindsight memory:

Figure 4: Benchmark performance across 10 sequential scans on multi-account AWS topologies containing untagged trap archetypes.
| Scan Iteration | Baseline Acceptance Rate | Hindsight Acceptance Rate | Baseline Trap Hits (Alerts) | Hindsight Trap Hits (Alerts) |
|---|---|---|---|---|
| Scan 1 (Cold Start) | 72.4% | 72.4% | 12.0 | 12.0 |
| Scan 2 | 71.1% | 84.2% | 13.0 | 5.3 |
| Scan 3 | 68.1% | 86.3% | 14.3 | 4.3 |
| Scan 4 | 73.7% | 88.1% | 11.3 | 4.0 |
| Scan 10 (Steady State) | 72.4% | 88.3% | 12.3 | 4.0 |
Acceptance Rate Progression:
Baseline: [72%] ---> [71%] ---> [68%] ---> [73%] ---> [72%] (Flatline)
Hindsight: [72%] ---> [84%] ---> [86%] ---> [88%] ---> [88%] (+16% Gain)
False-Positive Trap Hits:
Baseline: 12.0 ---> 13.0 ---> 14.3 ---> 11.3 ---> 12.3 (Continuous Noise)
Hindsight: 12.0 ---> 5.3 ---> 4.3 ---> 4.0 ---> 4.0 (-66.7% Noise Reduction)
What the Data Proves
- The Heuristic Ceiling: Without memory, the baseline agent is trapped at ~71% accuracy. It continues to flag the same standby instances and compliance archives every single cycle, spamming engineers with 12 to 14 bad recommendations.
- Rapid Generalization: Within two scan iterations, Hindsight consolidates isolated human rejections into active observations. Acceptance jumps to 84.2% on Scan 2 and plateaus at 88.3% by Scan 4.
- Tag-Independent Protection: Trap hits dropped by 66.7% even though the look-alike resources had randomized identifiers. The agent successfully matched topological role hints and traffic profiles against the organizational mental model.
Four Hard-Earned Engineering Lessons
Building this graph-memory bridge surfaced several non-obvious distributed systems lessons:
1. Dependency Graph Inversion for Security Groups
In AWS, Security Groups define inbound permissions. If Security Group A allows ingress from Security Group B, engineers naturally think "A connects to B." But in an infrastructure dependency graph, the relationship is inverted: B depends on A. If you terminate the instance in A, the services running in B lose their backend connection. Getting this edge orientation right was essential for calculating true blast radius.
2. Multi-Tenant Scoping with observation_scopes
In early prototypes, a rule learned in one AWS account failed to protect resources in another. We discovered that in the hindsight-client Python SDK, aretain_batch accepts a list of scopes. Passing [[], [f"account:{res.account_alias}"]] creates a dual-scoped memory: the inner empty list [] permits global organizational inheritance, while the account tag maintains local fidelity.
3. Curation via Source Invalidation
LLMs can occasionally learn incorrect rules if an engineer gives a flawed rejection reason. Because Hindsight's observations are derived artifacts, there is no direct endpoint to delete a learned rule. Instead, you invalidate its source facts using UpdateMemoryRequest(state="invalidated"). Hindsight's asynchronous consolidation engine then cleanly purges or recalculates the derived rules while maintaining an immutable audit log.
4. Untagged Mental Models Avoid Query Filtering
We initially tagged our organizational playbook mental model with tags=["playbook"]. However, Hindsight treats tagged mental models with strict tag-matching rules during recall. When our agent recalled memories with account-specific tags (account:prod-us-east), the playbook was filtered out of the response. Keeping the central playbook mental model untagged ensured it was always delivered as foundational system context.
Conclusion: Context is the Missing Link in AI Automation
The failure of early cloud optimization agents was never about LLM reasoning capability. It was about context blindness.
When an AI system operates without architectural awareness and without memory, every scan starts from zero. It asks the same questions, makes the same mistakes, and gradually destroys the trust of the engineers it was meant to assist.
By combining directed graph topology with Hindsight agent memory, we built a system that listens, remembers, and respects the unwritten rules of enterprise infrastructure.
Special thanks to @Code.in for supporting builders in the autonomous agent ecosystem.
For complete implementation details, visit the Hindsight documentation and the CloudSense repository.
Top comments (0)