A memory agent is only as effective as the underlying data model it uses to represent historical knowledge. When converting unstructured incident histories, post-mortems, and CI/CD error outputs into vector memory banks, raw unstructured text can lead to poor retrieval precision.
To solve this, our team engineered the ingestion pipeline in seed_memories.py to structure historical outage logs before seeding them into our Hindsight memory bank (devops-pipeline-agent).
Data Modeling for Incident Logs
Rather than storing raw, unformatted error messages, each historical incident memory is structured into a normalized pair:
-
Error Signature (Query Anchor): Core log pattern, error code (e.g.,
Exit Code 137,ENOSPC), or trace signature stripped of transient identifiers like timestamps or specific process IDs. - Remediation Context (Resolution Payloads): Concrete steps, configuration flags, or shell commands required to fix the underlying issue.
+-----------------------------------------------------------+
| Historical Incident |
+-----------------------------------------------------------+
|
Parse & Normalize Payload
|
v
+-----------------------------------------------------------+
| Content: "Error: Exit Code 137 (OOM Killed) |
| Resolution: Increase memory limits in pipeline" |
+-----------------------------------------------------------+
|
Hindsight Retention API
|
v
+-----------------------------------------------------------+
| Hindsight Bank: devops-pipeline-agent |
+-----------------------------------------------------------+
Seeding Historical Memories (seed_memories.py)
The seeding script reads historical records, formats them into dense representations, and commits them to the memory bank via the Hindsight SDK:
import os
from dotenv import load_dotenv
from hindsight_client import Hindsight
load_dotenv()
client = Hindsight(
base_url=os.getenv("HINDSIGHT_API_URL"),
api_key=os.getenv("HINDSIGHT_API_TOKEN")
)
BANK_ID = os.getenv("HINDSIGHT_BANK_ID", "devops-pipeline-agent")
HISTORICAL_INCIDENTS = [
{
"content": "Error: ENOSPC: no space left on device, write\nResolution: Run 'docker system prune -af --volumes' to clear unused Docker caches.",
"category": "disk_space"
},
{
"content": "Error: Command failed with exit code 137\nResolution: Out of Memory error. Increase runner memory in .gitlab-ci.yml or build config.",
"category": "resource_exhaustion"
}
]
def seed_bank():
print(f"🌱 Seeding memory bank: {BANK_ID}...")
for incident in HISTORICAL_INCIDENTS:
client.retain(
bank_id=BANK_ID,
content=incident["content"],
metadata={"category": incident["category"], "type": "historical_seed"}
)
print("✅ Memory bank seeding complete!")
if __name__ == "__main__":
seed_bank()
Key Data Modeling Strategies
- Noise Filtering: Strip out build IDs, specific commit hashes, and timestamps prior to vector embedding to prevent low-similarity scoring on reoccurring errors.
- Structured Document Pairing: Explicitly separate the error signature from the resolution in the content string so vector embeddings retain both problem context and solution context.
-
Metadata Tagging: Attach category metadata (
resource_exhaustion,disk_space) for fine-grained filtering during retrieval operations.
Resources & Links
-
GitHub Repository:
25wh1a05be/devops-pipeline-agent -
Hindsight Documentation:
hindsight.vectorize.io
Top comments (1)