Connecting Honeycomb Observability & SLOs into an AI Graph Memory with Cognee
When production goes down or latency spikes at 2 AM, autonomous SRE agents and on-call engineers need answers fast:
- "Which alert triggers monitor 5xx errors on the checkout service?"
- "What is our SLO error budget target for payments API and what queries measure it?"
- "Which dashboards visualize database lock wait durations?"
All of this operational knowledge lives inside Honeycomb. To make that metadata instantly queryable by AI agents and LLM tools, we built and contributed the Honeycomb Connector for Cognee.
1. Why Observability Belongs in a Knowledge Graph
Rather than indexing raw high-velocity span telemetry (which belongs in high-throughput timeseries storage), high-leverage contextual intelligence lives in Honeycomb’s configuration layer:
- Datasets: The telemetry boundaries and services being monitored.
- Boards / Dashboards: Curated graphs tracking p99 latencies, error budgets, and spans.
- Triggers: Threshold rules, evaluation frequencies, and runbook definitions.
- Service Level Objectives (SLOs): Target reliability percentages (e.g. 99.95%) and SLI formulas.
flowchart TD
HC[Honeycomb Management API] -->|Datasets / Boards / Triggers / SLOs| Client[HoneycombClient]
Client -->|yield Structured Markdown| DLT[DLT Pipeline]
DLT -->|DOCUMENT_SOURCE_ATTR| Cognee[Cognee Cognify]
Cognee -->|Build Topology & Dependency Graph| Memory[(Graph DB + Vector Store)]
Memory -->|Natural Language & Incident Automation| Agent[Autonomous SRE Agent]
2. Connector Architecture
Clean Markdown Entity Serialization
The connector normalizes triggers, SLO contracts, and board layouts into rich structured Markdown:
# Honeycomb Trigger: High 5xx HTTP Error Rate
- **Trigger ID:** `trig_api_5xx`
- **Dataset:** `api-gateway`
- **Status:** Active
- **Evaluation Frequency:** Every 60s
- **Threshold Rule:** Metric > 2.0
- **Last Updated:** 2026-10-02T11:00:00Z
### Description & Runbook
Fires when the percentage of 5xx HTTP responses exceeds 2% over a 5-minute rolling window. Check upstream ingress proxy health.
# Honeycomb Service Level Objective (SLO): Checkout Availability
- **SLO ID:** `slo_checkout_v1`
- **Dataset:** `payments-svc`
- **Target Reliability:** 99.9% over 30 days
- **SLI Metric Definition:** `checkout_success_rate`
### Target Rationale & Scope
99.9% of checkout requests must return HTTP 200 within 1000ms.
When Cognee processes this, it builds an operational knowledge graph linking Services, SLOs, Triggers, and Dashboards.
3. Quickstart Example
import asyncio
import cognee
from cognee_community_connector_honeycomb import honeycomb_source
async def main():
source = honeycomb_source(
api_key="YOUR_HONEYCOMB_API_KEY",
include_datasets=True,
include_boards=True,
include_triggers=True,
include_slos=True,
)
# Ingest and Cognify
await cognee.add(source, dataset_name="infra_knowledge")
await cognee.cognify(dataset_name="infra_knowledge")
# Ask operational questions
answers = await cognee.search(
"What are the alerting rules and runbooks for high 5xx error rates on api-gateway?",
dataset_name="infra_knowledge"
)
print(answers)
if __name__ == "__main__":
asyncio.run(main())
4. Offline Unit Testing
The package includes a comprehensive mock test suite verifying header authentication (X-Honeycomb-Team), incremental timestamp filtering (since), and pipeline extraction:
uv run pytest packages/connector/honeycomb/tests/test_honeycomb.py -v
Resources
- Pull Request: topoteretes/cognee-community#237
- Monorepo: github.com/topoteretes/cognee-community
Top comments (0)