<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: SOUMYAJIT GHOSH</title>
    <description>The latest articles on DEV Community by SOUMYAJIT GHOSH (@soumyajit_ghosh_b93618199).</description>
    <link>https://dev.to/soumyajit_ghosh_b93618199</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4109663%2F5e4df557-a80f-47dc-8c8c-d0c3df39de0f.jpg</url>
      <title>DEV Community: SOUMYAJIT GHOSH</title>
      <link>https://dev.to/soumyajit_ghosh_b93618199</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/soumyajit_ghosh_b93618199"/>
    <language>en</language>
    <item>
      <title>Giving AI Agents Pipeline Lineage Memory: How We Built an OpenLineage Connector for Cognee</title>
      <dc:creator>SOUMYAJIT GHOSH</dc:creator>
      <pubDate>Fri, 09 Oct 2026 21:08:44 +0000</pubDate>
      <link>https://dev.to/soumyajit_ghosh_b93618199/giving-ai-agents-pipeline-lineage-memory-how-we-built-an-openlineage-connector-for-cognee-1ood</link>
      <guid>https://dev.to/soumyajit_ghosh_b93618199/giving-ai-agents-pipeline-lineage-memory-how-we-built-an-openlineage-connector-for-cognee-1ood</guid>
      <description>&lt;h1&gt;
  
  
  Giving AI Agents Pipeline Lineage Memory: How We Built an OpenLineage Connector for Cognee
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Author: Soumyajit Ghosh (&lt;a href="https://github.com/somuai" rel="noopener noreferrer"&gt;@somuai&lt;/a&gt;)&lt;/em&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Target Publication: Dev.to / Substack&lt;/em&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Hackathon: Mergetober (WeMakeDevs x Cognee)&lt;/em&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Pull Request: &lt;a href="https://github.com/topoteretes/cognee-community/pull/351" rel="noopener noreferrer"&gt;topoteretes/cognee-community#351&lt;/a&gt; (Closes &lt;a href="https://github.com/topoteretes/cognee/issues/5554" rel="noopener noreferrer"&gt;topoteretes/cognee#5554&lt;/a&gt;)&lt;/em&gt;  &lt;/p&gt;


&lt;h2&gt;
  
  
  1. The 2:00 AM Incident: Why LLMs Need Pipeline Lineage Memory
&lt;/h2&gt;

&lt;p&gt;Picture this scenario: It is 2:15 AM, and your data platform's automated anomaly detector fires a P0 page. A core executive revenue dashboard is reading empty.&lt;/p&gt;

&lt;p&gt;You launch your internal AI operations assistant and ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Why is the &lt;code&gt;daily_revenue_summary&lt;/code&gt; table empty, and what failed upstream?"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If your assistant is backed by traditional RAG or a naive relational database connector, you will hit an immediate dead end. The agent might know the schema of &lt;code&gt;daily_revenue_summary&lt;/code&gt;, but it has no operational visibility into the pipeline that generated it. It cannot see that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An upstream Apache Spark job (&lt;code&gt;etl_clean_payments&lt;/code&gt;) threw an &lt;code&gt;OutOfMemoryError&lt;/code&gt; 45 minutes ago.&lt;/li&gt;
&lt;li&gt;That Spark job reads from an Apache Kafka topic whose schema was altered by an external microservice.&lt;/li&gt;
&lt;li&gt;A downstream dbt model was skipped because its dependency never finished writing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most autonomous AI agents suffer from a fundamental architectural blind spot: &lt;strong&gt;they possess zero operational lineage memory&lt;/strong&gt;. They know what data looks like at rest, but they have no cognitive map of how data moves, transforms, or fails across distributed compute engines.&lt;/p&gt;

&lt;p&gt;To solve this problem, we designed and built the &lt;strong&gt;OpenLineage and Marquez data-source connector for &lt;a href="https://github.com/topoteretes/cognee" rel="noopener noreferrer"&gt;Cognee&lt;/a&gt;&lt;/strong&gt; (&lt;a href="https://github.com/topoteretes/cognee-community/pull/351" rel="noopener noreferrer"&gt;&lt;code&gt;topoteretes/cognee-community#351&lt;/code&gt;&lt;/a&gt;), bridging open metadata standards into autonomous AI memory graphs.&lt;/p&gt;

&lt;p&gt;In this deep dive, we walk through the engineering design, the mechanics of cognitive memory extraction, the full-snapshot reconciliation pattern, and how you can run this in production.&lt;/p&gt;


&lt;h2&gt;
  
  
  2. Understanding the Foundation: OpenLineage &amp;amp; Cognee
&lt;/h2&gt;

&lt;p&gt;Before diving into code, let us look at the two systems powering this integration.&lt;/p&gt;
&lt;h3&gt;
  
  
  The OpenLineage Standard &amp;amp; Marquez
&lt;/h3&gt;

&lt;p&gt;Created under the Linux Foundation (LF AI &amp;amp; Data), &lt;strong&gt;OpenLineage&lt;/strong&gt; is the open-source industry standard for observational data lineage. OpenLineage defines an extensible specification for tracking:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Jobs&lt;/strong&gt;: Units of computation (Airflow tasks, Spark applications, dbt models, Flink jobs).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Datasets&lt;/strong&gt;: Data stores consumed or generated (PostgreSQL tables, Iceberg datasets, S3 Parquet paths).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runs&lt;/strong&gt;: Specific executions of jobs with state transitions (&lt;code&gt;START&lt;/code&gt;, &lt;code&gt;RUNNING&lt;/code&gt;, &lt;code&gt;COMPLETE&lt;/code&gt;, &lt;code&gt;FAIL&lt;/code&gt;, &lt;code&gt;ABORT&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Facets&lt;/strong&gt;: Granular metadata attachments, including schema definitions, SQL query text, data quality assertions, and failure stack traces.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Marquez&lt;/strong&gt; serves as the reference backend implementation for OpenLineage, storing and exposing lineage graphs via a clean REST API.&lt;/p&gt;
&lt;h3&gt;
  
  
  What is Cognee?
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Cognee&lt;/strong&gt; (&lt;code&gt;topoteretes/cognee&lt;/code&gt;) is an open-source memory engine engineered specifically for AI agents. Rather than treating information as disconnected vector chunks, Cognee:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ingests&lt;/strong&gt; multimodal data via declarative data-loading pipelines powered by &lt;code&gt;dlt&lt;/code&gt; (data load tool).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cognifies&lt;/strong&gt; incoming documents: extracting concepts, typed entities, and inter-entity relationships using language models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persists&lt;/strong&gt; the resulting knowledge graph (backed by FalkorDB, Neo4j, or NetworkX) coupled with vector index retrieval (LanceDB, Qdrant).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Executes Graph Completion Searches&lt;/strong&gt;: enabling LLM agents to traverse deep graph paths to answer complex queries that vector distance alone cannot resolve.&lt;/li&gt;
&lt;/ol&gt;


&lt;h2&gt;
  
  
  3. Architecture &amp;amp; Data Flow: From Pipeline Telemetry to Cognitive Graphs
&lt;/h2&gt;

&lt;p&gt;Connecting raw telemetry to an LLM memory graph requires a disciplined ingestion pipeline. The diagram below illustrates how lineage metadata flows from operational orchestrators into Cognee:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faey1a5e95a0xd1a2uyum.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faey1a5e95a0xd1a2uyum.png" alt="Cognee OpenLineage Architecture" width="799" height="436"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  The Architectural Problem: Ingesting Raw JSON vs. Semantic Projections
&lt;/h3&gt;

&lt;p&gt;An enterprise OpenLineage deployment ingests millions of raw JSON &lt;code&gt;RunEvent&lt;/code&gt; payloads every single day. A typical raw event payload contains over 500 lines of nested transport telemetry: socket addresses, producer client versions, heartbeat counters, and system metrics.&lt;/p&gt;

&lt;p&gt;Attempting to dump raw JSON &lt;code&gt;RunEvents&lt;/code&gt; into an LLM context creates severe issues:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context Window Flooding&lt;/strong&gt;: A single pipeline run can consume tens of thousands of tokens without providing actionable insight.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Degraded Entity Extraction&lt;/strong&gt;: General-purpose LLMs struggle to infer graph relationships when buried in deep nested JSON boilerplate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security Vulnerabilities&lt;/strong&gt;: Raw event facets frequently contain unredacted database connection URIs, service account tokens, or environment parameters.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Our Solution: High-Density Semantic Markdown Projections
&lt;/h3&gt;

&lt;p&gt;Rather than streaming raw JSON events, our connector queries the Marquez catalog (&lt;code&gt;/api/v1/namespaces&lt;/code&gt;, &lt;code&gt;/api/v1/jobs&lt;/code&gt;, &lt;code&gt;/api/v1/runs&lt;/code&gt;, &lt;code&gt;/api/v1/datasets&lt;/code&gt;) and synthesizes high-density, structured Markdown documentation for each job:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# OpenLineage Job: analytics.etl_clean_payments&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="gs"&gt;**Namespace**&lt;/span&gt;: analytics
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Job Name**&lt;/span&gt;: etl_clean_payments
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Job Type**&lt;/span&gt;: BATCH
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**Description**&lt;/span&gt;: Nightly payment cleanup and currency conversion job

&lt;span class="gu"&gt;## Input Datasets&lt;/span&gt;
| Dataset Name | Namespace | Physical Location | Schema Fields |
| :--- | :--- | :--- | :--- |
| raw_transactions | payment_gateway | s3://lakehouse/payments/raw | 14 columns |

&lt;span class="gu"&gt;## Output Datasets&lt;/span&gt;
| Dataset Name | Namespace | Physical Location | Schema Fields |
| :--- | :--- | :--- | :--- |
| daily_revenue_summary | analytics | s3://lakehouse/analytics/daily_revenue | 8 columns |

&lt;span class="gu"&gt;## Recent Execution Runs&lt;/span&gt;
| Run ID | Status | Started | Ended | Error Message |
| :--- | :--- | :--- | :--- | :--- |
| run_84920 | FAIL | 2026-10-09 23:30:00 | 2026-10-09 23:34:12 | OutOfMemoryError: Java heap space |
| run_84919 | COMPLETE | 2026-10-08 23:30:00 | 2026-10-08 23:42:01 | None |
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When Cognee processes this projection through its &lt;code&gt;cognify&lt;/code&gt; engine, the LLM immediately recognizes the relationships:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Entity &lt;code&gt;etl_clean_payments&lt;/code&gt; &lt;strong&gt;READS_FROM&lt;/strong&gt; &lt;code&gt;raw_transactions&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Entity &lt;code&gt;etl_clean_payments&lt;/code&gt; &lt;strong&gt;WRITES_TO&lt;/strong&gt; &lt;code&gt;daily_revenue_summary&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Entity &lt;code&gt;etl_clean_payments&lt;/code&gt; &lt;strong&gt;HAS_RUN&lt;/strong&gt; &lt;code&gt;run_84920&lt;/code&gt; with state &lt;code&gt;FAIL&lt;/code&gt; and error &lt;code&gt;OutOfMemoryError&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4. Connector Implementation: Document Mode &amp;amp; Defensive Engineering
&lt;/h2&gt;

&lt;p&gt;Let us inspect the implementation details of &lt;code&gt;packages/connector/openlineage/cognee_community_connector_openlineage/openlineage.py&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fij9xwahb9nrxzcf0x02t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fij9xwahb9nrxzcf0x02t.png" alt="OpenLineage Connector Core Implementation Snippet" width="799" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A. Routing via Cognee Document Mode
&lt;/h3&gt;

&lt;p&gt;In Cognee, structured tabular data is normally routed into relational SQL schemas. However, operational lineage graph models are richest when ingested through the unstructured document pipeline.&lt;/p&gt;

&lt;p&gt;We tag the &lt;code&gt;dlt&lt;/code&gt; source with Cognee's internal marker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;cognee.tasks.ingestion.dlt_utils&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DOCUMENT_SOURCE_ATTR&lt;/span&gt;

&lt;span class="n"&gt;source&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;_openlineage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;endpoint_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;endpoint_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;namespaces&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;namespaces&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;job_names&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;job_names&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;include_facets&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;include_facets&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_runs_per_job&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;max_runs_per_job&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;setattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;DOCUMENT_SOURCE_ATTR&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openlineage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;source&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This ensures Cognee routes the generated lineage documents directly into entity extraction, constructing typed nodes for jobs, datasets, schemas, and run histories.&lt;/p&gt;

&lt;h3&gt;
  
  
  B. Defeating Ghost Knowledge: Full-Snapshot Replacement (&lt;code&gt;write_disposition="replace"&lt;/code&gt;)
&lt;/h3&gt;

&lt;p&gt;In modern data infrastructure, pipelines evolve rapidly. Airflow DAGs are renamed, dbt staging models are consolidated, and deprecated jobs are turned off.&lt;/p&gt;

&lt;p&gt;A critical vulnerability in persistent AI memory systems is &lt;strong&gt;ghost knowledge&lt;/strong&gt;: an agent continuing to believe that an old pipeline exists weeks after data engineers deleted it.&lt;/p&gt;

&lt;p&gt;To solve this, our connector enforces a full-snapshot replacement strategy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@dlt.resource&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openlineage_jobs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;write_disposition&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;replace&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;openlineage_jobs&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="c1"&gt;# Emits current active jobs from the lineage backend
&lt;/span&gt;    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list_jobs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;namespace&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;namespace&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;render_job_markdown&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;metadata&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;namespace&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;namespace&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;job_name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openlineage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  How Forget-on-Delete Works:
&lt;/h4&gt;

&lt;ol&gt;
&lt;li&gt;Every sync run retrieves the authoritative state of the lineage catalog.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;write_disposition="replace"&lt;/code&gt; replaces the staging table with the current state.&lt;/li&gt;
&lt;li&gt;If an old job is deleted from Marquez, it disappears from staging.&lt;/li&gt;
&lt;li&gt;Cognee's internal &lt;code&gt;orphan_cleanup&lt;/code&gt; detects the vanished entity and reconciles it out of both the knowledge graph and vector indices.&lt;/li&gt;
&lt;li&gt;Unchanged jobs retain stable content hashes, preventing wasteful re-cognification and saving LLM API tokens.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  C. Defensive Secret Sanitization
&lt;/h3&gt;

&lt;p&gt;Lineage facets can accidentally expose database credentials, JDBC passwords, or Bearer tokens embedded in connection strings or run parameters.&lt;/p&gt;

&lt;p&gt;The connector applies defensive regex masking prior to document emission:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;_SECRET_PATTERN&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;(?i)(token|secret|password|passwd|api[_-]?key|access[_-]?key|auth|bearer)\s*[:=]\s*[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\']?([^&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\'\s]+)[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\']?&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_sanitize_string&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;_SECRET_PATTERN&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sub&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\1: [REDACTED]&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No raw credential ever reaches the knowledge graph or the LLM's prompt.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Live Execution &amp;amp; Test Verification
&lt;/h2&gt;

&lt;p&gt;All connector code was validated with end-to-end tests covering namespace discovery, facet parsing, secret redaction, and snapshot replacement:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuermlhznhuytipkf0r2b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuermlhznhuytipkf0r2b.png" alt="Execution &amp;amp; Test Verification Terminal" width="799" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Test Suite Highlights (&lt;code&gt;tests/test_openlineage.py&lt;/code&gt;):
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;10 of 10 Unit &amp;amp; Integration Tests Passing&lt;/strong&gt;: Executed offline in 7.62 seconds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero Network Flakiness&lt;/strong&gt;: Implements &lt;code&gt;FakeMarquezClient&lt;/code&gt;, allowing CI runners in GitHub Actions to test pagination, facet parsing, and error backoff without spinning up external Docker services or cloud dependencies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ruff Compliance&lt;/strong&gt;: 100% clean under &lt;code&gt;ruff check&lt;/code&gt; and &lt;code&gt;ruff format&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  6. Real-World Walkthrough: Diagnosing Failures with Graph Completion
&lt;/h2&gt;

&lt;p&gt;Let us look at how an AI agent uses this connector in production.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Ingesting Lineage Topologies
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cognee&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;cognee_community_connector_openlineage&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;openlineage_source&lt;/span&gt;

&lt;span class="n"&gt;DATASET_NAME&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pipeline_lineage_memory&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ingest&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="c1"&gt;# 1. Connect to OpenLineage / Marquez backend
&lt;/span&gt;    &lt;span class="n"&gt;source&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;openlineage_source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;endpoint_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENLINEAGE_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:5000&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENLINEAGE_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;namespaces&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;analytics&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;payment_gateway&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;include_facets&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_runs_per_job&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# 2. Sync and cognify into graph memory
&lt;/span&gt;    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Ingesting pipeline lineage into Cognee...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;cognee&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;remember&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dataset_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;DATASET_NAME&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Lineage memory graph initialized.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;ingest&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 2: Querying the Graph with an Autonomous Agent
&lt;/h3&gt;

&lt;p&gt;Once the memory graph is populated, the agent can answer complex operational questions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;query_agent&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;question&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which upstream jobs write to daily_revenue_summary, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;and did any recent execution runs fail?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;cognee&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;query_text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;query_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;cognee&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SearchType&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GRAPH_COMPLETION&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;datasets&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;DATASET_NAME&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;--- Agent Root Cause Analysis ---&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;query_agent&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Agent Response &amp;amp; Lineage Traversal
&lt;/h3&gt;

&lt;p&gt;Here is the exact query input and synthesized graph response:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft9cl1eadx2rxnux0mwru.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft9cl1eadx2rxnux0mwru.png" alt="Agent Lineage Query &amp;amp; Root Cause Diagnosis" width="799" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Notice how the agent traversed the graph:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It located the node &lt;code&gt;daily_revenue_summary&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;It traced the incoming &lt;code&gt;WRITES_TO&lt;/code&gt; edge backwards to find &lt;code&gt;etl_clean_payments&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;It inspected the &lt;code&gt;HAS_RUN&lt;/code&gt; relationships to surface run &lt;code&gt;run_84920&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;It pulled the exact error facet (&lt;code&gt;OutOfMemoryError: Java heap space&lt;/code&gt;) and warned the engineer about the downstream impact on executive dashboards.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The entire reasoning chain was derived purely from the structured knowledge graph in less than 2 seconds.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Key Engineering Takeaways
&lt;/h2&gt;

&lt;p&gt;Building this integration for Mergetober highlighted three core architectural lessons for agentic data infrastructure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Semantic Density Beats Raw Logs&lt;/strong&gt;: LLMs thrive on structured architectural Markdown. Ingesting curated metadata projections produces exponentially higher retrieval precision than dumping raw JSON logs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forget-on-Delete is Mandatory&lt;/strong&gt;: If your AI memory layer does not reconcile deletions, it will inevitably hallucinate deprecated systems. Full-snapshot replacement with orphan cleanup is the only sustainable strategy for production pipelines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic CI Runtimes&lt;/strong&gt;: CI test suites must never depend on external network services. Mocking the API boundary (&lt;code&gt;FakeMarquezClient&lt;/code&gt;) ensures fast, deterministic verification for maintainers.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  8. Get Involved &amp;amp; Try It Out
&lt;/h2&gt;

&lt;p&gt;The OpenLineage &amp;amp; Marquez connector is available in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Upstream Pull Request: &lt;a href="https://github.com/topoteretes/cognee-community/pull/351" rel="noopener noreferrer"&gt;topoteretes/cognee-community#351&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Fork Branch: &lt;a href="https://github.com/somuai/cognee-community/tree/feat/connector-openlineage" rel="noopener noreferrer"&gt;&lt;code&gt;somuai/cognee-community:feat/connector-openlineage&lt;/code&gt;&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Package Path: &lt;code&gt;packages/connector/openlineage/&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Associated Issue: &lt;a href="https://github.com/topoteretes/cognee/issues/5554" rel="noopener noreferrer"&gt;topoteretes/cognee#5554&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Have you integrated data lineage into your agent architectures, or are you exploring autonomous data platform operations? Drop your thoughts, questions, and feedback in the comments below!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>opensource</category>
      <category>data</category>
    </item>
    <item>
      <title>Giving AI Agents Lakehouse Memory: How We Built an Apache Iceberg Connector for Cognee</title>
      <dc:creator>SOUMYAJIT GHOSH</dc:creator>
      <pubDate>Fri, 09 Oct 2026 20:40:21 +0000</pubDate>
      <link>https://dev.to/soumyajit_ghosh_b93618199/giving-ai-agents-lakehouse-memory-how-we-built-an-apache-iceberg-connector-for-cognee-2110</link>
      <guid>https://dev.to/soumyajit_ghosh_b93618199/giving-ai-agents-lakehouse-memory-how-we-built-an-apache-iceberg-connector-for-cognee-2110</guid>
      <description>&lt;h1&gt;
  
  
  Giving AI Agents Lakehouse Memory: How We Built an Apache Iceberg Connector for Cognee
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Author: Soumyajit Ghosh (&lt;a href="https://github.com/somuai" rel="noopener noreferrer"&gt;@somuai&lt;/a&gt;)&lt;/em&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Target Publication: Dev.to / Substack&lt;/em&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Hackathon: Mergetober (WeMakeDevs x Cognee)&lt;/em&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Pull Request: &lt;a href="https://github.com/topoteretes/cognee-community/pull/348" rel="noopener noreferrer"&gt;topoteretes/cognee-community#348&lt;/a&gt; (Closes &lt;a href="https://github.com/topoteretes/cognee/issues/5553" rel="noopener noreferrer"&gt;topoteretes/cognee#5553&lt;/a&gt;)&lt;/em&gt;  &lt;/p&gt;


&lt;h2&gt;
  
  
  1. The Blind Spot in Enterprise AI: Lakehouse Memory
&lt;/h2&gt;

&lt;p&gt;Modern autonomous AI agents are rapidly moving from toy chat interfaces to mission-critical operational tools. In data platform engineering, agents are increasingly tasked with diagnosing failed ETL pipelines, performing data governance audits, and answering developer queries like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Which tables contain customer transaction data, how are they partitioned, and what schema migrations occurred in the last release?"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;However, most LLM applications remain completely blind to the storage foundation of the modern data stack: &lt;strong&gt;The Data Lakehouse&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Over the past three years, &lt;strong&gt;Apache Iceberg&lt;/strong&gt; has become the open table format standard powering enterprise lakehouses at Netflix, Apple, Snowflake, Databricks, BigQuery, Starburst, and AWS. Iceberg provides ACID transactions, scalable metadata, hidden partitioning, and snapshot time travel.&lt;/p&gt;

&lt;p&gt;Yet, when engineering teams plug AI agents into their infrastructure, agents have no semantic memory of these lakehouses. Traditional relational connectors attempt to dump raw table records into context windows (which blows up tokens and leaks sensitive customer PII) or require cumbersome manual documentation that goes stale the moment a pipeline runs.&lt;/p&gt;

&lt;p&gt;To solve this, I built the &lt;strong&gt;Apache Iceberg data-source connector for &lt;a href="https://github.com/topoteretes/cognee" rel="noopener noreferrer"&gt;Cognee&lt;/a&gt;&lt;/strong&gt; as part of the &lt;strong&gt;Mergetober Hackathon&lt;/strong&gt; (&lt;a href="https://github.com/topoteretes/cognee-community/pull/348" rel="noopener noreferrer"&gt;topoteretes/cognee-community#348&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;In this technical walkthrough, I will break down how Cognee's cognitive memory engine works, how we designed a zero-leakage metadata connector using &lt;code&gt;dlt&lt;/code&gt; and &lt;code&gt;pyiceberg&lt;/code&gt;, and how agents can use knowledge graphs to reason across evolving lakehouse schemas.&lt;/p&gt;


&lt;h2&gt;
  
  
  2. What is Cognee?
&lt;/h2&gt;

&lt;p&gt;LLMs are inherently stateless. Each API request starts from scratch. Retrieval-Augmented Generation (RAG) helps, but naive vector search over chunked text often fails to capture relational hierarchies, schema dependencies, and operational changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cognee&lt;/strong&gt; (&lt;code&gt;topoteretes/cognee&lt;/code&gt;) is an open-source memory engine for AI agents. Rather than treating data as isolated chunks, Cognee:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Ingests data through resilient pipelines powered by &lt;code&gt;dlt&lt;/code&gt; (data load tool).&lt;/li&gt;
&lt;li&gt;Runs &lt;strong&gt;&lt;code&gt;cognify()&lt;/code&gt;&lt;/strong&gt;: an entity-extraction and graph-construction process that links concepts, relationships, and metadata into an interconnected Knowledge Graph (backed by FalkorDB, Neo4j, or NetworkX) alongside vector embeddings (LanceDB, Qdrant).&lt;/li&gt;
&lt;li&gt;Provides graph-completion search (&lt;code&gt;cognee.search(..., query_type=SearchType.GRAPH_COMPLETION)&lt;/code&gt;), allowing agents to traverse complex relationships and recall historical context.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A connector in Cognee is what plugs upstream systems into this company brain.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft4b2wr8w6qg4yxayehqj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft4b2wr8w6qg4yxayehqj.png" alt="Cognee Apache Iceberg Connector Architecture" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  3. Connector Architecture: Why Metadata &amp;amp; Document Mode?
&lt;/h2&gt;

&lt;p&gt;When designing the Iceberg connector, we made two critical architectural decisions:&lt;/p&gt;
&lt;h3&gt;
  
  
  A. Metadata-First Semantic Ingestion (No Raw Data Dumps)
&lt;/h3&gt;

&lt;p&gt;A typical production Iceberg table contains petabytes of Parquet files. An AI agent does not need to read 500 million transaction rows to know how to write an optimal query or debug a pipeline.&lt;/p&gt;

&lt;p&gt;Instead, the connector reads &lt;strong&gt;table metadata&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Full table identifier and namespace hierarchy (e.g., &lt;code&gt;analytics.finance.orders&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Column definitions: field IDs, names, nested types, nullability, and docstrings.&lt;/li&gt;
&lt;li&gt;Partition specifications: source fields and transforms (&lt;code&gt;identity&lt;/code&gt;, &lt;code&gt;bucket&lt;/code&gt;, &lt;code&gt;truncate&lt;/code&gt;, &lt;code&gt;day&lt;/code&gt;, &lt;code&gt;month&lt;/code&gt;, &lt;code&gt;year&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Snapshot commit history: commit timestamps, operation types (&lt;code&gt;append&lt;/code&gt;, &lt;code&gt;overwrite&lt;/code&gt;, &lt;code&gt;delete&lt;/code&gt;), and record summaries.&lt;/li&gt;
&lt;li&gt;Table properties and ownership metadata (with automatic secret redaction).&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  B. Cognee Document-Mode Routing
&lt;/h3&gt;

&lt;p&gt;In Cognee, tabular SQL data is often routed into rigid relational schema tables. However, lakehouse metadata is richest when treated as &lt;strong&gt;architectural prose&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;By tagging our &lt;code&gt;dlt&lt;/code&gt; source with Cognee's document-mode marker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;cognee.tasks.ingestion.dlt_utils&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DOCUMENT_SOURCE_ATTR&lt;/span&gt;

&lt;span class="n"&gt;source&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;_iceberg&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;setattr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;DOCUMENT_SOURCE_ATTR&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;iceberg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cognee's ingestion pipeline routes each table into the full &lt;code&gt;cognify&lt;/code&gt; pipeline. The LLM extracts semantic relationships (e.g., &lt;em&gt;&lt;code&gt;orders&lt;/code&gt; table is partitioned by &lt;code&gt;created_at&lt;/code&gt; day&lt;/em&gt;, &lt;em&gt;&lt;code&gt;user_id&lt;/code&gt; foreign key connects to &lt;code&gt;users&lt;/code&gt;&lt;/em&gt;), storing them as explicit graph nodes.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. The Core Engineering Challenge: Full-Snapshot Sync and Forget-on-Delete
&lt;/h2&gt;

&lt;p&gt;In production data platforms, tables are frequently altered, dropped, or decommissioned. A major failure mode in AI memory layers is &lt;strong&gt;ghost knowledge&lt;/strong&gt;: an agent hallucinating that a table still exists weeks after data engineers dropped it.&lt;/p&gt;

&lt;p&gt;To prevent this, the Iceberg connector implements a &lt;strong&gt;full-snapshot replacement strategy&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4c2d6ry6rid52gpjer8l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4c2d6ry6rid52gpjer8l.png" alt="Iceberg Connector Core Implementation Snippet" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  How Forget-on-Delete Works:
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;write_disposition="replace"&lt;/code&gt; ensures each sync run replaces staging with the exact set of tables currently visible in the catalog.&lt;/li&gt;
&lt;li&gt;If a table &lt;code&gt;finance.temp_payroll&lt;/code&gt; is dropped upstream, it simply disappears from the catalog listing.&lt;/li&gt;
&lt;li&gt;On the next sync run, the table is absent from staging.&lt;/li&gt;
&lt;li&gt;Cognee's internal &lt;code&gt;orphan_cleanup&lt;/code&gt; detects the vanished entity and reconciles it out of both the knowledge graph and vector indices.&lt;/li&gt;
&lt;li&gt;Unchanged tables maintain a stable content hash (&lt;code&gt;data_id&lt;/code&gt;), ensuring they are &lt;strong&gt;not re-cognified&lt;/strong&gt;, saving LLM token costs.&lt;/li&gt;
&lt;li&gt;A render failure immediately aborts the run before staging commits, ensuring transient API blips never cause accidental data deletion.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  5. Live Execution &amp;amp; Verification
&lt;/h2&gt;

&lt;p&gt;To verify that the connector behaves correctly under real workloads, we executed an end-to-end sync against an Iceberg REST catalog containing multiple namespaces and partitioned tables, followed by a simulated table decommissioning:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2bll0aew3rm3oqa45zxv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2bll0aew3rm3oqa45zxv.png" alt="Terminal Execution &amp;amp; Verification" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What the Execution Shows:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Discovery&lt;/strong&gt;: Automatically discovers all tables across target namespaces (&lt;code&gt;analytics&lt;/code&gt;, &lt;code&gt;finance&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parsing&lt;/strong&gt;: Extracts column types, partition transforms, and snapshot commit histories into structured markdown.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cognify Execution&lt;/strong&gt;: Populates 42 graph entities and 68 semantic relations into the lakehouse knowledge graph in 1.24s.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Orphan Cleanup&lt;/strong&gt;: Dropping a table upstream triggers clean tombstone reconciliation on the next incremental sync, purging stale nodes and vector embeddings without trace.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  6. Querying Lakehouse Memory with Graph Completion
&lt;/h2&gt;

&lt;p&gt;Here is how simple it is to use the connector in an autonomous agent application:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;cognee&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;cognee_community_connector_iceberg&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;iceberg_source&lt;/span&gt;

&lt;span class="n"&gt;DATASET_NAME&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lakehouse_memory&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="c1"&gt;# 1. Configure Iceberg Catalog Source
&lt;/span&gt;    &lt;span class="n"&gt;source&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;iceberg_source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;catalog_properties&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rest&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;uri&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://catalog.lakehouse.prod:8181&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;warehouse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;s3://production-lakehouse/warehouse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;token&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ICEBERG_BEARER_TOKEN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="n"&gt;namespaces&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;analytics&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,),&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;finance&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,)],&lt;/span&gt;
        &lt;span class="n"&gt;include_snapshots&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# 2. Sync into Cognee Memory
&lt;/span&gt;    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Syncing Apache Iceberg lakehouse metadata into Cognee...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;cognee&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;remember&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dataset_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;DATASET_NAME&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# 3. Ask complex structural and architectural questions
&lt;/span&gt;    &lt;span class="n"&gt;query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which tables in the finance namespace are partitioned by day, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;and what are their write operations?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;cognee&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;query_text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;query_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;cognee&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SearchType&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GRAPH_COMPLETION&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;datasets&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;DATASET_NAME&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Agent Answer:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the agent queries lakehouse memory, it traverses the knowledge graph to synthesize a precise answer:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fti7bk65ntxwe00fkjqqk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fti7bk65ntxwe00fkjqqk.png" alt="Agent Lakehouse Query Input and Output" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Notice that the agent did not need to run expensive SQL &lt;code&gt;SELECT *&lt;/code&gt; table scans over petabyte-scale Parquet datasets. It recalled the exact schema specifications, partition transforms, and commit summaries directly from Cognee's knowledge graph memory.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Reliability and Testing in CI
&lt;/h2&gt;

&lt;p&gt;Open-source connectors often fail in CI because tests rely on live external cloud credentials or network connections.&lt;/p&gt;

&lt;p&gt;Our test suite (&lt;code&gt;tests/test_iceberg.py&lt;/code&gt;) solves this with &lt;strong&gt;in-memory catalog fakes&lt;/strong&gt; (&lt;code&gt;FakeIcebergCatalog&lt;/code&gt;, &lt;code&gt;FakeIcebergTable&lt;/code&gt;):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;All 8 unit and integration tests run deterministically offline without internet access.&lt;/li&gt;
&lt;li&gt;Tests verify:

&lt;ol&gt;
&lt;li&gt;Markdown table formatting for column IDs, types, and docstrings.&lt;/li&gt;
&lt;li&gt;Partition transforms (&lt;code&gt;identity&lt;/code&gt;, &lt;code&gt;bucket&lt;/code&gt;, &lt;code&gt;truncate&lt;/code&gt;, &lt;code&gt;day&lt;/code&gt;, &lt;code&gt;month&lt;/code&gt;, &lt;code&gt;year&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Transient vs. permanent error classification (&lt;code&gt;NoSuchTableError&lt;/code&gt; permanently drops; connection errors bubble up).&lt;/li&gt;
&lt;li&gt;Full-snapshot reconciliation: verifying via an embedded DuckDB/SQLite pipeline that dropping a table from the catalog removes it from staging on the subsequent run.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Linting and formatting are verified 100% clean under &lt;code&gt;ruff check&lt;/code&gt; and &lt;code&gt;ruff format&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  8. Summary &amp;amp; Key Takeaways
&lt;/h2&gt;

&lt;p&gt;By bridging Apache Iceberg into Cognee:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Lakehouse Awareness&lt;/strong&gt;: AI agents gain persistent, self-updating awareness of lakehouse architectures.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero Token Waste&lt;/strong&gt;: Ingests architectural schemas rather than dumping millions of raw rows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero Ghost Entities&lt;/strong&gt;: Decommissioned tables are automatically pruned from memory via full-snapshot reconciliation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise Security&lt;/strong&gt;: Sensitive properties matching tokens, passwords, and secrets are defensively redacted.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The code is available in pull request &lt;a href="https://github.com/topoteretes/cognee-community/pull/348" rel="noopener noreferrer"&gt;topoteretes/cognee-community#348&lt;/a&gt; and fork branch &lt;a href="https://github.com/somuai/cognee-community/tree/feat/connector-iceberg" rel="noopener noreferrer"&gt;&lt;code&gt;somuai/cognee-community:feat/connector-iceberg&lt;/code&gt;&lt;/a&gt; under &lt;code&gt;packages/connector/iceberg/&lt;/code&gt;, resolving &lt;a href="https://github.com/topoteretes/cognee/issues/5553" rel="noopener noreferrer"&gt;&lt;code&gt;topoteretes/cognee#5553&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you are building AI agents that interact with enterprise data infrastructure, try out &lt;a href="https://github.com/topoteretes/cognee" rel="noopener noreferrer"&gt;Cognee&lt;/a&gt; and let us know what you think in the comments!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>opensource</category>
      <category>data</category>
    </item>
  </channel>
</rss>
