<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gauarv Chaudhary</title>
    <description>The latest articles on DEV Community by Gauarv Chaudhary (@anamasgard).</description>
    <link>https://dev.to/anamasgard</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3921264%2F1a26b2c1-1fab-45ad-9085-446d9d3403ee.jpeg</url>
      <title>DEV Community: Gauarv Chaudhary</title>
      <link>https://dev.to/anamasgard</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/anamasgard"/>
    <language>en</language>
    <item>
      <title>I connected Contentful to Cognee. Importing content was the easy part.</title>
      <dc:creator>Gauarv Chaudhary</dc:creator>
      <pubDate>Thu, 08 Oct 2026 16:41:59 +0000</pubDate>
      <link>https://dev.to/anamasgard/i-connected-contentful-to-cognee-importing-content-was-the-easy-part-387k</link>
      <guid>https://dev.to/anamasgard/i-connected-contentful-to-cognee-importing-content-was-the-easy-part-387k</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpu5c77ipozzrko5s3lnb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpu5c77ipozzrko5s3lnb.png" alt=" " width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Built for the WeMakeDevs × Cognee Mergetober 2026.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Imagine running an online store where hundreds of products are managed in Contentful.&lt;/p&gt;

&lt;p&gt;You want an AI assistant that can answer questions about those products using Cognee's memory.&lt;/p&gt;

&lt;p&gt;Getting the content into Cognee is one problem. Keeping that memory correct when someone edits or deletes a product is another.&lt;/p&gt;

&lt;p&gt;Say a product's warranty changes from two years to one. Or the product gets discontinued entirely.&lt;/p&gt;

&lt;p&gt;If Cognee still remembers the old information, your assistant could confidently give customers the wrong answer.&lt;/p&gt;

&lt;p&gt;That was the problem behind &lt;a href="https://github.com/topoteretes/cognee/issues/4787" rel="noopener noreferrer"&gt;Cognee issue #4787&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I built a Contentful connector that imports published content into Cognee, keeps it synchronized, and removes outdated information when the original content disappears.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually does
&lt;/h2&gt;

&lt;p&gt;The connector lives in &lt;code&gt;cognee-community/packages/connector/contentful/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It does five things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Connects to Contentful&lt;/strong&gt; using a Content Delivery API token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Imports entries, asset metadata, and content models&lt;/strong&gt; as documents Cognee can understand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Syncs changes&lt;/strong&gt; without downloading everything again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Handles deletions&lt;/strong&gt; so removed content doesn't remain in AI memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keeps sources isolated&lt;/strong&gt; so changing one selection doesn't affect another.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Content models are especially useful here.&lt;/p&gt;

&lt;p&gt;Suppose an entry has a field called &lt;code&gt;batteryLife&lt;/code&gt; with the value &lt;code&gt;12&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Twelve what? Hours? Days?&lt;/p&gt;

&lt;p&gt;The content model provides the field's structure and description. Without that context, an AI system has less information to interpret the value correctly.&lt;/p&gt;

&lt;p&gt;So the connector imports content models separately instead of just dumping entry values into documents.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqy63wplvpa4t4wajkryp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqy63wplvpa4t4wajkryp.png" alt=" " width="800" height="1369"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;One synchronization run. Entries and assets use the Sync API, while content models are fetched separately.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The tricky part: knowing when to forget
&lt;/h2&gt;

&lt;p&gt;The obvious approach to synchronization is simple.&lt;/p&gt;

&lt;p&gt;Fetch everything once, then use &lt;code&gt;sys.updatedAt&lt;/code&gt; to find recently edited entries.&lt;/p&gt;

&lt;p&gt;That works for updates.&lt;/p&gt;

&lt;p&gt;But what about deletions?&lt;/p&gt;

&lt;p&gt;A deleted entry won't appear in the normal listing anymore. A timestamp filter alone can't reliably tell Cognee which information it should forget.&lt;/p&gt;

&lt;p&gt;That's why I used Contentful's native &lt;strong&gt;Sync API&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The first run imports the available content and saves a checkpoint. Later runs use that checkpoint to retrieve changes.&lt;/p&gt;

&lt;p&gt;More importantly, Contentful also returns deletion events like &lt;code&gt;DeletedEntry&lt;/code&gt; and &lt;code&gt;DeletedAsset&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The connector turns those events into dlt hard-delete markers, allowing Cognee to clean up obsolete documents and their associated graph and vector data.&lt;/p&gt;

&lt;p&gt;There's another detail I didn't want to get wrong: pagination.&lt;/p&gt;

&lt;p&gt;If Contentful returns five pages and page three fails, the connector must not save the final checkpoint.&lt;/p&gt;

&lt;p&gt;Otherwise, the next synchronization could skip changes it never processed.&lt;/p&gt;

&lt;p&gt;So it follows every page before accepting the terminal sync token. If the response is incomplete, the run fails instead of silently losing information.&lt;/p&gt;

&lt;h2&gt;
  
  
  One source shouldn't delete another source's memory
&lt;/h2&gt;

&lt;p&gt;Here's another situation.&lt;/p&gt;

&lt;p&gt;Suppose one integration imports English product descriptions, while another imports articles in multiple languages.&lt;/p&gt;

&lt;p&gt;Changing the first integration shouldn't touch the second.&lt;/p&gt;

&lt;p&gt;I added a &lt;code&gt;source_id&lt;/code&gt; to give each logical source its own synchronization state.&lt;/p&gt;

&lt;p&gt;Changing filters under the same source ID triggers a selection refresh. The connector removes documents that no longer belong to that selection while preserving other sources.&lt;/p&gt;

&lt;p&gt;Credentials are also kept separate from source identity.&lt;/p&gt;

&lt;p&gt;Rotating an API token shouldn't mean importing the entire collection as a new source.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure case I cared about
&lt;/h2&gt;

&lt;p&gt;A successful API request doesn't mean the whole synchronization succeeded.&lt;/p&gt;

&lt;p&gt;Think about this sequence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Contentful returns an update or deletion.&lt;/li&gt;
&lt;li&gt;dlt successfully loads the change.&lt;/li&gt;
&lt;li&gt;Cognee fails while processing or cleaning up its stored memory.&lt;/li&gt;
&lt;li&gt;The next Contentful sync returns no new changes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Now what?&lt;/p&gt;

&lt;p&gt;If the connector only trusts the latest API response, that earlier failed work could remain unfinished.&lt;/p&gt;

&lt;p&gt;The implementation relies on Cognee's retained dlt staging so a later run can reconcile existing records again, even when the provider returns an empty delta.&lt;/p&gt;

&lt;p&gt;I also covered individual cleanup failures that Cognee logs without raising an exception.&lt;/p&gt;

&lt;p&gt;That's the distinction that matters here: &lt;strong&gt;successfully loading a change isn't the same as successfully updating AI memory.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Proving it works locally
&lt;/h2&gt;

&lt;p&gt;I wanted to test more than whether the connector could parse some JSON.&lt;/p&gt;

&lt;p&gt;The important questions were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does updating an entry replace obsolete content?&lt;/li&gt;
&lt;li&gt;Does deleting the final entry remove its memory?&lt;/li&gt;
&lt;li&gt;Can two sources coexist without deleting each other's data?&lt;/li&gt;
&lt;li&gt;Can an interrupted ingestion recover on the next run?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The October 8 local verification report recorded &lt;strong&gt;94 passing tests&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tests&lt;/th&gt;
&lt;th&gt;Coverage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;87&lt;/td&gt;
&lt;td&gt;Contentful provider behavior, synchronization, and dlt&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Real local Cognee ingestion, storage, deletion, and recovery&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The same 94 tests passed on Python 3.11, 3.12, and 3.13, plus another Python 3.12 run with the newest permitted dependencies.&lt;/p&gt;

&lt;p&gt;Eight relevant upstream Cognee regression tests also passed, along with Ruff checks and package builds.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi2jabsv0w22y84bpaq7o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi2jabsv0w22y84bpaq7o.png" alt=" " width="799" height="562"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Local test evidence. Contentful responses and model operations were mocked; SQLite, Ladybug, and LanceDB storage paths were exercised for real.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That distinction is important.&lt;/p&gt;

&lt;p&gt;These tests provide evidence of local correctness, including deletion from relational, graph, and vector storage. They don't yet prove authenticated Contentful lifecycle behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try it
&lt;/h2&gt;

&lt;p&gt;The implementation is available in my &lt;a href="https://github.com/ANAMASGARD/cognee-community/tree/feat/4787-contentful-connector/packages/connector/contentful" rel="noopener noreferrer"&gt;Contentful connector branch&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;From the connector directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv &lt;span class="nb"&gt;sync&lt;/span&gt; &lt;span class="nt"&gt;--locked&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Configure your Contentful space and Delivery token:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CONTENTFUL_SPACE_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-space-id"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CONTENTFUL_DELIVERY_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-token"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CONTENTFUL_ENVIRONMENT_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"master"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;LLM_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-model-key"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv run python examples/example.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The example synchronizes published content into Cognee and runs a search query.&lt;/p&gt;

&lt;p&gt;You can select specific content types and languages, or disable asset ingestion. Subsequent runs synchronize changes for the same source.&lt;/p&gt;

&lt;p&gt;For example, a product-catalog application could import descriptions, search for products matching a customer's requirements, and refresh its memory after catalog edits.&lt;/p&gt;

&lt;p&gt;That's an illustrative use case. A live Contentful-to-Cognee demonstration is still needed to establish real-provider behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't do
&lt;/h2&gt;

&lt;p&gt;A few limitations worth mentioning:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It reads published content, not drafts.&lt;/li&gt;
&lt;li&gt;It imports asset metadata, not image or video contents.&lt;/li&gt;
&lt;li&gt;It doesn't recursively expand every linked entry.&lt;/li&gt;
&lt;li&gt;It doesn't include a built-in scheduler.&lt;/li&gt;
&lt;li&gt;Authenticated Contentful import, update, unpublish, and deletion are not yet live-verified.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An invalid synchronization checkpoint also requires deliberate recovery rather than an automatic destructive reset.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd take away from this
&lt;/h2&gt;

&lt;p&gt;Importing data once is straightforward.&lt;/p&gt;

&lt;p&gt;Keeping it correct when the source changes, an entry disappears, or processing fails halfway through is the harder problem.&lt;/p&gt;

&lt;p&gt;The useful lesson from this connector is that synchronization needs to account for both sides: what the source says changed, and what the destination actually finished processing.&lt;/p&gt;

&lt;p&gt;Otherwise, everything can appear to work while the AI is still remembering yesterday's information.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Code and references&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/topoteretes/cognee/issues/4787" rel="noopener noreferrer"&gt;Cognee issue #4787&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ANAMASGARD/cognee-community/tree/feat/4787-contentful-connector/packages/connector/contentful" rel="noopener noreferrer"&gt;Contentful connector implementation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ANAMASGARD/cognee-community/blob/feat/4787-contentful-connector/packages/connector/contentful/tests/test_contentful_ingestion.py" rel="noopener noreferrer"&gt;Integration and recovery tests&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.contentful.com/developers/docs/concepts/sync/" rel="noopener noreferrer"&gt;Contentful Sync API documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pull request:&lt;/strong&gt; &lt;a href="https://github.com/topoteretes/cognee-community/pull/308" rel="noopener noreferrer"&gt;https://github.com/topoteretes/cognee-community/pull/308&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Built as part of WeMakeDevs × Cognee Mergetober 2026.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>cognee</category>
    </item>
    <item>
      <title>I built a OneNote connector for Cognee. Importing notes was the easy part.</title>
      <dc:creator>Gauarv Chaudhary</dc:creator>
      <pubDate>Thu, 08 Oct 2026 11:19:46 +0000</pubDate>
      <link>https://dev.to/anamasgard/i-built-a-onenote-connector-for-cognee-importing-notes-was-the-easy-part-19np</link>
      <guid>https://dev.to/anamasgard/i-built-a-onenote-connector-for-cognee-importing-notes-was-the-easy-part-19np</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzc550hhhyrh1lttvhjp3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzc550hhhyrh1lttvhjp3.png" alt=" " width="800" height="336"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Built for the Cognee community connector hackathon → &lt;a href="https://github.com/topoteretes/cognee/issues/4727" rel="noopener noreferrer"&gt;issue #4727&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Say you're working on a project. You've got architecture decisions in one OneNote section, meeting notes in another, and a few pages explaining why you abandoned certain approaches.&lt;/p&gt;

&lt;p&gt;Two months later, someone asks a simple question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Why did we choose this approach instead of the other one?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You know you wrote it down somewhere.&lt;/p&gt;

&lt;p&gt;Now you have to remember which notebook it was in, which section you used, and what you called the page.&lt;/p&gt;

&lt;p&gt;That's the situation I had in mind when I picked up &lt;a href="https://github.com/topoteretes/cognee/issues/4727" rel="noopener noreferrer"&gt;Cognee issue #4727&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;OneNote already does a good job of storing and organizing notes. I wasn't trying to replace it. I wanted to make those notes available to Cognee, so someone could ask a question and retrieve information from the pages they'd already written.&lt;/p&gt;

&lt;p&gt;But there was another problem I hadn't wanted to ignore.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens when those notes change?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If someone edits a decision, moves a page, or deletes something entirely, the searchable memory shouldn't keep treating the old version as the truth.&lt;/p&gt;

&lt;p&gt;That turned out to be the more interesting part of the implementation.&lt;/p&gt;

&lt;p&gt;Here's what I built.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code:&lt;/strong&gt; &lt;a href="https://github.com/topoteretes/cognee-community/pull/301" rel="noopener noreferrer"&gt;OneNote connector implementation&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually does
&lt;/h2&gt;

&lt;p&gt;The connector lets you select specific OneNote notebooks and bring their content into Cognee's document-memory pipeline.&lt;/p&gt;

&lt;p&gt;The workflow looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Sign in using Microsoft Graph's delegated authentication.&lt;/li&gt;
&lt;li&gt;Choose the notebooks you want to use.&lt;/li&gt;
&lt;li&gt;The connector walks through notebooks, sections, nested section groups, and pages.&lt;/li&gt;
&lt;li&gt;It converts the page HTML into readable text while preserving notebook and section context.&lt;/li&gt;
&lt;li&gt;Cognee processes the documents so their content can be retrieved later.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The important part is that &lt;strong&gt;you choose what gets ingested&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;You don't have to import every notebook in your account just because you connected Microsoft.&lt;/p&gt;

&lt;p&gt;And the context matters.&lt;/p&gt;

&lt;p&gt;Imagine two pages both called "Database."&lt;/p&gt;

&lt;p&gt;One is inside &lt;code&gt;Prototype / Experiments&lt;/code&gt;. The other is inside &lt;code&gt;Production / Architecture&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Without that context, they're just two documents with the same title. The connector includes the notebook and section path in the searchable text, so the original location isn't lost.&lt;/p&gt;

&lt;p&gt;It also preserves existing image alt text and resource references. It doesn't download binary images or attachments.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjafcwi9kl5kc38fhwo8b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjafcwi9kl5kc38fhwo8b.png" alt="Architecture diagram showing selected OneNote notebooks passing through Microsoft Graph, document processing, Cognee memory, and retrieval, with separate handling for synchronization and deletion." width="800" height="755"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting it running
&lt;/h2&gt;

&lt;p&gt;I kept this as a Python connector rather than building another dashboard.&lt;/p&gt;

&lt;p&gt;That's also what the issue asked for: a package, documentation, a runnable example, and tests. No new frontend was needed.&lt;/p&gt;

&lt;p&gt;The implementation lives in &lt;code&gt;packages/connector/onenote/&lt;/code&gt; in &lt;a href="https://github.com/topoteretes/cognee-community" rel="noopener noreferrer"&gt;cognee-community&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;With Python 3.11–3.13 and &lt;code&gt;uv&lt;/code&gt;, the setup starts with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;packages/connector/onenote
uv &lt;span class="nb"&gt;sync&lt;/span&gt; &lt;span class="nt"&gt;--locked&lt;/span&gt; &lt;span class="nt"&gt;--group&lt;/span&gt; dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Connecting your Microsoft account
&lt;/h3&gt;

&lt;p&gt;There's one detail worth explaining here.&lt;/p&gt;

&lt;p&gt;Microsoft's OneNote API doesn't support app-only authentication. The connector needs delegated access, meaning it works on behalf of a signed-in user.&lt;/p&gt;

&lt;p&gt;The example uses MSAL's device-code login flow.&lt;/p&gt;

&lt;p&gt;You register an application in the &lt;a href="https://entra.microsoft.com/" rel="noopener noreferrer"&gt;Microsoft Entra admin center&lt;/a&gt;, enable public-client authentication, and add the delegated &lt;code&gt;Notes.Read&lt;/code&gt; permission. Depending on your organization's settings, administrator consent might be required.&lt;/p&gt;

&lt;p&gt;Microsoft documents the authentication requirement &lt;a href="https://learn.microsoft.com/en-us/graph/api/resources/onenote-api-overview?view=graph-rest-1.0" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;After registering the app:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MICROSOFT_CLIENT_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-application-id"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;MICROSOFT_TENANT_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"common"&lt;/span&gt;

uv run &lt;span class="nt"&gt;--frozen&lt;/span&gt; python examples/onenote_example.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first run is designed to guide you through authentication and list the notebooks available to your account.&lt;/p&gt;

&lt;p&gt;Then you choose one or more notebook IDs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ONENOTE_NOTEBOOK_IDS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-notebook-id"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ONENOTE_DATASET&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"onenote-project-demo"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;LLM_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-provider-key"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ONENOTE_QUESTION&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"Why did we choose this project approach?"&lt;/span&gt;

uv run &lt;span class="nt"&gt;--frozen&lt;/span&gt; python examples/onenote_example.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'll also need to configure Cognee's LLM and embedding providers appropriately.&lt;/p&gt;

&lt;p&gt;The example is designed to ingest the selected notebook, process its documents, and run the question.&lt;/p&gt;

&lt;p&gt;For a first experiment, I'd use a disposable notebook with a decision and its explanation, rather than importing personal notes immediately.&lt;/p&gt;

&lt;p&gt;For example, a hypothetical page might say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;We chose local storage for the first prototype because users needed to work without an internet connection.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then ask:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Why did we choose local storage?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The result should be grounded in that page, not in a generic explanation of storage options.&lt;/p&gt;

&lt;p&gt;This is an example of the workflow, not a result from a completed live Microsoft test.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that needed more thought: synchronization
&lt;/h2&gt;

&lt;p&gt;Importing documents once is only half the job.&lt;/p&gt;

&lt;p&gt;Suppose you wrote that your project uses SQLite. A week later, you switch to PostgreSQL and update the explanation in OneNote.&lt;/p&gt;

&lt;p&gt;If Cognee still retrieves the old explanation, having searchable notes hasn't really solved your problem.&lt;/p&gt;

&lt;p&gt;So I needed the connector to handle repeated synchronization.&lt;/p&gt;

&lt;p&gt;Here's how that works.&lt;/p&gt;

&lt;p&gt;Every sync checks the metadata of the selected notebooks. Each OneNote page has a &lt;code&gt;lastModifiedDateTime&lt;/code&gt; value, which helps determine whether its content needs to be fetched again.&lt;/p&gt;

&lt;p&gt;If the page hasn't changed, the connector can reuse its cached body.&lt;/p&gt;

&lt;p&gt;If it's new or modified, the connector fetches its HTML again.&lt;/p&gt;

&lt;p&gt;One distinction matters: &lt;strong&gt;this is incremental content fetching, not a delta-only metadata feed.&lt;/strong&gt; The connector still enumerates the selected notebooks to determine what's present.&lt;/p&gt;

&lt;p&gt;There's another edge case.&lt;/p&gt;

&lt;p&gt;What if you rename a section without changing the page itself?&lt;/p&gt;

&lt;p&gt;The page text might be identical, but its location has changed.&lt;/p&gt;

&lt;p&gt;So the connector caches the page body separately and rebuilds the notebook and section context from current metadata.&lt;/p&gt;

&lt;p&gt;That way, a rename doesn't require downloading an unchanged page body just to update its location.&lt;/p&gt;

&lt;p&gt;And if a timestamp changes without changing the actual emitted document, the document identity can stay the same.&lt;/p&gt;

&lt;p&gt;Less unnecessary processing, while keeping the stored context up to date.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deleting a page is where things get risky
&lt;/h2&gt;

&lt;p&gt;This was the part I wanted to get right.&lt;/p&gt;

&lt;p&gt;Imagine Cognee has already ingested 50 pages from a notebook.&lt;/p&gt;

&lt;p&gt;On the next sync, Microsoft Graph returns only 30 because a request failed partway through pagination.&lt;/p&gt;

&lt;p&gt;If the connector blindly assumes the missing 20 pages were deleted, it could remove perfectly valid information.&lt;/p&gt;

&lt;p&gt;That's a pretty bad outcome for something that's supposed to help you remember things.&lt;/p&gt;

&lt;p&gt;So the connector doesn't treat a missing listing entry as proof of deletion.&lt;/p&gt;

&lt;p&gt;It checks previously known missing pages directly and considers the API's error responses and the surrounding notebook context.&lt;/p&gt;

&lt;p&gt;Microsoft has &lt;a href="https://learn.microsoft.com/en-us/graph/onenote-error-codes" rel="noopener noreferrer"&gt;different OneNote error codes&lt;/a&gt; for deleted resources, nonexistent resources, invalid identifiers, and access problems.&lt;/p&gt;

&lt;p&gt;Those differences matter.&lt;/p&gt;

&lt;p&gt;A confirmed deletion can lead to removing obsolete memory after successful reconciliation.&lt;/p&gt;

&lt;p&gt;A permission failure or ambiguous response shouldn't.&lt;/p&gt;

&lt;p&gt;The connector also handles a case that's easy to overlook:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens when you delete the very last page?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An empty Python generator isn't necessarily enough to make a loading system replace an existing table with an empty one.&lt;/p&gt;

&lt;p&gt;For the final-page case, the implementation uses DLT's empty-table materialization marker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;notebook_rows&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="n"&gt;notebook_rows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;dlt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mark&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;materialize_table_schema&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That explicitly represents an empty notebook table instead of creating a fake blank document.&lt;/p&gt;

&lt;p&gt;There are also deliberate boundaries.&lt;/p&gt;

&lt;p&gt;Deselecting a notebook doesn't automatically delete its previous memory. A page moved outside the selected scope retains its earlier ingested copy with a warning.&lt;/p&gt;

&lt;p&gt;Those choices may sound conservative, but I'd rather have a connector report uncertainty than silently delete data because an API request failed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually tested
&lt;/h2&gt;

&lt;p&gt;I didn't want to stop after checking whether the connector could parse a page.&lt;/p&gt;

&lt;p&gt;The important cases were the ones where something changed or went wrong.&lt;/p&gt;

&lt;p&gt;For local verification, I tested the connector using mocked Microsoft Graph responses and mocked external model calls.&lt;/p&gt;

&lt;p&gt;The storage layer wasn't mocked.&lt;/p&gt;

&lt;p&gt;Five integration tests exercised real SQLite, Ladybug, and LanceDB persistence and retrieval.&lt;/p&gt;

&lt;p&gt;The tests covered things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Deleting the final page and checking that its owned stored artifacts disappear.&lt;/li&gt;
&lt;li&gt;Updating a document without retaining the obsolete version.&lt;/li&gt;
&lt;li&gt;Recovering from a processing failure on a later sync.&lt;/li&gt;
&lt;li&gt;Retrying cleanup after an earlier failure.&lt;/li&gt;
&lt;li&gt;Keeping shared graph information that still belongs to another document.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;66 package tests passed locally on Python 3.12.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flk6wh4d26hruvrookz7e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flk6wh4d26hruvrookz7e.png" alt="Python pytest output showing 66 OneNote connector tests passing locally, using mocked Microsoft Graph and model calls with real SQLite, Ladybug, and LanceDB storage." width="800" height="571"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Another &lt;strong&gt;24 nearby Cognee regression tests&lt;/strong&gt; also passed, along with lint, formatting, package-build, and wheel-import checks.&lt;/p&gt;

&lt;p&gt;The five persistence integration tests also passed separately.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7bi8lkbc3buqvehhk1ws.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7bi8lkbc3buqvehhk1ws.png" alt="Terminal test results showing five passing Cognee integration tests for final-page deletion, content updates, failure recovery, cleanup retries, and preservation of shared graph data." width="800" height="707"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The package suite can be run with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv run &lt;span class="nt"&gt;--frozen&lt;/span&gt; pytest tests &lt;span class="nt"&gt;-ra&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The captured local test runs also reported warnings, so the passing summaries shouldn't be read as proof that every possible environment is free of issues.&lt;/p&gt;

&lt;p&gt;And one limitation is important enough to say separately:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I haven't completed live Microsoft validation yet.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That means I haven't demonstrated the full sign-in, notebook selection, edit, and deletion workflow against a real Microsoft account.&lt;/p&gt;

&lt;p&gt;The tests exercise the behavior with controlled Graph responses and real local storage. They don't prove that Microsoft's live API will behave exactly like the mocks.&lt;/p&gt;

&lt;p&gt;Hosted GitHub Actions CI validation is also still pending in this report.&lt;/p&gt;

&lt;p&gt;I don't want to turn a successful local test run into a claim about something I haven't tested.&lt;/p&gt;

&lt;h2&gt;
  
  
  Things it doesn't do
&lt;/h2&gt;

&lt;p&gt;A few boundaries are worth knowing before using this on important notes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There's no dedicated OneNote dashboard.&lt;/strong&gt; This contribution is a Python data-source connector. Authentication and notebook selection are handled through the example and API.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It doesn't download attachments.&lt;/strong&gt; Existing image descriptions and resource references are preserved, but binary attachments aren't ingested.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It doesn't perform a delta-only metadata sync.&lt;/strong&gt; Each run enumerates the selected scope, while unchanged page bodies can come from the cache.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Processing isn't one atomic transaction.&lt;/strong&gt; Snapshot validation, staging, document ingestion, graph processing, and cleanup happen in separate stages. If later processing fails, rerunning the same selection is the recovery path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It needs persistent state.&lt;/strong&gt; Cached notebook bodies and sync state are stored locally. That state can contain sensitive notes and should be protected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Concurrent syncs are limited.&lt;/strong&gt; The current supported usage is one active sync per account and dataset scope.&lt;/p&gt;

&lt;p&gt;These aren't things I'd hide behind a nice screenshot. They're part of knowing what the connector can safely do.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd take away from building this
&lt;/h2&gt;

&lt;p&gt;The issue looked like a straightforward integration task at first: authenticate, fetch pages, convert HTML, send documents to Cognee.&lt;/p&gt;

&lt;p&gt;But a useful connector has to answer a harder question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When the source changes, what should happen to the information we've already stored?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Fetching a page is one operation.&lt;/p&gt;

&lt;p&gt;Knowing when to reuse it, update it, or safely forget it requires a lot more care.&lt;/p&gt;

&lt;p&gt;That's what most of this contribution ended up focusing on.&lt;/p&gt;

&lt;p&gt;The implementation is available in the community repository, with local tests passing and live Microsoft validation still to be completed.&lt;/p&gt;

&lt;p&gt;If you're interested in how it works, the code and tests are here:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Repository:&lt;/strong&gt; &lt;a href="https://github.com/topoteretes/cognee-community" rel="noopener noreferrer"&gt;https://github.com/topoteretes/cognee-community&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Implementation:&lt;/strong&gt; &lt;a href="https://github.com/ANAMASGARD/cognee-community/tree/feat/4727-onenote-connector/packages/connector/onenote" rel="noopener noreferrer"&gt;https://github.com/ANAMASGARD/cognee-community/tree/feat/4727-onenote-connector/packages/connector/onenote&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Issue:&lt;/strong&gt; &lt;a href="https://github.com/topoteretes/cognee/issues/4727" rel="noopener noreferrer"&gt;https://github.com/topoteretes/cognee/issues/4727&lt;/a&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>python</category>
      <category>showdev</category>
      <category>cognee</category>
    </item>
  </channel>
</rss>
