DEV Community

Gauarv Chaudhary
Gauarv Chaudhary

Posted on

I built a OneNote connector for Cognee. Importing notes was the easy part.

Built for the Cognee community connector hackathon → issue #4727.

Say you're working on a project. You've got architecture decisions in one OneNote section, meeting notes in another, and a few pages explaining why you abandoned certain approaches.

Two months later, someone asks a simple question:

"Why did we choose this approach instead of the other one?"

You know you wrote it down somewhere.

Now you have to remember which notebook it was in, which section you used, and what you called the page.

That's the situation I had in mind when I picked up Cognee issue #4727.

OneNote already does a good job of storing and organizing notes. I wasn't trying to replace it. I wanted to make those notes available to Cognee, so someone could ask a question and retrieve information from the pages they'd already written.

But there was another problem I hadn't wanted to ignore.

What happens when those notes change?

If someone edits a decision, moves a page, or deletes something entirely, the searchable memory shouldn't keep treating the old version as the truth.

That turned out to be the more interesting part of the implementation.

Here's what I built.

Code: OneNote connector implementation

What it actually does

The connector lets you select specific OneNote notebooks and bring their content into Cognee's document-memory pipeline.

The workflow looks like this:

  1. Sign in using Microsoft Graph's delegated authentication.
  2. Choose the notebooks you want to use.
  3. The connector walks through notebooks, sections, nested section groups, and pages.
  4. It converts the page HTML into readable text while preserving notebook and section context.
  5. Cognee processes the documents so their content can be retrieved later.

The important part is that you choose what gets ingested.

You don't have to import every notebook in your account just because you connected Microsoft.

And the context matters.

Imagine two pages both called "Database."

One is inside Prototype / Experiments. The other is inside Production / Architecture.

Without that context, they're just two documents with the same title. The connector includes the notebook and section path in the searchable text, so the original location isn't lost.

It also preserves existing image alt text and resource references. It doesn't download binary images or attachments.

Architecture diagram showing selected OneNote notebooks passing through Microsoft Graph, document processing, Cognee memory, and retrieval, with separate handling for synchronization and deletion.

Getting it running

I kept this as a Python connector rather than building another dashboard.

That's also what the issue asked for: a package, documentation, a runnable example, and tests. No new frontend was needed.

The implementation lives in packages/connector/onenote/ in cognee-community.

With Python 3.11–3.13 and uv, the setup starts with:

cd packages/connector/onenote
uv sync --locked --group dev
Enter fullscreen mode Exit fullscreen mode

Connecting your Microsoft account

There's one detail worth explaining here.

Microsoft's OneNote API doesn't support app-only authentication. The connector needs delegated access, meaning it works on behalf of a signed-in user.

The example uses MSAL's device-code login flow.

You register an application in the Microsoft Entra admin center, enable public-client authentication, and add the delegated Notes.Read permission. Depending on your organization's settings, administrator consent might be required.

Microsoft documents the authentication requirement here.

After registering the app:

export MICROSOFT_CLIENT_ID="your-application-id"
export MICROSOFT_TENANT_ID="common"

uv run --frozen python examples/onenote_example.py
Enter fullscreen mode Exit fullscreen mode

The first run is designed to guide you through authentication and list the notebooks available to your account.

Then you choose one or more notebook IDs.

export ONENOTE_NOTEBOOK_IDS="your-notebook-id"
export ONENOTE_DATASET="onenote-project-demo"
export LLM_API_KEY="your-provider-key"
export ONENOTE_QUESTION="Why did we choose this project approach?"

uv run --frozen python examples/onenote_example.py
Enter fullscreen mode Exit fullscreen mode

You'll also need to configure Cognee's LLM and embedding providers appropriately.

The example is designed to ingest the selected notebook, process its documents, and run the question.

For a first experiment, I'd use a disposable notebook with a decision and its explanation, rather than importing personal notes immediately.

For example, a hypothetical page might say:

We chose local storage for the first prototype because users needed to work without an internet connection.

Then ask:

"Why did we choose local storage?"

The result should be grounded in that page, not in a generic explanation of storage options.

This is an example of the workflow, not a result from a completed live Microsoft test.

The part that needed more thought: synchronization

Importing documents once is only half the job.

Suppose you wrote that your project uses SQLite. A week later, you switch to PostgreSQL and update the explanation in OneNote.

If Cognee still retrieves the old explanation, having searchable notes hasn't really solved your problem.

So I needed the connector to handle repeated synchronization.

Here's how that works.

Every sync checks the metadata of the selected notebooks. Each OneNote page has a lastModifiedDateTime value, which helps determine whether its content needs to be fetched again.

If the page hasn't changed, the connector can reuse its cached body.

If it's new or modified, the connector fetches its HTML again.

One distinction matters: this is incremental content fetching, not a delta-only metadata feed. The connector still enumerates the selected notebooks to determine what's present.

There's another edge case.

What if you rename a section without changing the page itself?

The page text might be identical, but its location has changed.

So the connector caches the page body separately and rebuilds the notebook and section context from current metadata.

That way, a rename doesn't require downloading an unchanged page body just to update its location.

And if a timestamp changes without changing the actual emitted document, the document identity can stay the same.

Less unnecessary processing, while keeping the stored context up to date.

Deleting a page is where things get risky

This was the part I wanted to get right.

Imagine Cognee has already ingested 50 pages from a notebook.

On the next sync, Microsoft Graph returns only 30 because a request failed partway through pagination.

If the connector blindly assumes the missing 20 pages were deleted, it could remove perfectly valid information.

That's a pretty bad outcome for something that's supposed to help you remember things.

So the connector doesn't treat a missing listing entry as proof of deletion.

It checks previously known missing pages directly and considers the API's error responses and the surrounding notebook context.

Microsoft has different OneNote error codes for deleted resources, nonexistent resources, invalid identifiers, and access problems.

Those differences matter.

A confirmed deletion can lead to removing obsolete memory after successful reconciliation.

A permission failure or ambiguous response shouldn't.

The connector also handles a case that's easy to overlook:

What happens when you delete the very last page?

An empty Python generator isn't necessarily enough to make a loading system replace an existing table with an empty one.

For the final-page case, the implementation uses DLT's empty-table materialization marker:

if notebook_rows:
    yield from notebook_rows.values()
else:
    yield dlt.mark.materialize_table_schema()
Enter fullscreen mode Exit fullscreen mode

That explicitly represents an empty notebook table instead of creating a fake blank document.

There are also deliberate boundaries.

Deselecting a notebook doesn't automatically delete its previous memory. A page moved outside the selected scope retains its earlier ingested copy with a warning.

Those choices may sound conservative, but I'd rather have a connector report uncertainty than silently delete data because an API request failed.

What I actually tested

I didn't want to stop after checking whether the connector could parse a page.

The important cases were the ones where something changed or went wrong.

For local verification, I tested the connector using mocked Microsoft Graph responses and mocked external model calls.

The storage layer wasn't mocked.

Five integration tests exercised real SQLite, Ladybug, and LanceDB persistence and retrieval.

The tests covered things like:

  • Deleting the final page and checking that its owned stored artifacts disappear.
  • Updating a document without retaining the obsolete version.
  • Recovering from a processing failure on a later sync.
  • Retrying cleanup after an earlier failure.
  • Keeping shared graph information that still belongs to another document.

66 package tests passed locally on Python 3.12.

Python pytest output showing 66 OneNote connector tests passing locally, using mocked Microsoft Graph and model calls with real SQLite, Ladybug, and LanceDB storage.

Another 24 nearby Cognee regression tests also passed, along with lint, formatting, package-build, and wheel-import checks.

The five persistence integration tests also passed separately.

Terminal test results showing five passing Cognee integration tests for final-page deletion, content updates, failure recovery, cleanup retries, and preservation of shared graph data.

The package suite can be run with:

uv run --frozen pytest tests -ra
Enter fullscreen mode Exit fullscreen mode

The captured local test runs also reported warnings, so the passing summaries shouldn't be read as proof that every possible environment is free of issues.

And one limitation is important enough to say separately:

I haven't completed live Microsoft validation yet.

That means I haven't demonstrated the full sign-in, notebook selection, edit, and deletion workflow against a real Microsoft account.

The tests exercise the behavior with controlled Graph responses and real local storage. They don't prove that Microsoft's live API will behave exactly like the mocks.

Hosted GitHub Actions CI validation is also still pending in this report.

I don't want to turn a successful local test run into a claim about something I haven't tested.

Things it doesn't do

A few boundaries are worth knowing before using this on important notes.

There's no dedicated OneNote dashboard. This contribution is a Python data-source connector. Authentication and notebook selection are handled through the example and API.

It doesn't download attachments. Existing image descriptions and resource references are preserved, but binary attachments aren't ingested.

It doesn't perform a delta-only metadata sync. Each run enumerates the selected scope, while unchanged page bodies can come from the cache.

Processing isn't one atomic transaction. Snapshot validation, staging, document ingestion, graph processing, and cleanup happen in separate stages. If later processing fails, rerunning the same selection is the recovery path.

It needs persistent state. Cached notebook bodies and sync state are stored locally. That state can contain sensitive notes and should be protected.

Concurrent syncs are limited. The current supported usage is one active sync per account and dataset scope.

These aren't things I'd hide behind a nice screenshot. They're part of knowing what the connector can safely do.

What I'd take away from building this

The issue looked like a straightforward integration task at first: authenticate, fetch pages, convert HTML, send documents to Cognee.

But a useful connector has to answer a harder question:

When the source changes, what should happen to the information we've already stored?

Fetching a page is one operation.

Knowing when to reuse it, update it, or safely forget it requires a lot more care.

That's what most of this contribution ended up focusing on.

The implementation is available in the community repository, with local tests passing and live Microsoft validation still to be completed.

If you're interested in how it works, the code and tests are here:

Repository: https://github.com/topoteretes/cognee-community

Implementation: https://github.com/ANAMASGARD/cognee-community/tree/feat/4727-onenote-connector/packages/connector/onenote

Issue: https://github.com/topoteretes/cognee/issues/4727

Top comments (0)