DEV Community

Zaid Kamil
Zaid Kamil

Posted on

Echo Shelf: An AI Knowledge Lake That Connects What You Save to What Matters Now

Sanity Challenge Path Two Submission

This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange

What I Built

I built Echo Shelf, an AI-assisted personal knowledge lake designed around one problem I kept running into:

I save useful articles, videos, repositories, screenshots, documents, and notes — but most of them disappear into a bookmarking graveyard and I rarely return to them.

Echo Shelf turns that passive collection into an active knowledge system.

The core idea is:

Capture → Connect → Resurface

Instead of only storing links, Echo Shelf:

  • captures knowledge from multiple source types
  • extracts useful context from each source
  • generates structured metadata with AI
  • finds meaningful relationships between saved items
  • groups related knowledge into thematic clusters
  • connects current news back to things I saved before

The application supports eight content types:

  • Articles
  • Videos
  • Repositories
  • URLs
  • Images
  • Documents
  • Notes
  • Other resources

Smart Capture

Smart Capture analyzes the actual source before generating metadata.

Different sources use different extraction pipelines:

  • Articles and URLs → Mozilla Readability + jsdom
  • GitHub/GitLab repositories → provider APIs + README content
  • YouTube videos → YouTube Data API metadata
  • PDF/DOCX/PPTX/XLSX documents → format-specific parsers
  • Images → Groq Vision using qwen/qwen3.8-27b
  • Notes → direct text context

The extracted context is then sent through Groq for editable:

  • title
  • description
  • tags

The user can always review and edit the AI output before saving.

Smart Connections

After an item is saved, Echo Shelf can discover relationships between it and existing knowledge.

I intentionally split this into two stages.

First, a deterministic metadata shortlist compares:

  • tags
  • title keywords
  • description keywords

Only the strongest candidates are sent to the LLM.

Then openai/gpt-oss-120b evaluates whether the relationship is genuinely useful and returns:

  • relationship type
  • strength
  • explanation

This also means the AI is allowed to reject superficial keyword matches instead of forcing a connection.

Knowledge Clusters

Echo Shelf can analyze lightweight metadata across the user's library and discover broader themes.

For the final demonstration library, it successfully discovered themes such as:

  • Retrieval-Augmented Generation & Vector Search
  • Next.js Server Actions
  • Docker & Kubernetes Networking
  • Developer Productivity & Deep Work
  • Ergonomics & Workplace Health
  • Sleep Hygiene & Recovery

An item can belong to more than one cluster when that relationship is genuinely useful.

Contextual Rediscovery

This is the feature that completes the original idea.

Echo Shelf asks:

What is happening now that makes something I saved before relevant again?

Knowledge Cluster metadata is converted into compact news queries and sent to the GNews API.

Recent articles are then compared against saved knowledge using Groq.

Only meaningful strong or moderate matches are retained.

Each result explains:

Why this matters to your shelf

and links the current news article back to the related saved item and Knowledge Cluster.

Other details

Echo Shelf also includes:

  • exact duplicate detection using canonical URLs and SHA-256 fingerprints
  • email/password authentication
  • Google OAuth
  • GitHub OAuth
  • SSR cookie sessions with Supabase
  • server-enforced per-user ownership
  • embedded Sanity Studio
  • responsive branded UI
  • loading and pending states for asynchronous actions
  • duplicate-click prevention

The final application is a custom Next.js interface built on top of the Sanity Content Lake rather than a traditional CMS frontend.


Demo

Live Application

👉 https://echo-shelf-three.vercel.app/

Demo Account

The account below is pre-populated with a multi-domain knowledge library so the full experience can be explored immediately.

Email: echoshelf@gmail.com

Password: password

Recommended Testing Flow

  1. Sign in using the demo account.
  2. Browse the Library and its mixed content types.
  3. Open saved items and inspect their Smart Connections.
  4. Open Clusters to explore AI-discovered knowledge themes.
  5. Open Rediscover to see current developments connected back to saved knowledge.
  6. Try Add Item to test Smart Capture yourself.

Screenshots

Library / Personal Knowledge Lake
Echo Shelf library showing a mixed personal knowledge collection with search, filters, and saved item cards

Smart Capture
Echo Shelf Smart Capture page for adding and enriching new knowledge items

Knowledge Clusters
Echo Shelf Knowledge Clusters page showing AI-generated themes across saved items

Contextual Rediscovery
Echo Shelf Rediscover page showing recent news articles matched with previously saved knowledge, including relevance explanations and links back to related shelf items

Authentication
Echo Shelf authentication page showing email sign-in fields and social login options for Google and GitHub


Code

Source code:

👉 https://github.com/sheda3838/echo-shelf

The repository also contains:

👉 BUILD_LOG.md

I kept the build log throughout development instead of reconstructing the process at the end.

It records the decisions, failed approaches, manual tests, architectural changes, debugging sessions, and production fixes that shaped the final application.


My Build Process

I built Echo Shelf using an AI-native development workflow with Antigravity IDE, but I did not treat generated code as automatically correct.

My workflow was generally:

define a small milestone → prompt the IDE → inspect the implementation → run automated checks → manually test real flows → document what failed → refine

The most useful prompts were the ones with very narrow constraints.

Instead of asking:

"Build an AI knowledge app"

I broke Echo Shelf into individual systems such as:

  • define the Sanity schema
  • create the dynamic Add Item flow
  • implement one source extractor at a time
  • build deterministic candidate shortlisting
  • add AI relationship validation
  • introduce first-class Knowledge Cluster documents
  • build Rediscovery around persisted cluster metadata
  • introduce authenticated per-user ownership
  • harden production behavior

Where the model got things wrong

There were several cases where the first generated approach worked in theory but failed against real usage.

1. OCR was the wrong approach for general images

My first Image Smart Capture implementation used Tesseract.js OCR.

It worked for screenshots containing lots of text.

Then I tested it with a normal photo.

The OCR produced meaningless fragments, and the text model generated metadata about the OCR noise instead of describing the image.

That changed the architecture completely.

I replaced:

Image → OCR → text model

with:

Image → Groq Vision → structured metadata

using qwen/qwen3.8-27b.

That immediately made photos, diagrams, screenshots, and mixed visual content much more useful.

2. Document extraction passed tests but broke in the real browser flow

The document parser initially looked correct in automated testing.

Manual testing exposed several framework-level issues:

  • pdf-parse package entry behavior during Next.js bundling
  • Server Action body limits for uploaded files
  • extraction failures being incorrectly reported as AI failures

Those were fixed individually instead of hiding them behind a generic error.

This was one of the biggest reminders during the project that passing static checks is not the same as testing an actual product flow.

3. Production Smart Capture failed after deployment

After deploying to Vercel, Article Smart Capture suddenly returned HTTP 500 even though the local production build succeeded.

Vercel logs revealed an ESM/CommonJS incompatibility inside the deployed jsdom dependency chain.

The failure happened before the requested article was even fetched.

Instead of rewriting Smart Capture, I traced the runtime dependency problem and pinned jsdom to a compatible version while keeping the Readability architecture unchanged.

4. Knowledge Clusters became unstable with realistic demo data

The small test library worked.

The final demonstration library contained 41 saved items, and suddenly cluster generation became inconsistent:

  • sometimes 4 clusters
  • sometimes 2
  • sometimes a generic error

I built a non-persisting diagnostic runner and repeated the same generation multiple times.

The actual problem was not Sanity or cluster validation.

openai/gpt-oss-120b was spending too much of its completion budget on reasoning, causing Groq's JSON output to be truncated before it could form a valid document.

I fixed this by:

  • truncating descriptions sent to clustering
  • limiting tags
  • removing redundant metadata
  • using compact JSON
  • explicitly controlling completion tokens
  • reducing reasoning effort
  • limiting the result to a maximum of six meaningful clusters

After the change, repeated runs consistently generated complete, valid clusters.

The only later failures were API quota limits rather than malformed cluster output.

Prompts that worked best

The prompts that produced the strongest results were usually explicit about what not to change.

For example:

Keep the current Sanity schema and persistence behavior. Diagnose why clustering fails before changing the algorithm. Run repeated dry-run generations without modifying stored data and report the exact failure mechanism.

Or:

Replace only the image-understanding layer. Preserve Smart Capture form state, manual metadata, stale-response protection, and the existing save flow.

Those constraints kept the AI from solving one problem by accidentally rewriting an unrelated part of the product.

Sanity became more than storage

The project started with the idea that Sanity would hold saved items.

It eventually became the structural backbone of the application.

The Content Lake stores not only knowledge assets, but relationships between them:

  • users own saved items
  • saved items reference connected saved items
  • Knowledge Clusters reference multiple saved items
  • saved items can participate in multiple clusters
  • Rediscovery results reference both saved items and clusters

GROQ then powers:

  • library filtering
  • item lookup
  • duplicate detection
  • candidate shortlisting
  • cluster expansion
  • Rediscovery loading
  • per-user ownership filtering

This relational structure is what allowed the later AI features to build on one another.

Smart Connections would be much less useful without structured item references.

Rediscovery would be much less useful without first-class Knowledge Clusters.

That was the most important architectural lesson from this build.


Sanity Project Details

Sanity Project ID: jcon1mtg

Dataset: production

Echo Shelf uses Sanity as a structured Content Lake with four primary document types.

savedItem

Represents an individual knowledge asset.

It stores structured fields such as:

  • title
  • description
  • content type
  • source URL/text/file
  • image asset
  • tags
  • saved timestamp
  • favorite state
  • canonical source fingerprint
  • owner reference
  • Smart Connection references

knowledgeCluster

Represents an AI-discovered conceptual theme.

It stores:

  • title
  • slug
  • summary
  • tags
  • generated timestamp
  • references to constituent savedItem documents
  • owner reference

Cluster membership is many-to-many, so a saved item can participate in more than one theme.

rediscoveryResult

Represents a connection between current news and previously saved knowledge.

It stores:

  • news article metadata
  • relevance
  • connection type
  • explanation
  • saved-item reference
  • Knowledge Cluster reference
  • owner reference

user

Echo Shelf authentication is handled by Supabase.

Sanity stores only a privacy-safe user projection containing:

  • deterministic opaque user document ID
  • display name
  • optional avatar
  • creation timestamp

Raw Supabase user IDs, passwords, authentication tokens, and user emails are not stored in the Sanity user document.

Embedded Studio

Sanity Studio is embedded directly into the application at:

/studio

The user-facing product remains a fully custom Next.js interface, while Studio provides direct inspection of the underlying structured content.


Agent Session

I did not include a public Agent Session for this submission.

Instead, I maintained a detailed chronological build record throughout development:

👉 https://github.com/sheda3838/echo-shelf/blob/main/BUILD_LOG.md

It includes successful prompts, failed approaches, architectural pivots, debugging discoveries, testing notes, and production fixes.


Thanks for checking out Echo Shelf.

The project started as an attempt to make saved links easier to revisit.

It ended up becoming a system where previously saved knowledge can explain its relationships, organize itself, and reappear when something happening today makes it useful again.

Top comments (0)