DEV Community

Zaid Kamil
Zaid Kamil

Posted on

Echo Shelf 2.0: Turning Saved Knowledge Into a Living Second Brain

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

We save useful knowledge every day.

An interesting article gets bookmarked. A useful YouTube video gets added to a playlist. A GitHub repository gets starred. A PDF gets downloaded. An idea gets written into a note.

And then, most of it disappears into the digital pile.

I built Echo Shelf 2.0, an intelligent personal knowledge vault, for a friend who, like many of us, constantly saves useful articles, YouTube videos, GitHub repositories, documents, images, and notesβ€”but rarely gets meaningful value from them again.

The problem wasn't saving information.

The real problem was finding it again, understanding how different saved resources connect, and remembering it when it becomes relevant.

That became the foundation of Echo Shelf 2.0:

Capture β†’ Connect β†’ Resurface

Instead of building another bookmark manager, I wanted to create a living second brain that understands what you save, discovers relationships between your knowledge, and brings old information back when it becomes useful again.

Capture

Save knowledge from different sources and let Echo Shelf understand what was actually saved.

Connect

Discover meaningful relationships between resources that may have been saved days or weeks apart.

Resurface

Bring old knowledge back when something happening now makes it relevant again.

Imagine saving an article about a technology, business strategy, research topic, or idea today.

Weeks later, something related happens in the world.

Instead of expecting you to remember that old bookmark, Echo Shelf can make that connection for you.


What Can Echo Shelf Save?

Echo Shelf supports eight content types:

  • πŸ“° Articles
  • πŸŽ₯ YouTube videos
  • πŸ’» GitHub/GitLab repositories
  • πŸ”— Generic URLs
  • πŸ–ΌοΈ Images
  • πŸ“„ Documents
  • πŸ“ Notes
  • 🧩 Multi-source "Other" entries

A multi-source entry can combine things such as a personal note, several reference URLs, documents, and images into a single knowledge asset.

But I didn't want the AI to blindly guess what a URL or document contained.

That led to one of the most important architectural decisions in the project.


Smart Capture: Ground the AI Before Asking It to Think

A simple implementation could have looked like this:

URL β†’ AI β†’ "Tell me what this is"
Enter fullscreen mode Exit fullscreen mode

I deliberately avoided that.

Instead, Echo Shelf follows a source-grounded pipeline:

User Input
    ↓
Source-Specific Extraction
    ↓
Normalized Context
    ↓
Gemma
    ↓
Structured Metadata
    ↓
User Review
    ↓
Persistence
Enter fullscreen mode Exit fullscreen mode

The key principle is:

Extraction first. AI reasoning second.

Different sources therefore have different extraction paths.

🌐 Articles and Web Pages

Web pages are fetched server-side and processed using Mozilla Readability + JSDOM.

This extracts the meaningful readable content while removing much of the surrounding navigation, advertising, and page clutter.

πŸŽ₯ YouTube Videos

YouTube URLs are resolved using the YouTube Data API v3.

Echo Shelf can retrieve grounded metadata such as:

  • Title
  • Description
  • Channel
  • Publication information
  • Duration
  • Thumbnail

Gemma therefore reasons over actual video metadata rather than trying to infer the contents from a URL.

πŸ’» GitHub and GitLab Repositories

Repository URLs are processed using the relevant repository APIs.

Echo Shelf gathers repository metadata and README content before passing normalized context to Gemma.

πŸ“„ Documents

PDF and supported Office documents are parsed into textual context before AI processing.

The extracted information is normalized and treated as untrusted source content.

πŸ–ΌοΈ Images

Images are handled using Gemma's multimodal capabilities.

This allows visual knowledge such as screenshots, diagrams, charts, and infographics to become part of the user's knowledge library.

πŸ“ Notes

Notes already contain user-provided source content, so they can move directly into the normalization and intelligence pipeline.

The result is that Gemma receives grounded information rather than being expected to invent the contents of an inaccessible resource.


Gemma Is the Intelligence Layer

Echo Shelf 2.0 uses:

Gemma 4 26B β€” gemma-4-26b-a4b-it

through Google's Gemini Developer API.

I didn't want Gemma to exist as a chatbot sitting beside the application.

Instead, I wanted the model to become part of the application's reasoning architecture.

Gemma powers four major intelligence layers:

Smart Capture
      ↓
Smart Connections
      ↓
Knowledge Clusters
      ↓
Contextual Rediscovery
Enter fullscreen mode Exit fullscreen mode

Each layer builds on the previous one.


✨ Smart Capture

After Echo Shelf extracts a source, Gemma analyzes the normalized context and generates useful structured information such as:

  • Title
  • Description
  • Tags
  • Relevant metadata

The generated information is not blindly persisted.

The user can review and edit the generated metadata before saving the item.

I also deliberately kept duplicate detection outside the AI layer.

Echo Shelf uses deterministic mechanisms such as:

  • Canonical URL normalization
  • SHA-256 content fingerprints

This creates an important separation:

AI handles semantic understanding. Deterministic software handles exact identity.


πŸ”Ž Potential Connections

Even before an item is saved, Echo Shelf can look for possible relationships with existing knowledge.

However, I didn't want every operation to require an expensive AI call.

So Potential Connections are deterministic.

Echo Shelf compares information such as:

  • Tags
  • Titles
  • Keywords
  • Descriptions

to cheaply shortlist potentially related resources.

This provides immediate feedback while reserving Gemma for the deeper semantic reasoning that actually benefits from AI.


πŸ”— Smart Connections

After an item has been saved, the user can run Check Connections.

This is where Gemma performs deeper semantic reasoning.

The process has two stages.

Stage 1 β€” Deterministic Candidate Selection

Echo Shelf searches the authenticated user's library and creates a small shortlist of promising candidates.

Stage 2 β€” Gemma Semantic Analysis

Only those candidates are passed to Gemma for semantic analysis.

Gemma can classify relationships such as:

  • prerequisite
  • extends
  • complementary
  • conceptual-overlap
  • practical-application
  • contrast
  • alternative-approach
  • implementation-detail

The relationship also includes a strength and an explanation.

The result is that the user's library gradually becomes an interconnected knowledge graph rather than a flat collection of bookmarks.

Gemma can also reject candidates when apparent keyword similarity does not represent a meaningful conceptual relationship.


🧠 Knowledge Clusters

Individual connections are useful, but I wanted Echo Shelf to answer a bigger question:

What larger subjects am I actually building knowledge around?

That's where Knowledge Clusters come in.

Echo Shelf creates lightweight representations of the user's saved items using information such as:

  • Item ID
  • Title
  • Description
  • Tags
  • Content type

Gemma then looks for meaningful conceptual groups.

A generated cluster can contain:

  • A thematic title
  • A summary
  • Related saved items
  • Semantic tags

The system intentionally avoids forcing every item into a cluster.

If something doesn't meaningfully belong anywhere, it can remain ungrouped.

This prevents the feature from becoming a simple AI-generated folder system.

Instead, the goal is to identify real conceptual patterns across the user's knowledge.


πŸ› οΈ When AI Reasoning Took Longer Than Expected

Knowledge Clusters also created one of the most interesting engineering problems during development.

During testing, cluster generation appeared to simply stop working.

The request would run for a long time and eventually fail.

After investigating the actual model response behavior, I found that Gemma could spend a significant portion of the generation budget on internal reasoning before producing the final structured response.

Earlier output limits could therefore terminate generation before the useful JSON was returned.

I had to account for:

  • Thought-token separation
  • Structured JSON responses
  • Output limits
  • Longer inference times
  • Serverless execution limits
  • Validation of generated IDs

Another production blocker appeared because complex clustering operations could take roughly 45–55 seconds.

The final solution included:

  • Increasing the Gemma timeout
  • Using structured JSON response mode
  • Limiting the candidate set
  • Trimming descriptions before inference
  • Allowing longer execution for the clustering route
  • Validating generated item IDs before persistence

After the changes, Knowledge Cluster generation successfully completed during browser testing.

This was an important lesson from the project:

Integrating AI isn't just about getting a successful model response. You also have to engineer around reasoning time, token budgets, latency, structured output, failure recovery, and the execution environment around the model.


🌍 Contextual Rediscovery

This became my favorite part of Echo Shelf.

Saving and organizing information is useful.

But I wanted Echo Shelf to do something more proactive:

Tell me when something I already know becomes relevant again.

Contextual Rediscovery starts with the user's Knowledge Clusters.

Knowledge Clusters
        ↓
High-Signal Topics
        ↓
GNews
        ↓
Current Articles
        ↓
Gemma Relevance Analysis
        ↓
Old Knowledge Resurfaced
Enter fullscreen mode Exit fullscreen mode

Echo Shelf derives compact topics from the user's clusters and queries GNews for current developments.

Gemma then compares those developments against the user's existing knowledge.

The central question becomes:

What is happening now that makes something I saved before relevant again?

Only meaningful matches are surfaced.

A Rediscovery result can connect:

  • πŸ“° A current article
  • πŸ“š Previously saved knowledge
  • 🧠 A relevant Knowledge Cluster
  • πŸ’‘ An explanation of why the connection matters

This completes the original idea:

Capture β†’ Connect β†’ Resurface

The shelf isn't just storing information anymore.

It is helping decide when that information matters again.


πŸ› Debugging Rediscovery

Rediscovery produced another interesting real-world problem.

The GNews API key was valid and direct requests worked, but running Rediscovery across multiple topics sometimes resulted in failures.

The problem turned out to be rate limiting.

Multiple searches were being sent too quickly, causing HTTP 429 responses.

Some generated search queries were also too specific and returned little useful information.

I changed the system to:

  • Pace GNews requests
  • Prefer shorter, high-signal thematic queries
  • Apply explicit request timeouts
  • Handle provider errors more precisely
  • Preserve previous valid Rediscovery results when a refresh fails

After those changes, Echo Shelf successfully retrieved current news and connected it to previously saved knowledge.

That was the point where the Resurface part of the project really came alive.


πŸ” Security Was Part of the Architecture

Echo Shelf processes arbitrary URLs, documents, AI-generated output, and private user knowledge.

Because of that, security couldn't simply be added at the end.

Server-Verified Identity

Authentication is handled using Supabase Auth.

Every protected server operation derives the user's identity from the verified session.

The application does not trust an arbitrary userId supplied by the browser.

MongoDB Tenant Isolation

Every user-owned MongoDB record is associated with a user.

Reads, updates, deletes, AI candidate selection, connections, clusters, and rediscovery operations are scoped to the authenticated owner.

Cross-user resources are treated as unavailable rather than revealing whether another user's record exists.

SSRF Protection

Fetching arbitrary URLs server-side introduces Server-Side Request Forgery (SSRF) risk.

The extraction layer therefore protects against unsafe targets including:

  • Private network ranges
  • Loopback addresses
  • Link-local and metadata endpoints
  • Unsafe protocols
  • Redirects toward internal addresses
  • Unsafe DNS resolutions

Prompt Injection Boundaries

Extracted webpages and documents are treated as untrusted data.

Their contents are separated from model instructions so text such as:

Ignore all previous instructions...
Enter fullscreen mode Exit fullscreen mode

is treated as part of the source material rather than a trusted instruction.

AI Output Validation

AI-generated output is also treated as untrusted.

Generated IDs, relationships, strengths, and other structured values are validated before they can affect persisted application data.

For me, this became another important principle:

Model output is another trust boundary.


Demo

πŸš€ Live Application:

https://echo-shelf-2.vercel.app/

The deployed application supports the complete workflow:

Authentication
      ↓
Smart Capture
      ↓
Personal Library
      ↓
Smart Connections
      ↓
Knowledge Clusters
      ↓
Contextual Rediscovery
Enter fullscreen mode Exit fullscreen mode

Code

πŸ’» GitHub Repository:

https://github.com/sheda3838/echo-shelf-2

The repository contains:

  • Complete application source
  • Architecture documentation
  • Automated tests
  • Security tests
  • Project README
  • Detailed BUILD_LOG.md

The build log documents not only what worked, but also the architectural decisions, bugs, failed approaches, security hardening, production blockers, and fixes encountered throughout development.


βš™οΈ Tech Stack

Layer Technology
Framework Next.js 16
Frontend React 19 + TypeScript + Tailwind CSS
AI Gemma 4 26B (gemma-4-26b-a4b-it)
AI Access Google Gemini Developer API
Database MongoDB Atlas + Mongoose
Authentication Supabase Auth
Current News GNews API
Video Metadata YouTube Data API v3
Web Extraction Mozilla Readability + JSDOM
Deployment Vercel

πŸ§ͺ Testing It

I built automated tests around the areas where failures would matter most, including:

  • Authentication
  • User isolation
  • Duplicate detection
  • Smart Capture
  • Smart Connections
  • Knowledge Clusters
  • Contextual Rediscovery
  • Security boundaries
  • Document processing
  • Image processing

Before deployment, the project also passed its core quality gates:

TypeScript        β†’ PASS
ESLint            β†’ PASS
Production Build  β†’ PASS
Smart Connections β†’ PASS
Knowledge Clusters β†’ PASS
Rediscovery       β†’ PASS
Enter fullscreen mode Exit fullscreen mode

I also performed browser testing of the actual intelligence workflow rather than relying entirely on mocked responses.


How I Built It

Echo Shelf 2.0 was built as a fresh project for the Hacktoberfest Weekend Challenge.

The development process was iterative:

Idea
 ↓
Secure Application Foundation
 ↓
Authentication + User Isolation
 ↓
Source Extraction
 ↓
Smart Capture
 ↓
Duplicate Prevention
 ↓
Potential Connections
 ↓
Smart Connections
 ↓
Knowledge Clusters
 ↓
Contextual Rediscovery
 ↓
Adversarial + Security Testing
 ↓
UI/UX Refinement
 ↓
Production Debugging
 ↓
Deployment
Enter fullscreen mode Exit fullscreen mode

One principle stayed consistent throughout the project:

Use deterministic software when the answer should be deterministic. Use Gemma when semantic reasoning genuinely adds value.

That means things such as:

  • Authentication
  • Authorization
  • Duplicate detection
  • Tenant isolation
  • Candidate filtering
  • URL security
  • Persistence
  • Validation

remain conventional deterministic software.

Gemma handles the parts that genuinely benefit from semantic intelligence.

That separation made Echo Shelf much easier to reason about, test, secure, and debug.


Why Does Open Innovation Matter?

The most interesting part of building Echo Shelf wasn't simply adding an AI API.

It was designing a system around an open-weight model.

Gemma isn't a decorative chatbot bolted onto the side of Echo Shelf.

It is responsible for the semantic intelligence connecting the product:

Capture β†’ Semantic Connections β†’ Knowledge Organization β†’ Contextual Rediscovery

An open model ecosystem makes it possible to think beyond a single opaque endpoint.

The intelligence layer can evolve with:

  • Different model versions
  • Different inference environments
  • Deployment strategies
  • Performance optimizations
  • Community improvements

while the deterministic architecture surrounding it can remain stable.

Open innovation also encouraged me to think carefully about what AI shouldn't do.

I didn't use Gemma for authentication.

I didn't use it to determine whether two URLs are identical.

I didn't let it decide whether a user owns a database record.

I didn't let generated IDs enter the database without validation.

Instead:

Deterministic Engineering
          +
Open-Weight Intelligence
          =
A More Reliable AI-Native Product
Enter fullscreen mode Exit fullscreen mode

For me, that's what makes open AI exciting.

The model becomes something developers can build around, experiment with, optimize, and learn fromβ€”not simply a black box behind a chat interface.


My Agent Session

I used DevRelay during the Hacktoberfest development workflow while building Echo Shelf 2.0.

My development process included:

  • Architecture design
  • Implementation
  • AI integration
  • Adversarial QA
  • Security hardening
  • Browser verification
  • Production debugging
  • Deployment preparation

The full engineering journey is also documented publicly in the repository's build log:

πŸ‘‰ View the Echo Shelf 2.0 BUILD_LOG


Prize Categories

🟒 Best Use of Gemma

I'm entering Best Use of Gemma because Gemma is the central semantic intelligence layer of Echo Shelf 2.0.

Gemma powers:

  • ✨ Smart Capture
  • πŸ–ΌοΈ Multimodal image understanding
  • πŸ”— Smart Connections
  • 🧠 Knowledge Clusters
  • 🌍 Contextual Rediscovery

The application is deliberately designed around Gemma's semantic reasoning capabilities rather than using the model as a standalone chatbot.

Without that reasoning layer, Echo Shelf would largely be another storage application.

Gemma is what turns the shelf into an intelligent knowledge system.

πŸƒ Best Use of MongoDB Atlas

I'm also entering Best Use of MongoDB Atlas.

MongoDB Atlas serves as Echo Shelf's production persistence layer for the user's knowledge system, including saved assets and the structures built around them.

The document model fits naturally with the different shapes of:

  • Articles
  • Videos
  • Repositories
  • Documents
  • Notes
  • Metadata
  • Connections
  • Knowledge Clusters
  • Rediscovery results

Every user-owned record is tenant-scoped, and ownership is derived from server-verified authentication.

This lets MongoDB Atlas act as the persistent foundation beneath Echo Shelf's AI intelligence layer.


Final Thoughts

Echo Shelf 2.0 started with a simple frustration:

We save far more knowledge than we ever use again.

I didn't want to solve that by creating another folder system.

I wanted the saved knowledge itself to become more useful.

Something that understands what you captured.

Something that notices when two ideas connect.

Something that recognizes the larger themes you're learning about.

And most importantly:

Something that knows when an old piece of knowledge matters again.

That's Echo Shelf 2.0.

Capture it. Connect it. Resurface it.

πŸš€ Live Demo:

https://echo-shelf-2.vercel.app/

πŸ’» Source Code:

https://github.com/sheda3838/echo-shelf-2

Built for the Hacktoberfest 2026 Weekend Challenge: Build for a Friend.

Top comments (0)