DEV Community

Sushant Singh Negi
Sushant Singh Negi

Posted on

Echoes of an Era: Building a Grounded AI Time Capsule from My Grandfather’s Voice

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend


What I Built

Echoes of an Era is an open-source, retrieval-grounded AI voice memory capsule built specifically for my grandfather. It preserves, organizes, indexes, and makes searchable his authentic life stories and oral history without turning the AI into a fictionalized, hallucinatory avatar of him.

The Problem It Solves

My grandfather lived through a world that no longer exists—growing up in the 1950s, communicating via weekly inland letters, experiencing the transformation of traditional family customs, and observing decades of societal change. Our family captured hours of his stories on voice recorders, but like most personal audio archives:

  1. Voice recordings are unsearchable: Audio files are linear and opaque. You cannot Ctrl+F a voice recording.
  2. Stories lose context over time: Unlabelled files sitting in folders become difficult for future generations to explore.
  3. The AI Hallucination Trap: Generic cloud chatbots are trained to be helpful, which often means they fabricate heartwarming but entirely fake personal memories when asked questions they don't know the answer to.

The Core Philosophy

"AI interprets the memories. It does not create the memories."

This application is not an AI pretending to be my grandfather. It is an AI-powered interface to my grandfather's authentic, recorded voice. His original voice recordings remain the single source of truth.

Key Features

  • 🎙️ Audio Memory Vaulting: Upload and store raw voice recordings (.mp3, .wav) up to 10MB per file with automatic PyAV audio metadata decoding.
  • 📝 Timestamped Speech-to-Text: Verbatim transcription using faster-whisper with voice activity detection (VAD) and auto-detected language support (Hindi, Hinglish, English).
  • 🧠 Structured Memory Extraction: Automatically extracts titles, summaries, historical eras, locations, people mentioned, emotions, verbatim quotes, and life advice using gemma3:1B-Q4_K_M running locally through Docker Model Runner.
  • ⚡ 1024-Dimensional Vector Search: High-density semantic indexing using local BAAI/bge-m3 via SentenceTransformers stored in PostgreSQL with pgvector HNSW vector indexes.
  • 🔍 Hybrid Retrieval Engine: Combines dense vector similarity with keyword/lexical field scoring re-ranked via Reciprocal Rank Fusion (RRF).
  • 💬 Grounded Q&A ("Ask Grandfather"): Answers user questions strictly using retrieved memory context, refusing to fabricate answers when evidence is insufficient.
  • 🎧 Interactive Audio Player & Timestamps: Direct, range-supported HTTP audio streaming (206 Partial Content) that jump-cuts to the exact second where grandfather spoke the retrieved memory.
  • 📜 Chronological Era Timeline: Explore memories grouped chronologically across eras (1940s to present) and filter by thematic categories.
  • ⏳ "Then vs Now" Reflections: Compare historical experiences described by grandfather (THEN) against modern realities (NOW), generating generational reflection questions and enduring values.

Screenshots


Home Page Dashboard & Audio Memory Vault


Grounded Q&A Interface ("Ask Grandfather")


Signature "Then vs Now" Historical Reflections Engine


Decade-based Memory Categorization


Demo


Code

git clone https://github.com/NegiSushant/Echoes-of-an-Era.git
Enter fullscreen mode Exit fullscreen mode

How I Built It

Architecture & System Flow

graph TB
    subgraph Frontend Client Layer
        React[React 18 UI - Tailwind CSS]
        Player[Docked Audio Player Component]
        SDK[Client API Layer - api.js]
    end

    subgraph Backend Layer FastAPI
        API[FastAPI Core Application]
        STT[TranscriptionService - faster-whisper]
        Embed[EmbeddingService - BAAI/bge-m3]
        Extract[MemoryExtractorService]
        Retriever[RetrievalService - Hybrid Engine]
        RAG[RAGService - Grounded QA]
        Comp[ComparisonService - Then vs Now]
    end

    subgraph Local LLM Engine
        DMR[Docker Model Runner]
        Gemma[Gemma 3 1B Q4_K_M]
    end

    subgraph Storage & Database Layer
        PG[(PostgreSQL 16)]
        PGVector[pgvector Extension - Vector 1024]
        HNSW[HNSW Vector Index]
        AudioStore[File Storage ./uploads/audio]
    end

    React --> SDK
    React --> Player
    SDK -->|REST API| API

    API --> STT
    API --> Extract
    API --> Embed
    API --> Retriever
    API --> RAG
    API --> Comp

    RAG -->|Grounded Context| DMR
    Extract -->|Extraction Prompt| DMR
    Comp -->|Comparison Prompt| DMR
    DMR --> Gemma

    Retriever --> Embed
    Retriever -->|Vector + Lexical Query| PGVector
    Embed -->|1024-dim Vector| PGVector
    PGVector --> HNSW
    Postgres --> PGVector

    Player -->|HTTP Range Stream| API
    API -->|Read Audio Bytes| AudioStore
    STT -->|Read Audio| AudioStore

1. Gemma 3: Local Reasoning & Memory Extraction

  • Model: gemma3:1B-Q4_K_M (Gemma 3 1B 4-bit quantized).
  • Runtime: Gemma 3 1B Q4_K_M running locally through Docker Model Runner (docker model run gemma3:1B-Q4_K_M) on port 12434.
  • Purpose: Structured JSON memory extraction (temperature: 0.1), grounded RAG answer generation (temperature: 0.2), and "Then vs Now" historical comparison synthesis.

2. BGE-M3: 1024-Dimensional Local Vector Indexing

  • Model: BAAI/bge-m3 via local SentenceTransformers (EMBEDDING_LOCAL_ONLY=true).
  • Dimensions: 1024-dimensional normalized dense vectors.
  • Composite Text Schema: Encodes TITLE + SUMMARY + TOPICS + TRANSCRIPT together so semantic queries like "school days" or "travelling by train" map directly to relevant memories.

3. Speech-to-Text (STT)

  • Model: faster-whisper (large-v3-turbo model default) with PyAV and CTranslate2.
  • Execution: Local transcription with Voice Activity Detection (vad_filter=True) and precise segment timestamps (start_time, end_time).

4. PostgreSQL + pgvector Hybrid Retrieval

  • Database: PostgreSQL 16 image (pgvector/pgvector:pg16) with pgvector extension and HNSW vector index (Vector(1024)).
  • Hybrid Fusion: Combines cosine vector similarity (1.0 - cosine_distance) with multi-field lexical keyword matching re-ranked via Reciprocal Rank Fusion (RRF $k=60$) blended with linear score weighting (65% vector, 35% keyword).

5. Evidence-Grounded Refusal Guardrail

If no retrieved memory candidate meets the confidence floor ($\ge 0.20$), the system refuses to guess and outputs:

"Grandfather has not spoken about this in his recorded memories."

Technology Stack Table

Layer Technology Purpose
Frontend React 18, Vite, Tailwind CSS, Lucide Icons Modern Web UI & docked audio player
Backend FastAPI, Python 3.10+, PyAV, Pydantic REST API & media streaming server
Database PostgreSQL 16 Relational storage for memories, transcripts, & eras
Vector Index pgvector (0.2.5+) 1024-dim HNSW vector similarity search
LLM Model Gemma 3 1B Q4_K_M (gemma3:1B-Q4_K_M) Local memory extraction, grounded QA, Then vs Now
LLM Runtime Docker Model Runner (docker model) Local LLM inference engine (port 12434)
Embeddings BAAI/bge-m3 (SentenceTransformers) 1024-dimensional normalized dense vectors
Speech-to-Text faster-whisper (large-v3-turbo) Local transcription & timestamp extraction
Containerization Docker & Docker Compose Containerized database and vector engine setup

Why Does Open Innovation Matter?

Open-weight models (Gemma 3, BGE-M3, faster-whisper) and open-source infrastructure (PostgreSQL + pgvector) are essential for personal family archives:

  1. Privacy & Data Sovereignty: Personal family history, private transcripts, and voice audio remain strictly on local hardware rather than being transmitted to third-party public LLM cloud providers.
  2. Long-Term Longevity: Closed cloud APIs can change terms of service, adjust pricing, or sunset models. Open-weight models running on standard host hardware guarantee that the family time capsule will remain accessible 10 or 20 years from now.
  3. Full Transparency & Control: Open infrastructure allows us to inspect and customize prompt templates, tune retrieval thresholds, and manage vector indices directly without relying on proprietary black-box pipelines.

Prize Categories

  • Build for a Friend: Built specifically for my grandfather to preserve his living oral history and share his recorded memories with our family.
  • Best Use of Gemma: Uses Gemma 3 1B Q4_K_M running locally through Docker Model Runner for structured memory extraction, evidence-grounded Q&A generation, and historical "Then vs Now" reflections.

Top comments (0)