<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Love Yadav</title>
    <description>The latest articles on DEV Community by Love Yadav (@attlar).</description>
    <link>https://dev.to/attlar</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4157120%2Fb160e0c1-62ef-44ba-b9bb-b59132082419.jpg</url>
      <title>DEV Community: Love Yadav</title>
      <link>https://dev.to/attlar</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/attlar"/>
    <language>en</language>
    <item>
      <title>I Built a Grounded RAG Assistant That Answers From Any GitHub README</title>
      <dc:creator>Love Yadav</dc:creator>
      <pubDate>Sat, 03 Oct 2026 17:10:43 +0000</pubDate>
      <link>https://dev.to/attlar/i-built-a-grounded-rag-assistant-that-answers-from-any-github-readme-3p1k</link>
      <guid>https://dev.to/attlar/i-built-a-grounded-rag-assistant-that-answers-from-any-github-readme-3p1k</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Repo Chatter&lt;/strong&gt; lets you paste any public GitHub repository URL and ask questions about its README in plain English. The app fetches the README, indexes it, and answers using only what the documentation says. Every answer shows the exact README chunks it was based on, so you can verify the source yourself. If the README doesn't cover your question, the app says so instead of guessing.&lt;/p&gt;

&lt;p&gt;The app also tracks commit activity for each added repo through an hourly background job, and the interface uses a custom black-and-white design with a cursor-following glow and a scroll-driven landing page.&lt;/p&gt;

&lt;p&gt;I built this for my teammate &lt;strong&gt;&lt;a class="mentioned-user" href="https://dev.to/prakhar_410"&gt;@prakhar_410&lt;/a&gt;&lt;/strong&gt;, who regularly digs through unfamiliar open-source projects to find a single install command or environment variable. Reading a long README for one detail is slow. With Repo Chatter, it's one question away.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;🔗 &lt;strong&gt;Live app:&lt;/strong&gt; &lt;a href="https://repo-chatter.netlify.app/" rel="noopener noreferrer"&gt;https://repo-chatter.netlify.app/&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;The full source is on GitHub:&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/loveyadav1015" rel="noopener noreferrer"&gt;
        loveyadav1015
      &lt;/a&gt; / &lt;a href="https://github.com/loveyadav1015/RepoChatter" rel="noopener noreferrer"&gt;
        RepoChatter
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Repo Chatter is an AI-powered RAG application that allows users to chat with public GitHub repository READMEs through a cinematic, monochrome interface, providing grounded answers with source citations and tracking commit activity.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Repo Chatter 🚀&lt;/h1&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Chat with any GitHub repository's README using AI-powered RAG — wrapped in a
cinematic, monochrome interface.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Repo Chatter lets developers paste a public GitHub repository URL and instantly ask
natural-language questions about its README — grounded strictly in the actual
documentation, with source citations. It tracks repository commit activity in the
background via a scheduled job, and presents it all through a custom black-and-white
interface with a cursor-reactive glow and an illustrated scroll-triggered landing hero.&lt;/p&gt;
&lt;p&gt;Built as a submission for &lt;strong&gt;OverEngineered&lt;/strong&gt; — the Web Development Wing selection process.&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;🎯 Features&lt;/h2&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;📚 Add Any Public Repo&lt;/strong&gt; — paste a GitHub URL, README is fetched and indexed automatically&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;💬 Grounded Q&amp;amp;A&lt;/strong&gt; — ask questions in plain English, get answers based strictly on the README (no hallucination — the model explicitly refuses to answer outside the provided context)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;📍 Source Citations&lt;/strong&gt; — every answer shows exactly which README chunks were…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/loveyadav1015/RepoChatter" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Frontend&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;React with Vite&lt;/li&gt;
&lt;li&gt;shadcn/ui components&lt;/li&gt;
&lt;li&gt;GSAP with ScrollTrigger for the scroll-driven hero&lt;/li&gt;
&lt;li&gt;Tailwind CSS and Axios&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Backend&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Node.js and Express (ESM), organized into routes, controllers, and services&lt;/li&gt;
&lt;li&gt;Raw PostgreSQL driver (&lt;code&gt;pg&lt;/code&gt;) with parameterized queries throughout&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;node-cron&lt;/code&gt; for scheduled jobs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Database&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;PostgreSQL 16 with the pgvector extension&lt;/li&gt;
&lt;li&gt;Exact k-NN similarity search, which is correct and fast enough at this scale&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;AI models&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Embeddings:&lt;/strong&gt; &lt;code&gt;sentence-transformers/all-MiniLM-L6-v2&lt;/code&gt;, an open-weight model that produces 384-dimension vectors, served through the HuggingFace Inference API&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generation:&lt;/strong&gt; Mixtral 8x7B, an open-weight mixture-of-experts model, served through Groq for fast inference&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both models are open-weight, but the current deployment calls them through hosted APIs rather than running them on my own servers. Each model call sits behind a service adapter, so moving to a self-hosted model is a change to one module.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Infrastructure&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Docker Compose for the local stack (database, backend, frontend)&lt;/li&gt;
&lt;li&gt;Netlify for the static frontend&lt;/li&gt;
&lt;li&gt;Render for the backend and the managed PostgreSQL database&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How the RAG Pipeline Works
&lt;/h2&gt;

&lt;p&gt;Repo Chatter has two flows. Ingestion runs when a repo is added. Retrieval and generation run when a question is asked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ingestion: when a repo is added&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[GitHub URL] --&amp;gt; B[Fetch README&amp;lt;br/&amp;gt;GitHub REST API]
    B --&amp;gt; C[Chunk&amp;lt;br/&amp;gt;~500 chars each]
    C --&amp;gt; D[Embed each chunk&amp;lt;br/&amp;gt;384-dim vectors]
    D --&amp;gt; E[(repo_chunks&amp;lt;br/&amp;gt;pgvector column)]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;&lt;strong&gt;Retrieval and generation: when a question is asked&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Q[User question] --&amp;gt; EQ[Embed question&amp;lt;br/&amp;gt;same model]
    EQ --&amp;gt; R[Retrieve top 5 chunks&amp;lt;br/&amp;gt;cosine distance]
    E[(repo_chunks)] --&amp;gt; R
    R --&amp;gt; G[Generate grounded answer&amp;lt;br/&amp;gt;Groq, Mixtral 8x7B]
    G --&amp;gt; A[Answer + source chunks]
    A --&amp;gt; H[(chat_history)]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;&lt;strong&gt;Step by step&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Fetch:&lt;/strong&gt; the backend retrieves the README through the authenticated GitHub API. Authentication raises the rate limit well above the 60 requests per hour allowed for anonymous calls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chunk:&lt;/strong&gt; the README is split into fixed-size chunks of about 500 characters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embed:&lt;/strong&gt; each chunk is converted into a 384-dimension vector.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Store:&lt;/strong&gt; chunks and vectors are saved to the &lt;code&gt;repo_chunks&lt;/code&gt; table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embed the question:&lt;/strong&gt; the question goes through the same model, so it lives in the same vector space as the chunks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieve:&lt;/strong&gt; pgvector returns the five chunks with the smallest cosine distance to the question.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generate:&lt;/strong&gt; the model receives the question and those chunks, with instructions to answer only from that context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Respond and log:&lt;/strong&gt; the answer and its source chunks are returned to the frontend and saved to &lt;code&gt;chat_history&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When the README doesn't cover the question, the response looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"answer"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"I don't know, this isn't covered in the README."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"sourceChunkIds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"sourceChunkTexts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Commit Tracking
&lt;/h2&gt;

&lt;p&gt;An hourly &lt;code&gt;node-cron&lt;/code&gt; job keeps commit data current for every tracked repository:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    T[Hourly tick] --&amp;gt; L[Load tracked repos]
    L --&amp;gt; F[Fetch commits since last_fetched&amp;lt;br/&amp;gt;GitHub API, paginated]
    F --&amp;gt; I[Insert into commit_logs&amp;lt;br/&amp;gt;ON CONFLICT DO NOTHING]
    I --&amp;gt; U[Update commit_count&amp;lt;br/&amp;gt;and last_fetched]&lt;/code&gt;&lt;/pre&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;First run:&lt;/strong&gt; fetches up to 100 recent commits, with pagination handled explicitly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Later runs:&lt;/strong&gt; fetches only commits newer than &lt;code&gt;last_fetched&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deduplication:&lt;/strong&gt; a unique constraint on &lt;code&gt;(repo_id, commit_hash)&lt;/code&gt; means reruns never create duplicate rows.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Database Design
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Table&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;tracked_repos&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Repo URL, owner, cached README, fetch and embed timestamps, commit count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;repo_chunks&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;README chunks with 384-dim pgvector embeddings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;commit_logs&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Commits fetched by the cron job, unique on &lt;code&gt;(repo_id, commit_hash)&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;chat_history&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Each question, its answer, and the source chunks it cited&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;Open models are what made this project practical.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cost:&lt;/strong&gt; the embedding model is open-weight, so re-indexing a repo doesn't add a per-chunk charge on someone else's bill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproducibility:&lt;/strong&gt; because I can pin a specific model version, retrieval behavior stays consistent over time. A closed API can change under you without notice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flexibility:&lt;/strong&gt; the service adapters let me swap in a different open model, or self-host one, without rewriting the retrieval or generation logic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inspectability:&lt;/strong&gt; with open weights, I can test how the model behaves and read exactly what the retrieval step sends to it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Best Use of Render ($200 USD) :&lt;/strong&gt; The Express backend (RAG ingestion, retrieval, generation, and the hourly commit job) and the managed PostgreSQL database both run on Render.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>Introduction</title>
      <dc:creator>Love Yadav</dc:creator>
      <pubDate>Fri, 02 Oct 2026 10:41:52 +0000</pubDate>
      <link>https://dev.to/attlar/introduction-j9l</link>
      <guid>https://dev.to/attlar/introduction-j9l</guid>
      <description>&lt;p&gt;Hi , I am new to the community. Happy to be part of it.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
