<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sree Sruthi Alur</title>
    <description>The latest articles on DEV Community by Sree Sruthi Alur (@sree_sruthialur_466b9631).</description>
    <link>https://dev.to/sree_sruthialur_466b9631</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3933175%2Fbf85ae9e-f5aa-4009-b871-e9ec6ef998eb.png</url>
      <title>DEV Community: Sree Sruthi Alur</title>
      <link>https://dev.to/sree_sruthialur_466b9631</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sree_sruthialur_466b9631"/>
    <language>en</language>
    <item>
      <title>OpenCopilot: An Open-Source RAG Assistant for Research Cramming</title>
      <dc:creator>Sree Sruthi Alur</dc:creator>
      <pubDate>Sun, 04 Oct 2026 14:50:16 +0000</pubDate>
      <link>https://dev.to/sree_sruthialur_466b9631/opencopilot-an-open-source-rag-assistant-for-research-cramming-4915</link>
      <guid>https://dev.to/sree_sruthialur_466b9631/opencopilot-an-open-source-rag-assistant-for-research-cramming-4915</guid>
      <description>&lt;h2&gt;
  
  
  What I Built &amp;amp; Who It's For
&lt;/h2&gt;

&lt;p&gt;For this weekend's &lt;strong&gt;"Build for a Friend"&lt;/strong&gt; challenge, I built &lt;strong&gt;OpenCopilot&lt;/strong&gt;—a lightweight, private Retrieval-Augmented Generation (RAG) assistant designed for my project teammate. &lt;/p&gt;

&lt;p&gt;During hackathons and semester project crunches, we juggle massive technical documentation, research papers, and API specs across dozens of tabs. Finding specific implementation constraints or database schemas under tight deadlines is a massive bottleneck. &lt;/p&gt;

&lt;p&gt;OpenCopilot solves this by letting my friend drop in research papers or project PDFs and immediately query them through a clean chat interface—getting precise answers grounded strictly in our project files without hallucinations.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Open-Source AI Core
&lt;/h2&gt;

&lt;p&gt;OpenCopilot was built around an entirely open ecosystem to ensure zero vendor lock-in, data privacy, and rapid iteration:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Embeddings:&lt;/strong&gt; Hugging Face's open-source &lt;code&gt;sentence-transformers/all-MiniLM-L6-v2&lt;/code&gt; model running locally in Python to vectorize document chunks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vector Storage:&lt;/strong&gt; &lt;strong&gt;ChromaDB&lt;/strong&gt;, an open-source vector database that manages embeddings and similarity search directly on the host machine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM Reasoning:&lt;/strong&gt; Powered by high-speed open-weights through Groq's open inference infrastructure, providing instant answers without subscription paywalls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interface &amp;amp; Pipeline:&lt;/strong&gt; Built using &lt;strong&gt;Streamlit&lt;/strong&gt; with a direct, transparent RAG retrieval pipeline without heavy, opaque abstractions.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why Open Innovation Matters Here
&lt;/h2&gt;

&lt;p&gt;When building tools for university projects and peer collaboration, open innovation is essential:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Academic Data Privacy:&lt;/strong&gt; Hackathon prototypes and university research proposals often contain unpublished work. By leveraging local open-source embeddings and self-contained vector storage, proprietary project materials aren't indexed by proprietary cloud model providers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero Cost &amp;amp; Accessible Collaboration:&lt;/strong&gt; Commercial LLM subscriptions and paid enterprise APIs are cost-prohibitive for students. An open stack allows our entire team to run and test the assistant without worrying about API quotas or monthly fees.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Transparency &amp;amp; Portability:&lt;/strong&gt; Unlike closed platforms where underlying models can be altered or deprecated overnight, open-source building blocks let us swap embedding models or switch the generation backend seamlessly.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Handing It Over: Teammate Feedback
&lt;/h2&gt;

&lt;p&gt;I sent OpenCopilot over to Hasini to test against our latest coursework documentation.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"This cut through a 40-page technical specification in seconds. Instead of searching keywords across five open PDFs, I asked for the exact architectural constraints and got back the exact paragraphs I needed. It's going to save us hours on our upcoming hackathon deck."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  How It Works
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ingest &amp;amp; Chunk:&lt;/strong&gt; The user uploads a PDF in the Streamlit sidebar. The document is chunked into 500-character segments with overlap via &lt;code&gt;RecursiveCharacterTextSplitter&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vector Indexing:&lt;/strong&gt; Embeddings are generated using &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; and indexed in a local Chroma store.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contextual Retrieval:&lt;/strong&gt; User queries trigger a top-$k$ similarity search in Chroma to gather the most relevant document chunks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grounded Generation:&lt;/strong&gt; The retrieved context is formatted directly into a strict prompt template that forces the open-weight model to base its reasoning only on the provided evidence.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Development Session &amp;amp; Code
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Repository:&lt;/strong&gt; &lt;a href="https://github.com/SreeSruthiAlur/OpenCopilot-RAG" rel="noopener noreferrer"&gt;https://github.com/SreeSruthiAlur/OpenCopilot-RAG&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Built and configured using Windsurf and DevRelay to track the iterative agent session.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>opensource</category>
    </item>
    <item>
      <title>OpenCopilot: An Open-Source RAG Assistant for My Teammate's Research Cramming</title>
      <dc:creator>Sree Sruthi Alur</dc:creator>
      <pubDate>Sun, 04 Oct 2026 14:46:23 +0000</pubDate>
      <link>https://dev.to/sree_sruthialur_466b9631/opencopilot-an-open-source-rag-assistant-for-my-teammates-research-cramming-57ha</link>
      <guid>https://dev.to/sree_sruthialur_466b9631/opencopilot-an-open-source-rag-assistant-for-my-teammates-research-cramming-57ha</guid>
      <description>&lt;h2&gt;
  
  
  What I Built &amp;amp; Who It's For
&lt;/h2&gt;

&lt;p&gt;For this weekend's &lt;strong&gt;"Build for a Friend"&lt;/strong&gt; challenge, I built &lt;strong&gt;OpenCopilot&lt;/strong&gt;—a lightweight, private Retrieval-Augmented Generation (RAG) assistant designed for my project teammate, Hasini. &lt;/p&gt;

&lt;p&gt;During hackathons and semester project crunches, we juggle massive technical documentation, research papers, and API specs across dozens of tabs. Finding specific implementation constraints or database schemas under tight deadlines is a massive bottleneck. &lt;/p&gt;

&lt;p&gt;OpenCopilot solves this by letting Hasini drop in research papers or project PDFs and immediately query them through a clean chat interface—getting precise answers grounded strictly in our project files without hallucinations.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Open-Source AI Core
&lt;/h2&gt;

&lt;p&gt;OpenCopilot was built around an entirely open ecosystem to ensure zero vendor lock-in, data privacy, and rapid iteration:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Embeddings:&lt;/strong&gt; Hugging Face's open-source &lt;code&gt;sentence-transformers/all-MiniLM-L6-v2&lt;/code&gt; model running locally in Python to vectorize document chunks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vector Storage:&lt;/strong&gt; &lt;strong&gt;ChromaDB&lt;/strong&gt;, an open-source vector database that manages embeddings and similarity search directly on the host machine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM Reasoning:&lt;/strong&gt; Powered by high-speed open-weights through Groq's open inference infrastructure, providing instant answers without subscription paywalls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interface &amp;amp; Pipeline:&lt;/strong&gt; Built using &lt;strong&gt;Streamlit&lt;/strong&gt; with a direct, transparent RAG retrieval pipeline without heavy, opaque abstractions.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why Open Innovation Matters Here
&lt;/h2&gt;

&lt;p&gt;When building tools for university projects and peer collaboration, open innovation is essential:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Academic Data Privacy:&lt;/strong&gt; Hackathon prototypes and university research proposals often contain unpublished work. By leveraging local open-source embeddings and self-contained vector storage, proprietary project materials aren't indexed by proprietary cloud model providers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero Cost &amp;amp; Accessible Collaboration:&lt;/strong&gt; Commercial LLM subscriptions and paid enterprise APIs are cost-prohibitive for students. An open stack allows our entire team to run and test the assistant without worrying about API quotas or monthly fees.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Transparency &amp;amp; Portability:&lt;/strong&gt; Unlike closed platforms where underlying models can be altered or deprecated overnight, open-source building blocks let us swap embedding models or switch the generation backend seamlessly.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Handing It Over: Teammate Feedback
&lt;/h2&gt;

&lt;p&gt;I sent OpenCopilot over to Hasini to test against our latest coursework documentation.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"This cut through a 40-page technical specification in seconds. Instead of searching keywords across five open PDFs, I asked for the exact architectural constraints and got back the exact paragraphs I needed. It's going to save us hours on our upcoming hackathon deck."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  How It Works
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ingest &amp;amp; Chunk:&lt;/strong&gt; The user uploads a PDF in the Streamlit sidebar. The document is chunked into 500-character segments with overlap via &lt;code&gt;RecursiveCharacterTextSplitter&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vector Indexing:&lt;/strong&gt; Embeddings are generated using &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; and indexed in a local Chroma store.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contextual Retrieval:&lt;/strong&gt; User queries trigger a top-$k$ similarity search in Chroma to gather the most relevant document chunks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grounded Generation:&lt;/strong&gt; The retrieved context is formatted directly into a strict prompt template that forces the open-weight model to base its reasoning only on the provided evidence.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Development Session &amp;amp; Code
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Repository:&lt;/strong&gt; &lt;a href="https://github.com/SreeSruthiAlur/OpenCopilot-RAG" rel="noopener noreferrer"&gt;https://github.com/SreeSruthiAlur/OpenCopilot-RAG&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Built and configured using Windsurf and DevRelay to track the iterative agent session.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
