<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Soham Babrekar</title>
    <description>The latest articles on DEV Community by Soham Babrekar (@sohambabrekar).</description>
    <link>https://dev.to/sohambabrekar</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4154505%2F3f08bdb5-4646-4340-8989-265f98bee995.jpg</url>
      <title>DEV Community: Soham Babrekar</title>
      <link>https://dev.to/sohambabrekar</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sohambabrekar"/>
    <language>en</language>
    <item>
      <title>Building DevPilot AI: An Open-Source AI Software Engineering Agent</title>
      <dc:creator>Soham Babrekar</dc:creator>
      <pubDate>Sat, 03 Oct 2026 16:42:44 +0000</pubDate>
      <link>https://dev.to/sohambabrekar/building-devpilot-ai-an-open-source-ai-software-engineering-agent-2djk</link>
      <guid>https://dev.to/sohambabrekar/building-devpilot-ai-an-open-source-ai-software-engineering-agent-2djk</guid>
      <description>&lt;p&gt;I'm building DevPilot AI, an open-source AI software engineering assistant designed to help developers understand, modify, test, and work with real codebases.&lt;/p&gt;

&lt;p&gt;The goal isn't just to build another chatbot.&lt;/p&gt;

&lt;p&gt;I want DevPilot to behave more like a software engineering assistant that can reason about a repository and use tools when necessary.&lt;/p&gt;

&lt;p&gt;What I'm building&lt;/p&gt;

&lt;p&gt;The current architecture has three main pieces:&lt;/p&gt;

&lt;p&gt;React — frontend&lt;br&gt;
Go — backend/API layer&lt;br&gt;
Python + FastAPI — AI/agent services&lt;/p&gt;

&lt;p&gt;For the AI layer, I'm experimenting with open-weight models running locally, including Qwen2.5-Coder through Ollama.&lt;/p&gt;

&lt;p&gt;This is important to me because I want the project to remain usable without requiring everyone to depend on a paid proprietary API.&lt;/p&gt;

&lt;p&gt;The agent workflow&lt;/p&gt;

&lt;p&gt;The direction I'm taking is:&lt;/p&gt;

&lt;p&gt;User request&lt;br&gt;
     ↓&lt;br&gt;
Agent reasoning&lt;br&gt;
     ↓&lt;br&gt;
Understand repository&lt;br&gt;
     ↓&lt;br&gt;
Select tools&lt;br&gt;
     ↓&lt;br&gt;
Search / inspect code&lt;br&gt;
     ↓&lt;br&gt;
Generate changes&lt;br&gt;
     ↓&lt;br&gt;
Run tests&lt;br&gt;
     ↓&lt;br&gt;
Inspect failures&lt;br&gt;
     ↓&lt;br&gt;
Iterate&lt;br&gt;
     ↓&lt;br&gt;
Return result&lt;/p&gt;

&lt;p&gt;Instead of treating an LLM as a simple text-generation API, I'm treating it as one component inside a larger software-engineering system.&lt;/p&gt;

&lt;p&gt;Open-source AI&lt;/p&gt;

&lt;p&gt;One of the things I'm particularly interested in is running the model locally.&lt;/p&gt;

&lt;p&gt;I'm currently experimenting with:&lt;/p&gt;

&lt;p&gt;Ollama&lt;br&gt;
Qwen2.5-Coder&lt;br&gt;
FastAPI&lt;br&gt;
LangChain/LangGraph&lt;br&gt;
Repository-aware code search&lt;br&gt;
Tool calling&lt;br&gt;
Automated testing&lt;br&gt;
Agent execution&lt;/p&gt;

&lt;p&gt;This lets me experiment with agentic workflows while keeping the model infrastructure accessible to developers who don't necessarily have access to expensive APIs.&lt;/p&gt;

&lt;p&gt;What I've learned so far&lt;/p&gt;

&lt;p&gt;The biggest lesson has been that building an AI agent is much more than connecting an LLM to a prompt.&lt;/p&gt;

&lt;p&gt;The difficult parts are around the model:&lt;/p&gt;

&lt;p&gt;How does the agent understand the repository?&lt;br&gt;
Which tool should it use?&lt;br&gt;
How much context should be provided?&lt;br&gt;
What happens when a tool fails?&lt;br&gt;
How does it verify its own changes?&lt;br&gt;
How do we prevent the agent from making destructive changes?&lt;br&gt;
How do we know whether the agent actually solved the problem?&lt;/p&gt;

&lt;p&gt;These questions are becoming more interesting to me than simply making the model produce better text.&lt;/p&gt;

&lt;p&gt;What's next&lt;/p&gt;

&lt;p&gt;I'm working toward giving DevPilot capabilities such as:&lt;/p&gt;

&lt;p&gt;🔎 Repository/code search&lt;br&gt;
🛠️ Tool execution&lt;br&gt;
💻 Terminal interaction&lt;br&gt;
🧪 Test execution&lt;br&gt;
🔄 Failure recovery&lt;br&gt;
🧠 Repository-aware reasoning&lt;br&gt;
📚 RAG for larger codebases&lt;br&gt;
📊 Agent execution traces&lt;br&gt;
🔐 Safer tool permissions&lt;/p&gt;

&lt;p&gt;I'm also interested in connecting this work with another project I'm building around evaluating coding-agent reliability.&lt;/p&gt;

&lt;p&gt;The broader idea is simple:&lt;/p&gt;

&lt;p&gt;AI coding agents shouldn't only be judged by whether they eventually produce a correct answer. We should also measure how they got there.&lt;/p&gt;

&lt;p&gt;I'm still building and experimenting, so this is very much a work in progress.&lt;/p&gt;

&lt;p&gt;If you're also building with open-weight models, local AI, coding agents, or open-source developer tools, I'd love to hear what you're working on.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>llm</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
