I'm building DevPilot AI, an open-source AI software engineering assistant designed to help developers understand, modify, test, and work with real codebases.
The goal isn't just to build another chatbot.
I want DevPilot to behave more like a software engineering assistant that can reason about a repository and use tools when necessary.
What I'm building
The current architecture has three main pieces:
React — frontend
Go — backend/API layer
Python + FastAPI — AI/agent services
For the AI layer, I'm experimenting with open-weight models running locally, including Qwen2.5-Coder through Ollama.
This is important to me because I want the project to remain usable without requiring everyone to depend on a paid proprietary API.
The agent workflow
The direction I'm taking is:
User request
↓
Agent reasoning
↓
Understand repository
↓
Select tools
↓
Search / inspect code
↓
Generate changes
↓
Run tests
↓
Inspect failures
↓
Iterate
↓
Return result
Instead of treating an LLM as a simple text-generation API, I'm treating it as one component inside a larger software-engineering system.
Open-source AI
One of the things I'm particularly interested in is running the model locally.
I'm currently experimenting with:
Ollama
Qwen2.5-Coder
FastAPI
LangChain/LangGraph
Repository-aware code search
Tool calling
Automated testing
Agent execution
This lets me experiment with agentic workflows while keeping the model infrastructure accessible to developers who don't necessarily have access to expensive APIs.
What I've learned so far
The biggest lesson has been that building an AI agent is much more than connecting an LLM to a prompt.
The difficult parts are around the model:
How does the agent understand the repository?
Which tool should it use?
How much context should be provided?
What happens when a tool fails?
How does it verify its own changes?
How do we prevent the agent from making destructive changes?
How do we know whether the agent actually solved the problem?
These questions are becoming more interesting to me than simply making the model produce better text.
What's next
I'm working toward giving DevPilot capabilities such as:
🔎 Repository/code search
🛠️ Tool execution
💻 Terminal interaction
🧪 Test execution
🔄 Failure recovery
🧠 Repository-aware reasoning
📚 RAG for larger codebases
📊 Agent execution traces
🔐 Safer tool permissions
I'm also interested in connecting this work with another project I'm building around evaluating coding-agent reliability.
The broader idea is simple:
AI coding agents shouldn't only be judged by whether they eventually produce a correct answer. We should also measure how they got there.
I'm still building and experimenting, so this is very much a work in progress.
If you're also building with open-weight models, local AI, coding agents, or open-source developer tools, I'd love to hear what you're working on.
Top comments (0)