What I Built
What if an AI coding assistant didn't stop at “I think I found the bug”?
What if it had to investigate an unfamiliar repository, build an evidence-backed explanation, propose a change, and then prove that the repository actually works after the change?
I built RepoMedic for a friend who was trying to contribute to an unfamiliar open-source repository and was stuck on a failing test without understanding the codebase well enough to confidently debug it.
RepoMedic is an evidence-driven AI field engineer for unfamiliar open-source repositories.
Its workflow is:
BUG REPORT
↓
REPOSITORY ANALYSIS
↓
FAILURE REPRODUCTION
↓
INVESTIGATION
↓
HYPOTHESES + EVIDENCE
↓
ROOT CAUSE
↓
PATCH
↓
REAL TEST EXECUTION
↓
VERIFICATION
↓
READY TO CONTRIBUTE
The core principle is:
AI proposes. Evidence decides.
The model can investigate, reason, and propose changes. It cannot declare its own fix successful.
Deterministic tooling reads the source, validates and applies changes, executes the tests, and uses the actual process exit codes to determine whether the fix worked.
A real debugging scenario
For the challenge demo, I used the open-source Python project neithere/argh.
Under Python 3.14, this test:
python -m pytest tests/test_integration.py::test_prog -v
fails.
But:
pytest tests/test_integration.py::test_prog -v
passes.
That difference is the important clue.
The repository isn't simply “broken.” The way the test suite is invoked changes the observed behavior.
The failing output contained:
usage: python -m pytest [-h] {cmd} ...
while the test expected:
usage: __main__.py [-h] {cmd} ...
RepoMedic investigated the relevant source and test files, generated competing hypotheses, gathered evidence, and traced the behavior to Python 3.14's handling of the default prog value in argparse during python -m invocation.
The test was independently reconstructing its expected program name from sys.argv[0], making that expectation brittle.
The useful result wasn't simply:
“Change this line.”
It was:
Why does this happen, which source establishes that explanation, and what evidence distinguishes it from the alternatives?
That's the kind of reasoning I wanted RepoMedic to expose.
Demo
Demo Video: https://drive.google.com/file/d/1djNOexyGXihN_YkK6-_dwoQo-N6qTXpU/view?usp=sharing
The demo walks through the complete workflow:
Bug Report
↓
Repository Analysis
↓
Failure Reproduction
↓
AI Investigation
↓
Root Cause
↓
Patch
↓
Verification
For the verified run:
Targeted test ✓ PASS
Alternate invocation ✓ PASS
Full test suite ✓ 170 passed
The verification shown in the demo is independently re-executed against the previously verified patched workspace and makes zero model calls.
The result comes from actual pytest execution and process exit codes.
Not from the LLM.
Not from a "success": true response.
Not from a generated screenshot.
The tests decide.
Code
GitHub Repository:
https://github.com/ayushshandilya-dev/RepoMedic
The repository contains the frontend, backend, investigation system, deterministic repository tools, patch validation and application pipeline, and verification system.
The project is intentionally built as a separate tool rather than modifying the repository being investigated. The target repository is cloned into an isolated workspace, changes are applied there, and verification is performed against that workspace.
How I Built It
RepoMedic is built around Gemma 3 running locally through Ollama.
Gemma is used for the investigation layer:
- interpreting the bug report
- examining repository evidence
- generating competing hypotheses
- evaluating explanations
- contributing to root-cause reasoning
The architecture deliberately separates AI reasoning from deterministic execution:
OBSERVATION
↓
DETERMINISTIC EVIDENCE
↓
AI REASONING
↓
DETERMINISTIC PATCHING
↓
DETERMINISTIC VERIFICATION
The model never receives unrestricted shell access.
Instead, it works through structured repository tools and evidence.
The stack is deliberately straightforward:
Frontend
- Next.js
- React
- Tailwind CSS
Backend
- Python
- FastAPI
AI
- Gemma 3
- Ollama
- provider abstraction
Repository intelligence
- tree-sitter
- structured source analysis
- deterministic repository tools
Execution
- Git
- pytest
- isolated workspaces
- deterministic patch application
- real test verification
There is no vector database, no multi-agent framework, and no unnecessary orchestration layer.
The goal was to build the actual engineering workflow rather than assemble a collection of AI buzzwords.
One of the most important design decisions was what happens when the AI is wrong.
An AI model can produce an invalid or ambiguous patch proposal. RepoMedic rejects it rather than silently converting it into a successful fix.
For example, one model-generated proposal produced a source span whose ending line did not exist in the file:
end_line 80 is outside file containing 79 lines
RepoMedic rejected the proposal.
That's not a failure of the safety architecture.
That's the safety architecture working.
During development, the small local Gemma 3 1B model was useful for investigation and root-cause reasoning, but structured patch generation was not consistently reliable on my older hardware.
I kept that limitation visible rather than hiding it.
Why Does Open Innovation Matter?
Open innovation matters here because RepoMedic explores what happens when the AI reasoning layer can be open, local, and replaceable.
Using Gemma through Ollama means repository code does not inherently have to be sent to a closed third-party inference API.
Local inference provides:
- greater control over source-code privacy
- offline capability
- lower marginal inference cost
- model flexibility
- inspectable infrastructure
The provider boundary also means that the rest of RepoMedic does not need to depend on a single model vendor.
The deterministic parts of the system — source inspection, patch validation, patch application, and test verification — remain independent of the model.
That creates a clear trust boundary:
The AI can reason about the repository, but the repository itself determines whether the proposed fix works.
This is particularly relevant to open source because developers may be working with sensitive or private code where sending an entire repository to a hosted AI API isn't desirable.
Open models make it possible to experiment with AI assistance that can run closer to the developer's own environment while remaining replaceable and inspectable.
Prize Categories
Best Use of Gemma
RepoMedic uses Gemma 3 through Ollama as a local investigation provider.
Gemma analyzes repository evidence, generates hypotheses, evaluates competing explanations, and contributes to root-cause reasoning.
The final patch validation and verification are handled independently by deterministic tooling and real test execution.
I am intentionally not claiming that Gemma generated the final verified patch. The model's role is the investigation and reasoning layer; the system surrounding it establishes whether the resulting change is valid and whether the repository actually passes.
What's Next?
If I continued RepoMedic beyond the challenge, I'd focus on:
- stronger sandboxing for arbitrary repositories
- persistent workspace management
- more reliable structured patch generation
- model-assisted code navigation
- repository-wide dependency reasoning
- incremental investigation instead of repeated scans
- richer contribution guidance
- stronger provenance tracking for every piece of evidence
The objective would remain the same:
Make AI-assisted software engineering more inspectable and more trustworthy.
AI proposes. Evidence decides.
Top comments (0)