I've been interested in local AI for a while, especially running AI models on hardware that isn't exactly a massive AI workstation.
At some point, I started wondering:
How useful could a coding agent be if it had to run on a tiny local model?
So I decided to build one.
That's how MINICODE started.
What is MINICODE?
MINICODE is a terminal-based AI coding agent that I built around small local language models.
The default setup uses Qwen3 1.7B through Ollama.
The goal isn't to compete with massive cloud models.
Instead, I wanted to experiment with making a small model actually useful by giving it tools and letting it interact with a coding environment.
It can work with files, run commands, search the web, maintain conversation context, and interact with the project it's being used in.
Why 1.7B?
I don't have some giant GPU server sitting under my desk.
I wanted something that could run locally and be relatively lightweight.
A 1.7B model is obviously limited compared with much larger models, but that's actually what makes the experiment interesting.
Instead of solving everything by throwing a bigger model at the problem, I wanted to see how much I could improve the experience through the agent itself.
The model doesn't have to do everything alone.
MINICODE can give it tools.
Building the agent
One of the first things I learned is that building an AI coding agent is very different from just making a chatbot.
A chatbot can generate a response.
An agent needs to interact with the world.
For example, if I ask MINICODE to modify a project, the model needs to figure out what files it needs, inspect them, make changes, and potentially run commands to test those changes.
That creates a loop something like:
User
↓
Model
↓
Tool call
↓
Tool executes
↓
Result goes back to model
↓
Model decides what to do next
Getting that loop working reliably was one of the most interesting parts of the project.
I didn't want the AI to have unlimited control
Another thing I cared about was safety.
Giving an AI access to a terminal means it could potentially execute commands that you didn't intend.
So MINICODE asks for confirmation before performing potentially dangerous actions.
Instead of:
AI → execute command
it's more like:
AI → "I want to execute this command"
↓
You approve?
↙ ↘
Yes No
↓ ↓
Execute Tell AI
If I reject an action, the model gets that information and can try another approach.
It's a simple system, but I think keeping a human in the loop is important when experimenting with autonomous tools.
I also built memory
I wanted coding sessions to be persistent rather than disappearing when I closed the terminal.
MINICODE can save conversations so I can come back to them later.
That means I can work on something, save the session, and continue working on it another time.
It's not some incredibly sophisticated long-term memory system, but it makes the agent much more practical to use.
Making installation easy
I also wanted someone to be able to try the project without spending an hour setting everything up.
So I made an installer that handles the initial setup.
The basic idea is:
curl -fsSL https://raw.githubusercontent.com/ANIRudH-lab-life/MINICODE/main/install.sh | bash
Then:
minicode --setup
and:
minicode
The default configuration uses Ollama and Qwen3 1.7B, so you can start experimenting with a local coding agent without needing a cloud API.
The funny part: small models are hard
Working with a 1.7B model has actually made me learn more about agents.
With a larger model, you can sometimes get away with a mediocre prompt or imperfect tool design.
With a small model, those problems become obvious.
If the tools aren't designed properly, it struggles.
If the context gets messy, it struggles.
If an error isn't explained clearly, it can get stuck.
That means I have to think much more carefully about how the agent architecture works.
And that's probably been one of the most valuable parts of building MINICODE.
It's still a work in progress
MINICODE is nowhere near finished.
There are still plenty of things I want to improve:
- Better tool calling
- Better context management
- More reliable coding
- Better error handling
- More model support
- Improved terminal UI
- Better memory
- More autonomous workflows
But I don't really see that as a problem.
The whole point of the project is experimentation.
I want to keep seeing how far I can push small local models.
Why I'm building this
For me, MINICODE is more than just another AI wrapper.
I'm interested in what happens when you take a relatively small model and give it the right environment, tools, and architecture.
Maybe the future of local AI isn't always about running the biggest model possible.
Maybe sometimes it's about making smaller models work better.
That's what I'm trying to find out.
If you want to check it out, the entire project is open source:
GitHub: https://github.com/ANIRudH-lab-life/MINICODE
I'd love to see what other people can do with it.
Top comments (2)
"The model doesn't have to do everything alone" is the right instinct, and a 1.7B is a great forcing function because it fails loudly the moment the harness is sloppy. The place these tiny-model agents usually break isn't code quality — it's the loop losing the plot: after four or five tool calls the model forgets what it was doing because the tool results have crowded out the original task. A couple of things that bought me a lot at small sizes: keep the task statement pinned at the top of every turn rather than trusting it to survive in history, and summarize tool output before it re-enters context (a 1.7B drowns in raw file dumps fast). Also worth measuring: how often does it re-read a file it already read? That rediscovery loop is pure wasted context at any size but fatal at 1.7B. How are you handling context when a single file is bigger than what the model can usefully hold — chunk-and-summarize, or hard truncate?
Building the confirmation loop before adding more capability is the right order, and "if I reject an action, the model gets that information and can try another approach" is the part most implementations get wrong , a rejected tool call usually just dies.
The 1.7B constraint is a good forcing function too. A small model with tight tools beats a large model with vague ones, because the tool surface is doing the reasoning.
One thing worth separating if you ever want this running unattended: the model and the execution environment don't have to live in the same place. Confirmation prompts are the right control when a human is watching. When nobody is, you want the shell itself to be disposable.
That's roughly what we built Cubes for at Krova Cloud :- clone a snapshot into a fresh box per run, full root inside, destroy it after. No GPU, so the model side stays where it is.
Are you planning an unattended mode, or is interactive the point?