This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Hawk is a debate room for AI models. You ask one question, and three models answer it side by side, each in its own terminal-style window. After the first round, the models read each other's answers, respond to them, and vote. You can run up to three rounds and watch the answers change.
I built it for Tamara Surla, who is a developer, basically he uses AI so much for his work the problem is that he does not know which model to trust as it can backfire his progress so to build trust I have build this project.
His workflow looked like this:
- He asked the same question in three different chat tabs, one for each model.
- He copied answers back and forth to compare them.
- When the models disagreed, he had no quick way to tell which one was wrong, because each sounded confident.
- Paid plans for several models weren't an option for him, so he was juggling free tiers and limits.
How Hawk fixes it:
- He asks once, and three models answer at the same time in separate terminal-style windows.
- In the next round, each model sees the other answers and can point out mistakes, so errors get challenged instead of going unnoticed.
- The models vote, so he can see where they agree and where they split. A split is the signal to double-check that part.
- He picks the providers and models for each seat, so he can run it on free models.
Demo
- Live app: https://project-hawk-frontend.onrender.com/hawk
- Backend health check: https://project-hawk-backend.onrender.com/health (to check your backend health)
It runs on Render's free tier, so the first load after a quiet period can take about a minute while the services wake up.
Code
Hawk
Multi-LLM panel debates for coding tasks - several models discuss, vote, and synthesize a consensus while you watch live.
Hawk runs an interactive CLI with two modes:
- Driver chat - talk to a primary model for day-to-day coding help
-
Hawk debate - summon a panel (
/hawk) that argues in rounds, casts Approve/Reject/Abstain votes, and produces a synthesized conclusion
No paid subscriptions required. A free OpenRouter key is enough to run a 3-model panel.
Features
- Live multi-pane debate UI (Textual) with streaming responses and vote tallies
- First-run setup wizard - walks you through provider options before anything else runs
- Free-tier friendly - Groq, Cerebras, Agy/Gemini, Mistral, NVIDIA NIM, Ollama (local or Ollama Cloud), OpenRouter
:freemodels - Subscription CLI support - Claude Code and Codex when you already have them
- Optional tool use for API providers - read/write files, grep, run commands when explicitly enabled
- Rate-limit resilience - automatic…
How I Built It
-
Frontend: Next.js. A server-side route (
/api/debate) talks to the backend, so the browser never calls the engine directly. -
Backend: FastAPI (Python).
POST /debate/roundstreams each round back as server-sent events, so answers appear live as the models write them. - Seats: each debate has three "seats". A seat is a provider, a model and an API key. The backend supports OpenRouter, Groq, Gemini, Mistral, NVIDIA, TokenHarbor, Ollama, Claude and Codex, so you can mix them freely.
- Bring your own key: keys are sent with each request and are not stored on the server.
- Safety limits: requests are validated, with exactly three seats, a capped question length and at most two earlier rounds of context.
-
Deployment: two Render web services (Node for the frontend, Python for the backend), both deployed from
main. - Open models: [list the models you used. For example, the demo uses free OpenRouter models, and Ollama can run open-weight models locally.
Inside the Discussion Room
Hawk is a place where you seat a panel of three AI models, ask one question, and watch them work through it together.
You start on a clean home screen. In the left sidebar you can open Models & keys, and on the right you can click Choose models. Once your panel of three is seated, the question box unlocks. Until then it stays locked, so you can't start half-set-up.
Click New discussion and ask your question once. Each model answers in its own terminal-style window, with the replies streaming in live. You can compare them side by side instead of switching between tabs. In the following rounds the models read each other's answers and respond to them, so mistakes get challenged and you can see where the panel agrees and where it splits. Your debates are kept in History in the sidebar.
I kept it to what your screenshots show. If voting or History works differently from what I described, change those lines to match your app.
Run it locally
Hawk runs on your own computer with no paid hosting. It has two parts that run side by side: the engine (the Python backend) and the website (the Next.js frontend).
REQUEST:
If you see "The Hawk engine failed", it almost always means a model provider rejected the request, not that Hawk is broken. First, check your API key: make sure it's pasted in full, matches the provider of that seat's model, and is still active in your provider's dashboard. Then check its limit, because free keys have usage caps and some models have no free allowance at all. Wait a minute and retry, or swap the failing seat for another model in Models & keys. If you see "The Hawk engine isn't running" instead, the app can't reach the engine, so start it (locally) or wait about a minute for it to wake up (hosted).
If you like the concept please give me a star ⭐️ on github Alejandro (Project-Hawk)
Prize categories which I am applying or falling under are :
Gemma
Gemma is one of the open-weight models I seat at the Hawk table. It takes one of the three seats in a debate, answers the same question as the other two models, then reads their answers and responds to them. Hawk is built to compare models side by side, so Gemma gets tested directly against other models on the same question, and it can be swapped in or out of any seat without changing the code.
Render
Both halves of Hawk run on Render as two web services. The Python (FastAPI) engine streams debate rounds live, and the Next.js website talks to it. Both services deploy straight from my repo, so every push to main rebuilds and redeploys them automatically. The live demo linked in this post runs on Render.
GitHub
The whole project lives in a public GitHub repository, from the engine in src/ to the website in frontend/. GitHub is the source of truth for the project, and pushing to main is what triggers each deployment on Render.





Top comments (0)