This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
A few weeks ago my friend spent weeks preparing for a software engineering interview. They worked really hard, and then didn't get it. The thing that hurt most wasn't the "no". It was that nobody told them what went wrong. Interview rejections come with no feedback, so you're left replaying a 45-minute performance in your head.
I built EdgeMate for them: a private, offline mock interviewer that runs on their laptop. It runs a real interview (understand the problem, discuss an approach, code, debug, discuss complexity, get a debrief) and it never makes them feel stupid for being stuck.
Three things I cared about:
- It's gentle by default. The default "Friendly" persona opens with what you did well, gives hints freely, and never says "wrong".
- It never spoils the answer. The model is never shown the reference solution, so it can't leak it.
- It stays on their machine. Rejection is a vulnerable moment. The code they write and the things they say while stuck shouldn't sit on someone else's server.
Demo
Code
EdgeMate
Local Open-Source AI Coding Interviewer
"The code runner checks whether the solution works. The open-source AI handles the human part of the interview."
flowchart TD
subgraph Frontend["Frontend Layer (Dark-First Developer Theme)"]
UI["Web SPA / UI (FastAPI & Uvicorn / Gradio)\n• Progressive Stage Panels\n• Live Python Editor & Examples Drawer\n• Problem Catalog & Progress Dashboard"]
end
subgraph API["Backend & Orchestration Layer"]
ROUTER["REST API Routes (/api/sessions/*)"]
STATE["Session State Machine (InterviewSession)\n• State: UNDERSTANDING ➔ APPROACH ➔ REVIEW ➔ CODING ➔ TESTING ➔ DEBUGGING ➔ COMPLEXITY ➔ DEBRIEF\n• Real-Time Attempt Tracking & Timestamps"]
AGENT["Interviewer Agent (InterviewerAgent)\n• Context-Bounded Prompt Builder\n• Active Persona Adaptation\n• Socratic Guidance (Zero Solution Leaks)"]
end
subgraph LocalAI["Local Open-Source LLM (Ollama)"]
OLLAMA["Ollama Engine (Local CPU/GPU Inference)\n• Default: qwen2.5-coder:3b (Q4_K_M)\n• Alternatives: llama3.2:3b, mistral:7b\n• 100% Offline & Private"]
end
subgraph Execution["Deterministic Execution Engine (Zero LLM Hallucinations)"]
KB_LOADER["KB Loader (interview-kb/loader.py)\n• 119 Problems • 20 DSA Patterns\n• Pre-Materialized Hidden Tests & Large Generators"]
RUNNER["Test Runner…How I Built It
The one design rule: code checks, the model talks.
My first instinct was to ask an LLM "is this code correct?". That's exactly where small models are worst. They can't reliably trace pointer loops, big inputs, or off-by-one bugs, and they love to dump the full solution. So I split the job:
Candidate code ──► deterministic subprocess runner ──► pass/fail + a failure tag
│
Candidate chat ◄── local LLM (Ollama) ◄── pre-written follow-up question for that tag
-
Correctness is deterministic. Code runs in an isolated subprocess with a timeout against hidden tests. The LLM only receives
tests_passed,tests_totaland afailing_tagsuch asempty,duplicatesorlarge. -
The model only phrases the question. Every problem in my knowledge base has pre-written follow-ups mapped to failure tags, so the model turns "failed on
duplicates" into a kind, Socratic question. - A small model is enough. Because the model never judges code, a 3B model on an 8GB laptop does the job. I use Qwen2.5-Coder 3B via Ollama by default.
The interview itself is a state machine with 7 stages: understanding → approach → approach review → coding → Socratic debugging → complexity → 4-pillar debrief. The editor and the optimal Big-O stay hidden until the right stage, like a real interview. There's also a 3-level hint ladder (question → nudge → concrete invariant).
The knowledge base is 119 problems across 20 patterns, stored as plain JSON. Each has a reference solution, a brute-force solution, tagged tests, hints, and planted "common bug" solutions with the follow-up question for each. I checked it two ways:
- The reference solution must agree with an independent brute-force solution on hand-written tests and 300 random inputs per problem.
- In the real sandbox, every planted bug must fail at least one test carrying the tag it claims.
That second check caught my own mistakes: a "bug" that turned out to be a valid alternative algorithm, and bugs my tests couldn't actually catch. I fixed the tests and re-verified.
Stack: Python, FastAPI, SQLite for session history and progress, Ollama for local inference, plain files for personas and problems. Personas (Friendly / Standard / Silent) are markdown files with word budgets.
Why Does Open Innovation Matter?
1. Privacy where it matters most.
Someone processing a rejection is typing "I froze on this question" and their half-working code into a tool. With a hosted API, all of that goes to a server they don't control. With Ollama, nothing leaves the laptop. There's no account, no telemetry, and it works with no internet.
2. It runs on the hardware my friend actually has.
They have an 8GB laptop, so I couldn't pick a big model. The architecture makes that fine: deterministic tools do the judging and a 3B open model does the talking.
3. I can change how the interviewer behaves.
The interviewer's personality is a markdown file. I can make it warmer, stricter, or nearly silent for stress practice, by editing a text file. The problem bank is the same: when my friend gets asked something new next time, they add one JSON entry. No retraining, no subscription.
4. I could compare models instead of trusting one.
EdgeMate has a benchmark endpoint that compares two local models on latency and on whether they respect the tone and word limits.
5. Zero cost to practise.
Fifty practice rounds cost the same as one: nothing. For someone job hunting, that matters.
Where open beat closed, and where it didn't. A hosted frontier model would write a smoother conversation. But it would also cost money per session, send a vulnerable person's code and feelings to a third party, and tempt me to let it judge correctness, which is the part models get wrong. Open weights plus deterministic tools gave me something more trustworthy, more private, and cheaper for this one person.
Top comments (0)