DEV Community

Cover image for My Friend Needed Interview Practice. The AI Needed Tests.
Md Hasan Raza
Md Hasan Raza

Posted on Fully Autonomous

My Friend Needed Interview Practice. The AI Needed Tests.

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.

What I Built

My friend is preparing for campus placements at Trilogy, Assurant, and Joveo. I built Sparr, a local interview practice app, to bring their resume, projects, coding exercises, and explanations into the same session.

While testing it, I caught the interviewer giving bad advice.

A two-sum solution used a map to remember values and their indices. The model suggested replacing it with a set. The problem required the indices. Following that advice would throw away information needed to return the answer.

That failure made an existing design decision very concrete: the model can comment on a solution; the application keeps the execution results. A convincing explanation cannot turn a failed test into a pass. A passing test cannot make every sentence of AI feedback correct, either.

For my friend, a practice session works like this:

  1. Bring your own material. Import a resume, review the extracted details, and attach a public GitHub project, research document, notebook, or financial model.
  2. Work through an interview. Choose a role, answer a question, run a coding solution where applicable, and explain your reasoning. The model receives the answer and observed results before asking a follow-up.
  3. Return to the attempt. Reopen the saved session to see the answer, checks, and feedback together.

The available tracks span software engineering, quant, finance, markets, and ML. That breadth matters because a project discussion or a finance case needs a different kind of practice from a DSA exercise.

Editable company templates start with my friend's three targets, plus JPMorganChase and Goldman Sachs. They provide study context that the candidate can change; they are not banks of actual company questions.

My friend's initial feedback has been positive. I haven't measured a placement outcome or improvement in interview performance yet.

Demo

57 seconds: a resume, a coding attempt, and a finance-project discussion.

Open the video player · Read the walkthrough transcript

At 0:16, a Python solution runs against six test cases. At 0:18, Qwen's feedback appears alongside the results. At 0:38, the session moves to explaining assumptions in an imported financial model.

The recording is silent and uses a synthetic resume and forecast. Code execution and local model responses are real. It shows the app's practice flow; SEB lockdown is documented separately.

A Python solution with six passing tests and local Qwen feedback in Sparr

The checks and the model's advice remain visible together, so the candidate can examine both.

Code

Source code, screenshots, and setup instructions on GitHub

Sparr is MIT-licensed and runs from source on macOS. Windows support is planned. With Node.js 24+, Python 3, and Git installed:

git clone https://github.com/Hasan72341/sparr.git
cd sparr
npm ci
npm run build
npm start
Enter fullscreen mode Exit fullscreen mode

Open http://127.0.0.1:4318. Guided practice works immediately without a model. For local AI feedback, install Ollama, run ollama pull qwen2.5:7b, and select it in Sparr's settings. The local-model guide records the tested hardware, setup, and limitations.

How I Built It

React provides the workspace; Express and SQLite handle the local service and saved sessions. Python and JavaScript answers run in a macOS sandbox. Qwen2.5 7B runs through Ollama in the demonstrated setup.

The responsibilities are explicit:

Part Responsibility
Interview controller Select questions, manage timing and difficulty, and preserve the session
Execution and numeric checks Run supported code tests and compare numeric answers with reference values
Model provider Use the answer, context, and observed results to return feedback and a follow-up

The model cannot execute project files or rewrite test results. Public repositories can be inspected, and supported project tests use dependency-free Python or Node test suites. Imported notebooks and spreadsheets are inspected without executing cells or recalculating formulas.

The first local-model failure was structural. Asking for JSON produced JSON—but sometimes no follow-up. That is a broken interview even if the HTTP request succeeds.

I added a regression test, sent Ollama an explicit output schema, and retained validation before saving the response. This is the schema from the provider implementation:

const feedbackFormat = {
  type: "object",
  properties: {
    feedback: { type: "string" },
    followup: { type: "string" },
    observation: { type: "string" },
  },
  required: ["feedback", "followup", "observation"],
  additionalProperties: false,
};
Enter fullscreen mode Exit fullscreen mode

The second failure was the bad coding advice. The smaller 1.5B model could satisfy that schema and still recommend losing the indices. I moved the documented setup to Qwen2.5 7B and repeated the live checks.

The larger model still made mistakes. In the finance demo, it referred to a 1% margin increase when the source moved from 20% to 22%—two percentage points. I kept that limitation in the model evidence. Tests establish behavior on their cases; model feedback still needs judgment.

Testing the application found different problems. A late request could restore deleted data. A discussion response could still say “connect a model” after the model had answered. GitHub CI caught a sandbox test tied to a file on my own Mac. Each became a fix with a regression check.

The recorded macOS CI run passed 102 unit/integration tests, 44 browser tests, and 54 native checks. A separate real-model check verifies feedback, follow-ups, and saved reports; it fails if the app silently falls back to Guided practice.

For exam-style practice, Sparr also generates Safe Exam Browser launch links and offers an optional native guardian. The strict-mode guide explains the local device policy and its limits. Ordinary practice needs neither SEB nor camera or microphone access.

Why Does Open Innovation Matter?

An interview session can contain someone's resume, their project work, and an answer they are still struggling to explain. I wanted my friend to be able to practise with that material on their own laptop.

In the demonstrated setup, Qwen runs locally with Ollama's cloud features disabled. After downloading the runtime and model weights, feedback needs no hosted inference account. Fetching a GitHub repository still needs internet access.

Model replacement also became a practical part of development. Moving from the smaller Qwen model to 7B left the exercises, code runner, and saved history in place. My friend can try another installed model through the same interface. That choice comes with real trade-offs: the tested 7B download is about 4.7 GB, and local inference uses the Mac's memory and compute.

The application is open too. Someone can inspect how an answer was checked, contribute an edge case, improve a finance explanation, or add a provider. Useful contributions can be as small as a test that catches misleading feedback.

The outcome I care about is a specific one: my friend comes back to an answer and can explain a decision more clearly than before. Sparr gives us a place to practise that—and enough of the record to question the interviewer as well.

My Agent Session

This curated Codex transcript contains 37 original prompts and build updates, including the local-model failures, SEB integration, and verification work. Shortened messages are labeled; intermediate results reflect that point in development. Credentials, private machine paths, and raw tool logs are omitted.

Building Sparr: resume interviews, local Qwen feedback, and macOS testing
You

make sure resume could be added and this parse resume clone proejcts run it etc and ask from that along with cp dsa

You

make sure it can tackle finance companies too not just dsa, its for iitians

Agent

Sparr should target IIT campus placements across software, quant, finance, and ML roles. The interview should adapt to the role, the candidate’s resume, and the recruiting context.

Finance needs several distinct tracks. A quant researcher, a bank’s software engineer, and an investment banking analyst should get substantially different interviews.

[Remaining paragraphs omitted from this excerpt.]

You

we build prod ready for macos first, windows support will be added later mention in readme

You

now dont ask me anyhing complete it then do end to end testing and then fix the bugs etc. it should usable from start

Agent

The question bank now covers all seven tracks, including finance and project discussions. I’m implementing evidence-based feedback: numeric answers get independent checks, coding answers get executed tests, and open-ended answers remain “needs review” unless an assessment provider can evaluate them.

Agent

The first execution tests caught a real macOS issue: Node tried to read a Homebrew OpenSSL configuration outside the sandbox. I’ve removed that dependency from the runner environment. Python tests already confirm that candidate code cannot read home files or open network connections.

Agent

Browser testing is underway. An independent review found issues I’m fixing before the final run: follow-up answers could be compared with the original numeric answer, an in-flight project test could undo deletion, and some parsing/provider operations needed tighter resource limits.

Agent

The review fixes are in place: follow-ups now have a separate response path, deletion is guarded against in-flight work, and parsing/execution have time, memory, storage, and output limits. I also replaced the unsafe Codex path with a Claude Code adapter that disables tools and customizations; Codex stays unavailable until its isolation is verified.

Agent

All 18 browser tests now pass in Chromium and WebKit, including resume confirmation, draft recovery, code execution, follow-up handling, reports, exports, and camera/microphone cleanup. I’m finishing the startup instructions and remaining live checks for repository import and the configured agent adapter.

Agent

Document imports now use the same isolated parser as resumes. Notebooks expose cell source without executing it, and financial models show formulas and cached values without treating them as verified results. I also fixed vulnerabilities in the new parser dependencies; the dependency audit is clean.

Agent

The new finance-document flow passes in Chromium and WebKit. Review also found a deletion bug: an upload already in progress could finish after “delete all data” and restore a file. I’ve fixed the request ordering and am adding a regression check.

You

now complete this project by removing ai slops. Proper readme etc

Agent

The dashboard now has real resume, project, and provider shortcuts instead of the decorative code panel. I’ve rewritten the README with the first-run steps and current limitations, and replaced the planning reports with architecture, contribution, roadmap, and troubleshooting docs.

You

make sure the site could generate seb link with so seb could be triggered

Agent

SEB supports a launch link that downloads its configuration from Sparr. I’m adding that to interview setup, with a download fallback. Your selected options will carry over, and the timer will start only after you click “Start practice” inside SEB.

Agent

The full flow passed: generated link → SEB setup → finance answer → live Claude Code feedback → saved report. Sparr is now running with the launch controls at port 4318, and the README includes the instructions and screenshot. I’m shutting down the test VM.

Agent

I found and fixed a storage-failure bug: an invalidated attempt could briefly be saved as completed. Termination is now atomic, and SEB cleanup still runs if saving fails. All 28 enforcement tests pass; next I’m checking the generated configuration against SEB itself.

Agent

The native guardian correctly identifies the guest as a VM and rejects its camera as unsupported. The real SEB run also exposed a configuration-key initialization issue that browser fixtures didn’t catch; I’m fixing that before finishing.

Agent

SEB now supplies a key that matches the server’s generated configuration, and admission and heartbeats work in the guest. The native fallback also successfully closed that verified SEB process. I’m checking one final extra-camera violation through the live SEB page.

Agent

The live SEB test passed: a simulated second camera ended the attempt, preserved the saved answer and reason, and closed SEB. The final browser rerun also exposed a separate history race—a late response restored a deleted session on screen. I reproduced it and am applying the fix before the last test run.

Agent

The final run passed: 97 unit/integration tests, 42 browser tests, and 54 native self-checks. The existing VM and lab servers are stopped; Sparr remains available locally with Claude Code configured. The verification document records the real SEB results and the remaining physical-hardware checks.

You

complete gap and friend feedback is good, they prepare for trilogy, assurant, joveo etc with add extra ur self and remove and add chalege notes etc to git ignore

Agent

I’ll add editable preparation templates for Trilogy, Assurant, and Joveo, plus JPMorganChase and Goldman Sachs for broader practice. These will be suggested study contexts, not claims about their actual interview questions.

[Remaining paragraphs omitted from this excerpt.]

Agent

The five company templates now work and pass browser checks in Chromium and WebKit. The local model is downloaded, with Ollama cloud features disabled; I’m checking real feedback and saved reports next. Your positive-feedback report is recorded, and the challenge notes are now ignored by Git.

Agent

The first real-model check exposed a bug: Ollama returned JSON but omitted the required follow-up field. I’m using the systematic-debugging skill to fix the request contract, and a regression test now reproduces the missing schema constraint.

Agent

The local model now returns valid feedback and follow-ups, and both test sessions saved their reports. The coding answer passed all six cases. The full suite also passes: 98 unit/integration tests and 44 browser tests.

[Remaining paragraphs omitted from this excerpt.]

Agent

Reviewing the demo revealed a more serious limitation: the 1.5B model suggested a code change that would lose the required indices. I’m checking a stronger 7B model before recommending it; its download is about 4.7 GB.

[Remaining paragraphs omitted from this excerpt.]

Agent

The stronger model passed the live checks. I also found that the prompt included the app’s stock “needs review” feedback, which could bias the model toward repeating it. I’m removing that canned judgment from the model input and keeping the execution evidence.

[Remaining paragraphs omitted from this excerpt.]

Agent

The demo exposed one more bug: discussion feedback still said “connect a model” after the model had responded. I’ve reproduced it in a test and am fixing that message before publishing the demo.

Agent

GitHub caught a portability bug in a test: it tried to read a file from this Mac’s home directory, which doesn’t exist on the CI runner. I’m changing it to create and clean up its own sample file, so it verifies sandbox denial on any Mac.

Agent

The full GitHub macOS run passed, including both browser engines, native checks, and the dependency audit. I’m adding that result to the verification record. The code and demo are published; the remaining submission step is DEV sign-in.

You

physical device passed

Agent

Recorded physical-device validation as passed based on your confirmation. Updated the README, verification record, and related docs, then pushed the changes.

You

make sure video works directly, embed vid in md

Agent

The hosted video plays and seeks successfully in both Chromium and WebKit, including at mobile width. I’ve updated the draft with the inline player and the README with the direct player link.

Agent

Your submission is live on DEV.

Updated and verified through DevRelay, with your cover, challenge tags, and AI disclosure. Embedded video playback passed in Chromium and WebKit.

Built and documented with coding-agent assistance.

Top comments (0)