What I Built
Falsify is a local, offline-first debugging tool for competitive programmers. You give it a problem statement, your C++ solution, and a sample test — it uses an on-device open-source LLM (Gemma via Ollama) to write a correct brute force and a random generator. It then stress-tests your code against the brute force until it finds the exact smallest input that breaks your solution. When it finds a bug, it gives you three progressive hints instead of just the answer.
Every competitive programmer knows the pain of passing all sample cases but getting a "Wrong Answer" on submission. The classic fix is stress testing, which means writing three separate programs every time you're stuck. I built this for a friend (and myself) who spends more time setting up stress tests than actually solving problems.
Demo
Since this is a local tool that runs heavy LLM models on your own machine, the best way to experience it is by running the web UI locally:
bash
Get the model
ollama pull gemma4:latest
Clone and run the UI
git clone https://github.com/Sohith2007/falsify.git
cd falsify
pip install -r requirements.txt
python -m counterexample.cli ui
Open http://localhost:8000 to use the tool.
Here is a look at the Web UI finding a deliberate integer overflow bug in real-time:
[UPLOAD IMAGE 1 HERE: ui_full_page_1791139197343.png]
And tracking recurring mistakes in the Bug Report modal:
[UPLOAD IMAGE 2 HERE: bug_report_modal_1791139005483.png]
Code
🔍 Counterexample
A local, offline coach that finds the smallest input that breaks your competitive-programming solution, then gives you hints instead of the answer.
Caution
Counterexample compiles and executes model-generated C++ code on your machine
All binaries run inside a temp directory with strict CPU-time (5 s) and memory (512 MB) limits
Never run as root or admin. For stronger isolation, pass --sandbox docker (see Safety).
💡 What It Does
A friend is stuck on "Wrong Answer on test 7" and can't see the test. Counterexample takes the problem statement, the friend's C++ code, and the sample tests. Gemma writes (a) a slow but obviously-correct brute force and (b) a random input generator. A Python script compiles all three programs, runs them against each other on thousands of small inputs, and stops at the first input where the friend's code disagrees with the brute force. Gemma then explains…
How I Built It
Falsify is built entirely around Gemma 4 (or any other open-weight model) running locally via Ollama.
The core engine is a Python pipeline that orchestrates the model. It asks the LLM to generate C++ code for a brute-force solution and a test case generator. If the generated code fails to compile, the pipeline feeds the compiler error back to the LLM and retries. Once compiled, it runs a massive stress-testing loop (smallest inputs first) to find a counterexample. Finally, it uses the LLM again to classify the bug and generate progressive hints for the user.
The frontend is a FastAPI server using Server-Sent Events (SSE) to stream the pipeline's progress to a vanilla HTML/JS dark-mode web UI.
Why Does Open Innovation Matter?
For a tool like this, open-source AI isn't just a nice-to-have; it's a requirement:
Privacy & Local Execution: Competitive programmers cannot send unpublished contest solutions or homework assignments to closed APIs like OpenAI or Anthropic. Everything in Falsify runs in a local process. The LLM output never touches a network.
Offline Capability: Competitive programming often happens in exam halls, on planes, or in places with restricted internet. Once ollama pull is done, Falsify works completely offline.
No Lock-in or Costs: Stress testing can require dozens of LLM generations per problem if retries are needed. Running this on a paid, closed API would be prohibitively expensive for students.
Model Swappability: Because it uses Ollama's open API, users can instantly swap gemma4 for qwen2.5-coder or deepseek-coder without changing the application code.
My Agent Session
I built this project with the help of Google's Antigravity agent in my local IDE. The agent helped scaffold the Python pipeline, handle the Windows-specific subprocess quirks for C++ compilation, and built the FastAPI backend and vanilla JS frontend from scratch.

Top comments (0)