<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: sohith</title>
    <description>The latest articles on DEV Community by sohith (@sohith2007).</description>
    <link>https://dev.to/sohith2007</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1390426%2F9e965c1d-fb10-4041-b83c-289c4aa685c1.png</url>
      <title>DEV Community: sohith</title>
      <link>https://dev.to/sohith2007</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sohith2007"/>
    <language>en</language>
    <item>
      <title>Falsify — Find the Smallest Input That Breaks Your C++ Solution (Local AI, No API Keys)</title>
      <dc:creator>sohith</dc:creator>
      <pubDate>Mon, 05 Oct 2026 05:51:37 +0000</pubDate>
      <link>https://dev.to/sohith2007/falsify-find-the-smallest-input-that-breaks-your-c-solution-local-ai-no-api-keys-5ab3</link>
      <guid>https://dev.to/sohith2007/falsify-find-the-smallest-input-that-breaks-your-c-solution-local-ai-no-api-keys-5ab3</guid>
      <description>&lt;p&gt;What I Built&lt;br&gt;
Falsify is a local, offline-first debugging tool for competitive programmers. You give it a problem statement, your C++ solution, and a sample test — it uses an on-device open-source LLM (Gemma via Ollama) to write a correct brute force and a random generator. It then stress-tests your code against the brute force until it finds the exact smallest input that breaks your solution. When it finds a bug, it gives you three progressive hints instead of just the answer.&lt;/p&gt;

&lt;p&gt;Every competitive programmer knows the pain of passing all sample cases but getting a "Wrong Answer" on submission. The classic fix is stress testing, which means writing three separate programs every time you're stuck. I built this for a friend (and myself) who spends more time setting up stress tests than actually solving problems.&lt;/p&gt;

&lt;p&gt;Demo&lt;br&gt;
Since this is a local tool that runs heavy LLM models on your own machine, the best way to experience it is by running the web UI locally:&lt;/p&gt;

&lt;p&gt;bash&lt;/p&gt;
&lt;h1&gt;
  
  
  Get the model
&lt;/h1&gt;

&lt;p&gt;ollama pull gemma4:latest&lt;/p&gt;
&lt;h1&gt;
  
  
  Clone and run the UI
&lt;/h1&gt;

&lt;p&gt;git clone &lt;a href="https://github.com/Sohith2007/falsify.git" rel="noopener noreferrer"&gt;https://github.com/Sohith2007/falsify.git&lt;/a&gt;&lt;br&gt;
cd falsify&lt;br&gt;
pip install -r requirements.txt&lt;br&gt;
python -m counterexample.cli ui&lt;br&gt;
Open &lt;a href="http://localhost:8000" rel="noopener noreferrer"&gt;http://localhost:8000&lt;/a&gt; to use the tool.&lt;/p&gt;

&lt;p&gt;Here is a look at the Web UI finding a deliberate integer overflow bug in real-time:&lt;/p&gt;

&lt;p&gt;[UPLOAD IMAGE 1 HERE: ui_full_page_1791139197343.png]&lt;/p&gt;

&lt;p&gt;And tracking recurring mistakes in the Bug Report modal:&lt;/p&gt;

&lt;p&gt;[UPLOAD IMAGE 2 HERE: bug_report_modal_1791139005483.png]&lt;/p&gt;

&lt;p&gt;Code&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Sohith2007" rel="noopener noreferrer"&gt;
        Sohith2007
      &lt;/a&gt; / &lt;a href="https://github.com/Sohith2007/falsify" rel="noopener noreferrer"&gt;
        falsify
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;🔍 Counterexample&lt;/h1&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;A local, offline coach that finds &lt;strong&gt;the smallest input that breaks your competitive-programming solution&lt;/strong&gt;, then gives you hints instead of the answer.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href="https://www.python.org/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/a12590304d4f5a572eaba3900cb6460f580cbe63756bc998fde4d0ac705ad551/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f707974686f6e2d332e31312532422d626c75653f6c6f676f3d707974686f6e266c6f676f436f6c6f723d7768697465" alt="Python"&gt;&lt;/a&gt;
&lt;a href="https://github.com/Sohith2007/falsify/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/784362b26e4b3546254f1893e778ba64616e362bd6ac791991d2c9e880a3a64e/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d4d49542d677265656e2e737667" alt="License: MIT"&gt;&lt;/a&gt;
&lt;a href="https://ollama.com/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/f7c817c808f6bc4aac97d91c74869e17f25331cdc750a31f67cad4541e5e027c/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f706f776572656425323062792d4f6c6c616d612d6f72616e67653f6c6f676f3d6f6c6c616d61" alt="Ollama"&gt;&lt;/a&gt;
&lt;a href="https://learn.microsoft.com/en-us/windows/wsl/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/11a2737f84667d909c0ddcab713c4a305140f697ffd466f16ab83d0c405216b5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f706c6174666f726d2d4c696e757825323025374325323057534c2d6c69676874677265793f6c6f676f3d6c696e7578" alt="Platform"&gt;&lt;/a&gt;
&lt;a href="https://hacktoberfest.com/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/a2b318140d58f56862ecd5581d006b961f44164e891dbd5fd6834b93f824756f/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4861636b746f626572666573742d323032362d626c756576696f6c6574" alt="Hacktoberfest"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/Sohith2007/falsify/demo/demo.gif"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2FSohith2007%2Ffalsify%2FHEAD%2Fdemo%2Fdemo.gif" alt="Counterexample demo"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-alert markdown-alert-caution"&gt;
&lt;p class="markdown-alert-title"&gt;Caution&lt;/p&gt;
&lt;p&gt;Counterexample compiles and executes &lt;strong&gt;model-generated C++ code&lt;/strong&gt; on your machine
All binaries run inside a temp directory with strict CPU-time (5 s) and memory (512 MB) limits
&lt;strong&gt;Never run as root or admin.&lt;/strong&gt; For stronger isolation, pass &lt;code&gt;--sandbox docker&lt;/code&gt; (see &lt;a href="https://github.com/Sohith2007/falsify#-safety" rel="noopener noreferrer"&gt;Safety&lt;/a&gt;).&lt;/p&gt;
&lt;/div&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;💡 What It Does&lt;/h2&gt;

&lt;/div&gt;
&lt;p&gt;A friend is stuck on "Wrong Answer on test 7" and can't see the test. &lt;strong&gt;Counterexample&lt;/strong&gt; takes the problem statement, the friend's C++ code, and the sample tests. Gemma writes (a) a slow but obviously-correct brute force and (b) a random input generator. A Python script compiles all three programs, runs them against each other on thousands of small inputs, and stops at the first input where the friend's code disagrees with the brute force. Gemma then explains…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Sohith2007/falsify" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;How I Built It&lt;br&gt;
Falsify is built entirely around Gemma 4 (or any other open-weight model) running locally via Ollama.&lt;/p&gt;

&lt;p&gt;The core engine is a Python pipeline that orchestrates the model. It asks the LLM to generate C++ code for a brute-force solution and a test case generator. If the generated code fails to compile, the pipeline feeds the compiler error back to the LLM and retries. Once compiled, it runs a massive stress-testing loop (smallest inputs first) to find a counterexample. Finally, it uses the LLM again to classify the bug and generate progressive hints for the user.&lt;/p&gt;

&lt;p&gt;The frontend is a FastAPI server using Server-Sent Events (SSE) to stream the pipeline's progress to a vanilla HTML/JS dark-mode web UI.&lt;/p&gt;

&lt;p&gt;Why Does Open Innovation Matter?&lt;br&gt;
For a tool like this, open-source AI isn't just a nice-to-have; it's a requirement:&lt;/p&gt;

&lt;p&gt;Privacy &amp;amp; Local Execution: Competitive programmers cannot send unpublished contest solutions or homework assignments to closed APIs like OpenAI or Anthropic. Everything in Falsify runs in a local process. The LLM output never touches a network.&lt;br&gt;
Offline Capability: Competitive programming often happens in exam halls, on planes, or in places with restricted internet. Once ollama pull is done, Falsify works completely offline.&lt;br&gt;
No Lock-in or Costs: Stress testing can require dozens of LLM generations per problem if retries are needed. Running this on a paid, closed API would be prohibitively expensive for students.&lt;br&gt;
Model Swappability: Because it uses Ollama's open API, users can instantly swap gemma4 for qwen2.5-coder or deepseek-coder without changing the application code.&lt;br&gt;
My Agent Session&lt;br&gt;
I built this project with the help of Google's Antigravity agent in my local IDE. The agent helped scaffold the Python pipeline, handle the Windows-specific subprocess quirks for C++ compilation, and built the FastAPI backend and vanilla JS frontend from scratch.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
