BreakMyQuery: find the row that breaks your SQL
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
I built BreakMyQuery for a colleague(also a friend🫠) who is a new graduate hire and is struggling with SQL joins or let's say he's scared of SQL joins.
A query can pass the sample data while quietly dropping a customer who has no orders. For someone learning joins, seeing the missing customer gives them a concrete place to start investigating.
BreakMyQuery offers eight SQLite exercises covering joins, NULLs, aggregation, ties, arithmetic, dates, and anti-joins. The learner writes a query, runs it, inspects any counterexample, and tries a repair. Each counterexample shows the breaking dataset, the learner's result, and the expected result. My traps saves their mistakes so they can revisit a query through Retry.
I have demonstrated the app to my friend(or colleague). The goal is to help them connect the query they wrote to the rows it keeps, loses, or duplicates.
Demo
The first exercise asks for every customer's name and their number of paid orders, including customers with zero paid orders.
That one row exposes the problem. The learner can trace why Alice disappears and change their query. The app keeps the reference SQL out of the interface.
The hosted verification record documents the tested counterexample, accepted-query, error, history, and Retry paths.
Code
Source code, exercises, tests, and setup instructions.
The application code uses the MIT license. The model weights have their own licenses.
How I Built It
The stack is Python, Streamlit, SQLite, and Pydantic. The hosted demo uses the open-weight Gemma 4 26B A4B model through Google's Gemini API.
The app first checks the sample. If the results match, it searches for a breaking case:
- Propose. The model receives the schema, exercise question, learner query, and reference query. It proposes small datasets intended to make the queries disagree.
- Verify. The app validates the proposed data and executes both queries in SQLite. It compares results while preserving duplicate-row counts.
- Shrink. It removes rows while preserving the mismatch and database constraints.
- Practice. The learner sees the remaining data and differing results, then edits their own query.
Seeded random fuzzing generates additional datasets and provides a fallback when the model returns no usable proposal or the API is unavailable. The hunter does not use the exercise's authored trap fixtures.
What the evaluations showed
I evaluated model proposals and fuzzing separately against 14 deliberately wrong queries and eight equivalent alternatives.
| Method | Wrong queries exposed | False positives on alternatives |
|---|---|---|
| Gemma 4 via API | 14/14 | 0/8 |
| Using Local Llama 3.2 | 9/14 | 0/8 |
| Seeded random fuzzing | 14/14 | 0/8 |
The Gemma catalog evaluation used live API calls from the local checkout; the deployed app was checked separately.
Gemma's figures combine the primary run with one explicit follow-up after an unusable response. Covering the 22 cases took 23 model attempts. The Gemma report preserves that failed call; the Llama report records the separate local trial.
On this catalog, random testing caught five mistakes Llama missed. Keeping both methods gives the app another chance to find a breaking case.
I also found a teaching problem: Llama could produce a valid counterexample and still explain its cause incorrectly. I disabled generated hints and mistake labels in the demo. Gemma's teaching quality remains unevaluated. The tables and query results give the learner evidence they can inspect.
These measurements cover a small authored catalog. A passing run means no counterexample was found within the search budget. It does not prove the query correct.
Why Does Open Innovation Matter?
Open weights give this project a local deployment option. Once the dependencies, runtime, and model are installed, the Ollama workflow runs inference and SQLite on the learner's computer without a hosted inference API.
The public demo offers an easier starting point: my friend can open a link without installing a model. That convenience has a different data path. Submitted SQL goes to Google, and SQLite verification runs on the app server. The interface discloses this.
I could evaluate local Llama and Gemma through the API using the same SQLite checker. Another developer can inspect that checker, swap the model, and repeat the evaluation. The model's role is explicit: it proposes cases, and SQLite checks whether they demonstrate a mistake.
The counterexample approach was inspired by XData at IIT Bombay. BreakMyQuery applies that idea to a small practice tool for my friend: find a concrete case, inspect the rows, and work out the repair.
Prize Categories
Best Use of Gemma: Gemma 4 proposes the breaking datasets used by the hosted demo. SQLite verifies each accepted counterexample before the learner sees it.
Top comments (0)