DEV Community

Onur Kesim
Onur Kesim

Posted on AI-assisted

I said SymPy recomputes my app's answer keys. Then I read the gate.

This week my maths practice app, Tuzak Avcısı: LGS Matematik, went public on Google Play (store page checked without signing in on 7 October 2026). It drills multiple-choice questions for the LGS, the national exam Turkish students sit at the end of 8th grade. Every question is generated from parameters, so there is no fixed answer sheet a teacher could read through once. Whether the marked answer is right has to be checked by code.

My CV, and a LinkedIn post from the same week, said:

Answer keys are recomputed by an independent third solver (SymPy) — 1,475 items.

I sat down to write that up properly, opened the gate to quote it, and found the sentence isn't true. This post is about what the gate actually checks, what it can't see, and where real recomputation does exist.

The repository is private, so I can't link it. All numbers below were measured on 7 October 2026 on commit 5fe5cef, whose app file is byte-identical to the one inside the released APK (SHA-256 f14c132c…), unless a line says otherwise.

Where the sentence came from

The test runner calls the gate under this comment (translated from Turkish):

independent sympy code gate — extract the choices (jsdom), then re-solve the key with sympy (exactly one correct + no distractor value-equal to it)

The first half of that comment became my CV. The part in brackets is what the code does.

What the gate actually checks

It has two steps. A Node script loads the app in jsdom, runs every multiple-choice mode and phase through 40 seeds (the CI setting; my local runner uses 16), and dumps each question's options with an ok flag. A Python script then parses the options with SymPy and asks two questions:

  1. Is exactly one option marked correct?
  2. Is any distractor equal in value to the correct option? (simplify(a - b) == 0 for expressions, a solveset comparison for inequalities.)

That's all. It never reads the question text, so it never solves the problem. If a generator computes 42 and marks 42 correct when the answer is 41, the gate has nothing to compare 41 to.

My run on the release commit, with SymPy 1.14.0:

Denetlenen: 1560 | UYGULANMAZ (semantik): 160 | KASITLI biçim-tuzagi (allowlist): 56 | GERCEK FAIL: 0
Enter fullscreen mode Exit fullscreen mode

That reads: 1,560 checked, 160 not applicable, 56 intentional form traps (counted inside the 1,560), 0 real failures. 1,720 questions came out of 42 modes. One multi-select phase yields no options at all and is skipped without a word; that's a second blind spot I found while writing this. 160 have answers SymPy can't compare by value (words, drawn options, side-length triples, letter orderings), so only the one-correct check applies to them. The 56 "intentional" ones have the right value in the wrong form: 13×10⁻⁴ instead of 1.3×10⁻³ (marked wrong), or 2√28 instead of 4√7 (deliberately not counted wrong). The 1,475 on my CV was an older count.

It still earns its place

Duplicate-value options are a real bug class in generated multiple choice: two buttons that are both right, one of which the app calls wrong. This gate covers that class across every mode it can parse, except the three form-trap modes, where any value-equal option is allowed.

It also once went red because of itself. The release's build report (30 September 2026) records an item whose options included "20 ve 9" and "12 ve 15". Ve is Turkish for "and". SymPy's implicit multiplication split ve into two unknowns, v·e: 20·v·e·9 = 180·v·e = 12·v·e·15. Two different answers, declared equal. The fix taught the parser that "a ve b" is an unordered pair of numbers, and that any other option containing ve, veya ("or") or ile ("with") is compared as plain text.

What it can't see, measured

In this app a wrong key doesn't crash anything. A student picks the right answer, the app marks it wrong and may open a card explaining a trap they didn't fall into. Nobody files a bug report. A 14-year-old just learns to distrust a method that works.

To put a number on the blind spot, I simulated a wrong answer key. For each mode and phase, I added 1 to every correct option that was an integer in the extracted file and re-ran the gate. 21 mode-phases had at least one integer key in the sample (from 6 to 40 of 40).

  • 6 stayed green: exponent rules, scientific notation, multiplying and dividing roots, rational and irrational numbers, cylinder volume and cylinder surface area. In the two cylinder modes all 40 of the 40 sampled keys were now wrong. In two of the six, scientific notation and the root mode, a collision would have been waved through anyway, because value-equal options are allowed traps there.
  • 15 went red, for the wrong reason. Many distractors are deliberate off-by-one traps, so "correct + 1" often landed exactly on a distractor, and the gate reported "distractor equals the correct answer". That's detection by coincidence, with a message that points at a different bug.

Counted per question, 315 of the 580 wrong keys (54%) passed silently.

This experiment edits the gate's input, not the generators, so on its own it doesn't show that any generator is wrong. But the release code records a case where one was. In the exponent-rules mode, a question showed (−b)ᵐ·bⁿ while the key used the parity of m+n. For m = n = 3 the real value is negative and the key said positive. A comment in the generator records that the SymPy gate missed it, because the options were consistent with each other and wrong only against the question. A manual visual review caught it.

Where recomputation does exist

Two places, both narrower than my CV implied.

The GCD/LCM gate. It carries its own Euclid and never calls the app's own GCD/LCM functions. With a mutant that shows the LCM where the GCD belongs, it fails on the first pair: (30,18): shown GCD=90 ≠ egcd=6. Without the mutant it passes on 3,000 generated pairs.

The GCD/LCM word problems. A separate solver reads each question's raw parameters and text and recomputes the answer by its own route; for the "when do the two bells ring together for the n-th time" type it brute-forces minute by minute. My run: 8,000 questions, 0 mismatches. It isn't in the release runner or CI yet; only its mutants run there.

Both are about GCD and LCM. One more gate is described in its comment as an independent validity check for triangle-inequality questions; I haven't read its assertions yet, so I'm not counting it. The other gates I read check structure, variety and trap behaviour, not the key. For the two cylinder modes I found no file in the test folder that recomputes an answer. I didn't read all of the gates line by line, so treat the rest as "not found", not "absent". There is also a golden-answer harness that lives outside the repository; I didn't re-run it for this post, so I'm not counting it.

The rest of that CV line

I checked the other numbers against the release commit and its build reports the same evening.

CV said The release commit says
25-layer audit chain 55 gate calls in the runner
8 mutation tests 224 mutant pairs: 201 code, 23 visual; 9 exempt, 0 survived (build report, 1 October 2026)
49 question generators Couldn't define it. 56 functions start with gen, and some are helpers. Dropped.
Ad/analytics scan with a positive control Holds (release scan, 1 October 2026): 13 SDK signatures, 0 found, no INTERNET permission. The scan must find the Capacitor core class path, or it refuses to report "clean".

Most of those numbers had gone stale in the safe direction: the chain grew after I wrote them. The SymPy sentence went the other way. It described more checking than exists.

What I changed

  • CV: "A SymPy gate checks 1,560 sampled multiple-choice questions (39 of 42 topics) for exactly one option marked correct and no distractor equal in value to it, deliberate notation traps excepted. For GCD/LCM word problems, an independent solver recomputed the answer keys (8,000 questions, 0 mismatches)." The "49 generators" line is gone.
  • LinkedIn: edited with the same correction.
  • A rule for myself: a gate's name and its comment are claims. Before I quote a gate anywhere, I read its assertions.

The gap itself is still open. Volume, surface area and the other modes need a solver that reads the question, not the options. I haven't built it yet.

A question for you

If you generate maths questions and recompute the keys independently, where does your solver start: from the rendered question text, or from the generator's parameters? Parameters are far easier to work with, but then the solver trusts the same inputs the generator used.

This post was written with AI assistance, and the app and its gates were built with an AI coding agent working to my specifications. Numbers: release commit 5fe5cef, measured 7 October 2026 unless dated otherwise.

Top comments (0)