AI disclosure: I used AI tools to help draft and edit this article. I checked every technical claim and command against the current MozgoQuest repository and reran the complete content pipeline before preparing this version. On DEV, this article should be marked AI-Assisted.
I work on MozgoQuest, a free math practice project for children in grades 1 through 6. The current library has 180 original problems in Russian and English.
The risky part of publishing educational content is easy to underestimate. A page can build successfully while the answer is wrong. A translation can read naturally while silently changing a number. Two problems can be almost identical even if their text is not byte-for-byte equal. A migration can insert good content and still expose only part of it.
We treat each problem as a small data release. The source lives in reviewed YAML, and a single command checks the structure, originality, answers, translations, SQL output, and public sitemap before deployment.
This article walks through that pipeline and the mistakes it is designed to catch.
1. Keep authored content separate from runtime output
The authored source is split into four YAML sets. Each document records the language, set name, database ID range, migration filenames, ownership metadata, and the questions themselves.
A shortened question looks like this:
slug: mq26-g4-21
grade: 4
topic: arithmetic
difficulty: 3
statement: "..."
answer_type: number
answer: "42"
explanation: "..."
hints:
- "..."
- "..."
verification:
method: expression
expression: "6 * 7"
YAML is the editable source of truth. SQL migrations and the JavaScript translation bundle are generated outputs. This avoids a common maintenance trap where an editor fixes the source but forgets to update one of several runtime copies.
The build also creates new questions in an inactive state. A separate migration activates the exact ID range only after the insert has been checked. Content deployment becomes a two-phase operation:
- Insert the new rows and hints.
- Verify row counts and inactive status.
- Activate only the intended IDs.
That separation is useful when the database is remote and rollback options are limited.
2. Validate the contract before checking the mathematics
Our structural validator rejects a question when any required field is missing or malformed. It checks:
- stable slug format;
- grade and difficulty ranges;
- a fixed topic vocabulary;
- statement and explanation length;
- exactly two substantial, different hints;
- numeric answer presence;
- original authorship metadata;
- unique slugs and database IDs;
- forbidden competition names inside the problem text.
The final rule protects both editorial clarity and rights boundaries. The site may explain common olympiad formats, but its question bank must not present copied or branded competition material as original work.
Schema validation runs first because later checks assume the fields exist. A precise early error is much easier to fix than a translation or migration failure caused by incomplete input.
3. Detect near duplicates, including old content
Exact duplicate detection is not enough for word problems. These two templates are functionally the same even if a few nouns and numbers change:
A shop had 40 notebooks and sold one quarter...
A library had 60 books and gave away one quarter...
The validator normalizes punctuation and case, then compares every new statement with other authored statements. A similarity ratio of 0.86 or higher fails the build.
It also loads the recovered legacy question bank into an in-memory SQLite database and compares new statements against those older records. That threshold is stricter at 0.70 because the old material is excluded from the current active library and should not quietly return in rewritten form.
Similarity scores are only a guardrail. They catch suspicious pairs for review; they do not prove originality. Ownership metadata and editorial review still matter.
4. Give every answer an executable second opinion
Each problem stores both the expected answer and an independent verification expression. The pipeline evaluates the expression and compares its result with the authored answer using exact numeric tolerance rules.
The evaluator does not pass expressions directly to unrestricted eval. It parses a Python AST and accepts only a small allowlist:
allowed = {
ast.Expression, ast.Constant, ast.BinOp, ast.UnaryOp,
ast.Add, ast.Sub, ast.Mult, ast.Div, ast.FloorDiv,
ast.Mod, ast.Pow, ast.Call, ast.Name, ast.GeneratorExp,
ast.comprehension, ast.Compare, ast.Tuple, ast.BoolOp,
}
Only sum, range, gcd, and lcm are exposed as callable names. Builtins are removed from the evaluation environment.
This is still a deliberately small internal expression language, not a general sandbox for untrusted users. Its job is to make the editor express the calculation twice: once as the answer and once as a reproducible computation. If the two disagree, the build stops.
The strongest version of this idea would also have a second person solve every problem from scratch. We do manual checks for child-facing content, while the expression layer catches mechanical mistakes on every run.
5. Treat translation as structured data
The English bank uses the same slugs as the Russian source. The translation validator requires an exact one-to-one slug set and checks that every version keeps the same:
- grade;
- topic;
- answer;
- two-hint structure;
- numbers used in the statement, explanation, and hints.
The last check has caught the class of error we worry about most: a fluent translation that changes 18 to 80, drops a quantity, or introduces a number that does not exist in the source.
Each translation also needs an explicit review status. A missing or unreviewed translation is not compiled into the public runtime bundle.
Number parity cannot judge language quality, so it complements editorial review instead of replacing it. It is good at detecting numerical drift, which is exactly the kind of error a normal spell checker will miss.
6. Rebuild SQL in a temporary database
The generator writes deterministic SQL migrations from the YAML source. It then applies them to an in-memory SQLite database and verifies:
- the expected number of problem rows;
- exactly two hints per problem;
- the intended ID range;
- inactive status before activation.
Generating SQL without executing it would leave quoting mistakes and constraint errors undiscovered until deployment. The temporary database turns migration generation into an executable test.
7. Validate the public surface too
Passing database checks does not guarantee that users or crawlers can reach the new content. The same pipeline rebuilds both language sitemaps and verifies reciprocal hreflang links.
At the time of this run, the public structure contains:
- 230 Russian sitemap URLs;
- 231 English sitemap URLs;
- 180 task pages per language;
- 16 populated grade-topic hubs per language.
The guest practice route is intentionally excluded from indexing, while library and individual problem pages remain discoverable.
The command we run
All stages are chained behind one command:
npm run content:check
The verified output for this article was:
Validated 4 YAML set(s), 180 original questions;
grades {1: 30, 2: 30, 3: 30, 4: 30, 5: 30, 6: 30};
no duplicate or legacy-like statements
Verified 180 numeric answers by independent expressions
Validated and compiled 180 self-reviewed English translations
Built 180 original questions and 360 hints
Built ru sitemap with 230 URLs and en sitemap with 231 URLs
Validated reciprocal hreflang
The command exits on the first broken invariant. Deployment checks continue with unit tests, browser scenarios, a Worker dry run, and public health checks, but those are outside the content-specific pipeline described here.
What this pipeline does not prove
Automation can prove that two stored numbers match. It cannot prove that a problem is interesting, age-appropriate, clearly worded, or pedagogically useful.
For child-facing content, the remaining review questions are human ones:
- Can a child understand what is being asked without hidden context?
- Does the first hint preserve the chance to solve the problem?
- Does the second hint reveal the method without simply giving the answer?
- Does the explanation teach a reusable idea?
- Does the English version sound natural to its intended reader?
The practical lesson is to automate the invariants that machines can check reliably. That gives reviewers more time for the parts that require judgment.
Top comments (0)