DEV Community

Nikolai | ExamReady
Nikolai | ExamReady

Posted on AI-assisted

I built an exam prep platform for German B1/B2 exams as a solo founder: what I learned

I'm a solo developer, and for a while now I've been building ExamReady, a platform for preparing for official German exams. Most of the people it's meant for are immigrants who need to pass a B1 or B2 exam.

This is a build-in-public summary: what the product is, how it's put together, and the things I got wrong along the way. It's not a success story. It's a list of decisions, and a few of them I'd now make differently.

ExamReady is an independent project. It is not official and is not endorsed by ÖSD, ÖIF or the Goethe-Institut.

The problem I wanted to solve

The exams I cover are the ÖIF Integration Exam B1, ÖSD B1 (ZDÖ), ÖSD B2 and the Goethe-Zertifikat B1. Each has its own format, task types and scoring rules. Preparing means learning German and learning the exam format at the same time.

My idea was simple: keep the exam material in German, because that's what the real exam looks like, but explain everything around it in the learner's own language. The explanation of a task type should not itself be a reading test.

What exists today

There are 4 courses with 104 short interactive lessons, 20 full model tests, and explanations in 7 languages: German, English, Russian, Ukrainian, Turkish, Serbian and Romanian. It's free to start (2 keys a day), and Pro is a one-time payment rather than a subscription.

The stack

I picked it to be boring on purpose:

  • Next.js frontend
  • NestJS backend
  • MongoDB
  • Docker Compose behind Caddy on one small VPS

I'm the only person who deploys and debugs it, so I optimised for being able to hold the whole system in my head. One server, one compose file, one reverse proxy. I'm not claiming this setup is the right one for everyone, only that it fits one person maintaining a product.

A content model that doesn't fight me

The first real design decision was how to store exam content. I ended up with a normalized model: Exam > model test > module > task. Every task has a type, and the answer key lives separately from the task.

type Exam = { id: string; tests: ModelTest[] };
type ModelTest = { id: string; modules: Module[] };
type Module = { id: string; tasks: Task[] };
type Task = { id: string; type: TaskType; content: unknown };

// stored separately from the task
type AnswerKey = { taskId: string; value: unknown };
Enter fullscreen mode Exit fullscreen mode

Keeping the key apart has two benefits. The client never has to receive answers it doesn't need to display, and I can change or regenerate a key without touching the content.

Scoring goes through a registry: one scorer per task type, looked up by the task's type.

const scorers: Record<TaskType, (answer: unknown, key: unknown) => Score> = {
  // multipleChoice: ..., trueFalse: ..., and so on
};

function score(task: Task, answer: unknown, key: AnswerKey): Score {
  return scorers[task.type](answer, key.value);
}
Enter fullscreen mode Exit fullscreen mode

Adding a new task type means adding a type and a scorer. That's the whole contract. It kept me from scattering ad hoc scoring logic across the codebase.

Lessons as decks

Lessons are full-screen decks: one idea per slide, in the style of Brilliant. About 60% of a lesson is teaching, there are 3-4 questions, and a realistic exam task comes at the end. The one-idea-per-slide constraint forces me to cut. If an explanation doesn't fit on one slide, it's usually two ideas, or I haven't understood it well enough yet.

Where I got things wrong

A few honest ones.

The answer keys were guessable. I noticed that answer patterns repeated between tests, and that true/false keys alternated too regularly. A learner could guess the pattern without understanding anything. Now key sequences are generated with constraints (different from other tests, no strict alternation) and a script checks them.

Manual testing wasn't enough. Clicking through lessons yourself mostly covers the path you already know. I now run AI agents acting as personas through every lesson, plus an automated "impatient user" sweep. They found real bugs: a second example that never showed up, a writing draft lost when going back, and links losing the language prefix. I'll write about that setup in a separate post.

Translations needed real review. Machine-like translations had to be reviewed and rewritten. And when I removed em dashes from the text, some Russian and Ukrainian sentences broke, because those languages use the dash in place of the verb "to be".

Where I use AI, and where I don't

Writing and speaking practice use AI. For writing, the LLM (Gemini) gives grades per official criterion and written feedback. My code counts the words and applies the official rules, for example that writing under 50% of the required words scores 0 points. Models are unreliable at counting and arithmetic, so I don't ask them to do it. Corrections are also verified against the learner's own text before they're shown.

For speaking, there is a live AI conversation partner built on Gemini Live. It costs me roughly ten cents per speaking task, which is a real number I have to keep in mind when I think about pricing.

Building with AI coding agents

I built all of this as a solo developer with heavy use of AI coding agents (Claude Code). It let me cover a lot of surface area: content tooling, frontend, backend and scripts. It didn't remove the need to decide what to build, to read the diffs, or to test like a real user. The amount of code I can produce went up. The cost of judgement didn't go down.

Takeaways

  1. Keep the stack boring when you're the only maintainer.
  2. Model your content properly early. Separating answer keys from tasks paid off.
  3. Don't let the LLM do arithmetic or apply rules. Let it judge, let code count.
  4. Test as an impatient, confused user, not as the author.
  5. Treat translations as content that needs review, not as a switch you flip.

If you've built exam prep or other structured-content products, I'd like to hear what you'd do differently.

I'm building this at examready.at, feedback welcome.

Top comments (0)