DEV Community

Yashwanth Gadagani
Yashwanth Gadagani

Posted on

Building an AI STEM Solver That Shows Its Work — and Double-Checks It

Ask any AI chatbot a calculus question and you'll get an answer. Maybe even the right one. But a student doesn't need an answer — they need to understand how to get there, and they need to know the answer is actually correct. Those are two different engineering problems, and I built Forge (repo: Asdfyash1/GPAI) to solve both: a full-stack STEM copilot that turns any problem — typed, photographed, or pasted from a URL — into a step-by-step derivation, then has a second model independently verify the result.

The stack: Next.js 16 (App Router, Turbopack), TypeScript (strict), NVIDIA NIM models (Nemotron, DeepSeek Flash, Llama 3.3) via the Vercel AI SDK, deployed on Vercel's Hobby tier. Here's how the solver works, step by step, with the real architecture.

Step 1: Stream the solution, don't dump it

Nobody reads a wall of math. Forge's solver streams the derivation step-by-step over Server-Sent Events, so steps appear as they're generated and the UI can reveal them one at a time (collapsible step reveal — "show me the next step" instead of the whole answer).

The streaming endpoint lives at /api/educate/stream:

// src/app/api/educate/stream/route.ts (simplified)
import { requireAuth } from "@/lib/api-guard";
import { streamSolverResponse } from "@/lib/orchestrator";

export async function POST(request: Request) {
  const guard = await requireAuth(request);
  if (!guard.ok) return guard.response;

  const body = await request.json();
  const handle = await streamSolverResponse({ problem: body.problem });
  return new Response(handle.textStream, {
    headers: { "Content-Type": "text/plain; charset=utf-8" },
  });
}
Enter fullscreen mode Exit fullscreen mode

On the client, a small useStream hook consumes the SSE stream and appends steps to state as they arrive. The response parser (src/lib/response-parser.ts) decomposes the LLM's markdown into structured steps using regex — each step becomes a collapsible card in SolverView.tsx. That decomposition matters: if the model returns one blob, you can't do progressive reveal, follow-up chips ("Why does this step work?"), or per-step quizzes later.

Step 2: Verify with a second model (cross-check)

This is the part most AI solvers skip. Forge runs a cross-check: after the primary model produces a solution, a second model independently solves the same problem, and the two are compared. The UI shows one of three verdicts: agree, minor difference, or disagree.

Why it matters: a single model's confident wrong answer is the most dangerous output in education. Two models arriving at the same answer independently is dramatically stronger evidence than one model's confidence score. The cross-check model is configurable via env var:

NVIDIA_SOLVER_MODEL=meta/llama-3.3-70b-instruct
NVIDIA_VERIFIER_MODEL=meta/llama-3.3-70b-instruct
Enter fullscreen mode Exit fullscreen mode

You can point the verifier at a completely different model family (or even a different provider via the ADDITIONAL_OPENAI_COMPATIBLE_* vars) so the check isn't just the same weights agreeing with themselves. When the verdict is "disagree," the student sees that upfront instead of copying a wrong derivation into their homework.

Step 3: Render math like math

A derivation full of x^2 + 2x + 1 = 0 in monospace is unreadable. Forge renders formulas with KaTeX via rehype-katex + remark-math, layered on react-markdown + remark-gfm. The MathMarkdown.tsx component handles the whole pipeline, so model output containing $...$ and $$...$$ becomes properly typeset math inline.

Small detail, big UX difference: students trust a solution that looks like their textbook. Plain-text math looks like a chatbot guessing; typeset math looks like a worked example.

Step 4: Accept photos, not just text

Half of real homework exists as a photo of a notebook or a textbook page. Forge's OCR pipeline:

  1. User uploads a photo of handwritten work or a printed problem.
  2. Client-side compression shrinks it to 1600px at quality 0.85 — this keeps the payload under the 1.7 MB Vercel API gateway limit (MAX_INLINE_IMAGE_BYTES). Do this in the browser, not the server; it saves bandwidth and avoids gateway rejections.
  3. NVIDIA Nemotron Omni 30B reads the image and extracts the text.
  4. The extracted text feeds straight into the solver pipeline from Step 1.

PDFs get the same treatment via unpdf — a pure-JS parser with no native binaries, which matters because Vercel serverless functions don't love native modules.

Step 5: Make it work with zero API keys (demo mode)

Here's a trick more AI apps should steal: Forge runs without any API keys at all. If NVIDIA_API_KEY isn't set, the app falls back to demo-solver.ts — deterministic, hand-written demo output that exercises the entire UI (step reveal, KaTeX, quiz, cross-check display) with no model calls.

Why bother? Three reasons: (1) anyone can clone and run it in 30 seconds, which is gold for a portfolio project; (2) frontend development doesn't burn API credits; (3) the demo doubles as a UI test fixture. The env table is honest about what's needed for what:

Variable Needed for
NVIDIA_API_KEY All real AI features
JWT_SECRET Auth (self-generated)
RESEND_API_KEY OTP emails (lazy-initialized — build succeeds without it)
TELEGRAM_BOT_TOKENS Cloud storage backend

Step 6: Secure the endpoints by default

Every AI endpoint sits behind a requireAuth() guard that runs four checks in order: payload size (→ 413), origin validation for CSRF (→ 403), JWT cookie verification (→ 401), and per-user rate limiting (→ 429). The API key lives only in Vercel env vars — the frontend calls your own /api/* routes with an HttpOnly SameSite=Lax cookie and never sees a secret.

The rate limits are tuned per endpoint: 30 req/min for most AI routes, 10/min for debate mode (which fires 3 model calls per request). In-memory sliding-window counters per Vercel isolate, pruned periodically.

Honest gotchas

A few things I'd do differently or that bit me:

  • The OTP store is in-memory. Serverless cold starts wipe pending OTPs. Acceptable for an MVP, not for production — move it to Redis or Postgres before real users.
  • Cloud storage has a ceiling. The Telegram-bot storage backend hits a ~4096-char pinned-message cap around ~50 users. Fine for a demo; a document-based registry is needed to scale.
  • Next.js 16 broke things vs. older training data. Read the framework docs in node_modules, not your memory, when something behaves oddly.
  • One model to avoid: mistralai/mistral-large-3-675b-instruct-2512 isn't in the NVIDIA NIM catalog — every call 404s. The README warns about it because I learned the hard way.

The takeaway

The pattern generalizes beyond homework: stream structured output, verify with an independent model, render it like the domain expects, and degrade gracefully without keys. Most AI wrappers stop at "call the API and print the text." The difference between a demo and a tool is everything around the model call — and that's where the interesting engineering lives.

The repo (Asdfyash1/GPAI) has the full source, and the live app is at gpai-jade.vercel.app — try photographing a math problem and watch the cross-check verdict.

Top comments (0)