DEV Community

Tal Aizikov
Tal Aizikov

Posted on AI-assisted

Gemini picks the stretch IDs. TypeScript fills the minutes.

Large language models are great at choosing what belongs in a workout. They are unreliable at how long that workout should last.

When we shipped AI sessions for MotionLab — personalized stretching routines built around how your body feels today — we deliberately split those two jobs. Gemini picks stretch IDs from a fixed catalogue. Our Next.js route validates those IDs and fills or trims the list until the session lands near the duration the user asked for. The model never invents an exercise, and it never owns the clock.

This post walks through the real pipeline in our MotionLab-Web repo: the catalogue we send the model, the JSON contract, the local duration loop, and how that fits next to the deterministic generator we already used for non-AI routines. Snippets are trimmed from production TypeScript.

Stack, for context: Next.js 16.2, React 19, Supabase, Firebase, Stripe, and Zustand. The AI route calls Gemini 2.5 Flash.

Why not ask the model for a full routine?

An early temptation with any fitness AI is: send the user prompt, get back names, durations, cues, and ordering. That fails in three boring ways.

  1. Hallucinated exercises. A model will happily invent a stretch that is not in your library, misspell an ID, or mix two moves into one name.
  2. Duration drift. "About 15 minutes" in natural language does not map cleanly to a sum of per-exercise seconds, especially when some moves are side-specific and count twice.
  3. Content you already wrote. We already store instructions, breathing cues, equipment, and target areas per exercise. Regenerating that text per request would be slower, less consistent, and harder to review.

So the model is a selector, not an author. It returns a short JSON object: a session name, a session type, and an array of IDs that must already exist in ALL_EXERCISES.

The catalogue the model actually sees

Every exercise in the library is compacted into one line before it goes into the system prompt: id|targetAreas|stretchType. No long descriptions. No video URLs. Just enough for the model to match "tight hips after desk work" to the right IDs.

const CATALOGUE = ALL_EXERCISES
  .map((e) => `${e.id}|${e.targetAreas.join(",")}|${e.stretchType}`)
  .join("\n");

const SYSTEM = `You are a mobility coach. Given user input and the exercise list, select 4-8 exercises.
Reply ONLY with valid JSON (no markdown, no code fences):
{"name":"<short session name>","type":"DAILY_MAINTENANCE|PRE_WORKOUT|POST_WORKOUT","ids":["id1","id2",...]}
Rules: no duplicate ids, ids must be from the list, prioritize areas user mentions, avoid redundant overlaps.

Exercises (id|targetAreas|stretchType):
${CATALOGUE}`;
Enter fullscreen mode Exit fullscreen mode

That keeps the prompt dense. The catalogue currently has on the order of seventy-five stretches and growing; sending full copy for each would burn tokens for little gain. The client only needs to send a free-text prompt and an optional durationMinutes (we default to 15).

Calling Gemini with thinking turned off

The route is a standard App Router POST handler. We require a non-empty prompt and a configured GEMINI_API_KEY, then hit Gemini 2.5 Flash with a low temperature and a hard cap on output tokens. For this classification-style job we also set thinkingBudget: 0 so the model does not spend tokens on an internal chain-of-thought we will throw away.

const geminiRes = await fetch(
  `https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent?key=${key}`,
  {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({
      contents: [{ role: "user", parts: [{ text: `${SYSTEM}\n\nUser: ${prompt}` }] }],
      generationConfig: {
        temperature: 0.3,
        maxOutputTokens: 512,
        thinkingConfig: {
          thinkingBudget: 0, // disable thinking — not needed for simple classification
        },
      },
    }),
  }
);
Enter fullscreen mode Exit fullscreen mode

After the response comes back we strip accidental markdown fences, JSON.parse the text, and map IDs through a real lookup table. Anything not in ALL_EXERCISES is dropped. If nothing valid remains, we return 422 instead of improvising.

const exerciseMap = Object.fromEntries(ALL_EXERCISES.map((e) => [e.id, e]));
let exercises = (parsed.ids ?? [])
  .filter((id) => exerciseMap[id])
  .map((id) => exerciseMap[id]);

if (!exercises.length) {
  return NextResponse.json({ error: "No valid exercises returned" }, { status: 422 });
}
Enter fullscreen mode Exit fullscreen mode

That single filter is the safety rail. The model can only point at rows we already ship.

Filling and trimming to the user's clock

Gemini is asked for 4–8 IDs. That set is rarely the exact length of a 7-, 15-, or 25-minute session. Side-specific stretches also double their durationSeconds because each side gets its own hold. So duration is computed locally:

const targetSeconds = durationMinutes * 60;

const calcDur = (e: (typeof ALL_EXERCISES)[0]) =>
  e.isSideSpecific ? e.durationSeconds * 2 : e.durationSeconds;

let filledSeconds = exercises.reduce((s, e) => s + calcDur(e), 0);

if (filledSeconds < targetSeconds - 60) {
  const aiIds = new Set(exercises.map((e) => e.id));
  const tightAreas = new Set(exercises.flatMap((e) => e.targetAreas));
  const supplementPool = shuffle(
    ALL_EXERCISES.filter(
      (e) => !aiIds.has(e.id) && e.targetAreas.some((a) => tightAreas.has(a))
    )
  );
  for (const ex of supplementPool) {
    if (filledSeconds >= targetSeconds - 60) break;
    exercises.push(ex);
    filledSeconds += calcDur(ex);
  }
}
Enter fullscreen mode Exit fullscreen mode

If the AI list is short, we supplement from exercises that share the same target areas the model already chose. That keeps the session thematically coherent without a second model call. Then we trim so the accumulated length does not overrun the target by more than a minute:

const finalExercises: typeof exercises = [];
let accSeconds = 0;
for (const ex of exercises) {
  const dur = calcDur(ex);
  if (accSeconds + dur <= targetSeconds + 60) {
    finalExercises.push(ex);
    accSeconds += dur;
  }
}
Enter fullscreen mode Exit fullscreen mode

The ±60 second buffer matches how people experience a timed session: one stretch longer or shorter than the label is fine; three minutes over is not. Session type is validated against a small allowlist (DAILY_MAINTENANCE, PRE_WORKOUT, POST_WORKOUT) with a safe default if the model returns something else.

The response is a normal Routine object: ID, name, duration, type, full exercise records, and a short description that quotes the user's prompt. The client already knows how to render guided timers from that shape.

The deterministic twin: non-AI generation

AI sessions are not the only path. Most daily routines still go through generateRoutine in src/lib/routine/generator.ts — no LLM involved. That function takes a fixed duration of 7, 15, or 25 minutes, a routine type, optional sport keys, optional tight areas, and optional exclusions. It mirrors the weighting logic we use on Android.

Sport keys expand into priority target areas (running → hamstrings, calves, ankles, hip flexors; swimming → shoulders, upper back, neck, chest; and so on). Tight areas are merged first, then sport areas. The pool is filtered to matching exercises when possible. Sport-related exercises are double-weighted in a candidate list, the list is shuffled, and we walk it until the second budget is filled — again with no repeats and the same side-specific duration math.

export function generateRoutine(options: {
  durationMinutes: 7 | 15 | 25;
  type: RoutineType;
  sport?: string | null;
  tightAreas?: TargetArea[];
  excludeIds?: string[];
  focusId?: string;
}): Routine {
  // ... build focusAreas from tightAreas + sport maps ...
  // weight sport-specific exercises 2x, shuffle, fill without repeats
}
Enter fullscreen mode Exit fullscreen mode

For 15- and 25-minute sessions we sometimes extend a non-sided stretch to 120 seconds so longer blocks do not need twice as many distinct moves. Daily focus IDs (desk job, wake-up, lower body, and so on) can further restrict the pool when the catalogue has enough tagged exercises.

The important design point: AI and non-AI paths share one exercise catalogue and one duration model. Gemini only replaces the "which IDs first" step. Everything else — validation, timing, UI — stays identical. That made shipping AI sessions a route and a prompt, not a second product surface.

Practical lessons

Keep the model inside a closed set. If your product has a curated library, send IDs (or compact rows) and reject anything else. Do not ask the model to invent catalogue entries.

Own duration in application code. Summing typed fields you control is cheaper and more predictable than asking an LLM to hit a minute target.

Reuse the domain object. Returning the same Routine shape as the deterministic generator meant the player, history, and analytics did not need an AI-specific branch.

Turn off features you do not need. For ID selection, low temperature and thinkingBudget: 0 cut latency and cost without hurting quality.

Supplement thematically, do not random-pad. When we need more minutes, we pull from the same target areas the model already emphasized. Padding with unrelated stretches would feel like a broken personalization promise.

Where this sits in the product

Users can still pick a classic session type and duration and get a weighted routine with no model call. When they want to describe a situation in plain language — post–leg day, pre-run warm-up, shoulders locked up from swimming — AI sessions take that sentence, pick IDs, and hand the same timer experience. We wrote more about the product framing on the MotionLab blog; the implementation above is what sits behind the generate button.

If you are wiring Gemini (or any LLM) into a domain with a real inventory — recipes, workouts, lesson plans, spare parts — the same split applies: let the model choose from your SKUs, and let your code enforce quantities, prices, and clocks.

This article was drafted with AI assistance and reviewed by the MotionLab team.

Top comments (0)