If you've tried to illustrate a long blog post with an image model, you know how it goes. You ask for five pictures and get five different art styles, three versions of your "mascot", and a stock-photo handshake you never asked for. Each image is fine by itself. Put them together in one post, though, and it looks like a ransom note.
I build InkDoo, a tool that takes an article and returns a full set of hand-drawn explainer illustrations, all starring the same character. This post covers the parts that actually made the set look consistent. Spoiler: most of it is boring prompt plumbing, not model magic.
The pipeline in one paragraph
Article in. A planner LLM reads the text and returns JSON: which ideas are worth a picture, which paragraph each picture goes after, and a scene description for each. Then an image model draws every scene from a strict prompt template. Finally, an exporter drops the hosted image URLs back into your Markdown at the right spots. Three stages, and the consistency work happens in all three.
1. Describe the character like a spec, not a vibe
The biggest single win was writing each character as a short, very literal visual spec, then pasting that exact string into every image prompt. Here is the default character, Mochi:
"Mochi", a round low squishy cat-like blob drawn ONLY as a hollow black
outline (white inside, never filled black): a wide soft dumpling-shaped body
sitting flat, two tiny triangle ears, two short sleepy horizontal line eyes,
a tiny 'w' mouth, small nub paws, one thin curly tail.
A few things I learned writing these:
- Count things. "Two tiny triangle ears" and "ONE single large round eye" drift much less than "cute ears" or "big eye".
- Repeat the negative. "Hollow outline, white inside, never filled black" appears in the description and in the style block. Without it the model kept turning line-art characters into solid silhouettes.
- Give the character a personality in three words ("lazy-looking but surprisingly competent, deadpan"). It nudges poses without changing the design.
InkDoo ships six of these (Mochi, Doo, Blot, Stub, Folio, Puff), and they all follow the same spec format.
2. Make the character do the work
A consistent character that just stands in the corner waving is decoration. The planner's system prompt has one rule that changed the output more than anything else:
${N} must PERFORM the core conceptual action of each image (pulling, carrying,
sieving, weighing, stitching, guarding, pushing, folding, unpacking...).
If the image still works without ${N}, ${N} is too decorative — rewrite.
Because the character is part of the metaphor, every image has to draw it at a size and angle where its features are readable. A tiny figure in the background is where likeness falls apart, so this rule helps consistency too.
3. Lock the style separately from the content
Every image prompt has the same fixed "visual DNA" block: pure white background, minimalist black hand-drawn wobbly line art, at least 35% empty space, and sparse handwritten notes in only three colors. Each color has a job: orange for the main flow and arrows, red for the key warning or result, blue for side notes. It also carries a list of things to avoid: gradients, shadows, PPT infographics, cute-poster energy.
Only a few fields change from image to image: theme, structure type, core idea, composition, and two to five short labels. The planner picks the structure from a small, fixed list (Workflow, Before/after, Concept metaphor, Route map, Mini comic…). It also has to invent a fresh, low-tech physical metaphor using one or two objects: a funnel, a scale, a drawer, a broken machine.
Keeping the variable surface small is the real trick. The less each prompt is allowed to vary, the more the set reads as one series.
(Credit where it's due: the illustration method, meaning one idea per image, white background and sparse colored annotations, is adapted from the MIT-licensed Ian Xiaohei Illustrations project. The characters are original.)
4. The planner never sees images, so give it a text twin
Users can upload their own character as a reference picture. That creates an asymmetry. The image call can take the reference (InkDoo sends it to the edits endpoint), but the planner LLM is text-only. So a custom character gets two descriptions:
-
desc, for the image model: "the user's own original character, exactly as shown in the attached reference image… keep its silhouette, proportions, face and distinctive features… but redraw it in this minimalist black hand-drawn line-art style… ignore the reference background." -
planDesc, for the planner: a text-only version built from the name and any look description the user typed.
User text is sanitized (control characters, quotes and braces stripped, length capped) before it gets near a prompt. If the custom character is empty, the resolver quietly falls back to Mochi instead of producing a character-less set.
5. Switching characters without re-planning
Planning is free in InkDoo, and generating images costs credits. So I didn't want people re-planning just because they picked a different character halfway through. The planner writes scenes in English with the character's name in them, so switching is a careful, Unicode-aware whole-word replace of the old name with the new one. Mochi's scene ("Mochi sits on the overflowing inbox") becomes Doo's scene, with no extra LLM call.
6. Put the images back where they belong
The planner returns an after anchor for each shot: the first ~20 characters of the paragraph the image should follow, copied verbatim. The exporter splits the article into blocks, normalizes whitespace, and matches each anchor to a block. Then it writes out Markdown (or WeChat-friendly HTML) with the hosted image URLs in place. Matching on a short prefix holds up much better than asking the model for paragraph numbers, which it miscounts.
Things that still bite
- Long jobs vs. request timeouts. Image calls can take minutes, so generation runs as a background job on a Cloudflare Queue. The client polls, and a failed or stuck job refunds its credit automatically.
- Single-image redraws. Sometimes one image in the set is off. Each card can redraw just that scene with an edited prompt, and the rest stay as they are.
- Labels in the right language. Handwritten notes follow the article language: Chinese articles get Chinese labels, English ones get English.
Takeaways
If you want a consistent character across many AI images:
- Write the character as a literal, countable spec and reuse the exact string.
- Make the character perform the key action so it's always drawn large enough to stay recognizable.
- Freeze the style block and keep the per-image variables few and structured.
- Keep a text twin for any reference image the planner can't see.
If you'd rather not build the plumbing yourself, you can try the whole pipeline at inkdoo.app. Paste a post or drop in a link or a .md/.docx file. New accounts get two free images. I'd love to hear how you're handling character consistency in your own projects.

Top comments (0)