I've been building CrayonInk, a small web app that turns a photo into a black-and-white coloring page you can print at home. Upload a picture of your dog, your grandparents, a wedding photo, and get back a page with clean, closed outlines that's ready for crayons.
The idea isn't new. This post is about the engineering: the pipeline, a model comparison with per-call billing data, the prompt changes that mattered, and the cost guardrails I added before letting strangers spend my inference budget.
The stack
All of it runs on Cloudflare:
-
Next.js 16 deployed to Workers via
@opennextjs/cloudflare - D1 for jobs, accounts, credits and a spend ledger
- R2 for source photos and finished pages
-
Queues for generation jobs (
max_batch_size = 1,max_retries = 3) - A cron trigger every 15 minutes that recovers stuck jobs (and refunds the user's page), deletes source photos older than 24h, and clears out expired results and PDFs
- Cloudflare Email Service for operator alerts
The AI step is a single image-edit call to fal, using fal-ai/bytedance/seedream/v5/lite/edit. One more mode, "outline", skips AI entirely: it's Canny edge detection running inside the Worker. It's instant and costs nothing, but the lines are rougher and often don't close.
How a photo becomes a page
- In the browser: the photo is downscaled to at most 2048px on the long edge, re-encoded as JPEG (quality 0.9), and its EXIF data is stripped. Nothing is uploaded until the user clicks Generate.
- Submit: the API checks the user's credits, creates a job in D1 and drops a message on the queue.
- Queue consumer: builds the prompt from the selected detail level and style and calls Seedream. The photo goes inline as a data URI.
- Post-processing: the model's output gets converted to grayscale, thresholded to a 1-bit PNG at 300 DPI, and padded so the page keeps the photo's aspect ratio. The pages are pure black and white, with no gray anti-aliasing to muddy a home print.
- Delivery: PNG, or a single-page PDF in A4 or US Letter. Several pages can also be assembled into a coloring book PDF with a cover.
Post-processing is what turns "a drawing" into "something that prints cleanly," so every model in the comparison below went through this exact step.
Two bugs my first smoke test caught
Early on I ran 8 jobs through the live API (6 test images, all AI-generated with FLUX schnell, including the two children, so no real kids' photos). All 8 succeeded, nothing was blocked, and AI pages averaged 46.6 s end to end. Two problems mattered more than speed.
1. The model recomposed my photos into squares. Without image_size, Seedream returned a square. It didn't stretch a 4:3 photo; it re-framed it, and on one dog photo it invented a porch above the dog. For a product that promises "your photo," that's a bug.
The fix: compute an output size that matches the input's aspect ratio. This is the version from the commit that fixed it (the current code adds an optional long-edge target for larger paid sizes):
/** Seedream 5 Lite accepts 2560×1440 … 4096×4096 total pixels. */
const SEEDREAM_MIN_PIXELS = 2560 * 1440;
const SEEDREAM_MAX_PIXELS = 4096 * 4096;
const SEEDREAM_TARGET_PIXELS = 2048 * 2048;
const MAX_ASPECT = 3;
export function falImageSize(width: number, height: number): FalImageSize | null {
if (!(width > 0 && height > 0)) return null;
const aspect = Math.min(MAX_ASPECT, Math.max(1 / MAX_ASPECT, width / height));
const round16 = (v: number) => Math.max(16, Math.round(v / 16) * 16);
let w = round16(Math.sqrt(SEEDREAM_TARGET_PIXELS * aspect));
let h = round16(Math.sqrt(SEEDREAM_TARGET_PIXELS / aspect));
while (w * h < SEEDREAM_MIN_PIXELS) { w += 16; h = round16(w / aspect); }
while (w * h > SEEDREAM_MAX_PIXELS) { w -= 16; h = round16(w / aspect); }
return { width: w, height: h };
}
2. sync_mode doesn't do what my code comment said. I'd written that sync_mode "keeps the output out of fal's request history." It doesn't; it only returns the output inline instead of as a CDN URL. Querying the request history, I could pull back the full input payload, base64 photo included, for 5 of 7 requests. That matters when parents upload photos of their kids.
What I do now:
export const FAL_RETENTION_HEADERS: Record<string, string> = {
"X-Fal-Store-IO": "0",
"X-Fal-Object-Lifecycle-Preference": JSON.stringify({ expiration_duration_seconds: 3600 }),
};
About 90 seconds after each request, a queue message checks fal's request history for it. If anything was kept, it tries to delete the payload, and it records the outcome on the job (not_stored / deleted / stored / pending / error). Caveat: deleting payloads needs an admin API key. The privacy page names fal and what it receives, and I deliberately don't promise "we never train on your photos," because that's not my policy to make.
The model bake-off: Seedream vs Nano Banana 2 vs GPT Image 2.5
Then I checked whether I'd picked the right model, using the exact production prompt and parameters and the production post-processing, all via fal.
Setup: 8 photos from Wikimedia Commons (2 adults, 2 multi-person family photos, 2 children, 2 pets) × 2 detail levels (Standard and Detailed) × 4 model configs = 64 calls, run 8 at a time. Total spend: $3.37.
| Seedream 5.0 Lite edit (production) | Nano Banana 2 edit (1K) | GPT Image 2.5 Flare (medium) | GPT Image 2.5 Sunburst (high) | |
|---|---|---|---|---|
| Cost per page* | $0.0350 | $0.0800 | $0.0266 | $0.0689 |
| End-to-end time, median (max) | 61.7 s (127.5 s) | 13.1 s (44.1 s) | 21.8 s (42.4 s) | 41.7 s (56.5 s) |
| Child photos blocked | 0 of 4 | 0 of 4 | 0 of 4 | 0 of 4 |
| Likeness, 1–5 (subjective) | 3.6 | 2.9 | 3.4 | 3.9 |
| Printability, 1–5 (Standard / Detailed) | 4.5 / 4 | 4 / 3.5 | 4 / 3 | 3.5 / 3 |
| Main problem | Slowest; children's eyes often left blank | Least like the person; invents things (added trees to a cat photo) | Cartoonish, pupils filled solid black | Ignores framing, often deletes the background; heavy hatching |
*Cost is fal's per-request x-fal-billable-units response header × unit price. My API key couldn't read the billing endpoints (they need an admin key), so these numbers aren't reconciled against an invoice. The GPT configs came out above fal's listed 1024px prices because auto size produced roughly 2048px outputs and my prompt is long.
A few notes on reading that table honestly:
- "Likeness" isn't a face-recognition metric. It's a 1–5 score from eyeballing side-by-side face crops (every image at Standard, three spot-checked at Detailed). With 8 photos, it's a hint, not a benchmark.
- "Blocked" means the model refused or returned an error. No model refused any of the 4 child-photo calls (or the 4 family-photo calls with kids), and Seedream never returned an NSFW flag. Don't expect the model's safety filter to set your policy on kids' photos; your product has to.
- The printability gap is mostly ink density. After post-processing, all four outputs are pure black and white. The difference is how much black: on GPT's Standard pages heavy strokes and solid black take about 10–14% of the page, versus about 3% for Seedream and about 1% for Nano Banana 2.
- Speeds were measured with 8 requests in flight, so they're comparative, not single-request latency.
What I decided: Seedream stays the default. Among the cheap options it holds the frame best (no cropping, deleting or inventing), and its Standard pages are the cleanest to color. The price is speed: roughly a minute per page, so I don't promise "seconds." GPT Image 2.5 Flare is the challenger: about 24% cheaper and roughly 3× faster, but more cartoonish, so it's my next A/B candidate on a small slice of real traffic. Sunburst had the best likeness but costs about twice as much and drops backgrounds, so it only fits a premium "refine" option. Nano Banana 2 costs 2.3× as much and scored lowest on likeness.
Prompt tuning: what actually changed
The prompt is assembled from parts: a base instruction, a detail-level line, and a block of shared rules. Three changes made a real difference.
1. Forbid recomposing, explicitly. The image_size fix handled the canvas shape, but the model still liked to "improve" the shot. The shared rules went from "Keep the same composition and pose" to:
Trace the photo exactly as framed: keep the same composition, camera angle, subject position, scale and pose, and the same aspect ratio. Do not crop, zoom, extend the scene, add new elements or recompose. Keep faces and key features recognizable and faithful to the photo.
Post-processing had the same bug in miniature: it used to crop to the ink's bounding box, which also changed the aspect ratio. Now it pads each side proportionally instead.
2. Say what a coloring page isn't. Image models love shading, so the shared rules spell out the negatives: "Black lines only: no shading, no gray tones, no color, no hatching, no gradients, no text, no watermark, no frame or border." The thresholding step would turn gray shading into black blobs anyway, so it's better if the model never draws it.
3. A preset for detail-heavy photos. For things like wedding photos, "Bold & detailed" asks for "the same thick, even, confident black stroke, like a thick felt-tip marker," with "facial features … drawn cleanly and recognizably with the eyes left open and not filled in" and "lace, embroidery and beading as simplified open motifs." Post-processing then thins the lines to a centerline and redraws them with a round pen: 9px in open areas, down to 5px in dense ones like hair and lace. It's still one Seedream call per page.
Eyes are still a weak spot: in the bake-off, Seedream left the eyes blank on one child photo at both detail levels. I haven't changed the default preset for that yet; it's next, as its own small experiment.
Abuse prevention and cost control
At $0.035 per page, an anonymous free tier is an invitation to drain my fal balance. Here's what's in place:
- AI pages need Google sign-in. Signed-out visitors can browse examples and use the zero-cost outline tool. Clicking Generate opens a sign-in modal.
- A one-time signup grant, not a monthly allowance. New accounts get a small number of free pages once, and it never resets. The grant is idempotent per user, and the same IP (or IPv6 /64) or the same normalized account name can claim at most 2 grants in 24 hours.
- One active AI job per account, enforced atomically when the job is created, so two tabs can't race each other.
- A USD circuit breaker on the backend. Every AI page is charged to a D1 ledger at a configured cost per page at the moment it's submitted to the provider, and the charge is never reversed, even if the job fails, gets blocked or is deleted. Grant-funded pages have their own daily budget. When it's hit, new grants pause and never-paid accounts can't start AI pages until 00:00 UTC, while paying users carry on. An optional overall breaker can stop everything. Alert emails go out at 80% and 100%.
// Pure admission rule: paid users ignore the grant breaker; the overall breaker blocks everyone.
export function admissionAllowed(state: BreakerState, userHasPaid: boolean): boolean {
if (state.overallTripped) return false;
if (state.grantTripped && !userHasPaid) return false;
return true;
}
Charging on submission and never refunding the ledger is the key choice: the provider bills you either way, so the breaker counts the same way. (Users still get their page credit back on failure; that's a separate counter.)
-
A private-beta switch.
GENERATION_ALLOWLISTis a comma-separated list of Google emails. When it's set, only those accounts can generate, receive the signup grant or open checkout, and everyone else sees a "coming soon" panel. I ran with just my own account for a while, then cleared the list to open sign-ups. - Turnstile is on in production (Managed mode, server-side verification and a 30-minute pass cookie). Most people never see a challenge; scripted requests without a valid token get a 403 before any model call.
Lessons and trade-offs
- Smoke-test the live API before tuning anything. Eight requests found a framing bug and a privacy bug.
- Read your provider's retention docs, not your own comments. The comment was confidently wrong.
- Compare models on post-processed output. Users print the thresholded page, not the raw generation.
- Label where cost numbers come from. Without an admin key I can't reconcile against an invoice, so I say whether a number is list price or billable units.
- Slow is a real trade-off. I'm accepting about a minute per page for better framing and cleaner lines.
- Small samples are small. 8 photos, 64 calls and one person's eyeball scores. Good enough to rule out a 2.3× more expensive model, not good enough to declare a winner between two close ones.
If you want to see what it does with one of your photos, try CrayonInk. Sign in with Google and your new account gets a one-time batch of free pages to test with. The outline tool works without signing in. I'd love feedback on where it fails: hair, glasses, busy backgrounds, whatever breaks it.
Questions about the Workers/OpenNext setup, the queue design or the breaker are welcome in the comments.
Comparison photos: adult portraits by Anthony Ginsbrook and Leroy Skalstad, Labrador by Dktue (all CC0, via Wikimedia Commons); cat by Anil Öztas, CC BY-SA 4.0, via Wikimedia Commons. The derived comparison grid is shared under CC BY-SA 4.0.



Top comments (0)