Most "AI photo cleanup" tools have a dirty secret: you upload a 24-megapixel photo, and you get back something closer to 1–2 megapixels. The model works at a fixed resolution, so the whole image gets resized, regenerated and quietly degraded — even the parts you never touched.
I wanted a tool to remove people from photo online without that trade-off, so I built one: People Remover. You brush over a tourist, a photobomber or an ex, the AI fills the area, and you download a PNG that is exactly the same width and height as your original.
This post covers the architecture and the one trick that makes it work.
The core idea: generate small, composite big
Inpainting models are good at inventing plausible background. They are not good at preserving 6000×4000 pixels of detail they didn't need to touch. So I split the job in two:
- The model gets the full frame plus a black-and-white mask (white = edit, black = keep) and generates at 1K (Standard) or 2K (HD Repair).
- The browser keeps the original decoded pixels, scales the model output back up to the original dimensions, and copies in only the masked pixels.
Everything outside the brush stays identical to what the browser decoded from your file. Only the area you painted changes.
Here's a simplified version of the compositing step:
async function composite(
original: ImageBitmap, // full-resolution source
generated: ImageBitmap, // model output at 1K/2K
mask: ImageBitmap // white = replace, black = keep
): Promise<Blob> {
const { width, height } = original;
const canvas = new OffscreenCanvas(width, height);
const ctx = canvas.getContext('2d')!;
// 1. Start from the untouched original.
ctx.drawImage(original, 0, 0);
// 2. Build a layer that is "generated pixels, clipped to the mask".
const layer = new OffscreenCanvas(width, height);
const lctx = layer.getContext('2d')!;
lctx.drawImage(generated, 0, 0, width, height); // upscale to original size
lctx.globalCompositeOperation = 'destination-in';
lctx.drawImage(mask, 0, 0, width, height); // keep only masked area
// 3. Paint that layer over the original.
ctx.drawImage(layer, 0, 0);
return canvas.convertToBlob({ type: 'image/png' });
}
The real version handles a few more things: the mask needs an alpha channel (not just luminance) for destination-in, antialiased edges give a narrow transition so you don't see a hard seam, and repeated edits build on the last full-size composite rather than the original.
Why HD Repair is not "AI upscaling"
A common question: if the output is full size anyway, what does the 2K mode buy you? The answer is detail inside the brushed area. At 1K, a person you removed from a 6000px-wide photo is regenerated at roughly a sixth of the native resolution, then stretched. At 2K you get twice the detail in each direction. On small photos (under ~1000px on the long side) Standard already matches the source, so HD makes no difference. It's never a whole-image upscale — untouched pixels are never regenerated.
The stack
-
Next.js 15 deployed to Cloudflare Workers via
@opennextjs/cloudflare - D1 (SQLite at the edge) for accounts, jobs and a credit ledger
- R2 for uploads, masks and results, accessed through the S3 SDK
- Stripe for one-time credit packs and monthly plans
- Google OAuth as the only sign-in
No image processing happens on the server. The Worker validates uploads, creates a job with the model provider, polls it, and copies the result into R2. The heavy pixel work runs in the user's browser.
Things that bit me
Idempotency matters more than you think. Mobile networks drop requests all the time. Every removal carries a client-generated UUID requestId. If the same ID arrives twice, the server returns the existing job instead of charging again or creating a second model task.
Refunds need to be transactional. A generation can fail at the provider after the credits are deducted. Confirmed failures refund the exact charge once, recorded in the ledger. Temporary polling or download errors, on the other hand, stay retryable and don't refund — otherwise a flaky connection turns into free credits.
Save the final composite separately. The browser uploads the merged, original-size PNG back to the server after compositing. If that upload fails, the browser keeps the composite in memory and offers a save-only retry that never charges another edit.
Canvas has memory limits, especially on phones. A 25-megapixel photo decoded into RGBA is 100 MB of raw pixels, and iOS Safari will happily refuse a canvas that big. I cap input at 25 MP / 20 MB and show an explicit error instead of silently producing a broken file.
Metadata is gone. Output is PNG, and EXIF/ICC data isn't kept. Some people see that as a privacy bonus; photographers who care about color profiles should know it up front.
What you can use it for
The same brush-and-fill pipeline works for more than people:
- Removing strangers from travel shots — I wrote a step-by-step guide on how to remove people from photos with real before/after masks, including why you should always brush the shadow too.
- Cleaning up clutter: bins, cables, signs, cars. That's the dedicated tool to remove object from photo.
- Date stamps, watermark-free captions on your own images, screenshot text: remove text from image.
If you mostly edit on your phone, I also compared the built-in options (iPhone Clean Up, Google Photos Magic Eraser, Samsung Galaxy AI) in this roundup of free remove people from photo apps. The short version: the built-in tools are great if your phone supports them; a browser tool helps when it doesn't, or when you need the full-resolution file.
Try it / tell me what breaks
People Remover gives you 2 free credits without signing in (one Standard edit), and 4 more on your first Google sign-in. No watermark, no app install, full-resolution download.
I'd genuinely like feedback from other devs, especially on:
- Mask edge blending — any better approaches than an antialiased brush + alpha clip?
- Handling very large images on mobile without hitting canvas limits (tiling?
createImageBitmapwith resize options?) - Anyone running similar async job flows on Workers + D1 — how are you handling polling after the user closes the tab?
Drop a comment. Happy to go deeper on any part of this.
Top comments (0)