DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Six percent of the box, with an eight point floor: padding a crop the user drew with a finger

When Munchable does not have a product, you can add it by photographing the label. That flow is the one part of the app where the user's hand does the hardest part of the work: they point a camera at a curved, glossy, badly lit ingredients panel and we have to read the text off it. You can see what the app does with the result at munchable.app/conditions, and try the app itself at app.munchable.app.

We went through several designs for that screen. The one that survived is the least clever: shoot the whole panel, freeze the frame exactly where the live preview was, and let the user draw a box around the bit that matters with their finger.

This post is about the geometry of that box, which is a small pile of pure functions and three constants that each took a while to earn.

Why a drawn box at all

Two reasons, neither about accuracy.

The first is that a whole photo of a pack contains several things that look like an ingredients list. Nutrition tables, storage instructions, a recycling panel, and on multilingual packs a second and third ingredients list in other languages. Reading the whole frame means deciding which of those is the one, and getting that wrong produces a product entry that is wrong in a way nobody can see.

The second is that the user knows the answer. They are looking at the pack. Asking them to circle the paragraph is less work than any interface we could build to disambiguate afterwards.

Nothing is pre-drawn. There is a corner guide on the live camera, and it is only a hint about where to hold the pack, not a crop region. The user draws the box, or does not, and taking the whole photo is a perfectly valid path.

The mapping, and the one fact that makes it easy

A preview is a "cover" fit of the sensor frame: scaled uniformly until it fills the view, then centred, with the overflow off the edges. That is true on both platforms (expo-camera uses resizeAspectFill on iOS and CameraX FILL_CENTER on Android, with preview and still capture sharing an aspect ratio), which means a point on screen maps into the photo through one uniform scale plus a centring offset. No per-axis stretch, no letterbox case.

// Cover fit: the photo is scaled uniformly until it fills the view, then centred.
const scale = Math.max(view.width / photo.width, view.height / photo.height);
const offsetX = (view.width - photo.width * scale) / 2;
const offsetY = (view.height - photo.height * scale) / 2;
Enter fullscreen mode Exit fullscreen mode

And the piece of interaction design that makes the whole thing honest: the frozen still is displayed with the same cover fit as the live preview, in the same card, at the same size. So at the moment the shutter fires, nothing moves. The picture simply stops. A box drawn on the frozen frame therefore maps into photo pixels through the same function as a box drawn on a live one.

Everything in the module is pure, which is the other reason this part of the app is pleasant to work on:

 * Everything here is pure so it can be reasoned about
 * without a device.
Enter fullscreen mode Exit fullscreen mode

Camera code is miserable to test. Geometry is not, as long as you keep the two apart. The functions take sizes and rects and return rects, and the only thing a device is needed for is whether the camera actually reports the sizes you think it does.

Two small hygiene details in the mapping. Origins get Math.floor and ends get Math.ceil, so rounding always grows the crop rather than shaving a row of pixels off a line of text. And a mapped area that collapses below 8 pixels on either axis returns null instead of a sliver, so the caller can say "draw a bigger box" rather than sending something unreadable off to be read.

Normalising the drag

People do not drag top-left to bottom-right. They drag from whichever corner their thumb started on.

export function rectFromDrag(startX, startY, endX, endY, view: Size): Rect {
  const ax = clamp(startX, 0, view.width);
  const ay = clamp(startY, 0, view.height);
  const bx = clamp(endX, 0, view.width);
  const by = clamp(endY, 0, view.height);
  return {
    x: Math.min(ax, bx),
    y: Math.min(ay, by),
    width: Math.abs(bx - ax),
    height: Math.abs(by - ay),
  };
}
Enter fullscreen mode Exit fullscreen mode

Clamp both points to the view first, then normalise. Clamping first matters because a drag that leaves the card, which happens constantly when the box you want ends at the edge of the screen, should end at the edge rather than producing a rect with negative coordinates that the clamp downstream then has to rescue.

There is also a minimum:

/** Smallest drag that counts as a deliberate selection rather than a tap. */
export const MIN_SELECTION: Size = { width: 44, height: 28 };
Enter fullscreen mode Exit fullscreen mode

Without it, every tap on the still is a 2 by 3 point selection, and the screen has to decide what an unreadable crop means. With it, a smudge is reported as a smudge, the box is cleared, and the user tries again with no error message anywhere.

The constant I changed three times

Now the interesting one. The drawn box needs padding before it becomes a crop, because a drag ends where the finger lifts, which is not where the user thought they were pointing.

The first version was a fixed 12 points on each side. That is the right answer for the other place this geometry is used, a Lens style selection over a restaurant menu, where the boxes are all roughly the same size.

It is the wrong answer for labels, because label boxes differ by an order of magnitude:

 * Percentage rather than a fixed number of points because the boxes people
 * draw here differ by an order of magnitude: a two line "Ingredients: oats,
 * water" on a porridge pot against the full wrap of a shampoo sized smoothie
 * bottle. A fixed 12pt pad is a third of the height of the small box and
 * invisible on the big one.
Enter fullscreen mode Exit fullscreen mode

So the pad became a fraction of the box's own size:

export const CROP_PAD_FRACTION = 0.06;
Enter fullscreen mode Exit fullscreen mode

Six percent, for a reason I can actually defend: a deliberate drag ends a couple of millimetres inside where the user meant on a thumb-sized target, and six percent of a box somebody drew on purpose is reliably more than that drift, while staying small enough that it never swallows the neighbouring panel and hands the reader a second ingredients list to merge into the first.

And then the floor, which is what happens when you write a percentage rule and forget the small end of the range:

/**
 * Floor under the percentage pad, in view points. A box only just over
 * MIN_SELECTION would otherwise get 44 * 0.06 = 2.6pt of margin, which is less
 * than the width of the box's own border and cuts the ascenders off the top
 * line of text.
 */
export const MIN_CROP_PAD = 8;
Enter fullscreen mode Exit fullscreen mode

2.6 points of margin is narrower than the border we draw around the box. The user sees a box whose edge sits comfortably outside their text and gets a crop with the tops of the letters shaved off. Eight points is about one line of cap height at the scale a label is drawn on screen.

The pad is applied in view points, before the mapping into photo pixels, and that ordering is deliberate:

 * Worked out in view points, before the mapping into photo pixels, so what the
 * user gets is the margin they would have drawn with their finger rather than
 * a margin that grows with the megapixels.
Enter fullscreen mode Exit fullscreen mode

Pad in photo pixels instead and the same rule gives a 12 megapixel phone twice the margin of an older one, for a user gesture that was identical. Anything derived from a human hand belongs in the units the hand was working in.

Then it clamps to the view, and the comment records the right attitude to that:

 * Clamping to the view is what stops a box drawn hard against the edge from
 * asking for pixels off the screen; the pad is a courtesy, not a promise, and
 * where there is nothing to spare the crop simply stops at the edge.
Enter fullscreen mode Exit fullscreen mode

The path with no drag at all

One last function, and it is the one I would most easily have forgotten:

/** The crop that keeps the whole photo: the accessible and no-drag path. */
export function wholePhotoRect(photo: Size): CropRect | null
Enter fullscreen mode Exit fullscreen mode

A drag is a gesture, and a gesture is an accessibility problem. Somebody using a screen reader cannot draw a box on a photo they cannot see. The still's accessibility label says what the screen is for, and the parent offers the whole photo instead, which goes down the same pipeline and reads the same way. It is also the fallback for anyone who simply taps the shutter and does not want to draw anything.

The composed version of all of this is four lines:

export function stillRectToPhotoRect(rect, view, photo, padFraction = CROP_PAD_FRACTION): CropRect | null {
  if (view.width <= 0 || view.height <= 0) return null;
  return viewRectToPhotoRect(padRect(rect, view, padFraction), view, photo, 0);
}
Enter fullscreen mode Exit fullscreen mode

Pad in the user's units, then map into the camera's. Everything above is the argument for why those two steps are in that order, and for why the three numbers in them are the numbers they are.

One constraint worth naming, since it is a product rule rather than a geometric one: this is framing, never editing. The box decides which pixels get read, and nothing in the flow lets anyone touch the text that comes back. That is why the padding matters more than it would in a photo editor: if the crop cuts a word, the fix is to draw the box again, not to correct the word afterwards.

Top comments (0)