DEV Community

Ishan Naik
Ishan Naik

Posted on Originally published at phasesapp.vercel.app

On-Device Computer Vision in React Native: Auto-Aligning Progress Photos with MediaPipe and Expo

A person can take thirty photos of the same face and still produce a timelapse that looks like camera shake. On Monday they stand closer to the mirror. On Tuesday they tilt their phone. By Friday they frame their eyes near the top of the image.

For fitness, skincare, and haircut timelines, viewers need to see the change in the person without chasing their face around the screen. In Phases, the useful visual reference is the eye line: keep both eyes at the same output coordinates across the sequence.

This article develops the React Native and Expo architecture for that workflow: detect landmarks on the device, compute a two-point transform, retain originals in local storage, and render a reel without a rendering server.

Implementation scope: the public Phases site currently describes a browser app that users can add to their home screen. It also describes consent-based analytics. The native Expo architecture and zero-telemetry policy below describe the reference design for this article, rather than a verified description of that deployed browser build. The project link is github.com/Ishannaik/phases; the implementation examples below stand on their own.

Separate the photograph from its alignment

A capture guide helps the user repeat their pose, but it cannot guarantee matching camera distance or roll. I would keep the original photo and store its alignment metadata beside it. That lets the user change the reel's framing without reshooting or accumulating JPEG compression from repeated edits.

flowchart TD
    A[Camera capture] --> B[Normalize orientation and mirroring]
    B --> C[On-device facial landmark detection]
    C --> D{Usable eye anchors?}
    D -->|Yes| E[Compute affine matrix]
    D -->|No| F[Ask user to retake or mark anchors]
    B --> G[Original photo in app filesystem]
    E --> H[SQLite metadata and transform]
    G --> I[Local Canvas or WebGL composite]
    H --> I
    I --> J[On-device video encoder]
    J --> K[Timelapse reel on device]
    G -. Optional user-enabled backup .-> L[Cloud backup]

The engineer should treat capture, detection, and rendering as separate steps. After a successful detection, the renderer needs the photo URI and six transform coefficients. It does not need another inference pass each time the user plays the timeline.

For a native app identified as io.phases.app, use its app-private filesystem for photos and SQLite for timestamps, dimensions, eye coordinates, and alignment settings. io.phases.app is an application identifier, not a portable filesystem path. Android and iOS choose the sandbox's physical location.

With Expo, expo-file-system provides document storage, and expo-sqlite provides a persistent database. Store large image files in the filesystem rather than putting base64 photo strings into database rows.

Choose eye anchors your detector can supply

With MediaPipe, you can derive eye anchors from a face mesh. The older Face Mesh API offers iris refinement through refine_landmarks; the current Face Landmarker task provides a denser landmark result. Google documents 478 landmarks for its current model bundle.

MediaPipe canonical face mesh with landmark indices and eye-region topology

Google's canonical face model UV visualization. This repository asset illustrates landmark topology, not a Phases screenshot or a detector result from this article.

ML Kit can provide eye landmarks through its face detection API, but an eye landmark is not equivalent to a measured pupil center. If you choose ML Kit, name those anchors eye centers. If you need iris-center estimates, choose a detector configuration that supplies iris landmarks and document your index mapping against that model version.

Even an iris center gives an estimate of the visible iris, not a clinical measurement of pupil position. For framing a selfie sequence, the distinction matters because eye movement can introduce jitter. A user looking left and then right can change the anchors without moving their head. A consistent gaze and capture guide help. Eye-contour centers offer another anchor convention when gaze changes dominate the result.

I would run detection on each captured still image, with one intended face. A low-resolution preview can guide capture, but I would calculate the saved transform from the final image. Preview crops and final camera output often use different coordinate systems.

Before measuring the eye vector:

  1. Decode the photograph into a consistent upright orientation.
  2. Apply the chosen mirror policy to both pixels and landmarks.
  3. Convert detector coordinates into source-image pixels.
  4. Preserve consistent eye identities across photographs.

For a detector that returns normalized coordinates, multiply x by image width and y by image height. Measuring an angle in normalized coordinates distorts it whenever width and height differ.

Two anchors determine rotation, scale, and translation

Let the source eye anchors be:

$$
p_1=(u_1,v_1),\qquad p_2=(u_2,v_2).
$$

Choose fixed normalized output anchors:

$$
q_1=(x_1,y_1),\qquad q_2=(x_2,y_2).
$$

For horizontal eyes, choose $y_1=y_2$. In an output with width $W$ and height $H$, convert those targets to pixels:

$$
Q_1=(Wx_1,Hy_1),\qquad Q_2=(Wx_2,Hy_2).
$$

For example, on a 1080 by 1920 portrait canvas, (0.38, 0.35) and (0.62, 0.35) place the eyes at (410.4, 672) and (669.6, 672). Their separation is 259.2 pixels. These are illustrative framing choices, not measured Phases defaults.

SOURCE PHOTO                         OUTPUT CANVAS

                 p2 *                Q1 *=============* Q2
                   /                     fixed eye axis
                  / d
             p1 *

       tilted eye axis               horizontal eye axis
       variable separation           fixed separation

             translate + rotate + uniform scale
Enter fullscreen mode Exit fullscreen mode

Define the source and target eye vectors:

$$
d=p_2-p_1,\qquad D=Q_2-Q_1.
$$

Then calculate:

$$
s=\frac{\lVert D\rVert}{\lVert d\rVert},\qquad
\theta=\operatorname{atan2}(D_y,D_x)-\operatorname{atan2}(d_y,d_x).
$$

Use the source midpoint $m=(p_1+p_2)/2$ and target midpoint $M=(Q_1+Q_2)/2$. For any source pixel $p$:

$$
T(p)=sR(\theta)(p-m)+M,
$$

where:

$$
R(\theta)=
\begin{bmatrix}
\cos\theta & -\sin\theta\
\sin\theta & \cos\theta
\end{bmatrix}.
$$

You can write this as a homogeneous affine matrix:

$$
\begin{bmatrix}
x'\y'\1

\end{bmatrix}

\begin{bmatrix}
a&c&e\b&d&f\0&0&1
\end{bmatrix}
\begin{bmatrix}
x\y\1
\end{bmatrix}.
$$

Two point pairs determine this similarity transform, which uses uniform scale, rotation, and translation. They do not determine an arbitrary six-parameter affine transform with shear and independent axis scales. A similarity transform belongs to the affine family and preserves the face's proportions.

For a horizontal target eye line, $\theta$ cancels the source eye angle. The coordinate system uses positive y downward, as Canvas does; calculate both angles in that same system.

Calculate the Canvas coefficients in TypeScript

The following reference implementation accepts source anchors in pixels and target anchors in normalized output coordinates. It returns the six values in the order that Canvas expects for setTransform(a, b, c, d, e, f).

type Point = Readonly<{ x: number; y: number }>;
type Affine = [number, number, number, number, number, number];

function eyeAlignment(
  source1: Point,
  source2: Point,
  target1: Point,
  target2: Point,
  width: number,
  height: number,
): Affine {
  const values = [
    source1.x, source1.y, source2.x, source2.y,
    target1.x, target1.y, target2.x, target2.y,
    width, height,
  ];
  if (!values.every(Number.isFinite) || width <= 0 || height <= 0) {
    throw new Error("Alignment needs finite coordinates and positive dimensions");
  }
  if ([target1.x, target1.y, target2.x, target2.y]
      .some(value => value < 0 || value > 1)) {
    throw new Error("Target anchors must lie within normalized output bounds");
  }

  const sx = source2.x - source1.x;
  const sy = source2.y - source1.y;
  const tx = (target2.x - target1.x) * width;
  const ty = (target2.y - target1.y) * height;
  const sourceDistance = Math.hypot(sx, sy);
  const targetDistance = Math.hypot(tx, ty);
  if (sourceDistance < 1 || targetDistance < 1) {
    throw new Error("Eye anchors must be distinct");
  }

  const scale = targetDistance / sourceDistance;
  const angle = Math.atan2(ty, tx) - Math.atan2(sy, sx);
  const a = scale * Math.cos(angle);
  const b = scale * Math.sin(angle);
  const c = -b;
  const d = a;
  const mx = (source1.x + source2.x) / 2;
  const my = (source1.y + source2.y) / 2;
  const qx = (target1.x + target2.x) * width / 2;
  const qy = (target1.y + target2.y) * height / 2;
  const e = qx - a * mx - c * my;
  const f = qy - b * mx - d * my;
  return [a, b, c, d, e, f];
}

function drawAlignedFrame(
  ctx: CanvasRenderingContext2D,
  image: CanvasImageSource,
  imageWidth: number,
  imageHeight: number,
  matrix: Affine,
): void {
  ctx.save();
  try {
    ctx.setTransform(1, 0, 0, 1, 0, 0);
    ctx.clearRect(0, 0, ctx.canvas.width, ctx.canvas.height);
    ctx.fillStyle = "#000";
    ctx.fillRect(0, 0, ctx.canvas.width, ctx.canvas.height);
    ctx.setTransform(...matrix);
    ctx.drawImage(image, 0, 0, imageWidth, imageHeight);
  } finally {
    ctx.restore();
  }
}
Enter fullscreen mode Exit fullscreen mode

Use imageWidth and imageHeight from the same oriented image coordinate system as the source anchors. The code assumes the caller owns the canvas, with no inherited clipping region.

The one-pixel check prevents division by a degenerate eye vector. Production acceptance needs a stronger, image-relative criterion: a two-pixel eye separation on a large image can still produce an unusable scale. Check confidence, face size, and acceptable scale before saving alignment metadata.

To check the matrix mathematically, apply it to each source anchor. You should recover the corresponding target pixel coordinate:

a*u1 + c*v1 + e = W*x1
b*u1 + d*v1 + f = H*y1

a*u2 + c*v2 + e = W*x2
b*u2 + d*v2 + f = H*y2
Enter fullscreen mode Exit fullscreen mode

These equations provide useful unit-test assertions with a floating-point tolerance. They describe invariants, not test results from a running Phases build.

React Native needs a rendering boundary

React Native does not expose a browser DOM Canvas as a built-in drawing surface. The CanvasRenderingContext2D example belongs in a browser or a WebView. In a native renderer, adapt the same coefficients to that renderer's matrix convention.

You can use a WebView-hosted Canvas for a shared web rendering path, or use a native GPU drawing surface. With WebGL, upload a photo texture and apply the transform while drawing a quad. Account for the renderer's origin, texture orientation, and matrix layout; copying Canvas coefficients into an unrelated matrix API can invert or transpose the result.

For MediaPipe or ML Kit integrations that require native code outside Expo Go, create an Expo development build. Installing a JavaScript wrapper cannot add its native implementation to the Expo Go binary. Expo documents this boundary in its development build guidance.

Keep decoded image data near the detector and renderer. Passing a full-resolution base64 image through repeated JavaScript bridge messages adds allocation and copying. Transfer a URI or native image handle where the integration supports it.

Render frames locally, then encode locally

Drawing aligned photos on a canvas gives you frames. To export a video file, you also need an encoder and a container writer.

In a native implementation, connect the rendering surface to an on-device encoder. In a browser implementation, select a supported local recording or encoding API and handle device-specific codec availability. Do not assume a Canvas screenshot sequence equals an MP4 export.

For an export with one photograph per time interval, calculate presentation timestamps from the intended playback cadence. Processing speed should not determine the video's timing. A slow decode must delay export work, not stretch one day into a longer frame.

Memory makes sequential processing important. One 1080 by 1920 RGBA frame takes 8,294,400 bytes, about 7.9 MiB before renderer and decoder overhead. Holding 365 such frames would consume about 2.8 GiB in pixel buffers alone. Decode a bounded number of photos, render a frame, submit it to the encoder, and release resources as the encoder permits.

You should also choose how to handle uncovered canvas areas. Rotation can expose triangles at the corners; scaling cannot recover pixels outside the original capture. Use a consistent background, crop, or capture margin. Avoid changing the crop per frame, since that reintroduces apparent zoom.

Local storage and an optional backup path

A useful local record links a photo URI to its capture time, oriented dimensions, source eye anchors, detector version, and alignment settings version. Retain the originals so you can recalculate transforms after changing the target eye positions.

Coordinate filesystem writes with database commits. Write the image into durable document storage before committing a row that references it. If the database commit fails, remove the new file or reconcile it on the next launch. SQLite cannot roll back a separate filesystem write.

The reference design uses a zero-telemetry policy: no analytics events, session replay, or automatic diagnostic uploads. Developers must apply that policy to crash-reporting SDKs and dependencies as well as application code. Local inference does not by itself guarantee zero network traffic.

For offline detection on first use, package the model with the app. A model download on first launch introduces a network prerequisite even if subsequent inference stays local.

Keep cloud backup outside the alignment and rendering path. A user who declines backup should still capture, align, preview, and export. If they enable backup, explain which originals and metadata you upload, request the provider's permissions, and make local deletion versus remote deletion explicit. App-private storage also does not imply encryption or exclusion from operating-system backups; configure those policies for the threat model you support.

The public Phases FAQ describes optional Google Drive backup and analytics that remain off until consent. Readers evaluating the deployed app should use that published policy rather than the stricter native reference policy above.

Eye alignment has a defined limit

A two-eye transform corrects translation, camera roll, and apparent face scale. It cannot correct a person turning their head, pitching their chin, changing expression, or photographing themselves under a different light source. On a fitness timeline, the eyes also cannot guarantee matching shoulder position or body pose.

I would keep those constraints visible in capture UX: guide the user toward the same gaze, pose, and lighting, then use alignment to absorb small framing errors. For a failed detection or an occluded eye, offer a retake or manual anchors. Do not invent a plausible transform and label it a successful alignment.

Phases' core visual goal is concrete: put the eye axis in the same place across days. With consistent image coordinates, a four-degree-of-freedom transform, and a local rendering pipeline, you can achieve that framing without sending a user's progress photos to an inference or rendering service.

References

Written by Ishan Naik.

Top comments (0)