DEV Community

Cover image for From MediaPipe to Adaptive AI: Building Mendly's Rehab CV Pipeline
Quoc Bao An Nguyen
Quoc Bao An Nguyen

Posted on

From MediaPipe to Adaptive AI: Building Mendly's Rehab CV Pipeline

How our Top 11 HopHacks project turned noisy MediaPipe pose data into reliable rehabilitation metrics for adaptive AI planning

What if a rehabilitation exercise could understand how a patient moved, instead of only knowing whether they finished?

That was one of the ideas behind Mendly, an AI-powered stroke rehabilitation platform my team built during HopHacks.

By the end of the hackathon, Mendly became a Top 11 finalist out of 87 projects and won Best Use of DigitalOcean.

I worked mainly on the computer vision side of the project: pose tracking, joint-angle and range-of-motion measurement, valid-repetition detection, tracking reliability, and the pipeline that turns movement into structured performance data for adaptive AI planning.

Our core flow looked roughly like this:

Camera
  ↓
MediaPipe Pose
  ↓
Joint Angles / Range of Motion
  ↓
Repetition Validation
  ↓
Structured Performance Data
  ↓
Future AI Planning
Enter fullscreen mode Exit fullscreen mode

The interesting part was not simply getting MediaPipe to detect a person.

It was figuring out when we should actually trust what the camera was telling us.


Missing tracking should stay missing

One bug changed how I thought about computer vision.

Suppose the tracker sees:

Frame 1: 40°
Frame 2: 55°
Frame 3: tracking fails
Enter fullscreen mode Exit fullscreen mode

What should Frame 3 become?

At first, two obvious choices are:

  • use 0°
  • reuse the previous 55°

Both are wrong.

If tracking failed, we do not know where the patient's arm actually was.

Using zero invents movement. Reusing the old value assumes the patient stopped moving.

So we changed failed inference into an explicit invalid observation.

Pose succeeds
→ process measurement

Pose fails
→ ignore the observation
→ do not update movement state
→ do not count a repetition
Enter fullscreen mode Exit fullscreen mode

The rule became:

Unknown is not zero. Unknown stays unknown.

That sounds simple, but it prevented missing camera evidence from becoming fake movement data.


Counting reps is more than crossing a threshold

Another challenge was repetition detection.

Real movement data does not look perfectly clean:

155°
148°
139°
121°
97°
76°
81°
95°
119°
142°
151°
Enter fullscreen mode Exit fullscreen mode

If we count a repetition every time an angle crosses a threshold, noise or partial motion can create false reps.

Instead, we treated a repetition as a full movement cycle:

Start
  ↓
Movement begins
  ↓
Required range reached
  ↓
Optional hold
  ↓
Return
  ↓
Valid repetition
Enter fullscreen mode Exit fullscreen mode

This also helped us separate attempts from valid repetitions.

At one point, an exercise could finish after enough failed attempts even if the patient had not reached the required number of valid reps.

We changed that so automatic completion depends only on:

validRepCount >= targetRepCount
Enter fullscreen mode Exit fullscreen mode

If the target is five valid reps, the patient needs five valid reps.

Failed attempts can still be recorded, but they do not magically complete the exercise.


Recorded data is not always trusted data

We ran into the same distinction with range-of-motion measurements.

A patient might complete an attempt while tracking confidence is poor.

That attempt can still be useful as history, but its movement measurement should not automatically become trusted ROM data.

So we separated:

Did an attempt happen?
Enter fullscreen mode Exit fullscreen mode

from:

Do we trust the measurement?
Enter fullscreen mode Exit fullscreen mode

Low-confidence repetitions could remain in the result history while being excluded from trusted movement statistics.

That ended up being one of the most important ideas in the project:

Recorded does not necessarily mean reliable.


The LLM never sees the raw video

One architectural decision I liked about Mendly was keeping computer vision separate from AI reasoning.

Gemini does not inspect the patient's camera feed or count repetitions.

The CV layer converts movement into structured performance information first.

Conceptually:

{
  "exercise": "arm_flexion",
  "attempts": 9,
  "valid_repetitions": 7,
  "range_of_motion": 83,
  "tracking_quality": "good"
}
Enter fullscreen mode Exit fullscreen mode

The computer vision layer answers:

What happened?

The AI planning layer answers:

How should that performance influence a future exercise set?

The patient's current session still follows a practitioner-approved plan.

When a future set is drafted, Mendly uses Backboard and Gemini with practitioner-defined constraints and accumulated performance data.

The proposal then goes through deterministic guardrails and practitioner review before it reaches the patient.

So the model assists with planning, but it does not autonomously decide treatment in real time.


What I took away from the project

Before Mendly, it was easy for me to think about computer vision as:

input → model → prediction
Enter fullscreen mode Exit fullscreen mode

After debugging this system, I started thinking more about:

input
  ↓
prediction
  ↓
confidence
  ↓
validation
  ↓
structured data
  ↓
application behavior
Enter fullscreen mode Exit fullscreen mode

The model is only one part of the system.

The engineering around it determines whether its output is actually useful.

For me, the biggest lesson was:

Good AI is not only about better models. It is also about giving those models reliable information to reason over.

That was the most interesting part of building Mendly - and probably the part I learned the most from.

Top comments (0)