For our final year project, we tackled one of the most tedious bottlenecks in medical AI: annotation. Here's how Meta's Segment Anything Model 2 helped us label a dental X-ray dataset in batches instead of one painful image at a time.
The problem nobody warns you about
Everyone talks about training models. Nobody warns you that before you train anything, you'll spend weeks drawing polygons around teeth.
For our final year project, we set out to build a data annotation platform for dental X-rays — panoramic radiographs where a dentist needs structures like cavities and wisdom teeth precisely segmented. Manual annotation is slow, expensive, and inconsistent: two annotators will draw two different boundaries around the same cavity. Doing this by hand, image after image, was not feasible in a single semester.
We needed leverage. That leverage turned out to be SAM 2.
Why SAM 2 fits annotation perfectly
The Segment Anything Model 2 (SAM 2), released by Meta, is a promptable segmentation model: give it a point, a box, or a rough mask, and it returns a precise segmentation mask. Its killer feature for our use case is propagation — SAM 2 was designed to track objects across video frames, which means it can carry a segmentation from one image to the next.
That gave us our core idea: annotate once, propagate to the batch.
Instead of drawing every mask by hand, the annotator:
- Selects a target class — e.g. cavity or wisdom tooth
- Provides a single prompt (a click or bounding box) on one representative X-ray
- Lets SAM 2 segment that structure and propagate the segmentation across the batch
- Reviews the results and corrects only the failures
The human shifts from drawing to reviewing — a fundamentally faster job.
System design
Our platform has three layers:
Frontend — the annotation workspace. A web-based UI where the annotator uploads a batch of X-rays, picks the target class from a panel (cavity, wisdom tooth, and others), and places a prompt on one image. Results render as editable overlays — every mask can be accepted, tweaked, or rejected before it counts.
Inference backend — SAM 2. The prompt plus the image goes to a Python backend serving SAM 2, which returns the mask. For batch mode, we propagate the segmentation across the remaining images in the batch, exploiting the anatomical consistency of dental radiographs: teeth sit in roughly the same regions across panoramic X-rays.
Export layer. Accepted annotations are exported in CSV format (alongside the mask data), so the labeled dataset can feed directly into a training pipeline — no conversion step, no lock-in to our tool.
Dental X-rays are a hard test case
This wasn't an easy dataset to show off on. Panoramic dental X-rays have:
- Low contrast between cavities and healthy enamel in early-stage decay
- Overlapping structures — teeth crowd each other, roots overlap the jaw
- High anatomical variation — missing teeth, fillings, implants, braces all break naive assumptions
What we found: SAM 2's promptable design handles this better than fully automatic segmentation, because the human prompt disambiguates which structure matters. A click on a cavity tells the model exactly what to segment; the model handles the precise boundary. The review step catches the genuinely ambiguous cases — which are also the cases where two human annotators would disagree anyway.
What changed in practice
The honest summary of the project isn't a benchmark table — it's a workflow change. Annotating a batch of X-rays went from an hours-long drawing exercise to a review task: prompt once, let SAM 2 do the repetitive segmentation, and spend human effort only where judgment is actually needed. The review-and-correct loop kept final quality at human-expert level while throughput multiplied, because the scarce resource — a trained eye knowing what a cavity looks like on an X-ray — was no longer spent tracing pixels.
Lessons learned
- Promptable > automatic for expert domains. In medical imaging, a fully automatic segmenter that silently mislabels a cavity is dangerous. Keeping a human prompt + review in the loop isn't a compromise — it's the correct design.
- Batch propagation exploits dataset structure. Medical image batches are far more consistent than natural-image collections. Designing around that consistency is free performance.
- Export format matters as much as the UI. An annotation tool that can't hand clean, portable data (in our case, CSV) to a training script is a demo, not a tool. We built the export path as a first-class feature, not an afterthought.
- SAM 2 is a force multiplier, not a replacement. The annotator's domain knowledge is still the scarce resource. SAM 2 just lets one expert's judgment cover far more data.
What's next
Ideas we didn't get to in a semester: active learning that routes only the lowest-confidence propagations to human review, extending beyond panoramic X-rays to periapical and bitewing views, and packaging the whole thing as a reusable annotation service for other medical imaging datasets.
Built as a final year project at UMT Lahore by Waqas Ahmad. If you're working on medical image annotation or promptable segmentation, I'd love to compare notes — drop a comment.
Top comments (1)
tr.ee/dev-to