DEV Community

Pariedolia System
Pariedolia System

Posted on

Your Segmentation Model Isn't Broken, Your Labels Might Be: Measuring Annotation Quality with Python"

You trained a segmentation model on a medical imaging dataset. The Dice score plateaus, validation is noisy, and some cases look inexplicably wrong.

Before you touch the learning rate, ask a different question: would two trained annotators agree with each other on these labels?

If they wouldn't, your model is being asked to match a target that doesn't exist. This post walks through how to measure that with a few lines of Python.

Why label agreement matters

Medical image labels are human judgments. Boundaries blur at organ edges, small lesions sit near visibility limits, and guidelines leave gaps. Every annotator resolves those gaps slightly differently.

That variation puts a ceiling on what any model can learn, and it is invisible unless you measure it.

Inter-annotator agreement (IAA) quantifies how closely independent annotators match on the same cases. Low agreement usually points to an unclear annotation guideline, not careless annotators.

Setup
bash
pip install numpy scipy scikit-learn nibabel

Assume you have binary masks from two annotators, saved as NIfTI files with identical shape and spacing.

python
import numpy as np
import nibabel as nib

def load_mask(path):
img = nib.load(path)
return img.get_fdata() > 0.5, img.header.get_zooms()

mask_a, spacing = load_mask("case001_annotator_A.nii.gz")
mask_b, _ = load_mask("case001_annotator_B.nii.gz")
Metric 1: Dice coefficient

Dice measures overlap, from 0 (none) to 1 (identical).

python
def dice(a, b):
a, b = a.astype(bool), b.astype(bool)
total = a.sum() + b.sum()
if total == 0:
return np.nan # both empty: undefined, handle explicitly
return 2.0 * np.logical_and(a, b).sum() / total

Returning nan for two empty masks, instead of silently returning 1 or 0, keeps your averages honest. Decide separately how "both annotators found nothing" should count in your project.

Metric 2: IoU (Jaccard index)
python
def iou(a, b):
a, b = a.astype(bool), b.astype(bool)
union = np.logical_or(a, b).sum()
if union == 0:
return np.nan
return np.logical_and(a, b).sum() / union

IoU is always less than or equal to Dice for the same pair of masks, so don't compare the two numbers directly. Pick one as your headline metric and stay consistent.

Metric 3: 95th percentile Hausdorff distance

Overlap metrics can look fine while boundaries still differ at specific spots. Surface distance catches that. The 95th percentile version (HD95) ignores the worst 5% of outliers, which makes it more stable than the raw maximum.

python
from scipy.ndimage import distance_transform_edt, binary_erosion

def surface(mask):
return mask & ~binary_erosion(mask)

def hd95(a, b, spacing=(1.0, 1.0, 1.0)):
a, b = a.astype(bool), b.astype(bool)
if not a.any() or not b.any():
return np.nan
sa, sb = surface(a), surface(b)
dist_to_b = distance_transform_edt(~sb, sampling=spacing)
dist_to_a = distance_transform_edt(~sa, sampling=spacing)
distances = np.concatenate([dist_to_b[sa], dist_to_a[sb]])
return np.percentile(distances, 95)

Because sampling=spacing is passed in, the result is in real-world units (typically millimeters), not voxels.

Metric 4: Cohen's kappa for classification labels

If annotators assign categories (for example, benign vs. suspicious per lesion), use kappa, which corrects for agreement expected by chance.

python
from sklearn.metrics import cohen_kappa_score

rater_a = ["benign", "suspicious", "benign", "suspicious", "benign"]
rater_b = ["benign", "suspicious", "suspicious", "suspicious", "benign"]

kappa = cohen_kappa_score(rater_a, rater_b)
print(f"Cohen's kappa: {kappa:.2f}")

For three or more raters, look at Fleiss' kappa (available in statsmodels) or Krippendorff's alpha.

Compare all annotator pairs
python
from itertools import combinations

def pairwise_report(masks, spacing):
# masks: {"A": ndarray, "B": ndarray, "C": ndarray}
rows = []
for x, y in combinations(masks, 2):
rows.append({
"pair": f"{x}-{y}",
"dice": dice(masks[x], masks[y]),
"iou": iou(masks[x], masks[y]),
"hd95_mm": hd95(masks[x], masks[y], spacing),
})
return rows

A pair that stands out from the rest often means one annotator interpreted a rule differently. That is a training or guideline issue worth a conversation.

Find where annotators disagree

A single number per case hides the useful detail. Slice-wise Dice shows where the disagreement happens.

python
def slicewise_dice(a, b, axis=2):
scores = {}
for i in range(a.shape[axis]):
sa = np.take(a, i, axis=axis)
sb = np.take(b, i, axis=axis)
if sa.any() or sb.any():
scores[i] = dice(sa, sb)
return scores

per_slice = slicewise_dice(mask_a, mask_b)
worst = sorted(per_slice.items(), key=lambda kv: kv[1])[:5]
print("Lowest-agreement slices:", worst)

Disagreement usually clusters at the first and last slices of a structure, where the anatomical start and end landmarks matter, and at boundaries with neighboring organs. Those patterns tell you exactly which rules to tighten.

Interpreting the numbers

There is no universal "good" score. Use these only as rough orientation:

Metric Rough reading
Dice above ~0.90 Strong; common for large, well-defined organs
Dice ~0.70 to 0.90 Reasonable; depends on structure size and contrast
Dice below ~0.70 Investigate the guidelines and training
Kappa above 0.80 Almost perfect agreement
Kappa 0.61 to 0.80 Substantial
Kappa 0.41 to 0.60 Moderate

Two important caveats:

Small structures and lesions score lower than large organs, even with excellent annotators, because a one-pixel shift changes overlap a lot. Set targets per label, not one global threshold.
Agreement is not accuracy. Annotators can agree on the same wrong convention. Clinical review is still needed.
Turning measurement into improvement

A simple loop that works well:

Double-annotate a pilot batch (a few dozen cases).
Compute the metrics above, per label and per annotator pair.
Inspect low-agreement slices and cases.
Clarify the guideline, add examples, and log decisions in an edge-case file.
Retrain annotators and re-measure on a fresh sample.
Keep double-annotating a small percentage during production to catch drift.

Treat your annotation guideline like code: version it, review it, and test it.

Practical tips
Check that masks share shape, spacing, and orientation before comparing. A silent resampling mismatch will wreck your numbers.
Use the same preprocessing (window/level, orientation) for every annotator.
Keep per-label results. A good average can hide one poorly defined structure.
Report metrics with sample sizes and variation, not only means.
Never share identifiable patient data in public repos, notebooks, or screenshots.
Final thoughts

Model performance usually gets the attention, but label quality sets the ceiling. A few hundred lines of measurement code can tell you whether your data is ready to train on, or whether your next improvement should come from the guidelines instead of the architecture.

I've put together open SOP, QC checklist, and agreement templates in a GitHub repo, which you're welcome to use and improve: medical-image-annotation-guidelines

I work with the team at Pareidolia Systems LLP, where we do medical image annotation, segmentation, and quality control for healthcare AI teams. If you handle annotator disagreement differently, I'd love to hear how in the comments.
Github-

Top comments (0)