If you build a classifier with a fixed set of labels, one of those labels usually ends up doing a job nobody assigned it: catching everything that fails to match the others. Face shape detection is a clean, low-stakes example of this, and the numbers that the face shape detector measureface published about its own classifier make the effect easy to see.
Six labels, one of them defined by absence
Consumer face shape tools almost all use the same six categories: oval, round, square, heart, diamond, oblong. Five of those are defined by something being unusual. Square has a hard corner at the jaw. Oblong is longer than it should be. Diamond has a forehead narrower than the cheekbones. Heart comes to a point at the chin. Round is about as wide as it is long.
Oval is defined differently. The usual description is "longer than wide, widest at the cheekbones, tapering gently to a rounded chin with no hard corners." Read each clause and it rules something out instead of pointing at a feature. No corner, so not square. Longer than wide, so not round. Tapering gently, so not heart or diamond. In classifier terms, oval is the residual class.
What the residual class does to the output distribution
The measureface write-up on oval faces includes the result distribution of the classifier it ships, which is unusual enough to be worth reading closely. The classifier measures four lengths plus the jaw angle and compares them to a prototype for each shape. Across 43 distinct faces, 15 came back oval, the most of any label and nearly four times the 4 that came back oblong.
That skew is what you would predict from the geometry. If the oval prototype sits near the average of the other five, then a face that is unremarkable in every direction is closest to oval by construction. Most faces do not have a standout feature, so most ambiguous inputs drain into that bucket.
Two more numbers from the same page make the point sharper:
-
Tie cases. When the two closest prototypes are within a set margin, the detector reports a pair such as
oval/roundinstead of forcing one label. Eight of the 43 faces came back as pairs, and every one of the eight contained oval: four with round, three with heart, one with diamond. A label that borders one neighbour is a category. A label that borders all of them is a gap between categories. - No dominant cause. For every other shape, one measurement does most of the ruling out: forehead width for diamond, the jaw for square, length against width for round and oblong. For oval there is no dominant measurement at all. The forehead and the jaw each account for exactly 16 of the 43, a dead tie. Nothing is deciding oval; it is what remains after everything else has been decided.
Why this matters for anyone shipping a classifier
The practical consequence is that the residual label carries less information than the others. A "square" result tells you a specific measurement crossed a threshold. An "oval" result mostly tells you that no threshold was crossed. If your UI presents all six labels with the same confidence styling, users will read the residual label as a positive finding when it is closer to "no strong signal."
This is not specific to faces. Any taxonomy with an implicit "none of the above" member behaves the same way: support ticket routers with a "general" queue, document classifiers with an "other" bucket, sentiment models with "neutral." The residual class soaks up ambiguous inputs, its precision looks fine because the inputs are genuinely ambiguous, and the label quietly becomes the most common output.
A few design choices reduce the damage:
- Report distances, not only the argmax. Showing how close the input sits to every prototype lets a user see that "oval" won by a hair over "round." measureface does this, and it is the main reason a flip-flopping result becomes explainable instead of confusing.
- Emit ties explicitly. A pair result within a margin is more honest than a coin flip between two nearly equal scores.
- Publish the output distribution. If one label wins far more often than its neighbours, say so. That page states plainly that the finding is "not a flattering thing for us to publish about our own classifier," which is exactly the kind of disclosure that makes the rest of the numbers credible.
- Track which feature drives each decision. A label with no dominant feature is a strong hint that it is a residual.
The limits of this example
The test set is synthetic: every one of those faces was generated by an image model, none belongs to a real person. That makes 15-of-43 a statement about one set of images and one set of thresholds, and it would be a mistake to read it as a population rate. There is also no peer-reviewed prevalence data for these six styling categories, so nobody can tell you how common oval "really" is. The boundaries between shapes are conventions, and a different set of prototypes would move the counts.
That caveat does not weaken the structural point. Any tool built on the same six labels inherits the residual behaviour whether or not it reports it. If you are building one, measure it the way measureface did, show users the distances, and treat your most frequent label with the most suspicion.
Top comments (1)
Oval isn't a shape here, it's the complement of the other five, so of course it absorbs the ambiguous cases. That 'defined by absence' framing is the key insight. The design advice is right, especially reporting distances instead of just the argmax. Have you tried whether a calibrated 'none of the above' threshold on the distance margin reduces the oval skew?