DEV Community

Parth Maniar
Parth Maniar

Posted on

Fine-tuning EfficientNetB0 for 104 flower classes: validation vs held-out test scores, honestly

I built a GPU transfer-learning workflow for 104-class flower recognition and tried to report it honestly, including the gap between validation and held-out test scores.

Repo: https://github.com/officialpm/flower-classification

104-class flower recognition: held-out test macro-F1 0.86708

This is an evaluated model workflow, not a deployed flower-identification product.

The approach

Training workflow: load, train, select, export

  1. Verify TFRecord feature names, counts and that all 104 classes appear in training and validation.
  2. Load an ImageNet-pretrained EfficientNetB0 at 224 pixels with flip, rotation and zoom augmentation.
  3. Train a frozen head for three epochs, then fine-tune at a lower learning rate for up to eight. BatchNorm layers stay frozen.
  4. Keep the best validation-loss checkpoints and report macro-F1, accuracy and log loss for every candidate.
  5. Match test IDs exactly to the sample order. Reject duplicates, missing IDs, nonfinite probabilities or labels outside 0-103.

The training set has 12,753 labeled images across 104 class IDs. All classes are present but counts are uneven, so I track macro-F1 rather than accuracy alone.

Training row counts across all 104 flower classes

Results

Model Validation macro-F1 Validation accuracy
Frozen head 0.800261 0.823276
Fine-tuned 0.855443 0.878502
Equal blend 0.855260 0.873653

Validation macro-F1 by model

Validation accuracy by model

The fine-tuned model scored 0.86708 macro-F1 on the held-out test set, scored independently of my training and validation. That is a different measurement from validation (0.855443), and I don't read the gap as a real improvement.

Limits

  • The validation split was reused for checkpointing and model selection.
  • Fine-tuned and blended results are nearly tied. I claim no meaningful margin between them.
  • I did not purge duplicate images or test on an independent unseen domain.
  • ImageNet pretraining may overlap with public flower imagery.
  • This dataset's test labels are easy to find online. I did not fetch or use them, but it means near-perfect scores on it say little. Treat any score here with suspicion, including mine.
  • The portable train.py passed syntax checks but was not retrained end to end. The measured run used TensorFlow 2.20 on two Tesla T4 GPUs.
  • No throughput, latency or accuracy on real user photos has been measured.

Code is Apache-2.0. No dataset, checkpoints or prediction files are published.

Top comments (0)