I built a GPU transfer-learning workflow for 104-class flower recognition and tried to report it honestly, including the gap between validation and held-out test scores.
Repo: https://github.com/officialpm/flower-classification
This is an evaluated model workflow, not a deployed flower-identification product.
The approach
- Verify TFRecord feature names, counts and that all 104 classes appear in training and validation.
- Load an ImageNet-pretrained EfficientNetB0 at 224 pixels with flip, rotation and zoom augmentation.
- Train a frozen head for three epochs, then fine-tune at a lower learning rate for up to eight. BatchNorm layers stay frozen.
- Keep the best validation-loss checkpoints and report macro-F1, accuracy and log loss for every candidate.
- Match test IDs exactly to the sample order. Reject duplicates, missing IDs, nonfinite probabilities or labels outside 0-103.
The training set has 12,753 labeled images across 104 class IDs. All classes are present but counts are uneven, so I track macro-F1 rather than accuracy alone.
Results
| Model | Validation macro-F1 | Validation accuracy |
|---|---|---|
| Frozen head | 0.800261 | 0.823276 |
| Fine-tuned | 0.855443 | 0.878502 |
| Equal blend | 0.855260 | 0.873653 |
The fine-tuned model scored 0.86708 macro-F1 on the held-out test set, scored independently of my training and validation. That is a different measurement from validation (0.855443), and I don't read the gap as a real improvement.
Limits
- The validation split was reused for checkpointing and model selection.
- Fine-tuned and blended results are nearly tied. I claim no meaningful margin between them.
- I did not purge duplicate images or test on an independent unseen domain.
- ImageNet pretraining may overlap with public flower imagery.
- This dataset's test labels are easy to find online. I did not fetch or use them, but it means near-perfect scores on it say little. Treat any score here with suspicion, including mine.
- The portable
train.pypassed syntax checks but was not retrained end to end. The measured run used TensorFlow 2.20 on two Tesla T4 GPUs. - No throughput, latency or accuracy on real user photos has been measured.
Code is Apache-2.0. No dataset, checkpoints or prediction files are published.
Top comments (0)