Training an on-device YOLO26n food detection model from scratch across multi-source datasets
I set out to build a food recognition system that could look at a photo, tell you what each item is, count how many there are, flag anything that looks spoiled or rotten, and pick out the things that are not food at all. The twist is that it had to do all of that on a phone or an iPad with no internet connection, which meant there was no cloud to lean on and no API to call when the model got confused.
You would think the hard part of a project like this is the model, but for me it was the data.
I had about 149k images collected from different sources, and once I set aside one large dataset that had no bounding boxes, I was left with 49,158 images from 11 datasets to work with. Every one of them had its own way of naming things, its own idea of what a good photo looks like, and plenty of the same pictures turning up again and again. So before I trained anything, my first real job was deciding what to throw away.
I wrote this as I went, following the pipeline in the order I ran it in Google Colab, so you will see the problems I ran into and how I fixed them. Some of the mistakes were mine, and I have left them in on purpose, because that is where most of what I learned came from.
If you are about to train a model on data from more than one place, I hope this saves you a few of the wrong turns I took.
How fast the data shrank
The part that surprised me most was how much smaller the dataset got at every single stage, so I have laid it out in a table before we get into the details.
| Stage | Images |
|---|---|
| Original collection | about 149k |
| After removing a large dataset with no bounding boxes | 49,158 |
| After blur, brightness and duplicate cleaning | 31,589 |
| Clean images that actually had usable labels | 24,306 |
| Final split | 20,660 train and 3,646 validation |
That means I did not train on 49k images at all, I trained on about 20k, and every one of those drops had a reason that I will walk through as we go.
Checking the GPU first
This job needs a GPU, because a full run means a lot of epochs over a big dataset after several rounds of cleaning, and on a CPU that would take days where a Colab T4 gets it done in hours, so the first thing I do in any session is check what I have been given.
import torch
print('GPU available:', torch.cuda.is_available())
if torch.cuda.is_available():
print('GPU name:', torch.cuda.get_device_name(0))
print('VRAM:', round(torch.cuda.get_device_properties(0).total_memory / 1e9, 1), 'GB')
Mine came back as a Tesla T4 with 15.6 GB of VRAM, and if yours says no GPU is available you will need to either pay for one or find a free platform that offers it.
Configuring the run
A handful of variables control the whole run, and I like setting them all in one place at the top so I never have to hunt for them later.
RUN_MODE = 'train' # 'baseline_only' or 'train'
SAMPLE_SIZE = None # None uses everything, 5000 gives a quick test run
BLUR_THRESHOLD = 50
HASH_DISTANCE = 5
I planned the work in three stages, starting with a baseline run to see what the off-the-shelf model knows, then a small 5k training run to prove the whole pipeline works, and finally the full run once the pilot looked sensible. RUN_MODE decides whether the work stops at the baseline or training goes ahead, and the later cells simply check that one variable with an if/else.
The pilot saved me a lot of pain, because the bugs I found on the small run would have cost me hours if I had only discovered them on the big one.
Getting the data into Colab
I compressed the annotated data into a zip and stored it on Google Drive, and since reading straight from Drive is slow, I copy the zip onto Colab's local disk and extract it there, which is where two things went wrong for me.
The first was a partial extraction, where after a disconnect I ended up with a half-unzipped folder that looked perfectly fine but was missing files, and I only caught it because I counted the images and compared the total with the zip, so now I delete the folder, unzip again and check that the count is 49,158 before moving on.
The second was an extra nesting level, because the zip extracted to data/data/datasets and a hardcoded path simply fails, so I use a glob to find the datasets folder wherever it happens to land.
matches = glob.glob(os.path.join(EXTRACT_DIR, '**/datasets'), recursive=True)
datasets_dir = matches[0]
Installing and importing
!pip install -q ultralytics imagehash tqdm pyyaml
import os, glob, subprocess, shutil, random, json
from collections import Counter
import cv2
import imagehash
import numpy as np
import torch
import yaml
from PIL import Image
from tqdm import tqdm
from ultralytics import YOLO
from google.colab import drive
Ultralytics handles training, inference and evaluation, imagehash finds near duplicates, tqdm gives you progress bars, and PyYAML reads and writes the dataset config files. Ultralytics runs on PyTorch, which is why torch is imported for the GPU check, while OpenCV and PIL do the image work.
Discovering the classes
Next I read every data.yaml and printed the classes from each dataset, which looks like a boring step but is the one that shows you just how messy your data really is.
What I found is probably what you will find too, which is that the same thing gets named in many different ways. One dataset called something tomato, another vine-tomato and a third beef-tomato, apples came as granny-smith, pink-lady and royal-gala, and bell peppers showed up by colour and sometimes as bell-pepper and sometimes as bell_pepper. On top of that there was a long tail of overly specific labels for packaged products with brand names and sizes baked in, and if you leave all of that alone, every variant becomes its own class and the total balloons far past the number of distinct things you actually care about.
The bigger problem was the numbering, because in one dataset egg was class 6 and in another it was class 5, and if you train on both without fixing that, the model is told that one object is two different things and that two different things are the same.
Then there was one more thing that caught me out, because when I globbed for label files I also picked up the README files, and a "label" I opened turned out to contain the text of an export note instead of box coordinates, so I started filtering out anything with README in the name and checked my counts by tallying annotations per class in a dataset I knew well.
Building the unified class list and remapping labels
This is the step that decides whether your model ends up good or bad, so it deserves the most care, and I broke it into three parts.
First I merge the aliases, using a dictionary that maps every variant to one canonical name. I only merged things that are really the same object at the level of detail the task needs, and I kept spoiled and fresh versions as separate classes.
MERGE_ALIASES = {
'granny-smith': 'apple', 'pink-lady': 'apple', 'royal-gala': 'apple',
'vine-tomato': 'tomato', 'beef-tomato': 'tomato',
'bell-pepper': 'bell_pepper', 'capsicum': 'bell_pepper',
# ...many more
}
How hard you merge depends on what your model is for, so if you need to tell two varieties apart, keep them separate.
Then I build the global list by resolving every name through the aliases and sorting the result alphabetically, so the order stays stable between runs.
GLOBAL_CLASSES = sorted(all_class_names)
That gave me 94 classes, numbered 0 through 93.
Finally I rewrite every label file, because the first number on each line of a YOLO label file is the class ID and I needed every class to have the same ID in every dataset. In total 48,763 label files were remapped, and only 139 annotations were dropped because they mapped to nothing useful.
Cleaning the images
Blurry, badly lit and duplicated images teach a model bad habits and waste GPU time, so I filtered on four things, starting with blur.
Blur
The blur check converts the image to grayscale and runs a Laplacian filter, which responds to edges, so sharp images have strong edges and a high variance, while blurry ones have a low variance and get rejected if they fall below the threshold.
def is_blurry(img):
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
return cv2.Laplacian(gray, cv2.CV_64F).var() < BLUR_THRESHOLD
The usual starting threshold is 100, but my pilot run is the reason I changed it, because at 100 the filter threw away 14,732 images, about 30% of everything, and only 44.3% of the data survived cleaning. Food photos are often close-ups with soft textures and shallow depth of field, which score low even when they are perfectly good, so I dropped the threshold to 50, which cut the blurry rejections to 6,326 and lifted overall retention to 64.3%.
Brightness
def is_bad_brightness(img):
m = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY).mean()
return m < 30 or m > 225
An average below 30 means the picture is too dark to see anything and above 225 means it is washed out, and that check removed 929 images.
Duplicates
The same images appear across several scraped datasets, so duplicates were a big deal, and a perceptual hash is a good way to catch them because it fingerprints what an image looks like and near-identical images get near-identical hashes, so if a new image is within a hash distance of 5 of anything already seen, I skip it.
h = imagehash.phash(Image.fromarray(cv2.cvtColor(img, cv2.COLOR_BGR2RGB)))
if any(h - sh <= HASH_DISTANCE for sh in seen_hashes):
dupes += 1
continue
seen_hashes.append(h)
This dropped 10,314 images, and it is also the slow part, since every new image is compared against every hash kept so far, which means the loop gets slower as it goes and the whole cleaning pass took about 42 minutes for 49,158 images.
Letterboxing
YOLO expects square input and squashing a photo into a square distorts the shapes of the objects, so letterboxing scales the image to fit inside 640 by 640 and pads the rest with grey (pixel value 114), and the label coordinates have to be recalculated to match, which is very easy to forget.
scale = TARGET_SIZE / max(h, w)
nw, nh = int(w * scale), int(h * scale)
canvas = np.full((TARGET_SIZE, TARGET_SIZE, 3), 114, dtype=np.uint8)
I letterboxed the images to 640 by 640 myself and saved them to disk, even though Ultralytics can do it on the fly with imgsz=640, and the reason was speed. Many of my source images were large camera captures, and decoding big JPEGs again on every epoch was keeping the Colab CPUs busy while the T4 waited for data, so resizing once up front made the files smaller and the extraction quicker, and it left me with the exact offset and scaling logic I needed later to prepare camera frames for the ONNX model on the device.
Where the labels went missing
This one surprised me, because of the 31,589 clean images only 24,306 had a usable label file, and in the pilot it was much worse, with just 7,758 of 21,780. Some images never had annotations in their original dataset and some lost all of theirs during class remapping, but either way an image without labels is no use for detection training, so they were dropped.
Splitting into train and validation
I split the data 85 to 15, which gave 20,660 training images and 3,646 validation images, and I kept a test set aside for later, for when I want to check the model on data it has truly never touched.
random.seed(42)
if SAMPLE_SIZE is None or SAMPLE_SIZE >= len(labeled):
sample = labeled
else:
random.shuffle(labeled)
sample = labeled[:SAMPLE_SIZE]
split = int(len(sample) * 0.85)
train_imgs = sample[:split]
val_imgs = sample[split:]
If slicing is new to you, sample[:split] means everything from the start up to the split point, so the first 85 percent, and sample[split:] is the rest.
There is something in this code that I only noticed while writing this post, which is that the seed only matters if something random happens after it. In the 5k pilot the shuffle ran, so the seed made the sample reproducible, but in the full run I used every labelled image, so the code took the first branch, the shuffle never ran, and the split was simply the first 85 percent of a sorted file list against the last 15 percent. If your images are saved in dataset order, your validation set may come mostly from the last dataset or two instead of being a fair mix, so I would shuffle before splitting, ideally stratified by source dataset, and compare class counts in train against validation before trusting any score.
The YAML file then tells YOLO where everything is.
path: /content/data
train: train/images
val: val/images
nc: 94
names: [apple, asparagus, avocado, ...]
Saving a checkpoint to Drive
Once the clean split was ready, I zipped it up and saved it to Drive as a checkpoint.
subprocess.run(['zip', '-q', '-r', CLEAN_ZIP_OUT, 'train', 'val', 'dataset.yaml'],
cwd='/content/data')
I also made a Fast Resume cell that mounts Drive, unzips this file and restores the variables, so after a disconnect I can skip everything above and go straight to evaluation or training, and if you have ever lost a Colab session in the middle of a long job, you will know exactly why that cell exists.
The baseline
With the data ready, I ran a baseline, which is simply the off-the-shelf model evaluated on my validation set.
baseline_model = YOLO('yolo26n.pt')
baseline_results = baseline_model.val(data=YAML_PATH, imgsz=640, batch=16, device=0)
YOLO26n is the nano variant, with 122 layers, about 2.4 million parameters and 5.4 GFLOPs, and it comes pretrained on COCO, which has 80 general classes like person, car and laptop.
On my pilot validation set the baseline scored almost nothing, with mAP@50 around 0.0004, and I should be honest about what that tells you, which is not that the pretrained model is bad at your objects. COCO's class IDs and my class IDs do not line up, so the model gets marked wrong for predicting labels from a completely different list, and what the baseline really confirms is that the two label systems do not match. A fairer baseline would map the handful of overlapping classes, like apple, banana, orange, broccoli and carrot, and score only those, and since I did not do that, I treat my baseline as a sanity check and not as a competitor.
The training run
Then came the training run itself.
results = train_model.train(
data=YAML_PATH,
epochs=epochs,
imgsz=640,
batch=16,
workers=2,
patience=patience,
device=0,
pretrained=True,
optimizer='AdamW',
lr0=0.001,
cos_lr=True,
augment=True,
mixup=0.1,
copy_paste=0.1,
)
I scaled the schedule to the data, so the 5k pilot used 20 epochs with a patience of 10 and anything larger used 100 epochs with a patience of 25, and here is what each of the main settings is doing.
epochs. One epoch is one full pass through the training set, so 100 epochs means the model gets to see every training image 100 times.
batch=16. Instead of looking at one image at a time, the model sees 16 together and averages the learning signal, which is steadier and makes better use of the GPU, and if you hit a CUDA out of memory error you can drop it to 8.
optimizer='AdamW'. The optimizer decides how the weights change after each batch, and AdamW adapts per weight, making bigger moves for weights that have barely changed and smaller moves for the ones that are already well tuned.
lr0=0.001. The learning rate sets the size of each adjustment, and if it is too high the model overshoots and never settles, while if it is too low it learns painfully slowly.
cos_lr=True. Rather than staying fixed, the learning rate follows a cosine curve from 0.001 down towards zero, so the early epochs make big moves and the later ones make small refinements.
augment, mixup and copy_paste. Ultralytics flips, crops, recolours and distorts images as it trains, mixup blends two images together, and copy-paste lifts objects out of one image and drops them into another, so the model is pushed to learn features that hold up instead of memorising specific photos.
patience. If validation does not improve for that many epochs in a row, training stops on its own.
pretrained=True. The model starts from the COCO weights, so it already knows edges, textures and shapes, and training teaches it to apply that knowledge to your own classes.
What the runs showed
The 5k pilot took about half an hour and ended with mAP@50 of roughly 0.25, and it was still improving when it stopped, which told me 20 epochs was too few. The full run on about 20k images took around eight minutes per epoch on the T4, climbed quickly at first and then slowed down, and after 35 epochs it had reached about two thirds on mAP@50, while a full 100 epochs would take well over twelve hours, which is more than a Colab session likes to give you, so checkpointing to Drive really matters.
I would not read too much into the gap between those two numbers, because the runs differ in validation set, number of classes, blur threshold and amount of data all at once, so I cannot say how much of the jump came from more data and how much from the cleaner setup, and the pilot was only ever built to prove the pipeline, which it did.
The per-class results from the pilot were uneven, with orange, meat, mango, apple and peach doing well at roughly 0.75 to 0.82 on mAP@50, while pizza, hamburger and cooking spray did badly, all below 0.1, and several rare classes had only one to three validation images, which is too few to score honestly.
It is also worth being clear about what this system can and cannot do, because it flags what it sees, so it can say an item looks spoiled or that something is not food, but it cannot read a use-by date or tell you an item is safe, and with recall in the region of two thirds it will miss some objects, which means counts will tend to run low, so I treat it as a helper and never as a final check.
Making it work without internet
Running on a phone or an iPad is the reason I went with the nano variant, since YOLO26n has few enough parameters to run on mobile hardware, and exporting at 640 by 640 gives the model a fixed and predictable input shape.
Once training is done, I export the best weights to ONNX, a portable format that runs locally with no server and no connection.
model = YOLO('runs/best.pt')
model.export(format='onnx', imgsz=640, dynamic=False, simplify=True)
Setting dynamic=False fixes the input size and simplify=True trims the graph, and both help on small devices, so after the export the app only needs the ONNX file and a runtime, and all of the inference happens on the phone or iPad itself. The app also has to prepare each camera frame the same way I prepared the training images, with the same letterboxing to 640 by 640 and the same grey padding, otherwise the boxes will not line up with what the camera sees.
What I would do differently
Looking back, there are a few things I would change. I would shuffle before splitting and split by source dataset so that every dataset shows up in both sets, I would build a fairer baseline using the classes COCO and my labels share, and I would replace the duplicate check with something faster than comparing against every earlier hash. I would also hold out a proper test set, because right now the validation set is doing two jobs, and I would change one thing at a time between runs so I can tell what actually helped.
The biggest lesson for me is that the cleaning and the class unification took far longer than the training and changed the result far more, so if you are about to train on data from several places, that is where your time is best spent.
Top comments (2)
Very good read ngl, this made me want to try this out
I'm glad to know