How we approached recognition accuracy when the "easy" conditions are the exception, not the rule
Most container OCR demos look great. Clean daylight, dry containers, camera angle perfectly perpendicular to the door. That's also almost never what a real port gate looks like.
In production, our system has to read container numbers off boxes that are rusted, mud-splattered, backlit by a truck's headlights at 2am, or half-obscured by rain streaking across the lens. This post is about the engineering decisions — and the mistakes — that came out of trying to make recognition reliable under those conditions, not just accurate in a benchmark.
Why "high accuracy" numbers are often misleading
A lot of OCR vendors quote a single accuracy number, usually measured on a curated dataset shot in good conditions. That number tells you almost nothing about gate performance, because the failure modes at a real terminal cluster hard around a small set of scenarios:
Heavy rain — water droplets on the lens housing distort character edges; wet container paint also increases specular reflection
Strong backlight — trucks arriving at dawn/dusk, or headlights at night, blow out the exposure on one side of the frame
Low light / dark nights — most terminals don't have gate-level lighting tuned for camera exposure, they have lighting tuned for human eyes
Dirty or corroded containers — mud, rust bleed, and faded paint reduce character-to-background contrast, sometimes to the point a human squints too
If your training data doesn't proportionally represent these conditions, your reported accuracy is measuring the wrong thing. The first real fix wasn't a model change — it was admitting our evaluation set was too clean.
Rebuilding the dataset around failure conditions, not average conditions
We restructured our data collection and augmentation pipeline around condition buckets rather than raw volume. Instead of "collect more images," the question became "collect more images in the specific conditions we're currently failing on."
Concretely, that meant:
Tagging every training and eval image with a condition label (rain / backlight / low-light / soiled / clean) instead of treating the dataset as one undifferentiated pool
Tracking accuracy per bucket, not just in aggregate — a model can look great overall while quietly failing 30% of rainy-night captures
Synthetic augmentation targeted at the weakest buckets: simulated rain streaks and lens droplets, exposure/backlight simulation, and paint-degradation overlays (rust bleed, mud spatter patterns) generated procedurally rather than hand-collected, since real-world dirty-container photos are slow to gather at scale
This sounds obvious in hindsight, but it changed our priorities: we stopped chasing marginal gains on the easy 70% of captures and started spending compute budget where the model was actually failing.
Architecture: splitting localization from recognition
Early on we ran a single end-to-end model doing detection + recognition in one pass. It worked fine in good conditions and degraded unpredictably in bad ones — when the model got the character localization even slightly wrong under glare or mud occlusion, the recognition stage had no way to recover.
We moved to a two-stage pipeline:
1.Localization stage — finds the character region on the container regardless of legibility, trained to be robust to occlusion and lighting rather than optimized for recognition accuracy
2.Recognition stage — takes the localized crop and reads the ISO 6346 code, with condition-aware preprocessing (adaptive contrast/exposure correction) applied per-crop rather than globally on the full frame
The key benefit: when recognition confidence is low, we can tell whether the problem is localization (wrong region) or legibility (right region, hard read) and route those failures differently — a wrong localization is a hard failure, but a low-confidence legibility case is a good candidate for multi-frame fusion (below) rather than an immediate reject.
Multi-frame fusion instead of single-shot recognition
Single-frame recognition has a hard ceiling in bad conditions — one frame with a rain streak across a character just doesn't have the information a clean frame does. Since gate cameras capture a short burst as the truck passes rather than a single still, we fuse predictions across frames instead of picking "the best" one.
At a high level:
for each frame in burst:
run localization + recognition
record per-character confidence scores
for each character position:
take confidence-weighted vote across all frames
flag position as uncertain if no frame clears threshold
This isn't novel in computer vision generally, but it mattered more than any single model architecture change for our rain/night failure buckets — a lot of "unreadable" single frames turn out to be very readable once you're not relying on just one of them.
Edge inference constraints we didn't anticipate
Terminal network conditions are inconsistent — gate cameras are often on infrastructure that predates the OCR system by a decade, and round-tripping every frame to the cloud for inference isn't reliable enough for a gate that needs to clear a truck in a few seconds. That pushed us toward running inference on edge hardware near the gate, which came with its own constraints we underestimated going in:
Model size and latency budgets are much tighter on edge compute than in a cloud environment, which directly limited how large our recognition backbone could be
Thermal and power constraints in outdoor gate enclosures ruled out some hardware options that looked fine on paper
Multi-frame fusion (above) has to happen with a bounded frame buffer on-device, not an arbitrarily large window, which meant tuning burst length as a real trade-off between accuracy and memory/latency, not just "more frames = better"
Where this leaves us
None of this makes the problem "solved" — dirty, low-light, wet-lens conditions are still where we spend the most engineering time, because they're where real terminals actually operate, especially in regions with heavy seasonal rain or minimal gate lighting infrastructure. But shifting the framing from "overall accuracy" to "accuracy per failure condition," splitting localization from recognition, and treating multi-frame fusion as a first-class part of the pipeline rather than an afterthought made the biggest measurable difference for us.
If you're building or evaluating container OCR — or any OCR system meant to run outdoors, unattended, in whatever weather shows up — I'd genuinely be curious how others have approached the same rain/backlight/low-light cluster of problems. Happy to compare notes in the comments.
We're a team building AI-powered port and logistics automation systems, including container OCR, smart gate, and terminal integration tooling. Follow for more engineering write-ups on the practical side of computer vision in industrial environments.



Top comments (0)