We update the operating system on our handheld scanners about four times a year. The update is around a gigabyte, it is distributed by our device management platform, and for three years it caused no trouble at all, because the platform staggers devices and every depot has a decent fibre line.
In May one of those lines was cut by roadworks at seven in the morning. The depot failed over to its mobile backup router, as designed, and carried on working. At nine our quarterly update was released to its first wave, which included that depot's eighty scanners. The platform saw devices online and an update pending. It had no way of knowing that online now meant a mobile connection with a monthly allowance of fifty gigabytes. By half past eleven the allowance was gone, the carrier throttled the line to a crawl, and a depot that had survived the cut perfectly well could no longer print a label. The excess charge arrived the following month.
What struck me afterwards was how sensible every part had been. The backup was sized for the warehouse system, which uses very little. The update was staged by device group, which is how we had been told to stage it. Nobody had connected the two, because the people who chose the backup lines and the people who scheduled updates sat in different teams, and each was right about their own half.
The fix had three parts. Each depot now has a single cache device that downloads an update once and serves it to the scanners over the local network, so the number of scanners at a site no longer multiplies the download. The routers set a site attribute in the device management platform when they fail over, and any site on backup is excluded from every distribution until it is back on fibre. And the backup routers raise an alert when they pass half of their monthly allowance, whatever the cause.
Last month two depots were on backup during a release. Both were skipped automatically and caught up the following night.
Our automation knew every device by name and knew nothing about the road between it and the device. Every distribution we run now has to ask what kind of connection it is about to use before it asks whether the device is online.
– Serguey Shinder
Top comments (0)