<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Serguey Shinder</title>
    <description>The latest articles on DEV Community by Serguey Shinder (@serguey_shinder_4ab9b87b1).</description>
    <link>https://dev.to/serguey_shinder_4ab9b87b1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4099211%2F490b857a-890d-407a-98a0-6e561f1e9ddc.png</url>
      <title>DEV Community: Serguey Shinder</title>
      <link>https://dev.to/serguey_shinder_4ab9b87b1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/serguey_shinder_4ab9b87b1"/>
    <language>en</language>
    <item>
      <title>Our Payslips Were Locked With Each Employee's Date of Birth</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Fri, 02 Oct 2026 07:11:35 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/our-payslips-were-locked-with-each-employees-date-of-birth-2ofa</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/our-payslips-were-locked-with-each-employees-date-of-birth-2ofa</guid>
      <description>&lt;p&gt;When we stopped printing payslips in 2021, the payroll team chose the simplest protection available. Each payslip went out as a PDF attached to an email, locked with a password, and the password was the employee's date of birth. Nobody had to be told it, nobody forgot it, and the service desk never had a single call about it.&lt;/p&gt;

&lt;p&gt;I came to it through a complaint. An employee had separated from her partner, who still used the family laptop, and her payslip, with her salary, her bank details and her national insurance number, had been opened in the shared mail client. He knew her date of birth. Of course he did.&lt;/p&gt;

&lt;p&gt;The problem was much wider than one household. A date of birth is not a secret. It is on the birthday list in most of our offices, in the HR system that a hundred people can read, and on the social media profiles of a good share of our staff. For a working adult there are only around twenty thousand plausible values, which a freely available tool will try against a PDF in seconds. And payslips travel. People forward them to mortgage brokers, letting agents and personal email accounts, so the attachment lives on in places we will never see, protected by something close to public.&lt;/p&gt;

&lt;p&gt;The scheme had been chosen because it never generated a support call. For the same reason, nobody had ever tested it by trying to get in.&lt;/p&gt;

&lt;p&gt;Payslips no longer leave us as attachments. Staff read them in the self service portal, behind the same sign in and second factor as everything else, and the email only says that a new payslip is available. Anyone who needs a copy for a lender downloads it themselves, and we can see that it happened. Warehouse staff and drivers without a company mailbox enrol a personal phone at induction. The old attachments were purged from our own mail system, which was the part our legal team cared about most.&lt;/p&gt;

&lt;p&gt;Calls to payroll rose for about six weeks and then settled below where they had been before.&lt;/p&gt;

&lt;p&gt;A password that can be read off a birthday card identifies a person. It does not protect anything. These days a control that has never caused anybody any friction is one of the first things I check, because friction is often the only evidence that a control is doing something.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>cybersecurity</category>
      <category>privacy</category>
      <category>payroll</category>
    </item>
    <item>
      <title>Our Electric Vans Will Turn a Midnight Outage Into a Morning Without Deliveries</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Fri, 02 Oct 2026 07:08:45 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/our-electric-vans-will-turn-a-midnight-outage-into-a-morning-without-deliveries-4ia3</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/our-electric-vans-will-turn-a-midnight-outage-into-a-morning-without-deliveries-4ia3</guid>
      <description>&lt;p&gt;We run just over a hundred electric vans in a fleet of about nine hundred, and the plan the board approved this year makes most of the fleet electric within eight years. I was not in the room for most of that decision, and I understand why. A van is a vehicle and vehicles belong to fleet. But the more of them we run, the more I expect the electric fleet to become one of the most important systems I am responsible for.&lt;/p&gt;

&lt;p&gt;A diesel van is ready in the morning because somebody filled it the afternoon before. An electric van is ready because software decided overnight how to share a limited grid connection between eighty chargers. The charge management platform at our largest yard knows each van's battery level, tomorrow's route from the transport planning system, and the electricity price by the half hour, and it schedules every charger accordingly. If it loses the route plan, it charges every van equally. If it loses contact with the chargers, they fall back to a slow default rate.&lt;/p&gt;

&lt;p&gt;In March a change to the planning system's interface broke that link just after midnight, and nobody noticed until six. The vans facing the longest routes had charged exactly as much as the ones doing short town runs, and twelve of them could not have completed their day. Nothing had gone wrong while anyone was working. The damage from a fault at night had simply waited for the morning shift to find it.&lt;/p&gt;

&lt;p&gt;That changes how I think about recovery. For most of our systems a few hours of downtime overnight barely registers. For the charging platform the night is the only time that matters, and an outage at one in the morning costs far more than the same outage at noon. The integration now has the same monitoring and on call cover as the warehouse system. The platform has a documented fallback that charges against yesterday's routes instead of none. And its recovery objective is written in kilometres of range lost by six o'clock, not in hours of downtime.&lt;/p&gt;

&lt;p&gt;Over the next decade I expect more of our operation to depend on software that works while nobody is watching and reveals its failures the following day. Our tiers of criticality were drawn around systems people use while they work. They will have to be redrawn around the systems that prepare the work.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>future</category>
      <category>electricvehicles</category>
      <category>itstrategy</category>
    </item>
    <item>
      <title>Our Marketing Team Waited Nine Weeks for Two Days of Work</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Fri, 02 Oct 2026 07:03:34 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/our-marketing-team-waited-nine-weeks-for-two-days-of-work-5f05</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/our-marketing-team-waited-nine-weeks-for-two-days-of-work-5f05</guid>
      <description>&lt;p&gt;In the spring our marketing director asked for a small change to the customer portal. She wanted the quote form to ask how a customer had heard about us, so that her team could stop guessing which campaigns brought in business. My team estimated two days. She received it nine weeks later, and in the meantime she ran an entire campaign she had no way to measure.&lt;/p&gt;

&lt;p&gt;When she raised it at the leadership meeting, I defended my team, because the estimate had been honest and the work, once started, took exactly two days. That was true, and it was beside the point. She had never asked how long it would take us to do. She had asked when she would have it.&lt;/p&gt;

&lt;p&gt;So we measured what she had measured. For every request completed in the previous six months, we compared the days of work with the days between asking and receiving. The median request needed three days of effort and took forty seven days to arrive. The gap was widest for the smallest requests, because a small change had nobody pushing it and was always the easiest thing to put back. It waited for triage, then for an estimate, then for room in a sprint, then for a release window. Each wait was short and each had a reason. Nobody owned the total.&lt;/p&gt;

&lt;p&gt;Our own reports showed effort, utilisation and delivery against estimate, and on all three we looked good. The business never saw any of those numbers. It saw how long it waited.&lt;/p&gt;

&lt;p&gt;We now publish the time from request to delivery every month, by department, beside the effort it took. Changes estimated at under three days go into a separate lane staffed by two people, are estimated once, and go out in the next weekly release. We limit how many large pieces of work can be in progress at the same time, which slowed the start of new projects and noticeably sped up their finish. Small requests from marketing now arrive in about eight days.&lt;/p&gt;

&lt;p&gt;The effort had always been reasonable. What we had never managed was the time a request spent waiting between people who were each doing their own job well. Most of what a department experiences as IT being slow is time in which nobody in IT is working on its request at all.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>itbusiness</category>
      <category>itmanagement</category>
      <category>delivery</category>
    </item>
    <item>
      <title>Our Lorries Were Serviced by the Mileage Drivers Typed at the Fuel Pump</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Fri, 02 Oct 2026 06:58:23 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/our-lorries-were-serviced-by-the-mileage-drivers-typed-at-the-fuel-pump-4b5a</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/our-lorries-were-serviced-by-the-mileage-drivers-typed-at-the-fuel-pump-4b5a</guid>
      <description>&lt;p&gt;Our fleet workshop schedules servicing and safety inspections for about four hundred heavy vehicles. A few checks are due by date, but most are due by distance, and the maintenance system works out when a vehicle needs to come in from the latest mileage it holds.&lt;/p&gt;

&lt;p&gt;That mileage came from our fuel cards. When a driver fills up, the pump terminal asks for the odometer reading, and the card company sends us a file of every transaction each night. The file already existed, and the workshop needed a number from somewhere.&lt;/p&gt;

&lt;p&gt;Last winter a tractor unit went past its brake inspection interval by nine thousand kilometres. Nothing failed, but our transport manager found it in a routine audit, the kind of finding that ends up in front of a regulator. The vehicle's recorded mileage had barely moved in two months. Drivers at the pump, often in the dark, with a queue behind them and a terminal that would not continue without a number, had been typing whatever got them through. The previous reading. A row of zeros. In one case, the driver's own staff number. The system simply took the latest value, so one low entry could quietly push a vehicle's next service months into the future.&lt;/p&gt;

&lt;p&gt;Across the fleet, roughly one transaction in seven carried a reading that could not be true, either going backwards or jumping further than a lorry can travel between two fills. Thirty one vehicles were overdue for something.&lt;/p&gt;

&lt;p&gt;Every one of those vehicles already reported its real distance through the telematics unit we had fitted for route tracking. Nobody had ever connected the two systems.&lt;/p&gt;

&lt;p&gt;The maintenance system now takes its mileage from telematics. The pump reading is still stored, but only as a cross check, and when the two disagree by more than a few hundred kilometres the workshop gets a task to look at the vehicle. A reading that goes backwards or implies an impossible distance is rejected and flagged rather than saved. When a telematics unit stops reporting, its vehicle is booked in by date until the unit is repaired.&lt;/p&gt;

&lt;p&gt;We had chosen that source because it was convenient, not because anyone entering it had a reason to get it right. A number typed by somebody who needs to get past a screen says very little about the thing it claims to measure.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>data</category>
      <category>dataquality</category>
      <category>fleet</category>
    </item>
    <item>
      <title>Switching Off an Old Feature Flag Sent Six Hundred Bookings Nowhere</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Fri, 02 Oct 2026 06:53:12 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/switching-off-an-old-feature-flag-sent-six-hundred-bookings-nowhere-53fg</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/switching-off-an-old-feature-flag-sent-six-hundred-bookings-nowhere-53fg</guid>
      <description>&lt;p&gt;Our collection booking site, where business customers arrange for us to pick up from their premises, is built and run by a team of nine developers. They release several times a week, and like most teams who release often they use feature flags, so new code can go out switched off and be turned on gradually.&lt;/p&gt;

&lt;p&gt;In June the site slowed down during the afternoon peak. The engineer on call opened the flag console, found a flag called new_booking_flow, and decided that switching it off was the quickest way to take a recent change out of the path. The site recovered. Bookings kept arriving, customers kept receiving confirmations, and the incident was closed in twenty minutes.&lt;/p&gt;

&lt;p&gt;The flag was not recent. It had been created in 2024 to move bookings from the old flow to the new one, and nobody had ever removed the old code behind it. With the flag off, the site went back to the old flow, which wrote bookings into a table the transport planning system had stopped reading when we migrated. For three hours, six hundred and twenty collections were accepted, confirmed by email and sent nowhere. We found out the next morning, when customers rang to ask where their drivers were.&lt;/p&gt;

&lt;p&gt;The console held four hundred and thirty flags. About three hundred had been fully on or fully off for more than six months. Some were finished releases with dead code still attached. Others were genuine operational switches, such as turning off card payments if the payment provider fails. In the console they looked identical, and nothing told the engineer which kind he was holding.&lt;/p&gt;

&lt;p&gt;Every flag now gets a type, an owner and a review date when it is created. A release flag must be removed, code and all, within six weeks of reaching every customer, and the build fails once that date has passed, so the clean up cannot be put off indefinitely. Operational switches are listed in the runbook with what they do and when they must not be used, and each one is exercised in staging every quarter. The console shows the type beside the name. We are down to sixty flags.&lt;/p&gt;

&lt;p&gt;A flag that is never removed stops being a release tool and becomes an untested second version of the product. Switching one off during an incident is a deployment, even when it takes a single click.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>devops</category>
      <category>featureflags</category>
      <category>automation</category>
    </item>
    <item>
      <title>Our Finance Server Ran Out of Credits on the Second Day of Month End</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Fri, 02 Oct 2026 06:48:01 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/our-finance-server-ran-out-of-credits-on-the-second-day-of-month-end-28p8</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/our-finance-server-ran-out-of-credits-on-the-second-day-of-month-end-28p8</guid>
      <description>&lt;p&gt;When we moved the finance reporting server to the cloud, we sized it from a year of monitoring. Its processor averaged about twelve percent and rarely passed forty, so the team chose one of the provider's burstable instance types, which is cheaper because it assumes you will be idle most of the time. For three weeks of every month that assumption was exactly right.&lt;/p&gt;

&lt;p&gt;Month end began on the first working day of April. The consolidation jobs started at six, the finance team started running reports at eight, and by the middle of the second day a report that normally took four minutes was taking forty. Overnight jobs ran into the morning. The finance director asked me whether the cloud was simply slower than our old server room, and I could not answer her, because every graph I looked at showed the processor at twenty percent.&lt;/p&gt;

&lt;p&gt;Twenty percent was the answer. A burstable instance earns credits while it is quiet and spends them when it is busy, and once they are gone it is held at its baseline, which for that size was a fifth of a processor. The graph was not showing a machine with room to spare. It was showing a machine pinned to its ceiling, and the ceiling looked like headroom. The provider published the credit balance the whole time, as a separate metric, and nobody had put it on a dashboard, because no server we had ever owned had such a thing.&lt;/p&gt;

&lt;p&gt;That month's bill came in lower than forecast, which in hindsight was the clearest signal of all.&lt;/p&gt;

&lt;p&gt;The reporting server now runs on a fixed performance instance for the last working day and first three working days of each month, and drops back afterwards on a schedule finance can see. Every remaining burstable instance alerts when its credit balance falls below forty percent. Our sizing template asks for the busiest consecutive day rather than the monthly average. Reviewing the rest of the estate found two more servers living on credits, and one database whose disk throughput had its own allowance that was used up every night by the backup.&lt;/p&gt;

&lt;p&gt;A cloud resource can run out of things other than money. For every service we size, I now ask what it is quietly spending besides our budget, and whether anyone can see the balance.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>cloud</category>
      <category>infrastructure</category>
      <category>capacity</category>
    </item>
    <item>
      <title>Every Ninety Days Our Service Desk Had the Same Bad Week</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Fri, 02 Oct 2026 06:37:08 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/every-ninety-days-our-service-desk-had-the-same-bad-week-6o6</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/every-ninety-days-our-service-desk-had-the-same-bad-week-6o6</guid>
      <description>&lt;p&gt;For about a year my service desk had a bad week roughly once a quarter. Calls doubled on the Monday, the contact centre started the early shift short of agents, and by Thursday everything had calmed down again. Each review blamed something different and named nothing in particular.&lt;/p&gt;

&lt;p&gt;It was a team leader in the contact centre who saw the pattern, because she kept a diary of the mornings she started short handed. Her bad mornings were thirteen weeks apart, almost to the day.&lt;/p&gt;

&lt;p&gt;The cause was a project we had finished and celebrated. In the spring of 2024 we moved every account into a new directory over one weekend, and on the Monday everybody set a new password. Our policy expired passwords after ninety days. So two thousand three hundred passwords had been set on the same day, expired on the same day, and kept expiring together every quarter after that. At head office that meant reset calls. The contact centre had it worse. Agents changed the password at the desktop, but the softphone and the customer system each held the old one and kept retrying it, and within minutes the account locked. An agent who changed hers at five to eight was locked out by eight, and so was most of her row.&lt;/p&gt;

&lt;p&gt;Each lockout was real and each reset was handled properly. What nobody saw was that one weekend had set a clock running on every account in the company.&lt;/p&gt;

&lt;p&gt;We stopped expiring passwords on a schedule for accounts with a second factor, which is what most of the guidance we follow now recommends, and a password is changed when there is reason to think it is known. Self service reset works from the sign in screen for anyone enrolled. A lockout caused by an application retrying an old password raises an alert that names the application, and the softphone now signs in through single sign on, so there is no second copy to go stale.&lt;/p&gt;

&lt;p&gt;The quarterly bad week has gone, and reset calls are down by about two thirds.&lt;/p&gt;

&lt;p&gt;A project that touches every account on the same day can leave behind a timer nobody set on purpose. After a large migration I now ask what it has synchronised, as well as what it has changed.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>itoperations</category>
      <category>servicedesk</category>
      <category>identity</category>
    </item>
    <item>
      <title>Our Software Bills Have Started Moving With the Number of Parcels</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:50:23 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/our-software-bills-have-started-moving-with-the-number-of-parcels-482m</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/our-software-bills-have-started-moving-with-the-number-of-parcels-482m</guid>
      <description>&lt;p&gt;For most of my career, software was the most predictable line in my budget. We bought a number of licences, usually named users or servers, and paid the same every quarter whether the business had a good month or a bad one. I could forecast it in September for the following year and be within a few percent.&lt;/p&gt;

&lt;p&gt;That is changing, and I expect it to keep changing. At its last renewal our transport management supplier moved from a fee per user to a fee per shipment planned. Our customer notification platform charges per message sent. The new slotting tool in the warehouse is priced per order line. The analytics product our finance team wants is priced on compute consumed, and the assistant features appearing in half of our software are sold as credits. None of this is unreasonable. Suppliers want to be paid in proportion to the value they deliver, and a parcel is closer to value than a seat is.&lt;/p&gt;

&lt;p&gt;What it does to an IT department is less comfortable. Last December, our busiest month, those lines together cost forty one percent more than in an average month, and I learned that from the invoices in January. A department that has always forecast from a list of contracts now has costs that move with volume, and the volume forecast belongs to operations and commercial, who have never had a reason to send it to me.&lt;/p&gt;

&lt;p&gt;So we are starting to plan these lines the way our transport team plans fuel. They are budgeted as a rate per parcel multiplied by the volume plan, and when the volume plan changes, my forecast changes with it. Usage is reported monthly beside the volume it served, so a rising rate shows up as a rate rather than hiding inside a growing total. And every new consumption contract has to include a ceiling or a band, because a supplier paid per message has no reason to tell us that our system has started sending each message twice.&lt;/p&gt;

&lt;p&gt;Within a decade I expect a large share of what IT spends to be a variable cost of trading. The board's question will stop being what IT costs per year and become what it costs per parcel, and I would like to have that answer ready before anybody asks.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>future</category>
      <category>saas</category>
      <category>itbudget</category>
    </item>
    <item>
      <title>Our Depots Ran on the Laptops Head Office Had Finished With</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:45:11 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/our-depots-ran-on-the-laptops-head-office-had-finished-with-9d2</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/our-depots-ran-on-the-laptops-head-office-had-finished-with-9d2</guid>
      <description>&lt;p&gt;For as long as anybody could remember, our laptop refresh worked as a cascade. Every year we bought around three hundred new machines and they went to head office first, roughly in order of grade. The machines they replaced were wiped and sent to the depots, and the depots' oldest machines were scrapped. It looked thrifty, and the finance director liked it, because every laptop worked for six or seven years before anybody threw it away.&lt;/p&gt;

&lt;p&gt;I only questioned it when I sorted a year of hardware tickets by where the device sat rather than by model. Depot machines produced three times as many tickets per device as head office machines, and the reason was not mysterious. A head office laptop is used by one person for eight hours and spends the evening in a bag. The transport office terminal at a depot is used by three planners across three shifts, is never switched off, sits beside a roller door, and is between six and seven years old because that is where the cascade ends. Out of warranty, a failed one waits for a spare from head office, and at a depot a planner without a terminal means trailers without a plan.&lt;/p&gt;

&lt;p&gt;So the oldest equipment in the company was doing the most hours, on the work that stops vehicles when it stops, while the newest was mostly running email and spreadsheets in meeting rooms. Nobody had decided that. It was a side effect of refreshing people by grade rather than machines by use.&lt;/p&gt;

&lt;p&gt;We now refresh by measured hours of use and by what stops when the device stops. Shared shift terminals at depots are first in the queue and are replaced at four years whatever their condition. Single user office laptops run to five, and the cascade has ended, because a machine that is too old for a director is too old for a planner. Two senior colleagues objected, and one of them withdrew the objection after spending a night shift in a transport office with us.&lt;/p&gt;

&lt;p&gt;Depot hardware tickets per device fell by about half in the first year, and the number of laptops we buy has hardly changed.&lt;/p&gt;

&lt;p&gt;The cascade sorted our people by seniority and our equipment by nothing at all. Allocating by use cost no more, and it put the new machines where the work is done.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>itbusiness</category>
      <category>endusercomputing</category>
      <category>itmanagement</category>
    </item>
    <item>
      <title>Our Morning Report Was Built Before the Rural Vans Had Sent Their Day</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:40:00 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/our-morning-report-was-built-before-the-rural-vans-had-sent-their-day-33c0</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/our-morning-report-was-built-before-the-rural-vans-had-sent-their-day-33c0</guid>
      <description>&lt;p&gt;Every morning at five our delivery performance report is built from the previous day's proof of delivery records, and by eight it is in front of every depot manager. It counts a delivery as on time if the handset recorded a signature or a photograph inside the customer's window.&lt;/p&gt;

&lt;p&gt;For most of a year, two of our depots in the hills sat at the bottom of that report, six or seven points behind the rest. One of their managers was put on an improvement plan. He kept saying his drivers were delivering on time, and we kept showing him the report.&lt;/p&gt;

&lt;p&gt;He was right. Our handsets store a delivery when there is no signal and send it when there is. On parts of those two depots' routes a driver can go four or five hours without coverage, and some vans come back late with a day's records still waiting on the handset, which only finishes sending once it is docked. Most records arrived before five in the morning. A meaningful share arrived after, and by then the report had already counted them as deliveries that never happened. Customer services, reading the same data at half past eight, were raising a few dozen redelivery requests a week for parcels that were already sitting on doormats.&lt;/p&gt;

&lt;p&gt;Nothing was wrong with any individual record. Each one held the correct time of the delivery. What we had never stored was the time it reached us, so nobody could see how much of yesterday was still in transit when the report was frozen.&lt;/p&gt;

&lt;p&gt;Every record now carries both times. The report shows, for each depot, what share of the day's expected records had arrived when it was built, and its figures are marked provisional for forty eight hours and final after that. Customer services do not raise a redelivery for a stop while that van's handset is still holding unsent records. And a handset that has sent nothing for six hours raises a ticket for its depot, because by then it is usually a fault rather than a valley.&lt;/p&gt;

&lt;p&gt;Once the late records were counted, both depots finished the year above the company average. The improvement plan was withdrawn, and I apologised to the manager in person.&lt;/p&gt;

&lt;p&gt;A report with a fixed cut off had been measuring how quickly our data travels as well as how well our drivers worked, and it presented the two as a single number.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>data</category>
      <category>reporting</category>
      <category>dataquality</category>
    </item>
    <item>
      <title>Our Scanner Update Went Out Over the Backup Line and Used a Month of Data by Lunchtime</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:34:49 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/our-scanner-update-went-out-over-the-backup-line-and-used-a-month-of-data-by-lunchtime-42hf</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/our-scanner-update-went-out-over-the-backup-line-and-used-a-month-of-data-by-lunchtime-42hf</guid>
      <description>&lt;p&gt;We update the operating system on our handheld scanners about four times a year. The update is around a gigabyte, it is distributed by our device management platform, and for three years it caused no trouble at all, because the platform staggers devices and every depot has a decent fibre line.&lt;/p&gt;

&lt;p&gt;In May one of those lines was cut by roadworks at seven in the morning. The depot failed over to its mobile backup router, as designed, and carried on working. At nine our quarterly update was released to its first wave, which included that depot's eighty scanners. The platform saw devices online and an update pending. It had no way of knowing that online now meant a mobile connection with a monthly allowance of fifty gigabytes. By half past eleven the allowance was gone, the carrier throttled the line to a crawl, and a depot that had survived the cut perfectly well could no longer print a label. The excess charge arrived the following month.&lt;/p&gt;

&lt;p&gt;What struck me afterwards was how sensible every part had been. The backup was sized for the warehouse system, which uses very little. The update was staged by device group, which is how we had been told to stage it. Nobody had connected the two, because the people who chose the backup lines and the people who scheduled updates sat in different teams, and each was right about their own half.&lt;/p&gt;

&lt;p&gt;The fix had three parts. Each depot now has a single cache device that downloads an update once and serves it to the scanners over the local network, so the number of scanners at a site no longer multiplies the download. The routers set a site attribute in the device management platform when they fail over, and any site on backup is excluded from every distribution until it is back on fibre. And the backup routers raise an alert when they pass half of their monthly allowance, whatever the cause.&lt;/p&gt;

&lt;p&gt;Last month two depots were on backup during a release. Both were skipped automatically and caught up the following night.&lt;/p&gt;

&lt;p&gt;Our automation knew every device by name and knew nothing about the road between it and the device. Every distribution we run now has to ask what kind of connection it is about to use before it asks whether the device is online.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>automation</category>
      <category>devops</category>
      <category>mobiledevices</category>
    </item>
    <item>
      <title>Two Thousand Devices Had Our Server's Address Typed Into Them</title>
      <dc:creator>Serguey Shinder</dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:29:38 +0000</pubDate>
      <link>https://dev.to/serguey_shinder_4ab9b87b1/two-thousand-devices-had-our-servers-address-typed-into-them-2h09</link>
      <guid>https://dev.to/serguey_shinder_4ab9b87b1/two-thousand-devices-had-our-servers-address-typed-into-them-2h09</guid>
      <description>&lt;p&gt;Moving the warehouse system out of our data centre was planned as a weekend. The application had been rehearsed twice, database replication was running, and the cutover plan fitted on two pages. We expected to change one DNS record on Saturday night and go home.&lt;/p&gt;

&lt;p&gt;Three weeks before the date, during a rehearsal at a single depot, we found out that a DNS record was not what most of our devices used. The scanners pointed at the server by its IP address, entered in a configuration profile in 2016. So did the label printers, which had the address typed into each one through its front panel. The weighbridge terminals at four depots kept it in a text file. Two conveyor controllers had it compiled into their software by the integrator who installed them, and the engineer who did it no longer worked there. A name can follow a server to the cloud. An address belongs to the network we were leaving.&lt;/p&gt;

&lt;p&gt;Nobody had chosen this. Each device had been set up by whoever installed it, an address was what the installation guide showed, it worked, and it kept working for ten years because the server never moved. Our inventory recorded the model, serial number and site of every device. It did not record what any of them talked to.&lt;/p&gt;

&lt;p&gt;The weekend became four months. The scanners were straightforward, one profile pushed from device management. The printers needed somebody at the panel of each one, a hundred and twelve times. The conveyor controllers needed the integrator back, at a day rate, in a maintenance window operations could only give us on a Sunday. For those four months we ran a forwarder from the old address to the new environment, which added a little delay to every transaction and a single point of failure we had not wanted.&lt;/p&gt;

&lt;p&gt;Every device now refers to services by name, and those names carry a short time to live. The inventory has a field for what each device connects to, filled in at installation, and the standard we give contractors says a name or nothing. Before any future move, a report lists every device whose configuration contains a raw address. Today it lists nine.&lt;/p&gt;

&lt;p&gt;The server took a weekend to move. The references to it had been written into places we did not know we owned, and finding them was the actual migration.&lt;/p&gt;

&lt;p&gt;– Serguey Shinder&lt;/p&gt;

</description>
      <category>sergueyshinder</category>
      <category>infrastructure</category>
      <category>migration</category>
      <category>networking</category>
    </item>
  </channel>
</rss>
