The last time a retry loop burned through our API quota, it didn't look like a crime scene at first. It looked like a loop that had given up thinking:
[14:03:12] sync start target=inventory, items=4800
[14:03:12] PUT /items/A-9931 → 404
[14:03:12] retry 1/3
[14:03:12] PUT /items/A-9931 → 404
[14:03:12] retry 2/3
[14:03:12] PUT /items/A-9931 → 404
[14:03:12] retry 3/3
[14:03:12] PUT /items/A-9931 → 404
[14:03:13] give up, next item
Read it like a detective. Four calls. Four identical answers. The server was consistent about what it was saying: A-9931 did not exist. Final answer. The client read the first three responses, filed them as interruptions, and asked again. Then it gave up, as if the last 404 had revealed something the first three hadn't.
That's the signature of a loop that never reads its own evidence.
The job pushed 4,800 inventory items into the catalog API. Most came back 200. The rest split into two piles. One pile was the server saying this resource is gone or this payload is wrong — 404 for items removed from the source, 400 for fields we'd let drift. Those answers don't change with time. Wait a minute, an hour, a day, the item is still gone and the payload is still wrong. The other pile was the server saying something went sideways on its end — the 5xx range, plus the throttle response the gateway sends when account headroom runs out. Those questions get different answers later. Those are the ones where a retry has a reason to exist.
We'd wrapped the whole client in a retry-on-error policy. One line, and it sounded like resilience. In practice that meant every non-2xx was treated as temporary, and the only open question was the retry count. Each dead item burned four requests — the original plus three retries — and every request spent quota. A discontinued SKU like A-9931 doesn't come back between retry 1 and retry 2, and it doesn't come back between retry 2 and retry 3. But the pattern held: four calls, four identical answers, three wasteful retries.
The timeline on this incident is short. Someone looked at the quota dashboard and saw a burn rate that didn't match the amount of actual work going through. The bill was the alert. The fix started the same afternoon.
We figured out one thing first, and it's barely a technical fact: a retry is a claim about the future. Every retry says this failure might resolve itself. For a 500, that's often true — the database restarts, the load balancer stops eating itself, the upstream hiccup passes. For the throttle response, it's true, but only after you've let the meter reset. For a 404, you're claiming the warehouse will restock itself between your first call and your second. For a 400, you're claiming the payload will change shape on its own. Those retries bought nothing and spent quota.
The first cut of the fix was a classifier: look at the status, ask whether this failure heals itself, then either retry or skip. The first version handled the 5xx range and skipped everything else. Review caught the hole
Top comments (0)