DEV Community

Sergey Shinder
Sergey Shinder

Posted on

A third of our calls came from an address our carrier had never been given

From the fourteenth of August, roughly a third of our booking calls to the carrier came back 403 with an HTML error page in the body. Not a pattern we could see: not one customer, not one route, not one time of day. Retrying usually worked. We raised it with their support, they found no rejections in their application logs, and for sixteen days it sat between two companies as somebody else's intermittent problem.

It was ours, and it had shipped in June. We added a third availability zone to the cluster for capacity. Each zone has its own NAT gateway, each gateway has its own elastic address, and the new one's address had never been sent to anybody. The carrier allowlists source addresses at their edge, which is why their application never saw a thing; the rejection happened before it reached them. Two of our three addresses had been agreed by email in 2022 and typed into a form on their side by an engineer who has since left.

About a third of our pods sit in the new zone, which is where the third of failing calls came from. A retry that happened to land on a pod in an older zone succeeded, and that single detail is what kept us looking at their reliability instead of our own network for a week and a half.

All partner traffic now leaves through one small egress proxy with a single static address, defined in Terraform next to the rest of the network, so the identity a partner sees is one value under change control rather than a side effect of how many zones we happen to run. A synthetic check calls each partner's simplest endpoint every five minutes from a pod pinned to each zone and asserts both the status and the source address it arrived from. And the egress module's README lists the partners who have that address configured at their end, which is the nearest thing we have to a warning label.

Part of your configuration lives inside somebody else's system, was typed there by hand years ago, and quietly becomes wrong every time you change your own infrastructure. Nothing in your repository knows it is there, and no test you own will fail when it breaks.

– Sergey Shinder

Top comments (0)