A customer incident usually starts with a support message in Slack: this store had this happen with an order. And from there the path is almost always the same.
The manual path
Find the pipeline logs, reconstruct which webhooks arrived and in what order, understand what happened, write a test that reproduces it and, finally, fix it. Almost all of it is mechanical, and the only thing that changes from one incident to the next is the case. That's why I built an agent that walks that path.
What the agent does
Starting from the Slack thread:
- It retrieves the pipeline logs.
- It reconstructs the real sequence of webhooks we received from the store.
- It generates a test that reproduces the case.
- It proposes the fix in a pull request.
When reconstructing the sequence you have to look at two clocks: the order in which each webhook arrived and the date on which the provider emitted it. Many incidents are exactly the difference between the two; I cover the layers I use to prevent them in Why your webhooks arrive duplicated.
First the test, then the fix
The order matters. In my area, every incident comes in with a test that stops it from happening again, so the agent doesn't start by fixing: it starts by reproducing.
A test that fails for the right reason is the proof that you have reproduced the problem, and reproducing it is the first step to understanding it. If the agent can't reproduce the case, there is nothing to fix yet. Sometimes the cause is that a piece of data is missing: a webhook that never arrived or an external result that was left in doubt. That is also a useful result, and it's better to know it before it proposes a blind patch.
For a webhook incident, the test has to be deterministic: the events in the order in which they happened, a fixed clock and fakes that respect the same constraints as production, like uniqueness and versions, because a fake without them makes a concurrency test lie. The payloads come from real cases, but they have to be scrubbed before they are left in the repository.
The panel that already existed
Before this I had built another piece in the same direction: an operations panel for the webhook pipeline, so that Customer Support can diagnose a case without escalating it to engineering. You search by store, provider and event type, you see the processing state live, and each order has its timeline with all its events chained by correlation ID.
Logs explain what a process did, but they can be sampled or have expired. The current state and its history of transitions live in a durable entity, and that's where it's best to read what happened. Logs and traces point to that state, they don't replace it.
It's the same idea from two sides: making the path between "something has failed" and "I know exactly what happened" short.
Top comments (0)