DEV Community

Cover image for Your integration's real bugs are in the events you can't trigger
lna_stub
lna_stub

Posted on

Your integration's real bugs are in the events you can't trigger

Every integration I have written against a third-party API has the same shape.
The request/response part is easy and I get it right on day one. The failures show up later, in the branches that only run on rare events: the chargeback that lands a month after a clean payment, the KYC rejection, the webhook whose signature my code accepted without really checking, the callback that arrives out of order.

Those branches are exactly the ones a provider sandbox cannot help me exercise.
I cannot ask a real sandbox to charge back a payment on cue, or to deliver a webhook twice, or to send completed before I finished handling pending. And a classic mock server does not help either: it answers one request at a time with no memory, so it cannot represent "this payment already succeeded, now it is disputed".

So I wrote a small tool that treats an integration as what it actually is: a state machine.

The idea

You describe a scenario in YAML. A session remembers where a client is in the flow, keyed by whatever identifies it in your API - a body field, a header, a query or a path parameter. The same endpoint then answers according to the current state:

GET /v1/orders/ord_42/payment  ->  processing
   (a few seconds pass, a webhook fires)
GET /v1/orders/ord_42/payment  ->  succeeded
   (later, another webhook fires)
GET /v1/orders/ord_42/payment  ->  chargeback
Enter fullscreen mode Exit fullscreen mode

Entering a state can schedule webhooks: delayed, HMAC-signed in the Stripe format, with retries, a delivery journal and replay. That is the part I care about most, because the webhook path is where my code has the most untested branches. Does the handler reject a forged signature? Does it answer 2xx so the sender stops retrying? Does the same event delivered twice settle once?

Two flags make it usable in CI. --time-scale compresses "30 days later" into seconds, so a full lifecycle runs in a test. --seed makes the template functions (ids, timestamps) reproducible, so snapshot assertions are stable.

What it is not

It does not record or proxy real traffic, and it has no GUI. If you need those, WireMock and Mockoon are the right tools. This one is aimed narrowly at deterministic, versionable simulation of event-driven APIs.

And there is one honest limitation worth stating up front: a twin is only as accurate as the scenario you write. If the YAML mismodels the real response, the twin will happily repeat your mistake. The workflow that works for me is to confirm the response shape once against the provider's own sandbox, then use the tool to drive the hundred edge cases the sandbox cannot reach.

Practical notes

It is a single Go binary. No cloud, no account, MIT licensed. Every miss returns a diagnostic 404 that lists the closest matchers and the reason each one was rejected, which turns "why didn't this match" from guesswork into reading. Hot reload does not break active sessions - they finish on the config version they started with. Outbound webhook delivery has an SSRF guard and never follows redirects.

Try it

TwinStub is one binary that talks HTTP, so your integration stays in whatever language it already is. Install it with Docker, a prebuilt binary, or Go:

docker run -p 8080:8080 -p 9090:9090 ghcr.io/twinstub/twinstub
# or: go install github.com/twinstub/twinstub/cmd/twinstub@latest
# or: download a binary from the Releases page

twinstub init demo && cd demo
twinstub serve
Enter fullscreen mode Exit fullscreen mode

The generated project is the chargeback flow above. Repo, docs and a catalog of fintech scenarios: https://github.com/twinstub/twinstub

If you integrate with payment, logistics or CRM APIs, I would like to know where this model helps and where it breaks down for you.

Top comments (6)

Collapse
 
muhammad_turnergane_7ddc profile image
Muhammad Turner Gane •

"This payment already succeeded, now it is disputed" is the real reason sandboxes fall short. You can fire a single event on cue with the Stripe CLI, but ordering and timing are much harder to control, and those are the cases that leave a paying customer stuck on the free plan.

The limitation you call out is the one I'd underline too. One assertion I'd add at the end of every scenario is to check what the user can actually access in the app once the whole lifecycle has run, not only that every webhook got a 2xx. A handler can answer 200 to all of them and still leave the account in the wrong state.

Collapse
 
lna_stub profile image
lna_stub •

Exactly - the end-state assertion catches the real bugs; a 2xx on every delivery says almost nothing. The provider is only the stand-in. What you're testing is whether your own order, ledger or plan landed in the right state after the full sequence, and a handler can ack everything and still leave the account on the wrong plan.

Ordering and timing are the hard part: a single event on cue is easy, but "succeeded then disputed", the duplicate retry, or completed arriving before pending is what a sandbox won't stage. That's the lane I built this for, and the delivery journal lets you replay a specific bad ordering deterministically, so the "what can the user access now" check stays stable in CI. Thanks for the sharp addition - I may put "assert the resulting app state, not the acks" in the docs as the way to end a scenario.

Collapse
 
muhammad_turnergane_7ddc profile image
Muhammad Turner Gane •

glad it landed. one thing that pairs well with replaying a specific bad ordering: write the expected end state for each ordering as a small table - same events, different order, plan the user should end on. snapshot assertions tend to lock in whatever the handler did the first time, bug included. a table says what should be true, not what happened last time.

Thread Thread
 
lna_stub profile image
lna_stub •

Yes - the table is the spec, the snapshot is just a recording. Snapshotting the whole response quietly enshrines the current behavior, bug and all, and then stays green forever. Writing "these events in this order -> this plan" as an explicit table is the version that can actually fail when the handler is wrong. TwinStub's job is the left column (produce each ordering deterministically); the right column is yours to assert. That pairing is exactly what I'll write up in the docs.

Collapse
 
challan116ux profile image
challan116-ux •

This is exactly where EDI integrations break too — the events you cannot trigger are usually the ones your trading partner invents, like a 999 acknowledgment rejecting one segment inside an otherwise valid file. What has worked for us is treating every real failed document as a permanent regression test: save the exact payload that failed, and that scenario can never surprise you twice. The untriggerable list shrinks fast once production incidents become fixtures instead of memories.

Collapse
 
lna_stub profile image
lna_stub •

That incident-to-fixture discipline is the whole game, and it ties straight back to the one honest limitation: a twin is only as accurate as the scenarios you write, so every real failure that becomes a permanent scenario makes it more accurate. "Fixtures instead of memories" is a great way to put it - the 999 rejecting one segment is exactly the thing nobody stages until prod does it for you. I kept scenarios as plain YAML for this reason: a payload that failed once gets committed and can't surprise you twice. (TwinStub speaks HTTP/JSON rather than raw X12, but the discipline is identical.)