Every integration demo ends at the happy path. The purchase order parses, the webhook returns a 200, the invoice posts, everyone goes home. Then 2 a.m. arrives, and the part of the system that was never demoed — the error handling — quietly becomes the only part that matters.
EDI learned this lesson decades before web APIs existed, and it learned it the expensive way: a failed document is not a failed HTTP request. It is a truck that does not ship, a shelf that goes empty, and a chargeback with your name on it. Here is how to think about errors and retries the way battle-tested EDI systems do, translated for API-first integrations.
Not all errors are the same error
The single most damaging habit in integration code is treating every failure as retryable. EDI systems classify failures before they act, because the wrong response to a failure is worse than the failure itself:
- Transport failures — the file or request never arrived intact. Connection drops, timeouts, TLS errors. These are safe to retry, because nothing was processed.
- Syntax failures — it arrived, but it cannot be read. A malformed X12 interchange, an invalid JSON body, a failed signature check. Retrying changes nothing; the sender must fix the payload. In EDI these come back as a TA1 interchange acknowledgment, and the correct behavior is to stop and alert, not to resend the same broken bytes.
- Business-rule failures — it parsed, but the content is wrong. A purchase order references a SKU you do not carry, a ship date in the past, a total that does not foot. EDI answers with a 997 or 999 functional acknowledgment carrying segment-level error codes. Retrying the identical document will fail identically, forever.
- Downstream failures — your side accepted the document but the ERP, warehouse, or billing system rejected it. This is the silent killer: the partner believes you accepted their order because you acknowledged it, and nothing ships.
If your retry logic cannot tell these apart, it will eventually do something unforgivable, like re-sending a valid order three times and shipping three trucks.
Retry like you mean it: backoff, jitter, and a budget
EDI transmissions historically ran on scheduled windows — every fifteen minutes, hourly, nightly — which accidentally produced excellent retry discipline. Modern webhook culture retries instantly and repeatedly, which produces thundering herds.
A sane retry policy has four parts. First, exponential backoff: each attempt waits meaningfully longer than the last. Second, jitter: randomize the wait so a hundred failed integrations do not all retry in the same second. Third, a hard attempt budget: after N failures, the document goes to a dead-letter queue and a human gets paged — it does not retry until the sun burns out. Fourth, idempotency keys, so a retry is a question ("did this already post?") rather than a duplicate instruction. We covered the key mechanics in an earlier piece; the retry policy is what makes them necessary.
The detail most teams miss: make the retry budget visible. Log every attempt with the document control number, the failure class, and the next scheduled retry. "It eventually worked" is not observability.
Acknowledgments are a contract, not a courtesy
EDI's TA1, 997, and 999 acknowledgments exist because senders refuse to guess whether a document landed. APIs deserve the same contract. Acknowledge receipt fast, separately from processing — return the 200 when the payload is safely stored, not when the business logic finishes. Then report processing outcomes asynchronously, with error detail a machine can act on: which segment, which field, which rule.
And watch the acknowledgment stream itself. A partner whose 997s suddenly stop arriving has not become more reliable; their system is down, and your orders are piling up somewhere dark. Reconciliation sweeps — periodically comparing what you sent against what was acknowledged — catch what per-message error handling cannot.
A practical checklist
Classify every failure before you retry it. Retry only transport failures automatically; route syntax and business errors to a human-visible queue with the raw payload attached. Cap attempts, then dead-letter with an alert. Keep the original payload immutable so a corrected resend can be diffed against it. Test your failure paths deliberately — send the malformed file, kill the downstream, expire the certificate — in a sandbox, before production does it for you.
This is the mindset we build on at SignalEDI (https://signaledi.com?utm_source=devto&utm_medium=article&utm_campaign=2026-10-06): every document tracked from receipt to acknowledgment to business outcome, with failures classified and surfaced instead of silently retried. I am Chris, founder of SignalEDI — the happy path is easy; I would rather talk about your 2 a.m. path. What does your ugliest integration failure look like?
Top comments (0)