DEV Community

EllisThornton7395
EllisThornton7395

Posted on

How to Audit Default Routing for Vendor Concentration Risk (Under a Spend Ceiling)

Route payment-platform events to the default supplier until its failure boundary is crossed, then admit fallback traffic only while a predeclared spend budget remains. Keep the second path warm with synthetic events, but never let a health check create a ledger entry. The governing rule is bounded failover, not unconditional availability: once the fallback allowance is exhausted, refuse new traffic with a retryable response and preserve the events already accepted.

Short answer: use one stable event ID across both suppliers, separate acceptance from downstream processing, and write an append-only routing decision before dispatch. A fallback can reduce concentration risk, but it cannot provide exactly-once delivery by itself; the backend must make repeated delivery harmless and leave enough evidence to reconcile every accepted, refused, retried, and completed event.

This architecture decision record treats a supplier outage as a normal, testable state. Its primary trade-off is explicit: a payment platform may buy continuity up to a fixed ceiling, after which controlled refusal is safer than an unbounded bill or an unauditable duplicate.

How can default routing keep vendor concentration risk bounded?

Identity first.

The producer assigns an immutable event ID before any routing decision, and every attempt carries that same ID. A supplier-specific request ID is useful diagnostic data, but it must not become the business idempotency key because changing routes would then change identity. The second invariant is durable acceptance: return success only after the event and its routing state are durably recorded, even though processing may happen later. This boundary prevents an ambiguous success from becoming an invisible payment mutation, while a unique constraint on the event ID makes a producer retry converge on the original record. The third invariant is an audit trail that records decisions rather than rewriting history. Each attempt needs the event ID, selected route, policy version, budget reservation, outcome, and timestamp; the useful record says why evt_01 moved from the default route to the second provider under dual-route-v1, not merely that some request happened. Secrets do not belong in that record. OWASP recommends centralizing secrets management, applying least privilege, rotating secrets, and auditing their use; independent credentials for each supplier keep a compromise or rotation from collapsing both paths. These three invariants must be evaluated together because stable identity without durable acceptance loses acknowledged work, durable acceptance without attempt history frustrates reconciliation, and detailed history without credential separation can turn an audit system into another source of exposure.

Exactly once is the mindset, not a transport promise. The observable business effect must occur once even when delivery is repeated, a response is lost, or the router restarts between dispatch and acknowledgement. Reconciliation therefore compares accepted events, attempt records, and committed ledger effects by stable ID. Silence is not success.

No guesswork.

The failure boundary is deliberately narrow: timeouts, connection failure, or a retryable upstream status may open fallback eligibility, while invalid input and authorization failures do not. Sending malformed or unauthorized traffic to another supplier spends money without changing the answer. The router also stops when its budget ledger cannot be read; guessing that capacity remains would violate the spend ceiling.

Record the decision before sending traffic

The critical path has two state machines. The event moves from accepted to completed or terminally rejected. Separately, each delivery attempt moves from reserved to sent and then acknowledged, retryable, or terminal. Keeping these states separate makes the unpleasant interval between an upstream side effect and a lost response visible.

Use an outbox in the same transaction as event acceptance. A worker claims the outbox row, asks the routing policy for a route, and records that decision before making the network call. If the worker crashes after the call, it retries with the same event ID. The downstream adapter, and ultimately the ledger mutation, must enforce that ID as unique.

The following Go program models the policy boundary without tying it to a commercial API. It is intentionally small enough to run, but the Budget interface represents a transactional reservation store, not an in-process counter in a multi-instance deployment.

package main

import (
    "context"
    "errors"
    "fmt"
    "sync"
)

var ErrFallbackBudgetExhausted = errors.New("fallback budget exhausted")

type Event struct {
    ID      string
    Account string
    Cents   int64
}

type Budget interface {
    Reserve(ctx context.Context, eventID string, units int64) (bool, error)
}

type MemoryBudget struct {
    mu       sync.Mutex
    limit    int64
    used     int64
    reserved map[string]bool
}

func (b *MemoryBudget) Reserve(_ context.Context, eventID string, units int64) (bool, error) {
    b.mu.Lock()
    defer b.mu.Unlock()
    if b.reserved[eventID] {
        return true, nil
    }
    if units < 0 || b.used > b.limit-units {
        return false, nil
    }
    b.used += units
    b.reserved[eventID] = true
    return true, nil
}

type Decision struct {
    EventID      string
    Route        string
    Policy       string
    ReservedUnit int64
}

func chooseRoute(ctx context.Context, e Event, primaryHealthy bool, b Budget) (Decision, error) {
    if primaryHealthy {
        return Decision{EventID: e.ID, Route: "primary", Policy: "dual-route-v1"}, nil
    }

    const estimatedUnits int64 = 1
    ok, err := b.Reserve(ctx, e.ID, estimatedUnits)
    if err != nil {
        return Decision{}, fmt.Errorf("read fallback budget: %w", err)
    }
    if !ok {
        return Decision{}, ErrFallbackBudgetExhausted
    }
    return Decision{
        EventID: e.ID, Route: "fallback", Policy: "dual-route-v1", ReservedUnit: estimatedUnits,
    }, nil
}

func main() {
    b := &MemoryBudget{limit: 2, reserved: make(map[string]bool)}
    e := Event{ID: "evt_01", Account: "acct_17", Cents: 12500}
    d, err := chooseRoute(context.Background(), e, false, b)
    fmt.Printf("decision=%+v error=%v\n", d, err)
}
Enter fullscreen mode Exit fullscreen mode

Run it as a policy test, then replace the memory implementation with a database transaction that atomically checks the ceiling, reserves capacity by event ID, and inserts the audit row. The overflow-safe comparison used > limit-units matters because budget arithmetic belongs on the correctness path. Units should be conservative accounting units established by contract and policy; reconciliation can later replace a reservation with the final attributable usage without allowing the admission decision to exceed its ceiling.

One nuance deserves emphasis. A reservation retry for the same event returns the original decision and consumes no additional unit. That idempotency property is more important than the particular storage engine.

Two states, one identity.

Compare the operating modes

The relevant choice is not “one supplier or two” in the abstract. It is which failure mode the organization is prepared to own, observe, and rehearse.

Mode Outage behavior Spend boundary Main operational burden Appropriate when
Single route Refuse or queue all affected events One contracted path Restore and drain safely The business can tolerate delayed acceptance
Cold secondary Activate credentials and configuration during an incident Manual approval can cap exposure Untested drift and slow recovery Outages are tolerable long enough for a controlled change
Warm, bounded secondary Admit until a transactional allowance is exhausted Enforced before dispatch Continuous probes, reconciliation, and two integrations Limited continuity is worth ongoing validation
Active split Send production traffic through both paths routinely Requires allocation on every request Permanent cross-route consistency Both paths must remain exercised by real workload

For the stated decision axis, the warm bounded mode is the recorded choice. It exposes a hard stopping condition rather than treating fallback as infinite capacity. The rejected active split does keep both integrations exercised, but it expands the everyday reconciliation surface and makes supplier divergence a routine condition. It remains valid when contractual or regulatory policy requires continuous distribution, or when a team can reconcile both routes as a normal operation rather than an emergency procedure.

The single-route design is also valid. Queuing before acceptance, or refusing with an explicit retry contract, can be the more honest design where a second processor would not preserve the same semantics. Availability that changes financial meaning is a bad bargain.

Exercise fallback without creating money movement

A warm route needs evidence at three layers: credentials can be retrieved, transport negotiation succeeds, and the adapter understands the current request and response schema. A probe that checks only DNS or TCP can stay green while expired credentials or schema drift wait for the outage.

Use a supplier-supported non-mutating validation operation or a dedicated test environment. Do not manufacture a nominal-value payment as a heartbeat; even a small financial mutation creates reconciliation, refund, and compliance obligations. Production credentials should remain in the secret manager, scoped independently, and fetched only by the worker that needs them.

The adapter contract below makes the probe and delivery paths explicit. Tests can verify that a health check never calls Deliver, while a fallback delivery preserves the event ID.

package routing

import (
    "context"
    "errors"
)

type Event struct {
    ID      string
    Payload []byte
}

type Receipt struct {
    EventID    string
    SupplierID string
}

type Adapter interface {
    Validate(ctx context.Context) error
    Deliver(ctx context.Context, event Event) (Receipt, error)
}

func Probe(ctx context.Context, adapter Adapter) error {
    if err := adapter.Validate(ctx); err != nil {
        return errors.New("fallback validation failed")
    }
    return nil
}

func DeliverOnce(ctx context.Context, adapter Adapter, event Event) (Receipt, error) {
    if event.ID == "" {
        return Receipt{}, errors.New("event ID is required")
    }
    receipt, err := adapter.Deliver(ctx, event)
    if err != nil {
        return Receipt{}, err
    }
    if receipt.EventID != event.ID {
        return Receipt{}, errors.New("receipt event ID mismatch")
    }
    return receipt, nil
}
Enter fullscreen mode Exit fullscreen mode

The test cadence should be derived from recovery objectives and change frequency, not copied from a generic calendar. Run the probe after credential rotation, adapter deployment, policy changes, and supplier schema changes; schedule it frequently enough that its maximum undetected drift is acceptable. Record the policy version and sanitized result in the audit log.

Then rehearse refusal. A useful exercise forces the default route unavailable, gives the fallback budget two units, submits three distinct event IDs plus a duplicate, and verifies four facts: the duplicate consumes no extra reservation, two distinct events are admitted, the third is refused, and every outcome is reconcilable. That dataset is small on purpose. It finds policy errors before load testing obscures them.

Make refusal an auditable product behavior

Once the allowance is gone, the API should communicate temporary inability to accept work and should not claim success. HTTP semantics permit a 503 Service Unavailable response for temporary overload or maintenance, and the response may include Retry-After. The contract must also state that clients retry with the same event ID. Do not return a successful acceptance code and quietly drop the event.

Refusal needs its own metric and ledger. Track accepted events by route, duplicate submissions, reservations, reservation reversals, exhausted-budget refusals, ambiguous attempts, and the age of unreconciled accepted events. Alerting on the ratio alone is weak; one unreconciled high-value mutation can matter even when aggregate traffic looks normal.

Compliance scope depends on the data actually handled and the applicable regime, so the design should minimize sensitive payloads and let counsel and compliance owners set retention and access rules. PCI DSS requirements apply when systems store, process, or transmit account data within its defined scope; tokenization or outsourcing does not justify guessing that a component is out of scope. Record identifiers and decisions needed for evidence, never credentials, and avoid placing sensitive authentication data in logs.

The final release gate is therefore concrete: demonstrate idempotent duplicate handling, transactional budget admission, credential isolation, non-mutating route validation, controlled refusal, and reconciliation from event ID to ledger effect. If any link is missing, the second route is merely configured. It is not warm.

References

Top comments (0)