A prepaid API can stop while auto recharge is enabled because the recharge policy and the payment instrument are separate state. The operational choice is to treat recharge as a state machine: verify the effective account, resolve one eligible default payment method, test the funding path, reconcile the ledger, and only then restore traffic. An enabled switch is intent, not proof that money can move.
Short answer: inspect the account that actually owns the API usage, not merely the user who opened the dashboard. Confirm that this account has an active default payment method eligible for the intended charge, that the recharge rule is active for the same balance, and that the last recharge attempt reached a terminal result. Do not loop retries while those facts are unknown; duplicate attempts make attribution and reconciliation harder.
The customer-support scenario raises the stakes. During an upstream outage, queued ticket summaries, classifications, and suggested replies may arrive after the prepaid balance has crossed its threshold. A clean recovery must preserve which tenant generated each event and which funding event paid for the resulting usage. Otherwise service may return while the billing record quietly becomes indefensible.
Why can auto recharge be on while the API is stopped?
The misleading mental model is a single boolean: auto_recharge=true, therefore capacity is available. A more useful model has five gates. The request must map to the intended billing account; the policy must watch that account's prepaid balance; a default instrument must exist and be eligible; the charge attempt must succeed; and the resulting credit must be posted to the ledger before admission control sees it. A green policy flag proves only one gate.
This invariant matters: control-plane configuration does not equal data-plane capacity. Dashboards may present configuration, identity, payment, ledger, and request admission separately. Those views can be individually accurate while an operator draws the wrong conclusion from them.
I would bound the investigation before changing anything. Pick one failed customer-support event, record its immutable event ID, tenant ID, billing account ID, observed balance state, and failure timestamp, then follow that tuple through the system. This is a diagnostic procedure, not a claim about a particular provider. It prevents a common category error: repairing the payment profile of an administrator while requests are charged to a workspace, project, or organization.
Stop guessing.
The first evidence bundle should answer five questions:
- Which billing account did the rejected API request resolve to?
- Which prepaid ledger and currency did its admission check read?
- Which recharge policy was effective at that timestamp?
- Which payment method did the policy resolve as default and eligible?
- What terminal result, if any, belongs to the recharge attempt?
A missing answer is itself a finding. It means the recovery path lacks observability, even if an operator can eventually click through the problem.
Follow one event through identity, payment, and ledger state
Start from the stopped request because it represents the data-plane decision. Work backward from its billing account identifier. Starting from a person's email or a browser session is weaker: membership can span several accounts, and administrative identity does not establish charge ownership.
For customer support, keep tenant attribution on every durable event before any paid processing begins. A replay after an outage should use the same event ID and tenant ID, while each processing attempt gets a distinct attempt ID. That separation lets the platform deduplicate business work without erasing operational history.
The payment-method check should be exact. Present is insufficient. Determine whether one instrument is attached to the effective billing account, selected as its default for this charge type, currently eligible, and compatible with the account's currency and payment flow. Do not print full instrument details into logs. The OWASP Secrets Management Cheat Sheet recommends limiting access to secrets, rotating them, and recording relevant lifecycle events; payment credentials and provider tokens deserve the same disciplined handling.
Then inspect the ledger boundary. A successful external charge and an available prepaid credit are different events. Recovery automation should wait for a posted ledger entry, or for the platform's documented equivalent, rather than inferring credit from an initiated or pending charge. This distinction is where SLO language helps: measure time from threshold crossing to spendable credit, and separately measure time from spendable credit to admitted request. Combining them hides the failing gate.
Use a small correlation record rather than free-form log text. This Go type deliberately excludes raw credentials and instrument numbers:
type RechargeTrace struct {
EventID string
TenantID string
BillingAccountID string
LedgerID string
PolicyID string
PaymentMethodRef string
RechargeAttempt string
ObservedAt time.Time
}
PaymentMethodRef should be an opaque internal reference with restricted access. It is useful for correlation; it is not permission to copy sensitive payment data into every telemetry sink.
Make the preventative path idempotent
The preventative code path should reject ambiguous configuration early, serialize recharge decisions per billing account, and use a stable idempotency key for a logical recharge attempt. It should also refuse to resume queued work until the ledger reports spendable credit. These constraints trade a little recovery latency for attribution accuracy, which is the right direction when invoices must survive an audit.
package billing
import (
"context"
"errors"
"fmt"
)
type Snapshot struct {
AccountID string
LedgerID string
PolicyEnabled bool
DefaultMethodID string
Spendable bool
}
type Store interface {
Snapshot(context.Context, string) (Snapshot, error)
BeginRecharge(context.Context, string, string) (string, error)
CreditPosted(context.Context, string) (bool, error)
}
func EnsureCapacity(ctx context.Context, store Store, accountID, eventID string) error {
s, err := store.Snapshot(ctx, accountID)
if err != nil {
return fmt.Errorf("read billing snapshot: %w", err)
}
if s.AccountID != accountID || s.LedgerID == "" {
return errors.New("billing attribution is incomplete")
}
if s.Spendable {
return nil
}
if !s.PolicyEnabled {
return errors.New("recharge policy is disabled")
}
if s.DefaultMethodID == "" {
return errors.New("default payment method is missing")
}
key := fmt.Sprintf("recharge:%s:%s", accountID, eventID)
attemptID, err := store.BeginRecharge(ctx, accountID, key)
if err != nil {
return fmt.Errorf("begin recharge: %w", err)
}
posted, err := store.CreditPosted(ctx, attemptID)
if err != nil {
return fmt.Errorf("verify ledger credit: %w", err)
}
if !posted {
return errors.New("credit is not spendable yet")
}
return nil
}
In a real service, account-level serialization can be a database transaction, a compare-and-swap transition, or a lease with a carefully defined failure model. The mechanism matters less than the invariant: two workers observing the same low balance must not create two logically independent recharges. The stable key should be retained long enough to cover delayed retries and outage replay.
Do not use the customer event ID alone as the key across all tenants. Event namespaces collide. Account ID plus event ID is safer, and a ledger-issued attempt ID should connect the funding result to the accounting entry.
Buy, build, or split the responsibility
The decision is an ownership allocation across payment execution, prepaid accounting, admission control, and incident response. A buy-versus-build table makes the hidden on-call cost harder to ignore.
| Approach | Attribution control | On-call burden | Lock-in boundary | Best fit |
|---|---|---|---|---|
| Managed billing and ledger | Constrained by exported identifiers and records | Lower for payment execution; integration incidents remain yours | Account, ledger, and export semantics | Teams that can reconcile external records to internal tenants |
| Self-hosted ledger with managed payment execution | Internal tenant mapping stays local | Higher for ledger correctness and replay | Payment API plus internal schema | Teams with strict attribution or custom credit rules |
| Fully self-managed payment and ledger stack | Maximum schema control | Highest operational and compliance surface | Mostly internal interfaces | Organizations already staffed for payment operations |
No row is universally safer. A managed system reduces machinery the team must operate, but only if its account model and export trail preserve the attribution keys required by finance and support. A self-hosted ledger increases control while adding backups, migrations, consistency testing, and a larger pager surface. Capacity planning must include operators, not just requests per second.
Before choosing, run an outage exercise with delayed funding confirmation, duplicated event delivery, a missing default method, and 10,000 queued support events spread across multiple tenants. Ten thousand is a test workload, not a performance claim. The pass condition is exact: each business event runs at most once, every processing attempt remains visible, every unit of usage maps to one tenant, and resumption never precedes spendable credit.
Recovery order and the limits of this advice
During an incident, freeze automatic retry amplification first. Preserve the queue, but pause consumers that would repeatedly hit admission control or initiate more funding attempts. Next, identify the effective billing account from a rejected request, establish the default method's eligibility without exposing sensitive values, and inspect the latest recharge attempt. After the ledger shows spendable credit, resume a small cohort and watch both admission success and tenant-level accounting before opening the full queue.
Define operational targets in advance. Useful indicators include the percentage of requests with complete billing attribution, duplicate recharge attempts per logical key, age of the oldest queued support event, and recharge-to-credit latency. Alert on broken invariants rather than a dashboard toggle. A policy marked enabled is poor paging evidence; a low-balance account with no eligible default instrument is actionable.
This approach has explicit limitations and trade-offs. It does not apply unchanged to postpaid accounts, invoice terms, grants, or systems where request admission is intentionally independent of balance. Waiting for posted credit improves accounting certainty but lengthens recovery; resuming on a pending charge lowers delay but accepts a reconciliation risk. The procedure also cannot diagnose a processor-side rejection without the processor's documented result and trace identifier. In those cases, preserve the same attribution chain but follow the actual funding contract instead of forcing a prepaid model onto it.
The lasting fix is modest: make account resolution explicit, validate the default payment method before the threshold is crossed, use idempotent recharge attempts, and wait for ledger truth. Restore service from verified state, not configuration intent.
Top comments (1)
Treating an enabled auto-recharge switch as intent, not proof that money can move, is the right prepaid-API boundary. Looping a fresh charge while the last recharge outcome is still unknown is how duplicate funding attempts open and attribution breaks.
When the last recharge attempt timed out or returned unknown, do you refuse a second charge until reconcile returns settled, failed, or a hard unknown a human must clear?