A network timeout is not an order rejection. A process restart is not proof that an account is flat. These two distinctions make recovery one of the more interesting engineering problems in an on-device trading app.
The useful question is not just whether the app can reconnect. It is whether it can establish what actually happened before taking another action.
This article describes a recovery design and review checklist for Android software connected to an exchange API. The examples are deliberately generic: they are not executable trading instructions, a published implementation, or a claim that every API offers the same guarantees.
The failure happens between two systems
Consider a request to create an order:
- The app records the intended action.
- The request reaches the exchange.
- The exchange accepts it.
- The response never reaches the app.
The phone sees a timeout. The exchange may already have an order.
Repeating the action with a new identity can create an additional order. Treating the timeout as a definitive failure can also leave the interface contradicting the exchange.
Keep three concepts separate: intent, acknowledgment, and observed exchange state. An intent explains what the app attempted. An acknowledgment records the response received. A later account read provides evidence about what exists now. None should silently stand in for the others.
Coinbase's Advanced Trade Create Order documentation specifies a unique client_order_id and describes the behavior when an existing identifier is reused. That contract is worth reading before implementing retries. It does not make every endpoint, cancellation, or protection update automatically idempotent.
Recovery should be a gate, not a screen refresh
A practical state model separates reconnecting from being ready to act:
Interrupted
-> Read account and outstanding orders
-> Reconcile unresolved actions
-> Check protection and market freshness
-> Ready, or paused with an actionable reason
This is a conceptual model, not a drop-in state machine. The implementation still needs to account for pagination, delayed updates, partial fills, conflicting reads, and requests that fail during recovery itself.
Do not infer that a position disappeared simply because one request failed. Do not infer that protection is confirmed merely because a submission was attempted. Equally, do not turn each temporary reconfirmation into a new alarming activity event if the underlying account state is unchanged.
A good recovery gate admits the difference between "not yet established" and "established to be missing." That difference matters both to execution and to the messages a user sees.
Two callers can be a bigger problem than one failed call
The service, a reconnect callback, a refresh, and an explicit user action can all arrive close together. If they independently decide to repair the same unresolved action, recovery becomes a concurrency problem.
Give execution decisions a single owner. Readers can update the interface, but they should not each create competing repair workflows. A cancellation or a paused session should also invalidate work that belonged to the earlier active session; a late response should not restart it accidentally.
Keep network waits and reconciliation off the main thread. The visible interface should remain usable while execution waits for trustworthy evidence.
Android background work is not an uptime promise
Android's foreground service documentation describes user-visible ongoing work and its notification requirements. A foreground service is not a promise that power, networking, credentials, or the process will remain available indefinitely.
Design recovery for loss of those conditions. Treat opening the app, starting execution, and resuming after an interruption as distinct events rather than assuming that a visible screen means automation is already ready.
Device-side checks also do not replace exchange-side protection. When a platform supports hosted protection, an app should distinguish the requested configuration from the configuration it has actually confirmed at the exchange.
Protect secrets without obscuring the diagnosis
Android Keystore can keep cryptographic key material non-exportable. That is different from claiming that decrypted API credentials can never exist in application memory or that every device has the same hardware protection.
Diagnostic records should preserve useful action identities, timestamps, and outcomes without including API secrets, authorization headers, or raw credential imports. A support export needs enough evidence to explain an unresolved action, not enough access to reproduce a customer's account privileges.
Test the uncomfortable boundaries
Happy-path reconnect tests are not enough. A small fault matrix is more revealing:
| Interrupted moment | Evidence to establish before another action |
|---|---|
| Request sent; response lost | Whether the intended order exists under its original identity |
| Fill occurs before local refresh | Current position and remaining order quantity |
| Protection update is unresolved | Actual protection state, not just the attempted update |
| User pauses during recovery | No continuation of work owned by the earlier active session |
| Process restarts with cached balances | Fresh account state before sizing a new action |
| Market stream reconnects | Quote freshness before evaluating an executable price |
In each case, inspect both behavior and reporting. A test can prevent a duplicate action while still leaving a misleading "ready" status or a noisy activity history.
Simulated trading is useful for strategy and accounting checks, but a local simulation alone cannot prove how a live exchange handles an ambiguous request. Recovery tests need controlled failures at the API boundary as well.
The invariant is more useful than the retry count
"Retry three times" is a policy. "Never create a second independent action while the first remains unresolved" is an invariant.
An invariant makes reviews and tests concrete. It identifies what must remain true during timeouts, restarts, delayed responses, and user cancellation. The retry schedule can then be chosen around that rule instead of becoming a substitute for it.
For an Android trading app, reconnecting is only the beginning. Reconciliation is what makes the next action defensible.
Top comments (0)