DEV Community

Erfan Moghaddam
Erfan Moghaddam

Posted on

Designing Safer API Failover in an Android App

An Android app can depend on more than one API origin.
A primary host may be temporarily unavailable while a secondary host can serve the same API. Switching hosts sounds like a URL replacement, but the difficult questions come afterward: Which requests may be repeated? What if several requests fail at once? When should the app try the primary again?

I encountered these questions while building a Kotlin Android client with Retrofit and OkHttp. This article walks through the design decisions, including a retry rule that deserves more care than it first appears to need. The examples below are simplified illustrations; they are not a copy of the application's private implementation.

Define the scope before switching anything

The client keeps a configured primary origin and an optional secondary origin. An origin consists of the URL scheme, host, and port. The failover layer rewrites only requests whose origins match one of those configured API origins. Calls to unrelated hosts must pass through unchanged.

When rewriting a request, the client replaces only the scheme, host, and port. It retains the path and query parameters. This matters when the API includes versioned paths, search parameters, or signed request data. Any server side signatures tied to the host would require a separate design.

The two origins must serve the same API contract and operate on compatible data. If the secondary has different authentication rules, stale state, or a different TLS identity, switching hosts cannot repair the request.

Keep selection state in one place

The application holds one selected origin for the current session. At startup it selects the primary. A qualifying failure on the primary selects the secondary, if configured. A qualifying failure on the secondary moves the manager into an unavailable state. After a short cooldown, the next request may try the primary again.

Primary selected
  ├─ qualifying failure + secondary configured → Secondary selected
  └─ qualifying failure + no secondary         → Unavailable

Secondary selected
  └─ qualifying failure                         → Unavailable

Unavailable
  └─ cooldown elapsed + new request             → Primary selected
Enter fullscreen mode Exit fullscreen mode

The state transition is protected by a lock so simultaneous failures do not independently advance the state. A failed request identifies the origin it actually used. If another request already changed the selected origin, the manager does not treat the old failure as a failure of the newly selected origin.

For a cooldown on Android, SystemClock.elapsedRealtime() is appropriate because it measures elapsed time monotonically and continues across device sleep. Wall clock time can change when the user or network adjusts the date and time. Android's SystemClock reference explains the distinction.

The cooldown is a retry gate, not a health check. It does not establish that either host is healthy. A production system may also need longer backoff, jitter, connectivity awareness, and telemetry. The chosen interval should reflect the service and user experience; there is no universal five second value.

Put the decision at the HTTP boundary

An OkHttp application interceptor can examine each request, rewrite managed API URLs, and decide whether a failed call is eligible for a single retry against the alternative origin. It should leave cancellation alone and close an earlier response before calling proceed again with a replacement request.

The policy is more important than the URL change. For example, a GET that receives a gateway 502, 503, or 504 can be a candidate for failover if the two origins are equivalent. An authentication 401, validation 400, or missing resource 404 does not normally indicate that another origin should receive the same call. The exact mapping still depends on the API contract.

private fun mayReplayAutomatically(method: String): Boolean {
    return method.equals("GET", ignoreCase = true) ||
        method.equals("HEAD", ignoreCase = true) ||
        method.equals("OPTIONS", ignoreCase = true)
}

private fun isCandidateGatewayFailure(code: Int): Boolean {
    return code == 502 || code == 503 || code == 504
}
Enter fullscreen mode Exit fullscreen mode

This is deliberately a narrow example, not a universal definition of idempotency. HTTP defines PUT and DELETE as idempotent methods too, but applications still need to check whether replay is suitable for their endpoints and request bodies. OPTIONS is safe by HTTP semantics, although many mobile APIs never use it.

The subtle risk: repeating a write

A timeout or broken connection does not always tell the client whether the server applied a write. If an app automatically repeats a payment or reservation POST on a different host, it could create the operation twice. Even some errors that appear to occur before an HTTP response are a poor basis for a broad, app wide promise that every write is safe to replay.

RFC 9110 says a client should not automatically retry a non-idempotent request unless it knows the operation is effectively idempotent or can determine that the original request was never applied. HTTP Semantics, Section 9.2.2 is the relevant reference.

The existing client implementation that motivated this article allows failover for GET, HEAD, and OPTIONS on selected network errors and 502–504 responses. It also permits some connection exception classes to trigger failover for other methods. That latter rule needs an endpoint by endpoint review before anyone can claim it prevents duplicate writes. A conservative default is to replay read requests automatically and surface ambiguous write failures to the caller. If a write must be retried automatically, coordinate an idempotency key or equivalent deduplication rule with the server and confirm that both origins share its state.

Test the transition, not just the helper methods

Before shipping a failover policy, exercise it with a controlled test server and the actual OkHttp client configuration. The most useful cases are behavioral:

Situation Expected behavior
Primary serves a successful read Keep using primary
Primary fails during an eligible read Try secondary once
Primary returns a validation or authentication error Return the response without switching
Secondary fails after selection Enter unavailable state
New request arrives during cooldown Avoid repeated immediate attempts
Several primary requests fail concurrently Do not accidentally mark the new secondary as failed because of an old primary error
Call is canceled Do not launch a fallback
Write has an ambiguous transport failure Do not replay without an explicit endpoint policy
Request targets an unrelated host Leave it untouched

Also inspect redirects, request bodies that cannot be sent twice, authorization headers, TLS certificates, and the interaction with OkHttp's own connection recovery. A failover interceptor is one layer of availability, not a substitute for server side consistency or a well defined retry contract.

What I took away

The reliable part of failover is the boundary: keep origin selection centralized, guard concurrent transitions, limit rewriting to your own API, and make replay eligibility explicit. Changing the host is easy. Knowing whether a call can safely happen again is the design problem.

Top comments (0)