DEV Community

Cover image for Handling a Cold-Starting Android Backend
Raylabs
Raylabs

Posted on Originally published at raylabs.app

Handling a Cold-Starting Android Backend

When an application backend sleeps between requests or takes time to start, an Android app can mistake a temporary cold start for a permanent outage. The opposite failure is also common: a loading spinner retries forever and gives the user no clear way forward. How should an Android app handle an unavailable or cold-starting backend without showing an infinite startup spinner? For a related implementation, see Viewmodel Partial Success Refresh Failure.

The solution is to treat startup health checks as a finite state machine rather than an endless loop. By performing a single pass across configured endpoint candidates during launch, the application can quickly determine if the service is reachable. If the backend fails to respond, the app moves to an explicit unavailable state instead of locking the user into a permanent loading loop. When the user chooses to try again, an explicit retry action refreshes the remote configuration and checks again on a fixed interval, bounded by a visible deadline.

The Startup State Machine and Endpoint Checks

A robust mobile architecture separates initial startup discovery from active user sessions. During the first launch phase, the ViewModel checks configured base URLs once and selects the first healthy endpoint. If none of the candidates respond within the initial pass, the UI exposes a distinct unavailable state. For a related implementation, see To Robust Viewmodel Unit Tests.

This design prevents wasted network traffic and gives the user immediate clarity. Instead of guessing whether the application is frozen or broken, the user sees that the backend is currently offline. The application can distinguish between ordinary unavailability, maintenance modes, and configuration mismatches by evaluating response codes and payload flags.

Bounded Retry with Cancellation

When a cold-starting backend requires a moment to wake up, a manual retry mechanism gives users control without risking runaway resource consumption. The implementation uses a fixed two-second interval and a 75-second total deadline. These specific values correspond to the backend cold-start profile and are not universal defaults.

To prevent overlapping network calls, every new check cancels the previous coroutine job. This cancellation guarantee ensures that if a user triggers multiple retries or if configuration changes mid-check, obsolete requests do not continue running in the background.

class BackendRetryViewModel : ViewModel() {
    private var retryJob: Job? = null

    fun startBoundedRetry() {
        retryJob?.cancel()
        retryJob = viewModelScope.launch {
            val startTime = System.currentTimeMillis()
            while (System.currentTimeMillis() - startTime < 75_000) {
                if (checkHealth()) {
                    _uiState.value = UiState.Ready
                    return@launch
                }
                delay(2_000)
            }
            _uiState.value = UiState.Unavailable
        }
    }

    private suspend fun checkHealth(): Boolean {
        // Perform lightweight health endpoint check
        return false
    }
}
Enter fullscreen mode Exit fullscreen mode

Testing Failure and Recovery

Verifying timeout and recovery logic requires deterministic test execution. Unit tests must cover the initial healthy endpoint selection, the transition to the unavailable state after all candidates fail, the forced configuration refresh on manual retry, and the successful recovery when a later attempt succeeds. Using virtual time control in coroutine tests allows developers to simulate the 75-second deadline and the two-second intervals instantly without slowing down the test suite.

Adapting Retry Strategies for Production

While a fixed interval and a manual trigger work well for predictable cold starts, production systems often require more sophisticated patterns. Developers should consider exponential backoff, jitter, and provider Retry-After headers when dealing with public APIs or heavily loaded microservices. Evaluating connectivity classification helps determine whether a retry should be attempted at all when the device is offline or on a metered connection.

Conclusion for Resilient Startup Flows

Designing a resilient mobile startup flow means balancing persistence with user experience. By implementing a finite health-check pass, offering a bounded manual retry, and cancelling superseded requests, you protect both the client application and the backend service from unnecessary strain. Grounding your retry intervals in actual service behavior rather than guesswork ensures a predictable experience when services wake up.

Top comments (0)