Your catch (e: Exception) Block Is Quietly Breaking Coroutine Cancellation
You've probably written this exact pattern to make a network call "resilient":
suspend fun loadProfile(userId: String): Profile {
var attempt = 0
while (true) {
try {
return api.fetchProfile(userId)
} catch (e: Exception) {
attempt++
if (attempt >= 3) throw e
delay(1000L * attempt)
}
}
}
Looks reasonable. Retries on failure, backs off, gives up after three tries. It'll pass
every unit test you throw at it. It will also silently break structured concurrency the
moment the coroutine calling it gets cancelled — and nothing in this function will ever
tell you that happened.
Why catch (e: Exception) catches cancellation too
CancellationException is a subclass of Exception in Kotlin. When a coroutine's Job
gets cancelled — the screen is destroyed and viewModelScope cancels, a parent
coroutineScope fails and cancels its siblings, a timeout fires — the suspension point
the coroutine is sitting at (here, delay() or the actual api.fetchProfile suspend
call) throws a CancellationException to unwind it.
A catch (e: Exception) block doesn't know the difference between "the server returned
a 500" and "this coroutine was told to stop existing." It catches both. In the code
above, a cancellation during delay(1000L * attempt) gets treated as attempt-failed,
and the loop just... retries again. The coroutine keeps running network calls on a
Job that's already cancelled, for as long as the retry budget allows — a small,
easy-to-miss cancellation leak that shows up in production as "why is this screen still
making requests after the user navigated away" and almost never as a crash you can
grep for.
Structured concurrency assumes cancellation actually propagates
The entire value proposition of structured concurrency — a parent Job cancelling all
its children, a coroutineScope waiting for every child before returning, cancellation
composing cleanly through suspend function calls — depends on CancellationException
reaching the top of the coroutine and actually finishing the job. Swallow it partway up
the call stack and you've quietly opted this one coroutine out of the entire cancellation
model, without changing a single type signature. Nothing about the function's shape
tells a caller this happened.
It gets worse in a coroutineScope { } with multiple children: if one child's exception
is supposed to cancel its siblings, and a sibling's catch block is eating the resulting
CancellationException instead of rethrowing it, that sibling keeps running past the
point where the whole scope should have unwound — defeating the "fail one, cancel all"
guarantee that's the actual reason to use coroutineScope instead of supervisorScope
in the first place.
The fix
Check for cancellation explicitly, and rethrow it before your generic handling runs:
suspend fun loadProfile(userId: String): Profile {
var attempt = 0
while (true) {
try {
return api.fetchProfile(userId)
} catch (e: CancellationException) {
throw e // never swallow this — let cancellation propagate
} catch (e: Exception) {
attempt++
if (attempt >= 3) throw e
delay(1000L * attempt)
}
}
}
Order matters: the CancellationException catch has to come before the broader
Exception catch, or it never gets reached. This is easy to get backwards if you add
the cancellation handling later as an afterthought rather than up front.
If you're using runCatching instead of try/catch — common in repository layers that
want a Result<T> return type — you have the same bug, less visibly:
// Buggy: runCatching wraps CancellationException into a failed Result
suspend fun loadProfileResult(userId: String): Result<Profile> =
runCatching { api.fetchProfile(userId) }
// Fixed: rethrow cancellation, let everything else become a Result
suspend fun loadProfileResult(userId: String): Result<Profile> =
try {
Result.success(api.fetchProfile(userId))
} catch (e: CancellationException) {
throw e
} catch (e: Exception) {
Result.failure(e)
}
runCatching's own implementation catches Throwable, not just Exception — so this
version of the bug is arguably worse, and it's the one that shows up constantly in
"clean architecture" repository layers that standardized on Result<T> for every
suspend function.
Writing a test that actually catches this
The scary part of this bug is that it's invisible to a test that doesn't specifically
check for it — loadProfile returning the right value on success and throwing on
repeated failure both look fine. You need a test that cancels the calling coroutine
mid-flight and asserts the function actually stops:
@Test
fun `loadProfile does not swallow cancellation`() = runTest {
val fakeApi = FakeProfileApi(neverResolves = true)
val job = launch {
loadProfile(fakeApi, "user-1")
}
advanceTimeBy(100)
job.cancel()
job.join()
assertTrue(job.isCancelled)
assertEquals(0, fakeApi.callCountAfter(job.cancel()))
// If the buggy version is under test, this either hangs (TestCoroutineScheduler
// never idles because the retry loop keeps scheduling more delay()s) or the
// call count keeps climbing after cancel() — either is the leak, caught.
}
With kotlinx-coroutines-test's runTest, a coroutine that keeps rescheduling work
after being told to cancel will either fail the test's idle-detection or keep
incrementing a call counter you can assert against — either way, this is the test that
would have actually caught the retry-loop bug above, where "does it return the right
value" tests wouldn't.
The pattern to watch for in review
Any catch (e: Exception), catch (e: Throwable), or runCatching inside a suspend
function is worth a second look — not because catching broadly is always wrong, but
because it's the one shape of bug that every other kind of testing (unit tests on
success/failure paths, manual QA, crash reporting) is structurally blind to. The
function keeps working correctly for every input except "the coroutine was cancelled
while this was running," which is exactly the case a fast reviewer skims past.
This is one specific case of a broader category — cooperative cancellation leaks,
structured concurrency guarantees, and the Job-hierarchy internals that explain why
this works the way it does — that comes up constantly in senior Android interviews and,
more importantly, in production. I went deep on this and 54 other coroutines/Kotlin
questions in Modern Kotlin & Concurrency: Ultimate Interview Prep
Blueprint if you want the fuller picture,
including Job state transitions and CoroutineExceptionHandler interactions this
post didn't have room for.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.