Part 4 of a 5-part series on automating a multi-app Android release pipeline. Part 3 → covered the storage-quota wall and the move to a self-hosted runner. This part covers the cleanup that followed — including a bug I introduced while fixing the exact class of bug it eventually became — and the first real production release after all of it.
TL;DR
Fixing the storage-quota problem meant deleting the GitHub-hosted copy of each release AAB right after it publishes to the Play Store. That fix introduced its own bug: the delete step's upload half wasn't allowed to fail safely, so a hiccup in a step that existed purely for debugging convenience could block the actual Play Store publish — the exact kind of failure the redesign was supposed to eliminate. Around the same bug class, a Discord notification system built on relaying results between jobs through an uploaded artifact was found to be silently reporting genuinely successful builds as failed. Both got fixed the same way: stop routing something through an intermediate hop that doesn't need to exist. Then, on the first real production dispatch after all of it, the pipeline behaved exactly right — and hit one last honest, one-time gotcha that had nothing to do with any of this code.
Killing the GitHub Copy for Good
Part 3 ended on retention numbers that had never been sized deliberately — 30 days on a 70MB file that exists purely so a human can grab it for debugging, when the actual deliverable already lives in the Play Store the moment the job succeeds. The real fix wasn't a shorter number. It was: delete it as soon as it's no longer needed, and only fall back to a short retention window as a safety net in case that delete step itself doesn't run.
permissions:
actions: write # needed to delete the AAB artifact once Play publish succeeds
- name: Upload AAB artifact to GitHub
id: upload_aab
uses: actions/upload-artifact@v4
with:
name: "release-aab-${{ matrix.route }}-${{ github.sha }}"
path: local-builds/android/**/*.aab
retention-days: 7 # fallback only — the delete step below is what actually keeps this clean
# ... Publish to Google Play Console happens here ...
- name: Delete GitHub AAB artifact now that Play has it
if: success()
run: |
gh api -X DELETE "repos/${{ github.repository }}/actions/artifacts/${{ steps.upload_aab.outputs.artifact-id }}" || true
The file now lives on GitHub for the length of one job run, not a month. Multiply that across every release, both apps, and the storage math stops being a slow-motion problem entirely.
The Bug I Built Into My Own Fix
Here's the part worth actually admitting to, because it's a more useful lesson than the fix itself: adding that upload-then-delete sequence introduced a new failure mode that hadn't existed before.
The "Upload AAB artifact to GitHub" step, and a smaller step that uploads version-mismatch diagnostics, both sat before "Publish to Google Play Console" in the job. Neither had continue-on-error: true. Which meant: if either of those backup/debugging uploads failed for any reason — a transient GitHub API hiccup, exactly the kind of thing you'd expect more of while an org is already mid-incident on Actions storage — the job would stop right there. The real Play Store publish, the actual point of the entire job, would never run. A step that exists purely as a debugging convenience had the power to block the one thing that actually mattered, and it was found live, on a real release, in the middle of diagnosing the very quota incident this design was supposed to be helping with.
The fix is one line, on each of those two steps:
- name: Upload AAB artifact to GitHub
continue-on-error: true
uses: actions/upload-artifact@v4
with: # ...
If the backup copy fails to upload, that's logged and the job moves on — the Play Store publish, the part with actual consequences, is no longer hostage to a convenience feature. The lesson isn't "always add continue-on-error" — plenty of steps should block a job. It's: when you add a step whose entire purpose is "nice to have for debugging," explicitly decide whether its failure should be allowed to block the step that actually matters. Skipping that question is easy to do while you're focused on the fix in front of you, which is exactly how this one got in.
The Fix for the Fix Was Also Wrong, at First
That two-step fix shipped, and PR-check runs on the branch kept failing anyway. The honest first instinct was to assume the fix just hadn't taken effect yet. The actual pushback that broke that assumption was blunt: the fix, as shipped, wasn't working — go find out why instead of re-asserting that it should be.
Going back to actually verify, rather than re-checking the same two steps again, turned up a third step doing the exact same unprotected thing: a small "Upload release result" step, further down the same job, that also had no continue-on-error. The two steps fixed first only ever run during a real workflow_dispatch release — they're structurally silent on a plain PR check. This third step runs on if: always(), no gate at all, on every single event including ordinary PR checks. It had been the actual, sole cause of every PR-check failure all week — a different step than either of the two the first fix had targeted, doing the same unprotected pattern in a place that mattered more, not less: on a real release run, it also sits after "Publish to Google Play Console," meaning a real, successful, already-live production release could still get reported as "failure" over nothing but a bookkeeping upload hiccup.
The corrected fix protected all three steps, and — since continue-on-error quietly hides a failed step behind a green checkmark unless something says otherwise — added an explicit warning line to both the job summary and the Discord message whenever a backup step fails, so a silenced failure doesn't just become invisible instead of blocking. The real lesson here isn't the YAML. It's that "I already fixed this" is a claim, not a verification — and the fastest way to find the actual bug was someone else refusing to accept the first answer at face value.
The Notification System That Was Lying to Us, Quietly
A related design flaw turned up in the same investigation, in the Discord build-status notifications. The original design had one job build each app, then a separate "notify" job download an artifact containing each build's result and post a combined message. It looked clean on a diagram. In practice:
- A run where an app genuinely built and shipped successfully — confirmed directly in the step logs, with a working download link — still posted a failure message to the team, because the small bookkeeping file recording that success hadn't uploaded cleanly. The build didn't fail. The relay did, and the relay's failure got reported as the build's failure.
- A mixed result (one app ships, one doesn't) sometimes lost one app's outcome entirely in the hop between jobs, depending on exactly which artifact made it through.
Both bugs point at the same root design problem: routing a result through an intermediate artifact upload/download, between two separate jobs, adds a failure mode that has nothing to do with whether the actual build succeeded. The fix wasn't a smarter relay — it was removing the relay:
# Before: one job builds, a separate "notify" job downloads an artifact
# and posts one combined Discord message for both apps.
#
# After: each build job posts its own Discord message directly, using
# only data it already has locally. No artifact upload, no download,
# no cross-job relay to fail independently of the thing it's reporting on.
Each app now reports its own real outcome, straight from the job that actually knows what happened, with nothing in between that can fail on its own and get misread as the build's fault.
The First Real Production Ship
After the runner migration, the artifact cleanup, the continue-on-error fix, and the notification redesign, the pipeline got its first real test: an actual production dispatch, for real, to ship both apps' first production release.
The build ran clean. The AAB uploaded successfully. The Play publish step reached Google — and then failed, on committing the release, with a message that looked, at first glance, like it might be another chapter of the exact same saga:
Error: Release in track targeting no countries
It wasn't. This time, the pipeline had done everything right. Google Play requires every production track to have a list of countries/regions configured before a release can go out on it — and because neither app had ever had a production release before, that list had simply never been set. It's a one-time Play Console setup step, not a code problem, not a pipeline bug, not a repeat of the wrong-track issue from Part 2 (that one was "publishing to the wrong, unused track"; this one was "the right track, correctly targeted, just never configured for its very first use"). The fix lived entirely in Play Console — add the countries, save, re-dispatch. Nothing about the workflow needed to change.
Where Part 4 Leaves Off
Everything else in this series, up to this point, was the pipeline lying to us or breaking silently — a version number that drifted without telling anyone, an error message that pointed at the wrong track, a workflow file that failed to parse with zero errors shown, a notification system that reported success as failure. This last one was the opposite: the pipeline told the truth, plainly, on the first real attempt that mattered — and the one thing left to fix wasn't in the code at all.
It felt, at this point, like the story was over. It wasn't quite. A few hours after that first real release went out, a message came in that had nothing to do with signing, tracks, storage, or notifications — both apps were showing the same login screen.
Next: Part 5 → — the bug that shipped with the first real release, and a comment that was true the day it was written and silently stopped being true nine minutes later.
I write about the debugging journeys nobody puts in the docs — more at cycy.is-a.dev.
Top comments (0)