Consider a team running Playwright on GitHub Actions with ten pull requests a day, each split into four shards for faster feedback. Every shard spins up its own runner, and each one downloads the same browser binaries from scratch. That adds up to forty cold installs a day, roughly sixty to ninety seconds apiece, or forty to sixty minutes of download time before a single test result comes back.
GitHub Actions pricing has also become cheaper and easier to audit. Hosted runner rates for private repositories fell by up to 39 percent on January 1, 2026. A proposed charge for self-hosted runners was postponed indefinitely after community pushback, so self-hosted usage remains free.
Even with lower rates, the same costs show up repeatedly in Playwright pipelines. Browser binaries get downloaded unnecessarily, full suites run against unchanged code, and flaky tests consume extra compute time. Reducing CI spend usually comes down to using those minutes more efficiently, a process covered in this GitHub Actions setup guide. The examples in this article focus on private repositories running on GitHub-hosted Linux runners, since public repositories remain free.
Understand What You're Actually Paying For
Most teams checking a GitHub Actions bill see one number, though the real cost comes from several factors that rarely get tracked together. Billing mechanics, runner type, storage, and flakiness each contribute to that number, and missing any one of them makes it harder to understand what a Playwright pipeline is actually costing.
How GitHub Bills Playwright Jobs
GitHub rounds every job up to the nearest whole minute, and each shard in a matrix bills as its own job. Four shards each running for 8 minutes 10 seconds round up to nine billed minutes apiece, for a combined total of 36 minutes. That per-job rounding adds up as shard counts increase, since each additional shard is another chance to lose up to 59 seconds to the round-up.
The macOS Multiplier
Ubuntu runners bill at $0.006 per minute and macOS at $0.062, close to ten times as much for the same runtime. Wherever macOS appears, it is usually one of the most expensive parts of the bill. Few Playwright suites need it on every pull request. Restricting it to scheduled jobs or release branches preserves coverage while reducing spend on routine pushes.
The Silent Cost of Artifact Storage
Default artifact retention lasts 90 days, with storage billed at $0.25 per GB each month. A team generating 5 GB of traces, videos, and screenshots a day accumulates roughly 450 GB under that 90-day window. Reducing retention-days to 7 days lowers that figure to about 35 GB, though actual numbers depend on how much data a pipeline produces.
Flakiness as a Cost Multiplier
A flaky test that forces a job rerun costs more than the retry time itself, because checkout, dependency installation, and browser setup run again for the whole shard, not just the test that failed. In a four-shard matrix, that means one flaky test can add a full extra billed job on top of the four already completed, since only the shard containing that test needs to start over.
Playwright's own retry setting works differently. It replays only the failed test inside the same job, without repeating checkout or install, which is why it costs far less than a job rerun. Playwright's --last-failed flag, covered in this breakdown of re-running only failed tests, restricts execution to the tests that actually failed, instead of the whole shard. In sharded CI, pair it with --last-failed-file to keep a per-shard record of the last run (Playwright 1.61 or later).
Billing mechanics, runner costs, storage, and flakiness account for most of a Playwright CI bill. The next six fixes focus on those areas: caching browser binaries, scoping triggers, setting job timeouts, sharding appropriately, treating flakiness as a billing concern, and tightening artifact retention. Browser caching comes first because it is one of the most common sources of avoidable CI spend.
Fix 1: Cache Playwright Browser Binaries
Most Playwright pipelines never cache browser binaries. The case for fixing that comes from what skipping it costs.
What Browser Install Actually Costs Per Run
npx playwright install --with-deps downloads fresh browser binaries on every job. Measurements put that download at 430 to 450 MB, with cold installs taking sixty to ninety seconds. Sharding multiplies that cost because each shard runs on its own clean runner.
Caching the Right Path With the Right Key
Caching stops that repeat download by storing binaries at Playwright's default Linux install path, ~/.cache/ms-playwright, through actions/cache@v4. PLAYWRIGHT_BROWSERS_PATH is the setting Playwright itself uses to decide where it installs and looks for browsers, so if it's set elsewhere in the workflow, the cache path needs to match that exact directory.
A cache hit only needs install-deps, since install --with-deps normally bundles that step with a browser download the cache has already covered. GitHub-hosted runners have passwordless sudo configured already, so that step runs unattended there. A self-hosted runner needs that same setup in place, or the step hangs waiting on a password prompt that never comes. A cache miss runs the full install and rebuilds the cache.
The key should track the Playwright version rather than the lockfile hash, since an ESLint or Tailwind update would invalidate it unnecessarily, a practice Playwright's own CI guidance also recommends. The path below is Linux only, scoped to ubuntu-latest.
# Assumes actions/checkout and a dependency install step (npm ci) already ran
- name: Get Playwright version
id: pw-version
# Swap '@playwright/test' for 'playwright' if using the standalone package
run: |
PW_VERSION=$(node -p "require('@playwright/test/package.json').version")
echo "version=$PW_VERSION" >> "$GITHUB_OUTPUT"
- name: Cache Playwright browsers
uses: actions/cache@v4
id: playwright-cache
with:
# Linux path only, scoped to ubuntu-latest
path: ~/.cache/ms-playwright
key: playwright-browsers-${{ runner.os }}-${{ steps.pw-version.outputs.version }}
- name: Install OS dependencies only (cache hit)
if: steps.playwright-cache.outputs.cache-hit == 'true'
run: npx playwright install-deps
- name: Install browsers and dependencies (cache miss)
if: steps.playwright-cache.outputs.cache-hit != 'true'
run: npx playwright install --with-deps
What the Cache Cannot Store
The cached path holds only the browser binaries. Libraries like libgtk, libnss3, ALSA, and the fonts Chromium needs live in the operating system rather than that folder, so a cache hit never restores them. Skipping install-deps on a runner that's actually missing those libraries can still produce a launch failure with a vague message that rarely identifies the missing library, since the cached binaries can also carry over Playwright's own record of an earlier successful check.
The Savings Estimate
Applied to the earlier example of ten PRs a day across four shards, correct caching skips that download once the cache is built, cutting most of the forty to sixty minutes it added each day. Once populated, the cache typically cuts Playwright setup from around a minute to a few seconds, a similar order of reduction. The same math applies to any team's own PR volume and shard count.
The 10 GB Cache Limit
GitHub caps Actions cache storage at 10 GB per repository (can be increased; billed to your account), evicting the least recently accessed entries once full, according to its caching documentation. Teams running many active branches, each pinned to a different Playwright patch, can reach that limit without anyone noticing and see cold installs return with no explanation. Often, the cache has simply run out of room, making it the first thing worth checking when hit rates fall.
When configured correctly, browser installs stop costing most pipelines time on every run. The two issues that still show up most often are a skipped install-deps step and a cache that has reached its limit.
Caching fixes what a run installs. Whether that run needs to happen on every push is the next question, and most trigger setups answer it badly.
Fix 2: Stop Running the Full Suite on Every Push
Many workflow files run the full Playwright suite more often than necessary, spending CI minutes on runs that produce no useful signal. Three levers fix this, and they can be used independently or combined.
Concurrency Cancellation
Many workflows define both push and pull_request triggers, so a single commit pushed to an open PR starts the same suite twice. Dropping the push trigger when pull_request already covers the same commits removes that duplication.
Concurrency cancellation addresses a related problem. When a new commit arrives before the previous run finishes, GitHub can cancel the older run and keep only the latest one. Teams that push several times before review avoid paying for runs nobody will examine.
# Place at the top level of the workflow file, not inside a single job
concurrency:
# Cancels the older run in this group when a new one starts
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
Deploy workflows usually need the opposite configuration, cancel-in-progress: false, because a production deployment should never be cancelled mid-run.
Trigger Scoping
Running the full suite on pull requests targeting main, a smoke suite on feature-branch pushes, and a regression suite on a nightly schedule keeps PR feedback fast while preserving broader coverage.
Ten to twenty-five Chromium-only tests are a reasonable starting point for a smoke suite, adjusted to whichever workflows carry the most risk.
Playwright's tagging system supports this cleanly. Mark tests with @smoke in the title, then run npx playwright test --grep @smoke to execute just that subset. Playwright's own CI uses this exact pattern, tagging tests with @smoke and filtering with --grep in its Docker test workflow.
Path Filters
Skipping Playwright for documentation or configuration-only changes works differently depending on whether the suite is required. Repositories without that requirement can use paths-ignore at the trigger level. Repositories that require Playwright before merge cannot: a workflow that never triggers has no status to report, so the required check sits stuck at "waiting for status to be reported" and blocks the merge queue indefinitely.
A common workaround keeps the workflow triggering on every push, but wraps the expensive test job in an if: condition driven by dorny/paths-filter's output, so the job is skipped rather than never started. GitHub counts a skipped job as a passing status, so a required check set on that job usually clears even when the real test steps never execute.
That workaround comes with a cost. Each invocation consumes CI time, and monorepos with many packages pay it repeatedly. A GitHub discussion describes one team spending 23 minutes per commit on path detection alone across more than twenty services. Currents also provides a GitHub Action for rerunning only failed tests, helping filtered runs avoid repeating work that already passed.
Each lever solves a different problem. Concurrency cancellation is usually the safest place to start. Trigger scoping requires a smoke suite worth trusting. Path filtering is a tradeoff that depends on whether the detection cost outweighs the execution time it saves.
None of the three help when a job simply hangs. That problem has its own setting.
Fix 3: Set a timeout-minutes on Every Job
Most engineers only discover that GitHub Actions defaults every job to a six-hour timeout, with no way to set it once for an entire workflow, after a stuck run adds those hours to the bill. A page.goto() that never resolves, or a test waiting on a server that never starts, can keep a runner active until the timeout finally stops it. On a macOS runner, the same tenfold rate turns those six hours into the equivalent of 60 Linux-hours of billing.
One line per job prevents that outcome. A suite that normally finishes in 12 minutes needs a limit closer to 20 or 25, not 360.
jobs:
test:
runs-on: ubuntu-latest
timeout-minutes: 25 # roughly 1.5 to 2 times the suite's normal runtime
Setting that limit does not require guesswork. Take the suite's typical runtime and set the timeout to 1.5 to 2 times that figure. Playwright's scaffolded GitHub Actions workflow uses timeout-minutes: 60 as a starting point, which is reasonable early on and worth tightening once runtimes become predictable.
This setting is separate from Playwright's timeout configuration in playwright.config.ts, which stops a single hanging test from running indefinitely. The job-level timeout covers everything around the tests, including checkout, dependency installation, and a server that never starts. A strict per-test timeout without a job-level timeout can still leave a runner active for hours before the first test executes.
Nothing here helps when a job runs long because work across shards was distributed unevenly. Sharding addresses that problem next.
Fix 4: Shard Correctly for Playwright Specifically
Playwright's --shard splits a suite by file order rather than execution time. One shard can end up with every slow spec while another finishes in three minutes and sits idle. Total runtime can only be as fast as the slowest shard.
The Real Cost of Shard Imbalance
That imbalance still costs money, just not the way it first appears. GitHub bills each shard separately for the time it actually runs, so a shard that finishes early doesn't keep accruing charges after the fact. What actually gets spent is wall-clock time, since the whole matrix waits on the slowest shard, delaying the pipeline and, on accounts already near their concurrent-job limit, holding one of a limited number of runner slots longer than necessary.
The Setup-Cost Math Teams Miss
Take a scenario many teams recognize. Two minutes spent on checkout, dependency installation, and browser setup means four shards pay that cost four times before a test runs. Ten pull requests a day turns that into eighty minutes spent preparing runners, rising to roughly one hundred sixty minutes with eight shards.
Private repositories run ubuntu-latest on two vCPUs, not four. Maxing out workers on a single machine first avoids multiplying that setup cost. Sharding usually starts paying off once a well-tuned runner still takes more than ten minutes to finish.
Reducing Static Imbalance
fullyParallel: true is the first fix worth trying for uneven shards, splitting work by test rather than file so one oversized spec no longer keeps a single shard running long after the others finish.
// playwright.config.ts
import { defineConfig } from '@playwright/test';
export default defineConfig({
fullyParallel: true,
});
Speedboard, a tab added to the merged HTML report in Playwright 1.57, sorts every executed test by duration so imbalance across shards is visible at a glance instead of inferred from wall-clock time. It doesn't rebalance anything on its own. Someone still has to read the numbers and act on them.
Playwright also supports manual shard weighting through an environment variable, PWTEST_SHARD_WEIGHTS, set to colon-separated numbers such as 3:2:3:3 for four shards. It requires Playwright 1.58 or later, and there's no command-line flag, and isn't documented on playwright.dev as of this writing, so treat it as an internal setting that could change without notice. Weights still need recalculating by hand whenever the suite changes, and Speedboard's duration breakdown is a reasonable way to work out what those numbers should be.
Why Not Go Straight to Orchestration
Static and manually weighted sharding can't redistribute work mid-run, keep newly freed machines productive, or survive spot-instance eviction, the way a live queue can. Orchestration becomes worthwhile when runtimes are unpredictable run-to-run rather than merely uneven, and that point usually comes only after worker tuning, sharding, and any hand-set weights have already been tried.
Chromium Only for the PR Gate
Running all three browsers on every pull request triples the shard count. Cross-browser coverage usually fits better in nightly regression runs, while Chromium alone handles the PR gate. Moving Firefox and WebKit to a scheduled workflow keeps feedback fast without reducing coverage.
Artifact Retention and Dynamic Shard Count
Each shard uploads its own report, trace, and video, and at the default ninety-day retention that adds up fast. Setting retention-days: 7 and uploading only on if: ${{ !cancelled() }} instead of always() keeps failures visible without keeping everything.
# One step inside the sharded test job, after tests run
- uses: actions/upload-artifact@v4
if: ${{ !cancelled() }}
with:
name: blob-report-${{ matrix.shardIndex }}
path: blob-report/
retention-days: 7
Suites that grow or shrink outgrow a hardcoded matrix. A dynamic sharding setup that counts tests and computes shard count at run time keeps that number accurate.
Correct sharding, right-sized workers, and honest retention settle most of the sharding bill. A closer side-by-side look at when workers alone are enough and when sharding actually earns its cost covers that decision in more depth for teams still deciding between the two. None of this touches a test that fails only sometimes, a separate billing problem of its own.
Fix 5: Treat Flakiness as a Billing Problem, Not Just a Reliability Problem
Flakiness usually gets treated as a quality issue to fix when time allows. It also increases CI spend, and the impact is easy to measure.
The Cost Mechanism
A flaky test that forces a job rerun costs more than the retry alone. It repeats checkout, install, and browser setup for the shard that failed, on top of what the other shards already consumed. On a four-shard matrix, a single flaky test can add another billed job to the total.
Playwright's in-process retry costs much less. It discards only the failed worker and its browser, then retries within the same job without repeating checkout or installation.
The Visibility Problem
Playwright's HTML report shows one run at a time. It won't show that a test has failed one run in eight for three weeks, or how much retry spend traces back to it. Spotting that needs data from many runs, not one. Currents tracks how often each test fails over time, closing that exact gap with an actual failure rate instead of a guess.
The Intervention
Track flakiness across runs instead of guessing from a single failure. A test that "just failed once" might be failing one time in eight, a rate nobody has tracked.
Pull confirmed flaky tests out of the PR-blocking gate immediately, through quarantine or removal.
Skip retries: 2 on a smoke suite. Retries hide flakiness instead of fixing it, a position this piece takes rather than an assumed standard. Save retries for regression suites already tracked for flakiness.
The cost reduction from removing retries: 2 is modest, since in-process retries don't repeat setup. The larger saving comes from reducing full job reruns triggered by flaky failures. A test failing one run in eight, with one job rerun per failure, adds roughly twelve and a half percent to that job's cost on the PR gate. That estimate assumes the whole shard reruns each time. Pipelines that rerun only the failed test will see a smaller increase and should recalculate using their own rerun strategy.
Caching, trigger scoping, timeouts, and correct sharding all reduce cost per run. Fixing flakiness reduces how many reruns happen, and the effect compounds with every fix already covered. Artifacts introduce another cost, which comes next.
Fix 6: Artifact Storage Is a Silent Cost Driver
Dropping retention to seven days cuts most storage costs with a single setting. Making that reduction stick depends on a few upload practices many pipelines get wrong, along with a version change that still catches teams copying older YAML.
Upload Only on Failure or Completion
if: ${{ !cancelled() }} is usually a better fit than if: always() for Playwright uploads. always() also runs when someone manually cancels a workflow, storing artifacts from a run that never finished and is unlikely to be investigated later.
Upload Only What You Need
Traces and the HTML report provide most of the debugging value. Videos take more storage and are often never reviewed afterward. Setting video: 'off' in playwright.config.ts is worth evaluating case by case because visual regressions and animation issues can still be easier to diagnose when video is available.
Split Retention by Artifact Type
Uploading test-results/ separately from playwright-report/ allows each artifact to have its own retention period. Traces are usually most useful while investigating recent failures, making seven days sufficient in many cases. Reports from runs that merged to main may justify a longer retention window for audit or compliance purposes.
The v4 Unique Naming Requirement
actions/upload-artifact@v4 rejects any attempt to upload a second artifact under a name that's already in use within the same workflow run. A sharded matrix uploads artifacts from each shard independently, and reusing that same name: playwright-report in every shard produces a conflict instead of a successful upload.
The fix is to parameterize the artifact name:
# Inside a job whose strategy.matrix defines shardIndex
- uses: actions/upload-artifact@v4
with:
name: playwright-report-${{ matrix.shardIndex }}
path: playwright-report/
retention-days: 7
The requirement comes from v4's immutable artifact model, introduced when GitHub released the version to general availability. Version 3 allowed uploads with the same name to be merged into a single archive. Version 4 rejects duplicate names, which breaks parallel uploads that share one artifact name.
Version 3 was retired on January 30, 2025, so sharding examples copied from older tutorials fail as soon as they run on a modern workflow.
Getting the cancellation condition, selective uploads, retention periods, and artifact naming right is what turns the earlier retention estimates into real savings. These changes reduce storage costs, but they do not shorten execution time. As suites continue to grow, that pressure shifts back to compute, where sharding eventually reaches its limits.
When to Add Orchestration
Sharding, when tuned well, gets most suites running efficiently. It stops carrying the load once shard imbalance becomes the main source of wasted CI time.
The Threshold for Considering Orchestration
Orchestration is worth considering when the fixes covered so far are already in place and uneven file durations are still driving runtime and cost. It belongs near the end of the optimization path, after the lower-cost fixes have already been applied.
What Changes With a Live Queue
Static sharding assigns files before a run begins, so the slowest shard determines the final runtime. A live queue changes how work is distributed. Runners request new work as they finish, longer tests can be scheduled earlier, and machines that become available mid-run stay productive instead of waiting idle.
The Currents Figure
Currents reports up to 50% reduction in CI time compared with static sharding. That figure reflects a best-case scenario, so most teams should expect a smaller gain in day-to-day runs. It also measures the gain over Playwright's own sharding, not over running tests with no parallelism. FundGuard's case study shows an 80-minute suite dropping to 40 minutes after adopting orchestration, though the company itself credits that result to a wider CI overhaul that also included a Cypress-to-Playwright migration and a move to spot instances.
The full mechanism, working YAML examples, and decision criteria are covered in a deeper guide on orchestration for GitHub Actions and GitLab CI. That completes the optimization path, from browser caching through orchestration. The next section pulls the six fixes, plus this one, into a single prioritized list.
The Sprint-Ready Rollout Plan
The six fixes, the sharding sequence, and the orchestration decision are summarized in the table below. Quick, independent wins come first, followed by fixes with real prerequisites, in the sequence those prerequisites require. Effort and impact sit in separate columns so each fix can be evaluated on its own merits.
| Fix | Effort | Impact | When It Applies |
|---|---|---|---|
| Add timeout-minutes to every job | 5 min | High, prevents runaway billing | Every job, always |
| Add concurrency cancellation | 5 min | Medium to high, cancels duplicate runs | PR workflows with active pushes |
| Cache Playwright browser binaries | 15 min | High, cuts 60 to 90 seconds per shard | Every private repo |
| Set retention-days to 7 | 5 min | Medium, compounds with shard count | Always |
| Scope triggers: smoke on PR, regression on merge | 30 to 60 min | High, removes the full suite from every push | After a smoke suite is defined |
| Add path filters for non-code changes | 30 min | Low to medium, depends on how often docs change | Monorepos and docs-heavy repos |
| Chromium only on the PR gate | 10 min | High, cuts cross-browser suites to roughly a third of their shard count | Suites running all three browsers on every PR |
| Quarantine flaky tests | Ongoing | High, removes the retry cost multiplier | Any suite where retries fire regularly |
| Tune workers before sharding | 30 min | High, avoids unnecessary shard costs | Before adding any shards |
| Manual shard weighting via PWTEST_SHARD_WEIGHTS | 1 hr | Medium, reduces imbalance | Suites with uneven file durations, requires Playwright 1.58 or later |
| Add orchestration | Half a day | High, up to 50 percent against static sharding | After the fixes above, on large suites with shard imbalance |
Worker count gets tuned first. Sharding follows after that, with orchestration reserved for the point where uneven shard runtimes remain a significant source of cost and delay. Following that sequence helps avoid paying for a problem an earlier, lower-cost fix already addresses.
Cost Is a Configuration Problem
Most Playwright CI costs come from configuration choices. Repeated downloads, unscoped triggers, missing timeouts, uneven shards, flaky retries, and long artifact retention all increase cost without adding coverage. Each fix in this piece targets one of those areas, making the pipeline cheaper to run without reducing confidence in the results.
GitHub's January 2026 pricing update lowered hosted runner rates, and that reduced costs for many teams, but pricing changes do not alter the underlying pattern. A pipeline that spends minutes efficiently will cost less regardless of runner type or future pricing adjustments.
Whether a team runs on GitHub-hosted infrastructure or manages its own runners, controlling cost still starts with how the pipeline is configured.
Top comments (0)