DEV Community

Cover image for CI pipeline caching: cutting a 12 minute build down to 3
Amit Shukla
Amit Shukla

Posted on Originally published at amitshuklabag.hashnode.dev

CI pipeline caching: cutting a 12 minute build down to 3

You push a one line fix and wait twelve minutes for CI to tell you it passed. Multiply that by every push, every developer, every day, and the pipeline is quietly the slowest part of your team. The frustrating part is that most of those twelve minutes are not spent testing your change. They are spent downloading the same packages, rebuilding the same Docker layers, and compiling the same code that the last run already handled.

Hosted CI runners start from a clean machine every time. Nothing from the previous run survives unless you explicitly save it and restore it. Caching is how you stop paying full price on every push, and done carefully it is the single biggest speedup most pipelines can get without changing a line of application code.

Measure before you cache anything

Before adding a single cache, open a recent run and write down how long each step takes. It is tempting to guess, and the guess is usually wrong. A typical Node service with a Docker image might look like this:

checkout                  0m 06s
setup node                0m 12s
npm ci                    3m 48s
lint                      0m 41s
test                      3m 55s
docker build              3m 22s
total                    12m 04s
Enter fullscreen mode Exit fullscreen mode

Three steps account for almost the whole run: installing dependencies, running tests, and building the image. Every one of them redoes work from scratch. That list tells you where to spend your effort, and it gives you a baseline so you can prove each change actually helped.

Cache dependencies keyed on the lockfile

The install step is the easiest win. Your dependencies only change when your lockfile changes, so the lockfile hash makes a perfect cache key. In GitHub Actions:

- uses: actions/cache@v4
  with:
    path: ~/.npm
    key: npm-${{ runner.os }}-${{ hashFiles('package-lock.json') }}
    restore-keys: |
      npm-${{ runner.os }}-

- run: npm ci
Enter fullscreen mode Exit fullscreen mode

Two details in there matter more than they look.

Cache the download folder, not node_modules. ~/.npm holds the package tarballs npm already fetched. npm ci still deletes node_modules and does a clean, correct install, it just pulls from the local cache instead of the registry. Caching node_modules directly skips that clean install, which means you can restore packages built for a different Node version or a different OS and not find out until something breaks at runtime.

restore-keys gives you a partial hit. When someone adds one package, the lockfile hash changes and the exact key misses. Without a fallback you are back to a full download. With restore-keys, the pipeline restores the most recent cache that matches the prefix, so npm only fetches the handful of packages that are actually new.

If you use actions/setup-node, you can get the same behavior with one line, cache: npm, which caches the right folder and keys it on your lockfile for you. The same pattern works everywhere: ~/.cache/pip keyed on requirements.txt, ~/go/pkg/mod keyed on go.sum, ~/.m2/repository keyed on pom.xml.

In the example above, this took npm ci from 3m 48s to about 20 seconds on an exact hit.

Cache Docker layers between runs

On your laptop, Docker reuses layers from the last build automatically. On a fresh CI runner there is no last build, so every layer gets rebuilt, including the expensive dependency install inside the image. Docker's layer cache needs somewhere to live between runs, and BuildKit can export it to the GitHub Actions cache backend:

- uses: docker/setup-buildx-action@v3

- uses: docker/build-push-action@v6
  with:
    context: .
    push: false
    tags: myapp:ci
    cache-from: type=gha
    cache-to: type=gha,mode=max
Enter fullscreen mode Exit fullscreen mode

mode=max stores cache for every layer, including intermediate stages in a multi stage build, not just the layers that end up in the final image. That is usually what you want, since the dependency install stage is exactly the layer you are trying to reuse.

This only pays off if your Dockerfile is ordered so that expensive, rarely changing steps come first. Copy the lockfile, install dependencies, then copy the rest of the source. If you copy the whole project before installing, every code change invalidates the install layer and the cache never hits. In our example, a well ordered Dockerfile plus layer caching took the image build from 3m 22s to around 40 seconds.

Stop running everything on every push

Caching makes each step faster. The next win is not running steps that have nothing to do.

Split slow test suites across parallel jobs. Test runners like Jest support sharding natively, so a matrix can run one third of the suite on each of three runners at the same time:

strategy:
  matrix:
    shard: [1, 2, 3]
steps:
  - run: npx jest --shard=${{ matrix.shard }}/3
Enter fullscreen mode Exit fullscreen mode

The wall clock time drops to roughly the slowest shard instead of the whole suite. You pay for more runner minutes in total, but the developer waiting on the result gets an answer much sooner, which is the number that actually matters day to day.

Cancel runs that are already outdated. If you push three commits in a row, the first two runs are wasted work. A concurrency group cancels them automatically:

concurrency:
  group: ci-${{ github.ref }}
  cancel-in-progress: true
Enter fullscreen mode Exit fullscreen mode

Skip pipelines that do not apply. A docs change does not need a Docker build. Path filters on the workflow trigger keep irrelevant changes from starting the heavy jobs at all.

Where caching goes wrong

Caching is a trade. You trade correctness risk for speed, and a few habits keep that trade safe.

Keys that are too broad serve stale results. A key like npm-cache with no lockfile hash never changes, so the pipeline keeps restoring an old cache forever. Always include the hash of whatever file actually defines the contents.

Huge caches can be slower than no cache. Restoring a cache means downloading and extracting it. A multi gigabyte cache full of old build artifacts can take longer to restore than a fresh install. GitHub also limits each repository's cache storage and evicts entries that have not been used in seven days, so bloated caches push out the useful ones. Keep cache paths narrow and specific.

Never cache anything that holds a secret. A .npmrc with an auth token or a config file with credentials does not belong in a cache path. Caches can be restored by other branches and workflows, and a secret stored there is a secret you no longer fully control.

Rebuild from scratch on a schedule. A nightly or weekly run with caching disabled catches the problems a warm cache can hide, like a dependency that only installs because an old version was still sitting in the cache.

The result

Putting it together on the example pipeline:

checkout                  0m 06s
setup node (cached)       0m 05s
npm ci (cache hit)        0m 20s
lint                      0m 41s
test (3 shards)           1m 18s
docker build (gha cache)  0m 40s
total                     3m 10s
Enter fullscreen mode Exit fullscreen mode

Nothing about the application changed. The pipeline just stopped redoing work it had already done.

Takeaway

A slow pipeline is almost always a pipeline that forgets everything between runs. Measure where the time goes, cache dependencies keyed on the lockfile, give Docker somewhere to keep its layers, and stop running jobs that have nothing to do. Then keep your cache keys honest, because a fast pipeline that serves stale results is worse than a slow one.

Top comments (1)

Collapse
 
launchgatecheck profile image
Launch Gate •

The scheduled cold run is a useful backstop. I'd add one more comparison: build the same commit with and without the layer cache, then run the same checks on the resulting images. A cold build passing on its own doesn't establish that the cached image is equivalent. This is especially useful for an install step that depends on something outside the lockfile, such as a base-image tag or a downloaded system package. How do you handle those inputs in the cache key or rebuild policy?