DEV Community

Kazu
Kazu

Posted on

Deploy weekly, and the previous version keeps running for 17 hours

After a deploy, the service that collects errors from users' browsers (Sentry or similar) starts showing this:

TypeError: Failed to fetch dynamically imported module: https://app.example.com/assets/feature-DHyZGV0s.js
Enter fullscreen mode Exit fullscreen mode

The browser tried to load a JS file with import() at runtime and could not fetch it. feature-DHyZGV0s.js in the URL is a chunk: part of the app's functionality, split out into its own file by the build. DHyZGV0s is a hash the build derived from the file's contents.

Around the same time, some users report that "the button does nothing." Someone on the team opens the same page and clicks the same button, and it works. Nobody can reproduce it locally, and the errors keep coming in, not only right after the deploy but for hours afterward.

Impact

Users who lose a feature this way have no way to fix it themselves. They don't know what would make it work again, so all they can do is file a report. On the receiving end, the team spends time answering reports that don't reproduce internally and chasing their cause. This repeats with every deploy. In the browsers that throw the error, the deployed change is not running. A deploy meant to fix a bug does not reach those users. The side serving the app does not control which version runs in which user's browser. After every deploy, some users stay off the fixed version for as long as the previous version keeps running.

Reproduction

I built a minimal app with Vite. The build produces three files:

  • index.html: the HTML the browser fetches first when it opens the site
  • Entry JS: loaded by index.html, renders the page
  • Chunk: loaded with import() when the button is clicked

I built it twice with slightly different contents: a previous version, served first, and a new version that replaces it in the deploy. Since the contents differ, the hashes in the JS file names differ too. index.html is the only file without a hash.

Each version's files were baked into an nginx container image. To check whether this happens without any special configuration, nginx ran with the default config that ships in the image. A deploy means replacing the previous version's container with the new version's container. The new image does not include the previous version's JS, so after the deploy those files are gone from the server. The previous version started being served 10 minutes before the browser opened the site.

The browser is Chromium driven by Playwright. Each trial uses a fresh profile (the on-disk storage for history and other data that survives closing the browser), and runs these steps:

  1. Open the site while it serves the previous version. Two cases: click the button here to fetch the chunk, or don't
  2. Close the browser
  3. Deploy the new version. This finishes within a few seconds of opening the site
  4. Once a set time has passed since the first open, relaunch the browser with the same profile, open the site again, and click the button

Only two things are recorded: which version ran after reopening (the app displays its own version on the page), and whether the chunk loaded.

The setup was Chromium headless shell (the build without a window) 153.0.8010.12, Playwright 1.63.0, nginx 1.30.5, and Vite 8.3.1. The reproduction setup and scripts are in show-your-work.

I tried reopening 30, 54, and 66 seconds after the first open.

Without clicking the button first

Each case ran 18 times in CI, with identical results every time.

First open to reopen Version on reopen Result of clicking the button
30.1 to 30.3 s Previous Error
54.1 to 54.2 s Previous Error
66.1 to 66.2 s New New version's feature shown

When the previous version ran, clicking the button failed to load the chunk with the following error, and nothing appeared on the page.

TypeError: Failed to fetch dynamically imported module: http://127.0.0.1:18000/assets/feature-DHyZGV0s.js
Enter fullscreen mode Exit fullscreen mode

The site was opened after the deploy, yet the previous version ran, tried to load its own chunk, and failed. This is the same error as at the top, and the same "button does nothing" behavior.

With the button clicked first

On the first open, I clicked the button so the previous version's chunk was loaded, then closed the browser. Each case ran 9 times in CI, with identical results every time.

First open to reopen Version on reopen Result of clicking the button
30.1 to 30.2 s Previous Previous version's feature shown
54.1 to 54.2 s Previous Previous version's feature shown
66.1 to 66.4 s New New version's feature shown

The version on reopen was the same as without the click. The only difference was what happened when the button was clicked while the previous version was running: the already-loaded chunk ran as is, and not a single error occurred.

In both cases, for a browser that opened the site 10 minutes after the previous version started being served, reopening 54 seconds after the first open still ran the previous version, and reopening at 66 seconds ran the new one. The error appears only when the user had not clicked that button under the previous version. Only part of the users running the previous version show up in the error collector. The rest never appear there.

Narrowing it down

Access log

Look at the nginx access log around the time of the error. For a browser that ran the previous version and hit the error, the new version's server logged only this one line:

"GET /assets/feature-DHyZGV0s.js HTTP/1.1" 404 153 "http://127.0.0.1:18000/assets/index-DB4b__oJ.js"
Enter fullscreen mode Exit fullscreen mode

The previous version's chunk is being requested by the same version's entry JS (index-DB4b__oJ.js). The new version's server doesn't have that file, so it returns 404.

After the deploy, the previous version's page is running. I checked the possible causes in order, starting from the server side.

Old and new mixed during the deploy

In a rolling deploy, some requests reach previous-version servers while the replacement is in progress. The reproduction replaces a single container, and that finishes within a few seconds of the first open. The request above reached the new version's server about 50 seconds after that. The previous version's page is running after the replacement has finished, so this is ruled out.

Server-side cache

Check whether the server is returning the previous version's index.html, including any proxy or cache in between if the setup has one. Look at the new version's access log per browser. From browsers that ran the new version, requests for index.html and the entry JS arrived, and both returned 200.

"GET / HTTP/1.1" 200 321 "-"
"GET /assets/index-D879X9_S.js HTTP/1.1" 200 2573 "http://127.0.0.1:18000/"
"GET /assets/feature-CpSIXz8z.js HTTP/1.1" 200 89 "http://127.0.0.1:18000/assets/index-D879X9_S.js"
Enter fullscreen mode Exit fullscreen mode

From browsers that ran the previous version, no request arrived for index.html or for the entry JS. The chunk request was the only line. The server isn't returning the previous index.html; it isn't being asked for index.html at all. Ruled out.

Tabs left open

A tab left open since before the deploy can also request the previous version's chunk. In the reproduction, the browser was closed after the first open and relaunched after the deploy. The page was loaded after the deploy, so it is not a tab left open. Ruled out. To check this in your own environment, attach the page load time (performance.timeOrigin) when reporting errors and compare it with the deploy time.

Browser HTTP cache

What remains is the browser using an index.html it already had. Whether a page load came from the server or from the browser's own copy shows up in transferSize in Navigation Timing, which the browser records for each page load. Run this in the browser DevTools console to see the value:

performance.getEntriesByType('navigation')[0].transferSize
Enter fullscreen mode Exit fullscreen mode

JS files have the same transferSize in Resource Timing, which you can list with:

performance.getEntriesByType('resource').map(r => [r.name, r.transferSize])
Enter fullscreen mode Exit fullscreen mode

The value is positive when bytes came over the network and 0 when the response came from the browser's HTTP cache. When the previous version ran, both index.html and the entry JS were 0. When the new version ran, index.html was 621 (larger than the byte count in the access log because it includes headers) and the entry JS was 2873. The previous version's page was rendered from the index.html and JS in the browser's HTTP cache. In the case where the button had been clicked first, the chunk was 0 too, and not a single request reached the new version's server.

The headers on the index.html received at first open had no Cache-Control, which specifies how long a response may be cached. nginx's default config doesn't add one. Instead, they had the time the response was sent (Date) and the time the file was last modified (Last-Modified).

Date: Tue, 29 Sep 2026 08:26:55 GMT
Last-Modified: Tue, 29 Sep 2026 08:16:53 GMT
Enter fullscreen mode Exit fullscreen mode

Last-Modified is about 10 minutes before Date, matching the reproduction's setup of starting to serve the previous version 10 minutes before opening the site.

The browser is using an index.html with no specified cache lifetime without asking the server. But when reopening 66 seconds after the first open, it does go back to the server for index.html.

Relationship to time since modification

Next, I measured what determines how long the browser uses index.html without asking the server. Three things were varied:

  • How long index.html had gone unmodified at the moment it was fetched: the difference between Date and Last-Modified at fetch time, called time since modification below. "Opened 10 minutes after the previous version started being served" in the reproduction is a time since modification of 10 minutes. The 30, 54, and 66 second reopens correspond to 5%, 9%, and 11% of that
  • The time from the first open to the reopen
  • How the page was reopened: relaunching the browser and opening the site, or reloading the open page

Everything else matched the reproduction, and the button was not clicked on the first open. Each condition ran 4 times in CI, for 12 trials per data point.

Time since modification

I set the time since modification to 10 minutes, 100 minutes, 1000 minutes, and 6 days, and reopened at 5%, 9%, 11%, and 20% of each (9% of 6 days is over 13 hours). None of these times were actually waited out; the appendix at the end explains how they were produced.

index.html modified to first open How long the previous version ran
10 min About 1 min
100 min About 10 min
1000 min (about 17 h) About 100 min
6 days About 14 h

Each was about 10% of the time from modification to first open. With 100 minutes, for example, reopening after 9 minutes ran the previous version, and after 11 minutes the new one. With 6 days, reopening after 13 hours still ran the previous version, and clicking the button produced the error.

Reload

With a time since modification of 100 minutes, I relaunched the browser, opened the site, and reloaded the page before clicking the button. In all 24 trials where the previous version was running when the site opened, the reload switched to the new version, and clicking the button showed the new version's feature. The reload fetched index.html from the server again.

What is happening

For responses without Cache-Control, the HTTP caching spec (RFC 9111, section 4.2.2) lets a cache guess how long it may use the response without checking, and gives 10% of the time since Last-Modified as a typical value. Chromium uses exactly this, with no upper limit (GetFreshnessLifetimes in 153.0.8010.12).

lifetimes.freshness = (date_value - last_modified_value.value()) / 10;
Enter fullscreen mode Exit fullscreen mode

When build artifacts are baked into an image, Last-Modified is the build time. On a site that deploys once a week, a browser that opened the site just before the next deploy uses index.html and the entry JS without checking for about 17 hours, 10% of 7 days.

If a deploy happens in that window, reopening the site still runs the previous version's page. Its chunks are gone from the server, so using a feature that hasn't been loaded yet returns 404. After about 17 hours, index.html is fetched again and the new version takes over.

Users who opened the site earlier in the week had a shorter time since modification, so the previous version lingers for less time after the deploy. With weekly deploys, the errors last at most about 17 hours. Browsers that open the site for the first time after the deploy, or that reload, get the new version, so checking internally may not reproduce it.

What decides it

Which version runs in the browser is decided not by when you deployed, but by when the HTML the browser holds gets fetched again.

Hashed files change name whenever their contents change. Which version's files get loaded is decided entirely by the HTML, whose name never changes. If you don't specify a cache lifetime for that HTML, the browser decides when to fetch it again. When a deployed version reaches users is then out of the serving side's hands.

Fixes

Set a cache lifetime for the HTML

Add no-cache to index.html

With Cache-Control: no-cache on index.html, the browser checks with the server every time it opens the page whether index.html has changed. If it hasn't, the server returns 304 and sends no body. Opening the site after a deploy gets the new version's index.html.

In the reproduction's nginx, I added this:

location = /index.html {
    root   /usr/share/nginx/html;
    add_header Cache-Control "no-cache";
}
Enter fullscreen mode Exit fullscreen mode

The cost is one extra round trip to the server each time the page opens. This fits when you want deployed versions to reach users right away.

Add max-age to index.html

With Cache-Control: max-age=<seconds>, the browser uses index.html without checking for that many seconds, regardless of time since modification. The window in which the previous version runs doesn't go away. You decide its length yourself instead of letting the browser guess.

If the seconds you specify are longer than what the browser would have guessed, the previous version actually runs longer. This fits when you want to avoid a round trip on every open but still want a ceiling on how long the previous version can run.

Reload when loading fails

Catch the chunk load failure and reload the page. A reload fetches index.html again, so it switches to the new version. In Vite, a load failure is signaled with the vite:preloadError event. In the reproduction, I reloaded only once, like this:

window.addEventListener('vite:preloadError', () => {
  if (!sessionStorage.getItem('reloaded-once')) {
    sessionStorage.setItem('reloaded-once', '1');
    window.location.reload();
  }
});
Enter fullscreen mode Exit fullscreen mode

The cost is that in-progress page state is lost on reload, and you need a guard so that it doesn't reload repeatedly if the new version also fails. The code above never clears its flag, so if the same tab fails again after the next deploy, it won't reload. The previous version still runs, and this does nothing for users who are running it without hitting an error. Don't use it alone. Combine it with no-cache or max-age, and use it to move users who still hit the error onto the new version.

Effect of the fixes

Using the same setup as the reproduction, I opened an index.html with a time since modification of 100 minutes, varied the time until reopening, and compared against no fix. Each condition ran 4 times in CI, for 12 trials per data point.

no-cache

Reopen after No fix no-cache
5 min Previous (error) New
9 min Previous (error) New
11 min New New
20 min New New

The new version ran at every point measured, and the window where the previous version runs disappeared. When reopening without a deploy, there was one request for index.html each time, which returned 304 with no body.

max-age

I added max-age=300 (5 minutes).

Setup How long the previous version ran
No fix About 10 min (previous at 9 min, new at 11 min)
max-age=300 About 5 min (previous at 4 min, new at 6 min)

The window where the previous version runs shrank from about 10 minutes to about 5. Reopening at 9 and 11 minutes also ran the new version. With a time since modification of 1000 minutes, the result was the same: previous at 4 minutes, new at 6. The window is the length you specify, regardless of time since modification.

Reload on failure

Reopen after No fix Reload on failure
5 min Previous (error) Previous, new after reload
9 min Previous (error) Previous, new after reload
11 min New New
20 min New New

The previous version ran for about 10 minutes, the same as with no fix. Clicking the button under the previous version failed to load the chunk, triggered one reload, and then showed the new version's feature.

Scope

These results assume that index.html has Last-Modified but no Cache-Control, that hashed files are removed by the deploy, and that the browser is Chromium. Change any of these and the length of the previous-version window or the way errors show up changes. Of the variations below, only two were measured: the setup that keeps the previous version's JS, and the setup that doesn't return 404. The rest are read from source code and how the mechanism works.

Browsers

Only Chromium headless shell 153.0.8010.12 was measured.

According to their source, Firefox and Safari also use index.html without checking for 10% of the time since Last-Modified, the same as Chromium (Firefox: nsHttpResponseHead.cpp; Safari: WebKit's CacheValidation.cpp). With weekly deploys, the maximum is about 17 hours in every browser.

Firefox is the only one that differs from Chromium: it caps this window at one week. The cap kicks in only when the time since modification exceeds 70 days, so if you deploy more often than every 70 days, there is no difference.

In the reproduction, the browser was closed and relaunched for each reopen. Opening the site in a new tab with the browser still running was not measured.

When the previous version's JS is kept

I included the previous version's JS in the new image as well and measured with a time since modification of 100 minutes. The previous version ran for about 10 minutes, the same as with no fix, and clicking the button during that time returned the previous version's chunk with 200 and showed the previous version's feature. Not a single error occurred.

With no errors, it is harder to notice that the previous version is running. The previous version's page keeps calling the newly deployed server-side code, so that code has to stay compatible with it. A setup that puts build artifacts in S3 and runs aws s3 sync without --delete is in this state, because the previous version's files are never removed.

When missing paths don't return 404

In the main results, the request for the previous version's chunk got a 404. With nginx configured to return index.html for missing paths (try_files $uri /index.html), the same request gets a 200 whose body is index.html. Searching the access log for 404 finds nothing. Instead, look for lines where a hashed file path returned 200 with 321 bytes, the same size as index.html.

"GET /assets/feature-DHyZGV0s.js HTTP/1.1" 200 321 "http://127.0.0.1:18024/assets/index-DB4b__oJ.js"
Enter fullscreen mode Exit fullscreen mode

The message that reached the error collector was the same Failed to fetch dynamically imported module as with the 404. The browser console also showed this error:

Failed to load module script: Expected a JavaScript-or-Wasm module script but the server responded with a MIME type of "text/html". Strict MIME type checking is enforced for module scripts per HTML spec.
Enter fullscreen mode Exit fullscreen mode

Other ways of serving

Whether you serve from S3, a CDN, or a hosting service, or render the HTML on the server per request, you can tell whether this applies by looking at the response headers for index.html.

curl -sI https://app.example.com/ | grep -iE '^(cache-control|expires|last-modified|age):'
Enter fullscreen mode Exit fullscreen mode

If there is neither Cache-Control nor Expires, and only Last-Modified, the same thing as in this article will happen. If there is an Age header, the response came from a CDN in between that was holding it.

When file timestamps are pinned

If the build pins file modification times to a fixed past time (SOURCE_DATE_EPOCH and similar), Last-Modified is that time, not the build time. The time since modification is then not the deploy interval but however long it has been since that fixed time. Chromium has no upper limit, so the further back the fixed time, the longer index.html is used without checking, potentially for years.

With a Service Worker

If a Service Worker serves index.html from its own cache, the same symptom appears. In that case, adding no-cache or max-age to index.html has no effect. To check whether one is registered, run navigator.serviceWorker.getRegistrations() in the console.

Built with webpack

The same failure surfaces as ChunkLoadError. The message starts with Loading chunk <chunkId> failed. (webpack v5.111.1 source).

Appendix: how the times were produced

The time since modification was produced without waiting. nginx returns a file's modification time as Last-Modified, so when starting the container, I shifted the modification time of the files on the server into the past.

The time until reopening was actually waited out in the reproduction, and produced without waiting in the conditions that varied the time since modification. When the previous version's server adds an Age: <seconds> header to index.html, the browser treats the response as already that many seconds old at the moment it receives it. The reopen happened right after the fetch, and the Age value plus the few seconds that actually passed was recorded as the time until reopening.

I confirmed that this method doesn't change the results by comparing it with actually waiting, at a time since modification of 10 minutes. With both methods, the switch to the new version happened once the time until reopening exceeded 10% of the time since modification.

Changing when the new version is deployed (right after the first open, halfway to the reopen, or right before the reopen) gave the same results. This was confirmed by actually waiting, at a time since modification of 10 minutes.

Top comments (1)

Collapse
 
_firelinks profile image
Mike Dabydeen •

I'd add one measurement, because the case that worries me most is the one you found with the previous JS kept: no errors, and the old page keeps calling the new server. The error collector only hears from users who click a chunk that wasn't loaded, so it can't tell you how many browsers are still on the old build.

Have the client send its build version on every API call, for example a header filled from a constant Vite injects at build time through define, and count requests by version on the server. After each deploy you get a curve of old-version traffic decaying. Its tail should follow your 10% rule, about 17 hours for a weekly deploy with no Cache-Control, plus whatever tabs people leave open, which the curve shows and the cache rule doesn't.

It also answers the question your kept-JS section leaves open: when you can remove the server-side compatibility the previous version still depends on. I'd wait for that curve to reach zero rather than for the deploy to finish.