Removing register() won't remove your old service worker
Our redesign dropped its service worker, but the old /sw.js would still serve a cached index.html, unseen in server logs. The fix had to live at the old URL.
If you’ve shipped a service worker, run git log --oneline -- public/sw.js | wc -l and count how many of those commits do nothing but bump a cache-name string. In my case it was most of them.
I treated that as process. It was a symptom. I had a component that could overrule the server for every page on the origin, and keeping it in line took a manual step on every deploy.
A stranded browser never asks the server
The redesign moved to / and was built to register no service worker at all. The old console had registered /sw.js at scope /. A browser that already has a root-scope worker installed doesn’t need the new app’s permission to keep running it. Every navigation on the origin goes to that worker first, and if it has a cached index.html, that’s what the user gets.
That is what makes this bug expensive. The worker answers the navigation inside the browser, so the request never goes out. There’s no access log entry, no 500 and no alert. The deploy looks clean because, from the server’s point of view, nobody is asking for anything.
My first fix was a cleanup script in the new index.html that unregistered any worker it found and deleted every cache. That script is still there and it does no harm. It also can’t help the browsers that need it, because a stranded browser never receives the new index.html. The old worker hands it the old one.
The worker’s own update check does work. The browser fetches the registered script URL directly, without going through the worker’s fetch handler, so that request reaches the server. So /sw.js stayed at the same URL, and I replaced its contents with a kill switch:
// /sw.js — served at the URL the old app registered. No fetch handler, on purpose.
self.addEventListener('install', () => self.skipWaiting());
self.addEventListener('activate', (event) => {
event.waitUntil((async () => {
const keys = await caches.keys();
await Promise.all(keys.map((key) => caches.delete(key)));
await self.registration.unregister();
// navigate() only works on clients this worker controls, and
// matchAll() returns only controlled clients by default.
const windows = await self.clients.matchAll({ type: 'window' });
for (const client of windows) client.navigate(client.url);
})());
});
skipWaiting() lets it replace the old worker without waiting for every tab to close. On activation, the open pages from the old worker’s registration are handed to this one, so matchAll({type: 'window'}) returns them and navigate() is allowed to reload them. Because there’s no fetch handler, the reload goes to the network, and so does every load after it.
Where the NetworkFirst advice stops
The standard advice is good. n3ary’s fix switches HTML navigations to NetworkFirst and registers with updateViaCache: 'none'. Daniel Joffe’s writeup traces stale chunks back to stale-while-revalidate on the document. Both are right about the worker you ship next. Both assume there will be a next worker. They don’t cover dropping the PWA, moving to a different frontend on the same origin, or deleting sw.js during a cleanup. In those cases the fix sits in a script the stranded browser never runs.
Deleting the file is worse than leaving it alone. The Update algorithm in the Service Workers spec rejects the update job when the script response isn’t OK, or when its MIME type isn’t a JavaScript MIME type. A rejected update changes nothing, so the registration and its active worker stay exactly as they were. With a 404, Chrome logs A bad HTTP response code (404) was received when fetching the script. If your SPA catch-all serves index.html for the missing file, Chrome logs The script has an unsupported MIME type ('text/html'). In both cases the old worker is still the one answering. progressivewebapps.com describes the same deadlock for renamed workers, and Chrome’s Workbox guidance says the no-op worker has to be served at the URL that was originally registered. I haven’t run that failure across browsers myself, so check it in yours: in Chrome, open chrome://serviceworker-internals after the failed update. If the registration is still listed as ACTIVATED with the old script, you’re seeing it.
Clear-Site-Data: "storage" looks like a shortcut, because a response carrying that header tells the browser to drop the origin’s service worker registrations and caches. The catch is the same one that killed my index.html script: the header only takes effect on a response the browser actually gets from the server, and a stranded browser’s navigations never leave the worker. It can help on an API endpoint the old app still calls over the network, but "storage" also wipes localStorage and IndexedDB, so the user gets logged out with it. I didn’t rely on it. The kill switch at the old URL is the one path that a stranded browser is guaranteed to request.
Every URL ever passed to register() still needs a JavaScript 200
This affects any app where a service worker once controlled the origin: Create React App and Workbox builds, Vite PWA plugin apps, Next.js with next-pwa, or an internal dashboard rebuilt on the same hostname. The new code can’t evict the old worker, because the old worker decides what code the browser runs.
Every script URL that was ever passed to navigator.serviceWorker.register() on an origin must keep returning a 2xx with a JavaScript MIME type, serving either a kill switch with no fetch handler or your current worker, for as long as any browser might still have it installed.
The difficult part is listing those URLs. CRA calls register(swUrl, config) with a variable built from PUBLIC_URL. vite-plugin-pwa injects registerSW.js at build time, and next-pwa generates its own registration, so the string literal is never in git. A grep that only matches register('...') comes back empty on exactly those apps, and an empty result looks like a pass. Cast a wider net:
# 1a. Every worker-shaped file your history ever contained
git log --all --name-only --format= \
| grep -iE "(^|/)(sw|service-?worker|registerSW)[^/]*\.js$" | sort -u
# 1b. Every register() call in history, literal or variable
git log -p --all -S "serviceWorker" | grep -oE "serviceWorker\.register\([^)]*\)" | sort -u
# 1c. Built output, where generated registrations actually live
grep -rhoE "serviceWorker\.register\([^)]*\)|registerSW[^\"' ]*" dist build out .next 2>/dev/null | sort -u
# 1d. What you actually served: an old index.html from a release tarball
# or the Wayback Machine, saved as old-index.html
grep -oE "[^\"' ]*(sw|service-?worker|registerSW)[^\"' ]*\.js" old-index.html | sort -u
# 2. What each candidate URL serves today (repeat per URL)
curl -s -o /tmp/sw.out -w "%{http_code} %{content_type}\n" https://your.app/sw.js
grep -cE "addEventListener\(['\"]fetch|onfetch|workbox" /tmp/sw.out
grep -oE "importScripts\([^)]*\)" /tmp/sw.out # fetch these too and run the same grep
For each URL, a pass is any 2xx status with a JavaScript type: application/javascript, text/javascript (the type WHATWG now prefers), either one with ; charset=utf-8, and so on. It fails if the status isn’t 2xx or the type is text/html, which usually means your SPA fallback is answering for a file you deleted. Step 2 is heuristic. A count of 0 suggests a kill switch, but minified bundles and files loaded through importScripts can hide a fetch handler from the grep. The ground truth is in the browser. Open a profile that used the old app, check chrome://serviceworker-internals, and run (await navigator.serviceWorker.getRegistrations()).map(r => r.active?.scriptURL) and await caches.keys() in the console. If you see an old cache name there, that browser is still running a build you thought you had replaced.
Count requests for /sw.js before you delete the kill switch
I got this part wrong, and the error is the useful part. I shipped the kill switch without counting the requests for it, so I can’t tell you how many browsers were stranded or how fast that number fell. That count is the evidence you want, and it’s cheap to get. Once the kill switch is live, each stranded browser requests /sw.js when its user next opens the app, runs the kill switch, unregisters, and doesn’t request it again. The new app never registers it, so every non-bot GET /sw.js in your access log is a browser that was still stranded until that moment. Count those per day, starting the day you deploy.
That gives you the exit rule I should have written down at the start: retire the kill switch once /sw.js requests from non-bot user agents have stayed at zero for N consecutive days, where N is longer than the longest gap your real users go between visits. A browser only checks for an update when someone opens the app, so a user who comes back quarterly needs a quarter-sized N. For an internal tool used daily, 30 days is generous. For anything with occasional visitors, I’d use 90. Until the count reaches zero and stays there, the URL is still in use, even if nothing in your current code references it.