Your old frontend's service worker will outlive your redesign
Moving a redesigned agent console to '/' broke two things no diff showed: a cached service worker, and stream cursors that outlived a server restart.
If you are building an agent system, you probably have a web console in front of it. It streams model output over a websocket, it shows tool runs, and it has picked up a few years of panels nobody remembers adding. At some point you rebuild it, and you have to swap the new console in for the old one at the same URL while people have tabs open.
Most writing about redesigns covers navigation and layout. My cutover went wrong in two places that had nothing to do with layout. Both came from state the browser held, which pointed at a server that no longer existed.
The redesign moved to / and the old tree went to /classic/ for one release
Vodou’s console had been living as a staging build at /next/ while I used it every day. The cutover, done on 2026-09-03, made three moves. public/next/ became public/ and is served at /. The tree it replaced became public/classic/ at /classic/, kept for exactly one release as an escape hatch and then deleted. /next/ and /next/index.html return a 301 to /. The icons/ and vendor/ directories were byte-identical in both trees, so a single copy stays at the root.
The rail went down to six destinations. I dropped the New and Search tiles because neither one is a place you go: new chat is the + in the thread column, and the command palette is ⌘K, which the composer footer shows as a hint. Two days later the rail got a seventh tile, Feedback, which links to our Discord.
In the plan this step was a directory rename. It turned into a rewrite.
86 absolute paths in one index, 79 /next/ paths in the other
The first problem showed up when I grepped. The old index.html had 86 absolute /js/... and /css/... references. The staging index had 79 references prefixed with /next/. Once the trees changed places, every one of those 165 references pointed at the wrong tree. Neither index would have failed loudly: each would have loaded the other console’s stylesheets and returned a 404 for the scripts the other tree lacked. I rewrote both indexes and then grepped the scripts too, which held no absolute paths. That second grep is the step people skip.
That one was tedious but expected. The next two were not.
The old console registered /sw.js at scope / and would have kept serving itself
The old console registered a service worker at scope / that cached its static assets. Its last cache name was vodou-v313. Think through what that means at cutover. A browser that loaded the old console still has that worker installed. On the next visit to /, the worker answers from cache and serves the old index.html. The new console is deployed and correct, and the user never sees it. The server logs show nothing wrong, because those requests never reach the server.
My first instinct was to delete sw.js. That is the wrong move, and I was close to making it. When the browser checks for a worker update and the script returns 404, the update fails and the installed worker keeps running. Deleting the file leaves the old cache in front of your users indefinitely.
The fix has three parts. The classic index stops registering the worker. The new root index looks up any existing registrations on every load and unregisters them; this is idempotent, and ?reset-sw forces the same thing. And sw.js remains at the same URL, but it is now a kill switch:
self.addEventListener('install', () => { self.skipWaiting(); });
self.addEventListener('activate', (event) => {
event.waitUntil(
caches.keys()
.then((names) => Promise.all(names.map((n) => caches.delete(n))))
.then(() => self.registration.unregister())
.then(() => self.clients.matchAll({ type: 'window' }))
.then((clients) => { clients.forEach((c) => { try { c.navigate(c.url); } catch (_) {} }); })
);
});
There is deliberately no fetch handler. A browser holding the old worker downloads this file on its next navigation, installs it right away, deletes every cache, unregisters itself, and reloads its open windows. After that, nothing sits between the page and the network. I also extended the no-cache header rule for view assets to cover the new paths.
One reply lost six characters, and the next started 150 characters in
The second problem surfaced during real use, on a console I had open for QA. In one assistant reply, six characters were missing mid-sentence. The reply after it began about 150 characters into its text. Both messages were complete in the database. The corruption was only on screen.
The cause turned out to be the gateway restart at 15:04, with my tab still open. The console streams tokens over a websocket with per-conversation sequence numbers. On reconnect, the client sends its last cursor and asks the server to resume. That works when the same process comes back. A restarted process has a new replay buffer and its own sequence counter, and the client’s cursor refers to a buffer that is gone. The server resumed from a position that meant something different in the new process, and the client discarded what it took for duplicates.
I moved the resume request from socket-open to the connected handshake, which now includes the gateway’s process epoch. If the epoch matches, the client resumes as before. If it differs, the client clears every per-conversation cursor, skips the resume because that buffer never existed in this process, and flags the event so the chat view clears its own sequence map too. That last part matters: there were two cursor maps, and fixing only one would have left the bug in place. A gateway too old to send an epoch gets the old behavior.
A Runs panel that said “No runs recorded yet” after fourteen runs
A third bug from the same week has the same structure. A skill’s Runs panel said “No runs recorded yet.” The skill had run fourteen times. The panel read the graph-run ledger, and scheduled skills write to a different one, the scheduler’s. Now, when the graph ledger is empty, the panel falls back to the scheduler’s runs for skill:<name> and maps each row into the shape it already renders. An empty result from the wrong store was being shown to me as a fact about the world.
The invariant: client state must carry the identity of the server that minted it
All three bugs are one failure class. Here it is as a property you can check:
Any state a client holds that refers to server state (cached assets, resume cursors, “nothing found” results) must record which server process, build or store produced it, and the client must discard it when that identity changes.
A service worker cache with no build identity and no removal path fails this. A resume cursor sent before the client knows which process it is talking to fails this. A panel that turns “this table has no rows” into “this never happened” without naming the table fails this. In each case the code is correct on its own terms. The bug sits in an assumption that the counterpart is the same one as last time.
Check your own console in five minutes
Start with the service worker. In Chrome, open chrome://serviceworker-internals and search for your console’s origin. If you see a registration and your current frontend doesn’t register one, you have the leftover-worker problem. Then check what the worker’s update check gets:
curl -s -o /dev/null -w '%{http_code} %{content_type}\n' https://your-console.example/sw.js
It passes if this returns 200 with a JavaScript content type and the body unregisters itself, or if you have never shipped a worker at that path. It fails if it returns 404 while browsers still hold a registration: those users are pinned to the old cache.
Next, check stream resume. Open your console, start a long streamed reply, and restart the backend process partway through. When it finishes, compare the rendered text with what you stored:
SELECT id, length(content) AS stored_chars
FROM messages
WHERE conversation_id = :conv
ORDER BY created_at DESC LIMIT 2;
In the browser, compare against document.querySelectorAll('.message')[n].innerText.length for the same messages. If the lengths match, or the client reloaded from scratch, you pass. If the rendered text is shorter, or a reply starts mid-sentence, your cursor outlived the process. Then grep your handshake:
grep -rnE "resume|lastSeq|last_seq|cursor" src/ | grep -iE "open|connect"
If the resume is sent inside an onopen handler before any server message includes a process or boot identifier, you have the same bug I had.
What the redesign writeups leave out
The public redesign writeups I read are good on information architecture. The VaultCrux console tour also settles on six rail destinations and applies the persisted rail state before first paint so the layout does not jump. The kontourai console’s overview slice moves a new landing view to / with “all existing views kept reachable during migration,” which is the same decision as my /classic/. The seamless operations-rail redesign notes “No handler or data-shape changes.”
All three treat a redesign as a change to what the server sends. None of them mention what the browser already holds from the previous version. Keeping old routes reachable deals with stale links. It does nothing about a stale worker that intercepts the new route, or a cursor issued by a process that no longer exists. An agent console makes this worse than a normal web app, because it is a long-lived tab against a backend you restart all the time.
Still open: the kill switch only reaches browsers that come back
The kill switch runs only when a browser holding the old worker navigates to the origin again. A laptop that stays closed for the release window still has vodou-v313 cached when it wakes up, and /classic/ will be gone by then. That browser picks up the kill switch on its first navigation, so it recovers, but its first load comes from the old cache. The Runs fallback is also partial: it reads the scheduler ledger only when the graph ledger is empty, so a skill with rows in both still shows half its history. I know about that gap and haven’t fixed it yet.