building vodou.

Stream resume cursors don't survive a server restart

After a restart, an in-memory replay buffer is gone but clients still resume against it. Seq dedupe silently drops text. Reject foreign cursors by epoch.

Chad Priest / / 6 min read

A resume cursor is a number that only means something to the process that issued it. If your server streams LLM tokens with resumable offsets and those offsets come from process state, a restart leaves every connected client holding a cursor into a buffer that no longer exists. That’s the whole bug. Nothing throws. The text on screen is just wrong.

Six characters missing, then a reply that started 150 characters in

After a gateway restart, one chat reply came up 6 characters short. The next one started roughly 150 characters in. I didn’t record the exact count, and I didn’t log the seq numbers either, which turned out to be its own lesson. The copies in the database were complete. For a while I blamed the renderer. That was wrong: the renderer drew exactly what it was given.

The client keeps a per-conversation high-water mark and drops replays it has already seen:

const prev = _lastSeq.get(msg.conversationId) || 0;
if (msg.seq <= prev) return;

On a single process that’s correct. Here’s what a restart does to it. Say the client’s mark for a conversation is N. The new process numbers that conversation’s events from wherever its own state says to start. Call it M. Ours hydrated from a buffer that a 10-minute retention trims, so after a quiet stretch M sat well below N. The new process emits M+1, M+2, …. Every event numbered ≤ N fails the check and gets silently dropped. The mark doesn’t move until an event finally clears N.

So the loss is exactly the text carried by the first N − M events of the new stream. Not every event carries text: tool status and other frames take seqs too. The same gap can cost a few characters or a few sentences, depending on what happened to sit in those slots. Each conversation has its own N and its own M, so each one loses a different amount. One loss of 6 characters and another of about 150 is the shape this mechanism produces. Without the seqs I can’t show you the actual N and M. That’s why I log them now.

The cursor also lived in two places. The socket layer had _lastSeq and the chat view had its own _seenSeq map doing the same dedupe, so the fix had to land twice.

On reconnect the client sent {type:'resume', conversationId, lastSeq} and the server took it without complaint.

What I changed, and the half I didn’t

Every connected handshake now carries a process epoch. The client compares it to the last epoch it saw. If they differ, it clears both cursor maps and does not send a resume at all. I also moved the resume from socket-open to after the handshake, because you can’t judge a cursor before you know who issued it.

The rule I’m holding myself to:

A resume cursor must name the generation that issued it, and a cursor from another generation must end in a reset, never a replay.

My fix meets that rule for my client. The server doesn’t yet. The handler still accepts a bare {type:'resume', lastSeq} with no epoch and replays any buffered events with seq > lastSeq. Going by the handler code, a stale cursor gets back either nothing or the wrong events, and never an error. Any other client, or an old tab running pre-fix JS, can still walk straight into the bug. Enforcing it on the server is simple: put the epoch in the resume, compare it to the server’s own, and answer a mismatch with {type:'reset', epoch} instead of a replay. I haven’t shipped that, so I won’t pretend I have.

Mine uses Date.now() at boot as the epoch. That works on a single instance. With several instances behind a load balancer, two can boot in the same millisecond and look like one generation. Use crypto.randomUUID() there, or anything else that can’t collide.

Recovering the reply that was mid-stream

Clearing cursors stops the corruption. It doesn’t bring back the reply that was streaming when the process died. Right after connected, the server sends a history snapshot from the database, and the chat repaints from that. Whatever was stored, the user sees. Nothing resumes generation, because the process doing the generating is gone. If the reply was cut off, the user sees it cut off and has to ask again. I haven’t tested exactly what a reply killed mid-generation leaves in its row. If you apply this fix, test that before your users find out for you: kill the server mid-reply, reconnect, and look at what the snapshot gives back.

Kafka shipped this invariant years ago

This is a stale cursor accepted against a different generation of the log, and Kafka has two production answers to it.

KIP-516 gave topics IDs because a topic name isn’t an identity. Delete a topic and recreate it under the same name, and a client holding an offset into the old one would otherwise read the new one as if nothing happened. That’s my restart exactly: same conversation ID, different buffer.

KIP-101 and KIP-320 brought in leader epochs because an offset can still be in range and point at a truncated, rewritten log. A consumer checks its offset together with the epoch it came from, and when the epoch doesn’t match it truncates and resets. A number in a valid range doesn’t make it the right number.

SSE has the same hole. Last-Event-ID against a per-process counter fails the same way, and it’s worse than the stateless case. A new instance with no buffer has nothing to replay. A new instance with a different buffer answers with plausible events.

Persisting the counter helps but isn’t enough. Ours effectively was persisted, since it hydrated from a stored buffer, and it reset anyway once retention emptied that buffer. Failovers, wiped volumes and restored backups do the same thing. A counter tells you how far into the stream you are. It doesn’t tell you which stream.

Restart mid-stream and check what your server says

Quote every URL. zsh treats an unquoted ? as a glob and aborts with no matches found.

SSE:

curl -sN 'http://localhost:3000/stream?conv=t1' | tee before.log   # Ctrl-C mid-reply
LAST=$(grep '^id:' before.log | tail -1 | cut -d' ' -f2)
# restart the server process
curl -si -N -H "Last-Event-ID: $LAST" 'http://localhost:3000/stream?conv=t1' | head -20

WebSocket (adjust the path and message shape to match your protocol):

websocat 'ws://localhost:3000/ws' | tee before.log   # send a prompt, Ctrl-C mid-reply
LAST=$(grep -o '"seq":[0-9]*' before.log | tail -1 | cut -d: -f2)
# restart the server process
(echo "{\"type\":\"resume\",\"conversationId\":\"t1\",\"lastSeq\":$LAST}"; sleep 5) \
  | websocat 'ws://localhost:3000/ws'

Pass: the server rejects or resets the stale cursor. For SSE that’s a 409, a 410 or an explicit reset event. For WS it’s a reset or error frame. In both cases the server carries a generation marker that visibly changed across the restart.

Fail: a 200 with id: 1, or WS frames with low seqs, or silence. Silence counts as a fail, because your client can’t tell “nothing to replay” from “wrong buffer”.

Not a pass: ids or seqs that continue above $LAST. That only means your persisted counter survived this restart. Run the test again after whatever trims or replaces that state: retention expiry, failover, a fresh volume. If there’s no generation marker, you’ve only put the failure off.

My own server fails the WebSocket check today. Going by the handler code, it would answer silently.

For the end-to-end symptom, send one more prompt through your real client after a restart and concatenate what rendered. Then compare it to the stored row (SELECT content FROM messages ORDER BY id DESC LIMIT 1;). If the lengths differ, you have this bug.