debugging
31 posts.
-
Stream resume cursors don't survive a server restart
After a restart, an in-memory replay buffer is gone but clients still resume against it. Seq dedupe silently drops text. Reject foreign cursors by epoch.
-
Removing register() won't remove your old service worker
Our redesign dropped its service worker, but the old /sw.js would still serve a cached index.html, unseen in server logs. The fix had to live at the old URL.
-
No ANTHROPIC_API_KEY set: two code paths disagreed on null
A gateway deleted a valid Anthropic key because its cleanup branch read only the DB provider row while its detector also read env. How to test for split readers.
-
Another extension's iframe can lock your chrome.debugger agent out of a tab
Chrome refuses a debugger attach to another extension's chrome-extension:// frame. If your CDP agent auto-attaches to every iframe, skip those targets first.
-
ENOTFOUND for one API host while your health check says 200
Our agent went silent with ENOTFOUND on api.anthropic.com while health checks passed. Why one stuck hostname hides from liveness probes, plus a 5-minute check.
-
Pointing HTTPS_PROXY at port 1 only tests the fast failure
ECONNREFUSED returns in milliseconds, so a proxy-at-port-1 offline test never runs your timeout path. Test DNS failure and black holes too, with a time budget.
-
A missing settings row deleted my ANTHROPIC_API_KEY at boot
An absent SQLite settings row made our LLM gateway delete an env var the operator set. Why env-over-DB advice misses it, and a five-minute check for your stack.
-
An empty RAG query should raise, not return confident junk
A failed content-script probe sent an empty query to memory search, which fell back to captured turns and returned confident, unrelated results. Reject it.
-
Your load test measured the watchdog, not the server
A bash wait-plus-backgrounded-curl harness reported 85s and two hangs for a server whose real 4-concurrent latency was 6.5s. Per-request timeouts found it.
-
Near miss: your deploy script ships your working tree, not a commit
A deploy script that rsyncs your working directory ships half-finished edits and runs live migrations. Here is the invariant, and a check you can run in five minutes.
-
rsync --delete from a stale build/ rolls production back
A local build/ folder three policy versions old nearly overwrote live files and deleted server-only ones. Why dry runs miss it, and a five-minute check.
-
Your intent router is reading the attachment, not the user
An attached doc made our router fire a web search nobody asked for. Why routing on inlined context is a prompt-injection bug, and a five-minute check for yours.
-
Your agent memory stores vendor metrics with no expiry date
A 12-month keyword average of 2,400 hid an August value of 8,100. Agent memory kept both as timeless facts. Here is the invariant and a five-minute SQL check.
-
My skill miner learned from my cron jobs: 8 of 26 drafts
A recurrence-mining skill proposer built 8 of 26 drafts from my own cron and heartbeat turns. The fix was a write-time actor column, not a higher threshold.
-
Your config editor saves the template, not the file that runs
A settings editor returned ok on every save while the scheduler read a different copy of the file. The bug class, the invariant, and a five-minute stat check.
-
Your prompt fence is hardcoded to one source, not nine
Nine context lanes declare untrusted provenance; only one gets a fence in the prompt. Why a per-source trust label goes decorative, and the grep that finds it.
-
Your commit message claimed a call path that doesn't exist
A tool fan went from 770ms wall to 633ms against 1398ms of work. The commit said it fixed the chat path. The changed code had two callers, both CLI-only.
-
25 budgets declared, 3 enforced: the fallback was Infinity
A prompt assembler resolved undeclared lanes to Infinity, so a 600-token cap in the registry never fired once. The fix is a test that joins the two lists.
-
Your token budget is a comment: 8,000 declared, 12,000 shipped
A registry declared an 8,000-char budget, the code enforced 12,000, and the log recorded 14,188 as proof it worked. Nothing ever joined the three numbers.
-
Your memory extractor tags boat repair as a codebase gotcha
A memory pipeline tagged a ChatGPT thread about a stuck throttle as an engineering GOTCHA and ranked it first. Root cause, the class, and a 5-minute SQL check.
-
The dedupe key was set before the fetch, so it never retried
One connection refusal during gateway boot silenced the side panel for that URL forever: the dedupe key was committed before the request settled.
-
Your source-trust score is grading the pipe, not the source
Two records about one announcement, two different dates, both stored at trust_mult 1.0. The trust multiplier keyed on ingest channel, not on who asserted the claim.
-
Every stdio MCP server you spawn inherits your whole .env
A host that loads .env into its own process hands every credential to every stdio tool server it spawns. Per-server env config was a no-op. How to check yours in five minutes.
-
Your reranker fell back to cosine and the floor kept firing
A cross-encoder that fails to load doesn't throw. It leaves raw cosine in place, and a relevance floor tuned on logits keeps passing junk. Here's the check.
-
Two writers, one Markdown file: your appends get misfiled
Two processes appended to the same daily Markdown file. Both writes were atomic, nothing was lost, and every later fact was filed under the wrong heading.
-
Your retired button is still mounted, still handling clicks
A refactor left an old injected control alive in the page. It called the retired handler, threw no error, and looked exactly like the button that replaced it.
-
Your isolated test lab isn't isolated: 3 rows, 8 daemons
My failure-injection harness wrote to the live database, health-checked a stranger's process, and leaked 8 daemons. Three bugs in the test, none in the code.
-
Node execFile was 2.7x slower: the spawn cost nothing
Two implementations of the same memory lookup, one over a Unix socket and one shelling out to a CLI. I blamed the subprocess. The timing block said otherwise.
-
Why My LLM Agent Fabricated Numbers From Stale Context
An agent replayed a ten-minute-old CPU reading as live. The bug wasn't the model: a guard checked whether cached tool output existed, not whether it was fresh.
-
cached_tokens is 0 because your system prompt isn't stable
Our provider reported cached_tokens 0 of ~16K prompt tokens every turn. The cause: per-turn memory glued into the system prompt. Freeze the prefix instead.
-
SQLite FTS5 quietly ANDs your search terms
Our search box returned zero results for any sentence longer than four words. The bug wasn't ranking or embeddings: it was one space in a join() call.