Building Vodou in public.
I'm Chad Priest. I'm building Vodou, an AI operating system with persistent local memory, MCP orchestration, and a Rust engine. These are the bugs I actually hit, the dead ends I actually walked into, and the numbers I actually measured.
-
Agent memory leaks PII your redaction regex can't see
An agent writing public posts from a personal memory store can paste in names, pets and phone numbers. Pattern-based PII scanners miss most of them. Here is the check.
-
Never declare a side-effect flag in your tool's inputSchema
A declared boolean `open` flag got auto-filled to true and opened browser tabs all day. Side-effect flags belong in the handler, never in the LLM-visible schema.
-
Your agent's safety gate guards an adapter, not the token
A safety gate inside one tool adapter only binds callers of that adapter. Any route holding the same token skips it. Here's how to find the ungated routes in yours.
-
One adapter holds your write gate; every route holds the key
An agent route with no booking tool queried a live API with curl and could read a usable token. A gate in one adapter binds one route; the credential binds all.
-
Stream resume cursors don't survive a server restart
After a restart, an in-memory replay buffer is gone but clients still resume against it. Seq dedupe silently drops text. Reject foreign cursors by epoch.
-
Removing register() won't remove your old service worker
Our redesign dropped its service worker, but the old /sw.js would still serve a cached index.html, unseen in server logs. The fix had to live at the old URL.
-
No ANTHROPIC_API_KEY set: two code paths disagreed on null
A gateway deleted a valid Anthropic key because its cleanup branch read only the DB provider row while its detector also read env. How to test for split readers.
-
Your approval gate lives in a tool your agent can skip
An agent route without our booking tool hit Resy's API with curl. Approval gates in one tool adapter don't bind other routes; put them at the action boundary.
-
Your old frontend's service worker will outlive your redesign
Moving a redesigned agent console to '/' broke two things no diff showed: a cached service worker, and stream cursors that outlived a server restart.
-
Another extension's iframe can lock your chrome.debugger agent out of a tab
Chrome refuses a debugger attach to another extension's chrome-extension:// frame. If your CDP agent auto-attaches to every iframe, skip those targets first.
-
ENOTFOUND for one API host while your health check says 200
Our agent went silent with ENOTFOUND on api.anthropic.com while health checks passed. Why one stuck hostname hides from liveness probes, plus a 5-minute check.
-
Pointing HTTPS_PROXY at port 1 only tests the fast failure
ECONNREFUSED returns in milliseconds, so a proxy-at-port-1 offline test never runs your timeout path. Test DNS failure and black holes too, with a time budget.
-
A missing settings row deleted my ANTHROPIC_API_KEY at boot
An absent SQLite settings row made our LLM gateway delete an env var the operator set. Why env-over-DB advice misses it, and a five-minute check for your stack.
-
An empty RAG query should raise, not return confident junk
A failed content-script probe sent an empty query to memory search, which fell back to captured turns and returned confident, unrelated results. Reject it.
-
Your load test measured the watchdog, not the server
A bash wait-plus-backgrounded-curl harness reported 85s and two hangs for a server whose real 4-concurrent latency was 6.5s. Per-request timeouts found it.
-
Near miss: your deploy script ships your working tree, not a commit
A deploy script that rsyncs your working directory ships half-finished edits and runs live migrations. Here is the invariant, and a check you can run in five minutes.
-
rsync --delete from a stale build/ rolls production back
A local build/ folder three policy versions old nearly overwrote live files and deleted server-only ones. Why dry runs miss it, and a five-minute check.
-
git commit ships the whole index, not the paths you staged
Parallel AI agents in one git worktree share .git/index, so a careful git add still commits another agent's work. How it happened twice, and a private-index fix.
-
Your intent router is reading the attachment, not the user
An attached doc made our router fire a web search nobody asked for. Why routing on inlined context is a prompt-injection bug, and a five-minute check for yours.
-
Our memory ceiling logged ABORT while the file kept growing
A size ceiling enforced in one of three writers logged ABORT every cycle while the file grew. Five memory bugs where the guard lived in the wrong place.
-
Your agent memory stores vendor metrics with no expiry date
A 12-month keyword average of 2,400 hid an August value of 8,100. Agent memory kept both as timeless facts. Here is the invariant and a five-minute SQL check.
-
Chrome refuses an iframe silently, and onload still fires
A shipped panel proxied to an opt-in process, then went blank behind frame-ancestors. Two failure classes, one five-minute check you can run on your own stack.
-
A 'new' badge is a lie unless something records when you looked
A sidebar of scheduled agents needs a per-viewer watermark, written when the list is actually on screen. Plus the surrogate-pair bug that ate an emoji title.
-
A trust pin with no undo is an outage you shipped on purpose
An extension-ID pin hardened our browser bridge and could refuse every other browser on the machine, permanently. What that lockout taught me about auto-trust.
-
My skill miner learned from my cron jobs: 8 of 26 drafts
A recurrence-mining skill proposer built 8 of 26 drafts from my own cron and heartbeat turns. The fix was a write-time actor column, not a higher threshold.
-
Your config editor saves the template, not the file that runs
A settings editor returned ok on every save while the scheduler read a different copy of the file. The bug class, the invariant, and a five-minute stat check.
-
A parked human-in-the-loop prompt needs a clock, or a restart is your only exit
A workflow menu waited three days for a digit until a gateway restart freed it. Parked asks need a lazy TTL, a mismatch counter, and control words everywhere.
-
Your prompt fence is hardcoded to one source, not nine
Nine context lanes declare untrusted provenance; only one gets a fence in the prompt. Why a per-source trust label goes decorative, and the grep that finds it.
-
Your commit message claimed a call path that doesn't exist
A tool fan went from 770ms wall to 633ms against 1398ms of work. The commit said it fixed the chat path. The changed code had two callers, both CLI-only.
-
25 budgets declared, 3 enforced: the fallback was Infinity
A prompt assembler resolved undeclared lanes to Infinity, so a 600-token cap in the registry never fired once. The fix is a test that joins the two lists.
-
Your MCP audit table probably has no producer
The audit table existed since migration 046 and held zero rows for its whole life. What I found instrumenting both MCP client stacks, and three checks for yours.
-
Your token budget is a comment: 8,000 declared, 12,000 shipped
A registry declared an 8,000-char budget, the code enforced 12,000, and the log recorded 14,188 as proof it worked. Nothing ever joined the three numbers.
-
Nine agents in one worktree and each one knew only the branch name
How a daemon tells a coding session when a peer touched the same files: keyed on the transcript not the pid, silent unless the sets intersect.
-
The trust label your registry declares never reaches the model
Two memory searches per turn with no shared id, a THIS BLOCK WINS block that a budget could evict, and a trust field only a commit guard read. What fixing all three took.
-
The runs panel said 'No runs recorded yet'. The ledger had every run.
A scheduler wrote a run ledger nobody served, a skill prompt dropped the user's words, and a spawnSync stalled the gateway 9.6 s. Three checks for your own stack.
-
A console leaked one CSS class and hid 11 days of receipts
Two ways an agent console lied during a redesign: a view's opt-out class nothing removed, and a receipt table no page read. With a five-minute check for each.
-
Your memory extractor tags boat repair as a codebase gotcha
A memory pipeline tagged a ChatGPT thread about a stuck throttle as an engineering GOTCHA and ranked it first. Root cause, the class, and a 5-minute SQL check.
-
Your first-run path is the only code that runs with defaults off
A liveness ping answered by the transport, and a demo riding a toggle that ships off. Two failures in one first-run path, plus the check to run on yours.
-
Two writers, one judge: how a review queue became 90% noise
A memory-conflict queue hit 1,377 rows on a vault with 33 real conflicts. One producer had an LLM judge and dismissed 97% of its own work. The other had none.
-
The dedupe key was set before the fetch, so it never retried
One connection refusal during gateway boot silenced the side panel for that URL forever: the dedupe key was committed before the request settled.
-
Your message table is storing your UI, not your transcript
Four producers wrote into one turn and two baked HTML into storage. What it took to make the turn receipt a real event, and the SQL to check your own store.
-
A skill that finishes into a tab nobody has open did not finish
Console-mode skills completed into a workbench nobody was looking at. How results got routed to the surface the user is on, and the sibling-tree trap that cost two days.
-
The renderer existed. Nothing on that surface ever called it.
A side panel drew agent plans perfectly and could not run one. Four presses found four bugs a green suite could not see, plus the delivery check that catches them.
-
Your source-trust score is grading the pipe, not the source
Two records about one announcement, two different dates, both stored at trust_mult 1.0. The trust multiplier keyed on ingest channel, not on who asserted the claim.
-
66,570 characters reached the model and nothing logged them
I built a per-turn event log that rebuilds the exact request we sent, then measured the bytes it could not account for. The first number was 66,570.
-
Prompt injection defense that survives to turn 40
Page text an agent reads can outlive the turn as a stored memory and come back as trusted context. Two invariants: separate fields, and a turn-scoped approval gate.
-
Your agent's tool manifest is a comment until something reads it
A scheduled agent declared six MCP tools, called none of them, and reported ok. Enforcing that declaration at fire time cut selection from 1-of-942 to 1-of-6.
-
Your MCP allowlist controls tool names, not what they return
A read-only MCP profile still returned my cwd and MEMORY.md, because the disclosure was in the result envelope, not the tool list. How to check yours.
-
Every stdio MCP server you spawn inherits your whole .env
A host that loads .env into its own process hands every credential to every stdio tool server it spawns. Per-server env config was a no-op. How to check yours in five minutes.
-
A rule at byte 37,367 is not a rule if the cap is 32 KiB
Five AI coding hosts read five rules files in one repo. Mine had drifted for months, and the rule that mattered most sat past Codex's 32 KiB context cap.
-
Your link checker read planned downtime as a broken URL
A link checker flagged 7 correct references to a host I stopped on purpose, then offered to rewrite them. The fix: declare intended state, with an expiry.
-
Two tables both called 'skill', and nothing knew which was which
Vodou had 160 file skills and 15 console skills sharing one word. Every feature picked a table and called it the truth. Here is the seam and how to find yours.
-
Your reranker fell back to cosine and the floor kept firing
A cross-encoder that fails to load doesn't throw. It leaves raw cosine in place, and a relevance floor tuned on logits keeps passing junk. Here's the check.
-
Two writers, one Markdown file: your appends get misfiled
Two processes appended to the same daily Markdown file. Both writes were atomic, nothing was lost, and every later fact was filed under the wrong heading.
-
Your MEMORY.md is a file nobody writes to. Render it instead.
An always-injected memory file got one bullet in four months while the store grew to 42k chunks. Rendering it per session from the DB, and what valid_at fixed.
-
A 400-byte cap crashed my memory daemon on one emoji
Two memory bugs with the same shape: a truncation that counted bytes, and a fact verifier whose 'approved everything' looked identical to 'never ran'. With a five-minute check for your own stack.
-
Your document chunker is a memory chunker, and it quadruples your store
One 15,869-char file became 181 chunks. Fixing that exposed a scoring floor applied to two incomparable scales. Two invariants, one SQL check, for any RAG stack.
-
Per-site memory off switches only work if every reader asks the same authority
Governing an agent's page memory took three enforcement points, a soft delete with a dry run, and a live test that found four defects the unit suites missed.
-
Your Agent's Memory Has No Idea Where It Was
Agent memory stores what was said and drops where it happened. Adding a page axis to a memory store took four live-only defects and a per-site permission model.
-
Your retired button is still mounted, still handling clicks
A refactor left an old injected control alive in the page. It called the retired handler, threw no error, and looked exactly like the button that replaced it.
-
Your agent's approval gate is probably just a warning label
We shipped plan cards, real parallel tool calls and an approval gate for an MCP agent. The gate was decoration for two days, and the fan was never parallel.
-
A silent failure looks exactly like a feature you never built
Shipping an agent surface into 22 chat sites we don't own: the Chrome gesture that doesn't survive an await, and three builds all claiming one version number.
-
Every copy of the rule agreed. That was the bug.
Four surfaces minted the same token four different ways, and all four agreed. How we found the drift, why the tests couldn't see it, and the guard that now can.
-
Your isolated test lab isn't isolated: 3 rows, 8 daemons
My failure-injection harness wrote to the live database, health-checked a stranger's process, and leaked 8 daemons. Three bugs in the test, none in the code.
-
Node execFile was 2.7x slower: the spawn cost nothing
Two implementations of the same memory lookup, one over a Unix socket and one shelling out to a CLI. I blamed the subprocess. The timing block said otherwise.
-
Why My LLM Agent Fabricated Numbers From Stale Context
An agent replayed a ten-minute-old CPU reading as live. The bug wasn't the model: a guard checked whether cached tool output existed, not whether it was fresh.
-
cached_tokens is 0 because your system prompt isn't stable
Our provider reported cached_tokens 0 of ~16K prompt tokens every turn. The cause: per-turn memory glued into the system prompt. Freeze the prefix instead.
-
SQLite FTS5 quietly ANDs your search terms
Our search box returned zero results for any sentence longer than four words. The bug wasn't ranking or embeddings: it was one space in a join() call.