llm
20 posts.
-
Never declare a side-effect flag in your tool's inputSchema
A declared boolean `open` flag got auto-filled to true and opened browser tabs all day. Side-effect flags belong in the handler, never in the LLM-visible schema.
-
One adapter holds your write gate; every route holds the key
An agent route with no booking tool queried a live API with curl and could read a usable token. A gate in one adapter binds one route; the credential binds all.
-
No ANTHROPIC_API_KEY set: two code paths disagreed on null
A gateway deleted a valid Anthropic key because its cleanup branch read only the DB provider row while its detector also read env. How to test for split readers.
-
Your approval gate lives in a tool your agent can skip
An agent route without our booking tool hit Resy's API with curl. Approval gates in one tool adapter don't bind other routes; put them at the action boundary.
-
Your intent router is reading the attachment, not the user
An attached doc made our router fire a web search nobody asked for. Why routing on inlined context is a prompt-injection bug, and a five-minute check for yours.
-
Your agent memory stores vendor metrics with no expiry date
A 12-month keyword average of 2,400 hid an August value of 8,100. Agent memory kept both as timeless facts. Here is the invariant and a five-minute SQL check.
-
Your token budget is a comment: 8,000 declared, 12,000 shipped
A registry declared an 8,000-char budget, the code enforced 12,000, and the log recorded 14,188 as proof it worked. Nothing ever joined the three numbers.
-
Nine agents in one worktree and each one knew only the branch name
How a daemon tells a coding session when a peer touched the same files: keyed on the transcript not the pid, silent unless the sets intersect.
-
The trust label your registry declares never reaches the model
Two memory searches per turn with no shared id, a THIS BLOCK WINS block that a budget could evict, and a trust field only a commit guard read. What fixing all three took.
-
The runs panel said 'No runs recorded yet'. The ledger had every run.
A scheduler wrote a run ledger nobody served, a skill prompt dropped the user's words, and a spawnSync stalled the gateway 9.6 s. Three checks for your own stack.
-
Your message table is storing your UI, not your transcript
Four producers wrote into one turn and two baked HTML into storage. What it took to make the turn receipt a real event, and the SQL to check your own store.
-
66,570 characters reached the model and nothing logged them
I built a per-turn event log that rebuilds the exact request we sent, then measured the bytes it could not account for. The first number was 66,570.
-
A rule at byte 37,367 is not a rule if the cap is 32 KiB
Five AI coding hosts read five rules files in one repo. Mine had drifted for months, and the rule that mattered most sat past Codex's 32 KiB context cap.
-
Two tables both called 'skill', and nothing knew which was which
Vodou had 160 file skills and 15 console skills sharing one word. Every feature picked a table and called it the truth. Here is the seam and how to find yours.
-
Your MEMORY.md is a file nobody writes to. Render it instead.
An always-injected memory file got one bullet in four months while the store grew to 42k chunks. Rendering it per session from the DB, and what valid_at fixed.
-
A 400-byte cap crashed my memory daemon on one emoji
Two memory bugs with the same shape: a truncation that counted bytes, and a fact verifier whose 'approved everything' looked identical to 'never ran'. With a five-minute check for your own stack.
-
Your document chunker is a memory chunker, and it quadruples your store
One 15,869-char file became 181 chunks. Fixing that exposed a scoring floor applied to two incomparable scales. Two invariants, one SQL check, for any RAG stack.
-
Every copy of the rule agreed. That was the bug.
Four surfaces minted the same token four different ways, and all four agreed. How we found the drift, why the tests couldn't see it, and the guard that now can.
-
Why My LLM Agent Fabricated Numbers From Stale Context
An agent replayed a ten-minute-old CPU reading as live. The bug wasn't the model: a guard checked whether cached tool output existed, not whether it was fresh.
-
cached_tokens is 0 because your system prompt isn't stable
Our provider reported cached_tokens 0 of ~16K prompt tokens every turn. The cause: per-turn memory glued into the system prompt. Freeze the prefix instead.