An empty RAG query should raise, not return confident junk
A failed content-script probe sent an empty query to memory search, which fell back to captured turns and returned confident, unrelated results. Reject it.
On 2026-07-29 I shipped a side panel for our Chrome extension. On its first live run I found what looked like two bugs. The panel could no longer insert text into the page. Memory search was also returning unrelated results. I had asked for my wife’s name, and I got a full list of well-formed memories about our dev work and nothing about her. I assumed the relevance problem was in retrieval, so I started there.
Retrieval was fine. When I ran the same question from the command line, the right fact came back at rank 1 with a score of 0.790. The two symptoms came from one bug, and the second symptom was the more dangerous one.
The bug was live for one run. The panel feature was committed at 23:04 that night and the fix at 23:13. I caught it quickly for one reason: I happened to ask a question whose answer I already knew. With a question I couldn’t check, I would have read the list and believed it.
I have to admit the gap in my evidence. I didn’t save the junk results or their scores before I fixed it. So I can’t set their numbers next to the 0.790. What I can tell you is what the panel showed. It asked for k=15 and got a full list, in the same layout and ranking as a real search. Nothing on screen marked it as different.
Insert dead, search wrong, and only one bug
The panel starts by probing the page’s content script. The probe returns the text around the cursor, and that text becomes the search seed. After an extension update, Chrome doesn’t re-inject content scripts into tabs that are already open, so the probe fails with Receiving end does not exist. That explained the dead insert.
The wrong search came from what happened next. A failed probe gave an empty seed, and the background script forwarded it as query: ''. This is the branch in the gateway that received it, trimmed:
// what the user typed > the conversation seed > the host
const seed = query.trim() || seedFromConversation(provider, convId); // last 6 captured turns, capped at 1500 chars
const q = seed || (host ? `context for ${host}` : '');
if (!q) return { ok: false, error: 'empty query' };
In generic form, the thing to look for in your own code is if (!query) query = recentTurns(). The painful part is the third line. The gateway did have an empty-query check. It came after two defaults, so it could never fire. An empty query became six turns of recent conversation, and those turns got embedded and searched as if the user had typed them. In a dev chat, that meant “memories about the dev chat,” returned with full confidence.
Nothing threw an error or wrote a warning to the log. The only visible symptom was bad relevance, and that is the symptom you blame on your embeddings.
Re-inject the content script once, and never send an empty query
First, the panel repairs itself when a probe fails. It re-injects the content script and retries once. I no longer tell the user to reload the tab, because that advice lets the bug come back after every update.
Second, and more important: the panel never sends an empty seed as a search. When the seed is empty, it says it couldn’t read the page and asks for a question. It doesn’t show 15 memories. The gateway branch above is still there. That was a deliberate feature for a picker opened with nothing typed. So the real rule isn’t “never default.” The rule is that a default nobody can see is a bug.
Elasticsearch made the same choice on purpose
The general failure: a retrieval layer turns missing input into a default query, so an upstream capture failure reaches the model as confident, wrong context instead of as an error. An LLM downstream can’t tell the difference. It writes a fluent answer on top of whatever it was handed.
Elasticsearch’s fix for an NPE in the standard retriever resolves an empty retriever body to a MatchAllDocsQuery, “as in standard search API.” That’s a reasonable call for a search box and a trap for an agent. I can cite one retriever that does this on purpose, not a survey. I’d bet it isn’t the only one.
RAGFlow shows the mirror image: passing [] instead of None for doc filters silently produced zero-recall retrieval. That gave empty output, not confident junk, so it’s the milder failure. I include it because the root cause is the same: a “no value” sentinel got read as an instruction, and nothing at the boundary asked which one the caller meant.
The standard advice, from pieces like How to Handle Fewer-Than-K Vector Search Results, is to “handle empty results explicitly.” That’s correct for empty output. It doesn’t cover empty input that produces full output. My results were never empty. The panel got the full k=15, every result was wrong, and a check on result count would have passed.
The invariant, stated so you can check it in any codebase: every retrieval call either receives a non-empty query from its caller, or fails at the boundary with a specific error. If any layer substitutes a default seed, the response marks it as defaulted (for example seed_source: "conversation"), and the caller can see that and act on it. System-initiated retrieval is fine. Unmarked substitution is the bug.
Checking your own retriever
1. Empty input must raise at the boundary. Call the function your agent actually uses on every turn, not a test helper:
import traceback
# swap `retrieve` for whatever your agent calls per turn
for q in ["", " ", None]:
try:
docs = retrieve(q)
print(f"FAIL {q!r}: returned {len(docs)} docs")
except ValueError as e:
where = traceback.extract_tb(e.__traceback__)[-1]
print(f"pass? {q!r}: ValueError({e}) raised in {where.name} ({where.filename}:{where.lineno})")
except Exception as e:
print(f"FAIL {q!r}: {type(e).__name__} is an accident, not validation")
A pass is a ValueError (or your own named validation error) raised in your retrieval entry point, so the printed function name is the one you called. A ValueError raised three frames down in a tokenizer doesn’t count. Neither does an AttributeError from None.strip() or a TypeError from an embedding client. Those only mean the code happened to crash, and the next refactor will turn them into a silent default. Any returned list is a fail, including an empty one, unless it carries a defaulted marker the caller checks. A result like FAIL '': returned 15 docs means you have a match_all path that nobody chose.
If you’re on LangChain or LlamaIndex, test the invoke your chain calls, not the vector store underneath it. Any substitution happens in the wrapper.
2. Look for the fallback in the source. Search near retrieval calls for the patterns that do this:
grep -rnE 'query *\|\||query or |if not query|if \(!query|\.trim\(\) *\|\||\.strip\(\) or ' src/ \
| grep -iE 'search|retriev|seed|context|embed|history|turns'
For each hit, ask whether the defaulted value is marked in the response. If it isn’t, that’s a bug.
3. On Elasticsearch, send the empty body yourself. First get the index size, then send a standard retriever with no query:
GET /your-index/_count
GET /your-index/_search
{
"track_total_hits": true,
"retriever": { "standard": {} }
}
If hits.total.value equals the count from _count, an empty query from your code reaches Elasticsearch as match_all. Then check whether your client can ever build that body when the user text is empty. Send it the empty string and log the request.
4. Separately, check for a score floor. Send "zq xv plomb 7731". A dense vector store with no score threshold returns top-k for any input, so most readers will fail this, including people who have no empty-query bug. That’s a different bug, and a real one: it’s the reason my junk list looked like a real one. A pass means everything comes back below a threshold you enforce, and gets dropped.
When check 1 fails, don’t tune relevance. Trace where the empty string came from. Mine came from a content script that wasn’t there.