building vodou.

ENOTFOUND for one API host while your health check says 200

Our agent went silent with ENOTFOUND on api.anthropic.com while health checks passed. Why one stuck hostname hides from liveness probes, plus a 5-minute check.

Chad Priest / / 4 min read

From 8:51 to 9:08 PM last night, every text I sent to my own assistant got the same answer back:

API Error: Can't reach the API server (ENOTFOUND)

The path is simple. A text comes in over iMessage, a relay forwards it to my Mac, and the gateway on the Mac spawns the Claude CLI to answer it. During those seventeen minutes the relay returned 200, other sites loaded fine in the browser, and the Mac was obviously online. The only thing that failed was the name api.anthropic.com.

ENOTFOUND for api.anthropic.com, 200 from the relay

I trusted the green check first, and that was the mistake. The relay health probe is a real request over a real network, so when it said 200 I read that as “the network is fine” and went looking for a bug in the reply path. The probe answered a smaller question than the one I was asking. It proved that the relay’s hostname resolved. It said nothing about the one hostname the answer depended on.

ENOTFOUND is the error from getaddrinfo, and on macOS getaddrinfo goes through mDNSResponder and its cache. My best explanation is that the lookup for that one name got stuck in that layer while every other name resolved normally. I’m calling it the likely cause rather than the proven one because I never caught the cache in the act. There’s one detail that matters, though: the gateway spawns a fresh CLI process for each message. No in-process cache survives from one attempt to the next, so whatever was stale lived below the process, in the OS.

Flushing mDNSResponder is a cure, not a check

The standard advice is in every thread on this error, including aws-sdk-js-v3 #5236: sudo dscacheutil -flushcache; sudo killall -HUP mDNSResponder. That thread is also honest that the flush often didn’t fix it, and that nobody could tell whether getaddrinfo or mDNSResponder was at fault. A flush is something you do after a human has already noticed. It doesn’t tell you which layer is lying, and it does nothing for an agent that’s failing at 9 PM while you’re away from your desk.

NousResearch/hermes-agent #49111 is the closest match I found. There, a gateway’s event loop held on to a failed lookup, the PID watchdog passed, the DNS watchdog passed, and only a restart helped. It gets the shape exactly right: a live process and dead providers, with no probe that notices. Where it doesn’t fit my case is that its stale state sat inside one long-running process. Mine sat in a layer that a fresh process walks straight into. cr0x’s piece on caching layers puts it best: you debug this by finding the cache that’s lying, and a liveness probe that uses a different hostname can’t find it.

A stuck name in the system resolver, behind a probe that resolves a different one

Here’s the class with my stack removed. A service calls a third-party API through the host’s system resolver. A single name fails to resolve while every other name works. Generic liveness checks hit a different host, or just check a PID, so they stay green while the service’s real job is dead. The same thing shows up in Python httpx workers under systemd-resolved, in Node agents running as launchd jobs, and in Docker containers that forward to the host’s DNS stub.

A dependency health check is only valid if it resolves the same hostname, through the same resolver path, from the same execution context as the real call.

To check your own system, run the script below as your agent runs: same user, same launchd or systemd unit, same container. Pass it every hostname your agent calls. dns.lookup goes through getaddrinfo and the OS cache, the same path your SDK uses. dns.resolve4 asks the configured nameservers directly and skips that cache.

// dns-split.js   usage: node dns-split.js api.anthropic.com api.openai.com example.com
const dns = require('node:dns');
for (const h of process.argv.slice(2)) {
  dns.lookup(h, { all: true }, (e1, sys) => {
    dns.resolve4(h, (e2, direct) => {
      console.log(h.padEnd(22),
        'system:', e1 ? e1.code : sys.map(a => a.address).join(','),
        '| direct:', e2 ? e2.code : direct.join(','));
    });
  });
}
// healthy:  api.anthropic.com  system: 160.79.x.x | direct: 160.79.x.x
// stuck:    api.anthropic.com  system: ENOTFOUND  | direct: 160.79.x.x
// real outage or bad config: both columns fail

If both columns resolve, you’re fine today. A row that says ENOTFOUND in the system column and returns addresses in the direct column is this bug: DNS is healthy, and the layer between your code and DNS isn’t. For Python, compare socket.getaddrinfo against dnspython. Then open your health check and look at which hostname it resolves. If that isn’t your LLM provider’s hostname, the probe is checking something else.

I don’t have a structural fix yet, and I’m not going to pretend I do. What changed today is the probe. It now resolves the provider’s hostname through the system path, from the same context the CLI runs in, and it alerts when that disagrees with a direct query. Seventeen minutes of silence came from a green check that tested the wrong hostname. Make your probe resolve the name your agent actually calls.