# ENOTFOUND for one API host while your health check says 200

> Our agent went silent with ENOTFOUND on api.anthropic.com while health checks passed. Why one stuck hostname hides from liveness probes, plus a 5-minute check.

- Author: Chad Priest
- Published: 2026-10-01
- Canonical URL: https://blog.vodou.ai/enotfound-for-one-api-host-while-your-health-check-says-200/
- Tags: dns, node, debugging, ai

---

From 8:51 to 9:08 PM last night, every text I sent to my own assistant got the same answer back:

`API Error: Can't reach the API server (ENOTFOUND)`

The path is simple. A text comes in over iMessage, a relay forwards it to my Mac, and the gateway on the Mac spawns the Claude CLI to answer it. During those seventeen minutes the relay returned `200`, other sites loaded fine in the browser, and the Mac was obviously online. The only thing that failed was the name `api.anthropic.com`.

## `ENOTFOUND` for api.anthropic.com, `200` from the relay

I trusted the green check first, and that was the mistake. The relay health probe is a real request over a real network, so when it said `200` I read that as "the network is fine" and went looking for a bug in the reply path. The probe answered a smaller question than the one I was asking. It proved that the relay's hostname resolved. It said nothing about the one hostname the answer depended on.

`ENOTFOUND` is the error from `getaddrinfo`, and on macOS `getaddrinfo` goes through mDNSResponder and its cache. My best explanation is that the lookup for that one name got stuck in that layer while every other name resolved normally. I'm calling it the likely cause rather than the proven one because I never caught the cache in the act. There's one detail that matters, though: the gateway spawns a fresh CLI process for each message. No in-process cache survives from one attempt to the next, so whatever was stale lived below the process, in the OS.

## Flushing mDNSResponder is a cure, not a check

The standard advice is in every thread on this error, including [aws-sdk-js-v3 #5236](https://github.com/aws/aws-sdk-js-v3/issues/5236): `sudo dscacheutil -flushcache; sudo killall -HUP mDNSResponder`. That thread is also honest that the flush often didn't fix it, and that nobody could tell whether `getaddrinfo` or mDNSResponder was at fault. A flush is something you do after a human has already noticed. It doesn't tell you which layer is lying, and it does nothing for an agent that's failing at 9 PM while you're away from your desk.

[NousResearch/hermes-agent #49111](https://github.com/NousResearch/hermes-agent/issues/49111) is the closest match I found. There, a gateway's event loop held on to a failed lookup, the PID watchdog passed, the DNS watchdog passed, and only a restart helped. It gets the shape exactly right: a live process and dead providers, with no probe that notices. Where it doesn't fit my case is that its stale state sat inside one long-running process. Mine sat in a layer that a fresh process walks straight into. [cr0x's piece on caching layers](https://cr0x.net/en/dns-works-apps-still-fail/) puts it best: you debug this by finding the cache that's lying, and a liveness probe that uses a different hostname can't find it.

## A stuck name in the system resolver, behind a probe that resolves a different one

Here's the class with my stack removed. A service calls a third-party API through the host's system resolver. A single name fails to resolve while every other name works. Generic liveness checks hit a different host, or just check a PID, so they stay green while the service's real job is dead. The same thing shows up in Python `httpx` workers under systemd-resolved, in Node agents running as launchd jobs, and in Docker containers that forward to the host's DNS stub.

**A dependency health check is only valid if it resolves the same hostname, through the same resolver path, from the same execution context as the real call.**

To check your own system, run the script below as your agent runs: same user, same launchd or systemd unit, same container. Pass it every hostname your agent calls. `dns.lookup` goes through `getaddrinfo` and the OS cache, the same path your SDK uses. `dns.resolve4` asks the configured nameservers directly and skips that cache.

```js
// dns-split.js   usage: node dns-split.js api.anthropic.com api.openai.com example.com
const dns = require('node:dns');
for (const h of process.argv.slice(2)) {
  dns.lookup(h, { all: true }, (e1, sys) => {
    dns.resolve4(h, (e2, direct) => {
      console.log(h.padEnd(22),
        'system:', e1 ? e1.code : sys.map(a => a.address).join(','),
        '| direct:', e2 ? e2.code : direct.join(','));
    });
  });
}
// healthy:  api.anthropic.com  system: 160.79.x.x | direct: 160.79.x.x
// stuck:    api.anthropic.com  system: ENOTFOUND  | direct: 160.79.x.x
// real outage or bad config: both columns fail
```

If both columns resolve, you're fine today. A row that says `ENOTFOUND` in the system column and returns addresses in the direct column is this bug: DNS is healthy, and the layer between your code and DNS isn't. For Python, compare `socket.getaddrinfo` against `dnspython`. Then open your health check and look at which hostname it resolves. If that isn't your LLM provider's hostname, the probe is checking something else.

I don't have a structural fix yet, and I'm not going to pretend I do. What changed today is the probe. It now resolves the provider's hostname through the system path, from the same context the CLI runs in, and it alerts when that disagrees with a direct query. Seventeen minutes of silence came from a green check that tested the wrong hostname. Make your probe resolve the name your agent actually calls.

---

Source: [ENOTFOUND for one API host while your health check says 200](https://blog.vodou.ai/enotfound-for-one-api-host-while-your-health-check-says-200/) by Chad Priest, from Building Vodou in Public.
