Pointing HTTPS_PROXY at port 1 only tests the fast failure
ECONNREFUSED returns in milliseconds, so a proxy-at-port-1 offline test never runs your timeout path. Test DNS failure and black holes too, with a time budget.
My offline test was green, and it only covered the kind of offline that shows up in a test.
The setup was the one you have probably written yourself. I wanted to know what the boot path does when the machine can’t reach the internet, so I set one env var:
https_proxy=http://127.0.0.1:1
Nothing listens on port 1. Every outbound HTTPS request goes to the proxy, the kernel answers with a reset, and the client throws connect ECONNREFUSED 127.0.0.1:1. The boot path caught it, fell back to its offline state, and the test passed. I wrote it down as the way to simulate offline, and I believed that note for longer than I’d like to admit.
ECONNREFUSED 127.0.0.1:1 comes back before any timer starts
A refused connection is the most polite failure a network has. The error arrives in milliseconds, before any connect timeout, retry backoff or request deadline has a chance to fire. So the test proved that the catch block works. It said nothing about how long the user waits.
Real offline mostly looks different. On hotel Wi-Fi with a captive portal, on a laptop that just woke up, on a VPN that dropped its route, you get one of two things. Either DNS stalls or fails, or the SYN leaves and nothing ever comes back. That second case hangs until some timer somewhere gives up, and that timer is the part of the code my test never reached. My note from today says it plainly: port 1 tests the fast-fail case, and a silent drop is a different code path.
Refused, unresolvable and black-holed are three different code paths
Any client that calls a network dependency can fail in at least three ways: a fast refusal, a DNS lookup that fails or stalls, and a connection that is silently dropped and sits there until a timeout. A test that aims the proxy at a closed local port covers only the first. You will find the same gap in Node fetch/undici agents behind HTTPS_PROXY, in LLM SDK wrappers that stack their own retries on top of a long default timeout, in MCP servers and agent tools that call remote APIs, and in Electron apps with an “offline mode” banner that shows up 40 seconds late.
An offline test covers a failure mode only if it makes the client wait, and it passes only if the user-visible fallback arrives within a stated time budget.
That rule can be checked. A test with no time assertion can’t catch a hang, however many failure modes you feed it.
Toxiproxy’s timeout toxic sits after DNS has already answered
The standard advice is Toxiproxy: put a TCP proxy in front of the dependency and add toxics. That advice is right about the black hole. Its timeout toxic stops data and holds the connection open, which is exactly the hang that port 1 skips. Where it falls short is that your app connects to Toxiproxy by a local address you configured, so the name was resolved before any toxic ran. A stalled resolver never gets tested. Comcast gets closer, since it drops packets at the OS level, but it needs root and packet-filter rules, and you won’t wire that into a 30-second unit run. You need a cheaper test that covers all three.
Time your client against refused, unresolvable and black-holed
Run this against your own dependency. It needs curl, plus Docker for the last case.
URL=https://api.openai.com/v1/models # any HTTPS dependency your app calls
run() { local s=$(date +%s); "$@" >/dev/null 2>&1; echo "exit=$? secs=$(( $(date +%s) - s ))"; }
# 1. Refused: what most offline tests simulate
run curl -sS -x http://127.0.0.1:1 "$URL" # exit=7 secs=0
# 2. Unresolvable name: fast NXDOMAIN
run curl -sS https://api.example.invalid/ # exit=6 secs=0
# 3. Black hole: non-routable proxy, SYN goes out, nothing returns
run timeout 30 curl -sS -x http://10.255.255.1:3128 "$URL" # exit=124 secs=30
# 4. Resolver that never answers
run timeout 30 docker run --rm --dns 10.255.255.1 curlimages/curl -sS "$URL"
# Now the real check: your app's offline path under case 3, with a budget
HTTPS_PROXY=http://10.255.255.1:3128 timeout 10 npm test -- --grep offline; echo "exit=$?"
Cases 1 and 2 come back instantly on every machine. That’s the baseline, and it’s what your existing test already knows. Cases 3 and 4 show how long a client with no deadline waits. The last line is the test that counts. A pass exits 0 inside the budget, because your code hit its own connect timeout and showed the offline state. A fail exits 124: timeout killed the run while your client was still waiting on a socket that will never answer. If you use an LLM SDK, check its defaults before you pick a budget. Several ship with request timeouts measured in minutes, with retries on top.
One more trap: Node’s native fetch doesn’t read HTTPS_PROXY unless you give it a proxy agent. If your port-1 test passed anyway, your requests may never have gone through the proxy at all, and the test proved even less than mine did.
So: when an offline test passes in under a second, check whether it ever made the client wait. If it didn’t, add a black-holed case and a time budget.