Skip to content

Stop one blocked address from reading as an internet outage - #7932

Open
johnpipi wants to merge 3 commits into
omacom:quattrofrom
johnpipi:network-status-probe-fallback
Open

johnpipi wants to merge 3 commits into
omacom:quattrofrom
johnpipi:network-status-probe-fallback

Conversation

@johnpipi

Copy link
Copy Markdown

Problem

omarchy-network-status measures internet reachability by pinging a single hardcoded address:

internet_probe=1.1.1.1

Some networks blackhole 1.1.1.1 outright. A number of ISPs and consumer routers used it internally before Cloudflare acquired it, and those blocks are still around. On such a network the address fails for every protocol, not just ICMP:

Target ICMP TCP 443 TCP 53 DNS
1.1.1.1 100% loss blocked blocked fails
1.0.0.1 (Cloudflare's own secondary) 0% loss, 3.9 ms open — —
8.8.8.8 0% loss, 4.5 ms — — ok
9.9.9.9 0% loss, 15.5 ms — — ok

The route is normal and nothing on the LAN claims the address, so this is an upstream block, not a local fault. The result is that the network panel permanently shows Ping: Timeout and Packet Loss: 100% on a working connection:

$ omarchy-network-status --verbose
router_ping_ms      1.05
internet_ping_ms                 <- empty, forever

Related: omarchy dns Cloudflare also sets 1.1.1.1 as the primary resolver, which is a dead server on these networks. That is a separate issue and not addressed here.

Change

Derive the probe targets from the resolvers the system is actually configured to use — via resolvectl dns, falling back to /etc/resolv.conf for hosts not running systemd-resolved — and fall back to a list of public addresses when none qualify.

Three details worth calling out:

Only public unicast IPv4 addresses qualify as probes. omarchy dns DHCP commonly yields a resolver inside the LAN. Probing that would report the internet as reachable whenever the router answered, which is a worse bug than the one being fixed. The systemd-resolved stub also listens on loopback. Private, loopback, link-local, CGNAT, and multicast/reserved ranges are all rejected.

Candidates are probed concurrently and the best answer wins. A sample still costs one ping timeout rather than one per address, and a single blocked address no longer blanks the reading.

Route lookup keeps its own fixed public address. internet_probe was doing double duty — ip route get used it to find the interface carrying default traffic. That lookup is answered from the local routing table, so reachability is irrelevant there, but pointing it at a configured resolver could return an on-link route instead of the default one. It is now a separate route_probe.

Result

On a network that blocks 1.1.1.1:

before:  internet_ping_ms                    (1.06s -- full timeout burned)
after:   internet_ping_ms      4.01          (0.07s)

Non-verbose output is byte-identical. A genuine outage still reports empty, verified with a stubbed ping that fails for every public address while the gateway answers.

Tests

Adds test/shell.d/network-status-probe-test.sh, covering:

  • public addresses accepted, including the RFC boundaries that must stay public (172.15.0.1, 172.32.0.1, 192.169.0.1, 100.63.0.1)
  • private, loopback, link-local, CGNAT, multicast, IPv6, and malformed input rejected
  • configured public resolvers used as probes
  • a LAN-only resolver falling back to public probes rather than being probed
  • IPv6-only resolvers falling back
  • more than one address probed
  • route lookup still using a dedicated fixed address
  • an unreachable internet still reporting no latency

The test fails against the current code and passes with the change. ./test/cli and the shell suite pass.

The network panel's reachability sample pinged a single hardcoded address,
1.1.1.1. Networks that blackhole that address -- ISPs and consumer routers
that used it internally before Cloudflare did -- left the panel reporting
"Ping: Timeout" and "Packet Loss: 100%" permanently, while the internet
worked fine.

Probe the resolvers the system is configured to use instead, falling back to
a list of public addresses. Only public unicast IPv4 addresses qualify: a
resolver on the LAN, which is what `omarchy dns DHCP` usually yields, would
report the internet as reachable whenever the router answered, and the
systemd-resolved stub listens on loopback.

Candidates are probed concurrently, so a sample still costs one ping timeout
rather than one per address, and reporting the best answer means a single
blocked address no longer blanks the reading.

Route lookup keeps its own fixed public address. It resolves against the
local routing table, where reachability is irrelevant, and pointing it at a
configured resolver could return an on-link route instead of the default one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings August 23, 2026 17:25

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Improves network status detection by probing multiple public DNS addresses concurrently instead of relying on one address.

Changes:

  • Selects configured public IPv4 resolvers with public fallbacks.
  • Separates route lookup from reachability probing.
  • Adds shell tests for probe selection and outage behavior.

Tip

If you aren't ready for review, convert to a draft PR.
Click "Convert to draft" or run gh pr ready --undo.
Click "Ready for review" or run gh pr ready to reengage.

Reviewed changes

Copilot reviewed 1 out of 2 changed files in this pull request and generated 2 comments.

File Description
bin/omarchy-network-status Adds probe selection and concurrent latency sampling.
test/shell.d/network-status-probe-test.sh Tests address filtering, fallback selection, and outage handling.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +42 to +45
probes=$(PATH="$stub:$PATH" bash -c "$(declare -f is_public_ipv4 configured_dns_servers internet_probes)
fallback_probes=(1.1.1.1 8.8.8.8 9.9.9.9)
max_probes=3
internet_probes" | tr '\n' ' ')
Comment on lines +89 to +92
eval "$(sed -n '/^ping_latency_ms()/,/^}$/p' "$STATUS")"
eval "$(sed -n '/^ping_internet_ms()/,/^}$/p' "$STATUS")"
workdir=$(mktemp -d)
result=$(PATH="$stub:$PATH" ping_internet_ms "$workdir")
omarchybot and others added 2 commits September 28, 2026 17:11
resolvectl prints a server as addr[:port][%ifname][#name], and omarchy dns writes its Cloudflare, Google and DNS-over-TLS custom servers with the #name, so is_public_ipv4 rejected every one and the configured resolvers were silently replaced by the fixed fallbacks. On a network that also blocks those, the user's own public resolver was never tried.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Codex Medium <noreply@openai.com>
Clash and Mihomo in fake-IP TUN mode hand out resolvers inside 198.18.0.0/15, and the proxy answers for them itself, so probing one would keep reporting the internet as reachable through an outage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Codex Medium <noreply@openai.com>
@omarchybot omarchybot added verified Omarchy Triage has verified that this issue is ready for final review ready Good to merge labels Sep 28, 2026
@omarchybot

Copy link
Copy Markdown
Collaborator

Reviewed and verified on a disposable Omarchy VM. The bug is real and this fixes it. I pushed two small fixes to your branch.

Reproduction. With nft dropping all output to 1.1.1.1 on the VM, omarchy-network-status --verbose on quattro (ef6d9e6) gave an empty internet_ping_ms on every sample. At this head it reads about 7 ms. With every public address blocked and only the gateway allowed, it still reads empty, so a real outage still shows as one. Unblocked, it reads normally, and the non-verbose status line is unchanged. At 90d958e, test/shell.d/network-status-probe-test.sh (9/9), test/shell.d/network-test.sh (104/104) and ./test/cli all pass on the VM.

Pushed:

  • 7c19f78: resolvectl dns prints a server as addr[:port][%ifname][#name], and omarchy dns writes its Cloudflare, Google and DNS-over-TLS custom servers as 1.1.1.1#cloudflare-dns.com. is_public_ipv4 rejected every one of those, so configured resolvers were silently swapped for the fixed fallbacks. On a network that also blocks 8.8.8.8 and 9.9.9.9, a user's own public resolver was never tried. Each token is now cut at the first :, % or # before it is classified, and IPv6 is still rejected. Test cases added.
  • 90d958e: 198.18.0.0/15 is now rejected as a probe. Clash and Mihomo in fake-IP TUN mode hand out resolvers in that range and answer for them locally, so probing one would report the internet as up through an outage. Boundary cases added.

Not changed, for the maintainer to weigh:

Second opinion. Codex Medium reviewed the change twice. The first round found the suffix and 198.18.0.0/15 problems fixed above and the three notes listed. It also agreed that the bug is real and that the other pull requests are different bugs; independence isn't guaranteed, since it can read this session's notes. The second round, over the two pushed commits, found nothing.

This now waits on the maintainer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working ready Good to merge verified Omarchy Triage has verified that this issue is ready for final review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants