Skip to content

Restart Bluetooth agent when BlueZ is replaced - #11937

Open
n0mahd wants to merge 1 commit into
omacom:quattrofrom
n0mahd:fix/bluetooth-agent-bluez-restart
Open

n0mahd wants to merge 1 commit into
omacom:quattrofrom
n0mahd:fix/bluetooth-agent-bluez-restart

Conversation

@n0mahd

@n0mahd n0mahd commented Sep 15, 2026

Copy link
Copy Markdown

Summary

  • supervise bt-agent against the current org.bluez owner
  • restart the agent whenever bluetoothd is replaced
  • wait for BlueZ at login instead of permanently skipping after a startup race
  • bound service shutdown even though bt-agent ignores SIGTERM

Problem

Agent registrations belong to one bluetoothd process. When BlueZ restarts, the existing bt-agent process remains alive and systemd considers the user service healthy, but the replacement daemon has no registered agent. Pairing then stalls at Connecting while BlueZ logs:

src/device.c:new_auth() No agent available for request type 2
device_confirm_passkey: Operation not permitted

The previous ExecCondition also made a login-time race permanent: if the user service started before the system BlueZ service, the agent was skipped until manually restarted.

Verification

  • ./test/shell.d/bluetooth-agent-test.sh
  • ./test/shell.d/systemd-test.sh
  • ./test/cli
  • bash -n bin/omarchy-bluetooth-agent-supervisor test/shell.d/bluetooth-agent-test.sh
  • git diff --check
  • live Apple MacBook Air M1 / BlueZ 5.87 test: after replacing bluetoothd, the supervisor replaced the agent within one polling interval and both a Logitech K850 and MX Master 3S reconnected within 14 seconds

./test/all reaches the new tests successfully. Its six failing files are pre-existing host/fixture failures; five reproduce exactly from an untouched f2b419d9 worktree, and the sixth (locate-test.sh) passes alone and is test-order contamination.

Fixes #11936

@llstrk

llstrk commented Sep 28, 2026

Copy link
Copy Markdown

Automated AI review

Community review: Independent automated community review, unaffiliated with the Omarchy team, intended to help prepare PRs for their review.

Verified: the supervisor fixes the stale-agent problem from #11936. With the real bluez-tools bt-agent against a fake BlueZ on a private bus, running the old unit's ExecStart command left the agent alive and registered with a daemon that no longer existed. The new daemon had no agent. Under the supervisor, the replacement was detected and the stale agent killed, and the restarted agent registered with the live daemon. The findings below concern cost, robustness and rollout; none stopped the fix from working in the tested scope. The main suggestion is to probe only the org.bluez owner instead of listing the whole bus. That also removes one cause of the second finding below.

old unit     bluetoothd B1 ──replaced──▶ B2
             bt-agent registered with B1, still running ──▶ B2 has 0 agents

this PR      poll sees org.bluez PID change ──▶ supervisor exits 1, kills bt-agent
             (measured ~2 s after replacement at the default 2 s poll)
             restart ──▶ new bt-agent registered with B2
             (~2.9 s after B2 appeared, with RestartSec=1 simulated)

Verified in the same setup:

  • BlueZ stopping with no replacement is also handled. The restarted supervisor waits and registers once BlueZ returns.
  • The startup wait does not D-Bus-activate org.bluez.
  • bt-agent does ignore SIGTERM (bluez-tools src/bt-agent.c wires SIGTERM to the SIGUSR1 handler), and the trap's SIGKILL cleans it up on stop.
  • bluetooth-agent-test.sh passes on the head and on the head merged onto current quattro, fails on the merge base, and caught four behavioural mutants.

Every session lists the entire system bus every 2 seconds

bluez_pid() runs busctl --system list, which lists every unique, well-known and activatable name and queries credentials for each owned name. On a private test bus with 124 names, one call made 188 D-Bus method calls and took about 10 ms. busctl --system call … GetNameOwner s org.bluez made 2 calls and took about 1.5 ms. The loop runs for the whole session in every logged-in user's manager, spawning busctl, awk and sleep each time.

The removed ExecCondition used to skip the unit when bluetooth.service was inactive. Now, on a machine where /sys/class/bluetooth exists but BlueZ is disabled or never starts, the unit stays active (running) and polls indefinitely with no agent.

Impact: a small but permanent background cost on every session where the unit runs, including machines with a Bluetooth adapter that never run BlueZ. The real system-bus size and dbus-broker timings were not measured; the call count scales with the number of names on the bus.

Suggested change: query only the org.bluez owner and compare its unique name (for example :1.2), which is never reused on a bus. On the test bus (reference dbus-daemon), GetNameOwner returned s ":1.2" when BlueZ was present, failed with no such name when it was absent, and did not activate the name in either case. dbus-broker's error text was not checked. An event-driven NameOwnerChanged watch would remove polling entirely (not tested here).

A failed probe is treated as a BlueZ replacement

current_bluez_pid=$(bluez_pid)
if [[ $current_bluez_pid != "$initial_bluez_pid" ]]; then   # "" or "-" also differs
  stop_agent
  exit 1

Any empty or - result is compared as a new owner. That includes a busctl error, or a row whose PID credential lookup failed (busctl prints - in that case). With a busctl wrapper that failed once while the owner never changed, the supervisor exited 1 and killed the working agent. The same - value also keeps the startup loop waiting while org.bluez is actually owned.

Impact: a transient probe failure causes an unnecessary restart with no agent registered for about 1 second. The persistent - case is conditional on BlueZ's PID credential being unavailable and was only reproduced with a stubbed busctl.

Suggested change: exit only when the probe succeeds with a different owner, or definitively reports the name absent. Treat other errors as "unknown, poll again". Comparing unique names as suggested above removes the dependence on the PID column, but a failed call still needs this handling.

Running sessions keep the old agent until the next login

The PR adds no migration. Arch's user daemon-reload hook reloads unit files after the package update but does not restart running units. So an already-running bt-agent.service keeps the old ExecStart=/usr/bin/bt-agent process and stays stale across a later BlueZ restart in that session. Sessions where the old ExecCondition skipped the unit stay inactive until the next login.

Impact: on existing sessions, the fix only takes effect after the next login. Until then a BlueZ restart can still leave pairing without an agent.

Suggested change (optional): add a migration that restarts the unit for existing sessions without enabling units the user disabled. systemctl --user try-restart bt-agent.service covers running units. Sessions the old ExecCondition skipped would also need a restart when the unit is enabled. migrations/1789130779.sh and migrations/1785608166.sh restart user units in a similar way.

Optional: KillMode=mixed sends SIGTERM only to the supervisor. bash runs the TERM trap only after the foreground sleep finishes, so stops took up to about 2 s. SIGTERM to the whole process group, standing in for the default KillMode=control-group (not run under systemd), stopped it in 1 to 3 ms with the agent killed. Either dropping KillMode=mixed or using sleep "$poll_interval" & wait $! avoids the delay. TimeoutStopSec=3 already bounds it.

Optional test improvement:

  • The busctl mock prints only an org.bluez row. A supervisor that drops the $1 == "org.bluez" filter from the awk still passes both cases. A mock that also prints another name before org.bluez, with its own PID, catches that mutant while the real supervisor still passes.
  • The new systemd-test.sh greps supervisor implementation strings (while :, initial_bluez_pid=$(bluez_pid), sleep "$poll_interval", the comparison expression). They add nothing over the behavioural test and can fail on a behaviour-preserving refactor such as the owner probe above.

Related open PRs: these change the same bt-agent.service lines and conflict with this one:

The owner supervision here still applies if bt-agent stays (#13258, #9131, #8929). #12897's agent registers from an org.bluez name watch, so going by its source it already re-registers on its own (not run).


Review information

Test scope: source review of the head 46b62ae against the merge base and current quattro, plus bluez-tools (the commit Arch packages), BlueZ 5.87 and systemd v262 source. The runtime tests used the real bt-agent, a fake BlueZ on a private reference dbus-daemon (not dbus-broker), and a simulated restart delay. There was no real systemd user manager, bluetoothd, adapter or pairing, and no hardware test. In the test sandbox, test/cli fails on current quattro at an unrelated environment-dependent assertion before its metadata check, so that check was replicated separately and passes.

AI process: Opus 5.5 Medium coordination and synthesis, Opus 5.5 Xhigh technical review and final fact check, GPT 6 Sol Xhigh search for related issues, Opus 5.5 Medium editorial check.

Opt out: To stop receiving these reviews, reply to this comment saying so.

@omarchybot omarchybot added the enhancement New feature or request label Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bluetooth agent stays stale after bluetoothd restarts

3 participants