Daemon Readiness Is Not Process Liveness

Daemon readiness is not process liveness

A daemon process can be alive and still not be ready for clients. That sounds
obvious until a CLI launcher reports "daemon failed to start" while the child is
quietly doing slow pre-socket work.

The readiness contract is what the client can actually use. For a local daemon,
that usually means "the IPC endpoint accepts a status request", not "the child
PID exists".

The rule

If startup work happens before the socket or named pipe is bound, the launcher
wait budget must cover the slowest legitimate pre-bind operation. If that budget
feels ugly, the architecture is telling you to bind a small status surface
earlier and report degraded subsystems after that.

Both approaches can be honest:

What does not work is a short launcher timeout in front of long pre-socket
initialization. That creates false failures: the daemon is not dead, but the
client has already told the user it is.

What spotuify taught

In Spotuify, the failing symptom was:

spotuify daemon restart
# error: spotuify daemon did not become stable within 5s

The daemon had not crashed. Current code still shows the shape:
run_daemon() constructs DaemonState, calls ensure_player_ready(), and only
then binds the IPC listener. On macOS, the packaged embedded player registration
has its own 30s timeout. The launcher timeout therefore moved to 60s, with a
test asserting it covers packaged player registration.

That fix is valid, but it also names the deeper follow-up: a music daemon should
probably bind a minimal status socket before the player finishes registering.
Then doctor, the macOS app, and one-shot CLI commands can say "daemon running,
player degraded" instead of forcing startup to look binary.

Why this generalizes

The same issue appears in language servers, local AI runtimes, sync daemons, and
desktop helper services. A child PID tells you only that a process exists. It
does not prove the protocol is accepting requests, the schema version is right,
or the subsystem the user wants is ready.

Startup docs and tests should therefore name the exact readiness signal:

If those are different things, they need different states.

See also