fix(providers): admission-only block reasons no longer evict a ready provider master
The 60s provider truth observation re-ran while the local model server was running and compared available system RAM against the 13.00 GiB MLX admission floor, after the provider's own approximately 9.74 GiB resident footprint had already been subtracted from that available-memory reading. mlx_install.py returned host-ineligible / ram_insufficient, _readiness_block_observation mapped it to a host-blocked observation, and _handle_provider_truth_result deferred an admission-exclusive stop, SIGTERMing a healthy, ready, health-probed server about 60s after it started. The model's success at loading is what evicted it. ADMISSION_ONLY_REASON_CODES now lives beside REASON_CODES in runtime_health.py and names the reasons that gate admission rather than liveness, currently just ram-insufficient. A new branch in _handle_provider_truth_result declines to request a stop when an unchanged-fingerprint host-blocked observation carries an admission-only reason and the provider is already ready or ready-proof-unavailable, publishing a stale-result-ignored record that preserves the process record. Every other block reason keeps evicting exactly as before. The floor value is not wrong; the place it was consulted was. The declining branch returns True deliberately. Merely skipping the stop would reach the handler's generic tail, which nulls latest_plan, latches latest_phase = host-blocked, and writes a record with no process, a record that lies about a running provider and one _submit_provider_start_if_needed never recovers from because host-blocked is not restartable. Three discriminating test legs pin this. Parakeet's host-admission-blocked stays in the liveness bucket and keeps evicting, pinned by a new regression test: it is overloaded as both a RAM-derived admission verdict and the unrecognised-readiness-code fallback, and parakeet does not flap on the 60s cadence because its latched fingerprint excludes the live RAM reading. It can still evict once on a config-change recompute while RAM is transiently low; that is stated, not fixed. Separately, nothing replaces the load-shedding this removes; the eviction was accidental and fired at 13 GiB available, nowhere near OOM, so the OS now handles memory pressure. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>