<!-- llms-explorer concept facts · https://llms-explorer.com/tree/omlx-health-liveness-versus-scheduler-wedge/ · pack 2026-10-05 · ~574 tokens -->

# oMLX /health liveness versus scheduler wedge

> `/health` is unauthenticated in its signature (no `verify_api_key` dependency), while `/api/status` depends on `verify_api_key`.

Parent: [Mac local LLMs: oMLX, Rapid-MLX and related internals](https://llms-explorer.com/tree/mac-local-llms-omlx-and-rapid-mlx-internals/) · 1 facets · 8 facts · page: https://llms-explorer.com/tree/omlx-health-liveness-versus-scheduler-wedge/

## Facts

- `/health` is unauthenticated in its signature (no `verify_api_key` dependency), while `/api/status` depends on `verify_api_key`. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/server.py)
- `/health` returns status code 503 and `"status": "loading"` only when `pinned_preload_complete` is false; otherwise 200 and `"status": "healthy"`. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/server.py)
- `/health` body fields are `status`, `default_model`, `engine_pool` (`model_count`, `loaded_count`, `final_ceiling`, `current_model_memory`) and `mcp`; no field reads scheduler, step or request state. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/server.py)
- If `enforcer.get_final_ceiling()` raises, `/health` logs a warning and reports ceiling 0 instead of failing. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/server.py)
- `/api/status` computes `active_requests` as the sum of `len(core._output_collectors)` over loaded engines and `waiting_requests` as the sum of `len(scheduler.waiting)`; it always returns `"status": "ok"`. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/server.py)
- A probe that detects a wedge must therefore either watch `waiting_requests > 0` with `active_requests == 0` over time, or send a real tiny completion; `/health` cannot detect it on this code. — source: `asserted`
- The server comment says binding the port before the pinned preload lets "port watchdogs see liveness instead of a closed port" and cites issue 2184. — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/server.py)
- A pinned preload that outlasts a port watchdog's timeout would otherwise get the process hard-killed mid-load (server.py comment near line 616). — [source](https://raw.githubusercontent.com/jundot/omlx/main/omlx/server.py)
