Dashboard decision (t_5508360a): 'stop and disable llama-swap and start
vLLM and its 3 models' -- explicit approval of a breaking change.
Applied:
- llama-swap stopped + disabled on astro-orbiter (systemd unit removed
from multi-user.target.wants, files left in place -- full teardown is
t_6dff1ecc, separate task)
- vllm_service_enabled/state flipped to true/started -- vLLM is now the
permanent, boot-persistent serving layer (was shadow-only/staged)
- Attempted enabling Qwen3-8B-AWQ (the 3rd model) -- does NOT fit.
Qwen2.5-32B-Instruct-AWQ (~18.6GB) + nomic-embed (~0.8GB) leaves only
~1.25GiB free on the 23.55GiB usable budget, below the 3.53GiB floor
Qwen3-8B-AWQ needs even at gpu_memory_utilization=0.15. Confirmed via
journalctl: identical ValueError on 7/7 consecutive restart attempts,
not a transient crash-loop. Reverted Qwen3-8B-AWQ to enabled: false.
2 of the 3 requested models fit permanently, not 3.
Verified live: /health 200 on both :8000 and :8020, live completion and
live embedding both returned correct real output, NRestarts=0 on both
services after a clean idempotent re-run (changed=0).