Dashboard decision (t_5508360a): 'stop and disable llama-swap and start vLLM and its 3 models' -- explicit approval of a breaking change. Applied: - llama-swap stopped + disabled on astro-orbiter (systemd unit removed from multi-user.target.wants, files left in place -- full teardown is t_6dff1ecc, separate task) - vllm_service_enabled/state flipped to true/started -- vLLM is now the permanent, boot-persistent serving layer (was shadow-only/staged) - Attempted enabling Qwen3-8B-AWQ (the 3rd model) -- does NOT fit. Qwen2.5-32B-Instruct-AWQ (~18.6GB) + nomic-embed (~0.8GB) leaves only ~1.25GiB free on the 23.55GiB usable budget, below the 3.53GiB floor Qwen3-8B-AWQ needs even at gpu_memory_utilization=0.15. Confirmed via journalctl: identical ValueError on 7/7 consecutive restart attempts, not a transient crash-loop. Reverted Qwen3-8B-AWQ to enabled: false. 2 of the 3 requested models fit permanently, not 3. Verified live: /health 200 on both :8000 and :8020, live completion and live embedding both returned correct real output, NRestarts=0 on both services after a clean idempotent re-run (changed=0).
12 KiB
12 KiB