feat(llm-inference-multimodel): bump router --models-max 1->4 on astro-orbiter (t_33acbb2e)
Changes: - host_vars/astro-orbiter/vars.yml: add llm_router_models_max: 4 (overrides conservative default of 1). Detailed VRAM OOM risk note included inline: worst-case 3-model co-residency ~31GB > 24GB RTX 3090. LRU eviction mitigates in single-user operation; flagged for Ryan's review. - playbooks/day2_bump_router_models_max.yml: new targeted playbook; deploys updated router unit, restarts the live service, verifies /health 200 and /v1/models lists all 3 GGUFs post-restart. - group_vars/all/semaphore.yml: add llm_router_update_unit template pointing at the new playbook. - roles/llm-inference-multimodel/defaults/main.yml: update comment to reflect the var is now overridden in host_vars rather than 'hardcoded to 1'. - roles/llm-inference-multimodel/templates/llama-server-router.service.j2: correct stale 'HARDCODED TO 1' comment — value is variable-driven. Constraints honored: - --parallel 1 left untouched (not in scope, not modified anywhere) - No ad-hoc SSH/systemctl/curl state mutation; all execution via Semaphore - No installed/vendored code patched
This commit is contained in:
@@ -28,9 +28,12 @@ ExecStart={{ llm_binary_path }} \
|
||||
# - NO -m/--model flag: this is what enables llama-server router/supervisor mode.
|
||||
# Without -m, llama-server discovers all .gguf files in --models-dir, spawning
|
||||
# each as its own child process on demand (LRU-eviction when over models-max).
|
||||
# - --models-max {{ llm_router_models_max }} is HARDCODED TO 1.
|
||||
# Default cap is 4 simultaneous — OOM on 24GB with a 20GB model.
|
||||
# Do not increase without a VRAM budget review (see defaults/main.yml comment).
|
||||
# - --models-max {{ llm_router_models_max }} is driven by llm_router_models_max
|
||||
# (default 1 in defaults/main.yml; overridden to 4 in host_vars/astro-orbiter
|
||||
# as of t_33acbb2e after VRAM budget review — see host_vars for OOM risk note).
|
||||
# Default llama-server cap is 4 simultaneous — OOM on 24GB if all 3 current
|
||||
# GGUFs load at once. LRU eviction mitigates in practice but review before adding
|
||||
# models. See host_vars/astro-orbiter/vars.yml for full VRAM breakdown.
|
||||
# - --models-dir /opt/models: auto-discovers all .gguf files. Keep that directory
|
||||
# clean (Qwen-only) to avoid spurious extra entries in /v1/models.
|
||||
# - Clients select a model via "model": "<gguf-basename-without-.gguf>" in their
|
||||
|
||||
Reference in New Issue
Block a user