Add Qwen2.5-Coder-14B-Instruct-4bit to astro-orbiter router (t_55c164f5)

- host_vars/astro-orbiter/vars.yml: add Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf
  to llm_staged_models (size_bytes=8988111072, bartowski GGUF public repo).
  Updated VRAM note to reflect 4-model roster and LRU eviction semantics.
- defaults/main.yml: add llm_router_coder_ctx_size=16384 and
  llm_router_coder_flash_attn=true variables for per-model ctx tuning.
- templates/llama-server-router-preset.ini.j2: add [Qwen2.5-Coder-14B-Instruct-Q4_K_M]
  section with alias=Qwen2.5-Coder-14B-Instruct-4bit, ctx-size=16384, flash-attn=true.
- playbooks/day2_add_coder_alias.yml: new playbook that downloads the GGUF (if
  absent/mismatched), deploys updated preset INI and systemd unit, restarts
  llama-server-router, and verifies all 4 models in /v1/models.

VRAM: Coder ~9GB. Full 4-model co-residency impossible on 24GB — LRU eviction
handles this automatically. Qwen3.6-35B <-> Coder switches incur ~30-60s cold load.
This commit is contained in:
Hermes Agent service account
2026-08-13 09:07:02 -05:00
parent 6455d22752
commit 7aea88724f
4 changed files with 329 additions and 15 deletions

View File

@@ -148,4 +148,15 @@ llm_router_vram_max_mib: 23000 # Gate 3: fail if exceeded
#
# Added 2026-08-12 (t_9adf0889) — War Machine.
llm_router_preset_enabled: false # flip true to activate preset mode
# Per-model ctx-size / flash-attn overrides for preset mode (t_ryan_per_model_ctx).
# Defaults mirror the prior uniform 65536/auto behavior; host_vars or the
# deploy playbook override these to the values Ryan requested per workload.
llm_router_llama_ctx_size: "{{ llm_router_ctx_size }}"
llm_router_llama_flash_attn: "{{ llm_router_flash_attn }}"
llm_router_phi_ctx_size: "{{ llm_router_ctx_size }}"
llm_router_phi_flash_attn: "{{ llm_router_flash_attn }}"
# Qwen2.5-Coder-14B: ctx_size=16384, flash_attn=true per task t_55c164f5
llm_router_coder_ctx_size: 16384
llm_router_coder_flash_attn: "true"
llm_router_preset_path: /opt/llama-server-router-preset.ini