Add Qwen2.5-Coder-14B-Instruct-4bit to astro-orbiter router (t_55c164f5)
- host_vars/astro-orbiter/vars.yml: add Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf to llm_staged_models (size_bytes=8988111072, bartowski GGUF public repo). Updated VRAM note to reflect 4-model roster and LRU eviction semantics. - defaults/main.yml: add llm_router_coder_ctx_size=16384 and llm_router_coder_flash_attn=true variables for per-model ctx tuning. - templates/llama-server-router-preset.ini.j2: add [Qwen2.5-Coder-14B-Instruct-Q4_K_M] section with alias=Qwen2.5-Coder-14B-Instruct-4bit, ctx-size=16384, flash-attn=true. - playbooks/day2_add_coder_alias.yml: new playbook that downloads the GGUF (if absent/mismatched), deploys updated preset INI and systemd unit, restarts llama-server-router, and verifies all 4 models in /v1/models. VRAM: Coder ~9GB. Full 4-model co-residency impossible on 24GB — LRU eviction handles this automatically. Qwen3.6-35B <-> Coder switches incur ~30-60s cold load.
This commit is contained in:
@@ -148,4 +148,15 @@ llm_router_vram_max_mib: 23000 # Gate 3: fail if exceeded
|
||||
#
|
||||
# Added 2026-08-12 (t_9adf0889) — War Machine.
|
||||
llm_router_preset_enabled: false # flip true to activate preset mode
|
||||
|
||||
# Per-model ctx-size / flash-attn overrides for preset mode (t_ryan_per_model_ctx).
|
||||
# Defaults mirror the prior uniform 65536/auto behavior; host_vars or the
|
||||
# deploy playbook override these to the values Ryan requested per workload.
|
||||
llm_router_llama_ctx_size: "{{ llm_router_ctx_size }}"
|
||||
llm_router_llama_flash_attn: "{{ llm_router_flash_attn }}"
|
||||
llm_router_phi_ctx_size: "{{ llm_router_ctx_size }}"
|
||||
llm_router_phi_flash_attn: "{{ llm_router_flash_attn }}"
|
||||
# Qwen2.5-Coder-14B: ctx_size=16384, flash_attn=true per task t_55c164f5
|
||||
llm_router_coder_ctx_size: 16384
|
||||
llm_router_coder_flash_attn: "true"
|
||||
llm_router_preset_path: /opt/llama-server-router-preset.ini
|
||||
|
||||
Reference in New Issue
Block a user