- host_vars/astro-orbiter/vars.yml: - Stage Qwen3-8B-Q4_K_M GGUF (bartowski/Qwen_Qwen3-8B-GGUF, 5,027,784,224 bytes) - Fix stale VRAM comments: KV @ 128K ctx -> 65536 ctx (rolled back t_c9fed26c 2026-08-18) - Update worst-case VRAM table to include new 6th model - defaults/main.yml: - Add Qwen3-8B-Q4_K_M to llm_swapmode_models (port 8106, GPU, 32K ctx, q4_0 KV) - Add row5 to llm_swapmode_matrix_rows (Qwen3-8B & nomic-embed) - templates/llama-server-router-preset.ini.j2: - Add [Qwen3-8B-Q4_K_M] section (32K ctx, GPU, flash-attn, q4_0 KV, 60s idle evict) - Document thinking-mode handling: enabled by default; /no_think at call time for aux tasks VRAM math: Qwen3-8B ~4.68GB weights + ~0.5GB KV @ 32K = ~5.2GB. Fits with nomic-embed (~84MB) well within 24GB. Cannot co-reside with Qwen3.8-27B (17.8GB); LRU eviction applies. Note: war-machine profile config.yaml also updated (Qwen3.6 stale entry -> Qwen3.8-27B; added Qwen3-8B-Q4_K_M context_length: 32768). config.yaml not tracked in git.
6.9 KiB
6.9 KiB