Hermes Agent service account
bafd76a0b4
feat(astro-orbiter): add Qwen3-8B-Q4_K_M to inference stack (t_c5cef2b2)
- host_vars/astro-orbiter/vars.yml:
- Stage Qwen3-8B-Q4_K_M GGUF (bartowski/Qwen_Qwen3-8B-GGUF, 5,027,784,224 bytes)
- Fix stale VRAM comments: KV @ 128K ctx -> 65536 ctx (rolled back t_c9fed26c 2026-08-18)
- Update worst-case VRAM table to include new 6th model
- defaults/main.yml:
- Add Qwen3-8B-Q4_K_M to llm_swapmode_models (port 8106, GPU, 32K ctx, q4_0 KV)
- Add row5 to llm_swapmode_matrix_rows (Qwen3-8B & nomic-embed)
- templates/llama-server-router-preset.ini.j2:
- Add [Qwen3-8B-Q4_K_M] section (32K ctx, GPU, flash-attn, q4_0 KV, 60s idle evict)
- Document thinking-mode handling: enabled by default; /no_think at call time for aux tasks
VRAM math: Qwen3-8B ~4.68GB weights + ~0.5GB KV @ 32K = ~5.2GB. Fits with nomic-embed
(~84MB) well within 24GB. Cannot co-reside with Qwen3.8-27B (17.8GB); LRU eviction applies.
Note: war-machine profile config.yaml also updated (Qwen3.6 stale entry -> Qwen3.8-27B;
added Qwen3-8B-Q4_K_M context_length: 32768). config.yaml not tracked in git.
2026-08-19 11:17:07 -05:00
..
2025-10-19 17:02:16 -05:00
2026-08-19 11:17:07 -05:00
2025-07-30 12:41:18 -05:00
2025-11-21 05:48:43 -08:00
2026-05-26 11:48:23 -05:00
2025-08-17 22:48:39 -05:00
2025-07-30 12:41:18 -05:00
2026-03-14 21:08:35 -05:00
2025-11-21 05:48:43 -08:00
2025-07-30 12:41:18 -05:00
2026-03-07 22:47:37 -06:00
2026-08-01 21:06:07 -05:00
2025-07-30 12:41:18 -05:00
2025-07-30 12:41:18 -05:00
2025-11-21 05:48:43 -08:00
2025-07-30 12:41:18 -05:00
2025-07-30 12:41:18 -05:00
2025-07-30 12:41:18 -05:00
2025-11-21 05:48:43 -08:00
2025-11-21 05:48:43 -08:00
2025-11-21 05:48:43 -08:00
2025-07-30 12:41:18 -05:00
2025-09-15 22:27:42 -05:00
2025-09-26 11:55:32 -05:00
2026-03-21 14:00:37 -05:00
2025-11-21 05:48:43 -08:00