Hermes Agent service account
782cbe33d1
llm-inference: size ctx-size/parallel for aux task offload
Previous ctx-size=8192/parallel=4 gave 2048 tokens/slot, too small for
context compression inputs (observed live rejection at 3826 tokens).
Measured VRAM on astro-orbiter (RTX 3090 24GB): weights ~17GB resident,
~294KiB/token pool-wide for KV cache+buffers at prior sizing.
New: ctx-size=16384, parallel=2 -> 8192 tokens/slot (matches model's
native n_ctx_train max). Projected VRAM ~21.8GB, ~2.7GB headroom.
Applied directly via ansible-playbook (Semaphore currently broken --
fix tracked separately).
2026-08-05 12:14:11 -05:00
..
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-28 21:54:34 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-30 23:08:35 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-06-02 08:43:49 -05:00
2026-06-02 10:28:11 -05:00
2026-05-26 22:22:08 -05:00
2026-08-01 22:15:11 -05:00
2026-05-29 20:56:38 -05:00
2026-08-05 12:14:11 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-06-07 16:48:45 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00
2026-05-26 22:22:08 -05:00