Previous ctx-size=8192/parallel=4 gave 2048 tokens/slot, too small for context compression inputs (observed live rejection at 3826 tokens). Measured VRAM on astro-orbiter (RTX 3090 24GB): weights ~17GB resident, ~294KiB/token pool-wide for KV cache+buffers at prior sizing. New: ctx-size=16384, parallel=2 -> 8192 tokens/slot (matches model's native n_ctx_train max). Projected VRAM ~21.8GB, ~2.7GB headroom. Applied directly via ansible-playbook (Semaphore currently broken -- fix tracked separately).
2.5 KiB
2.5 KiB