Default 64000 exceeded vLLM Qwen2.5-32B-Instruct-AWQ's --max-model-len 8192, causing every retain call to 500 with 'max_tokens=64000 cannot be greater than max_model_len=8192'. llama-swap's Qwen3.8-27B ran at ctx=65536 so this never surfaced before the vLLM cutover. Lowered to 4096.
9.5 KiB
9.5 KiB