astro-orbiter's vLLM primary model changed 2026-09-01 from DeepSeek-R1-Distill-Qwen-32B-AWQ to Gemma-4-26B-A4B-it-AWQ (Google, Apache 2.0, US-origin). Same endpoint/API key — only the served model name changed. max_model_len also bumped to 65536 (was 32768). Gemma 4 does not emit an always-on <think> reasoning trace like DeepSeek-R1 did, so HINDSIGHT_API_RETAIN_MAX_COMPLETION_TOKENS=4096 should have more effective headroom for real extraction output than before, not less.
11 KiB
11 KiB