Each /metrics?model=<id> request on the llama.cpp router wakes the GPU sub-server to P2 (~110W). At 15s with two active jobs (llama3 + phi35), combined scrape frequency (~7-8s effective) keeps the GPU continuously at P2 despite zero real inference requests. At 90s: each scrape wakes GPU for ~5-10s then it drops to P8 (~20W) for ~80s. Verified via iptables block test on 2026-08-13 (t_e7d547ea): Before block: 110-115W P2 continuously After block: 19-21W P8 consistently After unblock: returned to 110W P2 within seconds qwen3 interval also set to 90s (model not loaded so moot, but consistent). Ref: t_e7d547ea
14 KiB
14 KiB