/metrics?model=Qwen3.6 forces the router to attempt loading Qwen3.6 every 90s scrape cycle, triggering a CUDA OOM error since VRAM is already consumed by the resident Llama3+Phi3.5 models. This produces real GPU power spikes (~110W load-attempt), not a benign counter read like the Llama3/Phi3.5 jobs. nvidia_gpu_exporter (:9835) already provides GPU power/VRAM/utilization at zero wake cost. No Grafana dashboard panel references Qwen3.6 model-specific llama-server metrics. Ryan approved full removal. Job is commented out (not deleted) for easy revert if Qwen3.6 is ever re-added as a resident model. Llama3/Phi3.5 scrape jobs untouched.
14 KiB
14 KiB