Merge origin/main — integrate monitoring/Phase3 updates with Qwen3-8B no-think deployment
Resolved add/add conflicts in:
- defaults/main.yml: kept our version (5 original models + Qwen3-8B x2 + rows 5-6)
- tasks/swapmode.yml: kept our version (7-model GATE 2 assert)
- templates/llama-server-router-preset.ini.j2: kept our version (+Qwen3-8B sections)
- templates/llama-swap-config.yaml.j2: kept our version (+chat_template_file support)
Remote changes incorporated from origin/main (14 commits):
- Ansible Phase 3 integration (llama-swap.service.j2, tasks/monitoring.yml)
- Prometheus monitoring: PrometheusRule, Grafana dashboard, scrape config
- VRAM exporter script, llama-swap-phase3 cutover results
- Day2 playbooks: nomic_embed, cpu_offload_aux, per_model_ctx, qwen38_ctx128k
- Router: CPU-offload Coder-14B + Llama-3.1-8B
- host_vars/astro-orbiter/vars.yml updates