Resolved conflicts between localbafd76a(t_c5cef2b2, Qwen3-8B single-variant) and origin/main5cf4468(t_664289a0, Qwen3-8B dual-thinking deployment). Conflict resolution strategy: took origin/main version throughout — it is the authoritative result from t_664289a0 which deployed the live no_think config to astro-orbiter and already reflects the correct production state. Changes incorporated from origin/main: - defaults/main.yml: Qwen3-8B-Q4_K_M-no_think model entry (port 8107, n_gpu_layers=99, chat_template_file), row6 matrix entry, dual-thinking comment block. - templates/llama-server-router-preset.ini.j2: [Qwen3-8B-Q4_K_M] with sleep-idle-seconds=60 plus new [Qwen3-8B-Q4_K_M-no_think] section with chat-template-file directive. - templates/qwen3-no-think.jinja.j2: new file — Qwen3 template with enable_thinking=false. - tasks/models.yml: template deploy task for qwen3-no-think.jinja (chat_templates tag). - tasks/swapmode.yml: GATE 2 assert updated to 7 models. - templates/llama-swap-config.yaml.j2: chat_template_file flag support. Also pulled in monitoring defaults (llm_monitoring_enabled, VRAM exporter settings, Grafana dashboard vars, Prometheus scrape config) from origin/main monitoring branch.
15 KiB
15 KiB