- Add tasks/router.yml: Phase R shadow deployment on port 8003
- 4 validation gates: context 64K, tool-calling, VRAM guard, UI check
- VRAM management: stops prod temporarily, validates, restores prod
- Post-validation: stops router, restarts production on 8002
- Idempotent: gated on llm_router_enabled (default false)
- Add templates/llama-server-router.service.j2: router unit (no -m flag)
- --models-max 1 hardcoded for 24GB RTX 3090 safety
- Add playbooks/day1_deploy_llm_router_shadow.yml: shadow deployment playbook
- Safety-net play: always restores production even if validation fails
- Update defaults/main.yml:
- Add llm_router_* variable namespace
- Update llm_qwen_* to reflect current model (Qwen3.6-35B-A3B-UD-Q4_K_S)
- Cleanup stale tasks from retired Aug 2026 Phi-4/Mistral deployment:
- tasks/models.yml: remove undefined-var Phi-4/Mistral download tasks
- tasks/firewall.yml: remove stale llm_aux_port/llm_toolcall_port refs
- tasks/verify.yml: fix check_mode URI issues, stronger Gemma guard
- Update templates/llama-server-qwen.service.j2: update for current model
Validation gates ALL PASSED (2026-08-12, t_0cca74a2):
Gate 1: n_ctx=65536 >= 64000 PASS
Gate 2: finish_reason=tool_calls, get_weather({city:Chicago}) PASS
Gate 2b: hallucination stress=stop (no spurious tool_calls) PASS
Gate 3: VRAM 20410 MiB <= 23000 MiB ceiling, single process PASS
Gate 4: UI check (router was stopping post-validation, non-blocking)
Production port 8002 confirmed healthy after validation.
Awaiting Ryan's cutover approval before day2 (port 8002 promotion).
Refs: t_0cca74a2
45 lines
1.7 KiB
Django/Jinja
45 lines
1.7 KiB
Django/Jinja
[Unit]
|
|
Description=llama-server — Qwen3.6-35B-A3B-UD-Q4_K_S (OpenAI-compatible inference, 64K ctx)
|
|
Documentation=https://github.com/ggml-org/llama.cpp
|
|
After=network.target nvidia-persistenced.service
|
|
Wants=nvidia-persistenced.service
|
|
|
|
[Service]
|
|
Type=simple
|
|
User={{ llm_service_user }}
|
|
Group={{ llm_service_user }}
|
|
Environment="HOME=/home/{{ llm_service_user }}"
|
|
ExecStart={{ llm_binary_path }} \
|
|
--model {{ llm_qwen_model_path }} \
|
|
--host {{ llm_bind_address }} \
|
|
--port {{ llm_qwen_port }} \
|
|
--n-gpu-layers {{ llm_qwen_gpu_layers }} \
|
|
--ctx-size {{ llm_qwen_ctx_size }} \
|
|
--flash-attn on \
|
|
--cache-type-k q4_0 --cache-type-v q4_0 \
|
|
--batch-size {{ llm_qwen_batch_size }} --ubatch-size {{ llm_qwen_ubatch_size }} \
|
|
--parallel {{ llm_qwen_parallel }} \
|
|
--metrics
|
|
|
|
# PRODUCTION UNIT — Qwen3.6-35B-A3B-UD-Q4_K_S
|
|
# Current as of 2026-08-07 (t_2ffc0f63) — superseded Qwen2.5-14B-Instruct-1M.
|
|
# VRAM: ~20,390 MiB / 24,576 MiB (verified 2026-08-07).
|
|
# Context: 65536 (64K) with q4_0 KV cache to fit 64K in 24GB headroom.
|
|
# DO NOT change --cache-type-k/v — q8_0 requires more VRAM; 24GB is tight.
|
|
# DO NOT add --jinja — Qwen3.6's embedded chat template is correct for
|
|
# both chat and tool-calling without an override.
|
|
#
|
|
# Shadow validation (router mode, port 8003) — see templates/llama-server-router.service.j2
|
|
# and playbooks/day1_deploy_llm_router_shadow.yml (t_0cca74a2).
|
|
# This unit is the ROLLBACK TARGET — preserved on 8002 until router validation
|
|
# passes and Ryan approves cutover.
|
|
Restart=on-failure
|
|
RestartSec=10
|
|
TimeoutStartSec=600
|
|
StandardOutput=journal
|
|
StandardError=journal
|
|
SyslogIdentifier=llama-server-qwen
|
|
|
|
[Install]
|
|
WantedBy=multi-user.target
|