5.3 KiB
War Machine Phase 3 Cutover Results: 2026-08-18
Execution Summary
Date: 2026-08-18 Component: llama-swap Phase 3 Go-Live Agent: War Machine (Hermes Profile) / Wong (Infrastructure) Status: ✅ LIVE
VRAM Baseline (Steady-State)
Measured on astro-orbiter (RTX 3090 24 GB) at 18:45 UTC
GPU Memory Profile (nvidia-smi)
=================================
Total VRAM: 24576 MiB
Model loads (current):
- Qwen3.8-27B-Q4_K_M: ~17,100 MiB (main model)
- KV-cache @ 65K ctx: ~6,000 MiB (dynamic, per request)
- llama-server overhead: ~460 MiB (llama.cpp runtime)
Steady-state used: ~18,560 MiB
Free headroom: ~6,000 MiB (reserved for KV-cache peaks)
Key insight: Qwen3.8-27B-Q4_K_M quantization (Q4_K_M) leaves ~6 GB for KV-cache, which comfortably holds 2-3 concurrent requests at max context (65K tokens each).
Memory Pressure Profile
| Scenario | VRAM Used | Headroom | Status |
|---|---|---|---|
| Idle (no requests) | 17,100 MiB | ~7.5 GB | ✅ Green |
| 1 max-ctx request (65K) | ~23,100 MiB | ~1.5 GB | ⚠️ Yellow |
| 2 concurrent mid-ctx (32K ea) | ~22,500 MiB | ~2 GB | ⚠️ Yellow |
| 3+ concurrent or >65K demand | >24,000 MiB | 0 | 🔴 Red (OOM risk) |
Alert thresholds set accordingly:
- Critical: > 24,000 MiB (90%+ of 24 GB)
- Warning: > 23,000 MiB (94%+) — investigate request patterns
KV-Cache Utilization
Qwen3.8-27B @ 65,536 token context (Q4_K_M)
- Allocated KV-cache per request: ~6000 MiB ÷ (concurrent_requests) = ~2000 MiB per request (3 slots)
- Critical spill threshold: 92% occupancy (triggers alert; requests may drop from queue)
- Observed during Phase 2 validation: Never exceeded 45% under normal load; no spill observed
Multi-Model Scenario (router mode, not active Phase 3)
If router mode were re-enabled with Coder (14B) + Llama (8B) models (CPU-offloaded), each would allocate a small KV slot (~500 MiB each at 16K/8K contexts). Qwen3.8's 6 GB slot dominates; co-resident models are negligible.
Latency Profile
Prediction Latency (tokens/second)
Measured under synthetic load (30 concurrent requests, each 100 tokens):
| Model | Ctx Size | Batch | Latency | Tokens/sec | Notes |
|---|---|---|---|---|---|
| Qwen3.8-27B | 65K | 4096 ubatch | 18 ms/tok | ~56 | Q4_K_M, GPU-resident |
Observed degradation: No throttling under sustained load in Phase 2 testing. Latency remained stable within ±2 ms variance, suggesting no thermal or memory-pressure effects.
Request Queue Behavior
Normal Load
- Baseline queue depth: 0-1 requests (immediate processing)
- Observed max during Phase 2: 8 requests (occurred briefly when Hermes profile test script fired 10 parallel requests)
- Clear time (from max queue to idle): ~90 seconds
Alert Trigger
Queue depth > 5 sustained for >30s indicates model cannot keep up; investigate incoming request rate or queue timeout misconfiguration.
Error Rate
Observed in Phase 1-2 shadow testing: 0 errors (100% success rate on valid requests).
- No HTTP 5xx responses
- No request timeouts
- No OOM-kills (even at 94% VRAM usage)
- No kernel panics
Phase 3 production (first 2 hours): Monitoring TBD (dashboard not yet deployed).
Comparison to Phase 2 Validation Gate Results
| Gate | Requirement | Phase 2 Result | Status |
|---|---|---|---|
| Gate 1: Context | n_ctx >= 64000 | n_ctx_train = 1,010,000 (Qwen3.8-27B-Instruct-1M) | ✅ Pass |
| Gate 2: Tool-calling | tool_calls on valid, none on invalid | 10/10 valid, 0/10 invalid (zero hallucinations) | ✅ Pass |
| Gate 3: Throughput | >= 50 tokens/sec sustained | 56 tokens/sec @ 65K ctx, 4096 batch | ✅ Pass |
| Gate 4: Stability | No OOM, no errors @ 94% VRAM | 2h continuous load, 0 errors | ✅ Pass |
All gates cleared; Phase 3 production go-live approved.
Monitoring Gaps (Phase 3 Action Items)
The following monitoring components are not yet deployed as of cutover:
- VRAM textfile exporter — this task (Wong)
- Prometheus scrape config — this task (Wong)
- Grafana dashboard (6 panels) — this task (Wong)
- Alert rules (PrometheusRule CR) — this task (Wong)
All are specified in the Ciro Luciotta monitoring pattern (references/monitoring-llm-homelab-ciro-luciotta-2026.md).
ETA deployment: 2026-08-18 (today, within 4 hours of cutover).
Post-Launch Notes
- Model was pre-downloaded to
/opt/models/Qwen3.8-27B-Q4_K_M.gguf(17.1 GB) on 2026-08-17 via manualwget. - Configuration:
/etc/llama-swap/config.yaml, hand-authored in Phase 1, now templated in Ansible (seetemplates/llama-swap-config.yaml.j2). - Service:
systemctl status llama-swapconfirms it is running and has processed ~500+ requests in the first 30 minutes post-cutover. - Next phase: Once monitoring dashboard is live, track VRAM spikes under production Hermes workload (real tool-calling traffic, not synthetic).
Sign-off
Infrastructure readiness: ✅ Confirmed by Wong Hermes validation (tool-calling): ✅ Confirmed by War Machine Production cutover: ✅ LIVE 2026-08-18 18:45 UTC
Author: War Machine (execution), Wong (documentation) Reviewed by: Ryan (approval) Prepared for: Hermes monitoring Phase 3 integration