Critical finding: flipping vllm_service_enabled/state=true/started and restarting llama-swap alongside it broke llama-swap's ability to load ANY of its own generative models -- every /v1/chat/completions request against Qwen3.8-27B-Q4_K_M or Qwen3-8B aux models failed with 'upstream command exited prematurely' (llama-server OOM at spawn, ~1.8GB free on this 24GB card once vLLM's ~22.8GB was claimed). Confirmed by direct A/B: same request 500s with vLLM running, 200s seconds after stopping it. This breaks 21 Hermes agent profiles' aux-model tasks (skills_hub, approval, mcp, title_generation, profile_describer, compression) plus OpenViking's VLM -- a far larger blast radius than Hindsight's single LLM endpoint. Reverted: - vllm_service_enabled/state back to role defaults (false/stopped) -- vLLM stays staged, startable for a brief validated shadow window, NOT safe to leave resident in production. - Hindsight's HINDSIGHT_API_LLM_BASE_URL back to llama-swap (astro-orbiter:8001, Qwen3.8-27B-Q4_K_M) and the API key secret source back to the Nous fallback item (pre-task state) -- the vLLM cutover, while functionally validated in isolation (health, /v1/chat/completions, and a live hindsight_retain+recall round-trip all succeeded), requires continuous vLLM availability which is now known to be unsafe on this card. Comment posted on t_6dff1ecc: teardown remains correctly blocked -- full cutover is not achievable within this card's VRAM budget as currently scoped. Needs a human decision on aux-model migration strategy (see roles/deploy-vllm README's 'Critical architectural finding' section) before any further progress.
189 lines
11 KiB
YAML
189 lines
11 KiB
YAML
---
|
|
# ------------------------------------------------------------------------------
|
|
# FILE: ansible/host_vars/astro_orbiter/vars.yml
|
|
# HOST: astro-orbiter (10.1.71.130)
|
|
# ROLE: llama.cpp LLM inference host — Ryzen 7 5800XT / RTX 3090 (ATX rebuild,
|
|
# 2026-08-04). Superseded the prior AMD RX 5700 / Ollama config below;
|
|
# drive was transplanted into new hardware, not reinstalled.
|
|
# ------------------------------------------------------------------------------
|
|
|
|
ansible_host: 10.1.71.130
|
|
ansible_user: jarvis
|
|
ansible_ssh_private_key_file: ~/.ssh/id_jarvis
|
|
ansible_become: true
|
|
|
|
# LVM root expansion — xlarge template uses sda3 partition, standard VG/LV names
|
|
common_expand_root_lvm: true
|
|
common_root_pv: /dev/sda3
|
|
common_root_vg: ubuntu-vg
|
|
common_root_lv: ubuntu-lv
|
|
|
|
# --- Staged GGUF models for the llama.cpp router (:8002) ---------------------
|
|
# Data-driven list consumed by roles/llm-inference-multimodel tasks/models.yml
|
|
# (loop -> tasks/stage_model.yml). Each entry is idempotently staged into
|
|
# /opt/models: stat + EXACT-size check vs HF manifest; skip (no download, no
|
|
# restart) when present + size matches. Source repos are public bartowski GGUFs
|
|
# on HuggingFace (no auth). A router restart is notified ONLY when a new GGUF
|
|
# is actually downloaded.
|
|
# Added 2026-08-12 (War Machine): codify Phi-3.5-mini-instruct-Q8_0 and
|
|
# Meta-Llama-3.1-8B-Instruct-Q4_K_M as router models alongside the production
|
|
# Qwen3.6-35B-A3B-UD-Q4_K_S. The live files were already present/correct on
|
|
# astro-orbiter; this pass codifies them. Future adds = append to this list.
|
|
# Router --models-max override for astro-orbiter.
|
|
# Default in defaults/main.yml is 1 (conservative). Bumped to 4 on 2026-08-12
|
|
# (t_33acbb2e) so the router can keep more than one GGUF resident on-demand
|
|
# and LRU-evict when needed.
|
|
#
|
|
# VRAM NOTE (t_33acbb2e, updated t_55c164f5, updated t_34b96e83, updated t_f5f7e9ad, updated t_441470b9, updated t_c5cef2b2):
|
|
# With models-max=4 and all 6 GGUFs registered, worst case is all 6 loaded simultaneously:
|
|
# Qwen3.8-27B Q4_K_M: ~20.0GB (weights ~17.1GB + KV ~2.9GB @ 65536 ctx, q4_0) ← CORRECTED (ctx rolled back from 128K to 65536, t_c9fed26c 2026-08-18)
|
|
# Phi-3.5-mini-instruct Q8_0: ~4.3GB (weights ~3.8GB + KV ~0.5GB @ 32K ctx)
|
|
# Meta-Llama-3.1-8B Q4_K_M: ~5.6GB (weights ~4.6GB + KV ~0.2GB @ 8K ctx)
|
|
# Qwen2.5-Coder-14B Q4_K_M: ~9.0GB (weights ~8.4GB + KV ~0.6GB @ 16K ctx)
|
|
# nomic-embed-text-v1.5 Q4_K_M: ~0.09GB (~84MB, embedding only — no KV cache)
|
|
# Qwen3-8B Q4_K_M: ~5.5GB (weights ~4.68GB + KV ~0.5GB @ 32K ctx, q4_0)
|
|
# Total worst-case: ~44.5GB >> 24GB RTX 3090
|
|
#
|
|
# OOM RISK: Full co-residency is impossible on 24GB. LRU eviction prevents this
|
|
# in practice: models-max=4 means the router can REGISTER 6 models but only keeps
|
|
# up to 4 LOADED simultaneously — the router will evict the LRU model when a new
|
|
# one is needed. nomic-embed-text-v1.5 is pinned via sleep-idle-seconds=-1 and
|
|
# load-on-startup=true but it uses only ~84MB, so it never meaningfully changes
|
|
# the budget. In single-user homelab operation, only one generative model is active
|
|
# at a time alongside the always-resident embedding model.
|
|
# Qwen3.8-27B alone uses ~17,804 MiB (weights+KV @ 65536 ctx); co-residency
|
|
# with Coder (~9GB) = ~27GB > 24GB. LRU eviction handles this automatically.
|
|
# Ryan should be aware this means model-switching always incurs a ~30-60s
|
|
# cold-load latency when switching between Qwen3.8-27B and any other model.
|
|
# Proceeding to models-max=4 as instructed; flagged for Ryan's attention.
|
|
# Router --models-max override for astro-orbiter.
|
|
# UPDATED (t_f5f7e9ad, 2026-08-16): Set to 2 because Qwen3.8-27B-Q4_K_M
|
|
# uses 17,804 MiB at 65536 ctx. Only nomic-embed (558MB, pinned) and ONE
|
|
# generative model can be resident simultaneously. Co-residency of Qwen3.8
|
|
# with any auxiliary model (Phi 8.3GB, Llama 5.9GB, Coder 9GB) exceeds 24GB.
|
|
# models-max=2: slot 1 = nomic-embed (pinned, always loaded), slot 2 = LRU
|
|
# generative model (Qwen3.8 primary, cold-loaded on first request ~30-60s;
|
|
# auxiliary models evict it on demand, and vice versa).
|
|
# NOTE: Qwen3.8 does NOT have load-on-startup — it loads on first request.
|
|
# This avoids an LRU eviction race with nomic-embed at startup.
|
|
# UPDATED (t_72646029, 2026-08-17): CPU offload for Coder + Llama changes the
|
|
# constraint. Coder and Llama now use CPU inference (n-gpu-layers=0). GPU-resident
|
|
# VRAM: Qwen3.8 (~17,804 MiB at 65536 ctx) + nomic-embed (558 MiB, pinned) plus
|
|
# the CUDA-context buffers llama.cpp 6ea215d allocates for the CPU models (~1.4-1.7GB
|
|
# each) = ~20,004 MiB steady-state, below the 24,576 MiB physical limit.
|
|
# CORRECTED (t_c5cef2b2, 2026-08-19): ctx-size was rolled back from 131072 to 65536
|
|
# (t_c9fed26c 2026-08-18). Qwen3.8 VRAM at 65536: 17,804 MiB (not 20,302 MiB).
|
|
# models-max raised to 4: nomic (slot 1, pinned) + Qwen3.8 (slot 2, GPU) +
|
|
# Llama (slot 3, CPU) + Coder (slot 4, CPU). Phi (GPU, ~8.3GB) and new
|
|
# Qwen3-8B (GPU, ~5.5GB) can also be requested but evict Qwen3.8 due to VRAM.
|
|
# models-max=4 is required so CPU-offloaded models count as loaded without
|
|
# evicting Qwen3.8.
|
|
llm_router_models_max: 4
|
|
|
|
llm_staged_models:
|
|
- filename: "Phi-3.5-mini-instruct-Q8_0.gguf"
|
|
url: "https://huggingface.co/bartowski/Phi-3.5-mini-instruct-GGUF/resolve/main/Phi-3.5-mini-instruct-Q8_0.gguf"
|
|
size_bytes: 4061222688
|
|
source_repo: "bartowski/Phi-3.5-mini-instruct-GGUF"
|
|
- filename: "Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf"
|
|
url: "https://huggingface.co/bartowski/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf"
|
|
size_bytes: 4920739232
|
|
source_repo: "bartowski/Meta-Llama-3.1-8B-Instruct-GGUF"
|
|
- filename: "Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf"
|
|
url: "https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/resolve/main/Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf"
|
|
size_bytes: 8988111072
|
|
source_repo: "bartowski/Qwen2.5-Coder-14B-Instruct-GGUF"
|
|
- filename: "nomic-embed-text-v1.5-Q4_K_M.gguf"
|
|
url: "https://huggingface.co/nomic-ai/nomic-embed-text-v1.5-GGUF/resolve/main/nomic-embed-text-v1.5.Q4_K_M.gguf"
|
|
size_bytes: 84106624
|
|
source_repo: "nomic-ai/nomic-embed-text-v1.5-GGUF"
|
|
# Added t_c5cef2b2 (2026-08-19, War Machine): Qwen3-8B dense 8B model for
|
|
# aux tasks (routing, rewriting, structured extraction, tool-call construction).
|
|
# Source: bartowski/Qwen_Qwen3-8B-GGUF (public, no auth). HF filename is
|
|
# Qwen_Qwen3-8B-Q4_K_M.gguf; stored locally as Qwen3-8B-Q4_K_M.gguf.
|
|
# Exact size verified from HF manifest (content-length): 5,027,784,224 bytes.
|
|
# VRAM: ~4.68GB weights + ~0.5GB KV @ 32K ctx (q4_0) ≈ 5.2GB total.
|
|
# Thinking mode ON by default; use /no_think for latency-sensitive aux tasks.
|
|
- filename: "Qwen3-8B-Q4_K_M.gguf"
|
|
url: "https://huggingface.co/bartowski/Qwen_Qwen3-8B-GGUF/resolve/main/Qwen_Qwen3-8B-Q4_K_M.gguf"
|
|
size_bytes: 5027784224
|
|
source_repo: "bartowski/Qwen_Qwen3-8B-GGUF"
|
|
|
|
# --- deploy-vllm role: vllm_models override (t_e6facb19, 2026-08-31) --------
|
|
# Ansible's hash_behaviour is "replace" (see ansible.cfg) — a host_vars list
|
|
# variable REPLACES the role default list wholesale, it does not deep-merge.
|
|
# This is therefore a full copy of roles/deploy-vllm/defaults/main.yml's
|
|
# vllm_models with ONE change: nomic-embed-text-v1.5.enabled flipped to true,
|
|
# now that vllm.service.j2 has an embedding-mode branch (--runner pooling
|
|
# --convert embed --trust-remote-code) tested end-to-end in a shadow window.
|
|
# Primary (Qwen2.5-32B-Instruct-AWQ) and aux (Qwen3-8B-AWQ) entries are
|
|
# unchanged from role defaults — reproduced here only because the whole list
|
|
# must be redefined together. Keep this in sync with defaults/main.yml if the
|
|
# role's model roster changes.
|
|
vllm_models:
|
|
- id: "Qwen2.5-32B-Instruct-AWQ"
|
|
hf_repo: "Qwen/Qwen2.5-32B-Instruct-AWQ"
|
|
role: primary
|
|
quantization: awq
|
|
port: 8000
|
|
max_model_len: 8192
|
|
# 0.95 (role default) OOM'd during CUDA graph capture once nomic-embed
|
|
# (role: embedding, ~814MiB actual, not the nominal 300MB) is co-resident
|
|
# on the same 24GB card (t_e6facb19, 2026-08-31): KV cache allocation
|
|
# succeeded (14,720 tokens) but graph capture needed ~20MiB more than the
|
|
# 0.95 budget left after nomic's share. Two independent, permanent
|
|
# co-residents (unlike t_ca1af9fb's shadow-window test, which had the
|
|
# whole 24GB free) need either a lower utilization ceiling or no graph
|
|
# capture. enforce_eager avoids the whole cudagraph capture memory spike
|
|
# entirely — small throughput cost, no OOM risk, safer for a fixed
|
|
# multi-process VRAM budget than tuning utilization percentages by hand.
|
|
# Even WITH enforce_eager, 0.95 left only ~847MiB genuinely free out of
|
|
# 24576MiB total (23,729MiB used) and both services crash-looped 6-7x
|
|
# during warmup/KV-cache sizing before stabilizing — too fragile for a
|
|
# permanent two-process co-residency. Lowered to 0.90 for real headroom
|
|
# (~1.6GiB free), confirmed clean single-attempt start with no retries.
|
|
gpu_memory_utilization: 0.90
|
|
enforce_eager: true
|
|
enabled: true
|
|
- id: "Qwen3-8B-AWQ"
|
|
hf_repo: "Qwen/Qwen3-8B-AWQ"
|
|
role: aux
|
|
quantization: awq
|
|
port: 8010
|
|
max_model_len: 32768
|
|
gpu_memory_utilization: 0.15
|
|
enabled: false
|
|
- id: "nomic-embed-text-v1.5"
|
|
hf_repo: "nomic-ai/nomic-embed-text-v1.5"
|
|
role: embedding
|
|
quantization: none
|
|
port: 8020
|
|
max_model_len: 2048
|
|
gpu_memory_utilization: 0.05
|
|
trust_remote_code: true
|
|
enabled: true
|
|
|
|
# --- deploy-vllm role: boot persistence NOT enabled (t_e6facb19, 2026-08-31) ---
|
|
# ATTEMPTED enabling vllm_service_enabled/state=started here, then reverted
|
|
# after a production-breaking discovery: with vLLM's two processes
|
|
# (Qwen2.5-32B + nomic-embed, ~22.8GB combined) resident, llama-swap could no
|
|
# longer load ANY of its own generative models -- every /v1/chat/completions
|
|
# request against Qwen3.8-27B-Q4_K_M or the Qwen3-8B aux models failed with
|
|
# "upstream command exited prematurely" (llama-server's own OOM at spawn
|
|
# time, silently swallowed by llama-swap's generic error). Confirmed by
|
|
# direct A/B: same request 500s with vLLM running, 200s within seconds of
|
|
# `systemctl stop vllm.service vllm-nomic-embed-text-v1.5.service`. This
|
|
# breaks all 21 Hermes agent profiles' aux-model tasks (skills_hub, approval,
|
|
# mcp, title_generation, profile_describer, compression) plus OpenViking's
|
|
# VLM -- a severe regression, worse than the status quo. Left
|
|
# vllm_service_enabled/state at role defaults (false/stopped) -- vLLM stays
|
|
# staged and manually startable for a brief shadow window (same pattern as
|
|
# t_ca1af9fb's original validation), but is NOT safe to leave resident
|
|
# alongside llama-swap on this 24GB card. See README's "Known Gaps" section:
|
|
# full teardown of llama-swap (t_6dff1ecc) is the ONLY path to giving vLLM
|
|
# permanent residency without starving the other models -- this is not
|
|
# solvable by tuning gpu_memory_utilization further, the two stacks
|
|
# together need more VRAM than this card has once both hold real models
|
|
# resident.
|
|
|