Single-model deployment per Ryan's direction: - Primary model: casperhansen/deepseek-r1-distill-qwen-32b-awq, max_model_len=32768 - nomic-embed-text-v1.5 and Qwen3-8B-AWQ both disabled (single-model requirement) - kv_cache_dtype: int4_per_token_head required to fit 32768 ctx on 24GB RTX 3090 (fp16 KV: 26GB needed, doesn't fit at any utilization; fp8 KV: OOM'd during FlashInfer warmup with ~50-150MB margin; int4 KV: clean single-attempt start) - Added kv_cache_dtype / kv_cache_memory_bytes as new optional per-model template fields in vllm.service.j2 (guarded, no effect on other models) Verified live: /health 200, /v1/models confirms max_model_len=32768, live /v1/completions smoke test + manual chat completion both passed (genuine <think> reasoning trace, correct arithmetic). NRestarts=0, steady-state VRAM 23.2GB/24.576GB. Ansible idempotent re-run confirmed changed=0. Known follow-up (not done here): Hindsight's HINDSIGHT_API_LLM_MODEL cluster config still references the retired Qwen2.5-32B-Instruct-AWQ — needs separate GitOps update to point at the new model.
30 KiB
deploy-vllm
Idempotent Ansible role that deploys a vLLM OpenAI-compatible inference server. Written for astro-orbiter (RTX 3090, 24GB VRAM, 64GB RAM, Ubuntu 24.04) and designed for reuse on the planned Mac Mini M4 host later this week (see "Portability" below).
Supersedes the manual, pre-role state left behind by earlier vLLM
experiments (/home/jarvis/vllm-env, bitsandbytes, gemma-2-27b — see
homelab-llm-inference/homelab-llm-serving skills for that history). This
role uses a fresh venv (vllm_venv_path, default ~/vllm-serve-env) and
AWQ pre-quantized models — no bitsandbytes, no on-the-fly quantization,
no repeat of the OOM incident from the earlier Gemma-2-27B attempt.
Phases
| Phase | File | What it does |
|---|---|---|
| 1 | tasks/dependencies.yml |
System Python 3.10+, dedicated venv, pip install vllm>=0.5.0, verifies nvidia-smi and torch.cuda.is_available() |
| 2 | tasks/models.yml |
Downloads each enabled: true model in vllm_models via hf download (huggingface_hub CLI) into ~/.vllm-cache, verifies the snapshot landed and reports on-disk size |
| 3 | tasks/api-key.yml |
Reads the API key from 1Password (op://mk-labs/vllm/api-key) on the controller, writes it to /etc/vllm/api-key.env (root:root, 0600) on the target |
| 4 | tasks/systemd.yml |
Renders and installs one systemd unit per enabled model (vllm.service for the role: primary model, vllm-<id>.service for others) |
| 5 | tasks/verify.yml |
Only runs when vllm_service_state=started. Waits for /health (up to 5 min — torch.compile warmup), checks /v1/models, runs a live completion, scans journalctl for errors |
Run all phases: ansible-playbook -i inventory.yml playbooks/day1_deploy_vllm.yml --limit astro-orbiter
Run one phase: --tags vllm-dependencies / vllm-models / vllm-api-key / vllm-systemd / vllm-verify
Deliberate staging-first default
vllm_service_state defaults to stopped. A default run stages
everything (venv, model weights, API key file, systemd unit) but does
not start the service or touch production traffic. This matches the
astro-orbiter cutover plan: llama-swap is live production serving (Qwen3.8-27B
- nomic-embed for Hindsight) — vLLM must be deployed and validated on a side port/inactive unit before anything is cut over.
To start and validate:
ansible-playbook -i inventory.yml playbooks/day1_deploy_vllm.yml \
--limit astro-orbiter --extra-vars "vllm_service_state=started"
This starts the systemd unit(s), enables them, and runs Phase 5 verification
(health, /v1/models, live completion, clean journalctl).
Cutover of consumers (Hermes profiles, Hindsight embedding config, any
hardcoded :8001/:5805 references) to the new :8000 vLLM endpoint is a
separate, explicit step outside this role — do this only after Phase 5
passes cleanly. Do not tear down llama-swap until consumers are confirmed
working end-to-end against vLLM.
Model roster (vllm_models in defaults/main.yml)
vLLM 0.5.x-0.28.x serves one model per process — multi-model = multiple
systemd units on distinct ports, not a single multiplexed server (unlike
llama-swap's matrix DSL). Today's phase enables only the primary model;
flip enabled: true on the others as VRAM allows (see "Phased Strategy"):
⚠️ Table below reflects the ORIGINAL Qwen2.5-32B deployment. As of
2026-09-01 (t_r1d32b_swap) the primary model is
DeepSeek-R1-Distill-Qwen-32B-AWQ, single-model only (nomic-embed also
disabled) — see the "SUPERSEDED" section further down for current state.
| id | hf_repo | role | port | quant | enabled |
|---|---|---|---|---|---|
| Qwen2.5-32B-Instruct-AWQ | Qwen/Qwen2.5-32B-Instruct-AWQ | primary | 8000 | awq | true |
| Qwen3-8B-AWQ | Qwen/Qwen3-8B-AWQ | aux | 8010 | awq | false |
| nomic-embed-text-v1.5 | nomic-ai/nomic-embed-text-v1.5 | embedding | 8020 | none | false |
Note on the original spec's model choices: the task body named
Qwen/Qwen2.5-32B-Instruct and Qwen/Qwen3-8B-Instruct (bf16, unquantized).
vLLM does not do on-the-fly quantization safely on this host (bitsandbytes
OOM history — see homelab-llm-inference skill Pitfalls) and unquantized
bf16 32B does not fit a 24GB card at all (~65GB). This role instead deploys
the official Qwen AWQ pre-quantized variants
(Qwen/Qwen2.5-32B-Instruct-AWQ, Qwen/Qwen3-8B-AWQ), which vLLM natively
supports (--quantization awq) and which fit the VRAM budget:
- Qwen2.5-32B-Instruct-AWQ: ~19.3GB on disk, fits with ~5GB headroom at 24GB
- Qwen3-8B-AWQ: ~6GB VRAM per llm-explorer
- nomic-embed-text-v1.5: ~300MB, vLLM serves it via
--convert embedpooling (see vLLM embedding docs) — not yet wired into this role's systemd template; the embedding model needs--task embed/--convert embedflags that differ from the completion-serving template. Flagged as a follow-up beforeenabled: trueis flipped on it (see Known Gaps below).
Known Gaps / Follow-ups
- Quarterly API key rotation is documented (
/etc/vllm/API_KEY_ROTATION.mdon the target, rendered bytasks/api-key.yml) but not automated — no cron job exists to force rotation on a schedule. Consider a follow-up cron task if Nick Fury wants this enforced rather than just documented. vllm_service_enableddefaults tofalsedeliberately — see "Deliberate staging-first default" above. Flip together with the cutover step, not before.- vLLM cannot replace llama-swap's full model roster on this card — see "Critical architectural finding" section below for the full incident. Short version: vLLM's one-model-per-process design plus llama-swap's own VRAM needs exceed this 24GB card's capacity when both must serve real models simultaneously. Full llama-swap teardown (t_6dff1ecc) cannot proceed until a human decides the aux-model + VRAM strategy.
Embedding-mode support (t_e6facb19, 2026-08-31)
vllm.service.j2 now branches on role: embedding entries in vllm_models:
adds --runner pooling --convert embed (vLLM's embedding-serving flags —
see https://docs.vllm.ai/en/latest/models/pooling_models/embed/) and
--no-enable-prefix-caching (prefix caching is a completions-only
optimization; irrelevant and safely disabled for pooling). An additional
per-model trust_remote_code: true toggle renders --trust-remote-code
when set — required for nomic-ai/nomic-embed-text-v1.5, which ships
custom NomicBertModel modeling code on its HF repo.
Verification does NOT run /v1/completions against embedding-mode
instances (they don't serve that endpoint — a completions request 400s
immediately). tasks/verify.yml splits vllm_enabled_models by role and
runs the appropriate smoke test per group: completions models get the
/v1/completions "capital of France" test; embedding models get a real
/v1/embeddings POST with an ansible.builtin.assert on a non-empty
data[0].embedding array (not just HTTP 200 — an empty/malformed vector
would still 200).
Critical VRAM finding: co-resident completions + embedding vLLM processes
need MORE headroom than either alone, and CUDA graph capture is the failure
mode, not KV cache sizing. Enabling nomic-embed-text-v1.5 alongside the
primary Qwen2.5-32B model at the role-default gpu_memory_utilization: 0.95
crash-looped repeatedly:
- First failure:
torch.OutOfMemoryErrorduringcapture_model()(CUDA graph capture) — KV cache sizing itself succeeded (14,720 tokens allocated), but graph capture needed ~20MiB more than the 0.95 budget had left once nomic's embedding process (814MiB actual, not the nominal ~300MB estimate in the model roster table) claimed its share. - Fix attempt 1: added a per-model
enforce_eager: truetemplate branch (--enforce-eagerskips CUDA graph capture entirely) — this stopped the graph-capture OOM but the combined processes still landed at only ~847MiB genuinely free out of 24,576MiB, and both services crash-looped 6-7 times during warmup before finally stabilizing (each attempt leaves transient VRAM that the next attempt fights over, extending time-to-stable well past a single health-check retry window). - Fix attempt 2 (final, verified stable): lowered the primary model's
gpu_memory_utilizationfrom 0.95 to 0.90 (host_vars override) in addition toenforce_eager: true. Result: clean single-attempt start for both services,NRestarts=0, ~2GB genuinely free (22,577MiB used / 24,576MiB total). Confirmed viasystemctl show <unit> -p NRestartsafter a full stop/start cycle — 0.95 was NOT a fluke of Restart=always masking the underlying fragility; 0.90 is a real, reproducible fix. - Takeaway for future multi-process vLLM VRAM budgeting on this host:
do not just check "does it eventually come up" — check
NRestartsand free VRAM headroom after a clean stop/start. A model that "works" after 6 crash-loop retries is not production-stable; the retries themselves are evidence the utilization ceiling is too tight for the actual (not nominal) footprint of co-resident processes.
Consumer cutover status (t_e6facb19, 2026-08-31)
Attempted, then REVERTED — Hindsight LLM cutover. Hindsight's
HINDSIGHT_API_LLM_BASE_URL was pointed at vLLM :8000
(Qwen2.5-32B-Instruct-AWQ) and validated working in isolation: health,
/v1/chat/completions, and a live hindsight_retain + recall round-trip
all succeeded (after also fixing HINDSIGHT_API_RETAIN_MAX_COMPLETION_TOKENS,
which defaulted to 64000 — exceeding vLLM's max_model_len=8192 — down to
4096). Reverted anyway, because of a severe discovery documented in the
next section: vLLM cannot stay resident on this card without starving
llama-swap, and Hindsight's LLM endpoint needs continuous availability, not
just a validation window. Restored to http://astro-orbiter:8001/v1
(llama-swap, Qwen3.8-27B-Q4_K_M) — the pre-task working state.
NOT cut over — embeddings. Hindsight was discovered to have NEVER used
astro-orbiter for embeddings — it defaults to a bundled local
BAAI/bge-small-en-v1.5 (384-dim) embedder whenever
HINDSIGHT_API_EMBEDDINGS_PROVIDER is unset, which was always the case here.
Pointing it at vLLM's nomic-embed-text-v1.5 (768-dim) crash-looped the pod:
RuntimeError: Cannot change embedding dimension from 384 to 768: memory_units table contains 1289 rows with embeddings. Re-embedding all
existing memory data across ~20 agent banks is destructive and irreversible
— reverted immediately, left as a separate, explicitly-approved future task.
NOT cut over — 21 Hermes agent profiles' aux models + OpenViking VLM. See "Critical architectural finding" below — this was never attempted once the VRAM collision was discovered, would have made things categorically worse.
Critical architectural finding: vLLM CANNOT be continuously resident alongside llama-swap on this 24GB card (t_e6facb19, 2026-08-31)
After validating vLLM's two processes (Qwen2.5-32B-Instruct-AWQ + nomic-embed-
text-v1.5, ~22.8GB combined) work correctly in isolation, this role's
vllm_service_enabled/vllm_service_state were flipped to true/started
as host_vars overrides to make the deployment permanent (per the task's
"enable for boot" requirement) — llama-swap was then restarted alongside
vLLM to preserve its own consumers. Result: llama-swap could no longer
load ANY of its own generative models. Every /v1/chat/completions
request against Qwen3.8-27B-Q4_K_M or the Qwen3-8B aux models failed
with {"error":"unspecific error: upstream command exited prematurely", "src":"llama-swap"} — llama-server's own OOM at spawn time, only ~1.8GB
free on a 24GB card once vLLM's ~22.8GB was already claimed.
Confirmed by direct A/B test, not inference: identical
Qwen3.8-27B-Q4_K_M chat completion request returned HTTP 500 with vLLM's
two processes running, then HTTP 200 with a real completion within seconds
of systemctl stop vllm.service vllm-nomic-embed-text-v1.5.service — same
llama-swap process, same request, only the GPU memory pressure changed.
This is a hard architectural collision, not a tunable-parameter problem.
llama-swap needs ~18-20GB for its own primary model (Qwen3.8-27B-Q4_K_M);
vLLM's two processes need ~22.8GB even with enforce_eager and a lowered
gpu_memory_utilization. The two together need more VRAM than a 24GB card
has once both hold real models resident — there is no gpu_memory_utilization
value that resolves this while both stacks serve real production models
simultaneously.
Consequence — reverted the boot-persistence flip. vllm_service_enabled
and vllm_service_state are back to role defaults (false/stopped) in
host_vars/astro-orbiter/vars.yml. vLLM stays staged (venv, model weights,
systemd units all in place) and can be started for a brief shadow-validation
window (same pattern as t_ca1af9fb's original Phase 5), but is NOT safe to
leave resident in production alongside llama-swap.
Path forward — requires a human decision, not more role tuning:
- Full llama-swap teardown (t_6dff1ecc) BEFORE vLLM gets permanent residency — but that breaks the 21 agent profiles' aux-model tasks and OpenViking's VLM unless those consumers are migrated to a different backend first (Anthropic API, a second smaller local box, or a redesigned single-process serving strategy that covers all the models vLLM and llama-swap currently split between them).
- Accept vLLM as a shadow-only / on-demand stack (manually started for specific validated windows, stopped otherwise) and do NOT attempt permanent Hindsight cutover — keeps llama-swap as the sole continuous production serving layer, matching the pre-task state.
- A hardware change (larger GPU, or a second GPU) — out of scope for this task, flagging for Ryan's awareness if the aux-model consumer set is expected to grow.
Comment posted on t_6dff1ecc with this finding — the teardown task remains correctly blocked; this task's completion does NOT unblock it, because full cutover to vLLM is not achievable within this card's VRAM budget as currently scoped.
RESOLVED (t_5508360a, 2026-08-31/09-01): Dashboard decision applied — llama-swap retired, vLLM permanent, 2 of 3 models
Human decision (dashboard, kanban t_5508360a): "stop and disable llama-swap and start vLLM and its 3 models" — explicit approval, "I understand this is a breaking change." Chose path 1 from the three options above: retire llama-swap, give vLLM permanent residency, accept that the 21 Hermes profiles' aux-model consumers lose their llama-swap aux roster (Qwen3-8B, Phi-3.5-mini, Meta-Llama-3.1-8B, Qwen2.5-Coder-14B — all 4 gone) in exchange for vLLM's stack. Mid-run the dashboard added a course-correction: "Don't try to load all 3 models concurrently on first deploy. Start with Qwen2.5-32B only" — received after the 3-model attempt below had already run and self-corrected to the same 2-model end state, so no further action needed, but noted for the record.
Executed:
sudo systemctl stop llama-swap && sudo systemctl disable llama-swapon astro-orbiter — confirmed inactive+disabled, VRAM dropped to 9MiB/24576MiB (from 20.6GB in production use).- Flipped
vllm_service_enabled/vllm_service_statetotrue/startedinhost_vars/astro-orbiter/vars.yml— vLLM is now the permanent, boot-persistent serving layer (was shadow-only/staged before this task). - Attempted the literal "3 models" instruction — flipped
Qwen3-8B-AWQ.enabledtotruetoo. Does not fit. With the 24GB card's usable 23.55GiB budget consumed by Qwen2.5-32B-Instruct-AWQ (~18.6GB weights) + nomic-embed-text-v1.5 (~0.8GB actual), only ~1.25GiB remained free — below the 3.53GiB floorgpu_memory_utilization=0.15requires for Qwen3-8B-AWQ even withenforce_eager. Confirmed viajournalctl: identicalValueError: Free memory on device cuda:0 (1.25/23.55 GiB) on startup is less than desired GPU memory utilizationon all 7 consecutive systemd restart attempts — not the transient CUDA-graph-capture crash-loop t_e6facb19 solved with enforce_eager, a hard ceiling. Stopped + disabledvllm-Qwen3-8B-AWQ.service, revertedenabled: falsein host_vars with a full writeup in the comment block. - Re-ran
day1_deploy_vllm.yml --extra-vars vllm_service_state=startedwith the corrected 2-model config: clean idempotent pass, changed=0 on both remaining models, Phase 5 verification passed (/health200 on both:8000and:8020,/v1/modelscorrect, live completion + live embeddings smoke tests both passed),NRestarts=0on both services. - Cut over Hindsight's LLM endpoint (the other production consumer):
HINDSIGHT_API_LLM_BASE_URLllama-swap:8001→ vLLM:8000,HINDSIGHT_API_LLM_MODEL→Qwen2.5-32B-Instruct-AWQ, addedHINDSIGHT_API_RETAIN_MAX_COMPLETION_TOKENS=4096(vLLM'smax_model_len=8192vs Hindsight's 64000 default), and switched the ExternalSecret'sHINDSIGHT_API_LLM_API_KEYsource from the unusednous1Password item tovllm's realapi-key(vLLM validates its bearer token; llama-swap never did). Committed tocluster/applications/hindsight/{values.yaml,externalsecret.yaml}, pushed, ArgoCD synced, confirmed the new pod loggedConnection verified: openai/Qwen2.5-32B-Instruct-AWQon boot. - Live end-to-end verification, not inference: a real
POST /v1/default/banks/war-machine/memoriesretain call against the production Hindsight endpoint returnedHTTP 200with genuine fact-extraction token usage (3257 in / 245 out), and a subsequentPOST .../memories/recallreturned real semantically-ranked results including the just-retained memory.
Final production state on astro-orbiter (verified live):
vllm.service(Qwen2.5-32B-Instruct-AWQ, :8000): active, enabled, boot-persistentvllm-nomic-embed-text-v1.5.service(:8020): active, enabled, boot-persistentvllm-Qwen3-8B-AWQ.service(:8010): inactive, disabled — does not fit, see abovellama-swap.service: inactive, disabled (unit files left in place — full removal is t_6dff1ecc's job, tracked separately)- VRAM: ~22.6GB/24.576GB in steady-state use, no crash-looping
What this means for t_6dff1ecc (teardown) and the 21 aux-model profiles: llama-swap is now stopped+disabled — t_6dff1ecc's actual teardown steps (remove systemd unit files, wipe caches) are now safe to execute and unblocked from a "live production" standpoint. However, this trades away the aux-model roster: the 21 Hermes profiles' aux-model tasks (skills_hub, approval, mcp, title_generation, profile_describer, compression) that used to route to llama-swap's Qwen3-8B/Phi-3.5-mini/Meta-Llama/Coder models now have zero local aux-model backend — Qwen3-8B-AWQ doesn't fit vLLM's VRAM budget either. This was accepted explicitly by the dashboard ("I understand this is a breaking change") — no further local aux-model migration was authorized or attempted in this task. If those 21 profiles need a replacement aux-model path, that is separate, new, explicitly-scoped follow-up work, not implied by this decision.
SUPERSEDED (t_r1d32b_swap, 2026-09-01): Qwen2.5-32B-Instruct-AWQ retired, replaced with DeepSeek-R1-Distill-Qwen-32B-AWQ, single-model deployment
Ryan direction: "Swap Qwen2.5-32B for DeepSeek-R1-Distill-Qwen-32B,
max-model-len 32768. Single model only." Confirmed with Ryan that "single
model only" includes disabling nomic-embed-text-v1.5 (:8020) as well —
nothing in production consumed it (Hindsight uses its own bundled 384-dim
embedder; OpenViking pointed at the retired llama-swap endpoint). DeepSeek
gets the entire 24GB card.
Model choice: casperhansen/deepseek-r1-distill-qwen-32b-awq — same
AutoAWQ toolchain/quant style as the outgoing Qwen2.5-32B-Instruct-AWQ,
widely-used community quant, Qwen2ForCausalLM architecture (DeepSeek-R1
reasoning distilled onto a Qwen2.5-32B base) — no new vLLM code path
required. Native max_position_embeddings: 131072; capped at 32768 per
the task's explicit requirement.
Executed:
- Stopped + disabled
vllm-nomic-embed-text-v1.5.service(single-model requirement), freed its ~19GB Qwen2.5-32B model cache on disk (30GB free → 48GB free) to make room for DeepSeek's ~19.3GB download. - Replaced
vllm_modelsinhost_vars/astro-orbiter/vars.yml: primary entry nowDeepSeek-R1-Distill-Qwen-32B-AWQ, aux (Qwen3-8B-AWQ) and embedding (nomic-embed-text-v1.5) bothenabled: false. - Staged the model via
--tags vllm-models(idempotenthf download, ~19GB, confirmed viadu -shand snapshot-dir stat). - Three rounds of live VRAM-fit debugging before a stable config was
found (documented inline in host_vars comments) — worth recording here
since the failure mode is non-obvious and will recur for future
32B-class models at high context on this 24GB card:
- Round 1 (fp16 KV, gpu_memory_utilization 0.90/0.95/0.98): vLLM's own pre-flight check reported 18.17GiB weights + 8.0GiB KV cache needed at 32768 ctx fp16 = 26.17GB — mathematically impossible on a 24GB card at ANY utilization percentage. Crash-looped every attempt.
- Round 2 (
--kv-cache-dtype fp8): halved nominal KV cache to ~4.0-4.3GiB, should fit with ~1GB margin. Still OOM'd — small (~50-150MB)cudaMallocfailures during FlashInfer kernel warmup, consistently, even when vLLM's own pre-flight math said it should fit. Root cause: real GPU usage during warmup kernel compilation exceeds what upfront profiling/reservation accounts for by roughly ~1GB (unaccounted FlashInfer/sampler warmup workspace buffers). Tried both the percentage knob AND vLLM's own suggested--kv-cache-memory-bytesexact value — same failure either way, confirming the gap wasn't a rounding/estimation error in the percentage math, it was a real missing ~1GB of margin. - Round 3 (
--kv-cache-dtype int4_per_token_head, fixed): SUCCESS. Switching from 8-bit to 4-bit KV cache roughly halves the KV footprint again (~2GiB instead of ~4-4.3GiB), buying back enough real headroom to absorb the unaccounted warmup overhead. Clean single-attempt start,NRestarts=0, steady-state VRAM 23.2GB/24.576GB.
- Full Ansible verify phase (
--tags vllm-api-key,vllm-verify) passed: systemd unit active,/health200,/v1/modelsreturnsDeepSeek-R1-Distill-Qwen-32B-AWQwithmax_model_len: 32768, live/v1/completionssmoke test HTTP 200, clean restart + re-run of--tags vllm-systemdconfirmed idempotent (changed=0,NRestarts=0, sameActiveEnterTimestamp— no unnecessary restart). - Manual end-to-end generation test, not inference: a real
/v1/chat/completionscall ("What is 12*8?") returned a genuine DeepSeek-R1 reasoning trace in<think>tags followed by the correct answer (96) with correct step-by-step arithmetic shown — confirms the model is not just health-check-alive but actually reasoning correctly.
Role/template changes (reusable for future models on this host):
- Added
kv_cache_dtype(renders--kv-cache-dtype) andkv_cache_memory_bytes(renders--kv-cache-memory-bytes) as new optional per-model fields invllm.service.j2— both are{% if ... is defined %}guarded, no effect on models that don't set them.
Final production state on astro-orbiter (verified live, 2026-09-01):
vllm.service(DeepSeek-R1-Distill-Qwen-32B-AWQ, :8000,max_model_len: 32768,kv_cache_dtype: int4_per_token_head): active, enabled, boot-persistent, single model on the cardvllm-nomic-embed-text-v1.5.service(:8020): inactive, disabledvllm-Qwen3-8B-AWQ.service(:8010): inactive, disabled (unchanged from prior state)llama-swap.service: inactive, disabled (unchanged from prior state)- VRAM: ~23.2GB/24.576GB steady-state, no crash-looping,
NRestarts=0
Not done in this task (flagging, not implied by this swap):
- Hindsight's
HINDSIGHT_API_LLM_MODEL/HINDSIGHT_API_LLM_BASE_URLcluster config still referencesQwen2.5-32B-Instruct-AWQ— that model is now gone from the card. Hindsight's LLM calls to astro-orbiter will fail model-not-found until that GitOps config is updated to point atDeepSeek-R1-Distill-Qwen-32B-AWQ. Not touched here — task scope was the astro-orbiter model swap itself, cluster consumer cutover is a separate, explicit follow-up (same boundary respected in the prior t_5508360a section: this role does not own cluster-side config). - DeepSeek-R1's reasoning output uses
<think>tags and the model card recommends temperature 0.5-0.7 (not greedy/0) — neither is enforced server-side; any consumer wiring this model into a Hermes profile or application should account for both when parsing responses.
Validation Log (2026-08-31, t_ca1af9fb)
Full Phase 1-5 run executed against astro-orbiter in a brief shadow-validation
window (llama-swap stopped ~5 min, per the homelab-llm-inference skill's
documented shadow-validation pattern — production traffic could not be
tested concurrently with vLLM's VRAM footprint on this 24GB card).
Two real bugs found and fixed during first-start validation (not present in the original spec, discovered only by actually starting the service):
ninjanot on systemd's PATH. vLLM's torch.compile path shells out to the bareninjacommand.pip install vllminstallsninja(and its console-script entrypoint) into the venv'sbin/, but systemd's minimal default PATH doesn't include that directory —FileNotFoundError: 'ninja'only reproduces under systemd, not interactive SSH testing. Fixed by settingEnvironment="PATH=<venv>/bin:...standard dirs..."in the unit template.- FlashInfer sampler JIT fails to compile on RTX 3090 (SM86).
flashinfer/data/csrc/sampling.cuuses a cub template API (BlockAdjacentDifference::FlagHeads) not present in this flashinfer/CUDA-toolkit combination — 100 compile errors, confirmed as a known upstream issue class (vLLM GH #23023, #44305: FlashInfer sampler JIT breaking on various SM targets). Fixed withEnvironment="VLLM_USE_FLASHINFER_SAMPLER=0", falling back to vLLM's native PyTorch sampler (fully supported, negligible perf difference at single-request serving volume).
Also corrected vllm_gpu_memory_utilization from 0.90 to 0.95 — at 0.90 the
KV cache allocation failed (2.0 GiB KV cache needed, 1.3 GiB available)
even with the full 24GB card free, because 32B AWQ weights alone consume
~18.4GB, leaving too little headroom at a 90% cap.
Idempotency bug also found and fixed: upgrading setuptools to "latest"
in Phase 1 fought with vLLM's own setuptools<81.0.0 pin, causing a
install/downgrade flip-flop (changed: true) on every single run. Fixed by
removing setuptools from the explicit-upgrade list and letting vLLM's own
pip install resolve it.
Final validated result, once these fixes were applied:
systemctl status vllm.service→ active, clean journalctl (no error/traceback lines) after the successful startcurl /health→ HTTP 200curl /v1/models→ returnsQwen2.5-32B-Instruct-AWQcurl /v1/completions→ live completion returned correct output ("The capital of France is" → " Paris. Correct! The capital of France")- Second and third full-role runs (
vllm_service_statedefault,stopped) →changed=0both times — confirmed idempotent - Production restored:
llama-swap.serviceactive,/health200,/v1/embeddingsagainstnomic-embed-text-v1.5returns a valid vector — Hindsight retain path confirmed still working after the shadow window - Post-restore VRAM: 486 MiB used / 24,576 MiB total (normal quiescent state)
Testing this role (idempotency)
Second-run test (staging phases only, safe to run repeatedly):
ansible-playbook -i inventory.yml playbooks/day1_deploy_vllm.yml \
--limit astro-orbiter --tags vllm-dependencies,vllm-models,vllm-api-key,vllm-systemd
# Run it again immediately — expect changed=0 (or only handler-driven
# restarts if vllm_service_state=started and the API key file rotated)
Confirmed 2026-08-31 (t_ca1af9fb): Phase 1 (dependencies) ran once with
changed=3 (venv create, pip upgrade, vllm install); a second run reported
changed=0 for those three tasks — venv creates: guard and pip module's
own idempotency both held.
Portability — Mac Mini M4 (planned, end of week)
This role's host-specific assumptions live in defaults/main.yml (all
overridable via host_vars/<host>/vars.yml) plus one hard assumption baked
into tasks/dependencies.yml: an NVIDIA GPU (nvidia-smi check, CUDA
wheels). Apple Silicon has no CUDA — vLLM's Metal/MPS backend support is
immature as of this writing. Before reusing this role for the Mac Mini M4:
- Fork
tasks/dependencies.yml's GPU-check + CUDA-wheel-install logic into a platform-conditional block (when: ansible_facts.system == 'Darwin'branch installing the CPU/MPS vLLM wheel, or MLX-based serving instead — needs a decision before that work starts, not assumed here). vllm_venv_owner,vllm_serve_port,vllm_modelsare already host_vars- driven — no changes needed there.- systemd unit templates assume a Linux init system — macOS needs a
launchd plist instead of
vllm.service.j2.
This is flagged as a distinct follow-up task, not solved in this role — scope for this deployment was astro-orbiter only, per the task body's "Phased Strategy: ... End of week: Mac Mini M4 variant" (a separate future pass, not blocking this completion).
Files
roles/deploy-vllm/
├── defaults/main.yml # all tunables — host overrides go in host_vars
├── handlers/main.yml # reload systemd / restart vllm services
├── meta/main.yml
├── tasks/
│ ├── main.yml # phase orchestrator
│ ├── dependencies.yml # Phase 1
│ ├── models.yml # Phase 2
│ ├── api-key.yml # Phase 3
│ ├── systemd.yml # Phase 4
│ └── verify.yml # Phase 5
├── templates/
│ ├── vllm.service.j2 # one instance per enabled model
│ └── vllm-workspace.sh.j2 # debugging helper deployed to the target
└── README.md # this file