Hermes Agent service account
782cbe33d1
llm-inference: size ctx-size/parallel for aux task offload
...
Previous ctx-size=8192/parallel=4 gave 2048 tokens/slot, too small for
context compression inputs (observed live rejection at 3826 tokens).
Measured VRAM on astro-orbiter (RTX 3090 24GB): weights ~17GB resident,
~294KiB/token pool-wide for KV cache+buffers at prior sizing.
New: ctx-size=16384, parallel=2 -> 8192 tokens/slot (matches model's
native n_ctx_train max). Projected VRAM ~21.8GB, ~2.7GB headroom.
Applied directly via ansible-playbook (Semaphore currently broken --
fix tracked separately).
2026-08-05 12:14:11 -05:00
Hermes Agent service account
aa8e229e64
fix(llm-inference): switch serve phase from vLLM+bitsandbytes to llama.cpp+GGUF
...
bitsandbytes peak RAM ~54GB (bf16 load before quantize) — kills 40GB OptiPlex.
llama.cpp Q4_K_M GGUF loads pre-quantized: peak RAM ~15.5GB, fits cleanly.
Changes:
- serve.yml: build llama.cpp with CUDA, download Q4_K_M GGUF from bartowski,
disable vllm-serve, deploy llama-server.service
- llama-server.service.j2: OpenAI-compatible server on same port 8000,
--n-gpu-layers 99 (full GPU offload), --parallel 4, gemma chat template
- defaults: llm_gguf_dir, llm_gguf_path, llm_gpu_layers, llm_parallel_slots
- handlers: restart llama-server, vllm-serve failed_when=false (may not exist)
GGUF: bartowski/gemma-2-27b-it-Q4_K_M.gguf (15.5GB, 24GB VRAM fits w/ ~8GB headroom)
2026-08-03 12:37:03 -05:00
Hermes Agent service account
22a020e4c7
fix(llm-inference): bitsandbytes int4 OOM — pending switch to llama.cpp+GGUF
...
bitsandbytes quantizes on-the-fly: loads full bf16 weights (~54GB RAM peak)
before compressing to int4. Kills the 40GB OptiPlex on torch.compile warmup.
Fix in next commit: switch serve phase to llama.cpp + GGUF Q4_K_M.
Pre-quantized weights load directly — peak RAM ~16GB, no compile overhead.
2026-08-03 12:35:46 -05:00
Hermes Agent service account
e879cf73d3
fix(llm-inference): gpu_exporter version 1.2.2 → 1.13.1 (correct release tag)
2026-08-03 11:57:42 -05:00
Hermes Agent service account
423891001c
feat(llm-inference): Phase 7 — Prometheus monitoring + Grafana dashboard
...
- Phase 7 task file: monitoring.yml
- node_exporter (port 9100) via apt, systemd managed
- nvidia_gpu_exporter v1.2.2 (port 9835) — GPU util, VRAM, temp, power
- Patches kube-prometheus additionalScrapeConfigs secret with 3 new jobs:
node-astro-orbiter, gpu-astro-orbiter, vllm-astro-orbiter
- Deploys Grafana dashboard ConfigMap via kubectl apply
- Grafana dashboard (11 panels):
- Row 1: GPU util %, VRAM used, GPU temp gauge
- Row 2: GPU power draw, vLLM token throughput, request queue depth
- Row 3: vLLM e2e latency p50/p95/p99, KV cache utilization %
- Row 4: System CPU %, memory, root disk gauge
- defaults/main.yml: llm_gpu_exporter_version, llm_gpu_exporter_port
- handlers/main.yml: restart nvidia-gpu-exporter
2026-08-03 11:53:36 -05:00
Hermes Agent service account
dda6b91330
feat(llm-inference): Day 1 playbook for RTX 3090 vLLM stack on astro-orbiter
...
- nvidia-driver-595-open (already installed 2026-08-03, idempotent)
- Python venv + vLLM 0.26.0 (already installed, idempotent)
- Gemma 2 27B model download via HuggingFace hub
- systemd vllm-serve.service on port 8000
- Hermes provider integration on carousel-of-progress
- vault_hf_token added to group_vars/all/vault
- ansible.cfg: vault_password_file set to absolute path
- inventory: astro_orbiter group added
Run with: env -u ANSIBLE_VAULT_PASSWORD_FILE ansible-playbook -i inventory.yml playbooks/day1_deploy_llm_inference.yml
2026-08-03 11:51:34 -05:00