Ryan Blundon rblundon
  • Joined on 2026-05-07
rblundon pushed to main at rblundon/homelab 2026-08-25 10:51:05 -05:00
173d00504c hindsight: swap LLM astro-orbiter Qwen3.8-27B -> Nous free-tier stepfun/step-3.7-flash:free (t_90261bb1)
rblundon pushed to main at rblundon/homelab 2026-08-24 19:03:33 -05:00
7cdcc984a5 hindsight: Phase C manifests (multi-source app wave 8, external pgvector PG, ES from 1Password, chart-native ingress)
rblundon pushed to main at rblundon/homelab 2026-08-20 13:47:11 -05:00
e301770adc openviking: repoint embedding+vlm api_base :8002->:8001
rblundon pushed to main at rblundon/homelab 2026-08-19 12:43:21 -05:00
ab1e32711d Merge origin/main: sync Qwen3-8B no_think variant to Ansible repo (t_36e8ba68)
bafd76a0b4 feat(astro-orbiter): add Qwen3-8B-Q4_K_M to inference stack (t_c5cef2b2)
Compare 2 commits »
rblundon pushed to main at rblundon/homelab 2026-08-19 11:37:32 -05:00
5c0df8c73c Merge origin/main — integrate monitoring/Phase3 updates with Qwen3-8B no-think deployment
5cf4468754 Add Qwen3-8B no-think variant — dual thinking deployment (t_664289a0)
Compare 2 commits »
rblundon pushed to main at rblundon/homelab 2026-08-18 23:18:44 -05:00
24735f7e5c fix: correct metric names in llama-swap monitoring (llamacpp_* -> llamaswap_*), update alerts + dashboard + scrape config
rblundon pushed to main at rblundon/homelab 2026-08-18 22:23:21 -05:00
7867be688a monitoring: llama-swap GPU/LLM stack (v250) — PrometheusRule, Grafana dashboard, scrape config, VRAM exporter
03b3ce9dee llm-router: CPU-offload Coder-14B + Llama-3.1-8B (t_72646029)
Compare 2 commits »
rblundon pushed to main at rblundon/homelab 2026-08-16 22:39:46 -05:00
a2994bf55d feat(astro-orbiter): bump Qwen3.8-27B ctx-size 32768->131072 (128K) [t_441470b9]
rblundon pushed to main at rblundon/homelab 2026-08-16 20:41:01 -05:00
7b44a41da3 feat(llm): swap astro-orbiter primary model Qwen3.6 -> Qwen3.8-27B-Q4_K_M
rblundon pushed to main at rblundon/homelab 2026-08-15 00:13:02 -05:00
efaff340a4 openviking: fix VLM model alias and lower max_input_tokens to 1024
rblundon pushed to main at rblundon/homelab 2026-08-14 23:22:41 -05:00
48536f2615 fix(openviking): cap embedding max_input_tokens at 1536 to stay under llama.cpp nomic-bert 2048 ctx limit
rblundon pushed to main at rblundon/homelab 2026-08-14 12:49:54 -05:00
170a31d090 feat(openviking): deploy maelstrom-ui Web Studio frontend
rblundon pushed to main at rblundon/homelab 2026-08-14 00:16:29 -05:00
aa2730efd5 fix: OpenViking ingress TLS issuer from letsencrypt-internal to letsencrypt-prod
rblundon pushed to main at rblundon/homelab 2026-08-13 23:49:36 -05:00
0dbb77b023 Fix OpenViking: 1Password item mismatch + invalid embedding config fields
rblundon pushed to main at rblundon/homelab 2026-08-13 23:43:13 -05:00
fee9965d0a fix: OpenViking sync-wave deadlock - move ExternalSecret ordering inside Application
rblundon pushed to main at rblundon/homelab 2026-08-13 23:37:14 -05:00
d0f3ddba0d OpenViking application.yaml: fix invalid Helm chart source (chart -> path)
rblundon pushed to main at rblundon/homelab 2026-08-13 23:33:58 -05:00
d9e41118f8 feat(openviking): pilot deployment to fastpass (wave 8)
rblundon pushed to main at rblundon/homelab 2026-08-13 23:18:43 -05:00
ad70b3439c feat(llm): add nomic-embed-text-v1.5-Q4_K_M to astro-orbiter router
rblundon pushed to main at rblundon/homelab 2026-08-13 18:47:18 -05:00
a04435ee9b fix(monitoring): drop Qwen3.6 scrape job — causes CUDA OOM on each scrape (t_02c15dae)
rblundon pushed to main at rblundon/homelab 2026-08-13 15:44:13 -05:00
a2ddb65425 fix(monitoring): reduce llama-server scrape_interval 15s -> 90s to allow GPU P8 idle