Commit Graph

  • 7e4b103e68 Clean up duplicate OPENAI_BASE_URL env vars; keep only one main Hermes Agent service account 2026-09-01 13:58:44 -05:00
  • 404ff3d91d Set OPENAI_API_BASE_URLS (plural) for model discovery Hermes Agent service account 2026-09-01 13:49:20 -05:00
  • 3c6f6dfe2e Fix: Use local astro-orbiter endpoint, not OpenAI API Hermes Agent service account 2026-09-01 13:48:37 -05:00
  • 56db1b94ee Disable OpenAI API key validation for local vLLM endpoint Hermes Agent service account 2026-09-01 13:44:41 -05:00
  • f0387c1033 Simplify: use vllm-api-key for WEBUI_SECRET_KEY Hermes Agent service account 2026-09-01 13:39:14 -05:00
  • dd5ce10910 Fix ExternalSecret template: remove webui-secret-key reference Hermes Agent service account 2026-09-01 13:38:11 -05:00
  • 1e537cf5f7 Add storageClassName to PVC manifest for explicit NFS declaration Hermes Agent service account 2026-09-01 13:31:18 -05:00
  • f4b1fc9e71 Fix ArgoCD Application destination namespace Hermes Agent service account 2026-09-01 13:30:24 -05:00
  • f51d1c16ac Fix Open WebUI storage and model detection Hermes Agent service account 2026-09-01 13:28:21 -05:00
  • e80a1dc088 Temporarily disable webui-secret-key from ExternalSecret Hermes Agent service account 2026-09-01 13:19:18 -05:00
  • 3dc58cf644 Fix Open WebUI auth: correct WEBUI_SECRET_KEY and remove Ollama config Hermes Agent service account 2026-09-01 13:18:55 -05:00
  • d95477fc3b Fix Open WebUI connectivity to astro-orbiter Hermes Agent service account 2026-09-01 13:14:03 -05:00
  • 274ce1fd8a fix http route Ryan Blundon 2026-09-01 13:00:22 -05:00
  • eed2fcb7c7 feat: add HTTPRoute and remove old nginx Ingress for GitOps Hermes Agent service account 2026-09-01 12:53:10 -05:00
  • 266b6c7be1 chore: apply all changes Hermes Agent service account 2026-09-01 12:28:16 -05:00
  • e9924a2524 rename open-webui aplication/namespace Ryan Blundon 2026-09-01 11:20:19 -05:00
  • 5ee8309d32 Remove readiness probe (too aggressive); keep liveness probe Hermes Agent service account 2026-08-31 22:52:45 -05:00
  • 261f6be7db Fix readiness probe to use /health endpoint (no auth needed) Hermes Agent service account 2026-08-31 22:51:05 -05:00
  • 8ea19dbf70 Fix ExternalSecret API format and deployment env vars Hermes Agent service account 2026-08-31 22:47:34 -05:00
  • e6cb187f8e Deploy Body Wars Observability WebUI (Open WebUI → astro-orbiter vLLM) Hermes Agent service account 2026-08-31 22:46:06 -05:00
  • 56f19af578 feat(hindsight): cut over LLM model to Gemma-4-26B-A4B-it-AWQ (t_gemma4_swap) Hermes Agent service account 2026-08-31 21:54:04 -05:00
  • a3c92f70bf feat(deploy-vllm): swap DeepSeek-R1-Distill-Qwen-32B for Gemma 4 26B A4B AWQ (t_gemma4_swap) Hermes Agent service account 2026-08-31 21:28:48 -05:00
  • f907acde95 feat(hindsight): cut over LLM model to DeepSeek-R1-Distill-Qwen-32B-AWQ (t_r1d32b_swap) Hermes Agent service account 2026-08-31 20:23:38 -05:00
  • 53a55e7317 feat(deploy-vllm): swap Qwen2.5-32B for DeepSeek-R1-Distill-Qwen-32B-AWQ (t_r1d32b_swap) Hermes Agent service account 2026-08-31 20:21:27 -05:00
  • 2c0db1c7a1 docs: record t_5508360a resolution in deploy-vllm README Hermes Agent service account 2026-08-31 19:10:53 -05:00
  • 39c5fdca69 hindsight: cut over LLM endpoint to vLLM (t_5508360a) Hermes Agent service account 2026-08-31 19:06:11 -05:00
  • 6bfcc76845 vllm: cutover to permanent residency, retire llama-swap (t_5508360a) Hermes Agent service account 2026-08-31 19:03:54 -05:00
  • 1af645d272 REVERT: vLLM cannot be continuously resident alongside llama-swap (t_e6facb19) Hermes Agent service account 2026-08-31 18:36:04 -05:00
  • f3a5687adf hindsight: cap RETAIN_MAX_COMPLETION_TOKENS for vLLM's 8192 ctx Hermes Agent service account 2026-08-31 18:29:22 -05:00
  • 9d6869ad9d hindsight: revert embeddings cutover after dimension-mismatch crash Hermes Agent service account 2026-08-31 18:17:25 -05:00
  • 2cc9370f3d deploy-vllm: add embedding-mode support, cut over Hindsight to vLLM (t_e6facb19) Hermes Agent service account 2026-08-31 18:15:22 -05:00
  • 60220e18b6 feat(astro-orbiter): add deploy-vllm Ansible role (t_ca1af9fb) Hermes Agent service account 2026-08-31 17:37:37 -05:00
  • b3b925ff77 hindsight: cap LLM concurrency to 1 + raise client/ingress timeout to 600s (fix 502s on serial astro-orbiter) Hermes Agent service account 2026-08-29 16:56:52 -05:00
  • 3d8eb1bf1c hindsight: restore LLM to local astro-orbiter Qwen3.8-27B (Nous retain broken) (t_e0e6f7ca) Hermes Agent service account 2026-08-29 12:44:01 -05:00
  • 7f8ba8b859 hindsight: swap LLM stepfun/step-3.7-flash:free -> upstage/solar-pro4:free (t_8516dba2) Hermes Agent service account 2026-08-28 23:35:54 -05:00
  • aee61d4511 hindsight: interim swap LLM to Nous stepfun/step-3.7-flash:free (astro-orbiter down) Hermes Agent service account 2026-08-28 23:12:48 -05:00
  • 9bc29508d7 hindsight: move LLM to local astro-orbiter Qwen3.8-27B (off Nous free tier) Hermes Agent service account 2026-08-28 21:00:47 -05:00
  • 9bfc9384e4 hindsight: re-promote upstage/solar-pro4:free as primary LLM (ingress timeout fixed) Hermes Agent service account 2026-08-25 12:05:54 -05:00
  • 152230c100 hindsight: raise nginx ingress proxy read/send timeout to 300s (fixes reflect 504) Hermes Agent service account 2026-08-25 12:04:49 -05:00
  • 13df80ab43 Revert "hindsight: promote solar-pro4:free (tool-calling) to primary LLM — fixes reflect 500 (t_d0dffc3d)" Hermes Agent service account 2026-08-25 11:42:06 -05:00
  • a2123819b3 hindsight: promote solar-pro4:free (tool-calling) to primary LLM — fixes reflect 500 (t_d0dffc3d) peter-parker 2026-08-25 11:34:14 -05:00
  • 173d00504c hindsight: swap LLM astro-orbiter Qwen3.8-27B -> Nous free-tier stepfun/step-3.7-flash:free (t_90261bb1) Hermes Agent service account 2026-08-25 10:50:52 -05:00
  • 7cdcc984a5 hindsight: Phase C manifests (multi-source app wave 8, external pgvector PG, ES from 1Password, chart-native ingress) Hermes Agent service account 2026-08-24 19:03:07 -05:00
  • e301770adc openviking: repoint embedding+vlm api_base :8002->:8001 Maria Hill 2026-08-20 13:46:41 -05:00
  • ab1e32711d Merge origin/main: sync Qwen3-8B no_think variant to Ansible repo (t_36e8ba68) Hermes Agent service account 2026-08-19 12:42:46 -05:00
  • 5c0df8c73c Merge origin/main — integrate monitoring/Phase3 updates with Qwen3-8B no-think deployment Hermes Agent service account 2026-08-19 11:36:46 -05:00
  • 5cf4468754 Add Qwen3-8B no-think variant — dual thinking deployment (t_664289a0) Hermes Agent service account 2026-08-19 11:35:08 -05:00
  • bafd76a0b4 feat(astro-orbiter): add Qwen3-8B-Q4_K_M to inference stack (t_c5cef2b2) Hermes Agent service account 2026-08-19 11:17:07 -05:00
  • 24735f7e5c fix: correct metric names in llama-swap monitoring (llamacpp_* -> llamaswap_*), update alerts + dashboard + scrape config Hermes Agent service account 2026-08-18 23:18:22 -05:00
  • 7867be688a monitoring: llama-swap GPU/LLM stack (v250) — PrometheusRule, Grafana dashboard, scrape config, VRAM exporter Hermes Agent service account 2026-08-18 22:22:53 -05:00
  • 03b3ce9dee llm-router: CPU-offload Coder-14B + Llama-3.1-8B (t_72646029) Hermes Agent service account 2026-08-17 17:06:37 -05:00
  • a2994bf55d feat(astro-orbiter): bump Qwen3.8-27B ctx-size 32768->131072 (128K) [t_441470b9] Hermes Agent service account 2026-08-16 22:39:12 -05:00
  • 7b44a41da3 feat(llm): swap astro-orbiter primary model Qwen3.6 -> Qwen3.8-27B-Q4_K_M Hermes Agent service account 2026-08-16 20:40:31 -05:00
  • efaff340a4 openviking: fix VLM model alias and lower max_input_tokens to 1024 Hermes Agent service account 2026-08-15 00:12:40 -05:00
  • 48536f2615 fix(openviking): cap embedding max_input_tokens at 1536 to stay under llama.cpp nomic-bert 2048 ctx limit Hermes Agent service account 2026-08-14 23:22:12 -05:00
  • 170a31d090 feat(openviking): deploy maelstrom-ui Web Studio frontend Hermes Agent service account 2026-08-14 12:49:00 -05:00
  • aa2730efd5 fix: OpenViking ingress TLS issuer from letsencrypt-internal to letsencrypt-prod Peter Parker 2026-08-14 00:16:06 -05:00
  • 0dbb77b023 Fix OpenViking: 1Password item mismatch + invalid embedding config fields Hermes Agent service account 2026-08-13 23:48:59 -05:00
  • fee9965d0a fix: OpenViking sync-wave deadlock - move ExternalSecret ordering inside Application Hermes Agent service account 2026-08-13 23:42:45 -05:00
  • d0f3ddba0d OpenViking application.yaml: fix invalid Helm chart source (chart -> path) Hermes Agent service account 2026-08-13 23:36:50 -05:00
  • d9e41118f8 feat(openviking): pilot deployment to fastpass (wave 8) Hermes Agent service account 2026-08-13 23:33:36 -05:00
  • ad70b3439c feat(llm): add nomic-embed-text-v1.5-Q4_K_M to astro-orbiter router Hermes Agent service account 2026-08-13 23:18:08 -05:00
  • a04435ee9b fix(monitoring): drop Qwen3.6 scrape job — causes CUDA OOM on each scrape (t_02c15dae) Hermes Agent service account 2026-08-13 18:46:54 -05:00
  • a2ddb65425 fix(monitoring): reduce llama-server scrape_interval 15s -> 90s to allow GPU P8 idle Hermes Agent service account 2026-08-13 15:43:43 -05:00
  • a87da82ebd fix: correct astro-orbiter llama-server scrape target for router cutover Hermes Agent service account 2026-08-13 09:08:01 -05:00
  • 7aea88724f Add Qwen2.5-Coder-14B-Instruct-4bit to astro-orbiter router (t_55c164f5) Hermes Agent service account 2026-08-13 09:07:02 -05:00
  • 6455d22752 feat(llm-router): add Meta-Llama-3.1-8B-Instruct-4bit alias; document Phi json_schema limitation (t_9adf0889) Hermes Agent service account 2026-08-12 23:33:20 -05:00
  • a47b29d49f feat(llm-router): switch to --models-preset mode; add Phi-3.5-mini-instruct-8bit alias (t_9adf0889) Hermes Agent service account 2026-08-12 22:58:12 -05:00
  • 9c969f783d feat(llm-inference-multimodel): bump router --models-max 1->4 on astro-orbiter (t_33acbb2e) Hermes Agent service account 2026-08-12 22:26:49 -05:00
  • 081156ecab feat(llm-inference-multimodel): codify Phi-3.5-mini + Llama-3.1-8B GGUF staging (t_730f9584) Hermes Agent service account 2026-08-12 22:19:34 -05:00
  • 3783ded62a fix: update router unit template comment — no longer a shadow deployment (t_cd0d5388) Hermes Agent service account 2026-08-12 20:43:31 -05:00
  • 5a2246a540 feat: add day2_cutover_qwen_to_router.yml playbook (t_cd0d5388) Hermes Agent service account 2026-08-12 20:41:11 -05:00
  • ba311a3ec6 feat(llm-inference): add llama.cpp router mode shadow deployment Hermes Agent service account 2026-08-12 20:23:00 -05:00
  • d1f97ad5ac Phase 2 revised: consolidate astro-orbiter to single Qwen2.5-14B-1M model (port 8002) Hermes Agent service account 2026-08-06 11:42:34 -05:00
  • b4bdb63e4a llm-inference-multimodel: correct stale VRAM estimate for qwen-1m shadow slot Hermes Agent service account 2026-08-06 10:40:13 -05:00
  • b741f9b20b llm-inference-multimodel: repoint qwen shadow slot to Qwen2.5-14B-Instruct-1M (base Qwen disqualified, n_ctx_train=32768) Hermes Agent service account 2026-08-06 10:39:45 -05:00
  • a3c1342837 llm-inference-multimodel: reset qwen shadow unit to disabled by default -- model disqualified (n_ctx_train=32768, not 64K+), leaving enabled would crash-loop on next playbook run Hermes Agent service account 2026-08-06 09:38:32 -05:00
  • d4ff2681ac llm-inference-multimodel: fix qwen unit -- llama.cpp requires --flash-attn <on|off|auto>, not bare flag Hermes Agent service account 2026-08-06 09:24:42 -05:00
  • 75cb93f25c llm-inference-multimodel: enable Qwen2.5-14B shadow instance (port 8002) for shadow-test window Hermes Agent service account 2026-08-06 09:10:06 -05:00
  • d10255297c llm-inference-multimodel: add Qwen2.5-14B shadow instance (port 8002, gated off — VRAM co-residency not yet confirmed) Hermes Agent service account 2026-08-06 09:08:33 -05:00
  • 79edb8f4e1 llm-inference-multimodel: log Run 2 validation PASS (tool-calling + hallucination), preserve procedure doc Hermes Agent service account 2026-08-05 17:25:24 -05:00
  • 5dc76a8348 llm-inference-multimodel: fix tool-calling support (jinja template + gpu-layers=20 for VRAM fit) Hermes Agent service account 2026-08-05 17:13:32 -05:00
  • a76ad3195c llm-inference-multimodel: fix verify.yml losing Gemma-stop gate when run with --tags verify Hermes Agent service account 2026-08-05 16:34:13 -05:00
  • 73ef806dd6 llm-inference-multimodel: stop pre-existing Gemma service before Phase 4 starts new instances Hermes Agent service account 2026-08-05 16:28:34 -05:00
  • 628dae06a8 llm-inference-multimodel: fix Phase 2 unexpectedly restarting both services Hermes Agent service account 2026-08-05 16:21:50 -05:00
  • c3755aa29e llm-inference-multimodel: role + day1 playbook (phase 0 discover approved) master Hermes Agent service account 2026-08-05 15:53:31 -05:00
  • 782cbe33d1 llm-inference: size ctx-size/parallel for aux task offload Hermes Agent service account 2026-08-05 12:14:11 -05:00
  • aff792a061 feat(llm-inference): move astro-orbiter monitoring to GitOps (values.yaml + dashboards.yaml) Hermes Agent service account 2026-08-05 09:43:54 -05:00
  • aa8e229e64 fix(llm-inference): switch serve phase from vLLM+bitsandbytes to llama.cpp+GGUF Hermes Agent service account 2026-08-03 12:37:03 -05:00
  • 22a020e4c7 fix(llm-inference): bitsandbytes int4 OOM — pending switch to llama.cpp+GGUF Hermes Agent service account 2026-08-03 12:35:46 -05:00
  • e879cf73d3 fix(llm-inference): gpu_exporter version 1.2.2 → 1.13.1 (correct release tag) Hermes Agent service account 2026-08-03 11:57:42 -05:00
  • 423891001c feat(llm-inference): Phase 7 — Prometheus monitoring + Grafana dashboard Hermes Agent service account 2026-08-03 11:53:36 -05:00
  • dda6b91330 feat(llm-inference): Day 1 playbook for RTX 3090 vLLM stack on astro-orbiter Hermes Agent service account 2026-08-03 11:51:34 -05:00
  • 265d3f8fd6 jmri: remove one-shot xpra migration task (idempotency fix) Hermes Agent service account 2026-08-01 22:15:11 -05:00
  • b61d19cb91 jmri: add udev rule for LCC buffer (Microchip CDC -> jmri-lcc) Hermes Agent service account 2026-08-01 22:09:19 -05:00
  • 00be18b1f1 jmri: move udev symlinks to /dev/jmri-* (flat, JMRI-enumerable) Hermes Agent service account 2026-08-01 21:58:42 -05:00
  • 6c7ec507ef jmri: fix NCE udev rule — FTDI FT232 (ttyUSB), not Microchip CDC (ttyACM) Hermes Agent service account 2026-08-01 21:55:43 -05:00
  • 63b0bc72fe jmri: deploy udev rules for stable /dev/jmri/* symlinks Hermes Agent service account 2026-08-01 21:18:06 -05:00
  • 02af5d26dc jmri: fix xpra remove task idempotency (skip if already from upstream repo) Hermes Agent service account 2026-08-01 21:07:32 -05:00
  • 3eb38b74bd jmri: add rblundon@laptop SSH key for xpra access Hermes Agent service account 2026-08-01 21:06:07 -05:00