OpenViking Phase 1b (t_34b96e83) — Ryan-approved implementation. Changes: - roles/llm-inference-multimodel/templates/llama-server-router-preset.ini.j2: Add [nomic-embed-text-v1.5] section with embedding=true, n-gpu-layers=99, ctx-size=8192, load-on-startup=true, sleep-idle-seconds=-1. No flash-attn or KV cache params (embedding models use bidirectional forward pass, not autoregressive KV cache). Var: llm_router_nomic_ctx_size. - roles/llm-inference-multimodel/defaults/main.yml: Add llm_router_nomic_ctx_size: 8192. - host_vars/astro-orbiter/vars.yml: Add nomic-embed-text-v1.5-Q4_K_M.gguf to llm_staged_models list (size_bytes: 84106624, source: nomic-ai/nomic-embed-text-v1.5-GGUF). Update VRAM note to reflect 5 registered models (nomic adds ~84MB, negligible given sleep-idle-seconds=-1 / load-on-startup=true pinning). - playbooks/day2_add_nomic_embed.yml: New day2 playbook following the coder-alias pattern: Phase 1: idempotent GGUF download (exact size check) Phase 2: redeploy preset INI Phase 3: redeploy + restart systemd unit Phase 4: /v1/models gate (all 5 models present) Phase 5: /v1/embeddings smoke test (vector returned, not empty) VRAM: ~84MB, always pinned. No impact on generative model LRU behavior. peter-parker Helm values already point at :8002 for the embedding endpoint.
168 lines
9.5 KiB
YAML
168 lines
9.5 KiB
YAML
---
|
|
# ------------------------------------------------------------------------------
|
|
# FILE: roles/llm-inference-multimodel/defaults/main.yml
|
|
# DESCRIPTION: Overridable defaults for the llm-inference-multimodel role.
|
|
# Deploy target: astro-orbiter (10.1.71.130, RTX 3090 24GB).
|
|
# Built ALONGSIDE roles/llm-inference (not a replacement) — that
|
|
# role's CUDA/build/driver phases are the prerequisite; this role
|
|
# assumes /opt/llama.cpp/build/bin/llama-server already exists.
|
|
#
|
|
# See /home/hermes/astro-orbiter-multi-model-plan.md for the full
|
|
# approved design (VRAM math, rationale, rollback story).
|
|
# ------------------------------------------------------------------------------
|
|
|
|
# Shared
|
|
llm_service_user: jarvis
|
|
llm_binary_path: /opt/llama.cpp/build/bin/llama-server
|
|
llm_models_dir: /opt/models
|
|
|
|
# Bind address — deliberately NOT 0.0.0.0 (see plan §5). Default to the private
|
|
# LAN interface so both instances are reachable from Hermes but not the world.
|
|
# Override to 127.0.0.1 if even LAN-wide reachability is unwanted and a reverse
|
|
# proxy/localhost-only tunnel is used instead.
|
|
llm_bind_address: "10.1.71.130"
|
|
|
|
# Firewall scoping (Phase 3) — subnet/hosts allowed to reach the ports above.
|
|
# Override per-environment; default assumes Hermes runs somewhere on this /24.
|
|
llm_allowed_source_cidr: "10.1.70.0/24"
|
|
|
|
# --- RETIRED (2026-08-06): Aux / classification instance (port 8000, Phi-4-14B)
|
|
# and Tool-calling instance (port 8001, Mistral-Small-24B) --------------------
|
|
# Consolidated down to a single production model (Qwen2.5-14B-Instruct-1M,
|
|
# port 8002) serving BOTH the friday and war-machine Hermes profiles. Ryan
|
|
# explicitly accepted the tradeoffs (single model for chat + tool-calling +
|
|
# aux duties) over keeping the aux/toolcall split running.
|
|
# Both llama-server-aux and llama-server-toolcall services were stopped,
|
|
# disabled, and had their unit files removed from astro-orbiter; their GGUF
|
|
# weights (phi-4-14b-instruct-Q4_K_M.gguf, mistral-small-24b-instruct-2501-
|
|
# Q3_K_M.gguf) were deleted from /opt/models (~45GB reclaimed). The
|
|
# templates/tasks that deployed them have been removed from this role — see
|
|
# git log for the prior variable definitions and unit templates if a future
|
|
# rollback needs them restored.
|
|
|
|
# --- Production instance (port 8002, Qwen2.5-14B-Instruct-1M) ----------------
|
|
# History (2026-08-06): Qwen2.5-14B-Instruct (base) was deployed to this slot
|
|
# and DISQUALIFIED — live /v1/models meta reported n_ctx_train=32768, well
|
|
# under the 64K Hermes floor (the model card's "128K" figure conflated
|
|
# YaRN-extended inference-time scaling with actual trained context; disabled
|
|
# by default, not baked in). Llama-3.1-8B-Instruct was tried next — cleared
|
|
# the context gate (verified live n_ctx_train=131072) but failed the
|
|
# tool-calling validation harness badly (8/10 hallucination-stress prompts
|
|
# triggered spurious tool_calls even at temp=0.1 with the correct official
|
|
# chat template) — purged from disk and Ansible entirely, see git log.
|
|
# Current model: Qwen2.5-14B-Instruct-1M (bartowski GGUF) — distinct
|
|
# checkpoint with genuine additional long-context pretraining, NOT the same
|
|
# weights as the disqualified base model above. Live-verified 2026-08-06:
|
|
# /v1/models reports n_ctx=65536, n_ctx_train=1010000 (well over the 64K
|
|
# floor). Tool-calling verified live via a /v1/chat/completions probe with a
|
|
# tools= payload — returned a well-formed tool_calls response (finish_reason
|
|
# "tool_calls", valid JSON arguments), no hallucinated calls observed.
|
|
# PROMOTED TO PRODUCTION (2026-08-06): llm_qwen_service_enabled now defaults
|
|
# to true — this is the sole model serving both Hermes profiles. Ports
|
|
# 8000/8001 are permanently freed; no co-residency VRAM gate applies anymore.
|
|
llm_qwen_service_enabled: true
|
|
llm_qwen_port: 8002
|
|
llm_qwen_model_path: "{{ llm_models_dir }}/Qwen3.6-35B-A3B-UD-Q4_K_S.gguf"
|
|
llm_qwen_model_min_bytes: 19000000000 # guard threshold; complete file ~20GB
|
|
llm_qwen_ctx_size: 65536
|
|
llm_qwen_parallel: 1
|
|
llm_qwen_gpu_layers: 99
|
|
llm_qwen_batch_size: 2048
|
|
llm_qwen_ubatch_size: 512
|
|
llm_qwen_service_name: llama-server-qwen
|
|
llm_qwen_model_id: Qwen3.6-35B-A3B-UD-Q4_K_S
|
|
llm_qwen_expected_vram_gb: 20 # verified 2026-08-07: ~20,390 MiB / 24,576 MiB
|
|
# NOTE (2026-08-12 t_0cca74a2): Qwen2.5-14B-Instruct-1M was superseded by
|
|
# Qwen3.6-35B-A3B-UD-Q4_K_S (task t_2ffc0f63, 2026-08-07). Defaults updated
|
|
# to reflect the current production model. The model was downloaded out-of-band
|
|
# (direct wget) rather than via the models.yml get_url pattern.
|
|
# llm_qwen_model_url is intentionally not set — see models.yml WARN task for
|
|
# the HuggingFace URL if a re-download is ever needed.
|
|
|
|
# --- Staged GGUF models (data-driven, idempotent staging) --------------------
|
|
# Additional GGUFs to ensure are present in llm_models_dir, alongside the
|
|
# production Qwen3.6-35B. Consumed by tasks/models.yml (loop over
|
|
# tasks/stage_model.yml). Each entry:
|
|
# filename: target filename in llm_models_dir
|
|
# url: HuggingFace resolve URL (public repos; no auth needed)
|
|
# size_bytes: EXACT expected byte size (HF manifest) — guard: download only
|
|
# if the file is missing OR its size != this value (idempotent;
|
|
# never re-pulls a correct file, never needlessly restarts).
|
|
# source_repo: upstream HF repo (audit/lineage)
|
|
# The REAL list is defined per-host in host_vars/astro-orbiter/vars.yml (NOT
|
|
# hardcoded here) so the role stays generic and reusable for future model adds.
|
|
# Empty default = nothing staged (safe no-op).
|
|
llm_staged_models: []
|
|
|
|
# --- Existing Gemma baseline (rollback target — never modified by this role) -
|
|
# Populated by Phase 0 discovery (tasks/discover.yml) if not already known.
|
|
# Set here only as a fallback name to search for; discovery is authoritative.
|
|
llm_existing_gemma_service_name_guess: llama-server
|
|
|
|
# --- Router mode shadow deployment (port 8003) --------------------------------
|
|
# Deploy llama-server in router/supervisor mode (no -m flag) on a shadow port.
|
|
# Production unit (llama-server-qwen, port 8002) is UNCHANGED until validation
|
|
# gates pass and Ryan explicitly approves cutover.
|
|
#
|
|
# Default: llm_router_enabled: false — all router tasks are no-ops until you
|
|
# flip this to true (either in host_vars, extra-vars, or the shadow playbook).
|
|
#
|
|
# CRITICAL: llm_router_models_max default is 1 here for safety. It is
|
|
# overridden to 4 in host_vars/astro-orbiter/vars.yml (t_33acbb2e) with
|
|
# a full VRAM budget note. DO NOT raise it without a VRAM budget review.
|
|
# Default llama-server cap is 4 simultaneous — that would OOM a 24GB card
|
|
# immediately when Qwen3.6-35B (20GB) is the resident model.
|
|
#
|
|
# Added 2026-08-12 (t_0cca74a2): router mode migration — War Machine.
|
|
llm_router_enabled: false
|
|
llm_router_port: 8003
|
|
llm_router_service_name: llama-server-router
|
|
llm_router_models_dir: "{{ llm_models_dir }}" # /opt/models — same dir as production
|
|
llm_router_models_max: 1 # CRITICAL: RTX 3090 24GB, single model only
|
|
llm_router_ctx_size: 65536 # 64K — must match production (Hermes floor)
|
|
llm_router_parallel: 1
|
|
llm_router_gpu_layers: 99
|
|
llm_router_batch_size: 2048
|
|
llm_router_ubatch_size: 512
|
|
llm_router_cache_type_k: q4_0 # required to fit 64K KV in 24GB
|
|
llm_router_cache_type_v: q4_0
|
|
llm_router_flash_attn: "auto"
|
|
llm_router_bind_address: "{{ llm_bind_address }}" # 10.1.71.130
|
|
llm_router_allowed_source_cidr: "{{ llm_allowed_source_cidr }}" # 10.1.70.0/24
|
|
llm_router_expected_model_id: "Qwen3.6-35B-A3B-UD-Q4_K_S" # verified at Gate 1
|
|
llm_router_vram_max_mib: 23000 # Gate 3: fail if exceeded under load
|
|
|
|
# --- Router preset mode (--models-preset INI) ---------------------------------
|
|
# Set llm_router_preset_enabled: true to switch from --models-dir to
|
|
# --models-preset. Preset mode is REQUIRED to support model aliases.
|
|
# The template at llama-server-router-preset.ini.j2 defines all 3 router models:
|
|
# - Qwen3.6-35B-A3B-UD-Q4_K_S (no alias — primary ID unchanged)
|
|
# - Phi-3.5-mini-instruct-Q8_0 (alias: Phi-3.5-mini-instruct-8bit) <-- t_9adf0889
|
|
# - Meta-Llama-3.1-8B-Instruct-Q4_K_M (no alias — primary ID unchanged)
|
|
#
|
|
# llm_router_preset_path: on-disk path where the rendered INI is deployed.
|
|
# Default: /opt/llama-server-router-preset.ini (owned by root, readable by all).
|
|
#
|
|
# GH #22364 note: --models-preset causes an extra "default" entry in /v1/models.
|
|
# This is cosmetic and does not affect model selection by name. Accept it.
|
|
#
|
|
# Added 2026-08-12 (t_9adf0889) — War Machine.
|
|
llm_router_preset_enabled: false # flip true to activate preset mode
|
|
|
|
# Per-model ctx-size / flash-attn overrides for preset mode (t_ryan_per_model_ctx).
|
|
# Defaults mirror the prior uniform 65536/auto behavior; host_vars or the
|
|
# deploy playbook override these to the values Ryan requested per workload.
|
|
llm_router_llama_ctx_size: "{{ llm_router_ctx_size }}"
|
|
llm_router_llama_flash_attn: "{{ llm_router_flash_attn }}"
|
|
llm_router_phi_ctx_size: "{{ llm_router_ctx_size }}"
|
|
llm_router_phi_flash_attn: "{{ llm_router_flash_attn }}"
|
|
# Qwen2.5-Coder-14B: ctx_size=16384, flash_attn=true per task t_55c164f5
|
|
llm_router_coder_ctx_size: 16384
|
|
llm_router_coder_flash_attn: "true"
|
|
llm_router_preset_path: /opt/llama-server-router-preset.ini
|
|
# nomic-embed-text-v1.5: embedding model, ctx-size=8192 per task t_34b96e83
|
|
# No flash_attn or KV cache params — embedding models use bidirectional forward pass,
|
|
# not autoregressive KV cache. load-on-startup=true / sleep-idle-seconds=-1 keep it
|
|
# always warm at negligible VRAM cost (~84MB).
|
|
llm_router_nomic_ctx_size: 8192
|