; ------------------------------------------------------------------------------ ; FILE: roles/llm-inference-multimodel/templates/llama-server-router-preset.ini.j2 ; DESCRIPTION: llama.cpp --models-preset INI for llama-server-router. ; ; Purpose: define all router-served GGUFs as named model entries so that: ; - Each model is explicitly named and configured (no auto-discovery surprises) ; - Aliases can be added per model (impossible with --models-dir alone) ; - The Phi-3.5-mini-instruct-Q8_0 entry carries the alias ; "Phi-3.5-mini-instruct-8bit" — Ryan's Hermes auxiliary.title_generation ; already references this friendlier name; both names resolve to the same ; GGUF child process. ; ; INI format notes (llama.cpp preset.md): ; - Section header (e.g. [Phi-3.5-mini-instruct-Q8_0]) is the primary model ID ; that appears in /v1/models and that clients send in the "model" field. ; - `alias` adds an ADDITIONAL name — both the section name and the alias work. ; - The `model` key is the absolute path to the GGUF file. ; - All other keys map directly to llama-server CLI flags (underscores or hyphens). ; ; Known upstream issues (Aug 2026): ; - GH #22364: --models-preset creates an extra "default" model entry in ; /v1/models. This is cosmetic — it has no functional effect on model ; selection by name. Document and move on. ; - GH #23460: can't pass per-model samplers via --models-preset in router ; mode. Non-issue: Hermes always sends sampling params in the request body. ; ; Added 2026-08-12 (t_9adf0889): Phi alias — War Machine. ; All per-model settings carry over unchanged from the --models-dir baseline ; (ctx_size=65536, n_gpu_layers=99, cache=q4_0 for both K and V, models-max=4). ; ; UPDATED (t_ryan_per_model_ctx, per Ryan/JARVIS request): Llama-3.1-8B and ; Phi-3.5-mini now get PER-MODEL ctx-size/flash-attn matched to actual ; workload instead of the uniform 65536 used by every model previously: ; - Llama-3.1-8B-Instruct-Q4_K_M: ctx-size 8192 (tool-routing/micro-tasks) ; - Phi-3.5-mini-instruct-Q8_0: ctx-size 32768 (long web scrapes/logs) ; Both now request explicit flash-attn=true (was "auto"). Qwen3.6-35B is ; INTENTIONALLY left untouched at ctx-size 65536 / flash-attn auto — not part ; of this change. Existing aliases (Meta-Llama-3.1-8B-Instruct-4bit, ; Phi-3.5-mini-instruct-8bit) are PRESERVED unchanged to avoid breaking live ; Hermes custom_providers routing — see role README / deployment report for ; the alias-naming ambiguity flag (Ryan's pasted TOML used different alias ; strings: "llama-3.1-8b" / "phi-3.5-mini"). ; ------------------------------------------------------------------------------ ; --- Production model: Qwen3.6-35B-A3B-UD-Q4_K_S ---------------------------- ; Primary model ID: Qwen3.6-35B-A3B-UD-Q4_K_S (unchanged from --models-dir) ; ~20GB, primary Hermes production LLM. Context: 64K with q4_0 KV cache. [Qwen3.6-35B-A3B-UD-Q4_K_S] model = {{ llm_models_dir }}/Qwen3.6-35B-A3B-UD-Q4_K_S.gguf n-gpu-layers = {{ llm_router_gpu_layers }} ctx-size = {{ llm_router_ctx_size }} cache-type-k = {{ llm_router_cache_type_k }} cache-type-v = {{ llm_router_cache_type_v }} batch-size = {{ llm_router_batch_size }} ubatch-size = {{ llm_router_ubatch_size }} parallel = {{ llm_router_parallel }} ; --- Auxiliary model: Phi-3.5-mini-instruct-Q8_0 ---------------------------- ; Primary model ID: Phi-3.5-mini-instruct-Q8_0 (unchanged from --models-dir) ; Alias: Phi-3.5-mini-instruct-8bit (NEW — Ryan's config target) ; Both names resolve to this GGUF child process. ; ~3.8GB, auxiliary.title_generation consumer in Ryan's Hermes config. ; The alias is deployed — "model not found" errors are resolved. ; ; KNOWN ISSUE (2026-08-12, t_9adf0889): ; json_schema response_format fails for Phi-3.5-mini in llama.cpp router ; mode due to chat template grammar sampler incompatibility (GH #23460). ; The grammar sampler generates root ::= "assistant|>\n" ... which fails ; to initialize. This means Hermes auxiliary.title_generation still errors ; with "HTTP 400: Failed to initialize samplers" when json_schema format ; is requested. Without response_format (plain text), Phi works fine. ; ; Resolution options: ; a) Update auxiliary.title_generation.model in ~/.hermes/config.yaml to ; Meta-Llama-3.1-8B-Instruct-Q4_K_M (which supports json_schema) — Ryan ; needs to approve this config.yaml write (protected file). ; b) Rebuild llama.cpp from a newer commit if this bug is fixed upstream. ; c) Accept title generation degradation for Phi-specific structured output. [Phi-3.5-mini-instruct-Q8_0] model = {{ llm_models_dir }}/Phi-3.5-mini-instruct-Q8_0.gguf alias = Phi-3.5-mini-instruct-8bit n-gpu-layers = {{ llm_router_gpu_layers }} ctx-size = {{ llm_router_phi_ctx_size }} flash-attn = {{ llm_router_phi_flash_attn }} cache-type-k = {{ llm_router_cache_type_k }} cache-type-v = {{ llm_router_cache_type_v }} batch-size = {{ llm_router_batch_size }} ubatch-size = {{ llm_router_ubatch_size }} parallel = {{ llm_router_parallel }} ; --- Auxiliary model: Meta-Llama-3.1-8B-Instruct-Q4_K_M -------------------- ; Primary model ID: Meta-Llama-3.1-8B-Instruct-Q4_K_M (unchanged from --models-dir) ; Alias: Meta-Llama-3.1-8B-Instruct-4bit (NEW — friendlier name) ; Both names resolve to this GGUF child process. ; ~4.6GB, general-purpose small model. Works with json_schema structured output. [Meta-Llama-3.1-8B-Instruct-Q4_K_M] model = {{ llm_models_dir }}/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf alias = Meta-Llama-3.1-8B-Instruct-4bit n-gpu-layers = {{ llm_router_gpu_layers }} ctx-size = {{ llm_router_llama_ctx_size }} flash-attn = {{ llm_router_llama_flash_attn }} cache-type-k = {{ llm_router_cache_type_k }} cache-type-v = {{ llm_router_cache_type_v }} batch-size = {{ llm_router_batch_size }} ubatch-size = {{ llm_router_ubatch_size }} parallel = {{ llm_router_parallel }} ; --- Coder model: Qwen2.5-Coder-14B-Instruct-Q4_K_M ------------------------- ; Primary model ID: Qwen2.5-Coder-14B-Instruct-Q4_K_M (filename-derived) ; Alias: Qwen2.5-Coder-14B-Instruct-4bit (friendlier name) ; Both names resolve to this GGUF child process. ; ~8.4GB weights + ~0.6GB KV @ 16K ctx = ~9.0GB VRAM. ; ctx-size=16384, flash-attn=true per task t_55c164f5 / Ryan's request. ; Source: bartowski/Qwen2.5-Coder-14B-Instruct-GGUF (public, no auth) ; Added 2026-08-13 (t_55c164f5) — War Machine. [Qwen2.5-Coder-14B-Instruct-Q4_K_M] model = {{ llm_models_dir }}/Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf alias = Qwen2.5-Coder-14B-Instruct-4bit n-gpu-layers = {{ llm_router_gpu_layers }} ctx-size = {{ llm_router_coder_ctx_size }} flash-attn = {{ llm_router_coder_flash_attn }} cache-type-k = {{ llm_router_cache_type_k }} cache-type-v = {{ llm_router_cache_type_v }} batch-size = {{ llm_router_batch_size }} ubatch-size = {{ llm_router_ubatch_size }} parallel = {{ llm_router_parallel }}