Hermes Agent service account a3c92f70bf feat(deploy-vllm): swap DeepSeek-R1-Distill-Qwen-32B for Gemma 4 26B A4B AWQ (t_gemma4_swap)
Retired DeepSeek-R1-Distill-Qwen-32B after confirming its auto tool-choice
reliability is a known, documented DeepSeek-R1-distillation limitation
(trained on pure reasoning traces, no function-calling data — upstream
GitHub-confirmed, not a config gap). Model choice moved to Gemma 4 26B A4B
(Google, Apache 2.0, US-origin, matches Ryan's model-origin preference):

- cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit — MoE (25.2B total / 3.8B active),
  chosen over the dense 31B variant for smaller on-disk footprint
  (~17.2GB vs ~20.9GB), buying more KV-cache headroom on this 24GB card
- max_model_len=65536 (comfortably over Hermes's 64K floor; native
  context is 256K, no extension trick needed)
- Native gemma4 tool-call-parser + gemma4 reasoning-parser (both
  registered in this host's vLLM 0.28.0) — purpose-built for this
  model's actual output format, not a same-family approximation
- kv_cache_dtype: int4_per_token_head from the outset (learned from the
  DeepSeek swap's fp16->fp8->int4 trial-and-error escalation)

Bug found and fixed during deployment: the repo's config.json declares
quant_method 'compressed-tensors' (llm-compressor output) despite the
repo name saying 'AWQ-4bit'. Passing --quantization awq explicitly
caused a hard pydantic ValidationError on every startup attempt.
Fix: omit the quantization field entirely and let vLLM auto-detect from
the model's own config.json — confirmed clean single-attempt start,
NRestarts=0, once removed.

Verified live:
- /health 200, /v1/models confirms max_model_len=65536
- Live completion: correct answer, no unwanted reasoning trace by default
- tool_choice=auto with a clear trigger prompt: correct tool_calls
  response with valid JSON args — the exact test DeepSeek-R1-Distill
  failed (it either answered in plain text or burned tokens reasoning
  about how to call the tool instead of calling it)
- tool_choice=auto with an irrelevant tool present: correctly answered
  in plain text, did not over-trigger the tool
- Ansible idempotent re-run confirmed: changed=0, NRestarts=0, clean
  journalctl (zero error/traceback lines) after a fresh restart

Known follow-up (not done here): Hindsight's HINDSIGHT_API_LLM_MODEL
cluster config still references the retired DeepSeek-R1-Distill-Qwen-32B
(itself a follow-up from the prior Qwen2.5-32B swap) — needs another
GitOps update to point at Gemma-4-26B-A4B-it-AWQ.
2026-08-31 21:28:48 -05:00
2026-05-26 11:48:23 -05:00
2026-03-06 23:10:56 -06:00
2026-05-25 20:21:24 -05:00
2026-05-25 20:21:24 -05:00
2026-06-29 20:55:25 -05:00
2025-03-16 19:28:01 -05:00
2026-06-29 20:49:57 -05:00

mk-labs

Automated infrastructure provisioning and configuration for a personal homelab, built on GitOps practices with clear tool responsibility boundaries.

Architecture

A single operator action — setting a VM record's status to Staged in NetBox — triggers a fully automated provisioning pipeline:

NetBox (webhook) → n8n (validate & orchestrate) → Terraform (create VM + DHCP)
                                                 → Ansible (OS config + DNS + status update)
Tool Host IP Responsibility
NetBox fire-station 10.1.71.102 Source of truth — VM records, IP allocation, VLAN data
n8n tiki-room 10.1.71.23 Event orchestration, validation, pipeline sequencing
Terraform city-hall 10.1.71.35 Proxmox VM lifecycle, Unifi DHCP reservations
Ansible / Semaphore imagineering 10.1.71.22 OS configuration, DNS records, NetBox status updates
Proxmox fantasyland 10.1.71.13 Target hypervisor

All systems on the Server Trusted VLAN (10.1.71.0/24).

Repository Structure

homelab/
├── ansible/
│   ├── inventory/          # NetBox dynamic inventory + static
│   ├── playbooks/          # Runnable playbooks (vm-provision, DNS, OS updates)
│   ├── roles/              # vm-baseline, dns-manager, common, haproxy, n8n, observer, etc.
│   ├── tasks/              # Shared includable task files
│   ├── group_vars/         # Group variable definitions
│   ├── host_vars/          # Per-host variable definitions
│   ├── templates/          # Jinja2 templates
│   └── ansible.cfg
│
├── terraform/
│   ├── proxmox/vm/         # bpg/proxmox provider — VM creation from templates
│   ├── unifi/dhcp/         # Unifi provider — DHCP static reservations on UDM Pro
│   └── dns/                # DNS record management
│
├── packer/
│   ├── ubuntu-24.04/       # Ubuntu 24.04 VM template (small → xlarge-plus sizes)
│   └── fedora-42/          # Fedora 42 VM template
│
├── n8n/
│   └── workflows/          # Exported n8n workflow JSON (vm-provisioning)
│
├── netbox/
│   └── initializers/       # Custom fields, VLANs, IP prefixes as code
│
└── docs/
    └── decisions/          # Architecture decision records

Pipeline Flow

# System Action
1 NetBox Operator sets VM status to Staged → webhook fires
2 n8n Validates payload (hostname, IP, VLAN, template, proxmox_node)
3 n8n → city-hall SSH + terraform apply — creates VM on Proxmox
4 n8n Queries Proxmox API for MAC address
5 n8n → NetBox Writes MAC to VM interface record
6 n8n → city-hall SSH + terraform apply — creates DHCP reservation on UDM Pro
7 n8n → imagineering Triggers Ansible via Semaphore API
8 Ansible OS baseline, SSH hardening, Technitium DNS A record
9 Ansible → NetBox Sets VM status to Active

On any failure, NetBox status is set to Failed. No auto-retry — operator investigates.

Hardware

  • 7× Dell 7050 SFF
  • 3× Minisforum TH60
  • 2× Minisforum MS01
  • Synology DS1621+
  • Ubiquiti UDM Pro

Software Stack

  • Virtualization: Proxmox
  • Automation: Terraform, Ansible, n8n, Semaphore
  • DNS: Technitium (authoritative), Unbound (recursive)
  • IPAM/DCIM: NetBox
  • Networking: Ubiquiti UDM Pro
  • Templates: Packer (Ubuntu 24.04, Fedora 42)

Getting Started

See docs/decisions/vm-provisioning-flow.md for the full architecture decision record.

Previous OpenShift/ACM/Fastpass content is preserved in the archive/pre-mk-labs branch.

Security

No sensitive data is stored in this repository. Secrets are managed via Ansible Vault and environment variables on pipeline hosts.


Status: 🚧 Active Development — VM Provisioning Pipeline

Last Updated: February 2026

Description
No description provided
Readme MIT 24 MiB
Languages
Jinja 50.7%
HCL 29.8%
Shell 18.8%
Dockerfile 0.7%