fix(llm-inference): bitsandbytes int4 OOM — pending switch to llama.cpp+GGUF
bitsandbytes quantizes on-the-fly: loads full bf16 weights (~54GB RAM peak) before compressing to int4. Kills the 40GB OptiPlex on torch.compile warmup. Fix in next commit: switch serve phase to llama.cpp + GGUF Q4_K_M. Pre-quantized weights load directly — peak RAM ~16GB, no compile overhead.
This commit is contained in:
@@ -22,9 +22,11 @@
|
||||
become: true
|
||||
become_user: "{{ llm_venv_owner }}"
|
||||
|
||||
- name: Install vLLM
|
||||
- name: Install vLLM and bitsandbytes
|
||||
ansible.builtin.pip:
|
||||
name: vllm
|
||||
name:
|
||||
- vllm
|
||||
- bitsandbytes
|
||||
state: present
|
||||
virtualenv: "{{ llm_venv_path }}"
|
||||
become: true
|
||||
|
||||
Reference in New Issue
Block a user