Topics
- llama.cpp
- gguf
- huggingface
- quantization
- local-inference
- cpu
- cuda
- metal
- rocm
- model-serving
What it does
Finds GGUF-quantized models on Hugging Face Hub compatible with llama.cpp, selects appropriate quantization levels, and launches them locally via llama-cli or llama-server with CPU, Metal, CUDA, or ROCm acceleration. Covers model discovery, quant selection, exact file lookup, format conversion, and OpenAI-compatible local serving.
Generated from the current SKILL.md.
Frequently asked
Does this skill work with GPU acceleration?
Yes. The skill covers llama.cpp builds for Metal (Mac), CUDA (Nvidia), and ROCm (AMD), plus CPU inference. Hardware-specific setup is detailed in the bundled hardware.md reference.
What if a model repo doesn't have GGUF files pre-quantized?
The skill includes a conversion workflow: download Transformers weights with `hf download`, convert to GGUF using `convert_hf_to_gguf.py`, then quantize with `llama-quantize`.
How do I choose the right quantization level?
Start with the quant the Hugging Face local-app page recommends, default to Q4_K_M otherwise, and prefer Q5_K_M or Q6_K for code workloads if memory allows. The bundled quantization.md reference has format tables and tradeoff details.
Can I run the model as an OpenAI-compatible server?
Yes. Use `llama-server` instead of `llama-cli` to launch an OpenAI-compatible API on localhost:8080 that accepts `/v1/chat/completions` requests.
Does this work with gated models?
Yes, if you authenticate first with `hf auth login`. The skill assumes you have Hugging Face Hub access for the repo you want to run.
Generated from the current SKILL.md. These answers refresh after source changes.