All skills
huggingface avatar

/huggingface-local-models

@f4ddab4 official
by Hugging Facehuggingface/skills11k stars
753

Use to select models to run locally with llama.cpp and GGUF on CPU, Mac Metal, CUDA, or ROCm. Covers finding GGUFs, quant selection, running servers, exact GGUF file lookup, conversion, and OpenAI-compatible local serving.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-local-models

This session only. Nothing lands on disk.

referenceshardware.md

≈156 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Hardware Acceleration

Apple Silicon (Metal)

make clean && make GGML_METAL=1
llama-cli -m model.gguf -ngl 99 -p "Hello"

NVIDIA (CUDA)

make clean && make GGML_CUDA=1
llama-cli -m model.gguf -ngl 35 -p "Hello"

# Hybrid for large models
llama-cli -m llama-70b.Q4_K_M.gguf -ngl 20

# Multi-GPU split
llama-cli -m large-model.gguf --tensor-split 0.5,0.5 -ngl 60

AMD (ROCm)

make LLAMA_HIP=1
llama-cli -m model.gguf -ngl 999

CPU

# Match physical cores, not logical threads
llama-cli -m model.gguf -t 8 -p "Hello"

# BLAS acceleration
make LLAMA_OPENBLAS=1

Source: SKILL.md on GitHub

1 warning16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides comprehensive instructions for selecting, downloading, and running AI models locally using llama.cpp and Hugging Face. It utilizes standard community tools and trusted platforms for model distribution. While the workflow involves external downloads and command execution, these are essential components of the skill's intended functionality.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

Signed by skilld at f4ddab4. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 5 months ago
  • llama.cpp
  • gguf
  • huggingface
  • quantization
  • local-inference
  • cpu
  • cuda
  • metal
  • rocm
  • model-serving

README badge

README badge for huggingface/skills/huggingface-local-models

Finds GGUF-quantized models on Hugging Face Hub compatible with llama.cpp, selects appropriate quantization levels, and launches them locally via llama-cli or llama-server with CPU, Metal, CUDA, or ROCm acceleration. Covers model discovery, quant selection, exact file lookup, format conversion, and OpenAI-compatible local serving.

Generated from the current SKILL.md.

Does this skill work with GPU acceleration?
Yes. The skill covers llama.cpp builds for Metal (Mac), CUDA (Nvidia), and ROCm (AMD), plus CPU inference. Hardware-specific setup is detailed in the bundled hardware.md reference.
What if a model repo doesn't have GGUF files pre-quantized?
The skill includes a conversion workflow: download Transformers weights with `hf download`, convert to GGUF using `convert_hf_to_gguf.py`, then quantize with `llama-quantize`.
How do I choose the right quantization level?
Start with the quant the Hugging Face local-app page recommends, default to Q4_K_M otherwise, and prefer Q5_K_M or Q6_K for code workloads if memory allows. The bundled quantization.md reference has format tables and tradeoff details.
Can I run the model as an OpenAI-compatible server?
Yes. Use `llama-server` instead of `llama-cli` to launch an OpenAI-compatible API on localhost:8080 that accepts `/v1/chat/completions` requests.
Does this work with gated models?
Yes, if you authenticate first with `hf auth login`. The skill assumes you have Hugging Face Hub access for the repo you want to run.

Generated from the current SKILL.md. These answers refresh after source changes.