All skills
av avatar

/run-llms

@e0627a9
by Ivan Charapanauav/skills16 stars
2

Set up and run local LLMs with Harbor. Use when the user wants to run models locally, install Harbor, start Open WebUI, llama.cpp, Ollama, vLLM, Docker Model Runner, MLX, or oMLX, pull GGUF or HuggingFace models, add SearXNG web search, Speaches TTS/STT, or Open Terminal code execution, launch Codex/Claude/Grok/OpenCode against a Harbor backend, or troubleshoot GPU, VRAM, and service startup.

Use this Skill: https://skilld.dev/gh/av/skills/run-llms

This session only. Nothing lands on disk.

referencestroubleshooting.md

≈778 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Troubleshooting

Read logs with docker logs harbor.<service> (optionally --tail 200). Do not run harbor logs from an agent shell; it follows forever.

Services will not start

harbor ps
docker logs harbor.<service>
docker ps -a | grep harbor
harbor doctor
docker info
harbor fixfs                    # Linux volume ACLs
harbor down && harbor up

harbor doctor checks Docker, Compose 2.23.1+, disk, registry, Harbor files, NVIDIA/ROCm when present, and WSL version.

harbor up pre-checks host port conflicts (webui 33801, llamacpp 33831, ollama 33821, ...). Inspect with ss -tlnp / lsof -i. Skip the check only with harbor up --skip-port-check.

WebUI first boot can sit in starting for several minutes while embedding/Whisper weights download. Wait for healthy before treating it as a crash.

No GPU / CUDA / ROCm

nvidia-smi
docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi
cat /etc/docker/daemon.json     # nvidia runtime
sudo systemctl restart docker
harbor down && harbor up

AMD: /dev/kfd and /dev/dri present, ROCm capability on (harbor config get capabilities.default). Apple Silicon: do not debug missing CUDA; use dmr, mlx, or omlx.

Model will not load / OOM

nvidia-smi                      # or equivalent
  • Ollama: smaller quant, harbor pull model:q4_k_m, or harbor ollama ctx 4096
  • llama.cpp: harbor llamacpp args '-c 2048 --n-gpu-layers 20' then harbor restart llamacpp
  • vLLM: --max-model-len 4096, bitsandbytes quant, --cpu-offload-gb 4, --enforce-eager
  • Router shows no GGUFs: pull one first (references/models.md), confirm specifier is empty
harbor restart <backend>

Cannot open the UI

harbor ps
harbor url webui
ss -tlnp | grep 33801
harbor open

Direct URL is http://localhost:33801. Admin account issues: incognito / clear site data, then docker logs harbor.webui.

Model missing in the UI

harbor models ls
harbor ollama list              # if using Ollama
# Browser refresh
# Open WebUI → Settings → Connections
docker logs harbor.webui
harbor restart webui

llama.cpp models appear after a GGUF pull while the service is in router mode. vLLM serves the configured harbor vllm model, not every file in the HF cache.

Web search missing

harbor ps | grep searxng
docker logs harbor.searxng
harbor restart webui
harbor url searxng

Slow / hanging inference

harbor top                      # nvtop
docker logs harbor.<backend>    # first load is slow
harbor ollama ps                # several models resident

vLLM compiles CUDA graphs on first start; wait for Application startup complete, or --enforce-eager. Check logs for CPU fallback.

llama.cpp pull container dies on a custom GPU image

Ephemeral GGUF pulls run without GPU devices. Switch the capability image to the official CPU server, pull, restore (see references/models.md).

Open Terminal sandbox reset

harbor down openterminal
rm -rf "$(harbor home)/services/openterminal/data"
harbor up openterminal

Source: SKILL.md on GitHub

1 alert1d4 checks · Risk HIGH
  • Gen Agent Trust Hub1d

    This skill provides a playbook for managing local LLM stacks using Harbor. It includes instructions to download and execute setup scripts and agent instructions directly from the vendor's GitHub repository, uses administrative commands for Docker configuration, and manages various local services via shell commands.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: MEDIUM · 1 issue

  • Runlayer7mo

    2/2 files flagged

Signed by skilld at e0627a9. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 4 weeks ago

README badge

README badge for av/skills/run-llms