Official agent skill

Hugging Face Local Models

by huggingface in huggingface/skills

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Hugging Face Local Models

skills CLI
$ npx skills add huggingface/skills --skill huggingface-local-models -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huggingface/skills huggingface-local-models --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/huggingface-local-models .claude/skills/huggingface-local-models && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
huggingface-local-models
GitHub stars
11k
Used in
3 other repos
Token cost
~945 tokens
SKILL.md length
265 words
Files
4 (incl. references)
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

  • Works in 7 steps: Search the Hub with apps=llama.cpp. → Open… → Prefer the exact HF local-app snippet… → …
  • Picking a GGUF model that fits your laptop or GPU memory
  • SKILL.md covers Default Workflow, Quick Start, Quant Choice and Load References, plus 1 more section
  • Calls hf, brew and winget; reaches huggingface.co and github.com

What it does

This skill searches the Hugging Face Hub for repositories that ship GGUF files usable with llama.cpp, helps choose a quantization, and starts the model with llama-cli or llama-server on CPU, Mac Metal, CUDA or ROCm. Its default workflow searches with the llama.cpp app filter, opens the repo's local-app page to read the recommended snippet and quant, confirms exact .gguf filenames through the Hub API, and launches with the repo and quant name, falling back to explicit repo and file flags when a repo names its files unusually.

Quant guidance keeps repo-native labels, defaults to Q4_K_M and prefers Q5_K_M or Q6_K for code or technical work when memory allows. Conversion from Transformers weights is a last resort for repositories without GGUF files. The skill also covers installing llama.cpp with Homebrew, winget or a source build, logging in with hf auth for gated repos, and smoke-testing the OpenAI-compatible server on localhost port 8080 with curl. Reference notes cover hardware, Hub discovery and quantization.

When your agent uses it

  • Picking a GGUF model that fits your laptop or GPU memory
  • Launching a local OpenAI-compatible server with llama-server
  • Finding the exact GGUF file in a Hugging Face repository
  • Converting a Transformers model to GGUF when no quantized files exist

Example prompts

  • “Find a coding model on Hugging Face that runs under llama.cpp on my MacBook with Metal.”
  • “Start llama-server with a Q4_K_M quant of this repo and check that it answers on port 8080.”
  • “List the GGUF files in this Hugging Face repo and tell me which quant suits a CUDA card.”
  • “Convert this model, which has no GGUF files, so I can run it locally.”

Requirements

  • llama.cpp installed (Homebrew, winget or a source build)
  • Hugging Face login for gated repositories
  • Network access to the Hugging Face Hub

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Search the Hub with apps=llama.cpp.
  2. Open https://huggingface.co/?local-app=llama.cpp.
  3. Prefer the exact HF local-app snippet and quant recommendation when it is visible.
  4. Confirm exact .gguf filenames with https://huggingface.co/api/models//tree/main?recursive=true.
  5. Launch with llama-cli -hf : or llama-server -hf :.
  6. Fall back to --hf-repo plus --hf-file when the repo uses custom file naming.
  7. Convert from Transformers weights only if the repo does not already expose GGUF files.

What it can do on your machine

Read from SKILL.md and the folder at commit ca0325b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • hf
    • brew
    • winget
    • git
    • make
    • python
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hugging Face Local Models loads about 945 tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 265 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~945
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from huggingface/skills at commit ca0325b, republished under its Apache-2.0 licence (© huggingface). 265 words, ~945 tokens.

Download SKILL.mdSave it as .claude/skills/huggingface-local-models/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
huggingface-local-models
description
Use to select models to run locally with llama.cpp and GGUF on CPU, Mac Metal, CUDA, or ROCm. Covers finding GGUFs, quant selection, running servers, exact GGUF file lookup, conversion, and OpenAI-compatible local serving.

Hugging Face Local Models

Search the Hugging Face Hub for llama.cpp-compatible GGUF repos, choose the right quant, and launch the model with llama-cli or llama-server.

Default Workflow

  1. Search the Hub with apps=llama.cpp.
  2. Open https://huggingface.co/<repo>?local-app=llama.cpp.
  3. Prefer the exact HF local-app snippet and quant recommendation when it is visible.
  4. Confirm exact .gguf filenames with https://huggingface.co/api/models/<repo>/tree/main?recursive=true.
  5. Launch with llama-cli -hf <repo>:<QUANT> or llama-server -hf <repo>:<QUANT>.
  6. Fall back to --hf-repo plus --hf-file when the repo uses custom file naming.
  7. Convert from Transformers weights only if the repo does not already expose GGUF files.

Quick Start

Install llama.cpp
bash
brew install llama.cpp
winget install llama.cpp
bash
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
make
Authenticate for gated repos
bash
hf auth login
Search the Hub
text
https://huggingface.co/models?apps=llama.cpp&sort=trending
https://huggingface.co/models?search=Qwen3.6&apps=llama.cpp&sort=trending
https://huggingface.co/models?search=<term>&apps=llama.cpp&num_parameters=min:0,max:24B&sort=trending
Run directly from the Hub
bash
llama-cli -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M
llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M
Run an exact GGUF file
bash
llama-server \
    --hf-repo unsloth/Qwen3.6-35B-A3B-GGUF \
    --hf-file Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \
    -c 4096
Convert only when no GGUF is available
bash
hf download <repo-without-gguf> --local-dir ./model-src
python convert_hf_to_gguf.py ./model-src \
    --outfile model-f16.gguf \
    --outtype f16
llama-quantize model-f16.gguf model-q4_k_m.gguf Q4_K_M
Smoke test a local server
bash
llama-server -hf unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M
bash
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer no-key" \
  -d '{
    "messages": [
      {"role": "user", "content": "Write a limerick about exception handling"}
    ]
  }'

Quant Choice

  • Prefer the exact quant that HF marks as compatible on the ?local-app=llama.cpp page.
  • Keep repo-native labels such as UD-Q4_K_M instead of normalizing them.
  • Default to Q4_K_M unless the repo page or hardware profile suggests otherwise.
  • Prefer Q5_K_M or Q6_K for code or technical workloads when memory allows.
  • Consider Q3_K_M, Q4_K_S, or repo-specific IQ / UD-* variants for tighter RAM or VRAM budgets.
  • Treat mmproj-*.gguf files as projector weights, not the main checkpoint.

Load References

  • Read hub-discovery.md for URL-first workflows, model search, tree API extraction, and command reconstruction.
  • Read quantization.md for format tables, model scaling, quality tradeoffs, and imatrix.
  • Read hardware.md for Metal, CUDA, ROCm, or CPU build and acceleration details.

Resources

  • llama.cpp: https://github.com/ggml-org/llama.cpp
  • Hugging Face GGUF + llama.cpp docs: https://huggingface.co/docs/hub/gguf-llamacpp
  • Hugging Face Local Apps docs: https://huggingface.co/docs/hub/main/local-apps
  • Hugging Face Local Agents docs: https://huggingface.co/docs/hub/agents-local
  • GGUF converter Space: https://huggingface.co/spaces/ggml-org/gguf-my-repo

© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/huggingface-local-models of huggingface/skills.

  • SKILL.md
  • references/hardware.md
  • references/hub-discovery.md
  • references/quantization.md

Open the folder on GitHubat commit ca0325b

Used in 3 other repositories

We found 8 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Hugging Face Local Models next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hugging Face Local Models compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hugging Face Local Models this skillhuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Hf Quant And Layer Package JobsMesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Test Modelguoqingbao/xinfer333—~2.6kAutomated safety check: PassMIT

Similar skills

  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when running quantization of a BF16/FP16 GGUF repo and Skippy layer-package creation as one local or Hugging Face Jobs workflow, publishing both artifacts to Hugging Face.

    3.5k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Test Model

    guoqingbao/xinfer

    Test LLM models served by xinfer for correctness, output quality, and performance.

    333 GitHub stars~2.6k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check passed
  • Aqua Model Lifecycle

    oracle/accelerated-data-science

    Official

    Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

    125 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from huggingface/skills

All 25 skills in this repo
  • Official

    Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.

    11k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    Auto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    Auto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed

Questions about Hugging Face Local Models

What does Hugging Face Local Models do?

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server. cpp, helps choose a quantization, and starts the model with llama-cli or llama-server on CPU, Mac Metal, CUDA or ROCm.gguf filenames through the Hub API, and launches with the repo and quant name, falling back to explicit repo and file flags when a repo names its files unusually.

When should I use Hugging Face Local Models?

Hugging Face Local Models fits situations like: picking a GGUF model that fits your laptop or GPU memory; launching a local OpenAI-compatible server with llama-server; finding the exact GGUF file in a Hugging Face repository; converting a Transformers model to GGUF when no quantized files exist.

How do I install Hugging Face Local Models in Claude Code?

Run `npx skills add huggingface/skills --skill huggingface-local-models -a claude-code`. Or copy the skill folder (skills/huggingface-local-models in huggingface/skills) into .claude/skills/huggingface-local-models in your project. Claude Code loads it when a task matches its description.

How do I install Hugging Face Local Models in Codex?

Run `npx skills add huggingface/skills --skill huggingface-local-models -a codex`. Or copy the skill folder (skills/huggingface-local-models in huggingface/skills) into .agents/skills/huggingface-local-models in your project. Codex loads it when a task matches its description.

Can I use Hugging Face Local Models in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill huggingface-local-models -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/huggingface-local-models, .gemini/skills/huggingface-local-models, .github/skills/huggingface-local-models and .opencode/skills/huggingface-local-models in your project.

What does Hugging Face Local Models need to run?

Going by SKILL.md and its folder, Hugging Face Local Models needs the command-line tools its instructions call (hf, brew, winget, git, make and python). Our summary lists: llama.cpp installed (Homebrew, winget or a source build); Hugging Face login for gated repositories; Network access to the Hugging Face Hub.

Does Hugging Face Local Models access the network?

SKILL.md names 2 domains. In commands or code: huggingface.co and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Hugging Face Local Models safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hugging Face Local Models use?

Hugging Face Local Models is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hugging Face Local Models use?

About 945 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Hugging Face Local Models?

Skills that share tags, products or a category with Hugging Face Local Models: Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Hf Quant And Layer Package Jobs (Mesh-LLM/mesh-llm, 3.5k stars), Add Model (guoqingbao/xinfer, 333 stars) and Resolve (alexziskind1/model-shelf, 130 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hugging Face Local Models?

huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,148 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 1, 2026.

Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.