Agent skill

Llama Cpp

by Tommy-yw in Tommy-yw/RunbookHermes

llama.cpp local GGUF inference + HF Hub model discovery. An agent skill from Tommy-yw/RunbookHermes.

MITAuto-check passedAI & LLM Engineering

Install Llama Cpp

skills CLI
$ npx skills add Tommy-yw/RunbookHermes --skill llama-cpp -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Tommy-yw/RunbookHermes llama-cpp --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Tommy-yw/RunbookHermes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mlops/inference/llama-cpp .claude/skills/llama-cpp && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llama-cpp
GitHub stars
546
Used in
4 other repos
Token cost
~2.2k tokens
SKILL.md length
698 words
Files
7 (incl. references)
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

llama.cpp local GGUF inference + HF Hub model discovery. An agent skill from Tommy-yw/RunbookHermes.

  • Works in 7 steps: Search for candidate repos on the Hub → Open the repo with the llama.cpp… → Treat the local-app snippet as the… → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers When to use, Model Discovery workflow, Quick start and Python bindings…, plus 6 more sections
  • Calls cmake, pip and brew; reaches huggingface.co and github.com

What it does

Llama Cpp is an agent skill from Tommy-yw/RunbookHermes. llama.cpp local GGUF inference + HF Hub model discovery.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/advanced-usage.md`, `references/hub-discovery.md` and `references/optimization.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with llama.cpp, Hugging Face and Python. The repository describes itself as: Hermes-native AIOps agent for evidence-driven incident response, approval-gated remediation, and runbook learning. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/llama-cpp”

Requirements

  • Python 3
  • Docker

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Search for candidate repos on the Hub
  2. Open the repo with the llama.cpp local-app view
  3. Treat the local-app snippet as the source of truth when it is visible
  4. Read the same ?local-app=llama.cpp URL as page text or HTML and extract the section under Hardware compatibility
  5. Query the tree API to confirm what actually exists
  6. If the local-app snippet is not text-visible, reconstruct the command from the repo plus the chosen quant
  7. Only suggest conversion from Transformers weights if the repo does not already expose GGUF files.

What it can do on your machine

Read from SKILL.md and the folder at commit 7fd2b9a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cmake
    • pip
    • brew
    • winget
    • git
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Llama Cpp loads about 2.2k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 17 tokens; SKILL.md has 698 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~17
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Tommy-yw/RunbookHermes at commit 7fd2b9a, republished under its MIT licence (© Tommy-yw). 698 words, ~2,209 tokens.

Download SKILL.mdSave it as .claude/skills/llama-cpp/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
llama-cpp
description
llama.cpp local GGUF inference + HF Hub model discovery.
version
2.1.2
author
Orchestra Research
license
MIT
dependencies
llama-cpp-python>=0.2.0

llama.cpp + GGUF

Use this skill for local GGUF inference, quant selection, or Hugging Face repo discovery for llama.cpp.

When to use

  • Run local models on CPU, Apple Silicon, CUDA, ROCm, or Intel GPUs
  • Find the right GGUF for a specific Hugging Face repo
  • Build a llama-server or llama-cli command from the Hub
  • Search the Hub for models that already support llama.cpp
  • Enumerate available .gguf files and sizes for a repo
  • Decide between Q4/Q5/Q6/IQ variants for the user's RAM or VRAM

Model Discovery workflow

Prefer URL workflows before asking for hf, Python, or custom scripts.

  1. Search for candidate repos on the Hub:
    • Base: https://huggingface.co/models?apps=llama.cpp&sort=trending
    • Add search=<term> for a model family
    • Add num_parameters=min:0,max:24B or similar when the user has size constraints
  2. Open the repo with the llama.cpp local-app view:
    • https://huggingface.co/<repo>?local-app=llama.cpp
  3. Treat the local-app snippet as the source of truth when it is visible:
    • copy the exact llama-server or llama-cli command
    • report the recommended quant exactly as HF shows it
  4. Read the same ?local-app=llama.cpp URL as page text or HTML and extract the section under Hardware compatibility:
    • prefer its exact quant labels and sizes over generic tables
    • keep repo-specific labels such as UD-Q4_K_M or IQ4_NL_XL
    • if that section is not visible in the fetched page source, say so and fall back to the tree API plus generic quant guidance
  5. Query the tree API to confirm what actually exists:
    • https://huggingface.co/api/models/<repo>/tree/main?recursive=true
    • keep entries where type is file and path ends with .gguf
    • use path and size as the source of truth for filenames and byte sizes
    • separate quantized checkpoints from mmproj-*.gguf projector files and BF16/ shard files
    • use https://huggingface.co/<repo>/tree/main only as a human fallback
  6. If the local-app snippet is not text-visible, reconstruct the command from the repo plus the chosen quant:
    • shorthand quant selection: llama-server -hf <repo>:<QUANT>
    • exact-file fallback: llama-server --hf-repo <repo> --hf-file <filename.gguf>
  7. Only suggest conversion from Transformers weights if the repo does not already expose GGUF files.

Quick start

Install llama.cpp
bash
# macOS / Linux (simplest)
brew install llama.cpp
bash
winget install llama.cpp
bash
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release
Run directly from the Hugging Face Hub
bash
llama-cli -hf bartowski/Llama-3.2-3B-Instruct-GGUF:Q8_0
bash
llama-server -hf bartowski/Llama-3.2-3B-Instruct-GGUF:Q8_0
Run an exact GGUF file from the Hub

Use this when the tree API shows custom file naming or the exact HF snippet is missing.

bash
llama-server \
    --hf-repo microsoft/Phi-3-mini-4k-instruct-gguf \
    --hf-file Phi-3-mini-4k-instruct-q4.gguf \
    -c 4096
OpenAI-compatible server check
bash
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [
      {"role": "user", "content": "Write a limerick about Python exceptions"}
    ]
  }'

Python bindings (llama-cpp-python)

pip install llama-cpp-python (CUDA: CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python --force-reinstall --no-cache-dir; Metal: CMAKE_ARGS="-DGGML_METAL=on" ...).

Basic generation
python
from llama_cpp import Llama

llm = Llama(
    model_path="./model-q4_k_m.gguf",
    n_ctx=4096,
    n_gpu_layers=35,     # 0 for CPU, 99 to offload everything
    n_threads=8,
)

out = llm("What is machine learning?", max_tokens=256, temperature=0.7)
print(out["choices"][0]["text"])
Chat + streaming
python
llm = Llama(
    model_path="./model-q4_k_m.gguf",
    n_ctx=4096,
    n_gpu_layers=35,
    chat_format="llama-3",   # or "chatml", "mistral", etc.
)

resp = llm.create_chat_completion(
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is Python?"},
    ],
    max_tokens=256,
)
print(resp["choices"][0]["message"]["content"])

# Streaming
for chunk in llm("Explain quantum computing:", max_tokens=256, stream=True):
    print(chunk["choices"][0]["text"], end="", flush=True)
Embeddings
python
llm = Llama(model_path="./model-q4_k_m.gguf", embedding=True, n_gpu_layers=35)
vec = llm.embed("This is a test sentence.")
print(f"Embedding dimension: {len(vec)}")

You can also load a GGUF straight from the Hub:

python
llm = Llama.from_pretrained(
    repo_id="bartowski/Llama-3.2-3B-Instruct-GGUF",
    filename="*Q4_K_M.gguf",
    n_gpu_layers=35,
)
Show full SKILL.md (308 more words)Show less

Choosing a quant

Use the Hub page first, generic heuristics second.

  • Prefer the exact quant that HF marks as compatible for the user's hardware profile.
  • For general chat, start with Q4_K_M.
  • For code or technical work, prefer Q5_K_M or Q6_K if memory allows.
  • For very tight RAM budgets, consider Q3_K_M, IQ variants, or Q2 variants only if the user explicitly prioritizes fit over quality.
  • For multimodal repos, mention mmproj-*.gguf separately. The projector is not the main model file.
  • Do not normalize repo-native labels. If the page says UD-Q4_K_M, report UD-Q4_K_M.

Extracting available GGUFs from a repo

When the user asks what GGUFs exist, return:

  • filename
  • file size
  • quant label
  • whether it is a main model or an auxiliary projector

Ignore unless requested:

  • README
  • BF16 shard files
  • imatrix blobs or calibration artifacts

Use the tree API for this step:

  • https://huggingface.co/api/models/<repo>/tree/main?recursive=true

For a repo like unsloth/Qwen3.6-35B-A3B-GGUF, the local-app page can show quant chips such as UD-Q4_K_M, UD-Q5_K_M, UD-Q6_K, and Q8_0, while the tree API exposes exact file paths such as Qwen3.6-35B-A3B-UD-Q4_K_M.gguf and Qwen3.6-35B-A3B-Q8_0.gguf with byte sizes. Use the tree API to turn a quant label into an exact filename.

Search patterns

Use these URL shapes directly:

text
https://huggingface.co/models?apps=llama.cpp&sort=trending
https://huggingface.co/models?search=<term>&apps=llama.cpp&sort=trending
https://huggingface.co/models?search=<term>&apps=llama.cpp&num_parameters=min:0,max:24B&sort=trending
https://huggingface.co/<repo>?local-app=llama.cpp
https://huggingface.co/api/models/<repo>/tree/main?recursive=true
https://huggingface.co/<repo>/tree/main

Output format

When answering discovery requests, prefer a compact structured result like:

text
Repo: <repo>
Recommended quant from HF: <label> (<size>)
llama-server: <command>
Other GGUFs:
- <filename> - <size>
- <filename> - <size>
Source URLs:
- <local-app URL>
- <tree API URL>

References

  • hub-discovery.md - URL-only Hugging Face workflows, search patterns, GGUF extraction, and command reconstruction
  • advanced-usage.md — speculative decoding, batched inference, grammar-constrained generation, LoRA, multi-GPU, custom builds, benchmark scripts
  • quantization.md — quant quality tradeoffs, when to use Q4/Q5/Q6/IQ, model size scaling, imatrix
  • server.md — direct-from-Hub server launch, OpenAI API endpoints, Docker deployment, NGINX load balancing, monitoring
  • optimization.md — CPU threading, BLAS, GPU offload heuristics, batch tuning, benchmarks
  • troubleshooting.md — install/convert/quantize/inference/server issues, Apple Silicon, debugging

Resources

© Tommy-yw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in skills/mlops/inference/llama-cpp of Tommy-yw/RunbookHermes.

  • SKILL.md
  • references/advanced-usage.md
  • references/hub-discovery.md
  • references/optimization.md
  • references/quantization.md
  • references/server.md
  • references/troubleshooting.md

Open the folder on GitHubat commit 7fd2b9a

Used in 4 other repositories

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in Tommy-yw/RunbookHermes, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Llama Cpp next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Llama Cpp compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Llama Cpp this skillTommy-yw/RunbookHermes5464 repos~2.2kAutomated safety check: PassMIT
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Outlines Structured GenerationOrchestra-Research/AI-Research-SKILLs13k10 repos~4kAutomated safety check: PassMIT
Aqua Model Lifecycleoracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.0
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0

Similar skills

  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Outlines Structured Generation

    Orchestra-Research/AI-Research-SKILLs

    Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models.

    13k GitHub starsUsed in 10 repos~4k tokens
    AI & LLM EngineeringAuto-check passed
  • Aqua Model Lifecycle

    oracle/accelerated-data-science

    Official

    Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

    125 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check: notes
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from Tommy-yw/RunbookHermes

All 38 skills in this repo
  • Fastmcp

    Tommy-yw/RunbookHermes

    Build, test, inspect, install, and deploy MCP servers with FastMCP in Python.

    546 GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Drug Discovery

    Tommy-yw/RunbookHermes

    Pharmaceutical research assistant for drug discovery workflows.

    546 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Youtube Content

    Tommy-yw/RunbookHermes

    Fetch YouTube video transcripts and transform them into structured content (chapters, summaries, threads, blog posts).

    546 GitHub starsUsed in 1 repo~785 tokens
    Auto-check passed
  • Oss Forensics

    Tommy-yw/RunbookHermes

    Supply chain investigation, evidence recovery, and forensic analysis for GitHub repositories.

    546 GitHub starsUsed in 3 repos~5k tokens
    Auto-check passed
  • P5js

    Tommy-yw/RunbookHermes

    Production pipeline for interactive and generative visual art using p5.js.

    546 GitHub starsUsed in 1 repo~6.8k tokens
    Auto-check passed
  • Touchdesigner MCP

    Tommy-yw/RunbookHermes

    Control a running TouchDesigner instance via twozero MCP — create operators, set parameters, wire connections, execute Python, build real-time visuals.

    546 GitHub starsUsed in 2 repos~3.4k tokens
    Auto-check passed

Questions about Llama Cpp

What does Llama Cpp do?

llama.cpp local GGUF inference + HF Hub model discovery. An agent skill from Tommy-yw/RunbookHermes. Llama Cpp is an agent skill from Tommy-yw/RunbookHermes.cpp local GGUF inference + HF Hub model discovery.

When should I use Llama Cpp?

Llama Cpp fits situations like: tasks that involve LLM inference and serving.

How do I install Llama Cpp in Claude Code?

Run `npx skills add Tommy-yw/RunbookHermes --skill llama-cpp -a claude-code`. Or copy the skill folder (skills/mlops/inference/llama-cpp in Tommy-yw/RunbookHermes) into .claude/skills/llama-cpp in your project. Claude Code loads it when a task matches its description.

How do I install Llama Cpp in Codex?

Run `npx skills add Tommy-yw/RunbookHermes --skill llama-cpp -a codex`. Or copy the skill folder (skills/mlops/inference/llama-cpp in Tommy-yw/RunbookHermes) into .agents/skills/llama-cpp in your project. Codex loads it when a task matches its description.

Can I use Llama Cpp in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Tommy-yw/RunbookHermes --skill llama-cpp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llama-cpp, .gemini/skills/llama-cpp, .github/skills/llama-cpp and .opencode/skills/llama-cpp in your project.

What does Llama Cpp need to run?

Going by SKILL.md and its folder, Llama Cpp needs the command-line tools its instructions call (cmake, pip, brew, winget, git and curl). Our summary lists: Python 3; Docker.

Does Llama Cpp access the network?

SKILL.md names 2 domains. In commands or code: huggingface.co and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Llama Cpp safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Llama Cpp use?

Llama Cpp is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Llama Cpp use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.8k tokens, read only when the agent opens those files.

What are the alternatives to Llama Cpp?

Skills that share tags, products or a category with Llama Cpp: Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Outlines Structured Generation (Orchestra-Research/AI-Research-SKILLs, 13k stars), Aqua Model Lifecycle (oracle/accelerated-data-science, 125 stars) and Add Model (guoqingbao/xinfer, 333 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Llama Cpp?

Tommy-yw (a GitHub user) maintains it in Tommy-yw/RunbookHermes, which has 546 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on May 18, 2026.

Source: Tommy-yw/RunbookHermes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.