Agent skill

Llama Cpp

by magnus919 in magnus919/agent-skills

Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

MITAuto-check passedAI & LLM Engineering

Install Llama Cpp

skills CLI
$ npx skills add magnus919/agent-skills --skill llama-cpp -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install magnus919/agent-skills llama-cpp --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/llama-cpp .claude/skills/llama-cpp && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llama-cpp
GitHub stars
111
Token cost
~2.3k tokens
SKILL.md length
1,055 words
Files
12 (incl. references)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

  • Works in 5 steps: Select and verify the installation → Select and inspect the model → Prove local inference → …
  • Building llama.cpp
  • SKILL.md covers Operating contract, When not to use, Read-only preflight and Route the task, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Llama Cpp is an agent skill from magnus919/agent-skills. Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems. Use when installing or building llama.cpp, selecting or inspecting GGUF models, running llama-cli, serving an OpenAI-compatible API with llama-server, tuning memory and performance, or diagnosing backend, context, template, and API failures. Do not use for model training or fine-tuning, general inference-framework selection, llama-cpp-python or other bindings, LlamaIndex…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files (for example `EVIDENCE-LEDGER.md`, `README.md` and `evals/evals.json`). Compatibility notes: Requires a supported llama.cpp binary or a build environment. Model use requires a compatible GGUF file and sufficient disk and memory; accelerator paths…

It sits in AI & LLM Engineering, covering LLM inference and serving and Fine-tuning. It works with llama.cpp, LlamaIndex, Ollama and CUDA. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.

When your agent uses it

  • Building llama.cpp
  • Inspecting GGUF models
  • Running llama-cli
  • Serving an OpenAI-compatible API with llama-server

Example prompts

  • “/llama-cpp”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires a supported llama.cpp binary or a build environment. Model use requires a compatible GGUF file and sufficient disk and memory; accelerator paths require the matching driver and SDK.

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Select and verify the installation
  2. Select and inspect the model
  3. Prove local inference
  4. Prove serving
  5. Tune one dimension at a time

What it can do on your machine

Read from SKILL.md and the folder at commit 96fbe07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires a supported llama.cpp binary or a build environment. Model use requires a compatible GGUF file and sufficient disk and memory; accelerator paths require the matching driver and SDK.

    From compatibility in the SKILL.md frontmatter.

Context cost

Llama Cpp loads about 2.3k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 139 tokens; SKILL.md has 1,055 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~139
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from magnus919/agent-skills at commit 96fbe07, republished under its MIT licence (© magnus919). 1,055 words, ~2,288 tokens.

Download SKILL.mdSave it as .claude/skills/llama-cpp/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
llama-cpp
description
Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems. Use when installing or building llama.cpp, selecting or inspecting GGUF models, running llama-cli, serving an OpenAI-compatible API with llama-server, tuning memory and performance, or diagnosing backend, context, template, and API failures. Do not use for model training or fine-tuning, general inference-framework selection, llama-cpp-python or other bindings, LlamaIndex, Ollama, or LM Studio operation.
compatibility
Requires a supported llama.cpp binary or a build environment. Model use requires a compatible GGUF file and sufficient disk and memory; accelerator paths require the matching driver and SDK.
license
MIT
metadata.source
https://github.com/ggml-org/llama.cpp
metadata.source_index
references/source-index.md
metadata.research_checked
2026-07-25

llama.cpp Operations

Treat every launch recipe as a hypothesis about a specific build, model, host, and workload. Discover capabilities from the installed binary, inspect the model and startup logs, then measure the requested boundary.

Operating contract

  1. Record the exact llama.cpp version or commit, installation method, OS and architecture, CPU and RAM, accelerator and memory, driver/toolkit, available devices, model provenance and quantization, intended context, concurrency, and workload.
  2. Read the installed command's --help before using a flag from documentation. llama.cpp flags, defaults, binary names, and REST behavior change frequently.
  3. Confirm the target, scope, and rollback path before acting. Read-only discovery may proceed without confirmation.
  4. Verify the backend from --list-devices and model-load logs. A successful build or an accepted GPU flag does not prove acceleration is active.
  5. Start with a bounded CLI smoke test on loopback or local input. Establish a measured baseline before changing threads, batches, context, cache types, offload, or split mode.
  6. Call work complete only at the requested boundary: binary, model load, generated output, API response, benchmark comparison, or diagnosed failure with evidence.

When not to use

Use ml-engineering for model training, fine-tuning, broad quantization methodology, evaluation design, or choosing among llama.cpp, vLLM, TGI, and other engines. Use the relevant product skill for Ollama, LM Studio, or LlamaIndex. Use binding-specific documentation for llama-cpp-python, node-llama-cpp, or other language wrappers.

Read-only preflight

Run only commands that exist in the installed build:

sh
llama-cli --version
llama-cli --help
llama-cli --list-devices
llama-server --version
llama-server --help
llama-bench --help

Also inspect host memory and accelerator state with native OS/vendor tools. Record results in the operation record. If no binary exists, choose an installation path only after reading installation and backends.

Route the task

NeedRead first
Install, build, choose CPU/Metal/CUDA/HIP/Vulkan/SYCL, Docker, or prove backend useinstallation and backends
Acquire, convert, inspect, license-check, quantize, or fit a GGUF modelmodels, GGUF, and memory
Run llama-cli, expose llama-server, call compatible APIs, use templates, structured output, embeddings, reranking, or toolsinference and serving
Tune threads, batches, context, cache, offload, concurrency, or multi-GPU and compare resultsperformance and benchmarking
Diagnose load, backend, OOM, speed, context, template, output, or API failurestroubleshooting
Check the evidence, research date, upstream revision, or refresh rule behind a claimsource index

Safe workflow

1. Select and verify the installation

Prefer a supported package or release binary when its compiled backend matches the target. Build from a pinned revision when backend options, portability, or reproducibility require it. Use Docker when host isolation is useful and device passthrough is understood. After installation, capture version, help, device listing, and a model-load log before claiming success.

2. Select and inspect the model

Accept a user-specified local GGUF path or Hugging Face repository. Before downloading, record the repository, revision, file, size, model card, license, base-model lineage, and quantizer when available. Inspect GGUF metadata and model-load output for architecture, quantization, context, tokenizer, chat template, and sidecars. Plan capacity from actual file size plus KV cache, context, batch/concurrency, compute buffers, and backend overhead; parameter count alone is insufficient.

Do not call one quantization universally best. Start from workload quality and capacity constraints, avoid requantizing an already quantized model when a higher-precision source is available, and compare candidate quants with the same task-quality and performance workload.

3. Prove local inference

Use a short, fixed prompt and bounded token count. Record the exact command, seed or sampling settings, startup log, output, timings, and whether the expected backend loaded. If the model has a chat template, test the template path required by the intended workload rather than treating plain completion as chat proof.

4. Prove serving

Bind to 127.0.0.1 for the first launch. Wait for /health to report ready, query /v1/models, then make a representative request using a reported model identifier. A listening process or HTTP 200 from a shallow endpoint is not inference proof. External exposure requires an explicit decision about bind address, API keys, TLS or reverse proxy, firewall, CORS, rate limits, logging, and whether experimental built-in tools are disabled.

Show full SKILL.md (406 more words)Show less
5. Tune one dimension at a time

Preserve a baseline before changing context size, generation and batch threads, logical or physical batch size, GPU layers, KV cache type/offload, Flash Attention, parallel slots, or multi-GPU split. Use llama-bench for prompt-processing and token-generation comparisons, and an end-to-end client or server benchmark for TTFT and request latency. Record each comparison in the benchmark template.

Hard boundaries

  • Do not infer accelerator use from the command line alone; require device and load-log evidence.
  • Do not expose an unauthenticated server beyond loopback by accident. Authentication is not a substitute for network and TLS controls.
  • Do not print or tee unredacted unit definitions, process environments, environment-file contents, API keys, or other credential-bearing configuration. Prefer redacted metadata; never retain secrets merely as evidence. If exact rollback requires a secret-bearing backup, keep it temporarily outside the repository with mode 0600, minimum retention, and explicit cleanup. Do not publish raw or low-entropy secret hashes.
  • Do not enable llama-server built-in filesystem or shell tools in an untrusted environment.
  • Do not override a chat template until model metadata, the original model card, and rendered behavior have been inspected.
  • For finish_reason: "length" with populated reasoning_content and empty content, inspect the access-controlled raw response but report only sanitized field state, lengths, finish_reason, and usage. From the same baseline, run separate one-variable probes: a bounded output-budget increase and, when the exact template supports it, chat_template_kwargs.enable_thinking: false. In streaming, inspect the documented schema for choices[].delta.reasoning_content, choices[].delta.content, and terminal choices[].finish_reason; record each as absent, null, empty, or populated rather than assuming presence. Treat reasoning_effort: "none" and server-side reasoning flags as version-sensitive, source-verified alternatives, not portable defaults.
  • Do not compare benchmark numbers from different models, quants, commits, backends, contexts, batches, thermal states, or workloads as if only one variable changed.
  • Do not treat llama-bench tokens per second as TTFT; its measurements exclude tokenization and sampling.
  • Do not claim a larger configured context preserves quality unless the model and scaling behavior support it and the workload was evaluated.

Exit criteria

The task is complete when the requested boundary is evidenced: the expected binary and backend are observed; the selected model's provenance and fit are recorded; a bounded prompt returns usable output; a server reaches readiness and completes a representative API request; a tuning change beats or preserves the declared metrics under matched conditions; or a failure is reduced to a supported cause with a safe next action. List any stronger boundary that was not tested.

© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (references) in llama-cpp of magnus919/agent-skills.

  • SKILL.md
  • EVIDENCE-LEDGER.md
  • README.md
  • evals/evals.json
  • references/inference-and-serving.md
  • references/installation-and-backends.md
  • references/models-gguf-and-memory.md
  • references/performance-and-benchmarking.md
  • references/source-index.md
  • references/troubleshooting.md
  • templates/benchmark-comparison.md
  • templates/operation-record.md

Open the folder on GitHubat commit 96fbe07

Compare with similar skills

Llama Cpp next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Llama Cpp compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Llama Cpp this skillmagnus919/agent-skills111—~2.3kAutomated safety check: PassMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
Mesh APImr-tbot/mesh-api179—~1.8kAutomated safety check: PassGPL-3.0
Vllm Deploy Simplevllm-project/vllm-skills103—~1.6kAutomated safety check: PassApache-2.0
Model Discoveryaiskillstore/marketplace4301 repos~1.9kAutomated safety check: PassNone

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Mesh API

    mr-tbot/mesh-api

    Interact with a Meshtastic LoRa mesh network through MESH-API — list nodes, read messages, send texts, and check connection status.

    179 GitHub stars~1.8k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Simple

    vllm-project/vllm-skills

    Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Model Discovery

    aiskillstore/marketplace

    Fetch current model names from AI providers (Anthropic, OpenAI, Gemini, Ollama), classify them into tiers (fast/default/heavy), and detect new models.

    430 GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Local Models

    glebis/claude-skills

    Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama.

    388 GitHub stars~1.4k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed

More from magnus919/agent-skills

All 129 skills in this repo
  • Artifact Pyramids

    magnus919/agent-skills

    Organize durable agent research outputs as summaries, analysis, and evidence dossiers.

    111 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Ascii City Engine

    magnus919/agent-skills

    Build portable, first-person colored ASCII city engines and small GIS-derived city packs.

    111 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Color Management

    magnus919/agent-skills

    Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.

    111 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    111 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Docker Compose

    magnus919/agent-skills

    Use Docker Compose to define, run, debug, and harden multi-container applications.

    111 GitHub stars~2k tokensUpdated yesterday
    Auto-check: notes
  • Fpga Development

    magnus919/agent-skills

    Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.

    111 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed

Questions about Llama Cpp

What does Llama Cpp do?

Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems. Llama Cpp is an agent skill from magnus919/agent-skills.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

When should I use Llama Cpp?

Llama Cpp fits situations like: building llama.cpp; inspecting GGUF models; running llama-cli; serving an OpenAI-compatible API with llama-server.

How do I install Llama Cpp in Claude Code?

Run `npx skills add magnus919/agent-skills --skill llama-cpp -a claude-code`. Or copy the skill folder (llama-cpp in magnus919/agent-skills) into .claude/skills/llama-cpp in your project. Claude Code loads it when a task matches its description.

How do I install Llama Cpp in Codex?

Run `npx skills add magnus919/agent-skills --skill llama-cpp -a codex`. Or copy the skill folder (llama-cpp in magnus919/agent-skills) into .agents/skills/llama-cpp in your project. Codex loads it when a task matches its description.

Can I use Llama Cpp in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill llama-cpp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llama-cpp, .gemini/skills/llama-cpp, .github/skills/llama-cpp and .opencode/skills/llama-cpp in your project.

What does Llama Cpp need to run?

SKILL.md names no scripts, command-line tools or credentials: Llama Cpp is instructions for the agent only. Our summary lists: Python 3; Docker. Compatibility (from SKILL.md): Requires a supported llama.cpp binary or a build environment. Model use requires a compatible GGUF file and sufficient disk and memory; accelerator paths require the matching driver and SDK..

Does Llama Cpp access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Llama Cpp safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Llama Cpp use?

Llama Cpp is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Llama Cpp use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.5k tokens, read only when the agent opens those files.

What are the alternatives to Llama Cpp?

Skills that share tags, products or a category with Llama Cpp: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), Mesh API (mr-tbot/mesh-api, 179 stars) and Vllm Deploy Simple (vllm-project/vllm-skills, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Llama Cpp?

magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 111 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 6, 2026.

Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.