Agent skill

Local Models

by glebis in glebis/claude-skills

Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama.

MITAuto-check passedAI & LLM Engineering

Install Local Models

skills CLI
$ npx skills add glebis/claude-skills --skill local-models -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install glebis/claude-skills local-models --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/local-models .claude/skills/local-models && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
local-models
GitHub stars
388
Token cost
~1.4k tokens
SKILL.md length
502 words
Files
4 (incl. scripts, references)
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama.

  • Cheap/bulk text work (summarize
  • SKILL.md covers When to use this skill, The core trick: reuse Ollama's…, Usage and Critical gotchas, plus 2 more sections
  • Runs Python scripts from its folder; calls brew
  • Local embeddings

What it does

Local Models is an agent skill from glebis/claude-skills. Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama. Use for cheap/bulk text work (summarize, classify, extract JSON, anonymize PII, translate, proofread, keywords), local embeddings, and offline image description — and prefer it over a cloud API whenever a task is privacy-sensitive, must run offline, is high-volume/low-stakes, or just needs a fast throwaway answer. Provides an lm CLI wrapper plus an OpenAI-compatible local server.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/serving-and-embeddings.md` and `scripts/ollama_blob.py`).

It sits in AI & LLM Engineering, covering LLM inference and serving and Embeddings. It works with Ollama, llama.cpp and OpenAI. The repository describes itself as: Collection of Claude Code skills for enhanced AI workflows. The licence is MIT.

When your agent uses it

  • Cheap/bulk text work (summarize
  • Local embeddings
  • Offline image description — and prefer it over a cloud API whenever a task is privacy-sensitive
  • Must run offline

Example prompts

  • “/local-models”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 7524dff. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Local Models loads about 1.4k tokens when it runs, and up to ~2.2k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 502 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from glebis/claude-skills at commit 7524dff, republished under its MIT licence (© glebis). 502 words, ~1,374 tokens.

Download SKILL.mdSave it as .claude/skills/local-models/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
local-models
description
Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama. Use for cheap/bulk text work (summarize, classify, extract JSON, anonymize PII, translate, proofread, keywords), local embeddings, and offline image description — and prefer it over a cloud API whenever a task is privacy-sensitive, must run offline, is high-volume/low-stakes, or just needs a fast throwaway answer. Provides an `lm` CLI wrapper plus an OpenAI-compatible local server.

local-models

Quick access to local LLMs through llama.cpp, reusing the GGUF models already pulled by Ollama (no re-download for text and embeddings). Everything runs on the machine — no API key, no network, no per-token cost.

When to use this skill

Reach for local models instead of a cloud API when the task is:

  • Privacy-sensitive — redacting PII, processing personal notes, health data, secrets-adjacent text. The data never leaves the machine.
  • Offline — no network available, or the user explicitly wants local-only.
  • High-volume / low-stakes — classifying or tagging hundreds of items, where a small model is good enough and cloud cost/latency would add up.
  • A fast throwaway — a quick summary, translation, or "what is this" where round-tripping to a frontier model is overkill.

Prefer a frontier (Claude) model when the task needs strong reasoning, long context, careful code, or high accuracy — these local models are small (0.6–4B).

The core trick: reuse Ollama's models

Ollama stores model weights as extension-less GGUF blobs under ~/.ollama/models/blobs/. These are ordinary GGUF files — llama.cpp loads them directly. scripts/ollama_blob.py reads Ollama's manifests and resolves a friendly name (e.g. qwen2.5:3b) to its weights blob path. No conversion, no duplicate downloads.

Usage

The entry point is scripts/lm. Run scripts/lm help for the full list. Invoke it with an absolute path, e.g. ~/ai_projects/claude-skills/local-models/scripts/lm.

bash
lm models                       # list local models (text / vision / embed)
lm ask [MODEL] "PROMPT"         # one-shot prompt (default qwen2.5:3b)
lm chat [MODEL]                 # interactive REPL

# Text presets — accept a file path, inline text, OR stdin:
lm summarize  report.md
cat notes.txt | lm tldr
lm keywords   article.txt
lm anonymize  transcript.txt         # → [NAME] [EMAIL] [PHONE] [ADDRESS] ...
lm proofread  draft.md
lm translate  German "Good morning"
lm classify   "praise,complaint,question"  feedback.txt   # → one label
lm extract    "invoice_number, total, due_date"  invoice.txt   # → JSON

# Vision (downloads model+projector once via HuggingFace — see note below):
lm describe-image photo.jpg
lm tag-image      screenshot.png
lm vision photo.jpg "What brand is the shoe?"

# Embeddings & serving:
lm embed "text to embed"             # → OpenAI-style JSON vector
lm serve qwen2.5:3b 8080             # OpenAI-compatible server on :8080

Output is clean (just the answer) — the wrapper drives llama-completion in single-turn mode and strips the chat-template scaffolding and llama.cpp logs.

Choosing a model

Defaults are tuned for clean, fast output and can be overridden per call:

  • General text presets → qwen2.5:3b (LM_TEXT_MODEL)
  • Classify / extract → qwen2.5:3b (LM_REASON_MODEL), run at temperature 0
  • Embeddings → jeffh/intfloat-multilingual-e5-large:f16 (LM_EMBED_MODEL)
  • Other envs: LM_NTOK (max tokens), LM_VISION_HF (vision repo), LM_DEBUG=1 (show llama.cpp logs)

Pass an explicit model as the first argument to ask/chat/embed/serve (e.g. lm ask qwen3:4b "...").

Show full SKILL.md (213 more words)Show less

Critical gotchas

  • Ollama's gemma3 GGUF does NOT load in stock llama.cpp. It fails with key not found in model: gemma3.attention.layer_norm_rms_epsilon because Ollama writes custom metadata keys mainline llama.cpp doesn't read. Use a qwen* model instead, or pull a community gemma3 GGUF via -hf. This is why the defaults are qwen, not gemma3.
  • qwen3:4b emits <think>…</think> reasoning blocks before its answer. Fine for ask/chat, but it pollutes preset output (JSON, labels) — the presets default to qwen2.5:3b to avoid this.
  • Vision has no Ollama blob to reuse. Ollama did not store an mmproj (vision projector) for qwen2.5vl, and llama.cpp needs one. So the vision commands use llama-mtmd-cli -hf ggml-org/Qwen2.5-VL-3B-Instruct-GGUF, which downloads model+projector (~2–3 GB) into ~/.cache/llama.cpp on first use, then runs offline. Warn the user before the first vision call.
  • Each one-shot call reloads the model (a few seconds for these small models). For many sequential calls, start a server once with lm serve and hit http://localhost:8080/v1/chat/completions — see references/serving-and-embeddings.md.

Reference material

  • references/serving-and-embeddings.md — running llama-server as an OpenAI-compatible endpoint (and pointing the llm CLI or any OpenAI client at it), plus local embeddings / RAG patterns with llama-embedding.

Requirements

  • llama.cpp installed (brew install llama.cpp) — provides llama-completion, llama-mtmd-cli, llama-embedding, llama-server.
  • Ollama with at least one pulled model (for the blob-reuse path). python3 for the resolver. No API keys.

© glebis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in local-models of glebis/claude-skills.

  • SKILL.md
  • references/serving-and-embeddings.md
  • scripts/lm
  • scripts/ollama_blob.py

Open the folder on GitHubat commit 7524dff

Compare with similar skills

Local Models next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Local Models compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Local Models this skillglebis/claude-skills388—~1.4kAutomated safety check: PassMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Gemma4 Local Deploymajiayu000/spellbook286—~875Automated safety check: NotesMIT
Vllmmagnus919/agent-skills111—~4.1kAutomated safety check: NotesMIT
Llama Cppmagnus919/agent-skills111—~2.3kAutomated safety check: PassMIT
Local AI App Integrationamd/skills395—~6kAutomated safety check: PassMIT

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Gemma4 Local Deploy

    majiayu000/spellbook

    在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4KM、64K/128K 长上下文、QAT Q40 @ 256K、左右对比演示之间选择,配置 tmux…

    286 GitHub stars~875 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Vllm

    magnus919/agent-skills

    Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

    111 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Llama Cpp

    magnus919/agent-skills

    Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

    111 GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Integrates local AI capabilities into applications using Embeddable Lemonade.

    395 GitHub stars~6k tokensUpdated today
    Media & CreativeAuto-check passed
  • Perfup

    raullenchai/Rapid-MLX

    Autonomous performance optimization: research, PoC, benchmark, implement, review, PR

    3.9k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from glebis/claude-skills

All 91 skills in this repo
  • Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.

    388 GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Automates a dedicated, logged-in Chrome instance per profile without ever closing the user's own open tabs or browser windows.

    388 GitHub stars~973 tokensUpdated 11 days ago
    Auto-check passed
  • Deep Research

    glebis/claude-skills

    This skill should be used when conducting comprehensive research on any topic using the OpenAI Deep Research API.

    388 GitHub stars~2.6k tokensUpdated 11 days ago
    Auto-check: notes
  • Elimination Research

    glebis/claude-skills

    This skill should be used for elimination-style research where the user wants to choose from a shortlist of products, tools, services, vendors, or other options using explicit criteria, numeric…

    388 GitHub stars~1.6k tokensUpdated 11 days ago
    Auto-check passed
  • Narrated HTML Presentations

    glebis/claude-skills

    Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations.

    388 GitHub stars~2.3k tokensUpdated 11 days ago
    Auto-check: notes
  • Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.

    388 GitHub stars~2.9k tokensUpdated 11 days ago
    Auto-check passed

Questions about Local Models

What does Local Models do?

Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama. Local Models is an agent skill from glebis/claude-skills.cpp, reusing models already downloaded by Ollama.

When should I use Local Models?

Local Models fits situations like: cheap/bulk text work (summarize; local embeddings; offline image description — and prefer it over a cloud API whenever a task is privacy-sensitive; must run offline.

How do I install Local Models in Claude Code?

Run `npx skills add glebis/claude-skills --skill local-models -a claude-code`. Or copy the skill folder (local-models in glebis/claude-skills) into .claude/skills/local-models in your project. Claude Code loads it when a task matches its description.

How do I install Local Models in Codex?

Run `npx skills add glebis/claude-skills --skill local-models -a codex`. Or copy the skill folder (local-models in glebis/claude-skills) into .agents/skills/local-models in your project. Codex loads it when a task matches its description.

Can I use Local Models in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add glebis/claude-skills --skill local-models -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/local-models, .gemini/skills/local-models, .github/skills/local-models and .opencode/skills/local-models in your project.

What does Local Models need to run?

Going by SKILL.md and its folder, Local Models needs Python for the scripts in its folder and the command-line tools its instructions call (brew). Our summary lists: Python 3.

Does Local Models access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Local Models safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Local Models use?

Local Models is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Local Models use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 779 tokens, read only when the agent opens those files.

What are the alternatives to Local Models?

Skills that share tags, products or a category with Local Models: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Gemma4 Local Deploy (majiayu000/spellbook, 286 stars), Vllm (magnus919/agent-skills, 111 stars) and Llama Cpp (magnus919/agent-skills, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Local Models?

glebis (a GitHub user) maintains it in glebis/claude-skills, which has 388 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on September 26, 2026.

Source: glebis/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.