Agent skill

Ollama

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when running open-weight LLMs locally with Ollama — pulling and tagging models, calling the local API, picking a quantization or GGUF, writing Modelfiles, and sizing VRAM and…

MITAuto-check passedAI & LLM Engineering

Install Ollama

skills CLI
$ npx skills add ericrisco/rsc-harness --skill ollama -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness ollama --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ollama .claude/skills/ollama && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ollama
GitHub stars
167
Token cost
~2.8k tokens
SKILL.md length
1,234 words
Files
6 (incl. scripts, references)
Skills in repo
227
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when running open-weight LLMs locally with Ollama — pulling and tagging models, calling the local API, picking a quantization or GGUF, writing Modelfiles, and sizing VRAM and…

  • Running open-weight LLMs locally with Ollama — pulling and tagging models
  • SKILL.md covers When to use / when not, Quickstart, Pick a model + quant and The API, plus 5 more sections
  • Runs Shell scripts from its folder; calls ollama and curl
  • Calling the local API

What it does

Ollama is an agent skill from ericrisco/rsc-harness. Use when running open-weight LLMs locally with Ollama — pulling and tagging models, calling the local API, picking a quantization or GGUF, writing Modelfiles, and sizing VRAM and RAM for the machine at hand. NOT remote or managed GPU serving and autoscaling (that is runpod), NOT downloading raw weights or datasets (that is huggingface), NOT retrieval pipeline design (that is rag).

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/api.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving and Retrieval-augmented generation. It works with Ollama, llama.cpp and Hugging Face. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Running open-weight LLMs locally with Ollama — pulling and tagging models
  • Calling the local API
  • Picking a quantization
  • Writing Modelfiles

Example prompts

  • “/ollama”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • ollama
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • ollama.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ollama loads about 2.8k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 1,234 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,234 words, ~2,843 tokens.

Download SKILL.mdSave it as .claude/skills/ollama/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
ollama
description
Use when running open-weight LLMs locally with Ollama — pulling and tagging models, calling the local API, picking a quantization or GGUF, writing Modelfiles, and sizing VRAM and RAM for the machine at hand. NOT remote or managed GPU serving and autoscaling (that is `runpod`), NOT downloading raw weights or datasets (that is `huggingface`), NOT retrieval pipeline design (that is `rag`).
tags
ollama, local-llm, gguf, quantization, self-hosted-inference
recommends
huggingface, runpod, modal, llm-pipeline, rag
origin
risco

Ollama — run open-weight LLMs on one box

Ollama serves GGUF models from a local daemon at http://localhost:11434, exposing both a native HTTP API and an OpenAI-compatible layer. Your job: reach for the right command, the right endpoint, and the right quant for the hardware in front of you — and recognize when the model does not fit and the work belongs on a remote GPU instead.

This skill owns: install/serve, pull/tag, the local API (native + OpenAI-compat), Modelfiles, quantization choice, and VRAM/RAM sizing on a single machine.

When to use / when not

Use when the model runs on this machine: pulling/running a model, fixing an OOM, choosing Q4 vs Q8, authoring a Modelfile, or wiring an app to localhost:11434.

Go elsewhere when:

  • Hosting behind a managed/remote GPU, autoscaling, or serverless inference → runpod, modal, replicate, together-fireworks, fal. Ollama is local, single-box, no autoscale.
  • Downloading raw weights, datasets, hf/transformers, repo management → huggingface.
  • Designing chunking / retrieval / reranking around a model → rag or embeddings-search.
  • Orchestrating multi-step calls, routing, pipeline evals → llm-pipeline / agent-eval.
  • Writing the prompt/system-message content itself → prompt-engineering.

(Those siblings live in the catalog by id; link them only once their SKILL.md exists on disk.)

Quickstart

bash
ollama serve                 # start the daemon (a desktop install already runs it)
ollama pull qwen3:8b         # download a model + tag; :8b is explicit — avoid bare :latest
ollama run qwen3:8b          # interactive REPL, or: ollama run qwen3:8b "summarize this"
ollama ps                    # what is LOADED in VRAM right now + when it unloads (keep_alive)
ollama list                  # what is on disk (pulled), not what is loaded
ollama show qwen3:8b         # template, params, context length, quant of a model
ollama rm qwen3:8b           # free disk; ollama stop qwen3:8b unloads from memory

ps vs list is the OOM-debug split: list is disk, ps is memory. A model only eats VRAM once a request loads it; it unloads after keep_alive (default 5m).

Pick a model + quant

Quantization trades VRAM for quality. The everyday default is Q4_K_M: roughly half the memory of fp16 for ~3–5% quality loss. Q8_0 is near-lossless at ~1 byte/param. fp16 is the unquantized ceiling at 2 bytes/param.

Sizing formula (weights only) — a rule of thumb, not a per-model spec sheet:

text
weights_GB ≈ params(B) × bytes_per_param × 1.2   # ×1.2 = runtime overhead
bytes_per_param:  Q4_K_M ≈ 0.5   Q8_0 ≈ 1.0   fp16 = 2.0
# then ADD the KV cache (see below) — it is NOT in this number.

These bytes/param are conservative round-downs of the measured k-quant rates: llama.cpp's quantize benchmark reports Q4_K_M ≈ 4.89 bits/weight (~0.6 byte/param) and Q8_0 ≈ 8.5 bits/weight (~1.06 byte/param) on Llama-3.1-8B (llama.cpp quantize README, accessed 2026-06-02). Rounding to 0.5 / 1.0 keeps the estimate on the safe side; the per-row GB figures in the table below are derived from this formula, not vendor-published numbers — verify with ollama show.

VRAM / unified memComfortable choice (Q4_K_M)Notes
8 GB7–8B Q4_K_M (~5–6 GB)leave headroom for KV cache + the OS
12 GBup to ~14B Q4_K_M (~9–10 GB)7–8B at Q8_0 also fits
16 GB14B Q4_K_M comfortably; 32B is tight32B Q4_K_M ≈ 20 GB — won't fit
24 GB32B Q4_K_M (~20 GB)70B does not fit at any usable quant
48 GB+ / 2×24 GB70B Q4_K_M (~40–48 GB)needs the full budget; long context pushes over
Mac unified (e.g. 64 GB)weights share RAM with everything elsebudget against total unified memory

KV cache is the trap. It grows ~linearly with num_ctx and lives in VRAM on top of the weights. At long context (e.g. 128K) a 70B can add tens of GB of cache — often more than people budget for. If you are tight: cap num_ctx, or shrink the cache with OLLAMA_KV_CACHE_TYPE=q8_0 (or q4_0). See references/hardware-sizing.md for the KV math and a per-context table.

Ollama runs a llama.cpp-backed engine (GGUF) by default, with a scheduler that reduces OOM crashes and improves multi-GPU placement. On Apple Silicon it can use an MLX backend (shipped in Ollama 0.19, per ollama.com/blog/mlx, 2026-03-30), but only on Macs with >32 GB of unified memory — below that gate it stays on the llama.cpp engine. None of this invents memory you don't have: when the box can't hold the model, that's a runpod/modal job, not a quant downgrade.

The API

Two surfaces, same daemon. Use native /api/chat when you want Ollama-specific fields (keep_alive, format as a JSON schema, think); use the OpenAI-compat /v1 layer to reuse an existing OpenAI SDK unchanged.

Native chat (/api/chat), non-streaming:

bash
curl http://localhost:11434/api/chat -d '{
  "model": "qwen3:8b",
  "messages": [{"role": "user", "content": "Name three primes."}],
  "stream": false,
  "options": {"temperature": 0.2, "num_ctx": 8192},
  "keep_alive": "10m"
}'

stream defaults to true (NDJSON, one object per line, final object has done: true + timing stats). options.num_ctx sets the context window for this request — it does not persist; bake it into a Modelfile if you want it permanent.

OpenAI-compatible — point any OpenAI SDK at localhost:11434/v1 with a dummy key:

python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")  # key is ignored
resp = client.chat.completions.create(
    model="qwen3:8b",
    messages=[{"role": "user", "content": "Name three primes."}],
    temperature=0.2,
)
print(resp.choices[0].message.content)

Structured output — pass a JSON schema as format (native) so the model is constrained to valid JSON:

bash
curl http://localhost:11434/api/chat -d '{
  "model": "qwen3:8b",
  "messages": [{"role": "user", "content": "Extract name and age from: Ana is 30."}],
  "stream": false,
  "format": {
    "type": "object",
    "properties": {"name": {"type": "string"}, "age": {"type": "integer"}},
    "required": ["name", "age"]
  }
}'

Tool calling (tools), multimodal (images as base64), embeddings (/api/embed), and the full field tables live in references/api.md. Endpoint map at a glance: /api/generate, /api/chat, /api/embed, /api/create, /api/pull, /api/show, /api/ps, /api/tags.

Show full SKILL.md (535 more words)Show less

Modelfiles

A Modelfile bakes a base model + system prompt + parameters into a new named model. Build with ollama create.

dockerfile
FROM qwen3:8b
SYSTEM "You are a terse senior code reviewer. Answer in bullet points."
PARAMETER num_ctx 16384
PARAMETER temperature 0.2
PARAMETER stop "<|im_end|>"
bash
ollama create reviewer -f Modelfile     # now: ollama run reviewer
  • FROM is required — a model tag or a local file (FROM ./model.gguf to import a raw GGUF).
  • PARAMETER num_ctx makes the context window permanent (vs the per-request options.num_ctx).
  • SYSTEM, TEMPLATE, LICENSE, ADAPTER (LoRA) round out the instruction set.

Quantize on create from an fp16/fp32 source:

bash
ollama create reviewer --quantize q4_K_M -f Modelfile   # FROM must be an fp16/fp32 model

--quantize only works when the FROM source is full-precision; you cannot re-quantize an already-Q4 model. To go from Hugging Face weights to a GGUF in the first place, that conversion is a huggingface job — Ollama imports the result.

When to leave the box

If the comfortable-choice row for your VRAM can't hold the model you actually need (e.g. you need 70B quality on a 12 GB laptop), stop downgrading quant — quality collapses below Q4 and you'll still OOM at real context. Move it to a remote GPU: runpod (rent a GPU), modal (serverless container + GPU autoscale), or a hosted endpoint (replicate, together-fireworks, fal). Ollama is the right tool until the weights + KV cache exceed the single box.

Anti-patterns

BadGoodWhy
Pull fp16 on a box that only fits Q4Pull Q4_K_M (or Q8_0 if it fits)fp16 is 4× the VRAM of Q4 for ~3–5% quality; you'll OOM for nothing
num_ctx: 128000 on a 12 GB GPUCap num_ctx to what fits; OLLAMA_KV_CACHE_TYPE=q8_0KV cache scales with context and sits on top of weights — long context dwarfs the model
/api/generate for a chat with history/api/chat with a messages arraygenerate is single-turn; you'd hand-concatenate history and break the chat template
ollama pull mistral:latest, assume it's smallPin an explicit tag (:7b, a quant tag) and ollama show it:latest size/quant drifts release to release; sizing breaks silently
Treat Ollama as a multi-tenant prod serverUse it local/single-box; scale → runpod/modalone daemon, limited parallelism (OLLAMA_NUM_PARALLEL); not built for fleet serving
Hardcode api.openai.com when target is localbase_url="http://localhost:11434/v1", dummy keythe OpenAI SDK works unchanged against the compat layer; no remote calls, no key leak
Downgrade to Q2 to force a 70B onto 12 GBPick a model that fits, or move to a remote GPUsub-Q4 quality drops sharply and it still won't fit at real context
Assume ollama list means it's loadedollama ps for memory, list for diska pulled model uses 0 VRAM until a request loads it

Verify

Run scripts/verify.sh [TARGET] from your project root (or a dir holding a Modelfile). Static by default — it needs neither Ollama installed nor a running daemon. It lints a Modelfile (FAIL if no FROM; WARN on unknown instructions or a num_ctx so high it will OOM consumer GPUs), notes whether app code points at the local localhost:11434 / /v1 endpoint vs only-remote hosts, and — only if ollama is on PATH — best-effort confirms a model is present (WARN, not FAIL). It exits non-zero only on a real FAIL; an empty/clean target passes.

References

  • references/api.md — full endpoint catalog, request/response field tables, OpenAI-compat path mapping, structured output, tool calling, streaming, embeddings (curl + Python).
  • references/hardware-sizing.md — the full quant ladder, VRAM formula derivation, KV-cache math + per-context table, per-model chart, Apple Silicon unified-memory notes, and the env knobs (OLLAMA_KV_CACHE_TYPE, OLLAMA_FLASH_ATTENTION, OLLAMA_NUM_PARALLEL, OLLAMA_MAX_LOADED_MODELS) for fitting tight boxes.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/ollama of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/api.md
  • references/hardware-sizing.md
  • scripts/verify.sh

Open the folder on GitHubat commit e3d5b33

Compare with similar skills

Ollama next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ollama compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ollama this skillericrisco/rsc-harness167—~2.8kAutomated safety check: PassMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Aider DelegateamElnagdy/delegate-skills2.3k3 repos~3kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT

Similar skills

  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 3 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed

More from ericrisco/rsc-harness

All 227 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    167 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    167 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    167 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    167 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    167 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    167 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed

Questions about Ollama

What does Ollama do?

A skill your agent uses when running open-weight LLMs locally with Ollama — pulling and tagging models, calling the local API, picking a quantization or GGUF, writing Modelfiles, and sizing VRAM and…. Ollama is an agent skill from ericrisco/rsc-harness. Use when running open-weight LLMs locally with Ollama — pulling and tagging models, calling the local API, picking a quantization or GGUF, writing Modelfiles, and sizing VRAM and RAM for the machine at hand.

When should I use Ollama?

Ollama fits situations like: running open-weight LLMs locally with Ollama — pulling and tagging models; calling the local API; picking a quantization; writing Modelfiles.

How do I install Ollama in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill ollama -a claude-code`. Or copy the skill folder (skills/ollama in ericrisco/rsc-harness) into .claude/skills/ollama in your project. Claude Code loads it when a task matches its description.

How do I install Ollama in Codex?

Run `npx skills add ericrisco/rsc-harness --skill ollama -a codex`. Or copy the skill folder (skills/ollama in ericrisco/rsc-harness) into .agents/skills/ollama in your project. Codex loads it when a task matches its description.

Can I use Ollama in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill ollama -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ollama, .gemini/skills/ollama, .github/skills/ollama and .opencode/skills/ollama in your project.

What does Ollama need to run?

Going by SKILL.md and its folder, Ollama needs a shell for the scripts in its folder and the command-line tools its instructions call (ollama and curl). Our summary lists: Python 3; A Bash shell.

Does Ollama access the network?

SKILL.md names 2 domains. As links in the text: github.com and ollama.com. This is read from the text; nothing was executed.

Is Ollama safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ollama use?

Ollama is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ollama use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Ollama?

Skills that share tags, products or a category with Ollama: Resolve (alexziskind1/model-shelf, 130 stars), Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ollama?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.