Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Ollama

skills CLI
$ npx skills add Prism-Shadow/penguin-harness --skill ollama -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Prism-Shadow/penguin-harness ollama --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/model-development/skills/ollama .claude/skills/ollama && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ollama
GitHub stars
2.5k
Token cost
~839 tokens
SKILL.md length
305 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.

  • Works in 6 steps: Ask the user which model to run; with no… → Pick the engine the user prefers: Ollama… → Install Ollama if missing, then ollama… → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Before you start, Suggested workflow, Install and Pull and run, plus 3 more sections
  • Calls ollama, curl and sh; reaches ollama.com

What it does

Ollama is an agent skill from Prism-Shadow/penguin-harness. Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.

Its SKILL.md is about 840 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Ollama, OpenAI and Qwen. The repository describes itself as: 🐧 Unified and Stable RSI Platform. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/ollama”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Ask the user which model to run; with no preference, recommend Qwen/Qwen3.5-0.8B (qwen3.5:0.8b).
  2. Pick the engine the user prefers: Ollama is the default; vLLM covers high-throughput GPU serving.
  3. Install Ollama if missing, then ollama pull qwen3.5:0.8b.
  4. Verify with curl http://localhost:11434/v1/models.
  5. Register the endpoint: penguin config model add ... --client-type openai-chat --base-url http://localhost:11434/v1 — a pulled Ollama model…
  6. Confirm the new entry with penguin config model list.

What it can do on your machine

Read from SKILL.md and the folder at commit d56d9ce. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ollama
    • curl
    • sh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ollama.com

    Also links to:

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ollama loads about 839 tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 305 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~839

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePipes a well-known installer script into a shellSKILL.md:37
    curl -fsSL https://ollama.com/install.sh | sh   # Linux; macOS/Windows use the desktop app

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Prism-Shadow/penguin-harness at commit d56d9ce, republished under its Apache-2.0 licence (© Prism-Shadow). 305 words, ~839 tokens.

Download SKILL.mdSave it as .claude/skills/ollama/SKILL.md (or your agent's skills folder).
name
ollama
description
Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.

Ollama Serving

Ollama runs open-weight models locally with automatic GPU detection and an OpenAI-compatible API on http://localhost:11434.

Before you start

If the user's message only invokes this skill (e.g. "use ollama skill") without a concrete request, ask the user what they want. Do not run any command until the goal is clear.

Ask the user which model to run; if they have no preference, recommend the small default Qwen/Qwen3.5-0.8B (ollama pull qwen3.5:0.8b). The model must fit the machine's RAM/VRAM.

Ollama runs everywhere — macOS, Linux and Windows, on CPUs as well as NVIDIA/AMD GPUs — so engine choice follows the user's preference: Ollama is the simple default, while vLLM targets high-throughput GPU serving. Check the current state first:

bash
ollama --version   # is Ollama installed?
ollama ps          # is the service already serving models?

If port 11434 is already serving, reuse that instance — never kill an existing Ollama process.

Suggested workflow

  1. Ask the user which model to run; with no preference, recommend Qwen/Qwen3.5-0.8B (qwen3.5:0.8b).
  2. Pick the engine the user prefers: Ollama is the default; vLLM covers high-throughput GPU serving.
  3. Install Ollama if missing, then ollama pull qwen3.5:0.8b.
  4. Verify with curl http://localhost:11434/v1/models.
  5. Register the endpoint: penguin config model add ... --client-type openai-chat --base-url http://localhost:11434/v1 — a pulled Ollama model is not visible to Penguin until added.
  6. Confirm the new entry with penguin config model list.

Install

bash
curl -fsSL https://ollama.com/install.sh | sh   # Linux; macOS/Windows use the desktop app

The service then listens on http://localhost:11434.

Pull and run

bash
ollama pull qwen3.5:0.8b   # download a model
ollama run qwen3.5:0.8b    # interactive chat (pulls first if missing)
ollama list                # downloaded models
ollama ps                  # models loaded in memory
ollama stop qwen3.5:0.8b   # unload a model

OpenAI-compatible endpoint

The endpoint is http://localhost:11434/v1; any non-empty API key is accepted (conventionally ollama):

bash
curl http://localhost:11434/v1/models

Context length

The default context window is small, and agent sessions need a large one. Raise it in the server's environment:

bash
OLLAMA_CONTEXT_LENGTH=32768 ollama serve   # systemd service: set it via `systemctl edit ollama`

Or bake it into a model variant with a Modelfile:

FROM qwen3.5:0.8b
PARAMETER num_ctx 32768
bash
ollama create qwen3.5-32k -f Modelfile

Register with PenguinHarness

Model configuration is the penguin CLI's job — penguin config model add registers an endpoint and penguin config model list shows what has been registered. A pulled Ollama model is not visible to Penguin until you add it:

bash
penguin config model add --provider custom --client-type openai-chat \
  --base-url http://localhost:11434/v1 --model-id qwen3.5:0.8b --api-key ollama
penguin config model list   # the new entry should now be listed

© Prism-Shadow, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/model-development/skills/ollama of Prism-Shadow/penguin-harness.

Open the folder on GitHubat commit d56d9ce

Compare with similar skills

Ollama next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ollama compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ollama this skillPrism-Shadow/penguin-harness2.5k—~839Automated safety check: NotesApache-2.0
Page AgentTommy-yw/RunbookHermes5463 repos~2.3kAutomated safety check: NotesMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Perfupraullenchai/Rapid-MLX3.9k—~1.6kAutomated safety check: NotesCustom licence
Ideer Daily Paper ChatbotAI45Lab/iDeer416—~3kAutomated safety check: NotesAGPL-3.0
Agent Frameworkjihadkhawaja/Egroo178—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Page Agent

    Tommy-yw/RunbookHermes

    Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…

    546 GitHub starsUsed in 3 repos~2.3k tokens
    Productivity & AutomationAuto-check: notes
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Perfup

    raullenchai/Rapid-MLX

    Autonomous performance optimization: research, PoC, benchmark, implement, review, PR

    3.9k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Use iDeer as a daily paper-reading workflow for chatbot-first users such as Codex, Gemini, or ChatGPT.

    416 GitHub stars~3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Framework

    jihadkhawaja/Egroo

    Build, extend, and debug AI agents in Egroo using the Microsoft Agent Framework (C .NET).

    178 GitHub stars~1.9k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from Prism-Shadow/penguin-harness

All 31 skills in this repo
  • A2ui

    Prism-Shadow/penguin-harness

    Make a reply easier to read and act on with rich blocks inside ordinary Markdown — a choice the user picks from, a form that collects several answers, a procedure as steps with warnings in place, a…

    2.5k GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed
  • Penguin Harness Dev

    Prism-Shadow/penguin-harness

    A skill your agent uses when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release…

    2.5k GitHub stars~3.4k tokensUpdated 2 days ago
    Auto-check passed
  • Bento Slides

    Prism-Shadow/penguin-harness

    Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON.

    2.5k GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed
  • Penguin Harness Manual Test

    Prism-Shadow/penguin-harness

    A skill your agent uses when standing PenguinHarness up to try a change by hand — launching the Web App, the desktop shell, the landing page, the docs site or the component gallery to click through…

    2.5k GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Penguin Harness Frontend

    Prism-Shadow/penguin-harness

    A skill your agent uses when changing the PenguinHarness Web App (packages/web) or the shared UI package — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a…

    2.5k GitHub stars~6.4k tokensUpdated 2 days ago
    Auto-check passed
  • Browser Automation

    Prism-Shadow/penguin-harness

    Drive the PenguinHarness agent browser — the desktop app's built-in browser or the user's own Chrome — from the shell with penguin browser: open pages, read them as simplified HTML or text, act with…

    2.5k GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check: warnings

Questions about Ollama

What does Ollama do?

Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents. Ollama is an agent skill from Prism-Shadow/penguin-harness. Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.

When should I use Ollama?

Ollama fits situations like: tasks that involve LLM inference and serving.

How do I install Ollama in Claude Code?

Run `npx skills add Prism-Shadow/penguin-harness --skill ollama -a claude-code`. Or copy the skill folder (plugins/model-development/skills/ollama in Prism-Shadow/penguin-harness) into .claude/skills/ollama in your project. Claude Code loads it when a task matches its description.

How do I install Ollama in Codex?

Run `npx skills add Prism-Shadow/penguin-harness --skill ollama -a codex`. Or copy the skill folder (plugins/model-development/skills/ollama in Prism-Shadow/penguin-harness) into .agents/skills/ollama in your project. Codex loads it when a task matches its description.

Can I use Ollama in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Prism-Shadow/penguin-harness --skill ollama -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ollama, .gemini/skills/ollama, .github/skills/ollama and .opencode/skills/ollama in your project.

What does Ollama need to run?

Going by SKILL.md and its folder, Ollama needs the command-line tools its instructions call (ollama, curl and sh).

Does Ollama access the network?

SKILL.md names 2 domains. In commands or code: ollama.com; the agent is likely to contact it when it follows the instructions. As links in the text: huggingface.co. This is read from the text; nothing was executed.

Is Ollama safe to install?

Our automated static check of SKILL.md found notes only (pipes a well-known installer script into a shell), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Ollama use?

Ollama is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ollama use?

About 839 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ollama?

Skills that share tags, products or a category with Ollama: Page Agent (Tommy-yw/RunbookHermes, 546 stars), Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Perfup (raullenchai/Rapid-MLX, 3.9k stars) and Ideer Daily Paper Chatbot (AI45Lab/iDeer, 416 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ollama?

Prism-Shadow (a GitHub organization) maintains it in Prism-Shadow/penguin-harness, which has 2,464 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: Prism-Shadow/penguin-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.