Agent skill

Engine Ollama

by autonomous-ai in autonomous-ai/openharness

Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama.

MITAuto-check: notesAI & LLM Engineering

Install Engine Ollama

skills CLI
$ npx skills add autonomous-ai/openharness --skill engine-ollama -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/openharness engine-ollama --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-ollama .claude/skills/engine-ollama && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
engine-ollama
GitHub stars
1.2k
Token cost
~2.2k tokens
SKILL.md length
1,097 words
Files
1
Skills in repo
99
Repo updated
First seen
Licence
MIT

At a glance

Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama.

  • Works in 3 steps: The PID is alive. 2. GET / answers… → One bounded request through… → GET /api/ps shows the model with…
  • Tasks that involve LLM inference and serving
  • SKILL.md covers When to use it, MLX inside Ollama (Apple…, Where its models live / how to… and Installed? Running?, plus 7 more sections
  • Calls ollama and brew; reaches raw.githubusercontent.com

What it does

Engine Ollama is an agent skill from autonomous-ai/openharness. Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama. Load before touching Ollama, its models or its settings.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Ollama and llama.cpp. The repository describes itself as: The ultimate harness for coding agents and beyond. All your agents. All your machines. One command center. Start with code, then follow your curiosity and build across… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/engine-ollama”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The PID is alive. 2. GET / answers "Ollama is running" [run]. 3. GET /api/version → 200 [run].
  2. One bounded request through /v1/chat/completions with the model id, max_tokens 16 and
  3. GET /api/ps shows the model with context_length ≥ 65536 [doc]. FAIL context means this Ollama

What it can do on your machine

Read from SKILL.md and the folder at commit 74c2733. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ollama
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • raw.githubusercontent.com

    Also links to:

    • docs.ollama.com
    • github.com
    • ollama.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Engine Ollama loads about 2.2k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 1,097 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:126
    cOS app: quit from the menu bar; Linux: `sudo systemctl stop ollama` [doc].

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/openharness at commit 74c2733, republished under its MIT licence (© autonomous-ai). 1,097 words, ~2,171 tokens.

Download SKILL.mdSave it as .claude/skills/engine-ollama/SKILL.md (or your agent's skills folder).
name
engine-ollama
description
Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama. Load before touching Ollama, its models or its settings.

Ollama

Official docs, read 2026-09-29 (source files in github.com/ollama/ollama/tree/main/docs): FAQ · Context length · OpenAI compatibility · Thinking · CLI · macOS · Linux · Import · API reference Tested: Ollama 0.34.4 (brew install ollama), MacBook Pro M1 Pro 32 GB, macOS 26.6, 2026-09-29. Tags: [doc] official page above, [run] seen on the tested machine, [?] unverified.

When to use it

  • A model in Ollama's store (fleet models → START WITH ollama): Ollama runs it. It downloaded the file and ships its own engine for new architectures, which Grid's engine can refuse [?].
    • Answering: adopt it as it is; it is the person's app, so never restart or reconfigure it.
    • (start it): start one yourself (below). An Ollama that is off is the normal case, not a reason to switch engines [run].
  • Ollama not installed: its blobs are GGUF files, so fleet models names Grid's llama.cpp — link the blob into ~/.grid/models and join --serve [run].
  • A GGUF from another app enters Ollama only by ollama create (a Modelfile with FROM <file>), and that copies the whole file into Ollama's store (+624 MB for a 640 MB GGUF) [run]. So never do it unasked; when the person wants Ollama and it has nothing suitable, offer it with the size it adds on disk, beside the no-copy choice (the file with its own app). Never ollama pull without the go-ahead: it downloads.

MLX inside Ollama (Apple silicon)

Since 0.19 Ollama has its own MLX engine on Apple silicon, announced as a preview on 2026-03-30 (blog) [doc]. It runs on MLX only the architectures registered in its source — read the list live, never from memory: https://raw.githubusercontent.com/ollama/ollama/main/mlxrunner/model/architectures/architectures.go (one import per architecture folder under mlxrunner/model/) [doc]. Compare with model_type from fleet model-facts. Everything else keeps running on Ollama's GGML engine. The announced models are Ollama-library tags in their own quantization, and the post asks for more than 32 GB of unified memory [doc]; importing other MLX models is not yet documented [?].

Where its models live / how to list them

  • Default dir: macOS ~/.ollama/models, Linux service /usr/share/ollama/.ollama/models, Windows C:\Users\%username%\.ollama\models; OLLAMA_MODELS moves it [doc].
  • Layout: manifests/registry.ollama.ai/library/<model>/<tag> (a JSON file) and blobs/sha256-<hex>; the manifest layer application/vnd.ollama.image.model names the weights blob, a GGUF [run]. Model id is <model>:<tag>; other namespaces are <user>/<model>:<tag> [?].
  • fleet models already reads the manifests: entries with source: ollama, the blob path, and alsoAt when another app links the same file [run].
  • From a running server: GET /api/tags (downloaded), GET /api/ps (loaded, with context_length and size_vram) [doc], GET /v1/models (owned_by is the Ollama user, library by default) [doc][run]. POST /api/show {"model":ID} lists capabilities (completion, tools, thinking, vision) [doc].

Installed? Running?

  • Installed: command -v ollama or /Applications/Ollama.app. ollama --version with no server prints "Warning: could not connect to a running Ollama instance" and the client version [run].
  • The macOS app starts the server at login (a login item) [doc]; brew install ollama starts nothing [run]; the Linux installer creates a systemd service ollama [doc].
  • Running: fleet models → engines[] with kind: ollama. Default bind is 127.0.0.1:11434 [doc].

Start an already-downloaded model (Ollama not running, or running with too small a window)

Port P from outside machine.listeningPorts, and never 11434: that is the Ollama app's own, and it must still start when the person opens it [run]. The port is set only through OLLAMA_HOST (no --port) [doc]:

"$GRID_FLEET" serve ollama-P --env OLLAMA_HOST=127.0.0.1:P --env OLLAMA_CONTEXT_LENGTH=65536 -- ollama serve

(fleet serve keeps it alive after your shell returns; log run/ollama-P.log, PID run/ollama-P.pid.)

  • Foreground process; the log says Listening on 127.0.0.1:P (version …) when ready [run].
  • Every other ollama command must carry the same OLLAMA_HOST, or it talks to 11434 [run].
  • Nothing downloads unless you pull; the model loads on its first request [run].
  • Context: the default depends on GPU memory — 4K below 24 GiB, 32K at 24–48 GiB, 256K at 48 GiB or more [doc] (the FAQ still says 4096 [doc]). Agents need at least 64000 [doc], and every model here is used by an agent: always set OLLAMA_CONTEXT_LENGTH to 65536 or more. The OpenAI API cannot set context per request [doc].
Show full SKILL.md (429 more words)Show less

Ready means

Run "$GRID_FLEET" verify --at http://127.0.0.1:P/v1 --model <model>:<tag> --kind ollama — it performs these checks with deadlines and prints each one. What it checks:

  1. The PID is alive. 2. GET / answers "Ollama is running" [run]. 3. GET /api/version → 200 [run].
  2. One bounded request through /v1/chat/completions with the model id, max_tokens 16 and "reasoning_effort":"none" for a thinking model → non-empty content [doc][run].
  3. GET /api/ps shows the model with context_length ≥ 65536 [doc]. FAIL context means this Ollama runs every model with that window [run]: leave the grid, then leave their Ollama as it is and start a second one with OLLAMA_CONTEXT_LENGTH (above) [run].

Join Harness Compute

"$GRID_FLEET" run -- join GRID --at http://127.0.0.1:P/v1 -m <model>:<tag> --advertise-as ALIAS

/v1 is required: without it /models and /chat/completions answer 404, and Grid's capability probe records JSON output as unsupported without any error [run]. Grid's own detector finds Ollama only on 11434 [run].

Tool calls, JSON output, thinking

  • /v1/chat/completions supports tools, response_format and reasoning_effort [doc]. reasoning_effort: "none" asks for no thinking; the native API uses "think": false [doc].
  • A model's thinking values and default: /api/show → thinking.values, thinking.default [doc].
  • Tool support is per model: check capabilities contains tools before offering it for coding [doc].

Memory and speed knobs (server environment, apply to all models)

VariableDefaultEffect
OLLAMA_CONTEXT_LENGTHby GPU memory (above)context for every model [doc]
OLLAMA_NUM_PARALLEL1requests at once per model; memory scales with parallel × context [doc]
OLLAMA_KV_CACHE_TYPEf16q8_0 ≈ half the KV memory, q4_0 ≈ a quarter; needs flash attention [doc]
OLLAMA_FLASH_ATTENTIONautomatic1 forces on, 0 off [doc]
OLLAMA_KEEP_ALIVE5mhow long an idle model stays loaded [doc]
OLLAMA_MAX_LOADED_MODELS3 × GPUs (3 on CPU)models loaded at once [doc]

brew services start ollama sets OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0 [run: brew caveat]. ollama ps → PROCESSOR must read 100% GPU; a CPU/GPU split is slow [doc].

Stop

  • A server you started: "$GRID_FLEET" stop ollama-P.
  • Unload one model but keep the server: OLLAMA_HOST=… ollama stop <model> [doc].
  • The person's app or service: ask first. macOS app: quit from the menu bar; Linux: sudo systemctl stop ollama [doc].

Known failures → what to do

SignDo
404 on /models or /chat/completionsthe URL lacks /v1 [run]
content empty, thinking fulladd reasoning_effort: "none" [doc]
/api/ps context below 65536the person's Ollama runs its default (4K or 32K); start a second Ollama with OLLAMA_CONTEXT_LENGTH instead of changing their app
503 "server is overloaded"queue full (OLLAMA_MAX_QUEUE, default 512) [doc]; wait, do not retry in a loop
PROCESSOR shows CPU sharemodel plus context does not fit the GPU; smaller context or model [doc]
model not foundnot downloaded; offer the download as a slow step, never pull silently

© autonomous-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in store/agents/autonomous-grid/skills/engine-ollama of autonomous-ai/openharness.

Open the folder on GitHubat commit 74c2733

Compare with similar skills

Engine Ollama next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Engine Ollama compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Engine Ollama this skillautonomous-ai/openharness1.2k—~2.2kAutomated safety check: NotesMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Ollama Optimizerluongnv89/skills131—~4.1kAutomated safety check: NotesMIT
Local LLM Expertsickn33/agentic-awesome-skills47k2 repos~1.6kAutomated safety check: PassMIT
Jetson LLM BenchmarkNVIDIA/skills3.5k1 repos~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Ollama Optimizer

    luongnv89/skills

    Optimize Ollama configuration for the current machine's hardware.

    131 GitHub stars~4.1k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Local LLM Expert

    sickn33/agentic-awesome-skills

    Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio.

    47k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

    3.5k GitHub starsUsed in 1 repo~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Cost Local

    ruvnet/ruflo

    Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a…

    74k GitHub stars~336 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from autonomous-ai/openharness

All 99 skills in this repo
  • G-code Slicer Tool

    autonomous-ai/openharness

    Slices 3D mesh files into printer-profiled plain G-code through real slicer CLIs, with backend discovery, input inspection, dry runs and static validation.

    1.2k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Home Assistant Automation Builder

    autonomous-ai/openharness

    Turns a home-automation request into standard, testable automations.yaml, run against Home Assistant Core's real triggers and verified with its own trace tool.

    1.2k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Score Music Composer

    autonomous-ai/openharness

    Turns a musical brief into LilyPond concert-pitch music, checked parts for each instrument and a playable practice pack.

    1.2k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • OrcaSlicer 3MF and G-code Workflow

    autonomous-ai/openharness

    Turns an STL and explicit printer and material requirements into compared OrcaSlicer plans, an editable 3MF project, checked G-code and a portable handoff.

    1.2k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Sheets and Docs Report Builder

    autonomous-ai/openharness

    Builds an editable DOCX report, a formula-driven XLSX workbook and a fresh LibreOffice PDF preview from one structured source file, then checks them together.

    1.2k GitHub stars~708 tokensUpdated today
    Auto-check passed
  • Bambu Labs

    autonomous-ai/openharness

    Dry-run, upload, and cautiously initiate local Bambu Lab print jobs from validated plain .gcode, using Bambu LAN FTPS/MQTT handoffs.

    1.2k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check: warnings

Works with

Questions about Engine Ollama

What does Engine Ollama do?

Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama. Engine Ollama is an agent skill from autonomous-ai/openharness. Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama.

When should I use Engine Ollama?

Engine Ollama fits situations like: tasks that involve LLM inference and serving.

How do I install Engine Ollama in Claude Code?

Run `npx skills add autonomous-ai/openharness --skill engine-ollama -a claude-code`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-ollama in autonomous-ai/openharness) into .claude/skills/engine-ollama in your project. Claude Code loads it when a task matches its description.

How do I install Engine Ollama in Codex?

Run `npx skills add autonomous-ai/openharness --skill engine-ollama -a codex`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-ollama in autonomous-ai/openharness) into .agents/skills/engine-ollama in your project. Codex loads it when a task matches its description.

Can I use Engine Ollama in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/openharness --skill engine-ollama -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/engine-ollama, .gemini/skills/engine-ollama, .github/skills/engine-ollama and .opencode/skills/engine-ollama in your project.

What does Engine Ollama need to run?

Going by SKILL.md and its folder, Engine Ollama needs the command-line tools its instructions call (ollama and brew).

Does Engine Ollama access the network?

SKILL.md names 4 domains. In commands or code: raw.githubusercontent.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.ollama.com, github.com and ollama.com. This is read from the text; nothing was executed.

Is Engine Ollama safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Engine Ollama use?

Engine Ollama is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Engine Ollama use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Engine Ollama?

Skills that share tags, products or a category with Engine Ollama: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Resolve (alexziskind1/model-shelf, 130 stars), Ollama Optimizer (luongnv89/skills, 131 stars) and Local LLM Expert (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Engine Ollama?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/openharness, which has 1,194 GitHub stars. The repository holds 99 skills in this directory. The repository was last updated on October 9, 2026.

Source: autonomous-ai/openharness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.