Agent skill

Run Local Model

by autonomous-ai in autonomous-ai/openharness

Details behind AGENTS.md's 'Starting a model on this computer' flow: reading fleet models, sizing a start, choosing file and engine, picking a model for an engine with none, what fleet verify…

MITAuto-check passedAgent Workflows

Install Run Local Model

skills CLI
$ npx skills add autonomous-ai/openharness --skill run-local-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/openharness run-local-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/store/agents/autonomous-grid/skills/run-local-model .claude/skills/run-local-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-local-model
GitHub stars
1.1k
Token cost
~1.8k tokens
SKILL.md length
1,028 words
Files
1
Skills in repo
100
Repo updated
First seen
Licence
MIT

At a glance

Details behind AGENTS.md's 'Starting a model on this computer' flow: reading fleet models, sizing a start, choosing file and engine, picking a model for an engine with none, what fleet verify…

  • Works in 4 steps: "$GRID_FLEET" recipe vllm|sglang… → "$GRID_FLEET" model-facts ORG/NAME|DIR —… → Agent indexes: docs.ollama.com/llms.txt, → …
  • Tasks that involve Agent instruction files
  • SKILL.md covers Purpose → settings (the only…, Reading fleet models, Size and Choose the file, then the engine, plus 4 more sections
  • Calls ollama

What it does

Run Local Model is an agent skill from autonomous-ai/openharness. Details behind AGENTS.md's 'Starting a model on this computer' flow: reading fleet models, sizing a start, choosing file and engine, picking a model for an engine with none, what fleet verify checks, and where to read when unsure. Open it when a step of that flow needs more than the flow says.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Agent instruction files. The repository describes itself as: The ultimate harness for coding agents and beyond. All your agents. All your machines. One command center. Start with code, then follow your curiosity and build across… The licence is MIT.

When your agent uses it

  • Tasks that involve Agent instruction files

Example prompts

  • “Starting a model on this computer”
  • “Use the run-local-model skill to detail behind AGENTS.md's 'Starting a model on this computer' flow: reading fleet models, sizing a start, choosing…”
  • “/run-local-model”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. "$GRID_FLEET" recipe vllm|sglang ORG/NAME — the official recipe for that exact model (exit 3: none).
  2. "$GRID_FLEET" model-facts ORG/NAME|DIR — architecture, context, sampling defaults, the template's
  3. Agent indexes: docs.ollama.com/llms.txt,
  4. Nothing says it: tell the person this model cannot be set up reliably here and offer the next one.

What it can do on your machine

Read from SKILL.md and the folder at commit 50da5db. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ollama

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • recipes.vllm.ai
    • docs.ollama.com
    • lmstudio.ai
    • docs.sglang.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Local Model loads about 1.8k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 1,028 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/openharness at commit 50da5db, republished under its MIT licence (© autonomous-ai). 1,028 words, ~1,807 tokens.

Download SKILL.mdSave it as .claude/skills/run-local-model/SKILL.md (or your agent's skills folder).
name
run-local-model
description
Details behind AGENTS.md's 'Starting a model on this computer' flow: reading `fleet models`, sizing a start, choosing file and engine, picking a model for an engine with none, what `fleet verify` checks, and where to read when unsure. Open it when a step of that flow needs more than the flow says.

Run a local model — details

The flow itself is in AGENTS.md. Tags: [run] seen on a real machine (M1 Pro 32 GB, macOS 26.6, 2026-09-29), [doc] official docs (see the engine-* skill), [code] read in the tool's source.

Purpose → settings (the only thing asked)

Coding · Chat and writing · Reading images · Just something fast. Context is never below 64K: every model here is used by an agent whose own prompt fills a small window; Ollama's docs set the same floor for agents [doc]. When 64K does not fit, take a smaller model, never a smaller window.

PurposeContextAt onceAlso needs
Coding128K when it fits, else 64K1tool calls
Chat and writing, or fast64K1—
Reading images64K1a projector beside the file

Thinking is off for everyday use.

Reading fleet models

  • machine.accelerators[]: a GPU counts only with active: true (its own tool answered). active: false comes with an error (e.g. no driver): say it in one line; plan no GPU engine on it.
  • machine.engines[]: installed engines, running or not. machine.canRun: the kinds this hardware can run at all — never propose one outside it.
  • machine.memory.availableBytes, swapUsedBytes: room right now. Metal and device-info do not see other apps (Metal said 25 GiB free while macOS was 11 GB into swap [run]).
  • machine.accelerators[].totalBytes: the GPU ceiling, below total RAM on a Mac — read it here, never assume it; device-info reports a different figure there [run] — use the smaller.
  • engines[]: answering engines and their exact --at URL. openai-compatible whose "models" are not models is another app: leave it alone.
  • models[]: one entry per real file (alsoAt = other apps holding the same file), with format, bytes, projector, and for GGUF contextLength, kvBytesPerToken, toolCalls, unsupportedTensorTypes.

Size

need = weights + context × kvBytesPerToken × slots + 0.5 GB (+ projector when vision is on)

kvBytesPerToken matched what llama.cpp allocated exactly [run]; files of similar size needed from 20 KiB to 160 KiB per token, so never skip this. Null (latent attention): start at 64K and read the engine's memory report. Fits when need ≤ availableBytes + 3 GB and ≤ the GPU ceiling (macOS moves idle pages to swap once; swap still climbing a minute after start means too big). Too big: a smaller file, or ask the one trade-off — "close other apps first, or a lighter model beside your work?".

Choose the file, then the engine

Drop: unsupportedTensorTypes not empty (llama.cpp refuses them [run]); context below 64K; no tool calls for coding; no projector for images; anything that does not fit. Prefer newer families, more parameters at 4-bit over fewer at 8-bit, and mixture-of-experts models when memory allows.

The fileMac (Apple silicon)Linux + active NVIDIA/AMD GPUCPU only
served by an engine already answeringjoin --at URL/v1 -m ID --advertise-as ALIASsamesame
GGUF from any appGrid's engine (link into ~/.grid/models)samesame
MLX foldermlx-lm (engine-mlx-lm)——
Hugging Face safetensorsmlx-lmvLLM or SGLang from the model's recipefind a GGUF

Grid only serves from ~/.grid/models: its launcher keeps just the file name of --serve and looks there; a projector must sit beside it [code: grid shared/engine/launcher.py]. A symlink costs no disk. Never ollama create to reuse a file (it copies it [run]); never vLLM or SGLang on a Mac (CPU-only / no macOS build [run]); never download a second copy only to switch engines. On Apple silicon, when a new download is needed and mlx-lm or LM Studio is installed, an MLX build is sound — Ollama itself moved its Apple engine to MLX [doc: ollama.com/blog/mlx].

On Apple silicon with mlx-lm (or LM Studio) installed, MLX first: when the same model exists as an MLX folder and as a GGUF and the MLX one fits, serve the MLX one. The choice is still model-first — never a bigger MLX model that does not fit over a smaller GGUF that does. The report names what was passed over and why, one line each ("<model> (<format>): needs <N> GB, <M> GB free"), so "why not MLX?" is answered before anyone asks.

Show full SKILL.md (368 more words)Show less

An engine installed, nothing to serve

Offer 2–3 that fit (plain name, GB, what it is good at) plus "none of these"; the download waits for the go-ahead.

  • Mac with mlx-lm or LM Studio: "$GRID_FLEET" candidates mlx [--search WORDS] [--sort downloads|trending|recent] — mlx-community models sized from their real files, cache at 64K, fits against the GPU ceiling.
  • Linux GPU with vLLM or SGLang: pick families from recipes.vllm.ai/llms.txt or the SGLang cookbook, then fleet recipe each; vramMinimumGb against the GPU decides.
  • Grid's engine or Ollama with nothing: Grid's catalog (grid-operations step 4), pulled into Grid.
  • The person names an engine outside canRun: say why in one line (no active GPU, not a Mac).

This computer already on that grid

In remote mode a computer joins a grid as one identity, and Grid's --serve engine cannot share it: a second model is refused with "can't join a multi-engine identity. Run grid leave, then re-join every engine as external --at <url> -m <model>" [run]. Ask Replace X with Y · Keep X; never leave an engine the person did not agree to.

What fleet verify checks

Engine: ready (/models within 180 s, every 3 s), listed, answer (max_tokens 16, thinking off; reasoning-only output fails), tool call (read_file must come back as tool_calls), speed. With --grid: relay listed (every 10 s up to 300 s) and relay answer (up to 420 s, a "still waiting" line every 15 s). Exit 0 only when all pass. Seen end to end: an mlx-lm engine joined a local grid with --at …/v1 in 6 s and passed all seven in 10 s; a Grid-engine GGUF passed at 19 tok/s [run].

When unsure: read, never guess

  1. "$GRID_FLEET" recipe vllm|sglang ORG/NAME — the official recipe for that exact model (exit 3: none).
  2. "$GRID_FLEET" model-facts ORG/NAME|DIR — architecture, context, sampling defaults, the template's tool-call syntax, the authors' serve commands, and whether each engine lists the architecture (listed: null = cannot tell).
  3. Agent indexes: docs.ollama.com/llms.txt, lmstudio.ai/llms.txt, recipes.vllm.ai/llms.txt, docs.sglang.io/llms.txt; vLLM and mlx-lm docs as raw markdown on GitHub.
  4. Nothing says it: tell the person this model cannot be set up reliably here and offer the next one.

fleet models reads this computer only; a Harness-linked machine's disk is not visible yet.

© autonomous-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in store/agents/autonomous-grid/skills/run-local-model of autonomous-ai/openharness.

Open the folder on GitHubat commit 50da5db

Compare with similar skills

Run Local Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Local Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Local Model this skillautonomous-ai/openharness1.1k—~1.8kAutomated safety check: PassMIT
Using Agent Skillsaddyosmani/agent-skills103k4 repos~2.4kAutomated safety check: PassMIT
Claude ReflectBayramAnnakov/claude-reflect1.7k2 repos~627Automated safety check: PassMIT
Writing For Agentsbestofjs/bestofjs3.1k19 repos~2.7kAutomated safety check: PassMIT
Neat-Freak Knowledge CloseoutKKKKhazix/khazix-skills21k—~1.9kAutomated safety check: PassMIT
Task Observerrebelytics/one-skill-to-rule-them-all3.2k1 repos~12kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Using Agent Skills

    addyosmani/agent-skills

    Meta-skill for choosing which workflow skill fits the task at hand, plus always-on habits: surface assumptions, stop on confusion, push back, keep it simple and stay in scope.

    103k GitHub starsUsed in 4 repos~2.4k tokens
    Agent WorkflowsAuto-check passed
  • Claude Reflect

    BayramAnnakov/claude-reflect

    Self-learning system that captures corrections during sessions and reminds users to run /reflect to update CLAUDE.md.

    1.7k GitHub starsUsed in 2 repos~627 tokens
    Agent WorkflowsAuto-check passed
  • Writing For Agents

    bestofjs/bestofjs

    Writing documents for agents. An agent skill from bestofjs/bestofjs.

    3.1k GitHub starsUsed in 19 repos~2.7k tokens
    Agent WorkflowsAuto-check passed
  • Neat-Freak Knowledge Closeout

    KKKKhazix/khazix-skills

    Brings project docs, agent rule files, authorized memory and leftover workspace files back in line with what the code and runtime actually do at the end of a work session.

    21k GitHub stars~1.9k tokensUpdated 7 days ago
    Agent WorkflowsAuto-check passed
  • Task Observer

    rebelytics/one-skill-to-rule-them-all

    Monitors task execution for skill improvement opportunities.

    3.2k GitHub starsUsed in 1 repo~12k tokens
    Agent WorkflowsAuto-check passed
  • SkillOpt Sleep Cycle

    microsoft/SkillOpt

    Official

    Runs an on-demand or nightly sleep cycle that reviews past Claude Code sessions and proposes validated updates to CLAUDE.md and skills.

    18k GitHub stars~2.3k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed

More from autonomous-ai/openharness

All 100 skills in this repo
  • G-code Slicer Tool

    autonomous-ai/openharness

    Slices 3D mesh files into printer-profiled plain G-code through real slicer CLIs, with backend discovery, input inspection, dry runs and static validation.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Home Assistant Automation Builder

    autonomous-ai/openharness

    Turns a home-automation request into standard, testable automations.yaml, run against Home Assistant Core's real triggers and verified with its own trace tool.

    1.1k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Score Music Composer

    autonomous-ai/openharness

    Turns a musical brief into LilyPond concert-pitch music, checked parts for each instrument and a playable practice pack.

    1.1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • OrcaSlicer 3MF and G-code Workflow

    autonomous-ai/openharness

    Turns an STL and explicit printer and material requirements into compared OrcaSlicer plans, an editable 3MF project, checked G-code and a portable handoff.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Sheets and Docs Report Builder

    autonomous-ai/openharness

    Builds an editable DOCX report, a formula-driven XLSX workbook and a fresh LibreOffice PDF preview from one structured source file, then checks them together.

    1.1k GitHub stars~708 tokensUpdated today
    Auto-check passed
  • Bambu Labs

    autonomous-ai/openharness

    Dry-run, upload, and cautiously initiate local Bambu Lab print jobs from validated plain .gcode, using Bambu LAN FTPS/MQTT handoffs.

    1.1k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check: warnings

Categories

Questions about Run Local Model

What does Run Local Model do?

Details behind AGENTS.md's 'Starting a model on this computer' flow: reading fleet models, sizing a start, choosing file and engine, picking a model for an engine with none, what fleet verify…. Run Local Model is an agent skill from autonomous-ai/openharness.md's 'Starting a model on this computer' flow: reading fleet models, sizing a start, choosing file and engine, picking a model for an engine with none, what fleet verify checks, and where to read when unsure.

When should I use Run Local Model?

Run Local Model fits situations like: tasks that involve Agent instruction files.

How do I install Run Local Model in Claude Code?

Run `npx skills add autonomous-ai/openharness --skill run-local-model -a claude-code`. Or copy the skill folder (store/agents/autonomous-grid/skills/run-local-model in autonomous-ai/openharness) into .claude/skills/run-local-model in your project. Claude Code loads it when a task matches its description.

How do I install Run Local Model in Codex?

Run `npx skills add autonomous-ai/openharness --skill run-local-model -a codex`. Or copy the skill folder (store/agents/autonomous-grid/skills/run-local-model in autonomous-ai/openharness) into .agents/skills/run-local-model in your project. Codex loads it when a task matches its description.

Can I use Run Local Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/openharness --skill run-local-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-local-model, .gemini/skills/run-local-model, .github/skills/run-local-model and .opencode/skills/run-local-model in your project.

What does Run Local Model need to run?

Going by SKILL.md and its folder, Run Local Model needs the command-line tools its instructions call (ollama).

Does Run Local Model access the network?

SKILL.md names 4 domains. As links in the text: recipes.vllm.ai, docs.ollama.com, lmstudio.ai and docs.sglang.io. This is read from the text; nothing was executed.

Is Run Local Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Local Model use?

Run Local Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run Local Model use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Run Local Model?

Skills that share tags, products or a category with Run Local Model: Using Agent Skills (addyosmani/agent-skills, 103k stars), Claude Reflect (BayramAnnakov/claude-reflect, 1.7k stars), Writing For Agents (bestofjs/bestofjs, 3.1k stars) and Neat-Freak Knowledge Closeout (KKKKhazix/khazix-skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Local Model?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/openharness, which has 1,149 GitHub stars. The repository holds 100 skills in this directory. The repository was last updated on October 8, 2026.

Source: autonomous-ai/openharness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.