Agent skill

Overmind Inference

by overmind-core in overmind-core/overmind

Inspect Overmind model deployments, serving metrics, worker state and live routing; test inference and activate an approved model.

AGPL-3.0Auto-check passedAI & LLM Engineering

Install Overmind Inference

skills CLI
$ npx skills add overmind-core/overmind --skill overmind-inference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install overmind-core/overmind overmind-inference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/overmind-core/overmind.git skills-src && mkdir -p .claude/skills && cp -r skills-src/overmind/skills/overmind-inference .claude/skills/overmind-inference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
overmind-inference
GitHub stars
597
Token cost
~753 tokens
SKILL.md length
369 words
Files
3 (incl. assets)
Skills in repo
20
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Inspect Overmind model deployments, serving metrics, worker state and live routing; test inference and activate an approved model.

  • Inference and serving operations
  • SKILL.md covers Inspect serving and Test or activate
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Benchmark selection

What it does

Overmind Inference is an agent skill from overmind-core/overmind. Inspect Overmind model deployments, serving metrics, worker state and live routing; test inference and activate an approved model. Use for Inference and serving operations, not training or benchmark selection.

Its SKILL.md is about 750 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including assets (for example `agents/openai.yaml`).

It sits in AI & LLM Engineering, covering Deployment. The repository describes itself as: The platform for continuously improving AI agents. The licence is AGPL-3.0.

When your agent uses it

  • Inference and serving operations
  • Benchmark selection

Example prompts

  • “/overmind-inference”

What it can do on your machine

Read from SKILL.md and the folder at commit 3dec73c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Overmind Inference loads about 753 tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 369 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~753

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from overmind-core/overmind at commit 3dec73c, republished under its AGPL-3.0 licence (© overmind-core). 369 words, ~753 tokens.

Download SKILL.mdSave it as .claude/skills/overmind-inference/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
overmind-inference
description
Inspect Overmind model deployments, serving metrics, worker state and live routing; test inference and activate an approved model. Use for Inference and serving operations, not training or benchmark selection.

Overmind Inference

Start with list_projects and choose the intended accessible project. For an account connection, pass its project_id on every project tool and resource URI query; follow returned links. Project API keys retain their narrower access.

Determine whether a model is ready, warm, selected and actually used. These are separate states. Use the chosen MCP project and returned deployment identities; do not invent model IDs or infer a live alias from training success.

Inspect serving

Read overmind://deployments/{deployment} and the relevant capability resource. For metrics, the deployment resource accepts period=1h|24h|7d|30d|all and source=application|all. State the selected filters; all-time/all-traffic is the default. Use application traffic when determining whether the user's application is connected.

Separate deployment readiness, current worker warmth, capability routing and successful application traffic. Worker measurements can be unavailable: null counts are unknown, not zero. Reading metrics does not wake a model. Historical traffic does not prove the worker is currently warm.

Describe failures and latency percentiles with their period and traffic source. Do not let platform evaluations or smoke calls stand in for application usage.

Show full SKILL.md (195 more words)Show less

Test or activate

Use run_inference only for a requested test on a ready deployment, with the user's supplied messages and output budget. Preserve explicit max_tokens; omission uses the production default. Report finish reason and truncation. An oversized context request needs a deliberate input/budget change, not silent clipping. A test may incur inference spend.

For an authorized rollout, prefer ship-model when capability, deployment and fine-tuning job are known. Otherwise inspect readiness, then use set_active_model for the approved capability/deployment. Poll the returned model_activation job. The previous alias remains selected until verification succeeds; a requested switch is not a completed switch.

Use retry_deployment only for a failed or deleted deployment. Clearing routing is a separate requested action; omitting deployment from set_active_model clears it and cancels a pending switch. Do not change the benchmark to activate serving.

For a repository rollout, get_model_swap_prompt supplies a local handoff. Apply code changes only within the requested scope and verify final capability and deployment state. Successful application API-key calls establish connection independently of a completed activation.

Report readiness, worker state, selected alias and application evidence separately. Open inference under the project's Console base, preserving projectId, when the user wants the visual serving view.

© overmind-core, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (assets) in overmind/skills/overmind-inference of overmind-core/overmind.

  • SKILL.md
  • agents/openai.yaml
  • assets/icon.png

Open the folder on GitHubat commit 3dec73c

Compare with similar skills

Overmind Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Overmind Inference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Overmind Inference this skillovermind-core/overmind597—~753Automated safety check: PassAGPL-3.0
Dynamo Interconnect CheckNVIDIA/skills3.5k1 repos~1.6kAutomated safety check: PassApache-2.0
Cohere Migration Deep Divejeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMIT
Claude Cost Optimizationmajiayu000/claude-skill-registry6661 repos~3.1kAutomated safety check: PassMIT
Anth Prod Checklistjeremylongshore/tons-of-skills-marketplace2.8k—~1.8kAutomated safety check: PassMIT
Coreweave Local Dev Loopjeremylongshore/tons-of-skills-marketplace2.8k—~927Automated safety check: NotesMIT

Similar skills

  • Official

    Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink.

    3.5k GitHub starsUsed in 1 repo~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Cohere Migration Deep Dive

    jeremylongshore/tons-of-skills-marketplace

    Migrate an application to or from Cohere with a provider adapter, parallel embedding index, quality evaluation, canary traffic, and rollback.

    2.8k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Claude Cost Optimization

    majiayu000/claude-skill-registry

    Comprehensive cost tracking and optimization for production Claude deployments.

    666 GitHub starsUsed in 1 repo~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Anth Prod Checklist

    jeremylongshore/tons-of-skills-marketplace

    Execute production deployment checklist for Claude API integrations.

    2.8k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Coreweave Local Dev Loop

    jeremylongshore/tons-of-skills-marketplace

    Set up local development workflow for CoreWeave GPU deployments.

    2.8k GitHub stars~927 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Tensorflow Savedmodel Creator

    jeremylongshore/tons-of-skills-marketplace

    Create tensorflow savedmodel creator operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~593 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from overmind-core/overmind

All 20 skills in this repo
  • API Endpoints

    overmind-core/overmind

    End-to-end workflow for adding or changing a backend API endpoint — which module the serializer and view belong in, URL registration, OpenAPI client regeneration, and typed consumption from the…

    597 GitHub stars~830 tokensUpdated today
    Auto-check: notes
  • Finetuning Model Onboarding

    overmind-core/overmind

    Rules for adding a new model or model family to the finetuning pipeline, or changing finetuning behavior for an existing one — engine-agnostic customization via family hooks instead of if/else in…

    597 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Frontend Design

    overmind-core/overmind

    Overmind Console design system — semantic tokens, shared primitives, geometry and icons, the border-contrast floor, the duplicated table implementations, and the verification scripts.

    597 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • MCP

    overmind-core/overmind

    End-to-end workflow for adding or changing Overmind MCP tools, resources, prompts, authentication, or result contracts — server layers, catalog registration, MCP-impact classification, and required…

    597 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • PR Etiquette

    overmind-core/overmind

    How to open a complete pull request on overmind-core/overmind — the CI gates, the cross-cutting surfaces a change must carry with it (MCP, blast radius, the docs repo), gh pr edit being broken here…

    597 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Seed Demo Data

    overmind-core/overmind

    Run or modify the seeddemo management command (the one-project Support Copilot demo) without breaking the beat-safety invariants that keep celery workers from re-driving seeded rows.

    597 GitHub stars~973 tokensUpdated today
    Auto-check passed

Questions about Overmind Inference

What does Overmind Inference do?

Inspect Overmind model deployments, serving metrics, worker state and live routing; test inference and activate an approved model. Overmind Inference is an agent skill from overmind-core/overmind. Inspect Overmind model deployments, serving metrics, worker state and live routing; test inference and activate an approved model.

When should I use Overmind Inference?

Overmind Inference fits situations like: inference and serving operations; benchmark selection.

How do I install Overmind Inference in Claude Code?

Run `npx skills add overmind-core/overmind --skill overmind-inference -a claude-code`. Or copy the skill folder (overmind/skills/overmind-inference in overmind-core/overmind) into .claude/skills/overmind-inference in your project. Claude Code loads it when a task matches its description.

How do I install Overmind Inference in Codex?

Run `npx skills add overmind-core/overmind --skill overmind-inference -a codex`. Or copy the skill folder (overmind/skills/overmind-inference in overmind-core/overmind) into .agents/skills/overmind-inference in your project. Codex loads it when a task matches its description.

Can I use Overmind Inference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add overmind-core/overmind --skill overmind-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/overmind-inference, .gemini/skills/overmind-inference, .github/skills/overmind-inference and .opencode/skills/overmind-inference in your project.

What does Overmind Inference need to run?

SKILL.md names no scripts, command-line tools or credentials: Overmind Inference is instructions for the agent only.

Does Overmind Inference access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Overmind Inference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Overmind Inference use?

Overmind Inference is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Overmind Inference use?

About 753 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Overmind Inference?

Skills that share tags, products or a category with Overmind Inference: Dynamo Interconnect Check (NVIDIA/skills, 3.5k stars), Cohere Migration Deep Dive (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Claude Cost Optimization (majiayu000/claude-skill-registry, 666 stars) and Anth Prod Checklist (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Overmind Inference?

overmind-core (a GitHub organization) maintains it in overmind-core/overmind, which has 597 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 8, 2026.

Source: overmind-core/overmind on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.