Agent skill

Run Tests

by overmind-core in overmind-core/overmind

How to run the platform test suites without wasting minutes — log-to-file pattern, when to skip tests entirely.

AGPL-3.0Auto-check passedTesting & QA

Install Run Tests

skills CLI
$ npx skills add overmind-core/overmind --skill run-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install overmind-core/overmind run-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/overmind-core/overmind.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/run-tests .claude/skills/run-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-tests
GitHub stars
544
Token cost
~914 tokens
SKILL.md length
436 words
Files
1
Skills in repo
20
Repo updated
First seen
Licence
AGPL-3.0

At a glance

How to run the platform test suites without wasting minutes — log-to-file pattern, when to skip tests entirely.

  • Works in 2 steps: Decide whether to run at all → Run once, tee to a log, grep the log
  • Tasks that involve Test generation
  • SKILL.md covers 1. Decide whether to run at all, 2. Run once, tee to a log,…, Journeys and Where a test belongs, plus 1 more section
  • Calls uv, bun and make; needs CURSOR_API_KEY

What it does

Run Tests is an agent skill from overmind-core/overmind. How to run the platform test suites without wasting minutes — log-to-file pattern, when to skip tests entirely. Use before running pytest or the frontend test suite.

Its SKILL.md is about 910 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test generation and Unit testing. It works with pytest and Redis. The repository describes itself as: The platform for continuously improving AI agents. The licence is AGPL-3.0.

When your agent uses it

  • Tasks that involve Test generation
  • Tasks that involve Unit testing

Example prompts

  • “/run-tests”

Requirements

  • Docker
  • A credential in CURSOR_API_KEY

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Decide whether to run at all
  2. Run once, tee to a log, grep the log

What it can do on your machine

Read from SKILL.md and the folder at commit 2c65378. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • bun
    • make
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CURSOR_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Tests loads about 914 tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 436 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~914

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from overmind-core/overmind at commit 2c65378, republished under its AGPL-3.0 licence (© overmind-core). 436 words, ~914 tokens.

Download SKILL.mdSave it as .claude/skills/run-tests/SKILL.md (or your agent's skills folder).
name
run-tests
description
How to run the platform test suites without wasting minutes — log-to-file pattern, when to skip tests entirely. Use before running pytest or the frontend test suite.

Running tests on this repo

Test runs here are expensive (backend suite ~90s+, frontend ~40s). Two rules:

1. Decide whether to run at all

For simple, self-evident fixes (renames, copy changes, small mechanical edits): do not run the suite. Typecheck + lint is the bar:

bash
cd frontend && bun run typecheck && bun run lint   # frontend
uv run ruff check .                                # backend

Say plainly that tests weren't run. Reserve full runs for changes whose behavior you can't reason through.

2. Run once, tee to a log, grep the log

Never re-run a suite just to filter its output differently:

bash
uv run pytest tests/ 2>&1 | tee "$SCRATCHPAD/pytest.log"
grep -E "FAILED|ERROR|passed|failed" "$SCRATCHPAD/pytest.log"

($SCRATCHPAD = the session scratchpad directory.) Re-run only after actually changing code, scoped to the failing files/tests:

bash
uv run pytest tests/test_foo.py::test_bar 2>&1 | tee "$SCRATCHPAD/pytest-rerun.log"

Journeys

make test-journeys runs tests/journeys/ against the live stack. It needs the compose postgres and redis services, and uses Redis DB 15. Each journey writes a run record to tests/journeys/.runs/<timestamp>/; read it for the LLM requests, background task failures and the error. make test never collects journeys: they need TEST_REDIS_URL and a serial run. The journeys take about six minutes, and CI runs them as their own job beside test; make test-journeys test_args="-k <name>" runs one. test_kit_*.py files run one journey per vendor (connectors, workshop LLM engines, GPU training outcomes): a new vendor joins its kit's parametrize list.

Where a test belongs

  • Journey (tests/journeys/): a customer promise across CLI, MCP, REST, SDK, worker and Console on the live stack. tests/journeys/console/ builds the self-hosted Console (bun run build, ~4 s) and drives it with Playwright on its own thread; uv run playwright install chromium-headless-shell once per machine.
  • Contract kit (tests/test_*_contract.py, test_otlp_dialects.py): one table over every vendor or dialect, the real client talking to a network fake.
  • Unit test: pure logic with many edge cases, through a public entry point.
Show full SKILL.md (161 more words)Show less

The unit suite runs behind the same outbound guard as the journeys: sockets reach loopback only, and every HTTP call must be answered by tests/fakes (fake_llm is autouse; fake_modal, sft, serving, clerk, stripe_api, scripted(host), slept on request). FakeLLM serves chat, streaming (stream_rounds), the OpenRouter catalog (catalog_payload, limits, prices), and Jev decisions (decide; the Jev path needs Redis, so point CACHES at TEST_REDIS_URL). The Cursor SDK dials out from a bridge process the guard cannot see, so CURSOR_API_KEY is removed unless a test asks for the cursor fake.

A known product defect gets a pytest.mark.xfail(strict=True, reason=...) control that asserts the correct behaviour; it turns red when the fix lands, and the marker goes.

Gotchas

  • Backend tests need the compose postgres service on localhost:5432 (TEST_POSTGRES_HOST/TEST_POSTGRES_PORT override it).
  • uv add/uv remove resync the venv WITHOUT dev/test groups — pytest vanishes. Restore with uv sync --group dev --group test.
  • Celery-dependent behavior needs docker compose restart of the worker to pick up backend changes; API containers hot-reload.

© overmind-core, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/run-tests of overmind-core/overmind.

Open the folder on GitHubat commit 2c65378

Compare with similar skills

Run Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Tests this skillovermind-core/overmind544—~914Automated safety check: PassAGPL-3.0
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence
Test Coverage Reviewareed1192/finance-news-aggregator149—~2.6kAutomated safety check: PassMIT
Write Test389ds/389-ds-base294—~1.8kAutomated safety check: PassCustom licence
MoAI TDD Workflowmodu-ai/moai-adk1.2k—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Coverage Review

    areed1192/finance-news-aggregator

    Audit, plan, write, and verify unit tests for Python projects using pytest.

    149 GitHub stars~2.6k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Write Test

    389ds/389-ds-base

    Add or extend a pytest integration test for 389 Directory Server under dirsrvtests/.

    294 GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • MoAI TDD Workflow

    modu-ai/moai-adk

    Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.

    1.2k GitHub stars~3.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Testing

    meleantonio/ChernyCode

    Testing conventions using pytest. An agent skill from meleantonio/ChernyCode.

    516 GitHub stars~408 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from overmind-core/overmind

All 20 skills in this repo
  • API Endpoints

    overmind-core/overmind

    End-to-end workflow for adding or changing a backend API endpoint — which module the serializer and view belong in, URL registration, OpenAPI client regeneration, and typed consumption from the…

    544 GitHub stars~830 tokensUpdated yesterday
    Auto-check: notes
  • Finetuning Model Onboarding

    overmind-core/overmind

    Rules for adding a new model or model family to the finetuning pipeline, or changing finetuning behavior for an existing one — engine-agnostic customization via family hooks instead of if/else in…

    544 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Frontend Design

    overmind-core/overmind

    Overmind Console design system — semantic tokens, shared primitives, geometry and icons, the border-contrast floor, the duplicated table implementations, and the verification scripts.

    544 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • MCP

    overmind-core/overmind

    End-to-end workflow for adding or changing Overmind MCP tools, resources, prompts, authentication, or result contracts — server layers, catalog registration, MCP-impact classification, and required…

    544 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • PR Etiquette

    overmind-core/overmind

    How to open a complete pull request on overmind-core/overmind — the CI gates, the cross-cutting surfaces a change must carry with it (MCP, blast radius, the docs repo), gh pr edit being broken here…

    544 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Seed Demo Data

    overmind-core/overmind

    Run or modify the seeddemo management command (the one-project Support Copilot demo) without breaking the beat-safety invariants that keep celery workers from re-driving seeded rows.

    544 GitHub stars~973 tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Run Tests

What does Run Tests do?

How to run the platform test suites without wasting minutes — log-to-file pattern, when to skip tests entirely. Run Tests is an agent skill from overmind-core/overmind. How to run the platform test suites without wasting minutes — log-to-file pattern, when to skip tests entirely.

When should I use Run Tests?

Run Tests fits situations like: tasks that involve Test generation; tasks that involve Unit testing.

How do I install Run Tests in Claude Code?

Run `npx skills add overmind-core/overmind --skill run-tests -a claude-code`. Or copy the skill folder (.agents/skills/run-tests in overmind-core/overmind) into .claude/skills/run-tests in your project. Claude Code loads it when a task matches its description.

How do I install Run Tests in Codex?

Run `npx skills add overmind-core/overmind --skill run-tests -a codex`. Or copy the skill folder (.agents/skills/run-tests in overmind-core/overmind) into .agents/skills/run-tests in your project. Codex loads it when a task matches its description.

Can I use Run Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add overmind-core/overmind --skill run-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-tests, .gemini/skills/run-tests, .github/skills/run-tests and .opencode/skills/run-tests in your project.

What does Run Tests need to run?

Going by SKILL.md and its folder, Run Tests needs the command-line tools its instructions call (uv, bun, make and docker) and credentials named CURSOR_API_KEY. Our summary lists: Docker; A credential in CURSOR_API_KEY.

Does Run Tests access the network?

SKILL.md contains no URLs. Its commands use uv and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Run Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Tests use?

Run Tests is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run Tests use?

About 914 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Run Tests?

Skills that share tags, products or a category with Run Tests: Adk Verify Snippets (google/adk-python, 22k stars), Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars), Test Coverage Review (areed1192/finance-news-aggregator, 149 stars) and Write Test (389ds/389-ds-base, 294 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Tests?

overmind-core (a GitHub organization) maintains it in overmind-core/overmind, which has 544 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 6, 2026.

Source: overmind-core/overmind on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.