Agent skill

Testing Validation

by AbdelStark in AbdelStark/worldforge

A skill your agent uses when selecting, running, or fixing WorldForge validation: pytest, coverage, ruff, generated provider docs, MkDocs strict build, package contract, CI failures, and release…

MITAuto-check passedTesting & QA

Install Testing Validation

skills CLI
$ npx skills add AbdelStark/worldforge --skill testing-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AbdelStark/worldforge testing-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AbdelStark/worldforge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/testing-validation .claude/skills/testing-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-validation
GitHub stars
108
Token cost
~871 tokens
SKILL.md length
274 words
Files
2
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when selecting, running, or fixing WorldForge validation: pytest, coverage, ruff, generated provider docs, MkDocs strict build, package contract, CI failures, and release…

  • Fixing WorldForge validation: pytest
  • SKILL.md covers Choose The Gate, Standard Commands, Rules and Definition Of Done, plus 1 more section
  • Calls uv, bash and uvx
  • Generated provider docs

What it does

Testing Validation is an agent skill from AbdelStark/worldforge. Use when selecting, running, or fixing WorldForge validation: pytest, coverage, ruff, generated provider docs, MkDocs strict build, package contract, CI failures, and release gates. Produces the smallest credible command set first, then escalates to full validation when public behavior changes.

Its SKILL.md is about 870 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Testing & QA, covering Linting and formatting, Unit testing and Failing and flaky tests. It works with pytest and Ruff. The repository describes itself as: Harness framework to build world model based workflows for physical AI systems. The licence is MIT.

When your agent uses it

  • Fixing WorldForge validation: pytest
  • Generated provider docs
  • MkDocs strict build
  • Package contract

Example prompts

  • “/testing-validation”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 30b65da. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • bash
    • uvx
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, uvx and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing Validation loads about 871 tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 274 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~871

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AbdelStark/worldforge at commit 30b65da, republished under its MIT licence (© AbdelStark). 274 words, ~871 tokens.

Download SKILL.mdSave it as .claude/skills/testing-validation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
testing-validation
description
Use when selecting, running, or fixing WorldForge validation: pytest, coverage, ruff, generated provider docs, MkDocs strict build, package contract, CI failures, and release gates. Produces the smallest credible command set first, then escalates to full validation when public behavior changes.

Testing And Validation

Choose The Gate

ChangeMinimum useful validation
Python logicuv run pytest tests/test_target.py -q plus ruff
Provider behaviorprovider-focused pytest, fixtures, contract helper, provider-doc check
CLI help/outputtargeted CLI tests and help snapshots
Docs/provider catalogprovider-doc check and uv run mkdocs build --strict
Agent context/skillsskill quick_validate.py, symlink check, and git diff --check
Public API/package surfacefull public gate below
TUI/harnessfocused harness tests plus --extra harness coverage when relevant

Standard Commands

Focused gate:

bash
uv run ruff check src tests examples scripts
uv run ruff format --check src tests examples scripts
uv run pytest tests/test_target.py -q

Docs gate:

bash
uv run python scripts/generate_provider_docs.py --check
uv run mkdocs build --strict

Public behavior/package gate:

bash
uv lock --check
uv run ruff check src tests examples scripts
uv run ruff format --check src tests examples scripts
uv run python scripts/generate_provider_docs.py --check
uv run mkdocs build --strict
uv run pytest
uv run --extra harness pytest --cov=src/worldforge --cov-report=term-missing --cov-fail-under=90
bash scripts/test_package.sh
uv build --out-dir dist --clear --no-build-logs

Dependency audit for release work:

bash
tmp_req="$(mktemp requirements-audit.XXXXXX)"
uv export --frozen --all-groups --no-emit-project --no-hashes -o "$tmp_req" >/dev/null
uvx --from pip-audit pip-audit -r "$tmp_req" --no-deps --disable-pip --progress-spinner off
rm -f "$tmp_req"

Rules

  • Reproduce the failing command before broad edits.
  • Add regression tests for bug fixes and documented failure modes.
  • Keep src tests examples scripts in Ruff targets.
  • Keep --cov-fail-under=90; add tests instead of lowering it.
  • Do not replace deterministic tests with live-service requirements.
  • Report skipped gates with the concrete blocker.
  • Match the gate to the claim: a narrow passing test never proves a broad public-release or agentic-context quality claim.

Definition Of Done

  • The final report names exact commands, pass/fail status, and any unverified surfaces.
  • Validation covers the files actually changed and the public contract they affect.
  • Skill changes pass quick_validate.py for every edited skill and preserve .agents/skills plus .claude/skills symlinks.
  • Broad gates are escalated when behavior, packaging, docs navigation, or release evidence changes.

Sharp Edges

SymptomCauseFix
Provider docs check failsGenerated README/provider catalog staleRun generator without --check, inspect diff
Coverage barely failsNew branch lacks focused testsAdd direct failure-path tests instead of weakening gate
Package contract fails only in isolated venvMissing package data or import pathInspect pyproject.toml hatch settings and scripts/test_package.sh
MkDocs strict failsNav/SUMMARY/docs link driftSync mkdocs.yml, docs/src/SUMMARY.md, and links

© AbdelStark, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .codex/skills/testing-validation of AbdelStark/worldforge.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 30b65da

Compare with similar skills

Testing Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing Validation this skillAbdelStark/worldforge108—~871Automated safety check: PassMIT
Simple Modern Uvjlevy/simple-modern-uv301—~1.9kAutomated safety check: PassMIT
Run Testsartcc/freelingo158—~976Automated safety check: PassAGPL-3.0
Checkav1155/houndarr292—~366Automated safety check: PassAGPL-3.0
Validatejuliepy/AI-Engineer-from-scrach430—~500Automated safety check: PassNone
Gateguardana/guardana152—~776Automated safety check: PassApache-2.0

Similar skills

  • Simple Modern Uv

    jlevy/simple-modern-uv

    Start, selectively modernize, fully migrate, or update Python projects using simple-modern-uv practices: uv, ruff, BasedPyright, pytest, GitHub Actions CI, and tag-driven PyPI publishing.

    301 GitHub stars~1.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Run Tests

    artcc/freelingo

    A skill your agent uses when the user asks to run tests, ejecutar tests, lanzar tests, pytest, vitest, check types, typecheck, lint, or verify the codebase.

    158 GitHub stars~976 tokensUpdated today
    Testing & QAAuto-check passed
  • Check

    av1155/houndarr

    Run Houndarr's full quality gate (ruff lint, ruff format check, mypy, bandit, pytest) and report results in a single table.

    292 GitHub stars~366 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Validate

    juliepy/AI-Engineer-from-scrach

    Run the full quality gate (ruff + mypy + pytest + tsc + vitest) and report PASS/FAIL for each command.

    430 GitHub stars~500 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Gate

    guardana/guardana

    Run this project's verification — the full local CI mirror (ruff, mypy, import contract, pytest with PostgreSQL, coverage floors, dogfood, generated docs and site, the isolated example suites, the…

    152 GitHub stars~776 tokensUpdated today
    Testing & QAAuto-check passed
  • Python Helper

    shepherdjerred/monorepo

    Current Python development guidance for versions, uv and pip, packaging, typing, asyncio, pytest, Ruff, security, and runtime boundaries.

    112 GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed

More from AbdelStark/worldforge

  • Public Docs Release

    AbdelStark/worldforge

    A skill your agent uses for WorldForge README, docs, changelog, generated provider docs, MkDocs navigation, version/release metadata, public positioning, and release or publish readiness checks.

    108 GitHub stars~679 tokensUpdated 18 days ago
    Auto-check passed
  • Evaluation Benchmarking

    AbdelStark/worldforge

    A skill your agent uses for WorldForge evaluation suites, benchmark harness changes, benchmark input fixtures, budget gates, report rendering, metrics semantics, and any claims based on benchmark or…

    108 GitHub stars~732 tokensUpdated 18 days ago
    Auto-check passed
  • Optional Runtime Smokes

    AbdelStark/worldforge

    A skill your agent uses for LeWorldModel, GR00T, LeRobot, PushT robotics showcase, real-checkpoint smoke scripts, checkpoint building, and host-owned optional runtime dependencies.

    108 GitHub stars~855 tokensUpdated 18 days ago
    Auto-check passed
  • Persistence State

    AbdelStark/worldforge

    A skill your agent uses for WorldForge local state: run workspaces and artifacts under .worldforge/ (run manifests, evidence bundles, retention/prune), the WorldForge(statedir=...) directory, and…

    108 GitHub stars~921 tokensUpdated 18 days ago
    Auto-check passed
  • Provider Adapter Development

    AbdelStark/worldforge

    A skill your agent uses for WorldForge provider work: adding adapters, changing capability declarations, promoting scaffolds, debugging provider failures, updating provider catalog docs, or touching…

    108 GitHub stars~990 tokensUpdated 18 days ago
    Auto-check: notes
  • Tui Development

    AbdelStark/worldforge

    A skill your agent uses for the robotics showcase Textual UI: report panes, launch helpers, screenshots, visual tests, and changes under src/worldforge/harness/tui.py or robotics view/rendering…

    108 GitHub stars~792 tokensUpdated 18 days ago
    Auto-check passed

Works with

Categories

Questions about Testing Validation

What does Testing Validation do?

A skill your agent uses when selecting, running, or fixing WorldForge validation: pytest, coverage, ruff, generated provider docs, MkDocs strict build, package contract, CI failures, and release…. Testing Validation is an agent skill from AbdelStark/worldforge. Use when selecting, running, or fixing WorldForge validation: pytest, coverage, ruff, generated provider docs, MkDocs strict build, package contract, CI failures, and release gates.

When should I use Testing Validation?

Testing Validation fits situations like: fixing WorldForge validation: pytest; generated provider docs; mkDocs strict build; package contract.

How do I install Testing Validation in Claude Code?

Run `npx skills add AbdelStark/worldforge --skill testing-validation -a claude-code`. Or copy the skill folder (.codex/skills/testing-validation in AbdelStark/worldforge) into .claude/skills/testing-validation in your project. Claude Code loads it when a task matches its description.

How do I install Testing Validation in Codex?

Run `npx skills add AbdelStark/worldforge --skill testing-validation -a codex`. Or copy the skill folder (.codex/skills/testing-validation in AbdelStark/worldforge) into .agents/skills/testing-validation in your project. Codex loads it when a task matches its description.

Can I use Testing Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AbdelStark/worldforge --skill testing-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-validation, .gemini/skills/testing-validation, .github/skills/testing-validation and .opencode/skills/testing-validation in your project.

What does Testing Validation need to run?

Going by SKILL.md and its folder, Testing Validation needs the command-line tools its instructions call (uv, bash, uvx and git). Our summary lists: Python 3.

Does Testing Validation access the network?

SKILL.md contains no URLs. Its commands use uv, uvx and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Testing Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing Validation use?

Testing Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing Validation use?

About 871 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing Validation?

Skills that share tags, products or a category with Testing Validation: Simple Modern Uv (jlevy/simple-modern-uv, 301 stars), Run Tests (artcc/freelingo, 158 stars), Check (av1155/houndarr, 292 stars) and Validate (juliepy/AI-Engineer-from-scrach, 430 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing Validation?

AbdelStark (a GitHub user) maintains it in AbdelStark/worldforge, which has 108 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 19, 2026.

Source: AbdelStark/worldforge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.