Agent skill

Review Docs

by Kaggle in Kaggle/kaggle-environments

Audit a game environment's README.md and AGENTS.md against its engine implementation.

Apache-2.0Auto-check passedAgent Workflows

Install Review Docs

skills CLI
$ npx skills add Kaggle/kaggle-environments --skill review-docs -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kaggle/kaggle-environments review-docs --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/review-docs .claude/skills/review-docs && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review-docs
GitHub stars
454
Token cost
~2.7k tokens
SKILL.md length
1,409 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit a game environment's README.md and AGENTS.md against its engine implementation.

  • The user asks to review
  • SKILL.md covers Core rule, Method: simulate, don't reason, Checklist and Style, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Fact-check environment docs

What it does

Review Docs is an agent skill from Kaggle/kaggle-environments. Audit a game environment's README.md and AGENTS.md against its engine implementation. Use when the user asks to "review", "audit", "check", or "fact-check" environment docs, when players report doc/engine discrepancies, or before a competition launch. Also use proactively after changing an environment's engine (kaggleenvironments/envs/<game/<game.py) when a sibling README.md or AGENTS.md exists — rebalances and mechanic changes routinely leave those docs stale. The engine is always the source of truth — docs get…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Fact-checking and source verification, Agent instruction files and Technical documentation. It works with Kaggle. The licence is Apache-2.0.

When your agent uses it

  • The user asks to review
  • Fact-check environment docs
  • Players report doc/engine discrepancies
  • Before a competition launch

Example prompts

  • “review”
  • “fact-check”
  • “/review-docs”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bac60f5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Review Docs loads about 2.7k tokens when it runs. Until then it costs about 139 tokens; SKILL.md has 1,409 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~139
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Kaggle/kaggle-environments at commit bac60f5, republished under its Apache-2.0 licence (© Kaggle). 1,409 words, ~2,708 tokens.

Download SKILL.mdSave it as .claude/skills/review-docs/SKILL.md (or your agent's skills folder).
name
review-docs
description
Audit a game environment's README.md and AGENTS.md against its engine implementation. Use when the user asks to "review", "audit", "check", or "fact-check" environment docs, when players report doc/engine discrepancies, or before a competition launch. Also use proactively after changing an environment's engine (`kaggle_environments/envs/<game>/<game>.py`) when a sibling README.md or AGENTS.md exists — rebalances and mechanic changes routinely leave those docs stale. The engine is always the source of truth — docs get fixed, not the engine.

Review Environment Docs Against the Engine

Audit kaggle_environments/envs/<game>/README.md and AGENTS.md for claims that disagree with the engine.

Core rule

The engine is the source of truth. Fix the docs, not the engine. Agents play against the implementation, and changing engine behavior late invalidates existing submissions and replays.

Exception: if a mechanic is clearly a bug rather than an undocumented quirk, flag it and ask. Don't silently enshrine a bug as intended.

Method: simulate, don't reason

Never verify a claim by reading code and reasoning about it. Execute the engine and observe. Cheapest entry point first:

python
# 1. Internal helpers — isolates one rule
from kaggle_environments.envs.<game> import <game> as G
G._new_<entity>(...)      # constructor: initial field values
G._apply_<action>(...)    # one action
G._<periodic_hook>(...)   # end-of-turn / round / phase transition

# 2. Full env — catches validation the helpers skip
from kaggle_environments import make
env = make("<game>", configuration={...}); env.reset(2); env.step([a0, a1])

# 3. OpenSpiel games
import pyspiel; state = pyspiel.load_game("<name>").new_initial_state()

Step through the entire lifecycle — creation to terminal state — recording observable state at every step, not just the endpoints. Most doc bugs live in the middle: a value that plateaus early, a counter starting at the wrong number, a transition firing a step sooner than the prose implies.

Gotchas:

  • Don't name scratch files after stdlib modules (numbers.py breaks the import chain).
  • Helpers may replace an object rather than mutate it (one entity type becoming another). Re-read from the container after each call.
  • Verify action- and transaction-related claims through make() too.

Checklist

Skip sections that don't apply. Each item is a real bug class found in a prior audit.

Constants, tables, and formulas
  • Every value in a doc table matches the engine's constants. Check every row, not a sample.
  • Derived columns (rates, per-unit costs) use one consistent formula, and the doc says which. Mixed formulas across rows is the tell.
  • "Max" values are reachable. A cap needing an optional booster needs that qualifier; a cap unreachable in a variant lacking the booster is wrong.
  • Column headers mean what the values show. If they disagree, ask the user whether to change the number or the header.
  • Formulas match the code symbol for symbol — exponents, log bases, clamps, order of operations. Recompute any worked example.
State initialization and lifecycle
  • Initial value of every field in each constructor. A counter starting at 1 instead of 0 silently removes a grace period the prose implies.
  • "Ongoing" / "indefinite" / "repeating" claims: find the cap (a count-vs-max guard, a lifespan field being armed). If one exists, the doc must say so.
  • Exact step an entity or episode ends, and what stays usable on the way out.
  • Ordering inside periodic hooks: which mutations run before which checks. That ordering is often what makes a rule surprising.
  • Cooldowns: when the counter ticks relative to the action, whether a failed attempt consumes it, turns-between vs. turns-until.
  • Rates that ramp or decay: recompute at start, middle, and end. Watch for rounding or a clamp that makes the stated endpoint arrive early.
  • Entities that change type in place (transforming, upgrading, being captured) — what carries over, what resets.
Actions and preconditions
  • Walk each documented action against its handler. Every early return/continue is an undocumented precondition.
  • Partial success: an action that works in one state but no-ops in a neighbouring one needs the qualifier.
  • No-op semantics: does a rejected action still consume the resource, turn, or cooldown?
  • Legality rules for movement, placement, targeting — and whether spawn/setup logic obeys the same rules the player does.
  • Malformed input: wrong arity, wrong type, out-of-range. Silently dropping one move from a submitted list is a rule players need.
  • Per-turn action limits and what happens to the excess.
Turn pipeline and interaction resolution
  • If the doc lists a numbered turn order, walk the engine top to bottom and confirm each stage's position. Late-inserted stages (cleanup, spawning, expiry) get omitted.
  • Where an entity is mutated relative to where it's read. Spawned after the action phase → unusable until next turn; removed before it → never usable.
  • Contested outcomes (combat, collision, simultaneous claims): work ties, three-way contests, and zero-survivor cases by hand. Docs describe only the two-party win case.
  • Continuous vs. endpoint-sampled detection (swept paths vs. position checks). Changes which interactions are possible; state it outright.
  • Precedence when several removal conditions fire on one entity in the same step.
Procedural generation and randomness
  • Documented ranges match the generator's bounds, including any retry cap that can silently under-deliver.
  • Guarantees ("at least one of X", "always symmetric") are enforced by code, not merely likely. Find the loop that makes them true, or drop "guaranteed".
  • Distribution claims ("skewed low", "weighted toward") imply specific sampling — verify rather than restate.
  • What the seed controls, and whether it's scrubbed from the observation before agents see it.
Hidden state and partial observability
  • Fields agents are told to use for prediction contain what's claimed, and stay valid as the episode progresses.
  • Anything the doc says is derivable, derive it: recompute from the observation alone and compare to engine state.
  • Information agents should not have is genuinely absent — check what the agent receives, not the shared state object. Engines routinely keep global state on state[0].observation and strip it per player. Run an agent that dumps its own observation keys.
  • If visibility is asymmetric, confirm the players' observations actually differ and each sees only its share.
  • Per-category persistence. In a "visible now / remembered later" table every row is a separate claim, and "remembered" may mean last-known-value rather than live.
  • Sentinels for unknown data (-1, None, absent key) are documented and mean unknown, not a legitimate zero.
Show full SKILL.md (552 more words)Show less
Termination and scoring
  • Every termination condition, not just the step limit. Elimination, stalemate, and early-exit branches get missed.
  • The step the episode actually ends on. Off-by-one against episodeSteps is common; run to completion.
  • The reward the interpreter assigns, including ties and all-players-eliminated. Does the doc's "winner" language match a win/loss constant or a raw score?
  • Rewards that change meaning mid-episode (running score during play, win/loss constant at the end) need both forms documented.
  • Multi-stage tiebreakers: each stage exists, in that order, with the stated final fallback.
  • What counts toward the final score and what's excluded — value in intermediate state often doesn't count.
Observation and action format
  • Every documented field exists, with that type and meaning.
  • Field comments describing when something is set: verify the trigger. "Set after X" is often wrong when the engine sets it unconditionally.
  • Boolean vs. counter. A boolean means it doesn't accumulate — say so if a player might expect stockpiling.
  • Anything the docs imply lives in the state structure but doesn't. Agents will search and find nothing; document the coordinates or accessor instead.
  • Documented action strings and argument shapes match what the handler parses.
Configuration
  • Doc defaults match the .json specification.
  • Every documented knob is actually read. Diff the doc table against the .json keys, then grep each in the engine. A knob ignored in favour of a module constant is worse than undocumented — players tune it and see no effect.
  • Claims holding only at default config are labelled as such.
  • Derived figures quoted in prose recompute from the stated defaults.
Resource and economy systems (if present)
  • Which items/actions are permitted on each side of a transaction. Check membership tests (x in SOME_LIST) separately from explicit whitelists — an unfiltered membership test permits more than the author may have intended.
  • Calibration constants: confirm the stated derivation reproduces them. If not, they're stale from a rebalance or the explanation is wrong.
  • Any horizon or reference quantity differing from the episode length needs a reason in the doc.
Cross-file consistency
  • README and AGENTS.md agree. When they conflict, both may still be wrong — check the engine.
  • Beginner/advanced or base/arena variants have their own engines. A fix in one usually applies to the sibling, but verify against that sibling's engine — constants often differ.
  • Module-level comments and docstrings in the engine are docs too, and drift.

Style

Docs are declarative rulebooks, not strategy guides.

  • State the rule; don't advise. Say what the mechanic does, not what the player should do about it.
  • No "you should", "plan on", "budget for", "make sure", "worth noting".
  • No bold lead-ins to make a rule feel urgent.
  • Leave pre-existing prose alone unless it's factually wrong. Rewriting a doc's voice is out of scope; if a paragraph is hard to read, a paragraph break beats an edit.

Reporting

Group findings as confirmed / overstated / rejected. For each: file:line, what the engine does, how you verified it. Separate doc bugs from engine bugs — engine bugs need the user's decision, not your fix.

Surface explicitly:

  • Bugs the reporter missed. A user-filed list is a starting point, not a scope.
  • Judgment calls, where the engine's behavior looks unintentional. Documenting it is the default, but say so plainly and give the alternative fix.

Before finishing: re-verify every number you wrote with one audit script asserting doc value == engine value, and run the env's test suite.

© Kaggle, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/review-docs of Kaggle/kaggle-environments.

Open the folder on GitHubat commit bac60f5

Compare with similar skills

Review Docs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Review Docs compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Review Docs this skillKaggle/kaggle-environments454—~2.7kAutomated safety check: PassApache-2.0
Neat-Freak Knowledge CloseoutKKKKhazix/khazix-skills21k—~1.9kAutomated safety check: PassMIT
Dsh Web Documentationzhu1090093659/dsh-web8.5k—~479Automated safety check: PassApache-2.0
Newprojectscunning1975/MixtapeTools4731 repos~850Automated safety check: PassNone
Sync Public DocsCaldis/react-zmage946—~3.3kAutomated safety check: PassMIT
Documentationsamchon/typia5.9k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Neat-Freak Knowledge Closeout

    KKKKhazix/khazix-skills

    Brings project docs, agent rule files, authorized memory and leftover workspace files back in line with what the code and runtime actually do at the end of a work session.

    21k GitHub stars~1.9k tokensUpdated 8 days ago
    Agent WorkflowsAuto-check passed
  • Dsh Web Documentation

    zhu1090093659/dsh-web

    A skill your agent uses when adding or editing dsh-web README files, docs, AGENTS.md instructions, user-facing configuration text, or bilingual documentation pairs.

    8.5k GitHub stars~479 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Newproject

    scunning1975/MixtapeTools

    Scaffold a new research project with standard directory structure, CLAUDE.md template, and documented README.

    473 GitHub starsUsed in 1 repo~850 tokens
    Agent WorkflowsAuto-check passed
  • Sync Public Docs

    Caldis/react-zmage

    A skill your agent uses when modifying public API in packages/core (types/global.ts, types/default.ts, index.ts, or package.json exports field), adding/renaming/removing props, changing default…

    946 GitHub stars~3.3k tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed
  • Documentation

    samchon/typia

    Defines README, website-guide, and agent-instruction structure, audience, prose formatting, and voice for typia.

    5.9k GitHub stars~1.1k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Technical Documentation

    openclaw/openclaw

    Build and review high-quality technical docs as well as agent instruction files in your repository.

    392k GitHub stars~1.5k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from Kaggle/kaggle-environments

  • Create Harness

    Kaggle/kaggle-environments

    Create or update an LLM harness that lets a language model play a kaggle-environments game.

    454 GitHub stars~8.9k tokensUpdated today
    Auto-check passed
  • Review Harness

    Kaggle/kaggle-environments

    Review an existing LLM harness for correctness and gameplay-impacting bugs.

    454 GitHub stars~24k tokensUpdated today
    Auto-check passed
  • Run Ablation

    Kaggle/kaggle-environments

    Run a prompt-ablation study on a kaggle-environments game's LLM harness.

    454 GitHub stars~6.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Review Docs

What does Review Docs do?

Audit a game environment's README.md and AGENTS.md against its engine implementation. Review Docs is an agent skill from Kaggle/kaggle-environments.md against its engine implementation.

When should I use Review Docs?

Review Docs fits situations like: the user asks to review; fact-check environment docs; players report doc/engine discrepancies; before a competition launch.

How do I install Review Docs in Claude Code?

Run `npx skills add Kaggle/kaggle-environments --skill review-docs -a claude-code`. Or copy the skill folder (.agents/skills/review-docs in Kaggle/kaggle-environments) into .claude/skills/review-docs in your project. Claude Code loads it when a task matches its description.

How do I install Review Docs in Codex?

Run `npx skills add Kaggle/kaggle-environments --skill review-docs -a codex`. Or copy the skill folder (.agents/skills/review-docs in Kaggle/kaggle-environments) into .agents/skills/review-docs in your project. Codex loads it when a task matches its description.

Can I use Review Docs in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kaggle/kaggle-environments --skill review-docs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-docs, .gemini/skills/review-docs, .github/skills/review-docs and .opencode/skills/review-docs in your project.

What does Review Docs need to run?

SKILL.md names no scripts, command-line tools or credentials: Review Docs is instructions for the agent only. Our summary lists: Python 3.

Does Review Docs access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Review Docs safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Review Docs use?

Review Docs is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Review Docs use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Review Docs?

Skills that share tags, products or a category with Review Docs: Neat-Freak Knowledge Closeout (KKKKhazix/khazix-skills, 21k stars), Dsh Web Documentation (zhu1090093659/dsh-web, 8.5k stars), Newproject (scunning1975/MixtapeTools, 473 stars) and Sync Public Docs (Caldis/react-zmage, 946 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Review Docs?

Kaggle (a GitHub organization) maintains it in Kaggle/kaggle-environments, which has 454 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 8, 2026.

Source: Kaggle/kaggle-environments on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.