Agent skill

Run Ablation

by Kaggle in Kaggle/kaggle-environments

Run a prompt-ablation study on a kaggle-environments game's LLM harness.

Apache-2.0Auto-check passed

Install Run Ablation

skills CLI
$ npx skills add Kaggle/kaggle-environments --skill run-ablation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kaggle/kaggle-environments run-ablation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/run-ablation .claude/skills/run-ablation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-ablation
GitHub stars
454
Token cost
~6.9k tokens
SKILL.md length
3,362 words
Files
3
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run a prompt-ablation study on a kaggle-environments game's LLM harness.

  • Works in 9 steps: Identify the env → Inventory existing variants → Read the production prompt → …
  • The user mentions ablation
  • SKILL.md covers Step 0: Identify the env, Step 1: Inventory existing…, Step 2: Read the production… and Step 3: ALWAYS propose new…, plus 8 more sections
  • Calls python and uv; needs MODEL_PROXY_KEY

What it does

Run Ablation is an agent skill from Kaggle/kaggle-environments. Run a prompt-ablation study on a kaggle-environments game's LLM harness. Use when the user mentions "ablation", "prompt sensitivity", "test prompt variants", "compare prompts", "ablate the prompt", "prompt rewrite study", or asks whether their prompt wording is doing real work. Bootstraps a promptvariants.py if missing, proposes new variants interactively, then runs a paired-seat tournament and reports the leaderboards.

Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files.

It works with Kaggle. The licence is Apache-2.0.

When your agent uses it

  • The user mentions ablation
  • Prompt sensitivity
  • Test prompt variants
  • Compare prompts

Example prompts

  • “s LLM harness. Use when the user mentions”
  • “prompt sensitivity”
  • “test prompt variants”
  • “/run-ablation”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Identify the env
  2. Inventory existing variants
  3. Read the production prompt
  4. ALWAYS propose new variants
  5. Write the variants
  6. Verify with ablation check
  7. Confirm the run plan
  8. Run the tournament
  9. Report findings

What it can do on your machine

Read from SKILL.md and the folder at commit 1c8acf1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MODEL_PROXY_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Ablation loads about 6.9k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 3,362 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Kaggle/kaggle-environments at commit 1c8acf1, republished under its Apache-2.0 licence (© Kaggle). 3,362 words, ~6,881 tokens.

Download SKILL.mdSave it as .claude/skills/run-ablation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
run-ablation
description
Run a prompt-ablation study on a kaggle-environments game's LLM harness. Use when the user mentions "ablation", "prompt sensitivity", "test prompt variants", "compare prompts", "ablate the prompt", "prompt rewrite study", or asks whether their prompt wording is doing real work. Bootstraps a prompt_variants.py if missing, proposes new variants interactively, then runs a paired-seat tournament and reports the leaderboards.

Run a Prompt Ablation Study

This skill drives the end-to-end prompt-sensitivity workflow for one game's harness: discover or bootstrap variants, propose additional ones in conversation with the user, run the paired-seat tournament, and report the findings. It is completely standalone — it does not depend on or modify the create-harness, review-harness, or create-environment skills.

Two CLI tools work together:

bash
# Verify baseline parity + variant rendering before spending any budget
python -m kaggle_environments.ablation check --env <env_name>

# Run the tournament. Auto-preflight, auto-abort on runaway crashes, and
# --redo-crashed for recovery are all on by default; see Step 7.
python -m kaggle_environments.ablation run --env <env_name> --models <csv> --games <N> --out <dir>

# Post-hoc statistical analysis (permutation test against the null variant's noise floor)
python -m kaggle_environments.ablation_analysis --csv <dir>/games.csv --baseline baseline --null null

The runner contract:

  • Every game's harness directory contributes a prompt_variants.py exposing VARIANTS: dict[str, GameHarness]. Each variant is a GameHarness (implements get_legal_moves / make_prompt / parse_response), so create_agent_fn(variant) works directly.
  • The baseline variant must be byte-identical to the production harness.py. The check subcommand enforces this against seeded observations.
  • The null variant must be a byte-identical second copy of baseline. The runner schedules independent cells for it; the difference between null-vs-baseline rankings is pure LLM-sampling noise, which calibrates the noise floor for the permutation test.
  • For each variant, the runner plays a round-robin among the M models with both players using that variant. Every {A, B} matchup is scheduled in seat-flipped pairs sharing one chance seed — the same instance is played twice with seats swapped. --games must be even.

Step 0: Identify the env

Confirm with the user which env they want to ablate. Acceptable forms: open_spiel_bargaining, werewolf, tictactoe, etc. The env name is what gets passed to kaggle_environments.make(env_name).

Resolve the harness directory:

  • open_spiel_<game> → kaggle_environments/envs/open_spiel_env/games/<game>/
  • <env> → kaggle_environments/envs/<env>/

Confirm harness.py exists at that path. If not, stop and tell the user: the production harness has to exist before you can ablate it. Suggest they run create-harness first.

Step 1: Inventory existing variants

Read prompt_variants.py if it exists at the harness directory. List the names found and one-line summaries:

Found prompt_variants.py with: baseline, compact, minimal, no_accept_preview, generic_names

If the file doesn't exist, you'll bootstrap it in Step 4. Tell the user you're going to:

No prompt_variants.py yet — I'll bootstrap one from harness.py.
The 'baseline' variant will be byte-identical to the production prompt;
I'll also propose a few ablations for your review.

Step 2: Read the production prompt

Open harness.py and read the make_prompt (or generate_prompt) function plus any prompt-template constants it uses. Note the prompt's structural features — these are what your variant proposals will target:

  • What rules / mechanics paragraphs are in there?
  • What helper text (legal-move list, current-state summary, payoff hints, "you would receive...", "you may agree")?
  • What domain-specific labels (item names, square notation, color names)?
  • Is there a goal/strategy nudge ("end the game with a higher reward than your opponent")?
  • What's the JSON schema?

You can't propose a good ablation without understanding the prompt's joints. Spend a turn here.

Step 3: ALWAYS propose new variants

This step is mandatory. Even if prompt_variants.py already has variants, propose new ones. Skipping this step defeats the purpose of the skill — the point is to force a "what am I actually testing?" beat before any LLM money is spent.

Generate 3–5 candidate variants with a one-line rationale each. Cover at least three of these generic axes:

  • Compact — same content as baseline, just tighter prose. Drops worked examples, repeated reminders, loss-threat sentences, flavor padding. Keeps every computed helper baseline provides — per-turn counters, accept-previews like "you would receive [...]", rendered complements in history rows, end-of-prompt constraint reminders. Tests whether the verbose prose padding is load-bearing while holding the helper surface constant.
  • Minimal — strips both prose and derived/helper info. Keeps mechanics (private-vs-public state, action semantics, reward formula, terminal conditions) but removes anything the model can compute itself from the basic state: per-turn counters (countable from history), accept-previews (subtractable from the last offer), complement-rendered history (subtractable from pool), constraint reminders. Tests whether the model can derive what baseline spoon-feeds.

Compact and Minimal are distinct ablation axes, not "more terse" vs "even more terse." If you're proposing both, make sure they actually test different things — Compact = "do we need the prose?", Minimal = "do we need the helpers?" The mechanics keep is non-negotiable in both (see Step 4).

  • No goal hint — removes any "your goal is to..." sentence. Tests whether the goal nudge is doing work.
  • Generic labels — swaps real names for letters (Book/Hat/Basketball → A/B/C, North/East/South/West → 0/1/2/3). Tests for semantic priors that bias decisions away from stated valuations.
  • No legal-move list — removes the enumeration of legal moves (if the prompt includes one). Tests whether the model can derive legality from rules.
  • Raw observation — uses the env's raw observation string instead of any reformatting the prompt does. Tests whether the prompt's restructuring is helping.
  • Drop a specific helper — pick one piece of computed help in the prompt (e.g. Bargaining's "you would receive [...]" accept preview, Chess's notation explainer, Go's territory hint) and remove it. Tests whether that helper is load-bearing.

Then propose 1–2 game-specific axes that you identify from reading the prompt in Step 2. These are the most interesting ablations because they're tailored to the actual prompt design choices the author made.

Each proposal should include:

  • A short name (snake_case, e.g. no_payoff_hint)
  • One sentence of rationale (what hypothesis it tests)
  • The concrete change vs. baseline (one or two bullet points)

Use the AskUserQuestion tool to surface the proposals, with one question per group of related variants, options for accept/skip/rewrite/replace. Concrete usage pattern:

AskUserQuestion({
  questions: [
    {
      question: "Which of these structural ablations should we include?",
      header: "Structural",
      multiSelect: true,
      options: [
        {label: "compact",       description: "Same info, terser — drops worked example + loss threat"},
        {label: "minimal",       description: "Strip rules/goal/helpers, just state + JSON schema"},
        {label: "no_goal_hint",  description: "Remove the 'win by ending with higher reward' sentence"},
      ]
    },
    {
      question: "Which game-specific ablations should we include?",
      header: "Game-specific",
      multiSelect: true,
      options: [
        {label: "no_accept_preview", description: "Drop 'you would receive [...]' — model must compute complement itself"},
        {label: "generic_names",     description: "Book/Hat/Basketball → A/B/C — strip semantic priors"},
      ]
    },
  ]
})

After the user responds, ask one follow-up if anything is ambiguous (a rewritten rationale, a swap of one axis for another). It's fine to loop two or three times. You may not skip the proposal — even if the user accepts everything immediately, you must have surfaced concrete proposals first.

Step 4: Write the variants

Bootstrap prompt_variants.py if it doesn't exist; otherwise update it.

Pick the right template based on env type:

  • OpenSpiel games (open_spiel_*) → start from templates/prompt_variants_openspiel.py.tmpl
  • Anything else → start from templates/prompt_variants_generic.py.tmpl

For each variant:

  1. baseline must be byte-identical to harness.py. Port the existing prompt template and helper logic into a BaselineVariant class. Do not paraphrase — the next step (ablation check) will fail if even one character differs.
  2. null must be byte-identical to baseline but registered under a separate name. The simplest implementation is just NULL = PromptVariant(name="null", ..., build_body=<same as BASELINE>, ...) — share every field with BASELINE except name. This guarantees no accidental drift between the two. The runner will schedule independent cells for null, and the resulting null-vs-baseline Σ|Δrank| measures pure LLM-sampling noise that the analysis step uses as the noise floor. Always include null in VARIANTS — without it, the permutation test has to estimate the noise floor by resampling, which overestimates noise at small N and drowns real effects.
  3. Each new variant is a subclass of BaselineVariant (or sibling class) that overrides make_prompt (and parse_response if the schema changed — e.g. generic_names uses {a, b, c} JSON keys and needs an alias map).
  4. Expose VARIANTS = {"baseline": BaselineVariant(), "null": NullVariant(), ...} at module level. The runner enforces a "baseline" entry; the analysis script defaults to looking for "null" and degrades gracefully if missing.

Keep the variant code in one file. Do not import shared prompt fragments from a sibling helper module — harness.py is manually deployed and versioned separately, so a shared-module import would be a deploy-isolation footgun.

Hard rule: never strip game-mechanics information

Decoration is fair game (worked examples, flavor sentences, repeated reminders, loss-threat language, computed per-turn helpers). Mechanics is not. Every variant — including the most aggressive minimal — must convey the information the model needs to understand the game:

  • What's private vs. public. If some piece of state is hidden from the opponent (Bargaining: per-player valuations; Werewolf: roles; Avalon: alignment), say so. Models default to assuming symmetric information and play very differently when wrong.
  • Action semantics. What does each move-type mean? What does each field in the JSON schema represent? ("keep" = items YOU retain; "agree" = accept opponent's last offer.)
  • Reward / win condition. How is the score computed?
  • Terminal / timeout behavior. What happens if no resolution is reached?

If you find yourself stripping any of these, stop — you're testing comprehension, not prompt design. The variant's poor performance can't be attributed to the prompt-design choice you wanted to ablate. Audit each variant by reading it as a model with no prior game knowledge would. Can you play correctly from this prompt alone? If not, restore the missing mechanics.

Hard rule: every prompt must ask for reasoning before JSON

Every variant's template — including the most aggressive minimal strip — must explicitly instruct the model to output its reasoning before outputting the final action JSON. The wording must make clear that the reasoning belongs in the response, not just in the model's head. Two safe phrasings:

  • baseline-style: "Respond with your reasoning, then conclude with a JSON block of EITHER form:"
  • terse-style: "Respond with your reasoning, then end your response with JSON, one of:"

Avoid weak verbs like "Reason briefly through your move, then respond with JSON" — models often interpret this as "think about it internally, then output JSON" and skip writing reasoning into the response. The instruction must start with an output verb (Respond, Write, Explain, Output) applied to the reasoning, not just to the JSON.

This is non-negotiable. Models perform noticeably worse on these games when they skip writing out reasoning; the chain-of-thought is load-bearing, not a stylistic preference. When you propose new variants in Step 3, when you write the variant classes in this step, and when you review the final prompt_variants.py before running, check that every template's final-output instruction puts an output verb on the reasoning, not just on the JSON. If a variant's whole point is "strip everything", strip rules and helpers and examples — but keep the reason-first instruction.

Step 5: Verify with ablation check

bash
uv run python -m kaggle_environments.ablation check --env <env_name>

Required output:

env: <env_name>
variants: ['baseline', 'compact', ...]
  baseline parity obs[0]: ok (NNNN chars)
  baseline parity obs[1]: ok (NNNN chars)
  ...
  variant 'compact': rendered ok across 5 observations
  ...
OK

If baseline parity fails, stop and fix the BaselineVariant port before proceeding. The most common cause is a paraphrased docstring or a stripped trailing newline. Show the user the diff (use harness.generate_prompt(obs, []) vs VARIANTS["baseline"].make_prompt(obs, [])) so they can confirm what changed.

If a variant errors on render (template KeyError, missing field), fix it. Do not move on with broken variants.

Step 6: Confirm the run plan

Before spending API budget, confirm the run with the user. Use AskUserQuestion or open prose. Surface:

  • The env, the variant list (always including null), the model list, --games (paired games per matchup — must be even).
  • The total cell count: K · M(M−1)/2 · games, where K includes the null variant. Add M · games per leaderboard if --self-play.
  • A rough cost estimate. For Bargaining-class games (≤10 prompt-response rounds, ~1k–3k prompt tokens), assume ~20 LLM calls per game. For longer-horizon or simultaneous-move games (e.g. Markov Soccer at horizon 100), assume 40+.
  • The output directory.

Default suggestion if the user hasn't specified: --games 30, all variants (including null), --concurrency 8. Don't go below N=20 unless smoke-testing — at N=10, the permutation test has poor power and most variants come back inside noise even when their qualitative rank shifts look real. Wait for explicit go before invoking the runner.

Set expectations from existing leaderboard CIs

If the user has an existing Bradley-Terry leaderboard for this game, look at the CI widths across models before promising results. If the middle tiers have CIs of ±30 Elo or more (overlapping bands), most prompt-ablation effects will be inside noise regardless of how well the ablation is run — the game is not distinguishing prompts because it's not distinguishing anything below the extremes. Say this upfront: the ablation can still tell them "is it safe to simplify?" (a valuable answer) but is unlikely to tell them "which prompt wins" (because no prompt does, at that noise level). Steer them toward asking specific behavioral questions ("does model X drop when helper Y is stripped?") rather than expecting wholesale rank shifts.

Step 7: Run the tournament

bash
MODEL_PROXY_KEY=$KEY MODEL_PROXY_URL=$URL \
uv run python -m kaggle_environments.ablation run \
  --env <env_name> \
  --models <csv of model names> \
  --games <N> \
  --concurrency 8 \
  --out results/<env>_<slug>/

Prefer a stable slug over a date in the output directory (e.g. open_spiel_markov_soccer_v1, not 20260706). Long runs can cross date boundaries and stale dates in results paths are confusing.

Stream the runner's progress to the user. It writes games.csv (one row per game) and summary.md (per-variant leaderboards + cross-variant rank shifts + anomalies) into --out.

Built-in safety nets

The runner protects budget with three mechanisms — all on by default, all overridable:

  • Preflight probe. Before scheduling any game, the runner calls each model with a one-token prompt. If any model fails auth/quota/connectivity, the run aborts with the error message and zero games spent. Override with --skip-preflight only when you're sure your models work (CI, cached configs).
  • Auto-abort on crash rate. After every completed cell, the runner checks per-variant crash rate. If a variant crosses --max-crash-rate (default 0.6, i.e. 60%) after --min-cells-per-variant (default 8) completions, the run halts and prints the failure category. Combined with the interleaved scheduling (variants round-robin, not sequential), a broken variant or a dying proxy is caught in the first few minutes rather than after burning the whole budget on one leaderboard.
  • Interleaved scheduling. Cells are scheduled round-robin across variants, so each variant gets some completions early. Even without auto-abort, this means a systemic failure surfaces on the first pass through the variants rather than after 3/4 of the run.

Disable auto-abort with --max-crash-rate 0 for smoke tests.

Show full SKILL.md (1,272 more words)Show less
Recovering from a partial run

If the run halts mid-way — auto-abort, network hiccup, ctrl-C, quota exhaustion — games.csv is preserved with everything completed so far. To resume:

  • Straight resume (skip anything already in games.csv): re-run with --resume. Good when the interruption was clean and the completed cells are trustworthy.
  • Redo the crashed cells (proxy died, some cells contaminated): re-run with --resume --redo-crashed. This strips rows where crash_p0 or crash_p1 is True from games.csv and re-schedules only those cells. Cleanly-completed cells are preserved. Use this when the crashes were driven by a transient environmental cause (quota, network) rather than a real variant bug.

Before choosing between these, inspect the failure_reason column in games.csv to diagnose:

bash
awk -F',' 'NR>1 && $NF!="" {print $NF}' games.csv | sort | uniq -c

Categories: quota (proxy budget rejected the call), timeout (LLM call or game watchdog fired), parse_failure (model produced illegal/unparseable output after retries), agent_error (uncategorized agent-side exception), env_error (env.run itself raised). A concentration of quota or timeout is fixable with --redo-crashed once the environment is healthy; a concentration of parse_failure on one variant usually means the variant's prompt is broken and needs a fix in prompt_variants.py first.

Step 8: Report findings

First run the permutation-test analysis. This is the headline output — it tells you which rank shifts are real vs. noise:

bash
uv run python -m kaggle_environments.ablation_analysis \
  --csv <out>/games.csv \
  --baseline baseline --null null \
  --permutations 2000 \
  --out <out>/analysis.md

The script writes a Markdown table per real variant with:

  • observed Σ|Δrank| — how much the leaderboard reshuffled vs. baseline
  • null-floor Σ|Δrank| — what the byte-identical null variant produced (pure sampling noise)
  • permutation p-value — fraction of label-shuffles producing a shift this large or larger

Lead with the analysis result, not the raw summary.md tables. A variant whose observed Σ|Δrank| is comfortably above the null floor and has p < 0.05 is a real prompt-sensitivity finding. A variant whose observed Σ|Δrank| is near or below the null floor isn't moving the leaderboard meaningfully, even if the raw rank table looks suggestive.

Then surface the supporting context from summary.md:

  • The cross-variant rank-shift table — for the variants the analysis flagged as significant, show what moved (which model swapped with which).
  • The anomalies section. Variants with >5% crash rate or hard errors usually indicate the prompt is under-specified for some model — investigate before treating the variant's rankings as comparable to others.
  • Mean-score deltas for context on the direction of effect (variant helps or hurts each model).

Make one or two concrete recommendations: which variant(s) look like statistically-supported upgrades to baseline, which look like clear regressions to avoid. Don't just dump the tables — interpret them through the lens of "did this clear the noise floor?"

When the noise floor swallows everything

If every real variant has Σ|Δrank| at or below the null floor, the honest answer is the prompt doesn't meaningfully affect rankings at this sample size with these models on this game. Don't massage marginal results — say so plainly, and recommend either (a) running at higher N if the user has budget, (b) trying more aggressive variants (the current ones may be too close to baseline), or (c) accepting that prompt scaffolding is mostly decoration for this task.

Common pitfalls

  • Stripping mechanics, not decoration. "Minimal" doesn't mean "remove everything." Variants must preserve private-vs-public splits, action semantics, reward computation, and terminal conditions. If a model couldn't play the game from the variant's prompt alone, you're not isolating a prompt-design effect — you're measuring comprehension failure. See the hard-rule subsection in Step 4.
  • Omitting the reason-first instruction from a variant. Every prompt (baseline + all variants, however terse) must ask the model to reason before outputting JSON. A variant that drops it is testing a confounded prompt — the lost performance from skipping reasoning will dominate whatever else you were trying to ablate. See the hard-rule subsection in Step 4.
  • Forgetting baseline parity. If the control arm isn't byte-identical to production, every comparison is contaminated. Run ablation check after every edit to prompt_variants.py, not just at the end.
  • Skipping the null variant. Without it, you have no calibrated noise floor and the permutation test has to estimate one by resampling (which overestimates noise at small N). The cost is one extra variant in the matrix — always include it.
  • Reporting raw rank tables as findings. The rank table in summary.md is unweighted by significance. A #1↔#2 swap looks dramatic but at N=10 it happens regularly from sampling noise. Always run ablation_analysis.py and lead with its p-values; the rank table is supporting evidence, not the headline.
  • Asking before proposing. "What variants would you like?" is the wrong question. The user is invoking this skill because they want concrete suggestions. Always propose 3–5 concrete variants first, then iterate.
  • Picking too few games per matchup. With N=2 (one pair per matchup), confidence intervals on win-rate span nearly the full [0%, 100%] range. Default to N=30 (15 pairs) for the permutation test to have real power; drop to N=4 only for smoke tests. At N=10 most real prompt effects come back inside the noise envelope.
  • Self-play by default. --self-play doubles cost and rarely changes conclusions. Off unless the user asks.
  • Modifying harness.py. The production prompt stays in harness.py. All experimental prompts live in prompt_variants.py. If a variant turns out to be a win, promote it to harness.py in a separate PR (see the promotion workflow section below).
  • Per-game extra metrics. The runner ships generic columns (score, winner, length, crash, duration, failure_reason). If you want game-specific columns (Bargaining's agreement_step, chess's mate_in_n), add them to games.csv in a follow-up post-processing step — the runner's schema is intentionally minimal.
  • Ignoring the failure_reason column. If crashes cluster in quota or timeout, the environment is the problem, not the variant — recover with --resume --redo-crashed. If they cluster in parse_failure on one variant, the variant's prompt is under-specified for at least one model — fix it and re-run just that variant with --variants X --resume --redo-crashed. Never conflate the two: rerunning a broken variant burns budget on data you'll discard again.
  • Skipping the preflight probe habitually. --skip-preflight exists for CI where the model list has been validated by a prior run. In interactive use it's the "trust me, they work" button and it will bite you. Cost of the probe is ~30s and one token per model.

Promotion workflow: from variant win to production

The ablation's prompt_variants.py stays checked in permanently — it's the experimental surface for future re-testing. When an ablation identifies a winning variant that should replace baseline in production, do NOT edit prompt_variants.py to make the winner the new baseline. Instead, promote in a separate PR:

  1. This PR (the ablation itself): commit prompt_variants.py, the results directory (games.csv, summary.md, analysis.md), and any notes. This is the evidence trail.
  2. Follow-up PR (the promotion): modify harness.py (and its test) only, applying the winning variant's specific changes. Reference the ablation PR in the description for justification. Do not touch prompt_variants.py in this PR.

Canonical example: Bargaining. PR #1279 introduced the ablation tool + skill; PR #1273 ("Save Bargaining prompt changes from experiment") promoted the winning variant into harness.py in an isolated 16-line diff that only touched harness + test.

Why the separation:

  • The ablation stays reproducible — future runs can re-verify or challenge the finding.
  • The promotion PR is easy to review — it's just the diff to production.
  • Reverting production is a one-commit revert, not a partial rollback.

If the ablation produced no statistically supported winner, just do step 1. Don't ship prompt changes on directional-but-inside-noise signal; add a note to the results dir describing what was tried so future work doesn't retread the same axes.

Reference: the contract enforced by the runner

python
# prompt_variants.py
from kaggle_environments.core_harness import GameHarness, ParseResult

class BaselineVariant:  # implements GameHarness
    def get_legal_moves(self, observation): ...
    def make_prompt(self, observation, move_history,
                    previous_response=None, previous_action=None): ...
    def parse_response(self, response, legal_action_strings,
                       *, observation=None): ...

VARIANTS: dict[str, GameHarness] = {
    "baseline": BaselineVariant(),
    "null":     NullVariant(),       # byte-identical duplicate of baseline
    "compact":  CompactVariant(),
    # ...
}

The runner imports this module, picks variants by name, and passes each to create_agent_fn(variant, model_override=...) per game. The analysis step (ablation_analysis.py) reads the resulting games.csv and uses the null variant as the noise floor for permutation tests against every other variant. No further wiring needed.

© Kaggle, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .agents/skills/run-ablation of Kaggle/kaggle-environments.

  • SKILL.md
  • templates/prompt_variants_generic.py.tmpl
  • templates/prompt_variants_openspiel.py.tmpl

Open the folder on GitHubat commit 1c8acf1

Compare with similar skills

Run Ablation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Ablation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Ablation this skillKaggle/kaggle-environments454—~6.9kAutomated safety check: PassApache-2.0
Nvidia Kaggle SkillNVIDIA/nvidia-kaggle336—~2.3kAutomated safety check: NotesMIT
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT
Lilly Community Researchssaaffaakk/Lilly171—~1.4kAutomated safety check: PassMIT
Kaggle LearnerGalaxy-Dawn/claude-scholar5.7k2 repos~940Automated safety check: PassMIT
Kaggle Researchbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~868Automated safety check: PassCustom licence

Similar skills

  • Nvidia Kaggle Skill

    NVIDIA/nvidia-kaggle

    Official

    A skill your agent uses for Kaggle competition overview fetches, writeups, discussion/kernel research, submissions, and dataset uploads.

    336 GitHub stars~2.3k tokensUpdated 2 mo ago
    Auto-check: notes
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Lilly Community Research

    ssaaffaakk/Lilly

    Lilly community-research skill. An agent skill from ssaaffaakk/Lilly.

    171 GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Kaggle Learner

    Galaxy-Dawn/claude-scholar

    This skill should be used when the user asks to "learn from Kaggle", "study Kaggle solutions", "analyze Kaggle competitions", or mentions Kaggle competition URLs.

    5.7k GitHub starsUsed in 2 repos~940 tokens
    Data & AnalyticsAuto-check passed
  • Kaggle Research

    brycewang-stanford/Auto-Empirical-Research-Skills

    A skill your agent uses when a research task needs reproducible Kaggle discovery, metadata inspection, bounded public-data downloads, competition or kernel discovery, model discovery, or an…

    4.5k GitHub stars~868 tokensUpdated 3 days ago
    Auto-check passed
  • Annotation Data

    SharpAI/DeepCamera

    Dataset annotation management — COCO labels, sequences, export, and Kaggle upload

    3.1k GitHub stars~495 tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed

More from Kaggle/kaggle-environments

  • Review Docs

    Kaggle/kaggle-environments

    Audit a game environment's README.md and AGENTS.md against its engine implementation.

    454 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Create Harness

    Kaggle/kaggle-environments

    Create or update an LLM harness that lets a language model play a kaggle-environments game.

    454 GitHub stars~8.9k tokensUpdated today
    Auto-check passed
  • Review Harness

    Kaggle/kaggle-environments

    Review an existing LLM harness for correctness and gameplay-impacting bugs.

    454 GitHub stars~24k tokensUpdated today
    Auto-check passed

Works with

Questions about Run Ablation

What does Run Ablation do?

Run a prompt-ablation study on a kaggle-environments game's LLM harness. Run Ablation is an agent skill from Kaggle/kaggle-environments. Run a prompt-ablation study on a kaggle-environments game's LLM harness.

When should I use Run Ablation?

Run Ablation fits situations like: the user mentions ablation; prompt sensitivity; test prompt variants; compare prompts.

How do I install Run Ablation in Claude Code?

Run `npx skills add Kaggle/kaggle-environments --skill run-ablation -a claude-code`. Or copy the skill folder (.agents/skills/run-ablation in Kaggle/kaggle-environments) into .claude/skills/run-ablation in your project. Claude Code loads it when a task matches its description.

How do I install Run Ablation in Codex?

Run `npx skills add Kaggle/kaggle-environments --skill run-ablation -a codex`. Or copy the skill folder (.agents/skills/run-ablation in Kaggle/kaggle-environments) into .agents/skills/run-ablation in your project. Codex loads it when a task matches its description.

Can I use Run Ablation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kaggle/kaggle-environments --skill run-ablation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-ablation, .gemini/skills/run-ablation, .github/skills/run-ablation and .opencode/skills/run-ablation in your project.

What does Run Ablation need to run?

Going by SKILL.md and its folder, Run Ablation needs the command-line tools its instructions call (python and uv) and credentials named MODEL_PROXY_KEY. Our summary lists: Python 3.

Does Run Ablation access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Run Ablation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Ablation use?

Run Ablation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run Ablation use?

About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Run Ablation?

Skills that share tags, products or a category with Run Ablation: Nvidia Kaggle Skill (NVIDIA/nvidia-kaggle, 336 stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars), Lilly Community Research (ssaaffaakk/Lilly, 171 stars) and Kaggle Learner (Galaxy-Dawn/claude-scholar, 5.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Ablation?

Kaggle (a GitHub organization) maintains it in Kaggle/kaggle-environments, which has 454 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 8, 2026.

Source: Kaggle/kaggle-environments on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.