Nvidia Kaggle Skill
NVIDIA/nvidia-kaggle
A skill your agent uses for Kaggle competition overview fetches, writeups, discussion/kernel research, submissions, and dataset uploads.
Run a prompt-ablation study on a kaggle-environments game's LLM harness.
$ npx skills add Kaggle/kaggle-environments --skill run-ablation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Kaggle/kaggle-environments run-ablation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/run-ablation .claude/skills/run-ablation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "run-ablation" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/run-ablation into .claude/skills/run-ablation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-ablation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/run-ablationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Kaggle/kaggle-environments --skill run-ablation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Kaggle/kaggle-environments run-ablation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/run-ablation .agents/skills/run-ablation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "run-ablation" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/run-ablation into .agents/skills/run-ablation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-ablation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Kaggle/kaggle-environments --skill run-ablation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Kaggle/kaggle-environments run-ablation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/run-ablation .cursor/skills/run-ablation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "run-ablation" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/run-ablation into .cursor/skills/run-ablation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-ablation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Kaggle/kaggle-environments.git --path .agents/skills/run-ablation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Kaggle/kaggle-environments --skill run-ablation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Kaggle/kaggle-environments run-ablation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/run-ablation .gemini/skills/run-ablation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "run-ablation" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/run-ablation into .gemini/skills/run-ablation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-ablation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Kaggle/kaggle-environments run-ablationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Kaggle/kaggle-environments --skill run-ablation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/run-ablation .github/skills/run-ablation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "run-ablation" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/run-ablation into .github/skills/run-ablation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-ablation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Kaggle/kaggle-environments --skill run-ablation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Kaggle/kaggle-environments run-ablation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/run-ablation .opencode/skills/run-ablation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "run-ablation" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/run-ablation into .opencode/skills/run-ablation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-ablation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
run-ablationRun a prompt-ablation study on a kaggle-environments game's LLM harness.
Run Ablation is an agent skill from Kaggle/kaggle-environments. Run a prompt-ablation study on a kaggle-environments game's LLM harness. Use when the user mentions "ablation", "prompt sensitivity", "test prompt variants", "compare prompts", "ablate the prompt", "prompt rewrite study", or asks whether their prompt wording is doing real work. Bootstraps a promptvariants.py if missing, proposes new variants interactively, then runs a paired-seat tournament and reports the leaderboards.
Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files.
It works with Kaggle. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1c8acf1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonuvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
MODEL_PROXY_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Run Ablation loads about 6.9k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 3,362 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Kaggle/kaggle-environments at commit 1c8acf1, republished under its Apache-2.0 licence (© Kaggle). 3,362 words, ~6,881 tokens.
.claude/skills/run-ablation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.This skill drives the end-to-end prompt-sensitivity workflow for one game's harness: discover or bootstrap variants, propose additional ones in conversation with the user, run the paired-seat tournament, and report the findings. It is completely standalone — it does not depend on or modify the create-harness, review-harness, or create-environment skills.
Two CLI tools work together:
# Verify baseline parity + variant rendering before spending any budget
python -m kaggle_environments.ablation check --env <env_name>
# Run the tournament. Auto-preflight, auto-abort on runaway crashes, and
# --redo-crashed for recovery are all on by default; see Step 7.
python -m kaggle_environments.ablation run --env <env_name> --models <csv> --games <N> --out <dir>
# Post-hoc statistical analysis (permutation test against the null variant's noise floor)
python -m kaggle_environments.ablation_analysis --csv <dir>/games.csv --baseline baseline --null nullThe runner contract:
prompt_variants.py exposing VARIANTS: dict[str, GameHarness]. Each variant is a GameHarness (implements get_legal_moves / make_prompt / parse_response), so create_agent_fn(variant) works directly.baseline variant must be byte-identical to the production harness.py. The check subcommand enforces this against seeded observations.null variant must be a byte-identical second copy of baseline. The runner schedules independent cells for it; the difference between null-vs-baseline rankings is pure LLM-sampling noise, which calibrates the noise floor for the permutation test.{A, B} matchup is scheduled in seat-flipped pairs sharing one chance seed — the same instance is played twice with seats swapped. --games must be even.Confirm with the user which env they want to ablate. Acceptable forms: open_spiel_bargaining, werewolf, tictactoe, etc. The env name is what gets passed to kaggle_environments.make(env_name).
Resolve the harness directory:
open_spiel_<game> → kaggle_environments/envs/open_spiel_env/games/<game>/<env> → kaggle_environments/envs/<env>/Confirm harness.py exists at that path. If not, stop and tell the user: the production harness has to exist before you can ablate it. Suggest they run create-harness first.
Read prompt_variants.py if it exists at the harness directory. List the names found and one-line summaries:
Found prompt_variants.py with: baseline, compact, minimal, no_accept_preview, generic_namesIf the file doesn't exist, you'll bootstrap it in Step 4. Tell the user you're going to:
No prompt_variants.py yet — I'll bootstrap one from harness.py.
The 'baseline' variant will be byte-identical to the production prompt;
I'll also propose a few ablations for your review.Open harness.py and read the make_prompt (or generate_prompt) function plus any prompt-template constants it uses. Note the prompt's structural features — these are what your variant proposals will target:
You can't propose a good ablation without understanding the prompt's joints. Spend a turn here.
This step is mandatory. Even if prompt_variants.py already has variants, propose new ones. Skipping this step defeats the purpose of the skill — the point is to force a "what am I actually testing?" beat before any LLM money is spent.
Generate 3–5 candidate variants with a one-line rationale each. Cover at least three of these generic axes:
Compact and Minimal are distinct ablation axes, not "more terse" vs "even more terse." If you're proposing both, make sure they actually test different things — Compact = "do we need the prose?", Minimal = "do we need the helpers?" The mechanics keep is non-negotiable in both (see Step 4).
Then propose 1–2 game-specific axes that you identify from reading the prompt in Step 2. These are the most interesting ablations because they're tailored to the actual prompt design choices the author made.
Each proposal should include:
no_payoff_hint)Use the AskUserQuestion tool to surface the proposals, with one question per group of related variants, options for accept/skip/rewrite/replace. Concrete usage pattern:
AskUserQuestion({
questions: [
{
question: "Which of these structural ablations should we include?",
header: "Structural",
multiSelect: true,
options: [
{label: "compact", description: "Same info, terser — drops worked example + loss threat"},
{label: "minimal", description: "Strip rules/goal/helpers, just state + JSON schema"},
{label: "no_goal_hint", description: "Remove the 'win by ending with higher reward' sentence"},
]
},
{
question: "Which game-specific ablations should we include?",
header: "Game-specific",
multiSelect: true,
options: [
{label: "no_accept_preview", description: "Drop 'you would receive [...]' — model must compute complement itself"},
{label: "generic_names", description: "Book/Hat/Basketball → A/B/C — strip semantic priors"},
]
},
]
})After the user responds, ask one follow-up if anything is ambiguous (a rewritten rationale, a swap of one axis for another). It's fine to loop two or three times. You may not skip the proposal — even if the user accepts everything immediately, you must have surfaced concrete proposals first.
Bootstrap prompt_variants.py if it doesn't exist; otherwise update it.
Pick the right template based on env type:
open_spiel_*) → start from templates/prompt_variants_openspiel.py.tmpltemplates/prompt_variants_generic.py.tmplFor each variant:
baseline must be byte-identical to harness.py. Port the existing prompt template and helper logic into a BaselineVariant class. Do not paraphrase — the next step (ablation check) will fail if even one character differs.null must be byte-identical to baseline but registered under a separate name. The simplest implementation is just NULL = PromptVariant(name="null", ..., build_body=<same as BASELINE>, ...) — share every field with BASELINE except name. This guarantees no accidental drift between the two. The runner will schedule independent cells for null, and the resulting null-vs-baseline Σ|Δrank| measures pure LLM-sampling noise that the analysis step uses as the noise floor. Always include null in VARIANTS — without it, the permutation test has to estimate the noise floor by resampling, which overestimates noise at small N and drowns real effects.BaselineVariant (or sibling class) that overrides make_prompt (and parse_response if the schema changed — e.g. generic_names uses {a, b, c} JSON keys and needs an alias map).VARIANTS = {"baseline": BaselineVariant(), "null": NullVariant(), ...} at module level. The runner enforces a "baseline" entry; the analysis script defaults to looking for "null" and degrades gracefully if missing.Keep the variant code in one file. Do not import shared prompt fragments from a sibling helper module — harness.py is manually deployed and versioned separately, so a shared-module import would be a deploy-isolation footgun.
Decoration is fair game (worked examples, flavor sentences, repeated reminders, loss-threat language, computed per-turn helpers). Mechanics is not. Every variant — including the most aggressive minimal — must convey the information the model needs to understand the game:
If you find yourself stripping any of these, stop — you're testing comprehension, not prompt design. The variant's poor performance can't be attributed to the prompt-design choice you wanted to ablate. Audit each variant by reading it as a model with no prior game knowledge would. Can you play correctly from this prompt alone? If not, restore the missing mechanics.
Every variant's template — including the most aggressive minimal strip — must explicitly instruct the model to output its reasoning before outputting the final action JSON. The wording must make clear that the reasoning belongs in the response, not just in the model's head. Two safe phrasings:
"Respond with your reasoning, then conclude with a JSON block of EITHER form:""Respond with your reasoning, then end your response with JSON, one of:"Avoid weak verbs like "Reason briefly through your move, then respond with JSON" — models often interpret this as "think about it internally, then output JSON" and skip writing reasoning into the response. The instruction must start with an output verb (Respond, Write, Explain, Output) applied to the reasoning, not just to the JSON.
This is non-negotiable. Models perform noticeably worse on these games when they skip writing out reasoning; the chain-of-thought is load-bearing, not a stylistic preference. When you propose new variants in Step 3, when you write the variant classes in this step, and when you review the final prompt_variants.py before running, check that every template's final-output instruction puts an output verb on the reasoning, not just on the JSON. If a variant's whole point is "strip everything", strip rules and helpers and examples — but keep the reason-first instruction.
ablation checkuv run python -m kaggle_environments.ablation check --env <env_name>Required output:
env: <env_name>
variants: ['baseline', 'compact', ...]
baseline parity obs[0]: ok (NNNN chars)
baseline parity obs[1]: ok (NNNN chars)
...
variant 'compact': rendered ok across 5 observations
...
OKIf baseline parity fails, stop and fix the BaselineVariant port before proceeding. The most common cause is a paraphrased docstring or a stripped trailing newline. Show the user the diff (use harness.generate_prompt(obs, []) vs VARIANTS["baseline"].make_prompt(obs, [])) so they can confirm what changed.
If a variant errors on render (template KeyError, missing field), fix it. Do not move on with broken variants.
Before spending API budget, confirm the run with the user. Use AskUserQuestion or open prose. Surface:
null), the model list, --games (paired games per matchup — must be even).K · M(M−1)/2 · games, where K includes the null variant. Add M · games per leaderboard if --self-play.Default suggestion if the user hasn't specified: --games 30, all variants (including null), --concurrency 8. Don't go below N=20 unless smoke-testing — at N=10, the permutation test has poor power and most variants come back inside noise even when their qualitative rank shifts look real. Wait for explicit go before invoking the runner.
If the user has an existing Bradley-Terry leaderboard for this game, look at the CI widths across models before promising results. If the middle tiers have CIs of ±30 Elo or more (overlapping bands), most prompt-ablation effects will be inside noise regardless of how well the ablation is run — the game is not distinguishing prompts because it's not distinguishing anything below the extremes. Say this upfront: the ablation can still tell them "is it safe to simplify?" (a valuable answer) but is unlikely to tell them "which prompt wins" (because no prompt does, at that noise level). Steer them toward asking specific behavioral questions ("does model X drop when helper Y is stripped?") rather than expecting wholesale rank shifts.
MODEL_PROXY_KEY=$KEY MODEL_PROXY_URL=$URL \
uv run python -m kaggle_environments.ablation run \
--env <env_name> \
--models <csv of model names> \
--games <N> \
--concurrency 8 \
--out results/<env>_<slug>/Prefer a stable slug over a date in the output directory (e.g. open_spiel_markov_soccer_v1, not 20260706). Long runs can cross date boundaries and stale dates in results paths are confusing.
Stream the runner's progress to the user. It writes games.csv (one row per game) and summary.md (per-variant leaderboards + cross-variant rank shifts + anomalies) into --out.
The runner protects budget with three mechanisms — all on by default, all overridable:
--skip-preflight only when you're sure your models work (CI, cached configs).--max-crash-rate (default 0.6, i.e. 60%) after --min-cells-per-variant (default 8) completions, the run halts and prints the failure category. Combined with the interleaved scheduling (variants round-robin, not sequential), a broken variant or a dying proxy is caught in the first few minutes rather than after burning the whole budget on one leaderboard.Disable auto-abort with --max-crash-rate 0 for smoke tests.
If the run halts mid-way — auto-abort, network hiccup, ctrl-C, quota exhaustion — games.csv is preserved with everything completed so far. To resume:
--resume. Good when the interruption was clean and the completed cells are trustworthy.--resume --redo-crashed. This strips rows where crash_p0 or crash_p1 is True from games.csv and re-schedules only those cells. Cleanly-completed cells are preserved. Use this when the crashes were driven by a transient environmental cause (quota, network) rather than a real variant bug.Before choosing between these, inspect the failure_reason column in games.csv to diagnose:
awk -F',' 'NR>1 && $NF!="" {print $NF}' games.csv | sort | uniq -cCategories: quota (proxy budget rejected the call), timeout (LLM call or game watchdog fired), parse_failure (model produced illegal/unparseable output after retries), agent_error (uncategorized agent-side exception), env_error (env.run itself raised). A concentration of quota or timeout is fixable with --redo-crashed once the environment is healthy; a concentration of parse_failure on one variant usually means the variant's prompt is broken and needs a fix in prompt_variants.py first.
First run the permutation-test analysis. This is the headline output — it tells you which rank shifts are real vs. noise:
uv run python -m kaggle_environments.ablation_analysis \
--csv <out>/games.csv \
--baseline baseline --null null \
--permutations 2000 \
--out <out>/analysis.mdThe script writes a Markdown table per real variant with:
Lead with the analysis result, not the raw summary.md tables. A variant whose observed Σ|Δrank| is comfortably above the null floor and has p < 0.05 is a real prompt-sensitivity finding. A variant whose observed Σ|Δrank| is near or below the null floor isn't moving the leaderboard meaningfully, even if the raw rank table looks suggestive.
Then surface the supporting context from summary.md:
Make one or two concrete recommendations: which variant(s) look like statistically-supported upgrades to baseline, which look like clear regressions to avoid. Don't just dump the tables — interpret them through the lens of "did this clear the noise floor?"
If every real variant has Σ|Δrank| at or below the null floor, the honest answer is the prompt doesn't meaningfully affect rankings at this sample size with these models on this game. Don't massage marginal results — say so plainly, and recommend either (a) running at higher N if the user has budget, (b) trying more aggressive variants (the current ones may be too close to baseline), or (c) accepting that prompt scaffolding is mostly decoration for this task.
baseline parity. If the control arm isn't byte-identical to production, every comparison is contaminated. Run ablation check after every edit to prompt_variants.py, not just at the end.null variant. Without it, you have no calibrated noise floor and the permutation test has to estimate one by resampling (which overestimates noise at small N). The cost is one extra variant in the matrix — always include it.summary.md is unweighted by significance. A #1↔#2 swap looks dramatic but at N=10 it happens regularly from sampling noise. Always run ablation_analysis.py and lead with its p-values; the rank table is supporting evidence, not the headline.--self-play doubles cost and rarely changes conclusions. Off unless the user asks.harness.py. The production prompt stays in harness.py. All experimental prompts live in prompt_variants.py. If a variant turns out to be a win, promote it to harness.py in a separate PR (see the promotion workflow section below).agreement_step, chess's mate_in_n), add them to games.csv in a follow-up post-processing step — the runner's schema is intentionally minimal.quota or timeout, the environment is the problem, not the variant — recover with --resume --redo-crashed. If they cluster in parse_failure on one variant, the variant's prompt is under-specified for at least one model — fix it and re-run just that variant with --variants X --resume --redo-crashed. Never conflate the two: rerunning a broken variant burns budget on data you'll discard again.--skip-preflight exists for CI where the model list has been validated by a prior run. In interactive use it's the "trust me, they work" button and it will bite you. Cost of the probe is ~30s and one token per model.The ablation's prompt_variants.py stays checked in permanently — it's the experimental surface for future re-testing. When an ablation identifies a winning variant that should replace baseline in production, do NOT edit prompt_variants.py to make the winner the new baseline. Instead, promote in a separate PR:
prompt_variants.py, the results directory (games.csv, summary.md, analysis.md), and any notes. This is the evidence trail.harness.py (and its test) only, applying the winning variant's specific changes. Reference the ablation PR in the description for justification. Do not touch prompt_variants.py in this PR.Canonical example: Bargaining. PR #1279 introduced the ablation tool + skill; PR #1273 ("Save Bargaining prompt changes from experiment") promoted the winning variant into harness.py in an isolated 16-line diff that only touched harness + test.
Why the separation:
If the ablation produced no statistically supported winner, just do step 1. Don't ship prompt changes on directional-but-inside-noise signal; add a note to the results dir describing what was tried so future work doesn't retread the same axes.
# prompt_variants.py
from kaggle_environments.core_harness import GameHarness, ParseResult
class BaselineVariant: # implements GameHarness
def get_legal_moves(self, observation): ...
def make_prompt(self, observation, move_history,
previous_response=None, previous_action=None): ...
def parse_response(self, response, legal_action_strings,
*, observation=None): ...
VARIANTS: dict[str, GameHarness] = {
"baseline": BaselineVariant(),
"null": NullVariant(), # byte-identical duplicate of baseline
"compact": CompactVariant(),
# ...
}The runner imports this module, picks variants by name, and passes each to create_agent_fn(variant, model_override=...) per game. The analysis step (ablation_analysis.py) reads the resulting games.csv and uses the null variant as the noise floor for permutation tests against every other variant. No further wiring needed.
© Kaggle, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in .agents/skills/run-ablation of Kaggle/kaggle-environments.
Open the folder on GitHubat commit 1c8acf1
Run Ablation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Run Ablation this skillKaggle/kaggle-environments | 454 | — | ~6.9k | Automated safety check: Pass | Apache-2.0 | |
| Nvidia Kaggle SkillNVIDIA/nvidia-kaggle | 336 | — | ~2.3k | Automated safety check: Notes | MIT | |
| Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill | 188 | — | ~4k | Automated safety check: Pass | MIT | |
| Lilly Community Researchssaaffaakk/Lilly | 171 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Kaggle LearnerGalaxy-Dawn/claude-scholar | 5.7k | 2 repos | ~940 | Automated safety check: Pass | MIT | |
| Kaggle Researchbrycewang-stanford/Auto-Empirical-Research-Skills | 4.5k | — | ~868 | Automated safety check: Pass | Custom licence |
NVIDIA/nvidia-kaggle
A skill your agent uses for Kaggle competition overview fetches, writeups, discussion/kernel research, submissions, and dataset uploads.
FrankS-IntelLab/agentic-kaggle-skill
Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.
ssaaffaakk/Lilly
Lilly community-research skill. An agent skill from ssaaffaakk/Lilly.
Galaxy-Dawn/claude-scholar
This skill should be used when the user asks to "learn from Kaggle", "study Kaggle solutions", "analyze Kaggle competitions", or mentions Kaggle competition URLs.
brycewang-stanford/Auto-Empirical-Research-Skills
A skill your agent uses when a research task needs reproducible Kaggle discovery, metadata inspection, bounded public-data downloads, competition or kernel discovery, model discovery, or an…
SharpAI/DeepCamera
Dataset annotation management — COCO labels, sequences, export, and Kaggle upload
Kaggle/kaggle-environments
Audit a game environment's README.md and AGENTS.md against its engine implementation.
Kaggle/kaggle-environments
Create or update an LLM harness that lets a language model play a kaggle-environments game.
Kaggle/kaggle-environments
Review an existing LLM harness for correctness and gameplay-impacting bugs.
Works with
Run a prompt-ablation study on a kaggle-environments game's LLM harness. Run Ablation is an agent skill from Kaggle/kaggle-environments. Run a prompt-ablation study on a kaggle-environments game's LLM harness.
Run Ablation fits situations like: the user mentions ablation; prompt sensitivity; test prompt variants; compare prompts.
Run `npx skills add Kaggle/kaggle-environments --skill run-ablation -a claude-code`. Or copy the skill folder (.agents/skills/run-ablation in Kaggle/kaggle-environments) into .claude/skills/run-ablation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Kaggle/kaggle-environments --skill run-ablation -a codex`. Or copy the skill folder (.agents/skills/run-ablation in Kaggle/kaggle-environments) into .agents/skills/run-ablation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kaggle/kaggle-environments --skill run-ablation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-ablation, .gemini/skills/run-ablation, .github/skills/run-ablation and .opencode/skills/run-ablation in your project.
Going by SKILL.md and its folder, Run Ablation needs the command-line tools its instructions call (python and uv) and credentials named MODEL_PROXY_KEY. Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Run Ablation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Run Ablation: Nvidia Kaggle Skill (NVIDIA/nvidia-kaggle, 336 stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars), Lilly Community Research (ssaaffaakk/Lilly, 171 stars) and Kaggle Learner (Galaxy-Dawn/claude-scholar, 5.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Kaggle (a GitHub organization) maintains it in Kaggle/kaggle-environments, which has 454 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 8, 2026.
Source: Kaggle/kaggle-environments on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.