Annotation Data
SharpAI/DeepCamera
Dataset annotation management — COCO labels, sequences, export, and Kaggle upload
Create or update an LLM harness that lets a language model play a kaggle-environments game.
$ npx skills add Kaggle/kaggle-environments --skill create-harness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Kaggle/kaggle-environments create-harness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/create-harness .claude/skills/create-harness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "create-harness" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/create-harness into .claude/skills/create-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-harness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/create-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Kaggle/kaggle-environments --skill create-harness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Kaggle/kaggle-environments create-harness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/create-harness .agents/skills/create-harness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "create-harness" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/create-harness into .agents/skills/create-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-harness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Kaggle/kaggle-environments --skill create-harness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Kaggle/kaggle-environments create-harness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/create-harness .cursor/skills/create-harness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "create-harness" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/create-harness into .cursor/skills/create-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-harness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Kaggle/kaggle-environments.git --path .agents/skills/create-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Kaggle/kaggle-environments --skill create-harness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Kaggle/kaggle-environments create-harness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/create-harness .gemini/skills/create-harness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "create-harness" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/create-harness into .gemini/skills/create-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-harness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Kaggle/kaggle-environments create-harnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Kaggle/kaggle-environments --skill create-harness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/create-harness .github/skills/create-harness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "create-harness" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/create-harness into .github/skills/create-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-harness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Kaggle/kaggle-environments --skill create-harness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Kaggle/kaggle-environments create-harness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kaggle/kaggle-environments.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/create-harness .opencode/skills/create-harness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "create-harness" agent skill from https://github.com/Kaggle/kaggle-environments/tree/master/.agents/skills/create-harness into .opencode/skills/create-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-harness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
create-harnessCreate or update an LLM harness that lets a language model play a kaggle-environments game.
Create Harness is an agent skill from Kaggle/kaggle-environments. Create or update an LLM harness that lets a language model play a kaggle-environments game. Use this skill whenever the user wants to write a harness, LLM agent, or game-playing prompt for any kaggle-environments game — including OpenSpiel games, word games, or custom environments. Also use it when the user mentions coreharness.py, GameHarness, ParseResult, or asks how to connect an LLM to a game.
Its SKILL.md is about 8.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering. It works with Kaggle. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1c8acf1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Create Harness loads about 8.9k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 3,883 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Kaggle/kaggle-environments at commit 1c8acf1, republished under its Apache-2.0 licence (© Kaggle). 3,883 words, ~8,854 tokens.
.claude/skills/create-harness/SKILL.md (or your agent's skills folder).A harness bridges a game environment and an LLM. It translates game observations into prompts, sends them to the model, and parses the model's response back into a legal game action. The framework in kaggle_environments/core_harness.py handles the LLM call, retry loop, telemetry, and agent lifecycle — you only implement the game-specific logic.
A harness module that implements three functions (the GameHarness protocol):
| Method | Purpose |
|---|---|
get_legal_moves(observation) | Extract legal actions from the observation |
make_prompt(observation, move_history, ...) | Build the LLM prompt |
parse_response(response, legal_action_strings) | Extract the chosen action from the LLM's text |
Production wires these three module-level functions into an agent via an external wrapper template — do not write an adapter class or a create_agent_fn(...) line in harness.py. Tests construct their own local adapter when they want end-to-end coverage; see Step 6 for the pattern.
Before writing any code, understand the game you're building a harness for:
*_proxy.py) — if one exists, the observation will be structured JSON rather than raw OpenSpiel text. The proxy's state_dict() method shows exactly what fields the LLM will see. Non-OpenSpiel environments already provide structured JSON observations directly, so no proxy is needed.get_legal_moves returns dict[int, str].get_legal_moves returns None.observationString — the game state (often JSON from a proxy)legalActions — list of action IDs (ints)legalActionStrings — list of human-readable action stringscurrentPlayer — whose turn it is (-2 means simultaneous)playerId — this agent's player IDisTerminal — whether the game is overserializedGameAndState — full pyspiel state (OpenSpiel games only)get_legal_movesThis function extracts legal actions from the observation and returns them as {action_id: action_string}.
from typing import Any, Mapping, Sequence
from kaggle_environments.core_harness import ParseResult
def get_legal_moves(observation: Mapping[str, Any]) -> dict[int, str] | None:
"""Return {action_id: action_string} for enumerable games, or None for free-form."""
legal_actions = observation.get("legalActions")
legal_action_strings = observation.get("legalActionStrings")
if legal_actions and legal_action_strings:
return dict(zip(legal_actions, legal_action_strings))
return {}For OpenSpiel games with a fallback (when the environment doesn't always provide legalActions):
def get_legal_moves(observation: Mapping[str, Any]) -> dict[int, str]:
legal_actions = observation.get("legalActions")
legal_action_strings = observation.get("legalActionStrings")
if legal_actions and legal_action_strings:
return dict(zip(legal_actions, legal_action_strings))
# Fallback: deserialize pyspiel state
serialized = observation.get("serializedGameAndState", "")
if not serialized:
return {}
_, state = pyspiel.deserialize_game_and_state(serialized)
return {a: state.action_to_string(a) for a in state.legal_actions()}For mixed action spaces (like word_association), return None on free-form turns:
def get_legal_moves(observation: Mapping[str, Any]) -> dict[int, str] | None:
if observation.get("isCluemaster"):
return None # Free-form: LLM submits any valid clue
# Enumerable: return word choices
words = observation.get("words", [])
return {i: f"{i}: {words[i]}" for i in range(len(words))}The prompt is the most important part of a harness — it determines how well the LLM plays. The make_prompt function (sometimes named generate_prompt in older code) builds the full text sent to the model.
def generate_prompt(
observation: Mapping[str, Any],
move_history: list[str],
previous_response: str | None = None,
previous_action: str | None = None,
) -> str:observation — the current game statemove_history — list of this agent's past action strings (persists across turns)previous_response — the LLM's last response if this is a retry after an illegal moveprevious_action — the illegal move that was attemptedA good prompt follows this pattern:
observation["observationString"]Do not enumerate the legal moves in the prompt. The framework already
validates the LLM's response against legalActionStrings and triggers a
retry with the rethink suffix on an illegal move, so listing every legal
action just bloats the prompt, encourages the model to pick mechanically
from the list instead of reasoning about the position, and trivializes the
benchmark for games whose challenge is partly finding the legal moves
(e.g. checkers' forced captures). Instead, describe the move-legality rules
clearly and let the model derive legality from the state. The lone
exception is when the legal-action set is not derivable from the visible
state at all (e.g. a hidden hand of cards the LLM owns but the rules text
cannot enumerate) — in that case, list only what the model cannot infer.
Read kaggle_environments/envs/open_spiel_env/games/checkers/harness.py —
it's the canonical shape: a single *_PROMPT_TEMPLATE constant with
explicit {field} placeholders, plus two named rethink templates
(RETHINK_ILLEGAL, RETHINK_UNPARSABLE) selected by
render_rethink_suffix from core_harness.
The required pieces:
For example: line.RETHINK_ILLEGAL — fires when the parser extracted a move that wasn't
legal. Leads with {previous_action} ("You suggested … but this is
not a legal move"). Does NOT include the previous response.RETHINK_UNPARSABLE — fires when the parser extracted nothing. Leads
with {previous_response} and restates the JSON format with a clean
placeholder + concrete example.generate_prompt that builds the main prompt and appends
render_rethink_suffix(RETHINK_ILLEGAL, RETHINK_UNPARSABLE, previous_response, previous_action). The helper returns "" on the
first attempt, picks the right branch otherwise, and truncates
previous_response to the last 500 chars.For an imperfect-information variant (per-player observation, custom
matcher=), use dark_hex/harness.py as the model instead.
observation["observationString"] is JSON — parse it and present the state clearly rather than dumping raw JSON."move") works well.previous_action was extracted. Each case wants a different signal back at the model:previous_action is set (parser pulled a value from the JSON) → the action string itself is the most useful signal. Lead with "You suggested move {previous_action} but this is not legal." Do NOT also include the full previous response — that's noise the model has to skim past to find the actual correction signal. Tail with a brief reminder to keep using the same JSON format.previous_action is None (parser found nothing) → there's no action string to show. Lead with the previous response (last 500 chars, not first 500 — the model's conclusion is at the end, not the preamble) so the model can see what it tried, then restate the JSON format with a concrete example. Tail with a brief reminder that the move must also be legal.previous_action. A single one-size-fits-all suffix that always restates the format teaches the wrong fix on illegal-move retries; a single suffix that never restates the format leaves the model in the dark on unparseable retries.```json `{"move": "<notation>, e.g. 24/23 24/22"}` ``` reads like the value should literally be that whole string. Instead, use a clean placeholder in the fenced block and put a concrete example on its own line right after:```json
{"move": "<your_move>"}
```
For example: `{"move": "24/23 24/22"}`state_dict() key produces it? (b) which engine call produces THAT? (c) what params is the env actually loaded with? If you can't trace a claim to a source, it's a phantom — either implement the missing path or delete the claim. The two recurring failure shapes are drift (the field was renamed or the env switched params underneath the prompt) and aspirational copy (the author described a feature they planned but never wired up — e.g. oshi-zumo's docstring once claimed opponent coin counts were "encoded as a hidden suffix" of move_history entries, but no code surfaced them). A prompt that lies is worse than one that says less.load_game params. Don't trust the game's name or docstring — read the env factory. Gin rummy's prompt described the "Oklahoma variant" because the author assumed the default, but the env loads gin_rummy with oklahoma=false. Whenever the prompt says "in this variant…" or names a ruleset, confirm the params at the pyspiel.load_game(name, params) call site match.player_id. Sentences like "lower is your goal", "toward the top edge", "first move", or "your stones are at the bottom" are almost always asymmetric — correct for one player and wrong for the other. Render the prompt for player_id=0 AND player_id=1 and diff them; any directional text that's byte-identical between the two is probably the bug. Factor through a per-player helper (e.g. _player_info(player_id) -> (label, code, direction_text) like dark_hex does) so the asymmetry is in one place. Oshi-zumo shipped with "lower is your goal" baked in once; 7.7% of P0 turns echoed the wrong direction.{Pass, Knock}-shaped legals (gin rummy's Wall and Layoff) or any other coincidental legal-action signature, a legals-based fallback will silently misroute. Build a {phase_name: template} table keyed on the engine's own phase string/enum, assert at construction time that every engine phase is covered, and raise on an unknown phase instead of falling through to a default — a new phase should fail loudly, not silently get the wrong prompt.board_size, dark_hex's num_rows/num_cols, amazons' build-dependent defaults). A prompt that bakes in "10x10 grid", hardcodes column letters a–j, or assumes a fixed coordinate range will silently lie to the model whenever the env is loaded with a different size. Source dims from the parsed observationString (proxy state_dict() usually exposes board plus num_rows/num_cols or board_size), or — as a fallback — deserialize serializedGameAndState and read state.get_game().get_parameters(). See amazons/harness.py:113 (_board_dims) for the canonical pattern: prefer the actual board grid, fall back to explicit dimension fields, only then a default. Interpolate num_rows / num_cols into the prompt template ("on a {num_rows}x{num_cols} grid") and derive any coordinate-system text from those dims (e.g. compute the column-letter range from num_cols rather than literal "a-j"). Render the prompt at every configured size the env supports and verify each renders correctly.move_history argument is this agent's past actions only — it omits the opponent's moves entirely. A prompt that uses only this is showing the model half the game.move_number (or moves_played, turn_count) and the prompt interpolates that — "Moves played so far: 14" — instead of the actual move list. This was the clobber bug: the prompt rendered the count and the model had no way to reconstruct what had been played. A count tells the model "we're 14 moves in" and nothing else; the model needs the actual sequence (a1b1, b3a3, c2c3, ...) to reason about threats, repetitions, and what the opponent has been doing.
Sources for the full move list, in order of preference:state_dict() exposes a move_history (or action_history, moves, played_moves) field covering both players — see coin_game/harness.py and ant_foraging_arena/harness.py for proxies that surface this and harnesses that render it.serializedGameAndState and walk state.history() / state.full_history() to reconstruct it (see chess/harness.py:36 _build_pgn_movetext for a worked example that emits PGN-style movetext from the pyspiel state). For games with no chance phase, play actions alternate from player 0, so per-move player_ids fall out of index parity.move_history field (clobber's proxy did this — see clobber_proxy.py's state_dict()). Don't fall back to the per-agent move_history argument and call it "history" — that's the "Move history framing wrong" anti-pattern.
Render the moves themselves ("a1b1, b3a3, c2c3, ..." or PGN-style for chess), readably and with player labels if alternation isn't obvious from order, and label the line accurately — "Moves played so far this game (both players, oldest first): a1b1, b3a3, ..." reads true; "Moves played so far: 14" does not. Keep move_number as a separate field if useful, but never as a substitute for the move list.move_history_str = ", ".join(move_history) if move_history else "None"Shorter prompts have matched or beaten longer ones across games; padding buries the signal. After the first working draft, do a dedicated compaction pass at a real rendered observation:
Start compact; only add clarifying prose or worked examples back in when a later evaluation shows the model actually struggling.
parse_responseThe parser extracts the LLM's chosen move from its text response and matches it to a legal action. This is where robustness matters — LLMs don't always follow instructions perfectly.
def parse_response(
response: str,
legal_action_strings: Sequence[str] | None,
) -> ParseResult:ParseResult is a frozen dataclass with three fields:
@dataclasses.dataclass(frozen=True)
class ParseResult:
legal_action: str | None = None # Matched legal move string (enumerable)
raw_action: str | None = None # What the model actually said (for rethink context)
submission: Any = None # Parsed object (free-form only)For enumerable actions: set legal_action to the matched string from legal_action_strings, and raw_action to what the model originally said.
For free-form actions: set submission to the parsed object (e.g., a dict), and raw_action to a string representation.
On failure: return ParseResult(raw_action=<what_was_attempted>) — the framework will retry with the rethink prompt.
parse_json_action from core_harnessFor an enumerable harness, your parse_response should be one line —
delegate to parse_json_action:
from kaggle_environments.core_harness import ParseResult, parse_json_action
def parse_response(
response: str, legal_action_strings: Sequence[str],
) -> ParseResult:
"""Trust the model's JSON answer; let the rethink loop fix anything else."""
return parse_json_action(response, legal_action_strings)parse_json_action enforces the design rule that the parser has exactly
one intent surface: the model's structured JSON answer. It uses
extract_last_json_object under the hood (last-block-wins, both fenced
and bare), and:
legal_action=None, raw_action=None → rethink loop asks for one.legal_action=None, raw_action=<what_the_model_said> → rethink loop shows the model what
it tried and asks it to pick a legal move.legal_action=<matched>, raw_action=<raw> →
submit.Do not write a prose-scan fallback. Coord regex, keyword regex,
response.rfind(legal), anything that picks a token mentioned in the
reasoning text — these silently substitute moves the model never
explicitly chose (almost always a rejected option from the prose). The
review-harness skill's "Ghost-fallback / prose-scan rescue" entry has
the empirical case: across 12 harnesses, 7,477 substitutions in one
game's replay archive touched 74% of episodes.
If your game uses a key other than "move", pass json_key=:
return parse_json_action(response, legal_action_strings, json_key="bid")The default matcher is case-insensitive and strips whitespace. For
games that need notation tolerance, alias handling, or canonicalization
(e.g. 'A7' → 'a7', 'b1-c2' → 'b1xc2'), pass matcher=:
def _match_move_to_legal(
raw: str, legal_action_strings: Sequence[str],
) -> str | None:
"""Game-specific normalization, e.g. canonicalize a coord."""
canonical = _canonicalize(raw)
return canonical if canonical in set(legal_action_strings) else None
def parse_response(
response: str, legal_action_strings: Sequence[str],
) -> ParseResult:
return parse_json_action(
response, legal_action_strings, matcher=_match_move_to_legal,
)The matcher operates on the raw value the model wrote inside the JSON, never the full response. This is the one place where game-specific parsing logic belongs.
Pre-flight check before shipping the default matcher. Print a handful
of state.action_to_string(...) outputs at representative states
(initial, mid-game, post-capture, multi-component turns). For every
non-alphanumeric marker that appears — * (backgammon hit), x
(checkers/chess capture), trailing Pass (backgammon per-die filler),
- (move separator), O-O (castling) — ask: "would a model naturally
omit or add this?" If yes, the default _default_match won't tolerate
it and the rethink loop pays for every drift. Backgammon's audit found
97.3% of episodes forfeited; adding just * and Pass tolerance via
matcher= recovered 34.2% of forfeit turns. Build the tolerant
normalization once, keep it in the matcher.
For free-form turns (where legal_action_strings is None), don't use
parse_json_action — write the extraction by hand using
extract_last_json_object, since the resulting object goes into
submission rather than legal_action:
import json
from kaggle_environments.core_harness import (
ParseResult, extract_last_json_object, parse_json_action,
)
def parse_response(response, legal_action_strings):
if legal_action_strings is None:
data = extract_last_json_object(response, required_keys=("clue",))
if data and "clue" in data:
return ParseResult(
submission=data,
raw_action=json.dumps(data),
)
return ParseResult(raw_action=response[:200])
return parse_json_action(response, legal_action_strings)When the model writes the structured answer surface more than once — a draft, then a revision — the last occurrence is the intent. Models almost always enumerate options ("considered a1, then b2, but going with e5") before stating the final answer. Forward iteration silently picks the rejected first candidate.
parse_json_action and extract_last_json_object already handle this
for JSON answers. If your harness uses a different stage-1 surface
(e.g. a tagged Final Answer: <x> line), apply the same rule:
| Pattern | Wrong | Right |
|---|---|---|
| Multiple JSON blocks | _JSON_BLOCK_RE.search(response) | parse_json_action(response, legal_action_strings) (preferred) or extract_last_json_object(response, required_keys=(...)) |
| Single-match regex extract for an answer tag | m = r.search(response) | matches = list(r.finditer(response)); m = matches[-1] if matches else None |
| Substring of a fixed answer tag | response.find("Final Answer:") | response.rfind("Final Answer:") |
Do not scan the response for unstructured candidates. Iterating a
regex (finditer / findall) over the response for coords or keywords,
or iterating legal_action_strings and checking response.rfind(legal),
is the prose-scan rescue antipattern even with the "reverse-iter" fix:
the model often discusses several options it doesn't choose, and any
loop over those mentions silently substitutes a non-chosen move. If the
structured answer is missing or illegal, return legal_action=None and
let the rethink loop ask the model to fix its format. See the
review-harness "Ghost-fallback / prose-scan rescue" entry for the
empirical impact.
parse_json_action over rolling your own. It's the audited single-stage default and removes the temptation to add a "smart" fallback later.legal_action=None and let the rethink loop ask for one. Prose-scan fallbacks reliably substitute moves the model never chose (rejected options it discussed in its reasoning).matcher= only when you need notation tolerance, alias handling, or canonicalization.legal_action_strings; the default matcher will pick them up.json_key=: "move" for board games (default), "bid" for auction games, "action" for generic games — whatever matches your prompt.Harness tests follow a consistent 4-class structure. Create your test file at the same relative path as the harness, under tests/.
For example:
kaggle_environments/envs/open_spiel_env/games/myg/harness.pytests/envs/open_spiel_env/games/myg/harness_test.pyCopy tests/envs/open_spiel_env/games/checkers/harness_test.py as the
starting point. It's the canonical layout: a _make_observation helper
that builds a harness-style obs dict from a proxy state, and four test
classes that together cover the surface area.
| Class | What it covers |
|---|---|
ParseResponseTest | Parser in isolation. Cover at minimum: fenced JSON, bare JSON, case-insensitive match, illegal-move-returns-raw, prose-only-returns-None (no ghost fallback), multiple-JSON-last-wins. Add a test_illegal_json_does_not_ghost_substitute_from_prose regression — the model writes a legal token in prose, then commits to an illegal one in JSON; the parser must NOT silently substitute the prose token. |
GeneratePromptTest | Prompt contents from a real proxy state. Cover: rules keywords present, board orientation, player-asymmetric text differs between player_id=0 and player_id=1, captures/phase flags render correctly, rethink suffixes appear under the right conditions, the JSON example format is unambiguous. If the harness has multi-branch prompts (roles/phases), assert each branch contains its required rules. |
GetLegalMovesTest | Round-trip from legalActions + legalActionStrings, fallback from serializedGameAndState, empty-obs returns {}. |
AgentIntegrationTest | Full harness through create_agent_fn with litellm.completion patched. Cover: successful move, retry-on-bad-parse, raise-after-two-failures, terminal-step-returns-inactive, and a short scripted game (first-legal-each-turn) that round-trips through pyspiel without raising. Define a small test-local _MyGameHarness adapter (see the snippet below) at the top of the test file and pass it to create_agent_fn; do NOT import an adapter from harness.py (there isn't one). |
Test-local adapter pattern (the only place an adapter class should live):
class _MyGameHarness:
"""Test-local GameHarness adapter; mirrors the prod wrapper shape."""
def get_legal_moves(self, observation):
return get_legal_moves(observation)
def make_prompt(self, observation, move_history,
previous_response=None, previous_action=None):
return generate_prompt(
observation, move_history, previous_response, previous_action,
)
def parse_response(self, response, legal_action_strings, *, observation=None):
# Most parsers ignore `observation`. If yours forwards it to a
# module-level parse_response that needs the env state, pass it
# through here.
return parse_response(response, legal_action_strings)The *, observation keyword on parse_response is required by the GameHarness protocol — core_harness always passes the current turn's observation. Most parsers can match the model's output against legal_action_strings alone and can ignore it; for parsers that genuinely need the env state at parse time (e.g. repeated_poker's bet-size soft-matching), declare an opt-in *, observation: Mapping[str, Any] | None = None kwarg on the module-level parse_response too and forward it through. See kaggle_environments/envs/open_spiel_env/games/repeated_poker/harness.py for that pattern.
Mock helpers (_StreamDelta, _StreamChoice, _StreamChunk,
_make_mock_response, _ENV) live at the top of checkers'
harness_test.py — copy them verbatim; they're game-independent.
test_llm_game.pyEvery harness should include a test_llm_game.py script that runs a
full game locally with a real LLM — catches issues mocked unit tests
miss (bad prompts, unparseable responses, env interaction bugs).
Use run_llm_game from kaggle_environments.local_harness_runner. The
helper handles API-key discovery, env-var defaults, --model /
--replay-path CLI flags, game execution, per-step printing, and
replay save. Per-game files are 3 lines plus a docstring:
"""Run a full Checkers game with LLM agents for local integration testing."""
from kaggle_environments.local_harness_runner import run_llm_game
if __name__ == "__main__":
run_llm_game("open_spiel_checkers", caller_file=__file__)For games that need extra config, pass configuration=,
replay_filename=, num_agents=, or agent_module=. See
havannah/test_llm_game.py (custom board size), word_association/test_llm_game.py
(4 agents, custom post-run output), and
python_ant_foraging/test_llm_game.py (custom replay filename) for
real examples. Capture the returned env if you need to print
game-specific results after the run.
Place it next to the harness module:
kaggle_environments/envs/open_spiel_env/games/<name>/test_llm_game.pykaggle_environments/envs/<name>/test_llm_game.pykaggle_environments/envs/open_spiel_env/games/<name>/
├── __init__.py
├── <name>_proxy.py # proxy (if not already created)
├── harness.py # <-- your harness
└── test_llm_game.py # local LLM integration test
tests/envs/open_spiel_env/games/<name>/
└── harness_test.pykaggle_environments/envs/<name>/
├── harness.py # <-- your harness
├── test_llm_game.py # local LLM integration test
└── ...
tests/envs/<name>/
└── harness_test.pyuv sync && uv run pytest tests/envs/open_spiel_env/games/<name>/harness_test.py -vget_legal_moves returns dict[int, str] or None as appropriateload_game params)state_dict() or reconstructed from serializedGameAndState), not just the per-agent move_history argument — and the prompt renders the actual moves (e.g. "a1b1, b3a3, ..."), not just a count like "Moves played so far: 14"player_id=0 and player_id=1 — directional/orientation language mirrors correctlyprevious_response and previous_action)parse_response delegates to parse_json_action (uses last-mention-wins JSON extraction; no prose-scan fallback — that's the ghost-fallback anti-pattern)parse_response does case-insensitive matching (default matcher) or passes a custom matcher= for notation tolerance (check state.action_to_string outputs for *, x, trailing Pass, -, etc. that models drop or add)ParseResult fields are set correctly (enumerable: legal_action; free-form: submission)test_llm_game.py script runs a full game with real LLM agentsuv run ruff check --fix . && uv run ruff format .| File | What to learn from it |
|---|---|
kaggle_environments/core_harness.py | The framework — GameHarness protocol, ParseResult, create_agent_fn, parse_json_action, render_rethink_suffix |
| Canonical templates — start here: | |
kaggle_environments/envs/open_spiel_env/games/checkers/harness.py | Modern enumerable shape: delegates to parse_json_action, branches render_rethink_suffix, demonstrates a phase-conditional prompt section (multi-jump continuation) |
kaggle_environments/envs/open_spiel_env/games/dark_hex/harness.py | Same modern shape with a custom matcher= callable for notation tolerance; per-player rendering for imperfect-information games |
kaggle_environments/envs/word_association/harness.py | Mixed free-form + enumerable harness (non-OpenSpiel) |
| Test patterns: | |
tests/envs/open_spiel_env/games/checkers/harness_test.py | Full enumerable test suite (parser stress, prompt assertions, create_agent_fn integration with mocked LLM) |
tests/core_harness_test.py | Framework test patterns |
kaggle_environments/envs/open_spiel_env/games/checkers/test_llm_game.py | LLM integration test script (OpenSpiel) |
kaggle_environments/envs/word_association/test_llm_game.py | LLM integration test script (non-OpenSpiel) |
© Kaggle, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/create-harness of Kaggle/kaggle-environments.
Open the folder on GitHubat commit 1c8acf1
Create Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Create Harness this skillKaggle/kaggle-environments | 454 | — | ~8.9k | Automated safety check: Pass | Apache-2.0 | |
| Annotation DataSharpAI/DeepCamera | 3.1k | — | ~495 | Automated safety check: Pass | MIT | |
| Dataset FinderLeoYeAI/openclaw-master-skills | 2.2k | — | ~5.4k | Automated safety check: Pass | Proprietary | |
| Kaggle Finetunemajiayu000/claude-skill-registry | 666 | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Lilly Community Researchssaaffaakk/Lilly | 171 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Dataset Finder Guidewentorai/research-plugins | 298 | 1 repos | ~2.2k | Automated safety check: Pass | MIT |
SharpAI/DeepCamera
Dataset annotation management — COCO labels, sequences, export, and Kaggle upload
LeoYeAI/openclaw-master-skills
A skill your agent uses when users need to search for datasets, download data files, or explore data repositories.
majiayu000/claude-skill-registry
End-to-end workflow for fine-tuning LLMs using Kaggle datasets.
ssaaffaakk/Lilly
Lilly community-research skill. An agent skill from ssaaffaakk/Lilly.
wentorai/research-plugins
Search and download research datasets from Kaggle, HuggingFace, and repos
FrankS-IntelLab/agentic-kaggle-skill
Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.
Kaggle/kaggle-environments
Audit a game environment's README.md and AGENTS.md against its engine implementation.
Kaggle/kaggle-environments
Review an existing LLM harness for correctness and gameplay-impacting bugs.
Kaggle/kaggle-environments
Run a prompt-ablation study on a kaggle-environments game's LLM harness.
Works with
Categories
Create or update an LLM harness that lets a language model play a kaggle-environments game. Create Harness is an agent skill from Kaggle/kaggle-environments. Create or update an LLM harness that lets a language model play a kaggle-environments game.
Create Harness fits situations like: the user wants to write a harness; game-playing prompt for any kaggle-environments game — including OpenSpiel games; custom environments; the user mentions coreharness.py.
Run `npx skills add Kaggle/kaggle-environments --skill create-harness -a claude-code`. Or copy the skill folder (.agents/skills/create-harness in Kaggle/kaggle-environments) into .claude/skills/create-harness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Kaggle/kaggle-environments --skill create-harness -a codex`. Or copy the skill folder (.agents/skills/create-harness in Kaggle/kaggle-environments) into .agents/skills/create-harness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kaggle/kaggle-environments --skill create-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-harness, .gemini/skills/create-harness, .github/skills/create-harness and .opencode/skills/create-harness in your project.
Going by SKILL.md and its folder, Create Harness needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Create Harness is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.9k tokens (SKILL.md is roughly 35k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Create Harness: Annotation Data (SharpAI/DeepCamera, 3.1k stars), Dataset Finder (LeoYeAI/openclaw-master-skills, 2.2k stars), Kaggle Finetune (majiayu000/claude-skill-registry, 666 stars) and Lilly Community Research (ssaaffaakk/Lilly, 171 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Kaggle (a GitHub organization) maintains it in Kaggle/kaggle-environments, which has 454 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 8, 2026.
Source: Kaggle/kaggle-environments on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.