Agent skill

Pokemon Goal Benchmark

by shepherdjerred in shepherdjerred/monorepo

Run and interpret the real-model Pokémon Emerald goal benchmark.

GPL-3.0Auto-check passedDevelopment

Install Pokemon Goal Benchmark

skills CLI
$ npx skills add shepherdjerred/monorepo --skill pokemon-goal-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install shepherdjerred/monorepo pokemon-goal-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/shepherdjerred/monorepo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/discord-plays-pokemon/.agents/skills/pokemon-goal-benchmark .claude/skills/pokemon-goal-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pokemon-goal-benchmark
GitHub stars
112
Token cost
~1.2k tokens
SKILL.md length
522 words
Files
2
Skills in repo
63
Repo updated
First seen
Licence
GPL-3.0

At a glance

Run and interpret the real-model Pokémon Emerald goal benchmark.

  • Works in 4 steps: Select the exact implementation commit… → Copy the source Emerald save outside the… → Build or obtain the target's… → …
  • Measuring goal-agent reliability
  • SKILL.md covers Prepare a clean measurement, Run, Read the evidence and Interpret exit status
  • Calls bun, mise and bunx

What it does

Pokemon Goal Benchmark is an agent skill from shepherdjerred/monorepo. Run and interpret the real-model Pokémon Emerald goal benchmark. Use when measuring goal-agent reliability, comparing baseline and candidate runs, diagnosing benchmark artifacts, or deciding whether a run is a valid game failure versus invalid provider or harness evidence.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Development. It works with WebAssembly. The repository describes itself as: Monorepo for all of my projects. The licence is GPL-3.0.

When your agent uses it

  • Measuring goal-agent reliability
  • Comparing baseline and candidate runs
  • Diagnosing benchmark artifacts
  • Deciding whether a run is a valid game failure versus invalid provider

Example prompts

  • “/pokemon-goal-benchmark”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Select the exact implementation commit and prepare a clean detached
  2. Copy the source Emerald save outside the target checkout. Require exactly
  3. Build or obtain the target's pokeemerald.wasm. From the package root,
  4. Choose a new output path outside the target checkout. The harness refuses

What it can do on your machine

Read from SKILL.md and the folder at commit d66c497. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun
    • mise
    • bunx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use bunx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pokemon Goal Benchmark loads about 1.2k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 522 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from shepherdjerred/monorepo at commit d66c497, republished under its GPL-3.0 licence (© shepherdjerred). 522 words, ~1,172 tokens.

Download SKILL.mdSave it as .claude/skills/pokemon-goal-benchmark/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pokemon-goal-benchmark
description
Run and interpret the real-model Pokémon Emerald goal benchmark. Use when measuring goal-agent reliability, comparing baseline and candidate runs, diagnosing benchmark artifacts, or deciding whether a run is a valid game failure versus invalid provider or harness evidence.

Pokémon Goal Benchmark

Measure general gameplay with the strict catch evaluator. Do not turn the benchmark into a scripted solver or count model prose as success.

Prepare a clean measurement

  1. Select the exact implementation commit and prepare a clean detached worktree or clone. Install the root toolchain and workspace dependencies and run generation before measuring:

    bash
    mise install
    bun install --frozen-lockfile
    bunx turbo run generate
  2. Copy the source Emerald save outside the target checkout. Require exactly 131,072 bytes; remove a trailing 16-byte RTC block from the copy, not the live save. Require a decodable active slot and at least one empty party slot so a catch creates independent party-identity evidence. Treat this source copy as immutable.

  3. Build or obtain the target's pokeemerald.wasm. From the package root, bun scripts/build-wasm.ts writes packages/backend/assets/pokeemerald.wasm. Confirm that it belongs to the implementation being measured: the harness hashes the external file but records its relationship to wasm-src/upstream.json as not-verified.

  4. Choose a new output path outside the target checkout. The harness refuses an existing output directory and requires both the benchmark runner and target implementation worktrees to be clean.

Run

From packages/discord-plays-pokemon/packages/backend in the clean candidate:

bash
bun run benchmark:goal \
  --save /absolute/path/to/copied-emerald.sav \
  --wasm ./assets/pokeemerald.wasm \
  --output /absolute/path/to/artifacts/candidate-<commit> \
  --runs 3 \
  --goal "get me a pokeman"

Use --implementation-root /path/to/clean/copy when the harness runner and target are separate. It accepts either the monorepo root or packages/discord-plays-pokemon. The target must be clean and fully set up. For comparable repeated trials, keep the immutable source save, goal, model, reasoning, runtime, and target WASM fixed. Each run receives a fresh harness-owned save copy and a distinct control port.

Run bun run benchmark:goal --help for model, reasoning, runtime, port, Codex binary, and boot-timeout options.

Show full SKILL.md (266 more words)Show less

Read the evidence

Read the printed summary.json first. It is written only after the run series:

  • completedRuns, validRuns, successfulRuns, failedRuns, and invalidRuns show the measurement population.
  • successRate uses valid runs only. Never include provider or harness failures in its denominator.
  • stoppedEarly plus stopReason: external-provider-failure means a provider failure stopped the remaining requested runs.
  • allSucceeded is true only when every requested run completed successfully.

Then inspect each run-NNN/result.json:

  • outcome is the authoritative classification.
  • For success or game-failure, inspect evaluation.evidence, evaluation.failures, and evaluation.verifiedCaughtSpecies. Success requires one exact caught species to correlate across a post-start event, post-event state, final live state, and a complete save persisted after the catch.
  • For invalid-provider, inspect providerFailure and provider-failure.json. This is external provider evidence, not gameplay evidence; evaluation is null.
  • For harness-error, inspect error, lifecycle fields, and worker logs. Treat the run as invalid and rerun only after fixing the cause.
  • Use telemetry for turns, controls, repeated-position loops, ignored inputs, screenshots, knowledge queries, tokens, and cost. Use provenance to compare source-save, WASM, target, runner, evaluator, and Codex identities.

Preserve input.flash, persisted.flash, codex.jsonl, screenshots, worker logs, and any provider-failure files. Do not claim success from the last Codex message, process exit, screenshots, or a catch event alone.

Interpret exit status

  • 0: every requested run passed the strict evaluator.
  • 1: at least one valid run was a game-failure, with no invalid run.
  • 2: at least one invalid-provider or harness-error, or command arguments/preflight failed. Preflight failures may occur before summary.json exists.

Count only success and game-failure as valid model measurements. Repeat invalid runs after resolving provider or harness faults; do not relabel them as game failures.

© shepherdjerred, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in packages/discord-plays-pokemon/.agents/skills/pokemon-goal-benchmark of shepherdjerred/monorepo.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit d66c497

Compare with similar skills

Pokemon Goal Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pokemon Goal Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pokemon Goal Benchmark this skillshepherdjerred/monorepo112—~1.2kAutomated safety check: PassGPL-3.0
Connection Import WasmfeigeCode/navop1.8k—~2.7kAutomated safety check: PassCustom licence
Changesetwhitphx/stlite1.7k—~1.9kAutomated safety check: PassApache-2.0
Debug Php Wasm Main ModuleWordPress/wordpress-playground2k—~3.1kAutomated safety check: PassGPL-2.0
Code Reviewziggy42/epsilon439—~1.7kAutomated safety check: PassApache-2.0
Debug Php Wasm Side ModulesWordPress/wordpress-playground2k—~2.5kAutomated safety check: PassGPL-2.0

Similar skills

  • Connection Import Wasm

    feigeCode/navop

    A skill your agent uses when implementing, debugging, packaging, or host-enabling onetcli WASM connection importers such as DBeaver, Navicat, Navicat Lite, Termius, connection-import.wit components…

    1.8k GitHub stars~2.7k tokensUpdated today
    DevelopmentAuto-check passed
  • Changeset

    whitphx/stlite

    Create or update a changeset fragment (.changeset/.md) reflecting the changes made in the current session or branch.

    1.7k GitHub stars~1.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Debug Php Wasm Main Module

    WordPress/wordpress-playground

    Debug PHP.wasm main module crashes including Asyncify errors (unreachable, memory access out of bounds), JSPI errors (SuspendError, trying to suspend JS frames), WASM memory growth bugs, and runtime…

    2k GitHub stars~3.1k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Code Review

    ziggy42/epsilon

    A skill your agent uses when the user asks for a code review.

    439 GitHub stars~1.7k tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Debug Php Wasm Side Modules

    WordPress/wordpress-playground

    Debug WASM side modules (dynamic PHP extensions) including dlopen failures, SIDEMODULE loading, JSPI suspension crashes in extensions, C++ weak symbol issues, and extension runtime errors.

    2k GitHub stars~2.5k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Merjs

    justrach/merjs

    Work with the merjs Zig web framework. An agent skill from justrach/merjs.

    356 GitHub stars~4k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes

More from shepherdjerred/monorepo

All 63 skills in this repo
  • Bun Runtime Best Practices

    shepherdjerred/monorepo

    Bun runtime APIs and current operational patterns for files, processes, modules, networking, databases, tests, and deployment.

    112 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Bun Test Patterns

    shepherdjerred/monorepo

    Current Bun test runner guidance for discovery, isolation, parallelism, sharding, changed tests, mocks, timers, snapshots, coverage, DOM Testing Library, and integration teardown.

    112 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Bun Workspaces

    shepherdjerred/monorepo

    Current Bun workspace guidance for isolated and hoisted linkers, catalogs, filters, scripts, dependency classes, lockfiles, lifecycle trust, caches, publishing, TypeScript package exports, and…

    112 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Deep Research

    shepherdjerred/monorepo

    This skill should be used when the user asks to "deep research", "research this topic", "investigate thoroughly", "do a deep dive on", "comprehensive research on", "find everything about", "survey…

    112 GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Figma Use

    shepherdjerred/monorepo

    This skill should be used when the user asks to "create a Figma design", "design in Figma", "make a Figma mockup", "create an app icon", "design UI", "render JSX to Figma", "export from Figma"…

    112 GitHub stars~946 tokensUpdated today
    Auto-check passed
  • Fish Helper

    shepherdjerred/monorepo

    Current Fish shell scripting, functions, abbreviations, completions, variables, events, configuration, plugins, testing, and safety guidance.

    112 GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Pokemon Goal Benchmark

What does Pokemon Goal Benchmark do?

Run and interpret the real-model Pokémon Emerald goal benchmark. Pokemon Goal Benchmark is an agent skill from shepherdjerred/monorepo. Run and interpret the real-model Pokémon Emerald goal benchmark.

When should I use Pokemon Goal Benchmark?

Pokemon Goal Benchmark fits situations like: measuring goal-agent reliability; comparing baseline and candidate runs; diagnosing benchmark artifacts; deciding whether a run is a valid game failure versus invalid provider.

How do I install Pokemon Goal Benchmark in Claude Code?

Run `npx skills add shepherdjerred/monorepo --skill pokemon-goal-benchmark -a claude-code`. Or copy the skill folder (packages/discord-plays-pokemon/.agents/skills/pokemon-goal-benchmark in shepherdjerred/monorepo) into .claude/skills/pokemon-goal-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Pokemon Goal Benchmark in Codex?

Run `npx skills add shepherdjerred/monorepo --skill pokemon-goal-benchmark -a codex`. Or copy the skill folder (packages/discord-plays-pokemon/.agents/skills/pokemon-goal-benchmark in shepherdjerred/monorepo) into .agents/skills/pokemon-goal-benchmark in your project. Codex loads it when a task matches its description.

Can I use Pokemon Goal Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shepherdjerred/monorepo --skill pokemon-goal-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pokemon-goal-benchmark, .gemini/skills/pokemon-goal-benchmark, .github/skills/pokemon-goal-benchmark and .opencode/skills/pokemon-goal-benchmark in your project.

What does Pokemon Goal Benchmark need to run?

Going by SKILL.md and its folder, Pokemon Goal Benchmark needs the command-line tools its instructions call (bun, mise and bunx).

Does Pokemon Goal Benchmark access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Pokemon Goal Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pokemon Goal Benchmark use?

Pokemon Goal Benchmark is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pokemon Goal Benchmark use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pokemon Goal Benchmark?

Skills that share tags, products or a category with Pokemon Goal Benchmark: Connection Import Wasm (feigeCode/navop, 1.8k stars), Changeset (whitphx/stlite, 1.7k stars), Debug Php Wasm Main Module (WordPress/wordpress-playground, 2k stars) and Code Review (ziggy42/epsilon, 439 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pokemon Goal Benchmark?

shepherdjerred (a GitHub user) maintains it in shepherdjerred/monorepo, which has 112 GitHub stars. The repository holds 63 skills in this directory. The repository was last updated on October 10, 2026.

Source: shepherdjerred/monorepo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.