Agent skill

Gm Evaluate

by RandallLiuXin in RandallLiuXin/GodotMaker

Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for…

Custom licenceAuto-check passedTesting & QA

Install Gm Evaluate

skills CLI
$ npx skills add RandallLiuXin/GodotMaker --skill gm-evaluate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RandallLiuXin/GodotMaker gm-evaluate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RandallLiuXin/GodotMaker.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/core/gm-evaluate .claude/skills/gm-evaluate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gm-evaluate
GitHub stars
550
Token cost
~4.7k tokens
SKILL.md length
2,092 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
Custom licence

At a glance

Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for…

  • Works in 5 steps: Understand Requirements → Maintain the e2e/ suite → Mandatory Checks → …
  • Tasks that involve End-to-end testing
  • SKILL.md covers Session Setup, Resume Check, Resolve godot binary and Evaluation Process, plus 2 more sections
  • Calls git and python

What it does

Gm Evaluate is an agent skill from RandallLiuXin/GodotMaker. Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for mechanics this tag deliberately removed), and reason about gameplay quality. Independent from the build process — fresh perspective on the final product. Explicit invocation only — use /gm-evaluate.

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing and Game development. The repository describes itself as: Autonomous text-to-game pipeline for Godot, powered by Claude Code,Codex,Opencode.

When your agent uses it

  • Tasks that involve End-to-end testing
  • Tasks that involve Game development

Example prompts

  • “/gm-evaluate”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Understand Requirements
  2. Maintain the e2e/ suite
  3. Mandatory Checks
  4. Gameplay Reasoning
  5. Final Assessment (Pass/Fail)

What it can do on your machine

Read from SKILL.md and the folder at commit 1d4702c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gm Evaluate loads about 4.7k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 2,092 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 2,092 words (~4,656 tokens).

name
gm-evaluate
disable-model-invocation
true

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/core/gm-evaluate of RandallLiuXin/GodotMaker.

Open the folder on GitHubat commit 1d4702c

Compare with similar skills

Gm Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gm Evaluate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gm Evaluate this skillRandallLiuXin/GodotMaker550—~4.7kAutomated safety check: PassCustom licence
Test Playable Web Gamesnirholas/three.ws2291 repos~429Automated safety check: PassApache-2.0
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
Uloop Replay Inputkurotu/VRCQuestTools3733 repos~615Automated safety check: PassMIT
Ui4 Convert Testspayloadcms/payload45k—~3.5kAutomated safety check: PassMIT

Similar skills

  • Test Playable Web Games

    nirholas/three.ws

    Test a playable browser game end to end with deterministic fixtures and real browser evidence.

    229 GitHub starsUsed in 1 repo~429 tokens
    Testing & QAAuto-check passed
  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • Uloop Replay Input

    kurotu/VRCQuestTools

    Replay recorded PlayMode keyboard and mouse input. An agent skill from kurotu/VRCQuestTools.

    373 GitHub starsUsed in 3 repos~615 tokens
    Testing & QAAuto-check passed
  • Ui4 Convert Tests

    payloadcms/payload

    A skill your agent uses when UI changes are complete and e2e tests need updating.

    45k GitHub stars~3.5k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • E2E

    stackia/rtp2httpd

    Write, run, review, or debug rtp2httpd E2E tests and their harness in e2e/ and scripts/run-e2e.sh.

    2.2k GitHub stars~517 tokensUpdated 8 days ago
    Testing & QAAuto-check passed

More from RandallLiuXin/GodotMaker

All 41 skills in this repo
  • Godot E2E

    RandallLiuXin/GodotMaker

    Write and run E2E (end-to-end) game tests using the godot-e2e framework.

    550 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • MCP Driver

    RandallLiuXin/GodotMaker

    Runtime debugging and live project inspection via godot-mcp.

    550 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Gecs

    RandallLiuXin/GodotMaker

    gecs ECS framework API reference — the Entity-Component-System addon for Godot 4.x used by this project.

    550 GitHub stars~2.7k tokensUpdated 23 days ago
    Auto-check passed
  • Gdtoolkit

    RandallLiuXin/GodotMaker

    Lint and format GDScript files using gdtoolkit (gdlint + gdformat).

    550 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Project Scaffold

    RandallLiuXin/GodotMaker

    Generate a new Godot game project with standardized ECS structure and tooling.

    550 GitHub stars~2.1k tokensUpdated 23 days ago
    Auto-check passed
  • Gdunit Driver

    RandallLiuXin/GodotMaker

    Run gdUnit4 unit tests and parse results into structured output.

    550 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed

Questions about Gm Evaluate

What does Gm Evaluate do?

Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for…. Gm Evaluate is an agent skill from RandallLiuXin/GodotMaker. Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for mechanics this tag deliberately removed), and reason about gameplay quality.

When should I use Gm Evaluate?

Gm Evaluate fits situations like: tasks that involve End-to-end testing; tasks that involve Game development.

How do I install Gm Evaluate in Claude Code?

Run `npx skills add RandallLiuXin/GodotMaker --skill gm-evaluate -a claude-code`. Or copy the skill folder (skills/core/gm-evaluate in RandallLiuXin/GodotMaker) into .claude/skills/gm-evaluate in your project. Claude Code loads it when a task matches its description.

How do I install Gm Evaluate in Codex?

Run `npx skills add RandallLiuXin/GodotMaker --skill gm-evaluate -a codex`. Or copy the skill folder (skills/core/gm-evaluate in RandallLiuXin/GodotMaker) into .agents/skills/gm-evaluate in your project. Codex loads it when a task matches its description.

Can I use Gm Evaluate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RandallLiuXin/GodotMaker --skill gm-evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gm-evaluate, .gemini/skills/gm-evaluate, .github/skills/gm-evaluate and .opencode/skills/gm-evaluate in your project.

What does Gm Evaluate need to run?

Going by SKILL.md and its folder, Gm Evaluate needs the command-line tools its instructions call (git and python).

Does Gm Evaluate access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Gm Evaluate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gm Evaluate use?

Gm Evaluate has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Gm Evaluate use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gm Evaluate?

Skills that share tags, products or a category with Gm Evaluate: Test Playable Web Games (nirholas/three.ws, 229 stars), Web Application Testing (anthropics/skills, 180k stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars) and Uloop Replay Input (kurotu/VRCQuestTools, 373 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gm Evaluate?

RandallLiuXin (a GitHub user) maintains it in RandallLiuXin/GodotMaker, which has 550 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on September 17, 2026.

Source: RandallLiuXin/GodotMaker on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.