Agent skill

Ear Test Scoresheet

by scragnog in scragnog/HOT-Step-CPP

The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…

MITAuto-check passedMarketing & SEO

Install Ear Test Scoresheet

skills CLI
$ npx skills add scragnog/HOT-Step-CPP --skill ear-test-scoresheet -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scragnog/HOT-Step-CPP ear-test-scoresheet --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ear-test-scoresheet .claude/skills/ear-test-scoresheet && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ear-test-scoresheet
GitHub stars
174
Token cost
~1.9k tokens
SKILL.md length
839 words
Files
4
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…

  • Works in 4 steps: Lay out the study folder → Write study.json and generate the page → Hand it to Rob → …
  • Renders need judging by ear - checkpoint ladders
  • SKILL.md covers Running a test, What the first study taught…, Rendering ladders for YuE2 and Traps, plus 1 more section
  • Runs JavaScript scripts from its folder; calls node

What it does

Ear Test Scoresheet is an agent skill from scragnog/HOT-Step-CPP. The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score groups by rung so the point where likeness and quality cross is visible. Use whenever renders need judging by ear - checkpoint ladders, recipe A/B tests, sampler or quant comparisons - and whenever you are about to ask Rob to "listen to these files and tell me".

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files.

It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: Turn dials. Summon bangers! NOW WITH MORE C++! Local AI music generation powered by GGML. The licence is MIT.

When your agent uses it

  • Renders need judging by ear - checkpoint ladders
  • Recipe A/B tests
  • Quant comparisons - and whenever you are about to ask Rob to listen to these files and tell me

Example prompts

  • “listen to these files and tell me”
  • “/ear-test-scoresheet”

Requirements

  • Node.js

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Lay out the study folder
  2. Write study.json and generate the page
  3. Hand it to Rob
  4. Read the results

What it can do on your machine

Read from SKILL.md and the folder at commit eeeded6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ear Test Scoresheet loads about 1.9k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 839 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scragnog/HOT-Step-CPP at commit eeeded6, republished under its MIT licence (© scragnog). 839 words, ~1,914 tokens.

Download SKILL.mdSave it as .claude/skills/ear-test-scoresheet/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
ear-test-scoresheet
description
The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score groups by rung so the point where likeness and quality cross is visible. Use whenever renders need judging by ear - checkpoint ladders, recipe A/B tests, sampler or quant comparisons - and whenever you are about to ask Rob to "listen to these files and tell me".

Ear-test score sheet

Rob's verdict (2026-09-23): "the best way so far we've had to score tests like this". Use it instead of asking for free-text impressions of a folder of WAVs.

What it gives you:

  • Every render plays inside the page, straight from disk (WAV is fine).
  • Scores 1-5 on named criteria, a Keep / Borderline / Over the line verdict, a notes box per render, and a pick per group (the "last good rung").
  • A chart per group that averages the criteria into two lines, one for what you're pushing for (likeness) and one for what breaks (quality), plotted by rung.
  • Rows for renders that don't exist yet say "Rendering…" and fill in on their own when the file appears, so Rob can start while the rest render.
  • Scores save to scores.json beside the page, which is how you read them.

Files in this folder:

  • template.html — the page (local file; the default).
  • make-scoresheet.mjs — writes index.html from a study.json.
  • template-artifact.html — the claude.ai-hosted variant, only for when Rob can't open files on this machine (see the end).

Running a test

1. Lay out the study folder

Put renders under _experiments/_LISTENING/<date>-<study>/ (the listening-hub convention), e.g. round1/<group>/NN-<label>.wav. The page lives at the top of that folder and refers to tracks by relative path.

2. Write study.json and generate the page
json
{
  "id": "yue2-decoder-long-2026-09-23",
  "title": "YuE2 Decoder Ladder",
  "intro": "Play each render, score it 1–5 on each criterion. Scores save as you go.",
  "criteria": [
    {"key": "voice", "name": "Voice", "desc": "the singer sounds like the artist", "series": "likeness"},
    {"key": "writing", "name": "Songwriting", "desc": "melodies, hooks and structure feel like theirs", "series": "likeness"},
    {"key": "sound", "name": "Sound", "desc": "guitar, drum and production tone match the album", "series": "likeness"},
    {"key": "diction", "name": "Diction", "desc": "every word intelligible, nothing garbled", "series": "quality"},
    {"key": "audio", "name": "Audio", "desc": "clean: no hiss, phasing, crackle or clipping", "series": "quality"},
    {"key": "coherence", "name": "Coherence", "desc": "holds together, no loops, ends properly", "series": "quality"}
  ],
  "series": {"likeness": {"name": "Likeness", "color": "var(--accent)"},
             "quality":  {"name": "Quality",  "color": "var(--warn)"}},
  "roundNames": {"1": "Round 1 · decoder sweep"},
  "axis": {"1": "Decoder step"},
  "pickLabel": "Last good rung",
  "tracks": [
    {"id": "r1-rbf-01", "round": 1, "group": "rbf_whyrockhard", "order": 1,
     "label": "Base", "sublabel": "no adapter", "reference": true, "file": "round1/rbf_whyrockhard/01-base.wav"},
    {"id": "r1-rbf-02", "round": 1, "group": "rbf_whyrockhard", "order": 2,
     "label": "Decoder 300", "sublabel": "planner frozen at 240", "x": 300, "file": "round1/rbf_whyrockhard/02-nar300.wav"}
  ]
}
node .claude/skills/ear-test-scoresheet/make-scoresheet.mjs <study-folder>/study.json
  • group becomes a tab; x is the chart position; order sorts rows.
  • The six criteria above (three per series) are the proven set: about 40 s of scoring per render. Keep two series, "pushing for" and "what breaks".
  • reference: true marks a control (base model, no adapter): scorable, but off the chart and out of the picker. Always include one; it anchors the scale.
  • List every planned render up front, rendered or not. To add a round later, append tracks and re-run the generator. Scores are keyed by track id, so they survive.
3. Hand it to Rob

Give him the path to index.html. On first use he clicks Save scores to a file… and saves scores.json next to the page. After that every change writes to it; the browser remembers the file and asks once per session to reconnect. Scores are also kept in the browser either way. Browsers without file saving (Firefox) get Export scores instead: he exports and drops scores.json in the folder.

4. Read the results

scores.json: {scores: {<track id>: {<criterion>: 1-5, verdict?, note?}}, picks: {"r<round>-<group>": {lastGood, note}}}.

Scores alone are enough. On the first study Rob skipped verdicts and picks and the criteria still answered the question. Average each series per rung, lay the groups side by side, and read the trend, not single rungs.

What the first study taught (read before interpreting)

  • Renders are not reproducible from the seed. The same checkpoints and seed through /api/generate gave different songs of different lengths. Every rung is a fresh take, so plan two renders per rung when a decision rests on it.
  • Rob's scoring noise is about ±1 per criterion. Measured by putting the same combination in two rounds without saying so. Treat a difference under ~0.5 on a series average as noise. A single low rung between two good ones is noise, not a line.
  • Keep automated metrics off the page. They bias the ear. Compare them afterwards; the forced-aligner diction score correlated only 0.37 with Rob's diction scores.
  • Note boxes carry the surprises ("there are no vocals in this track"). Read every note.
Show full SKILL.md (275 more words)Show less

Rendering ladders for YuE2

server/scripts/yue2-ladder.mjs <config.json> renders planner (AR) and decoder (NAR) checkpoints from different steps through the app's own generate path, scores diction with /yue2/align, writes numbered WAVs plus results.jsonl, and restores the user's adapter picks at the end. It skips WAVs that already exist, so re-running after a failure only fills the gaps. Config: {runDir, outDir, caption, lyrics, seed, pairs: [[arStep, narStep], ...]}; [0, 0] is the base-model reference.

Traps

  • /api/generate needs a bearer token: GET /api/auth/auto returns one. /api/backends/models does not, so a script can change the picks and then fail to render. The ladder script handles both.
  • Editing server/src restarts the dev server and kills any running training batch. Put tools in server/scripts/ (tracked, not watched); tools/yue2-* is gitignored.
  • Watch free space on D: before a long ladder. A full disk cut one render off mid-write (a 0-byte WAV); delete it and re-run the ladder.
  • Prompts: use Lyric Studio lyrics the adapter never trained on, with their own caption (generations table in server/data/hotstep.db: caption, lyrics). The same prompt and seed for every rung of a group.

Hosted variant (only when needed)

template-artifact.html is the same page published as a claude.ai artifact, for scoring away from this machine. It keeps rows and scores in the artifact's database (capabilities: {"db": {}, "assets": {}}; config in a CONFIG block at the top of its script; rows seeded with ArtifactData into tracks). Audio must be uploaded: the asset store refuses .wav/.mp3 but takes an audio-only AAC .mp4 (ffmpeg -i in.wav -vn -codec:a aac -b:a 256k -movflags +faststart out.mp4), then each track gets {url: "/_blob/<id>"} via a pinned (if_version) ArtifactData batch update. Much more work than the local page.

© scragnog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in .claude/skills/ear-test-scoresheet of scragnog/HOT-Step-CPP.

  • SKILL.md
  • make-scoresheet.mjs
  • template-artifact.html
  • template.html

Open the folder on GitHubat commit eeeded6

Compare with similar skills

Ear Test Scoresheet next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ear Test Scoresheet compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ear Test Scoresheet this skillscragnog/HOT-Step-CPP174—~1.9kAutomated safety check: PassMIT
Ab Testingcoreyhaines31/marketingskills54k3 repos~3.1kAutomated safety check: PassMIT
AnalyticsNexus-JPF/note-companion8707 repos~2.2kAutomated safety check: PassMIT
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab Test Analyzeririnabuht12-oss/marketing-skills4.1k—~1.4kAutomated safety check: PassNone
Ab Test Store Listingappeeky/aso-skills2.2k—~1.8kAutomated safety check: PassMIT

Similar skills

  • Ab Testing

    coreyhaines31/marketingskills

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

    54k GitHub starsUsed in 3 repos~3.1k tokens
    Marketing & SEOAuto-check passed
  • Analytics

    Nexus-JPF/note-companion

    When the user wants to set up, improve, or audit analytics tracking and measurement.

    870 GitHub starsUsed in 7 repos~2.2k tokens
    Marketing & SEOAuto-check passed
  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    4.1k GitHub stars~1.4k tokensUpdated 17 days ago
    Marketing & SEOAuto-check passed
  • Ab Test Store Listing

    appeeky/aso-skills

    When the user wants to A/B test App Store product page elements to improve conversion rate.

    2.2k GitHub stars~1.8k tokensUpdated 4 days ago
    Marketing & SEOAuto-check passed
  • Ab Test Setup

    freekmurze/dotfiles

    When the user wants to plan, design, or implement an A/B test or experiment.

    1k GitHub starsUsed in 14 repos~1.8k tokens
    Marketing & SEOAuto-check passed

More from scragnog/HOT-Step-CPP

All 18 skills in this repo
  • Engine Performance

    scragnog/HOT-Step-CPP

    Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.

    174 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed
  • Mm3 Backend

    scragnog/HOT-Step-CPP

    Maps HOT-Step's native MiniMax-Music3 backend — engine port modules, endpoints, server/UI integration, parity/fixture infrastructure, and the hard-won trap list.

    174 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • Mm3 Lm Adapter Training

    scragnog/HOT-Step-CPP

    The validated recipe for training MiniMax-Music3 planner-LM style adapters (artist/album clones) with ace-train mm3-lm-train and the Training Studio.

    174 GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Release Process

    scragnog/HOT-Step-CPP

    Runbook for cutting and publishing a HOT-Step CPP release via a v git tag that triggers the multi-platform CI build and drafts a GitHub Release.

    174 GitHub stars~5.1k tokensUpdated 2 days ago
    Auto-check passed
  • Upstream Sync

    scragnog/HOT-Step-CPP

    Safely pulls upstream acestep.cpp changes into the HOT-Step engine fork without destroying its integration hooks.

    174 GitHub stars~5k tokensUpdated 2 days ago
    Auto-check passed
  • Debugging Runtime

    scragnog/HOT-Step-CPP

    Diagnoses HOT-Step CPP generation failures, engine crashes, hangs, and startup problems from the logs/ session folders.

    174 GitHub stars~6k tokensUpdated 2 days ago
    Auto-check: notes

Categories

Questions about Ear Test Scoresheet

What does Ear Test Scoresheet do?

The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…. Ear Test Scoresheet is an agent skill from scragnog/HOT-Step-CPP. The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score groups by rung so the point where likeness and quality cross is visible.

When should I use Ear Test Scoresheet?

Ear Test Scoresheet fits situations like: renders need judging by ear - checkpoint ladders; recipe A/B tests; quant comparisons - and whenever you are about to ask Rob to listen to these files and tell me.

How do I install Ear Test Scoresheet in Claude Code?

Run `npx skills add scragnog/HOT-Step-CPP --skill ear-test-scoresheet -a claude-code`. Or copy the skill folder (.claude/skills/ear-test-scoresheet in scragnog/HOT-Step-CPP) into .claude/skills/ear-test-scoresheet in your project. Claude Code loads it when a task matches its description.

How do I install Ear Test Scoresheet in Codex?

Run `npx skills add scragnog/HOT-Step-CPP --skill ear-test-scoresheet -a codex`. Or copy the skill folder (.claude/skills/ear-test-scoresheet in scragnog/HOT-Step-CPP) into .agents/skills/ear-test-scoresheet in your project. Codex loads it when a task matches its description.

Can I use Ear Test Scoresheet in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scragnog/HOT-Step-CPP --skill ear-test-scoresheet -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ear-test-scoresheet, .gemini/skills/ear-test-scoresheet, .github/skills/ear-test-scoresheet and .opencode/skills/ear-test-scoresheet in your project.

What does Ear Test Scoresheet need to run?

Going by SKILL.md and its folder, Ear Test Scoresheet needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node). Our summary lists: Node.js.

Does Ear Test Scoresheet access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ear Test Scoresheet safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ear Test Scoresheet use?

Ear Test Scoresheet is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ear Test Scoresheet use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ear Test Scoresheet?

Skills that share tags, products or a category with Ear Test Scoresheet: Ab Testing (coreyhaines31/marketingskills, 54k stars), Analytics (Nexus-JPF/note-companion, 870 stars), Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars) and Ab Test Analyzer (irinabuht12-oss/marketing-skills, 4.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ear Test Scoresheet?

scragnog (a GitHub user) maintains it in scragnog/HOT-Step-CPP, which has 174 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 9, 2026.

Source: scragnog/HOT-Step-CPP on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.