Agent skill

Listening To Music

by oaustegard in oaustegard/claude-skills

Listen to generated music by measuring it and by reading it as sheet music: render Strudel code or record a Web Audio page through the real engine in headless Chromium, then read spectrograms, pYIN…

MITAuto-check passedMedia & Creative

Install Listening To Music

skills CLI
$ npx skills add oaustegard/claude-skills --skill listening-to-music -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills listening-to-music --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/listening-to-music .claude/skills/listening-to-music && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
listening-to-music
GitHub stars
150
Token cost
~2.5k tokens
SKILL.md length
1,232 words
Files
17 (incl. scripts)
Skills in repo
67
Repo updated
First seen
Licence
MIT

At a glance

Listen to generated music by measuring it and by reading it as sheet music: render Strudel code or record a Web Audio page through the real engine in headless Chromium, then read spectrograms, pYIN…

  • Works in 4 steps: Render → Look → Measure → …
  • Check whether generated music sounds right
  • SKILL.md covers 1. Render, 2. Look, 3. Measure and Sheet music: notation out,…, plus 2 more sections
  • Runs Python and JavaScript scripts from its folder; calls python3, node and pip

What it does

Listening To Music is an agent skill from oaustegard/claude-skills. Listen to generated music by measuring it and by reading it as sheet music: render Strudel code or record a Web Audio page through the real engine in headless Chromium, then read spectrograms, pYIN pitch lines, chroma, roughness and semitone-rub meters; engrave notes as a score with beat-by-beat harmony; transcribe a melody from audio; read MusicXML/MIDI/ABC or a score image into notes and into Strudel code. Use to check whether generated music sounds right or better (before/after a change), to match a reference…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including scripts (for example `CHANGELOG.md`, `scripts/clashes.py` and `scripts/compare_notes.py`).

It sits in Media & Creative, covering Transcription. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • Check whether generated music sounds right
  • Better (before/after a change)
  • Match a reference track
  • Spectrogram image

Example prompts

  • “does it sound better”
  • “listen to”
  • “render and check”
  • “/listening-to-music”

Requirements

  • Python 3
  • Node.js

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Render
  2. Look
  3. Measure
  4. Change one thing, re-render, re-measure

What it can do on your machine

Read from SKILL.md and the folder at commit 90b0f1b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 15 files in scripts/ (Python and JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • node
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Listening To Music loads about 2.5k tokens when it runs. Until then it costs about 255 tokens; SKILL.md has 1,232 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~255
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 90b0f1b, republished under its MIT licence (© oaustegard). 1,232 words, ~2,500 tokens.

Download SKILL.mdSave it as .claude/skills/listening-to-music/SKILL.md (or your agent's skills folder). This skill also uses 16 other files; get the full folder from GitHub.
name
listening-to-music
description
Listen to generated music by measuring it and by reading it as sheet music: render Strudel code or record a Web Audio page through the real engine in headless Chromium, then read spectrograms, pYIN pitch lines, chroma, roughness and semitone-rub meters; engrave notes as a score with beat-by-beat harmony; transcribe a melody from audio; read MusicXML/MIDI/ABC or a score image into notes and into Strudel code. Use to check whether generated music sounds right or better (before/after a change), to match a reference track, spectrogram image or score, or to find discordant chords, a buried melody, clipping or muddy mixes. Triggers on 'does it sound better', 'listen to', 'render and check', 'spectrogram', 'sfft/stft', 'discordant', 'clashing chords', 'sounds off', 'compare to the original', 'reproduce this track in Strudel', 'sheet music', 'score', 'notation', 'MusicXML', 'MIDI', 'ABC', 'transcribe', 'OMR'. Pairs with strudeling (writing Strudel); for format conversion use processing-video.
metadata.version
0.2.0

Listening to music

Claude hears nothing, but it can render audio in the real engine and measure it. Set up the loop render → look → measure → change one thing → re-render, and keep a table of the numbers per version. A version is better when a number you chose beforehand moves by more than the take-to-take noise, and the spectrogram shows the same thing.

Requirements: Node with playwright and Chromium (preinstalled in Claude Code on the web), Python with librosa soundfile scipy matplotlib pillow, and for sheet music music21 verovio mir_eval cairosvg (pip install --break-system-packages librosa music21 verovio mir_eval cairosvg). The first render installs @strudel/web into ~/.cache/listening-to-music/. All scripts live in scripts/; run them with --help or read the docstring for options.

1. Render

bash
S=/path/to/listening-to-music/scripts
node $S/render_strudel.mjs loop.js --out v1.wav --seconds 40 --warmup 4        # Strudel code
node $S/record_page.mjs page.html --click "#power" --out v1.wav --seconds 40 --warmup 8 \
     --query "station=house" --eval "window.__fm.setSeed(7)"                   # any Web Audio page

Both print peak level and flag clipping. A peak above 1.0 means the listener hears hard clipping; fix levels before judging anything else.

For a before/after comparison, fix the randomness (a seed hook via --eval) so both versions play the same material, and record each version twice.

2. Look

bash
python3 $S/spectrogram.py v1.wav --out v1.png --from-onset --pyin-floor G3 --label v1

Then open v1.png with the image viewer. The top panel is a note-axis CQT spectrogram with the pYIN lead line, and the bottom panel is chroma. Compare it with the reference picture or the previous version side by side. The .npz next to it holds the numbers for step 3.

3. Measure

QuestionToolReads
Do simultaneous notes rub? (discordant chords)render_strudel.mjs --events → clashes.pyminor 2nd/9th overlap per cycle, by layer pair; worst moments
Do they rub in the recording?rubs.py, --pair for replicated takesshare of tonal peak energy in semitone/minor-9th pairs
Is it rougher or smoother overall?roughness.py, --pair for replicated takesSethares roughness median/p90
Does it match a reference image?reference.py ticks/decode/compareper-semitone profile, melody agreement, overtone offsets, chroma agreement
Is the melody audible over the mix?spectrogram.py pYIN voiced %, reference.py compare melody %

The clash scan is exact for the notes as written. rubs.py confirms them in the audio, where release tails and reverb add overlaps the note data does not show. Record harmony checks with percussion and noise textures muted in both versions: drum partials form their own peaks and hold a full mix at a rub share near 0.28 whatever the chords do. Use --warmup, 40 s or more, and two takes per version; --pair prints the take-to-take spread and says when a change is inside it.

Summed roughness (roughness.py) is a coarse overall measure. On Strudel FM (2026-09-24) it could not tell the harmony fix apart from take-to-take noise in sleep and house. On the same takes, the rub meter measured sleep −51%, ambient −33%, lofi −20% and house −13%, with spreads of 0.006–0.015. A loud bass drone dominates the roughness normalisation and hides a pad's rubs. Don't treat an unmoved roughness figure as proof that nothing changed.

Sheet music: notation out, notation in

Notation is the other way to hear. A score shows voicings, clusters (noteheads pushed sideways are seconds), register and rhythm at a glance, and Claude reads clean engraving accurately. All these scripts share one events JSON (layer, t0, t1, midi; times in bars).

bash
node $S/render_strudel.mjs loop.js --events ev.json --cycles 8          # notes as written
python3 $S/score.py ev.json --out score.png --layers LEAD,KEYS,BASS --harmony   # engrave + analyse
python3 $S/transcribe.py take.wav --out mel.json --bpm 115              # notes as heard (melody)
python3 $S/compare_notes.py ev.json mel.json --ref-layer LEAD --cand-layer melody --align 1
python3 $S/read_score.py tune.musicxml --out ref.json                   # .mxl .mid .abc .krn, corpus:
python3 $S/to_strudel.py ref.json --out tune.js                         # score -> Strudel code

score.py --harmony prints each chord change with its pitches, music21's chord name, a Roman numeral in the detected key, and a rub flag. On Strudel FM the same four lofi bars went from 15 of 28 chord changes with a rub to 0 of 28 after the voicing fix, and the engraving showed why: the old Fmaj7 had its E and F side by side.

Measured on 2026-09-25 (compare_notes.py, onset within 0.05 bar and pitch within 50 cents):

PathTestF1
read_score → to_strudel → render eventsBach BWV 66.6, 4 parts, 163 notes100% every part
transcribe.py, one linesame chorale, soprano alone, piano, not used for tuning87%
transcribe.py --no-splitsynth lead with delay echoes82% (65% with splitting)
transcribe.pyfull four-part chorale0%: monophonic only
OMR (oemer 0.1.8)clean engraved soprano line6%; key read as flats, quarters as whole notes
Claude reading the image → ABCunseen chorale, soprano + bass, 38 notes100% (one clean engraving)

So:

  • Score images: read them yourself. Open the page and zoom into dense staves by cropping a 3× render or scan. Write ABC (K:, M:, L:, one V: per staff), convert it with read_score.py, re-engrave with score.py, and compare the two pictures bar by bar before using the notes. Use oemer only as a rough first pass, and never trust it unchecked.
  • Transcription is monophonic. For chords and inner voices, read the generating code's events or the score. Splitting at re-attacks is on by default, because a re-struck piano note has no gap; switch it off with --no-split for leads with delay or echo.
  • Unknown downbeat: a recording that starts mid-loop needs compare_notes.py --align N. It searches shifts up to N bars and prints the one used.
Show full SKILL.md (434 more words)Show less

4. Change one thing, re-render, re-measure

Change one parameter or rule at a time, and after each change render and rerun the same measurements. A metric can reward the wrong thing, so check it against the spectrogram each round.

Known traps (each cost a round on 2026-09-24/25)

  • Mini-notation strings are patterns, not values. .delaytime("3/8") means "3, over 8 cycles", so the delay was 3 s (clamped to 1 s). Pass a number: .delaytime(0.39). .gain("0.9, 0.3") on a stacked pattern is itself a stack, so every hit fired twice and the drums peaked at 2.9× full scale. Give each layer its own scalar gain.
  • AudioWorklet synths (supersaw, …) need a secure context. Pages served from http:// get no audioWorklet, and those synths are silent while everything else plays. The scripts serve local pages from https://local.test/ for this reason. initStrudel also loads the worklets only on the first document click, so the scripts click.
  • pYIN floor. With a C3 floor, a loud bass made pYIN pick the bass's upper partials an octave below the tune. Tracking fell from 60% to 25% even though the melody was unchanged. Set the floor just under the melody's lowest note.
  • Lines above a melody are often its overtones. In a reference, lines at the melody's pitch +12, +19 and +24 semitones are usually its harmonics, not a second part. Test it with reference.py compare (the overtone rows against the control rows) before voicing a pad from them. A misread gave a Gmaj7 that the source never played.
  • Measure in the engine. An offline numpy model of Strudel's FM mispredicted the rendered harmonic spectrum. The low-passed saw it favoured came out darker in the engine than FM.
  • A metric can be confounded. A "lead prominence" score rewarded a pad that doubled the melody's notes. Before trusting a metric, ask what else could raise it.
  • Check reproducibility once. Render the same version twice and confirm the numbers repeat, so differences between versions can be trusted.
  • music21 respelling. chordify() discards pitch spelling, and a Key's inferred tonic transposes to the wrong letter (a leading tone of F instead of E♯ in F♯ minor). notes.spell re-spells after every chordify from the key's own scale plus the raised 6th and 7th. Without that, the score reads as wrong notes.
  • Quantise to a beat subdivision, not a bar fraction. A 16-step bar in 3/4 produced 64th notes and triple dots; the grid is now 16ths in any metre.
  • Decoding reference images. Bin centres sit on the tick rows. Get exact ticks with reference.py ticks rather than by eye; a half-bin error splits every note across two semitones.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 16 other files (scripts) in listening-to-music of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • scripts/clashes.py
  • scripts/common.mjs
  • scripts/compare_notes.py
  • scripts/notes.py
  • scripts/read_score.py
  • scripts/record_page.mjs
  • scripts/reference.py
  • scripts/render_strudel.mjs
  • scripts/roughness.py
  • scripts/rubs.py
  • scripts/score.py
  • scripts/spectrogram.py
  • scripts/tap.js
  • scripts/to_strudel.py
  • scripts/transcribe.py

Open the folder on GitHubat commit 90b0f1b

Compare with similar skills

Listening To Music next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Listening To Music compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Listening To Music this skilloaustegard/claude-skills150—~2.5kAutomated safety check: PassMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Native Subtitle Quote Imagechengyi-ai/native-subtitle-quote-image2.6k—~2.4kAutomated safety check: PassMIT
Edu Chem Videowy51ai/edulab1.4k—~2.1kAutomated safety check: NotesApache-2.0
Transcription Memory ReconstructionNxcoreAI/EverRoom3k—~714Automated safety check: PassCustom licence
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0

Similar skills

  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Native Subtitle Quote Image

    chengyi-ai/native-subtitle-quote-image

    将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…

    2.6k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Edu Chem Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a chemistry problem (化学题: 氧化还原配平 双线桥 电子守恒, 物质的量计算, 化学平衡 三段式 平衡常数 转化率 反应速率, 离子反应, 电化学, 溶液 滴定…

    1.4k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Reconstruct a complete, searchable memory from an untrusted meeting or conversation transcript.

    3k GitHub stars~714 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Transcribe

    JetBrains/skills

    Official

    Transcribe audio files to text with optional diarization and known-speaker hints.

    366 GitHub starsUsed in 4 repos~776 tokens
    Media & CreativeAuto-check passed

More from oaustegard/claude-skills

All 67 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Deciding With Confidence

    oaustegard/claude-skills

    Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds…

    150 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Listening To Music

What does Listening To Music do?

Listen to generated music by measuring it and by reading it as sheet music: render Strudel code or record a Web Audio page through the real engine in headless Chromium, then read spectrograms, pYIN…. Listening To Music is an agent skill from oaustegard/claude-skills. Listen to generated music by measuring it and by reading it as sheet music: render Strudel code or record a Web Audio page through the real engine in headless Chromium, then read spectrograms, pYIN pitch lines, chroma, roughness and semitone-rub meters; engrave notes as a score with beat-by-beat harmony; transcribe a melody from audio; read MusicXML/MIDI/ABC or a score image into notes and into Strudel code.

When should I use Listening To Music?

Listening To Music fits situations like: check whether generated music sounds right; better (before/after a change); match a reference track; spectrogram image.

How do I install Listening To Music in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill listening-to-music -a claude-code`. Or copy the skill folder (listening-to-music in oaustegard/claude-skills) into .claude/skills/listening-to-music in your project. Claude Code loads it when a task matches its description.

How do I install Listening To Music in Codex?

Run `npx skills add oaustegard/claude-skills --skill listening-to-music -a codex`. Or copy the skill folder (listening-to-music in oaustegard/claude-skills) into .agents/skills/listening-to-music in your project. Codex loads it when a task matches its description.

Can I use Listening To Music in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill listening-to-music -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/listening-to-music, .gemini/skills/listening-to-music, .github/skills/listening-to-music and .opencode/skills/listening-to-music in your project.

What does Listening To Music need to run?

Going by SKILL.md and its folder, Listening To Music needs Python and JavaScript for the scripts in its folder and the command-line tools its instructions call (python3, node and pip). Our summary lists: Python 3; Node.js.

Does Listening To Music access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Listening To Music safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Listening To Music use?

Listening To Music is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Listening To Music use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Listening To Music?

Skills that share tags, products or a category with Listening To Music: HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Native Subtitle Quote Image (chengyi-ai/native-subtitle-quote-image, 2.6k stars), Edu Chem Video (wy51ai/edulab, 1.4k stars) and Transcription Memory Reconstruction (NxcoreAI/EverRoom, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Listening To Music?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 67 skills in this directory. The repository was last updated on October 9, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.