Agent skill

Lyria

by calesthio in calesthio/OpenMontage

Generate and validate music with Google Lyria 3 through the Gemini Interactions API.

AGPL-3.0Auto-check passedMedia & Creative

Install Lyria

skills CLI
$ npx skills add calesthio/OpenMontage --skill lyria -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/OpenMontage lyria --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/lyria .claude/skills/lyria && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
lyria
GitHub stars
66k
Token cost
~2k tokens
SKILL.md length
966 words
Files
3 (incl. references)
Skills in repo
41
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Generate and validate music with Google Lyria 3 through the Gemini Interactions API.

  • Works in 8 steps: Confirm the intended role: underscore,… → Confirm vocals, language, target… → Announce the provider, exact model,… → …
  • Tasks that involve Music and audio generation
  • SKILL.md covers Required Workflow, Choose The Model Deliberately, Build The Prompt and Direct Vocals Deliberately, plus 4 more sections
  • Calls ffprobe; needs GEMINI_API_KEY and GOOGLE_API_KEY

What it does

Lyria is an agent skill from calesthio/OpenMontage. Generate and validate music with Google Lyria 3 through the Gemini Interactions API. Use before calling OpenMontage googlemusic, designing Lyria 3 Clip or Pro prompts, using image-to-music or custom lyrics, choosing between Lyria 3 and Lyria RealTime, diagnosing Google music-generation failures, or preparing exact-duration music for a video.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/api-and-prompting.md`).

It sits in Media & Creative, covering Music and audio generation. The repository describes itself as: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant… The licence is AGPL-3.0.

When your agent uses it

  • Tasks that involve Music and audio generation

Example prompts

  • “/lyria”

Requirements

  • A credential in GEMINI_API_KEY
  • A credential in GOOGLE_API_KEY

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the intended role: underscore, loop, instrumental cue, or full song.
  2. Confirm vocals, language, target duration, musical structure, and delivery format.
  3. Announce the provider, exact model, estimated per-request cost, and whether the call is exploratory or final.
  4. Write one structured prompt using the contract below.
  5. Make one approved generation call. Treat a retry as another potentially billable stochastic generation.
  6. Preserve the returned provider source unchanged.
  7. Probe the actual file and listen before recording duration, format, or approval metadata.
  8. Derive a separate production master when the edit requires an exact duration.

What it can do on your machine

Read from SKILL.md and the folder at commit 9327439. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ffprobe

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY
    • GOOGLE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Lyria loads about 2k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 88 tokens; SKILL.md has 966 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/OpenMontage at commit 9327439, republished under its AGPL-3.0 licence (© calesthio). 966 words, ~2,002 tokens.

Download SKILL.mdSave it as .claude/skills/lyria/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
lyria
description
Generate and validate music with Google Lyria 3 through the Gemini Interactions API. Use before calling OpenMontage `google_music`, designing Lyria 3 Clip or Pro prompts, using image-to-music or custom lyrics, choosing between Lyria 3 and Lyria RealTime, diagnosing Google music-generation failures, or preparing exact-duration music for a video.

Google Lyria 3

Use Lyria 3 as a single-turn music generator. Keep it distinct from Lyria RealTime, which is an experimental WebSocket model for continuously steered instrumental performance.

Read references/api-and-prompting.md when choosing a model, designing vocals or custom lyrics, using image inputs, debugging the response, or checking current limits and pricing.

Required Workflow

  1. Confirm the intended role: underscore, loop, instrumental cue, or full song.
  2. Confirm vocals, language, target duration, musical structure, and delivery format.
  3. Announce the provider, exact model, estimated per-request cost, and whether the call is exploratory or final.
  4. Write one structured prompt using the contract below.
  5. Make one approved generation call. Treat a retry as another potentially billable stochastic generation.
  6. Preserve the returned provider source unchanged.
  7. Probe the actual file and listen before recording duration, format, or approval metadata.
  8. Derive a separate production master when the edit requires an exact duration.

Choose The Model Deliberately

NeedModelContract
Prompt iteration, preview, loop, exact 30-second sourcelyria-3-clip-previewAlways generates a 30-second MP3; currently $0.04/request
Full song, vocals, longer structure, image-conditioned scorelyria-3-pro-previewPrompt-influenced duration up to roughly three minutes; currently $0.08/request
Live, continuously steered instrumental performancelyria-realtime-expSeparate WebSocket workflow; do not route through google_music

The current OpenMontage google_music adapter is locked to lyria-3-pro-preview. It does not expose Clip, WAV response selection, multiple images, or RealTime controls. Surface that limitation rather than implying those options are available through the adapter.

Do not change models silently. For a 30-second video, either obtain approval for Pro plus exact-duration mastering or use Clip through an explicitly supported path.

Build The Prompt

Specify, in this order:

  1. Purpose and duration — what the music supports and the requested length.
  2. Genre and era — use musical vocabulary, not a living artist imitation.
  3. Tempo and harmony — BPM or tempo range, meter, key or tonal center.
  4. Instrumentation and texture — name lead, rhythm, bass, and ambient layers.
  5. Structure — timestamped sections or [Intro], [Verse], [Chorus], [Bridge], [Outro].
  6. Dynamics and synchronization — entrances, rests, builds, hits, and holds tied to edit times.
  7. Vocal policy — instrumental-only constraints, a vocal profile, or clearly separated custom lyrics.
  8. Mix and ending — density, foreground/background role, headroom character, and final decay.
  9. Exclusions — unwanted vocals, instruments, gestures, clichés, abrupt endings, or copyrighted material.

For video underscore, use timestamp windows that cover the full requested duration. Ask for one primary change per window and identify the exact synchronization moment.

For instrumental-only output, say all of the following when they matter:

text
Instrumental only. No lead or backing vocals, speech, choir, humming,
vocal chops, spoken samples, lyrical fragments, or recognizable quotations.

For custom lyrics, put performance direction before a separate Lyrics: block and use section labels. Prompt in the language the singer should use.

Direct Vocals Deliberately

When vocals are requested, define these before writing the prompt:

  1. Vocal role — solo lead, duet, call-and-response, backing ensemble, or vocal texture.
  2. Language and script — name the sung language and keep the custom lyrics in one intentional script; do not silently transliterate or code-switch.
  3. Singer profile — voice type or range, timbre, intensity, diction, ornamentation, and emotional distance. Do not imitate a named artist.
  4. Section behavior — state where the lead enters, where harmonies or echoes appear, and which sections remain instrumental.
  5. Lyric contract — separate directions from a Lyrics: block, use [Verse], [Chorus], [Bridge], and [Outro], and reserve parentheses for intentional backing-vocal echoes.

Treat the returned vocal as untrusted until auditioned. Check lyric adherence, language drift, pronunciation, intelligibility, unwanted backing vocals, vocal/instrument balance, and whether the performance follows the requested emotional arc. A technically valid file with poor diction or altered lyrics is not an approved vocal result.

Show full SKILL.md (374 more words)Show less

Treat Duration As Untrusted Until Probed

Lyria 3 Pro duration is controlled through prompt instructions and timestamps, not an exact API parameter. The OpenMontage adapter appends a target-duration instruction, but its returned duration_seconds field is the request, not a media probe.

Always inspect the generated file:

bash
ffprobe -v error -show_entries \
  format=duration,format_name,bit_rate:stream=codec_name,sample_rate,channels \
  -of json output.mp3

If exact duration is required:

  • keep the provider source untouched;
  • record requested and measured durations separately;
  • derive a new master by trimming at a musically sensible boundary and applying a short fade;
  • do not stretch, loop, or regenerate without the approved production plan;
  • record the derivation and probe the master again.

Authenticate And Diagnose Safely

The Gemini API commonly uses GEMINI_API_KEY. OpenMontage also supports GOOGLE_API_KEY and Vertex service-account credentials.

  • Use one known credential path per run.
  • Never print keys or edit credential files while debugging.
  • Do not assume a rejected first key will fall through to a second configured key.
  • In the current OpenMontage resolver, GOOGLE_API_KEY takes precedence over GEMINI_API_KEY when both are non-empty.
  • Treat 403, API_KEY_SERVICE_BLOCKED, and project/service restrictions as authentication or Google-project configuration failures, not prompt-quality failures.
  • Do not spend retries on permission failures. Resolve the credential/project path first.
  • Retry only transient rate-limit or timeout failures within the approved retry and budget policy.

Parse And Record The Result

Prefer interaction.output_audio. For interleaved responses, traverse model_output steps and select the audio block; capture output text separately if lyrics or a structure description are relevant.

Record:

  • provider and exact model;
  • original prompt and any image provenance;
  • requested duration and probed duration;
  • actual codec, sample rate, channels, and file path;
  • cost per call and total attempts;
  • whether vocals were requested and whether any were detected by listening;
  • the provider source and any separately derived production master.

Quality Checklist

  • The prompt states purpose, tempo, instruments, structure, dynamics, vocal policy, and ending.
  • Timestamp windows cover the intended runtime without contradictions.
  • No artist impersonation or copyrighted lyrics are requested.
  • The output file exists, is non-empty, decodes, and contains an audio stream.
  • Actual duration and technical properties come from a probe, not request metadata.
  • Instrumental output is checked for accidental vocal material.
  • The opening, synchronization moments, transitions, and ending are auditioned.
  • The untouched source is preserved and any production master has explicit provenance.
  • SynthID watermarking and preview-model instability are acknowledged where provenance matters.

© calesthio, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/lyria of calesthio/OpenMontage.

  • SKILL.md
  • agents/openai.yaml
  • references/api-and-prompting.md

Open the folder on GitHubat commit 9327439

Compare with similar skills

Lyria next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Lyria compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Lyria this skillcalesthio/OpenMontage66k—~2kAutomated safety check: PassAGPL-3.0
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT
Sound Effectstadaspetra/loop2962 repos~1.1kAutomated safety check: PassMIT
Characteristic VoiceNoizAI/skills526—~1.8kAutomated safety check: PassNone
Text To Sfxsonilo-ai/skills1151 repos~1.6kAutomated safety check: NotesMIT

Similar skills

  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Media & CreativeAuto-check passed
  • A skill your agent uses whenever the user wants speech to sound more human, companion-like, or emotionally expressive.

    526 GitHub stars~1.8k tokensUpdated 11 days ago
    Media & CreativeAuto-check passed
  • Text To Sfx

    sonilo-ai/skills

    Generate a sound effect from a text description using Sonilo — a UI chime, a whoosh, an impact, ambience, a stylized cue — when there is no video to match.

    115 GitHub starsUsed in 1 repo~1.6k tokens
    Media & CreativeAuto-check: notes
  • Music Caption Rewriter

    T8mars/T8-penguin-canvas

    Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware Arrangement.

    615 GitHub stars~2.2k tokensUpdated today
    Media & CreativeAuto-check passed

More from calesthio/OpenMontage

All 41 skills in this repo
  • Video Understand

    calesthio/OpenMontage

    Understand video content locally using ffmpeg frame extraction and Whisper transcription.

    66k GitHub stars~841 tokensUpdated 6 days ago
    Auto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 6 days ago
    Auto-check passed
  • D3 Viz

    calesthio/OpenMontage

    Creating interactive data visualisations using d3.js. An agent skill from calesthio/OpenMontage.

    66k GitHub starsUsed in 3 repos~5.4k tokens
    Auto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    66k GitHub stars~1.3k tokensUpdated 6 days ago
    Auto-check passed
  • Threejs World Generation

    calesthio/OpenMontage

    Build deterministic, editable, free-viewpoint Three.js worlds from text or structured briefs.

    66k GitHub stars~2k tokensUpdated 6 days ago
    Auto-check passed
  • Video Edit

    calesthio/OpenMontage

    Edit videos locally using ffmpeg. An agent skill from calesthio/OpenMontage.

    66k GitHub stars~855 tokensUpdated 6 days ago
    Auto-check: notes

Questions about Lyria

What does Lyria do?

Generate and validate music with Google Lyria 3 through the Gemini Interactions API. Lyria is an agent skill from calesthio/OpenMontage. Generate and validate music with Google Lyria 3 through the Gemini Interactions API.

When should I use Lyria?

Lyria fits situations like: tasks that involve Music and audio generation.

How do I install Lyria in Claude Code?

Run `npx skills add calesthio/OpenMontage --skill lyria -a claude-code`. Or copy the skill folder (.agents/skills/lyria in calesthio/OpenMontage) into .claude/skills/lyria in your project. Claude Code loads it when a task matches its description.

How do I install Lyria in Codex?

Run `npx skills add calesthio/OpenMontage --skill lyria -a codex`. Or copy the skill folder (.agents/skills/lyria in calesthio/OpenMontage) into .agents/skills/lyria in your project. Codex loads it when a task matches its description.

Can I use Lyria in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/OpenMontage --skill lyria -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lyria, .gemini/skills/lyria, .github/skills/lyria and .opencode/skills/lyria in your project.

What does Lyria need to run?

Going by SKILL.md and its folder, Lyria needs the command-line tools its instructions call (ffprobe) and credentials named GEMINI_API_KEY and GOOGLE_API_KEY. Our summary lists: A credential in GEMINI_API_KEY; A credential in GOOGLE_API_KEY.

Does Lyria access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Lyria safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Lyria use?

Lyria is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Lyria use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Lyria?

Skills that share tags, products or a category with Lyria: HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Music (tadaspetra/loop, 296 stars), Sound Effects (tadaspetra/loop, 296 stars) and Characteristic Voice (NoizAI/skills, 526 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Lyria?

calesthio (a GitHub user) maintains it in calesthio/OpenMontage, which has 65,930 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 3, 2026.

Source: calesthio/OpenMontage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.