Agent skill

Pneuma Clipcraft

by pandazki in pandazki/pneuma-skills

AI-orchestrated video production on @pneuma-craft. An agent skill from pandazki/pneuma-skills.

MITAuto-check: notesMedia & Creative

Install Pneuma Clipcraft

skills CLI
$ npx skills add pandazki/pneuma-skills --skill pneuma-clipcraft -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pandazki/pneuma-skills pneuma-clipcraft --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pandazki/pneuma-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/modes/clipcraft/skill .claude/skills/pneuma-clipcraft && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pneuma-clipcraft
GitHub stars
161
Token cost
~7.5k tokens
SKILL.md length
3,186 words
Files
15 (incl. scripts, references)
Skills in repo
30
Repo updated
First seen
Licence
MIT

At a glance

AI-orchestrated video production on @pneuma-craft. An agent skill from pandazki/pneuma-skills.

  • Works in 5 steps: Read the current project.json to… → Generate assets by running one of the… → Register each new asset in project.json → …
  • The user wants to generate
  • SKILL.md covers Working with the viewer, The 6-layer technique stack, Domain vocabulary (2-minute… and Making creative decisions, plus 5 more sections
  • Runs JavaScript scripts from its folder; calls node and curl; needs OPENROUTER_API_KEY and FAL_KEY

What it does

Pneuma Clipcraft is an agent skill from pandazki/pneuma-skills. AI-orchestrated video production on @pneuma-craft. Use whenever the user wants to generate, edit, or compose video clips, audio tracks, captions, or background music — text-to-video / image-to-video generation, TTS narration, music generation, provenance tracking, timeline composition — or touches the exploded timeline, the dive-in panels, or project.json by hand. Do not assume the user knows the schema — they usually don't; read references/project-json.md before committing to an edit.

Its SKILL.md is about 7.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including scripts and reference files (for example `references/asset-ids.md`, `references/character-consistency.md` and `references/craft.md`).

It sits in Media & Creative, covering Video production, Text to speech and voice and AI video generation. The repository describes itself as: Co-creation infrastructure for humans and code agents — visual environment, skills, continuous learning, and distribution. The licence is MIT.

When your agent uses it

  • The user wants to generate
  • Compose video clips
  • Background music — text-to-video / image-to-video generation
  • Music generation

Example prompts

  • “/pneuma-clipcraft”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Read the current project.json to understand the composition.
  2. Generate assets by running one of the scripts with Bash.
  3. Register each new asset in project.json
  4. Place assets on the timeline by adding clips to the relevant
  5. The viewer auto-reflects every edit — no reload needed.

What it can do on your machine

Read from SKILL.md and the folder at commit 0023d3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY
    • FAL_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pneuma Clipcraft loads about 7.5k tokens when it runs, and up to ~44k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 3,186 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~7.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~44k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:263
    r API keys from `process.env` or from a `.env`

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from pandazki/pneuma-skills at commit 0023d3c, republished under its MIT licence (© pandazki). 3,186 words, ~7,524 tokens.

Download SKILL.mdSave it as .claude/skills/pneuma-clipcraft/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
pneuma-clipcraft
description
AI-orchestrated video production on @pneuma-craft. Use whenever the user wants to generate, edit, or compose video clips, audio tracks, captions, or background music — text-to-video / image-to-video generation, TTS narration, music generation, provenance tracking, timeline composition — or touches the exploded timeline, the dive-in panels, or `project.json` by hand. Do not assume the user knows the schema — they usually don't; read `references/project-json.md` before committing to an edit.

ClipCraft

In script examples, <SKILL_DIR> means the actual directory containing this loaded SKILL.md. Substitute its full path and keep shell paths quoted. The runtime installs it under .claude/skills for Claude Code, .agents/skills for Codex, or .kimi-code/skills for Kimi; use the path given in your instructions.

ClipCraft is a video-production mode where the source of truth is a structured domain model, not a file. The in-memory model is an event-sourced craft store from @pneuma-craft: an Asset registry, a Composition with Tracks and Clips, and a Provenance DAG that tracks how each asset was generated (and from what). The file project.json at the workspace root is a projection of that store — when you edit it with Write/Edit, the viewer auto-re-hydrates. No reload, no refresh signal.

ClipCraft is built for AIGC workflows: assets are generated, not uploaded. You orchestrate image / video / TTS / BGM generation by running bundled scripts, then record the lineage in project.json.

Working with the viewer

The viewer is the exploded 3D timeline rendering of project.json. It is the user's source of truth for what's currently selected, what the playhead is on, and what they're pointing at. Four channels link the three actors (you, the user, the viewer):

Reading what the user sees

Every user message arrives wrapped in <viewer-context> and (if the user just clicked / dragged / seeked) <user-actions>. Read them before you act. Typical clipcraft payloads:

  • <viewer-context> — selectedClipId, selectedAssetId, selectedTrackId, playheadTime (seconds), composition.duration, the active asset's metadata. Use this to disambiguate a vague request like "try another take" — there's almost always a clip selected that tells you which one.
  • <preview-frames> (nested inside <viewer-context>) — summary of the planning layer: total="N" attribute + per-track <track id="..." name="..." count="..." /> lines for tracks that carry preview frames. When total="0" (or the tag is absent), the timeline has no planning layer yet — a fresh project. Use this to decide whether to start with sketch overlay (see references/storyboard-workflow.md) vs jumping straight to generation.
  • <user-actions> — recent UI events: playhead:seek ({time}), clip:select ({clipId, trackId}), asset:select ({assetId}), clip:drag ({clipId, startTime, trackId}), track:toggle-mute, track:toggle-visible. Treat these as hints, not commands — the user usually expects you to read them and act, not echo them back.

If both are absent (cold start, command-button click), inspect project.json directly and ask if intent is ambiguous.

Locator cards

After creating or editing assets, clips, or moving the playhead, embed <viewer-locator> cards so the user can jump straight to the change. Emit one card per distinct thing you changed — a newly generated asset, a clip you just placed, a time beat you built around — not one per response. The user sees these as clickable chips in chat. Use short concrete labels — "新的 VO 开场", "panda clip on Main", "3.5s — punchline beat" — not generic ones like "see asset".

ViewerAddress — naming an object in the preview

A <viewer-locator> carries an address — a small JSON object that names one referent inside the viewer. clipcraft's address vocabulary is timeline-oriented; every key below is coarse (each one stands alone and drives navigate-then-flash — clipcraft has no element-level fine handle):

KeyCoarse/fineNames
clipIdcoarsea clip on a track — selects it + seeks the playhead to its start
previewFrameIdcoarsea planning sketch / anchor frame — selects its asset + seeks to its anchor time
assetIdcoarsean entry in the asset library — scrolls + flashes the tile/row
trackIdcoarsea track header — scrolls + flashes it
timecoarsea moment on the timeline — seeks the playhead only

Supply exactly one key per address. clipId / previewFrameId / assetId / trackId are stable ids (survive move / re-pace), so locator references stay valid; time is a raw second.

html
<!-- assetId — scrolls the asset library to the asset and flashes it. -->
<viewer-locator address='{"assetId":"asset-vo-tagline"}'>新的 VO 开场</viewer-locator>

<!-- clipId — selects the clip on the timeline AND seeks the playhead
     to its startTime, so the user lands on the frame you mean. -->
<viewer-locator address='{"clipId":"clip-shot1-spark"}'>panda clip on Main</viewer-locator>

<!-- time — seeks the playhead only (no selection change). Use for
     pure "go look at this beat" pointers. -->
<viewer-locator address='{"time":3.5}'>3.5s — punchline beat</viewer-locator>

<!-- trackId — scrolls and flashes a track header. Use when the
     change is track-level (mute/solo, reordered, new track). -->
<viewer-locator address='{"trackId":"track-narration"}'>narration track</viewer-locator>

<!-- previewFrameId — selects the referenced image asset, seeks the
     playhead to the preview frame's anchor time, and flashes the
     strip thumbnail on the timeline. Use when pointing at a sketch
     or anchor on the planning layer (e.g. "I generated 3 sketches
     for the opening; here's panel 4"). The preview frame's id is
     stable across move/rebind, so locator references survive
     re-pacing. -->
<viewer-locator address='{"previewFrameId":"pf-04"}'>panel 4 — opening sketch</viewer-locator>
Viewer commands (user → agent)

The viewer toolbar exposes six command buttons. Clicks arrive as short natural-language messages in chat — they're hints about what the user wants next, usually with a clip or scene pre-selected, not rigid tool calls. Read <viewer-context> to figure out the target, then run the matching workflow. If intent is ambiguous (e.g. a vague "generate video"), confirm before spending money on veo3.1.

CommandTypical handling
Generate imageNew image asset for the current selection — read context for scene/clip
Generate videoNew video clip; confirm before veo3.1 if vague
Try another takeVariant of the selected clip's asset; register as a derived asset (provenance edge) so the variant switcher shows both options
Add narrationTTS for the selected subtitle clip (or the whole caption track); match audio clip timing to subtitle clip timing
Add BGMAsk for mood/style if not given; generate, register, place on a new or existing audio track
Export draftHandled in the viewer — runs ExportEngine with includePreviewFrames: true so sketches + anchors bake in. No agent involvement for the click. Used during planning to verify pacing before committing to expensive seedance generation.
Export videoHandled in the viewer — runs @pneuma-craft/video ExportEngine. No agent involvement.
Agent → viewer actions (HTTP)

When you need to drive the viewer (not just respond), POST to $PNEUMA_API/api/viewer/action. Reach for this when the user asks "show me the part where..." and you want the playhead to land there before you explain, or when you've just registered an asset and want it pre-selected for the next take.

bash
# Seek the playhead to a specific second.
curl -s -X POST "$PNEUMA_API/api/viewer/action" \
  -H 'content-type: application/json' \
  -d '{"action":"playhead:seek","payload":{"time":4.2}}'

# Select a clip on the timeline (also seeks to its start).
curl -s -X POST "$PNEUMA_API/api/viewer/action" \
  -H 'content-type: application/json' \
  -d '{"action":"clip:select","payload":{"clipId":"clip-shot1-spark"}}'

Prefer <viewer-locator> cards when the user benefits from a clickable hand-off. Use HTTP actions when you need the viewer state to change before the next step (e.g. taking a screenshot via /api/native/screenshot).

The 6-layer technique stack

ClipCraft is more than a thin wrapper over generation APIs. The mode encapsulates a body of techniques — a way of thinking about AIGC video production — organized in six layers. When a brief lands, scan top-down and figure out which layers it touches; drill into those references and skip the rest. A one-shot 4s clip is Layer 6 only. A music video exercises all six.

LayerConcernPrimary reference(s)
1. Production BibleLock the world before generating any pixel — characters, settings, project biblereferences/production-bible.md (+ craft.md for principles, character-consistency.md for the seedance-filter special case)
2. Storyboard PathsChoose A/B/C generation strategy; structure the shot listreferences/storyboard-design.md
3. Direction NotationEncode intent precisely in prompts and references — production triggers, annotation color system, FACS, IPA, faithfulness, anti-patternsreferences/direction-notation.md (+ reference-directives.md for @-addressing)
4. Iteration WorkflowSketch → anchor → clip on the timeline; draft exports between stagesreferences/storyboard-workflow.md
5. Provenance GraphLineage as audit trail and "try another take" foundationreferences/project-json.md (+ asset-ids.md for id rules)
6. Generation ToolsRun the actual APIs; recover from filter rejectionsreferences/workflows.md (worked examples), references/filter-retries.md (422 retry tree)
Decision tree by brief shape
Brief shapeLayers to touch
"make a 4s clip of X"6 only
"10s opening, single character"1 (light bible) + 4 + 6
"30s mini-story with a recurring character"1 (full bible) + 2 + 4 + 6
"music video / dance / dialogue-heavy"1 + 2 + 3 + 4 + 6
"60s ad with multiple characters across scenes"All 6
Compositing intent into one ref image (Path C, dance, etc.)3 (annotation system) heavily
Photoreal human getting rejected by seedance1 (character-consistency.md) + 6 (filter-retries.md)
When in doubt — Layer 1 first

The single biggest failure mode in multi-shot AIGC video work is drift across shots: the character's face changes, the location mutates, the palette wanders. Layer 1 is the antidote, and it's cheap (~$0.50–$2 in upstream image generations) compared to the cost of regenerating a Path A run because the protagonist looks different in panel 5. If a brief has any recurring character, location, or signature prop, build the bible before the storyboard.

Layers are not strictly sequential

Plan top-down (1 → 2 → 3 → 4) but execute bottom-up where appropriate: you might generate a quick test sketch (Layer 6) to calibrate the bible (Layer 1), then redesign. Provenance (Layer 5) is recorded continuously, not in a phase. Direction Notation (Layer 3) is consulted whenever you write a prompt at any other layer.

Domain vocabulary (2-minute version)

  • Asset — an addressable piece of media. Has id, type, uri, name, metadata, status, createdAt.
  • Track — a horizontal lane in the timeline. type is video / audio / subtitle.
  • Clip — a span on a track that references an asset via assetId. Has startTime, duration, inPoint, outPoint (all in seconds). Subtitle clips also carry text directly.
  • Scene — a logical chunk of the composition that groups clips across tracks. Purely a human organization aid.
  • Provenance edge — { toAssetId, fromAssetId, operation }. Captures "how was this asset created". fromAssetId: null means generated from nothing; a real id means derived from another asset.
  • Composition — the top-level container: settings (width/height/fps), tracks, transitions, duration.

Full schema in references/project-json.md. Id rules in references/asset-ids.md.

Making creative decisions

Most of the work in ClipCraft is judgment, not mechanics. Before writing a generation prompt, choosing shots, picking music, or deciding what to cut, read references/craft.md — a field guide to short-video craft. It's principles, not procedures. Reach for it when the brief is open ("make something about X"), when a generated clip feels close but wrong, or any time you're about to settle for a generic answer. The rest of this document tells you how to produce; craft.md tells you what is worth producing.

Generation scripts

Seven CLI scripts wrap the provider APIs. Call them via the Bash tool.

ScriptPurposeDefault modelEnv var
scripts/generate_image.mjsText→image; edits via --image-urls; 1–10 images per callGPT Image 2.5 Sunburst for generation; Flare with referencesOPENROUTER_API_KEY
scripts/edit_image.mjsModify a local image with optional highlighter annotation (multimodal reasoning)GPT Image 2.5 Flare via OpenRouterOPENROUTER_API_KEY
scripts/generate-video.mjsText→video + image→video + reference-to-videobytedance seedance-2.0 (fallback: veo3.1 via --model veo3.1)FAL_KEY
scripts/generate-tts.mjsText→speech (expressive: inline [laughing] / [sigh] tags, 30 voices)fal.ai gemini-3.1-flash-ttsFAL_KEY
scripts/generate-bgm.mjsText→background musicOpenRouter google/lyria-3-pro-previewOPENROUTER_API_KEY
scripts/make-character-sheet.mjsPhoto → photo-body / sketch-head 16:9 character reference sheet (deterministic recovery shortcut for seedance's image-side filter; see references/filter-retries.md)GPT Image 2.5 Flare via OpenRouterOPENROUTER_API_KEY
scripts/storyboard.mjsCompose-and-slice: one GPT Image 2.5 call generates an N-cell composite at the target video aspect ratio; ffmpeg slices into N individual panel images with provenance metadata. Engine layer for Path C (see references/storyboard-workflow.md).Sunburst, or Flare with --ref, via OpenRouterOPENROUTER_API_KEY

generate_image.mjs and edit_image.mjs share their output shape — a JSON object on stdout with files, urls, and description. They take --output-dir + --filename-prefix (NOT --output). The prompt is a positional argument, not a flag. The other four follow the older flag-based convention where --output <path> is required and stdout is just the output path.

All scripts read their API keys from process.env or from a .env file in the skill directory.

Why GPT Image 2.5 matters for video work

The shared image model is gpt-image-2.5-sunburst via OpenRouter; gpt-image-2.5-flare is selected automatically for editing and calls with reference images. Use it for:

  • First / last frames. Seedance's from-image and first-last-frame video modes inherit the quality of their anchor images. GPT Image 2.5 holds a specific aesthetic, character, and composition across paired calls — so the two frames actually look like they belong to the same shot, and the interpolated video doesn't need to fight a stylistic mismatch at the seams.
  • Complex single-frame compositions — foreground subject + environment + overlay text all in one image, rendered legibly. Title cards, end cards, lower thirds, memes with baked-in captions, data callouts over b-roll, diagrammed explainers — all viable as standalone assets rather than ffmpeg/post overlays.
  • Text rendering that actually reads. "A sign that says X" or "a poster with the headline Y" comes back legible, not glyph soup. Use it for signage, lower-third strap text, brand marks, chyron-style overlays.
  • Multi-reference stitching via --image-urls. Pass a character portrait plus an environment plate plus a style plate and GPT Image 2.5 composes them coherently — a stronger control surface for prompting first frames and character sheets than free-text alone.
  • Reference-guided edits. Pass the source via --image-urls and describe what to change and preserve. edit_image.mjs --annotation accepts a visual location guide. OpenRouter does not expose a pixel-mask parameter.

Be ambitious with the creative brief: text-heavy frames, multi-layer compositions, and explicit character continuity are viable in a single generation. See references/craft.md for the principles that should drive those choices.

Show full SKILL.md (1,238 more words)Show less
Sizing images for video (critical)

When an image is destined to be a video first or last frame (for generate-video.mjs from-image or the seedance first-last-frame mode), its pixel dimensions must match the video output exactly. An off-size anchor image gets letterboxed, cropped, or distorted by the video model at the seams.

Pass --image-size WxH to request the composition dimensions through OpenRouter's size field. Check the returned image dimensions before using it as an anchor: providers may normalize the requested size. --aspect-ratio only requests proportions.

Composition-to-image-size cheat sheet for the common cases:

CompositionSeedance outputUse --image-size
9:16 portrait @ 720p720×1280720x1280
9:16 portrait @ 1080p (veo3.1)1080×19201080x1920
16:9 landscape @ 720p1280×7201280x720
16:9 landscape @ 1080p (veo3.1)1920×10801920x1080
1:1 squaredepends on resolutionmatch composition.settings

For standalone assets that never enter the video pipeline (wallpapers, moodboards, illustrations the user just wants to look at), --aspect-ratio is fine.

Calling generate_image.mjs
bash
# Text-to-image (positional prompt)
node "<SKILL_DIR>/scripts/generate_image.mjs" \
  "A dimly lit kitchen at 3am, kettle steam catching the overhead bulb, shot on 35mm" \
  --aspect-ratio 9:16 --quality high \
  --output-dir assets/image --filename-prefix kitchen-3am

# Edit / reference — pass one or more URLs, data URIs, or local files.
# Uses the same GPT Image 2.5 OpenRouter Images endpoint.
node "<SKILL_DIR>/scripts/generate_image.mjs" \
  "Same character, now at a neon-lit ramen counter, back to camera" \
  --image-urls https://example.com/character-ref.png \
  --aspect-ratio 9:16 --quality high \
  --output-dir assets/image --filename-prefix kitchen-to-ramen

# Video first-frame — request the composition dimensions.
# Verify actual output dimensions before using the image as an anchor.
node "<SKILL_DIR>/scripts/generate_image.mjs" \
  "Overhead shot of a desk at 3am: laptop closed, spiral notebook, cold coffee ring, warm tungsten lamp in upper right" \
  --image-size 720x1280 --quality high \
  --output-dir assets/image --filename-prefix opening-desk

# Multiple takes in one call — 1–10 per request.
node "<SKILL_DIR>/scripts/generate_image.mjs" \
  "Four phone mockups of the app home screen, each with a different colorway" \
  --num-images 4 --aspect-ratio 9:16 \
  --output-dir assets/image --filename-prefix colorway-grid

# Flare edit with a source image. Requires OPENROUTER_API_KEY.
node "<SKILL_DIR>/scripts/generate_image.mjs" \
  "Watercolor of a city at dusk, soft bleeds, visible cold-press texture" \
  --image-urls assets/image/kitchen-3am.png --aspect-ratio 16:9 \
  --output-dir assets/image --filename-prefix dusk-watercolor

edit_image.mjs is the sibling for local file + highlighter- annotation edits. Use it when the source is a file on disk (not a URL) and the user has circled a region they want changed — the highlighter annotation is sent as a second image so the model knows which part to modify.

Video subcommands

generate-video.mjs has three subcommands — one per scenario:

bash
# 1. Text-to-video — default subcommand.
# Routes to bytedance/seedance-2.0/reference-to-video with zero refs
# (= pure t2v). Add --model veo3.1 to fall back to Google Veo 3.1.
node scripts/generate-video.mjs \
  --prompt "a serene bamboo forest with gentle wind" \
  --duration 4 --aspect-ratio 16:9 \
  --output assets/video/forest.mp4

# 2. First / last frame continuity — the `from-image` subcommand.
# Routes to bytedance/seedance-2.0/image-to-video by default.
# --image-url is the START frame; --end-image-url (seedance only)
# is an optional END frame for frame-to-frame interpolation.
node scripts/generate-video.mjs from-image \
  --prompt "the panda slowly rolls over and looks at the camera" \
  --image-url assets/images/panda-rolling.png \
  --duration 4 --aspect-ratio 16:9 \
  --output assets/video/panda-rolls.mp4

# 3. Multi-reference — the `reference` subcommand, seedance only.
# A compositional directing system: each ref is an addressable asset
# you assign a specific role in the prompt — character, first frame,
# destination environment, camera motion, style, audio bed.
# Addressing: 1-indexed by the order of the flag. First --image-url
# is @image1, second --image-url is @image2; videos and audios are
# numbered separately (@video1, @audio1, ...).
# Slots: up to 9 --image-url, 3 --video-url, 3 --audio-url (total ≤12).
# Audio refs require at least one image or video ref.
node scripts/generate-video.mjs reference \
  --prompt "Replace the character in @video1 with @image1, with @image1 as the first frame. Match the camera movement of @video1. Travel into the environment of @image2." \
  --image-url assets/image/hero.jpg         `# @image1: character` \
  --image-url assets/image/destination.jpg  `# @image2: destination` \
  --video-url assets/video/dolly-shot.mp4   `# @video1: camera grammar` \
  --duration 8 --aspect-ratio 16:9 \
  --output assets/video/shot.mp4

Full directive vocabulary (character / first frame / destination / camera transfer / style / prop / POV / audio) lives in references/reference-directives.md. Use it whenever more than one visual intent needs to be pinned down — a single image plus long prose prompt does not constrain seedance enough.

Shared flags across subcommands: --duration (required; 4–15 seconds or auto; veo3.1 only allows 4/6/8), --aspect-ratio (seedance: auto | 21:9 | 16:9 | 4:3 | 1:1 | 3:4 | 9:16; veo3.1: 16:9 | 9:16), --resolution (seedance: 480p | 720p; veo3.1: 720p | 1080p), --no-audio (disables generated audio — use when the content policy rejects auto-audio), --seed (integer, seedance only), --model seedance | veo3.1.

Seedance minimum duration is 4 seconds. Any beat shorter than that — a two-second reaction, a three-second punchline, a half-second sting — must still be generated at --duration 4 and then trimmed on the timeline by setting clip.outPoint lower than the source asset's full length. Plan beats in multiples of 4s when possible, and treat "I need 2 seconds here" as "I need 4 seconds of which I'll use the first 2". Never try to request --duration 2; the API will reject it, not silently clamp.

Content-policy retries

Seedance has two distinct content-filter failure modes. The script surfaces the full API response to stderr; match the error signature and apply the matching recovery:

  • loc:["body","image_urls"] + partner_validation_failed — image-side rejection. A photorealistic human face was detected in a reference. Run scripts/make-character-sheet.mjs on the photo to produce a photo-body / sketch-head 16:9 sheet, replace the --image-url with the sheet, and retry (add --no-audio on the retry too).
  • loc:["body","generated_video"] + "Output audio has sensitive content" — audio-side rejection. Video frames are fine; just retry the exact same command with --no-audio appended.

Full decision tree, fallback to --model veo3.1, and hard-limit notes (when to stop retrying and surface to the user) are in references/filter-retries.md. The character-consistency.md doc covers the sheet anatomy, honest limits, and prompt rules for the image-side case.

Why this shape: the script only does what a Bash subprocess does best — call a provider API and save bytes. Schema knowledge lives in this file + references/project-json.md, not in the scripts. You compose provenance yourself via Edit on project.json.

Audio layering — video tracks carry their own audio

Video tracks play their clips' embedded audio alongside audio-track clips. A seedance clip on a video track and a TTS clip on an audio track will both be audible at once. One-shot videos (generate once, get both picture and sound) work, but you have to plan whether a video's auto-audio should survive into the mix:

  • Picture only (common for b-roll where you'll add narration or BGM separately): pass --no-audio when generating with generate-video.mjs, OR mute the video track after the fact via the eye-beside-speaker icon on the track label (dispatches composition:toggle-track-mute, writes muted: true).
  • Picture + ambient (standalone clip, no competing audio): keep seedance's auto-audio.
  • Picture + dialogue from seedance: don't mute, but skip a separate narration track for that segment.

muted and visible on a track are orthogonal — muted governs audio only, visible governs picture only. Hiding a video track's picture means visible: false; silencing its audio means muted: true.

Typical workflow

  1. Read the current project.json to understand the composition.
  2. Generate assets by running one of the scripts with Bash.
  3. Register each new asset in project.json:
    • Add an entry to assets[] with a stable semantic id.
    • Add a matching edge to provenance[] with operation.type: "generate" and operation.params filled out.
  4. Place assets on the timeline by adding clips to the relevant track.
  5. The viewer auto-reflects every edit — no reload needed.

Full worked examples for the three most common flows are in references/workflows.md. When the user asks for a generation task, pattern-match the closest example there first, then adapt.

Character consistency (photorealistic humans)

When a specific human character appears — especially photorealistic, and especially across multiple shots — follow the protocol in references/character-consistency.md. Do not pass a photorealistic headshot or all-photo character sheet directly to generate-video.mjs reference: seedance 2.0's image-side filter rejects photorealistic human faces at input with partner_validation_failed, and prompt-side "virtual character" / "not a real person" phrasing does not defeat it (the filter does not read the prompt).

The verified-passing shape: a 16:9 sheet of 4 vertical panels, three of them photographic full-body views with the heads replaced by white-line pencil sketches, and a fourth panel holding a detailed pencil portrait plus typewriter-style OUTFIT / CHARACTER notes. Pass that sheet as the sole --image-url, use plain photorealistic prompt words (no "CG render" / "virtual character" — they degrade output quality), and always include --no-audio (seedance's output-audio filter rejects these generations at the second gate). Full recipe and honest-limits disclosures in the reference doc.

Gotchas

  • Metadata is for physical properties only. Put width, height, duration, fps, codec, sampleRate, channels in asset.metadata. Put prompt, model, seed, cost etc. in provenance.operation.params.
  • createdAt must be stable. When editing an existing asset, keep its createdAt unchanged — hydration relies on it.
  • Empty uri is legal for pending or generating assets. Set the uri when the script finishes and the file exists.
  • Never edit $schema. It's always "pneuma-craft/project/v1".
  • Time is in seconds. Not frames. fps only matters for playback/export.
  • fromAssetId: null means "from nothing", not "no lineage known". If the asset was generated from a text prompt alone, null is the correct value.
  • Clip ids are unique across all tracks, not per-track. Use semantic names so collisions are easy to avoid.

See also

Organized by the 6-layer technique stack:

Layer 1 — Production Bible (lock the world before generating)

  • references/production-bible.md — character cards (director-grade template), setting cards, project bible
  • references/character-consistency.md — seedance filter recovery: photo-body / sketch-head sheet for photoreal humans (orthogonal to bible)
  • references/craft.md — broader principles of short-video craft

Layer 2 — Storyboard Paths (choose A/B/C, structure the shot list)

  • references/storyboard-design.md — five-layer pre-pro design + per-panel template + delivery options A/B/C

Layer 3 — Direction Notation (precision in prompts and references)

  • references/direction-notation.md — production triggers, annotation color system, FACS, IPA, faithfulness directives, anti-patterns
  • references/reference-directives.md — @-addressing, multi-ref role vocabulary

Layer 4 — Iteration Workflow (sketch → anchor → clip)

  • references/storyboard-workflow.md — Path A/B/C, density recipes, draft exports, locator cards, atomic-edit rules

Layer 5 — Provenance Graph (lineage + audit trail)

  • references/project-json.md — full schema
  • references/asset-ids.md — id naming and stability rules

Layer 6 — Generation Tools (run the APIs, recover from rejection)

  • references/workflows.md — three end-to-end worked examples
  • references/filter-retries.md — decision tree for the two seedance 422 signatures
  • scripts/ — the seven bundled generator CLIs (generate_image.mjs, edit_image.mjs, generate-video.mjs, generate-tts.mjs, generate-bgm.mjs, make-character-sheet.mjs, storyboard.mjs)

© pandazki, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (scripts, references) in modes/clipcraft/skill of pandazki/pneuma-skills.

  • SKILL.md
  • references/asset-ids.md
  • references/character-consistency.md
  • references/craft.md
  • references/direction-notation.md
  • references/filter-retries.md
  • references/production-bible.md
  • references/project-json.md
  • references/reference-directives.md
  • references/storyboard-design.md
  • references/storyboard-workflow.md
  • references/workflows.md
  • scripts/generate-bgm.mjs
  • scripts/generate-video.mjs
  • scripts/make-character-sheet.mjs

Open the folder on GitHubat commit 0023d3c

Compare with similar skills

Pneuma Clipcraft next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pneuma Clipcraft compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pneuma Clipcraft this skillpandazki/pneuma-skills161—~7.5kAutomated safety check: NotesMIT
Suggest Sfxhassancs91/claude-youtube-editor328—~3kAutomated safety check: PassMIT
Ergo Remotion Videoitwanger/toBeBetterJavaer18k—~1.1kAutomated safety check: PassNone
Paper Collage Explainer Generatortl2012tl/comfyUI-llama-TE2414 repos~5.2kAutomated safety check: PassNone
RunninghubHM-RunningHub/OpenClaw_RH_Skills142—~1.6kAutomated safety check: PassApache-2.0
AI Video Production Assistantwanghui2323/ai-video-maker101—~924Automated safety check: PassMIT

Similar skills

  • Suggest Sfx

    hassancs91/claude-youtube-editor

    Step 4 of the AI Video Editor pipeline — the SFX pass. An agent skill from hassancs91/claude-youtube-editor.

    328 GitHub stars~3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Ergo Remotion Video

    itwanger/toBeBetterJavaer

    把口播稿做成二哥风格的 Remotion 视频,包括整理视频用稿、火山 TTS 配音、音画对齐、逐章动画预览和导出带配音的 MP4。用户说“做视频”“口播稿转视频”“Remotion”“继续做下一章”“出片”“渲染”“改读音”“配音读错了”,或给出 docs/src/ai/video/ 下的稿子要做成视频时使用。共享工具、配置和素材在…

    18k GitHub stars~1.1k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Paper Collage Explainer Generator

    tl2012tl/comfyUI-llama-TE

    For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics.

    241 GitHub starsUsed in 4 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Runninghub

    HM-RunningHub/OpenClaw_RH_Skills

    Generate images, videos, audio, and 3D models via RunningHub API (420 endpoints) and run any RunningHub AI Application (custom ComfyUI workflow) by webappId.

    142 GitHub stars~1.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • AI Video Production Assistant

    wanghui2323/ai-video-maker

    Turns an idea, article, outline or audio file into a sourced, reviewable AI video, tracking whether narration uses a human, synthetic or cloned voice.

    101 GitHub stars~924 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Media Gen

    clacky-ai/openclacky

    Generate or edit images, videos, or audio in the current task.

    1.2k GitHub stars~7.5k tokensUpdated today
    Media & CreativeAuto-check passed

More from pandazki/pneuma-skills

All 30 skills in this repo
  • Pneuma Bansho

    pandazki/pneuma-skills

    Explain something by writing it on a board. An agent skill from pandazki/pneuma-skills.

    161 GitHub stars~6.9k tokensUpdated 2 days ago
    Auto-check passed
  • Pneuma Lucid

    pandazki/pneuma-skills

    Pneuma Lucid Mode workspace guidelines. An agent skill from pandazki/pneuma-skills.

    161 GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check: warnings
  • Pneuma Plotwise

    pandazki/pneuma-skills

    Pneuma Plotwise workspace guidelines. An agent skill from pandazki/pneuma-skills.

    161 GitHub stars~8.9k tokensUpdated 2 days ago
    Auto-check passed
  • Pneuma Sprite

    pandazki/pneuma-skills

    Pneuma Sprite Mode workspace guidelines. An agent skill from pandazki/pneuma-skills.

    161 GitHub stars~16k tokensUpdated 2 days ago
    Auto-check passed
  • Pneuma Webcraft

    pandazki/pneuma-skills

    Pneuma WebCraft Mode workspace guidelines with Impeccable.style design intelligence.

    161 GitHub stars~7.5k tokensUpdated 2 days ago
    Auto-check: notes
  • Pneuma Wordtaste

    pandazki/pneuma-skills

    A goal-driven Chinese long-form writing partner. An agent skill from pandazki/pneuma-skills.

    161 GitHub stars~7.5k tokensUpdated 2 days ago
    Auto-check passed

Questions about Pneuma Clipcraft

What does Pneuma Clipcraft do?

AI-orchestrated video production on @pneuma-craft. An agent skill from pandazki/pneuma-skills. Pneuma Clipcraft is an agent skill from pandazki/pneuma-skills. AI-orchestrated video production on @pneuma-craft.

When should I use Pneuma Clipcraft?

Pneuma Clipcraft fits situations like: the user wants to generate; compose video clips; background music — text-to-video / image-to-video generation; music generation.

How do I install Pneuma Clipcraft in Claude Code?

Run `npx skills add pandazki/pneuma-skills --skill pneuma-clipcraft -a claude-code`. Or copy the skill folder (modes/clipcraft/skill in pandazki/pneuma-skills) into .claude/skills/pneuma-clipcraft in your project. Claude Code loads it when a task matches its description.

How do I install Pneuma Clipcraft in Codex?

Run `npx skills add pandazki/pneuma-skills --skill pneuma-clipcraft -a codex`. Or copy the skill folder (modes/clipcraft/skill in pandazki/pneuma-skills) into .agents/skills/pneuma-clipcraft in your project. Codex loads it when a task matches its description.

Can I use Pneuma Clipcraft in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pandazki/pneuma-skills --skill pneuma-clipcraft -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pneuma-clipcraft, .gemini/skills/pneuma-clipcraft, .github/skills/pneuma-clipcraft and .opencode/skills/pneuma-clipcraft in your project.

What does Pneuma Clipcraft need to run?

Going by SKILL.md and its folder, Pneuma Clipcraft needs JavaScript for the scripts in its folder, the command-line tools its instructions call (node and curl) and credentials named OPENROUTER_API_KEY and FAL_KEY. Our summary lists: Node.js.

Does Pneuma Clipcraft access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Pneuma Clipcraft safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Pneuma Clipcraft use?

Pneuma Clipcraft is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pneuma Clipcraft use?

About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 37k tokens, read only when the agent opens those files.

What are the alternatives to Pneuma Clipcraft?

Skills that share tags, products or a category with Pneuma Clipcraft: Suggest Sfx (hassancs91/claude-youtube-editor, 328 stars), Ergo Remotion Video (itwanger/toBeBetterJavaer, 18k stars), Paper Collage Explainer Generator (tl2012tl/comfyUI-llama-TE, 241 stars) and Runninghub (HM-RunningHub/OpenClaw_RH_Skills, 142 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pneuma Clipcraft?

pandazki (a GitHub user) maintains it in pandazki/pneuma-skills, which has 161 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on October 9, 2026.

Source: pandazki/pneuma-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.