Agent skill

Wjs Overlaying Video

by jianshuo in jianshuo/claude-skills

A skill your agent uses when the user has one or more video clips and wants to add post-production on top — AI-generated cover as first frame, HTML/CSS captions synced to SRT, kinetic illustration…

MITAuto-check passedMedia & Creative

Install Wjs Overlaying Video

skills CLI
$ npx skills add jianshuo/claude-skills --skill wjs-overlaying-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jianshuo/claude-skills wjs-overlaying-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jianshuo/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/wjs-overlaying-video .claude/skills/wjs-overlaying-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wjs-overlaying-video
GitHub stars
131
Token cost
~6.9k tokens
SKILL.md length
2,290 words
Files
7 (incl. scripts, references)
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user has one or more video clips and wants to add post-production on top — AI-generated cover as first frame, HTML/CSS captions synced to SRT, kinetic illustration…

  • Works in 10 steps: cover — full-frame AI image as first frame → caption — 关键词高亮 captions (字幕风格 03)… → chapter — top-left chapter chip (4s… → …
  • The user has one
  • SKILL.md covers When to use, What this skill IS — and IS NOT, The pipeline and Color: tone-map HLG/HDR source…, plus 8 more sections
  • Runs Python scripts from its folder; calls npx, python3 and ffmpeg; reaches fonts.googleapis.com and fonts.gstatic.com

What it does

Wjs Overlaying Video is an agent skill from jianshuo/claude-skills. Use when the user has one or more video clips and wants to add post-production on top — AI-generated cover as first frame, HTML/CSS captions synced to SRT, kinetic illustration overlays at hook moments, chapter chips, end-card CTA, or any other timed motion graphics. Most often used as the downstream of /wjs-segmenting-video — pick up where that skill stopped (raw cropped clip + per-clip SRT) and produce the upload-ready MP4. Backed by HyperFrames so everything compiles to ONE final encode — no cascade of…

Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/build_hf_clips.py`, `references/custom_overlay_recipes.md` and `references/example_spec.json`).

It sits in Media & Creative, covering Motion graphics, Video production and Transcription. It works with HeyGen. The repository describes itself as: 13 Claude Code skills for video production (transcribe / translate / dub / multicam / subtitles / reframe) + WeChat publishing. Compatible with Claude Code, OpenAI Codex CLI… The licence is MIT.

When your agent uses it

  • The user has one
  • More video clips and wants to add post-production on top — AI-generated cover as first frame
  • HTML/CSS captions synced to SRT
  • Kinetic illustration overlays at hook moments

Example prompts

  • “post-production”
  • “title card”
  • “kinetic captions”
  • “/wjs-overlaying-video”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. cover — full-frame AI image as first frame
  2. caption — 关键词高亮 captions (字幕风格 03) synced to SRT
  3. chapter — top-left chapter chip (4s reveal then fade)
  4. stack illustration — top-right vertical list card
  5. hammer illustration — center-frame big equation/text overlay
  6. cta — end-card with channel CTA
  7. Generate AI covers at the right aspect
  8. For each clip, scaffold a HyperFrames project
  9. Define illustrations per clip
  10. Build + render

What it can do on your machine

Read from SKILL.md and the folder at commit b2690f5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • python3
    • ffmpeg
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • fonts.googleapis.com
    • fonts.gstatic.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wjs Overlaying Video loads about 6.9k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 165 tokens; SKILL.md has 2,290 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~165
When it runs · the whole SKILL.md, loaded when a task matches
~6.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jianshuo/claude-skills at commit b2690f5, republished under its MIT licence (© jianshuo). 2,290 words, ~6,864 tokens.

Download SKILL.mdSave it as .claude/skills/wjs-overlaying-video/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
wjs-overlaying-video
description
Use when the user has one or more video clips and wants to add post-production on top — AI-generated cover as first frame, HTML/CSS captions synced to SRT, kinetic illustration overlays at hook moments, chapter chips, end-card CTA, or any other timed motion graphics. Most often used as the downstream of `/wjs-segmenting-video` — pick up where that skill stopped (raw cropped clip + per-clip SRT) and produce the upload-ready MP4. Backed by HyperFrames so everything compiles to ONE final encode — no cascade of re-encodes. Triggers — "加封面", "加字幕", "加动画", "加 CTA", "做后期", "post-production", "title card", "kinetic captions", "end card".

wjs-overlaying-video

Post-production for a video clip: cover, captions, illustrations, CTA, custom motion graphics — all composed in ONE HyperFrames project and rendered in a SINGLE final encode. No cascade of decodes/re-encodes (each cascade pass degrades quality and burns time).

When to use

  • Downstream of /wjs-segmenting-video — the segmentation skill hands you cropped clips + per-clip SRTs; this skill turns them into upload-ready MP4s with cover/captions/illustrations/CTA.
  • User has a finished video and wants to dress it up with motion graphics: opening hook, key-quote callout, closing slogan, chapter cards, AI-generated cover as first frame.
  • User wants HTML/CSS-quality captions on a video (kinetic word-by-word highlighting, custom fonts, large outlined text, seekable per cue).
  • User wants illustration overlays at specific hook moments — diagrams, big text emphasis, flow charts.

Don't use for:

  • Splitting one long video into clips → use /wjs-segmenting-video.
  • Creating the source SRT → use /wjs-transcribing-audio (then /wjs-translating-subtitles if you need a different language).
  • Full HyperFrames productions where the source isn't a fixed video → use hyperframes directly.
  • 微信视频号 / 抖音 upload (no public API for those) → this skill produces the MP4; upload is manual.

What this skill IS — and IS NOT

IsIs not
Everything that goes ON TOP of a video clip: cover, caption, chapter, illustration, CTACutting / cropping a video (that's /wjs-segmenting-video + /wjs-reframing-video)
One HyperFrames composition per clip = ONE final encodeA multi-step decode/encode cascade
cover is the literal first frame of the output (platforms auto-pick it as thumbnail)A separate thumbnail file the user uploads alongside
Captions are HTML/CSS — -webkit-text-stroke for white-on-anything readabilitylibass burn-in (deprecated)
Illustrations: re-usable stack / hammer patterns + custom escape hatchOne bespoke HTML/CSS per illustration without re-use
AI covers regenerated at native target aspect (1024×1792 for vertical, 1536×1024 for horizontal)Single 1024×1536 default that letterboxes or crops on the platform

The pipeline

clip.mp4 + clip.zh-CN.burn.srt   (from /wjs-segmenting-video hand-off)
   ↓
1. (Optional) Generate AI cover via gpt-image-2
   make_cover.py --segments S.json --out output/ --size 1024x1792
   cover_NN_slug.png

2. Scaffold a HyperFrames project per clip
   hf_clip_NN/1080/{index.html, clip.mp4, cover.png, captions.json}

3. Compose: cover scene + body video + caption track + chapter chip
            + 1-2 illustrations at hook moments + CTA scene

4. npm run check (lint + validate + visual inspect)
   npm run render → upload-ready MP4

A 2-minute vertical 1080×1920 composition renders in ~2-3 min on M-series Mac.

Color: tone-map HLG/HDR source → SDR BEFORE compositing

Only tone-map genuinely HLG/HDR sources. If the body clip is ALREADY Rec.709 SDR — e.g. a graded multicam render, or polysync output where an S-Log3→709 LUT was already applied — running the HLG tone-map recipe on it washes/darkens the already-correct color. build_hf_clips.py's tonemap_to_sdr now probes color_transfer (_is_hlg_hdr): HLG/PQ → tone-map; otherwise a straight re-encode with dense keyframes (no tone-map). Either way you still get the -g 30 dense-keyframe encode HyperFrames needs.

iPhone / modern-camera footage is often HLG HDR (bt2020 / arib-std-b67). If you feed that straight into HyperFrames it either renders washed-out ("发白") or, with a naive --sdr, too dark ("发黑"); and the HDR x265 path can hang the renderer. Pre-convert the body clip to SDR (bt709) 30fps h264 with a locked zscale tone-map, then composite the SDR clip.

The verified recipe (tonemap_to_sdr() in build_hf_clips.py). npl=203 matches macOS-native (qlmanage) reference brightness; hable keeps contrast; this preserves the ORIGINAL look (natural skin / foliage / brick), no wash, no darkening:

python
# zscale-capable ffmpeg — Homebrew's lacks zscale/tonemap.
# imageio-ffmpeg ships one: .../imageio_ffmpeg/binaries/ffmpeg-macos-aarch64-v7.1
TONEMAP_VF = ("zscale=tin=arib-std-b67:min=bt2020nc:pin=bt2020:t=linear:npl=203,"
              "format=gbrpf32le,tonemap=tonemap=hable:desat=0,"
              "zscale=t=bt709:m=bt709:p=bt709:r=tv,format=yuv420p,fps=30")
# encode: libx264 -crf 18 -color_primaries/-trc/-colorspace bt709
#         -g 30 -keyint_min 30 -movflags +faststart   ← see gotcha below

Dense-keyframe gotcha. HyperFrames seeks the body video frame-by-frame. A clip with sparse keyframes (long GOP) makes it freeze on stale frames — the render log warns Video "video" has sparse keyframes. Always encode the SDR clip with -g 30 -keyint_min 30 (one keyframe per frame-second) so every seek lands clean.

Verify the render log says No HDR sources detected — rendering SDR. If it says HDR detected, your clip wasn't tone-mapped — fix that first.

Version stamp (every output)

Stamp 「skill名字 + 版本号」 bottom-right, shown during the END/CTA scene, so every render is traceable to the pipeline version that made it. Bump VERSION in build_hf_clips.py on each pipeline change.

css
#ver-stamp { position: absolute; right: 28px; bottom: 28px; z-index: 30;
  font-size: 20px; color: rgba(150,150,156,0.55); letter-spacing: 0.06em; }
html
<div id="ver-stamp" class="clip" data-start="{cta_start}" data-duration="{cta_dur}"
     data-track-index="2">wjs-overlaying-video v1.3</div>

Standard overlay types (the 6 building blocks)

Every clip's final composition is built from some combination of these. The agent picks the right ones per clip — typically all 6 for a podcast highlight, or just 1-2 for a single annotation overlay.

1. cover — full-frame AI image as first frame

The cover IS the first frame (no animation, no zoom) so platforms that auto-pick the first frame as the thumbnail get your designed cover by default. Always verify with ffmpeg -ss 0 -vframes 1 — frame 0 must NOT be black or platform thumbnails will be black.

HTML:

html
<div id="cover" class="clip" data-start="0" data-duration="1.6"
     data-track-index="1" data-layout-allow-overflow>
  <img src="cover.png" alt="" data-layout-allow-overflow />
</div>

CSS:

css
#cover { position: absolute; inset: 0; background: #0c0d10; overflow: hidden; }
#cover img { position: absolute; inset: 0; width: 100%; height: 100%; object-fit: cover; }

Generation: use /wjs-segmenting-video/scripts/make_cover.py (wraps gpt-image-2 images edit with the midpoint frame as ref):

bash
# For 1080×1920 vertical output (视频号 / 抖音):
make_cover.py --segments S.json --out output/ --size 1024x1792 [--single N]

# For 1920×1080 horizontal output (YouTube / B站):
make_cover.py --segments S.json --out output/ --size 1536x1024

Aspect must match output frame. --size 1024x1536 (2:3, the script default) gets letterboxed or cropped on 9:16 output — always pass 1024x1792 for vertical. The cover image's aspect is what the viewer sees full-frame, so mismatch is visible. Re-roll one with --single N; codex provider can transient-fail mid-batch.

Codex auth required: the script calls codex CLI via gpt-image-2-skill. If ~/.codex/auth.json is missing, the script errors. See gpt-image-2-skill for setup.

Reference frame must match the OUTPUT orientation. make_cover reads output/frame_NN_slug.jpg as the photographic background it keeps. For a vertical clip that came from a horizontal two-person source, the default frame_NN is the horizontal two-shot — feeding that to a 1024x1792 cover crams both people into portrait awkwardly. Replace frame_NN_slug.jpg with a vertical single-speaker frame pulled from the already-cropped body clip first (ffmpeg -ss <t> -i clip_vert.mp4 -frames:v 1 frame_NN_slug.jpg), then run make_cover. The cover then matches the body framing.

Baked-title cover ⇒ drop the animated #hook opener. make_cover stamps the segment title into the cover image (white fill + heavy black stroke, placed clear of faces). That cover IS the title card. Do NOT also run the animated #hook opener over it (overlay type below) — you'd double-stamp the title. Pick one: either a make_cover baked-title cover (then leave HOOK empty), or a plain video-frame cover + animated hook. The house default the user approved is the make_cover baked-title cover (a clean video frame with the title burned in, no AI painting).

2. caption — 关键词高亮 captions (字幕风格 03) synced to SRT

Chosen style for 王建硕 (user-approved): 字幕风格 03「关键词高亮」+ 思源宋体 Noto Serif SC. Serif white text with a black stroke, and punchy QUANTITATIVE keywords (倍数 / 大数量级 / 百分比) wrapped in a small gold gradient block. Captions are vertically centered in a fixed zone (so 1-line vs 2-line cues don't make the visual center jump up and down).

There were 4 candidate styles (描边白字 / 质感底条 / 关键词高亮 / 逐字点亮); the user picked 03 关键词高亮 with serif sc font. Use that. The plain-stroke style (-webkit-text-stroke: 5px #000, no gold block, sans font) is the fallback if a clip has no quantitative keywords to highlight.

Font — load Noto Serif SC from Google Fonts in <head> (the HyperFrames compiler fetches & inlines requested Google font families automatically; verify the render log says Fetched … Noto Serif SC):

html
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Noto+Serif+SC:wght@600;700;900&display=swap" rel="stylesheet">

HTML:

html
<div id="caption" class="clip" data-start="{body_start}"
     data-duration="{body_dur}" data-track-index="4"></div>

CSS (vertical 1080×1920) — 字幕风格 03:

css
#caption {
  position: absolute; left: 0; right: 0; bottom: 240px;
  height: 240px; z-index: 10; overflow: visible;
}
#caption .bubble {
  position: absolute; top: 50%; left: 50%;
  display: inline-block; padding: 0 24px;
  font-family: "Noto Serif SC", "Songti SC", "STSong", serif;
  font-size: 52px; line-height: 1.32; font-weight: 700;
  color: #fff; max-width: 980px; text-align: center;
  -webkit-text-stroke: 2.5px rgba(0,0,0,0.9);
  paint-order: stroke fill;
  text-shadow: 0 2px 8px rgba(0,0,0,0.7), 0 0 2px rgba(0,0,0,0.9);
  letter-spacing: 0.01em;
}
#caption .bubble .hot {        /* gold keyword block */
  color: #1a1206; -webkit-text-stroke: 0;
  background: linear-gradient(180deg, #f3c877, #c79655);
  padding: 2px 12px; border-radius: 9px; margin: 0 3px;
  box-shadow: 0 3px 10px -3px rgba(232,176,99,0.6);
}

Keyword auto-selection (sparse on purpose). Wrap only genuinely emphatic magnitudes so the gold block stays meaningful, not noisy. Deliberately EXCLUDE generic 个/年 ("一个", "20年"). Handles thousands-commas ("1,000万"). build_hf_clips.py does this in mark_keywords():

python
_NUM = r"[0-90-9,,一二三四五六七八九十百千两零几]+"
_HOT_RE = re.compile(rf"(?:翻了?{_NUM}?[倍番]|{_NUM}\s*(?:[倍番]|万亿?|亿|%|%))")
# → highlights: 一倍 五六倍 十倍 10倍 50万 800万 1,000万 50% 翻一倍
# render the cue with b.innerHTML = g.html (HTML-escape the non-keyword text)

JS (one bubble per cue + GSAP fade in/out, all centered at container midpoint):

js
// SRT cues are loaded as inline JSON. Each cue's start/end is offset
// by the cover-scene duration (e.g., 1.5s) so the timing aligns with
// the composition timeline (not the body's own t=0).
const captionEl = document.getElementById("caption");
const groups = JSON.parse(document.getElementById("captions-data").textContent);
const bubbles = groups.map((g, i) => {
  const b = document.createElement("span");
  b.className = "bubble"; b.id = "cap-" + i;
  b.innerHTML = g.html || g.text;   // g.html has <span class="hot"> keyword blocks
  b.style.opacity = "0";
  captionEl.appendChild(b);
  return b;
});
// GSAP xPercent/yPercent for centering (CSS transform would get
// overwritten the moment we tween y).
gsap.set(bubbles, { xPercent: -50, yPercent: -50 });
groups.forEach((g, i) => {
  const el = bubbles[i];
  tl.fromTo(el, { opacity: 0, y: 12 }, { opacity: 1, y: 0, duration: 0.18, ease: "power2.out" }, g.start);
  const exitStart = Math.max(g.start + 0.18, g.end - 0.12);
  tl.to(el, { opacity: 0, duration: 0.12, ease: "power2.in" }, exitStart);
  tl.set(el, { opacity: 0 }, g.end);
});

Source SRT — slice + shift before inlining. Prefer the word-timed .asr.srt built by /wjs-transcribing-audio (火山 streaming ASR → build_srt_from_asr.py) — its per-word timing means cues sit exactly on the spoken audio with no drift. Parse each cue, add the cover duration to every start/end, run mark_keywords() to produce the html field, and inline as JSON in a <script id="captions-data" type="application/json"> block.

MarginV / position notes:

  • Vertical (1080×1920): bottom: 240px keeps captions clear of the 视频号/抖音 bottom UI overlay (likes/comments/share buttons).
  • Horizontal (1920×1080): bottom: 100px, font-size: 48px, -webkit-text-stroke: 4px is a reasonable default.

Caption length cap. If a single cue exceeds ~18 Chinese chars on 1080-wide at 56px, it wraps to 2 lines awkwardly. This is upstream discipline — /wjs-translating-subtitles should cap cues at ~18 chars using word-gap split + punctuation split. If you receive longer cues, either reduce font-size to 48px or accept the wrap.

3. chapter — top-left chapter chip (4s reveal then fade)

A subtle badge identifying the segment. Enters at body start, fades after a few seconds so it doesn't compete with the rest of the composition.

HTML:

html
<div id="chapter" class="clip" data-start="{body_start}"
     data-duration="{body_dur}" data-track-index="3">
  <span class="dot"></span>
  <span class="text">第一段 · 自然语言才是新代码</span>
</div>

CSS:

css
#chapter {
  position: absolute; top: 80px; left: 60px; z-index: 9;
  display: inline-flex; align-items: center; gap: 12px;
  padding: 12px 20px;
  background: rgba(12,13,16,0.78);
  border: 1px solid rgba(199,150,85,0.4);
  border-radius: 999px;
}
#chapter .dot { width: 10px; height: 10px; border-radius: 999px; background: #e8b063; }
#chapter .text {
  font-size: 24px; color: #f4f4f5; letter-spacing: 0.04em; font-weight: 600;
}

GSAP:

js
tl.from("#chapter", { x: -40, opacity: 0, duration: 0.5, ease: "expo.out" }, body_start + 0.4);
tl.to("#chapter", { opacity: 0, duration: 0.4, ease: "power2.in" }, body_start + 4.0);
4. stack illustration — top-right vertical list card

A list of items (e.g., language hierarchy, workflow steps, levels) in a dark card at the top-right. One item can be accented in amber to highlight the relevant level/step.

Use for: showing a hierarchy or list while the speaker explains it. Card stays visible 8-50s.

HTML:

html
<div id="ill-stack" class="clip" data-start="{start}" data-duration="{dur}" data-track-index="5">
  <div class="ill-card">
    <div class="ill-card-label">我们写的层级</div>
    <div class="ill-row"><span class="ill-tag accent">自然语言</span></div>
    <div class="ill-row"><span class="ill-tag">Python</span></div>
    <div class="ill-row"><span class="ill-tag">C</span></div>
    <div class="ill-row"><span class="ill-tag">Assembly</span></div>
  </div>
</div>

CSS: (see references/illustration_patterns.md for the full canonical CSS — copy verbatim)

GSAP — slide in from right + stagger rows:

js
tl.fromTo("#ill-stack", { x: 360, opacity: 0 }, { x: 0, opacity: 1, duration: 0.6, ease: "expo.out" }, start + 0.2);
tl.from("#ill-stack .ill-row", { y: 20, opacity: 0, duration: 0.4, stagger: 0.12, ease: "power2.out" }, start + 0.4);
tl.to("#ill-stack", { x: 360, opacity: 0, duration: 0.5, ease: "power2.in" }, end - 0.5);
5. hammer illustration — center-frame big equation/text overlay

A BIG center-frame text/equation that visually "hammers" a key claim. Best for the single most quotable moment in a clip (e.g., "LLM = 编译器", "Token = 新 GDP", "AI ≠ 更快的轿子"). Visible 4–8s.

HTML:

html
<div id="ill-hammer" class="clip" data-start="{start}" data-duration="{dur}" data-track-index="6">
  <div class="ill-h-content">
    <div class="ill-h-eq">
      <span class="ill-h-left">LLM</span>
      <span class="ill-h-equals">=</span>
      <span class="ill-h-right">新编译器</span>
    </div>
    <div class="ill-h-foot">自然语言 → Python → 汇编</div>
  </div>
</div>

GSAP — scale-pop entrance + stagger each piece + scale-fade exit:

js
tl.fromTo("#ill-hammer", { scale: 0.85, opacity: 0 },
  { scale: 1.0, opacity: 1, duration: 0.45, ease: "back.out(1.6)" }, start);
tl.from("#ill-hammer .ill-h-left", { x: -40, opacity: 0, duration: 0.4, ease: "expo.out" }, start + 0.2);
tl.from("#ill-hammer .ill-h-equals", { scale: 0, opacity: 0, duration: 0.4, ease: "back.out(2)" }, start + 0.4);
tl.from("#ill-hammer .ill-h-right", { x: 40, opacity: 0, duration: 0.4, ease: "expo.out" }, start + 0.6);
tl.from("#ill-hammer .ill-h-foot", { y: 20, opacity: 0, duration: 0.4, ease: "power2.out" }, start + 0.8);
tl.to("#ill-hammer", { scale: 1.05, opacity: 0, duration: 0.45, ease: "power2.in" }, end - 0.45);

(see references/illustration_patterns.md for full canonical CSS)

Show full SKILL.md (894 more words)Show less
6. cta — end-card with channel CTA

A branded outro for the final 3 seconds. Use 王建硕 as the channel name (per global instructions) — never put a guest's name in the CTA slot.

HTML:

html
<div id="cta" class="clip" data-start="{cta_start}" data-duration="3.24" data-track-index="1">
  <div class="cta-line-1">关注王建硕</div>
  <div class="arrow">↓</div>
  <div class="cta-line-2">微信公众号 · 视频号</div>
  <div class="cta-foot">聊 AI · 聊创业 · 持续更新</div>
</div>

CSS / GSAP: see references/illustration_patterns.md.

Legacy types (for one-off overlays on a single video)

The spec.json + scaffold.py workflow also supports these older overlay types — useful when you want to dress up ONE existing video without going through the full post-production workflow above:

  • quote — full-width kinetic typography, top or bottom gradient. Best for opening hooks and key-quote callouts.
  • slogan — alias for quote with position: bottom and larger type. Best for closing slogans.
  • callout — small annotation panel in a corner. Best for chapter labels, lower-thirds, "as seen in" notes.
  • custom — escape hatch. Claude writes the overlay's HTML/CSS/GSAP inside an overlays/<name>.html fragment file. See references/custom_overlay_recipes.md.

Workflow A — Post-segmentation preset (most common)

Use this when you're coming directly from /wjs-segmenting-video and want the standard cover + caption + chapter + illustrations + CTA treatment for each clip.

Step 1 — Generate AI covers at the right aspect
bash
# For vertical 9:16 output (视频号 / 抖音):
python3 ~/.claude/skills/wjs-segmenting-video/scripts/make_cover.py \
    --segments segments.json --out output/ --size 1024x1792 --single 1
# Verify segment 1's cover; then batch:
python3 ~/.claude/skills/wjs-segmenting-video/scripts/make_cover.py \
    --segments segments.json --out output/ --size 1024x1792
Step 2 — For each clip, scaffold a HyperFrames project

hf_clip_NN/1080/ with:

  • index.html — the composition (from template; see references/post_segmentation_template.html)
  • clip.mp4 — copied from output/clip_NN_slug.mp4
  • cover.png — copied from output/cover_NN_slug.png
  • captions.json — generated from output/clip_NN_slug.zh-CN.burn.srt with every cue's start/end shifted by +cover_duration (so cues align with the composition timeline, not the body's own clock)

The build script at references/build_hf_clips.py does this for all segments in one pass. It reads segments.json + an ILLUSTRATIONS dict (illustrations per clip, see Step 3) + the template, and emits 5 ready-to-render projects.

Step 3 — Define illustrations per clip

For each clip, identify 1-2 hook moments and pick stack or hammer:

python
ILLUSTRATIONS = {
    1: [
        # The language hierarchy as a stack card during the opening
        {"key": "stack", "pattern": "stack", "body_start": 0.3, "body_end": 9.0,
         "label": "我们写的层级",
         "rows": [
             {"text": "自然语言", "accent": True},
             {"text": "Python",   "accent": False},
             {"text": "C",        "accent": False},
             {"text": "Assembly", "accent": False},
         ]},
        # The hammer at the most quotable moment
        {"key": "hammer", "pattern": "hammer", "body_start": 10.8, "body_end": 14.6,
         "left": "LLM", "equals": "=", "right": "新编译器",
         "foot": "自然语言 → Python → 汇编"},
    ],
    # ... clips 2-5
}

Timestamps are body-relative (after the cover-scene duration); the build script adds the cover offset when emitting GSAP positions.

Step 4 — Build + render
bash
python3 references/build_hf_clips.py    # scaffolds all projects
for n in 01 02 03 04 05; do
  cd "hf_clip_$n/1080"
  npx hyperframes lint
  npx hyperframes validate
  npx hyperframes render
  cd ../..
done

A 2:30 clip renders in ~3 min. Output: hf_clip_NN/1080/renders/*.mp4.

Workflow B — Custom overlays on a single video (legacy spec.json)

Use this when you have ONE existing video and want to add a few ad-hoc overlays (title cards, annotations, lower-thirds).

spec.json schema
json
{
  "source_video": "../path/to/source.mp4",
  "duration": 135.4,
  "size": "1920x1080",
  "name": "clip_01_animated",
  "overlays": [
    {"id": "o1", "type": "quote", "start": 8.0, "duration": 6.0,
     "position": "top", "lines": ["代码不存在错误", "只存在意图错配"],
     "accent": [false, true]},
    {"id": "o2", "type": "callout", "start": 30.0, "duration": 5.0,
     "anchor": "top-right", "text": "FRP 概念"},
    {"id": "o3", "type": "slogan", "start": 122.0, "duration": 13.4,
     "lines": ["改 prompt", "不改 AI 生成的代码"], "accent": [false, true]}
  ]
}
FieldRequiredNotes
source_videoYesPath to source MP4. Symlinked into the project as source.mp4.
durationYesTotal composition length in seconds — match the source video.
sizeNoWIDTHxHEIGHT (default 1920x1080).
overlays[].typeYesquote, slogan, callout, or custom.
overlays[].startYesStart time in seconds.
overlays[].durationYesHow long the overlay is on screen.
Scaffold + render
bash
python3 ~/.claude/skills/wjs-overlaying-video/scripts/scaffold.py spec.json
cd <name> && npm run check && npm run render

Output checklist

Before considering a clip done:

  • Frame 0 is the cover (not black) — ffmpeg -ss 0 -vframes 1 out.mp4
  • Captions are synced with audio (lint a few seconds with audio playback)
  • All illustrations enter and exit at the speech moments they support
  • CTA renders correctly (关注王建硕, not a guest's name)
  • npx hyperframes lint && npx hyperframes validate both pass
  • npx hyperframes inspect shows no layout overflow
  • Total duration matches the source clip + cover + CTA durations

Common mistakes

  • Cover aspect ≠ output aspect. 1024x1536 (the default make_cover.py size) is 2:3 and gets letterboxed or cropped on 9:16 output. Always pass --size 1024x1792 for vertical.
  • Caption alignment jumps with line count. Anchor by CENTER (translate(-50%, -50%)) inside a fixed-height container so 1-line vs 2-line cues share the same visual midline. NOT anchored from bottom (causes growth-upward).
  • GSAP overwrites CSS transform centering. If you set transform: translate(-50%, -50%) in CSS and then tween y, GSAP replaces the transform and centering breaks. Use gsap.set(el, { xPercent: -50, yPercent: -50 }) instead so xPercent/yPercent compose with subsequent y/x tweens.
  • Burning libass subs on top of HTML/CSS captions. Pick ONE caption system per output video. If you're using this skill's HTML/CSS captions, do NOT also burn subs in /wjs-segmenting-video — request the raw clip via the hand-off package.
  • Frame 0 is black. If your cover scene has an opacity fade-in starting from 0, the literal first frame is black and the platform thumbnail will be black. Place the cover statically (no opacity tween) and verify with ffmpeg -ss 0 -vframes 1.
  • Channel name in CTA = guest's name. Always use 王建硕. Guests belong in description text inside the metadata, not in the on-screen CTA.
  • Cover image cropped because of object-fit: cover on mismatched aspect. Either regenerate the cover at the right aspect (see Step
    1. or letterbox with object-fit: contain + dark background.

Integration with other skills

  • /wjs-segmenting-video — the typical upstream. After it cuts
    • crops + slices SRTs, this skill picks up. The hand-off package is clip_NN.mp4 + clip_NN.zh-CN.burn.srt + segments.json.
  • /wjs-transcribing-audio + /wjs-translating-subtitles — if no SRT exists, run them first. The word-level Whisper or Volcano/豆包 ASR output is preferred for accurate cue timing.
  • hyperframes — the underlying composition framework. This skill is a thin wrapper that encodes the proven post-production patterns; everything in the hyperframes skill applies (preview, render, transitions, audio-reactive, etc.). Read it whenever you write custom overlays.
  • hyperframes-cli — the CLI commands the project uses (init, lint, validate, inspect, render).
  • gpt-image-2-skill — the cover generator. make_cover.py invokes it via the codex CLI; the codex auth in ~/.codex/auth.json is required.
  • /wjs-uploading-video — the next downstream after this skill produces an MP4. Uploads the renders to YouTube with title / description / tags from a metadata file.

Files & references

  • scripts/scaffold.py — Workflow B scaffolder (legacy spec.json for ad-hoc overlays)
  • references/post_segmentation_template.html — Workflow A template: the canonical cover + caption + chapter + illustration + CTA composition shape, with placeholder substitutions
  • references/build_hf_clips.py — Workflow A multi-clip builder. Reads segments.json + per-clip illustrations dict, scaffolds and populates one project per clip
  • references/illustration_patterns.md — canonical CSS / GSAP for the stack and hammer illustration patterns
  • references/custom_overlay_recipes.md — reusable custom overlay recipes (terminal demo, layer-stack diagram, callout with arrow)
  • references/example_spec.json — Workflow B example

© jianshuo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in wjs-overlaying-video of jianshuo/claude-skills.

  • SKILL.md
  • references/build_hf_clips.py
  • references/custom_overlay_recipes.md
  • references/example_spec.json
  • references/illustration_patterns.md
  • references/illustrations.py
  • scripts/scaffold.py

Open the folder on GitHubat commit b2690f5

Compare with similar skills

Wjs Overlaying Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wjs Overlaying Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wjs Overlaying Video this skilljianshuo/claude-skills131—~6.9kAutomated safety check: PassMIT
Short Form Editnateherkai/hyperframes-student-kit1.3k—~5.3kAutomated safety check: PassCustom licence
Yuv Viral Videohoodini/ai-agents-skills282—~7.5kAutomated safety check: NotesNone
3D Camera Captions for HyperFramesheygen-com/hyperframes-community-skills187—~2.3kAutomated safety check: PassApache-2.0
Edit Videonateherkai/hyperframes-student-kit1.3k—~673Automated safety check: PassCustom licence
Embedded Video Captionsheygen-com/hyperframes60k3 repos~8.6kAutomated safety check: PassApache-2.0

Similar skills

  • Short Form Edit

    nateherkai/hyperframes-student-kit

    Turn talking-head footage into a finished reel, YouTube Short, or short advertisement with curiosity-led openings, earned payoffs, story-driven cuts, transcript-synced motion graphics, moving…

    1.3k GitHub stars~5.3k tokensUpdated 13 days ago
    Media & CreativeAuto-check passed
  • Yuv Viral Video

    hoodini/ai-agents-skills

    Edit any selfie or screen-share footage into a viral short-form video in YUV.AI's signature style — Apple-style liquid-glass cards (real CSS backdrop-filter), dark-mode polish, MrBeast-paced cuts…

    282 GitHub stars~7.5k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • 3D Camera Captions for HyperFrames

    heygen-com/hyperframes-community-skills

    Builds captions that move in 3D space around a talking head: a virtual camera flies past words at different depths, hero words hide behind the speaker, and text steps at 15 fps with motion blur.

    187 GitHub stars~2.3k tokensUpdated 12 days ago
    Media & CreativeAuto-check passed
  • Edit Video

    nateherkai/hyperframes-student-kit

    Edit a raw talking-head video through transcription, silence trimming, mistake review, visual storytelling, motion graphics, and verified HyperFrames rendering.

    1.3k GitHub stars~673 tokensUpdated 13 days ago
    Media & CreativeAuto-check passed
  • Embedded Video Captions

    heygen-com/hyperframes

    Adds captions to a single-subject talking-head video without editing the footage, from plain subtitles to cinematic text placed behind the speaker.

    60k GitHub starsUsed in 3 repos~8.6k tokens
    Media & CreativeAuto-check passed
  • Talking-Head Video Pipeline

    naive-kun/naive-video-skill

    Turns raw or rough-cut talking-head footage into a captioned, animated final video through a staged, resumable production pipeline.

    132 GitHub stars~3.3k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed

More from jianshuo/claude-skills

All 38 skills in this repo
  • Wjs Segmenting Video

    jianshuo/claude-skills

    A skill your agent uses when the user has a long-form video (interview / lecture / podcast / conversation) and a transcript SRT, and wants to extract 3–6 stand-alone topical short clips from it.

    131 GitHub stars~3.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Uploading Video

    jianshuo/claude-skills

    Upload one or many videos to YouTube. An agent skill from jianshuo/claude-skills.

    131 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Converting Wp To Hugo

    jianshuo/claude-skills

    A skill your agent uses when migrating a WordPress site to a Hugo static site on GitHub Pages from a WXR export (.xml) plus the wp-content/uploads folder — preserving /archives/<id/ URLs, localizing…

    131 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Burning Subtitles

    jianshuo/claude-skills

    A skill your agent uses when the user has a video + an SRT and wants the subtitles either burned into the pixels (libass, always-visible) or soft-muxed as a togglable track.

    131 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Cleaning Spam

    jianshuo/claude-skills

    A skill your agent uses when the user complains about spam on his X/Twitter posts — 同城面付 / 寻固炮 / 线下上门 / 免费破处 这类引流号在他推文下刷的 emoji 垃圾回复 — and wants them removed.

    131 GitHub stars~532 tokensUpdated 1 mo ago
    Auto-check passed
  • Wjs Creating Video Book

    jianshuo/claude-skills

    A skill your agent uses when the user wants a book turned into YouTube chapter videos — 每章用 VoiceDrop 读书的有声书 mp3 做音轨,配 GPT Image 2 画面和中心思想大字,输出 1920×1080 横屏视频发 YouTube。Triggers — "把这本书做成视频"…

    131 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Wjs Overlaying Video

What does Wjs Overlaying Video do?

A skill your agent uses when the user has one or more video clips and wants to add post-production on top — AI-generated cover as first frame, HTML/CSS captions synced to SRT, kinetic illustration…. Wjs Overlaying Video is an agent skill from jianshuo/claude-skills. Use when the user has one or more video clips and wants to add post-production on top — AI-generated cover as first frame, HTML/CSS captions synced to SRT, kinetic illustration overlays at hook moments, chapter chips, end-card CTA, or any other timed motion graphics.

When should I use Wjs Overlaying Video?

Wjs Overlaying Video fits situations like: the user has one; more video clips and wants to add post-production on top — AI-generated cover as first frame; HTML/CSS captions synced to SRT; kinetic illustration overlays at hook moments.

How do I install Wjs Overlaying Video in Claude Code?

Run `npx skills add jianshuo/claude-skills --skill wjs-overlaying-video -a claude-code`. Or copy the skill folder (wjs-overlaying-video in jianshuo/claude-skills) into .claude/skills/wjs-overlaying-video in your project. Claude Code loads it when a task matches its description.

How do I install Wjs Overlaying Video in Codex?

Run `npx skills add jianshuo/claude-skills --skill wjs-overlaying-video -a codex`. Or copy the skill folder (wjs-overlaying-video in jianshuo/claude-skills) into .agents/skills/wjs-overlaying-video in your project. Codex loads it when a task matches its description.

Can I use Wjs Overlaying Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jianshuo/claude-skills --skill wjs-overlaying-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wjs-overlaying-video, .gemini/skills/wjs-overlaying-video, .github/skills/wjs-overlaying-video and .opencode/skills/wjs-overlaying-video in your project.

What does Wjs Overlaying Video need to run?

Going by SKILL.md and its folder, Wjs Overlaying Video needs Python for the scripts in its folder and the command-line tools its instructions call (npx, python3, ffmpeg and npm). Our summary lists: Python 3.

Does Wjs Overlaying Video access the network?

SKILL.md names 2 domains. In commands or code: fonts.googleapis.com and fonts.gstatic.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Wjs Overlaying Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Wjs Overlaying Video use?

Wjs Overlaying Video is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wjs Overlaying Video use?

About 6.9k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.

What are the alternatives to Wjs Overlaying Video?

Skills that share tags, products or a category with Wjs Overlaying Video: Short Form Edit (nateherkai/hyperframes-student-kit, 1.3k stars), Yuv Viral Video (hoodini/ai-agents-skills, 282 stars), 3D Camera Captions for HyperFrames (heygen-com/hyperframes-community-skills, 187 stars) and Edit Video (nateherkai/hyperframes-student-kit, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wjs Overlaying Video?

jianshuo (a GitHub user) maintains it in jianshuo/claude-skills, which has 131 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on August 20, 2026.

Source: jianshuo/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.