Agent skill

Clipify

by ZJU-REAL in ZJU-REAL/Easel

从长视频中自动提取精彩片段,切成独立短视频,支持 16:9→9:16 竖版转制和逐字字幕烧录. An agent skill from ZJU-REAL/Easel.

Apache-2.0Auto-check passedMedia & Creative

Install Clipify

skills CLI
$ npx skills add ZJU-REAL/Easel --skill clipify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZJU-REAL/Easel clipify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZJU-REAL/Easel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/openclaw/clipify .claude/skills/clipify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
clipify
GitHub stars
3.4k
Token cost
~2.2k tokens
SKILL.md length
888 words
Files
7 (incl. scripts, assets)
Skills in repo
114
Repo updated
First seen
Licence
Apache-2.0

At a glance

从长视频中自动提取精彩片段,切成独立短视频,支持 16:9→9:16 竖版转制和逐字字幕烧录. An agent skill from ZJU-REAL/Easel.

  • Works in 6 steps: Find the funniest parts → Trim each chosen clip → Decide the output format → …
  • Media & Creative work in your project
  • SKILL.md covers Inputs, Tooling (use only the fastest…, Workflow and Pitfalls (lessons from prior…
  • Runs Python scripts from its folder; calls ffmpeg, whisper and python3

What it does

Clipify is an agent skill from ZJU-REAL/Easel. 从长视频中自动提取精彩片段,切成独立短视频,支持 16:9→9:16 竖版转制和逐字字幕烧录。 当用户说"视频切片""提取精彩片段""长视频切短""切成短视频""高光剪辑""逐字字幕""转竖版短视频"时使用。 和 video-highlights 的区别:clipify 专做英文口播找笑点+动态人脸 pan;video-highlights 更通用(中文/直播皆可),静态转竖版更稳。

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and assets (for example `EASEL-META.md`, `scripts/analyze.py` and `scripts/audio_align.py`).

It sits in Media & Creative. It works with FFmpeg. The repository describes itself as: An open-source AI agent for social media — discover trends, create content, publish everywhere, and learn what works across Xiaohongshu, Douyin, Zhihu, Bilibili, and more.🎨一个开源的… The licence is Apache-2.0.

When your agent uses it

  • Media & Creative work in your project

Example prompts

  • “提取精彩片段”
  • “转竖版短视频”
  • “/clipify”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Find the funniest parts
  2. Trim each chosen clip
  3. Decide the output format
  4. If 16:9 → 9:16: pan-between-faces vs split-screen
  5. Add subtitles
  6. Deliver

What it can do on your machine

Read from SKILL.md and the folder at commit ede33b8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • ffmpeg
    • whisper
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Clipify loads about 2.2k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 888 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ZJU-REAL/Easel at commit ede33b8, republished under its Apache-2.0 licence (© ZJU-REAL). 888 words, ~2,210 tokens.

Download SKILL.mdSave it as .claude/skills/clipify/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
clipify
description
从长视频中自动提取精彩片段,切成独立短视频,支持 16:9→9:16 竖版转制和逐字字幕烧录。 当用户说"视频切片""提取精彩片段""长视频切短""切成短视频""高光剪辑""逐字字幕""转竖版短视频"时使用。 和 video-highlights 的区别:clipify 专做英文口播找笑点+动态人脸 pan;video-highlights 更通用(中文/直播皆可),静态转竖版更稳。
layer
produce

Clipify

Find the funniest moments in a video, cut them as standalone clips, optionally reformat 16:9 → 9:16 (face-pan or split-screen), and burn opus-style word-by-word captions.

Inputs

  • A video file path (the user will provide it; otherwise ask)
  • Optional: requested format (9:16, 16:9, 1:1) — if not given, ask after candidates are picked
  • Optional: subtitle style preference — if not given, ask before captioning

Tooling (use only the fastest path)

  • Whisper: whisper --model tiny.en --word_timestamps True --output_format json (≈10× faster than small.en; quality fine for English). For non-English: --model base (drop --language).
  • ffmpeg: hardware decode is optional and platform-specific — use -hwaccel auto, or omit it (macOS: videotoolbox; Linux: vaapi/cuda/none). Add -preset ultrafast for renders. Use -c:v libx264 -crf 20 for the final master.
  • Numpy for audio alignment (FFT cross-correlation). No scipy/cv2 needed.
  • Scripts: <skill-dir>/scripts/ (where <skill-dir> is the directory containing this SKILL.md — typically ~/.claude/skills/clipify/)
    • analyze.py — speaker timeline from two ROI motion files
    • build_pan.py — ffmpeg crop x-expression with hard cuts
    • build_ass.py — opus-style ASS captions from whisper JSON
    • audio_align.py — find offset of a sub-clip in a longer source

Working dir: /tmp/clipify/ (mkdir at start, leave artifacts for debugging).


Workflow

Step 1 — Find the funniest parts
bash
mkdir -p /tmp/clipify
ffmpeg -y -i "$VIDEO" -vn -ac 1 -ar 16000 /tmp/clipify/audio.wav
whisper /tmp/clipify/audio.wav --model tiny.en --word_timestamps True --output_format json --output_dir /tmp/clipify --language en

Read the resulting JSON (or .txt) and pick 3–5 candidate clips. Funny signals to scan for:

  • Punchlines and reactions: words like "what", "wait", "no way", laughter, "haha", swearing
  • Reversal moments: setup question → unexpected answer
  • Awkward pauses: Whisper segment with long gap, or filler ("uh", "um")
  • Self-roast / quotable one-liners: short declarative sentences that stand alone
  • Audio peaks: detect via ffmpeg -af volumedetect or look for rapid back-and-forth (alternating short Whisper segments)

For each candidate, propose: [start, end, why-it's-funny, suggested title]. Aim for 10–25s clips. Show the list and let the user confirm/pick.

Step 2 — Trim each chosen clip
bash
ffmpeg -y -ss "$START" -t "$DURATION" -i "$VIDEO" -c copy /tmp/clipify/clip_$N.mp4

(Use -c copy for instant trim. Re-encode only if cuts must be frame-accurate.)

Step 3 — Decide the output format

Ask the user (skip if they already specified): "9:16 (TikTok / Reels), 16:9 (YouTube), or 1:1 (Insta feed)?"

Step 4 — If 16:9 → 9:16: pan-between-faces vs split-screen

Detect source aspect with ffprobe. If source is 16:9 and target is 9:16, ask:

"Two options: (a) hard-cut pan that follows whoever is speaking (single face on screen at a time), or (b) split-screen stack with both faces visible. Which do you want?"

Skip the question if there's only one face (single-talker clip). For single-talker, just center-crop.

  1. Locate the two face ROIs. Sample one frame: ffmpeg -ss <middle> -i <clip> -frames:v 1 /tmp/clipify/probe.jpg. Read it. Eyeball each face's mouth+chin area as x,y,w,h in the source's pixel space. (No cv2 needed — camera is static within a clip; one frame is enough.) Verify by drawing boxes:

    bash
    ffmpeg -i probe.jpg -vf "drawbox=x=$LX:y=$LY:w=$LW:h=$LH:color=cyan@0.9:t=4,drawbox=x=$RX:y=$RY:w=$RW:h=$RH:color=magenta@0.9:t=4" verify.jpg

    Iterate at most twice. Boxes should cover mouth + chin and avoid hands/mics. Don't over-tune — frame differencing is forgiving.

  2. Extract per-frame motion energy in each ROI:

    bash
    ffmpeg -y -i clip.mp4 -filter_complex "
    [0:v]split=2[a][b];
    [a]crop=$LW:$LH:$LX:$LY,format=gray,tblend=all_mode=difference,signalstats,metadata=mode=print:key=lavfi.signalstats.YAVG:file=/tmp/clipify/L.txt[la];
    [b]crop=$RW:$RH:$RX:$RY,format=gray,tblend=all_mode=difference,signalstats,metadata=mode=print:key=lavfi.signalstats.YAVG:file=/tmp/clipify/R.txt[ra]
    " -map "[la]" -f null - -map "[ra]" -f null -
  3. Build speaker timeline (min dwell 1.0s — short interjections merge into the prior speaker):

    bash
    python3 <skill-dir>/scripts/analyze.py /tmp/clipify/L.txt /tmp/clipify/R.txt 1.0 > /tmp/clipify/segments.json
  4. Pick pan x-coordinates for a 9:16 vertical strip from the source. With source W=1920 and target W=1080, crop strip width = 608.

    • LEFT_X = face_left_center_x - 304 (clamp ≥ 0)
    • RIGHT_X = face_right_center_x - 304 (clamp ≤ source_W - 608)
  5. Generate the hard-cut x expression and render:

    bash
    EXPR=$(python3 <skill-dir>/scripts/build_pan.py /tmp/clipify/segments.json $LEFT_X $RIGHT_X)
    ffmpeg -y -i clip.mp4 -filter_complex \
      "[0:v]crop=608:1080:x='$EXPR':y=0,scale=1080:1920:flags=lanczos[v]" \
      -map "[v]" -map 0:a -c:v libx264 -preset fast -crf 20 -pix_fmt yuv420p \
      -c:a aac -b:a 192k /tmp/clipify/clip_panned.mp4

    Source 1920×1080 assumed; for 4K source either downscale first or double all coordinates.

Show full SKILL.md (350 more words)Show less
Step 4b — Split-screen (both faces always visible)

Two stacked tiles, 1080×960 each. The active speaker's tile is on top — overlay flips at speaker changes.

[0:v]split=2[a0][a1];
[a0]crop=Wcrop:Hcrop:LX_tile:LY_tile,scale=1080:960,split=2[lt0][lt1];
[a1]crop=Wcrop:Hcrop:RX_tile:RY_tile,scale=1080:960,split=2[rt0][rt1];
[lt0][rt0]vstack[layoutL];
[rt1][lt1]vstack[layoutR];
[layoutL][layoutR]overlay=0:0:enable='<RIGHT_SPEAKER_ENABLE>'[v]

Build <RIGHT_SPEAKER_ENABLE> from segments.json as between(t,a,b)+between(t,a,b)+... over the right-speaker segments. Tile crops should target ~720×640 around each face (1.125:1 to match 1080×960).

Step 5 — Add subtitles

Ask once (only if user hasn't already specified a style):

"Three subtitle styles: opus (big bold white, yellow active-word highlight), karaoke (4-word chunks, green highlight), minimal (clean Helvetica, no highlight). Or paste an example you like."

If they paste a reference image/example: match the font, size, weight, color, position, and animation as closely as possible — write a custom ASS by hand or extend build_ass.py.

Else use the preset:

bash
# Re-run whisper on the trimmed clip for accurate timestamps relative to clip start
whisper /tmp/clipify/clip_panned.mp4 --model tiny.en --word_timestamps True --output_format json --output_dir /tmp/clipify --language en
python3 <skill-dir>/scripts/build_ass.py /tmp/clipify/clip_panned.json /tmp/clipify/captions.ass opus

Burn captions:

bash
ffmpeg -y -i /tmp/clipify/clip_panned.mp4 -vf "subtitles=/tmp/clipify/captions.ass" \
  -c:v libx264 -preset fast -crf 20 -c:a copy "$OUTPUT.mp4"
Step 6 — Deliver
  • Save each output to <source_dir>/clipify_out/ (mkdir if missing)
  • Print one line per clip: name, duration, what was funny, output path
  • Print the first output path (or open it — Linux xdg-open <path>, macOS open <path>) so the user can check it
  • Offer to iterate (different style, different ROI, swap to split-screen, retime captions)

Pitfalls (lessons from prior runs — don't repeat)

  • Don't over-tune ROIs. Two iterations max. Motion-diff is forgiving — wider ROIs covering mouth+chin work fine even if not perfectly mouth-centered.
  • Watch out for scene cuts inside a clip. Run ffmpeg -filter:v "select='gt(scene,0.3)',showinfo" -f null - to count cuts. If a 16:9→9:16 clip has many cuts, the fixed face ROIs only work for the dominant scene; warn the user, and offer to either pick a single-take clip or accept off-center framing during cuts.
  • Source resolution matters. If source is 4K, either downscale to 1920×1080 first (faster, fine for 9:16 output) or multiply all ROI/pan coordinates by 2.
  • Burned-in subtitles in source. Some "raw" clips still have subtitles. If so, find the no-subs master via audio cross-correlation (audio_align.py) and trim from there.
  • Don't run whisper on the full feature-length source if a short clip suffices. Whisper the trimmed clip after Step 2; only whisper the full source in Step 1 if you need a transcript to find funny moments.
  • State the plan in one line, then act. Don't narrate every iteration.

© ZJU-REAL, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, assets) in skills/openclaw/clipify of ZJU-REAL/Easel.

  • SKILL.md
  • EASEL-META.md
  • assets/preview.png
  • scripts/analyze.py
  • scripts/audio_align.py
  • scripts/build_ass.py
  • scripts/build_pan.py

Open the folder on GitHubat commit ede33b8

Compare with similar skills

Clipify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clipify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clipify this skillZJU-REAL/Easel3.4k—~2.2kAutomated safety check: PassApache-2.0
Podcastzarazhangrui/personalized-podcast438—~2.3kAutomated safety check: NotesNone
Book Sales VideoKianzzz/book-sales-video218—~3.2kAutomated safety check: PassMIT
Frames CLIviticci/frames-cli405—~6.1kAutomated safety check: PassMIT
Video Assemblezenstory-ai/video-recap-skills561—~1.7kAutomated safety check: PassMIT
Extract Video Framesqdhenry/Claude-Command-Suite1.3k—~1.6kAutomated safety check: PassNone

Similar skills

  • Podcast

    zarazhangrui/personalized-podcast

    Generate a podcast episode from content you provide. An agent skill from zarazhangrui/personalized-podcast.

    438 GitHub stars~2.3k tokensUpdated 6 mo ago
    Media & CreativeAuto-check: notes
  • Book Sales Video

    Kianzzz/book-sales-video

    从书名或飞书多维表格中的成稿文案出发,结合微信读书资料与公开点评创作图书带货/书评短视频,并用豆包 TTS、Pexels、Codex 生图和本机 OpenChatCut 完成配音、配图、双语字幕、音效、动效、BGM、可编辑初稿与按需导出。用户提出“根据一本书做带货视频”“读取飞书文案制作图书视频”“写书评口播并自动剪成抖音视频”“仿参考样式做图书推荐短视频”时使用;仅查书、仅写普通书评或无关剪辑…

    218 GitHub stars~3.2k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Frames CLI

    viticci/frames-cli

    Frame screenshots and screen recordings with the frames CLI.

    405 GitHub stars~6.1k tokensUpdated 11 days ago
    Media & CreativeAuto-check passed
  • Video Assemble

    zenstory-ai/video-recap-skills

    合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

    561 GitHub stars~1.7k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Extract Video Frames

    qdhenry/Claude-Command-Suite

    Extracts frames and timestamped audio segments from video files (GIF, MP4, MOV) at configurable intervals and stores them in a directory with a manifest file.

    1.3k GitHub stars~1.6k tokensUpdated 7 mo ago
    Media & CreativeAuto-check passed
  • Video Cut

    zenstory-ai/video-recap-skills

    把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clipplan.json 与源视频, 输出 editedsource.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。

    561 GitHub stars~1.6k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed

More from ZJU-REAL/Easel

All 114 skills in this repo
  • Gzh Design

    ZJU-REAL/Easel

    微信公众号文章排版引擎:把 Markdown / Word(.docx) / PDF / 纯文本转成可直接粘贴进公众号编辑器的 HTML,自动章节编号、关键词标记、引言卡、目录、代码块、图片/GIF、作者签名;主题从 references/theme-index.md…

    3.4k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • 微信公众号文章自动创作与发布工具。给定参考文章、文字或文档,自动搜索整理全网相关信息、生成图文并茂的公众号文章,并发布到微信公众号草稿箱。特别强调反 AI 检测写作。

    3.4k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Card Design

    ZJU-REAL/Easel

    社媒卡片视觉设计系统:提供配色、中文字体层级、满画幅布局、品类骨架和死空白/密度质检,避免模板化 PPT 与廉价 AI 感。

    3.4k GitHub stars~657 tokensUpdated today
    Auto-check passed
  • Ecom Details Image

    ZJU-REAL/Easel

    生成电商商品视觉方案:主图概念、场景图、详情页视觉方向和 AI 生图 Prompt. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1.1k tokensUpdated today
    Auto-check: notes
  • Infographic

    ZJU-REAL/Easel

    将数据或文字内容转化为可视化信息图,支持静态(AntV)和动画 GIF 两种模式。当用户需要制作信息图、数据可视化、流程图、对比图、动画图表、GIF 图表、思维导图、SWOT 分析图时调用。本地渲染信息图/GIF 动画;要单张静态图片 URL 用 chart-visualization,要 CSV/JSON→整页报告用 data-report

    3.4k GitHub stars~643 tokensUpdated today
    Auto-check passed
  • Novel Writer

    ZJU-REAL/Easel

    长篇小说/网文连载创作:从世界观、人设和三级大纲写到逐章正文,并用文件化状态维护伏笔、前情和跨章一致性. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1k tokensUpdated today
    Auto-check passed

Works with

Questions about Clipify

What does Clipify do?

从长视频中自动提取精彩片段,切成独立短视频,支持 16:9→9:16 竖版转制和逐字字幕烧录. An agent skill from ZJU-REAL/Easel. Clipify is an agent skill from ZJU-REAL/Easel.

When should I use Clipify?

Clipify fits situations like: media & Creative work in your project.

How do I install Clipify in Claude Code?

Run `npx skills add ZJU-REAL/Easel --skill clipify -a claude-code`. Or copy the skill folder (skills/openclaw/clipify in ZJU-REAL/Easel) into .claude/skills/clipify in your project. Claude Code loads it when a task matches its description.

How do I install Clipify in Codex?

Run `npx skills add ZJU-REAL/Easel --skill clipify -a codex`. Or copy the skill folder (skills/openclaw/clipify in ZJU-REAL/Easel) into .agents/skills/clipify in your project. Codex loads it when a task matches its description.

Can I use Clipify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZJU-REAL/Easel --skill clipify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/clipify, .gemini/skills/clipify, .github/skills/clipify and .opencode/skills/clipify in your project.

What does Clipify need to run?

Going by SKILL.md and its folder, Clipify needs Python for the scripts in its folder and the command-line tools its instructions call (ffmpeg, whisper and python3). Our summary lists: Python 3.

Does Clipify access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Clipify safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Clipify use?

Clipify is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clipify use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Clipify?

Skills that share tags, products or a category with Clipify: Podcast (zarazhangrui/personalized-podcast, 438 stars), Book Sales Video (Kianzzz/book-sales-video, 218 stars), Frames CLI (viticci/frames-cli, 405 stars) and Video Assemble (zenstory-ai/video-recap-skills, 561 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clipify?

ZJU-REAL (a GitHub organization) maintains it in ZJU-REAL/Easel, which has 3,441 GitHub stars. The repository holds 114 skills in this directory. The repository was last updated on October 11, 2026.

Source: ZJU-REAL/Easel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.