Triage
TalAter/annyang
Triage and close GitHub issues on TalAter/annyang. An agent skill from TalAter/annyang.
Tighten a long recording aggressively — remove silences, fillers, hedges, false starts, and repetitions while preserving laughs and comedic pauses.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add louisedesadeleer/cut-video --skill cut-video -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install louisedesadeleer/cut-video cut-video --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cut-video" agent skill from https://github.com/louisedesadeleer/cut-video/tree/main into .claude/skills/cut-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cut-video", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add louisedesadeleer/cut-video --skill cut-video -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install louisedesadeleer/cut-video cut-video --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cut-video" agent skill from https://github.com/louisedesadeleer/cut-video/tree/main into .agents/skills/cut-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cut-video", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add louisedesadeleer/cut-video --skill cut-video -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install louisedesadeleer/cut-video cut-video --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cut-video" agent skill from https://github.com/louisedesadeleer/cut-video/tree/main into .cursor/skills/cut-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cut-video", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add louisedesadeleer/cut-video --skill cut-video -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install louisedesadeleer/cut-video cut-video --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cut-video" agent skill from https://github.com/louisedesadeleer/cut-video/tree/main into .gemini/skills/cut-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cut-video", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install louisedesadeleer/cut-video cut-videoInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add louisedesadeleer/cut-video --skill cut-video -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cut-video" agent skill from https://github.com/louisedesadeleer/cut-video/tree/main into .github/skills/cut-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cut-video", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add louisedesadeleer/cut-video --skill cut-video -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install louisedesadeleer/cut-video cut-video --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cut-video" agent skill from https://github.com/louisedesadeleer/cut-video/tree/main into .opencode/skills/cut-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cut-video", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cut-videoTighten a long recording aggressively — remove silences, fillers, hedges, false starts, and repetitions while preserving laughs and comedic pauses.
Cut Video is an agent skill from louisedesadeleer/cut-video. Tighten a long recording aggressively — remove silences, fillers, hedges, false starts, and repetitions while preserving laughs and comedic pauses. Cuts are driven by Montreal Forced Aligner (MFA) word boundaries (auto-installs on first run), not raw Whisper timestamps. Use when the user pastes a video and says "cut this", "tighten this video", "remove silences", "strip ums", "clean this up", or any variant of "make this video shorter without losing the good parts".
Its SKILL.md is about 8.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `README.md` and `make_review.py`).
It sits in AI & LLM Engineering, covering Speech recognition and synthesis. The repository describes itself as: Claude Code skill: tighten long recordings — remove silences, ums, dead air. Preserves laughs and comedic pauses. Fast on Apple Silicon. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a5a87a6. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
ffmpegcondabrewpython3whisperkindffprobeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cut Video loads about 8.1k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 4,057 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
**Do NOT wait for approval — print the summary and start the render in the same turn** (changed 2026-07-09: the approvalAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from louisedesadeleer/cut-video at commit a5a87a6, republished under its MIT licence (© louisedesadeleer). 4,057 words, ~8,096 tokens.
.claude/skills/cut-video/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Tighten a long-form recording aggressively: remove silences, fillers, hedges, weak transitions, false starts, and repetitions. Preserves laughs and comedic pauses. Outputs a cleaned MP4 ready to drop into CapCut for layouts/zooms/memes.
This skill is more accurate than cutting from Whisper alone. Cuts are driven by Montreal Forced Aligner (MFA) word boundaries (~10–20ms precision, with true inter-word silences as explicit intervals), not Whisper timestamps (±100–300ms, with pauses embedded inside word durations). Whisper is used only to produce the transcript text that MFA aligns to the audio. The result: tighter cuts, no clipped word onsets/tails, and reliable silence detection.
⚠️ But MFA times are a DRAFT, not ground truth (burned 2026-07-02,
AI is bad at jokesDJI run). MFA can only align the transcript whisper gave it. On retake-heavy footage whisper COLLAPSES repeated lines, so MFA smears one transcribed instance across several spoken takes — measured drift on that run was 1–6 seconds in the back half ("the average of all working code": MFA said 217.2s, real onset 225.4s; the final "what this means" take: MFA said 114.7s, real 130.4s). Every keep boundary must be ground-truthed with isolated-window re-transcription (Step 2.5) before rendering. Skipping that step on that video would have produced garbage cuts for the entire second half.
Style target (calibrated from AI mogging my dad.mp4):
aggressive / balanced / sentimental / documentary) — default: aggressive/tmp/cut-video/<basename>/ — mkdir at start, leave artifacts for debugging.
If the source is HEVC, >500MB, or 4K+, transcode to a 1080p H.264 working copy FIRST. Every subsequent step runs against the proxy, not the source.
Check orientation first (ffprobe … width,height) — DJI/phone footage is often vertical 9:16, and a hard-coded scale=1920:1080 would squash it. Use scale=1920:1080 only for landscape; use scale=1080:1920 for portrait (or the orientation-proof scale=-2:1080 / scale=1080:-2).
ffmpeg -y -hwaccel videotoolbox -i "$SRC" \
-vf scale=1920:1080 \ # portrait sources: scale=1080:1920
-c:v libx264 -preset fast -crf 20 -pix_fmt yuv420p \
-c:a aac -b:a 192k \
/tmp/cut-video/$NAME/proxy.mp4Critical flags:
-hwaccel videotoolbox on the INPUT — hardware-decodes HEVC, ~5–10× faster on Apple Silicon. Skip this and you'll wait minutes instead of seconds.-preset fast — the libx264 default is medium, which is ~3× slower for no useful quality gain on a working copy.Hardware encode alternative (even faster on M-series, slightly larger file):
-c:v h264_videotoolbox -b:v 8MMFA is a forced aligner, not a transcriber — it needs a transcript to align. Whisper's only job here is to produce that text; whisper word timestamps are NOT used for cutting (they're ±100–300ms off and turbo embeds pauses inside word durations). MFA (Step 1.5) supplies all timing.
ffmpeg -y -hwaccel videotoolbox -i /tmp/cut-video/$NAME/proxy.mp4 \
-vn -ac 1 -ar 16000 /tmp/cut-video/$NAME/audio.wav
whisper /tmp/cut-video/$NAME/audio.wav \
--model tiny.en --word_timestamps True --output_format json \
--output_dir /tmp/cut-video/$NAME --language enModel choice (revised 2026-07-02 — "we only need the words" was WRONG): the transcript text IS the alignment input, so transcript errors become timing errors. On the AI is bad at jokes run, tiny.en dropped exactly the words the old warning predicted ("bland", "slop"), misheard "Compare this to" as "comparative is", and collapsed full-line retakes — and MFA, aligning that wrong text, drifted 1–6s across the back half. Use tiny.en only for a first fast pass to see the take structure; if the transcript shows retakes/repetitions, or the footage is noisy/outdoor, the per-region ground-truth pass (Step 2.5) with small.en supplies the real cut times anyway. (A full-pass small.en transcript is a reasonable upgrade for the MFA input too — ~1–2 min on a 5-min video — but it ALSO collapses retakes, so it does not remove the need for Step 2.5.) For non-English use --model base and drop --language. We keep --word_timestamps True only so whisper times survive as a cross-check/fallback (Step 1.5 caveats, 2a-legacy) — never as the primary cut source.
This is the spine of the skill. MFA forced-aligns the whisper transcript text to the audio and returns word boundaries at ~10–20ms precision, with true inter-word silences as explicit empty intervals. All cut decisions (gaps, fillers, false starts, retakes) use MFA word times.
Auto-install MFA (the skill does this itself — don't make the user do it). Run this idempotent bootstrap at the start of every run; it's a no-op once everything's present (a few seconds), and a one-time ~2–3 min install on a fresh machine. Tell the user "installing MFA (one-time)…" only when it actually installs something.
# 1. Ensure conda exists (MFA's only supported install path). Install miniforge via brew if missing.
if ! command -v conda >/dev/null 2>&1; then
echo "conda not found — installing miniforge (one-time)…"
brew install --cask miniforge || brew install miniforge
# make conda available in this shell
eval "$("$(brew --prefix)/bin/conda" shell.bash hook 2>/dev/null || conda shell.bash hook)"
fi
source "$(conda info --base)/etc/profile.d/conda.sh"
# 2. Ensure the mfa env exists with MFA installed
if ! conda env list | grep -q '^mfa\b'; then
echo "creating mfa conda env (one-time)…"
conda create -n mfa -c conda-forge montreal-forced-aligner -y
fi
# 3. Ensure the English acoustic model + dictionary are downloaded (idempotent; skips if present)
conda run -n mfa mfa model download acoustic english_mfa 2>/dev/null || true
conda run -n mfa mfa model download dictionary english_mfa 2>/dev/null || trueIf brew itself is missing, MFA can't be auto-installed — fall back to the whisper+silencedetect path (see the fallback bullet below) and tell the user, with the one-line brew install miniforge they'd need to unlock MFA precision.
Per run:
# corpus = a dir with audio.wav + audio.txt (the whisper transcript as PLAIN TEXT, no timestamps)
mkdir -p /tmp/cut-video/$NAME/corpus
cp /tmp/cut-video/$NAME/audio.wav /tmp/cut-video/$NAME/corpus/
python3 -c "import json;d=json.load(open('/tmp/cut-video/$NAME/audio.json'));open('/tmp/cut-video/$NAME/corpus/audio.txt','w').write(d['text'])"
conda run -n mfa mfa align --clean /tmp/cut-video/$NAME/corpus english_mfa english_mfa /tmp/cut-video/$NAME/aligned
# → aligned/audio.TextGrid — "words" tier: (start, end, word); EMPTY-label intervals = true silencesParse the TextGrid words tier (pip praatio, or a 20-line parser — interval tiers are plain text). What MFA buys:
silencedetect pass (2a) as a CROSS-CHECK, and treat disagreements > 0.3s as suspect regions to re-inspect. On retake-heavy self-shot footage the drift is not local — it's cumulative and can reach seconds (see the warning at the top and Step 2.5). And when the noise floor kills silencedetect too (low-dynamic-range guard), you have NO automatic cross-check — Step 2.5 is then mandatory, not optional.Compute a list of (start, end) intervals to KEEP from the MFA word intervals (whisper JSON only as fallback). Aggressive cutting removes content in four categories:
With MFA (Step 1.5), true silences are the TextGrid's empty-label intervals — apply the gap table below to those directly. Cross-check with audio silence detection — but CALIBRATE the threshold to the recording's own noise floor; never hard-code -30dB (added 2026-06-16, burned on a DJI-mic to-camera take whose floor sat at ~-24dB: a fixed -30dB saw the ENTIRE take as non-silent, so every "silent" gap measured "above threshold = content, keep" and ~8s dead-air pauses survived the cut):
# 1. Measure the noise floor F and speech level S from per-0.5s-window RMS
# (slice the wav into 0.5s windows, run volumedetect on each, collect mean_volume;
# F ≈ 10th-percentile window RMS (room tone), S ≈ 90th-percentile (speech))
# 2. Set the silence threshold RELATIVE to the floor, not absolute:
THRESH=$(python3 -c "print(f'{F + 6:.0f}')") # 6 dB above the measured floor
ffmpeg -i proxy.mp4 -af silencedetect=noise=${THRESH}dB:d=0.14 -f null - 2>&1 | grep silence_⚠️ Low-dynamic-range guard (the real lesson, 2026-06-16): if S − F < ~8dB (noisy mic — outdoor/DJI/lav with AC hum, where the per-second RMS reads the SAME during speech and silence), NO energy threshold can separate speech from silence. Do not trust silencedetect/volumedetect at all in that case — drive every cut from MFA word boundaries, or in the no-MFA fallback from small.en word ONSETS + stretched-word pause detection (a whisper word whose duration ≫ a normal word IS a hidden pause: keep ~0.10 + 0.105·len(word)s of onset, then cut to the next word's start). Say which path you used in the plan summary.
Without MFA, the silencedetect pass is the trustworthy gap source ONLY when the dynamic-range guard passes: mlx_whisper/whisper-turbo (and tiny.en) EMBED pauses inside word durations — a 3s "word" is really "word [long pause]", inter-word gaps read ~0.00, and gap-based trimming does NOTHING. Use small.en (not tiny.en) for the fallback path — tiny stretches words across silence even worse and silently drops words ("bland", "slop") and whole retakes. Parse silence_start/silence_end, remove those intervals (keep ~0.04s pad; MFA-precision boundaries allow ~0.02s). For "almost no pauses" → d=0.12, pad 0.03. Combine with the dedup (2c/2d) by removing silences ∪ dropped-word-ranges. This is how you hit median ~1.1–1.3s.
| Gap | aggressive (default) | balanced | sentimental | documentary |
|---|---|---|---|---|
< 0.25s | keep | keep | keep | keep |
0.25–1s | trim to 0.1s | trim to 0.3s | trim to 0.5s | keep |
1–3s | trim to 0.2s | trim to 0.8s | trim to 1.5s | trim to 2s |
> 3s | trim to 0.2s | trim to 0.5s | trim to 1s | trim to 3s |
The aggressive thresholds match the reference video's pacing (median cut 1.13s).
Always cut: um, uh, umm, uhh, er, erm, ah, mhm, hmm (standalone)
Cut these when they're delivered as filler (not as substantive content). Mark for cut, then verify against context in the review step:
like (as filler, not as comparison), you know, I mean, I guess, I think (when hedging, not asserting), kind of, sort of, basically, literally, actually (as filler)so (sentence-opener), and so, and then, but um, but uh, okay so, right so, wellkind of like, sort of like, or whatever, or somethinganyway, so anyway, the thing is, what I'm saying is, to be honest, honestlyIn balanced / sentimental / documentary modes, only cut hedges that precede a re-statement (false start, see 2d). Keep them in conversational moments.
Detect by scanning for:
"I want — I wanted to...", "It's a — it's an interesting...". Drop the earlier attempt + the gap."so basically what I'm saying is...", "the point is...", "let me start over...". Drop the preamble, keep the actual point.Implementation hint: build a sliding-window fuzzy-match on consecutive word n-grams. When 3-gram similarity > 0.8 within a 3-second window, flag the earlier instance for removal.
Self-shot intros/promos contain whole-sentence retakes separated by 5–60s ("The next person is Carmen who works. … The next person, the next guest I invited is Carmen who's social and content lead at Slate") — far outside 2d's 3-second window. Detect and resolve them:
line → takes at [t1, t2, t3] → keeping t3) so the user can override which take survives.ffmpeg -i audio.wav -af "volumedetect" -f null - 2>&1 | grep mean_volumeF + ~10dB it contains audible content (laugh, breath, reaction) — keep it. A hard-coded -30dB/-25dB cutoff fails on noisy mics whose floor already sits above it — it then "protects" pure room tone as if it were a laugh, which is exactly how dead air survives. Use ffmpeg -i audio.wav -af "astats=metadata=1:reset=1,ametadata=print:key=lavfi.astats.Overall.RMS_level" for per-window RMS. (And remember the low-dynamic-range guard in 2a: when S − F < ~8dB, RMS can't tell a laugh from room tone either — fall back to MFA/word-onset timing and surface anything ambiguous in the plan for the user to judge.)aggressive mode, still trim to 0.4–0.6s (don't kill it, just tighten).After building the keep list, run a machine scan — do NOT trust your eyes on the transcript:
MFA gives you the take STRUCTURE (what was said, roughly where). It does not reliably give you cut times on retake-heavy or noisy footage. Before rendering, re-transcribe every region you plan to keep in an isolated ±few-second window with small.en — within a short window whisper's word times are accurate, and it hears words the full pass dropped ("bland", "slop") and retakes the full pass collapsed ("It turns out— it turns out…").
# windows: name:start:duration — start ≥0.5s BEFORE the expected onset
for n in "r1:5.5:8.4" "r2:17.0:6.0" "r3:30.5:6.0"; do
f=${n%%:*}; rest=${n#*:}; o=${rest%%:*}; t=${rest##*:}
ffmpeg -y -hide_banner -loglevel error -ss $o -t $t -i audio.wav gt/$f.wav
whisper gt/$f.wav --model small.en --word_timestamps True --output_format json --output_dir gt --language en >/dev/null 2>&1
python3 -c "
import json
d=json.load(open('gt/$f.json'))
print('=== $f (+$o) ===')
for s in d['segments']:
for w in s.get('words',[]):
print(f\"{$o+w['start']:7.2f}-{$o+w['end']:7.2f} {w['word']}\")"
done(That loop is zsh-safe — don't use set -- $var, zsh doesn't word-split unquoted variables.)
Rules for reading the windows:
small.en collapses within-window repeats just like the full pass ("to" spanning 93.1–97.9 contained an entire abandoned take). Re-window tighter around any word ≳1s. Sometimes what's inside is a complete clean take that neither full-pass ASR surfaced — check before assembling a splice from fragments.[gap_start+0.05, gap_end−0.05]). Leave deliberate rhetorical beats ("median… or average") no shorter than ~0.15s.slop 285.08–286.32). Trimming what looks like the pause half CLIPS THE WORD mid-vowel. Before trimming into any stretched word at a keep boundary, re-window it tightly to find where the voice actually stops — and if the post-render verify transcript is missing a word that was there before, your trim ate it: push the boundary back out.Cost on a 5-min video: ~10–15 windows × a few seconds of small.en each ≈ 2 minutes. It is the difference between a clean cut and re-doing the whole back half.
Do NOT wait for approval — print the summary and start the render in the same turn (changed 2026-07-09: the approval gate cost Louise a wasted wait when she didn't notice the question; the render is non-destructive and cheap to redo, so render-first is strictly better). Print this summary, then go straight to Step 4:
Original duration: 7m 25s
Cleaned duration: 3m 58s
Cut: 3m 27s (47%)
Total keep intervals: 312
Median cut duration: 1.2s (target: 1.1–1.5s)
Hard cuts (dead air > 3s removed): 23
Filler hits: 47
Hedge/transition cuts: 89
False-start drops: 12
Notable preserved long pauses: list timestamps + context (laughs, punchlines)Self-check vs benchmark:
Run the self-check yourself and fix violations before rendering — that's the quality gate now, not the user. Then render immediately. The user reviews the result; if a cut killed a laugh or an intentional pause, adjust the keep list and re-render (fast — trim/concat, not re-transcription).
(Demoted from always-on 2026-07-02: Louise found the editor fiddly in practice and prefers the pipeline to just cut tighter automatically — the intra-take pause trimming in Step 2.5 came out of that. Generate the editor only if she asks to review/adjust by hand.)
A timeline editor so she can trim silences and drop anything she doesn't want, before (or after) the render:
# keeps.json must be [start, end, "transcript text"] triplets in the working dir;
# needs proxy.mp4 (and ideally audio.wav) in the same dir. Generates wave.png once.
python3 ~/.claude/skills/cut-video/make_review.py /tmp/cut-video/$NAME
open /tmp/cut-video/$NAME/review.htmlWhat the page gives her (self-contained HTML next to proxy.mp4, no server needed — waveform peaks and silence suggestions are embedded in the file, so it works on file://):
S), drop/restore a block (D/⌫ or click its card), double-click a gap to resurrect cut footage, undo (Z, 60 levels).{"keeps": [[a,b], ...]} — the full edited keep list in source-proxy seconds.Applying her decisions: the pasted keeps array REPLACES the old keep list — carry transcript text over by time-overlap with the previous keeps.json (blocks she created from gaps have no text; label them "(restored)"), rebuild filter.txt, re-render. The proxy is already there, so a revision costs seconds.
Boundary hygiene: her hand-dragged edges are intentional — do NOT re-snap them to MFA/ASR word boundaries. Only warn if an edge lands mid-word per the ground-truth words (say which word and offer the nearest clean boundary).
trim + concat, NOT selectCritical: Use the trim+concat filtergraph pattern, NOT select/aselect. The select filter compresses video frames but does NOT properly compress audio PTS, leaving audio and video misaligned.
Build a filtergraph with one trim+atrim pair per keep interval, then concat them:
[0:v]trim=A1:B1,setpts=PTS-STARTPTS[v0];
[0:a]atrim=A1:B1,asetpts=PTS-STARTPTS[a0];
[0:v]trim=A2:B2,setpts=PTS-STARTPTS[v1];
[0:a]atrim=A2:B2,asetpts=PTS-STARTPTS[a1];
...
[v0][a0][v1][a1]...concat=n=N:v=1:a=1[outv][outa]Write the filtergraph to a file (it'll get long — aggressive cuts mean 200–400 intervals on a 7-min source) and use -filter_complex_script. Render call:
ffmpeg -y -hwaccel videotoolbox -i /tmp/cut-video/$NAME/proxy.mp4 \
-filter_complex_script /tmp/cut-video/$NAME/filter.txt \
-map "[outv]" -map "[outa]" \
-c:v libx264 -preset fast -crf 20 -pix_fmt yuv420p \
-c:a aac -b:a 192k \
/tmp/cut-video/$NAME/cleaned.mp4Faster alternative for many cuts: if filtergraph exceeds ~500 trims, switch to per-segment extract + lossless concat:
ffmpeg -f concat -safe 0 -i concat.txt -c copy cleaned.mp4This avoids re-encoding entirely on the trim pass.
<source_dir>/cut_out/<source_name>_cut.mp4 (mkdir if missing)open the file so the user can review immediately-hwaccel videotoolbox on the proxy step. HEVC software decode on a 7-min 4K source can take 5+ minutes. With the flag, ~30s.select filter for cuts. It misaligns audio. Use trim+concat.-preset medium (the libx264 default). It's ~3× slower than -preset fast for no quality gain on a working copy.whisper the entire raw source if a proxy exists. Run whisper against audio.wav extracted from the proxy.-c:a copy to skip an unnecessary AAC pass.aggressive. The reference cut style is fast — if the output feels "safe", it's not matching Louise's CapCut pacing.h264_videotoolbox -b:v 8M is fine for the final render too, not just the proxy — the whole 4K-HEVC→proxy→18-segment render finished in ~1 min total on M-series.AI mogging my dad.mp4 (2026-05-03, CapCut project 0501):
Use this as ground-truth for "aggressive" mode tuning.
Keep this skill focused on one thing: produce a tighter MP4 from a long-form recording, fast and aggressive.
© louisedesadeleer, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in the repository root of louisedesadeleer/cut-video.
Open the folder on GitHubat commit a5a87a6
Cut Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cut Video this skilllouisedesadeleer/cut-video | 108 | — | ~8.1k | Automated safety check: Warn | MIT | |
| TriageTalAter/annyang | 6.8k | 1 repos | ~810 | Automated safety check: Notes | MIT | |
| Yichen Asrmcncarl/yichen-skills | 4.3k | — | ~780 | Automated safety check: Pass | Custom licence | |
| Dingtalk MinutesDingTalk-Real-AI/dingtalk-workspace-cli | 3.2k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Youtube FetcherJimmySadek/youtube-fetcher-to-markdown | 485 | — | ~3.1k | Automated safety check: Pass | MIT | |
| Yichen Web Researchmcncarl/yichen-skills | 4.3k | — | ~1.9k | Automated safety check: Pass | Custom licence |
TalAter/annyang
Triage and close GitHub issues on TalAter/annyang. An agent skill from TalAter/annyang.
mcncarl/yichen-skills
逸尘自用的统一音视频转写入口,在 StepFun Step ASR 与火山引擎豆包 ASR 之间按输出需求、安全边界和可用状态路由。用于本地音频或视频的纯文本转写、时间戳、SRT 字幕、口播粗剪,以及转写前体检;用户明确指定服务商时不得静默切换。Use when a local audio or video file needs transcription and the correct…
DingTalk-Real-AI/dingtalk-workspace-cli
钉钉 AI 听记。Use when 查询或修改听记摘要、完整逐字稿、关键词、标签、行动项、录音、上传、思维导图、发言人洞察、ASR 热词/识别词配置或分享权限。写文档走 dingtalk-doc;建待办走 dingtalk-todo;日程走 dingtalk-calendar。命令前缀:dws minutes。
JimmySadek/youtube-fetcher-to-markdown
Retrieve transcripts from YouTube, Instagram, TikTok, X, Vimeo and other video sites, summarize or analyze what was said (and shown on screen), or save an Obsidian-ready Markdown knowledge-base note…
mcncarl/yichen-skills
逸尘自用的互联网研究总入口。用于跨平台且跨阶段、用户尚未确定工具,或明确要求对公司、产品、人物、技术、行业和领域做横纵分析、发展史加现状对比或有来源约束的系统深度研究;先生成有截止日期和证据闸门的计划,再把搜索发现、候选核验、有限归档、按需转写和证据综合路由到…
ysyecust/lecture-to-notes
Transcribe local audio or video with Volcengine Doubao file ASR, including BigASR 1.0 Turbo direct upload and asynchronous 1.0 standard, 1.0 idle, or 2.0 standard jobs through TOS.
Categories
Tighten a long recording aggressively — remove silences, fillers, hedges, false starts, and repetitions while preserving laughs and comedic pauses. Cut Video is an agent skill from louisedesadeleer/cut-video. Tighten a long recording aggressively — remove silences, fillers, hedges, false starts, and repetitions while preserving laughs and comedic pauses.
Cut Video fits situations like: the user pastes a video and says cut this; tighten this video; remove silences; any variant of make this video shorter without losing the good parts.
Run `npx skills add louisedesadeleer/cut-video --skill cut-video -a claude-code`. Or copy the skill folder (the louisedesadeleer/cut-video repository) into .claude/skills/cut-video in your project. Claude Code loads it when a task matches its description.
Run `npx skills add louisedesadeleer/cut-video --skill cut-video -a codex`. Or copy the skill folder (the louisedesadeleer/cut-video repository) into .agents/skills/cut-video in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add louisedesadeleer/cut-video --skill cut-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cut-video, .gemini/skills/cut-video, .github/skills/cut-video and .opencode/skills/cut-video in your project.
Going by SKILL.md and its folder, Cut Video needs Python for the scripts in its folder and the command-line tools its instructions call (ffmpeg, conda, brew, python3, whisper and kind). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.
Cut Video is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.1k tokens (SKILL.md is roughly 32k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cut Video: Triage (TalAter/annyang, 6.8k stars), Yichen Asr (mcncarl/yichen-skills, 4.3k stars), Dingtalk Minutes (DingTalk-Real-AI/dingtalk-workspace-cli, 3.2k stars) and Youtube Fetcher (JimmySadek/youtube-fetcher-to-markdown, 485 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
louisedesadeleer (a GitHub user) maintains it in louisedesadeleer/cut-video, which has 108 GitHub stars. The repository was last updated on July 9, 2026.
Source: louisedesadeleer/cut-video on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.