Agent skill

Auto Cut

by notivn in notivn/AIEV

Cut silences and dead weight (fillers, repeated takes, false starts) out of a talking-head video BEFORE building the edit - call the measured /auto-trim API instead of hand-rolling ffmpeg, review…

MITAuto-check passedMedia & Creative

Install Auto Cut

skills CLI
$ npx skills add notivn/AIEV --skill auto-cut -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install notivn/AIEV auto-cut --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/notivn/AIEV.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/auto-cut .claude/skills/auto-cut && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
auto-cut
GitHub stars
126
Token cost
~3.2k tokens
SKILL.md length
1,818 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Cut silences and dead weight (fillers, repeated takes, false starts) out of a talking-head video BEFORE building the edit - call the measured /auto-trim API instead of hand-rolling ffmpeg, review…

  • Works in 7 steps: Prerequisite: a transcript with word… → Analyze (free, repeatable, no encoding) → REVIEW every candidate (your job, not… → …
  • Tasks that involve Video production
  • SKILL.md covers Principles, Step 0 - Prerequisite: a…, Step 1 - Analyze (free,… and Step 2 - REVIEW every…, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Auto Cut is an agent skill from notivn/AIEV. Cut silences and dead weight (fillers, repeated takes, false starts) out of a talking-head video BEFORE building the edit - call the measured /auto-trim API instead of hand-rolling ffmpeg, review the dead-weight candidates it returns, and do the one job only a human/AI can do (spotting repeated POINTS). Read this when the brief enables "Tự động cắt ngắn video" (autoCut) or the user complains that the video still has dead weight/silences.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Video production and Motion graphics. It works with FFmpeg. The repository describes itself as: Automatic AI video editing. Claude directs HyperFrames (HTML + GSAP motion graphics) and Remotion (timeline assembly) to turn raw footage into a finished MP4 - transcript… The licence is MIT.

When your agent uses it

  • Tasks that involve Video production
  • Tasks that involve Motion graphics

Example prompts

  • “Tự động cắt ngắn video”
  • “/auto-cut”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Prerequisite: a transcript with word timestamps
  2. Analyze (free, repeatable, no encoding)
  3. REVIEW every candidate (your job, not the machine's)
  4. REPEATED POINTS (only you can find these - mandatory)
  5. Apply (goes through the render queue)
  6. Read the report (the job is not done until you do)
  7. Use the cut version everywhere

What it can do on your machine

Read from SKILL.md and the folder at commit 1a4c2b0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Auto Cut loads about 3.2k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 1,818 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from notivn/AIEV at commit 1a4c2b0, republished under its MIT licence (© notivn). 1,818 words, ~3,175 tokens.

Download SKILL.mdSave it as .claude/skills/auto-cut/SKILL.md (or your agent's skills folder).
name
auto-cut
description
Cut silences and dead weight (fillers, repeated takes, false starts) out of a talking-head video BEFORE building the edit - call the measured /auto-trim API instead of hand-rolling ffmpeg, review the dead-weight candidates it returns, and do the one job only a human/AI can do (spotting repeated POINTS). Read this when the brief enables "Tự động cắt ngắn video" (autoCut) or the user complains that the video still has dead weight/silences.

Auto-Cut - measured silence & dead-weight trimming

Principles

  1. Cut BEFORE building scenes/captions. Cutting afterwards throws off every timestamp (captions, zooms, SFX). Output of this step: assets/<source>.cut.mp4 + assets/transcript.cut.json - every later step uses the cut version.
  2. Cutting is MANDATORY when the brief enables autoCut - it is not a suggestion. If nothing can be cut, you must state a concrete reason (the video was already tight) in the report.
  3. The thresholds live in the server, not in your head. Do NOT run silencedetect yourself and do NOT pick a dB threshold by feel. Two measured facts killed that workflow:
    • The right threshold is a property of the FILE, not a constant. On one real file, -40dB found 0 silences, -30dB found 13, -25dB found 21. A hardcoded number is either useless or eats speech.
    • Loudness alone cannot tell "pausing" from "speaking quietly". On the source file measured, silencedetect at -30dB reported 48.6s of "silence" but only 18.2s of it sat in a real gap between words; the other 30.4s was inside words (syllable breaks, unvoiced Vietnamese finals c/t/p/ch). Cutting on sound level alone swallows speech. The server does the measuring (autoTrim.ts) and the candidate generation (deadWeight.ts). You do the reviewing. That split is the whole point.
  4. Two kinds of dead weight get cut: silences (machine-measured, guarded by the transcript) and content dead weight (fillers, stutters, restated takes - proposed by the server, approved by you; plus repeated POINTS, which only you can find - see Step 3).

Step 0 - Prerequisite: a transcript with word timestamps

The whole guard depends on it. The server looks for assets/transcript.raw.json, then assets/transcript.json. Without one, analysis still runs but comes back guarded: false, the dead-weight list is empty, and the numbers are guesswork in both directions (it can pass a bad cut and fail a good one). Transcribe first - never trim a Vietnamese talking head unguarded.

Step 1 - Analyze (free, repeatable, no encoding)

POST http://localhost:6869/api/projects/<id>/auto-trim/analyze
body: {}                                   # or { "source": "assets/face.mp4", "level": "tight" }

level is natural | default | tight and defaults to brief.autoCutLevel. Response:

  • silence - durationSec, the chosen thresholdDb + thresholdNote (why that threshold won), measured noiseFloorDb, the silences that will be cut, keepRanges, removedSec, the full sweep table, and wordGuard (how many seconds the transcript vetoed).
  • deadWeight - candidates[] with kind (filler | stutter | repeat-take | hesitation), start/end, text, confidence, reason, context; plus totalSec and byKind.
  • guarded - true only when a transcript was found. If false, fix that before cutting.

Nothing is encoded, so call it as often as you like (a 217s source analyzes in about a second).

Step 2 - REVIEW every candidate (your job, not the machine's)

The list is deterministic: same transcript, same candidates, same confidence. What it cannot do is understand meaning. Go through them one by one and keep only the ones you would defend:

  • Low confidence means "read the context first", not "probably fine". Vietnamese fillers almost always collide with real words. đó, ấy, thế, mà, là are both filler particles and real demonstratives/conjunctions.
  • Connector phrases (hoặc là, tức là, bởi vì là, với lại là) are the trap. Measured on a real transcript: "…có thể là ok ứng dụng nó / Hoặc là / Tham khảo để tìm cách…" is a genuine alternative - cutting it destroys one branch of the sentence. But "Còn trường hợp mà mọi người có thể nghe / Hoặc là / Còn trường hợp mà người AI không ứng dụng được…" is an abandoned sentence and should go. The surface form is identical; only the meaning separates them. That is exactly why these come back with a low base confidence and a reason that says READ THE CONTEXT.
  • repeat-take candidates: check that the later take really is the fuller one before approving.
  • hesitation candidates are pure silence between words - usually safe, but check you are not removing a deliberate beat before a punchline.

Reject freely. A filler left in costs 0.4s; a real word cut makes the sentence nonsense and the viewer hears it immediately.

Step 3 - REPEATED POINTS (only you can find these - mandatory)

No detector catches this, and it is the dead weight users complain about most. Read the transcript as a piece of speech and analyze it semantically:

  1. Group sentences that make the SAME POINT - the words need not match, only the content. A speaker often restates one point 2-3 times: short/stumbling first, more complete later. Consider NON-adjacent sentences too (makes point A, rambles, comes back to point A in more detail).
  2. Keep exactly ONE version per group - the MOST COMPLETE one:
    • Earlier short version, later fuller restatement -> keep the later, cut the earlier.
    • Later sentence is only a short echo ("đúng vậy, như tôi nói…") -> keep the earlier, cut the echo.
    • Two equivalent versions -> keep the smoother delivery (fewer stumbles, no fillers).
  3. Check the flow after cutting: re-read the kept text end to end. If a kept sentence refers back ("như vừa nói") to one you cut, either keep both or pick the version without the back-reference.
  4. Turn each decision into a {start, end} range and send it with the approved candidates in Step 4.
  5. Build a table before cutting (put it in the report): timestamp | sentence cut | reason | sentence kept - so the user can review every decision.

Step 4 - Apply (goes through the render queue)

POST http://localhost:6869/api/projects/<id>/auto-trim/apply
body: { "cutCandidates": [{ "start": 51.81, "end": 52.25 }, …] }   # ONLY what you approved
-> 202 { job }

Send an empty list if you approved nothing - the measured silences still get cut. Then poll GET /api/jobs/<jobId> until it finishes. The job:

  1. Re-analyzes with the transcript guard, merges the silence ranges with your approved ranges (overlaps merged, so nothing is double counted).
  2. Drops any approved range that would swallow a word midpoint - a last safety net over your review, and it logs every range it rejects. Rejected ranges in the log mean you approved something that sits on real speech; go back and re-read that spot.
  3. Cuts in ONE ffmpeg pass (trim/atrim + setpts/asetpts + concat), writing assets/<stem>.cut.mp4.
  4. Remaps every timestamp into the cut timeline -> assets/transcript.cut.json.
  5. Verifies the OUTPUT with the remapped words and writes assets/auto-trim-report.json.

If nothing is worth cutting, the job deliberately does NOT produce a .cut.mp4 (a re-encoded identical copy only loses quality) - output in the report is null and you keep using the source.

Step 5 - Read the report (the job is not done until you do)

assets/auto-trim-report.json holds: before/after duration, removed.silenceSec vs removed.approvedSec, the chosen threshold and why, rejected candidates, and verification.

  • verdict: "pass" - the result meets the profile for that level.
  • verdict: "fail" - the job still finished and the file is usable, but it did NOT meet the profile. The log says so explicitly. Approve more dead weight (Step 2/3) and run apply again, or state in the final report why you are accepting it. Never report the cut as done while ignoring a fail.

The pass criteria are not the old "no silence > 0.8s" rule - measurement showed that rule is far too lax. A real file passed it while carrying 13 residual silences of 0.45-0.67s totalling 6.7s (4.2% of the runtime), which sounds obviously draggy. The profiles now cap BOTH the longest single silence and the total ratio (default: 0.5s / 3%).

Show full SKILL.md (656 more words)Show less

Step 6 - Use the cut version everywhere

From here on, captions, key layout, SFX timing and zooms read assets/transcript.cut.json and the .cut.mp4. The QC check dead-air re-measures the assembled video against the transcript and FAILS when brief.autoCut is on and dead air is still over the profile - that is the backstop, not the plan.

Final report must include: the cut table from Step 3, seconds removed split into silence vs approved dead weight, and the verification verdict. Take the numbers from auto-trim-report.json.

Known issues

  • Using the OLD transcript after cutting -> captions drift further out of sync toward the end. Always switch to transcript.cut.json.
  • Cutting only silences and ignoring repeated points - what users call "dead weight" is mostly Step 3.
  • Cutting the rendered draft instead of the source -> quality drops from a second encode. Always cut from the source (the server refuses to overwrite the source, and skips files already named *.cut.* when picking a default).
  • Analyzing a .cut.mp4 while the only transcript is the raw one -> the guard is aligned to the wrong timeline. Match the source file to the transcript that describes it.
  • SFX/narration files with their own lead silence -> use data-media-start to trim inside HyperFrames (see "SFX loudness and lead silence" in the noti-tiktok-vn skill), do not re-encode an audio file just for 0.3s of leading silence.
  • Joining cut segments without a fade -> a click at every join. Fixed in autoTrim.ts; the rule is below, and it applies to any new code that cuts or concatenates audio.

Audio: fade 30ms at every cut edge (hard rule)

A cut lands mid-waveform. The last sample of one segment and the first sample of the next are unrelated, so the join is a vertical step, and a step is a click. The cause is not the encoder and not the source - it is the join itself, so it survives any re-encode downstream.

How much it matters depends on WHERE the cut lands, and the difference is large. Measured losslessly on a real 278s talking-head, comparing the same cut with and without the fade (largest sample-to-sample step at the join; anything above ~0.02 is audible):

Kind of joinBeforeAfter
Silence trim, 40 joins (balanced)median 0.0040, worst 0.0113 - none audiblemedian 0.0000
Mid-speech, 4 joins (dead weight / repeated points)0.1195, 0.0744, 0.0232, 0.0064 - 3 of 4 audibleall ≤ 0.0008

For scale, the loudest step anywhere in normal speech in those files was 0.06-0.13. So a mid-speech join without the fade can be a bigger jump than anything the content itself produces, while a silence join is nowhere near audible.

Do not skip the fade on the strength of that first row. Plain silence trimming is the case where it happens not to matter; Step 3 (repeated points) and clip extraction cut straight through speech, and those are the joins that click.

Fix: fade both edges of every segment over 30ms. Long enough to kill the step, far too short to hear as a fade. Use audioCutFade() in apps/server/src/util.ts; do not hand-roll the filter.

[0:a]atrim=start=S:end=E,asetpts=PTS-STARTPTS,afade=t=in:st=0:d=0.030,afade=t=out:st=D-0.030:d=0.030[aN]

Three things that make it silently do nothing:

  • Order. afade must come AFTER asetpts=PTS-STARTPTS. afade reads st off the stream's own clock; on original timestamps st=0 is in the past and the fade-in never fires.
  • Length. D is the segment duration, so the fade-out start is D - 0.030. Clamp the fade to D/2 or the two fades overlap on a short segment and swallow it.
  • Extracting a clip, not just joining. A clip cut out of a longer video has the same two raw edges even though nothing is being joined - reframe.ts fades them via -af.

One note on measuring this yourself: do not compare two AAC encodes. Re-encoding perturbs samples everywhere, which buried the 40 real joins among 294 spurious difference regions and put a 54ms offset between the two files. Render both variants straight to pcm_s16le instead - then the two are sample-aligned and the joins sit exactly at the cumulative segment lengths.

© notivn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/auto-cut of notivn/AIEV.

Open the folder on GitHubat commit 1a4c2b0

Compare with similar skills

Auto Cut next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Auto Cut compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Auto Cut this skillnotivn/AIEV126—~3.2kAutomated safety check: PassMIT
HyperFrames Video Entry Pointheygen-com/hyperframes59k3 repos~5.2kAutomated safety check: PassApache-2.0
Cut Silencesnateherkai/hyperframes-student-kit1.2k—~1.2kAutomated safety check: PassCustom licence
Anime Cel Video Makeredenfunf/reelmimic1.8k—~1.6kAutomated safety check: PassMIT
Paper Cut-out Animationedenfunf/reelmimic1.8k—~1.5kAutomated safety check: PassMIT
p5 Paint Animationheygen-com/hyperframes-community-skills183—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    59k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Cut Silences

    nateherkai/hyperframes-student-kit

    Agent 1 of the video editing pipeline. An agent skill from nateherkai/hyperframes-student-kit.

    1.2k GitHub stars~1.2k tokensUpdated 11 days ago
    Media & CreativeAuto-check passed
  • Anime Cel Video Maker

    edenfunf/reelmimic

    Builds Japanese TV-anime style MP4 videos in code, with cel-shaded characters, painted backgrounds and staging effects, rendered in headless Chrome and encoded with ffmpeg.

    1.8k GitHub stars~1.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Paper Cut-out Animation

    edenfunf/reelmimic

    Builds paper cut-out, stop-motion style MP4 videos in code, with jointed paper puppets, torn or scissor-cut edges and soft shadows, rendered frame by frame in headless Chrome.

    1.8k GitHub stars~1.5k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • p5 Paint Animation

    heygen-com/hyperframes-community-skills

    Turns a text prompt, photo or short video into hand-made looking p5.js animation: self-writing handwriting, brushstroke repaints and living paintings rendered offline.

    183 GitHub stars~1.8k tokensUpdated 10 days ago
    Media & CreativeAuto-check passed
  • Pixel Art Video Engine

    edenfunf/reelmimic

    Makes animated pixel-art videos in an indie-game style: a 480×270 canvas upscaled four times, limited palettes with dithering, procedural characters and dialogue captions, encoded as MP4.

    1.8k GitHub stars~1.4k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed

More from notivn/AIEV

All 14 skills in this repo
  • Background Music

    notivn/AIEV

    Pick background music from the assets/music/ library and configure auto-ducking (music dips automatically under speech) via meta.json audio.music for the Remotion assembly layer.

    126 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Color Grading

    notivn/AIEV

    Color grading video in the AI Edit Video system - delog/tonemap HDR-HLG-log footage, apply the color preset the user approved in the UI, and the visual verification workflow.

    126 GitHub stars~932 tokensUpdated yesterday
    Auto-check passed
  • Build a Vietnamese vertical TikTok explainer in the "MỔ XẺ PAPER AI" (AI paper dissection) format with HyperFrames (HTML/CSS/GSAP → MP4), Noti.vn style.

    126 GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Skill Authoring

    notivn/AIEV

    The standard for writing new skills for the AI Edit Video system - file structure, frontmatter, tone of voice, and how to accumulate production lessons into skills.

    126 GitHub stars~925 tokensUpdated yesterday
    Auto-check passed
  • Noti Tiktok Vn

    notivn/AIEV

    Edit a Vietnamese vertical TikTok video (9:16) with HyperFrames following the Noti.vn/GĐT standard - talking-head + kinetic typography + karaoke captions + zoom/punch-in camera + timestamp-synced…

    126 GitHub stars~5.4k tokensUpdated yesterday
    Auto-check passed
  • Build a Vietnamese landscape 16:9 YouTube video (1920×1080) with HyperFrames (HTML/CSS/GSAP → MP4), keeping the Noti.vn/GĐT branding inherited from noti-tiktok-vn.

    126 GitHub stars~5.6k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Auto Cut

What does Auto Cut do?

Cut silences and dead weight (fillers, repeated takes, false starts) out of a talking-head video BEFORE building the edit - call the measured /auto-trim API instead of hand-rolling ffmpeg, review…. Auto Cut is an agent skill from notivn/AIEV. Cut silences and dead weight (fillers, repeated takes, false starts) out of a talking-head video BEFORE building the edit - call the measured /auto-trim API instead of hand-rolling ffmpeg, review the dead-weight candidates it returns, and do the one job only a human/AI can do (spotting repeated POINTS).

When should I use Auto Cut?

Auto Cut fits situations like: tasks that involve Video production; tasks that involve Motion graphics.

How do I install Auto Cut in Claude Code?

Run `npx skills add notivn/AIEV --skill auto-cut -a claude-code`. Or copy the skill folder (.claude/skills/auto-cut in notivn/AIEV) into .claude/skills/auto-cut in your project. Claude Code loads it when a task matches its description.

How do I install Auto Cut in Codex?

Run `npx skills add notivn/AIEV --skill auto-cut -a codex`. Or copy the skill folder (.claude/skills/auto-cut in notivn/AIEV) into .agents/skills/auto-cut in your project. Codex loads it when a task matches its description.

Can I use Auto Cut in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add notivn/AIEV --skill auto-cut -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/auto-cut, .gemini/skills/auto-cut, .github/skills/auto-cut and .opencode/skills/auto-cut in your project.

What does Auto Cut need to run?

SKILL.md names no scripts, command-line tools or credentials: Auto Cut is instructions for the agent only.

Does Auto Cut access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Auto Cut safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Auto Cut use?

Auto Cut is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Auto Cut use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Auto Cut?

Skills that share tags, products or a category with Auto Cut: HyperFrames Video Entry Point (heygen-com/hyperframes, 59k stars), Cut Silences (nateherkai/hyperframes-student-kit, 1.2k stars), Anime Cel Video Maker (edenfunf/reelmimic, 1.8k stars) and Paper Cut-out Animation (edenfunf/reelmimic, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Auto Cut?

notivn (a GitHub organization) maintains it in notivn/AIEV, which has 126 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 8, 2026.

Source: notivn/AIEV on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.