Agent skill

Freecut

by Moh4696 in Moh4696/freecut

Edit any video by conversation. An agent skill from Moh4696/freecut.

MITAuto-check: warningsMedia & Creative

Install Freecut

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add Moh4696/freecut --skill freecut -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Moh4696/freecut freecut --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
freecut
GitHub stars
280
Token cost
~5.9k tokens
SKILL.md length
2,616 words
Files
17
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Edit any video by conversation. An agent skill from Moh4696/freecut.

  • Works in 7 steps: LLM reasons from raw transcript +… → Audio is primary, visuals follow. Cut… → Ask → confirm → execute → iterate →… → …
  • Tasks that involve Transcription
  • SKILL.md covers Principle, Hard Rules (production…, Directory layout and Setup, plus 12 more sections
  • Runs Python scripts from its folder; calls npx, pip and uv; needs ELEVENLABS_API_KEY

What it does

Freecut is an agent skill from Moh4696/freecut. Edit any video by conversation. Transcribe (locally, free, by default), cut, color grade, generate overlay animations, burn subtitles — for talking heads, montages, tutorials, travel, interviews. No presets, no menus. Ask questions, confirm the plan, execute, iterate, persist. Production-correctness rules are hard; everything else is artistic freedom.

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files (for example `README.md`, `helpers/grade.py` and `helpers/pack_transcripts.py`).

It sits in Media & Creative, covering Transcription. The repository describes itself as: Fork of browser-use/video-use with the paid ElevenLabs dependency replaced by free, pluggable transcription (local Whisper by default; VibeVoice-ASR for diarization). The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription

Example prompts

  • “/freecut”

Requirements

  • Python 3
  • Node.js
  • A credential in ELEVENLABS_API_KEY

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. LLM reasons from raw transcript + on-demand visuals. The only derived artifact that earns its keep is a packed phrase-level transcript…
  2. Audio is primary, visuals follow. Cut candidates come from speech boundaries and silence gaps. Drill into visuals only at decision points.
  3. Ask → confirm → execute → iterate → persist. Never touch the cut until the user has confirmed the strategy in plain English.
  4. Generalize. Do not assume what kind of video this is. Look at the material, ask the user, then edit.
  5. Artistic freedom is the default. Every specific value, preset, font, color, duration, pitch structure, and technique in this document is a…
  6. Invent freely. If the material calls for a technique not described here — split-screen, picture-in-picture, lower-third identity cards…
  7. Verify your own output before showing it to the user. If you wouldn't ship it, don't present it.

What it can do on your machine

Read from SKILL.md and the folder at commit f1d4334. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • pip
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, pip and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ELEVENLABS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Freecut loads about 5.9k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 2,616 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningTells the agent its actions are pre-authorized / not to stop for confirmationSKILL.md:15
    ey can do anything the format supports. Do not wait for permission.
  • NoteMentions a .env fileSKILL.md:62
    EVENLABS_API_KEY` in the environment or `.env` at the freecut repo root.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Moh4696/freecut at commit f1d4334, republished under its MIT licence (© Moh4696). 2,616 words, ~5,854 tokens.

Download SKILL.mdSave it as .claude/skills/freecut/SKILL.md (or your agent's skills folder). This skill also uses 16 other files; get the full folder from GitHub.
name
freecut
description
Edit any video by conversation. Transcribe (locally, free, by default), cut, color grade, generate overlay animations, burn subtitles — for talking heads, montages, tutorials, travel, interviews. No presets, no menus. Ask questions, confirm the plan, execute, iterate, persist. Production-correctness rules are hard; everything else is artistic freedom.

freecut

Principle

  1. LLM reasons from raw transcript + on-demand visuals. The only derived artifact that earns its keep is a packed phrase-level transcript (takes_packed.md). Everything else — filler tagging, retake detection, shot classification, emphasis scoring — you derive at decision time.
  2. Audio is primary, visuals follow. Cut candidates come from speech boundaries and silence gaps. Drill into visuals only at decision points.
  3. Ask → confirm → execute → iterate → persist. Never touch the cut until the user has confirmed the strategy in plain English.
  4. Generalize. Do not assume what kind of video this is. Look at the material, ask the user, then edit.
  5. Artistic freedom is the default. Every specific value, preset, font, color, duration, pitch structure, and technique in this document is a worked example from one proven video — not a mandate. Read them to understand what's possible and why each worked. Then make your own taste calls based on what the material actually is and what the user actually wants. The only things you MUST do are in the Hard Rules section below. Everything else is yours.
  6. Invent freely. If the material calls for a technique not described here — split-screen, picture-in-picture, lower-third identity cards, reaction cuts, speed ramps, freeze frames, crossfades, match cuts, L-cuts, J-cuts, speed ramps over breath, whatever — build it. The helpers are ffmpeg and PIL. They can do anything the format supports. Do not wait for permission.
  7. Verify your own output before showing it to the user. If you wouldn't ship it, don't present it.

Hard Rules (production correctness — non-negotiable)

These are the things where deviation produces silent failures or broken output. They are not taste, they are correctness. Memorize them.

  1. Subtitles are applied LAST in the filter chain, after every overlay. Otherwise overlays hide captions. Silent failure.
  2. Per-segment extract → lossless -c copy concat, not single-pass filtergraph. Otherwise you double-encode every segment when overlays are added.
  3. 30ms audio fades at every segment boundary (afade=t=in:st=0:d=0.03,afade=t=out:st={dur-0.03}:d=0.03). Otherwise audible pops at every cut.
  4. Overlays use setpts=PTS-STARTPTS+T/TB to shift the overlay's frame 0 to its window start. Otherwise you see the middle of the animation during the overlay window.
  5. Master SRT uses output-timeline offsets: output_time = word.start - segment_start + segment_offset. Otherwise captions misalign after segment concat.
  6. Never cut inside a word. Snap every cut edge to a word boundary from the transcript.
  7. Pad every cut edge. Working window: 30–200ms. ASR timestamps drift 50–100ms — padding absorbs the drift. Tighter for fast-paced, looser for cinematic.
  8. Word-level verbatim ASR only. Never SRT/phrase mode (loses sub-second gap data). Never normalized fillers (loses editorial signal).
  9. Cache transcripts per source. Never re-transcribe unless the source file itself changed.
  10. Parallel sub-agents for multiple animations. Never sequential. Spawn N at once via the Agent tool; total wall time ≈ slowest one.
  11. Strategy confirmation before execution. Never touch the cut until the user has approved the plain-English plan.
  12. All session outputs in <videos_dir>/edit/. Never write inside the freecut/ project directory.

Everything else in this document is a worked example. Deviate whenever the material calls for it.

Directory layout

The skill lives in freecut/. User footage lives wherever they put it. All session outputs go into <videos_dir>/edit/.

<videos_dir>/
├── <source files, untouched>
└── edit/
    ├── project.md               ← memory; appended every session
    ├── takes_packed.md          ← phrase-level transcripts, the LLM's primary reading view
    ├── edl.json                 ← cut decisions
    ├── transcripts/<name>.json  ← cached normalized ASR JSON
    ├── animations/slot_<id>/    ← per-animation source + render + reasoning
    ├── clips_graded/            ← per-segment extracts with grade + fades
    ├── master.srt               ← output-timeline subtitles
    ├── downloads/               ← yt-dlp outputs
    ├── verify/                  ← debug frames / timeline PNGs
    ├── preview.mp4
    └── final.mp4

Setup

First-time install lives in install.md (clone, deps, ffmpeg, skill registration, API key). Don't re-run it every session; on cold start just verify:

  • Transcription backend is available. The default is whisper (free, local) — verify mlx-whisper (Apple Silicon) or faster-whisper is importable. If neither is installed, tell the user to pip install mlx-whisper (or faster-whisper) rather than falling through to a paid backend silently. For the optional vibevoice / elevenlabs backends, check VIBEVOICE_ASR_URL / ELEVENLABS_API_KEY in the environment or .env at the freecut repo root.
  • ffmpeg + ffprobe on PATH.
  • Python deps installed (uv sync or pip install -e . inside the repo).
  • Node.js + npm available if the session needs HyperFrames or Remotion slots. HyperFrames currently requires Node.js 22+.
  • yt-dlp, HyperFrames, Remotion, Manim installed only on first use.
  • First-use animation setup happens inside the slot directory, never at the freecut repo root. HyperFrames can be invoked with npx --yes hyperframes ...; Remotion can be scaffolded with npx create-video@latest or installed as a project-local dependency before using its remotion render command.
  • This skill vendors skills/manim-video/. Read its SKILL.md when building a Manim slot.

Helpers (helpers/transcribe.py, helpers/render.py, etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at ~/.claude/skills/freecut/ or ~/.codex/skills/freecut/.

Helpers

  • transcribe.py <video> — single-file transcription. Pluggable backend via --backend {whisper,vibevoice,elevenlabs} (default whisper, local + free, single speaker). vibevoice uses VIBEVOICE_ASR_URL for real diarization; elevenlabs uses ELEVENLABS_API_KEY and supports --num-speakers. Cached.
  • transcribe_batch.py <videos_dir> — 4-worker parallel transcription. Same --backend flag. Use for multi-take.
  • pack_transcripts.py --edit-dir <dir> — transcripts/*.json → takes_packed.md (phrase-level, break on silence ≥ 0.5s or speaker change). Works identically across backends because the JSON shape is normalized.
  • timeline_view.py <video> <start> <end> — filmstrip + waveform PNG. On-demand visual drill-down. Not a scan tool — use it at decision points, not constantly.
  • render.py <edl.json> -o <out> — per-segment extract → concat → overlays (PTS-shifted) → subtitles LAST. --preview for 720p fast. --build-subtitles to generate master.srt inline.
  • grade.py <in> -o <out> — ffmpeg filter chain grade. Presets + --filter '<raw>' for custom.

For animations, create <edit>/animations/slot_<id>/ with Bash and spawn a sub-agent via the Agent tool.

The process

  1. Inventory. ffprobe every source. transcribe_batch.py on the directory. pack_transcripts.py to produce takes_packed.md. Sample one or two timeline_views for a visual first impression.

  2. Pre-scan for problems. One pass over takes_packed.md to note verbal slips, obvious mis-speaks, or phrasings to avoid. Plain list, feed into the editor brief.

  3. Converse. Describe what you see in plain English. Ask questions shaped by the material. Collect: content type, target length/aspect, aesthetic/brand direction, pacing feel, must-preserve moments, must-cut moments, animation and grade preferences, subtitle needs. Do not use a fixed checklist — the right questions are different every time.

  4. Propose strategy. 4–8 sentences: shape, take choices, cut direction, animation plan, grade direction, subtitle style, length estimate. Wait for confirmation.

  5. Execute. Produce edl.json via the editor sub-agent brief. Drill into timeline_view at ambiguous moments. Build animations in parallel sub-agents. Apply grade per-segment. Compose via render.py.

  6. Preview. render.py --preview.

  7. Self-eval (before showing the user). Run timeline_view on the rendered output (not the sources) at every cut boundary (±1.5s window). Check each image for:

    • Visual discontinuity / flash / jump at the cut
    • Waveform spike at the boundary (audio pop that slipped past the 30ms fade)
    • Subtitle hidden behind an overlay (Rule 1 violation)
    • Overlay misaligned or showing wrong frames (Rule 4 violation)

    Also sample: first 2s, last 2s, and 2–3 mid-points — check grade consistency, subtitle readability, overall coherence. Run ffprobe on the output to verify duration matches the EDL expectation.

    If anything fails: fix → re-render → re-eval. Cap at 3 self-eval passes — if issues remain after 3, flag them to the user rather than looping forever. Only present the preview once the self-eval passes.

  8. Iterate + persist. Natural-language feedback, re-plan, re-render. Never re-transcribe. Final render on confirmation. Append to project.md.

Cut craft (techniques)

  • Audio-first. Candidate cuts from word boundaries and silence gaps.
  • Preserve peaks. Laughs, punchlines, emphasis beats. Extend past punchlines to include reactions — the laugh IS the beat.
  • Speaker handoffs benefit from air between utterances. Common values: 400–600ms. Less for fast-paced, more for cinematic. Taste call.
  • Audio events as signals. (laughs), (sighs), (applause) mark beats. Extend past them.
  • Silence gaps are cut candidates. Silences ≥400ms are usually the cleanest. 150–400ms phrase boundaries are usable with a visual check. <150ms is unsafe (mid-phrase).
  • Example cut padding (the launch video shipped with this): 50ms before the first kept word, 80ms after the last. Tighter for montage energy, looser for documentary. Stay in the 30–200ms working window (Hard Rule 7).
  • Never reason audio and video independently. Every cut must work on both tracks.

The packed transcript (primary reading view)

pack_transcripts.py reads all transcripts/*.json and produces one markdown file where each take is a list of phrase-level lines, each prefixed with its [start-end] time range. Phrases break on any silence ≥ 0.5s OR speaker change. This is the artifact the editor sub-agent reads to pick cuts — it gives word-boundary precision from text alone at 1/10 the tokens of raw JSON.

Example line:

## C0103  (duration: 43.0s, 8 phrases)
  [002.52-005.36] S0 Ninety percent of what a web agent does is completely wasted.
  [006.08-006.74] S0 We fixed this.

Editor sub-agent brief (for multi-take selection)

When the task is "pick the best take of each beat across many clips," spawn a dedicated sub-agent with a brief shaped like this. The structure is load-bearing; the pitch-shape example is not.

You are editing a <type> video. Pick the best take of each beat and 
assemble them chronologically by beat, not by source clip order.

INPUTS:
  - takes_packed.md (time-annotated phrase-level transcripts of all takes)
  - Product/narrative context: <2 sentences from the user>
  - Speaker(s): <name, role, delivery style note>
  - Expected structure: <pick an archetype or invent one>
  - Verbal slips to avoid: <list from the pre-scan pass>
  - Target runtime: <seconds>

Common structural archetypes (pick, adapt, or invent):
  - Tech launch / demo:   HOOK → PROBLEM → SOLUTION → BENEFIT → EXAMPLE → CTA
  - Tutorial:             INTRO → SETUP → STEPS → GOTCHAS → RECAP
  - Interview:            (QUESTION → ANSWER → FOLLOWUP) repeat
  - Travel / event:       ARRIVAL → HIGHLIGHTS → QUIET MOMENTS → DEPARTURE
  - Documentary:          THESIS → EVIDENCE → COUNTERPOINT → CONCLUSION
  - Music / performance:  INTRO → VERSE → CHORUS → BRIDGE → OUTRO
  - Or invent your own.

RULES:
  - Start/end times must fall on word boundaries from the transcript.
  - Pad cut boundaries (working window 30–200ms).
  - Prefer silences ≥ 400ms as cut targets.
  - Unavoidable slips are kept if no better take exists. Note them in "reason".
  - If over budget, revise: drop a beat or trim tails. Report total and self-correct.

OUTPUT (JSON array, no prose):
  [{"source": "C0103", "start": 2.42, "end": 6.85, "beat": "HOOK",
    "quote": "...", "reason": "..."}, ...]

Return the final EDL and a one-line total runtime check.

Color grade (when requested)

Your job is to reason about the image, not apply a preset. Look at a frame (via timeline_view), decide what's wrong, adjust one thing, look again.

Mental model is ASC CDL. Per channel: out = (in * slope + offset) ** power, then global saturation. slope → highlights, offset → shadows, power → midtones.

Example filter chains (grade.py has --list-presets; use them as starting points or mix your own):

  • warm_cinematic — retro/technical, subtle teal/orange split, desaturated. Shipped in a real launch video. Safe for talking heads.
  • neutral_punch — minimal corrective: contrast bump + gentle S-curve. No hue shifts.
  • none — straight copy. Default when the user hasn't asked.

For anything else — portraiture, nature, product, music video, documentary — invent your own chain. grade.py --filter '<raw ffmpeg>' accepts any filter string.

Hard rules: apply per-segment during extraction (not post-concat, which re-encodes twice). Never go aggressive without testing skin tones.

Show full SKILL.md (1,090 more words)Show less

Subtitles (when requested)

Subtitles have three dimensions worth reasoning about: chunking (1/2/3/sentence per line), case (UPPER/Title/Natural), and placement (margin from bottom). The right combo depends on content.

Worked styles — pick, adapt, or invent:

bold-overlay — short-form tech launch, fast-paced social. 2-word chunks, UPPERCASE, break on punctuation, Helvetica 18 Bold, white-on-outline, MarginV=35. render.py ships with this as SUB_FORCE_STYLE.

FontName=Helvetica,FontSize=18,Bold=1,
PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,BackColour=&H00000000,
BorderStyle=1,Outline=2,Shadow=0,
Alignment=2,MarginV=35

natural-sentence (if you invent this mode) — narrative, documentary, education. 4–7 word chunks, sentence case, break on natural pauses, MarginV=60–80, larger font for readability, slightly wider max-width. No shipped force_style — design one if you need it.

Invent a third style if neither fits. Hard rules: subtitles LAST (Rule 1), output-timeline offsets (Rule 5).

Animations (when requested)

Animations match the content and the brand. Get the palette, font, and visual language from the conversation — never assume a default. If the user hasn't told you, propose a palette in the strategy phase and wait for confirmation before building anything.

Tool options:

Pick the engine per animation slot. Do not default to Remotion just because the animation is web-adjacent.

  • HyperFrames — Browser-native HTML/CSS/GSAP video compositions: product UI motion, website-to-video or mockup-to-video captures, kinetic typography, landing-page/storyboard promos, data-driven UI states, transparent WebM overlays, and clips that need deterministic frame capture plus HyperFrames lint/validate/render checks. Best when the animation should be authored and verified like a web composition instead of a React component tree.
  • Remotion — React/CSS compositions with component state, reusable React primitives, or an existing Remotion brand system. Best when the user specifically asks for React/Remotion or when React composition is the simpler authoring model.
  • Manim — formal diagrams, state machines, equation derivations, graph morphs. Read skills/manim-video/SKILL.md and its references for depth.
  • PIL + PNG sequence + ffmpeg — simple overlay cards: counters, typewriter text, single bar reveals, progressive draws. Fast to iterate, any aesthetic you want. The launch video used this.

For HyperFrames slots, scaffold the slot inside edit/animations/slot_<id>/ with npx --yes hyperframes init . --example blank --non-interactive --skip-skills, build the HTML composition there, run the HyperFrames checks that fit the slot (lint, validate, and a draft render when practical), then produce the final overlay video with npx --yes hyperframes render . -o render.mp4 or --format webm -o render.webm when alpha is required. Point the EDL overlay file at the actual rendered path.

For Remotion slots, keep the Remotion project isolated inside the same slot directory, scaffold with npx create-video@latest or install Remotion locally there, render the composition to render.mp4 with the project-local remotion render command, and verify duration and dimensions with ffprobe.

None is mandatory. Invent hybrids if useful (e.g., PIL background with a HyperFrames or Remotion layer on top).

Duration rules of thumb, context-dependent:

  • Sync-to-narration explanations. A viewer needs to parse the content at 1×. Rough floor 3s, typical 5–7s for simple cards, 8–14s for complex diagrams. The launch video shipped at 5–7s per simple card.
  • Beat-synced accents (music video, fast montage). 0.5–2s is fine — they're visual accents, not information. The "readable at 1×" rule becomes "recognizable at 1×", not "fully parseable."
  • Hold the final frame ≥ 1s before the cut (universal).
  • Over voiceover: total duration ≥ narration_length + 1s (universal).
  • Never parallel-reveal independent elements — the eye can't track two new things at once. One thing, pause, next thing.

Animation payoff timing (rule for sync-to-narration): get the payoff word's timestamp. Start the overlay reveal_duration seconds earlier so the landing frame coincides with the spoken payoff word. Without this sync the animation feels disconnected.

Easing (universal — never linear, it looks robotic):

python
def ease_out_cubic(t):    return 1 - (1 - t) ** 3
def ease_in_out_cubic(t):
    if t < 0.5: return 4 * t ** 3
    return 1 - (-2 * t + 2) ** 3 / 2

ease_out_cubic for single reveals (slow landing). ease_in_out_cubic for continuous draws.

Typing text anchor trick: center on the FULL string's width, not the partial-string width — otherwise text slides left during reveal.

Example palette (the launch video — one aesthetic among infinite):

  • Background (10, 10, 10) near-black
  • Accent #FF5A00 / (255, 90, 0) orange
  • Labels (110, 110, 110) dim gray
  • Font: Menlo Bold at /System/Library/Fonts/Menlo.ttc (index 1)
  • ≤ 2 accent colors, ~40% empty space, minimal chrome
  • Result: terminal / retro tech feel

This is one style. If the brand is warm and serif, use that. If it's colorful and playful, use that. If the user handed you a style guide, follow it. If they didn't, propose one and confirm.

Parallel sub-agent brief — each animation is one sub-agent spawned via the Agent tool. Each prompt is self-contained (sub-agents have no parent context). Include:

  1. One-sentence goal: "Build ONE animation: [spec]. Nothing else."
  2. Absolute output path (<edit>/animations/slot_<id>/render.mp4)
  3. Exact technical spec: resolution, fps, codec, pix_fmt, CRF, duration
  4. Style palette as concrete values (RGB tuples, hex, or reference to a design system)
  5. Font path with index
  6. Frame-by-frame timeline (what happens when, with easing)
  7. Anti-list ("no chrome, no extras, no titles unless specified")
  8. Code pattern reference (copy helpers inline, don't import across slots)
  9. Deliverable checklist (script, render, verify duration via ffprobe, report)
  10. "Do not ask questions. If anything is ambiguous, pick the most obvious interpretation and proceed."

One sub-agent = one file (unique filenames, parallel agents don't overwrite each other).

Output spec

Match the source unless the user asked for something specific. Common targets: 1920×1080@24 cinematic, 1920×1080@30 screen content, 1080×1920@30 vertical social, 3840×2160@24 4K cinema, 1080×1080@30 square. render.py defaults the scale to 1080p from any source; pass --filter or edit the extract command for other targets. Worth asking the user which delivery format matters.

EDL format

json
{
  "version": 1,
  "sources": {"C0103": "/abs/path/C0103.MP4", "C0108": "/abs/path/C0108.MP4"},
  "ranges": [
    {"source": "C0103", "start": 2.42, "end": 6.85,
     "beat": "HOOK", "quote": "...", "reason": "Cleanest delivery, stops before slip at 38.46."},
    {"source": "C0108", "start": 14.30, "end": 28.90,
     "beat": "SOLUTION", "quote": "...", "reason": "Only take without the false start."}
  ],
  "grade": "warm_cinematic",
  "overlays": [
    {"file": "edit/animations/slot_1/render.mp4", "start_in_output": 0.0, "duration": 5.0}
  ],
  "subtitles": "edit/master.srt",
  "total_duration_s": 87.4
}

grade is a preset name or raw ffmpeg filter. overlays are rendered animation clips. subtitles is optional and applied LAST.

Memory — project.md

Append one section per session at <edit>/project.md:

markdown
## Session N — YYYY-MM-DD

**Strategy:** one paragraph describing the approach
**Decisions:** take choices, cuts, grades, animations + why
**Reasoning log:** one-line rationale for non-obvious decisions
**Outstanding:** deferred items

On startup, read project.md if it exists and summarize the last session in one sentence before asking whether to continue.

Anti-patterns

Things that consistently fail regardless of style:

  • Hierarchical pre-computed codec formats with USABILITY / tone tags / shot layers. Over-engineering. Derive from the transcript at decision time.
  • Hand-tuned moment-scoring functions. The LLM picks better than any heuristic you'll write.
  • Whisper SRT / phrase-level output. Loses sub-second gap data. Always word-level verbatim.
  • Slow CPU-only Whisper on a big machine. Use mlx-whisper on Apple Silicon, or faster-whisper with GPU/CoreML acceleration where available. Never fall back to a paid backend silently just because local is slow — tell the user.
  • Burning subtitles into base before compositing overlays. Overlays hide them. (Hard Rule 1.)
  • Single-pass filtergraph when you have overlays. Double re-encodes. Use per-segment extract → concat.
  • Linear animation easing. Looks robotic. Always cubic.
  • Hard audio cuts at segment boundaries. Audible pops. (Hard Rule 3.)
  • Typing text centered on the partial string. Text slides left as it grows.
  • Sequential sub-agents for multiple animations. Always parallel.
  • Editing before confirming the strategy. Never.
  • Re-transcribing cached sources. Immutable outputs of immutable inputs.
  • Assuming what kind of video it is. Look first, ask second, edit last.

© Moh4696, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 16 other files in the repository root of Moh4696/freecut.

  • SKILL.md
  • .env.example
  • .gitignore
  • LICENSE
  • README.md
  • helpers/grade.py
  • helpers/pack_transcripts.py
  • helpers/render.py
  • helpers/timeline_view.py
  • helpers/transcribe.py
  • helpers/transcribe_batch.py
  • install.md
  • poster.html
  • pyproject.toml
  • skills
  • static/timeline-view.svg
  • static/video-use-banner.png

Open the folder on GitHubat commit f1d4334

Compare with similar skills

Freecut next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Freecut compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Freecut this skillMoh4696/freecut280—~5.9kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Native Subtitle Quote Imagechengyi-ai/native-subtitle-quote-image2.6k—~2.4kAutomated safety check: PassMIT
Edu Chem Videowy51ai/edulab1.4k—~2.1kAutomated safety check: NotesApache-2.0
Transcription Memory ReconstructionNxcoreAI/EverRoom3k—~714Automated safety check: PassCustom licence
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0

Similar skills

  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Native Subtitle Quote Image

    chengyi-ai/native-subtitle-quote-image

    将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…

    2.6k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Edu Chem Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a chemistry problem (化学题: 氧化还原配平 双线桥 电子守恒, 物质的量计算, 化学平衡 三段式 平衡常数 转化率 反应速率, 离子反应, 电化学, 溶液 滴定…

    1.4k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Reconstruct a complete, searchable memory from an untrusted meeting or conversation transcript.

    3k GitHub stars~714 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Transcribe

    JetBrains/skills

    Official

    Transcribe audio files to text with optional diarization and known-speaker hints.

    366 GitHub starsUsed in 4 repos~776 tokens
    Media & CreativeAuto-check passed

Questions about Freecut

What does Freecut do?

Edit any video by conversation. An agent skill from Moh4696/freecut. Freecut is an agent skill from Moh4696/freecut. Edit any video by conversation.

When should I use Freecut?

Freecut fits situations like: tasks that involve Transcription.

How do I install Freecut in Claude Code?

Run `npx skills add Moh4696/freecut --skill freecut -a claude-code`. Or copy the skill folder (the Moh4696/freecut repository) into .claude/skills/freecut in your project. Claude Code loads it when a task matches its description.

How do I install Freecut in Codex?

Run `npx skills add Moh4696/freecut --skill freecut -a codex`. Or copy the skill folder (the Moh4696/freecut repository) into .agents/skills/freecut in your project. Codex loads it when a task matches its description.

Can I use Freecut in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Moh4696/freecut --skill freecut -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/freecut, .gemini/skills/freecut, .github/skills/freecut and .opencode/skills/freecut in your project.

What does Freecut need to run?

Going by SKILL.md and its folder, Freecut needs Python for the scripts in its folder, the command-line tools its instructions call (npx, pip and uv) and credentials named ELEVENLABS_API_KEY. Our summary lists: Python 3; Node.js; A credential in ELEVENLABS_API_KEY.

Does Freecut access the network?

SKILL.md contains no URLs. Its commands use npx, pip and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Freecut safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Freecut use?

Freecut is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Freecut use?

About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Freecut?

Skills that share tags, products or a category with Freecut: HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Native Subtitle Quote Image (chengyi-ai/native-subtitle-quote-image, 2.6k stars), Edu Chem Video (wy51ai/edulab, 1.4k stars) and Transcription Memory Reconstruction (NxcoreAI/EverRoom, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Freecut?

Moh4696 (a GitHub user) maintains it in Moh4696/freecut, which has 280 GitHub stars. The repository was last updated on July 1, 2026.

Source: Moh4696/freecut on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.