Agent skill

Video Voiceover And SFX Generator

by 0xsline in 0xsline/OpenChatCut

Generates text-to-speech narration and custom sound effects for a video timeline, keeping existing voiceover in sync after visual retiming edits.

AGPL-3.0Auto-check passedMedia & Creative

Install Video Voiceover And SFX Generator

skills CLI
$ npx skills add 0xsline/OpenChatCut --skill voice -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install 0xsline/OpenChatCut voice --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/0xsline/OpenChatCut.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/agent/skills/voice .claude/skills/voice && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice
GitHub stars
2.2k
Token cost
~4.4k tokens
SKILL.md length
1,759 words
Files
4 (incl. references)
Skills in repo
31
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Generates text-to-speech narration and custom sound effects for a video timeline, keeping existing voiceover in sync after visual retiming edits.

  • Works in 6 steps: Filter references/voices.md by target… → If no preset matches all explicit… → Pick 2-4 matching curated presets. → …
  • Generating narration or voiceover audio from a script for a video
  • SKILL.md covers When to Use, TTS (Text-to-Speech), Voice Audition Before Generation and Sound Effects, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

This skill generates voiceover audio and sound effects for a video timeline, requiring a concrete provider and voice to be chosen before calling its generation tool — configured provider options can include several speech vendors, each used only when shown as available. A curated voice catalog covers only a few of them by name; any other provider needs a concrete voice id supplied directly.

When narration targets an existing visual sequence — a screen recording, slide animation, product demo, or edited clip — it reads a dedicated sync reference before drafting new narration or placing audio, even if the request never says the word sync, since timing and meaning may need to track what's on screen. The same reference applies when visuals get trimmed, reordered, or replaced and existing narration needs to stay aligned without being regenerated from scratch. For sound effects, it checks the existing library before generating a new one from a description.

When your agent uses it

  • Generating narration or voiceover audio from a script for a video
  • Keeping an existing voiceover aligned after retiming or reordering a timeline
  • Creating a custom sound effect not already in the effects library

Example prompts

  • “Generate voiceover for this product demo script.”
  • “Keep the narration in sync after I retimed the intro section.”
  • “Create a custom whoosh sound effect for this transition.”

Requirements

  • A configured text-to-speech provider and API access for it

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Filter references/voices.md by target narration language / provider and
  2. If no preset matches all explicit requirements, say there is no exact match
  3. Pick 2-4 matching curated presets.
  4. Load widget-forms, then call ask_followup_questions with voice options
  5. Wait for the user to choose.
  6. Call submit_voice with the selected preset ID as voiceId.

What it can do on your machine

Read from SKILL.md and the folder at commit 2e6f4a2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript and html).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Voiceover And SFX Generator loads about 4.4k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 1,759 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from 0xsline/OpenChatCut at commit 2e6f4a2, republished under its AGPL-3.0 licence (© 0xsline). 1,759 words, ~4,424 tokens.

Download SKILL.mdSave it as .claude/skills/voice/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
voice
description
Text-to-Speech (TTS), voiceover, narration placement/sync, and custom sound effects (SFX) generator. Use when the user wants generated speech from text, wants to add/replace/align narration or voiceover for an existing video/timeline, wants to keep existing voiceover synced after visual retiming edits, needs voice audition/selection, or explicitly wants a newly generated/custom sound effect that is not available in the Sound Effects library.
user-invocable
true

Voice & Sound Effects Generator

Generate voiceovers (TTS) and sound effects. For TTS, choose a concrete provider and voice before calling submit_voice.

When to Use

  • Generate voiceover/narration from text
  • Create text-to-speech audio for videos
  • Add, replace, or redo narration/voiceover for an existing video, timeline, screen recording, slide animation, product demo, B-roll edit, MG explainer, or other visual sequence
  • Keep existing narration/voiceover aligned after trimming, speeding up, slowing down, moving, reordering, or replacing the visuals it describes
  • Offer and audition TTS voice choices when the user has not picked a concrete voice
  • Generate custom sound effects from text descriptions only after checking the Sound Effects library first

TTS (Text-to-Speech)

If the current request has an existing visual target and the user wants narration, voiceover, dubbing, or replacement speech for that target, read references/video-sync.md before drafting new narration, using existing narration text to generate TTS, or placing audio. Do this even when the user did not explicitly say "sync" or "match the visuals"; the existence of a visual target means narration timing and meaning may need to follow on-screen content. Use the normal standalone TTS path only when there is no visual target or the user just wants an audio asset from text.

Also read references/video-sync.md when the timeline already has narration/voiceover and the user asks to change the visuals while keeping that voiceover aligned. This is a sync maintenance task even if no new TTS is needed.

Use submit_voice to create a TTS audio asset. The current MCP tool contract is:

  • provider is required. Configured choices may be doubao, elevenlabs, minimax, inworld, fishaudio, speechify, openai, gemini, mistral, or cartesia. All providers are opt-in; use only providers shown as configured in the capabilities prompt.
  • voiceId is required, concrete, and provider-specific. The only exception is deliberate MiniMax timbreWeights mixing, where voiceId must be empty. Do not mix catalogs.
  • The curated catalog in references/voices.md covers only Doubao, ElevenLabs, and MiniMax. Other providers have no bundled preset or sample catalog in OpenChatCut. Require a concrete voice ID from the user or their provider account; never invent a preset or /voice-samples/... URL.
  • AI SDK-backed fields are provider-specific: OpenAI supports modelId, speed, outputFormat, and instructions; Gemini supports modelId, outputFormat, and instructions; Mistral supports modelId and outputFormat; Cartesia supports modelId, speed, languageCode, and outputFormat. Omit unsupported or unrequested fields.
  • Inworld, Fish Audio, and Speechify accept only voiceId plus optional modelId. Do not pass expressive, speed, language, or output controls to these providers.
  • submit_voice creates an audio asset only. Timeline placement, replacement, trimming, and alignment happen later with timeline tools.
  • For long narration, multiple submit_voice calls can be useful: split at natural pauses, sentence groups, or script beat boundaries when the workflow benefits from separately timed or placed voice clips.
  • Doubao supports speedRatio, loudnessRatio, pitch, emotion, emotionScale, performancePrompt, and explicitDialect, but not every voice supports every expressive control. Check references/voices.md before using them.
  • ElevenLabs retains its official voice settings, language, seed, output, normalization, pronunciation-dictionary, continuity, logging, and latency controls. MiniMax retains its dedicated controls documented in references/minimax-tts.md.

Doubao control support for current curated voices:

  • vivi, xiaohe, yunzhou, xiaotian, naiqimengwa, yingtaowanzi, wenroumama, zhixingnv, dayi, jitangnv, liuchang, ruyayichen, morgan, qingcang, huiben, popo, yuanboxiaoshu, baqiqingshu, and tangseng support explicit emotion / emotionScale, performancePrompt, and ASMR-style prompt directions.
  • shuanglangshaonian supports performancePrompt and COT/QA-style instruction following, but does not support explicit emotion / emotionScale or ASMR-style control.
  • explicitDialect is only supported by vivi and can be dongbei, shaanxi, or sichuan.

ElevenLabs control support for current curated voices:

  • amelia, brittney, hope, jessica, arabella, jane, maria, mark, frederick, peter, james, jon, sully, david, and alex all support the same request-level controls; model-specific support is still validated by ElevenLabs.
  • These controls are not per-voice guarantees of a specific acting style. Use the preset tags/samples to pick a naturally suitable voice, then use the controls for moderate delivery changes.
  • For ElevenLabs eleven_v3, inline audio tags are available when the user asks for expressive delivery such as emotion, tone, nonverbal cues, accent hints, or local pacing. Official examples fit these useful TTS categories: emotion/tone tags such as [happy], [sad], [angry], [excited], [curious], [sarcastic], [crying], [annoyed], [appalled], [thoughtful], [surprised], and [mischievously]; vocal delivery and nonverbal cue tags such as [whispers], [laughs], [sighs], [exhales], [inhales deeply], [clears throat], [snorts], [swallows], [wheezing], and [coughs]; pacing/pause/local speed tags such as [slowly], [pause], [short pause], [long pause], [rushed], and [drawn out]; and accent/special-performance tags such as [strong X accent], for example [strong French accent], plus [sings], [singing], [woo], and [pirate voice]. Official examples are non-exhaustive; similar auditory tags can be tried when the user explicitly asks for that delivery and the tag describes how the voice should sound, not a visual action. Write tags directly in text, close to the short phrase they should affect. Treat tags as local guidance, not paragraph-wide controls.
  • For pauses and pacing in eleven_v3, use punctuation, text structure, shorter generated segments, or local audio tags such as [short pause] and [slowly] when needed.
ts
// English / multilingual via ElevenLabs
submit_voice({
  provider: "elevenlabs",
  text: "Hello world",
  voiceId: "peter",
});

// Chinese via Doubao
submit_voice({
  provider: "doubao",
  text: "你好世界",
  voiceId: "liuchang",
});

// With speed adjustment (Doubao only)
submit_voice({
  provider: "doubao",
  text: "这是一段稍快的中文旁白。",
  voiceId: "liuchang",
  speedRatio: 1.5,
});

// With expressive Doubao controls
submit_voice({
  provider: "doubao",
  text: "这次事故提醒我们,安全永远不能侥幸。",
  voiceId: "liuchang",
  emotion: "sad",
  emotionScale: 3,
  performancePrompt: "痛心但克制,语速稍慢,像新闻专题旁白",
  pitch: -1,
  speedRatio: 0.92,
});

// With ElevenLabs delivery controls
submit_voice({
  provider: "elevenlabs",
  text: "The launch changed how teams plan their daily work.",
  voiceId: "peter",
  speed: 0.95,
  stability: 0.4,
  similarityBoost: 0.8,
  outputFormat: "wav_44100",
});

// MiniMax TTS (when configured) — see references/minimax-tts.md
submit_voice({
  provider: "minimax",
  text: "欢迎使用视频编辑助手。",
  voiceId: "female-yujie",
  speed: 1,
  name: "VO · welcome",
});

// Cartesia shape after the user confirms the exact account voice ID.
// confirmedCartesiaVoiceId represents that supplied value, not a preset.
submit_voice({
  provider: "cartesia",
  text: "A concise product introduction.",
  voiceId: confirmedCartesiaVoiceId,
  modelId: "sonic-3",
  speed: 1,
  languageCode: "en",
  outputFormat: "mp3",
});

Voice Audition Before Generation

When the user needs TTS and has not already chosen a concrete voice, first separate providers with curated OpenChatCut choices from providers that require an account-specific voice ID.

For Doubao, ElevenLabs, or MiniMax, read references/voices.md before recommending, rendering, or submitting an option. Use it as the only source for curated preset IDs, provider choice, display labels, tags, and bundled sample URLs. Do not create voice options from memory, translated names, or broad user descriptions.

For Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, or Cartesia, do not offer an invented audition list or sample URL. Ask the user for the concrete voice ID from that configured provider. A broad description such as "warm female" is not a valid voiceId.

First determine two separate languages:

  • User conversation language: the language the user used to talk to you. Use this for the surrounding reply, form-visual label, visual-option name, and summary.
  • Target narration language: the language of the text being synthesized. Use this only to choose provider and voice catalog.

The audition widget's submit button is fixed to the default label in this build (submitLabel is accepted but not rendered); keep the question label and option labels in the user conversation language, not the target narration language. For example: English users see submit_label="Submit", Chinese users see submit_label="提交", and Spanish users see submit_label="Enviar".

"help me generate ... voice over in Chinese" is an English conversation asking for Chinese narration, so the audition widget copy stays in English while the voice candidates come from Doubao.

For a curated provider:

  1. Filter references/voices.md by target narration language / provider and explicit requirements such as gender, age range, tone, and use case.
  2. If no preset matches all explicit requirements, say there is no exact match and offer the closest supported presets with a clear caveat.
  3. Pick 2-4 matching curated presets.
  4. Load widget-forms, then call ask_followup_questions with voice options and real bundled audio samples.
  5. Wait for the user to choose.
  6. Call submit_voice with the selected preset ID as voiceId.

For a provider without a curated OpenChatCut catalog, ask for a free-text, concrete provider voice ID instead. Do not add media or synthesize a /voice-samples/... path. Wait for the user to supply/confirm the exact ID before calling submit_voice.

For each curated audition option, keep value, display label, media, and summary tied to the same preset row from references/voices.md. Use only the sample URLs recorded there. Keep value as the preset ID and media as its matching sample URL. Write name and summary in the user's conversation language. The target narration language only decides the provider/voice catalog. After submission, map the display name back to the preset ID from the same candidate list.

English request for Chinese narration:

html
<widget submit_label="Submit">
  <form-visual
    id="voiceId"
    label="For Chinese voiceover, I recommend a few voices to try:"
    required="true"
  >
    <visual-option
      value="vivi"
      name="Vivi"
      media="/voice-samples/doubao-vivi.mp3"
      aspect-ratio="16:5"
      summary="Female / young / friendly, general"
    />
    <visual-option
      value="xiaohe"
      name="Xiaohe"
      media="/voice-samples/doubao-xiaohe.mp3"
      aspect-ratio="16:5"
      summary="Female / young / soft, clear"
    />
    <visual-option
      value="yunzhou"
      name="Yunzhou"
      media="/voice-samples/doubao-yunzhou.mp3"
      aspect-ratio="16:5"
      summary="Male / young / neutral, business"
    />
  </form-visual>
</widget>

Chinese request for Chinese narration:

html
<widget submit_label="提交">
  <form-visual
    id="voiceId"
    label="我推荐这几个中文旁白音色,先试听一下:"
    required="true"
  >
    <visual-option
      value="morgan"
      name="Morgan"
      media="/voice-samples/doubao-morgan.mp3"
      aspect-ratio="16:5"
      summary="男 / 中年 / 低沉知识解说"
    />
    <visual-option
      value="zhixingnv"
      name="知性女声"
      media="/voice-samples/doubao-zhixingnv.mp3"
      aspect-ratio="16:5"
      summary="女 / 中年 / 冷静知识讲解"
    />
    <visual-option
      value="vivi"
      name="Vivi"
      media="/voice-samples/doubao-vivi.mp3"
      aspect-ratio="16:5"
      summary="女 / 年轻 / 亲切通用口播"
    />
  </form-visual>
</widget>
Show full SKILL.md (489 more words)Show less

Sound Effects

For ordinary editing sound effects (SFX), do not generate first. Use the built-in Sound Effects library before generating:

  1. Call browse_library with category:"sound-effects" and a query such as "whoosh", "camera shutter", "notification", "censor beep", or "record scratch".
  2. Inspect the returned library:sound:<id>.
  3. Place it with edit_item, using fromFrame as the sound's anchor/editorial moment frame:
ts
browse_library({
  category: "sound-effects",
  query: "short whoosh transition",
});

edit_item({
  adds: [
    {
      type: "audio",
      assetId: "library:sound:whoosh-short",
      fromFrame: 120,
      trackId: "A1",
    },
  ],
});

Only generate sound effects from text descriptions with submit_sound when:

  • The user explicitly asks for a generated/original/custom sound.
  • The requested sound is too specific for the existing Sound Effects library.
  • browse_library({ category:"sound-effects", query }) returns no suitable match.
ts
// Custom/generated sound effect after the library has no suitable match
submit_sound({ prompt: "A dog barking in the distance" });

// With custom duration (0.5-22 seconds)
submit_sound({
  prompt: "Thunder and heavy rain",
  durationSeconds: 15,
});

// High prompt adherence
submit_sound({
  prompt: "Sci-fi laser gun firing",
  promptInfluence: 0.8,
});

Tips for better results:

  • Be specific: "A dog barking loudly" vs just "dog"
  • Include context: "Footsteps on wooden floor in an empty room"
  • Specify style: "Cinematic whoosh" or "8-bit game sound"

Parameters

TTS
FieldDescriptionNotes
providerdoubao, elevenlabs, minimax, inworld, fishaudio, speechify, openai, gemini, mistral, or cartesiaRequired; configured choices only
textText to synthesizeRequired
voiceIdConcrete provider-specific voice IDRequired except MiniMax timbre mix
modelIdProvider model overrideElevenLabs, Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, Cartesia
speedSpeech speedElevenLabs, MiniMax, OpenAI, Cartesia
languageCodeLanguage hint/codeElevenLabs, Cartesia
outputFormatProvider-supported output formatElevenLabs, OpenAI, Gemini, Mistral, Cartesia
instructionsNatural-language delivery directionOpenAI, Gemini
speedRatioSpeech speedDoubao only
nameMedia-pool asset nameOptional
Sound Effects
FieldDescriptionNotes
promptSound descriptionRequired
durationSecondsDuration0.5-22 seconds
promptInfluencePrompt adherence0-1
nameAsset nameOptional

Voices

Use the submit_voice voiceId guide and references/voices.md for the current curated preset list, display labels, tags, and sample URLs.

Voice IDs are provider-specific — do NOT mix them

The curated catalog contains separate Doubao, ElevenLabs, and MiniMax IDs. vivi / dayi are only Doubao; mark / amelia / james are only ElevenLabs; female-yujie is only MiniMax. Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, and Cartesia require a concrete provider-specific ID confirmed by the user and have no bundled OpenChatCut samples.

Provider choice:

  • Honor an explicit configured provider first.
  • For Chinese narration, prefer a matching curated Doubao voice (or MiniMax when configured/requested). Use another provider only after the user chooses it and confirms its voice ID.
  • For English / multilingual narration, prefer a curated ElevenLabs voice. Use another provider only after the user chooses it and confirms its voice ID.
  • Never offer a provider that is not shown as configured in capabilities.

Hard rules — what you must NOT do

  1. Never use a voice ID from a different provider.
  2. Never submit TTS while the voice is only described broadly; require a concrete provider-specific ID confirmed by the user.
  3. Never recommend or render a curated TTS option before checking references/voices.md.
  4. Never invent presets or sample URLs for Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, or Cartesia.
  5. Never pass provider-specific fields to a provider that does not support them.
  6. Never claim stable age, regional accent, pronunciation dictionary, or exact duration controls unless the selected provider exposes them.
  7. Never replace original recorded speech with TTS unless the user asks.

© 0xsline, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in src/agent/skills/voice of 0xsline/OpenChatCut.

  • SKILL.md
  • references/minimax-tts.md
  • references/video-sync.md
  • references/voices.md

Open the folder on GitHubat commit 2e6f4a2

Compare with similar skills

Video Voiceover And SFX Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Voiceover And SFX Generator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Voiceover And SFX Generator this skill0xsline/OpenChatCut2.2k—~4.4kAutomated safety check: PassAGPL-3.0
HyperFrames Media Useheygen-com/hyperframes59k—~2.4kAutomated safety check: PassApache-2.0
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT
Sound Effectstadaspetra/loop2962 repos~1.1kAutomated safety check: PassMIT
Characteristic VoiceNoizAI/skills526—~1.8kAutomated safety check: PassNone
Sound FxNoizAI/skills526—~1.4kAutomated safety check: PassNone

Similar skills

  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    59k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Media & CreativeAuto-check passed
  • A skill your agent uses whenever the user wants speech to sound more human, companion-like, or emotionally expressive.

    526 GitHub stars~1.8k tokensUpdated 11 days ago
    Media & CreativeAuto-check passed
  • Sound Fx

    NoizAI/skills

    A skill your agent uses whenever the user wants to generate sound effects, ambient audio, or short audio clips from a text description.

    526 GitHub stars~1.4k tokensUpdated 11 days ago
    Media & CreativeAuto-check passed
  • Z Qwen Audio Studio

    tjxj/z-skills

    A skill your agent uses when creating complete generated audio with qwen-audio-3.1-tts-next, including podcasts, radio drama, advertisements, multiple speakers, reference voices, ambience, sound…

    547 GitHub stars~716 tokensUpdated 17 days ago
    Media & CreativeAuto-check passed

More from 0xsline/OpenChatCut

All 31 skills in this repo
  • OpenChatCut Video Editing

    0xsline/OpenChatCut

    Connects an MCP-capable agent to the local OpenChatCut video editor to inspect and edit projects through draft edit sessions, with manual approval by default.

    2.2k GitHub starsUsed in 1 repo~655 tokens
    Auto-check passed
  • Video Shader Generator

    0xsline/OpenChatCut

    Generates WebGL shaders for video effects, transitions, masks and color grades in the OpenChatCut editor, trying built-in catalog effects such as zoom before making anything new.

    2.2k GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • AI Image Generation

    0xsline/OpenChatCut

    Generates still images through the submit_image tool, choosing among Fal.ai, gpt-image-2, nano-banana, MiniMax image-01 and Grok Imagine by configured keys.

    2.2k GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Livestream to Clips

    0xsline/OpenChatCut

    Cuts a livestream recording into evidence-backed, platform-ready clips by combining transcript, visual, audio and genre-specific signals.

    2.2k GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed
  • Music Generation

    0xsline/OpenChatCut

    Generates instrumentals, songs, soundtracks and covers through Mureka, MiniMax, Atlas Cloud or Sonilo using the `submit_music` tool.

    2.2k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • AI Video Generation

    0xsline/OpenChatCut

    Submits AI video generation jobs to Fal.ai, Seedance, Kling, MiniMax Hailuo, xAI Grok Imagine or OFox for text-to-video, image-to-video, transitions and clip extension.

    2.2k GitHub stars~4.3k tokensUpdated 2 days ago
    Auto-check passed

Questions about Video Voiceover And SFX Generator

What does Video Voiceover And SFX Generator do?

Generates text-to-speech narration and custom sound effects for a video timeline, keeping existing voiceover in sync after visual retiming edits. This skill generates voiceover audio and sound effects for a video timeline, requiring a concrete provider and voice to be chosen before calling its generation tool — configured provider options can include several speech vendors, each used only when shown as available. A curated voice catalog covers only a few of them by name; any other provider needs a concrete voice id supplied directly.

When should I use Video Voiceover And SFX Generator?

Video Voiceover And SFX Generator fits situations like: generating narration or voiceover audio from a script for a video; keeping an existing voiceover aligned after retiming or reordering a timeline; creating a custom sound effect not already in the effects library.

How do I install Video Voiceover And SFX Generator in Claude Code?

Run `npx skills add 0xsline/OpenChatCut --skill voice -a claude-code`. Or copy the skill folder (src/agent/skills/voice in 0xsline/OpenChatCut) into .claude/skills/voice in your project. Claude Code loads it when a task matches its description.

How do I install Video Voiceover And SFX Generator in Codex?

Run `npx skills add 0xsline/OpenChatCut --skill voice -a codex`. Or copy the skill folder (src/agent/skills/voice in 0xsline/OpenChatCut) into .agents/skills/voice in your project. Codex loads it when a task matches its description.

Can I use Video Voiceover And SFX Generator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add 0xsline/OpenChatCut --skill voice -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice, .gemini/skills/voice, .github/skills/voice and .opencode/skills/voice in your project.

What does Video Voiceover And SFX Generator need to run?

SKILL.md names no scripts, command-line tools or credentials: Video Voiceover And SFX Generator is instructions for the agent only. Our summary lists: A configured text-to-speech provider and API access for it.

Does Video Voiceover And SFX Generator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Video Voiceover And SFX Generator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Video Voiceover And SFX Generator use?

Video Voiceover And SFX Generator is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Voiceover And SFX Generator use?

About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.8k tokens, read only when the agent opens those files.

What are the alternatives to Video Voiceover And SFX Generator?

Skills that share tags, products or a category with Video Voiceover And SFX Generator: HyperFrames Media Use (heygen-com/hyperframes, 59k stars), Music (tadaspetra/loop, 296 stars), Sound Effects (tadaspetra/loop, 296 stars) and Characteristic Voice (NoizAI/skills, 526 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Voiceover And SFX Generator?

0xsline (a GitHub user) maintains it in 0xsline/OpenChatCut, which has 2,211 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: 0xsline/OpenChatCut on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.