The Descript craft skill — edit talk content (podcasts, interviews, talking-head video) by editing the transcript instead of the timeline.

MITAuto-check passedMedia & Creative

Install Descript

skills CLI
$ npx skills add social-media-skills/skills --skill descript -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install social-media-skills/skills descript --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/social-media-skills/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/descript .claude/skills/descript && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
descript
GitHub stars
134
Token cost
~2k tokens
SKILL.md length
844 words
Files
6 (incl. references)
Skills in repo
106
Repo updated
First seen
Licence
MIT

At a glance

The Descript craft skill — edit talk content (podcasts, interviews, talking-head video) by editing the transcript instead of the timeline.

  • Works in 2 steps: The recording's content skill —… → brand-profile + voice-builder (written…
  • Someone wants to edit in Descript
  • SKILL.md covers The POV: the transcript is the…, Read these first, The framework: WORDS and The reality (verify-quarterly), plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Descript is an agent skill from social-media-skills/skills. The Descript craft skill — edit talk content (podcasts, interviews, talking-head video) by editing the transcript instead of the timeline. Use when someone wants to edit in Descript, edit a podcast or interview, remove filler words/silences, clean up audio (Studio Sound), fix a flubbed word without re-recording (Overdub), auto-cut between speakers, turn one recording into clips + show notes + chapters, or asks about Descript's plans, credits, or Underlord. Uses the WORDS framework plus the interview rule…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `evals/evals.json`, `references/descript-2026-reality.md` and `references/scope-and-connections.md`).

It sits in Media & Creative, covering Podcasting and Text to speech and voice. It works with Model Context Protocol. The repository describes itself as: 106 social media skills for AI agents - strategy, writing, video, design, platform growth, publishing, and analytics. Works with Claude, Cursor, OpenClaw, Hermes & 40+ agents. The licence is MIT.

When your agent uses it

  • Someone wants to edit in Descript
  • Remove filler words/silences
  • Clean up audio (Studio Sound)
  • Fix a flubbed word without re-recording (Overdub)

Example prompts

  • “/descript”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. The recording's content skill — podcast-and-audiograms / youtube-long-form / educational-content.
  2. brand-profile + voice-builder (written outputs) + design-and-templates (captions/layout).

What it can do on your machine

Read from SKILL.md and the folder at commit 6e30eeb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Descript loads about 2k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 252 tokens; SKILL.md has 844 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~252
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from social-media-skills/skills at commit 6e30eeb, republished under its MIT licence (© social-media-skills). 844 words, ~2,043 tokens.

Download SKILL.mdSave it as .claude/skills/descript/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
descript
description
The Descript craft skill — edit talk content (podcasts, interviews, talking-head video) by editing the transcript instead of the timeline. Use when someone wants to edit in Descript, edit a podcast or interview, remove filler words/silences, clean up audio (Studio Sound), fix a flubbed word without re-recording (Overdub), auto-cut between speakers, turn one recording into clips + show notes + chapters, or asks about Descript's plans, credits, or Underlord. Uses the WORDS framework plus the interview rule: concision yes, meaning-flips never. Reads the recording's content skill + brand-profile/voice-builder first. The agent plans the edit (API/MCP where connected); the HUMAN verifies by ear and approves; WoopSocial publishes the exports. Overdub is consent-verified own-voice-only; tiers/credits are verified in-app. Distinct from capcut (visual short-form), captions-and-clipping/opus-clip (clip selection at scale), ai-voiceover (dedicated TTS), and podcast-and-audiograms (the strategy).
version
1.0.0

descript

The talk-content editing tool skill — write the edit in the transcript, Overdub with consent, refine the sound, dress the visuals, and ship the cuts. The agent plans (and can drive Underlord via API/MCP where connected); the human verifies by ear and approves; WoopSocial publishes the exports. (Ships with tools/integrations/descript.md.)

The POV: the transcript is the timeline — decide in text, verify by ear

For dialogue-heavy content, editing the transcript beats scrubbing a timeline: delete the sentence, the clip disappears; move the paragraph, the footage follows — reviews report ~60–70% editing-time cuts for talk content. But the paradigm has two sharp edges the top 1% respect. (1) The voice spine: Overdub's consent-verified, own-voice-only design is the model, not an obstacle — it exists so nobody types words into someone else's mouth; and the craft truth is it shines on flubbed words, not paragraphs (long Overdub drifts synthetic — re-record those). (2) The meaning spine: text-editing makes it dangerously easy to rearrange a guest into saying something they didn't — concision yes, meaning-flips never, and the human owns the final cut of anyone else's words. Operationally: the accuracy pass is mandatory (transcript errors become wrong edits AND wrong captions), and since the Sept 2025 overhaul, the workflow must be credit-aware — media minutes count everything you import, and formerly-unlimited AI features are metered.

Read these first

  1. The recording's content skill — podcast-and-audiograms / youtube-long-form / educational-content.
  2. brand-profile + voice-builder (written outputs) + design-and-templates (captions/layout).

The framework: WORDS

(Depth: references/the-words-framework.md.)

  • W — Write the edit: accuracy pass first; then cut tangents/bad takes, one-step filler+silence removal (Underlord), restructure by moving paragraphs — decide in text, verify by ear.
  • O — Overdub with consent: own-voice-only, consent-verified; single words/short phrases (paragraphs = re-record); vocabulary/credit limits; disclose synthetic speech where required.
  • R — Refine the sound: Studio Sound once per source (credits; cleaner audio also improves the transcript); level speakers; extreme noise is a re-record, not a rescue.
  • D — Dress the visuals: Automatic Multicam (record separate tracks on purpose), captions from the corrected transcript, human-approved B-roll, Eye Contact used honestly; beat-sync/color route elsewhere.
  • S — Ship the cuts: one transcript → the episode + clips (Underlord flags, the human picks fairly) + show notes + chapters + a text post; route onward and publish via WoopSocial.

The reality (verify-quarterly)

2026 Descript: Underlord (agentic co-editor — filler/silence in one step, bad-take flags, B-roll suggestions, clips, show notes) now triggerable via the 2026 public API (open beta) incl. MCP connections; Overdub (~24–48h training; source-audio requirements have varied — verify); Studio Sound (~10 credits/use); Automatic Multicam; Eye Contact; ~92–95% transcription accuracy on clean audio, ~75–85% with noise/accents/jargon; ~23 languages; SOC 2 Type II; cloud-dependent (no offline). Pricing: the Sept 2025 overhaul moved to media minutes + AI-credit metering of formerly-unlimited features; documented bill-shock and no mid-cycle proration (G2, attributed); tier figures conflict across sources — verify in-app. Full detail: references/descript-2026-reality.md. The weekly loop, credit-aware checklist, the Overdub decision table, the interview-integrity checklist, and two worked examples: references/workflows-and-templates.md.

Show full SKILL.md (370 more words)Show less

Honest scope (never violate)

  • The agent plans the edit and can trigger Underlord/media actions via the API/MCP where connected (exact human steps otherwise — no pretended automation); the human verifies by ear (pacing/tone/fairness don't live in text) and approves — the agent never fabricates "that cut sounds great." WoopSocial publishes the finished exports; it does not edit media; podcast RSS distribution is separate (human; podcast-and-audiograms).
  • Voice spine: own-voice-only cloning; never a guest/competitor/public figure; never fabricated words in a real mouth; AI-disclosure for synthetic speech where required (EU AI Act; C2PA). Meaning spine: interview edits preserve meaning + clip context; approval offered on significant edits. Never fabricate tiers, credits, or metrics — verify in-app. (Full scope: references/scope-and-connections.md.)

Distinct from its siblings (route correctly)

descript (this) = text-based talk-content editing · capcut = beat-synced visual short-form (the hybrid: master here, style cuts there) · captions-and-clipping / opus-clip = clip selection at scale (this feeds them the master) · podcast-and-audiograms = the strategy this tool serves · ai-voiceover (elevenlabs) = dedicated TTS/narration (Overdub = own-voice corrections) · talking-head-and-piece-to-camera = the performance (Eye Contact patches a read, doesn't replace delivery) · youtube-long-form = the structure the recording follows.

Where this connects

Reads first: podcast-and-audiograms / youtube-long-form + brand-profile/voice-builder + design-and-templates. Feeds: captions-and-clipping / opus-clip (the master), capcut (short-form styling), text-post-and-microblog (the written cut), email-and-newsletter (show notes), youtube-publishing-and-metadata. Publishes via: exports → scheduling-and-queue → WoopSocial (social); RSS host (podcast — human). Tool file: tools/integrations/descript.md. Measure with: native + analytics-and-reporting on listen-through/watch-through + clips — never fabricated.

Definition of done

A talk-content edit made at the speed of text and verified by ear: transcript corrected first (names/jargon — errors become wrong edits and captions), tangents/bad takes cut and filler+silences removed in batched Underlord passes, restructure done in text and the result listened through; flubs fixed with consent-verified own-voice Overdub at word/phrase length only (paragraphs re-recorded; synthetic speech disclosed where required); Studio Sound run once per source; multicam/captions/layout dressed on-brand; the interview rule held (meaning + clip context preserved, approval offered, the human owning the final cut of anyone else's words); one corrected transcript shipped as the episode + human-picked clips + show notes + chapters, routed onward and published via WoopSocial; the workflow credit-aware under the post-Sept-2025 model (import only what you'll edit; verify tiers in-app); API/MCP automation only where actually connected; no cloned third-party voices, no meaning-flips, no fabricated tiers/credits/metrics; and correctly distinguished from capcut, captions-and-clipping/opus-clip, ai-voiceover, and podcast-and-audiograms.

© social-media-skills, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/descript of social-media-skills/skills.

  • SKILL.md
  • evals/evals.json
  • references/descript-2026-reality.md
  • references/scope-and-connections.md
  • references/the-words-framework.md
  • references/workflows-and-templates.md

Open the folder on GitHubat commit 6e30eeb

Compare with similar skills

Descript next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Descript compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Descript this skillsocial-media-skills/skills134—~2kAutomated safety check: PassMIT
Podcastzarazhangrui/personalized-podcast438—~2.3kAutomated safety check: NotesNone
Z Qwen Audio Studiotjxj/z-skills548—~716Automated safety check: PassNone
Podcastteam-attention/plugins-for-claude-natives827—~1.5kAutomated safety check: PassMIT
Audio Duckingsonilo-ai/skills115—~1.8kAutomated safety check: NotesMIT
Auto Dubbingsonilo-ai/skills115—~4.5kAutomated safety check: NotesMIT

Similar skills

  • Podcast

    zarazhangrui/personalized-podcast

    Generate a podcast episode from content you provide. An agent skill from zarazhangrui/personalized-podcast.

    438 GitHub stars~2.3k tokensUpdated 6 mo ago
    Media & CreativeAuto-check: notes
  • Z Qwen Audio Studio

    tjxj/z-skills

    A skill your agent uses when creating complete generated audio with qwen-audio-3.1-tts-next, including podcasts, radio drama, advertisements, multiple speakers, reference voices, ambience, sound…

    548 GitHub stars~716 tokensUpdated 19 days ago
    Media & CreativeAuto-check passed
  • Podcast

    team-attention/plugins-for-claude-natives

    Generate Korean podcast episodes from any source (URLs, tweets, articles, PDFs) — analyzes content, writes a script, generates audio via OpenAI TTS, converts to MP4, and auto-uploads to YouTube.

    827 GitHub stars~1.5k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Audio Ducking

    sonilo-ai/skills

    Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps.

    115 GitHub stars~1.8k tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • Auto Dubbing

    sonilo-ai/skills

    Dub a video into one or more other languages using Sonilo, translating and re-voicing the speech into a new video per language.

    115 GitHub stars~4.5k tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • Proofread

    sonilo-ai/skills

    Transcribe a video with Sonilo and translate the transcript into editable .srt files — one per target language, plus the detected source language — so the wording can be read and corrected before…

    115 GitHub stars~4.4k tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes

More from social-media-skills/skills

All 106 skills in this repo
  • AI Image Editing

    social-media-skills/skills

    The AI image-editing router — inpainting/object removal, background removal, upscaling, outpainting, old-photo restoration, and retouch, routed task-first to the right engine.

    134 GitHub stars~2.1k tokensUpdated 9 days ago
    Auto-check passed
  • AI Music And Sound

    social-media-skills/skills

    The AI music + sound-design skill for social -- original/licensed audio beds and sound design for Reels/TikToks/Shorts/videos.

    134 GitHub stars~1.9k tokensUpdated 9 days ago
    Auto-check passed
  • AI Search Optimization

    social-media-skills/skills

    A skill your agent uses to get a brand and its content CITED and RECOMMENDED by AI answer engines — the GEO (Generative Engine Optimization) / AI-search-visibility skill.

    134 GitHub stars~2k tokensUpdated 9 days ago
    Auto-check passed
  • AI Video

    social-media-skills/skills

    The model-agnostic AI-video router and brief — the counterpart to image-prompt.

    134 GitHub stars~1.3k tokensUpdated 9 days ago
    Auto-check passed
  • AI Voiceover

    social-media-skills/skills

    The AI narration / voiceover mini-skill (ElevenLabs-led). An agent skill from social-media-skills/skills.

    134 GitHub stars~1.1k tokensUpdated 9 days ago
    Auto-check passed
  • Analytics And Reporting

    social-media-skills/skills

    Social media analytics and reporting — read native platform data honestly and turn it into next actions.

    134 GitHub stars~1.4k tokensUpdated 9 days ago
    Auto-check passed

Questions about Descript

What does Descript do?

The Descript craft skill — edit talk content (podcasts, interviews, talking-head video) by editing the transcript instead of the timeline. Descript is an agent skill from social-media-skills/skills. The Descript craft skill — edit talk content (podcasts, interviews, talking-head video) by editing the transcript instead of the timeline.

When should I use Descript?

Descript fits situations like: someone wants to edit in Descript; remove filler words/silences; clean up audio (Studio Sound); fix a flubbed word without re-recording (Overdub).

How do I install Descript in Claude Code?

Run `npx skills add social-media-skills/skills --skill descript -a claude-code`. Or copy the skill folder (skills/descript in social-media-skills/skills) into .claude/skills/descript in your project. Claude Code loads it when a task matches its description.

How do I install Descript in Codex?

Run `npx skills add social-media-skills/skills --skill descript -a codex`. Or copy the skill folder (skills/descript in social-media-skills/skills) into .agents/skills/descript in your project. Codex loads it when a task matches its description.

Can I use Descript in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add social-media-skills/skills --skill descript -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/descript, .gemini/skills/descript, .github/skills/descript and .opencode/skills/descript in your project.

What does Descript need to run?

SKILL.md names no scripts, command-line tools or credentials: Descript is instructions for the agent only.

Does Descript access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Descript safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Descript use?

Descript is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Descript use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.2k tokens, read only when the agent opens those files.

What are the alternatives to Descript?

Skills that share tags, products or a category with Descript: Podcast (zarazhangrui/personalized-podcast, 438 stars), Z Qwen Audio Studio (tjxj/z-skills, 548 stars), Podcast (team-attention/plugins-for-claude-natives, 827 stars) and Audio Ducking (sonilo-ai/skills, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Descript?

social-media-skills (a GitHub organization) maintains it in social-media-skills/skills, which has 134 GitHub stars. The repository holds 106 skills in this directory. The repository was last updated on October 1, 2026.

Source: social-media-skills/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.