Agent skill

Media Transform

by swyxio in swyxio/skills

Generic media transformation orchestrator — download videos from any source (X/Twitter, Zoom, YouTube, web embeds), upload to YouTube, transcribe with timestamps, generate chapters, create…

MITAuto-check passedMedia & Creative

Install Media Transform

skills CLI
$ npx skills add swyxio/skills --skill media-transform -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swyxio/skills media-transform --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/media-transform .claude/skills/media-transform && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
media-transform
GitHub stars
175
Token cost
~2.6k tokens
SKILL.md length
975 words
Files
1
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Generic media transformation orchestrator — download videos from any source (X/Twitter, Zoom, YouTube, web embeds), upload to YouTube, transcribe with timestamps, generate chapters, create…

  • Works in 4 steps: Stage selection — which steps to run… → Preferences — battle-tested defaults and… → Learnings — what worked, what didn't,… → …
  • The user wants to move a video from one platform to another
  • SKILL.md covers Architecture, Atomic Skills, Pipelines and Title Generation, plus 4 more sections
  • Calls python3, yt-dlp and brew; reaches x.com

What it does

Media Transform is an agent skill from swyxio/skills. Generic media transformation orchestrator — download videos from any source (X/Twitter, Zoom, YouTube, web embeds), upload to YouTube, transcribe with timestamps, generate chapters, create thumbnails with GPT-Image-2, and A/B test titles. Use when the user wants to move a video from one platform to another, or asks to "download and upload this video to YouTube", "publish this recording", "save and transcribe this", or any video pipeline task. Encodes learned best practices and preferences for each stage.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription, A/B testing and Image generation. It works with YouTube and X (Twitter). The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.

When your agent uses it

  • The user wants to move a video from one platform to another
  • Asks to download and upload this video to YouTube
  • Publish this recording
  • Save and transcribe this

Example prompts

  • “download and upload this video to YouTube”
  • “publish this recording”
  • “save and transcribe this”
  • “/media-transform”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Stage selection — which steps to run based on source/destination
  2. Preferences — battle-tested defaults and known gotchas
  3. Learnings — what worked, what didn't, what to avoid
  4. Checkpoints — present plan, get confirmation, then execute

What it can do on your machine

Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • yt-dlp
    • brew
    • gcloud
    • pipx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • x.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Media Transform loads about 2.6k tokens when it runs. Until then it costs about 131 tokens; SKILL.md has 975 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~131
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 975 words, ~2,589 tokens.

Download SKILL.mdSave it as .claude/skills/media-transform/SKILL.md (or your agent's skills folder).
name
media-transform
description
Generic media transformation orchestrator — download videos from any source (X/Twitter, Zoom, YouTube, web embeds), upload to YouTube, transcribe with timestamps, generate chapters, create thumbnails with GPT-Image-2, and A/B test titles. Use when the user wants to move a video from one platform to another, or asks to "download and upload this video to YouTube", "publish this recording", "save and transcribe this", or any video pipeline task. Encodes learned best practices and preferences for each stage.

Media Transform

Generic orchestrator for media transformation pipelines. Chains atomic skills based on source and destination, with stage-by-stage checkpoints.

Architecture

Each pipeline stage is handled by a dedicated atomic skill. This orchestrator provides:

  1. Stage selection — which steps to run based on source/destination
  2. Preferences — battle-tested defaults and known gotchas
  3. Learnings — what worked, what didn't, what to avoid
  4. Checkpoints — present plan, get confirmation, then execute

Atomic Skills

StageSkillNotes
Download (X/Twitter)download-x-videoyt-dlp, --print after_move:filepath
Download (Zoom)zoom-downloadBrowser-based, gallery view preferred
Download (web embeds)download-videoHandles Vimeo, YouTube embeds, referer headers
Download (generic URL)yt-dlp directlybrew install yt-dlp
Upload to YouTubeyoutube-apiOAuth, resumable upload, tags, metadata
Update metadatayoutube-apiupdate_metadata.py — title, description, tags
Set thumbnailyoutube-apiset_thumbnail.py — upload custom thumbnail
Transcribetranscribe-anythingMulti-backend, auto-selects best
Chapters (LLM titles)podcast-publishing-assistantHigh-quality chapter summaries

Pipelines

Pipeline A: X/Twitter → YouTube + Chapters
download-x-video → youtube-api (upload) → transcribe-anything → youtube-api (update description)
bash
# 1. Download
python3 download-x-video/scripts/download_x_video.py "https://x.com/user/status/123/video/1" /tmp

# 2. Upload (unlisted)
python3 youtube-api/scripts/upload_video.py \
    --file /tmp/x_video_<id>.mp4 \
    --title "Video Title" \
    --privacy unlisted

# 3. Transcribe (prefer mlx_whisper on Apple Silicon)
mlx_whisper /tmp/x_video_<id>.mp4 \
    --model mlx-community/whisper-turbo \
    --output-dir /tmp --output-format json \
    --word-timestamps True

# 4. Generate chapters + update description
# See Chapter Generation section below
Pipeline B: Zoom → YouTube + Thumbnails
zoom-download → youtube-api (upload + metadata + thumbnail)

Zoom recordings typically have built-in transcripts. Focus on proper titling, playlist assignment, and thumbnails.

Pipeline C: Generic Video → YouTube + Transcription
yt-dlp download → youtube-api (upload) → transcribe-anything

For any video URL that yt-dlp supports (YouTube, Vimeo, etc.), download and re-publish.

Title Generation

Generate 3-5 title candidates using the LLM. Evaluate against these heuristics:

What makes a good YouTube title:

  • Curiosity gap: Implies something the viewer doesn't know yet
  • Specificity: Names, numbers, concrete claims beat vague ones
  • Pattern interrupt: Unexpected framing or contradiction
  • Under 70 chars: Avoids truncation in search results
  • Front-load keywords: Most important words first
  • No clickbait: Title must match content (retention matters more than CTR)

Title generation prompt template:

Generate 5 YouTube title candidates for a video about [topic].
The video is [duration] and [brief content description].

Requirements:
- Under 70 characters each
- Different angles: (1) curiosity-driven, (2) how-to/value, (3) controversial/contrarian,
  (4) specific/numbers-driven, (5) question-based
- No ALL CAPS, no emoji overuse
- Titles must accurately reflect the content
A/B Testing Titles

YouTube Studio has native "Test & Compare" (tests up to 3 titles/thumbnails, runs up to 2 weeks, winner based on watch time share). This is NOT available via the YouTube Data API directly.

Programmatic DIY A/B testing:

Use youtube-api/scripts/update_metadata.py to rotate titles on a schedule, then analyze performance via YouTube Analytics:

bash
# Start test: set title A
python3 youtube-api/scripts/update_metadata.py --video-id <ID> --title "Title A"

# After 24-48h: rotate to title B
python3 youtube-api/scripts/update_metadata.py --video-id <ID> --title "Title B"

# After 24-48h more: check analytics to determine winner
# Winner = higher CTR * average view duration (or just CTR for early tests)

A/B testing schedule:

  • Rotate every 24-48 hours (YouTube needs time to collect impressions)
  • Test 2-3 titles per video
  • Run for 1-2 weeks total
  • Winner based on: CTR (click-through rate) × retention, not just view count

Thumbnail Generation

GPT-Image-2 (openai/gpt-image-2) via the image_generate tool is the preferred thumbnail generator:

Key capabilities relevant to thumbnails:

  • Near-perfect text rendering: Can include readable text on thumbnails (previously impossible with AI)
  • Thinking mode: Plans composition before rendering — ensures faces, text, and layout are coherent
  • Up to 2K resolution: 2048px, perfect for 1280×720 thumbs with room to crop
  • Aspect ratio 16:9: Native YouTube thumbnail ratio
  • Multilingual text: Works across scripts (Latin, CJK, etc.)
  • Multi-variant generation: Up to 4-8 coherent variations from one prompt

Thumbnail prompt template:

YouTube thumbnail for a video titled "[TITLE]". Style: [clean/bold/minimalist/tech].
[Specific visual elements: faces, diagrams, text overlays].
Aspect ratio: 16:9. High contrast, eye-catching. No clutter. 
Text on image (if any): "[KEY PHRASE]" in [position].

Post-generation:

  • Use youtube-api/scripts/set_thumbnail.py to upload
  • Compress if >2MB: convert -resize 1280x720 -quality 85 input.png output.jpg
Thumbnail A/B Testing

YouTube's native "Test & Compare" supports up to 3 thumbnails. Generate 3 distinct concepts:

  1. Text-heavy: Key phrase or number in large font
  2. Face/emotion: Expressive reaction, eye contact
  3. Concept/abstract: Visual metaphor for the topic

Stage-by-Stage Preferences & Learnings

Download

yt-dlp path detection:

  • Use --print after_move:filepath for reliable final path (don't parse stdout for [download] Destination)
  • HLS streams from X/Twitter use fragmented filenames during download; only the after_move path is the final merged file

X/Twitter auth:

  • Some videos require authentication: yt-dlp --cookies-from-browser chrome
Upload

OAuth token caching:

  • youtube-api skill handles this: ~/.config/youtube-api/token.pickle (or Cowork path)
  • First run opens browser for consent; cached for subsequent runs
  • On Mac → local config; in Cowork VM → mounted Downloads folder (persists across resets)

Privacy default:

  • Always default to unlisted unless user explicitly asks for public

Resumable uploads:

  • Google API client supports resumable uploads — large files (100MB+) upload smoothly
Show full SKILL.md (388 more words)Show less
Transcription

Prefer mlx_whisper on Apple Silicon (10x faster):

  • mlx_whisper (pipx install mlx-whisper): ~1300 frames/s → ~2 min for 27 min audio
  • openai-whisper CLI: ~95 frames/s → ~28 min for same audio
  • openai-whisper with --device mps produces NaN errors with turbo/large models — avoid, use mlx_whisper instead

Turbo model is the sweet spot:

  • Fast enough for real-time use
  • Quality nearly as good as large
  • Small is too inaccurate for chapter generation

Diarization is aspirational:

  • Requires whisperX + pyannote + HuggingFace token
  • Adds 5-10 min processing
  • Quality varies with audio clarity
  • Use transcribe-anything with --diarize flag when available
Chapter Generation

Garbage filtering is essential:

  • Filter out pure filler segments: "Yeah.", "Cool.", "Mm-hmm.", "Right."
  • Filter repetitive filler: "Yeah. Yeah. Yeah." (3+ garbage words in a row)
  • Null segments (empty text, zero duration) at the end are common

Word-boundary truncation:

  • Don't truncate chapter titles mid-word
  • "What areas of data do you feel are underserved by now that l" → truncate at last space

LLM titles when quality matters:

  • Raw transcript chapters are functional but ugly
  • For polished output, use podcast-publishing-assistant or feed segments to an LLM
  • Prompt: "Generate concise chapter titles (<60 chars) for these transcript segments with timestamps"

Interval tuning:

  • Default 30s gives ~46 chapters for 27 min video — good for navigation
  • 60s gives ~27 chapters — cleaner but less granular
  • 10s is too granular for YouTube (chapter limit is ~100)

Checkpoint Pattern

Before each action phase, present a summary and get confirmation. This catches mismatches early:

  1. Pre-flight: Scan source (tweet, Zoom recordings, etc.) → list what's available
  2. Title check: Present 3-5 title candidates, user picks
  3. Thumbnail check: Generate 3 thumbnail variants, user picks
  4. Download complete: Confirm file, title, duration
  5. Upload complete: Confirm URL, privacy, playlist
  6. Transcription complete: Confirm segment count, quality
  7. Final: Present all results, offer title A/B test setup

Troubleshooting

YouTube API not enabled
bash
gcloud services enable youtube.googleapis.com --project=<PROJECT_ID>
OAuth redirect fails (ERR_CONNECTION_REFUSED)
  • Ensure port is free: lsof -i :8080
  • GCP OAuth must have http://localhost in redirect URIs
mlx_whisper "Failed to load audio"
  • brew install ffmpeg
Chapter quality is poor
  • Raw transcript chapters work for quick navigation but look unprofessional
  • For publication-quality, use podcast-publishing-assistant or LLM post-processing
  • Garbage filtering catches most bad chapters but may miss edge cases ("I mean", "you know")
Thumbnail too large
  • YouTube max is 2MB. Compress: convert -resize 1280x720 -quality 85 input.png output.jpg
  • GPT-Image-2 outputs may need compression for multi-variant uploads

© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in media-transform of swyxio/skills.

Open the folder on GitHubat commit 038ef34

Compare with similar skills

Media Transform next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Media Transform compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Media Transform this skillswyxio/skills175—~2.6kAutomated safety check: PassMIT
Yt Dlp DownloaderMapleShaw/yt-dlp-downloader-skill2091 repos~1.5kAutomated safety check: PassNone
Videohub Youtubecacity/VideoHub168—~550Automated safety check: PassMIT
Watchmathiaschu/watch142—~4kAutomated safety check: WarnMIT
Generate Youtube Thumbnailkrusemediallc/arcads-claude-code1.6k—~2.4kAutomated safety check: NotesMIT
Youtube Thumbnailhassancs91/claude-youtube-editor325—~2.1kAutomated safety check: NotesMIT

Similar skills

  • Yt Dlp Downloader

    MapleShaw/yt-dlp-downloader-skill

    Download videos from YouTube, Bilibili, Twitter, and thousands of other sites using yt-dlp.

    209 GitHub starsUsed in 1 repo~1.5k tokens
    Media & CreativeAuto-check passed
  • Videohub Youtube

    cacity/VideoHub

    处理 YouTube、Twitter(X)、Bilibili 和本地音视频/文本的转写、字幕、翻译与总结。优先复用 src/youtubetranscriber.py 现有 CLI。

    168 GitHub stars~550 tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Watch

    mathiaschu/watch

    Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

    142 GitHub stars~4k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: warnings
  • Generate Youtube Thumbnail

    krusemediallc/arcads-claude-code

    Generate high-CTR YouTube thumbnails using Nano Banana 2 via the Arcads external API.

    1.6k GitHub stars~2.4k tokensUpdated 16 days ago
    Media & CreativeAuto-check: notes
  • Youtube Thumbnail

    hassancs91/claude-youtube-editor

    Dedicated YouTube thumbnail generator — interviews you for exactly the style elements you want (environment, text budget, extras, accent color), then renders high-contrast, vibrant, face-consistent…

    325 GitHub stars~2.1k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Watchless

    chenzixin1/watchless

    A skill your agent uses when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML…

    144 GitHub stars~3.7k tokensUpdated 23 days ago
    Media & CreativeAuto-check: warnings

More from swyxio/skills

All 89 skills in this repo
  • Programmatic Agents

    swyxio/skills

    Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.

    175 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed
  • Design, implement, audit, or refresh protected username and handle namespaces for public products.

    175 GitHub stars~1.1k tokensUpdated 4 days ago
    Auto-check passed
  • New Mac Setup

    swyxio/skills

    Fully automated new Mac setup for fullstack web developers and AI engineers.

    175 GitHub stars~4.3k tokensUpdated 4 days ago
    Auto-check passed
  • Youtube API

    swyxio/skills

    Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…

    175 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed
  • Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.

    175 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check: warnings
  • Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

    175 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed

Questions about Media Transform

What does Media Transform do?

Generic media transformation orchestrator — download videos from any source (X/Twitter, Zoom, YouTube, web embeds), upload to YouTube, transcribe with timestamps, generate chapters, create…. Media Transform is an agent skill from swyxio/skills. Generic media transformation orchestrator — download videos from any source (X/Twitter, Zoom, YouTube, web embeds), upload to YouTube, transcribe with timestamps, generate chapters, create thumbnails with GPT-Image-2, and A/B test titles.

When should I use Media Transform?

Media Transform fits situations like: the user wants to move a video from one platform to another; asks to download and upload this video to YouTube; publish this recording; save and transcribe this.

How do I install Media Transform in Claude Code?

Run `npx skills add swyxio/skills --skill media-transform -a claude-code`. Or copy the skill folder (media-transform in swyxio/skills) into .claude/skills/media-transform in your project. Claude Code loads it when a task matches its description.

How do I install Media Transform in Codex?

Run `npx skills add swyxio/skills --skill media-transform -a codex`. Or copy the skill folder (media-transform in swyxio/skills) into .agents/skills/media-transform in your project. Codex loads it when a task matches its description.

Can I use Media Transform in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill media-transform -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/media-transform, .gemini/skills/media-transform, .github/skills/media-transform and .opencode/skills/media-transform in your project.

What does Media Transform need to run?

Going by SKILL.md and its folder, Media Transform needs the command-line tools its instructions call (python3, yt-dlp, brew, gcloud and pipx). Our summary lists: Python 3.

Does Media Transform access the network?

SKILL.md names 1 domain. In commands or code: x.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Media Transform safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Media Transform use?

Media Transform is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Media Transform use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Media Transform?

Skills that share tags, products or a category with Media Transform: Yt Dlp Downloader (MapleShaw/yt-dlp-downloader-skill, 209 stars), Videohub Youtube (cacity/VideoHub, 168 stars), Watch (mathiaschu/watch, 142 stars) and Generate Youtube Thumbnail (krusemediallc/arcads-claude-code, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Media Transform?

swyxio (a GitHub user) maintains it in swyxio/skills, which has 175 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.

Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.