Agent skill

Podcast Transcribe

by chubbyguan in chubbyguan/chubbyskills

播客/小宇宙 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills.

MITAuto-check: notesMedia & Creative

Install Podcast Transcribe

skills CLI
$ npx skills add chubbyguan/chubbyskills --skill podcast-transcribe -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chubbyguan/chubbyskills podcast-transcribe --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/chubbyguan/chubbyskills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/podcast-transcribe .claude/skills/podcast-transcribe && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
podcast-transcribe
GitHub stars
1.2k
Token cost
~1.1k tokens
SKILL.md length
199 words
Files
8 (incl. scripts)
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

播客/小宇宙 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills.

  • Tasks that involve Transcription
  • SKILL.md covers 环境要求, 单集与批量, 本地模型选择:SenseVoice-Small vs… and 可选云端转录, plus 4 more sections
  • Runs Python scripts from its folder; calls python3 and pip; reaches xiaoyuzhoufm.com; needs DASHSCOPE_API_KEY and GROQ_API_KEY
  • Tasks that involve Speech recognition and synthesis

What it does

Podcast Transcribe is an agent skill from chubbyguan/chubbyskills. 播客/小宇宙 → 下载 → 转录 → 存为 Markdown 的完整工作流。 支持 RSS 批量下载、单集链接转录。

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts (for example `scripts/batch_transcribe.py`, `scripts/cloud_transcribe.py` and `scripts/local_qwen_asr.py`).

It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. It works with Qwen and FFmpeg. The repository describes itself as: 把中文全渠道内容(抖音 / B站 / 小红书 / 公众号 / X / 播客)采集进个人知识库的 14 个 AI Skill:图文存图、视频转文字稿、字幕优先免 GPU、RSS/YouTube 订阅调度与每日情报简报,附带知识库 MCP server。| Ingest Chinese content into your personal knowledge… The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/podcast-transcribe”

Requirements

  • Python 3
  • A credential in DASHSCOPE_API_KEY
  • A credential in GROQ_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 1b759b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • xiaoyuzhoufm.com

    Also links to:

    • github.com
    • help.aliyun.com
    • console.groq.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DASHSCOPE_API_KEY
    • GROQ_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Podcast Transcribe loads about 1.1k tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 199 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:28
    # Ubuntu: sudo apt install ffmpeg

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from chubbyguan/chubbyskills at commit 1b759b1, republished under its MIT licence (© chubbyguan). 199 words, ~1,122 tokens.

Download SKILL.mdSave it as .claude/skills/podcast-transcribe/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
podcast-transcribe
description
播客/小宇宙 → 下载 → 转录 → 存为 Markdown 的完整工作流。 支持 RSS 批量下载、单集链接转录。
metadata.category
media
metadata.triggers
["用户发送小宇宙/播客链接", "帮我转录这个播客", "下载播客", "批量转录播客"]
metadata.version
1.1.0
metadata.tags
["media", "audio", "podcast", "transcription", "xiaoyuzhou"]

播客转录 Skill

将播客音频下载并转录为 Markdown。支持小宇宙、喜马拉雅、直接音频地址、本地音频和 RSS 批量流程。

默认在本地使用 SenseVoice-Small(与视频类技能共用的 chubby_common/funasr.py 封装)。可选云端后端为阿里云百炼 DashScope 的 qwen3-asr-flash 和 Groq 的 whisper-large-v3-turbo,用户明确选择后才启用;音频会发送至云端并可能计费。

环境要求

以下命令在本 skill 目录运行,建议使用 Python 3.11 或更新版本:

bash
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -r requirements.txt
# macOS: brew install ffmpeg
# Ubuntu: sudo apt install ffmpeg

本地模型首次使用时需要下载。仅使用云端后端不需要 funasr 本地依赖;不支持的音频容器转换仍可能需要 ffmpeg。云端密钥通过安全环境配置,不能写入命令参数、转录稿或版本库。

单集与批量

bash
python3 scripts/transcribe.py "https://www.xiaoyuzhoufm.com/episode/xxxxx" ./output \
  --provider local
python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output \
  --provider local --language zh
python3 scripts/batch_transcribe.py --rss-url "替换为实际 RSS 地址" \
  --output ./output --count 10 --provider local

将示例地址替换为实际来源。单集页面会尝试提取音频链接,提取失败时改用直接音频地址或本地文件。标准输出最后一行是成功生成的 Markdown 路径。

自动下载仅接受公网 HTTP(S) 直连地址,禁用代理和重定向,拒绝本地/私网地址。需要跳转或代理的来源,请先自行下载音频,再传入本地文件路径。

本地模型选择:SenseVoice-Small vs Qwen3-ASR-0.6B

--provider local 支持两个本地模型:SenseVoiceSmall(默认)和 qwen3-asr-0.6b:

bash
python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output --provider local
python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output --provider local --model qwen3-asr-0.6b

qwen3-asr-0.6b 是可选重依赖,不在默认安装内,需自行 pip install qwen-asr transformers torch。

以下对比数据来自 2026-10 在 MacBook Pro(Apple M3 Pro,纯 CPU)上对 5 分钟中文播客的实测:

SenseVoice-Small(默认)Qwen3-ASR-0.6B(可选)
速度RTF 0.11(5 分钟音频约 33 秒)RTF 0.80(约 4 分钟,慢约 7 倍)
专有名词一般("岩茶"误作"盐茶",人名前后不一致)更稳("岩茶"、人名识别一致)
主要风险输出混入情感标签,正式文稿需清洗有幻觉式改写风险("黄金加工厂"→"皇帝家族");无 ITN,数字输出为全文字
长音频VAD 自动分段,稳定整段进模型;本后端已把 max_new_tokens 调到 4096 避免截断
适用场景CPU 默认选择,长播客友好GPU 机器,或对专名/人名准确性敏感的内容

两者质量互有胜负、没有代差。CPU 场景请保持默认 SenseVoice-Small;有 GPU 或专名敏感时再选 Qwen3-ASR-0.6B。

可选云端转录

后端凭据环境变量默认模型
local无SenseVoiceSmall(仅用于元数据记录)
dashscopeDASHSCOPE_API_KEYqwen3-asr-flash
groqGROQ_API_KEYwhisper-large-v3-turbo

配置 DASHSCOPE_API_KEY 或 GROQ_API_KEY 后:

bash
python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output \
  --provider dashscope --cloud-timeout 1800 \
  --state-dir "$HOME/.local/state/chubbyskills/podcast"
python3 scripts/transcribe.py "/你的音频目录/episode.mp3" ./output \
  --provider groq --cloud-timeout 1800 \
  --state-dir "$HOME/.local/state/chubbyskills/podcast"

--provider 优先于 PODCAST_TRANSCRIBE_PROVIDER,都未指定时使用 local。批量入口支持同样的 provider、模型、语言、等待和状态目录参数,并把最终选项显式传给单集进程。

两个云端后端都是同步接口。DashScope 限制为编码后不超过 10MB、时长不超过 5 分钟(客户端在原始文件超过 7 MiB 时拒绝并提示改用本地 SenseVoice)。Groq 免费层单文件上限 25MB,达到上限的长音频会自动分片:ffmpeg 切成 20 分钟一段(约 9.6MB/段),逐段转录后按顺序拼接,时间戳自动累加偏移;分片进度逐段落盘,中断后重跑同一命令断点续传,不重复提交已完成分片;遇 429 按 Retry-After/指数退避等待。Groq 返回的分段时间戳会作为附录保留。分片和容器转换需要本机 ffmpeg。

云端任务恢复

默认状态目录为 ~/.local/state/chubbyskills/podcast;设置了 XDG_STATE_HOME 时使用其中的 chubbyskills/podcast。还可用 CHUBBY_PODCAST_STATE_DIR 或 --state-dir 指定。目录包含任务与转录内容,应按音频内容的隐私要求保存。

相同音频、provider、模型、语言和服务地址再次运行时,已有任务继续查询;已完成结果可以重新导出。提交结果不明确或服务端报告失败时,不自动重新提交。普通网络重试保留状态目录并重跑原命令即可。

--resubmit 明确创建新任务,可能重复计费。不要用它解决单纯的轮询超时。删除状态目录或改变输入配置也可能失去复用条件。客户端超时不等于服务端取消,也不代表没有计费。

完整仓库使用说明、统一入库和验证范围见云端转录说明。

产物与限制

产物为带 frontmatter、来源和转录后端标记的 Markdown。后端返回可用分段时保留时间戳;没有时间信息时不编造时间轴。

  • 本地 SenseVoice-Small 只输出纯文本全文,不生成逐段时间戳(段数记为 1)。
  • 本地 CPU 推理耗时受音频长度和机器配置影响。
  • 转录可能有专有名词、数字或断句错误,引用前核对原始音频。
  • 云端真实可用性、账号权限、音频兼容性和费用尚需独立验收。
  • 本 skill 不提供通用说话人分离保证。

贡献与参考

可选云端转录需求最初来自 binyangzhu000-sudo 的 PR #3 和 Anil-matcha 的 PR #5(Atlas / MuAPI 实验后端,现已被 DashScope 后端取代,归属保留)。

合规声明

请遵守来源平台条款并尊重内容版权,控制请求频率。云端处理前确认自己有权向所选服务提交音频;下载和转录不会改变原内容的版权归属。

© chubbyguan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts) in podcast-transcribe of chubbyguan/chubbyskills.

  • SKILL.md
  • requirements.txt
  • scripts/batch_transcribe.py
  • scripts/cloud_transcribe.py
  • scripts/local_qwen_asr.py
  • scripts/provider_config.py
  • scripts/safe_download.py
  • scripts/transcribe.py

Open the folder on GitHubat commit 1b759b1

Compare with similar skills

Podcast Transcribe next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Podcast Transcribe compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Podcast Transcribe this skillchubbyguan/chubbyskills1.2k—~1.1kAutomated safety check: NotesMIT
Claude Real VideoHUANGCHIHHUNGLeo/claude-real-video2.2k—~639Automated safety check: PassMIT
Claude Real Video For AgentsHUANGCHIHHUNGLeo/claude-real-video2.2k—~2kAutomated safety check: NotesMIT
Video Clip Extractorlinzzzzzz/openclip569—~2.8kAutomated safety check: WarnMIT
Summarize Callreysu/ai-life-skills270—~3.8kAutomated safety check: NotesMIT
Watchmathiaschu/watch142—~4kAutomated safety check: WarnMIT

Similar skills

  • Claude Real Video

    HUANGCHIHHUNGLeo/claude-real-video

    Watch a video for the user. An agent skill from HUANGCHIHHUNGLeo/claude-real-video.

    2.2k GitHub stars~639 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Claude Real Video For Agents

    HUANGCHIHHUNGLeo/claude-real-video

    Install and use crv (claude-real-video) — a tool that lets any AI agent watch videos by extracting scene-aware keyframes, deduplicating them, and transcribing audio.

    2.2k GitHub stars~2k tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • Video Clip Extractor

    linzzzzzz/openclip

    Processes videos to identify engaging moments, generate transcripts, and create highlight clips with artistic titles and custom cover images.

    569 GitHub stars~2.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: warnings
  • Summarize Call

    reysu/ai-life-skills

    Transcribe a call recording with speaker diarization, summarize it, and create Obsidian vault notes (call note, transcript, person notes for participants).

    270 GitHub stars~3.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Watch

    mathiaschu/watch

    Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

    142 GitHub stars~4k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: warnings
  • Bggg Tiktok Readvideo

    binggandata/bggg-skills

    把 TikTok、Reels、YouTube Shorts、UGC 广告、本地 MP4/MOV/WebM 等视频拆成 Codex 可读的视频上下文。

    605 GitHub stars~1.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from chubbyguan/chubbyskills

All 14 skills in this repo
  • Bilibili Transcribe

    chubbyguan/chubbyskills

    哔哩哔哩视频 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills.

    1.2k GitHub stars~578 tokensUpdated 3 days ago
    Auto-check: notes
  • Douyin Transcribe

    chubbyguan/chubbyskills

    抖音视频 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills.

    1.2k GitHub stars~551 tokensUpdated 3 days ago
    Auto-check: notes
  • Industry Intelligence Radar

    chubbyguan/chubbyskills

    行业情报雷达:多源扫描(X/即刻/V2EX/HN) → 关键词过滤 → 趋势检测 → 每日情报简报。触发词:行业情报、竞品监控、热点扫描、情报雷达

    1.2k GitHub stars~806 tokensUpdated 3 days ago
    Auto-check passed
  • Knowledge Base Management

    chubbyguan/chubbyskills

    管理本地 Markdown/Obsidian 知识库:素材入库、健康检查、增量索引、关键词与 semantic-lite 检索、逐字原文资料包、归档和 MCP 连接。用于搜索知识库、整理资料及为 Agent 准备可定位的引用。

    1.2k GitHub stars~3.3k tokensUpdated 3 days ago
    Auto-check passed
  • Learning Notes Automation

    chubbyguan/chubbyskills

    学习笔记自动化:视频/播客转录 → 知识点提取 → 闪卡生成 → 知识图谱更新。触发词:学习笔记、闪卡、Anki、知识提取、视频学习

    1.2k GitHub stars~898 tokensUpdated 3 days ago
    Auto-check passed
  • Tiktok Transcribe

    chubbyguan/chubbyskills

    TikTok 视频 → 下载 → 转录 → 存为 Markdown. An agent skill from chubbyguan/chubbyskills.

    1.2k GitHub stars~536 tokensUpdated 3 days ago
    Auto-check: notes

Works with

Questions about Podcast Transcribe

What does Podcast Transcribe do?

播客/小宇宙 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills. Podcast Transcribe is an agent skill from chubbyguan/chubbyskills.

When should I use Podcast Transcribe?

Podcast Transcribe fits situations like: tasks that involve Transcription; tasks that involve Speech recognition and synthesis.

How do I install Podcast Transcribe in Claude Code?

Run `npx skills add chubbyguan/chubbyskills --skill podcast-transcribe -a claude-code`. Or copy the skill folder (podcast-transcribe in chubbyguan/chubbyskills) into .claude/skills/podcast-transcribe in your project. Claude Code loads it when a task matches its description.

How do I install Podcast Transcribe in Codex?

Run `npx skills add chubbyguan/chubbyskills --skill podcast-transcribe -a codex`. Or copy the skill folder (podcast-transcribe in chubbyguan/chubbyskills) into .agents/skills/podcast-transcribe in your project. Codex loads it when a task matches its description.

Can I use Podcast Transcribe in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chubbyguan/chubbyskills --skill podcast-transcribe -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/podcast-transcribe, .gemini/skills/podcast-transcribe, .github/skills/podcast-transcribe and .opencode/skills/podcast-transcribe in your project.

What does Podcast Transcribe need to run?

Going by SKILL.md and its folder, Podcast Transcribe needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and pip) and credentials named DASHSCOPE_API_KEY and GROQ_API_KEY. Our summary lists: Python 3; A credential in DASHSCOPE_API_KEY; A credential in GROQ_API_KEY.

Does Podcast Transcribe access the network?

SKILL.md names 4 domains. In commands or code: xiaoyuzhoufm.com; the agent is likely to contact it when it follows the instructions. As links in the text: github.com, help.aliyun.com and console.groq.com. This is read from the text; nothing was executed.

Is Podcast Transcribe safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Podcast Transcribe use?

Podcast Transcribe is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Podcast Transcribe use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Podcast Transcribe?

Skills that share tags, products or a category with Podcast Transcribe: Claude Real Video (HUANGCHIHHUNGLeo/claude-real-video, 2.2k stars), Claude Real Video For Agents (HUANGCHIHHUNGLeo/claude-real-video, 2.2k stars), Video Clip Extractor (linzzzzzz/openclip, 569 stars) and Summarize Call (reysu/ai-life-skills, 270 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Podcast Transcribe?

chubbyguan (a GitHub user) maintains it in chubbyguan/chubbyskills, which has 1,217 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 8, 2026.

Source: chubbyguan/chubbyskills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.