Agent skill

Audio2txt

by CoderWanFeng in CoderWanFeng/python-office

从音频文件中提取文字(语音转文字),需要腾讯云 API 凭据。当用户提到音频转文字、语音识别、ASR、音频转字幕时使用。

Apache-2.0Auto-check passedAI & LLM Engineering

Install Audio2txt

skills CLI
$ npx skills add CoderWanFeng/python-office --skill audio2txt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install CoderWanFeng/python-office audio2txt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/CoderWanFeng/python-office.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video/audio2txt .claude/skills/audio2txt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audio2txt
GitHub stars
1.4k
Token cost
~230 tokens
SKILL.md length
47 words
Files
2
Skills in repo
39
Repo updated
First seen
Licence
Apache-2.0

At a glance

从音频文件中提取文字(语音转文字),需要腾讯云 API 凭据。当用户提到音频转文字、语音识别、ASR、音频转字幕时使用。

  • Tasks that involve Speech recognition and synthesis
  • SKILL.md covers 功能描述, 所属分类, 调用方式 and 参数说明, plus 3 more sections
  • Runs Python scripts from its folder

What it does

Audio2txt is an agent skill from CoderWanFeng/python-office. 从音频文件中提取文字(语音转文字),需要腾讯云 API 凭据。当用户提到音频转文字、语音识别、ASR、音频转字幕时使用。

Its SKILL.md is about 230 tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `__init__.py`).

It sits in AI & LLM Engineering, covering Speech recognition and synthesis. The repository describes itself as: pip install python-office 自动化办公专用库. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/audio2txt”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit dd6fe56. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio2txt loads about 230 tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 47 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~230

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from CoderWanFeng/python-office at commit dd6fe56, republished under its Apache-2.0 licence (© CoderWanFeng). 47 words, ~230 tokens.

Download SKILL.mdSave it as .claude/skills/audio2txt/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
audio2txt
description
从音频文件中提取文字(语音转文字),需要腾讯云 API 凭据。当用户提到音频转文字、语音识别、ASR、音频转字幕时使用。

audio2txt Skill

从音频里提取文字

功能描述

使用语音识别 API 从音频中提取文字。

⚠️ 注意:本地语音文件不能大于 5MB。

所属分类

office/skills/video/audio2txt/

调用方式

python
from skills.video import audio2txt

audio2txt(
    audio_path='./test.mp3',
    appid='your_appid',
    secret_id='your_secret_id',
    secret_key='your_secret_key'
)

参数说明

参数类型必填默认值说明
audio_pathstr是-音频文件路径
appidstr是-语音识别 API 的应用 ID
secret_idstr是-语音识别 API 的密钥 ID
secret_keystr是-语音识别 API 的密钥

返回值

None

使用示例

python
from skills.video import audio2txt
audio2txt(
    audio_path='./audio.mp3',
    appid='your_appid',
    secret_id='your_secret_id',
    secret_key='your_secret_key'
)

原始函数

office.api.video.audio2txt

© CoderWanFeng, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/video/audio2txt of CoderWanFeng/python-office.

  • SKILL.md
  • __init__.py

Open the folder on GitHubat commit dd6fe56

Compare with similar skills

Audio2txt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio2txt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio2txt this skillCoderWanFeng/python-office1.4k—~230Automated safety check: PassApache-2.0
TriageTalAter/annyang6.8k1 repos~810Automated safety check: NotesMIT
Yichen Asrmcncarl/yichen-skills4.3k—~780Automated safety check: PassCustom licence
Dingtalk MinutesDingTalk-Real-AI/dingtalk-workspace-cli3.2k—~2.3kAutomated safety check: PassApache-2.0
Youtube FetcherJimmySadek/youtube-fetcher-to-markdown485—~3.1kAutomated safety check: PassMIT
Yichen Web Researchmcncarl/yichen-skills4.3k—~1.9kAutomated safety check: PassCustom licence

Similar skills

  • Triage

    TalAter/annyang

    Triage and close GitHub issues on TalAter/annyang. An agent skill from TalAter/annyang.

    6.8k GitHub starsUsed in 1 repo~810 tokens
    AI & LLM EngineeringAuto-check: notes
  • Yichen Asr

    mcncarl/yichen-skills

    逸尘自用的统一音视频转写入口,在 StepFun Step ASR 与火山引擎豆包 ASR 之间按输出需求、安全边界和可用状态路由。用于本地音频或视频的纯文本转写、时间戳、SRT 字幕、口播粗剪,以及转写前体检;用户明确指定服务商时不得静默切换。Use when a local audio or video file needs transcription and the correct…

    4.3k GitHub stars~780 tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Dingtalk Minutes

    DingTalk-Real-AI/dingtalk-workspace-cli

    钉钉 AI 听记。Use when 查询或修改听记摘要、完整逐字稿、关键词、标签、行动项、录音、上传、思维导图、发言人洞察、ASR 热词/识别词配置或分享权限。写文档走 dingtalk-doc;建待办走 dingtalk-todo;日程走 dingtalk-calendar。命令前缀:dws minutes。

    3.2k GitHub stars~2.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Youtube Fetcher

    JimmySadek/youtube-fetcher-to-markdown

    Retrieve transcripts from YouTube, Instagram, TikTok, X, Vimeo and other video sites, summarize or analyze what was said (and shown on screen), or save an Obsidian-ready Markdown knowledge-base note…

    485 GitHub stars~3.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Yichen Web Research

    mcncarl/yichen-skills

    逸尘自用的互联网研究总入口。用于跨平台且跨阶段、用户尚未确定工具,或明确要求对公司、产品、人物、技术、行业和领域做横纵分析、发展史加现状对比或有来源约束的系统深度研究;先生成有截止日期和证据闸门的计划,再把搜索发现、候选核验、有限归档、按需转写和证据综合路由到…

    4.3k GitHub stars~1.9k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Volcengine Asr

    ysyecust/lecture-to-notes

    Transcribe local audio or video with Volcengine Doubao file ASR, including BigASR 1.0 Turbo direct upload and asynchronous 1.0 standard, 1.0 idle, or 2.0 standard jobs through TOS.

    270 GitHub stars~783 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

More from CoderWanFeng/python-office

All 39 skills in this repo
  • Decrypt4pdf

    CoderWanFeng/python-office

    对加密的 PDF 文件进行解密。当用户提到 PDF 解密、解除 PDF 密码、解 PDF 时使用. An agent skill from CoderWanFeng/python-office.

    1.4k GitHub stars~242 tokensUpdated 1 mo ago
    Auto-check passed
  • Del4pdf

    CoderWanFeng/python-office

    删除 PDF 文件中指定的页面。当用户提到 PDF 删除页、PDF 删页、删 PDF 页 时使用. An agent skill from CoderWanFeng/python-office.

    1.4k GitHub stars~209 tokensUpdated 1 mo ago
    Auto-check passed
  • Doc2docx

    CoderWanFeng/python-office

    将旧版 Word 文档(.doc)转换为新版(.docx)。当用户提到 doc 转 docx、旧 Word 转新 Word 时使用。

    1.4k GitHub stars~269 tokensUpdated 1 mo ago
    Auto-check passed
  • Docx2doc

    CoderWanFeng/python-office

    将新版 Word 文档(.docx)转换为旧版(.doc)。当用户提到 docx 转 doc、新 Word 转旧 Word 时使用。

    1.4k GitHub stars~196 tokensUpdated 1 mo ago
    Auto-check passed
  • Docx2pdf

    CoderWanFeng/python-office

    将 Word 文档(.docx)转换为 PDF,支持单文件和整个文件夹批量转换。当用户提到 Word 转 PDF、docx 转 PDF 时使用。

    1.4k GitHub stars~197 tokensUpdated 1 mo ago
    Auto-check passed
  • Docx4imgs

    CoderWanFeng/python-office

    从 Word 文档中提取所有图片并保存到指定目录。当用户提到 Word 提取图片、docx 抠图、Word 取图时使用。

    1.4k GitHub stars~185 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Audio2txt

What does Audio2txt do?

从音频文件中提取文字(语音转文字),需要腾讯云 API 凭据。当用户提到音频转文字、语音识别、ASR、音频转字幕时使用。. Audio2txt is an agent skill from CoderWanFeng/python-office.

When should I use Audio2txt?

Audio2txt fits situations like: tasks that involve Speech recognition and synthesis.

How do I install Audio2txt in Claude Code?

Run `npx skills add CoderWanFeng/python-office --skill audio2txt -a claude-code`. Or copy the skill folder (skills/video/audio2txt in CoderWanFeng/python-office) into .claude/skills/audio2txt in your project. Claude Code loads it when a task matches its description.

How do I install Audio2txt in Codex?

Run `npx skills add CoderWanFeng/python-office --skill audio2txt -a codex`. Or copy the skill folder (skills/video/audio2txt in CoderWanFeng/python-office) into .agents/skills/audio2txt in your project. Codex loads it when a task matches its description.

Can I use Audio2txt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add CoderWanFeng/python-office --skill audio2txt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio2txt, .gemini/skills/audio2txt, .github/skills/audio2txt and .opencode/skills/audio2txt in your project.

What does Audio2txt need to run?

Going by SKILL.md and its folder, Audio2txt needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Audio2txt access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audio2txt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audio2txt use?

Audio2txt is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio2txt use?

About 230 tokens (SKILL.md is roughly 920 characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audio2txt?

Skills that share tags, products or a category with Audio2txt: Triage (TalAter/annyang, 6.8k stars), Yichen Asr (mcncarl/yichen-skills, 4.3k stars), Dingtalk Minutes (DingTalk-Real-AI/dingtalk-workspace-cli, 3.2k stars) and Youtube Fetcher (JimmySadek/youtube-fetcher-to-markdown, 485 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio2txt?

CoderWanFeng (a GitHub user) maintains it in CoderWanFeng/python-office, which has 1,359 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on August 19, 2026.

Source: CoderWanFeng/python-office on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.