Agent skill

Transcribe

by greatSumini in greatSumini/cc-system

로컬 오디오 파일(m4a 등)을 한국어로 전사하는 skill. An agent skill from greatSumini/cc-system.

MITAuto-check passedMedia & Creative

Install Transcribe

skills CLI
$ npx skills add greatSumini/cc-system --skill transcribe -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install greatSumini/cc-system transcribe --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/greatSumini/cc-system.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/transcribe .claude/skills/transcribe && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
transcribe
GitHub stars
438
Token cost
~909 tokens
SKILL.md length
491 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

로컬 오디오 파일(m4a 등)을 한국어로 전사하는 skill. An agent skill from greatSumini/cc-system.

  • Works in 6 steps: 파일 경로 확인 → 도메인 힌트 구성 → 오디오 전처리 → …
  • Tasks that involve Transcription
  • SKILL.md covers 플로우, 화자 분리 요청 처리, Invariants and 의존성
  • Calls ffmpeg, uv and brew

What it does

Transcribe is an agent skill from greatSumini/cc-system. 로컬 오디오 파일(m4a 등)을 한국어로 전사하는 skill. "전사해줘", "받아쓰기", "음성 파일 텍스트로", "transcribe", "음성 받아적어줘", "녹음 파일 변환" 등 오디오 → 텍스트 변환 요청에 트리거된다.

Its SKILL.md is about 910 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription. The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription

Example prompts

  • “transcribe”
  • “/transcribe”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. 파일 경로 확인
  2. 도메인 힌트 구성
  3. 오디오 전처리
  4. 전사 실행
  5. 결과 미리보기 + 저장 위치 확인
  6. 정리

What it can do on your machine

Read from SKILL.md and the folder at commit 172bdd4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ffmpeg
    • uv
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Transcribe loads about 909 tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 491 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~909

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from greatSumini/cc-system at commit 172bdd4, republished under its MIT licence (© greatSumini). 491 words, ~909 tokens.

Download SKILL.mdSave it as .claude/skills/transcribe/SKILL.md (or your agent's skills folder).
name
transcribe
description
로컬 오디오 파일(m4a 등)을 한국어로 전사하는 skill. "전사해줘", "받아쓰기", "음성 파일 텍스트로", "transcribe", "음성 받아적어줘", "녹음 파일 변환" 등 오디오 → 텍스트 변환 요청에 트리거된다.

transcribe — 로컬 오디오 전사

Apple Silicon 네이티브 mlx-whisper(Whisper large-v3-turbo)로 로컬 오디오 파일을 한국어 텍스트로 전사한다.

플로우

1. 파일 경로 확인

사용자가 제공한 오디오 파일 경로를 확인한다. 상대경로면 현재 작업 디렉토리 기준 절대경로로 변환한다.

지원 포맷: m4a 기본, mp3/wav/mp4 등 ffmpeg이 디코딩 가능한 포맷 모두 동작.

2. 도메인 힌트 구성

현재 대화 맥락에서 다음 카테고리를 적극적으로 끌어모아 한국어 한두 문장으로 조합해 --initial-prompt에 넣는다. 맥락이 있으면 가능한 한 풍부하게 채운다(없는 정보를 지어내지는 않는다).

수집할 카테고리:

  • 참석자/화자 이름 (가능하면 전부)
  • 회사명·제품명·서비스명
  • 자주 쓰일 영문 약어·전문용어의 한글 표기 예시 (예: "토스페이먼츠를 'TPay'로도 부른다", "API를 '에이피아이'로 발음한다")
  • 도메인 키워드 5~20개

예시:

  • 대화에서 "최수민, 바이브마피아클럽, 구글, B2B 제안서"가 언급되었다면
    • --initial-prompt "이 녹음은 바이브마피아클럽 최수민과 구글의 B2B 제안 회의다. 등장 용어: 바이브마피아클럽(VMC), 구글, B2B 제안서, 견적, 라이선스."

맥락이 전혀 없으면 이 단계는 건너뛴다. 단, 사용자에게 "전사 품질을 올리려면 등장 인물·고유명사를 알려주세요"라고 한 번 권유한다.

3. 오디오 전처리

작은 목소리 화자 누락과 볼륨 편차로 인한 환각을 줄이기 위해 ffmpeg loudnorm으로 정규화한 뒤 전사한다. 이 단계는 항상 수행한다(추가 비용 거의 없음, 효과 큼).

bash
NORMALIZED="$(mktemp -t transcribe_norm).wav"
ffmpeg -y -i "<AUDIO_PATH>" \
  -af loudnorm=I=-16:TP=-1.5:LRA=11 \
  -ar 16000 -ac 1 \
  "$NORMALIZED"
  • -ar 16000 -ac 1: Whisper 내부 표현(16kHz 모노)에 맞춰 미리 변환. 모델 입력 변환 비용 절감.
  • 사용자가 "잡음이 심하다", "BGM 깔려 있다"고 명시한 경우에만 추가로 demucs(음성 분리) 또는 arnndn(RNNoise) 단계를 적용한다. 기본 플로우에는 넣지 않는다 — 모델 다운로드/실행 비용이 크고 깨끗한 녹음에선 오히려 음성을 깎는다.
4. 전사 실행

임시 출력 디렉토리에 txt 포맷으로 내보낸다. 환각 방지 플래그를 기본 적용한다.

bash
OUTDIR="$(mktemp -d -t transcribe)"
ORIG_BASE="$(basename "<AUDIO_PATH>")"; ORIG_BASE="${ORIG_BASE%.*}"
mlx_whisper \
  --model mlx-community/whisper-large-v3-turbo \
  --language ko \
  --output-format txt \
  --output-dir "$OUTDIR" \
  --output-name "$ORIG_BASE" \
  --condition-on-previous-text False \
  --temperature 0 \
  --no-speech-threshold 0.6 \
  --verbose False \
  [--initial-prompt "<DOMAIN_HINT>"] \
  "$NORMALIZED"

환각 방지 플래그 의도:

  • --condition-on-previous-text False: 한 번 잘못 인식한 텍스트가 뒤 구간으로 전염되는 문제 차단. 회의 녹음에선 거의 항상 이득.
  • --temperature 0: 샘플링 무작위성 제거. 동일 입력에 동일 출력 보장.
  • --no-speech-threshold 0.6: 무음/BGM 구간에서 "구독과 좋아요 부탁드립니다" 같은 환각 생성 방지. 기본값(0.6)을 명시적으로 박아둔다.

기타:

  • 첫 실행 시 모델(~1.5GB)이 HuggingFace에서 자동 다운로드된다. 사용자에게 미리 알린다.
  • 60분 오디오 기준 M 시리즈 Mac에서 3~8분 소요 예상.
  • 출력 파일명은 {원본 오디오 basename}.txt 형태로 $OUTDIR 아래에 생성된다 (--output-name 덕분에 정규화 임시파일 이름이 새지 않음).
Show full SKILL.md (188 more words)Show less
5. 결과 미리보기 + 저장 위치 확인

생성된 txt 파일을 읽어 사용자에게 미리보기를 제시한다 (긴 경우 앞/뒤 일부만).

그다음 저장 위치를 사용자에게 반드시 질의한다. 저장 위치 후보:

  • 원본 오디오 파일과 같은 디렉토리에 {basename}.txt
  • logs/transcripts/YYYY-MM-DD_{slug}.md (frontmatter 포함)
  • meeting-logs/YYMMDD_{title}.md (전사 후 회의록 가공까지 요청한 경우)
  • projects/{project_id}/meeting-logs/... (프로젝트 관련 녹음)
  • 저장하지 않고 화면 출력만

사용자가 저장 위치를 지정하면 해당 위치로 이동(mv) 또는 복사(cp)한다. 마크다운으로 저장하는 경우 ## Document Frontmatter 규칙에 따른 frontmatter를 추가한다.

6. 정리

임시 출력 디렉토리($OUTDIR)와 정규화된 임시 오디오 파일($NORMALIZED)은 저장 완료 후 삭제한다.

bash
rm -rf "$OUTDIR" "$NORMALIZED"

화자 분리 요청 처리

사용자가 "화자 구분", "누가 말했는지", "speaker diarization" 등을 명시적으로 요청하면:

현재 skill은 화자 분리를 지원하지 않습니다. 추가하려면 pyannote.audio + HuggingFace 토큰 설정이 필요합니다. 진행하시겠습니까?

라고 안내하고 사용자 결정을 기다린다. 임의로 진행하지 않는다.

Invariants

  • 언어는 한국어(ko) 고정. 다른 언어 요청 시 사용자에게 확인 후 --language 변경.
  • 모델은 기본 mlx-community/whisper-large-v3-turbo. 품질 불만 시에만 whisper-large-v3 (turbo 아님, 더 느리고 품질 살짝 높음) 로 교체.
  • 결과 저장 위치는 항상 사용자에게 확인한다. 임의로 리포지토리 안에 커밋될 파일을 만들지 않는다.
  • 원본 오디오 파일은 절대 수정/삭제하지 않는다.

의존성

사전 설치 완료된 상태를 가정한다:

  • mlx_whisper (uv tool install mlx-whisper)
  • ffmpeg (brew install ffmpeg)

둘 중 하나라도 없으면 사용자에게 설치 필요함을 알리고 중단한다.

© greatSumini, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/transcribe of greatSumini/cc-system.

Open the folder on GitHubat commit 172bdd4

Compare with similar skills

Transcribe next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Transcribe compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Transcribe this skillgreatSumini/cc-system438—~909Automated safety check: PassMIT
HyperFrames Media Useheygen-com/hyperframes59k—~2.4kAutomated safety check: PassApache-2.0
Native Subtitle Quote Imagechengyi-ai/native-subtitle-quote-image2.4k—~1.8kAutomated safety check: PassMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Transcription Memory ReconstructionNxcoreAI/EverRoom3k—~714Automated safety check: PassCustom licence
TranscribeJetBrains/skills3664 repos~776Automated safety check: PassApache-2.0

Similar skills

  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    59k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Native Subtitle Quote Image

    chengyi-ai/native-subtitle-quote-image

    将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…

    2.4k GitHub stars~1.8k tokensUpdated today
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated 11 days ago
    Media & CreativeAuto-check: notes
  • Reconstruct a complete, searchable memory from an untrusted meeting or conversation transcript.

    3k GitHub stars~714 tokensUpdated today
    Media & CreativeAuto-check passed
  • Transcribe

    JetBrains/skills

    Official

    Transcribe audio files to text with optional diarization and known-speaker hints.

    366 GitHub starsUsed in 4 repos~776 tokens
    Media & CreativeAuto-check passed
  • Bilibili Transcribe

    chubbyguan/chubbyskills

    哔哩哔哩视频 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills.

    1.2k GitHub stars~578 tokensUpdated yesterday
    Media & CreativeAuto-check: notes

More from greatSumini/cc-system

All 12 skills in this repo
  • Slash Command Creator

    greatSumini/cc-system

    Guide for creating Claude Code slash commands. An agent skill from greatSumini/cc-system.

    438 GitHub stars~773 tokensUpdated 4 mo ago
    Auto-check passed
  • Subagent Creator

    greatSumini/cc-system

    Create specialized Claude Code sub-agents with custom system prompts and tool configurations.

    438 GitHub stars~958 tokensUpdated 4 mo ago
    Auto-check passed
  • Cc Usage Audit

    greatSumini/cc-system

    한 프로젝트에서 사용자가 Claude Code에 입력한 프롬프트·작업 이력을 분석해 (1) 사용 패턴 정량화, (2) 사용자의 개발 철학 추출(근거 인용), (3) 그 철학을 렌즈로 한 메타 시스템(하니스·게이트·CI·프로세스) audit, (4) 우선순위가 매겨진 개선점 발굴을 수행한다.

    438 GitHub stars~586 tokensUpdated 4 mo ago
    Auto-check passed
  • Hook Creator

    greatSumini/cc-system

    Create and configure Claude Code hooks for customizing agent behavior.

    438 GitHub stars~662 tokensUpdated 4 mo ago
    Auto-check: notes
  • Youtube Collector

    greatSumini/cc-system

    유튜브 채널을 등록하고 새 컨텐츠를 수집하여 자막 기반 요약을 생성하는 skill. An agent skill from greatSumini/cc-system.

    438 GitHub stars~846 tokensUpdated 4 mo ago
    Auto-check passed
  • Ship

    greatSumini/cc-system

    지금까지의 작업을 한 번에 출하한다 — 커밋 → push → PR 생성 → squash merge. An agent skill from greatSumini/cc-system.

    438 GitHub stars~697 tokensUpdated 4 mo ago
    Auto-check passed

Questions about Transcribe

What does Transcribe do?

로컬 오디오 파일(m4a 등)을 한국어로 전사하는 skill. An agent skill from greatSumini/cc-system. Transcribe is an agent skill from greatSumini/cc-system. 로컬 오디오 파일(m4a 등)을 한국어로 전사하는 skill.

When should I use Transcribe?

Transcribe fits situations like: tasks that involve Transcription.

How do I install Transcribe in Claude Code?

Run `npx skills add greatSumini/cc-system --skill transcribe -a claude-code`. Or copy the skill folder (.claude/skills/transcribe in greatSumini/cc-system) into .claude/skills/transcribe in your project. Claude Code loads it when a task matches its description.

How do I install Transcribe in Codex?

Run `npx skills add greatSumini/cc-system --skill transcribe -a codex`. Or copy the skill folder (.claude/skills/transcribe in greatSumini/cc-system) into .agents/skills/transcribe in your project. Codex loads it when a task matches its description.

Can I use Transcribe in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add greatSumini/cc-system --skill transcribe -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/transcribe, .gemini/skills/transcribe, .github/skills/transcribe and .opencode/skills/transcribe in your project.

What does Transcribe need to run?

Going by SKILL.md and its folder, Transcribe needs the command-line tools its instructions call (ffmpeg, uv and brew).

Does Transcribe access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Transcribe safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Transcribe use?

Transcribe is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Transcribe use?

About 909 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Transcribe?

Skills that share tags, products or a category with Transcribe: HyperFrames Media Use (heygen-com/hyperframes, 59k stars), Native Subtitle Quote Image (chengyi-ai/native-subtitle-quote-image, 2.4k stars), Edu Math Video (wy51ai/edulab, 1.4k stars) and Transcription Memory Reconstruction (NxcoreAI/EverRoom, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Transcribe?

greatSumini (a GitHub user) maintains it in greatSumini/cc-system, which has 438 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on June 9, 2026.

Source: greatSumini/cc-system on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.