Agent skill

Transcribe Meeting

by amd in amd/gaia

Transcribe and summarize a meeting recording — speaker-attributed transcript, corrected mis-hearings, then a brief with decisions and action items.

MITAuto-check passedMedia & Creative

Install Transcribe Meeting

skills CLI
$ npx skills add amd/gaia --skill transcribe-meeting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/gaia transcribe-meeting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/hub/skills/transcribe-meeting .claude/skills/transcribe-meeting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
transcribe-meeting
GitHub stars
1.6k
Token cost
~1.9k tokens
SKILL.md length
1,004 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Transcribe and summarize a meeting recording — speaker-attributed transcript, corrected mis-hearings, then a brief with decisions and action items.

  • Works in 7 steps: Confirm the file exists, and set… → Transcribe. transcribe_media(file_path).… → Refine — one call, and it is not your… → …
  • The user points at an audio
  • SKILL.md covers Procedure, Never build the brief from a…, When speaker identification… and Speaker names are inferred,…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Transcribe Meeting is an agent skill from amd/gaia. Transcribe and summarize a meeting recording — speaker-attributed transcript, corrected mis-hearings, then a brief with decisions and action items. Use whenever the user points at an audio or video file (.mp4, .mkv, .mov, .m4a, .mp3, .wav) or a Teams/Zoom transcript export and asks to transcribe it, summarize it, take notes, write minutes, say what was discussed, who said what, what was decided, or what the action items and follow-ups are.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription. The repository describes itself as: Build AI agents for your PC. The licence is MIT.

When your agent uses it

  • The user points at an audio
  • Video file (.mp4
  • A Teams/Zoom transcript export and asks to transcribe it
  • Say what was discussed

Example prompts

  • “/transcribe-meeting”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the file exists, and set expectations. If the path is not there,
  2. Transcribe. transcribe_media(file_path). Pass output_path to write
  3. Refine — one call, and it is not your work to redo.
  4. Report both paths, every time. State the raw transcript_path and the
  5. Summarize. refine_transcript already indexed the file, so go straight
  6. Remember the outcome, not the transcript. remember a short record of
  7. Invite the follow-up. Close by telling the user they can ask questions

What it can do on your machine

Read from SKILL.md and the folder at commit 05fb50b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Transcribe Meeting loads about 1.9k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 1,004 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/gaia at commit 05fb50b, republished under its MIT licence (© amd). 1,004 words, ~1,923 tokens.

Download SKILL.mdSave it as .claude/skills/transcribe-meeting/SKILL.md (or your agent's skills folder).
name
transcribe-meeting
description
Transcribe and summarize a meeting recording — speaker-attributed transcript, corrected mis-hearings, then a brief with decisions and action items. Use whenever the user points at an audio or video file (.mp4, .mkv, .mov, .m4a, .mp3, .wav) or a Teams/Zoom transcript export and asks to transcribe it, summarize it, take notes, write minutes, say what was discussed, who said what, what was decided, or what the action items and follow-ups are.
license
MIT
version
1.1.0

Transcribe Meeting

A recording becomes a usable brief in three tool calls. Each one produces a file the next one reads — nothing is passed inline, because a 46-minute meeting is ~135,000 characters against a 60,000-character tool-result budget, and a summary built from the truncated half looks finished rather than truncated.

#CallProduces
1transcribe_media(file_path)Raw transcript at ~/.gaia/transcripts/<name>.txt. Returns transcript_path, never the text.
2refine_transcript(transcript_path)A speaker-labelled markdown transcript, already indexed. Returns refined_path, speakers, turns.
3summarize_document(refined_path, summary_type='detailed')The brief.

Procedure

  1. Confirm the file exists, and set expectations. If the path is not there, say so and quote it — never describe a recording from its filename. Measured on a 46-minute recording: ~6 min to transcribe, ~4 min to separate the voices, ~1-2 min to name them and summarize — about 11-12 minutes, or roughly 4x realtime. Say that before you start, not after the user has waited.

  2. Transcribe. transcribe_media(file_path). Pass output_path to write elsewhere, language to skip auto-detection. The transcript is written before anything downstream runs, so it survives a later failure; if the call errors, report its message as given — it names the missing component and the command that installs it.

    preview is the first 1200 characters, for identifying the recording only. low_confidence_spans and low_confidence_span_count are informational — step 3 does the correcting. Do not repair them yourself.

  3. Refine — one call, and it is not your work to redo. refine_transcript(transcript_path) groups the transcript by the voices already separated from the audio, puts names to them where the conversation supports it, writes the result and indexes it. You do not work out the speakers yourself, do not rewrite wording, and do not write a transcript file. Summarizing the raw transcript instead produces a brief with no owners on the action items.

    It needs the timings saved alongside the raw transcript. If it reports they are missing, re-run transcribe_media on the original media rather than trying to segment the text yourself — turn boundaries guessed from prose are invented, not observed.

  4. Report both paths, every time. State the raw transcript_path and the refined_path in your reply — even when the user only asked for action items, even if a later step fails. Transcription costs minutes of compute; a user who does not know the files exist pays for it twice.

  5. Summarize. refine_transcript already indexed the file, so go straight to summarize_document(refined_path, summary_type='detailed'), which folds the whole transcript forward in sections so the brief covers the entire meeting.

  6. Remember the outcome, not the transcript. remember a short record of the meeting: what it was, when, who was in it, the decisions, and any action item the user personally owes. Two or three sentences.

    The transcript is already indexed — that is what detailed questions read from. Memory is for the durable facts that should surface without being asked, weeks later, when the user says "what did I commit to?" or a related topic comes up in a different conversation.

    Do not store the transcript, long quotes, or the full brief in memory. That duplicates the index, crowds out everything else the user asked to be remembered, and gets recalled in conversations it has nothing to do with. Store nothing when the recording turned out to be something other than a meeting.

  7. Invite the follow-up. Close by telling the user they can ask questions about the meeting and you will answer from the indexed transcript. They will not discover this on their own.

Show full SKILL.md (432 more words)Show less

Never build the brief from a fragment

  • Not from preview. 1200 characters tells you which recording this is and nothing else — not the topic, not the participants, not the decisions, and certainly not the action items, which cluster at the end.
  • Not from query_documents. It returns only the top few matching chunks, so a summary built from it silently omits most of the meeting. query_documents is the right tool for a later follow-up question about a specific detail.

If a stage did not run, say which one and what the brief is therefore missing. Do not present a partial pass as the finished product.

When speaker identification did not run

transcribe_media returns a speaker_identification field ONLY when voices could not be separated — the engine was missing, or the audio defeated it. When it is there, say so in your reply and give its reason. A transcript with no speaker labels reads exactly like a meeting with one participant, so silence here is not neutral: it is a wrong answer the user has no way to catch.

Carry on and summarize anyway. A brief without owners still beats no brief — just do not present it as though you knew who spoke.

Speaker names are inferred, not identified

Voices ARE separated from the audio, by acoustic diarization — that part is measured, not guessed. What stays inferred is the name attached to each voice: refine_transcript reads self-introductions, direct address and role cues. So present the names as best-effort while the separation is not. Voices it could not name stay Speaker 1, Speaker 2; never upgrade one to a real name yourself. A confidently wrong name propagates into the brief, into who owns each action item, and into the indexed transcript that answers questions weeks later.

Output shape

refine_transcript writes:

markdown
# Transcript — staff-meeting

Source: C:\recordings\staff-meeting.mp4

## Speakers

- Priya Raman
- Dan Okafor
- Speaker 3

## Transcript

Priya Raman: Thanks everyone for joining...
Dan Okafor: The migration finished Tuesday...

Brief shape

Shape your brief from summarize_document's output as: what the meeting was for and what came out of it (2–3 sentences); key facts stated rather than inferred; action items with an owner each, unknown when the transcript does not say; and anything unresolved or blocked.

This section is the one to change when a reader wants their briefs a different way — a different order, a section dropped, a section added. Everything above it is about producing a trustworthy transcript and is not a matter of taste.

Fork this

The three calls stay as they are — the file-not-inline pipeline, the speaker honesty, and refinement before summarization are what make any output trustworthy. Change only Brief shape: an incident review wants Timeline, Root Cause, Impact and Remediation; a customer call wants Asks, Objections, Commitments and Next Steps.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in hub/skills/transcribe-meeting of amd/gaia.

Open the folder on GitHubat commit 05fb50b

Compare with similar skills

Transcribe Meeting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Transcribe Meeting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Transcribe Meeting this skillamd/gaia1.6k—~1.9kAutomated safety check: PassMIT
HyperFrames Media Useheygen-com/hyperframes59k—~2.4kAutomated safety check: PassApache-2.0
Native Subtitle Quote Imagechengyi-ai/native-subtitle-quote-image2.4k—~1.8kAutomated safety check: PassMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Transcription Memory ReconstructionNxcoreAI/EverRoom3k—~714Automated safety check: PassCustom licence
TranscribeJetBrains/skills3664 repos~776Automated safety check: PassApache-2.0

Similar skills

  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    59k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Native Subtitle Quote Image

    chengyi-ai/native-subtitle-quote-image

    将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…

    2.4k GitHub stars~1.8k tokensUpdated today
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated 11 days ago
    Media & CreativeAuto-check: notes
  • Reconstruct a complete, searchable memory from an untrusted meeting or conversation transcript.

    3k GitHub stars~714 tokensUpdated today
    Media & CreativeAuto-check passed
  • Transcribe

    JetBrains/skills

    Official

    Transcribe audio files to text with optional diarization and known-speaker hints.

    366 GitHub starsUsed in 4 repos~776 tokens
    Media & CreativeAuto-check passed
  • Bilibili Transcribe

    chubbyguan/chubbyskills

    哔哩哔哩视频 → 下载 → 转录 → 存为 Markdown 的完整工作流. An agent skill from chubbyguan/chubbyskills.

    1.2k GitHub stars~578 tokensUpdated yesterday
    Media & CreativeAuto-check: notes

More from amd/gaia

All 44 skills in this repo
  • Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

    1.6k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.

    1.6k GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.

    1.6k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.

    1.6k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Questions about Transcribe Meeting

What does Transcribe Meeting do?

Transcribe and summarize a meeting recording — speaker-attributed transcript, corrected mis-hearings, then a brief with decisions and action items. Transcribe Meeting is an agent skill from amd/gaia. Transcribe and summarize a meeting recording — speaker-attributed transcript, corrected mis-hearings, then a brief with decisions and action items.

When should I use Transcribe Meeting?

Transcribe Meeting fits situations like: the user points at an audio; video file (.mp4; A Teams/Zoom transcript export and asks to transcribe it; say what was discussed.

How do I install Transcribe Meeting in Claude Code?

Run `npx skills add amd/gaia --skill transcribe-meeting -a claude-code`. Or copy the skill folder (hub/skills/transcribe-meeting in amd/gaia) into .claude/skills/transcribe-meeting in your project. Claude Code loads it when a task matches its description.

How do I install Transcribe Meeting in Codex?

Run `npx skills add amd/gaia --skill transcribe-meeting -a codex`. Or copy the skill folder (hub/skills/transcribe-meeting in amd/gaia) into .agents/skills/transcribe-meeting in your project. Codex loads it when a task matches its description.

Can I use Transcribe Meeting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill transcribe-meeting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/transcribe-meeting, .gemini/skills/transcribe-meeting, .github/skills/transcribe-meeting and .opencode/skills/transcribe-meeting in your project.

What does Transcribe Meeting need to run?

SKILL.md names no scripts, command-line tools or credentials: Transcribe Meeting is instructions for the agent only.

Does Transcribe Meeting access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Transcribe Meeting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Transcribe Meeting use?

Transcribe Meeting is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Transcribe Meeting use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Transcribe Meeting?

Skills that share tags, products or a category with Transcribe Meeting: HyperFrames Media Use (heygen-com/hyperframes, 59k stars), Native Subtitle Quote Image (chengyi-ai/native-subtitle-quote-image, 2.4k stars), Edu Math Video (wy51ai/edulab, 1.4k stars) and Transcription Memory Reconstruction (NxcoreAI/EverRoom, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Transcribe Meeting?

amd (a GitHub organization) maintains it in amd/gaia, which has 1,579 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 8, 2026.

Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.