Agent skill

Interview Transcription

by jamditis in jamditis/claude-skills-journalism

Transcription, recording management, and quote extraction. An agent skill from jamditis/claude-skills-journalism.

MITAuto-check passedMedia & Creative

Install Interview Transcription

skills CLI
$ npx skills add jamditis/claude-skills-journalism --skill interview-transcription -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jamditis/claude-skills-journalism interview-transcription --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jamditis/claude-skills-journalism.git skills-src && mkdir -p .claude/skills && cp -r skills-src/journalism-core/skills/interview-transcription .claude/skills/interview-transcription && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
interview-transcription
GitHub stars
417
Token cost
~3.7k tokens
SKILL.md length
647 words
Files
2
Skills in repo
53
Repo updated
First seen
Licence
MIT

At a glance

Transcription, recording management, and quote extraction. An agent skill from jamditis/claude-skills-journalism.

  • Processing audio/video
  • SKILL.md covers When to activate, Recording setup for…, Transcription workflows and Quote extraction and…, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Generating timestamped transcripts

What it does

Interview Transcription is an agent skill from jamditis/claude-skills-journalism. Transcription, recording management, and quote extraction. Use when processing audio/video or generating timestamped transcripts.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Media & Creative, covering Transcription. It works with Whisper. The repository describes itself as: Claude Code skills for journalism, media, and academia - verification, FOIA, data journalism, academic writing, and more. The licence is MIT.

When your agent uses it

  • Processing audio/video
  • Generating timestamped transcripts

Example prompts

  • “/interview-transcription”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit e3e2172. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Interview Transcription loads about 3.7k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 647 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jamditis/claude-skills-journalism at commit e3e2172, republished under its MIT licence (© jamditis). 647 words, ~3,659 tokens.

Download SKILL.mdSave it as .claude/skills/interview-transcription/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
interview-transcription
description
Transcription, recording management, and quote extraction. Use when processing audio/video or generating timestamped transcripts.

Interview transcription and management

Practical workflows for journalists managing interviews from preparation through publication.

When to activate

  • Preparing questions for an interview
  • Processing audio/video recordings
  • Creating or managing transcripts
  • Organizing notes from multiple sources
  • Building a source relationship database
  • Generating timestamped quotes for fact-checking
  • Converting recordings to publishable quotes

Recording setup for transcription

For pre-interview research, question design, attribution agreements, and consent scripts, use the interview-prep skill. The notes here cover only the recording configuration that affects transcription quality.

python
# Standard recording configuration for clean transcription
RECORDING_SETTINGS = {
    'format': 'wav',           # Lossless for transcription
    'sample_rate': 16000,      # Whisper resamples to 16k anyway; 16k saves disk
    'channels': 1,             # Mono is fine for speech; stereo only if mics are positionally distinct
    'backup': True,            # Always run a backup recorder
}

# File naming convention
# YYYY-MM-DD_source-lastname_topic.wav
# Example: 2026-05-08_smith_budget-hearing.wav

Two-device rule. Always record on two devices. Phone as backup minimum. If using a wireless lav mic, the recorder built into the lav unit is one device; the phone running a backup app is the second.

Mono is preferred unless each speaker has their own dedicated microphone routed to a distinct channel. Stereo with both speakers bleeding into both channels is worse for diarization than clean mono.

Transcription workflows

Automated transcription pipeline

Vanilla OpenAI Whisper transcribes audio to text but does not assign speaker labels. To get diarized output ("Speaker 1:" / "Speaker 2:" / etc.) you need a tool that combines Whisper with a diarization model, typically WhisperX (m-bain/whisperX), which wraps faster-whisper transcription with pyannote.audio diarization and produces word-level timestamps with speaker IDs in one pass.

python
from pathlib import Path
import subprocess
import json

def transcribe_interview(
    audio_path: str,
    output_dir: str = "./transcripts",
    diarize: bool = True,
    hf_token: str | None = None,
    min_speakers: int = 2,
    max_speakers: int = 2,
) -> dict:
    """
    Transcribe an interview using WhisperX (Whisper + pyannote diarization).
    Returns a transcript with word-level timestamps and speaker labels.

    Diarization needs a Hugging Face token with access to the pyannote
    speaker-diarization-3.1 model. Accept the model EULA at
    huggingface.co/pyannote/speaker-diarization-3.1 once, then pass the token.
    """
    Path(output_dir).mkdir(exist_ok=True)

    cmd = [
        'whisperx', audio_path,
        '--model', 'large-v3',
        '--output_format', 'json',
        '--output_dir', output_dir,
        '--language', 'en',
        '--compute_type', 'int8',     # CPU-friendly; use 'float16' on GPU
        '--min_speakers', str(min_speakers),
        '--max_speakers', str(max_speakers),
    ]

    if diarize:
        cmd.append('--diarize')
        if hf_token:
            cmd += ['--hf_token', hf_token]

    subprocess.run(cmd, check=True, capture_output=True)

    json_path = Path(output_dir) / f"{Path(audio_path).stem}.json"
    with open(json_path) as f:
        return json.load(f)

def format_for_editing(transcript: dict) -> str:
    """Convert to journalist-friendly format with timestamps."""
    lines = []
    for segment in transcript.get('segments', []):
        timestamp = format_timestamp(segment['start'])
        text = segment['text'].strip()
        lines.append(f"[{timestamp}] {text}")
    return '\n\n'.join(lines)

def format_timestamp(seconds: float) -> str:
    """Convert seconds to HH:MM:SS format."""
    h = int(seconds // 3600)
    m = int((seconds % 3600) // 60)
    s = int(seconds % 60)
    return f"{h:02d}:{m:02d}:{s:02d}"

Falling back to plain Whisper. If diarization is overkill or you can't get a Hugging Face token, drop the --diarize flag, the model still produces accurate timestamped transcription and you label speakers manually based on context. faster-whisper (CTranslate2 backend) is the speed-optimized variant and works the same way at the CLI. whisper.cpp is the C++ port for resource-constrained machines (Raspberry Pi, older laptops); it doesn't include diarization but runs the small/medium models on CPU comfortably.

Manual transcription template

For sensitive interviews or when AI transcription fails:

markdown
## Transcript: [Source] - [Date]

**Recording file**: [filename]
**Duration**: [XX:XX]
**Transcribed by**: [name]
**Verified against recording**: [ ] Yes / [ ] No

---

[00:00:15] **Q**: [Your question]

[00:00:45] **A**: [Source response - verbatim, including ums, pauses noted as (...)]

[00:01:30] **Q**: [Follow-up]

[00:01:42] **A**: [Response]

---

## Notes
- [Anything not captured in audio: gestures, documents shown, etc.]

## Potential quotes
- [00:01:42] "Quote that stands out" - context: [why it matters]

Quote extraction and verification

Pull quotes workflow
python
from dataclasses import dataclass
from typing import Optional
import re

@dataclass
class Quote:
    text: str
    timestamp: str
    speaker: str
    context: str
    verified: bool = False
    used_in: Optional[str] = None

class QuoteBank:
    """Manage quotes from interview transcripts."""

    def __init__(self):
        self.quotes = []

    def extract_quote(self, transcript: str, start_time: str,
                      end_time: str, speaker: str, context: str) -> Quote:
        """Extract and store a quote with metadata."""
        # Pull text between timestamps
        pattern = rf'\[{re.escape(start_time)}\](.+?)(?=\[\d|$)'
        match = re.search(pattern, transcript, re.DOTALL)

        if match:
            text = match.group(1).strip()
            quote = Quote(
                text=text,
                timestamp=start_time,
                speaker=speaker,
                context=context
            )
            self.quotes.append(quote)
            return quote
        return None

    def verify_quote(self, quote: Quote, audio_path: str) -> bool:
        """Mark quote as verified against original recording."""
        # In practice: listen to audio at timestamp, confirm accuracy
        quote.verified = True
        return True

    def export_for_story(self) -> str:
        """Export verified quotes ready for publication."""
        output = []
        for q in self.quotes:
            if q.verified:
                output.append(f'"{q.text}"\n- {q.speaker}\n[Timestamp: {q.timestamp}]')
        return '\n\n'.join(output)
Quote accuracy checklist

Before publishing any quote:

markdown
- [ ] Listened to original recording at timestamp
- [ ] Quote is verbatim (or clearly marked as paraphrased)
- [ ] Context preserved (not cherry-picked to change meaning)
- [ ] Speaker identified correctly
- [ ] Timestamp documented for fact-checker
- [ ] Source approved quote (if agreement made)

Source management database

Interview tracking schema
python
from dataclasses import dataclass, field
from datetime import datetime
from typing import List, Optional
from enum import Enum

class SourceStatus(Enum):
    ACTIVE = "active"           # Currently engaged
    DORMANT = "dormant"         # Not recently contacted
    DECLINED = "declined"       # Refused to participate
    OFF_RECORD = "off_record"   # Background only

class InterviewType(Enum):
    ON_RECORD = "on_record"
    BACKGROUND = "background"
    DEEP_BACKGROUND = "deep_background"
    OFF_RECORD = "off_record"

@dataclass
class Source:
    name: str
    organization: str
    contact_info: dict  # email, phone, signal, etc.
    beat: str
    status: SourceStatus = SourceStatus.ACTIVE
    interviews: List['Interview'] = field(default_factory=list)
    notes: str = ""

    # Relationship tracking
    first_contact: Optional[datetime] = None
    trust_level: int = 1  # 1-5 scale

@dataclass
class Interview:
    source: str
    date: datetime
    interview_type: InterviewType
    recording_path: Optional[str] = None
    transcript_path: Optional[str] = None
    story_slug: Optional[str] = None
    key_quotes: List[str] = field(default_factory=list)
    follow_up_needed: bool = False
    notes: str = ""
Quick source lookup
python
def find_sources_for_story(sources: List[Source], topic: str,
                           beat: str = None) -> List[Source]:
    """Find relevant sources for a new story."""
    matches = []
    for source in sources:
        # Filter by beat if specified
        if beat and source.beat != beat:
            continue
        # Only suggest active sources
        if source.status != SourceStatus.ACTIVE:
            continue
        # Check if they've spoken on similar topics
        for interview in source.interviews:
            if topic.lower() in interview.notes.lower():
                matches.append(source)
                break

    # Sort by trust level
    return sorted(matches, key=lambda s: s.trust_level, reverse=True)

Audio/video processing

Batch processing multiple recordings
python
from pathlib import Path
from concurrent.futures import ProcessPoolExecutor
import json

def batch_transcribe(recordings_dir: str, output_dir: str) -> dict:
    """Process all recordings in a directory."""
    recordings = list(Path(recordings_dir).glob('*.wav')) + \
                 list(Path(recordings_dir).glob('*.mp3')) + \
                 list(Path(recordings_dir).glob('*.m4a'))

    results = {}

    with ProcessPoolExecutor(max_workers=4) as executor:
        futures = {
            executor.submit(transcribe_interview, str(rec), output_dir): rec
            for rec in recordings
        }

        for future in futures:
            rec = futures[future]
            try:
                transcript = future.result()
                results[rec.name] = {
                    'status': 'success',
                    'transcript': transcript
                }
            except Exception as e:
                results[rec.name] = {
                    'status': 'error',
                    'error': str(e)
                }

    return results
Video interview extraction
python
import subprocess

def extract_audio_from_video(video_path: str, output_path: str = None) -> str:
    """Extract audio track from video for transcription."""
    if output_path is None:
        output_path = video_path.rsplit('.', 1)[0] + '.wav'

    subprocess.run([
        'ffmpeg', '-i', video_path,
        '-vn',  # No video
        '-acodec', 'pcm_s16le',  # WAV format
        '-ar', '44100',  # Sample rate
        '-ac', '1',  # Mono
        output_path
    ], check=True)

    return output_path
markdown
## Recording consent record

**Date**:
**Source name**:
**Recording type**: [ ] Audio [ ] Video
**Interview type**: [ ] On record [ ] Background [ ] Off record

### Consent obtained:
- [ ] Verbal consent recorded at start of interview
- [ ] Written consent form signed
- [ ] Email confirmation of consent

### Jurisdiction notes:
- Interview location state/country:
- One-party or two-party consent jurisdiction:
- Any specific restrictions agreed:

### Agreed terms:
- [ ] Full attribution allowed
- [ ] Organization attribution only
- [ ] Anonymous source
- [ ] Review quotes before publication
- [ ] Embargo until [date]:

For the per-state breakdown of one-party vs. all-party consent, hidden-recording rules, and federal preemption, use the interview-prep skill (which points to the Reporters Committee for Freedom of the Press Reporter's Recording Guide, the authoritative continuously-updated source).

Always get explicit consent on recording regardless of jurisdiction. Note the consent verbatim at the head of every transcript file (timestamp, speaker, response). This protects you legally everywhere and gives the fact-checker a clean starting point.

Show full SKILL.md (245 more words)Show less

Tools and resources

ToolPurposeNotes
OpenAI WhisperLocal transcription, no diarizationFree, runs offline. large-v3 is the current best model
WhisperXWhisper + speaker diarizationm-bain/whisperX. Free. Word-level timestamps with speaker IDs. Needs a Hugging Face token for the pyannote model
faster-whisperSpeed-optimized WhisperCTranslate2 backend. ~4x faster than vanilla Whisper at the same accuracy. Used internally by WhisperX
whisper.cppCPU-friendly Whisper portC++ implementation. Runs the small/medium models on a Raspberry Pi
pyannote.audioStandalone speaker diarizationUse directly when you already have transcripts from another source
MacWhisper / BuzzGUI wrappers for WhispermacOS / cross-platform GUIs for journalists who don't want a CLI
Otter.aiCloud transcription, real-timeVerify privacy posture before using with sensitive sources, Otter Pilot has historically joined meetings unannounced and indexed transcripts; check current settings
DescriptEdit audio like textGood for pulling clips. Cloud-hosted
Rev (human + AI)Human transcription for sensitive materialSlower, more accurate. Cloud-hosted
TrintJournalist-focused, collaborationCloud-hosted. Has team features
oTranscribeFree web-based manual transcription aidLocal-only (browser); no upload. Good for off-the-record material you can't hand to a cloud service
  • interview-prep, Pre-interview research, question design, consent scripts, and recording-law jurisdiction
  • source-verification, Verify source credentials before interview
  • fact-check-workflow, Verify quotes against the recording before publication
  • foia-requests, Get documents to inform interview questions
  • data-journalism, Analyze data sources mentioned in interviews
  • newsroom-style, Convert verbatim quotes into AP-style copy for publication

Skill metadata

FieldValue
version1.0.0
created2025-12-26
updated2026-05-08
authorJoe Amditis
domainjournalism, research
complexityintermediate

© jamditis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in journalism-core/skills/interview-transcription of jamditis/claude-skills-journalism.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit e3e2172

Compare with similar skills

Interview Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Interview Transcription compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Interview Transcription this skilljamditis/claude-skills-journalism417—~3.7kAutomated safety check: PassMIT
Transcribe Mdhrescak/transcribe-md104—~474Automated safety check: NotesMIT
Wjs Transcribing Audiojianshuo/claude-skills131—~4.4kAutomated safety check: NotesMIT
Whisper Transcriptionbenchflow-ai/skillsbench1.8k—~1.1kAutomated safety check: PassApache-2.0
Voice Memo SyncLeoYeAI/openclaw-master-skills2.2k—~5.1kAutomated safety check: PassMIT
WhisperAlexAI-MCP/hermes-CCC135—~1.9kAutomated safety check: PassMIT

Similar skills

  • Transcribe Md

    hrescak/transcribe-md

    Record and transcribe audio to a markdown file using whisper.cpp (mic + system audio)

    104 GitHub stars~474 tokensUpdated 6 mo ago
    Media & CreativeAuto-check: notes
  • Wjs Transcribing Audio

    jianshuo/claude-skills

    A skill your agent uses when the user has audio or video and wants a timestamped transcript (SRT) in the source language.

    131 GitHub stars~4.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Whisper Transcription

    benchflow-ai/skillsbench

    Transcribe audio/video to text with word-level timestamps using OpenAI Whisper.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Voice Memo Sync

    LeoYeAI/openclaw-master-skills

    Sync, transcribe, and intelligently organize voice memos, audio/video files, and URLs.

    2.2k GitHub stars~5.1k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Whisper

    AlexAI-MCP/hermes-CCC

    OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.

    135 GitHub stars~1.9k tokensUpdated 6 mo ago
    Media & CreativeAuto-check passed
  • Transcription

    MadAppGang/claude-code

    Audio/video transcription using OpenAI Whisper. An agent skill from MadAppGang/claude-code.

    285 GitHub stars~1.7k tokensUpdated 6 mo ago
    Media & CreativeAuto-check passed

More from jamditis/claude-skills-journalism

All 53 skills in this repo
  • Web Design Picker

    jamditis/claude-skills-journalism

    A skill your agent uses when creating distinct website directions, a client review picker, asset catalog, previews, and Cloudflare-ready handoffs.

    417 GitHub stars~3.1k tokensUpdated 4 days ago
    Auto-check passed
  • Okf Wiki

    jamditis/claude-skills-journalism

    Builds an Open Knowledge Format (OKF) knowledge base from existing docs, notes, or a repo.

    417 GitHub stars~4.7k tokensUpdated 4 days ago
    Auto-check passed
  • Private Secret Scanning

    jamditis/claude-skills-journalism

    Local Gitleaks scans for staged changes, push ranges, and full history in private repos, with redacted reports.

    417 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed
  • Data Journalism

    jamditis/claude-skills-journalism

    Acquire, clean, analyze, verify, visualize, and explain data for journalism.

    417 GitHub stars~1.6k tokensUpdated 4 days ago
    Auto-check passed
  • Document Design

    jamditis/claude-skills-journalism

    Creates print-ready HTML that exports to PDF. An agent skill from jamditis/claude-skills-journalism.

    417 GitHub stars~1.9k tokensUpdated 4 days ago
    Auto-check passed
  • Using Superjawn

    jamditis/claude-skills-journalism

    Establishes how to find and use skills, requiring Skill tool invocation before any response.

    417 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Interview Transcription

What does Interview Transcription do?

Transcription, recording management, and quote extraction. An agent skill from jamditis/claude-skills-journalism. Interview Transcription is an agent skill from jamditis/claude-skills-journalism. Transcription, recording management, and quote extraction.

When should I use Interview Transcription?

Interview Transcription fits situations like: processing audio/video; generating timestamped transcripts.

How do I install Interview Transcription in Claude Code?

Run `npx skills add jamditis/claude-skills-journalism --skill interview-transcription -a claude-code`. Or copy the skill folder (journalism-core/skills/interview-transcription in jamditis/claude-skills-journalism) into .claude/skills/interview-transcription in your project. Claude Code loads it when a task matches its description.

How do I install Interview Transcription in Codex?

Run `npx skills add jamditis/claude-skills-journalism --skill interview-transcription -a codex`. Or copy the skill folder (journalism-core/skills/interview-transcription in jamditis/claude-skills-journalism) into .agents/skills/interview-transcription in your project. Codex loads it when a task matches its description.

Can I use Interview Transcription in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jamditis/claude-skills-journalism --skill interview-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/interview-transcription, .gemini/skills/interview-transcription, .github/skills/interview-transcription and .opencode/skills/interview-transcription in your project.

What does Interview Transcription need to run?

SKILL.md names no scripts, command-line tools or credentials: Interview Transcription is instructions for the agent only. Our summary lists: Python 3.

Does Interview Transcription access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Interview Transcription safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Interview Transcription use?

Interview Transcription is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Interview Transcription use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Interview Transcription?

Skills that share tags, products or a category with Interview Transcription: Transcribe Md (hrescak/transcribe-md, 104 stars), Wjs Transcribing Audio (jianshuo/claude-skills, 131 stars), Whisper Transcription (benchflow-ai/skillsbench, 1.8k stars) and Voice Memo Sync (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Interview Transcription?

jamditis (a GitHub user) maintains it in jamditis/claude-skills-journalism, which has 417 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on October 4, 2026.

Source: jamditis/claude-skills-journalism on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.