Agent skill

Lecture To Notes

by drpwchen in drpwchen/lecture-to-notes

Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or…

MITAuto-check: notesMedia & Creative

Install Lecture To Notes

skills CLI
$ npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install drpwchen/lecture-to-notes lecture-to-notes --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
lecture-to-notes
GitHub stars
108
Token cost
~4.9k tokens
SKILL.md length
1,913 words
Files
87 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or…

  • Works in 12 steps: Ask the language (mandatory, no command) → Set up the lecture directory → GPU pre-flight → …
  • Tasks that involve Transcription
  • SKILL.md covers HARD RULES, Input types and routing, Pipeline and Edge cases, plus 3 more sections
  • Calls python, ollama and pip

What it does

Lecture To Notes is an agent skill from drpwchen/lecture-to-notes. Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or a dropped media file. Handles batch runs.

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 91 other files, including scripts (for example `.github/workflows/secret-scan.yml`, `CHANGELOG.md` and `README.md`).

It sits in Media & Creative, covering Transcription and Slides and decks. The repository describes itself as: Lecture recordings → structured grounded notes + a synced HTML viewer: video, timestamped transcript and curated summary on one page. Local GPU pipeline (Whisper ASR · slide… The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Slides and decks

Example prompts

  • “/lecture-to-notes”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, Glob, Grep, Agent

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Ask the language (mandatory, no command)
  2. Set up the lecture directory
  3. GPU pre-flight
  4. Transcribe
  5. Stage A: frame extraction (Path A only)
  6. Stage B: quick OCR + entropy (Path A only)
  7. alt — Path B / B-images bridge
  8. Stage C: semantic dedup (Path A only)
  9. Stage B2: high-quality OCR (Surya)
  10. Stage D: VLM signals
  11. Stage E: transcript grounding
  12. Flag suspect ASR tokens

What it can do on your machine

Read from SKILL.md and the folder at commit 79053a3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • Glob
    • Grep
    • Agent

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • ollama
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Lecture To Notes loads about 4.9k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 1,913 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, Glob, Grep, Agent

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from drpwchen/lecture-to-notes at commit 79053a3, republished under its MIT licence (© drpwchen). 1,913 words, ~4,914 tokens.

Download SKILL.mdSave it as .claude/skills/lecture-to-notes/SKILL.md (or your agent's skills folder). This skill also uses 86 other files; get the full folder from GitHub.
name
lecture-to-notes
description
Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or a dropped media file. Handles batch runs.
allowed-tools
Read, Write, Edit, Bash, Glob, Grep, Agent

Lecture-to-Notes

Turn a lecture recording (video or audio-only) into structured notes. Every heavy stage runs locally at 0 Claude tokens; Claude only does the final synthesis. This page is the map; detail lives in reference/, one topic per file.

FileWhat is in it
reference/pipeline.mdPer-stage flags, thresholds, JSON schemas, timeouts, observability
reference/note-spec.mdNote quality spec, tier scoring, width table, synthesis prompt requirements
reference/segmented-mode.mdMulti-talk workshop folders → per-segment L2/L3 + Hub + web viewer
reference/multi-camera.mdOne long recording + many phone clips/photos → one timeline
reference/decisions.mdPost-mortems, benchmarks, wrong turns, VRAM measurements

HARD RULES

  1. ==ASK the user what language the speaker(s) used== (English / Mandarin / bilingual code-switching) before transcribing. There is no default and transcribe_video.py exits without --lang. A wrong guess makes Whisper hallucinate Chinese from accented English and the transcript is unusable.
  2. ==Never skip Stage D (VLM) or Stage E (grounding)== for speed or for a deadline. ==The user has not set a deadline; do not invent one.== If a stage really is too slow (>2 h ETA), report the ETA and ask.
  3. ==Never auto-correct the transcript.== Flag suspects, let synthesis resolve them. Both auto-correction passes ever built were measured and retired — see reference/decisions.md#asr-auto-correction. Do not add an auto-apply mode.
  4. ==Do not bypass the collapse auto-retry.== transcribe_video.py auto-runs retranscribe_segment.py --auto on detected token-collapse. If collapses survive that, escalate (wider beam + --no-repeat-ngram-size + a glossary), never skip.
  5. ==The VLM does not do OCR.== Stage D asks only for semantic signals. Text comes from Stage B quick_text, Stage B2 clean_text, or pdf_text.json. ocr.vlm_text is an always-empty compatibility field.
  6. ==Serialize all GPU work.== Whisper and the VLM may not run concurrently on an 8 GB card, and frame extraction must not run alongside transcription.
  7. ==On an 8 GB card, never exceed --batch-size 4 with --beam-size 10==, and never combine --beam-size 15 with sequential mode — that crashes (0xC0000005). The measured sweet spot is --batch-size 3 --beam-size 10.
  8. ==Director batch dispatch==: for a multi-lecture batch, dispatch Steps 1–9 as one subagent and Step 10 synthesis as a separate fresh subagent, spawned only after slides_grounded.json exists. A single subagent bounces during the long GPU waits and burns 30+ min of wall time per lecture.
  9. ==Order material by REAL CAPTURE TIME — never by filename, never by the printed agenda.== Filenames are labels, not clocks: a camcorder counter restarts across days (a two-day shoot has two 00000), a recorder's 240526_1119.mp3 sorts into the middle of the video files, and on-site -1-/-2- labels get stuck on the wrong file. Run scripts/batch/course_timeline.py BEFORE segmenting any multi-source course; manifest.json clip order and L1_coarse.md section order must both be built from it. An agenda is not a clock either — the 2024-05 Conference-Y conference ran ~25 min early on day 1 and ~35 min late on day 2, while its break gaps matched to the minute.
  10. ==🚫 PHI red line==: if the recording contains patient-identifiable content (case discussion, ward rounds, named patients), transcribe LOCAL ONLY — drop --engine groq. When unsure, ask; default to local.

Input types and routing

==Start here. route_inputs.py is the front door== — it classifies a folder and prints the ordered commands plus the questions a human must answer. It is plan-only: it never runs anything and never writes a file.

bash
python <skill-dir>/scripts/route_inputs.py <material_dir> [--recursive] [--out-dir DIR] [--json]
What is in the folderSlide sourceRoute
Video, no deckframes from the videoPath A — Steps 5–7
Audio/video + PDF deck (==preferred==)PDF text + page rendersPath B — build_slides_from_pdf.py
Audio/video + loose slide images (≥3)the images themselvesPath B-images — build_slides_from_images.py
Audio only, no decknonePath C — transcript-only note
N-up handout PDFcropped tilesPath B-multi — crop_multiup_pdf.py first
Multi-talk workshop folderper segmentreference/segmented-mode.md
One long recording + many phone clips/photosper sourcereference/multi-camera.md
.pptx / .docx / .key—convert to PDF yourself first; there is no conversion step here

==Multi-source contract==: when two or more independent sources are present, establish the timeline BEFORE anything else — course_timeline.py <course_dir> for a course folder with a manifest (it writes _seg/real_timeline.json, maps photos onto the recordings, and with --reorder-manifest fixes clip order at the root), or media_capture_index.py --emit-alignment alignment.json for a loose material folder. ==Capture timestamps are HYPOTHESES; transcript cross-correlation (xcorr_media_offsets.py) is EVIDENCE.== A source whose reliable flag is false got its start from mtime or has none — it must not be aligned on. Nothing is ever auto-corrected: a claimed-vs-measured disagreement >5 s is flagged "conflict": true for a human to judge. Details in reference/multi-camera.md.

Pipeline

One command plus its purpose per step; flags, thresholds and outputs are in reference/pipeline.md.

Step 1 — Ask the language (mandatory, no command)

English / Mandarin / bilingual? Accented speakers? Code-switching mid-sentence? Use AskUserQuestion if the user has not said. HARD RULE 1.

Step 2 — Set up the lecture directory

One directory per lecture holds every intermediate; name it {date}_{speaker}_{topic}, the shape finalize_to_vault.py parses.

Step 3 — GPU pre-flight
bash
python <skill-dir>/scripts/gpu_check.py --out-dir "$OUT_DIR" --min-free-mb 6000

Gate before transcription and again before Stage D. Exit 0 proceed, 1 warn and proceed, 2 blocked — surface it, ==do not retry in a loop==. A card whose total VRAM is under the threshold (2–4 GB laptops) is GPU_TOO_SMALL, exit 0: not contention, nothing will free up — proceed with the CPU path (transcribe_video.py --device cpu --model small, or --engine groq). → reference/pipeline.md#gpu-check

Step 4 — Transcribe
bash
python <skill-dir>/scripts/transcribe_video.py "<media>" \
    --output-dir "$OUT_DIR" --lang <zh|en|bilingual|auto> \
    --batch-size 3 --beam-size 10

Local faster-whisper by default; --engine groq is an optional offload (HARD RULE 9). Default model alias is breeze25 (needs a local model dir); on a machine without one, pass --model large-v3, which faster-whisper downloads. Recordings over ~30 min go through the chunked runner instead. → reference/pipeline.md#transcription

Step 5 — Stage A: frame extraction (Path A only)
bash
python <skill-dir>/scripts/extract_slides.py "<video>" --output-dir "$OUT_DIR" --interval 15

Writes slides/frame_NNNN.jpg + slides/timestamps.json, phash-deduping adjacent near-identical frames. Path B/B-images skip this. → reference/pipeline.md#stage-a

Step 6 — Stage B: quick OCR + entropy (Path A only)
bash
python <skill-dir>/scripts/quick_ocr.py "$OUT_DIR"

RapidOCR on every frame → slides_raw.json. ==Required==: without it every slide looks decorative to the Stage D gate. → reference/pipeline.md#stage-b

Step 6-alt — Path B / B-images bridge
bash
python <skill-dir>/scripts/build_slides_from_pdf.py    "$OUT_DIR"   [--audio-duration-sec N]
python <skill-dir>/scripts/build_slides_from_images.py "<img_dir>" -o "$OUT_DIR" [--audio-duration-sec N]

Either bridge emits slides_raw.json + slides_dedup.json directly, replacing Steps 5–7. → reference/pipeline.md#path-b

Step 7 — Stage C: semantic dedup (Path A only)
bash
python <skill-dir>/scripts/dedup_semantic.py "$OUT_DIR"

Merges adjacent frames by text-subset or layout similarity, marks dedup.is_canonical. Output slides_dedup.json. → reference/pipeline.md#stage-c

Step 8 — Stage B2: high-quality OCR (Surya)
bash
python <skill-dir>/scripts/ocr_surya.py "$OUT_DIR" [--resume]

Surya in its own venv on canonical text-bearing slides, RapidOCR as the shallow fallback. Adds ocr.clean_text / ocr_engine / ocr_confidence. ==Updates slides_dedup.json in place== (one-time backup slides_dedup.pre_b2.json) and writes slides_ocr.json. Path B skips it — pdf_text is already clean. Without a Surya venv it warns and routes everything to RapidOCR rather than failing. → reference/pipeline.md#stage-b2

Step 9 — Stage D: VLM signals
bash
python <skill-dir>/scripts/vlm_signals.py "$OUT_DIR" --model minicpm-v:8b --num-ctx 4096

Semantic signals per canonical slide, behind a 4-condition pre-skip gate for decorative frames. Re-check the GPU first (Step 3). Output slides_vlm.json. scripts/ocr_slides.py is a deprecated shim forwarding here, same argv and outputs. → reference/pipeline.md#stage-d

Step 10 — Stage E: transcript grounding
bash
python <skill-dir>/scripts/ground_slides.py "$OUT_DIR"

Pure Python, 0 LLM calls. Ties each canonical slide to the words spoken over it. Output slides_grounded.json — the input to synthesis. → reference/pipeline.md#stage-e

Step 11 — Flag suspect ASR tokens
bash
python <skill-dir>/scripts/flag_asr_suspects.py --dir "$OUT_DIR"

Runs HERE, after Stage E: the slide glossary it needs comes from slides_grounded.json. Writes asr_suspects.txt; ==the transcript is left byte-identical==. Treat each line as a question, never a substitution. → reference/pipeline.md#asr-suspects

Show full SKILL.md (766 more words)Show less
Step 12 — Chunked pre-summarization (long lectures only)

Over ~30 min / 25 k tokens of transcript, offload chunk summaries to a Sonnet subagent instead of reading the whole transcript into main context. Coverage guards ([CHUNK_END], [CONTINUE_NEEDED], expected-chunk count) are mandatory. → reference/pipeline.md#chunked-summarization

Step 13 — Stage F: synthesis (Claude)

Two passes for batches and long lectures — ==Tier-pass then Write-pass==:

  • Tier-pass subagent reads slides_grounded.json + transcript.txt + pdf_text.json, applies the tier scoring rules, writes only slides_final.json (integer tier, attachment_name, embed_width, section_suggestion). This file is the frozen tier authority.
  • Write-pass subagent reads the frozen slides_final.json + transcript + slide text, writes note_draft.md with [[EMBED sN]] placeholders only — no paths, widths or callouts.

One pass is fine for one short lecture; splitting them stops the writer from simplifying structure to make its own embed audit pass. → reference/note-spec.md (mandatory: quality spec, tier rules, prompt requirements)

Step 14 — Render, finalize, audit
bash
python <skill-dir>/scripts/render_embeds.py    "$OUT_DIR" --note note_draft.md --in-place
python <skill-dir>/scripts/finalize_to_vault.py "$OUT_DIR" [--vault-root PATH]
python <skill-dir>/scripts/audit_note.py "<note path>" --mode lecture --grounding "$OUT_DIR"

render_embeds.py expands placeholders to col-0 callouts with path + width and audits Tier-1/2 coverage; finalize_to_vault.py copies cited slides + the note into the vault; the auditor is the gate. ==Always pass --grounding== — without it the caption↔frame check only warns. → reference/note-spec.md Draft review exemption: this output is machine-transcribed and synthesized — write to the inbox without showing a draft; the user reviews in Obsidian.

Step 15 — Web viewer export (default whenever there is video)
bash
python <skill-dir>/scripts/build_single_talk_web.py "$OUT_DIR" --plan seg_plan.json
# … Stage F writes the L3/ files … then:
python <skill-dir>/scripts/build_single_talk_web.py "$OUT_DIR" --plan seg_plan.json --export

A single talk gets the synced HTML viewer by default, not just workshops. ==Stage F authors the segment plan itself== (JSON list: seg/start/end/slug/title_zh, derived from the transcript + slide topics — no human segmentation needed); ==content sections = more segments, not more headings==, the viewer builds exactly two chapters per segment. The script assembles manifest + segments.json + L2 slices + HUB skeleton from transcript.json, and ==auto-appends an L3-only 全場總整理 overview segment== (sorted first — many segments still need one place listing everything; the HUB never renders in the viewer). Stage F then writes one L3_segNN_<slug>.md per segment (==image embeds by basename only, timecodes `(V1 MM:SS)`==) plus the overview note — content shape picked by ==one question: is this a single body of knowledge?== (single talk → pearls + per-segment one-liners; same-lecturer arc → thematic reorganization with sources; multi-speaker workshop → ==catalog only, never forced synthesis== — spec in reference/pipeline.md#web-export); --export refuses while any L3 is missing. Compression default is ==H.265 CRF 24== (self-use; --codec h264 when sharing to machines you can't verify). Multi-talk workshops keep their own flow → reference/segmented-mode.md. → reference/pipeline.md#web-export

Edge cases

  • Audio only, no deck → transcript-only note using # 逐段筆記 instead of # 逐投影片筆記.
  • N-up handout PDF → render one mid page and ==look at it== before deciding the grid; heuristics are unreliable on slide-heavy PDFs. Then crop_multiup_pdf.py <pdf> <out_dir> --expected-rows R --expected-cols C. Pages that are genuinely 1-up (title pages) are handled per page, not forced into the consensus grid.
  • Very long lecture (>90 min) → chunked runner for transcription, Step 12 for reading it. Batch of recordings → transcribe strictly sequentially; ~4.5 GB RAM per faster-whisper instance.
  • One talk split across several files → one note, not several.
  • Slides English, speaker Mandarin → keep both; the deck gives terms, the transcript gives the explanation.
  • Speaker asked not to be recorded → exclude that content.
  • Dense text handout, not a slide deck → primary source, but drop the slide-by-slide structure.
  • CJK path failures (exit 127 / 3221226505) → extract audio with ffmpeg separately first; the script reuses a validated audio.wav.

Optional infrastructure

Everything below is ==this machine's setup, not a requirement== — nothing in the pipeline depends on any of it, and the generic alternative is inline.

Used forGeneric alternative
job_runner.py wrapping long GPU jobs (tree-kills children on timeout)plain timeout <n> <cmd>, or run in the foreground
gpu_lease.py / a pause-flag file between concurrent batchesrun GPU stages one at a time; leave paths.pause_flag empty in config
vault-search / OpenEvidence / Zotero lookups during synthesisskip; cite only what the lecture itself provided
ntfy completion pingsskip
external batch control plane (run_queue, rerun_batch, clip_order, dashboard)course-specific, not shipped — see reference/segmented-mode.md

Vault paths (99Attachment/lecture_{slug}, the inbox folder) are ==a private vault convention== and configurable: render_embeds.py --attach-root / --attach-dir, finalize_to_vault.py --vault-root.

Dependencies

bash
ollama pull minicpm-v:8b               # ~5.5 GB Q4_K_M, Stage D
pip install rapidfuzz rapidocr-onnxruntime scikit-image pyyaml json_repair pillow numpy

ffmpeg + ffprobe on PATH. Surya (Stage B2) lives in its own venv; point ocr_engine.surya_python at it in config.yaml, or leave it blank to fall back to RapidOCR. Copy config.example.yaml → config.yaml on a new machine; every machine-specific value there is blank by default and env-overridable. ==Optional dependencies degrade loudly, not silently== — a missing scikit-image or RapidOCR is reported and gated, because a silent degrade produced wrong output rather than less output (reference/decisions.md#optional-dependency-degradation).

File locations

<skill-dir>/
├── SKILL.md, config.yaml, config.example.yaml
├── reference/  pipeline.md · note-spec.md · segmented-mode.md · multi-camera.md · decisions.md
├── data/       real_words.txt, real_acronyms.txt   (regenerated, not committed)
├── ocr_bench/  engine benchmark harness (bring your own fixtures)
└── scripts/
    ├── route_inputs.py           front door — classifies material, prints the plan
    ├── transcribe_video.py  retranscribe_segment.py  gpu_check.py  groq_asr.py
    ├── extract_slides.py  quick_ocr.py  dedup_semantic.py  ocr_surya.py
    ├── vlm_signals.py            (ocr_slides.py = deprecated shim → here)
    ├── build_slides_from_pdf.py  build_slides_from_images.py  crop_multiup_pdf.py
    ├── ground_slides.py  flag_asr_suspects.py  make_glossary.py  build_real_words.py
    ├── render_embeds.py  finalize_to_vault.py  audit_note.py  export_web.py
    ├── media_capture_index.py  xcorr_media_offsets.py  query_near_field.py
    ├── _mdpm.py                  AVCHD .MTS recording clock (no container tag)
    ├── adapters/    surya_adapter.py        (production OCR adapters)
    ├── batch/       course_timeline (real-time ordering authority) ·
    │                audit_segmentation (is the proposal still valid?) ·
    │                build_L1 · split_segments · split_L1_by_segment · add_dhash ·
    │                vlm_cache · detect_language · detect_language_audio ·
    │                phi_mask · process_slide_deck     (generic batch layer)
    ├── layout2/     viewer.css, viewer.js   (web viewer assets, edited verbatim)
    └── _common.py  _log.py  _paths.py

Per-lecture output directory:

{lecture}/
├── metadata.json          run_id + media fingerprint + per-stage status
├── transcript.json/.txt   timestamped segments; .txt is [MM:SS] text, H:MM:SS past an hour
├── asr_suspects.txt       flagged tokens — flags only, never a rewrite
├── alignment.json         multi-source capture-start hypotheses (when applicable)
├── slides/                frame_NNNN.jpg | page_NN.jpg | original photo names
├── slides_raw.json        Stage B    quick text + density + entropy
├── slides_dedup.json      Stage C    canonical markers (Stage B2 updates in place)
├── slides_dedup.pre_b2.json          one-time pre-Stage-B2 snapshot
├── slides_ocr.json        Stage B2   Surya result, for inspection
├── slides_vlm.json        Stage D    VLM signals + vlm_skip + skip_metrics
├── slides_grounded.json   Stage E    transcript grounding + retrieval fields
├── slides_final.json      Stage F    tier + score + attachment_name + width
└── logs/progress_*.jsonl  per-stage event streams (not every stage emits one)

Each stage output is a superset of the previous, so you can re-run one stage without redoing transcription or frame extraction. runs.jsonl one level up carries one summary line per stage run, joined by run_id. Keep the intermediates — they are how tier decisions get debugged.

© drpwchen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 86 other files (scripts) in the repository root of drpwchen/lecture-to-notes.

  • SKILL.md
  • .gitattributes
  • .github/workflows/secret-scan.yml
  • .gitignore
  • CHANGELOG.md
  • LICENSE
  • README.md
  • README.zh-TW.md
  • SYNC-LEDGER.md
  • config.example.yaml
  • data/.gitignore
  • data/real_acronyms.txt
  • data/real_words.txt
  • docs/AUDIT_SUMMARY.md
  • docs/assets/hero.png
  • docs/assets/hero.zh.png
  • … and 71 more

Open the folder on GitHubat commit 79053a3

Compare with similar skills

Lecture To Notes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Lecture To Notes compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Lecture To Notes this skilldrpwchen/lecture-to-notes108—~4.9kAutomated safety check: NotesMIT
Generate Doc Templatejeremylongshore/tons-of-skills-marketplace2.8k—~1.3kAutomated safety check: PassMIT
Youtube Notetakersickn33/agentic-awesome-skills47k1 repos~2.3kAutomated safety check: PassMIT
Multimodal Extractionswyxio/skills175—~922Automated safety check: PassMIT
Explain Videolimin112/min-skill412—~2.4kAutomated safety check: PassNone
Orange Line Illustrationorange2ai/orange-line-illustration440—~2.5kAutomated safety check: PassProprietary

Similar skills

  • Generate Doc Template

    jeremylongshore/tons-of-skills-marketplace

    Generates on-brand document and deck templates — letterhead, slide, and one-pager — as SVG from the active brand profile, with editable title, subtitle, and body zones.

    2.8k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Youtube Notetaker

    sickn33/agentic-awesome-skills

    Turn YouTube talks into local study notes with slides, transcripts, editable annotations, and a markdown-backed viewer.

    47k GitHub starsUsed in 1 repo~2.3k tokens
    Documents & OfficeAuto-check passed
  • Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the…

    175 GitHub stars~922 tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Explain Video

    limin112/min-skill

    Build a narrated explainer video from a concept — discussion → structure → HTML slide deck → narration script → TTS voice → subtitles → background music → Playwright screen recording → ffmpeg…

    412 GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Orange Line Illustration

    orange2ai/orange-line-illustration

    Generate New Yorker-style minimalist editorial illustrations — thin black ink lines on white, vast negative space, a single orange accent (F97316) — for articles, principles, covers, concept…

    440 GitHub stars~2.5k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Strategy Consulting Visualization

    kgraph57/mckinsey-style-visualization-skill

    A skill your agent uses when turning any content into clear, professional visualizations - board slides, reports, proposals, research summaries, training materials, technical diagrams, infographics…

    128 GitHub stars~3.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

Questions about Lecture To Notes

What does Lecture To Notes do?

Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or…. Lecture To Notes is an agent skill from drpwchen/lecture-to-notes. Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or a dropped media file.

When should I use Lecture To Notes?

Lecture To Notes fits situations like: tasks that involve Transcription; tasks that involve Slides and decks.

How do I install Lecture To Notes in Claude Code?

Run `npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a claude-code`. Or copy the skill folder (the drpwchen/lecture-to-notes repository) into .claude/skills/lecture-to-notes in your project. Claude Code loads it when a task matches its description.

How do I install Lecture To Notes in Codex?

Run `npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a codex`. Or copy the skill folder (the drpwchen/lecture-to-notes repository) into .agents/skills/lecture-to-notes in your project. Codex loads it when a task matches its description.

Can I use Lecture To Notes in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lecture-to-notes, .gemini/skills/lecture-to-notes, .github/skills/lecture-to-notes and .opencode/skills/lecture-to-notes in your project.

What does Lecture To Notes need to run?

Going by SKILL.md and its folder, Lecture To Notes needs the command-line tools its instructions call (python, ollama and pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Glob, Grep, Agent.

Does Lecture To Notes access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Lecture To Notes safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Lecture To Notes use?

Lecture To Notes is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Lecture To Notes use?

About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Lecture To Notes?

Skills that share tags, products or a category with Lecture To Notes: Generate Doc Template (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Youtube Notetaker (sickn33/agentic-awesome-skills, 47k stars), Multimodal Extraction (swyxio/skills, 175 stars) and Explain Video (limin112/min-skill, 412 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Lecture To Notes?

drpwchen (a GitHub user) maintains it in drpwchen/lecture-to-notes, which has 108 GitHub stars. The repository was last updated on September 1, 2026.

Source: drpwchen/lecture-to-notes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.