Generate Doc Template
jeremylongshore/tons-of-skills-marketplace
Generates on-brand document and deck templates — letterhead, slide, and one-pager — as SVG from the active brand profile, with editable title, subtitle, and body zones.
Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or…
$ npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install drpwchen/lecture-to-notes lecture-to-notes --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "lecture-to-notes" agent skill from https://github.com/drpwchen/lecture-to-notes/tree/main into .claude/skills/lecture-to-notes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lecture-to-notes", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install drpwchen/lecture-to-notes lecture-to-notes --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "lecture-to-notes" agent skill from https://github.com/drpwchen/lecture-to-notes/tree/main into .agents/skills/lecture-to-notes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lecture-to-notes", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install drpwchen/lecture-to-notes lecture-to-notes --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "lecture-to-notes" agent skill from https://github.com/drpwchen/lecture-to-notes/tree/main into .cursor/skills/lecture-to-notes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lecture-to-notes", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install drpwchen/lecture-to-notes lecture-to-notes --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "lecture-to-notes" agent skill from https://github.com/drpwchen/lecture-to-notes/tree/main into .gemini/skills/lecture-to-notes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lecture-to-notes", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install drpwchen/lecture-to-notes lecture-to-notesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "lecture-to-notes" agent skill from https://github.com/drpwchen/lecture-to-notes/tree/main into .github/skills/lecture-to-notes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lecture-to-notes", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install drpwchen/lecture-to-notes lecture-to-notes --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "lecture-to-notes" agent skill from https://github.com/drpwchen/lecture-to-notes/tree/main into .opencode/skills/lecture-to-notes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lecture-to-notes", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
lecture-to-notesTurn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or…
Lecture To Notes is an agent skill from drpwchen/lecture-to-notes. Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or a dropped media file. Handles batch runs.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 91 other files, including scripts (for example `.github/workflows/secret-scan.yml`, `CHANGELOG.md` and `README.md`).
It sits in Media & Creative, covering Transcription and Slides and decks. The repository describes itself as: Lecture recordings → structured grounded notes + a synced HTML viewer: video, timestamped transcript and curated summary on one page. Local GPU pipeline (Whisper ASR · slide… The licence is MIT.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 79053a3. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBashGlobGrepAgentFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
pythonollamapipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Lecture To Notes loads about 4.9k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 1,913 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Write, Edit, Bash, Glob, Grep, AgentAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from drpwchen/lecture-to-notes at commit 79053a3, republished under its MIT licence (© drpwchen). 1,913 words, ~4,914 tokens.
.claude/skills/lecture-to-notes/SKILL.md (or your agent's skills folder). This skill also uses 86 other files; get the full folder from GitHub.Turn a lecture recording (video or audio-only) into structured notes. Every heavy
stage runs locally at 0 Claude tokens; Claude only does the final synthesis. This
page is the map; detail lives in reference/, one topic per file.
| File | What is in it |
|---|---|
reference/pipeline.md | Per-stage flags, thresholds, JSON schemas, timeouts, observability |
reference/note-spec.md | Note quality spec, tier scoring, width table, synthesis prompt requirements |
reference/segmented-mode.md | Multi-talk workshop folders → per-segment L2/L3 + Hub + web viewer |
reference/multi-camera.md | One long recording + many phone clips/photos → one timeline |
reference/decisions.md | Post-mortems, benchmarks, wrong turns, VRAM measurements |
transcribe_video.py exits without --lang. A wrong guess makes Whisper
hallucinate Chinese from accented English and the transcript is unusable.reference/decisions.md#asr-auto-correction. Do not add an auto-apply mode.transcribe_video.py auto-runs
retranscribe_segment.py --auto on detected token-collapse. If collapses
survive that, escalate (wider beam + --no-repeat-ngram-size + a glossary),
never skip.quick_text, Stage B2 clean_text, or pdf_text.json.
ocr.vlm_text is an always-empty compatibility field.--batch-size 4 with --beam-size 10==, and
never combine --beam-size 15 with sequential mode — that crashes
(0xC0000005). The measured sweet spot is --batch-size 3 --beam-size 10.slides_grounded.json exists. A single subagent bounces during the long
GPU waits and burns 30+ min of wall time per lecture.00000), a recorder's
240526_1119.mp3 sorts into the middle of the video files, and on-site
-1-/-2- labels get stuck on the wrong file. Run
scripts/batch/course_timeline.py BEFORE segmenting any multi-source course;
manifest.json clip order and L1_coarse.md section order must both be
built from it. An agenda is not a clock either — the 2024-05 Conference-Y
conference ran ~25 min early on day 1 and ~35 min late on day 2, while its
break gaps matched to the minute.--engine groq. When unsure, ask; default to local.==Start here. route_inputs.py is the front door== — it classifies a folder and
prints the ordered commands plus the questions a human must answer. It is
plan-only: it never runs anything and never writes a file.
python <skill-dir>/scripts/route_inputs.py <material_dir> [--recursive] [--out-dir DIR] [--json]| What is in the folder | Slide source | Route |
|---|---|---|
| Video, no deck | frames from the video | Path A — Steps 5–7 |
| Audio/video + PDF deck (==preferred==) | PDF text + page renders | Path B — build_slides_from_pdf.py |
| Audio/video + loose slide images (≥3) | the images themselves | Path B-images — build_slides_from_images.py |
| Audio only, no deck | none | Path C — transcript-only note |
| N-up handout PDF | cropped tiles | Path B-multi — crop_multiup_pdf.py first |
| Multi-talk workshop folder | per segment | reference/segmented-mode.md |
| One long recording + many phone clips/photos | per source | reference/multi-camera.md |
.pptx / .docx / .key | — | convert to PDF yourself first; there is no conversion step here |
==Multi-source contract==: when two or more independent sources are present,
establish the timeline BEFORE anything else — course_timeline.py <course_dir>
for a course folder with a manifest (it writes _seg/real_timeline.json, maps
photos onto the recordings, and with --reorder-manifest fixes clip order at
the root), or media_capture_index.py --emit-alignment alignment.json for a
loose material folder.
==Capture timestamps are HYPOTHESES; transcript cross-correlation
(xcorr_media_offsets.py) is EVIDENCE.== A source whose reliable flag is false
got its start from mtime or has none — it must not be aligned on. Nothing is ever
auto-corrected: a claimed-vs-measured disagreement >5 s is flagged
"conflict": true for a human to judge. Details in reference/multi-camera.md.
One command plus its purpose per step; flags, thresholds and outputs are in
reference/pipeline.md.
English / Mandarin / bilingual? Accented speakers? Code-switching mid-sentence? Use AskUserQuestion if the user has not said. HARD RULE 1.
One directory per lecture holds every intermediate; name it
{date}_{speaker}_{topic}, the shape finalize_to_vault.py parses.
python <skill-dir>/scripts/gpu_check.py --out-dir "$OUT_DIR" --min-free-mb 6000Gate before transcription and again before Stage D. Exit 0 proceed, 1 warn
and proceed, 2 blocked — surface it, ==do not retry in a loop==. A card whose
total VRAM is under the threshold (2–4 GB laptops) is GPU_TOO_SMALL, exit
0: not contention, nothing will free up — proceed with the CPU path
(transcribe_video.py --device cpu --model small, or --engine groq).
→ reference/pipeline.md#gpu-check
python <skill-dir>/scripts/transcribe_video.py "<media>" \
--output-dir "$OUT_DIR" --lang <zh|en|bilingual|auto> \
--batch-size 3 --beam-size 10Local faster-whisper by default; --engine groq is an optional offload (HARD
RULE 9). Default model alias is breeze25 (needs a local model dir); on a
machine without one, pass --model large-v3, which faster-whisper downloads.
Recordings over ~30 min go through the chunked runner instead.
→ reference/pipeline.md#transcription
python <skill-dir>/scripts/extract_slides.py "<video>" --output-dir "$OUT_DIR" --interval 15Writes slides/frame_NNNN.jpg + slides/timestamps.json, phash-deduping adjacent
near-identical frames. Path B/B-images skip this. → reference/pipeline.md#stage-a
python <skill-dir>/scripts/quick_ocr.py "$OUT_DIR"RapidOCR on every frame → slides_raw.json. ==Required==: without it every slide
looks decorative to the Stage D gate. → reference/pipeline.md#stage-b
python <skill-dir>/scripts/build_slides_from_pdf.py "$OUT_DIR" [--audio-duration-sec N]
python <skill-dir>/scripts/build_slides_from_images.py "<img_dir>" -o "$OUT_DIR" [--audio-duration-sec N]Either bridge emits slides_raw.json + slides_dedup.json directly, replacing
Steps 5–7. → reference/pipeline.md#path-b
python <skill-dir>/scripts/dedup_semantic.py "$OUT_DIR"Merges adjacent frames by text-subset or layout similarity, marks
dedup.is_canonical. Output slides_dedup.json.
→ reference/pipeline.md#stage-c
python <skill-dir>/scripts/ocr_surya.py "$OUT_DIR" [--resume]Surya in its own venv on canonical text-bearing slides, RapidOCR as the shallow
fallback. Adds ocr.clean_text / ocr_engine / ocr_confidence. ==Updates
slides_dedup.json in place== (one-time backup slides_dedup.pre_b2.json) and
writes slides_ocr.json. Path B skips it — pdf_text is already clean. Without
a Surya venv it warns and routes everything to RapidOCR rather than failing.
→ reference/pipeline.md#stage-b2
python <skill-dir>/scripts/vlm_signals.py "$OUT_DIR" --model minicpm-v:8b --num-ctx 4096Semantic signals per canonical slide, behind a 4-condition pre-skip gate for
decorative frames. Re-check the GPU first (Step 3). Output slides_vlm.json.
scripts/ocr_slides.py is a deprecated shim forwarding here, same argv and
outputs. → reference/pipeline.md#stage-d
python <skill-dir>/scripts/ground_slides.py "$OUT_DIR"Pure Python, 0 LLM calls. Ties each canonical slide to the words spoken over it.
Output slides_grounded.json — the input to synthesis.
→ reference/pipeline.md#stage-e
python <skill-dir>/scripts/flag_asr_suspects.py --dir "$OUT_DIR"Runs HERE, after Stage E: the slide glossary it needs comes from
slides_grounded.json. Writes asr_suspects.txt; ==the transcript is left
byte-identical==. Treat each line as a question, never a substitution.
→ reference/pipeline.md#asr-suspects
Over ~30 min / 25 k tokens of transcript, offload chunk summaries to a Sonnet
subagent instead of reading the whole transcript into main context. Coverage
guards ([CHUNK_END], [CONTINUE_NEEDED], expected-chunk count) are mandatory.
→ reference/pipeline.md#chunked-summarization
Two passes for batches and long lectures — ==Tier-pass then Write-pass==:
slides_grounded.json + transcript.txt +
pdf_text.json, applies the tier scoring rules, writes only
slides_final.json (integer tier, attachment_name, embed_width,
section_suggestion). This file is the frozen tier authority.slides_final.json + transcript +
slide text, writes note_draft.md with [[EMBED sN]] placeholders only — no
paths, widths or callouts.One pass is fine for one short lecture; splitting them stops the writer from
simplifying structure to make its own embed audit pass. → reference/note-spec.md
(mandatory: quality spec, tier rules, prompt requirements)
python <skill-dir>/scripts/render_embeds.py "$OUT_DIR" --note note_draft.md --in-place
python <skill-dir>/scripts/finalize_to_vault.py "$OUT_DIR" [--vault-root PATH]
python <skill-dir>/scripts/audit_note.py "<note path>" --mode lecture --grounding "$OUT_DIR"render_embeds.py expands placeholders to col-0 callouts with path + width and
audits Tier-1/2 coverage; finalize_to_vault.py copies cited slides + the note
into the vault; the auditor is the gate. ==Always pass --grounding== — without
it the caption↔frame check only warns. → reference/note-spec.md
Draft review exemption: this output is machine-transcribed and synthesized —
write to the inbox without showing a draft; the user reviews in Obsidian.
python <skill-dir>/scripts/build_single_talk_web.py "$OUT_DIR" --plan seg_plan.json
# … Stage F writes the L3/ files … then:
python <skill-dir>/scripts/build_single_talk_web.py "$OUT_DIR" --plan seg_plan.json --exportA single talk gets the synced HTML viewer by default, not just workshops.
==Stage F authors the segment plan itself== (JSON list:
seg/start/end/slug/title_zh, derived from the transcript + slide topics — no
human segmentation needed); ==content sections = more segments, not more
headings==, the viewer builds exactly two chapters per segment. The script assembles manifest + segments.json + L2 slices + HUB
skeleton from transcript.json, and ==auto-appends an L3-only 全場總整理
overview segment== (sorted first — many segments still need one place listing
everything; the HUB never renders in the viewer). Stage F then writes one
L3_segNN_<slug>.md per segment (==image embeds by basename only, timecodes
`(V1 MM:SS)`==) plus the overview note — content shape picked by ==one
question: is this a single body of knowledge?== (single talk → pearls +
per-segment one-liners; same-lecturer arc → thematic reorganization with
sources; multi-speaker workshop → ==catalog only, never forced synthesis== —
spec in reference/pipeline.md#web-export); --export refuses while any L3
is missing.
Compression default is ==H.265 CRF 24== (self-use; --codec h264 when sharing
to machines you can't verify). Multi-talk workshops keep their own flow →
reference/segmented-mode.md. → reference/pipeline.md#web-export
# 逐段筆記 instead of
# 逐投影片筆記.crop_multiup_pdf.py <pdf> <out_dir> --expected-rows R --expected-cols C.
Pages that are genuinely 1-up (title pages) are handled per page, not forced
into the consensus grid.audio.wav.Everything below is ==this machine's setup, not a requirement== — nothing in the pipeline depends on any of it, and the generic alternative is inline.
| Used for | Generic alternative |
|---|---|
job_runner.py wrapping long GPU jobs (tree-kills children on timeout) | plain timeout <n> <cmd>, or run in the foreground |
gpu_lease.py / a pause-flag file between concurrent batches | run GPU stages one at a time; leave paths.pause_flag empty in config |
vault-search / OpenEvidence / Zotero lookups during synthesis | skip; cite only what the lecture itself provided |
ntfy completion pings | skip |
external batch control plane (run_queue, rerun_batch, clip_order, dashboard) | course-specific, not shipped — see reference/segmented-mode.md |
Vault paths (99Attachment/lecture_{slug}, the inbox folder) are ==a private vault
convention== and configurable: render_embeds.py --attach-root / --attach-dir,
finalize_to_vault.py --vault-root.
ollama pull minicpm-v:8b # ~5.5 GB Q4_K_M, Stage D
pip install rapidfuzz rapidocr-onnxruntime scikit-image pyyaml json_repair pillow numpyffmpeg + ffprobe on PATH. Surya (Stage B2) lives in its own venv; point
ocr_engine.surya_python at it in config.yaml, or leave it blank to fall back
to RapidOCR. Copy config.example.yaml → config.yaml on a new machine; every
machine-specific value there is blank by default and env-overridable.
==Optional dependencies degrade loudly, not silently== — a missing scikit-image or
RapidOCR is reported and gated, because a silent degrade produced wrong output
rather than less output (reference/decisions.md#optional-dependency-degradation).
<skill-dir>/
├── SKILL.md, config.yaml, config.example.yaml
├── reference/ pipeline.md · note-spec.md · segmented-mode.md · multi-camera.md · decisions.md
├── data/ real_words.txt, real_acronyms.txt (regenerated, not committed)
├── ocr_bench/ engine benchmark harness (bring your own fixtures)
└── scripts/
├── route_inputs.py front door — classifies material, prints the plan
├── transcribe_video.py retranscribe_segment.py gpu_check.py groq_asr.py
├── extract_slides.py quick_ocr.py dedup_semantic.py ocr_surya.py
├── vlm_signals.py (ocr_slides.py = deprecated shim → here)
├── build_slides_from_pdf.py build_slides_from_images.py crop_multiup_pdf.py
├── ground_slides.py flag_asr_suspects.py make_glossary.py build_real_words.py
├── render_embeds.py finalize_to_vault.py audit_note.py export_web.py
├── media_capture_index.py xcorr_media_offsets.py query_near_field.py
├── _mdpm.py AVCHD .MTS recording clock (no container tag)
├── adapters/ surya_adapter.py (production OCR adapters)
├── batch/ course_timeline (real-time ordering authority) ·
│ audit_segmentation (is the proposal still valid?) ·
│ build_L1 · split_segments · split_L1_by_segment · add_dhash ·
│ vlm_cache · detect_language · detect_language_audio ·
│ phi_mask · process_slide_deck (generic batch layer)
├── layout2/ viewer.css, viewer.js (web viewer assets, edited verbatim)
└── _common.py _log.py _paths.pyPer-lecture output directory:
{lecture}/
├── metadata.json run_id + media fingerprint + per-stage status
├── transcript.json/.txt timestamped segments; .txt is [MM:SS] text, H:MM:SS past an hour
├── asr_suspects.txt flagged tokens — flags only, never a rewrite
├── alignment.json multi-source capture-start hypotheses (when applicable)
├── slides/ frame_NNNN.jpg | page_NN.jpg | original photo names
├── slides_raw.json Stage B quick text + density + entropy
├── slides_dedup.json Stage C canonical markers (Stage B2 updates in place)
├── slides_dedup.pre_b2.json one-time pre-Stage-B2 snapshot
├── slides_ocr.json Stage B2 Surya result, for inspection
├── slides_vlm.json Stage D VLM signals + vlm_skip + skip_metrics
├── slides_grounded.json Stage E transcript grounding + retrieval fields
├── slides_final.json Stage F tier + score + attachment_name + width
└── logs/progress_*.jsonl per-stage event streams (not every stage emits one)Each stage output is a superset of the previous, so you can re-run one stage
without redoing transcription or frame extraction. runs.jsonl one level up
carries one summary line per stage run, joined by run_id. Keep the intermediates
— they are how tier decisions get debugged.
© drpwchen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 86 other files (scripts) in the repository root of drpwchen/lecture-to-notes.
Open the folder on GitHubat commit 79053a3
Lecture To Notes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Lecture To Notes this skilldrpwchen/lecture-to-notes | 108 | — | ~4.9k | Automated safety check: Notes | MIT | |
| Generate Doc Templatejeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Youtube Notetakersickn33/agentic-awesome-skills | 47k | 1 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Multimodal Extractionswyxio/skills | 175 | — | ~922 | Automated safety check: Pass | MIT | |
| Explain Videolimin112/min-skill | 412 | — | ~2.4k | Automated safety check: Pass | None | |
| Orange Line Illustrationorange2ai/orange-line-illustration | 440 | — | ~2.5k | Automated safety check: Pass | Proprietary |
jeremylongshore/tons-of-skills-marketplace
Generates on-brand document and deck templates — letterhead, slide, and one-pager — as SVG from the active brand profile, with editable title, subtitle, and body zones.
sickn33/agentic-awesome-skills
Turn YouTube talks into local study notes with slides, transcripts, editable annotations, and a markdown-backed viewer.
swyxio/skills
Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the…
limin112/min-skill
Build a narrated explainer video from a concept — discussion → structure → HTML slide deck → narration script → TTS voice → subtitles → background music → Playwright screen recording → ffmpeg…
orange2ai/orange-line-illustration
Generate New Yorker-style minimalist editorial illustrations — thin black ink lines on white, vast negative space, a single orange accent (F97316) — for articles, principles, covers, concept…
kgraph57/mckinsey-style-visualization-skill
A skill your agent uses when turning any content into clear, professional visualizations - board slides, reports, proposals, research summaries, training materials, technical diagrams, infographics…
Categories
Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or…. Lecture To Notes is an agent skill from drpwchen/lecture-to-notes. Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or a dropped media file.
Lecture To Notes fits situations like: tasks that involve Transcription; tasks that involve Slides and decks.
Run `npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a claude-code`. Or copy the skill folder (the drpwchen/lecture-to-notes repository) into .claude/skills/lecture-to-notes in your project. Claude Code loads it when a task matches its description.
Run `npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a codex`. Or copy the skill folder (the drpwchen/lecture-to-notes repository) into .agents/skills/lecture-to-notes in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add drpwchen/lecture-to-notes --skill lecture-to-notes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lecture-to-notes, .gemini/skills/lecture-to-notes, .github/skills/lecture-to-notes and .opencode/skills/lecture-to-notes in your project.
Going by SKILL.md and its folder, Lecture To Notes needs the command-line tools its instructions call (python, ollama and pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Glob, Grep, Agent.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Lecture To Notes is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Lecture To Notes: Generate Doc Template (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Youtube Notetaker (sickn33/agentic-awesome-skills, 47k stars), Multimodal Extraction (swyxio/skills, 175 stars) and Explain Video (limin112/min-skill, 412 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
drpwchen (a GitHub user) maintains it in drpwchen/lecture-to-notes, which has 108 GitHub stars. The repository was last updated on September 1, 2026.
Source: drpwchen/lecture-to-notes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.