Agent skill

Tutorial Videos

by matthiasn in matthiasn/lotti

Build, debug, extend, localize, or publish the automated tutorial videos — the workbench that drives the real Linux app under Xvfb (virtual mic, live Voxtral transcription, task-agent proposals)…

GPL-3.0Auto-check: notesMedia & Creative

Install Tutorial Videos

skills CLI
$ npx skills add matthiasn/lotti --skill tutorial-videos -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install matthiasn/lotti tutorial-videos --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/matthiasn/lotti.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tutorial-videos .claude/skills/tutorial-videos && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tutorial-videos
GitHub stars
1.2k
Token cost
~2.2k tokens
SKILL.md length
920 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
GPL-3.0

At a glance

Build, debug, extend, localize, or publish the automated tutorial videos — the workbench that drives the real Linux app under Xvfb (virtual mic, live Voxtral transcription, task-agent proposals)…

  • Works in 4 steps: The run log tail shows the failing step… → Look at the failure screenshot first → Reproduce faster without video: run the… → …
  • Asked to build/regenerate the tutorial video(s)
  • SKILL.md covers Build videos, Publish to R2, The pipeline in one paragraph and Debugging a failing run, plus 4 more sections
  • Calls make, python3 and flutter; needs GEMINI_API_KEY and MELIOUS_API_KEY

What it does

Tutorial Videos is an agent skill from matthiasn/lotti. Build, debug, extend, localize, or publish the automated tutorial videos — the workbench that drives the real Linux app under Xvfb (virtual mic, live Voxtral transcription, task-agent proposals), records the screen, narrates with Gemini TTS, fast-forwards waits, composes MP4s via OpenMontage, and uploads finished videos to Cloudflare R2 for docs-site embedding. Use when asked to "build/regenerate the tutorial video(s)", add a tutorial scenario, add a tutorial locale, fix a broken tutorial run, adapt the pipeline…

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice, Static sites and blogs and Proposals and quotes. It works with Cloudflare R2, Linux and Flutter. The repository describes itself as: A private logbook with a staff of personal AI assistants. Agents read what you record and propose what to do next — you approve the changes. End-to-end encrypted sync between… The licence is GPL-3.0.

When your agent uses it

  • Asked to build/regenerate the tutorial video(s)
  • Add a tutorial scenario
  • Add a tutorial locale
  • Fix a broken tutorial run

Example prompts

  • “build/regenerate the tutorial video(s)”
  • “/tutorial-videos”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • A credential in MELIOUS_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. The run log tail shows the failing step and a Diagnostics: line
  2. Look at the failure screenshot first
  3. Reproduce faster without video: run the same test via flutter drive
  4. Common causes, in order of historical frequency: a finder matching an

What it can do on your machine

Read from SKILL.md and the folder at commit e3dfbf7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make
    • python3
    • flutter
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY
    • MELIOUS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tutorial Videos loads about 2.2k tokens when it runs. Until then it costs about 140 tokens; SKILL.md has 920 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~140
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:25
    - `.env` at repo root with `GEMINI_API_KEY`, `MELIOUS_API_KEY`,
  • NoteMentions a .env fileSKILL.md:53
    the bucket's public r2.dev URL) and the `.env` variable list are in

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from matthiasn/lotti at commit e3dfbf7, republished under its GPL-3.0 licence (© matthiasn). 920 words, ~2,214 tokens.

Download SKILL.mdSave it as .claude/skills/tutorial-videos/SKILL.md (or your agent's skills folder).
name
tutorial-videos
description
Build, debug, extend, localize, or publish the automated tutorial videos — the workbench that drives the real Linux app under Xvfb (virtual mic, live Voxtral transcription, task-agent proposals), records the screen, narrates with Gemini TTS, fast-forwards waits, composes MP4s via OpenMontage, and uploads finished videos to Cloudflare R2 for docs-site embedding. Use when asked to "build/regenerate the tutorial video(s)", add a tutorial scenario, add a tutorial locale, fix a broken tutorial run, adapt the pipeline, or upload/publish a video.
argument-hint
<what to do, e.g. 'build de+en', 'add scenario X', 'debug the failing run'>

Tutorial Video Workbench

Architecture and rationale live in tools/tutorial_videos/README.md — read it first. This skill is the operational runbook.

Build videos

sh
make tutorial_video TUTORIAL_LOCALE=de                       # one locale (~5-8 min)
make tutorial_video TUTORIAL_LOCALE=de TUTORIAL_DEVICE=mobile # phone-shaped window, `_mobile`-suffixed output
make tutorial_videos_all                                     # all TUTORIAL_LOCALES

TUTORIAL_DEVICE (desktop|mobile, default desktop) picks the capture window size — see the README's "Desktop vs. mobile" section for the breakpoint rationale and the tutorialDeviceIsMobile() test hook.

Preconditions (fail fast if missing):

  • .env at repo root with GEMINI_API_KEY, MELIOUS_API_KEY, MELIOUS_BASE_URL. If absent, the keys can be extracted from the dev app's ai_config.sqlite provider rows (never print or commit them).
  • Sibling ../OpenMontage checkout at the commit pinned in tools/tutorial_videos/config/openmontage.pin, bootstrapped via make -C ../OpenMontage setup.
  • Xvfb, ffmpeg, pactl, paplay, x11-utils installed; Flutter via fvm.

Output lands in build/tutorial_videos/<scenario>_<locale>.mp4. Verify with frame extraction (ffmpeg -ss <t> -frames:v 1) and check: cursor visible and gliding, HUD clock top-center, transcript/proposals on screen, narration audible at step starts (volumedetect), duration ≈ warped timeline total.

Publish to R2

sh
make tutorial_video_publish TUTORIAL_SCENARIO=create_task_from_audio TUTORIAL_LOCALE=de
make tutorial_video_publish TUTORIAL_SCENARIO=create_task_from_audio TUTORIAL_LOCALE=de TUTORIAL_DEVICE=mobile

Uploads build/tutorial_videos/<scenario>_<locale>.mp4 (or, with TUTORIAL_DEVICE=mobile, <scenario>_<locale>_mobile.mp4) to Cloudflare R2 at the matching tutorial-videos/... key and prints its public r2.dev URL — this is the URL the docs-site TutorialVideo component builds from TUTORIAL_VIDEO_BASE_URL/tutorialVideoBaseUrl. TUTORIAL_DEVICE must match whatever --device the video was built with, or publish will look for a file that was never produced. Full one-time Cloudflare setup (API token, enabling the bucket's public r2.dev URL) and the .env variable list are in the README's "Publishing to Cloudflare R2" section — read it before the first publish.

boto3 (needed only by publish, not by build/tts/validate) is not installed into the system Python on this host (PEP 668). Either run publish through tools/tutorial_videos/.venv (.venv/bin/python3 -m tutorial_videos publish --scenario ... --locale ..., created via python3 -m venv .venv && .venv/bin/pip install boto3 pyyaml), or install boto3 into whatever Python make tutorial_video_publish resolves to.

The pipeline in one paragraph

python3 -m tutorial_videos build (run from tools/tutorial_videos/) does: Gemini-TTS pre-pass (cached by content hash → durations manifest) → Xvfb + virtual-mic null sink + ffmpeg x11grab capture → fvm flutter drive runs integration_test/tutorial/<scenario>_tutorial_test.dart against the real app (penguin demo world + Melious/Voxtral/Qwen agent config seeded, UI language via MANUAL_LANGUAGE settings row) → the test emits timeline.json (step boundaries, wait spans, zero epoch) → timewarp.py plans fast-forward segments for the waits (narration never overlaps) → compose.py warps the capture and drives the pinned OpenMontage headlessly (narration placement + mux) → ffprobe validation.

Debugging a failing run

  1. The run log tail shows the failing step and a Diagnostics: line (current route, detail stack, key widget presence).
  2. Look at the failure screenshot first: build/tutorial_videos/failure_<context>.png is captured at the moment of every timeout — it almost always explains the failure instantly.
  3. Reproduce faster without video: run the same test via flutter drive directly (see the doc comment in integration_test/tutorial/tutorial_harness.dart for the env contract); Xvfb + sink setup must match session.py.
  4. Common causes, in order of historical frequency: a finder matching an offstage duplicate (scope to the visible page + .hitTestable()), a scroll helper given a non-vertical scrollable, a modal/dialog blocking the UI (onboarding, discard-recording), the agent's automaticUpdatesEnabled flag not sticking (set-and-verify), cloud latency beyond a pumpUntil timeout.

Adding a locale

  1. Add the locale's text to tools/tutorial_videos/config/scenarios/<scenario>.yaml: title, dictionary, every step's narration, the dictation step's dictation_text. Write informal register (du/tu) per repo l10n rules; keep the penguin vocabulary.
  2. Add per-stream style lines for the locale in config/voices.yaml (style prompts in the target language — Voxtral/Gemini behave better).
  3. python3 -m tutorial_videos validate --scenario <s> --locale <l> must pass, then build. The app UI localizes itself via LOTTI_MANUAL_LOCALE (same mechanism as manual screenshots; for a wholly new app locale, run the add-flutter-docusaurus-locale skill first).
  4. Localized finder strings in the test (localized(en:, de:) / manualScreenshotText) need the new locale only if its UI labels are used as finders — prefer keys over text finders where possible.
Show full SKILL.md (322 more words)Show less

Adding a scenario

  1. New YAML in config/scenarios/<name>.yaml — schema documented in create_task_from_audio.yaml (unique step ids, exactly one dictation: true step, per-locale text, min_duration floors).
  2. New integration_test/tutorial/<name>_tutorial_test.dart built on TutorialAppHarness + TutorialDriver. Copy the structure of create_task_from_audio_tutorial_test.dart:
    • one driver.step(id, action) per YAML step, ids matching exactly;
    • interactions via driver.tapLikeUser (cursor glide), waits via driver.pumpUntil* (records fast-forwardable wait spans), scrolling via driver.scrollIntoView with a vertical scrollable finder;
    • anything that must be SEEN gets asserted via visible widgets, not just the DB;
    • wire the diagnostics and onTimeout hooks for failure screenshots.
  3. Iterate with builds; expect the first runs to fail — read the failure screenshots, they carry the answer.

Non-negotiable gotchas

The full list with rationale is in the README's "Hard-won constraints" — the ones that bite hardest: flutter drive not flutter test (no window otherwise); GDK_BACKEND=x11 + empty WAYLAND_DISPLAY (or the app opens on the real desktop); Voxtral prompts in the audio's language (else it translates); automaticUpdatesEnabled set-and-verify (initial wake clobbers early writes); 16ms pump cadence during holds (else animations render choppy); timeline starts after startup settles; assert transcript keywords, not exact strings.

Adapting / extending

  • App Store preview: config/scenarios/app_store_preview.yaml is the narration of make store_preview_ios, not a tutorial — never make tutorial_video it. Its step ids must stay the _beatIds of integration_test/store_preview_test.dart (tests/test_app_preview.py guards it), and all four lines together must leave the beats under 30 s (python3 -m tutorial_videos.app_preview pacing fails first if not). See the README's "The App Store preview reuses the narration".
  • Compositor changes: all OpenMontage API usage is isolated in tutorial_videos/om_compose_driver.py; bump the pin in config/openmontage.pin deliberately and re-verify determinism (compose twice, compare ffmpeg -f framemd5 of the streams — never file hashes).
  • TTS vendor: implement the TtsEngine protocol (tts/base.py) next to tts/gemini.py; voices/styles per locale in config/voices.yaml.
  • Speed/pacing knobs: timewarp.py (MAX_SPEED, lead-in, narration gap); per-step floors in the scenario YAML.
  • Character overlay (future): render transparent PNG frames via the character film-strip harness and composite as an extra layer keyed to timeline.json step ids.

© matthiasn, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/tutorial-videos of matthiasn/lotti.

Open the folder on GitHubat commit e3dfbf7

Compare with similar skills

Tutorial Videos next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tutorial Videos compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tutorial Videos this skillmatthiasn/lotti1.2k—~2.2kAutomated safety check: NotesGPL-3.0
Local AI App Integrationamd/skills398—~6kAutomated safety check: PassMIT
AI SDK Developmenttrypostit/trypost6782 repos~3.5kAutomated safety check: PassMIT
Audio To SubtitlesbozhouDev/video-skills-toolkit150—~1.9kAutomated safety check: NotesMIT
Xsaimoeru-ai/airi50k1 repos~1.3kAutomated safety check: PassMIT
Local AI Useamd/skills398—~5kAutomated safety check: NotesMIT

Similar skills

  • Integrates local AI capabilities into applications using Embeddable Lemonade.

    398 GitHub stars~6k tokensUpdated today
    Media & CreativeAuto-check passed
  • AI SDK Development

    trypostit/trypost

    TRIGGER when working with ai-sdk which is Laravel official first-party AI SDK.

    678 GitHub starsUsed in 2 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Audio To Subtitles

    bozhouDev/video-skills-toolkit

    Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API.

    150 GitHub stars~1.9k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Xsai

    moeru-ai/airi

    A skill your agent uses when the user is building with xsai or any @xsai/ package, or is evaluating xsAI for a small OpenAI-compatible workflow with text generation, streaming, tool calling…

    50k GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    398 GitHub stars~5k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Azure AI

    microsoft/GitHub-Copilot-for-Azure

    Official

    A skill your agent uses for Azure AI: Search, Speech, OpenAI, Document Intelligence.

    255 GitHub starsUsed in 2 repos~852 tokens
    Media & CreativeAuto-check passed

More from matthiasn/lotti

All 10 skills in this repo
  • Add, complete, or audit a locale across a Flutter application's ARB catalogs and a localized Docusaurus manual, including generated localization code, locale selectors, native platform declarations…

    1.2k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Maintain a Docusaurus product manual whose prose, navigation, localized page trees, screenshots, coverage metadata, and release build must stay aligned with the application.

    1.2k GitHub stars~930 tokensUpdated today
    Auto-check passed
  • App Screenshots

    matthiasn/lotti

    Capture in-app screenshots of a widget or flow at mobile + desktop sizes — offline, deterministic, real fonts and icons — using the reusable screenshot harness.

    1.2k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Beads

    matthiasn/lotti

    A skill your agent uses for authorized Lotti maintainer work that needs private durable task tracking, issue dependencies, blocker management, multi-session handoff, or shared agent memory.

    1.2k GitHub stars~589 tokensUpdated today
    Auto-check passed
  • Design Review Panel

    matthiasn/lotti

    Run a multi-agent design review on a UI surface — capture reproducible baseline screenshots, then rate them with a panel of design experts (one agent per craft dimension) and, optionally, a panel of…

    1.2k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Release

    matthiasn/lotti

    Cut a Lotti release — assemble the changelog.d/ fragments into CHANGELOG.md and the Flathub metainfo, bump the version, open the release PR, then tag.

    1.2k GitHub stars~2.6k tokensUpdated today
    Auto-check passed

Questions about Tutorial Videos

What does Tutorial Videos do?

Build, debug, extend, localize, or publish the automated tutorial videos — the workbench that drives the real Linux app under Xvfb (virtual mic, live Voxtral transcription, task-agent proposals)…. Tutorial Videos is an agent skill from matthiasn/lotti. Build, debug, extend, localize, or publish the automated tutorial videos — the workbench that drives the real Linux app under Xvfb (virtual mic, live Voxtral transcription, task-agent proposals), records the screen, narrates with Gemini TTS, fast-forwards waits, composes MP4s via OpenMontage, and uploads finished videos to Cloudflare R2 for docs-site embedding.

When should I use Tutorial Videos?

Tutorial Videos fits situations like: asked to build/regenerate the tutorial video(s); add a tutorial scenario; add a tutorial locale; fix a broken tutorial run.

How do I install Tutorial Videos in Claude Code?

Run `npx skills add matthiasn/lotti --skill tutorial-videos -a claude-code`. Or copy the skill folder (.claude/skills/tutorial-videos in matthiasn/lotti) into .claude/skills/tutorial-videos in your project. Claude Code loads it when a task matches its description.

How do I install Tutorial Videos in Codex?

Run `npx skills add matthiasn/lotti --skill tutorial-videos -a codex`. Or copy the skill folder (.claude/skills/tutorial-videos in matthiasn/lotti) into .agents/skills/tutorial-videos in your project. Codex loads it when a task matches its description.

Can I use Tutorial Videos in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add matthiasn/lotti --skill tutorial-videos -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tutorial-videos, .gemini/skills/tutorial-videos, .github/skills/tutorial-videos and .opencode/skills/tutorial-videos in your project.

What does Tutorial Videos need to run?

Going by SKILL.md and its folder, Tutorial Videos needs the command-line tools its instructions call (make, python3, flutter and ffmpeg) and credentials named GEMINI_API_KEY and MELIOUS_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY; A credential in MELIOUS_API_KEY.

Does Tutorial Videos access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tutorial Videos safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tutorial Videos use?

Tutorial Videos is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tutorial Videos use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tutorial Videos?

Skills that share tags, products or a category with Tutorial Videos: Local AI App Integration (amd/skills, 398 stars), AI SDK Development (trypostit/trypost, 678 stars), Audio To Subtitles (bozhouDev/video-skills-toolkit, 150 stars) and Xsai (moeru-ai/airi, 50k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tutorial Videos?

matthiasn (a GitHub user) maintains it in matthiasn/lotti, which has 1,198 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 7, 2026.

Source: matthiasn/lotti on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.