Agent skill

Book Video Factory

by jaxxchen003 in jaxxchen003/book-video-factory

Create or operate a portable, auditable Chinese book-review short-video workflow from a clean local workspace.

MITAuto-check passedWriting & Content

Install Book Video Factory

skills CLI
$ npx skills add jaxxchen003/book-video-factory --skill book-video-factory -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaxxchen003/book-video-factory book-video-factory --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaxxchen003/book-video-factory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/book-video-factory .claude/skills/book-video-factory && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
book-video-factory
GitHub stars
105
Token cost
~2.4k tokens
SKILL.md length
1,062 words
Files
93 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Create or operate a portable, auditable Chinese book-review short-video workflow from a clean local workspace.

  • Works in 4 steps: Resolve SKILL_ROOT as the directory… → In an empty or new workspace, run → Create a project only after the user… → …
  • Starting a book-video factory
  • SKILL.md covers First use, Working model, Style profiles and Production sequence, plus 4 more sections
  • Calls python3; needs GEMINI_API_KEY

What it does

Book Video Factory is an agent skill from jaxxchen003/book-video-factory. Create or operate a portable, auditable Chinese book-review short-video workflow from a clean local workspace. Use when starting a book-video factory, choosing between the original 3:4 bilingual editorial style and the new 9:16 VOX-style book-video workflow, importing dbs-content-system source documents and QST/CON/OPI/CAS/SOL content units, linking script claims to evidence and scenes, using approved Gemini API or Google Flow assets, collecting rights-aware media, recording run cost, or preparing a local master…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 99 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/first-run.md` and `references/paper-collage-explainer.md`).

It sits in Writing & Content, covering Translation and LLM API integration. It works with Google Gemini. The repository describes itself as: Portable Codex skill for auditable, rights-aware Chinese book-review short-video workflows. The licence is MIT.

When your agent uses it

  • Starting a book-video factory
  • Choosing between the original 3:4 bilingual editorial style and the new 9:16 VOX-style book-video workflow
  • Importing dbs-content-system source documents and QST/CON/OPI/CAS/SOL content units
  • Linking script claims to evidence and scenes

Example prompts

  • “/book-video-factory”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Resolve SKILL_ROOT as the directory containing this SKILL.md; do not assume a particular user home directory, operating system, brand…
  2. In an empty or new workspace, run
  3. Create a project only after the user approves the book/topic
  4. Read references/first-run.md before generating the first production package. Read references/quality-gates.md before declaring a release…

What it can do on your machine

Read from SKILL.md and the folder at commit af296f9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Book Video Factory loads about 2.4k tokens when it runs, and up to ~7.9k if it reads all its reference files. Until then it costs about 143 tokens; SKILL.md has 1,062 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~143
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jaxxchen003/book-video-factory at commit af296f9, republished under its MIT licence (© jaxxchen003). 1,062 words, ~2,373 tokens.

Download SKILL.mdSave it as .claude/skills/book-video-factory/SKILL.md (or your agent's skills folder). This skill also uses 92 other files; get the full folder from GitHub.
name
book-video-factory
description
Create or operate a portable, auditable Chinese book-review short-video workflow from a clean local workspace. Use when starting a book-video factory, choosing between the original 3:4 bilingual editorial style and the new 9:16 VOX-style book-video workflow, importing dbs-content-system source documents and QST/CON/OPI/CAS/SOL content units, linking script claims to evidence and scenes, using approved Gemini API or Google Flow assets, collecting rights-aware media, recording run cost, or preparing a local master for optional ChatCut fine editing.

图书视频工厂

First use

  1. Resolve SKILL_ROOT as the directory containing this SKILL.md; do not assume a particular user home directory, operating system, brand, voice, font, or credential.

  2. In an empty or new workspace, run:

    bash
    python3 <SKILL_ROOT>/scripts/bootstrap_workspace.py --workspace .
    python3 <SKILL_ROOT>/scripts/doctor.py --profile planning

    Bootstrap copies the bundled deterministic runtime into book_video_factory/; it does not download hidden media or depend on the maintainer's machine.

  3. Create a project only after the user approves the book/topic:

    bash
    python3 <SKILL_ROOT>/scripts/bootstrap_workspace.py \
      --workspace . --slug <slug> --book-title '<title>' --author '<author>'

    If the topic comes from a structured content asset system, add --mode content-system-backed. This mode requires a validated content package and traceability map before assets can pass their gate.

  4. Read references/first-run.md before generating the first production package. Read references/quality-gates.md before declaring a release ready.

Working model

  • Keep reusable workflow/configuration in book_video_factory/ and project-specific evidence/media in book_video_warehouse/.
  • Treat all book covers, quotations, BGM, voice references, reference videos, generated assets, and credentials as user-owned/project-local inputs. Never ship, download, clone, or reuse a hidden default asset.
  • Create a new release directory for every revision. Preserve the local master as the source of truth; ChatCut is an optional editable polish layer.
  • Keep human gates for topic approval, script approval, source/rights approval, native-language review, and publish approval.

Style profiles

Select one style profile before making assets; never silently mix their frame, typography, or review contracts.

  • 双语编辑模板图书视频(原风格) — machine ID book-editorial-bilingual-v2. Deterministic 3:4 bilingual template with a real cover, 12 approved stills, local narration, rendered captions, and an optional 9:16 derivative.
  • VOX风格图书视频(新增风格) — machine ID paper-collage-explainer-v1. A 9:16, voice-led editorial paper-collage explainer made from 4–8-second visual beats. It is a descriptive public label, not an affiliation claim or permission to imitate a named publisher/program. Read references/paper-collage-explainer.md before selecting it.

The original style is the backward-compatible default. For VOX style, create the project with --style-profile paper-collage-explainer-v1 and explicitly choose one generation lane:

  • gemini-api — programmatic lane. Requires a user-authorized GEMINI_API_KEY, the current Google Gen AI SDK, provider cost/quota approval, and an immutable operation/output record. Use gemini-omni-flash-preview through the Interactions API by default; use veo-3.1-generate-preview through generate_videos when first/last-frame control or video extension is required.
  • google-flow — manual creative lane. Requires current account/region eligibility, an eligible Google AI plan, available credits, and a user-operated browser session. Do not assume Flow exposes a programmable API; import only user-authorized exports and record prompts, visible credits, hashes, and manual-run provenance.

The VOX-style path is currently an orchestration-and-import workflow. Do not claim this repository bundles a one-click Google video generator. The current dbs-content-system scene-traceability contract is V4-specific, so content-system-backed is supported by the original style only; VOX style is fail-closed to single-book until a manifest-based traceability contract is implemented and tested.

Production sequence

  1. Topic and evidence — collect public, attributable book metadata. If a WeRead credential or another data source is unavailable, record the limitation and use user-supplied/publicly attributable evidence; do not bypass logins or platform restrictions.
  2. Script, voice, and timing — create and approve a concise Chinese script plus any English production draft. Generate/record narration only with authorization; derive timestamps from actual audio via permitted ASR or an editor transcript. Never invent timing. VOX style locks this timing before full clip generation.
  3. Style-specific assets — the original style obtains a real cover plus 12 approved text-free stills. VOX style approves one metaphor, then a first/last frame or contact sheet, then each generated visual beat; prompts use non-branded editorial descriptors.
  4. Picture lock — map approved assets to the locked narration. Normalize VOX clips to silent 720×1280 H.264 at the local 30 fps timeline without speed changes; inspect cadence and keep optical flow off by default.
  5. Audio and render — add project-owned narration, local captions, permitted BGM and SFX. BGM rights approval and creative bgm_review are separate decisions. Never copy a reference video's audio.
  6. QC and delivery — run technical and human gates, bind approvals to hashes, write manifests and cost events, then optionally import the passed local master into ChatCut for fine editing.
Show full SKILL.md (428 more words)Show less

Workflow contracts

  • project.json.workflow.style_profile_id selects the visual workflow; its mapped release_profile_id selects the renderer/output contract. Existing projects without a style ID fall back to the original style only for compatibility.
  • Use config/release_profiles/book-v4-bilingual-3x4.json for the original style and config/release_profiles/book-vox-vertical-9x16-v1.json for VOX style; never treat either style's dimensions or assets as universal constants.
  • Use scripts/workflow.py evaluate --project <project> --release-id <release-id> to resolve the project profile and derive a release-scoped state. project.json.status is not an approval mechanism, and approvals from different releases are never combined.
  • Record human decisions with scripts/workflow.py approve; approvals bind to the reviewed file hash and become stale after edits.
  • Use scripts/workflow.py manifest-stage for immutable stage manifests with input/output hashes.
  • Long titles are pixel-measured, semantically wrapped to at most two lines, and fail closed if they cannot fit the configured safe area.
  • Keep dbs-content-system upstream: it owns source audits, QST / CON / OPI / CAS / SOL, theme maps, relationships, deduplication, canonical versions, and assembly. Do not reproduce those algorithms in this Skill.
  • For content-system-backed, use scripts/content_bridge.py export-dbs, validate-package, import-package, attach-traceability, and status. export-dbs is serialization only; it does not perform upstream semantic work. Imports and active-version changes are append-only and hash-bound.
  • A valid bridge package contains source_document / content_unit / claim / assembly_brief; the traceability map connects every script line to reviewed Claim evidence or an explicit editorial exemption, and to the renderer's actual scene contract. A human traceability approval bound to that map is required before assets_ready.

Required release gates

  • Do not use a generated imitation as a book cover. Record the actual cover source and rights/usage status.
  • Do not reuse copyrighted music, sound effects, reference-video audio, a person's voice, public-figure likeness, or account branding without explicit rights.
  • Require 12 non-duplicate numbered stills (S01–S12) only for the original V4-style delivery. VOX style uses its approved clip manifest and does not impose a universal scene count.
  • For VOX style, keep source_audit, gate_1_metaphor, gate_2_still, gate_3_clip_qa, bgm_review, audio_qc, local_master_review, and publish as distinct hash-bound gates. Apply cover_rights and english_native only when those elements are delivered.
  • Accept invisible provider provenance such as SynthID; do not remove or tamper with it. Reject visible third-party watermarks.
  • Never invent token counts. Record only usage values exposed by the relevant provider.

Cost ledger

Use the bundled append-only ledger for known usage and operational facts:

bash
python3 <SKILL_ROOT>/scripts/run_cost.py record \
  --warehouse book_video_warehouse --project <slug> --stage assets.generate \
  --images 12 --note 'Approved scene images'
python3 <SKILL_ROOT>/scripts/run_cost.py report --warehouse book_video_warehouse

— in a report means the usage was not available; it is not zero cost.

ChatCut handoff

Use ChatCut only after local QC passes. Import the local master and subtitle file, make scoped editorial changes, and export v4-chatcut-<revision> without overwriting the local master. Record the project ID, edits, reviewer decision, and export path in the project delivery manifest.

© jaxxchen003, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 92 other files (scripts, references) in skills/book-video-factory of jaxxchen003/book-video-factory.

  • SKILL.md
  • agents/openai.yaml
  • references/first-run.md
  • references/paper-collage-explainer.md
  • references/quality-gates.md
  • runtime/book_video_factory/DEPENDENCIES.md
  • runtime/book_video_factory/README.md
  • runtime/book_video_factory/config/brand_voice_profile.json
  • runtime/book_video_factory/config/music_sources.json
  • runtime/book_video_factory/config/release_profiles/book-v4-bilingual-3x4.json
  • runtime/book_video_factory/config/release_profiles/book-vox-vertical-9x16-v1.json
  • runtime/book_video_factory/config/style_profiles/book-editorial-bilingual-v2.json
  • runtime/book_video_factory/config/style_profiles/paper-collage-explainer-v1.json
  • runtime/book_video_factory/config/topic_scoring.json
  • … and 79 more

Open the folder on GitHubat commit af296f9

Compare with similar skills

Book Video Factory next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Book Video Factory compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Book Video Factory this skilljaxxchen003/book-video-factory105—~2.4kAutomated safety check: PassMIT
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0
Roo Translationzgsm-ai/costrict4.5k—~1.8kAutomated safety check: PassApache-2.0
Thumbnail Creatormohitagw15856/pm-claude-skills1.4k—~6.3kAutomated safety check: PassMIT
ModLens Image Vision Bridgeliustack/modlens4.2k—~1.3kAutomated safety check: NotesMIT
AI Daily DigestvigorX777/ai-daily-digest1.6k—~1.3kAutomated safety check: PassNone

Similar skills

  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated 4 days ago
    Backend & APIsAuto-check passed
  • Roo Translation

    zgsm-ai/costrict

    Provides comprehensive guidelines for translating and localizing CoStrict extension strings.

    4.5k GitHub stars~1.8k tokensUpdated 11 days ago
    Writing & ContentAuto-check passed
  • Thumbnail Creator

    mohitagw15856/pm-claude-skills

    Generate article or newsletter thumbnail candidates using the Gemini API from inside Claude Code.

    1.4k GitHub stars~6.3k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.2k GitHub stars~1.3k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check: notes
  • AI Daily Digest

    vigorX777/ai-daily-digest

    Fetches RSS feeds from 90 top Hacker News blogs (curated by Karpathy), uses AI to score and filter articles, and generates a daily digest in Markdown with Chinese-translated titles, category…

    1.6k GitHub stars~1.3k tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Gemini API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice…

    4.3k GitHub stars~5.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

Works with

Questions about Book Video Factory

What does Book Video Factory do?

Create or operate a portable, auditable Chinese book-review short-video workflow from a clean local workspace. Book Video Factory is an agent skill from jaxxchen003/book-video-factory. Create or operate a portable, auditable Chinese book-review short-video workflow from a clean local workspace.

When should I use Book Video Factory?

Book Video Factory fits situations like: starting a book-video factory; choosing between the original 3:4 bilingual editorial style and the new 9:16 VOX-style book-video workflow; importing dbs-content-system source documents and QST/CON/OPI/CAS/SOL content units; linking script claims to evidence and scenes.

How do I install Book Video Factory in Claude Code?

Run `npx skills add jaxxchen003/book-video-factory --skill book-video-factory -a claude-code`. Or copy the skill folder (skills/book-video-factory in jaxxchen003/book-video-factory) into .claude/skills/book-video-factory in your project. Claude Code loads it when a task matches its description.

How do I install Book Video Factory in Codex?

Run `npx skills add jaxxchen003/book-video-factory --skill book-video-factory -a codex`. Or copy the skill folder (skills/book-video-factory in jaxxchen003/book-video-factory) into .agents/skills/book-video-factory in your project. Codex loads it when a task matches its description.

Can I use Book Video Factory in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaxxchen003/book-video-factory --skill book-video-factory -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/book-video-factory, .gemini/skills/book-video-factory, .github/skills/book-video-factory and .opencode/skills/book-video-factory in your project.

What does Book Video Factory need to run?

Going by SKILL.md and its folder, Book Video Factory needs the command-line tools its instructions call (python3) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY.

Does Book Video Factory access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Book Video Factory safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Book Video Factory use?

Book Video Factory is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Book Video Factory use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.

What are the alternatives to Book Video Factory?

Skills that share tags, products or a category with Book Video Factory: Gemini Live API Dev (google-gemini/gemini-skills, 4.3k stars), Roo Translation (zgsm-ai/costrict, 4.5k stars), Thumbnail Creator (mohitagw15856/pm-claude-skills, 1.4k stars) and ModLens Image Vision Bridge (liustack/modlens, 4.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Book Video Factory?

jaxxchen003 (a GitHub user) maintains it in jaxxchen003/book-video-factory, which has 105 GitHub stars. The repository was last updated on August 17, 2026.

Source: jaxxchen003/book-video-factory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.