Stage Edit
Orkas-AI/Orkas-VideoStudio
Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…
A skill your agent uses when the user wants a finished video out of gflow rather than a single clip — a scripted scene, a talking-head or dialogue piece, an explainer, a product montage, a story…
$ npx skills add ffroliva/gflow-cli --skill video-production -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ffroliva/gflow-cli video-production --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ffroliva/gflow-cli.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-production .claude/skills/video-production && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "video-production" agent skill from https://github.com/ffroliva/gflow-cli/tree/develop/skills/video-production into .claude/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ffroliva/gflow-cli/tree/develop/skills/video-productionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ffroliva/gflow-cli --skill video-production -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ffroliva/gflow-cli video-production --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ffroliva/gflow-cli.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/video-production .agents/skills/video-production && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "video-production" agent skill from https://github.com/ffroliva/gflow-cli/tree/develop/skills/video-production into .agents/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ffroliva/gflow-cli --skill video-production -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ffroliva/gflow-cli video-production --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ffroliva/gflow-cli.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/video-production .cursor/skills/video-production && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "video-production" agent skill from https://github.com/ffroliva/gflow-cli/tree/develop/skills/video-production into .cursor/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ffroliva/gflow-cli.git --path skills/video-production--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ffroliva/gflow-cli --skill video-production -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ffroliva/gflow-cli video-production --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ffroliva/gflow-cli.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/video-production .gemini/skills/video-production && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "video-production" agent skill from https://github.com/ffroliva/gflow-cli/tree/develop/skills/video-production into .gemini/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ffroliva/gflow-cli video-productionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ffroliva/gflow-cli --skill video-production -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ffroliva/gflow-cli.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/video-production .github/skills/video-production && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "video-production" agent skill from https://github.com/ffroliva/gflow-cli/tree/develop/skills/video-production into .github/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ffroliva/gflow-cli --skill video-production -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ffroliva/gflow-cli video-production --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ffroliva/gflow-cli.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/video-production .opencode/skills/video-production && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "video-production" agent skill from https://github.com/ffroliva/gflow-cli/tree/develop/skills/video-production into .opencode/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
video-productionA skill your agent uses when the user wants a finished video out of gflow rather than a single clip — a scripted scene, a talking-head or dialogue piece, an explainer, a product montage, a story…
Video Production is an agent skill from ffroliva/gflow-cli. Use when the user wants a finished video out of gflow rather than a single clip — a scripted scene, a talking-head or dialogue piece, an explainer, a product montage, a story sequence, an audition or rehearsal reference, a short film — or asks for consistent actors, a consistent location, a specific prop that must not change, several camera angles, captions or subtitles, or joining clips into one file. Also use when clips came back wrong: a film-strip border, a room that changes between shots, a prop that morphs…
Its SKILL.md is about 7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `clip_qa.py`, `composition.md` and `consistency.md`).
It sits in Media & Creative, covering Video production, Transcription and AI video generation. It works with FFmpeg. The repository describes itself as: Drive Google Flow from the command line: Veo video and Imagen images, scripted, batched and pipeline-ready. Ships an MCP server so coding agents can drive it too, giving you and… The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit eb1b1ec. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonffmpegFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
labs.googlegstatic.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Video Production loads about 7k tokens when it runs. Until then it costs about 163 tokens; SKILL.md has 3,699 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ffroliva/gflow-cli at commit eb1b1ec, republished under its MIT licence (© ffroliva). 3,699 words, ~7,036 tokens.
.claude/skills/video-production/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Core principle: control is everything. Every guard-rail you put in front of the engine is drift you do not pay for later. Lock the shape, the cast, the location and the words before spending, then gate every clip on evidence rather than impression.
This skill covers composing gflow into a finished video. It does not restate the command surface — that is the gflow-cli skill, which owns per-command syntax, flags and single-shot recipes. Load that one for "how do I call t2v", this one for "how do I turn a script into a film that holds together".
Every non-obvious statement is tagged. This is one calibrated approach, not the only one. Other people drive this engine with techniques not tracked here; absence from this document means untested, never forbidden.
| Tag | Meaning | You may |
|---|---|---|
| [CONSTRAINT] | the engine refuses, fails, or silently drops the request | not deviate |
| [CALIBRATED] | measured here, sample size and conditions stated | deviate with evidence |
| [CONVENTION] | one shape that works; alternatives exist | deviate freely |
| [UNEXPLORED] | known to exist, not tested here | try it, then add a task |
A single clip with no continuity requirement, one image, or a pure command-syntax question — use the gflow-cli skill. Editing or grading existing footage — this is generation, not post. Any engine that is not Flow.
Everything gflow — Python 3.11+, uv, an installed gflow, Playwright Chromium, a signed-in profile, Flow access — belongs to the gflow-cli skill's Prerequisites. Run its checks; do not restate them.
This skill adds three:
ffmpeg and ffprobe on PATH, 5.0 or newer — ffmpeg -version. The assembly step uses -fps_mode, which does not exist before 5.0.libass and libfreetype — the same banner lists enabled libraries. Burned subtitles (subtitles) and title cards (drawtext) are absent from minimal or "essentials" builds. Usual Windows trip.faster-whisper, only if you want the transcript gate — python -c "import faster_whisper". Nothing in gflow installs it. First run pulls base.en, ~75 MB, once. Without it you lose the word-hit check and keep every other gate.clip_qa.py, beside this file, needs only ffmpeg, ffprobe and the standard library — so the fluidity, lip-sync and A/V-drift gates still run where the transcript gate cannot.
The plan changes completely with the answers. Ask, or read them from the brief; do not assume.
| Question | Why it changes the plan |
|---|---|
| What is the deliverable, and who watches it? | a rehearsal reference tolerates flaws a client cut does not |
| Does anyone speak on camera? | decides dialogue budgeting, lip-sync gating, and i2v-vs-r2v |
| Is audio generated, added later, or discarded? | Veo always generates audio [CONSTRAINT] — see below |
| Captions: none, burned-in, or a sidecar file? | burned-in needs libass and a timing source |
| One location or several? How many camera angles each? | drives the plate-chaining work in consistency.md |
| Recurring people? How many in frame at once? | one face-bearing reference per generation [CONSTRAINT] |
| A prop that must not change? | needs its own sheet with a scale anchor |
| Aspect, total length, credit ceiling | 16:9 or 9:16; clips are 4/6/8, or 10 on omni-flash |
| One-off, or a repeatable pipeline? | decides the production shape in composition.md |
Veo always generates audio [CONSTRAINT]. There is no silent mode and omitting sound from the prompt does not produce silence. A "silent" montage returns invented room tone under every shot. If the audio is unwanted, strip it at assembly (-an) rather than hoping for a quiet clip.
Load https://labs.google/fx/tools/flow/project/<id> in the profile's Chrome and read the final URL.
| Lands on | Lane | What runs there |
|---|---|---|
labs.google/fx/… | A, full | everything |
flow.google.com/project/… | B, partial | video t2v --project, video i2v --initial-frame <local file> --project, video r2v --ref <local file> --project (ported in v0.70.0, #683), character create (v0.70.0), and image t2i / image i2i with local --ref files (#639). Everything else — image batch, extend, scene create, movie run, r2v by @Name or --reference-entity, i2v by media UUID / @Name, and Imagen 4 — exits 36 [CONSTRAINT]. A local --end-frame IS ported (start+end interpolation, #639) |
Lane B grows as forms are ported, so confirm the row rather than trusting it: the
maintained list is CONFIGURATION § GFLOW_CLI_FLOW_HOST
and the #639 entry in KNOWN_ISSUES. Exit 36 on a form the table
says is ported is a regression worth filing, not the environment.
An unported command on lane B exits 36, non-retryable, with a message naming the
migration. A RecaptchaError instead means one of two other things: on gflow ≤ 0.68.0 the
migration guard ran after the reCAPTCHA mint on the image path, so image commands died
as exit 1 there (gflow-cli#673, fixed); on any build, a RecaptchaError right after
gflow auth login is the cookie harvest keyed on the old host (gflow-cli#644). Either
way the first move is the same: read the final host URL before anything else. Do not
plan entities or plates before this check.
There is a gflow project create [CONSTRAINT]. gflow project create --name <piece> --json > out.json, redirected to a file, never piped through head, which truncates the process before the JSON prints. Do not mint a project by burning a placeholder generation, and do not scrape the id from a browser URL.
Pick one; each has a worked command chain in composition.md.
| Shape | When | Driver |
|---|---|---|
| One-off sequence | a scene, an audition, a montage | shell, clip per beat |
| Manifest-driven | a script that will be re-run as it is edited | movie run with a stable scene id per slot |
| Continuous shot | one camera move longer than 8 s | video extend |
| Deliberate cut with continuity | a new angle that must match the last frame | video chain |
| Assembly only | clips already exist | scene create (free) |
A person in a shot is anchored by exactly one of these. Take the highest rung the host and path allow, and write down which rung you used and what blocked the one above. This is not a preference order, it is the production record: a film whose log does not say how each shot was anchored cannot be debugged when a face drifts.
| # | Anchor | What it carries | Use when |
|---|---|---|---|
| 1 | Character entity — @Name or --reference-entity <id> | face and wardrobe and voice, server-side, durable | always, unless a rung-1 blocker is recorded |
| 2 | That entity's own generated plate — --ref <its portrait or body crop> | face and wardrobe, as a flat image | rung 1 refused by the host/path |
| 3 | A plate cut from an approved take — --ref <frame> | face, plus that take's grade, light and artefacts | no entity exists |
| 4 | Prose canon only | a type, never an identity | nothing else is available |
Rung 4 does not hold a person. Each generation invents someone new who merely matches the description. Two shots on rung 4 are two different actors. Never plan a multi-shot piece on it and never let it be the silent default.
Rung 3 carries contamination. v1's opening take came back with 72 px letterbox bars and its plate carried those bars into every shot that referenced it. A rung-2 plate is generated clean by the character editor; a rung-3 plate is only as clean as the take it was cut from.
Descending a rung is a decision that gets written down, in production.json next to the
shot, naming the blocker. "I used an image" is not a record; "rung 2, because
--reference-entity exits 36 on this host (#639)" is.
Written from a run that skipped its own ladder. On 2026-09-07 a five-shot film created two real character entities and then anchored every shot on rung 2, because gflow refuses entity references on the migrated host. The refusal was never questioned — and a $0 probe the same day showed the host takes an entity mention perfectly well (
data-reference-type="entity"with the realentity_id); it is gflow that has not ported the gesture. The run went a rung lower than it had to and recorded it only in passing.
Flow's @ picker offers character entities and media assets in one list, and a query
matching both can resolve to either — measured 2026-09-07: the same @Kael returned
reference_type="entity" on one gesture and reference_type="media" (a JPEG that happened
to be named after him) on another. Flow does not rank them.
So you rank them. When a name matches a character and a file, the character wins, and if
you meant the file, reference it by path (--ref <path>) rather than by name. A shot that
silently binds a still image where you asked for a character produces footage that looks
right and drifts on the next cut.
A character entity is the only thing that carries a voice. A plate does not: it carries face and wardrobe as pixels, and the engine invents a new voice for every clip. Measured 2026-09-07 on one character across three plate-bound takes: 88 Hz, 103 Hz, 118 Hz — three different actors — against an engine noise floor of 4.3 Hz on an identical prompt repeated three times. No re-shoot fixes that, because nothing in a plate was ever carrying the voice.
So if anyone speaks more than once, you need rung 1. Here is the whole path.
1 — pick the voice before you create anyone. gflow character voices lists 29 presets.
Each has a public sample you can actually listen to, so audition rather than trust the
descriptor — five of the 29 descriptors disagree with the measured pitch of their own sample:
gflow character voices --json
# every voice has: https://gstatic.com/aitestkitchen/voices/samples/<Name>.wav2 — create the character with the voice bound. Face, body and voice in one call:
gflow character create --project "$PROJECT" \
--name "<UniqueName>" \
--face-prompt "a man, <face description with a GENDER WORD>" \
--body-prompt "<outfit; no print, no logo>" \
--voice "<VoiceName>" \
--personality "<how they behave>" \
--json > cast_<name>.jsonThree traps, each measured:
@ picker searches
characters and media together and does not rank them (4b), so an uploaded kael_ref.jpg
wins the query @Kael and your character becomes unreachable by name. Never name a
plate after a character; plate_a.png, not <name>_ref.jpg.character create — a two-hander whose unbound actor is
described without one renders the wrong person.entity_id yourself. Read it from --json at creation. Do not plan on
recovering it from the catalog afterwards.3 — attach the character to every shot they appear in.
gflow video t2v "<prompt>" --project "$PROJECT" \
--reference-entity "<entity_id>" \
--reference-entity-name "<UniqueName>" \
--aspect 16:9 --duration 8One face-bearing reference per generation [CONSTRAINT] — so in a two-hander, bind the person who speaks and carry the other in prose. Getting that backwards is what produced the 88 Hz stranger above: the beat bound the silent actor, so the speaker fell to rung 4.
4 — expect the submit to outlive your patience, and do not read a timeout as a failure.
Flow allows five concurrent generations and throttles per-minute throughput after heavy
daily use, so a queued job routinely outlives gflow's submit-reply budget. An
entity-bound run can exit 9 TransportTimeoutError while the video is rendering
normally — the job is in Flow's queue and will finish. Check the project before you
re-submit, or you will double-spend on a generation you already have. (gflow-cli #723,
#741.)
5 — verify the voice actually landed, relatively. Never assert an absolute band: a character legitimately speaks differently in an action beat than in a quiet one, and the pitch follows the performance. The two sound comparisons are:
And stage every dialogue beat in still air. An energy gate is not a voicing gate: on this production's own wordless clips the detector reported a confident 145 Hz and 280 Hz with no speech present at all, and periodicity did not separate them either. A beat shot in wind cannot be checked.
A Flow character entity bundles face and wardrobe — the body reference fixes the outfit. There is no wardrobe axis inside one entity.
So a character who changes clothes is two entities sharing a face prompt, with different body prompts, named for the costume state:
Kael_ridge face_prompt=<the canonical face> body_prompt=<dust-brown canvas jacket, sand scarf>
Kael_coat face_prompt=<the same canonical face verbatim> body_prompt=<heavy oiled coat, hood down>The face prompt must be byte-identical across costume states; only the body prompt moves. Then every scene names the costume-state entity, not the character, and continuity becomes a lookup instead of a hope.
This is not what movie.toml does today [CONSTRAINT]. Character.variants is a
Mapping[str, str] and resolve_variant() appends a text delta to the prose appearance
(composition.py:67-80), while the runner creates exactly one entity per character name
(cli_movie.py:589-597). On an identity = "entity" character a variant therefore changes
the words while the entity's body plate keeps the original outfit, and the two argue inside
one generation. Until that is fixed, express costume states as separate entries in the
manifest's characters array, each with identity = "entity" and a shared face prompt —
not as variants. See MOVIE.md for the TOML.
Full method in consistency.md. The rest of the short form:
--reference-entity or @Name. Identity.--ref. Look.t2i, every other angle by i2i --ref <anchor>. Independent calls from the same paragraph produce different rooms.omni-flash 7, veo-lite / veo-fast / veo-lite-lp 3, veo-quality 0 — it accepts no references at all. So a shot carrying any --ref or entity cannot use veo-quality however much you want its quality. For a single generation needing references and quality together, that makes omni-flash the pick; a veo-lite variant when 3 refs is enough. video chain is the exception [CONSTRAINT] — it refuses omni-flash outright (its i2v is wire-verified for single generations only) and exits with a model/mode incompatibility, so chained links take a Veo 3.1 model and its cap of 3. Full table, image models included: consistency.md.One row per clip: id, camera setup, action, lines with delivery, duration, model, references.
--duration requires an explicit --model; omitted, it binds veo-lite, which renders no duration control, and exits 2.consistency.md.style → setting → geometry → cast → setup → action → dialogue → avoid.
cropdetect reported nothing on the bar-carrying clip, so check the frame's own dark-row extents instead. A plate cut from a barred clip carries the bars into every shot that references it.NAME says, weary: …. Levers, most to least reliable: volume, emotional state, pace, register, physical condition, accent.Generate one clip, run the gates, fix the template, then continue two beats per foreground call [CALIBRATED]. A detached background run lost a clip mid-poll and a template bug repeats once per clip at full price. Never loop a retry into a refusal.
Nothing is accepted on impression. Run clip_qa.py; add the transcript check when speech matters.
python clip_qa.py <clips_dir> # fluidity, lip sync, A/V drift, per clip
python clip_qa.py final.mp4 # the assembled cut
python clip_qa.py --selftest <clip.mp4> # prove the detector before believing itDirectory mode matches exactly [a-z]{2}\d{2}.mp4 — two letters then two digits, e.g.
ka01.mp4. Descriptive names and 0100.mp4 both silently match nothing and report
"no clips matched", which reads like a path error.
On a clip with no speech, run --selftest before acting on a sync verdict [CALIBRATED].
A wordless close-up flagged DRIFT sync=+0.500s r=0.7 — correlation above the 0.3 threshold,
so the gate applied. --selftest on that same clip then failed to recover a known injected
0.2 s shift (saw -0.500s), which disqualifies the reading rather than confirming it. The
detector correlates face-region motion against audio; with only wind on the track there is
nothing for it to lock onto, and it locks onto noise. This is the skill's own listed weak spot
("a lip-sync detector trusted without first proving it against a known delay") reproduced.
| Gate | Threshold | Tier |
|---|---|---|
| Stream lengths on the cut | video and audio within 0.1 s | CONSTRAINT (a mismatch is drift) |
| Speech onset after the cut | ≤ 1.6 s | CALIBRATED |
| Face-region motion floor | 10th percentile > 0.15 | CALIBRATED, 25 clips @ 24 fps 720p |
| Lip-sync lag | −0.045 s to +0.125 s, when correlation ≥ 0.3 | ITU-R BT.1359 detectability |
| Transcript word-hit | ≥ 70 % of scripted words | CALIBRATED |
| Mean volume | > −40 dB | CALIBRATED |
| Frames, by eye, 1 fps | identity, wardrobe, geometry, no text, no extra person, no border | judgment |
The whole-frame motion median does not work for dialogue [CALIBRATED]. Calibrated on moving scenes it reads above 1.0, but a locked-off talking head sits at 0.3–0.9 while performing normally. Gating on it condemns good work.
No metric can tell good motion from bad. A hallucinated object is motion, so it raises every score; the highest-scoring take of five was the broken one. The eye stays in the loop.
Failure → delete the clip, change one thing, re-check. A second identical failure means the diagnosis is wrong: restage or delete the beat rather than rewrite the prompt again.
Lane A joins with scene create and per-clip trims, free and server-side. Otherwise ffmpeg, and join with the concat filter, not the demuxer — mixed frame rates through the demuxer produced 164 s of video under 171 s of audio, lips running 4 % ahead [CALIBRATED]. Title cards need one drawtext per line; a newline inside one renders literally as "nn". Captions come from the script, not the transcript, so the reader sees the correct line even where the engine fluffed it.
Ship a review page beside the cut: the final video, and every clip with its prompt, lines, transcript, metrics and frames.
--reference-entity flags, or characters = [A, B] on one manifest scene.--duration with no --model; --ref with veo-quality (it takes 0 — reach for omni-flash).t2i calls.| Rationalisation | Reality |
|---|---|
| "Both faces must stay consistent, so both entities go in" | the second entity is the 400. One entity plus a role noun beats a refused generation. |
| "The room is described in every prompt, that is enough" | without an anchored plate chain, the first new angle invents a different room. |
| "Trim the speech so it fits 8 s" | split it across beats. The words are the deliverable. |
| "Reverse-angle drift is unavoidable" | it is avoidable: chain the angle off the anchor plate with i2i. |
| "Kick the batch off and review in the morning" | one template bug bills once per clip. Trial, gate, then pairs. |
| "It looks fine" | run the gates. Two clips that looked fine carried audible lag. |
Untested here, not discouraged. If you try one and it works, add a scored task and say so in optimization_notes.
[UNEXPLORED] seed locking across generations for consistency · a single-image storyboard sheet fed to a video model as a generation input rather than a review artefact · first-and-last-frame interpolation for transitions · custom voices and voice references · agent-mode brief cards · character archetype generation · manifest runs at large scale · reference-to-video at high reference counts · non-English delivery, which Google documents as unevaluated · and on the tooling side, a real face detector in place of clip_qa.py's fixed crop, plus frame-rate normalisation in the motion metric.
The repo ships a scored harness. Measure before and after any edit:
python scripts/dev/skillopt/harness.py --skill skills/video-production/SKILL.md \
--tasks skills/video-production/tasks.json --dry-runtasks.json beside this file holds the scored scenarios; every entry exists because an agent got it wrong in a rollout. When you find a new failure, add a task first, confirm it fails, then edit the skill until it passes. A rule added without a failing task is a guess.
consistency.md — character sheets, environment sets, prop sheets, film grammarcomposition.md — reference budget and ordering, command chains per production shapefailure-modes.md — symptom to cause to fixclip_qa.py — the gatestasks.json — the scored set© ffroliva, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in skills/video-production of ffroliva/gflow-cli.
Open the folder on GitHubat commit eb1b1ec
Video Production next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Video Production this skillffroliva/gflow-cli | 278 | — | ~7k | Automated safety check: Pass | MIT | |
| Stage EditOrkas-AI/Orkas-VideoStudio | 499 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Cassette Video EditCassette-Editor/oh-my-cassette | 119 | 1 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Media ProductionWrongStack/WrongStack | 371 | — | ~1k | Automated safety check: Pass | MIT | |
| Video Understandcalesthio/OpenMontage | 66k | — | ~841 | Automated safety check: Pass | AGPL-3.0 | |
| HyperFrames Video Entry Pointheygen-com/hyperframes | 60k | 3 repos | ~5.2k | Automated safety check: Pass | Apache-2.0 |
Orkas-AI/Orkas-VideoStudio
Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…
Cassette-Editor/oh-my-cassette
Edit, trim, cut, caption, subtitle, reframe, combine, add background music to, or export video, audio, and image files through Cassette.
WrongStack/WrongStack
Create and process finished videos with Remotion, Motion Canvas, Manim, FFmpeg or an available AI video provider.
calesthio/OpenMontage
Understand video content locally using ffmpeg frame extraction and Whisper transcription.
heygen-com/hyperframes
Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.
eternityspring/reelbench-skills
拉片:把一条成片拆成逐镜头的分析表——每个镜头的时长、景别、类别、运镜、画面. An agent skill from eternityspring/reelbench-skills.
ffroliva/gflow-cli
A skill your agent uses when the user wants to drive Google Flow (Veo image-to-video, Veo text-to-video, Imagen / Nano Banana image generation) from the terminal or a script — including…
ffroliva/gflow-cli
A skill your agent uses when triaging a GitHub issue for gflow-cli — a reporter's bug claim, a freshly-filed issue, or deciding whether and how to act on one.
ffroliva/gflow-cli
A skill your agent uses when an assessed gflow-cli issue (verdict CONFIRMED-BUG or LIKELY-BUG) has localized, verifiable scope and should be driven to a fix.
ffroliva/gflow-cli
Two-part gate for gflow-cli feature/fix work. An agent skill from ffroliva/gflow-cli.
ffroliva/gflow-cli
Auto-fix lint and formatting, then report types and tests. An agent skill from ffroliva/gflow-cli.
ffroliva/gflow-cli
Use before cutting any gflow-cli release or after a major documentation change — systematic council-driven audit that combines a mechanical 7-section checklist with a 3-agent parallel review…
Works with
Categories
A skill your agent uses when the user wants a finished video out of gflow rather than a single clip — a scripted scene, a talking-head or dialogue piece, an explainer, a product montage, a story…. Video Production is an agent skill from ffroliva/gflow-cli. Use when the user wants a finished video out of gflow rather than a single clip — a scripted scene, a talking-head or dialogue piece, an explainer, a product montage, a story sequence, an audition or rehearsal reference, a short film — or asks for consistent actors, a consistent location, a specific prop that must not change, several camera angles, captions or subtitles, or joining clips into one file.
Video Production fits situations like: the user wants a finished video out of gflow rather than a single clip — a scripted scene; A product montage; A story sequence; rehearsal reference.
Run `npx skills add ffroliva/gflow-cli --skill video-production -a claude-code`. Or copy the skill folder (skills/video-production in ffroliva/gflow-cli) into .claude/skills/video-production in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ffroliva/gflow-cli --skill video-production -a codex`. Or copy the skill folder (skills/video-production in ffroliva/gflow-cli) into .agents/skills/video-production in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ffroliva/gflow-cli --skill video-production -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-production, .gemini/skills/video-production, .github/skills/video-production and .opencode/skills/video-production in your project.
Going by SKILL.md and its folder, Video Production needs Python for the scripts in its folder and the command-line tools its instructions call (python and ffmpeg). Our summary lists: Python 3.
SKILL.md names 2 domains. In commands or code: labs.google and gstatic.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Video Production is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Video Production: Stage Edit (Orkas-AI/Orkas-VideoStudio, 499 stars), Cassette Video Edit (Cassette-Editor/oh-my-cassette, 119 stars), Media Production (WrongStack/WrongStack, 371 stars) and Video Understand (calesthio/OpenMontage, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ffroliva (a GitHub user) maintains it in ffroliva/gflow-cli, which has 278 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 11, 2026.
Source: ffroliva/gflow-cli on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.