Agent skill

Scenario Gemini Omni

by scenario-labs in scenario-labs/skills

A skill your agent uses when generating, extending, or editing video with Gemini Omni models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video keeping a…

MITAuto-check passedMedia & Creative

Install Scenario Gemini Omni

skills CLI
$ npx skills add scenario-labs/skills --skill scenario-gemini-omni -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scenario-labs/skills scenario-gemini-omni --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario-gemini-omni .claude/skills/scenario-gemini-omni && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scenario-gemini-omni
GitHub stars
946
Token cost
~1.8k tokens
SKILL.md length
965 words
Files
1
Skills in repo
146
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when generating, extending, or editing video with Gemini Omni models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video keeping a…

  • Works in 6 steps: search with target="models",… → model_schema_get with that id: required… → upload_asset two or three angles of the… → …
  • Editing video with Gemini Omni models on Scenario via MCP: text-to-video
  • SKILL.md covers Overview, Quick reference, Sound is prompted, not switched and Identity from images, change…, plus 2 more sections
  • Calls npx

What it does

Scenario Gemini Omni is an agent skill from scenario-labs/skills. Use when generating, extending, or editing video with Gemini Omni models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video keeping a character, product, or place consistent, appending footage to a clip, restyling footage by prompt (season, wardrobe, art style), native audio with dialogue and ambience, or picking between first-frame and reference conditioning. Keywords: Gemini Omni Flash, 1.1 Flash, Google, T2V, I2V, V2V, extend, subject consistency.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering AI video generation. It works with Google Gemini and Model Context Protocol. The repository describes itself as: Get production-ready images, video, audio, and 3D from any AI agent: skills that pick the right model, price before spending, and keep characters and brands consistent through… The licence is MIT.

When your agent uses it

  • Editing video with Gemini Omni models on Scenario via MCP: text-to-video
  • Image-to-video from a first frame
  • Reference-to-video keeping a character
  • Place consistent

Example prompts

  • “/scenario-gemini-omni”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. search with target="models", query="gemini omni", public=true. Pick the member by mode, here model_google-omni-flash-r2v (a live hit at…
  2. model_schema_get with that id: required fields, caps, and defaults before anything else.
  3. upload_asset two or three angles of the mascot (see the scenario skill) to get asset ids.
  4. model_run with that model_id, dry_run=true, and parameters={"prompt": "The character skips across a rain-slick plaza, catches a falling…
  5. Repeat model_run with wait=false, then jobs_wait with the returned job id, re-called with pending_job_ids on timeout, never a second…
  6. asset_display the output and review it with sound: the audio is part of the deliverable.

What it can do on your machine

Read from SKILL.md and the folder at commit f6f8ab7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scenario Gemini Omni loads about 1.8k tokens when it runs. Until then it costs about 129 tokens; SKILL.md has 965 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~129
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scenario-labs/skills at commit f6f8ab7, republished under its MIT licence (© scenario-labs). 965 words, ~1,845 tokens.

Download SKILL.mdSave it as .claude/skills/scenario-gemini-omni/SKILL.md (or your agent's skills folder).
name
scenario-gemini-omni
description
Use when generating, extending, or editing video with Gemini Omni models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video keeping a character, product, or place consistent, appending footage to a clip, restyling footage by prompt (season, wardrobe, art style), native audio with dialogue and ambience, or picking between first-frame and reference conditioning. Keywords: Gemini Omni Flash, 1.1 Flash, Google, T2V, I2V, V2V, extend, subject consistency.
license
MIT

Scenario Gemini Omni Video

Overview

Gemini Omni, Google's video family on Scenario, splits its modes across catalog members instead of folding them into one model: a text and first-frame generator (ranked first for text-to-video in public arena voting at authoring time), a reference-to-video member for subject consistency, an edit member that restyles existing footage, and on the newer 1.1 Flash line an extend member that appends footage to a clip. Every member generates audio in the same pass. Pick the line and the member by mode with search, reading ids and descriptions rather than rank (the newer 1.1 members ranked below the original line at authoring time), then treat model_schema_get as the contract. Gemini image models are the scenario-gemini-image skill's domain; Gemini TTS belongs to audio.

Connection and the core loop: see the scenario skill; model-agnostic video work: the scenario-video skill. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference

Two lines at authoring time, Flash and 1.1 Flash, one member per mode (names from the live schemas):

MemberModeInputs
Gemini Omnitext or first frameprompt and/or image, plus optional referenceImages (up to 7); on 1.1 an optional lastFrameImage beside image
Reference-to-Videoconsistent subjectsreferenceImages required (1 to 7), prompt optional; on 1.1 a referenceVideo of 3 s or less can replace the images or join up to 5 of them
Editrestyle a clipvideo and a change prompt, both required; optional referenceImages (1 to 5)
Extend (1.1 only)continue a clipvideo required, prompt optional

At authoring time both generators took duration 3 to 10 seconds (default 8) and aspectRatio 16:9 or 9:16. Edit exposes neither knob: length and shape follow the source clip. Extend inherits the shape too (no aspectRatio) but takes duration 3 to 10 (default 6) as footage appended after the source, so an 8 s clip plus 6 returns about 14 s. The original Flash members render 720p only; the 1.1 Flash line adds resolution on every member: 360p a faster draft tier, 720p the default, 1080p and 4k upscaled. The tier does not move price (at authoring time a generation quoted the same at 360p and 1080p, and so did Extend), so a cheap draft is a shorter duration, not a lower tier. Seed and negative prompt exist nowhere in the family. Cost moves with duration, a first-frame image, Reference-to-Video's references, and Edit's and Extend's source clips; Edit spanned the family's widest cost range at authoring time and its jobs ran about twice as long, so dry_run before editing anything long.

Sound is prompted, not switched

No member has an audio parameter: every clip arrives with generated dialogue, ambience, and effects, and the prompt is the only lever. On the generators, end the prompt with an explicit audio line naming what should be heard ("crackling campfire, distant owls") and put spoken words in quotes; leave it out and the model chooses for you. On Edit, never prompt audio: it regenerates to match the new look on its own.

Show full SKILL.md (442 more words)Show less

Identity from images, change from words

Reference-to-Video takes the subject's look from referenceImages, and re-describing it in the prompt fights the images. Call the subject "the character" or "the product" and spend the words on action, setting, camera, and sound; several angles of one subject tighten the hold, distinct subjects share one scene, and with no prompt at all the subjects still appear. There is no first-frame anchor here: when the clip must open on an exact composition, pass that still as image on the base member, which also accepts references alongside it and, on 1.1, a lastFrameImage to interpolate towards (it needs the first frame).

Edit is the inverse: motion, camera path, and timing are locked from the source video, only the look moves. Lead the prompt with the change ("make it winter", "swap the red car for a vintage blue Beetle"), name the target concretely, and never ask for new choreography, cuts, or camera moves: those instructions will not take.

Worked example: a mascot short with a consistent character

  1. search with target="models", query="gemini omni", public=true. Pick the member by mode, here model_google-omni-flash-r2v (a live hit at authoring time: re-discover each session).
  2. model_schema_get with that id: required fields, caps, and defaults before anything else.
  3. upload_asset two or three angles of the mascot (see the scenario skill) to get asset ids.
  4. model_run with that model_id, dry_run=true, and parameters={"prompt": "The character skips across a rain-slick plaza, catches a falling leaf, and holds it up in triumph. Overcast soft light, low tracking shot. Audio: light rain, footsteps on wet stone, one bright chirp.", "referenceImages": ["asset_a", "asset_b"], "duration": 8, "aspectRatio": "16:9"}; references and duration move the cost, so re-estimate after changing either.
  5. Repeat model_run with wait=false, then jobs_wait with the returned job id, re-called with pending_job_ids on timeout, never a second model_run.
  6. asset_display the output and review it with sound: the audio is part of the deliverable.

Common mistakes

  • Describing the reference subject's appearance in the prompt: identity comes from the images; write action and setting around "the character".
  • Asking Edit for new motion, cuts, or re-timing: they are locked to the source; prompt only the look change.
  • Passing duration or aspectRatio to Edit, or aspectRatio to Extend: not in their schemas; the output follows the source clip.
  • Hunting for an audio toggle: none exists; steer sound in the prompt or accept the model's choice.
  • Restating a first-frame image as a static scene: describe the motion continuing from it.
  • Packing a multi-scene story into one run: 3 to 10 seconds holds one continuous beat; carry a beat past 10 seconds with Extend, or sequence shots with the scenario-video-assembly skill.

© scenario-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scenario-gemini-omni of scenario-labs/skills.

Open the folder on GitHubat commit f6f8ab7

Compare with similar skills

Scenario Gemini Omni next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scenario Gemini Omni compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scenario Gemini Omni this skillscenario-labs/skills946—~1.8kAutomated safety check: PassMIT
Fal AI Mediaaffaan-m/ECC277k4 repos~1.9kAutomated safety check: PassMIT
Fal AI Mediaaffaan-m/ECC277k2 repos~1.2kAutomated safety check: PassMIT
Fal AI Mediaaffaan-m/ECC276k—~1.4kAutomated safety check: PassMIT
Higgsfield Marketing StudioOSideMedia/higgsfield-ai-prompt-skill713—~15kAutomated safety check: PassMIT
Avatar Videocalesthio/OpenMontage66k—~1.6kAutomated safety check: PassAGPL-3.0

Similar skills

  • Fal AI Media

    affaan-m/ECC

    Unified media generation via fal.ai MCP — image, video, and audio.

    277k GitHub starsUsed in 4 repos~1.9k tokens
    Media & CreativeAuto-check passed
  • Fal AI Media

    affaan-m/ECC

    通过 fal.ai MCP 实现统一的媒体生成——图像、视频和音频。涵盖文本到图像(Nano Banana)、文本/图像到视频(Seedance、Kling、Veo 3)、文本到语音(CSM-1B),以及视频到音频(ThinkSound)。当用户想要使用 AI 生成图像、视频或音频时使用。

    277k GitHub starsUsed in 2 repos~1.2k tokens
    Media & CreativeAuto-check passed
  • Fal AI Media

    affaan-m/ECC

    fal.ai MCPによる統合メディア生成(画像、動画、音声)。テキストから画像(Nano Banana)、テキスト/画像から動画(Seedance、Kling、Veo 3)、テキストから音声(CSM-1B)、動画から音声(ThinkSound)をカバーします。ユーザーがAIで画像、動画、音声を生成したい場合に使用します。

    276k GitHub stars~1.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Higgsfield Marketing Studio

    OSideMedia/higgsfield-ai-prompt-skill

    A skill your agent uses when the user mentions Marketing Studio, DTC Ads, ad video, UGC video, the marketingstudiovideo MCP model, or wants to generate one of the 9 Marketing Studio ad presets (UGC…

    713 GitHub stars~15k tokensUpdated 14 days ago
    Media & CreativeAuto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    66k GitHub stars~1.3k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed

More from scenario-labs/skills

All 146 skills in this repo
  • Scenario Blender Grease Pencil

    scenario-labs/skills

    A skill your agent uses when drawing or animating with Grease Pencil in Blender 5.x from Python: 2D or 2.5D illustration, frame-by-frame animation, a cutout or part-based 2D character, strokes with…

    946 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Blender Hair

    scenario-labs/skills

    A skill your agent uses when grooming hair or fur in Blender with hair curves, such as a character hairstyle, animal fur, procedural fur in geometry nodes, or hair cards and mesh hair for games.

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when lighting, rendering or compositing in Blender: light a character, product or hero shot, interior at dusk or night, three-point or motivated lighting, sun and sky, HDRI…

    946 GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Chatgpt Pet Create

    scenario-labs/skills

    A skill your agent uses when creating a ChatGPT pet or Codex pet with Scenario: hatching an animated companion from a text idea, a character, mascot or brand cue, or reference photos and art; making…

    946 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Animation

    scenario-labs/skills

    A skill your agent uses when animating characters or scenes in Godot 4.7: AnimationPlayer clips and RESET, AnimationTree state machines and blend spaces built in code, Mixamo or glTF import, loop…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Audio

    scenario-labs/skills

    A skill your agent uses when adding or fixing sound in Godot 4.7: audio buses and effects, volume sliders, 'too many sounds', combat audio with hundreds of enemies, sounds clipping or distorting, 3D…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed

Questions about Scenario Gemini Omni

What does Scenario Gemini Omni do?

A skill your agent uses when generating, extending, or editing video with Gemini Omni models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video keeping a…. Scenario Gemini Omni is an agent skill from scenario-labs/skills. Use when generating, extending, or editing video with Gemini Omni models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video keeping a character, product, or place consistent, appending footage to a clip, restyling footage by prompt (season, wardrobe, art style), native audio with dialogue and ambience, or picking between first-frame and reference conditioning.

When should I use Scenario Gemini Omni?

Scenario Gemini Omni fits situations like: editing video with Gemini Omni models on Scenario via MCP: text-to-video; image-to-video from a first frame; reference-to-video keeping a character; place consistent.

How do I install Scenario Gemini Omni in Claude Code?

Run `npx skills add scenario-labs/skills --skill scenario-gemini-omni -a claude-code`. Or copy the skill folder (skills/scenario-gemini-omni in scenario-labs/skills) into .claude/skills/scenario-gemini-omni in your project. Claude Code loads it when a task matches its description.

How do I install Scenario Gemini Omni in Codex?

Run `npx skills add scenario-labs/skills --skill scenario-gemini-omni -a codex`. Or copy the skill folder (skills/scenario-gemini-omni in scenario-labs/skills) into .agents/skills/scenario-gemini-omni in your project. Codex loads it when a task matches its description.

Can I use Scenario Gemini Omni in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scenario-labs/skills --skill scenario-gemini-omni -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-gemini-omni, .gemini/skills/scenario-gemini-omni, .github/skills/scenario-gemini-omni and .opencode/skills/scenario-gemini-omni in your project.

What does Scenario Gemini Omni need to run?

Going by SKILL.md and its folder, Scenario Gemini Omni needs the command-line tools its instructions call (npx). Our summary lists: Node.js.

Does Scenario Gemini Omni access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Scenario Gemini Omni safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scenario Gemini Omni use?

Scenario Gemini Omni is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scenario Gemini Omni use?

About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scenario Gemini Omni?

Skills that share tags, products or a category with Scenario Gemini Omni: Fal AI Media (affaan-m/ECC, 277k stars), Fal AI Media (affaan-m/ECC, 277k stars), Fal AI Media (affaan-m/ECC, 276k stars) and Higgsfield Marketing Studio (OSideMedia/higgsfield-ai-prompt-skill, 713 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scenario Gemini Omni?

scenario-labs (a GitHub organization) maintains it in scenario-labs/skills, which has 946 GitHub stars. The repository holds 146 skills in this directory. The repository was last updated on October 10, 2026.

Source: scenario-labs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.