Agent skill

Scenario Grok Imagine Video

by scenario-labs in scenario-labs/skills

A skill your agent uses when generating or editing video with Grok Imagine models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video with @image tags for…

MITAuto-check passedMedia & Creative

Install Scenario Grok Imagine Video

skills CLI
$ npx skills add scenario-labs/skills --skill scenario-grok-imagine-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scenario-labs/skills scenario-grok-imagine-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario-grok-imagine-video .claude/skills/scenario-grok-imagine-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scenario-grok-imagine-video
GitHub stars
946
Token cost
~1.6k tokens
SKILL.md length
836 words
Files
1
Skills in repo
146
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when generating or editing video with Grok Imagine models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video with @image tags for…

  • Works in 6 steps: search with target="models", query="grok… → model_schema_get with that id: fields,… → upload_asset the hero still (see the… → …
  • Editing video with Grok Imagine models on Scenario via MCP: text-to-video
  • SKILL.md covers Overview, Quick reference, First frame or references and Sound is prompted, not switched, plus 2 more sections
  • Calls npx

What it does

Scenario Grok Imagine Video is an agent skill from scenario-labs/skills. Use when generating or editing video with Grok Imagine models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video with @image tags for consistent people, products, or clothing, prompt-based editing, extending a clip from its last frame, native audio with lip-synced dialogue and sound effects, or choosing between first-frame and reference conditioning. Keywords: Grok Imagine Video 1.5, R2V, Grok Edit Video, Grok Extend Video, xAI, T2V, I2V, V2V.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering AI video generation, Video production and Music and audio generation. It works with Model Context Protocol. The repository describes itself as: Get production-ready images, video, audio, and 3D from any AI agent: skills that pick the right model, price before spending, and keep characters and brands consistent through… The licence is MIT.

When your agent uses it

  • Editing video with Grok Imagine models on Scenario via MCP: text-to-video
  • Image-to-video from a first frame
  • Reference-to-video with @image tags for consistent people
  • Prompt-based editing

Example prompts

  • “/scenario-grok-imagine-video”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. search with target="models", query="grok imagine video", public=true. Match the member by name: first-frame animation wants e.g…
  2. model_schema_get with that id: fields, caps, and defaults before anything else.
  3. upload_asset the hero still (see the scenario skill) to get its asset id.
  4. model_run with that model_id, dry_run=true, and the exact parameters={"prompt": "The frame shows a knight on a cliff at dusk. She lowers…
  5. Repeat model_run with wait=false, then jobs_wait with the returned job id, re-called with pending_job_ids on timeout, never a second…
  6. asset_display the output and watch it with sound: the audio direction is judged by ear.

What it can do on your machine

Read from SKILL.md and the folder at commit f6f8ab7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scenario Grok Imagine Video loads about 1.6k tokens when it runs. Until then it costs about 129 tokens; SKILL.md has 836 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~129
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scenario-labs/skills at commit f6f8ab7, republished under its MIT licence (© scenario-labs). 836 words, ~1,591 tokens.

Download SKILL.mdSave it as .claude/skills/scenario-grok-imagine-video/SKILL.md (or your agent's skills folder).
name
scenario-grok-imagine-video
description
Use when generating or editing video with Grok Imagine models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video with @image tags for consistent people, products, or clothing, prompt-based editing, extending a clip from its last frame, native audio with lip-synced dialogue and sound effects, or choosing between first-frame and reference conditioning. Keywords: Grok Imagine Video 1.5, R2V, Grok Edit Video, Grok Extend Video, xAI, T2V, I2V, V2V.
license
MIT

Scenario Grok Imagine Video

Overview

Grok Imagine, xAI's video family on Scenario, splits its modes across single-purpose members: text and first-frame generation, reference-to-video, prompt editing, and clip extension each live in their own model, so picking the member is picking the mode. Discover them with search and treat model_schema_get as the contract: members agree on prompt style and disagree on every cap. The Grok Imagine image models belong to the scenario-grok-imagine-image skill.

Connection and the core loop: see the scenario skill in this repo; model-agnostic video work: the scenario-video skill. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference

Pick the member by intent (input names from the live schema):

IntentMemberInputs
Text-to-videoVideo 1.5prompt alone; aspectRatio auto lands on 16:9
Animate a stillVideo 1.5image + prompt; auto follows the image ratio
Identity, opening freeVideo 1.5 R2VreferenceImages + prompt tagging @image1
Restyle or swap in placeEdit Videovideo + a prompt naming only the changes
Continue a clipExtend Videovideo + a continuation prompt; duration is new seconds

Caps are per member, so read them off model_schema_get. At authoring time the generation members took 1 to 15 seconds, 1 to 4 numOutputs, and eight aspectRatio values including auto; 1080p existed only on Video 1.5's text and first-frame modes, while R2V stopped at 720p and defaulted to 480p. The earlier unversioned Grok Imagine Video takes the same inputs as Video 1.5, also caps at 720p, and runs cheaper per dry_run. R2V took up to 7 referenceImages; Extend generated 2 to 10 new seconds; Edit exposed no duration or resolution and preprocessed its source down to 8.7 seconds at 720p, so trim to the segment first (see scenario-video). No seed, negative prompt, or camera parameter exists anywhere in the family.

First frame or references

image locks frame one: the video opens on that exact composition and animates out of it. referenceImages (its own member, R2V) carries people, products, and clothing across the shot without deciding how it opens. Tags bind by array order: @image1 is referenceImages[0] (<IMAGE_1> also works), and one clean subject per reference keeps control; a busy reference dilutes it. If the opening must match a composition exactly, render that still and pass it as image; when only a face, product, or outfit must stay consistent, use R2V and tag each reference where it acts.

Sound is prompted, not switched

Generation members produce native audio and no parameter controls it: the prompt is the whole mixing desk. Left unaddressed, the track tends toward generic background music, so end every prompt with named sounds ("Audio: rain on glass, distant traffic") or "no music". Dialogue written in quotes with a delivery verb lip-syncs: She says calmly: "We're live." Structure the rest like a director: scene, camera, style, motion, audio, in present tense, one continuous shot, one dominant camera move, physical actions instead of named emotions. 80 to 150 words touching at least three of those layers beat a one-line scene; editing prompts run shorter, naming only what changes. Extension prompts describe the next beat, not a new scene: open with a bridge ("the shot continues"), keep the established light and cast, and restate the audio.

Show full SKILL.md (283 more words)Show less

Worked example: dialogue over an animated hero still

  1. search with target="models", query="grok imagine video", public=true. Match the member by name: first-frame animation wants e.g. model_xai-grok-imagine-video-1-5 (a live hit at authoring time: re-discover each session).
  2. model_schema_get with that id: fields, caps, and defaults before anything else.
  3. upload_asset the hero still (see the scenario skill) to get its asset id.
  4. model_run with that model_id, dry_run=true, and the exact parameters={"prompt": "The frame shows a knight on a cliff at dusk. She lowers her sword, turns to the camera, and says quietly: 'It ends tonight.' Slow dolly-in, wind lifting her cloak. Audio: wind, distant thunder, her line clear. No music.", "image": "asset_abc", "duration": 8, "resolution": "1080p"} for the cost estimate; re-estimate after any change to duration, resolution, or numOutputs.
  5. Repeat model_run with wait=false, then jobs_wait with the returned job id, re-called with pending_job_ids on timeout, never a second model_run.
  6. asset_display the output and watch it with sound: the audio direction is judged by ear.

Common mistakes

  • Leaving audio to the default: generic music appears; direct the sound or state "no music" in every prompt.
  • Prompting an exact opening frame in R2V: references guide identity, not frame one; pass the still as image on Video 1.5.
  • Expecting 1080p everywhere: at authoring time it existed only on Video 1.5 text and first-frame runs, and R2V silently defaults to 480p.
  • Feeding Edit Video a long clip and expecting full-length output: the source is preprocessed down (8.7 seconds at authoring time); trim to the segment first.
  • Reading Extend's duration as total length: it counts only new footage.
  • Stacking camera moves or contradictory directions ("zoom in as the camera pulls back"): one dominant move per shot.

© scenario-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scenario-grok-imagine-video of scenario-labs/skills.

Open the folder on GitHubat commit f6f8ab7

Compare with similar skills

Scenario Grok Imagine Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scenario Grok Imagine Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scenario Grok Imagine Video this skillscenario-labs/skills946—~1.6kAutomated safety check: PassMIT
Tesseract Videomirage-hq/Tesseract130—~1.9kAutomated safety check: PassCustom licence
Audio And Videoglifxyz/glif-mcp-server213—~1.6kAutomated safety check: PassMIT
Pneuma Clipcraftpandazki/pneuma-skills161—~7.5kAutomated safety check: NotesMIT
Edit VideoArcReel/ArcReel5.4k—~1.6kAutomated safety check: PassAGPL-3.0
Avatar Videocalesthio/OpenMontage66k—~1.6kAutomated safety check: PassAGPL-3.0

Similar skills

  • Tesseract Video

    mirage-hq/Tesseract

    Edit existing footage into finished videos locally with Tesseract.

    130 GitHub stars~1.9k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Audio And Video

    glifxyz/glif-mcp-server

    Make or edit audio and video with Glif, from a text brief or from a reference image, video or audio file.

    213 GitHub stars~1.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Pneuma Clipcraft

    pandazki/pneuma-skills

    AI-orchestrated video production on @pneuma-craft. An agent skill from pandazki/pneuma-skills.

    161 GitHub stars~7.5k tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • Edit Video

    ArcReel/ArcReel

    在剪辑时间线上剪辑一集,并按要求出成片或导出剪映草稿。本集视频已齐、要剪一版时使用;用户按片段编号提出修改(如「c3 太拖」),或要求配 BGM、出成片、导出剪映草稿、复制或重命名剪辑时间线、回到旧修订时,也使用本 skill。

    5.4k GitHub stars~1.6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • VRGDG H3 Short Film Pipeline

    vrgamegirl19/comfyui-vrgamedevgirl

    Builds an AI short film in a local ComfyUI with the VRGDG Video Builder, MiniMax H3 scenes, reference images, a music score, QA and a final edit.

    765 GitHub stars~4.2k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from scenario-labs/skills

All 146 skills in this repo
  • Scenario Blender Grease Pencil

    scenario-labs/skills

    A skill your agent uses when drawing or animating with Grease Pencil in Blender 5.x from Python: 2D or 2.5D illustration, frame-by-frame animation, a cutout or part-based 2D character, strokes with…

    946 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Blender Hair

    scenario-labs/skills

    A skill your agent uses when grooming hair or fur in Blender with hair curves, such as a character hairstyle, animal fur, procedural fur in geometry nodes, or hair cards and mesh hair for games.

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when lighting, rendering or compositing in Blender: light a character, product or hero shot, interior at dusk or night, three-point or motivated lighting, sun and sky, HDRI…

    946 GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Chatgpt Pet Create

    scenario-labs/skills

    A skill your agent uses when creating a ChatGPT pet or Codex pet with Scenario: hatching an animated companion from a text idea, a character, mascot or brand cue, or reference photos and art; making…

    946 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Animation

    scenario-labs/skills

    A skill your agent uses when animating characters or scenes in Godot 4.7: AnimationPlayer clips and RESET, AnimationTree state machines and blend spaces built in code, Mixamo or glTF import, loop…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Audio

    scenario-labs/skills

    A skill your agent uses when adding or fixing sound in Godot 4.7: audio buses and effects, volume sliders, 'too many sounds', combat audio with hundreds of enemies, sounds clipping or distorting, 3D…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed

Questions about Scenario Grok Imagine Video

What does Scenario Grok Imagine Video do?

A skill your agent uses when generating or editing video with Grok Imagine models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video with @image tags for…. Scenario Grok Imagine Video is an agent skill from scenario-labs/skills. Use when generating or editing video with Grok Imagine models on Scenario via MCP: text-to-video, image-to-video from a first frame, reference-to-video with @image tags for consistent people, products, or clothing, prompt-based editing, extending a clip from its last frame, native audio with lip-synced dialogue and sound effects, or choosing between first-frame and reference conditioning.

When should I use Scenario Grok Imagine Video?

Scenario Grok Imagine Video fits situations like: editing video with Grok Imagine models on Scenario via MCP: text-to-video; image-to-video from a first frame; reference-to-video with @image tags for consistent people; prompt-based editing.

How do I install Scenario Grok Imagine Video in Claude Code?

Run `npx skills add scenario-labs/skills --skill scenario-grok-imagine-video -a claude-code`. Or copy the skill folder (skills/scenario-grok-imagine-video in scenario-labs/skills) into .claude/skills/scenario-grok-imagine-video in your project. Claude Code loads it when a task matches its description.

How do I install Scenario Grok Imagine Video in Codex?

Run `npx skills add scenario-labs/skills --skill scenario-grok-imagine-video -a codex`. Or copy the skill folder (skills/scenario-grok-imagine-video in scenario-labs/skills) into .agents/skills/scenario-grok-imagine-video in your project. Codex loads it when a task matches its description.

Can I use Scenario Grok Imagine Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scenario-labs/skills --skill scenario-grok-imagine-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-grok-imagine-video, .gemini/skills/scenario-grok-imagine-video, .github/skills/scenario-grok-imagine-video and .opencode/skills/scenario-grok-imagine-video in your project.

What does Scenario Grok Imagine Video need to run?

Going by SKILL.md and its folder, Scenario Grok Imagine Video needs the command-line tools its instructions call (npx). Our summary lists: Node.js.

Does Scenario Grok Imagine Video access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Scenario Grok Imagine Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scenario Grok Imagine Video use?

Scenario Grok Imagine Video is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scenario Grok Imagine Video use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scenario Grok Imagine Video?

Skills that share tags, products or a category with Scenario Grok Imagine Video: Tesseract Video (mirage-hq/Tesseract, 130 stars), Audio And Video (glifxyz/glif-mcp-server, 213 stars), Pneuma Clipcraft (pandazki/pneuma-skills, 161 stars) and Edit Video (ArcReel/ArcReel, 5.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scenario Grok Imagine Video?

scenario-labs (a GitHub organization) maintains it in scenario-labs/skills, which has 946 GitHub stars. The repository holds 146 skills in this directory. The repository was last updated on October 10, 2026.

Source: scenario-labs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.