Agent skill

Scenario Model Comparison

by scenario-labs in scenario-labs/skills

A skill your agent uses when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a…

MITAuto-check passedMarketing & SEO

Install Scenario Model Comparison

skills CLI
$ npx skills add scenario-labs/skills --skill scenario-model-comparison -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scenario-labs/skills scenario-model-comparison --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario-model-comparison .claude/skills/scenario-model-comparison && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scenario-model-comparison
GitHub stars
931
Token cost
~2.6k tokens
SKILL.md length
1,454 words
Files
1
Skills in repo
143
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a…

  • Works in 7 steps: Frame: "a rusty iron key with a… → recommend with that brief as prompt and… → model_schema_get on each: prompt cap,… → …
  • Comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs
  • SKILL.md covers Overview, Quick reference, Normalization and Budget, plus 4 more sections
  • Calls npx

What it does

Scenario Model Comparison is an agent skill from scenario-labs/skills. Use when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a default model, testing a new release against the incumbent, or finding which model handles a style, edit, or reference image best. Keywords: model comparison, benchmark, bake-off, A/B test, contact sheet, cost per asset, latency.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Marketing & SEO, covering A/B testing. It works with Model Context Protocol. The repository describes itself as: Get production-ready images, video, audio, and 3D from any AI agent: skills that pick the right model, price before spending, and keep characters and brands consistent through… The licence is MIT.

When your agent uses it

  • Comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs
  • Generation time and quality measured side by side
  • Choosing a default model
  • Testing a new release against the incumbent

Example prompts

  • “/scenario-model-comparison”

Requirements

  • Node.js

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Frame: "a rusty iron key with a skull-shaped bow, hand-painted style, plain background", a 1024 square, criteria: skull bow present…
  2. recommend with that brief as prompt and limit: 5, then search with target="models", public=true for the two models the user named; keep…
  3. model_schema_get on each: prompt cap, size fields, sample-count field, any prompt-expansion flag.
  4. model_run with dry_run: true (top-level) on each, the same prompt, the count field at one, the closest 1024 square each allows, the…
  5. model_run three times with wait: false, then one jobs_wait with the three job ids, re-called with pending_job_ids on a timeout; read…
  6. model_scenario-grid-maker with images as the three asset ids in candidate order and columns: 3; asset_display the sheet, then each asset…
  7. Deliver the table (model, parameters that differed, cost, seconds, dimensions, criteria passed, notes) and file the three outputs and the…

What it can do on your machine

Read from SKILL.md and the folder at commit 91caa01. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scenario Model Comparison loads about 2.6k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 1,454 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scenario-labs/skills at commit 91caa01, republished under its MIT licence (© scenario-labs). 1,454 words, ~2,635 tokens.

Download SKILL.mdSave it as .claude/skills/scenario-model-comparison/SKILL.md (or your agent's skills folder).
name
scenario-model-comparison
description
Use when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a default model, testing a new release against the incumbent, or finding which model handles a style, edit, or reference image best. Keywords: model comparison, benchmark, bake-off, A/B test, contact sheet, cost per asset, latency.
license
MIT

Scenario Model Comparison

Overview

A comparison is one brief run unchanged across several models, then judged on criteria written down before the first paid call. Every candidate has its own contract, so the comparison lives in the normalization: what is held fixed, what each schema forces to differ, and where each number came from. recommend shortlists, dry_run prices, the jobs_wait rows carry the billed cost, the grid tool puts the results side by side, and a table delivers. Connection and the core loop: see the scenario skill; a rubric for judging output against a brief: scenario-refine-loop. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference

StepDo
1. FrameOne brief, its inputs (prompt text, reference asset ids), three to five pass/fail criteria, all written before running
2. Shortlistrecommend with the brief as prompt (limit up to 10); search for candidates the user named; keep three to five
3. Normalizemodel_schema_get each candidate: the shared fields, the caps that differ, the flags that rewrite prompts
4. Pricemodel_run with dry_run: true (top-level, beside model_id, never inside parameters) per candidate and utility model; apply the Budget gate
5. Runmodel_run with top-level wait: false per candidate, launched back to back, then one jobs_wait over all the job ids
6. MeasurecuCost from the jobs_wait rows; createdAt to updatedAt from job_get for the seconds
7. JudgeThe sheet from model_scenario-grid-maker, then each asset at full size, scored against the criteria
8. DeliverOne table row per candidate; the assets filed in one collection and tagged by model id

Normalization

  • Hold the brief fixed: the same prompt text byte for byte, the same reference asset ids, the same output count (one, unless every candidate exposes the same batch field; a model whose default sample count is above one multiplies its cost and must be pinned).
  • Let the schemas differ only where they must. Sizes are enums or pixel pairs per model: pick the closest each allows to the brief's target and record the delivered width and height from asset_get beside the result, since a larger output costs more and reads sharper. Prompt caps differ (max_length): an overrun is a 400, never an implicit trim. Preserve an exact user brief: replace a candidate whose cap is too small, or ask before changing the brief for every candidate.
  • Turn off what rewrites the prompt: a prompt-expansion or optimize flag left on for one candidate compares its rewrite, not the brief. Where a flag cannot be turned off, say so in the table.
  • Seeds do not compare across models: a seed reproduces one model's own sampling and transfers nothing. Leave seed unset, or set it only for repeat runs of the same candidate.
  • A Turbo and a Quality member of one family are two candidates, not one. A LoRA runs through its base (runs_as and run_with, see scenario), and its cost and time are the base run's.
  • Where one family's skill prescribes tuning (a reference tag syntax, a strength dial), fairness means two rows for that candidate, defaults and tuned, rather than a tuned one against untuned rivals.

Budget

Price generations and required grid, extraction, or analysis steps. For a hard total cap, count each launched job once, including in-flight commitments. Before the first and every later paid call, require: committed spend + the next call's quote + verified reservations for remaining required work <= the cap. A substitute-input quote cannot verify a reservation unless documented pricing proves it bounds the real payload. If the sheet needs outputs that do not yet exist and has no verified bound, do not launch the generations: propose staged spending or a revised plan and wait for the user's agreement. Do not silently drop requested candidates or artifacts to fit. usage reports consumed CU, not remaining credits; it cannot certify a balance for this budget gate.

Measurement

recommend reports modeled cost and latency, right for the shortlist and wrong as a result: dry_run prices the exact payload, and the jobs_wait row's cuCost is what was billed. Time comes from job_get, which returns createdAt and updatedAt; their difference on a finished job is the wall-clock latency including queue time, comparable across candidates launched within the same minute. Launch every candidate before waiting on any, within the team's concurrency ceiling (the 429 rows in scenario). Use the returned actionLimit on a concurrency 429; never infer the ceiling from this comparison's batch size. Re-call jobs_wait with pending_job_ids until every row is terminal. Record per candidate: model id, the exact parameters sent, cuCost, seconds, delivered dimensions, asset ids, and any normalization it forced (a size, a cap, a flag).

Video candidates are compared on contact sheets, never on a first frame: sweep each clip with model_scenario-video-to-image-seq (a fixed first-party id, Scenario's single deterministic frame extractor, so discovery would only re-derive it), wait for the extraction job, then sheet the frames per candidate; the extractor's frame-order and stride contract and the sheet's 100-image cap are in scenario-video-editing. Audio and 3D candidates are compared on the assets themselves through asset_display.

Show full SKILL.md (596 more words)Show less

Judging

Pre-registered criteria are pass/fail statements about the output ("the label text is legible", "the scar is on the left cheek", "no extra fingers", "the background is plain"), so a result is scored, not admired, and a surprising winner cannot rewrite the test after the fact. Score the literal criterion: a related feature or partial match is not a pass; ambiguous evidence is unverified. Build the sheet with model_scenario-grid-maker (a fixed first-party id, Scenario's single deterministic grid tool, so discovery would only re-derive it): images in candidate order, wrapped as an array even for one, since the array order is the legend (the tool has no labels), columns equal to the candidate count so one row is one brief, cellRatio matching the outputs, padding for a gutter. Then asset_display each output at full size: a sheet hides fine text, edge halos and small anatomy. Score every criterion for every candidate, and put cost and seconds in the same table so the trade-off is read in one place. For a large set, asset_analyze (catalog, cost-bearing, and write-class: scenario_tools_search then scenario_tool_execute_write, see scenario) can score a batch: the criteria list verbatim as its instruction, the outputs as images, one verdict per criterion per image, and a criterion the image cannot settle at output resolution (small lettering, a fine edge) recorded as unverified, never as a pass; pre-register that instruction with the criteria.

Worked example: three image models on one prop brief

  1. Frame: "a rusty iron key with a skull-shaped bow, hand-painted style, plain background", a 1024 square, criteria: skull bow present, exactly one key, no lettering, plain background.
  2. recommend with that brief as prompt and limit: 5, then search with target="models", public=true for the two models the user named; keep three. Never hardcode the ids: catalogs differ per team.
  3. model_schema_get on each: prompt cap, size fields, sample-count field, any prompt-expansion flag.
  4. model_run with dry_run: true (top-level) on each, the same prompt, the count field at one, the closest 1024 square each allows, the expansion flag off; write the three prices down and apply the Budget gate, including the required sheet.
  5. model_run three times with wait: false, then one jobs_wait with the three job ids, re-called with pending_job_ids on a timeout; read cuCost off each row and createdAt and updatedAt off job_get for each job.
  6. model_scenario-grid-maker with images as the three asset ids in candidate order and columns: 3; asset_display the sheet, then each asset; score the four criteria.
  7. Deliver the table (model, parameters that differed, cost, seconds, dimensions, criteria passed, notes) and file the three outputs and the sheet in a collection (collection_create, then collection_add_assets), tagging each output with its model id through asset_add_tags.

Common mistakes

  • Quoting recommend's numbers as results: they are modeled. Bill from jobs_wait, time from job_get.
  • Different sizes or sample counts across candidates, or one candidate's prompt expansion left on: the comparison then measures the payload difference.
  • One sample per candidate on a brief whose output varies wildly: run two or three where the budget allows, and say how many in the table.
  • Reading a winner as a constant: the ranking is per brief and per team catalog. Re-run when either changes, and never write the winning id into a skill or a pipeline as a fixed value.
  • Comparing video by first frame or a single still: sweep the clips into contact sheets.
  • Shipping the sheet as the deliverable: the table with costs and criteria is the result; the sheet is evidence.
  • Pasting signed download URLs into the report: file the assets and share the asset ids.

© scenario-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scenario-model-comparison of scenario-labs/skills.

Open the folder on GitHubat commit 91caa01

Compare with similar skills

Scenario Model Comparison next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scenario Model Comparison compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scenario Model Comparison this skillscenario-labs/skills931—~2.6kAutomated safety check: PassMIT
Ab Test Analyzeririnabuht12-oss/marketing-skills4k—~1.4kAutomated safety check: PassNone
Sealeap Dijiang Amazon Review Mined Image Iterationxjli360/sealeap-amazon-skills247—~748Automated safety check: PassMIT
Sealeap Taotie Amazon Image Script And Ab Testingxjli360/sealeap-amazon-skills247—~818Automated safety check: PassMIT
Continue Projectglifxyz/glif-mcp-server212—~734Automated safety check: PassMIT
Ab Testingcoreyhaines31/marketingskills54k3 repos~3.1kAutomated safety check: PassMIT

Similar skills

  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    4k GitHub stars~1.4k tokensUpdated 16 days ago
    Marketing & SEOAuto-check passed
  • Mine competitor reviews with AI assistance to surface recurring pain points and desired improvements, benchmark the main image against top competitors, validate candidate images with quick…

    247 GitHub stars~748 tokensUpdated 12 days ago
    Marketing & SEOAuto-check passed
  • Build a listing image script from competitor review insights and traffic keywords—relevance, value proposition and call to action in the first frames, then scenes, details, care and sizing with…

    247 GitHub stars~818 tokensUpdated 12 days ago
    Marketing & SEOAuto-check passed
  • Continue Project

    glifxyz/glif-mcp-server

    Find the user's saved Glif project, show its media, and add new work to it without starting a separate project.

    212 GitHub stars~734 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Ab Testing

    coreyhaines31/marketingskills

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

    54k GitHub starsUsed in 3 repos~3.1k tokens
    Marketing & SEOAuto-check passed
  • Analytics

    Nexus-JPF/note-companion

    When the user wants to set up, improve, or audit analytics tracking and measurement.

    870 GitHub starsUsed in 7 repos~2.2k tokens
    Marketing & SEOAuto-check passed

More from scenario-labs/skills

All 143 skills in this repo
  • Scenario Blender Grease Pencil

    scenario-labs/skills

    A skill your agent uses when drawing or animating with Grease Pencil in Blender 5.x from Python: 2D or 2.5D illustration, frame-by-frame animation, a cutout or part-based 2D character, strokes with…

    931 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Blender Hair

    scenario-labs/skills

    A skill your agent uses when grooming hair or fur in Blender with hair curves, such as a character hairstyle, animal fur, procedural fur in geometry nodes, or hair cards and mesh hair for games.

    931 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when lighting, rendering or compositing in Blender: light a character, product or hero shot, interior at dusk or night, three-point or motivated lighting, sun and sky, HDRI…

    931 GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Animation

    scenario-labs/skills

    A skill your agent uses when animating characters or scenes in Godot 4.7: AnimationPlayer clips and RESET, AnimationTree state machines and blend spaces built in code, Mixamo or glTF import, loop…

    931 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Audio

    scenario-labs/skills

    A skill your agent uses when adding or fixing sound in Godot 4.7: audio buses and effects, volume sliders, 'too many sounds', combat audio with hundreds of enemies, sounds clipping or distorting, 3D…

    931 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Multiplayer

    scenario-labs/skills

    A skill your agent uses when a Godot 4 game goes online or co-op: host and join with ENet, WebSocket for a web build, RPCs (@rpc, rpcid, anypeer), MultiplayerSpawner and MultiplayerSynchronizer…

    931 GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed

Questions about Scenario Model Comparison

What does Scenario Model Comparison do?

A skill your agent uses when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a…. Scenario Model Comparison is an agent skill from scenario-labs/skills. Use when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a default model, testing a new release against the incumbent, or finding which model handles a style, edit, or reference image best.

When should I use Scenario Model Comparison?

Scenario Model Comparison fits situations like: comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs; generation time and quality measured side by side; choosing a default model; testing a new release against the incumbent.

How do I install Scenario Model Comparison in Claude Code?

Run `npx skills add scenario-labs/skills --skill scenario-model-comparison -a claude-code`. Or copy the skill folder (skills/scenario-model-comparison in scenario-labs/skills) into .claude/skills/scenario-model-comparison in your project. Claude Code loads it when a task matches its description.

How do I install Scenario Model Comparison in Codex?

Run `npx skills add scenario-labs/skills --skill scenario-model-comparison -a codex`. Or copy the skill folder (skills/scenario-model-comparison in scenario-labs/skills) into .agents/skills/scenario-model-comparison in your project. Codex loads it when a task matches its description.

Can I use Scenario Model Comparison in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scenario-labs/skills --skill scenario-model-comparison -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-model-comparison, .gemini/skills/scenario-model-comparison, .github/skills/scenario-model-comparison and .opencode/skills/scenario-model-comparison in your project.

What does Scenario Model Comparison need to run?

Going by SKILL.md and its folder, Scenario Model Comparison needs the command-line tools its instructions call (npx). Our summary lists: Node.js.

Does Scenario Model Comparison access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Scenario Model Comparison safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scenario Model Comparison use?

Scenario Model Comparison is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scenario Model Comparison use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scenario Model Comparison?

Skills that share tags, products or a category with Scenario Model Comparison: Ab Test Analyzer (irinabuht12-oss/marketing-skills, 4k stars), Sealeap Dijiang Amazon Review Mined Image Iteration (xjli360/sealeap-amazon-skills, 247 stars), Sealeap Taotie Amazon Image Script And Ab Testing (xjli360/sealeap-amazon-skills, 247 stars) and Continue Project (glifxyz/glif-mcp-server, 212 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scenario Model Comparison?

scenario-labs (a GitHub organization) maintains it in scenario-labs/skills, which has 931 GitHub stars. The repository holds 143 skills in this directory. The repository was last updated on October 8, 2026.

Source: scenario-labs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.