Ab Test Analyzer
irinabuht12-oss/marketing-skills
Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.
A skill your agent uses when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a…
$ npx skills add scenario-labs/skills --skill scenario-model-comparison -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install scenario-labs/skills scenario-model-comparison --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario-model-comparison .claude/skills/scenario-model-comparison && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scenario-model-comparison" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-model-comparison into .claude/skills/scenario-model-comparison/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-model-comparison", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/scenario-labs/skills/tree/main/skills/scenario-model-comparisonType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add scenario-labs/skills --skill scenario-model-comparison -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install scenario-labs/skills scenario-model-comparison --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/scenario-model-comparison .agents/skills/scenario-model-comparison && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scenario-model-comparison" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-model-comparison into .agents/skills/scenario-model-comparison/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-model-comparison", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scenario-labs/skills --skill scenario-model-comparison -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install scenario-labs/skills scenario-model-comparison --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/scenario-model-comparison .cursor/skills/scenario-model-comparison && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scenario-model-comparison" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-model-comparison into .cursor/skills/scenario-model-comparison/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-model-comparison", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/scenario-labs/skills.git --path skills/scenario-model-comparison--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add scenario-labs/skills --skill scenario-model-comparison -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install scenario-labs/skills scenario-model-comparison --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/scenario-model-comparison .gemini/skills/scenario-model-comparison && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scenario-model-comparison" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-model-comparison into .gemini/skills/scenario-model-comparison/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-model-comparison", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install scenario-labs/skills scenario-model-comparisonInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add scenario-labs/skills --skill scenario-model-comparison -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/scenario-model-comparison .github/skills/scenario-model-comparison && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scenario-model-comparison" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-model-comparison into .github/skills/scenario-model-comparison/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-model-comparison", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scenario-labs/skills --skill scenario-model-comparison -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install scenario-labs/skills scenario-model-comparison --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/scenario-model-comparison .opencode/skills/scenario-model-comparison && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scenario-model-comparison" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-model-comparison into .opencode/skills/scenario-model-comparison/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-model-comparison", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scenario-model-comparisonA skill your agent uses when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a…
Scenario Model Comparison is an agent skill from scenario-labs/skills. Use when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a default model, testing a new release against the incumbent, or finding which model handles a style, edit, or reference image best. Keywords: model comparison, benchmark, bake-off, A/B test, contact sheet, cost per asset, latency.
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Marketing & SEO, covering A/B testing. It works with Model Context Protocol. The repository describes itself as: Get production-ready images, video, audio, and 3D from any AI agent: skills that pick the right model, price before spending, and keep characters and brands consistent through… The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 91caa01. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scenario Model Comparison loads about 2.6k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 1,454 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from scenario-labs/skills at commit 91caa01, republished under its MIT licence (© scenario-labs). 1,454 words, ~2,635 tokens.
.claude/skills/scenario-model-comparison/SKILL.md (or your agent's skills folder).A comparison is one brief run unchanged across several models, then judged on criteria written down before the first paid call. Every candidate has its own contract, so the comparison lives in the normalization: what is held fixed, what each schema forces to differ, and where each number came from. recommend shortlists, dry_run prices, the jobs_wait rows carry the billed cost, the grid tool puts the results side by side, and a table delivers. Connection and the core loop: see the scenario skill; a rubric for judging output against a brief: scenario-refine-loop. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.
| Step | Do |
|---|---|
| 1. Frame | One brief, its inputs (prompt text, reference asset ids), three to five pass/fail criteria, all written before running |
| 2. Shortlist | recommend with the brief as prompt (limit up to 10); search for candidates the user named; keep three to five |
| 3. Normalize | model_schema_get each candidate: the shared fields, the caps that differ, the flags that rewrite prompts |
| 4. Price | model_run with dry_run: true (top-level, beside model_id, never inside parameters) per candidate and utility model; apply the Budget gate |
| 5. Run | model_run with top-level wait: false per candidate, launched back to back, then one jobs_wait over all the job ids |
| 6. Measure | cuCost from the jobs_wait rows; createdAt to updatedAt from job_get for the seconds |
| 7. Judge | The sheet from model_scenario-grid-maker, then each asset at full size, scored against the criteria |
| 8. Deliver | One table row per candidate; the assets filed in one collection and tagged by model id |
width and height from asset_get beside the result, since a larger output costs more and reads sharper. Prompt caps differ (max_length): an overrun is a 400, never an implicit trim. Preserve an exact user brief: replace a candidate whose cap is too small, or ask before changing the brief for every candidate.seed unset, or set it only for repeat runs of the same candidate.runs_as and run_with, see scenario), and its cost and time are the base run's.Price generations and required grid, extraction, or analysis steps. For a hard total cap, count each launched job once, including in-flight commitments. Before the first and every later paid call, require: committed spend + the next call's quote + verified reservations for remaining required work <= the cap. A substitute-input quote cannot verify a reservation unless documented pricing proves it bounds the real payload. If the sheet needs outputs that do not yet exist and has no verified bound, do not launch the generations: propose staged spending or a revised plan and wait for the user's agreement. Do not silently drop requested candidates or artifacts to fit. usage reports consumed CU, not remaining credits; it cannot certify a balance for this budget gate.
recommend reports modeled cost and latency, right for the shortlist and wrong as a result: dry_run prices the exact payload, and the jobs_wait row's cuCost is what was billed. Time comes from job_get, which returns createdAt and updatedAt; their difference on a finished job is the wall-clock latency including queue time, comparable across candidates launched within the same minute. Launch every candidate before waiting on any, within the team's concurrency ceiling (the 429 rows in scenario). Use the returned actionLimit on a concurrency 429; never infer the ceiling from this comparison's batch size. Re-call jobs_wait with pending_job_ids until every row is terminal. Record per candidate: model id, the exact parameters sent, cuCost, seconds, delivered dimensions, asset ids, and any normalization it forced (a size, a cap, a flag).
Video candidates are compared on contact sheets, never on a first frame: sweep each clip with model_scenario-video-to-image-seq (a fixed first-party id, Scenario's single deterministic frame extractor, so discovery would only re-derive it), wait for the extraction job, then sheet the frames per candidate; the extractor's frame-order and stride contract and the sheet's 100-image cap are in scenario-video-editing. Audio and 3D candidates are compared on the assets themselves through asset_display.
Pre-registered criteria are pass/fail statements about the output ("the label text is legible", "the scar is on the left cheek", "no extra fingers", "the background is plain"), so a result is scored, not admired, and a surprising winner cannot rewrite the test after the fact. Score the literal criterion: a related feature or partial match is not a pass; ambiguous evidence is unverified. Build the sheet with model_scenario-grid-maker (a fixed first-party id, Scenario's single deterministic grid tool, so discovery would only re-derive it): images in candidate order, wrapped as an array even for one, since the array order is the legend (the tool has no labels), columns equal to the candidate count so one row is one brief, cellRatio matching the outputs, padding for a gutter. Then asset_display each output at full size: a sheet hides fine text, edge halos and small anatomy. Score every criterion for every candidate, and put cost and seconds in the same table so the trade-off is read in one place. For a large set, asset_analyze (catalog, cost-bearing, and write-class: scenario_tools_search then scenario_tool_execute_write, see scenario) can score a batch: the criteria list verbatim as its instruction, the outputs as images, one verdict per criterion per image, and a criterion the image cannot settle at output resolution (small lettering, a fine edge) recorded as unverified, never as a pass; pre-register that instruction with the criteria.
recommend with that brief as prompt and limit: 5, then search with target="models", public=true for the two models the user named; keep three. Never hardcode the ids: catalogs differ per team.model_schema_get on each: prompt cap, size fields, sample-count field, any prompt-expansion flag.model_run with dry_run: true (top-level) on each, the same prompt, the count field at one, the closest 1024 square each allows, the expansion flag off; write the three prices down and apply the Budget gate, including the required sheet.model_run three times with wait: false, then one jobs_wait with the three job ids, re-called with pending_job_ids on a timeout; read cuCost off each row and createdAt and updatedAt off job_get for each job.model_scenario-grid-maker with images as the three asset ids in candidate order and columns: 3; asset_display the sheet, then each asset; score the four criteria.collection_create, then collection_add_assets), tagging each output with its model id through asset_add_tags.recommend's numbers as results: they are modeled. Bill from jobs_wait, time from job_get.© scenario-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/scenario-model-comparison of scenario-labs/skills.
Open the folder on GitHubat commit 91caa01
Scenario Model Comparison next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scenario Model Comparison this skillscenario-labs/skills | 931 | — | ~2.6k | Automated safety check: Pass | MIT | |
| Ab Test Analyzeririnabuht12-oss/marketing-skills | 4k | — | ~1.4k | Automated safety check: Pass | None | |
| Sealeap Dijiang Amazon Review Mined Image Iterationxjli360/sealeap-amazon-skills | 247 | — | ~748 | Automated safety check: Pass | MIT | |
| Sealeap Taotie Amazon Image Script And Ab Testingxjli360/sealeap-amazon-skills | 247 | — | ~818 | Automated safety check: Pass | MIT | |
| Continue Projectglifxyz/glif-mcp-server | 212 | — | ~734 | Automated safety check: Pass | MIT | |
| Ab Testingcoreyhaines31/marketingskills | 54k | 3 repos | ~3.1k | Automated safety check: Pass | MIT |
irinabuht12-oss/marketing-skills
Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.
xjli360/sealeap-amazon-skills
Mine competitor reviews with AI assistance to surface recurring pain points and desired improvements, benchmark the main image against top competitors, validate candidate images with quick…
xjli360/sealeap-amazon-skills
Build a listing image script from competitor review insights and traffic keywords—relevance, value proposition and call to action in the first frames, then scenes, details, care and sizing with…
glifxyz/glif-mcp-server
Find the user's saved Glif project, show its media, and add new work to it without starting a separate project.
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
Nexus-JPF/note-companion
When the user wants to set up, improve, or audit analytics tracking and measurement.
scenario-labs/skills
A skill your agent uses when drawing or animating with Grease Pencil in Blender 5.x from Python: 2D or 2.5D illustration, frame-by-frame animation, a cutout or part-based 2D character, strokes with…
scenario-labs/skills
A skill your agent uses when grooming hair or fur in Blender with hair curves, such as a character hairstyle, animal fur, procedural fur in geometry nodes, or hair cards and mesh hair for games.
scenario-labs/skills
A skill your agent uses when lighting, rendering or compositing in Blender: light a character, product or hero shot, interior at dusk or night, three-point or motivated lighting, sun and sky, HDRI…
scenario-labs/skills
A skill your agent uses when animating characters or scenes in Godot 4.7: AnimationPlayer clips and RESET, AnimationTree state machines and blend spaces built in code, Mixamo or glTF import, loop…
scenario-labs/skills
A skill your agent uses when adding or fixing sound in Godot 4.7: audio buses and effects, volume sliders, 'too many sounds', combat audio with hundreds of enemies, sounds clipping or distorting, 3D…
scenario-labs/skills
A skill your agent uses when a Godot 4 game goes online or co-op: host and join with ENet, WebSocket for a web build, RPCs (@rpc, rpcid, anypeer), MultiplayerSpawner and MultiplayerSynchronizer…
Works with
Categories
A skill your agent uses when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a…. Scenario Model Comparison is an agent skill from scenario-labs/skills. Use when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a default model, testing a new release against the incumbent, or finding which model handles a style, edit, or reference image best.
Scenario Model Comparison fits situations like: comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs; generation time and quality measured side by side; choosing a default model; testing a new release against the incumbent.
Run `npx skills add scenario-labs/skills --skill scenario-model-comparison -a claude-code`. Or copy the skill folder (skills/scenario-model-comparison in scenario-labs/skills) into .claude/skills/scenario-model-comparison in your project. Claude Code loads it when a task matches its description.
Run `npx skills add scenario-labs/skills --skill scenario-model-comparison -a codex`. Or copy the skill folder (skills/scenario-model-comparison in scenario-labs/skills) into .agents/skills/scenario-model-comparison in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scenario-labs/skills --skill scenario-model-comparison -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-model-comparison, .gemini/skills/scenario-model-comparison, .github/skills/scenario-model-comparison and .opencode/skills/scenario-model-comparison in your project.
Going by SKILL.md and its folder, Scenario Model Comparison needs the command-line tools its instructions call (npx). Our summary lists: Node.js.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Scenario Model Comparison is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Scenario Model Comparison: Ab Test Analyzer (irinabuht12-oss/marketing-skills, 4k stars), Sealeap Dijiang Amazon Review Mined Image Iteration (xjli360/sealeap-amazon-skills, 247 stars), Sealeap Taotie Amazon Image Script And Ab Testing (xjli360/sealeap-amazon-skills, 247 stars) and Continue Project (glifxyz/glif-mcp-server, 212 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
scenario-labs (a GitHub organization) maintains it in scenario-labs/skills, which has 931 GitHub stars. The repository holds 143 skills in this directory. The repository was last updated on October 8, 2026.
Source: scenario-labs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.