Agent skill

Scenario

by ffroliva in ffroliva/gflow-cli

Pre-implementation edge-case and scenario explorer for gflow-cli.

MITAuto-check passedTesting & QA

Install Scenario

skills CLI
$ npx skills add ffroliva/gflow-cli --skill scenario -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ffroliva/gflow-cli scenario --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ffroliva/gflow-cli.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario .claude/skills/scenario && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scenario
GitHub stars
264
Token cost
~3.1k tokens
SKILL.md length
1,315 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Pre-implementation edge-case and scenario explorer for gflow-cli.

  • Tasks that involve Browser testing
  • SKILL.md covers When to invoke, Invocation, The 13 dimensions and Artifact location, plus 2 more sections
  • Reaches github.com
  • Tasks that involve AI video generation

What it does

Scenario is an agent skill from ffroliva/gflow-cli. Pre-implementation edge-case and scenario explorer for gflow-cli. Decomposes a proposed feature or change across 13 structured dimensions specific to gflow-cli's failure surfaces (WAF/reCAPTCHA, Playwright selectors, auth token lifecycle, batch resume, data layer safety, cross-platform paths). Produces a severity-ranked scenario table to feed into the PLAN spec and the test matrix before EXECUTE begins.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Browser testing and AI video generation. It works with Playwright. The repository describes itself as: Drive Google Flow from the command line: Veo video and Imagen images, scripted, batched and pipeline-ready. Ships an MCP server so coding agents can drive it too, giving you and… The licence is MIT.

When your agent uses it

  • Tasks that involve Browser testing
  • Tasks that involve AI video generation

Example prompts

  • “/scenario”

What it can do on your machine

Read from SKILL.md and the folder at commit cb6d501. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scenario loads about 3.1k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 1,315 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ffroliva/gflow-cli at commit cb6d501, republished under its MIT licence (© ffroliva). 1,315 words, ~3,092 tokens.

Download SKILL.mdSave it as .claude/skills/scenario/SKILL.md (or your agent's skills folder).
name
scenario
description
Pre-implementation edge-case and scenario explorer for gflow-cli. Decomposes a proposed feature or change across 13 structured dimensions specific to gflow-cli's failure surfaces (WAF/reCAPTCHA, Playwright selectors, auth token lifecycle, batch resume, data layer safety, cross-platform paths). Produces a severity-ranked scenario table to feed into the PLAN spec and the test matrix before EXECUTE begins.

scenario — Edge Case & Scenario Explorer

Systematic pre-implementation scenario analysis. For a given feature or change, produces a severity-ranked table of test scenarios across 13 dimensions tuned to gflow-cli's known failure surfaces. Feed the output into PLAN.md tasks and tests/features/ BDD scenarios before entering EXECUTE mode.


When to invoke

Use before implementing any of:

  • A new generation path (T2V, I2V, R2V, batch manifest runner)
  • Transport changes (any new HTTP call against aisandbox-pa.googleapis.com)
  • Auth or session changes (new strategy, cookie extraction, SAPISIDHASH wiring)
  • Selector cascade changes (ONBOARDING_SELECTORS, NEW_PROJECT_SELECTORS, FRAME_SLOTS_STRUCT, IMAGE_MODEL_OPTION_SELECTORS)
  • Data layer changes (schema migration, new OperationRecorder callsite, redaction change)
  • New CLI subcommand, flag, or exit code

Skip for: pure doc changes, CHANGELOG/version bumps, scripts/ tooling with no production callpath.


Invocation

/gflow:scenario <feature or change description>

<feature or change description> is a brief summary of what you're about to implement. Examples:

  • "SAPISIDHASH auth header wired into _post_json for aisandbox-pa routes"
  • "gflow video batch manifest ledger (skip already-completed rows)"
  • "Image model picker converted from English has-text to structural anchor"
  • "CDP-attach transport as opt-in alongside ui_automation"

The 13 dimensions

For each dimension, enumerate scenarios that are non-obvious — do not list things that a basic happy-path test already covers. Focus on things that break in production but pass in unit tests.

D1 — Auth & session lifecycle

The SAPISID cookie expires. The user re-runs gflow auth login mid-batch. A profile is created but the Flow session was never verified. The session is valid for labs.google tRPC but not for aisandbox-pa. Two profiles are in use simultaneously (Chromium profile-lock).

D2 — WAF / reCAPTCHA scoring

A profile's WAF heat score is elevated from prior automation runs. reCAPTCHA Enterprise detects navigator.webdriver=true despite the --disable-blink-features stealth flag. The grecaptcha.execute() call times out or returns a challenge that requires human interaction. The same token is submitted twice (single-use token reuse). A batch run fires multiple rapid token mints within one session.

D3 — Selector cascade drift & Locale Invariance (Flow UI updates)

Google Flow ships a UI update that renames or restructures a selector anchor. FRAME_SLOTS_STRUCT matches zero elements (the PR #70 regression). The structural-first tier matches but selects the wrong element (e.g., two elements matching div[type='button'][aria-haspopup='dialog'] after Flow adds a new dialog button). A non-English locale is used (e.g. GFLOW_CLI_LOCALE=pt-BR or a Brazilian Google account) and single-language English text selectors fail because they lack Tier 1 structural anchors (a[href*='changelog'], button:has(i:text('close'))) or Tier 2 multi-locale text cascades. A new locale is used whose CMP dialog is not in _ONBOARDING_TEXT_SELECTORS. The --lang=en-US Chromium arg is removed before IMAGE_MODEL_OPTION_SELECTORS is converted to structural anchors.

D4 — Batch manifest & resume

A 50-row TSV manifest dies at row 23 (auth expiry). On resume, rows 1–22 are re-submitted and double-billed. A row has an empty start_image field (T2V) while the column parser expected a path. The output path already exists on disk from a prior run (overwrite vs skip decision). Two gflow video batch invocations against the same profile run concurrently (Chromium profile-lock).

D5 — Concurrency & Page pool

GFLOW_CLI_CONCURRENCY=16 fans out 16 simultaneous page.evaluate() reCAPTCHA token mints — do they share site key state safely? A Page is returned to the pool while still in a modal dialog (state contamination for the next checkout). A QueueFull double-checkin race where two coroutines both attempt to return the same Page object. asyncio.gather propagates one failure but leaves the other coroutines in a half-completed state with no cleanup path.

D6 — Data layer (SQLite / DataStore)

DataStore migration runs on a schema that's one version ahead (newer-schema detection should raise DataStoreError exit 16, not silently corrupt). An OperationRecorder.on_started callback raises inside a generation loop — does the batch continue or crash? GFLOW_CLI_HISTORY_PROMPTS=redacted is set — does the new callsite respect the redaction gate? A Windows path with a drive letter is stored in the DB and retrieved on Linux (or vice versa). The DB is on a network filesystem and the write locks time out.

D7 — Error propagation & exit codes

A new exception class is added but not registered in EXIT_CODE_MAP. A FlowApiError subclass is raised inside a retry loop — does reraise=True propagate the right typed exception or does tenacity eat it? WireFormatError discovery payload logs the route body prefix — does it redact reCAPTCHA tokens and auth headers? An unhandled exception escapes run_with_handlers — is it SHA-256-hashed before logging (never raw message)?

D8 — Cross-platform paths

A Windows user has a Unicode character in their %LOCALAPPDATA% path. The output directory path contains spaces. platformdirs returns a different base path on macOS vs Linux vs Windows — does the feature hardcode any of these? PYTHONUTF8=1 is not set — are there cp1252 encoding traps on Windows for prompt text or file names?

D9 — Transport edge cases

The HTTP response body is valid JSON but has an unexpected top-level key (should emit WireFormatError with discovery payload, not crash). A video generation response omits operations[0].operation.name (the omni-flash NULL flow_operation_id issue). A status-poll response arrives before the listener is attached (_attach_status_response_listener race). A video download URL has a googleapis.com host not in the SSRF allowlist.

Show full SKILL.md (503 more words)Show less
D10 — Headless vs headed environment

The feature is run in a CI environment with no display server (Linux, no Xvfb). The Playwright context is launched headed (user has GFLOW_CLI_HEADLESS=false) and the window is closed manually mid-batch. gflow auth login --browser internal is used on a machine where Playwright bundled Chromium is flagged by Google's G12 block.

D11 — Input validation & boundary values

A prompt string is 4001 characters (Flow's apparent limit). A manifest TSV has zero rows (after the header). --seed receives a negative integer or a string. -n 5 is passed to gflow image t2i (Flow's UI cap is x4). An aspect ratio of 3:2 is passed (valid-looking but not in the whitelist). A --profile name contains path-separator characters.

D12 — Observability & structured log contract

A new structlog event is emitted — is its key name stable and documented in docs/ARCHITECTURE.md? A new error_raised path: does it carry the full RFC 9457 Problem Details shape (type, title, status, detail, instance, remediation_hint)? correlation_id is bound at the process boundary — does it propagate into the new code path or is it missing from the event?

D13 — MCP surface parity

gflow ships the same capability twice: as a CLI command and as an MCP tool. Every edge case above therefore has an MCP twin — enumerate it, because the paths are not the same code. The MCP tool can run direct or queued through worker/codec.py, so a param that works via the CLI can be accepted and silently dropped on the queued path. Ask: what does this feature look like invoked as an MCP tool? Which errors reach the agent as a Problem Details envelope rather than an exit code? Does the tool's docstring still describe the behaviour truthfully? The six mirror axes are enumerated once, in skills/check/SKILL.md step 1b — do not restate them here, scenario's job is only to name the MCP cases that need scenarios.


Artifact location

Write the analysis to docs/superpowers/plans/<YYYY-MM-DD>-<feature-slug>/SCENARIO.md (same directory the subsequent /gflow:plan writes its PLAN.md into) and commit it with the plan on the feature branch. Established by the image-upscale (#171) and locale-selectors (#170) cycles — do not leave the analysis only in conversation.

Output format

# Scenario: <feature short title>

## Coverage map
<Which of the 13 dimensions are relevant for this feature? List active dimensions and why skipped dimensions are skipped.>

## Scenario table

| # | Dimension | Scenario | Severity | Expected behaviour | Test category |
|---|---|---|---|---|---|
| 1 | D2 WAF/reCAPTCHA | WAF score elevated on profile from prior run | Critical | WafRejectionError raised with remediation hint "let score decay or use a different profile"; exit code 403-mapped | E2E live (`@pytest.mark.live`) |
| … | … | … | … | … | … |

Severity: **Critical** (data loss / billed twice / unrecoverable) · **High** (feature broken, workaround exists) · **Medium** (degraded UX, explicit error) · **Low** (cosmetic or edge-only)

Test category: **Unit** (no I/O) · **Integration** (mocked HTTP/Playwright) · **BDD** (Gherkin feature file) · **E2E smoke** (`@pytest.mark.smoke`) · **E2E live** (`@pytest.mark.live`, opt-in)

## Must-cover before merge (Critical + High)
1. …

## Deferred (Medium + Low — log as issues, not blockers)
1. …

## Suggested BDD scenarios
```gherkin
@e2e @e2e_auth
Feature: <feature name>
  Scenario: <scenario title>
    Given …
    When …
    Then …

Tag by surface — the tag is the binding, not decoration. pytest-bdd turns each Gherkin tag into a pytest marker, and addopts' -m 'not e2e …' filters on exactly that. So the tag decides where the scenario runs and who must run it:

The scenario can only happen…Feature-level tagsBound from
in a real browser / against real Flow@e2e + one cost tier (@e2e_auth, @e2e_image, @e2e_video, …)tests/e2e/test_<slug>_bdd.py
in our own code (parsing, routing, exit codes, redaction)nonetests/features/test_<slug>_steps.py

One feature file is bound by exactly one module — bound twice, its scenarios run twice. tests/features/test_e2e_binding_guard.py enforces this offline in four directions (orphan, missing tier, untagged live binding, double binding), so an @e2e scenario nobody wrote a test for fails normal CI. Mechanics: docs/E2E_TESTING.md § BDD-bound e2e.

Known-issues cross-reference

<Any scenario that maps to an existing KNOWN_ISSUES entry — link and note whether the proposed implementation resolves, mitigates, or is blocked by it.>


---

## Integration & Pipeline Continuation (Next Step Handoff)

1. Run `/gflow:predict` first to validate the approach (GO/CAUTION/STOP).
2. Run `/gflow:scenario` to enumerate edge cases and build the test matrix.
3. Use the "Must-cover before merge" list as the acceptance criteria for the PLAN.md task.
4. Add BDD scenarios **before** coding (TDD is non-negotiable per AGENTS.md), each
   tagged by its surface per the table above — a browser-only scenario is bound from
   `tests/e2e/`, and a mocked stand-in does not discharge it
   ([`skills/issue-resolve`](../issue-resolve/SKILL.md) § The Bug Lane, step 5).
5. **Next step:** Proactively announce: **"BDD Scenarios generated. Next step: Phase 4 Implementation Plan (`/gflow:plan <feature>`)."**

---

## Provenance

Adapted from `vc-scenario` in [vibecode-pro-max-kit](https://github.com/withkynam/vibecode-pro-max-kit) (assessment 2026-05-28).
13 dimensions re-scoped to gflow-cli's specific failure surfaces: Google WAF/reCAPTCHA, Playwright selector drift, auth token lifecycle, batch resume idempotency, SQLite data layer, RFC 9457 error propagation, cross-platform Windows/macOS/Linux paths.

© ffroliva, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scenario of ffroliva/gflow-cli.

Open the folder on GitHubat commit cb6d501

Compare with similar skills

Scenario next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scenario compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scenario this skillffroliva/gflow-cli264—~3.1kAutomated safety check: PassMIT
UI Demosanity-io/ui1764 repos~3.8kAutomated safety check: PassMIT
Demo Videohmislk/hmis236—~3.9kAutomated safety check: PassGPL-3.0
Record E2E Giflablup/backend.ai-webui133—~907Automated safety check: NotesLGPL-3.0
Render 3D Product Showcasegooseworks-ai/goose-skills1.2k—~1.2kAutomated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0

Similar skills

  • UI Demo

    sanity-io/ui

    Official

    Record polished UI demo videos using Playwright. An agent skill from sanity-io/ui.

    176 GitHub starsUsed in 4 repos~3.8k tokens
    Testing & QAAuto-check passed
  • Demo Video

    hmislk/hmis

    A skill your agent uses when asked to make a demo, training, how-to or tutorial video with sound or voice-over showing an HMIS function or configuration (e.g.

    236 GitHub stars~3.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Record E2E Gif

    lablup/backend.ai-webui

    Record Playwright e2e tests as one GIF per test case (video → ffmpeg palette GIF) and return a markdown table for a PR description.

    133 GitHub stars~907 tokensUpdated today
    Testing & QAAuto-check: notes
  • Render 3D Product Showcase

    gooseworks-ai/goose-skills

    Assemble a premium 3D product-showcase ad from a config — four beat clips (an orbiting hero rotation, a macro push-in, a physics reveal, a typographic close) normalized to the chosen backdrop…

    1.2k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Playwright CLI

    sanity-io/sanity

    Official

    Automates browser interactions for web testing, form filling, screenshots, and data extraction.

    6.4k GitHub starsUsed in 18 repos~1.9k tokens
    Testing & QAAuto-check passed

More from ffroliva/gflow-cli

All 17 skills in this repo
  • Gflow CLI

    ffroliva/gflow-cli

    A skill your agent uses when the user wants to drive Google Flow (Veo image-to-video, Veo text-to-video, Imagen / Nano Banana image generation) from the terminal or a script — including…

    264 GitHub stars~4.8k tokensUpdated yesterday
    Auto-check: notes
  • Issue Assessment

    ffroliva/gflow-cli

    A skill your agent uses when triaging a GitHub issue for gflow-cli — a reporter's bug claim, a freshly-filed issue, or deciding whether and how to act on one.

    264 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Issue Resolve

    ffroliva/gflow-cli

    A skill your agent uses when an assessed gflow-cli issue (verdict CONFIRMED-BUG or LIKELY-BUG) has localized, verifiable scope and should be driven to a fix.

    264 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Live Verify

    ffroliva/gflow-cli

    Two-part gate for gflow-cli feature/fix work. An agent skill from ffroliva/gflow-cli.

    264 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Video Production

    ffroliva/gflow-cli

    A skill your agent uses when the user wants a finished video out of gflow rather than a single clip — a scripted scene, a talking-head or dialogue piece, an explainer, a product montage, a story…

    264 GitHub stars~7k tokensUpdated yesterday
    Auto-check passed
  • Check

    ffroliva/gflow-cli

    Auto-fix lint and formatting, then report types and tests. An agent skill from ffroliva/gflow-cli.

    264 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Scenario

What does Scenario do?

Pre-implementation edge-case and scenario explorer for gflow-cli. Scenario is an agent skill from ffroliva/gflow-cli. Pre-implementation edge-case and scenario explorer for gflow-cli.

When should I use Scenario?

Scenario fits situations like: tasks that involve Browser testing; tasks that involve AI video generation.

How do I install Scenario in Claude Code?

Run `npx skills add ffroliva/gflow-cli --skill scenario -a claude-code`. Or copy the skill folder (skills/scenario in ffroliva/gflow-cli) into .claude/skills/scenario in your project. Claude Code loads it when a task matches its description.

How do I install Scenario in Codex?

Run `npx skills add ffroliva/gflow-cli --skill scenario -a codex`. Or copy the skill folder (skills/scenario in ffroliva/gflow-cli) into .agents/skills/scenario in your project. Codex loads it when a task matches its description.

Can I use Scenario in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ffroliva/gflow-cli --skill scenario -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario, .gemini/skills/scenario, .github/skills/scenario and .opencode/skills/scenario in your project.

What does Scenario need to run?

SKILL.md names no scripts, command-line tools or credentials: Scenario is instructions for the agent only.

Does Scenario access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Scenario safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scenario use?

Scenario is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scenario use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scenario?

Skills that share tags, products or a category with Scenario: UI Demo (sanity-io/ui, 176 stars), Demo Video (hmislk/hmis, 236 stars), Record E2E Gif (lablup/backend.ai-webui, 133 stars) and Render 3D Product Showcase (gooseworks-ai/goose-skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scenario?

ffroliva (a GitHub user) maintains it in ffroliva/gflow-cli, which has 264 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 7, 2026.

Source: ffroliva/gflow-cli on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.