Agent skill

Argent QA Flows

by bbplayer-app in bbplayer-app/BBPlayer

Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria.

MITAuto-check passedTesting & QA

Install Argent QA Flows

skills CLI
$ npx skills add bbplayer-app/BBPlayer --skill argent-qa-flows -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bbplayer-app/BBPlayer argent-qa-flows --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bbplayer-app/BBPlayer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/argent-qa-flows .claude/skills/argent-qa-flows && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
argent-qa-flows
GitHub stars
1.1k
Token cost
~3.2k tokens
SKILL.md length
1,660 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
MIT

At a glance

Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria.

  • Works in 6 steps: Define the test contract → Record the scenario → Make evidence discriminating → …
  • The user asks to generate
  • SKILL.md covers Definition of done, 1. Define the test contract, 2. Record the scenario and 3. Make evidence discriminating, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Argent QA Flows is an agent skill from bbplayer-app/BBPlayer. Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria. Use when the user asks to generate or preserve an automated regression scenario, with deterministic setup, stable targets, executable structural or visual evidence, and two consecutive full passes. For one-off UI checks or replayable paths without acceptance criteria, use argent-test-ui-flow or argent-create-flow. Supports iOS, Android, Chromium, and Vega (Fire TV), where recorded tv-remote steps replace…

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering User stories, Test generation and End-to-end testing. It works with Android and iOS. The repository describes itself as: 一款简约、好用的 BiliBili 音乐播放器。 The licence is MIT.

When your agent uses it

  • The user asks to generate
  • Preserve an automated regression scenario
  • With deterministic setup
  • Executable structural

Example prompts

  • “/argent-qa-flows”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Define the test contract
  2. Record the scenario
  3. Make evidence discriminating
  4. Finish and audit
  5. Prove two consecutive passes
  6. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 1e1ae9f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Argent QA Flows loads about 3.2k tokens when it runs. Until then it costs about 162 tokens; SKILL.md has 1,660 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~162
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from bbplayer-app/BBPlayer at commit 1e1ae9f, republished under its MIT licence (© bbplayer-app). 1,660 words, ~3,245 tokens.

Download SKILL.mdSave it as .claude/skills/argent-qa-flows/SKILL.md (or your agent's skills folder).
name
argent-qa-flows
description
Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria. Use when the user asks to generate or preserve an automated regression scenario, with deterministic setup, stable targets, executable structural or visual evidence, and two consecutive full passes. For one-off UI checks or replayable paths without acceptance criteria, use argent-test-ui-flow or argent-create-flow. Supports iOS, Android, Chromium, and Vega (Fire TV), where recorded tv-remote steps replace touch directives. Apple TV and Android TV are unsupported; use argent-tv-interact there and report the limitation.

Create a QA regression flow

Load argent-create-flow as the authoring engine. Follow its required references for recorder syntax, selectors, polish, platform exceptions, and repair. This skill adds the QA contract and completion gate.

Vega supports every item below: launch: { vega: ... }, await:/assert: selectors, snapshot:, and idle all run there. Only the touch directives are missing, because Vega is remote-driven. Navigate with recorded tool: tv-remote steps and type with tool: keyboard, which leaves item 5 with nothing to govern. A D-pad path is relative to where focus already is, so gate every move with item 4's identity check rather than assuming the cursor landed. Read argent-tv-interact for focus reading and remote navigation.

Apple TV and Android TV are out of scope. The runner does not reject touch directives there, so they fail at the gesture layer instead of with authoring guidance. Use argent-tv-interact and report the limitation.

Physical iPhones run QA flows, with three hardware limits: replay never auto-binds a phone, so pass its udid as device (CLI --device) and keep it connected; pinch/rotate steps fail there like the live tools, so drive the app's own zoom UI instead; the flow tree is the describe tree (same ids and roles), so a selector authored on a simulator can miss there. Read argent-ios-device-interact for the app-scoped contract before recording.

Definition of done

A QA flow is complete only when:

  1. The first step that is not echo: or script: is launch:. In-flow setup proves a deterministic data baseline. Repeated runs do not accumulate artifacts or require manual cleanup.
  2. The first walkthrough recorded every action and live structural check. Only the three documented polish insertions are unrecorded.
  3. Every requirement maps to a hard await:, assert:, or reviewed snapshot:. Echoes and screenshots are not verdicts. A negative check needs the same stable selector established as visible earlier.
  4. Every screen change has destination identity followed by idle readiness.
  5. Targets satisfy the stable-selector and coordinate-fallback rules. QA keeps coordinates only for genuinely unlabeled targets. Vacuous on Vega, which has no coordinate targets.
  6. The unchanged YAML passes twice with the same runner. Pass 1 starts with fresh mobile Argent services, and pass 2 follows immediately.

1. Define the test contract

Before touching the app, write a compact table. Restate it in the final report. Include:

  • App, platform, and named start state.
  • Ordered user actions.
  • One row for each expected outcome, persistence rule, or absence claim.
  • Stable executable evidence for each row.
  • Required data and side effects.

Use structural checks for semantic state, snapshots for pixels, and both for mixed requirements. One behavioral scenario becomes one qa-<area>-<behavior> flow.

Do not invent a material value or weaken ambiguity. Choose the strongest UI-verifiable reading and report it. Ask when the choice changes test meaning.

Make repeated runs deterministic:

  1. Inspect the required baseline without mutation.
  2. If the account is dirty, record a safe reset or seed flow. Alternatively, include safe normalization in setup.
  3. After setup navigation, echo the named baseline and hard-check it before the first scenario mutation. Use assert: or a destination await: that fully proves the baseline.
  4. Prefer to restore the baseline at the end.

A flow has two fixture mechanisms:

  • run: replays a separately recorded reset or seed flow.
  • script: runs requested local setup or cleanup. Record it with flow-add-script where it belongs in the walkthrough.

Ask before cleanup that creates or deletes meaningful user data outside the request.

Compact example

Ticket: select Dark in Settings. Verify Dark is selected, Light is absent, and the screen renders in dark mode.

Contract rowActionEvidenceState effect
Signed-in HomeLaunchawait: { visible: { id: home-screen } }, then await: { idle: true }Existing account
Open SettingsTap settings-tabawait: { visible: { id: settings-screen } }, then await: { idle: true }None
Prove Light selectedInspect Settingsassert: { visible: { id: theme-light-selected } }Fails if already Dark
Prove Dark selectedTap theme-dark-optionawait: { visible: { id: theme-dark-selected } }Theme becomes Dark
Prove Light absentInspect settled screenassert: { hidden: { id: theme-light-selected } }None
Verify dark renderingInspect settled screensnapshot: settings-darkNone
Restore baselineTap theme-light-optionawait: { visible: { id: theme-light-selected } }Next run starts clean

The initial Light check establishes the selector used by the later hidden check. The final restore makes pass 2 independent.

2. Record the scenario

Follow argent-create-flow's start order and live-authoring cycle. Record each structural contract check when its state appears.

A snapshot has no recorder form. Inspect its stable state during the walkthrough, then add the planned snapshot during polish. If direct recovery changes state, re-record the affected behavior. A recovered walkthrough is not proof.

3. Make evidence discriminating

  • State change: prove the new state and the old state's absence when both can otherwise match.
  • Cancel/persistence: cross the commit boundary. After cancel or save, leave, re-enter, then verify the stored state: unchanged after cancel or updated after save.
  • Absence: prove the containing screen and record the same stable selector as visible, then the action, then hidden. Do not add an unestablished hidden check only to strengthen a positive baseline. In a collection, viewport absence is not global absence. Use fixed seeded position, count, empty state, or other collection-wide evidence.
  • Overlays: use the create-flow obscured-target procedure.
  • Repeated controls: prefer an id. Otherwise use flow-only within with a stable container. Use text.in to prove rendered membership inside that container.
  • Dynamic content: assert controlled state or stable app chrome. Use anchored structure for unavoidable dynamic values and disclose the dependency.
  • Visual state: snapshot only a correct, settled, deterministic screen. Use full screen for global changes and cropOn for one component.

Never put acceptance evidence inside when:. Use when: only for optional setup that reconverges to the required path.

4. Finish and audit

Complete the create-flow polish and blocking audit. Then:

  1. Map every contract row to an executed action or hard check.
  2. Build a navigation table with one row per screen change, naming both the identity gate and the readiness gate. A row missing either is a blocking defect.
  3. Confirm setup and end state permit an immediate second run.
ActionDestinationIdentityReadiness
Tap settings-tabSettingssettings-screen visibleidle

The two are repaired differently. A missing identity check must be recorded live on the restored screen. A missing idle check is added in YAML, because await: { idle: true } has no recorder form and is one of argent-create-flow's three permitted polish insertions. Re-record any missing action or other structural check.

Show full SKILL.md (603 more words)Show less

5. Prove two consecutive passes

After the last edit and audit, set the streak to zero:

  1. Choose one runner for both passes. Use flow-execute locally or argent flow run <name> --platform <platform> for CI. Switching runners resets the streak.
  2. Seed, review, and freeze snapshot baselines. Baseline updates do not count as passes.
  3. Before mobile pass 1, recycle Argent services for this flow's device: two warm passes are correlated evidence, because a fixed timing margin can pass twice simply because environment speed did not change. Scope stop-all-simulator-servers to devices: [<device>]. Never omit the scope — a bare call is the machine-wide sweep, and step 7 restarts this proof often enough to reap every other agent's devices repeatedly. Use the MCP call for flow-execute, or argent run stop-all-simulator-servers --devices <device> from the standalone runner's install. The reset must not change app or account data. For Chromium, let the runner boot the declared app and omit device. Vega owns no recyclable Argent services, so the teardown is a no-op there and both passes are warm.
  4. Run from the flow's launch and setup without baseline-update mode. Count a pass only when ok: true and every acceptance check executed. A false when: can skip optional setup only. An errored step does not advance the streak, and the count mixes two kinds — read each reason. One that could not run (an unreadable tree under idle, an unresolvable run: target) is environment: fix it and rerun. A failed launch: also scores errored, and it is a verdict about the app — an app that no longer installs or starts is the regression this test exists to catch, so report it instead of rerunning.
  5. Resolve every passing-step warning before completion. Also resolve recorded-wait warnings from flow-finish-recording. Follow Live waits and checks. For runner warnings, await: { idle: true } raises six different warnings, so read which one it is first. Two say the screen was moving. One says the wait ran out mid-hold and needs a larger timeout:. One says the tree stayed empty. One — settled on the UI tree alone — says the hierarchy did hold still and only the screenshot pairs were missing, so inspect the capture path rather than the app's rendering. One says the step ended with no evidence either way. Inspect the screen, disclose the cause, and verify that surrounding acceptance checks use stable elements rather than stillness. A selector-less gesture — a coordinate tap/long-press/swipe, or a pinch/rotate with no on: — warns in a different shape: a tree-source outage left it unsettled, so it dispatched blind and the green says only that the gesture was sent. Restore the tree source, usually by relaunching the app so the instrumentation loads, and rerun. Accepting that warning needs an app that serves no tree, which cannot satisfy this contract anyway.
  6. Run the same YAML again immediately with the same runner. Do not manually reset app or account data.
  7. Reset the streak after any failure, edit, re-recording, baseline update, or state-changing manual recovery. Repair through argent-create-flow, audit again, and restart with fresh services.

Finish only when the streak reaches two. If the intended runner is unavailable, report proof as blocked. If product behavior fails, keep the strong check and report the regression. Never weaken it to obtain green output.

Record the runner and fresh-service setup used.

6. Report

Report:

  • Flow name, path, platform, and standalone command.
  • Contract rows mapped to actions and checks.
  • Navigation table.
  • Baseline setup, end-state restoration, and accepted data dependencies.
  • Both pass results, runner, fresh-service setup, and resolved warnings.
  • Snapshot scope, reviewed baseline status, and mismatch tolerance.
  • Coordinate or raw-gesture exceptions.
  • Remaining manual judgment or blocker.

© bbplayer-app, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/argent-qa-flows of bbplayer-app/BBPlayer.

Open the folder on GitHubat commit 1e1ae9f

Compare with similar skills

Argent QA Flows next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Argent QA Flows compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Argent QA Flows this skillbbplayer-app/BBPlayer1.1k—~3.2kAutomated safety check: PassMIT
Kane CLI Browser TestingLambdaTest/kane-cli247—~8.4kAutomated safety check: PassApache-2.0
MAUI UI Test Writerdotnet/maui23k—~3kAutomated safety check: PassMIT
Test Scenariosphuryn/pm-skills27k—~866Automated safety check: PassMIT
Engine E2Ewix/react-native-navigation13k—~1.1kAutomated safety check: PassMIT
E2Egronxb/hot-updater1.8k—~1.6kAutomated safety check: PassCustom licence

Similar skills

  • Kane CLI Browser Testing

    LambdaTest/kane-cli

    Drives a real browser through the kane-cli tool and designs requirement-linked test suites from a PRD or a plain description, with mobile and cloud-grid runs.

    247 GitHub stars~8.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Official

    Writes UI tests that reproduce a GitHub issue in .NET MAUI and keeps iterating until the tests actually fail, proving they catch the bug.

    23k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Scenarios

    phuryn/pm-skills

    Create comprehensive test scenarios from user stories with test objectives, starting conditions, user roles, step-by-step actions, and expected outcomes.

    27k GitHub stars~866 tokensUpdated 22 days ago
    Testing & QAAuto-check passed
  • Engine E2E

    wix/react-native-navigation

    Official

    Run Wix Engine (mobile-apps-engine) iOS E2E tests locally to validate RNN changes.

    13k GitHub stars~1.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • E2E

    gronxb/hot-updater

    Run end-to-end OTA verification for examples/v0.85.0 with agent-device.

    1.8k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • E2E Current PR

    gronxb/hot-updater

    Generate and execute PR-aware OTA E2E scenarios for examples/v0.85.0 by diffing the checked-out branch against its PR base branch or default branch, inferring the affected runtime, rollout, and…

    1.8k GitHub stars~1k tokensUpdated today
    Testing & QAAuto-check passed

More from bbplayer-app/BBPlayer

All 19 skills in this repo
  • Expo Module

    bbplayer-app/BBPlayer

    Guide for creating and writing Expo native modules and views using the Expo Modules API (Swift, Kotlin, TypeScript).

    1.1k GitHub starsUsed in 3 repos~1.4k tokens
    Auto-check passed
  • Argent Metro Debugger

    bbplayer-app/BBPlayer

    Debug a JS runtime via CDP using argent debugger tools. An agent skill from bbplayer-app/BBPlayer.

    1.1k GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Optimizes a React Native app by profiling first to find real bottlenecks, then sweeping for mechanical issues.

    1.1k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Argent Lens

    bbplayer-app/BBPlayer

    Propose multiple visual design variants for on-screen elements and let the human pick in the Argent Lens window.

    1.1k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Argent Screenshot Diff

    bbplayer-app/BBPlayer

    Compare saved or live app screenshots with the argent screenshot-diff tool.

    1.1k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Argent Device Interact

    bbplayer-app/BBPlayer

    Interact with an iOS simulator, Android emulator, or Chromium (CDP) app using argent MCP tools.

    1.1k GitHub stars~6k tokensUpdated yesterday
    Auto-check: warnings

Works with

Questions about Argent QA Flows

What does Argent QA Flows do?

Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria. Argent QA Flows is an agent skill from bbplayer-app/BBPlayer. Create repeatable QA regression E2E tests as Argent flows from test cases, tickets, or acceptance criteria.

When should I use Argent QA Flows?

Argent QA Flows fits situations like: the user asks to generate; preserve an automated regression scenario; with deterministic setup; executable structural.

How do I install Argent QA Flows in Claude Code?

Run `npx skills add bbplayer-app/BBPlayer --skill argent-qa-flows -a claude-code`. Or copy the skill folder (.agents/skills/argent-qa-flows in bbplayer-app/BBPlayer) into .claude/skills/argent-qa-flows in your project. Claude Code loads it when a task matches its description.

How do I install Argent QA Flows in Codex?

Run `npx skills add bbplayer-app/BBPlayer --skill argent-qa-flows -a codex`. Or copy the skill folder (.agents/skills/argent-qa-flows in bbplayer-app/BBPlayer) into .agents/skills/argent-qa-flows in your project. Codex loads it when a task matches its description.

Can I use Argent QA Flows in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bbplayer-app/BBPlayer --skill argent-qa-flows -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/argent-qa-flows, .gemini/skills/argent-qa-flows, .github/skills/argent-qa-flows and .opencode/skills/argent-qa-flows in your project.

What does Argent QA Flows need to run?

SKILL.md names no scripts, command-line tools or credentials: Argent QA Flows is instructions for the agent only.

Does Argent QA Flows access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Argent QA Flows safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Argent QA Flows use?

Argent QA Flows is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Argent QA Flows use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Argent QA Flows?

Skills that share tags, products or a category with Argent QA Flows: Kane CLI Browser Testing (LambdaTest/kane-cli, 247 stars), MAUI UI Test Writer (dotnet/maui, 23k stars), Test Scenarios (phuryn/pm-skills, 27k stars) and Engine E2E (wix/react-native-navigation, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Argent QA Flows?

bbplayer-app (a GitHub organization) maintains it in bbplayer-app/BBPlayer, which has 1,126 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 6, 2026.

Source: bbplayer-app/BBPlayer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.