Agent skill

Expect

by yonatangross in yonatangross/orchestkit

Diff-aware AI browser testing — reads the git diff, maps changes to affected pages via the route map, generates a targeted test plan, and executes it via agent-browser (Rust daemon + CDP…

MITAuto-check: notesTesting & QA

Install Expect

skills CLI
$ npx skills add yonatangross/orchestkit --skill expect -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit expect --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/expect .claude/skills/expect && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
expect
GitHub stars
289
Token cost
~4.8k tokens
SKILL.md length
1,323 words
Files
37 (incl. scripts, references)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Diff-aware AI browser testing — reads the git diff, maps changes to affected pages via the route map, generates a targeted test plan, and executes it via agent-browser (Rust daemon + CDP…

  • Works in 7 steps: MCP Probe + Prerequisite Check → Fingerprint Check → Diff Scan → …
  • Testing UI changes
  • SKILL.md covers Argument Resolution, STEP 0: MCP Probe +…, CRITICAL: Task Management and Pipeline Overview, plus 14 more sections
  • Calls vault, git and npx

What it does

Expect is an agent skill from yonatangross/orchestkit. Diff-aware AI browser testing — reads the git diff, maps changes to affected pages via the route map, generates a targeted test plan, and executes it via agent-browser (Rust daemon + CDP, ARIA-tree-first) with pass/fail reporting. Use when testing UI changes, verifying PRs before merge, or running regression checks on changed components.

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 38 other files, including scripts and reference files (for example `flows/crud.md`, `flows/login.md` and `flows/navigation.md`). Compatibility notes: Claude Code 2.1.277+. Requires agent-browser = 0.31.1 (Rust-native, no Playwright; 0.27.1 documented broken on prod pages).

It sits in Testing & QA, covering Test generation, Browser automation and Browser testing. It works with Git and Rust. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • Testing UI changes
  • Verifying PRs before merge
  • Running regression checks on changed components

Example prompts

  • “/expect”

Requirements

  • Python 3
  • Node.js
  • Compatibility (from SKILL.md): Claude Code 2.1.277+. Requires agent-browser >= 0.31.1 (Rust-native, no Playwright; 0.27.1 documented broken on prod pages).
  • Pre-approved tools (allowed-tools): AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, ToolSearch, WebFetch, Monitor, PushNotification, mcp__memory__search_nodes

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. MCP Probe + Prerequisite Check
  2. Fingerprint Check
  3. Diff Scan
  4. Route Map
  5. Test Plan Generation
  6. Execution
  7. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 0ef71d2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • AskUserQuestion
    • Bash
    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Agent
    • TaskCreate
    • TaskUpdate

    …and 6 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • vault
    • git
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+. Requires agent-browser >= 0.31.1 (Rust-native, no Playwright; 0.27.1 documented broken on prod pages).

    From compatibility in the SKILL.md frontmatter.

Context cost

Expect loads about 4.8k tokens when it runs, and up to ~18k if it reads all its reference files. Until then it costs about 87 tokens; SKILL.md has 1,323 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~18k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, ToolS

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit 0ef71d2, republished under its MIT licence (© yonatangross). 1,323 words, ~4,824 tokens.

Download SKILL.mdSave it as .claude/skills/expect/SKILL.md (or your agent's skills folder). This skill also uses 36 other files; get the full folder from GitHub.
name
expect
description
Diff-aware AI browser testing — reads the git diff, maps changes to affected pages via the route map, generates a targeted test plan, and executes it via agent-browser (Rust daemon + CDP, ARIA-tree-first) with pass/fail reporting. Use when testing UI changes, verifying PRs before merge, or running regression checks on changed components.
allowed-tools
AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, ToolSearch, WebFetch, Monitor, PushNotification, mcp__memory__search_nodes
compatibility
Claude Code 2.1.277+. Requires agent-browser >= 0.31.1 (Rust-native, no Playwright; 0.27.1 documented broken on prod pages).
license
MIT
argument-hint
[-m <instruction>] [--target unstaged|branch|commit] [--flow <slug>] [-y]
context
fork
background
false
user-invocable
true
skills
testing-e2e, chain-patterns, memory
effort
high
model
sonnet
metadata.category
testing
metadata.milestone
M99

Expect — Diff-Aware AI Browser Testing

Analyze git changes, generate targeted test plans, and execute them via AI-driven browser automation.

Note: If disableSkillShellExecution is enabled (CC 2.1.91), the agent-browser install check won't run. Verify it's installed: npx agent-browser --version.

bash
expect                              # Auto-detect changes, test affected pages
expect -m "test the checkout flow"  # Specific instruction
expect --flow login                 # Replay a saved test flow
expect --target branch              # Test all changes on current branch vs main
expect -y                           # Skip plan review, run immediately

Core principle: Only test what changed. Git diff drives scope — no wasted cycles on unaffected pages.

Argument Resolution

python
ARGS = "[-m <instruction>] [--target unstaged|branch|commit] [--flow <slug>] [-y]"

# Parse from full argument string
import re
raw = ""  # Full argument string from CC

INSTRUCTION = None
TARGET = "unstaged"  # Default: test unstaged changes
FLOW = None
SKIP_REVIEW = False

# Extract -m "instruction"
m_match = re.search(r'-m\s+["\']([^"\']+)["\']|-m\s+(\S+)', raw)
if m_match:
    INSTRUCTION = m_match.group(1) or m_match.group(2)

# Extract --target
t_match = re.search(r'--target\s+(unstaged|branch|commit)', raw)
if t_match:
    TARGET = t_match.group(1)

# Extract --flow
f_match = re.search(r'--flow\s+(\S+)', raw)
if f_match:
    FLOW = f_match.group(1)

# Extract -y
if '-y' in raw.split():
    SKIP_REVIEW = True

STEP 0: MCP Probe + Prerequisite Check

python
# memory is alwaysLoad in .mcp.json (CC 2.1.121+, #1541) — probe below kept as fallback for older CC:
ToolSearch(query="select:mcp__memory__search_nodes")

# Verify agent-browser is available (Rust-native, no Playwright)
Bash("command -v agent-browser || npx agent-browser --version")
# If missing: "Install agent-browser: npm i -g agent-browser"

# Load agent-browser's version-matched workflow guide (ships with the CLI).
# NOT `skills get agent-browser` — that resolves but returns only the thin
# top-level router; `core --full` is the actual 2,800+ line command reference.
Bash("agent-browser skills get core --full")

CRITICAL: Task Management

python
# 1. Create main task IMMEDIATELY
TaskCreate(
  subject="Expect: test changed code",
  description="Diff-aware browser testing pipeline",
  activeForm="Running diff-aware browser tests"
)

# 2. Create subtasks for each pipeline phase
TaskCreate(subject="Check fingerprint (skip if unchanged)", activeForm="Checking fingerprint")  # id=2
TaskCreate(subject="Scan git diff and classify changes", activeForm="Scanning diff")            # id=3
TaskCreate(subject="Map changes to routes/URLs", activeForm="Mapping routes")                   # id=4
TaskCreate(subject="Generate AI test plan", activeForm="Generating test plan")                   # id=5
TaskCreate(subject="Execute tests via agent-browser", activeForm="Executing browser tests")     # id=6
TaskCreate(subject="Compile test report", activeForm="Compiling report")                        # id=7

# 3. Set dependencies for sequential phases
TaskUpdate(taskId="3", addBlockedBy=["2"])  # Diff scan needs fingerprint check
TaskUpdate(taskId="4", addBlockedBy=["3"])  # Route map needs diff results
TaskUpdate(taskId="5", addBlockedBy=["4"])  # Test plan needs route map
TaskUpdate(taskId="6", addBlockedBy=["5"])  # Execution needs test plan
TaskUpdate(taskId="7", addBlockedBy=["6"])  # Report needs execution results

# 4. Update status as you progress
TaskUpdate(taskId="2", status="in_progress")  # When starting
TaskUpdate(taskId="2", status="completed")    # When done — repeat for each subtask

Pipeline Overview

Git Diff → Route Map → Fingerprint Check → Test Plan → Execute → Report
PhaseWhatOutputReference
1. FingerprintSHA-256 hash of changed filesSkip if unchanged since last runreferences/fingerprint.md
2. Diff ScanParse git diff, classify changesChangesFor data (files, components, routes)references/diff-scanner.md
3. Route MapMap changed files to affected pages/URLsScoped page listreferences/route-map.md
4. Test PlanGenerate AI test plan from diff + route mapMarkdown test plan with stepsreferences/test-plan.md
5. ExecuteRun test plan via agent-browserPass/fail per step, screenshotsreferences/execution.md
6. ReportAggregate results, artifacts, exit codeStructured report + artifactsreferences/report.md

Phase 1: Fingerprint Check

Check if the current changes have already been tested:

python
Read(".expect/fingerprints.json")  # Previous run hashes
# Compare SHA-256 of changed files against stored fingerprints
# If match: "No changes since last test run. Use --force to re-run."
# If no match or --force: continue to Phase 2

Load: Read("references/fingerprint.md")

Phase 2: Diff Scan

Analyze git changes based on --target:

python
if TARGET == "unstaged":
    diff = Bash("git diff")
    files = Bash("git diff --name-only")
elif TARGET == "branch":
    diff = Bash("git diff main...HEAD")
    files = Bash("git diff main...HEAD --name-only")
elif TARGET == "commit":
    diff = Bash("git diff HEAD~1")
    files = Bash("git diff HEAD~1 --name-only")

Classify each changed file into 3 levels:

  1. Direct — the file itself changed
  2. Imported — a file that imports the changed file
  3. Routed — the page/route that renders the changed component

Load: Read("references/diff-scanner.md")

Phase 3: Route Map

Map changed files to testable URLs using .expect/config.yaml:

yaml
# .expect/config.yaml
base_url: http://localhost:3000
route_map:
  "src/components/Header.tsx": ["/", "/about", "/pricing"]
  "src/app/auth/**": ["/login", "/signup", "/forgot-password"]
  "src/app/dashboard/**": ["/dashboard"]

If no route map exists, infer from Next.js App Router / Pages Router conventions.

Load: Read("references/route-map.md")

Phase 4: Test Plan Generation

Build an AI test plan scoped to the diff, using the scope strategy for the current target:

python
scope_strategy = get_scope_strategy(TARGET)  # See references/scope-strategy.md

prompt = f"""
{scope_strategy}

Changes: {diff_summary}
Affected pages: {affected_urls}
Instruction: {INSTRUCTION or "Test that the changes work correctly"}

Generate a test plan with:
1. Page-level checks (loads, no console errors, correct content)
2. Interaction tests (forms, buttons, navigation affected by the diff)
3. Visual regression (compare ARIA snapshots if saved)
4. Accessibility (axe-core scan on affected pages)
"""

If --flow specified, load saved flow from .expect/flows/{slug}.yaml instead of generating.

If NOT --y, present plan to user via AskUserQuestion for review before executing.

Load: Read("references/test-plan.md")

Phase 5: Execution

agent-browser Quick Primer

Floor is >= 0.31.1 (0.27.1 is documented broken on prod pages); current tested release is 0.38.1 (see upstream-version-tested). Commands below hold across this range. 0.30+ adds agent-browser read (agent-readable text extraction) and the --restore / --namespace session-restore workflow for stable, isolated browser state across agent runs. 0.33.0 adds agent-browser a11y [url], an embedded axe-core audit (WCAG tag filtering, selector scoping, iframe-aware text/JSON output) available as both a CLI command and an MCP tool. 0.34.0 adds persistent session-to-tab binding for shared Chrome sessions, with --pin-tab making the binding strict so an externally closed tab yields a stable tab_gone error. 0.35.0 adds --ca-cert <path> for a trusted local CA behind an SSL-inspecting proxy (Linux-only; the targeted alternative to --ignore-https-errors). 0.35.1 changes behaviour the Diff row below relies on: diff snapshot now resets ref numbering per diff and invalidates refs across navigations, so refs captured before a navigation must be re-read rather than reused. 0.36.0 adds experimental WebMCP tool discovery (webmcp list/invoke/result/cancel, brief untrusted catalog summaries on open). 0.37.0 makes record capture the active page at 30 fps and inherits session setup into new tabs. 0.38.0 adds snapshot --delta (full baseline, then unchanged or a compact structural change), screenshot --if-changed --threshold, persistent refs that survive same-document DOM changes, --human / --input-mode pointer realism, and auth login --no-navigate. Full command surface: references/upstream.md and agent-browser skills get core --full.

AreaCommandNotes
Snapshotagent-browser snapshot -iARIA tree w/ @eN refs. -C/--cursor was removed in 0.22
Snapshot delta (0.38+)agent-browser snapshot --deltaFull tree once, then unchanged or a compact structural diff; refs persist across same-document changes
Semantic locatoragent-browser find role button click --name "Continue"Grammar: find <locator> <value> [action]; stable alternative to @eN refs
Interactionfill @e1 "...", click @e2, press Enter, drag @e1 @e2, upload @e1 file.pdfAll take ARIA refs. Add --human for a curved approach instead of an instant jump
Waitswait --load networkidle, wait --text "Success", wait --fn "window.ready"Event-driven, never sleep-based
Networknetwork route "*analytics*" --abort, network route "https://api/*" --body '{...}'Intercept + stub
Statestate save/load auth.json, --session-name <name>Persist auth across runs
Vaultvault store github_pat, vault load github_patEncrypted credential store
Diffdiff snapshot, diff screenshot --baseline /tmp/x.pngARIA + pixel diffing across separate runs; prefer snapshot --delta for within-run before/after on the same page
Capturescreenshot --annotate, screenshot --if-changed --threshold <n> (0.38+), pdf, record start/stop --fps <n>--if-changed skips re-encoding an unchanged screenshot; record defaults to 30 fps (0.37+)
Dashboardagent-browser dashboard start (0.25+)Browser-side runtime inspector on :4848
Run the test plan
python
expect_task = Agent(
  subagent_type="ork:expect-agent",
  prompt=f"""Execute this test plan:
  {test_plan}

  For each step:
  1. Navigate to the URL
  2. Execute the test action
  3. Take a screenshot on failure
  4. Report PASS/FAIL with evidence
  """,
  run_in_background=True,
  model="sonnet",
  max_turns=50
)

# Stream agent-browser progress line-by-line instead of polling (CC 2.1.98+)
# Each stdout line from agent-browser arrives as a notification — useful for
# catching a failing step early rather than waiting for the full plan.
# Full pattern: Read("${CLAUDE_PLUGIN_ROOT}/skills/chain-patterns/references/monitor-patterns.md")
Monitor(pid=expect_task.agent_id)

# For long test plans (>3 min typical), notify on completion — requires
# Remote Control + "Push when Claude decides" config (CC 2.1.110+).
# Skip silently if the user doesn't have Remote Control enabled.
if test_plan_duration_estimate > 180:
    PushNotification(
        message=f"ork:expect complete — {passed}/{total} steps passed on {len(affected_urls)} pages",
        status="proactive"
    )

Load: Read("references/execution.md")

Phase 6: Report

expect Report
═══════════════════════════════════════
Target: unstaged (3 files changed)
Pages tested: 4
Duration: 45s

Results:
  ✓ /login — form renders, submit works
  ✓ /signup — validation triggers on empty fields
  ✗ /dashboard — chart component crashes (TypeError)
  ✓ /settings — preferences save correctly

3 passed, 1 failed

Artifacts:
  .expect/reports/2026-03-26T16-30-00.json
  .expect/screenshots/dashboard-error.png

Load: Read("references/report.md")

Saved Flows

Reusable test sequences stored in .expect/flows/:

yaml
# .expect/flows/login.yaml
name: Login Flow
steps:
  - navigate: /login
  - fill: { selector: "#email", value: "test@example.com" }
  - fill: { selector: "#password", value: "password123" }
  - click: button[type="submit"]
  - assert: { url: "/dashboard" }
  - assert: { text: "Welcome back" }

Run with: expect --flow login

Auto-trigger after UI edits (M125 #2)

When the dev stack is live (dev), saving any .tsx, .jsx, .css, or .scss file (and Next.js route files like app/**/page.tsx, pages/**/*.tsx) emits a nudge to run expect <route>. The hook (posttool/ui-change-detector) is default-on and:

  • skips silently if dev hasn't booted (no agent-browser session to attach to);
  • enforces a 30-second cooldown per route to prevent spam on rapid saves;
  • honors .claude/state/expect-skip.<sessionId> as a per-session opt-out (write any content);
  • honors ORK_EXPECT_AUTO=0 for an env-level kill switch.

Route resolution: app/dashboard/page.tsx → /dashboard, pages/settings.tsx → /settings, component / global-style edits → / (home as proxy). Route groups like app/(marketing)/pricing/page.tsx strip to /pricing.

Show full SKILL.md (500 more words)Show less

ARIA snapshot recording (M125 #6)

After a passing run, the posttool/expect/snapshot-recorder hook persists the captured ARIA tree to .claude/state/expect-snapshots/<route-slug>/<parent-commit>.json. Subsequent expect <route> --diff runs compare against the most recent prior snapshot for that route — surfaces structural regressions (added/removed buttons, label changes, hierarchy shifts) without needing a baseline screenshot.

For the snapshot recorder to fire, the expect run output must contain RUN_COMPLETED|passed, ROUTE|<route>, and ARIA|<json-summary> tags. The agent-browser-driven flow already emits these.

Jev step judge (opt-in, three modes)

ORK_EXPECT_JEV selects the mode. Unset or falsey is off, and off means zero network. shadow (also selected by a truthy ORK_EXPECT_JEV_SHADOW for compatibility) logs one Jev request per step beside the agent's pick: a Choice over the bounded legal-action set built from the interactive ARIA elements, plus three verify Nouls, with probabilities and per-step agreement in the report. 1 or act lets the Jev pick drive the step instead, failing closed to the agent's own pick on any error, timeout, empty or malformed answer, or a confidence below act_confidence_floor. Tunables live in .expect/config.yaml jev_shadow:. Load: Read("references/jev-shadow.md")

When NOT to Use

  • Unit tests — use cover instead
  • API-only changes — no browser UI to test
  • Generated files — skip build artifacts, lock files
  • Docs-only changes — unless you want to verify docs site rendering

Quality Bar

Done means all of these hold:

  • Test-plan scope is derived from the git diff for the chosen --target — no unaffected page appears in the plan.
  • Every changed file maps to at least one tested route (via .expect/config.yaml or inferred convention) OR is explicitly excluded as API-only, generated, or docs-only.
  • Fingerprint check runs first; a diff unchanged since the last run skips execution instead of re-testing.
  • Unless -y is passed, the plan is presented for review before any browser action runs.
  • Each executed step reports PASS or FAIL with evidence (screenshot on failure), and the report's pass/fail totals match the steps actually run.
  • The report's exit code is non-zero whenever any step failed.

Budget: at most 30 minutes total and 5 minutes per page (a slower page is skipped, per references/test-plan.md; defaults, .expect/config.yaml overrides), 50 executor turns (max_turns=50), one sequential ork:expect-agent, never parallel browsers (rules/no-parallel-browsers.md); stop and report partial results at the finish line or the first cap, whichever comes first.

  • agent-browser — Browser automation engine (required dependency)
  • ork:cover — Test suite generation (unit/integration/e2e)
  • ork:verify — Grade existing test quality
  • testing-e2e — Playwright patterns and best practices

References

Load on demand with Read("references/<file>"):

FileContent
fingerprint.mdSHA-256 gating logic
diff-scanner.mdGit diff parsing + 3-level classification
route-map.mdFile-to-URL mapping conventions
test-plan.mdAI test plan generation prompt templates
execution.mdagent-browser orchestration patterns
report.mdReport format + artifact storage
config-schema.md.expect/config.yaml full schema
aria-diffing.mdARIA snapshot comparison for semantic diffing
scope-strategy.mdTest depth strategy per target mode
saved-flows.mdMarkdown+YAML flow format, adaptive replay
rrweb-recording.mdrrweb DOM replay integration
human-review.mdAskUserQuestion plan review gate
ci-integration.mdGitHub Actions workflow + pre-push hooks
research.mdmillionco/expect architecture analysis
jev-shadow.mdOpt-in Jev step judge per step: bounded-action Choice + verify Nouls; shadow logs, act drives with fail-closed fallback

Version: 1.0.0 (March 2026) — Initial scaffold, M99 milestone

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 36 other files (scripts, references) in src/skills/expect of yonatangross/orchestkit.

  • SKILL.md
  • flows/crud.md
  • flows/login.md
  • flows/navigation.md
  • jev-shadow.defaults.yaml
  • references/aria-diffing.md
  • references/ci-integration.md
  • references/config-schema.md
  • references/diff-scanner.md
  • references/execution.md
  • references/fingerprint.md
  • references/human-review.md
  • references/jev-shadow.md
  • references/report.md
  • references/research.md
  • references/route-map.md
  • references/rrweb-recording.md
  • references/saved-flows.md
  • references/scope-strategy.md
  • … and 18 more

Open the folder on GitHubat commit 0ef71d2

Compare with similar skills

Expect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Expect compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Expect this skillyonatangross/orchestkit289—~4.8kAutomated safety check: NotesMIT
Kane CLI Browser TestingLambdaTest/kane-cli247—~8.4kAutomated safety check: PassApache-2.0
Quality Engineering Playwright CLIHoangNguyen0403/agent-skills-standard571—~1.2kAutomated safety check: PassMIT
Playwright Component Testingmellowagain/gitarena1151 repos~2.6kAutomated safety check: PassMIT
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone
Playwright Rs Usagepadamson/playwright-rust153—~5.3kAutomated safety check: PassApache-2.0

Similar skills

  • Kane CLI Browser Testing

    LambdaTest/kane-cli

    Drives a real browser through the kane-cli tool and designs requirement-linked test suites from a PRD or a plain description, with mobile and cloud-grid runs.

    247 GitHub stars~8.4k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Quality Engineering Playwright CLI

    HoangNguyen0403/agent-skills-standard

    Standardizes token-efficient browser automation via playwright-cli, with Playwright MCP as the fallback driver.

    571 GitHub stars~1.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Playwright Component Testing

    mellowagain/gitarena

    Set up component testing with Playwright using a story gallery — scaffold stories and a gallery dev page driven by the built-in mount fixture, no dedicated component-testing runtime.

    115 GitHub starsUsed in 1 repo~2.6k tokens
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Playwright Rs Usage

    padamson/playwright-rust

    Procedural reference for using playwright-rs in Rust browser-automation code — object model (Browser/Context/Page/Locator), the locator!() macro, builder pattern for options, auto-wait semantics…

    153 GitHub stars~5.3k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Playwright Trace

    mellowagain/gitarena

    Inspect Playwright trace files from the command line — list actions, view requests, console, errors, snapshots and screenshots.

    115 GitHub starsUsed in 1 repo~1.2k tokens
    Testing & QAAuto-check passed

More from yonatangross/orchestkit

All 108 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    289 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    289 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    289 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    289 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    289 GitHub stars~4.5k tokensUpdated today
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    289 GitHub stars~3.9k tokensUpdated today
    Auto-check: notes

Works with

Questions about Expect

What does Expect do?

Diff-aware AI browser testing — reads the git diff, maps changes to affected pages via the route map, generates a targeted test plan, and executes it via agent-browser (Rust daemon + CDP…. Expect is an agent skill from yonatangross/orchestkit. Diff-aware AI browser testing — reads the git diff, maps changes to affected pages via the route map, generates a targeted test plan, and executes it via agent-browser (Rust daemon + CDP, ARIA-tree-first) with pass/fail reporting.

When should I use Expect?

Expect fits situations like: testing UI changes; verifying PRs before merge; running regression checks on changed components.

How do I install Expect in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill expect -a claude-code`. Or copy the skill folder (src/skills/expect in yonatangross/orchestkit) into .claude/skills/expect in your project. Claude Code loads it when a task matches its description.

How do I install Expect in Codex?

Run `npx skills add yonatangross/orchestkit --skill expect -a codex`. Or copy the skill folder (src/skills/expect in yonatangross/orchestkit) into .agents/skills/expect in your project. Codex loads it when a task matches its description.

Can I use Expect in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill expect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/expect, .gemini/skills/expect, .github/skills/expect and .opencode/skills/expect in your project.

What does Expect need to run?

Going by SKILL.md and its folder, Expect needs the command-line tools its instructions call (vault, git and npx). Our summary lists: Python 3; Node.js. Its frontmatter pre-approves these tools: AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, ToolSearch, WebFetch, Monitor, PushNotification, mcp__memory__search_nodes. Compatibility (from SKILL.md): Claude Code 2.1.277+. Requires agent-browser >= 0.31.1 (Rust-native, no Playwright; 0.27.1 documented broken on prod pages)..

Does Expect access the network?

SKILL.md contains no URLs. Its commands use git and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Expect safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Expect use?

Expect is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Expect use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to Expect?

Skills that share tags, products or a category with Expect: Kane CLI Browser Testing (LambdaTest/kane-cli, 247 stars), Quality Engineering Playwright CLI (HoangNguyen0403/agent-skills-standard, 571 stars), Playwright Component Testing (mellowagain/gitarena, 115 stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Expect?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 289 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.