Agent skill

Evidence-Driven Testing

by michaelshimeles in michaelshimeles/skills

Records an annotated screen recording of the agent testing an app hands-on, then posts the video and a results summary to the PR and tracker issue.

No licenceAuto-check passedTesting & QA

Install Evidence-Driven Testing

skills CLI
$ npx skills add michaelshimeles/skills --skill evidence-driven-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install michaelshimeles/skills evidence-driven-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/michaelshimeles/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/evidence-driven-testing .claude/skills/evidence-driven-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evidence-driven-testing
GitHub stars
1.3k
Used in
1 other repo
Token cost
~3.9k tokens
SKILL.md length
1,819 words
Files
2 (incl. scripts)
Skills in repo
4
Repo updated
First seen
Licence
None found

At a glance

Records an annotated screen recording of the agent testing an app hands-on, then posts the video and a results summary to the PR and tracker issue.

  • Works in 5 steps: Prepare the screen → Start recording → Test via computer use, annotating as you… → …
  • Proving a UI change works before merging, with a video instead of prose
  • SKILL.md covers Inputs, The recorder, Instructions and Guardrails, plus 4 more sections
  • Runs Python scripts from its folder; calls python3, git and npx

What it does

The agent drives the app itself through computer use while `scripts/evidence.py` records the display with FFmpeg. Each assertion added during testing is timestamped, burned into the video when the recording stops and checked with ffprobe, and the session folder receives `evidence.mp4`, `report.md` and `manifest.json`. A recording only counts as evidence if it shows the live, interactive session.

You supply test targets phrased as testable statements and, optionally, a PR or issue to post to. A `doctor` command checks FFmpeg, ffprobe, libx264 and the ass filter, then reports which capture source works: x11grab on Linux X11, wf-recorder on supported Wayland compositors, avfoundation on macOS and gdigrab on Windows. For headless setups and non-UI changes the skill falls back to scripted screenshots, measured numbers or output pairs, and posting uses the GitHub CLI or an equivalent.

When your agent uses it

  • Proving a UI change works before merging, with a video instead of prose
  • Attaching test evidence to a pull request and its tracker issue
  • Documenting a non-UI change with measured numbers or before-and-after output

Example prompts

  • “Test the new checkout flow hands-on, record it with annotations and post the video to the PR.”
  • “Record evidence that the dark-mode toggle survives a page reload and attach it to the issue.”
  • “This change has no UI. Capture before-and-after output for the parser fix as evidence for the PR.”

Requirements

  • Python 3 with ffmpeg and ffprobe (libx264 and the ass filter)
  • A GUI environment the agent can drive, or the cua-driver CLI
  • An authenticated browser session for the app under test
  • The GitHub CLI or an equivalent to post evidence
  • Compatibility (from SKILL.md): Screen-recording path requires a GUI environment the agent can drive — built-in computer use, or the cua-driver CLI (trycua/cua) when the harness has no computer-use tools — plus an authenticated browser session for the app under test. The bundled recorder (scripts/evidence.py) runs on Linux (X11 via x11grab, Wayland via wf-recorder), macOS (avfoundation, needs Screen Recording permission) and Windows (gdigrab) and needs Python 3 plus ffmpeg + ffprobe built with libx264 and the ass filter. The headless path requires only a running app and a scriptable browser (e.g. Playwright via npx). Posting evidence requires gh (GitHub CLI) or equivalent.

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Prepare the screen
  2. Start recording
  3. Test via computer use, annotating as you go
  4. Stop and review
  5. Post the evidence

What it can do on your machine

Read from SKILL.md and the folder at commit 4b72f46. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • git
    • npx
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Screen-recording path requires a GUI environment the agent can drive — built-in computer use, or the cua-driver CLI (trycua/cua) when the harness has no computer-use tools — plus an authenticated browser session for the app under test. The bundled recorder (scripts/evidence.py) runs on Linux (X11 via x11grab, Wayland via wf-recorder), macOS (avfoundation, needs Screen Recording permission) and Windows (gdigrab) and needs Python 3 plus ffmpeg + ffprobe built with libx264 and the ass filter. The headless path requires only a running app and a scriptable browser (e.g. Playwright via npx). Posting evidence requires gh (GitHub CLI) or equivalent.

    From compatibility in the SKILL.md frontmatter.

Context cost

Evidence-Driven Testing loads about 3.9k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 1,819 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,819 words (~3,880 tokens).

“Record annotated proof of behavior, then attach it to the PR and tracker issue.”

— opening of SKILL.md by michaelshimeles
name
evidence-driven-testing
compatibility
Screen-recording path requires a GUI environment the agent can drive — built-in computer use, or the cua-driver CLI (trycua/cua) when the harness has no computer-use tools — plus an authenticated browser session for the app under test. The bundled recorder (scripts/evidence.py) runs on Linux (X11 via x11grab, Wayland via wf-recorder), macOS (avfoundation, needs Screen Recording permission) and Windows (gdigrab) and needs Python 3 plus ffmpeg + ffprobe built with libx264 and the ass filter. The headless path requires only a running app and a scriptable browser (e.g. Playwright via npx). Posting evidence requires gh (GitHub CLI) or equivalent.
metadata.version
1.2

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file (scripts) in evidence-driven-testing of michaelshimeles/skills.

  • SKILL.md
  • scripts/evidence.py

Open the folder on GitHubat commit 4b72f46

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in michaelshimeles/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Evidence-Driven Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evidence-Driven Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evidence-Driven Testing this skillmichaelshimeles/skills1.3k1 repos~3.9kAutomated safety check: PassNone
Electron App Screenshotkeybase/client9.3k—~476Automated safety check: PassBSD-3-Clause
Site Bug Huntdebs-obrien/debbie.codes142—~1.2kAutomated safety check: PassNone
Windows QA EngineerCodeAlive-AI/ai-driven-development157—~1.4kAutomated safety check: PassMIT
Agentic Browser Testingpetrkindlmann/qa-skills165—~4.5kAutomated safety check: PassMIT
Supporthomebridge-plugins/homebridge-eufy223—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Takes a screenshot of a running Electron desktop app through playwright-cli over remote debugging, shrinks it and shows it so you can check the UI visually.

    9.3k GitHub stars~476 tokensUpdated today
    Testing & QAAuto-check passed
  • Site Bug Hunt

    debs-obrien/debbie.codes

    Exploratory QA for debbie.codes. An agent skill from debs-obrien/debbie.codes.

    142 GitHub stars~1.2k tokensUpdated 7 days ago
    Testing & QAAuto-check passed
  • Windows QA Engineer

    CodeAlive-AI/ai-driven-development

    A skill your agent uses when testing Windows 11 desktop apps (WinForms/WPF/UWP) via UFO UIA/Win32 automation MCP.

    157 GitHub stars~1.4k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    165 GitHub stars~4.5k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Support

    homebridge-plugins/homebridge-eufy

    Triage a GitHub issue using diagnostics archives and logs. An agent skill from homebridge-plugins/homebridge-eufy.

    223 GitHub stars~1.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Record E2E Gif

    lablup/backend.ai-webui

    Record Playwright e2e tests as one GIF per test case (video → ffmpeg palette GIF) and return a markdown table for a PR description.

    133 GitHub stars~907 tokensUpdated today
    Testing & QAAuto-check: notes

More from michaelshimeles/skills

  • Before and After Screenshots

    michaelshimeles/skills

    Captures before and after screenshots of two URLs or images with the Vercel before-and-after CLI, uploads them and can add a markdown comparison to a pull request.

    1.3k GitHub starsUsed in 2 repos~1k tokens
    Auto-check passed
  • Greploop Apps

    michaelshimeles/skills

    Loops on a large pull request, merge request or Perforce changelist, fixing Greptile findings until it scores 5/5 with no unresolved comments.

    1.3k GitHub starsUsed in 1 repo~3.6k tokens
    Auto-check passed
  • New Feature

    michaelshimeles/skills

    Start a new task in an isolated Git worktree branched from origin/main so multiple agents can work on the same repo in parallel without conflicts.

    1.3k GitHub starsUsed in 1 repo~699 tokens
    Auto-check passed

Questions about Evidence-Driven Testing

What does Evidence-Driven Testing do?

Records an annotated screen recording of the agent testing an app hands-on, then posts the video and a results summary to the PR and tracker issue. py` records the display with FFmpeg.json`.

When should I use Evidence-Driven Testing?

Evidence-Driven Testing fits situations like: proving a UI change works before merging, with a video instead of prose; attaching test evidence to a pull request and its tracker issue; documenting a non-UI change with measured numbers or before-and-after output.

How do I install Evidence-Driven Testing in Claude Code?

Run `npx skills add michaelshimeles/skills --skill evidence-driven-testing -a claude-code`. Or copy the skill folder (evidence-driven-testing in michaelshimeles/skills) into .claude/skills/evidence-driven-testing in your project. Claude Code loads it when a task matches its description.

How do I install Evidence-Driven Testing in Codex?

Run `npx skills add michaelshimeles/skills --skill evidence-driven-testing -a codex`. Or copy the skill folder (evidence-driven-testing in michaelshimeles/skills) into .agents/skills/evidence-driven-testing in your project. Codex loads it when a task matches its description.

Can I use Evidence-Driven Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add michaelshimeles/skills --skill evidence-driven-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evidence-driven-testing, .gemini/skills/evidence-driven-testing, .github/skills/evidence-driven-testing and .opencode/skills/evidence-driven-testing in your project.

What does Evidence-Driven Testing need to run?

Going by SKILL.md and its folder, Evidence-Driven Testing needs Python for the scripts in its folder and the command-line tools its instructions call (python3, git, npx and gh). Our summary lists: Python 3 with ffmpeg and ffprobe (libx264 and the ass filter); A GUI environment the agent can drive, or the cua-driver CLI; An authenticated browser session for the app under test; The GitHub CLI or an equivalent to post evidence. Compatibility (from SKILL.md): Screen-recording path requires a GUI environment the agent can drive — built-in computer use, or the cua-driver CLI (trycua/cua) when the harness has no computer-use tools — plus an authenticated browser session for the app under test. The bundled recorder (scripts/evidence.py) runs on Linux (X11 via x11grab, Wayland via wf-recorder), macOS (avfoundation, needs Screen Recording permission) and Windows (gdigrab) and needs Python 3 plus ffmpeg + ffprobe built with libx264 and the ass filter. The headless path requires only a running app and a scriptable browser (e.g. Playwright via npx). Posting evidence requires gh (GitHub CLI) or equivalent..

Does Evidence-Driven Testing access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Evidence-Driven Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evidence-Driven Testing use?

No licence was found for Evidence-Driven Testing or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Evidence-Driven Testing use?

About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evidence-Driven Testing?

Skills that share tags, products or a category with Evidence-Driven Testing: Electron App Screenshot (keybase/client, 9.3k stars), Site Bug Hunt (debs-obrien/debbie.codes, 142 stars), Windows QA Engineer (CodeAlive-AI/ai-driven-development, 157 stars) and Agentic Browser Testing (petrkindlmann/qa-skills, 165 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evidence-Driven Testing?

michaelshimeles (a GitHub user) maintains it in michaelshimeles/skills, which has 1,267 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 4, 2026.

Source: michaelshimeles/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.