Agent skill

Diff-Driven QA

by Skyvern-AI in Skyvern-AI/skyvern

Reads your git diff, decides whether the change needs browser QA, API checks or repo tests, runs that validation and reports pass or fail with evidence.

AGPL-3.0Auto-check: warningsTesting & QA

Install Diff-Driven QA

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add Skyvern-AI/skyvern --skill qa -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Skyvern-AI/skyvern qa --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Skyvern-AI/skyvern.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skyvern/cli/skills/qa .claude/skills/qa && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qa
GitHub stars
23k
Token cost
~4.7k tokens
SKILL.md length
1,860 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Reads your git diff, decides whether the change needs browser QA, API checks or repo tests, runs that validation and reports pass or fail with evidence.

  • Works in 5 steps: Understand the Changes → Classify the Diff → Choose the Validation Strategy → …
  • Validating code changes before opening a pull request
  • SKILL.md covers Quick Start, How It Works, Step 1: Understand the Changes and Step 2: Classify the Diff, plus 7 more sections
  • Calls git, curl and ngrok

What it does

The /qa command starts from git diff, comparing against the last commit or the working tree, and reads every changed file that affects behavior: routes, components, visible text, forms, schemas, endpoints, validators and the tests that changed. It then classifies the diff as frontend or browser, backend API, backend-internal or mixed.

Frontend changes get browser QA against the dev server, backend API changes get the backend started locally and targeted requests run, backend-internal changes use the repo's own validation, and mixed changes get both. It accepts an explicit frontend URL or a short instruction such as validating a particular API. It is not a general site crawler and should not invent API checks unrelated to the diff; when the diff is empty there is nothing to QA. The result is a pass or fail report with concrete evidence.

When your agent uses it

  • Validating code changes before opening a pull request
  • Checking a UI change in the browser against the local dev server
  • Running targeted API requests for changed route handlers
  • Deciding which kind of validation a mixed frontend and backend diff needs

Example prompts

  • “/qa http://localhost:3000 after my change to the checkout form.”
  • “QA my uncommitted changes and tell me whether the new filters API still behaves.”
  • “Run the right validation for the last commit and report pass or fail with evidence.”

Requirements

  • A git repository with a diff to test
  • A locally runnable dev server or backend for the changed code

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Understand the Changes
  2. Classify the Diff
  3. Choose the Validation Strategy
  4. Report Results
  5. Post Evidence to PR

What it can do on your machine

Read from SKILL.md and the folder at commit 94c7ee1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • curl
    • ngrok
    • npm
    • go

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, curl and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Diff-Driven QA loads about 4.7k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 1,860 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a paste, webhook or tunnelling service often used to send data outSKILL.md:149
    onnect(cdp_url="wss://<ngrok-subdomain>.ngrok-free.app/devtools/browser/<id>")

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Skyvern-AI/skyvern at commit 94c7ee1, republished under its AGPL-3.0 licence (© Skyvern-AI). 1,860 words, ~4,700 tokens.

Download SKILL.mdSave it as .claude/skills/qa/SKILL.md (or your agent's skills folder).
name
qa
description
QA test your code changes by reading your git diff, choosing the right validation path for frontend/browser and backend changes, and reporting pass/fail with evidence.

QA — Validate Frontend and Backend Changes

Read the diff, classify what changed, and run the right validation path: browser QA for frontend/browser changes, API validation for backend surface changes, repo-native validation for backend-internal changes, and both for mixed changes.

<!-- NOTE: .agents/skills/qa/SKILL.md is the repository canonical source.
     Keep both synchronized copies in sync with it:
     1. skyvern/cli/skills/qa/SKILL.md  (bundled with the pip package)
     2. skyvern/cli/mcp_tools/prompts.py (QA_TEST_CONTENT for the MCP prompt) -->

You changed code. This skill is diff-driven first: it reads what changed, understands the affected behavior, and validates that behavior with the right tools. It is not a generic website crawler, and it should not invent random API checks that are unrelated to the diff.

Quick Start

text
/qa                              # Diff-based: choose the right validation path automatically
/qa http://localhost:3000        # Same, explicit frontend URL
/qa -- validate the workflow filters API

How It Works

  1. Read the code changes from git diff
  2. Read the changed files to understand behavior, routes, schemas, and UI
  3. Classify the diff as frontend/browser, backend API, backend-internal, or mixed
  4. Run the right validation flow
  5. Report pass/fail with concrete evidence

Step 1: Understand the Changes

Get the diff
bash
# What files changed?
git diff --name-only HEAD~1     # vs last commit (if changes are committed)
git diff --name-only            # vs working tree (if uncommitted)

# Full diff for context
git diff HEAD~1                 # or git diff for uncommitted

Pick whichever diff has content. If both are empty, there is nothing diff-driven to QA.

Read the changed files

Read the full contents of every changed file that affects behavior:

  • Frontend files: .tsx, .jsx, .ts, .js, .css, .html
  • Backend/API files: routes, controllers, request/response schemas, serializers, handlers
  • Backend-internal files: services, workers, business logic, validators, data-layer code
  • Tests that changed alongside the implementation

Look for:

  • route paths and page entry points
  • component names, visible text, forms, buttons, error states
  • API endpoints, request params, response fields, auth requirements
  • validation logic, branching behavior, feature flags, empty states
  • tests that describe the expected behavior

Step 2: Classify the Diff

ModeTriggerPrimary validation
Frontend/browserUI/routes/components/styles changedBrowser QA against the dev server
Backend APIRoute handlers, request/response schemas, or externally visible API behavior changedStart backend locally and run targeted API requests
Backend-internalServices/workers/business logic changed without public API surface changesRepo-native fast checks plus targeted tests
MixedFrontend/browser and backend changed togetherBackend validation first, then frontend/browser QA

Use these rules:

  • If both frontend/browser and backend changed, treat it as Mixed.
  • If only backend internals changed, do not invent unrelated browser tests or random API calls.
  • If a backend change might affect the public contract, inspect routes, schemas, and tests before choosing backend-internal.
  • If the diff is mostly documentation or comments, keep QA lightweight and report that no behavioral validation was warranted.

Step 3: Choose the Validation Strategy

Frontend/browser mode

Use browser automation against the dev server. Validate the specific UI changes plus 1-2 adjacent regression checks.

Backend API mode

Use the repo's documented local startup and auth instructions, start the backend if needed, identify the changed endpoint(s), and run targeted HTTP requests to validate the changed contract.

Backend-internal mode

Run the repo's fast verification commands first, then targeted unit/integration/scenario tests for the changed logic. Only start the backend and do live API calls if the change affects exposed behavior.

Mixed mode

Validate the backend first, then run frontend/browser QA against the flow that depends on it. If the backend contract is broken, frontend results are not trustworthy.

Step 4A: Frontend/Browser QA

Find the dev server

If the user provided a URL, use it. Otherwise auto-detect common local ports:

text
5173, 3000, 3001, 8080, 8000, 4200

If none respond, start the most direct repo-documented local command for the changed surface. If the diff needs both frontend and backend running together and the repo provides a combined frontend/backend dev script, prefer that. Only ask the user to start something manually if the repo has no documented command or startup fails.

Connect to a browser

Try these in order:

Option A: Local browser (fastest)
text
skyvern_browser_session_create(local=true, headless=false, timeout=15)

Use local=true so the browser can reach localhost.

Option B: Local browser via tunnel

If local session creation fails because the MCP server is remote, the cloud browser cannot reach localhost. Tell the user to run:

bash
# Terminal 1: Launch a local browser with CDP exposed
skyvern browser serve --port 9222

# Terminal 2: Tunnel it to the internet
ngrok http 9222

Then connect:

text
skyvern_browser_session_connect(cdp_url="wss://<ngrok-subdomain>.ngrok-free.app/devtools/browser/<id>")

The user can get the browser ID from the skyvern browser serve output or by calling the ngrok URL's /json endpoint.

Option C: Cloud browser
text
skyvern_browser_session_create(timeout=15)

Only works for publicly reachable URLs. localhost URLs will not work here.

Generate frontend/browser test cases

For each changed frontend file, create targeted checks. Examples:

text
Test 1: Settings page renders the new "Retry failed run" button
  - Navigate to /settings/runs
  - Assert: button with text "Retry failed run" exists
  - Click it
  - Assert: success toast appears

Test 2: Adjacent regression
  - Verify the existing "Delete run" action still works or is still visible

Be specific. Do not write "verify the page works."

Run the frontend/browser tests

For each test case:

text
skyvern_navigate(url="http://localhost:<port>/<route>")

Health gate after navigation:

In extension mode, use this health gate. Do not call skyvern_evaluate. Before each skyvern_navigate to a route under test, call skyvern_get_errors(clear=True), skyvern_console_messages(clear=True), skyvern_handle_dialog(clear=True), and skyvern_network_requests(clear=True); discard what they return.

text
skyvern_tab_list()
skyvern_get_errors()
skyvern_console_messages(level="error")
skyvern_get_html(selector="body")
skyvern_find(by="role", value="alert")
skyvern_find(by="role", value="dialog")
skyvern_handle_dialog()

PASS requires the active tab URL to match the expected route, with no unexpected login redirect. The body must contain the expected page content, not only scripts or an empty app container. Require no unexpected error text, visible alerts, visible dialogs, JavaScript errors, console errors, or JavaScript dialog events. skyvern_handle_dialog reads dialog history; JavaScript dialogs are auto-dismissed by default. If a call fails or evidence is incomplete, record FAIL and stop this test. Do not treat missing evidence as PASS.

Outside extension mode, use this health gate:

text
skyvern_evaluate(expression="(() => {
  const errors = [];
  const body = document.body?.innerText || '';
  if (body.includes('Something went wrong')) errors.push('error_message');
  if (body.includes('Cannot read properties')) errors.push('js_error_in_ui');
  if (/\\bundefined\\b/.test(body) && !/\\bif\\b|\\btypeof\\b|\\bdocument|tutorial|example/i.test(body) && body.length < 5000) errors.push('undefined_text');
  if (body.includes('connection refused')) errors.push('connection_refused');
  if (/sign.?in|log.?in|auth/i.test(window.location.pathname)) errors.push('auth_redirect');
  if (document.querySelector('[role=\"alert\"]')) errors.push('alert_element');
  if (!document.querySelector('main, [role=\"main\"], nav, header, h1, h2, [class*=\"layout\" i], [class*=\"page\" i], [class*=\"app\" i]'))
    errors.push('blank_page');
  return JSON.stringify({ pass: errors.length === 0, errors });
})()")

In extension mode, use skyvern_find, skyvern_get_html, skyvern_get_value, or skyvern_tab_list for assertions.

text
skyvern_find(by="role", value="button")
skyvern_get_html(selector="h1")
skyvern_tab_list()

Outside extension mode, prefer deterministic DOM assertions:

text
skyvern_evaluate(expression="!!document.querySelector('button')")
skyvern_evaluate(expression="document.querySelector('h1')?.textContent?.trim()")
skyvern_evaluate(expression="window.location.pathname")

Use interaction tools when needed:

text
skyvern_act(prompt="Click the 'Retry failed run' button")
skyvern_act(prompt="Fill the email field with 'test@example.com' and click Submit")
skyvern_validate(prompt="The page shows the success toast and the form is no longer loading")
skyvern_screenshot()

In extension mode, call skyvern_network_requests() and inspect captured requests for failures or HTTP status codes of 400 or greater. If capture is unavailable or incomplete, report the network check as unverified. Outside extension mode, check for failed network requests once per page:

text
skyvern_evaluate(expression="(() => {
  const entries = performance.getEntriesByType('resource').filter(e => e.responseStatus >= 400);
  return JSON.stringify({ failed: entries.map(e => ({ url: e.name, status: e.responseStatus })).slice(0, 5) });
})()")

Step 4B: Backend API QA

Gather repo-local context first

Before starting the server or sending requests, read the repo's local instructions:

  • README, AGENTS.md, CLAUDE.md, Makefile, package.json, pyproject.toml
  • existing test files for the changed endpoints
  • any docs that describe auth, local ports, or startup commands

Do not guess the startup command if the repo already documents one.

Start the backend if needed

If the backend is not already responding on the expected local port:

  1. Start it with the most direct repo-documented local command for the changed surface. If the validation needs both frontend and backend and the repo documents a combined dev environment command, prefer that over inventing separate startup steps.
  2. Wait for readiness
  3. Confirm the API is reachable before sending validation requests

If the repo requires background processes, start them in the background and keep notes on how you did it.

Identify the changed API surface

Use the diff to answer:

  • Which endpoints changed?
  • Which request parameters, headers, or bodies changed?
  • Which response fields or status codes changed?
  • Did auth requirements change?
  • Is there a create/update/delete side effect that needs follow-up verification?

Do not stop at the route file. Read the full handler, schema, and any changed tests.

Generate backend API test cases

For each changed endpoint, create targeted checks:

  • Happy path with valid parameters
  • Empty result or not-found path where applicable
  • Invalid input or validation error path
  • Combined filters or sorting behavior when relevant
  • Follow-up read after mutation endpoints

Examples:

text
Test 1: GET /api/runs returns the new field in the response body
Test 2: GET /api/runs?status=missing returns an empty list, not a 500
Test 3: POST /api/runs rejects invalid payload with a 4xx validation error
Test 4: PATCH /api/runs/:id updates the record and a follow-up GET shows the change
Show full SKILL.md (757 more words)Show less
Execute the API requests

Use the repo's documented auth scheme and local base URL. Use curl, the repo SDK, or a small one-off client if that is clearer than shell quoting. Prefer simple, inspectable commands.

Examples:

bash
curl -sS -H "Authorization: Bearer <token>" \
  "http://localhost:<port>/api/..."

curl -sS -X POST \
  -H "Content-Type: application/json" \
  -H "<auth-header>: <token>" \
  -d '{"example":"value"}' \
  "http://localhost:<port>/api/..."

Capture:

  • request being tested
  • status code
  • response body snippet or parsed result
  • whether the changed field/behavior is present

If the endpoint is authenticated and you cannot obtain local credentials from repo docs, say so clearly and stop rather than faking coverage.

Step 4C: Backend-Internal QA

If the diff is backend-only but does not change an exposed endpoint or UI flow:

  1. Run the repo's fastest compile/type/lint checks for the changed files
  2. Run targeted unit/integration/scenario tests that cover the changed logic
  3. Add live API calls only if the internal change affects exposed behavior

Examples of appropriate checks:

  • compile/type-check the changed files
  • targeted pytest, npm test, go test, or equivalent
  • scenario/integration tests around changed services or workflows

Examples of inappropriate checks:

  • hitting an unrelated health endpoint and calling it "validated"
  • browsing the UI when no frontend behavior changed
  • calling random APIs just because the backend changed

Step 5: Report Results

markdown
## QA Report

### Validation Mode
- Mode: Backend API
- Scope: `routes/runs.py`, `schemas/run_response.py`

### Changes Tested
- Added `retryable` field to run responses
- Updated `status` filter handling

### Results
| # | Test | Result | Evidence |
|---|------|--------|----------|
| 1 | GET /api/runs returns `retryable` for valid runs | PASS | HTTP 200, field present in response |
| 2 | GET /api/runs?status=missing returns empty list | PASS | HTTP 200, `[]` |
| 3 | GET /api/runs?status=invalid returns validation error | PASS | HTTP 422 |
| 4 | Frontend runs page still renders filter state | PASS | screenshot_3 |

### Issues Found
1. `retryable` is missing from one branch of the response serializer.

### Verdict
3/4 tests passed. 1 issue found.

Report the evidence that actually matters:

  • screenshots for frontend/browser results
  • status codes and response snippets for backend API results
  • command + failing assertion for unit/integration tests

Step 6: Post Evidence to PR

After generating the QA report, persist it to the pull request as a sticky comment so the evidence survives beyond the conversation.

Save and post the sticky comment

Write the full report markdown from Step 5 to .qa/latest-report.md with a filesystem editing tool. Do not place report text in a shell command, variable assignment, heredoc, or command substitution.

Then run the fixed command below. It reads .qa/latest-report.md and passes it to gh as a literal argument vector without shell evaluation, updating this user's own <!-- skyvern-qa-report --> comment when one already exists:

bash
skyvern skill post-qa-report

If no PR exists for the current branch, the command leaves the report at .qa/latest-report.md. Tell the user to run /qa again after creating a PR. Do not create a PR just to post a QA report.

Screenshot handling

Screenshots taken during QA (via skyvern_screenshot()) are saved locally for the agent's verification. They are not uploaded to the PR comment because GitHub's API does not support image uploads in issue comments. The text report describes what was observed.

If the user asks to preserve screenshots, save them to .qa/screenshots/ and tell the user the local path. Do not include local file paths in the PR comment — they are meaningless to other reviewers.

Rules
  • Use skyvern skill post-qa-report; do not reconstruct its gh calls in a shell.
  • The command includes the <!-- skyvern-qa-report --> marker, short commit hash, and UTC timestamp.
  • Do not create a PR just to post a QA report — that is the user's decision.
  • If gh is not available or not authenticated, fall back to saving the report locally and tell the user.

Error Handling

ProblemAction
No git diff foundAsk what behavior to validate, then fall back to explore mode
Frontend dev server not runningStart the most direct repo-documented local command for the changed surface; prefer a combined dev command only when the validation needs both frontend and backend; only ask the user if no documented command exists or startup fails
Backend server not runningStart the most direct repo-documented local command for the changed surface; prefer a combined dev environment command only when the validation needs both sides
Cannot identify changed endpointRead changed routes, schemas, and tests before proceeding
Auth required but no local creds availableReport the blocker clearly; do not fake coverage
Component does not renderCapture screenshot and specific UI error
API returns unexpected 5xxSave request/response evidence and report the regression

Session Recording

Before closing, fetch the session recording so you can include it in the QA report.

Via MCP tools
text
skyvern_browser_session_get(session_id="pbs_xxx")
→ Returns app_url (watch in browser) and recordings (download URLs)

Or simply close — skyvern_browser_session_close() now returns recording data too:

text
skyvern_browser_session_close()
→ { session_id, closed, app_url, recordings: [{url, filename}], downloaded_files: [{url, filename}] }
Via CLI
bash
skyvern browser session get --session pbs_xxx --json
# → { "app_url": "https://...", "recordings": [...] }

skyvern browser session close --session pbs_xxx --json
# → { "session_id": "pbs_xxx", "closed": true, "app_url": "https://...", "recordings": [...] }

Include in the QA report:

  • Watch recording: the app_url value
  • Download recording: the first entry in recordings[].url

Session Cleanup

Always close browser sessions when done:

text
skyvern_browser_session_close()

If you started local servers or background processes, leave the user a clear note about what is still running.

Fallback: Explore Mode

If there is no useful diff, fall back to explicit exploration:

  1. Ask what behavior should be validated
  2. If it is a frontend flow, use browser QA
  3. If it is a backend/API flow, run targeted local API checks
  4. Report findings with the same evidence standard

The primary mode is still diff-driven. Always try to understand the code changes first.

© Skyvern-AI, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skyvern/cli/skills/qa of Skyvern-AI/skyvern.

Open the folder on GitHubat commit 94c7ee1

Compare with similar skills

Diff-Driven QA next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Diff-Driven QA compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Diff-Driven QA this skillSkyvern-AI/skyvern23k—~4.7kAutomated safety check: WarnAGPL-3.0
QA Test and Fixgarrytan/gstack136k—~13kAutomated safety check: NotesMIT
QA Report Onlygarrytan/gstack136k—~11kAutomated safety check: NotesMIT
Playwright Regression Testingfugazi/test-automation-skills-agents247—~1.4kAutomated safety check: PassMIT
Playwright Screen Recordingliaohch3/claude-tap3.3k—~714Automated safety check: PassMIT
Reprovaadin/web-components582—~1.3kAutomated safety check: PassNone

Similar skills

  • QA Test and Fix

    garrytan/gstack

    Tests a site or service for bugs, fixes what it finds with one atomic commit per fix, and reports fix evidence and a ship-readiness summary at one of three depths.

    136k GitHub stars~13k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • QA Report Only

    garrytan/gstack

    Tests a browser app, API, CLI, job, worker or webhook and writes a structured bug report with repro steps, without changing any code.

    136k GitHub stars~11k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Playwright Regression Testing

    fugazi/test-automation-skills-agents

    Govern Playwright TypeScript regression suites across many tests.

    247 GitHub stars~1.4k tokensUpdated 6 days ago
    Testing & QAAuto-check passed
  • Playwright Screen Recording

    liaohch3/claude-tap

    Records headless Playwright sessions as .webm videos to show a bug fix working or to give pull request reviewers visual evidence.

    3.3k GitHub stars~714 tokensUpdated 17 days ago
    Testing & QAAuto-check passed
  • Repro

    vaadin/web-components

    Reproduce a Vaadin web component bug from a GitHub issue in vaadin/web-components.

    582 GitHub stars~1.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Whole-App Health Sweep

    reticlehq/reticle

    Sweeps a running web app by clicking every reachable control, then reports dead buttons, console errors, failed requests and mismatches between API data and the screen.

    1.2k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed

More from Skyvern-AI/skyvern

  • Skyvern Browser Automation

    Skyvern-AI/skyvern

    Picks the right Skyvern CLI command for a web task, from quick yes/no checks to reusable multi-page workflows, instead of falling back to plain page fetching.

    23k GitHub starsUsed in 1 repo~2.9k tokens
    Auto-check passed
  • Skyvern Version Bump

    Skyvern-AI/skyvern

    Walks through a Skyvern open-source release bump: update the version, rebuild the Python and TypeScript SDKs with Fern, commit, and open a pull request.

    23k GitHub stars~1k tokensUpdated today
    Auto-check: notes
  • Skyvern Browser Automation

    Skyvern-AI/skyvern

    Automates websites with Skyvern's AI browser agent to fill forms, extract data, download files, log in and run multi-step workflows through SDKs, REST, MCP or a CLI.

    23k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Smoke-tests a Skyvern deployment by checking the backend API, frontend rendering, browser session provisioning and workflow execution in sequence.

    23k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Diff-Driven Smoke Tests

    Skyvern-AI/skyvern

    Reads your git diff, writes a handful of happy-path browser smoke tests, runs them with Skyvern or Chrome DevTools MCP and posts screenshot evidence to the PR.

    23k GitHub stars~5.2k tokensUpdated today
    Auto-check passed

Works with

Questions about Diff-Driven QA

What does Diff-Driven QA do?

Reads your git diff, decides whether the change needs browser QA, API checks or repo tests, runs that validation and reports pass or fail with evidence. The /qa command starts from git diff, comparing against the last commit or the working tree, and reads every changed file that affects behavior: routes, components, visible text, forms, schemas, endpoints, validators and the tests that changed. It then classifies the diff as frontend or browser, backend API, backend-internal or mixed.

When should I use Diff-Driven QA?

Diff-Driven QA fits situations like: validating code changes before opening a pull request; checking a UI change in the browser against the local dev server; running targeted API requests for changed route handlers; deciding which kind of validation a mixed frontend and backend diff needs.

How do I install Diff-Driven QA in Claude Code?

Run `npx skills add Skyvern-AI/skyvern --skill qa -a claude-code`. Or copy the skill folder (skyvern/cli/skills/qa in Skyvern-AI/skyvern) into .claude/skills/qa in your project. Claude Code loads it when a task matches its description.

How do I install Diff-Driven QA in Codex?

Run `npx skills add Skyvern-AI/skyvern --skill qa -a codex`. Or copy the skill folder (skyvern/cli/skills/qa in Skyvern-AI/skyvern) into .agents/skills/qa in your project. Codex loads it when a task matches its description.

Can I use Diff-Driven QA in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Skyvern-AI/skyvern --skill qa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa, .gemini/skills/qa, .github/skills/qa and .opencode/skills/qa in your project.

What does Diff-Driven QA need to run?

Going by SKILL.md and its folder, Diff-Driven QA needs the command-line tools its instructions call (git, curl, ngrok, npm and go). Our summary lists: A git repository with a diff to test; A locally runnable dev server or backend for the changed code.

Does Diff-Driven QA access the network?

SKILL.md contains no URLs. Its commands use git, curl and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Diff-Driven QA safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): mentions a paste, webhook or tunnelling service often used to send data out. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Diff-Driven QA use?

Diff-Driven QA is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Diff-Driven QA use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Diff-Driven QA?

Skills that share tags, products or a category with Diff-Driven QA: QA Test and Fix (garrytan/gstack, 136k stars), QA Report Only (garrytan/gstack, 136k stars), Playwright Regression Testing (fugazi/test-automation-skills-agents, 247 stars) and Playwright Screen Recording (liaohch3/claude-tap, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Diff-Driven QA?

Skyvern-AI (a GitHub organization) maintains it in Skyvern-AI/skyvern, which has 23,162 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 9, 2026.

Source: Skyvern-AI/skyvern on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.