Write Tests
grafana/synthetic-monitoring-app
Write Jest integration and unit tests for the Grafana Synthetic Monitoring app using React Testing Library, MSW, and src/test helpers.
Define, track, and act on QA metrics: test coverage percentage, flakiness rate, defect escape rate, MTTR, test execution time trends, automation ROI, quality gates, and SLAs for test suites.
$ npx skills add petrkindlmann/qa-skills --skill qa-metrics -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install petrkindlmann/qa-skills qa-metrics --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qa-metrics .claude/skills/qa-metrics && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "qa-metrics" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-metrics into .claude/skills/qa-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-metrics", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-metricsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add petrkindlmann/qa-skills --skill qa-metrics -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install petrkindlmann/qa-skills qa-metrics --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/qa-metrics .agents/skills/qa-metrics && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "qa-metrics" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-metrics into .agents/skills/qa-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-metrics", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill qa-metrics -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install petrkindlmann/qa-skills qa-metrics --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/qa-metrics .cursor/skills/qa-metrics && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "qa-metrics" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-metrics into .cursor/skills/qa-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-metrics", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/petrkindlmann/qa-skills.git --path skills/qa-metrics--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add petrkindlmann/qa-skills --skill qa-metrics -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install petrkindlmann/qa-skills qa-metrics --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/qa-metrics .gemini/skills/qa-metrics && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "qa-metrics" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-metrics into .gemini/skills/qa-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-metrics", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install petrkindlmann/qa-skills qa-metricsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add petrkindlmann/qa-skills --skill qa-metrics -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/qa-metrics .github/skills/qa-metrics && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "qa-metrics" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-metrics into .github/skills/qa-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-metrics", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill qa-metrics -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install petrkindlmann/qa-skills qa-metrics --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/qa-metrics .opencode/skills/qa-metrics && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "qa-metrics" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-metrics into .opencode/skills/qa-metrics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-metrics", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
qa-metricsDefine, track, and act on QA metrics: test coverage percentage, flakiness rate, defect escape rate, MTTR, test execution time trends, automation ROI, quality gates, and SLAs for test suites.
QA Metrics is an agent skill from petrkindlmann/qa-skills. Define, track, and act on QA metrics: test coverage percentage, flakiness rate, defect escape rate, MTTR, test execution time trends, automation ROI, quality gates, and SLAs for test suites. Includes metric formulas, realistic targets by company stage, and the action to take when each metric goes red. Use when: "QA metrics," "test metrics," "quality KPIs," "test health," "flakiness rate," "defect escape rate." Not for: building the dashboard UI (Allure/Grafana) — use qa-dashboard; measuring coverage gaps and…
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/dashboards.md` and `references/rollout.md`).
It sits in Testing & QA, covering Test coverage, Failing and flaky tests and CI/CD. It works with Grafana. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
dora.devFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
QA Metrics loads about 5.3k tokens when it runs, and up to ~6.7k if it reads all its reference files. Until then it costs about 166 tokens; SKILL.md has 2,712 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,712 words, ~5,274 tokens.
.claude/skills/qa-metrics/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.<objective>
A suite with 4,000 tests where 800 are disabled and 200 are flaky looks healthy on a slide and lies in production. This skill defines the metrics that actually change behavior — each one with a formula, a target, an owner, and a concrete action when it goes red — so the dashboard becomes a feedback loop instead of decoration.
</objective>
| Your situation | Metric to reach for | Section / reference |
|---|---|---|
| No metrics yet, where to start | code coverage, flakiness, defect escape | Core Principles + references/rollout.md |
| Bugs escaping despite high coverage | defect escape rate + mutation score | Coverage Metrics, Defect Metrics |
| Tests flaky, devs ignore CI | flakiness rate, pass rate | Test Health Metrics |
| Slow pipeline, stale feedback | suite duration, parallelism efficiency | Execution Metrics |
| Justifying automation spend | automation ROI | Process Metrics |
| Picking targets for our stage | company-stage table | Setting Realistic Targets |
| Leadership wants delivery + quality | DORA + escape rate | Engineering Quality Metrics (DORA) |
| Building the dashboard view | (go to qa-dashboard) | — |
Check .agents/qa-project-context.md first — if it exists, use it as the foundation and skip anything already answered there. Then:
A metric without an action plan is decoration. For every metric, define: what threshold triggers action, what the action is, and who takes it. If flakiness crosses 5%, the on-call engineer investigates the top 3 flaky tests that week. No ambiguity.
Defect escape rate and MTTR are lagging: you learn after users were hurt. Flakiness, coverage delta, and skipped-test count are leading: they predict escapes before they happen. Lagging metrics tell leadership whether quality is moving; leading metrics are where engineers spend their daily attention because they can still change the outcome.
A single number is nearly useless. Coverage at 72% means nothing; coverage trending 68% → 72% over three sprints tells a story. Display metrics as time series and evaluate direction, not absolute position.
A target makes a metric actionable; an owner makes it accountable. Without both it becomes background noise. Set targets from your team's maturity (see the targets table) and assign owners who can actually move the number.
"We have 4,000 tests" sounds impressive until 800 are disabled and 200 flaky. Count what matters: tests that run, pass reliably, and catch real bugs. The ultimate metric is whether users hit bugs that hurt the business — work backward from there. Defect escape rate connects directly to user experience; lines of test code connects to nothing.
Each metric: definition, formula, recommended target, why it matters, and the action when it goes red.
How much of our system is verified by automated tests? (leading)
The percentage of code exercised by automated tests (line, branch, or statement level).
Line coverage = (lines executed by tests / total executable lines) × 100
Branch coverage = (branches executed by tests / total branches) × 100Targets: Line: 70-85% for app code. Branch: 60-75%. Critical paths (payments, auth): 90%+.
Why it matters: Coverage identifies blind spots — code never exercised by tests is where bugs hide undetected.
When it goes red: Coverage drops on PR: block merge or flag. Low in critical module: create targeted tasks. Plateaus: check for dead code vs. genuinely untested logic.
Warning: Coverage measures execution, not assertion quality. Pair it with mutation testing (StrykerJS v9.6+, mutmut) for a truer picture. Stryker's Vitest runner tracks recent Vitest releases — check the runner's peer-dependency before pinning a Vitest version. Use incremental: true in monorepo CI to keep mutation runs cheap. For deep coverage/mutation analysis, see coverage-analysis.
Requirement coverage = (features with at least one test / total features) × 100Targets: 100% for P0/P1 features, 80%+ for P2.
Why it matters: A feature can have zero tests even if surrounding code is well-covered. Use test tags (@feature:checkout, @story:PROJ-1234) for traceability.
Risk coverage = (high-risk areas with automated tests / total high-risk areas) × 100Targets: 95%+ for high-risk areas. Maintain a risk register and cross-reference against coverage data per module.
Can we trust our test suite? (leading)
The percentage of test runs producing inconsistent results without code changes.
Flakiness rate = (test runs with flaky results / total test runs) × 100Targets: Acceptable: <2%. Warning: 2-5%. Critical: >5%.
Why it matters: Flaky tests erode trust. Once developers think "probably just flaky," they stop paying attention to results at all. Flakiness is the single biggest threat to a suite's credibility.
When it goes red: Quarantine flaky tests immediately. Investigate the top 3 weekly — most flakiness comes from a small number of tests. Common causes: timing/race conditions, shared state, external dependencies, order-dependent tests. Tests flaky for 30+ days should be deleted or rewritten.
Detection: Buildkite Test Analytics, Datadog Test Optimization (formerly Datadog CI Visibility — ships Flaky Test Management, Auto Test Retries, Early Flake Detection, Failed Test Replay, and Test Impact Analysis; its Bits AI Dev Agent can auto-open PRs to fix flaky tests), Trunk Flaky Tests, or a script comparing results across runs on the same commit.
Pass rate = (green CI runs / total CI runs) × 100 [rolling 7 days]Targets: Healthy: >95%. Warning: 90-95%. Broken: <90%.
The 7-day rolling average smooths daily noise and reveals the real trend. A consistently red build means developers ignore the pipeline.
Total tests marked skip, disabled, pending, xit, xdescribe or equivalent.
Target: Trend toward zero. Skipped tests older than 2 sprints: fix or delete. Add a CI step that fails if skipped count exceeds 5% of total tests.
Why it matters: Skipped tests are invisible coverage gaps. A suite with 500 passing and 150 skipped tests has a 150 / (500 + 150) = 23% gap that dashboards hide.
Wall-clock time from suite start to completion.
Targets: Unit: <5 min. Integration: <10 min. E2E: <15 min. Full pipeline: <30 min (the stage targets are the real budget; <30 min is the sum, not a separate looser bar).
Why it matters: Slow tests break the feedback loop. 45-minute results mean developers have already context-switched.
When it goes red: Profile slowest tests (10% often account for 50% of runtime). Increase parallelism. Move slow tests post-merge. Check for unnecessary setup/teardown. Plot duration over time and alert on step changes (>20% increase in a week).
Are we catching bugs before users do? (lagging)
The percentage of defects found in production relative to all defects found.
Defect escape rate = (defects found in production / total defects found) × 100Targets: Excellent: <5%. Acceptable: 5-10%. Needs work: 10-20%. Critical: >20%.
Worked example: A release surfaces 15 total defects, 2 of them in production → 2 / 15 × 100 = 13.3% escape rate, which lands in "needs work."
Why it matters: The single most important quality metric. It directly measures whether testing catches bugs before users do.
When it goes red: Classify escaped defects by layer (unit? integration? E2E? review?). Write a retrospective test for each. Identify if escapes cluster in specific areas — those need targeted investment.
How to track: Tag production bugs (escaped-defect label). Count escaped vs. pre-release defects at sprint retros.
MTTR = sum(resolution_time for each defect) / number of defectsTargets: P0: <4 hours. P1: <24 hours. P2: <1 sprint. P3: <2 sprints.
High MTTR often signals process bottlenecks (slow review, unclear ownership, complex deploys) rather than technical difficulty. Break into phases (triage, assign, fix, deploy) to find the bottleneck. (This is your QA defect resolution time — distinct from the DORA recovery metric below.)
Defect density = defects found / KLOC (or per feature shipped)Targets: Track your own baseline; industry benchmarks are 1-10 defects per KLOC. If one module has 5x the density of others, it needs refactoring or better coverage.
Healthy: P0 <5%, P1 10-15%, P2 40-50%, P3 30-40%. Visualize as a stacked bar over time. Heavy P0/P1 concentration means testing misses critical issues.
Is our CI pipeline fast, reliable, and cost-effective? (leading)
Total wall-clock time by stage. Targets: Lint: <2 min. Unit: <5 min. Integration: <10 min. E2E: <15 min. Full: <30 min (= sum of stages above, reconciled with the suite-duration budget).
When it goes red: Optimize the slowest stage first. Split fast checks (every push) from slow checks (PR merge). Profile setup time vs. execution time. Parallelize sequential stages.
CI cost per run = (compute minutes × cost per minute) + fixed costs50 builds/day at $0.50 each = $750/month. Optimize by caching dependencies, using spot instances, right-sizing runners, and skipping unchanged suites.
Parallelism efficiency = (total sequential time / (wall-clock time × workers)) × 100Target: >80%. Example: 4 workers finishing in 3, 3, 3, and 12 minutes → wall-clock 12, sequential total 21, so efficiency = 21 / (12 × 4) = 44% — well under target because three workers idled 9 minutes each. Fix by splitting by estimated duration (not file count), breaking up slow test files, and using dynamic splitting (Playwright sharding, Jest --shard).
Is our QA process improving over time? (mix of leading and lagging)
Automation rate = (automated test cases / total test cases) × 100Targets: Regression: 90%+. Smoke: 100%. Exploratory: 0% (by definition). Overall: 70-85%.
Manual testing does not scale. Automation compounds — once written, a test runs thousands of times. Automate the most frequently executed scenarios first.
New automated tests added per sprint (net new, excluding refactors).
Target: At least 3-5 automated tests per user story shipped. A sprint with 20 features and 0 new tests signals a growing coverage gap.
Manual cost = (manual time per cycle × cycles per year) × hourly rate
Automation cost = (write time + annual maintenance) × hourly rate
ROI = (manual cost - automation cost) / automation cost × 100Example: Manual regression: 8 hrs/release × 26 releases × $75/hr = $15,600/yr. Automation: 120 hrs to write + 40 hrs/yr maintenance × $75/hr = $12,000 year 1, $3,000/yr after. Year 1 ROI: (15,600 - 12,000) / 12,000 = 30%. Year 2 ROI: (15,600 - 3,000) / 3,000 = 420%. Use this to justify investment to stakeholders.
DORA metrics are the standard vocabulary for leadership delivery dashboards. They pair with defect escape rate — DORA tracks delivery throughput; QA metrics track delivery quality.
| Metric | What it measures | Benchmark (top-15%) |
|---|---|---|
| Lead Time for Changes | Commit → production | < 1 day |
| Deployment Frequency | How often you ship | Multiple per day |
| Failed Deployment Recovery Time (formerly MTTR) | Time to restore service after a change-induced failure | < 1 hour |
| Change Failure Rate | % of deploys that cause incidents | ~5% (older "high performer" bar was <15%) |
| Rework Rate | % of deploys that are unplanned fixes for a prior bad deploy | Low and falling |
DORA 2025 formalized Rework Rate as the fifth metric and regrouped the set: the first three above are throughput (recovery time moved here because fast teams just ship the fix), Change Failure Rate + Rework Rate are instability. Reliability (availability, latency, error budget against your SLOs) is tracked alongside as a separate dimension — it is where QA/escape framing meets SRE, since error rate and availability are the user-facing tail of escaped defects. DORA 2025 also retired the named Elite/High/Medium/Low tiers in favor of percentile distributions and seven team archetypes — read the column as a percentile benchmark, not a tier you "are." The Failed Deployment Recovery Time rename (2023/2024 reports) separates change-induced failures from external outages; it is a delivery metric distinct from your QA defect MTTR above.
Source: https://dora.dev/research/. Tools that surface DORA from Git/CI data: Sleuth, Faros, LinearB, Jellyfish, Swarmia.
TIA selects which tests to run based on which code changed (using coverage data). It trades test breadth for CI cost. Track:
Hosted: Datadog Test Optimization (TIA), CloudBees Smart Tests, NCrunch (in-IDE). Self-built: derive from coverage data + git diff. See coverage-analysis.
Targets should match your team's maturity. Chasing enterprise metrics at a seed startup wastes effort.
| Metric | Startup (seed-Series A) | Growth (Series B-C) | Enterprise (public/large) |
|---|---|---|---|
| Unit test coverage | 60% | 75% | 85% |
| Branch coverage | 45% | 60% | 75% |
| E2E coverage (critical paths) | Top 5 flows | Top 15 flows | All P0/P1 flows |
| Flakiness rate | <5% | <3% | <1% |
| Pass rate (7-day) | >90% | >95% | >98% |
| Defect escape rate | <20% | <10% | <5% |
| MTTR (P0) | <8 hours | <4 hours | <2 hours |
| CI pipeline duration | <20 min | <15 min | <10 min |
| Automation rate | 50% | 75% | 90% |
| Metrics tracked | 3-5 core | 8-10 with dashboards | Full suite with alerting |
Progression path:
For the full phased rollout (week-by-week), the three dashboard layouts, and the data-source extraction table, see references/rollout.md and references/dashboards.md.
references/rollout.md).The whole point is a working feedback loop — prove the data actually flows before declaring done:
coverage/coverage-final.json, lcov.info) — not just an HTML report a human reads.escaped-defect (or equivalent) label and confirm at least the tagging convention is in place..skip and confirm CI fails it — a gate that never fires proves nothing.references/)© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/qa-metrics of petrkindlmann/qa-skills.
Open the folder on GitHubat commit b3bb61b
QA Metrics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| QA Metrics this skillpetrkindlmann/qa-skills | 170 | — | ~5.3k | Automated safety check: Pass | MIT | |
| Write Testsgrafana/synthetic-monitoring-app | 171 | — | ~1.2k | Automated safety check: Pass | AGPL-3.0 | |
| Matlab Run Testsmatlab/matlab-agentic-toolkit | 1.1k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Dd Unblock PRDataDog/pup | 1k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Azsdk Common Pipeline AnalysisAzure/azure-sdk-tools | 134 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Unblock PRdatadog-labs/agent-skills | 177 | — | ~2.7k | Automated safety check: Pass | MIT |
grafana/synthetic-monitoring-app
Write Jest integration and unit tests for the Grafana Synthetic Monitoring app using React Testing Library, MSW, and src/test helpers.
matlab/matlab-agentic-toolkit
Run MATLAB test suites and analyze results, including test filtering, parallel execution, result diagnostics, code coverage collection, gap analysis, and CI/CD with buildtool.
DataDog/pup
Load when investigating a failing PR CI pipeline or checking PR health.
Azure/azure-sdk-tools
Analyze Azure SDK CI/CD pipeline failures into a structured diagnosis, and define the required output format.
datadog-labs/agent-skills
Load when investigating a failing PR CI pipeline or checking PR health.
HoangNguyen0403/agent-skills-standard
Audit test coverage health, gaps, and QE debt for Jira stories or epics.
petrkindlmann/qa-skills
Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
petrkindlmann/qa-skills
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.
petrkindlmann/qa-skills
Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.
petrkindlmann/qa-skills
Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
Works with
Categories
Define, track, and act on QA metrics: test coverage percentage, flakiness rate, defect escape rate, MTTR, test execution time trends, automation ROI, quality gates, and SLAs for test suites. QA Metrics is an agent skill from petrkindlmann/qa-skills. Define, track, and act on QA metrics: test coverage percentage, flakiness rate, defect escape rate, MTTR, test execution time trends, automation ROI, quality gates, and SLAs for test suites.
QA Metrics fits situations like: defect escape rate. Not for: building the dashboard UI (Allure/Grafana) — use qa-dashboard; measuring coverage gaps and mutation score — use coverage-analysis.
Run `npx skills add petrkindlmann/qa-skills --skill qa-metrics -a claude-code`. Or copy the skill folder (skills/qa-metrics in petrkindlmann/qa-skills) into .claude/skills/qa-metrics in your project. Claude Code loads it when a task matches its description.
Run `npx skills add petrkindlmann/qa-skills --skill qa-metrics -a codex`. Or copy the skill folder (skills/qa-metrics in petrkindlmann/qa-skills) into .agents/skills/qa-metrics in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill qa-metrics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa-metrics, .gemini/skills/qa-metrics, .github/skills/qa-metrics and .opencode/skills/qa-metrics in your project.
SKILL.md names no scripts, command-line tools or credentials: QA Metrics is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: dora.dev. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
QA Metrics is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with QA Metrics: Write Tests (grafana/synthetic-monitoring-app, 171 stars), Matlab Run Tests (matlab/matlab-agentic-toolkit, 1.1k stars), Dd Unblock PR (DataDog/pup, 1k stars) and Azsdk Common Pipeline Analysis (Azure/azure-sdk-tools, 134 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 170 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.
Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.