Megatron-LM CI Failure Triage
NVIDIA/Megatron-LM
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
Build and visualize QA dashboards and reports with Allure Report, Grafana, and ReportPortal.
$ npx skills add petrkindlmann/qa-skills --skill qa-dashboard -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install petrkindlmann/qa-skills qa-dashboard --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qa-dashboard .claude/skills/qa-dashboard && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "qa-dashboard" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-dashboard into .claude/skills/qa-dashboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-dashboard", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-dashboardType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add petrkindlmann/qa-skills --skill qa-dashboard -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install petrkindlmann/qa-skills qa-dashboard --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/qa-dashboard .agents/skills/qa-dashboard && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "qa-dashboard" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-dashboard into .agents/skills/qa-dashboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-dashboard", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill qa-dashboard -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install petrkindlmann/qa-skills qa-dashboard --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/qa-dashboard .cursor/skills/qa-dashboard && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "qa-dashboard" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-dashboard into .cursor/skills/qa-dashboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-dashboard", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/petrkindlmann/qa-skills.git --path skills/qa-dashboard--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add petrkindlmann/qa-skills --skill qa-dashboard -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install petrkindlmann/qa-skills qa-dashboard --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/qa-dashboard .gemini/skills/qa-dashboard && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "qa-dashboard" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-dashboard into .gemini/skills/qa-dashboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-dashboard", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install petrkindlmann/qa-skills qa-dashboardInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add petrkindlmann/qa-skills --skill qa-dashboard -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/qa-dashboard .github/skills/qa-dashboard && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "qa-dashboard" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-dashboard into .github/skills/qa-dashboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-dashboard", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill qa-dashboard -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install petrkindlmann/qa-skills qa-dashboard --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/qa-dashboard .opencode/skills/qa-dashboard && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "qa-dashboard" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/qa-dashboard into .opencode/skills/qa-dashboard/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "qa-dashboard", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
qa-dashboardBuild and visualize QA dashboards and reports with Allure Report, Grafana, and ReportPortal.
QA Dashboard is an agent skill from petrkindlmann/qa-skills. Build and visualize QA dashboards and reports with Allure Report, Grafana, and ReportPortal. Covers test execution visualization, stakeholder-facing quality reports, trend/flakiness panels, release-readiness gates, alerting, and CI integration for automated report generation. Use when: "test dashboard," "Allure," "test report," "quality dashboard," "Grafana," "ReportPortal," "test results visualization." Not for: defining which KPIs to measure or how to interpret them — use qa-metrics (this skill builds the…
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/allure.md` and `references/grafana.md`).
It sits in DevOps & Cloud, covering Monitoring and alerting, Failing and flaky tests and Issue triage. It works with Grafana. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxbrewcurldockerjqnpmFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
raw.githubusercontent.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
RP_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
QA Dashboard loads about 4.7k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 158 tokens; SKILL.md has 2,194 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,194 words, ~4,742 tokens.
.claude/skills/qa-dashboard/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.<objective>
Build dashboards that drive decisions, not dashboards that display data. The failure
mode this prevents: a wall of panels showing "5 failures" with no link to the failures,
no target, and no trend — a notification dressed up as a tool that everyone stops opening
by week three. This skill delivers audience-specific QA reporting — rich HTML reports
(Allure), real-time trend panels with regression alerts (Grafana/InfluxDB), self-hosted
AI-assisted aggregation (ReportPortal), and automated stakeholder summaries — each panel
mapped to a question someone actually asks and every red indicator drilling down to the
failing test.
</objective>
| Situation | Go to |
|---|---|
| Already on Grafana for infra metrics | Grafana Dashboards — add test metrics alongside prod |
| Test runner has a hosted dashboard (Cypress/Playwright) | SaaS-Native Dashboards — least plumbing |
| Need rich HTML report, no infra to run | Allure Report — generate in CI, publish artifact |
| Need self-hosted aggregation across frameworks + AI triage | ReportPortal |
| Need a single view across multiple runners | Grafana or Allure (cross-runner aggregation) |
| Need a weekly summary / release verdict for stakeholders | Stakeholder Reports |
Check .agents/qa-project-context.md first — if it exists, use it and skip anything answered there. Then:
1. Dashboards answer questions, they do not display numbers. Every panel must map to a question someone actually asks. "What is the flakiness rate?" is a question. "Total test count" is trivia.
2. Different audiences need different views. A developer debugging a CI failure needs stack traces, screenshots, and traces. A VP needs a single number: "Are we ready to release?" Do not force both through the same dashboard.
3. Real-time for CI, trends for leadership. CI dashboards update on every pipeline run. Leadership dashboards aggregate weekly or per-sprint. Mixing cadences confuses both audiences.
4. Drill-down to action. Every red indicator must link to the specific failing test, the specific flaky test, or the specific coverage gap. A dashboard that shows "5 failures" but does not link to the failures is a notification, not a tool.
5. Automate report generation. Reports that require manual effort (running scripts, copying data, formatting slides) will not survive the first busy sprint. Generate reports from CI pipelines automatically.
Allure generates rich HTML reports from test results with history, categories, and retries built in. It works with Playwright, Jest, Vitest, pytest, and most frameworks.
Allure 2 vs Allure 3. Two paths, and they are easy to mix up because the framework adapters (allure-playwright, allure-vitest, allure-jest) emit the same Allure 2 result files; only the reader differs:
allure-commandline (2.42.1, Jun 2026). brew install allure installs this, not v3. Commands: allure generate / allure open / allure serve; categories via a categories.json dropped into allure-results/. Stable, frozen, dependency-bump-only now.allure npm package + allurerc.mjs (3.9.0, May 2026). TypeScript rewrite: plugins, single-file config, real-time allure watch, project-wide quality gates, multi-environment reports, and Allure Service for server-side history. Commands: npx allure run / npx allure generate / npx allure watch; categories move into allurerc.mjs (the dropped-in categories.json is a v2 concept).Choose Allure 3 for new projects. If you follow only the brew install allure / allure generate commands you are on the v2 path — that is fine and fully supported; just know which one you are running.
Minimal Playwright reporter wiring — register the adapter as reporter: ["allure-playwright", { outputFolder, environmentInfo }] so every run drops results into allure-results/ with the environment captured:
// playwright.config.ts
reporter: [
["list"],
["allure-playwright", {
outputFolder: "allure-results",
environmentInfo: {
Environment: process.env.TEST_ENV ?? "local",
BaseURL: process.env.BASE_URL ?? "http://localhost:3000",
},
}],
],Add per-test metadata (allure.severity / feature / story / tag) to drive grouping, define failure
categories to split product bugs from infra breakage, and preserve history/ across CI runs so trends
exist at all. Without history an Allure report is a single snapshot — no trends, no intermittent-failure
detection. See references/allure.md for the full Playwright/Vitest configs, the v2 categories.json,
the v3 allurerc.mjs + allure run runnable path, and the GitHub Actions history-preservation steps.
Grafana gives real-time dashboards with alerting. Best when the team already runs Grafana for infrastructure and wants test metrics next to production metrics.
Data pipeline: a post-test CI step parses results (JUnit XML, coverage JSON, timing) and pushes points to a time-series DB (InfluxDB or a Prometheus pushgateway); Grafana queries those. The push script writes two measurements — test_execution (one point per test, tagged by suite/test_name/status/branch/run_id) and test_run_summary (one point per run with pass_rate, total, failed, avg_duration_ms). Tag every point with branch and run_id so panels can filter to main and link back to a specific run.
The script runs under if: always(), so wrap the write loop in try/flush/close — a throw mid-loop otherwise loses every buffered point. See references/grafana.md for the full push-test-metrics.ts (with the flush guard) and the GitHub Actions step.
| Panel | Question | Query shape |
|---|---|---|
| Pass Rate Trend (time series) | Is quality improving? | SELECT mean("pass_rate") FROM "test_run_summary" WHERE "branch"='main' GROUP BY time(1d), thresholds at 95% (yellow) / 99% (green) |
| Release Readiness (stat) | Is main ready to release? | SELECT last("pass_rate") FROM "test_run_summary" WHERE "branch"='main', red <95 / yellow 95–99 / green ≥99 — pair with a coverage stat ≥80 |
| Flakiness Top 10 (table) | Which tests waste the most time? | SELECT "test_name", count("retries") AS retry_count FROM "test_execution" WHERE "retries">0 AND time>now()-14d GROUP BY "test_name" ORDER BY retry_count DESC LIMIT 10 |
| CI Duration Trend | Is the pipeline getting slower? | avg_duration_ms over time with a target line at 600s |
Full queries plus Coverage Trend and Duration Distribution panels are in references/grafana.md.
Provision alert rules as YAML under provisioning/alerting/ (Grafana 11+). The one the Done When
requires: main pass rate drops more than 2 percentage points in a single day. Build it from two
queries (mean pass_rate over the last 1d vs the day before) feeding a math expression
$yesterday - $today > 2, then route via a Slack contact point (incoming-webhook URL). The full
provisioned rule + contact point is in references/grafana.md. Also alert on: pass rate below 95%
(10m window), CI duration above 15 min, coverage drop >2% in a week. Dashboards are for
investigation; alerts are for detection — a dashboard no one opens catches nothing.
Self-hosted test reporting platform with ML-powered failure analysis, cross-framework aggregation, and real-time dashboards.
# Pin to a tagged release (current 26.0.x line: 26.0.3). The `master` branch may not
# match the supported 26.x line.
curl -LO https://raw.githubusercontent.com/reportportal/reportportal/26.0.3/docker-compose.yml
docker compose up -d
# Access at http://localhost:8080 (default: superadmin/erebus)
# ML auto-analysis lives in the `service-auto-analyzer` container — confirm it is running.As of 26.0.3 ReportPortal also ingests agentic test results (launches carry an AGENTIC vs
AUTOMATION execution-type badge) — relevant if part of your suite runs through Claude Code or another agent.
// playwright.config.ts
reporter: [
["list"],
["@reportportal/agent-js-playwright", {
apiKey: process.env.RP_API_KEY,
endpoint: process.env.RP_ENDPOINT ?? "http://localhost:8080/api/v1",
project: "my-project",
launch: `E2E Tests - ${process.env.CI ? "CI" : "local"}`,
attributes: [
{ key: "branch", value: process.env.GITHUB_HEAD_REF ?? "local" },
{ key: "build", value: process.env.GITHUB_RUN_ID ?? "dev" },
],
}],
],Install with npm i -D @reportportal/agent-js-playwright.
| Feature | What it does |
|---|---|
| Auto-analysis | ML failure classification: product bug, test bug, system issue, or to-investigate |
| Defect type mapping | Custom defect categories with sub-types for your project |
| Flaky test detection | Tests that flip pass/fail across launches |
| Merge launches | Combine sharded CI runs into one unified view |
| Quality gates | Pass/fail criteria per launch (max failures, min pass rate) |
| Comparison | Side-by-side of two launches to spot regressions |
Quality gates are queryable after a run (GET /api/v1/$PROJECT/launch/$LAUNCH_ID/quality-gate) — use the
status as a CI gate and fail the pipeline if it is not PASSED.
If self-hosting feels heavy, Allure TestOps (26.2.x line, 2026) is the SaaS path: Allure 3 quality gates, named environments, global attachments, and Allure 3-style flaky detection (flags a test once it shows ≥3 status transitions across its last 10 runs). Its MCP server is in public beta (26.1.1), letting AI agents query launches and quality gates directly — relevant when your QA workflow runs through Claude Code / Cursor.
If your test runner has a first-class hosted dashboard, prefer it over Allure/Grafana for that runner's native data — less plumbing, more retention, built-in PR comments. Cross-pollinate with Allure/Grafana only for cross-runner aggregation.
| Platform | Test runner | Native data + PR comments |
|---|---|---|
| Cypress Cloud | Cypress | Test replay, parallelization, flake detection; AI add-on (Auto Heal, Bug Triage) |
| Currents.dev | Cypress, Playwright | OSS-friendly Cypress Cloud alternative; lower price point |
Playwright HTML + --reporter=blob | Playwright | Free, self-hosted; combine shards with merge-reports |
| Datadog Test Optimization | Any (CI-side) | Flaky Test Management (now with Bits AI auto-fix), TIA, native APM |
| Allure TestOps | Any | Allure 3 quality gates, named environments, MCP server beta |
Combining sharded Playwright runs (free, native). Have each shard emit a blob report, then merge into one HTML report — the no-cost answer to "combine sharded CI runs":
# each shard: npx playwright test --reporter=blob (uploads blob-report/ as an artifact)
npx playwright merge-reports --reporter html ./blob-reportsUse Allure or Grafana when you need one dashboard across multiple runners, or when a SaaS option's pricing/data-residency does not fit. Otherwise the SaaS-native dashboard is usually the cheapest path to PR-level signal.
Weekly QA Summary — automate via a scheduled CI job. Include: pass rate + trend, new vs fixed failures, top 5 flaky tests, coverage delta, avg CI duration. Classify health: STABLE (>= 98%), NEEDS ATTENTION (>= 95%), CRITICAL (< 95%). Post to Slack automatically.
Release Quality Report — generate before each release. Gate on: E2E pass rate >= 99%, unit pass rate 100%, branch coverage >= 80%, zero critical bugs, major bugs <= 2, and the Core Web Vitals budget. Output a READY / NOT READY verdict with a per-gate pass/fail breakdown.
Core Web Vitals "good" thresholds for the perf gate (current as of 2026): LCP < 2500ms, INP < 200ms, CLS < 0.1. Gate on INP, not FID — INP replaced FID as a Core Web Vital on 2024-03-12 and FID was fully retired on 2024-09-09.
A practical set covering the most common questions teams ask.
| Panel | Question It Answers | Data Source | Audience |
|---|---|---|---|
| Pass/Fail Trend | Is quality improving or degrading? | CI test results over time | Everyone |
| Flakiness Top 10 | Which tests waste the most time? | Tests with retries in last 14 days | Developers, QA |
| Coverage Heatmap | Where are we blind? | Coverage by module/directory | Developers |
| Defect Escape Trend | Are bugs reaching production? | Incidents tagged as test escapes | QA leads, Leadership |
| CI Duration | Is the pipeline getting slower? | Pipeline duration over time | DevOps, Developers |
| Test Velocity | Tests proportional to features? | New tests added per sprint | QA leads |
| Failure Categories | Product bugs or test infra? | Categorized failure reasons | QA leads |
| Release Readiness | Can we ship? | Composite score from all gates | Leadership |
Dashboard with 30 panels. No one reads it. Start with 5–6 panels that answer the most urgent questions; add panels only when someone asks one the dashboard cannot answer.
Metrics without context. "Pass rate: 97%" means nothing without "target: 99%" and "last week: 98.5%." Every metric needs a target and a trend to be actionable.
Manual report generation. If the weekly summary needs someone to SSH in, run queries, and paste into slides, it stops happening by week 3. Automate it into CI.
Same dashboard for developers and leadership. Developers need failure details, stack traces, and repro steps; leadership needs one traffic light. Build separate views.
Reporting test counts as progress. "We added 200 tests" says nothing about quality. Report critical-path coverage, defect escape rate, and mean time to detect regressions instead.
No alerting on regressions. A dashboard no one checks is useless. Alert on pass-rate drops, coverage decreases, and CI-duration increases. Dashboards are for investigation; alerts are for detection.
Allure without history. A single Allure report is a snapshot — no trends, no intermittent-failure detection, no measure of improvement. Always preserve history/ across CI runs (or use Allure 3 / Allure Service for server-side history).
Mixing the Allure 2 and 3 code paths. Following the v3 callout in prose but copying brew install allure + allure generate + a dropped-in categories.json lands you on v2 with v2 categories. Pick one path and use its commands end to end.
Prove the report/dashboard actually renders before calling it done. Smallest checks first:
# Allure: a report builds and the trend widget shows >1 run (history preserved)
npx allure generate allure-results --clean -o allure-report && npx allure open allure-report
# -> Overview page loads; the "Trend" widget shows more than one run.
# Grafana push: the summary point landed in InfluxDB
influx query 'from(bucket:"test-results") |> range(start:-1d) |> filter(fn:(r)=> r._measurement=="test_run_summary")'
# -> returns at least one row with a pass_rate field for branch=main.
# Grafana alert: the provisioned rule loaded
curl -s http://localhost:3000/api/v1/provisioning/alert-rules -u admin:admin | jq '.[].title'
# -> includes "Main pass rate dropped >2pp in a day".references/)categories.json, the v3 allurerc.mjs + allure run runnable path, and the CI history-preservation workflow.push-test-metrics.ts script (with flush guard), all panel queries, and the provisioned ">2pp pass-rate drop" alert rule + Slack contact point.© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/qa-dashboard of petrkindlmann/qa-skills.
Open the folder on GitHubat commit b3bb61b
QA Dashboard next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| QA Dashboard this skillpetrkindlmann/qa-skills | 168 | — | ~4.7k | Automated safety check: Pass | MIT | |
| Megatron-LM CI Failure TriageNVIDIA/Megatron-LM | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Admingrafana/skills | 281 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Cloud Infra Supply Chainzhaji2333/CkSKILLS | 114 | — | ~688 | Automated safety check: Warn | MIT | |
| Monitoring Agents Mdhardisgroupcom/sfdx-hardis | 402 | — | ~3.7k | Automated safety check: Notes | AGPL-3.0 | |
| Monitoring APIsjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT |
NVIDIA/Megatron-LM
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
grafana/skills
Manage Grafana Cloud accounts — organizations, stacks, RBAC roles and assignments, SSO/SAML/OAuth/GitHub auth, service accounts for CI/CD, user invites, team membership, and API-driven provisioning.
zhaji2333/CkSKILLS
当目标涉及云资产(对象存储/云元数据/Serverless)、容器/K8s、运维面板(宝塔/Grafana/Zabbix/Jenkins/GitLab/Nacos等)、消息队列/缓存中间件、CI/CD流水线、第三方回调集成、依赖组件CVE、信息泄露配置时调用。负责未授权访问、弱口令、云配置错误、供应链漏洞与敏感信息挖掘。
hardisgroupcom/sfdx-hardis
The AGENTS.md file that sf hardis:org:monitor:backup writes at the root of every monitoring repository, so a coding agent can answer questions about the org, its CI/CD deployments and its Grafana…
jeremylongshore/tons-of-skills-marketplace
Build real-time API monitoring dashboards with metrics, alerts, and health checks.
grafana/skills
Reduces JavaScript dependency footprint with pnpm while preserving lockfile, workspace layout, and dependency range style.
petrkindlmann/qa-skills
Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
petrkindlmann/qa-skills
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.
petrkindlmann/qa-skills
Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.
petrkindlmann/qa-skills
Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
Works with
Categories
Build and visualize QA dashboards and reports with Allure Report, Grafana, and ReportPortal. QA Dashboard is an agent skill from petrkindlmann/qa-skills. Build and visualize QA dashboards and reports with Allure Report, Grafana, and ReportPortal.
QA Dashboard fits situations like: : test dashboard; quality dashboard; test results visualization. Not for: defining which KPIs to measure; how to interpret them — use qa-metrics (this skill builds the panels.
Run `npx skills add petrkindlmann/qa-skills --skill qa-dashboard -a claude-code`. Or copy the skill folder (skills/qa-dashboard in petrkindlmann/qa-skills) into .claude/skills/qa-dashboard in your project. Claude Code loads it when a task matches its description.
Run `npx skills add petrkindlmann/qa-skills --skill qa-dashboard -a codex`. Or copy the skill folder (skills/qa-dashboard in petrkindlmann/qa-skills) into .agents/skills/qa-dashboard in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill qa-dashboard -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa-dashboard, .gemini/skills/qa-dashboard, .github/skills/qa-dashboard and .opencode/skills/qa-dashboard in your project.
Going by SKILL.md and its folder, QA Dashboard needs the command-line tools its instructions call (npx, brew, curl, docker, jq and npm) and credentials named RP_API_KEY. Our summary lists: Node.js; Docker; A credential in RP_API_KEY.
SKILL.md names 1 domain. In commands or code: raw.githubusercontent.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
QA Dashboard is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with QA Dashboard: Megatron-LM CI Failure Triage (NVIDIA/Megatron-LM, 18k stars), Admin (grafana/skills, 281 stars), Cloud Infra Supply Chain (zhaji2333/CkSKILLS, 114 stars) and Monitoring Agents Md (hardisgroupcom/sfdx-hardis, 402 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 168 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.
Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.