Megatron-LM CI Failure Triage
NVIDIA/Megatron-LM
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation.
$ npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install petrkindlmann/qa-skills ai-bug-triage --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-bug-triage .claude/skills/ai-bug-triage && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ai-bug-triage" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/ai-bug-triage into .claude/skills/ai-bug-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-bug-triage", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/petrkindlmann/qa-skills/tree/main/skills/ai-bug-triageType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install petrkindlmann/qa-skills ai-bug-triage --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ai-bug-triage .agents/skills/ai-bug-triage && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ai-bug-triage" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/ai-bug-triage into .agents/skills/ai-bug-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-bug-triage", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install petrkindlmann/qa-skills ai-bug-triage --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ai-bug-triage .cursor/skills/ai-bug-triage && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ai-bug-triage" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/ai-bug-triage into .cursor/skills/ai-bug-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-bug-triage", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/petrkindlmann/qa-skills.git --path skills/ai-bug-triage--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install petrkindlmann/qa-skills ai-bug-triage --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ai-bug-triage .gemini/skills/ai-bug-triage && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ai-bug-triage" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/ai-bug-triage into .gemini/skills/ai-bug-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-bug-triage", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install petrkindlmann/qa-skills ai-bug-triageInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ai-bug-triage .github/skills/ai-bug-triage && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ai-bug-triage" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/ai-bug-triage into .github/skills/ai-bug-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-bug-triage", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install petrkindlmann/qa-skills ai-bug-triage --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ai-bug-triage .opencode/skills/ai-bug-triage && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ai-bug-triage" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/ai-bug-triage into .opencode/skills/ai-bug-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-bug-triage", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ai-bug-triageHybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation.
AI Bug Triage is an agent skill from petrkindlmann/qa-skills. Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation. Normalizes CI logs, creates stable fingerprints, clusters near-duplicates, then uses LLM for severity classification and ticket writing. Includes bug reporting templates and severity/priority matrix. Use when: "bug triage," "classify bugs," "failure analysis," "auto-classify," "CI failures," "bug report," "defect template." Not for: runtime self-healing of one flaky locator — use test-reliability. Not for: designing…
Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/ci-failure-analysis.md`, `references/classification-taxonomy.md` and `references/pipeline-prompts-and-integration.md`).
It sits in Testing & QA, covering Issue triage, CI/CD and QA and bug reports. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AI Bug Triage loads about 5.2k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 169 tokens; SKILL.md has 2,229 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,229 words, ~5,241 tokens.
.claude/skills/ai-bug-triage/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.<objective>
A hybrid pipeline for bug classification, deduplication, and ticket generation. Deterministic fingerprinting handles deduplication (what LLMs are bad at); LLM handles explanation, severity assessment, and ticket writing (what LLMs are good at).
Key reframe: The LLM is best at explaining and routing, not deduplication. Teach agents to DESIGN the pipeline, not BE the pipeline.
</objective>
Check .agents/qa-project-context.md first — it carries tech stack, component mapping, and known flaky areas that improve classification accuracy. Use it and skip anything already answered there. Then clarify:
What is the failure source?
What is the ticket destination?
What is the deduplication scope?
What approval workflow is needed?
What historical data exists?
Deterministic first, LLM second. Use stable, reproducible fingerprinting for deduplication and clustering. Use LLM only for tasks requiring understanding: severity classification, root cause hypothesis, and human-readable ticket writing.
Normalize before comparing. Raw CI logs are full of timestamps, port numbers, process IDs, and random suffixes that make identical failures look different. Strip all noise before fingerprinting.
Fingerprints are anchored to stable elements. Exception type, top stack frames, test name, error message template, and URL pattern are stable. Timestamps, request IDs, and ephemeral ports are not.
Human approval before destructive actions. Auto-closing a ticket as duplicate or auto-merging reports requires human confirmation. False deduplication wastes more time than manual triage.
Classification drives routing. The value of triage is not the label itself but the routing decision it enables: which team, what priority, what SLA.
Track triage accuracy. Measure how often auto-classification matches human judgment. Below 85% accuracy, the pipeline needs tuning.
CI Log / Error Report
│
▼
Step 1: NORMALIZE
Strip timestamps, process IDs, ports, random suffixes, ANSI codes
│
▼
Step 2: EXTRACT STABLE ANCHORS
Exception type, top N stack frames, test name, error message template, URL pattern
│
▼
Step 3: HASH CANONICAL FORM
Deterministic fingerprint from ordered anchors
│
▼
Step 4: CLUSTER NEAR-DUPLICATES
Similarity scoring for non-identical but related failures
│
▼
Step 5: LLM CLASSIFY
Severity, component, suspected root cause, failure category
│
▼
Step 6: LLM GENERATE TICKET
Title, description, repro steps, evidence, suggested assignee
│
▼
Step 7: HUMAN APPROVAL
Review before create/close/mergeStrip noise that makes identical failures look different.
Normalization rules (apply in order):
1. Strip ANSI color codes: \x1b\[[0-9;]*m → ""
2. Strip timestamps: \d{4}-\d{2}-\d{2}[T ]\d{2}:\d{2}:\d{2}[.\d]*Z? → "<TIMESTAMP>"
3. Strip UUIDs: [0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12} → "<UUID>"
4. Strip process IDs: pid[=: ]\d+ → "pid=<PID>"
5. Strip port numbers: :\d{4,5}(?=[\s/,)\]]|$) → ":<PORT>"
6. Strip temp file paths: /tmp/[^\s]+ → "<TMPPATH>"
7. Strip memory addresses: 0x[0-9a-f]{8,16} → "<ADDR>"
8. Strip random suffixes: [-_][a-z0-9]{6,8}(?=\.) → "<RAND>"
9. Strip request IDs: (?:request[_-]?id|trace[_-]?id|correlation[_-]?id)[=: ]["']?[a-zA-Z0-9-]+ → "<REQ_ID>"
10. Collapse whitespace: \s+ → " "Example:
Before: 2025-03-22T14:32:01.456Z [pid=42891] Error: Connection refused at 127.0.0.1:54321
request_id=abc-123-def-456
After: <TIMESTAMP> [pid=<PID>] Error: Connection refused at 127.0.0.1:<PORT>
<REQ_ID>Rule 5 strips the port but not the literal loopback IP — 127.0.0.1 stays in the fingerprint. That's fine for same-host failures, but two runners that bind different hosts (e.g. 127.0.0.1 vs 0.0.0.0) will split into separate fingerprints. If you run heterogeneous hosts, add a rule to normalize bind addresses too.
From the normalized log, extract elements that identify the failure regardless of environment or timing.
Anchor types (in priority order):
| Anchor | Example | Stability |
|---|---|---|
| Exception type | TypeError, AssertionError, HTTP 500 | Very high |
| Error message template | Cannot read property 'X' of undefined | High |
| Top 3 stack frames | at processOrder (order.ts:142) | High |
| Test name | checkout.spec.ts > completes payment | Very high |
| URL pattern | POST /api/orders | High |
| HTTP status code | 500, 429, 503 | Very high |
| Exit code | exit code 1, SIGKILL | High |
| Assertion diff | Expected: 200, Received: 500 | Medium |
Extraction rules:
/api/orders/<ID>)Create a deterministic fingerprint from the extracted anchors.
Algorithm:
1. Sort anchors alphabetically by type
2. Concatenate: exception_type + "|" + message_template + "|" + top_frames + "|" + test_name
3. SHA-256 hash the concatenated string
4. Take first 16 hex characters as fingerprintFingerprint properties:
Example:
Anchors:
exception_type: "TypeError"
message_template: "Cannot read property 'vendorId' of undefined"
top_frames: "processOrder|groupByVendor|checkout"
test_name: "checkout.spec.ts > multi-vendor checkout"
Canonical: "TypeError|Cannot read property 'vendorId' of undefined|processOrder|groupByVendor|checkout|checkout.spec.ts > multi-vendor checkout"
Fingerprint: a3f8b2c1e9d04567Exact fingerprint matching catches identical failures. Similarity scoring catches related failures that differ slightly (same root cause, different manifestation).
Similarity dimensions:
| Dimension | Weight | Match Criteria |
|---|---|---|
| Exception type | 0.30 | Exact match |
| Error message | 0.25 | Levenshtein distance < 20% of message length |
| Stack frames | 0.25 | Jaccard similarity of top 5 frames > 0.6 |
| Component/file | 0.10 | Same directory or module |
| Test name | 0.10 | Same describe block or test file |
Clustering threshold: similarity score > 0.75 = likely duplicate, suggest merge.
Human review required for:
After deterministic fingerprinting and clustering, use the LLM to classify the failure. The prompt feeds in exception, message, top 5 stack frames, test name, and CI context, and asks for five fields:
test bug | application bug | environment issue | flaky test | build failurecritical | major | minor | trivial (see the severity matrix below)high | medium | low; when low, the LLM states what extra information would resolve itRoute low-confidence classifications to human review rather than auto-acting. See references/pipeline-prompts-and-integration.md for the full prompt text and references/classification-taxonomy.md for the bug-category, severity, and component-mapping definitions the prompt should reference.
Failure categories (see references/ci-failure-analysis.md for detail):
| Category | Description | Typical Action |
|---|---|---|
| Application bug | The app is broken | File bug ticket, assign to owning team |
| Test bug | The test is wrong | Fix the test, no app change needed |
| Environment issue | CI infra / network / service down | Retry, notify infra team |
| Flaky test | Intermittent, non-deterministic | Quarantine, investigate root cause |
| Build failure | Compilation, dependency, config | Fix build, usually blocking |
Once classified, use the LLM to generate a human-quality bug ticket. The prompt takes the classification plus the normalized error, a log excerpt, and related cluster fingerprints, and produces:
[component, severity, failure-category, fingerprint]The fingerprint belongs on the ticket (label and Fingerprint field) so future dedup can match. See references/pipeline-prompts-and-integration.md for the full prompt and the bug report template.
No automated action without review. The pipeline suggests; humans decide.
Approval decisions:
Severity measures impact. Priority measures urgency. They are independent dimensions.
| Severity | Definition | Examples |
|---|---|---|
| Critical | System unusable, data loss, security breach, no workaround | Payment processing fails, user data exposed, app crashes on launch |
| Major | Core feature broken, degraded experience, workaround exists | Search returns wrong results, checkout requires page reload, form data lost on back-button |
| Minor | Non-core feature affected, cosmetic with functional impact | Sorting does not persist, tooltip clipped on mobile, secondary action fails |
| Trivial | Cosmetic only, no functional impact | Typo in label, 1px alignment, inconsistent capitalization |
| Priority | Definition | SLA (example) |
|---|---|---|
| P0 | Fix immediately, blocks release or production | Same day |
| P1 | Fix this sprint, significant user impact | This sprint |
| P2 | Fix next sprint, moderate impact | Next sprint |
| P3 | Fix when convenient, low impact | Backlog |
| Critical | Major | Minor | Trivial | |
|---|---|---|---|---|
| Affects all users | P0 | P0 | P1 | P2 |
| Affects segment (>10%) | P0 | P1 | P2 | P3 |
| Affects few users (<10%) | P1 | P1 | P2 | P3 |
| Edge case only | P1 | P2 | P3 | P3 |
Use the same template for any bug report, whether auto-generated or human-written. It carries the defect heading, severity/priority/component/environment/fingerprint/reporter metadata, then Description, Steps to Reproduce, Expected/Actual Behavior, Evidence, Frequency, Suggested Root Cause, and Related Issues. See references/pipeline-prompts-and-integration.md for the full copy-paste Markdown template.
| Pattern | Detection | Action |
|---|---|---|
| Exact duplicate | Same fingerprint | Merge into existing ticket, add evidence |
| Near-duplicate | Same cluster (similarity > 0.75) | Link tickets, suggest merge for human review |
| Same root cause, different symptom | Same exception type + overlapping frames in different tests | Create parent ticket linking symptom tickets |
| Regression of fixed bug | Fingerprint matches closed ticket | Reopen ticket, flag as regression, increase priority |
| Flaky recurrence | Same fingerprint intermittently across CI runs | Tag as flaky, quarantine if rate > 10% |
See references/ci-failure-analysis.md for comprehensive patterns. Key decision: consistent failure = test bug or app bug; intermittent failure = flaky test or environment; multiple failures at once = environment or shared component; build failure = code or dependency issue.
The pipeline output is tracker-agnostic: Step 6 produces title, description, labels, severity, and component that map to any tracker's fields. See references/pipeline-prompts-and-integration.md for the gh issue create / fingerprint-dedup commands, the GitHub Actions "triage on failure" workflow, and notes on Jira/Linear/Azure DevOps REST/GraphQL integration.
Before implementing the full pipeline, check whether a hosted platform already covers the work you'd be doing. Several tools now ship AI-driven test triage that overlaps directly with Steps 4-6.
| Platform | Covers | Notes |
|---|---|---|
| Trunk Flaky Tests | Fingerprinting, clustering, severity routing, native PR comments + webhooks | Dedicated Agents feature for triage; documented Quarantining workflow — the closest off-the-shelf analog to this skill's pipeline |
| CloudBees Smart Tests | Fingerprinting, ML-based prioritization, Test Impact Analysis | Formerly Launchable — agents searching old docs may find the old name |
| Datadog Test Optimization | Flaky Test Management (Auto Retries, Early Flake Detection, Failed Test Replay), Test Impact Analysis | Bits AI Dev Agent now auto-generates fix PRs and Flaky Test Policies auto-quarantine-then-disable after 30 days; pairs with Datadog APM if you're already on Datadog |
| Sealights | Quality intelligence and test-impact gating | Enterprise; strongest in regulated industries |
Use the in-skill pipeline when (a) you need on-prem or air-gapped deployment, (b) your tracker integration is exotic, or (c) you want an explicit AI-prompt audit trail for compliance. Otherwise, buying is usually cheaper than rebuilding fingerprinting + clustering.
Use Sonnet 4.6 / Haiku 4.5 for classification (Step 5) — it's cheap and capable enough. Escalate to Opus 4.8 only when the cluster is novel, the failure is ambiguous, or the suggested root cause has low confidence. Burning Opus on every triage is wasteful.
LLMs are non-deterministic. The same two errors compared twice may get different similarity scores. Use deterministic fingerprinting for deduplication; use LLM only for explaining and classifying.
Automatically closing a ticket as "duplicate" based on fingerprint matching can merge distinct issues. Always require human confirmation for close/merge actions.
If everything is "critical," nothing is. Follow the severity matrix strictly. A cosmetic typo is trivial even if it annoys someone.
Labeling all failures as "app bug" when many are CI infrastructure issues (Docker OOM, network timeout, disk full). Classify environment issues separately -- they need different remediation.
Building the pipeline once and never measuring accuracy. Track: auto-classification accuracy, false duplicate rate, ticket quality ratings from developers.
Pasting 500 lines of raw CI output into a bug ticket. Normalize, extract relevant lines, and present the 5-10 lines that matter.
Hashing raw log lines produces unstable fingerprints that change every run — the same failure gets a different fingerprint each time, so dedup never fires. Normalization is mandatory: always normalize before fingerprinting (Step 1), never fingerprint raw logs.
Classification without routing is useless. Maintain a component-to-team mapping so that classified bugs reach the right people.
The fingerprinter is the load-bearing piece — prove it before trusting any dedup. Run the 5 stability assertions from references/ci-failure-analysis.md against your implementation:
TypeError vs RangeError) → different fingerprints.'name' vs 'email') → different fingerprints.Tests 1-3 must collapse to one hash; tests 4-5 must produce distinct hashes. A failure here means normalization (Step 1) is leaking noise into the hash or stripping a stable anchor — fix that before tuning clustering. The clustering weights (0.30 / 0.25 / 0.25 / 0.10 / 0.10) must sum to 1.00.
failures.json has non-null severity, component, and category fields.qa-metrics).qa-metrics — Track triage accuracy, duplicate rates, mean time to classification, and defect escape rates.ci-cd-integration — Pipeline configuration for running triage on test failures, parallel execution, and reporting.test-reliability — Runtime per-test healing and quarantine for a single flaky locator. Triage classifies failures; test-reliability fixes one test live.observability-driven-testing — Goes the other direction: turns production telemetry into new test designs. Use it when prod errors should spawn tests, not tickets.qa-project-context — Project context that improves classification accuracy: component map, known issues, ownership.ai-test-generation — Generate regression tests from triaged bug reports.references/)© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/ai-bug-triage of petrkindlmann/qa-skills.
Open the folder on GitHubat commit b3bb61b
AI Bug Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AI Bug Triage this skillpetrkindlmann/qa-skills | 170 | — | ~5.2k | Automated safety check: Pass | MIT | |
| Megatron-LM CI Failure TriageNVIDIA/Megatron-LM | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb | 6.7k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| MAUI CI Investigatordotnet/maui | 23k | — | ~2k | Automated safety check: Pass | MIT | |
| Diagnose a Red Rundifferent-ai/openwork | 24k | — | ~779 | Automated safety check: Pass | Custom licence | |
| CI Triagecanton-network/splice | 117 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 |
NVIDIA/Megatron-LM
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
GreptimeTeam/greptimedb
Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.
dotnet/maui
Adds dotnet/maui-specific context for investigating failing PR checks and broken nightly builds: pipelines, Helix logs, binlogs and merge-readiness verdicts.
different-ai/openwork
Classifies a failing test, typecheck or CI job before any code changes, by recording the failure and running a clean control to show whether it was already broken.
canton-network/splice
Triage a failed splice GitHub Actions job (cn-test-failures ref) into a reproducible evidence packet - fetch job log and artifact, isolate the flagged lines, check the known flake families for…
cognesy/instructor-php
Run multi-dimensional quality assurance for InstructorPHP. An agent skill from cognesy/instructor-php.
petrkindlmann/qa-skills
Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
petrkindlmann/qa-skills
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.
petrkindlmann/qa-skills
Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.
petrkindlmann/qa-skills
Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
Categories
Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation. AI Bug Triage is an agent skill from petrkindlmann/qa-skills. Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation.
AI Bug Triage fits situations like: failure analysis; defect template. Not for: runtime self-healing of one flaky locator — use test-reliability.
Run `npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a claude-code`. Or copy the skill folder (skills/ai-bug-triage in petrkindlmann/qa-skills) into .claude/skills/ai-bug-triage in your project. Claude Code loads it when a task matches its description.
Run `npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a codex`. Or copy the skill folder (skills/ai-bug-triage in petrkindlmann/qa-skills) into .agents/skills/ai-bug-triage in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-bug-triage, .gemini/skills/ai-bug-triage, .github/skills/ai-bug-triage and .opencode/skills/ai-bug-triage in your project.
Going by SKILL.md and its folder, AI Bug Triage needs the command-line tools its instructions call (gh).
SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
AI Bug Triage is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with AI Bug Triage: Megatron-LM CI Failure Triage (NVIDIA/Megatron-LM, 18k stars), GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars), MAUI CI Investigator (dotnet/maui, 23k stars) and Diagnose a Red Run (different-ai/openwork, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 170 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.
Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.