Agent skill

AI Bug Triage

by petrkindlmann in petrkindlmann/qa-skills

Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation.

MITAuto-check passedTesting & QA

Install AI Bug Triage

skills CLI
$ npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install petrkindlmann/qa-skills ai-bug-triage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-bug-triage .claude/skills/ai-bug-triage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-bug-triage
GitHub stars
170
Token cost
~5.2k tokens
SKILL.md length
2,229 words
Files
4 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation.

  • Works in 12 steps: Normalize → Extract Stable Anchors → Hash Canonical Form → …
  • Failure analysis
  • SKILL.md covers Discovery Questions, Core Principles, The Pipeline and Severity/Priority Matrix, plus 6 more sections
  • Calls gh

What it does

AI Bug Triage is an agent skill from petrkindlmann/qa-skills. Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation. Normalizes CI logs, creates stable fingerprints, clusters near-duplicates, then uses LLM for severity classification and ticket writing. Includes bug reporting templates and severity/priority matrix. Use when: "bug triage," "classify bugs," "failure analysis," "auto-classify," "CI failures," "bug report," "defect template." Not for: runtime self-healing of one flaky locator — use test-reliability. Not for: designing…

Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/ci-failure-analysis.md`, `references/classification-taxonomy.md` and `references/pipeline-prompts-and-integration.md`).

It sits in Testing & QA, covering Issue triage, CI/CD and QA and bug reports. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.

When your agent uses it

  • Failure analysis
  • Defect template. Not for: runtime self-healing of one flaky locator — use test-reliability

Example prompts

  • “bug triage,”
  • “classify bugs,”
  • “failure analysis,”
  • “/ai-bug-triage”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Normalize
  2. Extract Stable Anchors
  3. Hash Canonical Form
  4. Cluster Near-Duplicates
  5. LLM Classify
  6. LLM Generate Ticket
  7. Human Approval
  8. Using LLM for Deduplication
  9. Auto-Closing Without Review
  10. Over-Classifying Severity
  11. Ignoring Environment Failures
  12. No Feedback Loop

What it can do on your machine

Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Bug Triage loads about 5.2k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 169 tokens; SKILL.md has 2,229 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~169
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,229 words, ~5,241 tokens.

Download SKILL.mdSave it as .claude/skills/ai-bug-triage/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
ai-bug-triage
description
Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation. Normalizes CI logs, creates stable fingerprints, clusters near-duplicates, then uses LLM for severity classification and ticket writing. Includes bug reporting templates and severity/priority matrix. Use when: "bug triage," "classify bugs," "failure analysis," "auto-classify," "CI failures," "bug report," "defect template." Not for: runtime self-healing of one flaky locator — use test-reliability. Not for: designing new tests from production telemetry — use observability-driven-testing. Related: qa-metrics, qa-dashboard, ci-cd-integration, qa-project-context.
license
MIT
metadata.author
kindlmann
metadata.version
2.0
metadata.category
ai-qa
<objective>
A hybrid pipeline for bug classification, deduplication, and ticket generation. Deterministic fingerprinting handles deduplication (what LLMs are bad at); LLM handles explanation, severity assessment, and ticket writing (what LLMs are good at).

Key reframe: The LLM is best at explaining and routing, not deduplication. Teach agents to DESIGN the pipeline, not BE the pipeline. </objective>


Discovery Questions

Check .agents/qa-project-context.md first — it carries tech stack, component mapping, and known flaky areas that improve classification accuracy. Use it and skip anything already answered there. Then clarify:

  1. What is the failure source?

    • CI pipeline logs (GitHub Actions, GitLab CI, Jenkins, CircleCI)
    • Test framework output (Playwright, Jest, pytest, Vitest)
    • Production error monitoring (Sentry, Datadog, Bugsnag)
    • Manual bug reports from QA or users
  2. What is the ticket destination?

    • Jira, Linear, GitHub Issues, Azure DevOps, Shortcut
    • What fields are required? (component, severity, priority, labels)
    • What workflows exist? (triage board, auto-assignment rules)
  3. What is the deduplication scope?

    • Same test run? Same sprint? Same release? All time?
    • Do you already have fingerprinting? What is the current duplicate rate?
  4. What approval workflow is needed?

    • Auto-create tickets with human review?
    • Suggest tickets for human approval before creation?
    • Auto-close duplicates? (dangerous -- require approval)
  5. What historical data exists?

    • Past bug reports with resolution data?
    • Flaky test history? Known environment issues?
    • Component ownership mapping?

Core Principles

  1. Deterministic first, LLM second. Use stable, reproducible fingerprinting for deduplication and clustering. Use LLM only for tasks requiring understanding: severity classification, root cause hypothesis, and human-readable ticket writing.

  2. Normalize before comparing. Raw CI logs are full of timestamps, port numbers, process IDs, and random suffixes that make identical failures look different. Strip all noise before fingerprinting.

  3. Fingerprints are anchored to stable elements. Exception type, top stack frames, test name, error message template, and URL pattern are stable. Timestamps, request IDs, and ephemeral ports are not.

  4. Human approval before destructive actions. Auto-closing a ticket as duplicate or auto-merging reports requires human confirmation. False deduplication wastes more time than manual triage.

  5. Classification drives routing. The value of triage is not the label itself but the routing decision it enables: which team, what priority, what SLA.

  6. Track triage accuracy. Measure how often auto-classification matches human judgment. Below 85% accuracy, the pipeline needs tuning.


The Pipeline

CI Log / Error Report
  │
  ▼
Step 1: NORMALIZE
  Strip timestamps, process IDs, ports, random suffixes, ANSI codes
  │
  ▼
Step 2: EXTRACT STABLE ANCHORS
  Exception type, top N stack frames, test name, error message template, URL pattern
  │
  ▼
Step 3: HASH CANONICAL FORM
  Deterministic fingerprint from ordered anchors
  │
  ▼
Step 4: CLUSTER NEAR-DUPLICATES
  Similarity scoring for non-identical but related failures
  │
  ▼
Step 5: LLM CLASSIFY
  Severity, component, suspected root cause, failure category
  │
  ▼
Step 6: LLM GENERATE TICKET
  Title, description, repro steps, evidence, suggested assignee
  │
  ▼
Step 7: HUMAN APPROVAL
  Review before create/close/merge
Step 1: Normalize

Strip noise that makes identical failures look different.

Normalization rules (apply in order):

1. Strip ANSI color codes:        \x1b\[[0-9;]*m → ""
2. Strip timestamps:              \d{4}-\d{2}-\d{2}[T ]\d{2}:\d{2}:\d{2}[.\d]*Z? → "<TIMESTAMP>"
3. Strip UUIDs:                   [0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12} → "<UUID>"
4. Strip process IDs:             pid[=: ]\d+ → "pid=<PID>"
5. Strip port numbers:            :\d{4,5}(?=[\s/,)\]]|$) → ":<PORT>"
6. Strip temp file paths:         /tmp/[^\s]+ → "<TMPPATH>"
7. Strip memory addresses:        0x[0-9a-f]{8,16} → "<ADDR>"
8. Strip random suffixes:         [-_][a-z0-9]{6,8}(?=\.) → "<RAND>"
9. Strip request IDs:             (?:request[_-]?id|trace[_-]?id|correlation[_-]?id)[=: ]["']?[a-zA-Z0-9-]+ → "<REQ_ID>"
10. Collapse whitespace:          \s+ → " "

Example:

Before: 2025-03-22T14:32:01.456Z [pid=42891] Error: Connection refused at 127.0.0.1:54321
        request_id=abc-123-def-456
After:  <TIMESTAMP> [pid=<PID>] Error: Connection refused at 127.0.0.1:<PORT>
        <REQ_ID>

Rule 5 strips the port but not the literal loopback IP — 127.0.0.1 stays in the fingerprint. That's fine for same-host failures, but two runners that bind different hosts (e.g. 127.0.0.1 vs 0.0.0.0) will split into separate fingerprints. If you run heterogeneous hosts, add a rule to normalize bind addresses too.

Step 2: Extract Stable Anchors

From the normalized log, extract elements that identify the failure regardless of environment or timing.

Anchor types (in priority order):

AnchorExampleStability
Exception typeTypeError, AssertionError, HTTP 500Very high
Error message templateCannot read property 'X' of undefinedHigh
Top 3 stack framesat processOrder (order.ts:142)High
Test namecheckout.spec.ts > completes paymentVery high
URL patternPOST /api/ordersHigh
HTTP status code500, 429, 503Very high
Exit codeexit code 1, SIGKILLHigh
Assertion diffExpected: 200, Received: 500Medium

Extraction rules:

  • Keep function names but strip line numbers (they change with edits)
  • Keep URL paths but strip query parameters and IDs in paths (/api/orders/<ID>)
  • Keep error message structure but replace dynamic values with placeholders
  • Keep test file and test name exactly as-is
Step 3: Hash Canonical Form

Create a deterministic fingerprint from the extracted anchors.

Algorithm:

1. Sort anchors alphabetically by type
2. Concatenate: exception_type + "|" + message_template + "|" + top_frames + "|" + test_name
3. SHA-256 hash the concatenated string
4. Take first 16 hex characters as fingerprint

Fingerprint properties:

  • Same failure always produces same fingerprint (deterministic)
  • Different failures produce different fingerprints (collision-resistant)
  • Minor log format changes do not change fingerprint (stable)
  • Fingerprint is short enough for Jira labels and GitHub tags

Example:

Anchors:
  exception_type: "TypeError"
  message_template: "Cannot read property 'vendorId' of undefined"
  top_frames: "processOrder|groupByVendor|checkout"
  test_name: "checkout.spec.ts > multi-vendor checkout"

Canonical: "TypeError|Cannot read property 'vendorId' of undefined|processOrder|groupByVendor|checkout|checkout.spec.ts > multi-vendor checkout"
Fingerprint: a3f8b2c1e9d04567
Step 4: Cluster Near-Duplicates

Exact fingerprint matching catches identical failures. Similarity scoring catches related failures that differ slightly (same root cause, different manifestation).

Similarity dimensions:

DimensionWeightMatch Criteria
Exception type0.30Exact match
Error message0.25Levenshtein distance < 20% of message length
Stack frames0.25Jaccard similarity of top 5 frames > 0.6
Component/file0.10Same directory or module
Test name0.10Same describe block or test file

Clustering threshold: similarity score > 0.75 = likely duplicate, suggest merge.

Human review required for:

  • Scores between 0.60 and 0.75 (ambiguous)
  • First occurrence of a new fingerprint (no history to compare)
  • Failures in components with known intermittent issues
Step 5: LLM Classify

After deterministic fingerprinting and clustering, use the LLM to classify the failure. The prompt feeds in exception, message, top 5 stack frames, test name, and CI context, and asks for five fields:

  1. Failure category — test bug | application bug | environment issue | flaky test | build failure
  2. Severity — critical | major | minor | trivial (see the severity matrix below)
  3. Component — inferred from stack trace and file paths
  4. Suspected root cause — 1-2 sentence hypothesis
  5. Confidence — high | medium | low; when low, the LLM states what extra information would resolve it

Route low-confidence classifications to human review rather than auto-acting. See references/pipeline-prompts-and-integration.md for the full prompt text and references/classification-taxonomy.md for the bug-category, severity, and component-mapping definitions the prompt should reference.

Failure categories (see references/ci-failure-analysis.md for detail):

CategoryDescriptionTypical Action
Application bugThe app is brokenFile bug ticket, assign to owning team
Test bugThe test is wrongFix the test, no app change needed
Environment issueCI infra / network / service downRetry, notify infra team
Flaky testIntermittent, non-deterministicQuarantine, investigate root cause
Build failureCompilation, dependency, configFix build, usually blocking
Step 6: LLM Generate Ticket

Once classified, use the LLM to generate a human-quality bug ticket. The prompt takes the classification plus the normalized error, a log excerpt, and related cluster fingerprints, and produces:

  • Title — concise, searchable, includes component name (under 80 chars)
  • Description — what happened, in plain language (never raw logs)
  • Steps to reproduce — derived from the test name and log context
  • Evidence — relevant log lines, assertion diffs, screenshots if available
  • Suggested labels — [component, severity, failure-category, fingerprint]
  • Suggested assignee — based on component ownership, if known

The fingerprint belongs on the ticket (label and Fingerprint field) so future dedup can match. See references/pipeline-prompts-and-integration.md for the full prompt and the bug report template.

Step 7: Human Approval

No automated action without review. The pipeline suggests; humans decide.

Approval decisions:

  • Create ticket — New failure, clear root cause, assign to team
  • Merge into existing — Duplicate of known issue, add evidence to existing ticket
  • Quarantine test — Flaky test, not an app bug, quarantine and schedule investigation
  • Retry and monitor — Environment issue, retry CI, alert if persists
  • Dismiss — Known issue already fixed in pending deploy, or test bug with obvious fix

Severity/Priority Matrix

Severity measures impact. Priority measures urgency. They are independent dimensions.

Severity Definitions
SeverityDefinitionExamples
CriticalSystem unusable, data loss, security breach, no workaroundPayment processing fails, user data exposed, app crashes on launch
MajorCore feature broken, degraded experience, workaround existsSearch returns wrong results, checkout requires page reload, form data lost on back-button
MinorNon-core feature affected, cosmetic with functional impactSorting does not persist, tooltip clipped on mobile, secondary action fails
TrivialCosmetic only, no functional impactTypo in label, 1px alignment, inconsistent capitalization
Priority Definitions
PriorityDefinitionSLA (example)
P0Fix immediately, blocks release or productionSame day
P1Fix this sprint, significant user impactThis sprint
P2Fix next sprint, moderate impactNext sprint
P3Fix when convenient, low impactBacklog
Severity x Priority Decision Guide
CriticalMajorMinorTrivial
Affects all usersP0P0P1P2
Affects segment (>10%)P0P1P2P3
Affects few users (<10%)P1P1P2P3
Edge case onlyP1P2P3P3

Bug Report Template

Use the same template for any bug report, whether auto-generated or human-written. It carries the defect heading, severity/priority/component/environment/fingerprint/reporter metadata, then Description, Steps to Reproduce, Expected/Actual Behavior, Evidence, Frequency, Suggested Root Cause, and Related Issues. See references/pipeline-prompts-and-integration.md for the full copy-paste Markdown template.


Show full SKILL.md (919 more words)Show less

Deduplication Patterns

PatternDetectionAction
Exact duplicateSame fingerprintMerge into existing ticket, add evidence
Near-duplicateSame cluster (similarity > 0.75)Link tickets, suggest merge for human review
Same root cause, different symptomSame exception type + overlapping frames in different testsCreate parent ticket linking symptom tickets
Regression of fixed bugFingerprint matches closed ticketReopen ticket, flag as regression, increase priority
Flaky recurrenceSame fingerprint intermittently across CI runsTag as flaky, quarantine if rate > 10%

CI Failure Analysis

See references/ci-failure-analysis.md for comprehensive patterns. Key decision: consistent failure = test bug or app bug; intermittent failure = flaky test or environment; multiple failures at once = environment or shared component; build failure = code or dependency issue.


Integration Patterns

The pipeline output is tracker-agnostic: Step 6 produces title, description, labels, severity, and component that map to any tracker's fields. See references/pipeline-prompts-and-integration.md for the gh issue create / fingerprint-dedup commands, the GitHub Actions "triage on failure" workflow, and notes on Jira/Linear/Azure DevOps REST/GraphQL integration.

Buy vs Build

Before implementing the full pipeline, check whether a hosted platform already covers the work you'd be doing. Several tools now ship AI-driven test triage that overlaps directly with Steps 4-6.

PlatformCoversNotes
Trunk Flaky TestsFingerprinting, clustering, severity routing, native PR comments + webhooksDedicated Agents feature for triage; documented Quarantining workflow — the closest off-the-shelf analog to this skill's pipeline
CloudBees Smart TestsFingerprinting, ML-based prioritization, Test Impact AnalysisFormerly Launchable — agents searching old docs may find the old name
Datadog Test OptimizationFlaky Test Management (Auto Retries, Early Flake Detection, Failed Test Replay), Test Impact AnalysisBits AI Dev Agent now auto-generates fix PRs and Flaky Test Policies auto-quarantine-then-disable after 30 days; pairs with Datadog APM if you're already on Datadog
SealightsQuality intelligence and test-impact gatingEnterprise; strongest in regulated industries

Use the in-skill pipeline when (a) you need on-prem or air-gapped deployment, (b) your tracker integration is exotic, or (c) you want an explicit AI-prompt audit trail for compliance. Otherwise, buying is usually cheaper than rebuilding fingerprinting + clustering.

Model selection cost note

Use Sonnet 4.6 / Haiku 4.5 for classification (Step 5) — it's cheap and capable enough. Escalate to Opus 4.8 only when the cluster is novel, the failure is ambiguous, or the suggested root cause has low confidence. Burning Opus on every triage is wasteful.


Anti-Patterns

1. Using LLM for Deduplication

LLMs are non-deterministic. The same two errors compared twice may get different similarity scores. Use deterministic fingerprinting for deduplication; use LLM only for explaining and classifying.

2. Auto-Closing Without Review

Automatically closing a ticket as "duplicate" based on fingerprint matching can merge distinct issues. Always require human confirmation for close/merge actions.

3. Over-Classifying Severity

If everything is "critical," nothing is. Follow the severity matrix strictly. A cosmetic typo is trivial even if it annoys someone.

4. Ignoring Environment Failures

Labeling all failures as "app bug" when many are CI infrastructure issues (Docker OOM, network timeout, disk full). Classify environment issues separately -- they need different remediation.

5. No Feedback Loop

Building the pipeline once and never measuring accuracy. Track: auto-classification accuracy, false duplicate rate, ticket quality ratings from developers.

6. Raw Logs in Tickets

Pasting 500 lines of raw CI output into a bug ticket. Normalize, extract relevant lines, and present the 5-10 lines that matter.

7. Fingerprinting Without Normalization

Hashing raw log lines produces unstable fingerprints that change every run — the same failure gets a different fingerprint each time, so dedup never fires. Normalization is mandatory: always normalize before fingerprinting (Step 1), never fingerprint raw logs.

8. No Component Ownership Mapping

Classification without routing is useless. Maintain a component-to-team mapping so that classified bugs reach the right people.


Verification

The fingerprinter is the load-bearing piece — prove it before trusting any dedup. Run the 5 stability assertions from references/ci-failure-analysis.md against your implementation:

  1. Same error, different timestamps → same fingerprint.
  2. Same error, different PIDs and ports → same fingerprint.
  3. Same error, different line numbers (code edited) → same fingerprint.
  4. Different errors (e.g. TypeError vs RangeError) → different fingerprints.
  5. Same exception type, different message property ('name' vs 'email') → different fingerprints.

Tests 1-3 must collapse to one hash; tests 4-5 must produce distinct hashes. A failure here means normalization (Step 1) is leaking noise into the hash or stripping a stable anchor — fix that before tuning clustering. The clustering weights (0.30 / 0.25 / 0.25 / 0.10 / 0.10) must sum to 1.00.


Done When

  • Every failure in failures.json has non-null severity, component, and category fields.
  • The 5 fingerprint stability assertions above pass (tests 1-3 same hash, tests 4-5 distinct).
  • Duplicates are merged or linked, each pointing to the canonical ticket's fingerprint label.
  • Every P0/P1 ticket has an assignee set.
  • Auto-classification accuracy is measured and recorded (target > 85%; see qa-metrics).

  • qa-metrics — Track triage accuracy, duplicate rates, mean time to classification, and defect escape rates.
  • ci-cd-integration — Pipeline configuration for running triage on test failures, parallel execution, and reporting.
  • test-reliability — Runtime per-test healing and quarantine for a single flaky locator. Triage classifies failures; test-reliability fixes one test live.
  • observability-driven-testing — Goes the other direction: turns production telemetry into new test designs. Use it when prod errors should spawn tests, not tickets.
  • qa-project-context — Project context that improves classification accuracy: component map, known issues, ownership.
  • ai-test-generation — Generate regression tests from triaged bug reports.

Reference Files (in references/)

  • pipeline-prompts-and-integration.md — LLM classification + ticket-generation prompts (Steps 5-6), the full bug report Markdown template, and tracker integration code (GitHub Issues, CI workflow, Jira/Linear notes).
  • classification-taxonomy.md — Bug categories, severity definitions, component mapping rules, and root cause categories.
  • ci-failure-analysis.md — CI log parsing patterns, failure category decision tree, fingerprinting algorithm detail.

© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/ai-bug-triage of petrkindlmann/qa-skills.

  • SKILL.md
  • references/ci-failure-analysis.md
  • references/classification-taxonomy.md
  • references/pipeline-prompts-and-integration.md

Open the folder on GitHubat commit b3bb61b

Compare with similar skills

AI Bug Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Bug Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Bug Triage this skillpetrkindlmann/qa-skills170—~5.2kAutomated safety check: PassMIT
Megatron-LM CI Failure TriageNVIDIA/Megatron-LM18k—~1.6kAutomated safety check: PassApache-2.0
GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb6.7k—~4.4kAutomated safety check: PassApache-2.0
MAUI CI Investigatordotnet/maui23k—~2kAutomated safety check: PassMIT
Diagnose a Red Rundifferent-ai/openwork24k—~779Automated safety check: PassCustom licence
CI Triagecanton-network/splice117—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.

    18k GitHub stars~1.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.

    6.7k GitHub stars~4.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Adds dotnet/maui-specific context for investigating failing PR checks and broken nightly builds: pipelines, Helix logs, binlogs and merge-readiness verdicts.

    23k GitHub stars~2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Diagnose a Red Run

    different-ai/openwork

    Classifies a failing test, typecheck or CI job before any code changes, by recording the failure and running a clean control to show whether it was already broken.

    24k GitHub stars~779 tokensUpdated today
    Testing & QAAuto-check passed
  • CI Triage

    canton-network/splice

    Triage a failed splice GitHub Actions job (cn-test-failures ref) into a reproducible evidence packet - fetch job log and artifact, isolate the flagged lines, check the known flake families for…

    117 GitHub stars~1.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Instructor QA

    cognesy/instructor-php

    Run multi-dimensional quality assurance for InstructorPHP. An agent skill from cognesy/instructor-php.

    328 GitHub stars~1.3k tokensUpdated 3 days ago
    Testing & QAAuto-check passed

More from petrkindlmann/qa-skills

All 45 skills in this repo
  • Accessibility Testing

    petrkindlmann/qa-skills

    Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).

    170 GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    170 GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • AI Test Generation

    petrkindlmann/qa-skills

    Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.

    170 GitHub stars~4.8k tokensUpdated 4 mo ago
    Auto-check passed
  • API Testing

    petrkindlmann/qa-skills

    Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.

    170 GitHub stars~2.7k tokensUpdated 4 mo ago
    Auto-check passed
  • CI CD Integration

    petrkindlmann/qa-skills

    Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.

    170 GitHub stars~4.8k tokensUpdated 4 mo ago
    Auto-check passed
  • Compliance Testing

    petrkindlmann/qa-skills

    Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…

    170 GitHub stars~4.6k tokensUpdated 4 mo ago
    Auto-check passed

Questions about AI Bug Triage

What does AI Bug Triage do?

Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation. AI Bug Triage is an agent skill from petrkindlmann/qa-skills. Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation.

When should I use AI Bug Triage?

AI Bug Triage fits situations like: failure analysis; defect template. Not for: runtime self-healing of one flaky locator — use test-reliability.

How do I install AI Bug Triage in Claude Code?

Run `npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a claude-code`. Or copy the skill folder (skills/ai-bug-triage in petrkindlmann/qa-skills) into .claude/skills/ai-bug-triage in your project. Claude Code loads it when a task matches its description.

How do I install AI Bug Triage in Codex?

Run `npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a codex`. Or copy the skill folder (skills/ai-bug-triage in petrkindlmann/qa-skills) into .agents/skills/ai-bug-triage in your project. Codex loads it when a task matches its description.

Can I use AI Bug Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill ai-bug-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-bug-triage, .gemini/skills/ai-bug-triage, .github/skills/ai-bug-triage and .opencode/skills/ai-bug-triage in your project.

What does AI Bug Triage need to run?

Going by SKILL.md and its folder, AI Bug Triage needs the command-line tools its instructions call (gh).

Does AI Bug Triage access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is AI Bug Triage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Bug Triage use?

AI Bug Triage is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Bug Triage use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.4k tokens, read only when the agent opens those files.

What are the alternatives to AI Bug Triage?

Skills that share tags, products or a category with AI Bug Triage: Megatron-LM CI Failure Triage (NVIDIA/Megatron-LM, 18k stars), GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars), MAUI CI Investigator (dotnet/maui, 23k stars) and Diagnose a Red Run (different-ai/openwork, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Bug Triage?

petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 170 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.

Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.