Agent skill

Triage Failure Reports

by WolframResearch in WolframResearch/Chatbook

Triage Chatbook's auto-generated bug-report issues on GitHub by grouping them on their Failure Data signature (the failing Wolfram Language function plus the confirmed expression/pattern), finding…

MITAuto-check passedTesting & QA

Install Triage Failure Reports

skills CLI
$ npx skills add WolframResearch/Chatbook --skill triage-failure-reports -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install WolframResearch/Chatbook triage-failure-reports --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/WolframResearch/Chatbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/triage-failure-reports .claude/skills/triage-failure-reports && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triage-failure-reports
GitHub stars
124
Token cost
~2.7k tokens
SKILL.md length
1,291 words
Files
2 (incl. scripts)
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Triage Chatbook's auto-generated bug-report issues on GitHub by grouping them on their Failure Data signature (the failing Wolfram Language function plus the confirmed expression/pattern), finding…

  • Works in 6 steps: Establish the reference signature → Search broadly for candidates → Confirm each candidate against the… → …
  • Wants to dedupe
  • SKILL.md covers How a Chatbook failure report…, Workflow, Worked example and Gotchas
  • Runs Python scripts from its folder; calls gh and python

What it does

Triage Failure Reports is an agent skill from WolframResearch/Chatbook. Triage Chatbook's auto-generated bug-report issues on GitHub by grouping them on their Failure Data signature (the failing Wolfram Language function plus the confirmed expression/pattern), finding the canonical issue or the PR that resolves each cluster, and closing the rest with an explanatory comment. Use this whenever the user wants to dedupe, triage, or clean up GitHub issues - for example "find duplicates of 1550", "are any of these crash reports the same bug?", "close the issues already fixed by PR 906"…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/extract_signatures.py`).

It sits in Testing & QA, covering QA and bug reports and Mobile testing and debugging. It works with GitHub, Wolfram Alpha and OpenAI. The repository describes itself as: Wolfram Notebooks + LLMs. The licence is MIT.

When your agent uses it

  • Wants to dedupe
  • Clean up GitHub issues - for example find duplicates of 1550
  • Are any of these crash reports the same bug?
  • Close the issues already fixed by PR 906

Example prompts

  • “find duplicates of 1550”
  • “are any of these crash reports the same bug?”
  • “close the issues already fixed by PR 906”
  • “/triage-failure-reports”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Establish the reference signature
  2. Search broadly for candidates
  3. Confirm each candidate against the signature
  4. Identify and verify the resolver
  5. Verify the timeline - by version, not just date
  6. Close with an explanatory comment

What it can do on your machine

Read from SKILL.md and the folder at commit a601484. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • gh
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triage Failure Reports loads about 2.7k tokens when it runs. Until then it costs about 237 tokens; SKILL.md has 1,291 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~237
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from WolframResearch/Chatbook at commit a601484, republished under its MIT licence (© WolframResearch). 1,291 words, ~2,731 tokens.

Download SKILL.mdSave it as .claude/skills/triage-failure-reports/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
triage-failure-reports
description
Triage Chatbook's auto-generated bug-report issues on GitHub by grouping them on their Failure Data signature (the failing Wolfram Language function plus the confirmed expression/pattern), finding the canonical issue or the PR that resolves each cluster, and closing the rest with an explanatory comment. Use this whenever the user wants to dedupe, triage, or clean up GitHub issues - for example "find duplicates of #1550", "are any of these crash reports the same bug?", "close the issues already fixed by PR #906", "triage the open failure reports", or after landing a fix, "which open issues does this close?". It distinguishes true duplicates from look-alikes that fail in a different code path, verifies a candidate fixing PR actually touches the relevant code, and checks each report's Version/ReleaseID against the fix so reports filed on outdated versions are closed correctly and possible regressions are flagged.

Triage duplicate issues by Failure Data signature

Chatbook posts machine-generated bug reports to GitHub. They pile up because the same underlying bug gets reported many times across different services, versions, and users. This skill turns that pile into a small number of signature-grouped clusters, ties each cluster to the issue or PR that resolves it, and closes the duplicates with a comment that tells the reporter where the real fix lives.

The hard part is not the closing - it's being sure two reports are actually the same bug before you act on someone else's issue. Most of this skill is about earning that confidence.

How a Chatbook failure report is structured

Each report has an auto-generated <details> block. The fields that matter:

  • Debug Data table - Version and ReleaseID (the git SHA the user was running). These drive the timeline check; do not skip them.
  • Settings table - Model, e.g. <|"Service" -> "OpenAI", "Name" -> Automatic|>. Tells you which service/model triggered it.
  • Failure Data - the heart of the signature:
    • "Evaluation" - the failing call, e.g. Wolfram`Chatbook`Common`resolveFullModelSpec[<|"Service" -> "OpenAI", "Name" -> Automatic|>]. The short function name (resolveFullModelSpec) is the primary key.
    • "Expression" - the value that broke a Confirm* check, e.g. Missing["NoModelList"]. (Unhandled-definition failures don't have this field; the bad value is a call argument instead - see the script's fallback.)
    • "Pattern" - what the value failed to match, e.g. _List | Missing["NotConnected"].
    • "Information" - Tag@@path:line,col. Useful for the file; ignore the line number - it drifts between versions for the same bug.
  • Stack Data - the call stack, e.g. resolveAutoSettings0 -> EvaluateChatInput.
What makes two reports the same bug

Same failing function + same confirmed expression (and pattern, when present). That's it. Everything else is allowed to vary:

  • Different service/model (OpenAI vs Anthropic vs LocalEvaluator) - same bug.
  • Different line number in Information - same bug, different version.
  • Different version/ReleaseID - same bug reported over time.

The trap to avoid: the same expression surfacing in a different function is a sibling bug, not a duplicate. For example Missing["NoModelList"] thrown inside resolveFullModelSpec (chat evaluation) is a different bug from Missing["NoModelList"] passed into makeServiceModelMenu (the model submenu UI), even though both mention NoModelList. They get fixed by different PRs. Keyword co-occurrence is a candidate filter, never a conclusion.

Workflow

1. Establish the reference signature

If the user named a reference issue, read it (gh issue view <N>) and pull the failing function + expression + pattern from its Failure Data. If they instead handed you a set of reports or a fix, derive the signature from the cluster or from the code the fix touches.

2. Search broadly for candidates

Search the distinctive symbols from the signature - typically the head of the expression (NoModelList) and the failing function (resolveFullModelSpec):

bash
gh issue list --state all --search "NoModelList" --json number,title,state
gh issue list --state all --search "resolveFullModelSpec" --json number,title,state

Two rules that matter:

  • Search --state all. You want closed siblings too: the canonical issue may already be closed, and you must not re-close or miss it.
  • Use clean alphanumeric tokens. GitHub's search tokenizer garbles backticks, brackets, and quotes - search NoModelList, not Missing["NoModelList"].

Strong candidates appear in every term's results (intersection). The helper script does this intersection for you.

3. Confirm each candidate against the signature

Do not trust keyword co-occurrence. Open each candidate's Failure Data and verify the failing function and confirmed expression actually match. The helper script extracts and groups these for you:

bash
# Intersect the searches and print every candidate grouped by signature:
python .claude/skills/triage-failure-reports/scripts/extract_signatures.py \
  --search NoModelList --search resolveFullModelSpec

# Or inspect a hand-picked set:
python .claude/skills/triage-failure-reports/scripts/extract_signatures.py 1550 1545 547

Run it from the repo working directory so gh resolves the right repository. It prints, per signature group, every issue with its state, version, service, created date, and ReleaseID - exactly the columns you need for the next two steps. Read scripts/extract_signatures.py if you need to adjust parsing for a report format it doesn't recognize.

Confirm the groups make sense: the cluster you care about shares one function + expression; anything in a different function is a sibling to set aside.

4. Identify and verify the resolver

Every cluster needs a single thing it is "resolved by". Two cases:

  • Duplicates of another issue. Pick the canonical issue - usually the earliest, or the one with the clearest title/discussion. The rest are duplicates of it.
  • Fixed by a merged PR. The user may name it ("these are fixed by #906"), or you find it (gh pr list --search, or the PR that last touched the failing function). Verify it - don't take it on faith. A #NNNN can be an issue or a PR (they share one number space), so confirm with gh pr view <N>, then check the diff actually addresses this signature:
bash
gh pr view 906 --json number,title,state,mergedAt,mergeCommit
gh pr diff 906 | grep -iE 'makeServiceModelMenu|NoModelList'

You want to see the diff add or change handling for the failing function and the expression (e.g. a new makeServiceModelMenu[..., Missing["NoModelList"]] definition). If the diff doesn't touch the signature's code path, it is not the fix.

Show full SKILL.md (534 more words)Show less
5. Verify the timeline - by version, not just date

This is the step that prevents wrong closes. A report can be filed after a fix merged and still be the old bug, because the user was on a stale paclet. So compare the report's Version/ReleaseID (from Debug Data) against the fix, not the issue's creation date:

  • Report version predates the fix -> the fix resolves it. Close it.
  • Report version is newer than the fix and still shows the signature -> it is not resolved (possible regression). Do not close as fixed; flag it for a human and keep it open.
  • An issue created after the merge but on an older version is still fixed - the version is what counts.

For an open canonical issue (not a merged PR), there's no version gate; you're just folding duplicates into the one that stays open.

6. Close with an explanatory comment

Print the grouped findings first - what will be closed, under which canonical issue or PR, and why - so each gh write you then make is an informed action. (Closing and commenting are irreversible and land on other people's issues; the per-write approval is the safety gate, and your verification in steps 3-5 is what makes that approval meaningful.)

Comment, then close. Lead the comment with the GitHub-recognized phrasing so the link is obvious, and pick the close reason that matches the situation:

SituationComment leads withClose reason
Duplicate of another issueDuplicate of #<canonical>.--reason "not planned"
Fixed by a merged PRFixed by #<pr>.--reason completed

gh has no native "duplicate" reason; not planned is the honest fit because the work is tracked elsewhere. A real merged fix is completed.

Quoting matters. WL signatures contain backticks, brackets, and quotes that a shell will mangle (backticks trigger command substitution). Write the comment to a file and use --body-file:

bash
# comment.md  ->  "Duplicate of #1550 - same `resolveFullModelSpec` failure with
#                  `Missing[\"NoModelList\"]` ... Fixed by #1558. Closing as a duplicate."
gh issue comment 1545 --body-file comment.md
gh issue close   1545 --reason "not planned"

Reuse one comment file across a cluster. Use the scratchpad for fetched bodies and comment files. Leave already-closed issues alone (just note them). Leave sibling clusters open unless they have their own resolver.

Worked example

The case this skill was built from:

  • Reference #1550: resolveFullModelSpec[<|... "Name" -> Automatic|>] fails a ConfirmMatch with Missing["NoModelList"] against _List | Missing["NotConnected"].
  • Searching NoModelList x resolveFullModelSpec and grouping yielded 8 issues with that exact signature (1086, 1137, 1303, 1468, 1514, 1530, 1545, 1550) across OpenAI/Anthropic/DeepSeek/GoogleGemini/LocalEvaluator - all duplicates, fixed by PR #1558. Closed with Duplicate of #1550 / not planned.
  • Three more (#547, #969, #892) shared Missing["NoModelList"] but in makeServiceModelMenu - a sibling, not a duplicate. Verified PR #906 added the makeServiceModelMenu[..., Missing["NoModelList"]] overload; their versions (1.4.1, 1.4.6, 1.5.2) all predate it, so closed with Fixed by #906 / completed. #969 was filed after #906 merged but on old v1.4.6 - the version check, not the date, is what confirmed it.

Gotchas

  • Issues and PRs share one number sequence - confirm which a #NNNN is.
  • GitHub search ignores/garbles backticks, brackets, quotes - search symbols.
  • Information line numbers drift across versions - never match on them.
  • Verify a claimed fixing PR's diff; "should be fixed by #X" is a hypothesis.
  • Timeline check is by version/ReleaseID, not issue creation date.
  • A shared expression in a different function is a sibling bug - close it only under its own resolver, never fold it into the wrong cluster.

© WolframResearch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .claude/skills/triage-failure-reports of WolframResearch/Chatbook.

  • SKILL.md
  • scripts/extract_signatures.py

Open the folder on GitHubat commit a601484

Compare with similar skills

Triage Failure Reports next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triage Failure Reports compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triage Failure Reports this skillWolframResearch/Chatbook124—~2.7kAutomated safety check: PassMIT
Community Triageroryeckel/wyoming_openai219—~2.9kAutomated safety check: PassApache-2.0
Weavebench Cua ReproduceAMAP-ML/LongHorizon-Harness1.7k—~1.6kAutomated safety check: PassMIT
Evidence-Driven Testingmichaelshimeles/skills1.3k1 repos~3.9kAutomated safety check: PassNone
Create GitHub IssueNVIDIA/OpenShell16k—~1.7kAutomated safety check: PassApache-2.0
Triage IssuesClickHouse/clickhouse-java1.6k—~904Automated safety check: PassApache-2.0

Similar skills

  • Community Triage

    roryeckel/wyoming_openai

    GitHub issues, pull requests, bug reports, scope questions, and support threads.

    219 GitHub stars~2.9k tokensUpdated 5 days ago
    DevelopmentAuto-check passed
  • Weavebench Cua Reproduce

    AMAP-ML/LongHorizon-Harness

    Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.

    1.7k GitHub stars~1.6k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Evidence-Driven Testing

    michaelshimeles/skills

    Records an annotated screen recording of the agent testing an app hands-on, then posts the video and a results summary to the PR and tracker issue.

    1.3k GitHub starsUsed in 1 repo~3.9k tokens
    Testing & QAAuto-check passed
  • Create GitHub Issue

    NVIDIA/OpenShell

    Official

    Create GitHub issues using the gh CLI. An agent skill from NVIDIA/OpenShell.

    16k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Triage Issues

    ClickHouse/clickhouse-java

    Analyzes a single GitHub issue at a time. An agent skill from ClickHouse/clickhouse-java.

    1.6k GitHub stars~904 tokensUpdated today
    Testing & QAAuto-check passed
  • Gentle AI Issue Creation

    Gentleman-Programming/gentle-shell

    Create and triage GitHub issues from repository evidence. An agent skill from Gentleman-Programming/gentle-shell.

    1.2k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check passed

More from WolframResearch/Chatbook

  • Drive Frontend

    WolframResearch/Chatbook

    Drive a real Wolfram front end (WolframNB) to test Chatbook's actual UI: launch a fresh front end on a private Xvfb display (optionally with the development Chatbook loaded) or attach to one that is…

    124 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Test Skill

    WolframResearch/Chatbook

    Verifies that Chatbook's agent skill support is working. An agent skill from WolframResearch/Chatbook.

    124 GitHub stars~186 tokensUpdated today
    Auto-check passed

Categories

Questions about Triage Failure Reports

What does Triage Failure Reports do?

Triage Chatbook's auto-generated bug-report issues on GitHub by grouping them on their Failure Data signature (the failing Wolfram Language function plus the confirmed expression/pattern), finding…. Triage Failure Reports is an agent skill from WolframResearch/Chatbook. Triage Chatbook's auto-generated bug-report issues on GitHub by grouping them on their Failure Data signature (the failing Wolfram Language function plus the confirmed expression/pattern), finding the canonical issue or the PR that resolves each cluster, and closing the rest with an explanatory comment.

When should I use Triage Failure Reports?

Triage Failure Reports fits situations like: wants to dedupe; clean up GitHub issues - for example find duplicates of 1550; are any of these crash reports the same bug?; close the issues already fixed by PR 906.

How do I install Triage Failure Reports in Claude Code?

Run `npx skills add WolframResearch/Chatbook --skill triage-failure-reports -a claude-code`. Or copy the skill folder (.claude/skills/triage-failure-reports in WolframResearch/Chatbook) into .claude/skills/triage-failure-reports in your project. Claude Code loads it when a task matches its description.

How do I install Triage Failure Reports in Codex?

Run `npx skills add WolframResearch/Chatbook --skill triage-failure-reports -a codex`. Or copy the skill folder (.claude/skills/triage-failure-reports in WolframResearch/Chatbook) into .agents/skills/triage-failure-reports in your project. Codex loads it when a task matches its description.

Can I use Triage Failure Reports in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add WolframResearch/Chatbook --skill triage-failure-reports -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triage-failure-reports, .gemini/skills/triage-failure-reports, .github/skills/triage-failure-reports and .opencode/skills/triage-failure-reports in your project.

What does Triage Failure Reports need to run?

Going by SKILL.md and its folder, Triage Failure Reports needs Python for the scripts in its folder and the command-line tools its instructions call (gh and python). Our summary lists: Python 3.

Does Triage Failure Reports access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Triage Failure Reports safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Triage Failure Reports use?

Triage Failure Reports is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triage Failure Reports use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triage Failure Reports?

Skills that share tags, products or a category with Triage Failure Reports: Community Triage (roryeckel/wyoming_openai, 219 stars), Weavebench Cua Reproduce (AMAP-ML/LongHorizon-Harness, 1.7k stars), Evidence-Driven Testing (michaelshimeles/skills, 1.3k stars) and Create GitHub Issue (NVIDIA/OpenShell, 16k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triage Failure Reports?

WolframResearch (a GitHub organization) maintains it in WolframResearch/Chatbook, which has 124 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 9, 2026.

Source: WolframResearch/Chatbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.