Agent skill

Triage Integration Test Failures

by newrelic in newrelic/newrelic-dotnet-agent

A skill your agent uses when a CI integration, unbounded, or container test job fails intermittently, passes on a rerun, fails identically across every matrix variant, or shows an infra-shaped error…

Apache-2.0Auto-check passedTesting & QA

Install Triage Integration Test Failures

skills CLI
$ npx skills add newrelic/newrelic-dotnet-agent --skill triage-integration-test-failures -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install newrelic/newrelic-dotnet-agent triage-integration-test-failures --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/newrelic/newrelic-dotnet-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/triage-integration-test-failures .claude/skills/triage-integration-test-failures && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triage-integration-test-failures
GitHub stars
117
Token cost
~2.5k tokens
SKILL.md length
1,420 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when a CI integration, unbounded, or container test job fails intermittently, passes on a rerun, fails identically across every matrix variant, or shows an infra-shaped error…

  • Works in 4 steps: is it even a flake? → classify by evidence, not guess → pick the remedy for the category → …
  • A CI integration
  • SKILL.md covers Overview, Hard rule: delegate the reading, Step 1: is it even a flake? and Step 2: classify by evidence,…, plus 4 more sections
  • Calls az, gh and kubectl

What it does

Triage Integration Test Failures is an agent skill from newrelic/newrelic-dotnet-agent. Use when a CI integration, unbounded, or container test job fails intermittently, passes on a rerun, fails identically across every matrix variant, or shows an infra-shaped error (connection timeout, account lock, native crash, docker daemon not ready) rather than a clear assertion diff -- the cause may be environmental, tooling, or a real product defect exposed only sometimes. Classify the failure before proposing any fix.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Integration testing. It works with Docker. The repository describes itself as: The New Relic .NET language agent. The licence is Apache-2.0.

When your agent uses it

  • A CI integration
  • Container test job fails intermittently
  • Passes on a rerun
  • Fails identically across every matrix variant

Example prompts

  • “/triage-integration-test-failures”

Requirements

  • Docker

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. is it even a flake?
  2. classify by evidence, not guess
  3. pick the remedy for the category
  4. record the decision

What it can do on your machine

Read from SKILL.md and the folder at commit aca7658. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • az
    • gh
    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use az, gh and kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triage Integration Test Failures loads about 2.5k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 1,420 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from newrelic/newrelic-dotnet-agent at commit aca7658, republished under its Apache-2.0 licence (© newrelic). 1,420 words, ~2,496 tokens.

Download SKILL.mdSave it as .claude/skills/triage-integration-test-failures/SKILL.md (or your agent's skills folder).
name
triage-integration-test-failures
description
Use when a CI integration, unbounded, or container test job fails intermittently, passes on a rerun, fails identically across every matrix variant, or shows an infra-shaped error (connection timeout, account lock, native crash, docker daemon not ready) rather than a clear assertion diff -- the cause may be environmental, tooling, or a real product defect exposed only sometimes. Classify the failure before proposing any fix.

Triage integration test failures

Overview

A failure investigation ends with one of five remedies, not a patch guessed under time pressure. Classify the failure first. The category decides the remedy; the wrong category produces a fix that treats a symptom.

REQUIRED BACKGROUND: superpowers:systematic-debugging governs root-cause work in general. This skill adds the classification step and the toolkit specific to this repo's CI, agent, and test infrastructure.

Hard rule: delegate the reading

A CI job log, an az monitor/kubectl dump, and an agent log are evidence, not the deliverable. None of them belongs in the main session.

  • Run every command in the toolkit below inside a subagent (Agent tool -- general-purpose or Explore), not in the main session. Give the subagent the exact command and the exact question ("did this job fail on a connection error before any agent code ran, or on an assertion diff?"); it returns a short verdict with the one or two evidence lines that prove it, never the raw log or the full command output.
  • This applies even to a "quick look" -- a CI log or agent log runs to hundreds of KB or has single lines tens of KB wide, and once that text lands in the main session it is re-billed on every later turn for the rest of the investigation.
  • For agent-log evidence specifically, dispatch a subagent that follows the analyze-dotnet-agent-logs skill; do not hand-parse the log yourself even inside that subagent.
  • The main session's job is Steps 2-4: collect verdicts, classify, decide, and record the decision. It holds conclusions, not evidence bytes.

Step 1: is it even a flake?

  • Same PR, same test, one run failed and a rerun passed with no code change: flake. Continue.
  • Fails on every run, or only after a real code change: this is a regression, not a flake. Use superpowers:systematic-debugging instead.
  • Check gh run list for the workflow: has this exact test or job failed before on unrelated PRs? Count occurrences of the exact signature across recent scheduled runs, and state the rate as per-run, not per-call -- one occurrence in 100 runs is not proof of absence. A recurring signature across unrelated changes is strong evidence of an environmental cause, not a code bug.

Step 2: classify by evidence, not guess

Before matching a row, trace the failing message or exception string to the code that emits it, and note whether the assertion sits in the test body or in fixture setup/teardown (e.g. RemoteApplicationFixture.TestForKnownProblems is a health check, not the test's own logic) -- this decides which row applies.

CategorySignatureEvidence to pull
Infra / clusterWhole class of tests fails identically across every framework/OS variant in the same run; error is a connection, timeout, or account/lock error thrown before any agent code runsaz monitor metrics list on the AKS node and load balancer; kubectl get pods -n unbounded-services -o wide for restarts/uptime; does a bare retry pass (transient) or fail identically (persistent server state)?
Timing / raceAssertion is a count or an event that "usually" arrives, on a harvest or aggregator cycle; failure shows the right data, late or short by one cycleThe agent log, via analyze-dotnet-agent-logs -- confirm the exact interleaving (e.g. Seen vs Sent counts) before touching the test
Tooling / dependency bugCrash signature is inside a third-party tool's native code, not in test or agent code; all tests report Passed, only the collector/runner process diesThe crash stack (module name, access-violation code); the tool's own changelog/issue tracker for a fix version already in flight
External / upstreamError surfaces in CI infrastructure setup (runner boot, docker daemon, VM image), not in the app or the agent at allSearch the CI vendor's own issue tracker (e.g. actions/runner-images) for the exact error string before assuming it's local
Product defect, exposed nondeterministicallyThe traced assertion is a health check or invariant (fixture setup/teardown), or the emitting code shows a structural hole (missing try/finally, unguarded async race)The suspect code path traced above; a repro run that proves the path executed (see Toolkit)

Never conclude "just flaky" without pulling at least one row's evidence. See [[feedback_no_guessing]] -- a root cause claim needs a log line, a metric, or a matching upstream issue behind it, not a hunch. An assertion that exists to catch a product defect is never reclassified as timing/race and its threshold is never widened -- trace and fix the defect instead.

Show full SKILL.md (695 more words)Show less

Step 3: pick the remedy for the category

CategoryRemedyDo NOT
Infra, transient (network blip)CI-level retry-once on the job, plus raise the client-side timeout that actually stalled; document as an accepted riskSilently widen unrelated timeouts repo-wide
Infra, persistent server state (account lock, stuck pod)Operational fix (restart the pod/service); decide explicitly whether to harden the exposure, and record that decision even if the answer is "no, too costly"Write code to retry around infra state that a retry cannot clear
Timing / race in event-harvest assertionsA race-free wait keyed on the exact asserted evidence (e.g. an AgentLogBase.WaitForMetricAggregateCallCount-style helper polling the real aggregate)Bump a harvest-cycle magic number again -- each prior bump on the same test is evidence the race is still there, only the window changed
Tooling bug with no correctness impactWait for the upstream version bump already in flight; record the specific version and re-check after it landsChange unrelated project settings (e.g. PDB format) as a workaround before confirming the bump doesn't already fix it
External/upstream CI regressionApply the vendor's own recommended workaround inside the repo's CI composite action/step; link the upstream issue in a comment; remove the workaround once fixed upstreamInvent a bespoke workaround when the vendor already published one
Product defect, exposed nondeterministicallyTrace the code path, file a ticket, leave the check alone; hand off to superpowers:systematic-debugging with the gathered evidence as the starting hypothesis and stop triageReclassify as timing/race, widen the assertion or its timeout, or suppress it to keep CI green

Step 4: record the decision

Every one of the categories above ends in a decision someone will ask about again ("why don't we just fix the LB exposure", "did coverlet ever get bumped"). Capture, in a project memory or handoff note:

  1. The failure signature and the run/job that showed it.
  2. The evidence that ruled out the other categories.
  3. The remedy chosen, and what was explicitly rejected and why.
  4. Anything left open (a version to watch for, a follow-up not started).

Toolkit quick reference

Each row runs inside a subagent per the hard rule above; the main session sees only the verdict it returns.

  • Both gh run view --job <id> --log-failed and --log truncate silently past ~889 lines / 168KB, cutting off mid agent-log dump. For the complete text: gh api repos/<org>/<repo>/actions/jobs/<jobid>/logs. Copy any log to a file before a step that clears it -- lost evidence cannot be re-created. The pristine agent log also survives as CI artifact integration-test-results-<matrix> (all_solutions.yml uploads C:\IntegrationTestWorkingDirectory\**\*.log unconditionally).
  • az monitor metrics list --resource <aks-or-lb-id> --metric <name> -- cluster/LB health at the failure timestamp (MSYS_NO_PATHCONV=1 needed for the slash-prefixed resource ID in Git Bash).
  • kubectl get pods -n unbounded-services -o wide -- pod restarts/uptime after az aks get-credentials.
  • analyze-dotnet-agent-logs skill -- agent-side evidence (harvest timing, Seen/Sent counts, connect sequence) from the test's own log.
  • run-integration-tests skill -- reproduce the failing test locally before trusting a fix. A null result only counts if the code path is proven to have executed (log entry/exit counts); when the agent itself is a suspect, rerun once with the agent detached to isolate the confound. Delegate the run itself, per that skill's own subagent guidance; take back the summary, not the console output.

Common mistakes

  • Treating "all variants failed the same way" as proof of a code bug -- it is usually the opposite: identical failure across unrelated variants points at a shared external dependency, not the code under test.
  • This pattern covers whole-class failures across unrelated variants only -- a single test failing once, or a fixture health-check assertion firing, is not covered by it and may be the product-defect row instead.
  • Fixing the assertion's timeout/retry count instead of the mechanism that makes the wait racy. If a similar bump has already happened once on this test, that is the signal to stop bumping and find the mechanism.
  • Closing the investigation without writing down why the infra issue was not hardened -- the same question comes back next time the flake recurs.
  • Reading a CI log or agent log directly in the main session "just this once" -- delegate it every time; a single raw read is how a multi-step investigation quietly fills the context window.

© newrelic, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/triage-integration-test-failures of newrelic/newrelic-dotnet-agent.

Open the folder on GitHubat commit aca7658

Compare with similar skills

Triage Integration Test Failures next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triage Integration Test Failures compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triage Integration Test Failures this skillnewrelic/newrelic-dotnet-agent117—~2.5kAutomated safety check: PassApache-2.0
Remote Executor Integration Testsopeninterpreter/openinterpreter69k2 repos~842Automated safety check: PassApache-2.0
MongoDB Source Connector E2E Harnessairbytehq/airbyte22k—~1.9kAutomated safety check: PassCustom licence
Integration Tests for pRESTprest/prest4.6k—~1.1kAutomated safety check: PassMIT
Go Redis Client Test Runnerredis/go-redis22k—~786Automated safety check: PassBSD-2-Clause
Test Weave Router with Codexweave-os/router5.6k—~4.7kAutomated safety check: NotesApache-2.0

Similar skills

  • Remote Executor Integration Tests

    openinterpreter/openinterpreter

    Explains how to run agent integration tests against remote executors, using Docker for Linux or Wine for Windows, and how to opt tests in or skip them.

    69k GitHub starsUsed in 2 repos~842 tokens
    Testing & QAAuto-check passed
  • Official

    Starts a throwaway MongoDB 7.0 replica set and runs the Airbyte spec, check, discover and read commands against source-mongodb-v2 images for local end-to-end testing.

    22k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Guides writing and reviewing pREST Docker-based integration tests so every HTTP request is explained by step comments or table-driven descriptions.

    4.6k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Explains how to run go-redis tests: the Docker Compose stack, make targets, focusing a single Ginkgo spec, the e2e suite and the version environment variables.

    22k GitHub stars~786 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Local test harness for the Weave router: a docker compose stack plus codex exec runs that confirm how Codex requests are routed, translated and marked.

    5.6k GitHub stars~4.7k tokensUpdated today
    Testing & QAAuto-check: notes
  • Official

    Starts a local PostgreSQL 16 container, loads SQL fixtures and runs the Airbyte spec, check, discover and read commands against a chosen source-postgres image.

    22k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check passed

More from newrelic/newrelic-dotnet-agent

  • Analyze Dotnet Agent Logs

    newrelic/newrelic-dotnet-agent

    Parse and diagnose New Relic .NET agent logs (newrelicagent.log) and profiler logs (NewRelic.Profiler.<pid.log) from a support ticket.

    117 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Angler Dotnet Release

    newrelic/newrelic-dotnet-agent

    Open an Angler PR that adds the latest .NET agent release to metricnames.txt.

    117 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Build Dotnet Agent

    newrelic/newrelic-dotnet-agent

    Build the New Relic .NET agent locally and validate that code changes compile.

    117 GitHub stars~871 tokensUpdated today
    Auto-check passed
  • Run Integration Tests

    newrelic/newrelic-dotnet-agent

    Run the New Relic .NET agent integration tests locally. An agent skill from newrelic/newrelic-dotnet-agent.

    117 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Update Dotty Docs

    newrelic/newrelic-dotnet-agent

    A skill your agent uses when the user mentions a Dotty PR, dotty (package/dependency) updates, or asks to update the .NET agent compatibility docs / net-agent-compatibility-requirements after tested…

    117 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Dotnet Unit Test Coverage

    newrelic/newrelic-dotnet-agent

    Mandatory unit-test and code-coverage rules for the .NET agent.

    117 GitHub stars~1k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Triage Integration Test Failures

What does Triage Integration Test Failures do?

A skill your agent uses when a CI integration, unbounded, or container test job fails intermittently, passes on a rerun, fails identically across every matrix variant, or shows an infra-shaped error…. Triage Integration Test Failures is an agent skill from newrelic/newrelic-dotnet-agent. Use when a CI integration, unbounded, or container test job fails intermittently, passes on a rerun, fails identically across every matrix variant, or shows an infra-shaped error (connection timeout, account lock, native crash, docker daemon not ready) rather than a clear assertion diff -- the cause may be environmental, tooling, or a real product defect exposed only sometimes.

When should I use Triage Integration Test Failures?

Triage Integration Test Failures fits situations like: A CI integration; container test job fails intermittently; passes on a rerun; fails identically across every matrix variant.

How do I install Triage Integration Test Failures in Claude Code?

Run `npx skills add newrelic/newrelic-dotnet-agent --skill triage-integration-test-failures -a claude-code`. Or copy the skill folder (.claude/skills/triage-integration-test-failures in newrelic/newrelic-dotnet-agent) into .claude/skills/triage-integration-test-failures in your project. Claude Code loads it when a task matches its description.

How do I install Triage Integration Test Failures in Codex?

Run `npx skills add newrelic/newrelic-dotnet-agent --skill triage-integration-test-failures -a codex`. Or copy the skill folder (.claude/skills/triage-integration-test-failures in newrelic/newrelic-dotnet-agent) into .agents/skills/triage-integration-test-failures in your project. Codex loads it when a task matches its description.

Can I use Triage Integration Test Failures in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add newrelic/newrelic-dotnet-agent --skill triage-integration-test-failures -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triage-integration-test-failures, .gemini/skills/triage-integration-test-failures, .github/skills/triage-integration-test-failures and .opencode/skills/triage-integration-test-failures in your project.

What does Triage Integration Test Failures need to run?

Going by SKILL.md and its folder, Triage Integration Test Failures needs the command-line tools its instructions call (az, gh and kubectl). Our summary lists: Docker.

Does Triage Integration Test Failures access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Triage Integration Test Failures safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triage Integration Test Failures use?

Triage Integration Test Failures is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triage Integration Test Failures use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triage Integration Test Failures?

Skills that share tags, products or a category with Triage Integration Test Failures: Remote Executor Integration Tests (openinterpreter/openinterpreter, 69k stars), MongoDB Source Connector E2E Harness (airbytehq/airbyte, 22k stars), Integration Tests for pREST (prest/prest, 4.6k stars) and Go Redis Client Test Runner (redis/go-redis, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triage Integration Test Failures?

newrelic (a GitHub organization) maintains it in newrelic/newrelic-dotnet-agent, which has 117 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: newrelic/newrelic-dotnet-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.