Agent skill

Triage CI Flake

by payloadcms in payloadcms/payload

A skill your agent uses when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests

MITAuto-check passedTesting & QA

Install Triage CI Flake

skills CLI
$ npx skills add payloadcms/payload --skill triage-ci-flake -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install payloadcms/payload triage-ci-flake --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/payloadcms/payload.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/triage-ci-flake .claude/skills/triage-ci-flake && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triage-ci-flake
GitHub stars
45k
Token cost
~4.4k tokens
SKILL.md length
1,323 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests

  • Works in 9 steps: Fetch PR Status Checks → Aggregate Failing Tests → Extract Failure Details from Each Check → …
  • CI tests fail on main branch after PR merge
  • SKILL.md covers Overview, When to Use, PR-Based Workflow (When PR… and MANDATORY First Steps, plus 12 more sections
  • Calls pnpm; reaches github.com

What it does

Triage CI Flake is an agent skill from payloadcms/payload. Use when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests

Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Failing and flaky tests. It works with Payload CMS and GitHub. The repository describes itself as: Payload is the open-source, fullstack Next.js framework, giving you instant backend superpowers. Get a full TypeScript backend and admin panel instantly. Use Payload as a… The licence is MIT.

When your agent uses it

  • CI tests fail on main branch after PR merge
  • Investigating flaky test failures
  • User provides a PR URL/number to aggregate all failing tests

Example prompts

  • “/triage-ci-flake”

Requirements

  • Node.js
  • Pre-approved tools (allowed-tools): Write, Bash(date:*), Bash(mkdir -p *)

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Fetch PR Status Checks
  2. Aggregate Failing Tests
  3. Extract Failure Details from Each Check
  4. Identify Common Patterns
  5. Proceed with Triage
  6. Extract CI Details
  7. Reproduce with Dev Code
  8. Reproduce with Bundled Code
  9. Unable to Reproduce

What it can do on your machine

Read from SKILL.md and the folder at commit c66a7be. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Write
    • Bash(date:*)
    • Bash(mkdir -p *)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triage CI Flake loads about 4.4k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 1,323 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from payloadcms/payload at commit c66a7be, republished under its MIT licence (© payloadcms). 1,323 words, ~4,400 tokens.

Download SKILL.mdSave it as .claude/skills/triage-ci-flake/SKILL.md (or your agent's skills folder).
name
triage-ci-flake
description
Use when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests
allowed-tools
Write, Bash(date:*), Bash(mkdir -p *)

Triage CI Failure

Overview

Systematic workflow designed specifically for triaging and fixing end-to-end (e2e) test failures in CI, especially flaky Playwright tests that pass locally but fail in CI. E2E tests that made it to main are usually flaky due to timing, bundling, or environment differences.

CRITICAL RULE: You MUST run the reproduction workflow before proposing any fixes. No exceptions.

When to Use

  • CI test fails on main branch after PR was merged
  • Test passes locally but fails in CI
  • Test failure labeled as "flaky" or intermittent
  • E2E test timing out in CI only
  • User provides a PR URL/number with failing checks

PR-Based Workflow (When PR URL/Number Provided)

When the user provides a PR URL (e.g., https://github.com/payloadcms/payload/pull/16701) or PR number:

Step 1: Fetch PR Status Checks

First, use tool_search with query "github pull request status checks" to load the GitHub tools.

Then use the github-pull-request_pullRequestStatusChecks tool to get all failing checks:

Tool: github-pull-request_pullRequestStatusChecks
Parameters:
  pullRequestNumber: <extracted from URL or provided>
  repo: { owner: "payloadcms", name: "payload" }
Step 2: Aggregate Failing Tests

Parse the status checks response and create a summary table:

SuiteCheck StatusTarget URL
admine2elist-viewfailed[link]
plugin-import-exportfailed[link]

For each failing check, the context field contains the test suite name (e.g., admin__e2e__list-view).

Step 3: Extract Failure Details from Each Check

For each failing check:

  1. Visit the targetUrl to get detailed failure logs
  2. Extract the specific test name and error message
  3. Add to the aggregated failure list

Present a consolidated summary:

markdown
## Failing Tests Summary

### 1. admin**e2e**list-view (3/4)

- **Test**: "should use custom pagination limit"
- **Error**: Locator `.per-page__button` not found
- **File**: test/admin/e2e/list-view/e2e.spec.ts:1495

### 2. plugin-import-export

- **Test**: "should inherit limit from list view URL"
- **Error**: Locator `.per-page__button` not found
- **File**: test/plugin-import-export/e2e.spec.ts:150
Step 4: Identify Common Patterns

Look for patterns across failures:

  • Same selector errors → likely a component change needing test updates
  • Same test file → localized issue
  • Different errors in same suite → may be test pollution
Step 5: Proceed with Triage

For each unique failure, follow the standard reproduction workflow below.

MANDATORY First Steps

YOU MUST EXECUTE THESE COMMANDS. Reading code or analyzing logs does NOT count as reproduction.

  1. Extract suite name, test name, and error from CI logs
  2. EXECUTE: Kill port 3000 to avoid conflicts
  3. EXECUTE: pnpm dev $SUITE_NAME (use run_in_background=true)
  4. EXECUTE: Wait for server to be ready (check with curl or sleep)
  5. EXECUTE: Run the specific failing test with Playwright directly (npx playwright test test/TEST_SUITE_NAME/e2e.spec.ts:31:3 --headed -g "TEST_DESCRIPTION_TARGET_GOES_HERE")
  6. If test passes, EXECUTE: pnpm prepare-run-test-against-prod
  7. EXECUTE: pnpm dev:prod $SUITE_NAME and run test again

Only after EXECUTING these commands and seeing their output can you proceed to analysis and fixes.

"Analysis from logs" is NOT reproduction. You must RUN the commands.

Core Workflow

dot
digraph triage_ci {
    "CI failure reported" [shape=box];
    "Extract details from CI logs" [shape=box];
    "Identify suite and test name" [shape=box];
    "Run dev server: pnpm dev $SUITE" [shape=box];
    "Run specific test by name" [shape=box];
    "Did test fail?" [shape=diamond];
    "Debug with dev code" [shape=box];
    "Run prepare-run-test-against-prod" [shape=box];
    "Run: pnpm dev:prod $SUITE" [shape=box];
    "Run specific test again" [shape=box];
    "Did test fail now?" [shape=diamond];
    "Debug bundling issue" [shape=box];
    "Unable to reproduce - check logs" [shape=box];
    "Fix and verify" [shape=box];

    "CI failure reported" -> "Extract details from CI logs";
    "Extract details from CI logs" -> "Identify suite and test name";
    "Identify suite and test name" -> "Run dev server: pnpm dev $SUITE";
    "Run dev server: pnpm dev $SUITE" -> "Run specific test by name";
    "Run specific test by name" -> "Did test fail?";
    "Did test fail?" -> "Debug with dev code" [label="yes"];
    "Did test fail?" -> "Run prepare-run-test-against-prod" [label="no"];
    "Run prepare-run-test-against-prod" -> "Run: pnpm dev:prod $SUITE";
    "Run: pnpm dev:prod $SUITE" -> "Run specific test again";
    "Run specific test again" -> "Did test fail now?";
    "Did test fail now?" -> "Debug bundling issue" [label="yes"];
    "Did test fail now?" -> "Unable to reproduce - check logs" [label="no"];
    "Debug with dev code" -> "Fix and verify";
    "Debug bundling issue" -> "Fix and verify";
}

Step-by-Step Process

1. Extract CI Details

From CI logs or GitHub Actions URL, identify:

  • Suite name: Directory name (e.g., i18n, fields, lexical)
  • Test file: Full path (e.g., test/i18n/e2e.spec.ts)
  • Test name: Exact test description
  • Error message: Full stack trace
  • Test type: E2E (Playwright)
2. Reproduce with Dev Code

CRITICAL: Always run the specific test by name, not the full suite.

SERVER MANAGEMENT RULES:

  1. ALWAYS kill all servers before starting a new one
  2. NEVER assume ports are free
  3. ALWAYS wait for server ready confirmation before running tests
bash
# ========================================
# STEP 2A: STOP ALL SERVERS
# ========================================
lsof -ti:3000 | xargs kill -9 2>/dev/null || echo "Port 3000 clear"

# ========================================
# STEP 2B: START DEV SERVER
# ========================================
# Start dev server with the suite (in background with run_in_background=true)
pnpm dev $SUITE_NAME

# ========================================
# STEP 2C: WAIT FOR SERVER READY
# ========================================
# Wait for server to be ready (REQUIRED - do not skip)
until curl -s http://localhost:3000/admin > /dev/null 2>&1; do sleep 1; done && echo "Server ready"

# ========================================
# STEP 2D: RUN SPECIFIC TEST
# ========================================
# Run ONLY the specific failing test using Playwright directly
# For E2E tests (DO NOT use pnpm test:e2e as it spawns its own server):
pnpm exec playwright test test/$SUITE_NAME/e2e.spec.ts -g "exact test name"

Did the test fail?

  • ✅ YES: You reproduced it! Proceed to debug with dev code.
  • ❌ NO: Continue to step 3 (bundled code test).
3. Reproduce with Bundled Code

If test passed with dev code, the issue is likely in bundled/production code.

IMPORTANT: You MUST stop the dev server before starting prod server.

bash
# ========================================
# STEP 3A: STOP ALL SERVERS (INCLUDING DEV SERVER FROM STEP 2)
# ========================================
lsof -ti:3000 | xargs kill -9 2>/dev/null || echo "Port 3000 clear"

# ========================================
# STEP 3B: BUILD AND PACK FOR PROD
# ========================================
# Build all packages and pack them (this takes time - be patient)
pnpm prepare-run-test-against-prod

# ========================================
# STEP 3C: START PROD SERVER
# ========================================
# Start prod dev server (in background with run_in_background=true)
pnpm dev:prod $SUITE_NAME

# ========================================
# STEP 3D: WAIT FOR SERVER READY
# ========================================
# Wait for server to be ready (REQUIRED - do not skip)
until curl -s http://localhost:3000/admin > /dev/null 2>&1; do sleep 1; done && echo "Server ready"

# ========================================
# STEP 3E: RUN SPECIFIC TEST
# ========================================
# Run the specific test again using Playwright directly
pnpm exec playwright test test/$SUITE_NAME/e2e.spec.ts -g "exact test name"

Did the test fail now?

  • ✅ YES: Bundling or production build issue. Look for:
    • Missing exports in package.json
    • Build configuration problems
    • Code that behaves differently when bundled
  • ❌ NO: Unable to reproduce locally. Proceed to step 4.
4. Unable to Reproduce

If you cannot reproduce locally after both attempts:

  • Review CI logs more carefully for environment differences
  • Check for race conditions (run test multiple times: for i in {1..10}; do pnpm test:e2e...; done)
  • Look for CI-specific constraints (memory, CPU, timing)
  • Consider if it's a true race condition that's highly timing-dependent

Common Flaky Test Patterns

Race Conditions
  • Page navigating while assertions run
  • Network requests not settled before assertions
  • State updates not completed

Fix patterns:

  • Use Playwright's web-first assertions (toBeVisible(), toHaveText())
  • Wait for specific conditions, not arbitrary timeouts
  • Use waitForFunction() with condition checks
Test Pollution
  • Tests leaving data in database
  • Shared state between tests
  • Missing cleanup in afterEach

Fix patterns:

  • Track created IDs and clean up in afterEach
  • Use isolated test data
  • Don't use deleteAll that affects other tests
Timing Issues
  • setTimeout/sleep instead of condition-based waiting
  • Not waiting for page stability
  • Animations/transitions not complete

Fix patterns:

  • Use waitForPageStability() helper
  • Wait for specific DOM states
  • Use Playwright's built-in waiting mechanisms

Linting Considerations

When fixing e2e tests, be aware of these eslint rules:

  • playwright/no-networkidle - Avoid waitForLoadState('networkidle') (use condition-based waiting instead)
  • payload/no-wait-function - Avoid custom wait() functions (use Playwright's built-in waits)
  • payload/no-flaky-assertions - Avoid non-retryable assertions
  • playwright/prefer-web-first-assertions - Use built-in Playwright assertions

Existing code may violate these rules - when adding new code, follow the rules even if existing code doesn't.

Verification

After fixing:

bash
# Ensure dev server is running on port 3000
# Run test multiple times to confirm stability
for i in {1..10}; do
  pnpm exec playwright test test/$SUITE_NAME/e2e.spec.ts -g "exact test name" || break
done

# Run full suite
pnpm exec playwright test test/$SUITE_NAME/e2e.spec.ts

# If you modified bundled code, test with prod build
lsof -ti:3000 | xargs kill -9 2>/dev/null
pnpm prepare-run-test-against-prod
pnpm dev:prod $SUITE_NAME
until curl -s http://localhost:3000/admin > /dev/null; do sleep 1; done
pnpm exec playwright test test/$SUITE_NAME/e2e.spec.ts
Show full SKILL.md (531 more words)Show less

The Iron Law

NO FIX WITHOUT REPRODUCTION FIRST

If you propose a fix before completing steps 1-3 of the workflow, you've violated this skill.

This applies even when:

  • The fix seems obvious from the logs
  • You've seen this error before
  • Time pressure from the team
  • You're confident about the root cause
  • The logs show clear stack traces

No exceptions. Run the reproduction workflow first.

Rationalization Table

Every excuse for skipping reproduction, and why it's wrong:

RationalizationReality
"The logs show the exact error"Logs show symptoms, not root cause. Reproduce.
"I can see the problem in the code"You're guessing. Reproduce to confirm.
"This is obviously a race condition"Maybe. Reproduce to be sure.
"I've seen this error before"This might be different. Reproduce.
"The stack trace is clear"Stack trace shows where, not why. Reproduce.
"Time pressure - need to fix fast"Reproducing IS fast. Guessing wastes time.
"The test file shows the issue"Reading ≠ running. Execute the commands.
"I'll analyze the code first"Code analysis comes AFTER reproduction.
"Let me investigate the root cause"Reproduction IS the investigation.
"I need to understand the error"Understanding comes from seeing it fail.

Red Flags - STOP

These thoughts mean you're about to violate the skill:

  • "Let me just fix the obvious issue" → STOP. Execute reproduction first.
  • "I'll increase the timeout" → STOP. Find root cause via reproduction.
  • "This looks like X from the logs" → STOP. Run the test.
  • "I'll run the full suite" → STOP. Run specific test first.
  • "Let me check the code" → STOP. Run reproduction first.
  • "I can see the problem in the stack trace" → STOP. Reproduce to confirm.
  • "The logs are clear enough" → STOP. Execute the commands.
  • "Let me analyze..." → STOP. Analysis comes after reproduction.

All of these mean: Stop immediately. Execute the reproduction commands. No analysis without execution.

Example Session

CORRECT approach - executes commands first:

User: CI failed on i18n suite - https://github.com/.../actions/runs/XXX

Agent:
1. Extracted details from CI logs:
   - Suite: i18n
   - Test: "ensure translations update correctly when switching language"
   - File: test/i18n/e2e.spec.ts:124
   - Error: "Error loading language: en" TypeError: network error

2. Executing reproduction workflow:
   [Uses Bash tool to kill ports]
   $ lsof -ti:3000 | xargs kill -9 2>/dev/null || echo "Port 3000 clear"

   [Uses Bash tool with run_in_background=true]
   $ pnpm dev i18n

   [Waits for server to start with curl check]
   $ until curl -s http://localhost:3000/admin > /dev/null; do sleep 1; done && echo "Server ready"

   [Uses Bash tool to run Playwright directly]
   $ pnpm exec playwright test test/i18n/e2e.spec.ts -g "ensure translations update correctly when switching language"

   Output shows: Test PASSED locally

3. Test passed with dev code, trying bundled code:
   [Kills ports again]
   $ lsof -ti:3000 | xargs kill -9 2>/dev/null

   [Uses Bash tool]
   $ pnpm prepare-run-test-against-prod

   [Uses Bash tool with run_in_background=true]
   $ pnpm dev:prod i18n

   [Waits for server]
   $ until curl -s http://localhost:3000/admin > /dev/null; do sleep 1; done

   [Uses Bash tool]
   $ pnpm exec playwright test test/i18n/e2e.spec.ts -g "ensure translations update correctly when switching language"

   Output shows: Test FAILED - reproduced!

4. Now that I've reproduced it, analyzing the failure...
   [Code analysis, root cause identification]

5. Implementing fix: [specific changes]

6. Verification:
   [Uses Bash tool]
   $ for i in {1..10}; do pnpm test:e2e i18n -g "exact test"; done
   All 10 runs passed!

INCORRECT approach - analyzes without executing:

User: CI failed on i18n suite

Agent: Looking at the error logs, I can see this is a race condition...
[Analyzes code, proposes fix without running any commands]

❌ WRONG - This violates the skill. Must execute reproduction commands first.

Common Mistakes

MistakeFix
Running full test suite firstRun specific test by name
Skipping dev code reproductionAlways try dev code first
Not testing with bundled codeIf dev passes, test with prepare-run-test-against-prod
Proposing fix without reproducingFollow the workflow - reproduce first
Using networkidle in new codeUse condition-based waiting with waitForFunction()
Adding arbitrary wait() callsUse Playwright's built-in assertions and waits

Key Principles

  1. Reproduce before fixing: Never propose a fix without reproducing the issue
  2. Test specifically: Run the exact failing test, not the full suite
  3. Dev first, prod second: Check dev code before bundled code
  4. Follow the workflow: No shortcuts - the steps exist to save time
  5. Verify stability: Run tests multiple times to confirm fix

Completion: Creating a PR

After you have:

  1. ✅ Reproduced the issue
  2. ✅ Implemented a fix
  3. ✅ Verified the fix passes locally (multiple runs)
  4. ✅ Tested with prod build (if applicable)

You MUST prompt the user to create a PR:

The fix has been verified and is ready for review. Would you like me to create a PR with these changes?

Summary of changes:
- [List files modified]
- [Brief description of the fix]
- [Verification results]

IMPORTANT:

  • DO NOT automatically create a PR - always ask the user first
  • Provide a clear summary of what was changed and why
  • Include verification results (number of test runs, pass rate)
  • Let the user decide whether to create the PR immediately or make additional changes first

This ensures the user has visibility and control over what gets submitted for review.

© payloadcms, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/triage-ci-flake of payloadcms/payload.

Open the folder on GitHubat commit c66a7be

Compare with similar skills

Triage CI Flake next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triage CI Flake compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triage CI Flake this skillpayloadcms/payload45k—~4.4kAutomated safety check: PassMIT
GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb6.7k—~4.4kAutomated safety check: PassApache-2.0
Fix Issueremix-run/remix33k—~1.8kAutomated safety check: PassMIT
Issue To Regression Testbrunosabot/streamline-card269—~529Automated safety check: PassMIT
Detect Flaky Testsagent-substrate/substrate4.5k—~3kAutomated safety check: PassApache-2.0
Fix Random CI Test Failuredotnet/macios2.9k—~1.3kAutomated safety check: PassCustom licence

Similar skills

  • Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.

    6.7k GitHub stars~4.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Fix Issue

    remix-run/remix

    Fix a reported issue in Remix from a GitHub issue. An agent skill from remix-run/remix.

    33k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Issue To Regression Test

    brunosabot/streamline-card

    A skill your agent uses when the user asks to fix a bug, references a GitHub issue number, or describes an issue and wants a fix.

    269 GitHub stars~529 tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Detect Flaky Tests

    agent-substrate/substrate

    Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM…

    4.5k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Investigate and fix flaky/random CI test failures in dotnet/macios.

    2.9k GitHub stars~1.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Adds dotnet/maui-specific context for investigating failing PR checks and broken nightly builds: pipelines, Helix logs, binlogs and merge-readiness verdicts.

    23k GitHub stars~2k tokensUpdated today
    Testing & QAAuto-check passed

More from payloadcms/payload

All 9 skills in this repo
  • Record PR Demo

    payloadcms/payload

    A skill your agent uses when a Payload pull request needs a concise visual walkthrough for reviewers.

    45k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Payload

    payloadcms/payload

    A skill your agent uses when working with Payload projects (payload.config.ts, collections, fields, hooks, access control, Payload API).

    45k GitHub starsUsed in 5 repos~6.2k tokens
    Auto-check passed
  • Audit Dependencies

    payloadcms/payload

    A skill your agent uses when fixing dependency vulnerabilities, running pnpm audit, or when the audit-dependencies CI check fails

    45k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Generate Translations

    payloadcms/payload

    A skill your agent uses when new translation keys are added to packages to generate new translations strings

    45k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Ui4 Convert Tests

    payloadcms/payload

    A skill your agent uses when UI changes are complete and e2e tests need updating.

    45k GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Payload Accessibility

    payloadcms/payload

    A skill your agent uses when changing or reviewing rendered Payload UI, interaction or focus behavior, semantic markup, accessibility tests, or WCAG/VPAT evidence.

    45k GitHub stars~808 tokensUpdated today
    Auto-check passed

Categories

Questions about Triage CI Flake

What does Triage CI Flake do?

A skill your agent uses when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests. Triage CI Flake is an agent skill from payloadcms/payload.

When should I use Triage CI Flake?

Triage CI Flake fits situations like: CI tests fail on main branch after PR merge; investigating flaky test failures; user provides a PR URL/number to aggregate all failing tests.

How do I install Triage CI Flake in Claude Code?

Run `npx skills add payloadcms/payload --skill triage-ci-flake -a claude-code`. Or copy the skill folder (.agents/skills/triage-ci-flake in payloadcms/payload) into .claude/skills/triage-ci-flake in your project. Claude Code loads it when a task matches its description.

How do I install Triage CI Flake in Codex?

Run `npx skills add payloadcms/payload --skill triage-ci-flake -a codex`. Or copy the skill folder (.agents/skills/triage-ci-flake in payloadcms/payload) into .agents/skills/triage-ci-flake in your project. Codex loads it when a task matches its description.

Can I use Triage CI Flake in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add payloadcms/payload --skill triage-ci-flake -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triage-ci-flake, .gemini/skills/triage-ci-flake, .github/skills/triage-ci-flake and .opencode/skills/triage-ci-flake in your project.

What does Triage CI Flake need to run?

Going by SKILL.md and its folder, Triage CI Flake needs the command-line tools its instructions call (pnpm). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Write, Bash(date:*), Bash(mkdir -p *).

Does Triage CI Flake access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Triage CI Flake safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triage CI Flake use?

Triage CI Flake is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triage CI Flake use?

About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triage CI Flake?

Skills that share tags, products or a category with Triage CI Flake: GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars), Fix Issue (remix-run/remix, 33k stars), Issue To Regression Test (brunosabot/streamline-card, 269 stars) and Detect Flaky Tests (agent-substrate/substrate, 4.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triage CI Flake?

payloadcms (a GitHub organization) maintains it in payloadcms/payload, which has 45,140 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 8, 2026.

Source: payloadcms/payload on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.