Triage CI Flake
payloadcms/payload
A skill your agent uses when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests
Monitors or repairs an open GitHub PR: CI failures, conflicts, review threads, and merge readiness, reporting state changes.
$ npx skills add mblode/agent-skills --skill pr-babysitter -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mblode/agent-skills pr-babysitter --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mblode/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pr-babysitter .claude/skills/pr-babysitter && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pr-babysitter" agent skill from https://github.com/mblode/agent-skills/tree/main/skills/pr-babysitter into .claude/skills/pr-babysitter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pr-babysitter", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mblode/agent-skills/tree/main/skills/pr-babysitterType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mblode/agent-skills --skill pr-babysitter -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mblode/agent-skills pr-babysitter --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mblode/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/pr-babysitter .agents/skills/pr-babysitter && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pr-babysitter" agent skill from https://github.com/mblode/agent-skills/tree/main/skills/pr-babysitter into .agents/skills/pr-babysitter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pr-babysitter", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mblode/agent-skills --skill pr-babysitter -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mblode/agent-skills pr-babysitter --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mblode/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/pr-babysitter .cursor/skills/pr-babysitter && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pr-babysitter" agent skill from https://github.com/mblode/agent-skills/tree/main/skills/pr-babysitter into .cursor/skills/pr-babysitter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pr-babysitter", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mblode/agent-skills.git --path skills/pr-babysitter--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mblode/agent-skills --skill pr-babysitter -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mblode/agent-skills pr-babysitter --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mblode/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/pr-babysitter .gemini/skills/pr-babysitter && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pr-babysitter" agent skill from https://github.com/mblode/agent-skills/tree/main/skills/pr-babysitter into .gemini/skills/pr-babysitter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pr-babysitter", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mblode/agent-skills pr-babysitterInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mblode/agent-skills --skill pr-babysitter -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mblode/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/pr-babysitter .github/skills/pr-babysitter && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pr-babysitter" agent skill from https://github.com/mblode/agent-skills/tree/main/skills/pr-babysitter into .github/skills/pr-babysitter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pr-babysitter", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mblode/agent-skills --skill pr-babysitter -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mblode/agent-skills pr-babysitter --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mblode/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/pr-babysitter .opencode/skills/pr-babysitter && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pr-babysitter" agent skill from https://github.com/mblode/agent-skills/tree/main/skills/pr-babysitter into .opencode/skills/pr-babysitter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pr-babysitter", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pr-babysitterMonitors or repairs an open GitHub PR: CI failures, conflicts, review threads, and merge readiness, reporting state changes.
PR Babysitter is an agent skill from mblode/agent-skills. Monitors or repairs an open GitHub PR: CI failures, conflicts, review threads, and merge readiness, reporting state changes. Use when asked to "watch this PR", "fix CI", "resolve conflicts", or "address review comments".
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `evals/evals.json`, `references/bot-patterns.md` and `references/ci-platforms.md`). Compatibility notes: Requires a Git checkout, authenticated GitHub CLI, and jq. Continuous monitoring also needs a supported scheduler or event subscription.
It sits in Testing & QA, covering Failing and flaky tests. It works with GitHub. The repository describes itself as: Nobody ships AI slop on purpose. These skills make sure you don’t. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cef4cfa. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
ghgitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires a Git checkout, authenticated GitHub CLI, and jq. Continuous monitoring also needs a supported scheduler or event subscription.
From compatibility in the SKILL.md frontmatter.
PR Babysitter loads about 3.4k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 1,840 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from mblode/agent-skills at commit cef4cfa, republished under its MIT licence (© mblode). 1,840 words, ~3,402 tokens.
.claude/skills/pr-babysitter/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.pr-creator), reviewing or fixing the diff itself (tidy), or npm release PRs (autoship watches its own release CI; never babysit a release or Version Packages PR it drives).| Invocation | Mode |
|---|---|
| "babysit", "watch this PR", "monitor", "keep it green" | Monitor: Phase 1 once, then phases 2-5 on every event or tick |
| "fix CI", "why is CI red", "loop on CI" | One-shot Phase 3 loop |
| "resolve conflicts", "rebase onto main", "update the branch" | One-shot Phase 2 |
| "address the comments", "reply to the reviewers", "triage review comments" | One-shot Comment Triage Workflow |
| "is it ready", "what is blocking the merge" | One-shot Phase 5 report |
Standing rules, every mode:
Monitoring or fixing code does not by itself authorize posting replies. Post, resolve threads, or request reviews only when the user authorized that communication; otherwise prepare replies and report them.
Resolve scripts/fetch-comments.sh relative to this installed SKILL.md. ${CLAUDE_SKILL_DIR} below is a Claude Code adapter, not a portable environment variable.
No setup questions. Auto-detect the PR, the CI platforms, and the defaults (poll every 2 minutes, auto-resolve noise, no auto-merge), then start. Overrides arrive inline: "poll every 5 minutes", "enable auto-merge".
Skip closed or merged PRs. Skip drafts (isDraft) unless asked.
Comment triage runs autonomously; the plan file is an audit trail, not an approval gate.
Speak only on transitions. A quiet poll says nothing.
| File | Read when |
|---|---|
references/monitoring-setup.md | Phase 1 and Stopping: watch ladder, Monitor watch script, cron fallback, state file format, defaults, stop and lifecycle |
references/merge-conflicts.md | Phase 2: mergeStateStatus table, rebase workflow, lockfile and generated-file resolution, abort criteria |
references/ci-platforms.md | Phase 3: gh pr checks fields and exit codes, per-platform log and retry commands, Buildkite auth chain, failure classification |
scripts/fetch-comments.sh | Comment triage: run ${CLAUDE_SKILL_DIR}/scripts/fetch-comments.sh {N} first. One JSON document of every review, thread, and issue comment; --help prints the output shape |
references/github-api.md | Comment triage: script output contract, thread accounting, anchor ladder, awaiting-reply rule, reply and resolve |
references/bot-patterns.md | Comment triage: reviewer detection, severity mapping, merge gates, noise markers, dedup, false positives |
references/fix-plan-template.md | Comment triage: plan file format and the legal ignore reasons |
references/verification-gate.md | Before any commit: lint, type-check, test, knip, stray-artifact sweep |
references/git-resilience.md | A git command hangs or fails transiently (fsmonitor wedge, stale index.lock, network blip) |
evals/evals.json | Only when changing this skill; never during a PR task |
Phase 1 runs once in the foreground and starts the watch. Every event or tick then runs phases 2-5, diffs against the state file, and speaks only when something changed.
Copy this checklist to track progress:
PR babysit progress:
- [ ] Phase 1: Initialize (detect PR, pick watch mechanism, snapshot state)
- [ ] Phase 2: Conflict check
- [ ] Phase 3: CI check (diagnose, fix, gate, push)
- [ ] Phase 4: Comment check (triage new comments)
- [ ] Phase 5: Readiness check (report transitions, write state file)gh pr view [N] --json number,url,title,state,isDraft,headRefName,baseRefName,headRefOid,mergeable,mergeStateStatus,reviewDecision. No PR for the branch: say so and stop.gh repo view --json owner,name for the calls that need owner/repo.gh pr checks --json name,link (pattern table in references/ci-platforms.md).references/monitoring-setup.md that applies (harness PR subscription, Monitor tool, cron, none). With no rung, do not claim monitor mode: run the matching one-shot mode or say this runtime cannot keep polling..claude/pr-babysitter/babysit-pr-{N}.md: mechanism and ID, head SHA, mergeability, check states, open and awaiting-reply thread counts, review decision. This folder is never staged.gh pr view --json mergeable,mergeStateStatus. DIRTY resolves; BEHIND updates; UNKNOWN means GitHub is still computing, recheck next tick; anything else moves on.
git fetch origin {base} && git rebase origin/{base}
git push --force-with-lease --force-if-includesgit rebase --abort, then notify with the files and what each side changed. Human intent decides those.Bare --force is never used. A refused lease means someone else pushed: abort and notify rather than overwrite their commits. More than one author on the branch means a rebase rewrites their commits: merge origin/{base} instead.
gh pr checks --json name,state,bucket,link,workflow. bucket is pass, fail, pending, skipping, or cancel; link is the details URL.pending: wait. Diagnosing a half-finished run fixes the wrong thing.fail: fetch logs with the platform's commands in references/ci-platforms.md (GitHub Actions, Buildkite auth chain, Vercel, Fly.io).knip (delete dead code or configure), infrastructure (notify; not fixable from code).One-shot loop ("fix CI"): after each push, gh pr checks --watch --fail-fast (exit 0 green, 1 a check failed, 8 still pending). Stop and summarize when checks are green, the failure is infrastructure, or the same check fails twice with the same error after a fix. Two identical failures is the signal to stop pushing, not to try a third variant.
updated_at across review and issue comments against the state file. An edited-in-place bot comment and a reply on a resolved thread both have to register.mergeable == MERGEABLE, every required check pass, reviewDecision == APPROVED from a review whose commit_id is the head SHA, zero open blocking threads, zero threads awaiting my reply, every merge gate satisfied.gh pr merge --auto with the repo's merge method, and only when the user opted in.Inline from Phase 4 or one-shot. Autonomous: no approval gate, the plan file is the audit trail.
Run ${CLAUDE_SKILL_DIR}/scripts/fetch-comments.sh {N}. It pages every thread and every thread's comments, recovers anchors, buckets threads, and computes owedReply against your own login. A non-zero exit prints one sentence on stderr saying why; fix that cause and re-run rather than hand-writing queries.
Check reviewers[] before classifying: every login that spoke must appear in the output with findings, a verdict, or an explicit "no content". A reviewer with reviews but zero comments is a fetch that lost something. anchor.source == "needs-translation" means finish the anchor ladder against the working tree before judging that finding.
Early exit only when open threads, awaiting-reply threads, actionable reviews, and actionable issue comments are all zero.
github-actions[bot] is shared by reviewers and noise alike.references/fix-plan-template.md) to .claude/pr-babysitter/pr-{N}-review-plan.md, print the counts, proceed."Stop babysitting" or "cancel the monitor": stop the mechanism recorded in the state file per references/monitoring-setup.md (Stopping, Session Lifecycle), then report polls or events handled, fixes applied, conflicts resolved, comments triaged, and current state.
isResolved == false: GitHub collapses resolved threads, so a human reply posted after the resolve is the comment most likely to go unread.comments(first: 20): thread comments arrive oldest first, so a truncated page hides the newest comment, the one that decides whether you owe a reply. The script pages every thread to the end.viewerDidAuthor returned false on the viewer's own comments. Compare author.login to gh api user --jq .login.line means outdated or multi-line, not PR-level. Only a null path is PR-level. Recover the anchor before deciding anything.id, so a state diff on ids sees nothing. Compare updated_at.commit_id is not the head SHA as an approval: branch protection with "dismiss stale approvals" drops it on the next push, and the PR reads ready until then.git add -A after a fix commits hook artifacts (a root schema.gql) into the PR. Sweep git status --porcelain and stage paths.pr-creator: opens or edits the PR; babysitting starts after it existsplanning: writes plans a fresh session executes. The fix plan this skill writes is an audit trail for one PR, not a planning deliverabletidy: reviews and fixes the diff itself; this skill applies GitHub review comments. Run it on monitor-authored fixes beyond a trivial patchautoship: npm release pipelines; it watches its own release CI, so never babysit a release PR it drives© mblode, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (scripts, references) in skills/pr-babysitter of mblode/agent-skills.
Open the folder on GitHubat commit cef4cfa
PR Babysitter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PR Babysitter this skillmblode/agent-skills | 143 | — | ~3.4k | Automated safety check: Pass | MIT | |
| Triage CI Flakepayloadcms/payload | 45k | — | ~4.4k | Automated safety check: Pass | MIT | |
| GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb | 6.7k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| Fix Issueremix-run/remix | 33k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Issue To Regression Testbrunosabot/streamline-card | 269 | — | ~529 | Automated safety check: Pass | MIT | |
| Detect Flaky Testsagent-substrate/substrate | 4.5k | — | ~3k | Automated safety check: Pass | Apache-2.0 |
payloadcms/payload
A skill your agent uses when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests
GreptimeTeam/greptimedb
Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.
remix-run/remix
Fix a reported issue in Remix from a GitHub issue. An agent skill from remix-run/remix.
brunosabot/streamline-card
A skill your agent uses when the user asks to fix a bug, references a GitHub issue number, or describes an issue and wants a fix.
agent-substrate/substrate
Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM…
dotnet/macios
Investigate and fix flaky/random CI test failures in dotnet/macios.
mblode/agent-skills
Implements agent-readiness on public sites and docs from Mintlify Agent Score, AFDocs, Is Agentic, Is It Agent Ready, or url-discovery-bench reports, or from server logs of agents 404ing on guessed…
mblode/agent-skills
Creates and improves portable Agent Skills with a validator, routing scenarios, and evidence-based keep, cut, merge, or retire decisions.
mblode/agent-skills
Recovers decisions, previous fixes, research, and what followed a prompt from past AI conversations, with source evidence.
mblode/agent-skills
Cuts the wait from push to green by measuring a pipeline's critical path from run timestamps, then splitting, sharding, trimming setup and sharing test module state, with a before/after ledger.
mblode/agent-skills
Builds and maintains a repo's own verification harness (verify CLI, doctor, worktree isolation, feature map, seed data) and a reproduce-first bug handoff.
mblode/agent-skills
Builds, reviews, and measures UI motion, including springs, gestures, scroll effects, curve fitting from recordings, and sparse interface sound.
Works with
Categories
Monitors or repairs an open GitHub PR: CI failures, conflicts, review threads, and merge readiness, reporting state changes. PR Babysitter is an agent skill from mblode/agent-skills. Monitors or repairs an open GitHub PR: CI failures, conflicts, review threads, and merge readiness, reporting state changes.
PR Babysitter fits situations like: asked to watch this PR; resolve conflicts; address review comments.
Run `npx skills add mblode/agent-skills --skill pr-babysitter -a claude-code`. Or copy the skill folder (skills/pr-babysitter in mblode/agent-skills) into .claude/skills/pr-babysitter in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mblode/agent-skills --skill pr-babysitter -a codex`. Or copy the skill folder (skills/pr-babysitter in mblode/agent-skills) into .agents/skills/pr-babysitter in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mblode/agent-skills --skill pr-babysitter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pr-babysitter, .gemini/skills/pr-babysitter, .github/skills/pr-babysitter and .opencode/skills/pr-babysitter in your project.
Going by SKILL.md and its folder, PR Babysitter needs a shell for the scripts in its folder and the command-line tools its instructions call (gh and git). Our summary lists: A Bash shell. Compatibility (from SKILL.md): Requires a Git checkout, authenticated GitHub CLI, and jq. Continuous monitoring also needs a supported scheduler or event subscription..
SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
PR Babysitter is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with PR Babysitter: Triage CI Flake (payloadcms/payload, 45k stars), GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars), Fix Issue (remix-run/remix, 33k stars) and Issue To Regression Test (brunosabot/streamline-card, 269 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mblode (a GitHub user) maintains it in mblode/agent-skills, which has 143 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 6, 2026.
Source: mblode/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.