Agent skill

Babysit

by nubjs in nubjs/nub

Bring one nubjs/nub pull request to merge-readiness — pull the inline reviews, verify each finding against the code, re-verify every fix round locally, fix CI, and loop.

MITAuto-check passedDevelopment

Install Babysit

skills CLI
$ npx skills add nubjs/nub --skill babysit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nubjs/nub babysit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nubjs/nub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/babysit .claude/skills/babysit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
babysit
GitHub stars
4.4k
Token cost
~1.9k tokens
SKILL.md length
897 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Bring one nubjs/nub pull request to merge-readiness — pull the inline reviews, verify each finding against the code, re-verify every fix round locally, fix CI, and loop.

  • Tasks that involve Pull requests
  • SKILL.md covers Reviews, Triage, Fix rounds and CI, plus 2 more sections
  • Calls gh, git and make
  • Tasks that involve Failing and flaky tests

What it does

Babysit is an agent skill from nubjs/nub. Bring one nubjs/nub pull request to merge-readiness — pull the inline reviews, verify each finding against the code, re-verify every fix round locally, fix CI, and loop. Invoke (via the Skill tool) when asked to babysit a PR, to get one merge-ready, or when you own a PR and are waiting on reviews or CI. Scope — babysit ENDS at merge-readiness and hands back; only an instruction to babysit it TO MERGE authorizes the merge, and that phrasing also delegates the readiness judgment, so decide rather than ask. Carries…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Pull requests and Failing and flaky tests. It works with GitHub. The repository describes itself as: The fast all-in-one Node.js toolkit. The licence is MIT.

When your agent uses it

  • Tasks that involve Pull requests
  • Tasks that involve Failing and flaky tests

Example prompts

  • “/babysit”

Requirements

  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 568e73a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git
    • make
    • jq
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, git and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Babysit loads about 1.9k tokens when it runs. Until then it costs about 224 tokens; SKILL.md has 897 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~224
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nubjs/nub at commit 568e73a, republished under its MIT licence (© nubjs). 897 words, ~1,866 tokens.

Download SKILL.mdSave it as .claude/skills/babysit/SKILL.md (or your agent's skills folder).
name
babysit
description
Bring one nubjs/nub pull request to merge-readiness — pull the inline reviews, verify each finding against the code, re-verify every fix round locally, fix CI, and loop. Invoke (via the Skill tool) when asked to babysit a PR, to get one merge-ready, or when you own a PR and are waiting on reviews or CI. Scope — babysit ENDS at merge-readiness and hands back; only an instruction to babysit it TO MERGE authorizes the merge, and that phrasing also delegates the readiness judgment, so decide rather than ask. Carries the traps that each cost a round trip: the GitHub CLI's PR-view JSON silently omits the inline review comments where the real findings live, the CI gate aggregator check registers LAST so a partial green rollup is not done, and PR CI is opt-in, so a pull request nobody labelled has no checks at all and a re-based or re-targeted branch keeps the checks of its old head.

Babysit a PR (nubjs/nub)

Loop: pull reviews → verify each finding → fix → re-verify locally → push → repeat. CI comes ONCE, at the end: PR CI is opt-in, so a push starts nothing and the run is requested with gh pr edit <n> --add-label ci when the head is final.

Babysit ends at merge-ready. Report that state and stop; the merge is the maintainer's. Only an instruction to babysit it to merge authorizes one — and that phrasing delegates the readiness judgment too, so decide against the criteria below instead of asking.

Reviews

Two bots review within minutes of every push: pullfrog, which posts a top-level review plus inline threads with a <details> write-up per finding, and copilot-pull-request-reviewer. The maintainer outranks both. A PR carrying no review was merged before they ran rather than approved by them, and only the review list distinguishes those.

The gh pr view --json reviews,comments form silently omits inline pulls/{n}/comments, which is where the actionable findings are. Use scripts/pr-reviews.ts — one GraphQL round trip for reviews, inline threads (id, resolved/outdated state, file:line) and PR conversation:

bash
nub scripts/pr-reviews.ts <pr> | jq '.reviewThreads[] | select(.isResolved | not)'

Triage

Skip resolved and outdated threads first — an outdated one anchors a line a later push already rewrote.

Every finding is a hypothesis. Verify it against the code before acting; a confident write-up is the one that gets acted on unchecked. Null guards, "just in case" branches, and try/catch around code that cannot throw are the slop shape AGENTS.md names — reject those by default.

Reject with one factual line on the thread saying what you checked. Never ignore silently, and never reply conversationally to a bot. Load the prose-writing skill before writing any comment. Resolve each thread you address:

bash
gh api graphql -f query='mutation($id:ID!){resolveReviewThread(input:{threadId:$id}){thread{isResolved}}}' -F id=<thread-id>

Fix rounds

A review-driven fix earns the same gate as the original change. Run make verify. For a behavior change, sweep adversarial fixtures against the branch's MERGE-BASE build (ad-hoc-test) — a run that only re-confirms the reported case tests your intent, not your fix. Escalate to the impact-analysis skill when the fix widened the blast radius. A one-line typo fix skips both.

Then push. Pushing is free — PR CI is opt-in and nothing fires until it is asked for — so push each round as it lands and hold the CI request until the review loop is done. A requested run is roughly five workflows across eight platforms, and asking round after round is what starves the shared runner pool.

CI

Look at gh pr checks <N> before asking. The label is stripped within a minute of being added, so an empty label list says nothing about whether a run was requested, and a second request for the same head cancels the run already in flight — a full matrix thrown away (2026-09-19, #959: the maintainer had labelled it half an hour earlier, and the babysitter's request cancelled fifteen jobs). Checks already pending or running on the current head mean the run exists; watch that one.

Ask for the run, then watch it. Both steps — the request is not optional, and until it lands the pull request has no checks at all:

bash
gh pr edit <N> --add-label ci      # requests CI against the current head; the label is removed again immediately
nub scripts/ci-watch.ts --pr <N> --required "CI gate" --timeout 90

Push a further commit after that and the run you were watching is stale — its checks belong to the old head, so ask again. ci-watch will not cover for a forgotten request: an empty or all-skipped rollup reports as pending, never green.

Show full SKILL.md (352 more words)Show less

The CI gate job is an if: always() aggregator over every path-gated job, so it registers LAST: fifty green checks with it still pending means the run is not done. A raw gh pr checks --watch exits 0 while the run is queued with no jobs registered, so watch through scripts/ci-watch.ts.

Diagnose red with ci-triage; a run whose rollup is cancelled can hold a job that failed, so the run-level conclusion misses real red.

Fix only what this PR caused, and never edit a workflow to make a check pass. A merge-blocking failure that looks unrelated is often a stale base — merge main in and re-run. The heavy legs fail for environment reasons too, so read the job log before calling one a regression:

Docker smoke · Windows matrix · daily-driver · native-deps

Stacked PR. After a rebase plus gh pr edit --base main, a LOW check count is a FALSE green. The --base edit fires no synchronize event, so every workflow gated on pull_request: branches: [main] never runs. Force the real matrix with gh pr close <N> && gh pr reopen <N>, then wait for it — a Rust PR carries around sixty checks.

Merge-ready

  • Mergeable: no conflicts, base current.
  • The CI gate check is green, or every failure is investigated and written up as unrelated.
  • Every valid finding is addressed, or rejected with reasoning on the thread.
  • Threads resolved.
  • The body carries Closes #N if the PR resolves an issue. Without it the issue stays silently open after the merge; gh pr edit <N> --body fixes it.

Report that state and stop here, unless the instruction was to merge.

Merging

The Main ruleset blocks update on main, so a plain merge fails with the base branch policy prohibits the merge, and auto-merge is disabled repo-wide. That ruleset sets no required status checks either, so nothing server-side stops a red merge. The flag below bypasses the protection rule, not the judgment — confirm CI gate yourself first.

bash
gh pr merge <N> --squash --admin
git -C <shared-tree> fetch origin && git -C <shared-tree> merge --ff-only origin/main  # never pull --ff-only
make install-dev                                                                       # skip only if the diff was docs-only
git push origin --delete <branch> && git branch -D <branch>
git worktree remove <dir> --force

The shared tree always carries a sibling's uncommitted work, so pull --ff-only runs rebase's precondition check and aborts on any dirty file. For a queue of approved PRs, scripts/merge-cascade.ts does all of the above per entry.

© nubjs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/babysit of nubjs/nub.

Open the folder on GitHubat commit 568e73a

Compare with similar skills

Babysit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Babysit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Babysit this skillnubjs/nub4.4k—~1.9kAutomated safety check: PassMIT
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Exposed Bug Fix WorkflowJetBrains/Exposed9.3k—~3.8kAutomated safety check: PassApache-2.0
GitHub PR Imagesbikeindex/bike_index308—~1.8kAutomated safety check: PassAGPL-3.0
Debug Os Failure On GitHubstrands-agents/box110—~1.3kAutomated safety check: NotesApache-2.0
Gh Address Commentstech-leads-club/agent-skills7k—~353Automated safety check: PassApache-2.0

Similar skills

  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Exposed Bug Fix Workflow

    JetBrains/Exposed

    Official

    Takes a GitHub or YouTrack issue for the Exposed project through reproduction, a failing test, a fix, validation and a pull request.

    9.3k GitHub stars~3.8k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • GitHub PR Images

    bikeindex/bike_index

    Embed a local image file into an existing GitHub PR — either in the PR body or as a comment.

    308 GitHub stars~1.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Debug Os Failure On GitHub

    strands-agents/box

    Debug a CI failure on an OS you are not on (you are on Linux, it fails on macos-latest, or the reverse) without opening a pull request per attempt.

    110 GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check: notes
  • Gh Address Comments

    tech-leads-club/agent-skills

    Address review and issue comments on the open GitHub PR for the current branch using gh CLI.

    7k GitHub stars~353 tokensUpdated 18 days ago
    DevelopmentAuto-check passed
  • GitHub Gh CLI

    huangruiteng/CS-Notes

    Inspect GitHub pull requests, issues, workflow runs, and API data with the gh CLI.

    4k GitHub stars~377 tokensUpdated 2 days ago
    DevelopmentAuto-check passed

More from nubjs/nub

All 31 skills in this repo
  • Cpu Reduction

    nubjs/nub

    Diagnose and clear CPU, memory, and disk contention on the maintainer's dev host.

    4.4k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Reclaim disk on the maintainer's Mac when the volume is full or filling — ENOSPC, "no space left on device", a failed build or agent harness, or a routine sweep of Rust build residue.

    4.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Nub Charts

    nubjs/nub

    Build a performance chart for nubjs.com — the SVG bar figures in blog posts, docs pages and social posts (a runtime augmentation against plain node, an install or dispatch comparison, a cross-tool…

    4.4k GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Audit Thread

    nubjs/nub

    A skill your agent uses when running a compatibility/parity AUDIT — enumerating where nub diverges from a reference it claims parity with (pnpm CLI grammar, a lockfile format, a Node behavior, a…

    4.4k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Linux Vm Test

    nubjs/nub

    Run ad-hoc Nub tests and debugging probes on real local Linux guests.

    4.4k GitHub stars~986 tokensUpdated yesterday
    Auto-check passed
  • Performance-trace Nub package-manager installs using the existing phase timings, structured diagnostics, and sampling-profiler workflow.

    4.4k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Babysit

What does Babysit do?

Bring one nubjs/nub pull request to merge-readiness — pull the inline reviews, verify each finding against the code, re-verify every fix round locally, fix CI, and loop. Babysit is an agent skill from nubjs/nub. Bring one nubjs/nub pull request to merge-readiness — pull the inline reviews, verify each finding against the code, re-verify every fix round locally, fix CI, and loop.

When should I use Babysit?

Babysit fits situations like: tasks that involve Pull requests; tasks that involve Failing and flaky tests.

How do I install Babysit in Claude Code?

Run `npx skills add nubjs/nub --skill babysit -a claude-code`. Or copy the skill folder (.claude/skills/babysit in nubjs/nub) into .claude/skills/babysit in your project. Claude Code loads it when a task matches its description.

How do I install Babysit in Codex?

Run `npx skills add nubjs/nub --skill babysit -a codex`. Or copy the skill folder (.claude/skills/babysit in nubjs/nub) into .agents/skills/babysit in your project. Codex loads it when a task matches its description.

Can I use Babysit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nubjs/nub --skill babysit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/babysit, .gemini/skills/babysit, .github/skills/babysit and .opencode/skills/babysit in your project.

What does Babysit need to run?

Going by SKILL.md and its folder, Babysit needs the command-line tools its instructions call (gh, git, make, jq and docker). Our summary lists: Docker.

Does Babysit access the network?

SKILL.md contains no URLs. Its commands use gh, git and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Babysit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Babysit use?

Babysit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Babysit use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Babysit?

Skills that share tags, products or a category with Babysit: PR Babysitter (openinterpreter/openinterpreter, 69k stars), Exposed Bug Fix Workflow (JetBrains/Exposed, 9.3k stars), GitHub PR Images (bikeindex/bike_index, 308 stars) and Debug Os Failure On GitHub (strands-agents/box, 110 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Babysit?

nubjs (a GitHub organization) maintains it in nubjs/nub, which has 4,372 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: nubjs/nub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.