Agent skill

Agent Watchdog

by BuilderIO in BuilderIO/skills

A skill your agent uses when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log…

MITAuto-check passed

Install Agent Watchdog

skills CLI
$ npx skills add BuilderIO/skills --skill agent-watchdog -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BuilderIO/skills agent-watchdog --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BuilderIO/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-watchdog .claude/skills/agent-watchdog && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-watchdog
GitHub stars
4.5k
Token cost
~2k tokens
SKILL.md length
1,007 words
Files
3
Skills in repo
25
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log…

  • Works in 4 steps: Identify every artifact the user… → Use the host's native thread/history… → If the artifact is still running and the… → …
  • Fix another agents work from a Codex session ID
  • SKILL.md covers Choose The Mode, Resolve The Target, Reconstruct The Contract and Investigate Independently, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Watchdog is an agent skill from BuilderIO/skills. Use when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log, or pasted run summary. Monitor until the other agent is done or blocked, reconstruct what the user asked, independently investigate the same problem to form your own hypotheses and approach, inspect what the agent actually changed and verified, compare the two investigations, report gaps and add-on directions, and optionally make scoped…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `README.md` and `agents/openai.yaml`).

The licence is MIT.

When your agent uses it

  • Fix another agents work from a Codex session ID
  • Claude Code session/transcript
  • Chat/thread link
  • Pasted run summary

Example prompts

  • “/agent-watchdog”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Identify every artifact the user supplied: session ID, transcript path,
  2. Use the host's native thread/history tools, local transcript files, repo
  3. If the artifact is still running and the user asked to watch, poll at a
  4. If the artifact cannot be resolved, ask for the missing identifier or path.

What it can do on your machine

Read from SKILL.md and the folder at commit 530d9ee. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Watchdog loads about 2k tokens when it runs. Until then it costs about 143 tokens; SKILL.md has 1,007 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~143
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from BuilderIO/skills at commit 530d9ee, republished under its MIT licence (© BuilderIO). 1,007 words, ~1,952 tokens.

Download SKILL.mdSave it as .claude/skills/agent-watchdog/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
agent-watchdog
description
Use when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log, or pasted run summary. Monitor until the other agent is done or blocked, reconstruct what the user asked, independently investigate the same problem to form your own hypotheses and approach, inspect what the agent actually changed and verified, compare the two investigations, report gaps and add-on directions, and optionally make scoped fixes when the user authorizes repair.

Agent Watchdog

Watch another agent's work like a reviewer with a pager who is also a second investigator: wait for completion when needed, reconstruct the request, run your own independent investigation of the same problem, verify the evidence, and close the gap between what was asked and what actually happened. You are not just grading their homework — you are a second solver whose findings get diffed against theirs so nothing is missed.

Choose The Mode

Infer the mode from the user's wording:

  • Watch only: monitor a session, PR, branch, CI run, or transcript until it reaches a terminal state. Do not edit files.
  • Audit: read the prompt, transcript, diff, tests, CI, comments, screenshots, or final claims, run the independent investigation below, and return a gap report plus add-on directions. Do not edit files.
  • Audit and fix: audit first, then make narrow fixes for clear gaps. Avoid broad rewrites, branch movement, or speculative changes.
  • Compare: when given multiple sessions or agents, compare their work against the same original request and reconcile the important differences.

If authority is unclear, default to audit-only and say what you would fix.

Resolve The Target

  1. Identify every artifact the user supplied: session ID, transcript path, thread URL, PR, branch, commit, CI run, issue, Slack link, or pasted summary.
  2. Use the host's native thread/history tools, local transcript files, repo logs, GitHub tools, or pasted content to resolve the artifact. Prefer the most direct source over summaries.
  3. If the artifact is still running and the user asked to watch, poll at a reasonable interval until it is done, blocked, stale, or clearly waiting on a human/external system.
  4. If the artifact cannot be resolved, ask for the missing identifier or path.

Reconstruct The Contract

Build a compact contract before judging the work:

  • Original user request and any later changes in scope.
  • Explicit constraints: branch rules, no-edit requests, deadlines, package versions, validation expectations, design requirements, or security/privacy limits.
  • Implied acceptance criteria: user-visible behavior, tests, CI, docs, deploys, screenshots, review replies, or status updates.
  • The other agent's final claims and any "could not do" caveats.

Treat the user's request as the source of truth, not the other agent's summary.

Investigate Independently

Act as if the original prompt had been given to you, in parallel with auditing the watched agent. Anchoring is the failure mode: reading their work first and nodding along. Your value comes from a genuinely separate second pass.

  1. Form your own hypotheses about root causes and the approach you would take — ideally before reading the watched agent's conclusions, and if you have already seen them, still reason from first principles rather than from their framing.
  2. Explore the code, data, logs, production state, and docs yourself, directly or via subagents. While the watched agent is still running, use the wait to pre-map the problem domains so hypotheses are ready before their diff lands.
  3. Prioritize evidence the watched agent may not have looked at: production run ledgers or databases, session replays, error trackers, user-supplied screenshots, deploy/version state, other worktrees, memory of past incidents in the same area.
  4. Verify the watched agent's key claims against primary sources, and verify your own leads the same way — reopen the cited files and line refs before asserting either side is right. Subagent reports are leads, not facts.
  5. Diff the two investigations: what they found that you missed, what you found that they missed, where the approaches diverge, and any product or design fork they took silently. Convert the diff into concrete, actionable add-on directions (exact files, guards, test cases) — not vague concerns — and offer them as a paste-ready note the user can relay.
Show full SKILL.md (400 more words)Show less

Relay Sparingly

Finding something is not a reason to send it yet. A watched agent that is still building re-plans around every note it receives, so the cost of a relay is their attention and their sequencing, not your tokens.

  • Interrupt a live run only for what changes their current step: a defect in code they are touching now, a correction to something you told them earlier, or an answer they are blocked on.
  • Everything else — coverage gaps, reference material, conventions, polish — waits for their next checkpoint and travels as one batched note.
  • Never hand a mid-run agent an unranked list of what is missing. Ranked, and labelled as a map rather than a to-do list, or not at all: an unranked backlog reliably produces several half-finished surfaces instead of a few complete ones.
  • Say what you verified as correct, not only what is wrong. It stops them re-litigating settled ground and keeps the relationship peer-to-peer.
  • If the watched agent has not yet acted on your previous note, do not send another one.

Audit The Evidence

Inspect evidence, not vibes:

  • Read changed files and relevant unchanged files around the touched paths.
  • Check git status/diff without reverting unrelated work.
  • Compare commands the agent claimed to run with actual output when available.
  • Inspect failed or skipped tests, CI logs, browser screenshots, review comments, deploy output, and error traces.
  • For PR/review work, verify unresolved threads and CI state from the source system when tools are available.
  • For UI work, prefer screenshots or browser checks over prose claims.

Classify each issue as:

  • Gap: requested behavior is missing or incomplete.
  • Bug: the implementation likely fails or regresses behavior.
  • Verification miss: the work may be right but the evidence is weak.
  • Scope drift: the agent changed something unrelated or skipped a constraint.
  • No issue: the concern is already handled, with evidence.

Fix Narrowly

When the user authorized repair:

  1. Fix only gaps with clear evidence.
  2. Preserve unrelated local changes and do not move branches unless explicitly asked for that branch operation.
  3. Use existing repo patterns and targeted tests.
  4. Re-run the smallest useful validation after each meaningful fix.
  5. If a fix would require a product decision, credential, destructive action, or broad rewrite, stop and report the decision instead of guessing.

Report

Lead with the outcome. Keep the report short enough to scan:

md
Status
- Done, blocked, stale, or still running.

Requested
- What the user asked the watched agent to do.

Observed
- What the watched agent changed, claimed, and verified.

Gaps
- Missing behavior, bugs, weak verification, or scope drift.

Independent findings
- What your own investigation surfaced that the watched agent missed, where
  your approach would have differed, and concrete add-on directions to relay.

Fixes made
- Files changed and validation run. Omit this section for audit-only work.

Remaining risk
- Anything still unverified or waiting on CI/review/deploy/human input.

Name exact files, commands, PRs, or thread IDs when they matter.

© BuilderIO, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/agent-watchdog of BuilderIO/skills.

  • SKILL.md
  • README.md
  • agents/openai.yaml

Open the folder on GitHubat commit 530d9ee

Compare with similar skills

Agent Watchdog next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Watchdog compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Watchdog this skillBuilderIO/skills4.5k—~2kAutomated safety check: PassMIT
Babysitsimstudioai/sim30k—~2.9kAutomated safety check: PassApache-2.0
Watchdogmvschwarz/openrig5.5k—~1.5kAutomated safety check: PassApache-2.0
Canary Watchaffaan-m/ECC274k1 repos~770Automated safety check: PassMIT
Canary Watchaffaan-m/ECC274k—~405Automated safety check: PassMIT
Canary Watchaffaan-m/ECC274k—~414Automated safety check: PassMIT

Similar skills

  • Babysit

    simstudioai/sim

    Drive a PR to a clean review (Greptile 5/5, zero open threads) — ships if needed, keeps it mergeable against staging, re-triggers both Greptile and cubic, fixes real findings, replies to and…

    30k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Watchdog

    mvschwarz/openrig

    A skill your agent uses when configuring rig watchdog policies, authoring wake/refocus/alignment-checkpoint messages, or choosing the right intervention level for a stale-owner situation.

    5.5k GitHub stars~1.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Canary Watch

    affaan-m/ECC

    A skill your agent uses to monitor and verify a deployed URL after releases — checks HTTP endpoints, SSE streams, static assets, console errors, and performance regressions after deploys, merges, or…

    274k GitHub starsUsed in 1 repo~770 tokens
    DevOps & CloudAuto-check passed
  • Canary Watch

    affaan-m/ECC

    このスキルを使用して、デプロイメント、マージ、または依存関係アップグレード後にデプロイされたURLの回帰を監視します. An agent skill from affaan-m/ECC.

    274k GitHub stars~405 tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Canary Watch

    affaan-m/ECC

    使用此技能在部署、合并或依赖升级后监控已部署的URL是否存在回归问题。

    274k GitHub stars~414 tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Babysit

    nubjs/nub

    Bring one nubjs/nub pull request to merge-readiness — pull the inline reviews, verify each finding against the code, re-verify every fix round locally, fix CI, and loop.

    4.4k GitHub stars~1.9k tokensUpdated today
    DevelopmentAuto-check passed

More from BuilderIO/skills

All 25 skills in this repo
  • Plan Arbiter

    BuilderIO/skills

    A skill your agent uses when asked to compare, cross-review, merge, judge, choose, or arbitrate competing plans from multiple agents such as Codex and Claude Code; when given two or more proposed…

    4.5k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Plow Ahead

    BuilderIO/skills

    A skill your agent uses when the user explicitly wants autonomous progress without routine clarification stops: "plow ahead", "do not stop", "use your best judgment", "keep going until done"…

    4.5k GitHub stars~973 tokensUpdated today
    Auto-check passed
  • Read The Damn Docs

    BuilderIO/skills

    A skill your agent uses when implementing, integrating, upgrading, debugging, or answering anything involving third-party APIs, libraries, frameworks, CLIs, cloud services, model/provider SDKs…

    4.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Stay Within Limits

    BuilderIO/skills

    A skill your agent uses when long-running or parallel agent work must respect 5-hour and weekly usage limits by checking usage between waves, pausing near the cap, and resuming only when the window…

    4.5k GitHub stars~857 tokensUpdated today
    Auto-check passed
  • Efficient Fable

    BuilderIO/skills

    A skill your agent uses when running Claude Fable on codebase-heavy or token-heavy work and the user wants Fable to orchestrate research, coding, and testing while cheaper subagents do bounded heavy…

    4.5k GitHub stars~996 tokensUpdated today
    Auto-check passed
  • Quick Recap

    BuilderIO/skills

    A skill your agent uses when adding or following the red/yellow/green final status block convention for agent responses, especially by installing managed AGENTS.md or CLAUDE.md instructions.

    4.5k GitHub stars~352 tokensUpdated today
    Auto-check passed

Questions about Agent Watchdog

What does Agent Watchdog do?

A skill your agent uses when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log…. Agent Watchdog is an agent skill from BuilderIO/skills. Use when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log, or pasted run summary.

When should I use Agent Watchdog?

Agent Watchdog fits situations like: fix another agents work from a Codex session ID; Claude Code session/transcript; chat/thread link; pasted run summary.

How do I install Agent Watchdog in Claude Code?

Run `npx skills add BuilderIO/skills --skill agent-watchdog -a claude-code`. Or copy the skill folder (skills/agent-watchdog in BuilderIO/skills) into .claude/skills/agent-watchdog in your project. Claude Code loads it when a task matches its description.

How do I install Agent Watchdog in Codex?

Run `npx skills add BuilderIO/skills --skill agent-watchdog -a codex`. Or copy the skill folder (skills/agent-watchdog in BuilderIO/skills) into .agents/skills/agent-watchdog in your project. Codex loads it when a task matches its description.

Can I use Agent Watchdog in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BuilderIO/skills --skill agent-watchdog -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-watchdog, .gemini/skills/agent-watchdog, .github/skills/agent-watchdog and .opencode/skills/agent-watchdog in your project.

What does Agent Watchdog need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Watchdog is instructions for the agent only.

Does Agent Watchdog access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Watchdog safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Watchdog use?

Agent Watchdog is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Watchdog use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Watchdog?

Skills that share tags, products or a category with Agent Watchdog: Babysit (simstudioai/sim, 30k stars), Watchdog (mvschwarz/openrig, 5.5k stars), Canary Watch (affaan-m/ECC, 274k stars) and Canary Watch (affaan-m/ECC, 274k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Watchdog?

BuilderIO (a GitHub organization) maintains it in BuilderIO/skills, which has 4,528 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 7, 2026.

Source: BuilderIO/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.