Agent skill

Babysit Runs

by swyxio in swyxio/skills

Operate unattended runs through completion, from a single detached pilot to a large batch.

MITAuto-check passed

Install Babysit Runs

skills CLI
$ npx skills add swyxio/skills --skill babysit-runs -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swyxio/skills babysit-runs --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/babysit-runs .claude/skills/babysit-runs && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
babysit-runs
GitHub stars
176
Token cost
~2.9k tokens
SKILL.md length
1,425 words
Files
1
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Operate unattended runs through completion, from a single detached pilot to a large batch.

  • Supervised pilots
  • SKILL.md covers Choose the operating mode, Adopt, detach and schedule, Saturate useful work and Recover without losing…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Overnight runs and bounded autonomous optimization

What it does

Babysit Runs is an agent skill from swyxio/skills. Operate unattended runs through completion, from a single detached pilot to a large batch. Instrument execution, inspect traces and output quality, recover failures, and improve demonstrated performance or reliability problems using the existing workflow and temporary monitoring. Use for supervised pilots, overnight runs and bounded autonomous optimization, not a one-time status check or concurrency advice alone.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.

When your agent uses it

  • Supervised pilots
  • Overnight runs and bounded autonomous optimization
  • Not a one-time status check
  • Concurrency advice alone

Example prompts

  • “/babysit-runs”

What it can do on your machine

Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Babysit Runs loads about 2.9k tokens when it runs. Until then it costs about 107 tokens; SKILL.md has 1,425 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 1,425 words, ~2,870 tokens.

Download SKILL.mdSave it as .claude/skills/babysit-runs/SKILL.md (or your agent's skills folder).
name
babysit-runs
description
Operate unattended runs through completion, from a single detached pilot to a large batch. Instrument execution, inspect traces and output quality, recover failures, and improve demonstrated performance or reliability problems using the existing workflow and temporary monitoring. Use for supervised pilots, overnight runs and bounded autonomous optimization, not a one-time status check or concurrency advice alone.

Babysit Runs

Operate existing runners toward the authorized outcome at the agreed quality. Prefer targeted repairs over a new scheduler or pipeline rewrite. Use the selected provider/runtime skill when invocation or access needs troubleshooting; use live-ai-pipelines only when recovery architecture needs implementation.

Choose the operating mode

Completion mode: finish the authorized workload using the established implementation, repairing defects as needed.

Pilot-and-improve mode: exercise an existing or modified workflow on representative inputs, inspect execution and output quality, and make bounded improvements before broader rollout. State what is being tested, the quality baseline, resource/spending limits and the completion boundary. A pilot is complete when the requested behavior is demonstrated and material findings are resolved or clearly reported; optional optimization ideas do not keep it running indefinitely.

Adopt, detach and schedule

Inspect actual processes, owners, logs and retained results. Adopt a matching live run rather than launch a duplicate. Discover routine facts from existing metadata and task history instead of re-interviewing the user.

Keep a compact continuation brief in the existing run metadata or monitor prompt:

  • Goal, operating mode, scope, checkout, command, owner and run/log/artifact locations.
  • Requested model/provider, shared concurrency ceiling, independent transfer limits and applicable resource/cost bounds.
  • Agreed quality references, authorized repairs, completion evidence and release or handoff boundary.
  • Current progress, retry state, next recovery time and monitoring/reporting cadence.

Check changed or uncertain dependencies before submitting work: launcher permissions, authentication/model access, actual tool configuration and capacity. Reuse recent comparable successful evidence. Record the effective rules and versions where acceptance or replay depends on them; stale historical instructions must not override the current authorized contract. Do not silently substitute models.

Detach substantial runs using the existing manager or simplest reliable mechanism. Verify that execution survives the launching shell/tool and logs remain discoverable.

Create or update one temporary monitor. In Codex, use the automation tool and a heartbeat in the main thread, normally every 2–5 minutes during a pilot or unstable run; reuse a matching automation. Use another scheduler when requested or appropriate. Disclose unavailable scheduling rather than promise future check-ins. Make the saved prompt self-contained using the continuation brief, including intervention authority and stopping conditions, and update it when the run moves or resumes. Persist decisions so wakeups do not depend on conversation memory or a stale PID.

Each wakeup inspects actual ownership, progress, recent traces, errors and representative new results. Periodically use a stronger model within the authorized model and spending scope to spot-check substantive outputs and diagnose trace patterns, especially after model substitutions, implementation changes, repeated failures or unexpected slowdowns. Reuse prior assessments when their inputs are unchanged. Process health and valid schemas do not establish output quality.

Reduce monitoring frequency once execution is healthy. Notify the user of meaningful milestones, material regressions, required external action and completion; stay quiet when unchanged.

Saturate useful work

Fill slots with ready work and advance independent items without whole-batch barriers. Parallelize preparation, LLM calls, transfers and downstream stages where dependencies permit. Do not add inference or busywork merely to occupy capacity.

Apply the allowance across all runners sharing the same constrained resource, rather than giving each runner the full ceiling. Keep independently justified transfer/media/provider caps. Retain established healthy concurrency instead of repeating ramp experiments on every resume.

Identify the current bottleneck, commonly LLM calls or data transfer. Adjust parallelism using comparable accepted throughput, latency, retries, throttling and resource pressure. Downshift when useful performance degrades and recover capacity when healthy. Increase capacity at the demonstrated bottleneck; active-call count alone is not success.

Recover without losing successful work

Each check compares accepted/remaining counts, active/queued stages, oldest work, ownership, recent logs, deliveries and retry state with the previous observation. Quiet logs alone do not establish a stall; use representative durations and other progress signals.

Isolate failed items or surfaces while independent work continues. Before replacement, reconcile retained results and cancel or fence a stale owner. Drain affected active work before activating incompatible runner changes. Preserve successful stages and repair only the failed dimension.

  • Established intermittent failure: recover automatically with bounded backoff and jitter. A status code alone does not establish quota exhaustion or permanent denial.
  • Explicit authentication, entitlement or quota failure: isolate the affected surface and use normal authorized recovery. If external action is needed, report the concrete requirement.
  • Unknown delivery: inspect retained output, request IDs and available provider status/idempotency before resubmission. Preserve attempt history; missing local output does not establish that no result was produced.
  • Runner or content defect: fix the demonstrated cause and resume unfinished work, rather than regenerate accepted results.

Use bounded immediate retries, followed by persisted cooldowns for periodic outages. Respect existing budgets. Wakeups must not reset retry history. When repeated recovery produces no useful progress, diagnose or pause the affected work and report evidence instead of maintaining a hot retry loop. Routine recoverable failures should not require another approval interruption.

Show full SKILL.md (627 more words)Show less

Authority to intervene

Within the authorized outcome and resource/spending limits, the supervisor may autonomously pause admission, drain or stop affected work, repair or refactor the existing implementation, resume retained stages and increase useful concurrency. Routine diagnosis, targeted refactoring, recovery and concurrency adjustment are part of the delegated job and do not require repeated approval.

Intervene when evidence shows incorrect output, a stall, excessive latency, repeated unreliability or substantial wasted work. Preserve original evidence and accepted results, reconcile ambiguous deliveries, and change executable code only at safe boundaries.

Prefer one coordinated repair over repeated small interruptions. Record the finding, intervention and observed result. Stop repeating an intervention that produces no useful improvement. Do not weaken quality requirements, invent a replacement pipeline or expand the workload merely to keep workers busy.

Preserve quality; remove accidental gates

Use agreed examples and past transcripts to judge quality and unnecessary blockers. Local repairs may address runner bugs, contradictory stage rules, performance and finishing defects within the authorized goal. Throughput improvements must preserve substantive quality.

Separate content acceptance, execution validity and metadata completeness. Suggestions and model self-check flags do not become blockers without evidence of a defect. Retain required blockers for demonstrated unsupported claims, identity conflicts or substantive loss where those affect the task. Declared tool access should match actual execution; do not allow normal tool use and reject it downstream.

Sample after changes that could materially affect agreed quality, using existing accepted calibration when sufficient. Test the affected interaction when a repair risks acceptance, replay or output integrity. Neither sampling nor broad verification is a prerequisite for every resume. Stop gathering proof once the relevant risk and authorized outcome are resolved. Preserve applicable privacy, security and release boundaries.

Instrument performance

Instrument enough of the critical path to distinguish useful execution, queue waits, retries/backoff, release waits and duplicated work. Use existing structured logs/status artifacts, not a new telemetry service. Record run/item/stage/attempt IDs, effective model/configuration, input size, active/queued counts, concurrency, queue/start/end timestamps, outcome, retry reason, output disposition and receipt where available. For transfers, record bytes and duration. Exclude credentials and unnecessary source content. Instrumentation should answer an operational question; incomplete optional observability should not block a useful pilot.

Look for missed opportunities: idle capacity despite ready work, unnecessary stage barriers, serial independent I/O, repeated validation of unchanged artifacts, redundant model calls and repeated builds. Prioritize the largest measured delay. Measure accepted output delivered per elapsed hour, and record user interventions as an operational cost; worker activity alone is not progress.

Diagnose the largest unexpected delay on the critical path: slow calls, starvation, transfers, repairs or gating. Record performance changes with their hypothesis and before/after useful throughput. Variance is a diagnostic signal, not an automatic blocker.

Report progress and ETA

Use report-progress-eta-analyze for whole-task summaries, remaining-stage forecasts, comparisons with prior estimates and analysis of measured inefficiencies. Keep evidence in the existing run artifacts. Apply findings through the intervention authority above; reporting does not create a second monitor or completion ledger.

Normally report routine meaningful progress at most every fifteen minutes, with material regressions, completion and required action reported promptly. Stay quiet when progress, forecast and risk are unchanged; answer explicit status requests directly.

Stop and hand off

Completion is the authorized outcome: validated outputs for generation, or deployment/live evidence when shipping was authorized. Optional unavailable checks are disclosed; a blocked required dependency receives a precise incomplete handoff.

For pilot-and-improve runs, report output quality, timing, reliability, significant interventions and remaining optimization opportunities. Update the existing workflow documentation to reflect verified behavior, distinguishing measured results from proposals.

Link outcomes and remaining qualifications. Pause the monitor after verified completion or explicit external handoff. Preserve reusable results and receipts; clean known disposable resources and avoid orphaned workers or duplicate monitors. Babysitting does not itself authorize additional scope or publication.

© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in babysit-runs of swyxio/skills.

Open the folder on GitHubat commit 038ef34

Compare with similar skills

Babysit Runs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Babysit Runs compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Babysit Runs this skillswyxio/skills176—~2.9kAutomated safety check: PassMIT
Babysitsimstudioai/sim30k—~2.9kAutomated safety check: PassApache-2.0
It Operationsdavila7/claude-code-templates33k1 repos~3.7kAutomated safety check: PassMIT
Operator Approval Loopaffaan-m/ECC277k—~3.3kAutomated safety check: PassMIT
Business Operations Skillsalirezarezvani/claude-skills28k—~2.3kAutomated safety check: PassMIT
Babysitnubjs/nub4.4k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Babysit

    simstudioai/sim

    Drive a PR to a clean review (Greptile 5/5, zero open threads) — ships if needed, keeps it mergeable against staging, re-triggers both Greptile and cubic, fixes real findings, replies to and…

    30k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • It Operations

    davila7/claude-code-templates

    Manages IT infrastructure, monitoring, incident response, and service reliability.

    33k GitHub starsUsed in 1 repo~3.7k tokens
    DevOps & CloudAuto-check passed
  • Operator approval contract with internal filing notices for agent-drafted outbound messages, hashed drafts, epoch-keyed decisions, durable delivery claims and receipts, and a pre-draft baseline gate.

    277k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Business Operations Skills

    alirezarezvani/claude-skills

    A skill your agent uses when running, diagnosing, or designing internal business operations — process documentation, vendor SLAs, capacity planning, internal comms, SOP/runbook authoring…

    28k GitHub stars~2.3k tokensUpdated 1 mo ago
    Business, Finance & HRAuto-check passed
  • Babysit

    nubjs/nub

    Bring one nubjs/nub pull request to merge-readiness — pull the inline reviews, verify each finding against the code, re-verify every fix round locally, fix CI, and loop.

    4.4k GitHub stars~1.9k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Verification Before Completion

    foryourhealth111-pixel/Vibe-Skills

    Completion-evidence route used before claiming work is complete, fixed, passing, committed, or PR-ready.

    3.6k GitHub stars~1.1k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed

More from swyxio/skills

All 89 skills in this repo
  • Programmatic Agents

    swyxio/skills

    Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.

    176 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Design, implement, audit, or refresh protected username and handle namespaces for public products.

    176 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • New Mac Setup

    swyxio/skills

    Fully automated new Mac setup for fullstack web developers and AI engineers.

    176 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Youtube API

    swyxio/skills

    Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…

    176 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.

    176 GitHub stars~1.5k tokensUpdated today
    Auto-check: warnings
  • Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

    176 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Questions about Babysit Runs

What does Babysit Runs do?

Operate unattended runs through completion, from a single detached pilot to a large batch. Babysit Runs is an agent skill from swyxio/skills. Operate unattended runs through completion, from a single detached pilot to a large batch.

When should I use Babysit Runs?

Babysit Runs fits situations like: supervised pilots; overnight runs and bounded autonomous optimization; not a one-time status check; concurrency advice alone.

How do I install Babysit Runs in Claude Code?

Run `npx skills add swyxio/skills --skill babysit-runs -a claude-code`. Or copy the skill folder (babysit-runs in swyxio/skills) into .claude/skills/babysit-runs in your project. Claude Code loads it when a task matches its description.

How do I install Babysit Runs in Codex?

Run `npx skills add swyxio/skills --skill babysit-runs -a codex`. Or copy the skill folder (babysit-runs in swyxio/skills) into .agents/skills/babysit-runs in your project. Codex loads it when a task matches its description.

Can I use Babysit Runs in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill babysit-runs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/babysit-runs, .gemini/skills/babysit-runs, .github/skills/babysit-runs and .opencode/skills/babysit-runs in your project.

What does Babysit Runs need to run?

SKILL.md names no scripts, command-line tools or credentials: Babysit Runs is instructions for the agent only.

Does Babysit Runs access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Babysit Runs safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Babysit Runs use?

Babysit Runs is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Babysit Runs use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Babysit Runs?

Skills that share tags, products or a category with Babysit Runs: Babysit (simstudioai/sim, 30k stars), It Operations (davila7/claude-code-templates, 33k stars), Operator Approval Loop (affaan-m/ECC, 277k stars) and Business Operations Skills (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Babysit Runs?

swyxio (a GitHub user) maintains it in swyxio/skills, which has 176 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.

Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.