Agent skill

CI Adhoc Test

by nubjs in nubjs/nub

Run ad-hoc / exploratory tests on a real OS or platform via CI when the behavior CANNOT be reproduced on the local host or in Docker — macOS Seatbelt / sandbox-exec / codesigning, Windows cmd.exe /…

MITAuto-check passedDevOps & Cloud

Install CI Adhoc Test

skills CLI
$ npx skills add nubjs/nub --skill ci-adhoc-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nubjs/nub ci-adhoc-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nubjs/nub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ci-adhoc-test .claude/skills/ci-adhoc-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ci-adhoc-test
GitHub stars
4.4k
Token cost
~1.7k tokens
SKILL.md length
805 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Run ad-hoc / exploratory tests on a real OS or platform via CI when the behavior CANNOT be reproduced on the local host or in Docker — macOS Seatbelt / sandbox-exec / codesigning, Windows cmd.exe /…

  • Asks to test this on macOS/Windows CI
  • SKILL.md covers The key fact: no PR is required, The harness shape, Test every candidate FIX in… and Run and watch, plus 3 more sections
  • Calls gh, git and cargo
  • Probe this across operating systems / platforms

What it does

CI Adhoc Test is an agent skill from nubjs/nub. Run ad-hoc / exploratory tests on a real OS or platform via CI when the behavior CANNOT be reproduced on the local host or in Docker — macOS Seatbelt / sandbox-exec / codesigning, Windows cmd.exe / --script-shell shell selection / .cmd resolution / Authenticode, musl-vs-glibc, Linux-arm64, a specific Node floor. Invoke (via the Skill tool) whenever the user asks to "test this on macOS/Windows CI", "probe this across operating systems / platforms", "run an ad-hoc cross-platform check", or any validation that needs…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Pull requests, Containers and CI/CD. It works with macOS, Docker, Linux and GitHub Actions. The repository describes itself as: The fast all-in-one Node.js toolkit. The licence is MIT.

When your agent uses it

  • Asks to test this on macOS/Windows CI
  • Probe this across operating systems / platforms
  • Run an ad-hoc cross-platform check
  • Any validation that needs a real macOS

Example prompts

  • “test this on macOS/Windows CI”
  • “probe this across operating systems / platforms”
  • “run an ad-hoc cross-platform check”
  • “/ci-adhoc-test”

Requirements

  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 568e73a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git
    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CI Adhoc Test loads about 1.7k tokens when it runs. Until then it costs about 232 tokens; SKILL.md has 805 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~232
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nubjs/nub at commit 568e73a, republished under its MIT licence (© nubjs). 805 words, ~1,691 tokens.

Download SKILL.mdSave it as .claude/skills/ci-adhoc-test/SKILL.md (or your agent's skills folder).
name
ci-adhoc-test
description
Run ad-hoc / exploratory tests on a real OS or platform via CI when the behavior CANNOT be reproduced on the local host or in Docker — macOS Seatbelt / sandbox-exec / codesigning, Windows cmd.exe / --script-shell shell selection / .cmd resolution / Authenticode, musl-vs-glibc, Linux-arm64, a specific Node floor. Invoke (via the Skill tool) whenever the user asks to "test this on macOS/Windows CI", "probe this across operating systems / platforms", "run an ad-hoc cross-platform check", or any validation that needs a real macOS or Windows runner. THE KEY FACT this skill exists to carry: a pull request is NOT required — a branch-scoped GitHub Actions workflow (workflow_dispatch + push to the branch) runs the probe with no PR open. Pairs with `ad-hoc-test` (local host probing), `dev-loop` (build), `ci-watch` (await the run), and AGENTS.md's Docker section (Linux-only; this skill covers what Docker can't).
metadata.internal
true

Ad-hoc testing on a real OS/platform via CI

Some behavior is only observable on a real target platform: macOS Seatbelt / sandbox-exec / Gatekeeper / codesigning, Windows cmd.exe / --script-shell selection / .cmd/.bat resolution / Authenticode / SmartScreen, musl-vs-glibc detection, Linux-arm64, a pinned Node floor. Docker closes the Linux corners cheaply but runs Linux containers only — it is not a substitute for macOS or Windows.

The key fact: no PR is required

A GitHub Actions workflow does not need an open pull request. Trigger it on the branch:

yaml
on:
  push:
    branches: [<your-probe-branch>]  # THE trigger: runs on every push to THIS branch
    paths:                           # only when the harness itself changes
      - 'tests/<probe-name>/**'
      - '.github/workflows/<wf>.yml'
  workflow_dispatch:                 # INERT until the file lands on the default branch — see below
  • Push to the branch → the probe runs. Open no PR, or close one you opened.
  • Re-run with another push: git commit --allow-empty -m rerun && git push. Or re-run a finished run with gh run rerun <run-id> (gh run list --branch <branch> for the id).
  • push is load-bearing; workflow_dispatch alone will NOT work for a branch-only probe. GitHub only registers a workflow_dispatch workflow if the file is on the default branch, so gh workflow run <wf>.yml --ref <branch> errors ("no workflows found") and it never appears in the Actions UI. Keep the entry for when the probe graduates to main; don't rely on it before then.
  • Do NOT open a PR just to get CI. A PR signals "ready to land," which a prototype is not.
  • Omit any pull_request: trigger — it ties runs to PR state, which is what you're avoiding.

The harness shape

Keep the probe self-contained under tests/<probe-name>/, mirroring the existing ones:

  • A generator / runner (e.g. an SBPL profile generator + a sandbox-exec runner; a .cmd resolver harness).
  • Fast unit + smoke tests asserting the enforcement and the bypass/fail-closed cases.
  • A README.md (what it validates, how to reproduce locally) and a results.md (findings, plus heavy runs reproduced on demand).
  • The branch-scoped workflow .github/workflows/<wf>.yml.

Mirror tests/sandbox-macos-writeconfine/ + .github/workflows/sandbox-macos-writeconfine.yml, or its Windows counterpart tests/sandbox-win-probes/ (each on its own probe branch, not main).

Keep CI lean — fast, deterministic core only: unit tests + the enforcement/bypass smoke matrix, no network, no mega-fixture. Heavy or combinatorial runs are documented in results.md and reproduced on demand. CI capacity is shared; a 22-minute-per-push probe job is already a lot.

Test every candidate FIX in one build-free run, not one per push

When the probe fails and you have several plausible repairs, do not guess-and-push: each loop costs a full build plus queue time. Write a job that tries all candidates at once against a stand-in, with no compile step — a .cmd/shell quoting question needs cmd.exe and a two-line stub, not the real binary. Minutes instead of half an hour, and the losing candidates are results you keep.

Two rules that make the output trustworthy:

  • Include the CURRENT broken form as a control. It must fail. If it passes, the reproduction is wrong and every other row is meaningless — say so in the output rather than reading the winners.
  • A losing alternative is a result, not a job failure. Exit non-zero only when the control passes or when NO candidate works. Otherwise the job sits permanently red over a candidate you never intended to ship, and readers learn to ignore it.
Show full SKILL.md (304 more words)Show less

Run and watch

  • Kick a run by pushing to the branch; list with gh run list --workflow <wf>.yml --branch <branch>.
  • Await a specific run with the ci-watch skill (scripts/ci-watch.ts) rather than raw gh run watch — it waits for the run to exist, polls authoritative terminal status, fails fast, and exits 0 only on confirmed success.
  • A failure is immediately actionable — read the job log, fix the harness, push again.

Lifecycle

  • The branch is the durable home of the probe while it's exploratory. Push, run, iterate.
  • When it graduates into a permanent regression check, fold it into main through the normal flow (it's tests/** + a workflow file — a content/CI change routed straight to main). Decide its steady-state trigger then.
  • If you only needed the one-time answer, leave the branch as the record (or delete it once results.md captures the findings) — never open a PR to "preserve" a throwaway probe.

A cross-compile CHECK is not a test run

cargo check -p nub-cli --all-targets --target x86_64-pc-windows-gnu proves the code COMPILES for Windows. It says nothing about whether the tests PASS there, and the gap is exactly where platform behavior lives: .cmd shims vs #!/usr/bin/env node, path separators, NODE_OPTIONS tokenizing, process.title, mode bits that are no-ops.

A cross-compile is a necessary pre-flight, never the evidence. If you are un-gating tests from #[cfg(unix)], or writing anything whose behavior could differ by platform, run a branch probe BEFORE pushing — it needs no PR and costs one workflow run.

When to reach for this vs the alternatives

  • Local host probe (ad-hoc-test) — the behavior reproduces on your dev machine. Default for anything not platform-gated.
  • Docker (AGENTS.md) — a Linux corner: musl/glibc, a Node floor, a clean dependency-free environment, first-run install.
  • This skill (CI branch probe) — a macOS or Windows behavior, or a real multi-runner matrix, that neither the host nor Docker can show.

© nubjs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/ci-adhoc-test of nubjs/nub.

Open the folder on GitHubat commit 568e73a

Compare with similar skills

CI Adhoc Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CI Adhoc Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CI Adhoc Test this skillnubjs/nub4.4k—~1.7kAutomated safety check: PassMIT
Releasing MarchatCod-e-Codes/marchat137—~801Automated safety check: PassMIT
Swig CI Reproswig/swig6.3k—~1.2kAutomated safety check: PassCustom licence
Gh Localankitvgupta/exo496—~581Automated safety check: PassCustom licence
GitHub Runnermagnus919/agent-skills113—~1.7kAutomated safety check: PassMIT
.NET Crash Dump Collectiondotnet/skills5.6k2 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Releasing Marchat

    Cod-e-Codes/marchat

    Prepares marchat releases: version bumps, CHANGELOG, packaging checksums, GitHub Actions release workflow, and Docker tags.

    137 GitHub stars~801 tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed
  • Swig CI Repro

    swig/swig

    Reproduce a GitHub Actions Linux CI failure locally when it does not happen on your machine: a podman/docker image that mirrors the ubuntu-22.04 runner by reusing the real Tools/CI-linux-.sh install…

    6.3k GitHub stars~1.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Gh Local

    ankitvgupta/exo

    Run GitHub Actions CI workflows locally using nektos/act in Docker.

    496 GitHub stars~581 tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • GitHub Runner

    magnus919/agent-skills

    Deploy, manage, and troubleshoot self-hosted GitHub Actions runners.

    113 GitHub stars~1.7k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Official

    Configures automatic crash dumps or captures dumps from running processes for modern .NET apps on Linux, macOS and Windows, including Docker and Kubernetes.

    5.6k GitHub starsUsed in 2 repos~1.1k tokens
    DevOps & CloudAuto-check passed
  • PR Review

    kimdre/doco-cd

    Review pull request diffs for correctness and regressions, and provide actionable feedback when asked to review a PR.

    1.7k GitHub stars~311 tokensUpdated today
    DevOps & CloudAuto-check passed

More from nubjs/nub

All 31 skills in this repo
  • Cpu Reduction

    nubjs/nub

    Diagnose and clear CPU, memory, and disk contention on the maintainer's dev host.

    4.4k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Reclaim disk on the maintainer's Mac when the volume is full or filling — ENOSPC, "no space left on device", a failed build or agent harness, or a routine sweep of Rust build residue.

    4.4k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Nub Charts

    nubjs/nub

    Build a performance chart for nubjs.com — the SVG bar figures in blog posts, docs pages and social posts (a runtime augmentation against plain node, an install or dispatch comparison, a cross-tool…

    4.4k GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Audit Thread

    nubjs/nub

    A skill your agent uses when running a compatibility/parity AUDIT — enumerating where nub diverges from a reference it claims parity with (pnpm CLI grammar, a lockfile format, a Node behavior, a…

    4.4k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Linux Vm Test

    nubjs/nub

    Run ad-hoc Nub tests and debugging probes on real local Linux guests.

    4.4k GitHub stars~986 tokensUpdated today
    Auto-check passed
  • Performance-trace Nub package-manager installs using the existing phase timings, structured diagnostics, and sampling-profiler workflow.

    4.4k GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Questions about CI Adhoc Test

What does CI Adhoc Test do?

Run ad-hoc / exploratory tests on a real OS or platform via CI when the behavior CANNOT be reproduced on the local host or in Docker — macOS Seatbelt / sandbox-exec / codesigning, Windows cmd.exe /…. CI Adhoc Test is an agent skill from nubjs/nub.cmd resolution / Authenticode, musl-vs-glibc, Linux-arm64, a specific Node floor.

When should I use CI Adhoc Test?

CI Adhoc Test fits situations like: asks to test this on macOS/Windows CI; probe this across operating systems / platforms; run an ad-hoc cross-platform check; any validation that needs a real macOS.

How do I install CI Adhoc Test in Claude Code?

Run `npx skills add nubjs/nub --skill ci-adhoc-test -a claude-code`. Or copy the skill folder (.claude/skills/ci-adhoc-test in nubjs/nub) into .claude/skills/ci-adhoc-test in your project. Claude Code loads it when a task matches its description.

How do I install CI Adhoc Test in Codex?

Run `npx skills add nubjs/nub --skill ci-adhoc-test -a codex`. Or copy the skill folder (.claude/skills/ci-adhoc-test in nubjs/nub) into .agents/skills/ci-adhoc-test in your project. Codex loads it when a task matches its description.

Can I use CI Adhoc Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nubjs/nub --skill ci-adhoc-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ci-adhoc-test, .gemini/skills/ci-adhoc-test, .github/skills/ci-adhoc-test and .opencode/skills/ci-adhoc-test in your project.

What does CI Adhoc Test need to run?

Going by SKILL.md and its folder, CI Adhoc Test needs the command-line tools its instructions call (gh, git and cargo). Our summary lists: Docker.

Does CI Adhoc Test access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is CI Adhoc Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does CI Adhoc Test use?

CI Adhoc Test is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CI Adhoc Test use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to CI Adhoc Test?

Skills that share tags, products or a category with CI Adhoc Test: Releasing Marchat (Cod-e-Codes/marchat, 137 stars), Swig CI Repro (swig/swig, 6.3k stars), Gh Local (ankitvgupta/exo, 496 stars) and GitHub Runner (magnus919/agent-skills, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CI Adhoc Test?

nubjs (a GitHub organization) maintains it in nubjs/nub, which has 4,372 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: nubjs/nub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.