Aisafetyhot
wuyoscar/AISafetyHot-Hub
Query AI Safety HOT news, research papers, incidents, hot topics, and daily/weekly/monthly reports through its public read-only MCP service.
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including…
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses guardrail-tester --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/guardrail-tester .claude/skills/guardrail-tester && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "guardrail-tester" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/guardrail-tester into .claude/skills/guardrail-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "guardrail-tester", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/guardrail-testerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses guardrail-tester --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/guardrail-tester .agents/skills/guardrail-tester && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "guardrail-tester" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/guardrail-tester into .agents/skills/guardrail-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "guardrail-tester", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses guardrail-tester --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/guardrail-tester .cursor/skills/guardrail-tester && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "guardrail-tester" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/guardrail-tester into .cursor/skills/guardrail-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "guardrail-tester", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/RyanAlberts/best-of-Agent-Harnesses.git --path skills/guardrail-tester--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses guardrail-tester --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/guardrail-tester .gemini/skills/guardrail-tester && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "guardrail-tester" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/guardrail-tester into .gemini/skills/guardrail-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "guardrail-tester", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses guardrail-testerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/guardrail-tester .github/skills/guardrail-tester && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "guardrail-tester" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/guardrail-tester into .github/skills/guardrail-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "guardrail-tester", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses guardrail-tester --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/guardrail-tester .opencode/skills/guardrail-tester && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "guardrail-tester" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/guardrail-tester into .opencode/skills/guardrail-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "guardrail-tester", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
guardrail-testerGuardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including…
Guardrail Tester is an agent skill from RyanAlberts/best-of-Agent-Harnesses. Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including wrapped, reordered, and full-path forms that slip past prefix rules, and measures prompt friction on the latest tool calls. Use when the user asks whether deny rules or hooks block force pushes, rm -rf, secret reads, or downloads piped into a shell; wants to test or audit guardrails, or find a bypass in permission…
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts and reference files (for example `README.md`, `references/bypass-forms.md` and `references/harness-rules.md`). Compatibility notes: Python 3.9+, standard library only. Reading Codex config.toml and Gemini CLI policy files needs Python 3.11+. Makes no network calls; the only commands it…
It sits in AI & LLM Engineering, covering LLM guardrails. The repository describes itself as: 🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4fa20bc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 8 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3bashgitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Python 3.9+, standard library only. Reading Codex config.toml and Gemini CLI policy files needs Python 3.11+. Makes no network calls; the only commands it runs are the user's own hook commands, fed test JSON, and only with --run-hooks after the user says yes.
From compatibility in the SKILL.md frontmatter.
Guardrail Tester loads about 2.9k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 186 tokens; SKILL.md has 1,571 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from RyanAlberts/best-of-Agent-Harnesses at commit 4fa20bc, republished under its MIT licence (© RyanAlberts). 1,571 words, ~2,883 tokens.
.claude/skills/guardrail-tester/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Permission rules match the text of a command, so git push origin +main, rm -r -f build, and
bash -c '...' can walk past a deny rule written for the plain form. This skill checks the rules
and hooks the user already has against about 90 dangerous commands in those forms, then replays
their recent tool calls to count how often the rules would ask or block. The user gets a headline
("Your Claude Code guardrails block 58 of 90 dangerous commands outright. 24 more stop at a
prompt..."), every miss with a tested fix, and their friction numbers. Everything is read and
simulated locally; nothing is sent anywhere, and it runs the user's hook commands only after they
say yes.
sandbox-check.rules-to-guards.runaway-guard.templates/claude-code-safe-settings/ in this
repository, then test it here.Say this to the user in short before the first run with hooks. It is the whole contract.
./guardrail-tester-probe/
folder, its remotes and branches (probe-remote, probe-branch) are made up, and its hosts end
in .invalid, a name reserved for addresses that resolve nowhere.references/harness-rules.md says what each harness simulation covers.--run-hooks flag). Each matching
PreToolUse hook command then runs once per battery case, one run at a time, with a 10-second
limit, from the folder its harness uses (usually the project), with the session's environment
variables plus CLAUDE_PROJECT_DIR. A hook that writes a log or keeps state records those test
calls. A hook that times out twice is not run again.--replay-hooks (which needs --run-hooks) it also runs this
project's hooks on replayed calls from this project. Calls from other project folders are
checked against their own rules; their hooks never run.scripts/battery.json holds the dangerous commands on purpose: they are the test. Each line
carries the marker skillscan:allow, which tells this repository's security scanner that the line
is test data, not a command the skill runs.
<skill-dir> means the folder that holds this SKILL.md (Claude Code shows it as the skill's base
directory). Run every command from the user's project folder, and give each Bash call a 10-minute
timeout (600000 ms): hooks run one at a time, so a slow hook makes the run long.
Run the rules-only test, which reads settings and runs no hook:
python3 "<skill-dir>/scripts/test_guards.py" --project . --harness claude-codeAdd --harness for the harness you are running in (claude-code, codex, gemini-cli,
opencode, or cursor) unless the user asks about all; without it the tester covers every
harness with settings on this machine. Done when the output starts with a bold headline, or
you have told the user the error (exit code 2 means a bad argument or an unreadable battery
file).
Tell the user what was found: the settings files and every hook command from the report's "What was found" section, quoted exactly. Done when the user has seen the list.
Read each hook script, then ask before running hooks. Open the script each hook command
runs. If a script runs, evaluates, or sends its input anywhere (eval, bash -c "$cmd", a
curl with the input, a queue), or has other side effects such as writing a log, tell the user
exactly that and recommend testing without hooks. Otherwise name the hook commands and say each
will get test JSON for about 90 battery calls. On a clear yes, run this with a 10-minute Bash
timeout:
python3 "<skill-dir>/scripts/test_guards.py" --project . --harness claude-code --run-hooks --replay 500Add --replay-hooks only when the user also agrees that this project's hooks see up to 500 of
their recent real calls. When the user declines, or a hook has side effects, run the same
command without --run-hooks. Done when the headline matches the choice: it starts "Your ...
guardrails block" after hooks ran, and "Without running your N hooks" when hooks exist and were
skipped.
Pick the fixes. For each row in "Misses, worst first", take the Fix column: a deny rule
the simulation proved catches that case and leaves everyday commands and files such as
.env.example alone, or a named hook check (an extended regular expression printed below the
tables). Use references/bypass-forms.md to explain why a form slips. Done when every miss you
report has its fix or the note "use the sandbox".
Offer the change. Draft the settings edit or hook lines as a diff against the user's file, show it, and apply it only after a clear yes. Then rerun step 3 to show the new count.
defaultMode in the settings, the tester simulates auto mode, the built-in
default on Claude Code 2.1.283 and later, and adds a "Claude Code in Manual mode" row.
--mode simulates another mode. Plan mode lasts only until a plan is approved.block cases count as stopped only when blocked outright, because people approve
prompts by reflex. ask cases count when they prompt or are blocked. --fail-on-miss uses this
count.blocked, asks first, runs without asking, left to the auto-mode classifier, or unknown. The part in parentheses names the layer that decided: a hook, a rule,
a built-in check (protected paths, critical-path removals, the read-only command set), the
mode, or the sandbox. With a sandbox on, a command that needs the network or writes outside the
project shows asks first (needs network) or stopped by the sandbox.--hook-timeout to settle them.Bash matcher that Cursor never fires, a Codex hook nobody trusted in /hooks,
a shadowed OpenCode deny, or bypass mode left available.--json has every case with verdict, bucket, decided_by, rules_verdict, and
fix, plus hook_checks, every hook check by name.templates/claude-code-safe-settings/ rather than writing a new guard.scripts/test_guards.py: the tester. Flags: --harness, --project, --run-hooks,
--no-hooks (the default), --hook-workers N (default 1), --replay N, --replay-hooks,
--since DAYS, --mode, --battery, --hook-timeout, --home, --json, --out,
--fail-on-miss (exit 1 when a case is not blocked as expected).scripts/battery.json: the dangerous commands, one case per line, each with what it needs to
do its harm (network, write-outside, write-git, write-project).scripts/claude_rules.py, scripts/harness_rules.py: rule simulation per harness.scripts/hook_runner.py: runs one hook command with test JSON and reads its decision.scripts/shell_split.py: splits a command line into the commands it runs.scripts/transcripts.py: reads session transcripts for replay (shared copy).scripts/safe.py: masks secrets and puts text from settings, hooks, and transcripts in inline
code for the report (shared copy).references/bypass-forms.md: each form in the battery and why rules miss it.references/harness-rules.md: matching rules per harness, with sources and the date checked.© RyanAlberts, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts, references) in skills/guardrail-tester of RyanAlberts/best-of-Agent-Harnesses.
Open the folder on GitHubat commit 4fa20bc
Guardrail Tester next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Guardrail Tester this skillRyanAlberts/best-of-Agent-Harnesses | 1.1k | — | ~2.9k | Automated safety check: Pass | MIT | |
| Aisafetyhotwuyoscar/AISafetyHot-Hub | 827 | — | ~1.4k | Automated safety check: Pass | Custom licence | |
| ObliteratusRedWoodOG/Hermes-Desktop | 177 | 5 repos | ~3.8k | Automated safety check: Pass | MIT | |
| Lemonade Router Builderamd/skills | 408 | — | ~4k | Automated safety check: Pass | MIT | |
| Writing Eval Scenariosopen-bias/open-bias | 143 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Wp Project Triagegambitph/Stackable | 351 | 3 repos | ~371 | Automated safety check: Pass | GPL-3.0 |
wuyoscar/AISafetyHot-Hub
Query AI Safety HOT news, research papers, incidents, hot topics, and daily/weekly/monthly reports through its public read-only MCP service.
RedWoodOG/Hermes-Desktop
Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…
amd/skills
Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.
open-bias/open-bias
Guide for writing eval conversation JSONs and running them through policy engines
gambitph/Stackable
A skill your agent uses when you need a deterministic inspection of a WordPress repository (plugin/theme/block theme/WP core/Gutenberg/full site) including tooling/tests/version hints, and a…
aws-samples/sample-well-architected-skills-and-steering
Generate preventive Well-Architected guardrails — AWS Config rules, Service Control Policies, permission boundaries, CloudWatch alarms, and IaC policy checks (CDK Aspects, cfn-guard, OPA/Sentinel) —…
RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot instructions) each coding agent loads from a repo, what gets cut or skipped, and whether the commands those…
RyanAlberts/best-of-Agent-Harnesses
Claim checker that audits a coding agent's statements that tests pass or a build is clean against its own session transcripts: whether a matching run happened before the claim, whether it passed…
RyanAlberts/best-of-Agent-Harnesses
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests…
RyanAlberts/best-of-Agent-Harnesses
Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where…
RyanAlberts/best-of-Agent-Harnesses
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding agent keeps breaking, counts every violation in recent Claude Code, Codex, Gemini CLI, and OpenCode…
RyanAlberts/best-of-Agent-Harnesses
Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on the same tool call, keeps failing, or exceeds a dollar cap.
Categories
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including…. Guardrail Tester is an agent skill from RyanAlberts/best-of-Agent-Harnesses. Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including wrapped, reordered, and full-path forms that slip past prefix rules, and measures prompt friction on the latest tool calls.
Guardrail Tester fits situations like: the user asks whether deny rules; hooks block force pushes; downloads piped into a shell; audit guardrails.
Run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester -a claude-code`. Or copy the skill folder (skills/guardrail-tester in RyanAlberts/best-of-Agent-Harnesses) into .claude/skills/guardrail-tester in your project. Claude Code loads it when a task matches its description.
Run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester -a codex`. Or copy the skill folder (skills/guardrail-tester in RyanAlberts/best-of-Agent-Harnesses) into .agents/skills/guardrail-tester in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill guardrail-tester -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/guardrail-tester, .gemini/skills/guardrail-tester, .github/skills/guardrail-tester and .opencode/skills/guardrail-tester in your project.
Going by SKILL.md and its folder, Guardrail Tester needs Python for the scripts in its folder and the command-line tools its instructions call (python3, bash and git). Our summary lists: Python 3; Docker. Compatibility (from SKILL.md): Python 3.9+, standard library only. Reading Codex config.toml and Gemini CLI policy files needs Python 3.11+. Makes no network calls; the only commands it runs are the user's own hook commands, fed test JSON, and only with --run-hooks after the user says yes..
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Guardrail Tester is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Guardrail Tester: Aisafetyhot (wuyoscar/AISafetyHot-Hub, 827 stars), Obliteratus (RedWoodOG/Hermes-Desktop, 177 stars), Lemonade Router Builder (amd/skills, 408 stars) and Writing Eval Scenarios (open-bias/open-bias, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
RyanAlberts (a GitHub user) maintains it in RyanAlberts/best-of-Agent-Harnesses, which has 1,133 GitHub stars. The repository was last updated on October 9, 2026.
Source: RyanAlberts/best-of-Agent-Harnesses on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.