Aisafetyhot
wuyoscar/AISafetyHot-Hub
Query AI Safety HOT news, research papers, incidents, hot topics, and daily/weekly/monthly reports through its public read-only MCP service.
A skill your agent uses when a Scenario generation is blocked, refused, or returns a moderation or sensitive-content error, when a video is rejected after rendering or its audio track is flagged…
$ npx skills add scenario-labs/skills --skill scenario-moderation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install scenario-labs/skills scenario-moderation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario-moderation .claude/skills/scenario-moderation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scenario-moderation" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-moderation into .claude/skills/scenario-moderation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-moderation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/scenario-labs/skills/tree/main/skills/scenario-moderationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add scenario-labs/skills --skill scenario-moderation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install scenario-labs/skills scenario-moderation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/scenario-moderation .agents/skills/scenario-moderation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scenario-moderation" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-moderation into .agents/skills/scenario-moderation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-moderation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scenario-labs/skills --skill scenario-moderation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install scenario-labs/skills scenario-moderation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/scenario-moderation .cursor/skills/scenario-moderation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scenario-moderation" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-moderation into .cursor/skills/scenario-moderation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-moderation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/scenario-labs/skills.git --path skills/scenario-moderation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add scenario-labs/skills --skill scenario-moderation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install scenario-labs/skills scenario-moderation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/scenario-moderation .gemini/skills/scenario-moderation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scenario-moderation" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-moderation into .gemini/skills/scenario-moderation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-moderation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install scenario-labs/skills scenario-moderationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add scenario-labs/skills --skill scenario-moderation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/scenario-moderation .github/skills/scenario-moderation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scenario-moderation" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-moderation into .github/skills/scenario-moderation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-moderation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scenario-labs/skills --skill scenario-moderation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install scenario-labs/skills scenario-moderation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/scenario-moderation .opencode/skills/scenario-moderation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scenario-moderation" agent skill from https://github.com/scenario-labs/skills/tree/main/skills/scenario-moderation into .opencode/skills/scenario-moderation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scenario-moderation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scenario-moderationA skill your agent uses when a Scenario generation is blocked, refused, or returns a moderation or sensitive-content error, when a video is rejected after rendering or its audio track is flagged…
Scenario Moderation is an agent skill from scenario-labs/skills. Use when a Scenario generation is blocked, refused, or returns a moderation or sensitive-content error, when a video is rejected after rendering or its audio track is flagged, when an output is flagged for likeness to a real person or for IP and copyright detection, when the same prompt passes on one model and fails on another, or when a team's own characters, weapons, or props get flagged. Keywords: blocked prompt, content moderation, refused generation, provider filter, false positive.
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM guardrails. The repository describes itself as: Get production-ready images, video, audio, and 3D from any AI agent: skills that pick the right model, price before spending, and keep characters and brands consistent through… The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f6f8ab7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scenario Moderation loads about 3.6k tokens when it runs. Until then it costs about 128 tokens; SKILL.md has 1,794 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from scenario-labs/skills at commit f6f8ab7, republished under its MIT licence (© scenario-labs). 1,794 words, ~3,610 tokens.
.claude/skills/scenario-moderation/SKILL.md (or your agent's skills folder).Three gates can stop a generation, and only two of them read the prompt. Content filters run on the model provider's side, not on Scenario, so a provider block is a property of the model that was picked and the same prompt usually passes elsewhere in the catalog. IP Detection is the team's own gate: it screens the prompt and the input images before the run, an organization admin switches it on, and no model switch clears it. Plan and blocklist restrictions remove models before any prompt is judged. Treat a block as a routing problem first and a wording problem second. This skill is about false positives on content a team is entitled to make; it is not a way to produce content a provider prohibits. Core loop: see the scenario skill. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.
| Step | Call |
|---|---|
| Read the actual error | job_get with job_id: the row carries error and hint, plus modelId (the model to exclude) and cuCost (what the failed run charged); verbose=true adds metadata.input with the exact prompt the job ran. Or read error and hint off the jobs_wait row |
| Confirm the charge | job_get's cuCost first; a policy block is typically refunded and failed jobs are reimbursed except xAI generations stopped by moderation (see scenario), but the reservation is not always released at once, so confirm later with usage bounded by start_date and end_date to the job's day and include=["usages"], reading the per-type totals (the headline is project-lifetime; a team-gate screen shows as its own IP Detection line), and file the job_id with support if a charge stands |
| Read the team's gate | teams_list: the team row carries ipDetectionEnforcement; anything but disabled means a refusal naming intellectual-property risk is the team's own policy, screened before the run |
| Find alternative models | recommend with the failed job's capability (txt2img for a text-to-image block, img2img with a reference in play, img2video for a clip animated from a still) plus the user's own words; set max_cost_cu a little above the failed row's cuCost per asset to stay in the cost band, and drop the failed modelId from the ranking yourself, since recommend has no exclusion argument |
| Price an alternative | model_run with dry_run=true |
| Re-test the same intent | One model_run per candidate, prompt unchanged, so the model stays the only variable |
Four failures look alike and only the first is about wording: a provider moderation block, an IP Detection block from the team's own settings, a 403 Forbidden error (the plan does not include that model), and a model a team has put on its own blocklist. Read the error before rewriting anything.
Three triggers stack, and each alone often sits under the threshold:
With two present the prompt sits near the threshold, which is why the same intent passes one run and fails the next: a small rewording tips it over. An upstream LLM step that rewrites prompts is a frequent cause, because it leans harder on intensifiers to solve an unrelated problem and re-introduces the block on every run.
Scenario's public content policy guide ranks the providers: Google's models (Gemini, Veo) block the most brand and character references, ByteDance, OpenAI, Ideogram and Hunyuan sit in the middle, and FLUX and Recraft are the most permissive. recommend ranks by measured performance, not by filter strictness, so an alternative from the same provider inherits the same filter: take the next candidate from a different provider, and read modelId on the failed row to know which one that was: the provider is usually legible in the id, and when it is not, model_get (catalog-only, read lane) returns the record with complianceMetadata.modelProvider.
Video changes the economics of a block. Several providers moderate the finished render, so a rejected clip has already been rendered in full, and a retry of the identical payload renders it again. Read cuCost on the failed row before anything else and confirm the charge per the quick reference, then switch provider as above, prompt unchanged.
The audio track is judged on its own. OutputAudioSensitiveContentDetected on an audio-enabled video run means the soundtrack tripped the filter, and the usual cause is a named instrument or genre, even inside an exclusion: "no music" and "a distant guitar" have both failed where "room tone, footsteps, a single voice" passed. Describe diegetic sound positively and name nothing to exclude; when the clip needs no sound, turn the audio field off where the schema exposes one (generateAudio on many families) rather than prompting silence.
Teams on the Enterprise plan can switch on IP Detection, Scenario's own gate, independent of any provider's filter: it reads the prompt and the input images before the run against the filters an organization admin enabled (fictional characters, brands and trademarks, celebrity likeness, artist styles, and custom filters), and a match fails the job at once with a message naming intellectual-property risk, nothing generated and no generation cost charged. The screen itself bills 1 CU per screened generation plus 1 CU per input image, charged even when the result is a block and reported as its own IP Detection line in usage; video, audio and 3D inputs are neither analyzed nor charged. The setting rides the teams_list row as ipDetectionEnforcement: while it reads disabled, every IP or copyright refusal is a provider's; the active levels differ in one thing only, whether a job is let through or blocked when the check itself cannot run, so report the literal value and that distinction. When it is active, read which gate the error names: the team's gate names intellectual-property risk, a provider's names a content policy violation, and when the text says neither, run the unchanged prompt once on one alternative provider: a provider filter clears with the switch, the team's gate repeats on any model. A team-gate block earns one retry with the design described visually and no protected mark in the prompt or in any input image (the brands filter reads a logo shown in a reference like a name, while generic product descriptions are not flagged). If that repeats, it is a conversation, not a call: tell the user which gate fired; an organization admin owns the filters and can opt a single project out of screening for licensed work, and verified IP holders who keep hitting false positives on their own licensed property have an escalation path through their Scenario account manager. A copyright warning attached to a completed output is the provider's flag and informational: the asset was delivered, and reviewing it before commercial use is the team's call.
scenario-consistency). Drop real-person comparisons the same way.Then stop. If every provider refuses and one honest rewrite has not cleared it, the filter is reading something real: say so and hand it back to the user. Grinding out variants until one slips through is evasion, not art direction.
The studio's own character is named Onyx, and "Onyx's oversized war hammer, huge spiked head" comes back flagged.
job_get with job_id (verbose=true for the exact prompt the job ran), or the error and hint fields on the jobs_wait row. It names moderation, so this is a provider filter, not a plan restriction, a team blocklist, or IP Detection (teams_list shows ipDetectionEnforcement: "disabled").recommend with the failed job's capability (txt2img here) and the user's own words, take two alternatives from providers other than the failed modelId's, price each with model_run and dry_run=true, then run the unchanged prompt on each. One passes: done, the filter belonged to the first provider.scenario-consistency).scenario-report.© scenario-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/scenario-moderation of scenario-labs/skills.
Open the folder on GitHubat commit f6f8ab7
Scenario Moderation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scenario Moderation this skillscenario-labs/skills | 946 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Aisafetyhotwuyoscar/AISafetyHot-Hub | 708 | — | ~1.4k | Automated safety check: Pass | Custom licence | |
| ObliteratusRedWoodOG/Hermes-Desktop | 177 | 5 repos | ~3.8k | Automated safety check: Pass | MIT | |
| Lemonade Router Builderamd/skills | 408 | — | ~4k | Automated safety check: Pass | MIT | |
| Execution Guardrailsmrtooher/fable-mode | 873 | — | ~1k | Automated safety check: Pass | None | |
| Writing Eval Scenariosopen-bias/open-bias | 143 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 |
wuyoscar/AISafetyHot-Hub
Query AI Safety HOT news, research papers, incidents, hot topics, and daily/weekly/monthly reports through its public read-only MCP service.
RedWoodOG/Hermes-Desktop
Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…
amd/skills
Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.
mrtooher/fable-mode
Always-on operational guardrails, model-independent. An agent skill from mrtooher/fable-mode.
open-bias/open-bias
Guide for writing eval conversation JSONs and running them through policy engines
gambitph/Stackable
A skill your agent uses when you need a deterministic inspection of a WordPress repository (plugin/theme/block theme/WP core/Gutenberg/full site) including tooling/tests/version hints, and a…
scenario-labs/skills
A skill your agent uses when drawing or animating with Grease Pencil in Blender 5.x from Python: 2D or 2.5D illustration, frame-by-frame animation, a cutout or part-based 2D character, strokes with…
scenario-labs/skills
A skill your agent uses when grooming hair or fur in Blender with hair curves, such as a character hairstyle, animal fur, procedural fur in geometry nodes, or hair cards and mesh hair for games.
scenario-labs/skills
A skill your agent uses when lighting, rendering or compositing in Blender: light a character, product or hero shot, interior at dusk or night, three-point or motivated lighting, sun and sky, HDRI…
scenario-labs/skills
A skill your agent uses when creating a ChatGPT pet or Codex pet with Scenario: hatching an animated companion from a text idea, a character, mascot or brand cue, or reference photos and art; making…
scenario-labs/skills
A skill your agent uses when animating characters or scenes in Godot 4.7: AnimationPlayer clips and RESET, AnimationTree state machines and blend spaces built in code, Mixamo or glTF import, loop…
scenario-labs/skills
A skill your agent uses when adding or fixing sound in Godot 4.7: audio buses and effects, volume sliders, 'too many sounds', combat audio with hundreds of enemies, sounds clipping or distorting, 3D…
Categories
A skill your agent uses when a Scenario generation is blocked, refused, or returns a moderation or sensitive-content error, when a video is rejected after rendering or its audio track is flagged…. Scenario Moderation is an agent skill from scenario-labs/skills. Use when a Scenario generation is blocked, refused, or returns a moderation or sensitive-content error, when a video is rejected after rendering or its audio track is flagged, when an output is flagged for likeness to a real person or for IP and copyright detection, when the same prompt passes on one model and fails on another, or when a team's own characters, weapons, or props get flagged.
Scenario Moderation fits situations like: A Scenario generation is blocked; returns a moderation; sensitive-content error; A video is rejected after rendering.
Run `npx skills add scenario-labs/skills --skill scenario-moderation -a claude-code`. Or copy the skill folder (skills/scenario-moderation in scenario-labs/skills) into .claude/skills/scenario-moderation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add scenario-labs/skills --skill scenario-moderation -a codex`. Or copy the skill folder (skills/scenario-moderation in scenario-labs/skills) into .agents/skills/scenario-moderation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scenario-labs/skills --skill scenario-moderation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-moderation, .gemini/skills/scenario-moderation, .github/skills/scenario-moderation and .opencode/skills/scenario-moderation in your project.
Going by SKILL.md and its folder, Scenario Moderation needs the command-line tools its instructions call (npx). Our summary lists: Node.js.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Scenario Moderation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Scenario Moderation: Aisafetyhot (wuyoscar/AISafetyHot-Hub, 708 stars), Obliteratus (RedWoodOG/Hermes-Desktop, 177 stars), Lemonade Router Builder (amd/skills, 408 stars) and Execution Guardrails (mrtooher/fable-mode, 873 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
scenario-labs (a GitHub organization) maintains it in scenario-labs/skills, which has 946 GitHub stars. The repository holds 146 skills in this directory. The repository was last updated on October 10, 2026.
Source: scenario-labs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.