Obliteratus
RedWoodOG/Hermes-Desktop
Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…
Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL.
$ npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install softspark/ai-toolkit content-moderation-patterns --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/app/skills/content-moderation-patterns .claude/skills/content-moderation-patterns && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "content-moderation-patterns" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/content-moderation-patterns into .claude/skills/content-moderation-patterns/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-patterns", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/softspark/ai-toolkit/tree/main/app/skills/content-moderation-patternsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install softspark/ai-toolkit content-moderation-patterns --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/app/skills/content-moderation-patterns .agents/skills/content-moderation-patterns && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "content-moderation-patterns" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/content-moderation-patterns into .agents/skills/content-moderation-patterns/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-patterns", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install softspark/ai-toolkit content-moderation-patterns --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/app/skills/content-moderation-patterns .cursor/skills/content-moderation-patterns && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "content-moderation-patterns" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/content-moderation-patterns into .cursor/skills/content-moderation-patterns/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-patterns", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/softspark/ai-toolkit.git --path app/skills/content-moderation-patterns--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install softspark/ai-toolkit content-moderation-patterns --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/app/skills/content-moderation-patterns .gemini/skills/content-moderation-patterns && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "content-moderation-patterns" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/content-moderation-patterns into .gemini/skills/content-moderation-patterns/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-patterns", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install softspark/ai-toolkit content-moderation-patternsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/app/skills/content-moderation-patterns .github/skills/content-moderation-patterns && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "content-moderation-patterns" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/content-moderation-patterns into .github/skills/content-moderation-patterns/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-patterns", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install softspark/ai-toolkit content-moderation-patterns --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/app/skills/content-moderation-patterns .opencode/skills/content-moderation-patterns && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "content-moderation-patterns" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/content-moderation-patterns into .opencode/skills/content-moderation-patterns/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-patterns", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
content-moderation-patternsContent moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL.
Content Moderation Patterns is an agent skill from softspark/ai-toolkit. Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL. Triggers: moderation, safety filter, policy enforcement, content classifier.
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM guardrails. The repository describes itself as: Professional-grade AI coding toolkit: 94 skills, 44 agents, multi-platform (Claude, Cursor, Windsurf, Copilot, Gemini, Cline, Roo Code, Aider, Augment, Antigravity, Codex CLI… The licence is Apache-2.0.
Read from SKILL.md and the folder at commit d64db2b. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
platform.claude.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Content Moderation Patterns loads about 1.7k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 409 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from softspark/ai-toolkit at commit d64db2b, republished under its Apache-2.0 licence (© softspark). 409 words, ~1,659 tokens.
.claude/skills/content-moderation-patterns/SKILL.md (or your agent's skills folder).Apply a versioned product policy with deterministic checks, a structured classifier, and a review path. Select the model using labeled workload results. No model family has a universal accuracy or cost advantage for moderation.
input → size/format checks → policy checks → structured classifier → decision
├─ allow
├─ reject
└─ human reviewTreat submitted text as data, including any instructions it contains. Keep the classification policy in the system message. Request a short policy-grounded reason, not hidden reasoning.
Use configured size limits and exact parsed hostname checks for URL policies.
A prefix regex can mistakenly accept allowed.example.attacker.test.
from urllib.parse import urlsplit
def is_allowed_url(value, allowed_hosts):
try:
parsed = urlsplit(value)
host = parsed.hostname
port = parsed.port
except ValueError:
return False
return (
parsed.scheme == "https"
and parsed.username is None
and parsed.password is None
and host is not None
and host.casefold() in allowed_hosts
and port in (None, 443)
)This checks an already-extracted URL against normalized exact hostnames. It is not a general URL extractor or an SSRF defense. Evaluate false positives from keyword filters instead of assuming a fixed percentage of input should be blocked.
Use native output_config.format. Supply the selected model, policy and output
budget from application configuration. The following taxonomy is an example;
change its enum and routing thresholds together to match the product policy.
import json
MODERATION_SCHEMA = {
"type": "object",
"properties": {
"categories": {"type": "array", "items": {
"type": "string", "enum": ["clean", "needs_review", "spam", "harassment"],
}},
"confidence": {"type": "number"},
"reason": {"type": "string"},
},
"required": ["categories", "confidence", "reason"],
"additionalProperties": False,
}
def classify(client, model, policy, text, max_tokens):
response = client.messages.create(
model=model,
max_tokens=max_tokens,
system=policy,
output_config={"format": {"type": "json_schema", "schema": MODERATION_SCHEMA}},
messages=[{"role": "user", "content": text}],
)
if response.stop_reason != "end_turn":
raise ValueError(f"Classification incomplete: {response.stop_reason}")
blocks = [block.text for block in response.content if block.type == "text"]
if len(blocks) != 1:
raise ValueError("Expected one classification")
return json.loads(blocks[0])Apply local validation before routing. Refusal, truncation, invalid JSON or an
API failure produces a review/error outcome, never an implicit allow.
See json-mode-patterns for schema limitations and response checks.
A repeated system string is not automatically cached. If policy size and reuse
justify it, explicitly configure caching as in prompt-caching-patterns.
Do not generate heartbeat traffic to keep a cache warm.
Define categories and blocking behavior in the product policy. Keep clean
exclusive: a result containing both clean and a violation is inconsistent.
Use needs_review for uncertainty. Thresholds come from calibration and policy,
not the model's claim that its confidence is reliable.
import math
def route(classification, block_thresholds, allow_threshold):
if not isinstance(classification, dict) or set(classification) != {"categories", "confidence", "reason"}:
return "human_review"
if not isinstance(classification["reason"], str):
return "human_review"
confidence = classification.get("confidence")
categories = classification.get("categories")
if (type(confidence) not in (int, float)
or not 0 <= confidence <= 1 or not math.isfinite(confidence)):
return "human_review"
if not isinstance(categories, list) or not categories or not all(isinstance(c, str) for c in categories):
return "human_review"
categories = {category.casefold() for category in categories}
if categories - (set(block_thresholds) | {"clean", "needs_review"}):
return "human_review"
if "needs_review" in categories or ("clean" in categories and len(categories) != 1):
return "human_review"
if categories == {"clean"}:
return "pass" if confidence >= allow_threshold else "human_review"
if any(confidence >= block_thresholds[category] for category in categories):
return "reject"
return "human_review"Validate configuration thresholds as finite numbers in [0, 1] at startup. The example's category thresholds are policy-specific; it does not decide which categories your product must reject.
Use held-out labeled examples covering language, context, quoted material, benign mentions and adversarial inputs. Track precision, recall, appeal outcomes and per-category error cost. Neither false positives nor false negatives are always cheaper; the product policy determines that trade-off.
Send ambiguous cases to human review. Store decision metadata, policy/model versions and the minimum evidence needed for review under the application's retention and access controls. Do not indiscriminately log raw sensitive input.
Refresh evaluations when the policy, model or input distribution changes. Run an offline comparison before deploying a new route or threshold.
Reviewed 2026-09-23:
Use security-patterns for application input security, model-routing-patterns
for model evaluation and prompt-caching-patterns for policy caching.
© softspark, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in app/skills/content-moderation-patterns of softspark/ai-toolkit.
Open the folder on GitHubat commit d64db2b
Content Moderation Patterns next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Content Moderation Patterns this skillsoftspark/ai-toolkit | 179 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| ObliteratusRedWoodOG/Hermes-Desktop | 177 | 6 repos | ~3.8k | Automated safety check: Pass | MIT | |
| Aisafetyhotwuyoscar/AISafetyHot-Hub | 175 | — | ~1.2k | Automated safety check: Pass | Custom licence | |
| Lemonade Router Builderamd/skills | 395 | — | ~4k | Automated safety check: Pass | MIT | |
| Persona Designkangarooking/system-prompt-skills | 205 | 1 repos | ~956 | Automated safety check: Pass | MIT | |
| Execution Guardrailsmrtooher/fable-mode | 870 | — | ~1k | Automated safety check: Pass | None |
RedWoodOG/Hermes-Desktop
Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…
wuyoscar/AISafetyHot-Hub
Read AI Safety HOT daily digests, search recent AI safety research and incidents, and follow current hot topics.
amd/skills
Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.
kangarooking/system-prompt-skills
当需要为 AI 产品定义核心身份、角色声明和能力边界时调用此 skill。典型场景包括:设计新 AI 产品的 system prompt 首段、为不同场景创建差异化角色(如教学助手 vs 编程代理)、重新定义 AI 与用户的关系框架。
mrtooher/fable-mode
Always-on operational guardrails, model-independent. An agent skill from mrtooher/fable-mode.
open-bias/open-bias
Guide for writing eval conversation JSONs and running them through policy engines
softspark/ai-toolkit
Prepare or verify a project QA environment with source identity, readiness, browser access, evidence paths and owned cleanup.
softspark/ai-toolkit
Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. An agent skill from softspark/ai-toolkit.
softspark/ai-toolkit
Analyzes code quality, complexity, patterns across codebase.
softspark/ai-toolkit
Drives a brief, specification, issue or existing PR through implementation, review, tests and QA to a ready PR.
softspark/ai-toolkit
Direct technical voice for docs, README, user-facing text. An agent skill from softspark/ai-toolkit.
softspark/ai-toolkit
Detect/generate/debug CI pipeline config (GitHub Actions, GitLab CI).
Categories
Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL. Content Moderation Patterns is an agent skill from softspark/ai-toolkit. Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL.
Content Moderation Patterns fits situations like: tasks that involve LLM guardrails.
Run `npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a claude-code`. Or copy the skill folder (app/skills/content-moderation-patterns in softspark/ai-toolkit) into .claude/skills/content-moderation-patterns in your project. Claude Code loads it when a task matches its description.
Run `npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a codex`. Or copy the skill folder (app/skills/content-moderation-patterns in softspark/ai-toolkit) into .agents/skills/content-moderation-patterns in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/content-moderation-patterns, .gemini/skills/content-moderation-patterns, .github/skills/content-moderation-patterns and .opencode/skills/content-moderation-patterns in your project.
SKILL.md names no scripts, command-line tools or credentials: Content Moderation Patterns is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read.
SKILL.md names 1 domain. As links in the text: platform.claude.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Content Moderation Patterns is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Content Moderation Patterns: Obliteratus (RedWoodOG/Hermes-Desktop, 177 stars), Aisafetyhot (wuyoscar/AISafetyHot-Hub, 175 stars), Lemonade Router Builder (amd/skills, 395 stars) and Persona Design (kangarooking/system-prompt-skills, 205 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
softspark (a GitHub user) maintains it in softspark/ai-toolkit, which has 179 GitHub stars. The repository holds 112 skills in this directory. The repository was last updated on October 7, 2026.
Source: softspark/ai-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.