Agent skill

Content Moderation Patterns

by softspark in softspark/ai-toolkit

Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Content Moderation Patterns

skills CLI
$ npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install softspark/ai-toolkit content-moderation-patterns --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/app/skills/content-moderation-patterns .claude/skills/content-moderation-patterns && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
content-moderation-patterns
GitHub stars
179
Token cost
~1.7k tokens
SKILL.md length
409 words
Files
1
Skills in repo
112
Repo updated
First seen
Licence
Apache-2.0

At a glance

Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL.

  • Tasks that involve LLM guardrails
  • SKILL.md covers Architecture, Deterministic checks, Structured classifier and Categories and decision routing, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Content Moderation Patterns is an agent skill from softspark/ai-toolkit. Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL. Triggers: moderation, safety filter, policy enforcement, content classifier.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM guardrails. The repository describes itself as: Professional-grade AI coding toolkit: 94 skills, 44 agents, multi-platform (Claude, Cursor, Windsurf, Copilot, Gemini, Cline, Roo Code, Aider, Augment, Antigravity, Codex CLI… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM guardrails

Example prompts

  • “/content-moderation-patterns”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read

What it can do on your machine

Read from SKILL.md and the folder at commit d64db2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Content Moderation Patterns loads about 1.7k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 409 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from softspark/ai-toolkit at commit d64db2b, republished under its Apache-2.0 licence (© softspark). 409 words, ~1,659 tokens.

Download SKILL.mdSave it as .claude/skills/content-moderation-patterns/SKILL.md (or your agent's skills folder).
name
content-moderation-patterns
description
Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL. Triggers: moderation, safety filter, policy enforcement, content classifier.
allowed-tools
Read
effort
medium
user-invocable
false

Content Moderation Patterns

Apply a versioned product policy with deterministic checks, a structured classifier, and a review path. Select the model using labeled workload results. No model family has a universal accuracy or cost advantage for moderation.

Architecture

text
input → size/format checks → policy checks → structured classifier → decision
                                                              ├─ allow
                                                              ├─ reject
                                                              └─ human review

Treat submitted text as data, including any instructions it contains. Keep the classification policy in the system message. Request a short policy-grounded reason, not hidden reasoning.

Deterministic checks

Use configured size limits and exact parsed hostname checks for URL policies. A prefix regex can mistakenly accept allowed.example.attacker.test.

python
from urllib.parse import urlsplit


def is_allowed_url(value, allowed_hosts):
    try:
        parsed = urlsplit(value)
        host = parsed.hostname
        port = parsed.port
    except ValueError:
        return False
    return (
        parsed.scheme == "https"
        and parsed.username is None
        and parsed.password is None
        and host is not None
        and host.casefold() in allowed_hosts
        and port in (None, 443)
    )

This checks an already-extracted URL against normalized exact hostnames. It is not a general URL extractor or an SSRF defense. Evaluate false positives from keyword filters instead of assuming a fixed percentage of input should be blocked.

Structured classifier

Use native output_config.format. Supply the selected model, policy and output budget from application configuration. The following taxonomy is an example; change its enum and routing thresholds together to match the product policy.

python
import json

MODERATION_SCHEMA = {
    "type": "object",
    "properties": {
        "categories": {"type": "array", "items": {
            "type": "string", "enum": ["clean", "needs_review", "spam", "harassment"],
        }},
        "confidence": {"type": "number"},
        "reason": {"type": "string"},
    },
    "required": ["categories", "confidence", "reason"],
    "additionalProperties": False,
}


def classify(client, model, policy, text, max_tokens):
    response = client.messages.create(
        model=model,
        max_tokens=max_tokens,
        system=policy,
        output_config={"format": {"type": "json_schema", "schema": MODERATION_SCHEMA}},
        messages=[{"role": "user", "content": text}],
    )
    if response.stop_reason != "end_turn":
        raise ValueError(f"Classification incomplete: {response.stop_reason}")
    blocks = [block.text for block in response.content if block.type == "text"]
    if len(blocks) != 1:
        raise ValueError("Expected one classification")
    return json.loads(blocks[0])

Apply local validation before routing. Refusal, truncation, invalid JSON or an API failure produces a review/error outcome, never an implicit allow. See json-mode-patterns for schema limitations and response checks.

A repeated system string is not automatically cached. If policy size and reuse justify it, explicitly configure caching as in prompt-caching-patterns. Do not generate heartbeat traffic to keep a cache warm.

Show full SKILL.md (193 more words)Show less

Categories and decision routing

Define categories and blocking behavior in the product policy. Keep clean exclusive: a result containing both clean and a violation is inconsistent. Use needs_review for uncertainty. Thresholds come from calibration and policy, not the model's claim that its confidence is reliable.

python
import math


def route(classification, block_thresholds, allow_threshold):
    if not isinstance(classification, dict) or set(classification) != {"categories", "confidence", "reason"}:
        return "human_review"
    if not isinstance(classification["reason"], str):
        return "human_review"
    confidence = classification.get("confidence")
    categories = classification.get("categories")
    if (type(confidence) not in (int, float)
            or not 0 <= confidence <= 1 or not math.isfinite(confidence)):
        return "human_review"
    if not isinstance(categories, list) or not categories or not all(isinstance(c, str) for c in categories):
        return "human_review"
    categories = {category.casefold() for category in categories}
    if categories - (set(block_thresholds) | {"clean", "needs_review"}):
        return "human_review"
    if "needs_review" in categories or ("clean" in categories and len(categories) != 1):
        return "human_review"
    if categories == {"clean"}:
        return "pass" if confidence >= allow_threshold else "human_review"
    if any(confidence >= block_thresholds[category] for category in categories):
        return "reject"
    return "human_review"

Validate configuration thresholds as finite numbers in [0, 1] at startup. The example's category thresholds are policy-specific; it does not decide which categories your product must reject.

Evaluation and review

Use held-out labeled examples covering language, context, quoted material, benign mentions and adversarial inputs. Track precision, recall, appeal outcomes and per-category error cost. Neither false positives nor false negatives are always cheaper; the product policy determines that trade-off.

Send ambiguous cases to human review. Store decision metadata, policy/model versions and the minimum evidence needed for review under the application's retention and access controls. Do not indiscriminately log raw sensitive input.

Refresh evaluations when the policy, model or input distribution changes. Run an offline comparison before deploying a new route or threshold.

Reviewed 2026-09-23:

Use security-patterns for application input security, model-routing-patterns for model evaluation and prompt-caching-patterns for policy caching.

© softspark, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in app/skills/content-moderation-patterns of softspark/ai-toolkit.

Open the folder on GitHubat commit d64db2b

Compare with similar skills

Content Moderation Patterns next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Content Moderation Patterns compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Content Moderation Patterns this skillsoftspark/ai-toolkit179—~1.7kAutomated safety check: PassApache-2.0
ObliteratusRedWoodOG/Hermes-Desktop1776 repos~3.8kAutomated safety check: PassMIT
Aisafetyhotwuyoscar/AISafetyHot-Hub175—~1.2kAutomated safety check: PassCustom licence
Lemonade Router Builderamd/skills395—~4kAutomated safety check: PassMIT
Persona Designkangarooking/system-prompt-skills2051 repos~956Automated safety check: PassMIT
Execution Guardrailsmrtooher/fable-mode870—~1kAutomated safety check: PassNone

Similar skills

  • Obliteratus

    RedWoodOG/Hermes-Desktop

    Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…

    177 GitHub starsUsed in 6 repos~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Aisafetyhot

    wuyoscar/AISafetyHot-Hub

    Read AI Safety HOT daily digests, search recent AI safety research and incidents, and follow current hot topics.

    175 GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.

    395 GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Persona Design

    kangarooking/system-prompt-skills

    当需要为 AI 产品定义核心身份、角色声明和能力边界时调用此 skill。典型场景包括:设计新 AI 产品的 system prompt 首段、为不同场景创建差异化角色(如教学助手 vs 编程代理)、重新定义 AI 与用户的关系框架。

    205 GitHub starsUsed in 1 repo~956 tokens
    AI & LLM EngineeringAuto-check passed
  • Execution Guardrails

    mrtooher/fable-mode

    Always-on operational guardrails, model-independent. An agent skill from mrtooher/fable-mode.

    870 GitHub stars~1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Writing Eval Scenarios

    open-bias/open-bias

    Guide for writing eval conversation JSONs and running them through policy engines

    143 GitHub stars~1.5k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed

More from softspark/ai-toolkit

All 112 skills in this repo
  • Prepare Test Env

    softspark/ai-toolkit

    Prepare or verify a project QA environment with source identity, readiness, browser access, evidence paths and owned cleanup.

    179 GitHub stars~1.8k tokensUpdated today
    Auto-check: notes
  • A11y Validate

    softspark/ai-toolkit

    Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~3.8k tokensUpdated today
    Auto-check: notes
  • Analyze

    softspark/ai-toolkit

    Analyzes code quality, complexity, patterns across codebase.

    179 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Autonomous Dev

    softspark/ai-toolkit

    Drives a brief, specification, issue or existing PR through implementation, review, tests and QA to a ready PR.

    179 GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Brand Voice

    softspark/ai-toolkit

    Direct technical voice for docs, README, user-facing text. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • CI

    softspark/ai-toolkit

    Detect/generate/debug CI pipeline config (GitHub Actions, GitLab CI).

    179 GitHub stars~1.1k tokensUpdated today
    Auto-check: notes

Questions about Content Moderation Patterns

What does Content Moderation Patterns do?

Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL. Content Moderation Patterns is an agent skill from softspark/ai-toolkit. Content moderation with Claude: pre-filter vs LLM-classify, categories, thresholds, HITL.

When should I use Content Moderation Patterns?

Content Moderation Patterns fits situations like: tasks that involve LLM guardrails.

How do I install Content Moderation Patterns in Claude Code?

Run `npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a claude-code`. Or copy the skill folder (app/skills/content-moderation-patterns in softspark/ai-toolkit) into .claude/skills/content-moderation-patterns in your project. Claude Code loads it when a task matches its description.

How do I install Content Moderation Patterns in Codex?

Run `npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a codex`. Or copy the skill folder (app/skills/content-moderation-patterns in softspark/ai-toolkit) into .agents/skills/content-moderation-patterns in your project. Codex loads it when a task matches its description.

Can I use Content Moderation Patterns in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add softspark/ai-toolkit --skill content-moderation-patterns -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/content-moderation-patterns, .gemini/skills/content-moderation-patterns, .github/skills/content-moderation-patterns and .opencode/skills/content-moderation-patterns in your project.

What does Content Moderation Patterns need to run?

SKILL.md names no scripts, command-line tools or credentials: Content Moderation Patterns is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read.

Does Content Moderation Patterns access the network?

SKILL.md names 1 domain. As links in the text: platform.claude.com. This is read from the text; nothing was executed.

Is Content Moderation Patterns safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Content Moderation Patterns use?

Content Moderation Patterns is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Content Moderation Patterns use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Content Moderation Patterns?

Skills that share tags, products or a category with Content Moderation Patterns: Obliteratus (RedWoodOG/Hermes-Desktop, 177 stars), Aisafetyhot (wuyoscar/AISafetyHot-Hub, 175 stars), Lemonade Router Builder (amd/skills, 395 stars) and Persona Design (kangarooking/system-prompt-skills, 205 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Content Moderation Patterns?

softspark (a GitHub user) maintains it in softspark/ai-toolkit, which has 179 GitHub stars. The repository holds 112 skills in this directory. The repository was last updated on October 7, 2026.

Source: softspark/ai-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.