Agent skill

Content Moderation Gate

by mrmps in mrmps/classifier-dev

Check user-generated text against a written policy before it is published.

MITAuto-check passedAI & LLM Engineering

Install Content Moderation Gate

skills CLI
$ npx skills add mrmps/classifier-dev --skill content-moderation-gate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mrmps/classifier-dev content-moderation-gate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/content-moderation-gate .claude/skills/content-moderation-gate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
content-moderation-gate
GitHub stars
424
Token cost
~1.5k tokens
SKILL.md length
622 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Check user-generated text against a written policy before it is published.

  • Works in 6 steps: Policy document to labels → One call, multi-label → The gate → …
  • Chat product needs moderation
  • SKILL.md covers When not to use it, 1. Policy document to labels, 2. One call, multi-label and 3. The gate, plus 5 more sections
  • Calls curl and jq; reaches classifier.dev

What it does

Content Moderation Gate is an agent skill from mrmps/classifier-dev. Check user-generated text against a written policy before it is published. One multi-label call scores every policy category independently, the band sets severity, only the middle band reaches a human queue, and hard-block categories are held whatever the score says. Use when a comment box, feed or chat product needs "moderation", "policy enforcement", "flag bad posts", or an answer to "what do we do with the borderline ones".

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM guardrails. The repository describes itself as: Zero-shot text classification over plain HTTP — no API key, no account. One Cloudflare Worker, a CLI, and an MCP server. https://classifier.dev. The licence is MIT.

When your agent uses it

  • Chat product needs moderation
  • Policy enforcement
  • An answer to what do we do with the borderline ones

Example prompts

  • “moderation”
  • “policy enforcement”
  • “flag bad posts”
  • “/content-moderation-gate”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Policy document to labels
  2. One call, multi-label
  3. The gate
  4. Gate on scores, not labels
  5. Carve-outs go in instructions
  6. What you log

What it can do on your machine

Read from SKILL.md and the folder at commit b9211dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • classifier.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Content Moderation Gate loads about 1.5k tokens when it runs. Until then it costs about 114 tokens; SKILL.md has 622 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mrmps/classifier-dev at commit b9211dd, republished under its MIT licence (© mrmps). 622 words, ~1,499 tokens.

Download SKILL.mdSave it as .claude/skills/content-moderation-gate/SKILL.md (or your agent's skills folder).
name
content-moderation-gate
description
Check user-generated text against a written policy before it is published. One multi-label call scores every policy category independently, the band sets severity, only the middle band reaches a human queue, and hard-block categories are held whatever the score says. Use when a comment box, feed or chat product needs "moderation", "policy enforcement", "flag bad posts", or an answer to "what do we do with the borderline ones".
license
MIT

Gate a post against your policy

A keyword list misses anything written politely, and a model asked to "check this post" returns a paragraph you cannot branch on. classifier.dev returns a probability per category, so the branch is arithmetic.

When not to use it

  • As the only control. It gates publication; you still need reporting, appeals and someone who owns the policy.
  • For images, audio, video or links: it reads text.
  • For legal determinations or age verification.

1. Policy document to labels

One label per enforceable category, written as a moderator would say it rather than as a section number: harassment or threats aimed at a person classifies, policy 3.2(b) does not.

  • Say what the category is, not what it is near: "unsafe advice that could injure someone who follows it" beats "dangerous content".
  • Keep categories disjoint; overlapping labels split the score and both land in the middle.
  • Add none of these, or a clean post lands in your nearest violation.
  • Fix the list; changing it mid-week voids last week's numbers.

2. One call, multi-label

policy.json:

{
  "multi": true,
  "max_labels": 3,
  "labels": [
    "harassment or threats aimed at a person",
    "spam or commercial link bait",
    "unsafe advice that could injure someone who follows it",
    "sexual content involving a minor",
    "self-harm or suicidal intent",
    "none of these"
  ],
  "instructions": "Apply the policy to the post itself, not to the topic it discusses. Quoting or reporting a violation is not a violation. Judge the post, not the author.",
  "inputs": [
    "everyone in this thread should log off, you are all worthless and I know where you work",
    "MAKE $4000/WEEK FROM HOME >> click my profile link, limited spots",
    "just double the dose if it is not working after an hour, the label is only a guideline",
    "i do not see the point of any of this anymore and i have been thinking about ending it",
    "you clearly have no idea what you are talking about and your code is bad",
    "i reported a user yesterday for threatening another member, how long does moderation take",
    "the sauce needs more acid, try a splash of vinegar at the end"
  ]
}
curl -s https://classifier.dev/v1/classify -H 'content-type: application/json' --data @policy.json \
  | jq -r '.results | to_entries[] | "post \(.key)  " + (.value.scores | to_entries
      | map(select(.key != "none of these")) | max_by(.value) | "\(.value)  \(.key)")'

Real output (up to 1,000 posts a call):

post 0  0.98  harassment or threats aimed at a person
post 1  0.99  spam or commercial link bait
post 2  0.96  unsafe advice that could injure someone who follows it
post 3  0.98  self-harm or suicidal intent
post 4  0.69  harassment or threats aimed at a person
post 5  0.03  harassment or threats aimed at a person
post 6  0.06  unsafe advice that could injure someone who follows it

A real queue: four obvious, one arguable, two fine.

3. The gate

Take the highest policy score per post, ignoring none of these.

  • 0.9 and above — act. Severity follows the category: remove for harassment and unsafe advice, hold for spam, route self-harm to whatever support flow your product has rather than to a punishment.
  • 0.5 to 0.9 — human queue. Post 4 only. Staff moderators from how many posts land in this band per day.
  • Below 0.5 — publish, keeping the score for when a report arrives.

Hard block, above all of it. If sexual content involving a minor clears the floor you set — start at 0.2, not 0.5 — the post is held for a human and this gate never publishes it, whatever else scored. Same for a credible threat of violence if your policy names one: a low floor, a one-way door, and none of the bands above.

4. Gate on scores, not labels

The labels array holds only categories at 0.7 and up, so post 4 reads clean through it: score=0.67 labels=[]. Gate on scores, which carries every category. none of these is a sanity check, not a gate: it ran from 0.20 on the worst post to 0.56 on the cleanest and never reached 0.7.

Show full SKILL.md (232 more words)Show less

5. Carve-outs go in instructions

The highest-value edit is saying that reporting a violation is not one. Post 5 scored harassment 0.03 with that sentence in instructions and 0.34 with it removed, same batch.

Keep it to one or two sentences. Pasting the policy document in there flattens every score toward the middle.

6. What you log

Counts and bands, never posts. One row per batch — day, label, band, count — and one per held post: post id, category, score, band, moderator verdict. The body stays in your own store under your own retention rule; it does not belong in a metrics table or an alert email.

It is also the calibration check: if weekly agreement between moderators and the 0.9 band falls under about 85%, rewrite the labels before you touch the thresholds.

Pitfalls

  • Scores near a band edge move between runs. Post 4 came back 0.63, 0.64, 0.65 and 0.71 over four runs while post 0 held at 0.98. The band is the decision; no rule should turn on 0.69 against 0.71.
  • Quoted abuse scores like abuse unless instructions says it does not.
  • A 429 means hold the posts, not publish them.

What done looks like

Every post has a per-category score, a band and a decision. Hard-block categories have their own floor and never auto-publish, the queue holds only the middle band, and the week's agreement rate is on record.

© mrmps, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/content-moderation-gate of mrmps/classifier-dev.

Open the folder on GitHubat commit b9211dd

Compare with similar skills

Content Moderation Gate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Content Moderation Gate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Content Moderation Gate this skillmrmps/classifier-dev424—~1.5kAutomated safety check: PassMIT
ObliteratusRedWoodOG/Hermes-Desktop1776 repos~3.8kAutomated safety check: PassMIT
Aisafetyhotwuyoscar/AISafetyHot-Hub175—~1.2kAutomated safety check: PassCustom licence
Lemonade Router Builderamd/skills395—~4kAutomated safety check: PassMIT
Persona Designkangarooking/system-prompt-skills2051 repos~956Automated safety check: PassMIT
Execution Guardrailsmrtooher/fable-mode870—~1kAutomated safety check: PassNone

Similar skills

  • Obliteratus

    RedWoodOG/Hermes-Desktop

    Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…

    177 GitHub starsUsed in 6 repos~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Aisafetyhot

    wuyoscar/AISafetyHot-Hub

    Read AI Safety HOT daily digests, search recent AI safety research and incidents, and follow current hot topics.

    175 GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.

    395 GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Persona Design

    kangarooking/system-prompt-skills

    当需要为 AI 产品定义核心身份、角色声明和能力边界时调用此 skill。典型场景包括:设计新 AI 产品的 system prompt 首段、为不同场景创建差异化角色(如教学助手 vs 编程代理)、重新定义 AI 与用户的关系框架。

    205 GitHub starsUsed in 1 repo~956 tokens
    AI & LLM EngineeringAuto-check passed
  • Execution Guardrails

    mrtooher/fable-mode

    Always-on operational guardrails, model-independent. An agent skill from mrtooher/fable-mode.

    870 GitHub stars~1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Writing Eval Scenarios

    open-bias/open-bias

    Guide for writing eval conversation JSONs and running them through policy engines

    143 GitHub stars~1.5k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed

More from mrmps/classifier-dev

All 21 skills in this repo
  • Bulk Classify

    mrmps/classifier-dev

    Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.

    424 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Computer Use Action Picker

    mrmps/classifier-dev

    Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Label each context chunk keep, drop or replace-with-a-pointer and pass the survivors through byte for byte instead of summarising, with key-shaped chunks decided locally and never sent, and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Document Intake Routing

    mrmps/classifier-dev

    Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Headline Filter Map Reduce

    mrmps/classifier-dev

    Filter hundreds or thousands of headlines, search results or feed items against a written brief before opening any of them, using a two-stage cascade that spends a fast model on everything and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Type candidate (subject, sentence, object) triples against a fixed relation schema and flag triples that contradict each other, batched, with a calibrated confidence per edge so only confident edges…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Questions about Content Moderation Gate

What does Content Moderation Gate do?

Check user-generated text against a written policy before it is published. Content Moderation Gate is an agent skill from mrmps/classifier-dev. Check user-generated text against a written policy before it is published.

When should I use Content Moderation Gate?

Content Moderation Gate fits situations like: chat product needs moderation; policy enforcement; an answer to what do we do with the borderline ones.

How do I install Content Moderation Gate in Claude Code?

Run `npx skills add mrmps/classifier-dev --skill content-moderation-gate -a claude-code`. Or copy the skill folder (skills/content-moderation-gate in mrmps/classifier-dev) into .claude/skills/content-moderation-gate in your project. Claude Code loads it when a task matches its description.

How do I install Content Moderation Gate in Codex?

Run `npx skills add mrmps/classifier-dev --skill content-moderation-gate -a codex`. Or copy the skill folder (skills/content-moderation-gate in mrmps/classifier-dev) into .agents/skills/content-moderation-gate in your project. Codex loads it when a task matches its description.

Can I use Content Moderation Gate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mrmps/classifier-dev --skill content-moderation-gate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/content-moderation-gate, .gemini/skills/content-moderation-gate, .github/skills/content-moderation-gate and .opencode/skills/content-moderation-gate in your project.

What does Content Moderation Gate need to run?

Going by SKILL.md and its folder, Content Moderation Gate needs the command-line tools its instructions call (curl and jq).

Does Content Moderation Gate access the network?

SKILL.md names 1 domain. In commands or code: classifier.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Content Moderation Gate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Content Moderation Gate use?

Content Moderation Gate is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Content Moderation Gate use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Content Moderation Gate?

Skills that share tags, products or a category with Content Moderation Gate: Obliteratus (RedWoodOG/Hermes-Desktop, 177 stars), Aisafetyhot (wuyoscar/AISafetyHot-Hub, 175 stars), Lemonade Router Builder (amd/skills, 395 stars) and Persona Design (kangarooking/system-prompt-skills, 205 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Content Moderation Gate?

mrmps (a GitHub user) maintains it in mrmps/classifier-dev, which has 424 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 6, 2026.

Source: mrmps/classifier-dev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.