Obliteratus
RedWoodOG/Hermes-Desktop
Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…
Check user-generated text against a written policy before it is published.
$ npx skills add mrmps/classifier-dev --skill content-moderation-gate -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mrmps/classifier-dev content-moderation-gate --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/content-moderation-gate .claude/skills/content-moderation-gate && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "content-moderation-gate" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/content-moderation-gate into .claude/skills/content-moderation-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-gate", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mrmps/classifier-dev/tree/main/skills/content-moderation-gateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mrmps/classifier-dev --skill content-moderation-gate -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mrmps/classifier-dev content-moderation-gate --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/content-moderation-gate .agents/skills/content-moderation-gate && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "content-moderation-gate" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/content-moderation-gate into .agents/skills/content-moderation-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-gate", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mrmps/classifier-dev --skill content-moderation-gate -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mrmps/classifier-dev content-moderation-gate --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/content-moderation-gate .cursor/skills/content-moderation-gate && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "content-moderation-gate" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/content-moderation-gate into .cursor/skills/content-moderation-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-gate", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mrmps/classifier-dev.git --path skills/content-moderation-gate--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mrmps/classifier-dev --skill content-moderation-gate -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mrmps/classifier-dev content-moderation-gate --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/content-moderation-gate .gemini/skills/content-moderation-gate && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "content-moderation-gate" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/content-moderation-gate into .gemini/skills/content-moderation-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-gate", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mrmps/classifier-dev content-moderation-gateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mrmps/classifier-dev --skill content-moderation-gate -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/content-moderation-gate .github/skills/content-moderation-gate && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "content-moderation-gate" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/content-moderation-gate into .github/skills/content-moderation-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-gate", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mrmps/classifier-dev --skill content-moderation-gate -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mrmps/classifier-dev content-moderation-gate --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/content-moderation-gate .opencode/skills/content-moderation-gate && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "content-moderation-gate" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/content-moderation-gate into .opencode/skills/content-moderation-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "content-moderation-gate", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
content-moderation-gateCheck user-generated text against a written policy before it is published.
Content Moderation Gate is an agent skill from mrmps/classifier-dev. Check user-generated text against a written policy before it is published. One multi-label call scores every policy category independently, the band sets severity, only the middle band reaches a human queue, and hard-block categories are held whatever the score says. Use when a comment box, feed or chat product needs "moderation", "policy enforcement", "flag bad posts", or an answer to "what do we do with the borderline ones".
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM guardrails. The repository describes itself as: Zero-shot text classification over plain HTTP — no API key, no account. One Cloudflare Worker, a CLI, and an MCP server. https://classifier.dev. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b9211dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curljqFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
classifier.devFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Content Moderation Gate loads about 1.5k tokens when it runs. Until then it costs about 114 tokens; SKILL.md has 622 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mrmps/classifier-dev at commit b9211dd, republished under its MIT licence (© mrmps). 622 words, ~1,499 tokens.
.claude/skills/content-moderation-gate/SKILL.md (or your agent's skills folder).A keyword list misses anything written politely, and a model asked to "check
this post" returns a paragraph you cannot branch on. classifier.dev returns
a probability per category, so the branch is arithmetic.
One label per enforceable category, written as a moderator would say it rather
than as a section number: harassment or threats aimed at a person classifies,
policy 3.2(b) does not.
none of these, or a clean post lands in your nearest violation.policy.json:
{
"multi": true,
"max_labels": 3,
"labels": [
"harassment or threats aimed at a person",
"spam or commercial link bait",
"unsafe advice that could injure someone who follows it",
"sexual content involving a minor",
"self-harm or suicidal intent",
"none of these"
],
"instructions": "Apply the policy to the post itself, not to the topic it discusses. Quoting or reporting a violation is not a violation. Judge the post, not the author.",
"inputs": [
"everyone in this thread should log off, you are all worthless and I know where you work",
"MAKE $4000/WEEK FROM HOME >> click my profile link, limited spots",
"just double the dose if it is not working after an hour, the label is only a guideline",
"i do not see the point of any of this anymore and i have been thinking about ending it",
"you clearly have no idea what you are talking about and your code is bad",
"i reported a user yesterday for threatening another member, how long does moderation take",
"the sauce needs more acid, try a splash of vinegar at the end"
]
}curl -s https://classifier.dev/v1/classify -H 'content-type: application/json' --data @policy.json \
| jq -r '.results | to_entries[] | "post \(.key) " + (.value.scores | to_entries
| map(select(.key != "none of these")) | max_by(.value) | "\(.value) \(.key)")'Real output (up to 1,000 posts a call):
post 0 0.98 harassment or threats aimed at a person
post 1 0.99 spam or commercial link bait
post 2 0.96 unsafe advice that could injure someone who follows it
post 3 0.98 self-harm or suicidal intent
post 4 0.69 harassment or threats aimed at a person
post 5 0.03 harassment or threats aimed at a person
post 6 0.06 unsafe advice that could injure someone who follows itA real queue: four obvious, one arguable, two fine.
Take the highest policy score per post, ignoring none of these.
Hard block, above all of it. If sexual content involving a minor clears
the floor you set — start at 0.2, not 0.5 — the post is held for a human and
this gate never publishes it, whatever else scored. Same for a credible threat
of violence if your policy names one: a low floor, a one-way door, and none of
the bands above.
scores, not labelsThe labels array holds only categories at 0.7 and up, so post 4 reads clean
through it: score=0.67 labels=[]. Gate on scores, which carries every
category. none of these is a sanity check, not a gate: it ran from 0.20 on
the worst post to 0.56 on the cleanest and never reached 0.7.
instructionsThe highest-value edit is saying that reporting a violation is not one. Post 5
scored harassment 0.03 with that sentence in instructions and 0.34 with it
removed, same batch.
Keep it to one or two sentences. Pasting the policy document in there flattens every score toward the middle.
Counts and bands, never posts. One row per batch — day, label, band, count — and one per held post: post id, category, score, band, moderator verdict. The body stays in your own store under your own retention rule; it does not belong in a metrics table or an alert email.
It is also the calibration check: if weekly agreement between moderators and the 0.9 band falls under about 85%, rewrite the labels before you touch the thresholds.
instructions says it does not.Every post has a per-category score, a band and a decision. Hard-block categories have their own floor and never auto-publish, the queue holds only the middle band, and the week's agreement rate is on record.
© mrmps, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/content-moderation-gate of mrmps/classifier-dev.
Open the folder on GitHubat commit b9211dd
Content Moderation Gate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Content Moderation Gate this skillmrmps/classifier-dev | 424 | — | ~1.5k | Automated safety check: Pass | MIT | |
| ObliteratusRedWoodOG/Hermes-Desktop | 177 | 6 repos | ~3.8k | Automated safety check: Pass | MIT | |
| Aisafetyhotwuyoscar/AISafetyHot-Hub | 175 | — | ~1.2k | Automated safety check: Pass | Custom licence | |
| Lemonade Router Builderamd/skills | 395 | — | ~4k | Automated safety check: Pass | MIT | |
| Persona Designkangarooking/system-prompt-skills | 205 | 1 repos | ~956 | Automated safety check: Pass | MIT | |
| Execution Guardrailsmrtooher/fable-mode | 870 | — | ~1k | Automated safety check: Pass | None |
RedWoodOG/Hermes-Desktop
Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…
wuyoscar/AISafetyHot-Hub
Read AI Safety HOT daily digests, search recent AI safety research and incidents, and follow current hot topics.
amd/skills
Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.
kangarooking/system-prompt-skills
当需要为 AI 产品定义核心身份、角色声明和能力边界时调用此 skill。典型场景包括:设计新 AI 产品的 system prompt 首段、为不同场景创建差异化角色(如教学助手 vs 编程代理)、重新定义 AI 与用户的关系框架。
mrtooher/fable-mode
Always-on operational guardrails, model-independent. An agent skill from mrtooher/fable-mode.
open-bias/open-bias
Guide for writing eval conversation JSONs and running them through policy engines
mrmps/classifier-dev
Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.
mrmps/classifier-dev
Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.
mrmps/classifier-dev
Label each context chunk keep, drop or replace-with-a-pointer and pass the survivors through byte for byte instead of summarising, with key-shaped chunks decided locally and never sent, and a…
mrmps/classifier-dev
Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.
mrmps/classifier-dev
Filter hundreds or thousands of headlines, search results or feed items against a written brief before opening any of them, using a two-stage cascade that spends a fast model on everything and a…
mrmps/classifier-dev
Type candidate (subject, sentence, object) triples against a fixed relation schema and flag triples that contradict each other, batched, with a calibrated confidence per edge so only confident edges…
Categories
Check user-generated text against a written policy before it is published. Content Moderation Gate is an agent skill from mrmps/classifier-dev. Check user-generated text against a written policy before it is published.
Content Moderation Gate fits situations like: chat product needs moderation; policy enforcement; an answer to what do we do with the borderline ones.
Run `npx skills add mrmps/classifier-dev --skill content-moderation-gate -a claude-code`. Or copy the skill folder (skills/content-moderation-gate in mrmps/classifier-dev) into .claude/skills/content-moderation-gate in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mrmps/classifier-dev --skill content-moderation-gate -a codex`. Or copy the skill folder (skills/content-moderation-gate in mrmps/classifier-dev) into .agents/skills/content-moderation-gate in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mrmps/classifier-dev --skill content-moderation-gate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/content-moderation-gate, .gemini/skills/content-moderation-gate, .github/skills/content-moderation-gate and .opencode/skills/content-moderation-gate in your project.
Going by SKILL.md and its folder, Content Moderation Gate needs the command-line tools its instructions call (curl and jq).
SKILL.md names 1 domain. In commands or code: classifier.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Content Moderation Gate is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Content Moderation Gate: Obliteratus (RedWoodOG/Hermes-Desktop, 177 stars), Aisafetyhot (wuyoscar/AISafetyHot-Hub, 175 stars), Lemonade Router Builder (amd/skills, 395 stars) and Persona Design (kangarooking/system-prompt-skills, 205 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mrmps (a GitHub user) maintains it in mrmps/classifier-dev, which has 424 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 6, 2026.
Source: mrmps/classifier-dev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.