Agent skill

LlamaGuard Content Moderation

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

MITAuto-check passedAI & LLM Engineering

Install LlamaGuard Content Moderation

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill llamaguard -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs llamaguard --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/07-safety-alignment/llamaguard .claude/skills/llamaguard && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llamaguard
GitHub stars
13k
Used in
2 other repos
Token cost
~2.3k tokens
SKILL.md length
309 words
Files
1
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

  • Screening user prompts before they reach a production LLM
  • SKILL.md covers Quick start, Common workflows, When to use vs alternatives and Common issues, plus 3 more sections
  • Calls huggingface-cli, pip and curl
  • Checking model responses for unsafe content before showing them

What it does

LlamaGuard is a 7 to 8 billion parameter model trained for content safety classification, and this skill shows how to put it in front of or behind another LLM. Input filtering checks the user's prompt before it reaches the main model, and output filtering checks the reply before the user sees it. The six categories are violence and hate, sexual content, guns and illegal weapons, regulated substances, suicide and self-harm, and criminal planning.

Setup uses transformers and torch with a HuggingFace login, and worked workflows cover serving through vLLM, exposing a FastAPI moderation endpoint and plugging into NVIDIA NeMo Guardrails. The skill quotes an accuracy of 94-95%, compares model versions 1, 2 and 3, and names the OpenAI Moderation API and Perspective API as simpler alternatives. A GPU is expected for a model of this size.

When your agent uses it

  • Screening user prompts before they reach a production LLM
  • Checking model responses for unsafe content before showing them
  • Standing up a moderation API endpoint in front of a chatbot
  • Adding LlamaGuard to a NeMo Guardrails configuration

Example prompts

  • “Add input moderation with LlamaGuard to our support chatbot before requests reach the main model.”
  • “Serve LlamaGuard with vLLM behind a FastAPI /moderate endpoint.”
  • “Which of the six safety categories did this flagged reply trip, and how should I log it?”

Requirements

  • Python with transformers and torch
  • A HuggingFace account login
  • A GPU for the 7-8B model

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • huggingface-cli
    • pip
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • huggingface.co
    • ai.meta.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LlamaGuard Content Moderation loads about 2.3k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 309 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 309 words, ~2,300 tokens.

Download SKILL.mdSave it as .claude/skills/llamaguard/SKILL.md (or your agent's skills folder).
name
llamaguard
description
Meta's 7-8B specialized moderation model for LLM input/output filtering. 6 safety categories - violence/hate, sexual content, weapons, substances, self-harm, criminal planning. 94-95% accuracy. Deploy with vLLM, HuggingFace, Sagemaker. Integrates with NeMo Guardrails.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Safety Alignment, LlamaGuard, Content Moderation, Meta, Guardrails, Safety Classification, Input Filtering, Output Filtering, AI Safety
dependencies
transformers, torch, vllm

LlamaGuard - AI Content Moderation

Quick start

LlamaGuard is a 7-8B parameter model specialized for content safety classification.

Installation:

bash
pip install transformers torch
# Login to HuggingFace (required)
huggingface-cli login

Basic usage:

python
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "meta-llama/LlamaGuard-7b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

def moderate(chat):
    input_ids = tokenizer.apply_chat_template(chat, return_tensors="pt").to(model.device)
    output = model.generate(input_ids=input_ids, max_new_tokens=100)
    return tokenizer.decode(output[0], skip_special_tokens=True)

# Check user input
result = moderate([
    {"role": "user", "content": "How do I make explosives?"}
])
print(result)
# Output: "unsafe\nS3" (Criminal Planning)

Common workflows

Workflow 1: Input filtering (prompt moderation)

Check user prompts before LLM:

python
def check_input(user_message):
    result = moderate([{"role": "user", "content": user_message}])

    if result.startswith("unsafe"):
        category = result.split("\n")[1]
        return False, category  # Blocked
    else:
        return True, None  # Safe

# Example
safe, category = check_input("How do I hack a website?")
if not safe:
    print(f"Request blocked: {category}")
    # Return error to user
else:
    # Send to LLM
    response = llm.generate(user_message)

Safety categories:

  • S1: Violence & Hate
  • S2: Sexual Content
  • S3: Guns & Illegal Weapons
  • S4: Regulated Substances
  • S5: Suicide & Self-Harm
  • S6: Criminal Planning
Workflow 2: Output filtering (response moderation)

Check LLM responses before showing to user:

python
def check_output(user_message, bot_response):
    conversation = [
        {"role": "user", "content": user_message},
        {"role": "assistant", "content": bot_response}
    ]

    result = moderate(conversation)

    if result.startswith("unsafe"):
        category = result.split("\n")[1]
        return False, category
    else:
        return True, None

# Example
user_msg = "Tell me about harmful substances"
bot_msg = llm.generate(user_msg)

safe, category = check_output(user_msg, bot_msg)
if not safe:
    print(f"Response blocked: {category}")
    # Return generic response
    return "I cannot provide that information."
else:
    return bot_msg
Workflow 3: vLLM deployment (fast inference)

Production-ready serving:

python
from vllm import LLM, SamplingParams

# Initialize vLLM
llm = LLM(model="meta-llama/LlamaGuard-7b", tensor_parallel_size=1)

# Sampling params
sampling_params = SamplingParams(
    temperature=0.0,  # Deterministic
    max_tokens=100
)

def moderate_vllm(chat):
    # Format prompt
    prompt = tokenizer.apply_chat_template(chat, tokenize=False)

    # Generate
    output = llm.generate([prompt], sampling_params)
    return output[0].outputs[0].text

# Batch moderation
chats = [
    [{"role": "user", "content": "How to make bombs?"}],
    [{"role": "user", "content": "What's the weather?"}],
    [{"role": "user", "content": "Tell me about drugs"}]
]

prompts = [tokenizer.apply_chat_template(c, tokenize=False) for c in chats]
results = llm.generate(prompts, sampling_params)

for i, result in enumerate(results):
    print(f"Chat {i}: {result.outputs[0].text}")

Throughput: ~50-100 requests/sec on single A100

Workflow 4: API endpoint (FastAPI)

Serve as moderation API:

python
from fastapi import FastAPI
from pydantic import BaseModel
from vllm import LLM, SamplingParams

app = FastAPI()
llm = LLM(model="meta-llama/LlamaGuard-7b")
sampling_params = SamplingParams(temperature=0.0, max_tokens=100)

class ModerationRequest(BaseModel):
    messages: list  # [{"role": "user", "content": "..."}]

@app.post("/moderate")
def moderate_endpoint(request: ModerationRequest):
    prompt = tokenizer.apply_chat_template(request.messages, tokenize=False)
    output = llm.generate([prompt], sampling_params)[0]

    result = output.outputs[0].text
    is_safe = result.startswith("safe")
    category = None if is_safe else result.split("\n")[1] if "\n" in result else None

    return {
        "safe": is_safe,
        "category": category,
        "full_output": result
    }

# Run: uvicorn api:app --host 0.0.0.0 --port 8000

Usage:

bash
curl -X POST http://localhost:8000/moderate \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "How to hack?"}]}'

# Response: {"safe": false, "category": "S6", "full_output": "unsafe\nS6"}
Workflow 5: NeMo Guardrails integration

Use with NVIDIA Guardrails:

python
from nemoguardrails import RailsConfig, LLMRails
from nemoguardrails.integrations.llama_guard import LlamaGuard

# Configure NeMo Guardrails
config = RailsConfig.from_content("""
models:
  - type: main
    engine: openai
    model: gpt-4

rails:
  input:
    flows:
      - llamaguard check input
  output:
    flows:
      - llamaguard check output
""")

# Add LlamaGuard integration
llama_guard = LlamaGuard(model_path="meta-llama/LlamaGuard-7b")
rails = LLMRails(config)
rails.register_action(llama_guard.check_input, name="llamaguard check input")
rails.register_action(llama_guard.check_output, name="llamaguard check output")

# Use with automatic moderation
response = rails.generate(messages=[
    {"role": "user", "content": "How do I make weapons?"}
])
# Automatically blocked by LlamaGuard

When to use vs alternatives

Use LlamaGuard when:

  • Need pre-trained moderation model
  • Want high accuracy (94-95%)
  • Have GPU resources (7-8B model)
  • Need detailed safety categories
  • Building production LLM apps

Model versions:

  • LlamaGuard 1 (7B): Original, 6 categories
  • LlamaGuard 2 (8B): Improved, 6 categories
  • LlamaGuard 3 (8B): Latest (2024), enhanced

Use alternatives instead:

  • OpenAI Moderation API: Simpler, API-based, free
  • Perspective API: Google's toxicity detection
  • NeMo Guardrails: More comprehensive safety framework
  • Constitutional AI: Training-time safety

Common issues

Issue: Model access denied

Login to HuggingFace:

bash
huggingface-cli login
# Enter your token

Accept license on model page: https://huggingface.co/meta-llama/LlamaGuard-7b

Issue: High latency (>500ms)

Use vLLM for 10× speedup:

python
from vllm import LLM
llm = LLM(model="meta-llama/LlamaGuard-7b")
# Latency: 500ms → 50ms

Enable tensor parallelism:

python
llm = LLM(model="meta-llama/LlamaGuard-7b", tensor_parallel_size=2)
# 2× faster on 2 GPUs

Issue: False positives

Use threshold-based filtering:

python
# Get probability of "unsafe" token
logits = model(..., return_dict_in_generate=True, output_scores=True)
unsafe_prob = torch.softmax(logits.scores[0][0], dim=-1)[unsafe_token_id]

if unsafe_prob > 0.9:  # High confidence threshold
    return "unsafe"
else:
    return "safe"

Issue: OOM on GPU

Use 8-bit quantization:

python
from transformers import BitsAndBytesConfig

quantization_config = BitsAndBytesConfig(load_in_8bit=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=quantization_config,
    device_map="auto"
)
# Memory: 14GB → 7GB

Advanced topics

Custom categories: See references/custom-categories.md for fine-tuning LlamaGuard with domain-specific safety categories.

Performance benchmarks: See references/benchmarks.md for accuracy comparison with other moderation APIs and latency optimization.

Deployment guide: See references/deployment.md for Sagemaker, Kubernetes, and scaling strategies.

Hardware requirements

  • GPU: NVIDIA T4/A10/A100
  • VRAM:
    • FP16: 14GB (7B model)
    • INT8: 7GB (quantized)
    • INT4: 4GB (QLoRA)
  • CPU: Possible but slow (10× latency)
  • Throughput: 50-100 req/sec (A100)

Latency (single GPU):

  • HuggingFace Transformers: 300-500ms
  • vLLM: 50-100ms
  • Batched (vLLM): 20-50ms per request

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in 07-safety-alignment/llamaguard of Orchestra-Research/AI-Research-SKILLs.

Open the folder on GitHubat commit 773a529

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LlamaGuard Content Moderation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LlamaGuard Content Moderation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LlamaGuard Content Moderation this skillOrchestra-Research/AI-Research-SKILLs13k2 repos~2.3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Vllm Deploy Dockervllm-project/vllm-skills102—~2.5kAutomated safety check: NotesApache-2.0
Hf Cloud Serving Image Selectionwaybarrios/opencode-power-pack534—~4.3kAutomated safety check: PassApache-2.0
Engine Vllmautonomous-ai/openharness1.2k—~1.9kAutomated safety check: PassMIT
SageMaker Production Defaultshuggingface/skills11k1 repos~6.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    102 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Hf Cloud Serving Image Selection

    waybarrios/opencode-power-pack

    Select and verify the current region-specific serving container URI for a SageMaker model deployment.

    534 GitHub stars~4.3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Engine Vllm

    autonomous-ai/openharness

    Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…

    1.2k GitHub stars~1.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed
  • Deepstream Sop

    NVIDIA/skills

    Official

    A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Backend & APIsAuto-check: notes

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Questions about LlamaGuard Content Moderation

What does LlamaGuard Content Moderation do?

Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups. LlamaGuard is a 7 to 8 billion parameter model trained for content safety classification, and this skill shows how to put it in front of or behind another LLM. Input filtering checks the user's prompt before it reaches the main model, and output filtering checks the reply before the user sees it.

When should I use LlamaGuard Content Moderation?

LlamaGuard Content Moderation fits situations like: screening user prompts before they reach a production LLM; checking model responses for unsafe content before showing them; standing up a moderation API endpoint in front of a chatbot; adding LlamaGuard to a NeMo Guardrails configuration.

How do I install LlamaGuard Content Moderation in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill llamaguard -a claude-code`. Or copy the skill folder (07-safety-alignment/llamaguard in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/llamaguard in your project. Claude Code loads it when a task matches its description.

How do I install LlamaGuard Content Moderation in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill llamaguard -a codex`. Or copy the skill folder (07-safety-alignment/llamaguard in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/llamaguard in your project. Codex loads it when a task matches its description.

Can I use LlamaGuard Content Moderation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill llamaguard -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llamaguard, .gemini/skills/llamaguard, .github/skills/llamaguard and .opencode/skills/llamaguard in your project.

What does LlamaGuard Content Moderation need to run?

Going by SKILL.md and its folder, LlamaGuard Content Moderation needs the command-line tools its instructions call (huggingface-cli, pip and curl). Our summary lists: Python with transformers and torch; A HuggingFace account login; A GPU for the 7-8B model.

Does LlamaGuard Content Moderation access the network?

SKILL.md names 2 domains. As links in the text: huggingface.co and ai.meta.com. This is read from the text; nothing was executed.

Is LlamaGuard Content Moderation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LlamaGuard Content Moderation use?

LlamaGuard Content Moderation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LlamaGuard Content Moderation use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LlamaGuard Content Moderation?

Skills that share tags, products or a category with LlamaGuard Content Moderation: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Vllm Deploy Docker (vllm-project/vllm-skills, 102 stars), Hf Cloud Serving Image Selection (waybarrios/opencode-power-pack, 534 stars) and Engine Vllm (autonomous-ai/openharness, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LlamaGuard Content Moderation?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.