Agent skill

Implementing LLM Guardrails For Security

by mukul975 in mukul975/Anthropic-Cybersecurity-Skills

Implements input/output validation guardrails for LLM applications using NVIDIA NeMo Guardrails (Colang), custom Python validators for PII detection, and the Guardrails AI framework, intercepting…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Implementing LLM Guardrails For Security

skills CLI
$ npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill implementing-llm-guardrails-for-security -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mukul975/Anthropic-Cybersecurity-Skills implementing-llm-guardrails-for-security --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/implementing-llm-guardrails-for-security .claude/skills/implementing-llm-guardrails-for-security && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
implementing-llm-guardrails-for-security
GitHub stars
34k
Token cost
~2.2k tokens
SKILL.md length
599 words
Files
4 (incl. scripts, references)
Skills in repo
644
Repo updated
First seen
Licence
Apache-2.0

At a glance

Implements input/output validation guardrails for LLM applications using NVIDIA NeMo Guardrails (Colang), custom Python validators for PII detection, and the Guardrails AI framework, intercepting…

  • Works in 6 steps: Install Guardrail Frameworks → Run the Guardrails Security Agent → Configure Content Policies → …
  • Adding safety controls to an LLM app/chatbot/RAG pipeline
  • SKILL.md covers When to Use, Prerequisites, Workflow and Verification, plus 2 more sections
  • Runs Python scripts from its folder; calls python and pip; needs OPENAI_API_KEY

What it does

Implementing LLM Guardrails For Security is an agent skill from mukul975/Anthropic-Cybersecurity-Skills. Implements input/output validation guardrails for LLM applications using NVIDIA NeMo Guardrails (Colang), custom Python validators for PII detection, and the Guardrails AI framework, intercepting user inputs (prompt injection, PII, off-topic queries) and model outputs (hallucinations, toxic content, schema compliance). Use when adding safety controls to an LLM app/chatbot/RAG pipeline or validating outputs conform to expected schemas.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/api-reference.md` and `scripts/agent.py`).

It sits in AI & LLM Engineering, covering LLM guardrails. It works with NVIDIA AI Platform and Python. The repository describes itself as: 817 structured cybersecurity skills for AI agents · Mapped to 6 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND, NIST AI RMF & MITRE F3 (Fight Fraud) · agentskills.io…. The licence is Apache-2.0.

When your agent uses it

  • Adding safety controls to an LLM app/chatbot/RAG pipeline
  • Validating outputs conform to expected schemas

Example prompts

  • “Use the implementing-llm-guardrails-for-security skill to implement input/output validation guardrails for LLM applications using NVIDIA NeMo…”
  • “/implementing-llm-guardrails-for-security”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Install Guardrail Frameworks
  2. Run the Guardrails Security Agent
  3. Configure Content Policies
  4. Integrate NeMo Guardrails with Colang
  5. Deploy as a Validation Middleware
  6. Monitor Guardrail Effectiveness

What it can do on your machine

Read from SKILL.md and the folder at commit 54a7988. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Implementing LLM Guardrails For Security loads about 2.2k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 599 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mukul975/Anthropic-Cybersecurity-Skills at commit 54a7988, republished under its Apache-2.0 licence (© mukul975). 599 words, ~2,223 tokens.

Download SKILL.mdSave it as .claude/skills/implementing-llm-guardrails-for-security/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
implementing-llm-guardrails-for-security
description
Implements input/output validation guardrails for LLM applications using NVIDIA NeMo Guardrails (Colang), custom Python validators for PII detection, and the Guardrails AI framework, intercepting user inputs (prompt injection, PII, off-topic queries) and model outputs (hallucinations, toxic content, schema compliance). Use when adding safety controls to an LLM app/chatbot/RAG pipeline or validating outputs conform to expected schemas.
domain
cybersecurity
subdomain
ai-security
tags
LLM-guardrails, NeMo-Guardrails, input-validation, output-filtering, AI-safety
version
1.0.0
author
mukul975
license
Apache-2.0
atlas_techniques
AML.T0051, AML.T0054, AML.T0056, AML.T0057, AML.T0062
nist_ai_rmf
GOVERN-1.1, GOVERN-6.1, MEASURE-2.7, MEASURE-2.5, MANAGE-2.4
d3fend_techniques
Content Validation, Content Filtering, Content Excision, Application Hardening, Execution Isolation
nist_csf
GV.OC-03, ID.RA-01, PR.PS-01, DE.AE-02

Implementing LLM Guardrails for Security

When to Use

  • Deploying a new LLM-powered application that processes user input and needs input/output safety controls
  • Adding content policy enforcement to an existing chatbot or AI agent to comply with organizational policies
  • Implementing PII detection and redaction in LLM pipelines handling sensitive customer data
  • Building topic-restricted AI assistants that must refuse off-topic or disallowed queries
  • Validating that LLM responses conform to expected schemas before they reach downstream systems or users
  • Protecting RAG pipelines from indirect prompt injection in retrieved documents

Do not use as a replacement for proper authentication, authorization, and network security controls. Guardrails are a defense-in-depth layer, not a perimeter defense. Not suitable for real-time content moderation of user-to-user communication without LLM involvement.

Prerequisites

  • Python 3.10+ with pip for installing guardrail dependencies
  • An OpenAI API key or local LLM endpoint for NeMo Guardrails self-check rails (set as OPENAI_API_KEY environment variable)
  • The nemoguardrails package for Colang-based guardrail definitions
  • The guardrails-ai package for structured output validation (optional, for JSON schema enforcement)
  • Familiarity with YAML configuration and basic Colang 2.0 syntax for defining rail flows

Workflow

Step 1: Install Guardrail Frameworks

Install the required Python packages:

bash
# Core NeMo Guardrails library
pip install nemoguardrails

# Guardrails AI for structured output validation (optional)
pip install guardrails-ai

# Additional dependencies for PII detection and content analysis
pip install presidio-analyzer presidio-anonymizer spacy
python -m spacy download en_core_web_lg
Step 2: Run the Guardrails Security Agent

The agent implements a complete input/output validation pipeline:

bash
# Analyze a single input through all guardrail layers
python agent.py --input "Tell me how to hack into a system"

# Analyze input with a custom content policy file
python agent.py --input "Some text" --policy policy.json

# Scan a file of prompts through the guardrail pipeline
python agent.py --file prompts.txt --mode full

# Input-only validation (no LLM call, just check if input is safe)
python agent.py --input "Some text" --mode input-only

# Output validation mode (validate a pre-generated LLM response)
python agent.py --input "User question" --response "LLM response to validate" --mode output-only

# PII detection and redaction mode
python agent.py --input "My SSN is 123-45-6789 and email john@example.com" --mode pii

# JSON output for pipeline integration
python agent.py --file prompts.txt --output json
Step 3: Configure Content Policies

Create a JSON policy file defining allowed topics, blocked patterns, and PII categories:

json
{
  "allowed_topics": ["customer_support", "product_info", "billing"],
  "blocked_topics": ["politics", "violence", "illegal_activities", "competitor_products"],
  "blocked_patterns": ["how to hack", "create malware", "bypass security"],
  "pii_categories": ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER", "US_SSN", "CREDIT_CARD"],
  "max_output_length": 2000,
  "require_grounded_response": true
}
Step 4: Integrate NeMo Guardrails with Colang

Create a NeMo Guardrails configuration directory with config.yml and Colang flow files:

yaml
# config.yml
models:
  - type: main
    engine: openai
    model: gpt-4o-mini

rails:
  input:
    flows:
      - self check input
      - check jailbreak
      - mask sensitive data on input
  output:
    flows:
      - self check output
      - check hallucination
colang
# rails.co - Colang 2.0 flow definitions
define user ask about hacking
  "How do I hack into a system"
  "Tell me how to break into a network"
  "How to exploit vulnerabilities"

define bot refuse hacking request
  "I cannot provide instructions on unauthorized hacking or security exploitation.
   If you are interested in cybersecurity, I can suggest legitimate learning resources
   and ethical hacking certifications."

define flow
  user ask about hacking
  bot refuse hacking request
Step 5: Deploy as a Validation Middleware

Integrate the guardrails into your application as middleware:

python
from agent import GuardrailsPipeline

pipeline = GuardrailsPipeline(policy_path="policy.json")

# Pre-LLM input validation
input_result = pipeline.validate_input("user message here")
if not input_result["safe"]:
    return input_result["blocked_reason"]

# Post-LLM output validation
llm_response = your_llm.generate(input_result["sanitized_input"])
output_result = pipeline.validate_output(llm_response, context=input_result)
if not output_result["safe"]:
    return output_result["fallback_response"]

return output_result["validated_response"]
Step 6: Monitor Guardrail Effectiveness

Review guardrail logs to track block rates, false positives, and bypass attempts:

bash
# Generate a summary report from guardrail logs
python agent.py --file interaction_logs.txt --mode full --output json > guardrail_audit.json

Verification

  • Input guardrails correctly block known prompt injection patterns (system override, role-play escape, delimiter injection)
  • PII detection identifies and redacts email addresses, phone numbers, SSNs, and credit card numbers in user inputs
  • Topic restriction guardrails refuse off-policy queries and allow on-policy queries without false positives
  • Output guardrails detect and flag responses containing toxic content, PII leakage, or off-topic material
  • The guardrails pipeline adds less than 200ms of latency to the request/response cycle for input-only validation
  • JSON output mode produces valid, parseable JSON suitable for downstream monitoring dashboards
Show full SKILL.md (237 more words)Show less

Key Concepts

TermDefinition
Input RailA guardrail that intercepts and validates user input before it reaches the LLM, blocking injection attempts and redacting sensitive data
Output RailA guardrail that validates LLM-generated output before it reaches the user, filtering toxic content and enforcing schema compliance
ColangNVIDIA's domain-specific language for defining conversational guardrail flows, with Python-like syntax for specifying user intent patterns and bot responses
PII RedactionThe process of detecting and masking personally identifiable information (names, emails, SSNs) in text before processing
Content PolicyA configuration file defining which topics, patterns, and content categories are allowed or blocked by the guardrail system
Self-Check RailA NeMo Guardrails technique where the LLM itself evaluates whether its input or output violates defined policies
Hallucination DetectionOutput validation that checks whether the LLM response is grounded in the provided context, flagging fabricated claims

Tools & Systems

  • NVIDIA NeMo Guardrails: Open-source toolkit for adding programmable input, dialog, and output rails to LLM applications using Colang flow definitions and YAML configuration
  • Guardrails AI: Python framework for structured output validation with a hub of pre-built validators for PII, toxicity, JSON schema compliance, and more
  • Microsoft Presidio: Open-source PII detection and anonymization engine supporting 30+ entity types with configurable NLP backends
  • Colang 2.0: Event-driven interaction modeling language for defining guardrail flows with Python-like syntax, supporting multi-turn dialog control
  • OpenAI Guardrails Python: OpenAI's client-side guardrails library for prompt injection detection and content policy enforcement

© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/implementing-llm-guardrails-for-security of mukul975/Anthropic-Cybersecurity-Skills.

  • SKILL.md
  • LICENSE
  • references/api-reference.md
  • scripts/agent.py

Open the folder on GitHubat commit 54a7988

Compare with similar skills

Implementing LLM Guardrails For Security next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Implementing LLM Guardrails For Security compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Implementing LLM Guardrails For Security this skillmukul975/Anthropic-Cybersecurity-Skills34k—~2.2kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Sponsio Agent Safety SetupSponsioLabs/Sponsio454—~12kAutomated safety check: PassApache-2.0
Onboard Jetpack5 Inference BackendsEGalahad/sim2real146—~1.1kAutomated safety check: PassNone
Sglang Diffusion Modelopt Quantsgl-project/sglang37k2 repos~5kAutomated safety check: PassApache-2.0

Similar skills

  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Sponsio Agent Safety Setup

    SponsioLabs/Sponsio

    Installs, tunes and enforces Sponsio contracts that block unsafe tool calls in LLM agents, covering setup, auditing, observe mode and flipping to enforce.

    454 GitHub stars~12k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

    146 GitHub stars~1.1k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

    37k GitHub starsUsed in 2 repos~5k tokens
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from mukul975/Anthropic-Cybersecurity-Skills

All 644 skills in this repo
  • Campaign Attribution Evidence Analysis

    mukul975/Anthropic-Cybersecurity-Skills

    Weighs infrastructure, TTP, malware code and timing evidence with the Diamond Model and competing hypotheses to reach a confidence-rated attribution.

    34k GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Go Malware Analysis in Ghidra

    mukul975/Anthropic-Cybersecurity-Skills

    Walks through reverse engineering Go-compiled malware in Ghidra: parsing buildinfo and pclntab, recovering stripped function names and extracting dependencies.

    34k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • LNK and Jump List Forensics

    mukul975/Anthropic-Cybersecurity-Skills

    Guides forensic analysis of Windows LNK shortcut files and Jump Lists with LECmd, JLECmd and manual parsing to show file access and program execution.

    34k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Malware Persistence Analysis with Autoruns

    mukul975/Anthropic-Cybersecurity-Skills

    Hunts Windows malware persistence with Sysinternals Autoruns, covering run keys, services, scheduled tasks and drivers, with baseline comparison.

    34k GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • NTFS MFT Deleted File Recovery

    mukul975/Anthropic-Cybersecurity-Skills

    Guides a Windows forensic examination of the NTFS Master File Table to recover deleted-file evidence, build timelines and spot timestomping.

    34k GitHub stars~2.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Network Covert Channel Analysis

    mukul975/Anthropic-Cybersecurity-Skills

    Detects DNS tunneling, ICMP exfiltration and HTTP-based covert channels in packet captures and DNS logs when hunting for hidden command-and-control traffic.

    34k GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Implementing LLM Guardrails For Security

What does Implementing LLM Guardrails For Security do?

Implements input/output validation guardrails for LLM applications using NVIDIA NeMo Guardrails (Colang), custom Python validators for PII detection, and the Guardrails AI framework, intercepting…. Implementing LLM Guardrails For Security is an agent skill from mukul975/Anthropic-Cybersecurity-Skills. Implements input/output validation guardrails for LLM applications using NVIDIA NeMo Guardrails (Colang), custom Python validators for PII detection, and the Guardrails AI framework, intercepting user inputs (prompt injection, PII, off-topic queries) and model outputs (hallucinations, toxic content, schema compliance).

When should I use Implementing LLM Guardrails For Security?

Implementing LLM Guardrails For Security fits situations like: adding safety controls to an LLM app/chatbot/RAG pipeline; validating outputs conform to expected schemas.

How do I install Implementing LLM Guardrails For Security in Claude Code?

Run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill implementing-llm-guardrails-for-security -a claude-code`. Or copy the skill folder (skills/implementing-llm-guardrails-for-security in mukul975/Anthropic-Cybersecurity-Skills) into .claude/skills/implementing-llm-guardrails-for-security in your project. Claude Code loads it when a task matches its description.

How do I install Implementing LLM Guardrails For Security in Codex?

Run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill implementing-llm-guardrails-for-security -a codex`. Or copy the skill folder (skills/implementing-llm-guardrails-for-security in mukul975/Anthropic-Cybersecurity-Skills) into .agents/skills/implementing-llm-guardrails-for-security in your project. Codex loads it when a task matches its description.

Can I use Implementing LLM Guardrails For Security in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill implementing-llm-guardrails-for-security -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/implementing-llm-guardrails-for-security, .gemini/skills/implementing-llm-guardrails-for-security, .github/skills/implementing-llm-guardrails-for-security and .opencode/skills/implementing-llm-guardrails-for-security in your project.

What does Implementing LLM Guardrails For Security need to run?

Going by SKILL.md and its folder, Implementing LLM Guardrails For Security needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Implementing LLM Guardrails For Security access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Implementing LLM Guardrails For Security safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Implementing LLM Guardrails For Security use?

Implementing LLM Guardrails For Security is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Implementing LLM Guardrails For Security use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to Implementing LLM Guardrails For Security?

Skills that share tags, products or a category with Implementing LLM Guardrails For Security: Dstack Prototyping (dstackai/dstack, 2.3k stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars), Sponsio Agent Safety Setup (SponsioLabs/Sponsio, 454 stars) and Onboard Jetpack5 Inference Backends (EGalahad/sim2real, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Implementing LLM Guardrails For Security?

mukul975 (a GitHub user) maintains it in mukul975/Anthropic-Cybersecurity-Skills, which has 34,116 GitHub stars. The repository holds 644 skills in this directory. The repository was last updated on August 31, 2026.

Source: mukul975/Anthropic-Cybersecurity-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.