Agent skill

AI Security Hardening

by sickn33 in sickn33/agentic-awesome-skills

Harden AI/LLM deployments against prompt injection, data exfiltration, model theft, and supply chain attacks.

MITAuto-check passedSecurity

Install AI Security Hardening

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill ai-security-hardening -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills ai-security-hardening --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-security-hardening .claude/skills/ai-security-hardening && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-security-hardening
GitHub stars
47k
Used in
1 other repo
Token cost
~2.7k tokens
SKILL.md length
306 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Harden AI/LLM deployments against prompt injection, data exfiltration, model theft, and supply chain attacks.

  • Tasks that involve Prompt injection and agent security
  • SKILL.md covers When to Use This Skill, AI-Specific Threat Model, Prompt Injection Defense and Guardrails with NeMo Guardrails, plus 9 more sections
  • Calls pip; needs SECRET_KEY
  • Tasks that involve Supply chain security

What it does

AI Security Hardening is an agent skill from sickn33/agentic-awesome-skills. Harden AI/LLM deployments against prompt injection, data exfiltration, model theft, and supply chain attacks.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant security tooling (scanners, vault CLIs) and an authorized scope for any active assessment. Docs-only; helper scripts and templates not…

It sits in Security, covering Prompt injection and agent security and Supply chain security. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Prompt injection and agent security
  • Tasks that involve Supply chain security

Example prompts

  • “/ai-security-hardening”

Requirements

  • Python 3
  • A credential in SECRET_KEY
  • Compatibility (from SKILL.md): Requires the relevant security tooling (scanners, vault CLIs) and an authorized scope for any active assessment. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SECRET_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant security tooling (scanners, vault CLIs) and an authorized scope for any active assessment. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

AI Security Hardening loads about 2.7k tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 306 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 306 words, ~2,728 tokens.

Download SKILL.mdSave it as .claude/skills/ai-security-hardening/SKILL.md (or your agent's skills folder).
name
ai-security-hardening
description
Harden AI/LLM deployments against prompt injection, data exfiltration, model theft, and supply chain attacks.
compatibility
Requires the relevant security tooling (scanners, vault CLIs) and an authorized scope for any active assessment. Docs-only; helper scripts and templates not bundled.
category
security
risk
safe
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

AI Security Hardening

Secure LLM and AI systems against prompt injection, jailbreaks, data leakage, and supply chain threats in production environments.

When to Use This Skill

Use this skill when:

  • Deploying an LLM-powered application handling sensitive user data
  • Protecting against prompt injection attacks in AI agents
  • Implementing output filtering and content moderation
  • Securing model weights and API endpoints from theft
  • Achieving SOC2 or ISO 27001 compliance for AI systems

AI-Specific Threat Model

Threat                    Risk                          Control
─────────────────────────────────────────────────────────────────────
Prompt injection          System prompt override         Input sanitization, separate context
Data exfiltration         PII in model outputs           Output filtering, DLP scanning
Jailbreaking             Policy bypass                  Content moderation, guardrails
Model theft               Weight extraction via API      Rate limiting, access controls
Training data poisoning   Backdoored fine-tuned model    Dataset validation, provenance
Supply chain attack       Malicious model weights        Signature verification, scanning
Insecure output           XSS/SQLi from LLM response     Output encoding, parameterized queries

Prompt Injection Defense

python
import re
from typing import Optional

INJECTION_PATTERNS = [
    r"ignore\s+(all\s+)?(previous|prior|above)\s+instructions",
    r"you\s+are\s+now\s+",
    r"new\s+instructions?:",
    r"system\s+prompt",
    r"forget\s+everything",
    r"act\s+as\s+",
    r"jailbreak",
    r"dan\s+mode",
    r"<\s*system\s*>",
    r"\[INST\]",
]

def detect_prompt_injection(user_input: str) -> tuple[bool, Optional[str]]:
    """Return (is_suspicious, matched_pattern)."""
    normalized = user_input.lower().strip()
    for pattern in INJECTION_PATTERNS:
        if re.search(pattern, normalized, re.IGNORECASE):
            return True, pattern
    return False, None

def sanitize_user_input(user_input: str, max_length: int = 4000) -> str:
    """Sanitize input before passing to LLM."""
    # Truncate
    user_input = user_input[:max_length]

    # Remove null bytes and control characters
    user_input = re.sub(r'[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]', '', user_input)

    # Check for injection
    suspicious, pattern = detect_prompt_injection(user_input)
    if suspicious:
        raise ValueError(f"Potential prompt injection detected: {pattern}")

    return user_input

Guardrails with NeMo Guardrails

python
# guardrails.yaml
from nemoguardrails import RailsConfig, LLMRails

config = RailsConfig.from_path("./guardrails-config")
rails = LLMRails(config)

async def safe_llm_call(user_message: str) -> str:
    response = await rails.generate_async(
        messages=[{"role": "user", "content": user_message}]
    )
    return response["content"]
yaml
# guardrails-config/config.yml
models:
  - type: main
    engine: openai
    model: gpt-4o-mini

rails:
  input:
    flows:
      - check jailbreak
      - check sensitive data
  output:
    flows:
      - check output for PII
      - check output for harmful content

Output Filtering & PII Scrubbing

python
import re
from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine

analyzer = AnalyzerEngine()
anonymizer = AnonymizerEngine()

PII_ENTITIES = ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER", "CREDIT_CARD",
                "US_SSN", "IBAN_CODE", "IP_ADDRESS", "LOCATION"]

def scrub_pii_from_output(text: str) -> str:
    """Remove PII from LLM output before returning to user."""
    results = analyzer.analyze(text=text, entities=PII_ENTITIES, language="en")
    if not results:
        return text
    anonymized = anonymizer.anonymize(text=text, analyzer_results=results)
    return anonymized.text

def validate_output_safety(output: str) -> bool:
    """Check output doesn't contain prompt injection artifacts."""
    dangerous_patterns = [
        r"<\s*script\s*>",         # XSS
        r"javascript:",             # XSS
        r";\s*(DROP|DELETE|INSERT)",# SQLi
        r"\$\{.*\}",               # template injection
        r"`.*`",                   # command injection in some contexts
    ]
    for pattern in dangerous_patterns:
        if re.search(pattern, output, re.IGNORECASE):
            return False
    return True

API Security for LLM Endpoints

python
from fastapi import FastAPI, HTTPException, Depends, Request
from fastapi.security import HTTPBearer, HTTPAuthorizationCredentials
import jwt
import time
from collections import defaultdict

app = FastAPI()
security = HTTPBearer()

# Rate limiting (per API key)
request_counts = defaultdict(list)

def rate_limit(api_key: str, max_requests: int = 100, window_seconds: int = 60):
    now = time.time()
    requests = request_counts[api_key]
    # Remove old requests outside window
    request_counts[api_key] = [t for t in requests if now - t < window_seconds]
    if len(request_counts[api_key]) >= max_requests:
        raise HTTPException(status_code=429, detail="Rate limit exceeded")
    request_counts[api_key].append(now)

async def verify_token(
    credentials: HTTPAuthorizationCredentials = Depends(security)
) -> dict:
    try:
        payload = jwt.decode(credentials.credentials, SECRET_KEY, algorithms=["HS256"])
        rate_limit(payload["sub"])
        return payload
    except jwt.ExpiredSignatureError:
        raise HTTPException(status_code=401, detail="Token expired")
    except jwt.InvalidTokenError:
        raise HTTPException(status_code=401, detail="Invalid token")

@app.post("/v1/chat/completions")
async def chat(request: Request, token: dict = Depends(verify_token)):
    body = await request.json()

    # Input validation
    user_msg = body.get("messages", [{}])[-1].get("content", "")
    try:
        safe_input = sanitize_user_input(user_msg)
    except ValueError as e:
        raise HTTPException(status_code=400, detail=str(e))

    # Call LLM and scrub output
    response = await call_llm(safe_input, token["scope"])
    response["choices"][0]["message"]["content"] = scrub_pii_from_output(
        response["choices"][0]["message"]["content"]
    )
    return response

Model Weight Security

bash
# Verify model weights with SHA-256 hash before loading
MODEL_DIR="./models/llama-3.1-8b"
EXPECTED_HASH="sha256:abc123..."

# Generate hash of downloaded model
actual_hash=$(find "$MODEL_DIR" -name "*.safetensors" | sort | xargs sha256sum | sha256sum)
echo "Model hash: $actual_hash"

# Compare (automate in CI/CD)
if [ "$actual_hash" != "$EXPECTED_HASH" ]; then
  echo "ERROR: Model hash mismatch — possible tampering!"
  exit 1
fi

# Scan model files for embedded malware (ModelScan)
pip install modelscan
modelscan scan -p "$MODEL_DIR"

Network Isolation for AI Services

yaml
# Kubernetes NetworkPolicy — isolate LLM API
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: llm-api-isolation
  namespace: ai-services
spec:
  podSelector:
    matchLabels:
      app: vllm
  policyTypes:
  - Ingress
  - Egress
  ingress:
  - from:
    - namespaceSelector:
        matchLabels:
          name: backend           # only backend can call LLM
    ports:
    - protocol: TCP
      port: 8000
  egress:
  - to:
    - namespaceSelector:
        matchLabels:
          name: monitoring        # metrics only
    ports:
    - protocol: TCP
      port: 9090
  # Block egress to internet — prevent data exfiltration
  # (allow only internal cluster traffic)

Audit Logging

python
import structlog
from datetime import datetime, timezone

audit_log = structlog.get_logger("ai.audit")

def log_llm_interaction(
    user_id: str,
    session_id: str,
    model: str,
    prompt_tokens: int,
    completion_tokens: int,
    was_filtered: bool,
    injection_detected: bool,
):
    audit_log.info(
        "llm_interaction",
        timestamp=datetime.now(timezone.utc).isoformat(),
        user_id=user_id,
        session_id=session_id,
        model=model,
        prompt_tokens=prompt_tokens,
        completion_tokens=completion_tokens,
        was_filtered=was_filtered,
        injection_detected=injection_detected,
        # DO NOT log prompt/completion content — PII risk
    )

Common Issues

IssueCauseFix
False positive injection blocksOverly broad regexTune patterns; use ML-based classifier for high-traffic
PII in model outputsModel trained on PII dataAdd Presidio scrubbing to output layer
API key leakageKeys in logs or responsesMask keys in logging; use vault for key storage
Model weight tamperingUnverified downloadsAlways verify SHA-256; use modelscan
Rate limit bypassPer-IP not per-userRate limit on authenticated user ID, not IP

Best Practices

  • Never log raw prompts or completions — they may contain PII or sensitive data.
  • Treat LLM output as untrusted input — always encode before rendering in HTML.
  • Use network policies to prevent LLM pods from making outbound internet calls.
  • Rotate API keys quarterly; use short-lived JWT tokens for service-to-service auth.
  • Run modelscan on any model downloaded from the internet before serving.
  • hashicorp-vault (hashicorp-vault) - Secrets management for API keys
  • [network-security] (../../network/) - Network-level controls
  • linux-hardening (linux-hardening) - Host hardening
  • agent-observability (agent-observability) - AI audit logging
  • llm-gateway (llm-gateway) - Centralized access control

Limitations

  • Apply guidance only within authorized scope; test destructive steps in non-production first.
  • Docs-only import: upstream scripts and templates not bundled.
Example
bash
# Read-only first: inventory before any active step.
which <tool> && <tool> --help | head -n 20

Adapted from BagelHole/DevOps-Security-Agent-Skills (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ai-security-hardening of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

AI Security Hardening next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Security Hardening compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Security Hardening this skillsickn33/agentic-awesome-skills47k1 repos~2.7kAutomated safety check: PassMIT
Skill Scannergetsentry/skills1k4 repos~2.5kAutomated safety check: WarnApache-2.0
Kesekit Checkcdppcorp/KESE-KIT360—~1.3kAutomated safety check: PassMIT
Kesekit Fixcdppcorp/KESE-KIT360—~1.1kAutomated safety check: PassMIT
Kesekit Guidecdppcorp/KESE-KIT360—~1.4kAutomated safety check: PassMIT
Plugin Scanneriflytek/skillhub5.2k2 repos~1.1kAutomated safety check: NotesApache-2.0

Similar skills

  • Skill Scanner

    getsentry/skills

    Official

    Scan agent skills for security issues. An agent skill from getsentry/skills.

    1k GitHub starsUsed in 4 repos~2.5k tokens
    SecurityAuto-check: warnings
  • Kesekit Check

    cdppcorp/KESE-KIT

    Run a pre-deployment security compliance checklist based on KISA guidelines.

    360 GitHub stars~1.3k tokensUpdated 6 mo ago
    SecurityAuto-check passed
  • Kesekit Fix

    cdppcorp/KESE-KIT

    Auto-fix security vulnerabilities found in CII, AI, robot, space, and supply chain systems.

    360 GitHub stars~1.1k tokensUpdated 6 mo ago
    SecurityAuto-check passed
  • Kesekit Guide

    cdppcorp/KESE-KIT

    Generate secure coding prompts and guides for AI tools (Claude, ChatGPT, Cursor, Copilot).

    360 GitHub stars~1.4k tokensUpdated 6 mo ago
    SecurityAuto-check passed
  • Plugin Scanner

    iflytek/skillhub

    Scan AI agent skills, plugins, MCP servers, and agent tooling for prompt injection, unsafe commands, secret exposure, and supply-chain risks before installing or trusting them.

    5.2k GitHub starsUsed in 2 repos~1.1k tokens
    SecurityAuto-check: notes
  • Kesekit Start

    cdppcorp/KESE-KIT

    Run a security vulnerability assessment based on KISA guidelines.

    360 GitHub stars~2.3k tokensUpdated 6 mo ago
    SecurityAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Categories

Questions about AI Security Hardening

What does AI Security Hardening do?

Harden AI/LLM deployments against prompt injection, data exfiltration, model theft, and supply chain attacks. AI Security Hardening is an agent skill from sickn33/agentic-awesome-skills. Harden AI/LLM deployments against prompt injection, data exfiltration, model theft, and supply chain attacks.

When should I use AI Security Hardening?

AI Security Hardening fits situations like: tasks that involve Prompt injection and agent security; tasks that involve Supply chain security.

How do I install AI Security Hardening in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill ai-security-hardening -a claude-code`. Or copy the skill folder (skills/ai-security-hardening in sickn33/agentic-awesome-skills) into .claude/skills/ai-security-hardening in your project. Claude Code loads it when a task matches its description.

How do I install AI Security Hardening in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill ai-security-hardening -a codex`. Or copy the skill folder (skills/ai-security-hardening in sickn33/agentic-awesome-skills) into .agents/skills/ai-security-hardening in your project. Codex loads it when a task matches its description.

Can I use AI Security Hardening in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill ai-security-hardening -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-security-hardening, .gemini/skills/ai-security-hardening, .github/skills/ai-security-hardening and .opencode/skills/ai-security-hardening in your project.

What does AI Security Hardening need to run?

Going by SKILL.md and its folder, AI Security Hardening needs the command-line tools its instructions call (pip) and credentials named SECRET_KEY. Our summary lists: Python 3; A credential in SECRET_KEY. Compatibility (from SKILL.md): Requires the relevant security tooling (scanners, vault CLIs) and an authorized scope for any active assessment. Docs-only; helper scripts and templates not bundled..

Does AI Security Hardening access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is AI Security Hardening safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Security Hardening use?

AI Security Hardening is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Security Hardening use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Security Hardening?

Skills that share tags, products or a category with AI Security Hardening: Skill Scanner (getsentry/skills, 1k stars), Kesekit Check (cdppcorp/KESE-KIT, 360 stars), Kesekit Fix (cdppcorp/KESE-KIT, 360 stars) and Kesekit Guide (cdppcorp/KESE-KIT, 360 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Security Hardening?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.