Prompt Guard
Orchestra-Research/AI-Research-SKILLs
Meta's 86M prompt injection and jailbreak detector. An agent skill from Orchestra-Research/AI-Research-SKILLs.
Agent skill
Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM…
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-indirect-prompt-injection -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mukul975/Anthropic-Cybersecurity-Skills detecting-indirect-prompt-injection --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/detecting-indirect-prompt-injection .claude/skills/detecting-indirect-prompt-injection && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "detecting-indirect-prompt-injection" agent skill from https://github.com/mukul975/Anthropic-Cybersecurity-Skills/tree/main/skills/detecting-indirect-prompt-injection into .claude/skills/detecting-indirect-prompt-injection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "detecting-indirect-prompt-injection", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mukul975/Anthropic-Cybersecurity-Skills/tree/main/skills/detecting-indirect-prompt-injectionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-indirect-prompt-injection -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mukul975/Anthropic-Cybersecurity-Skills detecting-indirect-prompt-injection --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/detecting-indirect-prompt-injection .agents/skills/detecting-indirect-prompt-injection && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "detecting-indirect-prompt-injection" agent skill from https://github.com/mukul975/Anthropic-Cybersecurity-Skills/tree/main/skills/detecting-indirect-prompt-injection into .agents/skills/detecting-indirect-prompt-injection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "detecting-indirect-prompt-injection", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-indirect-prompt-injection -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mukul975/Anthropic-Cybersecurity-Skills detecting-indirect-prompt-injection --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/detecting-indirect-prompt-injection .cursor/skills/detecting-indirect-prompt-injection && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "detecting-indirect-prompt-injection" agent skill from https://github.com/mukul975/Anthropic-Cybersecurity-Skills/tree/main/skills/detecting-indirect-prompt-injection into .cursor/skills/detecting-indirect-prompt-injection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "detecting-indirect-prompt-injection", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git --path skills/detecting-indirect-prompt-injection--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-indirect-prompt-injection -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mukul975/Anthropic-Cybersecurity-Skills detecting-indirect-prompt-injection --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/detecting-indirect-prompt-injection .gemini/skills/detecting-indirect-prompt-injection && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "detecting-indirect-prompt-injection" agent skill from https://github.com/mukul975/Anthropic-Cybersecurity-Skills/tree/main/skills/detecting-indirect-prompt-injection into .gemini/skills/detecting-indirect-prompt-injection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "detecting-indirect-prompt-injection", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mukul975/Anthropic-Cybersecurity-Skills detecting-indirect-prompt-injectionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-indirect-prompt-injection -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/detecting-indirect-prompt-injection .github/skills/detecting-indirect-prompt-injection && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "detecting-indirect-prompt-injection" agent skill from https://github.com/mukul975/Anthropic-Cybersecurity-Skills/tree/main/skills/detecting-indirect-prompt-injection into .github/skills/detecting-indirect-prompt-injection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "detecting-indirect-prompt-injection", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-indirect-prompt-injection -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mukul975/Anthropic-Cybersecurity-Skills detecting-indirect-prompt-injection --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mukul975/Anthropic-Cybersecurity-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/detecting-indirect-prompt-injection .opencode/skills/detecting-indirect-prompt-injection && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "detecting-indirect-prompt-injection" agent skill from https://github.com/mukul975/Anthropic-Cybersecurity-Skills/tree/main/skills/detecting-indirect-prompt-injection into .opencode/skills/detecting-indirect-prompt-injection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "detecting-indirect-prompt-injection", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
detecting-indirect-prompt-injectionDetect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM…
Detecting Indirect Prompt Injection is an agent skill from mukul975/Anthropic-Cybersecurity-Skills. Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM Guard's PromptInjection scanner or Hugging Face Prompt Guard 2. Use when an agent ingests untrusted external content and you need to screen it for injected instructions before the LLM processes it.
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/api-reference.md`, `references/standards.md` and `scripts/agent.py`).
It sits in Security, covering Prompt injection and agent security, LLM guardrails and Database schema design. It works with Hugging Face. The repository describes itself as: 817 structured cybersecurity skills for AI agents · Mapped to 6 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND, NIST AI RMF & MITRE F3 (Fight Fraud) · agentskills.io…. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 54a7988. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pippythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comhuggingface.cocrummy.comatlas.mitre.orggenai.owasp.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Detecting Indirect Prompt Injection loads about 2.8k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 107 tokens; SKILL.md has 834 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
# Debian/Ubuntu: sudo apt-get install -y tesseract-ocrZERO_WIDTH = dict.fromkeys(map(ord, "⟨U+200B⟩⟨U+200C⟩⟨U+200D⟩⟨U+2060⟩⟨U+FEFF⟩"), None)Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from mukul975/Anthropic-Cybersecurity-Skills at commit 54a7988, republished under its Apache-2.0 licence (© mukul975). 834 words, ~2,767 tokens.
.claude/skills/detecting-indirect-prompt-injection/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Authorized-use-only notice: Scripts in this skill scan untrusted content for injection payloads and run detector models. Run scanning only on data you are authorized to process, and treat any extracted payloads as live untrusted input — never paste them back into a privileged LLM context.
Indirect prompt injection (MITRE ATLAS AML.T0051.001, OWASP LLM01:2025) occurs when an LLM-powered agent ingests external content — a web page it browses, a PDF or email it summarizes, an image it OCRs, a tool result it reads — and that content contains hidden instructions the model then follows as if they came from the developer or user. Because the agent treats all tokens in its context window as equally authoritative, an attacker who controls any consumed artifact can hijack the agent's behavior: exfiltrate conversation history, redirect tool calls, leak secrets, or pivot through connected systems.
Unlike direct injection (the user types the attack), indirect injection arrives through a trusted-looking data channel, which is why naive input filtering misses it. Payloads hide in many forms: HTML comments and display:none/zero-width text on web pages, white-on-white or tiny-font text in PDFs, alt-text and EXIF metadata in images, text rendered into pixels (invisible to OCR-light filters but read by multimodal models), Unicode tag/zero-width characters, and Base64/ROT13 obfuscation. This skill builds a detection pipeline that normalizes and scans every artifact before it reaches the model, combining heuristic/regex detection, dedicated detector models (Meta Prompt Guard 2, ProtectAI's deberta-v3 prompt-injection classifier via LLM Guard), and multimodal extraction for images, and then defines response actions and detection telemetry.
python -m venv .venv && source .venv/bin/activate
# LLM Guard — input/output scanners incl. PromptInjection
pip install llm-guard
# Hugging Face transformers for Prompt Guard 2 / deberta classifiers
pip install transformers torch
# Content extraction: HTML, PDF, images
pip install beautifulsoup4 pypdf pillow pytesseract
# pytesseract requires the Tesseract OCR engine:
# Debian/Ubuntu: sudo apt-get install -y tesseract-ocr
# macOS: brew install tesseract
# Windows: choco install tesseractmeta-llama/Llama-Prompt-Guard-2-86M on Hugging Face, or use the open protectai/deberta-v3-base-prompt-injection-v2 classifier.| ID | Official Name | Relevance |
|---|---|---|
| AML.T0051.001 | LLM Prompt Injection: Indirect | The exact technique this skill detects and mitigates |
| AML.T0051 | LLM Prompt Injection | Parent technique covering all prompt-injection variants |
| AML.T0057 | LLM Data Leakage | Common objective of an indirect injection that this detection prevents |
| AML.T0053 | LLM Plugin Compromise | Injected instructions frequently target the agent's tools/plugins |
Pull comments, hidden elements, and metadata that a human never sees but the model does.
# extract_html.py
from bs4 import BeautifulSoup, Comment
def extract_hidden(html: str):
soup = BeautifulSoup(html, "html.parser")
hidden = []
for c in soup.find_all(string=lambda t: isinstance(t, Comment)):
hidden.append(("comment", c.strip()))
for el in soup.select('[style*="display:none"],[style*="visibility:hidden"],[hidden]'):
hidden.append(("css-hidden", el.get_text(strip=True)))
for img in soup.find_all("img"):
if img.get("alt"):
hidden.append(("alt-text", img["alt"]))
return [h for h in hidden if h[1]]Strip zero-width / Unicode-tag characters and decode common encodings so detectors see the real payload.
# normalize.py
import base64, codecs, re, unicodedata
ZERO_WIDTH = dict.fromkeys(map(ord, ""), None)
TAG_RANGE = range(0xE0000, 0xE0080) # Unicode tag chars used to smuggle text
def normalize(text: str) -> str:
text = text.translate(ZERO_WIDTH)
text = "".join(ch for ch in text if ord(ch) not in TAG_RANGE)
text = unicodedata.normalize("NFKC", text)
for token in re.findall(r"[A-Za-z0-9+/=]{20,}", text):
try:
decoded = base64.b64decode(token).decode("utf-8", "ignore")
if decoded.isprintable():
text += f"\n[decoded-b64] {decoded}"
except Exception:
pass
text += "\n[decoded-rot13] " + codecs.decode(text, "rot_13")
return textLLM Guard wraps a transformer classifier and returns a risk score per input.
# scan_llmguard.py
from llm_guard.input_scanners import PromptInjection
from llm_guard.input_scanners.prompt_injection import MatchType
scanner = PromptInjection(threshold=0.5, match_type=MatchType.FULL)
def scan(text: str):
sanitized, is_valid, risk = scanner.scan(text)
return {"is_valid": is_valid, "risk": risk} # is_valid=False => injection detectedRun Meta Prompt Guard 2 (or the open ProtectAI deberta classifier) for a second opinion.
# detector_model.py
from transformers import pipeline
# Open classifier (no gating); swap to meta-llama/Llama-Prompt-Guard-2-86M if licensed
clf = pipeline("text-classification",
model="protectai/deberta-v3-base-prompt-injection-v2")
def is_injection(text: str, threshold: float = 0.5) -> bool:
out = clf(text[:512])[0]
return out["label"].upper() == "INJECTION" and out["score"] >= thresholdMultimodal agents read text painted into pixels; OCR it and run the same scanners.
# scan_image.py
from PIL import Image
import pytesseract
def ocr(path: str) -> str:
return pytesseract.image_to_string(Image.open(path))
# Feed ocr(path) through normalize() + scan() + is_injection()Combine signals into block / sanitize / allow, and log a structured event for the SIEM.
# decide.py
import json, hashlib
from datetime import datetime, timezone
def decide(source, raw, normalized, llmguard_invalid, model_flag):
flagged = llmguard_invalid or model_flag
event = {
"ts": datetime.now(timezone.utc).isoformat(),
"source": source,
"sha256": hashlib.sha256(raw.encode("utf-8", "ignore")).hexdigest(),
"atlas": "AML.T0051.001",
"llmguard_injection": llmguard_invalid,
"model_injection": model_flag,
"decision": "block" if flagged else "allow",
}
print(json.dumps(event))
return event["decision"]Run the pipeline over a labeled set of clean + injected artifacts, measure precision/recall, and tune threshold to balance false positives against missed injections. Re-test whenever the agent's model or ingestion sources change.
| Tool | Purpose | Source |
|---|---|---|
| LLM Guard | Input/output scanners incl. PromptInjection | https://github.com/protectai/llm-guard |
| Meta Prompt Guard 2 | Dedicated jailbreak/injection classifier | https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-86M |
| ProtectAI deberta-v3 | Open prompt-injection classifier | https://huggingface.co/protectai/deberta-v3-base-prompt-injection-v2 |
| BeautifulSoup4 | HTML parsing / hidden-element extraction | https://www.crummy.com/software/BeautifulSoup/ |
| pytesseract / Tesseract | OCR text from images | https://github.com/madmaze/pytesseract |
| MITRE ATLAS | AI threat technique taxonomy | https://atlas.mitre.org/ |
| OWASP LLM01:2025 | Prompt Injection reference | https://genai.owasp.org/llmrisk/llm01-prompt-injection/ |
| Surface | Hiding technique | Extraction step |
|---|---|---|
| Web page | HTML comments, display:none, alt-text | BeautifulSoup hidden-element pass |
| white/tiny font, off-page text | pypdf text extraction + normalize | |
| Image | rendered pixels, EXIF, alt-text | OCR + metadata read |
| Any text | zero-width / Unicode-tag chars | normalize() de-obfuscation |
| Any text | Base64 / ROT13 encoding | decode pass in normalize() |
© mukul975, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. 5 hidden characters (zero-width or bidirectional) removed. Raw file
SKILL.md and 4 other files (scripts, references) in skills/detecting-indirect-prompt-injection of mukul975/Anthropic-Cybersecurity-Skills.
Open the folder on GitHubat commit 54a7988
Detecting Indirect Prompt Injection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Detecting Indirect Prompt Injection this skillmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~2.8k | Automated safety check: Warn | Apache-2.0 | |
| Prompt GuardOrchestra-Research/AI-Research-SKILLs | 13k | 1 repos | ~2.4k | Automated safety check: Warn | MIT | |
| AI Agent ActivitySCStelz/security-investigator | 249 | — | ~17k | Automated safety check: Pass | MIT | |
| Security GuidejnMetaCode/shellward | 140 | — | ~644 | Automated safety check: Warn | Apache-2.0 | |
| Prompt Injection Defensesickn33/agentic-awesome-skills | 47k | 2 repos | ~4.2k | Automated safety check: Warn | MIT | |
| AI Securityalirezarezvani/claude-skills | 28k | — | ~4.5k | Automated safety check: Warn | MIT |
Orchestra-Research/AI-Research-SKILLs
Meta's 86M prompt injection and jailbreak detector. An agent skill from Orchestra-Research/AI-Research-SKILLs.
SCStelz/security-investigator
Report/investigate RUNTIME ACTIVITY of AI agents (Agent 365 / Copilot Studio / M365 Copilot / Work IQ) — agents used, tools/connectors, channels, tokens, prompt/reply content, and Prompt Shield…
jnMetaCode/shellward
OpenClaw 安全部署指南 / Security deployment guide — help users secure their OpenClaw installation
sickn33/agentic-awesome-skills
Defend AI systems against prompt injection and indirect prompt attacks using input controls, tool permissions, output validation, and isolation boundaries.
alirezarezvani/claude-skills
A skill your agent uses when assessing AI/ML systems for prompt injection, jailbreak vulnerabilities, model inversion risk, data poisoning exposure, or agent tool abuse.
jnMetaCode/shellward
按中国法规(网安法 / PIPL / 等保2.0 / 数据出境 / AI生成内容标识)审计一个 AI 项目的代码仓库,产出每条都带 文件:行 取证、经独立复核、经脚本校验的合规报告。当用户问「这个项目上线合不合规」「调用了 OpenAI/Claude 算不算数据出境」「要不要做 AI 标识」「帮我做合规自查/等保/PIPL 检查」时使用。Audit an AI project's…
mukul975/Anthropic-Cybersecurity-Skills
Weighs infrastructure, TTP, malware code and timing evidence with the Diamond Model and competing hypotheses to reach a confidence-rated attribution.
mukul975/Anthropic-Cybersecurity-Skills
Walks through reverse engineering Go-compiled malware in Ghidra: parsing buildinfo and pclntab, recovering stripped function names and extracting dependencies.
mukul975/Anthropic-Cybersecurity-Skills
Guides forensic analysis of Windows LNK shortcut files and Jump Lists with LECmd, JLECmd and manual parsing to show file access and program execution.
mukul975/Anthropic-Cybersecurity-Skills
Hunts Windows malware persistence with Sysinternals Autoruns, covering run keys, services, scheduled tasks and drivers, with baseline comparison.
mukul975/Anthropic-Cybersecurity-Skills
Guides a Windows forensic examination of the NTFS Master File Table to recover deleted-file evidence, build timelines and spot timestomping.
mukul975/Anthropic-Cybersecurity-Skills
Detects DNS tunneling, ICMP exfiltration and HTTP-based covert channels in packet captures and DNS logs when hunting for hidden command-and-control traffic.
Works with
Categories
Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM…. Detecting Indirect Prompt Injection is an agent skill from mukul975/Anthropic-Cybersecurity-Skills. Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM Guard's PromptInjection scanner or Hugging Face Prompt Guard 2.
Detecting Indirect Prompt Injection fits situations like: an agent ingests untrusted external content and you need to screen it for injected instructions before the LLM processes it; tasks that involve Prompt injection and agent security; tasks that involve LLM guardrails.
Run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-indirect-prompt-injection -a claude-code`. Or copy the skill folder (skills/detecting-indirect-prompt-injection in mukul975/Anthropic-Cybersecurity-Skills) into .claude/skills/detecting-indirect-prompt-injection in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-indirect-prompt-injection -a codex`. Or copy the skill folder (skills/detecting-indirect-prompt-injection in mukul975/Anthropic-Cybersecurity-Skills) into .agents/skills/detecting-indirect-prompt-injection in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mukul975/Anthropic-Cybersecurity-Skills --skill detecting-indirect-prompt-injection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/detecting-indirect-prompt-injection, .gemini/skills/detecting-indirect-prompt-injection, .github/skills/detecting-indirect-prompt-injection and .opencode/skills/detecting-indirect-prompt-injection in your project.
Going by SKILL.md and its folder, Detecting Indirect Prompt Injection needs Python for the scripts in its folder and the command-line tools its instructions call (pip and python). Our summary lists: Python 3.
SKILL.md names 5 domains. As links in the text: github.com, huggingface.co, crummy.com, atlas.mitre.org and genai.owasp.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): contains zero-width characters. Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Detecting Indirect Prompt Injection is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 906 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Detecting Indirect Prompt Injection: Prompt Guard (Orchestra-Research/AI-Research-SKILLs, 13k stars), AI Agent Activity (SCStelz/security-investigator, 249 stars), Security Guide (jnMetaCode/shellward, 140 stars) and Prompt Injection Defense (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mukul975 (a GitHub user) maintains it in mukul975/Anthropic-Cybersecurity-Skills, which has 34,116 GitHub stars. The repository holds 644 skills in this directory. The repository was last updated on August 31, 2026.
Source: mukul975/Anthropic-Cybersecurity-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.