Agent skill

Moai Ref LLM Security

by modu-ai in modu-ai/moai-adk

AI/LLM defensive security reference: prompt-injection defense, OWASP LLM Top 10 defensive mapping, MCP and agentic tool-call hardening, training-data poisoning detection, model-output validation and…

Apache-2.0Auto-check passedSecurity

Install Moai Ref LLM Security

skills CLI
$ npx skills add modu-ai/moai-adk --skill moai-ref-llm-security -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install modu-ai/moai-adk moai-ref-llm-security --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/modu-ai/moai-adk.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/moai-ref-llm-security .claude/skills/moai-ref-llm-security && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
moai-ref-llm-security
GitHub stars
1.2k
Token cost
~4.5k tokens
SKILL.md length
2,074 words
Files
1
Skills in repo
48
Repo updated
First seen
Licence
Apache-2.0

At a glance

AI/LLM defensive security reference: prompt-injection defense, OWASP LLM Top 10 defensive mapping, MCP and agentic tool-call hardening, training-data poisoning detection, model-output validation and…

  • Tasks that involve Prompt injection and agent security
  • SKILL.md covers Target Use, Trust Boundaries in an LLM…, OWASP LLM Top 10 — Defensive… and Prompt-Injection Defense, plus 10 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Web application vulnerabilities

What it does

Moai Ref LLM Security is an agent skill from modu-ai/moai-adk. AI/LLM defensive security reference: prompt-injection defense, OWASP LLM Top 10 defensive mapping, MCP and agentic tool-call hardening, training-data poisoning detection, model-output validation and guardrails, MITRE ATLAS defensive correlation, and NIST AI RMF governance. Agent-extending skill that amplifies backend, security, and AI-application engineering with production-grade defensive patterns for LLM-backed systems. NOT for: offensive techniques (jailbreak authoring, attack-payload crafting, red-team…

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Security, covering Prompt injection and agent security, Web application vulnerabilities and LLM guardrails. It works with Model Context Protocol. The repository describes itself as: Agentic development harness for Claude Code — SPEC-driven plan/run/sync, TRUST 5 quality gates, model+effort routing, and Claude×GLM multi-LLM cost control. Single Go binary, 16… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Prompt injection and agent security
  • Tasks that involve Web application vulnerabilities
  • Tasks that involve LLM guardrails

Example prompts

  • “/moai-ref-llm-security”

What it can do on your machine

Read from SKILL.md and the folder at commit a2a184a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Moai Ref LLM Security loads about 4.5k tokens when it runs. Until then it costs about 183 tokens; SKILL.md has 2,074 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~183
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from modu-ai/moai-adk at commit a2a184a, republished under its Apache-2.0 licence (© modu-ai). 2,074 words, ~4,526 tokens.

Download SKILL.mdSave it as .claude/skills/moai-ref-llm-security/SKILL.md (or your agent's skills folder).
name
moai-ref-llm-security
description
AI/LLM defensive security reference: prompt-injection defense, OWASP LLM Top 10 defensive mapping, MCP and agentic tool-call hardening, training-data poisoning detection, model-output validation and guardrails, MITRE ATLAS defensive correlation, and NIST AI RMF governance. Agent-extending skill that amplifies backend, security, and AI-application engineering with production-grade defensive patterns for LLM-backed systems. NOT for: offensive techniques (jailbreak authoring, attack-payload crafting, red-team exploitation), model training or fine-tuning methodology, prompt optimization for capability, web-app OWASP Top 10 (see moai-ref-owasp-checklist), or general API design (see moai-ref-api-patterns).
when_to_use
Use when hardening an LLM-backed application or agent against prompt injection, designing guardrails or output validation, scoping MCP/tool permissions…
user-invocable
false
metadata.version
1.0.0
metadata.category
domain
metadata.status
active
metadata.updated
2026-06-24
metadata.tags
llm, ai-security, prompt-injection, guardrails, mcp, agentic, owasp-llm, mitre-atlas, nist-ai-rmf, reference
progressive_disclosure.enabled
true
progressive_disclosure.level1_tokens
100
progressive_disclosure.level2_tokens
3000

LLM / AI Defensive Security Reference

Defensive practitioner reference for hardening LLM-backed applications and agents. Every section is framed as defense, hardening, detection, or verification — it describes the misconfiguration, how to detect it, and how to prevent it, never how to exploit it. Cross-domain web-app vulnerabilities live in moai-ref-owasp-checklist; API design lives in moai-ref-api-patterns.

Target Use

Apply when reviewing or building an LLM-backed system — a chat product, a retrieval-augmented application, an autonomous agent, or an MCP server. The material assumes an untrusted-input threat model: any text the model reads (user turns, retrieved documents, tool results, file contents) may carry adversarial instructions, and any text the model emits may be acted on downstream.

Trust Boundaries in an LLM System

The core defensive insight: an LLM does not distinguish "data" from "instructions" the way a parser does. Treat every text channel that reaches the model as a boundary where adversarial instructions can enter.

ChannelEntry riskPrimary defense
End-user promptDirect prompt injectionInstruction-hierarchy enforcement, input screening
Retrieved documents (RAG)Indirect prompt injectionProvenance tagging, content isolation, retrieval allowlist
Tool / function resultsInjected instructions in tool outputTreat tool output as untrusted data, re-validate before re-prompting
System / developer promptLeakage, overrideMinimize secrets in prompt, assume prompt is recoverable
Model outputImproper downstream handlingSchema validation, encoding, never auto-execute raw output
Training / fine-tuning dataPoisoningData lineage, provenance verification, canary detection

OWASP LLM Top 10 — Defensive Mapping

The OWASP Top 10 for LLM Applications (2025 edition) is the canonical risk index for LLM systems. Each row below states the risk, the defensive check, and the hardening control. Citations are for defensive correlation only.

IDRiskDefensive checkHardening control
LLM01Prompt InjectionCan external text override system instructions?Instruction-hierarchy enforcement, input/output screening, content isolation
LLM02Sensitive Information DisclosureCan the model reveal secrets, PII, or system prompt?Output filtering, minimize secrets in context, response redaction
LLM03Supply ChainAre model weights, datasets, and plugins from verified sources?Provenance verification, model/dataset signing, dependency pinning
LLM04Data and Model PoisoningCould training/fine-tuning data carry adversarial content?Data lineage, canary artifacts, anomaly detection on training sets
LLM05Improper Output HandlingIs model output executed/rendered without validation?Schema validation, context-aware encoding, no auto-execution
LLM06Excessive AgencyDoes the agent hold more capability/permission than needed?Least-privilege tool design, human-in-loop for high-impact actions
LLM07System Prompt LeakageDoes the system prompt hold secrets that leak on extraction?Never store secrets in the prompt; assume the prompt is recoverable
LLM08Vector and Embedding WeaknessesCan embedding/RAG stores be poisoned or leak cross-tenant?Per-tenant isolation, embedding-source validation, access control
LLM09MisinformationIs unverified model output presented as authoritative?Grounding, citation requirements, confidence signalling
LLM10Unbounded ConsumptionCan a request exhaust tokens, cost, or compute?Rate limiting, token budgets, request-size caps, timeout enforcement

Prompt-Injection Defense

Prompt injection is the highest-frequency LLM risk (LLM01). Defense is defense-in-depth — no single control is sufficient, so layer the controls below.

Direct injection (the user is the attacker)
LayerControlPurpose
Instruction hierarchyMark system instructions as highest-priority; instruct the model that downstream text cannot override themReduce override success rate
Input screeningScan incoming prompts for known injection markers before they reach the modelDetect obvious override attempts
Privilege separationRun a high-trust planning model separately from a low-trust task model that touches untrusted contentLimit blast radius of a successful injection
Output gatingValidate the model's response against an allowlist of expected shapes before acting on itBlock injected instructions from taking effect downstream
Indirect injection (the attacker poisons retrieved content)

Indirect injection arrives through documents, web pages, emails, or tool results that the model reads. The defense is to treat all retrieved content as untrusted data, never as instructions.

  • Provenance tagging: wrap retrieved content in clearly delimited, labeled blocks so the model is instructed to treat the enclosed text as data to analyze, not commands to obey.
  • Content isolation: keep retrieved content in a separate channel from the instruction channel; do not concatenate untrusted text directly after a system instruction.
  • Retrieval allowlist: restrict RAG sources to verified corpora; reject or quarantine documents from unverified origins.
  • Re-validation after tool calls: when a tool returns text that re-enters the prompt, re-screen it — a tool result is an untrusted channel, not a trusted one.

This class of attack maps to MITRE ATLAS AML.T0051 (LLM Prompt Injection); the defenses above are how you detect and prevent it, not an attack procedure.

MCP / Agentic Tool-Call Hardening

Agentic systems that call tools (including MCP servers) expand the attack surface: a successful injection can now drive real actions. The defense is least-privilege tool design plus output re-validation (LLM06 Excessive Agency).

Tool-permission scoping
ControlDefensive rationale
Least-privilege tool setExpose only the tools the task needs; an agent that cannot delete cannot be tricked into deleting
Scoped credentials per toolA tool holds the minimum credential for its function, never a broad admin token
Human-in-the-loop gatesHigh-impact actions (payments, deletions, external posts, force-push) require explicit confirmation regardless of the model's confidence
Allowlist of callable targetsConstrain which hosts, paths, or resources a tool may reach; block internal/metadata endpoints
Idempotency + dry-runPrefer reversible or preview-able actions; confirm before irreversible ones
Tool-output validation

Treat every tool result as untrusted input on its way back into the model:

  • Schema-validate tool results before re-prompting — reject malformed or oversized payloads.
  • Strip or neutralize instruction-shaped content in tool output (a tool that returns a web page may carry an injection in that page).
  • Bound tool-call loops — cap the number of tool invocations per task to prevent an injected instruction from driving an unbounded action chain (LLM10 Unbounded Consumption).
  • Log every tool call with its arguments and result for audit and anomaly detection.

This hardening defends against MITRE ATLAS AML.T0053 (LLM Plugin Compromise) class concerns by minimizing what a compromised tool path can reach.

Training-Data and Model Poisoning Detection

Poisoning (LLM04) corrupts model behavior by inserting adversarial samples into training or fine-tuning data, or by substituting a tampered model artifact. The defenses are provenance and detection, not retraining recipes.

DefenseWhat it detectsHow to apply
Data lineageUntracked or unverified data entering the training setRecord the source, hash, and acquisition path of every dataset
Canary artifactsWhether a known marker sample influenced the model unexpectedlyInsert known canary records; verify the model's behavior on them post-training
Provenance verificationA substituted or tampered model/datasetVerify cryptographic signatures on model weights and datasets before use
Distribution monitoringAnomalous shifts in training-data statisticsCompare new data batches against the established distribution baseline
Source allowlistingData from unverified originsRestrict training corpora to vetted, signed sources

Poisoning maps to MITRE ATLAS AML.T0020 (Poison Training Data); the detection controls above are defensive correlation, never an attack method.

Model-Output Validation and Guardrails

Improper output handling (LLM05) is when downstream systems trust raw model output. The defense is to treat model output as untrusted until validated.

Output validation layers
LayerControlPrevents
Structured outputEnforce a schema (typed object, constrained grammar) on the responseFree-form output carrying injected commands
Content filteringScreen output for policy violations, leaked secrets, or disallowed contentSensitive disclosure (LLM02), system-prompt leakage (LLM07)
Context-aware encodingEncode output for its sink (HTML escape, shell-quote, SQL parameter)Output-handling injection into a downstream interpreter
No auto-executionNever feed raw model output directly into a shell, eval, or queryRemote code execution via the model
Grounding + citationRequire the model to cite sources for factual claimsMisinformation (LLM09) presented as authoritative
Show full SKILL.md (846 more words)Show less
Guardrail placement

Guardrails belong on both sides of the model — an input guardrail screens what enters, an output guardrail screens what leaves. An output-only guardrail misses injection that has already changed the model's plan; an input-only guardrail misses sensitive disclosure in the response. Layer both.

MITRE ATLAS — Defensive Correlation

MITRE ATLAS catalogs adversarial techniques against AI systems. The skills cite ATLAS technique IDs to correlate a defense with the technique it counters — never to provide an attack recipe. Representative defensive correlations:

ATLAS techniqueTechnique nameDefensive control that counters it
AML.T0051LLM Prompt InjectionInstruction-hierarchy enforcement, content isolation, input/output screening
AML.T0020Poison Training DataData lineage, canary artifacts, provenance verification
AML.T0054LLM JailbreakOutput filtering, refusal-policy enforcement, response gating
AML.T0057LLM Data LeakageOutput redaction, minimize context secrets, per-tenant isolation

For each technique, the skill states how to detect and defend, not how to execute. Treating an ATLAS technique ID as an attack recipe is the anti-pattern this skill explicitly forbids.

NIST AI RMF — Governance Mapping

The NIST AI Risk Management Framework (AI RMF 1.0) organizes AI risk into four functions. Map LLM-security controls onto them for defensive governance:

FunctionDefensive activity
GOVERNEstablish AI-security policy, accountability, and incident-response ownership for LLM systems
MAPInventory LLM components, data sources, tool permissions, and trust boundaries; identify the threat model
MEASURETest for prompt-injection resistance, output-validation coverage, and guardrail effectiveness; track findings
MANAGEPrioritize and remediate identified risks; monitor production behavior; maintain the response plan

The four functions give a defensive governance scaffold; the OWASP LLM Top 10 and MITRE ATLAS supply the specific risks to map, measure, and manage.

Cross-References

  • moai-ref-owasp-checklist — web-application OWASP Top 10, authentication patterns, input validation, HTTP security headers (the dev-time application security surface; this skill is the LLM-specific surface).
  • moai-ref-api-patterns — REST/GraphQL API design, error handling, rate limiting (the API-design surface that an LLM-backed service exposes).
  • moai-ref-supply-chain — dependency and artifact provenance, signing, SBOMs (the broader supply-chain surface that LLM03/LLM04 model and data provenance draws on).

Defensive Severity Levels

LevelLabelActionExample
P0CRITICALBlock releaseRaw model output executed in a shell; no output validation on an agent with destructive tools
P1HIGHFix before mergeNo instruction-hierarchy enforcement on a system handling untrusted documents
P2MEDIUMFix within iterationTool-call loop has no upper bound; missing input screening
P3LOWTrack in backlogSystem prompt slightly over-discloses internal structure (no secret leaked)
<!-- moai:evolvable-start id="rationalizations" -->

Common Rationalizations

RationalizationReality
"The model is smart enough to ignore injected instructions"Models do not reliably separate data from instructions. A capable model can still be steered by injected text; defense-in-depth, not model capability, is the control.
"We only read internal documents, so RAG content is trusted"Internal documents can be edited by a compromised account or an upstream feed. Indirect injection does not require an external attacker — treat all retrieved content as untrusted.
"Output validation slows the response, users will not notice the risk"Unvalidated output is how model responses become shell commands and XSS. Validation is the boundary between a chat reply and remote code execution.
"The agent needs broad permissions to be useful"Excessive agency is an OWASP LLM Top 10 risk. Scope tools to the task; a capability the agent does not hold cannot be abused via injection.
"We put guardrails on the output, that is enough"An output-only guardrail misses injection that already changed the model's plan and tool calls. Guardrails belong on both input and output.
"Prompt injection is a research problem, not a production one"Indirect prompt injection is exploited in production RAG and agentic systems today. It is a current operational risk, not a theoretical one.

Untrusted-by-default: every text channel reaching the model — user input, retrieved documents, tool results — is untrusted until screened. The model is not a trust boundary; your validation layer is.

<!-- moai:evolvable-end -->
<!-- moai:evolvable-start id="red-flags" -->

Red Flags

  • Model output fed directly into a shell, eval, query, or innerHTML without schema validation or encoding
  • Retrieved RAG content concatenated directly after a system instruction with no isolation or provenance tagging
  • An agent holds destructive or high-privilege tools with no human-in-the-loop gate on high-impact actions
  • Tool results re-enter the prompt without being re-screened as untrusted input
  • Secrets, API keys, or credentials embedded in the system prompt (assume the prompt is recoverable)
  • No upper bound on tool-call loops or token consumption per request
  • Guardrails present on only one side (input-only or output-only) of the model
<!-- moai:evolvable-end -->
<!-- moai:evolvable-start id="verification" -->

Verification

  • Every text channel reaching the model is classified as trusted or untrusted, and untrusted channels are screened
  • System instructions use instruction-hierarchy enforcement against override by downstream text
  • Retrieved/RAG content is isolated and provenance-tagged, not concatenated after instructions
  • Model output is schema-validated and context-encoded before any downstream use; nothing auto-executes raw output
  • Agent tools follow least-privilege scoping; high-impact actions have a human-in-the-loop gate
  • Tool results are re-validated as untrusted input before re-prompting; tool-call loops are bounded
  • The design is mapped against the OWASP LLM Top 10 (show which items were evaluated) and governed under the NIST AI RMF functions
  • No secrets reside in the system prompt; output filtering screens for sensitive disclosure
<!-- moai:evolvable-end -->

© modu-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/moai-ref-llm-security of modu-ai/moai-adk.

Open the folder on GitHubat commit a2a184a

Compare with similar skills

Moai Ref LLM Security next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Moai Ref LLM Security compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Moai Ref LLM Security this skillmodu-ai/moai-adk1.2k—~4.5kAutomated safety check: PassApache-2.0
Hunt LLM AIelementalsouls/Claude-BugHunter4.8k—~4kAutomated safety check: WarnMIT
Defending Applicationstelagod/code-abyss243—~777Automated safety check: PassMIT
Securitytelagod/code-abyss243—~907Automated safety check: PassMIT
China AI Compliance AuditjnMetaCode/shellward140—~1.1kAutomated safety check: PassApache-2.0
Security GuidejnMetaCode/shellward140—~644Automated safety check: WarnApache-2.0

Similar skills

  • Hunt LLM AI

    elementalsouls/Claude-BugHunter

    Hunt LLM/AI feature bugs — prompt injection, indirect injection, exfiltration via tool-use/markdown, ASCII smuggling, agentic AI security (OWASP Agentic Apps 2026, ASI01-ASI10).

    4.8k GitHub stars~4k tokensUpdated yesterday
    SecurityAuto-check: warnings
  • Defending Applications

    telagod/code-abyss

    Application security defense knowledge for builders. An agent skill from telagod/code-abyss.

    243 GitHub stars~777 tokensUpdated 2 mo ago
    SecurityAuto-check passed
  • Security

    telagod/code-abyss

    Defensive security engineering judgment, distilled from a stronger model - invoke when THREAT MODELING a system or feature; making security-relevant design decisions (auth, crypto, trust boundaries…

    243 GitHub stars~907 tokensUpdated 2 mo ago
    SecurityAuto-check passed
  • China AI Compliance Audit

    jnMetaCode/shellward

    按中国法规(网安法 / PIPL / 等保2.0 / 数据出境 / AI生成内容标识)审计一个 AI 项目的代码仓库,产出每条都带 文件:行 取证、经独立复核、经脚本校验的合规报告。当用户问「这个项目上线合不合规」「调用了 OpenAI/Claude 算不算数据出境」「要不要做 AI 标识」「帮我做合规自查/等保/PIPL 检查」时使用。Audit an AI project's…

    140 GitHub stars~1.1k tokensUpdated 9 days ago
    SecurityAuto-check passed
  • Security Guide

    jnMetaCode/shellward

    OpenClaw 安全部署指南 / Security deployment guide — help users secure their OpenClaw installation

    140 GitHub stars~644 tokensUpdated 9 days ago
    SecurityAuto-check: warnings
  • Reins Runtime Security

    pegasi-ai/reins

    Installs hooks that check each agent action against security policies before it runs, blocking destructive commands and logging every decision.

    392 GitHub stars~1.4k tokensUpdated 4 mo ago
    SecurityAuto-check: warnings

More from modu-ai/moai-adk

All 48 skills in this repo
  • Builds hand-editable SVG diagrams from computed layout coordinates, lints the source and renders a 2x PNG, with rules for when mermaid is the better choice.

    1.2k GitHub stars~5.2k tokensUpdated today
    Auto-check: notes
  • MoAI Foundation Core

    modu-ai/moai-adk

    Reference for MoAI-ADK's core development principles: TRUST 5 quality gates, SPEC-first domain-driven workflow, agent delegation and token budgeting.

    1.2k GitHub stars~5k tokensUpdated today
    Auto-check passed
  • MoAI SPEC Workflow

    modu-ai/moai-adk

    Manages SPEC documents for MoAI-ADK development, with GEARS or EARS requirement notation, acceptance criteria and a link into the Plan-Run-Sync workflow.

    1.2k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • MoAI TDD Workflow

    modu-ai/moai-adk

    Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.

    1.2k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • MoAI Worktree Management

    modu-ai/moai-adk

    Gives each SPEC its own Git worktree with a registry of active workspaces, base-branch sync and cleanup of merged ones, inside the MoAI-ADK workflow.

    1.2k GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Watches a pull request's CI checks after creation, separates required from auxiliary failures, applies limited safe fixes and escalates anything semantic to you.

    1.2k GitHub stars~2.4k tokensUpdated today
    Auto-check: notes

Questions about Moai Ref LLM Security

What does Moai Ref LLM Security do?

AI/LLM defensive security reference: prompt-injection defense, OWASP LLM Top 10 defensive mapping, MCP and agentic tool-call hardening, training-data poisoning detection, model-output validation and…. Moai Ref LLM Security is an agent skill from modu-ai/moai-adk. AI/LLM defensive security reference: prompt-injection defense, OWASP LLM Top 10 defensive mapping, MCP and agentic tool-call hardening, training-data poisoning detection, model-output validation and guardrails, MITRE ATLAS defensive correlation, and NIST AI RMF governance.

When should I use Moai Ref LLM Security?

Moai Ref LLM Security fits situations like: tasks that involve Prompt injection and agent security; tasks that involve Web application vulnerabilities; tasks that involve LLM guardrails.

How do I install Moai Ref LLM Security in Claude Code?

Run `npx skills add modu-ai/moai-adk --skill moai-ref-llm-security -a claude-code`. Or copy the skill folder (.claude/skills/moai-ref-llm-security in modu-ai/moai-adk) into .claude/skills/moai-ref-llm-security in your project. Claude Code loads it when a task matches its description.

How do I install Moai Ref LLM Security in Codex?

Run `npx skills add modu-ai/moai-adk --skill moai-ref-llm-security -a codex`. Or copy the skill folder (.claude/skills/moai-ref-llm-security in modu-ai/moai-adk) into .agents/skills/moai-ref-llm-security in your project. Codex loads it when a task matches its description.

Can I use Moai Ref LLM Security in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add modu-ai/moai-adk --skill moai-ref-llm-security -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/moai-ref-llm-security, .gemini/skills/moai-ref-llm-security, .github/skills/moai-ref-llm-security and .opencode/skills/moai-ref-llm-security in your project.

What does Moai Ref LLM Security need to run?

SKILL.md names no scripts, command-line tools or credentials: Moai Ref LLM Security is instructions for the agent only.

Does Moai Ref LLM Security access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Moai Ref LLM Security safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Moai Ref LLM Security use?

Moai Ref LLM Security is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Moai Ref LLM Security use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Moai Ref LLM Security?

Skills that share tags, products or a category with Moai Ref LLM Security: Hunt LLM AI (elementalsouls/Claude-BugHunter, 4.8k stars), Defending Applications (telagod/code-abyss, 243 stars), Security (telagod/code-abyss, 243 stars) and China AI Compliance Audit (jnMetaCode/shellward, 140 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Moai Ref LLM Security?

modu-ai (a GitHub organization) maintains it in modu-ai/moai-adk, which has 1,230 GitHub stars. The repository holds 48 skills in this directory. The repository was last updated on October 8, 2026.

Source: modu-ai/moai-adk on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.