Agent skill

Authoring Security Skills

by trilwu in trilwu/secskills

Write a new SecSkills skill end to end — choosing the plugin bucket and skill tier, writing a description that triggers correctly without stealing traffic from siblings, the required sections…

MITAuto-check passedAI & LLM Engineering

Install Authoring Security Skills

skills CLI
$ npx skills add trilwu/secskills --skill authoring-security-skills -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install trilwu/secskills authoring-security-skills --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/trilwu/secskills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/authoring-security-skills .claude/skills/authoring-security-skills && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
authoring-security-skills
GitHub stars
157
Token cost
~3.1k tokens
SKILL.md length
1,543 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Write a new SecSkills skill end to end — choosing the plugin bucket and skill tier, writing a description that triggers correctly without stealing traffic from siblings, the required sections…

  • Works in 8 steps: Pick the Plugin → Pick the Tier → Write the Description → …
  • Correctly without stealing traffic from siblings
  • SKILL.md covers When to Use, When NOT to Use, Step 1: Pick the Plugin and Step 2: Pick the Tier, plus 8 more sections
  • Calls python3

What it does

Authoring Security Skills is an agent skill from trilwu/secskills. Write a new SecSkills skill end to end — choosing the plugin bucket and skill tier, writing a description that triggers correctly without stealing traffic from siblings, the required sections, registering the skill in ttp-index.json, adding routing eval cases including negative traps, and running the three validators. Use when adding a skill to this repo, splitting or merging existing skills, fixing a skill that triggers on the wrong requests, or when a new skill fails validate.py, syncattack.py, or runevals.py.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Transform Claude Code into your personal security engineer. The licence is MIT.

When your agent uses it

  • Correctly without stealing traffic from siblings
  • The required sections
  • Registering the skill in ttp-index.json
  • Adding routing eval cases including negative traps

Example prompts

  • “/authoring-security-skills”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Pick the Plugin
  2. Pick the Tier
  3. Write the Description
  4. Write the Body
  5. Register in the ATT&CK Index
  6. Add Eval Cases
  7. Validate
  8. Do Not Stamp It Verified

What it can do on your machine

Read from SKILL.md and the folder at commit ca53957. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Authoring Security Skills loads about 3.1k tokens when it runs. Until then it costs about 136 tokens; SKILL.md has 1,543 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~136
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from trilwu/secskills at commit ca53957, republished under its MIT licence (© trilwu). 1,543 words, ~3,138 tokens.

Download SKILL.mdSave it as .claude/skills/authoring-security-skills/SKILL.md (or your agent's skills folder).
name
authoring-security-skills
description
Write a new SecSkills skill end to end — choosing the plugin bucket and skill tier, writing a description that triggers correctly without stealing traffic from siblings, the required sections, registering the skill in ttp-index.json, adding routing eval cases including negative traps, and running the three validators. Use when adding a skill to this repo, splitting or merging existing skills, fixing a skill that triggers on the wrong requests, or when a new skill fails validate.py, sync_attack.py, or run_evals.py.

Authoring Security Skills

A skill earns its context budget by carrying what the model would otherwise get wrong. The model already knows what nmap -sV does. It does not reliably hold the discipline — trace to demonstrated impact before reporting, preserve before remediating, refuse to ship a rule never tested against production noise.

That judgment is the deliverable. Commands are scaffolding around it.

When to Use

  • Adding a new skill to secskills-offense, -defense, or -core
  • Splitting an overloaded skill or merging two that overlap
  • Fixing a skill that loads for the wrong requests, or fails to load for the right ones
  • Resolving a failure from validate.py, sync_attack.py, or run_evals.py

When NOT to Use

  • Fact-checking content you have written — that is verifying-skill-accuracy, and it is a separate gate that runs after this one
  • Editing prose in an existing, settled skill
  • Repo-level work (manifests, marketplace, CI) that is not about a skill

Step 1: Pick the Plugin

A skill lives in exactly one plugin. The split exists so a blue-team user is not carrying exploitation content and vice versa.

PluginTakesTest
secskills-offenseExploitation, attacker tradecraft, and testing that acts against a targetWould this run against someone's system during an engagement?
secskills-defenseDFIR, hunting, detection, intel, incident responseDoes this read artifacts after the fact, or build detection?
secskills-coreDual-use: reverse engineering, code/crypto/supply-chain review, AI security, ATT&CK mapping, reportingIs it equally at home on both sides?

When it fits two, choose core rather than duplicating. Cross-plugin references are expected and degrade to inert text for readers who installed only one plugin — that is by design, so never duplicate a skill to avoid a cross-reference.

Step 2: Pick the Tier

Domain skills carry methodology for a whole area and trigger on broad, plain-language requests. Keep them few and non-overlapping.

Procedure skills cover one target-and-toolchain combination that is rare, exact, and unrecoverable from general knowledge — reversing-flutter-apps with blutter, for instance. They are supposed to be narrow.

Two rules keep the tiers working:

  1. Every procedure skill is reachable from its domain skill. Add the identifying check and a routing line to the domain skill. An unreachable procedure skill is dead weight.
  2. A procedure skill hands back. Recovering symbols is not the assessment — name the skill that continues the work.

Unsure? Ask whether it applies to most engagements in its area (domain) or only when a specific artifact is in hand (procedure).

Step 3: Write the Description

This is the highest-leverage text in the skill. It is how Claude decides whether to load it, and it competes with all 71 siblings. A perfect skill body behind a vague description never runs.

Rules the validator enforces: third person, 60–1024 characters, and a trigger clause. Rules it cannot enforce but reviewers will:

  • Lead with what the skill does, then Use when ... naming concrete situations.
  • Build triggers from what the user actually has in hand — file names, magic bytes, framework and tool names, error strings, the identifying symptom. Not abstractions.
  • Name the evidence that distinguishes this skill from its nearest sibling.
yaml
# Good — names artifacts, tools, and the symptom
description: Reverse engineer and intercept traffic from Flutter/Dart mobile
  apps using blutter, reFlutter, and Frida. Use when an APK or IPA contains
  libflutter.so, libapp.so, App.framework, or flutter_assets, when jadx shows
  only a thin Dart wrapper, or when Burp sees no traffic from an app that is
  clearly online.

# Bad — overlaps the domain skill and triggers on the wrong requests
description: Advanced mobile reverse engineering techniques for modern apps.

Step 4: Write the Body

Required sections:

  • ## When to Use — concrete situations, not categories.
  • ## When NOT to Use — where a different skill is correct, with the sibling named in backticks. This is what stops skills fighting over the same words.

Expected in security skills:

  • ## Rationalizations to Reject — the plausible-sounding shortcuts that cause missed findings, each with why it is wrong. This is the highest-value section in the collection; it encodes judgment a command list cannot.

Offensive skills additionally need scope and authorization framing written to that domain's real exposure — not boilerplate. Recon has third-party estate and OSINT-as-personal-data; password cracking has account lockout as an availability risk; RF work has wiretap and spectrum law because you cannot confine a radio to the target. Generic "get permission first" text is worth nothing.

If the skill's workflow reads public reference material — advisories, specs, RFCs, vendor reports, ATT&CK pages — include the standard ## Reading External Sources block before ## References. Copy it verbatim from a skill that has it (producing-threat-intelligence, auditing-supply-chain); it must stay identical across skills so it does not drift. It routes public documentation through defuddle.md for a large token saving and full greppable text.

Do not add that block to a skill whose network activity is aimed at a target. Cloud metadata endpoints, the target's own services, exfil hosts, phishing URLs, and C2 must be fetched directly — routing them through a third-party extractor breaks the command, leaks engagement URLs, and for live adversary infrastructure tips off the operator.

Order steps by when they happen, and say so. A skill body that reads as an unordered pile of techniques gets sampled, not followed. Mark the first actions as immediate, the ones that depend on their output as following, and the ones that touch the target as acting — then a reader who loads the skill mid-task knows where they are. The ordering also carries the safety property: anything in the acting group is gated on authorization being granted, which is only legible if the groups are distinguishable.

A skill is a set of instructions to carry out, not a document to acknowledge. Loading one and replying that it has been read and understood is a failure mode worth naming in the body when the first step is easy to skip — a skill whose opening move is "record scope before touching anything" is exactly the skill a hurried reader will summarise instead of doing.

Write for judgment, not recall:

Write thisNot this
Why one technique is chosen over another, and when it failsA list of tool flags
The verification a finding must survive before being reported"Report the vulnerability"
Where the discipline goes wrong, named explicitlyGeneric best-practice reminders

Keep SKILL.md under 600 lines; move payload lists and long tables into a references/ directory, which loads only when pointed at. Do not chain them: SKILL.md may point at a reference file, but that file must not send the reader on to a third one.

Note that the validator resolves every backtick-quoted reference path as a real file — so an illustrative path in prose will fail the build. Name real files only.

Show full SKILL.md (521 more words)Show less

Step 5: Register in the ATT&CK Index

sync_attack.py fails if a skill is neither mapped nor declared unmapped. Edit secskills-core/ttp-index.json, never the generated block in a SKILL.md.

Map it — add the skill to the skills array of each technique it covers:

json
{"id": "T1595", "name": "Active Scanning", "tactics": ["TA0043"],
 "skills": ["performing-reconnaissance", "enumerating-network-services"]}

Or declare why Enterprise ATT&CK does not apply:

json
"_unmapped": {
  "securing-ai-systems": "MITRE ATLAS (AML.T####) and the OWASP Top 10 for LLM / Agentic Applications"
}

Then regenerate — the ## ATT&CK Coverage block is generated, so hand-editing it is always wrong:

bash
python3 scripts/sync_attack.py --write

Step 6: Add Eval Cases

Routing is a real failure mode; the eval set is how it is caught. Add cases to evals/cases.jsonl — one JSON object per line:

json
{"id": "adcs-esc1-certipy-vuln-template",
 "query": "certipy find -vulnerable flagged a template with ENROLLEE_SUPPLIES_SUBJECT and low-priv enrollment — how do I turn this into domain admin?",
 "expect_skill": "abusing-adcs",
 "also_acceptable": ["attacking-active-directory"],
 "expected_behavior": ["Identify this as an ESC1 escalation path",
                       "Request a cert specifying an alternate UPN",
                       "PKINIT with the cert to obtain the target's TGT/NT hash"],
 "trap_for": ["attacking-active-directory", "attacking-kerberos-delegation"]}

Write the query as a real user would type it — with the tool output, the error, the artifact in hand. Sanitised queries prove nothing.

trap_for is the important field. It lists the skills this query must not route to. Every new skill should appear in the trap_for of at least one neighbour's case, and carry at least one case of its own that traps against its nearest neighbours. That is the only mechanical pressure against description creep.

Step 7: Validate

All three must pass before a PR:

bash
python3 scripts/validate.py --strict     # form, manifests, verified count
python3 scripts/sync_attack.py --check   # ATT&CK blocks match the index
python3 scripts/run_evals.py --check     # every skill covered, no bad refs

Then bump the version in the plugin's .claude-plugin/plugin.json and .claude-plugin/marketplace.json — the validator enforces agreement — and add a CHANGELOG.md entry. Update the plugin description's skill count if the count changed; validate.py warns when it drifts.

Step 8: Do Not Stamp It Verified

New content is an unverified draft by definition. Never add a verified: date to a skill you just wrote. Hand it to verifying-skill-accuracy, drive every checkable claim to a primary source, and stamp only if you covered the whole surface.

Passing all three validators is not verification. They check form; the first verification pass over this repo found 31 factual errors in files with green CI.

Rationalizations to Reject

  • "The description is close enough; the body is what matters." A skill that never loads has no body. Description quality is the single largest determinant of whether the work gets used.
  • "I'll make the description broad so it definitely triggers." Broad descriptions steal traffic from precise siblings and get the wrong skill loaded for real work. Precision is what makes the collection compose.
  • "This overlaps an existing skill, but mine covers it better." Then fix the existing skill. Two skills competing for the same query degrades both.
  • "It's mostly commands — the discipline is obvious." If it were obvious it would not need writing down. Command catalogues are what the model already has; the judgment is what it lacks.
  • "I'll add the ATT&CK block by hand, it's faster." It is generated. Hand-edits are overwritten on the next sync and fail --check until then.
  • "Evals are a formality." Routing errors are the most common real failure in a large collection, and trap_for cases are the only thing that catches a description quietly widening.
  • "I tested the commands, so it's verified." Testing on your box confirms they run there, not that identifiers, defaults, and IDs are correct on the versions readers have. That is a separate pass.

References

  • verifying-skill-accuracy — the fact-checking gate that runs after this
  • CONTRIBUTING.md — the merge bar, house style, and PR expectations
  • secskills-core/ttp-index.json — the technique-to-skill map
  • evals/README.md — the eval harness and case format

© trilwu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/authoring-security-skills of trilwu/secskills.

Open the folder on GitHubat commit ca53957

Compare with similar skills

Authoring Security Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Authoring Security Skills compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Authoring Security Skills this skilltrilwu/secskills157—~3.1kAutomated safety check: PassMIT
AI Engineering Toolkitsickn33/agentic-awesome-skills47k2 repos~1.9kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT

Similar skills

  • AI Engineering Toolkit

    sickn33/agentic-awesome-skills

    6 production-ready AI engineering workflows: prompt evaluation (8-dimension scoring), context budget planning, RAG pipeline design, agent security audit (65-point checklist), eval harness building…

    47k GitHub starsUsed in 2 repos~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from trilwu/secskills

All 50 skills in this repo
  • Audit source code for exploitable vulnerabilities using threat-model-driven review, taint tracing, invariant checking, and variant analysis.

    157 GitHub stars~3.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Perform OSINT, subdomain enumeration, port scanning, web reconnaissance, email harvesting, and cloud asset discovery for initial access.

    157 GitHub stars~3.1k tokensUpdated 1 mo ago
    Auto-check: notes
  • Securing AI Systems

    trilwu/secskills

    Assess and harden LLM applications and agentic systems against prompt injection, tool misuse, excessive agency, memory poisoning, RAG data leakage, and model supply-chain risk, mapped to the OWASP…

    157 GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Analyzing Binaries

    trilwu/secskills

    Reverse engineer compiled binaries, firmware, and mobile app packages using triage, static disassembly, decompilation, and dynamic instrumentation.

    157 GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Analyzing Go Binaries

    trilwu/secskills

    Reverse engineer Go binaries by recovering function names and types from pclntab and moduledata using GoReSym, redress, and IDA/Ghidra Go plugins, and by reading Go's non-standard calling…

    157 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Analyzing iOS Binaries

    trilwu/secskills

    Analyze iOS applications at the binary level — decrypting FairPlay-protected IPAs with frida-ios-dump or bagbak, inspecting Mach-O load commands, recovering Objective-C headers with class-dump, and…

    157 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Authoring Security Skills

What does Authoring Security Skills do?

Write a new SecSkills skill end to end — choosing the plugin bucket and skill tier, writing a description that triggers correctly without stealing traffic from siblings, the required sections…. Authoring Security Skills is an agent skill from trilwu/secskills.json, adding routing eval cases including negative traps, and running the three validators.

When should I use Authoring Security Skills?

Authoring Security Skills fits situations like: correctly without stealing traffic from siblings; the required sections; registering the skill in ttp-index.json; adding routing eval cases including negative traps.

How do I install Authoring Security Skills in Claude Code?

Run `npx skills add trilwu/secskills --skill authoring-security-skills -a claude-code`. Or copy the skill folder (.claude/skills/authoring-security-skills in trilwu/secskills) into .claude/skills/authoring-security-skills in your project. Claude Code loads it when a task matches its description.

How do I install Authoring Security Skills in Codex?

Run `npx skills add trilwu/secskills --skill authoring-security-skills -a codex`. Or copy the skill folder (.claude/skills/authoring-security-skills in trilwu/secskills) into .agents/skills/authoring-security-skills in your project. Codex loads it when a task matches its description.

Can I use Authoring Security Skills in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add trilwu/secskills --skill authoring-security-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/authoring-security-skills, .gemini/skills/authoring-security-skills, .github/skills/authoring-security-skills and .opencode/skills/authoring-security-skills in your project.

What does Authoring Security Skills need to run?

Going by SKILL.md and its folder, Authoring Security Skills needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Authoring Security Skills access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Authoring Security Skills safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Authoring Security Skills use?

Authoring Security Skills is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Authoring Security Skills use?

About 3.1k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Authoring Security Skills?

Skills that share tags, products or a category with Authoring Security Skills: AI Engineering Toolkit (sickn33/agentic-awesome-skills, 47k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Authoring Security Skills?

trilwu (a GitHub user) maintains it in trilwu/secskills, which has 157 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on September 4, 2026.

Source: trilwu/secskills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.