Agent skill

Vulnerability Triage

by provos in provos/ironcurtain

Reference vocabulary for interpreting vulnerability findings — detector-vs-impact distinction, severity anchoring on demonstrated evidence, the eleven-item interpretation rubric, delegation…

Apache-2.0Auto-check passedEducation

Install Vulnerability Triage

skills CLI
$ npx skills add provos/ironcurtain --skill vulnerability-triage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install provos/ironcurtain vulnerability-triage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/provos/ironcurtain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/workflow/workflows/vuln-discovery/skills/vulnerability-triage .claude/skills/vulnerability-triage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vulnerability-triage
GitHub stars
612
Token cost
~6.3k tokens
SKILL.md length
3,505 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reference vocabulary for interpreting vulnerability findings — detector-vs-impact distinction, severity anchoring on demonstrated evidence, the eleven-item interpretation rubric, delegation…

  • Works in 11 steps: Reproduction. Re-run the exact input.… → Effect realism. Classify the observation… → Attacker control surface. Is the… → …
  • Tasks that involve Bug bounty
  • SKILL.md covers Core rule, Detector evidence vs impact…, Interpretation rubric and Delegation transparency, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Vulnerability Triage is an agent skill from provos/ironcurtain. Reference vocabulary for interpreting vulnerability findings — detector-vs-impact distinction, severity anchoring on demonstrated evidence, the eleven-item interpretation rubric, delegation transparency, primitive-extent scaling, the disqualifier taxonomy (D-0..D-4), CVSS Achievable / Environmental framing, hedging-phrase elimination, and falsification asymmetry. Read when interpreting a finding to decide whether the demonstrated evidence supports the severity it would warrant. Surface-neutral; pulls in the…

Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Bug bounty and Quizzes and assessments. It works with C++ and Model Context Protocol. The repository describes itself as: A secure runtime for autonomous AI agents. Policy from plain-English constitutions. (https://ironcurtain.dev). The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bug bounty
  • Tasks that involve Quizzes and assessments

Example prompts

  • “/vulnerability-triage”

Workflow steps

11 steps, taken from the first numbered list in SKILL.md.

  1. Reproduction. Re-run the exact input. Confirm the same observation. Complete answer: same input, same observation, same artifact (crash…
  2. Effect realism. Classify the observation against the four categories below; the first two support an impact claim, the last two refute it.
  3. Attacker control surface. Is the effect's extent adversary-chosen, or bounded by structure the attacker cannot influence? Complete answer…
  4. Adjacency / payload model. What attacker-useful content plausibly reaches the affected location in production state? For memory-read…
  5. Channel reality. In deployed configurations, does the attacker actually observe the effect? Some channels realize impact (metadata APIs…
  6. Production hardening. What mitigations present in production (sanitizers, schema validators, CSP, ASLR, stack canaries, auth middleware…
  7. External exploitability trace. Does the input originate from an external / untrusted source, and through what delivery channel? Name it…
  8. Duplicate / CVE check. Search the project's issue tracker, CVE databases, and changelog for the code location. Note overlap with existing…
  9. Severity score. CVSS 3.1 anchored on demonstrated evidence under the Achievable / Environmental framing below. Justify each metric against…
  10. Delegation transparency. Skill section: Delegation transparency. If the finding depends on a delegate library, both linkage realism and…
  11. Primitive extent. Skill section: Primitive extent. The scaling axes for this finding's primitive class are enumerated. For each, discover…

What it can do on your machine

Read from SKILL.md and the folder at commit d51a346. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vulnerability Triage loads about 6.3k tokens when it runs. Until then it costs about 158 tokens; SKILL.md has 3,505 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~158
When it runs · the whole SKILL.md, loaded when a task matches
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from provos/ironcurtain at commit d51a346, republished under its Apache-2.0 licence (© provos). 3,505 words, ~6,335 tokens.

Download SKILL.mdSave it as .claude/skills/vulnerability-triage/SKILL.md (or your agent's skills folder).
name
vulnerability-triage
description
Reference vocabulary for interpreting vulnerability findings — detector-vs-impact distinction, severity anchoring on demonstrated evidence, the eleven-item interpretation rubric, delegation transparency, primitive-extent scaling, the disqualifier taxonomy (D-0..D-4), CVSS Achievable / Environmental framing, hedging-phrase elimination, and falsification asymmetry. Read when interpreting a finding to decide whether the demonstrated evidence supports the severity it would warrant. Surface-neutral; pulls in the relevant surface skill (e.g. `memory-safety-c-cpp`) for bug-class-specific exploitability factors.

Vulnerability Triage

Reference vocabulary for interpreting what an investigation actually demonstrated and translating it into threat-model language. Catalogs the rubric, disqualifiers, scoring conventions, and pitfalls a triage decision needs to reason about.

Core rule

Severity is anchored on what was actually demonstrated, never on the theoretical maximum implied by the vulnerable code. A detector firing on an isolated harness — sanitizer, fuzz crash, static analyzer hit, tainted-flow report, schema-validator alert, assertion failure — proves the detector caught something. It does NOT prove attacker-visible impact. Confidentiality, integrity, and availability claims require the corresponding effect to have been observed in production-equivalent conditions. When the demonstrated observation is weaker than the theoretical ceiling, score the demonstrated level and label the ceiling as such.

Detector evidence vs impact evidence

These two are routinely confused. Hold the line between them.

Detector evidence is anything a tool flagged: ASAN / UBSAN / TSAN / MSan crash, libFuzzer reproducer, static-analyzer warning, taint-flow trace, assertion site, schema-validator alert, AST pattern match. Detector evidence proves the detector caught an anomaly at the named site. It says nothing on its own about whether an attacker can reach that site, supply attacker-useful values, or observe the result.

Impact evidence is an observation of an attacker-visible CIA effect under production-equivalent conditions: reading bytes the attacker did not previously hold, mutating state the attacker did not previously control, denying service to a user other than the attacker, gaining a capability the attacker did not previously have. Impact evidence is anchored on observable artifacts — a crash, sanitizer trace, changed output, callback fired, file read, error message, measurable state change.

"Ran without error" is not impact evidence. If the expected effect was not observed, either the experiment was wrong or the bug is not triggered — diagnose which; do not paper over. A finding supported only by detector evidence is a detector hit awaiting demonstration, not a confirmed vulnerability.

Interpretation rubric

For each finding, the rubric below has eleven items. Skipping any one of them is a triage failure. Each item names a question, what counts as a complete answer, and what an incomplete answer looks like.

  1. Reproduction. Re-run the exact input. Confirm the same observation. Complete answer: same input, same observation, same artifact (crash signature, sanitizer trace, observed output). Incomplete: "should reproduce", "did not retry", or a different observation than the original. If reproduction fails, the finding is not yet triageable.

  2. Effect realism. Classify the observation against the four categories below; the first two support an impact claim, the last two refute it.

    • Attacker-uncontrolled target state observed (genuine impact).
    • Attacker-controlled effect on state the attacker did not previously hold (genuine impact).
    • Attacker's own input echoed or round-tripped back, or constant / neutral state (zeroed memory, schema-sanitized placeholder, default config, empty collection) — no impact.
    • Non-exploitable error path the attacker cannot steer — no impact.

    Reject impact claims when the observation falls in the last two categories.

  3. Attacker control surface. Is the effect's extent adversary-chosen, or bounded by structure the attacker cannot influence? Complete answer: explicitly state which input fields drive the effect's size, location, or content, and which are fixed by the surrounding code. Incomplete: "the attacker controls the input" without naming which subset of the input controls which dimension of the effect.

  4. Adjacency / payload model. What attacker-useful content plausibly reaches the affected location in production state? For memory-read findings, what target state sits adjacent to the affected region. For injection findings, what sink does the tainted value reach. For auth-bypass findings, what resources become reachable. State this as a reasoned model unless the experiment measured the specific payload.

  5. Channel reality. In deployed configurations, does the attacker actually observe the effect? Some channels realize impact (metadata APIs, attacker-rendered output, error messages echoed back). Others absorb or drop it (batch jobs, log-only sinks, sandboxed processes, observation channels not exposed to the caller). Complete answer: name the production channel through which the effect would be observed; trace it from the bug site to where the attacker sees it.

  6. Production hardening. What mitigations present in production (sanitizers, schema validators, CSP, ASLR, stack canaries, auth middleware, sandbox, rate limits, RELRO, FORTIFY_SOURCE) change the attacker's success rate? Did the experiment use the production profile or bypass it? A harness with hardening stripped tells you nothing about production outcomes — score what production does, not what the harness did.

    When a mitigation table asserts a runtime cap, flag, or mode is at a particular default value in production, cite the line that registers that default (GObject property spec, struct initializer, constant read at startup) — not a call site that happens to leave it unset. A harness that overrides the default via the library's public API (set_max_X(0), set_strict(false)) looks identical at the bug site to a real default of 0; the registration cite is what distinguishes them.

  7. External exploitability trace. Does the input originate from an external / untrusted source, and through what delivery channel? Name it concretely — network protocol, file upload, queue consumer, environment-variable propagation, IPC. Complete answer: a named delivery channel from an untrusted boundary to the bug site. Incomplete: "an attacker could supply input" without naming how.

  8. Duplicate / CVE check. Search the project's issue tracker, CVE databases, and changelog for the code location. Note overlap with existing fixes even when partial — a fix for one variant does not necessarily cover sibling variants.

  9. Severity score. CVSS 3.1 anchored on demonstrated evidence under the Achievable / Environmental framing below. Justify each metric against observations. If a metric is ambiguous, pick the lower value and note what work would support the higher.

  10. Delegation transparency. Skill section: Delegation transparency. If the finding depends on a delegate library, both linkage realism and delegate-side validation analysis are present. If either is missing, the score caps at Low regardless of other rubric items.

  11. Primitive extent. Skill section: Primitive extent. The scaling axes for this finding's primitive class are enumerated. For each, discover either demonstrated the axis or it is documented as un-explored. Un-explored axes that would change severity drive an insufficient verdict with directive — they do NOT raise the achievable score on extrapolation alone. For memory-safety findings, impact surface and heap layout reachability are explicitly analyzed.

Delegation transparency

When the violation site depends on values produced by a delegate library — any upstream parser, validator, decoder, or runtime not under the project's source control — severity cannot anchor beyond "latent code smell" until the delegate's own validation has been confirmed to allow the trigger values through. A harness that links against a project-internal stub of the delegate falsely demonstrates reachability: the trigger reaches the violation site because the stub's permissive behavior bypassed the real delegate's validation, not because the real delegate would.

Before scoring any severity above Low (CVSS ~3.x) on a finding that depends on a delegate library, two confirmations are required:

  1. Linkage realism. Cite evidence that the harness links against the real delegate library version production would link — not a project-internal stub. Acceptable evidence: dynamic linker output (ldd, otool -L), build manifest, package version pin. Project-internal stub directories (typically named *-stub/, *_mock/, mock-*/, fake-*/) resolved by the linker disqualify the evidence.
  2. Delegate-side validation analysis. Cite the specific delegate-side check that the trigger must pass through (function name, file, validation expression) and explain why it allows the trigger values through. If no such check exists in the delegate, state that explicitly with citation to the delegate's source.

Without both, the finding is a hardening opportunity at most. Score Low, recommend the upstream-side guard as defense-in-depth, and flag for harness redesign.

This rule sits above the rubric — failing it caps the achievable score regardless of the rubric's other items.

Primitive extent

The achievable CVSS score anchors on what discover actually demonstrated, including the scaling extent discover exercised — not on a larger scale the source suggests is reachable but no experiment ran. A primitive's theoretical ceiling implied by the surrounding code is not impact evidence; that is the same projection the Core rule forbids.

Triage's primitive-extent job is enumeration and gap-checking, not extrapolation. For each finding, list the scaling axes that apply to the primitive class. For each axis, classify whether discover demonstrated it (and at what value) or left it un-explored. An un-explored axis whose value would change severity if exercised is a triage gap, not a license to score on the projected value.

When an un-explored axis would change severity, the verdict is insufficient with a directive that names the axis, the boundary value to push toward, and what discover should run to characterize it. This is one of the concrete triggers for the insufficient-verdict discipline elsewhere in this skill — the achievable score does NOT rise on extrapolation alone. If every axis that matters has been demonstrated by discover (even at a non-maximum value), score on the demonstrated extent and note which axes were sampled where.

Common scaling axes per primitive class — for each, the entry names what discover should push to characterize the axis:

  • Out-of-bounds write/read. Discover should push distance past the allocation, total bytes accessed, attacker control over the written/read byte values, and stride between accesses. The minimum trigger may demonstrate one byte past one allocation with a fixed value; the axes ask whether that extends to attacker-chosen bytes across a larger span at a controllable stride.
  • Use-after-free. Discover should push reuse delay against any deferred-free queue or generation counter, and characterize the type set that can land in the freed slot under attacker-influenceable allocation pressure.
  • Type confusion. Discover should push the type-pair set the confusion lands on across the surrounding dispatch, and the depth of polymorphic dispatch reached in the confused state before the program faults or returns.
  • Integer overflow into undersized allocation. Discover should push the undersize ratio between requested and allocated size, and the extent of the downstream write or read that consumes the un-truncated value.
  • Unbounded iteration. Discover should push the iteration count against the buffer or frame the iteration writes into, and the attacker control over the value written at each step.
  • Resource exhaustion / DoS. For primitives whose impact is allocation/CPU/IO exhaustion rather than memory corruption, discover should push file-size-to-effect amplification (peak RSS, CPU time, or FD count divided by input file size on disk); wire-body-to-effect amplification when transport compression is in scope (same numerator divided by compressed POST body size — material for AV:N primitives, where gzip/br/zstd handed by upstream proxies can dominate cost-to-attacker); per-item-size × item-count multiplication when the bug multiplies via list/array containers (both ceilings are attacker-controlled — document independently); and the allocation-timing locus (eager parse before any application code runs vs. lazy on-demand — determines where in the call graph mitigations can be sited).

Score the achievable severity on the maximum extent discover demonstrated across these axes. If discover only ran the minimum trigger and an axis above plausibly changes severity, return insufficient with the axis named in the directive — do not silently raise the score, and do not silently cap it without flagging the gap.

When evaluating a heap OOB primitive, two additional analyses feed the impact axis. These reason over facts already in evidence (struct layouts, allocation order in the source, the crafted trigger's input shape) rather than projecting from one demonstrated trigger to a different one — they are static-analysis triage, not extrapolation:

  • Impact surface. Enumerate the adjacent allocation classes within the write stride: object vtables, library-internal struct fields, allocator metadata, and similar. Name what corruption produces, and which CVSS axis each corruption shape feeds (I, A, C). The enumeration must be concrete (named struct types, named fields) — "memory corruption could affect anything" is not enumeration.
  • Heap layout reachability. Is the heap state attacker-influenceable from the same input that triggers the bug? If yes, attack complexity drops a step (heap grooming is in-band rather than requiring a separate vector). If no, name the limiting factor (out-of-arena allocations, fresh process per request, deterministic allocator).

These two analyses are answerable from artifacts the workflow already produces (struct layouts, allocation traces, the crafted trigger input). They drive the CVSS Integrity and Attack Complexity scores; they are not optional for memory-safety findings.

Mitigation reachability. The abort site (function, file, line) where discover's demonstrated trigger actually faults or aborts is part of the primitive's characterization — record it alongside the scaling axes. Any mitigation triage recommends or accepts must sit at or upstream of that abort site in the call graph. A mitigation in code reachable only after the abort fires is unreachable in the demonstrated attack path and is not a valid recommendation, even if the static analysis of the surrounding code suggests it would catch the trigger. This check is most often violated when the abort fires inside a delegate library during eager parse and the proposed fix sits in application-side code that runs after parse completes.

Show full SKILL.md (1,398 more words)Show less

Disqualifier taxonomy

When the verdict is insufficient, the supporting notes carry one of the labels below. The labels disambiguate the gap and let downstream routing target it. The taxonomy is for justifying gaps in already-presented findings — not for seek-and-find dismissal of findings that meet the rubric.

  • D-0 — Evidence-synthesis failure. A prior experiment or analysis disproved the hypothesis the finding now rests on, but the finding was assembled without integrating that result. The contradiction is in the artifacts; the finding ignored it.

  • D-1 — Test, mock, example, or documentation code. The bug site lives in non-production code: unit test fixtures, example harnesses, sample inputs in documentation, mock implementations, debug-only utilities. Production callers do not reach it.

  • D-1.5 — Privilege tautology. The "attack" requires the attacker to already hold the capability the bug supposedly grants. A vulnerability must harm someone other than the attacker. Root reading a root-owned file, an admin reading admin-only data through an admin-only path, or a process reading its own memory through its own debug interface is not a vulnerability.

  • D-2 — Disqualifying preconditions. The bug requires conditions the attacker cannot supply: chaining with a separate unconfirmed vulnerability, a victim taking a specific action, physical access to the host, prior authentication where the bypass is not itself the finding, a non-default configuration the deployer would not reasonably enable.

  • D-3 — Hedged claim that did not survive verification. The finding's severity claim depends on language that hedges over the evidence rather than stating it (see "Hedging-language elimination" below). When the hedge is stripped, the claim collapses or contradicts the artifacts.

  • D-4 — No security impact. The bug class is real, but no CIA effect follows. Memory leak without amplification, resource exhaustion without disproportion, cosmetic state corruption that no consumer reads, an observation channel that absorbs the effect before any caller sees it, debug-only side effects.

CVSS Achievable / Environmental scoring

Triage emits one CVSS 3.1 vector and score per finding. That score is the Achievable / Environmental number — it reflects what was demonstrated under production-equivalent hardening, not the inherent worst case the bug class could imply on a hardening-stripped target.

Conventions for the metrics:

  • AV (Attack Vector). Decomposes how the attacker reaches the vulnerable code, not how they reach the host. A network-reachable service whose vulnerable function requires a local side-channel to trigger is AV:L for that finding, even though the service is on the network.

  • PR (Privileges Required). Decomposes the privileges needed to reach the vulnerable path, not the privileges needed to log into the system at all. If the vulnerable path is gated behind admin-only middleware, PR:H — even if the broader service has anonymous endpoints elsewhere.

  • UI / S / C / I / A. Anchored on what was observed under the production hardening profile. If production hardening absorbs the effect (a schema validator strips the payload; padding zeros the leak; a sandbox contains the side effect), the impact metric drops to None for the absorbed dimension and the absorption is noted.

  • Hardening-stripped harness contributes nothing here. A T1 harness compiled without canaries, ASLR, FORTIFY_SOURCE, or with custom permissive allocators may score very high in the abstract; that score is not the Achievable / Environmental score. If the production profile was not exercised, score what the production profile would do and call out the gap.

Hedging-language elimination

Hedged claims drift severity upward without evidence. Every hedge below must either be verified to a value-level statement or struck from the finding entirely. They are claims-to-verify, not claims-to-include.

  • "could potentially"
  • "may be exploitable"
  • "in theory"
  • "subject to specific conditions"
  • "if configured incorrectly"
  • "low likelihood"
  • "informational risk"
  • "cannot be ruled out"
  • "depending on environment"
  • "under certain circumstances"
  • "in some configurations"

A finding that depends on hedging language to assert its severity is a D-3 disqualifier candidate. Verified replacements name the specific value, configuration, or condition that turns the hedge into an observation.

Falsification asymmetry

One observed exploitation refutes "no bug." One failed exploitation does NOT refute "this bug." The burden is asymmetric: a single positive observation outweighs many failed attempts against the same evidence. Practical consequences:

  • If any approach succeeds in demonstrating the attacker-visible effect, the finding has impact evidence — even if other approaches against the same hypothesis failed.
  • If all plausible approaches failed against the same evidence, the finding is insufficient — not "ruled out." A different approach, a different harness tier, or a different production profile may still demonstrate impact.
  • Negative results constrain the demonstrated severity; they do not prove the bug-class is absent. State what was tested and what was not, and what range of severities the negatives rule out.

Common pitfalls

Patterns the rubric catches that get missed in practice. These are surface-neutral; per-surface restatements (e.g. heap-OOB-read adjacency rules, sanitizer-on-padding behavior) live in the relevant surface skill.

  • Detector evidence rebadged as impact evidence. A sanitizer fired on a Tier-1 isolated function. The finding writeup translates that into "remote code execution" without ever naming where the production code observes the effect. The detector says the detector caught something; that is not the same sentence.

  • Round-tripped attacker input mistaken for leaked target state. The bytes the bug "leaks" are the bytes the attacker just supplied, with possibly some encoding transformation. That is not a confidentiality finding. Any leak claim must demonstrate adjacency: what target state sits next to the affected region, and that the observation contained that target state.

  • Padding, zeroed memory, or sanitized placeholders mistaken for sensitive content. An OOB read returns NUL bytes, schema-sanitized fields, or default-config values. The detector fired; impact has not been demonstrated.

  • Hardening-stripped harness scored as production. The harness ran without ASLR, without canaries, without FORTIFY_SOURCE, with a permissive custom allocator, or with the validating middleware bypassed. The production binary intercepts the effect; the harness did not see that interception. The Achievable / Environmental score is what production does, not what the harness did.

  • Harness API override mistaken for production default. The harness called the library's public configuration API to lift a safety valve (set_max_X(0), disable strict mode, suppress validation) so a latent bug becomes reachable — a legitimate research move. The mitigations table then asserts that the production default is the lifted value, conflating "the harness set it to 0" with "the default is 0". Distinct from compiler-hardening stripping: this is the library's own runtime configuration, and the registration site (property spec, initializer) is the ground truth, not the call-site override.

  • Theoretical ceiling promoted to demonstrated severity by rewording. "This pattern could yield RCE" gets quoted, then rewritten without the hedge, then cited as "demonstrated RCE" two paragraphs later. The rubric's reproduction and effect-realism items refute this; the hedging-elimination rule names it.

  • Variant tunnel-vision. The harness reproduced one variant of the underlying pattern; the finding reports only that variant. When the same arithmetic / lifetime / dispatch shape reaches sibling variants on the same surface, those are findings in their own right. Document each, even when the experiment exercised only one.

  • Channel that absorbs the effect. The bug site exists, the value is attacker-controlled, but the production channel between the bug site and any observer is a logging path, a batch job that runs offline, a sandbox whose violations never escape. The effect happened; the attacker cannot observe it. Score impact at None for the dimension the channel absorbs.

  • Privilege tautology dressed as a finding. The "attacker" is the same principal as the "victim." A user reading their own data through their own session, an admin invoking an admin-only API, a process inspecting its own memory. D-1.5; not a vulnerability.

Insufficient-verdict discipline

When evidence does not support the severity the finding would warrant, the verdict is insufficient. The notes accompanying that verdict MUST name the specific gap concretely:

  • The parameter to sweep, with its named range and the boundary values not yet exercised.
  • The mitigation to exercise, with the production profile to rebuild under and the hardening flags to enable.
  • The adjacency to verify, with the target-state class that must be observed leaked rather than padding or attacker-input.
  • The hardening profile to rebuild under, with the specific compile flags or runtime configuration the prior experiment skipped.
  • The disqualifier label (D-0..D-4) that the finding currently triggers, when applicable.

Generic gap statements ("do more discovery", "needs more evidence", "investigate further") are forbidden. The downstream consumer of an insufficient verdict needs a directive concrete enough to act on without rederiving the gap.

A finding's insufficient verdict is not a closure — it is a redirection. The finding remains open; the gap statement is what closes it on the next round.

© provos, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/workflow/workflows/vuln-discovery/skills/vulnerability-triage of provos/ironcurtain.

Open the folder on GitHubat commit d51a346

Compare with similar skills

Vulnerability Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vulnerability Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vulnerability Triage this skillprovos/ironcurtain612—~6.3kAutomated safety check: PassApache-2.0
ChecksOtoDock/oto-dock190—~2.2kAutomated safety check: PassCustom licence
Nlm Skilliusztinpaul/ai-research-os-workshop1791 repos~6.9kAutomated safety check: PassMIT
Review Security ReportPrefectHQ/fastmcp28k—~1.2kAutomated safety check: PassApache-2.0
Osint Methodologyelementalsouls/Claude-OSINT2.8k—~8.7kAutomated safety check: NotesMIT
Openaso Asohubab1/OpenASO177—~946Automated safety check: PassMIT

Similar skills

  • Checks

    OtoDock/oto-dock

    How checks work on OtoDock — named units that judge an agent's work at the end of a turn (schema, script, handler and judge kinds), conditions on what the turn changed, rounds, where a check lives…

    190 GitHub stars~2.2k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Nlm Skill

    iusztinpaul/ai-research-os-workshop

    Expert guide for the NotebookLM CLI (nlm) and MCP server - interfaces for Google NotebookLM.

    179 GitHub starsUsed in 1 repo~6.9k tokens
    Knowledge ManagementAuto-check passed
  • Review Security Report

    PrefectHQ/fastmcp

    Review FastMCP vulnerability reports before accepting, rejecting, patching, scoring, or publishing them.

    28k GitHub stars~1.2k tokensUpdated today
    SecurityAuto-check passed
  • Osint Methodology

    elementalsouls/Claude-OSINT

    Comprehensive OSINT methodology for external red-team operations and authorized attack-surface assessments.

    2.8k GitHub stars~8.7k tokensUpdated 1 mo ago
    SecurityAuto-check: notes
  • Openaso Aso

    hubab1/OpenASO

    A skill your agent uses when performing App Store Optimization work with OpenASO MCP data: ASO audits, keyword research, metadata optimization, screenshot strategy, review analysis, competitor…

    177 GitHub stars~946 tokensUpdated 1 mo ago
    Marketing & SEOAuto-check passed
  • Project Settings Cascade

    samugit83/redamon

    Changing or adding a project setting / default value in RedAmon.

    3k GitHub stars~2.4k tokensUpdated yesterday
    SecurityAuto-check passed

More from provos/ironcurtain

  • Harness Design Fuzzing

    provos/ironcurtain

    Reference vocabulary for designing instrumented harnesses that drive vulnerability discovery — design classes (trigger-driven vs coverage-driven), tiered scope (T1 isolated function / T2…

    612 GitHub stars~5.7k tokensUpdated 2 days ago
    Auto-check passed
  • Email Formatting

    provos/ironcurtain

    Markdown formatting conventions for email summary documents — heading depth, list style, line length, emoji policy, and a mandatory provenance footer.

    612 GitHub stars~518 tokensUpdated 2 days ago
    Auto-check passed
  • Gmail Query Syntax

    provos/ironcurtain

    Reference for Gmail's search query syntax — operators like is:sent, newerthan:, from:, has:attachment, label:, and how they compose.

    612 GitHub stars~971 tokensUpdated 2 days ago
    Auto-check passed
  • JSON Schema Emails

    provos/ironcurtain

    Canonical shape of the .workflow/emails/emails.json file passed between the fetch and summarize states — required fields (sender, recipient, subject, date, body), types, and field semantics.

    612 GitHub stars~491 tokensUpdated 2 days ago
    Auto-check passed
  • Summary Style Guide

    provos/ironcurtain

    Tone and length conventions for email summaries — voice, verb tense, what to include vs omit, and target sentence count.

    612 GitHub stars~509 tokensUpdated 2 days ago
    Auto-check passed
  • Memory Safety C Cpp

    provos/ironcurtain

    Reference vocabulary for memory-safety vulnerabilities in native C/C++ code — bug-class taxonomy, common arithmetic patterns that lead to corruption, dispatch-family discipline, type-confusion…

    612 GitHub stars~5.6k tokensUpdated 2 days ago
    Auto-check passed

Questions about Vulnerability Triage

What does Vulnerability Triage do?

Reference vocabulary for interpreting vulnerability findings — detector-vs-impact distinction, severity anchoring on demonstrated evidence, the eleven-item interpretation rubric, delegation…. Vulnerability Triage is an agent skill from provos/ironcurtain.D-4), CVSS Achievable / Environmental framing, hedging-phrase elimination, and falsification asymmetry.

When should I use Vulnerability Triage?

Vulnerability Triage fits situations like: tasks that involve Bug bounty; tasks that involve Quizzes and assessments.

How do I install Vulnerability Triage in Claude Code?

Run `npx skills add provos/ironcurtain --skill vulnerability-triage -a claude-code`. Or copy the skill folder (src/workflow/workflows/vuln-discovery/skills/vulnerability-triage in provos/ironcurtain) into .claude/skills/vulnerability-triage in your project. Claude Code loads it when a task matches its description.

How do I install Vulnerability Triage in Codex?

Run `npx skills add provos/ironcurtain --skill vulnerability-triage -a codex`. Or copy the skill folder (src/workflow/workflows/vuln-discovery/skills/vulnerability-triage in provos/ironcurtain) into .agents/skills/vulnerability-triage in your project. Codex loads it when a task matches its description.

Can I use Vulnerability Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add provos/ironcurtain --skill vulnerability-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vulnerability-triage, .gemini/skills/vulnerability-triage, .github/skills/vulnerability-triage and .opencode/skills/vulnerability-triage in your project.

What does Vulnerability Triage need to run?

SKILL.md names no scripts, command-line tools or credentials: Vulnerability Triage is instructions for the agent only.

Does Vulnerability Triage access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vulnerability Triage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vulnerability Triage use?

Vulnerability Triage is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vulnerability Triage use?

About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vulnerability Triage?

Skills that share tags, products or a category with Vulnerability Triage: Checks (OtoDock/oto-dock, 190 stars), Nlm Skill (iusztinpaul/ai-research-os-workshop, 179 stars), Review Security Report (PrefectHQ/fastmcp, 28k stars) and Osint Methodology (elementalsouls/Claude-OSINT, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vulnerability Triage?

provos (a GitHub user) maintains it in provos/ironcurtain, which has 612 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 7, 2026.

Source: provos/ironcurtain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.