Security Alert Triage
elastic/agent-skills
Triage Elastic Security alerts — gather context, classify threats, create cases, and acknowledge.
GATES method validation for hunt-derived detections. An agent skill from Nebulock-Inc/agentic-threat-hunting-framework.
$ npx skills add Nebulock-Inc/agentic-threat-hunting-framework --skill gates -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Nebulock-Inc/agentic-threat-hunting-framework gates --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Nebulock-Inc/agentic-threat-hunting-framework.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/gates .claude/skills/gates && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gates" agent skill from https://github.com/Nebulock-Inc/agentic-threat-hunting-framework/tree/main/.claude/skills/gates into .claude/skills/gates/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gates", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Nebulock-Inc/agentic-threat-hunting-framework/tree/main/.claude/skills/gatesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Nebulock-Inc/agentic-threat-hunting-framework --skill gates -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Nebulock-Inc/agentic-threat-hunting-framework gates --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Nebulock-Inc/agentic-threat-hunting-framework.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/gates .agents/skills/gates && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gates" agent skill from https://github.com/Nebulock-Inc/agentic-threat-hunting-framework/tree/main/.claude/skills/gates into .agents/skills/gates/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gates", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Nebulock-Inc/agentic-threat-hunting-framework --skill gates -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Nebulock-Inc/agentic-threat-hunting-framework gates --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Nebulock-Inc/agentic-threat-hunting-framework.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/gates .cursor/skills/gates && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gates" agent skill from https://github.com/Nebulock-Inc/agentic-threat-hunting-framework/tree/main/.claude/skills/gates into .cursor/skills/gates/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gates", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Nebulock-Inc/agentic-threat-hunting-framework.git --path .claude/skills/gates--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Nebulock-Inc/agentic-threat-hunting-framework --skill gates -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Nebulock-Inc/agentic-threat-hunting-framework gates --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Nebulock-Inc/agentic-threat-hunting-framework.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/gates .gemini/skills/gates && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gates" agent skill from https://github.com/Nebulock-Inc/agentic-threat-hunting-framework/tree/main/.claude/skills/gates into .gemini/skills/gates/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gates", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Nebulock-Inc/agentic-threat-hunting-framework gatesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Nebulock-Inc/agentic-threat-hunting-framework --skill gates -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Nebulock-Inc/agentic-threat-hunting-framework.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/gates .github/skills/gates && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gates" agent skill from https://github.com/Nebulock-Inc/agentic-threat-hunting-framework/tree/main/.claude/skills/gates into .github/skills/gates/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gates", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Nebulock-Inc/agentic-threat-hunting-framework --skill gates -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Nebulock-Inc/agentic-threat-hunting-framework gates --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Nebulock-Inc/agentic-threat-hunting-framework.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/gates .opencode/skills/gates && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gates" agent skill from https://github.com/Nebulock-Inc/agentic-threat-hunting-framework/tree/main/.claude/skills/gates into .opencode/skills/gates/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gates", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gatesGATES method validation for hunt-derived detections. An agent skill from Nebulock-Inc/agentic-threat-hunting-framework.
Gates is an agent skill from Nebulock-Inc/agentic-threat-hunting-framework. GATES method validation for hunt-derived detections. Evaluates hunt KEEP phase against 5 BASE + 5 ADVANCED criteria to determine if findings are production-ready.
Its SKILL.md is about 12k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files (for example `AGENT_MEMORY_SCHEMA.md`, `README.md` and `examples/README.md`).
It sits in Security, covering Security operations. The repository describes itself as: ATHF is a framework for agentic threat hunting - building systems that can remember, learn, and act with increasing autonomy. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0ffe4db. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown, bash, yaml and sql).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gates loads about 12k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 5,812 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Nebulock-Inc/agentic-threat-hunting-framework at commit 0ffe4db, republished under its MIT licence (© Nebulock-Inc). 5,812 words, ~11,822 tokens.
.claude/skills/gates/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.Apply GATES (Generalizable, Additive, Tunable, Exposure-tested, Sustainable) method to hunt KEEP phase outputs.
/gates --hunt H-XXXX
GATES adapts to different hunt outcomes:
| Hunt Type | Output | GATES Action | Example |
|---|---|---|---|
| Behavioral detection hunt | Explicit detection rules proposed | Score those rules | H-0903 (ClickFix), H-0904 (MEGAsync) |
| Exploration hunt | Queries + observations, no rules | Extract detection candidates from patterns | H-0902 (Zeek inspection) |
| Negative hunt | 0 TPs, behavioral pattern tested | Evaluate for proactive deployment or HOLD | H-0901 (browser extension installer) |
| Risk assessment hunt | Configuration findings, no behaviors | Document as non-GATES (advisory output) | H-0907 (Camera infrastructure) |
| Multi-step investigation | Requires analyst correlation | Classify as RECURRING_HUNT playbook | Asset correlation hunts |
| Baseline/inventory hunt | Environmental understanding | Document as knowledge capture (no detection) | Asset inventory, normal behavior mapping |
Key insight: Not every hunt produces detections. GATES identifies which hunts should → detections vs. advisories vs. playbooks.
GATES 5+5: 5 BASE (assert from hunt) + 5 ADVANCED (demonstrate operationally)
The acronym expands differently per tier, on purpose. Each letter names a concern;
each tier asks a different question about it, so A and S take a different word in
ADVANCED than in BASE. The table below is the canonical naming — no other expansion of
GATES is current. In particular "Accuracy", "Telemetry" and "Executability" are not
GATES gates; if you see them, they're stale.
| Gate | BASE — assert | ADVANCED — demonstrate |
|---|---|---|
| G | Generalizable — repeatable behavior, or a one-off? | Generalizable — cross-fleet + companion rules that generalize the technique? |
| A | Additive — does it fill a coverage gap? | Actionable — is there a validated or automated playbook? |
| T | Tunable — tell attack from normal? 7- & 30-day look-back, FP rate in range? | Tunable — a reusable allowlist other rules can share? |
| E | Exposure-tested — covered the bypasses, or just the obvious path? | Exposure-tested — run emulation; did real telemetry reveal a gap? |
| S | Sustainable — juice worth the squeeze: can we reliably see it, is the upkeep fair? | Soaked — live 48h–7d: performance with concurrent detections, under the volume threshold? |
An unqualified gate name in this skill is always the BASE word, because BASE is what GATES scores. ADVANCED is assessed after deployment and is never scored here.
"BASE score" means the BASE tier's score and nothing else: the sum of those five gates, 0.0–5.0. ADVANCED contributes zero to it — which is exactly why the score carries a tier name rather than being called "the GATES score".
Failure modes, listed in the order they are evaluated — Step 4 holds the binding rule; when two gates fail, the one higher in this list decides:
Rows 2 and 4 both mean carry the rule anyway, on a cycle. If the hunt found zero instances of the behavior, neither cycle is worth its maintenance → HOLD. Row 6 is the carve-out: with no gate failed, 0 TPs over a clean baseline is a proactive PROMOTE.
Before scoring, classify the hunt:
Decision tree:
Does KEEP propose explicit detections?
├─ YES → Score those detections (Step 2)
└─ NO → Are findings behavioral patterns?
├─ YES → Extract detection candidates (Step 1.5)
└─ NO → Are findings configuration/risk issues?
├─ YES → Non-GATES workflow (advisory output)
└─ NO → Multi-step investigation?
├─ YES → RECURRING_HUNT playbook
└─ NO → Baseline/inventory (knowledge capture)Read hunt frontmatter status field:
planning → Generate planning assessment (not full GATES)
H-XXXX_PLANNING_ASSESSMENT.mdcompleted → Proceed to full GATES validation
in-progress → Wait for completion
Planning Assessment Output:
For planning-phase hunts, generate H-XXXX_PLANNING_ASSESSMENT.md:
# H-XXXX Planning Assessment
**Status:** Planning (not executed)
## GATES Applicability
NOT READY FOR GATES VALIDATION
Hunt is in planning phase. GATES requires completed CHECK/KEEP phases with findings.
## Pre-Assessment (Detection Potential)
Based on hypothesis, the *potential* of each gate once the hunt has run. This is a
pre-flight sketch, **not a score** — a planning hunt has no evidence, so it gets no
BASE score at all. Write what each gate is blocked on:
- G (Generalizable): [plausible | doubtful] — why
- A (Additive): [plausible | doubtful] — why
- T (Tunable): blocked on — needs a baseline
- E (Exposure-tested): blocked on — needs bypass testing
- S (Sustainable): blocked on — needs volume data
Do **not** write PASS/PARTIAL/FAIL here, and do not sum anything. Those are scoring
tokens and using them pre-execution invites someone to total the row into a verdict.
## Recommendations for Execution
1. Execute hunt queries
2. Document findings in KEEP phase
3. Return to GATES for full validationBefore scoring BASE criteria, verify telemetry can distinguish attack from legitimate activity.
Ask these questions:
Can telemetry differentiate malicious from benign?
Are critical attack phases observable?
Can FP rate be tuned below sustainable threshold?
If NO to Question 3 — the events exist and are visible, but no field separates
attack from benign at a workable rate — that is a T FAIL, not a telemetry gap. Score it
and let Step 4 return CONDITIONAL: the data is there, the baseline isn't.
Answer per behavior, not per hunt. A hunt's hypothesis is usually broader than any
one detectable behavior in it, so Q1 and Q2 are frequently PARTIAL at the hypothesis
level and PASS for some narrower behavior inside it. Example: the remote-control session
of a KVM-over-IP device is served off-box and invisible (Q2 PARTIAL for "detect misuse
of the device"), while the device connecting — T1200, a USB enumeration event — is
fully observable. That hunt is scoreable.
So before classifying a hunt as a telemetry gap, ask whether any scoped behavior in it is fully observable:
T scope-drift failure below.If NO to Question 1 or 2 for every candidate behavior:
.md assessment documenting visibility limitationsExamples:
Note: Example references public ATHF showcase hunt. Your results will vary based on environment telemetry coverage.
Why check first: a telemetry gap is uncommon but not rare, and catching it here avoids scoring five gates on a detection that cannot be built at all.
# Find and load the hunt file. Search every hunt location, not just production —
# hunts move test/ → production/, and a planning hunt (which routes to the
# not-ready-for-GATES path) is usually still in test/.
find hunts -name "${HUNT_ID}.md" -type fParse YAML frontmatter: hunt_id, techniques, tactics, findings_count, true_positives, false_positives (if present)
Load hunt content from all phases:
## KEEP: Findings & Response):## CHECK: Execute & Analyze):### Hunting Queries → #### Initial Query, #### Refined Query (SQL, process
searches, network patterns). The refined query is usually the better detection
candidate: it is the initial one after FP tuning#### Query Performance (result counts, time window, volumes)## OBSERVE: Expected Behaviors):## LEARN: Prepare the Hunt):Flexible evidence extraction:
S on the automation
check and Step 4 resolves that to RECURRING_HUNTIf hunt explicitly proposes detections in KEEP → skip to Step 2 (score those)
If hunt does NOT explicitly propose detections → identify opportunities:
Extract detection candidates from hunt content by analyzing:
Behavioral patterns observed:
Query patterns that returned results:
"What Suspicious Looks Like" section:
Findings without TP/FP counts:
Multi-step investigations:
S FAILs the automation check, which Step 4 turns into
RECURRING_HUNT. Score it rather than dropping it, so the reason is on recordFor each candidate, document:
Detection Decomposition (Hybrid Rules):
If detection mixes behavioral + IOC components, split into independent layers:
Identify hybrid pattern:
/api/installer/script (IOC)Split into layers:
Score each independently:
Output both detections — one candidate_id each, sharing a stem and
suffixed with the layer (-behavioral / -ioc). Never 1a/1b: ordinals are
not stable across runs, see "Candidate identity" in AGENT_MEMORY_SCHEMA.md.
detections:
- candidate_id: powershell-download-cradle-behavioral
name: "PowerShell Download Cradle (Behavioral)"
gates_assessment:
verdict: CONDITIONAL
- candidate_id: powershell-download-cradle-ioc
name: "Suspicious URI Pattern (IOC)"
gates_assessment:
verdict: TIME_BOX
refresh_cycle: 90_daysWhy split? IOC shelf life (90 days) ≠ behavioral logic (durable). Separate verdicts enable proper lifecycle management.
Hybrid patterns seen in practice:
Why this matters: a hybrid that is scored as one candidate gets one verdict, and that verdict is wrong for one of its two halves — either the behavioral logic inherits the IOC's expiry, or the IOC inherits the behavioral rule's permanence.
External Reference Extraction:
If hunt KEEP phase references external detection artifacts (commit hash, PR, file):
selection, condition, filtersExample: Hunt references "detection repo commit abc1234" → Agent must read commit, extract detection rules, document logic (selection criteria, conditions, filters), then score each rule independently.
When the reference is unreadable and it is the deployable candidate. "Don't score
placeholder references" and "detection_logic.query is REQUIRED for a deployable
candidate" collide exactly here: the hunt cites rule files that don't exist yet (local
uncommitted work, a private repo, a dead link) and that citation is the only statement of
the logic. Resolve it this way:
detection_logic.query from that evidence, not from the reference. A query
you derived from a documented hunt query is a real query. Never emit a placeholder,
a TODO, or the reference string itself as the query — a deployable candidate whose
query can't be run is the fictional coverage this assessment exists to prevent.T and S at PARTIAL — you have the behavior but not
the baseline it was tuned against.Output: List of 1-N detection candidates OR determination that hunt produces no detection artifacts (risk assessment, baseline study, inventory)
Fast path — IOC with zero prevalence:
An IOC-based candidate (G = FAIL) in a hunt that found 0 instances resolves to
HOLD by the zero-prevalence rule in Step 4: a watchlist for a campaign that has never
touched this environment is clutter, not early warning. You can recognize that case up
front.
Score the other four gates anyway. The verdict is already known, but the scores are the
record of why, and a candidate with no base_score can't be compared against the next
one that looks like it. Output is a .md assessment, and activation is contingent on
threat intel showing the campaign in scope.
Contrast H-0906_EXAMPLE.yaml: also an IOC watchlist, also a G FAIL, but the campaign
is present — so it is TIME_BOX, not HOLD. Prevalence is the whole difference. Note
the extension too: TIME_BOX is deployable, so it emits .yaml; the HOLD case above
emits .md.
For each detection candidate (explicit or identified), score all 5 BASE criteria as
✅ PASS, ⚠️ PARTIAL, or ❌ FAIL. There is no fourth value — an unevidenced gate is
PARTIAL, never UNKNOWN, because only these three have point values.
A (Additive) User Verification: If detection repository is not accessible (no ATHF/ADEF integration, external repo, etc.), prompt user:
A - Additive Verification:
Do you have existing detection coverage for:
- Technique: [T1XXX from hunt]
- Behavior: [detection pattern]
- Data source: [EDR/CloudTrail/etc.]
Questions:
1. Does this detection duplicate any existing rule? (Y/N)
2. Does this fill a gap in your current detection portfolio? (Y/N)
3. If similar detection exists, what's different about this one?
→ User input required to score A criterionIf detection repository IS accessible (ATHF/ADEF integrated), automatically check for overlaps.
| Gate | BASE Question | Evidence from Hunt (Flexible) |
|---|---|---|
| G - Generalizable | Repeatable behavior, or a one-off? | ✅ Available: Techniques (T1XXX), pattern vs. IOC, behavior chains<br>❌ IOC pattern: Domain literals, specific IPs, campaign hashes → G-fail |
| A - Additive | Does it fill a coverage gap? | ✅ Check: Technique coverage in detection repo<br>→ PROMPT USER if detection repository not available |
| T - Tunable | Tell attack from normal? Can FPs be managed? | ✅ If available: Hunt TP/FP counts from Findings table<br>⚠️ If unavailable: Assess from query volumes, documented exclusions, "normal" behaviors<br>✅ Tunable: Filters/allowlists documented in hunt<br>❌ High risk: Widespread legitimate use, no clear filters<br>🎯 Clean baseline: 0 suspicious over a large sample is evidence the gate is satisfiable, so it lifts T from PARTIAL to PASS. It is not a separate bonus — there is no +0.5 to add on top of a gate result<br>⏱️ The lift is only as good as its window. A hunt-length sample (≤7 days) shows the pattern can be tuned, not that it is: event volume is not the same as time coverage, and a clean week says nothing about monthly batch jobs, patch cycles or quarterly automation. Under 30 days the clean baseline holds T at PARTIAL and the candidate deploys TEST for a soak — record the window you measured, not just the event count |
| E - Exposure-tested | Covered the bypasses, or just the obvious path? | ✅ Multiple angles: Count of queries/detection layers<br>✅ Bypass testing: Documented evasion scenarios<br>⚠️ Single path: Only one query pattern tested |
| S - Sustainable | Juice worth the squeeze: can we see it, is upkeep fair? | ✅ If available: Hunt alert volumes from queries<br>⚠️ If unavailable: Project from query result counts<br>✅ Low burden: Static allowlist, reliable telemetry<br>❌ High burden: Dynamic allowlist, requires correlation<br>📊 VOLUME THRESHOLDS:<br>- <10/day: Sustainable (manual triage feasible)<br>- 10-100/day: Conditional (requires aggregation/throttling)<br>- >100/day: Likely unsustainable unless TP rate >1%<br>🤖 AUTOMATION CHECK: Can SOC act without per-alert human context?<br>- ✅ "Is user authorized?" (allowlist lookup = automatable)<br>- ❌ "Does policy allow this tool?" (business judgment = not automatable)<br>- If per-alert context required: S-FAIL → RECURRING_HUNT |
Each gate scores:
BASE Score = Sum of gate scores (0.0 to 5.0). Record it as a float —
4.5, not "4.5/5" — because consumers average and threshold it.
Verdict Mapping: see Step 4: Verdict & Output, which is the single authoritative statement of how a score and the gate results become a verdict. Do not score a verdict from this section; the score alone is not sufficient to decide one.
Examples:
5.04.53.54.0 — a high score with a FAIL. The FAIL decides the verdict,
not the 4.0; see Step 4.When to score PARTIAL:
Handling Missing Data:
When hunt doesn't provide explicit TP/FP counts or detection logic:
| Missing Data | How to Score | Guidance |
|---|---|---|
| No TP/FP table | Use query volumes + findings descriptions | T-score: PARTIAL if volumes manageable but unvalidated<br>S-score: PARTIAL if volume unknown, project from query results |
| No detection logic proposed | Identify from queries/observations (Step 1.5) | Score the behavioral patterns you extract |
| No query volumes | Project from hunt scope (days × tenants) | S-score: PARTIAL — volume is projected, not measured. Recommend TEST soak to confirm. Never UNKNOWN: it has no point value, so it would silently drop the score |
| No exclusion filters | Assess from "normal" behavior descriptions | T-score: PARTIAL if tuning possible, FAIL if no clear filters |
| Hunt found 0 TPs | Valid outcome, doesn't fail GATES | G/A/E can still PASS; T/S may be PARTIAL (unvalidated). Step 4's zero-prevalence rule decides: with a G or S FAIL it is HOLD; with no FAIL and a clean baseline it is a proactive PROMOTE |
| Multi-step correlation | S = FAIL (automation check) | Score all five gates; Step 4 resolves the S FAIL to RECURRING_HUNT |
| Risk assessment hunt | No detection artifacts | Document as non-GATES workflow (like H-0907) |
Key principle: missing data means PARTIAL, not an automatic FAIL. It does not
imply a verdict — the verdict is still Step 4's. Note that PARTIALs do not accumulate
into CONDITIONAL: an all-PARTIAL candidate scores 2.5, which is HOLD. Document what
is unmeasured and name the validation step that would resolve it.
Rule type impact on the G gate — rule type determines G and nothing else; the
other four gates and the verdict still come from the evidence:
G PASSG PASSG FAIL (→ TIME_BOX, or HOLD at zero prevalence)S FAILs the automation check → RECURRING_HUNT (quarterly hunt playbook, not a standing detection)ADVANCED criteria are PENDING until deployment. Provide recommendations:
| Gate | ADVANCED Question | Recommendation |
|---|---|---|
| G - Generalizable | Cross-fleet + companion rules? | Test cross-OS/provider, build companion rules to generalize |
| A - Actionable | Validated or automated playbook? | SOC walkthrough, wire auto-enrichment |
| T - Tunable | Reusable allowlist? | Build shared allowlist other rules can reference |
| E - Exposure-tested | Run emulation — gap revealed? | Atomic Red Team, purple team validation |
| S - Soaked | Performance with concurrent detections? | 48h–7d TEST soak, volume < threshold |
These are the ADVANCED words (A = Actionable, S = Soaked), per the Framework
Reference. They are recommendations only — ADVANCED is never scored into the BASE score.
Exactly one verdict per candidate, from the canonical enum:
PROMOTE | CONDITIONAL | TIME_BOX | RECURRING_HUNT | HOLD | DROP.
A gate FAIL names a specific defect and therefore dictates a specific remedy. A score band is only a summary. So a FAIL always overrides the band — otherwise a candidate can score into PROMOTE while failing the gate that says it isn't deployable at all, and two readers get two different answers.
Evaluate in this order and stop at the first rule that matches:
| # | Condition | Verdict | Why this precedence |
|---|---|---|---|
| 1 | A = FAIL | DROP (or merge into the existing detection) | Coverage already exists. Nothing to build regardless of how well it scores elsewhere — the strongest veto. |
| 2 | G = FAIL | TIME_BOX | Campaign- or IOC-specific. Useful now, worthless later — build it with an expiry and a refresh cycle. |
| 3 | E = FAIL | HOLD + expand the hunt | Not enough angles were tested to judge the candidate. This is an epistemic block: any other verdict would be a guess. |
| 4 | S = FAIL and the behavior occurs | RECURRING_HUNT | Real signal, but it cannot run as a standing detection. Re-run it periodically instead. |
| 5 | T = FAIL | CONDITIONAL | Good logic, too noisy today. Deployable once the baseline/allowlist exists. |
| 6 | no FAIL | score band below | Nothing is broken; the score decides readiness. |
Zero prevalence downgrades a deploy-anyway verdict to HOLD. Rows 2 and 4 both say
carry this rule anyway, on a cycle — TIME_BOX with a refresh, RECURRING_HUNT with a
cadence. Neither is worth doing when the hunt found zero instances of the behavior: a
campaign watchlist for a campaign that has never touched the environment, and a quarterly
re-run of a query that matches nothing, are both maintenance with no expected yield. In
either case the verdict is HOLD — preserve the logic, record why, re-assess when threat
intel moves.
This modifies only rows 2 and 4. It does not apply to row 1 (DROP means don't even
preserve it — the coverage exists), row 3 (HOLD already), or row 5 (a CONDITIONAL
prerequisite is worth building regardless of current prevalence). Critically, it does
not apply to a candidate with no FAIL at all: a generalizable, well-baselined, quiet
detection that found 0 TPs is a proactive PROMOTE — a clean baseline over a large
sample and a long enough window is exactly what buys you the confidence to deploy
ahead of the threat. Under 30 days it is still PROMOTE (4.5, T PARTIAL), but as a
TEST soak rather than a settled rule.
Two or more FAILs: the earliest matching row wins. It is the one that has to be solved first, and solving it changes the assessment of the rest.
Why G outranks E and S. A G FAIL is a statement about the candidate's
kind, not a defect in it: it is an IOC or campaign rule. Rules of that kind are
single-dimensional and short-lived by construction, so they will almost always
trip E and S too. Those aren't three independent problems to solve — they are
one fact reported three times, and TIME_BOX is already the mitigation for all
three. Reading E first would send every IOC watchlist to HOLD and the skill
would never emit TIME_BOX at all.
A still outranks G: if the coverage already exists, don't build a time-boxed
copy of it either.
Contiguous and total — every score from 0.0 to 5.0 maps to exactly one verdict:
| BASE score | Verdict |
|---|---|
| 4.0 – 5.0 | ✅ PROMOTE — deploy as a standing detection |
| 3.0 – 3.9 | ⚠️ CONDITIONAL — deploy after prerequisites |
| 0.0 – 2.9 | ❌ HOLD — record the reasoning, build nothing |
With no FAILs the floor is 2.5 (five PARTIALs), so the 0.0–2.9 row is reachable only
by an all-PARTIAL candidate. Score bands do not produce TIME_BOX,
RECURRING_HUNT or DROP — those come from the precedence table above, and only
from it.
| Criteria | Score | Band says | FAIL rule says | Verdict |
|---|---|---|---|---|
| G,A,T,E,S all PASS | 5.0 | PROMOTE | — | PROMOTE |
| 4 PASS + E PARTIAL | 4.5 | PROMOTE | — | PROMOTE |
| G FAIL, rest PASS | 4.0 | PROMOTE | #2 TIME_BOX | TIME_BOX |
| T FAIL, S PARTIAL, rest PASS | 3.5 | CONDITIONAL | #5 CONDITIONAL | CONDITIONAL |
| S FAIL, T PARTIAL, rest PASS | 3.5 | CONDITIONAL | #4 RECURRING_HUNT | RECURRING_HUNT |
| S FAIL, A+T PARTIAL, 0 hits found | 3.0 | CONDITIONAL | #4, zero prevalence | HOLD |
| G FAIL, rest PASS, 0 hits found | 4.0 | PROMOTE | #2, zero prevalence | HOLD |
| T+S FAIL, rest PASS | 3.0 | CONDITIONAL | #4 RECURRING_HUNT | RECURRING_HUNT |
| G+E+S FAIL, A+T PASS | 2.0 | HOLD | #2 TIME_BOX | TIME_BOX |
| A FAIL, rest PASS | 4.0 | PROMOTE | #1 DROP | DROP |
Every row except the first two is a case where the band and the gates disagree. Before this rule they were genuinely undecidable, and a consumer building from the document would take a different action depending on which section it read. The last two rows are the multi-FAIL cases — note that neither is decided by the highest count of FAILs or by the score, only by which row matches first.
Generate output based on verdict:
This table routes on the hunt-level verdict, not a candidate's. The hunt-level
verdict picks the file; each candidate keeps its own verdict inside that file. So a
PROMOTE hunt's .yaml legitimately contains TIME_BOX, HOLD or DROP candidates
alongside the deployable ones — that is not a contradiction, and a consumer must handle
every verdict it finds inside a .yaml, not only the deployable ones. See the
hunt-level vs candidate-level table in AGENT_MEMORY_SCHEMA.md.
| Hunt-level verdict | File Generated | Purpose | Contains |
|---|---|---|---|
| PROMOTE / CONDITIONAL / TIME_BOX | H-XXXX_GATES.yaml | Deployment + learning | Structured data, narrative (embedded), deployment templates, agent-queryable metadata |
| HOLD / DROP / RECURRING_HUNT | H-XXXX_GATES.md | Analysis only | Narrative verdict, reasoning, recommendations (no deployment) |
| Risk Assessment / Non-GATES | H-XXXX_ASSESSMENT.md | Advisory | Client advisory, remediation guidance (not detection) |
Result: exactly one file per hunt — never one per detection — and the extension tells a consumer whether there is anything deployable inside.
See AGENT_MEMORY_SCHEMA.md for complete schema. Key sections:
gates_assessment - Scoring rationale, patterns learneddeployment - Engine, query, entities, operational params (ready to deploy)activation_date, expiration_date, review_date, refresh_cycle_days to operational_parametersdeployment_prerequisites — a list of strings — to operational_parameters. Required. CONDITIONAL means "cleared to deploy once these are met", so a CONDITIONAL candidate without them hands the consumer a blocked detection and no way to unblock it. Write them per candidate: prerequisites differ between candidates in the same hunt, and the hunt-level gates_validation.prerequisites is a roll-up, not a substitute.The deployment.engine field MUST use one of these 4 values only. ADEF CI-enforces this set:
sigma # Sigma rule format
sql # SQL queries (including KQL, Elasticsearch DSL)
sch_sql # Scheduled SQL queries (including Splunk SPL)
composite # Multiple detection types combinedQuery Format → Engine Mapping:
| Query Format | Engine Value | Example |
|---|---|---|
| Sigma YAML | sigma | selection:/condition: blocks |
| Splunk SPL | sch_sql | index=... | stats ... |
| KQL (Kusto) | sql | SecurityEvent | where ... |
| SQL | sql | SELECT * FROM events WHERE ... |
| Elasticsearch DSL | sql | Query DSL in JSON |
| Multiple types | composite | Combination of above |
WRONG: engine: splunk, engine: kql, engine: elastic
RIGHT: engine: sch_sql, engine: sql, engine: sql
If unsure: Use sch_sql for scheduled queries with aggregation, sql for simple filters.
Lightweight narrative report:
# GATES Validation: H-XXXX
**Verdict:** <one of ❌ HOLD, ⏱️ TIME_BOX, 🔄 RECURRING_HUNT, 🚫 DROP>
**BASE Score:** X.X
## Summary
[One paragraph: why this verdict]
## Reasoning
[BASE criteria scoring with evidence]
## Recommendations
[What to do instead: recurring hunt, IOC refresh, merge with existing rule]Present summary:
GATES Validation Complete: H-XXXX
Verdict: <hunt-level verdict> (<N> candidates)
File: hunt-promotion-analysis/H-XXXX_GATES.yaml
(or H-XXXX_GATES.md if HOLD, or H-XXXX_ASSESSMENT.md if non-GATES)
Next: [Deploy to detection repository | Build baseline | Convert to recurring hunt]Detailed examples: examples/README.md
Quick pattern recognition:
| Rule Pattern | Typical Score | Common Verdict | Example |
|---|---|---|---|
| Behavioral (process) | 4.5-5.0 | ✅ PROMOTE | Process injection, credential-store access |
| Known-offensive-tool signature | 4.5-5.0 | ✅ PROMOTE | Named post-exploitation or C2 frameworks |
| IOC (domain/IP) | 1.0-2.0 | TIME_BOX, or HOLD at zero prevalence | Campaign C2 domains, single-campaign IPs |
| High-volume API | 2.0-3.0 | CONDITIONAL/HOLD | A cloud control-plane call at thousands/day |
| Hybrid (behavioral+IOC) | scored per layer | split into two candidates | Connection pattern + domain IOC — score each half separately; neither SPLIT LAYERS nor HYBRID is a verdict |
| Baseline-required | 2.0-3.0 | CONDITIONAL | Per-entity anomaly detection — a missing baseline is a T FAIL (row 5), not an S one |
| Duplicate coverage | Varies | DROP | Overlapping rules on the same technique |
Scores are typical, not prescriptive — score the evidence you actually have, then apply the Step 4 precedence rule.
Hunt: Investigated cleartext credential transmission via FTP/HTTP
KEEP phase finding (prose only):
"Found 12 FTP authentication events to partner file servers (benign), 3 HTTP basic auth to internal admin panels (sanctioned), 0 adversary credential theft."
Detection extraction:
GATES scoring:
BASE Score: 3.5 (G 1.0 + A 1.0 + T 0.0 + E 0.5 + S 1.0)
Verdict: ⚠️ CONDITIONAL — the T FAIL pins it (precedence #5). Build the 3-destination
allowlist, then re-score T.
CHECK phase query:
-- Found 47 MEGAsync process executions
SELECT endpoint.uid, actor.user.name, process.file.path
WHERE process.name = 'MEGAsync.exe'
AND time >= now() - INTERVAL 30 DAYDetection extraction:
GATES scoring:
BASE Score: 5.0 (all five PASS)
Verdict: ✅ PROMOTE
Hunt: Camera/NVR infrastructure security assessment
Findings: Flat network segmentation, prohibited-vendor hardware, no external exposure
Detection extraction: None - findings are configuration states, not behavioral events
GATES classification: NOT APPLICABLE (risk assessment hunt, not detection hunt)
Outcome: Client advisories + remediation plan (not GATES validation)
KEEP phase finding:
"Query 1: List users with MFA changes (37 users). Query 2: For those users, check concurrent sessions (analyst manually correlates). Query 3: For suspicious pairs, investigate login history."
Detection extraction: Attempted, but requires analyst correlation between steps
Verdict: 🔄 RECURRING_HUNT — the correlation between steps is analyst judgment, so
S FAILs (precedence #4) and the behavior does occur, so it is not HOLD. Convert to a
quarterly hunt playbook. There is no "not a detection" verdict: a candidate that can't be
automated is still recorded, with the reason.
The verdict itself comes from Step 4. This section is only about what the write-up has to say so the reader can act on it. Write prose — there is no template to fill in.
| Gate | Verdict (Step 4) | State in the narrative |
|---|---|---|
| G | TIME_BOX | Where the IOCs came from and their refresh cadence; the review date (creation + 90 days); the behavioral companion rule to build so coverage outlives the campaign |
| A | DROP | Which existing rule already covers this, and whether anything here is meaningfully different (data source, FP rate, tuning). If the two are better as one rule, recommend merging into the existing one rather than building this |
| T | CONDITIONAL | The FP sources observed; whether the allowlist is static or dynamic, its size, and who maintains it. Name the promotion criterion — the measurement that would move this to PROMOTE |
| E | HOLD | The specific bypasses not covered, how trivial each is to perform, and the emulation needed to close them (Atomic Red Team test IDs where they exist) |
| S | RECURRING_HUNT, or HOLD at zero prevalence | Alert volume per day, the maintenance burden, and whether throttling is possible. For RECURRING_HUNT: the execution cadence, and why analyst judgment is required instead of a standing rule |
Three bright lines make these calls non-arbitrary:
Volume (S). <10/day is sustainable — manual triage is feasible. 10–100/day needs aggregation or throttling first. >100/day is unsustainable as a standing detection unless the TP rate exceeds 1%.
Allowlist feasibility (T). Under ~10 entries and static is manageable; over ~100,
or needing monthly updates, is not. A dynamic allowlist is an ongoing maintenance
commitment rather than a one-time build, so it costs S as well as T.
Automation (S). Can the SOC act without per-alert human context? "Is this user
authorized?" is an allowlist lookup and automatable. "Does policy permit this tool?"
is business judgment and is not. If every alert needs the latter, S FAILs.
The check reads volume-independent but isn't: what S actually measures is triage
load, which is per-alert context times alert volume. With a zero or near-zero
baseline there is nothing to triage, so the identical unautomatable question can be
S PASS. Two candidates in one hunt asking the same judgment question may therefore
score differently — PASS on the one that fires rarely, FAIL on the one that fires a
few times a day — and that is correct, not an inconsistency. State the volume you
scored against whenever the automation check decides the gate, so the reader can tell
a judgment call from an arithmetic one. A zero-baseline PASS is only as durable as its
window: if the baseline came from a hunt-length sample, deploy TEST and say that the
gate re-scores if volume appears.
Problem: a detection scoped wider than what the hunt actually tested carries unvalidated FP risk. The hunt's baseline only covers the tested scope; anything wider introduces FP sources nobody has looked at.
Compare the queries executed in CHECK against the detection logic proposed in KEEP. Drift looks like:
process.name = 'specific.exe' → proposed LIKE '%specific%' (untested matches)Score the drift: minor (10–20% wider) is PARTIAL; moderate (2–5×) or major (10×+, or
fundamentally different) is FAIL. Either way, say so explicitly — "detection scope
exceeds hunt test scope" — and state that a baseline at the proposed scope is required
before deployment, after which T can be re-scored.
examples/outputs/H-0903_EXAMPLE.yaml is a complete conformant output: four candidates,
two verdicts, per-gate criteria that sum to the stated base_score, and a hunt-level
verdict that is deliberately not an aggregate of the candidate ones.
Read that file rather than a walkthrough reproduced here. Two copies of the same artifact drift — that is how three of the examples in this skill ended up with scores their own criteria didn't add up to.
A .yaml output is the handoff artifact for ADEF, whose FORGE lifecycle picks up where
GATES stops:
ATHF LOCK (Hunt) → GATES (Validate) → ADEF FORGE (Engineer)
KEEP → BASE assessment → F - FINDadef hunt-promote --gates ~/athf-workspace/hunt-promotion-analysis/H-0903_GATES.yaml
# --dry-run first to see what it would mintEach deployable candidate lands at the Find stage with its own D-XXXX, catalog record
and journal, keyed by hunt_ref: "<hunt_id>#<candidate_id>" — which is why
candidate_id has to be stable across re-runs (see AGENT_MEMORY_SCHEMA.md). Archival
candidates are reported as skipped rather than minted.
A .md output has nothing for ADEF to import by design: the verdict says build nothing
yet. It stays in the ATHF workspace as the record of why.
© Nebulock-Inc, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files in .claude/skills/gates of Nebulock-Inc/agentic-threat-hunting-framework.
Open the folder on GitHubat commit 0ffe4db
Gates next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gates this skillNebulock-Inc/agentic-threat-hunting-framework | 388 | — | ~12k | Automated safety check: Pass | MIT | |
| Security Alert Triageelastic/agent-skills | 592 | 1 repos | ~3.5k | Automated safety check: Notes | Apache-2.0 | |
| Kubernetes Network Security Auditkubeshark/kubeshark | 12k | — | ~7.3k | Automated safety check: Notes | Apache-2.0 | |
| Security Detection Rule Managementelastic/agent-skills | 592 | 1 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Chaitin CLIchaitin/chaitin-cli | 115 | — | ~15k | Automated safety check: Notes | GPL-3.0 | |
| Elasticsearch Auditaspectrr/deer | 405 | — | ~1.7k | Automated safety check: Pass | MIT |
elastic/agent-skills
Triage Elastic Security alerts — gather context, classify threats, create cases, and acknowledge.
kubeshark/kubeshark
Hunts for compromised workloads and malicious traffic in a Kubernetes cluster by sweeping network data through Kubeshark MCP, mapped to MITRE ATT&CK.
elastic/agent-skills
Create, tune, and manage Elastic Security detection rules (SIEM and Endpoint).
chaitin/chaitin-cli
A skill your agent uses when running chaitin-cli commands to manage Chaitin security products: SafeLine WAF (site management, IP blocking, ACL, policy rules, attack logs), X-Ray vulnerability…
aspectrr/deer
Enable, configure, and query Elasticsearch security audit logs.
wiz-sec-public/SITF
Generate SITF-compliant attack flow JSON files from attack descriptions or incident reports.
Categories
GATES method validation for hunt-derived detections. An agent skill from Nebulock-Inc/agentic-threat-hunting-framework. Gates is an agent skill from Nebulock-Inc/agentic-threat-hunting-framework. GATES method validation for hunt-derived detections.
Gates fits situations like: tasks that involve Security operations.
Run `npx skills add Nebulock-Inc/agentic-threat-hunting-framework --skill gates -a claude-code`. Or copy the skill folder (.claude/skills/gates in Nebulock-Inc/agentic-threat-hunting-framework) into .claude/skills/gates in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Nebulock-Inc/agentic-threat-hunting-framework --skill gates -a codex`. Or copy the skill folder (.claude/skills/gates in Nebulock-Inc/agentic-threat-hunting-framework) into .agents/skills/gates in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Nebulock-Inc/agentic-threat-hunting-framework --skill gates -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gates, .gemini/skills/gates, .github/skills/gates and .opencode/skills/gates in your project.
SKILL.md names no scripts, command-line tools or credentials: Gates is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gates is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 12k tokens (SKILL.md is roughly 47k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gates: Security Alert Triage (elastic/agent-skills, 592 stars), Kubernetes Network Security Audit (kubeshark/kubeshark, 12k stars), Security Detection Rule Management (elastic/agent-skills, 592 stars) and Chaitin CLI (chaitin/chaitin-cli, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Nebulock-Inc (a GitHub organization) maintains it in Nebulock-Inc/agentic-threat-hunting-framework, which has 388 GitHub stars. The repository was last updated on October 8, 2026.
Source: Nebulock-Inc/agentic-threat-hunting-framework on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.