Triaging Security Alerts
Triage is the highest-volume work in security and the least written about. The
output is a disposition you can defend later — most often to someone asking
why an alert that preceded a breach was closed.
Three properties make it hard. The overwhelming majority of alerts are not
incidents, so the prior is against you on every single one. The cost of the two
error types is wildly asymmetric — a wrongly escalated alert wastes hours, a
wrongly closed one costs the breach. And the queue keeps arriving, so unbounded
care on one alert is care stolen from the next.
When to Use
- Working an alert queue from a SIEM, EDR, or cloud security tool
- Deciding whether an alert becomes an incident
- Reworking a backlog, or a specific alert that keeps recurring
- Reviewing another analyst's disposition
- Determining whether a noisy detection should be tuned or removed
When NOT to Use
- The alert is already confirmed malicious — stop triaging and use
responding-to-incidents; triage ends where response begins
- Searching for compromise with no alert to start from — use
hunting-threats; a hunt is hypothesis-driven, triage is queue-driven
- Rewriting the rule — use
engineering-detections or writing-sigma-rules;
feed your triage findings to it rather than tuning in place mid-queue
- Analysing the sample an alert pointed at — use
analyzing-malware
- Investigating a specific compromised host in depth — use
investigating-windows-endpoints or the relevant cloud investigation skill
Three Dispositions, Not Two
The common failure is a binary true/false frame. There are three, and
conflating the last two destroys your detection programme:
An admin legitimately dumping LSASS for a memory test is a benign true
positive: the rule worked perfectly. Filing it as a false positive leads
someone to weaken a rule that is functioning exactly as designed. Over a year
that is how a detection programme quietly dies.
Base Rates Decide More Than Evidence Does
Most triage errors are not evidence-reading errors. They are prior-probability
errors.
Take a detection that is 99% accurate, firing across 10,000 hosts where 1 is
actually compromised. It produces roughly 100 false alerts and 1 true one. A
positive alert is about 1% likely to be a real compromise — even at 99%
accuracy. This is why "the tool flagged it" carries almost no weight on its
own, and why an analyst who escalates on tool severity alone will be wrong
almost every time.
The practical consequences:
- Severity is a property of the rule, not of the alert. It was assigned by
whoever wrote the detection, before your environment existed.
- Ask what else would produce this signal. If routine administration,
backup software, or a vulnerability scanner explains it, that explanation is
far more likely than compromise before you have contrary evidence.
- Corroboration beats confidence. Two weak independent signals pointing the
same way move the posterior much further than one strong signal, because
their benign explanations rarely coincide.
- Rare things are rare — but rarity is not innocence. The point is to make
the prior explicit so evidence has to actually overcome it, not to explain
every alert away.
Order Enrichment by Cheapest Discriminator
Work the question that most cheaply splits benign from malicious. Do not run a
fixed enrichment checklist.
- What is this asset and who uses it? A domain controller and a
developer's laptop generate different priors for identical activity.
- Is this normal for this host or user? Frequency and history first. An
action that ran daily for eight months is a baseline, not an event.
- Was it authorized? Change tickets, maintenance windows, deployment
pipelines. Most benign true positives resolve here.
- What is the parent and the chain? Provenance discriminates far better
than the artifact itself.
powershell.exe is meaningless; spawned by
winword.exe is not.
- Only then, external reputation. Hash and IOC lookups are the last
cheap step, not the first. A clean reputation proves nothing about targeted
activity, and a dirty one still needs the local context above.
Stop as soon as one of these settles it. Running every step on every alert is
how the queue wins.
Time-Boxing and Escalation
Set a bound before you start — commonly 15 minutes for a routine alert. When
it expires, you must choose, and the choice is not "keep digging":
- Enough to close → close with the evidence recorded.
- Enough to escalate → escalate.
- Neither → escalate anyway. An alert that resists a full time-box is
itself a signal. Ambiguity is not a reason to keep it in your queue; it is a
reason to give it more resources than you have.
Escalate immediately, without finishing triage, on any of: confirmed execution
on a crown-jewel asset, credential access on a domain controller or identity
provider, evidence of lateral movement, security tooling being disabled, or
anything touching backup infrastructure. These are too expensive to be wrong
about slowly.