Agent skill

Sentinel Ingestion Report

by SCStelz in SCStelz/security-investigator

Sentinel Ingestion Report — YAML-driven PowerShell pipeline gathers all data via az monitor/az rest/Graph API, writes a deterministic scratchpad, LLM renders the report.

MITAuto-check passedData & Analytics

Install Sentinel Ingestion Report

skills CLI
$ npx skills add SCStelz/security-investigator --skill sentinel-ingestion-report -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SCStelz/security-investigator sentinel-ingestion-report --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/SCStelz/security-investigator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/sentinel-ingestion-report .claude/skills/sentinel-ingestion-report && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sentinel-ingestion-report
GitHub stars
249
Token cost
~16k tokens
SKILL.md length
7,151 words
Files
59
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Sentinel Ingestion Report — YAML-driven PowerShell pipeline gathers all data via az monitor/az rest/Graph API, writes a deterministic scratchpad, LLM renders the report.

  • Works in 5 steps: Run Data Gathering → Load Rendering Context → Render Report (Single Write) → …
  • Tasks that involve Anomaly detection
  • SKILL.md covers Purpose, Architecture, Companion Files — When to Load and 📑 TABLE OF CONTENTS, plus 6 more sections
  • Runs PowerShell scripts from its folder; calls az and python

What it does

Sentinel Ingestion Report is an agent skill from SCStelz/security-investigator. Sentinel Ingestion Report — YAML-driven PowerShell pipeline gathers all data via az monitor/az rest/Graph API, writes a deterministic scratchpad, LLM renders the report. Covers table-level volume breakdown, tier classification (Analytics/Basic/Data Lake), SecurityEvent/Syslog/CommonSecurityLog deep dives, ingestion anomaly detection (24h and WoW), analytic rule inventory via REST API, rule health via SentinelHealth, detection coverage cross-reference, tier migration candidates with DL-eligibility lookup, license…

Its SKILL.md is about 16k tokens, which your agent loads only when the skill is triggered. The skill folder holds 62 other files (for example `SKILL-drilldown.md`, `SKILL-report.md` and `queries/phase1/Q1-UsageByDataType.yaml`).

It sits in Data & Analytics, covering Anomaly detection and REST APIs. It works with Microsoft 365 and PowerShell. The repository describes itself as: Automated security investigation tool using Microsoft MCP Servers, GitHub Copilot, Python Modules and custom copilot-instructions. The licence is MIT.

When your agent uses it

  • Tasks that involve Anomaly detection
  • Tasks that involve REST APIs

Example prompts

  • “/sentinel-ingestion-report”

Requirements

  • Python 3
  • PowerShell

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Run Data Gathering
  2. Load Rendering Context
  3. Render Report (Single Write)
  4. Initialization
  5. Render Output (LLM)

What it can do on your machine

Read from SKILL.md and the folder at commit 51e1385. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (PowerShell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • az
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • learn.microsoft.com
    • azure.microsoft.com
    • aka.ms
    • techcommunity.microsoft.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sentinel Ingestion Report loads about 16k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 7,151 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~161
When it runs · the whole SKILL.md, loaded when a task matches
~16k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from SCStelz/security-investigator at commit 51e1385, republished under its MIT licence (© SCStelz). 7,151 words, ~16,277 tokens.

Download SKILL.mdSave it as .claude/skills/sentinel-ingestion-report/SKILL.md (or your agent's skills folder). This skill also uses 58 other files; get the full folder from GitHub.
name
sentinel-ingestion-report
description
Sentinel Ingestion Report — YAML-driven PowerShell pipeline gathers all data via az monitor/az rest/Graph API, writes a deterministic scratchpad, LLM renders the report. Covers table-level volume breakdown, tier classification (Analytics/Basic/Data Lake), SecurityEvent/Syslog/CommonSecurityLog deep dives, ingestion anomaly detection (24h and WoW), analytic rule inventory via REST API, rule health via SentinelHealth, detection coverage cross-reference, tier migration candidates with DL-eligibility lookup, license benefit analysis (DfS P2 500MB/server/day, M365 E5 data grant). Inline chat and markdown file output.
threat_pulse_domains
cloud, admin
drill_down_prompt
Run a Sentinel ingestion report — data volume, table tiers, ingestion anomalies, analytic rule health, and cost optimization

Sentinel Ingestion Analysis Report — Instructions

Purpose

This skill generates a comprehensive Sentinel Ingestion Analysis Report covering workspace data volume, table-level breakdown, tier classification, ingestion anomalies, detection coverage, and optimization opportunities.

Entity Type: Sentinel workspace (from config.json)

ScopePrimary TablesUse Case
Workspace-wide (default)Usage, SentinelHealth, SentinelAuditFull ingestion and cost analysis
Per-table deep diveSecurityEvent, Syslog, CommonSecurityLog + any tableGranular breakdown of high-volume tables

What this report covers: Table-level volume breakdown with tier classification (Analytics/Basic/Data Lake), SecurityEvent/Syslog/CommonSecurityLog deep dives, ingestion anomaly detection (24h and week-over-week), analytic rule inventory with detection coverage cross-reference, rule health monitoring, tier migration candidates with DL-eligibility assessment, and license benefit analysis (DfS P2 and M365 E5).


Architecture

 ┌─────────────────────────────────────────────────────────────────┐
 │  YAML query files        PowerShell script        LLM render    │
 │  queries/phase1-5/  ──→  Invoke-IngestionScan  ──→  Phase 6     │
 │  (23 .yaml files)        .ps1 (~2600 lines)       (SKILL-       │
 │                          • az monitor (KQL)        report.md)   │
 │                          • az rest (REST API)                   │
 │                          • az monitor table list                │
 │                          • Invoke-MgGraphRequest                │
 │                          ↓                                      │
 │                     temp/ingest_scratch_<ts>.md                 │
 │                     (~50 KB, 64 sections)                       │
 └─────────────────────────────────────────────────────────────────┘

Execution model:

  • Phases 1-5 (data gathering): Fully automated by Invoke-IngestionScan.ps1. KQL queries run via az monitor log-analytics query. Non-KQL data (analytic rules, tier classifications, custom detections) is gathered via REST API, Azure CLI, and Microsoft Graph.
  • Phase 6 (rendering): LLM reads the scratchpad + SKILL-report.md and renders the report. This is the only phase requiring LLM involvement.

Design decision — TopRecommendations: The Top 3 Recommendations are computed by the LLM at render time (Phase 6), not pre-computed by PS1. Three of the seven Rule E categories (Data loss, DCR filter, Split ingestion) require cross-section reasoning that spans multiple scratchpad sections — this is precisely what the LLM excels at. The PS1 provides all the raw data; the LLM applies Rule E scoring across it.


Companion Files — When to Load

This skill spans 4 files. Load only the file(s) needed for the current phase:

FilePurposeWhen to Load
SKILL.md (this file)Architecture, workflow, rendering rules, domain referenceAlways — primary entry point
SKILL-report.mdReport templates (§1-§8), section-to-scratchpad mapping, formatting rulesPhase 6 rendering only
SKILL-drilldown.mdPost-report drill-down — rule cross-referencing (AR + CD via Graph API), ASIM parser verification, known pitfalls, error handlingAfter report is generated, when user asks follow-up questions (see §13 summary)
Invoke-IngestionScan.ps1PowerShell data-gathering pipeline (Phases 1-5)Execution only — no need to read unless debugging
slice_scratch.pyRead-only verbatim block slicer — extracts ## PRERENDERED tables/skeleton byte-for-byte so they aren't mangled by hand-copyPhase 6 rendering (optional but recommended)
render_dashboard.pyDeterministic SVG dashboard renderer — parses scratchpad + report + svg-widgets.yaml into the 7-row dashboard (no hardcoded run data)SVG Dashboard Generation (default — run this when asked to visualize)

📑 TABLE OF CONTENTS

  1. Quick Start - 3-step execution pattern
  2. Critical Workflow Rules - Prerequisites and prohibitions
  3. Execution Workflow - Phases 0-6
  4. Query File Reference - All 23 YAML files
  5. Output Modes - Inline chat vs. Markdown file
  6. Deterministic Rendering Rules - Rules A-G (mandatory for Phase 6)
  7. Domain Reference - SecurityEvent, Syslog, CommonSecurityLog interpretation
  8. Tier Classification - Analytics vs Basic vs Data Lake background
  9. Migration Classification - Zero-rule table categorization for §7a
  10. Reference: Data Lake Migration - DL-eligible tables, decision matrix, trade-off analysis
  11. Reference: License Benefits - DfS P2 / E5 pool calculations
  12. Report Template - JIT pointer → SKILL-report.md
  13. Post-Report Drill-Down Reference - Rule cross-referencing, Custom Detection API, ASIM verification, error handling
  14. SVG Dashboard Generation - Visual dashboard from completed report

Quick Start (TL;DR)

3-step execution pattern:

Step 1:  Run Invoke-IngestionScan.ps1 (Phases 1-5 — data gathering)
Step 2:  Read scratchpad + SKILL-report.md (Phase 6 prep)
Step 3:  Render full report (§1-§8) → create_file
Step 1: Run Data Gathering
powershell
# From workspace root — run all phases (default: 30 days):
& ".github/skills/sentinel-ingestion-report/Invoke-IngestionScan.ps1"

# Specify a custom window (1, 7, 30, 60, or 90 days):
& ".github/skills/sentinel-ingestion-report/Invoke-IngestionScan.ps1" -Days 7

# Or run a specific phase (for re-runs / debugging):
& ".github/skills/sentinel-ingestion-report/Invoke-IngestionScan.ps1" -Phase 3

# Synthetic mode — use pre-built test data (no Azure auth required):
& ".github/skills/sentinel-ingestion-report/Invoke-IngestionScan.ps1" -SyntheticDataDir ".github/skills/sentinel-ingestion-report/test-data/enterprise"

Synthetic mode: When the user asks to generate a report using "synthetic data" or "test data", use -SyntheticDataDir pointing to the enterprise test data directory. This bypasses all Azure/Sentinel queries and loads pre-built JSON files instead. Useful for testing report rendering without live workspace access.

Output: Scratchpad file at temp/ingest_scratch_<timestamp>.md (~50 KB, 64 sections).

Timing: Full run (Phase 0 = all phases) takes ~20-25 seconds. Individual phases: 3-8 seconds each.

Step 2: Load Rendering Context
  1. Read the scratchpad file (path printed by PS1 at completion)
  2. Read SKILL-report.md for rendering templates
Step 3: Render Report (Single Write)

Render the complete report (§1-§8) in a single create_file call. Apply SKILL-report.md templates to scratchpad data, following Rules A-G. Render all 8 sections (Executive Summary, Ingestion Overview, Deep Dives, Anomaly Detection, Detection Coverage, License Benefit Analysis, Optimization Recommendations, Appendix) and write to the report file.

⛔ Single-write requirement: The entire report MUST be rendered in one create_file call. Do NOT split rendering across multiple tool calls — splitting causes the LLM to lose template context for later sections (§5-§8), resulting in heading drift, column mutations, and invented content. The complete SKILL-report.md template must be active throughout the entire generation.

🔴 Verbatim table/skeleton blocks — use the deterministic slicer, never hand-copy. The PS1 pre-renders every table, the ASCII cost-waterfall, and the §-heading skeleton under ## PRERENDERED in the scratchpad (Headings, CostWaterfall, DailyChart, TopTables, DetectionPosture, AnomalyTable, CrossReference, SE_Computer, SE_EventID, SyslogHost/Facility/FacSev/Process, CSL_Vendor/Activity, Migration, HealthAlerts, BenefitSummary, DfSP2Detail, E5Tables, QueryTable, Footer). Copy them with the read-only helper instead of transcribing by hand:

powershell
python .github/skills/sentinel-ingestion-report/slice_scratch.py --scratch temp/ingest_scratch_<ts>.md --list
python .github/skills/sentinel-ingestion-report/slice_scratch.py --scratch temp/ingest_scratch_<ts>.md --section AnomalyTable

The slicer prefers the ## PRERENDERED copy when a section name also exists as a raw data block, folds the nested ### lines of the Headings skeleton into one block (so --section Headings returns the full §-heading lock list), strips pipeline scaffolding (<!-- … --> comments, SectionTitle: markers), preserves #### sub-headings, and collapses blank runs — so the output drops straight into the report as a valid markdown table or fenced block. Do NOT paste the raw Key | Value | … data blocks (the early raw sections with a <!-- header --> comment and no |---| separator row) — they render as plain text, not tables, and dumping the whole scratchpad tail into one section corrupts the report.


⚠️ CRITICAL WORKFLOW RULES - READ FIRST ⚠️

Before starting ANY ingestion report:

  1. Run Invoke-IngestionScan.ps1 — this single script handles ALL data gathering (Phases 1-5). The LLM does NOT run queries, transcribe output, or write scratchpad sections
  2. Read config.json for workspace ID, tenant, subscription, and Azure MCP parameters
  3. ALWAYS ask the user for output mode if not specified: inline chat summary, markdown file report, or both (default: both)
  4. ALWAYS ask the user for timeframe if not specified: supported values are 1, 7, 30 (default), 60, or 90 days. The -Days parameter controls the primary window; deep-dive and comparison windows are derived automatically
Date Window Model

The -Days parameter drives three time windows used across all queries:

WindowTokenDerivationPurpose
Primary{days}= -Days valueUsage overview (Q1-Q3), alert firing (Q12), license benefits (Q17/Q17b), tier summary (Q10b)
Deep-dive{deepDiveDays}≤7→Days, ≤30→7, ≤60→14, ≤90→30Table breakdowns (Q4-Q8), rule health (Q11/Q11d), cross-ref (Q13), migration candidates (Q16), WoW "this period" (Q15)
Comparison{wowTotalDays}= deepDiveDays × 2Period-over-period total lookback (Q15)

Example: -Days 60 → primary=60d, deep-dive=14d, comparison=28d

Dynamic period labels: Report column headers adapt automatically ("This Week"/"Last Week" for 7d deep-dive, "This Month"/"Last Month" for 30d, "This Period"/"Last Period" for 14d).

Exception: Q14 (24h anomaly detection) is unaffected by -Days — it uses fixed algorithmic constants (P30D lookback, 29-day weekday baseline). 5. ALWAYS use create_file for markdown reports (NEVER use PowerShell terminal commands) 6. ALWAYS sanitize PII from saved reports — use generic placeholders for real hostnames, workspace names, and tenant GUIDs in committed files 7. Read scratchpad + SKILL-report.md before rendering — the scratchpad is the sole data source for the report 8. Tier display convention — Azure CLI reports Data Lake tier tables as plan Auxiliary internally, but always refer to this tier as "Data Lake" in output — never use "Auxiliary"

Prerequisites
DependencyRequired BySetup
Azure CLI (az)All KQL queries (az monitor log-analytics query), analytic rule inventory (az rest), tier classification (az monitor log-analytics workspace table list)Install: aka.ms/installazurecli. Authenticate: az login --tenant <tenant_id> then az account set --subscription <subscription_id>
log-analytics extensionaz monitor log-analytics query (all KQL queries in Phases 1-5)Install: az extension add --name log-analytics. Verify: az extension list --query "[?name=='log-analytics']"
Azure RBACAzure CLI calls aboveLog Analytics Reader on the workspace (KQL queries + table list). Microsoft Sentinel Reader on the workspace (analytic rule inventory via az rest)
Microsoft.Graph PowerShellQ9b (Custom Detection rules via Invoke-MgGraphRequest)Install-Module Microsoft.Graph.Authentication -Scope CurrentUser. Required Graph scope: CustomDetection.Read.All (interactive consent on first run). PS1 skips gracefully if module not installed or auth fails
PowerShell 7.0+Parallel query executionForEach-Object -Parallel requires PS7+
🔴 PROHIBITED
  • ❌ Running KQL queries via MCP tools during data gathering — PS1 handles all queries
  • ❌ Writing or modifying scratchpad sections manually — PS1 is the sole writer
  • ❌ Reporting cost in dollar amounts — always use GB savings (e.g., "~78.7 GB/month savings")
  • ❌ Fabricating ingestion volumes, device names, or anomaly percentages
  • ❌ Overriding DL eligibility classification from PS1 output based on LLM knowledge
  • ❌ Rendering the report without first reading the scratchpad file

Execution Workflow

Phase 0: Initialization
  1. Read config.json for sentinel_workspace_id, subscription_id, Azure MCP parameters
  2. Confirm output mode and timeframe with user (pass -Days to PS1; default 30)
  3. Verify prerequisites: az login session active, correct subscription set
Phases 1-5: Data Gathering (automated by PS1)

Run Invoke-IngestionScan.ps1 — it handles all 5 phases automatically:

PhaseQueriesDescriptionExecution Type
1Q1, Q2, Q3Core ingestion overview — Usage by DataType, daily trend, workspace summaryKQL (parallel)
2Q4, Q5, Q6a, Q6b, Q6c, Q7, Q8Table deep dives — SecurityEvent, Syslog, CommonSecurityLog breakdownsKQL (parallel)
3Q9, Q9b, Q10, Q10bExternal data — analytic rule inventory (REST), custom detections (Graph), tier classification (CLI), tier summary (KQL)REST + Graph + CLI + KQL (sequential, with depends_on)
4Q11, Q11d, Q12, Q13Detection coverage — rule health (SentinelHealth), alert firing (SecurityAlert), cross-reference (all tables with data vs. rule inventory)KQL (parallel) + post-processing
5Q14, Q15, Q16, Q17, Q17bAnomaly detection + cost analysis — 24h anomaly, WoW comparison, migration candidates, license benefits, E5 per-tableKQL (parallel) + post-processing

Post-processing (automated by PS1, Phases 4-5):

TaskPhaseDescription
Table cross-reference4For each table with data (Q13), regex-search all enabled rule queries for that table name
ASIM parser detection4Search all rule queries for ASIM function patterns (_Im_, _ASim_, imDns, etc.)
Value-level rule verification4For each EventID/Facility/ProcessName/Activity/Vendor from deep dives, check if any rules reference it
Detection gap detection4Identify tables on DL/Basic tier that have enabled rules (🔴 critical finding)
Anomaly severity classification5Apply Rule A thresholds to Q14/Q15 results
DL eligibility classification5Classify all tables using hardcoded $dlYes/$dlNo reference arrays
Migration table assembly5Cross-reference volume × rule count × tier × DL eligibility → category assignment
License benefit computation5Compute DfS P2 pool, E5 grant breakdown

Scratchpad output: PS1 writes all results to temp/ingest_scratch_<timestamp>.md (~50 KB, ~64 named sections including PHASE_, PRERENDERED, and META blocks). See SKILL-report.md for the Section-to-Scratchpad Mapping.

Phase 6: Render Output (LLM)

🔴 MANDATORY — Load scratchpad + report template before rendering:

  1. Read the scratchpad file (path printed by PS1). This single file contains ALL data from Phases 1-5.
  2. Read SKILL-report.md for the complete rendering templates and formatting rules.

Pre-render validation:

  1. Verify scratchpad has all 5 phase sections (PHASE_1 through PHASE_5)
  2. Check that PHASE_5.DL_Script_Output is populated (proof of DL classification execution)
  3. Cross-validate: Q11 TotalRulesInHealth against Q9 AR_Enabled — if >10% gap, note it

Render — Section-by-Section Checklist:

Render the report section by section per SKILL-report.md templates. Do NOT skip any section. If a section's data returned 0 results, render the section header with a "✅ No anomalies/items found" note.

SectionData Source (scratchpad keys)Required
§1All phases✅ Workspace at a Glance, Cost Waterfall, Detection Posture, Top 3
§2PHASE_1.Tables, PHASE_3.TierSummary✅ Table breakdown + tier summary
§3PRERENDERED.SE_*, PRERENDERED.Syslog*, PRERENDERED.CSL_*✅ Deep dives (skip sub-section only if table not in top 20)
§4PHASE_5.Anomaly24h/AnomalyWoW, PHASE_1.DailyTrend✅ Anomaly table + daily trend chart
§5PHASE_3.RuleInventory, PHASE_4.*✅ Rule inventory + cross-ref + health
§6PHASE_5.LicenseBenefits/E5_Tables✅ DfS P2 + E5 analysis
§7PHASE_5.Migration, PHASE_4.CrossRef✅ Migration candidates + recommendations
§8All✅ Appendix (query reference, methodology)

Compute Top 3 Recommendations using Rule E: scan all scratchpad sections, score each candidate, select the top 3 by score.

Post-render:

  • Render inline chat executive summary (if requested)
  • Confirm markdown file path to user

Query File Reference

All queries are defined as YAML files in queries/phase1-5/. PS1 discovers, parses, and executes them automatically.

YAML Format
yaml
id: ingestion-q1                              # Unique identifier
name: Usage by DataType with Billing Breakdown # Human-readable name
description: Top 20 tables ranked by volume    # What it does
phase: 1                                       # Which phase (1-5)
type: kql                                      # kql | rest | cli | graph
timespan: P{days}D                             # Placeholder — PS1 substitutes at runtime
query: |                                       # KQL query (multiline block scalar)
  Usage
  | where TimeGenerated > ago({days}d)
  ...

Non-KQL types have additional fields:

TypeAdditional FieldsDescription
resturl, method, jmespathSentinel REST API via az rest
clicommandAzure CLI command (e.g., az monitor log-analytics workspace table list)
graphuri, methodMicrosoft Graph API via Invoke-MgGraphRequest
Complete Query Inventory
PhaseFileIDTypeDescription
1Q1-UsageByDataType.yamlingestion-q1kqlTop 20 tables by billable volume with solution mapping
1Q2-DailyIngestionTrend.yamlingestion-q2kqlDaily ingestion trend
1Q3-WorkspaceSummary.yamlingestion-q3kqlExecutive summary: table count, billable totals, daily average
2Q4-SecurityEventByComputer.yamlingestion-q4kqlSecurityEvent by Computer (top 25)
2Q5-SecurityEventByEventID.yamlingestion-q5kqlSecurityEvent by EventID (top 20)
2Q6a-SyslogByHost.yamlingestion-q6akqlSyslog by source host (top 25)
2Q6b-SyslogByFacilitySeverity.yamlingestion-q6bkqlSyslog by Facility × SeverityLevel (top 30)
2Q6c-SyslogByProcess.yamlingestion-q6ckqlSyslog top ProcessName by Facility (top 30)
2Q7-CSLByVendor.yamlingestion-q7kqlCommonSecurityLog by DeviceVendor/DeviceProduct (top 20)
2Q8-CSLByActivity.yamlingestion-q8kqlCommonSecurityLog by Activity/LogSeverity/DeviceAction (top 30)
3Q9-AnalyticRuleInventory.yamlingestion-q9restAnalytic rules (Scheduled + NRT) via Sentinel REST API
3Q9b-CustomDetectionRules.yamlingestion-q9bgraphCustom Detection rules via Microsoft Graph SDK
3Q10-TableTierClassification.yamlingestion-q10cliTable tier classification via Azure CLI
3Q10b-TierSummary.yamlingestion-q10bkqlPer-tier volume summary (depends_on: Q10)
4Q11-RuleHealthSummary.yamlingestion-q11kqlSentinelHealth — rule execution health summary
4Q11d-FailingRuleDetail.yamlingestion-q11dkqlSentinelHealth — top 20 failing rules detail
4Q12-SecurityAlertFiring.yamlingestion-q12kqlSecurityAlert — top 30 alert-producing rules
4Q13-AllTablesWithData.yamlingestion-q13kqlAll tables with data in deep-dive window (for cross-reference)
5Q14-IngestionAnomaly24h.yamlingestion-q14kql24h vs same-weekday avg anomaly detection (29d lookback, fallback to flat 7d, >50%, ≥0.01 GB)
5Q15-WeekOverWeek.yamlingestion-q15kqlPeriod-over-period volume comparison
5Q16-MigrationCandidates.yamlingestion-q16kqlBillable tables with deep-dive volume (for migration analysis)
5Q17-LicenseBenefitAnalysis.yamlingestion-q17kqlDfS P2 + E5 daily ingestion breakdown
5Q17b-E5PerTableBreakdown.yamlingestion-q17bkqlE5-eligible per-table volume

Output Modes

Mode 1: Inline Chat Summary (default for quick requests)

Compact executive summary rendered directly in chat.

Mode 2: Markdown File Report

Full detailed report saved to reports/sentinel/sentinel_ingestion_report_<YYYYMMDD_HHMMSS>.md.

Mode 3: Both (default when user says "report" or "generate report")

Inline chat executive summary + full markdown file.

Ask user if not specified:

"How would you like the report? I can provide:

  1. Inline chat summary — executive overview in chat
  2. Markdown file — detailed report saved to reports/sentinel/
  3. Both (recommended) — summary in chat + full report file"

Deterministic Rendering Rules

These rules eliminate LLM interpretation variance. Apply them EXACTLY during report rendering (Phase 6). No discretion allowed — the thresholds and formulas below are the sole authority.

Rule A: Anomaly Severity Classification

⚙️ Pre-computed by PS1 → PRERENDERED.AnomalyTable. Thresholds below retained for §8 methodology reference and manual verification.

Assign severity to each anomaly row deterministically based on absolute deviation AND volume.

Condition (both must be true)SeverityEmoji
abs(Deviation%) ≥ 200 AND max(Last24hGB, Avg7dGB) ≥ 0.05 GBHigh🟠
abs(Deviation%) ≥ 100 AND max(Last24hGB, Avg7dGB) ≥ 0.01 GBMedium🟡
abs(Deviation%) ≥ 50 AND max(Last24hGB, Avg7dGB) ≥ 0.01 GBLow⚪
Below thresholds OR both periods < 0.01 GB volumeExcluded—

Volume floor: The 0.01 GB minimum is enforced by the KQL queries. Tables below this floor are noise and MUST NOT appear in the anomaly table regardless of deviation percentage.

Override 1 — Rule-count: ANY table with ≥5 enabled rules AND an absolute change ≥40% (in either 24h or WoW) is automatically 🟠 regardless of base thresholds — a significant drop on a table feeding multiple rules signals potential connector or TI feed health issues that affect detection coverage. The 24h override catches same-day connector outages; the WoW override catches gradual multi-day degradation.

Override 2 — Near-zero: ANY table with deviation ≤ −95% AND max(volume) ≥ 0.05 GB is automatically 🟠 regardless of rule count — a near-complete signal loss on a significant table is an operational emergency (e.g., connector failure, API key expiry) even if no rules reference it directly.

⛔ PROHIBITED: Assigning severity based on "judgment", "context", or "this table is important" UNLESS the high-rule-count override above applies. Outside that specific override, the emoji MUST match the threshold table above — no discretionary overrides.

Rule B: Risk Rating Definition

In the Top 3 Recommendations table (§1), the "Risk" column means:

Risk = the security or operational impact of NOT acting on this recommendation.

Risk LevelDefinitionExamples
HighActive detection gap or data loss if not addressedRules silently failing on DL tier; connector dropping data; 0% detection coverage on critical table
MediumMissed optimization with measurable cost/posture impactZero-rule high-volume table on Analytics tier; noisy EventID with no detection value
LowMinor improvement, no immediate security or cost impactSmall-volume table tier change; informational tuning

⛔ PROHIBITED: Interpreting "Risk" as implementation difficulty, effort, or change management complexity. Those concerns belong in prose recommendations (§7b-d), NOT the Risk column.

Rule C: Weekday Average Computation

⚙️ Pre-computed by PS1 → PRERENDERED.DailyChart. Logic below retained for §8 methodology reference.

When computing per-weekday averages for the §4b daily trend chart:

  1. Exclude the report-generation day: If the last day in PHASE_1.DailyTrend matches META.Generated date, always exclude it from weekday averages — the report was generated mid-day so this is a partial day regardless of its volume. This prevents the partial day from non-deterministically dragging down whichever weekday it falls on.
  2. Exclude ingestion gaps: Any remaining day with total ingestion < 0.1 GB is also excluded. These are ingestion reporting gaps, not representative of normal patterns.
  3. Formula: Weekday Avg = sum(GB for that weekday, excluding days per rules 1–2) / count(qualifying days for that weekday)
  4. Round to 2 decimal places.

⛔ PROHIBITED: Including the report-generation partial day or days with < 0.1 GB in averages — they drag down specific weekdays non-deterministically.

Rule D: Cross-Validation Denominator

In §5b cross-validation (Q11 vs Q9), always use AR-only enabled count from Q9 as the denominator:

Gap% = (Q9_AR_Enabled - Q11_DistinctRules) / Q9_AR_Enabled × 100

Do NOT use Combined_Enabled (AR+CD) as the denominator. SentinelHealth only tracks AR executions (Scheduled + NRT), not Custom Detection executions. Comparing Q11 against combined AR+CD inflates the gap percentage.

Rule E: Top 3 Recommendation Ranking

Rank ALL candidate recommendations using this scoring formula. The top 3 by score become the Top 3 in §1. Computed by the LLM at render time by cross-referencing all scratchpad sections.

CategorySeverityWeightImpactValueScratchpad Source
🔴 Detection gap (rules on wrong tier)10Number of affected rulesPHASE_4.DetectionGaps — PS1 emits Detection gap (XDR) or Detection gap (non-XDR) in §7a Category column
🔴 Data loss / connector failure10Affected volume in GB/dayPHASE_5.Anomaly24h (large negative deviations)
🟠 DL-eligible migration (zero rules)5BillableGB from §2a (or deep-dive GB / deepDiveDays × Days if only in §7a)PHASE_5.Migration (Strong DL-eligible rows)
🟠 DL + KQL Job promotion4BillableGB (primary window)High-volume 🟣/🟢 table — can complement split ingestion or stand alone; present both options and note they are combinable
🟠 License benefit activation4Eligible unclaimed GB/dayPRERENDERED.BenefitSummary + PRERENDERED.E5Tables + PRERENDERED.DfSP2Detail (volume eligible but benefit not yet activated)
🟠 DCR filter / EventID pruning4Estimated saveable GB (deep dive % × table BillableGB)PHASE_2.SE_EventID + PHASE_4.ValueRef_EventID
🟠 Health fix (failing rules)4Number of failing rulesPHASE_4.FailingRules
🟡 Volume spike / cost anomaly3Spike GB on zero-rule tablesPHASE_5.Anomaly24h (large positive deviations on zero-rule tables — cost spike with no detection value)
🟡 Duplicate ingestion3Duplicate GBCross-ref PRERENDERED.SyslogFacility × PRERENDERED.CSL_Vendor (same-appliance overlap emitting both Syslog and CEF/ASA = double billing)
🟡 Split ingestion3BillableGB × estimated non-detection fractionPHASE_2 deep dives + PHASE_4.ValueRef_* (zero-rule values)
🟡 Tier review / unknown eligibility2BillableGBPHASE_5.Migration (Unknown rows)

Score = SeverityWeight × ImpactValue

Sorting: severity-first, then score. 🔴 items always rank above 🟠 items, which always rank above 🟡 items, regardless of score. Within the same severity tier, rank by descending score. This ensures detection gaps and data loss signals are never buried below cost optimizations.

Tie-breaking within same severity: higher score wins. If scores are equal, higher SeverityWeight wins. If still tied, higher ImpactValue wins.

⛔ PROHIBITED: Selecting Top 3 recommendations based on narrative variety, "one from each category", or subjective importance. The formula determines ranking — the LLM renders, it does not curate.

License benefit activation: Surfaces when PRERENDERED.BenefitSummary or PRERENDERED.E5Tables show E5-eligible or DfS-P2-eligible volume that is not yet being claimed (benefit shows 0 or is absent while eligible tables are ingesting). ImpactValue = the eligible GB/day that could be offset.

Volume spike / cost anomaly: Surfaces when PHASE_5.Anomaly24h shows a large positive deviation (>50% above baseline) on a table with zero detection rules (per PHASE_4.CrossRef). A spiking table with no rules has cost impact but no detection value — a strong signal to investigate and potentially filter or move to DL.

Duplicate ingestion: Surfaces when the same network appliance sends data via both Syslog and CommonSecurityLog (CEF/ASA). Compare appliance names/IPs in PRERENDERED.SyslogFacility/PRERENDERED.SyslogHost against PRERENDERED.CSL_Vendor — overlapping sources indicate double billing for the same data. ImpactValue = the smaller of the two streams (the duplicate portion).


Domain Reference

This section provides the domain knowledge needed during Phase 6 rendering. When writing deep dive sections (§3), anomaly analysis (§4), and recommendations (§7), consult these reference tables for interpretation guidance.

SecurityEvent — EventID Optimization

Which EventIDs generate the most volume and their detection vs. cost tradeoff:

EventIDDescriptionOptimization Potential
4663Object access (file auditing)🔴 High — often excessive. Consider DCR drop filter or scoping SACL
4624Successful logon🟡 Medium — valuable for hunting/forensics but rarely in analytic rules. Strong split ingestion candidate: send to Data Lake for retention, keep off Analytics tier
4688Process creation🟡 Medium — consider moving to MDE DeviceProcessEvents. If no rules reference it, split to Data Lake
4799Security group membership enumeration🟡 Medium — often noisy on domain controllers
4672Special privileges assigned🟡 Medium — high volume on DCs
4625Failed logon🟢 Low — usually valuable for security detection

🟣 Split ingestion tip: For any deep-dive table classified as 🟢 Keep Analytics (active detection rules), individual high-volume values with zero rule references (verified via Phase 4 value-level check) are strong candidates for sub-table split ingestion. Route those values to Data Lake via DCR transformation — they remain available for hunting while the detection-relevant values stay on Analytics tier. KQL jobs can also run against this split-routed DL data to surface aggregated insights back to Analytics if needed.

Syslog — Facility Reference

Optimization potential by Syslog facility:

FacilityDescriptionOptimization Potential
authAuthentication events (login, su, getty)🟢 Low — always security-relevant. Keep in Analytics tier
authprivPrivate authentication (PAM, sudo, sshd)🟢 Low — critical for security detection. Always keep
kernKernel messages (hardware, driver, critical system)🟡 Medium — security-relevant but can be noisy. Consider Error+ only for high-volume servers
cronScheduled task notifications🔴 High — rarely security-relevant at Info/Notice. Keep Warning+ only
daemonSystem daemon messages (systemd, sshd, named, httpd)🔴 High — typically largest Syslog contributor (50-80% of volume). Contains both security-critical processes (sshd) and noisy infrastructure (systemd). Drill down with Q6c to identify filterable processes
syslogInternal syslog daemon messages🟡 Medium — mostly operational. Keep Warning+ in Analytics
userUser-space application messages🟡 Medium — varies by application. Check ProcessName
mailMail subsystem (postfix, sendmail, dovecot)🟡 Medium — relevant if mail is in scope; otherwise DL candidate
local0–local7Custom application logs🔴 High — most common cost optimization targets. Custom apps often log at Debug/Info verbosity
ftpFTP daemon messages🟢 Low volume; keep for auditing if FTP in use
lprPrint subsystem🔴 High — almost never security-relevant. Set to None in DCR
newsNetwork news (NNTP)🔴 High — almost never security-relevant. Set to None in DCR
uucpUUCP subsystem🔴 High — almost never security-relevant. Set to None in DCR
markInternal timestamp marker🔴 High — operational only. Set to None in DCR
Syslog — DCR Severity-per-Facility Recommendations

The Data Collection Rule allows setting a minimum severity level per facility — the single most impactful cost control for Syslog:

FacilityRecommended MinimumRationale
auth, authprivDebug (collect all)Security-critical — never filter
kernNoticeKernel module loads (T1547.006) and promiscuous mode (T1040) are kern.notice. Volume impact is minimal
daemonWarning or ErrorMajor volume reduction. Note: sshd auth events go to auth/authpriv, not daemon. Trade-off: loses systemd service stop events at Info (security service tampering) — acceptable if EDR covers this
cronWarningTrade-off: cron job execution events are cron.info (T1053.003 persistence). Acceptable if auditd or MDE covers cron file monitoring
syslogWarningInternal operational messages are low-value at Info
userWarningUnless specific apps produce security telemetry
mailWarningInfo-level mail relay logs are very verbose
local0–local7Assess per-appNo safe default — network appliances, security tools, and databases commonly use local facilities. Review Q6c (Process by Facility) before setting severity filters
lpr, news, uucp, markNoneDisable collection entirely
Syslog — SeverityLevel Values
SeverityLevel (string)NumericMeaningRetention Priority
emerg0System unusable🔴 Always keep
alert1Immediate action required🔴 Always keep
crit2Critical condition🔴 Always keep
err3Error condition🟡 Keep for most facilities
warning4Warning condition🟡 Keep for security-relevant facilities
notice5Normal but significant🟡 Keep for auth/authpriv and kern
info6Informational🟢 Filter for high-volume facilities
debug7Debug-level detail🟢 Filter everywhere except auth/authpriv
Syslog — ProcessName Security Relevance
ProcessNameTypical FacilitySecurity RelevanceOptimization
systemddaemon🟡 Low-Medium — unit start/stop events🔴 Often 30-50% of daemon volume. Filter Info/Notice at DCR
systemd-loginddaemon🟡 Medium — session/seat trackingKeep Warning+
sshdauth, authpriv, daemon🟢 High — SSH login detection (brute force, lateral movement)🟢 Always keep
sudoauthpriv🟢 High — privilege escalation tracking🟢 Always keep
suauth, authpriv🟢 High — user switching🟢 Always keep
CRON / crondcron🟡 Low-Medium — scheduled tasksKeep Warning+ unless monitoring for T1053
named / binddaemon🟡 Medium — DNS. Relevant for DNS tunnelingKeep if DNS rules exist; otherwise Warning+
httpd / nginxdaemon🟡 Medium — web server logsAssess overlap with WAF/CSL data
postfix / sendmailmail🟡 Low-Medium — mail relayKeep Warning+
dhclient / NetworkManagerdaemon🟡 Low — DHCP/network changesFilter Info/Notice
kernelkern🟢 Medium-High — kernel events, module loadsKeep Warning+
auditddaemon, user🟢 High — Linux Audit Framework🟢 Always keep
polkitdauthpriv🟡 Medium — PolicyKit authorizationKeep Warning+
dbus-daemondaemon🟡 Low — IPC. Rarely security-relevantFilter all or keep Error+
rsyslogd / syslog-ngsyslog🟡 Low — internal syslog opsKeep Warning+

🟣 Split ingestion tip: If daemon facility accounts for >50% of Syslog and Q6c reveals systemd + systemd-logind + dbus-daemon dominate, consider a DCR transformation routing those processes to Data Lake while keeping sshd, auditd, and other security-critical processes in Analytics. KQL jobs can complement this by querying the DL-routed portion on schedule.

Syslog — Log Forwarding Architecture Note

In environments using centralized rsyslog/syslog-ng forwarders:

  • Computer = the log forwarder hostname (many servers collapse to 1-2 forwarders)
  • HostName = the actual originating device (from syslog header)
  • HostIP = the originating device's IP address

The Q6a query uses SourceHost = iff(isnotempty(HostName) and HostName != Computer, HostName, Computer) to prefer the original source. If Q6a shows only 1-2 hosts despite expecting 100+ servers, the environment uses forwarding.

CommonSecurityLog — Vendor Reference
DeviceVendorDeviceProductOptimization Potential
Palo Alto NetworksPAN-OS🔴 High — filter TRAFFIC activity, keep THREAT in Analytics
Check PointFirewall / VPN-1 & FireWall-1🔴 High — filter routine Accept actions
FortinetFortigate🔴 High — filter traffic subtype, keep utm and event
CiscoASA🟡 Medium — filter by message ID ranges
ZscalerNSSWeblog🟡 Medium — web proxy logs can be high volume
F5BIG-IP ASM / LTM🟡 Medium — WAF logs can spike during attacks
Trend MicroDeep Security🟢 Low — typically moderate volume

Firewall traffic/session logs often account for 60-80% of CSL volume. These are primarily TRAFFIC or Accept events with low detection value. Consider DCR transformation, Data Lake tier, split ingestion, or DL + KQL job promotion (these last two can be combined).

Show full SKILL.md (2,804 more words)Show less
CommonSecurityLog — LogSeverity Values
Value (string)Value (int)MeaningRetention Priority
Very-High9-10Critical security event🔴 Always keep in Analytics
High7-8Significant security event🔴 Keep in Analytics
Medium4-6Notable event🟡 Review — may be filterable
Low0-3Informational event🟢 Candidate for DL or DCR filter
(empty/Unknown)—Unmapped severity⚠️ Check vendor documentation

DeviceAction optimization: If >70% of events have DeviceAction = "Allow" or "Accept", the table is dominated by permitted traffic. Filter at DCR level or move to Data Lake, keeping only denied/blocked/threat events in Analytics.

Anomaly Interpretation (Q14/Q15)

24h anomalies (Q14): Flags tables where last-24h ingestion deviates >50% from the same-weekday daily average AND at least one period has ≥0.01 GB volume. Q14 uses a fixed 29-day lookback (algorithmic constant, not affected by -Days).

  • Positive spikes: May indicate attacks, misconfigured connectors, or bulk imports
  • Negative drops: May indicate connector failures, agent issues, or collection gaps

Period-over-period (Q15): Compares total volume per table between current and prior period (period length = deep-dive window).

  • New tables (100% change) → appeared only this period (new connector?)
  • Growing tables → expanding collection scope or increased activity
  • Shrinking tables → connector removal, collection changes, or seasonal patterns
  • Stable high-volume tables → included via ThisWeekMB > 100 filter for visibility

Tier Classification

Background

The Sentinel Usage table does NOT contain a TablePlan or Tier column. There is no KQL-native way to determine whether a table is on Analytics, Basic, or Data Lake tier.

PS1 handles this automatically: Q10 (CLI type) runs az monitor log-analytics workspace table list to fetch table plans, then Q10b (KQL, depends_on: Q10) computes per-tier volume summaries using the CLI output. The results are written to PHASE_3.Tiers and PHASE_3.TierSummary in the scratchpad.

Tier Display Convention

Azure CLI reports Data Lake tier tables as plan Auxiliary internally. Always refer to this tier as "Data Lake" in all output — never use "Auxiliary". The _CL suffix denotes a custom log table, not a copy — describe these as "Custom Data Lake table" (not "Auxiliary copy").

Q10b Cross-Reference Query

PS1 automatically populates the DataLakeTables and BasicTables arrays from CLI output and executes the tier summary KQL query. This computes per-tier TotalGB, BillableGB, TableCount, and PercentOfTotal using the full Usage table (not limited to Q1 top-20). These values are the authoritative source for PHASE_3.TierSummary and §2b rendering.


Migration Classification

Used when rendering §7a (Tier Migration Candidates). PS1 computes the Category column using these criteria; the LLM uses this reference for rendering interpretation and recommendation prose.

CategoryCriteriaAction
🔵 KQL Job outputTable name ends with _KQL_CLNEVER migrate — promoted data from Data Lake, essential for detection pipeline
🔵 Already on Data LakeQ10 tier = Data Lake AND zero rulesAlready migrated — no action needed
🟢 Keep Analytics≥1 enabled analytic rule AND healthy executionsActive detection coverage justifies Analytics cost
🟣 Split ingestion candidate1-2 enabled rules AND high-volume (≥5 GB/week) AND DL-eligibleFew rules need only a subset of events. Route detection-relevant subset to Analytics via DCR, rest to Data Lake
❗ Detection gap (non-XDR)≥1 enabled rule AND table is on Data Lake tier AND table is NOT an XDR tableCritical: Analytic rules cannot execute against DL tables — rules silently failing. Custom Detections also do NOT work because non-XDR tables are not available in Advanced Hunting on Data Lake. Remediation: (1) move table back to Analytics, OR (2) remove/disable the analytic rules referencing the table (accept DL tier). ⛔ PROHIBITED: Recommending "convert ARs to Custom Detections" for non-XDR tables — CDs run against Advanced Hunting which only retains Defender XDR tables for 30 days. Non-XDR tables on Data Lake are invisible to Advanced Hunting.
❗ Detection gap (XDR)≥1 enabled rule AND table is on Data Lake tier AND table IS an XDR tablePartial gap: Sentinel Analytic Rules (AR) cannot execute against DL tables — ARs silently failing. However, XDR-native tables (Device*, Email*, CloudAppEvents, UrlClickEvents) are ALWAYS available in Advanced Hunting for 30 days regardless of Sentinel tier. Custom Detection rules run against Advanced Hunting, so CD rules continue to work. Only ARs are broken. Remediation: (1) move table back to Analytics, (2) convert affected ARs to Custom Detections, OR (3) remove/disable the ARs. See Advanced Hunting data retention
🔴 Strong candidate (DL-eligible)0 rules AND DL classification = YesEvaluate DCR filtering to reduce unnecessary volume, then migrate remainder to Data Lake

LLM overlay checks (not separate emojis — flag as callout notes in §7b/7c prose):

  • Execution issues: If a 🟢 table's rules appear in PHASE_4.FailingRules with 0 executions or failures, add a ⚠️ note: "Rules targeting [table] have execution issues — see §5b. Fix rules before relying on this coverage."
  • ASIM dependency: If a 🔴 zero-rule table appears in PHASE_4.ASIM as consumed by ASIM parsers, add a ⚠️ note: "[table] is consumed by ASIM parsers ([parser names]) — migrating to Data Lake breaks these detections. Verify ASIM dependency before migrating." | 🟠 Not DL-eligible / unknown | 0 rules AND DL classification = No or Unknown | Optimize via DCR filtering or add analytic rules. Check MS docs for current eligibility |

Reference: Data Lake Migration

This section contains lookup tables and background guidance for DL migration classification. Consult when rendering §7a recommendations and explaining Data Lake trade-offs.

Known DL-Eligible Tables

PS1 uses these lists as hardcoded $dlYes/$dlNo arrays. Keep this reference in sync with the script.

CategoryDL-Eligible TablesNotes
Defender XDRCloudAppEvents, DeviceEvents, DeviceFileCertificateInfo, DeviceFileEvents, DeviceImageLoadEvents, DeviceInfo, DeviceLogonEvents, DeviceNetworkEvents, DeviceNetworkInfo, DeviceProcessEvents, DeviceRegistryEvents, EmailAttachmentInfo, EmailEvents, EmailPostDeliveryEvents, EmailUrlInfo, UrlClickEventsGA Feb 2025
Verified LA tablesAADManagedIdentitySignInLogs, AADNonInteractiveUserSignInLogs, AADProvisioningLogs, AADServicePrincipalSignInLogs, AADUserRiskEvents, AuditLogs, AWSCloudTrail, AzureDiagnostics, CommonSecurityLog, Event, GCPAuditLogs, LAQueryLogs, McasShadowItReporting, MicrosoftGraphActivityLogs, OfficeActivity, Perf, SecurityAlert, SecurityEvent, SecurityIncident, SentinelHealth, SigninLogs, StorageBlobLogs, Syslog, W3CIISLog, WindowsEvent, WindowsFirewall⚠️ Only these LA tables are verified DL-eligible. Unlisted → Unknown
Custom tablesAny table ending in _CL (except _KQL_CL)Custom log tables are workspace-managed → DL-eligible
Known DL-Ineligible Tables (as of Feb 2026)
CategoryIneligible TablesNotes
XDR — not yet supportedDeviceTvmSoftwareInventory, DeviceTvmSoftwareVulnerabilities, AlertEvidence, AlertInfo, IdentityDirectoryEvents, IdentityLogonEvents, IdentityQueryEventsMDI tables announced for future DL support
Entra ID ❌MicrosoftServicePrincipalSignInLogs, MicrosoftNonInteractiveUserSignInLogs, MicrosoftManagedIdentitySignInLogsNot yet DL-eligible
Threat Intelligence ❌ThreatIntelIndicators, ThreatIntelligenceIndicatorRequired on Analytics for TI matching rules. Never recommend migration
Log Analytics ❌AppDependencies, AppMetrics, AppPerformanceCounters, AppTraces, AzureActivity, AzureMetrics, ConfigurationChange, Heartbeat, SecurityRecommendationNot yet DL-eligible

Fallback rule: If a table is not in either list, the script classifies it as Unknown. Render as ❓ Unknown with note: "Verify at Manage data tiers before migrating."

Decision Matrix
Enabled RulesExecutions (Health)Alerts (Q12)DL-Eligible?VolumeRecommendation
0N/A0✅ Yes> 1 GB/week🔴 Evaluate DCR filtering to reduce volume, then migrate remainder to Data Lake (confirm no ASIM dependency)
0N/A0✅ Yes< 1 GB/week🔴 Migrate to Data Lake — minimal savings but cleaner tier alignment. DCR filtering optional at this volume
0N/A0❌ No / ❓ UnknownAny🟠 Not eligible or unknown — review ingestion necessity, apply DCR filtering
0N/A (on DL)0N/A — already DLAny🔵 Already on Data Lake — no action needed
0 (ASIM-dependent)N/A0AnyAny� Migrate — but LLM adds ⚠️ ASIM dependency callout in §7b
≥10 or failures0AnyAny🟢 Keep — but LLM adds ⚠️ execution issues callout in §7b
≥10 (on DL)AnyN/A — on DLAny🔴 Detection gap — ARs cannot execute against DL. PS1 emits Detection gap (XDR) or Detection gap (non-XDR). If XDR table: CDs still work via Advanced Hunting; recommend converting ARs→CDs or moving back to Analytics. If non-XDR table: move back to Analytics OR remove/disable rules. ⛔ NEVER recommend CD conversion for non-XDR tables
1-2> 0, healthyAny✅ Yes≥ 5 GB/week🟣 Split ingestion candidate
≥1> 0, healthy0AnyAny🟢 Keep Analytics — rules executing, no matches (normal for TI rules)
≥1> 0, healthy> 0AnyAny🟢 Keep Analytics — active detections generating alerts
Data Lake Trade-Off
CapabilityAnalytics TierData Lake Tier
Analytics rules, alerting, hunting✅ Full support❌ Not available (but see XDR exception below)
Custom Detection rules (Advanced Hunting)✅ Full support⚠️ XDR tables only: Still available — AH retains 30 days regardless of Sentinel tier. Non-XDR tables: ❌
Workbooks, playbooks, parsers, watchlists✅ Full support❌ Not available
KQL query performance✅ High-performance⚠️ Slower
Query cost✅ Included in ingestion price❌ Billed per query (data scanned)
KQL Jobs / Summary Rules / Search Jobs✅✅
Ingestion costStandardMinimal
Default retention90 days (Sentinel) / 30 days (XDR)Matches analytics, extendable to 12 years

Primary vs secondary security data: Primary security data (EDR alerts, auth logs, audit trails) belongs on Analytics. Secondary data (NetFlow, storage access logs, firewall traffic, IoT logs) is ideal for Data Lake.

Filter before you migrate: For high-volume zero-rule tables, DL migration and DCR filtering are complementary — not mutually exclusive. Evaluate whether all ingested data serves a hunting, forensic, or compliance purpose. If a portion is noise (e.g., verbose diagnostics, routine health checks, debug-level telemetry), apply DCR transformations to drop or reduce that portion first, then migrate the meaningful remainder to Data Lake. This avoids simply shifting cost from Analytics to Data Lake query charges on data nobody uses.

Even when a table has zero rules, consider whether it serves hunting/forensic purposes. Tables like SigninLogs or AuditLogs should generally remain on Analytics regardless.

Data Lake Promotion via KQL Jobs

For high-volume tables on Data Lake — whether fully migrated or partially routed via split ingestion — that still need detection coverage:

  1. Ingest raw logs into Data Lake tier (cheap)
  2. Create KQL jobs to query Data Lake on schedule, writing aggregated results to Analytics-tier _KQL_CL tables
  3. Point analytics rules at the _KQL_CL output table

KQL Job key facts: Full KQL (joins, unions, CTEs). Schedules: by-minute through monthly. Lookback up to 12 years. Limits: 3 concurrent / 100 enabled per tenant, 1hr query timeout. Data Lake has ~15-min ingestion latency — jobs should use now(-15m) as upper bound. TimeGenerated is overwritten if >2 days old — preserve source timestamps in a custom column.

Split Ingestion and/or DL + KQL Job Promotion

PS1 auto-classifies 🟣 Split candidates (1-2 rules, ≥5 GB/week, DL-eligible). For these tables (and high-volume 🟢 Keep tables), the report should present both optimization paths so the operator can choose — or combine them — based on their knowledge of the rule queries:

Split Ingestion (DCR)DL + KQL Job
How it worksDCR routes a detection-relevant subset to Analytics, bulk to DLAny data on DL (full table or split-routed portion); KQL job promotes aggregated results to _KQL_CL on Analytics
Detection latencyReal-time (subset stays on Analytics)15+ min (DL ingestion lag + job schedule)
Rule rewrite neededNo — rules keep targeting original tableYes — rules must target _KQL_CL output
Volume savingsModerate (bulk to DL, subset stays)Depends on scope — maximum if entire table goes to DL, incremental if applied to split-routed portion
Best whenRules filter on specific raw events (EventIDs, facilities)Rules use aggregation and tolerate latency

These approaches are complementary, not mutually exclusive. Split ingestion routes bulk data to DL while keeping detection-relevant events on Analytics. KQL jobs can then run against that DL portion to surface additional insights (e.g., aggregated anomalies) back to Analytics via _KQL_CL tables — giving you both real-time detection on the split subset AND scheduled analytics on the DL bulk.

Rendering guidance: The LLM does NOT have visibility into rule query text (aggregation vs raw filters), so it cannot definitively recommend one over the other. For 🟣 tables and high-volume 🟢 tables, present the comparison and note which approach fits which rule pattern. Do NOT change PS1's Category emoji in §7a — express as prose in §7b/7c.

References:


Reference: License Benefits

Defender for Servers P2 — 500MB/Server/Day Benefit
  • Each server protected by DfS P2 contributes 500 MB/day to a pooled daily allowance
  • Pool = (number of protected servers) × 500 MB — aggregate across subscription, not per-machine
  • Applies to security data types: SecurityAlert, SecurityBaseline, SecurityBaselineSummary, SecurityDetection, SecurityEvent, WindowsFirewall, MaliciousIPCommunication, SysmonEvent, ProtectionStatus, Update, UpdateSummary
  • Applied automatically at workspace level — shows as zero cost

Pool calculation from Q4:

Potential DfS P2 Pool = (Distinct servers from Q4) × 500 MB/day

Example: Q4 shows 12 servers → pool = 6 GB/day. If DFSP2-eligible avg is 4.2 GB/day → fully covered.

ScenarioConditionRecommendation
Pool far exceeds usageDfSP2_DailyGB < 50% of PoolGBHighlight the unused headroom and recommend increasing SecurityEvent logging levels (e.g., "All Events" instead of "Common") to broaden detection coverage at no additional ingestion cost. Note that increased data volume may affect retention storage costs
Pool covers usageDfSP2_DailyGB ≥ 50% and ≤ 100% of PoolGBPool covers current need — monitor growth and reference §3a if approaching ceiling
Usage exceeds poolDfSP2_DailyGB > PoolGBOverage is billed at standard rates — review §3a EventID breakdown for reduction opportunities, or consider onboarding more servers to DfS P2 to expand the pool
M365 E5 / Defender XDR Ingestion Benefit
  • M365 E5 (or E5 Security, A5, F5, G5) provides 5 MB per user per day pooled data grant (offer page)
  • Grant = (number of E5 licenses) × 5 MB/day
  • Covers: Entra ID sign-in/audit logs, MCAS shadow IT, Purview info protection, M365 advanced hunting data (29 tables in Q17/Q17b)
  • Applied automatically — Free Benefit - M365 Defender Data Ingestion
  • Always-free (all Sentinel users): Azure Activity, Office 365 Audit Logs, Defender alerts

⚠️ Ask user for E5 license count — not discoverable from Sentinel telemetry.

Example: 500 E5 licenses → grant = 2.5 GB/day. If E5-eligible avg exceeds grant, overage billed at standard rates.

References:


Report Template

📄 Just-in-time loading: Read SKILL-report.md at the start of Phase 6 rendering. It contains:

  • Inline Chat Executive Summary template — Workspace at a Glance, Cost Waterfall, Detection Posture, Overall Assessment, Top 3 Recommendations
  • Markdown File Structure — Complete §1–§8 rendering rules, mandatory format requirements, column specifications, validation checks
  • Section-to-Scratchpad Mapping — Which scratchpad keys feed each report section

Load ONLY when entering Phase 6 — NOT during Phases 1–5. Combine with scratchpad data for rendering.


Post-Report Drill-Down Reference

📄 Just-in-time loading: Read SKILL-drilldown.md for full instructions when any of these are needed.

Available Drill-Down Patterns

Use these when the user asks follow-up questions after a report is generated (e.g., "which rules use EventID 8002?", "look up custom detection rules", "do any ASIM parsers depend on this table?").

PatternPurposeTool / MethodTrigger Phrases
1. EventID cross-refWhich analytic rules reference a specific EventID?az rest (Sentinel REST API) + JMESPath contains()"which rules use EventID X", "does any rule need this EventID"
2. Syslog facility/processWhich rules reference a Syslog facility, source, or process?az rest + JMESPath"which rules use sshd", "any rules for authpriv"
3. CSL vendor/activityWhich rules reference a CEF vendor, product, or activity?az rest + JMESPath"rules for Palo Alto TRAFFIC", "which rules use CommonSecurityLog"
4. Full rule query dumpExport all enabled rule queries for manual analysisaz rest → JSON file"export all rule queries", "build EventID dependency map"
5. ASIM parser verificationWhich ASIM parsers consume a table slated for migration?az rest + regex match for _Im_/_ASim_ patterns"ASIM dependency", "do parsers use this table"
6. Custom Detection rulesInventory CD rules via Graph API (query text, schedule, last run)PowerShell Invoke-MgGraphRequest (NOT Graph MCP — scope CustomDetection.Read.All unavailable via MCP)"custom detection rules", "CD rules", "lookup custom detections"

⚠️ Graph MCP limitation: The Graph MCP server returns 403 for the Custom Detection endpoint (/beta/security/rules/detectionRules). Always use Invoke-MgGraphRequest via PowerShell terminal. See SKILL-drilldown.md and Q9b-CustomDetectionRules.yaml for the exact endpoint and select fields.

Also in SKILL-drilldown.md
SectionContents
Known PitfallsUsage table batching, _SPLT_CL naming, case-sensitive custom tables, LogSeverity types, value-level vs table-level coverage confusion
Error HandlingCommon errors from az rest, Graph API, az monitor; graceful degradation for missing tables; re-running individual PS1 phases
CloudAppEvents AppendixCustom Detection management audit trail (EditCustomDetection events) — distinct from execution telemetry
Additional ReferencesMicrosoft Learn links for cost optimization, DCR configuration, data tiers, ASIM parsers

SVG Dashboard Generation

After a report is generated, the user can request a visual SVG dashboard.

Trigger phrases: "generate SVG dashboard", "create a visual dashboard", "visualize this report", "SVG from the report"

✅ DEFAULT: run the deterministic renderer (render_dashboard.py)

Do this first — do NOT hand-author the SVG. render_dashboard.py produces the manifest-driven 7-row dashboard non-interactively, parsing every value from the scratchpad + report + svg-widgets.yaml (no hardcoded run data). It is faster, deterministic, and produces a known-good layout. Run it:

python .github/skills/sentinel-ingestion-report/render_dashboard.py \
  --scratch temp/ingest_scratch_<ts>.md \
  --manifest .github/skills/sentinel-ingestion-report/svg-widgets.yaml \
  --report reports/sentinel/sentinel_ingestion_report_<label>_<ts>.md \
  --out reports/sentinel/sentinel_ingestion_report_<label>_<ts>_dashboard.svg

It reads the posture gauge, ingestion KPI cards, daily-volume line chart, cost waterfall, tier donut, top-tables / detection-coverage tables, and WoW anomaly + alert-producing-rule tables from the scratchpad, and the header metadata + Overall Assessment + ### 🎯 Top 3 Recommendations cards from the report (--report is optional — the assessment banner and recommendation cards degrade gracefully if absent). The alert-rule subheader lookback suffix ((7d)/(30d)) is matched by prefix, so any reporting window parses. Output is self-contained SVG with explicit fill on every <text>.

ActionStatus
Running render_dashboard.py when the user asks to visualize/generate a dashboard✅ REQUIRED (default path)
Hand-authoring the SVG via the svg-dashboard skill instead of running the script❌ PROHIBITED unless the user explicitly asks for a bespoke/custom layout the renderer can't produce
Fallback — bespoke/interactive dashboards (svg-dashboard skill)

Only use this path when the user explicitly wants a custom layout, different widgets, or styling the deterministic renderer doesn't support. Edit svg-widgets.yaml first if the change is layout/field-level — the renderer reads it at generation time, so many "customizations" don't require hand-authoring. The YAML manifest is the single source of truth for layout, widgets, field mappings, colors, and data source documentation.

Step 1:  Read svg-widgets.yaml (this skill's widget manifest)
Step 2:  Read .github/skills/svg-dashboard/SKILL.md (rendering rules — Manifest Mode)
Step 3:  Read the completed report file (data source)
Step 4:  Render SVG → save to reports/sentinel/{report_name}_dashboard.svg

© SCStelz, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 58 other files in .github/skills/sentinel-ingestion-report of SCStelz/security-investigator.

  • SKILL.md
  • Invoke-IngestionScan.ps1
  • SKILL-drilldown.md
  • SKILL-report.md
  • queries/phase1/Q1-UsageByDataType.yaml
  • queries/phase1/Q2-DailyIngestionTrend.yaml
  • queries/phase1/Q3-WorkspaceSummary.yaml
  • queries/phase2/Q4-SecurityEventByComputer.yaml
  • queries/phase2/Q5-SecurityEventByEventID.yaml
  • queries/phase2/Q6a-SyslogByHost.yaml
  • queries/phase2/Q6b-SyslogByFacilitySeverity.yaml
  • queries/phase2/Q6c-SyslogByProcess.yaml
  • queries/phase2/Q7-CSLByVendor.yaml
  • queries/phase2/Q8-CSLByActivity.yaml
  • queries/phase3/Q10-TableTierClassification.yaml
  • queries/phase3/Q10b-TierSummary.yaml
  • queries/phase3/Q9-AnalyticRuleInventory.yaml
  • … and 42 more

Open the folder on GitHubat commit 51e1385

Compare with similar skills

Sentinel Ingestion Report next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sentinel Ingestion Report compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sentinel Ingestion Report this skillSCStelz/security-investigator249—~16kAutomated safety check: PassMIT
Kibana Anomaly Detectionelastic/agent-skills592—~6kAutomated safety check: PassApache-2.0
Time Series Analytics Useropen-edge-platform/edge-ai-libraries171—~3.1kAutomated safety check: PassApache-2.0
Entra Agent Idmicrosoft/GitHub-Copilot-for-Azure2552 repos~4kAutomated safety check: PassMIT
Govmap APILiorVainer/data-israel130—~1.7kAutomated safety check: PassNone
Elasticsearch Anomaly Detection Explainerelastic/agent-skills592—~4.5kAutomated safety check: PassApache-2.0

Similar skills

  • Kibana Anomaly Detection

    elastic/agent-skills

    Official

    Elastic ML anomaly detection — investigation/RCA, score explanation, job lifecycle troubleshooting, and job operations.

    592 GitHub stars~6k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    171 GitHub stars~3.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Entra Agent Id

    microsoft/GitHub-Copilot-for-Azure

    Official

    Provision Microsoft Entra Agent Identity Blueprints, BlueprintPrincipals, and per-instance Agent Identities via Microsoft Graph, and configure OAuth 2.0 token exchange (fmipath, OBO, cross-tenant)…

    255 GitHub starsUsed in 2 repos~4k tokens
    Backend & APIsAuto-check passed
  • Govmap API

    LiorVainer/data-israel

    This skill should be used when the user asks about 'GovMap layers', 'GovMap API', 'map layers', 'entitiesByPoint', 'layer metadata', 'govmap endpoints', 'what layers are available', 'query map data…

    130 GitHub stars~1.7k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check passed
  • Official

    Explain Elasticsearch ML anomaly detection scores, model behavior, and result interpretation.

    592 GitHub stars~4.5k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Fabric CLI

    data-goblin/power-bi-agentic-development

    Expert guidance for the Fabric CLI (fab) and the Fabric and Power BI REST APIs: workspaces, items, lakehouses, notebooks, pipelines, semantic models, reports, capacities, OneLake, deployment and…

    1k GitHub stars~10k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from SCStelz/security-investigator

All 22 skills in this repo
  • Ca Policy Investigation

    SCStelz/security-investigator

    A skill your agent uses when asked to investigate Conditional Access policy changes, sign-in failures related to CA policies (error codes 53000, 50074, 530032), or suspected policy…

    249 GitHub stars~3.8k tokensUpdated 2 days ago
    Auto-check passed
  • Context Memory Review

    SCStelz/security-investigator

    Weekly review of an investigation tenant-context memory file against the most recent SOC scan reports (e.g.

    249 GitHub stars~3.7k tokensUpdated 2 days ago
    Auto-check passed
  • Heatmap Visualization

    SCStelz/security-investigator

    A skill your agent uses when asked to create heatmaps, visualize patterns over time, show activity grids, or display aggregated data in a matrix format.

    249 GitHub stars~3.4k tokensUpdated 2 days ago
    Auto-check passed
  • AI Agent Activity

    SCStelz/security-investigator

    Report/investigate RUNTIME ACTIVITY of AI agents (Agent 365 / Copilot Studio / M365 Copilot / Work IQ) — agents used, tools/connectors, channels, tokens, prompt/reply content, and Prompt Shield…

    249 GitHub stars~17k tokensUpdated 2 days ago
    Auto-check passed
  • AI Agent Posture

    SCStelz/security-investigator

    Audit or report on AI agent security posture across Copilot Studio, Microsoft 365 Copilot, Microsoft Foundry, and third-party agents.

    249 GitHub stars~21k tokensUpdated 2 days ago
    Auto-check passed
  • App Registration Posture

    SCStelz/security-investigator

    Audit Entra ID app registration and service principal security posture.

    249 GitHub stars~21k tokensUpdated 2 days ago
    Auto-check passed

Questions about Sentinel Ingestion Report

What does Sentinel Ingestion Report do?

Sentinel Ingestion Report — YAML-driven PowerShell pipeline gathers all data via az monitor/az rest/Graph API, writes a deterministic scratchpad, LLM renders the report. Sentinel Ingestion Report is an agent skill from SCStelz/security-investigator. Sentinel Ingestion Report — YAML-driven PowerShell pipeline gathers all data via az monitor/az rest/Graph API, writes a deterministic scratchpad, LLM renders the report.

When should I use Sentinel Ingestion Report?

Sentinel Ingestion Report fits situations like: tasks that involve Anomaly detection; tasks that involve REST APIs.

How do I install Sentinel Ingestion Report in Claude Code?

Run `npx skills add SCStelz/security-investigator --skill sentinel-ingestion-report -a claude-code`. Or copy the skill folder (.github/skills/sentinel-ingestion-report in SCStelz/security-investigator) into .claude/skills/sentinel-ingestion-report in your project. Claude Code loads it when a task matches its description.

How do I install Sentinel Ingestion Report in Codex?

Run `npx skills add SCStelz/security-investigator --skill sentinel-ingestion-report -a codex`. Or copy the skill folder (.github/skills/sentinel-ingestion-report in SCStelz/security-investigator) into .agents/skills/sentinel-ingestion-report in your project. Codex loads it when a task matches its description.

Can I use Sentinel Ingestion Report in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SCStelz/security-investigator --skill sentinel-ingestion-report -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sentinel-ingestion-report, .gemini/skills/sentinel-ingestion-report, .github/skills/sentinel-ingestion-report and .opencode/skills/sentinel-ingestion-report in your project.

What does Sentinel Ingestion Report need to run?

Going by SKILL.md and its folder, Sentinel Ingestion Report needs PowerShell for the scripts in its folder and the command-line tools its instructions call (az and python). Our summary lists: Python 3; PowerShell.

Does Sentinel Ingestion Report access the network?

SKILL.md names 4 domains. As links in the text: learn.microsoft.com, azure.microsoft.com, aka.ms and techcommunity.microsoft.com. This is read from the text; nothing was executed.

Is Sentinel Ingestion Report safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sentinel Ingestion Report use?

Sentinel Ingestion Report is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sentinel Ingestion Report use?

About 16k tokens (SKILL.md is roughly 65k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sentinel Ingestion Report?

Skills that share tags, products or a category with Sentinel Ingestion Report: Kibana Anomaly Detection (elastic/agent-skills, 592 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 171 stars), Entra Agent Id (microsoft/GitHub-Copilot-for-Azure, 255 stars) and Govmap API (LiorVainer/data-israel, 130 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sentinel Ingestion Report?

SCStelz (a GitHub user) maintains it in SCStelz/security-investigator, which has 249 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 8, 2026.

Source: SCStelz/security-investigator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.