Agent skill

Monte Carlo Performance Diagnosis

by sickn33 in sickn33/agentic-awesome-skills

Diagnoses pipeline performance issues -- slow jobs, expensive queries, latency trends -- using Monte Carlo's cross-platform observability.

Apache-2.0Auto-check passedData & Analytics

Install Monte Carlo Performance Diagnosis

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill monte-carlo-performance-diagnosis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills monte-carlo-performance-diagnosis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/monte-carlo-performance-diagnosis .claude/skills/monte-carlo-performance-diagnosis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
monte-carlo-performance-diagnosis
GitHub stars
47k
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
995 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses pipeline performance issues -- slow jobs, expensive queries, latency trends -- using Monte Carlo's cross-platform observability.

  • Works in 5 steps: Identify the scope → Tier 1 -- Discovery → Bridge -- Job to Tables → …
  • Tasks that involve Root cause analysis
  • SKILL.md covers When to activate this skill, When NOT to activate this skill, Prerequisites and Workflow, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Monte Carlo Performance Diagnosis is an agent skill from sickn33/agentic-awesome-skills. Diagnoses pipeline performance issues -- slow jobs, expensive queries, latency trends -- using Monte Carlo's cross-platform observability. Uses a tiered investigation approach: discover problems, bridge to affected tables, then drill into root causes.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Root cause analysis, Observability and Data pipelines and ETL. It works with Model Context Protocol. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Root cause analysis
  • Tasks that involve Observability
  • Tasks that involve Data pipelines and ETL

Example prompts

  • “Use the monte-carlo-performance-diagnosis skill to diagnose pipeline performance issues -- slow jobs, expensive queries, latency trends -- using…”
  • “/monte-carlo-performance-diagnosis”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Identify the scope
  2. Tier 1 -- Discovery
  3. Bridge -- Job to Tables
  4. Tier 2 -- Diagnosis
  5. Present findings

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Monte Carlo Performance Diagnosis loads about 2k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 995 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its Apache-2.0 licence (© sickn33). 995 words, ~2,042 tokens.

Download SKILL.mdSave it as .claude/skills/monte-carlo-performance-diagnosis/SKILL.md (or your agent's skills folder).
name
monte-carlo-performance-diagnosis
description
Diagnoses pipeline performance issues -- slow jobs, expensive queries, latency trends -- using Monte Carlo's cross-platform observability. Uses a tiered investigation approach: discover problems, bridge to affected tables, then drill into root causes.
risk
critical
source
https://github.com/monte-carlo-data/mc-agent-toolkit/tree/main/skills/performance-diagnosis
source_repo
monte-carlo-data/mc-agent-toolkit
source_type
community
date_added
2026-07-01
license
Apache-2.0
license_source
https://github.com/monte-carlo-data/mc-agent-toolkit/blob/main/LICENSE

Monte Carlo Performance Diagnosis Skill

This skill helps diagnose data pipeline performance issues using Monte Carlo's cross-platform observability data. It works across Airflow, dbt, Databricks, and warehouse query engines to find bottlenecks, detect regressions, and identify root causes.

Monte Carlo tool routing (required): Always call Monte Carlo MCP tools through this plugin's bundled server, whose fully-qualified tool names are mcp__plugin_mc-agent-toolkit_monte-carlo-mcp__<tool> (e.g. mcp__plugin_mc-agent-toolkit_monte-carlo-mcp__get_alerts). Bare tool names used in this skill (get_alerts, search, get_table, …) refer to that bundled server. If the session also has a separately-configured monte-carlo-mcp server, do not route to it — it may point at a different endpoint or credentials.

Reference files live next to this skill file. Use the Read tool (not MCP resources) to access them:

  • Tiered investigation approach: references/investigation-tiers.md (relative to this file)
  • Query analysis patterns: references/query-analysis.md (relative to this file)

When to activate this skill

Activate when the user:

  • Asks about slow pipelines, jobs, or queries
  • Wants to find expensive or costly queries
  • Mentions performance regressions or degradation
  • Asks "why is this pipeline slow?" or "what's using the most compute?"
  • Wants to compare performance over time or find bottleneck tasks
  • Asks about failed or futile query patterns

When NOT to activate this skill

Do not activate when the user is:

  • Investigating data quality issues (use the prevent skill)
  • Looking at storage costs (use the storage-cost-analysis skill)
  • Creating monitors (use the monitoring-advisor skill)
  • Just querying data or exploring table contents

Prerequisites

The following MCP tools must be available (connect to Monte Carlo's MCP server):

Discovery tools (Tier 1):

  • get_jobs_performance -- find slow/failing jobs across Airflow, dbt, Databricks
  • get_top_slow_queries -- find slowest query groups by total runtime

Bridge tool:

  • get_tables_for_job -- convert job MCONs to table MCONs

Diagnosis tools (Tier 2):

  • get_tasks_performance -- drill into a job's individual tasks
  • get_change_timeline -- unified timeline of query changes, volume shifts, Airflow/dbt failures
  • get_query_rca -- root cause analysis for failed/futile queries
  • get_query_latency_distribution -- latency trend over time
  • get_asset_lineage -- trace upstream/downstream impact

Supporting tools:

  • get_warehouses -- list available warehouses

Workflow

Step 1: Identify the scope

Determine what the user wants to investigate:

  • Specific job/pipeline: User mentions a job name or pipeline
  • Specific table: User mentions a table that's slow to update
  • General discovery: User wants to find what's slow

Call get_warehouses to list available warehouses. Match the user's context to a warehouse.

Step 2: Tier 1 -- Discovery

If you don't have specific MCONs to investigate, start with discovery:

  1. Find slow jobs: Call get_jobs_performance with optional integration_type filter (AIRFLOW, DATABRICKS, DBT) if the user specifies a platform.

    • Results include: job name, average duration, trend (7-day), run count, failure rate
    • Look for: high avgDuration, negative runDurationTrend7d, high failure rates
  2. Find expensive queries: Call get_top_slow_queries with optional warehouse_id and query_type ("read" for SELECTs, "write" for INSERT/CREATE/MERGE).

    • Results include: query hash, total runtime, average runtime, run count
    • Look for: queries with high total runtime or high individual execution time

Present the top findings to the user before drilling deeper. A typical investigation needs only 3-7 tool calls.

If both discovery tools return no results: Tell the user no performance issues were found in the current time window. Suggest broadening the scope (different warehouse, longer time range, or a different platform filter).

Step 3: Bridge -- Job to Tables

After Tier 1 identifies problematic jobs, convert to table MCONs:

Call get_tables_for_job(job_mcon=..., integration_type=...) using the integration_type from the job performance results.

This gives you the table MCONs needed for Tier 2 investigation.

Show full SKILL.md (438 more words)Show less
Step 4: Tier 2 -- Diagnosis

Now drill into root causes using the MCONs from discovery or the bridge:

  1. Task bottleneck: Call get_tasks_performance to find which specific task in a job is the bottleneck.

  2. What changed? Call get_change_timeline -- this is your most powerful tool. It returns a unified timeline of:

    • Query text changes (schema modifications, new JOINs, filter changes)
    • Volume shifts (row count spikes/drops)
    • Airflow task failures
    • dbt model failures All in one call. Look for correlations: "query changed on day X, runtime doubled on day X+1."
  3. Why are queries failing? Call get_query_rca to get root cause analysis:

    • Failed queries: errors, timeouts, permission issues
    • Futile queries: queries that run but produce no useful output
    • Patterns are pre-computed -- the tool groups failures by cause
  4. Is latency degrading? Call get_query_latency_distribution to see the trend:

    • Compare p50 vs p95 -- if p95 >> p50 (>5x), the problem is outlier queries
    • Look for step-changes in latency (sudden increase = regression)
    • For step-change / regression-time-localization use cases, pass bucket="1h". The default downsamples to daily on windows ≥ 3 days, which hides hour-level steps.
  5. Trace impact: Call get_asset_lineage with direction="DOWNSTREAM" to see what's affected by a slow table, or direction="UPSTREAM" to find what feeds it.

Step 5: Present findings

Structure your response as:

  1. Problem summary: What's slow and by how much (with exact numbers from tools)
  2. Root cause: What changed or what's causing the issue
  3. Impact: What downstream systems are affected
  4. Recommendations: Specific actions to fix the issue
Important rules
  • Quote tool numbers exactly. If a tool returns "1282 runs, avg 22.5s", say exactly that. Never round, estimate, or fabricate numbers.
  • Always compare to baselines. Use 7-day trend data (runDurationTrend7d) to distinguish regressions from normal variance. Flag if trend data has less than 0.1 confidence.
  • Stop when you have a root cause. 3-7 tool calls is typical. More than 10 means you're over-investigating.
  • Read vs write queries: When the user asks about "reads" or "read queries", filter with query_type="read". When they ask about "writes", use query_type="write". Do NOT mix them.
  • Never expose MCONs, UUIDs, or internal identifiers to the user. Use human-readable names.
  • Cross-platform: This skill works across Airflow, dbt, and Databricks. Note which platform each finding comes from.

Example

User request:

Diagnose why this pipeline became slow, identify the bottleneck from the available telemetry, and propose the smallest verified fix.

Limitations

  • Use this skill only when the task clearly matches its upstream source and local project context.
  • Verify commands, generated code, dependencies, credentials, and external service behavior before applying changes.
  • Do not treat examples as a substitute for environment-specific tests, security review, or user approval for destructive or costly actions.

© sickn33, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/monte-carlo-performance-diagnosis of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Monte Carlo Performance Diagnosis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Monte Carlo Performance Diagnosis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Monte Carlo Performance Diagnosis this skillsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassApache-2.0
Upgrading Mwaa Environmentsaws/agent-toolkit-for-aws2.8k—~7.3kAutomated safety check: PassApache-2.0
Agent Observability Trace Rcadatadog-labs/agent-skills177—~10kAutomated safety check: PassMIT
Kubernetes Network Root Cause Analysiskubeshark/kubeshark12k—~5.3kAutomated safety check: PassApache-2.0
UModel Root Cause Analysisalibaba/UnifiedModel415—~1.9kAutomated safety check: PassCustom licence
Dagster Expertdagster-io/skills211—~2kAutomated safety check: PassApache-2.0

Similar skills

  • Upgrading Mwaa Environments

    aws/agent-toolkit-for-aws

    Official

    Upgrades an MWAA environment to a newer Airflow version — within 2.x, within 3.x, or across the 2.x-to-3.x boundary.

    2.8k GitHub stars~7.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Agent Observability Trace Rca

    datadog-labs/agent-skills

    Root cause analysis on production LLM traces. An agent skill from datadog-labs/agent-skills.

    177 GitHub stars~10k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.

    12k GitHub stars~5.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • UModel Root Cause Analysis

    alibaba/UnifiedModel

    Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.

    415 GitHub stars~1.9k tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed
  • Dagster Expert

    dagster-io/skills

    Official

    Expert guidance for working with Dagster and the dg CLI. An agent skill from dagster-io/skills.

    211 GitHub stars~2k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Erd Studio Setup

    liam-machine/erd-studio

    Friendly, step-by-step setup for ERD Studio in an existing dbt project, for people who may be new to dbt or data modelling.

    165 GitHub stars~8.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Monte Carlo Performance Diagnosis

What does Monte Carlo Performance Diagnosis do?

Diagnoses pipeline performance issues -- slow jobs, expensive queries, latency trends -- using Monte Carlo's cross-platform observability. Monte Carlo Performance Diagnosis is an agent skill from sickn33/agentic-awesome-skills. Diagnoses pipeline performance issues -- slow jobs, expensive queries, latency trends -- using Monte Carlo's cross-platform observability.

When should I use Monte Carlo Performance Diagnosis?

Monte Carlo Performance Diagnosis fits situations like: tasks that involve Root cause analysis; tasks that involve Observability; tasks that involve Data pipelines and ETL.

How do I install Monte Carlo Performance Diagnosis in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill monte-carlo-performance-diagnosis -a claude-code`. Or copy the skill folder (skills/monte-carlo-performance-diagnosis in sickn33/agentic-awesome-skills) into .claude/skills/monte-carlo-performance-diagnosis in your project. Claude Code loads it when a task matches its description.

How do I install Monte Carlo Performance Diagnosis in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill monte-carlo-performance-diagnosis -a codex`. Or copy the skill folder (skills/monte-carlo-performance-diagnosis in sickn33/agentic-awesome-skills) into .agents/skills/monte-carlo-performance-diagnosis in your project. Codex loads it when a task matches its description.

Can I use Monte Carlo Performance Diagnosis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill monte-carlo-performance-diagnosis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monte-carlo-performance-diagnosis, .gemini/skills/monte-carlo-performance-diagnosis, .github/skills/monte-carlo-performance-diagnosis and .opencode/skills/monte-carlo-performance-diagnosis in your project.

What does Monte Carlo Performance Diagnosis need to run?

SKILL.md names no scripts, command-line tools or credentials: Monte Carlo Performance Diagnosis is instructions for the agent only.

Does Monte Carlo Performance Diagnosis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Monte Carlo Performance Diagnosis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Monte Carlo Performance Diagnosis use?

Monte Carlo Performance Diagnosis is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Monte Carlo Performance Diagnosis use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Monte Carlo Performance Diagnosis?

Skills that share tags, products or a category with Monte Carlo Performance Diagnosis: Upgrading Mwaa Environments (aws/agent-toolkit-for-aws, 2.8k stars), Agent Observability Trace Rca (datadog-labs/agent-skills, 177 stars), Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars) and UModel Root Cause Analysis (alibaba/UnifiedModel, 415 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Monte Carlo Performance Diagnosis?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.