Analyze skill effectiveness across sessions. An agent skill from oliver-kriska/claude-elixir-phoenix.

MITAuto-check passed

Install Skill Monitor

skills CLI
$ npx skills add oliver-kriska/claude-elixir-phoenix --skill skill-monitor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oliver-kriska/claude-elixir-phoenix skill-monitor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oliver-kriska/claude-elixir-phoenix.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/skill-monitor .claude/skills/skill-monitor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-monitor
GitHub stars
565
Token cost
~2.2k tokens
SKILL.md length
599 words
Files
3 (incl. references)
Skills in repo
109
Repo updated
First seen
Licence
MIT

At a glance

Analyze skill effectiveness across sessions. An agent skill from oliver-kriska/claude-elixir-phoenix.

  • Works in 6 steps: Parse Arguments → Load Metrics → Compute Per-Skill Aggregates → …
  • SKILL.md covers Requirements, Usage, What Main Context Does and Iron Laws, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Skill Monitor is an agent skill from oliver-kriska/claude-elixir-phoenix. Analyze skill effectiveness across sessions. Computes per-skill metrics (action rate, friction, outcomes), identifies degrading skills, and generates improvement recommendations. Requires session-scan data in metrics.jsonl.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/effectiveness-metrics.md` and `references/improvement-template.md`).

The repository describes itself as: Claude Code plugin for Elixir/Phoenix/LiveView — 26 specialist agents, Iron Laws enforcement, and Tidewave MCP integration. Plan features with parallel research agents, execute… The licence is MIT.

Example prompts

  • “/skill-monitor”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Parse Arguments
  2. Load Metrics
  3. Compute Per-Skill Aggregates
  4. Display Dashboard
  5. Improvement Mode (--improve)
  6. Write Output

What it can do on your machine

Read from SKILL.md and the folder at commit 9767a82. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Monitor loads about 2.2k tokens when it runs, and up to ~5.5k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 599 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from oliver-kriska/claude-elixir-phoenix at commit 9767a82, republished under its MIT licence (© oliver-kriska). 599 words, ~2,198 tokens.

Download SKILL.mdSave it as .claude/skills/skill-monitor/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
skill-monitor
description
Analyze skill effectiveness across sessions. Computes per-skill metrics (action rate, friction, outcomes), identifies degrading skills, and generates improvement recommendations. Requires session-scan data in metrics.jsonl.
argument-hint
[--skill NAME] [--improve] [--window 7d|30d|all]
disable-model-invocation
true

Skill Monitor

Closed-loop skill effectiveness monitoring. Reads session metrics, computes per-skill signals, identifies what's working and what needs improvement.

Inspired by the deploy-monitor-evaluate-improve feedback loop: skills get better over time instead of staying static.

Requirements

Requires .claude/session-metrics/metrics.jsonl from /session-scan. If no data: suggest running /session-scan first.

Usage

/skill-monitor                     # Dashboard: all skills
/skill-monitor --skill review      # Deep-dive on one skill
/skill-monitor --improve           # Generate improvement recommendations
/skill-monitor --window 30d        # Change comparison window (default: 7d)

What Main Context Does

Step 1: Parse Arguments

Extract from $ARGUMENTS:

  • --skill NAME: Focus on one skill (e.g., review, plan, investigate)
  • --improve: Spawn analysis agent for improvement recommendations
  • --window PERIOD: Comparison window (7d, 30d, all; default: 7d)
Step 2: Load Metrics

Read .claude/session-metrics/metrics.jsonl. For each entry, extract the skill_effectiveness field (added by compute-metrics.py v2).

Filter by window period. Count sessions with and without skill usage.

If no skill_effectiveness data exists in metrics: "Metrics were computed before skill tracking was added. Run /session-scan --rescan to recompute."

OTel invocation_trigger (CC v2.1.126+): when compute-metrics.py ingests claude_code.skill_activated events, each invocation carries an invocation_trigger of "user-slash", "claude-proactive", or "nested-skill". If absent (older sessions), default to "unknown" — do NOT assume "user-slash".

Step 3: Compute Per-Skill Aggregates

For each skill found across all sessions, aggregate:

| Metric                  | Computation                                    |
|-------------------------|------------------------------------------------|
| Total invocations       | Sum of invocation_count across sessions        |
| Sessions used in        | Count of sessions containing this skill        |
| Action rate             | Weighted avg of per-session action_rate         |
| Avg post-errors         | Weighted avg of avg_post_errors                |
| Avg post-corrections    | Weighted avg of avg_post_corrections           |
| Outcome distribution    | Count of effective/friction/no_action/mixed    |
| Effectiveness score     | action_rate - (0.3 * avg_post_corrections)     |
| Adjusted score          | For analysis/check skills, use lower thresholds |
| Trigger distribution    | Counts of user-slash / claude-proactive / nested-skill / unknown |
| Proactive trigger rate  | claude-proactive / (user-slash + claude-proactive + nested-skill) |
| Auto-load gap           | Skills with 0 claude-proactive invocations across window |

Auto-load gap detection (CC v2.1.126+): Skills with auto-loaded behavior in their description (i.e., not disable-model-invocation: true) are EXPECTED to fire as claude-proactive. A skill that is ONLY ever invoked via user-slash is failing its description's routing intent. Flag any auto-loadable skill where proactive_trigger_rate == 0 over the window. This is the structural answer to the "zero skill auto-loading" gap from the 137-session analysis (see MEMORY.md). Confidence floor: only flag if total invocations >= 5 in window.

Skill type weighting: Analysis and check skills (verify, triage, perf, boundaries, pr-review, audit) have low action rates BY DESIGN — their success is "found issues" or "confirmed things pass". Apply adjusted thresholds:

Skill TypeFlag ThresholdExpected Action Rate
Execution (work, quick, full)< 0.5> 0.7
Analysis (perf, boundaries, audit, pr-review)< 0.30.3-0.5
Check (verify, triage)< 0.10.0-0.3
Knowledge (compound, learn, brief)< 0.5> 0.5

Also compute baseline friction (avg friction of sessions WITHOUT any skill usage) vs skill friction (avg friction of sessions WITH skill usage). Delta = skill_friction - baseline_friction. Negative delta = skills reduce friction (good).

Show full SKILL.md (257 more words)Show less
Step 4: Display Dashboard

Dashboard mode (no --skill):

## Skill Effectiveness Dashboard (last {window})

Baseline friction (no skills): 0.32 | With skills: 0.18 | Delta: -0.14

| Skill           | Uses | Sessions | Slash/Proactive/Nested | Action% | Errors | Corr | Outcome   | Score |
|-----------------|------|----------|------------------------|---------|--------|------|-----------|-------|
| /phx:review     | 12   | 8        |    8 /  3 /  1         | 92%     | 0.5    | 0.1  | effective | 0.89  |
| /phx:plan       | 9    | 7        |    9 /  0 /  0         | 100%    | 0.2    | 0.0  | effective | 1.00  |
| /phx:investigate| 5    | 5        |    5 /  0 /  0         | 80%     | 1.2    | 0.4  | mixed     | 0.68  |

Skills needing attention:
- /phx:investigate (high post-errors)
- /phx:plan (auto-load gap — 0/9 proactive; description not routing)

Flag skills using type-adjusted thresholds (see weighting table above). Also flag if avg_post_corrections > 1 or outcome is predominantly "friction". Also flag auto-load gap: auto-loadable skills (without disable-model-invocation: true) with proactive_trigger_rate == 0 and total invocations >= 5. This is a description/routing problem — the skill exists but Claude isn't loading it on its own.

When displaying flagged skills, note if the flag is "expected" for the skill type (e.g., verify at 0.24 is normal for a check skill).

Skill deep-dive (--skill NAME):

Show per-session breakdown for that skill, including session IDs, dates, individual outcome signals, AND invocation_trigger per invocation. If a skill is dominated by user-slash triggers, surface which 1-3 description keywords might unlock proactive routing — cross-reference against the skill's current description in plugins/elixir-phoenix/skills/{name}/SKILL.md. If session reports exist in .claude/session-analysis/, reference them.

Step 5: Improvement Mode (--improve)

Spawn skill-effectiveness-analyzer agent:

Agent(subagent_type="skill-effectiveness-analyzer", model="sonnet", prompt="""
Analyze skill effectiveness data and recommend improvements.

Metrics data: {aggregated_metrics_json}

Sessions with friction outcomes: {session_ids}

For each underperforming skill:
1. Identify failure patterns from outcome signals
2. Propose specific skill/agent changes
3. Suggest new Iron Laws if patterns are systematic

Write recommendations to: .claude/skill-metrics/recommendations-{date}.md
""")
Step 6: Write Output

Write aggregated metrics to .claude/skill-metrics/dashboard-{date}.json:

json
{
  "computed_at": "2026-03-03T14:00:00Z",
  "window": "7d",
  "baseline_friction": 0.32,
  "skill_friction": 0.18,
  "friction_delta": -0.14,
  "skills": {
    "/phx:plan": {
      "invocations": 9,
      "trigger_distribution": {
        "user-slash": 9,
        "claude-proactive": 0,
        "nested-skill": 0,
        "unknown": 0
      },
      "proactive_trigger_rate": 0.0,
      "auto_load_gap": true
    }
  },
  "flagged_skills": ["investigate", "plan:auto-load-gap"]
}

Append-only: never modify previous dashboard files.

Iron Laws

  1. NEVER modify metrics.jsonl — read-only from this skill
  2. Baseline comparison is mandatory — raw numbers without baseline are meaningless
  3. Flag, don't judge — surface data, let the human decide what to fix
  4. Evidence tags on recommendations — every suggestion needs session citations
  5. Trigger source must not be inferred — only treat invocations as user-slash / claude-proactive / nested-skill when the OTel invocation_trigger attribute is present (CC v2.1.126+). Older sessions use "unknown"; never silently bucket them as user-slash — it would hide the auto-load gap.

Integration

/session-scan → metrics.jsonl (with skill_effectiveness)
       ↓
/skill-monitor → dashboard + flagged skills
       ↓
/skill-monitor --improve → recommendations
       ↓
Developer updates skills/agents → deploy → repeat

References

  • references/effectiveness-metrics.md — Full metrics schema and evaluation criteria
  • references/improvement-template.md — Template for improvement recommendations

© oliver-kriska, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .claude/skills/skill-monitor of oliver-kriska/claude-elixir-phoenix.

  • SKILL.md
  • references/effectiveness-metrics.md
  • references/improvement-template.md

Open the folder on GitHubat commit 9767a82

Compare with similar skills

Skill Monitor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Monitor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Monitor this skilloliver-kriska/claude-elixir-phoenix565—~2.2kAutomated safety check: PassMIT
Cloud Monitoring Metric Selectiongoogle/skills21k—~2.4kAutomated safety check: PassApache-2.0
Gke AI Troubleshooting Tpu Metrics Monitoringgoogle/skills21k—~1.9kAutomated safety check: PassApache-2.0
Agent Session Monitorhigress-group/higress9.5k—~3.3kAutomated safety check: PassApache-2.0
Session Metricscentminmod/my-claude-code-setup2.7k—~8kAutomated safety check: PassMIT
CI Metricspytorch/pytorch104k—~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Official

    Retrieve, query, and identify relevant Cloud Monitoring metric descriptors on Google Cloud for a service or resource (such as Compute Engine, Spanner, BigQuery, Cloud Run, Cloud SQL, Pub/Sub, Cloud…

    21k GitHub stars~2.4k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.

    21k GitHub stars~1.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Agent Session Monitor

    higress-group/higress

    Real-time agent conversation monitoring - monitors Higress access logs, aggregates conversations by session, tracks token usage.

    9.5k GitHub stars~3.3k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Session Metrics

    centminmod/my-claude-code-setup

    Tally Claude Code session token usage and cost estimates from the raw JSONL conversation log.

    2.7k GitHub stars~8k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • CI Metrics

    pytorch/pytorch

    Query PyTorch CI, GitHub Actions, HUD, Grafana, and infrastructure metrics.

    104k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Audit Session Metrics

    centminmod/my-claude-code-setup

    Audit a session-metrics JSON export for token-usage waste and produce a plain-English findings report.

    2.7k GitHub stars~2.1k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed

More from oliver-kriska/claude-elixir-phoenix

All 109 skills in this repo
  • Codex Ab

    oliver-kriska/claude-elixir-phoenix

    Run an A/B codex review experiment — holistic codex review vs 3 focused dimension passes (security, ecto, liveview) on the branch diff, classify findings, report a panel-value verdict.

    565 GitHub stars~977 tokensUpdated 5 days ago
    Auto-check passed
  • Audit

    oliver-kriska/claude-elixir-phoenix

    Project health audit and health check — architecture, performance, tests, dependencies, code quality.

    565 GitHub stars~2k tokensUpdated 5 days ago
    Auto-check passed
  • Compound Docs

    oliver-kriska/claude-elixir-phoenix

    Searchable Elixir/Phoenix/Ecto solution documentation system with; Use when consulting past solutions…

    565 GitHub stars~528 tokensUpdated 5 days ago
    Auto-check passed
  • Compound Docs

    oliver-kriska/claude-elixir-phoenix

    Searchable Elixir/Phoenix/Ecto solution documentation system with YAML frontmatter.

    565 GitHub stars~547 tokensUpdated 5 days ago
    Auto-check passed
  • Deploy

    oliver-kriska/claude-elixir-phoenix

    Elixir/Phoenix deployment patterns — Dockerfile, fly.toml, runtime.exs, mix release, rel/ overlays.

    565 GitHub stars~1.1k tokensUpdated 5 days ago
    Auto-check passed
  • Deps Update

    oliver-kriska/claude-elixir-phoenix

    Bump outdated Hex deps — inventory, snapshot changelogs, update, fix breaks, split reviewable PRs (patches bundled, majors solo).

    565 GitHub stars~1.4k tokensUpdated 5 days ago
    Auto-check passed

Questions about Skill Monitor

What does Skill Monitor do?

Analyze skill effectiveness across sessions. An agent skill from oliver-kriska/claude-elixir-phoenix. Skill Monitor is an agent skill from oliver-kriska/claude-elixir-phoenix. Analyze skill effectiveness across sessions.

How do I install Skill Monitor in Claude Code?

Run `npx skills add oliver-kriska/claude-elixir-phoenix --skill skill-monitor -a claude-code`. Or copy the skill folder (.claude/skills/skill-monitor in oliver-kriska/claude-elixir-phoenix) into .claude/skills/skill-monitor in your project. Claude Code loads it when a task matches its description.

How do I install Skill Monitor in Codex?

Run `npx skills add oliver-kriska/claude-elixir-phoenix --skill skill-monitor -a codex`. Or copy the skill folder (.claude/skills/skill-monitor in oliver-kriska/claude-elixir-phoenix) into .agents/skills/skill-monitor in your project. Codex loads it when a task matches its description.

Can I use Skill Monitor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oliver-kriska/claude-elixir-phoenix --skill skill-monitor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-monitor, .gemini/skills/skill-monitor, .github/skills/skill-monitor and .opencode/skills/skill-monitor in your project.

What does Skill Monitor need to run?

SKILL.md names no scripts, command-line tools or credentials: Skill Monitor is instructions for the agent only.

Does Skill Monitor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skill Monitor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skill Monitor use?

Skill Monitor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Monitor use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.

What are the alternatives to Skill Monitor?

Skills that share tags, products or a category with Skill Monitor: Cloud Monitoring Metric Selection (google/skills, 21k stars), Gke AI Troubleshooting Tpu Metrics Monitoring (google/skills, 21k stars), Agent Session Monitor (higress-group/higress, 9.5k stars) and Session Metrics (centminmod/my-claude-code-setup, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Monitor?

oliver-kriska (a GitHub user) maintains it in oliver-kriska/claude-elixir-phoenix, which has 565 GitHub stars. The repository holds 109 skills in this directory. The repository was last updated on October 5, 2026.

Source: oliver-kriska/claude-elixir-phoenix on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.