Agent skill

Benchmark

by affaan-m in affaan-m/ECC

このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

MITAuto-check passed

Install Benchmark

skills CLI
$ npx skills add affaan-m/ECC --skill benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/ja-JP/skills/benchmark .claude/skills/benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark
GitHub stars
276k
Token cost
~412 tokens
SKILL.md length
26 words
Files
1
Skills in repo
673
Repo updated
First seen
Licence
MIT

At a glance

このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

  • SKILL.md covers 使用時期, 動作方法, 出力 and 統合
  • Calls bundle

What it does

Benchmark is an agent skill from affaan-m/ECC. このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します。

Its SKILL.md is about 410 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

Example prompts

  • “/benchmark”

Requirements

  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bundle

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark loads about 412 tokens when it runs. Until then it costs about 16 tokens; SKILL.md has 26 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~16
When it runs · the whole SKILL.md, loaded when a task matches
~412

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 26 words, ~412 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark/SKILL.md (or your agent's skills folder).
name
benchmark
description
このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します。
origin
ECC

ベンチマーク — パフォーマンスベースラインと回帰検出

使用時期

  • PR前後にパフォーマンスへの影響を測定
  • プロジェクトのパフォーマンスベースラインを設定
  • ユーザーが「遅く感じる」と報告したとき
  • ローンチ前 — パフォーマンスターゲットを満たしていることを確認
  • スタックを代替案と比較

動作方法

モード1:ページパフォーマンス

ブラウザMCPを介してリアルブラウザメトリクスを測定:

1. 各ターゲットURLに移動
2. Core Web Vitalsを測定:
   - LCP (Largest Contentful Paint) — ターゲット < 2.5s
   - CLS (Cumulative Layout Shift) — ターゲット < 0.1
   - INP (Interaction to Next Paint) — ターゲット < 200ms
   - FCP (First Contentful Paint) — ターゲット < 1.8s
   - TTFB (Time to First Byte) — ターゲット < 800ms
3. リソースサイズを測定:
   - 合計ページウェイト(ターゲット < 1MB)
   - JSバンドルサイズ(ターゲット < 200KBgzipped)
   - CSSサイズ
   - 画像ウェイト
   - サードパーティスクリプトウェイト
4. ネットワークリクエストをカウント
5. レンダリングブロッキングリソースをチェック
モード2:APIパフォーマンス

APIエンドポイントをベンチマーク:

1. 各エンドポイントに100回ヒット
2. 測定:p50、p95、p99レイテンシ
3. トラック:レスポンスサイズ、ステータスコード
4. ロード下でテスト:10個の同時リクエスト
5. SLAターゲットと比較
モード3:ビルドパフォーマンス

開発フィードバックループを測定:

1. コールドビルド時間
2. ホットリロード時間(HMR)
3. テストスイート期間
4. TypeScriptチェック時間
5. Lint時間
6. Dockerビルド時間
モード4:前後の比較

変更前後に実行して影響を測定:

/benchmark baseline    # 現在のメトリクスを保存
# ... 変更を加える ...
/benchmark compare     # ベースラインと比較

出力:

| Metric | Before | After | Delta | Verdict |
|--------|--------|-------|-------|---------|
| LCP | 1.2s | 1.4s | +200ms | WARNING: WARN |
| Bundle | 180KB | 175KB | -5KB | ✓ BETTER |
| Build | 12s | 14s | +2s | WARNING: WARN |

出力

.ecc/benchmarks/にJSONとしてベースラインを保存。Gitで追跡されるため、チームはベースラインを共有します。

統合

  • CI:すべてのPRで/benchmark compareを実行
  • /canary-watchとペアリングしてデプロイ後の監視
  • /browser-qaとペアリングして完全な出荷前チェックリスト

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/ja-JP/skills/benchmark of affaan-m/ECC.

Open the folder on GitHubat commit ef648e0

Compare with similar skills

Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark this skillaffaan-m/ECC276k—~412Automated safety check: PassMIT
Benchmarkandroidx/androidx6.1k—~1.1kAutomated safety check: PassApache-2.0
Benchmarksamchon/typia5.9k—~1.2kAutomated safety check: PassMIT
Gstack Performance Benchmarkgarrytan/gstack136k—~7.3kAutomated safety check: NotesMIT
Benchmarkzalando/skipper3.3k—~304Automated safety check: PassMIT
Agent Benchmark Suiteruvnet/ruflo74k2 repos~4.9kAutomated safety check: PassMIT

Similar skills

  • Benchmark

    androidx/androidx

    Benchmarking and improving the performance of Jetpack Compose.

    6.1k GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed
  • Benchmark

    samchon/typia

    Defines typia benchmark fixture integrity, result reporting, and publication safeguards.

    5.9k GitHub stars~1.2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Establishes page load, Core Web Vitals and resource-size baselines, then compares before and after on every pull request to track performance trends over time.

    136k GitHub stars~7.3k tokensUpdated today
    Frontend & DesignAuto-check: notes
  • Benchmark

    zalando/skipper

    Create or modify hot-path code or optimizations should be validated by a Go Benchmark

    3.3k GitHub stars~304 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Agent skill for benchmark-suite - invoke with $agent-benchmark-suite

    74k GitHub starsUsed in 2 repos~4.9k tokens
    Agent WorkflowsAuto-check passed
  • Agent skill for performance-benchmarker - invoke with $agent-performance-benchmarker

    74k GitHub starsUsed in 2 repos~6.8k tokens
    Auto-check passed

More from affaan-m/ECC

All 673 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    276k GitHub starsUsed in 5 repos~1.9k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    276k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    276k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    276k GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    276k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed
  • Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.

    276k GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check passed

Questions about Benchmark

What does Benchmark do?

このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC. Benchmark is an agent skill from affaan-m/ECC.

How do I install Benchmark in Claude Code?

Run `npx skills add affaan-m/ECC --skill benchmark -a claude-code`. Or copy the skill folder (docs/ja-JP/skills/benchmark in affaan-m/ECC) into .claude/skills/benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark in Codex?

Run `npx skills add affaan-m/ECC --skill benchmark -a codex`. Or copy the skill folder (docs/ja-JP/skills/benchmark in affaan-m/ECC) into .agents/skills/benchmark in your project. Codex loads it when a task matches its description.

Can I use Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark, .gemini/skills/benchmark, .github/skills/benchmark and .opencode/skills/benchmark in your project.

What does Benchmark need to run?

Going by SKILL.md and its folder, Benchmark needs the command-line tools its instructions call (bundle). Our summary lists: Docker.

Does Benchmark access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark use?

Benchmark is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark use?

About 412 tokens (SKILL.md is roughly 1.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark?

Skills that share tags, products or a category with Benchmark: Benchmark (androidx/androidx, 6.1k stars), Benchmark (samchon/typia, 5.9k stars), Gstack Performance Benchmark (garrytan/gstack, 136k stars) and Benchmark (zalando/skipper, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 275,546 GitHub stars. The repository holds 673 skills in this directory. The repository was last updated on October 5, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.