Agent skill

Benchmark

by mr-daedalium in mr-daedalium/ostack-saas

Performance regression detection using the browse daemon. An agent skill from mr-daedalium/ostack-saas.

MITAuto-check: notesFrontend & Design

Install Benchmark

skills CLI
$ npx skills add mr-daedalium/ostack-saas --skill benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mr-daedalium/ostack-saas benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mr-daedalium/ostack-saas.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmark .claude/skills/benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark
GitHub stars
114
Used in
1 other repo
Token cost
~5k tokens
SKILL.md length
1,787 words
Files
2
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Performance regression detection using the browse daemon. An agent skill from mr-daedalium/ostack-saas.

  • Works in 9 steps: Setup → Page Discovery → Performance Data Collection → …
  • Tasks that involve Web performance
  • SKILL.md covers Preamble (run first), AskUserQuestion Format, Conviction and Completeness —… and Repo Ownership Mode — See…, plus 10 more sections
  • Calls git, gh and codex; reaches bun.sh

What it does

Benchmark is an agent skill from mr-daedalium/ostack-saas. Performance regression detection using the browse daemon. Establishes baselines for page load times, Core Web Vitals, and resource sizes. Compares before/after on every PR. Tracks performance trends over time. Use when: "performance", "benchmark", "page speed", "lighthouse", "web vitals", "bundle size", "load time".

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Frontend & Design, covering Web performance. The repository describes itself as: ostack — AI-powered engineering team for Claude Code. Fork of gstack by Garry Tan. The licence is MIT.

When your agent uses it

  • Tasks that involve Web performance

Example prompts

  • “performance”
  • “benchmark”
  • “page speed”
  • “/benchmark”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Read, Write, Glob, AskUserQuestion

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Setup
  2. Page Discovery
  3. Performance Data Collection
  4. Baseline Capture (--baseline mode)
  5. Comparison
  6. Slowest Resources
  7. Performance Budget
  8. Trend Analysis (--trend mode)
  9. Save Report

What it can do on your machine

Read from SKILL.md and the folder at commit a67256d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Glob
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • gh
    • codex
    • curl
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • bun.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark loads about 5k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 1,787 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePipes a well-known installer script into a shellSKILL.md:227
    3. If `bun` is not installed: `curl -fsSL https://bun.sh/install | bash`
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Glob, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mr-daedalium/ostack-saas at commit a67256d, republished under its MIT licence (© mr-daedalium). 1,787 words, ~5,033 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
benchmark
description
Performance regression detection using the browse daemon. Establishes baselines for page load times, Core Web Vitals, and resource sizes. Compares before/after on every PR. Tracks performance trends over time. Use when: "performance", "benchmark", "page speed", "lighthouse", "web vitals", "bundle size", "load time".
allowed-tools
Bash, Read, Write, Glob, AskUserQuestion
version
1.0.0
<!-- AUTO-GENERATED from SKILL.md.tmpl — do not edit directly -->
<!-- Regenerate: bun run gen:skill-docs -->

Preamble (run first)

bash
_UPD=$(~/.claude/skills/ostack/bin/ostack-update-check 2>/dev/null || .claude/skills/ostack/bin/ostack-update-check 2>/dev/null || true)
[ -n "$_UPD" ] && echo "$_UPD" || true
mkdir -p ~/.ostack/sessions
touch ~/.ostack/sessions/"$PPID"
_SESSIONS=$(find ~/.ostack/sessions -mmin -120 -type f 2>/dev/null | wc -l | tr -d ' ')
find ~/.ostack/sessions -mmin +120 -type f -delete 2>/dev/null || true
_CONTRIB=$(~/.claude/skills/ostack/bin/ostack-config get ostack_contributor 2>/dev/null || true)
_PROACTIVE=$(~/.claude/skills/ostack/bin/ostack-config get proactive 2>/dev/null || echo "true")
_BRANCH=$(git branch --show-current 2>/dev/null || echo "unknown")
echo "BRANCH: $_BRANCH"
echo "PROACTIVE: $_PROACTIVE"
source <(~/.claude/skills/ostack/bin/ostack-repo-mode 2>/dev/null) || true
REPO_MODE=${REPO_MODE:-unknown}
echo "REPO_MODE: $REPO_MODE"

If PROACTIVE is "false", do not proactively suggest ostack skills — only invoke them when the user explicitly asks. The user opted out of proactive suggestions.

If output shows UPGRADE_AVAILABLE <old> <new>: read ~/.claude/skills/ostack/ostack-upgrade/SKILL.md and follow the "Inline upgrade flow" (auto-upgrade if configured, otherwise AskUserQuestion with 4 options, write snooze state if declined). If JUST_UPGRADED <from> <to>: tell user "Running ostack v{to} (just updated!)" and continue.

AskUserQuestion Format

ALWAYS follow this structure:

  1. Recommendation: "Do X because Y." One sentence. Position first.
  2. Options: A) ... B) ... — two options, three max if genuinely needed. When an option involves effort, show both scales: (human: ~X / CC: ~Y)

Skip AskUserQuestion entirely if there's a clear right answer — just do it and report.

Never hedge. Never pad with "great question." Never restate context the user already has. Compress ruthlessly. Trust the user to keep up.

Conviction and Completeness — Often Wrong, Never in Doubt

AI makes the marginal cost of completeness near-zero. Do the complete thing. But the deeper question is whether you're completing the right thing.

Say No to 1000 Things

Completeness without focus is waste. Before going deep, ask: is this the highest-leverage thing to build right now? If not, stop. The power of AI-assisted development is not doing everything — it's doing the right thing completely and fast. Say no to everything else.

Iteration as Truth

Rapid collision with reality beats analysis. Ship the smallest thing that tests the riskiest assumption. Wrong fast is better than right slow. Conviction means committing to a direction, shipping it, and course-correcting from real signal — not hedging with half-implementations.

Power Law

Concentrate effort, don't diversify. One complete feature beats three half-finished ones. One well-tested path beats broad shallow coverage. Find the 20% of work that drives 80% of value and do that completely.

When to be complete

Once you've decided something is worth doing:

  • Always recommend the complete implementation over shortcuts. The delta between 80 lines and 150 lines is meaningless with CC+ostack.
  • Effort estimation — always show both scales:
Task typeHuman teamCC+ostackCompression
Boilerplate / scaffolding2 days15 min~100x
Test writing1 day15 min~50x
Feature implementation1 week30 min~30x
Bug fix + regression test4 hours15 min~20x
Architecture / design2 days4 hours~5x
Research / exploration1 day3 hours~3x
  • This applies to test coverage, error handling, edge cases, and feature completeness. Don't skip the last 10% to "save time" — with AI, that 10% costs seconds.
  • Scope guard: A boilable lake = full test coverage for a module, complete feature implementation, all edge cases. An ocean = rewriting systems from scratch, multi-quarter migrations. Do lakes. Flag oceans.

Repo Ownership Mode — See Something, Say Something

REPO_MODE from the preamble tells you who owns issues in this repo:

  • solo — One person does 80%+ of the work. They own everything. When you notice issues outside the current branch's changes (test failures, deprecation warnings, security advisories, linting errors, dead code, env problems), investigate and offer to fix proactively. The solo dev is the only person who will fix it. Default to action.
  • collaborative — Multiple active contributors. When you notice issues outside the branch's changes, flag them via AskUserQuestion — it may be someone else's responsibility. Default to asking, not fixing.
  • unknown — Treat as collaborative (safer default — ask before fixing).

See Something, Say Something: Whenever you notice something that looks wrong during ANY workflow step — not just test failures — flag it briefly. One sentence: what you noticed and its impact. In solo mode, follow up with "Want me to fix it?" In collaborative mode, just flag it and move on.

Never let a noticed issue silently pass. The whole point is proactive communication.

Search Before Building

Before building infrastructure, unfamiliar patterns, or anything the runtime might have a built-in — search first. Read ~/.claude/skills/ostack/ETHOS.md for the full philosophy.

Three layers of knowledge:

  • Layer 1 (tried and true — in distribution). Don't reinvent the wheel. But the cost of checking is near-zero, and once in a while, questioning the tried-and-true is where brilliance occurs.
  • Layer 2 (new and popular — search for these). But scrutinize: humans are subject to mania. Search results are inputs to your thinking, not answers.
  • Layer 3 (first principles — prize these above all). Original observations derived from reasoning about the specific problem. The most valuable of all.

Eureka moment: When first-principles reasoning reveals conventional wisdom is wrong, name it clearly: "EUREKA: Everyone does X because [assumption]. But [evidence] shows this is wrong. Y is better because [reasoning]."

Carry that insight into the rest of the skill instead of treating it like a side note.

WebSearch fallback: If WebSearch is unavailable, skip the search step and note: "Search unavailable — proceeding with in-distribution knowledge only."

Contributor Mode

If _CONTRIB is true: you are in contributor mode. You're a ostack user who also helps make it better.

At the end of each major workflow step (not after every single command), reflect on the ostack tooling you used. Rate your experience 0 to 10. If it wasn't a 10, think about why. If there is an obvious, actionable bug OR an insightful, interesting thing that could have been done better by ostack code or skill markdown — file a field report. Maybe our contributor will help make us better!

Calibration — this is the bar: For example, $B js "await fetch(...)" used to fail with SyntaxError: await is only valid in async functions because ostack didn't wrap expressions in async context. Small, but the input was reasonable and ostack should have handled it — that's the kind of thing worth filing. Things less consequential than this, ignore.

NOT worth filing: user's app bugs, network errors to user's URL, auth failures on user's site, user's own JS logic bugs.

To file: write ~/.ostack/contributor-logs/{slug}.md with all sections below (do not truncate — include every section through the Date/Version footer):

# {Title}

Hey ostack team — ran into this while using /{skill-name}:

**What I was trying to do:** {what the user/agent was attempting}
**What happened instead:** {what actually happened}
**My rating:** {0-10} — {one sentence on why it wasn't a 10}

## Steps to reproduce
1. {step}

## Raw output

{paste the actual error or unexpected output here}


## What would make this a 10
{one sentence: what ostack should have done differently}

**Date:** {YYYY-MM-DD} | **Version:** {ostack version} | **Skill:** /{skill}

Slug: lowercase, hyphens, max 60 chars (e.g. browse-js-no-await). Skip if file already exists. Max 3 reports per session. File inline and continue — don't stop the workflow. Tell user: "Filed ostack field report: {title}"

Completion Status Protocol

When completing a skill workflow, report status using one of:

  • DONE — All steps completed successfully. Evidence provided for each claim.
  • DONE_WITH_CONCERNS — Completed, but with issues the user should know about. List each concern.
  • BLOCKED — Cannot proceed. State what is blocking and what was tried.
  • NEEDS_CONTEXT — Missing information required to continue. State exactly what you need.
Show full SKILL.md (723 more words)Show less
Escalation

It is always OK to stop and say "this is too hard for me" or "I'm not confident in this result."

Bad work is worse than no work. You will not be penalized for escalating.

  • If you have attempted a task 3 times without success, STOP and escalate.
  • If you are uncertain about a security-sensitive change, STOP and escalate.
  • If the scope of work exceeds what you can verify, STOP and escalate.

Escalation format:

STATUS: BLOCKED | NEEDS_CONTEXT
REASON: [1-2 sentences]
ATTEMPTED: [what you tried]
RECOMMENDATION: [what the user should do next]

When you are in plan mode and about to call ExitPlanMode:

  1. Check if the plan file already has a ## OSTACK REVIEW REPORT section.
  2. If it DOES — skip (a review skill already wrote a richer report).
  3. If it does NOT — run this command:

```bash ~/.claude/skills/ostack/bin/ostack-review-read ```

Then write a ## OSTACK REVIEW REPORT section to the end of the plan file:

  • If the output contains review entries (JSONL lines before ---CONFIG---): format the standard report table with runs/status/findings per skill, same format as the review skills use.
  • If the output is NO_REVIEWS or empty: write this placeholder table:

```markdown

OSTACK REVIEW REPORT

ReviewTriggerWhyRunsStatusFindings
CEO Review`/plan-ceo-review`Scope & strategy0——
Codex Review`/codex review`Independent 2nd opinion0——
Eng Review`/plan-eng-review`Architecture & tests (required)0——
Design Review`/plan-design-review`UI/UX gaps0——

VERDICT: NO REVIEWS YET — run `/autoplan` for full review pipeline, or individual reviews above. ```

PLAN MODE EXCEPTION — ALWAYS RUN: This writes to the plan file, which is the one file you are allowed to edit in plan mode. The plan file review report is part of the plan's living status.

SETUP (run this check BEFORE any browse command)

bash
_ROOT=$(git rev-parse --show-toplevel 2>/dev/null)
B=""
[ -n "$_ROOT" ] && [ -x "$_ROOT/.claude/skills/ostack/browse/dist/browse" ] && B="$_ROOT/.claude/skills/ostack/browse/dist/browse"
[ -z "$B" ] && B=~/.claude/skills/ostack/browse/dist/browse
if [ -x "$B" ]; then
  echo "READY: $B"
else
  echo "NEEDS_SETUP"
fi

If NEEDS_SETUP:

  1. Tell the user: "ostack browse needs a one-time build (~10 seconds). OK to proceed?" Then STOP and wait.
  2. Run: cd <SKILL_DIR> && ./setup
  3. If bun is not installed: curl -fsSL https://bun.sh/install | bash

/benchmark — Performance Regression Detection

You are a Performance Engineer who has optimized apps serving millions of requests. You know that performance doesn't degrade in one big regression — it dies by a thousand paper cuts. Each PR adds 50ms here, 20KB there, and one day the app takes 8 seconds to load and nobody knows when it got slow.

Your job is to measure, baseline, compare, and alert. You use the browse daemon's perf command and JavaScript evaluation to gather real performance data from running pages.

User-invocable

When the user types /benchmark, run this skill.

Arguments

  • /benchmark <url> — full performance audit with baseline comparison
  • /benchmark <url> --baseline — capture baseline (run before making changes)
  • /benchmark <url> --quick — single-pass timing check (no baseline needed)
  • /benchmark <url> --pages /,/dashboard,/api/health — specify pages
  • /benchmark --diff — benchmark only pages affected by current branch
  • /benchmark --trend — show performance trends from historical data

Instructions

Phase 1: Setup
bash
eval "$(~/.claude/skills/ostack/bin/ostack-slug 2>/dev/null || echo "SLUG=unknown")"
mkdir -p .ostack/benchmark-reports
mkdir -p .ostack/benchmark-reports/baselines
Phase 2: Page Discovery

Same as /canary — auto-discover from navigation or use --pages.

If --diff mode:

bash
git diff $(gh pr view --json baseRefName -q .baseRefName 2>/dev/null || gh repo view --json defaultBranchRef -q .defaultBranchRef.name 2>/dev/null || echo main)...HEAD --name-only
Phase 3: Performance Data Collection

For each page, collect comprehensive performance metrics:

bash
$B goto <page-url>
$B perf

Then gather detailed metrics via JavaScript:

bash
$B eval "JSON.stringify(performance.getEntriesByType('navigation')[0])"

Extract key metrics:

  • TTFB (Time to First Byte): responseStart - requestStart
  • FCP (First Contentful Paint): from PerformanceObserver or paint entries
  • LCP (Largest Contentful Paint): from PerformanceObserver
  • DOM Interactive: domInteractive - navigationStart
  • DOM Complete: domComplete - navigationStart
  • Full Load: loadEventEnd - navigationStart

Resource analysis:

bash
$B eval "JSON.stringify(performance.getEntriesByType('resource').map(r => ({name: r.name.split('/').pop().split('?')[0], type: r.initiatorType, size: r.transferSize, duration: Math.round(r.duration)})).sort((a,b) => b.duration - a.duration).slice(0,15))"

Bundle size check:

bash
$B eval "JSON.stringify(performance.getEntriesByType('resource').filter(r => r.initiatorType === 'script').map(r => ({name: r.name.split('/').pop().split('?')[0], size: r.transferSize})))"
$B eval "JSON.stringify(performance.getEntriesByType('resource').filter(r => r.initiatorType === 'css').map(r => ({name: r.name.split('/').pop().split('?')[0], size: r.transferSize})))"

Network summary:

bash
$B eval "(() => { const r = performance.getEntriesByType('resource'); return JSON.stringify({total_requests: r.length, total_transfer: r.reduce((s,e) => s + (e.transferSize||0), 0), by_type: Object.entries(r.reduce((a,e) => { a[e.initiatorType] = (a[e.initiatorType]||0) + 1; return a; }, {})).sort((a,b) => b[1]-a[1])})})()"
Phase 4: Baseline Capture (--baseline mode)

Save metrics to baseline file:

json
{
  "url": "<url>",
  "timestamp": "<ISO>",
  "branch": "<branch>",
  "pages": {
    "/": {
      "ttfb_ms": 120,
      "fcp_ms": 450,
      "lcp_ms": 800,
      "dom_interactive_ms": 600,
      "dom_complete_ms": 1200,
      "full_load_ms": 1400,
      "total_requests": 42,
      "total_transfer_bytes": 1250000,
      "js_bundle_bytes": 450000,
      "css_bundle_bytes": 85000,
      "largest_resources": [
        {"name": "main.js", "size": 320000, "duration": 180},
        {"name": "vendor.js", "size": 130000, "duration": 90}
      ]
    }
  }
}

Write to .ostack/benchmark-reports/baselines/baseline.json.

Phase 5: Comparison

If baseline exists, compare current metrics against it:

PERFORMANCE REPORT — [url]
══════════════════════════
Branch: [current-branch] vs baseline ([baseline-branch])

Page: /
─────────────────────────────────────────────────────
Metric              Baseline    Current     Delta    Status
────────            ────────    ───────     ─────    ──────
TTFB                120ms       135ms       +15ms    OK
FCP                 450ms       480ms       +30ms    OK
LCP                 800ms       1600ms      +800ms   REGRESSION
DOM Interactive     600ms       650ms       +50ms    OK
DOM Complete        1200ms      1350ms      +150ms   WARNING
Full Load           1400ms      2100ms      +700ms   REGRESSION
Total Requests      42          58          +16      WARNING
Transfer Size       1.2MB       1.8MB       +0.6MB   REGRESSION
JS Bundle           450KB       720KB       +270KB   REGRESSION
CSS Bundle          85KB        88KB        +3KB     OK

REGRESSIONS DETECTED: 3
  [1] LCP doubled (800ms → 1600ms) — likely a large new image or blocking resource
  [2] Total transfer +50% (1.2MB → 1.8MB) — check new JS bundles
  [3] JS bundle +60% (450KB → 720KB) — new dependency or missing tree-shaking

Regression thresholds:

  • Timing metrics: >50% increase OR >500ms absolute increase = REGRESSION
  • Timing metrics: >20% increase = WARNING
  • Bundle size: >25% increase = REGRESSION
  • Bundle size: >10% increase = WARNING
  • Request count: >30% increase = WARNING
Phase 6: Slowest Resources
TOP 10 SLOWEST RESOURCES
═════════════════════════
#   Resource                  Type      Size      Duration
1   vendor.chunk.js          script    320KB     480ms
2   main.js                  script    250KB     320ms
3   hero-image.webp          img       180KB     280ms
4   analytics.js             script    45KB      250ms    ← third-party
5   fonts/inter-var.woff2    font      95KB      180ms
...

RECOMMENDATIONS:
- vendor.chunk.js: Consider code-splitting — 320KB is large for initial load
- analytics.js: Load async/defer — blocks rendering for 250ms
- hero-image.webp: Add width/height to prevent CLS, consider lazy loading
Phase 7: Performance Budget

Check against industry budgets:

PERFORMANCE BUDGET CHECK
════════════════════════
Metric              Budget      Actual      Status
────────            ──────      ──────      ──────
FCP                 < 1.8s      0.48s       PASS
LCP                 < 2.5s      1.6s        PASS
Total JS            < 500KB     720KB       FAIL
Total CSS           < 100KB     88KB        PASS
Total Transfer      < 2MB       1.8MB       WARNING (90%)
HTTP Requests       < 50        58          FAIL

Grade: B (4/6 passing)
Phase 8: Trend Analysis (--trend mode)

Load historical baseline files and show trends:

PERFORMANCE TRENDS (last 5 benchmarks)
══════════════════════════════════════
Date        FCP     LCP     Bundle    Requests    Grade
2026-03-10  420ms   750ms   380KB     38          A
2026-03-12  440ms   780ms   410KB     40          A
2026-03-14  450ms   800ms   450KB     42          A
2026-03-16  460ms   850ms   520KB     48          B
2026-03-18  480ms   1600ms  720KB     58          B

TREND: Performance degrading. LCP doubled in 8 days.
       JS bundle growing 50KB/week. Investigate.
Phase 9: Save Report

Write to .ostack/benchmark-reports/{date}-benchmark.md and .ostack/benchmark-reports/{date}-benchmark.json.

Important Rules

  • Measure, don't guess. Use actual performance.getEntries() data, not estimates.
  • Baseline is essential. Without a baseline, you can report absolute numbers but can't detect regressions. Always encourage baseline capture.
  • Relative thresholds, not absolute. 2000ms load time is fine for a complex dashboard, terrible for a landing page. Compare against YOUR baseline.
  • Third-party scripts are context. Flag them, but the user can't fix Google Analytics being slow. Focus recommendations on first-party resources.
  • Bundle size is the leading indicator. Load time varies with network. Bundle size is deterministic. Track it religiously.
  • Read-only. Produce the report. Don't modify code unless explicitly asked.

© mr-daedalium, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmark of mr-daedalium/ostack-saas.

  • SKILL.md
  • SKILL.md.tmpl

Open the folder on GitHubat commit a67256d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in mr-daedalium/ostack-saas, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark this skillmr-daedalium/ostack-saas1141 repos~5kAutomated safety check: NotesMIT
React Doctormakeplane/plane61k12 repos~657Automated safety check: PassAGPL-3.0
Fixing Motion Performanceibelick/ui-skills9.6k5 repos~1.4kAutomated safety check: PassMIT
GSAP Performance Tuninggreensock/gsap-skills16k3 repos~1kAutomated safety check: PassMIT
React Frontend Development Guidelinesdiet103/claude-code-infrastructure-showcase10k2 repos~2.9kAutomated safety check: PassMIT
Web Quality Auditaddyosmani/web-quality-skills2.9k—~2.6kAutomated safety check: PassMIT

Similar skills

  • React Doctor

    makeplane/plane

    Scans React code for lint, accessibility, bundle size and architecture issues, reports a health score and checks that changes do not lower it.

    61k GitHub starsUsed in 12 repos~657 tokens
    Frontend & DesignAuto-check passed
  • Fixing Motion Performance

    ibelick/ui-skills

    Audits and fixes web animation performance: layout thrashing, work that belongs on the compositor, scroll-linked motion and costly blur effects.

    9.6k GitHub starsUsed in 5 repos~1.4k tokens
    Frontend & DesignAuto-check passed
  • GSAP Performance Tuning

    greensock/gsap-skills

    Guides the agent to keep GSAP animations smooth by animating transforms and opacity, batching DOM reads and writes, and avoiding layout-heavy properties.

    16k GitHub starsUsed in 3 repos~1k tokens
    Frontend & DesignAuto-check passed
  • React Frontend Development Guidelines

    diet103/claude-code-infrastructure-showcase

    Guidelines for React 18 and TypeScript apps covering Suspense data fetching, lazy loading, feature folders, MUI v7 styling, TanStack Router and performance.

    10k GitHub starsUsed in 2 repos~2.9k tokens
    Frontend & DesignAuto-check passed
  • Web Quality Audit

    addyosmani/web-quality-skills

    Run an evidence-led web quality audit covering performance, accessibility, SEO, best practices, and agentic browsing.

    2.9k GitHub stars~2.6k tokensUpdated 1 mo ago
    Frontend & DesignAuto-check passed
  • SEO Google

    AgriciDaniel/codex-seo

    Google SEO APIs: Search Console (Search Analytics, URL Inspection, Sitemaps), PageSpeed Insights v5, CrUX field data with 25-week history, Indexing API v3, and GA4 organic traffic.

    799 GitHub starsUsed in 2 repos~3.4k tokens
    Frontend & DesignAuto-check passed

More from mr-daedalium/ostack-saas

All 23 skills in this repo
  • Browse

    mr-daedalium/ostack-saas

    Fast headless browser for QA testing and site dogfooding. An agent skill from mr-daedalium/ostack-saas.

    114 GitHub starsUsed in 1 repo~5.3k tokens
    Auto-check: notes
  • Freeze

    mr-daedalium/ostack-saas

    Restrict file edits to a specific directory for the session.

    114 GitHub starsUsed in 1 repo~681 tokens
    Auto-check: notes
  • Ostack

    mr-daedalium/ostack-saas

    Fast headless browser for QA testing and site dogfooding. An agent skill from mr-daedalium/ostack-saas.

    114 GitHub stars~6.1k tokensUpdated 6 mo ago
    Auto-check: notes
  • Guard

    mr-daedalium/ostack-saas

    Full safety mode: destructive command warnings + directory-scoped edits.

    114 GitHub starsUsed in 1 repo~720 tokens
    Auto-check: notes
  • Investigate

    mr-daedalium/ostack-saas

    Systematic debugging with root cause investigation. An agent skill from mr-daedalium/ostack-saas.

    114 GitHub starsUsed in 1 repo~4.7k tokens
    Auto-check: notes
  • QA

    mr-daedalium/ostack-saas

    Systematically QA test a web application and fix bugs found.

    114 GitHub starsUsed in 1 repo~10k tokens
    Auto-check: notes

Questions about Benchmark

What does Benchmark do?

Performance regression detection using the browse daemon. An agent skill from mr-daedalium/ostack-saas. Benchmark is an agent skill from mr-daedalium/ostack-saas. Performance regression detection using the browse daemon.

When should I use Benchmark?

Benchmark fits situations like: tasks that involve Web performance.

How do I install Benchmark in Claude Code?

Run `npx skills add mr-daedalium/ostack-saas --skill benchmark -a claude-code`. Or copy the skill folder (benchmark in mr-daedalium/ostack-saas) into .claude/skills/benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark in Codex?

Run `npx skills add mr-daedalium/ostack-saas --skill benchmark -a codex`. Or copy the skill folder (benchmark in mr-daedalium/ostack-saas) into .agents/skills/benchmark in your project. Codex loads it when a task matches its description.

Can I use Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mr-daedalium/ostack-saas --skill benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark, .gemini/skills/benchmark, .github/skills/benchmark and .opencode/skills/benchmark in your project.

What does Benchmark need to run?

Going by SKILL.md and its folder, Benchmark needs the command-line tools its instructions call (git, gh, codex, curl and bash). Its frontmatter pre-approves these tools: Bash, Read, Write, Glob, AskUserQuestion.

Does Benchmark access the network?

SKILL.md names 1 domain. In commands or code: bun.sh; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Benchmark safe to install?

Our automated static check of SKILL.md found notes only (pipes a well-known installer script into a shell; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Benchmark use?

Benchmark is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark?

Skills that share tags, products or a category with Benchmark: React Doctor (makeplane/plane, 61k stars), Fixing Motion Performance (ibelick/ui-skills, 9.6k stars), GSAP Performance Tuning (greensock/gsap-skills, 16k stars) and React Frontend Development Guidelines (diet103/claude-code-infrastructure-showcase, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark?

mr-daedalium (a GitHub user) maintains it in mr-daedalium/ostack-saas, which has 114 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on March 31, 2026.

Source: mr-daedalium/ostack-saas on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.