Agent skill

Benchmark Update

by kdeldycke in kdeldycke/dotfiles

Create or update the competitive benchmark page (docs/benchmark.md) that compares this project against its alternatives.

BSD-2-ClauseAuto-check: notes

Install Benchmark Update

skills CLI
$ npx skills add kdeldycke/dotfiles --skill benchmark-update -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kdeldycke/dotfiles benchmark-update --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kdeldycke/dotfiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/dotfiles/.agents/skills/benchmark-update .claude/skills/benchmark-update && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-update
GitHub stars
173
Token cost
~3k tokens
SKILL.md length
1,374 words
Files
1
Skills in repo
25
Repo updated
First seen
Licence
BSD-2-Clause

At a glance

Create or update the competitive benchmark page (docs/benchmark.md) that compares this project against its alternatives.

  • Works in 11 steps: Identify the project and its domain → Discover competitors → Research features → …
  • SKILL.md covers Context and Instructions
  • Calls gh; reaches img.shields.io

What it does

Benchmark Update is an agent skill from kdeldycke/dotfiles. Create or update the competitive benchmark page (docs/benchmark.md) that compares this project against its alternatives. Check maintenance status, feature accuracy, new candidates and badge health.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code. Recommended model: Opus.

The repository describes itself as: 🍎 macOS dotfiles for Python developers. The licence is BSD-2-Clause.

Example prompts

  • “/benchmark-update”

Requirements

  • Compatibility (from SKILL.md): Designed for Claude Code. Recommended model: Opus.
  • Pre-approved tools (allowed-tools): Bash, Read, Grep, Glob, WebFetch, WebSearch, Agent

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. Identify the project and its domain
  2. Discover competitors
  3. Research features
  4. Gap analysis
  5. Build the page
  6. Maintenance status check
  7. Feature matrix accuracy
  8. Gap analysis freshness
  9. New candidates
  10. Badge health
  11. Star history chart

What it can do on your machine

Read from SKILL.md and the folder at commit 7947d0f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Grep
    • Glob
    • WebFetch
    • WebSearch
    • Agent

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • img.shields.io

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code. Recommended model: Opus.

    From compatibility in the SKILL.md frontmatter.

Context cost

Benchmark Update loads about 3k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 1,374 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Grep, Glob, WebFetch, WebSearch, Agent

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kdeldycke/dotfiles at commit 7947d0f, republished under its BSD-2-Clause licence (© kdeldycke). 1,374 words, ~2,996 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-update/SKILL.md (or your agent's skills folder).
name
benchmark-update
description
Create or update the competitive benchmark page (docs/benchmark.md) that compares this project against its alternatives. Check maintenance status, feature accuracy, new candidates and badge health.
allowed-tools
Bash, Read, Grep, Glob, WebFetch, WebSearch, Agent
compatibility
Designed for Claude Code. Recommended model: Opus.
argument-hint
[audit|init|add <project>|refresh-badges]

Context

![ -f pyproject.toml ] && grep '^name' pyproject.toml | head -1 || echo "No pyproject.toml" ![ -f docs/benchmark.md ] && head -15 docs/benchmark.md || echo "No docs/benchmark.md yet" ![ -f docs/benchmark.md ] && grep -c '^|' docs/benchmark.md || echo 0 ![ -f CLAUDE.md ] && grep -i 'ordering\|benchmark\|comparison' CLAUDE.md | head -5 || echo "No ordering conventions found"

Instructions

You create and maintain docs/benchmark.md, a competitive benchmark page comparing the current project against alternatives in the same space. The page follows a standard template with feature tables, GitHub activity badges, popularity charts, distribution info, and metadata.

Reference examples

Fetch these files as reference when building or auditing a benchmark page:

Scope selection
  • audit (default): Run all checks on an existing benchmark page and report findings. Do not edit unless the user confirms.
  • init: Create a benchmark page from scratch. Research competitors, build feature tables, and populate all sections.
  • add <project>: Research a specific project and add it to all tables in the correct position.
  • refresh-badges: Verify all badge URLs resolve. Fix broken ones (repos that moved, renamed packages).
Column ordering convention

If CLAUDE.md defines a benchmark ordering convention, follow it. Otherwise, place the current project first, its direct dependencies/foundation second, then remaining projects sorted by GitHub stars (descending).

Creating a benchmark from scratch (init)
1. Identify the project and its domain

Read pyproject.toml for the project name, description, and keywords. Determine what category of software this project is (CLI framework, package manager, static site generator, linter, etc.).

2. Discover competitors

Search for alternatives in the same space:

  • Search GitHub: gh search repos "<domain keywords>" --sort stars --limit 30
  • Search PyPI and other registries for similar tools.
  • Check "awesome" lists on GitHub.
  • Look at the project's own README or docs for mentions of alternatives.

For each candidate, collect: name, GitHub owner/repo, approximate star count, one-line description, and language/ecosystem.

Present the candidates and ask the user which to include.

3. Research features

For each included project, research its feature set by reading its documentation, README, and changelog. Build comparison tables relevant to the domain. Common table categories:

  • Features: core capabilities, grouped by audience (developer experience vs. end-user experience) when applicable.
  • Activity: GitHub badges (watchers, contributors, commit activity, commits since latest release, last release date, last commit, open issues, open PRs, forks, dependencies freshness).
  • Popularity: star history chart + badges (stars, SourceRank, dependent repos).
  • Distribution: package registry badges (PyPI, crates.io, npm, Homebrew, etc.) as applicable.
  • Metadata: license, main language, latest version, benchmark date.

Use ✅ for a supported feature, 🟡 for partial or opt-in support, ❌ for a feature the project lacks or rejects, and N/A where the row does not apply. Link every cell to its evidence: the documentation or a source line for ✅ and 🟡, and for ❌ a maintainer statement, a request closed as not planned or left open, or a documented limit. Absence of the feature is not evidence: with no citable source, leave the cell blank.

4. Gap analysis

After building the feature tables, analyze them to identify:

  • Unique strengths: features where this project leads or is the only one offering the capability. Summarize each with a brief explanation of why it matters.
  • Gaps and opportunities: features where competitors lead and this project could improve. For each gap:
    • Search the project's primary upstream dependency (if any) for related open issues and PRs using gh issue list --repo <upstream> --search "<feature>" --state open.
    • Link to specific upstream issues with their number, title, and demand signals (thumbs-up count, comment count).
    • Note whether upstream has declined the feature, has a pending PR, or has never been asked.
    • Describe what this project could do to close the gap (override, extend, integrate a third-party package, etc.).

This analysis turns the benchmark from a static comparison into an actionable roadmap.

5. Build the page

Follow this template structure:

markdown
# {octicon}`trophy` Benchmark

<intro paragraph about why this comparison exists>

## <Feature category 1>

<legend if needed>

| Feature | `this-project` | `competitor-1`[^1] | ... |
| ... | :---: | :---: | ... |

<brief prose highlighting key differences>

## <Feature category 2>

...

## Unique strengths

<bullet list of features where this project leads, with brief explanations>

## Gaps and opportunities

### <Gap title>

<description of the gap, with links to upstream issues like `[owner/repo#N](url)` including demand signals>

<what this project could do to close the gap>

### <Gap title>

...

## Activity

| Metrics | `this-project` | `competitor-1`[^1] | ... |
(GitHub badge rows: watchers, contributors, commit activity, etc.)

## Popularity

![Star history of `this-project` and its alternatives](assets/star-history-compared.svg)

<a sentence on the axis, when the chart is logarithmic, and one on how far back the store reaches>

| Metrics | `this-project` | `competitor-1`[^1] | ... |
(Stars, SourceRank, Dependent repos)

## Distribution

| Registry | `this-project` | `competitor-1`[^1] | ... |
(PyPI, crates.io, npm, Homebrew downloads)

## Metadata

| Metadata | `this-project` | `competitor-1`[^1] | ... |
(License, main language, latest version, benchmark date)

## Excluded projects

~~~{note}
<Project name> is not included because <reason>.
~~~

## Project URLs

[^1]: [<url>](<url>)

Badge format follows shields.io conventions: ![GitHub](https://img.shields.io/github/<metric>/<owner>/<repo>?label=%20&style=flat-square) for compact badges with no label text.

Star history chart: render it locally, never as a third-party embed. GitHub restricted its stargazer endpoints to a repository's own admins in June 2026, which left every third-party star chart rendering an error card. It reopened an anonymous star-history endpoint in September 2026. A history committed to the repository cannot be revoked upstream, whatever GitHub changes next. repomatic sample-metrics records the counts into a committed CSV and draws them as a themeable SVG: declare the projects in [tool.repomatic.metrics] subjects, declare a chart in [tool.repomatic.metrics] charts, and reference that chart's output path from this section. A comparison spanning projects of very different sizes wants scale = "logarithmic", and one comparing trajectories rather than dates wants mode = "relative".

A GitHub peer gets its whole curve, rebuilt from that endpoint back to its first star. That curve counts only the stars the repository still holds, so it understates every past week by the stars withdrawn since. A peer on any other forge is sampled forward, one reading per run, and that history cannot be backdated: it carries only two points at first, its creation date and the first reading, so a chart drawn early states a straight line between them. Both are worth a sentence under the chart rather than hiding.

Show full SKILL.md (506 more words)Show less
Auditing an existing benchmark (audit)
1. Maintenance status check

For each project in the benchmark, check via gh api:

  • Last commit date on the default branch.
  • Last GitHub release date.
  • Whether the repo is archived.
  • Any deprecation or sunsetting announcements in recent issues.

Classify each as: active, slow but maintained, stale (no activity in 12+ months), or abandoned/archived.

Stale or abandoned projects should be moved to the "Excluded projects" section with a {note} admonition explaining why.

2. Feature matrix accuracy

For each feature row, spot-check 2-3 projects against their current documentation or changelog. Flag cells that look wrong (features added or removed since the benchmark was written).

To audit the blank cells, research one project per subagent, and have it return the URL and the verbatim quote for each finding. Read every quote back from its source through the API before you fill a cell, and keep a cell blank when the finding contradicts a row that grades the same feature.

3. Gap analysis freshness

For each item in "Gaps and opportunities", check whether the linked upstream issues have changed status:

  • Closed issues: the gap may have been addressed upstream or in this project. Verify and update.
  • New upstream PRs: a fix may be in progress.
  • Features this project has since implemented: the gap should be removed and the feature moved to "Unique strengths" or the feature table updated.

Also check whether new gaps have appeared (competitors added features this project lacks).

4. New candidates

Search for projects in the same space with significant GitHub stars not already in the benchmark. Report candidates with: name, GitHub repo, stars, one-line description, and maintenance status.

5. Badge health

Verify the GitHub owner/repo in badge URLs matches the current canonical location. Repos get transferred. Check via gh api repos/{owner}/{repo} --jq '.full_name'.

6. Star history chart

Check the chart is a locally rendered asset rather than a third-party embed: every such embed served an error card once GitHub restricted stargazer access in June 2026, and a committed history cannot be revoked that way. Then check [tool.repomatic.metrics] subjects lists every project the page still compares and none it has excluded, since that one list feeds both the chart and the weekly sampler.

Adding a project (add)
  1. Research its features against every row in the existing feature tables.
  2. Determine the correct column position per the ordering convention.
  3. Add it to every table (features, activity, popularity, distribution, metadata).
  4. Add a footnote with the project URL.
  5. Add it to [tool.repomatic.metrics] subjects, which feeds the star history chart.
Output format

For audit, produce a summary table:

ProjectStatusIssues found
...active / stale / ...badge broken / feature changed / ...

Then list recommended actions. Do not edit the file until the user confirms.

Cross-referencing with docs/upstream.md

The gap analysis surfaces upstream issues and PRs that overlap with what docs/upstream.md tracks. When the benchmark discovers new upstream issues (reported, workaround in place, declined, or fixed), check whether they belong in docs/upstream.md too. Suggest running /upstream-audit sync-git afterward if new references were added to the benchmark.

© kdeldycke, BSD-2-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in dotfiles/.agents/skills/benchmark-update of kdeldycke/dotfiles.

Open the folder on GitHubat commit 7947d0f

Compare with similar skills

Benchmark Update next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Update compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Update this skillkdeldycke/dotfiles173—~3kAutomated safety check: NotesBSD-2-Clause
Benchmarkaffaan-m/ECC275k3 repos~654Automated safety check: PassMIT
Benchmarkaffaan-m/ECC275k—~412Automated safety check: PassMIT
Benchmarkaffaan-m/ECC275k—~330Automated safety check: PassMIT
Competitive Report Structureaffaan-m/ECC275k1 repos~2.1kAutomated safety check: PassMIT
Competitive Platform Analysisaffaan-m/ECC275k1 repos~3kAutomated safety check: PassMIT

Similar skills

  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    275k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed
  • Benchmark

    affaan-m/ECC

    このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

    275k GitHub stars~412 tokensUpdated 3 days ago
    Auto-check passed
  • Benchmark

    affaan-m/ECC

    使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。

    275k GitHub stars~330 tokensUpdated 3 days ago
    Auto-check passed
  • Assemble scored competitor profile cards (from benchmark-methodology) into a decision-grade competitive report with landscape map, competitor tiers, benchmarking matrix, white-space analysis…

    275k GitHub starsUsed in 1 repo~2.1k tokens
    Marketing & SEOAuto-check passed
  • A skill your agent uses when scoping a competitive landscape — identifying, categorising, and score-filtering a competitor set before any benchmarking begins.

    275k GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • Benchmark

    androidx/androidx

    Benchmarking and improving the performance of Jetpack Compose.

    6.1k GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed

More from kdeldycke/dotfiles

All 25 skills in this repo
  • Agent Config Self Tune

    kdeldycke/dotfiles

    Audit and tune the configuration of coding agents across Claude Code and pi - settings files (settings.json, settings.local.json), permission rules, instruction files (CLAUDE.md, AGENTS.md), skill…

    173 GitHub stars~3.4k tokensUpdated 4 days ago
    Auto-check: notes
  • Audit Repo Issues

    kdeldycke/dotfiles

    Analyze a GitHub repository's issues and PRs to find unaddressed feature requests, dismissed ideas, maintenance signals, and opportunities relevant to the current project.

    173 GitHub stars~2.5k tokensUpdated 4 days ago
    Auto-check passed
  • Brand Assets

    kdeldycke/dotfiles

    Create project logo and banner SVGs, then export them to light and dark PNG variants.

    173 GitHub stars~4.7k tokensUpdated 4 days ago
    Auto-check passed
  • Fill Web Form

    kdeldycke/dotfiles

    Fill a web form using data extracted from local documents (PDFs, images, spreadsheets).

    173 GitHub stars~2.3k tokensUpdated 4 days ago
    Auto-check passed
  • Rename With Dates

    kdeldycke/dotfiles

    Rename documents and files (PDFs, images, screenshots, etc.) by reading their content to extract the effective/publication date, then renaming them with a "YYYY-MM-DD - Clear descriptive title.ext"…

    173 GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check passed
  • Repomatic Test Matrix

    kdeldycke/dotfiles

    Choose what a repository's CI test matrix covers. An agent skill from kdeldycke/dotfiles.

    173 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check: notes

Questions about Benchmark Update

What does Benchmark Update do?

Create or update the competitive benchmark page (docs/benchmark.md) that compares this project against its alternatives. Benchmark Update is an agent skill from kdeldycke/dotfiles.md) that compares this project against its alternatives.

How do I install Benchmark Update in Claude Code?

Run `npx skills add kdeldycke/dotfiles --skill benchmark-update -a claude-code`. Or copy the skill folder (dotfiles/.agents/skills/benchmark-update in kdeldycke/dotfiles) into .claude/skills/benchmark-update in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Update in Codex?

Run `npx skills add kdeldycke/dotfiles --skill benchmark-update -a codex`. Or copy the skill folder (dotfiles/.agents/skills/benchmark-update in kdeldycke/dotfiles) into .agents/skills/benchmark-update in your project. Codex loads it when a task matches its description.

Can I use Benchmark Update in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kdeldycke/dotfiles --skill benchmark-update -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-update, .gemini/skills/benchmark-update, .github/skills/benchmark-update and .opencode/skills/benchmark-update in your project.

What does Benchmark Update need to run?

Going by SKILL.md and its folder, Benchmark Update needs the command-line tools its instructions call (gh). Its frontmatter pre-approves these tools: Bash, Read, Grep, Glob, WebFetch, WebSearch, Agent. Compatibility (from SKILL.md): Designed for Claude Code. Recommended model: Opus..

Does Benchmark Update access the network?

SKILL.md names 2 domains. In commands or code: img.shields.io; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Benchmark Update safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Update use?

Benchmark Update is published under the BSD-2-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Update use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Update?

Skills that share tags, products or a category with Benchmark Update: Benchmark (affaan-m/ECC, 275k stars), Benchmark (affaan-m/ECC, 275k stars), Benchmark (affaan-m/ECC, 275k stars) and Competitive Report Structure (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Update?

kdeldycke (a GitHub user) maintains it in kdeldycke/dotfiles, which has 173 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 4, 2026.

Source: kdeldycke/dotfiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.