Agent skill

Skillsbench

by benchflow-ai in benchflow-ai/skillsbench

SkillsBench contribution workflow. An agent skill from benchflow-ai/skillsbench.

Apache-2.0Auto-check passed

Install Skillsbench

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill skillsbench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench skillsbench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/skillsbench .claude/skills/skillsbench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skillsbench
GitHub stars
1.8k
Token cost
~426 tokens
SKILL.md length
63 words
Files
1
Skills in repo
189
Repo updated
First seen
Licence
Apache-2.0

At a glance

SkillsBench contribution workflow. An agent skill from benchflow-ai/skillsbench.

  • Creating benchmark tasks
  • SKILL.md covers Official Resources, Quick Workflow, Task Requirements and References
  • Calls uv
  • Understanding repo structure

What it does

Skillsbench is an agent skill from benchflow-ai/skillsbench. SkillsBench contribution workflow. Use when: (1) Creating benchmark tasks, (2) Understanding repo structure, (3) Preparing PRs for task submission.

Its SKILL.md is about 430 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Creating benchmark tasks
  • Understanding repo structure
  • Preparing PRs for task submission

Example prompts

  • “/skillsbench”

Requirements

  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skillsbench loads about 426 tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 63 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~426

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 63 words, ~426 tokens.

Download SKILL.mdSave it as .claude/skills/skillsbench/SKILL.md (or your agent's skills folder).
name
skillsbench
description
SkillsBench contribution workflow. Use when: (1) Creating benchmark tasks, (2) Understanding repo structure, (3) Preparing PRs for task submission.

SkillsBench

Benchmark evaluating how well AI agents use skills.

Official Resources

Quick Workflow

bash
# 1. Install and initialize a task
# v1.1 docs target BenchFlow 0.6.2; uv installs the current stable CLI.
uv tool install benchflow
bench tasks init <task-id>

# 2. Write files (see CONTRIBUTING.md for templates)

# 3. Validate
bench tasks check tasks/my-task
bench eval run --tasks-dir tasks/my-task --agent oracle --sandbox docker  # Must pass 100%

# 4. Test with agent (with skills)
bench eval run --tasks-dir tasks/my-task --agent claude-agent-acp \
  --model <model> --skill-mode with-skill \
  --skills-dir tasks/my-task/environment/skills/

# 5. Test WITHOUT skills
bench eval run --tasks-dir tasks/my-task --agent claude-agent-acp \
  --model <model> --skill-mode no-skill

# 6. Submit PR (see PR template)

Task Requirements

  • Realistic workflows people actually do
  • Measurably easier with skills than without
  • task.md prompt body and oracle/solve.sh must be human-authored
  • Deterministic, outcome-based verification
  • Skills must be generalizable and reusable

References

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/skillsbench of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Skillsbench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skillsbench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skillsbench this skillbenchflow-ai/skillsbench1.8k—~426Automated safety check: PassApache-2.0
Benchmarkaffaan-m/ECC276k3 repos~654Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k—~412Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k—~330Automated safety check: PassMIT
Contributemindfold-ai/Trellis15k—~2.7kAutomated safety check: PassAGPL-3.0
Benchmarkandroidx/androidx6.1k—~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    276k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed
  • Benchmark

    affaan-m/ECC

    このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

    276k GitHub stars~412 tokensUpdated 4 days ago
    Auto-check passed
  • Benchmark

    affaan-m/ECC

    使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。

    276k GitHub stars~330 tokensUpdated 4 days ago
    Auto-check passed
  • Contribute

    mindfold-ai/Trellis

    Guide for contributing to Trellis documentation and marketplace.

    15k GitHub stars~2.7k tokensUpdated 10 days ago
    DevelopmentAuto-check passed
  • Benchmark

    androidx/androidx

    Benchmarking and improving the performance of Jetpack Compose.

    6.1k GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed
  • Contribute

    superset-sh/superset

    Set up a Superset open-source contribution, from forking and cloning superset-sh/superset through local dev setup and the repo's rules to a merge-ready PR.

    15k GitHub stars~624 tokensUpdated today
    DevelopmentAuto-check passed

More from benchflow-ai/skillsbench

All 189 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Skillsbench

What does Skillsbench do?

SkillsBench contribution workflow. An agent skill from benchflow-ai/skillsbench. Skillsbench is an agent skill from benchflow-ai/skillsbench. SkillsBench contribution workflow.

When should I use Skillsbench?

Skillsbench fits situations like: creating benchmark tasks; understanding repo structure; preparing PRs for task submission.

How do I install Skillsbench in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill skillsbench -a claude-code`. Or copy the skill folder (.agents/skills/skillsbench in benchflow-ai/skillsbench) into .claude/skills/skillsbench in your project. Claude Code loads it when a task matches its description.

How do I install Skillsbench in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill skillsbench -a codex`. Or copy the skill folder (.agents/skills/skillsbench in benchflow-ai/skillsbench) into .agents/skills/skillsbench in your project. Codex loads it when a task matches its description.

Can I use Skillsbench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill skillsbench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skillsbench, .gemini/skills/skillsbench, .github/skills/skillsbench and .opencode/skills/skillsbench in your project.

What does Skillsbench need to run?

Going by SKILL.md and its folder, Skillsbench needs the command-line tools its instructions call (uv). Our summary lists: Docker.

Does Skillsbench access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Skillsbench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skillsbench use?

Skillsbench is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skillsbench use?

About 426 tokens (SKILL.md is roughly 1.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skillsbench?

Skills that share tags, products or a category with Skillsbench: Benchmark (affaan-m/ECC, 276k stars), Benchmark (affaan-m/ECC, 276k stars), Benchmark (affaan-m/ECC, 276k stars) and Contribute (mindfold-ai/Trellis, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skillsbench?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,834 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.