Agent skill

Skillhone Benchmark Optimization

by Tencent in Tencent/SkillHone

Run the additional paper-compatible optimization workflow with a frozen benchmark repository.

Custom licenceAuto-check passed

Install Skillhone Benchmark Optimization

skills CLI
$ npx skills add Tencent/SkillHone --skill skillhone-benchmark-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Tencent/SkillHone skillhone-benchmark-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Tencent/SkillHone.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skillhone-benchmark-optimization .claude/skills/skillhone-benchmark-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skillhone-benchmark-optimization
GitHub stars
168
Token cost
~1.5k tokens
SKILL.md length
739 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
Custom licence

At a glance

Run the additional paper-compatible optimization workflow with a frozen benchmark repository.

  • Works in 4 steps: Prepare the evaluation repository → Freeze the campaign → Measure and optimize → …
  • Asks to generate
  • SKILL.md covers 1. Prepare the evaluation…, 2. Freeze the campaign, 3. Measure and optimize and 4. Final measurement and review, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Skillhone Benchmark Optimization is an agent skill from Tencent/SkillHone. Run the additional paper-compatible optimization workflow with a frozen benchmark repository. Use only when the user asks to generate or reuse an eval set, establish a baseline, improve probe or PR-validation scores, compare a Skill against a benchmark, or reproduce the evaluation loop described in the SkillHone paper. Do not use for a single defect observed during normal Agent work; use skillhone-auto-optimization for the default path.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Continual agent skill evolution through persistent decision history. Whole-skill optimisation (SKILL.md + scripts + references) with every decision landing as a local Git issue /…

When your agent uses it

  • Asks to generate
  • Reuse an eval set
  • Establish a baseline
  • PR-validation scores

Example prompts

  • “/skillhone-benchmark-optimization”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Prepare the evaluation repository
  2. Freeze the campaign
  3. Measure and optimize
  4. Final measurement and review

What it can do on your machine

Read from SKILL.md and the folder at commit c613aa9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skillhone Benchmark Optimization loads about 1.5k tokens when it runs. Until then it costs about 118 tokens; SKILL.md has 739 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 739 words (~1,517 tokens).

“Use this additional workflow when the user explicitly wants a frozen evaluation dataset to discover which Skill failure to fix next. It is not a prerequisite for proactive maintenance. It feeds another evidence source into the same Issue-driven SkillHone review…”

— opening of SKILL.md by Tencent, Custom licence
name
skillhone-benchmark-optimization

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/skillhone-benchmark-optimization of Tencent/SkillHone.

Open the folder on GitHubat commit c613aa9

Compare with similar skills

Skillhone Benchmark Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skillhone Benchmark Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skillhone Benchmark Optimization this skillTencent/SkillHone168—~1.5kAutomated safety check: PassCustom licence
Benchmarkaffaan-m/ECC274k3 repos~654Automated safety check: PassMIT
Benchmarkaffaan-m/ECC274k—~412Automated safety check: PassMIT
Benchmarkaffaan-m/ECC274k—~330Automated safety check: PassMIT
Benchmarkandroidx/androidx6.1k—~1.1kAutomated safety check: PassApache-2.0
Benchmarksamchon/typia5.9k—~1.2kAutomated safety check: PassMIT

Similar skills

  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    274k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed
  • Benchmark

    affaan-m/ECC

    このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

    274k GitHub stars~412 tokensUpdated 2 days ago
    Auto-check passed
  • Benchmark

    affaan-m/ECC

    使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。

    274k GitHub stars~330 tokensUpdated 2 days ago
    Auto-check passed
  • Benchmark

    androidx/androidx

    Benchmarking and improving the performance of Jetpack Compose.

    6.1k GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed
  • Benchmark

    samchon/typia

    Defines typia benchmark fixture integrity, result reporting, and publication safeguards.

    5.9k GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Establishes page load, Core Web Vitals and resource-size baselines, then compares before and after on every pull request to track performance trends over time.

    136k GitHub stars~7.2k tokensUpdated today
    Frontend & DesignAuto-check: notes

More from Tencent/SkillHone

  • Skillhone

    Tencent/SkillHone

    Local Issue, pull-request, and Wiki workbench for agent skills.

    168 GitHub stars~3.4k tokensUpdated 18 days ago
    Auto-check passed
  • Mandatory interception workflow for a reproducible defect encountered while any Agent is using a skill.

    168 GitHub stars~2.1k tokensUpdated 18 days ago
    Auto-check: warnings

Questions about Skillhone Benchmark Optimization

What does Skillhone Benchmark Optimization do?

Run the additional paper-compatible optimization workflow with a frozen benchmark repository. Skillhone Benchmark Optimization is an agent skill from Tencent/SkillHone. Run the additional paper-compatible optimization workflow with a frozen benchmark repository.

When should I use Skillhone Benchmark Optimization?

Skillhone Benchmark Optimization fits situations like: asks to generate; reuse an eval set; establish a baseline; PR-validation scores.

How do I install Skillhone Benchmark Optimization in Claude Code?

Run `npx skills add Tencent/SkillHone --skill skillhone-benchmark-optimization -a claude-code`. Or copy the skill folder (skills/skillhone-benchmark-optimization in Tencent/SkillHone) into .claude/skills/skillhone-benchmark-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Skillhone Benchmark Optimization in Codex?

Run `npx skills add Tencent/SkillHone --skill skillhone-benchmark-optimization -a codex`. Or copy the skill folder (skills/skillhone-benchmark-optimization in Tencent/SkillHone) into .agents/skills/skillhone-benchmark-optimization in your project. Codex loads it when a task matches its description.

Can I use Skillhone Benchmark Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Tencent/SkillHone --skill skillhone-benchmark-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skillhone-benchmark-optimization, .gemini/skills/skillhone-benchmark-optimization, .github/skills/skillhone-benchmark-optimization and .opencode/skills/skillhone-benchmark-optimization in your project.

What does Skillhone Benchmark Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: Skillhone Benchmark Optimization is instructions for the agent only. Our summary lists: Python 3.

Does Skillhone Benchmark Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skillhone Benchmark Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skillhone Benchmark Optimization use?

Skillhone Benchmark Optimization has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Skillhone Benchmark Optimization use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skillhone Benchmark Optimization?

Skills that share tags, products or a category with Skillhone Benchmark Optimization: Benchmark (affaan-m/ECC, 274k stars), Benchmark (affaan-m/ECC, 274k stars), Benchmark (affaan-m/ECC, 274k stars) and Benchmark (androidx/androidx, 6.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skillhone Benchmark Optimization?

Tencent (a GitHub organization) maintains it in Tencent/SkillHone, which has 168 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 19, 2026.

Source: Tencent/SkillHone on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.