Agent skill

Pinchbench

by dixiyao in dixiyao/FoT

Run PinchBench benchmarks to evaluate OpenClaw agent performance across real-world tasks.

MITAuto-check passedWriting & Content

Install Pinchbench

skills CLI
$ npx skills add dixiyao/FoT --skill pinchbench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dixiyao/FoT pinchbench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dixiyao/FoT.git skills-src && mkdir -p .claude/skills && cp -r skills-src/experiment/pinchbench .claude/skills/pinchbench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pinchbench
GitHub stars
113
Used in
1 other repo
Token cost
~1.1k tokens
SKILL.md length
266 words
Files
262 (incl. scripts, assets)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Run PinchBench benchmarks to evaluate OpenClaw agent performance across real-world tasks.

  • Testing model capabilities
  • SKILL.md covers Prerequisites, Quick Start, Available Tasks (23) and Command Line Options, plus 4 more sections
  • Calls uv and jq
  • Comparing models

What it does

Pinchbench is an agent skill from dixiyao/FoT. Run PinchBench benchmarks to evaluate OpenClaw agent performance across real-world tasks. Use when testing model capabilities, comparing models, submitting benchmark results to the leaderboard, or checking how well your OpenClaw setup handles calendar, email, research, coding, and multi-step workflows.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 264 other files, including scripts and assets (for example `.github/benchmark-models.yml`, `.github/workflows/lint.yml` and `.github/workflows/release.yml`).

It sits in Writing & Content. The repository describes itself as: Federation over Text (FoT) is a federated-learning-like paradigm for multi-agent reasoning. The licence is MIT.

When your agent uses it

  • Testing model capabilities
  • Comparing models
  • Submitting benchmark results to the leaderboard
  • Checking how well your OpenClaw setup handles calendar

Example prompts

  • “/pinchbench”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9761942. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • pinchbench.com
    • docs.astral.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pinchbench loads about 1.1k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 266 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from dixiyao/FoT at commit 9761942, republished under its MIT licence (© dixiyao). 266 words, ~1,050 tokens.

Download SKILL.mdSave it as .claude/skills/pinchbench/SKILL.md (or your agent's skills folder). This skill also uses 261 other files; get the full folder from GitHub.
name
pinchbench
description
Run PinchBench benchmarks to evaluate OpenClaw agent performance across real-world tasks. Use when testing model capabilities, comparing models, submitting benchmark results to the leaderboard, or checking how well your OpenClaw setup handles calendar, email, research, coding, and multi-step workflows.
metadata.author
pinchbench
metadata.version
2.0.0-rc1
metadata.homepage
https://pinchbench.com
metadata.repository
https://github.com/pinchbench/skill

PinchBench Benchmark Skill

PinchBench measures how well LLM models perform as the brain of an OpenClaw agent. Results are collected on a public leaderboard at pinchbench.com.

Prerequisites

  • Python 3.10+
  • uv package manager
  • OpenClaw instance (this agent)

Quick Start

bash
cd <skill_directory>

# Run benchmark with a specific model
uv run benchmark.py --model anthropic/claude-sonnet-4

# Run only automated tasks (faster)
uv run benchmark.py --model anthropic/claude-sonnet-4 --suite automated-only

# Run specific tasks
uv run benchmark.py --model anthropic/claude-sonnet-4 --suite task_calendar,task_stock

# Skip uploading results
uv run benchmark.py --model anthropic/claude-sonnet-4 --no-upload

Available Tasks (23)

TaskCategoryDescription
task_sanityBasicVerify agent works
task_calendarProductivityCalendar event creation
task_stockResearchStock price lookup
task_blogWritingBlog post creation
task_weatherCodingWeather script
task_summaryAnalysisDocument summarization
task_eventsResearchConference research
task_emailWritingEmail drafting
task_memoryMemoryContext retrieval
task_filesFilesFile structure creation
task_workflowIntegrationMulti-step API workflow
task_clawdhubSkillsClawHub interaction
task_skill_searchSkillsSkill discovery
task_image_genCreativeImage generation
task_humanizerWritingText humanization
task_daily_summaryProductivityDaily digest
task_email_triageEmailInbox triage
task_email_searchEmailEmail search
task_market_researchResearchMarket analysis
task_spreadsheet_summaryAnalysisSpreadsheet analysis
task_eli5_pdf_summaryAnalysisPDF simplification
task_openclaw_comprehensionKnowledgeOpenClaw docs comprehension
task_second_brainMemoryKnowledge management

Command Line Options

OptionDescription
--modelModel identifier (e.g., anthropic/claude-sonnet-4)
--suiteall, automated-only, or comma-separated task IDs
--output-dirResults directory (default: results/)
--timeout-multiplierScale task timeouts for slower models
--runsNumber of runs per task for averaging
--no-uploadSkip uploading to leaderboard
--registerRequest new API token for submissions
--upload FILEUpload previous results JSON

Token Registration

To submit results to the leaderboard:

bash
# Register for an API token (one-time)
uv run benchmark.py --register

# Run benchmark (auto-uploads with token)
uv run benchmark.py --model anthropic/claude-sonnet-4

Results

Results are saved as JSON in the output directory:

bash
# View task scores
jq '.tasks[] | {task_id, score: .grading.mean}' results/0001_anthropic-claude-sonnet-4.json

# Show failed tasks
jq '.tasks[] | select(.grading.mean < 0.5)' results/*.json

# Calculate overall score
jq '{average: ([.tasks[].grading.mean] | add / length)}' results/*.json

Adding Custom Tasks

Create a markdown file in tasks/ following TASK_TEMPLATE.md. Each task needs:

  • YAML frontmatter (id, name, category, grading_type, timeout)
  • Prompt section
  • Expected behavior
  • Grading criteria
  • Automated checks (Python grading function)

Leaderboard

View results at pinchbench.com. The leaderboard shows:

  • Model rankings by overall score
  • Per-task breakdowns
  • Historical performance trends

© dixiyao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 261 other files (scripts, assets) in experiment/pinchbench of dixiyao/FoT.

  • SKILL.md
  • .github/benchmark-models.yml
  • .github/workflows/lint.yml
  • .github/workflows/release.yml
  • .github/workflows/update-task-count.yml
  • .gitignore
  • .pre-commit-config.yaml
  • BENCHMARK_VERSION
  • Dockerfile.benchmark
  • LICENSE
  • README.md
  • assets/GPT4.pdf
  • assets/OpenClaw Agent Use Cases and Gap Analysis for PinchBench.pdf
  • assets/ai_blog.txt
  • assets/broken_ci.yml
  • assets/broken_k8s_deployment.yml
  • assets/cncf-tag-runtime-notes.md
  • assets/company_expenses.xlsx
  • … and 244 more

Open the folder on GitHubat commit 9761942

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in dixiyao/FoT, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Pinchbench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pinchbench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pinchbench this skilldixiyao/FoT1131 repos~1.1kAutomated safety check: PassMIT
Notion To Blogwasp-lang/wasp19k—~922Automated safety check: PassMIT
Cuimao TranslatorCuimao777/cuimao-translator486—~3.2kAutomated safety check: PassNone
AI Daily Newsgeekjourneyx/ai-daily-skill235—~2.3kAutomated safety check: PassNone
China Travel Kittczyliu/china-travel-kit194—~1.3kAutomated safety check: PassMIT
Story Maintenancedanjdewhurst/story-skills2831 repos~5.4kAutomated safety check: NotesMIT

Similar skills

  • Notion To Blog

    wasp-lang/wasp

    Transfer a blog post from Notion to the Wasp blog. An agent skill from wasp-lang/wasp.

    19k GitHub stars~922 tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Cuimao Translator

    Cuimao777/cuimao-translator

    Translate English PDF books into natural Chinese with three quality modes.

    486 GitHub stars~3.2k tokensUpdated 5 mo ago
    Writing & ContentAuto-check passed
  • AI Daily News

    geekjourneyx/ai-daily-skill

    Fetches AI news from smol.ai RSS and generates structured markdown with intelligent summarization and categorization.

    235 GitHub stars~2.3k tokensUpdated today
    Writing & ContentAuto-check passed
  • China Travel Kit

    tczyliu/china-travel-kit

    Research and plan first-time independent trips in China with bilingual, source-aware city data and official live-check entry points.

    194 GitHub stars~1.3k tokensUpdated 1 mo ago
    Writing & ContentAuto-check passed
  • Story Maintenance

    danjdewhurst/story-skills

    This skill should be used when the user asks to "validate", "reindex", "repair registries", "check links", "run the continuity, pacing, clue, voice, or name checks", "count words", "summarize a…

    283 GitHub starsUsed in 1 repo~5.4k tokens
    Writing & ContentAuto-check: notes
  • Whitepaper Bare Section Prose Filler

    FlorianBruniaux/claude-code-ultimate-guide

    Adds two to four sentences of intro prose before each bare heading in the French and English whitepapers, splitting the work across nine parallel agents.

    6.1k GitHub stars~1.5k tokensUpdated 2 days ago
    Writing & ContentAuto-check passed

Questions about Pinchbench

What does Pinchbench do?

Run PinchBench benchmarks to evaluate OpenClaw agent performance across real-world tasks. Pinchbench is an agent skill from dixiyao/FoT. Run PinchBench benchmarks to evaluate OpenClaw agent performance across real-world tasks.

When should I use Pinchbench?

Pinchbench fits situations like: testing model capabilities; comparing models; submitting benchmark results to the leaderboard; checking how well your OpenClaw setup handles calendar.

How do I install Pinchbench in Claude Code?

Run `npx skills add dixiyao/FoT --skill pinchbench -a claude-code`. Or copy the skill folder (experiment/pinchbench in dixiyao/FoT) into .claude/skills/pinchbench in your project. Claude Code loads it when a task matches its description.

How do I install Pinchbench in Codex?

Run `npx skills add dixiyao/FoT --skill pinchbench -a codex`. Or copy the skill folder (experiment/pinchbench in dixiyao/FoT) into .agents/skills/pinchbench in your project. Codex loads it when a task matches its description.

Can I use Pinchbench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dixiyao/FoT --skill pinchbench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pinchbench, .gemini/skills/pinchbench, .github/skills/pinchbench and .opencode/skills/pinchbench in your project.

What does Pinchbench need to run?

Going by SKILL.md and its folder, Pinchbench needs the command-line tools its instructions call (uv and jq). Our summary lists: Python 3.

Does Pinchbench access the network?

SKILL.md names 2 domains. As links in the text: pinchbench.com and docs.astral.sh. This is read from the text; nothing was executed.

Is Pinchbench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Pinchbench use?

Pinchbench is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pinchbench use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pinchbench?

Skills that share tags, products or a category with Pinchbench: Notion To Blog (wasp-lang/wasp, 19k stars), Cuimao Translator (Cuimao777/cuimao-translator, 486 stars), AI Daily News (geekjourneyx/ai-daily-skill, 235 stars) and China Travel Kit (tczyliu/china-travel-kit, 194 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pinchbench?

dixiyao (a GitHub user) maintains it in dixiyao/FoT, which has 113 GitHub stars. The repository was last updated on September 29, 2026.

Source: dixiyao/FoT on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.