Agent skill

SWE Benchmark Task Adder

by ory in ory/lumen

Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

Custom licenceAuto-check passedAI & LLM Engineering

Install SWE Benchmark Task Adder

skills CLI
$ npx skills add ory/lumen --skill add-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ory/lumen add-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ory/lumen.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-benchmark .claude/skills/add-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-benchmark
GitHub stars
305
Token cost
~497 tokens
SKILL.md length
223 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
Custom licence

At a glance

Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

  • Works in 2 steps: Dispatch the task-curator agent with the… → Report the result including
  • Turning a real GitHub bug-fix into a new bench-swe benchmark task
  • SKILL.md covers Arguments, Repository selection criteria and Steps
  • Reaches github.com

What it does

You supply a GitHub issue or pull request URL and a language (go, python, typescript, javascript, rust, ruby, java, c, cpp, php or csharp). The agent hands the work to a task-curator agent, which validates the inputs, checks the repository's size and dependency count, resolves the fix PR, clones the repo, extracts the base and fix commits and generates the gold patch. It also works out the test command from the repo's conventions.

The task JSON is written under bench-swe/tasks in a folder for that language and the patch under bench-swe/patches. Good candidates are focused libraries with a clear bug, under 50 MB and 800 source files with fewer than 50 direct dependencies, and the agent rejects repositories beyond those limits. Five inline checks follow (patch applies, files match, no leaks, schema completeness, no test files in the patch), and any problems are fixed. The final report gives the task ID, repo, issue URL, files and lines changed, and a verification table.

When your agent uses it

  • Turning a real GitHub bug-fix into a new bench-swe benchmark task
  • Checking whether a candidate repository is small enough to serve as a benchmark
  • Verifying that a generated gold patch applies and does not include test files

Example prompts

  • “Add a benchmark task from https://github.com/gorilla/mux/issues/534 for Go.”
  • “Add https://github.com/gorilla/mux/pull/585 to bench-swe as a Go task.”
  • “Turn this bug-fix PR into a Python benchmark task: https://github.com/example-org/parser/pull/7.”

Requirements

  • A repository with the bench-swe pipeline
  • Network access to GitHub for cloning the target repo

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Dispatch the task-curator agent with the provided arguments. The agent
  2. Report the result including

What it can do on your machine

Read from SKILL.md and the folder at commit f60f9ec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

SWE Benchmark Task Adder loads about 497 tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 223 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~497

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 223 words (~497 tokens).

“Add a new benchmark task to the bench-swe pipeline from a real GitHub bug-fix. The human provides the GitHub issue or PR URL; the agent handles extraction, validation, and file creation.”

— opening of SKILL.md by ory, Custom licence
name
add-benchmark
argument-hint
<github-issue-or-pr-url> <language>
disable-model-invocation
true

Read the full SKILL.md on GitHub

Files

Just SKILL.md in .claude/skills/add-benchmark of ory/lumen.

Open the folder on GitHubat commit f60f9ec

Compare with similar skills

SWE Benchmark Task Adder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

SWE Benchmark Task Adder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
SWE Benchmark Task Adder this skillory/lumen305—~497Automated safety check: PassCustom licence
Octocode Benchmark Runnerbgauryy/octocode949—~2.1kAutomated safety check: PassMIT
Octocode Graph Eval Loopbgauryy/octocode949—~1.6kAutomated safety check: PassMIT
Windmill AI Evalswindmill-labs/windmill18k—~969Automated safety check: NotesCustom licence
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
GAIA Agent Benchmarkingamd/gaia1.6k—~1.8kAutomated safety check: PassMIT

Similar skills

  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Octocode Graph Eval Loop

    bgauryy/octocode

    Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

    949 GitHub stars~1.6k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • Windmill AI Evals

    windmill-labs/windmill

    Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

    18k GitHub stars~969 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • AI Project Copilot

    sun461941-hub/ai-project-copilot

    A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

    100 GitHub stars~3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from ory/lumen

  • Lumen Doctor

    ory/lumen

    Checks the Lumen semantic search setup for the current project: embedding service reachability, index freshness and storage totals, with remediation notes.

    305 GitHub stars~339 tokensUpdated 1 mo ago
    Auto-check passed
  • Refreshes or rebuilds the local Lumen semantic code index for the current project, preferring an MCP-driven refresh and reserving the CLI for an explicit clean rebuild.

    305 GitHub stars~391 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about SWE Benchmark Task Adder

What does SWE Benchmark Task Adder do?

Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch. You supply a GitHub issue or pull request URL and a language (go, python, typescript, javascript, rust, ruby, java, c, cpp, php or csharp). The agent hands the work to a task-curator agent, which validates the inputs, checks the repository's size and dependency count, resolves the fix PR, clones the repo, extracts the base and fix commits and generates the gold patch.

When should I use SWE Benchmark Task Adder?

SWE Benchmark Task Adder fits situations like: turning a real GitHub bug-fix into a new bench-swe benchmark task; checking whether a candidate repository is small enough to serve as a benchmark; verifying that a generated gold patch applies and does not include test files.

How do I install SWE Benchmark Task Adder in Claude Code?

Run `npx skills add ory/lumen --skill add-benchmark -a claude-code`. Or copy the skill folder (.claude/skills/add-benchmark in ory/lumen) into .claude/skills/add-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install SWE Benchmark Task Adder in Codex?

Run `npx skills add ory/lumen --skill add-benchmark -a codex`. Or copy the skill folder (.claude/skills/add-benchmark in ory/lumen) into .agents/skills/add-benchmark in your project. Codex loads it when a task matches its description.

Can I use SWE Benchmark Task Adder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ory/lumen --skill add-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-benchmark, .gemini/skills/add-benchmark, .github/skills/add-benchmark and .opencode/skills/add-benchmark in your project.

What does SWE Benchmark Task Adder need to run?

SKILL.md names no scripts, command-line tools or credentials: SWE Benchmark Task Adder is instructions for the agent only. Our summary lists: A repository with the bench-swe pipeline; Network access to GitHub for cloning the target repo.

Does SWE Benchmark Task Adder access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is SWE Benchmark Task Adder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does SWE Benchmark Task Adder use?

SWE Benchmark Task Adder has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does SWE Benchmark Task Adder use?

About 497 tokens (SKILL.md is roughly 2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to SWE Benchmark Task Adder?

Skills that share tags, products or a category with SWE Benchmark Task Adder: Octocode Benchmark Runner (bgauryy/octocode, 949 stars), Octocode Graph Eval Loop (bgauryy/octocode, 949 stars), Windmill AI Evals (windmill-labs/windmill, 18k stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains SWE Benchmark Task Adder?

ory (a GitHub organization) maintains it in ory/lumen, which has 305 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on August 11, 2026.

Source: ory/lumen on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.