Agent skill

Capability Horizon Estimator

by tokenbender in tokenbender/agent-guides

Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.

Apache-2.0Auto-check passed

Install Capability Horizon Estimator

skills CLI
$ npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tokenbender/agent-guides capability-horizon-estimator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude-skills/capability-horizon-estimator .claude/skills/capability-horizon-estimator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
capability-horizon-estimator
GitHub stars
367
Token cost
~1.5k tokens
SKILL.md length
656 words
Files
5 (incl. scripts, references)
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.

  • Scoping agent work
  • SKILL.md covers When to use, The core model, Output format and Calibration rules, plus 1 more section
  • Runs Python scripts from its folder
  • Deciding if a task is within reach

What it does

Capability Horizon Estimator is an agent skill from tokenbender/agent-guides. Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling. Use when scoping agent work, deciding if a task is within reach, planning retries/parallelism, estimating wall-clock time for SWE/MLE/math tasks, or answering "can you do this" / "how long will this take" for autonomous work.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `README.md`, `references/benchmark-catalog.md` and `references/horizon-data.md`).

The repository describes itself as: one page guides that i let my subscribed/customised agents consume to perform actions. The licence is Apache-2.0.

When your agent uses it

  • Scoping agent work
  • Deciding if a task is within reach
  • Planning retries/parallelism
  • Estimating wall-clock time for SWE/MLE/math tasks

Example prompts

  • “can you do this”
  • “how long will this take”
  • “/capability-horizon-estimator”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit a74dd9d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Capability Horizon Estimator loads about 1.5k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 656 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tokenbender/agent-guides at commit a74dd9d, republished under its Apache-2.0 licence (© tokenbender). 656 words, ~1,482 tokens.

Download SKILL.mdSave it as .claude/skills/capability-horizon-estimator/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
capability-horizon-estimator
description
Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling. Use when scoping agent work, deciding if a task is within reach, planning retries/parallelism, estimating wall-clock time for SWE/MLE/math tasks, or answering "can you do this" / "how long will this take" for autonomous work.

Capability & Time Horizon Estimator

Estimate what a model can do and how long it takes using the METR time-horizon framework: every task has a human-equivalent duration $t$, every model has a 50% time horizon $h_{50}$, and success probability follows a logistic curve in $\log(h/t)$.

When to use

  • "Can you do X autonomously?" — give a calibrated yes/maybe/no instead of vibes
  • "How long will this agent run take?" — wall-clock estimate with retry overhead
  • Scoping: should this be one task, or decomposed?
  • Deciding attempt budget (best-of-N) for a stretch task

The core model

Step 1 — Estimate human-equivalent task time $t$ (minutes). Anchor against known reference points (full table in references/benchmark-catalog.md):

Reference taskHuman time
SWE-bench Verified issue7 min – 2 h
HCAST task bands15 min / 1 h / 4 h / 8 h
RE-Bench ML research task8 h
Small bug fix, clear repro15–60 min
Multi-file feature2–8 h
Cross-repo refactor4–16 h
Kaggle competition (MLE-bench)days of human effort
FrontierMath T4 problemdays–weeks of expert time

Estimate for a low-context professional (new hire, contractor), not the resident expert — that is what the horizons are calibrated against.

Step 2 — Get the model's horizon $h_{50}$ (minutes). Look it up in references/horizon-data.md (METR TH v1.1, May 2026). If the model isn't listed, extrapolate from release date:

$$h_{50}(\text{date}) = h_{50}(\text{ref}) \times 2^{(\text{date} - \text{ref}) / D}, \quad D \approx 130\text{–}190 \text{ days}$$

Use $D = 150$ days as default; state the range. Post-2024 data supports faster ($\sim$90–130 days); all-time average is $\sim$190.

Step 3 — Success probability.

$$p = \sigma!\big(\beta \cdot \ln(h_{50}/t)\big), \quad \beta \approx 0.8 \text{ (range 0.6–0.9)}, \quad \sigma(x) = \frac{1}{1+e^{-x}}$$

Sanity anchors with $\beta = 0.8$: $t = h_{50} \Rightarrow p = 50%$ · $t = h_{50}/5.7 \Rightarrow p = 80%$ · $t = 2h_{50} \Rightarrow p \approx 36%$ · $t = 4h_{50} \Rightarrow p \approx 25%$.

Step 4 — Difficulty adjustments (multiply $t$ before Step 3).

ConditionMultiplier on $t$
Well-specified, auto-verifiable (benchmark-like)×1
Proprietary/large codebase, high prior context needed×2
Messy, underspecified, human-judged success×4 (range ×2–8)
GUI/computer-use (no API/DOM)×50–100 — different capability cluster
Parallelizable/additive work (N independent items)Not one task — split, apply model per piece

Rationale: METR's suite is clean and self-contained; models do measurably worse on messy tasks and under holistic (vs algorithmic) scoring. GUI horizons measured ~40–100× shorter than SWE/reasoning.

Step 5 — "Pushed itself" uplift.

  • Better scaffold/elicitation: ×1.5–2 on $h_{50}$ (scaffold choice moves measured horizons ~2×)
  • Best-of-N independent attempts: $P(\geq 1 \text{ success}) = 1 - (1-p)^N$
  • Both compose. A $p = 0.25$ task with a good scaffold ($\times 1.75 h$) and 6 attempts: $p \to \sigma(0.8 \ln(1.75 \cdot h/t)) \approx 0.37$ per attempt, $1-(1-0.37)^6 \approx 94%$.

Step 6 — Wall-clock time. Agents complete tasks they can do several times faster than the human baseline:

$$\text{wall-clock} \approx \frac{t_{\text{human}}}{S} \times \text{expected attempts}, \quad S \approx 3\text{–}8\times \text{ for SWE (default 5)}$$

Expected attempts $= 1/p$ for independent retries (cap at 3–5 in practice; if it fails 3×, the plan is wrong, not the luck). Report a range, not a point.

Show full SKILL.md (177 more words)Show less

Output format

TASK: <one line>
HUMAN-EQUIVALENT TIME (t): <estimate> (anchor: <reference task>)
MODEL HORIZON (h50): <value> (source: table / extrapolated from <date>, D=150d)
DIFFICULTY: ×<mult> (<reason>)  → effective t = <value>
SINGLE-ATTEMPT SUCCESS: <p> → <verdict>
PUSHED (scaffold ×1.75, N=<n>): <p_eff> per attempt, <P> overall
WALL-CLOCK ESTIMATE: <range> (t/S × attempts, S=5)

VERDICT: CONFIDENT (≥80%) | LIKELY (60–80%) | COIN-FLIP (35–60%) | STRETCH (15–35%) | OUT OF REACH (<15%)
CAVEATS: <extrapolation flags, messiness, jagged-domain warnings>
DECOMPOSITION ADVICE: <if stretch/out-of-reach: where to split>

Calibration rules

  • Always state the anchor. Every $t$ estimate names the reference task it was derived from. No anchor, no number.
  • Report ranges. Horizons have wide CIs (Opus 4.6 p50: 317–3634 min). Give the interval, pick the geometric mean for the headline.
  • Above 16 h human-equivalent, say so. METR's suite saturates at 16 h; estimates beyond it are extrapolation, not measurement.
  • Jaggedness is real. A model strong in SWE may be 50–100× weaker on GUI work, and different on math research. Never transfer a horizon across domains without the multiplier.
  • Long ≠ hard, hard ≠ long. The model applies to coherent, non-parallelizable tasks. 1000 independent 1-hour problems is a 1-hour task done 1000 times.
  • Failure mode asymmetry. Near the horizon, models fail by derailing (repeating failed approaches, losing state), not by slow progress. If $p < 50%$, recommend decomposition over retries.

Reference files

  • references/horizon-data.md — METR TH v1.1 per-model p50/p80 table (2019 → May 2026), doubling-time stats, methodology caveats
  • references/benchmark-catalog.md — SWE/MLE/math benchmark inventory with task times, human budgets, and July-2026 SOTA
  • scripts/estimate.py — calculator for Steps 3–6 (logistic p, best-of-N, wall-clock)

© tokenbender, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in claude-skills/capability-horizon-estimator of tokenbender/agent-guides.

  • SKILL.md
  • README.md
  • references/benchmark-catalog.md
  • references/horizon-data.md
  • scripts/estimate.py

Open the folder on GitHubat commit a74dd9d

Compare with similar skills

Capability Horizon Estimator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Capability Horizon Estimator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Capability Horizon Estimator this skilltokenbender/agent-guides367—~1.5kAutomated safety check: PassApache-2.0
Product Capabilityaffaan-m/ECC275k2 repos~1.1kAutomated safety check: PassMIT
Horizon Trackruvnet/ruflo74k—~744Automated safety check: NotesMIT
Artifact Capabilitiesasgeirtj/system_prompts_leaks69k—~4.3kAutomated safety check: PassCC0-1.0
Configuring Horizoncoollabsio/coolify63k4 repos~898Automated safety check: PassMIT
Task Effort EstimatorDonchitos/Claude-Code-Game-Studios26k—~1.2kAutomated safety check: PassMIT

Similar skills

  • Product Capability

    affaan-m/ECC

    Translate PRD intent, roadmap asks, or product discussions into an implementation-ready capability plan that exposes constraints, invariants, interfaces, and unresolved decisions before…

    275k GitHub starsUsed in 2 repos~1.1k tokens
    Product & Project ManagementAuto-check passed
  • Horizon Track

    ruvnet/ruflo

    Track long-horizon objectives across multiple sessions with milestone checkpoints, progress persistence, and drift detection

    74k GitHub stars~744 tokensUpdated yesterday
    Product & Project ManagementAuto-check: notes
  • Artifact Capabilities

    asgeirtj/system_prompts_leaks

    Runtime capabilities a published Artifact page can be granted — behavior static HTML cannot provide on its own, such as the page reading live or connected data, remembering what people do on it (a…

    69k GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Configuring Horizon

    coollabsio/coolify

    A skill your agent uses whenever the user mentions Horizon by name in a Laravel context.

    63k GitHub starsUsed in 4 repos~898 tokens
    Backend & APIsAuto-check passed
  • Task Effort Estimator

    Donchitos/Claude-Code-Game-Studios

    Estimates the effort for a game development task from code complexity, scope, risk and past sprint data, returning a range with a confidence level.

    26k GitHub stars~1.2k tokensUpdated 9 days ago
    Product & Project ManagementAuto-check passed
  • Product Capability

    affaan-m/ECC

    将PRD意图、路线图需求或产品讨论转化为可实施的方案计划,在开始多服务工作之前暴露约束、不变性、接口和未解决的决策。当用户需要ECC原生的PRD到SRS通道,而不是模糊的规划文本时使用。

    275k GitHub stars~439 tokensUpdated 4 days ago
    Product & Project ManagementAuto-check passed

More from tokenbender/agent-guides

All 11 skills in this repo
  • Technical Writing Workflow

    tokenbender/agent-guides

    A skill your agent uses for planning, researching, drafting, revising, or auditing technical write-ups, textbooks, papers, reports, READMEs, research notes, PR narratives, and public technical prose.

    367 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Epistemic Libido

    tokenbender/agent-guides

    Filter, compare, and rank papers, posts, captures, threads, bookmarks, product claims, or research ideas for high-entropy mechanistic insight and underpriced leverage.

    367 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Manim Math Explainer

    tokenbender/agent-guides

    Trigger when: (1) the user asks for Manim, Manim Community, or ManimCE, (2) code contains from manim import , or (3) the task is to build a mathematical explainer animation.

    367 GitHub stars~799 tokensUpdated 2 mo ago
    Auto-check passed
  • Worklog

    tokenbender/agent-guides

    Issue-led atomic work logging. An agent skill from tokenbender/agent-guides.

    367 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • X Thread Reader

    tokenbender/agent-guides

    A skill your agent uses when the user provides an X/Twitter status URL and needs the full thread, context beyond the first post, comparison, summary, intent analysis, title extraction, or reliable…

    367 GitHub stars~506 tokensUpdated 2 mo ago
    Auto-check passed
  • Audit Sft Data Quality

    tokenbender/agent-guides

    Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.

    367 GitHub stars~2.7k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Capability Horizon Estimator

What does Capability Horizon Estimator do?

Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling. Capability Horizon Estimator is an agent skill from tokenbender/agent-guides. Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.

When should I use Capability Horizon Estimator?

Capability Horizon Estimator fits situations like: scoping agent work; deciding if a task is within reach; planning retries/parallelism; estimating wall-clock time for SWE/MLE/math tasks.

How do I install Capability Horizon Estimator in Claude Code?

Run `npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a claude-code`. Or copy the skill folder (claude-skills/capability-horizon-estimator in tokenbender/agent-guides) into .claude/skills/capability-horizon-estimator in your project. Claude Code loads it when a task matches its description.

How do I install Capability Horizon Estimator in Codex?

Run `npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a codex`. Or copy the skill folder (claude-skills/capability-horizon-estimator in tokenbender/agent-guides) into .agents/skills/capability-horizon-estimator in your project. Codex loads it when a task matches its description.

Can I use Capability Horizon Estimator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/capability-horizon-estimator, .gemini/skills/capability-horizon-estimator, .github/skills/capability-horizon-estimator and .opencode/skills/capability-horizon-estimator in your project.

What does Capability Horizon Estimator need to run?

Going by SKILL.md and its folder, Capability Horizon Estimator needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Capability Horizon Estimator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Capability Horizon Estimator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Capability Horizon Estimator use?

Capability Horizon Estimator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Capability Horizon Estimator use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Capability Horizon Estimator?

Skills that share tags, products or a category with Capability Horizon Estimator: Product Capability (affaan-m/ECC, 275k stars), Horizon Track (ruvnet/ruflo, 74k stars), Artifact Capabilities (asgeirtj/system_prompts_leaks, 69k stars) and Configuring Horizon (coollabsio/coolify, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Capability Horizon Estimator?

tokenbender (a GitHub user) maintains it in tokenbender/agent-guides, which has 367 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on July 23, 2026.

Source: tokenbender/agent-guides on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.