Product Capability
affaan-m/ECC
Translate PRD intent, roadmap asks, or product discussions into an implementation-ready capability plan that exposes constraints, invariants, interfaces, and unresolved decisions before…
Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.
$ npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install tokenbender/agent-guides capability-horizon-estimator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude-skills/capability-horizon-estimator .claude/skills/capability-horizon-estimator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "capability-horizon-estimator" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/capability-horizon-estimator into .claude/skills/capability-horizon-estimator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "capability-horizon-estimator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/tokenbender/agent-guides/tree/main/claude-skills/capability-horizon-estimatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install tokenbender/agent-guides capability-horizon-estimator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .agents/skills && cp -r skills-src/claude-skills/capability-horizon-estimator .agents/skills/capability-horizon-estimator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "capability-horizon-estimator" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/capability-horizon-estimator into .agents/skills/capability-horizon-estimator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "capability-horizon-estimator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install tokenbender/agent-guides capability-horizon-estimator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/claude-skills/capability-horizon-estimator .cursor/skills/capability-horizon-estimator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "capability-horizon-estimator" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/capability-horizon-estimator into .cursor/skills/capability-horizon-estimator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "capability-horizon-estimator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/tokenbender/agent-guides.git --path claude-skills/capability-horizon-estimator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install tokenbender/agent-guides capability-horizon-estimator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/claude-skills/capability-horizon-estimator .gemini/skills/capability-horizon-estimator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "capability-horizon-estimator" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/capability-horizon-estimator into .gemini/skills/capability-horizon-estimator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "capability-horizon-estimator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install tokenbender/agent-guides capability-horizon-estimatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .github/skills && cp -r skills-src/claude-skills/capability-horizon-estimator .github/skills/capability-horizon-estimator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "capability-horizon-estimator" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/capability-horizon-estimator into .github/skills/capability-horizon-estimator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "capability-horizon-estimator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install tokenbender/agent-guides capability-horizon-estimator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokenbender/agent-guides.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/claude-skills/capability-horizon-estimator .opencode/skills/capability-horizon-estimator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "capability-horizon-estimator" agent skill from https://github.com/tokenbender/agent-guides/tree/main/claude-skills/capability-horizon-estimator into .opencode/skills/capability-horizon-estimator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "capability-horizon-estimator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
capability-horizon-estimatorEstimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.
Capability Horizon Estimator is an agent skill from tokenbender/agent-guides. Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling. Use when scoping agent work, deciding if a task is within reach, planning retries/parallelism, estimating wall-clock time for SWE/MLE/math tasks, or answering "can you do this" / "how long will this take" for autonomous work.
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `README.md`, `references/benchmark-catalog.md` and `references/horizon-data.md`).
The repository describes itself as: one page guides that i let my subscribed/customised agents consume to perform actions. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit a74dd9d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Capability Horizon Estimator loads about 1.5k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 656 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from tokenbender/agent-guides at commit a74dd9d, republished under its Apache-2.0 licence (© tokenbender). 656 words, ~1,482 tokens.
.claude/skills/capability-horizon-estimator/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Estimate what a model can do and how long it takes using the METR time-horizon framework: every task has a human-equivalent duration $t$, every model has a 50% time horizon $h_{50}$, and success probability follows a logistic curve in $\log(h/t)$.
Step 1 — Estimate human-equivalent task time $t$ (minutes).
Anchor against known reference points (full table in references/benchmark-catalog.md):
| Reference task | Human time |
|---|---|
| SWE-bench Verified issue | 7 min – 2 h |
| HCAST task bands | 15 min / 1 h / 4 h / 8 h |
| RE-Bench ML research task | 8 h |
| Small bug fix, clear repro | 15–60 min |
| Multi-file feature | 2–8 h |
| Cross-repo refactor | 4–16 h |
| Kaggle competition (MLE-bench) | days of human effort |
| FrontierMath T4 problem | days–weeks of expert time |
Estimate for a low-context professional (new hire, contractor), not the resident expert — that is what the horizons are calibrated against.
Step 2 — Get the model's horizon $h_{50}$ (minutes).
Look it up in references/horizon-data.md (METR TH v1.1, May 2026). If the model isn't listed, extrapolate from release date:
$$h_{50}(\text{date}) = h_{50}(\text{ref}) \times 2^{(\text{date} - \text{ref}) / D}, \quad D \approx 130\text{–}190 \text{ days}$$
Use $D = 150$ days as default; state the range. Post-2024 data supports faster ($\sim$90–130 days); all-time average is $\sim$190.
Step 3 — Success probability.
$$p = \sigma!\big(\beta \cdot \ln(h_{50}/t)\big), \quad \beta \approx 0.8 \text{ (range 0.6–0.9)}, \quad \sigma(x) = \frac{1}{1+e^{-x}}$$
Sanity anchors with $\beta = 0.8$: $t = h_{50} \Rightarrow p = 50%$ · $t = h_{50}/5.7 \Rightarrow p = 80%$ · $t = 2h_{50} \Rightarrow p \approx 36%$ · $t = 4h_{50} \Rightarrow p \approx 25%$.
Step 4 — Difficulty adjustments (multiply $t$ before Step 3).
| Condition | Multiplier on $t$ |
|---|---|
| Well-specified, auto-verifiable (benchmark-like) | ×1 |
| Proprietary/large codebase, high prior context needed | ×2 |
| Messy, underspecified, human-judged success | ×4 (range ×2–8) |
| GUI/computer-use (no API/DOM) | ×50–100 — different capability cluster |
| Parallelizable/additive work (N independent items) | Not one task — split, apply model per piece |
Rationale: METR's suite is clean and self-contained; models do measurably worse on messy tasks and under holistic (vs algorithmic) scoring. GUI horizons measured ~40–100× shorter than SWE/reasoning.
Step 5 — "Pushed itself" uplift.
Step 6 — Wall-clock time. Agents complete tasks they can do several times faster than the human baseline:
$$\text{wall-clock} \approx \frac{t_{\text{human}}}{S} \times \text{expected attempts}, \quad S \approx 3\text{–}8\times \text{ for SWE (default 5)}$$
Expected attempts $= 1/p$ for independent retries (cap at 3–5 in practice; if it fails 3×, the plan is wrong, not the luck). Report a range, not a point.
TASK: <one line>
HUMAN-EQUIVALENT TIME (t): <estimate> (anchor: <reference task>)
MODEL HORIZON (h50): <value> (source: table / extrapolated from <date>, D=150d)
DIFFICULTY: ×<mult> (<reason>) → effective t = <value>
SINGLE-ATTEMPT SUCCESS: <p> → <verdict>
PUSHED (scaffold ×1.75, N=<n>): <p_eff> per attempt, <P> overall
WALL-CLOCK ESTIMATE: <range> (t/S × attempts, S=5)
VERDICT: CONFIDENT (≥80%) | LIKELY (60–80%) | COIN-FLIP (35–60%) | STRETCH (15–35%) | OUT OF REACH (<15%)
CAVEATS: <extrapolation flags, messiness, jagged-domain warnings>
DECOMPOSITION ADVICE: <if stretch/out-of-reach: where to split>references/horizon-data.md — METR TH v1.1 per-model p50/p80 table (2019 → May 2026), doubling-time stats, methodology caveatsreferences/benchmark-catalog.md — SWE/MLE/math benchmark inventory with task times, human budgets, and July-2026 SOTAscripts/estimate.py — calculator for Steps 3–6 (logistic p, best-of-N, wall-clock)© tokenbender, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts, references) in claude-skills/capability-horizon-estimator of tokenbender/agent-guides.
Open the folder on GitHubat commit a74dd9d
Capability Horizon Estimator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Capability Horizon Estimator this skilltokenbender/agent-guides | 367 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Product Capabilityaffaan-m/ECC | 275k | 2 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Horizon Trackruvnet/ruflo | 74k | — | ~744 | Automated safety check: Notes | MIT | |
| Artifact Capabilitiesasgeirtj/system_prompts_leaks | 69k | — | ~4.3k | Automated safety check: Pass | CC0-1.0 | |
| Configuring Horizoncoollabsio/coolify | 63k | 4 repos | ~898 | Automated safety check: Pass | MIT | |
| Task Effort EstimatorDonchitos/Claude-Code-Game-Studios | 26k | — | ~1.2k | Automated safety check: Pass | MIT |
affaan-m/ECC
Translate PRD intent, roadmap asks, or product discussions into an implementation-ready capability plan that exposes constraints, invariants, interfaces, and unresolved decisions before…
ruvnet/ruflo
Track long-horizon objectives across multiple sessions with milestone checkpoints, progress persistence, and drift detection
asgeirtj/system_prompts_leaks
Runtime capabilities a published Artifact page can be granted — behavior static HTML cannot provide on its own, such as the page reading live or connected data, remembering what people do on it (a…
coollabsio/coolify
A skill your agent uses whenever the user mentions Horizon by name in a Laravel context.
Donchitos/Claude-Code-Game-Studios
Estimates the effort for a game development task from code complexity, scope, risk and past sprint data, returning a range with a confidence level.
affaan-m/ECC
将PRD意图、路线图需求或产品讨论转化为可实施的方案计划,在开始多服务工作之前暴露约束、不变性、接口和未解决的决策。当用户需要ECC原生的PRD到SRS通道,而不是模糊的规划文本时使用。
tokenbender/agent-guides
A skill your agent uses for planning, researching, drafting, revising, or auditing technical write-ups, textbooks, papers, reports, READMEs, research notes, PR narratives, and public technical prose.
tokenbender/agent-guides
Filter, compare, and rank papers, posts, captures, threads, bookmarks, product claims, or research ideas for high-entropy mechanistic insight and underpriced leverage.
tokenbender/agent-guides
Trigger when: (1) the user asks for Manim, Manim Community, or ManimCE, (2) code contains from manim import , or (3) the task is to build a mathematical explainer animation.
tokenbender/agent-guides
Issue-led atomic work logging. An agent skill from tokenbender/agent-guides.
tokenbender/agent-guides
A skill your agent uses when the user provides an X/Twitter status URL and needs the full thread, context beyond the first post, comparison, summary, intent analysis, title extraction, or reliable…
tokenbender/agent-guides
Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.
Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling. Capability Horizon Estimator is an agent skill from tokenbender/agent-guides. Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling.
Capability Horizon Estimator fits situations like: scoping agent work; deciding if a task is within reach; planning retries/parallelism; estimating wall-clock time for SWE/MLE/math tasks.
Run `npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a claude-code`. Or copy the skill folder (claude-skills/capability-horizon-estimator in tokenbender/agent-guides) into .claude/skills/capability-horizon-estimator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a codex`. Or copy the skill folder (claude-skills/capability-horizon-estimator in tokenbender/agent-guides) into .agents/skills/capability-horizon-estimator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tokenbender/agent-guides --skill capability-horizon-estimator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/capability-horizon-estimator, .gemini/skills/capability-horizon-estimator, .github/skills/capability-horizon-estimator and .opencode/skills/capability-horizon-estimator in your project.
Going by SKILL.md and its folder, Capability Horizon Estimator needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Capability Horizon Estimator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Capability Horizon Estimator: Product Capability (affaan-m/ECC, 275k stars), Horizon Track (ruvnet/ruflo, 74k stars), Artifact Capabilities (asgeirtj/system_prompts_leaks, 69k stars) and Configuring Horizon (coollabsio/coolify, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
tokenbender (a GitHub user) maintains it in tokenbender/agent-guides, which has 367 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on July 23, 2026.
Source: tokenbender/agent-guides on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.