Agent skill

Estimate Calibrator

by Mathews-Tom in Mathews-Tom/armory

Produces calibrated three-point PERT estimates (best/likely/worst) with confidence intervals, unknowns, and assumptions.

MITAuto-check passedData & Analytics

Install Estimate Calibrator

skills CLI
$ npx skills add Mathews-Tom/armory --skill estimate-calibrator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory estimate-calibrator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/estimate-calibrator .claude/skills/estimate-calibrator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
estimate-calibrator
GitHub stars
329
Token cost
~2.1k tokens
SKILL.md length
640 words
Files
6 (incl. references)
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

Produces calibrated three-point PERT estimates (best/likely/worst) with confidence intervals, unknowns, and assumptions.

  • Works in 5 steps: Decompose Work → Three-Point Estimate → Identify Unknowns → …
  • : estimate this
  • SKILL.md covers Reference Files, Prerequisites, Workflow and Output Format, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Estimate Calibrator is an agent skill from Mathews-Tom/armory. Produces calibrated three-point PERT estimates (best/likely/worst) with confidence intervals, unknowns, and assumptions. Triggers on: "estimate this", "how long will this take", "effort estimate", "confidence interval", "story points", "t-shirt sizing". NOT for task decomposition, use task-decomposer.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `evals/cases.yaml`, `references/calibration-tips.md` and `references/estimation-methods.md`).

It sits in Data & Analytics, covering Statistics and Task breakdown. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • : estimate this
  • How long will this take
  • Effort estimate
  • Confidence interval

Example prompts

  • “estimate this”
  • “how long will this take”
  • “effort estimate”
  • “/estimate-calibrator”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Decompose Work
  2. Three-Point Estimate
  3. Identify Unknowns
  4. Calculate Ranges
  5. Assign Confidence

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Estimate Calibrator loads about 2.1k tokens when it runs, and up to ~7.2k if it reads all its reference files. Until then it costs about 81 tokens; SKILL.md has 640 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 640 words, ~2,054 tokens.

Download SKILL.mdSave it as .claude/skills/estimate-calibrator/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
estimate-calibrator
description
Produces calibrated three-point PERT estimates (best/likely/worst) with confidence intervals, unknowns, and assumptions. Triggers on: "estimate this", "how long will this take", "effort estimate", "confidence interval", "story points", "t-shirt sizing". NOT for task decomposition, use task-decomposer.
metadata.version
1.1.1
metadata.category
development
metadata.tags
estimation, pert, confidence-interval, planning
metadata.difficulty
intermediate
metadata.phase
plan

Estimate Calibrator

Replaces single-point guesses with structured three-point estimates: decomposes work into atomic units, estimates best/likely/worst case for each, identifies unknowns and assumptions, calculates aggregate ranges using PERT, and assigns confidence levels with explicit rationale.

Reference Files

FileContentsLoad When
references/estimation-methods.mdPERT formula, three-point estimation, Monte Carlo basicsAlways
references/unknown-categories.mdTechnical, scope, external, and organizational uncertainty typesUnknown identification
references/calibration-tips.mdCognitive biases in estimation, historical calibration, buffer strategiesAlways
references/sizing-heuristics.mdCommon task size patterns, complexity indicators, reference class dataQuick sizing needed

Prerequisites

  • Work item description (feature, task, project)
  • Decomposed tasks (or use task-decomposer skill first)
  • Context: team familiarity, tech stack, existing codebase

Workflow

Phase 1: Decompose Work

If the work item is not already decomposed into atomic units:

  1. Break into tasks — Each task should be estimable independently.
  2. Right granularity — Tasks should be 1 hour to 3 days. Larger tasks have higher uncertainty; break them down further.
  3. Identify dependencies — Tasks on the critical path determine the minimum duration.
Phase 2: Three-Point Estimate

For each task, estimate three scenarios:

ScenarioDefinitionMindset
Best caseEverything goes right. No surprises."If I've done this exact thing before"
Likely caseNormal friction. Some minor obstacles."Realistic expectation with typical setbacks"
Worst caseSignificant problems. Not catastrophic."Murphy's law but not a disaster"

Key rule: Worst case is NOT "everything goes wrong." It's the realistic bad scenario (90th percentile), not the apocalyptic one (99th percentile).

Phase 3: Identify Unknowns

Categorize unknowns that affect estimates:

CategoryExampleImpact
Technical"Never used this library before"Likely case inflated, worst case much higher
Scope"Requirements may change"All estimates may shift
External"Depends on API access from partner"Blocking risk — could delay entirely
Integration"Haven't tested with production data"Hidden complexity at integration
Organizational"Need design approval"Calendar time, not effort time
Phase 4: Calculate Ranges

For individual tasks, use the PERT formula:

text
Expected = (Best + 4 × Likely + Worst) / 6
Std Dev = (Worst - Best) / 6

For aggregate (project) estimates:

  • Sum of expected values for total expected duration
  • Root sum of squares of std devs for aggregate uncertainty
Phase 5: Assign Confidence
ConfidenceMeaningWhen
HighLikely case within ±20%Well-understood task, team has done it before
MediumLikely case within ±50%Some unknowns, moderate familiarity
LowLikely case within ±100% or moreSignificant unknowns, new technology

Output Format

text
## Estimate: {Work Item}

### Summary
| Scenario | Duration |
|----------|----------|
| Best case | {time} |
| Likely case | {time} |
| Worst case | {time} |
| **PERT expected** | **{time}** |
| **Confidence** | **{High/Medium/Low}** |

### Task-Level Estimates

| # | Task | Best | Likely | Worst | PERT | Unknowns |
|---|------|------|--------|-------|------|----------|
| 1 | {task} | {time} | {time} | {time} | {time} | {key unknown or "None"} |
| 2 | {task} | {time} | {time} | {time} | {time} | {key unknown} |
| | **Total** | **{sum}** | **{sum}** | **{sum}** | **{pert}** | |

### Key Unknowns

| # | Unknown | Category | Impact on Estimate | Mitigation |
|---|---------|----------|-------------------|------------|
| 1 | {unknown} | {Technical/Scope/External} | +{time} if realized | {spike, prototype, early test} |

### Assumptions
- {Assumption 1 — what must be true for this estimate to hold}
- {Assumption 2}

### Risk Factors
- {Risk}: If realized, adds {time}. Likelihood: {High/Medium/Low}.

### Confidence Rationale
**{High/Medium/Low}** because:
- {Specific reason — e.g., "Team has built 3 similar features"}
- {Specific reason — e.g., "External API is a new integration"}

### Recommendation
{Commit to PERT expected with {X}% buffer, or spike the top unknown first.}
Show full SKILL.md (267 more words)Show less

Calibration Rules

  1. Three points, not one. Single-point estimates are always wrong. Three points communicate uncertainty — the most important part of any estimate.
  2. Worst case is the 90th percentile, not the 99th. "Asteroid hits the office" is not a useful worst case. "The API documentation is wrong and we need to reverse-engineer the protocol" is realistic worst case.
  3. Unknowns inflate estimates more than known difficulty. A hard but well-understood task is more predictable than an easy but novel one.
  4. Estimates are not commitments. Communicate ranges, not deadlines. If stakeholders need a single number, give the PERT expected plus a buffer for confidence level.
  5. Spike unknowns early. If a single unknown dominates the estimate range, invest 1-2 days spiking it before estimating the rest.

Error Handling

ProblemResolution
Work item not decomposedDecompose into 3-8 tasks first (or suggest task-decomposer skill).
No historical referenceEstimate relative to a known task: "This is about 2x the auth feature."
Stakeholder wants a single numberProvide PERT expected with buffer matching confidence level (High: +20%, Medium: +50%, Low: +100%).
Estimate seems too largeCheck for scope creep in task list. Remove non-essential tasks. Identify what can be deferred.
Team has never done this type of workMark confidence as Low. Recommend a spike before committing to an estimate.

When NOT to Estimate

Push back if:

  • The work is exploratory (research, spikes) — timebox instead of estimating
  • Requirements are completely undefined — define scope first
  • The user wants precision (hours) for a large project — provide ranges, not false precision
  • The estimate will be used as a commitment without acknowledging uncertainty

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/estimate-calibrator of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml
  • references/calibration-tips.md
  • references/estimation-methods.md
  • references/sizing-heuristics.md
  • references/unknown-categories.md

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

Estimate Calibrator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Estimate Calibrator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Estimate Calibrator this skillMathews-Tom/armory329—~2.1kAutomated safety check: PassMIT
Querying Indonesian Gov Datasuryast/indonesia-gov-apis172—~997Automated safety check: PassMIT
Sandbox Benchvercel/next.js143k—~4.1kAutomated safety check: PassMIT
Statistical Analysisspacering-net/codeg3.9k3 repos~5kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.7k15 repos~4.9kAutomated safety check: PassBSD-3-Clause
AI Daily DigestvigorX777/ai-daily-digest1.6k—~1.3kAutomated safety check: PassNone

Similar skills

  • Querying Indonesian Gov Data

    suryast/indonesia-gov-apis

    Query 57 Indonesian government APIs and data sources — BPJPH halal certification, BPOM food safety, OJK financial legality, BPS statistics, BMKG weather/earthquakes, Bank Indonesia exchange rates…

    172 GitHub stars~997 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Sandbox Bench

    vercel/next.js

    Official

    Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

    143k GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Statistical Analysis

    spacering-net/codeg

    Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting.

    3.9k GitHub starsUsed in 3 repos~5k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 15 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • AI Daily Digest

    vigorX777/ai-daily-digest

    Fetches RSS feeds from 90 top Hacker News blogs (curated by Karpathy), uses AI to score and filter articles, and generates a daily digest in Markdown with Chinese-translated titles, category…

    1.6k GitHub stars~1.3k tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Statistical Power

    spacering-net/codeg

    Sample-size and statistical power calculations for planning studies.

    3.9k GitHub starsUsed in 1 repo~3.6k tokens
    Data & AnalyticsAuto-check: notes

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    329 GitHub stars~4.6k tokensUpdated 4 days ago
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    329 GitHub stars~2.6k tokensUpdated 4 days ago
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    329 GitHub stars~2.8k tokensUpdated 4 days ago
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    329 GitHub stars~3.1k tokensUpdated 4 days ago
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    329 GitHub stars~4.9k tokensUpdated 4 days ago
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    329 GitHub stars~2.7k tokensUpdated 4 days ago
    Auto-check passed

Questions about Estimate Calibrator

What does Estimate Calibrator do?

Produces calibrated three-point PERT estimates (best/likely/worst) with confidence intervals, unknowns, and assumptions. Estimate Calibrator is an agent skill from Mathews-Tom/armory. Produces calibrated three-point PERT estimates (best/likely/worst) with confidence intervals, unknowns, and assumptions.

When should I use Estimate Calibrator?

Estimate Calibrator fits situations like: : estimate this; how long will this take; effort estimate; confidence interval.

How do I install Estimate Calibrator in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill estimate-calibrator -a claude-code`. Or copy the skill folder (skills/estimate-calibrator in Mathews-Tom/armory) into .claude/skills/estimate-calibrator in your project. Claude Code loads it when a task matches its description.

How do I install Estimate Calibrator in Codex?

Run `npx skills add Mathews-Tom/armory --skill estimate-calibrator -a codex`. Or copy the skill folder (skills/estimate-calibrator in Mathews-Tom/armory) into .agents/skills/estimate-calibrator in your project. Codex loads it when a task matches its description.

Can I use Estimate Calibrator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill estimate-calibrator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/estimate-calibrator, .gemini/skills/estimate-calibrator, .github/skills/estimate-calibrator and .opencode/skills/estimate-calibrator in your project.

What does Estimate Calibrator need to run?

SKILL.md names no scripts, command-line tools or credentials: Estimate Calibrator is instructions for the agent only.

Does Estimate Calibrator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Estimate Calibrator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Estimate Calibrator use?

Estimate Calibrator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Estimate Calibrator use?

About 2.1k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.1k tokens, read only when the agent opens those files.

What are the alternatives to Estimate Calibrator?

Skills that share tags, products or a category with Estimate Calibrator: Querying Indonesian Gov Data (suryast/indonesia-gov-apis, 172 stars), Sandbox Bench (vercel/next.js, 143k stars), Statistical Analysis (spacering-net/codeg, 3.9k stars) and Statsmodels (zLanqing/codex-claude-academic-skills, 4.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Estimate Calibrator?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 329 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.