Agent skill

Outcome Tracker

by mohitagw15856 in mohitagw15856/pm-claude-skills

Record the testable predictions inside a decision, then score them against reality later — so frameworks earn trust from outcomes, not vibes.

MITAuto-check passedProduct & Project Management

Install Outcome Tracker

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill outcome-tracker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills outcome-tracker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/outcome-tracker .claude/skills/outcome-tracker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
outcome-tracker
GitHub stars
1.4k
Token cost
~1.6k tokens
SKILL.md length
740 words
Files
2 (incl. scripts)
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Record the testable predictions inside a decision, then score them against reality later — so frameworks earn trust from outcomes, not vibes.

  • Committing to a prioritisation
  • SKILL.md covers What This Skill Produces, Required Inputs, Making Claims Falsifiable… and Scoring (review mode), plus 6 more sections
  • Runs Python scripts from its folder; calls python3
  • Plan (to log what it predicts)

What it does

Outcome Tracker is an agent skill from mohitagw15856/pm-claude-skills. Record the testable predictions inside a decision, then score them against reality later — so frameworks earn trust from outcomes, not vibes. Use when committing to a prioritisation, forecast, or plan (to log what it predicts), when asked to review what actually happened, or to compute how well-calibrated past RICE scores, forecasts, or bets have been. Produces a prediction record at decision time, and a calibration report with per-framework hit rates at review time.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/outcome_calibration.py`).

It sits in Product & Project Management, covering Prioritization frameworks and Performance reviews. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Committing to a prioritisation
  • Plan (to log what it predicts)
  • Asked to review what actually happened
  • Compute how well-calibrated past RICE scores

Example prompts

  • “/outcome-tracker”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Outcome Tracker loads about 1.6k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 740 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 740 words, ~1,564 tokens.

Download SKILL.mdSave it as .claude/skills/outcome-tracker/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
outcome-tracker
description
Record the testable predictions inside a decision, then score them against reality later — so frameworks earn trust from outcomes, not vibes. Use when committing to a prioritisation, forecast, or plan (to log what it predicts), when asked to review what actually happened, or to compute how well-calibrated past RICE scores, forecasts, or bets have been. Produces a prediction record at decision time, and a calibration report with per-framework hit rates at review time.

Outcome Tracker Skill

Every prioritisation, forecast, and launch plan makes predictions — then everyone forgets to check them. This skill closes the loop: extract the predictions at decision time, park them somewhere durable, and score them against reality on a schedule. Over time it answers the question no one can answer today: which of our frameworks actually predict outcomes?

What This Skill Produces

  • At decision time: a prediction record — each claim made falsifiable, with a metric, a direction/target, a check-by date, and a stated confidence
  • At review time: an outcome scoring of due predictions (hit / miss / partial / unresolvable), with what was learned
  • On demand: a calibration report — per-framework and per-confidence-band hit rates from the accumulated records

Required Inputs

Ask for (if not already provided):

  • Mode — record (new decision), review (score due predictions), or calibrate (analyse the history)
  • Record mode: the decision artifact (RICE table, forecast, launch plan, OKR set) and where records live (a predictions/ folder in the Brain, or a JSON/markdown file in the repo)
  • Review mode: the stored predictions plus current metric values for the due ones
  • Calibrate mode: the prediction history (the calculator below reads it as JSON)

Making Claims Falsifiable (record mode)

Walk the artifact and force each implicit claim into this shape — a prediction that can't fill the row doesn't get recorded, it gets flagged as untestable:

FieldRule
claimOne sentence, future tense, about a measurable effect ("onboarding redesign lifts activation")
metricThe exact instrumented metric, with today's baseline
predictedDirection + magnitude band ("+10-20% relative") — bands beat point estimates
confidence0.5–0.95, from the author, recorded before the outcome is knowable
check_byThe date the effect should be visible if real; also the review trigger
frameworkWhat produced the claim (rice-prioritisation, gut call, sales-forecasting-model…) — this is what calibration is about

Typical yields: a RICE table → one prediction per top-3 item (impact claims); a forecast → the quarter's number; a launch plan → its success metrics; an OKR set → each KR's target.

Scoring (review mode)

For each prediction past its check_by: hit (actual within the predicted band), partial (right direction, wrong magnitude), miss (wrong direction or no effect), unresolvable (metric never instrumented, or confounded by a simultaneous change — record why; a pile of unresolvables is itself a finding about how the team instruments its bets). Never rescore or reinterpret the original claim to make it a hit — the record is append-only.

Programmatic Helper

scripts/outcome_calibration.py (stdlib-only) computes the calibration report from a JSON array of prediction records:

bash
python3 scripts/outcome_calibration.py predictions.json
echo '[{"framework":"rice-prioritisation","confidence":0.8,"outcome":"hit"}]' | python3 scripts/outcome_calibration.py -

It reports per-framework hit rates (hits + half-credit partials over resolved), per-confidence-band calibration (do 80%-confidence claims land ~80% of the time?), and flags overconfident bands. Use the computed numbers; don't estimate them.

Show full SKILL.md (302 more words)Show less

Brain Integration

If a professional-brain (brain/) exists, records live in brain/predictions/<id>.md (one file per prediction, fields as frontmatter, [hunch]/[data] provenance on the baseline) and review mode starts by listing files with check_by in the past. Pair with schedule-recipe to run review mode monthly — outcome tracking only works as a ritual, not an intention.

Output Format

Record mode:

Predictions registered: [decision] — [date]
#ClaimMetric (baseline)PredictedConfidenceCheck byFramework
Untestable claims flagged: [claim → what instrumentation would make it testable]

Review mode:

Outcome review — [date]
#ClaimPredictedActualOutcomeLearning
Now due next: [next check_by dates]

Calibrate mode: the calculator's report plus 2-3 sentences of interpretation — which framework has earned trust, where the team is overconfident, and the single instrumentation fix that would resolve the most unresolvables.

Quality Checks

  • Every recorded prediction has all six fields — no "improve activation" without a metric, band, and date
  • Confidence was stated before the outcome was knowable, never backfilled
  • Review scored every due prediction, including the embarrassing ones — no silent skips
  • Unresolvables carry a reason, and the calibration report counts them separately from misses
  • Calibration numbers come from the calculator, not estimation

Anti-Patterns

  • Do not reinterpret a claim after the fact so it scores as a hit — the original wording is the contract
  • Do not record point estimates when the author thinks in ranges — bands are honest, points are theatre
  • Do not let a framework take credit for hits and blame "execution" for misses — score the prediction as made
  • Do not compute calibration on fewer than ~10 resolved predictions per framework — report "insufficient history" instead
  • Do not skip recording because the decision feels obvious — obvious bets that miss are the most valuable calibration data

Example Trigger Phrases

  • "Log what this plan predicts."
  • "Review what actually happened."
  • "Score last quarter's forecast against reality."
  • "Did our prioritisation framework work?"

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/outcome-tracker of mohitagw15856/pm-claude-skills.

  • SKILL.md
  • scripts/outcome_calibration.py

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Outcome Tracker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Outcome Tracker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Outcome Tracker this skillmohitagw15856/pm-claude-skills1.4k—~1.6kAutomated safety check: PassMIT
Bio Clinical Databases Variant PrioritizationGPTomics/bioSkills1.2k2 repos~6.3kAutomated safety check: PassMIT
Decision Matrixiflytek/skillhub5.2k—~2.3kAutomated safety check: PassMIT
High Output Management Frameworkamplitude/builder-skills159—~1.9kAutomated safety check: PassNone
Feature Investment Advisordeanpeters/Product-Manager-Skills7.2k1 repos~5.6kAutomated safety check: PassCustom licence
Customizeanthropics/claude-for-legal9.6k2 repos~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Prioritizes rare-disease variants from trio/quad WES/WGS with de novo (DeNovoGear, Triodenovo), compound-heterozygous phasing (WhatsHap), mosaic VAF tiering, phenotype-driven ranking (Exomiser…

    1.2k GitHub starsUsed in 2 repos~6.3k tokens
    Product & Project ManagementAuto-check passed
  • Decision Matrix

    iflytek/skillhub

    Compares options with pros and cons, weighted scoring, pre-mortems, opportunity costs and ICE prioritization, labeling assumptions instead of inventing numbers.

    5.2k GitHub stars~2.3k tokensUpdated today
    Product & Project ManagementAuto-check passed
  • High Output Management Framework

    amplitude/builder-skills

    Applies Andy Grove's High Output Management ideas to diagnose where a team loses output and to design meetings, processes and management decisions.

    159 GitHub stars~1.9k tokensUpdated 2 mo ago
    Product & Project ManagementAuto-check passed
  • Feature Investment Advisor

    deanpeters/Product-Manager-Skills

    Evaluates whether a feature deserves investment by weighing revenue connection, cost structure, ROI and strategic value, then gives a build or don't-build recommendation.

    7.2k GitHub starsUsed in 1 repo~5.6k tokens
    Product & Project ManagementAuto-check passed
  • Customize

    anthropics/claude-for-legal

    Official

    Guided customization of your product counsel practice profile — change one thing without re-running the whole cold-start interview.

    9.6k GitHub starsUsed in 2 repos~1.1k tokens
    Product & Project ManagementAuto-check passed
  • Research Finance

    borghei/Claude-Skills

    Research budgeting and funding operations — study budget construction, cost per participant and per insight, burn against milestones, and portfolio prioritisation by decision value.

    891 GitHub stars~3.4k tokensUpdated 3 days ago
    Product & Project ManagementAuto-check passed

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about Outcome Tracker

What does Outcome Tracker do?

Record the testable predictions inside a decision, then score them against reality later — so frameworks earn trust from outcomes, not vibes. Outcome Tracker is an agent skill from mohitagw15856/pm-claude-skills. Record the testable predictions inside a decision, then score them against reality later — so frameworks earn trust from outcomes, not vibes.

When should I use Outcome Tracker?

Outcome Tracker fits situations like: committing to a prioritisation; plan (to log what it predicts); asked to review what actually happened; compute how well-calibrated past RICE scores.

How do I install Outcome Tracker in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill outcome-tracker -a claude-code`. Or copy the skill folder (skills/outcome-tracker in mohitagw15856/pm-claude-skills) into .claude/skills/outcome-tracker in your project. Claude Code loads it when a task matches its description.

How do I install Outcome Tracker in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill outcome-tracker -a codex`. Or copy the skill folder (skills/outcome-tracker in mohitagw15856/pm-claude-skills) into .agents/skills/outcome-tracker in your project. Codex loads it when a task matches its description.

Can I use Outcome Tracker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill outcome-tracker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/outcome-tracker, .gemini/skills/outcome-tracker, .github/skills/outcome-tracker and .opencode/skills/outcome-tracker in your project.

What does Outcome Tracker need to run?

Going by SKILL.md and its folder, Outcome Tracker needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Outcome Tracker access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Outcome Tracker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Outcome Tracker use?

Outcome Tracker is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Outcome Tracker use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Outcome Tracker?

Skills that share tags, products or a category with Outcome Tracker: Bio Clinical Databases Variant Prioritization (GPTomics/bioSkills, 1.2k stars), Decision Matrix (iflytek/skillhub, 5.2k stars), High Output Management Framework (amplitude/builder-skills, 159 stars) and Feature Investment Advisor (deanpeters/Product-Manager-Skills, 7.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Outcome Tracker?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,434 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 9, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.