Agent skill

Reflect

by borghei in borghei/Claude-Skills

Turn reflection into decisions by scoring predictions against outcomes and tracking whether commitments held.

MITAuto-check passedBusiness, Finance & HR

Install Reflect

skills CLI
$ npx skills add borghei/Claude-Skills --skill reflect -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills reflect --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/personal-productivity/reflect .claude/skills/reflect && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
reflect
GitHub stars
881
Token cost
~3.6k tokens
SKILL.md length
1,778 words
Files
11 (incl. scripts, references, assets)
Skills in repo
349
Repo updated
First seen
Licence
MIT

At a glance

Turn reflection into decisions by scoring predictions against outcomes and tracking whether commitments held.

  • Works in 5 steps: Resolve every prediction past its date —… → Run the scorer with 5 buckets (10… → Read the calibration table first: gaps… → …
  • Running a quarterly review
  • SKILL.md covers When to use this skill, Inputs the skill expects, Clarify First and Workflows, plus 3 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Reflect is an agent skill from borghei/Claude-Skills. Turn reflection into decisions by scoring predictions against outcomes and tracking whether commitments held. Use when running a quarterly review, scoring forecast calibration, or reflection keeps producing notes instead of change.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts, reference files and assets (for example `assets/prediction-log-template.md`, `assets/quarterly-reflection-template.md` and `assets/sample_commitments.json`).

It sits in Business, Finance & HR, covering Performance reviews. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Running a quarterly review
  • Scoring forecast calibration
  • Reflection keeps producing notes instead of change

Example prompts

  • “/reflect”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Resolve every prediction past its date — true or false, no revising the
  2. Run the scorer with 5 buckets (10 buckets need roughly 50+ predictions to be
  3. Read the calibration table first: gaps above 10 points in buckets holding 5+
  4. Read the domain breakdown — bias is rarely uniform, and the worst domain is
  5. Pick exactly one correction and re-measure over a full quarter.

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Reflect loads about 3.6k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 60 tokens; SKILL.md has 1,778 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 1,778 words, ~3,559 tokens.

Download SKILL.mdSave it as .claude/skills/reflect/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
reflect
description
Turn reflection into decisions by scoring predictions against outcomes and tracking whether commitments held. Use when running a quarterly review, scoring forecast calibration, or reflection keeps producing notes instead of change.
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
personal-productivity
metadata.domain
personal-effectiveness
metadata.updated
2026-07-21
metadata.tags
reflection, calibration, brier-score, forecasting, commitments

Reflect

Structured reflection at daily, weekly, and quarterly cadence that ends in a changed behaviour rather than a paragraph of feelings. The mechanism is testing recorded beliefs against outcomes: written predictions scored for calibration, and commitments tracked for whether they actually held.

Boundary with weekly-review: that skill runs the weekly operating cadence — what happened, what is next, priorities and blockers. This one is the learning layer on top: was my judgement any good across weeks and quarters, and what should change as a result. Run them back to back, operating review first, using its output as this skill's raw material. Do not duplicate the wins/blockers synthesis here.

When to use this skill

  • Your reviews produce pleasant notes but nothing ever changes as a result
  • You want to know whether to trust your own confidence when making a call
  • Delivery dates keep slipping and you suspect the estimates, not the execution
  • The same commitment has been carried for months without progress or a decision
  • It is quarter end and you need to score judgement, not just report outcomes
  • You are starting a prediction log and want a scoring method rather than a journal

Inputs the skill expects

  • A prediction log — statement, confidence, and outcome once resolved
  • A commitment log — text, status (kept/missed/partial/open), optional due and carried_cycles
  • The cadence you are running: daily, weekly, or quarterly
  • A reference date for overdue calculations (passed explicitly — the tools never read the clock)
  • Optionally a domain per prediction, which is where the most actionable signal appears

Clarify First

Before generating, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Which cadence is being run — daily, weekly, and quarterly ask genuinely different questions; the wrong set produces either triviality or an hour-long session that gets skipped
  • Whether resolved predictions exist — below 20 resolved, calibration is too noisy to act on and the honest output is "keep logging," not a score
  • Whether a weekly operating review already runs — determines whether this layers on top or has to carry the operating cadence too
  • Carry-count history on open commitments — the three-cycle rule is the sharpest mechanic here and needs the count to fire

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Workflows

Workflow 1 — Score prediction calibration

Quarterly. The step that converts reflection from storytelling into measurement.

  1. Resolve every prediction past its date — true or false, no revising the original confidence.
  2. Run the scorer with 5 buckets (10 buckets need roughly 50+ predictions to be readable).
  3. Read the calibration table first: gaps above 10 points in buckets holding 5+ predictions are real bias, not noise.
  4. Read the domain breakdown — bias is rarely uniform, and the worst domain is usually your own delivery dates.
  5. Pick exactly one correction and re-measure over a full quarter.
bash
python3 personal-productivity/reflect/scripts/calibration_scorer.py \
  --input personal-productivity/reflect/assets/sample_predictions.json \
  --buckets 5

The decomposition is computed over distinct confidence values, so --buckets changes only the displayed table and never the statistics. Verify the identity BS = REL - RES + UNC and the other scoring invariants at any time:

bash
python3 personal-productivity/reflect/scripts/calibration_scorer.py --selftest

Isolate a single domain once you know where the bias lives:

bash
python3 personal-productivity/reflect/scripts/calibration_scorer.py \
  --input personal-productivity/reflect/assets/sample_predictions.json \
  --domain delivery --buckets 5 --format json
Workflow 2 — Generate a cadence-appropriate prompt set

Weekly, or at whichever cadence you are running.

  1. Update commitment statuses from the past cycle — kept, missed, partial, open.
  2. Increment carried_cycles on anything that rolled over again.
  3. Run the generator for the cadence, with today's date for overdue detection.
  4. Work the core prompts, then the accountability prompts — each of those needs a decision, not a note.
  5. Satisfy the closing requirement. If nothing changed, the session was journaling.
bash
python3 personal-productivity/reflect/scripts/reflection_prompt_generator.py \
  --input personal-productivity/reflect/assets/sample_commitments.json \
  --cadence weekly --as-of 2026-07-21

Quarterly, where structural change is allowed — note the quarterly log carries carried_cycles, which is what fires the three-cycle rule:

bash
python3 personal-productivity/reflect/scripts/reflection_prompt_generator.py \
  --input personal-productivity/reflect/assets/sample_commitments_quarterly.json \
  --cadence quarterly --as-of 2026-07-21 --format json
Workflow 3 — Run the quarterly reflection end to end

90 minutes, once a quarter. The only cadence where role, commitments, and method are on the table.

  1. Score the prediction log (20 min) — Workflow 1.
  2. Read the quarter's weekly reflections in one sitting (15 min). Individually unremarkable; in a batch they expose patterns invisible at weekly resolution.
  3. Work the quarterly prompts from assets/quarterly-reflection-template.md (20 min).
  4. Decide (20 min): kill one commitment, change one method, and give every chronic commitment an explicit date/delegate/kill decision.
  5. Write next quarter's predictions with confidence numbers (15 min).
bash
python3 personal-productivity/reflect/scripts/reflection_prompt_generator.py \
  --input personal-productivity/reflect/assets/sample_commitments_quarterly.json \
  --cadence quarterly --as-of 2026-07-21

Decision frameworks

Cadence selection
CadenceTimeQuestion it answersSkip it when
Daily5 minWhat did today prove me wrong about?Time is short — this is the optional layer
Weekly20 minDid my commitments hold, and what pattern explains the misses?Never — this is the load-bearing cadence
Quarterly90 minIs my judgement calibrated, and what structural thing must change?Never; it is the only place structural change happens

[PROVEN] Start with weekly only. The most common failure is starting daily because it looks smallest, missing three days, and abandoning everything. Weekly carries most of the value and survives a missed week without collapsing.

Reading a Brier score
BrierReading
< 0.10Excellent — or the predictions were too easy; check the skill score
0.10-0.15Strong
0.15-0.20Good; typical for a practised forecaster on genuinely uncertain questions
0.20-0.25Weak — approaching a coin flip
> 0.25Worse than always saying 50%. Your confidence is actively misleading you

The score decomposes into reliability (miscalibration — lower better) and resolution (discrimination — higher better). The common pattern is decent reliability with near-zero resolution: you have learned to hedge everything to the base rate, which is safe and useless. The reverse — sharp judgement, wrong numbers — is more valuable, because numeric calibration is easy to correct and directional judgement is not.

Commitment keep rate
Keep rateReadingAction
> 85%Under-committing; commitments are safe rather than execution strongCommit to something that might fail
60-85%HealthyContinue
< 60%Committing to more than you deliverCut the number of commitments before trying to improve execution

The instinct at a low keep rate is to try harder, which reliably fails — the cause is volume, not effort. Halve the commitments and the rate usually recovers on its own.

The three-cycle rule [PROVEN]

A commitment carried three cycles without progress is not waiting for time; it is waiting for a decision you keep declining to make. Carrying it a fourth time is the decision — to never do it — so make it explicitly: commit to a date, delegate it, or kill it.

Show full SKILL.md (717 more words)Show less
Where overconfidence concentrates
DomainTypical bias
Your own delivery datesStrongly overconfident — the most reliable bias in professional forecasting
Sales / deal closingOverconfident; role-required optimism leaks in
Competitor timingOverconfident on when, decent on what
Other teams' deliveryBetter calibrated — no inside view to be optimistic with
Metrics and trendsReasonably calibrated; anchored to observable history
Hiring outcomesOften underconfident; rejection memories are salient

If you log only one category, log your own delivery dates: largest bias, fastest resolution cycle, most immediate payoff in planning.

Anti-Patterns

Reflection as Journaling

Mistake: Writing a thoughtful account of how the period went, feeling clarified, and changing nothing. Why it happens: Writing is pleasant and feels productive; deciding is uncomfortable and can be wrong. Given a prompt with no wrong answer — "how did the week go?" — the session drifts to narration. Instead: Require every session to close with a named change, phrased as a rule rather than an intention. "No meetings before 11:00 on Tuesdays and Thursdays" is testable next week; "be better about deep work" can survive years unkept. If a session produces no rule, mark it as journaling in the log — after three consecutive entries, the prompts are wrong and need replacing.

Reflecting Against Memory Instead of Records

Mistake: Asking "was I right about that?" and consulting recollection for the answer. Why it happens: It does not feel like a failure mode. Memory reconstructs rather than replays, and it edits the prior belief toward the known outcome — so you sincerely remember having been less surprised, and having assigned more probability to what happened, than you did. Instead: Write predictions with explicit confidence numbers before outcomes are known, and treat the log as append-only. The number written in advance is the only thing later reflection cannot quietly rewrite. This is precisely why calibration scoring is the core of the practice rather than an optional extra.

Only Predicting Safe Things

Mistake: Filling the prediction log with claims you are already confident about, then reading the excellent Brier score as evidence of good judgement. Why it happens: A bad score feels like a grade, so the log drifts toward things that will score well. Predicting "the sun rises tomorrow" at 99% produces a superb Brier and teaches nothing. Instead: Watch the skill score, which compares you against always predicting the base rate. At or below zero, your forecasts carry no information beyond knowing how often things generally go your way. Deliberately log predictions you might be wrong about — finding your errors is the log's only job, and a log with no errors in it has failed at it.

Scoring Too Often

Mistake: Reviewing calibration monthly or after every significant miss, then adjusting the approach each time. Why it happens: A bad outcome creates an urge to fix something immediately, and the log is right there. Instead: Score quarterly. With 15-30 predictions per quarter, a month yields too few resolved items to distinguish bias from luck, and reacting to that noise produces exactly the thrashing the practice exists to eliminate. Apply one correction per quarter and hold it for a full cycle — changing several things at once means you learn nothing about which of them worked.

Files

FilePurpose
scripts/calibration_scorer.pyCLI entry point: assembles the report (Brier, Murphy decomposition, calibration table, per-domain breakdown, skill score vs base rate, worst-calls list), renders text/JSON, and runs --selftest asserting 12 scoring invariants
scripts/calibration_core.pyScoring internals imported by calibration_scorer.py: log loading, record normalisation, Murphy decomposition (exact — grouped on distinct forecast values, independent of --buckets), bucketing, per-domain stats, and the significance thresholds
scripts/reflection_prompt_generator.pyCadence-specific prompt sets with per-commitment accountability prompts, keep-rate interpretation, and a closing requirement
references/calibration-and-forecasting.mdWriting scoreable predictions, Brier and skill-score interpretation, Murphy decomposition, domain-specific bias table, correction protocols
references/reflection-cadences.mdDaily/weekly/quarterly prompt sets, the boundary with the weekly operating review, three-cycle rule, keep-rate bands, what makes reflection fail
assets/prediction-log-template.mdPrediction-log format, confidence conventions, weekly resolution ritual, quarterly scoring table
assets/quarterly-reflection-template.mdTimed 90-minute quarterly agenda with the kill list, chronic-commitment decisions, and next-quarter predictions
assets/sample_predictions.json30 predictions (27 resolved, 3 pending) showing realistic delivery-date overconfidence, so scoring runs out of the box
assets/sample_commitments.jsonOne week of commitments — 10 items spanning kept/missed/partial/open with two chronic carries, for the weekly cadence
assets/sample_commitments_quarterly.jsonOne quarter of commitments — 18 items with four chronic carries and three overdue, for the quarterly cadence

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts, references, assets) in personal-productivity/reflect of borghei/Claude-Skills.

  • SKILL.md
  • assets/prediction-log-template.md
  • assets/quarterly-reflection-template.md
  • assets/sample_commitments.json
  • assets/sample_commitments_quarterly.json
  • assets/sample_predictions.json
  • references/calibration-and-forecasting.md
  • references/reflection-cadences.md
  • scripts/calibration_core.py
  • scripts/calibration_scorer.py
  • scripts/reflection_prompt_generator.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Reflect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Reflect compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Reflect this skillborghei/Claude-Skills881—~3.6kAutomated safety check: PassMIT
Wp Performance Reviewelvismdev/claude-wordpress-skills2351 repos~4.5kAutomated safety check: PassMIT
Align Humanagentscope-ai/OpenJudge868—~3.1kAutomated safety check: PassApache-2.0
Performance ReportAffitor/affiliate-skills6991 repos~2.5kAutomated safety check: PassMIT
Run Mv Hoi Reconstructionnvidia-isaac/video_to_data850—~1.5kAutomated safety check: PassCustom licence
Company Analysiszhu1090093659/dsh-trading231—~4.2kAutomated safety check: PassCustom licence

Similar skills

  • Wp Performance Review

    elvismdev/claude-wordpress-skills

    WordPress performance code review and optimization analysis.

    235 GitHub starsUsed in 1 repo~4.5k tokens
    Business, Finance & HRAuto-check passed
  • Align Human

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

    868 GitHub stars~3.1k tokensUpdated 27 days ago
    Business, Finance & HRAuto-check passed
  • Performance Report

    Affitor/affiliate-skills

    Generate affiliate performance reports with KPIs and recommendations.

    699 GitHub starsUsed in 1 repo~2.5k tokens
    Business, Finance & HRAuto-check passed
  • Run Mv Hoi Reconstruction

    nvidia-isaac/video_to_data

    Run and validate the repository-local multi-view camera calibration and human-object reconstruction pipelines.

    850 GitHub stars~1.5k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Company Analysis

    zhu1090093659/dsh-trading

    A skill your agent uses when the user wants to analyze a listed company, stock, business, or investment target; challenge or revise an existing company report; compare A/H or primary-listing/ADR…

    231 GitHub stars~4.2k tokensUpdated 5 days ago
    Business, Finance & HRAuto-check passed
  • Windbg Diagnostic Method

    microsoft/win-dev-skills

    Official

    Use with every WinDbg plugin investigation to apply evidence-first reasoning, confidence calibration, contrarian review, structured reporting, and deterministic validation.

    462 GitHub stars~1.9k tokensUpdated today
    Business, Finance & HRAuto-check passed

More from borghei/Claude-Skills

All 349 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    881 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    881 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    881 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Reflect

What does Reflect do?

Turn reflection into decisions by scoring predictions against outcomes and tracking whether commitments held. Reflect is an agent skill from borghei/Claude-Skills. Turn reflection into decisions by scoring predictions against outcomes and tracking whether commitments held.

When should I use Reflect?

Reflect fits situations like: running a quarterly review; scoring forecast calibration; reflection keeps producing notes instead of change.

How do I install Reflect in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill reflect -a claude-code`. Or copy the skill folder (personal-productivity/reflect in borghei/Claude-Skills) into .claude/skills/reflect in your project. Claude Code loads it when a task matches its description.

How do I install Reflect in Codex?

Run `npx skills add borghei/Claude-Skills --skill reflect -a codex`. Or copy the skill folder (personal-productivity/reflect in borghei/Claude-Skills) into .agents/skills/reflect in your project. Codex loads it when a task matches its description.

Can I use Reflect in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill reflect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/reflect, .gemini/skills/reflect, .github/skills/reflect and .opencode/skills/reflect in your project.

What does Reflect need to run?

Going by SKILL.md and its folder, Reflect needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Reflect access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Reflect safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Reflect use?

Reflect is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Reflect use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.3k tokens, read only when the agent opens those files.

What are the alternatives to Reflect?

Skills that share tags, products or a category with Reflect: Wp Performance Review (elvismdev/claude-wordpress-skills, 235 stars), Align Human (agentscope-ai/OpenJudge, 868 stars), Performance Report (Affitor/affiliate-skills, 699 stars) and Run Mv Hoi Reconstruction (nvidia-isaac/video_to_data, 850 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Reflect?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 881 GitHub stars. The repository holds 349 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.