Agent skill

Reflection Coach

by paperclipai in paperclipai/paperclip

Reflect on another agent's recent execution record and propose the smallest durable instruction, skill, or tool-description change.

MITAuto-check passedSales & Support

Install Reflection Coach

skills CLI
$ npx skills add paperclipai/paperclip --skill reflection-coach -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install paperclipai/paperclip reflection-coach --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/paperclipai/paperclip.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/skills-catalog/catalog/bundled/paperclip-operations/reflection-coach .claude/skills/reflection-coach && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
reflection-coach
GitHub stars
99k
Token cost
~3k tokens
SKILL.md length
1,444 words
Files
1
Skills in repo
60
Repo updated
First seen
Licence
MIT

At a glance

Reflect on another agent's recent execution record and propose the smallest durable instruction, skill, or tool-description change.

  • Works in 10 steps: Confirm target and scope → Pull the recent record → Read the target's current guardrails → …
  • Evidence-backed coaching proposals
  • SKILL.md covers When to use, When not to use, Inputs and Hard guardrails, plus 3 more sections
  • Calls curl; needs PAPERCLIP_API_KEY

What it does

Reflection Coach is an agent skill from paperclipai/paperclip. Reflect on another agent's recent execution record and propose the smallest durable instruction, skill, or tool-description change. Use for evidence-backed coaching proposals, never hot-swaps.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Sales & Support, covering Proposals and quotes. The repository describes itself as: The open-source app everyone uses to manage agents at work. The licence is MIT.

When your agent uses it

  • Evidence-backed coaching proposals
  • Never hot-swaps

Example prompts

  • “/reflection-coach”

Requirements

  • A credential in PAPERCLIP_API_KEY

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Confirm target and scope
  2. Pull the recent record
  3. Read the target's current guardrails
  4. Cluster the failures
  5. Route each cluster to a target surface
  6. Draft the proposal document
  7. Write the actual drafts (files, not just prose)
  8. Benchmark-gate the proposal
  9. Publish and request acceptance
  10. Apply only after acceptance, in a follow-up run

What it can do on your machine

Read from SKILL.md and the folder at commit b9750b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • PAPERCLIP_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Reflection Coach loads about 3k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 1,444 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from paperclipai/paperclip at commit b9750b1, republished under its MIT licence (© paperclipai). 1,444 words, ~2,980 tokens.

Download SKILL.mdSave it as .claude/skills/reflection-coach/SKILL.md (or your agent's skills folder).
name
reflection-coach
description
Reflect on another agent's recent execution record and propose the smallest durable instruction, skill, or tool-description change. Use for evidence-backed coaching proposals, never hot-swaps.
key
paperclipai/bundled/paperclip-operations/reflection-coach
recommendedForRoles
manager, general
tags
paperclip, reflection, coaching, agents, skills

Reflection Coach

You are coaching another agent. You are not that agent. Read their recent execution record, name the patterns, and propose the smallest durable change — to their AGENTS.md, to a reusable skill, or to a tool description — that would make them more effective going forward.

This skill runs on a target agent and produces a reviewable proposal. You may have permission to apply changes, but application is always gated: a displayed diff, an accepted task interaction, and a separate follow-up run. You never propose and apply in the same run.

Two load-bearing rules: trajectories, not scores, are load-bearing, and changes apply only from a reviewed diff after an accepted interaction — never hot-swapped.

When to use

  • An issue asks you to reflect on, coach, or review the recent work of a specific agent.
  • A routine (e.g. recent-agent-reflection) hands you a bounded set of agents to review.
  • Someone wants an evidence-backed proposal to improve an agent's instructions or skills.

When not to use

  • The target agent id is your own. Refuse — no self-reflection.
  • You are asked to rewrite product code or shared infra. That is out of scope.
  • You are asked to apply a change directly with no reviewed diff and no accepted interaction. Refuse and name the gate.

Inputs

Required:

  • targetAgentId — the agent you are coaching. Never coach yourself.
  • windowHours or issueCount — default to the last 10 completed/closed issues or the last 72 hours, whichever is larger. Cap at 25 issues to stay within budget.

Optional:

  • focus — free-text hint ("verification misses", "late escalations"). Bias clustering toward this axis if given.
  • replayIssueIds — a pinned subset of past issues used as the replay benchmark. If absent, pick 3–5 representative recent issues from the window.

Hard guardrails

Every proposal must satisfy all of these:

  • No same-run apply. Discovery and application are separate runs. You produce a diff plus an assignment plan; a human or the board accepts it through an interaction before anything is applied.
  • Size caps. Skills ≤ 15KB. Tool descriptions ≤ 500 chars. AGENTS.md may grow by at most +20% per proposal. Want more? Split proposals.
  • Trajectory-backed or drop it. Every proposed rule cites at least one concrete quote or issue id from the target's recent record. No evidence, no rule.
  • Not your code. Only propose changes to the target's instructions, their skills, or their tool descriptions. Never to code they do not own or to shared infra.
  • Benchmark-gated. Name the replay cases the proposal must still resolve. If a rule would have broken a past success, drop it.
  • No reflection on yourself. If targetAgentId == PAPERCLIP_AGENT_ID, refuse and ask for another coach.

Procedure

1) Confirm target and scope
sh
curl -sS "$PAPERCLIP_API_URL/api/agents/<targetAgentId>" \
  -H "Authorization: Bearer $PAPERCLIP_API_KEY"

Record name, role, reportsTo, adapterType, adapterConfig.instructionsFilePath (where AGENTS.md lives), and current assigned skills via GET /api/agents/<targetAgentId>/skills. Refuse and exit if targetAgentId == $PAPERCLIP_AGENT_ID.

2) Pull the recent record
sh
curl -sS "$PAPERCLIP_API_URL/api/companies/$PAPERCLIP_COMPANY_ID/issues?assigneeAgentId=<targetAgentId>&status=done,in_review,blocked&limit=25" \
  -H "Authorization: Bearer $PAPERCLIP_API_KEY"

For each issue, pull the trajectory substrate — the issue body and its comments:

sh
curl -sS "$PAPERCLIP_API_URL/api/issues/<issueId>" -H "Authorization: Bearer $PAPERCLIP_API_KEY"
curl -sS "$PAPERCLIP_API_URL/api/issues/<issueId>/comments" -H "Authorization: Bearer $PAPERCLIP_API_KEY"

Keep status transitions, blocker reasons, reviewer comments, approval outcomes, human corrections, and PR-link comments. Comments are the closest thing Paperclip has to an execution trace — treat them as first-class evidence.

3) Read the target's current guardrails

Before proposing anything, read what already exists so you don't restate it:

  • Their AGENTS.md at adapterConfig.instructionsFilePath.
  • Their assigned skills (from step 1).
  • Any MEMORY.md / memory/ files in their cwd if the adapter uses para-memory-files.

If a rule you were about to propose is already present, drop it. A failure pattern despite an existing rule is a different finding — record it as "existing rule X is not being followed" and propose how to make it stick (move to a skill, add a negative example, strengthen the trigger), not a duplicate.

4) Cluster the failures

Name each cluster from this taxonomy:

  • verifier-miss — agent claimed done; reviewer rejected.
  • avoidable-rework — same issue reopened more than once.
  • stale-context — acted on an assumption already falsified in-thread.
  • instruction-miss — violated an existing rule in AGENTS.md.
  • late-escalation — stayed blocked too long without escalating.
  • human-correction — a user explicitly said to do X differently.
  • tool-misuse — hit the same tool-error pattern repeatedly.
  • scope-creep — changes beyond task scope.

For each cluster keep a list of (issueId, commentId, one-line evidence quote) tuples. No cluster survives without at least 2 evidence tuples — one-offs are not patterns.

5) Route each cluster to a target surface
  • Agent-specific, narrow, cheap to state → AGENTS.md update. E.g. "always re-run failing tests before marking in_review."
  • Generalizable, multi-step procedure with when-to-use logic → new or updated reusable skill.
  • Both → update/create the skill AND add a pointer line in AGENTS.md so the agent knows when to reach for it. Common case for non-obvious procedures.
  • Tool description → only if the failure was "agent didn't know when to use tool X" and a ≤500-char description change fixes it.

Sanity check reuse honestly: a rule that applies to all coders belongs in a shared skill; a "reusable skill" that only fits one role belongs in that agent's AGENTS.md.

6) Draft the proposal document

Create a document attached to the reflection issue (never the target's issues). One section per cluster:

markdown
## Cluster: <name>

**Pattern (1 sentence, quotable):**
**Root cause hypothesis:**
**Evidence (≥2):**
- [PAP-NNN](/PAP/issues/PAP-NNN) — "<verbatim fragment>"
- [PAP-MMM](/PAP/issues/PAP-MMM) — "<verbatim fragment>"

**Proposed change:**
- Target surface: AGENTS.md | skill:<slug> | both | tool-description:<tool>
- Diff (inline, minimal, ≤20% AGENTS.md growth / ≤15KB skill):
    ```diff
    ...
    ```

**Expected still-passes (replay):**
- [PAP-XXX](/PAP/issues/PAP-XXX), [PAP-YYY](/PAP/issues/PAP-YYY)

**Why this change, not something bigger:**
(1–2 sentences on why you didn't rewrite more.)
7) Write the actual drafts (files, not just prose)
  • Skill surface — draft a full SKILL.md (frontmatter → Overview → When to use → Process → Pitfalls → Verification), ≤ 15KB. Put it under drafts/<skill-slug>/SKILL.md and attach it to the reflection issue.
  • AGENTS.md surface — write a unified diff against the target's current AGENTS.md. Do not rewrite the whole file; quote 1–3 lines of context per change. Keep total growth ≤ +20%; split if you can't.
Show full SKILL.md (554 more words)Show less
8) Benchmark-gate the proposal

For each pinned replay issue, ask: "If this rule had been in effect, would the agent still have succeeded?" Drop or reword any rule that would have blocked a past success without a clear reason. Record the walk in "Expected still-passes." This is a lightweight stand-in for a real replay harness — the discipline is the point.

9) Publish and request acceptance

From a reflection issue (assigned to the target's manager or the requester):

  1. Attach the proposal document: PUT /api/issues/{issueId}/documents/reflection-proposal.
  2. If a draft skill was written, commit it under skills/<skill-slug>/ (or attach it) and link it in the proposal.
  3. Open the acceptance gate with a task interaction on the reflection issue. Mutations that change instructions, skills, or tool descriptions must use request_confirmation, show the diff in payload.detailsMarkdown, set continuationPolicy: wake_assignee_on_accept, and include the exact payload.target.key listed below.
  4. Leave a comment summarizing: target agent, window, clusters found, surfaces touched, link to the proposal, link to the interaction, and the next-step owner.

Server-enforced mutation target keys:

  • Agent instructions: agent:<agentId>:instructions
  • Agent/tool description fields: agent:<agentId>:profile
  • Existing company skill: skill:<skillId>
  • New local company skill by slug: skill-slug:<slug>
  • Imported or catalog skill source: skill-import:<source>
  • Project workspace skill scan: skills:scan-projects
10) Apply only after acceptance, in a follow-up run

When the interaction resolves accepted, apply the change in a separate run:

  • AGENTS.md — update the target's managed instruction file exactly as the accepted diff specified.
  • Skill — install/update the skill in the company library, then POST /api/agents/<targetAgentId>/skills/sync with {"mode":"add","desiredSkills":["<skill-ref>"]} when the target should receive it. Use remove only for the named assignments. Use replace only after explicit confirmation to overwrite the complete desired skill set.
  • Tool description — update the target agent's description/profile field that the accepted diff named.

The server rejects Reflection Coach mutations unless the accepted request_confirmation was created by Reflection Coach in a previous run, has a displayed diff, and is bound to the resource by one of the target keys above. If the interaction was rejected or is still pending, apply nothing. If you were asked to apply without a reviewed diff and an accepted interaction, refuse and name the gate — no-same-run-apply is load-bearing.

Pitfalls

  • Scoring without trajectories. Don't say "failed 3 times" without quoting the failures. Scores alone collapse improvement rate.
  • Proposing the bigger rewrite. Your job is the smallest change that would have prevented the cluster. Bigger feels impressive; it isn't.
  • Duplicating rules the agent already has. Read AGENTS.md + assigned skills first. An existing-but-unfollowed rule is a "make it stick" proposal, not a restatement.
  • Applying in the discovery run. Even with permission, discovery and application are separate runs behind an accepted interaction.
  • Silently expanding scope. The +20% cap exists because every new rule competes for attention. Four small proposals beat one big rewrite.
  • Promising runtime value. You are not improving the agent mid-session. This is offline, diff-reviewed, interaction-gated.

Verification (self-check before publishing)

  • targetAgentId != $PAPERCLIP_AGENT_ID
  • Each cluster has ≥2 evidence tuples with a linked issue + verbatim quote
  • Each proposal names the target surface explicitly and includes the diff (not just prose)
  • AGENTS.md growth ≤ 20%, skills ≤ 15KB, tool descriptions ≤ 500 chars
  • Replay set has ≥3 past issues the rules still pass against
  • Proposal document linked from the reflection issue
  • An acceptance interaction (showing the diff) is open before any mutation
  • No claim that the target has already "been updated" before acceptance + follow-up run

© paperclipai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/skills-catalog/catalog/bundled/paperclip-operations/reflection-coach of paperclipai/paperclip.

Open the folder on GitHubat commit b9750b1

Compare with similar skills

Reflection Coach next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Reflection Coach compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Reflection Coach this skillpaperclipai/paperclip99k—~3kAutomated safety check: PassMIT
Doc Coauthoringaws-samples/sample-strands-agent-with-agentcore19541 repos~3.2kAutomated safety check: PassMIT
Audit Onboarding Proposalhoangnb24/repository-harness1.2k—~4kAutomated safety check: PassMIT
No Negative EchoLB623/no-negative-echo897—~965Automated safety check: PassMIT
GEO Service Proposal Generatorzubair-trabzada/geo-seo-claude11k—~3kAutomated safety check: NotesMIT
Architectural ProposalsFritzAndFriends/SharpSite1452 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Doc Coauthoring

    aws-samples/sample-strands-agent-with-agentcore

    Official

    Guide users through a structured workflow for co-authoring documentation.

    195 GitHub starsUsed in 41 repos~3.2k tokens
    Sales & SupportAuto-check passed
  • Audit Onboarding Proposal

    hoangnb24/repository-harness

    Use only when the user explicitly invokes $audit-onboarding-proposal.

    1.2k GitHub stars~4k tokensUpdated 5 days ago
    Sales & SupportAuto-check passed
  • No Negative Echo

    LB623/no-negative-echo

    Prevent 此地无银三百两式 residue: finalize artifacts without echoing rejected session-only alternatives into labels, metadata, commits, PRs, or handoffs.

    897 GitHub stars~965 tokensUpdated 1 mo ago
    Sales & SupportAuto-check passed
  • GEO Service Proposal Generator

    zubair-trabzada/geo-seo-claude

    Builds a client-ready AI-search-optimization proposal from an existing GEO audit, with pricing tiers, an ROI estimate and a markdown document ready to send.

    11k GitHub stars~3k tokensUpdated today
    Sales & SupportAuto-check: notes
  • Architectural Proposals

    FritzAndFriends/SharpSite

    How to write comprehensive architectural proposals that drive alignment before code is written

    145 GitHub starsUsed in 2 repos~1.6k tokens
    Sales & SupportAuto-check passed
  • Counter Proposal Generator

    zubair-trabzada/ai-legal-claude

    Generates specific counter-proposals for every unfavorable clause, with replacement language, negotiation talking points, and a ready-to-send email template

    1.8k GitHub stars~2.4k tokensUpdated 6 mo ago
    Sales & SupportAuto-check passed

More from paperclipai/paperclip

All 60 skills in this repo
  • Garden Inbox

    paperclipai/paperclip

    Scan a Paperclip user's Mine inbox, classify reversible archive candidates, request checkbox confirmation, and archive only accepted selections.

    99k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Paperclip

    paperclipai/paperclip

    Interact with the Paperclip control plane API for task coordination and governance.

    99k GitHub stars~9.6k tokensUpdated today
    Auto-check passed
  • Paperclip

    paperclipai/paperclip

    A skill your agent uses for Paperclip-managed tasks and heartbeats: reading task context, delivering task documents or files, updating completion or blockers, coordinating or delegating work, and…

    99k GitHub stars~17k tokensUpdated today
    Auto-check passed
  • Design Guide

    paperclipai/paperclip

    Paperclip UI design system guide for building consistent, reusable frontend components.

    99k GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Paperclip Page

    paperclipai/paperclip

    Publish static HTML pages and asset folders to the Paperclip S3/CloudFront page host.

    99k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Paperclip Create Agent

    paperclipai/paperclip

    Create new agents in Paperclip with governance-aware hiring.

    99k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed

Categories

Questions about Reflection Coach

What does Reflection Coach do?

Reflect on another agent's recent execution record and propose the smallest durable instruction, skill, or tool-description change. Reflection Coach is an agent skill from paperclipai/paperclip. Reflect on another agent's recent execution record and propose the smallest durable instruction, skill, or tool-description change.

When should I use Reflection Coach?

Reflection Coach fits situations like: evidence-backed coaching proposals; never hot-swaps.

How do I install Reflection Coach in Claude Code?

Run `npx skills add paperclipai/paperclip --skill reflection-coach -a claude-code`. Or copy the skill folder (packages/skills-catalog/catalog/bundled/paperclip-operations/reflection-coach in paperclipai/paperclip) into .claude/skills/reflection-coach in your project. Claude Code loads it when a task matches its description.

How do I install Reflection Coach in Codex?

Run `npx skills add paperclipai/paperclip --skill reflection-coach -a codex`. Or copy the skill folder (packages/skills-catalog/catalog/bundled/paperclip-operations/reflection-coach in paperclipai/paperclip) into .agents/skills/reflection-coach in your project. Codex loads it when a task matches its description.

Can I use Reflection Coach in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add paperclipai/paperclip --skill reflection-coach -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/reflection-coach, .gemini/skills/reflection-coach, .github/skills/reflection-coach and .opencode/skills/reflection-coach in your project.

What does Reflection Coach need to run?

Going by SKILL.md and its folder, Reflection Coach needs the command-line tools its instructions call (curl) and credentials named PAPERCLIP_API_KEY. Our summary lists: A credential in PAPERCLIP_API_KEY.

Does Reflection Coach access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Reflection Coach safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Reflection Coach use?

Reflection Coach is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Reflection Coach use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Reflection Coach?

Skills that share tags, products or a category with Reflection Coach: Doc Coauthoring (aws-samples/sample-strands-agent-with-agentcore, 195 stars), Audit Onboarding Proposal (hoangnb24/repository-harness, 1.2k stars), No Negative Echo (LB623/no-negative-echo, 897 stars) and GEO Service Proposal Generator (zubair-trabzada/geo-seo-claude, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Reflection Coach?

paperclipai (a GitHub organization) maintains it in paperclipai/paperclip, which has 98,967 GitHub stars. The repository holds 60 skills in this directory. The repository was last updated on October 9, 2026.

Source: paperclipai/paperclip on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.