Official agent skill

Score Pyrefly Change

by facebook in facebook/pyrefly

Produces a scorecard evaluating a pyrefly change on correctness and quality.

OfficialMITAuto-check passedDevelopment

Install Score Pyrefly Change

skills CLI
$ npx skills add facebook/pyrefly --skill score-pyrefly-change -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebook/pyrefly score-pyrefly-change --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebook/pyrefly.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/score-pyrefly-change .claude/skills/score-pyrefly-change && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
score-pyrefly-change
GitHub stars
7.1k
Token cost
~1.5k tokens
SKILL.md length
827 words
Files
4 (incl. scripts, assets)
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Produces a scorecard evaluating a pyrefly change on correctness and quality.

  • Works in 4 steps: Explain the change → Fill out the scorecard → Summarize → …
  • Development work in your project
  • SKILL.md covers 1. Explain the change, 2. Fill out the scorecard, 3. Summarize and 4. Present
  • Runs Python scripts from its folder

What it does

Score Pyrefly Change is an agent skill from facebook/pyrefly, published by the product's own GitHub organization. Produces a scorecard evaluating a pyrefly change on correctness and quality.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and assets (for example `agents/openai.yaml`, `assets/template.md` and `scripts/score.py`).

It sits in Development. It works with Python. The repository describes itself as: A fast type checker and language server for Python. The licence is MIT.

When your agent uses it

  • Development work in your project

Example prompts

  • “Use the score-pyrefly-change skill to produce a scorecard evaluating a pyrefly change on correctness and quality”
  • “/score-pyrefly-change”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Explain the change
  2. Fill out the scorecard
  3. Summarize
  4. Present

What it can do on your machine

Read from SKILL.md and the folder at commit 6bc6ea9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Score Pyrefly Change loads about 1.5k tokens when it runs. Until then it costs about 24 tokens; SKILL.md has 827 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~24
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from facebook/pyrefly at commit 6bc6ea9, republished under its MIT licence (© facebook). 827 words, ~1,455 tokens.

Download SKILL.mdSave it as .claude/skills/score-pyrefly-change/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
score-pyrefly-change
description
Produces a scorecard evaluating a pyrefly change on correctness and quality.
disable-model-invocation
true

Evaluate this change, executing the following checklist in this exact order.

1. Explain the change

Record the exact change you are evaluating, and explain what problem it solves and how in plain language. Briefly describe the core building blocks, using full relative paths.

2. Fill out the scorecard

The scorecard consists of weighted sections with concrete dimensions. For each dimension, fill out:

  • A verdict of YES, NO, UNKNOWN, or N/A.
  • Evidence for a YES or NO verdict. For UNKNOWN, describe what data you're missing. Score N/A only when a dimension is structurally irrelevant, such as Code Quality for a documentation-only change.
Design (40%)
  • D1: Does the change address a real problem in Pyrefly?
    • A bug fix requires evidence that Pyrefly's behavior conflicts with the typing specification, Python runtime semantics, the LSP specification, or established behavior in other tools.
    • A new feature requires evidence that the feature is in-scope for Pyrefly, such as a linked GitHub issue created or confirmed by a maintainer.
    • Judge other changes, such as code refactors and documentation updates, on whether they improve the project on balance.
    • Never score this dimension as N/A.

If the problem cannot be demonstrated by a reproducer, then score N/A for the remaining design dimensions.

Otherwise, using D1's evidence, develop a minimal neutral reproducer for the problem that does not include the author's characterization or desired mechanism. Give this reproducer and the change's base revision to two fresh-context sub-agents, along with these restrictions:

  • Do not switch or modify the shared working copy.
  • Inspect only the base revision, using read-only revision-aware commands.
  • Do not read the change's metadata or contents. Ask each sub-agent to determine:
  1. The underlying logic bug, design flaw, or missing piece that causes the problem - i.e., the "root cause".
  2. What the correct behavior should be on the minimal reproducer and 1 additional representative example for the root cause.
  3. The core idea that a principled fix should use, independent of specific implementation strategy.
  4. Where in the type checker/language server life cycle the fix should live. Give each sub-agent the same prompt, and record that exact prompt.

Map the sub-agents' responses to 1-4 to D2-D5, respectively. Score UNKNOWN for any design dimension on which the sub-agents disagree.

  • D2: Does the change target the root cause?
  • D3: Does the change produce the correct behavior on the minimal reproducer?
    • Ignore the additional representative examples.
  • D4: Does the change use the correct core idea?
  • D5: Is the change implemented at the correct point(s) in the type checker/language server life cycle?
Show full SKILL.md (408 more words)Show less
Implementation (30%)

If the change modifies code, review the change adversarially.

  • Do not assume the description, tests, or apparent local correctness prove the approach.
  • Identify the change's central invariant, derive a matrix of expected behaviors, and actively try to falsify it with minimal counterexamples.
  • If sub-agents were used, include their additional representative examples here. Record all counterexamples discovered.

Always score I1 and I2. For non-code changes, interpret I1 as whether the changed content is accurate and internally consistent.

  • I1: Does the change implement its chosen semantics correctly?
    • Consider only counterexamples that demonstrate an incorrect behavior or side effect of changed or added code.
    • Distinguish incorrect from imprecise: score YES if the change faithfully implements what it claims to, even if more precise semantics are possible.
  • I2: Does the change fully implement its intended scope?
    • Consider only counterexamples that demonstrate an unintentionally incomplete implementation of the change description.
  • I3: Is the change's mypy_primer delta free of regressions?
Validation (10%)
  • V1: Is the change's core functionality validated?
    • Automated tests are preferable, but manual validation documented in the change description is acceptable when automated testing is infeasible.
  • V2: Is coverage robust?
    • Score YES if a majority of changed and new code paths are validated.
    • Exhaustive edge case coverage is not required.
Code Quality (10%)
  • C1: Is the control flow easy to follow?
  • C2: Does the code use existing abstractions and helpers instead of reinventing them or duplicating code?
  • C3: Does the code make appropriate use of concise, readable comments to explain non-obvious aspects?
    • Code without comments can score YES if the code does not need comments.
    • 2+ unnecessary or poorly written comments should score NO.
Reviewability (10%)
  • R1: Is the change focused?
    • Score NO if the change does multiple things that could be cleanly separated into self-contained changes.
  • R2: Is the change small?
    • Score YES if the change is split into pieces - for example, commits in a large PR - that each satisfy this dimension.
    • Each piece should be no more than approximately 150 LOC, not including tests.
  • R3: Does the change have a clear, informative title?
  • R4: Does the change have a clear, informative description?
    • An unnecessarily verbose description, such as raw LLM output, should score NO.

3. Summarize

Use scripts/score.py to compute per-section and overall quality scores. Report the scores, and briefly describe what the change does well and what can be improved.

4. Present

Present your findings using assets/template.md. Do not modify your findings while compiling them for presentation.

© facebook, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, assets) in .agents/skills/score-pyrefly-change of facebook/pyrefly.

  • SKILL.md
  • agents/openai.yaml
  • assets/template.md
  • scripts/score.py

Open the folder on GitHubat commit 6bc6ea9

Compare with similar skills

Score Pyrefly Change next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Score Pyrefly Change compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Score Pyrefly Change this skillfacebook/pyrefly7.1k—~1.5kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
Summarise Ecosystem Resultsastral-sh/ruff50k1 repos~2.2kAutomated safety check: PassMIT
Minimizing Ty Ecosystem Changesastral-sh/ruff50k—~4.6kAutomated safety check: PassMIT
Merge Dependabot PRsonyx-dot-app/onyx32k1 repos~2.2kAutomated safety check: PassMIT
Code Graph Mermaid Diagramstrailofbits/skills7.4k1 repos~1.7kAutomated safety check: PassCC-BY-SA-4.0

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Official

    A skill your agent uses when a user says "summarise ecosystem results", "summarize this ty ecosystem report", "what changed in this ecosystem run?", or asks to summarise or summarize ty ecosystem…

    50k GitHub starsUsed in 1 repo~2.2k tokens
    DevelopmentAuto-check passed
  • Official

    A skill your agent uses when a user says "minimize this ty ecosystem change", "reproduce this ecosystem result", "investigate a primer difference", "investigate a mypyprimer difference"…

    50k GitHub stars~4.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Merge Dependabot PRs

    onyx-dot-app/onyx

    Triages and lands a batch of open Dependabot PRs in the Onyx repo, where main is gated exclusively by GitHub's merge queue: approves and enqueues green PRs, closes superseded duplicates, fixes…

    32k GitHub starsUsed in 1 repo~2.2k tokens
    DevelopmentAuto-check passed
  • Code Graph Mermaid Diagrams

    trailofbits/skills

    Official

    Generates Mermaid diagrams from Trailmark code graphs, including call graphs, class hierarchies, module dependency maps, complexity heatmaps and attack surface data flows.

    7.4k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • CLI Developer

    Jeffallan/claude-skills

    Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.

    12k GitHub starsUsed in 2 repos~1.2k tokens
    DevelopmentAuto-check passed

More from facebook/pyrefly

  • Add Torch Shapes Example

    facebook/pyrefly

    Official

    A skill your agent uses when adding a new PyTorch model to Pyrefly's shape-tracking example corpus under tensor-shapes/pyrefly-torch-stubs/examples — i.e.

    7.1k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Benchmark Pyrefly

    facebook/pyrefly

    Official

    Run Pyrefly benchmarks locally via Buck or Cargo, including PyTorch real-world LSP benchmarks.

    7.1k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Modify Shaped Array Dsl

    facebook/pyrefly

    Official

    A skill your agent uses when Pyrefly computes a wrong tensor shape (or is missing one that can't be expressed in a stub signature) and you need to add or fix a shape-DSL rule.

    7.1k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Official

    Port a PyTorch model to use pyrefly's tensor shape type system (Tensor[[B, C, H, W]], Int[T]).

    7.1k GitHub stars~13k tokensUpdated today
    Auto-check passed
  • Review Pyrefly Diff

    facebook/pyrefly

    Official

    Reviews a comma separated pyrefly diff according to the pyrefly review best practices.

    7.1k GitHub stars~863 tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Score Pyrefly Change

What does Score Pyrefly Change do?

Produces a scorecard evaluating a pyrefly change on correctness and quality. Score Pyrefly Change is an agent skill from facebook/pyrefly, published by the product's own GitHub organization. Produces a scorecard evaluating a pyrefly change on correctness and quality.

When should I use Score Pyrefly Change?

Score Pyrefly Change fits situations like: development work in your project.

How do I install Score Pyrefly Change in Claude Code?

Run `npx skills add facebook/pyrefly --skill score-pyrefly-change -a claude-code`. Or copy the skill folder (.agents/skills/score-pyrefly-change in facebook/pyrefly) into .claude/skills/score-pyrefly-change in your project. Claude Code loads it when a task matches its description.

How do I install Score Pyrefly Change in Codex?

Run `npx skills add facebook/pyrefly --skill score-pyrefly-change -a codex`. Or copy the skill folder (.agents/skills/score-pyrefly-change in facebook/pyrefly) into .agents/skills/score-pyrefly-change in your project. Codex loads it when a task matches its description.

Can I use Score Pyrefly Change in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebook/pyrefly --skill score-pyrefly-change -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/score-pyrefly-change, .gemini/skills/score-pyrefly-change, .github/skills/score-pyrefly-change and .opencode/skills/score-pyrefly-change in your project.

What does Score Pyrefly Change need to run?

Going by SKILL.md and its folder, Score Pyrefly Change needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Score Pyrefly Change access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Score Pyrefly Change safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Score Pyrefly Change use?

Score Pyrefly Change is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Score Pyrefly Change use?

About 1.5k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Score Pyrefly Change?

Skills that share tags, products or a category with Score Pyrefly Change: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), Summarise Ecosystem Results (astral-sh/ruff, 50k stars), Minimizing Ty Ecosystem Changes (astral-sh/ruff, 50k stars) and Merge Dependabot PRs (onyx-dot-app/onyx, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Score Pyrefly Change?

facebook (a GitHub organization, an official publisher) maintains it in facebook/pyrefly, which has 7,054 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 8, 2026.

Source: facebook/pyrefly on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.