Official agent skill

Improve Skill Quality

by dotnet in dotnet/skills

Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement".

OfficialMITAuto-check passedDevelopment

Install Improve Skill Quality

skills CLI
$ npx skills add dotnet/skills --skill improve-skill-quality -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dotnet/skills improve-skill-quality --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dotnet/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/improve-skill-quality .claude/skills/improve-skill-quality && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
improve-skill-quality
GitHub stars
5.6k
Used in
1 other repo
Token cost
~3.6k tokens
SKILL.md length
1,908 words
Files
3 (incl. references)
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement".

  • Works in 9 steps: Get the evidence before forming a… → Classify the failure → Rule out harness and reliability causes → …
  • An evaluation verdict is a regression
  • SKILL.md covers When to Use, When Not to Use, Inputs and Workflow, plus 3 more sections
  • Calls python, dotnet and git

What it does

Improve Skill Quality is an agent skill from dotnet/skills, published by the product's own GitHub organization. Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Use when an evaluation verdict is a regression or underpowered, when a skill regressed after a change, when /evaluate reports no results, or when deciding whether a weak skill should be strengthened or retired. Do not use for scaffolding a brand-new skill (use create-skill) or a brand-new eval (use create-skill-test).

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/eval-triage.md` and `references/writing-for-baseline-delta.md`).

It sits in Development, covering Project scaffolding. It works with .NET. The repository describes itself as: Repository for skills to assist AI coding agents with .NET and C. The licence is MIT.

When your agent uses it

  • An evaluation verdict is a regression
  • A skill regressed after a change
  • /evaluate reports no results
  • Deciding whether a weak skill should be strengthened

Example prompts

  • “no credible improvement”
  • “Use the improve-skill-quality skill to diagnose and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate…”
  • “/improve-skill-quality”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Get the evidence before forming a hypothesis
  2. Classify the failure
  3. Rule out harness and reliability causes
  4. Verify the fixtures before touching the skill
  5. Check whether the eval could ever have passed
  6. Check whether the two arms differ at all
  7. Fix skill content against the losing trial
  8. Fix activation
  9. Re-validate

What it can do on your machine

Read from SKILL.md and the folder at commit 8d670fa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • dotnet
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Improve Skill Quality loads about 3.6k tokens when it runs, and up to ~8.1k if it reads all its reference files. Until then it costs about 125 tokens; SKILL.md has 1,908 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~125
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dotnet/skills at commit 8d670fa, republished under its MIT licence (© dotnet). 1,908 words, ~3,619 tokens.

Download SKILL.mdSave it as .claude/skills/improve-skill-quality/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
improve-skill-quality
description
Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Use when an evaluation verdict is a regression or underpowered, when a skill regressed after a change, when /evaluate reports no results, or when deciding whether a weak skill should be strengthened or retired. Do not use for scaffolding a brand-new skill (use create-skill) or a brand-new eval (use create-skill-test).

Improve Skill Quality

Turn a failing or unconvincing evaluation into a targeted fix. The single most common mistake in this repo is rewriting skill prose in response to a verdict whose real cause was the eval, the fixtures, or the harness. Classify first, then fix.

When to Use

  • An evaluation verdict is a regression, underpowered, or "no credible improvement".
  • A skill wins in the isolated arm but not in the plugin arm, or is reported "not activated".
  • /evaluate reports "Evaluation ran but produced no results".
  • A skill scores well but costs too much (tokens, turns, wall time, plugin menu budget).
  • Deciding whether to strengthen or retire a persistently weak skill.

When Not to Use

  • Creating a new skill from scratch — use create-skill.
  • Creating a new eval.yaml from scratch — use create-skill-test.
  • Changing the harness itself (eng/skill-validator, eng/vally-adapter, evaluation*.yml).

Inputs

InputRequiredDescription
Verdict evidenceYesThe /evaluate PR comment, or results.json from the run artifacts
Losing trial transcriptsYes for content fixesBaseline vs. skilled output plus the judge's stated reason
Stimulus-vote W/T/L and repeated-run W/T/LYesSeparates cross-task evidence from reliability
Activation status per armYesIsolated and plugin activation are different failures

Workflow

Step 1: Get the evidence before forming a hypothesis

Read InvestigatingResults.md for how to download artifacts and read results.json. Extract, per failing stimulus:

  • authoritative stimulus-vote W/T/L and separate repeated-run W/T/L
  • activation status in the isolated and plugin arms, separately
  • the judge's verbatim reason on each losing trial
  • whether any trial errored, timed out, or produced empty output

Do not change skill content until you can quote a losing trial and the judge's reason for it. For the other cause classes the evidence is different: harness failures are diagnosed from the job log and the spec, and power problems from the trial record — neither has a losing trial to quote, and demanding one is what sends people rewriting prose instead.

Step 2: Classify the failure

Work down this table and stop at the first row that matches. Rows are ordered by how often the symptom has been misdiagnosed as a skill-content problem — the fixture row is first because a fixture failure also presents as a setup or reliability failure and gets misfiled as one.

SymptomReal cause classGo to
A fixture does not build, is untracked by git, breaks for the wrong reason, or contradicts itselfFixtureStep 4
No results.json, "produced no results", or the spec never loadedHarness / spec-loadStep 3
Trials errored, timed out, or returned empty outputReliabilityStep 3
Trajectories unmatched, a trial errored, or the summary disagrees — verdict reported inconclusiveReliability (not power)Step 3
Positive record (e.g. 16W/8T/1L), comparison conclusive, verdict still not a passStatistical powerStep 5
Skilled arm equals baseline arm by constructionEval designStep 6
Activated and lost on quality, judge names a concrete defectSkill contentStep 7
Activated in isolation, not in pluginActivation / routingStep 8
Not activated in either armFrontmatter descriptionStep 8
Wins but costs far more than baselineScope and costStep 7

A verdict is only a measured result when the comparison was conclusive: adapt.mjs requires zero errored trials, zero unmatched trajectories, and an agreeing summary before it will report a pass or a regression. Confirm that before reading a record as a power problem.

Step 3: Rule out harness and reliability causes

See references/eval-triage.md for the full catalogue. The recurring ones:

  • The repository gate rejects the deprecated top-level config: alias, and Vally rejects a spec that declares both config: and defaults:. Replace the alias with one defaults: block.
  • An errored trial is not automatically a fixture problem — judge-side auth and session.idle failures look identical from the verdict and need harness fixes, not SDK pins.
  • expect_tools: [bash] on an advisory question forces a restore or build and turns an answer into a timeout with no quality gain.
  • Genuine code-generation stimuli need roughly 360s; a timeout yields empty output, which fails every grader and hides the real quality signal.
  • Unmatched trajectories, an errored trial, or a summary that disagrees make the comparison inconclusive: the remaining matched trials are biased, so the record is not a measured null and must not be read as a power or content problem.
  • A generic YAML parser is not the production loader. If Vally rejects a skill eval, or the native SDK lane rejects an agent eval's executable scenario fields, classify it as harness / spec-load before changing content. The native agent parser ignores golden references, so validate those separately with check_eval_quality.py and deterministic golden-workspace replay.
  • A pass with one worker or an enlarged local timeout is not normal execution evidence. Reproduce with the repository's normal concurrency and declared suite budget; a failure there is a reliability defect.
Step 4: Verify the fixtures before touching the skill

Run python eng/eval-quality/check_eval_quality.py — it blocks 22 defect classes that can cost a real result here. Then confirm by hand:

  • every fixture behaves as its stimulus assumes — a fixture meant to be healthy builds, and one meant to be broken fails for the exact reason the stimulus is about and no other;
  • every referenced fixture is in the git index (git ls-files), not merely on disk — .gitignore has silently swallowed committed coverage fixtures;
  • a fixture never states the same fact in two places that disagree — a Cobertura report whose declared line-rate, summary totals and <line> elements differ is the canonical case — or the two arms legitimately read different truths.
  • every preservation or scope assertion covers the complete in-scope file set, not one representative file.
  • the golden trajectory and patch pass the deterministic graders, and a realistic broken mutation fails the grader that is meant to protect the behavior.
Step 5: Check whether the eval could ever have passed

The gate has two independent bars, and confusing them is the usual misdiagnosis:

  1. Distinct stimuli ≥ 5. Below that the verdict is reported underpowered — never a pass, never a regression.
  2. The sign test must reach p ≤ 0.05 over the discordant (non-tie) stimulus votes. Ties are not discarded silently; they hold the discordant count down.
discordant stimulus votesrecords that passp
≤ 4none, however good the skill≥ 0.0625
5–7zero losses only (5W/0L)0.031
8one loss survivable (7W/1L)0.035

So at exactly 5 stimuli a single tie is fatal — it leaves 4 discordant. At 6 stimuli one tie is survivable (5W/1T/0L); at 7, up to two are (5W/2T/0L). A loss is not.

So a positive record with a failing verdict is a power problem, not a content problem. Fix it by adding discriminating stimuli. Raising runs measures reliability for the same task and cannot clear the floor.

Show full SKILL.md (808 more words)Show less
Step 6: Check whether the two arms differ at all

An eval that compares the skill against itself measures judge noise:

  • A dormancy guard (expect_activation: false) must not also set constraints.reject_skills. That makes the skilled arm skill-free, so the activation contract cannot observe a hijack. Schema version 4 retains the identical-arm comparison for diagnostics but excludes it from preference inference; unexpected isolated activation still blocks a pass.
  • A skill with disable-model-invocation: true is absent from the model-facing skilled arm, so its direct eval compares two identical arms regardless of whether graders inspect activation or answer content. Cover it through consumer outcomes instead; for example, filter-syntax is covered by run-tests and mtp-hot-reload.
  • A grader whose config is missing its required key enforces nothing, so the stimulus has one fewer assertion than it appears to.
Step 7: Fix skill content against the losing trial

Only now change the skill. Apply the patterns in references/writing-for-baseline-delta.md; the ones that most often flip a loss:

  • Replace reference prose the model already knows with decisions it would otherwise get wrong.
  • Add stop-conditions so a strong skill does not over-apply — but do not over-correct into answering more narrowly than the baseline did.
  • Scale output structure to input size; a dashboard for an 8-test suite loses to a direct answer.
  • Require truthful validation reporting; claiming "Build succeeded" after a failed restore is an automatic loss.
  • Verify load-bearing API claims by compiling or probing, not by reading source.
  • For cost regressions, gate rare or expensive paths behind references/ reads and size any orchestration to the user's scope.
Step 8: Fix activation

Activation failures are frontmatter and routing failures, not body failures. See references/eval-triage.md. Summary:

FailureFix
Not activated in any armPut the user's own words in description: symptoms, error codes, artifact names, quoted requests
A sibling skill wins the promptClaim the exact ambiguous words in description, and add matching exclusions on both siblings
Model answers with no skill at allRaise the stakes in the description, de-crowd the plugin menu, verify with the plugin arm
Boundary excludes real scenariosRe-read every "do not use for" clause against every eval prompt and real workflow phase
Description at the 1,024-char ceilingCut restated body content, not trigger phrases; check the plugin menu budget too
Step 9: Re-validate
bash
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/<plugin>
python eng/eval-quality/check_eval_quality.py
./eng/run-skill-evals.sh <plugin> <skill>

Use the production path at normal worker concurrency and with the declared defaults.timeout: Vally for skill evals, and skill-validator evaluate for agent evals. For an agent eval, separately run check_eval_quality.py, apply each golden patch to its materialized fixture, and run the applicable deterministic file, output, and command graders against the golden result. Do not use a serial-only pass or a larger ad hoc budget as completion evidence. For broad routing or behavior changes, collect separate GPT-family and Claude-family results. Do not pool model families into extra stimulus votes.

Then request the official run by submitting a PR review containing /evaluate (Files changed → Review changes), which binds the run to the reviewed commit. Before declaring a regression on the result, confirm the skill payload actually changed — reruns on byte-identical content have shifted 7W/2T/2L to 4W/5T/2L.

Validation

  • For a content fix, a losing trial and the judge's stated reason are quoted in the PR description.
  • The failure was classified before any content was edited.
  • check_eval_quality.py and skill-validator check both pass.
  • The production skill or agent runner accepts the executable spec.
  • Golden references pass the standalone checker and deterministic replay.
  • The eval completes under normal concurrency and its declared time budget.
  • Golden acceptance and mutation rejection have been demonstrated.
  • Distinct-stimulus count clears the power bar for the target effect and observed tie rate.
  • Isolated and plugin activation are both reported.
  • Broad changes have separate GPT-family and Claude-family evidence.
  • The PR body records root cause, fix, and validation so the lesson is reusable.

Common Pitfalls

PitfallSolution
Rewriting skill prose in response to an underpowered verdictUnderpowered means too few distinct stimuli; add discriminating stimuli instead
Using the deprecated top-level config: aliasRename it to defaults: and preserve its settings; the repository gate rejects the alias
Padding runs to clear the stimulus floorRepeats measure reliability for one task; add stimuli
Treating an errored trial as fixture nondeterminismRead the stderr first; judge-side auth failures need harness fixes
Fixing a "wrong" answer that the fixture actually made wrongCheck fixture self-consistency before blaming the response
Strengthening a skill nobody uses and nothing passesWeak eval signal plus thin telemetry is a valid retirement case
Landing a fix without re-runningVerify the invoked payload contains the fix; judge noise is real

References

© dotnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/improve-skill-quality of dotnet/skills.

  • SKILL.md
  • references/eval-triage.md
  • references/writing-for-baseline-delta.md

Open the folder on GitHubat commit 8d670fa

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in dotnet/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Improve Skill Quality next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Improve Skill Quality compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Improve Skill Quality this skilldotnet/skills5.6k1 repos~3.6kAutomated safety check: PassMIT
Scaffoldcodewithmukesh/dotnet-claude-kit751—~1.7kAutomated safety check: PassMIT
Revit Toolkit AnalyzersNice3point/RevitToolkit176—~3.8kAutomated safety check: PassMIT
Revit Toolkit InternalsNice3point/RevitToolkit176—~2.5kAutomated safety check: PassMIT
Corvus Benchmarkscorvus-dotnet/Corvus.JsonSchema199—~1.2kAutomated safety check: PassApache-2.0
Create Migrationfullstackhero/dotnet-starter-kit6.8k—~757Automated safety check: PassMIT

Similar skills

  • Scaffold

    codewithmukesh/dotnet-claude-kit

    Architecture-aware feature scaffolding for .NET 10 projects.

    751 GitHub stars~1.7k tokensUpdated 2 mo ago
    DevelopmentAuto-check passed
  • Revit Toolkit Analyzers

    Nice3point/RevitToolkit

    Author and extend the Roslyn tooling bundled in the Nice3point.Revit.Toolkit package: the incremental source generator that emits external-event boilerplate, the analyzers and code fixers that…

    176 GitHub stars~3.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Revit Toolkit Internals

    Nice3point/RevitToolkit

    Uphold the design contract of the Nice3point.Revit.Toolkit runtime library, that wraps the raw Revit add-in API.

    176 GitHub stars~2.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Corvus Benchmarks

    corvus-dotnet/Corvus.JsonSchema

    Run, interpret, and maintain BenchmarkDotNet benchmarks for JSON Schema validation and query languages.

    199 GitHub stars~1.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Create Migration

    fullstackhero/dotnet-starter-kit

    Create and apply an EF Core migration for a module's DbContext the FSH way (central Migrations project, per-module folder, correct --context).

    6.8k GitHub stars~757 tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Corvus Typescript Evaluator

    corvus-dotnet/Corvus.JsonSchema

    Work on the TypeScript port of the V5 standalone schema evaluator (src-ts/corvus-json-schema, npm package @corvus-dotnet/json-schema): loader, compiler, JavaScript code generator, results collector…

    199 GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed

More from dotnet/skills

All 91 skills in this repo
  • Official

    Resolves .NET runtime frames in Apple .ips crash logs to function names, source files and line numbers using dSYM symbols, atos and the Microsoft symbol server.

    5.6k GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Official

    Resolves native crash frames from .NET Android tombstones to function names, source files and line numbers using BuildIds, Microsoft's symbol server and llvm-symbolizer.

    5.6k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Official

    Scans C# and .NET code for about 50 performance anti-patterns and reports prioritized findings with concrete fixes, at a scan depth you choose.

    5.6k GitHub starsUsed in 3 repos~3.1k tokens
    Auto-check passed
  • Official

    Statically pairs source files with test files to list code that no test references, using Roslyn for C# or tree-sitter for many languages, with no build.

    5.6k GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed
  • Microbenchmarking

    dotnet/skills

    Official

    Activate this skill when BenchmarkDotNet (BDN) is involved in the task — creating, running, configuring, or reviewing BDN benchmarks.

    5.6k GitHub starsUsed in 3 repos~3.3k tokens
    Auto-check passed
  • Official

    Makes .NET projects compatible with Native AOT and trimming by resolving IL trim and AOT analyzer warnings through annotations rather than suppressions.

    5.6k GitHub starsUsed in 2 repos~4.2k tokens
    Auto-check passed

Works with

Categories

Questions about Improve Skill Quality

What does Improve Skill Quality do?

Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement". Improve Skill Quality is an agent skill from dotnet/skills, published by the product's own GitHub organization. Diagnoses and fixes skills in the dotnet/skills repository that lose to their own baseline, fail to activate, time out, or return "no credible improvement".

When should I use Improve Skill Quality?

Improve Skill Quality fits situations like: an evaluation verdict is a regression; A skill regressed after a change; /evaluate reports no results; deciding whether a weak skill should be strengthened.

How do I install Improve Skill Quality in Claude Code?

Run `npx skills add dotnet/skills --skill improve-skill-quality -a claude-code`. Or copy the skill folder (.agents/skills/improve-skill-quality in dotnet/skills) into .claude/skills/improve-skill-quality in your project. Claude Code loads it when a task matches its description.

How do I install Improve Skill Quality in Codex?

Run `npx skills add dotnet/skills --skill improve-skill-quality -a codex`. Or copy the skill folder (.agents/skills/improve-skill-quality in dotnet/skills) into .agents/skills/improve-skill-quality in your project. Codex loads it when a task matches its description.

Can I use Improve Skill Quality in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dotnet/skills --skill improve-skill-quality -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/improve-skill-quality, .gemini/skills/improve-skill-quality, .github/skills/improve-skill-quality and .opencode/skills/improve-skill-quality in your project.

What does Improve Skill Quality need to run?

Going by SKILL.md and its folder, Improve Skill Quality needs the command-line tools its instructions call (python, dotnet and git). Our summary lists: Python 3.

Does Improve Skill Quality access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Improve Skill Quality safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Improve Skill Quality use?

Improve Skill Quality is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Improve Skill Quality use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Improve Skill Quality?

Skills that share tags, products or a category with Improve Skill Quality: Scaffold (codewithmukesh/dotnet-claude-kit, 751 stars), Revit Toolkit Analyzers (Nice3point/RevitToolkit, 176 stars), Revit Toolkit Internals (Nice3point/RevitToolkit, 176 stars) and Corvus Benchmarks (corvus-dotnet/Corvus.JsonSchema, 199 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Improve Skill Quality?

dotnet (a GitHub organization, an official publisher) maintains it in dotnet/skills, which has 5,568 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on October 7, 2026.

Source: dotnet/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.