Agent skill

Skillgrade Setup

by mgechev in mgechev/skillgrade

Sets up and runs skillgrade evaluation pipelines for Agent Skills.

MITAuto-check passed

Install Skillgrade Setup

skills CLI
$ npx skills add mgechev/skillgrade --skill skillgrade-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mgechev/skillgrade skillgrade-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mgechev/skillgrade.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skillgrade-setup .claude/skills/skillgrade-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skillgrade-setup
GitHub stars
720
Token cost
~915 tokens
SKILL.md length
441 words
Files
3 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Sets up and runs skillgrade evaluation pipelines for Agent Skills.

  • Works in 2 steps: Verify Node.js 20+ and Docker are… → Run npm i -g skillgrade to install the…
  • Initializing eval configurations
  • SKILL.md covers Procedures and Error Handling
  • Calls npm; needs GEMINI_API_KEY and ANTHROPIC_API_KEY

What it does

Skillgrade Setup is an agent skill from mgechev/skillgrade. Sets up and runs skillgrade evaluation pipelines for Agent Skills. Use when initializing eval configurations, running trials, reviewing results, or integrating with CI. Don't use for writing grader scripts, general test authoring, or non-agentic documentation.

Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/ci-example.md` and `references/eval-yaml-spec.md`).

The repository describes itself as: "Unit tests" for your agent skills. The licence is MIT.

When your agent uses it

  • Initializing eval configurations
  • Reviewing results
  • Integrating with CI
  • Writing grader scripts

Example prompts

  • “/skillgrade-setup”

Requirements

  • Node.js
  • Docker
  • A credential in GEMINI_API_KEY
  • A credential in ANTHROPIC_API_KEY

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Verify Node.js 20+ and Docker are available.
  2. Run npm i -g skillgrade to install the CLI globally.

What it can do on your machine

Read from SKILL.md and the folder at commit bb213d4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY
    • ANTHROPIC_API_KEY
    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skillgrade Setup loads about 915 tokens when it runs, and up to ~2.1k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 441 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~915
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mgechev/skillgrade at commit bb213d4, republished under its MIT licence (© mgechev). 441 words, ~915 tokens.

Download SKILL.mdSave it as .claude/skills/skillgrade-setup/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
skillgrade-setup
description
Sets up and runs skillgrade evaluation pipelines for Agent Skills. Use when initializing eval configurations, running trials, reviewing results, or integrating with CI. Don't use for writing grader scripts, general test authoring, or non-agentic documentation.

Skillgrade Evaluation Setup

Procedures

Step 1: Install Skillgrade

  1. Verify Node.js 20+ and Docker are available.
  2. Run npm i -g skillgrade to install the CLI globally.

Step 2: Initialize an Eval Configuration

  1. Navigate to the skill directory (must contain a SKILL.md).
  2. Set the appropriate API key environment variable (GEMINI_API_KEY, ANTHROPIC_API_KEY, or OPENAI_API_KEY).
  3. Run skillgrade init to generate an eval.yaml with AI-powered tasks and graders.
  4. If an eval.yaml already exists, pass --force to overwrite: skillgrade init --force.
  5. Without an API key, a well-commented template is generated instead.

Step 3: Configure eval.yaml

  1. Read references/eval-yaml-spec.md for the full configuration schema.
  2. Define one or more tasks under the tasks: key. Each task requires:
    • name: unique task identifier
    • instruction: what the agent should accomplish
    • workspace: files to copy into the evaluation container
    • graders: one or more scoring mechanisms (see the skillgrade-graders skill)
  3. Optionally configure defaults: for agent, provider, trials, timeout, and threshold.

Step 4: Run Evaluations

  1. Select an appropriate preset based on the evaluation goal:
    • --smoke (5 trials): Quick capability check.
    • --reliable (15 trials): Reliable pass rate estimate.
    • --regression (30 trials): High-confidence regression detection.
  2. Run the evaluation: skillgrade --smoke.
  3. Run a specific eval by name: skillgrade --eval=fix-linting.
  4. Run multiple evals: skillgrade --eval=fix-linting,write-tests.
  5. Run only deterministic graders (skip LLM calls): skillgrade --grader=deterministic.
  6. Run only LLM rubric graders: skillgrade --grader=llm_rubric.
  7. The agent is auto-detected from the API key. Override with --agent=gemini|claude|codex|acp|opencode|command.
  8. For ACP, pass --acp-command="gemini --acp" or set defaults.acp.command.
  9. For OpenCode, pass --opencode-agent=build|plan|explore or --opencode-model=provider/model.
  10. For a custom agent, pass --agent=command --command="node mycli.js" or set defaults.command. The instruction is piped to the command's stdin.
  11. Override the provider with --provider=docker|local.
Show full SKILL.md (159 more words)Show less

Step 5: Review Results

  1. Run skillgrade preview for a CLI report.
  2. Run skillgrade preview browser to open the web UI at http://localhost:3847.
  3. Reports are saved to $TMPDIR/skillgrade/<skill-name>/results/. Override with --output=DIR.

Step 6: Integrate with CI

  1. Add a GitHub Actions step that installs skillgrade, navigates to the skill directory, and runs with --regression --ci --provider=local.
  2. Use --provider=local in CI — the runner is already an ephemeral sandbox, so Docker adds overhead without benefit.
  3. The --ci flag causes a non-zero exit code if the pass rate falls below --threshold (default: 0.8).
  4. Read references/ci-example.md for a complete workflow template.

Error Handling

  • If skillgrade init fails with "No SKILL.md found," verify the current directory contains a valid SKILL.md file.
  • If evaluation hangs, check Docker is running and the container has network access for API calls.
  • If all trials fail with "No API key," ensure the environment variable is exported, not just set inline for a different command.

© mgechev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/skillgrade-setup of mgechev/skillgrade.

  • SKILL.md
  • references/ci-example.md
  • references/eval-yaml-spec.md

Open the folder on GitHubat commit bb213d4

Compare with similar skills

Skillgrade Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skillgrade Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skillgrade Setup this skillmgechev/skillgrade720—~915Automated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k13 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
Opensource Pipelineaffaan-m/ECC275k1 repos~1.8kAutomated safety check: NotesMIT
Orch Pipelineaffaan-m/ECC275k1 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 13 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Open-source pipeline: fork, sanitize, and package private projects for safe public release.

    275k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check: notes
  • Orch Pipeline

    affaan-m/ECC

    Shared orchestration engine behind the orch- skill family — the gated Research-Plan-TDD-Review-Commit pipeline, size classifier, agent and command map, and two human gates (plan approval, commit…

    275k GitHub starsUsed in 1 repo~1.6k tokens
    Testing & QAAuto-check passed
  • Manage Settings

    asgeirtj/system_prompts_leaks

    Any explicit Muse Code setting question or change (model, reasoning effort, /settings) requires a silent readskill call for bundled:manage-settings as FIRST ACTION—no assistant text or other tool…

    69k GitHub stars~3.1k tokensUpdated today
    Auto-check passed

More from mgechev/skillgrade

  • Skillgrade Graders

    mgechev/skillgrade

    Authors deterministic and LLM rubric graders for skillgrade evaluations.

    720 GitHub stars~972 tokensUpdated 2 days ago
    Auto-check passed

Questions about Skillgrade Setup

What does Skillgrade Setup do?

Sets up and runs skillgrade evaluation pipelines for Agent Skills. Skillgrade Setup is an agent skill from mgechev/skillgrade. Sets up and runs skillgrade evaluation pipelines for Agent Skills.

When should I use Skillgrade Setup?

Skillgrade Setup fits situations like: initializing eval configurations; reviewing results; integrating with CI; writing grader scripts.

How do I install Skillgrade Setup in Claude Code?

Run `npx skills add mgechev/skillgrade --skill skillgrade-setup -a claude-code`. Or copy the skill folder (skills/skillgrade-setup in mgechev/skillgrade) into .claude/skills/skillgrade-setup in your project. Claude Code loads it when a task matches its description.

How do I install Skillgrade Setup in Codex?

Run `npx skills add mgechev/skillgrade --skill skillgrade-setup -a codex`. Or copy the skill folder (skills/skillgrade-setup in mgechev/skillgrade) into .agents/skills/skillgrade-setup in your project. Codex loads it when a task matches its description.

Can I use Skillgrade Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mgechev/skillgrade --skill skillgrade-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skillgrade-setup, .gemini/skills/skillgrade-setup, .github/skills/skillgrade-setup and .opencode/skills/skillgrade-setup in your project.

What does Skillgrade Setup need to run?

Going by SKILL.md and its folder, Skillgrade Setup needs the command-line tools its instructions call (npm) and credentials named GEMINI_API_KEY, ANTHROPIC_API_KEY and OPENAI_API_KEY. Our summary lists: Node.js; Docker; A credential in GEMINI_API_KEY; A credential in ANTHROPIC_API_KEY.

Does Skillgrade Setup access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Skillgrade Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skillgrade Setup use?

Skillgrade Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skillgrade Setup use?

About 915 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Skillgrade Setup?

Skills that share tags, products or a category with Skillgrade Setup: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Opensource Pipeline (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skillgrade Setup?

mgechev (a GitHub user) maintains it in mgechev/skillgrade, which has 720 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 5, 2026.

Source: mgechev/skillgrade on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.