Arize Evaluator
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Sets up and runs skillgrade evaluation pipelines for Agent Skills.
$ npx skills add mgechev/skillgrade --skill skillgrade-setup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mgechev/skillgrade skillgrade-setup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mgechev/skillgrade.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skillgrade-setup .claude/skills/skillgrade-setup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skillgrade-setup" agent skill from https://github.com/mgechev/skillgrade/tree/main/skills/skillgrade-setup into .claude/skills/skillgrade-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skillgrade-setup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mgechev/skillgrade/tree/main/skills/skillgrade-setupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mgechev/skillgrade --skill skillgrade-setup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mgechev/skillgrade skillgrade-setup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mgechev/skillgrade.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/skillgrade-setup .agents/skills/skillgrade-setup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skillgrade-setup" agent skill from https://github.com/mgechev/skillgrade/tree/main/skills/skillgrade-setup into .agents/skills/skillgrade-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skillgrade-setup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mgechev/skillgrade --skill skillgrade-setup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mgechev/skillgrade skillgrade-setup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mgechev/skillgrade.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/skillgrade-setup .cursor/skills/skillgrade-setup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skillgrade-setup" agent skill from https://github.com/mgechev/skillgrade/tree/main/skills/skillgrade-setup into .cursor/skills/skillgrade-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skillgrade-setup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mgechev/skillgrade.git --path skills/skillgrade-setup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mgechev/skillgrade --skill skillgrade-setup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mgechev/skillgrade skillgrade-setup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mgechev/skillgrade.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/skillgrade-setup .gemini/skills/skillgrade-setup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skillgrade-setup" agent skill from https://github.com/mgechev/skillgrade/tree/main/skills/skillgrade-setup into .gemini/skills/skillgrade-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skillgrade-setup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mgechev/skillgrade skillgrade-setupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mgechev/skillgrade --skill skillgrade-setup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mgechev/skillgrade.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/skillgrade-setup .github/skills/skillgrade-setup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skillgrade-setup" agent skill from https://github.com/mgechev/skillgrade/tree/main/skills/skillgrade-setup into .github/skills/skillgrade-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skillgrade-setup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mgechev/skillgrade --skill skillgrade-setup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mgechev/skillgrade skillgrade-setup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mgechev/skillgrade.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/skillgrade-setup .opencode/skills/skillgrade-setup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skillgrade-setup" agent skill from https://github.com/mgechev/skillgrade/tree/main/skills/skillgrade-setup into .opencode/skills/skillgrade-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skillgrade-setup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skillgrade-setupSets up and runs skillgrade evaluation pipelines for Agent Skills.
Skillgrade Setup is an agent skill from mgechev/skillgrade. Sets up and runs skillgrade evaluation pipelines for Agent Skills. Use when initializing eval configurations, running trials, reviewing results, or integrating with CI. Don't use for writing grader scripts, general test authoring, or non-agentic documentation.
Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/ci-example.md` and `references/eval-yaml-spec.md`).
The repository describes itself as: "Unit tests" for your agent skills. The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit bb213d4. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GEMINI_API_KEYANTHROPIC_API_KEYOPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skillgrade Setup loads about 915 tokens when it runs, and up to ~2.1k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 441 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mgechev/skillgrade at commit bb213d4, republished under its MIT licence (© mgechev). 441 words, ~915 tokens.
.claude/skills/skillgrade-setup/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Step 1: Install Skillgrade
npm i -g skillgrade to install the CLI globally.Step 2: Initialize an Eval Configuration
SKILL.md).GEMINI_API_KEY, ANTHROPIC_API_KEY, or OPENAI_API_KEY).skillgrade init to generate an eval.yaml with AI-powered tasks and graders.eval.yaml already exists, pass --force to overwrite: skillgrade init --force.Step 3: Configure eval.yaml
references/eval-yaml-spec.md for the full configuration schema.tasks: key. Each task requires:name: unique task identifierinstruction: what the agent should accomplishworkspace: files to copy into the evaluation containergraders: one or more scoring mechanisms (see the skillgrade-graders skill)defaults: for agent, provider, trials, timeout, and threshold.Step 4: Run Evaluations
--smoke (5 trials): Quick capability check.--reliable (15 trials): Reliable pass rate estimate.--regression (30 trials): High-confidence regression detection.skillgrade --smoke.skillgrade --eval=fix-linting.skillgrade --eval=fix-linting,write-tests.skillgrade --grader=deterministic.skillgrade --grader=llm_rubric.--agent=gemini|claude|codex|acp|opencode|command.--acp-command="gemini --acp" or set defaults.acp.command.--opencode-agent=build|plan|explore or --opencode-model=provider/model.--agent=command --command="node mycli.js" or set defaults.command. The instruction is piped to the command's stdin.--provider=docker|local.Step 5: Review Results
skillgrade preview for a CLI report.skillgrade preview browser to open the web UI at http://localhost:3847.$TMPDIR/skillgrade/<skill-name>/results/. Override with --output=DIR.Step 6: Integrate with CI
--regression --ci --provider=local.--provider=local in CI — the runner is already an ephemeral sandbox, so Docker adds overhead without benefit.--ci flag causes a non-zero exit code if the pass rate falls below --threshold (default: 0.8).references/ci-example.md for a complete workflow template.skillgrade init fails with "No SKILL.md found," verify the current directory contains a valid SKILL.md file.© mgechev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/skillgrade-setup of mgechev/skillgrade.
Open the folder on GitHubat commit bb213d4
Skillgrade Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skillgrade Setup this skillmgechev/skillgrade | 720 | — | ~915 | Automated safety check: Pass | MIT | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 2 repos | ~8.1k | Automated safety check: Notes | MIT | |
| LLM Evaluationdavila7/claude-code-templates | 32k | 13 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Agent Evaluationsickn33/agentic-awesome-skills | 47k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Opensource Pipelineaffaan-m/ECC | 275k | 1 repos | ~1.8k | Automated safety check: Notes | MIT | |
| Orch Pipelineaffaan-m/ECC | 275k | 1 repos | ~1.6k | Automated safety check: Pass | MIT |
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
sickn33/agentic-awesome-skills
Evaluate agent behavior with versioned cases and explicit verifiers.
affaan-m/ECC
Open-source pipeline: fork, sanitize, and package private projects for safe public release.
affaan-m/ECC
Shared orchestration engine behind the orch- skill family — the gated Research-Plan-TDD-Review-Commit pipeline, size classifier, agent and command map, and two human gates (plan approval, commit…
asgeirtj/system_prompts_leaks
Any explicit Muse Code setting question or change (model, reasoning effort, /settings) requires a silent readskill call for bundled:manage-settings as FIRST ACTION—no assistant text or other tool…
mgechev/skillgrade
Authors deterministic and LLM rubric graders for skillgrade evaluations.
Sets up and runs skillgrade evaluation pipelines for Agent Skills. Skillgrade Setup is an agent skill from mgechev/skillgrade. Sets up and runs skillgrade evaluation pipelines for Agent Skills.
Skillgrade Setup fits situations like: initializing eval configurations; reviewing results; integrating with CI; writing grader scripts.
Run `npx skills add mgechev/skillgrade --skill skillgrade-setup -a claude-code`. Or copy the skill folder (skills/skillgrade-setup in mgechev/skillgrade) into .claude/skills/skillgrade-setup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mgechev/skillgrade --skill skillgrade-setup -a codex`. Or copy the skill folder (skills/skillgrade-setup in mgechev/skillgrade) into .agents/skills/skillgrade-setup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mgechev/skillgrade --skill skillgrade-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skillgrade-setup, .gemini/skills/skillgrade-setup, .github/skills/skillgrade-setup and .opencode/skills/skillgrade-setup in your project.
Going by SKILL.md and its folder, Skillgrade Setup needs the command-line tools its instructions call (npm) and credentials named GEMINI_API_KEY, ANTHROPIC_API_KEY and OPENAI_API_KEY. Our summary lists: Node.js; Docker; A credential in GEMINI_API_KEY; A credential in ANTHROPIC_API_KEY.
SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Skillgrade Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 915 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Skillgrade Setup: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Opensource Pipeline (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mgechev (a GitHub user) maintains it in mgechev/skillgrade, which has 720 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 5, 2026.
Source: mgechev/skillgrade on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.