Official agent skill

Create Custom Grader

by NVIDIA in NVIDIA/SkillEvaluator

A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

OfficialApache-2.0Auto-check passedEducation

Install Create Custom Grader

skills CLI
$ npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/SkillEvaluator create-custom-grader --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/SkillEvaluator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skillevaluator/tier3/reference_skills/create-custom-grader .claude/skills/create-custom-grader && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
create-custom-grader
GitHub stars
548
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
950 words
Files
2
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

  • Works in 4 steps: Read the target skill, existing evals/,… → Choose default_plus_custom when custom… → Choose custom_only only when the user… → …
  • Converting an existing benchmark
  • SKILL.md covers Purpose, When To Use, Instructions and Examples, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Create Custom Grader is an agent skill from NVIDIA/SkillEvaluator, published by the product's own GitHub organization. Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).

It sits in Education, covering Quizzes and assessments, Agent evaluation and testing and LLM evaluation. The repository describes itself as: Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that… The licence is Apache-2.0.

When your agent uses it

  • Converting an existing benchmark
  • Domain check into SkillEvaluator BYOG/BYOT custom evaluation

Example prompts

  • “/create-custom-grader”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read the target skill, existing evals/, benchmark prompts, fixtures, and any verifier code.
  2. Choose default_plus_custom when custom metrics should complement default evaluator scoring.
  3. Choose custom_only only when the user wants the custom grader to own pass/fail semantics.
  4. Write or update evals/grader.py or evals/grader.sh, then validate the Harbor contract.

What it can do on your machine

Read from SKILL.md and the folder at commit f32c884. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Create Custom Grader loads about 2.1k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 950 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/SkillEvaluator at commit f32c884, republished under its Apache-2.0 licence (© NVIDIA). 950 words, ~2,147 tokens.

Download SKILL.mdSave it as .claude/skills/create-custom-grader/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
create-custom-grader
description
Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
metadata.author
SkillEvaluator Maintainers <maintainers@example.com>

Create Custom Grader

Convert team-owned benchmark definitions into runnable SkillEvaluator custom graders and, when needed, native Harbor tasks.

Purpose

Help an agent author valid SkillEvaluator BYOG/BYOT files from a user's benchmark instead of leaving the user with empty grader templates.

When To Use

Use this skill when the user wants to:

  • bring an existing benchmark into SkillEvaluator
  • turn a rubric into evals/grader.py or evals/grader.sh
  • add custom metrics beside the default evaluator metrics
  • convert task files such as task.yaml, task.json, pytest checks, or shell verifiers into BYOG or BYOT
  • prove a team can run its own benchmark through SkillEvaluator

Do not use this skill for ordinary evals/evals.json authoring when no custom grading logic is needed. Use the normal dataset authoring workflow for that.

Instructions

  1. Read the target skill, existing evals/, benchmark prompts, fixtures, and any verifier code.
  2. Choose default_plus_custom when custom metrics should complement default evaluator scoring.
  3. Choose custom_only only when the user wants the custom grader to own pass/fail semantics.
  4. Write or update evals/grader.py or evals/grader.sh, then validate the Harbor contract.

Examples

bash
skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom
skillevaluator tier3 validate <skill-dir>

Prerequisites

  • The target skill directory should contain SKILL.md.
  • The SkillEvaluator CLI should be available as skillevaluator.
  • Full E2E evaluation may need agent credentials, sandbox access, GPU access, or service credentials depending on the benchmark.

Core Choice

Choose one path before writing files:

User needEvaluator shape
Existing evals.json task plus extra domain checksTop-level BYOG: evals/grader.py or evals/grader.sh
Existing benchmark prompt/rubric that can run in the generated workspaceTop-level BYOG plus evals/evals.json and evals/files/
Benchmark owns task layout, setup, service lifecycle, or verifier harnessNative BYOT/BYOG: evals/harbor/<case>/...
User wants only custom reward/pass criteriagrading.mode: custom_only
User wants default evaluator dimensions plus custom metricsgrading.mode: default_plus_custom

Default to default_plus_custom unless the user explicitly wants the custom grader to replace the default evaluator metrics.

Workflow

  1. Resolve the target skill and benchmark source. Read the target SKILL.md, existing evals/, benchmark prompts, fixtures, rubric, reference solution, tags, and any expected trigger/non-trigger metadata.

  2. Map benchmark fields into evaluator inputs. Use benchmark prompts or prompt variants as question entries. Use the target skill as expected_skill. Put each case's required starter files under evals/files/<case-id>/, and declare files: ["evals/files/<case-id>"] on every corresponding eval entry. Do not omit files in a multi-case dataset, because omission intentionally stages the entire shared directory for legacy compatibility. Preserve benchmark-specific rubric text in the entry only when the grader needs to read it.

  3. Scaffold the evaluator contract. For generated tasks:

    bash
    skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom

    For shell checks:

    bash
    skillevaluator init-custom-grader <skill-dir> --language shell --mode default_plus_custom

    For native Harbor tasks:

    bash
    skillevaluator init-harbor-task <skill-dir> --case-id <case-id> --with-config
  4. Replace scaffold placeholders. The custom grader is real executable logic, not metadata. It must read available evidence, compute numeric scores, and write the evaluator reward contract.

  5. Validate before running.

    bash
    skillevaluator validate <skill-dir> --harbor-contract

    Fix missing files, invalid Python, missing reward output, and native Harbor ID mismatches before evaluation.

  6. Run the deepest practical proof. Prefer a real with-skill/baseline run. If services, credentials, GPU, or cost block full E2E, state exactly what was validated and what was not.

Grader Contract

Python and shell graders run inside the Harbor verifier context. They may read:

  • /logs/agent/trajectory.json for agent actions and final answer evidence
  • /tests/entry.json for the eval case metadata
  • /workspace/input/ for the entry's declared committed fixtures from evals/files/
  • /solution/ or other task outputs only when the task environment produces them

They must write:

  • /logs/verifier/reward.json
  • /logs/verifier/reward.txt with a numeric score from 0.0 to 1.0

Use this reward shape:

json
{
  "overall": 0.92,
  "custom_metrics": {
    "domain_repair": 1.0,
    "domain_verification": 0.8
  },
  "details": {
    "domain_repair": {
      "score": 1.0,
      "reason": "The solution repaired the required files."
    }
  }
}

In default_plus_custom, default evaluator scoring keeps its overall authoritative and adds the grader's custom_metrics into reports. In custom_only, the grader's overall is the pass/fail reward.

Never emit custom metric names that collide with reserved evaluator fields: security, skill_execution, skill_efficiency, accuracy, goal_accuracy, behavior_check, overall, details, metrics, metric_set, or entry_id.

Show full SKILL.md (347 more words)Show less

Translation Rules

  • Convert each rubric item into a deterministic check when possible.
  • If a rubric item requires judgment, encode observable proxies and explain the limits in details.
  • Keep metrics stable across baseline and with-skill runs.
  • Score only the generated task workspace. Do not accidentally score copied skill source files, reference fixtures, or grader templates.
  • Keep custom metric values clamped to 0.0 through 1.0.
  • Preserve benchmark prompt variants as separate eval entries only when they exercise meaningfully different behavior.
  • Convert expected trigger/non-trigger metadata into expected_skill, expected_behavior, negative cases, or custom metrics that inspect trajectory evidence.

RAPIDS-Style Example

For a benchmark task with task.yaml, code/, prompt variants, coverage, and a rubric:

  1. Copy code/ into evals/files/<case-id>/.
  2. Create one or more evals/evals.json entries from the prompt variants, and set files: ["evals/files/<case-id>"] on each corresponding entry.
  3. Set expected_skill to the benchmark's target skill.
  4. Implement evals/grader.py to inspect the agent trajectory and changed workspace files.
  5. Emit custom metrics for each rubric criterion, for example rapids_diagnosis, rapids_requirements_repair, rapids_repair_safety, and rapids_verification.
  6. Validate and run SkillEvaluator with and without the target skill, then report both default evaluator metrics and custom metric deltas.

Limitations

  • The skill can design and implement deterministic checks, but ambiguous rubric judgment still needs explicit observable proxies or a human-approved scoring policy.
  • init-custom-grader creates scaffolding only; the agent must replace the placeholder scoring logic.
  • Local validation proves file contracts, not live agent behavior. Do not call the benchmark proven until an evaluation run has produced real rewards.

Troubleshooting

ProblemFix
evals/evals.json missingCreate entries from the benchmark prompt or run init-custom-grader to seed one.
Custom metrics do not appearEnsure reward.json has numeric values under custom_metrics and no reserved-name collisions.
custom_only failsWrite numeric overall in reward.json or numeric reward.txt.
Grader scores copied fixturesRestrict file searches to generated workspace/output paths, not the skill package or grader source.

Final Response

When finished, report:

  • files created or changed
  • exact validation and evaluation commands
  • default evaluator metric results
  • custom metric results
  • whether the proof was full E2E or only static/local validation
  • any benchmark rubric criteria that remain partly judgment-based

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in src/skillevaluator/tier3/reference_skills/create-custom-grader of NVIDIA/SkillEvaluator.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit f32c884

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/SkillEvaluator, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Create Custom Grader next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Create Custom Grader compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Create Custom Grader this skillNVIDIA/SkillEvaluator5481 repos~2.1kAutomated safety check: PassApache-2.0
Task Creatorbenchflow-ai/benchflow353—~4.5kAutomated safety check: PassApache-2.0
Skill JudgeshareAI-lab/lab-skills314—~1.9kAutomated safety check: PassApache-2.0
Woo AI Smokewoocommerce/woocommerce-ios3581 repos~7.4kAutomated safety check: NotesGPL-2.0
Create Skill Testdotnet/skills5.6k1 repos~5.7kAutomated safety check: PassMIT
Agent Evaluationseb1n/awesome-ai-agent-skills206—~1.4kAutomated safety check: PassMIT

Similar skills

  • Task Creator

    benchflow-ai/benchflow

    SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

    353 GitHub stars~4.5k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Skill Judge

    shareAI-lab/lab-skills

    Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.

    314 GitHub stars~1.9k tokensUpdated 22 days ago
    EducationAuto-check passed
  • Woo AI Smoke

    woocommerce/woocommerce-ios

    Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

    358 GitHub starsUsed in 1 repo~7.4k tokens
    EducationAuto-check: notes
  • Create Skill Test

    dotnet/skills

    Official

    Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository.

    5.6k GitHub starsUsed in 1 repo~5.7k tokens
    EducationAuto-check passed
  • Agent Evaluation

    seb1n/awesome-ai-agent-skills

    Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis.

    206 GitHub stars~1.4k tokensUpdated 1 mo ago
    EducationAuto-check passed
  • Advanced Evaluation

    aiskillstore/marketplace

    This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…

    430 GitHub starsUsed in 4 repos~4.2k tokens
    EducationAuto-check passed

More from NVIDIA/SkillEvaluator

  • API Caller

    NVIDIA/SkillEvaluator

    Official

    Call any REST API dynamically. An agent skill from NVIDIA/SkillEvaluator.

    548 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Task List

    NVIDIA/SkillEvaluator

    Official

    Required for 4+ step requests; add tasks at start and update status after each step.

    548 GitHub stars~460 tokensUpdated yesterday
    Auto-check passed
  • Calculator

    NVIDIA/SkillEvaluator

    Official

    Evaluate mathematical expressions and unit conversions. An agent skill from NVIDIA/SkillEvaluator.

    548 GitHub stars~442 tokensUpdated yesterday
    Auto-check passed
  • Text Analyzer

    NVIDIA/SkillEvaluator

    Official

    Analyze text content and produce statistics including word count, line count, character count, most frequent words, and readability metrics.

    548 GitHub stars~453 tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Create Custom Grader

What does Create Custom Grader do?

A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. Create Custom Grader is an agent skill from NVIDIA/SkillEvaluator, published by the product's own GitHub organization. Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

When should I use Create Custom Grader?

Create Custom Grader fits situations like: converting an existing benchmark; domain check into SkillEvaluator BYOG/BYOT custom evaluation.

How do I install Create Custom Grader in Claude Code?

Run `npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a claude-code`. Or copy the skill folder (src/skillevaluator/tier3/reference_skills/create-custom-grader in NVIDIA/SkillEvaluator) into .claude/skills/create-custom-grader in your project. Claude Code loads it when a task matches its description.

How do I install Create Custom Grader in Codex?

Run `npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a codex`. Or copy the skill folder (src/skillevaluator/tier3/reference_skills/create-custom-grader in NVIDIA/SkillEvaluator) into .agents/skills/create-custom-grader in your project. Codex loads it when a task matches its description.

Can I use Create Custom Grader in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-custom-grader, .gemini/skills/create-custom-grader, .github/skills/create-custom-grader and .opencode/skills/create-custom-grader in your project.

What does Create Custom Grader need to run?

SKILL.md names no scripts, command-line tools or credentials: Create Custom Grader is instructions for the agent only. Our summary lists: Python 3.

Does Create Custom Grader access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Create Custom Grader safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Create Custom Grader use?

Create Custom Grader is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Create Custom Grader use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Create Custom Grader?

Skills that share tags, products or a category with Create Custom Grader: Task Creator (benchflow-ai/benchflow, 353 stars), Skill Judge (shareAI-lab/lab-skills, 314 stars), Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars) and Create Skill Test (dotnet/skills, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Create Custom Grader?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/SkillEvaluator, which has 548 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/SkillEvaluator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.