Task Creator
benchflow-ai/benchflow
SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.
A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
$ npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/SkillEvaluator create-custom-grader --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/SkillEvaluator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skillevaluator/tier3/reference_skills/create-custom-grader .claude/skills/create-custom-grader && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "create-custom-grader" agent skill from https://github.com/NVIDIA/SkillEvaluator/tree/main/src/skillevaluator/tier3/reference_skills/create-custom-grader into .claude/skills/create-custom-grader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-custom-grader", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/SkillEvaluator/tree/main/src/skillevaluator/tier3/reference_skills/create-custom-graderType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/SkillEvaluator create-custom-grader --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/SkillEvaluator.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/skillevaluator/tier3/reference_skills/create-custom-grader .agents/skills/create-custom-grader && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "create-custom-grader" agent skill from https://github.com/NVIDIA/SkillEvaluator/tree/main/src/skillevaluator/tier3/reference_skills/create-custom-grader into .agents/skills/create-custom-grader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-custom-grader", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/SkillEvaluator create-custom-grader --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/SkillEvaluator.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/skillevaluator/tier3/reference_skills/create-custom-grader .cursor/skills/create-custom-grader && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "create-custom-grader" agent skill from https://github.com/NVIDIA/SkillEvaluator/tree/main/src/skillevaluator/tier3/reference_skills/create-custom-grader into .cursor/skills/create-custom-grader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-custom-grader", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/SkillEvaluator.git --path src/skillevaluator/tier3/reference_skills/create-custom-grader--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/SkillEvaluator create-custom-grader --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/SkillEvaluator.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/skillevaluator/tier3/reference_skills/create-custom-grader .gemini/skills/create-custom-grader && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "create-custom-grader" agent skill from https://github.com/NVIDIA/SkillEvaluator/tree/main/src/skillevaluator/tier3/reference_skills/create-custom-grader into .gemini/skills/create-custom-grader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-custom-grader", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/SkillEvaluator create-custom-graderInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/SkillEvaluator.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/skillevaluator/tier3/reference_skills/create-custom-grader .github/skills/create-custom-grader && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "create-custom-grader" agent skill from https://github.com/NVIDIA/SkillEvaluator/tree/main/src/skillevaluator/tier3/reference_skills/create-custom-grader into .github/skills/create-custom-grader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-custom-grader", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/SkillEvaluator create-custom-grader --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/SkillEvaluator.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/skillevaluator/tier3/reference_skills/create-custom-grader .opencode/skills/create-custom-grader && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "create-custom-grader" agent skill from https://github.com/NVIDIA/SkillEvaluator/tree/main/src/skillevaluator/tier3/reference_skills/create-custom-grader into .opencode/skills/create-custom-grader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-custom-grader", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
create-custom-graderA skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
Create Custom Grader is an agent skill from NVIDIA/SkillEvaluator, published by the product's own GitHub organization. Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).
It sits in Education, covering Quizzes and assessments, Agent evaluation and testing and LLM evaluation. The repository describes itself as: Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that… The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f32c884. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and json).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Create Custom Grader loads about 2.1k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 950 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/SkillEvaluator at commit f32c884, republished under its Apache-2.0 licence (© NVIDIA). 950 words, ~2,147 tokens.
.claude/skills/create-custom-grader/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Convert team-owned benchmark definitions into runnable SkillEvaluator custom graders and, when needed, native Harbor tasks.
Help an agent author valid SkillEvaluator BYOG/BYOT files from a user's benchmark instead of leaving the user with empty grader templates.
Use this skill when the user wants to:
evals/grader.py or evals/grader.shtask.yaml, task.json, pytest checks, or shell
verifiers into BYOG or BYOTDo not use this skill for ordinary evals/evals.json authoring when no custom
grading logic is needed. Use the normal dataset authoring workflow for that.
evals/, benchmark prompts, fixtures, and any verifier code.default_plus_custom when custom metrics should complement default evaluator scoring.custom_only only when the user wants the custom grader to own pass/fail semantics.evals/grader.py or evals/grader.sh, then validate the Harbor contract.skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_custom
skillevaluator tier3 validate <skill-dir>SKILL.md.skillevaluator.Choose one path before writing files:
| User need | Evaluator shape |
|---|---|
Existing evals.json task plus extra domain checks | Top-level BYOG: evals/grader.py or evals/grader.sh |
| Existing benchmark prompt/rubric that can run in the generated workspace | Top-level BYOG plus evals/evals.json and evals/files/ |
| Benchmark owns task layout, setup, service lifecycle, or verifier harness | Native BYOT/BYOG: evals/harbor/<case>/... |
| User wants only custom reward/pass criteria | grading.mode: custom_only |
| User wants default evaluator dimensions plus custom metrics | grading.mode: default_plus_custom |
Default to default_plus_custom unless the user explicitly wants the custom
grader to replace the default evaluator metrics.
Resolve the target skill and benchmark source.
Read the target SKILL.md, existing evals/, benchmark prompts, fixtures,
rubric, reference solution, tags, and any expected trigger/non-trigger
metadata.
Map benchmark fields into evaluator inputs.
Use benchmark prompts or prompt variants as question entries. Use the
target skill as expected_skill. Put each case's required starter files
under evals/files/<case-id>/, and declare
files: ["evals/files/<case-id>"] on every corresponding eval entry. Do not
omit files in a multi-case dataset, because omission intentionally stages
the entire shared directory for legacy compatibility. Preserve
benchmark-specific rubric text in the entry only when the grader needs to
read it.
Scaffold the evaluator contract. For generated tasks:
skillevaluator init-custom-grader <skill-dir> --language python --mode default_plus_customFor shell checks:
skillevaluator init-custom-grader <skill-dir> --language shell --mode default_plus_customFor native Harbor tasks:
skillevaluator init-harbor-task <skill-dir> --case-id <case-id> --with-configReplace scaffold placeholders. The custom grader is real executable logic, not metadata. It must read available evidence, compute numeric scores, and write the evaluator reward contract.
Validate before running.
skillevaluator validate <skill-dir> --harbor-contractFix missing files, invalid Python, missing reward output, and native Harbor ID mismatches before evaluation.
Run the deepest practical proof. Prefer a real with-skill/baseline run. If services, credentials, GPU, or cost block full E2E, state exactly what was validated and what was not.
Python and shell graders run inside the Harbor verifier context. They may read:
/logs/agent/trajectory.json for agent actions and final answer evidence/tests/entry.json for the eval case metadata/workspace/input/ for the entry's declared committed fixtures from
evals/files//solution/ or other task outputs only when the task environment produces
themThey must write:
/logs/verifier/reward.json/logs/verifier/reward.txt with a numeric score from 0.0 to 1.0Use this reward shape:
{
"overall": 0.92,
"custom_metrics": {
"domain_repair": 1.0,
"domain_verification": 0.8
},
"details": {
"domain_repair": {
"score": 1.0,
"reason": "The solution repaired the required files."
}
}
}In default_plus_custom, default evaluator scoring keeps its overall
authoritative and adds the grader's custom_metrics into reports. In
custom_only, the grader's overall is the pass/fail reward.
Never emit custom metric names that collide with reserved evaluator fields:
security, skill_execution, skill_efficiency, accuracy,
goal_accuracy, behavior_check, overall, details, metrics,
metric_set, or entry_id.
details.0.0 through 1.0.expected_skill,
expected_behavior, negative cases, or custom metrics that inspect
trajectory evidence.For a benchmark task with task.yaml, code/, prompt variants, coverage, and a
rubric:
code/ into evals/files/<case-id>/.evals/evals.json entries from the prompt variants, and
set files: ["evals/files/<case-id>"] on each corresponding entry.expected_skill to the benchmark's target skill.evals/grader.py to inspect the agent trajectory and changed
workspace files.rapids_diagnosis, rapids_requirements_repair,
rapids_repair_safety, and rapids_verification.init-custom-grader creates scaffolding only; the agent must replace the
placeholder scoring logic.| Problem | Fix |
|---|---|
evals/evals.json missing | Create entries from the benchmark prompt or run init-custom-grader to seed one. |
| Custom metrics do not appear | Ensure reward.json has numeric values under custom_metrics and no reserved-name collisions. |
custom_only fails | Write numeric overall in reward.json or numeric reward.txt. |
| Grader scores copied fixtures | Restrict file searches to generated workspace/output paths, not the skill package or grader source. |
When finished, report:
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in src/skillevaluator/tier3/reference_skills/create-custom-grader of NVIDIA/SkillEvaluator.
Open the folder on GitHubat commit f32c884
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/SkillEvaluator, which our catalogue first saw on October 7, 2026.
Create Custom Grader next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Create Custom Grader this skillNVIDIA/SkillEvaluator | 548 | 1 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Task Creatorbenchflow-ai/benchflow | 353 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | |
| Skill JudgeshareAI-lab/lab-skills | 314 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Woo AI Smokewoocommerce/woocommerce-ios | 358 | 1 repos | ~7.4k | Automated safety check: Notes | GPL-2.0 | |
| Create Skill Testdotnet/skills | 5.6k | 1 repos | ~5.7k | Automated safety check: Pass | MIT | |
| Agent Evaluationseb1n/awesome-ai-agent-skills | 206 | — | ~1.4k | Automated safety check: Pass | MIT |
benchflow-ai/benchflow
SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.
shareAI-lab/lab-skills
Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.
woocommerce/woocommerce-ios
Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.
dotnet/skills
Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository.
seb1n/awesome-ai-agent-skills
Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis.
aiskillstore/marketplace
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…
NVIDIA/SkillEvaluator
Call any REST API dynamically. An agent skill from NVIDIA/SkillEvaluator.
NVIDIA/SkillEvaluator
Required for 4+ step requests; add tasks at start and update status after each step.
NVIDIA/SkillEvaluator
Evaluate mathematical expressions and unit conversions. An agent skill from NVIDIA/SkillEvaluator.
NVIDIA/SkillEvaluator
Analyze text content and produce statistics including word count, line count, character count, most frequent words, and readability metrics.
Categories
A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. Create Custom Grader is an agent skill from NVIDIA/SkillEvaluator, published by the product's own GitHub organization. Use when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.
Create Custom Grader fits situations like: converting an existing benchmark; domain check into SkillEvaluator BYOG/BYOT custom evaluation.
Run `npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a claude-code`. Or copy the skill folder (src/skillevaluator/tier3/reference_skills/create-custom-grader in NVIDIA/SkillEvaluator) into .claude/skills/create-custom-grader in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a codex`. Or copy the skill folder (src/skillevaluator/tier3/reference_skills/create-custom-grader in NVIDIA/SkillEvaluator) into .agents/skills/create-custom-grader in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/SkillEvaluator --skill create-custom-grader -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-custom-grader, .gemini/skills/create-custom-grader, .github/skills/create-custom-grader and .opencode/skills/create-custom-grader in your project.
SKILL.md names no scripts, command-line tools or credentials: Create Custom Grader is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Create Custom Grader is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Create Custom Grader: Task Creator (benchflow-ai/benchflow, 353 stars), Skill Judge (shareAI-lab/lab-skills, 314 stars), Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars) and Create Skill Test (dotnet/skills, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/SkillEvaluator, which has 548 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/SkillEvaluator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.