Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.

MITAuto-check passedAI & LLM Engineering

Install Nemo Evaluator SDK

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill nemo-evaluator-sdk -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs nemo-evaluator-sdk --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/11-evaluation/nemo-evaluator .claude/skills/nemo-evaluator-sdk && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-evaluator-sdk
GitHub stars
13k
Used in
2 other repos
Token cost
~3.1k tokens
SKILL.md length
477 words
Files
5 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.

  • Needing scalable evaluation on local Docker
  • SKILL.md covers Quick Start, Common Workflows, When to Use vs Alternatives and Supported Harnesses and Tasks, plus 7 more sections
  • Calls pip and docker; reaches integrate.api.nvidia.com; needs HF_TOKEN and NGC_API_KEY
  • Cloud platforms

What it does

Nemo Evaluator SDK is an agent skill from Orchestra-Research/AI-Research-SKILLs. Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/adapter-system.md`, `references/configuration.md` and `references/custom-benchmarks.md`).

It sits in AI & LLM Engineering, covering LLM evaluation and Containers. It works with Docker and NVIDIA AI Platform. The repository describes itself as: Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent… The licence is MIT.

When your agent uses it

  • Needing scalable evaluation on local Docker
  • Cloud platforms

Example prompts

  • “Use the nemo-evaluator-sdk skill to evaluate LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend…”
  • “/nemo-evaluator-sdk”

Requirements

  • Python 3
  • Docker
  • A credential in NGC_API_KEY
  • A credential in JUDGE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • integrate.api.nvidia.com

    Also links to:

    • github.com
    • build.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • NGC_API_KEY
    • JUDGE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Evaluator SDK loads about 3.1k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 477 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 477 words, ~3,057 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-evaluator-sdk/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
nemo-evaluator-sdk
description
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Evaluation, NeMo, NVIDIA, Benchmarking, MMLU, HumanEval, Multi-Backend, Slurm, Docker, Reproducible, Enterprise
dependencies
nemo-evaluator-launcher>=0.1.25, docker

NeMo Evaluator SDK - Enterprise LLM Benchmarking

Quick Start

NeMo Evaluator SDK evaluates LLMs across 100+ benchmarks from 18+ harnesses using containerized, reproducible evaluation with multi-backend execution (local Docker, Slurm HPC, Lepton cloud).

Installation:

bash
pip install nemo-evaluator-launcher

Set API key and run evaluation:

bash
export NGC_API_KEY=nvapi-your-key-here

# Create minimal config
cat > config.yaml << 'EOF'
defaults:
  - execution: local
  - deployment: none
  - _self_

execution:
  output_dir: ./results

target:
  api_endpoint:
    model_id: meta/llama-3.1-8b-instruct
    url: https://integrate.api.nvidia.com/v1/chat/completions
    api_key_name: NGC_API_KEY

evaluation:
  tasks:
    - name: ifeval
EOF

# Run evaluation
nemo-evaluator-launcher run --config-dir . --config-name config

View available tasks:

bash
nemo-evaluator-launcher ls tasks

Common Workflows

Workflow 1: Evaluate Model on Standard Benchmarks

Run core academic benchmarks (MMLU, GSM8K, IFEval) on any OpenAI-compatible endpoint.

Checklist:

Standard Evaluation:
- [ ] Step 1: Configure API endpoint
- [ ] Step 2: Select benchmarks
- [ ] Step 3: Run evaluation
- [ ] Step 4: Check results

Step 1: Configure API endpoint

yaml
# config.yaml
defaults:
  - execution: local
  - deployment: none
  - _self_

execution:
  output_dir: ./results

target:
  api_endpoint:
    model_id: meta/llama-3.1-8b-instruct
    url: https://integrate.api.nvidia.com/v1/chat/completions
    api_key_name: NGC_API_KEY

For self-hosted endpoints (vLLM, TRT-LLM):

yaml
target:
  api_endpoint:
    model_id: my-model
    url: http://localhost:8000/v1/chat/completions
    api_key_name: ""  # No key needed for local

Step 2: Select benchmarks

Add tasks to your config:

yaml
evaluation:
  tasks:
    - name: ifeval           # Instruction following
    - name: gpqa_diamond     # Graduate-level QA
      env_vars:
        HF_TOKEN: HF_TOKEN   # Some tasks need HF token
    - name: gsm8k_cot_instruct  # Math reasoning
    - name: humaneval        # Code generation

Step 3: Run evaluation

bash
# Run with config file
nemo-evaluator-launcher run \
  --config-dir . \
  --config-name config

# Override output directory
nemo-evaluator-launcher run \
  --config-dir . \
  --config-name config \
  -o execution.output_dir=./my_results

# Limit samples for quick testing
nemo-evaluator-launcher run \
  --config-dir . \
  --config-name config \
  -o +evaluation.nemo_evaluator_config.config.params.limit_samples=10

Step 4: Check results

bash
# Check job status
nemo-evaluator-launcher status <invocation_id>

# List all runs
nemo-evaluator-launcher ls runs

# View results
cat results/<invocation_id>/<task>/artifacts/results.yml
Workflow 2: Run Evaluation on Slurm HPC Cluster

Execute large-scale evaluation on HPC infrastructure.

Checklist:

Slurm Evaluation:
- [ ] Step 1: Configure Slurm settings
- [ ] Step 2: Set up model deployment
- [ ] Step 3: Launch evaluation
- [ ] Step 4: Monitor job status

Step 1: Configure Slurm settings

yaml
# slurm_config.yaml
defaults:
  - execution: slurm
  - deployment: vllm
  - _self_

execution:
  hostname: cluster.example.com
  account: my_slurm_account
  partition: gpu
  output_dir: /shared/results
  walltime: "04:00:00"
  nodes: 1
  gpus_per_node: 8

Step 2: Set up model deployment

yaml
deployment:
  checkpoint_path: /shared/models/llama-3.1-8b
  tensor_parallel_size: 2
  data_parallel_size: 4
  max_model_len: 4096

target:
  api_endpoint:
    model_id: llama-3.1-8b
    # URL auto-generated by deployment

Step 3: Launch evaluation

bash
nemo-evaluator-launcher run \
  --config-dir . \
  --config-name slurm_config

Step 4: Monitor job status

bash
# Check status (queries sacct)
nemo-evaluator-launcher status <invocation_id>

# View detailed info
nemo-evaluator-launcher info <invocation_id>

# Kill if needed
nemo-evaluator-launcher kill <invocation_id>
Workflow 3: Compare Multiple Models

Benchmark multiple models on the same tasks for comparison.

Checklist:

Model Comparison:
- [ ] Step 1: Create base config
- [ ] Step 2: Run evaluations with overrides
- [ ] Step 3: Export and compare results

Step 1: Create base config

yaml
# base_eval.yaml
defaults:
  - execution: local
  - deployment: none
  - _self_

execution:
  output_dir: ./comparison_results

evaluation:
  nemo_evaluator_config:
    config:
      params:
        temperature: 0.01
        parallelism: 4
  tasks:
    - name: mmlu_pro
    - name: gsm8k_cot_instruct
    - name: ifeval

Step 2: Run evaluations with model overrides

bash
# Evaluate Llama 3.1 8B
nemo-evaluator-launcher run \
  --config-dir . \
  --config-name base_eval \
  -o target.api_endpoint.model_id=meta/llama-3.1-8b-instruct \
  -o target.api_endpoint.url=https://integrate.api.nvidia.com/v1/chat/completions

# Evaluate Mistral 7B
nemo-evaluator-launcher run \
  --config-dir . \
  --config-name base_eval \
  -o target.api_endpoint.model_id=mistralai/mistral-7b-instruct-v0.3 \
  -o target.api_endpoint.url=https://integrate.api.nvidia.com/v1/chat/completions

Step 3: Export and compare

bash
# Export to MLflow
nemo-evaluator-launcher export <invocation_id_1> --dest mlflow
nemo-evaluator-launcher export <invocation_id_2> --dest mlflow

# Export to local JSON
nemo-evaluator-launcher export <invocation_id> --dest local --format json

# Export to Weights & Biases
nemo-evaluator-launcher export <invocation_id> --dest wandb
Workflow 4: Safety and Vision-Language Evaluation

Evaluate models on safety benchmarks and VLM tasks.

Checklist:

Safety/VLM Evaluation:
- [ ] Step 1: Configure safety tasks
- [ ] Step 2: Set up VLM tasks (if applicable)
- [ ] Step 3: Run evaluation

Step 1: Configure safety tasks

yaml
evaluation:
  tasks:
    - name: aegis              # Safety harness
    - name: wildguard          # Safety classification
    - name: garak              # Security probing

Step 2: Configure VLM tasks

yaml
# For vision-language models
target:
  api_endpoint:
    type: vlm  # Vision-language endpoint
    model_id: nvidia/llama-3.2-90b-vision-instruct
    url: https://integrate.api.nvidia.com/v1/chat/completions

evaluation:
  tasks:
    - name: ocrbench           # OCR evaluation
    - name: chartqa            # Chart understanding
    - name: mmmu               # Multimodal understanding

When to Use vs Alternatives

Use NeMo Evaluator when:

  • Need 100+ benchmarks from 18+ harnesses in one platform
  • Running evaluations on Slurm HPC clusters or cloud
  • Requiring reproducible containerized evaluation
  • Evaluating against OpenAI-compatible APIs (vLLM, TRT-LLM, NIMs)
  • Need enterprise-grade evaluation with result export (MLflow, W&B)

Use alternatives instead:

  • lm-evaluation-harness: Simpler setup for quick local evaluation
  • bigcode-evaluation-harness: Focused only on code benchmarks
  • HELM: Stanford's broader evaluation (fairness, efficiency)
  • Custom scripts: Highly specialized domain evaluation

Supported Harnesses and Tasks

HarnessTask CountCategories
lm-evaluation-harness60+MMLU, GSM8K, HellaSwag, ARC
simple-evals20+GPQA, MATH, AIME
bigcode-evaluation-harness25+HumanEval, MBPP, MultiPL-E
safety-harness3Aegis, WildGuard
garak1Security probing
vlmevalkit6+OCRBench, ChartQA, MMMU
bfcl6Function calling v2/v3
mtbench2Multi-turn conversation
livecodebench10+Live coding evaluation
helm15Medical domain
nemo-skills8Math, science, agentic
Show full SKILL.md (162 more words)Show less

Common Issues

Issue: Container pull fails

Ensure NGC credentials are configured:

bash
docker login nvcr.io -u '$oauthtoken' -p $NGC_API_KEY

Issue: Task requires environment variable

Some tasks need HF_TOKEN or JUDGE_API_KEY:

yaml
evaluation:
  tasks:
    - name: gpqa_diamond
      env_vars:
        HF_TOKEN: HF_TOKEN  # Maps env var name to env var

Issue: Evaluation timeout

Increase parallelism or reduce samples:

bash
-o +evaluation.nemo_evaluator_config.config.params.parallelism=8
-o +evaluation.nemo_evaluator_config.config.params.limit_samples=100

Issue: Slurm job not starting

Check Slurm account and partition:

yaml
execution:
  account: correct_account
  partition: gpu
  qos: normal  # May need specific QOS

Issue: Different results than expected

Verify configuration matches reported settings:

yaml
evaluation:
  nemo_evaluator_config:
    config:
      params:
        temperature: 0.0  # Deterministic
        num_fewshot: 5    # Check paper's fewshot count

CLI Reference

CommandDescription
runExecute evaluation with config
status <id>Check job status
info <id>View detailed job info
ls tasksList available benchmarks
ls runsList all invocations
export <id>Export results (mlflow/wandb/local)
kill <id>Terminate running job

Configuration Override Examples

bash
# Override model endpoint
-o target.api_endpoint.model_id=my-model
-o target.api_endpoint.url=http://localhost:8000/v1/chat/completions

# Add evaluation parameters
-o +evaluation.nemo_evaluator_config.config.params.temperature=0.5
-o +evaluation.nemo_evaluator_config.config.params.parallelism=8
-o +evaluation.nemo_evaluator_config.config.params.limit_samples=50

# Change execution settings
-o execution.output_dir=/custom/path
-o execution.mode=parallel

# Dynamically set tasks
-o 'evaluation.tasks=[{name: ifeval}, {name: gsm8k}]'

Python API Usage

For programmatic evaluation without the CLI:

python
from nemo_evaluator.core.evaluate import evaluate
from nemo_evaluator.api.api_dataclasses import (
    EvaluationConfig,
    EvaluationTarget,
    ApiEndpoint,
    EndpointType,
    ConfigParams
)

# Configure evaluation
eval_config = EvaluationConfig(
    type="mmlu_pro",
    output_dir="./results",
    params=ConfigParams(
        limit_samples=10,
        temperature=0.0,
        max_new_tokens=1024,
        parallelism=4
    )
)

# Configure target endpoint
target_config = EvaluationTarget(
    api_endpoint=ApiEndpoint(
        model_id="meta/llama-3.1-8b-instruct",
        url="https://integrate.api.nvidia.com/v1/chat/completions",
        type=EndpointType.CHAT,
        api_key="nvapi-your-key-here"
    )
)

# Run evaluation
result = evaluate(eval_cfg=eval_config, target_cfg=target_config)

Advanced Topics

Multi-backend execution: See references/execution-backends.md Configuration deep-dive: See references/configuration.md Adapter and interceptor system: See references/adapter-system.md Custom benchmark integration: See references/custom-benchmarks.md

Requirements

  • Python: 3.10-3.13
  • Docker: Required for local execution
  • NGC API Key: For pulling containers and using NVIDIA Build
  • HF_TOKEN: Required for some benchmarks (GPQA, MMLU)

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in 11-evaluation/nemo-evaluator of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/adapter-system.md
  • references/configuration.md
  • references/custom-benchmarks.md
  • references/execution-backends.md

Open the folder on GitHubat commit 773a529

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Nemo Evaluator SDK next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Evaluator SDK compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Evaluator SDK this skillOrchestra-Research/AI-Research-SKILLs13k2 repos~3.1kAutomated safety check: PassMIT
Convergence TestAMD-AGI/Primus131—~2.1kAutomated safety check: PassCustom licence
Vllm Deploy Dockervllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Generate Nemo Gym Envadithya-s-k/FineEnvs456—~2.1kAutomated safety check: PassApache-2.0
Setup Workshopbrevdev/workshop-build-an-agent146—~2.3kAutomated safety check: NotesApache-2.0
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Convergence Test

    AMD-AGI/Primus

    Run, monitor, stop and report Primus convergence tests -- training a model on a real corpus and checking that the loss curve is healthy -- from a plain-language request such as "run convergence test…

    131 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Generate Nemo Gym Env

    adithya-s-k/FineEnvs

    Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs.

    456 GitHub stars~2.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Setup Workshop

    brevdev/workshop-build-an-agent

    This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a.

    146 GitHub stars~2.3k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Upgrade Deps

    areal-project/AReaL

    Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~6k tokensUpdated today
    DevOps & CloudAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Questions about Nemo Evaluator SDK

What does Nemo Evaluator SDK do?

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Nemo Evaluator SDK is an agent skill from Orchestra-Research/AI-Research-SKILLs. Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.

When should I use Nemo Evaluator SDK?

Nemo Evaluator SDK fits situations like: needing scalable evaluation on local Docker; cloud platforms.

How do I install Nemo Evaluator SDK in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill nemo-evaluator-sdk -a claude-code`. Or copy the skill folder (11-evaluation/nemo-evaluator in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/nemo-evaluator-sdk in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Evaluator SDK in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill nemo-evaluator-sdk -a codex`. Or copy the skill folder (11-evaluation/nemo-evaluator in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/nemo-evaluator-sdk in your project. Codex loads it when a task matches its description.

Can I use Nemo Evaluator SDK in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill nemo-evaluator-sdk -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-evaluator-sdk, .gemini/skills/nemo-evaluator-sdk, .github/skills/nemo-evaluator-sdk and .opencode/skills/nemo-evaluator-sdk in your project.

What does Nemo Evaluator SDK need to run?

Going by SKILL.md and its folder, Nemo Evaluator SDK needs the command-line tools its instructions call (pip and docker) and credentials named HF_TOKEN, NGC_API_KEY and JUDGE_API_KEY. Our summary lists: Python 3; Docker; A credential in NGC_API_KEY; A credential in JUDGE_API_KEY.

Does Nemo Evaluator SDK access the network?

SKILL.md names 3 domains. In commands or code: integrate.api.nvidia.com; the agent is likely to contact it when it follows the instructions. As links in the text: github.com and build.nvidia.com. This is read from the text; nothing was executed.

Is Nemo Evaluator SDK safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Evaluator SDK use?

Nemo Evaluator SDK is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Evaluator SDK use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.9k tokens, read only when the agent opens those files.

What are the alternatives to Nemo Evaluator SDK?

Skills that share tags, products or a category with Nemo Evaluator SDK: Convergence Test (AMD-AGI/Primus, 131 stars), Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars), Generate Nemo Gym Env (adithya-s-k/FineEnvs, 456 stars) and Setup Workshop (brevdev/workshop-build-an-agent, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Evaluator SDK?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,374 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.