Agent skill

LLM Benchmarking with lm-evaluation-harness

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

MITAuto-check passedAI & LLM Engineering

Install LLM Benchmarking with lm-evaluation-harness

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-llms-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs evaluating-llms-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/11-evaluation/lm-evaluation-harness .claude/skills/evaluating-llms-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluating-llms-harness
GitHub stars
13k
Used in
8 other repos
Token cost
~3k tokens
SKILL.md length
495 words
Files
5 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

  • Benchmarking a model on MMLU, GSM8K or HumanEval
  • SKILL.md covers Quick start, Common workflows, When to use vs alternatives and Common issues, plus 3 more sections
  • Calls pip
  • Comparing two models on the same academic tasks

What it does

The skill installs `lm-eval` with pip and runs `lm_eval` with a model type, model arguments and a task list, using standardized prompts and metrics across more than 60 academic benchmarks. `lm_eval --tasks list` shows what is available. The standard workflow picks a suite, such as MMLU across 57 subjects, GSM8K, HellaSwag, TruthfulQA and ARC for reasoning, or HumanEval with 164 problems and MBPP for code, then configures the model, runs the evaluation and reads the JSON results.

Model setup covers a Hugging Face model, a quantized 4-bit or 8-bit load and a custom checkpoint with its own tokenizer. A second workflow tracks training progress by evaluating checkpoints at intervals on quick benchmarks. Reference files cover API-based evaluation, a benchmark guide, custom tasks and distributed evaluation, and the description lists Hugging Face, vLLM and hosted APIs as supported backends.

When your agent uses it

  • Benchmarking a model on MMLU, GSM8K or HumanEval
  • Comparing two models on the same academic tasks
  • Evaluating training checkpoints to track progress
  • Checking a quantized model's quality against the original

Example prompts

  • “Run MMLU and GSM8K on my fine-tuned checkpoint in ./checkpoints/final and save the results as JSON.”
  • “Compare Llama 2 7B against my quantized version on HellaSwag and TruthfulQA.”
  • “Set up evaluation every few training steps so I can track MMLU during fine-tuning.”
  • “Show me which tasks lm_eval has for code generation.”

Requirements

  • Python with the `lm-eval` package
  • A Hugging Face model, local checkpoint or model API to evaluate

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Benchmarking with lm-evaluation-harness loads about 3k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 495 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 495 words, ~2,974 tokens.

Download SKILL.mdSave it as .claude/skills/evaluating-llms-harness/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
evaluating-llms-harness
description
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Evaluation, LM Evaluation Harness, Benchmarking, MMLU, HumanEval, GSM8K, EleutherAI, Model Quality, Academic Benchmarks, Industry Standard
dependencies
lm-eval, transformers, vllm

lm-evaluation-harness - LLM Benchmarking

Quick start

lm-evaluation-harness evaluates LLMs across 60+ academic benchmarks using standardized prompts and metrics.

Installation:

bash
pip install lm-eval

Evaluate any HuggingFace model:

bash
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf \
  --tasks mmlu,gsm8k,hellaswag \
  --device cuda:0 \
  --batch_size 8

View available tasks:

bash
lm_eval --tasks list

Common workflows

Workflow 1: Standard benchmark evaluation

Evaluate model on core benchmarks (MMLU, GSM8K, HumanEval).

Copy this checklist:

Benchmark Evaluation:
- [ ] Step 1: Choose benchmark suite
- [ ] Step 2: Configure model
- [ ] Step 3: Run evaluation
- [ ] Step 4: Analyze results

Step 1: Choose benchmark suite

Core reasoning benchmarks:

  • MMLU (Massive Multitask Language Understanding) - 57 subjects, multiple choice
  • GSM8K - Grade school math word problems
  • HellaSwag - Common sense reasoning
  • TruthfulQA - Truthfulness and factuality
  • ARC (AI2 Reasoning Challenge) - Science questions

Code benchmarks:

  • HumanEval - Python code generation (164 problems)
  • MBPP (Mostly Basic Python Problems) - Python coding

Standard suite (recommended for model releases):

bash
--tasks mmlu,gsm8k,hellaswag,truthfulqa,arc_challenge

Step 2: Configure model

HuggingFace model:

bash
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf,dtype=bfloat16 \
  --tasks mmlu \
  --device cuda:0 \
  --batch_size auto  # Auto-detect optimal batch size

Quantized model (4-bit/8-bit):

bash
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf,load_in_4bit=True \
  --tasks mmlu \
  --device cuda:0

Custom checkpoint:

bash
lm_eval --model hf \
  --model_args pretrained=/path/to/my-model,tokenizer=/path/to/tokenizer \
  --tasks mmlu \
  --device cuda:0

Step 3: Run evaluation

bash
# Full MMLU evaluation (57 subjects)
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf \
  --tasks mmlu \
  --num_fewshot 5 \  # 5-shot evaluation (standard)
  --batch_size 8 \
  --output_path results/ \
  --log_samples  # Save individual predictions

# Multiple benchmarks at once
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf \
  --tasks mmlu,gsm8k,hellaswag,truthfulqa,arc_challenge \
  --num_fewshot 5 \
  --batch_size 8 \
  --output_path results/llama2-7b-eval.json

Step 4: Analyze results

Results saved to results/llama2-7b-eval.json:

json
{
  "results": {
    "mmlu": {
      "acc": 0.459,
      "acc_stderr": 0.004
    },
    "gsm8k": {
      "exact_match": 0.142,
      "exact_match_stderr": 0.006
    },
    "hellaswag": {
      "acc_norm": 0.765,
      "acc_norm_stderr": 0.004
    }
  },
  "config": {
    "model": "hf",
    "model_args": "pretrained=meta-llama/Llama-2-7b-hf",
    "num_fewshot": 5
  }
}
Workflow 2: Track training progress

Evaluate checkpoints during training.

Training Progress Tracking:
- [ ] Step 1: Set up periodic evaluation
- [ ] Step 2: Choose quick benchmarks
- [ ] Step 3: Automate evaluation
- [ ] Step 4: Plot learning curves

Step 1: Set up periodic evaluation

Evaluate every N training steps:

bash
#!/bin/bash
# eval_checkpoint.sh

CHECKPOINT_DIR=$1
STEP=$2

lm_eval --model hf \
  --model_args pretrained=$CHECKPOINT_DIR/checkpoint-$STEP \
  --tasks gsm8k,hellaswag \
  --num_fewshot 0 \  # 0-shot for speed
  --batch_size 16 \
  --output_path results/step-$STEP.json

Step 2: Choose quick benchmarks

Fast benchmarks for frequent evaluation:

  • HellaSwag: ~10 minutes on 1 GPU
  • GSM8K: ~5 minutes
  • PIQA: ~2 minutes

Avoid for frequent eval (too slow):

  • MMLU: ~2 hours (57 subjects)
  • HumanEval: Requires code execution

Step 3: Automate evaluation

Integrate with training script:

python
# In training loop
if step % eval_interval == 0:
    model.save_pretrained(f"checkpoints/step-{step}")

    # Run evaluation
    os.system(f"./eval_checkpoint.sh checkpoints step-{step}")

Or use PyTorch Lightning callbacks:

python
from pytorch_lightning import Callback

class EvalHarnessCallback(Callback):
    def on_validation_epoch_end(self, trainer, pl_module):
        step = trainer.global_step
        checkpoint_path = f"checkpoints/step-{step}"

        # Save checkpoint
        trainer.save_checkpoint(checkpoint_path)

        # Run lm-eval
        os.system(f"lm_eval --model hf --model_args pretrained={checkpoint_path} ...")

Step 4: Plot learning curves

python
import json
import matplotlib.pyplot as plt

# Load all results
steps = []
mmlu_scores = []

for file in sorted(glob.glob("results/step-*.json")):
    with open(file) as f:
        data = json.load(f)
        step = int(file.split("-")[1].split(".")[0])
        steps.append(step)
        mmlu_scores.append(data["results"]["mmlu"]["acc"])

# Plot
plt.plot(steps, mmlu_scores)
plt.xlabel("Training Step")
plt.ylabel("MMLU Accuracy")
plt.title("Training Progress")
plt.savefig("training_curve.png")
Workflow 3: Compare multiple models

Benchmark suite for model comparison.

Model Comparison:
- [ ] Step 1: Define model list
- [ ] Step 2: Run evaluations
- [ ] Step 3: Generate comparison table

Step 1: Define model list

bash
# models.txt
meta-llama/Llama-2-7b-hf
meta-llama/Llama-2-13b-hf
mistralai/Mistral-7B-v0.1
microsoft/phi-2

Step 2: Run evaluations

bash
#!/bin/bash
# eval_all_models.sh

TASKS="mmlu,gsm8k,hellaswag,truthfulqa"

while read model; do
    echo "Evaluating $model"

    # Extract model name for output file
    model_name=$(echo $model | sed 's/\//-/g')

    lm_eval --model hf \
      --model_args pretrained=$model,dtype=bfloat16 \
      --tasks $TASKS \
      --num_fewshot 5 \
      --batch_size auto \
      --output_path results/$model_name.json

done < models.txt

Step 3: Generate comparison table

python
import json
import pandas as pd

models = [
    "meta-llama-Llama-2-7b-hf",
    "meta-llama-Llama-2-13b-hf",
    "mistralai-Mistral-7B-v0.1",
    "microsoft-phi-2"
]

tasks = ["mmlu", "gsm8k", "hellaswag", "truthfulqa"]

results = []
for model in models:
    with open(f"results/{model}.json") as f:
        data = json.load(f)
        row = {"Model": model.replace("-", "/")}
        for task in tasks:
            # Get primary metric for each task
            metrics = data["results"][task]
            if "acc" in metrics:
                row[task.upper()] = f"{metrics['acc']:.3f}"
            elif "exact_match" in metrics:
                row[task.upper()] = f"{metrics['exact_match']:.3f}"
        results.append(row)

df = pd.DataFrame(results)
print(df.to_markdown(index=False))

Output:

| Model                  | MMLU  | GSM8K | HELLASWAG | TRUTHFULQA |
|------------------------|-------|-------|-----------|------------|
| meta-llama/Llama-2-7b  | 0.459 | 0.142 | 0.765     | 0.391      |
| meta-llama/Llama-2-13b | 0.549 | 0.287 | 0.801     | 0.430      |
| mistralai/Mistral-7B   | 0.626 | 0.395 | 0.812     | 0.428      |
| microsoft/phi-2        | 0.560 | 0.613 | 0.682     | 0.447      |
Workflow 4: Evaluate with vLLM (faster inference)

Use vLLM backend for 5-10x faster evaluation.

vLLM Evaluation:
- [ ] Step 1: Install vLLM
- [ ] Step 2: Configure vLLM backend
- [ ] Step 3: Run evaluation

Step 1: Install vLLM

bash
pip install vllm

Step 2: Configure vLLM backend

bash
lm_eval --model vllm \
  --model_args pretrained=meta-llama/Llama-2-7b-hf,tensor_parallel_size=1,dtype=auto,gpu_memory_utilization=0.8 \
  --tasks mmlu \
  --batch_size auto

Step 3: Run evaluation

vLLM is 5-10× faster than standard HuggingFace:

bash
# Standard HF: ~2 hours for MMLU on 7B model
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf \
  --tasks mmlu \
  --batch_size 8

# vLLM: ~15-20 minutes for MMLU on 7B model
lm_eval --model vllm \
  --model_args pretrained=meta-llama/Llama-2-7b-hf,tensor_parallel_size=2 \
  --tasks mmlu \
  --batch_size auto

When to use vs alternatives

Use lm-evaluation-harness when:

  • Benchmarking models for academic papers
  • Comparing model quality across standard tasks
  • Tracking training progress
  • Reporting standardized metrics (everyone uses same prompts)
  • Need reproducible evaluation

Use alternatives instead:

  • HELM (Stanford): Broader evaluation (fairness, efficiency, calibration)
  • AlpacaEval: Instruction-following evaluation with LLM judges
  • MT-Bench: Conversational multi-turn evaluation
  • Custom scripts: Domain-specific evaluation
Show full SKILL.md (181 more words)Show less

Common issues

Issue: Evaluation too slow

Use vLLM backend:

bash
lm_eval --model vllm \
  --model_args pretrained=model-name,tensor_parallel_size=2

Or reduce fewshot examples:

bash
--num_fewshot 0  # Instead of 5

Or evaluate subset of MMLU:

bash
--tasks mmlu_stem  # Only STEM subjects

Issue: Out of memory

Reduce batch size:

bash
--batch_size 1  # Or --batch_size auto

Use quantization:

bash
--model_args pretrained=model-name,load_in_8bit=True

Enable CPU offloading:

bash
--model_args pretrained=model-name,device_map=auto,offload_folder=offload

Issue: Different results than reported

Check fewshot count:

bash
--num_fewshot 5  # Most papers use 5-shot

Check exact task name:

bash
--tasks mmlu  # Not mmlu_direct or mmlu_fewshot

Verify model and tokenizer match:

bash
--model_args pretrained=model-name,tokenizer=same-model-name

Issue: HumanEval not executing code

Install execution dependencies:

bash
pip install human-eval

Enable code execution:

bash
lm_eval --model hf \
  --model_args pretrained=model-name \
  --tasks humaneval \
  --allow_code_execution  # Required for HumanEval

Advanced topics

Benchmark descriptions: See references/benchmark-guide.md for detailed description of all 60+ tasks, what they measure, and interpretation.

Custom tasks: See references/custom-tasks.md for creating domain-specific evaluation tasks.

API evaluation: See references/api-evaluation.md for evaluating OpenAI, Anthropic, and other API models.

Multi-GPU strategies: See references/distributed-eval.md for data parallel and tensor parallel evaluation.

Hardware requirements

  • GPU: NVIDIA (CUDA 11.8+), works on CPU (very slow)
  • VRAM:
    • 7B model: 16GB (bf16) or 8GB (8-bit)
    • 13B model: 28GB (bf16) or 14GB (8-bit)
    • 70B model: Requires multi-GPU or quantization
  • Time (7B model, single A100):
    • HellaSwag: 10 minutes
    • GSM8K: 5 minutes
    • MMLU (full): 2 hours
    • HumanEval: 20 minutes

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in 11-evaluation/lm-evaluation-harness of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/api-evaluation.md
  • references/benchmark-guide.md
  • references/custom-tasks.md
  • references/distributed-eval.md

Open the folder on GitHubat commit 773a529

Used in 8 other repositories

We found 9 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 8 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LLM Benchmarking with lm-evaluation-harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Benchmarking with lm-evaluation-harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Benchmarking with lm-evaluation-harness this skillOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Dataset FinderLeoYeAI/openclaw-master-skills2.2k—~5.4kAutomated safety check: PassProprietary
Aqua Model Lifecycleoracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.0
Hugging Face Community Evalshenryalouf/ruflow157—~1.6kAutomated safety check: PassMIT
Hugging Face Community Evalssickn33/agentic-awesome-skills47k1 repos~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dataset Finder

    LeoYeAI/openclaw-master-skills

    A skill your agent uses when users need to search for datasets, download data files, or explore data repositories.

    2.2k GitHub stars~5.4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Aqua Model Lifecycle

    oracle/accelerated-data-science

    Official

    Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

    125 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval.

    157 GitHub stars~1.6k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Community Evals

    sickn33/agentic-awesome-skills

    Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

    47k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface Community Evals

    sickn33/agentic-awesome-skills

    Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal.

    47k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes
  • FAISS Similarity Search

    Orchestra-Research/AI-Research-SKILLs

    Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.

    13k GitHub starsUsed in 6 repos~1.3k tokens
    Auto-check passed

Questions about LLM Benchmarking with lm-evaluation-harness

What does LLM Benchmarking with lm-evaluation-harness do?

Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints. The skill installs `lm-eval` with pip and runs `lm_eval` with a model type, model arguments and a task list, using standardized prompts and metrics across more than 60 academic benchmarks. `lm_eval --tasks list` shows what is available.

When should I use LLM Benchmarking with lm-evaluation-harness?

LLM Benchmarking with lm-evaluation-harness fits situations like: benchmarking a model on MMLU, GSM8K or HumanEval; comparing two models on the same academic tasks; evaluating training checkpoints to track progress; checking a quantized model's quality against the original.

How do I install LLM Benchmarking with lm-evaluation-harness in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-llms-harness -a claude-code`. Or copy the skill folder (11-evaluation/lm-evaluation-harness in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/evaluating-llms-harness in your project. Claude Code loads it when a task matches its description.

How do I install LLM Benchmarking with lm-evaluation-harness in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-llms-harness -a codex`. Or copy the skill folder (11-evaluation/lm-evaluation-harness in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/evaluating-llms-harness in your project. Codex loads it when a task matches its description.

Can I use LLM Benchmarking with lm-evaluation-harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evaluating-llms-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluating-llms-harness, .gemini/skills/evaluating-llms-harness, .github/skills/evaluating-llms-harness and .opencode/skills/evaluating-llms-harness in your project.

What does LLM Benchmarking with lm-evaluation-harness need to run?

Going by SKILL.md and its folder, LLM Benchmarking with lm-evaluation-harness needs the command-line tools its instructions call (pip). Our summary lists: Python with the `lm-eval` package; A Hugging Face model, local checkpoint or model API to evaluate.

Does LLM Benchmarking with lm-evaluation-harness access the network?

SKILL.md names 2 domains. As links in the text: github.com and huggingface.co. This is read from the text; nothing was executed.

Is LLM Benchmarking with lm-evaluation-harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Benchmarking with lm-evaluation-harness use?

LLM Benchmarking with lm-evaluation-harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Benchmarking with lm-evaluation-harness use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.

What are the alternatives to LLM Benchmarking with lm-evaluation-harness?

Skills that share tags, products or a category with LLM Benchmarking with lm-evaluation-harness: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Dataset Finder (LeoYeAI/openclaw-master-skills, 2.2k stars), Aqua Model Lifecycle (oracle/accelerated-data-science, 125 stars) and Hugging Face Community Evals (henryalouf/ruflow, 157 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Benchmarking with lm-evaluation-harness?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.