Agent skill

Evaluating LLMs Harness

by Luciole-Studio in Luciole-Studio/Misaka-Agent

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.

MITAuto-check passedAI & LLM Engineering

Install Evaluating LLMs Harness

skills CLI
$ npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Luciole-Studio/Misaka-Agent evaluating-llms-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness .claude/skills/evaluating-llms-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluating-llms-harness
GitHub stars
158
Used in
2 other repos
Token cost
~3.1k tokens
SKILL.md length
563 words
Files
5 (incl. references)
Skills in repo
77
Repo updated
First seen
Licence
MIT

At a glance

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.

  • Tasks that involve LLM evaluation
  • SKILL.md covers What's inside, Quick start, Common workflows and When to use vs alternatives, plus 4 more sections
  • Calls pip

What it does

Evaluating LLMs Harness is an agent skill from Luciole-Studio/Misaka-Agent. lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/api-evaluation.md`, `references/benchmark-guide.md` and `references/custom-tasks.md`).

It sits in AI & LLM Engineering, covering LLM evaluation. It works with vLLM, Hugging Face and Python. The repository describes itself as: A multi-agent research system for the humanities and social sciences. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “/evaluating-llms-harness”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 3bcf7a3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluating LLMs Harness loads about 3.1k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 19 tokens; SKILL.md has 563 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Luciole-Studio/Misaka-Agent at commit 3bcf7a3, republished under its MIT licence (© Luciole-Studio). 563 words, ~3,051 tokens.

Download SKILL.mdSave it as .claude/skills/evaluating-llms-harness/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
evaluating-llms-harness
description
lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).
version
1.0.1
author
Orchestra Research
license
MIT
dependencies
lm-eval, transformers, vllm
platforms
linux, macos

lm-evaluation-harness - LLM Benchmarking

What's inside

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

Quick start

lm-evaluation-harness evaluates LLMs across 60+ academic benchmarks using standardized prompts and metrics.

Installation:

bash
pip install lm-eval

Evaluate any HuggingFace model:

bash
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf \
  --tasks mmlu,gsm8k,hellaswag \
  --device cuda:0 \
  --batch_size 8

View available tasks:

bash
lm-eval ls tasks

Common workflows

Workflow 1: Standard benchmark evaluation

Evaluate model on core benchmarks (MMLU, GSM8K, HumanEval).

Copy this checklist:

Benchmark Evaluation:
- [ ] Step 1: Choose benchmark suite
- [ ] Step 2: Configure model
- [ ] Step 3: Run evaluation
- [ ] Step 4: Analyze results

Step 1: Choose benchmark suite

Core reasoning benchmarks:

  • MMLU (Massive Multitask Language Understanding) - 57 subjects, multiple choice
  • GSM8K - Grade school math word problems
  • HellaSwag - Common sense reasoning
  • TruthfulQA - Truthfulness and factuality
  • ARC (AI2 Reasoning Challenge) - Science questions

Code benchmarks:

  • HumanEval - Python code generation (164 problems)
  • MBPP (Mostly Basic Python Problems) - Python coding

Standard suite (recommended for model releases):

bash
--tasks mmlu,gsm8k,hellaswag,truthfulqa,arc_challenge

Step 2: Configure model

HuggingFace model:

bash
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf,dtype=bfloat16 \
  --tasks mmlu \
  --device cuda:0 \
  --batch_size auto  # Auto-detect optimal batch size

Quantized model (4-bit/8-bit):

bash
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf,load_in_4bit=True \
  --tasks mmlu \
  --device cuda:0

Custom checkpoint:

bash
lm_eval --model hf \
  --model_args pretrained=/path/to/my-model,tokenizer=/path/to/tokenizer \
  --tasks mmlu \
  --device cuda:0

Step 3: Run evaluation

bash
# Full MMLU evaluation (57 subjects)
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf \
  --tasks mmlu \
  --num_fewshot 5 \  # 5-shot evaluation (standard)
  --batch_size 8 \
  --output_path results/ \
  --log_samples  # Save individual predictions

# Multiple benchmarks at once
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf \
  --tasks mmlu,gsm8k,hellaswag,truthfulqa,arc_challenge \
  --num_fewshot 5 \
  --batch_size 8 \
  --output_path results/llama2-7b-eval.json

Step 4: Analyze results

Results saved to results/llama2-7b-eval.json:

json
{
  "results": {
    "mmlu": {
      "acc": 0.459,
      "acc_stderr": 0.004
    },
    "gsm8k": {
      "exact_match": 0.142,
      "exact_match_stderr": 0.006
    },
    "hellaswag": {
      "acc_norm": 0.765,
      "acc_norm_stderr": 0.004
    }
  },
  "config": {
    "model": "hf",
    "model_args": "pretrained=meta-llama/Llama-2-7b-hf",
    "num_fewshot": 5
  }
}
Workflow 2: Track training progress

Evaluate checkpoints during training.

Training Progress Tracking:
- [ ] Step 1: Set up periodic evaluation
- [ ] Step 2: Choose quick benchmarks
- [ ] Step 3: Automate evaluation
- [ ] Step 4: Plot learning curves

Step 1: Set up periodic evaluation

Evaluate every N training steps:

bash
#!/bin/bash
# eval_checkpoint.sh

CHECKPOINT_DIR=$1
STEP=$2

lm_eval --model hf \
  --model_args pretrained=$CHECKPOINT_DIR/checkpoint-$STEP \
  --tasks gsm8k,hellaswag \
  --num_fewshot 0 \  # 0-shot for speed
  --batch_size 16 \
  --output_path results/step-$STEP.json

Step 2: Choose quick benchmarks

Fast benchmarks for frequent evaluation:

  • HellaSwag: ~10 minutes on 1 GPU
  • GSM8K: ~5 minutes
  • PIQA: ~2 minutes

Avoid for frequent eval (too slow):

  • MMLU: ~2 hours (57 subjects)
  • HumanEval: Requires code execution

Step 3: Automate evaluation

Integrate with training script:

python
# In training loop
if step % eval_interval == 0:
    model.save_pretrained(f"checkpoints/step-{step}")

    # Run evaluation
    os.system(f"./eval_checkpoint.sh checkpoints step-{step}")

Or use PyTorch Lightning callbacks:

python
from pytorch_lightning import Callback

class EvalHarnessCallback(Callback):
    def on_validation_epoch_end(self, trainer, pl_module):
        step = trainer.global_step
        checkpoint_path = f"checkpoints/step-{step}"

        # Save checkpoint
        trainer.save_checkpoint(checkpoint_path)

        # Run lm-eval
        os.system(f"lm_eval --model hf --model_args pretrained={checkpoint_path} ...")

Step 4: Plot learning curves

python
import json
import matplotlib.pyplot as plt

# Load all results
steps = []
mmlu_scores = []

for file in sorted(glob.glob("results/step-*.json")):
    with open(file) as f:
        data = json.load(f)
        step = int(file.split("-")[1].split(".")[0])
        steps.append(step)
        mmlu_scores.append(data["results"]["mmlu"]["acc"])

# Plot
plt.plot(steps, mmlu_scores)
plt.xlabel("Training Step")
plt.ylabel("MMLU Accuracy")
plt.title("Training Progress")
plt.savefig("training_curve.png")
Workflow 3: Compare multiple models

Benchmark suite for model comparison.

Model Comparison:
- [ ] Step 1: Define model list
- [ ] Step 2: Run evaluations
- [ ] Step 3: Generate comparison table

Step 1: Define model list

bash
# models.txt
meta-llama/Llama-2-7b-hf
meta-llama/Llama-2-13b-hf
mistralai/Mistral-7B-v0.1
microsoft/phi-2

Step 2: Run evaluations

bash
#!/bin/bash
# eval_all_models.sh

TASKS="mmlu,gsm8k,hellaswag,truthfulqa"

while read model; do
    echo "Evaluating $model"

    # Extract model name for output file
    model_name=$(echo $model | sed 's/\//-/g')

    lm_eval --model hf \
      --model_args pretrained=$model,dtype=bfloat16 \
      --tasks $TASKS \
      --num_fewshot 5 \
      --batch_size auto \
      --output_path results/$model_name.json

done < models.txt

Step 3: Generate comparison table

python
import json
import pandas as pd

models = [
    "meta-llama-Llama-2-7b-hf",
    "meta-llama-Llama-2-13b-hf",
    "mistralai-Mistral-7B-v0.1",
    "microsoft-phi-2"
]

tasks = ["mmlu", "gsm8k", "hellaswag", "truthfulqa"]

results = []
for model in models:
    with open(f"results/{model}.json") as f:
        data = json.load(f)
        row = {"Model": model.replace("-", "/")}
        for task in tasks:
            # Get primary metric for each task
            metrics = data["results"][task]
            if "acc" in metrics:
                row[task.upper()] = f"{metrics['acc']:.3f}"
            elif "exact_match" in metrics:
                row[task.upper()] = f"{metrics['exact_match']:.3f}"
        results.append(row)

df = pd.DataFrame(results)
print(df.to_markdown(index=False))

Output:

| Model                  | MMLU  | GSM8K | HELLASWAG | TRUTHFULQA |
|------------------------|-------|-------|-----------|------------|
| meta-llama/Llama-2-7b  | 0.459 | 0.142 | 0.765     | 0.391      |
| meta-llama/Llama-2-13b | 0.549 | 0.287 | 0.801     | 0.430      |
| mistralai/Mistral-7B   | 0.626 | 0.395 | 0.812     | 0.428      |
| microsoft/phi-2        | 0.560 | 0.613 | 0.682     | 0.447      |
Workflow 4: Evaluate with vLLM (faster inference)

Use vLLM backend for 5-10x faster evaluation.

vLLM Evaluation:
- [ ] Step 1: Install vLLM
- [ ] Step 2: Configure vLLM backend
- [ ] Step 3: Run evaluation

Step 1: Install vLLM

bash
pip install vllm

Step 2: Configure vLLM backend

bash
lm_eval --model vllm \
  --model_args pretrained=meta-llama/Llama-2-7b-hf,tensor_parallel_size=1,dtype=auto,gpu_memory_utilization=0.8 \
  --tasks mmlu \
  --batch_size auto

Step 3: Run evaluation

vLLM is 5-10× faster than standard HuggingFace:

bash
# Standard HF: ~2 hours for MMLU on 7B model
lm_eval --model hf \
  --model_args pretrained=meta-llama/Llama-2-7b-hf \
  --tasks mmlu \
  --batch_size 8

# vLLM: ~15-20 minutes for MMLU on 7B model
lm_eval --model vllm \
  --model_args pretrained=meta-llama/Llama-2-7b-hf,tensor_parallel_size=2 \
  --tasks mmlu \
  --batch_size auto

When to use vs alternatives

Use lm-evaluation-harness when:

  • Benchmarking models for academic papers
  • Comparing model quality across standard tasks
  • Tracking training progress
  • Reporting standardized metrics (everyone uses same prompts)
  • Need reproducible evaluation

Use alternatives instead:

  • HELM (Stanford): Broader evaluation (fairness, efficiency, calibration)
  • AlpacaEval: Instruction-following evaluation with LLM judges
  • MT-Bench: Conversational multi-turn evaluation
  • Custom scripts: Domain-specific evaluation
Show full SKILL.md (209 more words)Show less

Common issues

Issue: Evaluation too slow

Use vLLM backend:

bash
lm_eval --model vllm \
  --model_args pretrained=model-name,tensor_parallel_size=2

Or reduce fewshot examples:

bash
--num_fewshot 0  # Instead of 5

Or evaluate subset of MMLU:

bash
--tasks mmlu_stem  # Only STEM subjects

Issue: Out of memory

Reduce batch size:

bash
--batch_size 1  # Or --batch_size auto

Use quantization:

bash
--model_args pretrained=model-name,load_in_8bit=True

Enable CPU offloading:

bash
--model_args pretrained=model-name,device_map=auto,offload_folder=offload

Issue: Different results than reported

Check fewshot count:

bash
--num_fewshot 5  # Most papers use 5-shot

Check exact task name:

bash
--tasks mmlu  # Not mmlu_direct or mmlu_fewshot

Verify model and tokenizer match:

bash
--model_args pretrained=model-name,tokenizer=same-model-name

Issue: HumanEval not executing code

Code-executing tasks (HumanEval, MBPP, etc.) are gated behind an explicit confirmation flag — you must pass --confirm_run_unsafe_code to run them:

bash
lm_eval --model hf \
  --model_args pretrained=model-name \
  --tasks humaneval \
  --confirm_run_unsafe_code  # Required to run tasks that execute generated code

Without this flag lm-eval refuses to run the task rather than silently skipping code execution.

Advanced topics

Benchmark descriptions: See references/benchmark-guide.md for detailed description of all 60+ tasks, what they measure, and interpretation.

Custom tasks: See references/custom-tasks.md for creating domain-specific evaluation tasks.

API evaluation: See references/api-evaluation.md for evaluating OpenAI, Anthropic, and other API models.

Multi-GPU strategies: See references/distributed-eval.md for data parallel and tensor parallel evaluation.

Hardware requirements

  • GPU: NVIDIA (CUDA 11.8+), works on CPU (very slow)
  • VRAM:
    • 7B model: 16GB (bf16) or 8GB (8-bit)
    • 13B model: 28GB (bf16) or 14GB (8-bit)
    • 70B model: Requires multi-GPU or quantization
  • Time (7B model, single A100):
    • HellaSwag: 10 minutes
    • GSM8K: 5 minutes
    • MMLU (full): 2 hours
    • HumanEval: 20 minutes

Resources

© Luciole-Studio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness of Luciole-Studio/Misaka-Agent.

  • SKILL.md
  • references/api-evaluation.md
  • references/benchmark-guide.md
  • references/custom-tasks.md
  • references/distributed-eval.md

Open the folder on GitHubat commit 3bcf7a3

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Luciole-Studio/Misaka-Agent, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Evaluating LLMs Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluating LLMs Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluating LLMs Harness this skillLuciole-Studio/Misaka-Agent1582 repos~3.1kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Outlines Structured GenerationOrchestra-Research/AI-Research-SKILLs13k9 repos~4kAutomated safety check: PassMIT
Aqua Model Lifecycleoracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.0
Hugging Face Community Evalshenryalouf/ruflow157—~1.6kAutomated safety check: PassMIT

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Outlines Structured Generation

    Orchestra-Research/AI-Research-SKILLs

    Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models.

    13k GitHub starsUsed in 9 repos~4k tokens
    AI & LLM EngineeringAuto-check passed
  • Aqua Model Lifecycle

    oracle/accelerated-data-science

    Official

    Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

    125 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval.

    157 GitHub stars~1.6k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Community Evals

    sickn33/agentic-awesome-skills

    Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

    47k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed

More from Luciole-Studio/Misaka-Agent

All 77 skills in this repo
  • Kanban Video Orchestrator

    Luciole-Studio/Misaka-Agent

    Plan and run multi-agent video production pipelines. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check: notes
  • Ast Grep

    Luciole-Studio/Misaka-Agent

    AST-aware structural code search and rewrite via ast-grep. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Drug Discovery

    Luciole-Studio/Misaka-Agent

    Drug discovery: ChEMBL search, drug-likeness, interactions. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Fitness Nutrition

    Luciole-Studio/Misaka-Agent

    Workout planning, macros, and body metrics via wger/USDA. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Hyperframes

    Luciole-Studio/Misaka-Agent

    Render MP4/WebM videos from HTML compositions. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • Osint Investigation

    Luciole-Studio/Misaka-Agent

    Follow the money via public records and sanctions data. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~2.9k tokens
    Auto-check passed

Questions about Evaluating LLMs Harness

What does Evaluating LLMs Harness do?

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent. Evaluating LLMs Harness is an agent skill from Luciole-Studio/Misaka-Agent.).

When should I use Evaluating LLMs Harness?

Evaluating LLMs Harness fits situations like: tasks that involve LLM evaluation.

How do I install Evaluating LLMs Harness in Claude Code?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a claude-code`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness in Luciole-Studio/Misaka-Agent) into .claude/skills/evaluating-llms-harness in your project. Claude Code loads it when a task matches its description.

How do I install Evaluating LLMs Harness in Codex?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a codex`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness in Luciole-Studio/Misaka-Agent) into .agents/skills/evaluating-llms-harness in your project. Codex loads it when a task matches its description.

Can I use Evaluating LLMs Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluating-llms-harness, .gemini/skills/evaluating-llms-harness, .github/skills/evaluating-llms-harness and .opencode/skills/evaluating-llms-harness in your project.

What does Evaluating LLMs Harness need to run?

Going by SKILL.md and its folder, Evaluating LLMs Harness needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Evaluating LLMs Harness access the network?

SKILL.md names 2 domains. As links in the text: github.com and huggingface.co. This is read from the text; nothing was executed.

Is Evaluating LLMs Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluating LLMs Harness use?

Evaluating LLMs Harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluating LLMs Harness use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.

What are the alternatives to Evaluating LLMs Harness?

Skills that share tags, products or a category with Evaluating LLMs Harness: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Outlines Structured Generation (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Aqua Model Lifecycle (oracle/accelerated-data-science, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluating LLMs Harness?

Luciole-Studio (a GitHub organization) maintains it in Luciole-Studio/Misaka-Agent, which has 158 GitHub stars. The repository holds 77 skills in this directory. The repository was last updated on October 8, 2026.

Source: Luciole-Studio/Misaka-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.