LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.
$ npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent evaluating-llms-harness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness .claude/skills/evaluating-llms-harness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluating-llms-harness" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness into .claude/skills/evaluating-llms-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms-harness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent evaluating-llms-harness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness .agents/skills/evaluating-llms-harness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluating-llms-harness" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness into .agents/skills/evaluating-llms-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms-harness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent evaluating-llms-harness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness .cursor/skills/evaluating-llms-harness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluating-llms-harness" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness into .cursor/skills/evaluating-llms-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms-harness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Luciole-Studio/Misaka-Agent.git --path misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent evaluating-llms-harness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness .gemini/skills/evaluating-llms-harness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluating-llms-harness" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness into .gemini/skills/evaluating-llms-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms-harness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Luciole-Studio/Misaka-Agent evaluating-llms-harnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness .github/skills/evaluating-llms-harness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluating-llms-harness" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness into .github/skills/evaluating-llms-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms-harness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent evaluating-llms-harness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness .opencode/skills/evaluating-llms-harness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluating-llms-harness" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness into .opencode/skills/evaluating-llms-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms-harness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluating-llms-harnesslm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.
Evaluating LLMs Harness is an agent skill from Luciole-Studio/Misaka-Agent. lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/api-evaluation.md`, `references/benchmark-guide.md` and `references/custom-tasks.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. It works with vLLM, Hugging Face and Python. The repository describes itself as: A multi-agent research system for the humanities and social sciences. The licence is MIT.
Read from SKILL.md and the folder at commit 3bcf7a3. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comhuggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluating LLMs Harness loads about 3.1k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 19 tokens; SKILL.md has 563 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Luciole-Studio/Misaka-Agent at commit 3bcf7a3, republished under its MIT licence (© Luciole-Studio). 563 words, ~3,051 tokens.
.claude/skills/evaluating-llms-harness/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
lm-evaluation-harness evaluates LLMs across 60+ academic benchmarks using standardized prompts and metrics.
Installation:
pip install lm-evalEvaluate any HuggingFace model:
lm_eval --model hf \
--model_args pretrained=meta-llama/Llama-2-7b-hf \
--tasks mmlu,gsm8k,hellaswag \
--device cuda:0 \
--batch_size 8View available tasks:
lm-eval ls tasksEvaluate model on core benchmarks (MMLU, GSM8K, HumanEval).
Copy this checklist:
Benchmark Evaluation:
- [ ] Step 1: Choose benchmark suite
- [ ] Step 2: Configure model
- [ ] Step 3: Run evaluation
- [ ] Step 4: Analyze resultsStep 1: Choose benchmark suite
Core reasoning benchmarks:
Code benchmarks:
Standard suite (recommended for model releases):
--tasks mmlu,gsm8k,hellaswag,truthfulqa,arc_challengeStep 2: Configure model
HuggingFace model:
lm_eval --model hf \
--model_args pretrained=meta-llama/Llama-2-7b-hf,dtype=bfloat16 \
--tasks mmlu \
--device cuda:0 \
--batch_size auto # Auto-detect optimal batch sizeQuantized model (4-bit/8-bit):
lm_eval --model hf \
--model_args pretrained=meta-llama/Llama-2-7b-hf,load_in_4bit=True \
--tasks mmlu \
--device cuda:0Custom checkpoint:
lm_eval --model hf \
--model_args pretrained=/path/to/my-model,tokenizer=/path/to/tokenizer \
--tasks mmlu \
--device cuda:0Step 3: Run evaluation
# Full MMLU evaluation (57 subjects)
lm_eval --model hf \
--model_args pretrained=meta-llama/Llama-2-7b-hf \
--tasks mmlu \
--num_fewshot 5 \ # 5-shot evaluation (standard)
--batch_size 8 \
--output_path results/ \
--log_samples # Save individual predictions
# Multiple benchmarks at once
lm_eval --model hf \
--model_args pretrained=meta-llama/Llama-2-7b-hf \
--tasks mmlu,gsm8k,hellaswag,truthfulqa,arc_challenge \
--num_fewshot 5 \
--batch_size 8 \
--output_path results/llama2-7b-eval.jsonStep 4: Analyze results
Results saved to results/llama2-7b-eval.json:
{
"results": {
"mmlu": {
"acc": 0.459,
"acc_stderr": 0.004
},
"gsm8k": {
"exact_match": 0.142,
"exact_match_stderr": 0.006
},
"hellaswag": {
"acc_norm": 0.765,
"acc_norm_stderr": 0.004
}
},
"config": {
"model": "hf",
"model_args": "pretrained=meta-llama/Llama-2-7b-hf",
"num_fewshot": 5
}
}Evaluate checkpoints during training.
Training Progress Tracking:
- [ ] Step 1: Set up periodic evaluation
- [ ] Step 2: Choose quick benchmarks
- [ ] Step 3: Automate evaluation
- [ ] Step 4: Plot learning curvesStep 1: Set up periodic evaluation
Evaluate every N training steps:
#!/bin/bash
# eval_checkpoint.sh
CHECKPOINT_DIR=$1
STEP=$2
lm_eval --model hf \
--model_args pretrained=$CHECKPOINT_DIR/checkpoint-$STEP \
--tasks gsm8k,hellaswag \
--num_fewshot 0 \ # 0-shot for speed
--batch_size 16 \
--output_path results/step-$STEP.jsonStep 2: Choose quick benchmarks
Fast benchmarks for frequent evaluation:
Avoid for frequent eval (too slow):
Step 3: Automate evaluation
Integrate with training script:
# In training loop
if step % eval_interval == 0:
model.save_pretrained(f"checkpoints/step-{step}")
# Run evaluation
os.system(f"./eval_checkpoint.sh checkpoints step-{step}")Or use PyTorch Lightning callbacks:
from pytorch_lightning import Callback
class EvalHarnessCallback(Callback):
def on_validation_epoch_end(self, trainer, pl_module):
step = trainer.global_step
checkpoint_path = f"checkpoints/step-{step}"
# Save checkpoint
trainer.save_checkpoint(checkpoint_path)
# Run lm-eval
os.system(f"lm_eval --model hf --model_args pretrained={checkpoint_path} ...")Step 4: Plot learning curves
import json
import matplotlib.pyplot as plt
# Load all results
steps = []
mmlu_scores = []
for file in sorted(glob.glob("results/step-*.json")):
with open(file) as f:
data = json.load(f)
step = int(file.split("-")[1].split(".")[0])
steps.append(step)
mmlu_scores.append(data["results"]["mmlu"]["acc"])
# Plot
plt.plot(steps, mmlu_scores)
plt.xlabel("Training Step")
plt.ylabel("MMLU Accuracy")
plt.title("Training Progress")
plt.savefig("training_curve.png")Benchmark suite for model comparison.
Model Comparison:
- [ ] Step 1: Define model list
- [ ] Step 2: Run evaluations
- [ ] Step 3: Generate comparison tableStep 1: Define model list
# models.txt
meta-llama/Llama-2-7b-hf
meta-llama/Llama-2-13b-hf
mistralai/Mistral-7B-v0.1
microsoft/phi-2Step 2: Run evaluations
#!/bin/bash
# eval_all_models.sh
TASKS="mmlu,gsm8k,hellaswag,truthfulqa"
while read model; do
echo "Evaluating $model"
# Extract model name for output file
model_name=$(echo $model | sed 's/\//-/g')
lm_eval --model hf \
--model_args pretrained=$model,dtype=bfloat16 \
--tasks $TASKS \
--num_fewshot 5 \
--batch_size auto \
--output_path results/$model_name.json
done < models.txtStep 3: Generate comparison table
import json
import pandas as pd
models = [
"meta-llama-Llama-2-7b-hf",
"meta-llama-Llama-2-13b-hf",
"mistralai-Mistral-7B-v0.1",
"microsoft-phi-2"
]
tasks = ["mmlu", "gsm8k", "hellaswag", "truthfulqa"]
results = []
for model in models:
with open(f"results/{model}.json") as f:
data = json.load(f)
row = {"Model": model.replace("-", "/")}
for task in tasks:
# Get primary metric for each task
metrics = data["results"][task]
if "acc" in metrics:
row[task.upper()] = f"{metrics['acc']:.3f}"
elif "exact_match" in metrics:
row[task.upper()] = f"{metrics['exact_match']:.3f}"
results.append(row)
df = pd.DataFrame(results)
print(df.to_markdown(index=False))Output:
| Model | MMLU | GSM8K | HELLASWAG | TRUTHFULQA |
|------------------------|-------|-------|-----------|------------|
| meta-llama/Llama-2-7b | 0.459 | 0.142 | 0.765 | 0.391 |
| meta-llama/Llama-2-13b | 0.549 | 0.287 | 0.801 | 0.430 |
| mistralai/Mistral-7B | 0.626 | 0.395 | 0.812 | 0.428 |
| microsoft/phi-2 | 0.560 | 0.613 | 0.682 | 0.447 |Use vLLM backend for 5-10x faster evaluation.
vLLM Evaluation:
- [ ] Step 1: Install vLLM
- [ ] Step 2: Configure vLLM backend
- [ ] Step 3: Run evaluationStep 1: Install vLLM
pip install vllmStep 2: Configure vLLM backend
lm_eval --model vllm \
--model_args pretrained=meta-llama/Llama-2-7b-hf,tensor_parallel_size=1,dtype=auto,gpu_memory_utilization=0.8 \
--tasks mmlu \
--batch_size autoStep 3: Run evaluation
vLLM is 5-10× faster than standard HuggingFace:
# Standard HF: ~2 hours for MMLU on 7B model
lm_eval --model hf \
--model_args pretrained=meta-llama/Llama-2-7b-hf \
--tasks mmlu \
--batch_size 8
# vLLM: ~15-20 minutes for MMLU on 7B model
lm_eval --model vllm \
--model_args pretrained=meta-llama/Llama-2-7b-hf,tensor_parallel_size=2 \
--tasks mmlu \
--batch_size autoUse lm-evaluation-harness when:
Use alternatives instead:
Issue: Evaluation too slow
Use vLLM backend:
lm_eval --model vllm \
--model_args pretrained=model-name,tensor_parallel_size=2Or reduce fewshot examples:
--num_fewshot 0 # Instead of 5Or evaluate subset of MMLU:
--tasks mmlu_stem # Only STEM subjectsIssue: Out of memory
Reduce batch size:
--batch_size 1 # Or --batch_size autoUse quantization:
--model_args pretrained=model-name,load_in_8bit=TrueEnable CPU offloading:
--model_args pretrained=model-name,device_map=auto,offload_folder=offloadIssue: Different results than reported
Check fewshot count:
--num_fewshot 5 # Most papers use 5-shotCheck exact task name:
--tasks mmlu # Not mmlu_direct or mmlu_fewshotVerify model and tokenizer match:
--model_args pretrained=model-name,tokenizer=same-model-nameIssue: HumanEval not executing code
Code-executing tasks (HumanEval, MBPP, etc.) are gated behind an explicit
confirmation flag — you must pass --confirm_run_unsafe_code to run them:
lm_eval --model hf \
--model_args pretrained=model-name \
--tasks humaneval \
--confirm_run_unsafe_code # Required to run tasks that execute generated codeWithout this flag lm-eval refuses to run the task rather than silently skipping code execution.
Benchmark descriptions: See references/benchmark-guide.md for detailed description of all 60+ tasks, what they measure, and interpretation.
Custom tasks: See references/custom-tasks.md for creating domain-specific evaluation tasks.
API evaluation: See references/api-evaluation.md for evaluating OpenAI, Anthropic, and other API models.
Multi-GPU strategies: See references/distributed-eval.md for data parallel and tensor parallel evaluation.
© Luciole-Studio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness of Luciole-Studio/Misaka-Agent.
Open the folder on GitHubat commit 3bcf7a3
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Luciole-Studio/Misaka-Agent, which our catalogue first saw on October 7, 2026.
Evaluating LLMs Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluating LLMs Harness this skillLuciole-Studio/Misaka-Agent | 158 | 2 repos | ~3.1k | Automated safety check: Pass | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Outlines Structured GenerationOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~4k | Automated safety check: Pass | MIT | |
| Aqua Model Lifecycleoracle/accelerated-data-science | 125 | — | ~1.4k | Automated safety check: Pass | UPL-1.0 | |
| Hugging Face Community Evalshenryalouf/ruflow | 157 | — | ~1.6k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Orchestra-Research/AI-Research-SKILLs
Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models.
oracle/accelerated-data-science
Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.
henryalouf/ruflow
Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval.
sickn33/agentic-awesome-skills
Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.
Luciole-Studio/Misaka-Agent
Plan and run multi-agent video production pipelines. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
AST-aware structural code search and rewrite via ast-grep. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
Drug discovery: ChEMBL search, drug-likeness, interactions. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
Workout planning, macros, and body metrics via wger/USDA. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
Render MP4/WebM videos from HTML compositions. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
Follow the money via public records and sanctions data. An agent skill from Luciole-Studio/Misaka-Agent.
Works with
Categories
lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent. Evaluating LLMs Harness is an agent skill from Luciole-Studio/Misaka-Agent.).
Evaluating LLMs Harness fits situations like: tasks that involve LLM evaluation.
Run `npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a claude-code`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness in Luciole-Studio/Misaka-Agent) into .claude/skills/evaluating-llms-harness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a codex`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/evaluation/evaluating-llms-harness in Luciole-Studio/Misaka-Agent) into .agents/skills/evaluating-llms-harness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Luciole-Studio/Misaka-Agent --skill evaluating-llms-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluating-llms-harness, .gemini/skills/evaluating-llms-harness, .github/skills/evaluating-llms-harness and .opencode/skills/evaluating-llms-harness in your project.
Going by SKILL.md and its folder, Evaluating LLMs Harness needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: github.com and huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evaluating LLMs Harness is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Evaluating LLMs Harness: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Outlines Structured Generation (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Aqua Model Lifecycle (oracle/accelerated-data-science, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Luciole-Studio (a GitHub organization) maintains it in Luciole-Studio/Misaka-Agent, which has 158 GitHub stars. The repository holds 77 skills in this directory. The repository was last updated on October 8, 2026.
Source: Luciole-Studio/Misaka-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.