Search

vLLM · LLM evaluation

8 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

Orchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT3 mo ago
2

Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

huggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.03 days ago
3

Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.

Leeroo-AI/superml195—~3.8kAutomated safety check: PassApache-2.06 mo ago
4

Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval.

henryalouf/ruflow157—~1.6kAutomated safety check: PassMIT4 mo ago
5

Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

sickn33/agentic-awesome-skills47k1 repo~1.9kAutomated safety check: PassApache-2.0yesterday
6

Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal.

sickn33/agentic-awesome-skills47k1 repo~1.7kAutomated safety check: PassMITyesterday
7

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1712 repos~3.1kAutomated safety check: PassMIT3 days ago
8

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

open-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.012 days ago