Search
Docker · LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Converts test suites from external eval frameworks into the Margin Eval suite format. | Margin-Lab/ | 160 | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | 2 mo ago |
| 2 | Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes. | benchflow-ai/ | 356 | — | ~4k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 3 | Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics. | Orchestra-Research/ | 13k | 4 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 4 | Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. | Orchestra-Research/ | 13k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 5 | Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 6 | Teaches how to write and run evals on the products/posthogai/evalharness/ harness — sandboxed agent suites that execute the real coding agent in a Docker or Modal sandbox against a seeded Hedgebox… | PostHog/ | 40k | — | ~4k | Automated safety check: Notes | Unknown | yesterday |