Search

Docker · LLM evaluation

6 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Converts test suites from external eval frameworks into the Margin Eval suite format.

Margin-Lab/evals160—~3.7kAutomated safety check: PassAGPL-3.02 mo ago
2

Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.

benchflow-ai/benchflow356—~4kAutomated safety check: PassApache-2.05 days ago
3

Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.9kAutomated safety check: PassMIT3 mo ago
4

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.

Orchestra-Research/AI-Research-SKILLs13k2 repos~3.1kAutomated safety check: PassMIT3 mo ago
5

Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report.

NVIDIA/skills3.6k—~1.3kAutomated safety check: NotesApache-2.0yesterday
6
6.Writing EvalsOfficial

Teaches how to write and run evals on the products/posthogai/evalharness/ harness — sandboxed agent suites that execute the real coding agent in a Docker or Modal sandbox against a seeded Hedgebox…

PostHog/posthog40k—~4kAutomated safety check: NotesUnknownyesterday