Search
NVIDIA AI Platform · LLM evaluation
4 skills found.
Category:
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. | Orchestra-Research/ | 13k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 3 | A skill your agent uses when integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own… | NVIDIA/ | 3.6k | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 4 | Filesystem RAG benchmarks: corpus/, train.json, evaluaterag.py (RAGAS quality). | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |