Search
arXiv · LLM evaluation
2 skills found.
Category:
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100… | open-thoughts/ | 301 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 2 | 2.Watch Monitor arXiv for new papers relevant to the GLIDE project (prediction-powered inference, active statistical inference, LLM evaluation debiasing, proxy annotation bias correction). | EmertonData/ | 119 | — | ~3.5k | Automated safety check: Pass | Unknown | 2 days ago |