Search
Bash · LLM evaluation
3 skills found.
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. | Devin-AXIS/ | 6.8k | — | ~851 | Automated safety check: Pass | Unknown | yesterday |
| 2 | Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. | vercel/ | 301 | — | ~3.6k | Automated safety check: Pass | Unknown | yesterday |
| 3 | Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit… | datadog-labs/ | 177 | — | ~25k | Automated safety check: Pass | MIT | 2 days ago |