Search
GitHub · LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Research, create, improve, migrate, evaluate, package, install-check, govern, and safely publish qiaomu-flavored agent skills from workflows, prompts, transcripts, docs, SOPs, runbooks, scripts, or… | joeseesun/ | 383 | — | ~2.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 2 | Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report. | bgauryy/ | 949 | — | ~2.1k | Automated safety check: Pass | MIT | 5 days ago |
| 3 | A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient… | sun461941-hub/ | 100 | — | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 4 | Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes. | frontier-harness-eval/ | 297 | — | ~8k | Automated safety check: Pass | No licence | 1 mo ago |
| 5 | Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch. | ory/ | 305 | — | ~497 | Automated safety check: Pass | Unknown | 1 mo ago |
| 6 | Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification. | bgauryy/ | 949 | — | ~1.6k | Automated safety check: Pass | MIT | 5 days ago |
| 7 | [Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). | pskoett/ | 311 | — | ~2.2k | Automated safety check: Pass | No licence | 3 days ago |
| 8 | Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs. | RConsortium/ | 118 | — | ~5.3k | Automated safety check: Pass | MIT | 4 days ago |
| 9 | INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith. | langchain-ai/ | 1.3k | — | ~8.7k | Automated safety check: Notes | MIT | 3 days ago |
| 10 | Author and validate Vally evals for Agent Skills under .github/skills. | Azure/ | 134 | — | ~946 | Automated safety check: Pass | MIT | today |