Search
GitHub · Agent evaluation and testing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers. | obra/ | 296k | 3 repos | ~1.7k | Automated safety check: Pass | MIT | today |
| 2 | Designs and verifies a deterministic grader that measures whether a GitHub Agentic Workflow run reached its real-world or repository outcome. | github/ | 5.4k | — | ~6.8k | Automated safety check: Pass | MIT | today |
| 3 | Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report. | bgauryy/ | 949 | — | ~2.1k | Automated safety check: Pass | MIT | 4 days ago |
| 4 | Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch. | ory/ | 305 | — | ~497 | Automated safety check: Pass | Unknown | 1 mo ago |
| 5 | A skill your agent uses when changing Deep Researcher Agent continuous integration, pre-commit, or contributor governance — editing .github/workflows/ (ci, ui, skills-eval, request-nvskills-ci)… | NVIDIA-AI-Blueprints/ | 883 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 6 | Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification. | bgauryy/ | 949 | — | ~1.6k | Automated safety check: Pass | MIT | 4 days ago |
| 7 | Analyzes a finished pull request that involved an agent to find what slowed or helped it, and recommends specific changes to instruction files, skills and docs. | dotnet/ | 23k | — | ~2.5k | Automated safety check: Pass | MIT | today |
| 8 | Investigates what went wrong in a superpowers session by reading its transcript, reports findings with path and line citations, and can draft a GitHub issue or redacted bundle. | jnMetaCode/ | 8.3k | — | ~858 | Automated safety check: Pass | MIT | 3 days ago |