Search
Development · Agent evaluation and testing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches. | QwenLM/ | 28k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Run and maintain Con's terminal-agent benchmark against a live app session. | nowledge-co/ | 626 | — | ~822 | Automated safety check: Pass | MIT | today |
| 3 | A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps… | microsoft/ | 138 | — | ~5.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 4 | Points agents to the right sources when changing the Smithers workspace graph, generated docs, benchmark and eval evidence, or durable flows in the Smithers repository. | smithersai/ | 431 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 5 | Analyzes a finished pull request that involved an agent to find what slowed or helped it, and recommends specific changes to instruction files, skills and docs. | dotnet/ | 23k | — | ~2.5k | Automated safety check: Pass | MIT | yesterday |
| 6 | A skill your agent uses when my-pi-agent tests, typecheck, lint, Pi CLI startup, OpenCode/MCP, Feishu channel, prompt rendering, memory/state, web-console/desktop build, or UI behavior fails… | skuramatata/ | 114 | — | ~777 | Automated safety check: Pass | No licence | 3 mo ago |
| 7 | A skill your agent uses when my-pi-agent tests, typecheck, lint, /code workflow, verifier probes, Feishu channel handling, MCP bootstrap, skill install, or task resume behavior fails unexpectedly. | skuramatata/ | 114 | — | ~829 | Automated safety check: Pass | No licence | 3 mo ago |
| 8 | 法律 Skill 分层质量评测工具。消费 skill-lint 的通用质量结论,再用三份测试材料、通用六维度、场景微调和律师 taste 评估法律产出,并定位最小修复单元。本技能应在审查、回归验证或发布验收法律 Skill 时使用。不要用于代替通用 Skill lint、正式法律意见或跨场景排名。 | cat-xierluo/ | 720 | — | ~2.2k | Automated safety check: Pass | CC-BY-NC-4.0 | yesterday |