Search
Development · LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Provides context about the CoStrict evals system structure in this monorepo. | zgsm-ai/ | 4.5k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 2 | Research, create, improve, migrate, evaluate, package, install-check, govern, and safely publish qiaomu-flavored agent skills from workflows, prompts, transcripts, docs, SOPs, runbooks, scripts, or… | joeseesun/ | 383 | — | ~2.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 3 | A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient… | sun461941-hub/ | 97 | — | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 4 | Plan a documentation eval for supabase/evals, where a docs guide is the subject under test. | supabase/ | 144 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | End-to-end GECX/CXAS/CES conversational agent lifecycle -- build agents from requirements (PRD-to-agent), create and run evals (goldens, simulations, tool tests, callback tests), debug failures, and… | GoogleCloudPlatform/ | 107 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 6 | Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars… | evaleval/ | 135 | — | ~2.5k | Automated safety check: Pass | MIT | 2 days ago |
| 7 | Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes. | microsoft/ | 3.1k | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 8 | This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs… | pifferologo/ | 129 | 1 repo | ~6.8k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 9 | Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager. | wso2/ | 108 | — | ~710 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | 10.Rde Eval Run a targeted local React Doctor Evals loop against an uncommitted rule change. | millionco/ | 15k | — | ~510 | Automated safety check: Pass | Unknown | today |
| 11 | Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository. | dotnet/ | 5.6k | 1 repo | ~6.1k | Automated safety check: Pass | MIT | today |
| 12 | eve framework guidance for durable AI agents and agent-powered applications. | vercel/ | 301 | 5 repos | ~1.2k | Automated safety check: Pass | Unknown | today |
| 13 | Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first. | leo-kuang-ai/ | 107 | — | ~13k | Automated safety check: Pass | MIT | yesterday |
| 14 | 14.Agenthub Multi-agent collaboration plugin that spawns N parallel subagents competing on the same task via git worktree isolation. | alirezarezvani/ | 28k | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 15 | 15.Testing Boss Author or review software tests and LLM/agent evals; choose test placement and mocks, diagnose flaky CI, or repair brittle suites. | pedronauck/ | 634 | 1 repo | ~636 | Automated safety check: Pass | No licence | 25 days ago |
| 16 | Build or refresh a product README showcase using a seeded Bag of Words workspace, polished in-product screenshots, and repository-ready visual assets. | bagofwords1/ | 459 | — | ~1.8k | Automated safety check: Pass | Unknown | today |
| 17 | A skill your agent uses when designing, auditing, refactoring, or explaining an agentic harness for any domain, especially when work must continue from a measured gap to verified completion. | AnastasiyaW/ | 154 | — | ~5.4k | Automated safety check: Pass | MIT | today |
| 18 | A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and… | tikalk/ | 141 | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 19 | Version tracking for Agent Skills bundles and their associated files across sessions, surfaces, and platforms. | LeoYeAI/ | 2.2k | — | ~4.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 20 | Design and apply QA methodology for software teams: test strategy, regression testing, CI failure triage, test automation, quality gates and metrics, risk-based testing, exploratory testing, test… | magnus919/ | 115 | — | ~3.2k | Automated safety check: Pass | MIT | today |