Search

Development · LLM evaluation

20 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Provides context about the CoStrict evals system structure in this monorepo.

zgsm-ai/costrict4.5k1 repo~1.9kAutomated safety check: PassApache-2.010 days ago
2

Research, create, improve, migrate, evaluate, package, install-check, govern, and safely publish qiaomu-flavored agent skills from workflows, prompts, transcripts, docs, SOPs, runbooks, scripts, or…

joeseesun/qiaomu-meta-skill383—~2.8kAutomated safety check: PassMIT2 mo ago
3

A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

sun461941-hub/ai-project-copilot97—~3kAutomated safety check: PassMIT1 mo ago
4

Plan a documentation eval for supabase/evals, where a docs guide is the subject under test.

supabase/evals144—~2.9kAutomated safety check: PassApache-2.0today
5

End-to-end GECX/CXAS/CES conversational agent lifecycle -- build agents from requirements (PRD-to-agent), create and run evals (goldens, simulations, tool tests, callback tests), debug failures, and…

GoogleCloudPlatform/cxas-scrapi107—~2.4kAutomated safety check: PassApache-2.0yesterday
6

Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…

evaleval/every_eval_ever135—~2.5kAutomated safety check: PassMIT2 days ago
7

Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

microsoft/skills3.1k—~2.8kAutomated safety check: PassMITtoday
8

This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs…

pifferologo/cloud-agents-cli1291 repo~6.8kAutomated safety check: PassApache-2.01 mo ago
9

Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.

wso2/agent-manager108—~710Automated safety check: PassApache-2.0yesterday
10

Run a targeted local React Doctor Evals loop against an uncommitted rule change.

millionco/react-doctor15k—~510Automated safety check: PassUnknowntoday
11

Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository.

dotnet/skills5.6k1 repo~6.1kAutomated safety check: PassMITtoday
12
12.EveOfficial

eve framework guidance for durable AI agents and agent-powered applications.

vercel/vercel-plugin3015 repos~1.2kAutomated safety check: PassUnknowntoday
13

Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.

leo-kuang-ai/spec-first107—~13kAutomated safety check: PassMITyesterday
14

Multi-agent collaboration plugin that spawns N parallel subagents competing on the same task via git worktree isolation.

alirezarezvani/claude-skills28k—~2kAutomated safety check: PassMIT1 mo ago
15

Author or review software tests and LLM/agent evals; choose test placement and mocks, diagnose flaky CI, or repair brittle suites.

pedronauck/skills6341 repo~636Automated safety check: PassNo licence25 days ago
16

Build or refresh a product README showcase using a seeded Bag of Words workspace, polished in-product screenshots, and repository-ready visual assets.

bagofwords1/bagofwords459—~1.8kAutomated safety check: PassUnknowntoday
17

A skill your agent uses when designing, auditing, refactoring, or explaining an agentic harness for any domain, especially when work must continue from a measured gap to verified completion.

AnastasiyaW/codex-claude-code-config154—~5.4kAutomated safety check: PassMITtoday
18

A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and…

tikalk/adlc-team-skills141—~1.5kAutomated safety check: PassMITtoday
19

Version tracking for Agent Skills bundles and their associated files across sessions, surfaces, and platforms.

LeoYeAI/openclaw-master-skills2.2k—~4.8kAutomated safety check: PassMIT2 mo ago
20

Design and apply QA methodology for software teams: test strategy, regression testing, CI failure triage, test automation, quality gates and metrics, risk-based testing, exploratory testing, test…

magnus919/agent-skills115—~3.2kAutomated safety check: PassMITtoday