Search

GitHub · LLM evaluation

10 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Research, create, improve, migrate, evaluate, package, install-check, govern, and safely publish qiaomu-flavored agent skills from workflows, prompts, transcripts, docs, SOPs, runbooks, scripts, or…

joeseesun/qiaomu-meta-skill383—~2.8kAutomated safety check: PassMIT2 mo ago
2

Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

bgauryy/octocode949—~2.1kAutomated safety check: PassMIT5 days ago
3

A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

sun461941-hub/ai-project-copilot100—~3kAutomated safety check: PassMIT1 mo ago
4

Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.

frontier-harness-eval/eval297—~8kAutomated safety check: PassNo licence1 mo ago
5

Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

ory/lumen305—~497Automated safety check: PassUnknown1 mo ago
6

Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

bgauryy/octocode949—~1.6kAutomated safety check: PassMIT5 days ago
7

[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).

pskoett/pskoett-ai-skills311—~2.2kAutomated safety check: PassNo licence3 days ago
8

Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.

RConsortium/pharma-skills118—~5.3kAutomated safety check: PassMIT4 days ago
9

INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.

langchain-ai/langchain-skills1.3k—~8.7kAutomated safety check: NotesMIT3 days ago
10

Author and validate Vally evals for Agent Skills under .github/skills.

Azure/azure-sdk-tools134—~946Automated safety check: PassMITtoday