Topic · AI & LLM Engineering

Best LLM evaluation skills, page 7

Skills #289–308 of 308, ranked by score.

LLM evaluation skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM evaluation skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
289

Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against…

AnastasiyaW/codex-claude-code-config154—~764Automated safety check: PassMITtoday
290

Evaluate and benchmark large language models for research applications

wentorai/research-plugins2981 repo~1.7kAutomated safety check: PassMIT3 mo ago
291

The first Rightbrain CLI that reaches past tasks — agents, approvals, evals, triggers and audit, plus a local mirror that makes credit spend and latency regressions queryable.

mvanhorn/printing-press-library2.1k—~9kAutomated safety check: NotesApache-2.0today
292

Make an AI agent or automation reliable enough to trust — the tests, checks, and guardrails that catch its failures before they reach anything real.

mohitagw15856/pm-claude-skills1.4k—~1.3kAutomated safety check: PassMITtoday
293

Design an evaluation plan for an LLM or AI feature before shipping it.

mohitagw15856/pm-claude-skills1.4k—~996Automated safety check: PassMITtoday
294

Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output.

mohitagw15856/pm-claude-skills1.4k—~1.1kAutomated safety check: PassMITtoday
295
295.Eval

Quality and performance evaluation with baseline comparison.

hashgraph-online/awesome-codex-plugins1.2k—~2kAutomated safety check: PassApache-2.0today
296

Implement a task step by step with automated LLM-as-Judge verification at the end of each phase

NeoLabHQ/context-engineering-kit1.7k—~20kAutomated safety check: PassGPL-3.01 mo ago
297

Design and apply QA methodology for software teams: test strategy, regression testing, CI failure triage, test automation, quality gates and metrics, risk-based testing, exploratory testing, test…

magnus919/agent-skills113—~3.2kAutomated safety check: PassMITyesterday
298

Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring

Abilityai/cornelius109—~2.4kAutomated safety check: NotesMIT15 days ago
299

Review evaluation history, distinguish system, data, suite, and evaluator changes, and apply an operating response.

ai-analyst-lab/ai-analyst304—~306Automated safety check: PassMIT7 days ago
300

Creates assessments with varied question types (MCQ, code-completion, debugging, projects) aligned to learning objectives with meaningful distractors based on common misconceptions.

aiskillstore/marketplace430—~4.3kAutomated safety check: PassNo licencetoday
301

Standardize and validate SKILL.md files against the Agent Skills specification (agentskills.io).

aiskillstore/marketplace430—~2.1kAutomated safety check: NotesNo licencetoday
302

A skill your agent uses when hardening a COLM paper's reproducibility story — pinning open-weight checkpoints and tokenizers, handling API-model drift and deprecation honestly, versioning evaluation…

brycewang-stanford/Awesome-Journal-Skills1.2k—~1.7kAutomated safety check: PassMIT10 days ago
303

A skill your agent uses when authoring a NEW rsc skill or editing an existing one — scoping it to one job, writing the description that decides whether it ever loads, splitting the body into…

ericrisco/rsc-harness167—~4.3kAutomated safety check: PassMITtoday
304

Reduces AI-detection signals in drafted text by routing every non-locked sentence through round-trip translation (English → intermediate language → English) and selecting the candidate that…

X-isdoingreat/canvas-pilot125—~10kAutomated safety check: NotesAGPL-3.02 mo ago
305

Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval.

majiayu000/claude-skill-registry6661 repo~1.7kAutomated safety check: PassMITtoday
306

A skill your agent uses when creating, reviewing, or editing Agent Skills-format skills, or when implementing skill discovery and loading in an agent client.

magnus919/agent-skills113—~3kAutomated safety check: PassMITyesterday
307

Build type-safe AI agents and graph-based workflows with PydanticAI and PydanticGraph.

magnus919/agent-skills113—~4kAutomated safety check: PassMITyesterday
308

Routes any legal task to the right LLM, like OpenRouter but for legal work and grounded in benchmarks instead of brand loyalty.

lawve-ai/awesome-legal-skills836—~1.5kAutomated safety check: PassAGPL-3.0-or-later5 days ago