Topic · AI & LLM Engineering
Best LLM evaluation skills, page 7
LLM evaluation skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 289 | Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against… | AnastasiyaW/ | 154 | — | ~764 | Automated safety check: Pass | MIT | today |
| 290 | Evaluate and benchmark large language models for research applications | wentorai/ | 298 | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 291 | 291.Pp Rightbrain The first Rightbrain CLI that reaches past tasks — agents, approvals, evals, triggers and audit, plus a local mirror that makes credit spend and latency regressions queryable. | mvanhorn/ | 2.1k | — | ~9k | Automated safety check: Notes | Apache-2.0 | today |
| 292 | Make an AI agent or automation reliable enough to trust — the tests, checks, and guardrails that catch its failures before they reach anything real. | mohitagw15856/ | 1.4k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 293 | 293.AI Eval Plan Design an evaluation plan for an LLM or AI feature before shipping it. | mohitagw15856/ | 1.4k | — | ~996 | Automated safety check: Pass | MIT | today |
| 294 | Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output. | mohitagw15856/ | 1.4k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 295 | 295.Eval Quality and performance evaluation with baseline comparison. | hashgraph-online/ | 1.2k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 296 | 296.Implement Task Implement a task step by step with automated LLM-as-Judge verification at the end of each phase | NeoLabHQ/ | 1.7k | — | ~20k | Automated safety check: Pass | GPL-3.0 | 1 mo ago |
| 297 | 297.QA Methodology Design and apply QA methodology for software teams: test strategy, regression testing, CI failure triage, test automation, quality gates and metrics, risk-based testing, exploratory testing, test… | magnus919/ | 113 | — | ~3.2k | Automated safety check: Pass | MIT | yesterday |
| 298 | 298.Benchmark Memory Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring | Abilityai/ | 109 | — | ~2.4k | Automated safety check: Notes | MIT | 15 days ago |
| 299 | 299.Monitor Evals Review evaluation history, distinguish system, data, suite, and evaluator changes, and apply an operating response. | ai-analyst-lab/ | 304 | — | ~306 | Automated safety check: Pass | MIT | 7 days ago |
| 300 | Creates assessments with varied question types (MCQ, code-completion, debugging, projects) aligned to learning objectives with meaningful distractors based on common misconceptions. | aiskillstore/ | 430 | — | ~4.3k | Automated safety check: Pass | No licence | today |
| 301 | Standardize and validate SKILL.md files against the Agent Skills specification (agentskills.io). | aiskillstore/ | 430 | — | ~2.1k | Automated safety check: Notes | No licence | today |
| 302 | A skill your agent uses when hardening a COLM paper's reproducibility story — pinning open-weight checkpoints and tokenizers, handling API-model drift and deprecation honestly, versioning evaluation… | brycewang-stanford/ | 1.2k | — | ~1.7k | Automated safety check: Pass | MIT | 10 days ago |
| 303 | 303.Author Skill A skill your agent uses when authoring a NEW rsc skill or editing an existing one — scoping it to one job, writing the description that decides whether it ever loads, splitting the body into… | ericrisco/ | 167 | — | ~4.3k | Automated safety check: Pass | MIT | today |
| 304 | 304.Canvas Humanizer Reduces AI-detection signals in drafted text by routing every non-locked sentence through round-trip translation (English → intermediate language → English) and selecting the candidate that… | X-isdoingreat/ | 125 | — | ~10k | Automated safety check: Notes | AGPL-3.0 | 2 mo ago |
| 305 | Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval. | majiayu000/ | 666 | 1 repo | ~1.7k | Automated safety check: Pass | MIT | today |
| 306 | 306.Agent Skills A skill your agent uses when creating, reviewing, or editing Agent Skills-format skills, or when implementing skill discovery and loading in an agent client. | magnus919/ | 113 | — | ~3k | Automated safety check: Pass | MIT | yesterday |
| 307 | 307.Pydanticai Build type-safe AI agents and graph-based workflows with PydanticAI and PydanticGraph. | magnus919/ | 113 | — | ~4k | Automated safety check: Pass | MIT | yesterday |
| 308 | Routes any legal task to the right LLM, like OpenRouter but for legal work and grounded in benchmarks instead of brand loyalty. | lawve-ai/ | 836 | — | ~1.5k | Automated safety check: Pass | AGPL-3.0-or-later | 5 days ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23