Topic · Education
Best quizzes and assessments skills, page 3
Quizzes and assessments skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Review rubric for the /review pull-request command. An agent skill from NVIDIA/Megatron-LM. | NVIDIA/ | 18k | — | ~839 | Automated safety check: Pass | Apache-2.0 | today |
| 98 | Write, extend, and debug PXI Playwright E2E tests for Phoenix. | Arize-ai/ | 12k | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 99 | Make a correct update to Parker's prompts, system docs, rubrics, knowledge docs, training corpus, or brand outputs — and propagate the change everywhere it needs to land. | real-simple-labs/ | 102 | — | ~4.6k | Automated safety check: Pass | Unknown | today |
| 100 | A skill your agent uses when creating, asking, grading, or explaining multiple choice questions for MATLAB programming practice, concept checks, quizzes, or tutoring exercises. | matlab/ | 184 | — | ~869 | Automated safety check: Pass | Unknown | today |
| 101 | Writes low-stakes retrieval-practice questions with answer notes from material you supply, mixing free recall, cued recall, recognition and application items. | iflytek/ | 5.2k | — | ~1.1k | Automated safety check: Pass | CC-BY-SA-4.0 | today |
| 102 | Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository. | dotnet/ | 5.6k | 1 repo | ~6.1k | Automated safety check: Pass | MIT | today |
| 103 | Patterns and techniques for evaluating and improving AI agent outputs. | github/ | 40k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 104 | 104.Eval Judge Decide whether ONE answer matches its golden, and say whether you believe the golden. | malloydata/ | 116 | — | ~3.4k | Automated safety check: Pass | MIT | today |
| 105 | 105.Canvas Humanizer A skill your agent uses when a local academic draft needs a meaning-preserving humanizing pass with less uniform syntax while retaining rubric, source, lock, voice, and length constraints. | X-isdoingreat/ | 125 | — | ~1.8k | Automated safety check: Pass | AGPL-3.0 | 2 mo ago |
| 106 | Reference vocabulary for interpreting vulnerability findings — detector-vs-impact distinction, severity anchoring on demonstrated evidence, the eleven-item interpretation rubric, delegation… | provos/ | 612 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 107 | 107.Skill Score Grade a Dex skill against the shape-aware quality rubric and report a ship/revise/no verdict with the exact fixes. | davekilleen/ | 494 | — | ~2.6k | Automated safety check: Pass | MIT | today |
| 108 | 108.01 Auto Arena Automatically evaluate and compare multiple AI models or agents without pre-existing test data. | agentscope-ai/ | 871 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 29 days ago |
| 109 | 109.Scoring 100-point NL artifact rubric: penalty tables per artifact type, calibration cases. | xiaolai/ | 150 | — | ~5.3k | Automated safety check: Pass | ISC | today |
| 110 | Score a scoped competitor set into comparable profile cards: nine weighted dimensions (positioning, voice, visual craft, offer packaging, evidence, enterprise-readiness, thought leadership, pricing… | affaan-m/ | 276k | 1 repo | ~2.5k | Automated safety check: Pass | MIT | today |
| 111 | Structured scholarly-work evaluation for papers, proposals, literature reviews, methods sections, evidence quality, citation support, and research-writing feedback. | affaan-m/ | 276k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 112 | 112.Eval Loop This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target. | jacob-dietle/ | 111 | — | ~5.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 113 | 113.Skill Test Validate skill files for structural compliance and behavioral correctness. | Donchitos/ | 26k | — | ~6k | Automated safety check: Pass | MIT | 2 days ago |
| 114 | 114.Evaluators Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output. | Arize-ai/ | 12k | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 115 | 115.Experiments Run, read, and compare dataset-backed experiments to find evidence that a prompt or pipeline is improving. | Arize-ai/ | 12k | — | ~1.8k | Automated safety check: Pass | Unknown | today |
| 116 | Bulk grading workflows for Canvas LMS assignments using rubrics. | vishalsachdev/ | 286 | — | ~2k | Automated safety check: Pass | MIT | today |
| 117 | 117.Cheat Status cheat-on-content 的状态看板。显示当前模式 / rubric 版本 / 校准进度 / 待复盘 / pool 状态 / 是否该升级 SQLite / 是否该 bump rubric。任何时候都可调,无副作用。触发词:"状态"/"看板"/"status"/"我现在该做什么"/"进度怎么样"。 | XBuilderLAB/ | 7.2k | — | ~1.5k | Automated safety check: Notes | MIT | 5 days ago |
| 118 | Place a goal-driven CALL-E call that collects specific structured answers, score those answers against a deterministic rubric you supply, and conditionally trigger a follow-up action — all runnable… | CALLE-AI/ | 107 | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 119 | 119.Vibes Brainstorm Lightweight requirements gathering before app generation. An agent skill from popmechanic/VibesOS. | popmechanic/ | 136 | — | ~2.2k | Automated safety check: Pass | MIT | 2 mo ago |
| 120 | 120.Lockedin Captures the user's work moments — a shipped feature, a meeting outcome, a learning, a decision — from inside their Claude Code session into structured local markdown. | daypunk/ | 128 | — | ~3.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 121 | Asks the user mid-run from inside an FSM action or Python type: user agent via the elicitation tools — requesttoolaccess for allowlist approvals and requesthumaninput for HITL (human-in-the-loop)… | friday-platform/ | 104 | — | ~2.4k | Automated safety check: Pass | Unknown | 1 mo ago |
| 122 | Interactive training for the GitHub Copilot CLI. An agent skill from github/awesome-copilot. | github/ | 40k | 2 repos | ~517 | Automated safety check: Pass | MIT | yesterday |
| 123 | A skill your agent uses when a local draft needs repeated canvas-humanizer passes with independent meaning, structure, citation, voice, and rubric-damage checks. | X-isdoingreat/ | 125 | — | ~1.7k | Automated safety check: Pass | AGPL-3.0 | 2 mo ago |
| 124 | 124.Hypothesis Gen A skill your agent uses when the user wants to generate and literature-vet a pool of novel, testable research hypotheses for a question or domain. | gaasher/ | 174 | — | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 125 | 维护者专用的短剧 know-how 学习、验证与生命周期治理。仅在维护者明确要求从其当前会话提供的授权只读文本源学习完整短剧项目链,并把私有观察逐步转成去标识、去复刻、经盲测与独立审查的公共 reference、rubric 或 synthetic fixture 候选时使用;不用于普通创作、公开运行时取数、媒体生成或粗略数据分析。 | zenstory-ai/ | 2.7k | — | ~914 | Automated safety check: Pass | MIT | 7 days ago |
| 126 | 126.Dhdna Profiler Applies the DHDNA framework as an exploratory rubric for reasoning and writing patterns in supplied text. | K-Dense-AI/ | 48k | 1 repo | ~2.8k | Automated safety check: Pass | MIT | 5 days ago |
| 127 | Decision rubric for when an LM agent should write-and-run code (Program-of-Thought / code interpreter) versus reason in natural language: classify each step as deterministic- computable (emit +… | agentsope/ | 436 | — | ~6.1k | Automated safety check: Pass | MIT | yesterday |
| 128 | 128.Agent Evaluation Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis. | seb1n/ | 206 | — | ~1.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 129 | 129.Study Quiz Use ao pedir quiz, teste, prova, simulado, questões de múltipla escolha, mini-desafio, gabarito, “me testa sobre X”, “gera um quiz de X” ou uma avaliação sobre um assunto, em um nível ou numa faixa… | kipperacademy/ | 149 | — | ~3.5k | Automated safety check: Pass | No licence | today |
| 130 | Investigate a topic against preserved sources and write a draft-status research article under research/ in a Knowledge Base project (the knowledge-base starter pack). | inkeep/ | 4.5k | — | ~5.6k | Automated safety check: Pass | GPL-3.0 | today |
| 131 | Expert in building shareable generator tools that go viral - name generators, quiz makers, avatar creators, personality tests, and calculator tools. | sickn33/ | 47k | 2 repos | ~2k | Automated safety check: Pass | MIT | yesterday |
| 132 | 132.Skill Doctor A skill your agent uses when the user wants their agent setup graded from real conversation history, asks which installed skills are actually working, or wants evidence-backed skill edits — scores… | alirezarezvani/ | 28k | — | ~1.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 133 | PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 5 days ago |
| 134 | 134.Exam Quiz 从 references/quizbank.json 抽取本章题目并按标准答案判分,支持选择、主观、画图、填空、判断、代码; 主观题按 keywords 要点覆盖判分,连续错两次提供提示/跳过/归档。禁止现场编题。用于阶段检查或模考。 | ZeKaiNie/ | 303 | — | ~1.9k | Automated safety check: Pass | MIT | 12 days ago |
| 135 | This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise… | aiskillstore/ | 433 | 3 repos | ~4.2k | Automated safety check: Pass | No licence | today |
| 136 | 136.Mock Exam This skill should be used when the user asks to "模拟考试", "出模拟卷", "做一套模拟题", "帮我批改", "批改试卷", "给我出题考试", "期末模拟", "模拟测试", "考考我", "出一套试卷", "模拟一下", "帮我评分", "我答完了帮我批", "出一套期末卷", "按考试标准出题", or when the user… | mingchen666/ | 244 | — | ~1.6k | Automated safety check: Notes | No licence | today |
| 137 | 137.Aer Referee Sim A skill your agent uses when a complete draft exists and needs an adversarial internal review before submission — simulating the AER desk screen and three referee reports with calibrated severity… | brycewang-stanford/ | 4.6k | 1 repo | ~2.7k | Automated safety check: Pass | Unknown | 4 days ago |
| 138 | Copilot left 14 review comments on your PR — half are nits. An agent skill from github/awesome-copilot. | github/ | 40k | — | ~3.4k | Automated safety check: Pass | MIT | yesterday |
| 139 | Draft and revise academic prose against a rubric, evidence, and citation requirements. | first-fluke/ | 1.3k | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 140 | 140.Comment Judge LLM-as-a-judge rubric for code comments (forbidden, false, stale, narration, noise, keep). | fmflurry/ | 171 | — | ~2.5k | Automated safety check: Pass | MIT | 3 days ago |
| 141 | A skill your agent uses when a user wants to build, launch, grade, or schedule a Claude Managed Agent (CMA) in their own Anthropic account — "build me an agent", "launch this as a managed agent"… | alirezarezvani/ | 28k | — | ~1.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 142 | 142.Grade Iterate Phase 3 of building a Claude Managed Agent — the bounded grade→iterate loop. | alirezarezvani/ | 28k | — | ~1k | Automated safety check: Pass | MIT | 1 mo ago |
| 143 | 143.Metric Design A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining… | agentscope-ai/ | 871 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 29 days ago |
| 144 | 144.Cookbook Audit Audit an Anthropic Cookbook notebook based on a rubric. An agent skill from Microck/ordinary-claude-skills. | Microck/ | 404 | 1 repo | ~3.1k | Automated safety check: Pass | Unknown | 1 mo ago |