Agent skill

Assessment Design

by cbrock84 in cbrock84/headcount

Builds questions and assessments that measure what they claim — matching item type to the thing being assessed, writing distractors that diagnose, and telling a question that measures learning from…

MITAuto-check passedEducation

Install Assessment Design

skills CLI
$ npx skills add cbrock84/headcount --skill assessment-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cbrock84/headcount assessment-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .claude/skills && cp -r skills-src/verticals/education/skills/education/assessment-design .claude/skills/assessment-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
assessment-design
GitHub stars
2k
Token cost
~1.3k tokens
SKILL.md length
738 words
Files
2 (incl. references)
Skills in repo
178
Repo updated
First seen
Licence
MIT

At a glance

Builds questions and assessments that measure what they claim — matching item type to the thing being assessed, writing distractors that diagnose, and telling a question that measures learning from…

  • Education work in your project
  • SKILL.md covers Name the inference before…, Match the item type to the claim, Distractors are the diagnostic and One item is not evidence, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Assessment Design is an agent skill from cbrock84/headcount. Builds questions and assessments that measure what they claim — matching item type to the thing being assessed, writing distractors that diagnose, and telling a question that measures learning from one that measures reading speed or test-taking. Use this to write or review assessment items, design a unit check or diagnostic, work out why scores do not match classroom performance, or decide how much evidence a claim about mastery actually needs.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/sources.md`).

It sits in Education. The repository describes itself as: An agent organization structured as a company — 15+ departments, 125+ skills, each independently installable, citing the standards and regulators that settle the question. Runs… The licence is MIT.

When your agent uses it

  • Education work in your project

Example prompts

  • “Use the assessment-design skill to build questions and assessments that measure what they claim — matching item type to the thing being assessed…”
  • “/assessment-design”

What it can do on your machine

Read from SKILL.md and the folder at commit 98d1c17. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Assessment Design loads about 1.3k tokens when it runs, and up to ~1.9k if it reads all its reference files. Until then it costs about 117 tokens; SKILL.md has 738 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cbrock84/headcount at commit 98d1c17, republished under its MIT licence (© cbrock84). 738 words, ~1,291 tokens.

Download SKILL.mdSave it as .claude/skills/assessment-design/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
assessment-design
description
Builds questions and assessments that measure what they claim — matching item type to the thing being assessed, writing distractors that diagnose, and telling a question that measures learning from one that measures reading speed or test-taking. Use this to write or review assessment items, design a unit check or diagnostic, work out why scores do not match classroom performance, or decide how much evidence a claim about mastery actually needs.

Assessment design

Every question measures something. The work is making sure it measures the thing you meant, because the alternatives — reading speed, familiarity with the format, willingness to guess — are always available and are usually easier for the student to use.

Name the inference before writing the item

An assessment is an argument: the student did this, therefore they can do that. The argument is where assessments fail, and it fails silently.

Write down the claim first — "can decompose a two-digit number into tens and ones" — then ask what performance would be evidence for it, and what performance would be evidence against. An item that a student who lacks the skill can still get right is not evidence. An item that a student who has the skill can still get wrong, for reasons unrelated to it, is worse: it produces a false negative that gets acted on.

The most common unrelated reason is reading. Any item whose stem is harder to read than the skill is to perform has quietly become a reading assessment.

Match the item type to the claim

  • Selected response (multiple choice, matching, true/false) is efficient and can only ever provide evidence of recognition. A student who can recognize the correct answer cannot be assumed to produce it.
  • Constructed response shows the path, which is what makes partial understanding visible. It costs scoring time and needs a rubric written before the responses arrive, not after.
  • Performance tasks are the only honest evidence for anything described as applying, investigating or designing — most NGSS performance expectations and every C3 inquiry, for instance, cannot be assessed by selected response at all.

Mismatch is the usual failure: a standard describing explanation assessed by a multiple-choice item, which measures whether the student can pick an explanation someone else wrote.

Distractors are the diagnostic

In a well-built multiple-choice item, each wrong answer is the result of a specific, predictable error. Then the pattern of wrong answers says what to reteach, and the item earns its place.

Distractors that are merely wrong — a random number, an obviously absurd option — turn a four-option item into a two-option one and tell you nothing beyond right or wrong.

The mechanical tells of a weak item, all of which students learn to exploit long before they learn the content:

  • The longest or most qualified option is correct.
  • One option is grammatically inconsistent with the stem.
  • "All of the above" appears, and is usually correct.
  • Two options are synonyms, so neither can be right.
  • The correct answer repeats wording from the stem.
Show full SKILL.md (315 more words)Show less

One item is not evidence

A single item carries noise — a misread word, a slip, a lucky guess — that swamps the signal for any individual student. Inferring mastery from one response is the most common measurement error in classroom material, and the most consequential, because it gets recorded.

Several items per claim, varied in surface form so that recognition of the format is not what is being measured. If a claim is worth recording against a student's name, it is worth three items.

Spacing matters as much as quantity: performance on a skill immediately after it is taught measures something closer to short-term recall than to learning. The same item a fortnight later measures more.

Formative and summative are different products

Formative assessment exists to change what happens next, which means it must be quick, frequent, low-stakes, and read immediately. An assessment that takes a week to score cannot be formative whatever it is called.

Summative assessment exists to record a judgment, which means it must be defensible: enough items, a rubric written in advance, and conditions that were the same for everyone.

Material sold as practice is usually formative in function and summative in appearance, which is worth stating plainly to a buyer. education:learning-materials-design covers the practice sequence this sits inside.

Sources

references/sources.md in this skill lists the outside authorities that settle the questions here — what each one is authoritative for, and what you may do with it. Check them before answering on anything they cover, and cite what you used. Most are free to read and not free to reproduce; the use note on each is binding.

Never

  • Write an item before naming the claim it is evidence for.
  • Assess an explanation standard with a recognition item.
  • Build distractors that are wrong without being diagnostic.
  • Record mastery from a single response.
  • Make the stem harder to read than the skill is to perform.

© cbrock84, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in verticals/education/skills/education/assessment-design of cbrock84/headcount.

  • SKILL.md
  • references/sources.md

Open the folder on GitHubat commit 98d1c17

Compare with similar skills

Assessment Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Assessment Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Assessment Design this skillcbrock84/headcount2k—~1.3kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Deep Reading Analystginobefun/deep-reading-analyst-skill3544 repos~3.6kAutomated safety check: PassMIT
OpenMAIC Setup and ExtensionTHU-MAIC/OpenMAIC40k—~1.7kAutomated safety check: NotesMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated 2 days ago
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Deep Reading Analyst

    ginobefun/deep-reading-analyst-skill

    Comprehensive framework for deep analysis of articles, papers, and long-form content using 10+ thinking models (SCQA, 5W2H, critical thinking, inversion, mental models, first principles, systems…

    354 GitHub starsUsed in 4 repos~3.6k tokens
    EducationAuto-check passed
  • Guides setup, classroom generation and secondary development for OpenMAIC, the multi-agent interactive classroom, one confirmed phase at a time.

    40k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check: notes
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • Zhang Xuefeng Perspective

    alchaincyf/zhangxuefeng-skill

    Answers education and career questions in the voice of Zhang Xuefeng, looking up current employment and admissions data before giving a direct verdict.

    10k GitHub stars~2.6k tokensUpdated 1 mo ago
    EducationAuto-check passed

More from cbrock84/headcount

All 178 skills in this repo
  • Agent Hierarchy

    cbrock84/headcount

    Designs orchestrator-and-subagent hierarchies for a repository — splitting agents by exclusive write surface, pairing every producer with an independent auditor, and enforcing the split with a…

    2k GitHub stars~1.2k tokensUpdated 22 days ago
    Auto-check passed
  • Access And Identity

    cbrock84/headcount

    Designs and audits who can reach what — authentication, authorization models, privileged access, service credentials, and joiner-mover-leaver process.

    2k GitHub stars~1.1k tokensUpdated 22 days ago
    Auto-check passed
  • Account Based Marketing

    cbrock84/headcount

    Concentrates marketing and sales effort on a named set of accounts rather than on volume — qualifying whether the model fits your economics at all, building the account list and the buying group…

    2k GitHub stars~1.2k tokensUpdated 22 days ago
    Auto-check passed
  • Activation

    cbrock84/headcount

    Gets new users from signup to first real value — signup flow, onboarding, time-to-value, and the early experience that determines whether someone becomes a user or a lapsed account.

    2k GitHub stars~865 tokensUpdated 22 days ago
    Auto-check passed
  • AI ML Governance

    cbrock84/headcount

    Governs models and AI systems in production — intended use, evaluation, monitoring, human oversight, documentation, and the decision to deploy or retire.

    2k GitHub stars~1k tokensUpdated 22 days ago
    Auto-check passed
  • AI Research Analyst

    cbrock84/headcount

    Produces executive-level research — market sizing, competitor mapping, trend analysis, and strategic intelligence — grounded in cited sources with the confidence in each claim made explicit.

    2k GitHub stars~916 tokensUpdated 22 days ago
    Auto-check passed

Categories

Questions about Assessment Design

What does Assessment Design do?

Builds questions and assessments that measure what they claim — matching item type to the thing being assessed, writing distractors that diagnose, and telling a question that measures learning from…. Assessment Design is an agent skill from cbrock84/headcount. Builds questions and assessments that measure what they claim — matching item type to the thing being assessed, writing distractors that diagnose, and telling a question that measures learning from one that measures reading speed or test-taking.

When should I use Assessment Design?

Assessment Design fits situations like: education work in your project.

How do I install Assessment Design in Claude Code?

Run `npx skills add cbrock84/headcount --skill assessment-design -a claude-code`. Or copy the skill folder (verticals/education/skills/education/assessment-design in cbrock84/headcount) into .claude/skills/assessment-design in your project. Claude Code loads it when a task matches its description.

How do I install Assessment Design in Codex?

Run `npx skills add cbrock84/headcount --skill assessment-design -a codex`. Or copy the skill folder (verticals/education/skills/education/assessment-design in cbrock84/headcount) into .agents/skills/assessment-design in your project. Codex loads it when a task matches its description.

Can I use Assessment Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cbrock84/headcount --skill assessment-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/assessment-design, .gemini/skills/assessment-design, .github/skills/assessment-design and .opencode/skills/assessment-design in your project.

What does Assessment Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Assessment Design is instructions for the agent only.

Does Assessment Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Assessment Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Assessment Design use?

Assessment Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Assessment Design use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 658 tokens, read only when the agent opens those files.

What are the alternatives to Assessment Design?

Skills that share tags, products or a category with Assessment Design: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Deep Reading Analyst (ginobefun/deep-reading-analyst-skill, 354 stars) and OpenMAIC Setup and Extension (THU-MAIC/OpenMAIC, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Assessment Design?

cbrock84 (a GitHub user) maintains it in cbrock84/headcount, which has 2,022 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on September 17, 2026.

Source: cbrock84/headcount on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.