Agent skill

Assessment Design Guide

by wentorai in wentorai/research-plugins

Psychometrics and educational assessment design for researchers

MITAuto-check passedEducation

Install Assessment Design Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill assessment-design-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins assessment-design-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/domains/education/assessment-design-guide .claude/skills/assessment-design-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
assessment-design-guide
GitHub stars
298
Used in
1 other repo
Token cost
~1.9k tokens
SKILL.md length
460 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Psychometrics and educational assessment design for researchers

  • Works in 5 steps: Content evidence: Expert review confirms… → Response process evidence: Think-aloud… → Internal structure evidence: Factor… → …
  • Education work in your project
  • SKILL.md covers Classical Test Theory, Item Response Theory, Validity Evidence and Computerized Adaptive Testing, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Assessment Design Guide is an agent skill from wentorai/research-plugins. Psychometrics and educational assessment design for researchers

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Education work in your project

Example prompts

  • “Use the assessment-design-guide skill to psychometric and educational assessment design for researchers”
  • “/assessment-design-guide”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Content evidence: Expert review confirms items represent the construct domain
  2. Response process evidence: Think-aloud protocols confirm examinees engage intended cognitive processes
  3. Internal structure evidence: Factor analysis confirms dimensionality matches the test blueprint
  4. Relations to other variables: Correlations with external criteria (convergent, discriminant, predictive)
  5. Consequences evidence: Test use leads to intended benefits without unintended harm

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Assessment Design Guide loads about 1.9k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 460 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 460 words, ~1,908 tokens.

Download SKILL.mdSave it as .claude/skills/assessment-design-guide/SKILL.md (or your agent's skills folder).
name
assessment-design-guide
description
Psychometrics and educational assessment design for researchers

Assessment Design Guide

A skill for designing, validating, and analyzing educational assessments using modern psychometric methods. Covers classical test theory, item response theory, test construction, validity evidence, and computerized adaptive testing.

Classical Test Theory

Reliability Analysis

Classical test theory (CTT) models observed scores as the sum of a true score and error:

X = T + E

Key reliability coefficients:

CoefficientMethodInterpretation
Cronbach's alphaInternal consistencyHomogeneity of items
Test-retestStability over timeTemporal consistency
Parallel formsEquivalent test versionsForm equivalence
Split-half (Spearman-Brown)Odd-even item splitInternal consistency
Inter-rater (Cohen's kappa)Multiple ratersScoring agreement
python
import numpy as np
import pandas as pd

def item_analysis(responses: pd.DataFrame, total_scores: pd.Series) -> pd.DataFrame:
    """
    Classical item analysis: difficulty, discrimination, point-biserial.
    responses: binary DataFrame (1=correct, 0=incorrect), items as columns.
    total_scores: total test score for each examinee.
    """
    results = []
    for item in responses.columns:
        scores = responses[item]
        difficulty = scores.mean()  # p-value (proportion correct)

        # Point-biserial correlation
        corr = scores.corr(total_scores)

        # Upper-lower discrimination (top/bottom 27%)
        n = len(total_scores)
        cutoff_high = total_scores.quantile(0.73)
        cutoff_low = total_scores.quantile(0.27)
        upper = scores[total_scores >= cutoff_high].mean()
        lower = scores[total_scores <= cutoff_low].mean()
        discrimination = upper - lower

        results.append({
            "item": item,
            "difficulty": round(difficulty, 3),
            "discrimination": round(discrimination, 3),
            "point_biserial": round(corr, 3),
            "flag": "review" if difficulty < 0.2 or difficulty > 0.9
                    or discrimination < 0.2 else "ok"
        })
    return pd.DataFrame(results)
Item Selection Guidelines
  • Difficulty: Aim for p-values between 0.30 and 0.80 for maximum discrimination
  • Discrimination: Items with D < 0.20 should be revised or removed
  • Distractors: Each distractor should attract at least 5% of examinees
  • Point-biserial: Should be positive and ideally above 0.25

Item Response Theory

The Three-Parameter Logistic Model

IRT provides a more rigorous framework than CTT by modeling the probability of a correct response as a function of ability and item parameters:

python
import numpy as np

def irt_3pl(theta: float, a: float, b: float, c: float) -> float:
    """
    Three-parameter logistic IRT model.
    theta: examinee ability (typically -3 to +3)
    a: discrimination parameter (slope, typically 0.5 to 2.5)
    b: difficulty parameter (location, same scale as theta)
    c: guessing parameter (lower asymptote, typically 0.0 to 0.35)
    Returns: probability of correct response
    """
    exponent = -a * (theta - b)
    return c + (1 - c) / (1 + np.exp(exponent))

# Item characteristic curves for three items
thetas = np.linspace(-3, 3, 100)
item_easy = [irt_3pl(t, a=1.0, b=-1.0, c=0.2) for t in thetas]
item_medium = [irt_3pl(t, a=1.5, b=0.0, c=0.2) for t in thetas]
item_hard = [irt_3pl(t, a=1.2, b=1.5, c=0.2) for t in thetas]
IRT Model Estimation
python
# Using the 'mirt' package in R (called via rpy2 or standalone)
# R code for fitting a 2PL model:
r_code = """
library(mirt)

# responses: binary matrix (examinees x items)
mod <- mirt(responses, model = 1, itemtype = "2PL")

# Item parameters
coef(mod, simplify = TRUE)

# Ability estimates (Expected A Posteriori)
theta_hat <- fscores(mod, method = "EAP")

# Model fit
M2(mod)  # limited-information fit statistic
itemfit(mod, fit_stats = "S_X2")
"""
Model Comparison
ModelParametersUse Case
Rasch (1PL)b onlyEqual discrimination assumed; measurement-focused
2PLa, bDifferent discrimination; general purpose
3PLa, b, cMultiple choice with guessing
Graded Responsea, b_kLikert-scale or partial credit items
Nominal Responsea_k, c_kMultiple choice with informative distractors

Validity Evidence

The Unified Validity Framework

Following the Standards for Educational and Psychological Testing (AERA/APA/NCME, 2014), validity is a unitary concept supported by five types of evidence:

  1. Content evidence: Expert review confirms items represent the construct domain
  2. Response process evidence: Think-aloud protocols confirm examinees engage intended cognitive processes
  3. Internal structure evidence: Factor analysis confirms dimensionality matches the test blueprint
  4. Relations to other variables: Correlations with external criteria (convergent, discriminant, predictive)
  5. Consequences evidence: Test use leads to intended benefits without unintended harm
python
from factor_analyzer import FactorAnalyzer

# Confirmatory approach: check dimensionality
fa = FactorAnalyzer(n_factors=3, rotation="promax")
fa.fit(item_responses)

# Eigenvalues for scree plot
eigenvalues, _ = fa.get_eigenvalues()
print("Eigenvalues:", eigenvalues[:10])

# Factor loadings
loadings = pd.DataFrame(
    fa.loadings_,
    columns=["Factor1", "Factor2", "Factor3"],
    index=item_names
)
print(loadings.round(3))
Show full SKILL.md (159 more words)Show less

Computerized Adaptive Testing

CAT Algorithm

Computerized adaptive testing selects items in real time to match examinee ability:

Initialize: theta_0 = 0 (prior mean)
For each item i = 1, 2, ..., until stopping rule met:
    1. Select item with maximum Fisher information at current theta
    2. Administer item, observe response
    3. Update theta estimate using maximum likelihood or Bayesian EAP
    4. Check stopping rule:
       - Fixed length (e.g., 30 items)
       - SE(theta) < threshold (e.g., 0.30)
       - Maximum time reached
Return: final theta estimate and standard error
Item Exposure Control

To prevent overuse of high-quality items and maintain test security:

  • Sympson-Hetter method: Set maximum exposure rates per item (e.g., 0.25)
  • a-stratified method: Divide item bank into strata by discrimination, sample within strata
  • Shadow test approach: Assemble full shadow tests at each step, administer the optimal item from the shadow test

Tools and Software

  • R mirt package: Full-featured IRT estimation, DIF analysis, CAT simulation
  • Python irt library (py-irt): Bayesian IRT models using PyTorch
  • jMetrik: Open-source Java application for classical and IRT analysis
  • TAO (Testing Assistee par Ordinateur): Open-source assessment delivery platform
  • Concerto: Open-source adaptive testing platform from Cambridge

Key References

  • Embretson, S.E. and Reise, S.P. (2000). Item Response Theory for Psychologists. Lawrence Erlbaum.
  • de Ayala, R.J. (2022). The Theory and Practice of Item Response Theory (2nd ed.). Guilford Press.
  • AERA, APA, and NCME (2014). Standards for Educational and Psychological Testing.

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/domains/education/assessment-design-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Assessment Design Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Assessment Design Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Assessment Design Guide this skillwentorai/research-plugins2981 repos~1.9kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Deep Reading Analystginobefun/deep-reading-analyst-skill3544 repos~3.6kAutomated safety check: PassMIT
OpenMAIC Setup and ExtensionTHU-MAIC/OpenMAIC40k—~1.7kAutomated safety check: NotesMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated 2 days ago
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Deep Reading Analyst

    ginobefun/deep-reading-analyst-skill

    Comprehensive framework for deep analysis of articles, papers, and long-form content using 10+ thinking models (SCQA, 5W2H, critical thinking, inversion, mental models, first principles, systems…

    354 GitHub starsUsed in 4 repos~3.6k tokens
    EducationAuto-check passed
  • Guides setup, classroom generation and secondary development for OpenMAIC, the multi-agent interactive classroom, one confirmed phase at a time.

    40k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check: notes
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • Zhang Xuefeng Perspective

    alchaincyf/zhangxuefeng-skill

    Answers education and career questions in the voice of Zhang Xuefeng, looking up current employment and admissions data before giving a direct verdict.

    10k GitHub stars~2.6k tokensUpdated 1 mo ago
    EducationAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Categories

Questions about Assessment Design Guide

What does Assessment Design Guide do?

Psychometrics and educational assessment design for researchers. Assessment Design Guide is an agent skill from wentorai/research-plugins.

When should I use Assessment Design Guide?

Assessment Design Guide fits situations like: education work in your project.

How do I install Assessment Design Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill assessment-design-guide -a claude-code`. Or copy the skill folder (skills/domains/education/assessment-design-guide in wentorai/research-plugins) into .claude/skills/assessment-design-guide in your project. Claude Code loads it when a task matches its description.

How do I install Assessment Design Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill assessment-design-guide -a codex`. Or copy the skill folder (skills/domains/education/assessment-design-guide in wentorai/research-plugins) into .agents/skills/assessment-design-guide in your project. Codex loads it when a task matches its description.

Can I use Assessment Design Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill assessment-design-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/assessment-design-guide, .gemini/skills/assessment-design-guide, .github/skills/assessment-design-guide and .opencode/skills/assessment-design-guide in your project.

What does Assessment Design Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: Assessment Design Guide is instructions for the agent only. Our summary lists: Python 3.

Does Assessment Design Guide access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Assessment Design Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Assessment Design Guide use?

Assessment Design Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Assessment Design Guide use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Assessment Design Guide?

Skills that share tags, products or a category with Assessment Design Guide: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Deep Reading Analyst (ginobefun/deep-reading-analyst-skill, 354 stars) and OpenMAIC Setup and Extension (THU-MAIC/OpenMAIC, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Assessment Design Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.