Design AI Benchmarking
Aperivue/medsci-skills
A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.
100-point NL artifact rubric: penalty tables per artifact type, calibration cases.
$ npx skills add xiaolai/nlpm --skill scoring -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install xiaolai/nlpm scoring --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/xiaolai/nlpm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nlpm/scoring .claude/skills/scoring && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scoring" agent skill from https://github.com/xiaolai/nlpm/tree/main/skills/nlpm/scoring into .claude/skills/scoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scoring", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/xiaolai/nlpm/tree/main/skills/nlpm/scoringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add xiaolai/nlpm --skill scoring -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install xiaolai/nlpm scoring --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xiaolai/nlpm.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nlpm/scoring .agents/skills/scoring && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scoring" agent skill from https://github.com/xiaolai/nlpm/tree/main/skills/nlpm/scoring into .agents/skills/scoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scoring", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add xiaolai/nlpm --skill scoring -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install xiaolai/nlpm scoring --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xiaolai/nlpm.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nlpm/scoring .cursor/skills/scoring && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scoring" agent skill from https://github.com/xiaolai/nlpm/tree/main/skills/nlpm/scoring into .cursor/skills/scoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scoring", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/xiaolai/nlpm.git --path skills/nlpm/scoring--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add xiaolai/nlpm --skill scoring -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install xiaolai/nlpm scoring --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xiaolai/nlpm.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nlpm/scoring .gemini/skills/scoring && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scoring" agent skill from https://github.com/xiaolai/nlpm/tree/main/skills/nlpm/scoring into .gemini/skills/scoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scoring", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install xiaolai/nlpm scoringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add xiaolai/nlpm --skill scoring -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/xiaolai/nlpm.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nlpm/scoring .github/skills/scoring && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scoring" agent skill from https://github.com/xiaolai/nlpm/tree/main/skills/nlpm/scoring into .github/skills/scoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scoring", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add xiaolai/nlpm --skill scoring -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install xiaolai/nlpm scoring --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/xiaolai/nlpm.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nlpm/scoring .opencode/skills/scoring && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scoring" agent skill from https://github.com/xiaolai/nlpm/tree/main/skills/nlpm/scoring into .opencode/skills/scoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scoring", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scoring100-point NL artifact rubric: penalty tables per artifact type, calibration cases.
Scoring is an agent skill from xiaolai/nlpm. 100-point NL artifact rubric: penalty tables per artifact type, calibration cases.
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/antigravity.md`, `references/calibration-examples.md` and `references/codex.md`).
It sits in Education, covering Quizzes and assessments and Performance reviews. The repository describes itself as: Natural-Language Programming Manager — scan, lint, and score NL artifacts with Claude-native quality scoring. The licence is ISC.
Read from SKILL.md and the folder at commit 6fdbd05. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scoring loads about 5.3k tokens when it runs, and up to ~8.8k if it reads all its reference files. Until then it costs about 23 tokens; SKILL.md has 2,544 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from xiaolai/nlpm at commit 6fdbd05, republished under its ISC licence (© xiaolai). 2,544 words, ~5,266 tokens.
.claude/skills/scoring/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.100-point quality scale for all NL programming artifacts. Apply penalties deterministically. Use calibration examples to anchor judgment on borderline cases.
base_score = 100
adjustments = sum of all applicable penalties (all penalties are negative)
final_score = max(0, min(100, base_score + adjustments))Penalties stack. The floor is 0; the ceiling is 100. No bonuses — the default assumption is that an artifact is well-formed, and quality is measured by what is missing or wrong.
The tables in this file cover every artifact type a Claude Code project contains. Tables for the remaining artifact types live in references/, one file per group. When the artifact you are scoring is listed below, Read that file and apply its tables as if they appeared here — they carry the same weight, and a finding cited to one of their rows is a rubric finding.
| Artifact being scored | Read |
|---|---|
Codex hook events; .codex-plugin/plugin.json; .agents/plugins/marketplace.json; agents/openai.yaml; .codex/config.toml | references/codex.md |
Antigravity / Gemini-lineage hook events; gemini-extension.json; .gemini/commands/*.toml | references/antigravity.md |
Memory files (~/.claude/projects/*/memory/*.md); agent workflow programs (project-root program.md-style files) | references/memory-and-workflow.md |
The universal Hooks checks below apply to Codex and Antigravity hook configs too.
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| -- | name present | Missing | -25 |
| -- | name matches parent directory | Frontmatter name: value does not equal parent directory name (per nlpm:conventions §5 — open spec MUST) | -15 |
| R04 | description present | Missing | -25 |
| R04 | Trigger quality | Description is generic (≤1 specific phrase) | -15 |
| R04 | Description length | Description 500–800 chars | -5 |
| R04 | Description length | Description >800 chars | -10 |
| R05 | Body length | 400–500 lines | -5 |
| R05 | Body length | >500 lines | -10 |
| R06 | Code examples | Complex concepts with no examples | -5 |
| R06 | Code examples | No examples at all in a technical skill | -10 |
| R06 | <example> blocks | Zero <example> blocks on a user-invocable skill (user-invocable absent or true) | -10 |
| R07 | Scope note | No scope note / cross-references | -3 |
Scope-note discipline: R07 means "scope note when related skills exist." Do NOT apply R07 to missing example blocks — that is the new R06 row above (penalty -10, not -15). The 2026-05-13 lijigang/ljg-skills audit applied R07 + −15 fourteen times for missing example blocks; both labels were wrong (R07 is not example-related, and -15 is the agents penalty, not the skills penalty). The validator at
auditor/scripts/validate-rule-ids.pycatches this kind of drift in CI.
<example>-block counting discipline (added 2026-08-01, origin: xiaolai/cc-suite v1.3.1 remediation): an<example>block counts only when it sits outside fenced code blocks — in the body or in a frontmatter description block scalar. Tags inside a fenced template (```markdown … ```) are illustrative content, and prose that names the string`<example>`is a mention, not a block. A 2026-07-31 scoring pass credited a skill with example blocks that existed only inside a fenced template, hiding a real R06 violation across 13 files. Verify by reading the file, not by grepping for the tag.
namematches parent directory (added 2026-05-25, audit: google/skills): the open Agent Skills spec at agentskills.io makes this a MUST. Mismatch is deterministic, high-confidence, and reproducible by single-line diff (frontmattername:vsbasename($(dirname FILE))). Mark such findingsconfidence: highper the manifest-vs-disk-diff principle inagents/scorer.mdstep 6. Note: this penalty did not exist before 2026-05-25 — re-scoring past audits will yield slightly lower scores for any corpus containing this defect, but no contribute outcomes are retroactively affected since no PRs were ever opened against a name-mismatch finding under the prior rubric.
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| R09 | description present | Missing | -25 |
| R09 | <example> blocks | Zero <example> blocks in the description | -15 |
| R09 | Exclusion clause | Description names no situation the agent is not for (a "Not for …"/"Do not use …" sentence, or an example where the assistant declines or routes elsewhere) | -5 |
| R09 | Description length | Description value over 1,200 characters, examples included | -5 |
| R10 | model declared | Not declared | -5 |
| R10 | model appropriate | Wrong tier for task (e.g. opus for parsing) | -5 |
| R11 | tools declared | Not declared | -5 |
| R11 | Unused tools | Each tool declared but not used in body | -3 each |
| R12 | Output format | No output format spec in body | -10 |
| R11 | Write on read-only | Audit/review/scan agent declares Write or Edit | -10 |
R09 example budget (changed in 1.4.0): one
<example>block is full credit. An agent'sdescriptionsits in the Agent tool's text on every turn and is the only thing Claude sees when choosing an agent, so examples stay in the description — moving them into the body hides them from routing. One well-chosen example plus an explicit exclusion ("Not for …") carries the routing signal at a fraction of the tokens; extra examples are allowed but cost always-on context, which is why R09 deducts 5 points when the description exceeds 1,200 characters.
min_examplesoverride (R09: { min_examples: N }in the project'snlpm.local.md, seenlpm:conventions§6): with N greater than 1, a description that has at least one but fewer than N<example>blocks costs -5 per missing block, capped at -15 so that a partial set never costs more than none; zero blocks stays -15. The exclusion-clause and length rows are unchanged. Without the override, N is 1 and the rows above apply as written.
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| -- | description present | Missing | -25 |
| R18 | argument-hint present | Command takes input but no hint | -5 |
| R14 | Steps numbered | Multi-step body with no numbered steps | -10 |
| R15 | Empty input handling | No handling for empty/missing input | -10 |
| R16 | Output format | No output format defined | -10 |
| R17 | Error paths | No error handling for missing files or bad data | -5 |
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| R19 | user-invocable: false | Missing or set to true | -25 |
| R20 | Purpose clear | Description doesn't state it's a partial | -10 |
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| R21 | description present | Missing frontmatter description | -10 |
| R21 | Format: bold imperative | No bold imperative opening | -5 |
| R21 | Format: rationale | No rationale following the imperative | -10 |
| R22 | Enforceability | Rule is not specific/testable | -10 |
| R23 | Budget | Rule file over 500 lines | -15 |
| R26 | Conflicts with other rules | Direct contradiction with another rule in same set | -20 |
| R24 | Duplicates tooling | Re-states what eslint/ruff/clippy already catches | -10 |
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| -- | Valid syntax | Hook config file fails to parse (JSON or TOML per tool) | -25 |
| R29 | Scripts exist | Referenced script file does not exist | -20 |
| -- | Command safety | Hook command contains dangerous patterns (rm -rf, git push --force, DROP TABLE) | -15 |
| -- | Matcher regex valid | Matcher pattern doesn't compile as valid regex | -10 |
| -- | Timeout reasonable | Hook specifies timeout > 30s (likely hangs) | -5 |
Authoritative event list: nlpm:conventions-claude §7 together with its extended allow-list in conventions-claude/reference.md. Per the multi-tool design (analysis/multi-tool-design-2026-05.md decision #4), Claude / Codex / Antigravity hook event vocabularies are NOT 1:1 mappable — three separate tables, no translation. The Codex and Antigravity tables are in references/codex.md and references/antigravity.md.
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| R27 | Event names valid (Claude) | Uses an event name that is in neither the nlpm:conventions-claude §7 table nor its extended allow-list (conventions-claude/reference.md, Hook Events). Those two lists together are the confirmed set (33 events, including SubagentStop, SubagentStart, PreCompact and Notification); a name missing from both still passes if code.claude.com/docs/en/hooks.md documents it | -15 |
| R27 | Case correct (Claude) | Event name has wrong case (e.g. pretooluse) | -10 |
| -- | Hook type valid (Claude) | Uses unrecognized type value — confirmed Claude types: command, http, mcp_tool, prompt, agent | -10 |
| -- | MCP matcher format (Claude) | Matcher targets MCP tool but doesn't use mcp__<server>__<tool> pattern | -5 |
.claude-plugin/plugin.json)| Check | Condition | Penalty |
|---|---|---|
name present | Missing | -25 |
version is semver | Present but not valid semver | -10 |
description present | Missing | -5 |
Schema reference: nlpm:conventions-claude §17 and its reference.md (Plugin Distribution).
| Check | Condition | Penalty |
|---|---|---|
| Valid JSON | File fails JSON parse | -25 |
name present | Missing | -25 |
owner.name present | owner missing, or it has no name | -10 |
plugins array present | Missing or empty | -10 |
Per-plugin name and source | Either missing | -10 each |
Per-plugin source valid | A relative path that doesn't start with ./ (bare names are valid only under metadata.pluginRoot), or an object whose source.source isn't one of the six object types in reference.md or that lacks that type's required field | -10 each |
Per-plugin version in step | Differs from the version in the plugin.json it describes, when that file is in the same repository (plugin.json wins at load, so the entry is stale) | -5 each |
Per-plugin description present | Missing (the /plugin browser shows it) | -3 each |
.mcp.json at repo root)| Check | Condition | Penalty |
|---|---|---|
| Valid JSON | File fails JSON parse | -25 |
Server command present | MCP server entry missing command field | -15 |
Schema details: nlpm:conventions-claude §12. Stable in 2026.
| Check | Condition | Penalty |
|---|---|---|
| Valid JSON | File fails JSON parse | -25 |
Schema details: nlpm:conventions-claude §13. Stable in 2026.
| Check | Condition | Penalty |
|---|---|---|
| Valid JSON | File fails JSON parse | -25 |
| Check | Condition | Penalty |
|---|---|---|
| Valid JSON | File fails JSON parse | -25 |
| No hardcoded secrets | Contains API keys, tokens, or passwords | -25 |
| Permission mode sanity | bypassPermissions enabled in a shared project settings file (not .local) | -15 |
| Recognized keys | Contains unknown top-level keys not in Claude Code schema | -5 each, cap -15 |
| Hook definitions valid | hooks key present — check event names valid and case-correct | -10 per invalid |
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| R49 | File exists | Neither AGENTS.md nor CLAUDE.md in plugin root | -10 |
| -- | Under 200 lines | CLAUDE.md exceeds 200 lines | -5 |
| R38 | Actionable content | CLAUDE.md has no actionable guidance (just filler) | -10 |
| R33 | Build/run command | No instructions for how to build or run the project | -10 |
| R34 | Test command | No instructions for how to run tests | -5 |
| R35 | Architecture overview | No structure/component description (what lives where) | -5 |
| R36 | Valid @ imports | Contains @ import syntax referencing a file that doesn't exist | -10 |
| R37 | No stale file references | Mentions files or functions that no longer exist in the repo | -10 |
| R38 | Actionability ratio | >60% of content is description rather than instructions | -5 |
| -- | Prerequisites section | No section covering required tools, versions, or setup steps | -5 |
| R39 | No rule conflicts | CLAUDE.md says X while a .claude/rules/ file says not-X | -15 |
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| R01 | Vague quantifier | Each occurrence of: "appropriate", "relevant", "as needed", "sufficient", "adequate", "reasonable", "properly", "correctly", "some", "several", "various" without measurable criteria | -2 each |
| R01 | Vague quantifier cap | Total vague quantifier penalty | max -20 |
Mention-versus-use exclusion (added 2026-08-01, origin: xiaolai/cc-suite audit-family false positives; design reviewed via Codex consultation): do not count a vague term when it is presented as a literal token AND the containing clause explicitly instructs the reader or a tool to detect, flag, reject, replace, avoid, or report that term — audit tooling must be able to name the words it hunts (e.g.
Flag uses of `some`, `several`, `various` without concrete criteriais R01's own job description, not a violation). Backtick or quotation formatting alone does NOT qualify: a term that still modifies an action, criterion, or requirement is counted even when backticked —handle errors `properly`remains a violation.
Applied only when R51: { enabled: true, vocabulary_skill: <path> } appears in .claude/nlpm.local.md. Without the opt-in, R51 contributes zero penalty regardless of artifact content. The configured vocabulary_skill must contain a registry.yaml listing canonical and deprecated terms; without it, R51 emits an advisory and contributes zero penalty.
| Rule | Check | Condition | Penalty |
|---|---|---|---|
| R51 | Deprecated synonym | Each occurrence of a term marked deprecated: in the project's registry.yaml, in the scope the artifact belongs to | -2 each |
| R51 | Drift cap | Total R51 penalty | max -10 per file |
| R51 | Missing registry | enabled: true but vocabulary_skill: not set or points to a directory with no registry.yaml | 0 (advisory only) |
Why opt-in: vocabulary discipline is high-leverage for projects with accumulated drift but premature for projects still discovering their domain. Each project decides when it has enough literary warrant (P6) to lock terms in. See
analysis/vocabulary-design-principles.mdfor the six principles R51 operationalizes.
Registry-declaration exclusion (added 2026-08-01, origin: xiaolai/cc-suite vocabulary-skill self-reference; design reviewed via Codex consultation): the file that declares a deprecation must name the deprecated term to do so. Within the configured
vocabulary_skillpath, do not count a deprecated term where it occurs in the declaration that registers it or maps it to its replacement —registry.yamldeprecated:lists and the SKILL.md deprecation tables. The exclusion covers ONLY those declaration term fields: deprecated terms in surrounding prose inside thevocabulary_skillpath are counted, and the path is not categorically exempt. The same mention-versus-use principle as R01's exclusion above, applied to R51.
Applied when linting an entire plugin rather than individual files.
| Check | Condition | Penalty |
|---|---|---|
| Broken partial refs | Command references commands/shared/X.md that doesn't exist | -20 |
| Broken skill refs | Agent references plugin:skill that isn't installed | -20 |
| Missing scripts | Hook references script that doesn't exist | -20 |
| Orphaned files | Agent/command/skill file not referenced by anything | -5 per file |
| Contradictions | Two rules/instructions in same plugin directly contradict each other | -15 per pair |
| Range | Label | Meaning |
|---|---|---|
| 90–100 | Excellent | Production-ready; minor or no findings |
| 80–89 | Good | Solid; one or two non-critical gaps |
| 70–79 | Adequate | Meets threshold; noticeable gaps to address |
| 60–69 | Weak | Below threshold; significant findings |
| <60 | Rewrite | Fundamental problems; recommend rewriting from scratch |
Default pass threshold: 70. Configurable in .claude/nlpm.local.md.
Four worked examples — Excellent Agent (97), Rewrite Agent (41), Excellent Rule (92), Weak Rule (41) — live in references/calibration-examples.md. Load that file on demand when scoring a borderline case (around band boundaries: 88-92, 68-72, 58-62) and you need an anchored reference.
The examples are not needed for routine scoring — the penalty tables (above, plus the per-tool reference files indexed under Penalty Tables) are self-contained. They were extracted from this file 2026-05-28 to keep the rubric under R05's 500-line body budget while preserving the calibration material verbatim.
This skill covers the NLPM scoring formula, penalty tables, score bands, and calibration examples. It does NOT cover:
nlpm:conventions (universal) and the tool-specific overlays nlpm:conventions-claude, nlpm:conventions-codex, nlpm:conventions-antigravity (the latter three created in PR-B; see analysis/multi-tool-design-2026-05.md)nlpm:patternscommands/score.mdnlpm now scores artifacts across three tool ecosystems — Claude Code, Codex CLI, and Antigravity (which absorbs Gemini CLI on 2026-06-18). The tier classification in agents/scorer.md separates open-spec (Tier 1), Tier 1.5 open-spec corpora, and per-tool Tier 2 overlays (2-Claude / 2-Codex / 2-Antigravity).
references/). The three tools' event vocabularies are not 1:1 mappable; no universal translation layer. See analysis/multi-tool-design-2026-05.md decision #4.references/codex.md): .codex-plugin/plugin.json, .agents/plugins/marketplace.json, .codex/config.toml, agents/openai.yaml sidecars.references/antigravity.md; advisory-only until spec stabilizes): gemini-extension.json, .gemini/commands/*.toml, Antigravity hook events..lsp.json, monitors/monitors.json (validate JSON-parse only until detailed schemas land). New SKILL.md fields documented in nlpm:conventions-claude.The following findings have historically been reported by the scorer despite having no backing in this rubric. They MUST NOT be penalized:
| Invalid finding | Why it is invalid |
|---|---|
Missing namespace: on skill | Not in the skill schema; conventions §5 does not list it |
Missing inline hooks:/skills: registration blocks in plugin.json | conventions §1 defines these as optional path strings |
AskUserQuestion / Task / WebFetch flagged as undocumented tool | Built-in per conventions-claude §16 |
Agent missing skills: when omission is documented in CLAUDE.md | Intentional architectural choice |
plugin.json missing engines: / minClaudeVersion: / main: | All optional per conventions §1 |
| plugin.json description shorter than sibling marketplace.json description | Desynchronization ≠ defect; only penalize if required field is absent |
When in doubt: if a finding cannot be cited to a specific row in the penalty tables above or in a reference file indexed under Penalty Tables, drop it.
© xiaolai, ISC. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/nlpm/scoring of xiaolai/nlpm.
Open the folder on GitHubat commit 6fdbd05
Scoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scoring this skillxiaolai/nlpm | 146 | — | ~5.3k | Automated safety check: Pass | ISC | |
| Design AI BenchmarkingAperivue/medsci-skills | 329 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Evaluator CalibrationArchive228/loopkit | 756 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Interview System Designerborghei/Claude-Skills | 881 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Advanced Evaluationguanyang/open-agent-hub | 975 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | |
| DeepTutor CLIHKUDS/DeepTutor | 41k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 |
Aperivue/medsci-skills
A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.
Archive228/loopkit
Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.
borghei/Claude-Skills
Design calibrated interview loops, competency-based question banks, and hiring calibration.
guanyang/open-agent-hub
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
rohitg00/ai-engineering-from-scratch
Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.
xiaolai/nlpm
Universal NL conventions: SKILL.md open spec, AGENTS.md, vague quantifiers, naming.
xiaolai/nlpm
Antigravity and Gemini CLI artifact schemas: .gemini/ paths, extensions, hooks.
xiaolai/nlpm
Codex CLI artifact schemas: config.toml, .codex-plugin, skills, hooks, AGENTS.md.
xiaolai/nlpm
Multi-agent workflow patterns: parallel dispatch, pipelines, QC gates, retries.
xiaolai/nlpm
NL artifact anti-patterns: vague quantifiers, bare prohibitions, oversized skills.
xiaolai/nlpm
NL artifact test specs for /nlpm:test: spec format, TDD for skills and agents.
Categories
100-point NL artifact rubric: penalty tables per artifact type, calibration cases. Scoring is an agent skill from xiaolai/nlpm. 100-point NL artifact rubric: penalty tables per artifact type, calibration cases.
Scoring fits situations like: tasks that involve Quizzes and assessments; tasks that involve Performance reviews.
Run `npx skills add xiaolai/nlpm --skill scoring -a claude-code`. Or copy the skill folder (skills/nlpm/scoring in xiaolai/nlpm) into .claude/skills/scoring in your project. Claude Code loads it when a task matches its description.
Run `npx skills add xiaolai/nlpm --skill scoring -a codex`. Or copy the skill folder (skills/nlpm/scoring in xiaolai/nlpm) into .agents/skills/scoring in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xiaolai/nlpm --skill scoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scoring, .gemini/skills/scoring, .github/skills/scoring and .opencode/skills/scoring in your project.
Going by SKILL.md and its folder, Scoring needs the command-line tools its instructions call (git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Scoring is published under the ISC licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Scoring: Design AI Benchmarking (Aperivue/medsci-skills, 329 stars), Evaluator Calibration (Archive228/loopkit, 756 stars), Interview System Designer (borghei/Claude-Skills, 881 stars) and Advanced Evaluation (guanyang/open-agent-hub, 975 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
xiaolai (a GitHub user) maintains it in xiaolai/nlpm, which has 146 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.
Source: xiaolai/nlpm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.