Agent skill

Scoring

by xiaolai in xiaolai/nlpm

100-point NL artifact rubric: penalty tables per artifact type, calibration cases.

ISCAuto-check passedEducation

Install Scoring

skills CLI
$ npx skills add xiaolai/nlpm --skill scoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xiaolai/nlpm scoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xiaolai/nlpm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nlpm/scoring .claude/skills/scoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scoring
GitHub stars
146
Token cost
~5.3k tokens
SKILL.md length
2,544 words
Files
5 (incl. references)
Skills in repo
15
Repo updated
First seen
Licence
ISC

At a glance

100-point NL artifact rubric: penalty tables per artifact type, calibration cases.

  • Tasks that involve Quizzes and assessments
  • SKILL.md covers Scoring Formula, Penalty Tables, Score Bands and Calibration Examples, plus 1 more section
  • Calls git
  • Tasks that involve Performance reviews

What it does

Scoring is an agent skill from xiaolai/nlpm. 100-point NL artifact rubric: penalty tables per artifact type, calibration cases.

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/antigravity.md`, `references/calibration-examples.md` and `references/codex.md`).

It sits in Education, covering Quizzes and assessments and Performance reviews. The repository describes itself as: Natural-Language Programming Manager — scan, lint, and score NL artifacts with Claude-native quality scoring. The licence is ISC.

When your agent uses it

  • Tasks that involve Quizzes and assessments
  • Tasks that involve Performance reviews

Example prompts

  • “/scoring”

What it can do on your machine

Read from SKILL.md and the folder at commit 6fdbd05. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scoring loads about 5.3k tokens when it runs, and up to ~8.8k if it reads all its reference files. Until then it costs about 23 tokens; SKILL.md has 2,544 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from xiaolai/nlpm at commit 6fdbd05, republished under its ISC licence (© xiaolai). 2,544 words, ~5,266 tokens.

Download SKILL.mdSave it as .claude/skills/scoring/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
scoring
description
100-point NL artifact rubric: penalty tables per artifact type, calibration cases.
version
0.4.0
user-invocable
false

NLPM Quality Scoring Rubric

100-point quality scale for all NL programming artifacts. Apply penalties deterministically. Use calibration examples to anchor judgment on borderline cases.


Scoring Formula

base_score = 100
adjustments = sum of all applicable penalties (all penalties are negative)
final_score = max(0, min(100, base_score + adjustments))

Penalties stack. The floor is 0; the ceiling is 100. No bonuses — the default assumption is that an artifact is well-formed, and quality is measured by what is missing or wrong.


Penalty Tables

The tables in this file cover every artifact type a Claude Code project contains. Tables for the remaining artifact types live in references/, one file per group. When the artifact you are scoring is listed below, Read that file and apply its tables as if they appeared here — they carry the same weight, and a finding cited to one of their rows is a rubric finding.

Artifact being scoredRead
Codex hook events; .codex-plugin/plugin.json; .agents/plugins/marketplace.json; agents/openai.yaml; .codex/config.tomlreferences/codex.md
Antigravity / Gemini-lineage hook events; gemini-extension.json; .gemini/commands/*.tomlreferences/antigravity.md
Memory files (~/.claude/projects/*/memory/*.md); agent workflow programs (project-root program.md-style files)references/memory-and-workflow.md

The universal Hooks checks below apply to Codex and Antigravity hook configs too.

Skills
RuleCheckConditionPenalty
--name presentMissing-25
--name matches parent directoryFrontmatter name: value does not equal parent directory name (per nlpm:conventions §5 — open spec MUST)-15
R04description presentMissing-25
R04Trigger qualityDescription is generic (≤1 specific phrase)-15
R04Description lengthDescription 500–800 chars-5
R04Description lengthDescription >800 chars-10
R05Body length400–500 lines-5
R05Body length>500 lines-10
R06Code examplesComplex concepts with no examples-5
R06Code examplesNo examples at all in a technical skill-10
R06<example> blocksZero <example> blocks on a user-invocable skill (user-invocable absent or true)-10
R07Scope noteNo scope note / cross-references-3

Scope-note discipline: R07 means "scope note when related skills exist." Do NOT apply R07 to missing example blocks — that is the new R06 row above (penalty -10, not -15). The 2026-05-13 lijigang/ljg-skills audit applied R07 + −15 fourteen times for missing example blocks; both labels were wrong (R07 is not example-related, and -15 is the agents penalty, not the skills penalty). The validator at auditor/scripts/validate-rule-ids.py catches this kind of drift in CI.

<example>-block counting discipline (added 2026-08-01, origin: xiaolai/cc-suite v1.3.1 remediation): an <example> block counts only when it sits outside fenced code blocks — in the body or in a frontmatter description block scalar. Tags inside a fenced template (```markdown … ```) are illustrative content, and prose that names the string `<example>` is a mention, not a block. A 2026-07-31 scoring pass credited a skill with example blocks that existed only inside a fenced template, hiding a real R06 violation across 13 files. Verify by reading the file, not by grepping for the tag.

name matches parent directory (added 2026-05-25, audit: google/skills): the open Agent Skills spec at agentskills.io makes this a MUST. Mismatch is deterministic, high-confidence, and reproducible by single-line diff (frontmatter name: vs basename($(dirname FILE))). Mark such findings confidence: high per the manifest-vs-disk-diff principle in agents/scorer.md step 6. Note: this penalty did not exist before 2026-05-25 — re-scoring past audits will yield slightly lower scores for any corpus containing this defect, but no contribute outcomes are retroactively affected since no PRs were ever opened against a name-mismatch finding under the prior rubric.


Agents
RuleCheckConditionPenalty
R09description presentMissing-25
R09<example> blocksZero <example> blocks in the description-15
R09Exclusion clauseDescription names no situation the agent is not for (a "Not for …"/"Do not use …" sentence, or an example where the assistant declines or routes elsewhere)-5
R09Description lengthDescription value over 1,200 characters, examples included-5
R10model declaredNot declared-5
R10model appropriateWrong tier for task (e.g. opus for parsing)-5
R11tools declaredNot declared-5
R11Unused toolsEach tool declared but not used in body-3 each
R12Output formatNo output format spec in body-10
R11Write on read-onlyAudit/review/scan agent declares Write or Edit-10

R09 example budget (changed in 1.4.0): one <example> block is full credit. An agent's description sits in the Agent tool's text on every turn and is the only thing Claude sees when choosing an agent, so examples stay in the description — moving them into the body hides them from routing. One well-chosen example plus an explicit exclusion ("Not for …") carries the routing signal at a fraction of the tokens; extra examples are allowed but cost always-on context, which is why R09 deducts 5 points when the description exceeds 1,200 characters.

min_examples override (R09: { min_examples: N } in the project's nlpm.local.md, see nlpm:conventions §6): with N greater than 1, a description that has at least one but fewer than N <example> blocks costs -5 per missing block, capped at -15 so that a partial set never costs more than none; zero blocks stays -15. The exclusion-clause and length rows are unchanged. Without the override, N is 1 and the rows above apply as written.


Commands
RuleCheckConditionPenalty
--description presentMissing-25
R18argument-hint presentCommand takes input but no hint-5
R14Steps numberedMulti-step body with no numbered steps-10
R15Empty input handlingNo handling for empty/missing input-10
R16Output formatNo output format defined-10
R17Error pathsNo error handling for missing files or bad data-5

Shared Partials
RuleCheckConditionPenalty
R19user-invocable: falseMissing or set to true-25
R20Purpose clearDescription doesn't state it's a partial-10

Rules
RuleCheckConditionPenalty
R21description presentMissing frontmatter description-10
R21Format: bold imperativeNo bold imperative opening-5
R21Format: rationaleNo rationale following the imperative-10
R22EnforceabilityRule is not specific/testable-10
R23BudgetRule file over 500 lines-15
R26Conflicts with other rulesDirect contradiction with another rule in same set-20
R24Duplicates toolingRe-states what eslint/ruff/clippy already catches-10

Hooks — universal checks (apply to all tools)
RuleCheckConditionPenalty
--Valid syntaxHook config file fails to parse (JSON or TOML per tool)-25
R29Scripts existReferenced script file does not exist-20
--Command safetyHook command contains dangerous patterns (rm -rf, git push --force, DROP TABLE)-15
--Matcher regex validMatcher pattern doesn't compile as valid regex-10
--Timeout reasonableHook specifies timeout > 30s (likely hangs)-5
Hooks (Claude Code — Tier 2-Claude only)

Authoritative event list: nlpm:conventions-claude §7 together with its extended allow-list in conventions-claude/reference.md. Per the multi-tool design (analysis/multi-tool-design-2026-05.md decision #4), Claude / Codex / Antigravity hook event vocabularies are NOT 1:1 mappable — three separate tables, no translation. The Codex and Antigravity tables are in references/codex.md and references/antigravity.md.

RuleCheckConditionPenalty
R27Event names valid (Claude)Uses an event name that is in neither the nlpm:conventions-claude §7 table nor its extended allow-list (conventions-claude/reference.md, Hook Events). Those two lists together are the confirmed set (33 events, including SubagentStop, SubagentStart, PreCompact and Notification); a name missing from both still passes if code.claude.com/docs/en/hooks.md documents it-15
R27Case correct (Claude)Event name has wrong case (e.g. pretooluse)-10
--Hook type valid (Claude)Uses unrecognized type value — confirmed Claude types: command, http, mcp_tool, prompt, agent-10
--MCP matcher format (Claude)Matcher targets MCP tool but doesn't use mcp__<server>__<tool> pattern-5

plugin.json (Claude Code — .claude-plugin/plugin.json)
CheckConditionPenalty
name presentMissing-25
version is semverPresent but not valid semver-10
description presentMissing-5

.claude-plugin/marketplace.json (Claude Code — Tier 2-Claude)

Schema reference: nlpm:conventions-claude §17 and its reference.md (Plugin Distribution).

CheckConditionPenalty
Valid JSONFile fails JSON parse-25
name presentMissing-25
owner.name presentowner missing, or it has no name-10
plugins array presentMissing or empty-10
Per-plugin name and sourceEither missing-10 each
Per-plugin source validA relative path that doesn't start with ./ (bare names are valid only under metadata.pluginRoot), or an object whose source.source isn't one of the six object types in reference.md or that lacks that type's required field-10 each
Per-plugin version in stepDiffers from the version in the plugin.json it describes, when that file is in the same repository (plugin.json wins at load, so the entry is stale)-5 each
Per-plugin description presentMissing (the /plugin browser shows it)-3 each

.mcp.json (Claude Code — .mcp.json at repo root)
CheckConditionPenalty
Valid JSONFile fails JSON parse-25
Server command presentMCP server entry missing command field-15

.lsp.json (Claude Code LSP — Tier 2-Claude)

Schema details: nlpm:conventions-claude §12. Stable in 2026.

CheckConditionPenalty
Valid JSONFile fails JSON parse-25

monitors/monitors.json (Claude Code monitors — Tier 2-Claude)

Schema details: nlpm:conventions-claude §13. Stable in 2026.

CheckConditionPenalty
Valid JSONFile fails JSON parse-25

Settings Files (.claude/settings.json, .claude/settings.local.json)
CheckConditionPenalty
Valid JSONFile fails JSON parse-25
No hardcoded secretsContains API keys, tokens, or passwords-25
Permission mode sanitybypassPermissions enabled in a shared project settings file (not .local)-15
Recognized keysContains unknown top-level keys not in Claude Code schema-5 each, cap -15
Hook definitions validhooks key present — check event names valid and case-correct-10 per invalid

Show full SKILL.md (1,060 more words)Show less
CLAUDE.md
RuleCheckConditionPenalty
R49File existsNeither AGENTS.md nor CLAUDE.md in plugin root-10
--Under 200 linesCLAUDE.md exceeds 200 lines-5
R38Actionable contentCLAUDE.md has no actionable guidance (just filler)-10
R33Build/run commandNo instructions for how to build or run the project-10
R34Test commandNo instructions for how to run tests-5
R35Architecture overviewNo structure/component description (what lives where)-5
R36Valid @ importsContains @ import syntax referencing a file that doesn't exist-10
R37No stale file referencesMentions files or functions that no longer exist in the repo-10
R38Actionability ratio>60% of content is description rather than instructions-5
--Prerequisites sectionNo section covering required tools, versions, or setup steps-5
R39No rule conflictsCLAUDE.md says X while a .claude/rules/ file says not-X-15

All Artifact Types: Vague Quantifiers
RuleCheckConditionPenalty
R01Vague quantifierEach occurrence of: "appropriate", "relevant", "as needed", "sufficient", "adequate", "reasonable", "properly", "correctly", "some", "several", "various" without measurable criteria-2 each
R01Vague quantifier capTotal vague quantifier penaltymax -20

Mention-versus-use exclusion (added 2026-08-01, origin: xiaolai/cc-suite audit-family false positives; design reviewed via Codex consultation): do not count a vague term when it is presented as a literal token AND the containing clause explicitly instructs the reader or a tool to detect, flag, reject, replace, avoid, or report that term — audit tooling must be able to name the words it hunts (e.g. Flag uses of `some`, `several`, `various` without concrete criteria is R01's own job description, not a violation). Backtick or quotation formatting alone does NOT qualify: a term that still modifies an action, criterion, or requirement is counted even when backticked — handle errors `properly` remains a violation.


All Artifact Types: Vocabulary Drift (R51 — opt-in, disabled by default)

Applied only when R51: { enabled: true, vocabulary_skill: <path> } appears in .claude/nlpm.local.md. Without the opt-in, R51 contributes zero penalty regardless of artifact content. The configured vocabulary_skill must contain a registry.yaml listing canonical and deprecated terms; without it, R51 emits an advisory and contributes zero penalty.

RuleCheckConditionPenalty
R51Deprecated synonymEach occurrence of a term marked deprecated: in the project's registry.yaml, in the scope the artifact belongs to-2 each
R51Drift capTotal R51 penaltymax -10 per file
R51Missing registryenabled: true but vocabulary_skill: not set or points to a directory with no registry.yaml0 (advisory only)

Why opt-in: vocabulary discipline is high-leverage for projects with accumulated drift but premature for projects still discovering their domain. Each project decides when it has enough literary warrant (P6) to lock terms in. See analysis/vocabulary-design-principles.md for the six principles R51 operationalizes.

Registry-declaration exclusion (added 2026-08-01, origin: xiaolai/cc-suite vocabulary-skill self-reference; design reviewed via Codex consultation): the file that declares a deprecation must name the deprecated term to do so. Within the configured vocabulary_skill path, do not count a deprecated term where it occurs in the declaration that registers it or maps it to its replacement — registry.yaml deprecated: lists and the SKILL.md deprecation tables. The exclusion covers ONLY those declaration term fields: deprecated terms in surrounding prose inside the vocabulary_skill path are counted, and the path is not categorically exempt. The same mention-versus-use principle as R01's exclusion above, applied to R51.


Cross-Artifact (--plugin flag)

Applied when linting an entire plugin rather than individual files.

CheckConditionPenalty
Broken partial refsCommand references commands/shared/X.md that doesn't exist-20
Broken skill refsAgent references plugin:skill that isn't installed-20
Missing scriptsHook references script that doesn't exist-20
Orphaned filesAgent/command/skill file not referenced by anything-5 per file
ContradictionsTwo rules/instructions in same plugin directly contradict each other-15 per pair

Score Bands

RangeLabelMeaning
90–100ExcellentProduction-ready; minor or no findings
80–89GoodSolid; one or two non-critical gaps
70–79AdequateMeets threshold; noticeable gaps to address
60–69WeakBelow threshold; significant findings
<60RewriteFundamental problems; recommend rewriting from scratch

Default pass threshold: 70. Configurable in .claude/nlpm.local.md.


Calibration Examples

Four worked examples — Excellent Agent (97), Rewrite Agent (41), Excellent Rule (92), Weak Rule (41) — live in references/calibration-examples.md. Load that file on demand when scoring a borderline case (around band boundaries: 88-92, 68-72, 58-62) and you need an anchored reference.

The examples are not needed for routine scoring — the penalty tables (above, plus the per-tool reference files indexed under Penalty Tables) are self-contained. They were extracted from this file 2026-05-28 to keep the rubric under R05's 500-line body budget while preserving the calibration material verbatim.


Scope Note

This skill covers the NLPM scoring formula, penalty tables, score bands, and calibration examples. It does NOT cover:

  • Artifact schemas and valid field values → see nlpm:conventions (universal) and the tool-specific overlays nlpm:conventions-claude, nlpm:conventions-codex, nlpm:conventions-antigravity (the latter three created in PR-B; see analysis/multi-tool-design-2026-05.md)
  • Patterns and anti-patterns catalog → see nlpm:patterns
  • How to run the score command → see commands/score.md
Multi-tool scoring (PR-B landed 2026-05-25)

nlpm now scores artifacts across three tool ecosystems — Claude Code, Codex CLI, and Antigravity (which absorbs Gemini CLI on 2026-06-18). The tier classification in agents/scorer.md separates open-spec (Tier 1), Tier 1.5 open-spec corpora, and per-tool Tier 2 overlays (2-Claude / 2-Codex / 2-Antigravity).

  • Hooks are scored per tool (the Claude table above; the Codex and Antigravity tables in references/). The three tools' event vocabularies are not 1:1 mappable; no universal translation layer. See analysis/multi-tool-design-2026-05.md decision #4.
  • Codex-specific artifacts scored (references/codex.md): .codex-plugin/plugin.json, .agents/plugins/marketplace.json, .codex/config.toml, agents/openai.yaml sidecars.
  • Antigravity-specific artifacts scored (references/antigravity.md; advisory-only until spec stabilizes): gemini-extension.json, .gemini/commands/*.toml, Antigravity hook events.
  • Claude-specific additions in 2026: .lsp.json, monitors/monitors.json (validate JSON-parse only until detailed schemas land). New SKILL.md fields documented in nlpm:conventions-claude.
Known False Positive Patterns

The following findings have historically been reported by the scorer despite having no backing in this rubric. They MUST NOT be penalized:

Invalid findingWhy it is invalid
Missing namespace: on skillNot in the skill schema; conventions §5 does not list it
Missing inline hooks:/skills: registration blocks in plugin.jsonconventions §1 defines these as optional path strings
AskUserQuestion / Task / WebFetch flagged as undocumented toolBuilt-in per conventions-claude §16
Agent missing skills: when omission is documented in CLAUDE.mdIntentional architectural choice
plugin.json missing engines: / minClaudeVersion: / main:All optional per conventions §1
plugin.json description shorter than sibling marketplace.json descriptionDesynchronization ≠ defect; only penalize if required field is absent

When in doubt: if a finding cannot be cited to a specific row in the penalty tables above or in a reference file indexed under Penalty Tables, drop it.

© xiaolai, ISC. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/nlpm/scoring of xiaolai/nlpm.

  • SKILL.md
  • references/antigravity.md
  • references/calibration-examples.md
  • references/codex.md
  • references/memory-and-workflow.md

Open the folder on GitHubat commit 6fdbd05

Compare with similar skills

Scoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scoring this skillxiaolai/nlpm146—~5.3kAutomated safety check: PassISC
Design AI BenchmarkingAperivue/medsci-skills329—~2.4kAutomated safety check: PassMIT
Evaluator CalibrationArchive228/loopkit756—~1.1kAutomated safety check: PassMIT
Interview System Designerborghei/Claude-Skills881—~1.7kAutomated safety check: PassMIT
Advanced Evaluationguanyang/open-agent-hub9752 repos~4.2kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • Design AI Benchmarking

    Aperivue/medsci-skills

    A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.

    329 GitHub stars~2.4k tokensUpdated 3 days ago
    EducationAuto-check passed
  • Evaluator Calibration

    Archive228/loopkit

    Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.

    756 GitHub stars~1.1k tokensUpdated 2 mo ago
    EducationAuto-check passed
  • Interview System Designer

    borghei/Claude-Skills

    Design calibrated interview loops, competency-based question banks, and hiring calibration.

    881 GitHub stars~1.7k tokensUpdated yesterday
    EducationAuto-check passed
  • Advanced Evaluation

    guanyang/open-agent-hub

    This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    AI & LLM EngineeringAuto-check passed
  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated today
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated yesterday
    EducationAuto-check passed

More from xiaolai/nlpm

All 15 skills in this repo
  • Conventions

    xiaolai/nlpm

    Universal NL conventions: SKILL.md open spec, AGENTS.md, vague quantifiers, naming.

    146 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Antigravity and Gemini CLI artifact schemas: .gemini/ paths, extensions, hooks.

    146 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Conventions Codex

    xiaolai/nlpm

    Codex CLI artifact schemas: config.toml, .codex-plugin, skills, hooks, AGENTS.md.

    146 GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Orchestration

    xiaolai/nlpm

    Multi-agent workflow patterns: parallel dispatch, pipelines, QC gates, retries.

    146 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Patterns

    xiaolai/nlpm

    NL artifact anti-patterns: vague quantifiers, bare prohibitions, oversized skills.

    146 GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Testing

    xiaolai/nlpm

    NL artifact test specs for /nlpm:test: spec format, TDD for skills and agents.

    146 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Categories

Questions about Scoring

What does Scoring do?

100-point NL artifact rubric: penalty tables per artifact type, calibration cases. Scoring is an agent skill from xiaolai/nlpm. 100-point NL artifact rubric: penalty tables per artifact type, calibration cases.

When should I use Scoring?

Scoring fits situations like: tasks that involve Quizzes and assessments; tasks that involve Performance reviews.

How do I install Scoring in Claude Code?

Run `npx skills add xiaolai/nlpm --skill scoring -a claude-code`. Or copy the skill folder (skills/nlpm/scoring in xiaolai/nlpm) into .claude/skills/scoring in your project. Claude Code loads it when a task matches its description.

How do I install Scoring in Codex?

Run `npx skills add xiaolai/nlpm --skill scoring -a codex`. Or copy the skill folder (skills/nlpm/scoring in xiaolai/nlpm) into .agents/skills/scoring in your project. Codex loads it when a task matches its description.

Can I use Scoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xiaolai/nlpm --skill scoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scoring, .gemini/skills/scoring, .github/skills/scoring and .opencode/skills/scoring in your project.

What does Scoring need to run?

Going by SKILL.md and its folder, Scoring needs the command-line tools its instructions call (git).

Does Scoring access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Scoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scoring use?

Scoring is published under the ISC licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scoring use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.

What are the alternatives to Scoring?

Skills that share tags, products or a category with Scoring: Design AI Benchmarking (Aperivue/medsci-skills, 329 stars), Evaluator Calibration (Archive228/loopkit, 756 stars), Interview System Designer (borghei/Claude-Skills, 881 stars) and Advanced Evaluation (guanyang/open-agent-hub, 975 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scoring?

xiaolai (a GitHub user) maintains it in xiaolai/nlpm, which has 146 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.

Source: xiaolai/nlpm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.