Agent skill

Testing

by xiaolai in xiaolai/nlpm

NL artifact test specs for /nlpm:test: spec format, TDD for skills and agents.

ISCAuto-check passedTesting & QA

Install Testing

skills CLI
$ npx skills add xiaolai/nlpm --skill testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xiaolai/nlpm testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xiaolai/nlpm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nlpm/testing .claude/skills/testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing
GitHub stars
146
Token cost
~1.4k tokens
SKILL.md length
337 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
ISC

At a glance

NL artifact test specs for /nlpm:test: spec format, TDD for skills and agents.

  • Works in 3 steps: Looking in .nlpm-test/ directory → Matching spec filename to artifact… → If artifact path doesn't exist → spec is…
  • Tasks that involve Test-driven development
  • SKILL.md covers The NL-TDD Cycle, Spec File Format, Test Results Format and Best Practices for Specs, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Testing is an agent skill from xiaolai/nlpm. NL artifact test specs for /nlpm:test: spec format, TDD for skills and agents.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test-driven development. The repository describes itself as: Natural-Language Programming Manager — scan, lint, and score NL artifacts with Claude-native quality scoring. The licence is ISC.

When your agent uses it

  • Tasks that involve Test-driven development

Example prompts

  • “/testing”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Looking in .nlpm-test/ directory
  2. Matching spec filename to artifact filename: my-agent.spec.md → agents/my-agent.md (uses the artifact: frontmatter field)
  3. If artifact path doesn't exist → spec is RED by default (artifact not yet created — this is the TDD "write test first" state)

What it can do on your machine

Read from SKILL.md and the folder at commit 6fdbd05. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown, yaml and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing loads about 1.4k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 337 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from xiaolai/nlpm at commit 6fdbd05, republished under its ISC licence (© xiaolai). 337 words, ~1,400 tokens.

Download SKILL.mdSave it as .claude/skills/testing/SKILL.md (or your agent's skills folder).
name
testing
description
NL artifact test specs for /nlpm:test: spec format, TDD for skills and agents.
version
0.1.0
user-invocable
false

The NL-TDD Cycle

1. Write spec (.nlpm-test/artifact-name.spec.md)     — define expectations
2. /nlpm:test                                         — RED: spec fails (artifact doesn't exist)
3. Write the artifact                                 — create the NL artifact
4. /nlpm:test                                         — check if it passes
5. /nlpm:score                                        — check quality score
6. Iterate until GREEN: all specs pass + score ≥ threshold

Spec File Format

Location: .nlpm-test/ directory in the project root (or alongside the artifact).

Filename convention: <artifact-name>.spec.md — matches the artifact filename without path.

yaml
---
artifact: agents/my-agent.md          # path to the artifact being tested
type: agent                           # agent | skill | command | rule | hook | prompt
min_score: 85                         # minimum /nlpm:score threshold for this artifact
---

Body sections (all optional — include what matters for this artifact):

triggers_on (skills and agents)
markdown
## Triggers On

Queries that SHOULD trigger this artifact:

- "review my database migrations before deploying"
- "check if these schema changes are safe"
- "audit the migration for breaking changes"
does_not_trigger_on (skills and agents)
markdown
## Does Not Trigger On

Queries that should NOT trigger this artifact:

- "write a migration for adding a users table"
- "help me with CSS styling"
- "deploy to production"
output_contains (agents and commands)
markdown
## Output Contains

Expected elements in the output:

- "## Migration Review" (heading present)
- "| Table | Change | Risk |" (table structure)
- severity classification (CRITICAL/HIGH/MEDIUM/LOW)
output_format (agents and commands)
markdown
## Output Format

The output should be a markdown report with:
1. Summary section with counts
2. Findings table with columns: File, Finding, Severity
3. Action items list
handles_input (commands)
markdown
## Handles Input

| Input | Expected Behavior |
|-------|------------------|
| (empty) | Score all artifacts in cwd |
| directory path | Score artifacts in that directory |
| nonexistent path | Error: "Directory not found: {path}" |
| file path | Score that single file |
follows_rules (rules)
markdown
## Follows Rules

Code that SHOULD comply:
```python
result: Result[User, AppError] = get_user(id)

Code that SHOULD violate:

python
user = get_user(id).unwrap()  # should be flagged

### frontmatter_valid (all types)

```markdown
## Frontmatter Valid

Required fields:
- description: present and trigger-style ("Use when...")
- model: sonnet
- tools: [Read, Glob, Grep]
- skills: [nlpm:conventions, nlpm:scoring]

Test Results Format

NLPM Test Report

Spec                              Artifact                    Result   Details
─────────────────────────────────────────────────────────────────────────────────
my-agent.spec.md                  agents/my-agent.md          PASS     5/5 checks
my-skill.spec.md                  skills/core/SKILL.md        FAIL     3/5 checks
  ✗ Trigger: "optimize React hooks" → predicted NO trigger (expected YES)
  ✗ Score: 68/100 (min: 85)

Overall: 1 passed, 1 failed (50%)

RED items (fix these):
  1. skills/core/SKILL.md — trigger gap: "optimize React hooks" not covered by description
  2. skills/core/SKILL.md — score 68 < min 85: missing <example> blocks (R06, -10)

Best Practices for Specs

  • Write specs BEFORE writing the artifact (TDD discipline)
  • 5-10 trigger queries for skills/agents (mix positive and negative)
  • Include edge cases in handles_input for commands
  • min_score should match your project's threshold (default 85 for new artifacts, 70 for legacy)
  • Specs are living documents — update when behavior requirements change

Spec File Discovery

The tester discovers specs by:

  1. Looking in .nlpm-test/ directory
  2. Matching spec filename to artifact filename: my-agent.spec.md → agents/my-agent.md (uses the artifact: frontmatter field)
  3. If artifact path doesn't exist → spec is RED by default (artifact not yet created — this is the TDD "write test first" state)

Worked Example: TDD Cycle

  1. RED — write the spec first. Create .nlpm-test/my-agent.spec.md with artifact: agents/my-agent.md and a triggers_on: block listing 5 user queries the agent should match. The artifact does not exist yet — /nlpm:test reports RED with one failure: artifact not found.

  2. GREEN — write the artifact. Create agents/my-agent.md with frontmatter (name, description, model, tools) and a body. The description must be specific enough that the listed trigger queries land on it. Re-run /nlpm:test: the agent now exists, the tester predicts triggers, and the spec passes.

  3. REFACTOR — change the body, re-run the spec. Edit the artifact's behavior. The same spec runs unchanged; if a refactor breaks a trigger query or violates a handles_input case, the test goes RED. Fix forward, re-run.

This is the natural-language analogue of pytest or jest — specs are the contract; artifacts are the implementation.

Scope Note

This skill covers the spec format and runner contract for NL-TDD. For the scoring rubric the tester compares against, see nlpm:scoring. For the schemas of artifact frontmatter the spec checks, see nlpm:conventions.

© xiaolai, ISC. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/nlpm/testing of xiaolai/nlpm.

Open the folder on GitHubat commit 6fdbd05

Compare with similar skills

Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing this skillxiaolai/nlpm146—~1.4kAutomated safety check: PassISC
TDDpietheinstrengholt/rssmonster56430 repos~906Automated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT
Test Driven Developmentfarm-fe/farm5.6k51 repos~2.5kAutomated safety check: PassMIT
Tapd Story PipelineTencentBlueKing/bk-bcs840—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • TDD

    pietheinstrengholt/rssmonster

    Test-driven development. An agent skill from pietheinstrengholt/rssmonster.

    564 GitHub starsUsed in 30 repos~906 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • A skill your agent uses when implementing any feature or bugfix, before writing implementation code

    5.6k GitHub starsUsed in 51 repos~2.5k tokens
    Testing & QAAuto-check passed
  • Tapd Story Pipeline

    TencentBlueKing/bk-bcs

    单需求实现流水线——把一个 TAPD 需求从零推进到代码提交。自动串联技术澄清、 开发计划、任务拆分、TDD 实现、架构/安全校验、代码提交六个阶段。

    840 GitHub stars~2.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Absolute Init

    maddhruv/absolute

    One-time setup for absolute: interview how you want it to behave (output style, autonomy, TDD strictness, spec dir, families) + detect the stack once, then write .absolute.config.json (project…

    218 GitHub starsUsed in 1 repo~3k tokens
    Testing & QAAuto-check passed

More from xiaolai/nlpm

All 15 skills in this repo
  • Conventions

    xiaolai/nlpm

    Universal NL conventions: SKILL.md open spec, AGENTS.md, vague quantifiers, naming.

    146 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Antigravity and Gemini CLI artifact schemas: .gemini/ paths, extensions, hooks.

    146 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Conventions Codex

    xiaolai/nlpm

    Codex CLI artifact schemas: config.toml, .codex-plugin, skills, hooks, AGENTS.md.

    146 GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Orchestration

    xiaolai/nlpm

    Multi-agent workflow patterns: parallel dispatch, pipelines, QC gates, retries.

    146 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Patterns

    xiaolai/nlpm

    NL artifact anti-patterns: vague quantifiers, bare prohibitions, oversized skills.

    146 GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Scoring

    xiaolai/nlpm

    100-point NL artifact rubric: penalty tables per artifact type, calibration cases.

    146 GitHub stars~5.3k tokensUpdated today
    Auto-check passed

Categories

Questions about Testing

What does Testing do?

NL artifact test specs for /nlpm:test: spec format, TDD for skills and agents. Testing is an agent skill from xiaolai/nlpm. NL artifact test specs for /nlpm:test: spec format, TDD for skills and agents.

When should I use Testing?

Testing fits situations like: tasks that involve Test-driven development.

How do I install Testing in Claude Code?

Run `npx skills add xiaolai/nlpm --skill testing -a claude-code`. Or copy the skill folder (skills/nlpm/testing in xiaolai/nlpm) into .claude/skills/testing in your project. Claude Code loads it when a task matches its description.

How do I install Testing in Codex?

Run `npx skills add xiaolai/nlpm --skill testing -a codex`. Or copy the skill folder (skills/nlpm/testing in xiaolai/nlpm) into .agents/skills/testing in your project. Codex loads it when a task matches its description.

Can I use Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xiaolai/nlpm --skill testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing, .gemini/skills/testing, .github/skills/testing and .opencode/skills/testing in your project.

What does Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Testing is instructions for the agent only. Our summary lists: Python 3.

Does Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing use?

Testing is published under the ISC licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing?

Skills that share tags, products or a category with Testing: TDD (pietheinstrengholt/rssmonster, 564 stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), TDD (sanity-io/sanity, 6.4k stars) and Test Driven Development (farm-fe/farm, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing?

xiaolai (a GitHub user) maintains it in xiaolai/nlpm, which has 146 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 8, 2026.

Source: xiaolai/nlpm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.