Agent skill

Plugin Test

by NanmiCoder in NanmiCoder/dsh-auto-mode

A skill your agent uses when writing or reviewing tests for DeepSeek Harness plugins, external DSH plugin packages, or package changes in the deepseek-harness repository.

MITAuto-check passedTesting & QA

Install Plugin Test

skills CLI
$ npx skills add NanmiCoder/dsh-auto-mode --skill plugin-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NanmiCoder/dsh-auto-mode plugin-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NanmiCoder/dsh-auto-mode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/plugin-test .claude/skills/plugin-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
plugin-test
GitHub stars
165
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
1,523 words
Files
7 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when writing or reviewing tests for DeepSeek Harness plugins, external DSH plugin packages, or package changes in the deepseek-harness repository.

  • Works in 4 steps: Read… → Add a targeted regression test for every… → Cold-start the exact target version… → …
  • Reviewing tests for DeepSeek Harness plugins
  • SKILL.md covers Test Harness Version Migrations, Run a Docker Release Smoke Test, Test Levels and Select Levels from the Change…, plus 4 more sections
  • Runs JavaScript scripts from its folder; calls pnpm; needs EXA_API_KEY and PERPLEXITY_API_KEY

What it does

Plugin Test is an agent skill from NanmiCoder/dsh-auto-mode. Use when writing or reviewing tests for DeepSeek Harness plugins, external DSH plugin packages, or package changes in the deepseek-harness repository. Select the minimum sufficient levels from unit tests, coverage, real-API end-to-end tests, snapshots, Web tests, real compositions, and built-artifact smoke tests. For Harness version migrations, use this Skill's built-in seven-touchpoint test workflow and validate the exact target-version runtime.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/README.md`, `references/docker-release-smoke.md` and `references/version-migration-testing.md`).

It sits in Testing & QA, covering QA and bug reports, Customer journey mapping and Unit testing. It works with DeepSeek and Docker. The repository describes itself as: Safe automatic permissions for DeepSeek Harness. The licence is MIT.

When your agent uses it

  • Reviewing tests for DeepSeek Harness plugins
  • External DSH plugin packages
  • Package changes in the deepseek-harness repository

Example prompts

  • “/plugin-test”

Requirements

  • Node.js
  • Docker
  • A credential in EXA_API_KEY
  • A credential in PERPLEXITY_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read references/version-migration-testing.md, build a migration ledger for the exact from/to versions, and scan all seven touchpoint…
  2. Add a targeted regression test for every applicable breaking or behavior change. A change-level test proves only that migration mapping…
  3. Cold-start the exact target version through the real product entry point and complete a full user turn. Typechecking, config parsing…
  4. Report unavailable credentials, providers, operating systems, browsers, PTYs, and destructive-migration boundaries honestly. Do not claim…

What it can do on your machine

Read from SKILL.md and the folder at commit 907d663. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • EXA_API_KEY
    • PERPLEXITY_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Plugin Test loads about 2.9k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 1,523 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NanmiCoder/dsh-auto-mode at commit 907d663, republished under its MIT licence (© NanmiCoder). 1,523 words, ~2,860 tokens.

Download SKILL.mdSave it as .claude/skills/plugin-test/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
plugin-test
description
Use when writing or reviewing tests for DeepSeek Harness plugins, external DSH plugin packages, or package changes in the deepseek-harness repository. Select the minimum sufficient levels from unit tests, coverage, real-API end-to-end tests, snapshots, Web tests, real compositions, and built-artifact smoke tests. For Harness version migrations, use this Skill's built-in seven-touchpoint test workflow and validate the exact target-version runtime.

Test DeepSeek Harness Plugins

Select the smallest set of test levels that proves the change is correct. Do not run the full suite by default or repeat checks that have already passed.

Test Harness Version Migrations

When adapting an existing plugin to a new DSH host version:

  1. Read references/version-migration-testing.md, build a migration ledger for the exact from/to versions, and scan all seven touchpoint classes.
  2. Add a targeted regression test for every applicable breaking or behavior change. A change-level test proves only that migration mapping; it does not prove the whole plugin is valid.
  3. Cold-start the exact target version through the real product entry point and complete a full user turn. Typechecking, config parsing, Loader smoke tests, and mock Contexts cannot replace this runtime proof.
  4. Report unavailable credentials, providers, operating systems, browsers, PTYs, and destructive-migration boundaries honestly. Do not claim comprehensive compatibility when any remain unverified.

The commands and repository paths below use test-level names from the official Harness monorepo. For external plugins, use equivalent scripts and paths from their own repositories. Do not add nonexistent Harness root commands or impose the monorepo layout on them.

Run a Docker Release Smoke Test

For a pre-release external plugin check, read references/docker-release-smoke.md and use the included runner on the packaged artifact. Pin one exact DSH version and a non-latest Node image, install the artifact into an isolated Profile, cold-start the real DSH entry, and add one argv-based functional probe when startup alone does not exercise the changed behavior.

Treat the generated JSON and Markdown as narrow evidence for that artifact and target. Report skipped providers, browsers, operating systems, security checks, and additional versions as unverified; a passing Docker smoke test does not expand its own coverage. Keep generated reports, raw logs, temporary Profiles, and Docker cache outside the Skill directory.

Test Levels

LevelCommandWhat it proves
Unit testspnpm run testRuns vitest cases under each package's tests/** and script tests under repository scripts/**/*.spec.ts. Cover edge cases, error paths, event ordering, concurrency races, and contract regressions. Every registry needs an HMR-safety test: dispose the fiber that contributed the registration and assert that the resource is removed.
Coverage gatepnpm run test:coverageIn the Harness monorepo, every file under packages/*/*/src must reach 100% coverage. In an external plugin, follow the coverage threshold declared by its repository. Coverage proves only that lines executed, not that the published feature actually works.
Real-API end-to-end testspnpm run test:e2eValidates behavior against real provider APIs, including DeepSeek models and provider smoke tests gated by their own credentials such as EXA_API_KEY and PERPLEXITY_API_KEY. Each suite skips itself when its credential is absent so credential-free CI remains green.
Snapshot testspnpm run test:snapshotValidates expected credential-free output, pinning transport contracts and presentation while using persisted logs to fix the assembled backend behavior.
Web browser snapshotspnpm run test:webCompares Chromium replay output with apps/web/tests/snapshots/. This is a required Linux PR gate. CI forces read-only DSH_SNAPSHOT=replay and never writes expected output. Record or refresh snapshots only locally and review every difference.

When provider credentials are available and real-API execution is authorized, run the corresponding end-to-end tests. Do not skip them merely to save inference spend. Credential-free tests prove only that the pipeline is connected; only credentialed runs prove that the Agent can work with a real model. Based on the change, cover file-writing prompts, multi-turn sessions, tool use, and cancellation during streaming output. The highest-value smoke test starts a real example, sends one prompt, and verifies the outcome from outside the Agent. Automatic skipping keeps credential-free CI unblocked. Record every skipped provider boundary, and never request, expose, or persist credentials merely to make tests pass.

Select Levels from the Change Surface

  • Pure logic or internal helpers → run unit tests only.
  • New or modified package source → run the repository's coverage gate. Apply the per-file 100% rule only when the Harness monorepo explicitly defines it.
  • Model-visible behavior such as prompts, tool Schemas, tool output, or Skill directories → add a credential-free snapshot to the owning example suite, then add a real-composition test.
  • Protocol-visible behavior such as ACP, JSON-RPC, or wire transport → add a credential-free snapshot to the owning example suite.
  • User-visible behavior such as CLI transcripts, interactive terminals, or GUI flows → in the Harness monorepo, use apps/cli/tests/snapshots/ or apps/web/tests/snapshots/; in an external plugin, use the product-entry test suite owned by that repository.
  • Provider behavior such as a new adapter or real provider feature → run real-API end-to-end tests when credentials are available and execution is authorized.
  • A plugin that users will actually run → execute a non-unit real-composition test as described below. Never test only a manually assembled ctx.plugin(...).

When Snapshot Tests Are Mandatory

Every nontrivial model-visible, protocol-visible, or user-visible change must add or update a credential-free scenario through a runnable composition owned by the repository. Package tests, end-to-end assertions, test-only mock compositions, and PR descriptions cannot replace the complete assembled record. In the Harness monorepo, ACP snapshots live under examples/<name>/tests/snapshots/, headless JSONL snapshots live under examples/headless-agent, terminal flows live under apps/cli/tests/snapshots/, and browser flows live under apps/web/tests/snapshots/. Use the record or refresh command provided by the exact target checkout and review every generated difference. External plugins should use their own snapshot framework. If they have none, add a minimal composition that runs through the real Loader.

Show full SKILL.md (647 more words)Show less

Test the Real Entry Path

  • User-visible plugins need a real-composition test. Start a test-only cordis.yml through the Loader and application or process entry point. Mock only external services or nondeterministic inputs. Assert the model-visible request or log, persisted state, or user-visible output. Do not add test options to published defaults.
  • A guard test is useful only if the regression truly makes it fail. When the exact target contract requires a bundled or composition module without inject to use named exports, add explicit expect('default' in mod).toBe(false) and unwrapExports round-trip assertions. Prove that the test turns red when the regression is introduced and green after the fix. Do not apply this guard to default plugin objects or Service classes supported by the target version. Validate those forms through their actual Loader contract.
  • The "real entry path" means the published artifact. A package bin must run its built entry under native Node to expose issues that tsx may hide, including shutdown races, module resolution, and swallowed load failures. Apply the same rule to runtime entries outside the default index and singleton modules shared across bundle compositions. In the Harness monorepo, keep its built-artifact smoke tests green, such as packages/examples/*/tests/built-bin.e2e.ts and packages/code-runtime/code-runtime-worker/tests/built-lib.e2e.ts when they exist in the target version. In an external plugin, smoke-test its packaged entry. Assert that the process exits nonzero when required configuration is truly missing.
  • In the Harness monorepo, normal tests resolve through the configured source plane, while only explicit built-entry smoke tests consume build artifacts. In an external plugin, follow its resolver but still add one explicit packaged-artifact consumer so source aliases cannot hide a missing export or a second runtime singleton.
  • For subprocess startup in the Harness monorepo, CI and built-test channels must use the shared launcher from the target checkout to run every example or Cordis config subprocess from built lib/. Never hand-write --import tsx for those subprocesses. Protocol and operating-system fixtures that do not load Cordis follow the current Node/TypeScript convention of the target repository. External plugins follow their own launcher but must cover the packaged artifact. Choose src only when source-path resolution itself is the test subject, and document that contract in the test.

Keep Tests Effective

  • Prefer real implementations over mocks. Mock only expensive or nondeterministic boundaries such as LLM adapters, networks, and clocks, while keeping everything downstream real. A handwritten fake proves only that a bridge moved bytes; it does not prove that the published tool delivers its claimed behavior. Tests for bridged tool calls should use a scripted mock model while keeping the tool and executor real.
  • Verify the external world instead of trusting the Agent's report. End-to-end assertions should rerun commands or reread files outside the Agent. Checking only keywords in Agent output lets a cheating Agent pass. Assert that untouched files remain byte-for-byte identical.
  • End-to-end tests must own their resources. Create the Harness inside the test and dispose it in afterEach, including after failure, retry, or timeout. Put shared fixtures in a normal tests/harness.ts, never another *.e2e.ts; importing a test file registers its describe again and duplicates real-API calls.
  • Recovery tests must separate failures before and after each chunk boundary and prove that a failed chunk derives no messages or tool side effects. Cover exhaustion, cancellation, policy composition, persistence, state, wire counts, transport-close idle timeouts, and the production Loader composition.

Commands

In the Harness monorepo, use commands that actually exist in the target checkout, such as pnpm run test, test:coverage, test:e2e, test:snapshot, and test:web. Confirm each script exists instead of assuming a historical command list is still valid. In an external plugin, use its own scripts, package the artifact, install it into an isolated Profile running the exact target version, and execute a cold-start plus core-path smoke test through the product entry. Run the smallest set that covers the changed surface only once. CI proves only the gates it actually defines.

See references/README.md for the reference index.

© NanmiCoder, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/plugin-test of NanmiCoder/dsh-auto-mode.

  • SKILL.md
  • references/README.md
  • references/docker-release-smoke.md
  • references/version-migration-testing.md
  • scripts/container-runner.mjs
  • scripts/docker-release-smoke.mjs
  • scripts/docker-release-smoke.test.mjs

Open the folder on GitHubat commit 907d663

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NanmiCoder/dsh-auto-mode, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Plugin Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Plugin Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Plugin Test this skillNanmiCoder/dsh-auto-mode1651 repos~2.9kAutomated safety check: PassMIT
Weavebench Cua ReproduceAMAP-ML/LongHorizon-Harness1.7k—~1.6kAutomated safety check: PassMIT
Issue WriterNVIDIA/container-canary309—~1.1kAutomated safety check: PassApache-2.0
DeerFlow Smoke Testbytedance/deer-flow83k—~2.5kAutomated safety check: NotesMIT
Codex Plugin QAcode-yeongyu/oh-my-openagent70k—~1.9kAutomated safety check: PassCustom licence
OpenCode QA Toolkitcode-yeongyu/oh-my-openagent70k—~2.9kAutomated safety check: PassCustom licence

Similar skills

  • Weavebench Cua Reproduce

    AMAP-ML/LongHorizon-Harness

    Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.

    1.7k GitHub stars~1.6k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Issue Writer

    NVIDIA/container-canary

    Official

    Draft and revise concise, human-focused GitHub issues for pytest-kind-ng.

    309 GitHub stars~1.1k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • DeerFlow Smoke Test

    bytedance/deer-flow

    Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.

    83k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • Codex Plugin QA

    code-yeongyu/oh-my-openagent

    Tests the omo Codex plugin in an isolated CODEX_HOME with a local mock model, proving hooks fired through app-server notifications without touching ~/.codex.

    70k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed
  • OpenCode QA Toolkit

    code-yeongyu/oh-my-openagent

    Tests the opencode coding agent itself: its CLI, server, plugin hooks and events, the terminal UI under tmux, and its SQLite session database, using tested helper scripts.

    70k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Write Tests

    grafana/synthetic-monitoring-app

    Official

    Write Jest integration and unit tests for the Grafana Synthetic Monitoring app using React Testing Library, MSW, and src/test helpers.

    171 GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed

More from NanmiCoder/dsh-auto-mode

All 10 skills in this repo
  • Dsh Upgrade Audit

    NanmiCoder/dsh-auto-mode

    Audit external compatibility between two DSH (DeepSeek Harness) versions and detect reverts, producing an upgrade-report directory; compares git tags with a source checkout, or published npm…

    165 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Plugin Workflow

    NanmiCoder/dsh-auto-mode

    Coordinate multiple DeepSeek Harness plugin Skills across inspection, migration, runtime debugging, heavy dependencies, testing, naming, and release.

    165 GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • Plugin Write

    NanmiCoder/dsh-auto-mode

    A skill your agent uses when creating a DeepSeek Harness plugin, choosing public names for a new external DSH plugin, validating a dsh-plugin.naming.json manifest, checking reviewed central…

    165 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed
  • Dsh Benchmark Case

    NanmiCoder/dsh-auto-mode

    A skill your agent uses when the user hands over a dsh plugin repository (or a real migration commit / version corridor) and wants its upgrade experience extracted into one auto-graded Harbor…

    165 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check: warnings
  • Plugin Release

    NanmiCoder/dsh-auto-mode

    Package, publish, and distribute DeepSeek Harness (DSH) plugins — npm pack artifact validation, GitHub/npm/hub release-track selection, tarball overrides installs for the unpublished cohort…

    165 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: warnings
  • Plugin Heavy Dep

    NanmiCoder/dsh-auto-mode

    A skill your agent uses when adding a heavyweight browser dependency (diagram/chart renderers like mermaid, code editors, big wasm-adjacent libs) to a lightweight DSH Web plugin that must stay…

    165 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed

Works with

Categories

Questions about Plugin Test

What does Plugin Test do?

A skill your agent uses when writing or reviewing tests for DeepSeek Harness plugins, external DSH plugin packages, or package changes in the deepseek-harness repository. Plugin Test is an agent skill from NanmiCoder/dsh-auto-mode. Use when writing or reviewing tests for DeepSeek Harness plugins, external DSH plugin packages, or package changes in the deepseek-harness repository.

When should I use Plugin Test?

Plugin Test fits situations like: reviewing tests for DeepSeek Harness plugins; external DSH plugin packages; package changes in the deepseek-harness repository.

How do I install Plugin Test in Claude Code?

Run `npx skills add NanmiCoder/dsh-auto-mode --skill plugin-test -a claude-code`. Or copy the skill folder (skills/plugin-test in NanmiCoder/dsh-auto-mode) into .claude/skills/plugin-test in your project. Claude Code loads it when a task matches its description.

How do I install Plugin Test in Codex?

Run `npx skills add NanmiCoder/dsh-auto-mode --skill plugin-test -a codex`. Or copy the skill folder (skills/plugin-test in NanmiCoder/dsh-auto-mode) into .agents/skills/plugin-test in your project. Codex loads it when a task matches its description.

Can I use Plugin Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NanmiCoder/dsh-auto-mode --skill plugin-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/plugin-test, .gemini/skills/plugin-test, .github/skills/plugin-test and .opencode/skills/plugin-test in your project.

What does Plugin Test need to run?

Going by SKILL.md and its folder, Plugin Test needs JavaScript for the scripts in its folder, the command-line tools its instructions call (pnpm) and credentials named EXA_API_KEY and PERPLEXITY_API_KEY. Our summary lists: Node.js; Docker; A credential in EXA_API_KEY; A credential in PERPLEXITY_API_KEY.

Does Plugin Test access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Plugin Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Plugin Test use?

Plugin Test is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Plugin Test use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.

What are the alternatives to Plugin Test?

Skills that share tags, products or a category with Plugin Test: Weavebench Cua Reproduce (AMAP-ML/LongHorizon-Harness, 1.7k stars), Issue Writer (NVIDIA/container-canary, 309 stars), DeerFlow Smoke Test (bytedance/deer-flow, 83k stars) and Codex Plugin QA (code-yeongyu/oh-my-openagent, 70k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Plugin Test?

NanmiCoder (a GitHub user) maintains it in NanmiCoder/dsh-auto-mode, which has 165 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on September 29, 2026.

Source: NanmiCoder/dsh-auto-mode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.