Agent skill

Agent Harness Testing Methodology

by huiliyi37 in huiliyi37/Tianshu-harness

Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type.

Apache-2.0Auto-check: notesTesting & QA

SKILL.md written in Chinese; this summary is our English description.

Install Agent Harness Testing Methodology

skills CLI
$ npx skills add huiliyi37/Tianshu-harness --skill agent-harness-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huiliyi37/Tianshu-harness agent-harness-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huiliyi37/Tianshu-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/skills/optional/agent-harness-testing .claude/skills/agent-harness-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-harness-testing
GitHub stars
1.1k
Token cost
~1k tokens
SKILL.md length
194 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type.

  • Works in 5 steps: 探测测试能力 → 按任务类型选择测试策略 → 探针管理 → …
  • Verifying a bugfix with a red-light test before changing any code
  • SKILL.md covers Stage 1: 探测测试能力, Stage 2: 按任务类型选择测试策略, Stage 3: 探针管理 and Stage 4: 环境模拟, plus 2 more sections
  • Calls npx, go and docker

What it does

Starts in a project the agent has never seen before by probing its capabilities rather than editing code: running a project inspector, reading the right config file per language (package.json for Node, pyproject.toml for Python, go.mod plus test files for Go, Cargo.toml for Rust), listing test files, and trying one test command to see what actually runs. Any failure during that trial, such as a missing dependency or a missing service, is recorded and reported rather than silently worked around.

Task type then picks the strategy: a bugfix must reproduce the failure as a genuine red-light test before any fix, confirming the failing assertion matches the bug description, and only moves to a documented alternative verification when local reproduction is truly impossible, such as a production-only race condition. A new feature needs tests covering the happy path and edge cases plus a typecheck and lint pass; a refactor needs the relevant regression tests and a typecheck; performance work needs a before-and-after benchmark; security work needs tests for authorization bypass, expired tokens and injection.

A third stage manages temporary diagnostic probes across three kinds — throwaway console logs that must be deleted after the fix, structured logs that may be kept for production diagnosis, and assertion probes that get promoted into real test assertions or removed once the fix is confirmed — each marked at insertion time with its location and purpose.

When your agent uses it

  • Verifying a bugfix with a red-light test before changing any code
  • Figuring out what test, lint and typecheck commands an unfamiliar project supports
  • Choosing the right testing strategy for a refactor, feature, or performance change
  • Deciding whether a diagnostic probe should be deleted or kept after a fix

Example prompts

  • “先探测一下这个项目能跑哪些测试命令,再开始修这个 bug。”
  • “Write a red-light test that reproduces this reported crash before we fix it.”
  • “这是个性能优化任务,帮我跑一下改动前后的benchmark对比。”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. 探测测试能力
  2. 按任务类型选择测试策略
  3. 探针管理
  4. 环境模拟
  5. 验证报告

What it can do on your machine

Read from SKILL.md and the folder at commit ce4b60a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • go
    • docker
    • npm
    • python
    • cargo
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, docker and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Harness Testing Methodology loads about 1k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 194 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:66
    如果试跑失败,看错误信息——可能缺依赖(`npm install`)、环境变量(`.env`)、或服务(Docker)。
  • NoteMentions a .env fileSKILL.md:192
    ### 注意 `.env` 和密钥
  • NoteMentions a .env fileSKILL.md:194
    - 不在对话中输出 `.env` 内容

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from huiliyi37/Tianshu-harness at commit ce4b60a, republished under its Apache-2.0 licence (© huiliyi37). 194 words, ~1,047 tokens.

Download SKILL.mdSave it as .claude/skills/agent-harness-testing/SKILL.md (or your agent's skills folder).
name
agent-harness-testing
description
在任意项目中执行开发任务时的测试方法论——包括测试能力探测、RED→GREEN 纪律、探针管理、环境模拟、测试策略选择。当面对 bugfix、功能开发、重构等需要验证的代码改动时使用。

Agent Harness Testing — 通用项目测试方法论

你在别人的项目中工作。你不知道他们的测试框架、不知道他们的构建系统、不知道他们的 CI。但你必须验证你的改动是正确且安全的——不能跳过测试、不能编造结果、不能假设一切正常。

总则(先复现再修复、先红灯再绿灯、结果必须来自真实执行、探针必须清理)已由运行时强制: 证据义务状态机跟踪 RED→GREEN、编辑门拦截无复现的源码修改、探针追踪扫描残留、 交付门禁核验验证证据。本 Skill 是操作手册——怎么探测测试能力、怎么构造红灯、 怎么模拟环境,不再复述总则。


Stage 1: 探测测试能力

进入陌生项目的第一件事——不是改代码,是了解怎么验证。

1.1 运行 inspect_project

这会返回项目摘要:语言、包管理器、scripts、入口文件、测试框架提示。先用它建立全局认知。

1.2 读项目配置文件

根据语言读对应配置:

语言读什么找什么
Node/TSpackage.jsonscripts.test, scripts.typecheck, scripts.lint, devDependencies 中的 vitest/jest/mocha
Pythonpyproject.toml / setup.cfg[tool.pytest], [tool.mypy], [tool.ruff]
Gogo.mod + 项目根 *_test.go测试文件存在即表示 go test ./... 可用
RustCargo.toml[dev-dependencies] 中的 test 相关 crate
1.3 列出测试文件
bash
# Node
ls **/*.test.ts **/*.spec.ts **/__tests__/*.ts 2>/dev/null
# Python
ls **/test_*.py **/*_test.py 2>/dev/null
# Go
ls **/*_test.go 2>/dev/null
# Rust
ls tests/ 2>/dev/null
1.4 试跑一条测试命令
bash
# Node 常见
npx vitest --run 2>&1 | head -5
npx jest --passWithNoTests 2>&1 | head -5

# Python
python -m pytest --co 2>&1 | head -5

# Go
go test ./... 2>&1 | head -5

# Rust
cargo test 2>&1 | head -5

如果试跑失败,看错误信息——可能缺依赖(npm install)、环境变量(.env)、或服务(Docker)。 不要跳过——记录障碍并告知用户,问是否需要帮助配置。

1.5 生成能力地图

探测完成后,心里形成一张表:

typecheck: 可用 (npx tsc --noEmit) / 不可用
lint:      可用 (npx eslint) / 不可用
unit test: 可用 (npx vitest --run) / 不可用
e2e:       可用 (npx playwright test) / 不可用
build:     可用 (npm run build) / 不可用
env sim:   可用 (docker compose up) / 不可用

Stage 2: 按任务类型选择测试策略

Bugfix
必须: RED 红灯测试 → 修复 → GREEN 绿灯测试 → 回归测试
如果无法写红灯测试: 说明原因 + 给替代验证方式

红灯测试构造方法:

  1. 从用户描述和错误日志提取失败场景
  2. 找现有测试文件,复制结构
  3. 写最简失败用例(最小数据、最少依赖)
  4. 运行 → 必须失败
  5. 确认失败断言与 Bug 描述一致
  6. 开始修复

如果无法构造(问题仅在生产环境/第三方回调/并发竞态复现):

  • 明确说明:为什么本地无法复现
  • 给出替代验证方式:staging 环境回放、日志对比、代码审查要点
  • 不跳过验证——只是换一种验证方式
Feature
必须: 新功能测试(覆盖 happy path + 边界)→ typecheck → lint
推荐: 集成测试(如果涉及多模块)
Refactor
必须: 相关回归测试 + typecheck
如果是缓存/不变量/前缀结构: 全量模块测试
Performance
必须: benchmark 对比(改动前后)
推荐: 压力测试、profile 数据
Security
必须: 安全测试(越权、过期令牌、注入)
推荐: staging smoke test

Stage 3: 探针管理

探针是临时诊断工具,不是永久代码。

三类探针
类型写法生命周期示例
临时日志console.log("[probe:name]", data)修复后必须删除console.log("[probe:filter]", candidates)
结构化日志logger.info({ event: "name", ... })可保留(用于线上诊断)logger.info({ event: "draw.select", id, stock })
断言探针assert(cond, "msg")修复确认后转为测试断言或删除assert(stock >= 0, "stock must not be negative")
探针纪律
  1. 插入前:在注释或 commit message 中标记位置和目的
  2. 使用中:保持探针干净——只输出必要字段,不打印整个对象
  3. 清理时:必须逐条检查 console.log / debugger / 临时 assert 是否残留
  4. 任务完成标记前,确认无临时探针残留。有残留 = 任务未完成。

Stage 4: 环境模拟

优先使用真实依赖而不是全 mock。

检查项目是否有 Docker 环境
bash
ls docker-compose.yml docker-compose.yaml Dockerfile Makefile 2>/dev/null

如果有 docker-compose.yml:

bash
# 启动依赖
docker compose up -d db redis
# 跑集成测试
npm run test:integration
# 关闭
docker compose down
如果只有 Makefile
bash
# 找 service/test 相关目标
grep -E '^(test|db|redis|service|up|down):' Makefile
如果什么都没有
  • 用 SQLite 文件做数据库测试(临时文件,测试完删除)
  • 用 node --experimental-test-runner 做最轻量测试
  • 说明:当前项目没有类生产环境,集成测试标记为"mock 验证"
注意 .env 和密钥
  • 不在对话中输出 .env 内容
  • 如果需要环境变量,让用户补充
  • 不在测试代码中硬编码密钥

Stage 5: 验证报告

任务完成时必须输出结构化验证报告,而非"已完成"。

最小报告模板
## 验证报告

### 改动
- 文件1: 改了什么
- 文件2: 改了什么

### 测试结果
- [PASS] 目标测试 (command)
- [PASS] typecheck (command)
- [SKIP] e2e (原因: 项目未配置)

### 未验证项
- 项目无 e2e 配置,手动验收路径: ...

### 风险
- 并发场景下的行为未验证

诚实报告(未跑=未验证、0 passed ≠ 通过、失败附错误信息)与反模式清单由 运行时诚实门禁 + 交付契约强制,不在此复述。


快速检查清单

任务完成前自问:

□ 我读了相关代码和测试吗?
□ Bugfix: 我构造了红灯测试(或说明了无法复现的原因)吗?
□ 我实际运行了测试并看了输出吗?
□ 测试结果能支撑"已验证"的结论吗?
□ typecheck/lint/build 通过了吗?
□ 临时探针清理了吗?
□ 我是否修改了无关文件?
□ 如果是高风险改动,我做了额外验证吗?
□ 我的验证报告是否诚实(不夸大、不推测)?

这 9 个问题全部能答"是",任务才算完成。

© huiliyi37, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/skills/optional/agent-harness-testing of huiliyi37/Tianshu-harness.

Open the folder on GitHubat commit ce4b60a

Compare with similar skills

Agent Harness Testing Methodology next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Harness Testing Methodology compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Harness Testing Methodology this skillhuiliyi37/Tianshu-harness1.1k—~1kAutomated safety check: NotesApache-2.0
Designing TestsCloudAI-X/claude-workflow-v21.4k1 repos~1.5kAutomated safety check: PassMIT
Testing Patternssoftspark/ai-toolkit179—~1.6kAutomated safety check: PassApache-2.0
MoAI TDD Workflowmodu-ai/moai-adk1.2k—~3.1kAutomated safety check: PassApache-2.0
Test-Driven Development Enforcerzereight/gitlab-mcp2k1 repos~904Automated safety check: PassMIT
Python Testing Patternsjh941213/my-cc-harness12617 repos~5.4kAutomated safety check: PassNone

Similar skills

  • Designing Tests

    CloudAI-X/claude-workflow-v2

    Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.

    1.4k GitHub starsUsed in 1 repo~1.5k tokens
    Testing & QAAuto-check passed
  • Testing Patterns

    softspark/ai-toolkit

    Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.

    179 GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • MoAI TDD Workflow

    modu-ai/moai-adk

    Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.

    1.2k GitHub stars~3.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Enforces strict red-green-refactor, with a failing test first, the minimum code to pass it, then cleanup, and a quick reference for common test runners.

    2k GitHub starsUsed in 1 repo~904 tokens
    Testing & QAAuto-check passed
  • Python Testing Patterns

    jh941213/my-cc-harness

    Implement comprehensive testing strategies with pytest, fixtures, mocking, and test-driven development.

    126 GitHub starsUsed in 17 repos~5.4k tokens
    Testing & QAAuto-check passed
  • TDD Guide

    alirezarezvani/claude-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from huiliyi37/Tianshu-harness

  • Visual Acceptance Review

    huiliyi37/Tianshu-harness

    Chinese-language final-check method for UI visual changes before delivery, using multi-theme screenshot matrices, pixel-level evidence and a CSS cascade checklist.

    1.1k GitHub stars~601 tokensUpdated today
    Auto-check passed
  • Frontend Prototype Workflow

    huiliyi37/Tianshu-harness

    Asks one clarifying question at a time, then builds two or three single-file HTML prototypes checked across phone, tablet, and desktop widths.

    1.1k GitHub stars~415 tokensUpdated today
    Auto-check passed
  • Word Document Generator

    huiliyi37/Tianshu-harness

    Generates and reads Word (.docx) documents such as reports, contracts and bids from heading, paragraph, table, code and list blocks, with a centered title.

    1.1k GitHub stars~252 tokensUpdated today
    Auto-check passed
  • Office Excel

    huiliyi37/Tianshu-harness

    Excel 电子表格读写与编辑最佳实践 — 用 xlsxread/xlsxwrite/xlsxedit 处理 .xlsx 时遵守的公式、数字格式与可维护性纪律

    1.1k GitHub stars~340 tokensUpdated today
    Auto-check passed
  • Office PDF Generator and Reader

    huiliyi37/Tianshu-harness

    Generates formal PDFs from structured content blocks with automatic CJK font handling and page numbers, and reads text back out of existing PDFs.

    1.1k GitHub stars~334 tokensUpdated today
    Auto-check passed
  • Office PPT Design Rules

    huiliyi37/Tianshu-harness

    Sets design rules for building PowerPoint decks with the pptx_create and pptx_read tools: a dominant color, varied layouts, readable type and a read-back check.

    1.1k GitHub stars~401 tokensUpdated today
    Auto-check passed

Questions about Agent Harness Testing Methodology

What does Agent Harness Testing Methodology do?

Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type. toml for Rust), listing test files, and trying one test command to see what actually runs. Any failure during that trial, such as a missing dependency or a missing service, is recorded and reported rather than silently worked around.

When should I use Agent Harness Testing Methodology?

Agent Harness Testing Methodology fits situations like: verifying a bugfix with a red-light test before changing any code; figuring out what test, lint and typecheck commands an unfamiliar project supports; choosing the right testing strategy for a refactor, feature, or performance change; deciding whether a diagnostic probe should be deleted or kept after a fix.

How do I install Agent Harness Testing Methodology in Claude Code?

Run `npx skills add huiliyi37/Tianshu-harness --skill agent-harness-testing -a claude-code`. Or copy the skill folder (docs/skills/optional/agent-harness-testing in huiliyi37/Tianshu-harness) into .claude/skills/agent-harness-testing in your project. Claude Code loads it when a task matches its description.

How do I install Agent Harness Testing Methodology in Codex?

Run `npx skills add huiliyi37/Tianshu-harness --skill agent-harness-testing -a codex`. Or copy the skill folder (docs/skills/optional/agent-harness-testing in huiliyi37/Tianshu-harness) into .agents/skills/agent-harness-testing in your project. Codex loads it when a task matches its description.

Can I use Agent Harness Testing Methodology in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huiliyi37/Tianshu-harness --skill agent-harness-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-harness-testing, .gemini/skills/agent-harness-testing, .github/skills/agent-harness-testing and .opencode/skills/agent-harness-testing in your project.

What does Agent Harness Testing Methodology need to run?

Going by SKILL.md and its folder, Agent Harness Testing Methodology needs the command-line tools its instructions call (npx, go, docker, npm, python and cargo).

Does Agent Harness Testing Methodology access the network?

SKILL.md contains no URLs. Its commands use npx, docker and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Agent Harness Testing Methodology safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Agent Harness Testing Methodology use?

Agent Harness Testing Methodology is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Harness Testing Methodology use?

About 1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Harness Testing Methodology?

Skills that share tags, products or a category with Agent Harness Testing Methodology: Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars), Testing Patterns (softspark/ai-toolkit, 179 stars), MoAI TDD Workflow (modu-ai/moai-adk, 1.2k stars) and Test-Driven Development Enforcer (zereight/gitlab-mcp, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Harness Testing Methodology?

huiliyi37 (a GitHub user) maintains it in huiliyi37/Tianshu-harness, which has 1,080 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: huiliyi37/Tianshu-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.