OpenHarness End-to-End Evals
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
A skill your agent uses when writing tests that hit a real database or broker, setting up Testcontainers for a project, testing HTTP endpoints end-to-end within the service boundary, or implementing…
$ npx skills add kid-sid/claude-spellbook --skill integration-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install kid-sid/claude-spellbook integration-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/integration-testing .claude/skills/integration-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "integration-testing" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/integration-testing into .claude/skills/integration-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integration-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/kid-sid/claude-spellbook/tree/main/skills/integration-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add kid-sid/claude-spellbook --skill integration-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install kid-sid/claude-spellbook integration-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/integration-testing .agents/skills/integration-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "integration-testing" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/integration-testing into .agents/skills/integration-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integration-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kid-sid/claude-spellbook --skill integration-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install kid-sid/claude-spellbook integration-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/integration-testing .cursor/skills/integration-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "integration-testing" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/integration-testing into .cursor/skills/integration-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integration-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/kid-sid/claude-spellbook.git --path skills/integration-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add kid-sid/claude-spellbook --skill integration-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install kid-sid/claude-spellbook integration-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/integration-testing .gemini/skills/integration-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "integration-testing" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/integration-testing into .gemini/skills/integration-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integration-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install kid-sid/claude-spellbook integration-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add kid-sid/claude-spellbook --skill integration-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/integration-testing .github/skills/integration-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "integration-testing" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/integration-testing into .github/skills/integration-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integration-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kid-sid/claude-spellbook --skill integration-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install kid-sid/claude-spellbook integration-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/integration-testing .opencode/skills/integration-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "integration-testing" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/integration-testing into .opencode/skills/integration-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integration-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
integration-testingA skill your agent uses when writing tests that hit a real database or broker, setting up Testcontainers for a project, testing HTTP endpoints end-to-end within the service boundary, or implementing…
Integration Testing is an agent skill from kid-sid/claude-spellbook. Use when writing tests that hit a real database or broker, setting up Testcontainers for a project, testing HTTP endpoints end-to-end within the service boundary, or implementing contract tests between two services.
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Integration testing. It works with Python. The repository describes itself as: A curated collection of skills, prompts, and workflows that extend Claude's capabilities — your personal grimoire for AI-powered development. The licence is MIT.
Read from SKILL.md and the folder at commit a7c2ac9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockerFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Integration Testing loads about 5k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 1,246 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
elopment or production. Keep a separate `.env.test` (or `config/test.yaml`) that is committed to the repo but contains o# .env.test — committed, non-sensitiveDB from dev — never `DATABASE_URL` from `.env`Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from kid-sid/claude-spellbook at commit a7c2ac9, republished under its MIT licence (© kid-sid). 1,246 words, ~5,015 tokens.
.claude/skills/integration-testing/SKILL.md (or your agent's skills folder).Integration tests verify that components work correctly together — real databases, real HTTP routing, real serialization — catching the bugs that unit tests cannot.
Integration tests add real I/O, real wiring, and real serialization. They catch mismatches between your code and the actual database schema, ORM behavior, HTTP middleware ordering, and message envelope formats. What they cost: seconds instead of milliseconds, a Docker dependency, and a higher flakiness risk if not isolated properly.
Testing pyramid starting point: 70% unit / 20% integration / 10% E2E.
| Situation | Use unit test | Use integration test |
|---|---|---|
| Pure function, no I/O | Yes | No |
| DB query logic | Mock is fine for query shape | Yes — real DB for index, constraint, join behavior |
| HTTP handler | Mock dependencies to test logic | Yes — real router for middleware, serialization |
| External API call | Mock HTTP client | Yes — use recorded fixtures or contract test |
| Message handler | Mock broker for handler logic | Yes — real broker for publish/consume wiring |
| Validation rules | Yes — fast feedback | No — unless schema is DB-enforced |
Testcontainers spins up a real Docker container (Postgres, MySQL, Redis, Kafka, etc.) per test suite, giving every developer and CI run an identical, isolated database. No shared test databases, no "works on my machine."
# pip install testcontainers[postgres] psycopg2-binary pytest sqlalchemy
import pytest
from testcontainers.postgres import PostgresContainer
from sqlalchemy import create_engine, text
@pytest.fixture(scope="session")
def postgres():
with PostgresContainer("postgres:16") as pg:
yield pg
@pytest.fixture(scope="session")
def engine(postgres):
engine = create_engine(postgres.get_connection_url())
# Run migrations before the suite
# alembic.config.main(argv=["upgrade", "head"])
return engine
@pytest.fixture
def db(engine):
# BAD: commit data and delete after — leaves residue if test crashes
# conn.execute(text("DELETE FROM users WHERE id = :id"), ...)
# GOOD: wrap in a transaction and roll back — zero cleanup needed
with engine.begin() as conn:
savepoint = conn.begin_nested()
yield conn
savepoint.rollback()// npm install testcontainers pg @types/pg
import { PostgreSqlContainer, StartedPostgreSqlContainer } from "testcontainers";
import { Pool } from "pg";
import { runMigrations } from "../db/migrate";
let container: StartedPostgreSqlContainer;
let pool: Pool;
beforeAll(async () => {
container = await new PostgreSqlContainer("postgres:16").start();
pool = new Pool({ connectionString: container.getConnectionUri() });
await runMigrations(pool); // run prisma migrate / knex migrate:latest
}, 60_000);
afterAll(async () => {
await pool.end();
await container.stop();
});
beforeEach(async () => {
await pool.query("BEGIN");
});
afterEach(async () => {
// GOOD: rollback keeps tests hermetic
await pool.query("ROLLBACK");
});// go get github.com/testcontainers/testcontainers-go/modules/postgres
package db_test
import (
"context"
"testing"
tcpostgres "github.com/testcontainers/testcontainers-go/modules/postgres"
"github.com/testcontainers/testcontainers-go/wait"
)
func TestMain(m *testing.M) {
ctx := context.Background()
container, err := tcpostgres.RunContainer(ctx,
tcpostgres.WithDatabase("testdb"),
tcpostgres.WithUsername("test"),
tcpostgres.WithPassword("test"),
tcpostgres.WithInitScripts("schema.sql"),
testcontainers.WithWaitStrategy(
wait.ForLog("database system is ready to accept connections"),
),
)
if err != nil {
panic(err)
}
defer container.Terminate(ctx)
connStr, _ := container.ConnectionString(ctx, "sslmode=disable")
// pass connStr to your repository layer
m.Run()
}Build test objects with a factory: sensible defaults, every field overridable.
# Python — factory_boy
import factory
from myapp.models import User
class UserFactory(factory.django.DjangoModelFactory):
class Meta:
model = User
email = factory.Sequence(lambda n: f"user{n}@example.com")
name = "Test User"
role = "member"
is_active = True
# In a test:
admin = UserFactory(role="admin")
inactive = UserFactory(is_active=False)// TypeScript — plain builder
function buildUser(overrides: Partial<User> = {}): User {
return {
id: crypto.randomUUID(),
email: `user-${Date.now()}@example.com`,
name: "Test User",
role: "member",
isActive: true,
...overrides,
};
}
const admin = buildUser({ role: "admin" });// Go — functional options
func NewUser(opts ...func(*User)) User {
u := User{
ID: uuid.New(),
Email: fmt.Sprintf("user-%d@example.com", time.Now().UnixNano()),
Name: "Test User",
Role: "member",
IsActive: true,
}
for _, opt := range opts {
opt(&u)
}
return u
}
func WithRole(role string) func(*User) {
return func(u *User) { u.Role = role }
}
admin := NewUser(WithRole("admin"))Always run migrations against the container before tests run — never against a pre-seeded snapshot.
alembic upgrade head pointed at the container URLprisma migrate deploy with DATABASE_URL set to container URLmigrate -path ./migrations -database $DSN upTest the full request/response cycle — routing, middleware, validation, serialization — without going over the network. The goal is to exercise real handler wiring, not mocked HTTP.
# pip install httpx pytest
from fastapi.testclient import TestClient
from myapp.main import app
client = TestClient(app)
def test_user_lifecycle(db): # db fixture provides rolled-back session
# Create
resp = client.post("/users", json={"email": "a@example.com", "name": "Alice"})
assert resp.status_code == 201
user_id = resp.json()["id"]
# Read back
resp = client.get(f"/users/{user_id}")
assert resp.status_code == 200
assert resp.json()["email"] == "a@example.com"
# Delete
resp = client.delete(f"/users/{user_id}")
assert resp.status_code == 204
resp = client.get(f"/users/{user_id}")
assert resp.status_code == 404// npm install supertest @types/supertest
import request from "supertest";
import { buildApp } from "../src/app";
const app = buildApp({ db: testPool });
test("user lifecycle", async () => {
const create = await request(app)
.post("/users")
.send({ email: "a@example.com", name: "Alice" })
.expect(201);
const { id } = create.body;
await request(app).get(`/users/${id}`).expect(200).expect((res) => {
expect(res.body.email).toBe("a@example.com");
});
await request(app).delete(`/users/${id}`).expect(204);
await request(app).get(`/users/${id}`).expect(404);
});import (
"net/http"
"net/http/httptest"
"testing"
"bytes"
"encoding/json"
)
func TestUserLifecycle(t *testing.T) {
handler := buildRouter(testDB)
// Create
body, _ := json.Marshal(map[string]string{"email": "a@example.com", "name": "Alice"})
w := httptest.NewRecorder()
handler.ServeHTTP(w, httptest.NewRequest(http.MethodPost, "/users", bytes.NewReader(body)))
if w.Code != http.StatusCreated { t.Fatalf("expected 201, got %d", w.Code) }
var created map[string]any
json.NewDecoder(w.Body).Decode(&created)
id := created["id"].(string)
// Read back
w = httptest.NewRecorder()
handler.ServeHTTP(w, httptest.NewRequest(http.MethodGet, "/users/"+id, nil))
if w.Code != http.StatusOK { t.Fatalf("expected 200, got %d", w.Code) }
// Delete
w = httptest.NewRecorder()
handler.ServeHTTP(w, httptest.NewRequest(http.MethodDelete, "/users/"+id, nil))
if w.Code != http.StatusNoContent { t.Fatalf("expected 204, got %d", w.Code) }
}Injecting real tokens makes tests slow and fragile. Instead:
# BAD
token = requests.post("/auth/login", json={"password": "..."}).json()["token"]
client.headers["Authorization"] = f"Bearer {token}"
# GOOD — test config bypasses auth middleware, or inject a pre-signed token
client = TestClient(app, headers={"X-Test-User-Id": str(user.id)})
# Middleware reads X-Test-User-Id only when TEST_MODE=trueConsumer-driven contract testing lets two services independently verify the API shape they agree on, without running both services simultaneously.
Consumer side: The consumer writes a Pact — a description of the interactions it expects. Pact runs a mock provider to record the contract.
Provider side: The provider fetches the Pact from the Pact Broker and replays each interaction against its real implementation.
Pact Broker: A shared registry where consumers publish contracts and providers pull them. Teams can see which consumers depend on which provider endpoints.
# Consumer (Python pact-python)
from pact import Consumer, Provider
pact = Consumer("OrderService").has_pact_with(Provider("UserService"))
pact.given("user 42 exists").upon_receiving("a request for user 42").with_request(
"GET", "/users/42"
).will_respond_with(200, body={"id": 42, "name": Like("Alice")})
with pact:
# call your actual client code against pact.uri
user = get_user(42, base_url=pact.uri)
assert user["id"] == 42// Consumer (TypeScript pact-js)
import { PactV3, MatchersV3 } from "@pact-foundation/pact";
const provider = new PactV3({ consumer: "OrderService", provider: "UserService" });
provider
.given("user 42 exists")
.uponReceiving("a request for user 42")
.withRequest({ method: "GET", path: "/users/42" })
.willRespondWith({
status: 200,
body: { id: MatchersV3.integer(42), name: MatchersV3.string("Alice") },
});
await provider.executeTest(async (mockServer) => {
const user = await getUser(42, mockServer.url);
expect(user.id).toBe(42);
});| Scenario | Use Pact | Skip Pact |
|---|---|---|
| Microservices, different teams | Yes | — |
| Consumer and provider deploy independently | Yes | — |
| Monolith with internal module calls | No | Unit/integration test |
| Same team owns both services | Optional — adds overhead | — |
| Internal library (not HTTP) | No | Unit test |
| Third-party external API | Use recorded fixtures instead | — |
Capture published messages in a list and assert on payload shape and schema.
# Python — fake Redis pub/sub with fakeredis
import fakeredis
from myapp.events import publish_order_created
def test_publish_order_created():
r = fakeredis.FakeRedis()
pubsub = r.pubsub()
pubsub.subscribe("orders")
publish_order_created(r, order_id="abc-123", amount=99.99)
message = pubsub.get_message(ignore_subscribe_messages=True, timeout=1)
payload = json.loads(message["data"])
assert payload["order_id"] == "abc-123"
assert payload["event"] == "order.created"// TypeScript — capture calls with a fake queue
const published: unknown[] = [];
const fakeQueue = {
add: (name: string, data: unknown) => {
published.push({ name, data });
return Promise.resolve();
},
};
await handleCheckout(fakeQueue, { orderId: "abc-123", amount: 99.99 });
expect(published).toHaveLength(1);
expect((published[0] as any).name).toBe("order.created");
expect((published[0] as any).data.orderId).toBe("abc-123");// Go — channel-based fake
type FakePublisher struct {
Messages []Event
}
func (f *FakePublisher) Publish(ctx context.Context, e Event) error {
f.Messages = append(f.Messages, e)
return nil
}
func TestPublishOrderCreated(t *testing.T) {
pub := &FakePublisher{}
HandleCheckout(pub, Order{ID: "abc-123", Amount: 99.99})
if len(pub.Messages) != 1 { t.Fatal("expected 1 message") }
if pub.Messages[0].Type != "order.created" { t.Fatal("wrong event type") }
}Inject a raw message into the handler and assert on side effects (DB row created, email sent, etc.).
# Python — Celery with CELERY_TASK_ALWAYS_EAGER
@pytest.fixture(autouse=True)
def eager_celery(settings):
settings.CELERY_TASK_ALWAYS_EAGER = True
settings.CELERY_TASK_EAGER_PROPAGATES = True
def test_consumer_creates_order(db):
send_order_created.delay({"order_id": "abc-123", "amount": 99.99})
order = db.query(Order).filter_by(id="abc-123").one()
assert order.amount == Decimal("99.99")A factory builds valid domain objects with sensible defaults. Every field is overridable. Never write raw SQL inserts in test bodies.
# BAD
cursor.execute("INSERT INTO users (id, email, role) VALUES ('1', 'a@b.com', 'member')")
# GOOD
user = UserFactory(role="admin")| Data type | Strategy |
|---|---|
| Reference data (countries, roles, plans) | Seeder — run once per suite |
| Test-specific domain objects | Factory — per test, rolled back |
| Large static lookup tables | Seeder — loaded from fixture file |
| Data with relationships under test | Factory — build full object graph |
| Strategy | When to use | Notes |
|---|---|---|
| Transaction rollback | Most cases | Fastest; requires single connection per test |
| Truncate after suite | Parallel workers with separate schemas | Slower than rollback |
| Test-specific schema | Parallel workers on shared DB | Drop schema after worker finishes |
| Delete by test ID | Legacy codebases only | Fragile — skip if possible |
Serialize a complex object to a file; fail if it changes. Useful for stable API response shapes.
# Python — syrupy
def test_user_response_shape(snapshot, client):
resp = client.get("/users/1")
assert resp.json() == snapshot # creates __snapshots__/test_users.ambr on first run// TypeScript — jest --updateSnapshot
test("user response shape", async () => {
const resp = await request(app).get("/users/1").expect(200);
expect(resp.body).toMatchSnapshot();
});Never share a test database URL with development or production. Keep a separate .env.test (or config/test.yaml) that is committed to the repo but contains only non-sensitive test-specific values.
# .env.test — committed, non-sensitive
DATABASE_URL=postgres://test:test@localhost:5433/apptest
REDIS_URL=redis://localhost:6380/1
MESSAGE_BROKER_URL=amqp://guest:guest@localhost:5673/| Approach | Speed | Isolation | Use when |
|---|---|---|---|
| Transaction rollback, shared schema | Fastest | Strong (single connection) | Default for most suites |
Schema-per-worker (search_path) | Fast | Good | Parallel pytest-xdist workers |
| Separate DB per worker | Slow to create | Strongest | Long parallel suites with DDL |
# BAD — two workers both insert user with email "admin@example.com"
user = UserFactory(email="admin@example.com")
# GOOD — unique per worker run
user = UserFactory(email=f"admin-{uuid.uuid4()}@example.com")
# Or use factory_boy sequences which are process-local
user = UserFactory() # email = factory.Sequence(lambda n: f"user{n}@example.com")ubuntu-latest does by default).docker pull postgres:16 as a warm-up step or via registry mirror).TESTCONTAINERS_RYUK_DISABLED=true only if your CI runner cannot start the Ryuk reaper container (rootless Docker environments).DELETE FROM users WHERE email LIKE 'test%' for cleanup — string-matching cleanup is fragile and misses rows created by factories with generated emails; use transaction rollback or truncate insteadsession or moduledb.query() is still a unit test; it cannot catch index mismatches, constraint violations, or ORM-generated SQL bugs"admin@example.com") — parallel workers both insert the same unique email and one fails with a constraint violation; use factories with sequences or UUIDsDATABASE_URL from .env© kid-sid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/integration-testing of kid-sid/claude-spellbook.
Open the folder on GitHubat commit a7c2ac9
Integration Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Integration Testing this skillkid-sid/claude-spellbook | 189 | — | ~5k | Automated safety check: Notes | MIT | |
| OpenHarness End-to-End EvalsHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT | |
| Create Flet Control Integration Testsflet-dev/flet | 17k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Godot E2ERandallLiuXin/GodotMaker | 549 | 1 repos | ~3.9k | Automated safety check: Pass | Custom licence | |
| JS-in-HTML Testingliaohch3/claude-tap | 3.3k | — | ~924 | Automated safety check: Pass | MIT | |
| Run Integration Testsvalkey-io/valkey-search | 145 | — | ~511 | Automated safety check: Pass | BSD-3-Clause |
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
flet-dev/flet
A skill your agent uses when asked to create or update integration tests for any Flet control in sdk/python/packages/flet/integrationtests, including visual goldens and interactive behavior tests.
RandallLiuXin/GodotMaker
Write and run E2E (end-to-end) game tests using the godot-e2e framework.
liaohch3/claude-tap
Tests JavaScript embedded in an HTML file in two layers: pytest checks of the logic ported to Python, and Playwright runs in a real browser for the DOM.
valkey-io/valkey-search
Run or troubleshoot Valkey Search C++ and Python integration tests, including focused development runs, full-suite verification, and packaging/build or runtime failure classification.
wshobson/agents
Test Temporal workflows with pytest, time-skipping, and mocking strategies.
kid-sid/claude-spellbook
A skill your agent uses when building or reviewing UI components for keyboard and screen reader compatibility, adding ARIA to custom widgets, auditing a page for WCAG AA conformance, or preparing…
kid-sid/claude-spellbook
A skill your agent uses when building, wiring, or debugging an Agentex agent — choosing agent type, configuring acp.py and manifest.yaml, using adk.messages or adk.state, or resolving…
kid-sid/claude-spellbook
A skill your agent uses when building production LLM applications — designing RAG pipelines, choosing vector databases, implementing agent orchestration, optimizing cost, or adding AI safety…
kid-sid/claude-spellbook
A skill your agent uses when building or refactoring Angular applications — choosing between signals, RxJS, and NgRx for state, configuring routing with guards and lazy loading, optimizing change…
kid-sid/claude-spellbook
A skill your agent uses when designing new REST endpoints, reviewing an existing API contract, adding pagination or filtering, planning a versioning strategy, or building a public or partner-facing…
kid-sid/claude-spellbook
A skill your agent uses when implementing login flows, issuing or validating JWTs, setting up OAuth2/OIDC with a provider, designing role-based or attribute-based access control, securing API…
Works with
Categories
A skill your agent uses when writing tests that hit a real database or broker, setting up Testcontainers for a project, testing HTTP endpoints end-to-end within the service boundary, or implementing…. Integration Testing is an agent skill from kid-sid/claude-spellbook. Use when writing tests that hit a real database or broker, setting up Testcontainers for a project, testing HTTP endpoints end-to-end within the service boundary, or implementing contract tests between two services.
Integration Testing fits situations like: writing tests that hit a real database; setting up Testcontainers for a project; testing HTTP endpoints end-to-end within the service boundary; implementing contract tests between two services.
Run `npx skills add kid-sid/claude-spellbook --skill integration-testing -a claude-code`. Or copy the skill folder (skills/integration-testing in kid-sid/claude-spellbook) into .claude/skills/integration-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add kid-sid/claude-spellbook --skill integration-testing -a codex`. Or copy the skill folder (skills/integration-testing in kid-sid/claude-spellbook) into .agents/skills/integration-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kid-sid/claude-spellbook --skill integration-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/integration-testing, .gemini/skills/integration-testing, .github/skills/integration-testing and .opencode/skills/integration-testing in your project.
Going by SKILL.md and its folder, Integration Testing needs the command-line tools its instructions call (docker). Our summary lists: Python 3; Node.js; Docker.
SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Integration Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Integration Testing: OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars), Create Flet Control Integration Tests (flet-dev/flet, 17k stars), Godot E2E (RandallLiuXin/GodotMaker, 549 stars) and JS-in-HTML Testing (liaohch3/claude-tap, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
kid-sid (a GitHub user) maintains it in kid-sid/claude-spellbook, which has 189 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on August 5, 2026.
Source: kid-sid/claude-spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.