Agent skill

Testability Scoring

by proffesor-for-testing in proffesor-for-testing/agentic-qe

AI-powered testability assessment using 10 principles of intrinsic testability with Playwright and optional Vibium integration.

MITAuto-check passedDevOps & Cloud

Install Testability Scoring

skills CLI
$ npx skills add proffesor-for-testing/agentic-qe --skill testability-scoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install proffesor-for-testing/agentic-qe testability-scoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/proffesor-for-testing/agentic-qe.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assets/skills/testability-scoring .claude/skills/testability-scoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testability-scoring
GitHub stars
494
Token cost
~3.2k tokens
SKILL.md length
699 words
Files
9 (incl. scripts)
Skills in repo
95
Repo updated
First seen
Licence
MIT

At a glance

AI-powered testability assessment using 10 principles of intrinsic testability with Playwright and optional Vibium integration.

  • Works in 5 steps: RUN assessment against target URL → ANALYZE all 10 principles automatically → GENERATE HTML report with radar chart → …
  • Assessing software testability
  • SKILL.md covers Browser engine, Quick Reference Card, Assessment Workflow and Integration Examples, plus 7 more sections
  • Runs JavaScript and Shell scripts from its folder; calls claude, npx and node

What it does

Testability Scoring is an agent skill from proffesor-for-testing/agentic-qe. AI-powered testability assessment using 10 principles of intrinsic testability with Playwright and optional Vibium integration. Evaluates web applications against Observability, Controllability, Algorithmic Simplicity, Transparency, Stability, Explainability, Unbugginess, Smallness, Decomposability, and Similarity. Use when assessing software testability, evaluating test readiness, identifying testability improvements, or generating testability reports.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts (for example `README.md`, `evals/testability-scoring.yaml` and `resources/templates/config.template.js`).

It sits in DevOps & Cloud, covering Observability and Browser testing. It works with Playwright. The repository describes itself as: Agentic QE Fleet is an open-source AI-powered QA/QE platform designed for use with Coding Agents (works best with Claude Code) featuring specialized agents and skills to support… The licence is MIT.

When your agent uses it

  • Assessing software testability
  • Evaluating test readiness
  • Identifying testability improvements
  • Generating testability reports

Example prompts

  • “/testability-scoring”

Requirements

  • Node.js
  • A Bash shell
  • Docker

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. RUN assessment against target URL
  2. ANALYZE all 10 principles automatically
  3. GENERATE HTML report with radar chart
  4. PRIORITIZE improvements by impact/effort
  5. INTEGRATE with QX Partner for holistic view

What it can do on your machine

Read from SKILL.md and the folder at commit 829d030. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (JavaScript and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • claude
    • npx
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • satisfice.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testability Scoring loads about 3.2k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 699 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from proffesor-for-testing/agentic-qe at commit 829d030, republished under its MIT licence (© proffesor-for-testing). 699 words, ~3,188 tokens.

Download SKILL.mdSave it as .claude/skills/testability-scoring/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
testability-scoring
description
AI-powered testability assessment using 10 principles of intrinsic testability with Playwright and optional Vibium integration. Evaluates web applications against Observability, Controllability, Algorithmic Simplicity, Transparency, Stability, Explainability, Unbugginess, Smallness, Decomposability, and Similarity. Use when assessing software testability, evaluating test readiness, identifying testability improvements, or generating testability reports.
category
testing-methodologies
priority
high
tokenEstimate
1100
agents
qe-quality-analyzer, qx-partner, qe-visual-tester
implementation_status
optimized
optimization_version
2.2
last_optimized
2025-12-12
quick_reference_card
true
tags
testability, scoring, playwright, vibium, assessment, 10-principles, intrinsic-testability, james-bach, michael-bolton
contributor
@fndlalit

Testability Scoring

Browser engine

Uses the qe-browser fleet skill (.claude/skills/qe-browser/) as the primary browser engine. Vibium is installed by aqe init. The legacy scripts/run-assessment.sh + Playwright path remains as a fallback when a team already has a Playwright test suite configured, but new runs should prefer:

bash
vibium go "$TARGET_URL"
vibium wait load
vibium a11y-tree --json > /tmp/testability/tree.json
vibium eval --stdin --json <<'EOF' > /tmp/testability/signals.json
JSON.stringify({
  headings: document.querySelectorAll('h1,h2,h3,h4,h5,h6').length,
  testIds: document.querySelectorAll('[data-testid]').length,
  forms: document.querySelectorAll('form').length,
  ariaLabels: document.querySelectorAll('[aria-label]').length,
  links: document.querySelectorAll('a[href]').length,
});
EOF
node .claude/skills/qe-browser/scripts/assert.js --checks '[
  {"kind": "no_console_errors"},
  {"kind": "no_failed_requests"}
]'

<default_to_action> When assessing testability:

  1. RUN assessment against target URL
  2. ANALYZE all 10 principles automatically
  3. GENERATE HTML report with radar chart
  4. PRIORITIZE improvements by impact/effort
  5. INTEGRATE with QX Partner for holistic view

Quick Assessment:

bash
# Run assessment on any URL
TEST_URL='https://example.com/' npx playwright test tests/testability-scoring/testability-scoring.spec.js --project=chromium --workers=1

# Or use shell script wrapper
.claude/skills/testability-scoring/scripts/run-assessment.sh https://example.com/

The 10 Principles at a Glance:

PrincipleWeightKey Question
Observability15%Can we see what's happening?
Controllability15%Can we control the application?
Algorithmic Simplicity10%Are behaviors predictable?
Algorithmic Transparency10%Can we understand what it does?
Algorithmic Stability10%Does behavior remain consistent?
Explainability10%Is the interface understandable?
Unbugginess10%How error-free is it?
Smallness10%Are components appropriately sized?
Decomposability5%Can we test parts in isolation?
Similarity5%Is the tech stack familiar?

Grade Scale:

  • A (90-100): Excellent testability
  • B (80-89): Good testability
  • C (70-79): Adequate testability
  • D (60-69): Below average
  • F (0-59): Poor testability </default_to_action>

Quick Reference Card

Running Assessments
MethodCommandWhen to Use
Shell Script./scripts/run-assessment.sh URLOne-time assessment
ENV OverrideTEST_URL='URL' npx playwright test...CI/CD integration
Config FileUpdate tests/testability-scoring/config.jsRepeated runs
Principle Details
High Weight (15% each)
PrincipleMeasuresIndicators
ObservabilityState visibility, logging, monitoringConsole output, network tracking, error visibility
ControllabilityInput control, state manipulationAPI access, test data injection, determinism
Medium Weight (10% each)
PrincipleMeasuresIndicators
SimplicityPredictable behaviorClear I/O relationships, low complexity
TransparencyUnderstanding what system doesVisible processes, readable code
StabilityConsistent behaviorChange resilience, maintainability
ExplainabilityInterface understandingGood docs, semantic structure, help text
UnbugginessError-free operationConsole errors, warnings, runtime issues
SmallnessComponent sizeElement count, script bloat, page complexity
Low Weight (5% each)
PrincipleMeasuresIndicators
DecomposabilityIsolation testingComponent separation, modular design
SimilarityTechnology familiarityStandard frameworks, known patterns

Assessment Workflow

1. Navigate to URL → 2. Collect Metrics → 3. Score Principles
                                              ↓
4. Generate JSON ← 5. Calculate Grades ← 6. Apply Weights
         ↓
7. Generate HTML Report with Radar Chart
         ↓
8. Open in Browser (auto-opens)
Output Files
tests/reports/
├── testability-results-<timestamp>.json  # Raw data
├── testability-report-<timestamp>.html   # Visual report
└── latest.json                           # Symlink

Integration Examples

CI/CD Integration
yaml
# GitHub Actions
- name: Testability Assessment
  run: |
    timeout 180 .claude/skills/testability-scoring/scripts/run-assessment.sh ${{ env.APP_URL }}

- name: Upload Reports
  uses: actions/upload-artifact@v3
  with:
    name: testability-reports
    path: tests/reports/testability-*.html
QX Partner Integration
typescript
// Combine testability with QX analysis
const qxAnalysis = await Task("QX Analysis", {
  target: 'https://example.com',
  integrateTestability: true
}, "qx-partner");

// Returns combined insights:
// - QX Score: 78/100
// - Testability Integration: Observability 72/100
// - Combined Insight: Low observability may mask UX issues
Programmatic Usage
typescript
import { runTestabilityAssessment } from './testability';

const results = await runTestabilityAssessment('https://example.com');
console.log(`Overall: ${results.overallScore}/100 (${results.grade})`);
console.log('Recommendations:', results.recommendations);

Agent Integration

typescript
// Run testability assessment
const assessment = await Task("Testability Assessment", {
  url: 'https://example.com',
  generateReport: true,
  openBrowser: true
}, "qe-quality-analyzer");

// Use with QX Partner for holistic analysis
const qxReport = await Task("Full QX Analysis", {
  target: 'https://example.com',
  integrateTestability: true,
  detectOracleProblems: true
}, "qx-partner");

Vibium Integration (Optional)

Overview

Vibium browser automation can be used alongside Playwright for enhanced testability assessment. While Playwright remains the primary engine, Vibium offers complementary capabilities for certain metrics.

Installation:

bash
claude mcp add vibium -- npx -y vibium
Vibium-Enhanced Metrics
PrincipleVibium EnhancementBenefit
ObservabilityAuto-wait duration trackingMeasures DOM stability (30s timeout, 100ms polling)
ControllabilityElement interaction success rateValidates automation readiness via MCP
StabilityScreenshot consistencyVisual regression detection for layout stability
ExplainabilityElement attribute extractionARIA labels, semantic HTML validation
Show full SKILL.md (286 more words)Show less
When to Use Vibium

✅ USE Vibium for:

  • Element stability metrics (auto-wait duration analysis)
  • Visual consistency checks (screenshot comparison)
  • MCP-native AI agent integration
  • Lightweight Docker images (400MB vs 1.2GB)

❌ USE Playwright for:

  • Console error detection (Vibium V1 lacks console API)
  • Network performance metrics (BiDi network APIs coming in V2)
  • Comprehensive browser coverage (Firefox, Safari)
  • Production-proven stability (Vibium V1 released Dec 2024)
Hybrid Assessment Example
typescript
// Testability assessment using both engines
const assessment = {
  // Playwright: Comprehensive metrics
  playwright: await runPlaywrightAssessment(url),

  // Vibium: Stability metrics
  vibium: {
    elementStability: await measureAutoWaitDuration(url),
    visualConsistency: await compareScreenshots(url),
    accessibilityAttributes: await extractARIALabels(url)
  }
};

// Enhanced Observability Score
const observability =
  (assessment.playwright.consoleErrors * 0.6) +
  (assessment.vibium.elementStability * 0.4);
Vibium MCP Tools for Testability
typescript
// 1. Element Stability Measurement
const browser = await browser_launch();
await browser_navigate({ url });
const startTime = Date.now();
const element = await browser_find({ selector: ".critical-element" });
const autoWaitDuration = Date.now() - startTime;
// Lower duration = better stability

// 2. Visual Consistency Check
const screenshot1 = await browser_screenshot();
await browser_navigate({ url }); // Reload
const screenshot2 = await browser_screenshot();
const visualDiff = compareImages(screenshot1.png, screenshot2.png);
// Lower diff = better stability

// 3. Accessibility Attribute Extraction
const elements = await browser_find({ selector: "button, a, input" });
const ariaLabels = elements.map(el => el.attributes["aria-label"]);
const semanticScore = (ariaLabels.filter(Boolean).length / elements.length) * 100;
Migration Strategy

Current (V2.2): Hybrid approach

  • Playwright: Primary engine for all 10 principles
  • Vibium: Optional enhancement for stability metrics

Future (V3.0): When Vibium V2 ships

  • Evaluate Vibium as primary engine if:
    • Console/Network APIs available
    • Production stability proven
    • Community adoption increases

Agent Coordination Hints

Memory Namespace
aqe/testability/
├── assessments/*       - Assessment results by URL
├── historical/*        - Historical scores for trend analysis
├── recommendations/*   - Improvement recommendations
├── integration/*       - QX integration data
└── vibium/*           - Vibium-specific metrics (optional)
Fleet Coordination
typescript
const testabilityFleet = await FleetManager.coordinate({
  strategy: 'testability-assessment',
  agents: [
    'qe-quality-analyzer',  // Primary assessment
    'qx-partner',           // UX integration
    'qe-visual-tester'      // Visual validation
  ],
  topology: 'sequential'
});

Common Issues & Solutions

IssueSolution
Tests timing outIncrease timeout: timeout 300 ./scripts/run-assessment.sh URL
Partial resultsCheck console errors, increase network timeout
Report not openingUse AUTO_OPEN=false, open manually
Config not updatingUse TEST_URL env var instead
Vibium not availableInstall via claude mcp add vibium -- npx -y vibium (optional)
Hybrid mode errorsVibium is optional; assessments work without it


Credits & References

Framework Origin
Implementation
Vibium Resources

Remember

Testability is an investment, not an afterthought.

Good testability:

  • Reduces debugging time
  • Enables faster feedback loops
  • Makes defects easier to find
  • Supports continuous testing

Low scores = High risk. Prioritize improvements by weight × impact.

© proffesor-for-testing, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts) in assets/skills/testability-scoring of proffesor-for-testing/agentic-qe.

  • SKILL.md
  • README.md
  • evals/testability-scoring.yaml
  • resources/templates/config.template.js
  • resources/templates/testability-scoring.spec.template.js
  • schemas/output.json
  • scripts/generate-html-report.js
  • scripts/run-assessment.sh
  • scripts/validate-config.json

Open the folder on GitHubat commit 829d030

Compare with similar skills

Testability Scoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testability Scoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testability Scoring this skillproffesor-for-testing/agentic-qe494—~3.2kAutomated safety check: PassMIT
Demo Cnpg Local Web Ngcarverauto/serviceradar921—~672Automated safety check: PassApache-2.0
Fieldsurvey Local Web Ngcarverauto/serviceradar921—~785Automated safety check: PassApache-2.0
Acarshub Tool Additionssdr-enthusiasts/docker-acarshub117—~707Automated safety check: PassGPL-3.0
Playwright CIzebbern/claude-code-guide4.7k2 repos~675Automated safety check: PassMIT
Release Processpadamson/playwright-rust153—~4.1kAutomated safety check: PassApache-2.0

Similar skills

  • Demo Cnpg Local Web Ng

    carverauto/serviceradar

    Run ServiceRadar web-ng locally against the live Kubernetes demo CNPG database for dashboard, SRQL, services, and UI testing.

    921 GitHub stars~672 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Fieldsurvey Local Web Ng

    carverauto/serviceradar

    Run ServiceRadar web-ng locally against the Kubernetes demo namespace FieldSurvey data, including CNPG NodePort access, NATS Object Store artifact access, authenticated browser checks, and…

    921 GitHub stars~785 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Acarshub Tool Additions

    sdr-enthusiasts/docker-acarshub

    Use ONLY when working in the docker-acarshub repository AND a task may require adding a system tool, npm package, or other dependency.

    117 GitHub stars~707 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Playwright CI

    zebbern/claude-code-guide

    Production-ready CI/CD configurations for Playwright — GitHub Actions, GitLab CI, CircleCI, Azure DevOps, Jenkins, Docker, parallel sharding, reporting, code coverage, and global setup/teardown.

    4.7k GitHub starsUsed in 2 repos~675 tokens
    DevOps & CloudAuto-check passed
  • Release Process

    padamson/playwright-rust

    End-to-end release runbook for playwright-rust — version bump, supply-chain refresh, per-crate CHANGELOGs, tag-prefix routing for the three workspace crates, the safer push-then-tag workflow that…

    153 GitHub stars~4.1k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Supercheck Execution Engine

    supercheck-io/supercheck

    Work on Supercheck BullMQ queues, schedulers, capacity management, Playwright or k6 execution, monitors, dynamic locations, cancellation, Redis, Kubernetes Jobs, gVisor, or app-worker execution…

    215 GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed

More from proffesor-for-testing/agentic-qe

All 95 skills in this repo
  • Contract Testing

    proffesor-for-testing/agentic-qe

    Consumer-driven contract testing for microservices using Pact, schema validation, API versioning, and backward compatibility testing.

    494 GitHub stars~1.8k tokensUpdated 3 days ago
    Auto-check passed
  • Mutation Testing

    proffesor-for-testing/agentic-qe

    Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

    494 GitHub stars~1.7k tokensUpdated 3 days ago
    Auto-check passed
  • Performance Testing

    proffesor-for-testing/agentic-qe

    Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates.

    494 GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check passed
  • Code Review Quality

    proffesor-for-testing/agentic-qe

    Conduct context-driven code reviews focusing on quality, testability, and maintainability.

    494 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Security Testing

    proffesor-for-testing/agentic-qe

    Scans for security vulnerabilities including XSS, SQL injection, CSRF, and auth flaws using OWASP Top 10 methodology.

    494 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check: notes
  • Database Testing

    proffesor-for-testing/agentic-qe

    Database schema validation, data integrity testing, migration testing, transaction isolation, and query performance.

    494 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed

Works with

Questions about Testability Scoring

What does Testability Scoring do?

AI-powered testability assessment using 10 principles of intrinsic testability with Playwright and optional Vibium integration. Testability Scoring is an agent skill from proffesor-for-testing/agentic-qe. AI-powered testability assessment using 10 principles of intrinsic testability with Playwright and optional Vibium integration.

When should I use Testability Scoring?

Testability Scoring fits situations like: assessing software testability; evaluating test readiness; identifying testability improvements; generating testability reports.

How do I install Testability Scoring in Claude Code?

Run `npx skills add proffesor-for-testing/agentic-qe --skill testability-scoring -a claude-code`. Or copy the skill folder (assets/skills/testability-scoring in proffesor-for-testing/agentic-qe) into .claude/skills/testability-scoring in your project. Claude Code loads it when a task matches its description.

How do I install Testability Scoring in Codex?

Run `npx skills add proffesor-for-testing/agentic-qe --skill testability-scoring -a codex`. Or copy the skill folder (assets/skills/testability-scoring in proffesor-for-testing/agentic-qe) into .agents/skills/testability-scoring in your project. Codex loads it when a task matches its description.

Can I use Testability Scoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add proffesor-for-testing/agentic-qe --skill testability-scoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testability-scoring, .gemini/skills/testability-scoring, .github/skills/testability-scoring and .opencode/skills/testability-scoring in your project.

What does Testability Scoring need to run?

Going by SKILL.md and its folder, Testability Scoring needs JavaScript and a shell for the scripts in its folder and the command-line tools its instructions call (claude, npx and node). Our summary lists: Node.js; A Bash shell; Docker.

Does Testability Scoring access the network?

SKILL.md names 2 domains. As links in the text: github.com and satisfice.com. This is read from the text; nothing was executed.

Is Testability Scoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Testability Scoring use?

Testability Scoring is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testability Scoring use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testability Scoring?

Skills that share tags, products or a category with Testability Scoring: Demo Cnpg Local Web Ng (carverauto/serviceradar, 921 stars), Fieldsurvey Local Web Ng (carverauto/serviceradar, 921 stars), Acarshub Tool Additions (sdr-enthusiasts/docker-acarshub, 117 stars) and Playwright CI (zebbern/claude-code-guide, 4.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testability Scoring?

proffesor-for-testing (a GitHub user) maintains it in proffesor-for-testing/agentic-qe, which has 494 GitHub stars. The repository holds 95 skills in this directory. The repository was last updated on October 4, 2026.

Source: proffesor-for-testing/agentic-qe on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.