Agent skill

Test Deduplicator

by ArabelaTso in ArabelaTso/Skills-4-SE

Analyzes test suites to identify redundant and duplicate test cases using coverage analysis, semantic similarity, and execution results.

Apache-2.0Auto-check passedTesting & QA

Install Test Deduplicator

skills CLI
$ npx skills add ArabelaTso/Skills-4-SE --skill test-deduplicator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ArabelaTso/Skills-4-SE test-deduplicator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ArabelaTso/Skills-4-SE.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-deduplicator .claude/skills/test-deduplicator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-deduplicator
GitHub stars
253
Token cost
~2.4k tokens
SKILL.md length
723 words
Files
6 (incl. scripts, references, assets)
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyzes test suites to identify redundant and duplicate test cases using coverage analysis, semantic similarity, and execution results.

  • Works in 5 steps: Collect Test Coverage Data → Analyze Semantic Similarity → Generate Deduplication Recommendations → …
  • You need to reduce test suite size
  • SKILL.md covers Overview, Workflow, Analysis Types and Configuration, plus 5 more sections
  • Runs Python scripts from its folder; calls python and pytest

What it does

Test Deduplicator is an agent skill from ArabelaTso/Skills-4-SE. Analyzes test suites to identify redundant and duplicate test cases using coverage analysis, semantic similarity, and execution results. Use this skill when you need to reduce test suite size, identify redundant tests, optimize test execution time, analyze test coverage overlap, find tests with identical behavior, or improve test suite maintainability. Triggers when users ask to deduplicate tests, find redundant test cases, reduce test suite size, identify duplicate tests, or optimize test coverage.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts, reference files and assets (for example `assets/deduplication_config.json`, `references/deduplication_methodology.md` and `scripts/coverage_analyzer.py`).

It sits in Testing & QA, covering Test generation and Test coverage. The repository describes itself as: A curated list of 180+ useful Claude Skills for Software Engineering and resources for customizing AI for SE workflows. The licence is Apache-2.0.

When your agent uses it

  • You need to reduce test suite size
  • Identify redundant tests
  • Optimize test execution time
  • Analyze test coverage overlap

Example prompts

  • “Use the test-deduplicator skill to analyz test suites to identify redundant and duplicate test cases using coverage analysis, semantic similarity…”
  • “/test-deduplicator”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Collect Test Coverage Data
  2. Analyze Semantic Similarity
  3. Generate Deduplication Recommendations
  4. Review Recommendations
  5. Apply Deduplication

What it can do on your machine

Read from SKILL.md and the folder at commit 4f38503. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Deduplicator loads about 2.4k tokens when it runs, and up to ~4.3k if it reads all its reference files. Until then it costs about 131 tokens; SKILL.md has 723 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~131
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ArabelaTso/Skills-4-SE at commit 4f38503, republished under its Apache-2.0 licence (© ArabelaTso). 723 words, ~2,350 tokens.

Download SKILL.mdSave it as .claude/skills/test-deduplicator/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
test-deduplicator
description
Analyzes test suites to identify redundant and duplicate test cases using coverage analysis, semantic similarity, and execution results. Use this skill when you need to reduce test suite size, identify redundant tests, optimize test execution time, analyze test coverage overlap, find tests with identical behavior, or improve test suite maintainability. Triggers when users ask to deduplicate tests, find redundant test cases, reduce test suite size, identify duplicate tests, or optimize test coverage.

Test Deduplicator

Overview

This skill analyzes test suites to identify redundant or duplicate tests by examining code coverage, semantic similarity, and execution behavior. It groups equivalent tests, explains deduplication rationale, and recommends which tests can be safely removed while preserving overall test effectiveness.

Workflow

1. Collect Test Coverage Data

First, gather coverage information for each test in the suite.

For Python projects with pytest:

bash
# Run tests with coverage for each test individually
pytest --cov=src --cov-report=json tests/

# Or use the coverage analyzer script with pre-collected data
python scripts/coverage_analyzer.py coverage_data.json --output coverage_analysis.json

Coverage data format:

json
{
  "test_module::test_function_1": {
    "coverage": {
      "src/module.py": [10, 11, 12, 15, 20],
      "src/utils.py": [5, 6, 7]
    }
  },
  "test_module::test_function_2": {
    "coverage": {
      "src/module.py": [10, 11, 12, 15, 20],
      "src/utils.py": [5, 6, 7]
    }
  }
}
2. Analyze Semantic Similarity

Parse test code to identify semantic similarities:

bash
python scripts/semantic_analyzer.py tests/*.py --output semantic_analysis.json --threshold 0.7

This analyzes:

  • Assertion patterns
  • Function call sequences
  • Code structure similarity
  • Variable usage patterns
3. Generate Deduplication Recommendations

Combine coverage and semantic analysis to produce recommendations:

bash
python scripts/deduplicator.py \
  --coverage coverage_analysis.json \
  --semantic semantic_analysis.json \
  --output recommendations.json \
  --threshold 0.8
4. Review Recommendations

Examine the generated recommendations:

json
{
  "redundant_groups": [
    {
      "type": "identical_coverage",
      "tests": ["test_a", "test_b", "test_c"],
      "reason": "Tests have identical code coverage",
      "confidence": 1.0,
      "keep": "test_a",
      "remove": ["test_b", "test_c"],
      "rationale": "test_a has descriptive name, fast execution"
    }
  ],
  "removal_candidates": [...],
  "coverage_impact": {
    "coverage_preserved": true,
    "tests_removed": 15
  }
}
5. Apply Deduplication

After manual review, remove approved redundant tests and verify coverage is preserved.

Analysis Types

Coverage-Based Analysis

Identifies tests with overlapping or identical coverage patterns.

Redundancy Types:

  • Identical Coverage: Tests covering exactly the same lines
  • Subsumption: One test's coverage is a subset of another
  • High Overlap: Tests with >80% coverage similarity

Script: scripts/coverage_analyzer.py

Output:

  • Identical coverage groups
  • Subsumed tests
  • Coverage similarity matrix
  • Unique coverage contributions
Semantic Analysis

Identifies tests with similar code structure and assertions.

Similarity Metrics:

  • Assertion similarity (40% weight)
  • Function call similarity (30% weight)
  • Source code similarity (30% weight)

Script: scripts/semantic_analyzer.py

Output:

  • Tests with identical assertions
  • Semantically similar test pairs
  • Test pattern groups
Integrated Analysis

Combines multiple signals for comprehensive recommendations.

Redundancy Score:

Redundancy = 0.5 × Coverage_Sim + 0.3 × Semantic_Sim + 0.2 × Execution_Sim

Script: scripts/deduplicator.py

Output:

  • Prioritized removal recommendations
  • Confidence scores
  • Coverage impact analysis
  • Rationale for each recommendation

Configuration

Customize analysis using the configuration template:

bash
cp assets/deduplication_config.json my_config.json
# Edit my_config.json

Key Settings:

Thresholds:

  • identical_coverage: 1.0 (exact match)
  • highly_similar_coverage: 0.9
  • semantic_similarity: 0.7
  • overall_redundancy: 0.8

Prioritization Weights:

  • unique_coverage_lines: 10 (most important)
  • execution_speed_bonus: 5
  • stability_bonus: 3
  • name_clarity_bonus: 2

Safety Checks:

  • preserve_coverage: true (must maintain coverage)
  • require_manual_review: true (human approval needed)
  • min_confidence_for_auto_removal: 0.95

Common Use Cases

Use Case 1: Reduce Test Suite Size
User: "Our test suite has grown to 500 tests and takes too long to run. Find redundant tests."
→ Run coverage analysis on all tests
→ Run semantic analysis on test files
→ Generate deduplication recommendations
→ Review and remove redundant tests
→ Verify coverage is preserved
Use Case 2: Identify Duplicate Tests After Merge
User: "We merged two branches and suspect there are duplicate tests now."
→ Analyze tests from both branches
→ Find tests with identical coverage and assertions
→ Recommend which duplicates to remove
→ Prioritize keeping tests with better names/documentation
Use Case 3: Optimize CI/CD Pipeline
User: "Tests take 30 minutes in CI. Which tests can we safely remove?"
→ Analyze coverage overlap
→ Find subsumed tests (tests fully covered by others)
→ Calculate time savings from removal
→ Generate removal plan with coverage guarantee
Use Case 4: Test Suite Refactoring
User: "Clean up our test suite and remove low-value tests."
→ Identify tests with zero unique coverage contribution
→ Find tests with identical assertions
→ Group semantically similar tests
→ Recommend merging or removing redundant tests

Understanding Results

Redundancy Types in Reports

Identical Coverage (Confidence: 1.0)

  • Tests execute exactly the same code paths
  • Safe to remove all but one
  • Keep the test with the best name/documentation

Subsumed Tests (Confidence: 0.95)

  • One test's coverage is completely contained in another
  • The subsumed test adds no unique value
  • Safe to remove

Highly Similar (Confidence: 0.85)

  • Tests have >80% redundancy score
  • Review manually before removal
  • Consider merging instead of removing

Semantically Identical (Confidence: 0.90)

  • Tests have identical assertions and logic
  • Different only in variable names or formatting
  • Safe to remove duplicates
Prioritization Rationale

Tests are prioritized to keep based on:

  1. Unique Coverage: Tests covering lines no other test covers
  2. Execution Speed: Faster tests preferred
  3. Stability: Tests that consistently pass
  4. Descriptiveness: Clear, well-named tests
  5. Recency: Recently updated tests
Show full SKILL.md (273 more words)Show less

Best Practices

  1. Start Conservative: Use high thresholds (≥0.9) initially

  2. Preserve Coverage: Always verify total coverage remains unchanged

  3. Manual Review: Review recommendations before removing tests

  4. Incremental Removal: Remove tests in small batches, verify after each

  5. Consider Test Intent: Don't remove tests that document important edge cases

  6. Monitor Impact: Track test suite effectiveness after deduplication

Limitations

  1. Parameterized Tests: May appear similar but test different values

  2. Test Intent: Cannot detect if tests serve different documentation purposes

  3. Flaky Tests: Inconsistent coverage may affect analysis

  4. Integration Tests: May intentionally duplicate unit test coverage

  5. Language Support: Currently optimized for Python (extensible to other languages)

Resources

scripts/coverage_analyzer.py

Analyzes test coverage to identify redundant tests:

  • Loads coverage data for each test
  • Calculates coverage similarity (Jaccard index)
  • Finds identical coverage groups
  • Identifies subsumed tests
  • Calculates unique coverage contributions
scripts/semantic_analyzer.py

Analyzes test code for semantic similarity:

  • Parses test files using AST
  • Extracts assertions and function calls
  • Calculates assertion similarity
  • Compares code structure
  • Identifies semantically identical tests
scripts/deduplicator.py

Integrated deduplication engine:

  • Combines coverage and semantic analysis
  • Calculates composite redundancy scores
  • Prioritizes tests to keep
  • Generates actionable recommendations
  • Validates coverage preservation
references/deduplication_methodology.md

Comprehensive methodology guide covering:

  • Redundancy types and detection methods
  • Similarity metrics and formulas
  • Deduplication strategies
  • Prioritization algorithms
  • Best practices and pitfalls
  • Coverage impact analysis
  • False positive mitigation

Read this reference when you need deeper understanding of the methodology, want to customize similarity metrics, or need to explain the approach to stakeholders.

assets/deduplication_config.json

Configuration template for customizing:

  • Analysis thresholds
  • Prioritization weights
  • Filtering rules
  • Safety checks
  • Output formats

Copy and modify this template to create custom configurations for specific projects or test suite characteristics.

© ArabelaTso, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references, assets) in skills/test-deduplicator of ArabelaTso/Skills-4-SE.

  • SKILL.md
  • assets/deduplication_config.json
  • references/deduplication_methodology.md
  • scripts/coverage_analyzer.py
  • scripts/deduplicator.py
  • scripts/semantic_analyzer.py

Open the folder on GitHubat commit 4f38503

Compare with similar skills

Test Deduplicator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Deduplicator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Deduplicator this skillArabelaTso/Skills-4-SE253—~2.4kAutomated safety check: PassApache-2.0
E2E Test ThinkerUniClipboard/UniClipboard1.9k—~1.7kAutomated safety check: PassAGPL-3.0
Ralph Coveragejvm-skills/jvm-skills140—~683Automated safety check: PassApache-2.0
Designing TestsCloudAI-X/opencode-workflow275—~2.9kAutomated safety check: PassMIT
Mutation Testingproffesor-for-testing/agentic-qe494—~1.7kAutomated safety check: PassMIT
Write Testsgnomeria/usbtree690—~622Automated safety check: PassMIT

Similar skills

  • E2E Test Thinker

    UniClipboard/UniClipboard

    Analyze the current branch's diff against main and determine which changes are testable via CLI-based end-to-end tests.

    1.9k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Ralph Coverage

    jvm-skills/jvm-skills

    Run Ralph in coverage mode — iteratively write tests for untested classes until coverage targets are met.

    140 GitHub stars~683 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Mutation Testing

    proffesor-for-testing/agentic-qe

    Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

    494 GitHub stars~1.7k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Write Tests

    gnomeria/usbtree

    Author tests that match the repo's stack and existing test style, at the cheapest level that catches the regression.

    690 GitHub stars~622 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Mutation Test

    jmagly/aiwg

    Run mutation testing to validate test quality beyond code coverage.

    220 GitHub stars~3.2k tokensUpdated 2 days ago
    Testing & QAAuto-check passed

More from ArabelaTso/Skills-4-SE

All 150 skills in this repo
  • Framework Migration Assistant

    ArabelaTso/Skills-4-SE

    Automatically migrate Python web applications between frameworks (Flask → FastAPI, Django → FastAPI).

    253 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Metamorphic Test Generator

    ArabelaTso/Skills-4-SE

    Generate test cases using metamorphic testing by applying transformations based on metamorphic properties.

    253 GitHub stars~798 tokensUpdated 1 mo ago
    Auto-check passed
  • Reproduction Trace Instrumenter

    ArabelaTso/Skills-4-SE

    Instruments programs to capture execution traces specifically for reproducing reported bugs, enabling consistent replay and diagnosis of failures.

    253 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spring Mvc To Boot Migrator

    ArabelaTso/Skills-4-SE

    Automatically migrate Spring MVC applications to Spring Boot.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • State Snapshot Instrumenter

    ArabelaTso/Skills-4-SE

    Instrument programs (Python, C/C++, Java) to capture snapshots of key program states at runtime, including variables, memory, and call stacks.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Test Deduplicator

What does Test Deduplicator do?

Analyzes test suites to identify redundant and duplicate test cases using coverage analysis, semantic similarity, and execution results. Test Deduplicator is an agent skill from ArabelaTso/Skills-4-SE. Analyzes test suites to identify redundant and duplicate test cases using coverage analysis, semantic similarity, and execution results.

When should I use Test Deduplicator?

Test Deduplicator fits situations like: you need to reduce test suite size; identify redundant tests; optimize test execution time; analyze test coverage overlap.

How do I install Test Deduplicator in Claude Code?

Run `npx skills add ArabelaTso/Skills-4-SE --skill test-deduplicator -a claude-code`. Or copy the skill folder (skills/test-deduplicator in ArabelaTso/Skills-4-SE) into .claude/skills/test-deduplicator in your project. Claude Code loads it when a task matches its description.

How do I install Test Deduplicator in Codex?

Run `npx skills add ArabelaTso/Skills-4-SE --skill test-deduplicator -a codex`. Or copy the skill folder (skills/test-deduplicator in ArabelaTso/Skills-4-SE) into .agents/skills/test-deduplicator in your project. Codex loads it when a task matches its description.

Can I use Test Deduplicator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ArabelaTso/Skills-4-SE --skill test-deduplicator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-deduplicator, .gemini/skills/test-deduplicator, .github/skills/test-deduplicator and .opencode/skills/test-deduplicator in your project.

What does Test Deduplicator need to run?

Going by SKILL.md and its folder, Test Deduplicator needs Python for the scripts in its folder and the command-line tools its instructions call (python and pytest). Our summary lists: Python 3.

Does Test Deduplicator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Deduplicator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Test Deduplicator use?

Test Deduplicator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Deduplicator use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Test Deduplicator?

Skills that share tags, products or a category with Test Deduplicator: E2E Test Thinker (UniClipboard/UniClipboard, 1.9k stars), Ralph Coverage (jvm-skills/jvm-skills, 140 stars), Designing Tests (CloudAI-X/opencode-workflow, 275 stars) and Mutation Testing (proffesor-for-testing/agentic-qe, 494 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Deduplicator?

ArabelaTso (a GitHub user) maintains it in ArabelaTso/Skills-4-SE, which has 253 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on August 21, 2026.

Source: ArabelaTso/Skills-4-SE on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.