Agent skill

Behavior Preservation Checker

by ArabelaTso in ArabelaTso/Skills-4-SE

Compare runtime behavior between original and migrated repositories to detect behavioral differences, regressions, and semantic changes.

Apache-2.0Auto-check passedDevelopment

Install Behavior Preservation Checker

skills CLI
$ npx skills add ArabelaTso/Skills-4-SE --skill behavior-preservation-checker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ArabelaTso/Skills-4-SE behavior-preservation-checker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ArabelaTso/Skills-4-SE.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/behavior-preservation-checker .claude/skills/behavior-preservation-checker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
behavior-preservation-checker
GitHub stars
253
Token cost
~2.3k tokens
SKILL.md length
618 words
Files
7 (incl. scripts, references)
Skills in repo
170
Repo updated
First seen
Licence
Apache-2.0

At a glance

Compare runtime behavior between original and migrated repositories to detect behavioral differences, regressions, and semantic changes.

  • Works in 4 steps: Setup Repositories → Run Behavior Comparison → Review Results → …
  • Validating code migrations
  • SKILL.md covers Overview, Core Workflow, Comparison Methods and Difference Detection, plus 4 more sections
  • Runs Python scripts from its folder; calls python, git and pytest

What it does

Behavior Preservation Checker is an agent skill from ArabelaTso/Skills-4-SE. Compare runtime behavior between original and migrated repositories to detect behavioral differences, regressions, and semantic changes. Use when validating code migrations, refactorings, language ports, framework upgrades, or any transformation that should preserve behavior. Automatically compares test results, execution traces, API responses, and observable outputs between two repository versions. Provides actionable guidance for fixing deviations and ensuring behavioral equivalence.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/comparison_techniques.md`, `references/difference_patterns.md` and `scripts/behavior_checker.py`).

It sits in Development, covering Refactoring and Code migrations. The repository describes itself as: A curated list of 180+ useful Claude Skills for Software Engineering and resources for customizing AI for SE workflows. The licence is Apache-2.0.

When your agent uses it

  • Validating code migrations
  • Framework upgrades
  • Any transformation that should preserve behavior

Example prompts

  • “/behavior-preservation-checker”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Setup Repositories
  2. Run Behavior Comparison
  3. Review Results
  4. Fix Deviations

What it can do on your machine

Read from SKILL.md and the folder at commit 4f38503. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • git
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Behavior Preservation Checker loads about 2.3k tokens when it runs, and up to ~7.7k if it reads all its reference files. Until then it costs about 130 tokens; SKILL.md has 618 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ArabelaTso/Skills-4-SE at commit 4f38503, republished under its Apache-2.0 licence (© ArabelaTso). 618 words, ~2,304 tokens.

Download SKILL.mdSave it as .claude/skills/behavior-preservation-checker/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
behavior-preservation-checker
description
Compare runtime behavior between original and migrated repositories to detect behavioral differences, regressions, and semantic changes. Use when validating code migrations, refactorings, language ports, framework upgrades, or any transformation that should preserve behavior. Automatically compares test results, execution traces, API responses, and observable outputs between two repository versions. Provides actionable guidance for fixing deviations and ensuring behavioral equivalence.

Behavior Preservation Checker

Overview

Validate that a migrated or refactored codebase preserves the original behavior by automatically comparing runtime behavior, test results, execution traces, and observable outputs between two repository versions.

Core Workflow

1. Setup Repositories

Prepare both repositories for comparison:

bash
# Clone or locate repositories
ORIGINAL_REPO=/path/to/original
MIGRATED_REPO=/path/to/migrated

# Ensure both are at comparable states
cd $ORIGINAL_REPO && git checkout main
cd $MIGRATED_REPO && git checkout main
2. Run Behavior Comparison

Use the comparison script to analyze behavioral differences:

bash
python scripts/behavior_checker.py \
    --original $ORIGINAL_REPO \
    --migrated $MIGRATED_REPO \
    --output behavior_report.json
3. Review Results

Examine the generated report for:

  • Test result differences
  • Execution trace divergences
  • Output mismatches
  • Performance regressions
  • API contract violations
4. Fix Deviations

Follow actionable guidance to resolve behavioral differences.

Comparison Methods

Method 1: Test-Based Comparison

Run the same test suite on both repositories and compare results:

Workflow:

  1. Identify common test suite (or create equivalent tests)
  2. Run tests on original repository
  3. Run tests on migrated repository
  4. Compare pass/fail status, assertions, and outputs

Example:

bash
# Run on original
cd $ORIGINAL_REPO
pytest tests/ --json-report --json-report-file=original_results.json

# Run on migrated
cd $MIGRATED_REPO
pytest tests/ --json-report --json-report-file=migrated_results.json

# Compare
python scripts/compare_test_results.py \
    original_results.json \
    migrated_results.json
Method 2: Execution Trace Comparison

Capture and compare execution traces:

Workflow:

  1. Instrument code to capture function calls, arguments, and return values
  2. Run identical inputs through both versions
  3. Compare execution traces for divergences

Example:

python
# Trace original
python scripts/trace_execution.py \
    --repo $ORIGINAL_REPO \
    --input test_inputs.json \
    --output original_trace.json

# Trace migrated
python scripts/trace_execution.py \
    --repo $MIGRATED_REPO \
    --input test_inputs.json \
    --output migrated_trace.json

# Compare traces
python scripts/compare_traces.py \
    original_trace.json \
    migrated_trace.json
Method 3: Observable Output Comparison

Compare program outputs for identical inputs:

Workflow:

  1. Define test inputs (API requests, CLI commands, function calls)
  2. Capture outputs from both versions (stdout, files, API responses)
  3. Compare outputs for differences

Example:

bash
# Test API endpoints
python scripts/compare_api_outputs.py \
    --original-url http://localhost:8000 \
    --migrated-url http://localhost:8001 \
    --test-cases api_test_cases.json
Method 4: Property-Based Testing

Use property-based testing to find behavioral differences:

Workflow:

  1. Define behavioral properties (invariants, contracts)
  2. Generate random inputs
  3. Verify properties hold for both versions
  4. Report any property violations

Example:

python
# Property: sorting should produce same result
from hypothesis import given, strategies as st

@given(st.lists(st.integers()))
def test_sort_equivalence(input_list):
    original_result = original_sort(input_list)
    migrated_result = migrated_sort(input_list)
    assert original_result == migrated_result

Difference Detection

Test Result Differences

What to check:

  • Tests that pass in original but fail in migrated
  • Tests that fail in original but pass in migrated
  • New test failures
  • Changed assertion messages

Severity levels:

  • Critical: Core functionality tests fail
  • High: Integration tests fail
  • Medium: Edge case tests fail
  • Low: Flaky tests or timing-dependent failures
Execution Trace Differences

What to check:

  • Different function call sequences
  • Different argument values
  • Different return values
  • Missing or extra function calls

Example divergence:

Original trace:
  calculate(x=10) -> 20
  validate(20) -> True
  save(20) -> Success

Migrated trace:
  calculate(x=10) -> 21  # ← Difference!
  validate(21) -> True
  save(21) -> Success
Output Differences

What to check:

  • Different stdout/stderr
  • Different file contents
  • Different API response bodies
  • Different status codes
  • Different error messages

Tolerance levels:

python
# Exact match required
assert original_output == migrated_output

# Numerical tolerance
assert abs(original_value - migrated_value) < 0.001

# Structural equivalence (ignore formatting)
assert json.loads(original) == json.loads(migrated)

Actionable Guidance

Pattern 1: Logic Error

Symptom: Different outputs for same inputs

Diagnosis:

bash
python scripts/isolate_difference.py \
    --original $ORIGINAL_REPO \
    --migrated $MIGRATED_REPO \
    --failing-test test_calculation

Guidance:

  1. Identify the diverging function
  2. Compare implementations side-by-side
  3. Check for off-by-one errors, operator changes, or logic inversions
  4. Add unit test for the specific case
Show full SKILL.md (238 more words)Show less
Pattern 2: Missing Functionality

Symptom: Tests pass in original but fail in migrated with "not implemented" or "attribute error"

Diagnosis:

bash
python scripts/find_missing_functions.py \
    --original $ORIGINAL_REPO \
    --migrated $MIGRATED_REPO

Guidance:

  1. List all missing functions/methods
  2. Implement missing functionality
  3. Verify with targeted tests
Pattern 3: API Contract Violation

Symptom: Different response structure or status codes

Diagnosis:

bash
python scripts/compare_api_contracts.py \
    --original-spec openapi_original.yaml \
    --migrated-spec openapi_migrated.yaml

Guidance:

  1. Document API contract differences
  2. Update migrated API to match original contract
  3. Add contract tests to prevent future violations
Pattern 4: Performance Regression

Symptom: Migrated version is significantly slower

Diagnosis:

bash
python scripts/benchmark_comparison.py \
    --original $ORIGINAL_REPO \
    --migrated $MIGRATED_REPO \
    --iterations 100

Guidance:

  1. Profile both versions to identify bottlenecks
  2. Check for algorithmic changes (O(n) → O(n²))
  3. Look for missing optimizations or caching
  4. Verify database query efficiency
Pattern 5: State Management Issues

Symptom: Tests fail intermittently or depend on execution order

Diagnosis:

bash
python scripts/detect_state_issues.py \
    --repo $MIGRATED_REPO \
    --test-suite tests/

Guidance:

  1. Identify shared state between tests
  2. Add proper setup/teardown
  3. Ensure test isolation
  4. Check for global variable usage

Report Format

The behavior checker generates a comprehensive JSON report:

json
{
  "summary": {
    "total_tests": 150,
    "passed_both": 140,
    "failed_both": 2,
    "passed_original_failed_migrated": 5,
    "failed_original_passed_migrated": 3,
    "behavioral_equivalence": "92.7%"
  },
  "differences": [
    {
      "type": "test_failure",
      "test_name": "test_user_authentication",
      "severity": "critical",
      "original_result": "passed",
      "migrated_result": "failed",
      "error_message": "AssertionError: Expected 200, got 401",
      "guidance": "Check authentication logic in migrated version",
      "affected_files": ["auth/login.py"]
    }
  ],
  "recommendations": [
    "Fix 5 critical test failures before deployment",
    "Review 3 output differences for correctness"
  ]
}

Best Practices

  1. Start with tests: Ensure comprehensive test coverage before migration
  2. Incremental validation: Check behavior after each migration step
  3. Document intentional changes: Mark expected behavioral differences
  4. Use multiple comparison methods: Combine tests, traces, and outputs
  5. Automate the process: Integrate into CI/CD pipeline
  6. Set tolerance thresholds: Define acceptable differences (e.g., timing, formatting)

Resources

  • references/comparison_techniques.md: Detailed comparison methodologies
  • references/difference_patterns.md: Common behavioral difference patterns
  • scripts/behavior_checker.py: Main comparison orchestrator
  • scripts/compare_test_results.py: Test result comparison
  • scripts/trace_execution.py: Execution trace capture
  • scripts/compare_traces.py: Trace comparison and analysis

© ArabelaTso, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/behavior-preservation-checker of ArabelaTso/Skills-4-SE.

  • SKILL.md
  • references/comparison_techniques.md
  • references/difference_patterns.md
  • scripts/behavior_checker.py
  • scripts/compare_test_results.py
  • scripts/compare_traces.py
  • scripts/trace_execution.py

Open the folder on GitHubat commit 4f38503

Compare with similar skills

Behavior Preservation Checker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Behavior Preservation Checker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Behavior Preservation Checker this skillArabelaTso/Skills-4-SE253—~2.3kAutomated safety check: PassApache-2.0
Migrate Core Code to Submodulestinyhumansai/openhuman42k—~2.6kAutomated safety check: PassGPL-3.0
ast-grep Structural Searchcode-yeongyu/oh-my-openagent70k—~3.3kAutomated safety check: PassMIT
ast-grep Codemod Referencewarp-drive-data/warp-drive3.2k—~2.6kAutomated safety check: PassMIT
Hai Ast Grephylarucoder/hai-stack386—~1.5kAutomated safety check: PassCustom licence
Java 21 Developer Guide for Grailsapache/grails-core2.9k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Migrate Core Code to Submodules

    tinyhumansai/openhuman

    Plans and carries out moving non-host-specific code and its tests from the OpenHuman core into vendored tiny submodule libraries, then releases the submodule and re-pins the host.

    42k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed
  • ast-grep Structural Search

    code-yeongyu/oh-my-openagent

    Searches and rewrites code by syntax-tree shape across 25 languages with ast-grep, for codemods, structural queries and YAML lint rules, using a Python wrapper script.

    70k GitHub stars~3.3k tokensUpdated today
    DevelopmentAuto-check passed
  • ast-grep Codemod Reference

    warp-drive-data/warp-drive

    Reference for writing and debugging TypeScript and JavaScript codemods with @ast-grep/napi: parsing, node queries, meta-variables, rule objects and editing.

    3.2k GitHub stars~2.6k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Hai Ast Grep

    hylarucoder/hai-stack

    Produces a ready-to-run ast-grep command or reusable YAML lint/codemod rule, validated against positive and negative fixtures.

    386 GitHub stars~1.5k tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Guide for writing modern Java 21 in a Grails and Groovy codebase: records, sealed classes, pattern matching, text blocks and how Java works alongside Groovy.

    2.9k GitHub stars~1.9k tokensUpdated today
    DevelopmentAuto-check passed
  • trace-mcp Codemods

    nikolai-vysotskyi/trace-mcp

    Replaces repeated hand edits with the trace-mcp apply_codemod tool, previewing matches in a dry run before bulk mechanical changes across one file or many.

    189 GitHub stars~729 tokensUpdated today
    DevelopmentAuto-check passed

More from ArabelaTso/Skills-4-SE

All 170 skills in this repo
  • Framework Migration Assistant

    ArabelaTso/Skills-4-SE

    Automatically migrate Python web applications between frameworks (Flask → FastAPI, Django → FastAPI).

    253 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Metamorphic Test Generator

    ArabelaTso/Skills-4-SE

    Generate test cases using metamorphic testing by applying transformations based on metamorphic properties.

    253 GitHub stars~798 tokensUpdated 1 mo ago
    Auto-check passed
  • Reproduction Trace Instrumenter

    ArabelaTso/Skills-4-SE

    Instruments programs to capture execution traces specifically for reproducing reported bugs, enabling consistent replay and diagnosis of failures.

    253 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spring Mvc To Boot Migrator

    ArabelaTso/Skills-4-SE

    Automatically migrate Spring MVC applications to Spring Boot.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • State Snapshot Instrumenter

    ArabelaTso/Skills-4-SE

    Instrument programs (Python, C/C++, Java) to capture snapshots of key program states at runtime, including variables, memory, and call stacks.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Behavior Preservation Checker

What does Behavior Preservation Checker do?

Compare runtime behavior between original and migrated repositories to detect behavioral differences, regressions, and semantic changes. Behavior Preservation Checker is an agent skill from ArabelaTso/Skills-4-SE. Compare runtime behavior between original and migrated repositories to detect behavioral differences, regressions, and semantic changes.

When should I use Behavior Preservation Checker?

Behavior Preservation Checker fits situations like: validating code migrations; framework upgrades; any transformation that should preserve behavior.

How do I install Behavior Preservation Checker in Claude Code?

Run `npx skills add ArabelaTso/Skills-4-SE --skill behavior-preservation-checker -a claude-code`. Or copy the skill folder (skills/behavior-preservation-checker in ArabelaTso/Skills-4-SE) into .claude/skills/behavior-preservation-checker in your project. Claude Code loads it when a task matches its description.

How do I install Behavior Preservation Checker in Codex?

Run `npx skills add ArabelaTso/Skills-4-SE --skill behavior-preservation-checker -a codex`. Or copy the skill folder (skills/behavior-preservation-checker in ArabelaTso/Skills-4-SE) into .agents/skills/behavior-preservation-checker in your project. Codex loads it when a task matches its description.

Can I use Behavior Preservation Checker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ArabelaTso/Skills-4-SE --skill behavior-preservation-checker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/behavior-preservation-checker, .gemini/skills/behavior-preservation-checker, .github/skills/behavior-preservation-checker and .opencode/skills/behavior-preservation-checker in your project.

What does Behavior Preservation Checker need to run?

Going by SKILL.md and its folder, Behavior Preservation Checker needs Python for the scripts in its folder and the command-line tools its instructions call (python, git and pytest). Our summary lists: Python 3.

Does Behavior Preservation Checker access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Behavior Preservation Checker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Behavior Preservation Checker use?

Behavior Preservation Checker is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Behavior Preservation Checker use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.

What are the alternatives to Behavior Preservation Checker?

Skills that share tags, products or a category with Behavior Preservation Checker: Migrate Core Code to Submodules (tinyhumansai/openhuman, 42k stars), ast-grep Structural Search (code-yeongyu/oh-my-openagent, 70k stars), ast-grep Codemod Reference (warp-drive-data/warp-drive, 3.2k stars) and Hai Ast Grep (hylarucoder/hai-stack, 386 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Behavior Preservation Checker?

ArabelaTso (a GitHub user) maintains it in ArabelaTso/Skills-4-SE, which has 253 GitHub stars. The repository holds 170 skills in this directory. The repository was last updated on August 21, 2026.

Source: ArabelaTso/Skills-4-SE on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.