Agent skill

Polylith Migrate Dedupe

by DavidVujic in DavidVujic/python-polylith

[Internal sub-skill of polylith-migrate-orchestrator (optional, runs only when opted in during polylith-migrate-discover).

MITAuto-check passedData & Analytics

Install Polylith Migrate Dedupe

skills CLI
$ npx skills add DavidVujic/python-polylith --skill polylith-migrate-dedupe -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install DavidVujic/python-polylith polylith-migrate-dedupe --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/DavidVujic/python-polylith.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/polylith/migrate-project/polylith-migrate-dedupe .claude/skills/polylith-migrate-dedupe && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
polylith-migrate-dedupe
GitHub stars
553
Token cost
~1.9k tokens
SKILL.md length
858 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

[Internal sub-skill of polylith-migrate-orchestrator (optional, runs only when opted in during polylith-migrate-discover).

  • Works in 4 steps: Identify Duplication Candidates → Present Candidates to the User → Execute Deduplication for Approved… → …
  • Tasks that involve Data cleaning
  • SKILL.md covers Goal, When to Use, Classification and When to parameterize vs. keep…, plus 7 more sections
  • Calls git

What it does

Polylith Migrate Dedupe is an agent skill from DavidVujic/python-polylith. [Internal sub-skill of polylith-migrate-orchestrator (optional, runs only when opted in during polylith-migrate-discover). Do not load directly — load polylith-migrate-orchestrator first.] Identify and execute controlled deduplication of code during migration (if the user opts in).

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: Tooling support for the Polylith Architecture in Python. The licence is MIT.

When your agent uses it

  • Tasks that involve Data cleaning

Example prompts

  • “/polylith-migrate-dedupe”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Identify Duplication Candidates
  2. Present Candidates to the User
  3. Execute Deduplication for Approved Candidates
  4. Verify Changes

What it can do on your machine

Read from SKILL.md and the folder at commit a7a80f2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Polylith Migrate Dedupe loads about 1.9k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 858 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from DavidVujic/python-polylith at commit a7a80f2, republished under its MIT licence (© DavidVujic). 858 words, ~1,853 tokens.

Download SKILL.mdSave it as .claude/skills/polylith-migrate-dedupe/SKILL.md (or your agent's skills folder).
name
polylith-migrate-dedupe
description
[Internal sub-skill of `polylith-migrate-orchestrator` (optional, runs only when opted in during `polylith-migrate-discover`). Do not load directly — load `polylith-migrate-orchestrator` first.] Identify and execute controlled deduplication of code during migration (if the user opts in).

Skill: polylith-migrate-dedupe

📐 Scope vs sibling skills. This skill is opportunistic deduplication that may be triggered any time during refactoring when duplication candidates surface. It is not the canonical place for the structural decompositions:

  • For "split this big component into smaller ones", use polylith-migrate-split-big-component (it already includes a dedup-analysis subsection — usually sufficient on a first migration).
  • For "this component's core.py mixes domains", use polylith-migrate-split-component-internals.
  • For "two projects have overlapping code, split shared from project-specific", use polylith-migrate-isolate-shared-and-project-logic.

Use polylith-migrate-dedupe when none of the above fits cleanly — e.g., duplication discovered across already-extracted components that don't map to a structural split.

Goal

Identify duplication candidates during the migration process and execute controlled deduplication for user-approved candidates.

When to Use

  • After splitting the big component or extracting standalone modules.
  • When potential duplication between components is suspected.

Classification

Use this table when deciding whether a candidate is a real duplicate:

ClassDefinitionAction
IdenticalSame logic, same control flow, only trivial differences (variable names, formatting, ordering of independent statements).Extract into a shared component. Both call sites import from it.
SimilarSame purpose, slightly different behaviour (e.g., different default arguments, project-specific fields on an otherwise shared model).Extract a parameterized shared component. Project-specific behaviour passes in as arguments or subclass hooks. Avoid forcing a one-size-fits-all signature.
CoincidentalLooks similar (same function name, same shape) but serves unrelated purposes.Leave alone. Sharing here would couple two domains that should evolve independently.

When to parameterize vs. keep separate

  • Parameterize when the core logic is identical and only data/config differs.
  • Keep separate when control flow or structure diverges (different frameworks, different patterns) — forcing a shared abstraction here creates a brittle "shared core" that grows project-specific flags over time.
  • Extract shared base + per-project wrappers when there's a significant shared core but non-trivial project-specific logic around it.
Worked example — logging

Two projects each had their own init_logging. The core (structlog setup, base log levels, JSON formatter) was identical; the differences were:

  • Project A added loggers for httpx, backoff.
  • Project B added a logger for confluent_kafka_helpers.

The shared component exposed an init(config, *, extra_loggers=None, cache_logger_on_first_use=False) function. Each project's base calls init with its own extra_loggers dict. No coincidental coupling, no version skew, and adding a third project requires only its own dict — not a change to the shared component.

Shared-component naming

When creating a shared component to deduplicate code, name it after the domain or capability it represents, never after how it's used. Good: logging, kafka_client, merchant_serializer. Bad: shared_utils, common, helpers, misc. Generic-named bricks attract more code over time and become the next thing that needs decomposing.

Inputs

From migration/<PROJECT>/state.md:

  • TARGET_TOP_NS
  • Verification commands (RUN_TEST_CMD, RUN_LINT_CMD, RUN_TYPECHECK_CMD).

From migration/<PROJECT>/manifest.md:

  • Module map of components.

All inputs from state.md are assumed to satisfy the validation rules in polylith-migrate-discover (### Validation rules). Validate before proceeding.

Steps

1. Identify Duplication Candidates
  • Use directory_tree and grep to scan for overlapping logic between components.
  • Classify candidates by type:
    • Identical: Code that is exactly the same.
    • Similar: Code that serves the same purpose but with minor differences.
    • Coincidental: Code that looks similar but serves unrelated purposes.
Show full SKILL.md (354 more words)Show less
2. Present Candidates to the User
  • Provide a list of duplication candidates, including:
    • Component names.
    • File paths.
    • Type of duplication (identical, similar, coincidental).
    • Risk assessment (low, medium, high).
  • Ask the user to approve or reject each candidate for deduplication.
3. Execute Deduplication for Approved Candidates
  • For each approved candidate:
    • Identical Code: Extract the shared logic into a new component and update imports.
    • Similar Code: Refactor to use shared logic or parameterize differences.
    • Coincidental Code: Leave as-is.
  • Update pyproject.toml to include any new components.
  • Run POLY_CMD_PREFIX sync to synchronize the workspace.
4. Verify Changes
  • Run RUN_TEST_CMD to ensure no regressions.
  • Run RUN_LINT_CMD and RUN_TYPECHECK_CMD if set.
  • Run POLY_CMD_PREFIX check to validate the workspace structure.

Verify

  • All tests pass (RUN_TEST_CMD).
  • Linting and type-checking pass (if set).
  • The workspace structure is valid (POLY_CMD_PREFIX check).

Common failure modes

SymptomLikely causeRemediation
Two pieces of code look identical but operate on different domains (e.g., both are validate(...) but one is for users, the other for transactions)Coincidental similarity, not real duplication.Classify as "coincidental"; leave both in place. Resist the urge to share.
The candidate shared component would pull in framework-specific dependencies (e.g., a "logging" shared brick that needs both confluent_kafka_helpers and httpx)Wrong shared abstraction — you're sharing the union of two project surfaces.Revert and parameterize instead: keep the shared core minimal and pass project-specific values as arguments. See the "Pattern: Parameterize the shared component" guidance in polylith-migrate-split-big-component.
Tests break after deduplication because mock.patch("<old.path>") no longer hits anythingPatch strings reference the pre-dedup module path.Update patch strings to the new shared module path. Validate by deliberately breaking the patched function and confirming the test fails.

Done When

  • Duplication candidates are identified and presented to the user.
  • User-approved candidates are deduplicated.
  • All tests and checks pass.
  • The workspace structure is valid.

Commit

After verification passes, commit this phase to the migration branch:

bash
git add -A && git commit -m "migrate(<PROJECT>): phase optional — dedupe"

Substitute <PROJECT> from state.md. This is an optional skill off the numbered main line, so the commit uses the literal phase optional label (no <N>). Do not proceed without a clean commit — the per-phase commit is the rollback point for the next phase's failure-mode tables.

© DavidVujic, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/polylith/migrate-project/polylith-migrate-dedupe of DavidVujic/python-polylith.

Open the folder on GitHubat commit a7a80f2

Compare with similar skills

Polylith Migrate Dedupe next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Polylith Migrate Dedupe compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Polylith Migrate Dedupe this skillDavidVujic/python-polylith553—~1.9kAutomated safety check: PassMIT
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 10 days ago
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from DavidVujic/python-polylith

All 37 skills in this repo
  • Polylith Base Creation

    DavidVujic/python-polylith

    Create a Polylith base with poly create base — the entry point of a deployable application (HTTP API, CLI, message-queue consumer, AWS Lambda handler, GCP Cloud Function, scheduled job).

    553 GitHub stars~757 tokensUpdated 3 days ago
    Auto-check passed
  • Polylith Check

    DavidVujic/python-polylith

    Validate a Polylith workspace with poly check — the canonical CI gate.

    553 GitHub stars~972 tokensUpdated 3 days ago
    Auto-check passed
  • Polylith Component Creation

    DavidVujic/python-polylith

    Create a Polylith component with poly create component — a reusable, isolated brick implementing business logic, a feature, a domain module, or a capability.

    553 GitHub stars~800 tokensUpdated 3 days ago
    Auto-check passed
  • Polylith Dependency Management

    DavidVujic/python-polylith

    Add or manage third-party dependencies in a Polylith workspace.

    553 GitHub stars~643 tokensUpdated 3 days ago
    Auto-check passed
  • Polylith Dependency Visualization

    DavidVujic/python-polylith

    Visualize brick × brick dependencies with poly deps — find circular dependencies, inspect a brick's public interface, and detect interface-bypass violations.

    553 GitHub stars~906 tokensUpdated 3 days ago
    Auto-check passed
  • Polylith Diff

    DavidVujic/python-polylith

    List Polylith bricks whose implementation changed since a git tag using poly diff.

    553 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed

Questions about Polylith Migrate Dedupe

What does Polylith Migrate Dedupe do?

[Internal sub-skill of polylith-migrate-orchestrator (optional, runs only when opted in during polylith-migrate-discover). Polylith Migrate Dedupe is an agent skill from DavidVujic/python-polylith. [Internal sub-skill of polylith-migrate-orchestrator (optional, runs only when opted in during polylith-migrate-discover).

When should I use Polylith Migrate Dedupe?

Polylith Migrate Dedupe fits situations like: tasks that involve Data cleaning.

How do I install Polylith Migrate Dedupe in Claude Code?

Run `npx skills add DavidVujic/python-polylith --skill polylith-migrate-dedupe -a claude-code`. Or copy the skill folder (.agents/skills/polylith/migrate-project/polylith-migrate-dedupe in DavidVujic/python-polylith) into .claude/skills/polylith-migrate-dedupe in your project. Claude Code loads it when a task matches its description.

How do I install Polylith Migrate Dedupe in Codex?

Run `npx skills add DavidVujic/python-polylith --skill polylith-migrate-dedupe -a codex`. Or copy the skill folder (.agents/skills/polylith/migrate-project/polylith-migrate-dedupe in DavidVujic/python-polylith) into .agents/skills/polylith-migrate-dedupe in your project. Codex loads it when a task matches its description.

Can I use Polylith Migrate Dedupe in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DavidVujic/python-polylith --skill polylith-migrate-dedupe -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/polylith-migrate-dedupe, .gemini/skills/polylith-migrate-dedupe, .github/skills/polylith-migrate-dedupe and .opencode/skills/polylith-migrate-dedupe in your project.

What does Polylith Migrate Dedupe need to run?

Going by SKILL.md and its folder, Polylith Migrate Dedupe needs the command-line tools its instructions call (git).

Does Polylith Migrate Dedupe access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Polylith Migrate Dedupe safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Polylith Migrate Dedupe use?

Polylith Migrate Dedupe is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Polylith Migrate Dedupe use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Polylith Migrate Dedupe?

Skills that share tags, products or a category with Polylith Migrate Dedupe: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Issues Deduplication (JetBrains/ideavim, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Polylith Migrate Dedupe?

DavidVujic (a GitHub user) maintains it in DavidVujic/python-polylith, which has 553 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on October 4, 2026.

Source: DavidVujic/python-polylith on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.