Official agent skill

Splitting Oversized Modules

by PostHog in PostHog/posthog-foss

Split an oversized Python module (a thousand-plus-line logic.py, models.py, api.py, or its test file) into a package of one module per concern, mechanically and provably without changing behavior.

OfficialMITAuto-check passedFrontend & Design

Install Splitting Oversized Modules

skills CLI
$ npx skills add PostHog/posthog-foss --skill splitting-oversized-modules -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PostHog/posthog-foss splitting-oversized-modules --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/splitting-oversized-modules .claude/skills/splitting-oversized-modules && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
splitting-oversized-modules
GitHub stars
721
Token cost
~2.2k tokens
SKILL.md length
1,104 words
Files
3 (incl. scripts)
Skills in repo
213
Repo updated
First seen
Licence
MIT

At a glance

Split an oversized Python module (a thousand-plus-line logic.py, models.py, api.py, or its test file) into a package of one module per concern, mechanically and provably without changing behavior.

  • Works in 6 steps: Map the symbols and assign them to… → Move the code → Retarget callers and mock patch targets → …
  • Frontend & Design work in your project
  • SKILL.md covers Is it worth doing here?, Method, Traps worth knowing and Left for you
  • Runs Python scripts from its folder; calls uv, ruff and git

What it does

Splitting Oversized Modules is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Split an oversized Python module (a thousand-plus-line logic.py, models.py, api.py, or its test file) into a package of one module per concern, mechanically and provably without changing behavior. Use on a request to split / break up / decompose a god module or move functions out of one, once a human has agreed to split one before some other change, or before restructuring code inside a module already over roughly a thousand lines — breaking up a long function or extracting helpers in place leaves everything in…

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/split_module.py` and `scripts/verify_pure_move.py`).

It sits in Frontend & Design. It works with Python. The repository describes itself as: PostHog FOSS is a read-only mirror of PostHog, with all proprietary code removed. NOTE: This repo is synced automatically from the main PostHog repo. Please raise any issues and… The licence is MIT.

When your agent uses it

  • Frontend & Design work in your project

Example prompts

  • “/splitting-oversized-modules”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Map the symbols and assign them to concerns
  2. Move the code
  3. Retarget callers and mock patch targets
  4. Split the tests the same way
  5. Prove it is a pure move
  6. Then the suite

What it can do on your machine

Read from SKILL.md and the folder at commit 2c48221. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • ruff
    • git
    • rg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Splitting Oversized Modules loads about 2.2k tokens when it runs. Until then it costs about 242 tokens; SKILL.md has 1,104 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~242
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from PostHog/posthog-foss at commit 2c48221, republished under its MIT licence (© PostHog). 1,104 words, ~2,242 tokens.

Download SKILL.mdSave it as .claude/skills/splitting-oversized-modules/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
splitting-oversized-modules
description
Split an oversized Python module (a thousand-plus-line logic.py, models.py, api.py, or its test file) into a package of one module per concern, mechanically and provably without changing behavior. Use on a request to split / break up / decompose a god module or move functions out of one, once a human has agreed to split one before some other change, or before restructuring code inside a module already over roughly a thousand lines — breaking up a long function or extracting helpers in place leaves everything in the same file, so check the worth-it gate and propose the split first. Covers that gate, assigning symbols to concerns with an acyclic dependency graph, the AST plus tokenize move script, and proving the result is a pure move. Python only — for frontend files use writing-ui-components. Not for extracting a shared helper into common/, and not for moving code between products, which is isolating-product-facade-contracts.

Splitting oversized modules

A module nobody wants to open taxes every change in it, and agents pay that tax on every task because they do not carry knowledge between them. A 3000-line logic.py beside a 3000-line test_logic.py is roughly 55k input tokens of reading before a line gets written. Splitting one measured case cut the read-set 81 to 97% depending on the concern touched, while total lines grew about 5% from the repeated import headers.

So the goal is a small read-set for a typical change, not tidiness.

Is it worth doing here?

All three must hold: over roughly a thousand lines, several concerns that change independently, and something still actively changing in it. A big cohesive frozen file buys nothing. Skip generated files, and skip a file several people are mid-change in (git log --oneline -20 -- <file>).

A complexity warning on one function is a symptom of this, not a separate job. Extracting that function into helpers leaves every helper in the same file, so the read-set is unchanged and the file gets longer. Measure the file first (wc -l), and when it clears the gate, propose the split instead of the in-place extraction.

Splitting is a separate PR from whatever you came to do. Land the move as its own base PR and branch your work on top — see /stacking-prs. Never bundle it into a feature diff, and say what you are doing in one line before you start: which file, and that the split lands separately. Nobody minds the base PR; they mind finding it inside a feature diff.

If a human has not asked for the split, propose it and let them decide.

Method

Commands assume the repo root. S=.agents/skills/splitting-oversized-modules/scripts.

1. Map the symbols and assign them to concerns
sh
uv run --no-project python $S/split_module.py <module> --skeleton > layout.json

Edit layout.json into modules named after concerns (baselines, quarantine, ci_status), not layers (helpers, utils, core). Put module-level state every module needs its own copy of — a logger, a compiled regex — under "__shared__".

Two rules while assigning:

Keep the dependency graph acyclic. If two modules need each other, the seam is wrong. Separating reads from writes fixes most cycles: a lifecycle module and a verification module that both need the same lookups should share a third leaf module holding those lookups.

Do not name a module after a common local variable. A module called artifacts or runs gets shadowed by artifacts = [...] inside a function, and Python binds the whole scope, so a call above the assignment fails too. ruff catches the reachable cases (F811, F823) but not one where the local is assigned before any module use. Pick a non-colliding name, or split finer until the name is specific.

2. Move the code
sh
uv run --no-project python $S/split_module.py <module> layout.json
ruff check <package>/ --fix && ruff format <package>/

Hand-editing a 3000-line move loses code and silently rewrites it. The script copies each symbol verbatim by AST line range, refuses to run unless the layout covers every symbol exactly once, and requalifies cross-module references with tokenize so it never rewrites a name inside a docstring or string literal.

Cross-module calls come out as from . import baselines plus baselines.foo(...), never from .baselines import foo. One binding per symbol means a mock patch on the definition reaches every caller, and from . import x also survives an accidental cycle by falling back to sys.modules.

ruff --fix leaves an emptied if TYPE_CHECKING: pass behind. Delete those, then re-run ruff check --fix so the orphaned TYPE_CHECKING import goes too. The # --- Section --- dividers from the monolith are usually redundant once the module name says it: rg -n '^#\s*-{2,}' <package>/.

3. Retarget callers and mock patch targets

Every logic.foo() becomes <module>.foo(), and every patch target moves to the module that defines the symbol, since that is the binding its callers resolve:

python
patch("products.foo.backend.logic._post_commit_status")   # -> ...logic.ci_status._post_commit_status
patch.object(logic, "_post_commit_status")                # -> patch.object(ci_status, "_post_commit_status")

Sweep conftest.py too, not just test_*.py — fixtures patch, and a pass matched on test files leaves conftest pointing at paths that no longer exist. Enumerate candidates with rg -n 'patch\(|patch\.object\(' <tests>/. Watch for forms a naive sweep misses: a combined from x import a, logic, a symbol import (from ..logic import SomeError), and a name the old module merely re-exported, which should now come from its real source.

Show full SKILL.md (426 more words)Show less
4. Split the tests the same way

A split module beside an untouched 3000-line test file solves half the problem. Mirror the package, so logic/comments.py pairs with tests/logic/test_comments.py:

sh
uv run --no-project python $S/split_module.py <test file> tests_layout.json \
    --package-dir <tests>/logic --init-doc "Unit tests for the logic package."

Assign whole test classes. Carving up a class stops being a pure move.

5. Prove it is a pure move
sh
git show HEAD:<module> > /tmp/before.py
uv run --no-project python $S/verify_pure_move.py /tmp/before.py <package>

It compares every top-level definition before and after, ignoring the module qualification and relative-import depth the split introduces, and reports anything missing, unexpected, duplicated with drift, or changed. This is what makes a 50-file diff reviewable: "every definition is identical" is checkable, unlike the diff. Always re-derive the before side from git, never from your own earlier output. Re-run it after any cleanup pass.

For a test split, pass --strip-package <package> so the source package's module names are normalized too.

6. Then the suite
sh
hogli test <tests>/

Static checks cover import wiring and equivalence, not behavior. Deferred imports inside function bodies fail at call time, so only running the tests catches a mistake there.

Traps worth knowing

  • A re-export shim in __init__.py gives every symbol two bindings, so patching one leaves internal callers on the real implementation: the patch applies, the test passes, and nothing intercepted. .semgrep/rules/devex/no-init-reexports.yaml already blocks eager re-exports repo-wide; the silent-mock failure is the extra reason not to add one.
  • Moving code makes its old semgrep findings new. Semgrep skips its baseline comparison for findings in files that did not exist in the baseline commit, so every pre-existing violation in the moved code comes back as blocking. Check before pushing: uv run --no-project --with semgrep semgrep --config .semgrep/rules/devex <package>/.
  • Relative imports break one level down. The script deepens them, including inside function bodies. A missed one is invisible until the function runs.
  • A models.py split is the one case that needs re-exports. Django imports an app's models package but does not recurse into it, so model classes in submodules never reach the app registry and every from ..models import Model caller breaks. Add aggregation imports to models/__init__.py; no-init-reexports.yaml exempts **/models/** for exactly this reason. The script warns when it writes a models/ package.
  • A top-level statement after the first definition stops the split. A module-level if, an assert, or a registration call cannot be attributed to a concern, and copying it into every module would run its side effect once per module. Move it into a function or above the definitions first.

Left for you

  • Decide on the section dividers the move leaves behind.
  • Rename locals that shadow a new module name.
  • Update any doc or skill that links the old path; a split leaves dead links behind.

© PostHog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in .agents/skills/splitting-oversized-modules of PostHog/posthog-foss.

  • SKILL.md
  • scripts/split_module.py
  • scripts/verify_pure_move.py

Open the folder on GitHubat commit 2c48221

Compare with similar skills

Splitting Oversized Modules next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Splitting Oversized Modules compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Splitting Oversized Modules this skillPostHog/posthog-foss721—~2.2kAutomated safety check: PassMIT
Claude Desktop Chinese Localizationjavaht/claude-desktop-zh-cn7.5k—~1.6kAutomated safety check: PassMIT
DocsPrefectHQ/fastmcp28k—~1kAutomated safety check: PassApache-2.0
Language Injectionmicrosoft/data-formulator18k—~1.2kAutomated safety check: PassMIT
Oil UIoil-oil/oil-ui926—~1.7kAutomated safety check: PassMIT
Jarvis Setupethanplusai/jarvis838—~2.5kAutomated safety check: NotesCustom licence

Similar skills

  • Claude Desktop Chinese Localization

    javaht/claude-desktop-zh-cn

    Adds missing Simplified and Traditional Chinese translations to the Claude Desktop Chinese patch across three layers, then checks how many mappings actually hit.

    7.5k GitHub stars~1.6k tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Docs

    PrefectHQ/fastmcp

    Write or revise a page under docs/ for gofastmcp.com. An agent skill from PrefectHQ/fastmcp.

    28k GitHub stars~1k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Language Injection

    microsoft/data-formulator

    Official

    LLM Agent 多语言注入规范。在修改 Agent 提示词、添加新的 Agent 端点、处理用户可见的后端消息(messagecode)时使用。

    18k GitHub stars~1.2k tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Oil UI

    oil-oil/oil-ui

    Design, improve, and review interfaces for websites, apps, dashboards, and components: explore distinct design directions, compare styles side by side, and refine visual hierarchy against real…

    926 GitHub stars~1.7k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Jarvis Setup

    ethanplusai/jarvis

    A skill your agent uses when helping someone install, configure, or debug a fresh clone of JARVIS (this repo) — especially "the mic doesn't work", "JARVIS says his language systems are down", any…

    838 GitHub stars~2.5k tokensUpdated 27 days ago
    Frontend & DesignAuto-check: notes
  • Ibm A11y Route Scan

    langflow-ai/langflow

    Batch-scan Langflow frontend routes for accessibility issues using the Python IBM Equal Access scanner (scripts/a11y/a11yscan.py) and produce JSON/Markdown/HTML reports.

    156k GitHub stars~1.6k tokensUpdated today
    Frontend & DesignAuto-check passed

More from PostHog/posthog-foss

All 213 skills in this repo
  • Authoring Log Alerts

    PostHog/posthog-foss

    Official

    Author useful, low-noise log alerts on services in a PostHog project.

    721 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Autoresolving PR Conflicts

    PostHog/posthog-foss

    Official

    Operating procedure for the conflict-autoresolver agent: sweep open PostHog/posthog PRs that conflict with master, resolve the trivial conflicts (generated artifacts deterministically, source…

    721 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Help users debug PostHog Error Tracking stack-trace symbolication for any supported platform — JavaScript/TypeScript web, React Native (Hermes), Android (Proguard / R8), or iOS / macOS (dSYM).

    721 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Exploring Apm Traces

    PostHog/posthog-foss

    Official

    Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP.

    721 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Exploring LLM Traces

    PostHog/posthog-foss

    Official

    Debug and inspect LLM/AI agent traces using PostHog's MCP tools.

    721 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Investigate Metric

    PostHog/posthog-foss

    Official

    Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries.

    721 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Splitting Oversized Modules

What does Splitting Oversized Modules do?

Split an oversized Python module (a thousand-plus-line logic.py, models.py, api.py, or its test file) into a package of one module per concern, mechanically and provably without changing behavior. Splitting Oversized Modules is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization.py, or its test file) into a package of one module per concern, mechanically and provably without changing behavior.

When should I use Splitting Oversized Modules?

Splitting Oversized Modules fits situations like: frontend & Design work in your project.

How do I install Splitting Oversized Modules in Claude Code?

Run `npx skills add PostHog/posthog-foss --skill splitting-oversized-modules -a claude-code`. Or copy the skill folder (.agents/skills/splitting-oversized-modules in PostHog/posthog-foss) into .claude/skills/splitting-oversized-modules in your project. Claude Code loads it when a task matches its description.

How do I install Splitting Oversized Modules in Codex?

Run `npx skills add PostHog/posthog-foss --skill splitting-oversized-modules -a codex`. Or copy the skill folder (.agents/skills/splitting-oversized-modules in PostHog/posthog-foss) into .agents/skills/splitting-oversized-modules in your project. Codex loads it when a task matches its description.

Can I use Splitting Oversized Modules in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PostHog/posthog-foss --skill splitting-oversized-modules -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/splitting-oversized-modules, .gemini/skills/splitting-oversized-modules, .github/skills/splitting-oversized-modules and .opencode/skills/splitting-oversized-modules in your project.

What does Splitting Oversized Modules need to run?

Going by SKILL.md and its folder, Splitting Oversized Modules needs Python for the scripts in its folder and the command-line tools its instructions call (uv, ruff, git and rg). Our summary lists: Python 3.

Does Splitting Oversized Modules access the network?

SKILL.md contains no URLs. Its commands use uv and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Splitting Oversized Modules safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Splitting Oversized Modules use?

Splitting Oversized Modules is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Splitting Oversized Modules use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Splitting Oversized Modules?

Skills that share tags, products or a category with Splitting Oversized Modules: Claude Desktop Chinese Localization (javaht/claude-desktop-zh-cn, 7.5k stars), Docs (PrefectHQ/fastmcp, 28k stars), Language Injection (microsoft/data-formulator, 18k stars) and Oil UI (oil-oil/oil-ui, 926 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Splitting Oversized Modules?

PostHog (a GitHub organization, an official publisher) maintains it in PostHog/posthog-foss, which has 721 GitHub stars. The repository holds 213 skills in this directory. The repository was last updated on October 7, 2026.

Source: PostHog/posthog-foss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.