Official agent skill

Maintaining Python Tests

by PostHog in PostHog/posthog-foss

Maintains existing pytest and Django test suites without weakening correctness.

OfficialMITAuto-check passedTesting & QA

Install Maintaining Python Tests

skills CLI
$ npx skills add PostHog/posthog-foss --skill maintaining-python-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PostHog/posthog-foss maintaining-python-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/maintaining-python-tests .claude/skills/maintaining-python-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
maintaining-python-tests
GitHub stars
721
Token cost
~2.5k tokens
SKILL.md length
1,314 words
Files
3 (incl. references)
Skills in repo
213
Repo updated
First seen
Licence
MIT

At a glance

Maintains existing pytest and Django test suites without weakening correctness.

  • Works in 10 steps: Define the result → Rank current work → State why the test exists → …
  • Asked to reduce Python test runtime
  • SKILL.md covers Principles, Workflow, Deletion rules and Boundaries, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Maintaining Python Tests is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Maintains existing pytest and Django test suites without weakening correctness. Use when asked to reduce Python test runtime or CI work, investigate slow pytest families, remove stale migration tests, consolidate repeated setup, improve Python test ownership, or measure whether a test optimization worked after merge. Ranks work by measured cost, applies the writing-tests value gate to existing coverage, preserves distinct behavior cases, validates isolation after shared-fixture changes, and separates testcase…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/measurement.md` and `references/optimization-patterns.md`).

It sits in Testing & QA, covering Failing and flaky tests, Unit testing and Test generation. It works with Python, pytest, PostHog and Django. The repository describes itself as: PostHog FOSS is a read-only mirror of PostHog, with all proprietary code removed. NOTE: This repo is synced automatically from the main PostHog repo. Please raise any issues and… The licence is MIT.

When your agent uses it

  • Asked to reduce Python test runtime
  • Investigate slow pytest families
  • Remove stale migration tests
  • Consolidate repeated setup

Example prompts

  • “Use the maintaining-python-tests skill to maintain existing pytest and Django test suites without weakening correctness”
  • “/maintaining-python-tests”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Define the result
  2. Rank current work
  3. State why the test exists
  4. Establish a baseline
  5. Find the cost center
  6. Select the smallest safe fix
  7. Preserve isolation
  8. Validate correctness and cost
  9. Report the local result accurately
  10. Verify after merge

What it can do on your machine

Read from SKILL.md and the folder at commit 2c48221. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Maintaining Python Tests loads about 2.5k tokens when it runs, and up to ~6.6k if it reads all its reference files. Until then it costs about 159 tokens; SKILL.md has 1,314 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PostHog/posthog-foss at commit 2c48221, republished under its MIT licence (© PostHog). 1,314 words, ~2,514 tokens.

Download SKILL.mdSave it as .claude/skills/maintaining-python-tests/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
maintaining-python-tests
description
Maintains existing pytest and Django test suites without weakening correctness. Use when asked to reduce Python test runtime or CI work, investigate slow pytest families, remove stale migration tests, consolidate repeated setup, improve Python test ownership, or measure whether a test optimization worked after merge. Ranks work by measured cost, applies the writing-tests value gate to existing coverage, preserves distinct behavior cases, validates isolation after shared-fixture changes, and separates testcase work from pytest-suite wall time. For an intermittent failure, use fixing-flaky-tests instead.

Maintaining Python tests

Before you propose a change to how the suite runs in CI, check things already tried. It records measured verdicts on test parallelism, sharding, and coverage-based selection, so a rejected approach is not rebuilt.

Use this skill for an existing Python test suite. Use /writing-tests before adding or substantially changing coverage. Use /fixing-flaky-tests when intermittent failure is the main problem.

The goal is not a smaller test count. The goal is a suite that catches the same realistic regressions with less compute, less waiting, and less maintenance.

Principles

  1. Measure before changing code. Rank tests by total observed work, not by one slow local run.
  2. Preserve behavior coverage. Keep cases that exercise different validation, persistence, integration, or output paths.
  3. Remove only expired or redundant coverage. Get explicit approval before deleting a test.
  4. Share expensive infrastructure, not mutable test state. Preserve isolation with unique IDs, schemas, tables, topics, or tenants.
  5. Measure after merge. Local results prove the mechanism. Fresh master data proves the result in CI.
  6. Separate testcase work from suite wall time. A change can reduce summed testcase time and not change the slowest pytest suite.

Read measurement.md before you query timing data or report an improvement. Read optimization-patterns.md when you select a fix.

Workflow

1. Define the result

Write down the user problem before choosing a test:

  • Reduce total test compute.
  • Reduce the slowest pytest suite.
  • Remove expired maintenance burden.
  • Restore test ownership.
  • Reduce repeated external-service setup.

These results need different measurements. Do not claim faster CI when only summed test call time decreased.

2. Rank current work

Use recent PostHog test spans from master when available. Start with a complete window after the latest relevant merge.

Rank at least three views:

  • Individual tests: execution count multiplied by duration.
  • Parameterized families: all cases under one test function or fixture family.
  • Pytest suites or shards: root-span wall time and the slowest suite per run attempt.

Use p50 to find steady cost. Use p95 to find contention or tail behavior. Use sampled observed hours to find repeated medium-cost tests. Use root-span testcase totals for complete pytest work.

Do not select a target from an old ranking after several fixes merge. Rebuild the ranking first.

Do not rank from .test_durations. It holds flat default values (0.01, 18.0, and 60.0) for tests that pytest-split could not time. These values are not measurements.

3. State why the test exists

Apply the /writing-tests gate to every target:

What realistic regression does this test catch that no existing test already catches?

Then classify each case:

  • Distinct behavior: keep it.
  • Same behavior with representative inputs: parameterize it.
  • Framework or implementation detail: replace it with an observable assertion, or propose to delete it.
  • Temporary migration coverage: check whether every supported environment completed the migration.
  • Runnable backfill or reusable migration system: keep active behavior coverage.

Do not infer redundancy from similar names. Read the setup, execution path, assertions, and production entry point.

4. Establish a baseline

Run the exact target with the same command that you will use after the change. Record:

  • Test call time.
  • Full command wall time.
  • Setup and teardown time when available.
  • Number of collected and executed cases.
  • Whether the run was cold or warm.

Run the surrounding class or file when fixtures can change the result. A single test can hide repeated setup that only appears across the family.

Do not compare full pytest wall time with summed call time. They measure different work.

5. Find the cost center

Profile before rewriting. Attribute time to one of these groups:

  • Test collection or environment boot.
  • Fixture setup or teardown.
  • Database creation, flush, or migration.
  • Worker, consumer, broker, or container startup.
  • Product code executed by the test.
  • Snapshot serialization or formatting.
  • Polling, retries, or real time.

Use a wall-clock profile for subprocesses and I/O. Use a CPU profile for Python or JavaScript work. A CPU profile can miss time spent in services or child processes.

6. Select the smallest safe fix

Prefer these options in order:

  1. Remove expired coverage with explicit approval.
  2. Replace incidental implementation assertions with stronger observable assertions.
  3. Move the test to a cheaper level.
  4. Reuse expensive immutable infrastructure across cases.
  5. Reuse production objects only when the production path repeats unnecessary work.
  6. Remove a distinct case only when another test proves the same regression through the same boundary.

Do not add timing assertions. CI timing is too noisy for a correctness test.

Show full SKILL.md (573 more words)Show less
7. Preserve isolation

A shared fixture must not create an ordering dependency.

Before sharing setup, identify all state it owns:

  • Tenant or team IDs.
  • Database rows and transactions.
  • ClickHouse tables.
  • Kafka topics and consumer offsets.
  • Temporal task queues and workflow IDs.
  • Object-storage prefixes.
  • Environment variables and global settings.
  • Mocks and patched functions.

Give each case unique mutable state. Reset shared clients when their local cache or offset can affect the next case.

Run cases alone, together, and in a different order when the framework permits it. Run the family more than once when shared setup has process-level state.

8. Validate correctness and cost

Use this validation ladder:

  1. Run the exact target.
  2. Run the parameter family, class, or file.
  3. Run related integration modes and pipeline versions.
  4. Repeat the family when setup is shared.
  5. Run lint, type checks, and repository preflight.

Keep exact result, response, persisted-state, or emitted-message assertions. Do not replace them with weaker row counts or truthiness checks to gain speed.

Use TDD when the change affects a helper or the contract of a test framework. Make the intended behavior fail first. For a pure runtime change, use a measured baseline instead of an unreliable timing test.

9. Report the local result accurately

Use this format:

text
Target:       <test, family, or shard>
Regression:   <what behavior remains protected>
Cost center:  <measured source of time>
Change:       <smallest fix>
Before:       <metric, command, cases, cold/warm state>
After:        <same metric and conditions>
Correctness:  <focused and surrounding tests>
Isolation:    <how shared state stays separate>
Follow-up:    <post-merge query or none>

Do not report a percentage from two different measurement types.

10. Verify after merge

Wait for fresh master runs that contain the merge commit. Compare equivalent windows and cohorts.

Check all relevant outcomes:

  • Exact test p50 and p95.
  • Family sampled time per workflow attempt.
  • Slowest affected suite per workflow attempt.
  • Complete JUnit testcase time per workflow attempt.
  • Slowest pytest-suite wall time.
  • Failure rate and test count.
  • Unowned test-span share when ownership changed.

If summed testcase time falls but suite wall time does not change, report both facts. Select the next target from the current slowest shard.

If the expected metric does not change, do not declare success from local data. Check whether the cost moved into fixture setup, teardown, or another test.

Deletion rules

Get explicit user approval before you delete tests. Then confirm all of these conditions:

  • You can name the behavior that the test covered.
  • Another named test covers it, or the production behavior no longer exists.
  • No supported upgrade, rollback, or runnable command needs the old state.
  • You keep migration source files and active backfill coverage.
  • The surrounding suite passes without hidden ordering dependencies.

Delete expired tests. Do not leave them skipped. A skip keeps unused code and can still add collection or service cost.

Boundaries

  • Do not remove cases to improve a count.
  • Do not use sleeps, retries, or larger timeouts as performance fixes.
  • Do not mock the behavior that the test exists to prove.
  • Do not convert integration coverage to a unit test unless an integration wiring guard remains.
  • Do not optimize production code only for a test. Confirm that the repeated work exists in production.
  • Do not edit CI workflows as part of a test-runtime fix.
  • Do not use PR runs as the post-merge result. Use fresh master runs.
  • Do not claim causality from a broad before-and-after window when unrelated changes also merged.
  • /writing-tests: decide whether new or changed coverage earns its cost.
  • /fixing-flaky-tests: reproduce and fix intermittent failures.
  • /debugging-ci-failures: classify a failing CI run before changing a test.
  • /django-migrations: change or remove migration-related code safely.
  • /establishing-code-ownership: add or correct ownership rules.
  • /querying-posthog-data: verify the trace schema and HogQL before reading CI timing data.

© PostHog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/maintaining-python-tests of PostHog/posthog-foss.

  • SKILL.md
  • references/measurement.md
  • references/optimization-patterns.md

Open the folder on GitHubat commit 2c48221

Compare with similar skills

Maintaining Python Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Maintaining Python Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Maintaining Python Tests this skillPostHog/posthog-foss721—~2.5kAutomated safety check: PassMIT
Pytestmathiasertl/django-ca158—~924Automated safety check: PassGPL-3.0
Pytestbobmatnyc/claude-mpm155—~8.2kAutomated safety check: PassMIT
Run And Verifyaropan/clist439—~461Automated safety check: PassApache-2.0
Python Devdoccker/cc-use-exp1.1k—~790Automated safety check: PassCustom licence
Django Test Profilinghashgraph-online/awesome-codex-plugins1.2k—~892Automated safety check: PassMIT

Similar skills

  • Pytest

    mathiasertl/django-ca

    Instructions for running, writing, and maintaining tests in this project

    158 GitHub stars~924 tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Pytest

    bobmatnyc/claude-mpm

    pytest - Python's most powerful testing framework with fixtures, parametrization, plugins, and framework integration for FastAPI, Django, Flask

    155 GitHub stars~8.2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Run And Verify

    aropan/clist

    Choose and run focused checks after changing CLIST Python code: Django tests, standalone pytest tests, offline parser fixtures, Ruff, or a relevant management-command check.

    439 GitHub stars~461 tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Python Dev

    doccker/cc-use-exp

    Python 开发规范。当用户操作 .py、pyproject.toml、requirements.txt、setup.py 文件, 或涉及 FastAPI、Django、Flask、pytest、asyncio 开发时触发。

    1.1k GitHub stars~790 tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed
  • Django Test Profiling

    hashgraph-online/awesome-codex-plugins

    Profile and measure slow Django test suites with Django's runner, pytest-django, shell timing, py-spy, cProfile, pytest durations, and profiler visualizations.

    1.2k GitHub stars~892 tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Python Rules

    softspark/ai-toolkit

    Python coding rules: style, patterns, security, testing. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~2.9k tokensUpdated today
    Backend & APIsAuto-check passed

More from PostHog/posthog-foss

All 213 skills in this repo
  • Authoring Log Alerts

    PostHog/posthog-foss

    Official

    Author useful, low-noise log alerts on services in a PostHog project.

    721 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Autoresolving PR Conflicts

    PostHog/posthog-foss

    Official

    Operating procedure for the conflict-autoresolver agent: sweep open PostHog/posthog PRs that conflict with master, resolve the trivial conflicts (generated artifacts deterministically, source…

    721 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Help users debug PostHog Error Tracking stack-trace symbolication for any supported platform — JavaScript/TypeScript web, React Native (Hermes), Android (Proguard / R8), or iOS / macOS (dSYM).

    721 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Exploring Apm Traces

    PostHog/posthog-foss

    Official

    Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP.

    721 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Exploring LLM Traces

    PostHog/posthog-foss

    Official

    Debug and inspect LLM/AI agent traces using PostHog's MCP tools.

    721 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Investigate Metric

    PostHog/posthog-foss

    Official

    Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries.

    721 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Categories

Questions about Maintaining Python Tests

What does Maintaining Python Tests do?

Maintains existing pytest and Django test suites without weakening correctness. Maintaining Python Tests is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Maintains existing pytest and Django test suites without weakening correctness.

When should I use Maintaining Python Tests?

Maintaining Python Tests fits situations like: asked to reduce Python test runtime; investigate slow pytest families; remove stale migration tests; consolidate repeated setup.

How do I install Maintaining Python Tests in Claude Code?

Run `npx skills add PostHog/posthog-foss --skill maintaining-python-tests -a claude-code`. Or copy the skill folder (.agents/skills/maintaining-python-tests in PostHog/posthog-foss) into .claude/skills/maintaining-python-tests in your project. Claude Code loads it when a task matches its description.

How do I install Maintaining Python Tests in Codex?

Run `npx skills add PostHog/posthog-foss --skill maintaining-python-tests -a codex`. Or copy the skill folder (.agents/skills/maintaining-python-tests in PostHog/posthog-foss) into .agents/skills/maintaining-python-tests in your project. Codex loads it when a task matches its description.

Can I use Maintaining Python Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PostHog/posthog-foss --skill maintaining-python-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/maintaining-python-tests, .gemini/skills/maintaining-python-tests, .github/skills/maintaining-python-tests and .opencode/skills/maintaining-python-tests in your project.

What does Maintaining Python Tests need to run?

SKILL.md names no scripts, command-line tools or credentials: Maintaining Python Tests is instructions for the agent only. Our summary lists: Python 3.

Does Maintaining Python Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Maintaining Python Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Maintaining Python Tests use?

Maintaining Python Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Maintaining Python Tests use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.1k tokens, read only when the agent opens those files.

What are the alternatives to Maintaining Python Tests?

Skills that share tags, products or a category with Maintaining Python Tests: Pytest (mathiasertl/django-ca, 158 stars), Pytest (bobmatnyc/claude-mpm, 155 stars), Run And Verify (aropan/clist, 439 stars) and Python Dev (doccker/cc-use-exp, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Maintaining Python Tests?

PostHog (a GitHub organization, an official publisher) maintains it in PostHog/posthog-foss, which has 721 GitHub stars. The repository holds 213 skills in this directory. The repository was last updated on October 7, 2026.

Source: PostHog/posthog-foss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.