Agent skill

Python Performance Optimization

by diegosouzapw in diegosouzapw/awesome-omni-skills

Python Performance Optimization workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

MITAuto-check passedDevelopment

Install Python Performance Optimization

skills CLI
$ npx skills add diegosouzapw/awesome-omni-skills --skill python-performance-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install diegosouzapw/awesome-omni-skills python-performance-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/diegosouzapw/awesome-omni-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills_omni/python-performance-optimization .claude/skills/python-performance-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
python-performance-optimization
GitHub stars
159
Token cost
~2.7k tokens
SKILL.md length
1,149 words
Files
20 (incl. scripts, references, assets)
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Python Performance Optimization workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

  • Works in 5 steps: define the user-visible performance… → collect a reproducible baseline, → choose the profiler that matches the… → …
  • The user needs to profile and optimize Python code using cProfile
  • SKILL.md covers Overview, When to Use, Operating Table and Workflow, plus 5 more sections
  • Calls python

What it does

Python Performance Optimization is an agent skill from diegosouzapw/awesome-omni-skills. Python Performance Optimization workflow skill. Use this skill when the user needs to profile and optimize Python code using cProfile, memory profilers, benchmark discipline, and performance best practices. Use when debugging slow Python code, isolating bottlenecks, or improving application performance with measurement-first evidence before and after each change.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files, including scripts, reference files and assets (for example `ATTRIBUTION.md`, `OMNI_ENHANCED.json` and `ORIGIN.md`).

It sits in Development, covering Performance optimization. It works with Python. The repository describes itself as: Public repository of AI coding skills, curated improved best-practice skills, and runtime surfaces for CLI, API, MCP, and A2A. The licence is MIT.

When your agent uses it

  • The user needs to profile and optimize Python code using cProfile
  • Memory profilers
  • Benchmark discipline
  • Performance best practices

Example prompts

  • “/python-performance-optimization”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. define the user-visible performance problem,
  2. collect a reproducible baseline,
  3. choose the profiler that matches the symptom,
  4. make one bounded change at a time,
  5. re-measure under the same conditions.

What it can do on your machine

Read from SKILL.md and the folder at commit c3af004. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Python Performance Optimization loads about 2.7k tokens when it runs, and up to ~4.4k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 1,149 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from diegosouzapw/awesome-omni-skills at commit c3af004, republished under its MIT licence (© diegosouzapw). 1,149 words, ~2,738 tokens.

Download SKILL.mdSave it as .claude/skills/python-performance-optimization/SKILL.md (or your agent's skills folder). This skill also uses 19 other files; get the full folder from GitHub.
name
python-performance-optimization
description
Python Performance Optimization workflow skill. Use this skill when the user needs to profile and optimize Python code using cProfile, memory profilers, benchmark discipline, and performance best practices. Use when debugging slow Python code, isolating bottlenecks, or improving application performance with measurement-first evidence before and after each change.
version
0.0.1
category
development
tags
python-performance-optimization, profile, optimize, python, cprofile, tracemalloc, benchmarking, performance, omni-enhanced
complexity
advanced
risk
caution
tools
codex-cli, claude-code, cursor, gemini-cli, opencode
source
omni-team
author
Omni Skills Team
date_added
2026-04-15
date_updated
2026-04-19

Python Performance Optimization

Overview

Use this skill to investigate and improve Python performance with a measurement-first workflow.

The core operating model is:

  1. define the user-visible performance problem,
  2. collect a reproducible baseline,
  3. choose the profiler that matches the symptom,
  4. make one bounded change at a time,
  5. re-measure under the same conditions.

This skill is for real diagnosis, not guesswork. Do not promise speedups from caching, vectorization, concurrency, or refactoring unless measurements on representative workloads show improvement.

The baseline workflow works with the Python standard library alone. Optional third-party profilers such as py-spy, pyperf, and Scalene can improve diagnosis when deterministic profiling is not enough, but they are not required.

Open these support files when needed:

  • references/runtime-practices.md for profiler selection, benchmark hygiene, and memory/concurrency rules.
  • examples/implementation-example.md for a concrete before/after optimization sequence.
  • scripts/validate-runtime.py for repeat timing, cProfile summaries, and optional tracemalloc allocation diffs.

When to Use

Use this skill when the task involves one or more of these:

  • Python code is slower than expected and the bottleneck is not yet known.
  • A recent change increased latency, CPU time, or memory allocation.
  • A user wants evidence before deciding whether to optimize algorithm, data structure, caching, I/O, or concurrency.
  • Memory growth needs investigation with allocation tracing rather than assumptions.
  • A proposed optimization needs before/after proof under the same interpreter, dependencies, and representative inputs.

Do not use this skill as the primary workflow when:

  • the issue is mainly database tuning, network architecture, operating-system tuning, or infrastructure scaling,
  • the operator cannot run representative workloads or gather measurements safely,
  • the task asks for speculative “make it faster” advice without code, workload details, or acceptance criteria.

Operating Table

SituationStart hereWhy it matters
You only know that the app is “slow”references/runtime-practices.mdHelps classify the symptom as CPU, wall-clock, memory, native-extension, or concurrency related before choosing tools
You need a baseline and reproducible evidencescripts/validate-runtime.pyCollects repeat timing, cProfile data, and optional tracemalloc diffs in one standard-library-first workflow
You need a worked example before touching user codeexamples/implementation-example.mdShows a complete baseline → profile → optimize → re-measure sequence
cProfile output is confusing or incompletereferences/runtime-practices.mdExplains cumulative vs total time and when to switch to sampling or mixed CPU/memory profilers
Memory appears to keep growingreferences/runtime-practices.mdDistinguishes Python allocation growth, cache growth, transient peaks, GC behavior, and RSS confusion

Workflow

  1. Define the performance question.

    • Capture the metric that matters: request latency, batch runtime, throughput, peak allocation, or allocation growth over time.
    • Record the relevant workload, input size, interpreter version, dependency set, and platform assumptions.
  2. Create a stable baseline before changing code.

    • Prefer repeated measurements over one-off timing.
    • Keep inputs, logging level, interpreter, and environment fixed.
    • If possible, test representative workloads instead of toy microbenchmarks.
    • Use scripts/validate-runtime.py or timeit/pyperf for repeatable timing.
  3. Choose the profiler that matches the symptom.

    • Use cProfile first for Python-call hotspot analysis.
    • Use tracemalloc when the problem is allocation growth, unexpected memory pressure, or peak Python allocations.
    • Use py-spy optionally when you need low-overhead sampling of a live or hard-to-instrument process.
    • Use Scalene optionally when you need stronger attribution for Python time vs native time or mixed CPU/memory behavior.
    • See references/runtime-practices.md for the symptom-to-tool matrix.
  4. Inspect the evidence carefully.

    • For cProfile, compare cumtime and tottime rather than scanning only call counts.
    • For memory, compare tracemalloc snapshots around the suspect operation.
    • Confirm whether the hotspot is algorithmic work, repeated conversion/parsing, I/O waiting, synchronization, serialization, or cache churn.
  5. Apply one bounded optimization at a time. Examples:

    • replace repeated O(n) lookups with a dict or set,
    • reduce unnecessary allocations or copies,
    • move repeated pure-function work behind a bounded cache,
    • batch I/O instead of performing many tiny operations,
    • consider process-based parallelism for CPU-bound work only after measuring task granularity and overhead.
  6. Re-measure under the same conditions.

    • Use the same dataset, interpreter, and environment.
    • Report baseline vs changed results with units and repetition count.
    • Note tradeoffs such as higher memory use, reduced readability, or changed concurrency behavior.
  7. Stop if evidence does not support the change.

    • Revert speculative optimizations.
    • Escalate to a different tool or a broader architecture review when the current measurement method cannot explain the bottleneck.

Troubleshooting

Show full SKILL.md (460 more words)Show less
Benchmarks are noisy or contradictory

Check these first:

  • input sizes changed between runs,
  • debug logging or tracing is enabled in only one run,
  • startup/import time is being mixed into the steady-state measurement,
  • too few repetitions were used,
  • the benchmark is too small and mostly measures noise,
  • the microbenchmark does not represent the real workload.

If available, use pyperf for stronger calibration and noise control. Otherwise, increase repetitions and keep the environment stable.

cProfile does not explain the observed latency

Possible reasons:

  • the process spends time blocked on I/O,
  • important work happens inside native extensions,
  • a live production-like process is difficult to instrument directly,
  • profiler overhead distorts extremely small hot paths.

In those cases, keep the cProfile evidence but consider optional tools such as py-spy or Scalene for sampling or mixed attribution.

Memory keeps growing

Do not assume “memory leak” immediately.

Check whether the growth is due to:

  • retained Python objects,
  • intentionally growing caches,
  • temporary peak allocations,
  • allocator behavior that affects RSS differently from Python-level allocation traces,
  • GC-sensitive object graphs.

Use tracemalloc snapshot comparisons first. If caching is involved, require bounded size and a measurable reason for the memory tradeoff.

The parallel version is slower

Common causes:

  • threads were used for CPU-bound Python work,
  • process startup and teardown costs dominate,
  • task size is too small,
  • serialization or pickling overhead is high,
  • shared resources create contention,
  • the benchmark measures startup rather than steady-state execution.

Validate concurrency changes with representative workloads. Threads are usually for I/O overlap; CPU-bound work may need algorithmic improvement or process-based parallelism, but only when the workload is large enough to amortize overhead.

A cache improved speed but memory got worse

That can be a valid tradeoff or a bad one. Confirm:

  • cache size is bounded,
  • the function is actually called with repeated inputs,
  • hit rate is high enough to matter,
  • memory growth remains acceptable,
  • invalidation or staleness risk is understood.

Examples

For a concrete end-to-end example, open examples/implementation-example.md.

A minimal local workflow looks like this:

bash
python scripts/validate-runtime.py --module target_module --callable run_case --repeat 7 --number 20 --sort cumtime --top 20

Memory-focused run:

bash
python scripts/validate-runtime.py --module target_module --callable run_case --repeat 5 --number 10 --tracemalloc --snapshot-diff-limit 15

Then compare the report before and after one code change. Do not stack multiple unrelated optimizations into the same measurement cycle.

Additional Resources

  • references/runtime-practices.md
  • Python cProfile / profile documentation
  • Python timeit documentation
  • Python tracemalloc documentation
  • Python functools documentation for bounded caching primitives
  • Python concurrent.futures and multiprocessing documentation
  • Optional: pyperf, py-spy, and Scalene project documentation

Use a different skill or workflow when the root issue is primarily:

  • database query planning or indexing,
  • distributed-system latency or queueing,
  • container or Kubernetes resource tuning,
  • front-end rendering performance,
  • OS-level or kernel-level profiling.

Limitations

  • This skill does not replace environment-specific validation.
  • Measurements taken on toy data or unstable environments may produce misleading conclusions.
  • Python allocation traces do not fully explain OS-level RSS behavior.
  • Optional third-party profilers are helpful, but the workflow must remain functional without them.

© diegosouzapw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 19 other files (scripts, references, assets) in skills_omni/python-performance-optimization of diegosouzapw/awesome-omni-skills.

  • SKILL.md
  • ATTRIBUTION.md
  • OMNI_ENHANCED.json
  • ORIGIN.md
  • agents/omni-import-router.md
  • assets/omni-import-source-manifest.json
  • examples/implementation-example.md
  • examples/omni-import-operator-packet.md
  • examples/omni-import-prompt-template.md
  • metadata.json
  • references/omni-import-checklist.md
  • references/omni-import-playbook.md
  • references/omni-import-rubric.md
  • references/omni-import-source-summary.md
  • references/runtime-practices.md
  • resources/implementation-playbook.md
  • … and 4 more

Open the folder on GitHubat commit c3af004

Compare with similar skills

Python Performance Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Python Performance Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Python Performance Optimization this skilldiegosouzapw/awesome-omni-skills159—~2.7kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
Pycrazyguitar/pysheeet8.2k—~886Automated safety check: PassMIT
Python Performance Optimizationwshobson/agents40k13 repos~814Automated safety check: PassMIT
Keybase RPC Log Analysiskeybase/client9.3k—~3kAutomated safety check: PassBSD-3-Clause
The Art of Debuggingstas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.0

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Py

    crazyguitar/pysheeet

    Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC.

    8.2k GitHub stars~886 tokensUpdated today
    DevelopmentAuto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    DevelopmentAuto-check passed
  • Captures a clean Keybase service log and analyzes it for redundant, duplicated or looping RPCs, then checks whether a caching fix reduced the calls.

    9.3k GitHub stars~3k tokensUpdated today
    DevelopmentAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Code Review Specialist

    luongnv89/claude-howto

    Reviews code for security, performance, quality and maintainability, using a checklist, a finding template and two metrics scripts.

    42k GitHub stars~764 tokensUpdated 7 days ago
    DevelopmentAuto-check passed

More from diegosouzapw/awesome-omni-skills

All 39 skills in this repo
  • Content Creator

    diegosouzapw/awesome-omni-skills

    Content Creator workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~4k tokensUpdated 3 mo ago
    Auto-check passed
  • Helm Chart Scaffolding

    diegosouzapw/awesome-omni-skills

    Helm Chart Scaffolding workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~2.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Prompt Engineering

    diegosouzapw/awesome-omni-skills

    Prompt Engineering Patterns workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~3.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Prompt Engineering Patterns

    diegosouzapw/awesome-omni-skills

    Prompt Engineering Patterns workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~4k tokensUpdated 3 mo ago
    Auto-check passed
  • Prompt Library

    diegosouzapw/awesome-omni-skills

    📝 Prompt Library workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~3.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Protocol Reverse Engineering

    diegosouzapw/awesome-omni-skills

    Protocol Reverse Engineering workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~3.8k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Categories

Questions about Python Performance Optimization

What does Python Performance Optimization do?

Python Performance Optimization workflow skill. An agent skill from diegosouzapw/awesome-omni-skills. Python Performance Optimization is an agent skill from diegosouzapw/awesome-omni-skills. Python Performance Optimization workflow skill.

When should I use Python Performance Optimization?

Python Performance Optimization fits situations like: the user needs to profile and optimize Python code using cProfile; memory profilers; benchmark discipline; performance best practices.

How do I install Python Performance Optimization in Claude Code?

Run `npx skills add diegosouzapw/awesome-omni-skills --skill python-performance-optimization -a claude-code`. Or copy the skill folder (skills_omni/python-performance-optimization in diegosouzapw/awesome-omni-skills) into .claude/skills/python-performance-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Python Performance Optimization in Codex?

Run `npx skills add diegosouzapw/awesome-omni-skills --skill python-performance-optimization -a codex`. Or copy the skill folder (skills_omni/python-performance-optimization in diegosouzapw/awesome-omni-skills) into .agents/skills/python-performance-optimization in your project. Codex loads it when a task matches its description.

Can I use Python Performance Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add diegosouzapw/awesome-omni-skills --skill python-performance-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/python-performance-optimization, .gemini/skills/python-performance-optimization, .github/skills/python-performance-optimization and .opencode/skills/python-performance-optimization in your project.

What does Python Performance Optimization need to run?

Going by SKILL.md and its folder, Python Performance Optimization needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Python Performance Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Python Performance Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Python Performance Optimization use?

Python Performance Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Python Performance Optimization use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Python Performance Optimization?

Skills that share tags, products or a category with Python Performance Optimization: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), Py (crazyguitar/pysheeet, 8.2k stars), Python Performance Optimization (wshobson/agents, 40k stars) and Keybase RPC Log Analysis (keybase/client, 9.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Python Performance Optimization?

diegosouzapw (a GitHub user) maintains it in diegosouzapw/awesome-omni-skills, which has 159 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on July 8, 2026.

Source: diegosouzapw/awesome-omni-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.