Agent skill

Fix Slow Endpoint

by vpcarlos in vpcarlos/profyle

Diagnose and fix a slow endpoint or request in a Python web app (FastAPI, Flask, Django, Tornado, any ASGI/WSGI framework) using real Profyle/VizTracer traces, then prove the fix by replaying the…

MITAuto-check passedBackend & APIs

Install Fix Slow Endpoint

skills CLI
$ npx skills add vpcarlos/profyle --skill fix-slow-endpoint -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vpcarlos/profyle fix-slow-endpoint --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vpcarlos/profyle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/profyle/claude .claude/skills/fix-slow-endpoint && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fix-slow-endpoint
GitHub stars
123
Token cost
~2.1k tokens
SKILL.md length
1,172 words
Files
2
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Diagnose and fix a slow endpoint or request in a Python web app (FastAPI, Flask, Django, Tornado, any ASGI/WSGI framework) using real Profyle/VizTracer traces, then prove the fix by replaying the…

  • Works in 7 steps: Check the setup → Find the endpoint and its traces → Get a trustworthy baseline → …
  • The user says an endpoint
  • SKILL.md covers 0. Check the setup, 1. Find the endpoint and its…, 2. Get a trustworthy baseline and 3. Diagnose, plus 3 more sections
  • Runs Python scripts from its folder; calls pip, uv and uvicorn

What it does

Fix Slow Endpoint is an agent skill from vpcarlos/profyle. Diagnose and fix a slow endpoint or request in a Python web app (FastAPI, Flask, Django, Tornado, any ASGI/WSGI framework) using real Profyle/VizTracer traces, then prove the fix by replaying the request and comparing traces. Use when the user says an endpoint, API call, page or request is slow, has high latency, times out, or asks where a bottleneck is.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `__init__.py`).

It sits in Backend & APIs, covering Backend development. It works with Python, Flask, FastAPI and Django. The repository describes itself as: Development tool for analysing and managing python traces. The licence is MIT.

When your agent uses it

  • The user says an endpoint
  • Request is slow
  • Has high latency
  • Asks where a bottleneck is

Example prompts

  • “/fix-slow-endpoint”

Requirements

  • Python 3
  • Docker

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Check the setup
  2. Find the endpoint and its traces
  3. Get a trustworthy baseline
  4. Diagnose
  5. Fix
  6. Verify
  7. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 60af2de. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • uv
    • uvicorn
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fix Slow Endpoint loads about 2.1k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 1,172 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vpcarlos/profyle at commit 60af2de, republished under its MIT licence (© vpcarlos). 1,172 words, ~2,076 tokens.

Download SKILL.mdSave it as .claude/skills/fix-slow-endpoint/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
fix-slow-endpoint
description
Diagnose and fix a slow endpoint or request in a Python web app (FastAPI, Flask, Django, Tornado, any ASGI/WSGI framework) using real Profyle/VizTracer traces, then prove the fix by replaying the request and comparing traces. Use when the user says an endpoint, API call, page or request is slow, has high latency, times out, or asks where a bottleneck is.

Fix a slow endpoint with Profyle traces

Work from measurements, not intuition: every claim you make about where time goes must come from a trace, and every fix must be verified with a new trace of the same request.

The profyle MCP server provides: doctor, slowest_endpoints, list_traces, analyze_trace, get_call_details, get_function_source, replay_request, compare_traces. All durations are in milliseconds.

0. Check the setup

Call doctor first. It reports which trace database is read, whether traces exist, which app wrote them and how (profyle run or middleware), and whether the app is running. If everything is ✓, go to step 1.

If there are no traces or the app is not running, get the app running under Profyle without editing the user's code:

  1. Make sure Profyle is installed in the project's environment (pip install profyle, uv add --dev profyle, ...). Ask before installing anything.
  2. Find how the project starts its dev server: README, Makefile, pyproject.toml scripts, Procfile, docker-compose, manage.py, or an app object such as main:app. Prefer the command with auto-reload: uvicorn main:app --reload, flask --app app run --debug, python manage.py runserver.
  3. Propose the command prefixed with profyle run, for example profyle run uvicorn main:app --reload, and ask the user to confirm it (and the port). profyle run adds tracing to FastAPI, Starlette, Flask, Django, Tornado (also under gunicorn's tornado worker) and any ASGI app served by uvicorn. If the app already uses ProfyleMiddleware, run the command as is.
  4. With the user's agreement, start it as a background process so it keeps running. It prints profyle ▸ tracing requests ... at start-up, then one line per request with the trace id and its main finding (profyle ▸ GET /orders 245 ms · #12 · repeated: list_orders → get_customer ×100 (83%)). Read those lines: they confirm tracing works.
  5. Trigger the slow request: with curl if it is a GET that needs no login, otherwise ask the user to do it in the app. Then call doctor again.

If the app runs where you cannot start it (a container, a remote machine), explain the setup to the user instead: profyle run <their command>, or ProfyleMiddleware from profyle.asgi / profyle.wsgi (or "profyle.django.ProfyleMiddleware" in MIDDLEWARE), with the traces database shared with this project (<project>/.profyle/profile.db, or PROFYLE_DB).

1. Find the endpoint and its traces

  • If the user named the endpoint, call list_traces(name_contains=...). Otherwise call slowest_endpoints and confirm with the user which one to work on.
  • Listings show each trace's main finding (for example repeated: a → b ×100 (83%)). It is a starting point, not the diagnosis: confirm it with analyze_trace. A first, much slower trace usually shows warm-up (imports, regex compiling) instead.

2. Get a trustworthy baseline

  • The first request after the server starts includes warm-up (imports, regex compiling, connection setup). Do not diagnose that trace if a later one exists.
  • For a GET request, call replay_request(trace_id, times=3) before changing anything. This gives warm traces and a median baseline, and confirms that replay works (the app is reachable and credentials are not missing). If the app is not reachable, start it as in step 0.
  • Never replay POST, PUT, PATCH or DELETE without the user's explicit permission. These requests may create or delete data. Ask first; only then pass allow_unsafe_method=true.
  • Auth headers and cookies are redacted when recording. If the replay returns 401 or 403, ask the user for a token or session for local testing and pass it in headers. Never print credentials back.
Show full SKILL.md (609 more words)Show less

3. Diagnose

Call analyze_trace on a representative warm trace and read it in this order:

  1. Critical path per thread. Sync endpoints often run in a worker thread, so the handler can appear under a second thread.
  2. Repeated calls from the same caller. parent → callee ×N together with different sample arguments (id=1 | id=2 | id=3) is the signature of an N+1.
  3. Top self time. Look for where the time actually goes: waiting on I/O (sleep, socket, DB driver, HTTP client) or computing.
  4. Your code by inclusive time. This is the code that can be changed.

Then drill down:

  • get_call_details(trace_id, function) shows the callers, callees and slowest invocations with arguments and return values.
  • Read the real file in the repository with your file tools. get_function_source shows the code as it was when the trace was recorded; the repository is the source of truth.

Typical root causes and fixes:

Signal in the traceLikely causeFix
Same query or fetch ×N with different idsN+1One batched query (IN, join), ORM selectinload/joinedload, select_related/prefetch_related, or a dataloader
time.sleep, requests, or a sync DB driver inside an async defBlocking the event loopUse an async client, or make the endpoint def / use run_in_threadpool
Several independent HTTP or DB calls in sequenceSequential I/ORun them concurrently (asyncio.gather), reuse a client/session for connection pooling
Same function with the same arguments and the same result many timesRedundant workHoist it out of the loop, memoize, or cache per request
Connection or client created on every requestMissing poolingCreate it once at startup and reuse it
Large time in serialization (jsonable_encoder, pydantic validation)Payload too bigPaginate, select fewer fields, return a response_model tailored to the endpoint
Pure-Python loop with a huge call countCPU hot loopBetter algorithm or data structure, builtins, or vectorization

Caveat: tracing adds about 1µs per function call, so functions with very large call counts look slower than they are untraced. Waits (sleep, network, DB) are measured accurately.

Before you edit anything, tell the user the root cause in two or three sentences, with the evidence: ms, % of the request, call counts and file:line.

4. Fix

  • Make the smallest change that removes the bottleneck. Keep the response identical: same status, same data, same order. Never make the endpoint "faster" by returning less, changing defaults such as page size, or skipping work the caller relies on.
  • Run the project's existing tests for the code you touched.

5. Verify

  • Make sure the app is running the new code. With auto-reload, wait a moment; if you started it without auto-reload, restart it; otherwise ask the user to restart it.
  • Call replay_request(baseline_trace_id, times=3). Read two columns before the timings:
    • status: a status change means the fix broke something.
    • body: identical is the goal. same structure, values differ is fine when the replay notes that values already vary between runs (timestamps, ids). DIFFERENT means the endpoint now returns other data: fix that before claiming any speed-up. If the baseline trace has no recorded body (Flask), compare against the first baseline replay instead.
  • Call compare_traces(baseline_id, new_id). Its response_body and status fields must hold. Then confirm that the bottleneck went away (for example, the N+1 callee drops from ×100 calls to ×1). Use analyze_trace on the new trace to see what dominates now.
  • If the improvement is small or the time moved somewhere else, go back to step 3.

6. Report

Finish with a short summary:

  • Root cause, with file:line.
  • The change.
  • Before → after, using median ms and the call counts that changed.
  • Behavior check: status and response body (identical, or same structure).
  • What dominates the request now, if more can be gained.

© vpcarlos, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in profyle/claude of vpcarlos/profyle.

  • SKILL.md
  • __init__.py

Open the folder on GitHubat commit 60af2de

Compare with similar skills

Fix Slow Endpoint next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fix Slow Endpoint compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fix Slow Endpoint this skillvpcarlos/profyle123—~2.1kAutomated safety check: PassMIT
Framework Migration AssistantArabelaTso/Skills-4-SE253—~1.9kAutomated safety check: PassApache-2.0
Sentry Python SDKgetsentry/sentry-for-ai268—~4.1kAutomated safety check: PassApache-2.0
Python Appservice Deploymicrosoft/GitHub-Copilot-for-Azure2551 repos~688Automated safety check: PassMIT
Python Devdoccker/cc-use-exp1.1k—~790Automated safety check: PassCustom licence
Fastapi Templatesjh941213/my-cc-harness126—~945Automated safety check: PassNone

Similar skills

  • Framework Migration Assistant

    ArabelaTso/Skills-4-SE

    Automatically migrate Python web applications between frameworks (Flask → FastAPI, Django → FastAPI).

    253 GitHub stars~1.9k tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed
  • Sentry Python SDK

    getsentry/sentry-for-ai

    Official

    Full Sentry SDK setup for Python. An agent skill from getsentry/sentry-for-ai.

    268 GitHub stars~4.1k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Python Appservice Deploy

    microsoft/GitHub-Copilot-for-Azure

    Official

    Deploy Python (Flask/Django/FastAPI) code to Azure App Service Linux.

    255 GitHub starsUsed in 1 repo~688 tokens
    Backend & APIsAuto-check passed
  • Python Dev

    doccker/cc-use-exp

    Python 开发规范。当用户操作 .py、pyproject.toml、requirements.txt、setup.py 文件, 或涉及 FastAPI、Django、Flask、pytest、asyncio 开发时触发。

    1.1k GitHub stars~790 tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed
  • Fastapi Templates

    jh941213/my-cc-harness

    Production-grade FastAPI project creation and setup guide. An agent skill from jh941213/my-cc-harness.

    126 GitHub stars~945 tokensUpdated 2 mo ago
    Backend & APIsAuto-check passed
  • Python Web App Security Audit

    aiskillstore/marketplace

    Run defensive pre-release security tests for Python web applications.

    430 GitHub stars~1.2k tokensUpdated today
    Backend & APIsAuto-check: notes

Categories

Questions about Fix Slow Endpoint

What does Fix Slow Endpoint do?

Diagnose and fix a slow endpoint or request in a Python web app (FastAPI, Flask, Django, Tornado, any ASGI/WSGI framework) using real Profyle/VizTracer traces, then prove the fix by replaying the…. Fix Slow Endpoint is an agent skill from vpcarlos/profyle. Diagnose and fix a slow endpoint or request in a Python web app (FastAPI, Flask, Django, Tornado, any ASGI/WSGI framework) using real Profyle/VizTracer traces, then prove the fix by replaying the request and comparing traces.

When should I use Fix Slow Endpoint?

Fix Slow Endpoint fits situations like: the user says an endpoint; request is slow; has high latency; asks where a bottleneck is.

How do I install Fix Slow Endpoint in Claude Code?

Run `npx skills add vpcarlos/profyle --skill fix-slow-endpoint -a claude-code`. Or copy the skill folder (profyle/claude in vpcarlos/profyle) into .claude/skills/fix-slow-endpoint in your project. Claude Code loads it when a task matches its description.

How do I install Fix Slow Endpoint in Codex?

Run `npx skills add vpcarlos/profyle --skill fix-slow-endpoint -a codex`. Or copy the skill folder (profyle/claude in vpcarlos/profyle) into .agents/skills/fix-slow-endpoint in your project. Codex loads it when a task matches its description.

Can I use Fix Slow Endpoint in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vpcarlos/profyle --skill fix-slow-endpoint -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fix-slow-endpoint, .gemini/skills/fix-slow-endpoint, .github/skills/fix-slow-endpoint and .opencode/skills/fix-slow-endpoint in your project.

What does Fix Slow Endpoint need to run?

Going by SKILL.md and its folder, Fix Slow Endpoint needs Python for the scripts in its folder and the command-line tools its instructions call (pip, uv, uvicorn and python). Our summary lists: Python 3; Docker.

Does Fix Slow Endpoint access the network?

SKILL.md contains no URLs. Its commands use pip and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fix Slow Endpoint safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fix Slow Endpoint use?

Fix Slow Endpoint is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fix Slow Endpoint use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fix Slow Endpoint?

Skills that share tags, products or a category with Fix Slow Endpoint: Framework Migration Assistant (ArabelaTso/Skills-4-SE, 253 stars), Sentry Python SDK (getsentry/sentry-for-ai, 268 stars), Python Appservice Deploy (microsoft/GitHub-Copilot-for-Azure, 255 stars) and Python Dev (doccker/cc-use-exp, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fix Slow Endpoint?

vpcarlos (a GitHub user) maintains it in vpcarlos/profyle, which has 123 GitHub stars. The repository was last updated on October 5, 2026.

Source: vpcarlos/profyle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.