Official agent skill

Improving MCP Tools

by PostHog in PostHog/posthog-foss

Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded…

OfficialMITAuto-check passedAgent Workflows

Install Improving MCP Tools

skills CLI
$ npx skills add PostHog/posthog-foss --skill improving-mcp-tools -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PostHog/posthog-foss improving-mcp-tools --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PostHog/posthog-foss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/improving-mcp-tools .claude/skills/improving-mcp-tools && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
improving-mcp-tools
GitHub stars
721
Token cost
~1.5k tokens
SKILL.md length
673 words
Files
2 (incl. references)
Skills in repo
213
Repo updated
First seen
Licence
MIT

At a glance

Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded…

  • Works in 6 steps: Measure. Run the harness for a baseline.… → Pick one issue. Rank by reach ×… → Fix, bounded. Only files inside the… → …
  • Asked to improve my MCP
  • SKILL.md covers The objective function, One iteration, Hard guardrails and Failure modes to expect
  • Calls pnpm; needs LIVE_MCP_TOKEN

What it does

Improving MCP Tools is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded fix, and keeps it only if before/after scores improve. Use when asked to "improve my MCP", run an MCP improvement campaign, fix tool discoverability or descriptions based on evidence, or prepare an eval-backed PR for a tool change. Every shipped change must carry eval evidence; guardrails below are hard rules.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/campaign-journal.md`).

It sits in Agent Workflows, covering MCP servers, LLM evaluation and Autonomous loops. It works with Model Context Protocol and PostHog. The repository describes itself as: PostHog FOSS is a read-only mirror of PostHog, with all proprietary code removed. NOTE: This repo is synced automatically from the main PostHog repo. Please raise any issues and… The licence is MIT.

When your agent uses it

  • Asked to improve my MCP
  • Run an MCP improvement campaign
  • Fix tool discoverability
  • Descriptions based on evidence

Example prompts

  • “improve my MCP”
  • “/improving-mcp-tools”

Requirements

  • A credential in LIVE_MCP_TOKEN

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Measure. Run the harness for a baseline. Pull production evidence with
  2. Pick one issue. Rank by reach × severity. Skip anything the journal
  3. Fix, bounded. Only files inside the allowlist (below). Typical fixes
  4. Validate. Re-run the affected benchmark slice plus a no-regression
  5. Ship. One PR per iteration with before/after scores in the body (format
  6. Journal. Append the iteration record before ending the pass.

What it can do on your machine

Read from SKILL.md and the folder at commit 2c48221. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LIVE_MCP_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Improving MCP Tools loads about 1.5k tokens when it runs, and up to ~2k if it reads all its reference files. Until then it costs about 133 tokens; SKILL.md has 673 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~133
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PostHog/posthog-foss at commit 2c48221, republished under its MIT licence (© PostHog). 673 words, ~1,458 tokens.

Download SKILL.mdSave it as .claude/skills/improving-mcp-tools/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
improving-mcp-tools
description
Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded fix, and keeps it only if before/after scores improve. Use when asked to "improve my MCP", run an MCP improvement campaign, fix tool discoverability or descriptions based on evidence, or prepare an eval-backed PR for a tool change. Every shipped change must carry eval evidence; guardrails below are hard rules.

Improving MCP tools

An MCP server gets better only in ways you can measure. This skill is the campaign procedure: score the current agent experience, fix the biggest problem, re-score, and only ship changes the numbers justify. It is the operating manual for the "improve my MCP" loop — one iteration per pass, journaled so a later iteration (or a different agent) can resume without repeating work.

The objective function

services/mcp/evals/ is the harness. benchmark/tasks.yaml is a fixed set of agent tasks with expected_tools and success_criteria; scores are only comparable across runs of the same benchmark version.

  • Probe mode (deterministic, no LLM): LIVE_MCP_URL=... LIVE_MCP_TOKEN=... pnpm exec tsx evals/runner/probe.ts --out score.json from services/mcp/. Reports tool-presence misses (discoverability), probe failures, and latency p50/p95. Non-zero exit = regression.
  • Agent mode (LLM replay + judge): scores task success and tool-selection accuracy. Use it for description/discoverability changes — probes cannot detect that an agent picks the wrong tool.

Run the harness against a seeded local or devbox stack, never against a customer project. Local recipe: NODE_ENV=development PORT=9876 POSTHOG_API_BASE_URL=http://localhost:8000 pnpm dev:hono, personal API key as LIVE_MCP_TOKEN.

One iteration

  1. Measure. Run the harness for a baseline. Pull production evidence with the MCP analytics tools (query-mcp-tool-stats, query-mcp-tool-failures, query-mcp-tool-descriptions, query-mcp-tool-sample-intents) and the lenses in the signals scout cookbook (products/signals/skills/signals-scout-mcp-tool-calls/references/queries.md): failure leaderboard, retry/struggle, latency, intents that matched no tool.
  2. Pick one issue. Rank by reach × severity. Skip anything the journal shows with two failed attempts. One issue per iteration — a PR that fixes three things can't be attributed to any of them when scores move.
  3. Fix, bounded. Only files inside the allowlist (below). Typical fixes: sharpen a tool description so the right intent finds it, tighten an input schema that agents keep getting wrong, fix an annotation, update a skill.
  4. Validate. Re-run the affected benchmark slice plus a no-regression sample. Keep the change only if the target metric improves and nothing else degrades. A discarded change is a normal outcome — journal it and move on.
  5. Ship. One PR per iteration with before/after scores in the body (format in references/campaign-journal.md). Keep it stampable: ≤400 changed lines, only files inside the allowlist below, request a stamphog review (MCP first, label fallback, see /merging-prs). Autonomy level comes from the campaign config — default is draft PR for human review; only arm auto-merge when the operator has explicitly enabled the self-driving experiment (see guardrails).
  6. Journal. Append the iteration record before ending the pass.
Show full SKILL.md (273 more words)Show less

Hard guardrails

These are not suggestions; violating any of them ends the campaign pass.

  • Allowlist — a campaign PR may only touch: products/*/mcp/tools.yaml, products/*/skills/**, services/mcp/evals/**, the codegen outputs of pnpm generate-tools / scaffold-yaml (services/mcp/src/tools/generated/** and services/mcp/schema/generated-tool-definitions.json), and docs. Anything else (handler code, package manifests, workflows, migrations, auth paths) → stop and hand the finding to a human as a draft PR or report instead.
  • Read-only against data. The harness and all production queries are read-only. Never create, mutate, or delete customer-visible objects while measuring.
  • Evidence or it didn't happen. No PR without a baseline score, an after score, and the exact harness commands used.
  • Benchmark integrity. Never edit benchmark/tasks.yaml in the same PR as a fix it validates — changing the exam and the answer together proves nothing. Benchmark changes are their own PR and bump version.
  • Budgets. Respect the operator's iteration/token/PR caps (default: stop after 3 open unmerged campaign PRs). Two failed attempts on an issue parks it permanently.
  • Kill switch. If the campaign config, its feature flag, or the operator says stop — stop mid-iteration, journal state, end cleanly.

Failure modes to expect

  • A description change that helps one intent can steal traffic from the right tool for another — that's why the no-regression sample is mandatory. The intent-cluster snapshot's tool_overlaps (see exploring-mcp-intent-clusters) lists exactly which pairs compete for which intents: snapshot it before a description rewrite and recompute after, and treat a capture shift in an overlapping pair as the regression signal.
  • Probe latency varies with stack warmth; compare medians across ≥3 runs before attributing a latency change to your fix.
  • Tool-presence misses can be feature-flag gating, not catalog absence — check getToolsForFeatures gating before "fixing" discoverability.

© PostHog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .agents/skills/improving-mcp-tools of PostHog/posthog-foss.

  • SKILL.md
  • references/campaign-journal.md

Open the folder on GitHubat commit 2c48221

Compare with similar skills

Improving MCP Tools next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Improving MCP Tools compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Improving MCP Tools this skillPostHog/posthog-foss721—~1.5kAutomated safety check: PassMIT
Conventions MCPstella/stella258—~2.7kAutomated safety check: PassApache-2.0
Agents Best PracticesDenisSergeevitch/agents-best-practices2.4k—~7.4kAutomated safety check: PassMIT
Ouroboros Evolve LoopQ00/ouroboros6.2k—~3.2kAutomated safety check: PassMIT
Hermes Mission ControlTh0rgal/sandboxed.sh515—~8.6kAutomated safety check: PassNone
Agent Native Architecturesandgardenhq/sgai137—~2kAutomated safety check: PassCustom licence

Similar skills

  • Conventions MCP

    stella/stella

    Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource.

    258 GitHub stars~2.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Agents Best Practices

    DenisSergeevitch/agents-best-practices

    A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.

    2.4k GitHub stars~7.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Starts, monitors or rewinds an evolutionary development loop that refines an ontology and acceptance criteria generation by generation until it converges, using the Ouroboros MCP tools.

    6.2k GitHub stars~3.2k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Hermes Mission Control

    Th0rgal/sandboxed.sh

    Teaches Hermes to monitor and steer long-running sandboxed.sh missions: spot where a model is stuck, switch backends or models between turns, and send targeted hints.

    515 GitHub stars~8.6k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Agent Native Architecture

    sandgardenhq/sgai

    Build AI agents using prompt-native architecture where features are defined in prompts, not code.

    137 GitHub stars~2k tokensUpdated 16 days ago
    Agent WorkflowsAuto-check passed
  • Reference guide to the Ouroboros commands and agents, covering interviews, seed specs, evaluation, lateral-thinking personas and the evolutionary loop.

    6.2k GitHub stars~1.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from PostHog/posthog-foss

All 213 skills in this repo
  • Authoring Log Alerts

    PostHog/posthog-foss

    Official

    Author useful, low-noise log alerts on services in a PostHog project.

    721 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Autoresolving PR Conflicts

    PostHog/posthog-foss

    Official

    Operating procedure for the conflict-autoresolver agent: sweep open PostHog/posthog PRs that conflict with master, resolve the trivial conflicts (generated artifacts deterministically, source…

    721 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Official

    Help users debug PostHog Error Tracking stack-trace symbolication for any supported platform — JavaScript/TypeScript web, React Native (Hermes), Android (Proguard / R8), or iOS / macOS (dSYM).

    721 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Exploring Apm Traces

    PostHog/posthog-foss

    Official

    Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP.

    721 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Exploring LLM Traces

    PostHog/posthog-foss

    Official

    Debug and inspect LLM/AI agent traces using PostHog's MCP tools.

    721 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Investigate Metric

    PostHog/posthog-foss

    Official

    Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries.

    721 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Improving MCP Tools

What does Improving MCP Tools do?

Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded…. Improving MCP Tools is an agent skill from PostHog/posthog-foss, published by the product's own GitHub organization. Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded fix, and keeps it only if before/after scores improve.

When should I use Improving MCP Tools?

Improving MCP Tools fits situations like: asked to improve my MCP; run an MCP improvement campaign; fix tool discoverability; descriptions based on evidence.

How do I install Improving MCP Tools in Claude Code?

Run `npx skills add PostHog/posthog-foss --skill improving-mcp-tools -a claude-code`. Or copy the skill folder (.agents/skills/improving-mcp-tools in PostHog/posthog-foss) into .claude/skills/improving-mcp-tools in your project. Claude Code loads it when a task matches its description.

How do I install Improving MCP Tools in Codex?

Run `npx skills add PostHog/posthog-foss --skill improving-mcp-tools -a codex`. Or copy the skill folder (.agents/skills/improving-mcp-tools in PostHog/posthog-foss) into .agents/skills/improving-mcp-tools in your project. Codex loads it when a task matches its description.

Can I use Improving MCP Tools in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PostHog/posthog-foss --skill improving-mcp-tools -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/improving-mcp-tools, .gemini/skills/improving-mcp-tools, .github/skills/improving-mcp-tools and .opencode/skills/improving-mcp-tools in your project.

What does Improving MCP Tools need to run?

Going by SKILL.md and its folder, Improving MCP Tools needs the command-line tools its instructions call (pnpm) and credentials named LIVE_MCP_TOKEN. Our summary lists: A credential in LIVE_MCP_TOKEN.

Does Improving MCP Tools access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Improving MCP Tools safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Improving MCP Tools use?

Improving MCP Tools is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Improving MCP Tools use?

About 1.5k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 502 tokens, read only when the agent opens those files.

What are the alternatives to Improving MCP Tools?

Skills that share tags, products or a category with Improving MCP Tools: Conventions MCP (stella/stella, 258 stars), Agents Best Practices (DenisSergeevitch/agents-best-practices, 2.4k stars), Ouroboros Evolve Loop (Q00/ouroboros, 6.2k stars) and Hermes Mission Control (Th0rgal/sandboxed.sh, 515 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Improving MCP Tools?

PostHog (a GitHub organization, an official publisher) maintains it in PostHog/posthog-foss, which has 721 GitHub stars. The repository holds 213 skills in this directory. The repository was last updated on October 7, 2026.

Source: PostHog/posthog-foss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.