---
name: skill-optimization-study
description: Re-runnable measurement loop for agent-driven JetBrains MPS work over mps_mcp_* tools — baseline headless worker runs on fixed scenarios, server call log + transcripts, hotspot ranking, remedy classification (docs / server tool / offline script / online script / template), optional A/B. Use when tool descriptions or mps-* skills changed, before a release, or when agents seem slow or retry-prone on MPS tasks.
---

# Skill optimisation study (MPS MCP)

A repeatable procedure for finding where agents waste turns and tokens when driving MPS through
`mps_mcp_*`, and for choosing the cheapest fix. This skill is self-sufficient: the 2026-09 study documents
(`plugins/mcp-tools/docs/skill-script-automation-study.md` and its runbook) are history and
evidence, not required reading. Scripts and scenario prompts live in `plugins/mcp-tools/study/`;
if that directory is removed, move `scripts/` here and the prompts into `assets/`.

## Conventions used below

Agent shells reset the working directory per call, so every command uses absolute paths through
two variables — set them at the start of each Bash call (or export them in a wrapper script):

```
STUDY=/Users/vaclav/work/MPS/myMPS-fix/plugins/mcp-tools/study   # adjust to the checkout
RUNS=$HOME/MPSProjects/mcp-study/runs                            # evidence dir, outside the repo
```
`$RUNS/inventory.json` is a load-bearing name: `run_worker.sh` records its sha in every run's meta.

## Roles

- **Observer** (this session, Opus-class): orchestrates, never performs the MPS task, never tells
  workers they are measured, evaluates results read-only, writes the report. Owns the whole
  lifecycle: **synthesizes** empty projects (`scripts/new_study_project.py`), **opens** them via CLI
  (`mps-project-management`), **closes** them with `mps_mcp_close_project`, and **starts, restarts
  and shuts MPS down** with `scripts/mps_control.sh`. Before each swap, tell the user the absolute
  path about to close (if any) and the absolute path about to open — do not ask them to perform the
  swap and do not wait for approval.
- **Workers**: headless CLI processes (`claude -p` or `junie --task`), one per (scenario, model, run),
  launched by `study/scripts/run_worker.sh`. Evidence = their transcript + the server call log.
- **Human**: answers the gate questions, approves pushes, and dismisses MPS dialogs when a call
  returns `MODAL_BLOCKED` or a close/exit hangs on a confirmation. Does **not** open, close or
  create projects and does **not** restart MPS — those are observer actions now.

## Gate questions to ask before starting (use them verbatim)

Before question 2, run `python3 $STUDY/scripts/list_worker_models.py` and present `models` as a
multi-select. The orchestrator model is first and marked; list it first with "(Recommended)" (the
picker cannot preselect). Extra ids the user types are allowed. Ask question 2a after question 2, as
one single-select question per selected model (batch them, at most 4 per AskUserQuestion call). The
recommended level is the one the user-level settings give that model today: its
`settingsEffort.perModel` entry for the full model id, else `settingsEffort.default`. List it first
with "(Recommended)". Fill the other options from `effortLevels`, at most 4 in all: on Claude leave
`max` to "Other", unless `max` is the recommended level, in which case drop `low` instead. With no
settings level (or on Junie) there is no recommendation; the CLI's built-in default is unknown, so
the user picks. Always pin it, because an unpinned worker takes the observer's last `/effort` for
that model, which changes between rounds without a trace (round 19). Ask question 3 only when the
detected harness is Claude; Junie non-interactive has no `bypassPermissions` equivalent (`--brave`
is interactive-only).

1. Instrumentation: server call log first (needs a plugin rebuild; the observer restarts MPS
   itself with `mps_control.sh restart`) or transcript-only?
2. Worker models (run list_worker_models.py; default: the orchestrator model from that list).
2a. Effort level for `<model>` (one per selected model; options from `effortLevels`; recommended:
   the level the user-level settings give that model today, if any). Every run of that model gets
   `EFFORT=<level>`.
3. Permission mode for workers (default: `bypassPermissions` on the developer's machine).
4. Scope of the first pass before gate 1 (default: S1 + S3 on the selected models).
Gate 1 (after the pilot): matrix size. Gate 2 (after the report): which remedies; A/B yes/no.

## Procedure (tick as you go; details in the references)

1. **Preflight** — MPS running with the MCP server enabled. The port belongs to the IDE selector
   (64343 on the 261 from-sources MPS, 64344 on 262); `mps_control.sh` and `run_worker.sh` detect
   it from the live launcher themselves (`mps_control.sh url --json` shows what they find). **Do
   not export `MPS_MCP_URL` during preflight**: MPS may not be up yet, and an exported unconfirmed
   value pins the whole round to the wrong port. Set it by hand only to override detection (e.g.
   several MPS processes), or after `mps_control.sh wait` has reported `confirmed: true`. SMOKE
   targets the harness project, never a developer checkout. Check the toolchain:
   `claude --version` (≥ 2.1; must accept `--output-format stream-json --strict-mcp-config`)
   when the detected harness is Claude, or `junie --version` when it is Junie,
   `python3 -c 'import sys; assert sys.version_info >= (3, 9)'`, `jq --version`.
   Preflight is self-healing, and every step of it is yours:
   - MPS not running → `mps_control.sh start` (or, with no capture on file, the IDEA `MPS` run
     configuration), then `mps_control.sh wait`.
   - Welcome screen → **synthesize** a *harness project* and open it:
     `python3 $STUDY/scripts/new_study_project.py --dir ~/MPSProjects/mcp-study/proj/harness`
     writes the three descriptor files (`migration.xml` derived from this MPS, so no Migration
     Assistant), then open it via CLI (`mps-project-management`). Announce the path first. Do not
     retry MCP until that open has landed — Welcome-screen calls are rejected before dispatch.
   - `mps_mcp_list_open_projects(projectPath=<harness>)` must then list it. A synthesized project is
     empty by construction (`mps_mcp_get_project_structure` returns no modules); nothing in the study
     depends on a hand-maintained one, though an existing empty project may be substituted if
     synthesis fails. **Every other project must be closed first** — including the developer's own
     checkout, the common case when MPS was started from the IDEA run configuration. Disjoint
     module names are not enough: with `confirmOpenNewProject2 = -1` (the default) the second open
     raises the modal New Window / This Window prompt and blocks the round (lesson 30). Announce
     the path you close; it is an observer action, not one to ask for.
   - Is the call log on? Ask the process, not the log:
   `ps -ww -p $(pgrep -f '[j]etbrains\.mps\.Launcher') -o args= | tr ' ' '\n' | grep calllog`
   must print the option, and the file must grow after a tool call. If it is off and gate question
   1 said call-log, step 2 turns it on; if gate 1 said transcript-only,
   expect 0-line `*-server.jsonl` slices and skip the call-log checks below.
   Record the tool inventory: `MPS_MCP_URL=$($STUDY/scripts/mps_control.sh url) python3
   $STUDY/scripts/tools_inventory.py --out $RUNS/inventory.json` (it does not detect the port itself).
   Run the harness's own unit tests once (`cd $STUDY/scripts && python3 -m unittest discover -s
   tests -p 'test_*.py'`) — a broken script is cheaper to find here than in the evidence.
   Then run the contamination guard yourself: `python3 $STUDY/scripts/check_user_agents.py`
   (exit 0 clean, 3 contaminated). It rejects MPS-related Markdown definitions below
   `~/.claude/agents` / `~/.junie/agents` — a filename matching `*mps*` or a body containing
   `mps_mcp`, both case-insensitively — **and** any `mps-*` folder in `~/.claude/skills` or
   `~/.junie/skills`, which would shadow the per-project catalog and silently replace the thing
   being measured (lessons 26, 32). `run_worker.sh` runs the same guard before any run side effect.
   The guard never modifies anything: move an offending user skill out of the skills directory
   for the round and restore it at wrap-up. Built-in `Explore` and `Task`
   agents are outside this pin and remain enabled.
2. **Instrument** — the plugin logs one JSON line per dispatched call when MPS runs with
   `-Dmps.mcp.calllog=<file>` (`McpCallLogListener`, off by default). Turn it on without a human and
   without touching a tracked file: `mps_control.sh capture` **while MPS is still alive**, then
   `mps_control.sh calllog $RUNS/server-calllog.jsonl` (writes the option into the capture), then
   `shutdown` (the close of the last project carries the exit), `start <project>`, `wait`, and one
   SMOKE run as the readiness gate. `capture` only preserves VM options the live process already
   carries, which is why `calllog` exists — adding the option to the `MPS` run configuration works
   too but is study-only and must be reverted at wrap-up (lesson 13), so prefer the capture route.
   Confirm the relaunched process actually carries it (`ps -ww -p <pid> -o args= | tr ' ' '\n' |
   grep calllog`) and that the file grows.
3. **Template** — the empty fixture is **synthesized, not snapshotted**
   (`scripts/new_study_project.py`): three descriptor files, no doc surface possible, and a
   `migration.xml` derived from the MPS that will open it. Module-bearing fixtures (`statechart`,
   `recipes*`) are still tarballs, snapshotted **without any agent doc surface**: exclude
   `.git`, `workspace.xml`, and also `.agents/`, `.claude/`, `AGENTS.md`, `CLAUDE.md`. A tarball is
   a point-in-time copy, so a catalog inside it is what every later round measures no matter how
   far the bundled skills have moved (lesson 20). Instead, `run_worker.sh` installs the **live**
   catalog into each run's project right before launching the worker — `scripts/install_skills.py`
   purges every `mps-*` folder plus both guides and calls `mps_mcp_initialize_project_for_agents`,
   then records `skillsSha256` in the meta. Verify every tarball:
   `tar -tzf <f>.tar.gz | grep -E '(^|/)(\.claude|\.agents|AGENTS\.md|CLAUDE\.md)'` must be empty
   (a synthesized project has nothing to verify).
   Do NOT put `.mcp.json` in the template; `run_worker.sh` generates the worker's MCP config per
   run into `$RUNS/<id>-mcp/` from the detected URL and passes it with `--strict-mcp-config`.
4. **Smoke** — `SMOKE` is a harness check, not a scenario: a read-only prompt that lists open
   projects and stops, so it runs against the harness project itself (no template copy, no
   evaluation, `pass` stays empty). It is also the **readiness gate after every MPS start or
   restart** — a live process is not readiness. If the harness project is not open, announce its
   path, synthesize it if needed and open it via CLI; do not close it afterwards unless the next
   run needs a different project.
   `RUNS=$RUNS EFFORT=<level for $MODEL> MAX_TURNS=6 PROJECT_SYNTHESIZED=1 $STUDY/scripts/run_worker.sh SMOKE $MODEL <n>
   <harness-project>` — bump `<n>` on every re-run (the harness refuses an existing run id);
   drop `PROJECT_SYNTHESIZED=1` if the harness project was not synthesized. The transcript
   must contain `tool_use`, `tool_result`, per-message `usage`; exactly one MCP server; and, when
   the call log is on, a `SMOKE-…-server.jsonl` slice of ≥ 1 line.
5. **Scenarios** — `study/scenarios/S1..S10/{worker_prompt.md,done_criteria.md}`. Which cells a
   changed skill actually forces is `study/scenarios.md` (brief list, then the directory). Look
   that up before picking the matrix; the gate-4 default (S1 + S3) is a pilot default, not that
   lookup. Add a scenario for whatever skill/tool the directory does not cover. **S10 (project
   lifecycle) runs last in a round**: it is the only
   scenario whose worker closes and opens projects, and a mistake in it can leave a modal dialog
   that blocks every later `mps_mcp_*` call. Prompts are developer-voice, fixed names, explicit "done", NO reporting
   requirements. Fixtures: `empty-project` (synthesized per run, not a tarball),
   `statechart` (Projectxx5), `recipes` (a passing S1) — regenerated per
   `study/fixtures/README.md`, not stored in git.
6. **Runs** — ONE scratch project open at a time (see lessons: shared module repository leaks across
   projects; S10 honours this by being sequential — it closes one project before opening the next).
   Restart MPS before a cell whose fixture language an earlier cell already loaded in the current
   process — any two of S3/S5/S6/S7/S9 on `recipes*`, S2/S8 on `statechart` (`shutdown` → `start`
   harness → `wait` → SMOKE; launch with `ISOLATION=per-shared-fixture-restart`, recorded per run
   beside `mpsPid`; lessons 40, 42). A read-only cell (S9) may precede a language-changing one in the
   same process; synthesized cells (S1, S10) need no restart.
   Per run: **synthesize** the empty project or copy the fixture tarball (`PROJECT_SYNTHESIZED=1`
   when synthesized) → announce the scratch path (and any path you will close first)
   → close a previous scratch with `mps_mcp_close_project` if one is still open → open the new copy
   via CLI (`mps-project-management`) → confirm with `list_open_projects` → launch
   detached (`run_worker.sh` first rejects MPS-related user agents, then installs the live skills;
   either guard failure aborts with exit 3) →
   poll the PID in bounded loops → evaluate with an Opus subagent using the `done_criteria.md`
   (read-only `mps_mcp_*`, always with `projectPath`) → record pass/evidence in `<id>.meta.json`
   → announce the path and close with `mps_mcp_close_project` (`force=false`; on `MODAL_BLOCKED`
   ask the user only to dismiss the dialog). Sequential, never two workers against one MPS. Pass the
   model's gate-2a level as `EFFORT` on every launch. Check that every meta's `effort` is that
   level and that every meta's
   `skillsSha256` and `guidesSha256` are each the same value before comparing runs; a differing one
   means the catalog or the installed `AGENTS.md` / `CLAUDE.md` moved mid-round. Open/close details: `references/harness.md`.
7. **Analyse** — `python3 $STUDY/scripts/analyze_runs.py $RUNS [--out DIR]` (default `$RUNS/analysis`)
   → `metrics.csv`, `tools.json`, `chains.json`, `errors.json`, `hotspots.md`; `pass` is filled from
   each run's meta after evaluation. Then
   `python3 $STUDY/scripts/families.py <baseline runs> <previous round> $RUNS > $RUNS/analysis/families.tsv`
   (per-run family counters, `references/analysis.md`). Filter chains containing `mps_mcp`, group into families, have an
   Opus reviewer inspect 3 instances per family with `study/scripts/show_steps.py` and assign
   determinism {1.0, 0.5, 0}. Rank by avoidable turns (fixed context ≈ 150 K cache-read tokens per
   turn dominates) as well as by the study formula.
8. **Classify** each hotspot D → S → P-off → P-on → T (first fit). Write `HOTSPOT_REPORT.md`:
   baseline table, ranked hotspots with `run:step` evidence, hypotheses, defects, remedies with
   owner/contract/saving/risk. Keep a separate `docs-defects.md` from day one.
9. **Treat** — parallel Opus implementers with DISJOINT file sets (skill docs / skill scripts +
   packaging + drift test / server batch). Never two agents in the same Kotlin toolset; the observer
   registers new tests in `McpToolsIntegrationTestSuite` and runs the suite between batches. A
   server-side change needs the plugin rebuilt and MPS restarted: `mps_control.sh restart` + SMOKE,
   no human step.
10. **A/B** (optional) — restart MPS onto the treated plugin (`mps_control.sh restart`, then
    SMOKE), then the same runs against the treated tools; success = ≥ 30 % fewer tool calls and
    ≥ 25 % fewer context tokens on treated scenarios, no drop in pass rate; delete remedies that do
    not pay.
11. **Wrap up** — fold conclusions into the study doc; revert the VM option; delete fixture tarballs
    (keep the SMOKE scenario — step 4 needs it); announce and close any remaining scratch with
    `mps_mcp_close_project`, **closing the last one with `shutdownWithLastProject=true`** if MPS
    should go down — the shutdown rides on that last close, because a Welcome-screen MPS cannot be
    stopped over MCP; delete any `<proj>-target` directory an S10 run left; clean
    `~/MPSProjects/mcp-study/`, `~/.claude.json` project entries, and
    `~/.claude/projects/-…-mcp-study-proj-*/` memory dirs; Junie workers may leave
    `~/.junie/sessions` — do not auto-delete them; keep the call-log listener. The capture
    file lives in `$TMPDIR`, outside that cleanup, so a later round can still relaunch.

## References

- `references/harness.md` — run_worker.sh, analyze_runs.py, show_steps.py, tools_inventory.py,
  new_study_project.py, mps_control.sh usage; clean-environment rule; the observer's project and
  MPS lifecycle protocols (create / open / close / shutdown + relaunch); per-run procedure card.
- `references/scenarios.md` — how to run and add a scenario, fixtures, orchestrator project swap,
  done-criteria style. Which scenario covers which skill: `plugins/mcp-tools/study/scenarios.md`.
- `references/analysis.md` — metrics, chain scoring, rubric, report template, thresholds.
- `references/lessons.md` — what went wrong the first time and the rule that came out of it.
