---
name: active-swe-eval
description: Prepare and run the complete local Docker evaluation for Active-SWE. Claude Code is the fixed executor for Recorded, Potential, and Judge.
---

# Active-SWE Evaluation With Claude Code

Use this skill to prepare an Active-SWE project checkout and evaluate one or
more models. Claude Code is both the outer orchestrator and the fixed executor
inside every task container.

## Locate the project

First look for an existing project root containing both `pyproject.toml` and
the `active_swe/` package. Reuse it without pulling, resetting, or replacing
local files. If no checkout exists, clone into a new relative workspace and
enter it:

```bash
git clone --depth 1 https://github.com/XLearning-SCU/Active-SWE.git Active-SWE
cd Active-SWE
```

Do not clone over a non-empty path. If Git or repository access is unavailable,
stop and ask the user for an existing checkout. After locating or cloning the
project, open `.claude/skills/active-swe-eval/SKILL.md` from that checkout and
use the local copy as the authority for all remaining steps; do not continue
from an older remote or cached copy. Then read these local files in order:

1. `.claude/skills/active-swe-eval/ENVIRONMENT.md`: Docker, Claude Code,
   dataset, images, model setup, and network boundaries.
2. `.claude/skills/active-swe-eval/EVALUATION.md`: Recorded, Potential, Judge,
   and the six paper metrics.

## Public Sources

```text
Source code: https://github.com/XLearning-SCU/Active-SWE
Dataset:     https://huggingface.co/datasets/XLearning-SCU/Active-SWE
```

When this skill is installed separately from the repository, its optional
`bin/bootstrap_active_swe.py` performs the same checkout check and clone.
Runtime tools belong under `./.tools`, and run artifacts belong under `./runs`;
neither is part of a source package.

## Required User Configuration

Check `config/evaluation.json`. If absent, copy
`config/evaluation.example.json` plus `config/.env.example`, then request the
missing values. One task file contains a non-empty `evaluation_models` list,
exactly one `judge_model`, plus `input`, `output`, and host-side `execution`
settings. Keep multiple evaluation tasks as separate JSON files and select one
with `--config`. Never overwrite an existing file silently. Put only credential
variable names in task files. Check
`config/.env`; if a required variable is absent from both that file and the
host environment, ask the user to populate it without sending the value in
chat. Keep `config/.env` at permission `0600`.
If the user already identifies a task JSON and dotenv file, use them directly
with `--config` and `--env-file`; do not ask for the same settings again.

```text
id:           stable, unique run/output identifier
input:        project-relative data path and row limit
output:       project-relative output root
execution:    host-side image-pull concurrency
model:        Anthropic-compatible model name
base_url:     fixed model API destination
credential:   API-key environment variable or local key file; otherwise prompt
concurrency:  optional per-model task concurrency; default 4
```

Run evaluation models sequentially. Recorded and Potential use the current
evaluation model; every evaluation model uses the same selected Judge. Defaults
are per-model concurrency 4,
maximum 300 turns, and 5,400 seconds per task.
When the user requests a subset, keep the shared config and pass one
`--only-model ID` per selected model. Do not edit away unselected entries.
Keep API-key values in project-local `config/.env`, host environment variables,
user-owned local key files, or hidden terminal input. The controller loads
`config/.env` automatically, while existing host variables take precedence.
Hidden input is only for a user launching the controller directly in an
interactive terminal. Never put key values directly in chat, `models.json`, a
skill command line, dataset, output, log, or generated result file.

## Workflow

1. Locate or bootstrap the project checkout.
2. Check standard local Docker and Claude Code. Install project-local tools
   under `./.tools` when needed.
3. Copy or check the tracked task example, update `config/evaluation.json`, then
   verify that `config/.env` or the host environment supplies each referenced
   credential variable without reading it back to the conversation.
4. Download the requested public dataset configuration to
   `data/Active-SWE.parquet`, convert it to `data/Active-SWE.jsonl`, copy the
   exact run input below `<output.root>/<id>/<timestamp>/inputs/`, validate it,
   and pull exactly the referenced images. Review the image preflight report
   before starting stages. The task's `input.limit` controls the default scope;
   pass `--limit N` only for an intentional one-run override.
5. Run each evaluation model through Recorded and Potential, use the selected
   fixed judge for Judge, then compute that evaluation model's metrics.
6. Report artifact locations, stage outcome counts, and computed metrics
   relative to the run root, without exposing host paths, credentials, or
   endpoint details. Treat controller `ok` as pipeline completion, not proof
   that every sample generated all expected artifacts.

Use the controller unless the user explicitly requests stage-level debugging:

```bash
python -m active_swe.run_evaluation \
  --config config/evaluation.json
```

## Isolation

Recorded and Potential containers always use Docker `--network none`. Model
traffic reaches only the configured fixed destination through the bundled
byte-transparent API tunnel. WebSearch, WebFetch, general network clients, and
Git history or network operations fail closed. The controller manages the
tunnel and must not print or persist its destination or payloads.

Judge is outside the Recorded/Potential isolation boundary. It retains the existing
host-network behavior and evaluates the image worktree without the generation
command policy. Never use Judge networking as a fallback for Recorded or
Potential.
