---
name: active-swe-eval
description: Prepare and orchestrate the complete local Docker evaluation for Active-SWE. Claude Code remains the fixed executor for Recorded, Potential, and Judge.
---

# Active-SWE Evaluation With Codex

Use this skill when Codex organizes Active-SWE on the host. Claude Code remains
the fixed executor inside every task container.

## Locate the project

First look for an existing project root containing both `pyproject.toml` and
the `active_swe/` package. Reuse it without pulling, resetting, or replacing
local files. If no checkout exists, clone into a new relative workspace and
enter it:

```bash
git clone --depth 1 https://github.com/XLearning-SCU/Active-SWE.git Active-SWE
cd Active-SWE
```

Do not clone over a non-empty path. If Git or repository access is unavailable,
stop and ask the user for an existing checkout. After locating or cloning the
project, open `.codex/skills/active-swe-eval/SKILL.md` from that checkout and
use the local copy as the authority for all remaining steps. Then read its
local `ENVIRONMENT.md` and `EVALUATION.md`; do not continue from an older remote
or cached copy of the skill.

Public sources are:

```text
Source code: https://github.com/XLearning-SCU/Active-SWE
Dataset:     https://huggingface.co/datasets/XLearning-SCU/Active-SWE
```

Check `config/evaluation.json`. If absent, copy `config/evaluation.example.json`
and `config/.env.example`, then ask for missing values without silently
overwriting files. One task JSON contains `evaluation_models`, exactly one
`judge_model`, plus `input`, `output`, and host-side `execution` settings; keep
multiple tasks as separate JSON files. A
credential may use `api_key_env` loaded from project-local `config/.env` or the
host environment; a permission-restricted `api_key_file` remains available.
If a referenced variable is missing, ask the user to populate `config/.env`
without sending the value in chat. Hidden input is only for a user launching
the controller directly in an interactive terminal. Models run sequentially;
per-model concurrency defaults to 4, maximum turns to 300, and timeout to 5,400
seconds per task. Never put an API-key value in chat, `models.json`, or a
command line.
When the user selects only part of the configured models, pass one
`--only-model ID` per selection; keep the full config intact.
When the user identifies an existing task JSON and dotenv file, reuse them with
`--config` and `--env-file`.

Use the public controller from the `active_swe` package:

```bash
python -m active_swe.run_evaluation \
  --config config/evaluation.json
```

Check standard local Docker and Claude Code, download the requested public
dataset configuration to
`data/Active-SWE.parquet`, convert it to `data/Active-SWE.jsonl`, copy the exact
run input below `<output.root>/<id>/<timestamp>/inputs/`, pull and validate
referenced images, review the preflight image report, then
run Recorded and Potential with each evaluation model, Judge with the same
selected fixed judge, then metrics. The task's `input.limit` controls the
default scope; pass `--limit N` only for an intentional one-run override. The
materialized input is also the exact source used to choose preflight images.
Project-local tools belong under `./.tools`; run artifacts belong under
`./runs`.

Recorded and Potential use Docker `--network none`; only their bundled
fixed-destination API tunnel may carry model traffic. Judge has a separate
network setting and retains host-network behavior. Report artifacts using
paths, stage outcome counts, and metrics relative to the run root, without
exposing credentials, endpoint details, or host paths. Treat controller `ok`
as pipeline completion, not proof that every sample produced all artifacts.
