Agent skill

Run Deep Swe

by sickn33 in sickn33/agentic-awesome-skills

Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.

MITAuto-check passedAI & LLM Engineering

Install Run Deep Swe

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill run-deep-swe -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills run-deep-swe --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/run-deep-swe .claude/skills/run-deep-swe && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-deep-swe
GitHub stars
47k
Used in
1 other repo
Token cost
~1.2k tokens
SKILL.md length
390 words
Files
1
Skills in repo
1,493
Repo updated
First seen
Licence
MIT

At a glance

Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.

  • Tasks that involve Model routing and gateways
  • SKILL.md covers When to Use, Prerequisites — state-check…, Setup and OpenRouter wiring (the part…, plus 6 more sections
  • Calls docker, git and uv; reaches github.com; needs OPENROUTER_API_KEY
  • Tasks that involve Agent evaluation and testing

What it does

Run Deep Swe is an agent skill from sickn33/agentic-awesome-skills. Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model routing and gateways and Agent evaluation and testing. It works with OpenRouter. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Model routing and gateways
  • Tasks that involve Agent evaluation and testing

Example prompts

  • “/run-deep-swe”

Requirements

  • Docker
  • A credential in OPENROUTER_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 680176d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • git
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Deep Swe loads about 1.2k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 390 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 680176d, republished under its MIT licence (© sickn33). 390 words, ~1,191 tokens.

Download SKILL.mdSave it as .claude/skills/run-deep-swe/SKILL.md (or your agent's skills folder).
name
run-deep-swe
description
Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.
category
agent-evaluation
risk
critical
source
community
source_repo
davidondrej/skills
source_type
community
date_added
2026-07-07
author
davidondrej
tags
benchmark, deepswe, openrouter, evaluation
tools
claude, codex
license
MIT

Run DeepSWE via OpenRouter

When to Use

  • Use when the user wants to benchmark a model on DeepSWE or mini-swe-agent tasks.
  • Use when you need a reproducible coding-agent evaluation plan and output artifacts.

DeepSWE (deepswe.datacurve.ai) is a 113-task Harbor-compatible coding-agent benchmark. It runs via Pier (Harbor fork) driving mini-swe-agent (model-agnostic). Any model reachable through OpenRouter can be scored.

Prerequisites — state-check first

bash
which uv git docker || echo "MISSING: install uv, git, docker"
docker info >/dev/null 2>&1 || echo "MISSING: Docker daemon not running (Pier's default sandbox)"
echo "OPENROUTER_API_KEY set? ${OPENROUTER_API_KEY:+YES}"

Docker must be running — Pier sandboxes each task in Docker by default (--env modal for cloud instead).

OPENROUTER_API_KEY must already be present in the environment. If it is unset, ask the user to configure their preferred secret-management path; do not read shell startup files, print secrets, or invent a key.

Setup

bash
git clone https://github.com/datacurve-ai/deep-swe && cd deep-swe
uv tool install datacurve-pier            # PyPI (preferred)
# or: uv tool install git+https://github.com/datacurve-ai/pier
# pier bundles mini-swe-agent as the --agent driver

Run all pier commands from inside deep-swe/, using relative -p tasks/....

OpenRouter wiring (the part the docs don't spell out)

mini-swe-agent has a native OpenRouter model class. Both routes below use OPENROUTER_API_KEY and the OpenRouter slug (vendor/model, e.g. minimax/minimax-m3):

Route A — native OpenRouter class (preferred, hits openrouter.ai/api/v1 directly):

bash
pier run -p deep-swe/tasks --agent mini-swe-agent \
  --model minimax/minimax-m3 --model-class openrouter

Route B — LiteLLM provider prefix (fallback; same key):

bash
pier run -p deep-swe/tasks --agent mini-swe-agent \
  --model openrouter/minimax/minimax-m3

Notes:

  • Slug = the exact OpenRouter slug. Verify it at openrouter.ai/models before running.
  • Free/zero-cost models: OpenRouter cost tracking can error. Set export MSWEA_COST_TRACKING=ignore_errors.
  • Flag spelling can vary by version — confirm with pier run --help and mini --help.

Smoke test FIRST (1 task — do this before any full run)

Always validate end-to-end wiring on a single task before spending tokens on the corpus:

bash
pier run -p deep-swe/tasks/<task-id> --agent mini-swe-agent \
  --model minimax/minimax-m3 --model-class openrouter
# list available task ids:
ls deep-swe/tasks

Pass criteria: run completes, model returns actions (not auth/format errors), a score/trajectory is emitted. If it 401s → key wrong. If "provider not provided"/"model not mapped" → fix slug or switch route.

Show full SKILL.md (131 more words)Show less

Subset run (deterministic sample)

bash
pier run -p deep-swe/tasks --agent mini-swe-agent \
  --model minimax/minimax-m3 --model-class openrouter \
  --n-tasks 10 --sample-seed 0

Full 113-task corpus (costs tokens + time — confirm with user first)

bash
pier run -p deep-swe/tasks --agent mini-swe-agent \
  --model minimax/minimax-m3 --model-class openrouter
# add `--env modal` to run in parallel Modal sandboxes (needs Modal configured)

Output & leaderboard

  • Trials land in jobs/<run>/<trial_id>/. Inspect with pier view jobs/<run>, pier analyze jobs/<run>, or pier critique run jobs/<run>.
  • Report: the exact command used, pass/fail, score, and any blockers.
  • Submit results for the official leaderboard to: <email-address>

Failure modes

SymptomCauseFix
HTTP 401bad/missing keyre-export OPENROUTER_API_KEY
"LLM Provider NOT provided"missing slug prefixuse Route B openrouter/... or Route A with --model-class openrouter
"model isn't mapped"/cost errorunknown cost for modelexport MSWEA_COST_TRACKING=ignore_errors
unknown flagversion driftcheck pier run --help

Limitations

  • Adapted from davidondrej/skills; verify local paths, tools, credentials, and agent features before acting.
  • For commands, remote access, scheduling, browser automation, or file-changing workflows, get explicit user approval and confirm the target environment first.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/run-deep-swe of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit 680176d

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Run Deep Swe next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Deep Swe compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Deep Swe this skillsickn33/agentic-awesome-skills47k1 repos~1.2kAutomated safety check: PassMIT
Openrouter Context Optimizationjeremylongshore/tons-of-skills-marketplace2.8k—~2.4kAutomated safety check: PassMIT
Openrouter Trending ModelsMadAppGang/claude-code285—~3.6kAutomated safety check: PassMIT
QuorumDetrol/quorum-cli119—~807Automated safety check: NotesCustom licence
OpenCode Agent Provider for NanoClawnanocoai/nanoclaw31k—~5kAutomated safety check: NotesMIT
Darwinian EvolverLuciole-Studio/Misaka-Agent1582 repos~2.1kAutomated safety check: WarnMIT

Similar skills

  • Openrouter Context Optimization

    jeremylongshore/tons-of-skills-marketplace

    Optimize context window usage for OpenRouter models to reduce cost and improve quality.

    2.8k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Openrouter Trending Models

    MadAppGang/claude-code

    Fetch trending programming models from OpenRouter rankings. An agent skill from MadAppGang/claude-code.

    285 GitHub stars~3.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Quorum

    Detrol/quorum-cli

    Run a structured debate between agent CLIs (claude, codex, agy, grok) and the user's configured API or local models (OpenAI, Anthropic, Google, xAI, OpenRouter, Ollama and more) through the Quorum…

    119 GitHub stars~807 tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check: notes
  • Installs OpenCode as an optional NanoClaw agent runtime, reaching OpenRouter, OpenAI, Google, DeepSeek and others through OpenCode's own configuration.

    31k GitHub stars~5k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check: notes
  • Darwinian Evolver

    Luciole-Studio/Misaka-Agent

    Evolve prompts/regex/SQL/code with Imbue's evolution loop. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 2 repos~2.1k tokens
    AI & LLM EngineeringAuto-check: warnings
  • Openrouter Load Balancing

    jeremylongshore/tons-of-skills-marketplace

    Distribute OpenRouter requests across multiple keys and models for high throughput.

    2.8k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,493 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Questions about Run Deep Swe

What does Run Deep Swe do?

Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent. Run Deep Swe is an agent skill from sickn33/agentic-awesome-skills. Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.

When should I use Run Deep Swe?

Run Deep Swe fits situations like: tasks that involve Model routing and gateways; tasks that involve Agent evaluation and testing.

How do I install Run Deep Swe in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill run-deep-swe -a claude-code`. Or copy the skill folder (skills/run-deep-swe in sickn33/agentic-awesome-skills) into .claude/skills/run-deep-swe in your project. Claude Code loads it when a task matches its description.

How do I install Run Deep Swe in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill run-deep-swe -a codex`. Or copy the skill folder (skills/run-deep-swe in sickn33/agentic-awesome-skills) into .agents/skills/run-deep-swe in your project. Codex loads it when a task matches its description.

Can I use Run Deep Swe in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill run-deep-swe -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-deep-swe, .gemini/skills/run-deep-swe, .github/skills/run-deep-swe and .opencode/skills/run-deep-swe in your project.

What does Run Deep Swe need to run?

Going by SKILL.md and its folder, Run Deep Swe needs the command-line tools its instructions call (docker, git and uv) and credentials named OPENROUTER_API_KEY. Our summary lists: Docker; A credential in OPENROUTER_API_KEY.

Does Run Deep Swe access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Run Deep Swe safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Deep Swe use?

Run Deep Swe is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run Deep Swe use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Run Deep Swe?

Skills that share tags, products or a category with Run Deep Swe: Openrouter Context Optimization (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Openrouter Trending Models (MadAppGang/claude-code, 285 stars), Quorum (Detrol/quorum-cli, 119 stars) and OpenCode Agent Provider for NanoClaw (nanocoai/nanoclaw, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Deep Swe?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,379 GitHub stars. The repository holds 1,493 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.