Agent skill

Sandbox Code Execution Fallback

by HKUDS in HKUDS/OpenSpace

Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

MITAuto-check passedAgent Workflows

Install Sandbox Code Execution Fallback

skills CLI
$ npx skills add HKUDS/OpenSpace --skill code-exec-fallback -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace code-exec-fallback --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/code-exec-fallback .claude/skills/code-exec-fallback && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-exec-fallback
GitHub stars
7.7k
Token cost
~588 tokens
SKILL.md length
200 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

  • Works in 3 steps: Write the Python Script → Execute via Shell → Handle Output
  • The code sandbox tool fails twice in a row on the same Python task
  • SKILL.md covers When to Use, The Pattern, Step-by-Step Instructions and Example, plus 2 more sections
  • Calls python3

What it does

The pattern applies after the execute_code_sandbox tool has failed repeatedly, typically twice or more, because of environment limits, timeouts, missing dependencies or sandbox restrictions. Instead of passing code to the sandbox, the agent saves the Python script with write_file, runs it with run_shell using python3, captures stdout and stderr, and optionally deletes the temporary file afterward.

A few tips keep the shell run reliable: put every import in the script because the shell environment may differ from the sandbox, use absolute paths or confirm the working directory, add error handling, raise the timeout from its 30 second default for long jobs, and print JSON when results must be parsed. It helps most when the sandbox lacks packages, restricts file input and output, times out on work a shell can finish, or when a task needs external commands or a multi-file layout.

When your agent uses it

  • The code sandbox tool fails twice in a row on the same Python task
  • A script needs packages the sandbox does not have installed
  • File reads or writes are blocked inside the sandbox
  • A long-running script times out in the sandbox but should work in a shell

Example prompts

  • “The sandbox keeps timing out on this CSV analysis, so save it as analyze.py and run it from the shell.”
  • “Retry the failed data-cleaning code by writing it to process_data.py and running python3.”
  • “The sandbox blocks file writes, so run this report script through the shell and print the result as JSON.”

Requirements

  • An agent with write_file and run_shell tools
  • Python 3 in the shell environment

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Write the Python Script
  2. Execute via Shell
  3. Handle Output

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sandbox Code Execution Fallback loads about 588 tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 200 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~588

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 200 words, ~588 tokens.

Download SKILL.mdSave it as .claude/skills/code-exec-fallback/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
code-exec-fallback
description
Fallback pattern for executing Python code when execute_code_sandbox fails

Code Execution Fallback

When to Use

Use this pattern when execute_code_sandbox fails repeatedly (typically 2+ attempts) due to environment limitations, timeouts, dependency issues, or sandbox restrictions.

The Pattern

Instead of executing code directly in the sandbox, write the Python script to a file and execute it via shell:

  1. Write the script using write_file
  2. Execute via shell using run_shell with python3 script.py
  3. Clean up (optional) remove the temporary file

Step-by-Step Instructions

Step 1: Write the Python Script
Use write_file to save your Python code:
- Path: Choose a descriptive name (e.g., "process_data.py", "analyze.py")
- Content: Your complete Python script with all imports and logic
Step 2: Execute via Shell
Use run_shell to execute:
- Command: "python3 <script_name>.py"
- Timeout: Set appropriately for your task (default 30s, increase if needed)
Step 3: Handle Output
- Capture stdout/stderr from run_shell
- Parse results as needed
- Optionally delete the script file after execution

Example

python
# Instead of this (which may fail):
execute_code_sandbox(code="import pandas as pd; df = pd.read_csv('data.csv')...")

# Do this:
write_file(path="analyze.py", content="""
import pandas as pd
import json

df = pd.read_csv('data.csv')
result = df.groupby('category').sum()
print(json.dumps(result.to_dict()))
""")

run_shell(command="python3 analyze.py", timeout=60)

Tips for Success

  1. Include all imports in the script file - the shell environment may differ from the sandbox
  2. Use absolute paths or ensure working directory is correct
  3. Add error handling to your script for better debugging
  4. Increase timeout for long-running operations (default is 30s)
  5. Print structured output (JSON) if you need to parse results
  6. Clean up temporary files after successful execution to avoid clutter

When This Helps

  • Sandbox has missing dependencies
  • Code execution times out in sandbox but would work in shell
  • File I/O operations are restricted in sandbox
  • Need to run external commands or system utilities
  • Complex multi-file projects that need proper file structure

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/code-exec-fallback of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

Sandbox Code Execution Fallback next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sandbox Code Execution Fallback compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sandbox Code Execution Fallback this skillHKUDS/OpenSpace7.7k—~588Automated safety check: PassMIT
AWS Lambda Durable Functionsawslabs/agent-plugins915—~2.3kAutomated safety check: PassApache-2.0
Fake Model Provider Faultsdifferent-ai/openwork24k—~642Automated safety check: PassCustom licence
Cross-Language Coding Standardszereight/gitlab-mcp2k1 repos~1.4kAutomated safety check: PassMIT
Python Devdatabricks-solutions/ai-dev-kit1.9k—~1.6kAutomated safety check: PassCustom licence
Python Script Runnercongchuanling-dot/Cohort199—~577Automated safety check: NotesMIT

Similar skills

  • AWS Lambda Durable Functions

    awslabs/agent-plugins

    Official

    Build resilient, long-running, multi-step applications with AWS Lambda durable functions with automatic state persistence, retry logic, and orchestration for long-running executions.

    915 GitHub stars~2.3k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Fake Model Provider Faults

    different-ai/openwork

    Makes the desktop app's model provider fail on demand, with refused connections, resets, stalls and HTTP 4xx and 5xx errors, so error and retry states can be reproduced.

    24k GitHub stars~642 tokensUpdated today
    Testing & QAAuto-check passed
  • Shared reference for naming, function size, complexity and error handling rules that reviewer agents apply across TypeScript, Python, Go, Rust, Java, C# and Swift.

    2k GitHub starsUsed in 1 repo~1.4k tokens
    DevelopmentAuto-check passed
  • Python Dev

    databricks-solutions/ai-dev-kit

    Python development guidance with code quality standards, error handling, testing practices, and environment management.

    1.9k GitHub stars~1.6k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Python Script Runner

    congchuanling-dot/Cohort

    Runs a Python script with checks first: confirms the interpreter version, compiles it for syntax errors, verifies imports, then explains any traceback in plain language.

    199 GitHub stars~577 tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • A skill your agent uses when writing shell scripts, Python automation, or any unattended batch job.

    3.5k GitHub stars~225 tokensUpdated 4 mo ago
    DevelopmentAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.7k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.7k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.7k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Generate documents with writefile when retrieval tools fail, with explicit guardrails against task context drift

    7.7k GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Categories

Questions about Sandbox Code Execution Fallback

What does Sandbox Code Execution Fallback do?

Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead. The pattern applies after the execute_code_sandbox tool has failed repeatedly, typically twice or more, because of environment limits, timeouts, missing dependencies or sandbox restrictions. Instead of passing code to the sandbox, the agent saves the Python script with write_file, runs it with run_shell using python3, captures stdout and stderr, and optionally deletes the temporary file afterward.

When should I use Sandbox Code Execution Fallback?

Sandbox Code Execution Fallback fits situations like: the code sandbox tool fails twice in a row on the same Python task; A script needs packages the sandbox does not have installed; file reads or writes are blocked inside the sandbox; A long-running script times out in the sandbox but should work in a shell.

How do I install Sandbox Code Execution Fallback in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill code-exec-fallback -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/code-exec-fallback in HKUDS/OpenSpace) into .claude/skills/code-exec-fallback in your project. Claude Code loads it when a task matches its description.

How do I install Sandbox Code Execution Fallback in Codex?

Run `npx skills add HKUDS/OpenSpace --skill code-exec-fallback -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/code-exec-fallback in HKUDS/OpenSpace) into .agents/skills/code-exec-fallback in your project. Codex loads it when a task matches its description.

Can I use Sandbox Code Execution Fallback in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill code-exec-fallback -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-exec-fallback, .gemini/skills/code-exec-fallback, .github/skills/code-exec-fallback and .opencode/skills/code-exec-fallback in your project.

What does Sandbox Code Execution Fallback need to run?

Going by SKILL.md and its folder, Sandbox Code Execution Fallback needs the command-line tools its instructions call (python3). Our summary lists: An agent with write_file and run_shell tools; Python 3 in the shell environment.

Does Sandbox Code Execution Fallback access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sandbox Code Execution Fallback safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sandbox Code Execution Fallback use?

Sandbox Code Execution Fallback is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sandbox Code Execution Fallback use?

About 588 tokens (SKILL.md is roughly 2.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sandbox Code Execution Fallback?

Skills that share tags, products or a category with Sandbox Code Execution Fallback: AWS Lambda Durable Functions (awslabs/agent-plugins, 915 stars), Fake Model Provider Faults (different-ai/openwork, 24k stars), Cross-Language Coding Standards (zereight/gitlab-mcp, 2k stars) and Python Dev (databricks-solutions/ai-dev-kit, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sandbox Code Execution Fallback?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,749 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.