Agent skill

Code Engineer

by openJiuwen-ai in openJiuwen-ai/sciencediscovery

A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…

Apache-2.0Auto-check passedData & Analytics

Install Code Engineer

skills CLI
$ npx skills add openJiuwen-ai/sciencediscovery --skill code-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openJiuwen-ai/sciencediscovery code-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/code-engineer .claude/skills/code-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-engineer
GitHub stars
151
Token cost
~2.8k tokens
SKILL.md length
1,188 words
Files
2 (incl. scripts)
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…

  • Works in 8 steps: Understand Requirements → Use the Bundled Executor → Inspect Available Data → …
  • You need to write and execute Python/R code to process
  • SKILL.md covers Overview, Core Capabilities, When to Use This Skill and Python Package Installation, plus 8 more sections
  • Runs Python scripts from its folder; calls python and pip; reaches pypi.tuna.tsinghua.edu.cn

What it does

Code Engineer is an agent skill from openJiuwen-ai/sciencediscovery. Use this skill when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology documentation. Supports statistical analysis, data transformation, visualization, method justification, and structured result output.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/execute.py`).

It sits in Data & Analytics, covering Statistics, Data analysis and Data cleaning. It works with Python. The repository describes itself as: ScienceDiscovery is an all‑in‑one agentic workbench built specifically for scientific research. The licence is Apache-2.0.

When your agent uses it

  • You need to write and execute Python/R code to process
  • Delivering reproducible computational results with complete code-level methodology documentation

Example prompts

  • “/code-engineer”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Understand Requirements
  2. Use the Bundled Executor
  3. Inspect Available Data
  4. Write and Execute Analysis Code
  5. Document and Return Results
  6. Inspect the data file
  7. Write analysis code and execute
  8. Document results per Output Schema

What it can do on your machine

Read from SKILL.md and the folder at commit 7a03242. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • pypi.tuna.tsinghua.edu.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Engineer loads about 2.8k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 1,188 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from openJiuwen-ai/sciencediscovery at commit 7a03242, republished under its Apache-2.0 licence (© openJiuwen-ai). 1,188 words, ~2,805 tokens.

Download SKILL.mdSave it as .claude/skills/code-engineer/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
code-engineer
description
Use this skill when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology documentation. Supports statistical analysis, data transformation, visualization, method justification, and structured result output.

Code Engineer Skill

Overview

This skill writes and executes Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology documentation. It focuses on code-level fidelity, computational result completeness, and methodology transparency.

Core Capabilities

  • Write and execute Python/R code for data processing, transformation, and statistical analysis
  • Inspect data file structure (sheets, columns, types, row counts) before analysis
  • Generate reproducible computational results with method justification
  • Document code-level methodology (libraries, data transformations, key function calls)
  • Adjust analysis per feedback and document changes
  • Specify data traceability (file name, sheet name, field/column name, row count)
  • Support both Python and R execution environments
  • Export execution output to CSV, JSON, or Markdown (supports both raw text and structured DataFrame export)

When to Use This Skill

Always load this skill when:

  • User asks for data processing, transformation, or statistical analysis that must be executed as Python or R code
  • User wants reproducible computational results with documented libraries, methods, and assumptions
  • User provides a data file (Excel/CSV/Parquet/etc.) and asks to inspect its schema, run analyses on it, or export structured results
  • User asks for code-level methodology documentation alongside results (method justification, data traceability, key function calls, limitations)
  • User wants a multi-step analysis pipeline with re-runnable code (Python or R scripts) rather than ad-hoc one-off answers

Python Package Installation

If you need to install new Python packages, install them through the Tsinghua PyPI mirror for reliability:

bash
pip install [python package] -i https://pypi.tuna.tsinghua.edu.cn/simple

Workflow

Step 1: Understand Requirements

Identify what the analysis task requires:

  • Analysis objectives: What computational results are expected
  • Available data: File paths, data descriptions, format details
Step 2: Use the Bundled Executor

The complete frozen package is already available read-only at $SCIENCEDISCOVERY_SKILLS_DIR/code-engineer when the execution sandbox starts. Invoke scripts/execute.py directly from that fixed package path and pass requested data paths and options as arguments. Do not load this large script into context, copy or rewrite it, search the filesystem for another copy, or execute it until the workflow requires an explicit inspect or run action.

Step 3: Inspect Available Data

Before writing analysis code, inspect the data to understand its schema and characteristics:

bash
python "$SCIENCEDISCOVERY_SKILLS_DIR/code-engineer/scripts/execute.py" \
  --action inspect \
  --files /path/to/data.xlsx

This returns:

  • Sheet names (for Excel) or filename (for CSV)
  • Column names, data types, and non-null counts
  • Row count per sheet/file
  • Sample data (first 5 rows)
Step 4: Write and Execute Analysis Code

Based on the analysis objectives and data schema, write Python/R code to perform the analysis.

Execute Python Code
bash
python "$SCIENCEDISCOVERY_SKILLS_DIR/code-engineer/scripts/execute.py" \
  --action run \
  --language python \
  --code-file /path/to/workspace/analysis_step1.py \
  --files /path/to/data.xlsx \
  --output-file /path/to/outputs/analysis_results.json
Execute R Code
bash
python "$SCIENCEDISCOVERY_SKILLS_DIR/code-engineer/scripts/execute.py" \
  --action run \
  --language r \
  --code-file /path/to/workspace/analysis_step1.R \
  --files /path/to/data.xlsx \
  --output-file /path/to/outputs/analysis_results.json
Run Inline Code Snippet
bash
python "$SCIENCEDISCOVERY_SKILLS_DIR/code-engineer/scripts/execute.py" \
  --action run \
  --language python \
  --code "import pandas as pd; df = pd.read_excel('/path/to/data.xlsx'); print(df.describe())" \
  --output-file /path/to/outputs/summary_stats.csv
Step 5: Document and Return Results

Structure your output per the Output Schema below. Ensure every result includes method justification, data traceability, assumptions, and code-level documentation.

When results may be evaluated downstream (e.g., by the result-evaluator skill), present a Result Package that includes both structured data AND methodology documentation. The --output-file exports only tabular data; the agent must also provide the following in conversation:

  • Methodology: libraries used, statistical methods, key function calls, and justification for method choices
  • Data traceability: source file names, sheet/column names, row counts, and any filtering or transformation applied
  • Assumptions & limitations: distributional assumptions, sample size considerations, known data quality issues
  • Analysis code: the complete code that produced the results (for reproducibility verification)
Parameters
ParameterRequiredDescription
--actionYesOne of: inspect, run
--languageFor runpython or r
--codeFor runInline code string to execute
--code-fileFor runPath to a Python/R script file to execute
--filesNoSpace-separated paths to data files (loaded into execution context)
--output-fileNoPath to export results (CSV/JSON/MD). If the code assigns a DataFrame to result, it is exported as structured tabular data; otherwise raw stdout/stderr is exported

[!NOTE] Do NOT read or copy the Python file. Call its fixed read-only package path with the parameters.

Variable Naming Rules

When using --files, each data file is automatically loaded into the execution context as a variable:

  • Excel files: The variable name is derived from the filename without extension (e.g., sales_2024.xlsx → sales_2024), loaded via pd.read_excel() in Python or read_excel() in R
  • CSV files: The variable name is derived from the filename without extension (e.g., data.csv → data), loaded via pd.read_csv() in Python or read.csv() in R
  • Special characters: Filenames with spaces or special characters are auto-sanitized (spaces → underscores). Names starting with digits are prefixed with t_ (e.g., 2024_data.csv → t_2024_data)
  • Multiple files: Each file creates a separate variable, enabling cross-file analysis
Show full SKILL.md (478 more words)Show less

Complete Example

Task: "Analyze the correlation between variable X and Y in dataset.csv, and test whether the correlation is statistically significant."

Step 1: Inspect the data file
bash
python "$SCIENCEDISCOVERY_SKILLS_DIR/code-engineer/scripts/execute.py" \
  --action inspect \
  --files /path/to/dataset.csv
Step 2: Write analysis code and execute
bash
python "$SCIENCEDISCOVERY_SKILLS_DIR/code-engineer/scripts/execute.py" \
  --action run \
  --language python \
  --code-file /path/to/workspace/correlation_analysis.py \
  --files /path/to/dataset.csv \
  --output-file /path/to/outputs/correlation_results.json

Where correlation_analysis.py contains:

python
import pandas as pd
from scipy import stats

data = pd.read_csv('/path/to/dataset.csv')
corr, p_value = stats.pearsonr(data['X'], data['Y'])

print(f"Pearson correlation: r={corr:.4f}, p={p_value:.6f}")
print(f"Sample size: n={len(data)}")
print(f"X stats: mean={data['X'].mean():.2f}, std={data['X'].std():.2f}")
print(f"Y stats: mean={data['Y'].mean():.2f}, std={data['Y'].std():.2f}")
Step 3: Document results per Output Schema

Return structured results with method justification, assumptions, limitations, and code documentation.

Output Handling

After code execution:

  • Present key findings directly in conversation — highlight the most important results, not just raw output
  • For large or multi-step results, export to file and share via present_files tool
  • Always explain computational findings in plain language with actionable takeaways
  • When code produces tables or statistics, format them clearly for readability
  • Suggest follow-up analyses or refinements when patterns are interesting or inconclusive
  • Offer to export results if the user wants to keep them
  • If execution fails, explain the error context (e.g., missing library, data schema mismatch) and suggest a corrected approach

Structured Export

When using --output-file, the script attempts to detect structured results automatically:

  • If your code assigns a pandas DataFrame (or a list of dicts) to a variable named result, the script captures it as structured tabular data and exports columns + rows properly
  • For CSV export: proper CSV with headers and rows
  • For JSON export: array of records [{col: val, ...}]
  • For MD export: Markdown table with | formatting
  • If no result variable is found, the export falls back to raw stdout/stderr text

Tip: To get structured output, simply assign your final DataFrame to result:

python
result = df.groupby('category').agg({'amount': 'sum'}).reset_index()

Caching (inspect only)

Caching applies only to --action inspect. The script stores loaded DuckDB tables to avoid re-parsing files on every inspect call:

  • On first inspect, files are loaded into a persistent DuckDB database under <tempdir>/.code-engineer-cache/
  • The cache key is a SHA256 hash of all input file contents — if files change, a new cache is created
  • Subsequent inspect calls with the same files reuse the cached database
  • Cache is transparent — no extra parameters needed

Note: --action run does not use this cache. The Python (pandas) and R (readxl/read.csv) subprocesses re-read the data files on every invocation. If you want run-time caching for an analysis pipeline, cache results yourself and reuse them.

For analyses where result quality matters, use the result-evaluator skill to evaluate output reliability and methodological rigor. This is especially recommended when:

  • Results inform decisions or will be presented to stakeholders
  • Statistical analyses where methodology correctness is critical
  • Complex multi-step analyses where errors can compound

To evaluate: load /mnt/skills/custom/result-evaluator/SKILL.md and provide the full Result Package (structured data + methodology documentation + data traceability + analysis code) as evaluation input.

Notes

  • Python execution uses the system Python environment with auto-installation of missing packages
  • R execution requires R to be installed on the system (auto-detected via Rscript or R)
  • For large datasets, DuckDB handles them efficiently without loading everything into memory
  • Column names with spaces are accessible using double quotes in SQL: "Column Name"

© openJiuwen-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/code-engineer of openJiuwen-ai/sciencediscovery.

  • SKILL.md
  • scripts/execute.py

Open the folder on GitHubat commit 7a03242

Compare with similar skills

Code Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Engineer this skillopenJiuwen-ai/sciencediscovery151—~2.8kAutomated safety check: PassApache-2.0
MatlabzLanqing/codex-claude-academic-skills4.6k9 repos~2.3kAutomated safety check: NotesGPL-3.0
Meridian MMM Model Buildinggoogle/meridian1.6k—~2.5kAutomated safety check: PassApache-2.0
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Statistical Data Analysislingzhi227/agent-research-skills384—~886Automated safety check: PassNone
Q-EDA Exploratory AnalysisTyrealQ/q-skills108—~1.1kAutomated safety check: PassMIT

Similar skills

  • Matlab

    zLanqing/codex-claude-academic-skills

    MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing.

    4.6k GitHub starsUsed in 9 repos~2.3k tokens
    Data & AnalyticsAuto-check: notes
  • Official

    Takes a user through building a Meridian marketing mix model, from loading CSV data and mapping columns to running EDA, fitting and saving the model.

    1.6k GitHub stars~2.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    384 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Runs exploratory data analysis on tabular data after you confirm each column's measurement level, then writes CSV tables and a narrative summary.

    108 GitHub stars~1.1k tokensUpdated 14 days ago
    Data & AnalyticsAuto-check passed
  • Data Analysis

    xiaoyuge886/aigc

    Perform data analysis tasks including data cleaning, statistical analysis, visualization, and insight generation.

    198 GitHub starsUsed in 1 repo~794 tokens
    Data & AnalyticsAuto-check passed

More from openJiuwen-ai/sciencediscovery

All 22 skills in this repo
  • Gitcode

    openJiuwen-ai/sciencediscovery

    Operate GitCode issues, PRs, wikis, code/MR refs, and cached org templates.

    151 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Structure Pocket Inspection

    openJiuwen-ai/sciencediscovery

    Inspect a local PDB structure, summarize chains and residue composition, and identify protein atoms near a user-specified ligand or pocket center.

    151 GitHub stars~624 tokensUpdated yesterday
    Auto-check passed
  • Antibody Design

    openJiuwen-ai/sciencediscovery

    Prepare, launch, monitor, and summarize the real RFdiffusion to ProteinMPNN to Protenix antibody pipeline on a local or remote ScienceDiscovery Runner with sandboxed Ascend NPUs.

    151 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Science Research Team

    openJiuwen-ai/sciencediscovery

    A skill your agent uses to orchestrate a multi-domain research team for literature/evidence research and data analysis.

    151 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Literature Searcher

    openJiuwen-ai/sciencediscovery

    A skill your agent uses when a research workflow needs verified academic source retrieval through literature-search MCP interfaces available in the current session before evidence extraction.

    151 GitHub stars~5.8k tokensUpdated yesterday
    Auto-check passed
  • Create GitHub PR

    openJiuwen-ai/sciencediscovery

    Open or update a pull request on GitHub's openJiuwen-ai/sciencediscovery: run the UT/ST/E2E layers locally, push the branch to the operator's own GitHub fork, write a body that says what was…

    151 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Code Engineer

What does Code Engineer do?

A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…. Code Engineer is an agent skill from openJiuwen-ai/sciencediscovery. Use this skill when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology documentation.

When should I use Code Engineer?

Code Engineer fits situations like: you need to write and execute Python/R code to process; delivering reproducible computational results with complete code-level methodology documentation.

How do I install Code Engineer in Claude Code?

Run `npx skills add openJiuwen-ai/sciencediscovery --skill code-engineer -a claude-code`. Or copy the skill folder (skills/code-engineer in openJiuwen-ai/sciencediscovery) into .claude/skills/code-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Code Engineer in Codex?

Run `npx skills add openJiuwen-ai/sciencediscovery --skill code-engineer -a codex`. Or copy the skill folder (skills/code-engineer in openJiuwen-ai/sciencediscovery) into .agents/skills/code-engineer in your project. Codex loads it when a task matches its description.

Can I use Code Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openJiuwen-ai/sciencediscovery --skill code-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-engineer, .gemini/skills/code-engineer, .github/skills/code-engineer and .opencode/skills/code-engineer in your project.

What does Code Engineer need to run?

Going by SKILL.md and its folder, Code Engineer needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3.

Does Code Engineer access the network?

SKILL.md names 1 domain. In commands or code: pypi.tuna.tsinghua.edu.cn; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Code Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Code Engineer use?

Code Engineer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Engineer use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Code Engineer?

Skills that share tags, products or a category with Code Engineer: Matlab (zLanqing/codex-claude-academic-skills, 4.6k stars), Meridian MMM Model Building (google/meridian, 1.6k stars), Pandas Pro (Jeffallan/claude-skills, 12k stars) and Statistical Data Analysis (lingzhi227/agent-research-skills, 384 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Engineer?

openJiuwen-ai (a GitHub organization) maintains it in openJiuwen-ai/sciencediscovery, which has 151 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 7, 2026.

Source: openJiuwen-ai/sciencediscovery on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.