Agent skill

Dask Parallel Computing

by jaechang-hits in jaechang-hits/SciAgent-Skills

Parallel/distributed computing for larger-than-RAM data. An agent skill from jaechang-hits/SciAgent-Skills.

BSD-3-ClauseAuto-check passedData & Analytics

Install Dask Parallel Computing

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill dask-parallel-computing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills dask-parallel-computing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scientific-computing/dask-parallel-computing .claude/skills/dask-parallel-computing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dask-parallel-computing
GitHub stars
370
Token cost
~4.1k tokens
SKILL.md length
917 words
Files
3 (incl. references)
Skills in repo
165
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Parallel/distributed computing for larger-than-RAM data. An agent skill from jaechang-hits/SciAgent-Skills.

  • Works in 5 steps: DataFrames — Parallel Pandas → Arrays — Parallel NumPy → Bags — Unstructured Data Processing → …
  • Tasks that involve DataFrames
  • SKILL.md covers Overview, When to Use, Prerequisites and Core API, plus 9 more sections
  • Calls pip

What it does

Dask Parallel Computing is an agent skill from jaechang-hits/SciAgent-Skills. Parallel/distributed computing for larger-than-RAM data. Components: DataFrames (parallel pandas), Arrays (parallel NumPy), Bags, Futures, Schedulers. Scales laptop to HPC cluster. For single-machine speed use polars; for out-of-core without cluster use vaex.

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/collections_guide.md` and `references/distributed_computing.md`).

It sits in Data & Analytics, covering DataFrames. It works with Dask, NumPy, pandas and Polars. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.

When your agent uses it

  • Tasks that involve DataFrames

Example prompts

  • “/dask-parallel-computing”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. DataFrames — Parallel Pandas
  2. Arrays — Parallel NumPy
  3. Bags — Unstructured Data Processing
  4. Futures — Task-Based Parallelism
  5. Schedulers & Configuration

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.dask.org
    • ml.dask.org
    • jobqueue.dask.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dask Parallel Computing loads about 4.1k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 917 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 917 words, ~4,064 tokens.

Download SKILL.mdSave it as .claude/skills/dask-parallel-computing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
dask-parallel-computing
description
Parallel/distributed computing for larger-than-RAM data. Components: DataFrames (parallel pandas), Arrays (parallel NumPy), Bags, Futures, Schedulers. Scales laptop to HPC cluster. For single-machine speed use polars; for out-of-core without cluster use vaex.
license
BSD-3-Clause

Dask — Parallel & Distributed Computing

Overview

Dask is a Python library for parallel and distributed computing that scales familiar pandas/NumPy APIs to larger-than-memory datasets. It provides five main components (DataFrames, Arrays, Bags, Futures, Schedulers) and scales from single-machine multi-core to multi-node HPC clusters.

When to Use

  • Processing datasets that exceed available RAM (10 GB–100 TB)
  • Parallelizing pandas or NumPy operations across multiple cores
  • Processing multiple files efficiently (CSV, Parquet, JSON, HDF5, Zarr)
  • Building custom parallel workflows with task dependencies
  • Distributing workloads across HPC clusters (SLURM, Kubernetes)
  • Streaming/ETL pipelines for unstructured data (logs, JSON records)
  • For in-memory single-machine speed: use polars instead
  • For out-of-core single-machine analytics: use vaex instead

Prerequisites

bash
pip install dask[complete]        # All components
pip install dask[dataframe]       # DataFrames only
pip install dask[distributed]     # Distributed scheduler + dashboard
pip install dask-jobqueue          # HPC cluster integration (SLURM, PBS)

Core API

1. DataFrames — Parallel Pandas
python
import dask.dataframe as dd

# Read multiple files as a single DataFrame
ddf = dd.read_csv('data/2024-*.csv')
ddf = dd.read_parquet('data/', columns=['id', 'value', 'category'])

# Operations are lazy until .compute()
filtered = ddf[ddf['value'] > 100]
result = filtered.groupby('category').agg({'value': ['mean', 'sum']}).compute()
print(result.shape)  # (n_categories, 2)

# Custom operations via map_partitions (preferred over apply)
def normalize_partition(df):
    df['norm_value'] = (df['value'] - df['value'].mean()) / df['value'].std()
    return df

ddf = ddf.map_partitions(normalize_partition)

# Joins
ddf_merged = ddf.merge(lookup_ddf, on='category', how='left')

# Write results
ddf.to_parquet('output/', engine='pyarrow')
python
# Repartitioning for optimal chunk sizes
ddf = ddf.repartition(npartitions=20)         # By count
ddf = ddf.repartition(partition_size='100MB')  # By size

# Index management for sorted operations
ddf = ddf.set_index('timestamp', sorted=True)

# Debugging
print(f"Partitions: {ddf.npartitions}")
print(f"Dtypes: {ddf.dtypes}")
sample = ddf.get_partition(0).compute()  # Inspect first partition
2. Arrays — Parallel NumPy
python
import dask.array as da
import numpy as np

# Create from various sources
x = da.random.random((100000, 1000), chunks=(10000, 1000))
x = da.from_array(np_array, chunks=(10000, 1000))
x = da.from_zarr('large_dataset.zarr')

# Standard operations (lazy)
y = (x - x.mean(axis=0)) / x.std(axis=0)  # Normalize
z = da.dot(x.T, x)                          # Matrix multiply
u, s, v = da.linalg.svd(x)                  # SVD

# Compute and persist
result = y.mean(axis=0).compute()
print(result.shape)  # (1000,)
python
# Custom operations with map_blocks
def custom_filter(block):
    from scipy.ndimage import gaussian_filter
    return gaussian_filter(block, sigma=2)

filtered = da.map_blocks(custom_filter, x, dtype=x.dtype)

# Rechunking for different access patterns
x_rechunked = x.rechunk({0: 5000, 1: 500})

# Save to disk
da.to_zarr(y, 'normalized.zarr')
3. Bags — Unstructured Data Processing
python
import dask.bag as db
import json

# Read unstructured data
bag = db.read_text('logs/*.json').map(json.loads)

# Functional operations
valid = bag.filter(lambda x: x['status'] == 'success')
ids = valid.pluck('user_id')
flat = bag.map(lambda x: x['tags']).flatten()

# Aggregation — use foldby instead of groupby (much faster)
counts = bag.foldby(
    key='category',
    binop=lambda total, x: total + x['amount'],
    initial=0,
    combine=lambda a, b: a + b,
    combine_initial=0
).compute()

# Convert to DataFrame for structured analysis
ddf = valid.to_dataframe(meta={'user_id': 'str', 'amount': 'float64', 'category': 'str'})
4. Futures — Task-Based Parallelism
python
from dask.distributed import Client

client = Client()  # Local cluster with all cores
print(client.dashboard_link)  # http://localhost:8787

# Submit individual tasks (executes immediately, not lazy)
def process(x, param):
    return x ** param

future = client.submit(process, 42, param=2)
print(future.result())  # 1764

# Map over many inputs
futures = client.map(process, range(100), param=2)
results = client.gather(futures)
print(len(results))  # 100
python
# Scatter large data to workers (avoids repeated transfers)
import numpy as np
big_data = np.random.random((10000, 1000))
data_future = client.scatter(big_data, broadcast=True)

# Submit tasks using scattered data
futures = [client.submit(process_chunk, data_future, i) for i in range(10)]
results = client.gather(futures)

# Progressive result processing
from dask.distributed import as_completed
for future in as_completed(futures):
    result = future.result()
    print(f"Completed: {result}")

# Coordination primitives
from dask.distributed import Lock, Queue, Event
lock = Lock('resource-lock')
with lock:
    # Thread-safe operation across workers
    pass

client.close()
5. Schedulers & Configuration
python
import dask

# Global scheduler setting
dask.config.set(scheduler='threads')       # Default: GIL-releasing numeric work
dask.config.set(scheduler='processes')     # Pure Python, GIL-bound work
dask.config.set(scheduler='synchronous')   # Debugging with pdb

# Context manager for temporary change
with dask.config.set(scheduler='synchronous'):
    result = computation.compute()  # Can use pdb here

# Per-compute override
result = ddf.mean().compute(scheduler='processes')

# Distributed scheduler with resource control
from dask.distributed import Client
client = Client(n_workers=4, threads_per_worker=2, memory_limit='4GB')
print(client.dashboard_link)
python
# HPC cluster integration
from dask_jobqueue import SLURMCluster
from dask.distributed import Client

cluster = SLURMCluster(
    cores=24, memory='100GB',
    walltime='02:00:00', queue='regular'
)
cluster.scale(jobs=10)  # Request 10 SLURM jobs
client = Client(cluster)

# Adaptive scaling
cluster.adapt(minimum=2, maximum=20)

result = computation.compute()
client.close()

Key Concepts

Component Selection Guide
Data TypeComponentWhen to Use
Tabular (CSV, Parquet)DataFramesStandard pandas-like operations at scale
Numeric arrays (HDF5, Zarr)ArraysNumPy operations, linear algebra, image processing
Text, JSON, logsBagsETL/cleaning → convert to DataFrame for analysis
Custom parallel tasksFuturesDynamic workflows, parameter sweeps, task dependencies
Any of aboveSchedulersControl execution backend (threads/processes/distributed)

Control level: DataFrames/Arrays/Bags = high-level lazy API. Futures = low-level immediate execution.

Lazy Evaluation Model

All DataFrames, Arrays, and Bags build a task graph — nothing executes until .compute() or .persist().

  • .compute() — execute and return result to local memory
  • .persist() — execute and keep result on workers (for reuse across multiple computations)
  • dask.compute(a, b, c) — compute multiple results in a single pass (shares intermediates)
Chunk Size Strategy

Target: ~100 MB per chunk (or 10 chunks per core in worker memory).

Chunk SizeEffect
Too large (>1 GB)Memory overflow, poor parallelization
Optimal (~100 MB)Good parallelism, manageable memory
Too small (<1 MB)Excessive scheduling overhead

Example: 8 cores, 32 GB RAM → target ~400 MB per chunk (32 GB / 8 cores / 10).

Scheduler Selection Guide
SchedulerOverheadBest ForGIL
threads (default)~10 µs/taskNumPy, pandas, scikit-learnAffected
processes~10 ms/taskPure Python, text processingNot affected
synchronous~1 µs/taskDebugging with pdbN/A
distributed~1 ms/taskDashboard, clusters, advanced featuresConfigurable

Common Workflows

Workflow 1: Multi-File ETL Pipeline
python
import dask.dataframe as dd
import dask

# Extract: Read all CSV files
ddf = dd.read_csv('raw_data/*.csv', dtype={'amount': 'float64'})

# Transform: Clean and process
ddf = ddf[ddf['status'] == 'valid']
ddf['amount'] = ddf['amount'].fillna(0)
ddf = ddf.dropna(subset=['category'])

# Aggregate
summary = ddf.groupby('category').agg({'amount': ['sum', 'mean', 'count']})

# Load: Save as Parquet (columnar, compressed)
summary.to_parquet('output/summary.parquet')
print(f"Processed {len(ddf)} rows across {ddf.npartitions} partitions")
Workflow 2: Large-Scale Array Processing
python
import dask.array as da

# Load large scientific dataset
x = da.from_zarr('experiment_data.zarr')  # e.g., (50000, 50000) float64
print(f"Shape: {x.shape}, Chunks: {x.chunks}")

# Normalize per-column
x_norm = (x - x.mean(axis=0)) / x.std(axis=0)

# Compute covariance matrix
cov = da.dot(x_norm.T, x_norm) / (x_norm.shape[0] - 1)

# SVD for dimensionality reduction (top-k)
u, s, v = da.linalg.svd_compressed(x_norm, k=50)

# Save results
da.to_zarr(u, 'pca_components.zarr')
print(f"Explained variance (top 5): {(s[:5]**2 / (s**2).sum()).compute()}")
Workflow 3: Unstructured Data to Analysis

This workflow is a simple combination of Bags (Section 3) → DataFrames (Section 1): read JSON logs with Bags, filter/transform, convert to DataFrame for groupby analysis. Each step maps directly to Core API examples above.

Key Parameters

ParameterModuleDefaultDescription
npartitionsDataFrameautoNumber of partitions (controls parallelism)
partition_sizeDataFrame—Target size per partition (e.g., '100MB')
chunksArrayrequiredChunk dimensions (e.g., (10000, 1000))
blocksizeBag'128 MiB'File read block size
schedulerAll'threads'Execution backend ('threads', 'processes', 'synchronous')
n_workersDistributedautoNumber of worker processes
threads_per_workerDistributedautoThreads per worker
memory_limitDistributedautoPer-worker memory limit (e.g., '4GB')
sortedset_indexFalseWhether data is pre-sorted (enables optimizations)
metamap_partitions—Output DataFrame/Series structure template

Best Practices

  1. Let Dask handle data loading — Never load data into pandas/numpy first then convert. Use dd.read_csv() / da.from_zarr() directly.

  2. Batch compute calls — Use dask.compute(a, b, c) instead of calling .compute() in loops. Allows sharing intermediates.

  3. Use map_partitions over apply — ddf.apply(func, axis=1) creates one task per row. ddf.map_partitions(func) creates one task per partition.

  4. Persist reused intermediates — Call .persist() on data accessed multiple times, then del when done.

  5. Use the dashboard — client.dashboard_link shows task progress, memory usage, worker states. Essential for diagnosing performance issues.

  6. Anti-pattern — Excessively large task graphs: If len(ddf.__dask_graph__()) returns millions, increase chunk sizes or use map_partitions/map_blocks to fuse operations.

  7. Anti-pattern — Wrong scheduler for workload: Using threads for pure Python text processing (GIL-bound) or processes for NumPy operations (unnecessary serialization overhead).

Show full SKILL.md (309 more words)Show less

Common Recipes

Recipe: Dask-ML Integration
python
from dask_ml.preprocessing import StandardScaler
from dask_ml.model_selection import train_test_split
import dask.array as da

X = da.random.random((100000, 50), chunks=(10000, 50))
y = da.random.randint(0, 2, size=100000, chunks=10000)

# Preprocessing
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# Train/test split
X_train, X_test, y_train, y_test = train_test_split(X_scaled, y, test_size=0.2)
print(f"Train: {X_train.shape}, Test: {X_test.shape}")
Recipe: Actors for Stateful Computation
python
from dask.distributed import Client

client = Client()

class RunningStats:
    def __init__(self):
        self.count = 0
        self.total = 0.0

    def add(self, value):
        self.count += 1
        self.total += value
        return self.total / self.count

# Create actor on a worker
stats = client.submit(RunningStats, actor=True).result()

# Call methods (~1ms roundtrip)
for v in [10, 20, 30]:
    mean = stats.add(v).result()
    print(f"Running mean: {mean}")

client.close()
Recipe: Kubernetes Cluster
python
from dask_kubernetes import KubeCluster
from dask.distributed import Client

cluster = KubeCluster()
cluster.adapt(minimum=2, maximum=50)
client = Client(cluster)

# Run computation on auto-scaling cluster
result = large_computation.compute()
client.close()

Troubleshooting

ProblemCauseSolution
MemoryError during .compute()Result too large for local memoryUse .to_parquet() or .persist() instead of .compute()
Slow computation startTask graph has millions of tasksIncrease chunk sizes; use map_partitions/map_blocks
Poor parallelizationGIL contention with threads schedulerSwitch to scheduler='processes' for Python-heavy code
TypeError in map_partitionsMissing or wrong meta parameterProvide meta=pd.DataFrame({'col': pd.Series(dtype='float64')})
Workers killed (OOM)Chunks exceed worker memoryDecrease chunk size; increase memory_limit
KilledWorker exceptionWorker process crashedCheck worker logs; reduce memory per task; increase memory_limit
Slow joins/mergesData not pre-sorted on join keyCall ddf.set_index('key', sorted=True) before join
NotImplementedErrorOperation not supported by DaskUse map_partitions with pandas equivalent
Dashboard not accessibleDistributed client not startedUse client = Client() to enable distributed scheduler
Data type mismatch across partitionsInconsistent CSV filesSpecify dtype explicitly in dd.read_csv()

Bundled Resources

  • references/collections_guide.md — Detailed DataFrames, Arrays, and Bags guide with comprehensive code examples for reading, transforming, aggregating, and writing data. Covers map_partitions patterns, meta parameter, chunking strategies, map_blocks, foldby, and collection conversion. Consolidated from original dataframes.md, arrays.md, and bags.md. Original best-practices.md content relocated to Best Practices section and Key Concepts inline.
  • references/distributed_computing.md — Futures API, distributed coordination primitives (Locks, Queues, Events, Variables), Actors, scheduler configuration, HPC cluster setup (SLURM, Kubernetes), adaptive scaling, dashboard monitoring, and performance profiling. Consolidated from original futures.md and schedulers.md.

Not migrated: Original had 6 reference files. best-practices.md content consolidated into Best Practices section and Key Concepts (chunk strategy, scheduler selection). Remaining content organized into 2 reference files covering the 5 main components.

  • polars-dataframes — In-memory single-machine DataFrame library; faster than Dask for data that fits in RAM
  • zarr-python — Chunked array storage format; primary Dask Array persistence backend
  • scikit-learn-machine-learning — ML library; Dask-ML provides distributed wrappers

References

© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/scientific-computing/dask-parallel-computing of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/collections_guide.md
  • references/distributed_computing.md

Open the folder on GitHubat commit 82c862c

Compare with similar skills

Dask Parallel Computing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dask Parallel Computing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dask Parallel Computing this skilljaechang-hits/SciAgent-Skills370—~4.1kAutomated safety check: PassBSD-3-Clause
Daskdavila7/claude-code-templates32k11 repos~3.5kAutomated safety check: PassMIT
DaskK-Dense-AI/scientific-agent-skills48k1 repos~4.4kAutomated safety check: NotesBSD-3-Clause
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Vaex Out-of-Core DataFramesdavila7/claude-code-templates32k12 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Dask

    davila7/claude-code-templates

    Parallel/distributed computing. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.5k tokens
    Data & AnalyticsAuto-check passed
  • Dask

    K-Dense-AI/scientific-agent-skills

    Scales pandas, NumPy, and custom Python research workflows beyond memory or across clusters with Dask.

    48k GitHub starsUsed in 1 repo~4.4k tokens
    Data & AnalyticsAuto-check: notes
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Vaex Out-of-Core DataFrames

    davila7/claude-code-templates

    Processes tabular datasets too large for RAM with Vaex: lazy DataFrames, fast aggregations, big-data plots and ML pipelines over CSV, HDF5, Arrow and Parquet.

    32k GitHub starsUsed in 12 repos~1.6k tokens
    Data & AnalyticsAuto-check passed
  • Hybrid-Engine Data Analysis

    code-yeongyu/oh-my-openagent

    Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.

    70k GitHub stars~1.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 165 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    370 GitHub stars~4k tokensUpdated 9 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    370 GitHub stars~3.2k tokensUpdated 9 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    370 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    370 GitHub stars~6.9k tokensUpdated 9 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub stars~2.3k tokensUpdated 9 days ago
    Auto-check passed

Questions about Dask Parallel Computing

What does Dask Parallel Computing do?

Parallel/distributed computing for larger-than-RAM data. An agent skill from jaechang-hits/SciAgent-Skills. Dask Parallel Computing is an agent skill from jaechang-hits/SciAgent-Skills. Parallel/distributed computing for larger-than-RAM data.

When should I use Dask Parallel Computing?

Dask Parallel Computing fits situations like: tasks that involve DataFrames.

How do I install Dask Parallel Computing in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill dask-parallel-computing -a claude-code`. Or copy the skill folder (skills/scientific-computing/dask-parallel-computing in jaechang-hits/SciAgent-Skills) into .claude/skills/dask-parallel-computing in your project. Claude Code loads it when a task matches its description.

How do I install Dask Parallel Computing in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill dask-parallel-computing -a codex`. Or copy the skill folder (skills/scientific-computing/dask-parallel-computing in jaechang-hits/SciAgent-Skills) into .agents/skills/dask-parallel-computing in your project. Codex loads it when a task matches its description.

Can I use Dask Parallel Computing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill dask-parallel-computing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dask-parallel-computing, .gemini/skills/dask-parallel-computing, .github/skills/dask-parallel-computing and .opencode/skills/dask-parallel-computing in your project.

What does Dask Parallel Computing need to run?

Going by SKILL.md and its folder, Dask Parallel Computing needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Dask Parallel Computing access the network?

SKILL.md names 3 domains. As links in the text: docs.dask.org, ml.dask.org and jobqueue.dask.org. This is read from the text; nothing was executed.

Is Dask Parallel Computing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dask Parallel Computing use?

Dask Parallel Computing is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dask Parallel Computing use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.

What are the alternatives to Dask Parallel Computing?

Skills that share tags, products or a category with Dask Parallel Computing: Dask (davila7/claude-code-templates, 32k stars), Dask (K-Dense-AI/scientific-agent-skills, 48k stars), Python Executor (cortega26/chile-hub, 113 stars) and Verified Data Analysis with pandas (pipeshub-ai/pipeshub-ai, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dask Parallel Computing?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 370 GitHub stars. The repository holds 165 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.