Agent skill

Data Distributed Storage

by Kilo-Org in Kilo-Org/kilo-marketplace

A skill your agent uses when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads.

MITAuto-check passedDatabases

Install Data Distributed Storage

skills CLI
$ npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kilo-Org/kilo-marketplace data-distributed-storage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-distributed-storage .claude/skills/data-distributed-storage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-distributed-storage
GitHub stars
190
Token cost
~4.9k tokens
SKILL.md length
1,064 words
Files
9 (incl. references)
Skills in repo
85
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads.

  • Works in 12 steps: Backend Selection → Storage Tiering → Lifecycle Policies → …
  • Designing distributed storage for HDFS
  • SKILL.md covers Purpose, Agent Protocol, Workflow and Rules, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Distributed Storage is an agent skill from Kilo-Org/kilo-marketplace. Use this skill when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads. This skill enforces: storage backend selection, data durability and replication, file format selection, partitioning and compression, lifecycle policies, storage tiering, and cost optimization. Do NOT use for: database storage engines, local filesystem tuning, or content delivery networks.

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `references/distributed-storage-advanced.md`, `references/distributed-storage-fundamentals.md` and `references/hdfs-architecture.md`).

It sits in Databases, covering File uploads and storage and Database administration. The repository describes itself as: Kilo Marketplace - A curated collection of Skills, MCP Servers, and Modes for enhancing AI agent capabilities across the Kilo ecosystem—including Kilo Code (VS Code extension)… The licence is MIT.

When your agent uses it

  • Designing distributed storage for HDFS
  • Any distributed file system for data workloads
  • : database storage engines
  • Local filesystem tuning

Example prompts

  • “/data-distributed-storage”

Requirements

  • Docker

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Backend Selection
  2. Storage Tiering
  3. Lifecycle Policies
  4. Durability and Replication
  5. Compression and Encoding
  6. Data Layout
  7. HDFS Architecture
  8. Security
  9. Cost Optimization
  10. Data Transfer and Migration
  11. Compliance Features
  12. Cost Modeling

What it can do on your machine

Read from SKILL.md and the folder at commit ff51758. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Distributed Storage loads about 4.9k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 1,064 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Kilo-Org/kilo-marketplace at commit ff51758, republished under its MIT licence (© Kilo-Org). 1,064 words, ~4,869 tokens.

Download SKILL.mdSave it as .claude/skills/data-distributed-storage/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
data-distributed-storage
description
Use this skill when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads. This skill enforces: storage backend selection, data durability and replication, file format selection, partitioning and compression, lifecycle policies, storage tiering, and cost optimization. Do NOT use for: database storage engines, local filesystem tuning, or content delivery networks.
metadata.category
data

Distributed Storage

Purpose

Design distributed storage systems for data workloads. Select the right storage backend (HDFS, S3, ADLS, GCS, MinIO), configure durability and replication, optimize file formats and compression, and implement lifecycle policies.

Agent Protocol

Trigger

Exact user phrases: "HDFS", "S3", "ADLS", "GCS", "MinIO", "distributed storage", "object storage", "file system", "data lake storage", "storage tier", "lifecycle policy", "replication factor", "erasure coding", "storage cost".

Input Context
  • Storage backend preference (S3, ADLS, GCS, HDFS, MinIO)
  • Data volume (current TB, annual growth)
  • Access patterns (frequent, infrequent, archive)
  • Durability and availability requirements
  • Budget and cost constraints
  • Compliance requirements (data residency, encryption)
  • Network bandwidth and latency constraints
Output Artifact

Distributed storage architecture with backend selection, tiering strategy, lifecycle policy, and cost model.

Response Format
yaml
# Storage backend
# Tier configuration
# Lifecycle rules
# Cost projection
Completion Criteria
  • Storage backend selected with rationale
  • Storage tiers defined (hot/warm/cold/archive)
  • Lifecycle policies configured for data movement
  • Data durability and replication strategy documented
  • Encryption (at rest, in transit) configured
  • Cost projection for current and 3-year growth
  • Backup and disaster recovery plan defined

Workflow

Step 1: Backend Selection
Object Store Comparison
FeatureAWS S3ADLS Gen2GCSMinIO
ConsistencyRead-after-writeStrongStrongStrong
AuthIAM roles, bucket policiesRBAC, SAS tokensIAM, service accountsJWT, OIDC, LDAP
EncryptionSSE-S3/KMS/CSESSE-AES/KMS/CMKGoogle-managed/CMEK/CSEKMS, auto-encryption
LifecycleTransition, expiry, versioningTiering, soft-deleteNearline/Coldline/ArchiveBucket lifecycle
Max object size5 TB4.75 TB5 TB100 TB (configurable)
S3 CompatibleNativeYes (via gateway)Yes (XML API)Native
Cost (per TB/mo)~$23 (standard)~$20 (hot)~$20 (standard)Hardware + ops
Durability99.999999999% (11 9s)99.999999999% (11 9s)99.999999999% (11 9s)Configurable
Availability SLA99.99%99.9%99.95%Depends on setup
Backend Selection Decision Tree
Cloud provider?
├── AWS → S3 (best integration with AWS ecosystem)
├── Azure → ADLS Gen2 (best with Azure AD, Active Directory)
├── GCP → GCS (strong consistency, best for Google ecosystem)
├── Multi-cloud → MinIO (S3-compatible layer across clouds)
└── On-premise / air-gapped → MinIO or HDFS

Primary workload?
├── Data lake with compute engines → Object store (S3/ADLS/GCS)
├── Legacy Hadoop/Spark without cloud → HDFS
├── Private cloud, edge, or air-gapped → MinIO
└── High-performance computing → Parallel file system (Lustre, GPFS)
Step 2: Storage Tiering
Tier Characteristics
TierStorage ClassLatencyCost/TB/moUse Case
HotS3 Standard, ADLS HotMillisecond~$23Active data, daily pipelines, ML training
WarmS3 Infrequent AccessMillisecond~$12.50Monthly queries, staging data
ColdS3 Glacier InstantMillisecond (instant)~$4Quarterly access, compliance retention
ArchiveS3 Glacier Deep Archive12 hours~$1Annual access, legal holds
Tiering Strategy

Hot: last 30-90 days of data. Warm: 90 days to 1 year. Cold: 1-3 years. Archive: 3-7+ years (retention-based). Intelligent tiering (S3) auto-moves objects between hot and warm based on access patterns.

Step 3: Lifecycle Policies
json
{
  "Rules": [
    {
      "Id": "TierAndExpire",
      "Status": "Enabled",
      "Filter": { "Prefix": "logs/" },
      "Transitions": [
        { "Days": 30, "StorageClass": "STANDARD_IA" },
        { "Days": 90, "StorageClass": "GLACIER" },
        { "Days": 365, "StorageClass": "DEEP_ARCHIVE" }
      ],
      "Expiration": { "Days": 2555 }
    },
    {
      "Id": "ExpireTempFiles",
      "Status": "Enabled",
      "Filter": { "Prefix": "tmp/" },
      "Expiration": { "Days": 7 }
    }
  ]
}
Step 4: Durability and Replication
Object Store Durability

S3/ADLS/GCS provide 11 9s of durability via erasure coding across multiple devices/facilities. Replication is handled by the cloud provider. For additional protection: cross-region replication (CRR) for disaster recovery, same-region replication (SRR) for compliance.

HDFS Replication

Default replication factor: 3 (1 primary, 2 replicas on different racks). Storage overhead: 3x. Erasure coding reduces overhead: RS-6-3 (1.33x overhead), RS-10-4 (1.5x overhead). EC recommended for cold data on HDFS (> 100TB).

Backup and DR

Cross-region backup for critical data. Versioning enabled on all buckets. Replication Rules: replicate Tier 1 data to DR region synchronously or asynchronously. Test restore quarterly.

Step 5: Compression and Encoding
CodecSpeedCompression RatioUse Case
ZSTD (level 3)Very fastGood (2-4x)Default for Parquet/ORC
GZIP (level 6)ModerateBest (3-6x)Archive, cold data
SnappyFastestFair (1.5-3x)Streaming, hot data
LZ4FastestFair (1.5-2x)Logs, real-time
BZIP2SlowBest (4-8x)Archive only

Compress data at the file format level (Parquet/ORC) rather than at the transport level. Avoid double compression. Test: compress a representative dataset with each codec, measure speed and ratio.

Step 6: Data Layout
Partitioning by File Structure

Organize storage by: source/system, data domain, date partition, file. Example: s3://data-lake/raw/salesforce/orders/year=2026/month=05/day=01/load_id=abc123/orders_001.parquet. This layout enables: partition pruning, easy lifecycle management, clear ownership boundaries.

File Size Targets

Small files cause performance problems for query engines. Target file size: 128MB-1GB. Use compaction jobs to merge small files. Monitor average file size per dataset.

Step 7: HDFS Architecture
NameNode and DataNode

Active NameNode manages filesystem metadata (namespace, blocks, locations). Standby NameNode for HA via QJM (Quorum Journal Manager). DataNodes store block data and report to NameNode. Block size: 128MB or 256MB default. Rack awareness for replica placement.

Show full SKILL.md (419 more words)Show less
HDFS Key Config
yaml
# hdfs-site.xml
dfs.replication: 3
dfs.block.size: 268435456  # 256MB
dfs.namenode.handler.count: 100
dfs.datanode.handler.count: 50
dfs.namenode.gc.time.threshold: 60s
dfs.namenode.checkpoint.period: 3600  # 1 hour
dfs.permissions.enabled: true
Step 8: Security
Encryption at Rest

Server-side encryption: SSE-S3 (AES-256, S3-managed keys), SSE-KMS (AWS KMS-managed keys), SSE-C (customer-provided keys). Client-side encryption: encrypt before upload, decrypt after download. For compliance: SSE-KMS with audit logging.

Encryption in Transit

TLS 1.2+ for all S3/ADLS/GCS API calls. Enforce HTTPS-only bucket policies. HDFS: enable SSL for RPC and data transfer.

json
{
  "Statement": [
    {
      "Effect": "Deny",
      "Principal": "*",
      "Action": "s3:*",
      "Resource": "arn:aws:s3:::data-lake/*",
      "Condition": {
        "Bool": { "aws:SecureTransport": "false" }
      }
    }
  ]
}
Step 9: Cost Optimization
Strategies

Use intelligent tiering for automatic cost savings on variable-access data. Request (S3) Reduced Redundancy for non-critical data (lower durability = lower cost). Reserved capacity for predictable usage. Monitor and alert on cost anomalies.

Cost Monitoring

Track: storage cost per bucket/container, data transfer costs (egress), request costs (PUT/GET/LIST), lifecycle transition costs. Allocate costs to teams/domains via tag-based cost allocation.

Decision Trees
Storage Backend
Cloud adoption strategy?
├── Single cloud → Native object store (S3/ADLS/GCS)
├── Multi-cloud → MinIO (abstraction layer)
├── On-premise
│   ├── Hadoop ecosystem → HDFS
│   └── Kubernetes-native → MinIO (S3-compatible)
└── Edge / IoT → MinIO (lightweight, S3 API)
Lifecycle Policy
Data access pattern?
├── Accessed frequently → Hot tier (delete when no longer needed)
├── Accessed occasionally → Intelligent tiering or manual tiering
├── Compliance retention → Write once, tier to cold → delete after retention
└── Temporary data → Hot tier → delete in 7-30 days
MinIO Deployment Architecture
Multi-Node Setup

MinIO runs as a distributed system across multiple nodes. Minimum 4 nodes for erasure coding protection. Each node: 4+ drives, SSD/NVMe preferred for performance. Deployed via: Docker Compose (dev/tiny), Kubernetes Operator (production), bare metal (HPC).

yaml
# MinIO Kubernetes Operator
apiVersion: minio.min.io/v2
kind: Tenant
metadata:
  name: data-lake-storage
spec:
  image: ${MINIO_IMAGE_REF:?Set MINIO_IMAGE_REF to quay.io/minio/minio@sha256:<reviewed-digest>}
  credsSecret:
    name: minio-creds
  pools:
    - servers: 4
      volumesPerServer: 4
      volumeClaimTemplate:
        spec:
          storageClassName: ssd
          accessModes: [ReadWriteOnce]
          resources:
            requests:
              storage: 4Ti
  mountPath: /export
  requestAutoCert: true
  buckets:
    - name: raw
    - name: staging
    - name: analytics
Erasure Coding

MinIO uses Reed-Solomon erasure coding. Default parity: N/2 drives (max protection). Configurable: standard (EC:4) for 8-drive setup, reduced (EC:2) for performance. Read tolerance: lose up to N/2 drives. Write tolerance: lose up to N/2 drives. Storage overhead: 2x for N/2 parity, 1.33x for N/4 parity.

yaml
# Erasure coding parity by deployment size
# 4 nodes × 4 drives = 16 drives
#   EC:8 → tolerate 8 drive failures, 2x overhead
#   EC:4 → tolerate 4 drive failures, 1.33x overhead
#   EC:2 → tolerate 2 drive failures, 1.14x overhead

# Set via environment variable
MINIO_STORAGE_CLASS_STANDARD: "EC:4"
MINIO_STORAGE_CLASS_RRS: "EC:2"     # Reduced redundancy for temp data
Performance Tuning
yaml
# Drive-level
# Use XFS filesystem with noatime,nodiratime mount options
# Separate write-intensive (metadata, WAL) from read-intensive drives

# Network-level
# MinIO uses S3-compatible HTTP/REST
# Enable HTTP/2 for multiplexing
# Use load balancer (nginx, HAProxy) for multi-node distribution

# Kernel tuning
# net.core.rmem_max: 134217728    # 128MB receive buffer
# net.core.wmem_max: 134217728    # 128MB send buffer
# net.ipv4.tcp_congestion_control: bbr  # BBR for better throughput
# vm.dirty_ratio: 30              # Page cache dirty ratio
# vm.dirty_background_ratio: 10
Step 10: Data Transfer and Migration
Transfer Acceleration
yaml
# AWS S3 Transfer Acceleration
# Uses CloudFront edge locations, TCP optimization
# Enable: s3.put_bucket_accelerate_configuration
# Best for: cross-region uploads, large objects > 1GB
# Cost: premium per GB transferred
# Test: s3-accelerate-speedtest.s3-accelerate.amazonaws.com

# Azure Data Box / Import-Export
# Physical transfer: Data Box (80TB), Data Box Disk (8TB)
# For: > 10TB initial loads, slow networks, air-gapped
# Timeline: order → receive → load → return → ingest (2-4 weeks)

# GCS Storage Transfer Service
# Online: from S3, HTTP, or GS URLs
# Schedule: one-time or recurring
# Best for: ongoing sync from S3 to GCS
Migration Strategy
yaml
migration_phases:
  phase_1_assessment:
    - inventory current storage (total TB, object count, avg size)
    - identify access patterns (hot vs cold, frequency)
    - document retention and compliance requirements
    - estimate egress costs and transfer time

  phase_2_pilot:
    - migrate 1-2 TB of non-critical data first
    - validate performance, cost, and access patterns
    - document issues and adjust strategy

  phase_3_bulk:
    - parallel multi-threaded copy (aws s3 sync, rclone, distcp)
    - validate data integrity checksums post-transfer
    - monitor for failures (rate limits, timeouts)

  phase_4_cutover:
    - final sync of delta changes
    - update application configs to point to new storage
    - redirect DNS / endpoint to new location
    - decommission old storage after 30-day validation window
Step 11: Compliance Features
Object Lock / WORM
yaml
# S3 Object Lock (compliance mode)
# Prevents object deletion or overwrite for retention period
# Governance mode: users with special permissions can override
# Compliance mode: no one (including root) can override
# Legal hold: prevents deletion regardless of retention

object_lock_examples:
  - name: "Financial records (SOX)"
    mode: COMPLIANCE
    retention_days: 2555  # 7 years
    bucket: "data-lake/finance/"

  - name: "Logs (internal policy)"
    mode: GOVERNANCE
    retention_days: 365  # 1 year
    bucket: "data-lake/logs/"

# ADLS Gen2: immutability policy
# Set at container level
# Time-based retention or legal hold
Data Residency
yaml
# Data residency requirements
# GDPR: data must stay in EU region
# Financial services: data within country borders
# Healthcare: HIPAA requires US-only storage (for US patients)

# Implementation:
# - Per-region buckets with IAM restrictions
# - S3 bucket policies denying cross-region replication for sensitive data
# - MinIO multi-tenant deployment per region
# - Regular audit to verify no data leaves allowed regions
Step 12: Cost Modeling
Cost Components
yaml
cost_components:
  storage:
    - per-GB-month by tier (hot, warm, cold, archive)
    - minimum storage duration charges (Glazer: 90 days min)
    - early deletion fees (cold/archive tiers)

  operations:
    - PUT/COPY/POST/LIST requests (per 1000)
    - GET/SELECT requests (per 1000)
    - lifecycle transition requests (per 1000)

  data_transfer:
    - upload (usually free)
    - download / egress (per GB, tiered pricing)
    - cross-region replication (per GB)
    - transfer acceleration (premium per GB)

  additional:
    - encryption (KMS: per key, per API call)
    - monitoring (CloudWatch, metrics per custom metric)
    - backup / replication (CRR storage in destination region)
Estimation Tool
yaml
# Monthly storage cost estimate
cost_estimate:
  hot_data:
    volume: 50 TB
    unit_cost: $23/TB/mo
    subtotal: $1,150/mo

  warm_data:
    volume: 150 TB
    unit_cost: $12.50/TB/mo
    subtotal: $1,875/mo

  cold_data:
    volume: 300 TB
    unit_cost: $4/TB/mo
    subtotal: $1,200/mo

  archive_data:
    volume: 500 TB
    unit_cost: $1/TB/mo
    subtotal: $500/mo

  operations:
    requests: $200/mo  # Based on 10M PUT + 100M GET

  data_transfer:
    egress: $300/mo    # Based on 10TB egress

  monitoring_backup:
    cross_region_replication: $500/mo  # 50TB replicated
    monitoring: $100/mo

  total_monthly: $5,825/mo
  total_annual: $69,900/yr
Step 13: Monitoring and Observability
Storage Metrics
yaml
# Object store monitoring (CloudWatch, Azure Monitor, GCS Ops)
metrics:
  storage:
    - BucketSizeBytes (by storage tier)
    - NumberOfObjects
    - AverageObjectSize

  requests:
    - AllRequests (count)
    - GetRequests / PutRequests / ListRequests
    - 4xxErrors / 5xxErrors
    - FirstByteLatency / TotalRequestLatency

  throughput:
    - BytesDownloaded / BytesUploaded
    - GetBandwidth / PutBandwidth (MB/s)

  cost:
    - StorageCost (by bucket/tag)
    - RequestCost
    - DataTransferCost
    - TotalCost
Alerting Thresholds
yaml
alerts:
  - name: "Storage growth anomaly"
    metric: BucketSizeBytes
    threshold: "> 20% week-over-week increase"
    severity: warning

  - name: "Error rate spike"
    metric: 5xxErrors
    threshold: "> 1% of total requests"
    severity: critical

  - name: "Latency degradation"
    metric: FirstByteLatency
    threshold: "p99 > 500ms"
    severity: warning

  - name: "Budget threshold"
    metric: TotalCost
    threshold: "> 80% of monthly budget"
    severity: warning

  - name: "Lifecycle failure"
    metric: LifecycleTransitions
    threshold: "transitions failed > 10 per hour"
    severity: warning
Decision Trees (continued)
Compression Codec Selection
Workload type?
├── Hot data, frequent queries → ZSTD (level 3) — fast decompression
├── Cold data, archive → GZIP (level 6) — best compression ratio
├── Streaming, low-latency → Snappy or LZ4 — minimal CPU overhead
├── Columnar formats (Parquet/ORC)
│   ├── Default → ZSTD (best balance)
│   └── Legacy compatibility → Snappy
└── JSON/CSV files
    ├── Compressed → GZIP (most tools support it)
    └── Splittable → BZIP2 or LZ4
Replication Strategy
Compliance requirement?
├── Disaster recovery (region outage)
│   ├── RTO < 1 hour → Synchronous CRR
│   └── RTO < 4 hours → Asynchronous CRR
├── Data sovereignty (must stay in region)
│   └── Same-region replication (SRR) + backup to another AZ
├── GDPR right to erasure
│   ├── Replicate selectively (no unnecessary copies)
│   └── Document replication topology for audit
└── No compliance requirement
    └── No replication — rely on cloud provider durability

Rules

  • Separate storage from compute for elasticity
  • Use lifecycle policies to automatically tier data
  • Compress all data at rest — storage is not free
  • Encrypt data at rest and in transit
  • Monitor storage growth and set budget alerts
  • Document retention policies for compliance requirements
  • Prefer object storage over HDFS for new deployments
  • Use versioning for data protection against accidental deletion
  • Target 128MB-1GB file sizes for query engine performance
  • Monitor and alert on storage cost anomalies by team/project
  • Test disaster recovery procedures annually
  • Allocate storage costs to teams for cost accountability
  • Run MinIO with at least 4 nodes and N/2 erasure coding parity for production
  • Benchmark storage performance before committing to architecture
  • Use Object Lock (WORM) for compliance-bound data
  • Plan data migration with pilot phase before bulk transfer
  • Model total cost of ownership including operations, transfer, and egress
  • Alert on storage growth anomalies and budget thresholds

References

© Kilo-Org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/data-distributed-storage of Kilo-Org/kilo-marketplace.

  • SKILL.md
  • LICENSE
  • local.patch
  • references/distributed-storage-advanced.md
  • references/distributed-storage-fundamentals.md
  • references/hdfs-architecture.md
  • references/s3-compatible-configs.md
  • references/s3-compatible.md
  • references/storage-deployment-guide.md

Open the folder on GitHubat commit ff51758

Compare with similar skills

Data Distributed Storage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Distributed Storage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Distributed Storage this skillKilo-Org/kilo-marketplace190—~4.9kAutomated safety check: PassMIT
Database Backupssickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: NotesMIT
Zarr PythonK-Dense-AI/scientific-agent-skills48k1 repos~2.2kAutomated safety check: NotesMIT
AWS S3majiayu000/claude-skill-registry6663 repos~3.1kAutomated safety check: PassMIT
Mail Timeveliovgroup/mail-time143—~1kAutomated safety check: PassBSD-3-Clause
Azure Storagemicrosoft/GitHub-Copilot-for-Azure2552 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • Database Backups

    sickn33/agentic-awesome-skills

    Implement database backup strategies. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    DatabasesAuto-check: notes
  • Zarr Python

    K-Dense-AI/scientific-agent-skills

    Stores and queries chunked N-D scientific arrays with Zarr-Python 3, including codecs, sharding, S3/GCS storage, and NumPy/Dask/Xarray integration.

    48k GitHub starsUsed in 1 repo~2.2k tokens
    DatabasesAuto-check: notes
  • AWS S3

    majiayu000/claude-skill-registry

    Configure S3 buckets, policies, and lifecycle rules. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 3 repos~3.1k tokens
    Backend & APIsAuto-check passed
  • Mail Time

    veliovgroup/mail-time

    A skill your agent uses when building, wiring, reviewing, or debugging MailTime and ostrio:mailer email queues for horizontally scaled Node.js, Bun, or Meteor apps.

    143 GitHub stars~1k tokensUpdated today
    DatabasesAuto-check passed
  • Azure Storage

    microsoft/GitHub-Copilot-for-Azure

    Official

    Azure Storage Services including Blob Storage, File Shares, Queue Storage, Table Storage, and Data Lake.

    255 GitHub starsUsed in 2 repos~1.3k tokens
    DatabasesAuto-check passed
  • AWS CLI Beast

    giuseppe-trisciuoglio/developer-kit

    Provides advanced AWS CLI patterns for managing EC2, Lambda, S3, DynamoDB, RDS, VPC, IAM, and CloudWatch.

    355 GitHub stars~1.7k tokensUpdated 28 days ago
    DatabasesAuto-check: notes

More from Kilo-Org/kilo-marketplace

All 85 skills in this repo
  • AzureML Project Scaffolding

    Kilo-Org/kilo-marketplace

    Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.

    190 GitHub stars~3.1k tokensUpdated 9 days ago
    Auto-check: notes
  • Jupyter Notebook Builder

    Kilo-Org/kilo-marketplace

    Creates, inspects, edits and runs Jupyter notebooks, scaffolding experiment or tutorial notebooks from templates and preferring a Jupyter MCP server over raw JSON edits.

    190 GitHub stars~1.3k tokensUpdated 9 days ago
    Auto-check passed
  • Tableau Dashboard Creator

    Kilo-Org/kilo-marketplace

    Takes a plain-language dashboard request through brand setup, data exploration, planning, an interactive HTML mock and a Tableau implementation spec.

    190 GitHub stars~3.8k tokensUpdated 9 days ago
    Auto-check: notes
  • Elasticsearch File Ingest

    Kilo-Org/kilo-marketplace

    Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    190 GitHub stars~2.8k tokensUpdated 9 days ago
    Auto-check passed
  • Nifi Flow Layout

    Kilo-Org/kilo-marketplace

    A skill your agent uses when arranging Apache NiFi processors, process groups, ports, comments, numbering, crossing connections, dense fan-in/fan-out, or reusable readable canvas layouts.

    190 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed
  • Splunk Ingest Processor Setup

    Kilo-Org/kilo-marketplace

    Render Cisco Data Fabric ingest-time routing workflows and Splunk Cloud Platform Ingest Processor setup plans with SPL2 pipelines, source types, destinations, lifecycle handoffs, queue and…

    190 GitHub stars~1.2k tokensUpdated 9 days ago
    Auto-check passed

Questions about Data Distributed Storage

What does Data Distributed Storage do?

A skill your agent uses when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads. Data Distributed Storage is an agent skill from Kilo-Org/kilo-marketplace. Use this skill when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads.

When should I use Data Distributed Storage?

Data Distributed Storage fits situations like: designing distributed storage for HDFS; any distributed file system for data workloads; : database storage engines; local filesystem tuning.

How do I install Data Distributed Storage in Claude Code?

Run `npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a claude-code`. Or copy the skill folder (skills/data-distributed-storage in Kilo-Org/kilo-marketplace) into .claude/skills/data-distributed-storage in your project. Claude Code loads it when a task matches its description.

How do I install Data Distributed Storage in Codex?

Run `npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a codex`. Or copy the skill folder (skills/data-distributed-storage in Kilo-Org/kilo-marketplace) into .agents/skills/data-distributed-storage in your project. Codex loads it when a task matches its description.

Can I use Data Distributed Storage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-distributed-storage, .gemini/skills/data-distributed-storage, .github/skills/data-distributed-storage and .opencode/skills/data-distributed-storage in your project.

What does Data Distributed Storage need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Distributed Storage is instructions for the agent only. Our summary lists: Docker.

Does Data Distributed Storage access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Distributed Storage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Distributed Storage use?

Data Distributed Storage is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Distributed Storage use?

About 4.9k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.4k tokens, read only when the agent opens those files.

What are the alternatives to Data Distributed Storage?

Skills that share tags, products or a category with Data Distributed Storage: Database Backups (sickn33/agentic-awesome-skills, 47k stars), Zarr Python (K-Dense-AI/scientific-agent-skills, 48k stars), AWS S3 (majiayu000/claude-skill-registry, 666 stars) and Mail Time (veliovgroup/mail-time, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Distributed Storage?

Kilo-Org (a GitHub organization) maintains it in Kilo-Org/kilo-marketplace, which has 190 GitHub stars. The repository holds 85 skills in this directory. The repository was last updated on September 28, 2026.

Source: Kilo-Org/kilo-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.