Database Backups
sickn33/agentic-awesome-skills
Implement database backup strategies. An agent skill from sickn33/agentic-awesome-skills.
A skill your agent uses when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads.
$ npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Kilo-Org/kilo-marketplace data-distributed-storage --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-distributed-storage .claude/skills/data-distributed-storage && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-distributed-storage" agent skill from https://github.com/Kilo-Org/kilo-marketplace/tree/main/skills/data-distributed-storage into .claude/skills/data-distributed-storage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-distributed-storage", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Kilo-Org/kilo-marketplace/tree/main/skills/data-distributed-storageType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Kilo-Org/kilo-marketplace data-distributed-storage --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/data-distributed-storage .agents/skills/data-distributed-storage && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-distributed-storage" agent skill from https://github.com/Kilo-Org/kilo-marketplace/tree/main/skills/data-distributed-storage into .agents/skills/data-distributed-storage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-distributed-storage", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Kilo-Org/kilo-marketplace data-distributed-storage --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/data-distributed-storage .cursor/skills/data-distributed-storage && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-distributed-storage" agent skill from https://github.com/Kilo-Org/kilo-marketplace/tree/main/skills/data-distributed-storage into .cursor/skills/data-distributed-storage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-distributed-storage", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Kilo-Org/kilo-marketplace.git --path skills/data-distributed-storage--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Kilo-Org/kilo-marketplace data-distributed-storage --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/data-distributed-storage .gemini/skills/data-distributed-storage && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-distributed-storage" agent skill from https://github.com/Kilo-Org/kilo-marketplace/tree/main/skills/data-distributed-storage into .gemini/skills/data-distributed-storage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-distributed-storage", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Kilo-Org/kilo-marketplace data-distributed-storageInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/data-distributed-storage .github/skills/data-distributed-storage && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-distributed-storage" agent skill from https://github.com/Kilo-Org/kilo-marketplace/tree/main/skills/data-distributed-storage into .github/skills/data-distributed-storage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-distributed-storage", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Kilo-Org/kilo-marketplace data-distributed-storage --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/data-distributed-storage .opencode/skills/data-distributed-storage && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-distributed-storage" agent skill from https://github.com/Kilo-Org/kilo-marketplace/tree/main/skills/data-distributed-storage into .opencode/skills/data-distributed-storage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-distributed-storage", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-distributed-storageA skill your agent uses when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads.
Data Distributed Storage is an agent skill from Kilo-Org/kilo-marketplace. Use this skill when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads. This skill enforces: storage backend selection, data durability and replication, file format selection, partitioning and compression, lifecycle policies, storage tiering, and cost optimization. Do NOT use for: database storage engines, local filesystem tuning, or content delivery networks.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `references/distributed-storage-advanced.md`, `references/distributed-storage-fundamentals.md` and `references/hdfs-architecture.md`).
It sits in Databases, covering File uploads and storage and Database administration. The repository describes itself as: Kilo Marketplace - A curated collection of Skills, MCP Servers, and Modes for enhancing AI agent capabilities across the Kilo ecosystem—including Kilo Code (VS Code extension)… The licence is MIT.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ff51758. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and json).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Distributed Storage loads about 4.9k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 1,064 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Kilo-Org/kilo-marketplace at commit ff51758, republished under its MIT licence (© Kilo-Org). 1,064 words, ~4,869 tokens.
.claude/skills/data-distributed-storage/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Design distributed storage systems for data workloads. Select the right storage backend (HDFS, S3, ADLS, GCS, MinIO), configure durability and replication, optimize file formats and compression, and implement lifecycle policies.
Exact user phrases: "HDFS", "S3", "ADLS", "GCS", "MinIO", "distributed storage", "object storage", "file system", "data lake storage", "storage tier", "lifecycle policy", "replication factor", "erasure coding", "storage cost".
Distributed storage architecture with backend selection, tiering strategy, lifecycle policy, and cost model.
# Storage backend
# Tier configuration
# Lifecycle rules
# Cost projection| Feature | AWS S3 | ADLS Gen2 | GCS | MinIO |
|---|---|---|---|---|
| Consistency | Read-after-write | Strong | Strong | Strong |
| Auth | IAM roles, bucket policies | RBAC, SAS tokens | IAM, service accounts | JWT, OIDC, LDAP |
| Encryption | SSE-S3/KMS/CSE | SSE-AES/KMS/CMK | Google-managed/CMEK/CSE | KMS, auto-encryption |
| Lifecycle | Transition, expiry, versioning | Tiering, soft-delete | Nearline/Coldline/Archive | Bucket lifecycle |
| Max object size | 5 TB | 4.75 TB | 5 TB | 100 TB (configurable) |
| S3 Compatible | Native | Yes (via gateway) | Yes (XML API) | Native |
| Cost (per TB/mo) | ~$23 (standard) | ~$20 (hot) | ~$20 (standard) | Hardware + ops |
| Durability | 99.999999999% (11 9s) | 99.999999999% (11 9s) | 99.999999999% (11 9s) | Configurable |
| Availability SLA | 99.99% | 99.9% | 99.95% | Depends on setup |
Cloud provider?
├── AWS → S3 (best integration with AWS ecosystem)
├── Azure → ADLS Gen2 (best with Azure AD, Active Directory)
├── GCP → GCS (strong consistency, best for Google ecosystem)
├── Multi-cloud → MinIO (S3-compatible layer across clouds)
└── On-premise / air-gapped → MinIO or HDFS
Primary workload?
├── Data lake with compute engines → Object store (S3/ADLS/GCS)
├── Legacy Hadoop/Spark without cloud → HDFS
├── Private cloud, edge, or air-gapped → MinIO
└── High-performance computing → Parallel file system (Lustre, GPFS)| Tier | Storage Class | Latency | Cost/TB/mo | Use Case |
|---|---|---|---|---|
| Hot | S3 Standard, ADLS Hot | Millisecond | ~$23 | Active data, daily pipelines, ML training |
| Warm | S3 Infrequent Access | Millisecond | ~$12.50 | Monthly queries, staging data |
| Cold | S3 Glacier Instant | Millisecond (instant) | ~$4 | Quarterly access, compliance retention |
| Archive | S3 Glacier Deep Archive | 12 hours | ~$1 | Annual access, legal holds |
Hot: last 30-90 days of data. Warm: 90 days to 1 year. Cold: 1-3 years. Archive: 3-7+ years (retention-based). Intelligent tiering (S3) auto-moves objects between hot and warm based on access patterns.
{
"Rules": [
{
"Id": "TierAndExpire",
"Status": "Enabled",
"Filter": { "Prefix": "logs/" },
"Transitions": [
{ "Days": 30, "StorageClass": "STANDARD_IA" },
{ "Days": 90, "StorageClass": "GLACIER" },
{ "Days": 365, "StorageClass": "DEEP_ARCHIVE" }
],
"Expiration": { "Days": 2555 }
},
{
"Id": "ExpireTempFiles",
"Status": "Enabled",
"Filter": { "Prefix": "tmp/" },
"Expiration": { "Days": 7 }
}
]
}S3/ADLS/GCS provide 11 9s of durability via erasure coding across multiple devices/facilities. Replication is handled by the cloud provider. For additional protection: cross-region replication (CRR) for disaster recovery, same-region replication (SRR) for compliance.
Default replication factor: 3 (1 primary, 2 replicas on different racks). Storage overhead: 3x. Erasure coding reduces overhead: RS-6-3 (1.33x overhead), RS-10-4 (1.5x overhead). EC recommended for cold data on HDFS (> 100TB).
Cross-region backup for critical data. Versioning enabled on all buckets. Replication Rules: replicate Tier 1 data to DR region synchronously or asynchronously. Test restore quarterly.
| Codec | Speed | Compression Ratio | Use Case |
|---|---|---|---|
| ZSTD (level 3) | Very fast | Good (2-4x) | Default for Parquet/ORC |
| GZIP (level 6) | Moderate | Best (3-6x) | Archive, cold data |
| Snappy | Fastest | Fair (1.5-3x) | Streaming, hot data |
| LZ4 | Fastest | Fair (1.5-2x) | Logs, real-time |
| BZIP2 | Slow | Best (4-8x) | Archive only |
Compress data at the file format level (Parquet/ORC) rather than at the transport level. Avoid double compression. Test: compress a representative dataset with each codec, measure speed and ratio.
Organize storage by: source/system, data domain, date partition, file. Example: s3://data-lake/raw/salesforce/orders/year=2026/month=05/day=01/load_id=abc123/orders_001.parquet. This layout enables: partition pruning, easy lifecycle management, clear ownership boundaries.
Small files cause performance problems for query engines. Target file size: 128MB-1GB. Use compaction jobs to merge small files. Monitor average file size per dataset.
Active NameNode manages filesystem metadata (namespace, blocks, locations). Standby NameNode for HA via QJM (Quorum Journal Manager). DataNodes store block data and report to NameNode. Block size: 128MB or 256MB default. Rack awareness for replica placement.
# hdfs-site.xml
dfs.replication: 3
dfs.block.size: 268435456 # 256MB
dfs.namenode.handler.count: 100
dfs.datanode.handler.count: 50
dfs.namenode.gc.time.threshold: 60s
dfs.namenode.checkpoint.period: 3600 # 1 hour
dfs.permissions.enabled: trueServer-side encryption: SSE-S3 (AES-256, S3-managed keys), SSE-KMS (AWS KMS-managed keys), SSE-C (customer-provided keys). Client-side encryption: encrypt before upload, decrypt after download. For compliance: SSE-KMS with audit logging.
TLS 1.2+ for all S3/ADLS/GCS API calls. Enforce HTTPS-only bucket policies. HDFS: enable SSL for RPC and data transfer.
{
"Statement": [
{
"Effect": "Deny",
"Principal": "*",
"Action": "s3:*",
"Resource": "arn:aws:s3:::data-lake/*",
"Condition": {
"Bool": { "aws:SecureTransport": "false" }
}
}
]
}Use intelligent tiering for automatic cost savings on variable-access data. Request (S3) Reduced Redundancy for non-critical data (lower durability = lower cost). Reserved capacity for predictable usage. Monitor and alert on cost anomalies.
Track: storage cost per bucket/container, data transfer costs (egress), request costs (PUT/GET/LIST), lifecycle transition costs. Allocate costs to teams/domains via tag-based cost allocation.
Cloud adoption strategy?
├── Single cloud → Native object store (S3/ADLS/GCS)
├── Multi-cloud → MinIO (abstraction layer)
├── On-premise
│ ├── Hadoop ecosystem → HDFS
│ └── Kubernetes-native → MinIO (S3-compatible)
└── Edge / IoT → MinIO (lightweight, S3 API)Data access pattern?
├── Accessed frequently → Hot tier (delete when no longer needed)
├── Accessed occasionally → Intelligent tiering or manual tiering
├── Compliance retention → Write once, tier to cold → delete after retention
└── Temporary data → Hot tier → delete in 7-30 daysMinIO runs as a distributed system across multiple nodes. Minimum 4 nodes for erasure coding protection. Each node: 4+ drives, SSD/NVMe preferred for performance. Deployed via: Docker Compose (dev/tiny), Kubernetes Operator (production), bare metal (HPC).
# MinIO Kubernetes Operator
apiVersion: minio.min.io/v2
kind: Tenant
metadata:
name: data-lake-storage
spec:
image: ${MINIO_IMAGE_REF:?Set MINIO_IMAGE_REF to quay.io/minio/minio@sha256:<reviewed-digest>}
credsSecret:
name: minio-creds
pools:
- servers: 4
volumesPerServer: 4
volumeClaimTemplate:
spec:
storageClassName: ssd
accessModes: [ReadWriteOnce]
resources:
requests:
storage: 4Ti
mountPath: /export
requestAutoCert: true
buckets:
- name: raw
- name: staging
- name: analyticsMinIO uses Reed-Solomon erasure coding. Default parity: N/2 drives (max protection). Configurable: standard (EC:4) for 8-drive setup, reduced (EC:2) for performance. Read tolerance: lose up to N/2 drives. Write tolerance: lose up to N/2 drives. Storage overhead: 2x for N/2 parity, 1.33x for N/4 parity.
# Erasure coding parity by deployment size
# 4 nodes × 4 drives = 16 drives
# EC:8 → tolerate 8 drive failures, 2x overhead
# EC:4 → tolerate 4 drive failures, 1.33x overhead
# EC:2 → tolerate 2 drive failures, 1.14x overhead
# Set via environment variable
MINIO_STORAGE_CLASS_STANDARD: "EC:4"
MINIO_STORAGE_CLASS_RRS: "EC:2" # Reduced redundancy for temp data# Drive-level
# Use XFS filesystem with noatime,nodiratime mount options
# Separate write-intensive (metadata, WAL) from read-intensive drives
# Network-level
# MinIO uses S3-compatible HTTP/REST
# Enable HTTP/2 for multiplexing
# Use load balancer (nginx, HAProxy) for multi-node distribution
# Kernel tuning
# net.core.rmem_max: 134217728 # 128MB receive buffer
# net.core.wmem_max: 134217728 # 128MB send buffer
# net.ipv4.tcp_congestion_control: bbr # BBR for better throughput
# vm.dirty_ratio: 30 # Page cache dirty ratio
# vm.dirty_background_ratio: 10# AWS S3 Transfer Acceleration
# Uses CloudFront edge locations, TCP optimization
# Enable: s3.put_bucket_accelerate_configuration
# Best for: cross-region uploads, large objects > 1GB
# Cost: premium per GB transferred
# Test: s3-accelerate-speedtest.s3-accelerate.amazonaws.com
# Azure Data Box / Import-Export
# Physical transfer: Data Box (80TB), Data Box Disk (8TB)
# For: > 10TB initial loads, slow networks, air-gapped
# Timeline: order → receive → load → return → ingest (2-4 weeks)
# GCS Storage Transfer Service
# Online: from S3, HTTP, or GS URLs
# Schedule: one-time or recurring
# Best for: ongoing sync from S3 to GCSmigration_phases:
phase_1_assessment:
- inventory current storage (total TB, object count, avg size)
- identify access patterns (hot vs cold, frequency)
- document retention and compliance requirements
- estimate egress costs and transfer time
phase_2_pilot:
- migrate 1-2 TB of non-critical data first
- validate performance, cost, and access patterns
- document issues and adjust strategy
phase_3_bulk:
- parallel multi-threaded copy (aws s3 sync, rclone, distcp)
- validate data integrity checksums post-transfer
- monitor for failures (rate limits, timeouts)
phase_4_cutover:
- final sync of delta changes
- update application configs to point to new storage
- redirect DNS / endpoint to new location
- decommission old storage after 30-day validation window# S3 Object Lock (compliance mode)
# Prevents object deletion or overwrite for retention period
# Governance mode: users with special permissions can override
# Compliance mode: no one (including root) can override
# Legal hold: prevents deletion regardless of retention
object_lock_examples:
- name: "Financial records (SOX)"
mode: COMPLIANCE
retention_days: 2555 # 7 years
bucket: "data-lake/finance/"
- name: "Logs (internal policy)"
mode: GOVERNANCE
retention_days: 365 # 1 year
bucket: "data-lake/logs/"
# ADLS Gen2: immutability policy
# Set at container level
# Time-based retention or legal hold# Data residency requirements
# GDPR: data must stay in EU region
# Financial services: data within country borders
# Healthcare: HIPAA requires US-only storage (for US patients)
# Implementation:
# - Per-region buckets with IAM restrictions
# - S3 bucket policies denying cross-region replication for sensitive data
# - MinIO multi-tenant deployment per region
# - Regular audit to verify no data leaves allowed regionscost_components:
storage:
- per-GB-month by tier (hot, warm, cold, archive)
- minimum storage duration charges (Glazer: 90 days min)
- early deletion fees (cold/archive tiers)
operations:
- PUT/COPY/POST/LIST requests (per 1000)
- GET/SELECT requests (per 1000)
- lifecycle transition requests (per 1000)
data_transfer:
- upload (usually free)
- download / egress (per GB, tiered pricing)
- cross-region replication (per GB)
- transfer acceleration (premium per GB)
additional:
- encryption (KMS: per key, per API call)
- monitoring (CloudWatch, metrics per custom metric)
- backup / replication (CRR storage in destination region)# Monthly storage cost estimate
cost_estimate:
hot_data:
volume: 50 TB
unit_cost: $23/TB/mo
subtotal: $1,150/mo
warm_data:
volume: 150 TB
unit_cost: $12.50/TB/mo
subtotal: $1,875/mo
cold_data:
volume: 300 TB
unit_cost: $4/TB/mo
subtotal: $1,200/mo
archive_data:
volume: 500 TB
unit_cost: $1/TB/mo
subtotal: $500/mo
operations:
requests: $200/mo # Based on 10M PUT + 100M GET
data_transfer:
egress: $300/mo # Based on 10TB egress
monitoring_backup:
cross_region_replication: $500/mo # 50TB replicated
monitoring: $100/mo
total_monthly: $5,825/mo
total_annual: $69,900/yr# Object store monitoring (CloudWatch, Azure Monitor, GCS Ops)
metrics:
storage:
- BucketSizeBytes (by storage tier)
- NumberOfObjects
- AverageObjectSize
requests:
- AllRequests (count)
- GetRequests / PutRequests / ListRequests
- 4xxErrors / 5xxErrors
- FirstByteLatency / TotalRequestLatency
throughput:
- BytesDownloaded / BytesUploaded
- GetBandwidth / PutBandwidth (MB/s)
cost:
- StorageCost (by bucket/tag)
- RequestCost
- DataTransferCost
- TotalCostalerts:
- name: "Storage growth anomaly"
metric: BucketSizeBytes
threshold: "> 20% week-over-week increase"
severity: warning
- name: "Error rate spike"
metric: 5xxErrors
threshold: "> 1% of total requests"
severity: critical
- name: "Latency degradation"
metric: FirstByteLatency
threshold: "p99 > 500ms"
severity: warning
- name: "Budget threshold"
metric: TotalCost
threshold: "> 80% of monthly budget"
severity: warning
- name: "Lifecycle failure"
metric: LifecycleTransitions
threshold: "transitions failed > 10 per hour"
severity: warningWorkload type?
├── Hot data, frequent queries → ZSTD (level 3) — fast decompression
├── Cold data, archive → GZIP (level 6) — best compression ratio
├── Streaming, low-latency → Snappy or LZ4 — minimal CPU overhead
├── Columnar formats (Parquet/ORC)
│ ├── Default → ZSTD (best balance)
│ └── Legacy compatibility → Snappy
└── JSON/CSV files
├── Compressed → GZIP (most tools support it)
└── Splittable → BZIP2 or LZ4Compliance requirement?
├── Disaster recovery (region outage)
│ ├── RTO < 1 hour → Synchronous CRR
│ └── RTO < 4 hours → Asynchronous CRR
├── Data sovereignty (must stay in region)
│ └── Same-region replication (SRR) + backup to another AZ
├── GDPR right to erasure
│ ├── Replicate selectively (no unnecessary copies)
│ └── Document replication topology for audit
└── No compliance requirement
└── No replication — rely on cloud provider durability© Kilo-Org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (references) in skills/data-distributed-storage of Kilo-Org/kilo-marketplace.
Open the folder on GitHubat commit ff51758
Data Distributed Storage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Distributed Storage this skillKilo-Org/kilo-marketplace | 190 | — | ~4.9k | Automated safety check: Pass | MIT | |
| Database Backupssickn33/agentic-awesome-skills | 47k | 2 repos | ~3.1k | Automated safety check: Notes | MIT | |
| Zarr PythonK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~2.2k | Automated safety check: Notes | MIT | |
| AWS S3majiayu000/claude-skill-registry | 666 | 3 repos | ~3.1k | Automated safety check: Pass | MIT | |
| Mail Timeveliovgroup/mail-time | 143 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | |
| Azure Storagemicrosoft/GitHub-Copilot-for-Azure | 255 | 2 repos | ~1.3k | Automated safety check: Pass | MIT |
sickn33/agentic-awesome-skills
Implement database backup strategies. An agent skill from sickn33/agentic-awesome-skills.
K-Dense-AI/scientific-agent-skills
Stores and queries chunked N-D scientific arrays with Zarr-Python 3, including codecs, sharding, S3/GCS storage, and NumPy/Dask/Xarray integration.
majiayu000/claude-skill-registry
Configure S3 buckets, policies, and lifecycle rules. An agent skill from majiayu000/claude-skill-registry.
veliovgroup/mail-time
A skill your agent uses when building, wiring, reviewing, or debugging MailTime and ostrio:mailer email queues for horizontally scaled Node.js, Bun, or Meteor apps.
microsoft/GitHub-Copilot-for-Azure
Azure Storage Services including Blob Storage, File Shares, Queue Storage, Table Storage, and Data Lake.
giuseppe-trisciuoglio/developer-kit
Provides advanced AWS CLI patterns for managing EC2, Lambda, S3, DynamoDB, RDS, VPC, IAM, and CloudWatch.
Kilo-Org/kilo-marketplace
Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.
Kilo-Org/kilo-marketplace
Creates, inspects, edits and runs Jupyter notebooks, scaffolding experiment or tutorial notebooks from templates and preferring a Jupyter MCP server over raw JSON edits.
Kilo-Org/kilo-marketplace
Takes a plain-language dashboard request through brand setup, data exploration, planning, an interactive HTML mock and a Tableau implementation spec.
Kilo-Org/kilo-marketplace
Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.
Kilo-Org/kilo-marketplace
A skill your agent uses when arranging Apache NiFi processors, process groups, ports, comments, numbering, crossing connections, dense fan-in/fan-out, or reusable readable canvas layouts.
Kilo-Org/kilo-marketplace
Render Cisco Data Fabric ingest-time routing workflows and Splunk Cloud Platform Ingest Processor setup plans with SPL2 pipelines, source types, destinations, lifecycle handoffs, queue and…
Categories
A skill your agent uses when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads. Data Distributed Storage is an agent skill from Kilo-Org/kilo-marketplace. Use this skill when designing distributed storage for HDFS, S3, ADLS, GCS, MinIO, NFS, or any distributed file system for data workloads.
Data Distributed Storage fits situations like: designing distributed storage for HDFS; any distributed file system for data workloads; : database storage engines; local filesystem tuning.
Run `npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a claude-code`. Or copy the skill folder (skills/data-distributed-storage in Kilo-Org/kilo-marketplace) into .claude/skills/data-distributed-storage in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a codex`. Or copy the skill folder (skills/data-distributed-storage in Kilo-Org/kilo-marketplace) into .agents/skills/data-distributed-storage in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kilo-Org/kilo-marketplace --skill data-distributed-storage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-distributed-storage, .gemini/skills/data-distributed-storage, .github/skills/data-distributed-storage and .opencode/skills/data-distributed-storage in your project.
SKILL.md names no scripts, command-line tools or credentials: Data Distributed Storage is instructions for the agent only. Our summary lists: Docker.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Data Distributed Storage is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Data Distributed Storage: Database Backups (sickn33/agentic-awesome-skills, 47k stars), Zarr Python (K-Dense-AI/scientific-agent-skills, 48k stars), AWS S3 (majiayu000/claude-skill-registry, 666 stars) and Mail Time (veliovgroup/mail-time, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Kilo-Org (a GitHub organization) maintains it in Kilo-Org/kilo-marketplace, which has 190 GitHub stars. The repository holds 85 skills in this directory. The repository was last updated on September 28, 2026.
Source: Kilo-Org/kilo-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.