Agent skill

Senior Data Engineer

by benchflow-ai in benchflow-ai/skillsbench

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

MITAuto-check passedData & Analytics

Install Senior Data Engineer

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill senior-data-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench senior-data-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/flink-query/environment/skills/senior-data-engineer .claude/skills/senior-data-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
senior-data-engineer
GitHub stars
1.8k
Token cost
~5.9k tokens
SKILL.md length
1,813 words
Files
14 (incl. scripts, references, assets)
Skills in repo
180
Repo updated
First seen
Licence
MIT

At a glance

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

  • Works in 5 steps: Building Production Data Pipelines → Data Quality Management → Data Modeling & Transformation → …
  • Designing data architectures
  • SKILL.md covers Core Capabilities, Key Workflows, Overview and Quick Start, plus 5 more sections
  • Runs Python scripts from its folder; calls python

What it does

Senior Data Engineer is an agent skill from benchflow-ai/skillsbench. World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, Flink, Kinesis, and modern data stack. Includes data modeling, pipeline orchestration, data quality, streaming quality monitoring, and DataOps. Use when designing data architectures, building batch or streaming data pipelines, optimizing data workflows, or implementing data governance.

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including scripts, reference files and assets (for example `HOW_TO_USE.md`, `assets/data-quality-checklist.md` and `assets/pipeline-design-template.md`). Compatibility notes: Python 3.8+; platforms: macos, linux, windows

It sits in Data & Analytics, covering Data pipelines and ETL. It works with Apache Kafka, Apache Airflow, dbt and Python. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is MIT.

When your agent uses it

  • Designing data architectures
  • Streaming data pipelines
  • Optimizing data workflows
  • Implementing data governance

Example prompts

  • “/senior-data-engineer”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Python 3.8+; platforms: macos, linux, windows

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Building Production Data Pipelines
  2. Data Quality Management
  3. Data Modeling & Transformation
  4. Performance Optimization
  5. Building Real-Time Streaming Pipelines

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Python 3.8+; platforms: macos, linux, windows

    From compatibility in the SKILL.md frontmatter.

Context cost

Senior Data Engineer loads about 5.9k tokens when it runs, and up to ~43k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 1,813 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~43k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its MIT licence (© benchflow-ai). 1,813 words, ~5,917 tokens.

Download SKILL.mdSave it as .claude/skills/senior-data-engineer/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
senior-data-engineer
description
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, Flink, Kinesis, and modern data stack. Includes data modeling, pipeline orchestration, data quality, streaming quality monitoring, and DataOps. Use when designing data architectures, building batch or streaming data pipelines, optimizing data workflows, or implementing data governance.
compatibility
Python 3.8+; platforms: macos, linux, windows
title
Senior Data Engineer Skill Package
domain
engineering
subdomain
data-engineering
difficulty
advanced
time-saved
TODO: Quantify time savings
frequency
TODO: Estimate usage frequency
use-cases
Designing data pipelines for ETL/ELT processes, Building data warehouses and data lakes, Implementing data quality and governance frameworks, Creating…
tech-stack
Python, SQL, Apache Spark, Airflow, dbt, Apache Kafka, Apache Flink, AWS Kinesis, Spark Structured Streaming, Kafka Streams, PostgreSQL, BigQuery, Snowflake…
stats.downloads
0
stats.stars
0

Senior Data Engineer

Core Capabilities

  • Batch Pipeline Orchestration - Design and implement production-ready ETL/ELT pipelines with Airflow, intelligent dependency resolution, retry logic, and comprehensive monitoring
  • Real-Time Streaming - Build event-driven streaming pipelines with Kafka, Flink, Kinesis, and Spark Streaming with exactly-once semantics and sub-second latency
  • Data Quality Management - Comprehensive batch and streaming data quality validation covering completeness, accuracy, consistency, timeliness, and validity
  • Streaming Quality Monitoring - Track consumer lag, data freshness, schema drift, throughput, and dead letter queue rates for streaming pipelines
  • Performance Optimization - Analyze and optimize pipeline performance with query optimization, Spark tuning, and cost analysis recommendations

Key Workflows

Workflow 1: Build ETL Pipeline

Time: 2-4 hours

Steps:

  1. Design pipeline architecture using Lambda, Kappa, or Medallion pattern
  2. Configure YAML pipeline definition with sources, transformations, targets
  3. Generate Airflow DAG with pipeline_orchestrator.py
  4. Define data quality validation rules
  5. Deploy and configure monitoring/alerting

Expected Output: Production-ready ETL pipeline with 99%+ success rate, automated quality checks, and comprehensive monitoring

Workflow 2: Build Real-Time Streaming Pipeline

Time: 3-5 days

Steps:

  1. Select streaming architecture (Kappa vs Lambda) based on requirements
  2. Configure streaming pipeline YAML (sources, processing, sinks, quality)
  3. Generate Kafka configurations with kafka_config_generator.py
  4. Generate Flink/Spark job scaffolding with stream_processor.py
  5. Deploy and monitor with streaming_quality_validator.py

Expected Output: Streaming pipeline processing 10K+ events/sec with P99 latency < 1s, exactly-once delivery, and real-time quality monitoring

World-class data engineering for production-grade data systems, scalable pipelines, and enterprise data platforms.

Overview

This skill provides comprehensive expertise in data engineering fundamentals through advanced production patterns. From designing medallion architectures to implementing real-time streaming pipelines, it covers the full spectrum of modern data engineering including ETL/ELT design, data quality frameworks, pipeline orchestration, and DataOps practices.

What This Skill Provides:

  • Production-ready pipeline templates (Airflow, Spark, dbt)
  • Comprehensive data quality validation framework
  • Performance optimization and cost analysis tools
  • Data architecture patterns (Lambda, Kappa, Medallion)
  • Complete DataOps CI/CD workflows

Best For:

  • Building scalable data pipelines for enterprise systems
  • Implementing data quality and governance frameworks
  • Optimizing ETL performance and cloud costs
  • Designing modern data architectures (lake, warehouse, lakehouse)
  • Production ML/AI data infrastructure

Quick Start

Pipeline Orchestration
bash
# Generate Airflow DAG from configuration
python scripts/pipeline_orchestrator.py --config pipeline_config.yaml --output dags/

# Validate pipeline configuration
python scripts/pipeline_orchestrator.py --config pipeline_config.yaml --validate

# Use incremental load template
python scripts/pipeline_orchestrator.py --template incremental --output dags/
Data Quality Validation
bash
# Validate CSV file with quality checks
python scripts/data_quality_validator.py --input data/sales.csv --output report.html

# Validate database table with custom rules
python scripts/data_quality_validator.py \
    --connection postgresql://user:pass@host/db \
    --table sales_transactions \
    --rules rules/sales_validation.yaml \
    --threshold 0.95
Performance Optimization
bash
# Analyze pipeline performance and get recommendations
python scripts/etl_performance_optimizer.py \
    --airflow-db postgresql://host/airflow \
    --dag-id sales_etl_pipeline \
    --days 30 \
    --optimize

# Analyze Spark job performance
python scripts/etl_performance_optimizer.py \
    --spark-history-server http://spark-history:18080 \
    --app-id app-20250115-001
Real-Time Streaming
bash
# Validate streaming pipeline configuration
python scripts/stream_processor.py --config streaming_config.yaml --validate

# Generate Kafka topic and client configurations
python scripts/kafka_config_generator.py \
    --topic user-events \
    --partitions 12 \
    --replication 3 \
    --output kafka/topics/

# Generate exactly-once producer configuration
python scripts/kafka_config_generator.py \
    --producer \
    --profile exactly-once \
    --output kafka/producer.properties

# Generate Flink job scaffolding
python scripts/stream_processor.py \
    --config streaming_config.yaml \
    --mode flink \
    --generate \
    --output flink-jobs/

# Monitor streaming quality
python scripts/streaming_quality_validator.py \
    --lag --consumer-group events-processor --threshold 10000 \
    --freshness --topic processed-events --max-latency-ms 5000 \
    --output streaming-health-report.html

Core Workflows

1. Building Production Data Pipelines

Steps:

  1. Design Architecture: Choose pattern (Lambda, Kappa, Medallion) based on requirements
  2. Configure Pipeline: Create YAML configuration with sources, transformations, targets
  3. Generate DAG: python scripts/pipeline_orchestrator.py --config config.yaml
  4. Add Quality Checks: Define validation rules for data quality
  5. Deploy & Monitor: Deploy to Airflow, configure alerts, track metrics

Pipeline Patterns: See frameworks.md for Lambda Architecture, Kappa Architecture, Medallion Architecture (Bronze/Silver/Gold), and Microservices Data patterns.

Templates: See templates.md for complete Airflow DAG templates, Spark job templates, dbt models, and Docker configurations.

2. Data Quality Management

Steps:

  1. Define Rules: Create validation rules covering completeness, accuracy, consistency
  2. Run Validation: python scripts/data_quality_validator.py --rules rules.yaml
  3. Review Results: Analyze quality scores and failed checks
  4. Integrate CI/CD: Add validation to pipeline deployment process
  5. Monitor Trends: Track quality scores over time

Quality Framework: See frameworks.md for complete Data Quality Framework covering all dimensions (completeness, accuracy, consistency, timeliness, validity).

Validation Templates: See templates.md for validation configuration examples and Python API usage.

3. Data Modeling & Transformation

Steps:

  1. Choose Modeling Approach: Dimensional (Kimball), Data Vault 2.0, or One Big Table
  2. Design Schema: Define fact tables, dimensions, and relationships
  3. Implement with dbt: Create staging, intermediate, and mart models
  4. Handle SCD: Implement slowly changing dimension logic (Type 1/2/3)
  5. Test & Deploy: Run dbt tests, generate documentation, deploy

Modeling Patterns: See frameworks.md for Dimensional Modeling (Kimball), Data Vault 2.0, One Big Table (OBT), and SCD implementations.

dbt Templates: See templates.md for complete dbt model templates including staging, intermediate, fact tables, and SCD Type 2 logic.

4. Performance Optimization

Steps:

  1. Profile Pipeline: Run performance analyzer on recent pipeline executions
  2. Identify Bottlenecks: Review execution time breakdown and slow tasks
  3. Apply Optimizations: Implement recommendations (partitioning, indexing, batching)
  4. Tune Spark Jobs: Optimize memory, parallelism, and shuffle settings
  5. Measure Impact: Compare before/after metrics, track cost savings

Optimization Strategies: See frameworks.md for performance best practices including partitioning strategies, query optimization, and Spark tuning.

Analysis Tools: See tools.md for complete documentation on etl_performance_optimizer.py with query analysis and Spark tuning.

5. Building Real-Time Streaming Pipelines

Steps:

  1. Architecture Selection: Choose Kappa (streaming-only) or Lambda (batch + streaming) architecture
  2. Configure Pipeline: Create YAML config with sources, processing engine, sinks, quality thresholds
  3. Generate Kafka Configs: python scripts/kafka_config_generator.py --topic events --partitions 12
  4. Generate Job Scaffolding: python scripts/stream_processor.py --mode flink --generate
  5. Deploy Infrastructure: Use Docker Compose for local dev, Kubernetes for production
  6. Monitor Quality: python scripts/streaming_quality_validator.py --lag --freshness --throughput

Streaming Patterns: See frameworks.md for stateful processing, stream joins, windowing, exactly-once semantics, and CDC patterns.

Templates: See templates.md for Flink DataStream jobs, Kafka Streams applications, PyFlink templates, and Docker Compose configurations.

Python Tools

pipeline_orchestrator.py

Automated Airflow DAG generation with intelligent dependency resolution and monitoring.

Key Features:

  • Generate production-ready DAGs from YAML configuration
  • Automatic task dependency resolution
  • Built-in retry logic and error handling
  • Multi-source support (PostgreSQL, S3, BigQuery, Snowflake)
  • Integrated quality checks and alerting

Usage:

bash
# Basic DAG generation
python scripts/pipeline_orchestrator.py --config pipeline_config.yaml --output dags/

# With validation
python scripts/pipeline_orchestrator.py --config config.yaml --validate

# From template
python scripts/pipeline_orchestrator.py --template incremental --output dags/

Complete Documentation: See tools.md for full configuration options, templates, and integration examples.

data_quality_validator.py

Comprehensive data quality validation framework with automated checks and reporting.

Capabilities:

  • Multi-dimensional validation (completeness, accuracy, consistency, timeliness, validity)
  • Great Expectations integration
  • Custom business rule validation
  • HTML/PDF report generation
  • Anomaly detection
  • Historical trend tracking

Usage:

bash
# Validate with custom rules
python scripts/data_quality_validator.py \
    --input data/sales.csv \
    --rules rules/sales_validation.yaml \
    --output report.html

# Database table validation
python scripts/data_quality_validator.py \
    --connection postgresql://host/db \
    --table sales_transactions \
    --threshold 0.95

Complete Documentation: See tools.md for rule configuration, API usage, and integration patterns.

etl_performance_optimizer.py

Pipeline performance analysis with actionable optimization recommendations.

Capabilities:

  • Airflow DAG execution profiling
  • Bottleneck detection and analysis
  • SQL query optimization suggestions
  • Spark job tuning recommendations
  • Cost analysis and optimization
  • Historical performance trending

Usage:

bash
# Analyze Airflow DAG
python scripts/etl_performance_optimizer.py \
    --airflow-db postgresql://host/airflow \
    --dag-id sales_etl_pipeline \
    --days 30 \
    --optimize

# Spark job analysis
python scripts/etl_performance_optimizer.py \
    --spark-history-server http://spark-history:18080 \
    --app-id app-20250115-001

Complete Documentation: See tools.md for profiling options, optimization strategies, and cost analysis.

stream_processor.py

Streaming pipeline configuration generator and validator for Kafka, Flink, and Kinesis.

Capabilities:

  • Multi-platform support (Kafka, Flink, Kinesis, Spark Streaming)
  • Configuration validation with best practice checks
  • Flink/Spark job scaffolding generation
  • Kafka topic configuration generation
  • Docker Compose for local streaming stacks
  • Exactly-once semantics configuration

Usage:

bash
# Validate configuration
python scripts/stream_processor.py --config streaming_config.yaml --validate

# Generate Kafka configurations
python scripts/stream_processor.py --config streaming_config.yaml --mode kafka --generate

# Generate Flink job scaffolding
python scripts/stream_processor.py --config streaming_config.yaml --mode flink --generate --output flink-jobs/

# Generate Docker Compose for local development
python scripts/stream_processor.py --config streaming_config.yaml --mode docker --generate

Complete Documentation: See tools.md for configuration format, validation checks, and generated outputs.

streaming_quality_validator.py

Real-time streaming data quality monitoring with comprehensive health scoring.

Capabilities:

  • Consumer lag monitoring with thresholds
  • Data freshness validation (P50/P95/P99 latency)
  • Schema drift detection
  • Throughput analysis (events/sec, bytes/sec)
  • Dead letter queue rate monitoring
  • Overall quality scoring with recommendations
  • Prometheus metrics export

Usage:

bash
# Monitor consumer lag
python scripts/streaming_quality_validator.py \
    --lag --consumer-group events-processor --threshold 10000

# Monitor data freshness
python scripts/streaming_quality_validator.py \
    --freshness --topic processed-events --max-latency-ms 5000

# Full quality validation
python scripts/streaming_quality_validator.py \
    --lag --freshness --throughput --dlq \
    --output streaming-health-report.html

Complete Documentation: See tools.md for all monitoring dimensions and integration patterns.

kafka_config_generator.py

Production-grade Kafka configuration generator with performance and security profiles.

Capabilities:

  • Topic configuration (partitions, replication, retention, compaction)
  • Producer profiles (high-throughput, exactly-once, low-latency, ordered)
  • Consumer profiles (exactly-once, high-throughput, batch)
  • Kafka Streams configuration with state store tuning
  • Security configuration (SASL-PLAIN, SASL-SCRAM, mTLS)
  • Kafka Connect source/sink configurations
  • Multiple output formats (properties, YAML, JSON)

Usage:

bash
# Generate topic configuration
python scripts/kafka_config_generator.py \
    --topic user-events --partitions 12 --replication 3 --retention-hours 168

# Generate exactly-once producer
python scripts/kafka_config_generator.py \
    --producer --profile exactly-once --transactional-id producer-001

# Generate Kafka Streams config
python scripts/kafka_config_generator.py \
    --streams --application-id events-processor --exactly-once

Complete Documentation: See tools.md for all profiles, security options, and Connect configurations.

Show full SKILL.md (702 more words)Show less

Reference Documentation

Frameworks (frameworks.md)

Comprehensive data engineering frameworks and patterns:

  • Architecture Patterns: Lambda, Kappa, Medallion, Microservices data architecture
  • Data Modeling: Dimensional (Kimball), Data Vault 2.0, One Big Table
  • ETL/ELT Patterns: Full load, incremental load, CDC, SCD, idempotent pipelines
  • Data Quality: Complete framework covering all quality dimensions
  • DataOps: CI/CD for data pipelines, testing strategies, monitoring
  • Orchestration: Airflow DAG patterns, backfill strategies
  • Real-Time Streaming: Stateful processing, stream joins, windowing strategies, exactly-once semantics, event time processing, watermarks, backpressure, Apache Flink patterns, AWS Kinesis patterns, CDC for streaming
  • Governance: Data catalog, lineage tracking, access control
Templates (templates.md)

Production-ready code templates and examples:

  • Airflow DAGs: Complete ETL DAG, incremental load, dynamic task generation
  • Spark Jobs: Batch processing, streaming, optimized configurations
  • dbt Models: Staging, intermediate, fact tables, dimensions with SCD Type 2
  • SQL Patterns: Incremental merge (upsert), deduplication, date spine, window functions
  • Python Pipelines: Data quality validation class, retry decorators, error handling
  • Real-Time Streaming: Apache Flink DataStream jobs (Java), Kafka Streams applications, PyFlink jobs, AWS Kinesis consumers, Docker Compose for streaming stack
  • Kafka Configs: Producer/consumer properties templates, topic configurations, security configurations
  • Docker: Dockerfiles for data pipelines, Docker Compose for local development including streaming stack (Kafka, Flink, Schema Registry)
  • Configuration: dbt project config, Spark configuration, Airflow variables, streaming pipeline YAML
  • Testing: pytest fixtures, integration tests, data quality tests
Tools (tools.md)

Python automation tool documentation:

  • pipeline_orchestrator.py: Complete usage guide, configuration format, DAG templates
  • data_quality_validator.py: Validation rules, dimension checks, Great Expectations integration
  • etl_performance_optimizer.py: Performance analysis, query optimization, Spark tuning
  • stream_processor.py: Streaming pipeline configuration, validation, job scaffolding generation
  • streaming_quality_validator.py: Consumer lag, data freshness, schema drift, throughput monitoring
  • kafka_config_generator.py: Topic, producer, consumer, Kafka Streams, and Connect configurations
  • Integration Patterns: Airflow, dbt, CI/CD, monitoring systems, Prometheus
  • Best Practices: Configuration management, error handling, performance, monitoring, streaming quality

Tech Stack

Core Technologies:

  • Languages: Python 3.8+, SQL, Scala (Spark), Java (Flink)
  • Orchestration: Apache Airflow, Prefect, Dagster
  • Batch Processing: Apache Spark, dbt, Pandas
  • Stream Processing: Apache Kafka, Apache Flink, Kafka Streams, Spark Structured Streaming, AWS Kinesis
  • Storage: PostgreSQL, BigQuery, Snowflake, Redshift, S3, GCS
  • Schema Management: Confluent Schema Registry, AWS Glue Schema Registry
  • Containerization: Docker, Kubernetes
  • Monitoring: Datadog, Prometheus, Grafana, Kafka UI

Data Platforms:

  • Cloud Data Warehouses: Snowflake, BigQuery, Redshift
  • Data Lakes: Delta Lake, Apache Iceberg, Apache Hudi
  • Streaming Platforms: Apache Kafka, AWS Kinesis, Google Pub/Sub, Azure Event Hubs
  • Stream Processing Engines: Apache Flink, Kafka Streams, Spark Structured Streaming
  • Workflow: Airflow, Prefect, Dagster

Integration Points

This skill integrates with:

  • Orchestration: Airflow, Prefect, Dagster for workflow management
  • Transformation: dbt for SQL transformations and testing
  • Quality: Great Expectations for data validation
  • Monitoring: Datadog, Prometheus for pipeline monitoring
  • BI Tools: Looker, Tableau, Power BI for analytics
  • ML Platforms: MLflow, Kubeflow for ML pipeline integration
  • Version Control: Git for pipeline code and configuration

See tools.md for detailed integration patterns and examples.

Best Practices

Pipeline Design:

  1. Idempotent operations for safe reruns
  2. Incremental processing where possible
  3. Clear data lineage and documentation
  4. Comprehensive error handling
  5. Automated recovery mechanisms

Data Quality:

  1. Define quality rules early
  2. Validate at every pipeline stage
  3. Automate quality monitoring
  4. Track quality trends over time
  5. Block bad data from downstream

Performance:

  1. Partition large tables by date/region
  2. Use columnar formats (Parquet, ORC)
  3. Leverage predicate pushdown
  4. Optimize for your query patterns
  5. Monitor and tune regularly

Operations:

  1. Version control everything
  2. Automate testing and deployment
  3. Implement comprehensive monitoring
  4. Document runbooks for incidents
  5. Regular performance reviews

Performance Targets

Batch Pipeline Execution:

  • P50 latency: < 5 minutes (hourly pipelines)
  • P95 latency: < 15 minutes
  • Success rate: > 99%
  • Data freshness: < 1 hour behind source

Streaming Pipeline Execution:

  • Throughput: 10K+ events/second sustained
  • End-to-end latency: P99 < 1 second
  • Consumer lag: < 10K records behind
  • Exactly-once delivery: Zero duplicates or losses

Data Quality (Batch):

  • Quality score: > 95%
  • Completeness: > 99%
  • Timeliness: < 2 hours data lag
  • Zero critical failures

Streaming Quality:

  • Data freshness: P95 < 5 minutes from event generation
  • Late data rate: < 5% outside watermark window
  • Dead letter queue rate: < 1%
  • Schema compatibility: 100% backward/forward compatible changes

Cost Efficiency:

  • Cost per GB processed: < $0.10
  • Cloud cost trend: Stable or decreasing
  • Resource utilization: > 70%

Resources


Version: 2.0.0 Last Updated: December 16, 2025 Documentation Structure: Progressive disclosure with comprehensive references Streaming Enhancement: Task #8 - Real-time streaming capabilities added

© benchflow-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (scripts, references, assets) in tasks/flink-query/environment/skills/senior-data-engineer of benchflow-ai/skillsbench.

  • SKILL.md
  • HOW_TO_USE.md
  • assets/.gitkeep
  • assets/data-quality-checklist.md
  • assets/pipeline-design-template.md
  • references/frameworks.md
  • references/templates.md
  • references/tools.md
  • scripts/data_quality_validator.py
  • scripts/etl_performance_optimizer.py
  • scripts/kafka_config_generator.py
  • scripts/pipeline_orchestrator.py
  • scripts/stream_processor.py
  • scripts/streaming_quality_validator.py

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Senior Data Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Senior Data Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Senior Data Engineer this skillbenchflow-ai/skillsbench1.8k—~5.9kAutomated safety check: PassMIT
Senior Data Engineeralirezarezvani/claude-skills28k3 repos~1.4kAutomated safety check: PassMIT
Senior Data Engineerdavila7/claude-code-templates32k1 repos~1.4kAutomated safety check: PassMIT
Transforming Dataancoleman/ai-design-components526—~3kAutomated safety check: PassMIT
Databricks JobsKilo-Org/kilo-marketplace1891 repos~3.1kAutomated safety check: PassCustom licence
Engineering Data Pipelinestelagod/code-abyss244—~236Automated safety check: PassMIT

Similar skills

  • Senior Data Engineer

    alirezarezvani/claude-skills

    Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

    28k GitHub starsUsed in 3 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    davila7/claude-code-templates

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

    32k GitHub starsUsed in 1 repo~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Transforming Data

    ancoleman/ai-design-components

    Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).

    526 GitHub stars~3k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed
  • Databricks Jobs

    Kilo-Org/kilo-marketplace

    Develop and deploy Lakeflow Jobs on Databricks via DABs, Python SDK, or the CLI.

    189 GitHub starsUsed in 1 repo~3.1k tokens
    Data & AnalyticsAuto-check passed
  • Engineering Data Pipelines

    telagod/code-abyss

    Data engineering knowledge reference covering Airflow, Dagster, Kafka Streams, Flink, dbt, and data quality patterns.

    244 GitHub stars~236 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    borghei/Claude-Skills

    Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka.

    874 GitHub stars~1.4k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from benchflow-ai/skillsbench

All 180 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed
  • Energy Calculator

    benchflow-ai/skillsbench

    Calculate per-second RMS energy from audio files. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~437 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Senior Data Engineer

What does Senior Data Engineer do?

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure. Senior Data Engineer is an agent skill from benchflow-ai/skillsbench. World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

When should I use Senior Data Engineer?

Senior Data Engineer fits situations like: designing data architectures; streaming data pipelines; optimizing data workflows; implementing data governance.

How do I install Senior Data Engineer in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill senior-data-engineer -a claude-code`. Or copy the skill folder (tasks/flink-query/environment/skills/senior-data-engineer in benchflow-ai/skillsbench) into .claude/skills/senior-data-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Senior Data Engineer in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill senior-data-engineer -a codex`. Or copy the skill folder (tasks/flink-query/environment/skills/senior-data-engineer in benchflow-ai/skillsbench) into .agents/skills/senior-data-engineer in your project. Codex loads it when a task matches its description.

Can I use Senior Data Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill senior-data-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/senior-data-engineer, .gemini/skills/senior-data-engineer, .github/skills/senior-data-engineer and .opencode/skills/senior-data-engineer in your project.

What does Senior Data Engineer need to run?

Going by SKILL.md and its folder, Senior Data Engineer needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3; Docker. Compatibility (from SKILL.md): Python 3.8+; platforms: macos, linux, windows.

Does Senior Data Engineer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Senior Data Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Senior Data Engineer use?

Senior Data Engineer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Senior Data Engineer use?

About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 37k tokens, read only when the agent opens those files.

What are the alternatives to Senior Data Engineer?

Skills that share tags, products or a category with Senior Data Engineer: Senior Data Engineer (alirezarezvani/claude-skills, 28k stars), Senior Data Engineer (davila7/claude-code-templates, 32k stars), Transforming Data (ancoleman/ai-design-components, 526 stars) and Databricks Jobs (Kilo-Org/kilo-marketplace, 189 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Senior Data Engineer?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,832 GitHub stars. The repository holds 180 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.