Agent skill

Streaming Data

by ancoleman in ancoleman/ai-design-components

Build event streaming and real-time data pipelines with Kafka, Pulsar, Redpanda, Flink, and Spark.

MITAuto-check passedBackend & APIs

Install Streaming Data

skills CLI
$ npx skills add ancoleman/ai-design-components --skill streaming-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ancoleman/ai-design-components streaming-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/streaming-data .claude/skills/streaming-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
streaming-data
GitHub stars
526
Token cost
~2.9k tokens
SKILL.md length
946 words
Files
16 (incl. references)
Skills in repo
75
Repo updated
First seen
Licence
MIT

At a glance

Build event streaming and real-time data pipelines with Kafka, Pulsar, Redpanda, Flink, and Spark.

  • Works in 3 steps: Choose a Message Broker → Choose a Stream Processor (if needed) → Implement Producer/Consumer Patterns
  • Tasks that involve Event-driven systems
  • SKILL.md covers When to Use This Skill, Core Concepts, Quick Start Guide and Common Patterns, plus 4 more sections
  • Runs Python and TypeScript scripts from its folder; calls python and bash

What it does

Streaming Data is an agent skill from ancoleman/ai-design-components. Build event streaming and real-time data pipelines with Kafka, Pulsar, Redpanda, Flink, and Spark. Covers producer/consumer patterns, stream processing, event sourcing, and CDC across TypeScript, Python, Go, and Java. When building real-time systems, microservices communication, or data integration pipelines.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including reference files (for example `examples/python/basic_consumer.py`, `examples/typescript/basic-producer.ts` and `outputs.yaml`).

It sits in Backend & APIs, covering Event-driven systems and Data pipelines and ETL. It works with Apache Kafka, TypeScript, Java and Python. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.

When your agent uses it

  • Tasks that involve Event-driven systems
  • Tasks that involve Data pipelines and ETL

Example prompts

  • “/streaming-data”

Requirements

  • Python 3
  • Node.js

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Choose a Message Broker
  2. Choose a Stream Processor (if needed)
  3. Implement Producer/Consumer Patterns

What it can do on your machine

Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and TypeScript), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Streaming Data loads about 2.9k tokens when it runs, and up to ~33k if it reads all its reference files. Until then it costs about 81 tokens; SKILL.md has 946 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~33k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 946 words, ~2,879 tokens.

Download SKILL.mdSave it as .claude/skills/streaming-data/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
streaming-data
description
Build event streaming and real-time data pipelines with Kafka, Pulsar, Redpanda, Flink, and Spark. Covers producer/consumer patterns, stream processing, event sourcing, and CDC across TypeScript, Python, Go, and Java. When building real-time systems, microservices communication, or data integration pipelines.

Streaming Data Processing

Build production-ready event streaming systems and real-time data pipelines using modern message brokers and stream processors.

When to Use This Skill

Use this skill when:

  • Building event-driven architectures and microservices communication
  • Processing real-time analytics, monitoring, or alerting systems
  • Implementing data integration pipelines (CDC, ETL/ELT)
  • Creating log or metrics aggregation systems
  • Developing IoT platforms or high-frequency trading systems

Core Concepts

Message Brokers vs Stream Processors

Message Brokers (Kafka, Pulsar, Redpanda):

  • Store and distribute event streams
  • Provide durability, replay capability, partitioning
  • Handle producer/consumer coordination

Stream Processors (Flink, Spark, Kafka Streams):

  • Transform and aggregate streaming data
  • Provide windowing, joins, stateful operations
  • Execute complex event processing (CEP)
Delivery Guarantees

At-Most-Once:

  • Messages may be lost, no duplicates
  • Lowest overhead
  • Use for: Metrics, logs where loss is acceptable

At-Least-Once:

  • Messages never lost, may have duplicates
  • Moderate overhead, requires idempotent consumers
  • Use for: Most applications (default choice)

Exactly-Once:

  • Messages never lost or duplicated
  • Highest overhead, requires transactional processing
  • Use for: Financial transactions, critical state updates

Quick Start Guide

Step 1: Choose a Message Broker

See references/broker-selection.md for detailed comparison.

Quick decision:

  • Apache Kafka: Mature ecosystem, enterprise features, event sourcing
  • Redpanda: Low latency, Kafka-compatible, simpler operations (no ZooKeeper)
  • Apache Pulsar: Multi-tenancy, geo-replication, tiered storage
  • RabbitMQ: Traditional message queues, RPC patterns
Step 2: Choose a Stream Processor (if needed)

See references/processor-selection.md for detailed comparison.

Quick decision:

  • Apache Flink: Millisecond latency, real-time analytics, CEP
  • Apache Spark: Batch + stream hybrid, ML integration, analytics
  • Kafka Streams: Embedded in microservices, no separate cluster
  • ksqlDB: SQL interface for stream processing
Step 3: Implement Producer/Consumer Patterns

Choose language-specific guide:

  • TypeScript/Node.js: references/typescript-patterns.md (KafkaJS)
  • Python: references/python-patterns.md (confluent-kafka-python)
  • Go: references/go-patterns.md (kafka-go)
  • Java/Scala: references/java-patterns.md (Apache Kafka Java Client)

Common Patterns

Basic Producer Pattern

Send events to a topic with error handling:

1. Create producer with broker addresses
2. Configure delivery guarantees (acks, retries, idempotence)
3. Send messages with key (for partitioning) and value
4. Handle delivery callbacks or errors
5. Flush and close producer on shutdown
Basic Consumer Pattern

Process events from topics with offset management:

1. Create consumer with broker addresses and group ID
2. Subscribe to topics
3. Poll for messages
4. Process each message
5. Commit offsets (auto or manual)
6. Handle errors (retry, DLQ, skip)
7. Close consumer gracefully
Error Handling Strategy

For production systems, implement:

  • Dead Letter Queue (DLQ): Send failed messages to separate topic
  • Retry Logic: Configurable retry attempts with backoff
  • Graceful Shutdown: Finish processing, commit offsets, close connections
  • Monitoring: Track consumer lag, error rates, throughput

Decision Frameworks

Framework: Message Broker Selection
START: What are requirements?

1. Need Kafka API compatibility?
   YES → Kafka or Redpanda
   NO → Continue

2. Is multi-tenancy critical?
   YES → Apache Pulsar
   NO → Continue

3. Operational simplicity priority?
   YES → Redpanda (single binary, no ZooKeeper)
   NO → Continue

4. Mature ecosystem needed?
   YES → Apache Kafka
   NO → Redpanda (better performance)

5. Task queues (not event streams)?
   YES → RabbitMQ or message-queues skill
   NO → Kafka/Redpanda/Pulsar
Framework: Stream Processor Selection
START: What is latency requirement?

1. Millisecond-level latency needed?
   YES → Apache Flink
   NO → Continue

2. Batch + stream in same pipeline?
   YES → Apache Spark Streaming
   NO → Continue

3. Embedded in microservice?
   YES → Kafka Streams
   NO → Continue

4. SQL interface for analysts?
   YES → ksqlDB
   NO → Flink or Spark

5. Python primary language?
   YES → Spark (PySpark) or Faust
   NO → Flink (Java/Scala)
Framework: Language Selection

TypeScript/Node.js:

  • API gateways, web services, real-time dashboards
  • KafkaJS library (827 code snippets, high reputation)

Python:

  • Data science, ML pipelines, analytics
  • confluent-kafka-python (192 snippets, score 68.8)

Go:

  • High-performance microservices, infrastructure tools
  • kafka-go (42 snippets, idiomatic Go)

Java/Scala:

  • Enterprise applications, Kafka Streams, Flink, Spark
  • Apache Kafka Java Client (683 snippets, score 76.9)

Advanced Patterns

Event Sourcing

Store state changes as immutable events. See references/event-sourcing.md for:

  • Event store design patterns
  • Event schema evolution
  • Snapshot strategies
  • Temporal queries and audit trails
Change Data Capture (CDC)

Capture database changes as events. See references/cdc-patterns.md for:

  • Debezium integration (MySQL, PostgreSQL, MongoDB)
  • Real-time data synchronization
  • Microservices data integration patterns
Exactly-Once Processing

Implement transactional guarantees. See references/exactly-once.md for:

  • Idempotent producers
  • Transactional consumers
  • End-to-end exactly-once pipelines
Error Handling

Production-grade error management. See references/error-handling.md for:

  • Dead letter queue patterns
  • Retry strategies with exponential backoff
  • Backpressure handling
  • Circuit breakers for downstream failures

Reference Files

Decision Guides
  • references/broker-selection.md - Kafka vs Pulsar vs Redpanda comparison
  • references/processor-selection.md - Flink vs Spark vs Kafka Streams
  • references/delivery-guarantees.md - At-least-once, exactly-once patterns
Language-Specific Implementation
  • references/typescript-patterns.md - KafkaJS patterns (producer, consumer, error handling)
  • references/python-patterns.md - confluent-kafka-python patterns
  • references/go-patterns.md - kafka-go patterns
  • references/java-patterns.md - Apache Kafka Java client patterns
Advanced Topics
  • references/event-sourcing.md - Event sourcing architecture
  • references/cdc-patterns.md - Change Data Capture with Debezium
  • references/exactly-once.md - Transactional processing
  • references/error-handling.md - DLQ, retries, backpressure
  • references/performance-tuning.md - Throughput optimization, partitioning strategies

Validation Scripts

Run these scripts for token-free validation and generation:

Show full SKILL.md (382 more words)Show less
Validate Kafka Configuration
bash
python scripts/validate-kafka-config.py --config producer.yaml
python scripts/validate-kafka-config.py --config consumer.yaml

Checks: broker connectivity, configuration validity, serialization format

Generate Schema Registry Templates
bash
python scripts/generate-schema.py --type avro --entity User
python scripts/generate-schema.py --type protobuf --entity Event

Creates: Avro/Protobuf schema definitions for Schema Registry

Benchmark Throughput
bash
bash scripts/benchmark-throughput.sh --broker localhost:9092 --topic test

Tests: Producer/consumer throughput, latency percentiles

Code Examples

TypeScript Example (KafkaJS)

See examples/typescript/ for:

  • basic-producer.ts - Simple event producer with error handling
  • basic-consumer.ts - Consumer with manual offset commits
  • transactional-producer.ts - Exactly-once producer pattern
  • consumer-with-dlq.ts - Dead letter queue implementation
Python Example (confluent-kafka-python)

See examples/python/ for:

  • basic_producer.py - Producer with delivery callbacks
  • basic_consumer.py - Consumer with error handling
  • async_producer.py - AsyncIO producer (aiokafka)
  • schema_registry.py - Avro serialization with Schema Registry
Go Example (kafka-go)

See examples/go/ for:

  • basic_producer.go - Idiomatic Go producer
  • basic_consumer.go - Consumer with manual commits
  • high_perf_consumer.go - Concurrent processing pattern
  • batch_producer.go - Batch message sending
Java Example (Apache Kafka)

See examples/java/ for:

  • BasicProducer.java - Producer with idempotence
  • BasicConsumer.java - Consumer with error recovery
  • TransactionalProducer.java - Exactly-once transactions
  • StreamsAggregation.java - Kafka Streams aggregation

Technology Comparison

Message Broker Comparison
FeatureKafkaPulsarRedpandaRabbitMQ
ThroughputVery HighHighVery HighMedium
LatencyMediumMediumLowLow
Event ReplayYesYesYesNo
Multi-TenancyManualNativeManualManual
Operational ComplexityMediumHighLowLow
Best ForEnterprise, big dataSaaS, IoTPerformance-criticalTask queues
Stream Processor Comparison
FeatureFlinkSparkKafka StreamsksqlDB
Processing ModelTrue streamingMicro-batchLibrarySQL engine
LatencyMillisecondSecondMillisecondSecond
DeploymentClusterClusterEmbeddedServer
Best ForReal-time analyticsBatch + streamMicroservicesAnalysts
Client Library Recommendations
LanguageLibraryTrust ScoreSnippetsUse Case
TypeScriptKafkaJSHigh827Web services, APIs
Pythonconfluent-kafka-pythonHigh (68.8)192Data pipelines, ML
Gokafka-goHigh42High-perf services
JavaKafka Java ClientHigh (76.9)683Enterprise, Flink/Spark

For authentication and security patterns, see the auth-security skill. For infrastructure deployment (Kubernetes operators, Terraform), see the infrastructure-as-code skill. For monitoring metrics and tracing, see the observability skill. For API design patterns, see the api-design-principles skill. For data architecture and warehousing, see the data-architecture skill.

Troubleshooting

Consumer Lag Issues
  • Check partition count vs consumer count (match for parallelism)
  • Increase consumer instances or reduce processing time
  • Monitor with Kafka consumer lag metrics
Message Loss
  • Verify producer acks=all configuration
  • Check broker replication factor (>1)
  • Ensure consumers commit offsets after processing
Duplicate Messages
  • Implement idempotent consumers (track message IDs)
  • Use exactly-once semantics (transactions)
  • Design for at-least-once delivery
Performance Bottlenecks
  • Increase partition count for parallelism
  • Tune batch size and linger time
  • Enable compression (GZIP, LZ4, Snappy)
  • See references/performance-tuning.md for details

© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (references) in skills/streaming-data of ancoleman/ai-design-components.

  • SKILL.md
  • examples/python/basic_consumer.py
  • examples/typescript/basic-producer.ts
  • outputs.yaml
  • references/broker-selection.md
  • references/cdc-patterns.md
  • references/delivery-guarantees.md
  • references/error-handling.md
  • references/event-sourcing.md
  • references/exactly-once.md
  • references/go-patterns.md
  • references/java-patterns.md
  • references/performance-tuning.md
  • references/processor-selection.md
  • references/python-patterns.md
  • references/typescript-patterns.md

Open the folder on GitHubat commit 76551b7

Compare with similar skills

Streaming Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Streaming Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Streaming Data this skillancoleman/ai-design-components526—~2.9kAutomated safety check: PassMIT
Databricks Zerobus Ingestdatabricks/databricks-agent-skills345—~3.1kAutomated safety check: PassCustom licence
Io ConnectorsKilo-Org/kilo-marketplace190—~1.3kAutomated safety check: PassApache-2.0
AWS Serverless Edazxkane/aws-skills3674 repos~3.2kAutomated safety check: PassMIT
Opensource Guide Coachcalf-ai/calfkit-sdk1491 repos~2.1kAutomated safety check: PassApache-2.0
Temporal Developertemporalio/skill-temporal-developer230—~2.5kAutomated safety check: PassMIT

Similar skills

  • Databricks Zerobus Ingest

    databricks/databricks-agent-skills

    Official

    Build Zerobus Ingest clients for near real-time data ingestion into Databricks Delta tables via gRPC.

    345 GitHub stars~3.1k tokensUpdated today
    Backend & APIsAuto-check passed
  • Io Connectors

    Kilo-Org/kilo-marketplace

    Guides development and usage of I/O connectors in Apache Beam.

    190 GitHub stars~1.3k tokensUpdated 9 days ago
    DatabasesAuto-check passed
  • AWS Serverless Eda

    zxkane/aws-skills

    AWS serverless and event-driven architecture expert based on Well-Architected Framework.

    367 GitHub starsUsed in 4 repos~3.2k tokens
    Backend & APIsAuto-check passed
  • Opensource Guide Coach

    calf-ai/calfkit-sdk

    A skill your agent uses when a user wants guidance on starting, contributing to, growing, governing, funding, securing, or sustaining an open source project, or asks about contributor onboarding…

    149 GitHub starsUsed in 1 repo~2.1k tokens
    Backend & APIsAuto-check passed
  • Temporal Developer

    temporalio/skill-temporal-developer

    Official

    Develop, debug, and manage Temporal applications across Python, TypeScript, Go, Java, .NET, Ruby, and Rust.

    230 GitHub stars~2.5k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Create Environment

    godatadriven/whirl

    Create a new Whirl environment in the envs/ directory. An agent skill from godatadriven/whirl.

    205 GitHub stars~1.9k tokensUpdated 6 days ago
    Backend & APIsAuto-check passed

More from ancoleman/ai-design-components

All 75 skills in this repo
  • Building AI Chat

    ancoleman/ai-design-components

    Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.

    526 GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed
  • Building Forms

    ancoleman/ai-design-components

    Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.

    526 GitHub stars~3.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Building Tables

    ancoleman/ai-design-components

    Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.

    526 GitHub stars~1.8k tokensUpdated 10 mo ago
    Auto-check passed
  • Creating Dashboards

    ancoleman/ai-design-components

    Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.

    526 GitHub stars~3.5k tokensUpdated 10 mo ago
    Auto-check passed
  • Designing Layouts

    ancoleman/ai-design-components

    Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.

    526 GitHub stars~1.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Displaying Timelines

    ancoleman/ai-design-components

    Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.

    526 GitHub stars~2.7k tokensUpdated 10 mo ago
    Auto-check passed

Categories

Questions about Streaming Data

What does Streaming Data do?

Build event streaming and real-time data pipelines with Kafka, Pulsar, Redpanda, Flink, and Spark. Streaming Data is an agent skill from ancoleman/ai-design-components. Build event streaming and real-time data pipelines with Kafka, Pulsar, Redpanda, Flink, and Spark.

When should I use Streaming Data?

Streaming Data fits situations like: tasks that involve Event-driven systems; tasks that involve Data pipelines and ETL.

How do I install Streaming Data in Claude Code?

Run `npx skills add ancoleman/ai-design-components --skill streaming-data -a claude-code`. Or copy the skill folder (skills/streaming-data in ancoleman/ai-design-components) into .claude/skills/streaming-data in your project. Claude Code loads it when a task matches its description.

How do I install Streaming Data in Codex?

Run `npx skills add ancoleman/ai-design-components --skill streaming-data -a codex`. Or copy the skill folder (skills/streaming-data in ancoleman/ai-design-components) into .agents/skills/streaming-data in your project. Codex loads it when a task matches its description.

Can I use Streaming Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill streaming-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/streaming-data, .gemini/skills/streaming-data, .github/skills/streaming-data and .opencode/skills/streaming-data in your project.

What does Streaming Data need to run?

Going by SKILL.md and its folder, Streaming Data needs Python and TypeScript for the scripts in its folder and the command-line tools its instructions call (python and bash). Our summary lists: Python 3; Node.js.

Does Streaming Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Streaming Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Streaming Data use?

Streaming Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Streaming Data use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 30k tokens, read only when the agent opens those files.

What are the alternatives to Streaming Data?

Skills that share tags, products or a category with Streaming Data: Databricks Zerobus Ingest (databricks/databricks-agent-skills, 345 stars), Io Connectors (Kilo-Org/kilo-marketplace, 190 stars), AWS Serverless Eda (zxkane/aws-skills, 367 stars) and Opensource Guide Coach (calf-ai/calfkit-sdk, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Streaming Data?

ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.

Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.