Agent skill

Tensorflow Data Pipelines

by TheBushidoCollective in TheBushidoCollective/han

Create efficient data pipelines with tf.data. An agent skill from TheBushidoCollective/han.

Custom licenceAuto-check: notesData & Analytics

Install Tensorflow Data Pipelines

skills CLI
$ npx skills add TheBushidoCollective/han --skill tensorflow-data-pipelines -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TheBushidoCollective/han tensorflow-data-pipelines --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TheBushidoCollective/han.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/specialized/tensorflow/skills/tensorflow-data-pipelines .claude/skills/tensorflow-data-pipelines && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tensorflow-data-pipelines
GitHub stars
198
Used in
2 other repos
Token cost
~4.4k tokens
SKILL.md length
643 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
Custom licence

At a glance

Create efficient data pipelines with tf.data. An agent skill from TheBushidoCollective/han.

  • Works in 12 steps: Always use prefetch() - Add… → Use num_parallel_calls=AUTOTUNE - Let… → Cache after expensive operations - Place… → …
  • Tasks that involve Data pipelines and ETL
  • SKILL.md covers Dataset Creation, Data Transformation, Batching and Shuffling and Performance Optimization, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tensorflow Data Pipelines is an agent skill from TheBushidoCollective/han. Create efficient data pipelines with tf.data

Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL and Deep learning. It works with TensorFlow. The repository describes itself as: A curated marketplace of Claude Code plugins that embody the principles of ethical and professional software development.

When your agent uses it

  • Tasks that involve Data pipelines and ETL
  • Tasks that involve Deep learning

Example prompts

  • “/tensorflow-data-pipelines”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Always use prefetch() - Add .prefetch(tf.data.AUTOTUNE) at the end of pipeline to overlap data loading with training
  2. Use num_parallel_calls=AUTOTUNE - Let TensorFlow automatically tune parallelism for map operations
  3. Cache after expensive operations - Place .cache() after preprocessing but before augmentation and shuffling
  4. Shuffle before batching - Call .shuffle() before .batch() to ensure random batches
  5. Use appropriate buffer sizes - Shuffle buffer should be >= dataset size for perfect shuffling, or at least several thousand
  6. Normalize data in pipeline - Apply normalization in map() function for consistency across train/val/test
  7. Batch after transformations - Apply .batch() after all element-wise transformations for efficiency
  8. Use drop_remainder for training - Set drop_remainder=True in batch() to ensure consistent batch sizes
  9. Leverage AUTOTUNE - Use tf.data.AUTOTUNE for automatic performance tuning instead of manual values
  10. Apply augmentation after caching - Cache deterministic preprocessing, apply random augmentation after
  11. Use interleave for file reading - Parallel file reading with interleave() for large multi-file datasets
  12. Repeat for infinite datasets - Use .repeat() for training datasets to avoid dataset exhaustion

What it can do on your machine

Read from SKILL.md and the folder at commit 19caa51. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • tensorflow.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tensorflow Data Pipelines loads about 4.4k tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 643 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 643 words (~4,438 tokens).

“Build efficient, scalable data pipelines using the tf.data API for optimal training performance. This skill covers dataset creation, transformations, batching, shuffling, prefetching, and advanced optimization techniques to maximize GPU/TPU utilization.”

— opening of SKILL.md by TheBushidoCollective, Custom licence
name
tensorflow-data-pipelines
allowed-tools
Bash, Read

Read the full SKILL.md on GitHub

Files

Just SKILL.md in plugins/specialized/tensorflow/skills/tensorflow-data-pipelines of TheBushidoCollective/han.

Open the folder on GitHubat commit 19caa51

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in TheBushidoCollective/han, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Tensorflow Data Pipelines next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tensorflow Data Pipelines compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tensorflow Data Pipelines this skillTheBushidoCollective/han1982 repos~4.4kAutomated safety check: NotesCustom licence
Technology Selectiondotnet/skills5.6k2 repos~2.1kAutomated safety check: PassMIT
Ray Data for ML PipelinesOrchestra-Research/AI-Research-SKILLs13k3 repos~1.8kAutomated safety check: PassMIT
ML Model Trainingsecondsky/claude-skills2271 repos~1.7kAutomated safety check: PassMIT
Umap Learnjaechang-hits/SciAgent-Skills370—~4.7kAutomated safety check: PassBSD-3-Clause
TensorBoard Training VisualizationOrchestra-Research/AI-Research-SKILLs13k3 repos~3.8kAutomated safety check: PassMIT

Similar skills

  • Official

    Guides technology selection and implementation of AI and ML features in .NET 8+ applications using ML.NET, Microsoft.Extensions.AI (MEAI), Microsoft Agent Framework (MAF), GitHub Copilot SDK, ONNX…

    5.6k GitHub starsUsed in 2 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Ray Data for ML Pipelines

    Orchestra-Research/AI-Research-SKILLs

    Uses Ray Data to read, transform and write large datasets across a cluster for ML training and batch inference, with streaming execution and optional GPU steps.

    13k GitHub starsUsed in 3 repos~1.8k tokens
    Data & AnalyticsAuto-check passed
  • ML Model Training

    secondsky/claude-skills

    Train ML models with scikit-learn, PyTorch, TensorFlow. An agent skill from secondsky/claude-skills.

    227 GitHub starsUsed in 1 repo~1.7k tokens
    Data & AnalyticsAuto-check passed
  • Umap Learn

    jaechang-hits/SciAgent-Skills

    UMAP dimensionality reduction for visualization, clustering prep, and feature engineering.

    370 GitHub stars~4.7k tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed
  • TensorBoard Training Visualization

    Orchestra-Research/AI-Research-SKILLs

    Covers logging and viewing training metrics, histograms, model graphs, embeddings and profiles with TensorBoard in PyTorch and TensorFlow projects.

    13k GitHub starsUsed in 3 repos~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Formatting

    brendanhasz/probflow

    Ensure consistent code formatting using the uv package manager and pre-commit.

    175 GitHub stars~381 tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed

More from TheBushidoCollective/han

All 12 skills in this repo
  • Android Jetpack Compose

    TheBushidoCollective/han

    A skill your agent uses when building Android UIs with Jetpack Compose, managing state with remember/mutableStateOf, or implementing declarative UI patterns.

    198 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check: notes
  • Graphql Inspector Validate

    TheBushidoCollective/han

    A skill your agent uses when validating GraphQL operations/documents against a schema, checking query depth, complexity, or fragment usage.

    198 GitHub starsUsed in 1 repo~2k tokens
    Auto-check: notes
  • Graphql Schema Design

    TheBushidoCollective/han

    A skill your agent uses when designing GraphQL schemas with type system, SDL patterns, field design, pagination, directives, and versioning strategies for maintainable and scalable APIs.

    198 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed
  • Maven Build Lifecycle

    TheBushidoCollective/han

    A skill your agent uses when working with Maven build phases, goals, profiles, or customizing the build process for Java projects.

    198 GitHub starsUsed in 1 repo~2.9k tokens
    Auto-check: notes
  • Maven Dependency Management

    TheBushidoCollective/han

    A skill your agent uses when managing Maven dependencies, resolving dependency conflicts, configuring BOMs, or optimizing dependency trees in Java projects.

    198 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check: notes
  • Maven Plugin Configuration

    TheBushidoCollective/han

    A skill your agent uses when configuring Maven plugins, setting up common plugins like compiler, surefire, jar, or creating custom plugin executions.

    198 GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check: notes

Works with

Questions about Tensorflow Data Pipelines

What does Tensorflow Data Pipelines do?

Create efficient data pipelines with tf.data. An agent skill from TheBushidoCollective/han. Tensorflow Data Pipelines is an agent skill from TheBushidoCollective/han.

When should I use Tensorflow Data Pipelines?

Tensorflow Data Pipelines fits situations like: tasks that involve Data pipelines and ETL; tasks that involve Deep learning.

How do I install Tensorflow Data Pipelines in Claude Code?

Run `npx skills add TheBushidoCollective/han --skill tensorflow-data-pipelines -a claude-code`. Or copy the skill folder (plugins/specialized/tensorflow/skills/tensorflow-data-pipelines in TheBushidoCollective/han) into .claude/skills/tensorflow-data-pipelines in your project. Claude Code loads it when a task matches its description.

How do I install Tensorflow Data Pipelines in Codex?

Run `npx skills add TheBushidoCollective/han --skill tensorflow-data-pipelines -a codex`. Or copy the skill folder (plugins/specialized/tensorflow/skills/tensorflow-data-pipelines in TheBushidoCollective/han) into .agents/skills/tensorflow-data-pipelines in your project. Codex loads it when a task matches its description.

Can I use Tensorflow Data Pipelines in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TheBushidoCollective/han --skill tensorflow-data-pipelines -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tensorflow-data-pipelines, .gemini/skills/tensorflow-data-pipelines, .github/skills/tensorflow-data-pipelines and .opencode/skills/tensorflow-data-pipelines in your project.

What does Tensorflow Data Pipelines need to run?

SKILL.md names no scripts, command-line tools or credentials: Tensorflow Data Pipelines is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read.

Does Tensorflow Data Pipelines access the network?

SKILL.md names 1 domain. As links in the text: tensorflow.org. This is read from the text; nothing was executed.

Is Tensorflow Data Pipelines safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tensorflow Data Pipelines use?

Tensorflow Data Pipelines has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Tensorflow Data Pipelines use?

About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tensorflow Data Pipelines?

Skills that share tags, products or a category with Tensorflow Data Pipelines: Technology Selection (dotnet/skills, 5.6k stars), Ray Data for ML Pipelines (Orchestra-Research/AI-Research-SKILLs, 13k stars), ML Model Training (secondsky/claude-skills, 227 stars) and Umap Learn (jaechang-hits/SciAgent-Skills, 370 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tensorflow Data Pipelines?

TheBushidoCollective (a GitHub organization) maintains it in TheBushidoCollective/han, which has 198 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 8, 2026.

Source: TheBushidoCollective/han on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.