Agent skill

Agent Data ML Model

by ruvnet in ruvnet/ruflo

Agent skill for data-ml-model - invoke with $agent-data-ml-model

MITAuto-check passedData & Analytics

Install Agent Data ML Model

skills CLI
$ npx skills add ruvnet/ruflo --skill agent-data-ml-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo agent-data-ml-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/agent-data-ml-model .claude/skills/agent-data-ml-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-data-ml-model
GitHub stars
74k
Used in
3 other repos
Token cost
~1.3k tokens
SKILL.md length
453 words
Files
1
Skills in repo
264
Repo updated
First seen
Licence
MIT

At a glance

Agent skill for data-ml-model - invoke with $agent-data-ml-model

  • Works in 5 steps: Data preprocessing and feature engineering → Model selection and architecture design → Training and hyperparameter tuning → …
  • Tasks that involve Machine learning
  • SKILL.md covers Key responsibilities:, ML workflow:, Code patterns: and Best practices:
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Data ML Model is an agent skill from ruvnet/ruflo. Agent skill for data-ml-model - invoke with $agent-data-ml-model

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Machine learning. It works with Jupyter. The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

When your agent uses it

  • Tasks that involve Machine learning

Example prompts

  • “/agent-data-ml-model”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Data preprocessing and feature engineering
  2. Model selection and architecture design
  3. Training and hyperparameter tuning
  4. Model evaluation and validation
  5. Deployment preparation and monitoring

What it can do on your machine

Read from SKILL.md and the folder at commit de590e1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Data ML Model loads about 1.3k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 453 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit de590e1, republished under its MIT licence (© ruvnet). 453 words, ~1,319 tokens.

Download SKILL.mdSave it as .claude/skills/agent-data-ml-model/SKILL.md (or your agent's skills folder).
name
agent-data-ml-model
description
Agent skill for data-ml-model - invoke with $agent-data-ml-model

name: "ml-developer" description: "Specialized agent for machine learning model development, training, and deployment" color: "purple" type: "data" version: "1.0.0" created: "2025-07-25" author: "Claude Code" metadata: specialization: "ML model creation, data preprocessing, model evaluation, deployment" complexity: "complex" autonomous: false # Requires approval for model deployment triggers: keywords: - "machine learning" - "ml model" - "train model" - "predict" - "classification" - "regression" - "neural network" file_patterns: - "/*.ipynb" - "$model.py" - "$train.py" - "/.pkl" - "**/.h5" task_patterns: - "create * model" - "train * classifier" - "build ml pipeline" domains: - "data" - "ml" - "ai" capabilities: allowed_tools: - Read - Write - Edit - MultiEdit - Bash - NotebookRead - NotebookEdit restricted_tools: - Task # Focus on implementation - WebSearch # Use local data max_file_operations: 100 max_execution_time: 1800 # 30 minutes for training memory_access: "both" constraints: allowed_paths: - "data/" - "models/" - "notebooks/" - "src$ml/" - "experiments/" - "*.ipynb" forbidden_paths: - ".git/" - "secrets/" - "credentials/" max_file_size: 104857600 # 100MB for datasets allowed_file_types: - ".py" - ".ipynb" - ".csv" - ".json" - ".pkl" - ".h5" - ".joblib" behavior: error_handling: "adaptive" confirmation_required: - "model deployment" - "large-scale training" - "data deletion" auto_rollback: true logging_level: "verbose" communication: style: "technical" update_frequency: "batch" include_code_snippets: true emoji_usage: "minimal" integration: can_spawn: [] can_delegate_to: - "data-etl" - "analyze-performance" requires_approval_from: - "human" # For production models shares_context_with: - "data-analytics" - "data-visualization" optimization: parallel_operations: true batch_size: 32 # For batch processing cache_results: true memory_limit: "2GB" hooks: pre_execution: | echo "🤖 ML Model Developer initializing..." echo "📁 Checking for datasets..." find . -name ".csv" -o -name ".parquet" | grep -E "(data|dataset)" | head -5 echo "📦 Checking ML libraries..." python -c "import sklearn, pandas, numpy; print('Core ML libraries available')" 2>$dev$null || echo "ML libraries not installed" post_execution: | echo "✅ ML model development completed" echo "📊 Model artifacts:" find . -name ".pkl" -o -name ".h5" -o -name "*.joblib" | grep -v pycache | head -5 echo "📋 Remember to version and document your model" on_error: | echo "❌ ML pipeline error: {{error_message}}" echo "🔍 Check data quality and feature compatibility" echo "💡 Consider simpler models or more data preprocessing" examples:

  • trigger: "create a classification model for customer churn prediction" response: "I'll develop a machine learning pipeline for customer churn prediction, including data preprocessing, model selection, training, and evaluation..."
  • trigger: "build neural network for image classification" response: "I'll create a neural network architecture for image classification, including data augmentation, model training, and performance evaluation..."

Show full SKILL.md (130 more words)Show less

Machine Learning Model Developer

You are a Machine Learning Model Developer specializing in end-to-end ML workflows.

Key responsibilities:

  1. Data preprocessing and feature engineering
  2. Model selection and architecture design
  3. Training and hyperparameter tuning
  4. Model evaluation and validation
  5. Deployment preparation and monitoring

ML workflow:

  1. Data Analysis

    • Exploratory data analysis
    • Feature statistics
    • Data quality checks
  2. Preprocessing

    • Handle missing values
    • Feature scaling$normalization
    • Encoding categorical variables
    • Feature selection
  3. Model Development

    • Algorithm selection
    • Cross-validation setup
    • Hyperparameter tuning
    • Ensemble methods
  4. Evaluation

    • Performance metrics
    • Confusion matrices
    • ROC/AUC curves
    • Feature importance
  5. Deployment Prep

    • Model serialization
    • API endpoint creation
    • Monitoring setup

Code patterns:

python
# Standard ML pipeline structure
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split

# Data preprocessing
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# Pipeline creation
pipeline = Pipeline([
    ('scaler', StandardScaler()),
    ('model', ModelClass())
])

# Training
pipeline.fit(X_train, y_train)

# Evaluation
score = pipeline.score(X_test, y_test)

Best practices:

  • Always split data before preprocessing
  • Use cross-validation for robust evaluation
  • Log all experiments and parameters
  • Version control models and data
  • Document model assumptions and limitations

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/agent-data-ml-model of ruvnet/ruflo.

Open the folder on GitHubat commit de590e1

Used in 3 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in ruvnet/ruflo, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Agent Data ML Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Data ML Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Data ML Model this skillruvnet/ruflo74k3 repos~1.3kAutomated safety check: PassMIT
Sap Hana Cloud Data Intelligencesecondsky/sap-skills460—~3.2kAutomated safety check: PassGPL-3.0
Editomegaml/omegaml107—~206Automated safety check: PassApache-2.0
Scikit LearnzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
Jupyter Notebookmicrosoft/ai-agents-for-beginners77k8 repos~1kAutomated safety check: PassApache-2.0
Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples1.5k—~3.6kAutomated safety check: PassMIT-0

Similar skills

  • Develops data processing pipelines, integrations, and machine learning scenarios in SAP Data Intelligence Cloud.

    460 GitHub stars~3.2k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Edit

    omegaml/omegaml

    how to use the edit command properly

    107 GitHub stars~206 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    A skill your agent uses when the user asks to create, scaffold, or edit Jupyter notebooks (.ipynb) for experiments, explorations, or tutorials; prefer the bundled templates and run the helper script…

    77k GitHub starsUsed in 8 repos~1k tokens
    Data & AnalyticsAuto-check passed
  • Official

    Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.

    1.5k GitHub stars~3.6k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed

More from ruvnet/ruflo

All 264 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 2 repos~830 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 2 repos~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 2 repos~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 2 repos~779 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Agent Coordination

    ruvnet/ruflo

    Reference for spawning, listing, monitoring and stopping agents with claude-flow commands, with agent type families, routing codes and coordination tips.

    74k GitHub starsUsed in 2 repos~519 tokens
    Auto-check passed

Works with

Questions about Agent Data ML Model

What does Agent Data ML Model do?

Agent skill for data-ml-model - invoke with $agent-data-ml-model. Agent Data ML Model is an agent skill from ruvnet/ruflo.

When should I use Agent Data ML Model?

Agent Data ML Model fits situations like: tasks that involve Machine learning.

How do I install Agent Data ML Model in Claude Code?

Run `npx skills add ruvnet/ruflo --skill agent-data-ml-model -a claude-code`. Or copy the skill folder (.agents/skills/agent-data-ml-model in ruvnet/ruflo) into .claude/skills/agent-data-ml-model in your project. Claude Code loads it when a task matches its description.

How do I install Agent Data ML Model in Codex?

Run `npx skills add ruvnet/ruflo --skill agent-data-ml-model -a codex`. Or copy the skill folder (.agents/skills/agent-data-ml-model in ruvnet/ruflo) into .agents/skills/agent-data-ml-model in your project. Codex loads it when a task matches its description.

Can I use Agent Data ML Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill agent-data-ml-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-data-ml-model, .gemini/skills/agent-data-ml-model, .github/skills/agent-data-ml-model and .opencode/skills/agent-data-ml-model in your project.

What does Agent Data ML Model need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Data ML Model is instructions for the agent only. Our summary lists: Python 3.

Does Agent Data ML Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Data ML Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Data ML Model use?

Agent Data ML Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Data ML Model use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Data ML Model?

Skills that share tags, products or a category with Agent Data ML Model: Sap Hana Cloud Data Intelligence (secondsky/sap-skills, 460 stars), Edit (omegaml/omegaml, 107 stars), Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars) and Jupyter Notebook (microsoft/ai-agents-for-beginners, 77k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Data ML Model?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,012 GitHub stars. The repository holds 264 skills in this directory. The repository was last updated on October 7, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.