Agent skill

Agentic Kaggle Workflow

by FrankS-IntelLab in FrankS-IntelLab/agentic-kaggle-skill

Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

MITAuto-check passedData & Analytics

Install Agentic Kaggle Workflow

skills CLI
$ npx skills add FrankS-IntelLab/agentic-kaggle-skill --skill agentic-kaggle-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FrankS-IntelLab/agentic-kaggle-skill agentic-kaggle-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentic-kaggle-skill
GitHub stars
188
Token cost
~4k tokens
SKILL.md length
1,840 words
Files
35 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

  • Works in 12 steps: Read the competition page, classify the… → If live competition intelligence tools… → Identify the task type: binary,… → …
  • Running a Kaggle competition through to a scored submission
  • SKILL.md covers Operating Loop, Resource Map, Default Kaggle Procedure and Definition Of Done, plus 4 more sections
  • Calls python3

What it does

The skill treats each competition as a validation problem first and a modeling problem second, defaulting to Kaggle-native notebooks, datasets and score receipts. It starts by classifying the submission mode (file or code competition) and reading the rules, data terms, sharing policy, metric and leakage warnings, then identifies the task type and designs folds before any feature work, preferably as a fold column saved in the training data.

From there it builds the simplest metric-correct baseline with out-of-fold predictions and a valid submission, then plans stronger architectures: diverse model families, producer notebooks that export private artifact datasets for downstream consumer notebooks, pseudo-labeling or distillation, calibration and an ensemble or stacker. Heavy training can be offloaded to Kaggle GPUs. Reference files cover code-competition debugging, competition intel from public notebooks within sharing policy, cross-validation and metrics, image and text workflows and score retrieval, with two case studies.

When your agent uses it

  • Running a Kaggle competition through to a scored submission
  • Designing cross-validation folds and metrics before modeling
  • Debugging a code competition that fails on hidden test data
  • Splitting work across producer and consumer notebooks with private artifact datasets
  • Offloading heavy training to Kaggle notebooks with GPUs

Example prompts

  • “Take the Titanic competition from data download to a scored Kaggle submission.”
  • “My code competition notebook scores locally but fails on the hidden test; help me debug it.”
  • “Set up a stacking ensemble from three producer notebooks and export the artifacts as a private dataset.”

Requirements

  • A Kaggle account with access to the competition
  • Kaggle CLI or kagglehub access for datasets and submissions

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Read the competition page, classify the submission mode as classic file submission or code/notebook scoring, then inspect the rules…
  2. If live competition intelligence tools are available, inspect top public open notebook solutions and relevant discussion activity before…
  3. Identify the task type: binary, multiclass, multilabel, regression, ranking, image, segmentation, text, time series, grouped entities, or…
  4. Design folds before feature engineering or modeling. Prefer a fold column saved into the training data so every experiment uses the same…
  5. Build the simplest metric-correct baseline and produce out-of-fold (OOF) predictions plus a valid submission.
  6. Proactively plan a stronger architecture once the baseline is trustworthy: diverse model families, feature/embedding producers…
  7. Iterate with validation gates: test one meaningful change at a time when possible, but launch several independent producer notebooks in…
  8. Offload heavy training, inference, embedding generation, image/text experiments, or memory-risky jobs to Kaggle notebooks/scripts when…
  9. For sophisticated architectures, split work into a small Kaggle pipeline: run several independent producer notebooks/scripts first, save…
  10. Submit the final submission-producing artifact to Kaggle for scoring and retrieve the resulting submission status/score.
  11. If Kaggle returns a code-competition error or vague scoring failure, enter the debugging loop: retrieve available logs, classify likely…
  12. Track local CV, remote Kaggle run status, Kaggle submission score, public LB, private-risk notes, seed, code version, data version, and…

What it can do on your machine

Read from SKILL.md and the folder at commit f07e4fa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentic Kaggle Workflow loads about 4k tokens when it runs, and up to ~32k if it reads all its reference files. Until then it costs about 143 tokens; SKILL.md has 1,840 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~143
When it runs · the whole SKILL.md, loaded when a task matches
~4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~32k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from FrankS-IntelLab/agentic-kaggle-skill at commit f07e4fa, republished under its MIT licence (© FrankS-IntelLab). 1,840 words, ~4,023 tokens.

Download SKILL.mdSave it as .claude/skills/agentic-kaggle-skill/SKILL.md (or your agent's skills folder). This skill also uses 34 other files; get the full folder from GitHub.
name
agentic-kaggle-skill
description
Kaggle-first end-to-end competition workflow for scored submissions. Use when Codex must run Kaggle or competitive ML workflows through scored submission, including code competitions, validation, metrics, policy-safe public notebook/discussion intel, tabular/text/image modeling, tuning, ensembling/stacking, proactive multi-notebook architectures, producer notebooks that train models and export private Kaggle artifact datasets, downstream consumer notebooks, Kaggle GPU offload, kagglehub access, hidden-test debugging, and public score retrieval.

Agentic Kaggle Skill

Operating Loop

Treat every competition as a validation problem first and a modeling problem second. The default target platform is Kaggle, so prefer Kaggle-native notebooks/scripts, datasets, model artifacts, competition submissions, and score receipts. For code competitions, assume the final notebook/kernel will be rerun by Kaggle against hidden data unless the competition docs prove otherwise.

  1. Read the competition page, classify the submission mode as classic file submission or code/notebook scoring, then inspect the rules, data-use terms, sharing policy, data dictionary, metric, submission format, train/test construction hints, and leakage warnings.
  2. If live competition intelligence tools are available, inspect top public open notebook solutions and relevant discussion activity before major architecture choices; treat them as clues, not authority.
  3. Identify the task type: binary, multiclass, multilabel, regression, ranking, image, segmentation, text, time series, grouped entities, or a hybrid.
  4. Design folds before feature engineering or modeling. Prefer a fold column saved into the training data so every experiment uses the same comparison surface.
  5. Build the simplest metric-correct baseline and produce out-of-fold (OOF) predictions plus a valid submission.
  6. Proactively plan a stronger architecture once the baseline is trustworthy: diverse model families, feature/embedding producers, augmentation/pseudo-label/distillation stages, calibration/postprocessing, and an ensemble or stacker.
  7. Iterate with validation gates: test one meaningful change at a time when possible, but launch several independent producer notebooks in parallel when they create diverse artifacts that can be compared by OOF score or ensemble diversity.
  8. Offload heavy training, inference, embedding generation, image/text experiments, or memory-risky jobs to Kaggle notebooks/scripts when local compute may OOM or take too long.
  9. For sophisticated architectures, split work into a small Kaggle pipeline: run several independent producer notebooks/scripts first, save each useful producer output as a private Kaggle dataset, then run one consumer notebook/script that attaches those datasets and creates the final OOF/test/submission outputs. When a producer trains a model, its checkpoint, tokenizer/config, fold metadata, OOF/test predictions, and manifest should be exported as a Kaggle dataset; downstream notebooks load the model from /kaggle/input/....
  10. Submit the final submission-producing artifact to Kaggle for scoring and retrieve the resulting submission status/score.
  11. If Kaggle returns a code-competition error or vague scoring failure, enter the debugging loop: retrieve available logs, classify likely failure mode, patch defensively, rerun the final kernel, resubmit, and repeat until scored or concretely blocked.
  12. Track local CV, remote Kaggle run status, Kaggle submission score, public LB, private-risk notes, seed, code version, data version, and artifact paths for every run.
  13. Ensemble only with OOF predictions generated without in-fold leakage.

Resource Map

  • Read references/method-map.md for the neutral workflow map behind the skill.
  • Read references/information-sharing-policy.md before publishing competition code, notebooks, datasets, models, artifacts, or reports outside the active team or workspace.
  • Read references/competition-intel.md when a Kaggle competition slug, URL, title, public leaderboard context, open solutions, or discussion activity may inform the approach.
  • Read references/cross-validation-and-metrics.md when choosing folds, metrics, thresholds, or leakage checks.
  • Read references/tabular-workflow.md for categorical variables, feature engineering, selection, and hyperparameter tuning.
  • Read references/image-text-workflow.md for image, segmentation, and NLP competition approaches.
  • Read references/kaggle-code-competition-pipeline.md when the target is a Kaggle code competition, hidden rerun, final notebook scoring flow, or model artifact handoff between producer and consumer notebooks.
  • Read references/advanced-notebook-architecture.md when a stronger solution may need multiple model families, staged feature/embedding/model producers, pseudo-labeling, distillation, postprocessing, blending, stacking, or parallel Kaggle GPU notebooks.
  • Read references/kaggle-offload.md when local runs are heavy, GPU is useful, Kaggle data access is needed, or remote notebook results must be collected.
  • Read references/kaggle-pipeline-datasets.md when using multiple Kaggle notebooks, intermediate datasets, notebook output sources, or kagglehub.
  • Read references/submission-endgame.md before stopping work; this skill is not done until Kaggle scoring has been attempted and the result has been collected or a concrete blocker is documented.
  • Read references/code-competition-debugging.md when a Kaggle code competition submission fails, times out, OOMs, produces no score, or reports a vague hidden-run/scoring error.
  • Read references/ensembling-and-reproducibility.md for project layout, OOF artifacts, stacking, blending, and repeatability.
  • Read references/research/ or examples/ only when the user asks for historical case studies, Hermes-era patterns, or concrete competition lessons from this repository.
  • Run scripts/scaffold_competition.py to create a competition workspace.
  • Run scripts/make_folds.py to add a fold column to a training CSV.
  • Run scripts/prepare_kaggle_kernel.py to create a Kaggle kernel folder with metadata and retrievable experiment logs.
  • Run scripts/prepare_kaggle_dataset.py to create metadata and commands for versioned intermediate artifact datasets.

Default Kaggle Procedure

Start with the metric and validation:

  • Reproduce the competition metric locally, including clipping, transforms, averaging mode, thresholds, and sample weights.
  • Check whether the competition allows public code, external data, pretrained weights, generated artifacts, and public repositories before sharing anything outside the team.
  • Classify the Kaggle submission path early. If it is a code competition or notebook-scored workflow, design the final consumer notebook as the scoring artifact from the start.
  • Use kaggle-competition-intel MCP tools when available: top_open_solutions for score-ranked plus latest high-vote notebooks, and competition_discussions for top-voted plus recently active discussion topics before choosing expensive baselines.
  • Choose folds to match the hidden test distribution. Use stratification for classification, binned stratification for regression, group splits for repeated entities, time-aware splits for temporal data, and multilabel-aware splits when label co-occurrence matters.
  • Save OOF predictions for every model. OOF files are the currency for error analysis, threshold tuning, blending, and stacking.
  • Compare mean fold score and fold variance, not only the best fold.
  • Treat a public LB jump that disagrees with CV as a validation question before treating it as a modeling victory.

Then build baselines in this order:

  1. Metric-only or constant baseline to verify scoring and submission format.
  2. Fast classical baseline: LightGBM/CatBoost/XGBoost for tabular, TF-IDF plus linear model for text, pretrained backbone for images.
  3. Strong baseline with stable folds, clean artifacts, and OOF predictions.
  4. Feature/model experiments with run logs and one primary metric.
  5. Ensembling or stacking after several diverse, individually validated models exist.

Do not wait for the user to explicitly ask for sophistication when the competition warrants it. After a stable baseline, propose and execute a stronger staged architecture if public solutions, discussion intel, data modality, metric pressure, or CV plateau suggests that a single notebook/model will underperform. Keep the architecture gated by OOF evidence, artifact manifests, and final Kaggle scoring.

When a run is likely to exceed local RAM/VRAM, require a GPU/TPU, or needs Kaggle-only data mounts, prepare a Kaggle kernel run instead of forcing it locally. The remote run must emit experiment_log.json, metrics.jsonl, and an artifact manifest so results can be retrieved and interpreted after kaggle kernels output.

When the solution has multiple heavy stages, keep the remote graph shallow and inspectable:

  • Use fewer than five independent producer notebooks/scripts per wave.
  • Give each producer a single responsibility such as feature generation, embedding extraction, fold training, checkpoint training, model distillation, or inference.
  • For code competitions, prefer producer notebooks that train models/checkpoints and export them as private Kaggle datasets; the final consumer should attach those datasets, load models from /kaggle/input/..., and perform hidden-test-safe inference.
  • Save producer notebook outputs as versioned private Kaggle datasets for durable reuse and as stable inputs to downstream notebooks; use direct kernel_sources only for short-lived chains where durability/versioning is unnecessary.
  • Run one final consumer notebook/script that attaches the produced datasets, reads their artifacts, writes the final OOF, test predictions, blend/stack report, and submission, then submit that final artifact/kernel to Kaggle for scoring.
  • Use kagglehub inside Python when it is more convenient to download datasets, competition files, notebook outputs, or upload dataset versions programmatically.
Show full SKILL.md (611 more words)Show less

Definition Of Done

Do not stop at a plan, scaffold, trained model, notebook run, downloaded output, or local validation score. Continue until one of these is true:

  • A competition submission has been sent to Kaggle, Kaggle has processed it, and the public score or submission status has been retrieved and recorded.
  • A code-competition final notebook/kernel has consumed all required producer artifacts, has been submitted with its kernel/version reference, and the resulting submission status/score has been retrieved and recorded.
  • A concrete external blocker prevents scoring: missing Kaggle credentials, rules not accepted, quota exhausted after retry attempts, competition closed, scoring disabled, required manual UI-only action, unavailable kernel version, or Kaggle service failure. Record the exact blocker and the next command/action needed.

Validation Choices

Use this quick mapping:

  • Balanced classification: StratifiedKFold.
  • Imbalanced classification: stratified folds plus metric-sensitive thresholding or class weights.
  • Regression: KFold for ordinary targets, binned stratified folds for skewed or multimodal targets.
  • Same user, patient, product, session, document, or image source appears multiple times: GroupKFold or stratified group splitting.
  • Ordered events or forecasting: time split, expanding-window validation, or competition-specific period split.
  • Images from same subject or scene: group by subject/scene before augmentation.
  • Text with duplicated sources, authors, questions, products, or prompts: group by source key.
  • Tiny data: repeated folds or more folds, but keep the final model selection discipline strict.

Modeling Priorities

  • For tabular competitions, prefer tree boosting first unless the data is mostly sparse text or high-cardinality categorical interactions. Add CatBoost when categorical columns are central.
  • For linear models and neural networks, scale numeric features and encode categoricals carefully.
  • For tree models, start with label/count/frequency encodings and sensible missing values; scaling is usually unnecessary.
  • For target encoding, always compute encodings out-of-fold, with smoothing and no validation-row target visibility.
  • For image tasks, get dataset, transforms, labels, masks, and metric correct before changing architecture.
  • For text tasks, keep a TF-IDF baseline even when using transformers; it is fast, interpretable, and often blends well.
  • For hyperparameter search, tune only after folds and baselines are trustworthy. Search learning rate, depth/leaves, regularization, sampling, and number of estimators with early stopping.

Competition Hygiene

  • Keep raw data read-only.
  • Keep competition data, private datasets, model checkpoints, downloaded outputs, submissions, credentials, and data-derived artifacts out of public Git repositories unless the competition rules and artifact licenses explicitly allow public redistribution.
  • Put all experiment-changing values in config or command-line args.
  • Seed Python, NumPy, framework, and model libraries when possible.
  • Save model files, OOF predictions, test predictions, fold scores, feature lists, and submission files with run IDs.
  • For Kaggle-offloaded jobs, write retrievable logs and artifacts into the notebook working directory, including run ID, git commit, config, fold, metric, hardware, elapsed time, OOM notes, and output file paths.
  • For multi-notebook pipelines, maintain a pipeline manifest with producer kernel refs, produced dataset refs, dataset versions, artifact schemas, metrics, and the final consumer run ID.
  • Keep producer datasets, kernels, and model artifacts private by default. Make them public only after verifying the target competition rules, data license, third-party IP, and user intent.
  • Record submission message, submission command or kernel/version, Kaggle submission timestamp, status, public score if available, and any error returned by Kaggle.
  • For vague code-competition failures, record the observed status, available logs, suspected failure class, patch attempted, and next retry command.
  • Never train a stacker on in-sample base predictions.
  • Never select features, target encoders, scalers, thresholds, or augmentations using validation targets outside the fold boundary.
  • Prefer scripts for repeatable training and notebooks for exploration and plots.

Useful Commands

Create a competition skeleton:

bash
python3 <skill-dir>/scripts/scaffold_competition.py --root .

Create stratified folds:

bash
python3 <skill-dir>/scripts/make_folds.py \
  --input input/train.csv \
  --output input/train_folds.csv \
  --target target \
  --strategy stratified \
  --n-splits 5

Create grouped folds:

bash
python3 <skill-dir>/scripts/make_folds.py \
  --input input/train.csv \
  --output input/train_folds.csv \
  --target target \
  --group-col patient_id \
  --strategy group

Prepare a Kaggle GPU kernel experiment folder:

bash
python3 <skill-dir>/scripts/prepare_kaggle_kernel.py \
  --output kaggle_kernels/exp_lgbm_gpu \
  --username kaggle-user \
  --slug exp-lgbm-gpu \
  --title "exp lgbm gpu" \
  --competition playground-series-sample \
  --accelerator NvidiaTeslaT4

Prepare a private Kaggle dataset folder for intermediate artifacts:

bash
python3 <skill-dir>/scripts/prepare_kaggle_dataset.py \
  --output kaggle_datasets/exp_features_v1 \
  --username kaggle-user \
  --slug exp-features-v1 \
  --title "exp features v1" \
  --description "OOF-safe feature artifacts for experiment exp_features_v1"

Prepare a private Kaggle dataset folder for trained model artifacts:

bash
python3 <skill-dir>/scripts/prepare_kaggle_dataset.py \
  --output kaggle_datasets/exp_model_v1 \
  --username kaggle-user \
  --slug exp-model-v1 \
  --title "exp model v1" \
  --description "Trained model artifacts for experiment exp_model_v1" \
  --artifact-kind model

© FrankS-IntelLab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 34 other files (scripts, references) in the repository root of FrankS-IntelLab/agentic-kaggle-skill.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.md
  • agents/openai.yaml
  • examples/audio-classification-case-study.md
  • examples/rl-game-case-study.md
  • references/advanced-notebook-architecture.md
  • references/code-competition-debugging.md
  • references/competition-intel.md
  • references/cross-validation-and-metrics.md
  • references/ensembling-and-reproducibility.md
  • references/image-text-workflow.md
  • references/information-sharing-policy.md
  • references/kaggle-code-competition-pipeline.md
  • references/kaggle-offload.md
  • references/kaggle-pipeline-datasets.md
  • references/method-map.md
  • … and 17 more

Open the folder on GitHubat commit f07e4fa

Compare with similar skills

Agentic Kaggle Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentic Kaggle Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentic Kaggle Workflow this skillFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT
Code Generatorliangdabiao/claude-data-analysis-ultra-main290—~513Automated safety check: PassNone
Data Scientistdavila7/claude-code-templates33k8 repos~2.6kAutomated safety check: PassMIT
Scientific Toolkit SkillzLanqing/codex-claude-academic-skills4.7k—~1.2kAutomated safety check: PassMIT
Data Sciencetravisjneuman/.claude100—~2.3kAutomated safety check: PassMIT
Research ML Practiceprobabl-ai/skills138—~1.2kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Code Generator

    liangdabiao/claude-data-analysis-ultra-main

    Generates production-ready analysis code in Python, R, SQL. An agent skill from liangdabiao/claude-data-analysis-ultra-main.

    290 GitHub stars~513 tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check passed
  • Data Scientist

    davila7/claude-code-templates

    Expert data scientist for advanced analytics, machine learning, and statistical modeling.

    33k GitHub starsUsed in 8 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Toolkit Skill

    zLanqing/codex-claude-academic-skills

    Research computing toolkit for optoelectronic information science and engineering, MATLAB/Octave, Python scientific analysis, signal processing, image processing, statistics, simulation…

    4.7k GitHub stars~1.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Data Science

    travisjneuman/.claude

    Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy.

    100 GitHub stars~2.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Research ML Practice

    probabl-ai/skills

    Literature and web research for an ML methodology concern (EDA extra measurements, leakage, transforms, feature engineering, learner family), or an EDA extra-analysis survey from JOURNAL plus…

    138 GitHub stars~1.2k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Fcr Data Analysis

    franklee16/academic-research-skills

    A skill your agent uses when executing and reporting the statistical analysis for a Field Crops Research (FCR) manuscript — mixed models for multi-environment and blocked/split-plot designs…

    223 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed

Works with

Questions about Agentic Kaggle Workflow

What does Agentic Kaggle Workflow do?

Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission. The skill treats each competition as a validation problem first and a modeling problem second, defaulting to Kaggle-native notebooks, datasets and score receipts. It starts by classifying the submission mode (file or code competition) and reading the rules, data terms, sharing policy, metric and leakage warnings, then identifies the task type and designs folds before any feature work, preferably as a fold column saved in the training data.

When should I use Agentic Kaggle Workflow?

Agentic Kaggle Workflow fits situations like: running a Kaggle competition through to a scored submission; designing cross-validation folds and metrics before modeling; debugging a code competition that fails on hidden test data; splitting work across producer and consumer notebooks with private artifact datasets.

How do I install Agentic Kaggle Workflow in Claude Code?

Run `npx skills add FrankS-IntelLab/agentic-kaggle-skill --skill agentic-kaggle-skill -a claude-code`. Or copy the skill folder (the FrankS-IntelLab/agentic-kaggle-skill repository) into .claude/skills/agentic-kaggle-skill in your project. Claude Code loads it when a task matches its description.

How do I install Agentic Kaggle Workflow in Codex?

Run `npx skills add FrankS-IntelLab/agentic-kaggle-skill --skill agentic-kaggle-skill -a codex`. Or copy the skill folder (the FrankS-IntelLab/agentic-kaggle-skill repository) into .agents/skills/agentic-kaggle-skill in your project. Codex loads it when a task matches its description.

Can I use Agentic Kaggle Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FrankS-IntelLab/agentic-kaggle-skill --skill agentic-kaggle-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentic-kaggle-skill, .gemini/skills/agentic-kaggle-skill, .github/skills/agentic-kaggle-skill and .opencode/skills/agentic-kaggle-skill in your project.

What does Agentic Kaggle Workflow need to run?

Going by SKILL.md and its folder, Agentic Kaggle Workflow needs the command-line tools its instructions call (python3). Our summary lists: A Kaggle account with access to the competition; Kaggle CLI or kagglehub access for datasets and submissions.

Does Agentic Kaggle Workflow access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentic Kaggle Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agentic Kaggle Workflow use?

Agentic Kaggle Workflow is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentic Kaggle Workflow use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 28k tokens, read only when the agent opens those files.

What are the alternatives to Agentic Kaggle Workflow?

Skills that share tags, products or a category with Agentic Kaggle Workflow: Code Generator (liangdabiao/claude-data-analysis-ultra-main, 290 stars), Data Scientist (davila7/claude-code-templates, 33k stars), Scientific Toolkit Skill (zLanqing/codex-claude-academic-skills, 4.7k stars) and Data Science (travisjneuman/.claude, 100 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentic Kaggle Workflow?

FrankS-IntelLab (a GitHub user) maintains it in FrankS-IntelLab/agentic-kaggle-skill, which has 188 GitHub stars. The repository was last updated on June 15, 2026.

Source: FrankS-IntelLab/agentic-kaggle-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.