Scikit Learn
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom…
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-custom-node-authoring --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/flowfile-custom-node-authoring .claude/skills/flowfile-custom-node-authoring && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "flowfile-custom-node-authoring" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-custom-node-authoring into .claude/skills/flowfile-custom-node-authoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-custom-node-authoring", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-custom-node-authoringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-custom-node-authoring --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/flowfile-custom-node-authoring .agents/skills/flowfile-custom-node-authoring && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "flowfile-custom-node-authoring" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-custom-node-authoring into .agents/skills/flowfile-custom-node-authoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-custom-node-authoring", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-custom-node-authoring --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/flowfile-custom-node-authoring .cursor/skills/flowfile-custom-node-authoring && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "flowfile-custom-node-authoring" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-custom-node-authoring into .cursor/skills/flowfile-custom-node-authoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-custom-node-authoring", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Edwardvaneechoud/Flowfile.git --path .claude/skills/flowfile-custom-node-authoring--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-custom-node-authoring --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/flowfile-custom-node-authoring .gemini/skills/flowfile-custom-node-authoring && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "flowfile-custom-node-authoring" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-custom-node-authoring into .gemini/skills/flowfile-custom-node-authoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-custom-node-authoring", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-custom-node-authoringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/flowfile-custom-node-authoring .github/skills/flowfile-custom-node-authoring && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "flowfile-custom-node-authoring" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-custom-node-authoring into .github/skills/flowfile-custom-node-authoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-custom-node-authoring", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-custom-node-authoring --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/flowfile-custom-node-authoring .opencode/skills/flowfile-custom-node-authoring && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "flowfile-custom-node-authoring" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-custom-node-authoring into .opencode/skills/flowfile-custom-node-authoring/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-custom-node-authoring", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
flowfile-custom-node-authoringTurn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom…
Flowfile Custom Node Authoring is an agent skill from Edwardvaneechoud/Flowfile. Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom node .py authored with the nodedesigner SDK (from flowfile import nodedesigner as nd, CustomNodeBase, NodeSettings/Section, ColumnSelector/SingleSelect/NumericInput/…). Use when a task says "generate/create/author/write a custom node", "make a node that does X", "a kernel node", "an sklearn/xgboost/lightgbm ML node", "a…
Its SKILL.md is about 8.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Machine learning, Plain language and style rules and Vulnerability scanning. It works with WebAssembly and scikit-learn. The repository describes itself as: Flowfile is a visual ETL tool and Python library combining drag-and-drop workflows with Polars dataframes. Build data pipelines visually, define flows programmatically with a… The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c054c90. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonpoetryFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Flowfile Custom Node Authoring loads about 8.6k tokens when it runs. Until then it costs about 228 tokens; SKILL.md has 2,608 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Edwardvaneechoud/Flowfile at commit c054c90, republished under its MIT licence (© Edwardvaneechoud). 2,608 words, ~8,562 tokens.
.claude/skills/flowfile-custom-node-authoring/SKILL.md (or your agent's skills folder).Generate one .py file that defines a user custom node via the node_designer SDK, from a description of what the node should do. The file is the whole node: the visual Node Designer round-trips it, the worker (or a Docker kernel) executes its process(), and the community registry publishes it as-is.
flowfile_core settings schema + FlowGraph + frontend registry + flowfile_frame + wasm) → flowfile-node-development. That is four registries agreeing by naming convention; this skill is one self-contained file.CLAUDE.md and flowfile_core/flowfile/community_nodes/.FLOWFILE_KERNEL_IMAGE_*, kernel↔core auth → flowfile-config-and-flags + flowfile_core/flowfile_core/kernel/.KernelRequiredError, exec-vs-AST scan failures → flowfile-debugging-playbook.flowfile-frame-and-codegen.All facts below verified in-repo on 2026-07-12, branch feature/add-custom-node-store. Re-verification commands are in §7 — run them before trusting a line number in a stale checkout.
my_node.py
├─ import polars as pl
├─ from flowfile import node_designer as nd # the ONLY flowfile import allowed
├─ class MyNodeSettings(nd.NodeSettings): ... # optional: the settings UI
└─ class MyNode(nd.CustomNodeBase): # exactly ONE such subclass per file
node_name / metadata fields ...
settings_schema: MyNodeSettings = MyNodeSettings()
def process(self, *inputs: pl.LazyFrame) -> pl.LazyFrame: ...from flowfile import node_designer as nd. The SDK lives in shared/node_designer/ and is import-pure (no flowfile_core, DB, or migrations). A node file may import only the SDK + stdlib + third-party libraries — any other flowfile/flowfile_core.* import raises a contract error in worker/dry-run contexts (shared/node_designer/loading.py:22-26,48-54).CustomNodeBase subclass per file — more raise NodeLoadError (loading.py:146-156).~/.flowfile/user_defined_nodes/. Community installs write a flat <id>.py there.environment="local"), or in an isolated Docker kernel (environment="kernel"). See §3.Follow these steps in order. Each maps to a section below.
environment="local" (omit the field). Needs ML libs or isolation → environment="kernel" + dependencies=[...] (plain PyPI specs only). Never set execution_location — that is a core-level concept, not an SDK field.settings_schema (§2.3–2.5). One control per knob, grouped into Sections. Keep every kwarg a designer literal (§4.4).process(self, *inputs) (§2.2, §2.6). Read settings via self.settings_schema.<section>.<component>.value; .collect() when a library needs eager data; int(...)-cast NumericInput/SliderInput values; for kernel nodes import heavy libs inside process and never reference SDK types; return a LazyFrame/DataFrame, or a dict keyed by output_names for multi-output.node_name (required), plus node_category/title/intro; author/version/tags for publishing. Add example_inputs + example_settings (both required to install/publish; both must be literals).CustomNodeBase subclass with node_name + process. PNG icon only for published bundles.dry_run.py; place it in ~/.flowfile/user_defined_nodes/.When you have a ~/.flowfile/user_defined_nodes/ reachable and the user wants it installed, write the file there directly; otherwise write to the path the user names (or the repo flowfile-community-nodes/nodes/<id>/node.py layout when they intend to publish).
CustomNodeBase fields (shared/node_designer/custom_node.py:401-444)A Pydantic model — declare attributes as annotated fields with defaults: node_name: str = "My Node".
| Field | Type | Default | Purpose |
|---|---|---|---|
node_name | str | required | Node name; slugged to the palette item. |
node_category | str | "Custom" | Palette group. |
node_icon | str | "user-defined-icon.png" | Bare icon filename (PNG for publishing — SVG is forbidden in bundles, §4.3). |
settings_schema | NodeSettings | None | None | The settings-UI object (a NodeSettings subclass instance). |
number_of_inputs | int | 1 | Input ports → inputs[0..n-1]. |
number_of_outputs | int | 1 | Output ports. |
output_names | list[str] | ["main"] | Keys for multi-output dict returns. |
environment | Literal["local","kernel"] | "local" | Execution environment (§3). |
dependencies | list[str] | [] | pip specs; auto-installed only when environment="kernel". |
example_inputs | list[dict[str,list]] | None | None | Dry-run/publish sample data: one {col: [values]} per input port. Each entry must actually build a pl.LazyFrame (§2.7). |
example_settings | dict[str,dict[str,Any]] | None | None | Dry-run/publish settings: {section: {component: value}}. |
title | str | None | "Custom Node" | Settings-drawer title. |
intro | str | None | "A custom node…" | Settings-drawer intro. |
author / version / tags | str|None / str|None / list[str] | None / None / [] | Publishing metadata. |
node_type | Literal["input","output","process"] | "process" | Node role. |
transform_type | Literal["narrow","wide","other"] | "wide" | Transform hint. |
requires_kernel: bool and kernel_id: str|None are deprecated — they map onto environment="kernel". Use environment.
process signature (custom_node.py:732-750)def process(self, *inputs: pl.LazyFrame) -> pl.LazyFrame | pl.DataFrame:pl.LazyFrame per connected input, in port order (inputs[0], inputs[1], …). .collect() inside process when you need eager data.LazyFrame or DataFrame (the framework normalizes).len(output_names) > 1): return dict[str, pl.LazyFrame | pl.DataFrame] keyed by the output_names.NodeSettings + Sectionclass MyNodeSettings(nd.NodeSettings):
my_section: nd.Section = nd.Section(
title="…", description="…", layout="vertical", # or "horizontal"
some_field=nd.TextInput(label="…"),
another=nd.NumericInput(label="…", default=1.0),
)Section(title=None, description=None, hidden=False, layout="vertical", **components) — every keyword after the first four is a component that becomes a field (ui_components.py:254-317).
FlowfileInComponent; import as nd.<Name>)| Component | Constructor args (defaults) |
|---|---|
TextInput | label, default="", placeholder="" |
FilePicker | label, default="", placeholder="", mode="open"|"create", file_types=[], allow_directory=False — path string via the standard file browser (server-side filesystem) |
NumericInput | label, default=None, min_value=None, max_value=None |
SliderInput | label, default=None, min_value=0, max_value=100, step=1 |
ToggleSwitch | label, default=False, description=None |
SingleSelect | label, options (required), default=None |
MultiSelect | label, options (required), default=[] |
ColumnSelector | label, required=False, multiple=False, data_types="ALL" |
ColumnActionInput | label, actions=[], output_name_template="{column}_{action}", show_group_by=False, show_order_by=False, data_types="ALL" |
SecretSelector | label, options=AvailableSecrets, required=False, description=None, name_prefix=None |
options= for the selects is a list of bare strings or (value, label) tuples, or a dynamic marker class passed by class (not instance):
nd.IncomingColumns — populate from the input frame's columns.nd.AvailableArtifacts — populate from upstream artifact names.nd.AvailableSecrets — the default for SecretSelector.data_types= (shared/node_designer/types.py)Groups: nd.Types.Numeric, .String, .AnyDate, .Boolean, .Binary, .Complex, .All. Specific: .Int/.Int64/…, .Float/.Float64, .Str, .Date, .Datetime, .Bool, .List, .Struct, etc. Accepts one value, a list, bare strings ("Numeric", "Int64", "int"/"float"/"str"), or a real pl.DataType. "ALL" = no filter. Both data_types=nd.Types.String and data_types=["Numeric"] are valid spellings.
process()Access: self.settings_schema.<section_attr>.<component_attr>.value. Gotchas:
NumericInput / SliderInput .value is a number but is not coerced — it arrives as whatever was supplied (an int or a float, e.g. example_settings with "n_clusters": 3 yields int). Cast with int(...)/float(...) when you need a specific type.ColumnSelector(multiple=True).value → list[str]; multiple=False → a single column-name str.SecretSelector → read .secret_value (a SecretStr; call .get_secret_value()), not .value (which is the secret name). Local-only — secrets are not available inside kernels yet (§3.3).ColumnActionInput.value → a ColumnActionValue (.rows of ColumnActionRow, .group_by_columns, .order_by_column).is_dry_run() and example_artifacts()A dry run runs your real process() against example_inputs, but sandboxes its side effects. Two paths do this: the designer Test tab (user_defined/dry_run.py) and the community-node CI harness (community_nodes/dry_run_local.py, run by the registry's PR checks).
example_inputs must build a real frame. Polars infers each column's dtype from its first value, so "x1": [5, 6.5, 4.7, 6.9] infers Int64 and then rejects 6.5. Write 5.0. Bundle validation constructs every entry and reports EXAMPLES_INVALID (validation.py::_check_example_inputs_build), so this fails at validate time rather than as a polars traceback mid-dry-run.
flowfile_ctx.is_dry_run() -> bool — True only in a dry run; False in a real flow run, a single-file export, and an exported project. Branch on it to skip work that only makes sense against real data:
if not flowfile_ctx.is_dry_run():
notify_downstream(result)example_artifacts() — a node that reads an artifact another node publishes (a predict node loading a train node's model) has nothing to read in a dry run, so it fails on a missing artifact instead of exercising its real path. Override the hook to supply one:
def example_artifacts(self) -> dict:
"""Objects the dry-run seeds into the artifact store before process() runs."""
from sklearn.ensemble import RandomForestClassifier # heavy imports inside, as in process()
...
return {"automl_model": {"pipeline": pipeline, "task": "classification", ...}}_check_examples only requires example_inputs/example_settings to be literals.{name: obj} dict. Each entry is seeded into both the flow-local store (read_artifact) and the global store (get_global), in the CI harness and the Test tab alike.fs_read / network to the node's consent disclosure (§4.2), and an inlined base64 model blob over 1024 chars is a hard FF-SEC-011 deny (§4.1).What the harness fakes. The dry-run flowfile_ctx is a faithful in-memory stand-in, not a no-op: publish_global returns a real positive id and versions on republish, get_global/read_artifact raise KeyError on a miss, publish_artifact raises ValueError on a duplicate name, and get_shared_location returns a writable temp path. Catalog APIs (read_catalog_table, list_catalogs, …) need a running server and raise a clear error — guard them with if not flowfile_ctx.is_dry_run():. Anything unmodelled falls through to a no-op returning None. A drift guard (tests/flowfile/community_nodes/test_dry_run_ctx_parity.py) fails CI if a new kernel API is added without a matching fake.
flowfile_ctxis only bound forenvironment="kernel"nodes. In alocalnode it is an undefined name at runtime.
environment="local" (default): runs in the Flowfile process/worker. Full SDK + polars available; secrets resolve; no Docker. Use for pure-polars/stdlib transforms.environment="kernel": runs in an isolated Docker kernel. Use when you need heavy libs (sklearn, xgboost, lightgbm, statsmodels) or isolation. Declare dependencies (plain PyPI specs) — auto-installed only for kernel nodes. A kernel node must be bound to a kernel before it runs; the user picks the kernel (and image flavour) in the UI. Unbound → KernelRequiredError (raised in flow_graph.py).A kernel node's source never runs directly. Core AST-generates a self-contained script (user_defined/kernel_codegen.py::generate_kernel_script, invoked from flow_graph.py). Consequences you must design around:
polars/json/logging/sys survive. Do not reference SDK types inside process.process + nested/helper defs are kept; all class-level attribute assignments are dropped. Put every bit of logic inside process — nothing computed at class-body scope exists at kernel runtime.self.settings_schema.<section>.<component>.value still works. Every settings value must be JSON-serializable or generation fails (KernelCodegenError).pl.scan_parquet per input); .collect() for eager work. Outputs are marshalled back via an injected flowfile_ctx.publish_output epilogue.flowfile_ctx is an injected runtime global (undefined in the plain file — that's expected; linters will flag it). Surface includes log_info(...) (surfaces in the Test panel), read_inputs/read_first, publish_artifact(name, obj) / read_artifact(name) (cloudpickle-backed) for persisting a trained model across nodes, publish_global / get_global for persisting one across flows, and is_dry_run() (§2.7).flowfile_ctx only.SecretSelector.secret_value raises there. Keep secret-using nodes local.process() is one script execution. A model trained and used within the same process() is just a local variable. To share a model between a train node and a predict node, flowfile_ctx.publish_artifact(...) / read_artifact(...).kernel_runtime/pyproject.toml + poetry.lock)polars, pyarrow, numpy, cloudpickle, joblib, deltalake (+ plumbing). No ML libs.scikit-learn 1.7.2, xgboost 2.1.4, lightgbm 4.6.0, statsmodels 0.14.6, polars-ds 0.12.0 (+ pandas transitive). This is the flavour with xgboost/sklearn.An XGBoost/sklearn kernel node works today with no code changes — provided the bound kernel uses the ML image, or the node declares the lib in dependencies for a per-kernel derived build.
environment: str = "kernel".dependencies: list[str] = ["xgboost"] (only strictly needed if not relying on the ML image).xgboost, sklearn, …) inside process(), never at module top. Core execs the whole module at placement (registry.ensure_class) to build the palette template, and core does not have the kernel-only libs installed — a module-top import xgboost raises ModuleNotFoundError and the node fails to load. Only polars + the nd import belong at module top. In-process third-party imports are still lifted verbatim into the kernel script (this is what the real kmeans node does — §5(b)).settings_schema (JSON-serializable), read via .value..collect() inputs inside process — sklearn/xgboost need eager numpy/pandas.flowfile_ctx.publish_artifact / read_artifact.The scanner (community_nodes/security_scan.py) is a conservative pre-filter; the community PR review + server-side re-scan are authoritative. Two outcomes: DENY (rejects the node) and FLAG (surfaces a capability for the consent dialog; never fails validation).
security_scan.py:15-37)eval/exec/compile; __import__/importlib.import_module/importlib.util.getattr/operator.attrgetter/operator.methodcaller/globals()/vars()/__builtins__ resolving a dangerous builtin, or any of them with a non-constant name.ctypes/cffi/_ctypes; os.system/os.popen*/os.exec*/os.spawn* (and posix.system/posix.popen*).subprocess with shell=True or non-literal args; pty/os.forkpty; importing both socket and subprocess.sys._getframe/inspect.stack/inspect.currentframe; importing pip/ensurepip.os.environ enumeration.network (socket/requests/httpx/urllib/boto3/…), subprocess (literal args), fs_write, fs_read, env_read (os.getenv/sys.argv), dynamic_code (setattr/delattr with non-constant name), serialization (pickle.load/yaml.load w/o SafeLoader), secrets (SecretSelector/.secret_value). Only trigger these when the node genuinely needs them.
ML libraries do not flag — xgboost, sklearn/scikit-learn, numpy, pandas, polars match no trigger set. (Caveat: pickle.load/yaml.load without SafeLoader flag serialization; joblib.load is not detected by the scanner but still deserializes untrusted data — prefer flowfile_ctx.read_artifact for models.)
community_nodes/validation.py:385-445)node.py (≤200 KB) + manifest.json (≤16 KB) required; optional icon.png (≤256 KB, ≤512×512), README.md, screenshots/ (≤5 PNGs). SVG forbidden anywhere. Total ≤6 MB.CustomNodeBase (no extra bases/decorators/keywords); defines node_name + process().example_inputs (literal list) and example_settings (literal dict), else EXAMPLES_REQUIRED. Every example_inputs entry must also construct a pl.LazyFrame, else EXAMPLES_INVALID (§2.7).dependencies requires environment="kernel"; each spec is a plain PyPI name + optional version specifier — no URLs, git+, or pip flags.manifest.id == slug of node_name; versions match and strictly increase on update.Keep all node attributes and settings/component kwargs as literals — constants, lists/tuples/dicts of scalars, nd.Types.*/nd.DataType.*, SDK marker classes. No f-strings, comprehensions, computed values, or **kwargs unpacking at that level, or the node degrades to code_only (still valid/installable, but the visual Form tab can't edit it). Logic inside process is unconstrained (node_designer/parsing.py::eval_literal).
docs/examples/custom_node.py (runs under the docs test harness).
import polars as pl
from flowfile import node_designer as nd
class GreetingSettings(nd.NodeSettings):
main_config: nd.Section = nd.Section(
title="Greeting Configuration",
description="Configure how to greet each row",
name_column=nd.ColumnSelector(
label="Name Column",
data_types=nd.Types.String,
required=True,
),
greeting=nd.SingleSelect(
label="Greeting",
options=[("formal", "Hello"), ("casual", "Hey")],
default="casual",
),
)
class GreetingNode(nd.CustomNodeBase):
node_name: str = "Greeting Generator"
node_category: str = "Text Processing"
title: str = "Add greetings"
intro: str = "Prefix a name column with a greeting."
settings_schema: GreetingSettings = GreetingSettings()
def process(self, *inputs: pl.LazyFrame) -> pl.LazyFrame:
lf = inputs[0]
name_col = self.settings_schema.main_config.name_column.value
style = self.settings_schema.main_config.greeting.value
word = "Hello" if style == "formal" else "Hey"
return lf.with_columns(
pl.concat_str([pl.lit(f"{word}, "), pl.col(name_col)]).alias("greeting")
)flowfile-community-nodes/nodes/kmeans_clustering/node.py (real, published). Note environment="kernel", dependencies, the import inside process, .collect() for eager numpy, the int(...) cast, and flowfile_ctx.log_info.
import polars as pl
from flowfile import node_designer as nd
class KmeansClusteringSettings(nd.NodeSettings):
clustering: nd.Section = nd.Section(
title="Clustering",
description="Settings for K-means cluster",
layout="horizontal",
feature_columns=nd.ColumnSelector(
label="Feature Columns", required=True, multiple=True, data_types=["Numeric"],
),
n_clusters=nd.NumericInput(label="Number of clusters", default=3.0, min_value=2.0, max_value=20.0),
cluster_column=nd.TextInput(label="Cluster column name", default="cluster"),
standardize=nd.ToggleSwitch(label="Standardize", default=True),
)
class KmeansClustering(nd.CustomNodeBase):
node_name: str = "kmeans clustering"
node_category: str = "ML"
node_icon: str = "kmeans.PNG"
title: str = "kmeans clustering"
intro: str = "Do a kmeans clustering with sklearn"
author: str = "Edwardvaneechoud"
version: str = "1.0.1"
tags: list[str] = ["machine learning", "data-science"]
environment: str = "kernel"
dependencies: list[str] = ["scikit-learn"]
example_inputs: list[dict[str, list]] = [
{"age": [26, 45, 47], "annual_income_k": [30.3, 95.6, 101.5], "spending_score": [34, 81, 86]},
]
example_settings: dict[str, dict] = {
"clustering": {
"feature_columns": ["age", "annual_income_k", "spending_score"],
"n_clusters": 3, "cluster_column": "cluster", "standardize": True,
},
}
settings_schema: KmeansClusteringSettings = KmeansClusteringSettings()
def process(self, *inputs: pl.LazyFrame) -> pl.LazyFrame:
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
cfg = self.settings_schema.clustering
feature_cols: list[str] = cfg.feature_columns.value
k = int(cfg.n_clusters.value) # .value is not coerced — cast to the type you need
label_col = cfg.cluster_column.value
df = inputs[0].collect() # sklearn needs eager data
features = df.select(feature_cols).to_numpy()
if cfg.standardize.value:
features = StandardScaler().fit_transform(features)
model = KMeans(n_clusters=k, n_init=10, random_state=42)
labels = model.fit_predict(features)
flowfile_ctx.log_info(f"Fitted {k} clusters over {df.height} rows") # injected at runtime
return df.with_columns(pl.Series(label_col, labels))Trains an XGBRegressor/XGBClassifier on labeled rows and appends a prediction column. Self-contained (no cross-node artifact needed); for a real train→serve split see the note after.
import polars as pl
from flowfile import node_designer as nd
class XGBoostPredictSettings(nd.NodeSettings):
model: nd.Section = nd.Section(
title="Model",
description="Train an XGBoost model on this data and predict a target column.",
feature_columns=nd.ColumnSelector(
label="Feature columns", required=True, multiple=True, data_types=["Numeric"],
),
target_column=nd.ColumnSelector(
label="Target column", required=True, data_types=["Numeric"],
),
task=nd.SingleSelect(
label="Task", options=[("regression", "Regression"), ("classification", "Classification")],
default="regression",
),
n_estimators=nd.NumericInput(label="Number of trees", default=200.0, min_value=10.0, max_value=2000.0),
max_depth=nd.NumericInput(label="Max tree depth", default=6.0, min_value=1.0, max_value=20.0),
prediction_column=nd.TextInput(label="Prediction column name", default="prediction"),
)
class XGBoostPredict(nd.CustomNodeBase):
node_name: str = "XGBoost Predict"
node_category: str = "ML"
title: str = "XGBoost predictions"
intro: str = "Train an XGBoost model on the input and append predictions."
author: str = "your-name"
version: str = "1.0.0"
tags: list[str] = ["machine learning", "xgboost", "prediction"]
environment: str = "kernel"
dependencies: list[str] = ["xgboost"]
example_inputs: list[dict[str, list]] = [
{"x1": [1.0, 2.0, 3.0, 4.0], "x2": [10.0, 8.0, 6.0, 4.0], "y": [12.0, 11.0, 10.0, 9.0]},
]
example_settings: dict[str, dict] = {
"model": {
"feature_columns": ["x1", "x2"], "target_column": "y", "task": "regression",
"n_estimators": 200, "max_depth": 6, "prediction_column": "prediction",
},
}
settings_schema: XGBoostPredictSettings = XGBoostPredictSettings()
def process(self, *inputs: pl.LazyFrame) -> pl.DataFrame:
import xgboost as xgb # heavy libs import INSIDE process (core execs the module at placement)
cfg = self.settings_schema.model
feature_cols: list[str] = cfg.feature_columns.value
target_col: str = cfg.target_column.value
pred_col: str = cfg.prediction_column.value
df = inputs[0].collect() # xgboost needs eager numpy
X = df.select(feature_cols).to_numpy()
y = df.get_column(target_col).to_numpy()
params = dict(n_estimators=int(cfg.n_estimators.value), max_depth=int(cfg.max_depth.value))
Model = xgb.XGBClassifier if cfg.task.value == "classification" else xgb.XGBRegressor
model = Model(**params)
model.fit(X, y)
preds = model.predict(X)
flowfile_ctx.log_info(f"Trained {cfg.task.value} on {len(feature_cols)} features, {df.height} rows")
return df.with_columns(pl.Series(pred_col, preds))Train → serve split (advanced). A train node fits the model and calls flowfile_ctx.publish_artifact("xgb_model", model); a separate predict node calls model = flowfile_ctx.read_artifact("xgb_model") (both kernel nodes). Use a SingleSelect/TextInput for the artifact name and nd.AvailableArtifacts to let the predict node pick from published artifacts.
import polars as pl
from flowfile import node_designer as nd
class SplitterSettings(nd.NodeSettings):
main: nd.Section = nd.Section(
title="Split",
split_column=nd.ColumnSelector(label="Split Column", data_types="Boolean", required=True),
)
class RowSplitter(nd.CustomNodeBase):
node_name: str = "Row Splitter"
node_category: str = "Custom"
number_of_outputs: int = 2
output_names: list[str] = ["pass", "fail"]
example_inputs: list[dict[str, list]] = [{"keep": [True, False, True], "v": [1, 2, 3]}]
example_settings: dict[str, dict] = {"main": {"split_column": "keep"}}
settings_schema: SplitterSettings = SplitterSettings()
def process(self, *inputs: pl.LazyFrame) -> dict:
col = self.settings_schema.main.split_column.value
return {"pass": inputs[0].filter(pl.col(col)), "fail": inputs[0].filter(~pl.col(col))}Multi-input: set number_of_inputs: int = 2 and read inputs[0], inputs[1] — e.g. pl.concat([inputs[0], inputs[1]], how="diagonal").
Lint: poetry run ruff check <file> (tests/generated node files are lint-exempt by config, but a clean parse matters). Confirm it imports: the file must import polars as pl and from flowfile import node_designer as nd and nothing else from flowfile.
Dry-run without the app: the designer Test tab drives flowfile_core/flowfile/user_defined/dry_run.py — it runs process against example_inputs/example_settings (kernel nodes go through the same AST-generated script path). This is the fastest way to confirm the node executes.
Install for real: write the file to ~/.flowfile/user_defined_nodes/<name>.py; the registry hot-reloads on save. A broken file stays visible-with-error rather than vanishing.
Publish-readiness: run the bundle validator / CLI (python -m flowfile_core.flowfile.community_nodes.cli) — this is the same code the community repo's CI uses:
python -m flowfile_core.flowfile.community_nodes.cli validate nodes/<id>
python -m flowfile_core.flowfile.community_nodes.cli dry-run nodes/<id> --install-deps --timeout 120 --row-limit 1000dry-run executes process() in a child process against example_inputs, with the in-memory flowfile_ctx of §2.7. A node that reads an artifact it does not publish needs example_artifacts() to pass.
Per the repo's working agreement, do not execute the user's tests or touch his running services — write the file and describe the verification the user can run.
Facts verified 2026-07-12 on feature/add-custom-node-store. To re-verify against a fresh checkout:
# SDK export surface (the exact nd.* names)
sed -n '40,73p' shared/node_designer/__init__.py
# Base-class fields, defaults, process() signature
sed -n '393,779p' shared/node_designer/custom_node.py
# Control catalog + Section
sed -n '49,553p' shared/node_designer/ui_components.py
# Type filters
cat shared/node_designer/types.py
# Kernel codegen (what survives into the kernel script) + kernel_ctx surface
sed -n '1,260p' flowfile_core/flowfile_core/flowfile/user_defined/kernel_codegen.py
cat kernel_runtime/kernel_runtime/flowfile_client.py
# Kernel image contents (which flavour has xgboost/sklearn)
sed -n '1,40p' kernel_runtime/pyproject.toml
# Scanner DENY/FLAG rules + bundle validation
sed -n '1,200p' flowfile_core/flowfile_core/flowfile/community_nodes/security_scan.py
sed -n '385,445p' flowfile_core/flowfile_core/flowfile/community_nodes/validation.py
# Real + tested examples
cat docs/examples/custom_node.py
cat flowfile-community-nodes/nodes/kmeans_clustering/node.py
ls flowfile_core/tests/flowfile/node_designer/corpus/User-facing docs to cross-reference: docs/users/visual-editor/creating-custom-nodes.md, node-designer.md, custom-node-tutorial.md, kmeans-kernel-node.md, kernels.md, kernel-api.md, community-nodes.md.
© Edwardvaneechoud, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/flowfile-custom-node-authoring of Edwardvaneechoud/Flowfile.
Open the folder on GitHubat commit c054c90
Flowfile Custom Node Authoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Flowfile Custom Node Authoring this skillEdwardvaneechoud/Flowfile | 370 | — | ~8.6k | Automated safety check: Pass | MIT | |
| Scikit LearnzLanqing/codex-claude-academic-skills | 4.6k | 17 repos | ~3.9k | Automated safety check: Pass | BSD-3-Clause | |
| Senior Data ScientistRaidriar7170/hermes-skilleval | 125 | 6 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Time Series Analytics Useropen-edge-platform/edge-ai-libraries | 168 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Estimate Online Covariancemicroprediction/precise | 336 | — | ~535 | Automated safety check: Pass | MIT | |
| Aeon Time Series Machine Learningdavila7/claude-code-templates | 32k | 14 repos | ~2.6k | Automated safety check: Pass | MIT |
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
open-edge-platform/edge-ai-libraries
Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…
microprediction/precise
Estimate a covariance / correlation / precision matrix incrementally with precise.
davila7/claude-code-templates
Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.
microprediction/precise
Online (incremental) covariance, correlation, and precision estimation in Python — the streaming complement to sklearn.covariance.
Edwardvaneechoud/Flowfile
Maps the /ai/ subsystem of flowfile_core, its three agent tiers, litellm seam, BYOK keys and rate limits, and sets rules for extending or debugging it safely.
Edwardvaneechoud/Flowfile
Maps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging.
Edwardvaneechoud/Flowfile
Recreates every Flowfile development and build environment from scratch, with exact version pins and an explanation of what each Makefile target really does.
Edwardvaneechoud/Flowfile
Explains how changes to the Flowfile monorepo are gated, versioned and released, including version sync, stub and docs drift checks, Alembic migrations and pinned dependencies.
Edwardvaneechoud/Flowfile
Runbook for closing gaps between a Flowfile visual flow's results and its exported Polars or FlowFrame Python code, measured by tests rather than by eye.
Edwardvaneechoud/Flowfile
Catalog of Flowfile's environment variables and runtime flags: what each does, where the code reads it, its default, and where the docs disagree with the code.
Works with
Categories
Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom…. Flowfile Custom Node Authoring is an agent skill from Edwardvaneechoud/Flowfile.py authored with the nodedesigner SDK (from flowfile import nodedesigner as nd, CustomNodeBase, NodeSettings/Section, ColumnSelector/SingleSelect/NumericInput/…).
Flowfile Custom Node Authoring fits situations like: A task says generate/create/author/write a custom node; make a node that does X; an sklearn/xgboost/lightgbm ML node; A node with settings for ….
Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a claude-code`. Or copy the skill folder (.claude/skills/flowfile-custom-node-authoring in Edwardvaneechoud/Flowfile) into .claude/skills/flowfile-custom-node-authoring in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a codex`. Or copy the skill folder (.claude/skills/flowfile-custom-node-authoring in Edwardvaneechoud/Flowfile) into .agents/skills/flowfile-custom-node-authoring in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flowfile-custom-node-authoring, .gemini/skills/flowfile-custom-node-authoring, .github/skills/flowfile-custom-node-authoring and .opencode/skills/flowfile-custom-node-authoring in your project.
Going by SKILL.md and its folder, Flowfile Custom Node Authoring needs the command-line tools its instructions call (python and poetry). Our summary lists: Python 3; Docker.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Flowfile Custom Node Authoring is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.6k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Flowfile Custom Node Authoring: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 168 stars) and Estimate Online Covariance (microprediction/precise, 336 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Edwardvaneechoud (a GitHub user) maintains it in Edwardvaneechoud/Flowfile, which has 370 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 6, 2026.
Source: Edwardvaneechoud/Flowfile on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.