Agent skill

Flowfile Custom Node Authoring

by Edwardvaneechoud in Edwardvaneechoud/Flowfile

Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom…

MITAuto-check passedData & Analytics

Install Flowfile Custom Node Authoring

skills CLI
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Edwardvaneechoud/Flowfile flowfile-custom-node-authoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/flowfile-custom-node-authoring .claude/skills/flowfile-custom-node-authoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
flowfile-custom-node-authoring
GitHub stars
370
Token cost
~8.6k tokens
SKILL.md length
2,608 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
MIT

At a glance

Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom…

  • Works in 8 steps: Mental model — one file, one class → The generation recipe (description → node) → The authoring contract → …
  • A task says generate/create/author/write a custom node
  • SKILL.md covers When NOT to use this skill, 0. Mental model — one file,…, 1. The generation recipe… and 2. The authoring contract, plus 5 more sections
  • Calls python and poetry

What it does

Flowfile Custom Node Authoring is an agent skill from Edwardvaneechoud/Flowfile. Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom node .py authored with the nodedesigner SDK (from flowfile import nodedesigner as nd, CustomNodeBase, NodeSettings/Section, ColumnSelector/SingleSelect/NumericInput/…). Use when a task says "generate/create/author/write a custom node", "make a node that does X", "a kernel node", "an sklearn/xgboost/lightgbm ML node", "a…

Its SKILL.md is about 8.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Machine learning, Plain language and style rules and Vulnerability scanning. It works with WebAssembly and scikit-learn. The repository describes itself as: Flowfile is a visual ETL tool and Python library combining drag-and-drop workflows with Polars dataframes. Build data pipelines visually, define flows programmatically with a… The licence is MIT.

When your agent uses it

  • A task says generate/create/author/write a custom node
  • Make a node that does X
  • An sklearn/xgboost/lightgbm ML node
  • A node with settings for …

Example prompts

  • “a node that runs on the kernel and does XGBoost predictions”
  • “a node that trims whitespace”
  • “an ML clustering node”
  • “/flowfile-custom-node-authoring”

Requirements

  • Python 3
  • Docker

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Mental model — one file, one class
  2. The generation recipe (description → node)
  3. The authoring contract
  4. Local vs kernel execution
  5. Guardrails — pass the scanner and validator
  6. Worked examples (verified in-repo)
  7. Verify the generated node
  8. Provenance and re-verification

What it can do on your machine

Read from SKILL.md and the folder at commit c054c90. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • poetry

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flowfile Custom Node Authoring loads about 8.6k tokens when it runs. Until then it costs about 228 tokens; SKILL.md has 2,608 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~228
When it runs · the whole SKILL.md, loaded when a task matches
~8.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Edwardvaneechoud/Flowfile at commit c054c90, republished under its MIT licence (© Edwardvaneechoud). 2,608 words, ~8,562 tokens.

Download SKILL.mdSave it as .claude/skills/flowfile-custom-node-authoring/SKILL.md (or your agent's skills folder).
name
flowfile-custom-node-authoring
description
Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom node `.py` authored with the `node_designer` SDK (`from flowfile import node_designer as nd`, `CustomNodeBase`, `NodeSettings`/`Section`, `ColumnSelector`/`SingleSelect`/`NumericInput`/…). Use when a task says "generate/create/author/write a custom node", "make a node that does X", "a kernel node", "an sklearn/xgboost/lightgbm ML node", "a node with settings for …", when picking `environment="local"` vs `"kernel"` and `dependencies`, when reading settings inside `process()`, or when a generated node must pass the security scanner / bundle validation to install or publish. NOT for adding a built-in node type across core/frontend/frame/wasm — that is `flowfile-node-development`.

Flowfile custom-node authoring

Generate one .py file that defines a user custom node via the node_designer SDK, from a description of what the node should do. The file is the whole node: the visual Node Designer round-trips it, the worker (or a Docker kernel) executes its process(), and the community registry publishes it as-is.

When NOT to use this skill

  • Adding a built-in node type (a first-party transform wired across flowfile_core settings schema + FlowGraph + frontend registry + flowfile_frame + wasm) → flowfile-node-development. That is four registries agreeing by naming convention; this skill is one self-contained file.
  • Publishing / browsing / installing shared nodes, the community registry, the index bot, device-flow PR publishing → see the community-nodes contract in the root CLAUDE.md and flowfile_core/flowfile/community_nodes/.
  • Kernel container lifecycle, image builds, FLOWFILE_KERNEL_IMAGE_*, kernel↔core auth → flowfile-config-and-flags + flowfile_core/flowfile_core/kernel/.
  • Debugging why an existing node opens in an error state, KernelRequiredError, exec-vs-AST scan failures → flowfile-debugging-playbook.
  • Deep FlowFrame/Expr / generated-code parity → flowfile-frame-and-codegen.

All facts below verified in-repo on 2026-07-12, branch feature/add-custom-node-store. Re-verification commands are in §7 — run them before trusting a line number in a stale checkout.


0. Mental model — one file, one class

my_node.py
├─ import polars as pl
├─ from flowfile import node_designer as nd          # the ONLY flowfile import allowed
├─ class MyNodeSettings(nd.NodeSettings):  ...        # optional: the settings UI
└─ class MyNode(nd.CustomNodeBase):                   # exactly ONE such subclass per file
      node_name / metadata fields ...
      settings_schema: MyNodeSettings = MyNodeSettings()
      def process(self, *inputs: pl.LazyFrame) -> pl.LazyFrame: ...
  • Canonical import: from flowfile import node_designer as nd. The SDK lives in shared/node_designer/ and is import-pure (no flowfile_core, DB, or migrations). A node file may import only the SDK + stdlib + third-party libraries — any other flowfile/flowfile_core.* import raises a contract error in worker/dry-run contexts (shared/node_designer/loading.py:22-26,48-54).
  • Exactly one CustomNodeBase subclass per file — more raise NodeLoadError (loading.py:146-156).
  • On disk: user nodes live in ~/.flowfile/user_defined_nodes/. Community installs write a flat <id>.py there.
  • Where it runs: in-process/worker by default (environment="local"), or in an isolated Docker kernel (environment="kernel"). See §3.

1. The generation recipe (description → node)

Follow these steps in order. Each maps to a section below.

  1. Classify the transform. From the description extract: how many inputs (default 1) and outputs (default 1); which knobs the user should be able to configure (columns, numbers, choices, toggles, secrets); whether it needs heavy/third-party libraries (sklearn, xgboost, lightgbm, statsmodels).
  2. Pick the environment (§3). Pure-polars / stdlib → environment="local" (omit the field). Needs ML libs or isolation → environment="kernel" + dependencies=[...] (plain PyPI specs only). Never set execution_location — that is a core-level concept, not an SDK field.
  3. Design settings_schema (§2.3–2.5). One control per knob, grouped into Sections. Keep every kwarg a designer literal (§4.4).
  4. Write process(self, *inputs) (§2.2, §2.6). Read settings via self.settings_schema.<section>.<component>.value; .collect() when a library needs eager data; int(...)-cast NumericInput/SliderInput values; for kernel nodes import heavy libs inside process and never reference SDK types; return a LazyFrame/DataFrame, or a dict keyed by output_names for multi-output.
  5. Add metadata + examples (§2.1). node_name (required), plus node_category/title/intro; author/version/tags for publishing. Add example_inputs + example_settings (both required to install/publish; both must be literals).
  6. Self-check against guardrails (§4). No DENY constructs. Import only the SDK + third-party. One bare CustomNodeBase subclass with node_name + process. PNG icon only for published bundles.
  7. Verify (§6). Ruff-clean the file; dry-run it via the designer Test tab or dry_run.py; place it in ~/.flowfile/user_defined_nodes/.

When you have a ~/.flowfile/user_defined_nodes/ reachable and the user wants it installed, write the file there directly; otherwise write to the path the user names (or the repo flowfile-community-nodes/nodes/<id>/node.py layout when they intend to publish).


2. The authoring contract

2.1 CustomNodeBase fields (shared/node_designer/custom_node.py:401-444)

A Pydantic model — declare attributes as annotated fields with defaults: node_name: str = "My Node".

FieldTypeDefaultPurpose
node_namestrrequiredNode name; slugged to the palette item.
node_categorystr"Custom"Palette group.
node_iconstr"user-defined-icon.png"Bare icon filename (PNG for publishing — SVG is forbidden in bundles, §4.3).
settings_schemaNodeSettings | NoneNoneThe settings-UI object (a NodeSettings subclass instance).
number_of_inputsint1Input ports → inputs[0..n-1].
number_of_outputsint1Output ports.
output_nameslist[str]["main"]Keys for multi-output dict returns.
environmentLiteral["local","kernel"]"local"Execution environment (§3).
dependencieslist[str][]pip specs; auto-installed only when environment="kernel".
example_inputslist[dict[str,list]] | NoneNoneDry-run/publish sample data: one {col: [values]} per input port. Each entry must actually build a pl.LazyFrame (§2.7).
example_settingsdict[str,dict[str,Any]] | NoneNoneDry-run/publish settings: {section: {component: value}}.
titlestr | None"Custom Node"Settings-drawer title.
introstr | None"A custom node…"Settings-drawer intro.
author / version / tagsstr|None / str|None / list[str]None / None / []Publishing metadata.
node_typeLiteral["input","output","process"]"process"Node role.
transform_typeLiteral["narrow","wide","other"]"wide"Transform hint.

requires_kernel: bool and kernel_id: str|None are deprecated — they map onto environment="kernel". Use environment.

2.2 process signature (custom_node.py:732-750)
python
def process(self, *inputs: pl.LazyFrame) -> pl.LazyFrame | pl.DataFrame:
  • One pl.LazyFrame per connected input, in port order (inputs[0], inputs[1], …). .collect() inside process when you need eager data.
  • Single output: return a LazyFrame or DataFrame (the framework normalizes).
  • Multi-output (len(output_names) > 1): return dict[str, pl.LazyFrame | pl.DataFrame] keyed by the output_names.
2.3 Settings container — NodeSettings + Section
python
class MyNodeSettings(nd.NodeSettings):
    my_section: nd.Section = nd.Section(
        title="…", description="…", layout="vertical",   # or "horizontal"
        some_field=nd.TextInput(label="…"),
        another=nd.NumericInput(label="…", default=1.0),
    )

Section(title=None, description=None, hidden=False, layout="vertical", **components) — every keyword after the first four is a component that becomes a field (ui_components.py:254-317).

2.4 Control catalog (all subclass FlowfileInComponent; import as nd.<Name>)
ComponentConstructor args (defaults)
TextInputlabel, default="", placeholder=""
FilePickerlabel, default="", placeholder="", mode="open"|"create", file_types=[], allow_directory=False — path string via the standard file browser (server-side filesystem)
NumericInputlabel, default=None, min_value=None, max_value=None
SliderInputlabel, default=None, min_value=0, max_value=100, step=1
ToggleSwitchlabel, default=False, description=None
SingleSelectlabel, options (required), default=None
MultiSelectlabel, options (required), default=[]
ColumnSelectorlabel, required=False, multiple=False, data_types="ALL"
ColumnActionInputlabel, actions=[], output_name_template="{column}_{action}", show_group_by=False, show_order_by=False, data_types="ALL"
SecretSelectorlabel, options=AvailableSecrets, required=False, description=None, name_prefix=None

options= for the selects is a list of bare strings or (value, label) tuples, or a dynamic marker class passed by class (not instance):

  • nd.IncomingColumns — populate from the input frame's columns.
  • nd.AvailableArtifacts — populate from upstream artifact names.
  • nd.AvailableSecrets — the default for SecretSelector.
2.5 Column type filters — data_types= (shared/node_designer/types.py)

Groups: nd.Types.Numeric, .String, .AnyDate, .Boolean, .Binary, .Complex, .All. Specific: .Int/.Int64/…, .Float/.Float64, .Str, .Date, .Datetime, .Bool, .List, .Struct, etc. Accepts one value, a list, bare strings ("Numeric", "Int64", "int"/"float"/"str"), or a real pl.DataType. "ALL" = no filter. Both data_types=nd.Types.String and data_types=["Numeric"] are valid spellings.

2.6 Reading settings inside process()

Access: self.settings_schema.<section_attr>.<component_attr>.value. Gotchas:

  • NumericInput / SliderInput .value is a number but is not coerced — it arrives as whatever was supplied (an int or a float, e.g. example_settings with "n_clusters": 3 yields int). Cast with int(...)/float(...) when you need a specific type.
  • ColumnSelector(multiple=True).value → list[str]; multiple=False → a single column-name str.
  • SecretSelector → read .secret_value (a SecretStr; call .get_secret_value()), not .value (which is the secret name). Local-only — secrets are not available inside kernels yet (§3.3).
  • ColumnActionInput.value → a ColumnActionValue (.rows of ColumnActionRow, .group_by_columns, .order_by_column).
2.7 Dry runs — is_dry_run() and example_artifacts()

A dry run runs your real process() against example_inputs, but sandboxes its side effects. Two paths do this: the designer Test tab (user_defined/dry_run.py) and the community-node CI harness (community_nodes/dry_run_local.py, run by the registry's PR checks).

example_inputs must build a real frame. Polars infers each column's dtype from its first value, so "x1": [5, 6.5, 4.7, 6.9] infers Int64 and then rejects 6.5. Write 5.0. Bundle validation constructs every entry and reports EXAMPLES_INVALID (validation.py::_check_example_inputs_build), so this fails at validate time rather than as a polars traceback mid-dry-run.

flowfile_ctx.is_dry_run() -> bool — True only in a dry run; False in a real flow run, a single-file export, and an exported project. Branch on it to skip work that only makes sense against real data:

python
if not flowfile_ctx.is_dry_run():
    notify_downstream(result)

example_artifacts() — a node that reads an artifact another node publishes (a predict node loading a train node's model) has nothing to read in a dry run, so it fails on a missing artifact instead of exercising its real path. Override the hook to supply one:

python
def example_artifacts(self) -> dict:
    """Objects the dry-run seeds into the artifact store before process() runs."""
    from sklearn.ensemble import RandomForestClassifier   # heavy imports inside, as in process()
    ...
    return {"automl_model": {"pipeline": pipeline, "task": "classification", ...}}
  • It is a method, not a class attribute — the objects (a fitted pipeline, a tokenizer) are not AST-extractable literals, and _check_examples only requires example_inputs/example_settings to be literals.
  • Return a flat {name: obj} dict. Each entry is seeded into both the flow-local store (read_artifact) and the global store (get_global), in the CI harness and the Test tab alike.
  • Build the objects in-process. Loading a fixture from disk or the network makes the security scanner attach fs_read / network to the node's consent disclosure (§4.2), and an inlined base64 model blob over 1024 chars is a hard FF-SEC-011 deny (§4.1).
  • Only dry runs call it; a real flow run never does.

What the harness fakes. The dry-run flowfile_ctx is a faithful in-memory stand-in, not a no-op: publish_global returns a real positive id and versions on republish, get_global/read_artifact raise KeyError on a miss, publish_artifact raises ValueError on a duplicate name, and get_shared_location returns a writable temp path. Catalog APIs (read_catalog_table, list_catalogs, …) need a running server and raise a clear error — guard them with if not flowfile_ctx.is_dry_run():. Anything unmodelled falls through to a no-op returning None. A drift guard (tests/flowfile/community_nodes/test_dry_run_ctx_parity.py) fails CI if a new kernel API is added without a matching fake.

flowfile_ctx is only bound for environment="kernel" nodes. In a local node it is an undefined name at runtime.


3. Local vs kernel execution

3.1 Choosing
  • environment="local" (default): runs in the Flowfile process/worker. Full SDK + polars available; secrets resolve; no Docker. Use for pure-polars/stdlib transforms.
  • environment="kernel": runs in an isolated Docker kernel. Use when you need heavy libs (sklearn, xgboost, lightgbm, statsmodels) or isolation. Declare dependencies (plain PyPI specs) — auto-installed only for kernel nodes. A kernel node must be bound to a kernel before it runs; the user picks the kernel (and image flavour) in the UI. Unbound → KernelRequiredError (raised in flow_graph.py).
3.2 What is different inside a kernel (the real path)

A kernel node's source never runs directly. Core AST-generates a self-contained script (user_defined/kernel_codegen.py::generate_kernel_script, invoked from flow_graph.py). Consequences you must design around:

  • SDK imports are stripped; only the node's own third-party imports + polars/json/logging/sys survive. Do not reference SDK types inside process.
  • Only process + nested/helper defs are kept; all class-level attribute assignments are dropped. Put every bit of logic inside process — nothing computed at class-body scope exists at kernel runtime.
  • Settings are baked as JSON and rebuilt behind a proxy, so self.settings_schema.<section>.<component>.value still works. Every settings value must be JSON-serializable or generation fails (KernelCodegenError).
  • Inputs arrive as parquet scans (pl.scan_parquet per input); .collect() for eager work. Outputs are marshalled back via an injected flowfile_ctx.publish_output epilogue.
  • flowfile_ctx is an injected runtime global (undefined in the plain file — that's expected; linters will flag it). Surface includes log_info(...) (surfaces in the Test panel), read_inputs/read_first, publish_artifact(name, obj) / read_artifact(name) (cloudpickle-backed) for persisting a trained model across nodes, publish_global / get_global for persisting one across flows, and is_dry_run() (§2.7).
Show full SKILL.md (971 more words)Show less
3.3 Kernel gotchas
  • No Flowfile SDK inside the kernel — polars + third-party libs + flowfile_ctx only.
  • Secrets are not available in kernels yet — SecretSelector.secret_value raises there. Keep secret-using nodes local.
  • Cross-call persistence: process() is one script execution. A model trained and used within the same process() is just a local variable. To share a model between a train node and a predict node, flowfile_ctx.publish_artifact(...) / read_artifact(...).
  • Docker required.
3.4 Libraries in the kernel images (kernel_runtime/pyproject.toml + poetry.lock)
  • BASE image — polars, pyarrow, numpy, cloudpickle, joblib, deltalake (+ plumbing). No ML libs.
  • ML image — BASE + scikit-learn 1.7.2, xgboost 2.1.4, lightgbm 4.6.0, statsmodels 0.14.6, polars-ds 0.12.0 (+ pandas transitive). This is the flavour with xgboost/sklearn.
  • LITE image — base packages baked, only polars pinned so user pip installs float.

An XGBoost/sklearn kernel node works today with no code changes — provided the bound kernel uses the ML image, or the node declares the lib in dependencies for a per-kernel derived build.

3.5 ML node shape
  1. environment: str = "kernel".
  2. dependencies: list[str] = ["xgboost"] (only strictly needed if not relying on the ML image).
  3. Import heavy libs (xgboost, sklearn, …) inside process(), never at module top. Core execs the whole module at placement (registry.ensure_class) to build the palette template, and core does not have the kernel-only libs installed — a module-top import xgboost raises ModuleNotFoundError and the node fails to load. Only polars + the nd import belong at module top. In-process third-party imports are still lifted verbatim into the kernel script (this is what the real kmeans node does — §5(b)).
  4. All config in settings_schema (JSON-serializable), read via .value.
  5. .collect() inputs inside process — sklearn/xgboost need eager numpy/pandas.
  6. Self-contained train+predict → keep the model local. Train node → predict node → flowfile_ctx.publish_artifact / read_artifact.

4. Guardrails — pass the scanner and validator

The scanner (community_nodes/security_scan.py) is a conservative pre-filter; the community PR review + server-side re-scan are authoritative. Two outcomes: DENY (rejects the node) and FLAG (surfaces a capability for the consent dialog; never fails validation).

4.1 DENY — never emit these (security_scan.py:15-37)
  • eval/exec/compile; __import__/importlib.import_module/importlib.util.
  • getattr/operator.attrgetter/operator.methodcaller/globals()/vars()/__builtins__ resolving a dangerous builtin, or any of them with a non-constant name.
  • ctypes/cffi/_ctypes; os.system/os.popen*/os.exec*/os.spawn* (and posix.system/posix.popen*).
  • subprocess with shell=True or non-literal args; pty/os.forkpty; importing both socket and subprocess.
  • sys._getframe/inspect.stack/inspect.currentframe; importing pip/ensurepip.
  • Opaque high-entropy string blobs (>4096 chars, or ≥1024 chars of near-base64); bulk os.environ enumeration.

network (socket/requests/httpx/urllib/boto3/…), subprocess (literal args), fs_write, fs_read, env_read (os.getenv/sys.argv), dynamic_code (setattr/delattr with non-constant name), serialization (pickle.load/yaml.load w/o SafeLoader), secrets (SecretSelector/.secret_value). Only trigger these when the node genuinely needs them.

ML libraries do not flag — xgboost, sklearn/scikit-learn, numpy, pandas, polars match no trigger set. (Caveat: pickle.load/yaml.load without SafeLoader flag serialization; joblib.load is not detected by the scanner but still deserializes untrusted data — prefer flowfile_ctx.read_artifact for models.)

4.3 Bundle validation to install/publish (community_nodes/validation.py:385-445)
  • Folder contract: node.py (≤200 KB) + manifest.json (≤16 KB) required; optional icon.png (≤256 KB, ≤512×512), README.md, screenshots/ (≤5 PNGs). SVG forbidden anywhere. Total ≤6 MB.
  • Class shape: exactly one class inheriting exactly CustomNodeBase (no extra bases/decorators/keywords); defines node_name + process().
  • Examples required: both example_inputs (literal list) and example_settings (literal dict), else EXAMPLES_REQUIRED. Every example_inputs entry must also construct a pl.LazyFrame, else EXAMPLES_INVALID (§2.7).
  • Dependencies: non-empty dependencies requires environment="kernel"; each spec is a plain PyPI name + optional version specifier — no URLs, git+, or pip flags.
  • Identity: folder name == manifest.id == slug of node_name; versions match and strictly increase on update.
4.4 The "designer literals" rule (stay visually editable)

Keep all node attributes and settings/component kwargs as literals — constants, lists/tuples/dicts of scalars, nd.Types.*/nd.DataType.*, SDK marker classes. No f-strings, comprehensions, computed values, or **kwargs unpacking at that level, or the node degrades to code_only (still valid/installable, but the visual Form tab can't edit it). Logic inside process is unconstrained (node_designer/parsing.py::eval_literal).


5. Worked examples (verified in-repo)

5(a) Settings-driven local transform — the auto-tested docs quick start

docs/examples/custom_node.py (runs under the docs test harness).

python
import polars as pl

from flowfile import node_designer as nd


class GreetingSettings(nd.NodeSettings):
    main_config: nd.Section = nd.Section(
        title="Greeting Configuration",
        description="Configure how to greet each row",
        name_column=nd.ColumnSelector(
            label="Name Column",
            data_types=nd.Types.String,
            required=True,
        ),
        greeting=nd.SingleSelect(
            label="Greeting",
            options=[("formal", "Hello"), ("casual", "Hey")],
            default="casual",
        ),
    )


class GreetingNode(nd.CustomNodeBase):
    node_name: str = "Greeting Generator"
    node_category: str = "Text Processing"
    title: str = "Add greetings"
    intro: str = "Prefix a name column with a greeting."

    settings_schema: GreetingSettings = GreetingSettings()

    def process(self, *inputs: pl.LazyFrame) -> pl.LazyFrame:
        lf = inputs[0]
        name_col = self.settings_schema.main_config.name_column.value
        style = self.settings_schema.main_config.greeting.value
        word = "Hello" if style == "formal" else "Hey"
        return lf.with_columns(
            pl.concat_str([pl.lit(f"{word}, "), pl.col(name_col)]).alias("greeting")
        )
5(b) Kernel ML node — K-means with scikit-learn

flowfile-community-nodes/nodes/kmeans_clustering/node.py (real, published). Note environment="kernel", dependencies, the import inside process, .collect() for eager numpy, the int(...) cast, and flowfile_ctx.log_info.

python
import polars as pl
from flowfile import node_designer as nd


class KmeansClusteringSettings(nd.NodeSettings):
    clustering: nd.Section = nd.Section(
        title="Clustering",
        description="Settings for K-means cluster",
        layout="horizontal",
        feature_columns=nd.ColumnSelector(
            label="Feature Columns", required=True, multiple=True, data_types=["Numeric"],
        ),
        n_clusters=nd.NumericInput(label="Number of clusters", default=3.0, min_value=2.0, max_value=20.0),
        cluster_column=nd.TextInput(label="Cluster column name", default="cluster"),
        standardize=nd.ToggleSwitch(label="Standardize", default=True),
    )


class KmeansClustering(nd.CustomNodeBase):
    node_name: str = "kmeans clustering"
    node_category: str = "ML"
    node_icon: str = "kmeans.PNG"
    title: str = "kmeans clustering"
    intro: str = "Do a kmeans clustering with sklearn"
    author: str = "Edwardvaneechoud"
    version: str = "1.0.1"
    tags: list[str] = ["machine learning", "data-science"]
    environment: str = "kernel"
    dependencies: list[str] = ["scikit-learn"]
    example_inputs: list[dict[str, list]] = [
        {"age": [26, 45, 47], "annual_income_k": [30.3, 95.6, 101.5], "spending_score": [34, 81, 86]},
    ]
    example_settings: dict[str, dict] = {
        "clustering": {
            "feature_columns": ["age", "annual_income_k", "spending_score"],
            "n_clusters": 3, "cluster_column": "cluster", "standardize": True,
        },
    }
    settings_schema: KmeansClusteringSettings = KmeansClusteringSettings()

    def process(self, *inputs: pl.LazyFrame) -> pl.LazyFrame:
        from sklearn.cluster import KMeans
        from sklearn.preprocessing import StandardScaler

        cfg = self.settings_schema.clustering
        feature_cols: list[str] = cfg.feature_columns.value
        k = int(cfg.n_clusters.value)                 # .value is not coerced — cast to the type you need
        label_col = cfg.cluster_column.value

        df = inputs[0].collect()                      # sklearn needs eager data
        features = df.select(feature_cols).to_numpy()
        if cfg.standardize.value:
            features = StandardScaler().fit_transform(features)

        model = KMeans(n_clusters=k, n_init=10, random_state=42)
        labels = model.fit_predict(features)
        flowfile_ctx.log_info(f"Fitted {k} clusters over {df.height} rows")   # injected at runtime
        return df.with_columns(pl.Series(label_col, labels))
5(c) Kernel ML node — XGBoost predictions (the user's example, self-contained)

Trains an XGBRegressor/XGBClassifier on labeled rows and appends a prediction column. Self-contained (no cross-node artifact needed); for a real train→serve split see the note after.

python
import polars as pl
from flowfile import node_designer as nd


class XGBoostPredictSettings(nd.NodeSettings):
    model: nd.Section = nd.Section(
        title="Model",
        description="Train an XGBoost model on this data and predict a target column.",
        feature_columns=nd.ColumnSelector(
            label="Feature columns", required=True, multiple=True, data_types=["Numeric"],
        ),
        target_column=nd.ColumnSelector(
            label="Target column", required=True, data_types=["Numeric"],
        ),
        task=nd.SingleSelect(
            label="Task", options=[("regression", "Regression"), ("classification", "Classification")],
            default="regression",
        ),
        n_estimators=nd.NumericInput(label="Number of trees", default=200.0, min_value=10.0, max_value=2000.0),
        max_depth=nd.NumericInput(label="Max tree depth", default=6.0, min_value=1.0, max_value=20.0),
        prediction_column=nd.TextInput(label="Prediction column name", default="prediction"),
    )


class XGBoostPredict(nd.CustomNodeBase):
    node_name: str = "XGBoost Predict"
    node_category: str = "ML"
    title: str = "XGBoost predictions"
    intro: str = "Train an XGBoost model on the input and append predictions."
    author: str = "your-name"
    version: str = "1.0.0"
    tags: list[str] = ["machine learning", "xgboost", "prediction"]
    environment: str = "kernel"
    dependencies: list[str] = ["xgboost"]
    example_inputs: list[dict[str, list]] = [
        {"x1": [1.0, 2.0, 3.0, 4.0], "x2": [10.0, 8.0, 6.0, 4.0], "y": [12.0, 11.0, 10.0, 9.0]},
    ]
    example_settings: dict[str, dict] = {
        "model": {
            "feature_columns": ["x1", "x2"], "target_column": "y", "task": "regression",
            "n_estimators": 200, "max_depth": 6, "prediction_column": "prediction",
        },
    }
    settings_schema: XGBoostPredictSettings = XGBoostPredictSettings()

    def process(self, *inputs: pl.LazyFrame) -> pl.DataFrame:
        import xgboost as xgb                          # heavy libs import INSIDE process (core execs the module at placement)

        cfg = self.settings_schema.model
        feature_cols: list[str] = cfg.feature_columns.value
        target_col: str = cfg.target_column.value
        pred_col: str = cfg.prediction_column.value

        df = inputs[0].collect()                      # xgboost needs eager numpy
        X = df.select(feature_cols).to_numpy()
        y = df.get_column(target_col).to_numpy()

        params = dict(n_estimators=int(cfg.n_estimators.value), max_depth=int(cfg.max_depth.value))
        Model = xgb.XGBClassifier if cfg.task.value == "classification" else xgb.XGBRegressor
        model = Model(**params)
        model.fit(X, y)
        preds = model.predict(X)

        flowfile_ctx.log_info(f"Trained {cfg.task.value} on {len(feature_cols)} features, {df.height} rows")
        return df.with_columns(pl.Series(pred_col, preds))

Train → serve split (advanced). A train node fits the model and calls flowfile_ctx.publish_artifact("xgb_model", model); a separate predict node calls model = flowfile_ctx.read_artifact("xgb_model") (both kernel nodes). Use a SingleSelect/TextInput for the artifact name and nd.AvailableArtifacts to let the predict node pick from published artifacts.

5(d) Multi-output — split rows by a boolean column
python
import polars as pl
from flowfile import node_designer as nd


class SplitterSettings(nd.NodeSettings):
    main: nd.Section = nd.Section(
        title="Split",
        split_column=nd.ColumnSelector(label="Split Column", data_types="Boolean", required=True),
    )


class RowSplitter(nd.CustomNodeBase):
    node_name: str = "Row Splitter"
    node_category: str = "Custom"
    number_of_outputs: int = 2
    output_names: list[str] = ["pass", "fail"]
    example_inputs: list[dict[str, list]] = [{"keep": [True, False, True], "v": [1, 2, 3]}]
    example_settings: dict[str, dict] = {"main": {"split_column": "keep"}}
    settings_schema: SplitterSettings = SplitterSettings()

    def process(self, *inputs: pl.LazyFrame) -> dict:
        col = self.settings_schema.main.split_column.value
        return {"pass": inputs[0].filter(pl.col(col)), "fail": inputs[0].filter(~pl.col(col))}

Multi-input: set number_of_inputs: int = 2 and read inputs[0], inputs[1] — e.g. pl.concat([inputs[0], inputs[1]], how="diagonal").


6. Verify the generated node

  1. Lint: poetry run ruff check <file> (tests/generated node files are lint-exempt by config, but a clean parse matters). Confirm it imports: the file must import polars as pl and from flowfile import node_designer as nd and nothing else from flowfile.

  2. Dry-run without the app: the designer Test tab drives flowfile_core/flowfile/user_defined/dry_run.py — it runs process against example_inputs/example_settings (kernel nodes go through the same AST-generated script path). This is the fastest way to confirm the node executes.

  3. Install for real: write the file to ~/.flowfile/user_defined_nodes/<name>.py; the registry hot-reloads on save. A broken file stays visible-with-error rather than vanishing.

  4. Publish-readiness: run the bundle validator / CLI (python -m flowfile_core.flowfile.community_nodes.cli) — this is the same code the community repo's CI uses:

    bash
    python -m flowfile_core.flowfile.community_nodes.cli validate nodes/<id>
    python -m flowfile_core.flowfile.community_nodes.cli dry-run nodes/<id> --install-deps --timeout 120 --row-limit 1000

    dry-run executes process() in a child process against example_inputs, with the in-memory flowfile_ctx of §2.7. A node that reads an artifact it does not publish needs example_artifacts() to pass.

Per the repo's working agreement, do not execute the user's tests or touch his running services — write the file and describe the verification the user can run.


7. Provenance and re-verification

Facts verified 2026-07-12 on feature/add-custom-node-store. To re-verify against a fresh checkout:

bash
# SDK export surface (the exact nd.* names)
sed -n '40,73p' shared/node_designer/__init__.py

# Base-class fields, defaults, process() signature
sed -n '393,779p' shared/node_designer/custom_node.py

# Control catalog + Section
sed -n '49,553p' shared/node_designer/ui_components.py

# Type filters
cat shared/node_designer/types.py

# Kernel codegen (what survives into the kernel script) + kernel_ctx surface
sed -n '1,260p' flowfile_core/flowfile_core/flowfile/user_defined/kernel_codegen.py
cat kernel_runtime/kernel_runtime/flowfile_client.py

# Kernel image contents (which flavour has xgboost/sklearn)
sed -n '1,40p' kernel_runtime/pyproject.toml

# Scanner DENY/FLAG rules + bundle validation
sed -n '1,200p' flowfile_core/flowfile_core/flowfile/community_nodes/security_scan.py
sed -n '385,445p' flowfile_core/flowfile_core/flowfile/community_nodes/validation.py

# Real + tested examples
cat docs/examples/custom_node.py
cat flowfile-community-nodes/nodes/kmeans_clustering/node.py
ls flowfile_core/tests/flowfile/node_designer/corpus/

User-facing docs to cross-reference: docs/users/visual-editor/creating-custom-nodes.md, node-designer.md, custom-node-tutorial.md, kmeans-kernel-node.md, kernels.md, kernel-api.md, community-nodes.md.

© Edwardvaneechoud, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/flowfile-custom-node-authoring of Edwardvaneechoud/Flowfile.

Open the folder on GitHubat commit c054c90

Compare with similar skills

Flowfile Custom Node Authoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flowfile Custom Node Authoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flowfile Custom Node Authoring this skillEdwardvaneechoud/Flowfile370—~8.6kAutomated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Time Series Analytics Useropen-edge-platform/edge-ai-libraries168—~3.1kAutomated safety check: PassApache-2.0
Estimate Online Covariancemicroprediction/precise336—~535Automated safety check: PassMIT
Aeon Time Series Machine Learningdavila7/claude-code-templates32k14 repos~2.6kAutomated safety check: PassMIT

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    168 GitHub stars~3.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Estimate Online Covariance

    microprediction/precise

    Estimate a covariance / correlation / precision matrix incrementally with precise.

    336 GitHub stars~535 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Aeon Time Series Machine Learning

    davila7/claude-code-templates

    Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.

    32k GitHub starsUsed in 14 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Precise

    microprediction/precise

    Online (incremental) covariance, correlation, and precision estimation in Python — the streaming complement to sklearn.covariance.

    336 GitHub stars~782 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from Edwardvaneechoud/Flowfile

All 19 skills in this repo
  • Flowfile AI Subsystem Guide

    Edwardvaneechoud/Flowfile

    Maps the /ai/ subsystem of flowfile_core, its three agent tiers, litellm seam, BYOK keys and rate limits, and sets rules for extending or debugging it safely.

    370 GitHub stars~7k tokensUpdated today
    Auto-check: notes
  • Flowfile Architecture Contract

    Edwardvaneechoud/Flowfile

    Maps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging.

    370 GitHub stars~9.9k tokensUpdated today
    Auto-check passed
  • Flowfile Build and Environment Setup

    Edwardvaneechoud/Flowfile

    Recreates every Flowfile development and build environment from scratch, with exact version pins and an explanation of what each Makefile target really does.

    370 GitHub stars~7.3k tokensUpdated today
    Auto-check: notes
  • Flowfile Change Control

    Edwardvaneechoud/Flowfile

    Explains how changes to the Flowfile monorepo are gated, versioned and released, including version sync, stub and docs drift checks, Alembic migrations and pinned dependencies.

    370 GitHub stars~7.3k tokensUpdated today
    Auto-check passed
  • Flowfile Codegen Parity Campaign

    Edwardvaneechoud/Flowfile

    Runbook for closing gaps between a Flowfile visual flow's results and its exported Polars or FlowFrame Python code, measured by tests rather than by eye.

    370 GitHub stars~7.5k tokensUpdated today
    Auto-check passed
  • Flowfile Config and Flags Catalog

    Edwardvaneechoud/Flowfile

    Catalog of Flowfile's environment variables and runtime flags: what each does, where the code reads it, its default, and where the docs disagree with the code.

    370 GitHub stars~12k tokensUpdated today
    Auto-check: notes

Questions about Flowfile Custom Node Authoring

What does Flowfile Custom Node Authoring do?

Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom…. Flowfile Custom Node Authoring is an agent skill from Edwardvaneechoud/Flowfile.py authored with the nodedesigner SDK (from flowfile import nodedesigner as nd, CustomNodeBase, NodeSettings/Section, ColumnSelector/SingleSelect/NumericInput/…).

When should I use Flowfile Custom Node Authoring?

Flowfile Custom Node Authoring fits situations like: A task says generate/create/author/write a custom node; make a node that does X; an sklearn/xgboost/lightgbm ML node; A node with settings for ….

How do I install Flowfile Custom Node Authoring in Claude Code?

Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a claude-code`. Or copy the skill folder (.claude/skills/flowfile-custom-node-authoring in Edwardvaneechoud/Flowfile) into .claude/skills/flowfile-custom-node-authoring in your project. Claude Code loads it when a task matches its description.

How do I install Flowfile Custom Node Authoring in Codex?

Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a codex`. Or copy the skill folder (.claude/skills/flowfile-custom-node-authoring in Edwardvaneechoud/Flowfile) into .agents/skills/flowfile-custom-node-authoring in your project. Codex loads it when a task matches its description.

Can I use Flowfile Custom Node Authoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-custom-node-authoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flowfile-custom-node-authoring, .gemini/skills/flowfile-custom-node-authoring, .github/skills/flowfile-custom-node-authoring and .opencode/skills/flowfile-custom-node-authoring in your project.

What does Flowfile Custom Node Authoring need to run?

Going by SKILL.md and its folder, Flowfile Custom Node Authoring needs the command-line tools its instructions call (python and poetry). Our summary lists: Python 3; Docker.

Does Flowfile Custom Node Authoring access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Flowfile Custom Node Authoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Flowfile Custom Node Authoring use?

Flowfile Custom Node Authoring is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Flowfile Custom Node Authoring use?

About 8.6k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Flowfile Custom Node Authoring?

Skills that share tags, products or a category with Flowfile Custom Node Authoring: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 168 stars) and Estimate Online Covariance (microprediction/precise, 336 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flowfile Custom Node Authoring?

Edwardvaneechoud (a GitHub user) maintains it in Edwardvaneechoud/Flowfile, which has 370 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 6, 2026.

Source: Edwardvaneechoud/Flowfile on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.