Agent skill

Flowfile Architecture Contract

by Edwardvaneechoud in Edwardvaneechoud/Flowfile

Maps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging.

MITAuto-check passedDevelopment

Install Flowfile Architecture Contract

skills CLI
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Edwardvaneechoud/Flowfile flowfile-architecture-contract --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/flowfile-architecture-contract .claude/skills/flowfile-architecture-contract && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
flowfile-architecture-contract
GitHub stars
375
Token cost
~9.9k tokens
SKILL.md length
4,434 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
MIT

At a glance

Maps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging.

  • Works in 12 steps: The system map → Load-bearing decision: core never… → The worker offload seam and wire protocol → …
  • Onboarding to the Flowfile codebase and learning how its services fit together
  • SKILL.md covers When NOT to use this skill, 1. The system map, 2. Load-bearing decision: core… and 3. The worker offload seam and…, plus 11 more sections
  • Calls git, python and curl; needs FLOWFILE_INTERNAL_TOKEN

What it does

This skill is a map of the system and the reasons behind its cross-service contracts. It explains who talks to whom, why core never collects a LazyFrame on the hot path and ships paths and query plans instead, the worker offload wire protocol, the $ffsec$ secrets format, dual node-execution state, node-hash caching discipline and the pending directions for decentralizing execution and the catalog.

It deliberately does not teach how to add a node, run the stack, set environment flags or write tests, and it points to sibling Flowfile skills for each, including debugging, failure archaeology and change control. One detail it records is that a single SQLite catalog database is the metadata source of truth, which the worker does not open, while the scheduler uses a hand-maintained lightweight mirror in shared/models.py.

When your agent uses it

  • Onboarding to the Flowfile codebase and learning how its services fit together
  • Deciding which service a new code path belongs in
  • Debugging contract breaks around secrets, worker offload, kernel containers or the scheduler
  • Judging whether a design change would block the maintainer's decentralization plans

Example prompts

  • “Tell me which Flowfile service should own a new data export step.”
  • “Walk me through why flowfile_core never calls collect on LazyFrames.”
  • “Secrets saved by core fail to decrypt in the worker; check the $ffsec$ contract.”
  • “Review this change to node-hash caching against the architecture contract.”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. The system map
  2. Load-bearing decision: core never materializes LazyFrames
  3. The worker offload seam and wire protocol
  4. Secrets: $ffsec$1$$ — never change one side alone
  5. Scheduler: no flowfile_core imports, ever
  6. Dual node-execution state — a classic bug source
  7. Node-hash caching, the _hash save/restore dance, and parameterized runs
  8. Flows live in process memory only
  9. Kernel container contract (sandboxed user Python)
  10. Group sharing is authorization-only
  11. Known weak points (state these plainly; they are not secrets)
  12. PENDING DIRECTION — unmerged, not yet true of main

What it can do on your machine

Read from SKILL.md and the folder at commit 13aa287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • python
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FLOWFILE_INTERNAL_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flowfile Architecture Contract loads about 9.9k tokens when it runs. Until then it costs about 173 tokens; SKILL.md has 4,434 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~173
When it runs · the whole SKILL.md, loaded when a task matches
~9.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Edwardvaneechoud/Flowfile at commit 13aa287, republished under its MIT licence (© Edwardvaneechoud). 4,434 words, ~9,886 tokens.

Download SKILL.mdSave it as .claude/skills/flowfile-architecture-contract/SKILL.md (or your agent's skills folder).
name
flowfile-architecture-contract
description
The system map and load-bearing design contracts of Flowfile's core/worker/frontend/kernel/scheduler/shared architecture — who talks to whom, why core never materializes LazyFrames, the worker offload wire protocol, the $ffsec$ secrets format, dual node-execution state, node-hash caching discipline, and the pending/vision directions for execution and catalog decentralization. Use when onboarding to the codebase, deciding which service a new code path belongs in, debugging cross-service contract breaks (secrets, worker offload, kernel containers, scheduler), or evaluating whether a design change would foreclose the maintainer's decentralization vision.

Flowfile architecture contract

This skill is the map of the system and the why behind its cross-service contracts. It does not teach you how to build a node, write a test, or run the app — it teaches you where the load-bearing walls are so you don't put a door through one by accident.

When NOT to use this skill

  • Adding/changing a node type (settings class, template, add_<type> method, frontend component) → flowfile-node-development.
  • Running/building/deploying the stack, ports, Docker Compose → flowfile-build-and-env / flowfile-run-and-operate.
  • Env vars and feature flags in detail (their full table, defaults, gotchas) → flowfile-config-and-flags.
  • Writing or running tests → flowfile-testing-and-validation.
  • The AI subsystem's internals (agents, providers, BYOK, prompt log) → flowfile-ai-subsystem.
  • Frontend/Vue/Tauri conventions → flowfile-frontend-conventions.
  • flowfile_frame Python API and code generation → flowfile-frame-and-codegen.
  • Debugging a live symptom step-by-step → flowfile-debugging-playbook (this skill explains why the failure mode exists; that one walks you through diagnosing this incident).
  • Known historical bugs and their fixes as a searchable archive → flowfile-failure-archaeology.
  • Change-control process (tests, drift gates, review requirements) → flowfile-change-control — nothing here overrides it.

1. The system map

Frontend (Tauri desktop / Vite web, :8080) ──HTTP(Axios)──► flowfile_core (:63578)
  FastAPI: JWT auth, catalog, secrets, AI, DAG engine
      │ HTTP(query-plan bytes) + WebSocket(streaming)   │ Docker SDK
      ▼                                                 ▼
  flowfile_worker (:63579)                    kernel_runtime containers
    spawned subprocess per job                  uvicorn :9999 in-container,
    (dataset memory lives here,                 host-mapped to 19000-19999,
     never in the FastAPI process)              call back to core on :63578

flowfile_scheduler — polls the shared catalog DB via its own models; embedded in
  core's process when FLOWFILE_SCHEDULER_ENABLED is set, or run standalone.
shared/ — cross-package utilities (storage paths, crypto, cloud_storage, kafka,
  ml, rest_api) used by core/worker/scheduler/kernel_runtime. Import direction
  is strictly downward: shared never imports core/worker/scheduler/frame.
flowfile_frame — Python API; imports flowfile_core in-process to build FlowGraph
  objects directly. Not a separate service.
WASM frontend (Pyodide) — runs fully in-browser, no core/worker/kernel at all.

One SQLite catalog DB (flowfile_catalog.db) is the metadata source of truth, resolved via shared/storage_config.get_database_url(). Correction to root CLAUDE.md: the worker does not open this DB — verified, no engine/session in flowfile_worker/ (only shared.sql_utils for user database connection URIs). Core has the full ORM (database/models.py); scheduler and the CLI run-completion path use a lightweight declarative mirror in shared/models.py (§5) — a hand-maintained copy, not core's own tables and not SQLAlchemy reflection.


2. Load-bearing decision: core never materializes LazyFrames

The rule: flowfile_core must never call .collect() on a real dataset's pl.LazyFrame on the hot path. Core ships paths and query plans; the worker holds dataset memory, and only inside subprocesses it spawns (mp_context = get_context("spawn") in flowfile_worker/__init__.py) — never in the worker's own FastAPI process. That's why the worker exists as a separate service at all: a crashed/OOM-killed job takes down a killable child process, never the API that's holding open connections for every other user.

Enforcement status — read this before assuming a lint catches violations: it is a convention, not a mechanically enforced rule. No lint or test forbids .collect() in core. Guardrails that do exist: bounded/preview collects are corralled inside FlowDataEngine itself (collect(n) → head(n).collect(engine="streaming"), collect_schema(), pl.len() for counts); a load-bearing comment in flow_node.py's _do_execute_remote warns to use is not None, never truthiness, on a FlowDataEngine because __len__ collects; NodeTemplate.laziness metadata plus FlowNode.check_upstream_laziness / FlowGraph.check_flow_laziness (GET /editor/laziness_check) report eager nodes upstream of a catalog virtual table — but that only gates one optimization path, not a general rule. Accepted, deliberate exceptions where core does collect a full frame: writing a Delta table for the catalog locally, the explore-data/Graphic Walker preview path, seeding a brand-new manual-input datasource node, and FlowDataEngine.to_arrow()/to_pylist()/to_dict() for previews.

Violation symptom: a .collect() on a real (non-preview, non-accepted-exception) LazyFrame inside flowfile_core moves a potentially multi-GB materialization into the FastAPI process every user's request goes through — the whole point of the worker split is defeated, and a bad file hangs or OOM-kills the API server itself instead of one job's subprocess.


3. The worker offload seam and wire protocol

Config (configs/settings.py): OFFLOAD_TO_WORKER is a live-mutable MutableBool read from FLOWFILE_OFFLOAD_TO_WORKER, default on. Gotcha: the check is os.environ.get("FLOWFILE_OFFLOAD_TO_WORKER", "1") == "1" — only the exact string "1" is truthy; "true"/"yes"/"on" all evaluate false, silently disabling offload, unlike the AI feature flags which accept the full true/1/yes/on set — don't assume symmetry. WORKER_URL resolves from FLOWFILE_WORKER_URL if set, else WORKER_HOST env, else 127.0.0.1 (Windows) / 0.0.0.0 (else), port FLOWFILE_WORKER_PORT (default 63579); appends /worker when FLOWFILE_SINGLE_FILE_MODE co-hosts the worker on the core port.

Wire protocol (flow_data_engine/subprocess_operations/subprocess_operations.py): serialization is pl.LazyFrame.serialize() raw bytes — the query plan, not data. POST {WORKER_URL}/submit_query/, Content-Type: application/octet-stream, headers X-Task-Id (= the node's hash, see §7), X-Operation-Type (store | calculate_number_of_records | …), X-Flow-Id, X-Node-Id. Sampling uses POST /store_sample/ with X-Sample-Size; fuzzy join uses POST /add_fuzzy_join. Preferred transport is a WebSocket stream, falling back to REST + polling on any connect/send error. GET /status/{file_ref} returns job status; a polars-op result is a base64 serialized LazyFrame — the worker hands back a lazy scan over its own cached parquet, still not materialized data, so core stays clean reading the result back too. results_exists(hash) (GET /status/{hash}) returns False on connection error (worker down), not an exception — a known weak point, see §11.

Where core calls the seam: FlowNode._do_execute_remote builds the lazy result, ships it via ExternalDfFetcher(lf=…, file_ref=self.hash), and reads back a lazy scan plus a row count from a second fetcher call. _do_execute_local_with_sampling computes narrow transforms lazily in core and only ships a 100-row ExternalSampler job to the worker for the preview. Source nodes doing network I/O (database, Kafka, Google Analytics, REST API) never fetch in core — they use an External*Fetcher/External*Writer (add_google_analytics_reader is the canonical template): the worker does all fetching. Sinks (output, api_response, flow_output, run_flow) stay in-core even in remote mode; the output node's own function routes the actual write through ExternalOutputWriter when not local. Every fetcher-based node also keeps an in-core execution_location == "local" branch (e.g. the database reader builds a local SqlSource) for CLI --run-flow / package mode where no worker process exists — that local branch is not a violation of "worker does all fetching," it only fires when there is no worker.


4. Secrets: $ffsec$1$<user_id>$<token> — never change one side alone

Format: $ffsec$1$<user_id>$<fernet_token> — the leading $1$ is the format version, <user_id> is embedded in plaintext, <fernet_token> is the Fernet-encrypted secret. Key derivation: HKDF-SHA256(master_key, length=32, salt=KEY_DERIVATION_VERSION, info=f"user-{user_id}") → base64-urlsafe → Fernet key, KEY_DERIVATION_VERSION = b"flowfile-secrets-v1".

Why the user_id is embedded: the worker re-derives the per-user key with zero context from core — no HTTP call back, no session, nothing but the shared master key and the id parsed out of the ciphertext prefix. That's what lets a spawned worker subprocess decrypt a DB/cloud/Kafka credential completely independently.

secret_manager/secret_manager.py (core) and secrets.py (worker) are byte-for-byte parallel implementations of this format (same SECRET_FORMAT_PREFIX, same KEY_DERIVATION_VERSION). Changing the prefix, salt, or HKDF info string on one side without the other is the single easiest way to break every DB/cloud/Kafka/GA connector in the app — and it's a confusing split-brain failure: core still works (it never re-derives), every worker job throws InvalidToken.

Legacy fallback: ciphertext not starting with the prefix is a raw Fernet token decrypted with the master key directly (pre-per-user-key secrets). Group sharing is authorization-only (§10), so this format needed zero changes when sharing shipped.

Serialized plans are a transport too. Polars inlines storage_options into a LazyFrame's serialized plan, so a cloud scan built in core with decrypted keys would ship them to the worker (and into logs/caches) in plaintext. Core-built cloud scans therefore pass CloudStorageReader.get_secure_scan_kwargs(...): the credentials ride in a shared/cloud_credential_provider.py::EncryptedCredentialProvider (a Polars credential_provider whose only state is a $ffsec$ ciphertext under the node's user), decrypted by whichever process executes the plan through the decryptor it registered (core: flow_data_engine/cloud_storage_reader.py; worker: flowfile_worker/utils.py). ADLS service-principal secrets and GCS service-account keys cannot go through a Polars provider and still travel in the options.


5. Scheduler: no flowfile_core imports, ever

flowfile_scheduler is a separate, dependency-light package. Verified: grep -rn "flowfile_core|flowfile_worker|flowfile_frame" flowfile_scheduler/flowfile_scheduler/ matches only two docstring sentences declaring the rule — zero actual imports. The scheduler's models.py is a backward-compat shim re-exporting from shared/models.py — independent declarative SQLAlchemy models on their own Base, a minimal hand-maintained mirror of flowfile_core/database/models.py, not SQLAlchemy table reflection. (Root CLAUDE.md's "polls the shared SQLite DB via reflected tables" is loose phrasing — treat "reflected" as informal, not literal Table(..., autoload_with=engine) reflection.) Dependency arrow is core → scheduler, never reverse.

Why it matters: a scheduler feature needing a new column means editing shared/models.py to keep the mirror in sync, and a real Alembic migration against flowfile_core/database/models.py (the canonical schema) — touching only one side either breaks the scheduler's queries or leaves core's migration history incomplete.

Worth knowing: single-leader via a SchedulerLock row (90s stale threshold); cron uses a naive local wall-clock cursor advanced to "now" (not the missed slot) on catch-up so a downed scheduler fires exactly once rather than backfilling; runs launch as a detached subprocess (python -m flowfile run flow <path> --run-id <id>), not an in-process call — this is why the scheduler can be embedded in core without core's own crash taking a running scheduled flow down with it.


6. Dual node-execution state — a classic bug source

Every FlowNode carries two parallel state representations:

  • node_stats: NodeStepStats — legacy. Still read by needs_run and get_predicted_resulting_data (schema prediction).
  • _execution_state: NodeExecutionState — new. Governs actual run/skip decisions via NodeExecutor._decide_execution.

They are synced one-way: NodeExecutor._sync_state_to_legacy copies the new state into the legacy one after every decision/run. There is no path that updates only node_stats and expects _execution_state to follow — if you add a new run-affecting flag, write it to _execution_state and let the sync propagate it; writing only node_stats is invisible to _decide_execution and writing only _execution_state without going through _sync_state_to_legacy leaves schema prediction (needs_run) looking at stale data. Symptom of getting this wrong: the UI's "needs run" badge and the actual run/skip behavior disagree — one lags the other by exactly one sync call.

NodeExecutor._decide_execution is the single source of truth for both whether to run and how (executor.py). Order matters — evaluated top to bottom, first match wins:

#ConditionResultReason tag
1node_template.node_group == "output"RUNOUTPUT_NODE — sinks always run
2node_type == "run_flow"RUNsubflow file can change with no parent-settings change, never "up to date"
3force_refresh (reset_cache)RUNFORCED_REFRESH
4cache_results=True, results_exists(hash)SKIPcache hit — checked before performance mode so a cached result survives even when upstream had no new data
4bcache_results=True, not results_existsRUNCACHE_MISSING
5performance_modeRUNPERFORMANCE_MODE
6not has_run_with_current_setupRUNNEVER_RAN
7read node, source file changed on diskRUNSOURCE_FILE_CHANGED
8elseSKIPresults already in memory from a previous run

Strategy (_determine_strategy, once RUN is decided): run_location == "local" → FULL_LOCAL; else cache_results → REMOTE (caching needs a fully materialized result); else transform_type == "narrow" → LOCAL_WITH_SAMPLING (compute lazily in core, ship only a 100-row sample job); else → REMOTE.


7. Node-hash caching, the _hash save/restore dance, and parameterized runs

FlowNode.calculate_hash folds: input-node hashes + hash(setting_input) + parent_uuid (a uuid1 stamped once per FlowGraph instance at construction) + _cache_epoch. The hash is the worker's cache key (file_ref in the wire protocol above) — results_exists(hash) and get_external_df_result(hash) both key off it.

Because parent_uuid is per-graph-instance, worker cache keys never survive a graph reload — restarting core or reopening a flow always recomputes from scratch even if nothing changed. This follows directly from §8's in-memory-only invariant but is easy to forget if you expect cache hits across sessions.

The _hash save/restore discipline (FlowGraph._execute_single_node): when a flow has ${param} placeholders, execution substitutes resolved parameter values into setting_input in place before running the node, then restores the original ${...} text afterward — and explicitly saves and restores node._hash around that substitution. Why: mutating setting_input in place, even temporarily, changes what calculate_hash would compute; without the explicit restore, the next real settings write after a parameterized run sees a hash mismatch, treats it as an external change, calls reset(), and silently loses example_data_generator / has_completed_last_run — a real bug this exact comment documents having fixed once. If you touch parameter substitution, preserve this save/restore or you reintroduce that regression.


8. Flows live in process memory only

FlowfileHandler keeps every open flow in an in-memory dict (process memory, not the DB — the DB only stores catalog registrations of flows, not live graph state). A core restart, or a "Save As" that changes the in-memory flow's identity, invalidates the frontend's flow_id. The run route has an explicit guard for this: hitting it with a stale id returns a targeted 404 telling the user to reload the flow, rather than a generic crash. If you add a new endpoint keyed by flow_id, expect the same failure mode and handle it the same way — don't assume a flow_id a client sends is still resolvable.


9. Kernel container contract (sandboxed user Python)

Each kernel is a Docker container running uvicorn kernel_runtime.main:app --host 0.0.0.0 --port 9999 inside the container (EXPOSE 9999, healthcheck curl -f http://localhost:9999/health). Local (electron/ package) topology: core maps that container port to a host port from 19000-19999 (_BASE_PORT = 19000, _PORT_RANGE = 1000); kernel URL is http://localhost:{allocated_port}. Docker-in-Docker (compose) topology: no host port mapping — kernels are reached by container name (http://flowfile-kernel-{id}:9999) on the shared Docker network, because core discovers the named volume covering the shared path and mounts the same volume at the same path in the kernel container, so paths line up identically across core/worker/kernel without translation.

Kernel → core auth: every kernel gets FLOWFILE_INTERNAL_TOKEN injected (via auth.jwt.get_internal_token()) plus FLOWFILE_CORE_URL; the kernel sends the token per-request, core verifies with secrets.compare_digest. Electron auto-generates the token if unset; docker mode hard-errors on a missing token — deliberate, electron is single-user and docker is multi-tenant. A valid token with no kernel id becomes a synthetic _internal_service principal that bypasses catalog access restriction (§10) but must never be treated as a real user for ownership — sharing code keys the synthetic check on username == "_internal_service", never on id, because that principal's id defaults to 1, a real user id.

Known asymmetry (gotcha): storage.shared_directory and the derived artifact/global-artifact dirs honor FLOWFILE_SHARED_DIR if set, but production's get_kernel_manager() hardcodes storage.temp_directory / "kernel_shared" and ignores that env var — setting it to a non-default path in a real deployment splits artifact staging away from the volume kernels actually mount. Tests dodge this by constructing KernelManager with the path passed explicitly.


10. Group sharing is authorization-only

Group-based resource sharing (12 shareable resource types: secrets, DB/cloud/Kafka/GA connections, catalog namespaces/tables/notebooks, flows, visualizations, dashboards, global artifacts) is a pure authorization layer bolted on top of existing ownership. The load-bearing property: a shared secret's ciphertext stays encrypted under the owner's user_id forever — sharing never re-encrypts it under the grantee's key. That's why flowfile_worker (§4) and flowfile_scheduler (§5) needed zero code changes to support sharing: a group-granted secret decrypts through the exact same owner-keyed-derivation path as an owned one.

Consequences of "authorization-only, never re-key": rotating a secret on a shared connection re-encrypts it under the owner's id, never the manage-grantee's who triggered the rotation. Changing a shared connection's target fields (host/endpoint/protocol) while it carries a bundled secret and no new credentials were supplied is rejected with 422 — an anti-repoint-harvest guard against a manage-grantee silently repointing the connection at a server they control to harvest the owner's credentials on the next connect. The catalog is private-by-default in docker mode (electron/package stay fully open); resolution is own-first (own rows, then group-granted, lowest row id wins on name collisions). /user-groups and /shares 404 in electron mode (require_sharing_enabled dependency, routes/user_groups.py), same status as a genuinely missing resource — no enumeration oracle. Every resource-delete path must call sharing.delete_grants_for_resource(...): SQLite reuses row ids, so a leftover grant on a deleted resource's old id would silently reattach to whatever unrelated resource claims that id next. An ORM after_delete backstop covers single-row deletes, but bulk query.delete() bypasses ORM events and must call the cleanup explicitly.


11. Known weak points (state these plainly; they are not secrets)

Weak pointWhat it means in practice
flow_graph.py god file (largest module in the repo)Holds the DAG engine, the add_* node builders, catalog Delta write helpers, ML train/apply plumbing, kernel execution, YAML serialization, groups, layout, history, and codegen entry all in one file. Navigate it with grep -n "def add_" flowfile_core/flowfile_core/flowfile/flow_graph.py — that's the practical index; don't try to read it top to bottom.
Skip-list shallownessThe pre-run skip pass (util/node_skipper.py) only expands one transitive level of "leads to" from incorrectly-configured nodes. Correctness for deeper chains relies on _execute_stages re-checking dependents of every failed/skipped node at run time — a genuine belt-and-suspenders design, not a hole, but don't assume the skip pre-pass alone is a complete skip set.
results_exists swallows worker downtimeReturns False on an HTTP connection error, same as a genuine cache miss. In Development mode this silently degrades "should skip, cached" decisions into full re-runs whenever the worker is briefly unreachable — no error surfaces, just unexpectedly slower runs.
HTTP 419 on add-node failurePOST /update_settings/ raises the non-standard status code 419 (not 422/500) when the node's add_<type> function itself raises — handle it explicitly in frontend/AI tool wrappers. Full dispatch-trap mechanics: flowfile-node-development §1.6.
FLOWFILE_OFFLOAD_TO_WORKER truthy parsingOnly the literal string "1" enables offload; "true" silently disables it. See §3.

Show full SKILL.md (1,803 more words)Show less

12. PENDING DIRECTION — unmerged, not yet true of main

As of 2026-07-03 there is an unmerged branch, claude/core-abstraction-flowgraph-001hrn, carrying 7 commits (5 substantive

  • a docs commit + a merge of main — matching flowfile-failure-archaeology's branch inventory) that begin decomposing the seams in §2, §3, and §6. None of this is on main or any release — treat it as a candidate design, not current behavior, until it lands:
  • WorkerTransport (flowfile/execution/transport.py) — proposed sole owner of worker URLs/HTTP/WS, with typed exceptions instead of §3's ad-hoc error-code scheme.
  • ExecutionBackend ABC with LocalBackend/RemoteWorkerBackend (flowfile/execution/backends/) — proposes replacing inline if execution_location == "local" branches (§3) with one backend.run_lazyframe/sample/count_records call per node. Adds a ratchet test pinning the count of remaining inline branches in flow_graph.py, failing if it goes up — worth reusing regardless of whether this branch lands.
  • NodeSpec registry (flowfile/node_registry/) — proposed single source of truth for built-in node types, unifying four independently-maintained catalogs (get_all_standard_nodes(), NODE_TYPE_TO_SETTINGS_CLASS, nodes_with_defaults, ai/tools/classification._NODE_CLASS_MAP — see flowfile-node-development for why keeping these in sync today is a known parity hazard) as derived views over one NodeSpec per type.

If asked to work near these seams, check whether the branch has merged (git log --oneline origin/main..origin/claude/core-abstraction-flowgraph-001hrn — empty means merged or superseded) before assuming either the old inline-branch shape or the new backend/registry shape is current truth.

Notebook kernel: canvas fallback everywhere, no database copy (planned 2026-10-02, landed 2026-10-03)

Agreed direction for the notebook kernel (flowfile_frame/notebook_kernel.py, flowfile_core/notebook/kernel_runner.py), built in phases 0–4a below. The kernel is a separate execution mode: node settings stay the contract that round-trips to the canvas on Push, the kernel runs natively what it can, and everything else goes to the canvas. The end state, reached with phase 3, is a kernel that opens no catalog database connection: the per-kernel SQLite copy (kernel/notebook_db.py, there because WAL did not cross the Docker Desktop VM) and the SQLite requirement of the gate (notebook/gate.py) are gone. Do not add SCD2, change-feed or SQL logic to kernel_runtime/flowfile_client.py, forward raw SQL to core, bring a database back into the kernel, or mount a host folder into it.

Phase 0 landed with the plan: a lineage run that holds no output node passes run_graph(commit_sources=False) (kernel_runner.lineage_commits), and a session takes the schemas of canvas nodes that ran (_Session.refresh).

Phase 1 landed 2026-10-03: held nodes run in core from their settings. Where the kernel's _Session._resolve finds a gate or deferred node without a canvas twin (or a read of a file it cannot see), it hands core the node as FlowfileData lists it plus its inputs as parquet under the session's results folder: notebook/held_run.py, POST /notebook/session/node_run, bound like node_result. Core builds the inputs as flow_input nodes fed those files (not add_dependency_on_polars_lazy_frame: a NodePromise node is never is_correct, so the planner skips it) plus the one node through populate_graph_from_flow_information, on a FlowGraph that lives for the call (_system_run, Performance mode, the canvas flow's source_registration_id and the session's parameters), runs it under KernelHold with commit_sources=False, answers every live output's path (a gate's dead handles as closed) or, for schema_only, its columns, and releases the graph's FlowLogger, log file and exchange folder. Closed allowlist HELD_NODE_TYPES plus installed non-output custom nodes; writers and other types still say "Push". A cell-built deferred node whose seed has no columns asks core for them at build (NotebookMode.schema_resolver, native.resolved_seed). A Python Script on another kernel runs this way too, with its artifacts on the call's graph only.

Phase 2 landed 2026-10-03: metadata lookups go through core. One hook, notebook/lookup.py::metadata_lookup (a ContextVar, the placement_check pattern), consulted by flow_graph._resolve_catalog_table_info / _resolve_catalog_sql_tables, storage_backend.resolve_for_namespace, subflow._registration (behind stamp_flow_reference and resolve_subflow_path), prechecks.placement_refusal and the frame's flowfile_frame/_metadata.py (behind catalog_reference.py, run_flow.py, kernels.py and the connection listings; every public function asks the hook when set, else runs its local_* twin on the database). The kernel installs a _metadata.SessionLookup for every op (_metadata.installed), which posts to POST /notebook/session/lookup, a closed lookup.KINDS each answered by the function the call site runs without the hook, as the kernel's owner, bound by kernel_runner._bound_kernel (the flow need not be open), metadata only: no plan, no storage credential, no password, no ciphertext. Secret-bearing nodes are held in a kernel session and take phase 1's path: every cloud_storage_reader (NOTEBOOK_DEFERRED_NODE_TYPES), a catalog reader of a cloud-backed table (_metadata.is_cloud_table, from the catalog_table answer), a custom node with a SecretSelector (custom_node._selects_secret); native._predicts_in_core keeps CONNECTION_SOURCE_TYPES from running their schema callback while a schema_resolver is set. The census is tests/notebook/test_kernel_database_census.py, at zero: a pool checkout listener under the sim's IN_KERNEL_OP marker over the whole corpus (conftest.kernel_db_opens), plus every kind answered without a $ffsec$ or a decrypt.

Phase 3 landed 2026-10-03: the copy is gone. A notebook kernel holds no catalog database: _notebook_env sets no FLOWFILE_DB_PATH (the kernel's get_database_url() falls to the never-mounted <storage>/database/), and every notebook_kernel.handle() first installs _refuse_database, a do_connect listener on the one cached engine behind connection.engine, SessionLocal and get_db_context (gated on FLOWFILE_KERNEL_ID, which only a kernel container has), so a cell that opens the database itself stops with NO_DATABASE, naming where it was opened (_opened_from: a cell, else the frame module of a regression) and the ff functions to use, before pysqlite could create a file. Deleted: kernel/notebook_db.py, POST /notebook/session/database, kernel_runner.refresh_database, _rearm and the refresh listeners, the hello schema-head handshake (migration.package_head, kernel_runner._schema_revision; the flowfile version check stays) and the gate's SQLite requirement (kernel_sessions_allowed is now not sharing_enabled(), which admits an electron app on a Postgres catalog without proving it). An uncached shared.database.create_catalog_engine() a cell calls by hand is not refused and creates an empty file in the container layer, as in any kernel with flowfile installed. Proof: test_kernel_database_census.py (no op connects), test_kernel_session.py::test_a_notebook_kernel_refuses_to_open_a_catalog_database (scratch engine) and the real-kernel test_kernel_notebook_docker.py (the default database file never appears in the container; a get_db_context() cell is refused; a namespace core creates shows in the next cell), which .github/workflows/test-notebook-kernel.yml now runs on every frame, notebook or kernel change (FLOWFILE_REQUIRE_NOTEBOOK_KERNEL makes a skip a failure).

Phase 4a landed 2026-10-03: no host folder is mounted. A notebook kernel's container mounts what every kernel mounts, /shared and /catalog_tables, and nothing else: not the Flowfile folders (flows, custom nodes and their mount directories, catalog tables) and not a folder its owner lists. Deleted: kernel/notebook_mounts.py, mounted_folders (models.MountedFolder, the kernel form's folders field, migration 034's column, dropped by 035; a pydantic model ignores it from an older client), FLOWFILE_NOTEBOOK_MOUNTS, notebook.kernel_path, input_schema.kernel_file_path, KernelManager.host_folders, the /host/<drive>/ Windows scheme, the key-store tmpfs, FLOWFILE_STORAGE_DIR in _notebook_env (the kernel's storage is its own ~/.flowfile, /root/.flowfile in the image) and _metadata.is_cloud_table / SessionLookup.cloud_tables. What replaced each use: a file a cell names is never opened in the kernel (notebook.paths_as_written sets only keep_paths_as_written, under which native._kernel_hidden_path hides every local read/list_files path, ${param} ones included) and core reads it (the canvas twin's rows, else a held run); whether a bare path is a folder is the is_directory lookup (flow_frame_methods._resolve_scan_mode); every catalog reader in a kernel session is deferred (native.notebook_defers) and its schema and rows come from core (native._predicts_in_core, CORE_PREDICTED_TYPES = CONNECTION_SOURCE_TYPES | {"catalog_reader", "run_flow"}); a run_flow reference's interface is the flow_interface lookup (_metadata.flow_interface, core-side resolve_subflow_path + get_subflow_interface); installed custom node files come through the custom_node_sources lookup, which notebook_kernel._mirror_custom_nodes writes as <node_key>.py into the kernel's own nodes folder before open/reset/execute/clean_run (it removes only files it wrote and rescans the registry once when anything changed; the kernel still runs a local node's process() itself); a push stores paths as written and notebook/validate.host_file_paths only recomputes abs_file_path on the host. Invariant: a kernel mounts exactly /shared and /catalog_tables; a file path in a cell is a path on the user's machine that only core opens. Plain Polars or open() on a host path in a cell finds no file, and a cell writes no file (an ff writer adds a writer node). Known cost: a catalog table a cell reads is handed over whole as parquet through /shared, even for display(). Proof: tests/kernel/test_notebook_kernel_env.py (exactly the two binds, no mounts or tmpfs, none of FLOWFILE_STORAGE_DIR, FLOWFILE_NOTEBOOK_MOUNTS, FLOWFILE_DB_PATH; an old client's mounted_folders ignored), the sim tests in tests/notebook/test_kernel_canvas_rows.py (a read is a held read run, a bare folder asks is_directory) and test_kernel_database_census.py (a catalog table, a flow file and a custom node reach the kernel through core; the three new kinds answer metadata only), and the real-kernel test_kernel_notebook_docker.py::test_an_installed_custom_node_is_mirrored_into_the_kernel (with test_a_file_a_cell_names_is_read_by_core, where pl.read_csv on the same path has no such file).

Remaining, each with its own plan: a Python Script on the notebook's own kernel (KernelHold refuses it during a fallback), a push that keeps the session's variables, a row-limited fallback for display() (which now also bounds a catalog table a cell reads, until then handed over whole), Postgres catalogs and docker-mode sessions (per-user access in the lookups, admin gating on node_run), session artifacts visible to the canvas.

Known gaps this leaves until then: a seeded canvas node below one that just ran keeps its seeded columns until it runs or the session is reset, and FlowFrame.collect() on a deferred frame in a script still commits sources (flow_frame.py calls run_graph(node_ids=...) with the default).


13. VISION (maintainer direction, 2026-07-03) — not implemented, do not foreclose

The maintainer's stated direction: decentralized execution on a centralized catalog. Global/shared components (catalog schema, definitions, sharing model), catalog metadata in PostgreSQL, catalog data in object storage (S3), and independent local Flowfile instances that each bring their own compute and mix local files with catalog-hosted remote files in the same flow.

This is a direction, not a spec, and not implemented anywhere in the repo today. Nothing above contradicts it — but two current seams are exactly what a future implementation would generalize, so do not add code that tightens their coupling to "local + SQLite" further:

  • shared/storage_config.get_database_url() (§1) is the only function deciding where the catalog metadata DB lives — a local SQLite file today (FLOWFILE_DB_PATH env, else TESTING temp path, else <base>/database/flowfile_catalog.db), no PostgreSQL branch yet. Route new catalog code through this seam (or storage/CatalogService above it) rather than hard-coding a SQLite path assumption inline.
  • catalog/storage_backend.py already demonstrates the pattern the vision needs for data: storage resolves per catalog (a level-0 namespace carries storage_uri + storage_connection_name; unset ⇒ local filesystem, set ⇒ object storage, credentials resolved as the catalog owner). This supports "data in S3" already; it does not yet support "metadata in Postgres" or "multiple compute instances sharing one catalog." Extend along this per-namespace-target pattern, not a single global env-var switch — the module already demoted the old FLOWFILE_CATALOG_STORAGE_URI/_CONNECTION env vars from "the mechanism" to "just a creation-time default," which is this vision's direction.

If asked to design a feature touching catalog storage, catalog identity, or cross-instance execution, treat "would this make a future Postgres-metadata/multi-instance-compute design harder" as a review question — not a blocker, but flag it rather than silently baking in a local-SQLite assumption.


Provenance and maintenance

All facts above were verified by reading source or running the listed read-only command against the repo at commit f6963c77 (branch feature/claude-skills), app version 0.12.7, dated 2026-07-03. Every volatile fact below should be re-checked with its command if this skill feels stale — architecture facts like these drift slower than config defaults, but line numbers and the pending-branch status will drift fastest.

bash
# system map / "worker doesn't touch the catalog DB" claim
grep -rn "create_engine|Session|shared.models" flowfile_worker/flowfile_worker/ | grep -v __pycache__

# core-never-collects: god-file sizing
wc -l flowfile_core/flowfile_core/flowfile/flow_graph.py \
      flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py \
      flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py

# worker offload truthy parsing ("1" only)
grep -n "FLOWFILE_OFFLOAD_TO_WORKER" flowfile_core/flowfile_core/configs/settings.py

# worker wire-protocol endpoints
grep -n "@router" flowfile_worker/flowfile_worker/routes.py

# secrets format parity between core and worker
grep -n "SECRET_FORMAT_PREFIX\|KEY_DERIVATION_VERSION" \
  flowfile_core/flowfile_core/secret_manager/secret_manager.py \
  flowfile_worker/flowfile_worker/secrets.py

# scheduler zero-core-imports invariant
grep -rn "flowfile_core|flowfile_worker|flowfile_frame" flowfile_scheduler/flowfile_scheduler/ | grep -v __pycache__

# NodeExecutor decision table (still ordered the same way?)
sed -n '/_decide_execution/,/_determine_strategy/p' flowfile_core/flowfile_core/flowfile/flow_node/executor.py

# kernel port range + internal token wiring
grep -n "_BASE_PORT\|_PORT_RANGE" flowfile_core/flowfile_core/kernel/manager.py
grep -n "get_internal_token\|verify_internal_token" flowfile_core/flowfile_core/auth/jwt.py

# group sharing 404-in-electron gate
grep -n "require_sharing_enabled" flowfile_core/flowfile_core/routes/user_groups.py

# HTTP 419 on add-node failure
grep -n "419" flowfile_core/flowfile_core/routes/routes.py

# catalog storage per-namespace resolution (vision seam)
sed -n '1,40p' flowfile_core/flowfile_core/catalog/storage_backend.py

# pending-branch status: has it merged / is it still distinct?
git log --oneline origin/main..origin/claude/core-abstraction-flowgraph-001hrn 2>/dev/null \
  || git log --oneline f6963c77..claude/core-abstraction-flowgraph-001hrn

# app version
grep -n '^version' pyproject.toml

© Edwardvaneechoud, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/flowfile-architecture-contract of Edwardvaneechoud/Flowfile.

Open the folder on GitHubat commit 13aa287

Compare with similar skills

Flowfile Architecture Contract next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flowfile Architecture Contract compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flowfile Architecture Contract this skillEdwardvaneechoud/Flowfile375—~9.9kAutomated safety check: PassMIT
Burla Internals Deep DiveBurla-Cloud/burla263—~2.4kAutomated safety check: PassCustom licence
Code Review Skillawesome-skills/code-review-skill2.1k—~2.8kAutomated safety check: NotesMIT
GitDiagram Repository Overviewahmedkhaleel2004/gitdiagram18k—~427Automated safety check: PassMIT
Deepwiki Rssopaco/deepwiki-rs3.1k—~748Automated safety check: PassMIT
Code Graph Mermaid Diagramstrailofbits/skills7.4k—~1.7kAutomated safety check: PassCC-BY-SA-4.0

Similar skills

  • Burla Internals Deep Dive

    Burla-Cloud/burla

    Reference for Burla internals: how a remote_parallel_map job flows between services, how clusters and nodes are managed, and where the head keeps its state.

    263 GitHub stars~2.4k tokensUpdated 16 days ago
    DevelopmentAuto-check passed
  • Code Review Skill

    awesome-skills/code-review-skill

    Provides comprehensive code review guidance for React 19, Vue 3, Angular 17+, Svelte 5, Rust, TypeScript, Java, Java 8, PHP, Ruby, Rails, Python, Django, FastAPI, Go, C/.NET, Kotlin, Swift, Dart…

    2.1k GitHub stars~2.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • GitDiagram Repository Overview

    ahmedkhaleel2004/gitdiagram

    Explains the architecture of a public GitHub repository through GitDiagram: how the code is organized, the main components with paths, and a Mermaid diagram.

    18k GitHub stars~427 tokensUpdated today
    DevelopmentAuto-check passed
  • Deepwiki Rs

    sopaco/deepwiki-rs

    AI-powered Rust documentation generation engine for comprehensive codebase analysis, C4 architecture diagrams, and automated technical documentation.

    3.1k GitHub stars~748 tokensUpdated 25 days ago
    DevelopmentAuto-check passed
  • Code Graph Mermaid Diagrams

    trailofbits/skills

    Official

    Generates Mermaid diagrams from Trailmark code graphs, including call graphs, class hierarchies, module dependency maps, complexity heatmaps and attack surface data flows.

    7.4k GitHub stars~1.7k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • GitDiagram Repo Architecture

    ahmedkhaleel2004/gitdiagram

    Explains how a public GitHub repository is built by fetching its GitDiagram architecture diagram, components and optional explainer video.

    18k GitHub stars~429 tokensUpdated today
    DevelopmentAuto-check passed

More from Edwardvaneechoud/Flowfile

All 19 skills in this repo
  • Flowfile AI Subsystem Guide

    Edwardvaneechoud/Flowfile

    Maps the /ai/ subsystem of flowfile_core, its three agent tiers, litellm seam, BYOK keys and rate limits, and sets rules for extending or debugging it safely.

    375 GitHub stars~7k tokensUpdated today
    Auto-check: notes
  • Flowfile Build and Environment Setup

    Edwardvaneechoud/Flowfile

    Recreates every Flowfile development and build environment from scratch, with exact version pins and an explanation of what each Makefile target really does.

    375 GitHub stars~7.3k tokensUpdated today
    Auto-check: notes
  • Flowfile Change Control

    Edwardvaneechoud/Flowfile

    Explains how changes to the Flowfile monorepo are gated, versioned and released, including version sync, stub and docs drift checks, Alembic migrations and pinned dependencies.

    375 GitHub stars~7.3k tokensUpdated today
    Auto-check passed
  • Flowfile Codegen Parity Campaign

    Edwardvaneechoud/Flowfile

    Runbook for closing gaps between a Flowfile visual flow's results and its exported Polars or FlowFrame Python code, measured by tests rather than by eye.

    375 GitHub stars~7.5k tokensUpdated today
    Auto-check passed
  • Flowfile Config and Flags Catalog

    Edwardvaneechoud/Flowfile

    Catalog of Flowfile's environment variables and runtime flags: what each does, where the code reads it, its default, and where the docs disagree with the code.

    375 GitHub stars~12k tokensUpdated today
    Auto-check: notes
  • Flowfile Custom Node Authoring

    Edwardvaneechoud/Flowfile

    Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom…

    375 GitHub stars~8.6k tokensUpdated today
    Auto-check passed

Questions about Flowfile Architecture Contract

What does Flowfile Architecture Contract do?

Maps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging. This skill is a map of the system and the reasons behind its cross-service contracts. It explains who talks to whom, why core never collects a LazyFrame on the hot path and ships paths and query plans instead, the worker offload wire protocol, the $ffsec$ secrets format, dual node-execution state, node-hash caching discipline and the pending directions for decentralizing execution and the catalog.

When should I use Flowfile Architecture Contract?

Flowfile Architecture Contract fits situations like: onboarding to the Flowfile codebase and learning how its services fit together; deciding which service a new code path belongs in; debugging contract breaks around secrets, worker offload, kernel containers or the scheduler; judging whether a design change would block the maintainer's decentralization plans.

How do I install Flowfile Architecture Contract in Claude Code?

Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a claude-code`. Or copy the skill folder (.claude/skills/flowfile-architecture-contract in Edwardvaneechoud/Flowfile) into .claude/skills/flowfile-architecture-contract in your project. Claude Code loads it when a task matches its description.

How do I install Flowfile Architecture Contract in Codex?

Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a codex`. Or copy the skill folder (.claude/skills/flowfile-architecture-contract in Edwardvaneechoud/Flowfile) into .agents/skills/flowfile-architecture-contract in your project. Codex loads it when a task matches its description.

Can I use Flowfile Architecture Contract in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flowfile-architecture-contract, .gemini/skills/flowfile-architecture-contract, .github/skills/flowfile-architecture-contract and .opencode/skills/flowfile-architecture-contract in your project.

What does Flowfile Architecture Contract need to run?

Going by SKILL.md and its folder, Flowfile Architecture Contract needs the command-line tools its instructions call (git, python and curl) and credentials named FLOWFILE_INTERNAL_TOKEN.

Does Flowfile Architecture Contract access the network?

SKILL.md contains no URLs. Its commands use git and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Flowfile Architecture Contract safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Flowfile Architecture Contract use?

Flowfile Architecture Contract is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Flowfile Architecture Contract use?

About 9.9k tokens (SKILL.md is roughly 40k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Flowfile Architecture Contract?

Skills that share tags, products or a category with Flowfile Architecture Contract: Burla Internals Deep Dive (Burla-Cloud/burla, 263 stars), Code Review Skill (awesome-skills/code-review-skill, 2.1k stars), GitDiagram Repository Overview (ahmedkhaleel2004/gitdiagram, 18k stars) and Deepwiki Rs (sopaco/deepwiki-rs, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flowfile Architecture Contract?

Edwardvaneechoud (a GitHub user) maintains it in Edwardvaneechoud/Flowfile, which has 375 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.

Source: Edwardvaneechoud/Flowfile on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.