Burla Internals Deep Dive
Burla-Cloud/burla
Reference for Burla internals: how a remote_parallel_map job flows between services, how clusters and nodes are managed, and where the head keeps its state.
Maps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-architecture-contract --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/flowfile-architecture-contract .claude/skills/flowfile-architecture-contract && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "flowfile-architecture-contract" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-architecture-contract into .claude/skills/flowfile-architecture-contract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-architecture-contract", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-architecture-contractType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-architecture-contract --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/flowfile-architecture-contract .agents/skills/flowfile-architecture-contract && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "flowfile-architecture-contract" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-architecture-contract into .agents/skills/flowfile-architecture-contract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-architecture-contract", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-architecture-contract --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/flowfile-architecture-contract .cursor/skills/flowfile-architecture-contract && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "flowfile-architecture-contract" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-architecture-contract into .cursor/skills/flowfile-architecture-contract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-architecture-contract", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Edwardvaneechoud/Flowfile.git --path .claude/skills/flowfile-architecture-contract--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-architecture-contract --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/flowfile-architecture-contract .gemini/skills/flowfile-architecture-contract && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "flowfile-architecture-contract" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-architecture-contract into .gemini/skills/flowfile-architecture-contract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-architecture-contract", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-architecture-contractInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/flowfile-architecture-contract .github/skills/flowfile-architecture-contract && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "flowfile-architecture-contract" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-architecture-contract into .github/skills/flowfile-architecture-contract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-architecture-contract", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-architecture-contract --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/flowfile-architecture-contract .opencode/skills/flowfile-architecture-contract && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "flowfile-architecture-contract" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-architecture-contract into .opencode/skills/flowfile-architecture-contract/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-architecture-contract", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
flowfile-architecture-contractMaps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging.
This skill is a map of the system and the reasons behind its cross-service contracts. It explains who talks to whom, why core never collects a LazyFrame on the hot path and ships paths and query plans instead, the worker offload wire protocol, the $ffsec$ secrets format, dual node-execution state, node-hash caching discipline and the pending directions for decentralizing execution and the catalog.
It deliberately does not teach how to add a node, run the stack, set environment flags or write tests, and it points to sibling Flowfile skills for each, including debugging, failure archaeology and change control. One detail it records is that a single SQLite catalog database is the metadata source of truth, which the worker does not open, while the scheduler uses a hand-maintained lightweight mirror in shared/models.py.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 13aa287. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitpythoncurlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git and curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
FLOWFILE_INTERNAL_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Flowfile Architecture Contract loads about 9.9k tokens when it runs. Until then it costs about 173 tokens; SKILL.md has 4,434 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Edwardvaneechoud/Flowfile at commit 13aa287, republished under its MIT licence (© Edwardvaneechoud). 4,434 words, ~9,886 tokens.
.claude/skills/flowfile-architecture-contract/SKILL.md (or your agent's skills folder).This skill is the map of the system and the why behind its cross-service contracts. It does not teach you how to build a node, write a test, or run the app — it teaches you where the load-bearing walls are so you don't put a door through one by accident.
add_<type> method,
frontend component) → flowfile-node-development.flowfile-build-and-env / flowfile-run-and-operate.flowfile-config-and-flags.flowfile-testing-and-validation.flowfile-ai-subsystem.flowfile-frontend-conventions.flowfile_frame Python API and code generation → flowfile-frame-and-codegen.flowfile-debugging-playbook (this skill explains why the failure mode exists; that one walks you through diagnosing this incident).flowfile-failure-archaeology.flowfile-change-control — nothing here overrides it.Frontend (Tauri desktop / Vite web, :8080) ──HTTP(Axios)──► flowfile_core (:63578)
FastAPI: JWT auth, catalog, secrets, AI, DAG engine
│ HTTP(query-plan bytes) + WebSocket(streaming) │ Docker SDK
▼ ▼
flowfile_worker (:63579) kernel_runtime containers
spawned subprocess per job uvicorn :9999 in-container,
(dataset memory lives here, host-mapped to 19000-19999,
never in the FastAPI process) call back to core on :63578
flowfile_scheduler — polls the shared catalog DB via its own models; embedded in
core's process when FLOWFILE_SCHEDULER_ENABLED is set, or run standalone.
shared/ — cross-package utilities (storage paths, crypto, cloud_storage, kafka,
ml, rest_api) used by core/worker/scheduler/kernel_runtime. Import direction
is strictly downward: shared never imports core/worker/scheduler/frame.
flowfile_frame — Python API; imports flowfile_core in-process to build FlowGraph
objects directly. Not a separate service.
WASM frontend (Pyodide) — runs fully in-browser, no core/worker/kernel at all.One SQLite catalog DB (flowfile_catalog.db) is the metadata source of truth,
resolved via shared/storage_config.get_database_url(). Correction to root
CLAUDE.md: the worker does not open this DB — verified, no
engine/session in flowfile_worker/ (only shared.sql_utils for user
database connection URIs). Core has the full ORM (database/models.py);
scheduler and the CLI run-completion path use a lightweight declarative
mirror in shared/models.py (§5) — a hand-maintained copy, not core's own
tables and not SQLAlchemy reflection.
The rule: flowfile_core must never call .collect() on a real dataset's
pl.LazyFrame on the hot path. Core ships paths and query plans; the
worker holds dataset memory, and only inside subprocesses it spawns
(mp_context = get_context("spawn") in flowfile_worker/__init__.py) —
never in the worker's own FastAPI process. That's why the worker exists as
a separate service at all: a crashed/OOM-killed job takes down a killable
child process, never the API that's holding open connections for every other
user.
Enforcement status — read this before assuming a lint catches violations:
it is a convention, not a mechanically enforced rule. No lint or test
forbids .collect() in core. Guardrails that do exist: bounded/preview
collects are corralled inside FlowDataEngine itself (collect(n) →
head(n).collect(engine="streaming"), collect_schema(), pl.len() for
counts); a load-bearing comment in flow_node.py's _do_execute_remote warns
to use is not None, never truthiness, on a FlowDataEngine because
__len__ collects; NodeTemplate.laziness metadata plus
FlowNode.check_upstream_laziness / FlowGraph.check_flow_laziness
(GET /editor/laziness_check) report eager nodes upstream of a catalog
virtual table — but that only gates one optimization path, not a general
rule. Accepted, deliberate exceptions where core does collect a full
frame: writing a Delta table for the catalog locally, the explore-data/Graphic
Walker preview path, seeding a brand-new manual-input datasource node, and
FlowDataEngine.to_arrow()/to_pylist()/to_dict() for previews.
Violation symptom: a .collect() on a real (non-preview,
non-accepted-exception) LazyFrame inside flowfile_core moves a potentially
multi-GB materialization into the FastAPI process every user's request goes
through — the whole point of the worker split is defeated, and a bad file
hangs or OOM-kills the API server itself instead of one job's subprocess.
Config (configs/settings.py): OFFLOAD_TO_WORKER is a live-mutable
MutableBool read from FLOWFILE_OFFLOAD_TO_WORKER, default on. Gotcha:
the check is os.environ.get("FLOWFILE_OFFLOAD_TO_WORKER", "1") == "1" —
only the exact string "1" is truthy; "true"/"yes"/"on" all evaluate
false, silently disabling offload, unlike the AI feature flags which
accept the full true/1/yes/on set — don't assume symmetry. WORKER_URL
resolves from FLOWFILE_WORKER_URL if set, else WORKER_HOST env, else
127.0.0.1 (Windows) / 0.0.0.0 (else), port FLOWFILE_WORKER_PORT
(default 63579); appends /worker when FLOWFILE_SINGLE_FILE_MODE co-hosts
the worker on the core port.
Wire protocol (flow_data_engine/subprocess_operations/subprocess_operations.py):
serialization is pl.LazyFrame.serialize() raw bytes — the query plan, not
data. POST {WORKER_URL}/submit_query/, Content-Type: application/octet-stream, headers X-Task-Id (= the node's hash, see §7),
X-Operation-Type (store | calculate_number_of_records | …), X-Flow-Id,
X-Node-Id. Sampling uses POST /store_sample/ with X-Sample-Size; fuzzy
join uses POST /add_fuzzy_join. Preferred transport is a WebSocket
stream, falling back to REST + polling on any connect/send error. GET /status/{file_ref} returns job status; a polars-op result is a base64
serialized LazyFrame — the worker hands back a lazy scan over its own
cached parquet, still not materialized data, so core stays clean reading the
result back too. results_exists(hash) (GET /status/{hash}) returns
False on connection error (worker down), not an exception — a known weak
point, see §11.
Where core calls the seam: FlowNode._do_execute_remote builds the lazy
result, ships it via ExternalDfFetcher(lf=…, file_ref=self.hash), and reads
back a lazy scan plus a row count from a second fetcher call.
_do_execute_local_with_sampling computes narrow transforms lazily in core
and only ships a 100-row ExternalSampler job to the worker for the preview.
Source nodes doing network I/O (database, Kafka, Google Analytics, REST API)
never fetch in core — they use an External*Fetcher/External*Writer
(add_google_analytics_reader is the canonical template): the worker does
all fetching. Sinks (output, api_response, flow_output, run_flow) stay
in-core even in remote mode; the output node's own function routes the actual
write through ExternalOutputWriter when not local. Every fetcher-based node
also keeps an in-core execution_location == "local" branch (e.g. the
database reader builds a local SqlSource) for CLI --run-flow / package
mode where no worker process exists — that local branch is not a violation of
"worker does all fetching," it only fires when there is no worker.
$ffsec$1$<user_id>$<token> — never change one side aloneFormat: $ffsec$1$<user_id>$<fernet_token> — the leading $1$ is the format
version, <user_id> is embedded in plaintext, <fernet_token> is the
Fernet-encrypted secret. Key derivation:
HKDF-SHA256(master_key, length=32, salt=KEY_DERIVATION_VERSION, info=f"user-{user_id}")
→ base64-urlsafe → Fernet key, KEY_DERIVATION_VERSION = b"flowfile-secrets-v1".
Why the user_id is embedded: the worker re-derives the per-user key with zero context from core — no HTTP call back, no session, nothing but the shared master key and the id parsed out of the ciphertext prefix. That's what lets a spawned worker subprocess decrypt a DB/cloud/Kafka credential completely independently.
secret_manager/secret_manager.py (core) and secrets.py (worker) are
byte-for-byte parallel implementations of this format (same
SECRET_FORMAT_PREFIX, same KEY_DERIVATION_VERSION). Changing the
prefix, salt, or HKDF info string on one side without the other is the
single easiest way to break every DB/cloud/Kafka/GA connector in the app —
and it's a confusing split-brain failure: core still works (it never
re-derives), every worker job throws InvalidToken.
Legacy fallback: ciphertext not starting with the prefix is a raw Fernet token decrypted with the master key directly (pre-per-user-key secrets). Group sharing is authorization-only (§10), so this format needed zero changes when sharing shipped.
Serialized plans are a transport too. Polars inlines storage_options
into a LazyFrame's serialized plan, so a cloud scan built in core with
decrypted keys would ship them to the worker (and into logs/caches) in
plaintext. Core-built cloud scans therefore pass
CloudStorageReader.get_secure_scan_kwargs(...): the credentials ride in a
shared/cloud_credential_provider.py::EncryptedCredentialProvider (a Polars
credential_provider whose only state is a $ffsec$ ciphertext under the
node's user), decrypted by whichever process executes the plan through the
decryptor it registered (core: flow_data_engine/cloud_storage_reader.py;
worker: flowfile_worker/utils.py). ADLS service-principal secrets and GCS
service-account keys cannot go through a Polars provider and still travel in
the options.
flowfile_core imports, everflowfile_scheduler is a separate, dependency-light package. Verified:
grep -rn "flowfile_core|flowfile_worker|flowfile_frame" flowfile_scheduler/flowfile_scheduler/ matches only two docstring sentences
declaring the rule — zero actual imports. The scheduler's models.py is a
backward-compat shim re-exporting from shared/models.py — independent
declarative SQLAlchemy models on their own Base, a minimal hand-maintained
mirror of flowfile_core/database/models.py, not SQLAlchemy table
reflection. (Root CLAUDE.md's "polls the shared SQLite DB via reflected
tables" is loose phrasing — treat "reflected" as informal, not literal
Table(..., autoload_with=engine) reflection.) Dependency arrow is core →
scheduler, never reverse.
Why it matters: a scheduler feature needing a new column means editing
shared/models.py to keep the mirror in sync, and a real Alembic migration
against flowfile_core/database/models.py (the canonical schema) — touching
only one side either breaks the scheduler's queries or leaves core's
migration history incomplete.
Worth knowing: single-leader via a SchedulerLock row (90s stale threshold);
cron uses a naive local wall-clock cursor advanced to "now" (not the
missed slot) on catch-up so a downed scheduler fires exactly once rather than
backfilling; runs launch as a detached subprocess
(python -m flowfile run flow <path> --run-id <id>), not an in-process call
— this is why the scheduler can be embedded in core without core's own crash
taking a running scheduled flow down with it.
Every FlowNode carries two parallel state representations:
node_stats: NodeStepStats — legacy. Still read by needs_run and
get_predicted_resulting_data (schema prediction)._execution_state: NodeExecutionState — new. Governs actual run/skip
decisions via NodeExecutor._decide_execution.They are synced one-way: NodeExecutor._sync_state_to_legacy copies the
new state into the legacy one after every decision/run. There is no path
that updates only node_stats and expects _execution_state to follow —
if you add a new run-affecting flag, write it to _execution_state and let
the sync propagate it; writing only node_stats is invisible to
_decide_execution and writing only _execution_state without going through
_sync_state_to_legacy leaves schema prediction (needs_run) looking at
stale data. Symptom of getting this wrong: the UI's "needs run" badge and
the actual run/skip behavior disagree — one lags the other by exactly one
sync call.
NodeExecutor._decide_execution is the single source of truth for both
whether to run and how (executor.py). Order matters — evaluated top to
bottom, first match wins:
| # | Condition | Result | Reason tag |
|---|---|---|---|
| 1 | node_template.node_group == "output" | RUN | OUTPUT_NODE — sinks always run |
| 2 | node_type == "run_flow" | RUN | subflow file can change with no parent-settings change, never "up to date" |
| 3 | force_refresh (reset_cache) | RUN | FORCED_REFRESH |
| 4 | cache_results=True, results_exists(hash) | SKIP | cache hit — checked before performance mode so a cached result survives even when upstream had no new data |
| 4b | cache_results=True, not results_exists | RUN | CACHE_MISSING |
| 5 | performance_mode | RUN | PERFORMANCE_MODE |
| 6 | not has_run_with_current_setup | RUN | NEVER_RAN |
| 7 | read node, source file changed on disk | RUN | SOURCE_FILE_CHANGED |
| 8 | else | SKIP | results already in memory from a previous run |
Strategy (_determine_strategy, once RUN is decided): run_location == "local" → FULL_LOCAL; else cache_results → REMOTE (caching needs a
fully materialized result); else transform_type == "narrow" →
LOCAL_WITH_SAMPLING (compute lazily in core, ship only a 100-row sample
job); else → REMOTE.
_hash save/restore dance, and parameterized runsFlowNode.calculate_hash folds: input-node hashes + hash(setting_input) +
parent_uuid (a uuid1 stamped once per FlowGraph instance at
construction) + _cache_epoch. The hash is the worker's cache key
(file_ref in the wire protocol above) — results_exists(hash) and
get_external_df_result(hash) both key off it.
Because parent_uuid is per-graph-instance, worker cache keys never
survive a graph reload — restarting core or reopening a flow always
recomputes from scratch even if nothing changed. This follows directly from
§8's in-memory-only invariant but is easy to forget if you expect cache hits
across sessions.
The _hash save/restore discipline (FlowGraph._execute_single_node):
when a flow has ${param} placeholders, execution substitutes resolved
parameter values into setting_input in place before running the node,
then restores the original ${...} text afterward — and explicitly saves
and restores node._hash around that substitution. Why: mutating
setting_input in place, even temporarily, changes what calculate_hash
would compute; without the explicit restore, the next real settings write
after a parameterized run sees a hash mismatch, treats it as an external
change, calls reset(), and silently loses example_data_generator /
has_completed_last_run — a real bug this exact comment documents having
fixed once. If you touch parameter substitution, preserve this save/restore
or you reintroduce that regression.
FlowfileHandler keeps every open flow in an in-memory dict (process
memory, not the DB — the DB only stores catalog registrations of flows, not
live graph state). A core restart, or a "Save As" that changes the
in-memory flow's identity, invalidates the frontend's flow_id. The run
route has an explicit guard for this: hitting it with a stale id returns a
targeted 404 telling the user to reload the flow, rather than a generic
crash. If you add a new endpoint keyed by flow_id, expect the same failure
mode and handle it the same way — don't assume a flow_id a client sends is
still resolvable.
Each kernel is a Docker container running uvicorn kernel_runtime.main:app --host 0.0.0.0 --port 9999 inside the container (EXPOSE 9999,
healthcheck curl -f http://localhost:9999/health). Local (electron/
package) topology: core maps that container port to a host port from
19000-19999 (_BASE_PORT = 19000, _PORT_RANGE = 1000); kernel URL is
http://localhost:{allocated_port}. Docker-in-Docker (compose)
topology: no host port mapping — kernels are reached by container name
(http://flowfile-kernel-{id}:9999) on the shared Docker network, because
core discovers the named volume covering the shared path and mounts the
same volume at the same path in the kernel container, so paths line up
identically across core/worker/kernel without translation.
Kernel → core auth: every kernel gets FLOWFILE_INTERNAL_TOKEN injected
(via auth.jwt.get_internal_token()) plus FLOWFILE_CORE_URL; the kernel
sends the token per-request, core verifies with secrets.compare_digest.
Electron auto-generates the token if unset; docker mode hard-errors
on a missing token — deliberate, electron is single-user and docker is
multi-tenant. A valid token with no kernel id becomes a synthetic
_internal_service principal that bypasses catalog access restriction (§10)
but must never be treated as a real user for ownership — sharing code keys
the synthetic check on username == "_internal_service", never on id,
because that principal's id defaults to 1, a real user id.
Known asymmetry (gotcha): storage.shared_directory and the derived
artifact/global-artifact dirs honor FLOWFILE_SHARED_DIR if set, but
production's get_kernel_manager() hardcodes
storage.temp_directory / "kernel_shared" and ignores that env var — setting
it to a non-default path in a real deployment splits artifact staging away
from the volume kernels actually mount. Tests dodge this by constructing
KernelManager with the path passed explicitly.
Group-based resource sharing (12 shareable resource types: secrets,
DB/cloud/Kafka/GA connections, catalog namespaces/tables/notebooks, flows,
visualizations, dashboards, global artifacts) is a pure authorization
layer bolted on top of existing ownership. The load-bearing property: a
shared secret's ciphertext stays encrypted under the owner's user_id
forever — sharing never re-encrypts it under the grantee's key. That's why
flowfile_worker (§4) and flowfile_scheduler (§5) needed zero code
changes to support sharing: a group-granted secret decrypts through the
exact same owner-keyed-derivation path as an owned one.
Consequences of "authorization-only, never re-key": rotating a secret on a
shared connection re-encrypts it under the owner's id, never the
manage-grantee's who triggered the rotation. Changing a shared connection's
target fields (host/endpoint/protocol) while it carries a bundled secret
and no new credentials were supplied is rejected with 422 — an
anti-repoint-harvest guard against a manage-grantee silently repointing the
connection at a server they control to harvest the owner's credentials on the
next connect. The catalog is private-by-default in docker mode
(electron/package stay fully open); resolution is own-first (own rows, then
group-granted, lowest row id wins on name collisions). /user-groups and
/shares 404 in electron mode (require_sharing_enabled dependency,
routes/user_groups.py), same status as a genuinely missing resource — no
enumeration oracle. Every resource-delete path must call
sharing.delete_grants_for_resource(...): SQLite reuses row ids, so a
leftover grant on a deleted resource's old id would silently reattach to
whatever unrelated resource claims that id next. An ORM after_delete
backstop covers single-row deletes, but bulk query.delete() bypasses ORM
events and must call the cleanup explicitly.
| Weak point | What it means in practice |
|---|---|
flow_graph.py god file (largest module in the repo) | Holds the DAG engine, the add_* node builders, catalog Delta write helpers, ML train/apply plumbing, kernel execution, YAML serialization, groups, layout, history, and codegen entry all in one file. Navigate it with grep -n "def add_" flowfile_core/flowfile_core/flowfile/flow_graph.py — that's the practical index; don't try to read it top to bottom. |
| Skip-list shallowness | The pre-run skip pass (util/node_skipper.py) only expands one transitive level of "leads to" from incorrectly-configured nodes. Correctness for deeper chains relies on _execute_stages re-checking dependents of every failed/skipped node at run time — a genuine belt-and-suspenders design, not a hole, but don't assume the skip pre-pass alone is a complete skip set. |
results_exists swallows worker downtime | Returns False on an HTTP connection error, same as a genuine cache miss. In Development mode this silently degrades "should skip, cached" decisions into full re-runs whenever the worker is briefly unreachable — no error surfaces, just unexpectedly slower runs. |
| HTTP 419 on add-node failure | POST /update_settings/ raises the non-standard status code 419 (not 422/500) when the node's add_<type> function itself raises — handle it explicitly in frontend/AI tool wrappers. Full dispatch-trap mechanics: flowfile-node-development §1.6. |
FLOWFILE_OFFLOAD_TO_WORKER truthy parsing | Only the literal string "1" enables offload; "true" silently disables it. See §3. |
mainAs of 2026-07-03 there is an unmerged branch,
claude/core-abstraction-flowgraph-001hrn, carrying 7 commits (5 substantive
main — matching flowfile-failure-archaeology's
branch inventory) that begin
decomposing the seams in §2, §3, and §6. None of this is on main or any
release — treat it as a candidate design, not current behavior, until it
lands:WorkerTransport (flowfile/execution/transport.py) — proposed sole owner
of worker URLs/HTTP/WS, with typed exceptions instead of §3's ad-hoc
error-code scheme.ExecutionBackend ABC with LocalBackend/RemoteWorkerBackend
(flowfile/execution/backends/) — proposes replacing inline
if execution_location == "local" branches (§3) with one
backend.run_lazyframe/sample/count_records call per node. Adds a
ratchet test pinning the count of remaining inline branches in
flow_graph.py, failing if it goes up — worth reusing regardless of
whether this branch lands.NodeSpec registry (flowfile/node_registry/) — proposed single source of
truth for built-in node types, unifying four independently-maintained
catalogs (get_all_standard_nodes(), NODE_TYPE_TO_SETTINGS_CLASS,
nodes_with_defaults, ai/tools/classification._NODE_CLASS_MAP — see
flowfile-node-development for why keeping these in sync today is a known
parity hazard) as derived views over one NodeSpec per type.If asked to work near these seams, check whether the branch has merged
(git log --oneline origin/main..origin/claude/core-abstraction-flowgraph-001hrn — empty
means merged or superseded) before assuming either the old inline-branch
shape or the new backend/registry shape is current truth.
Agreed direction for the notebook kernel (flowfile_frame/notebook_kernel.py,
flowfile_core/notebook/kernel_runner.py), built in phases 0–4a below. The
kernel is a separate execution mode: node settings stay the contract that
round-trips to the canvas on Push, the kernel runs natively what it can, and
everything else goes to the canvas. The end state, reached with phase 3, is a
kernel that opens no catalog database connection: the per-kernel SQLite
copy (kernel/notebook_db.py, there because WAL did not cross the Docker
Desktop VM) and the SQLite requirement of the gate (notebook/gate.py) are
gone. Do not add SCD2, change-feed or SQL logic to
kernel_runtime/flowfile_client.py, forward raw SQL to core, bring a
database back into the kernel, or mount a host folder into it.
Phase 0 landed with the plan: a lineage run that holds no output node
passes run_graph(commit_sources=False) (kernel_runner.lineage_commits),
and a session takes the schemas of canvas nodes that ran
(_Session.refresh).
Phase 1 landed 2026-10-03: held nodes run in core from their settings.
Where the kernel's _Session._resolve finds a gate or deferred node without
a canvas twin (or a read of a file it cannot see), it hands core the node as
FlowfileData lists it plus its inputs as parquet under the session's
results folder: notebook/held_run.py, POST /notebook/session/node_run,
bound like node_result. Core builds the inputs as flow_input nodes fed
those files (not add_dependency_on_polars_lazy_frame: a NodePromise node
is never is_correct, so the planner skips it) plus the one node through
populate_graph_from_flow_information, on a FlowGraph that lives for the
call (_system_run, Performance mode, the canvas flow's
source_registration_id and the session's parameters), runs it under
KernelHold with commit_sources=False, answers every live output's path
(a gate's dead handles as closed) or, for schema_only, its columns, and
releases the graph's FlowLogger, log file and exchange folder. Closed
allowlist HELD_NODE_TYPES plus installed non-output custom nodes; writers
and other types still say "Push". A cell-built deferred node whose seed has
no columns asks core for them at build (NotebookMode.schema_resolver,
native.resolved_seed). A Python Script on another kernel runs this way
too, with its artifacts on the call's graph only.
Phase 2 landed 2026-10-03: metadata lookups go through core. One hook,
notebook/lookup.py::metadata_lookup (a ContextVar, the placement_check
pattern), consulted by flow_graph._resolve_catalog_table_info /
_resolve_catalog_sql_tables, storage_backend.resolve_for_namespace,
subflow._registration (behind stamp_flow_reference and
resolve_subflow_path), prechecks.placement_refusal and the frame's
flowfile_frame/_metadata.py (behind catalog_reference.py, run_flow.py,
kernels.py and the connection listings; every public function asks the hook
when set, else runs its local_* twin on the database). The kernel installs a
_metadata.SessionLookup for every op (_metadata.installed), which posts to
POST /notebook/session/lookup, a closed lookup.KINDS each answered by the
function the call site runs without the hook, as the kernel's owner, bound by
kernel_runner._bound_kernel (the flow need not be open), metadata only: no
plan, no storage credential, no password, no ciphertext. Secret-bearing nodes
are held in a kernel session and take phase 1's path: every
cloud_storage_reader (NOTEBOOK_DEFERRED_NODE_TYPES), a catalog reader of a
cloud-backed table (_metadata.is_cloud_table, from the catalog_table
answer), a custom node with a SecretSelector (custom_node._selects_secret);
native._predicts_in_core keeps CONNECTION_SOURCE_TYPES from running their
schema callback while a schema_resolver is set. The census is
tests/notebook/test_kernel_database_census.py, at zero: a pool checkout
listener under the sim's IN_KERNEL_OP marker over the whole corpus
(conftest.kernel_db_opens), plus every kind answered without a $ffsec$ or
a decrypt.
Phase 3 landed 2026-10-03: the copy is gone. A notebook kernel holds no
catalog database: _notebook_env sets no FLOWFILE_DB_PATH (the kernel's
get_database_url() falls to the never-mounted <storage>/database/), and
every notebook_kernel.handle() first installs _refuse_database, a
do_connect listener on the one cached engine behind connection.engine,
SessionLocal and get_db_context (gated on FLOWFILE_KERNEL_ID, which only
a kernel container has), so a cell that opens the database itself stops with
NO_DATABASE, naming where it was opened (_opened_from: a cell, else the
frame module of a regression) and the ff functions to use, before pysqlite
could create a file. Deleted: kernel/notebook_db.py,
POST /notebook/session/database, kernel_runner.refresh_database, _rearm
and the refresh listeners, the hello schema-head handshake
(migration.package_head, kernel_runner._schema_revision; the flowfile
version check stays) and the gate's SQLite requirement
(kernel_sessions_allowed is now not sharing_enabled(), which admits an
electron app on a Postgres catalog without proving it). An uncached
shared.database.create_catalog_engine() a cell calls by hand is not
refused and creates an empty file in the container layer, as in any kernel
with flowfile installed. Proof: test_kernel_database_census.py (no op
connects), test_kernel_session.py::test_a_notebook_kernel_refuses_to_open_a_catalog_database
(scratch engine) and the real-kernel
test_kernel_notebook_docker.py (the default database file never appears in
the container; a get_db_context() cell is refused; a namespace core creates
shows in the next cell), which .github/workflows/test-notebook-kernel.yml
now runs on every frame, notebook or kernel change
(FLOWFILE_REQUIRE_NOTEBOOK_KERNEL makes a skip a failure).
Phase 4a landed 2026-10-03: no host folder is mounted. A notebook
kernel's container mounts what every kernel mounts, /shared and
/catalog_tables, and nothing else: not the Flowfile folders (flows, custom
nodes and their mount directories, catalog tables) and not a folder its owner
lists. Deleted: kernel/notebook_mounts.py, mounted_folders
(models.MountedFolder, the kernel form's folders field, migration 034's
column, dropped by 035; a pydantic model ignores it from an older client),
FLOWFILE_NOTEBOOK_MOUNTS, notebook.kernel_path,
input_schema.kernel_file_path, KernelManager.host_folders, the
/host/<drive>/ Windows scheme, the key-store tmpfs, FLOWFILE_STORAGE_DIR
in _notebook_env (the kernel's storage is its own ~/.flowfile,
/root/.flowfile in the image) and _metadata.is_cloud_table /
SessionLookup.cloud_tables. What replaced each use: a file a cell names is
never opened in the kernel (notebook.paths_as_written sets only
keep_paths_as_written, under which native._kernel_hidden_path hides every
local read/list_files path, ${param} ones included) and core reads it
(the canvas twin's rows, else a held run); whether a bare path is a folder is
the is_directory lookup (flow_frame_methods._resolve_scan_mode); every
catalog reader in a kernel session is deferred (native.notebook_defers) and
its schema and rows come from core (native._predicts_in_core,
CORE_PREDICTED_TYPES = CONNECTION_SOURCE_TYPES | {"catalog_reader", "run_flow"}); a run_flow reference's interface is the flow_interface
lookup (_metadata.flow_interface, core-side resolve_subflow_path +
get_subflow_interface); installed custom node files come through the
custom_node_sources lookup, which notebook_kernel._mirror_custom_nodes
writes as <node_key>.py into the kernel's own nodes folder before
open/reset/execute/clean_run (it removes only files it wrote and
rescans the registry once when anything changed; the kernel still runs a
local node's process() itself); a push stores paths as written and
notebook/validate.host_file_paths only recomputes abs_file_path on the
host. Invariant: a kernel mounts exactly /shared and /catalog_tables; a
file path in a cell is a path on the user's machine that only core opens.
Plain Polars or open() on a host path in a cell finds no file, and a cell
writes no file (an ff writer adds a writer node). Known cost: a catalog
table a cell reads is handed over whole as parquet through /shared, even
for display(). Proof: tests/kernel/test_notebook_kernel_env.py (exactly
the two binds, no mounts or tmpfs, none of FLOWFILE_STORAGE_DIR,
FLOWFILE_NOTEBOOK_MOUNTS, FLOWFILE_DB_PATH; an old client's
mounted_folders ignored), the sim tests in
tests/notebook/test_kernel_canvas_rows.py (a read is a held read run, a
bare folder asks is_directory) and test_kernel_database_census.py (a
catalog table, a flow file and a custom node reach the kernel through core;
the three new kinds answer metadata only), and the real-kernel
test_kernel_notebook_docker.py::test_an_installed_custom_node_is_mirrored_into_the_kernel
(with test_a_file_a_cell_names_is_read_by_core, where pl.read_csv on the
same path has no such file).
Remaining, each with its own plan: a Python Script on the notebook's own
kernel (KernelHold refuses it during a fallback), a push that keeps the
session's variables, a row-limited fallback for display() (which now also
bounds a catalog table a cell reads, until then handed over whole), Postgres
catalogs and docker-mode sessions (per-user access in the lookups, admin
gating on node_run), session artifacts visible to the canvas.
Known gaps this leaves until then: a seeded canvas node below one that just
ran keeps its seeded columns until it runs or the session is reset, and
FlowFrame.collect() on a deferred frame in a script still commits sources
(flow_frame.py calls run_graph(node_ids=...) with the default).
The maintainer's stated direction: decentralized execution on a centralized catalog. Global/shared components (catalog schema, definitions, sharing model), catalog metadata in PostgreSQL, catalog data in object storage (S3), and independent local Flowfile instances that each bring their own compute and mix local files with catalog-hosted remote files in the same flow.
This is a direction, not a spec, and not implemented anywhere in the repo today. Nothing above contradicts it — but two current seams are exactly what a future implementation would generalize, so do not add code that tightens their coupling to "local + SQLite" further:
shared/storage_config.get_database_url() (§1) is the only function
deciding where the catalog metadata DB lives — a local SQLite file today
(FLOWFILE_DB_PATH env, else TESTING temp path, else
<base>/database/flowfile_catalog.db), no PostgreSQL branch yet. Route new
catalog code through this seam (or storage/CatalogService above it)
rather than hard-coding a SQLite path assumption inline.catalog/storage_backend.py already demonstrates the pattern the vision
needs for data: storage resolves per catalog (a level-0 namespace
carries storage_uri + storage_connection_name; unset ⇒ local
filesystem, set ⇒ object storage, credentials resolved as the catalog
owner). This supports "data in S3" already; it does not yet support
"metadata in Postgres" or "multiple compute instances sharing one
catalog." Extend along this per-namespace-target pattern, not a single
global env-var switch — the module already demoted the old
FLOWFILE_CATALOG_STORAGE_URI/_CONNECTION env vars from "the mechanism"
to "just a creation-time default," which is this vision's direction.If asked to design a feature touching catalog storage, catalog identity, or cross-instance execution, treat "would this make a future Postgres-metadata/multi-instance-compute design harder" as a review question — not a blocker, but flag it rather than silently baking in a local-SQLite assumption.
All facts above were verified by reading source or running the listed
read-only command against the repo at commit f6963c77 (branch
feature/claude-skills), app version 0.12.7, dated 2026-07-03. Every
volatile fact below should be re-checked with its command if this skill feels
stale — architecture facts like these drift slower than config defaults, but
line numbers and the pending-branch status will drift fastest.
# system map / "worker doesn't touch the catalog DB" claim
grep -rn "create_engine|Session|shared.models" flowfile_worker/flowfile_worker/ | grep -v __pycache__
# core-never-collects: god-file sizing
wc -l flowfile_core/flowfile_core/flowfile/flow_graph.py \
flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py \
flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
# worker offload truthy parsing ("1" only)
grep -n "FLOWFILE_OFFLOAD_TO_WORKER" flowfile_core/flowfile_core/configs/settings.py
# worker wire-protocol endpoints
grep -n "@router" flowfile_worker/flowfile_worker/routes.py
# secrets format parity between core and worker
grep -n "SECRET_FORMAT_PREFIX\|KEY_DERIVATION_VERSION" \
flowfile_core/flowfile_core/secret_manager/secret_manager.py \
flowfile_worker/flowfile_worker/secrets.py
# scheduler zero-core-imports invariant
grep -rn "flowfile_core|flowfile_worker|flowfile_frame" flowfile_scheduler/flowfile_scheduler/ | grep -v __pycache__
# NodeExecutor decision table (still ordered the same way?)
sed -n '/_decide_execution/,/_determine_strategy/p' flowfile_core/flowfile_core/flowfile/flow_node/executor.py
# kernel port range + internal token wiring
grep -n "_BASE_PORT\|_PORT_RANGE" flowfile_core/flowfile_core/kernel/manager.py
grep -n "get_internal_token\|verify_internal_token" flowfile_core/flowfile_core/auth/jwt.py
# group sharing 404-in-electron gate
grep -n "require_sharing_enabled" flowfile_core/flowfile_core/routes/user_groups.py
# HTTP 419 on add-node failure
grep -n "419" flowfile_core/flowfile_core/routes/routes.py
# catalog storage per-namespace resolution (vision seam)
sed -n '1,40p' flowfile_core/flowfile_core/catalog/storage_backend.py
# pending-branch status: has it merged / is it still distinct?
git log --oneline origin/main..origin/claude/core-abstraction-flowgraph-001hrn 2>/dev/null \
|| git log --oneline f6963c77..claude/core-abstraction-flowgraph-001hrn
# app version
grep -n '^version' pyproject.toml© Edwardvaneechoud, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/flowfile-architecture-contract of Edwardvaneechoud/Flowfile.
Open the folder on GitHubat commit 13aa287
Flowfile Architecture Contract next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Flowfile Architecture Contract this skillEdwardvaneechoud/Flowfile | 375 | — | ~9.9k | Automated safety check: Pass | MIT | |
| Burla Internals Deep DiveBurla-Cloud/burla | 263 | — | ~2.4k | Automated safety check: Pass | Custom licence | |
| Code Review Skillawesome-skills/code-review-skill | 2.1k | — | ~2.8k | Automated safety check: Notes | MIT | |
| GitDiagram Repository Overviewahmedkhaleel2004/gitdiagram | 18k | — | ~427 | Automated safety check: Pass | MIT | |
| Deepwiki Rssopaco/deepwiki-rs | 3.1k | — | ~748 | Automated safety check: Pass | MIT | |
| Code Graph Mermaid Diagramstrailofbits/skills | 7.4k | — | ~1.7k | Automated safety check: Pass | CC-BY-SA-4.0 |
Burla-Cloud/burla
Reference for Burla internals: how a remote_parallel_map job flows between services, how clusters and nodes are managed, and where the head keeps its state.
awesome-skills/code-review-skill
Provides comprehensive code review guidance for React 19, Vue 3, Angular 17+, Svelte 5, Rust, TypeScript, Java, Java 8, PHP, Ruby, Rails, Python, Django, FastAPI, Go, C/.NET, Kotlin, Swift, Dart…
ahmedkhaleel2004/gitdiagram
Explains the architecture of a public GitHub repository through GitDiagram: how the code is organized, the main components with paths, and a Mermaid diagram.
sopaco/deepwiki-rs
AI-powered Rust documentation generation engine for comprehensive codebase analysis, C4 architecture diagrams, and automated technical documentation.
trailofbits/skills
Generates Mermaid diagrams from Trailmark code graphs, including call graphs, class hierarchies, module dependency maps, complexity heatmaps and attack surface data flows.
ahmedkhaleel2004/gitdiagram
Explains how a public GitHub repository is built by fetching its GitDiagram architecture diagram, components and optional explainer video.
Edwardvaneechoud/Flowfile
Maps the /ai/ subsystem of flowfile_core, its three agent tiers, litellm seam, BYOK keys and rate limits, and sets rules for extending or debugging it safely.
Edwardvaneechoud/Flowfile
Recreates every Flowfile development and build environment from scratch, with exact version pins and an explanation of what each Makefile target really does.
Edwardvaneechoud/Flowfile
Explains how changes to the Flowfile monorepo are gated, versioned and released, including version sync, stub and docs drift checks, Alembic migrations and pinned dependencies.
Edwardvaneechoud/Flowfile
Runbook for closing gaps between a Flowfile visual flow's results and its exported Polars or FlowFrame Python code, measured by tests rather than by eye.
Edwardvaneechoud/Flowfile
Catalog of Flowfile's environment variables and runtime flags: what each does, where the code reads it, its default, and where the docs disagree with the code.
Edwardvaneechoud/Flowfile
Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom…
Categories
Maps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging. This skill is a map of the system and the reasons behind its cross-service contracts. It explains who talks to whom, why core never collects a LazyFrame on the hot path and ships paths and query plans instead, the worker offload wire protocol, the $ffsec$ secrets format, dual node-execution state, node-hash caching discipline and the pending directions for decentralizing execution and the catalog.
Flowfile Architecture Contract fits situations like: onboarding to the Flowfile codebase and learning how its services fit together; deciding which service a new code path belongs in; debugging contract breaks around secrets, worker offload, kernel containers or the scheduler; judging whether a design change would block the maintainer's decentralization plans.
Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a claude-code`. Or copy the skill folder (.claude/skills/flowfile-architecture-contract in Edwardvaneechoud/Flowfile) into .claude/skills/flowfile-architecture-contract in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a codex`. Or copy the skill folder (.claude/skills/flowfile-architecture-contract in Edwardvaneechoud/Flowfile) into .agents/skills/flowfile-architecture-contract in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-architecture-contract -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flowfile-architecture-contract, .gemini/skills/flowfile-architecture-contract, .github/skills/flowfile-architecture-contract and .opencode/skills/flowfile-architecture-contract in your project.
Going by SKILL.md and its folder, Flowfile Architecture Contract needs the command-line tools its instructions call (git, python and curl) and credentials named FLOWFILE_INTERNAL_TOKEN.
SKILL.md contains no URLs. Its commands use git and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Flowfile Architecture Contract is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 9.9k tokens (SKILL.md is roughly 40k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Flowfile Architecture Contract: Burla Internals Deep Dive (Burla-Cloud/burla, 263 stars), Code Review Skill (awesome-skills/code-review-skill, 2.1k stars), GitDiagram Repository Overview (ahmedkhaleel2004/gitdiagram, 18k stars) and Deepwiki Rs (sopaco/deepwiki-rs, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Edwardvaneechoud (a GitHub user) maintains it in Edwardvaneechoud/Flowfile, which has 375 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.
Source: Edwardvaneechoud/Flowfile on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.