Hybrid-Engine Data Analysis
code-yeongyu/oh-my-openagent
Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.
Runbook for closing gaps between a Flowfile visual flow's results and its exported Polars or FlowFrame Python code, measured by tests rather than by eye.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-codegen-parity-campaign -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-codegen-parity-campaign --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/flowfile-codegen-parity-campaign .claude/skills/flowfile-codegen-parity-campaign && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "flowfile-codegen-parity-campaign" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-codegen-parity-campaign into .claude/skills/flowfile-codegen-parity-campaign/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-codegen-parity-campaign", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-codegen-parity-campaignType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-codegen-parity-campaign -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-codegen-parity-campaign --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/flowfile-codegen-parity-campaign .agents/skills/flowfile-codegen-parity-campaign && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "flowfile-codegen-parity-campaign" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-codegen-parity-campaign into .agents/skills/flowfile-codegen-parity-campaign/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-codegen-parity-campaign", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-codegen-parity-campaign -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-codegen-parity-campaign --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/flowfile-codegen-parity-campaign .cursor/skills/flowfile-codegen-parity-campaign && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "flowfile-codegen-parity-campaign" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-codegen-parity-campaign into .cursor/skills/flowfile-codegen-parity-campaign/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-codegen-parity-campaign", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Edwardvaneechoud/Flowfile.git --path .claude/skills/flowfile-codegen-parity-campaign--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-codegen-parity-campaign -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-codegen-parity-campaign --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/flowfile-codegen-parity-campaign .gemini/skills/flowfile-codegen-parity-campaign && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "flowfile-codegen-parity-campaign" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-codegen-parity-campaign into .gemini/skills/flowfile-codegen-parity-campaign/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-codegen-parity-campaign", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-codegen-parity-campaignInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-codegen-parity-campaign -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/flowfile-codegen-parity-campaign .github/skills/flowfile-codegen-parity-campaign && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "flowfile-codegen-parity-campaign" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-codegen-parity-campaign into .github/skills/flowfile-codegen-parity-campaign/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-codegen-parity-campaign", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-codegen-parity-campaign -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Edwardvaneechoud/Flowfile flowfile-codegen-parity-campaign --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/flowfile-codegen-parity-campaign .opencode/skills/flowfile-codegen-parity-campaign && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "flowfile-codegen-parity-campaign" agent skill from https://github.com/Edwardvaneechoud/Flowfile/tree/main/.claude/skills/flowfile-codegen-parity-campaign into .opencode/skills/flowfile-codegen-parity-campaign/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flowfile-codegen-parity-campaign", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
flowfile-codegen-parity-campaignRunbook for closing gaps between a Flowfile visual flow's results and its exported Polars or FlowFrame Python code, measured by tests rather than by eye.
Flowfile runs the same node settings through two executors: the live engine behind the Designer and a code generator that exports standalone Python through export_flow_to_polars and export_flow_to_flowframe. This skill is a decision-gated campaign to find and close cases where the exported code runs cleanly but returns a different answer.
Parity means the engine's result and the exported code's result are equal under polars.testing.assert_frame_equal, ignoring column and row order. The campaign reproduces a baseline, works through an inventory of divergences, fixes them from a ranked menu of solutions and promotes changes through change control. Success is binary: no xfailed or xpassed tests in the edge-case suite, the full test_code_generator.py corpus passing for both exporters, and make check_stubs clean.
A worked example is incident #544, where a sort used the string desc in one place and Descending in another. Adding FlowFrame methods or understanding how the emitter works belongs to a different skill, flowfile-frame-and-codegen.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 13aa287. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
poetrymakegitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Flowfile Codegen Parity Campaign loads about 7.5k tokens when it runs. Until then it costs about 151 tokens; SKILL.md has 3,159 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Edwardvaneechoud/Flowfile at commit 13aa287, republished under its MIT licence (© Edwardvaneechoud). 3,159 words, ~7,480 tokens.
.claude/skills/flowfile-codegen-parity-campaign/SKILL.md (or your agent's skills folder).Read this first. Flowfile ships two executors of the same node settings: the live DAG engine
(what the Designer runs) and the code generator (export_flow_to_polars /
export_flow_to_flowframe), which emits standalone Python. When they disagree, the exported code
runs clean but produces a different answer — silent, order-dependent, data-corrupting bugs. The
sort bug in incident #544 (a "desc" vs "Descending" string mismatch) shipped exactly because
generation was never tested against execution. This skill is the runbook to find and close those
gaps measurably — never "by eye".
The maintainer named this the single hardest live problem in the repo (2026-07-03).
Status 2026-09-12 (branch
fix/xfails): rows S1, X1, X2 and D1 are closed — the three edge-case markers are deleted,make_uniquetreatscolumns=[]as all-columns (keepingkeep=strategy), and concat aggregations emitstr.join(',')through the sharedtransform_schema.STRING_CONCAT_DELIMITER(the engine'sstring_concatuses the same constant). Scoreboard: edge-case49 passed, 0 xfailed, 0 xpassed, corpus692 passed, custom-nodes20 passed. Phases 3.1, 3.2 and 3.4 below are historical; D2 and F1–F4 remain open.
Parity holds for a flow when, for every output node:
collect( execute_flow(flow).node[N] ) == collect( exec(export_code(flow))["run_etl_pipeline"]() )i.e. same input flow → exported code → executed result equals the in-engine result, compared with
polars.testing.assert_frame_equal(..., check_column_order=False, check_row_order=False). The repo
encodes this exact check in assert_flow_result_matches_generated(flow, output_node_id, code)
(flowfile_core/tests/flowfile/test_code_generator_edge_cases.py:65). Every parity test builds a
flow, runs both executors, and asserts frame-equality.
Campaign success is binary and machine-checked: zero xfailed and zero xpassed in the codegen
edge-case suite, plus the full test_code_generator.py corpus (parameterized over both exporters)
still green, plus make check_stubs clean. If you cannot state the pass/xfail numbers, you are not done.
Expr._repr_str mechanics → use flowfile-frame-and-codegen (that is the home for emitter
internals; this skill only fixes divergences).flowfile-node-development.flowfile-failure-archaeology (this skill uses #544 only as the cautionary rule).flowfile-change-control (routed to, not duplicated).flowfile-config-and-flags and
flowfile-testing-and-validation (this skill states the one command it needs and cross-refs there).make stubs internals / build orchestration → flowfile-build-and-env.The code generator lives entirely in core. Importing flowfile_core auto-runs Alembic against the
catalog DB, so isolate the DB on every test invocation or you will corrupt the shared catalog.
# Pick a scratch DB path OUTSIDE the repo, one per parallel run. NEVER omit this.
export PARITY_DB=/tmp/ff_parity_$$.dbHard rules (see flowfile-config-and-flags for why):
FLOWFILE_DB_PATH=$PARITY_DB.FLOWFILE_DB_PATH with FLOWFILE_SKIP_STARTUP_MIGRATION=1 — the DB
will have no tables and you get no such table: users cascades.poetry run.Run these two commands now, before changing anything, and confirm you see the embedded numbers. If your numbers differ, the inventory below is stale — STOP and re-verify against the source (Phase 2 line refs) before proceeding.
FLOWFILE_DB_PATH=$PARITY_DB poetry run pytest \
flowfile_core/tests/flowfile/test_code_generator_edge_cases.py \
-q -p no:cacheprovider -rXExpected baseline (2026-09-12): 49 passed, with no xfailed or xpassed lines.
The campaign rule: an XPASS in this repo means "delete the marker". Non-strict markers keep CI green, so nobody notices the rot.
Branches:
no such table: users (or similar OperationalError cascade) → FLOWFILE_SKIP_STARTUP_MIGRATION=1
is set in your environment alongside a fresh DB. unset FLOWFILE_SKIP_STARTUP_MIGRATION, pick a new
$PARITY_DB, re-run. Never combine that flag with a fresh test DB.xpassed → a marker went stale; each one is a delete-the-marker row. -rX names them all.failed or fewer than 49 passed → the tree drifted from this frozen baseline; re-verify the
Phase 2 anchors with the Provenance greps before touching anything.FLOWFILE_DB_PATH=${PARITY_DB}.main poetry run pytest \
flowfile_core/tests/flowfile/test_code_generator.py \
-q -p no:cacheprovider
# Expected baseline (2026-09-12): 692 passed (no skips locally)The round-trip flow-builders are parameterized @pytest.mark.parametrize("export_func", [export_flow_to_polars, export_flow_to_flowframe], ids=["polars", "flowframe"]) — the parametrize
sites produce the bulk of the corpus — so the corpus is the sample-flow set for round-trip
equivalence across both export frameworks. (The un-parameterized remainder are exporter-specific
tests such as test_flowframe_* at :1326+.) That count is your floor. Any fix that drops it has
regressed parity.
Branches:
failed → re-run with -rf to name them; a fix that breaks the corpus is a parity regression,
full stop — revert or fix before proceeding.There is no directory of
.flowfilefixtures for parity — the corpus is these programmatic flow-builders. Add new sample flows there, next tocreate_sample_dataframe_node(test_code_generator.py:318), not as loose files.A third file exists:
test_code_generator_custom_nodes.py(baseline:21 passed, <1s). It covers custom-node emission on the Polars exporter and the FlowFrame exporter; run it too whenever you touchcode_generator.py's custom-node path.
Every known divergence, with a status column. Work top-down; each row links to a fix phase.
| ID | Divergence | Buggy side | Location (verify before editing) | Guard test | Status 2026-07-03 |
|---|---|---|---|---|---|
| S1 | Stale XPASS: IN-filter numeric quoting — flow now splits on , before the numeric check, matching the emitter | (already fixed in filter_expressions.py:_build_in_expression, splits at :155) | marker test_code_generator_edge_cases.py:201 | TestBasicFilterOperators::test_in_operator_numeric | closed 2026-09-12 |
| X1 | unique(columns=[]) (all columns): flow's unique uses group_by internally → "at least one key is required in a group_by operation"; emitted .unique(keep='first') is correct | flow engine (executor) | fix site flow_data_engine.py:2686 (make_unique, called from flow_graph.py:2415); marker test_code_generator_edge_cases.py:459 | TestUniqueOperationVariations::test_unique_without_columns | closed 2026-09-12 |
| X2 | group_by concat aggregation: emitter emits str.concat with default - delimiter; flow uses , → ['x-y'] vs ['x,y'] | emitter | map expression_helpers.py:152 ("concat": "str.concat"), emitted with no args at transform_handlers.py:28; engine truth: transform_schema.py:119 (string_concat, delimiter=",") | TestGroupByEdgeCases::test_groupby_with_concat_aggregation | closed 2026-09-12 |
| D1 | Emitter emits deprecated str.concat (Polars → str.join; note default delim differs: - vs "") | emitter | expression_helpers.py:152 | test_groupby_concat_uses_deprecated_str_concat:1004 (asserts deprecated form; TODO@:1054) | closed 2026-09-12 (emits str.join(',')) |
| D2 | Emitter emits deprecated with_row_count (Polars → with_row_index) | emitter | transform_handlers.py:370 | test_record_id_uses_deprecated_with_row_count:960 (TODO@:1002) | passing, documents debt |
| F1 | FlowFrame export: right joins emit .collect().lazy() → returns a pl.LazyFrame, breaks the FlowFrame chain | emitter (FlowFrame framework) | join_handlers.py:477 TODO(FlowFrame) | (no dedicated round-trip guard) | open |
| F2 | FlowFrame export: formula nodes emit pl.col/pl.lit without import polars as pl when framework == "ff" | emitter (FlowFrame framework) | transform_handlers.py:54 TODO(FlowFrame) (+ :326) | (partial via test_flowframe_formula_*) | re-verify — TODO(FlowFrame) marker no longer present |
| F3 | FlowFrame export: fuzzy-match serializes Polars Expr via repr → invalid code pl.lit(<Expr ['len()'] at 0x...>) | emitter (FlowFrame framework) | transform_handlers.py:259 TODO(FlowFrame) | (none) | open |
| F4 | FlowFrame export: polars-code nodes reference ff.LazyFrame, which the flowfile package does not export | emitter (FlowFrame framework) | code_generator.py:567 TODO(FlowFrame) | (none) | re-verify — TODO(FlowFrame) marker no longer present |
| S2 | Stale XPASS (adjacent subsystem, not codegen): node-designer "numeric" string alias | (fixed) | marker node_designer/test_node_designer.py:516 | TestNumericStringAliasBug | closed (marker already gone from main) |
Two secondary emitter defects that are not yet guarded by a parity test (open, lower priority): param
codegen accepts Python keywords as parameter names → invalid signature (param_codegen.py:38); typed
flow-parameter empty default exports "" for int/float params
(flowfile_core/flowfile_core/flowfile/param_types.py:96 — note: one level above code_generator/).
Add a failing parity test before fixing either.
The FlowFrame framework-prefix trap underlies F1–F4: the same converter emits both
pl.-prefixed Polars andff.-prefixed FlowFrame code.export_flow_to_flowframeis newer and less complete thanexport_flow_to_polars. When a bug is FlowFrame-only, fix the FlowFrame converter path — do not "fix" it by making the Polars path emit FlowFrame syntax.
Do them in order. Each phase names the command, the exact expected observation, and a
if you see X → do Y branch. Do not batch phases; land and verify one at a time.
Removing a stale xfail is pure debt-payoff: the test already passes, so deleting the decorator makes
it a normal green test that will now catch regressions.
test_code_generator_edge_cases.py:201 — delete the @pytest.mark.xfail(...) decorator on
test_in_operator_numeric (row S1). Leave the test body and docstring; optionally trim the
"will fail until fixed" line in the docstring.node_designer/test_node_designer.py:516
(row S2).Verify:
FLOWFILE_DB_PATH=$PARITY_DB poetry run pytest \
flowfile_core/tests/flowfile/test_code_generator_edge_cases.py \
-q -p no:cacheprovider -rX47 passed, 2 xfailed and NO xpassed line.1 xpassed still → you deleted the wrong decorator; the XPASS is test_in_operator_numeric.1 failed → the underlying fix regressed between baseline and now; re-run Phase 1a, and
if it reproduces, that IN-filter split (filter_expressions.py:155) has been reverted — treat as a
new bug, not a stale marker.The emitter must emit the flow's actual concat delimiter. Row X2 is a correctness bug; D1 is the
paired deprecation. Fix them together (the delimiter must be explicit either way, because
str.concat defaults to - and str.join defaults to "", but the flow uses ,).
transform_schema.string_concat (transform_schema.py:117-119: .str.concat(delimiter=",")),
selected by AggColl.agg_func when agg == "concat" (transform_schema.py:886). The emitter maps
"concat" → "str.concat" in _get_agg_function (expression_helpers.py:152) and the call site
emits the mapped name with no arguments (transform_handlers.py:28:
f"{self.framework}.col({old}).{agg_func}().alias({new})") — so Polars' default - wins. Fix: make the
concat aggregation emit str.join(",") (modern name, delimiter explicit to match the engine). The
map returns bare method names that get () appended, so special-case concat at the emit site or
let the map carry an argument string.test_groupby_concat_uses_deprecated_str_concat (:1004): flip its TODO at
:1054 — assert uses_modern and not uses_deprecated.@pytest.mark.xfail on test_groupby_with_concat_aggregation (:766).Verify (both classes — the D1 guard lives in TestDeprecatedMethodUsage, not TestGroupByEdgeCases):
FLOWFILE_DB_PATH=$PARITY_DB poetry run pytest \
"flowfile_core/tests/flowfile/test_code_generator_edge_cases.py::TestGroupByEdgeCases" \
"flowfile_core/tests/flowfile/test_code_generator_edge_cases.py::TestDeprecatedMethodUsage" \
-q -p no:cacheprovider -rX3 passed, 0 xfailed, 0 xpassed. (Baseline for this pair is 2 passed, 1 xfailed;
TestGroupByEdgeCases holds exactly one test today.)test_groupby_with_concat_aggregation still xfailed → your emitted call still lacks the
explicit delimiter. Print export_flow_to_polars(flow) inside the test to see the emitted
str.join/str.concat call.test_groupby_concat_uses_deprecated_str_concat fails → you changed the emission but did not
flip the guard's assertion at :1054 (or vice versa) — the two must land together.assert_frame_equal fails with values like 'y,x' vs 'x,y' → the element order inside the
concatenated string differs (the helper already ignores frame row/column order); ensure your emitted
expression preserves the flow's group-member order — do not add a sort the flow doesn't have.with_row_count → with_row_index (emitter fix, semantics-identical)with_row_index(name=, offset=) is the drop-in modern replacement. Swap it at
transform_handlers.py:370, then flip the D2 guard TODO at :1002.
Verify:
# D1 and D2 guards both live in TestDeprecatedMethodUsage (edge-case file:953).
# The round-trip guard for record-id semantics is TestRecordIdDeprecation (file:416).
FLOWFILE_DB_PATH=$PARITY_DB poetry run pytest \
"flowfile_core/tests/flowfile/test_code_generator_edge_cases.py::TestDeprecatedMethodUsage" \
"flowfile_core/tests/flowfile/test_code_generator_edge_cases.py::TestRecordIdDeprecation" \
-q -p no:cacheprovider3 passed, 0 failed (2 guards in TestDeprecatedMethodUsage + 1 round-trip in
TestRecordIdDeprecation; the pair is also 3 passed at baseline — the swap must keep it so).test_record_id_uses_deprecated_with_row_count fails → you swapped the emission but did not flip
the guard TODO at :1002 (they must land together).TestRecordIdDeprecation fails → with_row_index defaults differ in your
Polars pin — check offset handling before shipping.unique(columns=[]) (executor fix — the flow is wrong)This is the only row where the emitter is correct and the flow engine is buggy. The unique node
calls FlowDataEngine.make_unique (flow_graph.py:2415 → flow_data_engine.py:2686). Its guard only
special-cases columns is None; an empty list columns=[] falls into the subset branch and emits
.unique([], keep=...), which Polars rejects ("at least one key is required in a group_by operation").
Fix the executor: treat columns in (None, []) as all-columns → .unique(keep=unique_input.strategy)
(mirror the emitted code; note the current None branch also drops keep= — preserve the strategy).
Then delete the xfail at :459.
Because you are changing runtime engine semantics, you must keep the flow-vs-generated equivalence test as the regression guard (it already exists — that is the point). Do not fix the executor without it.
Verify:
FLOWFILE_DB_PATH=$PARITY_DB poetry run pytest \
"flowfile_core/tests/flowfile/test_code_generator_edge_cases.py::TestUniqueOperationVariations" \
-q -p no:cacheprovider -rX2 passed, 0 xfailed (baseline for this class: 1 passed, 1 xfailed).unique test in the class now fails → your executor change altered the non-empty-columns
path. Scope the fix to the columns == [] branch only.These are the TODO(FlowFrame) markers. They have no round-trip guard today, so the first step for
each is write a failing parity test parameterized on export_flow_to_flowframe, modeled on the
existing test_flowframe_* tests (test_code_generator.py:1326+). Only then fix the converter.
Recommended sub-order (leverage first): F2 (formula pl. import — most common node) → F4 (polars-code
ff.LazyFrame signature) → F1 (right-join .collect().lazy()) → F3 (fuzzy-match Expr repr — rarest).
Each TODO(FlowFrame) comment enumerates 2–3 sanctioned fix options; pick the one that keeps the
Polars path untouched (see the framework-prefix trap in Phase 2).
Verify per fix: run the new parity test (must go red→green) and the full corpus (Phase 1b, must
stay 692+). Do not delete any TODO(FlowFrame) comment until its parity test is green and committed.
For any divergence, pick exactly one and honor its obligation. Ranked by preference.
Fix the emitter — default and preferred when the flow engine produces the correct result and
the generated code is wrong (rows X2, D1, D2, F1–F4). Obligation: a round-trip parity test that
builds the flow, exports, executes, and assert_frame_equals against the engine. Cheapest, lowest
blast radius — you change only exported text.
Fix the FlowFrame executor / engine semantics — correct only when the engine is the buggy side and the emitted code is right (row X1). Obligation: (a) the flow-vs-generated equivalence test as the permanent guard, and (b) a full engine-suite run, because you altered runtime behavior for every flow, not just export. Never change engine semantics purely to make a codegen test green without a flow-level regression test — that is precisely how #544 shipped.
Quarantine with xfail(strict=True) + a real issue — only when the fix is genuinely out of
scope for this change (needs upstream Polars, or a large refactor). Obligation: the marker MUST be
strict=True (so a later fix that makes it pass fails CI and forces marker deletion — no more silent
XPASS rot), and the reason= MUST link a filed GitHub issue, not a flow_graph.py:1055 line ref
(those rot — the current stale reasons cite lines that have since moved). This is a last resort, not
a way to defer work quietly.
Enum-drift corollary (the #544 lesson): if a divergence is a settings string parsed two ways
("desc" vs "Descending", "asc" vs "Ascending", filter-operator spellings), the fix is a single
shared parser consumed by both the emitter and the engine — not two patched copies. The canonical
example is transform_schema.is_descending + SortByInput.descending (transform_schema.py:952-970),
which replaced the drifted how == "desc" comparison in both code_generator.py and the engine's
do_sort. Grep for other hand-rolled direction/enum comparisons before adding a third copy.
.pyi stubs (flowfile_frame/flowfile_frame/*.pyi,
**/*.pyi). They are machine-generated by make stubs and gated by make check_stubs in CI; hand
edits get clobbered and fail the drift gate. If your fix changes the FlowFrame public surface,
regenerate — see flowfile-frame-and-codegen / flowfile-build-and-env.check_dtypes=False/check_exact=False, do not swap assert_frame_equal for a shape check, do not
down-scope assert_flow_result_matches_generated. A weakened assertion is a parity regression that
reads as a pass — worse than the original bug.xfail(strict=True) + an issue link
(menu option 3), never a plain non-strict xfail.self.framework; keep the working path untouched.TODO(FlowFrame) comment before its parity test is green. The comment is the only
marker of a known-broken export.FLOWFILE_DB_PATH isolation — you will migrate/corrupt the
shared catalog DB and burn a later session.Do not open a PR until all of this passes. Success is the numbers, never a visual read.
Codegen edge-case suite is clean (the scoreboard — must show zero xfailed/xpassed for rows you closed):
FLOWFILE_DB_PATH=$PARITY_DB poetry run pytest \
flowfile_core/tests/flowfile/test_code_generator_edge_cases.py \
-q -p no:cacheprovider -rX
# Target after closing S1+X1+X2: 49 passed, 0 xfailed, 0 xpassed
# (46 baseline passed + 3 formerly-xfail/xpass now passing). Restate the exact number you observe.Full parity corpus still green over both exporters (regression floor):
FLOWFILE_DB_PATH=${PARITY_DB}.main poetry run pytest \
flowfile_core/tests/flowfile/test_code_generator.py \
-q -p no:cacheprovider
# Must be ≥ 692 passed (baseline). A drop = parity regression; do not promote.Core engine suite — mandatory whenever you took solution-menu option 2 (executor change; row X1), advisable otherwise:
FLOWFILE_DB_PATH=${PARITY_DB}.core poetry run pytest flowfile_core/tests -m "not kernel" \
-q -p no:cacheprovider
# Success criterion: 0 failed. ~5,079 tests collected as of 2026-07-03 (76 kernel-marked are
# deselected); long run. Autouse fixtures spawn the worker + Docker DB services — prerequisites
# and the SKIP_WORKER_TESTS=1 worker-less form are in flowfile-testing-and-validation.Frame suite — mandatory whenever you worked F-rows or touched anything under
flowfile_frame/ (exported FlowFrame code executes against this API):
FLOWFILE_DB_PATH=${PARITY_DB}.frame poetry run pytest flowfile_frame/tests \
-q -p no:cacheprovider
# Success criterion: 0 failed (~620 collected as of 2026-07-03). Cloud tests skip without
# MinIO on :9000 — prerequisites in flowfile-testing-and-validation.Stub drift gate — mandatory if you touched any FlowFrame public surface (F1–F4 can):
make check_stubs # regenerates and fails on any .pyi git diff; see flowfile-build-and-envLint:
poetry run ruff check flowfile_core flowfile_framePR through flowfile-change-control (branch off main, never force-push, never commit —
stage the diff and hand the maintainer the exact git add/git commit commands per
flowfile-change-control §6). The PR body MUST carry a baseline-vs-after table:
| Suite | Baseline (2026-09-12) | After |
|---|---|---|
| edge-case | 49 passed / 0 xfailed / 0 xpassed | (yours) |
| corpus | 692 passed | (yours) |
List each inventory ID you closed (S1, X2, D1…) and, for any deferred, the xfail(strict=True) +
issue link you added.
Definition of done: zero xfailed AND zero xpassed in the codegen edge-case suite for the rows in
scope, corpus pass-count not regressed, make check_stubs clean, PR shows baseline-vs-after. Anything
short of numbers is not done.
Everything below was verified by reading source or running the command on 2026-07-03 (v0.12.7,
commit f6963c77, branch feature/claude-skills). Re-verify volatile facts before trusting them.
| Fact | Re-verify command |
|---|---|
Baseline 49 passed, 0 xfailed, 0 xpassed (2026-09-12; was 46/2/1) | FLOWFILE_DB_PATH=/tmp/v.db poetry run pytest flowfile_core/tests/flowfile/test_code_generator_edge_cases.py -q -rX | tail -1 |
Corpus 692 passed (2026-09-12; was 653) | FLOWFILE_DB_PATH=/tmp/v2.db poetry run pytest flowfile_core/tests/flowfile/test_code_generator.py -q | tail -1 |
Custom-nodes file 21 passed (2026-09-26; was 20) | FLOWFILE_DB_PATH=/tmp/v5.db poetry run pytest flowfile_core/tests/flowfile/test_code_generator_custom_nodes.py -q | tail -1 |
Per-class gate baselines (2026-09-12): GroupBy 1 passed, Deprecated 2 passed, RecordId 1 passed, Unique 2 passed | class-scoped pytest exactly as written in each Phase 3 gate |
| Corpus dual-exporter parametrization (88 sites) | grep -c 'parametrize("export_func"' flowfile_core/tests/flowfile/test_code_generator.py |
Engine concat delimiter is hard-coded , (string_concat) | sed -n '113,120p' flowfile_core/flowfile_core/schemas/transform_schema.py |
| Emitter emits agg names with no args (X2 mechanism) | grep -n '_get_agg_function' -A4 flowfile_core/flowfile_core/flowfile/code_generator/transform_handlers.py |
X1 executor fix site (make_unique, empty-list branch) | grep -n 'def make_unique' flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py; grep -n 'make_unique' flowfile_core/flowfile_core/flowfile/flow_graph.py |
| The 3 xfail markers (edge-case) at lines 201/459/766 | grep -n 'pytest.mark.xfail' flowfile_core/tests/flowfile/test_code_generator_edge_cases.py |
TODO(FlowFrame) markers (grep for the live set) | grep -rn 'TODO(FlowFrame)' flowfile_core/flowfile_core/flowfile/code_generator/ |
Deprecated emissions: str.concat / with_row_count | grep -rn 'str.concat|with_row_count' flowfile_core/flowfile_core/flowfile/code_generator/ |
IN-filter fix that made S1 stale (splits on , first) | sed -n '144,160p' flowfile_core/flowfile_core/flowfile/filter_expressions.py |
| Shared enum parser (#544 pattern) | grep -n 'def is_descending|def descending' flowfile_core/flowfile_core/schemas/transform_schema.py |
| Export entry points | grep -n 'def export_flow_to_polars|def export_flow_to_flowframe' flowfile_core/flowfile_core/flowfile/code_generator/code_generator.py |
| Round-trip helper | grep -n 'def assert_flow_result_matches_generated' flowfile_core/tests/flowfile/test_code_generator_edge_cases.py |
| Stub gate | grep -n '^stubs:|^check_stubs:' Makefile |
Volatile caveats:
reason= strings themselves cite stale lines like
code_generator.py:1564 and flow_graph.py:1055 — the real sites are now expression_helpers.py:152
and filter_expressions.py:155.)main/feature/claude-skills at f6963c77; a merge that adds parity
tests moves them. Re-run Phase 1 to re-freeze before starting a fix.export_flow_to_flowframe is newer and less complete than export_flow_to_polars; the F-rows are
where that gap concentrates. Do not assume corpus-green means FlowFrame export is complete — F3/F4
have no guard test yet.Expr._repr_str contract) are owned by
flowfile-frame-and-codegen; if the emission path changed, re-read that skill before editing here.© Edwardvaneechoud, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/flowfile-codegen-parity-campaign of Edwardvaneechoud/Flowfile.
Open the folder on GitHubat commit 13aa287
Flowfile Codegen Parity Campaign next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Flowfile Codegen Parity Campaign this skillEdwardvaneechoud/Flowfile | 375 | — | ~7.5k | Automated safety check: Pass | MIT | |
| Hybrid-Engine Data Analysiscode-yeongyu/oh-my-openagent | 70k | — | ~1.4k | Automated safety check: Pass | Custom licence | |
| Optimuskgmims-harvard/OptimusKG | 146 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Narwhalsanam-org/metaxy | 124 | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| PolarsK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.3k | Automated safety check: Pass | MIT | |
| Ingesting Dataancoleman/ai-design-components | 526 | — | ~1.9k | Automated safety check: Pass | MIT |
code-yeongyu/oh-my-openagent
Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.
mims-harvard/OptimusKG
Guide for using OptimusKG, the biomedical knowledge graph, through the optimuskg Python client.
anam-org/metaxy
Effectively use Narwhals to write dataframe-agnostic code that works seamlessly across multiple Python dataframe libraries.
K-Dense-AI/scientific-agent-skills
High-performance DataFrame library for Python ETL, analytics, and pandas migration.
ancoleman/ai-design-components
Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.
ancoleman/ai-design-components
Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).
Edwardvaneechoud/Flowfile
Maps the /ai/ subsystem of flowfile_core, its three agent tiers, litellm seam, BYOK keys and rate limits, and sets rules for extending or debugging it safely.
Edwardvaneechoud/Flowfile
Maps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging.
Edwardvaneechoud/Flowfile
Recreates every Flowfile development and build environment from scratch, with exact version pins and an explanation of what each Makefile target really does.
Edwardvaneechoud/Flowfile
Explains how changes to the Flowfile monorepo are gated, versioned and released, including version sync, stub and docs drift checks, Alembic migrations and pinned dependencies.
Edwardvaneechoud/Flowfile
Catalog of Flowfile's environment variables and runtime flags: what each does, where the code reads it, its default, and where the docs disagree with the code.
Edwardvaneechoud/Flowfile
Turn a plain-English description ("a node that runs on the kernel and does XGBoost predictions", "a node that trims whitespace", "an ML clustering node") into a correct single-file Flowfile custom…
Categories
Runbook for closing gaps between a Flowfile visual flow's results and its exported Polars or FlowFrame Python code, measured by tests rather than by eye. Flowfile runs the same node settings through two executors: the live engine behind the Designer and a code generator that exports standalone Python through export_flow_to_polars and export_flow_to_flowframe. This skill is a decision-gated campaign to find and close cases where the exported code runs cleanly but returns a different answer.
Flowfile Codegen Parity Campaign fits situations like: fixing exported Python that disagrees with the flow result; closing xfail markers in the codegen edge-case tests; working in the code_generator folder of flowfile_core; deleting a stale xfail or XPASS marker.
Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-codegen-parity-campaign -a claude-code`. Or copy the skill folder (.claude/skills/flowfile-codegen-parity-campaign in Edwardvaneechoud/Flowfile) into .claude/skills/flowfile-codegen-parity-campaign in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-codegen-parity-campaign -a codex`. Or copy the skill folder (.claude/skills/flowfile-codegen-parity-campaign in Edwardvaneechoud/Flowfile) into .agents/skills/flowfile-codegen-parity-campaign in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-codegen-parity-campaign -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flowfile-codegen-parity-campaign, .gemini/skills/flowfile-codegen-parity-campaign, .github/skills/flowfile-codegen-parity-campaign and .opencode/skills/flowfile-codegen-parity-campaign in your project.
Going by SKILL.md and its folder, Flowfile Codegen Parity Campaign needs the command-line tools its instructions call (poetry, make and git). Our summary lists: A checkout of the Flowfile repository; Python with Polars and the project's test suite; make, for the check_stubs target.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Flowfile Codegen Parity Campaign is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Flowfile Codegen Parity Campaign: Hybrid-Engine Data Analysis (code-yeongyu/oh-my-openagent, 70k stars), Optimuskg (mims-harvard/OptimusKG, 146 stars), Narwhals (anam-org/metaxy, 124 stars) and Polars (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Edwardvaneechoud (a GitHub user) maintains it in Edwardvaneechoud/Flowfile, which has 375 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 9, 2026.
Source: Edwardvaneechoud/Flowfile on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.