Agent skill

Flowfile Testing And Validation

by Edwardvaneechoud in Edwardvaneechoud/Flowfile

Exact per-package pytest/vitest/playwright commands, the registered pytest markers and which need Docker, the testutils Docker fixture matrix, the shared-test-DB isolation model and its failure…

MITAuto-check: notesTesting & QA

Install Flowfile Testing And Validation

skills CLI
$ npx skills add Edwardvaneechoud/Flowfile --skill flowfile-testing-and-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Edwardvaneechoud/Flowfile flowfile-testing-and-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Edwardvaneechoud/Flowfile.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/flowfile-testing-and-validation .claude/skills/flowfile-testing-and-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
flowfile-testing-and-validation
GitHub stars
373
Token cost
~8.3k tokens
SKILL.md length
3,620 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
MIT

At a glance

Exact per-package pytest/vitest/playwright commands, the registered pytest markers and which need Docker, the testutils Docker fixture matrix, the shared-test-DB isolation model and its failure…

  • Works in 10 steps: The single pytest config, and why bare… → Run each suite — exact commands and… → test_utils/ — the Docker fixture matrix → …
  • Diagnosing no such table
  • SKILL.md covers When NOT to use this skill, 1. The single pytest config,…, 2. Run each suite — exact… and 3. test_utils/ — the Docker…, plus 8 more sections
  • Calls poetry, make and npm

What it does

Flowfile Testing And Validation is an agent skill from Edwardvaneechoud/Flowfile. Exact per-package pytest/vitest/playwright commands, the registered pytest markers and which need Docker, the testutils Docker fixture matrix, the shared-test-DB isolation model and its failure modes, xfail/skip discipline, and coverage/CI test-matrix mechanics for the Flowfile monorepo. Use when running or writing tests, diagnosing "no such table" or phantom test failures, deciding whether a change is "validated," seeing XPASS in test output, wiring a new Docker-backed fixture, or asking "which suite proves this…

Its SKILL.md is about 8.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Unit testing, Containers and Monorepo tooling. It works with pytest, Docker, Playwright and Vitest. The repository describes itself as: Flowfile is a visual ETL tool and Python library combining drag-and-drop workflows with Polars dataframes. Build data pipelines visually, define flows programmatically with a… The licence is MIT.

When your agent uses it

  • Diagnosing no such table
  • Phantom test failures
  • Deciding whether a change is validated
  • Seeing XPASS in test output

Example prompts

  • “no such table”
  • “validated,”
  • “which suite proves this change works.”
  • “/flowfile-testing-and-validation”

Requirements

  • Python 3
  • Node.js
  • Docker

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. The single pytest config, and why bare pytest is dangerous
  2. Run each suite — exact commands and prerequisites
  3. test_utils/ — the Docker fixture matrix
  4. State isolation model — and where it breaks
  5. Coverage
  6. Evidence bar — what "validated" means per change class
  7. xfail / XPASS discipline
  8. CI test matrix summary (test.yaml)
  9. History — read this before "fixing" CI test speed again
  10. Slow / flaky areas

What it can do on your machine

Read from SKILL.md and the folder at commit d98b76d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • poetry
    • make
    • npm
    • git
    • pytest
    • npx
    • pip
    • python
    • node
    • docker
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, git, npx, pip, docker and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flowfile Testing And Validation loads about 8.3k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 3,620 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~142
When it runs · the whole SKILL.md, loaded when a task matches
~8.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:13
    ndful that control test isolation (full `.env` catalog, feature flags) → `flowfile-config-and-flags`.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Edwardvaneechoud/Flowfile at commit d98b76d, republished under its MIT licence (© Edwardvaneechoud). 3,620 words, ~8,255 tokens.

Download SKILL.mdSave it as .claude/skills/flowfile-testing-and-validation/SKILL.md (or your agent's skills folder).
name
flowfile-testing-and-validation
description
Exact per-package pytest/vitest/playwright commands, the registered pytest markers and which need Docker, the test_utils Docker fixture matrix, the shared-test-DB isolation model and its failure modes, xfail/skip discipline, and coverage/CI test-matrix mechanics for the Flowfile monorepo. Use when running or writing tests, diagnosing "no such table" or phantom test failures, deciding whether a change is "validated," seeing XPASS in test output, wiring a new Docker-backed fixture, or asking "which suite proves this change works."

Flowfile Testing and Validation

When NOT to use this skill

  • Writing a new backend node or its tests → flowfile-node-development.
  • Debugging a specific failure you already reproduced (stack trace triage, log reading) → flowfile-debugging-playbook.
  • CI workflow authoring/release mechanics beyond test.yaml's test jobs (tagging, PyPI/Tauri release, Docker publish, branch protection) → flowfile-change-control (release/tag mechanics, branch protection, publish pipelines).
  • Env var reference beyond the handful that control test isolation (full .env catalog, feature flags) → flowfile-config-and-flags.
  • Known bugs / historical incidents not related to test infra itself → flowfile-failure-archaeology.
  • Local dev server startup, ports, Docker Compose for running the app (not testing it) → flowfile-run-and-operate / flowfile-build-and-env.

If you just need "how do I run the core tests" — that's this skill, keep reading.


1. The single pytest config, and why bare pytest is dangerous

pyproject.toml [tool.pytest.ini_options] is the only pytest config for the monorepo (exception: flowfile_wasm/pytest.ini, scoped so cd flowfile_wasm && pytest runs only tests/python). There is no addopts, no testpaths, no norecursedirs.

Consequence: running bare pytest from repo root collects everything, including slow Docker-compose E2E tests (tests/integration), Kafka tests (tests/kafka) and the cloud storage stack tests (tests/cloud_e2e). Any doc claiming "pytest excludes these by default" is wrong at the config level — they're only skipped in practice because everyone targets a specific directory.

Rule: always pass an explicit directory. poetry run pytest flowfile_core/tests, never bare poetry run pytest.

The registered markers (as of 2026-09-23; re-check pyproject.toml)
toml
markers = [
    "worker: Tests for the flowfile_worker package",
    "core: Tests for the flowfile_core package",
    "kernel: Integration tests requiring Docker kernel containers",
    "docker_integration: Full Docker-based E2E tests (require Docker, slow)",
    "kafka: Integration tests requiring a Kafka/Redpanda broker (Docker)",
    "lsp: Tests for the notebook LSP (Jedi) code-intelligence surface",
    "slow: Tests with a heavy workload or long runtime (deselect with -m 'not slow')",
    "cloud_e2e: Real core + worker processes driven over HTTP against MinIO (Docker)",
]
MarkerNeeds Docker?Where used
workerNo (marker only)flowfile_worker package tests
coreNo (marker only)flowfile_core package tests
kernelYesflowfile_core/tests -m kernel (76 tests); builds/runs kernel containers
docker_integrationYestests/integration — full docker-compose E2E
kafkaYestests/kafka, shared/tests/kafka — needs Redpanda
lspNoflowfile_core/tests/lsp/test_lsp_routes.py (hermetic); test_lsp_kernel_integration.py also carries @pytest.mark.kernel
slowNoflowfile_worker/tests/test_catalog_visualize.py, flowfile_core/tests/test_kernel_dependency_gate.py
cloud_e2eYes (MinIO)tests/cloud_e2e — spawns its own core + worker per session; skips locally without MinIO, fails when CI is set

requires_yaml (and a local slow) is registered only by tools/migrate/tests/conftest.py.


2. Run each suite — exact commands and prerequisites

All commands run from the repo root through the single Poetry env; there is no tox/nox.

SuiteCommandPrerequisites / notes
corepoetry run pytest flowfile_core/testsAutouse fixtures spawn worker + Postgres + MySQL. CI form: -m "not kernel". Worker-less: prefix SKIP_WORKER_TESTS=1. ~5,079 tests collected as of 2026-07-03 (v0.12.7), 76 of them kernel-marked.
workerpoetry run pytest flowfile_worker/testsconftest sets TEST_MODE=1; Postgres on :5433 autouse (hard-fails if Docker is present but the container won't start). Cloud tests need MinIO/GCS/Azurite up. ~311 collected.
framepoetry run pytest flowfile_frame/testsconftest registers a minio-flowframe-test cloud connection at import; cloud tests need MinIO on :9000. ~620 collected.
schedulerpoetry run pytest flowfile_scheduler/testsFully hermetic — tmp SQLite per test, spawn stubbed, clock pinned. ~13 collected.
shared (no kafka)poetry run pytest shared/tests --ignore=shared/tests/kafka~89 collected.
shared kafkapoetry run pytest shared/tests/kafka/ -vsession-autouse conftest starts/reuses Redpanda; skips without Docker.
kafka integrationpoetry run pytest tests/kafka -m kafkaNeeds Redpanda (auto-managed) + a running worker.
cloud storage E2Epoetry run pytest tests/cloud_e2e -m cloud_e2eMinIO on :9000 (started if Docker is up). Spawns its own core + worker stacks per session on free ports with an isolated DB, storage, secure store, HOME and AWS profile (safe next to a live stack; never imports flowfile_core). Connection-based tests run on a stack whose ambient AWS_ENDPOINT_URL is a dead port, so only a connection's own endpoint/allow-HTTP/profile reaches MinIO; "No connection" tests run on a second stack whose ambient profile and endpoint are MinIO's. Asserts the servers' working directories stay empty. A couple of minutes. POSIX only.
kernel (core-side)poetry run pytest flowfile_core/tests -m kernel -vBuilds flowfile-kernel image unless FLOWFILE_KERNEL_IMAGE is preset; needs Docker. ~76 tests.
kernel_runtime unitpoetry run pytest kernel_runtime/testsNo Docker — TestClient only. ~327 collected.
docker E2Epoetry run pytest tests/integration -m docker_integration -vDocker + compose v2; ports 63578/63579 must be free or tests skip; builds core+worker+kernel images (minutes).
auth Docker E2Epoetry run pytest flowfile_core/tests/test_auth_e2e.py -v -sBuilds real core image via the docker SDK directly (not compose); skips without Docker.
migrate toolpoetry run pytest tools/migrate/tests~64 collected; hermetic.
flowfile CLIpoetry run pytest flowfile/teststest_api.py boots a real server via start_flowfile_server_process(); needs free ports.
coveragemake test_coverageCore+worker sequential, --cov-append (see §5).
frontend unitcd flowfile_frontend && npm run test:unitVitest, node env, no jsdom/happy-dom. ~30 test files. Watch mode: npm run test:unit:watch.
frontend E2Ecd flowfile_frontend && npm run test:web (web-flow only) or npm run test:all (adds canvas-overlays)No webServer block in playwright.config.ts — core (:63578) and a Vite server must already be running. npx playwright install chromium first. Single worker, 2 retries in CI.
frontend E2E orchestratedmake test_e2e / make test_e2e_devBuilds/starts core + worker + preview(:4173) or dev(:8080), runs web-flow.spec.ts + csp.spec.ts, then stops whatever listens on 63578/63579/8080/4173. On macOS/Linux the target exits with Playwright's status; the Windows branch still ends in || true, so read the Playwright output there.
cloud storage E2E orchestratedmake test_e2e_cloud (or cd flowfile_frontend && API_URL=… TEST_URL=… npm run test:cloud against a disposable stack; it skips without API_URL)macOS/Linux + Docker. Starts and seeds MinIO (poetry run seed_cloud_e2e → s3://flowfile-test/cloud-e2e/source.parquet), an isolated core/worker/vite-preview on free ports, then tests/cloud_e2e and cloud-storage-flow.spec.ts; kills only its own PIDs and fails if a server wrote into its working dir. The spec's "No connection" test runs only with E2E_AWS_PROFILE_CONFIGURED=1.
wasm JScd flowfile_wasm && npm run test:runVitest, happy-dom, globals on. ~348 cases.
wasm Python enginecd flowfile_wasm && pip install -r tests/python/requirements.txt && python -m pytest tests/pythonOwn pytest.ini. Pins polars==1.18.0 / pydantic==2.10.5 / polars-expr-transformer==0.6.0 — the exact Pyodide 0.27.7 versions. Running through the monorepo Poetry env resolves a different Polars; use the pinned env for parity. ~85 test fns.
wasm Pyodide smokecd flowfile_wasm && npm install --no-save pyodide@0.27.7 parquet-wasm@0.7.1 && node tests/pyodide-smoke/smoke.cjsThe only guard for browser-namespace/bootstrap breakage — CPython tests can't catch it.

Worked example — run only core tests, isolated from any other pytest session, without needing a worker:

bash
FLOWFILE_DB_PATH=/tmp/ff_core_$$.db SKIP_WORKER_TESTS=1 \
  poetry run pytest flowfile_core/tests -m "not kernel" -q

3. test_utils/ — the Docker fixture matrix

Package layout: test_utils/{postgres,mysql,s3,gcs,azurite,kafka}/, each with fixtures.py + commands.py. Every start_*/stop_* command is a Poetry script (poetry run start_postgres, poetry run stop_postgres, etc.) and all start_* commands exit 0 when Docker is missing — "return success to allow pipeline to continue" — so downstream tests just skip rather than the CI step failing.

ServiceContainer nameHost port(s)Started bySkip behavior
Postgrestest-postgres-sample5433→5432poetry run start_postgres; core+worker conftest autouse (reuses if already listening)is_docker_available() False → tests skipif; if Docker present but start fails → pytest.fail (core), soft-skip (worker uses same pattern but less strict)
MySQLtest-mysql-sample3307→3306poetry run start_mysql; core conftest autouseCore: soft-fail with a log message, tests skip; image pull (only when the tag is absent locally, test_utils/docker_images.py) can take up to 300s first time
MinIO (S3)test-minio-s39000 API, 9001 consolepoetry run start_minio; poetry run seed_cloud_e2e (test_utils/s3/cloud_e2e_seed.py) idempotently seeds the cloud E2E source_minio_available() / requires_minio guards; frame conftest assumes :9000. Never write test data under the pre-existing sample-data/ bucket — use a unique prefix you delete
GCStest-fake-gcs4443poetry run start_gcsis_gcs_available() guard; also re-populates data if container is up but empty
Azuritetest-azurite10000 (blob)poetry run start_azuriteis_azurite_available() guard; well-known devstoreaccount1 creds hardcoded
Redpanda (Kafka)test-redpanda-kafka19092→9092poetry run start_redpanda; tests/kafka + shared/tests/kafka conftest autouseSkips without Docker; topics are UUID-suffixed per test so container reuse is safe

Shared skip logic (every service's is_docker_available()):

  1. On macOS or Windows when CI env is truthy → returns False (Docker treated as unavailable on non-Linux CI runners).
  2. Otherwise: shutil.which("docker") must exist and docker info must exit 0 within 5s.

KEEP_MINIO_RUNNING / KEEP_GCS_RUNNING / KEEP_AZURITE_RUNNING / KEEP_REDPANDA_RUNNING (=true) keep a container alive after the managed context exits, for debugging.

flowfile_core/tests/flowfile_core_test_utils.py::is_docker_available() additionally returns False on Windows unconditionally (not just in CI).


4. State isolation model — and where it breaks

One SQLite DB per mode, resolved by shared/storage_config.py::get_database_url() (priority order, verified in code):

  1. FLOWFILE_DATABASE_URL, else FLOWFILE_DB_PATH (path → sqlite:///…, :// values used as-is)
  2. TESTING=True env var (exact string) → sqlite:///<storage.temp_directory>/test_flowfile_catalog.db — a fixed shared path
  3. Default → sqlite:///<storage.database_directory>/flowfile_catalog.db (the live DB)

Core conftest sets os.environ['TESTING'] = 'True' at import (flowfile_core/tests/conftest.py:23), before any flowfile_core import, so the test DB resolves correctly. Frame conftest does the same (flowfile_frame/tests/conftest.py:3).

Fragile point #1 (the one that will burn you): fixed shared test-DB path

Because step 2's path is fixed (not per-process), two concurrent pytest sessions on the same machine share test_flowfile_catalog.db. The session-scoped autouse setup_test_db fixture (flowfile_core/tests/conftest.py) calls init_db() on setup and, on teardown, does Base.metadata.drop_all(engine) then deletes the DB file. If session A tears down while session B is mid-run, B's tables vanish out from under it.

Symptom: a cascade of sqlite3.OperationalError: no such table: <catalog_table_read_links|users|...> that looks like your change broke 35 unrelated tests. It didn't — a second concurrent pytest process (yours from an earlier terminal, a background CI-simulation run, anything) tore down the shared DB mid-test.

The fix — highest-priority override, use it for every isolated run:

bash
FLOWFILE_DB_PATH=/tmp/ff_$$.db poetry run pytest flowfile_core/tests -m "not kernel"

FLOWFILE_DB_PATH wins over TESTING in get_database_url(), so this fully isolates the DB per invocation. It also propagates to the worker subprocess the core conftest spawns (the worker inherits the pytest process env), so cross-boundary core↔worker tests stay isolated too — this is why the override is "complete," not partial.

Before blaming a code change for a wall of table-not-found errors, run ps aux | grep pytest and check for a second session.

Fragile point #2: never skip the startup migration for the test suite

flowfile_core/flowfile_core/database/init_db.py runs run_startup_migration() at import time unless FLOWFILE_SKIP_STARTUP_MIGRATION is set:

python
if not os.environ.get("FLOWFILE_SKIP_STARTUP_MIGRATION"):
    run_startup_migration()

The core test suite's setup_test_db fixture depends on that import-time migration to create the schema in the first place. If you set FLOWFILE_SKIP_STARTUP_MIGRATION=1 while running flowfile_core/tests against a fresh/isolated FLOWFILE_DB_PATH, every test errors with no such table: users — there was never a migration run to create the tables.

FLOWFILE_SKIP_STARTUP_MIGRATION=1 is for diagnostics only — e.g. importing flowfile_core in a scratch script to inspect something without touching a real DB's migration stamp. Pair it with its own throwaway FLOWFILE_DB_PATH in that case, and never use it for pytest flowfile_core/tests.

bash
# WRONG — will error "no such table: users" on every test
FLOWFILE_DB_PATH=/tmp/fresh.db FLOWFILE_SKIP_STARTUP_MIGRATION=1 poetry run pytest flowfile_core/tests

# RIGHT — isolated DB, migration allowed to run and build the schema
FLOWFILE_DB_PATH=/tmp/fresh.db poetry run pytest flowfile_core/tests
SKIP_WORKER_TESTS=1

flowfile_core/tests/conftest.py's session-autouse flowfile_worker fixture checks os.environ.get("SKIP_WORKER_TESTS") == "1" and no-ops if set (no worker spawned, no reuse-probe). The execution_location fixture (parametrized ["local", "remote"], used by catalog/flow-API/run-node tests) then auto-skips its remote half. Use this to run core tests fast without a worker; do not use it when validating a core↔worker contract change (§6 below).

FLOWFILE_TEST_REUSE_WORKER=1

With FLOWFILE_WORKER_PORT unset and something already answering on the default 63579 (a dev worker or the desktop app's sidecar), conftest.py::_claim_worker_port moves the session to a free port at import and the flowfile_worker fixture starts a worker from this checkout there with --port; pytest_report_header prints worker: port 63579 is taken, this session's worker uses N. Set FLOWFILE_TEST_REUSE_WORKER=1 (exactly "1") to reuse the running worker instead. An explicit FLOWFILE_WORKER_PORT is used as given, and a worker already listening on it is reused. The suite's worker still sends its logs to CORE_PORT (default 63578), so give the run a private CORE_PORT when a live core is up.

Other fragile points (verified in code, worth knowing)
  • Catalog seed erosion: many core test modules call a catalog_cleanup() helper that wipes all CatalogNamespace rows, including the init_db-seeded 'General' namespace. flowfile_core/tests/project/conftest.py has an autouse re-seed specifically to paper over this — a new suite depending on seeded catalog rows needs the same treatment.
  • Worker virtual-result cache: catalog_cleanup() also deletes .arrow files under the worker's virtual-results directory, because table-id recycling would otherwise let a stale cached file satisfy the next test's resolve.
  • Session-global services: Postgres (:5433), MySQL (:3307), Redpanda are reused if already listening — a dirty long-running instance from a previous session can leak state into a new run. The worker is the exception: see FLOWFILE_TEST_REUSE_WORKER=1 above.
  • Process-wide env mutation: the sharing test suite flips FLOWFILE_MODE=docker via monkeypatch.setenv per test only — several core test modules construct a TestClient and mint auth tokens at import time under electron mode; flipping the mode process-wide before those imports breaks them.
  • bcrypt monkeypatch: core conftest patches bcrypt.hashpw at import to truncate >72-byte passwords (passlib/bcrypt compat) — password-hashing tests behave differently from a prod bcrypt install without it.

5. Coverage

  • Source is core + worker only: [tool.coverage.run] source = ["flowfile_core/flowfile_core", "flowfile_worker/flowfile_worker"] — frame, scheduler, and shared are deliberately excluded from coverage.
  • fail_under = 0 — coverage never gates a build by threshold.
  • omit: */tests/*, */test_*, */__pycache__/*, */conftest.py.
  • No branch=true, no dynamic_context — deliberate, because CI sets COVERAGE_CORE=sysmon (PEP-669 sys.monitoring tracer), which doesn't support those options.

Local:

bash
make test_coverage
# = poetry run pytest flowfile_core/tests --cov --cov-report= --disable-warnings
#   poetry run pytest flowfile_worker/tests --cov --cov-append --cov-report= --disable-warnings
#   poetry run coverage report --show-missing

Core and worker run sequentially with --cov-append — the Makefile comment explains this avoids import collisions from both packages loading into one coverage-tracked process.

CI's dedicated coverage job (ubuntu, Python 3.12) sets COVERAGE_CORE: sysmon, runs core -m "not kernel" then worker, uploads coverage.xml/.coverage as an artifact, and pushes to Codecov (flags: backend, fail_ci_if_error: false — Codecov never blocks CI either). There is no codecov.yml in the repo; behavior is Codecov defaults.


Show full SKILL.md (1,563 more words)Show less

6. Evidence bar — what "validated" means per change class

A green run is only as strong as what actually executed. Docker-gated suites silently skip without Docker/MinIO/etc — always ask "did the Postgres/MinIO/kernel tests actually run, or did they skip?" (check the pytest summary line for skip counts, not just "passed").

Change classMinimum bar
Core/worker/shared/frame backend changeRun the owning package's suite with an isolated DB: FLOWFILE_DB_PATH=/tmp/ff_$$.db poetry run pytest <pkg>/tests. Check skip counts didn't balloon vs. a baseline run.
Core ↔ worker contract changeRun without SKIP_WORKER_TESTS — the real worker subprocess must be exercised; execution_location-parametrized tests need it for their remote half.
DB schema changeAdd flowfile_core/flowfile_core/alembic/versions/NNN_*.py, then poetry run pytest flowfile_core/tests/test_migration.py (builds DBs from scratch via FLOWFILE_DB_PATH).
Kernel-touching changepoetry run pytest flowfile_core/tests -m kernel locally with Docker running.
flowfile_frame public API changemake stubs and stage the .pyi diff (do not commit it yourself — hand off per flowfile-change-control's no-agent-commit policy) — make check_stubs is a hard CI gate (regenerates then git diff --exit-code).
Kafka path changepoetry run pytest tests/kafka -m kafka plus shared/tests/kafka.
Frontend renderer changenpm run test:unit + npm run build:web (lint + vue-tsc --noEmit run inside the build script) — that's what CI's test-web job enforces. Canvas/flow behavior changes additionally need make test_e2e (on Windows read its Playwright output — that branch still ignores the exit code, see §2).
Cloud storage node / connection changepoetry run pytest shared/tests/test_cloud_storage_options.py plus tests/cloud_e2e -m cloud_e2e (real core + worker over HTTP); UI-facing changes also make test_e2e_cloud. Tests must not reach real AWS: temp AWS_SHARED_CREDENTIALS_FILE/AWS_CONFIG_FILE, AWS_EC2_METADATA_DISABLED=true, and an explicit endpoint — test_utils/s3/aws_profiles.py::isolate_aws sets all of that up.
WASM engine changePinned-env pytest (tests/python) and the Pyodide smoke test — a green CPython run does not prove the browser namespace still works.
Full-stack / deploy-shaped changepoetry run pytest tests/integration -m docker_integration -v with ports 63578/63579 free.

7. xfail / XPASS discipline

Current inventory (verified live 2026-09-12 — re-run before trusting, see §9):

LocationWhat it claimsLive status
flowfile_wasm/tests/python/test_build_helpers.py::test_filter_advanced_expr_does_not_evaluate_pythonpolars-expr-transformer evals a crafted formula (standardize_quotes requotes 'a"b' unescaped, Classifier.get_pl_func evals it; to_polars_code's _validate_polars_code is a second sink)xfail(strict) — real upstream bug. Verified 2026-09-12 that 0.5.7 and 0.6.0 ship byte-identical standardize_quotes and the same eval, so a pin bump does not close it; the fix belongs in the upstream library (same maintainer).

Closed on 2026-09-12 (markers deleted, root causes fixed, branch fix/xfails): the three codegen markers in test_code_generator_edge_cases.py — test_in_operator_numeric (stale XPASS), test_unique_without_columns (engine make_unique now treats columns=[] as all-columns and keeps keep=strategy), test_groupby_with_concat_aggregation (emitter emits str.join(',') via the shared transform_schema.STRING_CONCAT_DELIMITER) — plus the two xfail(strict) scanner evasions in community_nodes/test_security_scan.py (scanner hardened: cross-method self.<attr> decode taint, operator.attrgetter/methodcaller rule; fixtures promoted from evade/ into deny/). The node-designer TestNumericStringAliasBug marker was already gone.

Rule for this repo: XPASS means the xfail is stale. Delete the marker (and the outdated bug description) — never leave it, never "celebrate" the pass. When you add a new xfail for a real known bug, prefer @pytest.mark.xfail(reason=..., strict=True) so a future fix turns into a hard CI failure demanding the marker's removal, instead of a silent XPASS nobody notices.

Verification recipe (the remaining marker lives in the DB-free WASM engine tests):

bash
poetry run pytest \
  "flowfile_wasm/tests/python/test_build_helpers.py::test_filter_advanced_expr_does_not_evaluate_python" \
  -q -p no:cacheprovider -rX
# → "1 xfailed"; an XPASS means the upstream fix landed — raise the pin everywhere and delete the marker

Other skip inventory:

  • flowfile_core/tests/flowfile/test_basic_filter.py:751,761 — unconditional pytest.mark.skip, "Manual input converts None to string; test requires actual null values from file sources." Product limitation, not test debt: is_null/is_not_null are effectively untested via manual-input at integration level.
  • flowfile_worker/tests/test_train_apply_model.py:159 — benign parametrize carve-out (logistic_regression/knn_classifier need 0/1 targets; covered by a dedicated round-trip test elsewhere).
  • No .skip/.todo/.fixme exist in any Playwright or Vitest spec (frontend or wasm) — the TS suites are clean of this pattern.
  • The dominant "skip" pattern by volume is environment skipif(not is_docker_available()) across core/worker/frame — dozens of uses. This is a coverage cliff, not debt: on a laptop without Docker+MinIO+emulators running, a large slice of integration surface silently skips, and a green local run is weak evidence of anything Docker-touching.

8. CI test matrix summary (test.yaml)

  • Concurrency: group ${{ github.workflow }}-${{ github.ref }}, cancel-in-progress: ${{ github.event_name == 'pull_request' }} — PR runs cancel their own superseded runs; main-branch runs are never cancelled (docker-publish/release pipelines key off completed main builds). This is the only workflow file in the repo with a concurrency block.
  • detect-changes (dorny/paths-filter) gates every downstream job by which paths changed; workflow_dispatch input run_all_tests: true forces everything regardless.
  • backend-tests matrix: fail-fast: false; ubuntu-latest × Python 3.10/3.11/3.12/3.13, plus macos-latest × 3.11. Starts Postgres/MySQL/MinIO/Azurite/GCS via the poetry run start_* scripts (no-op on macOS CI runners — Docker reports unavailable there). Core runs -m "not kernel". Linux entries then run tests/cloud_e2e -m cloud_e2e after the worker tests, while MinIO is still up (gated on core/worker/shared/tests/cloud_e2e/workflow changes).
  • coverage: separate ubuntu/3.12 job, COVERAGE_CORE=sysmon (see §5).
  • backend-tests-windows: windows-latest, Python 3.11 only, pwsh shell.
  • kernel-tests: ubuntu/3.11, 15-min timeout, builds the kernel image, runs kernel_runtime unit tests then flowfile_core/tests -m kernel.
  • check-stubs / check-formula-docs: drift gates — regenerate .pyi stubs / functions.md and fail on git diff.
  • test-web: Vitest unit + build:web + preview-server curl check.
  • docs-test: mkdocs build.
  • test-summary: if: always(), aggregates all jobs except version-sync, fails if any non-skipped job failed. Treat this as the real pass/fail signal for the whole run — but note it is not wired into required branch-protection checks (a CI-mechanics fact, not this skill's territory beyond flagging it).

Separate, path-filtered workflows cover what test.yaml doesn't: e2e-tests.yml (Playwright web E2E; also starts and seeds MinIO, runs core + worker from empty runner.temp dirs with a MinIO-only AWS profile, and fails if either wrote into its working dir), test-docker-auth.yml, test-kernel-integration.yml, test-docker-kernel-e2e.yml, test-kafka-integration.yml, flowfile-wasm-build.yml.


9. History — read this before "fixing" CI test speed again

The backend-tests (ubuntu, 3.12) job was ~56 minutes before 2026-06-20, caused by two compounding factors: coverage's default C-tracer roughly doubling runtime, and the core suite (~5k tests) running fully serially. Shipped fix (merged, do not re-litigate the same options without reading this first):

  1. COVERAGE_CORE=sysmon — switched the coverage job to the PEP-669 sys.monitoring tracer, near-zero overhead on 3.12 vs. the old C-tracer's ~2x tax.
  2. Coverage split into its own dedicated job — off the critical path of the functional matrix; the plain backend-tests matrix jobs run without --cov at all.
  3. Dropped redundant frontend builds from backend jobs.
  4. concurrency: cancel-in-progress for PR runs.

pytest-xdist was evaluated and explicitly deferred — do not casually re-propose it. Hard prerequisites that would need to be solved first, all rooted in the same fixed-DB-path problem as §4:

  • Each xdist worker needs its own FLOWFILE_DB_PATH, derived from PYTEST_XDIST_WORKER, and that derivation must happen before the first flowfile_core import — the DB engine binds to a URL at import time, so setting the env var after import is a no-op.
  • If used on the coverage job: parallel = true in [tool.coverage.run] plus a coverage combine step, neither of which exist today.
  • Expected speedup is ~1.4–1.8x, not 2–4x, because several fixtures are shared-service singletons (the worker on :63579, Postgres on :5433, MySQL on :3307) that don't parallelize cleanly across workers without further isolation work.

Without those prerequisites, naive -n auto reproduces exactly the "no such table" cascade from §4 fragile point #1, at coverage-job scale — an empty or garbaged coverage.xml is the typical failure mode.

Load-bearing sleeps — do not remove these thinking they're dead time:

  • Catalog test time.sleep(1.05) calls exist because SQLite's updated_at column has 1-second granularity; a faster sleep produces flaky ordering assertions.
  • 2-second sleeps around cancel-flow tests are similarly timing-load-bearing.

10. Slow / flaky areas

  • Core suite is the long pole (~5k tests, serial) even post-fix; expect the 3.12 matrix job around 28–31 minutes as of the 2026-06-20 fixes.
  • Docker image builds dominate -m kernel (mitigate by presetting FLOWFILE_KERNEL_IMAGE to skip the ~30s build) and -m docker_integration (builds core+worker+kernel; conftest uses 600s build timeouts).
  • MySQL container start can take up to 60s; first-time image pull up to 300s.
  • Worker viz tests are tagged @pytest.mark.slow (flowfile_worker/tests/test_catalog_visualize.py); deselect with -m 'not slow'.
  • Playwright: retries: 2 in CI plus trace/video on first retry — this retry budget can mask genuine flakes; workers: 1 avoids port conflicts, not a performance choice.
  • A trailing-slash axios/FastAPI mismatch has historically caused silent failures only in Docker (Vite's dev proxy and pytest's TestClient both mask a 307 redirect that Docker's real network path doesn't) — if a route "works locally but not in Docker," check for a slash mismatch in core logs, not test logic.

Provenance and maintenance

Volatile facts below need periodic re-verification — commands are copy-pasteable.

  • Pytest markers list (as of 2026-07-03, v0.12.7): grep -A8 '\[tool.pytest.ini_options\]' pyproject.toml
  • Coverage config: grep -A20 '\[tool.coverage.run\]' pyproject.toml
  • FLOWFILE_DB_PATH / TESTING priority order: sed -n '400,420p' shared/storage_config.py (look for get_database_url)
  • Startup-migration skip gate: sed -n '1,30p' flowfile_core/flowfile_core/database/init_db.py
  • SKIP_WORKER_TESTS wiring: grep -n "SKIP_WORKER_TESTS\|def flowfile_worker" flowfile_core/tests/conftest.py
  • Docker fixture ports: grep -n "_PORT = int(os.environ.get" test_utils/*/fixtures.py
  • Test-utils Poetry scripts: grep -n '^start_\|^stop_' pyproject.toml
  • xfail inventory + live status (re-run periodically — bugs get fixed and markers go stale silently, that's the whole point of §7; grep -rn "pytest.mark.xfail" --include=*.py . | grep -v "/.claude/" lists every marker):
    bash
    poetry run pytest \
      "flowfile_wasm/tests/python/test_build_helpers.py::test_filter_advanced_expr_does_not_evaluate_python" \
      -q -p no:cacheprovider -rX
  • Collect counts (~5,079 core / ~311 worker / ~620 frame / ~13 scheduler as of 2026-07-03): poetry run pytest flowfile_core/tests --collect-only -q | tail -3 (repeat per package)
  • CI job list, concurrency block, matrix versions: sed -n '1,120p' .github/workflows/test.yaml and grep -n "python-version:" .github/workflows/test.yaml
  • Coverage job env/steps: sed -n '216,296p' .github/workflows/test.yaml
  • 56-min → ~28-31-min history and xdist deferral: no single file encodes this — cross-check against .github/workflows/test.yaml coverage-job comments (sed -n '216,220p') which corroborate the sysmon rationale; the xdist-deferral reasoning is institutional knowledge captured here, verify by searching git log --oneline --all -- .github/workflows/test.yaml for the speed-fix commit if it needs re-confirming.
  • Required branch-protection checks / whether test-summary gates merges: out of this skill's scope — verify via gh api repos/<org>/Flowfile/branches/main/protection if needed for a CI-mechanics task.

© Edwardvaneechoud, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/flowfile-testing-and-validation of Edwardvaneechoud/Flowfile.

Open the folder on GitHubat commit d98b76d

Compare with similar skills

Flowfile Testing And Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flowfile Testing And Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flowfile Testing And Validation this skillEdwardvaneechoud/Flowfile373—~8.3kAutomated safety check: NotesMIT
Ckeditor5 TestingTriliumNext/Trilium38k—~3.3kAutomated safety check: PassAGPL-3.0
Running TestsNangoHQ/nango13k—~876Automated safety check: PassCustom licence
Slow TestsUKGovernmentBEIS/inspect_ai3k—~1.4kAutomated safety check: PassMIT
Testing Patternssoftspark/ai-toolkit179—~1.6kAutomated safety check: PassApache-2.0
Designing TestsCloudAI-X/claude-workflow-v21.4k1 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Ckeditor5 Testing

    TriliumNext/Trilium

    Testing CKEditor 5 plugins in the Trilium monorepo. An agent skill from TriliumNext/Trilium.

    38k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Running Tests

    NangoHQ/nango

    A skill your agent uses when running tests in the Nango monorepo - knows unit vs integration configs, vitest commands, Docker setup, and common test patterns

    13k GitHub stars~876 tokensUpdated today
    Testing & QAAuto-check passed
  • Slow Tests

    UKGovernmentBEIS/inspect_ai

    Run the gated test classes that plain pytest skips (slow Docker/sandbox tests, live model-provider API tests, flaky tests, trio variants).

    3k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Testing Patterns

    softspark/ai-toolkit

    Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.

    179 GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/claude-workflow-v2

    Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.

    1.4k GitHub starsUsed in 1 repo~1.5k tokens
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.1k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed

More from Edwardvaneechoud/Flowfile

All 19 skills in this repo
  • Flowfile AI Subsystem Guide

    Edwardvaneechoud/Flowfile

    Maps the /ai/ subsystem of flowfile_core, its three agent tiers, litellm seam, BYOK keys and rate limits, and sets rules for extending or debugging it safely.

    373 GitHub stars~7k tokensUpdated today
    Auto-check: notes
  • Flowfile Architecture Contract

    Edwardvaneechoud/Flowfile

    Maps Flowfile's core, worker, frontend, kernel, scheduler and shared services and the design contracts between them, for onboarding and cross-service debugging.

    373 GitHub stars~9.9k tokensUpdated today
    Auto-check passed
  • Flowfile Build and Environment Setup

    Edwardvaneechoud/Flowfile

    Recreates every Flowfile development and build environment from scratch, with exact version pins and an explanation of what each Makefile target really does.

    373 GitHub stars~7.3k tokensUpdated today
    Auto-check: notes
  • Flowfile Change Control

    Edwardvaneechoud/Flowfile

    Explains how changes to the Flowfile monorepo are gated, versioned and released, including version sync, stub and docs drift checks, Alembic migrations and pinned dependencies.

    373 GitHub stars~7.3k tokensUpdated today
    Auto-check passed
  • Flowfile Codegen Parity Campaign

    Edwardvaneechoud/Flowfile

    Runbook for closing gaps between a Flowfile visual flow's results and its exported Polars or FlowFrame Python code, measured by tests rather than by eye.

    373 GitHub stars~7.5k tokensUpdated today
    Auto-check passed
  • Flowfile Config and Flags Catalog

    Edwardvaneechoud/Flowfile

    Catalog of Flowfile's environment variables and runtime flags: what each does, where the code reads it, its default, and where the docs disagree with the code.

    373 GitHub stars~12k tokensUpdated today
    Auto-check: notes

Categories

Questions about Flowfile Testing And Validation

What does Flowfile Testing And Validation do?

Exact per-package pytest/vitest/playwright commands, the registered pytest markers and which need Docker, the testutils Docker fixture matrix, the shared-test-DB isolation model and its failure…. Flowfile Testing And Validation is an agent skill from Edwardvaneechoud/Flowfile. Exact per-package pytest/vitest/playwright commands, the registered pytest markers and which need Docker, the testutils Docker fixture matrix, the shared-test-DB isolation model and its failure modes, xfail/skip discipline, and coverage/CI test-matrix mechanics for the Flowfile monorepo.

When should I use Flowfile Testing And Validation?

Flowfile Testing And Validation fits situations like: diagnosing no such table; phantom test failures; deciding whether a change is validated; seeing XPASS in test output.

How do I install Flowfile Testing And Validation in Claude Code?

Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-testing-and-validation -a claude-code`. Or copy the skill folder (.claude/skills/flowfile-testing-and-validation in Edwardvaneechoud/Flowfile) into .claude/skills/flowfile-testing-and-validation in your project. Claude Code loads it when a task matches its description.

How do I install Flowfile Testing And Validation in Codex?

Run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-testing-and-validation -a codex`. Or copy the skill folder (.claude/skills/flowfile-testing-and-validation in Edwardvaneechoud/Flowfile) into .agents/skills/flowfile-testing-and-validation in your project. Codex loads it when a task matches its description.

Can I use Flowfile Testing And Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Edwardvaneechoud/Flowfile --skill flowfile-testing-and-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flowfile-testing-and-validation, .gemini/skills/flowfile-testing-and-validation, .github/skills/flowfile-testing-and-validation and .opencode/skills/flowfile-testing-and-validation in your project.

What does Flowfile Testing And Validation need to run?

Going by SKILL.md and its folder, Flowfile Testing And Validation needs the command-line tools its instructions call (poetry, make, npm, git, pytest and npx). Our summary lists: Python 3; Node.js; Docker.

Does Flowfile Testing And Validation access the network?

SKILL.md contains no URLs. Its commands use npm, git, npx, pip, docker and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Flowfile Testing And Validation safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Flowfile Testing And Validation use?

Flowfile Testing And Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Flowfile Testing And Validation use?

About 8.3k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Flowfile Testing And Validation?

Skills that share tags, products or a category with Flowfile Testing And Validation: Ckeditor5 Testing (TriliumNext/Trilium, 38k stars), Running Tests (NangoHQ/nango, 13k stars), Slow Tests (UKGovernmentBEIS/inspect_ai, 3k stars) and Testing Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flowfile Testing And Validation?

Edwardvaneechoud (a GitHub user) maintains it in Edwardvaneechoud/Flowfile, which has 373 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 8, 2026.

Source: Edwardvaneechoud/Flowfile on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.