Agent skill

Snowflake Snowpark Dbt

by Mindrally in Mindrally/skills

Best practices for Snowpark Python (DataFrames, UDFs, UDTFs, stored procedures) and dbt with the dbt-snowflake adapter.

Apache-2.0Auto-check passedData & Analytics

Install Snowflake Snowpark Dbt

skills CLI
$ npx skills add Mindrally/skills --skill snowflake-snowpark-dbt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mindrally/skills snowflake-snowpark-dbt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mindrally/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/snowflake-snowpark-dbt .claude/skills/snowflake-snowpark-dbt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
snowflake-snowpark-dbt
GitHub stars
271
Token cost
~2.5k tokens
SKILL.md length
753 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
Apache-2.0

At a glance

Best practices for Snowpark Python (DataFrames, UDFs, UDTFs, stored procedures) and dbt with the dbt-snowflake adapter.

  • Works in 8 steps: Snowpark: open a session — Build a… → Snowpark: express transforms with the… → Snowpark: push compute server-side — Use… → …
  • Writing server-side Snowpark pipelines
  • SKILL.md covers Workflow for a Snowpark or dbt…, Snowpark Python, dbt with the Snowflake Adapter and Best Practices, plus 1 more section
  • Calls dbt and pip; needs SNOWFLAKE_PASSWORD

What it does

Snowflake Snowpark Dbt is an agent skill from Mindrally/skills. Best practices for Snowpark Python (DataFrames, UDFs, UDTFs, stored procedures) and dbt with the dbt-snowflake adapter. Use when writing server-side Snowpark pipelines, registering UDFs or stored procedures, choosing dbt materializations, configuring incremental models, or setting up sources and tests for a Snowflake-backed dbt project.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL, Data warehousing and DataFrames. It works with Snowflake, dbt and Python. The repository describes itself as: 265+ Claude Code skills for every major framework and language. Install with: npx skills add Mindrally/skills. The licence is Apache-2.0.

When your agent uses it

  • Writing server-side Snowpark pipelines
  • Registering UDFs
  • Stored procedures
  • Choosing dbt materializations

Example prompts

  • “/snowflake-snowpark-dbt”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Snowpark: open a session — Build a Session from environment-scoped credentials, specifying role, warehouse, database, and schema explicitly.
  2. Snowpark: express transforms with the DataFrame API — Prefer .filter(), .select(), .group_by().agg(), and .join() over raw SQL strings for…
  3. Snowpark: push compute server-side — Use scalar UDFs for row-wise logic, vectorized (pandas) UDFs for ML inference, UDTFs when one input…
  4. dbt: model in layers — Staging models (stg_*) rename and type-cast; mart models express business logic on top of staging.
  5. dbt: choose a materialization — view for cheap logic, table only when reads are frequent, incremental for large fact tables, dynamic_table…
  6. dbt: define sources and tests — Declare sources in _sources.yml with freshness thresholds; add unique/not_null tests on key columns.
  7. dbt: run selectively — Use dbt run --select model+ (model and downstream) or +model (model and upstream) instead of full-project runs…
  8. dbt: build and validate — Run dbt build (run + test in dependency order) before merging, and dbt docs generate to keep documentation…

What it can do on your machine

Read from SKILL.md and the folder at commit 7682ca7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • dbt
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SNOWFLAKE_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Snowflake Snowpark Dbt loads about 2.5k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 753 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mindrally/skills at commit 7682ca7, republished under its Apache-2.0 licence (© Mindrally). 753 words, ~2,472 tokens.

Download SKILL.mdSave it as .claude/skills/snowflake-snowpark-dbt/SKILL.md (or your agent's skills folder).
name
snowflake-snowpark-dbt
description
Best practices for Snowpark Python (DataFrames, UDFs, UDTFs, stored procedures) and dbt with the dbt-snowflake adapter. Use when writing server-side Snowpark pipelines, registering UDFs or stored procedures, choosing dbt materializations, configuring incremental models, or setting up sources and tests for a Snowflake-backed dbt project.
metadata.maintainer
Mindrally
metadata.source
https://github.com/Mindrally/skills

Snowflake Snowpark Python & dbt

This skill covers building production data transformation pipelines with Snowpark Python (Snowflake's server-side Python API) and with dbt using the dbt-snowflake adapter.

Workflow for a Snowpark or dbt Transformation

  1. Snowpark: open a session — Build a Session from environment-scoped credentials, specifying role, warehouse, database, and schema explicitly.
  2. Snowpark: express transforms with the DataFrame API — Prefer .filter(), .select(), .group_by().agg(), and .join() over raw SQL strings for reusable pipeline code; DataFrames are lazily evaluated and only execute on .collect()/.show()/a write action.
  3. Snowpark: push compute server-side — Use scalar UDFs for row-wise logic, vectorized (pandas) UDFs for ML inference, UDTFs when one input row produces multiple output rows, and stored procedures for multi-step server-side orchestration.
  4. dbt: model in layers — Staging models (stg_*) rename and type-cast; mart models express business logic on top of staging.
  5. dbt: choose a materialization — view for cheap logic, table only when reads are frequent, incremental for large fact tables, dynamic_table for near-real-time freshness needs.
  6. dbt: define sources and tests — Declare sources in _sources.yml with freshness thresholds; add unique/not_null tests on key columns.
  7. dbt: run selectively — Use dbt run --select model+ (model and downstream) or +model (model and upstream) instead of full-project runs during iteration.
  8. dbt: build and validate — Run dbt build (run + test in dependency order) before merging, and dbt docs generate to keep documentation current.

Snowpark Python

Snowpark runs Python server-side inside a Snowflake warehouse — data never leaves Snowflake. Core abstractions: Session, DataFrame, UDF, UDTF, UDAF, and Stored Procedure.

Session
python
import os
from snowflake.snowpark import Session

session = Session.builder.configs({
    "account": os.environ["SNOWFLAKE_ACCOUNT"],
    "user": os.environ["SNOWFLAKE_USER"],
    "password": os.environ["SNOWFLAKE_PASSWORD"],
    "role": "my_role", "warehouse": "my_wh", "database": "my_db", "schema": "my_schema",
}).create()

Never hardcode credentials — always read them from environment variables or a secrets manager.

DataFrame API

DataFrames are lazily evaluated: Snowpark builds a query plan and only executes it on collect()/show() or a write action.

python
df = session.table("customers")
df_filtered = df.filter(df["region"] == "US").select("name", "email", "revenue")
df_agg = df.group_by("region").agg(sum("revenue").alias("total_revenue"))
df_agg.show()

Key operations: .filter(), .select(), .group_by().agg(), .join(), .sort(), .with_column(), .drop(), .distinct(), .limit(), .union_all(), .flatten(), .write.save_as_table().

Scalar UDFs
python
from snowflake.snowpark.functions import udf

@udf(name="normalize_email", replace=True)
def normalize_email(email: str) -> str:
    return email.strip().lower() if email else None
Vectorized UDFs

Vectorized (pandas) UDFs are 10-100x faster than scalar UDFs for ML inference because they batch rows instead of invoking Python per row.

python
import pandas as pd
from snowflake.snowpark.functions import udf

@udf(name="predict_score", packages=["scikit-learn", "pandas"], replace=True)
def predict_score(features: pd.Series) -> pd.Series:
    import pickle, sys
    model = pickle.load(open(sys.path[0] + "/model.pkl", "rb"))
    return pd.Series(model.predict(features.values.reshape(-1, 1)))
UDTFs (return multiple rows per input)
python
from snowflake.snowpark.types import StructType, StructField, StringType

class Tokenizer:
    def process(self, text: str):
        for token in text.split():
            yield (token,)

tokenize = session.udtf.register(
    Tokenizer,
    output_schema=StructType([StructField("token", StringType())]),
    input_types=[StringType()],
    name="tokenize",
    replace=True,
)
Stored procedures
python
from snowflake.snowpark import Session
from snowflake.snowpark.functions import sproc

@sproc(name="daily_etl", replace=True, packages=["snowflake-snowpark-python"])
def daily_etl(session: Session) -> str:
    raw = session.table("raw_events")
    cleaned = raw.filter(raw["event_type"].is_not_null())
    cleaned.write.mode("overwrite").save_as_table("cleaned_events")
    return f"Processed {cleaned.count()} rows"
Packages and file access
  • Add third-party packages with session.add_packages("pandas", "scikit-learn==1.3.0", "xgboost") — pin versions for production UDFs and stored procedures.
  • Attach static files (e.g., a pickled model) with session.add_import("@my_stage/model.pkl").
  • For pandas-on-Snowflake with no data movement to the client, use modin.pandas with the Snowpark plugin: import modin.pandas as pd; import snowflake.snowpark.modin.plugin; df = pd.read_snowflake("my_table").

dbt with the Snowflake Adapter

Install with pip install dbt-snowflake.

profiles.yml
yaml
my_project:
  target: dev
  outputs:
    dev:
      type: snowflake
      account: myaccount
      user: myuser
      password: "{{ env_var('SNOWFLAKE_PASSWORD') }}"
      role: transformer
      database: analytics
      warehouse: transforming
      schema: public
      threads: 4

Never commit real credentials into profiles.yml — always source secrets from env_var().

Materializations

Available materializations: view, table, incremental, ephemeral, dynamic_table. Default to view for cheap logic; reserve table for models read frequently enough to justify storage cost; use incremental for large fact tables and dynamic_table when near-real-time freshness matters.

Dynamic Tables in dbt
sql
{{ config(materialized='dynamic_table', snowflake_warehouse='transforming', target_lag='1 hour') }}
SELECT customer_id, SUM(amount) AS lifetime_value FROM {{ ref('stg_orders') }} GROUP BY 1
Show full SKILL.md (305 more words)Show less
Incremental models
sql
{{
  config(
    materialized='incremental',
    unique_key='event_id',
    incremental_strategy='merge',
    on_schema_change='sync_all_columns'
  )
}}
SELECT * FROM {{ ref('stg_events') }}
{% if is_incremental() %}
  WHERE event_timestamp > (SELECT MAX(event_timestamp) FROM {{ this }})
{% endif %}

Always guard {{ this }} with {% if is_incremental() %} — referencing it unconditionally breaks the first (full) run, when the target table doesn't exist yet.

Snowflake-specific configs
  • cluster_by=['col1', 'col2'] — clustering, large tables only (generally >1TB).
  • transient=true — no Fail-safe, lower storage cost; use for staging models.
  • query_tag='finance_daily' — workload attribution for cost tracking.
  • copy_grants=true — preserve access grants across a CREATE OR REPLACE.
  • snowflake_warehouse='lg_wh' — per-model warehouse override for heavy transforms.
  • secure=true — secure views, for models exposing sensitive columns.
Sources with freshness checks
yaml
sources:
  - name: raw
    database: raw_db
    schema: jaffle_shop
    tables:
      - name: customers
        loaded_at_field: _loaded_at
        freshness:
          warn_after: {count: 12, period: hour}
          error_after: {count: 24, period: hour}
Testing
yaml
models:
  - name: stg_customers
    columns:
      - name: customer_id
        tests: [unique, not_null]
Key commands
  • dbt run, dbt test, dbt build (run + test in dependency order), dbt compile
  • dbt run --select my_model+ — model and everything downstream
  • dbt run --select +my_model — model and everything upstream
  • dbt source freshness — check source staleness against configured thresholds
  • dbt docs generate && dbt docs serve — build and preview documentation
Custom schema naming
jinja
{% macro generate_schema_name(custom_schema_name, node) %}
  {% if custom_schema_name %}{{ custom_schema_name | trim }}{% else %}{{ target.schema }}{% endif %}
{% endmacro %}

Best Practices

  • Prefer vectorized (pandas) UDFs over scalar UDFs for ML inference.
  • Pin package versions in production UDFs and stored procedures.
  • Use the Snowpark DataFrame API over raw SQL strings in reusable Python pipelines.
  • Use staging models (stg_*) to rename and type-cast; keep business logic in mart models.
  • Use incremental materialization for fact tables and dynamic_table for near-real-time needs.
  • Set on_schema_change='sync_all_columns' on incremental models to handle upstream schema drift safely.
  • Use copy_grants=true to avoid permission churn on rebuild, and tag models for selective execution.
  • Use separate warehouses for dbt runs versus interactive analyst queries.

Anti-Patterns

  • Do not .collect() large Snowpark DataFrames to the client — keep processing server-side.
  • Do not use Python loops over rows in Snowpark — use DataFrame operations or vectorized UDFs.
  • Do not reference {{ this }} in an incremental model without an {% if is_incremental() %} guard.
  • Do not set cluster_by on small tables (under roughly 1TB) — the overhead outweighs the benefit.
  • Do not default every model to materialized='table' — views are free until queried.
  • Do not hardcode database/schema names in dbt models — use {{ ref() }} and {{ source() }}.

© Mindrally, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in snowflake-snowpark-dbt of Mindrally/skills.

Open the folder on GitHubat commit 7682ca7

Compare with similar skills

Snowflake Snowpark Dbt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Snowflake Snowpark Dbt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Snowflake Snowpark Dbt this skillMindrally/skills271—~2.5kAutomated safety check: PassApache-2.0
Snowflake Developmentsickn33/agentic-awesome-skills47k2 repos~2.1kAutomated safety check: PassMIT
Snowflake Developmentalirezarezvani/claude-skills28k—~3.2kAutomated safety check: PassMIT
Data Warehouse Experimentationrampstackco/claude-skills945—~7.3kAutomated safety check: PassMIT
Airflow State Storeastronomer/agents451—~6.1kAutomated safety check: PassApache-2.0
Migrating Dbt Project Across PlatformsKilo-Org/kilo-marketplace190—~3.9kAutomated safety check: PassApache-2.0

Similar skills

  • Snowflake Development

    sickn33/agentic-awesome-skills

    Comprehensive Snowflake development assistant covering SQL best practices, data pipeline design (Dynamic Tables, Streams, Tasks, Snowpipe), Cortex AI functions, Cortex Agents, Snowpark Python, dbt…

    47k GitHub starsUsed in 2 repos~2.1k tokens
    DatabasesAuto-check passed
  • Snowflake Development

    alirezarezvani/claude-skills

    A skill your agent uses when writing Snowflake SQL, building data pipelines with Dynamic Tables or Streams/Tasks, using Cortex AI functions, creating Cortex Agents, writing Snowpark Python…

    28k GitHub stars~3.2k tokensUpdated 1 mo ago
    DatabasesAuto-check passed
  • Data Warehouse Experimentation

    rampstackco/claude-skills

    Running experiments out of the data warehouse instead of via dedicated experiment platforms.

    945 GitHub stars~7.3k tokensUpdated 4 days ago
    DatabasesAuto-check passed
  • Airflow State Store

    astronomer/agents

    Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (taskstatestore, assetstatestore) and the crash-safe ResumableJobMixin.

    451 GitHub stars~6.1k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • A skill your agent uses when migrating a dbt project from one data platform or data warehouse to another (e.g., Snowflake to Databricks, Databricks to Snowflake) using dbt Fusion's real-time…

    190 GitHub stars~3.9k tokensUpdated 12 days ago
    Data & AnalyticsAuto-check passed
  • Transforming Data

    ancoleman/ai-design-components

    Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).

    525 GitHub stars~3k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed

More from Mindrally/skills

All 34 skills in this repo
  • Analytics Data Analysis

    Mindrally/skills

    Best practices for analytics, data analysis, and visualization using Python, pandas, matplotlib, seaborn, and Jupyter notebooks.

    271 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed
  • Best practices for AutoML and hyperparameter search with Optuna, Ray Tune, and PyCaret, covering search-space design, validation splits, and leakage prevention.

    271 GitHub stars~2.5k tokensUpdated 2 days ago
    Auto-check passed
  • Blender Python Addon

    Mindrally/skills

    Best practices for writing Blender Python add-ons using the bpy API, covering operators, panels, properties, registration, and API-safe scripting.

    271 GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed
  • Expert guidelines for Chrome extension development with Manifest V3, covering security, performance, and best practices.

    271 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check passed
  • Clean Code

    Mindrally/skills

    Clean, maintainable, human-readable code principles combined with anti-over-engineering discipline: naming, single responsibility, DRY, and scoping changes to exactly what was requested.

    271 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Design Systems

    Mindrally/skills

    Comprehensive design system guidelines for building consistent, accessible, and scalable component libraries.

    271 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed

Questions about Snowflake Snowpark Dbt

What does Snowflake Snowpark Dbt do?

Best practices for Snowpark Python (DataFrames, UDFs, UDTFs, stored procedures) and dbt with the dbt-snowflake adapter. Snowflake Snowpark Dbt is an agent skill from Mindrally/skills. Best practices for Snowpark Python (DataFrames, UDFs, UDTFs, stored procedures) and dbt with the dbt-snowflake adapter.

When should I use Snowflake Snowpark Dbt?

Snowflake Snowpark Dbt fits situations like: writing server-side Snowpark pipelines; registering UDFs; stored procedures; choosing dbt materializations.

How do I install Snowflake Snowpark Dbt in Claude Code?

Run `npx skills add Mindrally/skills --skill snowflake-snowpark-dbt -a claude-code`. Or copy the skill folder (snowflake-snowpark-dbt in Mindrally/skills) into .claude/skills/snowflake-snowpark-dbt in your project. Claude Code loads it when a task matches its description.

How do I install Snowflake Snowpark Dbt in Codex?

Run `npx skills add Mindrally/skills --skill snowflake-snowpark-dbt -a codex`. Or copy the skill folder (snowflake-snowpark-dbt in Mindrally/skills) into .agents/skills/snowflake-snowpark-dbt in your project. Codex loads it when a task matches its description.

Can I use Snowflake Snowpark Dbt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mindrally/skills --skill snowflake-snowpark-dbt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/snowflake-snowpark-dbt, .gemini/skills/snowflake-snowpark-dbt, .github/skills/snowflake-snowpark-dbt and .opencode/skills/snowflake-snowpark-dbt in your project.

What does Snowflake Snowpark Dbt need to run?

Going by SKILL.md and its folder, Snowflake Snowpark Dbt needs the command-line tools its instructions call (dbt and pip) and credentials named SNOWFLAKE_PASSWORD. Our summary lists: Python 3.

Does Snowflake Snowpark Dbt access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Snowflake Snowpark Dbt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Snowflake Snowpark Dbt use?

Snowflake Snowpark Dbt is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Snowflake Snowpark Dbt use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Snowflake Snowpark Dbt?

Skills that share tags, products or a category with Snowflake Snowpark Dbt: Snowflake Development (sickn33/agentic-awesome-skills, 47k stars), Snowflake Development (alirezarezvani/claude-skills, 28k stars), Data Warehouse Experimentation (rampstackco/claude-skills, 945 stars) and Airflow State Store (astronomer/agents, 451 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Snowflake Snowpark Dbt?

Mindrally (a GitHub organization) maintains it in Mindrally/skills, which has 271 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 8, 2026.

Source: Mindrally/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.