Agent skill

Smt E2E Dataflow Debugging

by GoogleCloudPlatform in GoogleCloudPlatform/DataflowTemplates

Debugs logical errors and data discrepancies in Dataflow templates by launching jobs via Terraform and comparing source (e.g.

Apache-2.0Auto-check passedDevelopment

Install Smt E2E Dataflow Debugging

skills CLI
$ npx skills add GoogleCloudPlatform/DataflowTemplates --skill smt-e2e-dataflow-debugging -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GoogleCloudPlatform/DataflowTemplates smt-e2e-dataflow-debugging --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GoogleCloudPlatform/DataflowTemplates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/smt-e2e-dataflow-debugging .claude/skills/smt-e2e-dataflow-debugging && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
smt-e2e-dataflow-debugging
GitHub stars
1.3k
Token cost
~1.8k tokens
SKILL.md length
700 words
Files
2
Skills in repo
12
Repo updated
First seen
Licence
Apache-2.0

At a glance

Debugs logical errors and data discrepancies in Dataflow templates by launching jobs via Terraform and comparing source (e.g.

  • Works in 5 steps: Deployment & Job Monitoring → Log Collection & Initial Analysis → Data Inspection & Validation → …
  • Debugging template startup/runtime crashes
  • SKILL.md covers Trigger, Scope, Goal and Prerequisites, plus 4 more sections
  • Calls gcloud; needs CSQL_PASSWORD

What it does

Smt E2E Dataflow Debugging is an agent skill from GoogleCloudPlatform/DataflowTemplates. Debugs logical errors and data discrepancies in Dataflow templates by launching jobs via Terraform and comparing source (e.g. Cloud SQL) vs destination (e.g. Spanner) data. Use ONLY when the pipeline launches and runs to completion (terminal state) but exhibits data discrepancies or logical issues. Do NOT use for debugging template startup/runtime crashes or staging/building new templates.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `TEST.md`).

It sits in Development, covering Debugging, End-to-end testing and Infrastructure as code. It works with Google Cloud, SQL, Terraform and Google BigQuery. The repository describes itself as: Cloud Dataflow Google-provided templates for solving in-Cloud data tasks. The licence is Apache-2.0.

When your agent uses it

  • Debugging template startup/runtime crashes
  • Staging/building new templates

Example prompts

  • “Use the smt-e2e-dataflow-debugging skill to debug logical errors and data discrepancies in Dataflow templates by launching jobs via Terraform and…”
  • “/smt-e2e-dataflow-debugging”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Deployment & Job Monitoring
  2. Log Collection & Initial Analysis
  3. Data Inspection & Validation
  4. Discrepancy Verification & Resolution
  5. Re-testing & Iteration

What it can do on your machine

Read from SKILL.md and the folder at commit c95daba. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gcloud

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gcloud, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CSQL_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Smt E2E Dataflow Debugging loads about 1.8k tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 700 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GoogleCloudPlatform/DataflowTemplates at commit c95daba, republished under its Apache-2.0 licence (© GoogleCloudPlatform). 700 words, ~1,804 tokens.

Download SKILL.mdSave it as .claude/skills/smt-e2e-dataflow-debugging/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
smt-e2e-dataflow-debugging
description
Debugs logical errors and data discrepancies in Dataflow templates by launching jobs via Terraform and comparing source (e.g. Cloud SQL) vs destination (e.g. Spanner) data. Use ONLY when the pipeline launches and runs to completion (terminal state) but exhibits data discrepancies or logical issues. Do NOT use for debugging template startup/runtime crashes or staging/building new templates.

Skill: Dataflow Template Logical Error Debugging

Trigger

/smt-e2e-dataflow-debugging <FILE_NAME>.tfvars

Scope

This skill is STRICTLY restricted to testing the following templates:

  • gcs-spanner-dv
  • sourcedb-to-spanner
  • datastream-to-spanner
  • spanner-to-sourcedb

Goal

To debug logical errors and data discrepancies by comparing source data (e.g., Cloud SQL) with destination data (e.g., Spanner).

Prerequisites

  • The Dataflow job must be able to launch and run to a terminal state (Succeeded, Failed) using the provided .tfvars file. This skill is NOT for debugging startup/runtime crashes.
  • terraform CLI installed and configured.
  • gcloud CLI installed and authenticated (gcloud auth login).
  • mvn (Maven) and a compatible JDK installed.
  • git installed.
  • Access to the source database (e.g., Cloud SQL) instance, database, and credentials.
  • Access to the destination Spanner instance and database.
  • Appropriate IAM permissions for Dataflow, Cloud SQL, Spanner, GCS, and Cloud Logging.
  • Database client tools installed (e.g., psql, mysql client, or willingness to use gcloud sql connect).

Variables

  • <FILE_NAME>.tfvars: The Terraform variables file.
  • <YOUR_PROJECT_ID>: The target Google Cloud Project ID.
  • <YOUR_REGION>: The Google Cloud region for the Dataflow job.
  • <JOB_ID>: The Dataflow Job ID.
  • <JOB_NAME>: The name of the Dataflow job.
  • Cloud SQL Connection Info (Extracted from .tfvars):
    • <CSQL_INSTANCE>: Cloud SQL instance name.
    • <CSQL_DATABASE>: Cloud SQL database name.
    • <CSQL_USER>: Cloud SQL username.
    • <CSQL_PASSWORD>: Cloud SQL password (if applicable).
    • <CSQL_PROJECT>: Project of the Cloud SQL instance.
  • Spanner Connection Info (Extracted from .tfvars):
    • <SPANNER_INSTANCE>: Spanner instance name.
    • <SPANNER_DATABASE>: Spanner database name.
    • <SPANNER_PROJECT>: Project of the Spanner instance.
  • <MAVEN_MODULE_PATH>: The relative path to the template's maven module (e.g. v2/sourcedb-to-spanner).
  • <TEMPLATE_NAME>: The name of the template.
  • <YOUR_STAGING_BUCKET>: Cloud Storage bucket for staging.

Global Agent Rules

  • Fail-Fast Protocol: If any executed terminal command returns an error or non-zero exit code, STOP IMMEDIATELY. Output the error to the user and ask for intervention. Do not attempt autonomous retries.
  • Data Privacy: Ensure no real PII, credentials, or production connection details are exposed in logs or command outputs.

Workflow Phases

Phase 1: Deployment & Job Monitoring
  1. Deploy the Dataflow Job:
    • Navigate to the directory containing <FILE_NAME>.tfvars.
    • Initialize Terraform:
      bash
      terraform init
    • Launch the job:
      bash
      terraform apply --var-file=<FILE_NAME>.tfvars -auto-approve
  2. Monitor Status: Wait for the job to reach a terminal state (Succeeded or Failed). Monitor status via Google Cloud Console or gcloud.
Phase 2: Log Collection & Initial Analysis
  1. Retrieve Job IDs: Infer <JOB_ID> and <JOB_NAME> from the Terraform output or by listing active jobs:
    bash
    gcloud dataflow jobs list --project=<YOUR_PROJECT_ID> --region=<YOUR_REGION>
  2. Inspect Job Logs: Retrieve job message logs to identify warnings or non-obvious issues:
    bash
    gcloud logging read 'resource.type="dataflow_step" AND resource.labels.job_id="<JOB_ID>" AND logName=~"projects/.*/logs/dataflow.googleapis.com%2Fjob-message"' \
      --project=<YOUR_PROJECT_ID> \
      --limit=200 \
      --format="table(timestamp, textPayload, severity)" --order=asc
  3. Inspect Worker Logs: Check worker execution logs for specific exceptions or stack traces:
    bash
    gcloud logging read 'resource.type="dataflow_step" AND logName=~"projects/.*/logs/dataflow.googleapis.com%2Fworker" AND resource.labels.job_id="<JOB_ID>"' \
      --project=<YOUR_PROJECT_ID> \
      --limit=500 \
      --format="table(timestamp, jsonPayload.message, severity)" --order=asc
Show full SKILL.md (252 more words)Show less
Phase 3: Data Inspection & Validation
  1. Inspect Source Data (Cloud SQL): Connect to the Cloud SQL instance and query source tables:
    bash
    # For PostgreSQL
    gcloud sql connect <CSQL_INSTANCE> --user=<CSQL_USER> --project=<CSQL_PROJECT>
    # For MySQL
    gcloud sql connect <CSQL_INSTANCE> --user=<CSQL_USER> --project=<CSQL_PROJECT>
    Execute validation queries:
    sql
    SELECT COUNT(*) FROM your_source_table;
    SELECT * FROM your_source_table LIMIT 10;
  2. Inspect Destination Data (Spanner): Execute SQL queries on the destination Spanner database:
    bash
    gcloud spanner databases execute-sql <SPANNER_DATABASE> \
      --instance=<SPANNER_INSTANCE> \
      --project=<SPANNER_PROJECT> \
      --sql="SELECT COUNT(*) FROM your_destination_table"
    
    gcloud spanner databases execute-sql <SPANNER_DATABASE> \
      --instance=<SPANNER_INSTANCE> \
      --project=<SPANNER_PROJECT> \
      --sql="SELECT * FROM your_destination_table LIMIT 10"
Phase 4: Discrepancy Verification & Resolution
  1. Analyze Differences:
    • Row Counts: Do source and destination row counts match?
    • Schema Mapping: Are column types and names mapped correctly?
    • Value Assertions: Check NULL values, string encoding, timestamps, and precision.
  2. Code Correction: Locate the transformation logic in the .java files under the GoogleCloudPlatform/DataflowTemplates repository and fix the bug.
  3. Re-stage Template: Rebuild and upload the updated Flex Template:
    bash
    mvn clean package -PtemplatesStage -DskipTests \
      -DprojectId="<YOUR_PROJECT_ID>" \
      -DbucketName="<YOUR_STAGING_BUCKET>" \
      -DstagePrefix="templates" \
      -DtemplateName="<TEMPLATE_NAME>" \
      -pl <MAVEN_MODULE_PATH> -am
Phase 5: Re-testing & Iteration
  1. Clean Destination Table: Delete destination records to prepare for a clean run:
    bash
    gcloud spanner databases execute-sql <SPANNER_DATABASE> \
      --instance=<SPANNER_INSTANCE> \
      --project=<SPANNER_PROJECT> \
      --sql="DELETE FROM your_destination_table WHERE true"
  2. Re-run Job: Deploy the job again with the newly built template (Phase 1) and verify the fix.

Important Considerations

  • Idempotency: Ensure clean test tables before re-running to avoid duplicate count errors.
  • Data Volume: Limit queries to subsets when working with large volumes.
  • Terraform State: Avoid modifying resources manually outside of Terraform to prevent state drift.

© GoogleCloudPlatform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/smt-e2e-dataflow-debugging of GoogleCloudPlatform/DataflowTemplates.

  • SKILL.md
  • TEST.md

Open the folder on GitHubat commit c95daba

Compare with similar skills

Smt E2E Dataflow Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Smt E2E Dataflow Debugging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Smt E2E Dataflow Debugging this skillGoogleCloudPlatform/DataflowTemplates1.3k—~1.8kAutomated safety check: PassApache-2.0
Google Cloud Storage Basicsgoogle/skills21k—~2.8kAutomated safety check: PassApache-2.0
Dd GCP Integrationdatadog-labs/agent-skills177—~8kAutomated safety check: NotesMIT
Test Monitor WorkflowGoogleCloudPlatform/magic-modules973—~1kAutomated safety check: PassCustom licence
Dd Azure Integrationdatadog-labs/agent-skills177—~7.1kAutomated safety check: NotesMIT
GCP To AWSaws/agent-toolkit-for-aws2.8k—~14kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Stores, retrieves, and manages data as objects in Cloud Storage on Google Cloud (also known colloquially as GCS) buckets.

    21k GitHub stars~2.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Dd GCP Integration

    datadog-labs/agent-skills

    Set up the Datadog Google Cloud integration with Terraform - creates a service account in the host project, lets Datadog's delegate principal impersonate it via roles/iam.serviceAccountTokenCreator…

    177 GitHub stars~8k tokensUpdated 5 days ago
    DevOps & CloudAuto-check: notes
  • Test Monitor Workflow

    GoogleCloudPlatform/magic-modules

    Workflow for fetching, triaging, analyzing, and reporting on nightly acceptance test results across Beta and GA Google Cloud Terraform providers.

    973 GitHub stars~1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Dd Azure Integration

    datadog-labs/agent-skills

    Set up the Datadog Azure integration with Terraform - creates an Entra ID app registration and service principal, assigns Monitoring Reader across the chosen subscriptions and management groups…

    177 GitHub stars~7.1k tokensUpdated 5 days ago
    DevOps & CloudAuto-check: notes
  • GCP To AWS

    aws/agent-toolkit-for-aws

    Official

    Migrate workloads from Google Cloud Platform to AWS — plus AI and agentic workloads from any provider.

    2.8k GitHub stars~14k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Cxas Configurable Dashboards

    GoogleCloudPlatform/cxas-scrapi

    Author, validate, and manage Contact Center AI (CCAI) Insights Configurable Dashboards.

    106 GitHub stars~1.8k tokensUpdated yesterday
    DatabasesAuto-check passed

More from GoogleCloudPlatform/DataflowTemplates

All 12 skills in this repo
  • Smt Functional Testing

    GoogleCloudPlatform/DataflowTemplates

    Functionally tests local Dataflow pipeline changes against the main branch using ephemeral GCP resources and gated approvals.

    1.3k GitHub stars~2.8k tokensUpdated today
    Auto-check: notes
  • Add Integ Tests Datastream To Spanner

    GoogleCloudPlatform/DataflowTemplates

    Specific runner skill that delegates to the Template-Agnostic Meta-Test Orchestrator for the datastream-to-spanner (CDC) template.

    1.3k GitHub stars~572 tokensUpdated today
    Auto-check passed
  • Add Integ Tests Gcs Spanner Dv

    GoogleCloudPlatform/DataflowTemplates

    Specific runner skill that creates integration tests for the gcs-spanner-dv (Data Validation) template.

    1.3k GitHub stars~517 tokensUpdated today
    Auto-check passed
  • Add Integ Tests Sourcedb To Spanner

    GoogleCloudPlatform/DataflowTemplates

    Specific runner skill that delegates to the Template-Agnostic Meta-Test Orchestrator for the sourcedb-to-spanner (Bulk) template.

    1.3k GitHub stars~475 tokensUpdated today
    Auto-check passed
  • Add Integ Tests Spanner To Sourcedb

    GoogleCloudPlatform/DataflowTemplates

    Specific runner skill that delegates to the Template-Agnostic Meta-Test Orchestrator for the spanner-to-sourcedb (Reverse Migration) template.

    1.3k GitHub stars~479 tokensUpdated today
    Auto-check passed
  • Add Source Datastream To Spanner

    GoogleCloudPlatform/DataflowTemplates

    Guide for implementing a database source connector in the v2/datastream-to-spanner forward migration Dataflow template.

    1.3k GitHub stars~1.7k tokensUpdated today
    Auto-check: notes

Questions about Smt E2E Dataflow Debugging

What does Smt E2E Dataflow Debugging do?

Debugs logical errors and data discrepancies in Dataflow templates by launching jobs via Terraform and comparing source (e.g. Smt E2E Dataflow Debugging is an agent skill from GoogleCloudPlatform/DataflowTemplates.g.

When should I use Smt E2E Dataflow Debugging?

Smt E2E Dataflow Debugging fits situations like: debugging template startup/runtime crashes; staging/building new templates.

How do I install Smt E2E Dataflow Debugging in Claude Code?

Run `npx skills add GoogleCloudPlatform/DataflowTemplates --skill smt-e2e-dataflow-debugging -a claude-code`. Or copy the skill folder (.agents/skills/smt-e2e-dataflow-debugging in GoogleCloudPlatform/DataflowTemplates) into .claude/skills/smt-e2e-dataflow-debugging in your project. Claude Code loads it when a task matches its description.

How do I install Smt E2E Dataflow Debugging in Codex?

Run `npx skills add GoogleCloudPlatform/DataflowTemplates --skill smt-e2e-dataflow-debugging -a codex`. Or copy the skill folder (.agents/skills/smt-e2e-dataflow-debugging in GoogleCloudPlatform/DataflowTemplates) into .agents/skills/smt-e2e-dataflow-debugging in your project. Codex loads it when a task matches its description.

Can I use Smt E2E Dataflow Debugging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GoogleCloudPlatform/DataflowTemplates --skill smt-e2e-dataflow-debugging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/smt-e2e-dataflow-debugging, .gemini/skills/smt-e2e-dataflow-debugging, .github/skills/smt-e2e-dataflow-debugging and .opencode/skills/smt-e2e-dataflow-debugging in your project.

What does Smt E2E Dataflow Debugging need to run?

Going by SKILL.md and its folder, Smt E2E Dataflow Debugging needs the command-line tools its instructions call (gcloud) and credentials named CSQL_PASSWORD.

Does Smt E2E Dataflow Debugging access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Smt E2E Dataflow Debugging safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Smt E2E Dataflow Debugging use?

Smt E2E Dataflow Debugging is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Smt E2E Dataflow Debugging use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Smt E2E Dataflow Debugging?

Skills that share tags, products or a category with Smt E2E Dataflow Debugging: Google Cloud Storage Basics (google/skills, 21k stars), Dd GCP Integration (datadog-labs/agent-skills, 177 stars), Test Monitor Workflow (GoogleCloudPlatform/magic-modules, 973 stars) and Dd Azure Integration (datadog-labs/agent-skills, 177 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Smt E2E Dataflow Debugging?

GoogleCloudPlatform (a GitHub organization) maintains it in GoogleCloudPlatform/DataflowTemplates, which has 1,315 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 7, 2026.

Source: GoogleCloudPlatform/DataflowTemplates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.