Official agent skill

Bigquery Pipeline Audit

by github in github/awesome-copilot

Audits Python + BigQuery pipelines for cost safety, idempotency, and production readiness.

OfficialMITAuto-check passedDatabases

Install Bigquery Pipeline Audit

skills CLI
$ npx skills add github/awesome-copilot --skill bigquery-pipeline-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot bigquery-pipeline-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bigquery-pipeline-audit .claude/skills/bigquery-pipeline-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bigquery-pipeline-audit
GitHub stars
40k
Used in
1 other repo
Token cost
~1.3k tokens
SKILL.md length
720 words
Files
1
Skills in repo
417
Repo updated
First seen
Licence
MIT

At a glance

Audits Python + BigQuery pipelines for cost safety, idempotency, and production readiness.

  • Works in 3 steps: A single set-based query with… → A staging table loaded with all dates… → Explicit chunks with a hard MAX_CHUNKS cap
  • Tasks that involve Data warehousing
  • SKILL.md covers A) COST EXPOSURE: What will…, B) DRY RUN AND EXECUTION MODES, C) BACKFILL AND LOOP DESIGN and D) QUERY SAFETY AND SCAN SIZE, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Bigquery Pipeline Audit is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Audits Python + BigQuery pipelines for cost safety, idempotency, and production readiness. Returns a structured report with exact patch locations.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Data warehousing. It works with Google BigQuery and Python. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is MIT.

When your agent uses it

  • Tasks that involve Data warehousing

Example prompts

  • “Use the bigquery-pipeline-audit skill to audit Python + BigQuery pipelines for cost safety, idempotency, and production readiness”
  • “/bigquery-pipeline-audit”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A single set-based query with GENERATE_DATE_ARRAY
  2. A staging table loaded with all dates then one join query
  3. Explicit chunks with a hard MAX_CHUNKS cap

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bigquery Pipeline Audit loads about 1.3k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 720 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 720 words, ~1,251 tokens.

Download SKILL.mdSave it as .claude/skills/bigquery-pipeline-audit/SKILL.md (or your agent's skills folder).
name
bigquery-pipeline-audit
description
Audits Python + BigQuery pipelines for cost safety, idempotency, and production readiness. Returns a structured report with exact patch locations.

BigQuery Pipeline Audit: Cost, Safety and Production Readiness

You are a senior data engineer reviewing a Python + BigQuery pipeline script. Your goals: catch runaway costs before they happen, ensure reruns do not corrupt data, and make sure failures are visible.

Analyze the codebase and respond in the structure below (A to F + Final). Reference exact function names and line locations. Suggest minimal fixes, not rewrites.


A) COST EXPOSURE: What will actually get billed?

Locate every BigQuery job trigger (client.query, load_table_from_*, extract_table, copy_table, DDL/DML via query) and every external call (APIs, LLM calls, storage writes).

For each, answer:

  • Is this inside a loop, retry block, or async gather?
  • What is the realistic worst-case call count?
  • For each client.query, is QueryJobConfig.maximum_bytes_billed set? For load, extract, and copy jobs, is the scope bounded and counted against MAX_JOBS?
  • Is the same SQL and params being executed more than once in a single run? Flag repeated identical queries and suggest query hashing plus temp table caching.

Flag immediately if:

  • Any BQ query runs once per date or once per entity in a loop
  • Worst-case BQ job count exceeds 20
  • maximum_bytes_billed is missing on any client.query call

B) DRY RUN AND EXECUTION MODES

Verify a --mode flag exists with at least dry_run and execute options.

  • dry_run must print the plan and estimated scope with zero billed BQ execution (BigQuery dry-run estimation via job config is allowed) and zero external API or LLM calls
  • execute requires explicit confirmation for prod (--env=prod --confirm)
  • Prod must not be the default environment

If missing, propose a minimal argparse patch with safe defaults.


C) BACKFILL AND LOOP DESIGN

Hard fail if: the script runs one BQ query per date or per entity in a loop.

Check that date-range backfills use one of:

  1. A single set-based query with GENERATE_DATE_ARRAY
  2. A staging table loaded with all dates then one join query
  3. Explicit chunks with a hard MAX_CHUNKS cap

Also check:

  • Is the date range bounded by default (suggest 14 days max without --override)?
  • If the script crashes mid-run, is it safe to re-run without double-writing?
  • For backdated simulations, verify data is read from time-consistent snapshots (FOR SYSTEM_TIME AS OF, partitioned as-of tables, or dated snapshot tables). Flag any read from a "latest" or unversioned table when running in backdated mode.

Suggest a concrete rewrite if the current approach is row-by-row.


D) QUERY SAFETY AND SCAN SIZE

For each query, check:

  • Partition filter is on the raw column, not DATE(ts), CAST(...), or any function that prevents pruning
  • No SELECT *: only columns actually used downstream
  • Joins will not explode: verify join keys are unique or appropriately scoped and flag any potential many-to-many
  • Expensive operations (REGEXP, JSON_EXTRACT, UDFs) only run after partition filtering, not on full table scans

Provide a specific SQL fix for any query that fails these checks.


Show full SKILL.md (253 more words)Show less

E) SAFE WRITES AND IDEMPOTENCY

Identify every write operation. Flag plain INSERT/append with no dedup logic.

Each write should use one of:

  1. MERGE on a deterministic key (e.g., entity_id + date + model_version)
  2. Write to a staging table scoped to the run, then swap or merge into final
  3. Append-only with a dedupe view: QUALIFY ROW_NUMBER() OVER (PARTITION BY <key>) = 1

Also check:

  • Will a re-run create duplicate rows?
  • Is the write disposition (WRITE_TRUNCATE vs WRITE_APPEND) intentional and documented?
  • Is run_id being used as part of the merge or dedupe key? If so, flag it. run_id should be stored as a metadata column, not as part of the uniqueness key, unless you explicitly want multi-run history.

State the recommended approach and the exact dedup key for this codebase.


F) OBSERVABILITY: Can you debug a failure?

Verify:

  • Failures raise exceptions and abort with no silent except: pass or warn-only
  • Each BQ job logs: job ID, bytes processed or billed when available, slot milliseconds, and duration
  • A run summary is logged or written at the end containing: run_id, env, mode, date_range, tables written, total BQ jobs, total bytes
  • run_id is present and consistent across all log lines

If run_id is missing, propose a one-line fix: run_id = run_id or datetime.utcnow().strftime('%Y%m%dT%H%M%S')


Final

1. PASS / FAIL with specific reasons per section (A to F). 2. Patch list ordered by risk, referencing exact functions to change. 3. If FAIL: Top 3 cost risks with a rough worst-case estimate (e.g., "loop over 90 dates x 3 retries = 270 BQ jobs").

© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/bigquery-pipeline-audit of github/awesome-copilot.

Open the folder on GitHubat commit 727ff2e

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in github/awesome-copilot, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bigquery Pipeline Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bigquery Pipeline Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bigquery Pipeline Audit this skillgithub/awesome-copilot40k1 repos~1.3kAutomated safety check: PassMIT
Analysis Artifactswarpdotdev/oz-skills825—~1.1kAutomated safety check: PassMIT
BigQuery Slot and Cost Optimizergoogle/skills21k—~2.3kAutomated safety check: PassApache-2.0
Bigquery Bigframesgoogle/skills21k1 repos~1.3kAutomated safety check: PassApache-2.0
Data Warehouse Experimentationrampstackco/claude-skills935—~7.3kAutomated safety check: PassMIT
Io ConnectorsKilo-Org/kilo-marketplace189—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Analysis Artifacts

    warpdotdev/oz-skills

    Generate reproducible analysis artifacts — SQL queries, Python visualizations, and summary tables — as you work through a BigQuery data analysis.

    825 GitHub stars~1.1k tokensUpdated 1 mo ago
    DatabasesAuto-check passed
  • Official

    Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.

    21k GitHub stars~2.3k tokensUpdated today
    DatabasesAuto-check passed
  • Bigquery Bigframes

    google/skills

    Official

    Generates Python code using BigQuery DataFrames (BigFrames).

    21k GitHub starsUsed in 1 repo~1.3k tokens
    DatabasesAuto-check passed
  • Data Warehouse Experimentation

    rampstackco/claude-skills

    Running experiments out of the data warehouse instead of via dedicated experiment platforms.

    935 GitHub stars~7.3k tokensUpdated today
    DatabasesAuto-check passed
  • Io Connectors

    Kilo-Org/kilo-marketplace

    Guides development and usage of I/O connectors in Apache Beam.

    189 GitHub stars~1.3k tokensUpdated 9 days ago
    Testing & QAAuto-check passed
  • Chdb SQL

    vemetric/vemetric

    A skill your agent uses when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse…

    394 GitHub starsUsed in 1 repo~1.2k tokens
    DatabasesAuto-check passed

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Categories

Questions about Bigquery Pipeline Audit

What does Bigquery Pipeline Audit do?

Audits Python + BigQuery pipelines for cost safety, idempotency, and production readiness. Bigquery Pipeline Audit is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Audits Python + BigQuery pipelines for cost safety, idempotency, and production readiness.

When should I use Bigquery Pipeline Audit?

Bigquery Pipeline Audit fits situations like: tasks that involve Data warehousing.

How do I install Bigquery Pipeline Audit in Claude Code?

Run `npx skills add github/awesome-copilot --skill bigquery-pipeline-audit -a claude-code`. Or copy the skill folder (skills/bigquery-pipeline-audit in github/awesome-copilot) into .claude/skills/bigquery-pipeline-audit in your project. Claude Code loads it when a task matches its description.

How do I install Bigquery Pipeline Audit in Codex?

Run `npx skills add github/awesome-copilot --skill bigquery-pipeline-audit -a codex`. Or copy the skill folder (skills/bigquery-pipeline-audit in github/awesome-copilot) into .agents/skills/bigquery-pipeline-audit in your project. Codex loads it when a task matches its description.

Can I use Bigquery Pipeline Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill bigquery-pipeline-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bigquery-pipeline-audit, .gemini/skills/bigquery-pipeline-audit, .github/skills/bigquery-pipeline-audit and .opencode/skills/bigquery-pipeline-audit in your project.

What does Bigquery Pipeline Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Bigquery Pipeline Audit is instructions for the agent only. Our summary lists: Python 3.

Does Bigquery Pipeline Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bigquery Pipeline Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bigquery Pipeline Audit use?

Bigquery Pipeline Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bigquery Pipeline Audit use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bigquery Pipeline Audit?

Skills that share tags, products or a category with Bigquery Pipeline Audit: Analysis Artifacts (warpdotdev/oz-skills, 825 stars), BigQuery Slot and Cost Optimizer (google/skills, 21k stars), Bigquery Bigframes (google/skills, 21k stars) and Data Warehouse Experimentation (rampstackco/claude-skills, 935 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bigquery Pipeline Audit?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.