Data Engineer
davila7/claude-code-templates
Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.
Official agent skill
by google in google/skills
Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine).
$ npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install google/skills google-cloud-solution-agentic-analytics-spark-knowledge-catalog --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog .claude/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "google-cloud-solution-agentic-analytics-spark-knowledge-catalog" agent skill from https://github.com/google/skills/tree/main/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog into .claude/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-cloud-solution-agentic-analytics-spark-knowledge-catalog", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/google/skills/tree/main/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalogType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install google/skills google-cloud-solution-agentic-analytics-spark-knowledge-catalog --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog .agents/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "google-cloud-solution-agentic-analytics-spark-knowledge-catalog" agent skill from https://github.com/google/skills/tree/main/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog into .agents/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-cloud-solution-agentic-analytics-spark-knowledge-catalog", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install google/skills google-cloud-solution-agentic-analytics-spark-knowledge-catalog --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog .cursor/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "google-cloud-solution-agentic-analytics-spark-knowledge-catalog" agent skill from https://github.com/google/skills/tree/main/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog into .cursor/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-cloud-solution-agentic-analytics-spark-knowledge-catalog", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/google/skills.git --path skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install google/skills google-cloud-solution-agentic-analytics-spark-knowledge-catalog --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog .gemini/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "google-cloud-solution-agentic-analytics-spark-knowledge-catalog" agent skill from https://github.com/google/skills/tree/main/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog into .gemini/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-cloud-solution-agentic-analytics-spark-knowledge-catalog", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install google/skills google-cloud-solution-agentic-analytics-spark-knowledge-catalogInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog .github/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "google-cloud-solution-agentic-analytics-spark-knowledge-catalog" agent skill from https://github.com/google/skills/tree/main/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog into .github/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-cloud-solution-agentic-analytics-spark-knowledge-catalog", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install google/skills google-cloud-solution-agentic-analytics-spark-knowledge-catalog --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/google/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog .opencode/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "google-cloud-solution-agentic-analytics-spark-knowledge-catalog" agent skill from https://github.com/google/skills/tree/main/skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog into .opencode/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "google-cloud-solution-agentic-analytics-spark-knowledge-catalog", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
google-cloud-solution-agentic-analytics-spark-knowledge-catalogDiscovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine).
Google Cloud Solution Agentic Analytics Spark Knowledge Catalog is an agent skill from google/skills, published by the product's own GitHub organization. Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine). Use when designing data science and analytics workflows across structured and unstructured distributed data (including in S3, Azure Blob, AlloyDB, and Iceberg), establishing metadata governance with Knowledge Catalog aspect types, or grounding agentic IDEs (VS Code, Antigravity) by using the Google Cloud Data Agent Kit. Don't use for provisioning…
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files and assets (for example `assets/output-template.md`, `references/design-recommendations.md` and `references/knowledge-catalog-documentation.md`).
It sits in Data & Analytics, covering Data warehousing and File uploads and storage. It works with Google Cloud, Microsoft Azure, Visual Studio Code and Apache Spark. The repository describes itself as: Agent Skills for Google products and technologies. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8a1ac05. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comAlso links to:
docs.cloud.google.comcodelabs.developers.google.comdevelopers.google.comdeveloperknowledge.googleapis.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Google Cloud Solution Agentic Analytics Spark Knowledge Catalog loads about 4.4k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 173 tokens; SKILL.md has 1,982 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from google/skills at commit 8a1ac05, republished under its Apache-2.0 licence (© google). 1,982 words, ~4,381 tokens.
.claude/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.This skill provides a workflow to design and implement a governed, secure pipeline for agentic analytics solution across structured and unstructured data that's distributed across Google Cloud, on-premises systems, and other cloud providers.
The workflow consists of the following phases:
Important notes about the workflow:
Request the user to describe the functional requirements (business processes, activities, and use cases) of their workload. Ask the user the following questions, one question at a time:
Request the user to describe the non-functional requirements of their workload.
The following are examples of questions you can ask to gather non-functional requirements:
Ask the user whether the workload currently runs on other cloud providers or on-premises.
Request the user to describe dependencies, if any, on other workloads, products, or tools. The following are examples of questions that you can ask to get information about the dependencies:
Review the input that the user has provided so far, and check whether there are any ambiguities or contradictions.
If you identify any ambiguities or contradictions in the requirements that the user has provided (e.g., zero-copy vs copying data to a repository), then do the following for each ambiguity or contradiction that you identify:
Critical: Until all the ambiguities and contradictions that you identify are resolved according to the preceding guidance, you must NOT recommend or generate any architecture design, technical decomposition, or Google Cloud product recommendations.
Important: DON'T start this step if there are unresolved contradictions or ambiguities from Step 5.
Generate a technical decomposition of the components of the workload.
Request the user to approve the generated technical decomposition.
If the user requests changes, then generate an updated technical decomposition.
Repeat steps 5 through 8 until the user approves the generated technical decomposition.
After the user approves the technical decomposition, proceed to Phase 2. Important: Don't proceed to the next phase until the user approves the generated technical decomposition of the workload.
For each task in this phase, to ensure that the generated content aligns with the latest and official Google Cloud guidance, ground the generated content by using the following resources:
developerknowledge:search_documentsdeveloperknowledge:get_documentsdeveloperknowledge:answer_queryreferences/product-selection-guidance.mdhttps://github.com/google/skills/blob/main/skills/cloud/google-cloud-solution-architecture/references/decision-making-guides.mdGenerate design recommendations and best practices to optimally configure each component in the architecture based on the workload's requirements.
Important:
references/design-recommendations.md.references/knowledge-catalog-documentation.mdgoogle-cloud-waf-securitygoogle-cloud-waf-reliabilitygoogle-cloud-waf-cost-optimizationgoogle-cloud-waf-operational-excellencegoogle-cloud-waf-performance-optimizationgoogle-cloud-waf-sustainabilityPresent the generated recommendations to the user and ask whether the user needs any changes.
If the user needs changes, then make the required changes.
Repeat steps 2 and 3 until the user confirms that the generated design recommendations meet their requirements.
Proceed to Task 2.5.
Generate deployment guidance, including code and instructions to enable the user to deploy the solution.
Important:
Present the generated deployment guidance to the user and ask whether the user needs any changes.
If the user requests changes, then make the required changes.
Repeat steps 2 and 3 until the user confirms that the generated deployment guidance meets their requirements.
Proceed to Phase 3.
curl or gcloud to perform
the steps in the approved validation plan.solution-architecture-guide.md, based on
the template in assets/output-template.md.© google, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references, assets) in skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog of google/skills.
Open the folder on GitHubat commit 8a1ac05
Google Cloud Solution Agentic Analytics Spark Knowledge Catalog next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Google Cloud Solution Agentic Analytics Spark Knowledge Catalog this skillgoogle/skills | 21k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| Data Engineerdavila7/claude-code-templates | 32k | 7 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Erd Studio Setupliam-machine/erd-studio | 165 | — | ~8.5k | Automated safety check: Pass | Custom licence | |
| Azure Storage Blob Pymicrosoft/skills | 3.1k | — | ~2.3k | Automated safety check: Pass | MIT | |
| Using Duckdbdata-goblin/power-bi-agentic-development | 1k | — | ~1.2k | Automated safety check: Pass | GPL-3.0 | |
| Cloud Retention Configmukul975/Privacy-Data-Protection-Skills | 295 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 |
davila7/claude-code-templates
Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.
liam-machine/erd-studio
Friendly, step-by-step setup for ERD Studio in an existing dbt project, for people who may be new to dbt or data modelling.
microsoft/skills
Azure Blob Storage SDK for Python. An agent skill from microsoft/skills.
data-goblin/power-bi-agentic-development
Query Fabric lakehouse and warehouse data using DuckDB, either locally or inside a Fabric notebook.
mukul975/Privacy-Data-Protection-Skills
Configures cloud storage retention policies across AWS S3, Azure Blob Storage, and Google Cloud Storage.
rohitg00/awesome-claude-code-toolkit
Data engineering patterns for ETL pipelines, data warehousing, Apache Spark, and data quality validation
google/skills
Manages Google Cloud Privileged Access Manager entitlements and grants: create and edit entitlements, request temporary access, and approve or deny pending grants.
google/skills
Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.
google/skills
Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.
google/skills
Searches, manages and scaffolds skills in the Gemini Enterprise Agent Platform Skill Registry using bundled Python scripts and Google Cloud credentials.
google/skills
Designs GCP infrastructure as local Terraform, validates and scans it against best practices, then imports it to Application Design Center for deployment and troubleshooting.
google/skills
Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans.
Categories
Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine). Google Cloud Solution Agentic Analytics Spark Knowledge Catalog is an agent skill from google/skills, published by the product's own GitHub organization. Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine).
Google Cloud Solution Agentic Analytics Spark Knowledge Catalog fits situations like: designing data science and analytics workflows across structured and unstructured distributed data (including in S3; establishing metadata governance with Knowledge Catalog aspect types; grounding agentic IDEs (VS Code; antigravity) by using the Google Cloud Data Agent Kit.
Run `npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog -a claude-code`. Or copy the skill folder (skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog in google/skills) into .claude/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog in your project. Claude Code loads it when a task matches its description.
Run `npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog -a codex`. Or copy the skill folder (skills/cloud/google-cloud-solution-agentic-analytics-spark-knowledge-catalog in google/skills) into .agents/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google/skills --skill google-cloud-solution-agentic-analytics-spark-knowledge-catalog -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog, .gemini/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog, .github/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog and .opencode/skills/google-cloud-solution-agentic-analytics-spark-knowledge-catalog in your project.
SKILL.md names no scripts, command-line tools or credentials: Google Cloud Solution Agentic Analytics Spark Knowledge Catalog is instructions for the agent only.
SKILL.md names 5 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.cloud.google.com, codelabs.developers.google.com, developers.google.com and developerknowledge.googleapis.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Google Cloud Solution Agentic Analytics Spark Knowledge Catalog is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Google Cloud Solution Agentic Analytics Spark Knowledge Catalog: Data Engineer (davila7/claude-code-templates, 32k stars), Erd Studio Setup (liam-machine/erd-studio, 165 stars), Azure Storage Blob Py (microsoft/skills, 3.1k stars) and Using Duckdb (data-goblin/power-bi-agentic-development, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
google (a GitHub organization, an official publisher) maintains it in google/skills, which has 20,994 GitHub stars. The repository holds 145 skills in this directory. The repository was last updated on October 6, 2026.
Source: google/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.