Agent skill

Source Baseline Matrix Research Skill

by GoogleCloudPlatform in GoogleCloudPlatform/spanner-migration-tool

Skill for helping achieve datatype and feature mapping matrix for any source against spanner.

Apache-2.0Auto-check passed

Install Source Baseline Matrix Research Skill

skills CLI
$ npx skills add GoogleCloudPlatform/spanner-migration-tool --skill source-baseline-matrix-research-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GoogleCloudPlatform/spanner-migration-tool source-baseline-matrix-research-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GoogleCloudPlatform/spanner-migration-tool.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/source_research_helper .claude/skills/source-baseline-matrix-research-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
source-baseline-matrix-research-skill
GitHub stars
155
Token cost
~3.4k tokens
SKILL.md length
1,760 words
Files
3
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Skill for helping achieve datatype and feature mapping matrix for any source against spanner.

  • Works in 5 steps: Goal → Research & Extraction Process → Creating the Logical Type Mapping Table → …
  • SKILL.md covers 1. Goal, 2. Research & Extraction Process, 3. Creating the Logical Type… and 4. Feature Support Discovery, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Source Baseline Matrix Research Skill is an agent skill from GoogleCloudPlatform/spanner-migration-tool. Skill for helping achieve datatype and feature mapping matrix for any source against spanner. It does research and provide baseline mapping matrix on which user can expand upto.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files.

It works with Google Cloud. The repository describes itself as: Spanner migration tool helps you migrate from your existing database/dump files to Spanner. The licence is Apache-2.0.

Example prompts

  • “/source-baseline-matrix-research-skill”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Goal
  2. Research & Extraction Process
  3. Creating the Logical Type Mapping Table
  4. Feature Support Discovery
  5. Dual-Format Artifact Generation & User Approval

What it can do on your machine

Read from SKILL.md and the folder at commit 345e09c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Source Baseline Matrix Research Skill loads about 3.4k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 1,760 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GoogleCloudPlatform/spanner-migration-tool at commit 345e09c, republished under its Apache-2.0 licence (© GoogleCloudPlatform). 1,760 words, ~3,407 tokens.

Download SKILL.mdSave it as .claude/skills/source-baseline-matrix-research-skill/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
source-baseline-matrix-research-skill
description
Skill for helping achieve datatype and feature mapping matrix for any source against spanner. It does research and provide baseline mapping matrix on which user can expand upto.

New Source DataType Mapping Research Skill

This skill instructs an AI Research Agent to discover, analyze, and compile an exhaustive, dialect-wide data type mapping matrix from a new database source dialect to Spanner standard SQL dialects.

This generated mapping file will act as the "Ground Truth" artifact during the Datatype Integration Testing phases.


1. Goal

Programmatically extract and research a comprehensive list of all native datatypes supported by the target database dialect, determine the native Spanner standard SQL mappings for them (both primary and alternative paths), and generate an exhaustive logical mapping matrix spreadsheet/markdown document for user approval.


2. Research & Extraction Process

[!IMPORTANT] Mandatory Tool-Based Raw Extraction Requirement: You MUST NOT rely on your internal LLM memory or knowledge base to list data types, as this inevitably suffers from token limits and summarization biases. Instead, you MUST write and execute a script (e.g., Python/Bash) or use web/file-reading tools to programmatically scrape/extract EVERY single data type from the official database documentation.

Anti-Summarization Rule: You are strictly forbidden from summarizing or limiting the list to "common" migration types. You must actively search for and include the database's "long tail" data types. Do not pre-filter types: Do not silently omit obscure, legacy, or system data types. You MUST include them in the mapping matrix and mark them as Unsupported (Scenario D) with a rationale if they are not supported in spanner. Exhaustive means exhaustive.

2.1 Programmatic Raw Extraction (First Principles Checklist)
  • Write and execute a script (or use available web/file tools) to scrape and extract the complete, unfiltered list of all native datatypes and their official aliases directly from the official database documentation.
  • Collection & UDT Mandate: You MUST actively search for and explicitly extract all Object-Relational Collection Types and hierarchical User-Defined Types. Do not restrict your extraction exclusively to simple scalar primitives.
  • Save this exhaustive list of extracted data types to a local text file named raw_types_extraction.txt. This file forms the baseline absolute ground truth checklist.
2.2 Hybrid Extraction Strategy: Undocumented/Experimental Research Pass
  • The official documentation is often incomplete. After establishing the programmatic baseline, you MUST perform broad LLM research (using web search, code search, or internal knowledge) to discover hidden, undocumented, deprecated, or experimental data types.
  • You must explicitly search across these avenues:
    1. Release Notes & Changelogs (for newly added but undocumented types).
    2. GitHub Issues & Pull Requests (for edge-cases developers complain about).
    3. Driver Source Code (JDBC, ODBC, or official language clients, which often reveal internal type enums like Types.OTHER).
    4. Database Engine Source Code (Internal type enumerations or system catalogs like pg_type / information_schema.columns).
    5. Migration Ecosystem (Debezium, Fivetran, Kafka Connect connector docs often document bizarre edge-case types).
    6. StackOverflow & Developer Blogs.
  • If you discover any types from these avenues that are not in your raw_types_extraction.txt baseline, you MUST append them to the file.
2.3 Programmatic Validation
  • Read the final augmented raw_types_extraction.txt. You MUST generate exactly one row in the mapping table for every single line in that file.
  • Reconcile and group the types into standard categories:
    • Numeric Types
    • Character/String Types
    • Binary/LOB Types
    • Temporal Types
    • Boolean Types
    • Array/Collection Types
    • JSON/XML Types
    • Geometric/Spatial Types
    • Network Address Types
    • UUID/Identifier Types
    • Range & Multirange Types
    • User-Defined & Custom Types
    • System/Internal Types
2.4 Discover Alternative Mappings in Conversion Code & Public Dialects
  • Analyze the Dataflow template's (repository: https://github.com/GoogleCloudPlatform/DataflowTemplates) generic conversion classes (e.g., GenericRecordTypeConvertor.java and AvroToValueMapper.java) to understand what alternative mappings are supported via user overrides.
  • Leverage Public Mapping References (Cheat Sheets): Research how mature open-source projects or JDBC drivers map the type:
    • JDBC Driver Type Mappings: Check how the JDBC driver handles the type (e.g., mapping IPADDRESS to binary format).
    • Beam SchemaUtil: Check org.apache.beam.sdk.io.jdbc.SchemaUtil to see how Beam maps standard JDBC types.
  • For every supported type, identify ALL viable, structurally valid alternative Spanner mappings (Scenario B) based on this research. Do not limit to just one alternative; list all possible fallback types (e.g., INT64, STRING, FLOAT64 for numerical fallbacks).

3. Creating the Logical Type Mapping Table

Construct a markdown table containing exactly these columns:

  • Mapping Category
  • Source Database Type / Alias
  • Is Source Datatype Supported as PK?
  • Spanner GoogleSQL Default Datatype
  • Spanner GoogleSQL Default Datatype If Column is PK
  • Spanner GoogleSQL Alternative Datatypes
  • Spanner PostgreSQL Default Datatype
  • Spanner PostgreSQL Default Datatype If Column is PK
  • Spanner PostgreSQL Alternative Datatypes
  • Datastream Support Status
  • Edge Cases for Smoke Test
  • Notes
Column Definitions & Constraints
  • Mapping Category & Source Database Type / Alias: Group the types logically and list the exact native type or alias from the source database.
  • Is Source Datatype Supported as PK?: You MUST explicitly verify whether the source database natively allows this exact datatype to be designated as a Primary Key (or part of a composite primary key) within the source architecture. Output 'Yes' or 'No'.
  • Spanner GoogleSQL & PostgreSQL Default Datatype: Identify the safest, most logical 1-to-1 default mapping for both dialects independently (observing syntax limits like NUMERIC vs PG.NUMERIC).
  • The "Default Datatype If Column is PK" Fallback Logic: You MUST explicitly execute a web search against the official Google Cloud documentation to verify if your proposed Spanner equivalent datatype is legally permitted in a Primary Key or Key Column. You must verify this for BOTH the GoogleSQL and PostgreSQL dialects independently because they deliberately contain divergent constraints. If the source datatype CAN be used as a PK, but your proposed Spanner Default Datatype is evaluated as illegal for a PK based on the official documentation, you MUST output a safe Spanner fallback type here (chosen from your alternatives). If the default is evaluated as safe for PKs, simply mirror the default type here. If the source datatype CANNOT be used as a PK natively, simply output 'N/A' (Not Applicable) for this column.
  • Spanner GoogleSQL & PostgreSQL Alternative Datatypes: List ALL viable, structurally valid alternative mappings for both dialects.
  • Datastream Support Status: Because our live templates rely on Datastream, you MUST explicitly verify each datatype against the official Google Cloud Datastream replication support documentation. If Datastream cannot currently pull this source datatype natively, we cannot migrate it in live pipelines. Mark this explicitly as 'Supported' or 'Unsupported' based purely on Datastream's capabilities.
  • Edge Cases for Smoke Test: You MUST explicitly perform research to definitively understand the specific architectural limits of the source database datatype first (evaluating the strict max precision bounds of a dialect's temporal type, or the max-byte boundary of its numeric type). Based purely on those discovered source limits, you MUST output an array of the literal, raw test values themselves that represent the extreme boundaries and edge cases. Do not write text instructions or hypotheses. These values will be extracted and directly fed into a downstream automated test runner. You MUST provide these test values for ALL datatypes, even if the type is marked as 'Unsupported' by Datastream or Spanner natively, so the testing pipeline can validate rejection handling or explicit workaround logic.
  • Notes: You MUST explicitly document in this column whether the Spanner mapping is a straightforward native 1-to-1 translation, OR document the precise workaround tricks required to achieve the mapping (e.g., storing a geometric shape as a JSON string).

Note: Do not put citations or sources inside the table itself. Instead, create a dedicated "Citations & Sources" section at the bottom of the artifact to list all URLs, JDBC docs, internal repos, and ecosystem tools you referenced during your research.

Show full SKILL.md (558 more words)Show less

4. Feature Support Discovery

[!IMPORTANT] Anti-Hallucination Rule: You are strictly forbidden from relying on your internal LLM memory to evaluate feature support. You MUST use search tools to locate official documentation for every single feature before mapping it.

4.1 The Master Validation List & Strict Categorization Taxonomy

You must explicitly research and evaluate the source database for all of the critical architectural features listed below.

Strict Category Names: When populating the Category column in the matrix, you MUST strictly use exactly these 5 bolded category strings for the items listed beneath them. Downstream automated templates rigidly parse this output expecting these exact category strings. However, you are NOT restricted to only these categories. You must research other core database functionalities (e.g., Stored Procedures, Views, Row-Level Security) or proprietary edge cases, map them to Others category, and append them to the report alongside these mandatory minimums.

  1. Constraints: Cascading Foreign Keys, Deferred Constraints, Check/Unique Constraints.
  2. Indexes: Partial/Filtered Indexes, Function-Based Indexes, Bitmap Indexes, Full-Text/Vector Indexes.
  3. Transactions: Transaction Isolation Levels (vs Strict Serializability), Pessimistic Locking (SELECT FOR UPDATE), Autonomous Transactions.
  4. Storage: Table Partitioning (Range/Hash/List), Temporary Tables (Global vs Session), Time Travel/Flashback, Change Data Capture (CDC), Object-Relational Collections.
  5. Identity: Sequences, Identity Columns / Auto-increment, UUID Generation.
4.2 Programmatic Citation Validation & Matrix Construction
  • You must create a feature matrix constructed with the following exact column headers to account for Spanner's dual-dialect nature:
    • Category
    • Feature Supported in Source
    • Spanner GoogleSQL Equivalent Feature (If supported in Spanner GoogleSQL natively or via workaround, output the explicit Spanner feature name. If not, output 'Unsupported')
    • Notes (Spanner GoogleSQL) (Dedicated architectural space to document limitations, workarounds, or application-tier refactoring requirements specific to Spanner GoogleSQL)
    • Spanner PostgreSQL Equivalent Feature (If supported in Spanner PostgreSQL natively or via workaround, output the explicit Spanner feature name. If not, output 'Unsupported')
    • Notes (Spanner PostgreSQL) (Dedicated architectural space to document limitations, workarounds, or application-tier refactoring requirements specific to Spanner PostgreSQL)
    • Source DB Documentation URL (Direct citation proving the source database's behavior)
    • Spanner Documentation URL (Direct citation proving Spanner's side of the support claim across dialects)
  • For every feature evaluated, you MUST populate both URL columns using search tools to prevent relying on internal LLM memory.
4.3 Open-Ended Proprietary Discovery
  • After completing the Master List, you must proactively execute searches for the source database's "Architecture Overview" or "Release Notes".
  • You must discover and append at least 3 wildly unique or proprietary features specific to this source dialect that were not mentioned in the Master List, ensuring we do not miss undocumented edge cases.

5. Dual-Format Artifact Generation & User Approval

  1. Store the Draft Baseline Artifacts (Dual Formats):
    • Datatype Matrix: Create the markdown version (<source_dialect>_datatype_mapping_matrix.md) and automatically export it as a strict CSV (<source_dialect>_datatype_mapping_matrix.csv).
    • Feature Matrix: Create the markdown version (<source_dialect>_feature_mapping_matrix.md) and automatically export it as a strict CSV (<source_dialect>_feature_mapping_matrix.csv).
    • Rule: Rigorously wrap CSV fields inside double-quotes ("") to prevent delimiter breakage from commas or newlines.
  2. User Confirmation Block:
    • Present the Markdown mapping tables visually to the user.
    • Inform the user that the corresponding .md and .csv formats have been saved to the workspace.
    • Stop and ask the user to confirm/approve the mappings or provide custom overrides.
  3. Apply Overrides & Verify:
    • If the user suggests any changes, update both the .md and .csv files immediately to reflect the feedback.
    • Once the user explicitly approves the artifacts, confirm they are ready to be passed downstream.

© GoogleCloudPlatform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .agents/skills/source_research_helper of GoogleCloudPlatform/spanner-migration-tool.

  • SKILL.md
  • sampleOutput/mysql_datatype_mapping_matrix.csv
  • sampleOutput/mysql_feature_mapping_matrix.csv

Open the folder on GitHubat commit 345e09c

Compare with similar skills

Source Baseline Matrix Research Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Source Baseline Matrix Research Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Source Baseline Matrix Research Skill this skillGoogleCloudPlatform/spanner-migration-tool155—~3.4kAutomated safety check: PassApache-2.0
Gmail Inbox Watchergoogleworkspace/cli31k1 repos~476Automated safety check: PassApache-2.0
Cloud Cost Optimizationwshobson/agents40k14 repos~1.7kAutomated safety check: PassMIT
Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit2606 repos~1.1kAutomated safety check: NotesCustom licence
Terravision Cloud Diagramspatrickchugh/terravision1.6k—~5.6kAutomated safety check: NotesAGPL-3.0-only
Thesvgglincker/thesvg2.8k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Gmail Inbox Watcher

    googleworkspace/cli

    Streams new Gmail messages as NDJSON from the gws command line tool using Google Pub/Sub, with label filters, batch settings and optional per-message files.

    31k GitHub starsUsed in 1 repo~476 tokens
    Productivity & AutomationAuto-check passed
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    DevOps & CloudAuto-check passed
  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    260 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes
  • Terravision Cloud Diagrams

    patrickchugh/terravision

    Draw cloud architecture diagrams for AWS, Azure or GCP with the official provider icon sets, using TerraVision.

    1.6k GitHub stars~5.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Thesvg

    glincker/thesvg

    Fetch brand SVG logos and cloud architecture icons (AWS, Azure, GCP) from theSVG.

    2.8k GitHub stars~1.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Check Kiln's model list for deprecated or sunset models across all providers.

    5.2k GitHub stars~2.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from GoogleCloudPlatform/spanner-migration-tool

  • Add Source E2E Integ Test Smt

    GoogleCloudPlatform/spanner-migration-tool

    Skill for independently generating Integration Tests and performing End-to-End QA validation for newly added Spanner Migration Tool database sources.

    155 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Add Source Support Smt

    GoogleCloudPlatform/spanner-migration-tool

    This skill provides comprehensive architectural instructions for agents to safely implement entirely new database source engines into the Spanner Migration Tool (SMT).

    155 GitHub stars~4.2k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Source Baseline Matrix Research Skill

What does Source Baseline Matrix Research Skill do?

Skill for helping achieve datatype and feature mapping matrix for any source against spanner. Source Baseline Matrix Research Skill is an agent skill from GoogleCloudPlatform/spanner-migration-tool. Skill for helping achieve datatype and feature mapping matrix for any source against spanner.

How do I install Source Baseline Matrix Research Skill in Claude Code?

Run `npx skills add GoogleCloudPlatform/spanner-migration-tool --skill source-baseline-matrix-research-skill -a claude-code`. Or copy the skill folder (.agents/skills/source_research_helper in GoogleCloudPlatform/spanner-migration-tool) into .claude/skills/source-baseline-matrix-research-skill in your project. Claude Code loads it when a task matches its description.

How do I install Source Baseline Matrix Research Skill in Codex?

Run `npx skills add GoogleCloudPlatform/spanner-migration-tool --skill source-baseline-matrix-research-skill -a codex`. Or copy the skill folder (.agents/skills/source_research_helper in GoogleCloudPlatform/spanner-migration-tool) into .agents/skills/source-baseline-matrix-research-skill in your project. Codex loads it when a task matches its description.

Can I use Source Baseline Matrix Research Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GoogleCloudPlatform/spanner-migration-tool --skill source-baseline-matrix-research-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/source-baseline-matrix-research-skill, .gemini/skills/source-baseline-matrix-research-skill, .github/skills/source-baseline-matrix-research-skill and .opencode/skills/source-baseline-matrix-research-skill in your project.

What does Source Baseline Matrix Research Skill need to run?

SKILL.md names no scripts, command-line tools or credentials: Source Baseline Matrix Research Skill is instructions for the agent only.

Does Source Baseline Matrix Research Skill access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Source Baseline Matrix Research Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Source Baseline Matrix Research Skill use?

Source Baseline Matrix Research Skill is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Source Baseline Matrix Research Skill use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Source Baseline Matrix Research Skill?

Skills that share tags, products or a category with Source Baseline Matrix Research Skill: Gmail Inbox Watcher (googleworkspace/cli, 31k stars), Cloud Cost Optimization (wshobson/agents, 40k stars), Senior DevOps Toolkit (maslennikov-ig/claude-code-orchestrator-kit, 260 stars) and Terravision Cloud Diagrams (patrickchugh/terravision, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Source Baseline Matrix Research Skill?

GoogleCloudPlatform (a GitHub organization) maintains it in GoogleCloudPlatform/spanner-migration-tool, which has 155 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 6, 2026.

Source: GoogleCloudPlatform/spanner-migration-tool on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.