Agent skill

Enrichment Data Sourcing

by gtmagents in gtmagents/gtm-agents

Chooses and orders data providers for email, phone, company and intent enrichment, with waterfall sequences and credit-saving tactics across 150+ sources.

Apache-2.0Auto-check passedSales & Support

Install Enrichment Data Sourcing

skills CLI
$ npx skills add gtmagents/gtm-agents --skill data-sourcing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gtmagents/gtm-agents data-sourcing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gtmagents/gtm-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/data-enrichment-master/skills/data-sourcing .claude/skills/data-sourcing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-sourcing
GitHub stars
414
Used in
2 other repos
Token cost
~2.4k tokens
SKILL.md length
629 words
Files
3 (incl. scripts, references)
Skills in repo
121
Repo updated
First seen
Licence
Apache-2.0

At a glance

Chooses and orders data providers for email, phone, company and intent enrichment, with waterfall sequences and credit-saving tactics across 150+ sources.

  • Works in 5 steps: Quality-Cost Balance: Optimize for… → Smart Routing: Route requests to… → Waterfall Logic: Use sequential provider… → …
  • Picking a provider stack for email, phone or company enrichment
  • SKILL.md covers When to Use, Framework, Templates and Tips
  • Runs Python scripts from its folder

What it does

This skill is a framework for selecting and routing among more than 150 enrichment providers so data quality stays high and credit spend stays low. Its principles are balancing quality against cost, routing by input type and likelihood of success, running providers in waterfall order, reusing cached data and batching similar requests for volume discounts. A provider selection matrix gives preferred orderings for email discovery by input type (a LinkedIn URL, a name and company, a domain only, or an email to validate) and for company data by type, such as firmographics, financials, technology stack, intent signals and news.

Email providers are grouped into premium, standard and budget tiers with success rates of 90%, 75% and 60% or better, and industry notes cover startups, enterprise, e-commerce, healthcare and financial services. A credit tier scheme starts with free cached or native operations and rises through 0.5-credit validations to standard enrichments at 1 to 2 credits. The folder includes references/provider_cheat_sheet.md and scripts/cost_calculator.py for estimating cost.

When your agent uses it

  • Picking a provider stack for email, phone or company enrichment
  • Building or tuning a waterfall sequence to raise match rates
  • Auditing credit consumption or provider performance
  • Designing enrichment logic for a RevOps or data engineering team

Example prompts

  • “We only have domains for these leads. Which providers should we try for emails, and in what order?”
  • “Our enrichment credits ran out early this month. Audit the waterfall and suggest cheaper routing.”
  • “Estimate the credit cost of enriching our whole account list with firmographics and technology stack.”

Requirements

  • Python to run scripts/cost_calculator.py
  • Accounts with the enrichment providers you decide to use

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Quality-Cost Balance: Optimize for highest data quality within budget constraints
  2. Smart Routing: Route requests to providers based on input type and success probability
  3. Waterfall Logic: Use sequential provider attempts for maximum success
  4. Caching Strategy: Leverage cached data to reduce redundant API calls
  5. Bulk Optimization: Process similar requests together for volume discounts

What it can do on your machine

Read from SKILL.md and the folder at commit 78e0419. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Enrichment Data Sourcing loads about 2.4k tokens when it runs, and up to ~2.5k if it reads all its reference files. Until then it costs about 33 tokens; SKILL.md has 629 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from gtmagents/gtm-agents at commit 78e0419, republished under its Apache-2.0 licence (© gtmagents). 629 words, ~2,420 tokens.

Download SKILL.mdSave it as .claude/skills/data-sourcing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
data-sourcing
description
Optimize provider selection, routing, and credit usage across 150+ enrichment sources for company/contact intelligence.

Data Sourcing & Provider Optimization Skill

When to Use

  • Selecting provider stacks for email, phone, company, or intent enrichment
  • Building or tuning waterfall sequences to improve success rates
  • Auditing credit consumption or provider performance
  • Designing enrichment logic for GTM ops, RevOps, or data engineering teams

Framework

You are an expert at selecting and optimizing data providers from 150+ available options to maximize data quality while minimizing credit costs. Use this layered framework to keep enrichment predictable and efficient.

Core Principles
  1. Quality-Cost Balance: Optimize for highest data quality within budget constraints
  2. Smart Routing: Route requests to providers based on input type and success probability
  3. Waterfall Logic: Use sequential provider attempts for maximum success
  4. Caching Strategy: Leverage cached data to reduce redundant API calls
  5. Bulk Optimization: Process similar requests together for volume discounts
Provider Selection Matrix
For Email Discovery

Best Input Scenarios:

  • Have LinkedIn URL: ContactOut → RocketReach → Apollo
  • Have Name + Company: Apollo → Hunter → RocketReach → FindyMail
  • Have Domain Only: Hunter → Apollo → Clearbit
  • Have Email (need validation): ZeroBounce → NeverBounce → Debounce

Quality Tiers:

  • Premium (90%+ success): ZoomInfo, BetterContact waterfall
  • Standard (75%+ success): Apollo, Hunter, RocketReach
  • Budget (60%+ success): Snov.io, Prospeo, ContactOut
For Company Intelligence

Data Type Priority:

  • Basic Firmographics: Clearbit (fastest) → Ocean.io → Apollo
  • Financial Data: Crunchbase → PitchBook → Dealroom
  • Technology Stack: BuiltWith → HG Insights → Clearbit
  • Intent Signals: B2D AI → ZoomInfo Intent → 6sense
  • News & Social: Google News → Social platforms → Owler

Industry Specialization:

  • Startups: Crunchbase, Dealroom, AngelList
  • Enterprise: ZoomInfo, D&B, HG Insights
  • E-commerce: Store Leads, BuiltWith, Shopify data
  • Healthcare: Definitive Healthcare + compliance providers
  • Financial Services: PitchBook, S&P Capital IQ
Credit Optimization Strategies
Cost Tiers
Tier 0 (Free): Native operations, cached data, manual inputs
Tier 1 (0.5 credits): Validation, verification, basic lookups
Tier 2 (1-2 credits): Standard enrichments (Apollo, Hunter, Clearbit)
Tier 3 (2-3 credits): Premium data (ZoomInfo, technographics, intent)
Tier 4 (3-5 credits): Enterprise intelligence (PitchBook, custom AI)
Tier 5 (5-10 credits): Specialized services (video generation, deep AI research)
Optimization Tactics

1. Cache Everything

  • Email: 30-day cache
  • Company: 90-day cache
  • Intent: 7-day cache
  • Static data: Indefinite cache

2. Batch Processing

python
# Process in batches for volume discounts
if record_count > 1000:
    use_provider("apollo_bulk")  # 10-30% discount
elif record_count > 100:
    use_parallel_processing()
else:
    use_standard_processing()

3. Smart Waterfalls

python
waterfall_sequence = [
    {"provider": "cache", "credits": 0},
    {"provider": "apollo", "credits": 1.5, "stop_if_success": True},
    {"provider": "hunter", "credits": 1.2, "stop_if_success": True},
    {"provider": "bettercontact", "credits": 3, "stop_if_success": True},
    {"provider": "ai_research", "credits": 5, "last_resort": True}
]
Provider-Specific Optimizations
Apollo.io
  • Strengths: US B2B, LinkedIn data, phone numbers
  • Weaknesses: International coverage, personal emails
  • Tips: Use bulk API for 10%+ discount, batch similar companies
ZoomInfo
  • Strengths: Enterprise data, org charts, intent signals
  • Weaknesses: Expensive, SMB coverage
  • Tips: Reserve for high-value accounts, negotiate enterprise deals
Hunter
  • Strengths: Domain searches, email patterns, API reliability
  • Weaknesses: Phone numbers, detailed contact info
  • Tips: Best for initial domain exploration, use pattern detection
Clearbit
  • Strengths: Real-time API, company data, speed
  • Weaknesses: Email discovery rates, phone numbers
  • Tips: Great for instant enrichment, combine with others for contacts
Show full SKILL.md (254 more words)Show less
BuiltWith
  • Strengths: Technology detection, historical data, e-commerce
  • Weaknesses: Contact information, company financials
  • Tips: Filter accounts by technology before enrichment
Waterfall Strategies
Maximum Success Waterfall
yaml
Priority: Success rate over cost
Sequence:
  1. BetterContact (aggregates 10+ sources)
  2. ZoomInfo (if enterprise)
  3. Apollo + Hunter + RocketReach
  4. AI web research
Expected Success: 95%+
Average Cost: 8-12 credits
Balanced Waterfall
yaml
Priority: Good success with reasonable cost
Sequence:
  1. Apollo.io
  2. Hunter (if domain match)
  3. RocketReach (if name match)
  4. Stop or continue based on confidence
Expected Success: 80%
Average Cost: 3-5 credits
Budget Waterfall
yaml
Priority: Minimize cost
Sequence:
  1. Cache check
  2. Hunter (domain only)
  3. Free sources (Google, LinkedIn public)
  4. Stop at first result
Expected Success: 60%
Average Cost: 1-2 credits
Quality Scoring Framework
python
def calculate_data_quality_score(data, sources):
    score = 0
    
    # Multi-source validation (30 points)
    if len(sources) > 1:
        score += min(len(sources) * 10, 30)
    
    # Data completeness (30 points)
    required_fields = ["email", "phone", "title", "company"]
    score += sum(10 for field in required_fields if data.get(field))
    
    # Verification status (20 points)
    if data.get("email_verified"):
        score += 10
    if data.get("phone_verified"):
        score += 10
    
    # Recency (20 points)
    days_old = get_data_age(data)
    if days_old < 30:
        score += 20
    elif days_old < 90:
        score += 10
    
    return score
Industry-Specific Provider Selection
SaaS/Technology
  • Primary: Apollo, Clearbit, BuiltWith
  • Secondary: ZoomInfo, HG Insights
  • Intent: G2, TrustRadius, 6sense
Financial Services
  • Primary: PitchBook, ZoomInfo
  • Compliance: LexisNexis, D&B
  • News: Bloomberg, Reuters
Healthcare
  • Primary: Definitive Healthcare
  • Compliance: NPPES, state boards
  • Standard: ZoomInfo with healthcare filters
E-commerce
  • Primary: Store Leads, BuiltWith
  • Platform-specific: Shopify, Amazon seller data
  • Standard: Clearbit with e-commerce signals
Troubleshooting Common Issues
Low Email Discovery Rate
  • Check email patterns with Hunter
  • Try personal email providers
  • Use AI research for executives
  • Consider LinkedIn outreach instead
High Credit Usage
  • Audit waterfall sequences
  • Increase cache TTL
  • Negotiate volume deals
  • Use native operations first
Poor Data Quality
  • Add verification steps
  • Cross-reference multiple sources
  • Set minimum confidence thresholds
  • Implement human review for critical data
Advanced Techniques
Hybrid Enrichment
python
# Combine AI and traditional providers
def hybrid_enrichment(company):
    # Fast, cheap base data
    base = clearbit_lookup(company)
    
    # AI for missing pieces
    if not base.get("description"):
        base["description"] = ai_generate_description(company)
    
    # Premium for high-value
    if is_enterprise_account(base):
        base.update(zoominfo_enrich(company))
    
    return base
Progressive Enrichment
python
# Enrich in stages based on engagement
def progressive_enrichment(lead):
    # Stage 1: Basic (on import)
    if lead.stage == "new":
        return basic_enrichment(lead)  # 1-2 credits
    
    # Stage 2: Engaged (opened email)
    elif lead.stage == "engaged":
        return standard_enrichment(lead)  # 3-5 credits
    
    # Stage 3: Qualified (booked meeting)
    elif lead.stage == "qualified":
        return comprehensive_enrichment(lead)  # 10+ credits

Templates

  • Provider Cheat Sheet: See references/provider_cheat_sheet.md for provider selection.
  • Cost Calculator: See scripts/cost_calculator.py for estimating credit usage.
  • Integration Code Templates:
javascript
// JavaScript/Node.js template
const enrichContact = async (name, company) => {
  // Check cache first
  const cached = await checkCache(name, company);
  if (cached) return cached;
  
  // Try providers in sequence
  const providers = ['apollo', 'hunter', 'rocketreach'];
  
  for (const provider of providers) {
    try {
      const result = await callProvider(provider, {name, company});
      if (result.email) {
        await saveToCache(result);
        return result;
      }
    } catch (error) {
      console.log(`${provider} failed, trying next...`);
    }
  }
  
  // Fallback to AI research
  return await aiResearch(name, company);
};

Tips

  • Pre-build waterfalls per motion so GTM teams can call a single orchestration command rather than juggling providers.
  • Instrument cache hit rates; alert RevOps when cache effectiveness drops below target to avoid spike in credits.
  • Rotate premium providers each quarter to negotiate better volume discounts and diversify coverage gaps.
  • Pair enrichment with QA hooks (e.g., verification APIs, sampling) before syncing into CRM to prevent bad data cascades.

Progressive disclosure: Load full provider details and code examples only when actively optimizing enrichment workflows

© gtmagents, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in plugins/data-enrichment-master/skills/data-sourcing of gtmagents/gtm-agents.

  • SKILL.md
  • references/provider_cheat_sheet.md
  • scripts/cost_calculator.py

Open the folder on GitHubat commit 78e0419

Used in 2 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in gtmagents/gtm-agents, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Enrichment Data Sourcing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Enrichment Data Sourcing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Enrichment Data Sourcing this skillgtmagents/gtm-agents4142 repos~2.4kAutomated safety check: PassApache-2.0
GEO Prospect Trackerzubair-trabzada/geo-seo-claude11k—~1.7kAutomated safety check: NotesMIT
Lead Gen Tool Builderexplorium-ai/gtm-skills185—~1.8kAutomated safety check: NotesMIT
B2B Lead Generationminhnv0807/ai-business-skills609—~1.2kAutomated safety check: PassMIT
Lead ScoringLeoYeAI/openclaw-master-skills2.2k—~4.6kAutomated safety check: PassMIT
Cold Outreachericrisco/rsc-harness180—~4.1kAutomated safety check: PassMIT

Similar skills

  • GEO Prospect Tracker

    zubair-trabzada/geo-seo-claude

    Tracks GEO agency leads and clients through a sales pipeline in a local JSON file, with notes, audit scores, deal values and a pipeline summary.

    11k GitHub stars~1.7k tokensUpdated today
    Sales & SupportAuto-check: notes
  • Lead Gen Tool Builder

    explorium-ai/gtm-skills

    Lead generation tool builder skill for Claude Code and Codex: scaffolds a complete, self-hostable, ZoomInfo-style B2B lead-generation web app — company & contact search UI, firmographic and…

    185 GitHub stars~1.8k tokensUpdated 2 days ago
    Sales & SupportAuto-check: notes
  • B2B Lead Generation

    minhnv0807/ai-business-skills

    Plans B2B pipeline work from ICP definition and prospecting through lead scoring, outbound sequences, sales assets and MQL to SQL handoff, aiming at qualified pipeline.

    609 GitHub stars~1.2k tokensUpdated 27 days ago
    Sales & SupportAuto-check passed
  • Lead Scoring

    LeoYeAI/openclaw-master-skills

    Set up and automate lead scoring for HubSpot and other CRMs.

    2.2k GitHub stars~4.6k tokensUpdated 2 mo ago
    Sales & SupportAuto-check passed
  • Cold Outreach

    ericrisco/rsc-harness

    A skill your agent uses when writing a cold email or LinkedIn DM to a stranger and its cadence: first-touch copy under a word ceiling, 4-7 step bump sequences, per-inbox volume and warm-up limits…

    180 GitHub stars~4.1k tokensUpdated today
    Sales & SupportAuto-check passed
  • Lead Gen

    ericrisco/rsc-harness

    A skill your agent uses when building and qualifying a prospect list before anyone reaches out — a falsifiable ICP, named accounts/contacts from Apollo/ZoomInfo/Clay, deduped against the CRM, tiered…

    180 GitHub stars~2.6k tokensUpdated today
    Sales & SupportAuto-check passed

More from gtmagents/gtm-agents

All 121 skills in this repo
  • Cold Email Personalization

    gtmagents/gtm-agents

    Writes personalized B2B cold emails from prospect research, using a gated workflow, a quality rubric and follow-up sequences.

    414 GitHub starsUsed in 1 repo~853 tokens
    Auto-check passed
  • Discovery Call Playbook

    gtmagents/gtm-agents

    Structures sales discovery calls with a PREP routine, a five-part call flow, a question bank, a qualification scorecard and a follow-up recap email.

    414 GitHub starsUsed in 1 repo~368 tokens
    Auto-check passed
  • Frames marketing automation around lifecycle stages, signals, touches and SLAs, with worksheets for onboarding, expansion, renewal and churn-prevention journeys.

    414 GitHub starsUsed in 1 repo~927 tokens
    Auto-check passed
  • Social Selling

    gtmagents/gtm-agents

    A skill your agent uses when engaging prospects through LinkedIn, communities, and social channels to spark warm conversations and meetings.

    414 GitHub starsUsed in 1 repo~431 tokens
    Auto-check passed
  • Drip Campaigns

    gtmagents/gtm-agents

    A skill your agent uses when you need to map sequenced nurture flows with pacing, storytelling arcs, and value ladders.

    414 GitHub starsUsed in 1 repo~310 tokens
    Auto-check passed
  • Editorial Ops

    gtmagents/gtm-agents

    A skill your agent uses when planning multi-channel editorial calendars, enforcing publishing cadences, and coordinating distribution workflows across GTM teams.

    414 GitHub starsUsed in 1 repo~821 tokens
    Auto-check passed

Questions about Enrichment Data Sourcing

What does Enrichment Data Sourcing do?

Chooses and orders data providers for email, phone, company and intent enrichment, with waterfall sequences and credit-saving tactics across 150+ sources. This skill is a framework for selecting and routing among more than 150 enrichment providers so data quality stays high and credit spend stays low. Its principles are balancing quality against cost, routing by input type and likelihood of success, running providers in waterfall order, reusing cached data and batching similar requests for volume discounts.

When should I use Enrichment Data Sourcing?

Enrichment Data Sourcing fits situations like: picking a provider stack for email, phone or company enrichment; building or tuning a waterfall sequence to raise match rates; auditing credit consumption or provider performance; designing enrichment logic for a RevOps or data engineering team.

How do I install Enrichment Data Sourcing in Claude Code?

Run `npx skills add gtmagents/gtm-agents --skill data-sourcing -a claude-code`. Or copy the skill folder (plugins/data-enrichment-master/skills/data-sourcing in gtmagents/gtm-agents) into .claude/skills/data-sourcing in your project. Claude Code loads it when a task matches its description.

How do I install Enrichment Data Sourcing in Codex?

Run `npx skills add gtmagents/gtm-agents --skill data-sourcing -a codex`. Or copy the skill folder (plugins/data-enrichment-master/skills/data-sourcing in gtmagents/gtm-agents) into .agents/skills/data-sourcing in your project. Codex loads it when a task matches its description.

Can I use Enrichment Data Sourcing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gtmagents/gtm-agents --skill data-sourcing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-sourcing, .gemini/skills/data-sourcing, .github/skills/data-sourcing and .opencode/skills/data-sourcing in your project.

What does Enrichment Data Sourcing need to run?

Going by SKILL.md and its folder, Enrichment Data Sourcing needs Python for the scripts in its folder. Our summary lists: Python to run scripts/cost_calculator.py; Accounts with the enrichment providers you decide to use.

Does Enrichment Data Sourcing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Enrichment Data Sourcing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Enrichment Data Sourcing use?

Enrichment Data Sourcing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Enrichment Data Sourcing use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 111 tokens, read only when the agent opens those files.

What are the alternatives to Enrichment Data Sourcing?

Skills that share tags, products or a category with Enrichment Data Sourcing: GEO Prospect Tracker (zubair-trabzada/geo-seo-claude, 11k stars), Lead Gen Tool Builder (explorium-ai/gtm-skills, 185 stars), B2B Lead Generation (minhnv0807/ai-business-skills, 609 stars) and Lead Scoring (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Enrichment Data Sourcing?

gtmagents (a GitHub user) maintains it in gtmagents/gtm-agents, which has 414 GitHub stars. The repository holds 121 skills in this directory. The repository was last updated on April 3, 2026.

Source: gtmagents/gtm-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.