Agent skill

Proprietary Data Generator

by Affitor in Affitor/affiliate-skills

Create original surveys, benchmarks, and aggregated data nobody else has.

MITAuto-check passedData & Analytics

Install Proprietary Data Generator

skills CLI
$ npx skills add Affitor/affiliate-skills --skill proprietary-data-generator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Affitor/affiliate-skills proprietary-data-generator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Affitor/affiliate-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/automation/proprietary-data-generator .claude/skills/proprietary-data-generator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
proprietary-data-generator
GitHub stars
699
Token cost
~2.8k tokens
SKILL.md length
1,008 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Create original surveys, benchmarks, and aggregated data nobody else has.

  • Works in 5 steps: Identify Data Opportunity → Design Data Collection → Create Collection Assets → …
  • : create original data
  • SKILL.md covers Stage, When to Use, Input Schema and Workflow, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Proprietary Data Generator is an agent skill from Affitor/affiliate-skills. Create original surveys, benchmarks, and aggregated data nobody else has. Automate data collection for content moats. Triggers on: "create original data", "proprietary data", "survey design", "benchmark study", "original research", "data-driven content", "create a survey", "industry benchmark", "aggregated data", "unique data", "first-party data", "data moat", "generate research data", "create a study", "original statistics", "data nobody else has", "competitive data advantage".

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Claude Code, ChatGPT, Gemini CLI, Cursor, Windsurf, OpenClaw, any AI agent

It sits in Data & Analytics, covering Statistics. The repository describes itself as: 50 AI agent skills for affiliate marketing. Research trending content, write data-backed posts, generate infographics, build landing pages, deploy — full flywheel with social… The licence is MIT.

When your agent uses it

  • : create original data
  • Proprietary data
  • Benchmark study
  • Original research

Example prompts

  • “create original data”
  • “proprietary data”
  • “survey design”
  • “/proprietary-data-generator”

Requirements

  • Compatibility (from SKILL.md): Claude Code, ChatGPT, Gemini CLI, Cursor, Windsurf, OpenClaw, any AI agent

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Identify Data Opportunity
  2. Design Data Collection
  3. Create Collection Assets
  4. Design Automation
  5. Self-Validation

What it can do on your machine

Read from SKILL.md and the folder at commit e43bfae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code, ChatGPT, Gemini CLI, Cursor, Windsurf, OpenClaw, any AI agent

    From compatibility in the SKILL.md frontmatter.

Context cost

Proprietary Data Generator loads about 2.8k tokens when it runs. Until then it costs about 128 tokens; SKILL.md has 1,008 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Affitor/affiliate-skills at commit e43bfae, republished under its MIT licence (© Affitor). 1,008 words, ~2,830 tokens.

Download SKILL.mdSave it as .claude/skills/proprietary-data-generator/SKILL.md (or your agent's skills folder).
name
proprietary-data-generator
description
Create original surveys, benchmarks, and aggregated data nobody else has. Automate data collection for content moats. Triggers on: "create original data", "proprietary data", "survey design", "benchmark study", "original research", "data-driven content", "create a survey", "industry benchmark", "aggregated data", "unique data", "first-party data", "data moat", "generate research data", "create a study", "original statistics", "data nobody else has", "competitive data advantage".
compatibility
Claude Code, ChatGPT, Gemini CLI, Cursor, Windsurf, OpenClaw, any AI agent
license
MIT
version
1.0.0
tags
affiliate-marketing, automation, scaling, workflow, data, original-research
metadata.author
affitor
metadata.version
1.0
metadata.stage
S7-Automation

Proprietary Data Generator

Create original surveys, benchmarks, and aggregated data that nobody else has. Proprietary data is the ultimate content moat — competitors can copy your writing style but they can't copy YOUR data. Automates the design and execution framework for data collection that feeds unique content angles.

Stage

S7: Automation & Scale — Generating data at scale requires automation. This skill designs the collection system, not just one data point. Creates repeatable data assets that compound over time.

When to Use

  • User wants to create content that can't be replicated by competitors
  • User asks about "original research", "surveys", "benchmarks", "proprietary data"
  • User says "data moat", "unique data", "first-party data", "original statistics"
  • After content-moat-calculator identifies the need for differentiated content
  • User wants to build authority through data-driven content
  • User wants to create linkable assets that earn backlinks naturally

Input Schema

yaml
niche: string                 # REQUIRED — topic area for data collection
                              # e.g., "AI video tools", "affiliate marketing"

data_type: string             # OPTIONAL — "survey" | "benchmark" | "aggregation" | "case_study"
                              # Default: recommend based on niche and resources

audience_access: string       # OPTIONAL — how you can reach respondents
                              # e.g., "email list of 500", "Reddit community", "Twitter followers"
                              # Default: suggest options

budget: string                # OPTIONAL — "zero" | "low" ($0-100) | "medium" ($100-500) | "high" ($500+)
                              # Default: "zero"

goal: string                  # OPTIONAL — "content_moat" | "backlink_magnet" | "authority" | "lead_gen"
                              # Default: "content_moat"

Chaining from S3 content-moat-calculator: Use competitive_advantages to identify data moat opportunities.

Workflow

Step 1: Identify Data Opportunity

Analyze the niche for data gaps:

  1. web_search: "[niche] statistics 2025" OR "[niche] survey" OR "[niche] benchmark" — what data already exists?
  2. Identify gaps: what questions does the industry ask that nobody has answered with data?
  3. web_search: "[niche] reddit" "I wish I knew" OR "does anyone know" — find unmet data needs
Step 2: Design Data Collection

Based on data_type (or recommend the best fit):

Survey Design:

  • 8-12 questions (shorter = higher completion)
  • Mix: 70% multiple choice, 20% scale (1-5), 10% open-ended
  • One "surprising" question that will generate headline-worthy data
  • Target sample size: 100+ for credibility
  • Distribution plan: where and how to reach respondents

Benchmark Study:

  • Define metrics to measure (3-5)
  • Data sources: public data, API calls, manual collection
  • Collection methodology: how often, what tools
  • Comparison framework: how to present findings

Data Aggregation:

  • Sources to aggregate from (public databases, APIs, web scraping targets)
  • Aggregation logic: how to combine and normalize
  • Update frequency: one-time or recurring
  • Visualization plan

Case Study Collection:

  • Template for collecting stories (5-7 structured questions)
  • Outreach template for requesting case studies
  • Anonymization rules
  • Minimum viable sample: 10+ cases
Step 3: Create Collection Assets

Produce ready-to-use assets:

  1. Survey questions (if survey) — complete question list with answer options
  2. Collection template — spreadsheet structure or form layout
  3. Outreach template — email/message to recruit respondents
  4. Data analysis plan — how to turn raw data into insights
  5. Content plan — how to present findings (blog post, infographic, report)
Step 4: Design Automation

Create a repeatable system:

  • Schedule: when to collect data (monthly, quarterly, annually)
  • Tools: recommended platforms (Google Forms, Typeform, Airtable)
  • Automation: how to automate collection and reporting
  • Update process: how to refresh and republish with new data
Step 5: Self-Validation
  • Data gap is real (verified by search — nobody else has this data)
  • Sample size is realistic given audience access
  • Questions are unbiased and well-structured
  • Collection method is feasible with stated budget
  • Output content plan is specific (not just "write a blog post")
  • Data is ethically collected (no scraping private data, survey has consent)

Output Schema

yaml
output_schema_version: "1.0.0"
proprietary_data:
  niche: string
  data_type: string
  data_gap: string              # What data doesn't exist yet
  headline_potential: string    # The "surprising finding" angle

  collection:
    method: string
    sample_target: number
    tools: string[]
    timeline: string
    budget_needed: string

  assets:
    survey_questions: object[]  # If survey type
    collection_template: string # Template description
    outreach_template: string   # Recruitment message
    analysis_plan: string

  content_outputs:              # Content to create from the data
    - type: string              # "blog" | "infographic" | "report" | "social"
      title: string
      skill_to_use: string     # Which skill creates this content

  data_assets: string[]        # Moat strengtheners for chaining

chain_metadata:
  skill_slug: "proprietary-data-generator"
  stage: "automation"
  timestamp: string
  suggested_next:
    - "affiliate-blog-builder"
    - "content-pillar-atomizer"
    - "content-moat-calculator"

Output Format

## Proprietary Data Plan: [Niche]

### The Data Gap
**Nobody has answered:** [the question]
**Why it matters:** [why people care]
**Headline potential:** "[Surprising finding template]"

### Collection Design

**Type:** [Survey / Benchmark / Aggregation / Case Study]
**Target sample:** XX responses
**Timeline:** X weeks
**Budget:** $XX
**Tools:** [tools list]

### Survey Questions (or Collection Template)
1. [Question] — [answer type] — [why this question]
2. [Question] — [answer type] — [why this question]
...

### Outreach Template
Subject: [subject line]
[email/message body]

### Content Plan (what to publish from this data)
1. **Blog post:** "[Title]" → build with `affiliate-blog-builder`
2. **Social thread:** Key findings → atomize with `content-pillar-atomizer`
3. **Lead magnet:** Full report PDF → distribute with `squeeze-page-builder`

### Automation Schedule
- **Collection:** [frequency]
- **Analysis:** [when after collection]
- **Publication:** [when after analysis]
- **Update:** [when to re-run with fresh data]

Error Handling

  • No niche provided: "Tell me your niche and I'll find data gaps nobody else is filling."
  • No audience access: Suggest free distribution channels: Reddit, Twitter, niche forums, ProductHunt. "You don't need an email list — Reddit alone can drive 100+ survey responses."
  • Zero budget: Design everything with free tools (Google Forms, Google Sheets, manual aggregation). "The best proprietary data costs $0 — just your time and curiosity."
  • Niche already well-researched: Dig deeper. "The broad stats exist, but nobody has [specific angle]. Let's own that."
Show full SKILL.md (435 more words)Show less

Examples

Example 1: "I want original data about AI video tools" → Design survey: "AI Video Tools Usage Survey 2025" — 10 questions about which tools, satisfaction, spend, use cases. Distribute on Reddit r/aivideo, Twitter, LinkedIn. Target 150 responses. Content plan: "State of AI Video 2025" blog post + infographic.

Example 2: "Create a benchmark for affiliate marketing earnings" → Aggregate public data from case studies, combine with original survey. Monthly recurring data collection. "Affiliate Marketing Earnings Benchmark Q1 2025."

Example 3: "Data moat for my content strategy" (after content-moat-calculator) → Identify that competitors have generic content but NO original data. Design case study collection: "How 50 Affiliate Marketers Made Their First $1,000." Instant authority.

Revenue & Action Plan

Expected Outcomes
  • Revenue potential: Original data content earns 5-10x more backlinks than generic content. Backlinks → higher domain authority → higher rankings for ALL your affiliate pages. One original data post can increase total site traffic by 20-50% over 6 months
  • Benchmark: Data-driven blog posts get 2x more shares and 3x more backlinks than opinion posts. "State of [Industry]" posts are the most linked-to content format in B2B niches
  • Key metric to track: Backlinks earned by the data content (check via Ahrefs, Semrush, or Google Search Console). Secondary: organic traffic increase to ALL affiliate pages (rising tide lifts all boats)
Do This Right Now (15 min)
  1. Launch the survey or start data collection TODAY — don't wait for the "perfect" survey. 80% good is enough to start
  2. Post the survey link in 3 places immediately: your email list, one relevant subreddit, and one social platform
  3. Set a 2-week deadline for data collection — urgency drives responses
  4. Pre-write the blog post outline using the Content Plan section — so you're ready to publish the moment data comes in
Track Your Results

After data collection: publish the findings as a blog post with affiliate-blog-builder. After 30 days: how many backlinks did the data post earn? After 90 days: did organic traffic to your money pages increase? If yes, plan your next data collection round — proprietary data compounds.

Next step — copy-paste this prompt: "Write a blog post presenting my original research findings about [topic]" → runs affiliate-blog-builder

Flywheel Connections

Feeds Into
  • affiliate-blog-builder (S3) — unique data angles for articles nobody else can write
  • content-pillar-atomizer (S2) — data findings to atomize across platforms
  • content-moat-calculator (S3) — proprietary data IS a moat strengthener
Fed By
  • content-moat-calculator (S3) — identifies need for differentiated content
  • performance-report (S6) — performance data to aggregate
Feedback Loop
  • Track backlinks and citations of your data → identify which data points get referenced most → double down on those angles in next collection

References

  • shared/references/case-studies.md — Real data-driven success examples
  • shared/references/flywheel-connections.md — Master connection map

© Affitor, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/automation/proprietary-data-generator of Affitor/affiliate-skills.

Open the folder on GitHubat commit e43bfae

Compare with similar skills

Proprietary Data Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Proprietary Data Generator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Proprietary Data Generator this skillAffitor/affiliate-skills699—~2.8kAutomated safety check: PassMIT
Meridian MMM Model Buildinggoogle/meridian1.6k—~2.5kAutomated safety check: PassApache-2.0
A/B Test Analysisphuryn/pm-skills27k—~893Automated safety check: PassMIT
Statistical Analystalirezarezvani/claude-skills28k1 repos~2.5kAutomated safety check: PassMIT
Experimentation Analyticsrampstackco/claude-skills9401 repos~8.9kAutomated safety check: PassMIT
Meta Results Forest Plot Analyzeraipoch/medical-research-skills2k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Official

    Takes a user through building a Meridian marketing mix model, from loading CSV data and mapping columns to running EDA, fitting and saving the model.

    1.6k GitHub stars~2.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • A/B Test Analysis

    phuryn/pm-skills

    Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop.

    27k GitHub stars~893 tokensUpdated 23 days ago
    Data & AnalyticsAuto-check passed
  • Statistical Analyst

    alirezarezvani/claude-skills

    Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Experimentation Analytics

    rampstackco/claude-skills

    How to read experiment results without fooling yourself. An agent skill from rampstackco/claude-skills.

    940 GitHub starsUsed in 1 repo~8.9k tokens
    Data & AnalyticsAuto-check passed
  • Meta Results Forest Plot Analyzer

    aipoch/medical-research-skills

    Analyzes forest plots for meta-analysis, generating detailed descriptions and formatting figure legends in Chinese or English.

    2k GitHub stars~1.9k tokensUpdated 21 days ago
    Data & AnalyticsAuto-check passed
  • Sandbox Bench

    vercel/next.js

    Official

    Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

    143k GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from Affitor/affiliate-skills

All 50 skills in this repo
  • Performance Report

    Affitor/affiliate-skills

    Generate affiliate performance reports with KPIs and recommendations.

    699 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Affiliate Check

    Affitor/affiliate-skills

    Live affiliate program data from openaffiliate.dev. An agent skill from Affitor/affiliate-skills.

    699 GitHub stars~808 tokensUpdated 23 days ago
    Auto-check: notes
  • Affiliate Program Search

    Affitor/affiliate-skills

    Research and evaluate affiliate programs to find the best ones to promote.

    699 GitHub stars~2.4k tokensUpdated 23 days ago
    Auto-check passed
  • Bio Link Deployer

    Affitor/affiliate-skills

    Create a Linktree-style bio link hub page as a single self-contained HTML file.

    699 GitHub stars~3k tokensUpdated 23 days ago
    Auto-check passed
  • Conversion Tracker

    Affitor/affiliate-skills

    Set up affiliate conversion tracking with UTM parameters and link tagging.

    699 GitHub stars~2.1k tokensUpdated 23 days ago
    Auto-check passed
  • Product Showcase Page

    Affitor/affiliate-skills

    Build a single-product deep-dive showcase page as a self-contained HTML file.

    699 GitHub stars~3.6k tokensUpdated 23 days ago
    Auto-check passed

Questions about Proprietary Data Generator

What does Proprietary Data Generator do?

Create original surveys, benchmarks, and aggregated data nobody else has. Proprietary Data Generator is an agent skill from Affitor/affiliate-skills. Create original surveys, benchmarks, and aggregated data nobody else has.

When should I use Proprietary Data Generator?

Proprietary Data Generator fits situations like: : create original data; proprietary data; benchmark study; original research.

How do I install Proprietary Data Generator in Claude Code?

Run `npx skills add Affitor/affiliate-skills --skill proprietary-data-generator -a claude-code`. Or copy the skill folder (skills/automation/proprietary-data-generator in Affitor/affiliate-skills) into .claude/skills/proprietary-data-generator in your project. Claude Code loads it when a task matches its description.

How do I install Proprietary Data Generator in Codex?

Run `npx skills add Affitor/affiliate-skills --skill proprietary-data-generator -a codex`. Or copy the skill folder (skills/automation/proprietary-data-generator in Affitor/affiliate-skills) into .agents/skills/proprietary-data-generator in your project. Codex loads it when a task matches its description.

Can I use Proprietary Data Generator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Affitor/affiliate-skills --skill proprietary-data-generator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/proprietary-data-generator, .gemini/skills/proprietary-data-generator, .github/skills/proprietary-data-generator and .opencode/skills/proprietary-data-generator in your project.

What does Proprietary Data Generator need to run?

SKILL.md names no scripts, command-line tools or credentials: Proprietary Data Generator is instructions for the agent only. Compatibility (from SKILL.md): Claude Code, ChatGPT, Gemini CLI, Cursor, Windsurf, OpenClaw, any AI agent.

Does Proprietary Data Generator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Proprietary Data Generator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Proprietary Data Generator use?

Proprietary Data Generator is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Proprietary Data Generator use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Proprietary Data Generator?

Skills that share tags, products or a category with Proprietary Data Generator: Meridian MMM Model Building (google/meridian, 1.6k stars), A/B Test Analysis (phuryn/pm-skills, 27k stars), Statistical Analyst (alirezarezvani/claude-skills, 28k stars) and Experimentation Analytics (rampstackco/claude-skills, 940 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Proprietary Data Generator?

Affitor (a GitHub organization) maintains it in Affitor/affiliate-skills, which has 699 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on September 15, 2026.

Source: Affitor/affiliate-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.