Agent skill

Substack Notes Scraper

by mohitagw15856 in mohitagw15856/pm-claude-skills

Scrapes a Substack Notes page and exports engagement data to a formatted .xlsx file.

MITAuto-check passedDocuments & Office

Install Substack Notes Scraper

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill substack-notes-scraper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills substack-notes-scraper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/substack-notes-scraper .claude/skills/substack-notes-scraper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
substack-notes-scraper
GitHub stars
1.4k
Token cost
~1.9k tokens
SKILL.md length
878 words
Files
1
Skills in repo
1,322
Repo updated
First seen
Licence
MIT

At a glance

Scrapes a Substack Notes page and exports engagement data to a formatted .xlsx file.

  • Works in 9 steps: Validate inputs → Fetch the Notes page → Paginate through all notes in the date… → …
  • Asked to download
  • SKILL.md covers Required Inputs, Output Structure, Instructions for Claude and Quality Checks, plus 2 more sections
  • Reaches substack.com

What it does

Substack Notes Scraper is an agent skill from mohitagw15856/pm-claude-skills. Scrapes a Substack Notes page and exports engagement data to a formatted .xlsx file. Use when asked to download, analyse, or export Substack Notes performance data including likes, comments, and restacks. Produces a formatted spreadsheet with conditional formatting, summary stats, and per-note engagement metrics.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Newsletters, Web scraping and Excel spreadsheets. It works with Substack and Microsoft Excel. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked to download
  • Export Substack Notes performance data including likes

Example prompts

  • “Use the substack-notes-scraper skill to scrape a Substack Notes page and exports engagement data to a formatted .xlsx file”
  • “/substack-notes-scraper”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Validate inputs
  2. Fetch the Notes page
  3. Paginate through all notes in the date window
  4. Parse each note
  5. Filter
  6. Calculate Total Engagement
  7. Identify top 20% by Likes
  8. Build the .xlsx file
  9. Report back

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • substack.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Substack Notes Scraper loads about 1.9k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 878 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 878 words, ~1,852 tokens.

Download SKILL.mdSave it as .claude/skills/substack-notes-scraper/SKILL.md (or your agent's skills folder).
name
substack-notes-scraper
description
Scrapes a Substack Notes page and exports engagement data to a formatted .xlsx file. Use when asked to download, analyse, or export Substack Notes performance data including likes, comments, and restacks. Produces a formatted spreadsheet with conditional formatting, summary stats, and per-note engagement metrics.

Substack Notes Scraper

Substack has no public API for Notes analytics. You can't see likes, comments, and restacks in one place without scrolling through your feed manually. This skill scrapes the rendered Notes page, filters to only your original content, and exports everything to a spreadsheet you can actually analyze.

Credit: Originally created by a Substack newsletter author — adapted and extended for this library.


Required Inputs

InputFormatExample
Notes URLFull URL to the Notes tabhttps://substack.com/@handle/notes
Author handle or nameExact handle or display name@handle or Jane Smith
Date rangePlain English or explicit rangelast 30 days or Jan 2026 – Mar 2026

Claude will ask for these if not provided upfront.


Output Structure

File
substack-notes-[handle]-[YYYY-MM-DD].xlsx
Sheet: "Notes Data"
ColumnDescription
DatePublication date (YYYY-MM-DD)
Text PreviewFirst 200 characters of the note
Full TextComplete note text
LikesLike count at time of scrape
CommentsComment count
RestacksRestack count
Total EngagementLikes + Comments + Restacks
LinkDirect URL to the note
Note Typeoriginal or restack

Formatting applied:

  • Row 1: frozen header row
  • Auto-filter enabled on all columns
  • Top 20% by Likes column: highlighted yellow (#FFF2CC)
  • Column widths: auto-fit to content, min 12, max 60
Sheet: "Summary"
Scrape Date:         [YYYY-MM-DD HH:MM UTC]
Author:              [handle]
Date Range:          [start] – [end]
Total Notes:         [n]
Original Notes:      [n]
Restacks Filtered:   [n]

Avg Likes:           [n.n]
Avg Comments:        [n.n]
Avg Restacks:        [n.n]
Avg Total Eng:       [n.n]

Best Note (Likes):   [date] — [first 80 chars] — [n] likes
Best Note (Eng):     [date] — [first 80 chars] — [n] total engagement

Instructions for Claude

Step 1: Validate inputs

Confirm the three required inputs are present. If any are missing, ask before proceeding. Parse the date range into a concrete start date and end date (convert relative ranges like "last 30 days" to explicit dates using today's date).

Step 2: Fetch the Notes page

Use WebFetch to load the Notes URL. Substack Notes pages are JavaScript-rendered — request the full rendered HTML. If WebFetch returns a skeleton page without note content, note this in your response and ask the user to paste the page HTML manually or confirm browser access is available.

Step 3: Paginate through all notes in the date window

Substack Notes load incrementally. Repeat fetching or scrolling until either:

  • A note's date falls outside the target date range (stop loading more), or
  • No new content loads on the next request.

Rate-limit: wait 2 seconds between each paginated request. Do not hammer the endpoint.

Step 4: Parse each note

For every note element found on the page, extract:

  • Date: the timestamp on the note (convert to YYYY-MM-DD)
  • Author: the display name or handle shown on the note
  • Full text: complete body text, stripping HTML tags
  • Text preview: first 200 characters of full text
  • Likes count: the number shown on the like/heart counter
  • Comments count: the number shown on the comment counter
  • Restacks count: the number shown on the restack counter
  • Link: the direct permalink to the note
  • Note type: original if the author matches the specified author; restack if it belongs to someone else
Step 5: Filter

Keep ALL rows in the data (restacks included as rows with Note Type = restack). The Summary sheet stats should count only original notes. Mark restacks clearly so the user can filter them out themselves in Excel if preferred.

Apply date filter: exclude any note outside the specified date range.

Step 6: Calculate Total Engagement

For each row: Total Engagement = Likes + Comments + Restacks

Show full SKILL.md (360 more words)Show less
Step 7: Identify top 20% by Likes

Sort original notes by Likes descending. Mark the top 20% (round up) for conditional formatting. These rows will be highlighted yellow in the output file.

Step 8: Build the .xlsx file

Use Python with openpyxl to generate the file. Structure:

python
# Required libraries
import openpyxl
from openpyxl.styles import PatternFill, Font, Alignment
from openpyxl.utils import get_column_letter
from datetime import datetime

# Sheet 1: Notes Data
# - Write header row, bold, freeze row 1
# - Write all data rows
# - Apply auto-filter: ws.auto_filter.ref = ws.dimensions
# - Apply yellow fill to top-20% rows by likes
# - Auto-size columns (iterate cells to find max length)

# Sheet 2: Summary
# - Write summary stats as key-value pairs, no table format

Name the file substack-notes-[handle]-[YYYY-MM-DD].xlsx using today's date.

Step 9: Report back

After generating the file, report:

  • File path
  • Total notes found, original vs. restacks
  • Date range actually covered
  • Top 3 notes by total engagement (date + preview + stats)
  • Any notes or warnings (e.g., page didn't fully load, some dates were ambiguous)

Quality Checks

  • All three required inputs were confirmed before starting
  • Rate limiting honored: 2-second delay between paginated requests
  • Author filter applied correctly — restacks are included as rows but flagged, not silently dropped
  • Date range filter applied — no notes outside the window appear in the data
  • Total Engagement column is Likes + Comments + Restacks (not hardcoded)
  • Top 20% highlight is based on the actual data distribution, not a fixed threshold
  • Header row is frozen and auto-filter is active
  • Summary sheet stats reference only original notes, not restacks
  • File is named with the author handle and today's date
  • If the page failed to load properly, the user was told — not silently given an empty file

Anti-Patterns

  • Do not proceed without a valid Substack handle or profile URL — scraping without a specific target cannot be completed
  • Do not ignore rate-limit responses from Substack — implement backoff and reduce request frequency before retrying
  • Do not export data without conditional formatting and summary stats — raw data without visualisation is not the expected output
  • Do not attempt to access private or subscriber-only notes — this skill is for public Notes content only
  • Do not produce output without a clear date range filter — undated exports make trend analysis impossible

Example Trigger Phrases

  • "Scrape my Substack Notes and export to Excel — my handle is @handle, last 60 days"
  • "Use the substack-notes-scraper skill on https://substack.com/@handle/notes for Q1 2026"
  • "Pull my notes engagement data into a spreadsheet"
  • "Export my Substack Notes stats with likes and restacks — author: Jane Smith, Jan–Mar 2026"
  • "Run the Substack scraper on my notes page and show me which posts performed best"

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/substack-notes-scraper of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Substack Notes Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Substack Notes Scraper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Substack Notes Scraper this skillmohitagw15856/pm-claude-skills1.4k—~1.9kAutomated safety check: PassMIT
Venice Augmentveniceai/skills143—~2.5kAutomated safety check: PassMIT
Firecrawl Parsefirecrawl/skills116—~441Automated safety check: PassISC
Anydocmagnus919/agent-skills113—~3.8kAutomated safety check: NotesMIT
Crawl Mentors To XLSXJunieXD/AutoEmailSender149—~429Automated safety check: PassGPL-3.0
Yao Doubao Crawleryaojingang/yao-geo-skills868—~475Automated safety check: PassMIT

Similar skills

  • Venice Augment

    veniceai/skills

    Venice augmentation endpoints for agent pipelines. An agent skill from veniceai/skills.

    143 GitHub stars~2.5k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Firecrawl Parse

    firecrawl/skills

    Convert a local file (PDF, DOCX, XLSX, HTML, …) to markdown, or answer questions about its content.

    116 GitHub stars~441 tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Anydoc

    magnus919/agent-skills

    Convert Word (.doc/.docx/.docm), PowerPoint (.ppt/.pps/.pot/.pptx/.pptm/.ppsx/.ppsm), Excel (.xls/.xlsx/.xlsm/.xlsb), OpenDocument (.odt/.ods/.odp), RTF, EPUB, CSV, and PDF documents to clean…

    113 GitHub stars~3.8k tokensUpdated yesterday
    Documents & OfficeAuto-check: notes
  • Crawl Mentors To XLSX

    JunieXD/AutoEmailSender

    从学校、学院、系所或实验室官网抓取公开导师/教师信息,核对个人主页与证据来源,并生成经过自动校验、可直接导入 Auto Email Sender 的 XLSX。Use when a user provides faculty, professor, mentor, supervisor, university, department, or lab directory URLs and…

    149 GitHub stars~429 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Yao Doubao Crawler

    yaojingang/yao-geo-skills

    A skill your agent uses when a user needs repeated Doubao AI-search collection from web or Android Appium into compatible JSON plus Markdown/Excel/HTML GEO reports.

    868 GitHub stars~475 tokensUpdated 7 days ago
    Data & AnalyticsAuto-check passed
  • Data Cleaning

    ericrisco/rsc-harness

    A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…

    167 GitHub stars~3.6k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from mohitagw15856/pm-claude-skills

All 1,322 skills in this repo
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Offer Comparison

    mohitagw15856/pm-claude-skills

    Compare two or more job offers as total-comp curves over four years — vesting cliffs, bonuses, 401(k) match, and the crossover year computed, not vibed.

    1.4k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Refinance Breakeven

    mohitagw15856/pm-claude-skills

    Compute the month a refinance actually starts saving money — payment delta, breakeven month, and total interest on both paths including the term-reset trap.

    1.4k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Rent Vs Buy

    mohitagw15856/pm-claude-skills

    Model rent-vs-buy honestly — year-by-year net position for both paths including the assumption everyone drops (the renter invests the difference), with a breakeven horizon instead of a verdict.

    1.4k GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Questions about Substack Notes Scraper

What does Substack Notes Scraper do?

Scrapes a Substack Notes page and exports engagement data to a formatted .xlsx file. Substack Notes Scraper is an agent skill from mohitagw15856/pm-claude-skills.xlsx file.

When should I use Substack Notes Scraper?

Substack Notes Scraper fits situations like: asked to download; export Substack Notes performance data including likes.

How do I install Substack Notes Scraper in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill substack-notes-scraper -a claude-code`. Or copy the skill folder (skills/substack-notes-scraper in mohitagw15856/pm-claude-skills) into .claude/skills/substack-notes-scraper in your project. Claude Code loads it when a task matches its description.

How do I install Substack Notes Scraper in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill substack-notes-scraper -a codex`. Or copy the skill folder (skills/substack-notes-scraper in mohitagw15856/pm-claude-skills) into .agents/skills/substack-notes-scraper in your project. Codex loads it when a task matches its description.

Can I use Substack Notes Scraper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill substack-notes-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/substack-notes-scraper, .gemini/skills/substack-notes-scraper, .github/skills/substack-notes-scraper and .opencode/skills/substack-notes-scraper in your project.

What does Substack Notes Scraper need to run?

SKILL.md names no scripts, command-line tools or credentials: Substack Notes Scraper is instructions for the agent only. Our summary lists: Python 3.

Does Substack Notes Scraper access the network?

SKILL.md names 1 domain. In commands or code: substack.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Substack Notes Scraper safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Substack Notes Scraper use?

Substack Notes Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Substack Notes Scraper use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Substack Notes Scraper?

Skills that share tags, products or a category with Substack Notes Scraper: Venice Augment (veniceai/skills, 143 stars), Firecrawl Parse (firecrawl/skills, 116 stars), Anydoc (magnus919/agent-skills, 113 stars) and Crawl Mentors To XLSX (JunieXD/AutoEmailSender, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Substack Notes Scraper?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,431 GitHub stars. The repository holds 1,322 skills in this directory. The repository was last updated on October 7, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.