Agent skill

Gallery Scraper

by jdrhyne in jdrhyne/agent-skills

Bulk download images from login-protected gallery websites using an attached browser session.

MITAuto-check passedData & Analytics

Install Gallery Scraper

skills CLI
$ npx skills add jdrhyne/agent-skills --skill gallery-scraper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jdrhyne/agent-skills gallery-scraper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jdrhyne/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clawdbot/gallery-scraper .claude/skills/gallery-scraper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gallery-scraper
GitHub stars
240
Token cost
~1.7k tokens
SKILL.md length
303 words
Files
3 (incl. scripts)
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

Bulk download images from login-protected gallery websites using an attached browser session.

  • Works in 6 steps: Attach Browser Tab → Discover Image URL Pattern → Extract Full-Size URLs → …
  • Asked to scrape
  • SKILL.md covers Safety Boundaries, Prerequisites, Workflow and Handling Lock Buttons, plus 2 more sections
  • Runs Shell and JavaScript scripts from its folder; calls curl

What it does

Gallery Scraper is an agent skill from jdrhyne/agent-skills. Bulk download images from login-protected gallery websites using an attached browser session. Use when asked to scrape, download, or save images from authenticated gallery pages, extract full-size images from thumbnails, or batch download from multi-page galleries.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/download_gallery.sh` and `scripts/extract_patterns.js`).

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: A collection of AI agent skills for Clawdbot, Claude Code, Codex. The licence is MIT.

When your agent uses it

  • Asked to scrape
  • Save images from authenticated gallery pages
  • Extract full-size images from thumbnails
  • Batch download from multi-page galleries

Example prompts

  • “/gallery-scraper”

Requirements

  • Python 3
  • Node.js
  • A Bash shell

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Attach Browser Tab
  2. Discover Image URL Pattern
  3. Extract Full-Size URLs
  4. Handle Pagination
  5. Check CDN Access
  6. Bulk Download

What it can do on your machine

Read from SKILL.md and the folder at commit 439cd3a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gallery Scraper loads about 1.7k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 303 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jdrhyne/agent-skills at commit 439cd3a, republished under its MIT licence (© jdrhyne). 303 words, ~1,661 tokens.

Download SKILL.mdSave it as .claude/skills/gallery-scraper/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
gallery-scraper
description
Bulk download images from login-protected gallery websites using an attached browser session. Use when asked to scrape, download, or save images from authenticated gallery pages, extract full-size images from thumbnails, or batch download from multi-page galleries.

Bulk download images from authenticated gallery websites via browser relay.

Safety Boundaries

  • Do not access gallery sites or user accounts that the user has not explicitly attached and authorized.
  • Do not download beyond the selected gallery, profile, or page range without confirmation.
  • Do not store cookies, tokens, or hidden form values in local output files.
  • Do not keep retrying blocked downloads indefinitely; surface rate limits or auth failures instead.

Prerequisites

  • User must have Chrome with OpenClaw Browser Relay extension
  • User must be logged into the target site
  • User must attach the browser tab (click relay toolbar button, badge ON)

Workflow

1. Attach Browser Tab

Ask user to:

  1. Log into the gallery site in Chrome
  2. Navigate to the target gallery/profile page
  3. Click the OpenClaw Browser Relay toolbar button (badge shows ON)
2. Discover Image URL Pattern

Most gallery sites store full-size URLs in data attributes. Common patterns:

javascript
// Extract via browser evaluate
() => {
  // Try common patterns
  const patterns = [
    'img[data-max]',           // data-max attribute
    'img[data-src]',           // lazy-load pattern
    'img[data-full]',          // full-size pattern
    'a[data-lightbox] img',    // lightbox galleries
    '.gallery-item img'        // generic gallery
  ];
  
  for (const sel of patterns) {
    const imgs = document.querySelectorAll(sel);
    if (imgs.length > 0) {
      return {
        selector: sel,
        count: imgs.length,
        sample: imgs[0].outerHTML.substring(0, 200)
      };
    }
  }
  return null;
}
3. Extract Full-Size URLs

Once pattern identified, extract all URLs:

javascript
// For data-max pattern (common)
() => Array.from(document.querySelectorAll('img[data-max]'))
  .map(img => img.dataset.max)

// For thumbnail→full conversion (replace path segment)
() => Array.from(document.querySelectorAll('.gallery img'))
  .map(img => img.src.replace('/thumb/', '/full/'))
4. Handle Pagination

Check for multiple pages:

javascript
() => {
  const pagination = document.querySelectorAll('.pagination a, [class*="page"] a');
  return Array.from(pagination).map(a => ({text: a.textContent, href: a.href}));
}

Navigate to each page and collect URLs.

4b. Batch scrape multiple galleries (iframe trick)

When you need multiple galleries quickly and can’t automate CDP, you can load each gallery in a hidden iframe and extract data-max URLs:

javascript
async () => {
  const urls = [
    'https://site.example/galleries/view/123',
    'https://site.example/galleries/view/456'
  ];
  const results = [];
  for (const url of urls) {
    const iframe = document.createElement('iframe');
    iframe.style.position = 'fixed';
    iframe.style.left = '-9999px';
    iframe.style.width = '800px';
    iframe.style.height = '600px';
    iframe.src = url;
    document.body.appendChild(iframe);
    await new Promise((resolve, reject) => {
      const t = setTimeout(() => reject(new Error('timeout load')), 20000);
      iframe.onload = () => { clearTimeout(t); resolve(); };
    });
    const doc = iframe.contentDocument;
    const start = Date.now();
    let imgs = [];
    while (Date.now() - start < 20000) {
      imgs = Array.from(doc.querySelectorAll('img[data-max]')).map(i => i.dataset.max);
      if (imgs.length) break;
      await new Promise(r => setTimeout(r, 500));
    }
    results.push({ id: url.split('/').pop(), urls: imgs });
    iframe.remove();
  }
  return results;
}
5. Check CDN Access

Test if CDN requires authentication or just Referer:

bash
# Test direct access
curl -I "CDN_URL" 2>/dev/null | head -3

# Test with Referer
curl -I -H "Referer: https://SITE_DOMAIN/" "CDN_URL" 2>/dev/null | head -3
6. Bulk Download

Collect the URLs into a text file, then parallel download:

bash
# Create output directory
mkdir -p ~/Downloads/gallery_name

# Download with Referer header (parallel)
cd ~/Downloads/gallery_name
while IFS= read -r url; do
  filename=$(basename "$url")
  curl -s -H "Referer: https://SITE_DOMAIN/" -o "$filename" "$url" &
  [ $(jobs -r | wc -l) -ge 8 ] && wait -n
done < urls.txt
wait

Python ThreadPool fallback (avoids shell quoting + wait -n issues):

python
import os
import requests
from concurrent.futures import ThreadPoolExecutor

outdir = os.path.expanduser('~/Downloads/gallery_name')
os.makedirs(outdir, exist_ok=True)
headers = {'Referer': 'https://SITE_DOMAIN/', 'User-Agent': 'Mozilla/5.0'}

with open('urls.txt') as f:
    urls = [line.strip() for line in f if line.strip()]

def download(url):
    filename = os.path.join(outdir, os.path.basename(url))
    if os.path.exists(filename) and os.path.getsize(filename) > 0:
        return
    r = requests.get(url, headers=headers, timeout=60)
    r.raise_for_status()
    with open(filename, 'wb') as f:
        f.write(r.content)

with ThreadPoolExecutor(max_workers=8) as ex:
    for url in urls:
        ex.submit(download, url)

Handling Lock Buttons

Some galleries have "lock" buttons to reveal hidden content. Look for:

javascript
// Find lock/unlock buttons
() => {
  const locks = document.querySelectorAll(
    '[class*="lock"], [class*="unlock"], ' +
    'button[title*="lock"], .premium-unlock'
  );
  return Array.from(locks).map(el => ({
    tag: el.tagName,
    class: el.className,
    text: el.innerText?.substring(0, 30)
  }));
}

Click each lock button before extracting URLs.

Output Organization

Optionally organize by gallery:

bash
# Derive a gallery-specific folder name from the selected URL
mkdir -p "gallery_<id>"

Troubleshooting

  • 403 Forbidden: Add Referer header or extract cookies from browser
  • Rate limited: Reduce parallel downloads, add delays
  • Missing images: Check for JavaScript-loaded content, may need scroll injection
  • Login required for CDN: Extract session cookies via document.cookie

© jdrhyne, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in clawdbot/gallery-scraper of jdrhyne/agent-skills.

  • SKILL.md
  • scripts/download_gallery.sh
  • scripts/extract_patterns.js

Open the folder on GitHubat commit 439cd3a

Compare with similar skills

Gallery Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gallery Scraper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gallery Scraper this skilljdrhyne/agent-skills240—~1.7kAutomated safety check: PassMIT
Tmuxtrpc-group/trpc-agent-go1.9k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch7001 repos~3.9kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.9k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    700 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 10 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from jdrhyne/agent-skills

All 20 skills in this repo
  • Gong

    jdrhyne/agent-skills

    Gong API for searching calls, transcripts, and conversation intelligence.

    240 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Elegant Reports

    jdrhyne/agent-skills

    Generate beautifully designed PDF reports with a Nordic/Scandinavian aesthetic.

    240 GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Ga4

    jdrhyne/agent-skills

    Read Google Analytics 4 reporting data through the GA4 Data API.

    240 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Gsc

    jdrhyne/agent-skills

    Read Google Search Console properties, Search Analytics, URL Inspection, and sitemaps.

    240 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Nudocs

    jdrhyne/agent-skills

    Upload, edit, and export documents via Nudocs.ai. An agent skill from jdrhyne/agent-skills.

    240 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Todo Tracker

    jdrhyne/agent-skills

    Persistent TODO scratch pad for tracking tasks across sessions.

    240 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Gallery Scraper

What does Gallery Scraper do?

Bulk download images from login-protected gallery websites using an attached browser session. Gallery Scraper is an agent skill from jdrhyne/agent-skills. Bulk download images from login-protected gallery websites using an attached browser session.

When should I use Gallery Scraper?

Gallery Scraper fits situations like: asked to scrape; save images from authenticated gallery pages; extract full-size images from thumbnails; batch download from multi-page galleries.

How do I install Gallery Scraper in Claude Code?

Run `npx skills add jdrhyne/agent-skills --skill gallery-scraper -a claude-code`. Or copy the skill folder (clawdbot/gallery-scraper in jdrhyne/agent-skills) into .claude/skills/gallery-scraper in your project. Claude Code loads it when a task matches its description.

How do I install Gallery Scraper in Codex?

Run `npx skills add jdrhyne/agent-skills --skill gallery-scraper -a codex`. Or copy the skill folder (clawdbot/gallery-scraper in jdrhyne/agent-skills) into .agents/skills/gallery-scraper in your project. Codex loads it when a task matches its description.

Can I use Gallery Scraper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jdrhyne/agent-skills --skill gallery-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gallery-scraper, .gemini/skills/gallery-scraper, .github/skills/gallery-scraper and .opencode/skills/gallery-scraper in your project.

What does Gallery Scraper need to run?

Going by SKILL.md and its folder, Gallery Scraper needs a shell and JavaScript for the scripts in its folder and the command-line tools its instructions call (curl). Our summary lists: Python 3; Node.js; A Bash shell.

Does Gallery Scraper access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Gallery Scraper safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gallery Scraper use?

Gallery Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gallery Scraper use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gallery Scraper?

Skills that share tags, products or a category with Gallery Scraper: Tmux (trpc-group/trpc-agent-go, 1.9k stars), Ketch (1broseidon/ketch, 700 stars), Crawl4AI Web Scraping (smallnest/goclaw, 599 stars) and Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gallery Scraper?

jdrhyne (a GitHub user) maintains it in jdrhyne/agent-skills, which has 240 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on August 30, 2026.

Source: jdrhyne/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.