Agent skill

Crawler Pep

by opensanctions in opensanctions/opensanctions

Scaffold a new PEP (Politically Exposed Persons) crawler — members of a parliament, legislature, senate, chamber of deputies, cabinet, judiciary, or an asset-declaration register — from a source URL…

MITAuto-check: notesData & Analytics

Install Crawler Pep

skills CLI
$ npx skills add opensanctions/opensanctions --skill crawler-pep -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install opensanctions/opensanctions crawler-pep --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/opensanctions/opensanctions.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/crawler-pep .claude/skills/crawler-pep && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
crawler-pep
GitHub stars
832
Token cost
~1.9k tokens
SKILL.md length
953 words
Files
4 (incl. scripts)
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Scaffold a new PEP (Politically Exposed Persons) crawler — members of a parliament, legislature, senate, chamber of deputies, cabinet, judiciary, or an asset-declaration register — from a source URL…

  • Works in 3 steps: Understand the source → Write the YAML and crawler → Validate
  • Members-of-parliament crawler
  • SKILL.md covers Step 1: Understand the source, Step 2: Write the YAML and… and Step 3: Validate
  • Runs Python scripts from its folder; calls python

What it does

Crawler Pep is an agent skill from opensanctions/opensanctions. Scaffold a new PEP (Politically Exposed Persons) crawler — members of a parliament, legislature, senate, chamber of deputies, cabinet, judiciary, or an asset-declaration register — from a source URL or GitHub issue. Creates the dataset .yml plus a crawler emitting Person, Position and Occupancy entities via makeposition/categorise/makeoccupancy. Use when asked to add, write or scaffold a PEP or members-of-parliament crawler.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `examples.md`, `scripts/pep_summary.py` and `validation.md`).

It sits in Data & Analytics, covering Web scraping. It works with GitHub. The repository describes itself as: An open database of international sanctions data, persons of interest and politically exposed persons. The licence is MIT.

When your agent uses it

  • Members-of-parliament crawler
  • Tasks that involve Web scraping

Example prompts

  • “/crawler-pep”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Edit, Write, Glob, Grep, Bash, WebFetch, WebSearch, Agent

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Understand the source
  2. Write the YAML and crawler
  3. Validate

What it can do on your machine

Read from SKILL.md and the folder at commit ce59ef9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Edit
    • Write
    • Glob
    • Grep
    • Bash
    • WebFetch
    • WebSearch
    • Agent

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Crawler Pep loads about 1.9k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 953 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Edit, Write, Glob, Grep, Bash, WebFetch, WebSearch, Agent

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from opensanctions/opensanctions at commit ce59ef9, republished under its MIT licence (© opensanctions). 953 words, ~1,879 tokens.

Download SKILL.mdSave it as .claude/skills/crawler-pep/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
crawler-pep
description
Scaffold a new PEP (Politically Exposed Persons) crawler — members of a parliament, legislature, senate, chamber of deputies, cabinet, judiciary, or an asset-declaration register — from a source URL or GitHub issue. Creates the dataset .yml plus a crawler emitting Person, Position and Occupancy entities via make_position/categorise/make_occupancy. Use when asked to add, write or scaffold a PEP or members-of-parliament crawler.
allowed-tools
Read, Edit, Write, Glob, Grep, Bash, WebFetch, WebSearch, Agent
argument-hint
[target path | source URL | GitHub issue URL]

New PEP Crawler

Create a new PEP crawler. The user will provide a target path, source data URL, and/or a GitHub issue URL: $ARGUMENTS

If given a GitHub issue URL, fetch it first to extract the data source URL and any context about the dataset.

Read upfront. These are the rules; this skill is the procedure for applying them, and repeats only the few rules the docs don't cover yet:

  1. zavod/docs/peps.md — the PEP model, properties, position naming, categorisation, occupancy dates and status, historical terms.
  2. .claude/docs/crawler-guide.md — shared crawler patterns, YAML template, lookups.
  3. .claude/skills/crawler-pep/examples.md → "Reference crawler" — a complete, reviewed crawler. Yours should look like it: the shortest code that still handles every source field explicitly. Each helper, constant, guard or extra request you add beyond it needs a reason a reviewer can see — first drafts fail review far more often from added machinery than from missing it.

Open the rest of examples.md when your source differs from the reference (mixed datasets, several positions, multi-term, subnational). Ground the crawler in these files, not in other crawlers: the codebase is old and many have drifted from current practice.

Step 1: Understand the source

Do this before writing any code. A wrong endpoint or a misread date produces a crawler that looks finished and is worthless.

  • Find the underlying JSON/XML endpoint before parsing HTML — page source, network calls, JS bundles. Parliament sites very often have an API behind the rendered page.
  • Prove the source is blocked before reaching for Zyte. In order: a browser http.user_agent; a language cookie or Accept-Language; a format suffix (.json) or Accept: header. Zyte costs you ci_test: false.
  • Establish the pagination contract from the response — the per-page cap, the total or next link. An API that silently caps page size truncates without error.
  • Enumerate every field the source returns and decide each one: emit, or ignore.
  • Find the terms the source exposes — a term switcher, an ElectionId parameter, a /legislaturas endpoint — and whether dates are per person or per term.
  • Wikidata QIDs: check on Wikidata that the item is instance of (P31): position and applies to jurisdiction (P1001) matches the country; a plausible label is not enough. Never pass one QID to two positions — it becomes the entity ID, and they'd collapse into one.
  • Citizenship: spawn a subagent (WebSearch/WebFetch) to find the legal document (constitution, electoral law) that requires citizenship for this specific position. Cite its URL in a comment next to person.add("citizenship", ...), or next to its omission if not required.
  • Term-bounded source (fixed mandates, per-term pages)? Note a structural signature (page URL, file name, term id) and fail in crawl() when it changes, so a new term can't go unnoticed.
  • Check the records against reality. Count them by role and by term, and compare with the seats the body actually has. Could a record's dates be stale (e.g. a re-elected member still carrying their previous term)?

Checkpoint: show the user a short recon note and wait before writing code — endpoints and pagination; every field with its decision; the terms exposed; counts by role and term against the seats; anything the source contradicts itself on, with the simplest options. A wrong assumption corrected here costs one message; found in review it costs a rewrite. In an automated run with no user, put the note at the top of your final report and proceed.

Show full SKILL.md (398 more words)Show less

Step 2: Write the YAML and crawler

  • Tag list.pep. For title, description and coverage.frequency apply /legislature-metadata (legislatures), and /dataset-metadata for the rest.
  • Base assertions bands on the entity counts the crawl actually emitted, not on the seat count: a multi-term crawl holds several cohorts and grows every election.
  • Constants (gender maps, headers, date formats, column labels) belong in the YAML, not the crawler — use /crawler-constants-to-yml if you've written one in code.
  • Pass topics= to make_position for positions the crawler names itself (["gov.national", "gov.legislative"], …); omit it for positions read from the source, where the review system decides.
  • Set every person property make_occupancy reads (birthDate, deathDate) before calling it, and emit the person after it — it adds role.pep to the person.
  • For judicial positions, also add role.judge to the person's topics.
  • Honorifics: zavod/docs/best_practices/name_titles.md. LLM-assisted or reviewed name cleaning is acceptable for PEP data (unlike sanctions): zavod/docs/extract/names.md.

Step 3: Validate

A crawler that has not completed a successful zavod crawl is not deliverable. If the source can't be fetched, stop and report the blocker with the evidence from your recon note — don't ship a parser validated against an archived copy.

bash
zavod crawl <path>               # then read the run's issues — they must be clean (see below)
zavod export <path>              # runs the dataset validators and assertions
contrib/lint_dataset.sh <path>   # ruff + mypy + pre-commit exactly as CI runs them

Use the lint script rather than bare ruff/mypy, which lack the repo config. A run's issues and statements are in data/datasets/<dataset>/_artifacts/<run version>/ (issues.json), or directly in data/datasets/<dataset>/ (issues.log) on older zavod versions.

Check that the status is true, not just well-formed. A wrong current/ended status is the costliest PEP defect, and no validator or assertion catches it:

bash
python .claude/skills/crawler-pep/scripts/pep_summary.py <dataset_name>

The current column should come close to each body's seats (for a multi-term source, the sitting term). A bigger gap means the dates don't mean what the crawler assumes — find which records are affected and why, from the source itself. When some of the source's dates are stale (e.g. re-elected members still carrying their previous term), keep them and let make_occupancy derive the status anyway; explain the gap in a maintainer comment in the YAML. Don't pass status= to paper over stale dates: an explicit status skips make_occupancy's checks entirely, so members who left long ago or have died are still emitted. Don't reconstruct status from other endpoints (votes, rosters): if you think a workaround is needed, bring the evidence and options to the user instead of building it.

Then run the integrity checks in .claude/skills/crawler-pep/validation.md (each should print nothing), and review your diff against zavod/docs/best_practices/merge_checklist.md. Include the pep_summary.py table in your final report.

© opensanctions, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in .claude/skills/crawler-pep of opensanctions/opensanctions.

  • SKILL.md
  • examples.md
  • scripts/pep_summary.py
  • validation.md

Open the folder on GitHubat commit ce59ef9

Compare with similar skills

Crawler Pep next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Crawler Pep compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Crawler Pep this skillopensanctions/opensanctions832—~1.9kAutomated safety check: NotesMIT
Multi Account Scrapingantibrow/anti-detect-browser-skills932—~3.7kAutomated safety check: WarnMIT
Octocode Scrapingbgauryy/octocode949—~1.3kAutomated safety check: PassMIT
Dev Pain Findertinyfish-io/tinyfish-cookbook2.2k—~2.6kAutomated safety check: PassMIT
Deepapidavidondrej/skills4.1k—~2.5kAutomated safety check: PassMIT
Performing Paste Site Monitoring For Credentialsmukul975/Anthropic-Cybersecurity-Skills34k—~3.7kAutomated safety check: WarnApache-2.0

Similar skills

  • Multi Account Scraping

    antibrow/anti-detect-browser-skills

    Run the same scrape or task across many accounts at once - each in its own browser profile with its own fingerprint, cookies and exit IP - and read data from sites that need a session or that answer…

    932 GitHub stars~3.7k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check: warnings
  • Octocode Scraping

    bgauryy/octocode

    A skill your agent uses when extracting or mapping public web content into a local cited corpus: scrape or crawl a URL/docs site, pull tables/pricing/product fields, diagnose blocked or thin pages…

    949 GitHub stars~1.3k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • Dev Pain Finder

    tinyfish-io/tinyfish-cookbook

    Scrape real developer pain points for any keyword, technology, or problem space from Reddit, Hacker News, dev.to, and GitHub Discussions simultaneously — then group complaints by theme, score them…

    2.2k GitHub stars~2.6k tokensUpdated 6 days ago
    Data & AnalyticsAuto-check passed
  • Deepapi

    davidondrej/skills

    Use DeepAPI for all web search, deep research, and web scraping (websites, LinkedIn, GitHub, X/Twitter, YouTube, Instagram) instead of built-in search, research, fetch, or browser tools.

    4.1k GitHub stars~2.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Performing Paste Site Monitoring For Credentials

    mukul975/Anthropic-Cybersecurity-Skills

    Monitor paste sites like Pastebin and GitHub Gists for leaked credentials, API keys, and sensitive data dumps using automated scraping and keyword matching to detect breaches early.

    34k GitHub stars~3.7k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check: warnings
  • GitHub Trending

    hoodini/ai-agents-skills

    Fetch and display GitHub trending repositories and developers.

    281 GitHub stars~2.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed

More from opensanctions/opensanctions

All 11 skills in this repo
  • Crawler Constants To Yml

    opensanctions/opensanctions

    Move hardcoded lookup/config constants (gender maps, header dicts, value translations, column-label maps, date formats) out of a crawler and into the dataset .yml — as datapatch lookups wherever…

    832 GitHub stars~1.4k tokensUpdated today
    Auto-check: notes
  • Dataset Metadata

    opensanctions/opensanctions

    Bring a dataset .yml's metadata in line with house conventions (title, summary, description, coverage, publisher, maintainer comments).

    832 GitHub stars~604 tokensUpdated today
    Auto-check: notes
  • Legislature Metadata

    opensanctions/opensanctions

    Refactor the title, description and coverage frequency of a legislature/parliament PEP dataset .yml into the house style.

    832 GitHub stars~953 tokensUpdated today
    Auto-check passed
  • Name Framework Migration First Step

    opensanctions/opensanctions

    Migrate ad-hoc name cleaning in a crawler to h.reviewnames (Step 1 of the name framework migration).

    832 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Refactor Crawler

    opensanctions/opensanctions

    Rewrite messy or AI-generated crawler code into clean, production-ready style that follows the zavod best practices.

    832 GitHub stars~1.1k tokensUpdated today
    Auto-check: notes
  • Release Datasets

    opensanctions/opensanctions

    Release one or more datasets by adding them to a topical collection, bumping coverage.start, and verifying.

    832 GitHub stars~566 tokensUpdated today
    Auto-check passed

Works with

Questions about Crawler Pep

What does Crawler Pep do?

Scaffold a new PEP (Politically Exposed Persons) crawler — members of a parliament, legislature, senate, chamber of deputies, cabinet, judiciary, or an asset-declaration register — from a source URL…. Crawler Pep is an agent skill from opensanctions/opensanctions. Scaffold a new PEP (Politically Exposed Persons) crawler — members of a parliament, legislature, senate, chamber of deputies, cabinet, judiciary, or an asset-declaration register — from a source URL or GitHub issue.

When should I use Crawler Pep?

Crawler Pep fits situations like: members-of-parliament crawler; tasks that involve Web scraping.

How do I install Crawler Pep in Claude Code?

Run `npx skills add opensanctions/opensanctions --skill crawler-pep -a claude-code`. Or copy the skill folder (.claude/skills/crawler-pep in opensanctions/opensanctions) into .claude/skills/crawler-pep in your project. Claude Code loads it when a task matches its description.

How do I install Crawler Pep in Codex?

Run `npx skills add opensanctions/opensanctions --skill crawler-pep -a codex`. Or copy the skill folder (.claude/skills/crawler-pep in opensanctions/opensanctions) into .agents/skills/crawler-pep in your project. Codex loads it when a task matches its description.

Can I use Crawler Pep in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add opensanctions/opensanctions --skill crawler-pep -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crawler-pep, .gemini/skills/crawler-pep, .github/skills/crawler-pep and .opencode/skills/crawler-pep in your project.

What does Crawler Pep need to run?

Going by SKILL.md and its folder, Crawler Pep needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Edit, Write, Glob, Grep, Bash, WebFetch, WebSearch, Agent.

Does Crawler Pep access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Crawler Pep safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Crawler Pep use?

Crawler Pep is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Crawler Pep use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Crawler Pep?

Skills that share tags, products or a category with Crawler Pep: Multi Account Scraping (antibrow/anti-detect-browser-skills, 932 stars), Octocode Scraping (bgauryy/octocode, 949 stars), Dev Pain Finder (tinyfish-io/tinyfish-cookbook, 2.2k stars) and Deepapi (davidondrej/skills, 4.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Crawler Pep?

opensanctions (a GitHub organization) maintains it in opensanctions/opensanctions, which has 832 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 8, 2026.

Source: opensanctions/opensanctions on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.