Agent skill

Crawler Sanctions

by opensanctions in opensanctions/opensanctions

Scaffold a new sanctions list crawler from a source URL or GitHub issue

MITAuto-check: notesData & Analytics

Install Crawler Sanctions

skills CLI
$ npx skills add opensanctions/opensanctions --skill crawler-sanctions -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install opensanctions/opensanctions crawler-sanctions --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/opensanctions/opensanctions.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/crawler-sanctions .claude/skills/crawler-sanctions && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
crawler-sanctions
GitHub stars
832
Token cost
~1.9k tokens
SKILL.md length
538 words
Files
2
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Scaffold a new sanctions list crawler from a source URL or GitHub issue

  • Works in 4 steps: Understand the source → YAML metadata — sanctions-specific parts → Write the crawler module → …
  • Tasks that involve Web scraping
  • SKILL.md covers Step 1: Understand the source, Step 2: YAML metadata —…, Step 3: Write the crawler module and Step 4: Sanctions-specific…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Crawler Sanctions is an agent skill from opensanctions/opensanctions. Scaffold a new sanctions list crawler from a source URL or GitHub issue

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `examples.md`).

It sits in Data & Analytics, covering Web scraping. It works with GitHub. The repository describes itself as: An open database of international sanctions data, persons of interest and politically exposed persons. The licence is MIT.

When your agent uses it

  • Tasks that involve Web scraping

Example prompts

  • “/crawler-sanctions”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Edit, Write, Glob, Grep, Bash, WebFetch, WebSearch, Agent

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Understand the source
  2. YAML metadata — sanctions-specific parts
  3. Write the crawler module
  4. Sanctions-specific validation checks

What it can do on your machine

Read from SKILL.md and the folder at commit ce59ef9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Edit
    • Write
    • Glob
    • Grep
    • Bash
    • WebFetch
    • WebSearch
    • Agent

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, yaml and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Crawler Sanctions loads about 1.9k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 538 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Edit, Write, Glob, Grep, Bash, WebFetch, WebSearch, Agent

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from opensanctions/opensanctions at commit ce59ef9, republished under its MIT licence (© opensanctions). 538 words, ~1,937 tokens.

Download SKILL.mdSave it as .claude/skills/crawler-sanctions/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
crawler-sanctions
description
Scaffold a new sanctions list crawler from a source URL or GitHub issue
allowed-tools
Read, Edit, Write, Glob, Grep, Bash, WebFetch, WebSearch, Agent

New Sanctions Crawler

Create a new sanctions list crawler. The user will provide a target path, source data URL, and/or a GitHub issue URL: $ARGUMENTS

If given a GitHub issue URL, fetch it first to extract the data source URL and any context about the dataset before proceeding.

Before writing any code, read these files — they contain everything you need:

  1. .claude/docs/crawler-guide.md — shared crawler patterns (YAML template, fetching data, entity creation, helpers, lookups, FTM schemata, qsv analysis)
  2. .claude/skills/crawler-sanctions/examples.md — full sanctions code examples

Do NOT search the repository for similar crawlers or patterns. The guide and examples above are the authoritative reference. Do not read datasets/CLAUDE.md or other crawler source files for patterns — use only the files listed above.

Step 1: Understand the source

Before writing any code, inspect the data source. In addition to the general checks (fields, date formats, language, record count), sanctions sources need:

  • Identify entity types present: persons, organizations, vessels, aircraft
  • Identify how sanctions programs are labeled in the source
  • Check if the source provides unique opaque IDs per entry (for slug-based IDs)
  • Check if relationships between entities are encoded (ownership, family, associates)
  • Identify the data structure: flat list vs nested XML vs paginated API

Step 2: YAML metadata — sanctions-specific parts

Use the generic YAML template from the crawler guide. Sanctions-specific additions:

yaml
tags:
  - list.sanction
  - issuer.west          # optional

assertions:
  min:
    schema_entities:
      Person: 1000       # band widths: see "Data assertions" in zavod/docs/metadata.md
      Organization: 200
      Sanction: 1000
    country_entities:
      cc: 100
  max:
    schema_entities:
      Person: 5000
      Organization: 1000
  • coverage.frequency: daily (house default for sanctions — see zavod/docs/metadata.md). Don't add a cron schedule: unless the run must follow the source's own publication time.
  • Assert Sanction entity counts alongside Person/Organization counts.
Sanctions-specific lookups

The most important sanctions lookup maps source program names to OpenSanctions keys:

yaml
lookups:
  # Entity type dispatch (when source uses custom type labels)
  type.entity:
    lowercase: true
    options:
      - match: [individual, person]
        value: Person
      - match: [entity, company, organization]
        value: Organization
      - match: [vessel, ship]
        value: Vessel

  # Map source program names to OpenSanctions program keys
  sanction.program:
    options:
      - match: "Executive Order 13224"
        value: US-EO13224

  # Date edge cases common in sanctions data
  type.date:
    options:
      - match: "1972-08-10 or 1972-08-11"
        values: ["1972-08-10", "1972-08-11"]
      - match: "1975-19-25"       # typo
        value: "1975"

type.* lookups are applied automatically by entity.add(). The sanction.program lookup must be called explicitly via h.lookup_sanction_program_key().

Step 3: Write the crawler module

Show full SKILL.md (256 more words)Show less
Sanction entity creation

Full reference: zavod/docs/programs.md

h.make_sanction() automatically sets country, authority, and sourceUrl from dataset metadata. The key parameters:

python
sanction = h.make_sanction(
    context,
    entity,                                    # the sanctioned entity (required)
    key=entry_id,                              # disambiguator when entity has multiple sanctions
    program_name=program,                      # human-readable program name
    source_program_key=program,                # raw value from source (preserved as original_value)
    program_key=h.lookup_sanction_program_key(  # OpenSanctions program key from yaml lookup
        context, program
    ),
    start_date=listing_date,                   # optional: when sanction began
    end_date=end_date,                         # optional: when sanction ended
)
  • key: Use when an entity appears on multiple sanctions lists/programs. The sanction ID is make_id("Sanction", entity.id, key), so key disambiguates multiple sanctions per entity.
  • program_key: Always go through h.lookup_sanction_program_key() which reads the sanction.program yaml lookup. Add entries to the lookup as you encounter new program names.
  • source_program_key: The raw program string from the source, preserved as original_value on the programId property for auditability.
  • Always also set entity.add("topics", "sanction") on the sanctioned entity.

For simple datasets with a single known program, you can skip the lookup:

python
sanction = h.make_sanction(context, entity, program_key="US-DOS-CU-PAL")
Checking if a sanction is active
python
if h.is_active(sanction):
    entity.add("topics", "sanction")
# Only mark as sanctioned if the sanction is currently active
Name handling in sanctions crawlers

Full reference: zavod/docs/extract/names.md

Sanctioned names are legal designations — do not use LLM-based name cleaning. Any normalisation must be human-reviewed via the stateful review system, or handled with explicit lookup entries.

Relationships between sanctioned entities

See the crawler guide for the generic Family and Ownership patterns. See examples.md for UnknownLink (sanctions-specific untyped relationships).

De-listing and modification tracking

When the source tracks modifications and de-listings, use sanction.add("endDate", ...) for de-listings and sanction.add("modifiedAt", ...) for amendments. See examples.md for the full pattern.

LLM extraction from free-text fields

Full reference: zavod/docs/data_reviews.md

For sources with unstructured "remarks" fields, use GPT extraction with the stateful review system. Requires ci_test: false. See examples.md for the pattern.

Step 4: Sanctions-specific validation checks

After running zavod crawl, use these sanctions-specific qsv checks (see the crawler guide for general qsv patterns):

bash
# Entity counts by schema
qsv search -s prop "^Person:id$" data/datasets/cc_dataset/statements.pack | qsv count
qsv search -s prop "^Organization:id$" data/datasets/cc_dataset/statements.pack | qsv count
qsv search -s prop "^Sanction:id$" data/datasets/cc_dataset/statements.pack | qsv count

# Sanction program distribution
qsv search -s prop "^Sanction:program$" data/datasets/cc_dataset/statements.pack | qsv frequency -s value

# Every Sanction:entity must point to a real entity
qsv search -s prop "^Sanction:entity$" data/datasets/cc_dataset/statements.pack | qsv select value | qsv behead | sort > /tmp/sanction_targets.txt && qsv search -s prop ":id$" data/datasets/cc_dataset/statements.pack | qsv select entity_id | qsv behead | sort -u > /tmp/all_entities.txt && comm -23 /tmp/sanction_targets.txt /tmp/all_entities.txt

# Check all entities have topics=sanction
qsv search -s prop ":id$" data/datasets/cc_dataset/statements.pack | qsv select entity_id | qsv behead | sort -u > /tmp/all_ids.txt && qsv search -s prop ":topics$" data/datasets/cc_dataset/statements.pack | qsv search -s value "^sanction$" | qsv select entity_id | qsv behead | sort -u > /tmp/sanctioned.txt && comm -23 /tmp/all_ids.txt /tmp/sanctioned.txt

Then run zavod export datasets/cc/dataset/cc_dataset.yml, which checks the dataset validators and assertions.

© opensanctions, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/crawler-sanctions of opensanctions/opensanctions.

  • SKILL.md
  • examples.md

Open the folder on GitHubat commit ce59ef9

Compare with similar skills

Crawler Sanctions next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Crawler Sanctions compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Crawler Sanctions this skillopensanctions/opensanctions832—~1.9kAutomated safety check: NotesMIT
Multi Account Scrapingantibrow/anti-detect-browser-skills932—~3.7kAutomated safety check: WarnMIT
Octocode Scrapingbgauryy/octocode949—~1.3kAutomated safety check: PassMIT
Dev Pain Findertinyfish-io/tinyfish-cookbook2.2k—~2.6kAutomated safety check: PassMIT
Deepapidavidondrej/skills4.1k—~2.5kAutomated safety check: PassMIT
Performing Paste Site Monitoring For Credentialsmukul975/Anthropic-Cybersecurity-Skills34k—~3.7kAutomated safety check: WarnApache-2.0

Similar skills

  • Multi Account Scraping

    antibrow/anti-detect-browser-skills

    Run the same scrape or task across many accounts at once - each in its own browser profile with its own fingerprint, cookies and exit IP - and read data from sites that need a session or that answer…

    932 GitHub stars~3.7k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check: warnings
  • Octocode Scraping

    bgauryy/octocode

    A skill your agent uses when extracting or mapping public web content into a local cited corpus: scrape or crawl a URL/docs site, pull tables/pricing/product fields, diagnose blocked or thin pages…

    949 GitHub stars~1.3k tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed
  • Dev Pain Finder

    tinyfish-io/tinyfish-cookbook

    Scrape real developer pain points for any keyword, technology, or problem space from Reddit, Hacker News, dev.to, and GitHub Discussions simultaneously — then group complaints by theme, score them…

    2.2k GitHub stars~2.6k tokensUpdated 6 days ago
    Data & AnalyticsAuto-check passed
  • Deepapi

    davidondrej/skills

    Use DeepAPI for all web search, deep research, and web scraping (websites, LinkedIn, GitHub, X/Twitter, YouTube, Instagram) instead of built-in search, research, fetch, or browser tools.

    4.1k GitHub stars~2.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Performing Paste Site Monitoring For Credentials

    mukul975/Anthropic-Cybersecurity-Skills

    Monitor paste sites like Pastebin and GitHub Gists for leaked credentials, API keys, and sensitive data dumps using automated scraping and keyword matching to detect breaches early.

    34k GitHub stars~3.7k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check: warnings
  • GitHub Trending

    hoodini/ai-agents-skills

    Fetch and display GitHub trending repositories and developers.

    281 GitHub stars~2.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed

More from opensanctions/opensanctions

All 11 skills in this repo
  • Crawler Pep

    opensanctions/opensanctions

    Scaffold a new PEP (Politically Exposed Persons) crawler — members of a parliament, legislature, senate, chamber of deputies, cabinet, judiciary, or an asset-declaration register — from a source URL…

    832 GitHub stars~1.9k tokensUpdated today
    Auto-check: notes
  • Crawler Constants To Yml

    opensanctions/opensanctions

    Move hardcoded lookup/config constants (gender maps, header dicts, value translations, column-label maps, date formats) out of a crawler and into the dataset .yml — as datapatch lookups wherever…

    832 GitHub stars~1.4k tokensUpdated today
    Auto-check: notes
  • Dataset Metadata

    opensanctions/opensanctions

    Bring a dataset .yml's metadata in line with house conventions (title, summary, description, coverage, publisher, maintainer comments).

    832 GitHub stars~604 tokensUpdated today
    Auto-check: notes
  • Legislature Metadata

    opensanctions/opensanctions

    Refactor the title, description and coverage frequency of a legislature/parliament PEP dataset .yml into the house style.

    832 GitHub stars~953 tokensUpdated today
    Auto-check passed
  • Name Framework Migration First Step

    opensanctions/opensanctions

    Migrate ad-hoc name cleaning in a crawler to h.reviewnames (Step 1 of the name framework migration).

    832 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Refactor Crawler

    opensanctions/opensanctions

    Rewrite messy or AI-generated crawler code into clean, production-ready style that follows the zavod best practices.

    832 GitHub stars~1.1k tokensUpdated today
    Auto-check: notes

Works with

Questions about Crawler Sanctions

What does Crawler Sanctions do?

Scaffold a new sanctions list crawler from a source URL or GitHub issue. Crawler Sanctions is an agent skill from opensanctions/opensanctions.

When should I use Crawler Sanctions?

Crawler Sanctions fits situations like: tasks that involve Web scraping.

How do I install Crawler Sanctions in Claude Code?

Run `npx skills add opensanctions/opensanctions --skill crawler-sanctions -a claude-code`. Or copy the skill folder (.claude/skills/crawler-sanctions in opensanctions/opensanctions) into .claude/skills/crawler-sanctions in your project. Claude Code loads it when a task matches its description.

How do I install Crawler Sanctions in Codex?

Run `npx skills add opensanctions/opensanctions --skill crawler-sanctions -a codex`. Or copy the skill folder (.claude/skills/crawler-sanctions in opensanctions/opensanctions) into .agents/skills/crawler-sanctions in your project. Codex loads it when a task matches its description.

Can I use Crawler Sanctions in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add opensanctions/opensanctions --skill crawler-sanctions -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crawler-sanctions, .gemini/skills/crawler-sanctions, .github/skills/crawler-sanctions and .opencode/skills/crawler-sanctions in your project.

What does Crawler Sanctions need to run?

SKILL.md names no scripts, command-line tools or credentials: Crawler Sanctions is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Edit, Write, Glob, Grep, Bash, WebFetch, WebSearch, Agent.

Does Crawler Sanctions access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Crawler Sanctions safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Crawler Sanctions use?

Crawler Sanctions is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Crawler Sanctions use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Crawler Sanctions?

Skills that share tags, products or a category with Crawler Sanctions: Multi Account Scraping (antibrow/anti-detect-browser-skills, 932 stars), Octocode Scraping (bgauryy/octocode, 949 stars), Dev Pain Finder (tinyfish-io/tinyfish-cookbook, 2.2k stars) and Deepapi (davidondrej/skills, 4.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Crawler Sanctions?

opensanctions (a GitHub organization) maintains it in opensanctions/opensanctions, which has 832 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 8, 2026.

Source: opensanctions/opensanctions on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.