Agent skill

Typechecker Fixes

by opensanctions in opensanctions/opensanctions

Fix mypy --strict type errors in crawler files. An agent skill from opensanctions/opensanctions.

MITAuto-check passedData & Analytics

Install Typechecker Fixes

skills CLI
$ npx skills add opensanctions/opensanctions --skill typechecker-fixes -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install opensanctions/opensanctions typechecker-fixes --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/opensanctions/opensanctions.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/typechecker-fixes .claude/skills/typechecker-fixes && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
typechecker-fixes
GitHub stars
831
Token cost
~2.6k tokens
SKILL.md length
822 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Fix mypy --strict type errors in crawler files. An agent skill from opensanctions/opensanctions.

  • Works in 10 steps: Add -> None return type to all functions… → Add type annotations to all untyped… → Replace bare dict with the narrowest… → …
  • The user asks to make the typechecker happy
  • SKILL.md covers Read first, Workflow, Patterns (most common first) and General principles
  • Calls mypy

What it does

Typechecker Fixes is an agent skill from opensanctions/opensanctions. Fix mypy --strict type errors in crawler files. Use when the user asks to make the typechecker happy, fix types, or add type annotations to a crawler.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping and Type safety. The repository describes itself as: An open database of international sanctions data, persons of interest and politically exposed persons. The licence is MIT.

When your agent uses it

  • The user asks to make the typechecker happy
  • Add type annotations to a crawler

Example prompts

  • “/typechecker-fixes”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Add -> None return type to all functions that don't return a value
  2. Add type annotations to all untyped parameters
  3. Replace bare dict with the narrowest value type the code actually uses
  4. Replace lxml .xpath(), .find() and .findall() calls with typed h.xpath_* helpers
  5. Replace .text_content() with h.element_text()
  6. Use from zavod.util import Element, ElementOrTree for lxml type annotations
  7. Use Iterator instead of Generator when only yielding
  8. Add return types to small helper functions
  9. Use keyword-only arguments for complex function signatures
  10. Using # type: ignore as a last resort

What it can do on your machine

Read from SKILL.md and the folder at commit 4499adf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • mypy

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Typechecker Fixes loads about 2.6k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 822 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from opensanctions/opensanctions at commit 4499adf, republished under its MIT licence (© opensanctions). 822 words, ~2,566 tokens.

Download SKILL.mdSave it as .claude/skills/typechecker-fixes/SKILL.md (or your agent's skills folder).
name
typechecker-fixes
description
Fix mypy --strict type errors in crawler files. Use when the user asks to make the typechecker happy, fix types, or add type annotations to a crawler.
argument-hint
[crawler.py path]

Typechecker Fixes for Crawlers

Fix mypy strict-mode errors in opensanctions crawler files. Run mypy --strict --explicit-package-bases on the target directory to find errors, then apply the patterns below.

Read first

  • zavod/zavod/helpers/html.py — typed HTML helpers (xpath_strings, xpath_elements, xpath_element, element_text)
  • zavod/zavod/util.py — Element and ElementOrTree type aliases

Workflow

  1. Run mypy --strict <crawler.py> to see current errors
  2. Apply fixes using the patterns below
  3. Run mypy --strict <crawler.py> again to verify errors resolved

Patterns (most common first)

1. Add -> None return type to all functions that don't return a value

This is the single most common fix. Every crawl(), crawl_row(), crawl_item(), parse_*(), and helper function that doesn't return needs -> None.

python
# Before
def crawl(context: Context):
def crawl_row(context: Context, row: dict):
def apply_identifier(context: Context, entity: Entity, id_number_line: str):

# After — pick the narrowest value type the row actually contains (see pattern #3)
def crawl(context: Context) -> None:
def crawl_row(context: Context, row: dict[str, str | None]) -> None:
def apply_identifier(context: Context, entity: Entity, id_number_line: str) -> None:
2. Add type annotations to all untyped parameters
python
# Before
def crawl_item(input_html, context: Context):
def crawl_term(context, link: HtmlElement, ...):

# After
def crawl_item(input_html: Element, context: Context) -> None:
def crawl_term(context: Context, link: HtmlElement, ...) -> None:
3. Replace bare dict with the narrowest value type the code actually uses

Pick the value type by looking at every assignment into the dict. Use Any only as a last resort. In order of preference:

  1. dict[str, str] — all values are strings (e.g. h.xpath_strings(...)[0], string literals, city.get("key", "")).
  2. dict[str, str | None] — some values can legitimately be None (e.g. element.text, element.get("attr"), record.get(key) with no default). At consumer sites that need str, narrow with a local + assert value is not None or relax the consumer's signature to accept None.
  3. dict[str, Any] — only when values are genuinely heterogeneous (mixed types that can't be expressed as a simple union, e.g. str, int, nested list/dict). Prefer a TypedDict if the dict has a fixed schema.
python
# Before
def crawl_item(input_dict: dict, context: Context):
json_data = { ... }

# After — values are all strings
def crawl_item(input_dict: dict[str, str], context: Context) -> None:
json_data: dict[str, str] = { ... }

# After — values include str | None from element.text
record: dict[str, str | None] = {}
record["name"] = row[0].text                       # str | None
record["url"] = urljoin(base, row[0].get("href"))  # str

# After — truly heterogeneous (resort to Any)
from typing import Any
item: dict[str, Any] = {"name": "x", "count": 3, "tags": [...]}

When you tighten to dict[str, str | None], expect two or three follow-up errors at call sites that expect strict str (e.g. context.fetch_html). Handle each by either:

  • Pulling the value into a local and asserting non-None: url = record["url"]; assert url is not None.
  • Widening the consumer's signature if None is a semantically valid input (e.g. is_valid(regno: str | None) returning False for None).

Use lowercase dict, list, set, tuple — not the deprecated Dict, List, Set, Tuple from typing. While fixing types, also migrate any existing typing.Dict etc. to builtins.

4. Replace lxml .xpath(), .find() and .findall() calls with typed h.xpath_* helpers

The raw lxml .xpath() returns Any. Use the zavod helpers instead:

python
# Before — returns Any
links = doc.xpath(".//a/@href")
elements = doc.xpath('.//div[@class="item"]')
text = doc.xpath(".//h1/text()")[0]

# After — properly typed
links = h.xpath_strings(doc, ".//a/@href")
elements = h.xpath_elements(doc, './/div[@class="item"]')
text = h.xpath_string(doc, ".//h1/text()")

When iterating over elements only to extract an attribute (e.g. .get("href")), move the attribute into the xpath and use h.xpath_strings instead:

python
# Before
for anchor in doc.xpath('//a[contains(@class, "name")]'):
    url = anchor.get("href")
    crawl_page(context, url)

# Also before (already migrated to xpath_elements but still using .get)
for anchor in h.xpath_elements(doc, '//a[contains(@class, "name")]'):
    url = anchor.get("href")
    crawl_page(context, url)

# After
for url in h.xpath_strings(doc, '//a[contains(@class, "name")]/@href'):
    crawl_page(context, url)

Use h.xpath_element() (singular) when you expect exactly one match:

python
# Before
divs = doc.xpath(divs_xpath)
assert len(divs) == 1
content = divs[0]

# After
content = h.xpath_element(doc, divs_xpath)
5. Replace .text_content() with h.element_text()

h.element_text() calls text_content() internally and applies collapse_spaces + strip. If the original code was calling squash_spaces or collapse_spaces on the result, that's now redundant and should be removed. If the extracted text is used for lookups (check the lookups: section in the crawler's .yml file) or exact comparisons, pass squash=False to preserve the original whitespace.

Watch out for code that splits the text on \n (or otherwise relies on line breaks) — the default squash=True collapses newlines into spaces, which silently breaks the splitting and merges all rows into one. Always pass squash=False in that case.

python
# Before
name_info = summary.text.strip()
body = body_els[0].text_content().strip()
category = squash_spaces(row.pop("category").text_content())

# After
name_info = h.element_text(summary)
body = h.element_text(body_els[0])
category = h.element_text(row.pop("category"))

# When exact text matters (used in lookups or comparisons):
label = h.element_text(el, squash=False)
6. Use from zavod.util import Element, ElementOrTree for lxml type annotations

Don't use lxml.etree._Element (private API) or xml.etree.ElementTree. Use the re-exported types:

python
# Before
from lxml import etree
def parse_record(context: Context, el: etree._Element):

# After
from zavod.util import Element
def parse_record(context: Context, el: Element) -> None:
Show full SKILL.md (323 more words)Show less
11. Use Iterator instead of Generator when only yielding
python
# Before
from typing import Generator
def parse_csv(context: Context, path: str) -> Generator[Item, None, None]:

# After
from typing import Iterator
def parse_csv(context: Context, path: str) -> Iterator[Item]:
12. Add return types to small helper functions
python
# Before
def clean_address(text):
def extract_passport_no(text):

# After
def clean_address(text: str | None) -> list[str] | None:
def extract_passport_no(text: str | None) -> list[str] | None:
13. Use keyword-only arguments for complex function signatures

When a function has many parameters, add * to force keyword arguments — this catches argument-order bugs at the call site:

python
# Before
def emit_linked_org(context, vessel_id, names, role, date):
    ...
emit_linked_org(context, vessel.id, related_ros, "Related Recognised Organization", start_date)

# After
def emit_linked_org(context: Context, *, vessel_id: str | None, names: str, role: str, date: str | None) -> None:
    ...
emit_linked_org(context, vessel_id=vessel.id, names=related_ros, role="Related Recognised Organization", date=start_date)
Add a comment to functions that return tuples

If the function returns a tuple, add a docstring comment to briefly describe the contents of the tuple. This is not a typechecker fix, but it helps readability since tuples don't have named fields.

14. Using # type: ignore as a last resort

# type: ignore[<code>] is only acceptable when the error originates from an external library with incomplete stubs (e.g. requests, operator.concat) and the runtime behaviour is demonstrably correct. Before reaching for it:

  1. Confirm no pattern above can fix the error without changing logic.
  2. Verify the runtime behaviour is correct independently of what mypy says.
  3. Always add a comment explaining why — describe the stub gap, not just the error code.
python
# Bad — silences the error without explaining it
context.http.cookies.set(name, value)  # type: ignore[no-untyped-call]

# Good — explanatory comment on the line before, type: ignore on the offending line
# requests stubs don't type RequestsCookieJar.set()
context.http.cookies.set(name, value)  # type: ignore[no-untyped-call]

# Good — explains the typing gap vs runtime reality
# operator.concat is typed Sequence→Sequence, not list→list; runtime is correct
aliases = reduce(concat, alias_lists, [])  # type: ignore[arg-type]

Never use # type: ignore to paper over errors in your own code — those should be fixed properly.

General principles

  • Never change logic. These are type-annotation-only fixes. Do not change control flow, data transformations, or output. Each fix must be resolved by exactly and narrowly applying one of the rules above. If a type error cannot be resolved that way without altering behavior, leave the error. It is fine to leave some errors unfixed rather than risk changing what the crawler emits.
  • Prefer the narrowest correct type. For dicts, follow the preference order in pattern #3: dict[str, str] > dict[str, str | None] > dict[str, Any]. Only fall back to Any when the values are genuinely heterogeneous.
  • Use str | None union syntax, not Optional[str], and lowercase dict/list/set not Dict/List/Set — but only when you're already editing the line for another reason. Do not make cosmetic-only changes to lines that have no type errors.
  • The context: Context parameter should always be first in crawler functions.

© opensanctions, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/typechecker-fixes of opensanctions/opensanctions.

Open the folder on GitHubat commit 4499adf

Compare with similar skills

Typechecker Fixes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Typechecker Fixes compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Typechecker Fixes this skillopensanctions/opensanctions831—~2.6kAutomated safety check: PassMIT
Tmuxtrpc-group/trpc-agent-go1.8k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch6961 repos~3.9kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5981 repos~2.5kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.8k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    696 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    598 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from opensanctions/opensanctions

All 11 skills in this repo
  • Crawler Pep

    opensanctions/opensanctions

    Scaffold a new PEP (Politically Exposed Persons) crawler — members of a parliament, legislature, senate, chamber of deputies, cabinet, judiciary, or an asset-declaration register — from a source URL…

    831 GitHub stars~1.9k tokensUpdated today
    Auto-check: notes
  • Crawler Constants To Yml

    opensanctions/opensanctions

    Move hardcoded lookup/config constants (gender maps, header dicts, value translations, column-label maps, date formats) out of a crawler and into the dataset .yml — as datapatch lookups wherever…

    831 GitHub stars~1.4k tokensUpdated today
    Auto-check: notes
  • Dataset Metadata

    opensanctions/opensanctions

    Bring a dataset .yml's metadata in line with house conventions (title, summary, description, coverage, publisher, maintainer comments).

    831 GitHub stars~604 tokensUpdated today
    Auto-check: notes
  • Legislature Metadata

    opensanctions/opensanctions

    Refactor the title, description and coverage frequency of a legislature/parliament PEP dataset .yml into the house style.

    831 GitHub stars~953 tokensUpdated today
    Auto-check passed
  • Name Framework Migration First Step

    opensanctions/opensanctions

    Migrate ad-hoc name cleaning in a crawler to h.reviewnames (Step 1 of the name framework migration).

    831 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Refactor Crawler

    opensanctions/opensanctions

    Rewrite messy or AI-generated crawler code into clean, production-ready style that follows the zavod best practices.

    831 GitHub stars~1.1k tokensUpdated today
    Auto-check: notes

Questions about Typechecker Fixes

What does Typechecker Fixes do?

Fix mypy --strict type errors in crawler files. An agent skill from opensanctions/opensanctions. Typechecker Fixes is an agent skill from opensanctions/opensanctions. Fix mypy --strict type errors in crawler files.

When should I use Typechecker Fixes?

Typechecker Fixes fits situations like: the user asks to make the typechecker happy; add type annotations to a crawler.

How do I install Typechecker Fixes in Claude Code?

Run `npx skills add opensanctions/opensanctions --skill typechecker-fixes -a claude-code`. Or copy the skill folder (.claude/skills/typechecker-fixes in opensanctions/opensanctions) into .claude/skills/typechecker-fixes in your project. Claude Code loads it when a task matches its description.

How do I install Typechecker Fixes in Codex?

Run `npx skills add opensanctions/opensanctions --skill typechecker-fixes -a codex`. Or copy the skill folder (.claude/skills/typechecker-fixes in opensanctions/opensanctions) into .agents/skills/typechecker-fixes in your project. Codex loads it when a task matches its description.

Can I use Typechecker Fixes in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add opensanctions/opensanctions --skill typechecker-fixes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/typechecker-fixes, .gemini/skills/typechecker-fixes, .github/skills/typechecker-fixes and .opencode/skills/typechecker-fixes in your project.

What does Typechecker Fixes need to run?

Going by SKILL.md and its folder, Typechecker Fixes needs the command-line tools its instructions call (mypy). Our summary lists: Python 3.

Does Typechecker Fixes access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Typechecker Fixes safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Typechecker Fixes use?

Typechecker Fixes is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Typechecker Fixes use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Typechecker Fixes?

Skills that share tags, products or a category with Typechecker Fixes: Tmux (trpc-group/trpc-agent-go, 1.8k stars), Ketch (1broseidon/ketch, 696 stars), Crawl4AI Web Scraping (smallnest/goclaw, 598 stars) and Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Typechecker Fixes?

opensanctions (a GitHub organization) maintains it in opensanctions/opensanctions, which has 831 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 7, 2026.

Source: opensanctions/opensanctions on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.