Agent skill

Crawler Scraper

by hiroppy in hiroppy/mf-dashboard

A skill your agent uses when adding new scraping targets to the crawler

MITAuto-check passedData & Analytics

Install Crawler Scraper

skills CLI
$ npx skills add hiroppy/mf-dashboard --skill crawler-scraper -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hiroppy/mf-dashboard crawler-scraper --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hiroppy/mf-dashboard.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/crawler-scraper .claude/skills/crawler-scraper && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
crawler-scraper
GitHub stars
418
Token cost
~1.3k tokens
SKILL.md length
420 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when adding new scraping targets to the crawler

  • Adding new scraping targets to the crawler
  • SKILL.md covers Checklist (MUST complete all), File Locations, Template and URL Registration, plus 2 more sections
  • Calls pnpm; reaches moneyforward.com
  • Tasks that involve Web scraping

What it does

Crawler Scraper is an agent skill from hiroppy/mf-dashboard. Use when adding new scraping targets to the crawler

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: マネーフォワードMeを自動化、保有資産の可視化を行います. The licence is MIT.

When your agent uses it

  • Adding new scraping targets to the crawler
  • Tasks that involve Web scraping

Example prompts

  • “/crawler-scraper”

What it can do on your machine

Read from SKILL.md and the folder at commit 21e4a61. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • moneyforward.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Crawler Scraper loads about 1.3k tokens when it runs. Until then it costs about 17 tokens; SKILL.md has 420 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~17
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hiroppy/mf-dashboard at commit 21e4a61, republished under its MIT licence (© hiroppy). 420 words, ~1,326 tokens.

Download SKILL.mdSave it as .claude/skills/crawler-scraper/SKILL.md (or your agent's skills folder).
name
crawler-scraper
description
Use when adding new scraping targets to the crawler

Crawler Scraper Skill

Checklist (MUST complete all)

  • Add URL to packages/meta/src/urls.ts
  • Create scraper function in apps/crawler/src/scrapers/
  • Define types in packages/db/src/types.ts
  • Add repository if new data storage needed
  • Integrate into apps/crawler/src/index.ts
  • Add DOM-independent unit tests for parser decisions and transformations
  • Add authenticated read-only E2E coverage when selectors, navigation, or page structure change

File Locations

PurposeLocation
URLspackages/meta/src/urls.ts
Scrapersapps/crawler/src/scrapers/*.ts
Typespackages/db/src/types.ts
Repositoriespackages/db/src/repositories/
Parsersapps/crawler/src/parsers.ts
Entry pointapps/crawler/src/index.ts

Template

typescript
// apps/crawler/src/scrapers/my-feature.ts
import type { MyData } from "@mf-dashboard/db/types";
import type { Page } from "playwright";
import { mfUrls } from "@mf-dashboard/meta";
import { debug } from "../logger.js";
import { parseJapaneseNumber } from "../parsers.js";

export async function getMyData(page: Page): Promise<MyData> {
  debug("Getting my data from /path...");

  await page.goto(mfUrls.myFeature, {
    waitUntil: "domcontentloaded",
  });
  await page.waitForTimeout(2000);

  // Scraping logic here...
  const rows = page.locator("table tbody tr");
  const count = await rows.count();

  const results: MyItem[] = [];
  for (let i = 0; i < count; i++) {
    const row = rows.nth(i);
    const text = await row
      .locator("td")
      .first()
      .textContent({ timeout: 1000 })
      .catch(() => "");

    results.push({
      // parsed data
    });
  }

  return { items: results };
}

URL Registration

typescript
// packages/meta/src/urls.ts
export const mfUrls = {
  // existing urls...
  myFeature: "https://moneyforward.com/path/to/feature",
};

Testing

Test Types and Priority
PriorityTypeWhen to UseLocation
1UnitPost-extraction parsing, normalization, comparison, mapping, and fail-closed decisions using anonymous strings or objects*.test.ts next to source
2E2ESelectors, navigation, and required HTML/DOM structure on the authenticated real servicetests/e2e/*.test.ts in the e2e Vitest project
3HTML fixture exceptionA failure branch that cannot be produced safely and deterministically against the real serviceThe narrowest applicable unit test
Rules (MUST follow)
  • Extract DOM values into strings or objects, then test the resulting pure transformation and decision logic without Playwright or embedded HTML.
  • Verify selector compatibility and page structure with authenticated read-only E2E tests. Do not trigger refresh, account updates, crawling, or database writes in structure-only E2E tests.
  • NEVER write assertions that depend on actual financial data. Do not assert or log real names, balances, account IDs, or other personal values.
  • E2E assertions may check navigation and structural properties only, including the presence and shape of headings, tables, rows, cells, attributes, and links.
  • Bound structure-only E2E navigation independently from production crawl coverage. If production scans every account for correctness, inspect at most one representative detail page in E2E, skip when no suitable candidate exists, and document the scope difference in the test and pull request.
  • Do not use embedded HTML fixtures merely to imitate the current service DOM. They are allowed only when a failure branch cannot be represented safely and deterministically in read-only E2E.
  • For every HTML fixture exception, keep the markup to the minimum needed and add a nearby comment explaining why authenticated read-only E2E cannot cover that branch.
  • Use anonymous hardcoded strings and objects for unit tests; never copy values from the production database or authenticated pages.
Show full SKILL.md (66 more words)Show less
Running Tests
  • Unit: pnpm --filter @mf-dashboard/crawler test
  • E2E: pnpm --filter @mf-dashboard/crawler test:e2e
  • Local manual testing: SKIP_REFRESH=true pnpm --filter @mf-dashboard/crawler start
  • Debug scripts go in debug/ directory
  • Screenshots saved to debug/ directory

Notes

  • Use parseJapaneseNumber() for Japanese currency format (e.g., "1,234円" → 1234)
  • Use debug() from logger for debug output
  • Handle missing elements gracefully with .catch(() => defaultValue)
  • Always use { timeout: 1000 } for individual element queries to avoid hanging

© hiroppy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/crawler-scraper of hiroppy/mf-dashboard.

Open the folder on GitHubat commit 21e4a61

Compare with similar skills

Crawler Scraper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Crawler Scraper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Crawler Scraper this skillhiroppy/mf-dashboard418—~1.3kAutomated safety check: PassMIT
Tmuxtrpc-group/trpc-agent-go1.8k23 repos~868Automated safety check: PassApache-2.0
Ketch1broseidon/ketch6961 repos~3.9kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5981 repos~2.5kAutomated safety check: PassMIT
Boss Zhipin Scrapereatmoreduck/boss-zhipin-scraper1.5k—~2.6kAutomated safety check: PassMIT
Axyusukebe/ax7191 repos~918Automated safety check: PassMIT

Similar skills

  • Tmux

    trpc-group/trpc-agent-go

    Remote-control tmux sessions for interactive CLIs by sending keystrokes and scraping pane output.

    1.8k GitHub starsUsed in 23 repos~868 tokens
    Data & AnalyticsAuto-check passed
  • Ketch

    1broseidon/ketch

    Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but…

    696 GitHub starsUsed in 1 repo~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    598 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Boss Zhipin Scraper

    eatmoreduck/boss-zhipin-scraper

    Scrape BOSS直聘 (job listing site) via Chrome CDP. An agent skill from eatmoreduck/boss-zhipin-scraper.

    1.5k GitHub stars~2.6k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed
  • Ax

    yusukebe/ax

    Use the ax CLI instead of curl + throwaway parsing scripts whenever you fetch a URL, explore an unknown web page, or extract structured data from HTML.

    719 GitHub starsUsed in 1 repo~918 tokens
    Data & AnalyticsAuto-check passed
  • Anakinscraper

    Anakin-Inc/anakin

    Scrape any website into clean markdown or structured JSON. An agent skill from Anakin-Inc/anakin.

    4.5k GitHub stars~859 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from hiroppy/mf-dashboard

  • Drizzle Erd

    hiroppy/mf-dashboard

    A skill your agent uses when needing to visualize database schema, generate ERD diagrams from Drizzle ORM schemas, or understand table relationships

    418 GitHub stars~627 tokensUpdated today
    Auto-check passed
  • UI Component

    hiroppy/mf-dashboard

    A skill your agent uses when creating new UI components under apps/web/src/components/

    418 GitHub stars~832 tokensUpdated today
    Auto-check passed
  • DB Schema Change

    hiroppy/mf-dashboard

    A skill your agent uses when adding/modifying database tables or columns in Drizzle ORM schema

    418 GitHub stars~362 tokensUpdated today
    Auto-check passed

Questions about Crawler Scraper

What does Crawler Scraper do?

A skill your agent uses when adding new scraping targets to the crawler. Crawler Scraper is an agent skill from hiroppy/mf-dashboard.

When should I use Crawler Scraper?

Crawler Scraper fits situations like: adding new scraping targets to the crawler; tasks that involve Web scraping.

How do I install Crawler Scraper in Claude Code?

Run `npx skills add hiroppy/mf-dashboard --skill crawler-scraper -a claude-code`. Or copy the skill folder (.agents/skills/crawler-scraper in hiroppy/mf-dashboard) into .claude/skills/crawler-scraper in your project. Claude Code loads it when a task matches its description.

How do I install Crawler Scraper in Codex?

Run `npx skills add hiroppy/mf-dashboard --skill crawler-scraper -a codex`. Or copy the skill folder (.agents/skills/crawler-scraper in hiroppy/mf-dashboard) into .agents/skills/crawler-scraper in your project. Codex loads it when a task matches its description.

Can I use Crawler Scraper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hiroppy/mf-dashboard --skill crawler-scraper -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crawler-scraper, .gemini/skills/crawler-scraper, .github/skills/crawler-scraper and .opencode/skills/crawler-scraper in your project.

What does Crawler Scraper need to run?

Going by SKILL.md and its folder, Crawler Scraper needs the command-line tools its instructions call (pnpm).

Does Crawler Scraper access the network?

SKILL.md names 1 domain. In commands or code: moneyforward.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Crawler Scraper safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Crawler Scraper use?

Crawler Scraper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Crawler Scraper use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Crawler Scraper?

Skills that share tags, products or a category with Crawler Scraper: Tmux (trpc-group/trpc-agent-go, 1.8k stars), Ketch (1broseidon/ketch, 696 stars), Crawl4AI Web Scraping (smallnest/goclaw, 598 stars) and Boss Zhipin Scraper (eatmoreduck/boss-zhipin-scraper, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Crawler Scraper?

hiroppy (a GitHub user) maintains it in hiroppy/mf-dashboard, which has 418 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.

Source: hiroppy/mf-dashboard on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.