Agent skill

Data Scraper Agent

by affaan-m in affaan-m/ECC

任意のパブリックソース(ジョブボード、価格、ニュース、GitHub、スポーツなど)用の完全自動化されたAI搭載データ収集エージェントを構築します。スケジュールでスクレイプし、無料LLM(Gemini Flash)でデータを豊かにし、Notion/Sheets/Supabaseに結果を保存し、ユーザーフィードバックから学習します。GitHub…

MITAuto-check passedDevOps & Cloud

Install Data Scraper Agent

skills CLI
$ npx skills add affaan-m/ECC --skill data-scraper-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC data-scraper-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/ja-JP/skills/data-scraper-agent .claude/skills/data-scraper-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-scraper-agent
GitHub stars
275k
Token cost
~377 tokens
SKILL.md length
69 words
Files
1
Skills in repo
645
Repo updated
First seen
Licence
MIT

At a glance

任意のパブリックソース(ジョブボード、価格、ニュース、GitHub、スポーツなど)用の完全自動化されたAI搭載データ収集エージェントを構築します。スケジュールでスクレイプし、無料LLM(Gemini Flash)でデータを豊かにし、Notion/Sheets/Supabaseに結果を保存し、ユーザーフィードバックから学習します。GitHub…

  • Works in 6 steps: ソースを定義 - どこからスクレイプするか、何を抽出するか → スクレイパーを構築 - BeautifulSoup または Playwright… → LLMを構成 - Gemini Flash でテキストをスコア付け/要約/分類 → …
  • Tasks that involve Web scraping
  • SKILL.md covers アクティベーション時期, コアコンセプト, ワークフロー and 例
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Scraper Agent is an agent skill from affaan-m/ECC. 任意のパブリックソース(ジョブボード、価格、ニュース、GitHub、スポーツなど)用の完全自動化されたAI搭載データ収集エージェントを構築します。スケジュールでスクレイプし、無料LLM(Gemini Flash)でデータを豊かにし、Notion/Sheets/Supabaseに結果を保存し、ユーザーフィードバックから学習します。GitHub Actions上で100%無料で実行。ユーザーがパブリックデータを自動的に監視、収集、または追跡したい場合に使用します。

Its SKILL.md is about 380 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Web scraping and CI/CD. It works with Supabase, GitHub Actions, Notion and GitHub. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Tasks that involve Web scraping
  • Tasks that involve CI/CD

Example prompts

  • “/data-scraper-agent”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. ソースを定義 - どこからスクレイプするか、何を抽出するか
  2. スクレイパーを構築 - BeautifulSoup または Playwright ベースのコレクタ
  3. LLMを構成 - Gemini Flash でテキストをスコア付け/要約/分類
  4. ストレージを設定 - Notion、Sheets、Supabase のいずれか
  5. GitHub Actions を設定 - 毎日/毎週実行するスケジュール
  6. フィードバックループを追加 - ユーザーの判断から学習

What it can do on your machine

Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Scraper Agent loads about 377 tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 69 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~377

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 69 words, ~377 tokens.

Download SKILL.mdSave it as .claude/skills/data-scraper-agent/SKILL.md (or your agent's skills folder).
name
data-scraper-agent
description
任意のパブリックソース(ジョブボード、価格、ニュース、GitHub、スポーツなど)用の完全自動化されたAI搭載データ収集エージェントを構築します。スケジュールでスクレイプし、無料LLM(Gemini Flash)でデータを豊かにし、Notion/Sheets/Supabaseに結果を保存し、ユーザーフィードバックから学習します。GitHub Actions上で100%無料で実行。ユーザーがパブリックデータを自動的に監視、収集、または追跡したい場合に使用します。
origin
community

データスクレイパーエージェント

任意のパブリックデータソース用の本番環境対応、AI搭載データ収集エージェントを構築。 スケジュールで実行され、無料LLMで結果を豊かにし、データベースに保存し、時間とともに改善されます。

スタック:Python · Gemini Flash(無料) · GitHub Actions(無料) · Notion / Sheets / Supabase

アクティベーション時期

  • ユーザーが任意のパブリックWebサイトまたはAPIをスクレイプまたは監視したい場合
  • ユーザーが「チェックするボットを構築」「Xを監視」「データを収集」と言う
  • ユーザーがジョブ、価格、ニュース、リポ、スポーツスコア、イベント、リストを追跡したい場合
  • ユーザーがホスティング用に支払わずにデータ収集を自動化する方法を尋ねる
  • ユーザーが決定に基づいて時間とともにより スマートになるエージェントを望む

コアコンセプト

3つのレイヤー

すべてのデータスクレイパーエージェントには3つのレイヤーがあります:

COLLECT → ENRICH → STORE
  │           │        │
Scraper    AI (LLM)  Database
runs on    scores/   Notion /
schedule   summarises Sheets /
           & classifies Supabase
無料スタック
LayerToolWhy
COLLECTPlaywright/BeautifulSoup無料のオープンソーススクレイピング
ENRICHGemini Flash無料で高速LLM
STORESupabase / Sheets無料データベースとスプレッドシート
SCHEDULEGitHub Actions無料クロンジョブ

ワークフロー

  1. ソースを定義 - どこからスクレイプするか、何を抽出するか
  2. スクレイパーを構築 - BeautifulSoup または Playwright ベースのコレクタ
  3. LLMを構成 - Gemini Flash でテキストをスコア付け/要約/分類
  4. ストレージを設定 - Notion、Sheets、Supabase のいずれか
  5. GitHub Actions を設定 - 毎日/毎週実行するスケジュール
  6. フィードバックループを追加 - ユーザーの判断から学習

例

  • ジョブボード監視:新しい公開

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/ja-JP/skills/data-scraper-agent of affaan-m/ECC.

Open the folder on GitHubat commit ef648e0

Compare with similar skills

Data Scraper Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Scraper Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Scraper Agent this skillaffaan-m/ECC275k—~377Automated safety check: PassMIT
Data Scraper Agentmajiayu000/claude-skill-registry6666 repos~6.3kAutomated safety check: NotesMIT
Apify CI Integrationjeremylongshore/tons-of-skills-marketplace2.8k—~1.7kAutomated safety check: PassMIT
Supabase CI Integrationjeremylongshore/tons-of-skills-marketplace2.8k—~2kAutomated safety check: PassMIT
Nushellccusage/ccusage19k—~938Automated safety check: PassCustom licence
Clawsweeperopenclaw/openclaw392k—~3kAutomated safety check: PassMIT

Similar skills

  • Data Scraper Agent

    majiayu000/claude-skill-registry

    Build a fully automated AI-powered data collection agent for any public source — job boards, prices, news, GitHub, sports, anything.

    666 GitHub starsUsed in 6 repos~6.3k tokens
    DevOps & CloudAuto-check: notes
  • Apify CI Integration

    jeremylongshore/tons-of-skills-marketplace

    Configure CI/CD pipelines for Apify Actor builds and deployments.

    2.8k GitHub stars~1.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Supabase CI Integration

    jeremylongshore/tons-of-skills-marketplace

    Configure Supabase continuous-integration and deployment pipelines with GitHub Actions: link projects, push migrations, deploy Edge Functions, generate types, and run tests against local Supabase…

    2.8k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Nushell

    ccusage/ccusage

    Guides ccusage Nushell scripts. An agent skill from ccusage/ccusage.

    19k GitHub stars~938 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Clawsweeper

    openclaw/openclaw

    A skill your agent uses for all ClawSweeper work: OpenClaw issue/PR sweep reports, repair jobs, cloud fix PRs, @clawsweeper maintainer mention commands, trusted ClawSweeper-reviewed…

    392k GitHub stars~3k tokensUpdated today
    DevOps & CloudAuto-check passed
  • AI News Radar

    LearnPrompt/ai-news-radar

    A skill your agent uses when working on AI News Radar, 24 小时 AI 更新雷达, AI 更新雷达, 伯乐Skill, or Scout Skill: finding high-signal AI/tech sources, adding RSS/OPML/GitHub feeds, checking source health…

    1.8k GitHub stars~2.5k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from affaan-m/ECC

All 645 skills in this repo
  • Videodb

    affaan-m/ECC

    Ingest, index, search, edit, and monitor video and audio with the VideoDB Python SDK — upload from files, URLs, or RTSP feeds, build spoken and scene indexes with timestamped search and playable…

    275k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    275k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    275k GitHub stars~2.9k tokensUpdated 3 days ago
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    275k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed
  • Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.

    275k GitHub stars~3.5k tokensUpdated 3 days ago
    Auto-check passed
  • Adds one optional external Codex critique that tries to break a council's decision draft, sent to OpenAI only after you consent.

    275k GitHub stars~1.5k tokensUpdated 3 days ago
    Auto-check passed

Questions about Data Scraper Agent

What does Data Scraper Agent do?

任意のパブリックソース(ジョブボード、価格、ニュース、GitHub、スポーツなど)用の完全自動化されたAI搭載データ収集エージェントを構築します。スケジュールでスクレイプし、無料LLM(Gemini Flash)でデータを豊かにし、Notion/Sheets/Supabaseに結果を保存し、ユーザーフィードバックから学習します。GitHub…. Data Scraper Agent is an agent skill from affaan-m/ECC.

When should I use Data Scraper Agent?

Data Scraper Agent fits situations like: tasks that involve Web scraping; tasks that involve CI/CD.

How do I install Data Scraper Agent in Claude Code?

Run `npx skills add affaan-m/ECC --skill data-scraper-agent -a claude-code`. Or copy the skill folder (docs/ja-JP/skills/data-scraper-agent in affaan-m/ECC) into .claude/skills/data-scraper-agent in your project. Claude Code loads it when a task matches its description.

How do I install Data Scraper Agent in Codex?

Run `npx skills add affaan-m/ECC --skill data-scraper-agent -a codex`. Or copy the skill folder (docs/ja-JP/skills/data-scraper-agent in affaan-m/ECC) into .agents/skills/data-scraper-agent in your project. Codex loads it when a task matches its description.

Can I use Data Scraper Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill data-scraper-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-scraper-agent, .gemini/skills/data-scraper-agent, .github/skills/data-scraper-agent and .opencode/skills/data-scraper-agent in your project.

What does Data Scraper Agent need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Scraper Agent is instructions for the agent only. Our summary lists: Python 3.

Does Data Scraper Agent access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Scraper Agent safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Scraper Agent use?

Data Scraper Agent is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Scraper Agent use?

About 377 tokens (SKILL.md is roughly 1.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Scraper Agent?

Skills that share tags, products or a category with Data Scraper Agent: Data Scraper Agent (majiayu000/claude-skill-registry, 666 stars), Apify CI Integration (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Supabase CI Integration (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Nushell (ccusage/ccusage, 19k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Scraper Agent?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 275,023 GitHub stars. The repository holds 645 skills in this directory. The repository was last updated on October 5, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.