Scraper Builder
jwynia/agent-skills
Guide AI agents to generate complete PageObject pattern web scraper projects using Playwright and TypeScript with Docker deployment.
URL과 수집 항목을 받아 사이트를 정찰하고 데이터를 수집하여 엑셀로 출력하는 범용 웹 크롤링 에이전트. An agent skill from byungjunjang/web-crawler.
$ npx skills add byungjunjang/web-crawler --skill web-crawler -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install byungjunjang/web-crawler web-crawler --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/byungjunjang/web-crawler.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/web-crawler .claude/skills/web-crawler && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "web-crawler" agent skill from https://github.com/byungjunjang/web-crawler/tree/master/.codex/skills/web-crawler into .claude/skills/web-crawler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-crawler", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/byungjunjang/web-crawler/tree/master/.codex/skills/web-crawlerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add byungjunjang/web-crawler --skill web-crawler -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install byungjunjang/web-crawler web-crawler --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/byungjunjang/web-crawler.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.codex/skills/web-crawler .agents/skills/web-crawler && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "web-crawler" agent skill from https://github.com/byungjunjang/web-crawler/tree/master/.codex/skills/web-crawler into .agents/skills/web-crawler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-crawler", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add byungjunjang/web-crawler --skill web-crawler -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install byungjunjang/web-crawler web-crawler --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/byungjunjang/web-crawler.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.codex/skills/web-crawler .cursor/skills/web-crawler && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "web-crawler" agent skill from https://github.com/byungjunjang/web-crawler/tree/master/.codex/skills/web-crawler into .cursor/skills/web-crawler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-crawler", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/byungjunjang/web-crawler.git --path .codex/skills/web-crawler--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add byungjunjang/web-crawler --skill web-crawler -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install byungjunjang/web-crawler web-crawler --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/byungjunjang/web-crawler.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.codex/skills/web-crawler .gemini/skills/web-crawler && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "web-crawler" agent skill from https://github.com/byungjunjang/web-crawler/tree/master/.codex/skills/web-crawler into .gemini/skills/web-crawler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-crawler", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install byungjunjang/web-crawler web-crawlerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add byungjunjang/web-crawler --skill web-crawler -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/byungjunjang/web-crawler.git skills-src && mkdir -p .github/skills && cp -r skills-src/.codex/skills/web-crawler .github/skills/web-crawler && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "web-crawler" agent skill from https://github.com/byungjunjang/web-crawler/tree/master/.codex/skills/web-crawler into .github/skills/web-crawler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-crawler", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add byungjunjang/web-crawler --skill web-crawler -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install byungjunjang/web-crawler web-crawler --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/byungjunjang/web-crawler.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.codex/skills/web-crawler .opencode/skills/web-crawler && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "web-crawler" agent skill from https://github.com/byungjunjang/web-crawler/tree/master/.codex/skills/web-crawler into .opencode/skills/web-crawler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-crawler", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
web-crawlerURL과 수집 항목을 받아 사이트를 정찰하고 데이터를 수집하여 엑셀로 출력하는 범용 웹 크롤링 에이전트. An agent skill from byungjunjang/web-crawler.
Web Crawler is an agent skill from byungjunjang/web-crawler. URL과 수집 항목을 받아 사이트를 정찰하고 데이터를 수집하여 엑셀로 출력하는 범용 웹 크롤링 에이전트. 사용자가 URL과 함께 데이터 수집/크롤링/스크래핑을 요청하거나, 웹사이트에서 정보를 추출하고 싶다고 할 때 반드시 이 스킬을 사용한다. "이 사이트에서 ~를 모아줘", "~를 크롤링해줘", "입찰공고를 수집해줘" 등의 요청에도 트리거된다.
Its SKILL.md is about 7.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/antibot-strategies.md`, `references/crawl-operations.md` and `references/fetcher-patterns.md`).
It sits in Data & Analytics, covering Web scraping and Browser automation. It works with Playwright. The repository describes itself as: 범용 웹 크롤링 에이전트 — URL과 수집 항목으로 사이트를 정찰·대량수집·엑셀 출력. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ced95d5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonyt-dlpcodexFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
r.jina.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Web Crawler loads about 7.2k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 51 tokens; SKILL.md has 3,575 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from byungjunjang/web-crawler at commit ced95d5, republished under its MIT licence (© byungjunjang). 3,575 words, ~7,161 tokens.
.claude/skills/web-crawler/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.requests, urllib, httpx, BeautifulSoup로 직접 수집하지 않는다 — 사다리 A 래퍼(plain_get/plain_session)만 지문·Referer를 통제하고 소프트블록 검사와 이어지기 때문이다.page.evaluate()로 데이터를 추출하거나 DOM을 파싱하여 수집하는 것은 금지. agent-browser는 구조 파악, 스크린샷, 네트워크 감시에만 사용한다.규칙 1 예외 — Playwright 직접 사용이 허용되는 경우: SPA 세션 보호 사이트(WebSquare, 내부 API 403 등)에서는 crawl_script.py 안에서 Playwright를 직접 사용하여 SPA를 로드하고,
page.on("response")로 XHR 응답을 인터셉트하여 데이터를 수집할 수 있다. 이 경우에도 agent-browser가 아닌 Playwright sync_api를 사용한다.
Step 1: 입력 파싱 → Step 1-A: 도메인 프로필 확인 → Step 1-B: Phase 0 공인 우회로 체크
↓ (재사용 Yes → Step 3) (해결되면 정찰 스킵 → Step 4)
Step 2: 정찰 (agent-browser) → Step 2-A: 인증 처리 (필요 시)
↓
Step 3: 사이트 분류 & 수집 전략 결정 ← 핵심 의사결정
↓
Step 4: 수집 코드 생성 & 실행
↓
Step 5: 데이터 검증 (Step 5.0 소프트블록 게이트 최우선)
↓
Step 5-A: 도메인 프로필 저장 (필수 게이트, 누락 시 파이프라인 미완료)
↓
Step 6: 엑셀 출력 & 보고사용자 메시지에서 추출:
불명확하면 되묻기. 최소 요건: URL 1개 + 수집 항목 1개.
프로필이 있으면 notes·fetcher_type·antibot_strategy를 그대로 채택하고 Phase 0(1-B)과 정찰을 건너뛰어 Step 3으로 간다. last_used가 3개월 넘었거나 사용자가 '최신 구조로'를 요청했으면 정찰을 추가한다. consent 기록이 없는 사다리 B 프로필이면 Step 3의 이음매 통지를 거친다.
from domain_profile import DomainProfile
profile_mgr = DomainProfile()
if profile_mgr.exists(domain):
profile = profile_mgr.load(domain)
# 묻지 않고 그대로 채택: Phase 0(1-B) 건너뛰고 바로 Step 3 (검증된 레시피 보유)
# last_used 가 3개월 초과이거나 사용자가 '최신 구조로' 를 요청했으면 정찰을 추가한다.
# consent 기록이 없는 사다리 B 프로필이면 이번이 최초 통과다 — Step 3 의 이음매 통지를 거친다.
# 프로필 없음: Step 1-B(Phase 0)부터 진행수집 성공 후 프로필 저장:
profile_mgr.save(domain, {
"domain": domain,
"capability": "<static|js_render|api|session>", # ★ SSOT — 능력 수준. 비워 두면 save() 가 fetcher_type 에서 채운다
"fetcher_type": "<Fetcher|FetcherSession|DynamicFetcher|DynamicSession|Spider|playwright_spa_intercept|curl_cffi_grid|StealthyFetcher|chrome_cdp|API_SESSION|yt-dlp|RSS|oEmbed|Jina>",
"antibot_strategy": "<none|playwright_intercept|impersonate|curl_cffi_grid|stealthy|chrome_cdp|naver_antibot|authenticated_browser>",
"site_type": "<static|csr|api|spa_session|akamai>",
# robots·ToS 사유로 배포에서 빼야 하면 명시한다 — 사다리 A 여도 이 선언은 무조건 인정된다
"distribution": "local", "distribution_reason": "<한 줄>",
"selectors": {<MAPPING>},
"pagination": {<CONFIG>},
"api_endpoints": [<LIST>],
"notes": "<특이사항>",
# 사다리 B(4단 이상)로 수집했을 때만. **실제로 통지했고 사용자가 '진행' 을 고른 경우에만 적는다** —
# 그 일이 없었으면 이 블록을 적지 않는다. 적으면 기록이 거짓이 되고, 이 기록의 유일한 쓸모가 사라진다.
# 근거가 아니라 선택을 적는다. 이미 `consent` 기록이 있는 프로필이면 자동으로 이어지므로(sticky) 생략해도 된다.
"consent": {"notified_at": "<통지한 실제 시각 ISO8601>", "choice": "proceed"},
})profile.json이 없거나(신규 도메인) 재사용 안 할 경우, 정찰(Step 2)로 가기 전에 그 플랫폼에 공개 API/피드/oEmbed가 있는지 먼저 본다. (insane-search Phase 0 차용) 있으면 HTML 정찰·긁기보다 빠르고 안정적이며 차단도 거의 없다 → 정찰 스킵하고 바로 수집.
| 플랫폼 유형 | 공인 우회로 | 비고 |
|---|---|---|
| YouTube/TikTok/SoundCloud 등 미디어 | yt-dlp --dump-json <URL> | 메타/자막, 1,800+ 사이트 |
| arXiv / Wikipedia / GitHub / CrossRef | Atom/REST 공개 API | 인증 불필요 |
| X(트위터) 단일 트윗 | cdn.syndication.twimg.com oEmbed | |
서브레딧/스레드 .rss (Atom) | ||
| Hacker News | Firebase JSON API | |
| 일반 사이트 (SPA 렌더·단발 본문 추출) | https://r.jina.ai/<URL> | 정찰·단발 본문용, 대량 수집엔 부적합 |
비공식 내부 엔드포인트는 Phase 0 이 아니다. 사이트가 문서화하지 않은 JSON 엔드포인트는 공인 경로가 아니라 사다리 2단(숨은 API) 이다 — 정찰로 찾아내는 것이고, 약관이 그 사용을 금지하는지는 별도로 확인해야 한다. 이 표는 제공자가 명시적으로 여는 경로만 담는다.
판정: Phase 0로 데이터가 충분히 나오면 → 정찰 스킵, Step 3 분류 트리도 건너뛰고 수집(Step 4)으로. 안 되면 → Step 2 정찰로 정상 진행.
정찰 요청을 보내기 전에 한 번 확인한다. 산문 지시가 아니라 실제 호출이다:
from utils import check_robots
verdict = check_robots(target_url)
if verdict["error"]:
# 가져오지 못한 것과 허용된 것은 다르다 — 사용자에게 '확인 못 함' 으로 알린다
...
elif not verdict["allowed"]:
# 차단 — 진행 여부를 사용자에게 묻는다. 임의로 진행하지 않는다
...
if verdict["crawl_delay"]:
limiter = RateLimiter(delay=max(verdict["crawl_delay"], 1.0))crawl_delay 가 있으면 RateLimiter 기본값보다 우선한다.error 가 있으면 "허용됨" 이 아니라 "확인 못 함" 이다 — 사용자에게 그대로 알린다.agent-browser는 이 프로젝트의 표준 정찰 도구다 (선택 아님). 단순 정적 사이트를 긁더라도 정찰 단계에서는 agent-browser를 먼저 사용한다.
시작 전 (양 host 공통): 우선
agent-browser skills get core --full을 실행해 agent-browser 사용법(snapshot-and-ref 워크플로우, 네트워크 캡처 등)을 로드한다. 설치된 구버전이Unknown command: skills를 반환하면agent-browser --help에서 snapshot/network 명령을 로드한다. 이 오류만으로 agent-browser 자체가 불능이라고 판정하거나 폴백으로 내려가지 않는다.CLI가 없거나 브라우저 실행이 막힌 제한 환경에서만 아래 정찰 폴백 티어로 내려간다. 환경 셋업·검증은
scripts/setup.ps1(Windows) 또는scripts/bootstrap.py+scripts/preflight.py.
티어 도구 host 표준 agent-browser양 host 공통 폴백 1 (Claude) Claude in Chrome ( mcp__claude-in-chrome__*)Claude Code / Cowork 전용 폴백 1 (Codex) ChatGPT Chrome 플러그인 Browser Use ( chrome:control-chrome)Codex 전용 (Chrome 확장 연결 시) 폴백 2 (공통) Scrapling DynamicFetcher/ Playwrightsync_api양 host 공통 host별 분기: Claude Code/Cowork는
agent-browser → Claude in Chrome → 폴백 2, Codex는agent-browser → ChatGPT Chrome Browser Use(연결 시) → 폴백 2다. Codex에chrome:control-chrome스킬이 없거나 ChatGPT Chrome 확장이 연결되지 않으면 Codex 폴백 1을 건너뛴다. Codex에서 Claude in Chrome을 찾지 않는다.폴백을 썼으면 어느 티어였는지 Step 5-A의 profile.json
notes에 남긴다.
agent-browser eval -b <base64> 가 페이로드를 UTF-8 이 아닌
인코딩으로 디코딩해 한글이 깨진다 — /입찰|공고/ 가 /?낆같|怨듦퀬/ 로 들어가
SyntaxError: Invalid regular expression 이 난다. 한글이 필요하면
'\uC785\uCC30'(= 입찰) 처럼 유니코드 이스케이프로 적는다. (파일을 만들 때도 Get-Content -Raw 는 UTF-8 을
ANSI 로 읽으므로 애초에 ASCII 로 생성하는 편이 안전하다.)agent-browser에서 네트워크 요청을 캡처하여 API를 식별한다.
API 식별 2단계:
/api/, /graphql/, /v1/ 포함, application/json 응답, 광고/분석 제외SPA 세션 보호 감지 (중요):
정찰 중 다음을 확인하면 "SPA 세션 보호 사이트"로 분류:
agent-browser를 못 쓸 때 Claude 계열 host에서 쓴다. 사용자의 실제 Chrome을 조종하므로 실제 쿠키·실제 IP가 그대로 붙는 것이 장점이다. 정찰 항목 5개는 그대로 대체된다.
| 정찰 항목 | 도구 |
|---|---|
| ① 스냅샷·DOM 구조 | read_page (a11y tree, filter:"interactive" / depth / ref_id로 축소) + computer 스크린샷 |
| ② 로딩 방식 판단 | javascript_tool — #__next/#root/[data-reactroot] 유무, 초기 DOM에 아이템이 있는지 |
| ③ CSS 셀렉터 | javascript_tool — 반복되는 tag.class 조합을 집계해 item_root 후보 도출 |
| ④ pagination | javascript_tool — li.next a/pager 텍스트/URL 패턴 |
| ⑤ 총 건수 | javascript_tool |
네트워크 감시는 제약이 있다 (실측). 아래 절차를 그대로 따르지 않으면 API를 못 찾는다.
read_network_requests는 처음 호출한 시점부터 추적을 시작한다. 페이지를 연 직후 한 번 호출해 무장한다.read_network_requests를 읽는다.javascript_tool에서 페이지 컨텍스트로 재호출한다:// 페이지 컨텍스트라 세션 쿠키·헤더가 그대로 탑승한다 (정찰용 1회 호출)
await fetch('<후보 API URL>', {cache:'no-store'}).then(r => r.json())fetch·XMLHttpRequest 둘 다 캡처 대상임은 확인됐다. 광고·분석 픽셀(criteo/adnxs/facebook)과 RUM(Datadog browser-intake-datadoghq.com/api/v2/rum)이 진짜 API보다 훨씬 많으므로, urlPattern을 /api/가 아니라 1st-party 도메인으로 걸어 읽는다.값 마스킹에 주의한다 (실측). Claude in Chrome은 자격증명처럼 보이는 값을 자동으로 가린다 — [BLOCKED: JWT token], [BLOCKED: Cookie/query string data], [BLOCKED: Base64 encoded data]. 원시 덩어리를 통째로 뽑으면 정작 필요한 부분이 가려진다:
el.outerHTML 통째 덤프, JSON.parse(...) 결과 전체 반환, 쿼리스트링 포함 URL 그대로 반환Object.keys(obj), el.className, el.getAttribute('href'), locator.count()Object.keys(j), Object.keys(j.data[0])즉 마스킹은 정찰을 막지 않는다. 정찰에 필요한 건 값이 아니라 구조이므로, 처음부터 키·개수·셀렉터만 뽑으면 걸리지 않는다.
여기서도 수집은 금지다.
javascript_tool은 구조 파악과 API 후보 1회 검증에만 쓴다. 페이지 안에서 루프 돌려 전량 추출하는 것은 절대 규칙 2 위반 — 수집은crawl_script.py로 한다.원격 전용 환경(Cowork)에서는 정찰까지만 가능하다. 샌드박스 egress가 기본 "package managers only"라 대상 사이트 직접 접속이 막히고, 통과시켜도 데이터센터 IP라 안티봇 프로필이 재현되지 않으며, VM에서 호스트 Chrome의 CDP 포트(9222)에 붙을 수 없어 Akamai 대응이 불가능하다. 원격에서는 정찰 → profile.json 갱신까지 하고, 수집은 로컬에서 이어서 실행한다.
agent-browser를 못 쓰고 현재 Codex 세션에 chrome:control-chrome 스킬이 있으며 사용자의 ChatGPT Chrome 확장이 연결돼 있을 때만 쓴다. 스킬을 먼저 읽고 그 Bootstrap·Chrome 선택 절차를 그대로 따른다. 반드시 agent.browsers.get("chrome")으로 Chrome을 명시 선택하고, 연결 후 chrome.nameSession(...)을 호출한 다음 탭을 생성하거나 사용자가 지정한 탭을 claim한다. 다른 브라우저 surface로 자동 대체하지 않는다.
이 경로는 사용자의 실제 Chrome을 조종하므로 기존 브라우저 상태·로그인 세션·실제 IP가 적용되는 것이 장점이다. 단 브라우저 쿠키·localStorage·프로필·비밀번호를 직접 조회하지 않는다.
| 정찰 항목 | Browser Use API |
|---|---|
| ① 스냅샷·DOM 구조 | tab.playwright.domSnapshot() + 필요 시 tab.screenshot() |
| ② 로딩 방식 판단 | tab.playwright.evaluate()의 read-only page scope에서 #__next/#root/[data-reactroot]·초기 아이템 유무 확인 |
| ③ CSS 셀렉터 | tab.playwright.locator()의 count() 또는 read-only evaluate()로 반복되는 tag.class 조합과 매칭 개수만 집계 |
| ④ pagination | locator()/evaluate()로 next/pager URL·텍스트 확인 후 expectNavigation() + click()으로 1회 검증 |
| ⑤ 총 건수 | read-only evaluate()로 표시 총계 또는 총 페이지 × 페이지당 건수 확인 |
네트워크 감시는 제한적이다 (실측).
tab.capabilities.list()에 pageAssets가 있으면 해당 capability의 documentation()을 먼저 읽고 pageAssets.list()로 현재 페이지에서 관찰된 script/image/stylesheet/font URL을 확인한다.evaluate() scope에서도 window.performance/document.defaultView.performance를 사용할 수 없었다./api/·/graphql/·/v1/ 후보나 JSON 응답 필드 매핑이 필요한 사이트는 네트워크 감시 부분에 한해 폴백 2의 Playwright sync_api(page.on("response"))를 병행한다. profile.json notes에는 Codex 폴백 1 Chrome Browser Use + 폴백 2 network 보조처럼 둘 다 기록한다.여기서도 수집은 금지다. Browser Use의
evaluate()/locator는 구조·셀렉터·페이지네이션·건수 판정에만 쓴다. DOM을 루프로 전량 추출하지 않는다. 수집은 반드시crawl_script.py안의 Scrapling 또는 Playwright로 한다.
로그인이 필요한 경우:
agent-browser close --all. 정찰로 이미 데몬이 떠 있으면 이후 호출의
--headed·--profile·--session 이 경고 한 줄만 남기고 무시된다
(⚠ --profile ignored: daemon already running). 창이 안 뜬 채로 "로그인해 주세요" 라고
말하게 되는 실패가 여기서 나온다.agent-browser --headed --profile <경로> open <로그인URL>.
사용자의 평소 Chrome 프로필을 붙이지 않는다(다른 사이트 세션까지 딸려온다).
경로를 주면 로그인이 그 디렉터리에 남아 다음 호출에서도 유지된다.NID_AUT/NID_SES, 인스타그램 sessionid/ds_user_id)agent-browser state save <스크래치경로> 로 받은 뒤 대상 도메인분만 골라
output/<도메인>/cookies.json 에 넣는다 — 프로필 전체 쿠키를 프로젝트 폴더로 옮기지 않는다.session.get(url, cookies=jar).
session.cookies.update(jar) 는 동작하지 않는다(_SyncSessionLogic 에 .cookies 없음).import json
from utils import plain_session
# 1. agent-browser로 수동 로그인 후 쿠키 추출 → output/<도메인>/cookies.json 저장
# ★ 수동 로그인 전에 `agent-browser close --all` — 데몬이 떠 있으면 --headed 가 무시된다
with open("output/<도메인>/cookies.json") as f:
cookies = json.load(f) # {"name": "value", ...}
# 2. 사다리 2단 세션에 주입 (위장 인자 없음 — 로그인 쿠키는 위장이 아니다)
# ★ 쿠키는 **요청별 인자**로 넘긴다. `session.cookies.update(...)` 는 동작하지 않는다 —
# plain_session() 이 돌려주는 _SyncSessionLogic 에는 .cookies 속성이 없다.
with plain_session() as session:
resp = session.get(url, cookies=cookies)쿠키 파일은 .gitignore의 **/cookies*.json 패턴으로 자동 차단된다. 전용 프로필 창에서
받은 상태를 저장했다면 대상 도메인분만 골라 넣는다 — 다른 사이트 세션까지 프로젝트
폴더로 들어오지 않게.
정찰 결과로 사이트를 분류하고 수집 전략을 고른다. 전체 워크플로우에서 가장 중요한 결정이다.
Phase 0 선행 확인됨 가정. 여기 오기 전 Step 1-B에서 공인 API/피드(yt-dlp·RSS·oEmbed·Jina)를 이미 확인했다. Phase 0로 해결됐으면 이 트리를 건너뛴다.
📖 도구 역할 분리 원칙과 Fetcher 선택 보충(래퍼, 프로필의
antibot_type)은references/crawl-operations.md를 읽는다.
사다리 B "상대가 나를 막고 있다" ← 돌파. 통지 후 진행
────────────────────────────────── ← 이음매: 성격이 바뀌는 지점
사다리 A "데이터가 어디 있나" ← 탐색. 자동1~3단에서 사이트는 나를 막은 적이 없다. 데이터가 있는 위치가 다를 뿐이다. 4단부터가 처음으로 "상대가 나를 식별하고 거절한" 상황이다. 통지 게이트가 정확히 이 이음매에 놓이는 것은 우연이 아니다.
| 칸 | 비유 | 판별 | 도구 | 페이지당 요청 |
|---|---|---|---|---|
| 0 공식 API·공개데이터 | 그냥 주는 것 | 개발자 문서 / data.go.kr | plain_session | 1 |
| 1 정적 HTML | 종이에 글자가 이미 있다 | Ctrl+U 소스에 보임 | plain_get | 1 |
| 2 숨은 API | 종이엔 없고 전화번호가 적혀 있다 | Network 탭 XHR 응답에 데이터 | plain_session | 1 |
| 3 렌더링 | 전화를 브라우저만 걸 수 있다 | 내부 주소 직접 호출 시 토큰·서명 부족 | DynamicFetcher | 수십~수백 |
page.on("response")). 브라우저는 띄우되 데이터는 JSON 으로 받으므로 정확하고 구조 변경에 강하다.plain_get/plain_session 이 impersonate·stealthy_headers 두 인자를 함께 꺼서 이 경계를 실제로 성립시킨다. 하나만 끄면 불일치 지문이 되어 오히려 악화된다.사다리 A 를 소진했고 다음이 사다리 B 라면, 자동으로 넘어가지 않는다. 사용자에게 한 번 알린다:
이 사이트는 자동 접근을 차단하고 있습니다 (<감지된 유형>).
다음 단계는 그 차단을 우회하는 것입니다.
· 수집 권한이나 정당한 사유가 있는지 확인하세요
· 참고: 공식 API·데이터 개방·제휴 경로나 대체 데이터원이 있으면 그쪽이 낫습니다
계속하시겠습니까? [진행 / 중단]consent 기록이다. 한 번 통과한 뒤 B 안에서 4→5→6 으로 더 올라가는 것은 다시 묻지 않는다(이음매는 한 곳이고 이미 넘었다).consent 블록에 남긴다 (Step 5-A). 이게 없으면 프로필 저장이 거부된다.consent 기록이 있는 프로필이면 통지하지 않는다. 그 기록 자체가 이 사용자가 이 도메인에서 통지받고 진행을 골랐다는 증거다(sticky). 프로필이 있어도 consent가 없다면(예: 사다리 A로만 수집돼 오다가 이번에 처음 사이트가 막은 경우) 이번이 최초로 이음매를 넘는 것이므로 그대로 통지한다.save() 가 consent 를 지운다 — 사용자가 언제 무엇을 통지받았는지를 배포되는 프로필에 실어 내보낼 수 없기 때문이다. 그래서 사이트가 나중에 새로 막기 시작하면 들고 있는 기록이 없고, 그대로 다시 통지한다. 그게 맞다 — 사이트가 새 보호를 건 것은 달라진 상황이고, 이음매를 다시 건너는 것은 새로운 사건이다.| 칸 | 무엇이 걸렸나 | 무엇을 하나 | 도구 |
|---|---|---|---|
| 4 지문 정렬 | 목소리 — "크롬입니다" 라고 말했는데 TLS 협상 지문이 파이썬 | 협상 지문을 실제 크롬과 동일하게. 브라우저 안 띄움 | curl_cffi 그리드 |
| 5 스텔스 브라우저 | 걸음걸이 — navigator.webdriver, 폰트·캔버스 지문, 직선 마우스 | 브라우저를 띄우되 자동화 흔적을 지움 | StealthyFetcher |
| 6 실제 크롬 | 위 전부가 안 통함 | 흉내가 아니라 실제 사용자 프로필 Chrome 을 띄우고 그 안에서 fetch | chrome_cdp |
B 는 순차가 아니다. 고급 WAF 는 4·5 단이 원리적으로 통하지 않아 바로 6 단으로 간다. WAF capability 라우팅은
references/antibot-strategies.md참조. 단 그 라우팅도 통지 이후에 일어난다.
통지에서 '진행' 을 받은 뒤 CAPTCHA 를 만나면 중단하지 않고 아래 순서로 이어간다. 이음매 안쪽의 일이라 CAPTCHA 때문에 다시 묻지 않는다.
StealthyFetcher().fetch(url, solve_cloudflare=True) 로 통과시킨다. 허용된 경로다 (references/antibot-strategies.md § Cloudflare).launch_chrome_cdp(url=<막힌 URL>, user_data_dir=".tmp/cdp_profile/<도메인>") 로 전용 프로필 Chrome 을 띄우고
"열린 창에서 CAPTCHA 를 풀고 알려 주세요" 라고 안내한 뒤 기다린다. 사용자가 끝났다고 하면
get_playwright_cdp_connection() 으로 그 브라우저에 붙어 같은 세션으로 수집을 이어간다.
통과 쿠키(cf_clearance 등)는 그 브라우저의 지문·IP 에 묶여 있어 plain_session 으로 옮기면 다시 막히기 쉽다.
수집 중 CAPTCHA 가 다시 뜨면 중간 저장 후 같은 창에서 다시 풀게 하고 이어간다.
Chrome CDP 를 띄울 수 없는 환경이면 Step 2-A 의 agent-browser --headed --profile <경로> 창으로 같은 핸드오프를 한다.프로필에는 실제로 쓴 값을 적는다: 1번으로 끝났으면 antibot_strategy: stealthy, 2번으로 이어갔으면 chrome_cdp 와 notes 에 "CAPTCHA 수동 풀이 필요". 코드는 references/antibot-strategies.md § Cloudflare → StealthyFetcher.
| 증상 | 무슨 뜻 | 칸 |
|---|---|---|
Ctrl+U 소스에 글자 있음 | 종이에 적혀 있다 | 1 |
| 소스엔 없는데 Network 에 JSON | 전화번호가 있다 | 2 |
| 내부 주소 호출 시 401·토큰 필요 | 전화는 브라우저만 건다 | 3 |
| 헤더 다 맞췄는데 403 | 지문에서 걸렸다 | 4 ⚠ |
| 챌린지 페이지 / "봇 감지" | 행동에서 걸렸다 | 5 ⚠ |
_abck 쿠키 · Access Denied | 고급 WAF | 6 ⚠ |
⚠ = 통지 대상. 이 표의 위 셋과 아래 셋은 성격이 다르다 — 위 셋은 내가 안 갖춘 것(고쳐라), 아래 셋은 상대가 안 주는 것(멈추고 물어라).
올리려면: ① 아래 칸이 실패했다는 '확인' (추측 아님)
② 4단 이상이면 사용자의 한 번의 진행 선택 (이미 `consent` 기록이 있는 프로필이면 면제)
내려오기: 6단으로 성공했어도 영구 자격이 아니다.
사이트 구조가 바뀌면 다시 1단부터 판별한다.📖 상세 코드 템플릿은
references/fetcher-patterns.md를 참조한다.
| 칸 | 전략 | 도구 | 참조 섹션 |
|---|---|---|---|
| 0·2 | API 직접 | plain_session | fetcher-patterns.md § API 수집 |
| 1 | 정적 HTML | plain_get | fetcher-patterns.md § 정적 HTML |
| 3 | JS 렌더링 | DynamicFetcher | fetcher-patterns.md § 동적 사이트 |
| 3 | SPA 세션 인터셉트 | Playwright 인터셉트 | antibot-strategies.md § SPA 세션 |
| 4 ⚠ | 경량 그리드 | curl_cffi 그리드 | antibot-strategies.md § curl_cffi 그리드 |
| 5 ⚠ | 스텔스 | StealthyFetcher | antibot-strategies.md § Cloudflare |
| 6 ⚠ | 실제 크롬 | Chrome CDP | antibot-strategies.md § 고급 WAF |
📖 상세 감지 로직은
references/antibot-strategies.md참조.
Akamai 계열: _abck/bm_sz/ak_bmsc 쿠키, Access Denied + errors.edgesuite.net
Cloudflare: cf_clearance 쿠키, 챌린지 페이지
SPA 세션 보호: 브라우저에서는 정상인데 API 직접 호출 시 403 (ErrorCode -801 등) — 이건 3단이지 우회 대상이 아니다
API 사용 시, 코드 생성 전에 샘플 5건으로 필드 매핑을 검증한다:
Step 3에서 결정한 전략에 맞는 코드 패턴을 references/fetcher-patterns.md에서 참조하여 crawl_script.py를 생성한다.
scripts/utils.py의 RateLimiter 사용raw_data.json 을 빈 배열로 밀어버리면 "이번 실패" 가 "지난 성공 소실" 이 된다. 쓰기 직전에 if not data: return 로 막거나, 기존 파일이 있으면 raw_data.json.bak 로 옮긴 뒤 쓴다except 블록은 그 페이지만 건너뛰고(
continue) 다음으로 간다. 루프를 끝내는 것은 필수 요소 2의 연속 실패 한도와 사다리 A 소진뿐이다.
📖 Fetcher 에스컬레이션 순서와 Rate Limiting 간격은
references/crawl-operations.md를 읽는다.
에이전트는 (profile.json + 정찰 결과 + 사용자 요청) 세 가지를 합성해 Python 수집 코드를 동적으로 생성한다.
selectors/api_endpoints/pagination이 있으면 그걸 기반으로 코드 골격을 짠다. 새로 정찰해서 코드를 처음부터 쓰지 않는다.output/<도메인>/ 의 이전 crawl_script.py 참조: profile.json에 안 박힌 미세 디테일(배치 사이즈, JS evaluate 패턴, 예외 처리)을 그대로 가져와 재사용. 단, raw_data.json은 PII 가능성 있으므로 구조만 확인하고 데이터는 읽지 않는다.scripts/utils.py를 import하여 RateLimiter, cookie 관리, 로깅 등 공통 기능 사용scripts/export_excel.py를 import하여 엑셀 출력scripts/chrome_cdp.py는 antibot_strategy: chrome_cdp(사다리 6단)로 기록된 도메인에서 사용 — 통지 게이트를 이미 넘은 경우references/output-layout.md 참조)storage_args={"storage_file": "./fingerprints/elements_storage.db"} 경로 사용last_used만 업데이트하지 말 것 (Step 5-A 게이트)python scripts/sync_domain_list.py 실행 — CLAUDE.md/README.md의 "알려진 도메인" 목록은 profile.json에서 생성된다. 손으로 고치지 말 것 (scripts/test_sync_domain_list.py가 어긋남을 잡는다)output/<도메인>/<주제_YYYYMMDD_HHMMSS>/
├── crawl_script.py # 생성된 수집 스크립트
├── raw_data.json # 원시 데이터
├── crawl_result.xlsx # 엑셀 결과
└── progress.json # 진행 상황📖 쿠키·도메인 프로필 위치와 폴더 이름 규칙은
references/output-layout.md를 읽는다.
# 첫 수집: 핑거프린트 저장
items = page.css("<SELECTOR>", auto_save=True,
storage_args={"storage_file": "./fingerprints/elements_storage.db"})
# 이후: 자가 치유
items = page.css("<SELECTOR>", adaptive=True, auto_save=True,
storage_args={"storage_file": "./fingerprints/elements_storage.db"})HTTP 200 = 성공이 아니라 "검증 시작"이다. Akamai/DataDome/PerimeterX는 200 OK로 가짜 챌린지 페이지나 빈 셸을 돌려준다. 0건이 아니라 "쓰레기 N건"으로 통과해 엑셀로 납품되는 사고를 막는 게 이 게이트의 목적이다. (insane-search R2 차용)
언제 실행하나: 첫 페이지 응답 직후(대량 루프 진입 전)와, Step 5 검증 시작 시 1회. 즉 수집 전·후 양쪽에서 건다.
from utils import detect_softblock
# 첫 페이지 본문 + status + 쿠키로 판별 (cookies는 session.cookies 등에서 dict로)
verdict = detect_softblock(
page.html_content, # 또는 resp.text
status=page.status,
cookies=dict(session.cookies) if hasattr(session, "cookies") else None,
selector_hit=bool(page.css("<ITEM_SELECTOR>")), # 핵심 콘텐츠 셀렉터 매칭 여부
)
if verdict["blocked"]:
logger.error(f"소프트블록 감지 — {verdict['verdict']}: {verdict['signals']}")
# 무엇이 감지됐든, 다음 단계가 사다리 B(4단 이상)라면 **먼저 Step 3 의 이음매 통지 게이트를 거친다.**
# 소프트블록 감지는 "상대가 나를 식별하고 거절했다" 는 신호다 — 즉 이음매에 도달했다는 뜻이지,
# 이음매를 건너뛰어도 된다는 뜻이 아니다.게이트 규칙:
blocked=True면 수집을 강행하지 않는다. 에스컬레이션 여부는 규칙 2(이음매 통지 게이트)를 따른다 — 게이트를 통과해 상위 단계로 가도 안 뚫리면 사용자에게 보고 후 중단.<감지된 유형> 에 넣는다.
사용자가 '진행' 을 고른 뒤에야 WAF capability 라우팅(4·5 를 건너뛸지 등)을 적용한다.weak_ok(셀렉터 미검증 통과)는 통과시키되, 수집 후 필드 채움률이 비정상적으로 낮으면 이 게이트를 의심한다.detect_pii(data) 실행95% 이상 유효 데이터면 통과. 미달 시 Step 4 재시도 (최대 2회).
📖 수집 실패 시 원인 진단은
references/troubleshooting.md를 참조한다.
📖 에러 유형별 대응(429·403·소프트블록·타임아웃·0건 등)은
references/crawl-operations.md§ 에러별 대응표를 읽는다.
건수와 null 비율만 보면 광고를 상품으로 가져온 경우를 통과시킨다. 자가치유 셀렉터는 "못 찾겠다" 고 말하지 않고 늘 무언가를 반환하기 때문이다.
from utils import validate_values
issues = validate_values(results, {
"상품명": {"type": "str", "required": True, "max_empty_ratio": 0.1},
"가격": {"type": "int", "required": True, "min": 1, "max": 100_000_000},
"카테고리": {"type": "str", "required": True, "allow_uniform": True}, # 한 페이지가 단일 카테고리일 수 있다 — 균일해도 정상
})
if issues:
logger.warning("값 검증 경고:\n" + "\n".join(issues))allow_uniform: True는 그 필드의 중복률 검사만 면제한다 — 카테고리·플래그·정액
배송비처럼 전부 같은 값이어도 정상인 필드에 쓴다. 명시적으로 걸지 않은 필드는 10건 이상
전부 동일하면 그대로 경고한다(셀렉터가 광고·머리글 등 고정 요소를 잡았을 가능성).adaptive=True 로 요소를 재탐색한 행에는 플래그 컬럼을 남겨 엑셀에서 구분되게 한다.검증을 통과한 직후, 반드시 fingerprints/<도메인>/profile.json을 저장하거나 갱신한다. 이걸 빼먹으면 다음 수집 시 정찰부터 다시 해야 하고, 다른 머신/세션에서는 노하우가 완전히 사라진다.
from domain_profile import DomainProfile
from datetime import date
profile_mgr = DomainProfile() # base_dir=./fingerprints
profile_mgr.save(domain, {
"domain": domain,
"capability": "<static|js_render|api|session>", # ★ SSOT — 능력 수준. 비워 두면 save() 가 fetcher_type 에서 채운다
"fetcher_type": "<yt-dlp|RSS|oEmbed|Jina|Fetcher|FetcherSession|DynamicFetcher|DynamicSession|Spider|playwright_spa_intercept|curl_cffi_grid|StealthyFetcher|chrome_cdp|API_SESSION>", # 파생 — 현재 엔진에서의 구현체. 앞 4개는 Step 1-B Phase 0 공인 우회로
"antibot_type": "<none|cloudflare|akamai|spa_session|naver_antibot|other>",
"antibot_strategy": "<none|playwright_intercept|impersonate|curl_cffi_grid|stealthy|chrome_cdp|naver_antibot|authenticated_browser>", # 실제로 쓴 대응. 사다리 B 를 썼으면 반드시 그 값을 적는다
"site_type": "<static|csr|api|spa_session|akamai>",
# robots·ToS 사유로 배포에서 빼야 하면 명시 — 사다리 A 여도 이 선언은 무조건 인정된다
"distribution": "local", "distribution_reason": "<한 줄>",
"selectors": {<필드: 셀렉터>},
"pagination": {<config — type/param/limit 등>},
"api_endpoints": [{<url, method, params, field_mapping>}],
"notes": "<다음 사람이 정찰 없이 바로 수집할 수 있는 결정적 한두 줄>",
"last_used": str(date.today()),
# 사다리 B(4단 이상)로 수집했을 때만. **실제로 통지했고 사용자가 '진행' 을 고른 경우에만 적는다** —
# 그 일이 없었으면 이 블록을 적지 않는다. 적으면 기록이 거짓이 되고, 이 기록의 유일한 쓸모가 사라진다.
# 근거가 아니라 선택을 적는다. 이미 `consent` 기록이 있는 프로필이면 자동으로 이어지므로(sticky) 생략해도 된다.
"consent": {"notified_at": "<통지한 실제 시각 ISO8601>", "choice": "proceed"},
})
fetcher_type/antibot_strategy는 위 목록 안의 값으로 적는다. 문서에 없는 값을 지어내면 분류기가 사다리 칸을 판별하지 못해 저장이 거부되고, 사다리 B 로 수집해 놓고 A 쪽 값을 적으면 그 레시피가 배포 대상으로 잘못 분류된다.여기 목록이 곧 분류기가 아는 전부는 아니다. 실제 판정표는
scripts/profile_policy.py의TOOLS하나뿐이고, 위 목록은 그중 흔히 쓰는 값을 추린 것이다. 적을 값이 애매하면 지어내지 말고 그 표를 직접 본다 — 표에 없는 값은save()가ConsentRequired로 거부한다(심사가 아니라 "분류기가 알아들을 수 있는 값을 달라" 는 뜻이다).
robots.txt·ToS 사유로 배포에서 빼야 하는 프로필은distribution: "local"을 명시한다. 사다리 A(1~3단)라 자동으로는public이 되는 프로필이라도 이 선언은 무조건 인정된다(조이는 방향은 항상 통과). 사유는distribution_reason에 한 줄로 남긴다 — 예:cafe.naver.com,www.instagram.com.
새 도메인 프로필을 처음 만들었다면 저장 직후 목록을 재생성한다 (CLAUDE.md 의 "알려진 도메인" 블록은 생성물):
python scripts/sync_domain_list.py # CLAUDE.md / README.md 목록 재생성
python scripts/sync_domain_list.py --check # 어긋나면 exit 1📖
fingerprints/의.gitignorewhitelist 정책은references/output-layout.md를 읽는다.
notes 필드는 비워두지 않는다. 다음 사람(미래의 나 포함)이 정찰 안 하고도 바로 수집할 수 있는 한두 줄의 결정적 정보를 적는다 — "API key는 OK, job_group_id=518이 일반 목록", "리스트는 SSR HTML, 상세는 XHR JSON — 2단으로 충분", "review API는 POST에 originProductNo 필요" 같은 형식..gitignore가 cookies*.json/*auth*.json/*token*.json/*secret*은 차단하지만 profile.json은 commit 대상이므로 평문 자격증명이 새지 않게 분리한다.antibot_strategy 에 그 사실을 적는다. none 으로 적으면 분류기가 탐색 단계로 오판해 그 레시피를 배포 대상에 넣고 consent 기록도 지운다. 실제로 쓴 것을 적을 것.
Spider 는 티어가 아니라 래퍼다 — 밑에서 실제로 쓴 티어를 적는다.consent 없이는 저장이 거부된다(ConsentRequired). 심사가 아니라 기록이다. 인식되지 않는 fetcher_type/antibot_strategy 값도 같은 예외로 저장을 막는다 — 이때는 통지를 기록할 게 아니라 값을 문서화된 것으로 고쳐야 한다.last_used만 갱신하지 말고, 이번 수집에서 새로 알아낸 게 있으면 notes와 endpoint/selector를 누적/수정한다.저장이 끝나면 Step 6으로 진행. profile.json 저장 실패 시 수집 결과는 살아있어도 **"파이프라인 미완료"**로 보고하고 사용자에게 원인을 알린다 (디스크 권한, 스키마 누락 등).
from export_excel import export_to_excel
export_to_excel(data, filepath)fingerprints/<도메인>/profile.json 저장 여부 (신규 / 갱신 / 실패) — 실패면 사유 명시| 파일 | 내용 | 언제 참조 |
|---|---|---|
references/fetcher-patterns.md | 모든 수집 패턴의 코드 템플릿 | Step 4에서 crawl_script.py 생성 시 |
references/antibot-strategies.md | Akamai, SPA 세션, Cloudflare 대응 전략 | Step 3에서 안티봇 감지 시 |
references/troubleshooting.md | 실패 사례와 해결책 | 수집 실패 시 원인 진단 |
references/crawl-operations.md | 도구 역할 분리, Fetcher 선택 보충, 에스컬레이션, 에러별 대응표, Rate Limiting | Step 3 전략 결정, Step 4 코드 생성, 수집 실패 시 |
references/output-layout.md | 출력·저장 디렉터리 구조, .gitignore whitelist 정책 | Step 4 출력 폴더 생성, Step 5-A 프로필 저장 시 |
scripts/scrapling_reference.md | Scrapling API 레퍼런스 (Fetcher 종류·Selector·Spider·세션) | 라이브러리 사용법이 헷갈릴 때 |
© byungjunjang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in .codex/skills/web-crawler of byungjunjang/web-crawler.
Open the folder on GitHubat commit ced95d5
Web Crawler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Web Crawler this skillbyungjunjang/web-crawler | 165 | — | ~7.2k | Automated safety check: Pass | MIT | |
| Scraper Builderjwynia/agent-skills | 169 | — | ~4k | Automated safety check: Pass | MIT | |
| Camofox Browserredf0x1/camofox-browser | 412 | — | ~4.6k | Automated safety check: Pass | MIT | |
| Browser Automationalirezarezvani/claude-skills | 28k | — | ~3.4k | Automated safety check: Notes | MIT | |
| Using Webctloaustegard/claude-skills | 150 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Skyvern Browser AutomationSkyvern-AI/skyvern | 23k | 1 repos | ~2.9k | Automated safety check: Pass | AGPL-3.0 |
jwynia/agent-skills
Guide AI agents to generate complete PageObject pattern web scraper projects using Playwright and TypeScript with Docker deployment.
redf0x1/camofox-browser
Anti-detection browser automation for AI agents. An agent skill from redf0x1/camofox-browser.
alirezarezvani/claude-skills
A skill your agent uses when the user asks to automate browser tasks, scrape websites, fill forms, capture screenshots, extract structured data from web pages, or build web automation workflows.
oaustegard/claude-skills
Browser automation via webctl CLI in Claude.ai containers with authenticated proxy support.
Skyvern-AI/skyvern
Picks the right Skyvern CLI command for a web task, from quick yes/no checks to reusable multi-page workflows, instead of falling back to plain page fetching.
ykdojo/claude-code-tips
Fetches Reddit posts, threads and search results as JSON through a browser session, using a DuckDuckGo redirect to get past Reddit's automated-access block.
Works with
URL과 수집 항목을 받아 사이트를 정찰하고 데이터를 수집하여 엑셀로 출력하는 범용 웹 크롤링 에이전트. An agent skill from byungjunjang/web-crawler. Web Crawler is an agent skill from byungjunjang/web-crawler. URL과 수집 항목을 받아 사이트를 정찰하고 데이터를 수집하여 엑셀로 출력하는 범용 웹 크롤링 에이전트.
Web Crawler fits situations like: tasks that involve Web scraping; tasks that involve Browser automation.
Run `npx skills add byungjunjang/web-crawler --skill web-crawler -a claude-code`. Or copy the skill folder (.codex/skills/web-crawler in byungjunjang/web-crawler) into .claude/skills/web-crawler in your project. Claude Code loads it when a task matches its description.
Run `npx skills add byungjunjang/web-crawler --skill web-crawler -a codex`. Or copy the skill folder (.codex/skills/web-crawler in byungjunjang/web-crawler) into .agents/skills/web-crawler in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add byungjunjang/web-crawler --skill web-crawler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-crawler, .gemini/skills/web-crawler, .github/skills/web-crawler and .opencode/skills/web-crawler in your project.
Going by SKILL.md and its folder, Web Crawler needs the command-line tools its instructions call (python, yt-dlp and codex). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: r.jina.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Web Crawler is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Web Crawler: Scraper Builder (jwynia/agent-skills, 169 stars), Camofox Browser (redf0x1/camofox-browser, 412 stars), Browser Automation (alirezarezvani/claude-skills, 28k stars) and Using Webctl (oaustegard/claude-skills, 150 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
byungjunjang (a GitHub user) maintains it in byungjunjang/web-crawler, which has 165 GitHub stars. The repository was last updated on October 6, 2026.
Source: byungjunjang/web-crawler on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.