Agent skill

Xrk Crawl

by xrkseek in xrkseek/XRK-AGT

当你需要开发/排查 HTTP 抓取、SSRF、Playwright 受控浏览器、本地字体增强截图,或判断 webfetch 与 browser 工作流如何选型时使用。

MITAuto-check passedSecurity

Install Xrk Crawl

skills CLI
$ npx skills add xrkseek/XRK-AGT --skill xrk-crawl -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xrkseek/XRK-AGT xrk-crawl --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xrkseek/XRK-AGT.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/xrk-crawl .claude/skills/xrk-crawl && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
xrk-crawl
GitHub stars
138
Token cost
~1.3k tokens
SKILL.md length
244 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

当你需要开发/排查 HTTP 抓取、SSRF、Playwright 受控浏览器、本地字体增强截图,或判断 webfetch 与 browser 工作流如何选型时使用。

  • Tasks that involve Web application vulnerabilities
  • SKILL.md covers 统一入口(业务优先), 能力分层(何时用谁), Playwright 启动(必读) and SSRF, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Browser testing

What it does

Xrk Crawl is an agent skill from xrkseek/XRK-AGT. 当你需要开发/排查 HTTP 抓取、SSRF、Playwright 受控浏览器、本地字体增强截图,或判断 webfetch 与 browser 工作流如何选型时使用。

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Security, covering Web application vulnerabilities, Browser testing and Web scraping. It works with Playwright. The repository describes itself as: 这是基于多语言所搭建的基于web,工作流,新型算法组件的智能体. The licence is MIT.

When your agent uses it

  • Tasks that involve Web application vulnerabilities
  • Tasks that involve Browser testing
  • Tasks that involve Web scraping

Example prompts

  • “/xrk-crawl”

What it can do on your machine

Read from SKILL.md and the folder at commit 63f004e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are javascript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Xrk Crawl loads about 1.3k tokens when it runs. Until then it costs about 24 tokens; SKILL.md has 244 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~24
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from xrkseek/XRK-AGT at commit 63f004e, republished under its MIT licence (© xrkseek). 244 words, ~1,271 tokens.

Download SKILL.mdSave it as .claude/skills/xrk-crawl/SKILL.md (or your agent's skills folder).
name
xrk-crawl
description
当你需要开发/排查 HTTP 抓取、SSRF、Playwright 受控浏览器、本地字体增强截图,或判断 web_fetch 与 browser 工作流如何选型时使用。

统一入口(业务优先)

src/infrastructure/crawl/index.js — 插件、workflow、HTTP 只从这里 import(#infrastructure/crawl/index.js)。

javascript
import {
  fetchWithPolicy,
  runWebFetch,
  buildWebFetchRuntime,
  assertUrlSafeForFetch,
  PlaywrightAgentSession,
  launchOptionsFromBrowserRuntime,
  toPlaywrightAgentLaunchOptions,
  createLocalFontScreenshotHelper,
  DEFAULT_DEVICE_SCALE_FACTOR,
  DOM_TWEAK_LABEL_COLON_HALF,
} from '#infrastructure/crawl/index.js'

fetchWithPolicy 实现在 #utils/fetch-with-retry.js,由 crawl 门面 re-export。

能力分层(何时用谁)

场景能力实现文件(均在 src/infrastructure/crawl/)
简单 API / 无需 JS 渲染fetchWithPolicy#utils/fetch-with-retry.js
正文提取、Readability、FirecrawlrunWebFetchweb-fetch-executor.js
开放域检索runWebSearchweb-search-executor.js + web-search-registry.js
零配置免费检索runParallelFreeSearchweb-search-parallel-free.js + web-search-mcp-client.js
浏览器运行时buildBrowserRuntimecrawl-config.js(ai-workflow.crawl + renderer.playwright)
runtime → launchtoPlaywrightAgentLaunchOptions / launchOptionsFromBrowserRuntimecrawl-config.js
MCP web_fetch / web_searchworkflow/web.js内部用 crawl
JS 渲染、交互、截图PlaywrightAgentSessionplaywright-session.js
MCP 受控浏览器workflow/browser.jsbrowser_* 工具
截图字体/样式与线上一致createLocalFontScreenshotHelperpage-screenshot-enhance.js

选型:HTTP 能拿正文 → runWebFetch;要渲染或 PNG → Playwright。

Playwright 启动(必读)

业务/插件禁止手抄 executablePath / launchArgs / 只挑部分字段。统一:

javascript
// 推荐:配置链 + 业务覆盖(viewport / DPR)
await PlaywrightAgentSession.using(
  launchOptionsFromBrowserRuntime({ deviceScaleFactor: 3, viewport: { width: 520, height: 960 } }),
  async (session) => { /* ... */ }
)

// 已有 runtime 对象时
const rt = buildBrowserRuntime()
await PlaywrightAgentSession.launch(toPlaywrightAgentLaunchOptions(rt))

映射会带上:navigationTimeoutMs、ssrfPolicy、closeTimeoutMs、pageCrashRetries、opTimeoutMs、系统浏览器路径、launchArgs。

稳定性(底层默认,插件勿再造):

能力说明
using(..., { crashRetries: 1 })整轮 Target/Page crashed → 换新浏览器再跑
goto / gotoAndCapture页崩溃 → recreatePage 后重试(后者含重新导航)
session.close()softClosePlaywrightTree(AbortSignal.timeout 竞速),超时不卡死
waitUntil经 normalizePlaywrightWaitUntil:load(默认)/domcontentloaded/networkidle/commit;非法值回落 load。官方 DISCOURAGED networkidle,优先 load 或交互后断言就绪

反模式:同一重型(高 DPR + 字体)会话连开多档位 goto;在插件里硬编码 Chrome 路径 / soft-close / 崩溃正则;默认依赖 networkidle。

SSRF

  • ssrf-policy.js:allowlist、legacy IP、DNS pinning、createPinnedDispatcher
  • ssrf-guard.js:对外 re-export assertUrlSafeForFetch
  • fetch-guard.js:fetchWithSsrFGuard(每跳 pin DNS + 重定向环)
  • browser-navigation-guard.js:gotoWithNavigationGuard(page.route 拦截)

PlaywrightAgentSession 常用 API

方法说明
roleSnapshot()ARIA ref 树 + storeRoleRefsOnPage
runAct({ kind, ref, ... })含 batch、scrollIntoView、fill fields
listTabs / newTab / closeTab / focusTab多标签
getConsoleMessages / getNetworkRequests页面观测
armDialog / respondDialog弹窗
goto(url)gotoWithNavigationGuard + 交互后 SSRF 复检

目录结构

src/infrastructure/crawl/
  index.js
  ssrf-*.js / fetch-guard.js / browser-navigation-guard.js
  playwright-session.js / pw-*.js / act-policy.js
  web-fetch-*.js / web-search-*.js / crawl-config.js
  page-screenshot-enhance.js

配置(commonconfig)

Schema:core/system-Core/commonconfig/system.js → ai-workflow.fields.crawl
默认模板:config/default_config/ai-workflow.yaml → crawl:
运行时数据:data/server_bots/{port}/ai-workflow.yaml → crawl.webFetch / crawl.webSearch / crawl.browser

优先级:调用方 overrides > ai-workflow.crawl > renderer.playwright(browser 启动)> 默认

单一实现:crawl-config.js — resolveWebFetchRuntime / resolveWebSearchConfig / buildBrowserRuntime
禁止在 crawl 模块内读取 process.env 做业务配置;凭据与参数一律写 data/server_bots/{port}/ai-workflow.yaml → crawl.*。

与工作流

  • core/system-Core/workflow/web.js → web_fetch / web_search
  • core/system-Core/workflow/browser.js → 全量 browser_*(见 docs/system-core.md)

常见陷阱

  • 扩展写在 src/infrastructure/crawl/ 内并在 index.js 导出,不要在 Core 内复制一份。
  • Core 业务只 import 门面,勿在 Core 写 SSRF/搜索驱动。
  • 勿 launch({ browserType: rt.browserType, ... }) 手抄字段——会丢掉 nav 超时与 SSRF。
  • 探活用轻量 launchOptionsFromBrowserRuntime();高清截图另开会话并尽量 一次 goto。

Node 26

  • fetch-with-retry.js、web-fetch-executor.js 已用全局 fetch + AbortSignal.timeout;扩展时沿用,禁止 node-fetch。
  • 正文/截图二进制:toBase64() / Uint8Array.fromBase64(),勿 toString('base64')。
  • catch 与 SSRF 错误:Error.isError / normalizeError(skill xrk-node-runtime)。

参考

  • docs/system-core.md(web / browser 章节)
  • 本地 vendor 插件:core/system-Core/plugin/ 下未写入 .gitignore 白名单的 .js 仍会加载,但不计入框架 baseline

© xrkseek, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/xrk-crawl of xrkseek/XRK-AGT.

Open the folder on GitHubat commit 63f004e

Compare with similar skills

Xrk Crawl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Xrk Crawl compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Xrk Crawl this skillxrkseek/XRK-AGT138—~1.3kAutomated safety check: PassMIT
Authenticated Session Acquisitiontransilienceai/communitytools563—~1.4kAutomated safety check: PassMIT
Extractactionbook/actionbook1.6k2 repos~3.4kAutomated safety check: PassApache-2.0
Brightdata Proxybrightdata/skills264—~5.1kAutomated safety check: PassMIT
Open BrowserJasonHonKL/Openbrowser114—~1.6kAutomated safety check: PassMIT
Playwright Bowserdisler/bowser265—~1.1kAutomated safety check: NotesNone

Similar skills

  • Authenticated Session Acquisition

    transilienceai/communitytools

    Acquire an authenticated session THROUGH MFA/OTP on an in-scope target and emit a reusable session artifact (Playwright storageState + Bearer) so executors can test the post-auth attack surface.

    563 GitHub stars~1.4k tokensUpdated 2 mo ago
    SecurityAuto-check passed
  • Extract

    actionbook/actionbook

    Extract structured data from websites and produce an executable Playwright script plus extracted data.

    1.6k GitHub starsUsed in 2 repos~3.4k tokens
    Sales & SupportAuto-check passed
  • Brightdata Proxy

    brightdata/skills

    Generate working code that routes HTTP requests through Bright Data proxy networks (Datacenter, ISP, Residential, Mobile) and help users decide which network and IP pool type to use (shared pool…

    264 GitHub stars~5.1k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Open Browser

    JasonHonKL/Openbrowser

    A skill your agent uses whenever the task involves browsing web pages, extracting page content, clicking forms, or completing web workflows.

    114 GitHub stars~1.6k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Playwright Bowser

    disler/bowser

    Headless browser automation using Playwright CLI. An agent skill from disler/bowser.

    265 GitHub stars~1.1k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check: notes
  • Official

    A skill your agent uses when writing Playwright tests, fixing flaky tests, debugging failures, implementing Page Object Model, configuring CI/CD, optimizing performance, mocking APIs, handling…

    6.4k GitHub starsUsed in 7 repos~6.4k tokens
    Testing & QAAuto-check passed

More from xrkseek/XRK-AGT

All 12 skills in this repo
  • Immersive Short Video

    xrkseek/XRK-AGT

    Produce immersive vertical short videos (口播/科普/产品讲解) without AI-slop aesthetics.

    138 GitHub stars~1k tokensUpdated 10 days ago
    Auto-check passed
  • Xrk HTTP API

    xrkseek/XRK-AGT

    当你需要开发或排查 HTTP API(core//http/.js)、理解 HttpApi 基类、HttpApiLoader、业务层约定时使用。

    138 GitHub stars~569 tokensUpdated 10 days ago
    Auto-check passed
  • Xrk Node Runtime

    xrkseek/XRK-AGT

    编写或审查 core/src 代码时,确保使用 Node 26 稳定 API,禁止旧写法(fetch/exec/错误/二进制)。AI 改 Core 前必读。

    138 GitHub stars~907 tokensUpdated 10 days ago
    Auto-check passed
  • Xrk GitHub Research

    xrkseek/XRK-AGT

    在 GitHub 搜索成熟开源实现、对比方案、读 issue/PR 时使用。架构选型、unfamiliar 领域、找业界最佳实践时主动调用。

    138 GitHub stars~202 tokensUpdated 10 days ago
    Auto-check passed
  • Xrk LLM

    xrkseek/XRK-AGT

    当你需要配置/新增/排查 LLM 提供商(OpenAI/Azure/Gemini/Anthropic/Ollama/各类兼容网关)时使用;确保 YAML/Schema/代码一致。

    138 GitHub stars~304 tokensUpdated 10 days ago
    Auto-check passed
  • Xrk Www Compat

    xrkseek/XRK-AGT

    编写或审查 core//www 静态页、校园 WebView 兼容、HttpResponse 前端解包时使用。Core www 走浏览器兼容层(xrk-www-compat)。含零配置静态与有 sign.json(纯静态/产物/反代,sign↔server 合并)挂载。

    138 GitHub stars~544 tokensUpdated 10 days ago
    Auto-check passed

Works with

Questions about Xrk Crawl

What does Xrk Crawl do?

当你需要开发/排查 HTTP 抓取、SSRF、Playwright 受控浏览器、本地字体增强截图,或判断 webfetch 与 browser 工作流如何选型时使用。. Xrk Crawl is an agent skill from xrkseek/XRK-AGT.

When should I use Xrk Crawl?

Xrk Crawl fits situations like: tasks that involve Web application vulnerabilities; tasks that involve Browser testing; tasks that involve Web scraping.

How do I install Xrk Crawl in Claude Code?

Run `npx skills add xrkseek/XRK-AGT --skill xrk-crawl -a claude-code`. Or copy the skill folder (.cursor/skills/xrk-crawl in xrkseek/XRK-AGT) into .claude/skills/xrk-crawl in your project. Claude Code loads it when a task matches its description.

How do I install Xrk Crawl in Codex?

Run `npx skills add xrkseek/XRK-AGT --skill xrk-crawl -a codex`. Or copy the skill folder (.cursor/skills/xrk-crawl in xrkseek/XRK-AGT) into .agents/skills/xrk-crawl in your project. Codex loads it when a task matches its description.

Can I use Xrk Crawl in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xrkseek/XRK-AGT --skill xrk-crawl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/xrk-crawl, .gemini/skills/xrk-crawl, .github/skills/xrk-crawl and .opencode/skills/xrk-crawl in your project.

What does Xrk Crawl need to run?

SKILL.md names no scripts, command-line tools or credentials: Xrk Crawl is instructions for the agent only.

Does Xrk Crawl access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Xrk Crawl safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Xrk Crawl use?

Xrk Crawl is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Xrk Crawl use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Xrk Crawl?

Skills that share tags, products or a category with Xrk Crawl: Authenticated Session Acquisition (transilienceai/communitytools, 563 stars), Extract (actionbook/actionbook, 1.6k stars), Brightdata Proxy (brightdata/skills, 264 stars) and Open Browser (JasonHonKL/Openbrowser, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Xrk Crawl?

xrkseek (a GitHub user) maintains it in xrkseek/XRK-AGT, which has 138 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 1, 2026.

Source: xrkseek/XRK-AGT on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.