Agent skill

Octocode Scraping

by bgauryy in bgauryy/octocode

A skill your agent uses when extracting or mapping public web content into a local cited corpus: scrape or crawl a URL/docs site, pull tables/pricing/product fields, diagnose blocked or thin pages…

MITAuto-check passedData & Analytics

Install Octocode Scraping

skills CLI
$ npx skills add bgauryy/octocode --skill octocode-scraping -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bgauryy/octocode octocode-scraping --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/octocode-scraping .claude/skills/octocode-scraping && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
octocode-scraping
GitHub stars
949
Token cost
~1.3k tokens
SKILL.md length
509 words
Files
43 (incl. scripts, references)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when extracting or mapping public web content into a local cited corpus: scrape or crawl a URL/docs site, pull tables/pricing/product fields, diagnose blocked or thin pages…

  • Mapping public web content into a local cited corpus: scrape
  • SKILL.md covers Route (pick one), Scripts (Node, no install;… and References
  • Runs JavaScript scripts from its folder
  • Crawl a URL/docs site

What it does

Octocode Scraping is an agent skill from bgauryy/octocode. Use when extracting or mapping public web content into a local cited corpus: scrape or crawl a URL/docs site, pull tables/pricing/product fields, diagnose blocked or thin pages, or answer from saved pages. Phrases like scrape this URL, crawl the docs, build a corpus, extract pricing. Prefer keyless fetch; ask before hosted spend. Live clicks/HAR/perf → octocode-chrome-devtools.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 45 other files, including scripts and reference files (for example `README.md`, `docs/ADDING_A_VENDOR.md` and `docs/PROVIDERS.md`).

It sits in Data & Analytics, covering Web scraping and Browser testing. It works with Chrome DevTools, Model Context Protocol and GitHub. The repository describes itself as: Code research platform for AI agents; find, understand, and prove context across your code and all of GitHub, in a fraction of the tokens. One toolset, MCP or CLI. The licence is MIT.

When your agent uses it

  • Mapping public web content into a local cited corpus: scrape
  • Crawl a URL/docs site
  • Pull tables/pricing/product fields
  • Diagnose blocked

Example prompts

  • “/octocode-scraping”

Requirements

  • Node.js

What it can do on your machine

Read from SKILL.md and the folder at commit c265e3f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (JavaScript, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Octocode Scraping loads about 1.3k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 509 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bgauryy/octocode at commit c265e3f, republished under its MIT licence (© bgauryy). 509 words, ~1,309 tokens.

Download SKILL.mdSave it as .claude/skills/octocode-scraping/SKILL.md (or your agent's skills folder). This skill also uses 42 other files; get the full folder from GitHub.
name
octocode-scraping
description
Use when extracting or mapping public web content into a local cited corpus: scrape or crawl a URL/docs site, pull tables/pricing/product fields, diagnose blocked or thin pages, or answer from saved pages. Phrases like scrape this URL, crawl the docs, build a corpus, extract pricing. Prefer keyless fetch; ask before hosted spend. Live clicks/HAR/perf → octocode-chrome-devtools.

Octocode Scraping

Flow: FRAME → POLICY → ROUTE → FETCH → CORPUS → SEARCH → CITE → RECOVER.

FRAME before the first fetch: fix target URL/domain, goal, depth, and output shape — vague ask → references/user-inputs.md.

Defaults: one public URL, --mode html, omit --provider (keyless cdp→direct), session .octocode/tmp/scrape/{sessionId}, compact stdout. Search corpus before refetch. Live interaction → chrome-devtools on one port, then har-ingest + corpus-run into the same session. Ask before auth, hosted spend, crawl widen, CAPTCHA/MFA, destructive actions. Cite paths + URL metadata — not raw dumps.

Stop when: two same-class failures (report evidence, route tried, sanitized status, next approval); hosted 403 (wrong key or credits gone — status only, no retry); CAPTCHA/MFA, auth wall, or cookie/profile transfer needed; still blocked after one cdp try (ask before --provider scrapingant); personal data, form submits, purchases, sends, deletes, or account changes in scope; the saved corpus already proves the claim (cite it, do not refetch); crawl widen before reports/summary.md is useful. Recovery table: references/failure-recovery.md.

Route (pick one)

NeedDoSkip
Vague scrape--mode html, omit --providermarkdown / auto hosted
See auto pickscripts/provider-check.mjsguessing
Fetch/crawlscripts/fetch.mjsdeprecated scrapingant-fetch
Corpus on diskcorpus-inspect → find helpers → corpus-runblind refetch
Live click/DOM/authchrome-devtools--provider cdp alone for interaction
Page healthchrome measure + measure-queryhosted scrape for scores
CDP → corpusscripts/har-ingest.mjs --from-cdp-dirnew sessionId
Prove fieldscripts/corpus-run.mjs --regex|--scriptreopen browser
Still blockedevidence + ask → --provider scrapingantsilent spend

Scripts (Node, no install; every one takes --help)

WhenRun
fetch / crawl / extract a URL — the owner of every network callscripts/fetch.mjs --url <u> [--mode html] [--crawl --same-domain --max-pages <n>] [--no-raw]
want the fetch plus an immediate corpus brief in one shotscripts/fetch-and-brief.mjs --url <u> (wraps fetch.mjs → corpus-inspect.mjs)
before routing or hosted spend: which provider auto-wins, credits leftscripts/provider-check.mjs [--provider <p>], scripts/provider-usage.mjs — sanitized, never prints the key
read a saved session first, then search itscripts/corpus-inspect.mjs --session-dir <d> [--page <n>], then scripts/corpus-find.mjs --session-dir <d> --query <t>
pull static DOM, assets, or graph paths out of the corpus (live DOM → chrome-devtools)scripts/dom-find.mjs --kind form/button/table, scripts/resource-list.mjs --kind asset/external, scripts/graph-navigate.mjs --from <nodeId> — each with --session-dir <d>
prove one field locally instead of reopening a browserscripts/corpus-run.mjs --session-dir <d> --roots cdp,extracts --regex <re> (or --script <file>, --concat-parts)
bridge chrome-devtools artifacts into this session, or hand a packet backscripts/har-ingest.mjs --session-dir <d> --from-cdp-dir <run>; reverse with --export-packet
need field names before an extractionscripts/schema-helper.mjs --intent "extract pricing and features"
an old transcript uses the legacy namesdeprecated shims scripts/scrapingant-fetch.mjs, scripts/scrapingant-check.mjs, scripts/scrapingant-usage.mjs forward to fetch.mjs / provider-check.mjs / provider-usage.mjs — call the new names
editing or extending a scriptshared modules in scripts/lib/ (providers registry, client fetch, corpus, analyzers, extractors, text, args, bridge); env/key resolution vendored in scripts/octocode-config.mjs; JSON contracts in scripts/schemas/
Show full SKILL.md (77 more words)Show less

Full flags, roles, and the library map: read scripts/README.md before adding a flag or a vendor.

References

WhenLoad
frame a vague scopereferences/user-inputs.md
legal/safety/privacy policyreferences/scraping-policy.md
cost / keyless vs hosted routereferences/route-selection.md
provider registry / add a vendorreferences/providers.md
hosted API, after approvalreferences/scrapingant.md
search a saved session corpusreferences/session-corpus.md
site graph / navigation / workflowsreferences/website-analysis.md
stdout and corpus file shapesreferences/data-contract.md
extract quality and citationsreferences/extraction-quality.md
live browser bridge playbookreferences/browser-scraping.md
blocked, thin, or oversized resultsreferences/failure-recovery.md

© bgauryy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 42 other files (scripts, references) in skills/octocode-scraping of bgauryy/octocode.

  • SKILL.md
  • README.md
  • docs/ADDING_A_VENDOR.md
  • docs/PROVIDERS.md
  • references/browser-scraping.md
  • references/data-contract.md
  • references/extraction-quality.md
  • references/failure-recovery.md
  • references/providers.md
  • references/route-selection.md
  • references/scraping-policy.md
  • references/scrapingant.md
  • references/session-corpus.md
  • references/user-inputs.md
  • references/website-analysis.md
  • scripts/README.md
  • scripts/corpus-find.mjs
  • scripts/corpus-inspect.mjs
  • … and 25 more

Open the folder on GitHubat commit c265e3f

Compare with similar skills

Octocode Scraping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Octocode Scraping compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Octocode Scraping this skillbgauryy/octocode949—~1.3kAutomated safety check: PassMIT
Chrome Devtoolseinverne/dotfiles1211 repos~1.6kAutomated safety check: NotesApache-2.0
Record E2E Giflablup/backend.ai-webui133—~907Automated safety check: NotesLGPL-3.0
Go Rod Masteraiskillstore/marketplace4334 repos~4.5kAutomated safety check: PassNone
Skill Seekers Builderyusufkaraaslan/Skill_Seekers15k—~760Automated safety check: PassMIT
Bdgszymdzum/browser-debugger-cli173—~3.8kAutomated safety check: NotesMIT

Similar skills

  • Chrome Devtools

    einverne/dotfiles

    Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

    121 GitHub starsUsed in 1 repo~1.6k tokens
    Data & AnalyticsAuto-check: notes
  • Record E2E Gif

    lablup/backend.ai-webui

    Record Playwright e2e tests as one GIF per test case (video → ffmpeg palette GIF) and return a markdown table for a PR description.

    133 GitHub stars~907 tokensUpdated today
    Testing & QAAuto-check: notes
  • Go Rod Master

    aiskillstore/marketplace

    Comprehensive guide for browser automation and web scraping with go-rod (Chrome DevTools Protocol) including stealth anti-bot-detection patterns.

    433 GitHub starsUsed in 4 repos~4.5k tokens
    Productivity & AutomationAuto-check passed
  • Skill Seekers Builder

    yusufkaraaslan/Skill_Seekers

    Detects the type of a knowledge source and uses the Skill Seekers MCP tools to turn docs, repos, PDFs or videos into packaged AI skills.

    15k GitHub stars~760 tokensUpdated 10 days ago
    Agent WorkflowsAuto-check passed
  • Bdg

    szymdzum/browser-debugger-cli

    Use bdg CLI to drive and debug a real Chrome via Chrome DevTools Protocol - navigate, click, fill and submit forms, check what an action changed (navigation, new messages, pending requests), inspect…

    173 GitHub stars~3.8k tokensUpdated today
    Frontend & DesignAuto-check: notes
  • Anti Detect Browser

    antibrow/anti-detect-browser-skills

    Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone…

    17 GitHub stars~9.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check: warnings

More from bgauryy/octocode

All 12 skills in this repo
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated 2 days ago
    Auto-check passed
  • Octocode Code Research

    bgauryy/octocode

    Researches code with evidence: traces callers, imports and cross-repo links, diagnoses failures and reports findings with exact file and line references and a confidence label.

    949 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Writes, repairs and copyedits project docs against the Google developer documentation style guide, verifying claims in the repository before stating them.

    949 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Octocode Mannequin

    bgauryy/octocode

    Poses and animates a 22-bone anatomical humanoid rig with joint range-of-motion limits, using a Node CLI, a Three.js viewer and WebMCP tools an agent can drive live.

    949 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Octocode Skills Manager

    bgauryy/octocode

    Finds, rates, reviews, creates, improves, installs and syncs Agent Skill folders from local workspaces, registries or remote sources, with a user gate before any write.

    949 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Octocode Brainstorming

    bgauryy/octocode

    Walks an idea through framing, diverging into options, researching evidence and stress-testing before converging on a build, prototype, narrow or park decision.

    949 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed

Questions about Octocode Scraping

What does Octocode Scraping do?

A skill your agent uses when extracting or mapping public web content into a local cited corpus: scrape or crawl a URL/docs site, pull tables/pricing/product fields, diagnose blocked or thin pages…. Octocode Scraping is an agent skill from bgauryy/octocode. Use when extracting or mapping public web content into a local cited corpus: scrape or crawl a URL/docs site, pull tables/pricing/product fields, diagnose blocked or thin pages, or answer from saved pages.

When should I use Octocode Scraping?

Octocode Scraping fits situations like: mapping public web content into a local cited corpus: scrape; crawl a URL/docs site; pull tables/pricing/product fields; diagnose blocked.

How do I install Octocode Scraping in Claude Code?

Run `npx skills add bgauryy/octocode --skill octocode-scraping -a claude-code`. Or copy the skill folder (skills/octocode-scraping in bgauryy/octocode) into .claude/skills/octocode-scraping in your project. Claude Code loads it when a task matches its description.

How do I install Octocode Scraping in Codex?

Run `npx skills add bgauryy/octocode --skill octocode-scraping -a codex`. Or copy the skill folder (skills/octocode-scraping in bgauryy/octocode) into .agents/skills/octocode-scraping in your project. Codex loads it when a task matches its description.

Can I use Octocode Scraping in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bgauryy/octocode --skill octocode-scraping -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/octocode-scraping, .gemini/skills/octocode-scraping, .github/skills/octocode-scraping and .opencode/skills/octocode-scraping in your project.

What does Octocode Scraping need to run?

Going by SKILL.md and its folder, Octocode Scraping needs JavaScript for the scripts in its folder. Our summary lists: Node.js.

Does Octocode Scraping access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Octocode Scraping safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Octocode Scraping use?

Octocode Scraping is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Octocode Scraping use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.4k tokens, read only when the agent opens those files.

What are the alternatives to Octocode Scraping?

Skills that share tags, products or a category with Octocode Scraping: Chrome Devtools (einverne/dotfiles, 121 stars), Record E2E Gif (lablup/backend.ai-webui, 133 stars), Go Rod Master (aiskillstore/marketplace, 433 stars) and Skill Seekers Builder (yusufkaraaslan/Skill_Seekers, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Octocode Scraping?

bgauryy (a GitHub user) maintains it in bgauryy/octocode, which has 949 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 9, 2026.

Source: bgauryy/octocode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.