Agent skill

Web Scraper API

by oxylabs in oxylabs/agent-skills

Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting.

MITAuto-check passedData & Analytics

Install Web Scraper API

skills CLI
$ npx skills add oxylabs/agent-skills --skill web-scraper-api -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oxylabs/agent-skills web-scraper-api --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oxylabs/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/web-scraper-api .claude/skills/web-scraper-api && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-scraper-api
GitHub stars
875
Token cost
~1.5k tokens
SKILL.md length
485 words
Files
3
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting.

  • Works in 3 steps: Use specific sources when available… → Use universal for unsupported sites -… → Enable parse: true for structured JSON…
  • The user needs to scrape web pages
  • SKILL.md covers Authentication, Endpoint, Core Parameters and Context Parameters, plus 7 more sections
  • Calls curl; reaches realtime.oxylabs.io and data.oxylabs.io; needs OXY_WSA_PASSWORD

What it does

Web Scraper API is an agent skill from oxylabs/agent-skills. Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting. Use when the user needs to scrape web pages, extract product data, get search results, or collect structured data from supported e-commerce and search platforms without worrying about getting blocked and when geo targeting is required.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `examples.md` and `sources.md`).

It sits in Data & Analytics, covering Web scraping. It works with Android and iOS. The repository describes itself as: Official Agent skills of Oxylabs products. The licence is MIT.

When your agent uses it

  • The user needs to scrape web pages
  • Extract product data
  • Get search results

Example prompts

  • “/web-scraper-api”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Use specific sources when available (amazon_product, google_search) - better parsing and reliability
  2. Use universal for unsupported sites - works with any URL
  3. Enable parse: true for structured JSON output on supported sources

What it can do on your machine

Read from SKILL.md and the folder at commit a410198. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • realtime.oxylabs.io
    • data.oxylabs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OXY_WSA_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Scraper API loads about 1.5k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 485 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from oxylabs/agent-skills at commit a410198, republished under its MIT licence (© oxylabs). 485 words, ~1,457 tokens.

Download SKILL.mdSave it as .claude/skills/web-scraper-api/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
web-scraper-api
description
Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting. Use when the user needs to scrape web pages, extract product data, get search results, or collect structured data from supported e-commerce and search platforms without worrying about getting blocked and when geo targeting is required.

Oxylabs Web Scraper API

Authentication

Requires HTTP Basic Auth with credentials from environment variables:

bash
curl -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" ...

Endpoint

POST https://realtime.oxylabs.io/v1/queries   # immediate response
POST https://data.oxylabs.io/v1/queries       # Push-Pull jobs, callbacks, storage
Content-Type: application/json

Core Parameters

ParameterRequiredDescription
sourceYesTarget scraper (e.g., universal, amazon_product, google_search)
urlConditionalURL to scrape (for universal and *_url sources)
queryConditionalSearch query or product ID (for *_search and *_product sources)
parseNoEnable structured data parsing (recommended for supported sources)
renderNoJavaScript rendering: html or png
geo_locationNoGeographic targeting: country/state/city, ZIP/postcode, coordinates, or Criteria ID where supported
session_idNoReuse the same proxy IP across multiple jobs
content_encodingNoSet to base64 when downloading image files via Realtime or Push-Pull
user_agent_typeNoDevice/browser preset, e.g., desktop_chrome, mobile_ios, tablet_android
localeNoInterface language / Accept-Language, e.g., de-DE
callback_urlNoPush-Pull callback endpoint
storage_type, storage_urlNoPush-Pull cloud upload target (gcs, s3, tos, s3_compatible)
markdown, xhrNoEnable markdown or captured XHR result types
browser_instructionsNoRendered browser actions; requires render: "html"
parsing_instructions, parser_presetNoCustom parser rules or saved preset; pair with parse: true
client_notesNoClient-side job tag saved with the job metadata
domain, subdomain, start_page, pages, limit, store_id, delivery_zip, fulfillment_typeSource-specificMarketplace/search/store localization and pagination fields

user_agent_type values: desktop, desktop_chrome, desktop_edge, desktop_firefox, desktop_opera, desktop_safari, mobile, mobile_android, mobile_ios, tablet, tablet_android, tablet_ios.

Context Parameters

Add these as { "key": "...", "value": ... } objects in context:

KeyUse
force_headers, headersMerge custom headers with managed headers
force_cookies, cookiesMerge custom cookies with managed cookies
http_method, contentUse post with Base64-encoded body content
follow_redirectsFollow 3xx redirect chains
successful_status_codesTreat specific non-standard HTTP codes as successful

For multi-format output, enable types in the payload (parse, markdown, xhr, render: "png") and request them with ?type=raw,parsed,png,markdown,xhr.

For batch Push-Pull jobs, use POST /v1/queries/batch with arrays only for query or url; keep all other parameters singular. Maximum batch size is 5,000 values.

Show full SKILL.md (191 more words)Show less

Quick Start

Scrape any URL:

bash
curl -X POST 'https://realtime.oxylabs.io/v1/queries' \
  -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \
  -H 'Content-Type: application/json' \
  -d '{"source": "universal", "url": "https://example.com"}'

Google search with parsing:

bash
curl -X POST 'https://realtime.oxylabs.io/v1/queries' \
  -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \
  -H 'Content-Type: application/json' \
  -d '{"source": "google_search", "query": "best laptops", "parse": true}'

Amazon product by ASIN:

bash
curl -X POST 'https://realtime.oxylabs.io/v1/queries' \
  -u "$OXY_WSA_USERNAME:$OXY_WSA_PASSWORD" \
  -H 'Content-Type: application/json' \
  -d '{"source": "amazon_product", "query": "B07FZ8S74R", "parse": true}'

Choosing the Right Source

  1. Use specific sources when available (amazon_product, google_search) - better parsing and reliability
  2. Use universal for unsupported sites - works with any URL
  3. Enable parse: true for structured JSON output on supported sources

Response Structure

json
{
  "results": [{
    "content": "...",
    "status_code": 200,
    "url": "https://..."
  }]
}

With parse: true, content contains structured data (title, price, reviews, etc.) instead of raw HTML.

Available Sources

For the complete list of 40+ supported sources organized by category, see sources.md.

More Examples

For detailed request/response examples including geo-location, JavaScript rendering, and custom headers, see examples.md.

Error Handling

CodeMeaning
200Success
400Invalid parameters
401Authentication failed
403Access denied
429Rate limit exceeded

Key Guidelines

  • Always set parse: true for supported sources to get structured data
  • Use ZIP codes for US e-commerce geo-location (e.g., "90210")
  • Use country/state format for search engines (e.g., "California,United States")
  • Add render: "html" for JavaScript-heavy pages
  • Use render: "" only to disable automatic forced rendering for force-rendered pages; set client timeouts near 180 seconds for rendered Realtime or Proxy Endpoint requests
  • Add content_encoding: "base64" when scraping image URLs, then decode results[0].content before saving the file

© oxylabs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/web-scraper-api of oxylabs/agent-skills.

  • SKILL.md
  • examples.md
  • sources.md

Open the folder on GitHubat commit a410198

Compare with similar skills

Web Scraper API next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Scraper API compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Scraper API this skilloxylabs/agent-skills875—~1.5kAutomated safety check: PassMIT
Apify App Store Intelligenceapify/awesome-skills264—~3.5kAutomated safety check: PassApache-2.0
Shotshypersocialinc/shots240—~2kAutomated safety check: PassNone
Monitor With HaolemeHaolemeApp/Haoleme157—~1.3kAutomated safety check: PassAGPL-3.0
Meowhub Browserzhaojiaqi/MeowHub111—~1.6kAutomated safety check: PassGPL-3.0
Yao Doubao Crawleryaojingang/yao-geo-skills868—~475Automated safety check: PassMIT

Similar skills

  • Apify App Store Intelligence

    apify/awesome-skills

    Official

    Pull structured Apple App Store and Google Play data — app metadata, price, rating, the 1–5★ ratings histogram, version, developer, and reviews — and watch it for changes over time.

    264 GitHub stars~3.5k tokensUpdated 15 days ago
    MobileAuto-check passed
  • Shots

    hypersocialinc/shots

    Generate, revise, translate, and manage App Store / Google Play marketing screenshots.

    240 GitHub stars~2k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Monitor With Haoleme

    HaolemeApp/Haoleme

    Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.

    157 GitHub stars~1.3k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Meowhub Browser

    zhaojiaqi/MeowHub

    Browse the web using Browserless.io cloud browser service. An agent skill from zhaojiaqi/MeowHub.

    111 GitHub stars~1.6k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Yao Doubao Crawler

    yaojingang/yao-geo-skills

    A skill your agent uses when a user needs repeated Doubao AI-search collection from web or Android Appium into compatible JSON plus Markdown/Excel/HTML GEO reports.

    868 GitHub stars~475 tokensUpdated 7 days ago
    Data & AnalyticsAuto-check passed
  • Anti Detect Browser

    antibrow/anti-detect-browser-skills

    Drive Chromium from standard Playwright APIs with a real-device fingerprint applied in the kernel, one persistent isolated profile per identity, and a per-profile proxy whose exit IP sets timezone…

    932 GitHub stars~9.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check: warnings

More from oxylabs/agent-skills

  • Agent Browser

    oxylabs/agent-skills

    Connects to Oxylabs remote agent browsers over the Chrome DevTools Protocol (CDP) with Playwright or Puppeteer.

    875 GitHub stars~3k tokensUpdated 7 days ago
    Auto-check passed
  • Proxies

    oxylabs/agent-skills

    Oxylabs proxy networks: Residential, Mobile, shared Datacenter/ISP, and Dedicated Datacenter/ISP proxies with geo-targeting, IP rotation, session persistence, and port-based sticky IPs.

    875 GitHub stars~1.7k tokensUpdated 7 days ago
    Auto-check passed
  • Video Data

    oxylabs/agent-skills

    YouTube data extraction API and high-bandwidth proxy downloads.

    875 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Web Unblocker

    oxylabs/agent-skills

    Bypasses anti-bot protections using Oxylabs Web Unblocker, an AI-powered proxy that handles fingerprinting, JavaScript rendering, and retries automatically.

    875 GitHub stars~865 tokensUpdated 7 days ago
    Auto-check passed

Works with

Questions about Web Scraper API

What does Web Scraper API do?

Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting. Web Scraper API is an agent skill from oxylabs/agent-skills. Production-grade web scraping with automatic anti-bot bypass, structured JSON parsing for 40+ targets, and geo-targeting.

When should I use Web Scraper API?

Web Scraper API fits situations like: the user needs to scrape web pages; extract product data; get search results.

How do I install Web Scraper API in Claude Code?

Run `npx skills add oxylabs/agent-skills --skill web-scraper-api -a claude-code`. Or copy the skill folder (skills/web-scraper-api in oxylabs/agent-skills) into .claude/skills/web-scraper-api in your project. Claude Code loads it when a task matches its description.

How do I install Web Scraper API in Codex?

Run `npx skills add oxylabs/agent-skills --skill web-scraper-api -a codex`. Or copy the skill folder (skills/web-scraper-api in oxylabs/agent-skills) into .agents/skills/web-scraper-api in your project. Codex loads it when a task matches its description.

Can I use Web Scraper API in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oxylabs/agent-skills --skill web-scraper-api -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-scraper-api, .gemini/skills/web-scraper-api, .github/skills/web-scraper-api and .opencode/skills/web-scraper-api in your project.

What does Web Scraper API need to run?

Going by SKILL.md and its folder, Web Scraper API needs the command-line tools its instructions call (curl) and credentials named OXY_WSA_PASSWORD.

Does Web Scraper API access the network?

SKILL.md names 2 domains. In commands or code: realtime.oxylabs.io and data.oxylabs.io; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Web Scraper API safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Web Scraper API use?

Web Scraper API is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Scraper API use?

About 1.5k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Scraper API?

Skills that share tags, products or a category with Web Scraper API: Apify App Store Intelligence (apify/awesome-skills, 264 stars), Shots (hypersocialinc/shots, 240 stars), Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars) and Meowhub Browser (zhaojiaqi/MeowHub, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Scraper API?

oxylabs (a GitHub organization) maintains it in oxylabs/agent-skills, which has 875 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 30, 2026.

Source: oxylabs/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.