Agent skill

Web Content Extractor

by OpenMinis in OpenMinis/MinisSkills

Clean and extract the main body content from a webpage URL. An agent skill from OpenMinis/MinisSkills.

MITAuto-check passedData & Analytics

Install Web Content Extractor

skills CLI
$ npx skills add OpenMinis/MinisSkills --skill web-content-extractor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenMinis/MinisSkills web-content-extractor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenMinis/MinisSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/web-content-extractor .claude/skills/web-content-extractor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
web-content-extractor
GitHub stars
444
Token cost
~564 tokens
SKILL.md length
270 words
Files
1
Skills in repo
49
Repo updated
First seen
Licence
MIT

At a glance

Clean and extract the main body content from a webpage URL. An agent skill from OpenMinis/MinisSkills.

  • Works in 2 steps: Defuddle (Default) → Jina AI Reader (Fallback)
  • The user asks to extract
  • SKILL.md covers How it works, Execution Steps and Notes
  • Calls curl; reaches defuddle.md and r.jina.ai

What it does

Web Content Extractor is an agent skill from OpenMinis/MinisSkills. Clean and extract the main body content from a webpage URL. Use this skill whenever the user asks to extract, scrape, or read the main article, text, or content from a webpage, especially if they mention wanting clean Markdown or text without ads and navigation.

Its SKILL.md is about 560 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Web scraping. The repository describes itself as: Skills collection for Minis. The licence is MIT.

When your agent uses it

  • The user asks to extract
  • Read the main article
  • Content from a webpage
  • Especially if they mention wanting clean Markdown

Example prompts

  • “/web-content-extractor”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Defuddle (Default)
  2. Jina AI Reader (Fallback)

What it can do on your machine

Read from SKILL.md and the folder at commit ae8c5db. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • defuddle.md
    • r.jina.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Web Content Extractor loads about 564 tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 270 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~564

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenMinis/MinisSkills at commit ae8c5db, republished under its MIT licence (© OpenMinis). 270 words, ~564 tokens.

Download SKILL.mdSave it as .claude/skills/web-content-extractor/SKILL.md (or your agent's skills folder).
name
web-content-extractor
description
Clean and extract the main body content from a webpage URL. Use this skill whenever the user asks to extract, scrape, or read the main article, text, or content from a webpage, especially if they mention wanting clean Markdown or text without ads and navigation.

Web Content Extractor

This skill helps you extract the clean main content (body text/Markdown) from a webpage URL by using Defuddle or Jina AI's reader API.

How it works

To extract the content of a target URL, you will prepend a specific service URL to the target URL and fetch it. This converts the messy webpage into clean Markdown containing only the main content.

Available Services
  1. Defuddle (Default)

    • Format: https://defuddle.md/<target-url>
    • Example: https://defuddle.md/https://example.com/article
    • Use this as the primary method.
  2. Jina AI Reader (Fallback)

    • Format: https://r.jina.ai/<target-url>
    • Example: https://r.jina.ai/https://example.com/article
    • Use this if Defuddle fails or returns an error.

Execution Steps

  1. Identify the target URL: Extract the full URL the user wants to read from their request. Ensure it includes the protocol (e.g., https://).
  2. Construct the fetch URL: Prepend https://defuddle.md/ to the target URL.
  3. Fetch the content: Use the shell_execute tool with curl -sL "FETCH_URL" to download the content.
    • Example command: curl -sL "https://defuddle.md/https://example.com/article"
  4. Handle Fallbacks: If the curl command fails, returns empty, or returns an error message indicating failure, try the Jina AI service instead: curl -sL "https://r.jina.ai/https://example.com/article"
  5. Process the output: The output will be in Markdown format.
    • If the user asked you to read it to answer a question, use the content to answer.
    • If the user asked you to extract or save it, present the Markdown to them or save it to a file as requested.

Notes

  • Always enclose the URL in quotes in the curl command to prevent shell interpretation of special characters like & or ?.
  • If the target URL is missing http:// or https://, prepend https:// before appending it to the service URL.

© OpenMinis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in web-content-extractor of OpenMinis/MinisSkills.

Open the folder on GitHubat commit ae8c5db

Compare with similar skills

Web Content Extractor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Web Content Extractor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Web Content Extractor this skillOpenMinis/MinisSkills444—~564Automated safety check: PassMIT
Optimize Button Patternsduckduckgo/tracker-radar-collector169—~1.6kAutomated safety check: PassCustom licence
Meowhub Browserzhaojiaqi/MeowHub111—~1.6kAutomated safety check: PassGPL-3.0
FirecrawlPrism-Shadow/penguin-harness2.5k—~902Automated safety check: PassApache-2.0
Yao Doubao Crawleryaojingang/yao-geo-skills868—~475Automated safety check: PassMIT
Apify Core Workflow Bjeremylongshore/tons-of-skills-marketplace2.8k—~1.8kAutomated safety check: PassMIT

Similar skills

  • Optimize Button Patterns

    duckduckgo/tracker-radar-collector

    Iteratively improves cookie-popup button regex patterns in button-patterns.js against labelled-button-texts.csv.

    169 GitHub stars~1.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Meowhub Browser

    zhaojiaqi/MeowHub

    Browse the web using Browserless.io cloud browser service. An agent skill from zhaojiaqi/MeowHub.

    111 GitHub stars~1.6k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Firecrawl

    Prism-Shadow/penguin-harness

    Search the web and scrape pages into clean markdown with the Firecrawl API — query-based discovery, single-URL extraction including public PDFs, driven by curl with a vault-stored API key.

    2.5k GitHub stars~902 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Yao Doubao Crawler

    yaojingang/yao-geo-skills

    A skill your agent uses when a user needs repeated Doubao AI-search collection from web or Android Appium into compatible JSON plus Markdown/Excel/HTML GEO reports.

    868 GitHub stars~475 tokensUpdated 7 days ago
    Data & AnalyticsAuto-check passed
  • Apify Core Workflow B

    jeremylongshore/tons-of-skills-marketplace

    Manage Apify datasets, key-value stores, and request queues programmatically, and orchestrate multi-Actor pipelines.

    2.8k GitHub stars~1.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Data Cleaning

    ericrisco/rsc-harness

    A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…

    167 GitHub stars~3.6k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from OpenMinis/MinisSkills

All 49 skills in this repo
  • Android UI Automation

    OpenMinis/MinisSkills

    Automate Android apps that have no public API or web version by driving the UI layer through the Accessibility Service.

    444 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Codex On Ish

    OpenMinis/MinisSkills

    Install and run OpenAI Codex CLI inside the Minis/iSH Alpine sandbox on iOS, where rustls TLS and async sockets are broken.

    444 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Evidence Chain Builder

    OpenMinis/MinisSkills

    Score a claim against the evidence behind it. An agent skill from OpenMinis/MinisSkills.

    444 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Hyperframes CLI

    OpenMinis/MinisSkills

    HyperFrames CLI and Minis rendering. An agent skill from OpenMinis/MinisSkills.

    444 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Whenpeak

    OpenMinis/MinisSkills

    Predict when a person's brain works best from their sleep, using the WhenPeak performance-intelligence API, and turn it into concrete scheduling advice.

    444 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Openstreetmap Marker

    OpenMinis/MinisSkills

    Generate a mobile-first, immersive, Amap(Gaode)-style custom landmark marker HTML map for travel itinerary planning, place showcasing, location sharing, etc.

    444 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Web Content Extractor

What does Web Content Extractor do?

Clean and extract the main body content from a webpage URL. An agent skill from OpenMinis/MinisSkills. Web Content Extractor is an agent skill from OpenMinis/MinisSkills. Clean and extract the main body content from a webpage URL.

When should I use Web Content Extractor?

Web Content Extractor fits situations like: the user asks to extract; read the main article; content from a webpage; especially if they mention wanting clean Markdown.

How do I install Web Content Extractor in Claude Code?

Run `npx skills add OpenMinis/MinisSkills --skill web-content-extractor -a claude-code`. Or copy the skill folder (web-content-extractor in OpenMinis/MinisSkills) into .claude/skills/web-content-extractor in your project. Claude Code loads it when a task matches its description.

How do I install Web Content Extractor in Codex?

Run `npx skills add OpenMinis/MinisSkills --skill web-content-extractor -a codex`. Or copy the skill folder (web-content-extractor in OpenMinis/MinisSkills) into .agents/skills/web-content-extractor in your project. Codex loads it when a task matches its description.

Can I use Web Content Extractor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenMinis/MinisSkills --skill web-content-extractor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-content-extractor, .gemini/skills/web-content-extractor, .github/skills/web-content-extractor and .opencode/skills/web-content-extractor in your project.

What does Web Content Extractor need to run?

Going by SKILL.md and its folder, Web Content Extractor needs the command-line tools its instructions call (curl).

Does Web Content Extractor access the network?

SKILL.md names 2 domains. In commands or code: defuddle.md and r.jina.ai; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Web Content Extractor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Web Content Extractor use?

Web Content Extractor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Web Content Extractor use?

About 564 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Web Content Extractor?

Skills that share tags, products or a category with Web Content Extractor: Optimize Button Patterns (duckduckgo/tracker-radar-collector, 169 stars), Meowhub Browser (zhaojiaqi/MeowHub, 111 stars), Firecrawl (Prism-Shadow/penguin-harness, 2.5k stars) and Yao Doubao Crawler (yaojingang/yao-geo-skills, 868 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Web Content Extractor?

OpenMinis (a GitHub organization) maintains it in OpenMinis/MinisSkills, which has 444 GitHub stars. The repository holds 49 skills in this directory. The repository was last updated on October 7, 2026.

Source: OpenMinis/MinisSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.