Agent skill

Cheerio Parsing

by Kilo-Org in Kilo-Org/kilo-marketplace

Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines.

Apache-2.0Auto-check passedData & Analytics

Install Cheerio Parsing

skills CLI
$ npx skills add Kilo-Org/kilo-marketplace --skill cheerio-parsing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kilo-Org/kilo-marketplace cheerio-parsing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cheerio-parsing .claude/skills/cheerio-parsing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cheerio-parsing
GitHub stars
190
Used in
1 other repo
Token cost
~2.2k tokens
SKILL.md length
213 words
Files
2
Skills in repo
86
Repo updated
First seen
Licence
Apache-2.0

At a glance

Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines.

  • Works in 10 steps: Always handle missing elements gracefully → Use specific selectors for better… → Cache jQuery-wrapped elements when reusing → …
  • Tasks that involve Web scraping
  • SKILL.md covers Core Expertise, Key Principles, Basic Setup and Selecting Elements, plus 9 more sections
  • Calls npm

What it does

Cheerio Parsing is an agent skill from Kilo-Org/kilo-marketplace. Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Data & Analytics, covering Web scraping and Document parsing. It works with Node.js. The repository describes itself as: Kilo Marketplace - A curated collection of Skills, MCP Servers, and Modes for enhancing AI agent capabilities across the Kilo ecosystem—including Kilo Code (VS Code extension)… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Web scraping
  • Tasks that involve Document parsing

Example prompts

  • “/cheerio-parsing”

Requirements

  • Node.js

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Always handle missing elements gracefully
  2. Use specific selectors for better performance
  3. Cache jQuery-wrapped elements when reusing
  4. Normalize extracted text (trim whitespace)
  5. Resolve relative URLs to absolute
  6. Validate extracted data types
  7. Implement rate limiting when scraping multiple pages
  8. Use appropriate User-Agent headers
  9. Handle character encoding issues
  10. Log extraction failures for debugging

What it can do on your machine

Read from SKILL.md and the folder at commit ff51758. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cheerio Parsing loads about 2.2k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 213 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Kilo-Org/kilo-marketplace at commit ff51758, republished under its Apache-2.0 licence (© Kilo-Org). 213 words, ~2,177 tokens.

Download SKILL.mdSave it as .claude/skills/cheerio-parsing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
cheerio-parsing
description
Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines.
metadata.category
development

Cheerio HTML Parsing

You are an expert in Cheerio, Node.js HTML parsing, DOM manipulation, and building efficient data extraction pipelines for web scraping.

Core Expertise

  • Cheerio API and jQuery-like syntax
  • CSS selector optimization
  • DOM traversal and manipulation
  • HTML/XML parsing strategies
  • Integration with HTTP clients (axios, got, node-fetch)
  • Memory-efficient processing of large documents
  • Data extraction patterns and best practices

Key Principles

  • Write clean, modular extraction functions
  • Use efficient selectors to minimize parsing overhead
  • Handle malformed HTML gracefully
  • Implement proper error handling for missing elements
  • Design reusable scraping utilities
  • Follow functional programming patterns where appropriate

Basic Setup

bash
npm install cheerio axios
Loading HTML
javascript
const cheerio = require('cheerio');
const axios = require('axios');

// Load from string
const $ = cheerio.load('<html><body><h1>Hello</h1></body></html>');

// Load with options
const $ = cheerio.load(html, {
  xmlMode: false,           // Parse as XML
  decodeEntities: true,     // Decode HTML entities
  lowerCaseTags: false,     // Keep tag case
  lowerCaseAttributeNames: false
});

// Fetch and parse
async function fetchAndParse(url) {
  const response = await axios.get(url);
  return cheerio.load(response.data);
}

Selecting Elements

CSS Selectors
javascript
// By tag
$('h1')

// By class
$('.article')

// By ID
$('#main-content')

// By attribute
$('[data-id="123"]')
$('a[href^="https://"]')  // Starts with
$('a[href$=".pdf"]')      // Ends with
$('a[href*="example"]')   // Contains

// Combinations
$('div.article > h2')     // Direct child
$('div.article h2')       // Any descendant
$('h2 + p')               // Adjacent sibling
$('h2 ~ p')               // General sibling

// Pseudo-selectors
$('li:first-child')
$('li:last-child')
$('li:nth-child(2)')
$('li:nth-child(odd)')
$('tr:even')
$('input:not([type="hidden"])')
$('p:contains("specific text")')
Multiple Selectors
javascript
// Select multiple types
$('h1, h2, h3')

// Chain selections
$('.article').find('.title')

Extracting Data

Text Content
javascript
// Get text (includes child text)
const text = $('h1').text();

// Get trimmed text
const text = $('h1').text().trim();

// Get HTML
const html = $('div.content').html();

// Get outer HTML
const outerHtml = $.html($('div.content'));
Attributes
javascript
// Get attribute
const href = $('a').attr('href');
const src = $('img').attr('src');

// Get data attributes
const id = $('div').data('id');  // data-id attribute

// Check if attribute exists
const hasClass = $('div').hasClass('active');
Multiple Elements
javascript
// Iterate with each
const items = [];
$('.product').each((index, element) => {
  items.push({
    name: $(element).find('.name').text().trim(),
    price: $(element).find('.price').text().trim(),
    url: $(element).find('a').attr('href')
  });
});

// Map to array
const titles = $('h2').map((i, el) => $(el).text()).get();

// Filter elements
const featured = $('.product').filter('.featured');

// First/Last
const first = $('li').first();
const last = $('li').last();

// Get by index
const third = $('li').eq(2);

DOM Traversal

Navigation
javascript
// Parent
$('span').parent()
$('span').parents()          // All ancestors
$('span').parents('.container')  // Specific ancestor
$('span').closest('.wrapper')    // Nearest ancestor matching selector

// Children
$('ul').children()           // Direct children
$('ul').children('li.active') // Filtered children
$('div').contents()          // Including text nodes

// Siblings
$('li').siblings()
$('li').next()
$('li').nextAll()
$('li').prev()
$('li').prevAll()
Filtering
javascript
// Filter by selector
$('li').filter('.active')

// Filter by function
$('li').filter((i, el) => $(el).data('price') > 100)

// Find within selection
$('.article').find('img')

// Check conditions
$('li').is('.active')  // Returns boolean
$('li').has('span')    // Has descendant matching selector

Data Extraction Patterns

Table Extraction
javascript
function extractTable(tableSelector) {
  const $ = this;
  const headers = [];
  const rows = [];

  // Get headers
  $(tableSelector).find('th').each((i, el) => {
    headers.push($(el).text().trim());
  });

  // Get rows
  $(tableSelector).find('tbody tr').each((i, row) => {
    const rowData = {};
    $(row).find('td').each((j, cell) => {
      rowData[headers[j]] = $(cell).text().trim();
    });
    rows.push(rowData);
  });

  return rows;
}
List Extraction
javascript
function extractList(selector, itemExtractor) {
  return $(selector).map((i, el) => itemExtractor($(el))).get();
}

// Usage
const products = extractList('.product', ($el) => ({
  name: $el.find('.name').text().trim(),
  price: parseFloat($el.find('.price').text().replace('$', '')),
  image: $el.find('img').attr('src'),
  link: $el.find('a').attr('href')
}));
javascript
function extractPaginationLinks() {
  return $('.pagination a')
    .map((i, el) => $(el).attr('href'))
    .get()
    .filter(href => href && !href.includes('#'));
}

Handling Missing Data

javascript
// Safe extraction with defaults
function safeText(selector, defaultValue = '') {
  const el = $(selector);
  return el.length ? el.text().trim() : defaultValue;
}

function safeAttr(selector, attr, defaultValue = null) {
  const el = $(selector);
  return el.length ? el.attr(attr) : defaultValue;
}

// Optional chaining pattern
const price = $('.price').first().text()?.trim() || 'N/A';

URL Resolution

javascript
const { URL } = require('url');

function resolveUrl(baseUrl, relativeUrl) {
  if (!relativeUrl) return null;
  try {
    return new URL(relativeUrl, baseUrl).href;
  } catch {
    return relativeUrl;
  }
}

// Usage
const baseUrl = 'https://example.com/products/';
$('a').each((i, el) => {
  const href = $(el).attr('href');
  const absoluteUrl = resolveUrl(baseUrl, href);
  console.log(absoluteUrl);
});

Performance Optimization

javascript
// Cache selections
const $products = $('.product');
$products.each((i, el) => {
  const $product = $(el);  // Wrap once
  // Use $product multiple times
});

// Limit parsing scope
const $article = $('.article');
const title = $article.find('.title').text();  // Searches only within article

// Use specific selectors
// Good
$('div.product > h2.title')
// Less efficient
$('div').find('.product').find('h2').filter('.title')

Complete Scraping Example

javascript
const cheerio = require('cheerio');
const axios = require('axios');

async function scrapeProducts(url) {
  const response = await axios.get(url, {
    headers: {
      'User-Agent': 'Mozilla/5.0 (compatible; MyScraper/1.0)'
    }
  });

  const $ = cheerio.load(response.data);
  const products = [];

  $('.product-card').each((index, element) => {
    const $el = $(element);

    products.push({
      name: $el.find('.product-title').text().trim(),
      price: parseFloat(
        $el.find('.price').text().replace(/[^0-9.]/g, '')
      ),
      rating: parseFloat($el.find('.rating').attr('data-rating')) || null,
      image: $el.find('img').attr('src'),
      url: new URL($el.find('a').attr('href'), url).href,
      inStock: !$el.find('.out-of-stock').length
    });
  });

  return products;
}

// With error handling
async function safeScrape(url) {
  try {
    return await scrapeProducts(url);
  } catch (error) {
    console.error(`Failed to scrape ${url}:`, error.message);
    return [];
  }
}

Key Dependencies

  • cheerio
  • axios (HTTP client)
  • got (alternative HTTP client)
  • node-fetch (fetch API for Node.js)
  • p-limit (concurrency control)

Best Practices

  1. Always handle missing elements gracefully
  2. Use specific selectors for better performance
  3. Cache jQuery-wrapped elements when reusing
  4. Normalize extracted text (trim whitespace)
  5. Resolve relative URLs to absolute
  6. Validate extracted data types
  7. Implement rate limiting when scraping multiple pages
  8. Use appropriate User-Agent headers
  9. Handle character encoding issues
  10. Log extraction failures for debugging

© Kilo-Org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/cheerio-parsing of Kilo-Org/kilo-marketplace.

  • SKILL.md
  • LICENSE

Open the folder on GitHubat commit ff51758

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Kilo-Org/kilo-marketplace, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Cheerio Parsing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cheerio Parsing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cheerio Parsing this skillKilo-Org/kilo-marketplace1901 repos~2.2kAutomated safety check: PassApache-2.0
Data Cleaningericrisco/rsc-harness174—~3.6kAutomated safety check: PassMIT
Data Scraperericrisco/rsc-harness174—~3.3kAutomated safety check: PassMIT
Pp Context Devmvanhorn/printing-press-library2.1k—~4.3kAutomated safety check: NotesApache-2.0
Web Access via Browser CDPeze-is/web-access9.1k4 repos~2.2kAutomated safety check: PassMIT
Brightdata SDK JSbrightdata/skills264—~3kAutomated safety check: PassMIT

Similar skills

  • Data Cleaning

    ericrisco/rsc-harness

    A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…

    174 GitHub stars~3.6k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Data Scraper

    ericrisco/rsc-harness

    A skill your agent uses when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable…

    174 GitHub stars~3.3k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Pp Context Dev

    mvanhorn/printing-press-library

    Printing Press CLI for Context.dev. An agent skill from mvanhorn/printing-press-library.

    2.1k GitHub stars~4.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check: notes
  • Routes every web task, from searching to logged-in browsing, through a tiered choice of search, fetch, curl or a real Chrome or Edge session driven over CDP.

    9.1k GitHub starsUsed in 4 repos~2.2k tokens
    Productivity & AutomationAuto-check passed
  • Brightdata SDK JS

    brightdata/skills

    Web data extraction and discovery using the Bright Data JavaScript/TypeScript SDK (@brightdata/sdk).

    264 GitHub stars~3k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed
  • Unipile Linkedin SDK

    LeoYeAI/openclaw-master-skills

    LinkedIn integration via Unipile's official Node.js SDK. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~4.1k tokensUpdated 2 mo ago
    Writing & ContentAuto-check: notes

More from Kilo-Org/kilo-marketplace

All 86 skills in this repo
  • AzureML Project Scaffolding

    Kilo-Org/kilo-marketplace

    Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.

    190 GitHub stars~3.1k tokensUpdated 11 days ago
    Auto-check: notes
  • Jupyter Notebook Builder

    Kilo-Org/kilo-marketplace

    Creates, inspects, edits and runs Jupyter notebooks, scaffolding experiment or tutorial notebooks from templates and preferring a Jupyter MCP server over raw JSON edits.

    190 GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Tableau Dashboard Creator

    Kilo-Org/kilo-marketplace

    Takes a plain-language dashboard request through brand setup, data exploration, planning, an interactive HTML mock and a Tableau implementation spec.

    190 GitHub stars~3.8k tokensUpdated 11 days ago
    Auto-check: notes
  • Elasticsearch File Ingest

    Kilo-Org/kilo-marketplace

    Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    190 GitHub stars~2.8k tokensUpdated 11 days ago
    Auto-check passed
  • Nifi Flow Layout

    Kilo-Org/kilo-marketplace

    A skill your agent uses when arranging Apache NiFi processors, process groups, ports, comments, numbering, crossing connections, dense fan-in/fan-out, or reusable readable canvas layouts.

    190 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed
  • Splunk Ingest Processor Setup

    Kilo-Org/kilo-marketplace

    Render Cisco Data Fabric ingest-time routing workflows and Splunk Cloud Platform Ingest Processor setup plans with SPL2 pipelines, source types, destinations, lifecycle handoffs, queue and…

    190 GitHub stars~1.2k tokensUpdated 11 days ago
    Auto-check passed

Works with

Questions about Cheerio Parsing

What does Cheerio Parsing do?

Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines. Cheerio Parsing is an agent skill from Kilo-Org/kilo-marketplace.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines.

When should I use Cheerio Parsing?

Cheerio Parsing fits situations like: tasks that involve Web scraping; tasks that involve Document parsing.

How do I install Cheerio Parsing in Claude Code?

Run `npx skills add Kilo-Org/kilo-marketplace --skill cheerio-parsing -a claude-code`. Or copy the skill folder (skills/cheerio-parsing in Kilo-Org/kilo-marketplace) into .claude/skills/cheerio-parsing in your project. Claude Code loads it when a task matches its description.

How do I install Cheerio Parsing in Codex?

Run `npx skills add Kilo-Org/kilo-marketplace --skill cheerio-parsing -a codex`. Or copy the skill folder (skills/cheerio-parsing in Kilo-Org/kilo-marketplace) into .agents/skills/cheerio-parsing in your project. Codex loads it when a task matches its description.

Can I use Cheerio Parsing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kilo-Org/kilo-marketplace --skill cheerio-parsing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cheerio-parsing, .gemini/skills/cheerio-parsing, .github/skills/cheerio-parsing and .opencode/skills/cheerio-parsing in your project.

What does Cheerio Parsing need to run?

Going by SKILL.md and its folder, Cheerio Parsing needs the command-line tools its instructions call (npm). Our summary lists: Node.js.

Does Cheerio Parsing access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Cheerio Parsing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cheerio Parsing use?

Cheerio Parsing is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cheerio Parsing use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cheerio Parsing?

Skills that share tags, products or a category with Cheerio Parsing: Data Cleaning (ericrisco/rsc-harness, 174 stars), Data Scraper (ericrisco/rsc-harness, 174 stars), Pp Context Dev (mvanhorn/printing-press-library, 2.1k stars) and Web Access via Browser CDP (eze-is/web-access, 9.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cheerio Parsing?

Kilo-Org (a GitHub organization) maintains it in Kilo-Org/kilo-marketplace, which has 190 GitHub stars. The repository holds 86 skills in this directory. The repository was last updated on September 28, 2026.

Source: Kilo-Org/kilo-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.