Deep-dive diagnostics on a low-quality or failed extraction.

MITAuto-check passedBusiness, Finance & HR

Install Diagnose Extraction

skills CLI
$ npx skills add RealEstateWebTools/property_web_scraper --skill diagnose-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RealEstateWebTools/property_web_scraper diagnose-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RealEstateWebTools/property_web_scraper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/diagnose-extraction .claude/skills/diagnose-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
diagnose-extraction
GitHub stars
132
Token cost
~1.2k tokens
SKILL.md length
483 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Deep-dive diagnostics on a low-quality or failed extraction.

  • Works in 6 steps: Run extraction with full diagnostics → Analyze content provenance → Analyze field traces → …
  • Business, Finance & HR work in your project
  • SKILL.md covers Inputs, Workflow, Key analysis patterns and MCP tools
  • Calls npx

What it does

Diagnose Extraction is an agent skill from RealEstateWebTools/property_web_scraper. Deep-dive diagnostics on a low-quality or failed extraction. Analyzes field traces, content provenance, fallback usage, and suggests mapping improvements.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Business, Finance & HR. It works with X (Twitter). The repository describes itself as: Web based UI to make processing scraped data from real estate websites super simple. The licence is MIT.

When your agent uses it

  • Business, Finance & HR work in your project

Example prompts

  • “/diagnose-extraction”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Run extraction with full diagnostics
  2. Analyze content provenance
  3. Analyze field traces
  4. Analyze the HTML structure
  5. Generate recommendations
  6. Offer to apply fixes

What it can do on your machine

Read from SKILL.md and the folder at commit 02088f3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Diagnose Extraction loads about 1.2k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 483 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from RealEstateWebTools/property_web_scraper at commit 02088f3, republished under its MIT licence (© RealEstateWebTools). 483 words, ~1,228 tokens.

Download SKILL.mdSave it as .claude/skills/diagnose-extraction/SKILL.md (or your agent's skills folder).
name
diagnose-extraction
description
Deep-dive diagnostics on a low-quality or failed extraction. Analyzes field traces, content provenance, fallback usage, and suggests mapping improvements.
argument-hint
scraper-name-or-url

Deep Extraction Diagnostics

Perform a thorough analysis of an extraction result to understand why quality is low or fields are missing. Goes beyond /extract by analyzing each field's extraction strategy, suggesting fixes, and identifying structural issues.

Inputs

$ARGUMENTS can be:

  • A scraper name (runs against existing fixture)
  • A URL (fetches and analyzes)
  • A file path to HTML

Workflow

Step 1: Run extraction with full diagnostics

Write and execute an inline script to get the complete diagnostic output:

bash
cd astro-app && npx tsx -e "
import { readFileSync } from 'fs';
import { extractFromHtml } from './src/lib/extractor/html-extractor.js';
const html = readFileSync('<fixture_path>', 'utf-8');
const result = extractFromHtml({
  html,
  sourceUrl: '<url>',
  scraperMappingName: '<name>',
});
const d = result.diagnostics;
console.log(JSON.stringify({
  grade: d?.qualityGrade,
  label: d?.qualityLabel,
  extractionRate: d?.extractionRate,
  weightedRate: d?.weightedExtractionRate,
  totalFields: d?.totalFields,
  populated: d?.populatedFields,
  extractable: d?.extractableFields,
  populatedExtractable: d?.populatedExtractableFields,
  criticalMissing: d?.criticalFieldsMissing,
  emptyFields: d?.emptyFields,
  contentAnalysis: d?.contentAnalysis,
  fieldTraces: d?.fieldTraces,
  splitSchema: result.splitSchema,
}, null, 2));
"
Step 2: Analyze content provenance

Check the contentAnalysis section:

  • appearsBlocked: true — The page was likely bot-blocked (captcha/verify page). The user needs to provide HTML from a real browser session.
  • appearsJsOnly: true — The page is a JS-only shell. The user needs to capture the rendered HTML (browser "Save As" after rendering).
  • jsonLdCount > 0 — JSON-LD structured data is available. Consider adding jsonLdPath strategies.
  • scriptJsonVarsFound — Known script variables detected (PAGE_MODEL, NEXT_DATA, etc). Consider adding scriptJsonPath strategies.
Step 3: Analyze field traces

For each empty or problematic field:

  1. Read the field trace — what strategy was attempted?
  2. Read the mapping — is the CSS selector still valid?
  3. Search the HTML fixture — where does the data actually live?
  4. Check for fallbacks — does the field have fallback strategies?
  5. Check the field importance — is it critical (title, price), important (coords, address), or optional?
Step 4: Analyze the HTML structure

Look at the fixture HTML for:

  • JSON-LD blocks (<script type="application/ld+json">) — often contain title, price, address, coordinates
  • Open Graph meta tags (og:title, og:image, og:description) — good fallback sources
  • Script variables (__NEXT_DATA__, PAGE_MODEL, __INITIAL_STATE__, dataLayer) — rich structured data
  • Microdata attributes (itemprop, itemtype) — semantic HTML markers
  • Twitter card meta tags (twitter:title, twitter:image) — another fallback source
Show full SKILL.md (210 more words)Show less
Step 5: Generate recommendations

Based on the analysis, provide specific recommendations:

  1. Selector updates — new CSS selectors for fields with broken selectors
  2. Fallback chains — add fallbacks arrays using alternative strategies
  3. Strategy switches — switch from fragile cssLocator to robust scriptJsonPath/jsonLdPath
  4. New fields — data available in HTML that isn't being extracted
  5. Mapping structural issues — fields in wrong sections, missing cssCountId, etc.
Step 6: Offer to apply fixes

Present the specific JSON changes needed and offer to:

  1. Edit the mapping file
  2. Update manifest expected values if needed
  3. Run validation tests
  4. Commit the changes

Key analysis patterns

Content SignalRecommendation
JSON-LD present, not usedAdd jsonLdPath strategies (most robust)
__NEXT_DATA__ presentAdd scriptJsonPath with scriptJsonVar: "__NEXT_DATA__"
PAGE_MODEL presentAdd scriptJsonPath with scriptJsonVar: "PAGE_MODEL"
Multiple CSS matchesAdd cssCountId: "0" to pick first element
CSS selector failsCheck if classes changed, try ID-based or microdata selectors
Critical fields missingPriority fix — grade capped at C until resolved
Fallback usedPrimary strategy is broken, should be updated

MCP tools

When the property-scraper MCP server is running, these tools can assist with diagnosis:

  • get_scraper_mapping — inspect the full mapping definition (selectors, regex, fallbacks)
  • list_supported_portals — check portal metadata and expected extraction rates
  • extract_property — re-run extraction with full diagnostics on modified HTML

© RealEstateWebTools, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/diagnose-extraction of RealEstateWebTools/property_web_scraper.

Open the folder on GitHubat commit 02088f3

Compare with similar skills

Diagnose Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Diagnose Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Diagnose Extraction this skillRealEstateWebTools/property_web_scraper132—~1.2kAutomated safety check: PassMIT
Stock Analysis24mlight/StockClaw1012 repos~2kAutomated safety check: NotesMIT
X Researchedinetdb/dexter-jp311—~789Automated safety check: PassMIT
GitHub Project Contributor Finder API Skillbrowser-act/skills6.1k1 repos~1.9kAutomated safety check: PassMIT
Okx Growth Competitioninternet-court/internet-court-skill6.5k2 repos~3kAutomated safety check: PassMIT
Finance Sentimenthimself65/finance-skills3.4k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Stock Analysis

    24mlight/StockClaw

    Analyze stocks and cryptocurrencies using Yahoo Finance data.

    101 GitHub starsUsed in 2 repos~2k tokens
    Business, Finance & HRAuto-check: notes
  • X Research

    edinetdb/dexter-jp

    X/Twitter public sentiment research. An agent skill from edinetdb/dexter-jp.

    311 GitHub stars~789 tokensUpdated yesterday
    Business, Finance & HRAuto-check passed
  • This skill helps users extract GitHub repository project details and contributor contact information using keywords, stars, and update dates.

    6.1k GitHub starsUsed in 1 repo~1.9k tokens
    Business, Finance & HRAuto-check passed
  • Okx Growth Competition

    internet-court/internet-court-skill

    List OKX Agentic Wallet exclusive trading competitions, register users for contests, track participation and leaderboard rankings, and claim won rewards.

    6.5k GitHub starsUsed in 2 repos~3k tokens
    Business, Finance & HRAuto-check passed
  • Finance Sentiment

    himself65/finance-skills

    Fetch normalized stock sentiment across Reddit, X.com, financial news, and Polymarket from the Adanos Finance API: buzz score, bullish percentage, mention or trade counts, and trend.

    3.4k GitHub stars~1.3k tokensUpdated 4 days ago
    Business, Finance & HRAuto-check passed
  • Bags

    alsk1992/CloddsBot

    Bags.fm - Complete Solana token launchpad with creator monetization

    2.9k GitHub stars~864 tokensUpdated 7 days ago
    Business, Finance & HRAuto-check passed

More from RealEstateWebTools/property_web_scraper

  • Add Scraper

    RealEstateWebTools/property_web_scraper

    Add support for scraping property listings from a new website.

    132 GitHub stars~1.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Fix Scraper

    RealEstateWebTools/property_web_scraper

    Diagnose and fix a broken scraper mapping. An agent skill from RealEstateWebTools/property_web_scraper.

    132 GitHub stars~1k tokensUpdated 2 mo ago
    Auto-check passed
  • Extract

    RealEstateWebTools/property_web_scraper

    Run a quick extraction test against a URL or HTML file. An agent skill from RealEstateWebTools/property_web_scraper.

    132 GitHub stars~787 tokensUpdated 2 mo ago
    Auto-check passed
  • Scraper Status

    RealEstateWebTools/property_web_scraper

    Quick overview of all scrapers — quality grades, extraction rates, and coverage gaps.

    132 GitHub stars~364 tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Diagnose Extraction

What does Diagnose Extraction do?

Deep-dive diagnostics on a low-quality or failed extraction. Diagnose Extraction is an agent skill from RealEstateWebTools/property_web_scraper. Deep-dive diagnostics on a low-quality or failed extraction.

When should I use Diagnose Extraction?

Diagnose Extraction fits situations like: business, Finance & HR work in your project.

How do I install Diagnose Extraction in Claude Code?

Run `npx skills add RealEstateWebTools/property_web_scraper --skill diagnose-extraction -a claude-code`. Or copy the skill folder (.claude/skills/diagnose-extraction in RealEstateWebTools/property_web_scraper) into .claude/skills/diagnose-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Diagnose Extraction in Codex?

Run `npx skills add RealEstateWebTools/property_web_scraper --skill diagnose-extraction -a codex`. Or copy the skill folder (.claude/skills/diagnose-extraction in RealEstateWebTools/property_web_scraper) into .agents/skills/diagnose-extraction in your project. Codex loads it when a task matches its description.

Can I use Diagnose Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RealEstateWebTools/property_web_scraper --skill diagnose-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/diagnose-extraction, .gemini/skills/diagnose-extraction, .github/skills/diagnose-extraction and .opencode/skills/diagnose-extraction in your project.

What does Diagnose Extraction need to run?

Going by SKILL.md and its folder, Diagnose Extraction needs the command-line tools its instructions call (npx). Our summary lists: Node.js.

Does Diagnose Extraction access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Diagnose Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Diagnose Extraction use?

Diagnose Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Diagnose Extraction use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Diagnose Extraction?

Skills that share tags, products or a category with Diagnose Extraction: Stock Analysis (24mlight/StockClaw, 101 stars), X Research (edinetdb/dexter-jp, 311 stars), GitHub Project Contributor Finder API Skill (browser-act/skills, 6.1k stars) and Okx Growth Competition (internet-court/internet-court-skill, 6.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Diagnose Extraction?

RealEstateWebTools (a GitHub organization) maintains it in RealEstateWebTools/property_web_scraper, which has 132 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on July 21, 2026.

Source: RealEstateWebTools/property_web_scraper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.