Crawl4AI Web Scraping
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
A skill your agent uses when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom…
$ npx skills add coralogix/cx-cli --skill cx-data-pipeline -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install coralogix/cx-cli cx-data-pipeline --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/coralogix/cx-cli.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cx-data-pipeline .claude/skills/cx-data-pipeline && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cx-data-pipeline" agent skill from https://github.com/coralogix/cx-cli/tree/master/skills/cx-data-pipeline into .claude/skills/cx-data-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cx-data-pipeline", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/coralogix/cx-cli/tree/master/skills/cx-data-pipelineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add coralogix/cx-cli --skill cx-data-pipeline -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install coralogix/cx-cli cx-data-pipeline --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/coralogix/cx-cli.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cx-data-pipeline .agents/skills/cx-data-pipeline && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cx-data-pipeline" agent skill from https://github.com/coralogix/cx-cli/tree/master/skills/cx-data-pipeline into .agents/skills/cx-data-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cx-data-pipeline", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add coralogix/cx-cli --skill cx-data-pipeline -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install coralogix/cx-cli cx-data-pipeline --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/coralogix/cx-cli.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cx-data-pipeline .cursor/skills/cx-data-pipeline && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cx-data-pipeline" agent skill from https://github.com/coralogix/cx-cli/tree/master/skills/cx-data-pipeline into .cursor/skills/cx-data-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cx-data-pipeline", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/coralogix/cx-cli.git --path skills/cx-data-pipeline--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add coralogix/cx-cli --skill cx-data-pipeline -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install coralogix/cx-cli cx-data-pipeline --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/coralogix/cx-cli.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cx-data-pipeline .gemini/skills/cx-data-pipeline && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cx-data-pipeline" agent skill from https://github.com/coralogix/cx-cli/tree/master/skills/cx-data-pipeline into .gemini/skills/cx-data-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cx-data-pipeline", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install coralogix/cx-cli cx-data-pipelineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add coralogix/cx-cli --skill cx-data-pipeline -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/coralogix/cx-cli.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cx-data-pipeline .github/skills/cx-data-pipeline && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cx-data-pipeline" agent skill from https://github.com/coralogix/cx-cli/tree/master/skills/cx-data-pipeline into .github/skills/cx-data-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cx-data-pipeline", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add coralogix/cx-cli --skill cx-data-pipeline -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install coralogix/cx-cli cx-data-pipeline --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/coralogix/cx-cli.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cx-data-pipeline .opencode/skills/cx-data-pipeline && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cx-data-pipeline" agent skill from https://github.com/coralogix/cx-cli/tree/master/skills/cx-data-pipeline into .opencode/skills/cx-data-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cx-data-pipeline", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cx-data-pipelineA skill your agent uses when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom…
Cx Data Pipeline is an agent skill from coralogix/cx-cli. Use this skill when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom enrichment table", "lookup table", "geo enrichment", "create metric from logs", "events to metrics", "convert logs to metrics", "generate metrics from events", "recording rule", "precomputed metrics", "PromQL recording", "configure data pipeline", "transform log data", "data processing rules", "rule group", "enrichment settings"…
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/e2m-schemas.md`).
It sits in Data & Analytics, covering Data pipelines and ETL. It works with Prometheus. The repository describes itself as: This is the Coralogix CLI. The licence is Apache-2.0.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c071372. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cx Data Pipeline loads about 3k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 223 tokens; SKILL.md has 1,097 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from coralogix/cx-cli at commit c071372, republished under its Apache-2.0 licence (© coralogix). 1,097 words, ~2,950 tokens.
.claude/skills/cx-data-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use this skill when configuring how Coralogix processes, enriches, and transforms data. It covers parsing rules (extract structured fields from raw logs), enrichments (add context from lookup tables), Events2Metrics (derive metrics from log/span events), and recording rules (precompute PromQL expressions).
| Command | Subcommands | Purpose |
|---|---|---|
cx parsing-rules | list, get, create, update, delete, bulk-delete, usage-limits | Manage log parsing rules |
cx enrichments | list, add, remove, overwrite, limit, settings | Manage enrichment rules |
cx enrichments custom | list, get, create, update, delete, search | Manage custom enrichment tables |
cx e2m | list, get, create, update, delete, labels-cardinality, limits | Manage Events2Metrics definitions |
cx recording-rules | list, get, create, update, delete | Manage Prometheus recording rule groups |
Key flags:
--from-file <path> (or - for stdin)-o json for structured output and -p <profile> for profile selectioncx parsing-rules update and cx recording-rules update require both --from-file and the rule group IDcx enrichments custom search requires --id <table-id> and --query <text>cx parsing-rules bulk-delete requires --ids <id1> <id2> ...These commands use complex JSON structures. Always template from an existing resource to avoid format errors:
# 1. Get an existing resource as a template
cx parsing-rules get <rule-group-id> -o json > template.json
# 2. Modify the template (change fields, remove the ID for create operations)
# 3. Create or update
cx parsing-rules create --from-file template.json
cx parsing-rules update --from-file template.json <rule-group-id>This pattern applies to all create/update operations across all 4 commands. It prevents payload format errors that are the #1 cause of failed attempts.
cx parsing-rules list -o json
cx parsing-rules list -o json | jq '[.[] | {id, name, enabled, rule_count: (.rules | length)}]'cx parsing-rules get <existing-rule-group-id> -o json > rule-template.jsonEdit the template for your new service, then:
cx parsing-rules create --from-file rule-template.jsonQuery recent logs to confirm fields are extracted (load cx-telemetry-querying for log querying):
cx logs 'source logs | filter $d.subsystem == "my-service" | limit 10' -o jsoncx parsing-rules usage-limits -o jsoncx enrichments list -o json
cx enrichments settings -o json
cx enrichments limit -o jsoncx enrichments custom list -o json
cx enrichments custom create --from-file table-definition.jsontable-definition.json must use the v5 JSON shape (inline file content, not multipart file=@...):
{
"name": "IP Lookup",
"description": "Maps IPs to locations",
"file": {
"textual": "ip,city\n1.2.3.4,London",
"extension": "csv",
"name": "lookup.csv",
"size": 24
}
}For updates, include customEnrichmentId (number) plus the same fields.
cx enrichments add --from-file enrichment-rules.jsonenrichment-rules.json must use requestEnrichments (not enrichments from list output). Each enrichmentType is an object, not a string:
{
"requestEnrichments": [
{
"fieldName": "sourceIPs",
"enrichmentType": { "geoIp": { "withAsn": true } }
}
]
}Other types: {"aws": {"resourceType": "ec2"}}, {"suspiciousIp": {}}, {"customEnrichment": {"id": 1}}.
cx enrichments custom search --id <table-id> --query "search term"Query logs on hot storage (FrequentSearch tier) to confirm enriched fields appear. Avoid querying archive for verification - ingestion delays can cause false negatives.
cx logs 'source logs | filter $d.enriched_field != null | limit 5' -o jsonE2M derives Prometheus metrics from log/span events. See references/e2m-schemas.md for the full JSON wire format, enums, and cardinality rules.
E2M aggregates events as they stream through the real-time ingestion pipeline into metric series (~1-min resolution). It is forward-only — metrics start from the moment the E2M is created; there is no backfill.
All ingested data flows through the pipeline; a TCO policy routes each stream into a tier, and the tier decides what's possible:
| TCO tier | Storage | E2M / alerts / dashboards |
|---|---|---|
| High | Frequent Search (hot, OpenSearch) | ✅ available |
| Medium | S3 archive (not hot storage) | ✅ available — still processed by the pipeline |
| Low | Compliance only | ❌ no aggregation features |
| Blocked | dropped | ❌ |
The axis is tier / processing level — NOT "Frequent Search vs archive" (Medium is archive and E2M works on it). Do not tell users to "point E2M at archive instead of Frequent Search" — that is incorrect.
Choose logs2metrics vs spans2metrics, the source field(s) + aggregations, and labels (with cardinality in mind — see references/e2m-schemas.md). To scope the E2M to a dataset, set the optional dataSource field to "<dataspace>/<dataset>"; this requires the account feature e2m_dataset_source_enabled (otherwise the API rejects it with "dataSource is not enabled for this company"). Omit it for the standard logs/spans stream.
cx e2m limits -o json # account E2M count limit + used
cx e2m labels-cardinality -o json # see caveat belowThe labels-cardinality endpoint is a draft forecast — given proposed labels + query it returns the per-day distinct-permutation count over the last 7 days, so you can size a design before creating it. But cx e2m labels-cardinality currently takes no arguments, so it sends no draft and returns an empty list (a CLI gap — it can't forecast yet). Until that's wired up, forecast via the UI or estimate permutations manually (product of distinct label values) and set permutationsLimit. Never use high-cardinality fields (IDs, raw URLs, IPs) as labels. Note the forecast only sees Frequent-Search (High-tier) data.
Only cx e2m get returns the full payload ({"e2m": {...}}); list prints a summary. Extract .e2m and drop read-only fields:
cx e2m get <existing-e2m-id> -o json | jq '.e2m | del(.id, .permutations, .createTime, .updateTime, .metricName)' > e2m.jsoncx e2m create --from-file e2m.jsonConfirm series are being produced (load cx-telemetry-querying for metrics querying):
cx metrics search --name "<targetBaseMetricName>"
cx metrics query "<target_metric_name>" --time nowcx tco list / cx-cost-optimization), not an E2M change.lucene filter as a live cx logs/cx spans query and confirm it returns recent results. Note cx logs queries Frequent-Search (High-tier) by default; for a Medium-tier (archive) source add --tier archive, since the data won't appear in a default Frequent-Search query even though E2M still produces series.When the aggregated/metric view is what the customer most cares about, convert High-tier logs → metrics, then downgrade the raw logs High → Medium. Medium still supports E2M/alerts/dashboards and costs less (S3 archive, no hot storage) — you keep cheap, detailed metrics while dropping expensive Frequent-Search retention.
cx usage summary / cx tco list (see cx-cost-optimization).cx-telemetry-querying).cx recording-rules list -o json
cx recording-rules list -o json | jq '[.[] | {id, name, rules: [.rules[]?.record]}]'cx recording-rules get <existing-id> -o json > recording-rule-template.jsoncx recording-rules create --from-file recording-rule-group.jsonConfirm the precomputed metric is available (load cx-telemetry-querying for metrics querying):
cx metrics query "new_precomputed_metric" --time nowcx <command> get <id> -o json > template.json before any create-o json - all payload inspection and creation should use JSON outputcx parsing-rules usage-limits and cx e2m limits before creating to avoid hitting capscx parsing-rules bulk-delete --ids for cleanup, not individual deletesreferences/e2m-schemas.md - Complete Events2Metrics JSON wire format: type/aggType enum values, logsQuery/spansQuery filters, metric labels & fields, the TCO-tier compute model, cardinality/permutations sizing, and gotchascx-telemetry-querying - discover what data is available before configuring pipeline, and verify parsing results, enriched fields, and E2M metric series via log/metrics queriescx-cost-optimization - find high-volume High-tier sources worth converting to metrics, and move the raw logs High→Medium (TCO) after the E2M is verified© coralogix, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/cx-data-pipeline of coralogix/cx-cli.
Open the folder on GitHubat commit c071372
Cx Data Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cx Data Pipeline this skillcoralogix/cx-cli | 121 | — | ~3k | Automated safety check: Pass | Apache-2.0 | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 599 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Glue 09 10 Migrationaws-samples/aws-glue-samples | 1.5k | — | ~2.4k | Automated safety check: Pass | MIT-0 | |
| Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | |
| Dbt Databricks PR Readydatabricks/dbt-databricks | 380 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Apache Spark EngineerJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT |
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
aws-samples/aws-glue-samples
Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.
aws-samples/aws-glue-samples
Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.
databricks/dbt-databricks
A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
MaterializeInc/materialize
Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.
coralogix/cx-cli
A skill your agent uses for any question or action about the user's AI/GenAI applications or agents — their behavior, prompts/responses, quality, hallucinations, guardrails, security, cost/tokens…
coralogix/cx-cli
This skill should be used when the user asks to "manage alerts", "create alert", "list alerts", "delete alert", "check alert status", "enable alert", "disable alert", "investigate firing alerts"…
coralogix/cx-cli
A skill your agent uses when the user asks about AI Center Coding Agents data, wants to reproduce or extend the Coding Agents dashboards, or asks questions about usage, cost, tokens, sessions…
coralogix/cx-cli
A skill your agent uses when the user asks to "check data usage", "list TCO policies", "reduce Coralogix costs", "optimize observability spend", "lower our logging bill", "data budget exceeded"…
coralogix/cx-cli
A skill your agent uses for any question involving telemetry data: "investigate an issue", "debug a problem", "find out why something is slow", "check error rates", "analyze user behavior"…
coralogix/cx-cli
Build and deploy a Coralogix dashboard for a given service from its logs, spans, metrics, and service specs.
Works with
Categories
A skill your agent uses when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom…. Cx Data Pipeline is an agent skill from coralogix/cx-cli.
Cx Data Pipeline fits situations like: the user asks to set up parsing; create parsing rule; extract fields from logs; regex extraction.
Run `npx skills add coralogix/cx-cli --skill cx-data-pipeline -a claude-code`. Or copy the skill folder (skills/cx-data-pipeline in coralogix/cx-cli) into .claude/skills/cx-data-pipeline in your project. Claude Code loads it when a task matches its description.
Run `npx skills add coralogix/cx-cli --skill cx-data-pipeline -a codex`. Or copy the skill folder (skills/cx-data-pipeline in coralogix/cx-cli) into .agents/skills/cx-data-pipeline in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add coralogix/cx-cli --skill cx-data-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cx-data-pipeline, .gemini/skills/cx-data-pipeline, .github/skills/cx-data-pipeline and .opencode/skills/cx-data-pipeline in your project.
Going by SKILL.md and its folder, Cx Data Pipeline needs the command-line tools its instructions call (jq).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cx Data Pipeline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Cx Data Pipeline: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Glue 09 10 Migration (aws-samples/aws-glue-samples, 1.5k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Dbt Databricks PR Ready (databricks/dbt-databricks, 380 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
coralogix (a GitHub organization) maintains it in coralogix/cx-cli, which has 121 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 7, 2026.
Source: coralogix/cx-cli on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.