Agent skill

Cx Data Pipeline

by coralogix in coralogix/cx-cli

A skill your agent uses when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom…

Apache-2.0Auto-check passedData & Analytics

Install Cx Data Pipeline

skills CLI
$ npx skills add coralogix/cx-cli --skill cx-data-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install coralogix/cx-cli cx-data-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/coralogix/cx-cli.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cx-data-pipeline .claude/skills/cx-data-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cx-data-pipeline
GitHub stars
121
Token cost
~3k tokens
SKILL.md length
1,097 words
Files
2 (incl. references)
Skills in repo
21
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom…

  • Works in 12 steps: List Existing Rules → Get a Template → Create New Rule Group → …
  • The user asks to set up parsing
  • SKILL.md covers CLI Commands, Working with JSON Payloads, Parsing Rules Workflow and Enrichment Workflow, plus 3 more sections
  • Calls jq

What it does

Cx Data Pipeline is an agent skill from coralogix/cx-cli. Use this skill when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom enrichment table", "lookup table", "geo enrichment", "create metric from logs", "events to metrics", "convert logs to metrics", "generate metrics from events", "recording rule", "precomputed metrics", "PromQL recording", "configure data pipeline", "transform log data", "data processing rules", "rule group", "enrichment settings"…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/e2m-schemas.md`).

It sits in Data & Analytics, covering Data pipelines and ETL. It works with Prometheus. The repository describes itself as: This is the Coralogix CLI. The licence is Apache-2.0.

When your agent uses it

  • The user asks to set up parsing
  • Create parsing rule
  • Extract fields from logs
  • Regex extraction

Example prompts

  • “set up parsing”
  • “create parsing rule”
  • “extract fields from logs”
  • “/cx-data-pipeline”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. List Existing Rules
  2. Get a Template
  3. Create New Rule Group
  4. Verify Parsing
  5. Check Usage Limits
  6. List Enrichment Rules
  7. Create Custom Enrichment Table (if needed)
  8. Add Enrichment Rules
  9. Search Custom Table Data
  10. Verify Enriched Fields
  11. Design the metric
  12. Size it: check limits & cardinality

What it can do on your machine

Read from SKILL.md and the folder at commit c071372. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cx Data Pipeline loads about 3k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 223 tokens; SKILL.md has 1,097 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~223
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from coralogix/cx-cli at commit c071372, republished under its Apache-2.0 licence (© coralogix). 1,097 words, ~2,950 tokens.

Download SKILL.mdSave it as .claude/skills/cx-data-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
cx-data-pipeline
description
Use this skill when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom enrichment table", "lookup table", "geo enrichment", "create metric from logs", "events to metrics", "convert logs to metrics", "generate metrics from events", "recording rule", "precomputed metrics", "PromQL recording", "configure data pipeline", "transform log data", "data processing rules", "rule group", "enrichment settings", "E2M definition", "labels cardinality", "bulk delete rules", "enrichment limits", "search enrichment table", "what should I convert to metrics", "E2M not producing metrics", "E2M no series", "reduce log cost with E2M", "logs to metrics aggregation", "spans to metrics", or wants to configure how Coralogix processes, enriches, or transforms ingested data.
metadata.version
0.1.0

Data Pipeline Skill

Use this skill when configuring how Coralogix processes, enriches, and transforms data. It covers parsing rules (extract structured fields from raw logs), enrichments (add context from lookup tables), Events2Metrics (derive metrics from log/span events), and recording rules (precompute PromQL expressions).


CLI Commands

CommandSubcommandsPurpose
cx parsing-ruleslist, get, create, update, delete, bulk-delete, usage-limitsManage log parsing rules
cx enrichmentslist, add, remove, overwrite, limit, settingsManage enrichment rules
cx enrichments customlist, get, create, update, delete, searchManage custom enrichment tables
cx e2mlist, get, create, update, delete, labels-cardinality, limitsManage Events2Metrics definitions
cx recording-ruleslist, get, create, update, deleteManage Prometheus recording rule groups

Key flags:

  • All create/update operations use --from-file <path> (or - for stdin)
  • All commands support -o json for structured output and -p <profile> for profile selection
  • cx parsing-rules update and cx recording-rules update require both --from-file and the rule group ID
  • cx enrichments custom search requires --id <table-id> and --query <text>
  • cx parsing-rules bulk-delete requires --ids <id1> <id2> ...

Working with JSON Payloads

These commands use complex JSON structures. Always template from an existing resource to avoid format errors:

bash
# 1. Get an existing resource as a template
cx parsing-rules get <rule-group-id> -o json > template.json

# 2. Modify the template (change fields, remove the ID for create operations)

# 3. Create or update
cx parsing-rules create --from-file template.json
cx parsing-rules update --from-file template.json <rule-group-id>

This pattern applies to all create/update operations across all 4 commands. It prevents payload format errors that are the #1 cause of failed attempts.


Parsing Rules Workflow

1. List Existing Rules
bash
cx parsing-rules list -o json
cx parsing-rules list -o json | jq '[.[] | {id, name, enabled, rule_count: (.rules | length)}]'
2. Get a Template
bash
cx parsing-rules get <existing-rule-group-id> -o json > rule-template.json
3. Create New Rule Group

Edit the template for your new service, then:

bash
cx parsing-rules create --from-file rule-template.json
4. Verify Parsing

Query recent logs to confirm fields are extracted (load cx-telemetry-querying for log querying):

bash
cx logs 'source logs | filter $d.subsystem == "my-service" | limit 10' -o json
5. Check Usage Limits
bash
cx parsing-rules usage-limits -o json

Enrichment Workflow

1. List Enrichment Rules
bash
cx enrichments list -o json
cx enrichments settings -o json
cx enrichments limit -o json
2. Create Custom Enrichment Table (if needed)
bash
cx enrichments custom list -o json
cx enrichments custom create --from-file table-definition.json

table-definition.json must use the v5 JSON shape (inline file content, not multipart file=@...):

json
{
  "name": "IP Lookup",
  "description": "Maps IPs to locations",
  "file": {
    "textual": "ip,city\n1.2.3.4,London",
    "extension": "csv",
    "name": "lookup.csv",
    "size": 24
  }
}

For updates, include customEnrichmentId (number) plus the same fields.

3. Add Enrichment Rules
bash
cx enrichments add --from-file enrichment-rules.json

enrichment-rules.json must use requestEnrichments (not enrichments from list output). Each enrichmentType is an object, not a string:

json
{
  "requestEnrichments": [
    {
      "fieldName": "sourceIPs",
      "enrichmentType": { "geoIp": { "withAsn": true } }
    }
  ]
}

Other types: {"aws": {"resourceType": "ec2"}}, {"suspiciousIp": {}}, {"customEnrichment": {"id": 1}}.

4. Search Custom Table Data
bash
cx enrichments custom search --id <table-id> --query "search term"
5. Verify Enriched Fields

Query logs on hot storage (FrequentSearch tier) to confirm enriched fields appear. Avoid querying archive for verification - ingestion delays can cause false negatives.

bash
cx logs 'source logs | filter $d.enriched_field != null | limit 5' -o json

Events2Metrics Workflow

E2M derives Prometheus metrics from log/span events. See references/e2m-schemas.md for the full JSON wire format, enums, and cardinality rules.

How E2M is computed (read this first)

E2M aggregates events as they stream through the real-time ingestion pipeline into metric series (~1-min resolution). It is forward-only — metrics start from the moment the E2M is created; there is no backfill.

All ingested data flows through the pipeline; a TCO policy routes each stream into a tier, and the tier decides what's possible:

TCO tierStorageE2M / alerts / dashboards
HighFrequent Search (hot, OpenSearch)✅ available
MediumS3 archive (not hot storage)✅ available — still processed by the pipeline
LowCompliance only❌ no aggregation features
Blockeddropped❌

The axis is tier / processing level — NOT "Frequent Search vs archive" (Medium is archive and E2M works on it). Do not tell users to "point E2M at archive instead of Frequent Search" — that is incorrect.

1. Design the metric

Choose logs2metrics vs spans2metrics, the source field(s) + aggregations, and labels (with cardinality in mind — see references/e2m-schemas.md). To scope the E2M to a dataset, set the optional dataSource field to "<dataspace>/<dataset>"; this requires the account feature e2m_dataset_source_enabled (otherwise the API rejects it with "dataSource is not enabled for this company"). Omit it for the standard logs/spans stream.

2. Size it: check limits & cardinality
bash
cx e2m limits -o json              # account E2M count limit + used
cx e2m labels-cardinality -o json  # see caveat below

The labels-cardinality endpoint is a draft forecast — given proposed labels + query it returns the per-day distinct-permutation count over the last 7 days, so you can size a design before creating it. But cx e2m labels-cardinality currently takes no arguments, so it sends no draft and returns an empty list (a CLI gap — it can't forecast yet). Until that's wired up, forecast via the UI or estimate permutations manually (product of distinct label values) and set permutationsLimit. Never use high-cardinality fields (IDs, raw URLs, IPs) as labels. Note the forecast only sees Frequent-Search (High-tier) data.

Show full SKILL.md (437 more words)Show less
3. Template from an existing definition

Only cx e2m get returns the full payload ({"e2m": {...}}); list prints a summary. Extract .e2m and drop read-only fields:

bash
cx e2m get <existing-e2m-id> -o json | jq '.e2m | del(.id, .permutations, .createTime, .updateTime, .metricName)' > e2m.json
4. Create the E2M
bash
cx e2m create --from-file e2m.json
5. Verify the metric

Confirm series are being produced (load cx-telemetry-querying for metrics querying):

bash
cx metrics search --name "<targetBaseMetricName>"
cx metrics query "<target_metric_name>" --time now
Troubleshooting: E2M produces no metric series
  1. Check the source data's TCO tier — if it's routed to Low/compliance (or blocked), E2M cannot run. Fix with a TCO change (cx tco list / cx-cost-optimization), not an E2M change.
  2. Verify the query matches streaming data — run the E2M's lucene filter as a live cx logs/cx spans query and confirm it returns recent results. Note cx logs queries Frequent-Search (High-tier) by default; for a Medium-tier (archive) source add --tier archive, since the data won't appear in a default Frequent-Search query even though E2M still produces series.
  3. Remember it's forward-only — no series exist for data ingested before the E2M was created.
Cost optimization: convert High-tier logs to metrics

When the aggregated/metric view is what the customer most cares about, convert High-tier logs → metrics, then downgrade the raw logs High → Medium. Medium still supports E2M/alerts/dashboards and costs less (S3 archive, no hot storage) — you keep cheap, detailed metrics while dropping expensive Frequent-Search retention.

  1. Find high-volume High-tier sources: cx usage summary / cx tco list (see cx-cost-optimization).
  2. Confirm which fields drive dashboards/alerts (see cx-telemetry-querying).
  3. Build + verify the E2M first (steps above).
  4. Then change the TCO policy to move the raw logs High → Medium. Keep data on High or Medium (both support E2M); do not drop it to Low/compliance if metrics or alerts are still needed.

Recording Rules Workflow

1. List Existing Recording Rules
bash
cx recording-rules list -o json
cx recording-rules list -o json | jq '[.[] | {id, name, rules: [.rules[]?.record]}]'
2. Get a Template
bash
cx recording-rules get <existing-id> -o json > recording-rule-template.json
3. Create Recording Rule Group
bash
cx recording-rules create --from-file recording-rule-group.json
4. Verify with PromQL

Confirm the precomputed metric is available (load cx-telemetry-querying for metrics querying):

bash
cx metrics query "new_precomputed_metric" --time now

Key Principles

  • Always template from existing - cx <command> get <id> -o json > template.json before any create
  • Verify after create - query logs/metrics to confirm the pipeline change took effect
  • Use -o json - all payload inspection and creation should use JSON output
  • Check limits first - cx parsing-rules usage-limits and cx e2m limits before creating to avoid hitting caps
  • Bulk operations - use cx parsing-rules bulk-delete --ids for cleanup, not individual deletes

Additional Resources

Reference Files
  • references/e2m-schemas.md - Complete Events2Metrics JSON wire format: type/aggType enum values, logsQuery/spansQuery filters, metric labels & fields, the TCO-tier compute model, cardinality/permutations sizing, and gotchas

  • cx-telemetry-querying - discover what data is available before configuring pipeline, and verify parsing results, enriched fields, and E2M metric series via log/metrics queries
  • cx-cost-optimization - find high-volume High-tier sources worth converting to metrics, and move the raw logs High→Medium (TCO) after the E2M is verified

© coralogix, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/cx-data-pipeline of coralogix/cx-cli.

  • SKILL.md
  • references/e2m-schemas.md

Open the folder on GitHubat commit c071372

Compare with similar skills

Cx Data Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cx Data Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cx Data Pipeline this skillcoralogix/cx-cli121—~3kAutomated safety check: PassApache-2.0
Crawl4AI Web Scrapingsmallnest/goclaw5991 repos~2.5kAutomated safety check: PassMIT
Glue 09 10 Migrationaws-samples/aws-glue-samples1.5k—~2.4kAutomated safety check: PassMIT-0
Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples1.5k—~3.6kAutomated safety check: PassMIT-0
Dbt Databricks PR Readydatabricks/dbt-databricks380—~2.8kAutomated safety check: PassApache-2.0
Apache Spark EngineerJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Glue 09 10 Migration

    aws-samples/aws-glue-samples

    Official

    Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.

    1.5k GitHub stars~2.4k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Official

    Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.

    1.5k GitHub stars~3.6k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Dbt Databricks PR Ready

    databricks/dbt-databricks

    Official

    A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.

    380 GitHub stars~2.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Apache Spark Engineer

    Jeffallan/claude-skills

    Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Data & AnalyticsAuto-check passed
  • Mz Dbt Release

    MaterializeInc/materialize

    Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.

    6.4k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from coralogix/cx-cli

All 21 skills in this repo
  • Cx AI Center

    coralogix/cx-cli

    A skill your agent uses for any question or action about the user's AI/GenAI applications or agents — their behavior, prompts/responses, quality, hallucinations, guardrails, security, cost/tokens…

    121 GitHub stars~2.5k tokensUpdated 4 days ago
    Auto-check passed
  • Cx Alerts

    coralogix/cx-cli

    This skill should be used when the user asks to "manage alerts", "create alert", "list alerts", "delete alert", "check alert status", "enable alert", "disable alert", "investigate firing alerts"…

    121 GitHub stars~2.5k tokensUpdated 4 days ago
    Auto-check passed
  • Cx Coding Agents

    coralogix/cx-cli

    A skill your agent uses when the user asks about AI Center Coding Agents data, wants to reproduce or extend the Coding Agents dashboards, or asks questions about usage, cost, tokens, sessions…

    121 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed
  • Cx Cost Optimization

    coralogix/cx-cli

    A skill your agent uses when the user asks to "check data usage", "list TCO policies", "reduce Coralogix costs", "optimize observability spend", "lower our logging bill", "data budget exceeded"…

    121 GitHub stars~3.4k tokensUpdated 4 days ago
    Auto-check passed
  • Cx Telemetry Querying

    coralogix/cx-cli

    A skill your agent uses for any question involving telemetry data: "investigate an issue", "debug a problem", "find out why something is slow", "check error rates", "analyze user behavior"…

    121 GitHub stars~2.6k tokensUpdated 4 days ago
    Auto-check passed
  • Cx Dashboards

    coralogix/cx-cli

    Build and deploy a Coralogix dashboard for a given service from its logs, spans, metrics, and service specs.

    121 GitHub stars~4.7k tokensUpdated 4 days ago
    Auto-check: warnings

Works with

Questions about Cx Data Pipeline

What does Cx Data Pipeline do?

A skill your agent uses when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom…. Cx Data Pipeline is an agent skill from coralogix/cx-cli.

When should I use Cx Data Pipeline?

Cx Data Pipeline fits situations like: the user asks to set up parsing; create parsing rule; extract fields from logs; regex extraction.

How do I install Cx Data Pipeline in Claude Code?

Run `npx skills add coralogix/cx-cli --skill cx-data-pipeline -a claude-code`. Or copy the skill folder (skills/cx-data-pipeline in coralogix/cx-cli) into .claude/skills/cx-data-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Cx Data Pipeline in Codex?

Run `npx skills add coralogix/cx-cli --skill cx-data-pipeline -a codex`. Or copy the skill folder (skills/cx-data-pipeline in coralogix/cx-cli) into .agents/skills/cx-data-pipeline in your project. Codex loads it when a task matches its description.

Can I use Cx Data Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add coralogix/cx-cli --skill cx-data-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cx-data-pipeline, .gemini/skills/cx-data-pipeline, .github/skills/cx-data-pipeline and .opencode/skills/cx-data-pipeline in your project.

What does Cx Data Pipeline need to run?

Going by SKILL.md and its folder, Cx Data Pipeline needs the command-line tools its instructions call (jq).

Does Cx Data Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cx Data Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cx Data Pipeline use?

Cx Data Pipeline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cx Data Pipeline use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Cx Data Pipeline?

Skills that share tags, products or a category with Cx Data Pipeline: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Glue 09 10 Migration (aws-samples/aws-glue-samples, 1.5k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Dbt Databricks PR Ready (databricks/dbt-databricks, 380 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cx Data Pipeline?

coralogix (a GitHub organization) maintains it in coralogix/cx-cli, which has 121 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 7, 2026.

Source: coralogix/cx-cli on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.