Agent skill

Chdb Datastore

by vemetric in vemetric/vemetric

A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

Apache-2.0Auto-check passedData & Analytics

Install Chdb Datastore

skills CLI
$ npx skills add vemetric/vemetric --skill chdb-datastore -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vemetric/vemetric chdb-datastore --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vemetric/vemetric.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/chdb-datastore .claude/skills/chdb-datastore && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chdb-datastore
GitHub stars
395
Used in
2 other repos
Token cost
~1.4k tokens
SKILL.md length
210 words
Files
6 (incl. scripts, references)
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

  • The user has tabular data (pandas DataFrame
  • SKILL.md covers The Key Insight, Decision Tree: Pick the Right…, Connect to Any Data Source —… and After Connecting — Full Pandas…, plus 4 more sections
  • Runs Python scripts from its folder; calls pip and python
  • Json) and wants to filter

What it does

Chdb Datastore is an agent skill from vemetric/vemetric. Use when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas. Provides chDB DataStore — same pandas API, ClickHouse engine underneath. Also handles reading from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake as DataFrames and joining across sources. TRIGGER when: user mentions DataFrame, parquet, csv, "fast pandas", "speed up pandas", or cross-source DataFrame joins; user imports chdb.datastore or from…

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `README.md`, `examples/examples.md` and `references/api-reference.md`). Compatibility notes: Requires Python 3.9+, macOS or Linux. pip install chdb.

It sits in Data & Analytics, covering DataFrames and Data warehousing. It works with pandas, ClickHouse, SQL and PostgreSQL. The repository describes itself as: Simple, yet powerful Web- & Product Analytics. The licence is Apache-2.0.

When your agent uses it

  • The user has tabular data (pandas DataFrame
  • Json) and wants to filter
  • Speed up slow pandas
  • : user mentions DataFrame

Example prompts

  • “fast pandas”
  • “speed up pandas”
  • “/chdb-datastore”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.9+, macOS or Linux. pip install chdb.

What it can do on your machine

Read from SKILL.md and the folder at commit 2352ee8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • clickhouse.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.9+, macOS or Linux. pip install chdb.

    From compatibility in the SKILL.md frontmatter.

Context cost

Chdb Datastore loads about 1.4k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 173 tokens; SKILL.md has 210 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~173
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from vemetric/vemetric at commit 2352ee8, republished under its Apache-2.0 licence (© vemetric). 210 words, ~1,381 tokens.

Download SKILL.mdSave it as .claude/skills/chdb-datastore/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
chdb-datastore
description
Use when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas. Provides chDB DataStore — same pandas API, ClickHouse engine underneath. Also handles reading from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake as DataFrames and joining across sources. TRIGGER when: user mentions DataFrame, parquet, csv, "fast pandas", "speed up pandas", or cross-source DataFrame joins; user imports `chdb.datastore` or `from datastore import DataStore`. SKIP this skill for raw SQL syntax (use chdb-sql instead), ClickHouse server administration, or non-Python DataStore API work.
compatibility
Requires Python 3.9+, macOS or Linux. pip install chdb.
license
Apache-2.0
metadata.author
chdb-io
metadata.version
4.1
metadata.homepage
https://clickhouse.com/docs/chdb

chdb DataStore — It's Just Faster Pandas

The Key Insight

python
# Change this:
import pandas as pd
# To this:
import chdb.datastore as pd
# Everything else stays the same.

DataStore is a lazy, ClickHouse-backed pandas replacement. Your existing pandas code works unchanged — but operations compile to optimized SQL and execute only when results are needed (e.g., print(), len(), iteration).

bash
pip install chdb

Decision Tree: Pick the Right Approach

1. "I have a file/database and want to analyze it with pandas"
   → DataStore.from_file() / from_mysql() / from_s3() etc.
   → See references/connectors.md

2. "I need to join data from different sources"
   → Create DataStores from each source, use .join()
   → See examples/examples.md #3-5

3. "My pandas code is too slow"
   → import chdb.datastore as pd — change one line, keep the rest

4. "I need raw SQL queries"
   → Use the chdb-sql skill instead

Connect to Any Data Source — One Pattern

python
from datastore import DataStore

# Local file (auto-detects .parquet, .csv, .json, .arrow, .orc, .avro, .tsv, .xml)
ds = DataStore.from_file("sales.parquet")

# Database
ds = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")

# Cloud storage
ds = DataStore.from_s3("s3://bucket/data.parquet", nosign=True)

# URI shorthand — auto-detects source type
ds = DataStore.uri("mysql://root:pass@db:3306/shop/orders")

All 16+ sources and URI schemes → connectors.md

After Connecting — Full Pandas API

python
result = ds[ds["age"] > 25]                                          # filter
result = ds[["name", "city"]]                                        # select columns
result = ds.sort_values("revenue", ascending=False)                  # sort
result = ds.groupby("dept")["salary"].mean()                         # groupby
result = ds.assign(margin=lambda x: x["profit"] / x["revenue"])     # computed column
ds["name"].str.upper()                                               # string accessor
ds["date"].dt.year                                                   # datetime accessor
result = ds1.join(ds2, on="id")                                      # join
result = ds.head(10)                                                 # preview
print(ds.to_sql())                                                   # see generated SQL

209 DataFrame methods supported. Full API → api-reference.md

Cross-Source Join — The Killer Feature

python
from datastore import DataStore

customers = DataStore.from_mysql(host="db:3306", database="crm", table="customers", user="root", password="pass")
orders = DataStore.from_file("orders.parquet")

result = (orders
    .join(customers, left_on="customer_id", right_on="id")
    .groupby("country")
    .agg({"amount": "sum", "rating": "mean"})
    .sort_values("sum", ascending=False))
print(result)

More join examples → examples.md

Writing Data

python
source = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")
target = DataStore("file", path="summary.parquet", format="Parquet")

target.insert_into("category", "total", "count").select_from(
    source.groupby("category").select("category", "sum(amount) AS total", "count() AS count")
).execute()

Troubleshooting

ProblemFix
ImportError: No module named 'chdb'pip install chdb
ImportError: cannot import 'DataStore'Use from datastore import DataStore or from chdb.datastore import DataStore
Database connection timeoutInclude port in host: host="db:3306" not host="db"
Join returns empty resultCheck key types match (both int or both string); use .to_sql() to inspect
Unexpected resultsCall ds.to_sql() to see the generated SQL and debug
Environment checkRun python scripts/verify_install.py (from skill directory)

References

Note: This skill teaches how to use chdb DataStore. For raw SQL queries, use the chdb-sql skill. For contributing to chdb source code, see CLAUDE.md in the project root.

© vemetric, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in .agents/skills/chdb-datastore of vemetric/vemetric.

  • SKILL.md
  • README.md
  • examples/examples.md
  • references/api-reference.md
  • references/connectors.md
  • scripts/verify_install.py

Open the folder on GitHubat commit 2352ee8

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in vemetric/vemetric, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Chdb Datastore next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chdb Datastore compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chdb Datastore this skillvemetric/vemetric3952 repos~1.4kAutomated safety check: PassApache-2.0
Analyzing Dataastronomer/agents451—~1.3kAutomated safety check: PassApache-2.0
Transforming Dataancoleman/ai-design-components526—~3kAutomated safety check: PassMIT
Using Timeseries Databasesancoleman/ai-design-components526—~1.7kAutomated safety check: PassMIT
Bigquery Bigframesgoogle/skills21k—~1.3kAutomated safety check: PassApache-2.0
Clickhouse Logs Queriessupabase/supabase111k—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    451 GitHub stars~1.3k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Transforming Data

    ancoleman/ai-design-components

    Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).

    526 GitHub stars~3k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed
  • Using Timeseries Databases

    ancoleman/ai-design-components

    Time-series database implementation for metrics, IoT, financial data, and observability backends.

    526 GitHub stars~1.7k tokensUpdated 10 mo ago
    DatabasesAuto-check passed
  • Bigquery Bigframes

    google/skills

    Official

    Generates Python code using BigQuery DataFrames (BigFrames).

    21k GitHub stars~1.3k tokensUpdated today
    DatabasesAuto-check passed
  • Clickhouse Logs Queries

    supabase/supabase

    Official

    Write, review, and migrate Supabase logs queries against the ClickHouse-backed logs table (the logs.all.otel analytics endpoint).

    111k GitHub stars~2.4k tokensUpdated today
    DatabasesAuto-check passed
  • Modeler

    sidequery/sidemantic

    Build, validate, and manage semantic models using Sidemantic.

    129 GitHub stars~4.2k tokensUpdated yesterday
    DatabasesAuto-check passed

More from vemetric/vemetric

  • Chdb SQL

    vemetric/vemetric

    A skill your agent uses when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse…

    395 GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • MUST USE when designing ClickHouse architectures, selecting between ingestion or modeling patterns, or translating best practices into workload-specific system designs.

    395 GitHub starsUsed in 2 repos~791 tokens
    Auto-check passed
  • Clickhouse Best Practices

    vemetric/vemetric

    MUST USE when reviewing ClickHouse schemas, queries, or configurations.

    395 GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed
  • Clickhouse JS Node Coding

    vemetric/vemetric

    Write idiomatic application code with the ClickHouse Node.js client (@clickhouse/client).

    395 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Troubleshoot and resolve common issues with the ClickHouse Node.js client (@clickhouse/client).

    395 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed

Questions about Chdb Datastore

What does Chdb Datastore do?

A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas. Chdb Datastore is an agent skill from vemetric/vemetric. Use when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

When should I use Chdb Datastore?

Chdb Datastore fits situations like: the user has tabular data (pandas DataFrame; json) and wants to filter; speed up slow pandas; : user mentions DataFrame.

How do I install Chdb Datastore in Claude Code?

Run `npx skills add vemetric/vemetric --skill chdb-datastore -a claude-code`. Or copy the skill folder (.agents/skills/chdb-datastore in vemetric/vemetric) into .claude/skills/chdb-datastore in your project. Claude Code loads it when a task matches its description.

How do I install Chdb Datastore in Codex?

Run `npx skills add vemetric/vemetric --skill chdb-datastore -a codex`. Or copy the skill folder (.agents/skills/chdb-datastore in vemetric/vemetric) into .agents/skills/chdb-datastore in your project. Codex loads it when a task matches its description.

Can I use Chdb Datastore in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vemetric/vemetric --skill chdb-datastore -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chdb-datastore, .gemini/skills/chdb-datastore, .github/skills/chdb-datastore and .opencode/skills/chdb-datastore in your project.

What does Chdb Datastore need to run?

Going by SKILL.md and its folder, Chdb Datastore needs Python for the scripts in its folder and the command-line tools its instructions call (pip and python). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Python 3.9+, macOS or Linux. pip install chdb..

Does Chdb Datastore access the network?

SKILL.md names 1 domain. As links in the text: clickhouse.com. This is read from the text; nothing was executed.

Is Chdb Datastore safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Chdb Datastore use?

Chdb Datastore is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chdb Datastore use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.1k tokens, read only when the agent opens those files.

What are the alternatives to Chdb Datastore?

Skills that share tags, products or a category with Chdb Datastore: Analyzing Data (astronomer/agents, 451 stars), Transforming Data (ancoleman/ai-design-components, 526 stars), Using Timeseries Databases (ancoleman/ai-design-components, 526 stars) and Bigquery Bigframes (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chdb Datastore?

vemetric (a GitHub organization) maintains it in vemetric/vemetric, which has 395 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 9, 2026.

Source: vemetric/vemetric on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.