Turn user CSV files and a question into a typed, joined, repeatable local SQLite analysis with traceable records, browser revisions and verified portable exports.

MITAuto-check passedDatabases

Install Data

skills CLI
$ npx skills add autonomous-ai/openharness --skill data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/openharness data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/store/agents/data-studio/skills/data .claude/skills/data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data
GitHub stars
1.1k
Token cost
~1.2k tokens
SKILL.md length
632 words
Files
5 (incl. scripts, references)
Skills in repo
99
Repo updated
First seen
Licence
MIT

At a glance

Turn user CSV files and a question into a typed, joined, repeatable local SQLite analysis with traceable records, browser revisions and verified portable exports.

  • Works in 5 steps: Write the question, method, source… → Make the queries answer the actual… → Add source-record links and a matching… → …
  • Tasks that involve CSV and tabular files
  • SKILL.md covers Understand the question, Build the useful workflow and Verify and hand off
  • Runs JavaScript and Shell scripts from its folder; calls bash and sh

What it does

Data is an agent skill from autonomous-ai/openharness. Turn user CSV files and a question into a typed, joined, repeatable local SQLite analysis with traceable records, browser revisions and verified portable exports.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/analysis-contract.md` and `scripts/update-verdict.sh`).

It sits in Databases, covering CSV and tabular files. It works with SQLite and SQL. The repository describes itself as: The ultimate harness for coding agents and beyond. All your agents. All your machines. One command center. Start with code, then follow your curiosity and build across… The licence is MIT.

When your agent uses it

  • Tasks that involve CSV and tabular files

Example prompts

  • “/data”

Requirements

  • Node.js
  • A Bash shell

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Write the question, method, source mappings, keys, joins, controls and
  2. Make the queries answer the actual question. Rates come from matched totals,
  3. Add source-record links and a matching detail query. Check the detail rows
  4. Use controls, exact tables, charts with truthful axes and empty states, and
  5. Rebuild, revise an input or query substantially, and independently check the

What it can do on your machine

Read from SKILL.md and the folder at commit 54a1f1b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (JavaScript and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • sh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data loads about 1.2k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 42 tokens; SKILL.md has 632 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from autonomous-ai/openharness at commit 54a1f1b, republished under its MIT licence (© autonomous-ai). 632 words, ~1,167 tokens.

Download SKILL.mdSave it as .claude/skills/data/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
data
description
Turn user CSV files and a question into a typed, joined, repeatable local SQLite analysis with traceable records, browser revisions and verified portable exports.

Data Studio

Deliver an answer the user can inspect, revise and rerun with their next data file. Do not stop at a starter screenshot or a watch-only visualization. Read the analysis contract before changing a schema, query, join, unit, missing-value rule or chart.

Understand the question

Use the user's files and known intent. Ask only for materially missing facts: what decision/question, field meanings, units/currency, period boundaries, join keys, duplicates and missing-data treatment. Never ask again for answers already supplied. If example data was approved, keep example=true and make the synthetic-data label visible in the UI and every report.

Inspect a bounded sample and counts. Preserve original CSV text; map columns explicitly in analysis.json. IDs that look numeric remain text. ISO dates are strictly parsed. For money, use decimal with a declared scale and unit: values such as 12.30 become integer cents, with no implicit rounding.

Build the useful workflow

  1. Write the question, method, source mappings, keys, joins, controls and limitations in analysis.json. Add readable queries/.sql and checks/.sql.
  2. Make the queries answer the actual question. Rates come from matched totals, not sums/averages of ratios. Missing remains missing unless the question explicitly defines an absent record as zero. Do not infer causality.
  3. Add source-record links and a matching detail query. Check the detail rows reconcile to the aggregate under the same filters. A valid record number alone does not prove the link supports the interpretation.
  4. Use controls, exact tables, charts with truthful axes and empty states, and visible calculation explanations. Changing SQL in the browser labels the old explanation as needing review. Update the description when formalizing that revision in the workspace.
  5. Rebuild, revise an input or query substantially, and independently check the new result. Include an alternative schema/question in package acceptance. Do not use the same aggregation helper as your only reference calculation.
sh
bash "$DATA_TOOLCHAIN/node.sh" "$HARNESS_WORKSPACE/tools/build.mjs" "$HARNESS_WORKSPACE"
sh "$DATA_SKILLS/data/scripts/update-verdict.sh"

The one-stop command rebuilds and runs the actual isolated browser proof. Node 22.16+ is supplied by the host or Harness-managed runtime. The build writes ready=false until a browser proof passes. Missing Playwright/viewer is a failed or unavailable check, never permission to manually write ready=true.

Show full SKILL.md (281 more words)Show less

Verify and hand off

Keep proof.json specific to the current data, actual controls and exact expected results. Its recipe must wait on the resulting state (not a fixed sleep). The default proof is for the synthetic sales project; replace it for a different question. The shared probe records loaded-file hashes and screenshots. Inspect desktop and 390px layouts, errors, nulls, zero denominators, no matches, negative changes, quoting, duplicate keys and invalid imports.

Exercise controls, keyboard drilldown, source records, revised SQL and CSV replacement. Save JSON, change the view, reopen the JSON and verify the edits, applied parameters and notes survive. Download the actual SQLite database and query it with an independent reader. Extract the original ZIP into another directory, rebuild with its own tools, and open its standalone server. A failed candidate must preserve the last good outputs and keep readiness false.

The result CSV prefixes formula-looking text and encodes null as an empty field; JSON/SQLite retain exact text and null. Explain decimal storage units. Do not substitute screenshots for usable exports. Reports must include applied controls, method, limitations, checks and the revision.

Browser edits are not filesystem autosaves. Save project captures current data/rules/queries/applied controls/notes; unapplied drafts are not saved. Original built project ZIP is the last CLI build. To create a ZIP of browser edits, restore the saved JSON to a new directory, then rebuild it. The portable PROJECT.md documents these commands and bounds.

Never overwrite a legacy workspace while upgrading. Never share, upload, redact or delete user data without authority. Only declared source files enter the portable project. Warn that saved JSON, SQLite and ZIP contain raw inputs. Do not edit vendored SQLite; its hashes and notices must remain intact.

© autonomous-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in store/agents/data-studio/skills/data of autonomous-ai/openharness.

  • SKILL.md
  • references/analysis-contract.md
  • scripts/perf.mjs
  • scripts/screenshot.mjs
  • scripts/update-verdict.sh

Open the folder on GitHubat commit 54a1f1b

Compare with similar skills

Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data this skillautonomous-ai/openharness1.1k—~1.2kAutomated safety check: PassMIT
Dummy Datasetkillvxk/pm-skills-zh167—~595Automated safety check: PassMIT
Duckdb EnaAAaqwq/AGI-Super-Team1051 repos~1.6kAutomated safety check: PassMIT
SQL Database Support for pRESTprest/prest4.6k—~1.6kAutomated safety check: PassMIT
Chdb SQLvemetric/vemetric3941 repos~1.2kAutomated safety check: PassApache-2.0
Cursor BYOK Database Schemaleookun/cursor-byok3.2k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Dummy Dataset

    killvxk/pm-skills-zh

    生成用于测试的逼真虚拟数据集,支持自定义列、约束条件及输出格式(CSV、JSON、SQL、Python 脚本)。适用于创建测试数据、构建模拟数据集,或为开发和演示生成示例数据。

    167 GitHub stars~595 tokensUpdated 6 mo ago
    DatabasesAuto-check passed
  • Duckdb En

    aAAaqwq/AGI-Super-Team

    DuckDB CLI specialist for SQL analysis, data processing and file conversion.

    105 GitHub starsUsed in 1 repo~1.6k tokens
    DatabasesAuto-check passed
  • Guides classifying, gap-analyzing and scaffolding support for a new SQL database in pREST, from Postgres-compatible variants to entirely new dialects.

    4.6k GitHub stars~1.6k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Chdb SQL

    vemetric/vemetric

    A skill your agent uses when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse…

    394 GitHub starsUsed in 1 repo~1.2k tokens
    DatabasesAuto-check passed
  • Cursor BYOK Database Schema

    leookun/cursor-byok

    Guides SQLite schema changes in the Cursor BYOK server, keeping SQLx migrations, the Rust store, API contracts and fixtures aligned.

    3.2k GitHub stars~1.3k tokensUpdated 9 days ago
    DatabasesAuto-check passed
  • Squix

    eduardofuncao/squix

    Run SQL queries across databases (Postgres, MySQL, SQLite, etc.) via the squix CLI.

    273 GitHub stars~784 tokensUpdated 15 days ago
    DatabasesAuto-check passed

More from autonomous-ai/openharness

All 99 skills in this repo
  • G-code Slicer Tool

    autonomous-ai/openharness

    Slices 3D mesh files into printer-profiled plain G-code through real slicer CLIs, with backend discovery, input inspection, dry runs and static validation.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Home Assistant Automation Builder

    autonomous-ai/openharness

    Turns a home-automation request into standard, testable automations.yaml, run against Home Assistant Core's real triggers and verified with its own trace tool.

    1.1k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Score Music Composer

    autonomous-ai/openharness

    Turns a musical brief into LilyPond concert-pitch music, checked parts for each instrument and a playable practice pack.

    1.1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • OrcaSlicer 3MF and G-code Workflow

    autonomous-ai/openharness

    Turns an STL and explicit printer and material requirements into compared OrcaSlicer plans, an editable 3MF project, checked G-code and a portable handoff.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Sheets and Docs Report Builder

    autonomous-ai/openharness

    Builds an editable DOCX report, a formula-driven XLSX workbook and a fresh LibreOffice PDF preview from one structured source file, then checks them together.

    1.1k GitHub stars~708 tokensUpdated today
    Auto-check passed
  • Bambu Labs

    autonomous-ai/openharness

    Dry-run, upload, and cautiously initiate local Bambu Lab print jobs from validated plain .gcode, using Bambu LAN FTPS/MQTT handoffs.

    1.1k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check: warnings

Works with

Questions about Data

What does Data do?

Turn user CSV files and a question into a typed, joined, repeatable local SQLite analysis with traceable records, browser revisions and verified portable exports. Data is an agent skill from autonomous-ai/openharness. Turn user CSV files and a question into a typed, joined, repeatable local SQLite analysis with traceable records, browser revisions and verified portable exports.

When should I use Data?

Data fits situations like: tasks that involve CSV and tabular files.

How do I install Data in Claude Code?

Run `npx skills add autonomous-ai/openharness --skill data -a claude-code`. Or copy the skill folder (store/agents/data-studio/skills/data in autonomous-ai/openharness) into .claude/skills/data in your project. Claude Code loads it when a task matches its description.

How do I install Data in Codex?

Run `npx skills add autonomous-ai/openharness --skill data -a codex`. Or copy the skill folder (store/agents/data-studio/skills/data in autonomous-ai/openharness) into .agents/skills/data in your project. Codex loads it when a task matches its description.

Can I use Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/openharness --skill data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data, .gemini/skills/data, .github/skills/data and .opencode/skills/data in your project.

What does Data need to run?

Going by SKILL.md and its folder, Data needs JavaScript and a shell for the scripts in its folder and the command-line tools its instructions call (bash and sh). Our summary lists: Node.js; A Bash shell.

Does Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data use?

Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Data?

Skills that share tags, products or a category with Data: Dummy Dataset (killvxk/pm-skills-zh, 167 stars), Duckdb En (aAAaqwq/AGI-Super-Team, 105 stars), SQL Database Support for pREST (prest/prest, 4.6k stars) and Chdb SQL (vemetric/vemetric, 394 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/openharness, which has 1,137 GitHub stars. The repository holds 99 skills in this directory. The repository was last updated on October 7, 2026.

Source: autonomous-ai/openharness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.