Agent skill

Data Formats

by vstorm-co in vstorm-co/pydantic-deepagents

Working with diverse data formats: binary, text, structured, and custom

MITAuto-check passedAI & LLM Engineering

Install Data Formats

skills CLI
$ npx skills add vstorm-co/pydantic-deepagents --skill data-formats -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vstorm-co/pydantic-deepagents data-formats --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vstorm-co/pydantic-deepagents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/apps/cli/skills/data-formats .claude/skills/data-formats && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-formats
GitHub stars
1.1k
Token cost
~714 tokens
SKILL.md length
316 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Working with diverse data formats: binary, text, structured, and custom

  • Works in 5 steps: Hex dump first 256 bytes: xxd file |… → Look for magic bytes, version numbers,… → Check file size — does it suggest a… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers Format Detection, Common Formats, Parsing Strategies and Common Pitfalls
  • Calls python3 and sqlite3

What it does

Data Formats is an agent skill from vstorm-co/pydantic-deepagents. Working with diverse data formats: binary, text, structured, and custom

Its SKILL.md is about 710 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with SQLite and Python. The repository describes itself as: Open-source, self-hosted Claude Code - a terminal AI assistant and the Python framework behind it. Tool-calling, sandboxed execution, multi-agent teams, skills, checkpoints… The licence is MIT.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/data-formats”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Hex dump first 256 bytes: xxd file | head -16
  2. Look for magic bytes, version numbers, string tables
  3. Check file size — does it suggest a pattern? (e.g., N * record_size)
  4. Look for documentation of the format online
  5. Write a minimal parser, test on known values

What it can do on your machine

Read from SKILL.md and the folder at commit 650b592. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • sqlite3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Formats loads about 714 tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 316 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~714

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vstorm-co/pydantic-deepagents at commit 650b592, republished under its MIT licence (© vstorm-co). 316 words, ~714 tokens.

Download SKILL.mdSave it as .claude/skills/data-formats/SKILL.md (or your agent's skills folder).
name
data-formats
description
Working with diverse data formats: binary, text, structured, and custom
tags
data, parsing, formats, benchmark
version
1.0.0

Data Formats

How to work with diverse and unknown data formats.

Format Detection

Always inspect before parsing:

bash
file <filename>                    # MIME type detection
xxd <filename> | head -5           # hex dump (first bytes)
head -3 <filename>                 # text preview
python3 -c "
with open('<filename>', 'rb') as f:
    h = f.read(16)
    print(h, h.hex())
"

Common Formats

Binary
  • Magic bytes: Most binary formats start with a signature (ELF: \x7fELF, PNG: \x89PNG)
  • Endianness: Check if little-endian or big-endian (struct.unpack('<I', ...) vs '>I')
  • Alignment: Fields are often aligned to 4 or 8 bytes
  • Offsets: Binary headers often contain offsets to other sections
Structured text
  • CSV/TSV: Check delimiter (comma, tab, pipe), quoting, header row
  • JSON: python3 -c "import json; json.load(open('f'))"
  • YAML: Check indentation, anchors/aliases
  • TOML: python3 -c "import tomllib; ..."
  • XML: Check encoding declaration, namespaces
Checkpoints / Model files
  • PyTorch: .pt, .pth → torch.load(f, map_location='cpu')
  • TensorFlow: .ckpt → index + data files, use tf.train.load_checkpoint()
  • NumPy: .npy, .npz → numpy.load()
  • HuggingFace: config.json + model.safetensors
  • ONNX: onnx.load()
Database files
  • SQLite: file says "SQLite 3.x database" → sqlite3 <file> ".tables"
  • WAL files: SQLite write-ahead log — recover with sqlite3 PRAGMA
  • CSV dumps: Often need schema inference

Parsing Strategies

Unknown binary format
  1. Hex dump first 256 bytes: xxd file | head -16
  2. Look for magic bytes, version numbers, string tables
  3. Check file size — does it suggest a pattern? (e.g., N * record_size)
  4. Look for documentation of the format online
  5. Write a minimal parser, test on known values
Large structured files
  1. Never load entirely — sample first: head, tail, shuf -n 10
  2. Check consistency: are all lines the same format?
  3. Count fields: head -1 file | awk -F',' '{print NF}'
  4. Watch for: mixed types, missing values, encoding issues
Multi-file datasets
  1. List all files and sizes
  2. Look for manifest/index files (often JSON or CSV)
  3. Check naming patterns — timestamps, sequence numbers, shards
  4. Process one file first, then generalize

Common Pitfalls

  • Assuming UTF-8 when the file is Latin-1 or binary
  • Assuming CSV when it's TSV (or vice versa)
  • Ignoring the header row
  • Not handling quoted fields with embedded delimiters
  • Reading binary files as text (corrupts data)
  • Endianness mismatch (x86 is little-endian, network byte order is big-endian)

© vstorm-co, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in apps/cli/skills/data-formats of vstorm-co/pydantic-deepagents.

Open the folder on GitHubat commit 650b592

Compare with similar skills

Data Formats next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Formats compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Formats this skillvstorm-co/pydantic-deepagents1.1k—~714Automated safety check: PassMIT
Cortexdb Memory Hermesliliang-cn/cortexdb274—~1.7kAutomated safety check: PassMIT
Cognee Session Memory and Improvetopoteretes/cognee32k—~3.2kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
Agent Memorytigerless-labs/agent-memory3k—~1.3kAutomated safety check: PassMIT
MoviePilot Database Operationjxxghp/MoviePilot12k—~7.3kAutomated safety check: PassGPL-3.0

Similar skills

  • Cortexdb Memory Hermes

    liliang-cn/cortexdb

    Give a Python agent (such as Hermes Agent by Nous Research) durable, local-first memory plus a queryable SPARQL knowledge graph, backed by CortexDB through its gRPC sidecar and the cortexdb-client…

    274 GitHub stars~1.7k tokensUpdated yesterday
    Knowledge ManagementAuto-check passed
  • Explains how cognee stores session memory by session_id and bridges it into the permanent graph with improve(), including the stages, results and settings.

    32k GitHub stars~3.2k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Agent Memory

    tigerless-labs/agent-memory

    Read and write the shared long-term memory store. An agent skill from tigerless-labs/agent-memory.

    3k GitHub stars~1.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Inspects, queries and carefully modifies the MoviePilot SQLite or PostgreSQL database through a bundled script that reads connection settings itself, without needing the password in the prompt.

    12k GitHub stars~7.3k tokensUpdated today
    DatabasesAuto-check passed
  • Nxv

    utensils/nxv

    Find any version of any Nix package across nixpkgs git history using the nxv CLI or HTTP API.

    136 GitHub stars~7.9k tokensUpdated 1 mo ago
    Backend & APIsAuto-check: notes

More from vstorm-co/pydantic-deepagents

All 14 skills in this repo
  • Diagram Design

    vstorm-co/pydantic-deepagents

    Best practices for creating research diagrams with Excalidraw MCP tools

    1.1k GitHub stars~884 tokensUpdated 2 days ago
    Auto-check passed
  • Research Methodology

    vstorm-co/pydantic-deepagents

    Best practices for systematic research, source evaluation, and evidence gathering

    1.1k GitHub stars~577 tokensUpdated 2 days ago
    Auto-check passed
  • Systematic Debugging

    vstorm-co/pydantic-deepagents

    Systematic approach to diagnosing and fixing errors. An agent skill from vstorm-co/pydantic-deepagents.

    1.1k GitHub stars~732 tokensUpdated 2 days ago
    Auto-check passed
  • Verification Strategy

    vstorm-co/pydantic-deepagents

    Thorough verification of completed work before declaring done

    1.1k GitHub stars~736 tokensUpdated 2 days ago
    Auto-check passed
  • Environment Discovery

    vstorm-co/pydantic-deepagents

    Systematic exploration of unknown environments before starting work

    1.1k GitHub stars~419 tokensUpdated 2 days ago
    Auto-check passed
  • Skill Creator

    vstorm-co/pydantic-deepagents

    Create new reusable skills from conversation context. An agent skill from vstorm-co/pydantic-deepagents.

    1.1k GitHub stars~331 tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Data Formats

What does Data Formats do?

Working with diverse data formats: binary, text, structured, and custom. Data Formats is an agent skill from vstorm-co/pydantic-deepagents.

When should I use Data Formats?

Data Formats fits situations like: AI & LLM Engineering work in your project.

How do I install Data Formats in Claude Code?

Run `npx skills add vstorm-co/pydantic-deepagents --skill data-formats -a claude-code`. Or copy the skill folder (apps/cli/skills/data-formats in vstorm-co/pydantic-deepagents) into .claude/skills/data-formats in your project. Claude Code loads it when a task matches its description.

How do I install Data Formats in Codex?

Run `npx skills add vstorm-co/pydantic-deepagents --skill data-formats -a codex`. Or copy the skill folder (apps/cli/skills/data-formats in vstorm-co/pydantic-deepagents) into .agents/skills/data-formats in your project. Codex loads it when a task matches its description.

Can I use Data Formats in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vstorm-co/pydantic-deepagents --skill data-formats -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-formats, .gemini/skills/data-formats, .github/skills/data-formats and .opencode/skills/data-formats in your project.

What does Data Formats need to run?

Going by SKILL.md and its folder, Data Formats needs the command-line tools its instructions call (python3 and sqlite3). Our summary lists: Python 3.

Does Data Formats access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Formats safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Formats use?

Data Formats is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Formats use?

About 714 tokens (SKILL.md is roughly 2.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Formats?

Skills that share tags, products or a category with Data Formats: Cortexdb Memory Hermes (liliang-cn/cortexdb, 274 stars), Cognee Session Memory and Improve (topoteretes/cognee, 32k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars) and Agent Memory (tigerless-labs/agent-memory, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Formats?

vstorm-co (a GitHub organization) maintains it in vstorm-co/pydantic-deepagents, which has 1,077 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 6, 2026.

Source: vstorm-co/pydantic-deepagents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.