Agent skill

Hypothesis

by anam-org in anam-org/metaxy

Use Hypothesis for property-based testing to automatically generate comprehensive test cases, find edge cases, and write more robust tests with minimal example shrinking.

Apache-2.0Auto-check passedData & Analytics

Install Hypothesis

skills CLI
$ npx skills add anam-org/metaxy --skill hypothesis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install anam-org/metaxy hypothesis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/anam-org/metaxy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/hypothesis .claude/skills/hypothesis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hypothesis
GitHub stars
124
Token cost
~1.7k tokens
SKILL.md length
283 words
Files
2
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Use Hypothesis for property-based testing to automatically generate comprehensive test cases, find edge cases, and write more robust tests with minimal example shrinking.

  • Tasks that involve DataFrames
  • SKILL.md covers Quick Start, Strategies, Settings and Helpers, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Test generation

What it does

Hypothesis is an agent skill from anam-org/metaxy. Use Hypothesis for property-based testing to automatically generate comprehensive test cases, find edge cases, and write more robust tests with minimal example shrinking. Includes Polars parametric testing integration.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EXAMPLES.md`).

It sits in Data & Analytics, covering DataFrames and Test generation. It works with Polars. The repository describes itself as: Pluggable metadata management framework for versioned incremental multimodal data/ML pipelines. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve DataFrames
  • Tasks that involve Test generation

Example prompts

  • “/hypothesis”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 8337842. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • hypothesis.readthedocs.io
    • docs.pola.rs

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hypothesis loads about 1.7k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 283 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from anam-org/metaxy at commit 8337842, republished under its Apache-2.0 licence (© anam-org). 283 words, ~1,721 tokens.

Download SKILL.mdSave it as .claude/skills/hypothesis/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
hypothesis
description
Use Hypothesis for property-based testing to automatically generate comprehensive test cases, find edge cases, and write more robust tests with minimal example shrinking. Includes Polars parametric testing integration.

Hypothesis - Property-Based Testing

Property-based testing framework that generates test cases automatically, finds minimal failing examples through shrinking, and verifies invariants.

Official Docs: https://hypothesis.readthedocs.io/en/latest/

Key Features:

  • Automatic test data generation from strategies
  • Minimal failing example shrinking
  • Stateful testing with rule-based state machines
  • pytest integration
  • Deterministic reproducibility

Quick Start

python
from hypothesis import given
from hypothesis import strategies as st


@given(st.integers())
def test_property(x):
    """Test properties that should always hold"""
    assert abs(x) >= 0


@given(st.lists(st.integers()))
def test_list_property(lst):
    sorted_lst = sorted(lst)
    assert len(sorted_lst) == len(lst)
    # Check monotonic property
    for i in range(len(sorted_lst) - 1):
        assert sorted_lst[i] <= sorted_lst[i + 1]

Strategies

Full reference: https://hypothesis.readthedocs.io/en/latest/data.html

Common strategies:

  • Primitives: st.integers(), st.floats(), st.text(), st.booleans()
  • Collections: st.lists(), st.dictionaries(), st.tuples(), st.sets()
  • Dates/Times: st.dates(), st.datetimes(), st.timedeltas()
  • Combinators: st.one_of(), st.sampled_from(), st.recursive()
  • Type-based: st.from_type(MyClass)
Composite Strategies
python
from hypothesis import strategies as st
from hypothesis.strategies import composite


@composite
def user_strategy(draw):
    age = draw(st.integers(min_value=18, max_value=100))
    name = draw(st.text(min_size=1))
    return {"name": name, "age": age, "is_adult": age >= 18}


@given(user_strategy())
def test_user(user):
    assert user["is_adult"] == (user["age"] >= 18)
Strategy Combinators
python
st.integers().filter(lambda x: x % 2 == 0)  # Filter
st.integers().map(str)  # Transform
st.one_of(st.integers(), st.text())  # Choose between strategies
st.sampled_from([1, 2, 3, 4, 5])  # Pick from collection
st.from_type(MyClass)  # Infer from type hints
st.builds(MyClass, arg1=st.integers())  # Build instances

Settings

python
from hypothesis import given, settings
from hypothesis import strategies as st


@given(st.integers())
@settings(
    max_examples=1000,  # Default: 100
    deadline=None,  # Remove time limit
    derandomize=True,  # Deterministic ordering
)
def test_example(x):
    pass


# Profiles for different environments
settings.register_profile("dev", max_examples=10)
settings.register_profile("ci", max_examples=1000, deadline=None)
# Activate: HYPOTHESIS_PROFILE=ci pytest

Full settings reference: https://hypothesis.readthedocs.io/en/latest/settings.html

Helpers

python
from hypothesis import given, assume, note, example, seed


@given(st.integers(), st.integers())
def test_division(x, y):
    assume(y != 0)  # Skip invalid cases (prefer .filter() instead)
    note(f"Testing {x} / {y}")  # Add debug info
    assert (x / y) * y == x


@given(st.integers())
@example(0)  # Always test specific cases
@seed(12345)  # Reproducible run
def test_something(x):
    pass

Stateful Testing

For testing complex stateful systems with rule-based state machines.

python
from hypothesis.stateful import RuleBasedStateMachine, rule, invariant
from hypothesis import strategies as st


class MyStateMachine(RuleBasedStateMachine):
    def __init__(self):
        super().__init__()
        self.data = []

    @rule(value=st.integers())
    def add(self, value):
        self.data.append(value)

    @invariant()
    def check_invariant(self):
        assert isinstance(self.data, list)


TestMachine = MyStateMachine.TestCase

Full stateful testing guide: https://hypothesis.readthedocs.io/en/latest/stateful.html

Polars Integration

Polars provides built-in parametric testing strategies for generating DataFrames.

Official docs: https://docs.pola.rs/api/python/stable/reference/api/polars.testing.parametric.dataframes.html

python
from hypothesis import given
import polars as pl
from polars.testing.parametric import dataframes, column


# Generate DataFrames with specific column schemas
@given(
    dataframes(
        cols=[
            column("id", dtype=pl.Int64),
            column("name", dtype=pl.String),
            column("value", dtype=pl.Float64),
        ],
        min_size=1,
        max_size=100,
    )
)
def test_dataframe_property(df: pl.DataFrame):
    """Test properties of DataFrame operations"""
    assert df.shape[0] >= 1
    assert set(df.columns) == {"id", "name", "value"}
    assert df["id"].dtype == pl.Int64


# With Narwhals wrapper
import narwhals as nw


@given(dataframes(cols=[column("a", dtype=pl.Int64)]))
def test_narwhals_operation(df: pl.DataFrame):
    nw_df = nw.from_native(df)
    result = nw_df.select(nw.col("a") * 2)
    assert result.shape[0] == nw_df.shape[0]

Key functions:

  • dataframes(): Generate DataFrames with specified columns
  • column(name, dtype, ...): Define column schemas with constraints
  • series(): Generate standalone Series

Column constraints:

  • null_probability: Control null value frequency
  • min_size/max_size: Control row count
  • allow_null: Enable/disable nulls
  • unique: Generate unique values
  • strategy: Custom strategy for column values

Best Practices

  • Use constraints over filters: st.integers(min_value=0) not st.integers().filter(lambda x: x >= 0)
  • Test properties, not examples: Focus on invariants that always hold
  • Combine with @example(): Test specific edge cases explicitly
  • Avoid assume() overuse: Makes tests slow; use filtered strategies
  • Document properties: Clear docstrings explain what invariant is tested
  • Set size limits: Always bound collection sizes to prevent memory issues
  • Use .hypothesis/ in .gitignore: Stores example database locally

Troubleshooting

Common issues and solutions:

  • HealthCheck failures: Too many rejected examples → use constrained strategies or suppress_health_check
  • Flaky tests: Non-deterministic code → use @seed() or @settings(derandomize=True)
  • Slow tests: Too many examples → reduce max_examples or use profiles
  • Deadline exceeded: Complex operations → increase deadline or set to None

Resources

© anam-org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/hypothesis of anam-org/metaxy.

  • SKILL.md
  • EXAMPLES.md

Open the folder on GitHubat commit 8337842

Compare with similar skills

Hypothesis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hypothesis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hypothesis this skillanam-org/metaxy124—~1.7kAutomated safety check: PassApache-2.0
DataframelyQuantco/dataframely619—~2.4kAutomated safety check: PassBSD-3-Clause
Hybrid-Engine Data Analysiscode-yeongyu/oh-my-openagent70k—~1.4kAutomated safety check: PassCustom licence
Polarsdavila7/claude-code-templates32k14 repos~2.3kAutomated safety check: PassMIT
Optimuskgmims-harvard/OptimusKG146—~1.9kAutomated safety check: PassMIT
Polars BioClawBio/ClawBio1.2k—~3.4kAutomated safety check: PassApache-2.0

Similar skills

  • Dataframely

    Quantco/dataframely

    Best practices for polars data processing with dataframely. An agent skill from Quantco/dataframely.

    619 GitHub stars~2.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Hybrid-Engine Data Analysis

    code-yeongyu/oh-my-openagent

    Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.

    70k GitHub stars~1.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Polars

    davila7/claude-code-templates

    Fast DataFrame library (Apache Arrow). An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 14 repos~2.3k tokens
    Data & AnalyticsAuto-check passed
  • Optimuskg

    mims-harvard/OptimusKG

    Guide for using OptimusKG, the biomedical knowledge graph, through the optimuskg Python client.

    146 GitHub stars~1.9k tokensUpdated 17 days ago
    Data & AnalyticsAuto-check passed
  • Polars Bio

    ClawBio/ClawBio

    Fast genomic interval operations (overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps), multi-format bioinformatics I/O, DataFusion SQL, and pileup on Polars DataFrames…

    1.2k GitHub stars~3.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Polars Bio

    K-Dense-AI/scientific-agent-skills

    Performs genomic interval overlap, nearest, merge, coverage, complement and subtraction on Polars DataFrames, and reads or writes BED, VCF, BCF, BAM, CRAM, GFF, GTF, FASTA and FASTQ data.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Data & AnalyticsAuto-check: notes

More from anam-org/metaxy

All 8 skills in this repo
  • Metaxy

    anam-org/metaxy

    This skill should be used when the user asks to "define a feature", "create a BaseFeature class", "track feature versions", "set up metadata store", "field-level lineage", "FieldSpec", "FeatureDep"…

    124 GitHub stars~1.5k tokensUpdated 8 days ago
    Auto-check passed
  • Claude Improve Config

    anam-org/metaxy

    Self-reflect on the current session to identify mistakes and propose improvements to .claude configuration (CLAUDE.md, hooks, skills).

    124 GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Tach

    anam-org/metaxy

    This skill should be used when the user asks to "add a tach module", "configure tach layers", "define module boundaries", "set up interfaces", "run tach check", "check module boundaries", "tach…

    124 GitHub stars~1.2k tokensUpdated 8 days ago
    Auto-check passed
  • Docs Page Frontmatter

    anam-org/metaxy

    Write YAML front matter for documentation pages with appropriate titles and descriptions for social cards.

    124 GitHub stars~948 tokensUpdated 8 days ago
    Auto-check passed
  • Narwhals

    anam-org/metaxy

    Effectively use Narwhals to write dataframe-agnostic code that works seamlessly across multiple Python dataframe libraries.

    124 GitHub stars~3.3k tokensUpdated 8 days ago
    Auto-check passed
  • Sybil

    anam-org/metaxy

    Use Sybil for testing code examples in documentation and docstrings.

    124 GitHub stars~1.6k tokensUpdated 8 days ago
    Auto-check passed

Works with

Questions about Hypothesis

What does Hypothesis do?

Use Hypothesis for property-based testing to automatically generate comprehensive test cases, find edge cases, and write more robust tests with minimal example shrinking. Hypothesis is an agent skill from anam-org/metaxy. Use Hypothesis for property-based testing to automatically generate comprehensive test cases, find edge cases, and write more robust tests with minimal example shrinking.

When should I use Hypothesis?

Hypothesis fits situations like: tasks that involve DataFrames; tasks that involve Test generation.

How do I install Hypothesis in Claude Code?

Run `npx skills add anam-org/metaxy --skill hypothesis -a claude-code`. Or copy the skill folder (.claude/skills/hypothesis in anam-org/metaxy) into .claude/skills/hypothesis in your project. Claude Code loads it when a task matches its description.

How do I install Hypothesis in Codex?

Run `npx skills add anam-org/metaxy --skill hypothesis -a codex`. Or copy the skill folder (.claude/skills/hypothesis in anam-org/metaxy) into .agents/skills/hypothesis in your project. Codex loads it when a task matches its description.

Can I use Hypothesis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add anam-org/metaxy --skill hypothesis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hypothesis, .gemini/skills/hypothesis, .github/skills/hypothesis and .opencode/skills/hypothesis in your project.

What does Hypothesis need to run?

SKILL.md names no scripts, command-line tools or credentials: Hypothesis is instructions for the agent only. Our summary lists: Python 3.

Does Hypothesis access the network?

SKILL.md names 2 domains. As links in the text: hypothesis.readthedocs.io and docs.pola.rs. This is read from the text; nothing was executed.

Is Hypothesis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hypothesis use?

Hypothesis is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hypothesis use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hypothesis?

Skills that share tags, products or a category with Hypothesis: Dataframely (Quantco/dataframely, 619 stars), Hybrid-Engine Data Analysis (code-yeongyu/oh-my-openagent, 70k stars), Polars (davila7/claude-code-templates, 32k stars) and Optimuskg (mims-harvard/OptimusKG, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hypothesis?

anam-org (a GitHub organization) maintains it in anam-org/metaxy, which has 124 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 30, 2026.

Source: anam-org/metaxy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.