Agent skill

Feature Engineering

by agiprolabs in agiprolabs/claude-trading-skills

Feature construction from market data for ML trading models including price, volume, on-chain, and microstructure features

MITAuto-check passedData & Analytics

Install Feature Engineering

skills CLI
$ npx skills add agiprolabs/claude-trading-skills --skill feature-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agiprolabs/claude-trading-skills feature-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agiprolabs/claude-trading-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/feature-engineering .claude/skills/feature-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
feature-engineering
GitHub stars
410
Token cost
~2.7k tokens
SKILL.md length
1,026 words
Files
5 (incl. scripts, references)
Skills in repo
68
Repo updated
First seen
Licence
MIT

At a glance

Feature construction from market data for ML trading models including price, volume, on-chain, and microstructure features

  • Works in 11 steps: Price Features → Volume Features → Technical Features → …
  • Tasks that involve Machine learning
  • SKILL.md covers Why Features Beat Models, Feature Categories, Stationarity and Normalization, plus 5 more sections
  • Runs Python scripts from its folder

What it does

Feature Engineering is an agent skill from agiprolabs/claude-trading-skills. Feature construction from market data for ML trading models including price, volume, on-chain, and microstructure features

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/feature_catalog.md`, `references/pitfalls.md` and `scripts/build_features.py`).

It sits in Data & Analytics, covering Machine learning, Trading and backtesting and Smart contracts. The repository describes itself as: 68 trading, DeFi, and quantitative finance Agent Skills. Works with Claude Code, Cursor, Codex, Gemini CLI, and 30+ other tools. The licence is MIT.

When your agent uses it

  • Tasks that involve Machine learning
  • Tasks that involve Trading and backtesting
  • Tasks that involve Smart contracts

Example prompts

  • “/feature-engineering”

Requirements

  • Python 3

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. Price Features
  2. Volume Features
  3. Technical Features
  4. Microstructure Features
  5. On-Chain Features
  6. Cross-Asset Features
  7. Time Features
  8. Remove Low-Variance Features
  9. Correlation Filter
  10. Feature Importance
  11. Mutual Information

What it can do on your machine

Read from SKILL.md and the folder at commit 981e1d7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Feature Engineering loads about 2.7k tokens when it runs, and up to ~6.9k if it reads all its reference files. Until then it costs about 36 tokens; SKILL.md has 1,026 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~36
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from agiprolabs/claude-trading-skills at commit 981e1d7, republished under its MIT licence (© agiprolabs). 1,026 words, ~2,679 tokens.

Download SKILL.mdSave it as .claude/skills/feature-engineering/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
feature-engineering
description
Feature construction from market data for ML trading models including price, volume, on-chain, and microstructure features

Feature Engineering for Trading ML

Feature engineering is the single highest-leverage activity in building ML trading models. Model selection (XGBoost vs. neural net vs. logistic regression) matters far less than the quality and diversity of input features. A simple model on great features will outperform a complex model on raw prices every time.

This skill covers constructing, validating, and selecting features from market data for use in classification (signal-classification) and regression models targeting crypto/Solana token trading.

Why Features Beat Models

Raw OHLCV data is non-stationary, noisy, and high-dimensional. Models trained directly on price series will overfit. Feature engineering transforms raw data into stationary, informative signals that capture distinct aspects of market behavior:

  • Compression: Reduce thousands of price bars to dozens of descriptive statistics
  • Stationarity: Convert non-stationary prices into stationary returns and ratios
  • Domain knowledge: Encode trader intuition (support/resistance, volume climax) as computable quantities
  • Regime awareness: Features that behave differently in trending vs. ranging markets help models adapt

Feature Categories

1. Price Features

Derived purely from OHLCV price columns. These capture trend, momentum, and volatility from the price series itself.

FeatureFormulaLookback
log_returnln(close_t / close_{t-1})1 bar
abs_returnabs(log_return)1 bar
return_volatilitystd(log_return, N)20 bars
momentum_Nclose_t / close_{t-N} - 15, 10, 20
accelerationmomentum_5 - momentum_5[5]10 bars
high_low_range(high - low) / close1 bar
close_position(close - low) / (high - low)1 bar
gapopen_t / close_{t-1} - 11 bar
rolling_skewskew(log_return, N)20 bars
rolling_kurtosiskurtosis(log_return, N)20 bars
2. Volume Features

Volume confirms or contradicts price movements. Divergences between price and volume are among the most reliable signals in short-term trading.

FeatureFormulaLookback
volume_ratiovolume_t / mean(volume, N)20 bars
volume_ma_ratiosma(volume, 5) / sma(volume, 20)20 bars
obv_slopeslope(OBV, N)10 bars
vwap_deviation(close - VWAP) / VWAPintraday
volume_accelerationvolume_ratio_t - volume_ratio_{t-1}21 bars
buy_volume_ratiobuy_volume / total_volume1 bar
dollar_volumeclose * volume1 bar
volume_cvstd(volume, N) / mean(volume, N)20 bars
3. Technical Features

Standard technical indicators computed via pandas-ta. Use the pandas-ta skill for full parameter documentation.

FeatureSourceLookback
rsiRSI(14)14 bars
macd_histogramMACD(12,26,9) histogram33 bars
bb_position(close - BB_lower) / (BB_upper - BB_lower)20 bars
bb_width(BB_upper - BB_lower) / BB_mid20 bars
atr_ratioATR(14) / close14 bars
adxADX(14)14 bars
stoch_kStochastic %K(14,3)14 bars
cciCCI(20)20 bars
mfiMFI(14)14 bars
supertrend_directionSupertrend direction (+1/-1)10 bars
4. Microstructure Features

Derived from trade-level data (individual swaps/transactions). Require on-chain or DEX API data.

FeatureDescription
trade_count_ratioTrades this bar / avg trades per bar
avg_trade_sizeMean trade size in USD
large_trade_pct% of volume from trades > $10k
unique_tradersCount of distinct wallet addresses
buy_count_ratioBuy trades / total trades
trade_size_entropyShannon entropy of trade size distribution
5. On-Chain Features

Derived from blockchain state changes. Require Helius or Solana RPC data.

FeatureDescription
holder_count_changeChange in unique holders over N periods
whale_net_flowNet tokens moved by top-10 holders
token_velocityTransfer volume / circulating supply
liquidity_changeChange in DEX liquidity pool TVL
6. Cross-Asset Features

Capture relationships between the target token and broader market.

FeatureDescription
sol_correlationRolling correlation with SOL price
btc_betaRolling beta to BTC returns
sector_momentumAverage return of tokens in same sector
7. Time Features

Cyclical encoding of calendar time. Use sin/cos encoding to preserve cyclical continuity (hour 23 is close to hour 0).

python
import numpy as np

hour_sin = np.sin(2 * np.pi * hour / 24)
hour_cos = np.cos(2 * np.pi * hour / 24)
day_of_week = np.sin(2 * np.pi * day / 7)

Stationarity

Non-stationary features will cause your model to fail on new data. A feature is stationary if its statistical properties (mean, variance) don't change over time.

Testing for Stationarity

Use the Augmented Dickey-Fuller (ADF) test:

python
from scipy.stats import adfuller

result = adfuller(feature_series.dropna())
p_value = result[1]
is_stationary = p_value < 0.05
Making Features Stationary
Non-StationaryStationary Transform
PriceLog return
VolumeVolume ratio (vol / avg vol)
OBVOBV slope (regression coefficient)
Holder countHolder count change
RSIAlready stationary (bounded 0-100)
Dollar volumeDollar volume / rolling mean

Rule: If a feature trends upward or downward over time, it is non-stationary. Transform it into a ratio, difference, or rate of change.

Show full SKILL.md (400 more words)Show less

Normalization

After computing features, normalize them so that all features have comparable scales. This is critical for distance-based models (KNN, SVM) and helpful for tree models.

MethodFormulaWhen to Use
Z-score(x - mean) / stdGaussian-like distributions
Min-max(x - min) / (max - min)Bounded features (RSI, BB position)
Rankrank(x) / len(x)Heavy-tailed distributions

Critical: Use rolling statistics for normalization. Never use full-sample mean/std — that introduces lookahead bias.

python
# CORRECT: rolling z-score
z = (feature - feature.rolling(60).mean()) / feature.rolling(60).std()

# WRONG: full-sample z-score (lookahead bias!)
z = (feature - feature.mean()) / feature.std()

No-Lookahead Guarantee

The most dangerous bug in trading ML is lookahead bias — using future information to compute features or targets. Follow these rules absolutely:

  1. Rolling calculations only: Never use .mean() or .std() on the full series. Always use .rolling(N).mean().
  2. Shift targets forward, not features backward: The target is close.shift(-N) / close - 1 (future return), not close / close.shift(N) - 1 (past return used as target).
  3. No future index alignment: When joining feature and target DataFrames, verify that feature row t is paired with target row t (where target already contains the forward shift).
  4. Train/test split by time: Never random split. Always train = data[:split_idx], test = data[split_idx:].

Feature Selection

After computing many features, select the most predictive and least redundant:

Step 1: Remove Low-Variance Features
python
from sklearn.feature_selection import VarianceThreshold
selector = VarianceThreshold(threshold=0.01)
X_filtered = selector.fit_transform(X)
Step 2: Correlation Filter

Remove features with > 0.9 correlation to another feature (keep the one with higher target correlation):

python
corr_matrix = X.corr().abs()
upper = corr_matrix.where(np.triu(np.ones(corr_matrix.shape), k=1).astype(bool))
to_drop = [col for col in upper.columns if any(upper[col] > 0.9)]
Step 3: Feature Importance

Train a random forest and rank by importance:

python
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
importances = pd.Series(rf.feature_importances_, index=X.columns).sort_values(ascending=False)
Step 4: Mutual Information

Non-linear alternative to correlation:

python
from sklearn.feature_selection import mutual_info_classif
mi = mutual_info_classif(X_train, y_train, random_state=42)
mi_scores = pd.Series(mi, index=X.columns).sort_values(ascending=False)

Label Creation

Labels (targets) define what the model learns to predict.

Binary Classification
python
forward_return = close.shift(-N) / close - 1
label = (forward_return > threshold).astype(int)  # 1 = up, 0 = not up

Typical thresholds: 1% for 1h bars, 3% for 4h bars, 5% for daily bars.

Multi-Class Classification
python
label = pd.cut(forward_return,
               bins=[-np.inf, -threshold, threshold, np.inf],
               labels=[0, 1, 2])  # 0=down, 1=flat, 2=up
Regression
python
target = forward_return  # Predict exact return magnitude

Binary classification is recommended for initial models — it's simpler and more robust to noise.

Integration with Other Skills

  • pandas-ta: Compute technical indicators that become features
  • birdeye-api: Fetch OHLCV and trade data for feature computation
  • helius-api: Fetch on-chain data for holder/whale features
  • signal-classification: Use engineered features as model inputs
  • regime-detection: Regime labels as features or for regime-conditional models
  • ohlcv-processing: Clean and resample raw data before feature computation

Files

References
  • references/feature_catalog.md — Complete catalog of ~40 features with formulas, lookbacks, stationarity status, and interpretation notes
  • references/pitfalls.md — Common mistakes in trading feature engineering: lookahead bias, overfitting, survivorship bias, data snooping, non-stationarity
Scripts
  • scripts/build_features.py — Compute 25+ features from OHLCV data with stationarity testing and quality reporting. Supports demo mode with synthetic data or live data via Birdeye API.
  • scripts/feature_importance.py — Rank features by predictive power using tree-based importance and permutation importance. Identifies redundant features via correlation analysis.

© agiprolabs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/feature-engineering of agiprolabs/claude-trading-skills.

  • SKILL.md
  • references/feature_catalog.md
  • references/pitfalls.md
  • scripts/build_features.py
  • scripts/feature_importance.py

Open the folder on GitHubat commit 981e1d7

Compare with similar skills

Feature Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Feature Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Feature Engineering this skillagiprolabs/claude-trading-skills410—~2.7kAutomated safety check: PassMIT
Metamask Agent Walletnirholas/three.ws226—~4.5kAutomated safety check: PassMIT
Crypto Price Data Guidenirholas/three.ws226—~2.7kAutomated safety check: PassMIT
QuantMind Training Config Generatorqusong0627/QuantMind1.7k—~1.5kAutomated safety check: PassAGPL-3.0
Machine Learning Trading StrategyHKUDS/Vibe-Trading35k—~3.2kAutomated safety check: PassMIT
Longbridge Quanthelsome/folio2691 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Metamask Agent Wallet

    nirholas/three.ws

    A skill your agent uses when the user asks anything about blockchain wallets, transactions, signing, token transfers, supported chains, wallet balances, perpetual futures trading, prediction…

    226 GitHub stars~4.5k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Crypto Price Data Guide

    nirholas/three.ws

    How cryptocurrency prices work — the complete data pipeline from CEX order books and DEX pools through oracle networks to aggregators.

    226 GitHub stars~2.7k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Turns a plain-language model training request into a validated QuantMind training config file that can be imported from the Model Training page.

    1.7k GitHub stars~1.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Trains scikit-learn models with walk-forward validation on features from OHLCV data to predict return direction and turn the predictions into trading signals.

    35k GitHub stars~3.2k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Longbridge Quant

    helsome/folio

    Quantitative strategy frameworks: pairs trading/cointegration, volatility regime strategies, seasonality/calendar effects, multi-factor models (IC/IR), factor research and screening, correlation…

    269 GitHub starsUsed in 1 repo~1.6k tokens
    Data & AnalyticsAuto-check passed
  • Quant Validation

    avelikiy/great_cto

    The methods a financial-ML result has to survive before it is evidence — purged cross-validation with an embargo, triple-barrier labelling, sample uniqueness under overlapping labels, fractional…

    103 GitHub starsUsed in 1 repo~1.9k tokens
    Business, Finance & HRAuto-check passed

More from agiprolabs/claude-trading-skills

All 68 skills in this repo
  • Backtrader

    agiprolabs/claude-trading-skills

    Event-driven backtesting with bar-by-bar execution, complex order types, multiple analyzers, and custom indicators

    410 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Birdeye API

    agiprolabs/claude-trading-skills

    Solana token market data via Birdeye — prices, OHLCV, trades, token metadata, security checks, and trader activity

    410 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Coingecko API

    agiprolabs/claude-trading-skills

    Broad crypto market data from CoinGecko covering 13,000+ tokens.

    410 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Cointegration Analysis

    agiprolabs/claude-trading-skills

    Cointegration testing for pairs trading using Engle-Granger, Johansen, and rolling stability analysis

    410 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Copy Trading

    agiprolabs/claude-trading-skills

    Wallet evaluation, monitoring, and copy-trade strategy design for Solana DEX trading

    410 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Correlation Analysis

    agiprolabs/claude-trading-skills

    Cross-asset correlation analysis including rolling correlation, hierarchical clustering, tail dependence, and regime-dependent correlation

    410 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Feature Engineering

What does Feature Engineering do?

Feature construction from market data for ML trading models including price, volume, on-chain, and microstructure features. Feature Engineering is an agent skill from agiprolabs/claude-trading-skills.

When should I use Feature Engineering?

Feature Engineering fits situations like: tasks that involve Machine learning; tasks that involve Trading and backtesting; tasks that involve Smart contracts.

How do I install Feature Engineering in Claude Code?

Run `npx skills add agiprolabs/claude-trading-skills --skill feature-engineering -a claude-code`. Or copy the skill folder (skills/feature-engineering in agiprolabs/claude-trading-skills) into .claude/skills/feature-engineering in your project. Claude Code loads it when a task matches its description.

How do I install Feature Engineering in Codex?

Run `npx skills add agiprolabs/claude-trading-skills --skill feature-engineering -a codex`. Or copy the skill folder (skills/feature-engineering in agiprolabs/claude-trading-skills) into .agents/skills/feature-engineering in your project. Codex loads it when a task matches its description.

Can I use Feature Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agiprolabs/claude-trading-skills --skill feature-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/feature-engineering, .gemini/skills/feature-engineering, .github/skills/feature-engineering and .opencode/skills/feature-engineering in your project.

What does Feature Engineering need to run?

Going by SKILL.md and its folder, Feature Engineering needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Feature Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Feature Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Feature Engineering use?

Feature Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Feature Engineering use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.

What are the alternatives to Feature Engineering?

Skills that share tags, products or a category with Feature Engineering: Metamask Agent Wallet (nirholas/three.ws, 226 stars), Crypto Price Data Guide (nirholas/three.ws, 226 stars), QuantMind Training Config Generator (qusong0627/QuantMind, 1.7k stars) and Machine Learning Trading Strategy (HKUDS/Vibe-Trading, 35k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Feature Engineering?

agiprolabs (a GitHub user) maintains it in agiprolabs/claude-trading-skills, which has 410 GitHub stars. The repository holds 68 skills in this directory. The repository was last updated on September 3, 2026.

Source: agiprolabs/claude-trading-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.