Agent skill

Factor Evaluation

by minihellboy in minihellboy/factorminer

Evaluate a factor library — recompute Information Coefficient (IC), ICIR, win rate, and turnover on held-out data, and surface train→test decay.

MITAuto-check passed

Install Factor Evaluation

skills CLI
$ npx skills add minihellboy/factorminer --skill factor-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install minihellboy/factorminer factor-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/minihellboy/factorminer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/integrations/factor-researcher/plugin/skills/factor-evaluation .claude/skills/factor-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
factor-evaluation
GitHub stars
123
Token cost
~605 tokens
SKILL.md length
250 words
Files
2 (incl. references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Evaluate a factor library — recompute Information Coefficient (IC), ICIR, win rate, and turnover on held-out data, and surface train→test decay.

  • Works in 4 steps: Recompute metrics → Read the table → Check decay → …
  • Judge how good a mined library actually is out of sample
  • SKILL.md covers Workflow, Interpreting the numbers and Guardrails
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Factor Evaluation is an agent skill from minihellboy/factorminer. Evaluate a factor library — recompute Information Coefficient (IC), ICIR, win rate, and turnover on held-out data, and surface train→test decay. Use to judge how good a mined library actually is out of sample. Triggers on "evaluate factors", "compute IC", "how good is this library", "factor metrics", "ICIR", "is this factor overfit", "out-of-sample".

Its SKILL.md is about 610 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/metrics.md`).

The repository describes itself as: A Self-Evolving Agent with Skills and Experience Memory for Financial Alpha Discovery. The licence is MIT.

When your agent uses it

  • Judge how good a mined library actually is out of sample
  • Evaluate factors
  • How good is this library
  • Is this factor overfit

Example prompts

  • “evaluate factors”
  • “compute IC”
  • “how good is this library”
  • “/factor-evaluation”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Recompute metrics
  2. Read the table
  3. Check decay
  4. Rank the survivors

What it can do on your machine

Read from SKILL.md and the folder at commit 75e0560. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Factor Evaluation loads about 605 tokens when it runs, and up to ~1.2k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 250 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~605
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from minihellboy/factorminer at commit 75e0560, republished under its MIT licence (© minihellboy). 250 words, ~605 tokens.

Download SKILL.mdSave it as .claude/skills/factor-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
factor-evaluation
description
Evaluate a factor library — recompute Information Coefficient (IC), ICIR, win rate, and turnover on held-out data, and surface train→test decay. Use to judge how good a mined library actually is out of sample. Triggers on "evaluate factors", "compute IC", "how good is this library", "factor metrics", "ICIR", "is this factor overfit", "out-of-sample".

Factor Evaluation

Mining proposes factors; evaluation decides whether to believe them. This skill recomputes a library's metrics on a chosen split and exposes overfitting.

See references/metrics.md for precise metric definitions (IC vs. paper-IC, ICIR, redundancy correlation).

Workflow

1. Recompute metrics
bash
factorminer evaluate output/run1/factor_library.json \
  --data path/to/market_data.csv \
  --period test

--period selects the split: train, test, or both. Always lead with test — in-sample IC is not evidence.

2. Read the table

The output table reports, per factor: IC Mean, Paper IC, Abs IC, Paper ICIR, Win%, and Turnover. The summary block gives library-level means and the IC range.

3. Check decay
bash
factorminer evaluate output/run1/factor_library.json --data market_data.csv --period both

--period both adds a decay table (train Paper IC → test Paper IC → delta). A large negative delta is the signature of an overfit factor. Report decay honestly; do not quote the train number as the headline.

4. Rank the survivors

To shortlist the strongest signals only:

bash
factorminer evaluate output/run1/factor_library.json --data market_data.csv --period test --top-k 10

The top-K-by-IC table is the signal shortlist — the natural handoff to a research-idea workflow that wants to know which quantitative signals are currently working. The MCP screen_factors tool returns this same shortlist directly.

Interpreting the numbers

  • IC ≈ 0.03–0.05 out of sample is a respectable single factor on liquid universes.
  • ICIR matters more than IC: a small but stable IC beats a large erratic one.
  • High turnover quietly erases IC once costs are applied — carry it into factor-backtest.

Guardrails

  • Never present train metrics as the result. The deliverable is the test number.
  • If every factor decays to ~0 on test, the library failed — say so. Do not search for a split that flatters it.

© minihellboy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in integrations/factor-researcher/plugin/skills/factor-evaluation of minihellboy/factorminer.

  • SKILL.md
  • references/metrics.md

Open the folder on GitHubat commit 75e0560

Compare with similar skills

Factor Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Factor Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Factor Evaluation this skillminihellboy/factorminer123—~605Automated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k12 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence
Factor Research with IC and IRHKUDS/Vibe-Trading35k—~2.1kAutomated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check passed
  • Evaluates factors across many instruments with IC and IR statistics and quantile backtests, then guides screening and weighting; uses the factor_analysis tool with factor and return CSVs.

    35k GitHub stars~2.1k tokensUpdated yesterday
    Business, Finance & HRAuto-check passed
  • Agent Evaluation Reporting

    sickn33/agentic-awesome-skills

    A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Agent WorkflowsAuto-check passed

More from minihellboy/factorminer

  • Factor Mining

    minihellboy/factorminer

    Discover alpha factors by running the FactorMiner research engine — the paper-faithful Ralph loop or the enhanced Helix loop (causal validation, regime conditioning, multi-specialist debate…

    123 GitHub stars~781 tokensUpdated 9 days ago
    Auto-check passed
  • Factor Backtest

    minihellboy/factorminer

    Combine a factor library into a composite signal and quintile-backtest it under transaction costs — long-short return, monotonicity, turnover, and tearsheets.

    123 GitHub stars~599 tokensUpdated 9 days ago
    Auto-check passed
  • Factor Benchmark

    minihellboy/factorminer

    Run FactorMiner benchmark workflows — the Table 1 Top-K freeze benchmark, memory and strategy ablations, transaction-cost pressure tests, and the full suite.

    123 GitHub stars~576 tokensUpdated 9 days ago
    Auto-check passed
  • Factor Data

    minihellboy/factorminer

    Validate, resample, and ingest market data for factor mining.

    123 GitHub stars~856 tokensUpdated 9 days ago
    Auto-check passed
  • Factor Report

    minihellboy/factorminer

    Generate static reports, tearsheets, and exports from FactorMiner artifacts — markdown/HTML research notes, plots, and library exports (JSON/CSV/formulas).

    123 GitHub stars~588 tokensUpdated 9 days ago
    Auto-check passed
  • Research Ingestion

    minihellboy/factorminer

    Absorb external research reports/papers into structured, retrievable hypothesis cues via FactorMiner's Report-to-Memory Absorption (RMA) service — an OHLCV-eligibility gate, a mechanism-family…

    123 GitHub stars~884 tokensUpdated 9 days ago
    Auto-check passed

Questions about Factor Evaluation

What does Factor Evaluation do?

Evaluate a factor library — recompute Information Coefficient (IC), ICIR, win rate, and turnover on held-out data, and surface train→test decay. Factor Evaluation is an agent skill from minihellboy/factorminer. Evaluate a factor library — recompute Information Coefficient (IC), ICIR, win rate, and turnover on held-out data, and surface train→test decay.

When should I use Factor Evaluation?

Factor Evaluation fits situations like: judge how good a mined library actually is out of sample; evaluate factors; how good is this library; is this factor overfit.

How do I install Factor Evaluation in Claude Code?

Run `npx skills add minihellboy/factorminer --skill factor-evaluation -a claude-code`. Or copy the skill folder (integrations/factor-researcher/plugin/skills/factor-evaluation in minihellboy/factorminer) into .claude/skills/factor-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Factor Evaluation in Codex?

Run `npx skills add minihellboy/factorminer --skill factor-evaluation -a codex`. Or copy the skill folder (integrations/factor-researcher/plugin/skills/factor-evaluation in minihellboy/factorminer) into .agents/skills/factor-evaluation in your project. Codex loads it when a task matches its description.

Can I use Factor Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add minihellboy/factorminer --skill factor-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/factor-evaluation, .gemini/skills/factor-evaluation, .github/skills/factor-evaluation and .opencode/skills/factor-evaluation in your project.

What does Factor Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Factor Evaluation is instructions for the agent only.

Does Factor Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Factor Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Factor Evaluation use?

Factor Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Factor Evaluation use?

About 605 tokens (SKILL.md is roughly 2.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 572 tokens, read only when the agent opens those files.

What are the alternatives to Factor Evaluation?

Skills that share tags, products or a category with Factor Evaluation: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Evaluators (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Factor Evaluation?

minihellboy (a GitHub user) maintains it in minihellboy/factorminer, which has 123 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 28, 2026.

Source: minihellboy/factorminer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.