Agent skill

Correlation and Cointegration Analysis

by HKUDS in HKUDS/Vibe-Trading

Finds co-moving assets and tests them for cointegration, with workflows for correlation studies, sector clustering, hedge ratios and pair-trading signals.

MITAuto-check passedBusiness, Finance & HR

Install Correlation and Cointegration Analysis

skills CLI
$ npx skills add HKUDS/Vibe-Trading --skill correlation-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/Vibe-Trading correlation-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/Vibe-Trading.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent/src/skills/correlation-analysis .claude/skills/correlation-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
correlation-analysis
GitHub stars
35k
Token cost
~10k tokens
SKILL.md length
1,274 words
Files
1
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Finds co-moving assets and tests them for cointegration, with workflows for correlation studies, sector clustering, hedge ratios and pair-trading signals.

  • Works in 8 steps: Prices vs returns: Use price series,… → Data alignment: Cross-market analysis… → Cointegration is not the same as high… → …
  • Screening a universe of stocks for candidates to pair with a target
  • SKILL.md covers Overview, Mode 1: Co-Movement Discovery, Mode 2: Deep… and Mode 3: Sector Clustering, plus 5 more sections
  • Calls pip

What it does

The skill describes four correlation modes: co-movement discovery, a deep return-correlation study of two assets, sector clustering and realized correlation. Discovery pulls daily returns for a target and a set of candidates, computes Pearson and Spearman correlations, ranks them and keeps the top few. Pairs above 0.8 go to cointegration testing, and correlations below 0.6 are usually judged unsuitable for pairs trading.

The deep study covers several correlation coefficients, beta and R-squared, rolling correlation and spread z-score, with a guide to when Pearson, Spearman or Kendall fits best. Further parts cover Engle-Granger and Johansen cointegration tests, half-life, a Kalman dynamic hedge ratio, cross-market linkage analysis and the path from analytics to pair-trading signals. The code examples use Python with pandas, numpy, scipy and statsmodels.

When your agent uses it

  • Screening a universe of stocks for candidates to pair with a target
  • Testing a pair for cointegration before building a spread strategy
  • Choosing between Pearson, Spearman and Kendall for return series
  • Estimating a dynamic hedge ratio and the half-life of a spread

Example prompts

  • “Find the ten assets most correlated with NVDA over the last two years and test the top ones for cointegration.”
  • “Run Engle-Granger and Johansen tests on KO and PEP and estimate the half-life of the spread.”
  • “Cluster these energy stocks by return correlation to see which sector groups form.”

Requirements

  • Python with pandas, numpy, scipy and statsmodels
  • Daily price or return data for the assets

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Prices vs returns: Use price series, which are non-stationary, for cointegration tests; use return series, which are stationary, for…
  2. Data alignment: Cross-market analysis must handle holiday mismatches with an inner join. Do not forward-fill missing trading days, or you…
  3. Cointegration is not the same as high correlation: Two series can have Pearson < 0.3 and still be cointegrated, and the reverse can also…
  4. Out-of-sample validation: If a pair is selected using cointegration on the first N years, you must verify whether the relationship…
  5. Crisis-period risk: Correlation jumps in crises, and both legs in a pair can crash together. Stop thresholds should be tighter than in…
  6. China A-share specifics: China A-shares contain many non-trading days due to holidays and suspensions. Date alignment is especially…
  7. Multiple testing: When testing N asset pairs simultaneously, use Benjamini-Hochberg FDR adjustment on p-values. Otherwise false positives…
  8. Kalman tuning: Tune delta with grid search plus out-of-sample validation. Do not rely blindly on the default value.

What it can do on your machine

Read from SKILL.md and the folder at commit 7f6908b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Correlation and Cointegration Analysis loads about 10k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 1,274 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/Vibe-Trading at commit 7f6908b, republished under its MIT licence (© HKUDS). 1,274 words, ~10,051 tokens.

Download SKILL.mdSave it as .claude/skills/correlation-analysis/SKILL.md (or your agent's skills folder).
name
correlation-analysis
description
Correlation and cointegration analysis — co-movement discovery, deep return-correlation analysis, sector clustering, realized correlation, Engle-Granger / Johansen cointegration, half-life, Kalman dynamic hedge ratio, cross-market linkage analysis, and pair-trading signal generation
category
analysis

Correlation and Cointegration Analysis

Overview

Correlation analysis is a foundational tool for pairs trading, portfolio construction, and risk management. This skill covers four analysis modes (co-movement discovery / return-correlation deep dive / sector clustering / realized correlation), a full cointegration-testing framework, cross-market linkage analysis, and the complete workflow from analytics to pair-trading signals.


Mode 1: Co-Movement Discovery

Use case: Given a target asset, scan a universe for highly correlated assets and build a candidate pool with similar industry or factor exposure, for use in pairs trading or substitute identification.

Workflow
1. Pull daily return series for the target asset and N candidates
2. Compute Pearson / Spearman correlations between the target and each candidate
3. Rank by correlation in descending order and keep Top-K (usually K=10-20)
4. Run cointegration tests on the Top-K set to retain pairs with real long-run equilibrium
5. Output the candidate pool and a correlation summary
python
import pandas as pd
import numpy as np
from scipy.stats import pearsonr, spearmanr

def scan_correlated_assets(
    target_returns: pd.Series,
    universe_returns: pd.DataFrame,
    top_k: int = 20,
    min_corr: float = 0.5,
    method: str = "pearson",
) -> pd.DataFrame:
    """Scan for assets that are highly correlated with the target asset.

    Args:
        target_returns: Daily return series for the target asset
        universe_returns: Candidate-universe return matrix, columns are symbols
        top_k: Number of top candidates to return
        min_corr: Minimum absolute-correlation threshold
        method: "pearson" or "spearman"

    Returns:
        A DataFrame containing symbol / corr / p_value / rank
    """
    aligned = universe_returns.dropna(axis=1, how="any")
    aligned, target_aligned = aligned.align(target_returns, join="inner", axis=0)

    results = []
    for col in aligned.columns:
        if method == "spearman":
            corr, p = spearmanr(target_aligned, aligned[col])
        else:
            corr, p = pearsonr(target_aligned, aligned[col])
        results.append({"symbol": col, "corr": corr, "p_value": p})

    df = pd.DataFrame(results)
    df = df[df["corr"].abs() >= min_corr].sort_values("corr", ascending=False)
    df["rank"] = range(1, len(df) + 1)
    return df.head(top_k).reset_index(drop=True)

Screening guidance:

CorrelationConclusionFollow-up Action
> 0.8Strong same-direction co-movementSend to the cointegration test queue
0.6 - 0.8Moderate co-movementCheck industry / factor alignment before cointegration
< 0.6Weak correlationUsually unsuitable for pairs trading
Negative and < -0.6Strong inverse co-movementCan be used in hedged portfolios, but be careful with spread direction

Mode 2: Deep Return-Correlation Analysis

Use case: Run a full bivariate correlation study on two assets, including multiple correlation coefficients, Beta / R², rolling correlation, and spread Z-Score.

Core Metrics
python
import statsmodels.api as sm
from scipy.stats import pearsonr, spearmanr, kendalltau

def bivariate_correlation_analysis(
    y: pd.Series,
    x: pd.Series,
    rolling_window: int = 60,
) -> dict:
    """Run deep correlation analysis for two assets.

    Args:
        y: Daily return series of asset A
        x: Daily return series of asset B
        rolling_window: Rolling-window length in trading days

    Returns:
        Dict of correlation statistics
    """
    # Align the two series.
    df = pd.concat([y.rename("y"), x.rename("x")], axis=1).dropna()
    y_clean, x_clean = df["y"], df["x"]

    # Static correlations.
    pearson_r, pearson_p = pearsonr(y_clean, x_clean)
    spearman_r, spearman_p = spearmanr(y_clean, x_clean)
    kendall_r, kendall_p = kendalltau(y_clean, x_clean)

    # OLS: y = α + β·x
    x_const = sm.add_constant(x_clean)
    ols = sm.OLS(y_clean, x_const).fit()
    beta = ols.params["x"]
    alpha = ols.params["const"]
    r_squared = ols.rsquared

    # Rolling Pearson correlation.
    rolling_corr = y_clean.rolling(rolling_window).corr(x_clean)

    # Spread and Z-Score using the hedge ratio.
    spread = y_clean - beta * x_clean
    spread_mean = spread.rolling(rolling_window).mean()
    spread_std = spread.rolling(rolling_window).std()
    z_score = (spread - spread_mean) / spread_std

    return {
        "pearson": {"r": round(pearson_r, 4), "p": round(pearson_p, 6)},
        "spearman": {"r": round(spearman_r, 4), "p": round(spearman_p, 6)},
        "kendall": {"r": round(kendall_r, 4), "p": round(kendall_p, 6)},
        "beta": round(beta, 4),
        "alpha": round(alpha, 6),
        "r_squared": round(r_squared, 4),
        "rolling_corr": rolling_corr,
        "spread": spread,
        "z_score": z_score,
        "spread_mean": spread_mean,
        "spread_std": spread_std,
    }
Correlation-Coefficient Selection Guide
CoefficientAssumptionBest Use CaseNot Suitable When
PearsonLinear, approximately normalReturn seriesHeavy tails / many outliers
SpearmanMonotonic relationshipRanking / quantile analysis, many outliersWhen magnitude information matters
KendallOrder consistencySmall samples, unknown distributionLarge samples due to slower computation

Practical rule in finance: Usually report all three coefficients. If Pearson and Spearman differ by more than 0.1, the relationship is likely nonlinear or heavy-tailed, and Spearman should carry more weight.


Mode 3: Sector Clustering

Use case: Run hierarchical clustering on the correlation matrix of N assets to discover sector structure, check portfolio diversification, and identify similar assets.

python
import numpy as np
import pandas as pd
from scipy.cluster.hierarchy import linkage, dendrogram, fcluster
from scipy.spatial.distance import squareform
import matplotlib.pyplot as plt
import seaborn as sns

def sector_clustering(
    returns: pd.DataFrame,
    method: str = "ward",
    n_clusters: int = 5,
    figsize: tuple = (12, 10),
) -> dict:
    """Run sector clustering analysis.

    Args:
        returns: Multi-asset daily return matrix, columns are symbols
        method: Linkage method: "ward" / "complete" / "average"
        n_clusters: Target number of clusters
        figsize: Heatmap size

    Returns:
        Dict containing the correlation matrix, cluster labels, and figure objects
    """
    # 1. Correlation matrix
    corr_matrix = returns.corr(method="pearson")

    # 2. Distance matrix where distance = 1 - |correlation|
    distance_matrix = 1 - corr_matrix.abs()
    condensed = squareform(distance_matrix.values, checks=False)

    # 3. Hierarchical clustering
    linkage_matrix = linkage(condensed, method=method)
    labels = fcluster(linkage_matrix, n_clusters, criterion="maxclust")
    cluster_df = pd.DataFrame({"symbol": corr_matrix.columns, "cluster": labels})

    # 4. Heatmap sorted by cluster
    order = cluster_df.sort_values("cluster").index
    sorted_corr = corr_matrix.iloc[order, order]

    fig_heatmap, ax = plt.subplots(figsize=figsize)
    sns.heatmap(
        sorted_corr,
        cmap="RdYlGn",
        center=0,
        vmin=-1,
        vmax=1,
        annot=len(corr_matrix) <= 20,
        fmt=".2f",
        ax=ax,
        cbar_kws={"label": "Pearson correlation"},
    )
    ax.set_title(f"Correlation Heatmap ({method.upper()} clustering order)")

    # 5. Dendrogram
    fig_dendro, ax2 = plt.subplots(figsize=(figsize[0], 6))
    dendrogram(
        linkage_matrix,
        labels=list(corr_matrix.columns),
        ax=ax2,
        leaf_rotation=90,
        color_threshold=0,
    )
    ax2.set_title(f"Hierarchical Dendrogram ({method.upper()} linkage)")
    ax2.set_ylabel("Distance")

    return {
        "corr_matrix": corr_matrix,
        "cluster_labels": cluster_df,
        "linkage_matrix": linkage_matrix,
        "fig_heatmap": fig_heatmap,
        "fig_dendrogram": fig_dendro,
        "n_clusters": n_clusters,
    }
Comparison of Three Linkage Methods
MethodFeatureBest Use CaseWeakness
WardMinimizes within-cluster variance, gives compact clustersDefault recommendation, stock-sector discoveryWorks best for spherical clusters, weaker for irregular shapes
CompleteUses maximum pairwise distance, conservativeWhen high within-cluster similarity is requiredCan produce elongated clusters
AverageUses average distance, compromise approachGeneral analysis where compactness is not the top prioritySensitive to noise

Mode 4: Realized Correlation

Authoritative Source: This section is the single source of truth for regime classification rules. Other files (e.g., agent/src/agent/context.py Layer 3, docs/trade-attribution-design.md) reference this definition.

Use case: Compute rolling correlation time series and analyze conditional correlation by market regime (bull / bear / high-volatility) to discover how correlation evolves dynamically.

python
def realized_correlation(
    y: pd.Series,
    x: pd.Series,
    benchmark: pd.Series,
    windows: list = [20, 60, 120],
    vol_window: int = 20,
    vol_threshold: float = 1.5,
) -> dict:
    """Rolling realized correlation plus regime-conditional correlation.

    Args:
        y, x: Daily return series of two assets
        benchmark: Daily return series of the benchmark index used for regime labeling
        windows: List of rolling windows in trading days
        vol_window: Volatility window
        vol_threshold: High-vol threshold as a multiple of average vol

    Returns:
        Rolling correlation series and conditional-correlation summary
    """
    df = pd.concat([y.rename("y"), x.rename("x"),
                    benchmark.rename("bm")], axis=1).dropna()

    # Rolling correlation time series.
    rolling_corrs = {}
    for w in windows:
        rolling_corrs[f"roll_{w}d"] = df["y"].rolling(w).corr(df["x"])

    # Regime labels.
    # N-day rolling mean (N = 252 for US equity; use 244 for A-share, 365 for crypto)
    bm_ret_252 = df["bm"].rolling(252).mean()
    bm_vol = df["bm"].rolling(vol_window).std()
    # N-day rolling mean of volatility (N = 252 for US equity; use 244 for A-share, 365 for crypto)
    bm_vol_mean = bm_vol.rolling(252).mean()

    df["regime"] = "sideways"
    df.loc[df["bm"] > bm_ret_252, "regime"] = "bull"
    df.loc[df["bm"] < -bm_ret_252.abs(), "regime"] = "bear"
    df.loc[bm_vol > bm_vol_mean * vol_threshold, "regime"] = "high_vol"

    # Conditional correlation.
    cond_corr = {}
    for regime in ["bull", "bear", "sideways", "high_vol"]:
        mask = df["regime"] == regime
        if mask.sum() >= 30:
            r, p = pearsonr(df.loc[mask, "y"], df.loc[mask, "x"])
            cond_corr[regime] = {"corr": round(r, 4), "p": round(p, 6), "n": int(mask.sum())}
        else:
            cond_corr[regime] = {"corr": None, "p": None, "n": int(mask.sum())}

    return {
        "rolling_corrs": pd.DataFrame(rolling_corrs),
        "regime_labels": df["regime"],
        "conditional_corr": cond_corr,
    }

Market-Adaptive Window (N): The rolling window for regime classification adapts to market type:

  • US equity: N = 252 (standard trading days per year)
  • A-share (China): N = 244 (fewer trading days due to holidays)
  • Crypto: N = 365 (24/7 market)

The default value 252 in the function signature targets US equity. Callers for other markets should adjust the rolling windows accordingly.

Fallback for Short Data: If available data length is shorter than N days (the market-equivalent annual window), regime classification falls back to fixed-threshold definitions:

  • Bull: cumulative benchmark return over the available period > +10%
  • Bear: cumulative benchmark return over the available period < -10%
  • High-vol / Sideways: not applicable in fallback mode
Typical Correlation Behavior by Market Regime
Market RegimeEquity-Equity CorrelationEquity-Bond CorrelationA-Share Characteristic
BullMedium (0.4-0.6)Low or negativeSmall-cap names tend to move together strongly
BearHigh (0.7-0.9)Negative (safe-haven effect)Broad selloff, correlation jumps sharply
High volatilityVery high (0.8+)NegativeIn crises, correlation often converges toward 1
SidewaysLow (0.2-0.4)Near zeroStock dispersion rises, ideal for pairs trading

Cointegration Analysis

Correlation measures the degree of co-movement. Cointegration measures whether a long-run equilibrium relationship exists. High correlation does not guarantee cointegration, and low correlation does not rule it out.

Engle-Granger Two-Step Method

Suitable for two-variable pairs, quick and intuitive.

python
from src.quantlib.timeseries import adf_test, cointegration_test, find_hedge_ratio

def engle_granger_coint(
    y: pd.Series,
    x: pd.Series,
    significance: float = 0.05,
) -> dict:
    """Run the Engle-Granger two-step cointegration test on one ordered pair.

    H0: No cointegration relationship exists (residuals contain a unit root).
    Both steps come from the tested quantlib layer rather than a local OLS
    fit: ``find_hedge_ratio`` estimates ``y = α + β·x`` (step 1) and
    ``cointegration_test`` runs the Engle-Granger test itself (step 2).

    Args:
        y, x: Two price series sharing one index. These must be
            non-stationary series, usually prices rather than returns.
        significance: Significance level

    Returns:
        Test results and spread series
    """
    fit = find_hedge_ratio(y, x)
    spread = y - fit["hedge_ratio"] * x - fit["intercept"]
    coint = cointegration_test(y, x, significance=significance)
    adf = adf_test(spread.dropna(), significance=significance)

    return {
        "method": "Engle-Granger",
        "is_cointegrated": coint["is_cointegrated"],
        "coint_p": round(coint["p_value"], 6),
        "coint_stat": round(coint["test_statistic"], 4),
        "critical_values": coint["critical_values"],
        "hedge_ratio": round(fit["hedge_ratio"], 6),
        "intercept": round(fit["intercept"], 6),
        "half_life": fit["half_life"],
        "spread": spread,
        "adf_on_spread": {"stat": round(adf["adf_statistic"], 4), "p": round(adf["p_value"], 6)},
    }

find_hedge_ratio and cointegration_test refuse two legs of different lengths or on different indices instead of zipping them positionally — inner-join the two price series on date before calling.

Note: Engle-Granger can detect only one cointegrating vector, and the test result depends on the ordering of y and x. In practice, test both directions and keep the direction with the smaller p-value.

Johansen Cointegration Test for Multiple Variables

Suitable for three or more assets and for estimating the number of cointegrating vectors (rank).

python
from statsmodels.tsa.vector_ar.vecm import coint_johansen

def johansen_coint(
    prices: pd.DataFrame,
    det_order: int = 0,
    k_ar_diff: int = 1,
) -> dict:
    """Run the Johansen cointegration test.

    Args:
        prices: Multi-asset price matrix, columns are symbols.
            The series must be non-stationary.
        det_order: Deterministic term. -1=no intercept, 0=intercept, 1=trend
        k_ar_diff: Number of lagged differences in the VAR, usually 1-5 chosen by AIC

    Returns:
        Trace-test and max-eigenvalue-test results
    """
    result = coint_johansen(prices.dropna(), det_order=det_order, k_ar_diff=k_ar_diff)
    n = prices.shape[1]

    # Trace test.
    trace_results = []
    for i in range(n):
        trace_results.append({
            "H0_rank_leq": i,
            "trace_stat": round(result.lr1[i], 4),
            "crit_10pct": result.cvt[i, 0],
            "crit_5pct": result.cvt[i, 1],
            "crit_1pct": result.cvt[i, 2],
            "reject_5pct": result.lr1[i] > result.cvt[i, 1],
        })

    # Max-eigenvalue test.
    maxeig_results = []
    for i in range(n):
        maxeig_results.append({
            "H0_rank_eq": i,
            "maxeig_stat": round(result.lr2[i], 4),
            "crit_10pct": result.cvm[i, 0],
            "crit_5pct": result.cvm[i, 1],
            "crit_1pct": result.cvm[i, 2],
            "reject_5pct": result.lr2[i] > result.cvm[i, 1],
        })

    # Cointegrating vectors, normalized.
    coint_vectors = pd.DataFrame(
        result.evec[:, :sum(r["reject_5pct"] for r in trace_results)],
        index=prices.columns,
    )

    return {
        "method": "Johansen",
        "n_coint_vectors_trace": sum(r["reject_5pct"] for r in trace_results),
        "trace_test": pd.DataFrame(trace_results),
        "maxeig_test": pd.DataFrame(maxeig_results),
        "coint_vectors": coint_vectors,
        "eigenvalues": result.eig,
    }

Johansen rank interpretation rules:

Start the trace test from H0: rank=0 and move upward.
The first rank that cannot be rejected is the estimated cointegration rank.

rank = 0   → no cointegration
rank = 1   → one cointegrating vector (most common, long-run equilibrium for a pair)
rank = k-1 → k-1 cointegrating vectors (system is tightly linked)
rank = k   → the series themselves are stationary, so cointegration is not needed
Half-Life Calculation

Half-life measures how long a spread takes to mean-revert after deviating from equilibrium. It is a practical reference for expected holding period in pairs trading.

python
from src.quantlib.timeseries import compute_half_life

half_life = compute_half_life(spread)
# Regresses ΔSpread_t on Spread_{t-1}; half-life = -ln(2) / λ, in observation
# periods (trading days for daily bars). Returns inf when λ >= 0, i.e. the
# spread does not mean-revert. Raises ValueError on fewer than 3 usable
# lag/delta pairs or on a spread that never moves.

find_hedge_ratio already reports this as half_life, so a pair screened through engle_granger_coint above carries it in its result.

Half-life reference ranges:

Half-LifeMeaningTrading Guidance
< 5 daysExtremely fast reversionIntraday or overnight trading, friction cost matters
5-20 daysFast reversionIdeal range for short-term pairs trading
20-60 daysMedium-speed reversionMedium-term holding, rolling windows 60-120 days
60-180 daysSlow reversionLong holding period, monitor cointegration stability
> 180 daysNear random walkHigh pairs-trading risk, use cautiously
Show full SKILL.md (487 more words)Show less
Kalman Filter Dynamic Hedge Ratio

Static OLS hedge ratios cannot capture gradual drift in the cointegration relationship. A Kalman filter provides a continuously updated dynamic hedge ratio.

python
import numpy as np

def kalman_hedge_ratio(
    y: pd.Series,
    x: pd.Series,
    delta: float = 1e-4,
    vt: float = 1.0,
) -> pd.DataFrame:
    """Estimate a dynamic hedge ratio with a Kalman filter.

    State equation:
        β_t = β_{t-1} + w_t,  w ~ N(0, Q)
    Observation equation:
        y_t = β_t · x_t + v_t,  v ~ N(0, R)

    Args:
        y: Price series of asset A
        x: Price series of asset B
        delta: State-noise intensity. Larger means faster hedge-ratio adaptation
        vt: Observation-noise variance

    Returns:
        DataFrame containing the dynamic hedge ratio and spread
    """
    n = len(y)
    # State: [β, α] = hedge ratio + intercept
    Wt = delta / (1 - delta) * np.eye(2)
    Vt = vt

    # Initialization
    theta = np.zeros((n, 2))
    P = np.zeros((n, 2, 2))
    P[0] = np.eye(2)

    spread = np.zeros(n)
    spread[0] = float("nan")

    for t in range(1, n):
        F = np.array([x.iloc[t], 1.0])

        # Predict
        theta_pred = theta[t - 1]
        P_pred = P[t - 1] + Wt

        # Innovation
        innovation = y.iloc[t] - F @ theta_pred
        S = F @ P_pred @ F.T + Vt

        # Kalman gain
        K = P_pred @ F.T / S

        # Update
        theta[t] = theta_pred + K * innovation
        P[t] = (np.eye(2) - np.outer(K, F)) @ P_pred

        spread[t] = y.iloc[t] - theta[t, 0] * x.iloc[t] - theta[t, 1]

    return pd.DataFrame({
        "hedge_ratio": theta[:, 0],
        "intercept": theta[:, 1],
        "spread": spread,
    }, index=y.index)

Static vs dynamic hedge-ratio comparison:

MethodStrengthWeaknessBest Use Case
OLSSimple, stableCannot capture time variationShort-term stable pairs
Rolling OLSTime-varying, intuitiveWindow-sensitive, endpoint effectMedium-term pairs
Kalman FilterReal-time, continuous updatedelta is harder to tuneLong-term or structurally shifting pairs

Cross-Market Correlation

Correlation Across China A-Share Sectors
python
# Typical China A-share sector-correlation patterns
ASHARE_SECTOR_PATTERNS = {
    "strong_pairs_gt_0_7": [
        "Banks & insurance",
        "Baijiu & consumer staples",
        "New energy & solar",
        "Defense & aerospace",
    ],
    "medium_pairs_0_4_to_0_7": [
        "Pharma & consumer",
        "Technology & semiconductors",
        "Real estate & building materials",
    ],
    "low_or_negative_lt_0_3": [
        "Gold & technology",
        "Utilities & cyclicals",
        "Consumer & cyclicals",
    ],
}
Cross-Market Linkage Analysis
python
def cross_market_correlation(
    markets: dict,  # {"China A-shares": series, "Hong Kong": series, "crypto": series, "US": series}
    rolling_window: int = 60,
    lag_days: list = [0, 1, 2, 3],
) -> dict:
    """Cross-market correlation plus lead-lag analysis.

    Args:
        markets: Daily return series for each market
        rolling_window: Rolling window
        lag_days: List of lags to test

    Returns:
        Correlation matrix, lead-lag analysis, and rolling correlation
    """
    df = pd.DataFrame(markets).dropna()

    # Static correlation matrix
    static_corr = df.corr()

    # Lead-lag correlation to detect cross-market transmission
    lead_lag = {}
    mkt_names = list(markets.keys())
    for i, m1 in enumerate(mkt_names):
        for m2 in mkt_names[i + 1:]:
            pair_key = f"{m1}_{m2}"
            lead_lag[pair_key] = {}
            for lag in lag_days:
                if lag == 0:
                    r, _ = pearsonr(df[m1], df[m2])
                else:
                    r, _ = pearsonr(df[m1].iloc[lag:], df[m2].iloc[:-lag])
                lead_lag[pair_key][f"lag_{lag}d"] = round(r, 4)

    # Rolling correlation
    rolling_corrs = {}
    for i, m1 in enumerate(mkt_names):
        for m2 in mkt_names[i + 1:]:
            key = f"{m1}_{m2}"
            rolling_corrs[key] = df[m1].rolling(rolling_window).corr(df[m2])

    return {
        "static_corr": static_corr,
        "lead_lag": pd.DataFrame(lead_lag).T,
        "rolling_corrs": pd.DataFrame(rolling_corrs),
    }
Empirical Cross-Market Linkage Patterns
Market PairAverage CorrelationTransmission DirectionLag
China A-shares ↔ Hong Kong0.5-0.7Two-way, Hong Kong slightly leads0-1 day
China A-shares ↔ U.S. equities0.2-0.4U.S. leads overnight1 day
BTC ↔ ETH0.7-0.9Highly synchronous< 1 hour
China A-shares ↔ BTC0.0-0.2Mostly independent, except correlation spikes in crisesUnstable
U.S. equities ↔ BTC0.1-0.4U.S. leads through institutional capital flowsWithin 1 day
RMB exchange rate ↔ China A-shares-0.2 - 0.3RMB weakness → foreign outflows → China A-share weakness0-2 days
Impact of FX Factors on Cross-Market Correlation

Cross-market correlation analysis must distinguish between local-currency returns and FX-adjusted returns. Otherwise, exchange-rate moves can create spurious correlation or hide the true one.

python
def fx_adjusted_correlation(
    foreign_price: pd.Series,   # foreign-market price, denominated in foreign currency
    domestic_price: pd.Series,  # domestic-market price
    fx_rate: pd.Series,         # foreign currency / domestic currency, e.g. USD/CNY
) -> dict:
    """Cross-market correlation adjusted for FX effects.

    Args:
        foreign_price: Foreign-market price series in foreign currency
        domestic_price: Domestic-market price series in domestic currency
        fx_rate: FX series expressed as foreign / domestic

    Returns:
        Raw correlation vs FX-adjusted correlation
    """
    # Domestic-currency foreign return = foreign return + FX return
    foreign_ret = foreign_price.pct_change(fill_method=None)
    fx_ret = fx_rate.pct_change(fill_method=None)
    foreign_ret_cny = (1 + foreign_ret) * (1 + fx_ret) - 1

    domestic_ret = domestic_price.pct_change(fill_method=None)

    df = pd.concat([foreign_ret.rename("foreign_raw"),
                    foreign_ret_cny.rename("foreign_domestic"),
                    domestic_ret.rename("domestic"),
                    fx_ret.rename("fx")], axis=1).dropna()

    raw_corr, _ = pearsonr(df["foreign_raw"], df["domestic"])
    adj_corr, _ = pearsonr(df["foreign_domestic"], df["domestic"])
    fx_corr, _ = pearsonr(df["fx"], df["domestic"])

    return {
        "raw_corr_foreign_domestic": round(raw_corr, 4),
        "fx_adjusted_corr": round(adj_corr, 4),
        "fx_domestic_corr": round(fx_corr, 4),
        "fx_contribution": round(adj_corr - raw_corr, 4),
        "note": "fx_contribution > 0 means FX amplified cross-market correlation",
    }
Correlation Breakdown During Crises

During crises, equity correlation converges toward 1 and diversification breaks down. This is one of the central challenges in portfolio risk management.

python
def correlation_breakdown_test(
    returns: pd.DataFrame,
    crisis_threshold: float = -0.02,  # one-day benchmark drop threshold for crisis days
    benchmark_col: str = None,
    window: int = 20,
) -> dict:
    """Detect jumps in correlation during crisis periods.

    Args:
        returns: Multi-asset daily return matrix
        crisis_threshold: Benchmark return below this level defines a crisis day
        benchmark_col: Benchmark column name. If None, use cross-sectional mean return
        window: Window for rolling average correlation

    Returns:
        Comparison of correlation in normal periods vs crisis periods
    """
    if benchmark_col:
        bm = returns[benchmark_col]
    else:
        bm = returns.mean(axis=1)

    crisis_mask = bm < crisis_threshold
    normal_mask = ~crisis_mask

    # Average pairwise correlation for each period
    def avg_corr(df_subset: pd.DataFrame) -> float:
        if len(df_subset) < 5:
            return float("nan")
        c = df_subset.corr()
        upper = c.where(np.triu(np.ones(c.shape), k=1).astype(bool))
        return float(upper.stack().mean())

    crisis_corr = avg_corr(returns[crisis_mask])
    normal_corr = avg_corr(returns[normal_mask])

    # Rolling average correlation to detect structural change
    rolling_avg_corr = pd.Series(dtype=float, index=returns.index)
    for i in range(window, len(returns)):
        sub = returns.iloc[i - window:i]
        rolling_avg_corr.iloc[i] = avg_corr(sub)

    return {
        "normal_avg_corr": round(normal_corr, 4),
        "crisis_avg_corr": round(crisis_corr, 4),
        "corr_jump": round(crisis_corr - normal_corr, 4),
        "crisis_days": int(crisis_mask.sum()),
        "normal_days": int(normal_mask.sum()),
        "rolling_avg_corr": rolling_avg_corr,
    }

Pair-Trading Signal Generation

Full Workflow From Correlation to Signal
Step 1: Asset screening
  - Run scan_correlated_assets and keep candidate pairs with Pearson > 0.6
  - Run engle_granger_coint and keep pairs with p < 0.05

Step 2: Spread quality assessment
  - Run compute_half_life and keep pairs with half-life between 5 and 60 days
  - Test spread stationarity with ADF and require p < 0.05
  - Measure how often the absolute Z-Score exceeds 1.5 over the last 12 months
    to estimate trading frequency

Step 3: Hedge-ratio selection
  - Use static OLS for stable pairs
  - Use Kalman Filter for long-lived or drifting pairs

Step 4: Signal generation
  - Compute rolling Z-Score with a lookback of 2-3× half-life
  - Generate long / short / exit signals based on thresholds

Step 5: Signal monitoring
  - Recompute Z-Score daily
  - Re-run cointegration monthly to avoid broken relationships
  - Warn if half-life exceeds 2× the original value
Z-Score Signal Generation
python
def generate_pair_signals(
    y_price: pd.Series,
    x_price: pd.Series,
    lookback: int = 60,
    entry_z: float = 2.0,
    exit_z: float = 0.5,
    stop_z: float = 3.5,
    use_kalman: bool = False,
) -> pd.DataFrame:
    """Generate pair-trading signals.

    Args:
        y_price, x_price: Two price series
        lookback: Rolling Z-Score lookback, usually 2-3× half-life
        entry_z: Entry threshold
        exit_z: Exit threshold, usually near mean reversion
        stop_z: Stop threshold. Crossing it suggests cointegration may have broken
        use_kalman: Whether to use a Kalman dynamic hedge ratio

    Returns:
        DataFrame containing signals, Z-Score, and positions
    """
    if use_kalman:
        kf = kalman_hedge_ratio(y_price, x_price)
        spread = kf["spread"]
    else:
        y_ret = y_price.pct_change(fill_method=None)
        x_ret = x_price.pct_change(fill_method=None)
        res = bivariate_correlation_analysis(y_ret, x_ret, lookback)
        hedge_ratio = abs(res["beta"])
        spread = np.log(y_price) - hedge_ratio * np.log(x_price)

    spread_mean = spread.rolling(lookback).mean()
    spread_std = spread.rolling(lookback).std()
    z_score = (spread - spread_mean) / spread_std

    # Signal state machine to avoid repeated re-entry.
    signal_y = pd.Series(0.0, index=y_price.index)
    signal_x = pd.Series(0.0, index=x_price.index)
    position = 0  # 0=flat, 1=long spread, -1=short spread

    for i in range(lookback, len(z_score)):
        z = z_score.iloc[i]
        if np.isnan(z):
            continue

        if position == 0:
            if z < -entry_z:
                position = 1   # spread is too low: buy y, sell x
            elif z > entry_z:
                position = -1  # spread is too high: sell y, buy x
        elif position == 1:
            if z > -exit_z or z > stop_z:
                position = 0
        elif position == -1:
            if z < exit_z or z < -stop_z:
                position = 0

        signal_y.iloc[i] = 0.5 * position
        signal_x.iloc[i] = -0.5 * position

    return pd.DataFrame({
        "spread": spread,
        "z_score": z_score,
        "spread_mean": spread_mean,
        "spread_std": spread_std,
        "signal_y": signal_y,
        "signal_x": signal_x,
        "position": signal_y * 2,  # 1=long spread, -1=short spread, 0=flat
    })
Z-Score Threshold Configuration Guide
ParameterConservativeStandardAggressiveNotes
entry_z2.52.01.5Higher threshold means fewer trades
exit_z0.30.50.8Higher threshold means shorter holding periods
stop_z3.03.54.0Beyond this level, cointegration may have broken
lookback906030Usually 2-3× half-life
Spread-Stability Monitoring
python
def monitor_spread_health(
    spread: pd.Series,
    original_half_life: float,
    original_corr: float,
    warning_hl_multiple: float = 2.0,
    warning_corr_drop: float = 0.2,
) -> dict:
    """Monitor spread stability and judge whether cointegration still holds.

    Args:
        spread: Live spread series
        original_half_life: Half-life at entry, in days
        original_corr: Correlation at entry
        warning_hl_multiple: Warn if half-life exceeds this multiple of the original
        warning_corr_drop: Warn if correlation drops by more than this amount

    Returns:
        Health-status report
    """
    recent = spread.iloc[-60:] if len(spread) > 60 else spread

    current_hl = compute_half_life(recent)
    current_adf = adfuller(recent.dropna())[1]

    hl_ratio = current_hl / original_half_life if original_half_life > 0 else float("inf")

    # Cointegration health score
    health_score = 100
    warnings = []

    if current_adf > 0.10:
        health_score -= 40
        warnings.append(f"Spread ADF p={current_adf:.3f} > 0.10, stationarity has weakened")

    if hl_ratio > warning_hl_multiple:
        health_score -= 30
        warnings.append(
            f"Half-life {current_hl:.1f}d is {hl_ratio:.1f}x the original {original_half_life:.1f}d"
        )

    if current_adf > 0.20:
        health_score -= 20
        warnings.append("Spread may no longer be stationary. Re-test cointegration immediately")

    status = "healthy" if health_score >= 70 else "warning" if health_score >= 40 else "danger"

    return {
        "health_score": health_score,
        "status": status,
        "current_half_life": round(current_hl, 1),
        "hl_ratio": round(hl_ratio, 2),
        "spread_adf_p": round(current_adf, 4),
        "warnings": warnings,
        "action": "hold" if status == "healthy" else "reduce" if status == "warning" else "exit_now",
    }

Visualization Templates

Rolling-Correlation Time-Series Plot
python
import matplotlib.pyplot as plt

def plot_rolling_correlation(
    rolling_corrs: pd.DataFrame,
    title: str = "Rolling Correlation",
    figsize: tuple = (14, 5),
) -> plt.Figure:
    """Plot rolling correlation time series across multiple windows."""
    fig, ax = plt.subplots(figsize=figsize)
    colors = ["#2196F3", "#FF9800", "#4CAF50"]
    for i, col in enumerate(rolling_corrs.columns):
        ax.plot(rolling_corrs.index, rolling_corrs[col],
                label=col, color=colors[i % len(colors)], alpha=0.8)
    ax.axhline(0, color="black", linestyle="--", linewidth=0.8)
    ax.axhline(0.6, color="green", linestyle=":", linewidth=0.8, label="high-correlation threshold (0.6)")
    ax.axhline(-0.6, color="red", linestyle=":", linewidth=0.8)
    ax.set_title(title)
    ax.set_ylabel("Correlation")
    ax.legend(loc="best")
    ax.grid(True, alpha=0.3)
    plt.tight_layout()
    return fig
Z-Score Signal Plot
python
def plot_zscore_signals(
    signal_df: pd.DataFrame,
    entry_z: float = 2.0,
    stop_z: float = 3.5,
    figsize: tuple = (14, 8),
) -> plt.Figure:
    """Plot spread Z-Score and pair-trading signals."""
    fig, axes = plt.subplots(2, 1, figsize=figsize, sharex=True)

    # Top chart: spread
    axes[0].plot(signal_df["spread"], label="Spread", color="#1565C0")
    axes[0].plot(signal_df["spread_mean"], label="Mean", color="orange", linestyle="--")
    axes[0].fill_between(signal_df.index,
                         signal_df["spread_mean"] - signal_df["spread_std"],
                         signal_df["spread_mean"] + signal_df["spread_std"],
                         alpha=0.2, color="orange", label="±1σ")
    axes[0].set_title("Spread and Mean")
    axes[0].legend()
    axes[0].grid(True, alpha=0.3)

    # Bottom chart: Z-Score + signals
    axes[1].plot(signal_df["z_score"], label="Z-Score", color="#1565C0")
    axes[1].axhline(entry_z, color="red", linestyle="--", label=f"entry threshold (±{entry_z})")
    axes[1].axhline(-entry_z, color="red", linestyle="--")
    axes[1].axhline(stop_z, color="darkred", linestyle=":", label=f"stop threshold (±{stop_z})")
    axes[1].axhline(-stop_z, color="darkred", linestyle=":")
    axes[1].axhline(0, color="black", linestyle="-", linewidth=0.8)

    # Annotate entry / exit points
    long_entry = signal_df["position"].diff() > 0
    short_entry = signal_df["position"].diff() < 0
    exit_pos = (signal_df["position"] == 0) & (signal_df["position"].shift(1) != 0)

    axes[1].scatter(signal_df.index[long_entry], signal_df["z_score"][long_entry],
                    color="green", marker="^", s=80, label="long spread", zorder=5)
    axes[1].scatter(signal_df.index[short_entry], signal_df["z_score"][short_entry],
                    color="red", marker="v", s=80, label="short spread", zorder=5)
    axes[1].scatter(signal_df.index[exit_pos], signal_df["z_score"][exit_pos],
                    color="gray", marker="o", s=40, label="exit", zorder=5)

    axes[1].set_title("Z-Score and Trading Signals")
    axes[1].legend(loc="best")
    axes[1].grid(True, alpha=0.3)

    plt.tight_layout()
    return fig

Dependencies

bash
pip install pandas numpy scipy statsmodels matplotlib seaborn

Output Format

markdown
## Correlation and Cointegration Analysis Report

### Pair: [Asset A] vs [Asset B] ([Start Date] - [End Date])

#### Correlation Statistics
| Metric | Value | Interpretation |
|------|------|------|
| Pearson r | 0.82 | Strong linear positive correlation |
| Spearman ρ | 0.80 | Consistent monotonic relationship |
| Beta (A/B) | 1.15 | Sensitivity of A to B |
| R² | 0.67 | 67% of return variance in A is explained by B |

#### Cointegration Tests
| Method | Statistic | p-value | Conclusion |
|------|--------|------|------|
| Engle-Granger | -4.12 | 0.008 | Cointegrated ** |
| Johansen trace test | 28.3 | — | 1 cointegrating vector |
| Spread ADF | -3.95 | 0.002 | Spread is stationary ** |

#### Mean-Reversion Characteristics
| Metric | Value |
|------|------|
| OLS hedge ratio | 1.23 |
| Half-life | 18.5 days |
| Suggested holding window | 10-30 days |
| Suggested lookback window | 40-60 days |

#### Conditional Correlation (Regime Analysis)
| Regime | Correlation | Sample Size |
|------|---------|--------|
| Bull | 0.76 | 312 days |
| Bear | 0.88 | 198 days |
| High volatility | 0.91 | 87 days |
| Sideways | 0.71 | 645 days |

#### Recommended Pair-Trading Signal Parameters
| Parameter | Value |
|------|-----|
| entry_z | 2.0 |
| exit_z | 0.5 |
| stop_z | 3.5 |
| lookback | 60 days |

#### Current Spread Status
| Metric | Value | Alert |
|------|-----|------|
| Current Z-Score | -2.3 | Near entry zone |
| Health score | 85/100 | Healthy |
| Half-life (last 60 days) | 21.2 days | Normal |

Notes

  1. Prices vs returns: Use price series, which are non-stationary, for cointegration tests; use return series, which are stationary, for correlation analysis. Mixing them is the most common mistake.
  2. Data alignment: Cross-market analysis must handle holiday mismatches with an inner join. Do not forward-fill missing trading days, or you will create fake correlation.
  3. Cointegration is not the same as high correlation: Two series can have Pearson < 0.3 and still be cointegrated, and the reverse can also happen.
  4. Out-of-sample validation: If a pair is selected using cointegration on the first N years, you must verify whether the relationship survives in later out-of-sample data to avoid overfitting.
  5. Crisis-period risk: Correlation jumps in crises, and both legs in a pair can crash together. Stop thresholds should be tighter than in normal periods.
  6. China A-share specifics: China A-shares contain many non-trading days due to holidays and suspensions. Date alignment is especially important in cross-market comparison.
  7. Multiple testing: When testing N asset pairs simultaneously, use Benjamini-Hochberg FDR adjustment on p-values. Otherwise false positives will be excessive.
  8. Kalman tuning: Tune delta with grid search plus out-of-sample validation. Do not rely blindly on the default value.

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in agent/src/skills/correlation-analysis of HKUDS/Vibe-Trading.

Open the folder on GitHubat commit 7f6908b

Compare with similar skills

Correlation and Cointegration Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Correlation and Cointegration Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Correlation and Cointegration Analysis this skillHKUDS/Vibe-Trading35k—~10kAutomated safety check: PassMIT
Statistical Data Analysislingzhi227/agent-research-skills383—~886Automated safety check: PassNone
Tooluniverse Epigenomicswu-yc/LabClaw1.1k2 repos~14kAutomated safety check: PassNone
Quant Analystmajiayu000/claude-skill-registry6661 repos~964Automated safety check: PassMIT
Tushare Datazillionare/zillionare3182 repos~2.3kAutomated safety check: PassNone
Quant Blog Writingzillionare/zillionare318—~895Automated safety check: PassNone

Similar skills

  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    383 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Production-ready genomics and epigenomics data processing for BixBench questions.

    1.1k GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed
  • Quant Analyst

    majiayu000/claude-skill-registry

    Expert in quantitative finance, algorithmic trading, and financial data analysis using Python (Pandas/NumPy), statistical modeling, and machine learning.

    666 GitHub starsUsed in 1 repo~964 tokens
    Data & AnalyticsAuto-check passed
  • Tushare Data

    zillionare/zillionare

    面向中文自然语言的 Tushare 数据研究技能。用于把“看看这只股票最近怎么样”“帮我查财报趋势”“最近哪个板块最强”“北向资金在买什么”“给我导出一份行情数据”这类请求,转成可执行的数据获取、清洗、对比、筛选、导出与简要分析流程。适用于 A 股、指数、ETF/基金、财务、估值、资金流、公告新闻、板块概念与宏观数据等研究场景。

    318 GitHub starsUsed in 2 repos~2.3k tokens
    Business, Finance & HRAuto-check passed
  • Quant Blog Writing

    zillionare/zillionare

    撰写文笔精炼、富有深度的量化交易博文,论点清晰、证据确凿、叙事层次更加丰富。适用于量化交易博文、因子研究、回测复盘、数据源排查、市场微观结构、策略原理、风险控制、职业观察、量化人物故事等选题。文章将聚焦具体角度,提供详实的大纲、证据规划及成稿,力求内容兼具思想深度与诚实性,而非单纯口号式宣传;同时,通过人物经历、引言、贡献及行业背景的融入,让文章更具可读性和吸引力。

    318 GitHub stars~895 tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Runs exploratory data analysis on tabular data after you confirm each column's measurement level, then writes CSV tables and a narrative summary.

    108 GitHub stars~1.1k tokensUpdated 13 days ago
    Data & AnalyticsAuto-check passed

More from HKUDS/Vibe-Trading

All 89 skills in this repo
  • Eastmoney Market Data

    HKUDS/Vibe-Trading

    Index of Eastmoney's free, no-token market data interfaces for China A-shares and Hong Kong stocks: fund flows, dragon-tiger lists, margin trading, reports and news.

    35k GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • OKX Market Data

    HKUDS/Vibe-Trading

    Retrieves public OKX cryptocurrency market data such as spot prices, candlesticks, funding rates and open interest through the OKX V5 REST API, with no authentication.

    35k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • SEC EDGAR Filings Fetcher

    HKUDS/Vibe-Trading

    Fetches U.S. SEC EDGAR data: resolves tickers to CIK numbers, lists recent 10-K, 10-Q and 8-K filings with document URLs, and pulls XBRL financial series.

    35k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • A-Share ST Risk Screener

    HKUDS/Vibe-Trading

    Predicts whether a mainland China A-share company risks an ST or *ST warning after its next annual report, using financial thresholds and Sina penalty records.

    35k GitHub stars~4.9k tokensUpdated yesterday
    Auto-check passed
  • Breaks a structural trend such as AI infrastructure into its physical supply chain and ranks lesser-known listed companies sitting on each bottleneck.

    35k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Plans and drafts an eight-part, roughly 120k-word investigative series on one company, built around a strict fact-check pass rather than fast drafting.

    35k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed

Questions about Correlation and Cointegration Analysis

What does Correlation and Cointegration Analysis do?

Finds co-moving assets and tests them for cointegration, with workflows for correlation studies, sector clustering, hedge ratios and pair-trading signals. The skill describes four correlation modes: co-movement discovery, a deep return-correlation study of two assets, sector clustering and realized correlation. Discovery pulls daily returns for a target and a set of candidates, computes Pearson and Spearman correlations, ranks them and keeps the top few.

When should I use Correlation and Cointegration Analysis?

Correlation and Cointegration Analysis fits situations like: screening a universe of stocks for candidates to pair with a target; testing a pair for cointegration before building a spread strategy; choosing between Pearson, Spearman and Kendall for return series; estimating a dynamic hedge ratio and the half-life of a spread.

How do I install Correlation and Cointegration Analysis in Claude Code?

Run `npx skills add HKUDS/Vibe-Trading --skill correlation-analysis -a claude-code`. Or copy the skill folder (agent/src/skills/correlation-analysis in HKUDS/Vibe-Trading) into .claude/skills/correlation-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Correlation and Cointegration Analysis in Codex?

Run `npx skills add HKUDS/Vibe-Trading --skill correlation-analysis -a codex`. Or copy the skill folder (agent/src/skills/correlation-analysis in HKUDS/Vibe-Trading) into .agents/skills/correlation-analysis in your project. Codex loads it when a task matches its description.

Can I use Correlation and Cointegration Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/Vibe-Trading --skill correlation-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/correlation-analysis, .gemini/skills/correlation-analysis, .github/skills/correlation-analysis and .opencode/skills/correlation-analysis in your project.

What does Correlation and Cointegration Analysis need to run?

Going by SKILL.md and its folder, Correlation and Cointegration Analysis needs the command-line tools its instructions call (pip). Our summary lists: Python with pandas, numpy, scipy and statsmodels; Daily price or return data for the assets.

Does Correlation and Cointegration Analysis access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Correlation and Cointegration Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Correlation and Cointegration Analysis use?

Correlation and Cointegration Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Correlation and Cointegration Analysis use?

About 10k tokens (SKILL.md is roughly 40k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Correlation and Cointegration Analysis?

Skills that share tags, products or a category with Correlation and Cointegration Analysis: Statistical Data Analysis (lingzhi227/agent-research-skills, 383 stars), Tooluniverse Epigenomics (wu-yc/LabClaw, 1.1k stars), Quant Analyst (majiayu000/claude-skill-registry, 666 stars) and Tushare Data (zillionare/zillionare, 318 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Correlation and Cointegration Analysis?

HKUDS (a GitHub organization) maintains it in HKUDS/Vibe-Trading, which has 34,884 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 6, 2026.

Source: HKUDS/Vibe-Trading on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.