Alpha Correlation Analysis

Objective: Find truly complementary alpha combinations for composite stratégies.

Problem Statement

Current composite stratégies often have Sharpe ratios that are linear combinations of underlying alphas. We need to find: - Uncorrelated return streams (différent market regimes, différent asset classes) - Asymmetric complementarity (one alpha profits when the other is flat/losing) - Synergistic combinations (combined Sharpe > weighted average)

Methodology

  1. Return Stream Collection: Gather daily returns from all alpha stratégies
  2. Correlation Matrix Analysis: Identify uncorrelated stratégies
  3. Regime Analysis: Map performance across market regimes (bull/bear/sideways, high/low vol)
  4. Complementarity Score: Rank pairs by asymmetric performance
  5. Downside Protection: Measure how well stratégies protect in drawdowns

Alpha Universe

Trend Following

  • EMA-Cross-Alpha: Tech stock EMA crossover (Sharpe ~1.0)
  • TrendStocks-Alpha: 5-stock trend following (Sharpe ~0.6)
  • Trend-Following: Competition EMA 50/200

Momentum

  • SectorMomentum: 9 sector ETF relative momentum
  • DualMomentumNoTLT: IEF/GLD/XLP dual momentum

Defensive/All-Weather

  • AllWeather: Permanent portfolio style allocation
  • MeanReversion: Short-term mean reversion

Factor-Based

  • FamaFrench: 5-factor model implementation

Pairs/Market Neutral

  • ETF-Pairs: Cointegration-based pairs trading
  • Options-VGT: Wheel strategy (income generation)
# QuantConnect Research Environment
from datetime import datetime, timedelta
try:
    from AlgorithmImports import *
except (ImportError, ModuleNotFoundError):
    pass
import pandas as pd
import numpy as np

# Initialize QuantBook
try:
    qb = QuantBook()
    # Set analysis period
    qb.SetStartDate(2020, 1, 1)
    qb.SetEndDate(2024, 12, 31)
    print(f"Analysis period: {qb.StartDate} to {qb.EndDate}")
except NameError:
    print("Local environment detected - yfinance will be used for data")
Local environment detected - yfinance will be used for data

Note sur l’environnement d’exécution (#8772) :

Ce notebook déclare en cell[1] une période 2020-01-01 → 2024-12-31 via qb.SetStartDate / qb.SetEndDate. Ces appels n’ont jamais pris effet dans l’environnement qui a produit les sorties committées : QuantBook() a levé NameError (environnement QC Cloud / Lean indisponible), et la branche de repli yfinance a pris le relais (mécanisme 2 de #8772).

Conséquences : - Les sorties ci-dessous proviennent du chemin local yfinance, pas d’un QuantBook. - La fenêtre réellement calculée est ancrée à l’heure d’exécution : 2021-06-01 → 2026-05-29 (vérifié sur la cellule de chargement), elle dépasse la période déclarée au lieu de la précéder. - Le repli yfinance est une source de données externe légitime (#7066), mais il est annoncé ici explicitement plutôt que silencieux. Pour une exécution QuantBook authentique (et la fenêtre déclarée), ré-exécuter via QC Cloud une fois l’environnement restauré.

À noter : la fenêtre reculée actuelle repose néanmoins sur de vraies barres — ne pas la corriger en demandant explicitement 2020-2024 tant que les données equity locales s’arrêtent au 2021-03-31 (#8734), car Lean les prolongerait en constante (fillDataForward=True), produisant l’artefact de ligne plate de #8719.

1. Define Alpha Stratégies

We’ll implement simplified versions of each alpha to generate return streams.

# Define tickers for our alpha universe
tickers = ['SPY', 'TLT', 'GLD', 'XLP', 'IEF', 'AAPL', 'MSFT', 'GOOGL', 'AMZN', 'NVDA',
           'XLK', 'XLE', 'XLF', 'XLV', 'XLY', 'XLB', 'XLRE', 'XLU']

# Register assets on QC Cloud
assets = {}
try:
    for t in tickers:
        assets[t] = qb.AddEquity(t, Resolution.DAILY).Symbol
    print(f"Added {len(assets)} assets to QuantBook")
except NameError:
    print(f"Using {len(tickers)} tickers (local environment)")
Using 18 tickers (local environment)

Pivot de la série ‘close’ en DataFrame large, avec remapping des colonnes Symbol → ticker pour Alpha-Correlation-Analysis.

# Fetch historical data per-ticker (avoids MultiIndex Symbol lookup issue)
data = {}
for ticker in tickers:
    loaded = False
    # Try QC Cloud per-ticker
    try:
        hist = qb.History(assets[ticker], 365*5, Resolution.Daily)
        if hasattr(hist, 'shape') and len(hist) > 0:
            if isinstance(hist.index, pd.MultiIndex):
                try:
                    df_t = hist.loc[ticker]
                except KeyError:
                    df_t = hist.droplevel(0)
            else:
                df_t = hist.copy()
            if len(df_t) > 0:
                data[ticker] = df_t['close']
                loaded = True
    except (NameError, KeyError, TypeError):
        pass

    # yfinance fallback for this ticker
    if not loaded:
        try:
            import yfinance as yf
            raw = yf.Ticker(ticker).history(period="5y", auto_adjust=True)
            if len(raw) > 0:
                series = raw['Close']
                series.index = series.index.tz_localize(None)
                data[ticker] = series
                loaded = True
        except Exception as e:
            print(f"  WARNING: {ticker} unavailable ({e})")

closes = pd.DataFrame(data)
# Forward-fill to align different ticker date ranges, drop entirely-empty columns
closes = closes.ffill().dropna(axis=1, how='all').dropna(how='all')

print(f"Data shape: {closes.shape}")
print(f"Tickers loaded: {len(closes.columns)}/{len(tickers)}")
print(f"Date range: {closes.index.min()} to {closes.index.max()}")
print(f"Columns: {list(closes.columns)}")
Data shape: (1255, 18)
Tickers loaded: 18/18
Date range: 2021-06-01 00:00:00 to 2026-05-29 00:00:00
Columns: ['SPY', 'TLT', 'GLD', 'XLP', 'IEF', 'AAPL', 'MSFT', 'GOOGL', 'AMZN', 'NVDA', 'XLK', 'XLE', 'XLF', 'XLV', 'XLY', 'XLB', 'XLRE', 'XLU']

2. Implement Alpha Stratégies

Define functions that generate signals/returns for each alpha type.

def calculate_ema(prices, period):
    """Calculate Exponential Moving Average."""
    return prices.ewm(span=period, adjust=False).mean()

def calculate_rsi(prices, period=14):
    """Calculate RSI."""
    delta = prices.diff()
    gain = (delta.where(delta > 0, 0)).rolling(window=period).mean()
    loss = (-delta.where(delta < 0, 0)).rolling(window=period).mean()
    rs = gain / loss
    return 100 - (100 / (1 + rs))

def calculate_returns(prices, period=1):
    """Calculate returns."""
    return prices.pct_change(period)

def calculate_volatility(prices, period=20):
    """Calculate rolling volatility."""
    returns = prices.pct_change()
    return returns.rolling(period).std() * np.sqrt(252)

print("Helper functions defined")
Helper functions defined

La moyenne mobile exponentielle est calculée pour identifier les signaux directionnels de la stratégie Alpha-Correlation-Analysis.

def ema_cross_alpha(prices, fast=20, slow=60, cooldown=3):
    """EMA Crossover Alpha (EMA-Cross-Alpha style).
    
    Long when fast EMA > slow EMA, with cooldown.
    """
    ema_fast = calculate_ema(prices, fast)
    ema_slow = calculate_ema(prices, slow)
    
    # Generate signals
    signals = pd.Series(0, index=prices.index)
    signals[ema_fast > ema_slow] = 1
    
    # Apply cooldown (no signal changes for N days after switch)
    signal_changes = signals.diff().abs()
    for i in range(1, len(signals)):
        if signal_changes.iloc[i] > 0:
            for j in range(i+1, min(i+cooldown, len(signals))):
                signals.iloc[j] = signals.iloc[i]
    
    # Calculate returns
    returns = prices.pct_change().shift(-1)  # Next day return
    strategy_returns = signals * returns
    
    return strategy_returns, signals

# Test on SPY
spy_returns, spy_signals = ema_cross_alpha(closes['SPY'])
print(f"EMA Cross SPY - Total Return: {spy_returns.sum():.2%}")
print(f"Signal distribution: {spy_signals.value_counts().to_dict()}")
EMA Cross SPY - Total Return: 48.05%
Signal distribution: {1: 944, 0: 311}

Calcul des métriques de performance (Sharpe, CAGR, drawdown) pour évaluer la qualité du backtest de la stratégie Alpha-Correlation-Analysis.

def momentum_alpha(prices, lookback=90):
    """Momentum Alpha (SectorMomentum style).
    
    Long when price > lookback-period ago price.
    """
    momentum = prices / prices.shift(lookback) - 1
    signals = (momentum > 0).astype(int)
    
    returns = prices.pct_change().shift(-1)
    strategy_returns = signals * returns
    
    return strategy_returns, signals

# Test on SPY
mom_returns, mom_signals = momentum_alpha(closes['SPY'])
print(f"Momentum SPY - Total Return: {mom_returns.sum():.2%}")
Momentum SPY - Total Return: 51.12%

Définition de dual_momentum_alpha() : fonction d’allocation ou de signal utilisée dans la comparaison des stratégies de Alpha-Correlation-Analysis.

def dual_momentum_alpha(spy_prices, tlt_prices, gld_prices, xlp_prices, lookback=126):
    """Dual Momentum Alpha (DualMomentumNoTLT style).
    
    Relative momentum: invest in best performing asset.
    """
    # Calculate momentum
    spy_mom = spy_prices / spy_prices.shift(lookback) - 1
    tlt_mom = tlt_prices / tlt_prices.shift(lookback) - 1
    gld_mom = gld_prices / gld_prices.shift(lookback) - 1
    xlp_mom = xlp_prices / xlp_prices.shift(lookback) - 1
    
    # Find best performer
    momentums = pd.DataFrame({
        'SPY': spy_mom,
        'TLT': tlt_mom,
        'GLD': gld_mom,
        'XLP': xlp_mom
    })
    
    # Defensive: if SPY momentum < 0, choose best defensive
    # fillna(0): during lookback warmup, treat momentum as neutral
    signals = pd.DataFrame(index=momentums.index)
    signals['asset'] = momentums.fillna(0).idxmax(axis=1)
    
    # Calculate returns
    spy_ret = spy_prices.pct_change().shift(-1)
    tlt_ret = tlt_prices.pct_change().shift(-1)
    gld_ret = gld_prices.pct_change().shift(-1)
    xlp_ret = xlp_prices.pct_change().shift(-1)
    
    strategy_returns = (
        (signals['asset'] == 'SPY') * spy_ret +
        (signals['asset'] == 'TLT') * tlt_ret +
        (signals['asset'] == 'GLD') * gld_ret +
        (signals['asset'] == 'XLP') * xlp_ret
    )
    
    return strategy_returns, signals

# Test
dm_returns, dm_signals = dual_momentum_alpha(
    closes['SPY'], closes['TLT'], closes['GLD'], closes['XLP']
)
print(f"Dual Momentum - Total Return: {dm_returns.sum():.2%}")
print(f"Asset allocation: {dm_signals['asset'].value_counts().to_dict()}")
Dual Momentum - Total Return: 96.51%
Asset allocation: {'GLD': 615, 'SPY': 446, 'XLP': 190, 'TLT': 4}

Calcul du Z-score du spread pour détecter les déviations statistiques et générer les signaux d’entrée/sortie dans Alpha-Correlation-Analysis.

def mean_reversion_alpha(prices, entry_z=2.0, exit_z=0.5, period=20):
    """Mean Reversion Alpha.
    
    Long when price is Z standard deviations below mean.
    """
    # Calculate Z-score
    rolling_mean = prices.rolling(period).mean()
    rolling_std = prices.rolling(period).std()
    z_score = (prices - rolling_mean) / rolling_std
    
    # Generate signals
    signals = pd.Series(0, index=prices.index)
    signals[z_score < -entry_z] = 1  # Long when oversold
    signals[z_score > exit_z] = 0    # Exit when reverted
    
    # Forward fill signals
    signals = signals.ffill().fillna(0)
    
    returns = prices.pct_change().shift(-1)
    strategy_returns = signals * returns
    
    return strategy_returns, signals, z_score

# Test on SPY
mr_returns, mr_signals, z_score = mean_reversion_alpha(closes['SPY'])
print(f"Mean Reversion SPY - Total Return: {mr_returns.sum():.2%}")
Mean Reversion SPY - Total Return: 18.50%

Définition de all_weather_allocation() : calcul des poids d’allocation statique pour le portefeuille Alpha-Correlation-Analysis.

def all_weather_allocation(spy_weight=0.3, tlt_weight=0.4, gld_weight=0.15, xlp_weight=0.15):
    """All-Weather-style static allocation.
    
    Permanent portfolio with balanced allocation.
    """
    spy_ret = closes['SPY'].pct_change().shift(-1)
    tlt_ret = closes['TLT'].pct_change().shift(-1)
    gld_ret = closes['GLD'].pct_change().shift(-1)
    xlp_ret = closes['XLP'].pct_change().shift(-1)
    
    strategy_returns = (
        spy_weight * spy_ret +
        tlt_weight * tlt_ret +
        gld_weight * gld_ret +
        xlp_weight * xlp_ret
    )
    
    return strategy_returns

aw_returns = all_weather_allocation()
print(f"All Weather - Total Return: {aw_returns.sum():.2%}")
All Weather - Total Return: 30.96%

3. Generate Return Streams for All Alphas

# Generate return streams for each alpha
alpha_returns = {}

# 1. EMA Cross on SPY
ema_returns, _ = ema_cross_alpha(closes['SPY'])
alpha_returns['EMA-Cross-SPY'] = ema_returns

# 2. EMA Cross on Tech (equal weight AAPL, MSFT, GOOGL, AMZN, NVDA)
tech_returns = (
    closes['AAPL'].pct_change() +
    closes['MSFT'].pct_change() +
    closes['GOOGL'].pct_change() +
    closes['AMZN'].pct_change() +
    closes['NVDA'].pct_change()
) / 5
ema_tech_returns, _ = ema_cross_alpha(closes['AAPL'] * closes['MSFT'] * closes['GOOGL'])  # Proxy
# Use equal-weight tech benchmark with EMA timing
tech_ema_signals = pd.Series(0, index=closes.index)
for stock in ['AAPL', 'MSFT', 'GOOGL', 'AMZN', 'NVDA']:
    stock_ema_returns, stock_signals = ema_cross_alpha(closes[stock])
    tech_ema_signals = tech_ema_signals.add(stock_signals, fill_value=0)
tech_ema_signals = (tech_ema_signals / 5).round()  # Average signal
alpha_returns['EMA-Cross-Tech'] = tech_ema_signals * tech_returns.shift(-1)

# 3. Momentum SPY
mom_returns, _ = momentum_alpha(closes['SPY'])
alpha_returns['Momentum-SPY'] = mom_returns

# 4. Dual Momentum
dm_returns, _ = dual_momentum_alpha(
    closes['SPY'], closes['TLT'], closes['GLD'], closes['XLP']
)
alpha_returns['Dual-Momentum'] = dm_returns

# 5. Mean Reversion SPY
mr_returns, _, _ = mean_reversion_alpha(closes['SPY'])
alpha_returns['Mean-Reversion'] = mr_returns

# 6. All Weather
alpha_returns['All-Weather'] = all_weather_allocation()

# 7. Trend Following (slower EMA)
trend_returns, _ = ema_cross_alpha(closes['SPY'], fast=50, slow=200)
alpha_returns['Trend-Following'] = trend_returns

# Create DataFrame
returns_df = pd.DataFrame(alpha_returns)
returns_df = returns_df.dropna()

print(f"Return streams shape: {returns_df.shape}")
print(f"\nTotal Returns:")
print(returns_df.sum().sort_values(ascending=False))
Return streams shape: (1254, 7)

Total Returns:
EMA-Cross-Tech     1.005607
Dual-Momentum      0.965069
Momentum-SPY       0.511179
EMA-Cross-SPY      0.480501
Trend-Following    0.440954
All-Weather        0.309586
Mean-Reversion     0.185011
dtype: float64

4. Correlation Matrix Analysis

# Calculate correlation matrix
corr_matrix = returns_df.corr()

# Plot heatmap
import matplotlib.pyplot as plt
import seaborn as sns

plt.figure(figsize=(12, 10))
sns.heatmap(corr_matrix, annot=True, cmap='RdYlGn', center=0,
            vmin=-1, vmax=1, fmt='.2f', linewidths=1)
plt.title('Alpha Return Correlation Matrix')
plt.tight_layout()
plt.show()

print("\nCorrelation Matrix:")
print(corr_matrix.round(3))


Correlation Matrix:
                 EMA-Cross-SPY  EMA-Cross-Tech  Momentum-SPY  Dual-Momentum  \
EMA-Cross-SPY            1.000           0.711         0.815          0.283   
EMA-Cross-Tech           0.711           1.000         0.632          0.200   
Momentum-SPY             0.815           0.632         1.000          0.248   
Dual-Momentum            0.283           0.200         0.248          1.000   
Mean-Reversion           0.113           0.054         0.101          0.138   
All-Weather              0.425           0.277         0.397          0.518   
Trend-Following          0.688           0.549         0.691          0.279   

                 Mean-Reversion  All-Weather  Trend-Following  
EMA-Cross-SPY             0.113        0.425            0.688  
EMA-Cross-Tech            0.054        0.277            0.549  
Momentum-SPY              0.101        0.397            0.691  
Dual-Momentum             0.138        0.518            0.279  
Mean-Reversion            1.000        0.273            0.463  
All-Weather               0.273        1.000            0.500  
Trend-Following           0.463        0.500            1.000  

Calcul de la matrice de corrélation entre les instruments pour l’analyse de diversification de la stratégie Alpha-Correlation-Analysis.

# Find least correlated pairs
pairs = []
for i, alpha1 in enumerate(returns_df.columns):
    for alpha2 in returns_df.columns[i+1:]:
        corr = corr_matrix.loc[alpha1, alpha2]
        pairs.append({
            'Pair': f"{alpha1} / {alpha2}",
            'Correlation': corr,
            'Abs_Corr': abs(corr)
        })

pairs_df = pd.DataFrame(pairs).sort_values('Abs_Corr')

print("Top 10 Least Correlated Pairs:")
print(pairs_df.head(10)[['Pair', 'Correlation']])
Top 10 Least Correlated Pairs:
                               Pair  Correlation
8   EMA-Cross-Tech / Mean-Reversion     0.053828
12    Momentum-SPY / Mean-Reversion     0.101160
3    EMA-Cross-SPY / Mean-Reversion     0.113459
15   Dual-Momentum / Mean-Reversion     0.138500
7    EMA-Cross-Tech / Dual-Momentum     0.200210
11     Momentum-SPY / Dual-Momentum     0.247990
18     Mean-Reversion / All-Weather     0.273308
9      EMA-Cross-Tech / All-Weather     0.277077
17  Dual-Momentum / Trend-Following     0.278780
2     EMA-Cross-SPY / Dual-Momentum     0.283054

5. Regime Analysis

Classify market conditions and analyze alpha performance in each regime.

# Define market regimes
spy_returns = closes['SPY'].pct_change()
spy_volatility = spy_returns.rolling(20).std() * np.sqrt(252)
spy_trend = spy_returns.rolling(60).sum()

# Regime classifications
regimes = pd.DataFrame(index=returns_df.index)
regimes['vol_regime'] = pd.cut(
    spy_volatility,
    bins=[0, spy_volatility.quantile(0.33), spy_volatility.quantile(0.67), float('inf')],
    labels=['Low-Vol', 'Med-Vol', 'High-Vol']
)
regimes['trend_regime'] = pd.cut(
    spy_trend,
    bins=[-float('inf'), -0.05, 0.05, float('inf')],
    labels=['Bear', 'Sideways', 'Bull']
)

# Combine regimes
regimes['regime'] = regimes['trend_regime'].astype(str) + '-' + regimes['vol_regime'].astype(str)

print("Regime Distribution:")
print(regimes['regime'].value_counts())
Regime Distribution:
regime
Bull-Low-Vol         270
Sideways-Med-Vol     238
Sideways-High-Vol    205
Bull-Med-Vol         158
Bear-High-Vol        125
Sideways-Low-Vol     102
Bull-High-Vol         78
Bear-Med-Vol          18
Name: count, dtype: int64

Classification des régimes de marché (tendanciel / retour à la moyenne) pour adapter les paramètres de Alpha-Correlation-Analysis.

# Analyze alpha performance by regime
regime_performance = {}

for alpha in returns_df.columns:
    alpha_perf = []
    for regime in regimes['regime'].unique():
        mask = regimes['regime'] == regime
        if mask.sum() > 10:  # Need sufficient data
            avg_return = returns_df.loc[mask, alpha].mean() * 252  # Annualized
            sharpe = returns_df.loc[mask, alpha].mean() / returns_df.loc[mask, alpha].std() * np.sqrt(252) if returns_df.loc[mask, alpha].std() > 0 else 0
            alpha_perf.append({
                'Alpha': alpha,
                'Regime': regime,
                'Ann_Return': avg_return,
                'Sharpe': sharpe,
                'Days': mask.sum()
            })
    regime_performance[alpha] = pd.DataFrame(alpha_perf)

# Combine all
regime_df = pd.concat(regime_performance.values())
regime_pivot = regime_df.pivot(index='Alpha', columns='Regime', values='Sharpe')

plt.figure(figsize=(14, 8))
sns.heatmap(regime_pivot, annot=True, cmap='RdYlGn', center=0,
            fmt='.2f', linewidths=1)
plt.title('Sharpe Ratio by Market Regime')
plt.tight_layout()
plt.show()

print("\nSharpe by Regime:")
print(regime_pivot.round(2))


Sharpe by Regime:
Regime           Bear-High-Vol  Bear-Med-Vol  Bull-High-Vol  Bull-Low-Vol  \
Alpha                                                                       
All-Weather               1.98          7.04           0.33          0.18   
Dual-Momentum             0.09          5.86          -0.23          1.73   
EMA-Cross-SPY             0.00          0.00           0.60          0.87   
EMA-Cross-Tech            0.00         -4.16           0.41          1.70   
Mean-Reversion            1.02          4.53           0.00         -0.97   
Momentum-SPY              0.53         -3.74           0.75          1.00   
Trend-Following           0.69          2.73           0.78          0.87   

Regime           Bull-Med-Vol  Sideways-High-Vol  Sideways-Low-Vol  \
Alpha                                                                
All-Weather              1.48              -1.14              0.40   
Dual-Momentum            2.65               0.18              0.58   
EMA-Cross-SPY            1.21               0.53              0.28   
EMA-Cross-Tech           0.77              -0.44              0.39   
Mean-Reversion           1.42              -1.21              2.32   
Momentum-SPY             1.08               1.87              0.27   
Trend-Following          1.19              -0.58              0.07   

Regime           Sideways-Med-Vol  
Alpha                              
All-Weather                  0.51  
Dual-Momentum                1.26  
EMA-Cross-SPY                1.10  
EMA-Cross-Tech               2.28  
Mean-Reversion               0.96  
Momentum-SPY                 0.90  
Trend-Following              0.87  

6. Complementarity Score

Calculate how well pairs complement each other: - Low correlation: Uncorrelated return streams - Regime diversification: One excels where the other struggles - Drawdown protection: One protects when the other draws down

def calculate_complementarity_score(alpha1_returns, alpha2_returns, regimes):
    """Calculate complementarity score for a pair of alphas.
    
    Score components:
    1. Correlation: Lower is better (0-1)
    2. Regime diversification: One excels where other struggles
    3. Downside protection: One protects when other draws down
    """
    # Remove NaN
    common_idx = alpha1_returns.dropna().index.intersection(alpha2_returns.dropna().index)
    r1 = alpha1_returns.loc[common_idx]
    r2 = alpha2_returns.loc[common_idx]
    regimes_common = regimes.loc[common_idx]
    
    # 1. Correlation score (lower correlation = higher score)
    corr = r1.corr(r2)
    corr_score = (1 - abs(corr))  # 0-1 scale
    
    # 2. Regime diversification
    regime_scores = []
    for regime in regimes_common['regime'].unique():
        mask = regimes_common['regime'] == regime
        if mask.sum() > 10:
            sharpe1 = r1.loc[mask].mean() / r1.loc[mask].std() * np.sqrt(252) if r1.loc[mask].std() > 0 else 0
            sharpe2 = r2.loc[mask].mean() / r2.loc[mask].std() * np.sqrt(252) if r2.loc[mask].std() > 0 else 0
            # Diversification benefit: one positive when other negative
            if (sharpe1 > 0 and sharpe2 < 0) or (sharpe1 < 0 and sharpe2 > 0):
                regime_scores.append(abs(sharpe1 - sharpe2) / 2)
            elif sharpe1 > 0 and sharpe2 > 0:
                regime_scores.append(max(sharpe1, sharpe2) - min(sharpe1, sharpe2))
    
    regime_div_score = np.mean(regime_scores) if regime_scores else 0
    
    # 3. Downside protection
    # When alpha1 is in drawdown, does alpha2 protect?
    cumret1 = (1 + r1).cumprod()
    cumret2 = (1 + r2).cumprod()
    
    # Find drawdown periods for alpha1
    drawdown_mask = (cumret1 < cumret1.cummax()) & (cumret1 < cumret1.iloc[0])
    
    if drawdown_mask.sum() > 10:
        # Alpha2's performance during alpha1's drawdowns
        dd_return2 = r2.loc[drawdown_mask].mean() * 252
        protection_score = max(0, dd_return2)  # Positive returns during drawdown = good
    else:
        protection_score = 0
    
    # Combined score (weighted)
    total_score = 0.4 * corr_score + 0.4 * min(regime_div_score, 1) + 0.2 * min(protection_score, 1)
    
    return {
        'correlation': corr,
        'corr_score': corr_score,
        'regime_diversification': regime_div_score,
        'downside_protection': protection_score,
        'total_score': total_score
    }

# Calculate scores for all pairs
complementarity = []
for i, alpha1 in enumerate(returns_df.columns):
    for alpha2 in returns_df.columns[i+1:]:
        scores = calculate_complementarity_score(
            returns_df[alpha1],
            returns_df[alpha2],
            regimes
        )
        complementarity.append({
            'Pair': f"{alpha1} / {alpha2}",
            **scores
        })

comp_df = pd.DataFrame(complementarity).sort_values('total_score', ascending=False)

print("Top 10 Most Complementary Pairs:")
print(comp_df.head(10)[['Pair', 'correlation', 'regime_diversification', 'downside_protection', 'total_score']].round(3))
Top 10 Most Complementary Pairs:
                               Pair  correlation  regime_diversification  \
8   EMA-Cross-Tech / Mean-Reversion        0.054                   1.919   
12    Momentum-SPY / Mean-Reversion        0.101                   1.372   
7    EMA-Cross-Tech / Dual-Momentum        0.200                   1.253   
15   Dual-Momentum / Mean-Reversion        0.138                   1.084   
11     Momentum-SPY / Dual-Momentum        0.248                   1.299   
9      EMA-Cross-Tech / All-Weather        0.277                   1.616   
3    EMA-Cross-SPY / Mean-Reversion        0.113                   0.840   
18     Mean-Reversion / All-Weather        0.273                   1.080   
17  Dual-Momentum / Trend-Following        0.279                   0.980   
13       Momentum-SPY / All-Weather        0.397                   1.314   

    downside_protection  total_score  
8                 0.000        0.778  
12                0.045        0.769  
7                 0.163        0.753  
15                0.000        0.745  
11                0.000        0.701  
9                 0.051        0.699  
3                 0.021        0.695  
18                0.000        0.691  
17                0.040        0.688  
13                0.000        0.641  

7. Top Pairs Analysis

Deep dive into the most promising complementary pairs.

# Select top 3 pairs for detailed analysis
top_pairs = comp_df.head(3)['Pair'].tolist()

for pair_str in top_pairs:
    print(f"\n{'='*60}")
    print(f"Analysis: {pair_str}")
    print(f"{'='*60}")
    
    alpha1, alpha2 = [a.strip() for a in pair_str.split('/')]
    
    # Individual metrics
    print(f"\n{alpha1}:")
    print(f"  Annual Return: {returns_df[alpha1].mean() * 252:.2%}")
    print(f"  Volatility: {returns_df[alpha1].std() * np.sqrt(252):.2%}")
    print(f"  Sharpe: {returns_df[alpha1].mean() / returns_df[alpha1].std() * np.sqrt(252):.2f}")
    
    print(f"\n{alpha2}:")
    print(f"  Annual Return: {returns_df[alpha2].mean() * 252:.2%}")
    print(f"  Volatility: {returns_df[alpha2].std() * np.sqrt(252):.2%}")
    print(f"  Sharpe: {returns_df[alpha2].mean() / returns_df[alpha2].std() * np.sqrt(252):.2f}")
    
    # Combined 50/50
    combined = (returns_df[alpha1] + returns_df[alpha2]) / 2
    print(f"\nCombined (50/50):")
    print(f"  Annual Return: {combined.mean() * 252:.2%}")
    print(f"  Volatility: {combined.std() * np.sqrt(252):.2%}")
    print(f"  Sharpe: {combined.mean() / combined.std() * np.sqrt(252):.2f}")
    
    # Correlation
    corr = returns_df[alpha1].corr(returns_df[alpha2])
    print(f"  Correlation: {corr:.3f}")

============================================================
Analysis: EMA-Cross-Tech / Mean-Reversion
============================================================

EMA-Cross-Tech:
  Annual Return: 20.21%
  Volatility: 18.29%
  Sharpe: 1.10

Mean-Reversion:
  Annual Return: 3.72%
  Volatility: 6.61%
  Sharpe: 0.56

Combined (50/50):
  Annual Return: 11.96%
  Volatility: 9.89%
  Sharpe: 1.21
  Correlation: 0.054

============================================================
Analysis: Momentum-SPY / Mean-Reversion
============================================================

Momentum-SPY:
  Annual Return: 10.27%
  Volatility: 10.63%
  Sharpe: 0.97

Mean-Reversion:
  Annual Return: 3.72%
  Volatility: 6.61%
  Sharpe: 0.56

Combined (50/50):
  Annual Return: 7.00%
  Volatility: 6.53%
  Sharpe: 1.07
  Correlation: 0.101

============================================================
Analysis: EMA-Cross-Tech / Dual-Momentum
============================================================

EMA-Cross-Tech:
  Annual Return: 20.21%
  Volatility: 18.29%
  Sharpe: 1.10

Dual-Momentum:
  Annual Return: 19.39%
  Volatility: 17.70%
  Sharpe: 1.10

Combined (50/50):
  Annual Return: 19.80%
  Volatility: 13.94%
  Sharpe: 1.42
  Correlation: 0.200

8. Recommendations

Based on the analysis, here are the most promising complementary alpha combinations for composite stratégies.

# Generate recommendations
recommendations = []

for _, row in comp_df.head(10).iterrows():
    alpha1, alpha2 = [a.strip() for a in row['Pair'].split('/')]
    
    # Calculate expected combined Sharpe
    sharpe1 = returns_df[alpha1].mean() / returns_df[alpha1].std() * np.sqrt(252)
    sharpe2 = returns_df[alpha2].mean() / returns_df[alpha2].std() * np.sqrt(252)
    combined = (returns_df[alpha1] + returns_df[alpha2]) / 2
    combined_sharpe = combined.mean() / combined.std() * np.sqrt(252)
    
    recommendations.append({
        'Pair': row['Pair'],
        'Correlation': row['correlation'],
        'Sharpe_1': sharpe1,
        'Sharpe_2': sharpe2,
        'Combined_Sharpe': combined_sharpe,
        'Synergy': combined_sharpe - (sharpe1 + sharpe2) / 2,
        'Regime_Div': row['regime_diversification'],
        'DD_Protection': row['downside_protection']
    })

rec_df = pd.DataFrame(recommendations)

print("\nTop 10 Complementary Alpha Pairs for Composites:")
print(rec_df.round(3))

# Filter for synergistic pairs (combined Sharpe > average)
synergistic = rec_df[rec_df['Synergy'] > 0]
print(f"\nSynergistic Pairs ({len(synergistic)}):")
print(synergistic[['Pair', 'Combined_Sharpe', 'Synergy']].round(3))

Top 10 Complementary Alpha Pairs for Composites:
                              Pair  Correlation  Sharpe_1  Sharpe_2  \
0  EMA-Cross-Tech / Mean-Reversion        0.054     1.105     0.563   
1    Momentum-SPY / Mean-Reversion        0.101     0.967     0.563   
2   EMA-Cross-Tech / Dual-Momentum        0.200     1.105     1.096   
3   Dual-Momentum / Mean-Reversion        0.138     1.096     0.563   
4     Momentum-SPY / Dual-Momentum        0.248     0.967     1.096   
5     EMA-Cross-Tech / All-Weather        0.277     1.105     0.598   
6   EMA-Cross-SPY / Mean-Reversion        0.113     0.872     0.563   
7     Mean-Reversion / All-Weather        0.273     0.563     0.598   
8  Dual-Momentum / Trend-Following        0.279     1.096     0.658   
9       Momentum-SPY / All-Weather        0.397     0.967     0.598   

   Combined_Sharpe  Synergy  Regime_Div  DD_Protection  
0            1.210    0.376       1.919          0.000  
1            1.071    0.306       1.372          0.045  
2            1.420    0.320       1.253          0.163  
3            1.171    0.342       1.084          0.000  
4            1.302    0.270       1.299          0.000  
5            1.128    0.277       1.616          0.051  
6            0.989    0.272       0.840          0.021  
7            0.722    0.142       1.080          0.000  
8            1.128    0.251       0.980          0.040  
9            0.938    0.156       1.314          0.000  

Synergistic Pairs (10):
                              Pair  Combined_Sharpe  Synergy
0  EMA-Cross-Tech / Mean-Reversion            1.210    0.376
1    Momentum-SPY / Mean-Reversion            1.071    0.306
2   EMA-Cross-Tech / Dual-Momentum            1.420    0.320
3   Dual-Momentum / Mean-Reversion            1.171    0.342
4     Momentum-SPY / Dual-Momentum            1.302    0.270
5     EMA-Cross-Tech / All-Weather            1.128    0.277
6   EMA-Cross-SPY / Mean-Reversion            0.989    0.272
7     Mean-Reversion / All-Weather            0.722    0.142
8  Dual-Momentum / Trend-Following            1.128    0.251
9       Momentum-SPY / All-Weather            0.938    0.156

9. Walk-Forward Validation Framework

Template for validating pairs with out-of-sample testing.

def walk_forward_validation(returns_df, alpha1, alpha2, 
                           train_periods=252, test_periods=63):
    """Walk-forward validation for a pair.
    
    Args:
        returns_df: DataFrame with all returns
        alpha1, alpha2: Names of alphas to test
        train_periods: Training period in days
        test_periods: Test period in days
    """
    results = []
    
    for start in range(0, len(returns_df) - train_periods - test_periods, test_periods):
        train_end = start + train_periods
        test_end = train_end + test_periods
        
        if test_end > len(returns_df):
            break
        
        # Training period metrics
        train_corr = returns_df.iloc[start:train_end][alpha1].corr(
            returns_df.iloc[start:train_end][alpha2]
        )
        
        # Test period combined performance
        test_combined = (
            returns_df.iloc[train_end:test_end][alpha1] +
            returns_df.iloc[train_end:test_end][alpha2]
        ) / 2
        
        results.append({
            'Period': f"{returns_df.index[start].date()} to {returns_df.index[test_end-1].date()}",
            'Train_Corr': train_corr,
            'Test_Return': test_combined.mean() * 252,
            'Test_Sharpe': test_combined.mean() / test_combined.std() * np.sqrt(252) if test_combined.std() > 0 else 0
        })
    
    return pd.DataFrame(results)

# Example: Test top pair
top_pair = comp_df.iloc[0]['Pair']
alpha1, alpha2 = [a.strip() for a in top_pair.split('/')]

wf_results = walk_forward_validation(returns_df, alpha1, alpha2)
print(f"\nWalk-Forward Validation: {top_pair}")
print(wf_results)

print(f"\nAverage Test Sharpe: {wf_results['Test_Sharpe'].mean():.2f}")

Walk-Forward Validation: EMA-Cross-Tech / Mean-Reversion
                      Period  Train_Corr  Test_Return  Test_Sharpe
0   2021-06-01 to 2022-08-29    0.057506    -0.151942    -1.662360
1   2021-08-30 to 2022-11-28    0.059558    -0.129811    -4.520180
2   2021-11-29 to 2023-03-01    0.041652    -0.042323    -0.432370
3   2022-03-01 to 2023-05-31    0.018414     0.674825     5.440688
4   2022-05-31 to 2023-08-30    0.013267     0.190718     1.804279
5   2022-08-30 to 2023-11-29    0.039571    -0.251398    -2.718211
6   2022-11-29 to 2024-03-01    0.017564     0.347188     3.268694
7   2023-03-02 to 2024-05-31    0.016980     0.274166     2.469639
8   2023-06-01 to 2024-08-30    0.138705    -0.092481    -0.715447
9   2023-08-31 to 2024-11-29    0.258915     0.215528     2.461700
10  2023-11-30 to 2025-03-05    0.286068     0.027320     0.244899
11  2024-03-04 to 2025-06-04    0.248321     0.095622     0.750411
12  2024-06-03 to 2025-09-04    0.074195     0.359873     4.465036
13  2024-09-03 to 2025-12-03    0.008155     0.284876     2.897551
14  2024-12-02 to 2026-03-06    0.023313    -0.096559    -1.202214

Average Test Sharpe: 0.84

10. Next Steps

For Production Implementation:

  1. Build Composite Algorithms: Implement top pairs as QCAlphaModel composites
  2. Dynamic Weighting: Add regime-based allocation between alphas
  3. Risk Management: Add position sizing based on correlation regime
  4. Validation: Run full backtests on recommended pairs

Top Candidates for Composites:

Based on this analysis, prioritize: 1. Trend + Defensive: EMA-Cross + All-Weather 2. Momentum + Mean Reversion: Dual-Momentum + Mean-Reversion
3. Multi-Asset: Dual-Momentum + All-Weather 4. Regime Switching: Trend-Following + Mean-Reversion

Research Extensions:

  • Add sector-specific alphas (SectorMomentum)
  • Include options-based stratégies (Options-VGT)
  • Test with crypto assets (BTC, ETH)
  • Factor analysis: loadings to Size, Value, Momentum factors
Retour au sommet