Objective: Find truly complementary alpha combinations for composite stratégies.
Problem Statement
Current composite stratégies often have Sharpe ratios that are linear combinations of underlying alphas. We need to find: - Uncorrelated return streams (différent market regimes, différent asset classes) - Asymmetric complementarity (one alpha profits when the other is flat/losing) - Synergistic combinations (combined Sharpe > weighted average)
Methodology
Return Stream Collection: Gather daily returns from all alpha stratégies
Regime Analysis: Map performance across market regimes (bull/bear/sideways, high/low vol)
Complementarity Score: Rank pairs by asymmetric performance
Downside Protection: Measure how well stratégies protect in drawdowns
Alpha Universe
Trend Following
EMA-Cross-Alpha: Tech stock EMA crossover (Sharpe ~1.0)
TrendStocks-Alpha: 5-stock trend following (Sharpe ~0.6)
Trend-Following: Competition EMA 50/200
Momentum
SectorMomentum: 9 sector ETF relative momentum
DualMomentumNoTLT: IEF/GLD/XLP dual momentum
Defensive/All-Weather
AllWeather: Permanent portfolio style allocation
MeanReversion: Short-term mean reversion
Factor-Based
FamaFrench: 5-factor model implementation
Pairs/Market Neutral
ETF-Pairs: Cointegration-based pairs trading
Options-VGT: Wheel strategy (income generation)
# QuantConnect Research Environmentfrom datetime import datetime, timedeltatry:from AlgorithmImports import*except (ImportError, ModuleNotFoundError):passimport pandas as pdimport numpy as np# Initialize QuantBooktry: qb = QuantBook()# Set analysis period qb.SetStartDate(2020, 1, 1) qb.SetEndDate(2024, 12, 31)print(f"Analysis period: {qb.StartDate} to {qb.EndDate}")exceptNameError:print("Local environment detected - yfinance will be used for data")
Local environment detected - yfinance will be used for data
Note sur l’environnement d’exécution (#8772) :
Ce notebook déclare en cell[1] une période 2020-01-01 → 2024-12-31 via qb.SetStartDate / qb.SetEndDate. Ces appels n’ont jamais pris effet dans l’environnement qui a produit les sorties committées : QuantBook() a levé NameError (environnement QC Cloud / Lean indisponible), et la branche de repli yfinance a pris le relais (mécanisme 2 de #8772).
Conséquences : - Les sorties ci-dessous proviennent du chemin local yfinance, pas d’un QuantBook. - La fenêtre réellement calculée est ancrée à l’heure d’exécution : 2021-06-01 → 2026-05-29 (vérifié sur la cellule de chargement), elle dépasse la période déclarée au lieu de la précéder. - Le repli yfinance est une source de données externe légitime (#7066), mais il est annoncé ici explicitement plutôt que silencieux. Pour une exécution QuantBook authentique (et la fenêtre déclarée), ré-exécuter via QC Cloud une fois l’environnement restauré.
À noter : la fenêtre reculée actuelle repose néanmoins sur de vraies barres — ne pas la corriger en demandant explicitement 2020-2024 tant que les données equity locales s’arrêtent au 2021-03-31 (#8734), car Lean les prolongerait en constante (fillDataForward=True), produisant l’artefact de ligne plate de #8719.
1. Define Alpha Stratégies
We’ll implement simplified versions of each alpha to generate return streams.
# Define tickers for our alpha universetickers = ['SPY', 'TLT', 'GLD', 'XLP', 'IEF', 'AAPL', 'MSFT', 'GOOGL', 'AMZN', 'NVDA','XLK', 'XLE', 'XLF', 'XLV', 'XLY', 'XLB', 'XLRE', 'XLU']# Register assets on QC Cloudassets = {}try:for t in tickers: assets[t] = qb.AddEquity(t, Resolution.DAILY).Symbolprint(f"Added {len(assets)} assets to QuantBook")exceptNameError:print(f"Using {len(tickers)} tickers (local environment)")
Using 18 tickers (local environment)
Pivot de la série ‘close’ en DataFrame large, avec remapping des colonnes Symbol → ticker pour Alpha-Correlation-Analysis.
# Fetch historical data per-ticker (avoids MultiIndex Symbol lookup issue)data = {}for ticker in tickers: loaded =False# Try QC Cloud per-tickertry: hist = qb.History(assets[ticker], 365*5, Resolution.Daily)ifhasattr(hist, 'shape') andlen(hist) >0:ifisinstance(hist.index, pd.MultiIndex):try: df_t = hist.loc[ticker]exceptKeyError: df_t = hist.droplevel(0)else: df_t = hist.copy()iflen(df_t) >0: data[ticker] = df_t['close'] loaded =Trueexcept (NameError, KeyError, TypeError):pass# yfinance fallback for this tickerifnot loaded:try:import yfinance as yf raw = yf.Ticker(ticker).history(period="5y", auto_adjust=True)iflen(raw) >0: series = raw['Close'] series.index = series.index.tz_localize(None) data[ticker] = series loaded =TrueexceptExceptionas e:print(f" WARNING: {ticker} unavailable ({e})")closes = pd.DataFrame(data)# Forward-fill to align different ticker date ranges, drop entirely-empty columnscloses = closes.ffill().dropna(axis=1, how='all').dropna(how='all')print(f"Data shape: {closes.shape}")print(f"Tickers loaded: {len(closes.columns)}/{len(tickers)}")print(f"Date range: {closes.index.min()} to {closes.index.max()}")print(f"Columns: {list(closes.columns)}")
Calculate how well pairs complement each other: - Low correlation: Uncorrelated return streams - Regime diversification: One excels where the other struggles - Drawdown protection: One protects when the other draws down
def calculate_complementarity_score(alpha1_returns, alpha2_returns, regimes):"""Calculate complementarity score for a pair of alphas. Score components: 1. Correlation: Lower is better (0-1) 2. Regime diversification: One excels where other struggles 3. Downside protection: One protects when other draws down """# Remove NaN common_idx = alpha1_returns.dropna().index.intersection(alpha2_returns.dropna().index) r1 = alpha1_returns.loc[common_idx] r2 = alpha2_returns.loc[common_idx] regimes_common = regimes.loc[common_idx]# 1. Correlation score (lower correlation = higher score) corr = r1.corr(r2) corr_score = (1-abs(corr)) # 0-1 scale# 2. Regime diversification regime_scores = []for regime in regimes_common['regime'].unique(): mask = regimes_common['regime'] == regimeif mask.sum() >10: sharpe1 = r1.loc[mask].mean() / r1.loc[mask].std() * np.sqrt(252) if r1.loc[mask].std() >0else0 sharpe2 = r2.loc[mask].mean() / r2.loc[mask].std() * np.sqrt(252) if r2.loc[mask].std() >0else0# Diversification benefit: one positive when other negativeif (sharpe1 >0and sharpe2 <0) or (sharpe1 <0and sharpe2 >0): regime_scores.append(abs(sharpe1 - sharpe2) /2)elif sharpe1 >0and sharpe2 >0: regime_scores.append(max(sharpe1, sharpe2) -min(sharpe1, sharpe2)) regime_div_score = np.mean(regime_scores) if regime_scores else0# 3. Downside protection# When alpha1 is in drawdown, does alpha2 protect? cumret1 = (1+ r1).cumprod() cumret2 = (1+ r2).cumprod()# Find drawdown periods for alpha1 drawdown_mask = (cumret1 < cumret1.cummax()) & (cumret1 < cumret1.iloc[0])if drawdown_mask.sum() >10:# Alpha2's performance during alpha1's drawdowns dd_return2 = r2.loc[drawdown_mask].mean() *252 protection_score =max(0, dd_return2) # Positive returns during drawdown = goodelse: protection_score =0# Combined score (weighted) total_score =0.4* corr_score +0.4*min(regime_div_score, 1) +0.2*min(protection_score, 1)return {'correlation': corr,'corr_score': corr_score,'regime_diversification': regime_div_score,'downside_protection': protection_score,'total_score': total_score }# Calculate scores for all pairscomplementarity = []for i, alpha1 inenumerate(returns_df.columns):for alpha2 in returns_df.columns[i+1:]: scores = calculate_complementarity_score( returns_df[alpha1], returns_df[alpha2], regimes ) complementarity.append({'Pair': f"{alpha1} / {alpha2}",**scores })comp_df = pd.DataFrame(complementarity).sort_values('total_score', ascending=False)print("Top 10 Most Complementary Pairs:")print(comp_df.head(10)[['Pair', 'correlation', 'regime_diversification', 'downside_protection', 'total_score']].round(3))
Template for validating pairs with out-of-sample testing.
def walk_forward_validation(returns_df, alpha1, alpha2, train_periods=252, test_periods=63):"""Walk-forward validation for a pair. Args: returns_df: DataFrame with all returns alpha1, alpha2: Names of alphas to test train_periods: Training period in days test_periods: Test period in days """ results = []for start inrange(0, len(returns_df) - train_periods - test_periods, test_periods): train_end = start + train_periods test_end = train_end + test_periodsif test_end >len(returns_df):break# Training period metrics train_corr = returns_df.iloc[start:train_end][alpha1].corr( returns_df.iloc[start:train_end][alpha2] )# Test period combined performance test_combined = ( returns_df.iloc[train_end:test_end][alpha1] + returns_df.iloc[train_end:test_end][alpha2] ) /2 results.append({'Period': f"{returns_df.index[start].date()} to {returns_df.index[test_end-1].date()}",'Train_Corr': train_corr,'Test_Return': test_combined.mean() *252,'Test_Sharpe': test_combined.mean() / test_combined.std() * np.sqrt(252) if test_combined.std() >0else0 })return pd.DataFrame(results)# Example: Test top pairtop_pair = comp_df.iloc[0]['Pair']alpha1, alpha2 = [a.strip() for a in top_pair.split('/')]wf_results = walk_forward_validation(returns_df, alpha1, alpha2)print(f"\nWalk-Forward Validation: {top_pair}")print(wf_results)print(f"\nAverage Test Sharpe: {wf_results['Test_Sharpe'].mean():.2f}")
Walk-Forward Validation: EMA-Cross-Tech / Mean-Reversion
Period Train_Corr Test_Return Test_Sharpe
0 2021-06-01 to 2022-08-29 0.057506 -0.151942 -1.662360
1 2021-08-30 to 2022-11-28 0.059558 -0.129811 -4.520180
2 2021-11-29 to 2023-03-01 0.041652 -0.042323 -0.432370
3 2022-03-01 to 2023-05-31 0.018414 0.674825 5.440688
4 2022-05-31 to 2023-08-30 0.013267 0.190718 1.804279
5 2022-08-30 to 2023-11-29 0.039571 -0.251398 -2.718211
6 2022-11-29 to 2024-03-01 0.017564 0.347188 3.268694
7 2023-03-02 to 2024-05-31 0.016980 0.274166 2.469639
8 2023-06-01 to 2024-08-30 0.138705 -0.092481 -0.715447
9 2023-08-31 to 2024-11-29 0.258915 0.215528 2.461700
10 2023-11-30 to 2025-03-05 0.286068 0.027320 0.244899
11 2024-03-04 to 2025-06-04 0.248321 0.095622 0.750411
12 2024-06-03 to 2025-09-04 0.074195 0.359873 4.465036
13 2024-09-03 to 2025-12-03 0.008155 0.284876 2.897551
14 2024-12-02 to 2026-03-06 0.023313 -0.096559 -1.202214
Average Test Sharpe: 0.84
10. Next Steps
For Production Implementation:
Build Composite Algorithms: Implement top pairs as QCAlphaModel composites
Dynamic Weighting: Add regime-based allocation between alphas
Risk Management: Add position sizing based on correlation regime
Validation: Run full backtests on recommended pairs
Top Candidates for Composites:
Based on this analysis, prioritize: 1. Trend + Defensive: EMA-Cross + All-Weather 2. Momentum + Mean Reversion: Dual-Momentum + Mean-Reversion
3. Multi-Asset: Dual-Momentum + All-Weather 4. Regime Switching: Trend-Following + Mean-Reversion
Research Extensions:
Add sector-specific alphas (SectorMomentum)
Include options-based stratégies (Options-VGT)
Test with crypto assets (BTC, ETH)
Factor analysis: loadings to Size, Value, Momentum factors