BLOG

Behind the numbers.

Code examples, market analysis, and data quality deep-dives.

How Far Apart Do S&P 500 Stocks Move? Cross-Sectional Return Dispersion in Python
Fama-French Factor Data: Download or Build Your Own?
How Much Does the Rebalance Date Change a Backtest? 21 Rebalance Days in Python
Does Trading Volume Predict Tomorrow's Volatility? Out-of-Sample Test in Python
Do Company Insiders Predict Their Own Stock's Returns? Form 4 Cross-Section in Python
How Many Independent Bets Are There in the S&P 500? Principal Component Analysis in Python
Can a Company's Revenue Be Forecast From Its Own History? Out-of-Sample Test in Python
What Is EBITDA and Why Do Sources Disagree?
Does the Turn-of-the-Month Effect Still Work? Calendar Anomaly Test in Python
Is the S&P 500 Getting More Capital Intensive? Capex Analysis in Python
Do High-Accrual Companies Underperform? Accruals Screening in Python
What Does a Financial Data API Cost?
Do Price Gaps Get Filled? Gap-Fill Rates Against a Random Walk in Python
What Expected Returns Does the S&P 500 Imply? Reverse Optimization in Python
Which Sectors Are Really Cyclical? Revenue Betas vs Stock Betas in Python
How to Build a Stock Dataset for Machine Learning
Comparing Companies With Different Fiscal Year Ends
How Much Has Corporate Debt Actually Repriced? Effective Interest Rates in Python
Does Deferred Revenue Predict Next Quarter's Sales? Leading Indicator Test in Python
How Much Debt Is Hidden in Operating Leases? Lease-Adjusted Leverage in Python
How Much Do Profits Move When Sales Move? Operating Leverage Regression in Python
How Concentrated Is the S&P 500? Index Weight Analysis in Python
How Long Is Cash Tied Up in a Business? Cash Conversion Cycle Analysis in Python
Broker API vs Data API for Historical Stock Data
Where to Get Free Cash Flow Data for Stocks in Python
Are Stock Returns Skewed? Return Skewness in Python
Do High-Margin Companies Trade at Higher Multiples? EV/Sales in Python
How Much Profit Becomes Cash? Free Cash Flow Conversion in Python
Which Sectors Lead Out of a Market Bottom? Sector Recovery Analysis in Python
Are Buybacks Funded by Cash Flow or by Debt? S&P 500 Payout Analysis in Python
What to Look for in Fundamentals Data
Does Cash on the Balance Sheet Cushion a Crash? Quintile Sorts in Python
Do Corporate Insiders Time the Market? S&P 500 Insider Buying Breadth in Python
Does Unstable Volatility Warn of Deeper Drawdowns? Vol-of-Vol Sorts in Python
Does Fast Asset Growth Predict Weak Stock Returns? Decile Sorts in Python
How to Replace yfinance in a Python Script
Does the Nasdaq-100 Index Effect Still Exist? Event Study in Python
Does the Piotroski F-Score Still Work? Quality Screening in Python
How Many S&P 500 Stocks Beat the Index? Return Breadth Analysis in Python
Why Beta Differs Between Data Sources
Where to Get Historical Dividend Data for Stocks
Does a High Dividend Yield Predict a Dividend Cut? Yield-Trap Screening in Python
Do Old Ticker Symbols Still Point to the Same Company? S&P 500 Ticker Recycling in Python
What Growth Is Priced Into the S&P 500? Reverse DCF in Python
Ticker vs CIK vs FIGI: Which Company ID to Use
Does Illiquidity Still Pay? Amihud Measure in the S&P 500 in Python
Where to Get R&D Spending Data for Public Companies
Does R&D Spending Predict Revenue Growth? Cross-Sectional Test in Python
What Actually Drives Return on Equity? DuPont Decomposition in Python
What Data Do You Need to Measure Portfolio Risk?
Does a 60/40 Portfolio Actually Cut Drawdowns? Stocks and Bonds in Python
Do Steady Margins Mean Calmer Stocks? Cross-Sectional Analysis in Python
Do Sectors Diversify When It Matters? Conditional Correlation in Python
How to Get Historical Market Cap Data in Python
How Much of the S&P 500 Survives 20 Years? Index Turnover Analysis in Python
Does Joining the S&P 500 Bring New Institutional Owners? 13F Event Study in Python
Is Volatility Seasonal? Calendar Month Analysis of Realized Volatility in Python
Does Fast Revenue Growth Force Companies to Borrow? Cash Funding Analysis in Python
How Far Back Does SEC EDGAR Data Go?
Are One-Time Charges Really One-Time? Charge Frequency Analysis in Python
Does Buying the Dip Work? Short-Term Reversal by Volatility Regime in Python
Alpha Vantage vs Massive vs xfinlink for Fundamentals
How Long Does a Stock Take to Recover From a 50% Fall? Drawdown Analysis in Python
Do Companies That Shrink Their Share Count Outperform? Net Buyback Yield in Python
How Much Does the Dow's Price Weighting Distort It? Index Weighting Analysis in Python
How to Get SEC Form 4 Insider Trading Data in Python
Can Anything Predict Next Month's Stock Returns? Out-of-Sample R-Squared Testing in Python
How Much of a Stock's Return Comes From Its Sector? Variance Decomposition in Python
Altman Z-Score: Where To Get It in Python
Do Value Screens Agree on Which Stocks Are Cheap? Multiple Overlap Analysis in Python
Annual vs Quarterly Financial Data: Which to Use
Do High Returns on Capital Persist? ROIC Fade Analysis in Python
Does Past Beta Predict Future Beta? Beta Stability Testing in Python
Do Defensive Sectors Actually Defend? Up and Down Capture in Python
How to Choose a Financial Data API
Do Small Caps Actually Beat Large Caps? Size Premium Test in Python
What If You Miss the Market's Best Days? Extreme-Day Analysis in Python
Does Rebalancing Add Return? Fixed-Weight vs Drift Portfolios in Python
Which S&P 500 Companies Are Closest to Default? Merton Distance-to-Default in Python
What Happens to Stocks Removed From the S&P 500? Replacement Pair Analysis in Python
Does the Golden Cross Work? 50/200 Moving Average Crossover Backtest in Python
Financial Data for Academic Finance Research
Does Skipping the Most Recent Month Improve Momentum? S&P 500 Decile Sorts in Python
Does Cointegration Survive Out of Sample? Pairs Trading Validation in Python
Which Dividends Are Not Covered by Cash? Free-Cash-Flow Coverage Screening in Python
GICS vs SIC vs NAICS: Which Industry Classification to Use
How Many Stocks Does It Take to Diversify? Random Portfolio Simulation in Python
How Many Days of Data Does a Volatility Estimate Need? Range-Based Estimators in Python
How Much of S&P 500 Cash Flow Is Stock Compensation? Cross-Sectional Analysis in Python
What Is a 13F Filing? Institutional Holdings Explained
Does Revenue Growth Explain Profit Growth? Cross-Sectional Decomposition in Python
How Much of the Nasdaq 100 Is Already in the S&P 500? Index Overlap Analysis in Python
How Often Does a 99% Value-at-Risk Limit Actually Break? VaR Backtesting in Python
Real-Time vs End-of-Day Market Data: Which Do You Need?
How Concentrated Are S&P 500 Earnings? Point-in-Time Index Analysis in Python
Does Volatility Scale With the Square Root of Time? Variance Ratio Test in Python
Does Goodwill Distort the Price-to-Book Screen? Goodwill-Adjusted Valuation in Python
How Are Shares Outstanding Reported (and Why They Disagree)
Do Low-Volatility Stocks Deliver Better Risk-Adjusted Returns? S&P 500 Quintile Sorts in Python
Does Trend Following Beat Buy and Hold? Time-Series Momentum in Python
Has the Stock-Bond Correlation Flipped? 60/40 Portfolio Risk in Python
What API to Use for a Stock Screener
Which Assets Hedge Inflation Shocks? Macro Factor Betas in Python
Does Covariance Shrinkage Beat the Sample Covariance? Minimum-Variance Portfolios in Python
Do Faster Inventory Turns Mean Thinner Margins? Gross Margin Return on Inventory in Python
SEC EDGAR API vs Fundamentals API: Which to Use
Does Fast Earnings Growth Persist? Rank Correlation Analysis in Python
Split Adjustment Explained: Adjusted Close vs Close
Which Trading Day of the Month Pays Best? Turn-of-the-Month Analysis in Python
Can You Use Yahoo Finance Data Commercially?
How Many Independent Bets Does a Nine-Sector Portfolio Give You? Eigenvalue Analysis in Python
Which Volatility Forecast Wins One Month Ahead? HAR vs EWMA in Python
How Concentrated Are Institutional Equity Portfolios? Form 13F Concentration Analysis in Python
Do Stocks Earn Their Returns Overnight or Intraday? Return Decomposition in Python
When Do Corporate Insiders Actually Trade? Form 4 Timing Analysis in Python
Data Requirements for Backtesting a Trading Strategy
What Is Survivorship Bias in Backtesting?
Do High Dividend Yields Come From Bigger Payouts or Falling Prices? Yield Decomposition in Python
Do Stocks Fall Harder Than They Rise? Downside Beta vs Upside Beta in Python
Does Volatility Targeting Improve Sharpe Ratios? Seven-Asset Backtest in Python
Free Stock Market Data APIs: What You Actually Get
How to Give an LLM Financial Data With an MCP Server
Does Post-Earnings Announcement Drift Survive Real Filing Dates? PEAD Event Study in Python
Do Insider Buying Clusters Predict Returns? Signal Testing in Python
Does the S&P 500 Index Effect Still Exist? Event Study in Python
Are Companies Leaving the S&P 500 Faster Than They Used To? Index Survival Analysis in Python
Does Gross Profitability Predict Stock Returns? Quintile Factor Test in Python
What Growth Rate Is the Market Pricing In? Reverse DCF in Python
Does a Strong Balance Sheet Cushion Drawdowns? Leverage and Downside Risk in Python
Does Ticker Recycling Corrupt a Mean-Reversion Backtest? Entity-Resolved Z-Scores in Python
How Much of a Growth Screen's Backtested Edge Is Survivorship Bias? Point-in-Time Index Testing in Python
Why Do Leveraged ETFs Decay? Measuring Volatility Drag in Python
Are Consumer Staples Margins Shrinking Under Inflation? Gross Margin Trend Analysis in Python
Do Weak Jobs Reports Predict Market Drawdowns? NFP Surprise Event Study in Python
Is the Rotation From Tech to Industrials Backed by Earnings? Relative EPS Growth Analysis in Python
Is the Semiconductor Rally Broadening Beyond NVIDIA? Return Dispersion Analysis in Python
Which Stocks Benefit Most When Oil Prices Fall? Oil Beta Screening in Python
Do Bond Returns Predict Stock Returns? Granger Causality Test in Python
Which Stocks Actually Drive Portfolio Returns? Shapley Value Attribution in Python
Does "Sell in May" Still Work? Calendar Anomaly Backtest in Python
How to Build Complete Price History Through Ticker Changes? Entity Resolution in Python
Are KO and PEP Cointegrated? Pairs Trading Signal Construction in Python
Which Commodities Have the Strongest Momentum? Rotation Backtest in Python
Which Commodity ETFs Have the Worst Tail Risk? Expected Shortfall in Python
Are Gold Miners Leveraged Gold Bets? Rolling Beta Analysis in Python
Does the Base-Metals-to-Gold Ratio Lead Cyclical Stocks? Signal Test in Python
Can Risk Parity Tame Commodity Volatility? Portfolio Optimization in Python
Are Power Stocks Becoming an AI Infrastructure Trade? Momentum Screening in Python
Which AI Chip Stocks Have Margin Momentum? Profitability Trend Analysis in Python
Which AI Stocks Are Cheapest Relative to Growth? Growth-Adjusted Valuation in Python
Does AI Stock Leadership Persist? Momentum Backtest in Python
Which AI Stocks Have the Cleanest Balance Sheets? Net Cash Screening in Python
Can Risk Parity Reduce Mega-Cap Drawdowns? Portfolio Optimization in Python
Which Growth Stocks Are Self-Funding? Cash-Flow Quality Screening in Python
Which Sectors Struggle When the Dollar Rallies? Sector Rotation Analysis in Python
Do Cheap Stocks Hold Up When Bonds Sell Off? Valuation Rotation in Python
Does the Nasdaq 100 Have Better Growth Quality Than the Dow? Index Constituent Analysis in Python
Do Healthcare Cash-Flow Margins Predict Returns? Signal Evaluation in Python
Which Dividend Stocks Survive a Cash-Flow Stress Test? Dividend Screening in Python
Does Heavy Insider Selling Predict Weak Returns? Insider Flow Test in Python
Can Quality Screens Reduce Small-Cap Balance-Sheet Risk? Russell 2000 Test in Python
Which Retailers Have Positive Operating Leverage? Margin Screening in Python
Is MSTR a Leveraged Bitcoin Proxy? Rolling Beta Analysis in Python
Is Micron's Memory Cycle Recovering? Inventory and Margin Forecasting in Python
Which Sectors Work When Bonds Rally? Rate-Sensitive Rotation in Python
Do One-Month Price Extremes Reverse? Signal Evaluation in Python
Do Low-Volatility S&P 500 Stocks Reduce Drawdowns? Factor Test in Python
Is AI Capex Paying Back Fast Enough? Revenue Hurdle Forecasting in Python
Could Shorter AI Asset Lives Hit Earnings? Depreciation Stress Test in Python
How Much AI Capex Risk Can a Portfolio Remove? Constrained Optimization in Python
Is the AI Capex Trade Crowded? Rolling Volatility and Sector Rotation in Python
Did the AI Boom Come From Existing S&P 500 Members? Point-in-Time Momentum Test in Python
Is AI Revenue Circular? Customer-Vendor Capex Loop Analysis in Python
Is the AI Trade Connected to Private Credit? Rolling Correlation Network in Python
Is Apollo More Balance-Sheet Sensitive Than Peers? Leverage Screen in Python
Are AI Earnings Supported by Cash Flow? Accrual and Capex Screen in Python
Can Defensive Stocks Hedge AI Drawdowns? Basket Regime Test in Python
How Fast Does the Market Price In Fed Decisions? FOMC Event Study in Python
How Much Are Options Sellers Overpaid? The Variance Risk Premium in Python
Which Companies Have the Worst Earnings Quality? Sloan Accrual Screen with Geographic Revenue Data in Python
Does the Oil-to-Gold Ratio Signal Recessions? XLE/GLD Backtest in Python
Is AI Spending Crowding Out Free Cash Flow? Capex Sustainability Across the Mag 7 in Python
Does a Long Energy / Short Bonds Portfolio Capture Inflation Surprises? Factor Construction in Python
Can a Hidden Markov Model Detect Oil Market Regimes? HMM Analysis in Python
Do Grain Prices Predict Food Inflation? Granger Causality Test in Python
Does the Corporate Credit Spread Predict Stock Market Crashes? BAA-AAA Spread Analysis in Python
Do Oil Stocks Hedge Inflation? Rolling Beta Analysis in Python
Which Stocks Are Most Rate-Sensitive? Equity Duration via Bond Beta in Python
Which Companies Have the Highest Accrual Ratios? Earnings Quality Screening in Python
Is Alpha Persistent or Decaying? Rolling Sharpe Ratio Analysis in Python
Are Markets Trending or Mean-Reverting? Hurst Exponent Analysis in Python
Is Consumer Discretionary vs Staples a Leading Indicator? XLY/XLP Ratio Analysis in Python
Does Heavy Capex Predict Future Stock Returns? Capital Expenditure Analysis in Python
How to Estimate Cost of Equity Using CAPM in Python
Is Volatility Predictable? Testing for Volatility Clustering in Python
Which Industrials Are Overleveraged? Net Debt to EBITDA Screening in Python
GM Before and After Bankruptcy: Why Entity Resolution Matters for Financial Data
What Is Adjusted Beta? Merrill Lynch Beta Shrinkage in Python
How Good Is a Stock Pick? Information Ratio and Tracking Error in Python
Do Stock Returns Follow a Normal Distribution? Testing for Fat Tails in Python
Which Large Caps Have the Highest Free Cash Flow Yield? FCF Screening in Python
Which Sectors Won Over 5 Years? Sector Rotation Analysis in Python
How to Forecast Stock Volatility with GARCH Models in Python
Are Stock Prices Mean-Reverting? Augmented Dickey-Fuller Test in Python
How to Calculate CAPM Alpha and Beta with Regression in Python
How to Compare Sector Sharpe Ratios and Sortino Ratios in Python
DELL: Why Stitching Historical Price Data Together Is Wrong
How to Analyze Drawdown and Recovery for Bank Stocks in Python
How to Screen SaaS Stocks by Revenue Growth and Cash Flow in Python
How to Screen REITs by Dividend Yield and Valuation in Python
How Correlated Are the Magnificent 7? Intra-Group Correlation in Python
AAPL vs XOM: Do Individual Stocks Have Seasonal Patterns?
How to Rank Large-Cap Stocks by Momentum in Python
How to Build a Multi-Endpoint Financial Dashboard in Python
How to Compare Volatility Across Energy Stocks in Python
How to Screen Healthcare Stocks by Valuation in Python
How to Build a Sector Correlation Matrix for Portfolio Diversification in Python
How to Find Oversold and Overbought Stocks Using Z-Scores in Python
How to Measure Earnings Quality: Cash Flow vs Net Income in Python
How to Build a Multi-Factor Stock Screen in Python (Value + Momentum + Quality)
How to Build a Simple DCF Model for Any Stock in Python
How to Screen Tech Stocks by Revenue Growth in Python
How to Screen Stocks by Balance Sheet Health in Python
Is "Sell in May" Real? SPY Monthly Seasonality Over 10 Years
How to Compare Sector Performance YTD Using Python
How to Screen Dividend Stocks by Yield and Quality in Python
How to Calculate Max Drawdown and Recovery Time for Any Stock in Python
How to Compare Profitability Across Mega-Cap Tech Stocks in Python
Why Ticker Symbols Are Unreliable: The Recycling Problem Every Quant Should Know
How to Calculate and Compare Stock Volatility in Python
How to Screen Blue-Chip Stocks by P/E Ratio in Python
How to Track Companies Through Ticker Changes, Bankruptcies, and Renames in Python
S&P 500 Turnover: How Much the Index Has Changed Since 2010
How to Calculate Stock Beta and Correlation in Python
← All articles

How Far Apart Do S&P 500 Stocks Move? Cross-Sectional Return Dispersion in Python

What’s the question?

Two numbers describe an equity market over a year, and they are routinely treated as one. Volatility measures how sharply the index moves from day to day. Dispersion measures how far apart its members finish from each other over the same window: the standard deviation of member returns taken across the index at a point in time rather than through time.

Dispersion sets the size of the prize in stock selection. If every member lands within a few points of the average, correct picks cannot add much and wrong picks cannot cost much; when the members finish a hundred points apart, the same amount of skill produces a far larger result in either direction.

Market commentary tends to assume the two travel together, and that a turbulent market rewards selection. That assumption is testable, along with the related idea that a violent year hands the next one a wide cross-section.

The approach

The sample is the 20 calendar years from 2006 to 2025.

  1. Rebuild S&P 500 membership on the first day of each year from the roster as it stood on that date, carrying every member by company identifier rather than ticker, so a company removed later still counts for the years it was a member.
  2. Pull daily total returns for each member and compound them into a full-year return inside the script, which keeps a company that changes its symbol mid-year intact.
  3. Apply two screens: a member needs a price series covering at least 95 percent of the year’s trading days, and its daily returns have to agree with its own price path day by day.
  4. Compute dispersion as the standard deviation of those annual returns across members, and the decile spread as the average return of the best tenth minus the average of the worst tenth.
  5. Measure the index itself over the same years through SPY, using annualised daily volatility.

That yields 9,532 member-years, between 443 and 495 members per year.

Code

import time
import numpy as np
import pandas as pd
from scipy import stats
import xfinlink as xfl

xfl.set_api_key("YOUR_API_KEY")  # free at https://xfinlink.com/signup

YEARS = list(range(2006, 2026))
CHUNK = 50


def fetch(**kwargs):
    for attempt in range(5):
        try:
            return xfl.prices(**kwargs)
        except xfl.XfinlinkError as exc:
            last = exc
            time.sleep(4 * (attempt + 1))
    raise last


rows = []
for year in YEARS:
    roster = xfl.index("sp500", as_of=f"{year}-01-01")
    ids = sorted({int(e) for e in roster["entity_id"].dropna()})
    px = pd.concat([fetch(entity_id=ids[i:i + CHUNK], start=f"{year}-01-01",
                          end=f"{year}-12-31", fields=["adj_close", "return_daily"],
                          max_rows=200_000)
                    for i in range(0, len(ids), CHUNK)], ignore_index=True)
    px = (px.drop_duplicates(["entity_id", "date"]).dropna(subset=["return_daily"])
            .sort_values(["entity_id", "date"]))

    days = px["date"].nunique()
    counts = px.groupby("entity_id")["return_daily"].size()
    step = np.log1p(px["return_daily"]) - np.log(px["adj_close"]).groupby(px["entity_id"]).diff()
    agrees = step.abs().groupby(px["entity_id"]).max()
    kept = counts[counts >= 0.95 * days].index.intersection(agrees[agrees <= 0.5].index)

    ann = px[px["entity_id"].isin(kept)].groupby("entity_id")["return_daily"].apply(
        lambda s: float(np.prod(1.0 + s.values) - 1.0))
    n10 = int(round(len(ann) * 0.1))
    rows.append(dict(year=year, members=len(ann), dispersion=ann.std(ddof=1),
                     top=ann.nlargest(n10).mean(), bottom=ann.nsmallest(n10).mean()))

t = pd.DataFrame(rows).set_index("year")
t["spread"] = t["top"] - t["bottom"]

spy = fetch(ticker="SPY", start="2006-01-01", end="2025-12-31",
            fields=["return_daily"], max_rows=200_000).dropna(subset=["return_daily"])
t["index_vol"] = spy.groupby(spy["date"].dt.year)["return_daily"].std(ddof=1) * np.sqrt(252)

print(t.round(3))
print(stats.pearsonr(t["dispersion"], t["index_vol"]))
print(stats.pearsonr(t["dispersion"].values[1:], t["index_vol"].values[:-1]))

Full script with formatting and visualisation: sp500-return-dispersion-python.py

Output

Two panels covering 2006 to 2025. The upper panel plots the spread between S&P 500 member returns against the volatility of the index itself; the two lines cross repeatedly, with dispersion peaking at 56 percent in 2009 while index volatility peaks at 41 percent in 2008. The lower panel shows the gap between the best and worst tenth of members each year, ranging from 76 points in 2018 to 188 points in 2009.
S&P 500 cross-sectional return dispersion, 2006-2025
point-in-time membership, 9,532 member-years, 250-253 trading days per year

year  members  dispersion  index vol  index return   top decile  bottom decile  decile spread
---------------------------------------------------------------------------------------------
2006      447       22.5%      10.0%       +15.8%       +57.1%        -22.0%         79.1pp
2007      443       34.1%      15.9%        +5.1%       +70.1%        -52.4%        122.5pp
2008      457       25.1%      41.3%       -36.8%        +5.1%        -82.0%         87.1pp
2009      464       55.8%      26.6%       +26.4%      +171.2%        -17.0%        188.2pp
2010      466       24.5%      17.9%       +15.1%       +68.7%        -17.1%         85.8pp
2011      469       24.7%      23.0%        +1.9%       +43.5%        -44.0%         87.5pp
2012      474       25.8%      12.7%       +16.0%       +68.9%        -24.8%         93.6pp
2013      479       32.7%      11.1%       +32.3%      +101.1%         -8.5%        109.7pp
2014      485       22.4%      11.2%       +13.5%       +53.2%        -27.6%         80.7pp
2015      474       26.0%      15.4%        +1.3%       +43.5%        -48.7%         92.2pp
2016      475       25.1%      13.1%       +12.0%       +59.7%        -25.7%         85.4pp
2017      484       26.9%       6.7%       +21.7%       +68.5%        -28.4%         97.0pp
2018      478       21.8%      17.0%        -4.6%       +32.3%        -43.8%         76.1pp
2019      486       25.4%      12.5%       +31.2%       +75.0%        -17.0%         92.0pp
2020      490       29.2%      33.4%       +18.4%       +68.6%        -36.5%        105.1pp
2021      493       29.2%      13.0%       +28.7%       +86.9%        -14.7%        101.6pp
2022      490       27.8%      24.2%       -18.2%       +45.0%        -52.4%         97.4pp
2023      495       31.7%      13.1%       +26.2%       +80.7%        -29.3%        110.0pp
2024      494       28.7%      12.6%       +24.9%       +68.8%        -33.0%        101.8pp
2025      489       35.4%      19.5%       +16.4%       +83.4%        -35.9%        119.3pp

widest dispersion    2009  55.8%   (index volatility 26.6%)
narrowest dispersion 2018  21.8%   (index volatility 17.0%)
loudest index        2008  index volatility 41.3%, dispersion 25.1%

average dispersion 28.7%, median 26.4%

dispersion against index volatility, same year   r=+0.22  p=0.351  rank r=+0.19
dispersion against index volatility, prior year  r=+0.58  p=0.009  rank r=+0.08
   leaving one year out, that r runs +0.02 to +0.64; without 2009 alone it is +0.02
dispersion against its own prior year            r=-0.20  p=0.400

decile spread divided by dispersion: mean 3.51, range 3.35-3.63  (3.51 if annual returns were normal)

What this tells us

Dispersion is never small. The narrowest year of the twenty, 2018, still put a standard deviation of 21.8 percent across the members, and the average year sits at 28.7 percent.

The comparison with index volatility is where the two ideas separate. Their correlation across the 20 years is +0.22 with a p-value of 0.35, and the rank correlation is +0.19; neither is distinguishable from zero. Two consecutive years show why. 2008 was the most violent year for the index in the sample at 41.3 percent volatility, yet its members finished 25.1 percent apart, below the 20-year average, because nearly everything fell at once. 2009 was calmer for the index at 26.6 percent volatility and produced the widest cross-section here, 55.8 percent, because the recovery was wildly uneven: the best tenth of members gained 171 percent while the worst tenth still lost 17 percent.

The lagged test looks more promising at first reading and then falls apart. Dispersion against the prior year’s index volatility gives r = +0.58 with p = 0.009, which would pass a casual significance check. The rank correlation on those pairs is +0.08, and refitting 19 times while holding out one year each time sends it as low as +0.02, which is what dropping 2009 alone produces. One pair of observations carries the result. Dispersion does not forecast its own next value either, at r = −0.20.

One relationship in the table is stable. The decile spread divided by dispersion averages 3.51 and stays between 3.35 and 3.63 in every year, and a normal distribution produces exactly 3.51. The gap between the best and worst tenth therefore carries no information beyond the standard deviation itself, which holds in 2009 with a member up 420 percent in it and in 2018 with nothing above 80 percent.

So what?

Dispersion converts directly into the size of a selection decision. Multiply the year’s dispersion by 3.5 and the result is the expected gap between a top-decile pick and a bottom-decile pick: 76 points in 2018, 188 points in 2009, 119 points in 2025. Any active risk budget or position limit is set against that number whether or not it has been measured.

Volatility is not a usable proxy for it. A calm index does not mean the members move together, which 2017 shows cleanly: the quietest index in the sample at 6.7 percent volatility, with a middling 26.9 percent dispersion beneath it. Waiting for a volatility spike before taking concentrated positions means waiting on an unrelated signal.

Since nothing tested here forecasts dispersion, it has to be measured as it happens. The same calculation runs on a rolling 60-day window of daily returns instead of a calendar year, giving a live reading of what the cross-section currently pays for being right. Without it, a 5-point win in 2018 and a 5-point win in 2009 look like the same achievement.

Built with xfinlink — free financial data API for Python. pip install -U xfinlink

Built with xfinlink — free financial data API for Python. pip install -U xfinlink
← All articles