Black-Scholes Delta Is Wrong: Hull-White’s Fix Beats SABR
The correction cuts S&P 500 hedging error up to 42% using three parameters versus SABR’s 87,000 — and most desks ignore it.
The practitioner Black-Scholes delta — the hedge ratio computed by substituting implied volatility into the BS formula and taking the partial derivative — systematically over-hedges S&P 500 call options and under-hedges put options relative to the position that actually minimizes P&L variance. The error follows from one structural feature of equity markets documented continuously since the 1970s: volatility and price move in opposite directions. The correction is a single-equation adjustment derivable from outputs every risk system already produces. Hull and White (2017), working with 1.3 million daily S&P 500 option observations across a data set spanning January 2004 to August 2015, find that switching to the minimum variance delta reduces out-of-sample hedging error variance by 25.7% for calls and 22.5% for puts. For deep out-of-the-money calls, the reduction reaches 42.1%. For actively traded strikes, the corrected hedge outperforms SABR stochastic volatility calibrated daily across every option maturity, using roughly one-third of one percent of the parameters. The practitioner delta remains the default output of standard risk systems. Most desks have not made the switch.
What the Consensus Gets Right
The practitioner Black-Scholes model, which prices each option at its own market-implied volatility and computes Greeks by differentiation, is not naive. It is internally calibrated to market prices by construction. A practitioner computing delta for an SPX put uses that put’s actual implied volatility, not a flat surface, so the delta reflects the option’s exact location on the skew. The implied BSM delta is better than constant-volatility BS delta for this reason.
The standard defense is clear: if practitioners already substitute market implied volatility into the BS formula, have they not already accounted for the smile? This is the question the empirical literature resolves. The answer is no, and the reason is structural rather than parametric.
The Missing Term
The practitioner BS delta is a partial derivative: it measures how an option’s price changes when spot moves while implied volatility is held fixed. The actual market value change of an option when spot moves by dS includes a second term:
dC = (∂C/∂S) · dS + (∂C/∂σ) · dσ
The practitioner delta captures only the first term. The second — vega times the concurrent change in implied volatility — is not zero for equity indices. The negative correlation between equity prices and their implied volatility has been empirically documented across every major equity index for decades, first established by Black (1976) and Christie (1982), confirmed in implied volatility terms by subsequent work. When S rises, σ falls; when S falls, σ rises.
The minimum variance (MV) delta — the hedge ratio that minimizes the daily variance of the hedged position — incorporates both terms:
Δ_MV = Δ_BS + ν × E(∂σ_imp/∂S)
where ν is the practitioner BS vega and E(∂σ_imp/∂S) is the expected change in implied volatility per unit change in spot. For equity indices this expectation is negative, making Δ_MV < Δ_BS for calls. Equity call options are systematically over-hedged when using the practitioner convention; equity put options are systematically under-hedged.
This identity is exact in a two-factor diffusion framework and an approximation under jump-diffusion or non-Markov processes. The approximation is tight for equity index options because the dominant source of variation in implied volatility is the price move, not idiosyncratic vol noise. Hull and White’s regression of implied vol changes on price changes, run across 2007 to 2015, finds that roughly 60% of the total variation in implied volatility changes for deep OTM S&P 500 calls is explained by concurrent index level changes.
The term ν × E(∂σ_imp/∂S) is the P&L contribution the practitioner delta convention ignores on every hedge rebalance. One clarification is worth stating explicitly: the 60% R² comes from a contemporaneous regression — it measures co-movement on the same day, not forecasting power. The correction does not require predicting vol changes in advance; it requires only that the historical relationship between price moves and vol moves, estimated from past data, is stable enough to serve as the expected conditional response going forward. Hull and White’s rolling estimation methodology is designed precisely to test whether that stability holds out-of-sample.
The Empirical Foundation
Precisely how the vol surface moves with spot determines the magnitude of the correction. Derman (1999), in a Goldman Sachs Quantitative Strategies paper also published in Risk Magazine, introduced the conceptual framework for equity index options: the surface can operate in a sticky-strike regime, where each fixed-strike option’s implied vol is independent of spot, or a sticky-delta regime, where moneyness determines implied vol and a spot move carries the entire surface. Neither extreme holds perfectly, but the data strongly favors models in which the surface moves with spot.
Daglish, Hull, and Suo tested multiple conventions against 47 months of S&P 500 over-the-counter consensus implied volatility surfaces from June 1998 to April 2002. The relative sticky-delta model fit the surface with an R² of 94.93% and an out-of-sample RMSE of 0.73 percentage points of implied volatility. The sticky-strike model achieved R² of 27% and RMSE of 5.25 percentage points — a 7.2× difference in out-of-sample RMSE. The paper’s F-statistic for equal explanatory power between the two models is 32.71, a decisive rejection at any conventional significance level. It is worth noting that Daglish et al.’s best-fitting model is actually a third option — the stochastic square-root-of-time rule, which achieves R²=97.12% — but all models that outperform sticky-strike share the same underlying implication: the SPX volatility surface moves with spot, not independently of it.
This is the foundation for the delta correction. When the surface shifts with spot, ∂σ_imp/∂S is consistently negative for fixed-strike options, and the practitioner delta — which assumes ∂σ_imp/∂S = 0 — is wrong in a predictable direction.
What the Numbers Say
Hull and White (2017) estimate E(∂σ_imp/∂S) empirically for S&P 500 options using rolling 36-month windows and find it is well approximated by a quadratic function of the option’s BS delta divided by the product of spot and the square root of time to maturity. This produces a correction formula with three free parameters, estimated once per month:
Δ_MV ≈ Δ_BS + ν × (a·Δ²_BS + b·Δ_BS + c) / (S·√T)
The three coefficients a, b, c are generally stable through time, though Hull and White note extreme parameter shifts during the 2008 credit crisis as the documented exception. They are re-estimated monthly using all strikes and maturities in the prior 36 months, and the results are not sensitive to the choice of window length between 12 and 60 months. One transparency note: Hull and White do not report sub-period performance, so the contribution of the 2008 crisis to the aggregate 25.7% figure is unobservable from the published results. The crisis period produced documented parameter instability; whether the model’s outperformance holds if that period is isolated is a question the paper does not answer. For a desk running tail-risk books, this gap in the published evidence is material.
The out-of-sample test runs from January 2007 to August 2015, covering the 2008 credit crisis. The Gain — Hull and White’s metric for percentage reduction in the sum of squared hedging errors, a variance measure before transaction costs — follows a clear gradient by moneyness: for deep OTM calls (BS delta near 0.1), 42.1%; for delta-0.2 calls, 35.8%; at ATM (delta near 0.5), 27.1%; for deep ITM calls (delta near 0.9), 16.6%. The average across all call strikes is 25.7%. For put options, the average gain is 22.5%, lower because idiosyncratic noise in put implied vol is higher and less of the vol variation is explained by price changes. Hull and White trace this put-call asymmetry to violations of put-call parity in the pre-2009 period; post-2008, the asymmetry narrows.
The gains hold across related index instruments: 23.0% for European-exercise S&P 100 calls (XEO), 16.7% for American-exercise S&P 100 calls (OEX), and 26.5% for DJIA calls. They collapse for individual stocks — 10.3% for calls on Dow components, a statistically negligible 2.5% for puts — and are minimal for interest rate ETFs (1.4%). The correction is an equity index phenomenon, driven by the systematic and stable negative vol-price relationship that characterizes large, liquid indices but is overwhelmed by idiosyncratic noise at the single-stock level.
The Objection: Stochastic Volatility Models Already Solve This
The natural rebuttal is that stochastic volatility models — Heston, SABR — already incorporate the vol-price correlation through the parameter ρ. A trader using SABR-derived delta is in principle computing something close to the minimum variance delta, because the model’s estimated ρ captures exactly E(∂σ/∂S). This is correct in theory. The empirical result is that it fails in practice.
Bakshi, Cao, and Chen (1997) examined three stochastic volatility specifications against S&P 500 options from June 1988 to May 1991 and found that stochastic volatility alone provides the best hedging performance among all models tested — adding jumps or stochastic interest rates does not further improve performance once stochastic vol is included. The model-implied delta with calibrated ρ does reduce hedging error. But the question is not whether SV models beat constant-vol BS; it is whether they match a simple empirical correction.
Hull and White’s abstract states the comparison explicitly: the empirical model outperforms stochastic volatility models “even when the latter are calibrated afresh each day for each option maturity.” The daily calibration of SABR is the deliberately favorable condition for SABR — not a methodological oversight. Hull and White do not test monthly SABR calibration, so the direct comparison is unavailable. What the paper’s own explanation implies, however, is that monthly SABR would likely perform no better. The stated mechanism for SABR’s underperformance is overfitting and model misspecification: “daily recalibration introduces noise into the estimated ρ that the monthly rolling regression avoids.” If overfitting is the cause, reducing calibration frequency would remove noise, potentially shrinking the gap — but the paper’s logic runs in the direction of the daily SABR already being suboptimal relative to a more stable estimator. Monthly SABR may do better or worse; the paper does not resolve this. What can be stated is that SABR with maximum calibration frequency, given every data advantage, still trails the empirical model.
SABR, calibrated daily for every eligible option maturity, requires roughly 87,000 total parameter estimates across the full test period — approximately 40 per trading day on average. The paper’s figure of 78 parameters per day refers to the maximum on days when all 13 tracked maturities pass Hull and White’s data quality filters; on average, roughly 6 to 7 qualifying maturities are available per side per day, not 13. The empirical model, by contrast, estimates three coefficients once per month. It achieves a hedging gain of 24.6% for calls and 19.0% for puts versus the empirical model’s 25.7% and 22.5%. The Newey-West adjusted t-statistics for the difference exceed 8 for all calls and 11 for all puts, each significant at any conventional threshold. The only buckets where SABR leads are deep-in-the-money options — a region of thin volume where Hull and White explicitly note the exception.
Alexander, Rubinov, Kalepky, and Leontsinis (2012) confirm the same direction in a different market: 16.5 years of FTSE 100 options data, where Markov-switching smile-adjusted deltas reduce hedging errors to roughly 50–60% of implied BSM hedging errors on average across all regimes, with substantially greater improvement during volatile periods. The vol-price elasticity is not constant — it is weaker in trending markets and stronger in volatile ones — and the rolling estimation window in Hull and White’s model captures this implicitly.
Why the Gap Persists
The persistence of the practitioner BS delta as the industry standard is not informational. The empirical literature establishing its inferiority spans two decades: Bakshi, Cao, and Chen (1997), Coleman, Kim, Li, and Verma (cited in Hull and White 2017) on S&P 500 options as early as 2001, Crépey (2004), Alexander et al. (2012), and Hull and White (2017). The gap between knowing and implementing — which has persisted across those two decades of cumulative evidence — has three institutional sources.
The practitioner delta is the default output of standard risk systems. Compliance tests, delta limits, and intraday P&L attribution are built against this number. Switching from the partial derivative convention to a minimum variance convention requires changing not just a model but a risk infrastructure, including audit trail requirements and regulatory approval for VaR frameworks.
The correction is index-specific. A mixed book of single-stock and index options captures diluted benefits: individual equity calls show 10.3% gain, individual puts 2.5%. For a desk with material single-stock optionality, the payoff-to-implementation ratio is lower than for a pure index book.
The corrected delta is lower than the practitioner delta for calls. In a rising market, a smaller delta hedge costs less to carry and generates better hedge P&L. In a flat market, the smaller hedge produces lower variance reduction. Managers who evaluate hedging quality on sharp down-days — when delta is clearly insufficient — are not the same managers who evaluate on variance reduction across all trading days. The improvement appears in the variance metric, not the crisis-day metric.
The result is a structural variance cost that accumulates in books using practitioner delta relative to those running smile-aware delta — not a direct extraction of P&L from counterparties with worse hedges, since the underlying’s move dominates any individual rebalance’s realized P&L, but a compounding statistical drag across thousands of daily hedges. Dealers running proprietary smile-aware implementations carry lower realized hedging variance over time; the economic value of that difference is not directly observable from the published evidence but shows up in lower residual risk per unit of notional.
What Would Change This View
Three conditions would reduce or eliminate the minimum variance correction.
The correction depends on a stable negative vol-price correlation for equity indices. If this correlation reverted to zero or turned positive — under a structural regime shift where volatility becomes demand-driven and disconnected from the leverage effect — the term E(∂σ_imp/∂S) would vanish and the practitioner delta would become optimal. The negative vol-price relationship has been stable across decades for large-cap equity indices, but it is not a mathematical law.
If sub-14-day options dominate the book, the correction loses efficacy. Hull and White exclude sub-14-day options from their test and note that including them worsens results because large near-the-money gamma generates P&L variance that a delta correction alone cannot address. The 2008 credit crisis also produced documented extreme parameter shifts in the correction coefficients, a caveat Hull and White flag explicitly. The model is stable under normal conditions but not immune to regime breaks.
If the correction became widely implemented, the systematic pricing asymmetry between buy-side and dealer books would narrow. The current state — that even SABR, which theoretically captures the correction through ρ, underperforms the simple empirical formula — suggests the correction is underexploited even at well-resourced institutions. That underexploitation is what keeps the gap alive.
The Actionable Implication
A PM running vanilla equity index options can compute the minimum variance delta directly from standard risk system outputs. The inputs are the practitioner BS delta (Δ_BS), the practitioner BS vega (ν), current spot (S), and time to maturity (T). The three coefficients must be estimated from historical data, but Hull and White’s published findings establish that estimates derived from any 36-month window of daily option closing data are robust across window-length choices between 12 and 60 months and apply consistently across all strikes and maturities simultaneously.
The correction matters most where it is most frequently ignored: for deep OTM index calls (BS delta near 0.1), the standard convention over-hedges by the amount that generates a 42.1% variance penalty. These are the instruments used by institutional desks as upside participation and by systematic vol strategies as short-gamma positions. The over-hedging of OTM calls and under-hedging of OTM puts is not noise — it is a structural feature of the practitioner convention applied to a vol surface that demonstrably moves with spot.
The 25.7% reduction in sum-of-squared hedging errors for S&P 500 call options, achieved out-of-sample from 2007 to 2015, is not sensitive to parameter choice or window length. It measures the difference between measuring what an option actually does when spot moves and assuming the vol surface is frozen while it moves. The former is a three-number estimate computable from three years of daily option closing prices — matching the rolling window Hull and White used to produce every headline figure in the paper. The latter is the industry default.
Primary sources: Hull, J. and White, A. (2017), “Optimal Delta Hedging for Options,” Journal of Banking and Finance, 82, 180–190. Daglish, T., Hull, J. and Suo, W., “Volatility Surfaces: Theory, Rules of Thumb, and Empirical Evidence,” working paper, University of Toronto. Bakshi, G., Cao, C. and Chen, Z. (1997), “Empirical Performance of Alternative Option Pricing Models,” Journal of Finance, 52(5), 2003–2049. Alexander, C., Rubinov, A., Kalepky, M. and Leontsinis, S. (2012), “Regime-Dependent Smile-Adjusted Delta Hedging,” Journal of Futures Markets, 32(3), 203–229. Crépey, S. (2004), “Delta-Hedging Vega Risk,” Quantitative Finance, 4, 559–579.
Connect with me on LinkedIn, or for video breakdowns of research like this, subscribe on YouTube.
📊 Want Deeper Quantitative Analysis?
This research took a long time of data collection, verification, and analysis. If you found value in this deep-dive, I publish exclusive quantitative research, trading strategies, and institutional-grade analysis on Patreon.
By joining, you’ll be supporting my work and motivating me to publish more content like this.


