§3.2 — Research
Point Shaving in NCAA
Retested Wolfers (2006)'s point-shaving test on 95,781 D-I games — the identifying divergence turns out to be mechanical, not corruption, once heteroskedasticity is modeled correctly.
95,781 D-I games, 2003-2025 · 42 spread models averaged into one line · z_absline coefficient drops 70.5% (0.668 → 0.197) once heteroskedasticity is modeled
joint work with a classmate (Kellen) · data from Todd Beck's prediction tracker (thepredictiontracker.com) · not a solo project
1. why the standard point-shaving test is broken
›a mechanical fact, not a judgment call
2. the real trial log — 70,207 favorite-won games, four models
›Model 1 — standard probit (Wolfers baseline)LL -39438.4
If Wolfers is right, standardized absolute spread (z_absline) should positively predict winning-without-covering.
P(Y=1) = Φ(β0 + β1·z_absline + β2·z_absline²), Y = 1 if favorite won but didn't cover, conditional on winning.
z_absline: 0.668*** (t=87.09) — large and significant, exactly what Wolfers' reading would predict.
›Model 2 — heteroskedastic probitLL -39013.4
Large-spread blowouts (garbage time, pulled starters) should inflate outcome variance, not just shift the mean — conflating the two is exactly the flaw in Model 1.
Same mean equation, but the latent variance is now itself a function of spread: P(Y=1) = Φ((β0+β1·z)/exp(γ0+γ1·z+γ2·z²)).
z_absline in the mean equation drops from 0.668 to 0.197*** (t=9.11) — a 70% reduction. The variance equation confirms why: z_absline predicts outcome dispersion with coefficient 1.009*** (t=23.45).
›Model 3 — skew-probit, constant αLL -39166.7
Even after correcting variance, the outcome distribution might not be symmetric — the heteroskedastic probit still assumes it is.
Azzalini (1985) skew-normal link with a single shape parameter α estimated jointly with the mean equation.
α = −0.728*** (t=−25.99) — statistically significant asymmetry the first two models structurally can't represent.
›Model 4 — skew-probit, α varying by spreadLL -38962.4
The direction and degree of asymmetry might itself change across the spread distribution — close games and blowouts could be skewed differently.
α = γ0 + γ1·z + γ2·z² + γ3·z³, a cubic in standardized spread, estimated jointly with the mean equation.
Best fit of all four models. Every α-equation coefficient significant at p<0.001 — the shape of asymmetry genuinely shifts with spread size, not just its presence.
3. what the correction actually looks like
σ(z) — outcome variance rises with spread
α(z) — skewness shifts with spread, not constant
Both curves computed directly from the paper's own real fitted coefficients (Table 2, Models 2 & 4), plotted over the standardized spread range (z_absline) most favorite-won games actually fall in. Neither is flat — variance genuinely grows with spread size, and the skew genuinely changes sign and magnitude across the distribution rather than sitting at one constant value.
headline numbers
conclusion
Wolfers' identification strategy is fundamentally flawed — the divergence it relies on is mechanical, reproduced exactly by a zero-manipulation simulation. There is genuine, statistically significant asymmetry in outcomes (the skew-probit's real improvement in fit proves that), but the paper is explicit that asymmetry ≠ point shaving: it's equally consistent with garbage-time substitution patterns, coaches resting starters in blowouts, or betting markets pricing more accurately over time. No concrete evidence of manipulation either way.