§3.1 — Research
March Madness Predictor
Bradley-Terry point-margin ratings over 57k D-I games, tuned by real backtested MAE — real 2026 Final Four predictions.
57,136 games, 2015–2025 · tuned MAE 9.03 pts on 1,250 real held-out postseason games
real 2026 Final Four predictions below — generated once from the actual model, not illustrative
1. the model — Bradley-Terry point-margin ratings
›the regression setup
One dummy regressor per D-I team (+1 home, −1 away, 0 otherwise), regressed against real point differential, plus a neutral-site indicator — the classic Bradley-Terry / Massey setup. Fit each season's regular season, predict that season's tournament games, score against the real final margin. Three knobs on top of the base regression: point-capping extreme blowouts before fitting, time-decay weighting so recent games count more, and shrinkage of each prediction toward the league-average margin.
2. tuning each knob — real backtested MAE, not assumed values
point cap — chosen 38
time decay (r) — chosen 0.990
shrinkage — chosen 0.084
Each knob grid-searched one at a time (holding the other two at their already-chosen values), real MAE on 1,250 real postseason games across 10 tournaments. Untuned baseline: 9.0552 pts. Fully tuned: 9.0277 pts — a modest, real ~0.3-point improvement, not a dramatic one.
3. real 2026 Final Four predictions
Semifinal 1
Semifinal 2
Championship (if UConn/Arizona)
Championship (if UConn/Michigan)
Championship (if Illinois/Arizona)
Championship (if Illinois/Michigan)
Real output of the tuned model applied to the actual 2025–26 season (through Mar 19) and a real bracket template — not a worked example. Both semifinals are on a neutral court, so "team1/team2" is just the template's slot order, not a home team.
honest limitations
Backtest MAE (~9 points) is on the final margin of games that were often already close by tip-off in the tournament — a naive "always predict the season-average margin" baseline wouldn't be dramatically worse, so this model's real edge over a trivial baseline is modest, not large. No injury or lineup-change information, and the tuning grid was searched sequentially (cap, then decay, then shrinkage) rather than jointly, so it may not be the true joint optimum.