§1.2 — Trading
Kalshi Weather Model
Skew-t fair-value model for NYC temperature contracts, driving the live Kalshi execution stack.
Kalshi's KXHIGHNY markets pay $1 if NYC's daily high temperature lands in a given 1°F bucket (or above/below a threshold) — the model below estimates the true probability of each bucket to find the mispriced ones.
v1 → v4: −5.3% NLL end-to-end, each step LR-tested and walk-forward validated
see it trading → PnL Dashboard
the academic version of this question → Senior Honors Thesis
1. why skew-t, not Normal
Real fitted densities for one actual day (2022-12-27) — the skew-t curve is visibly shifted and asymmetric with a fatter left tail, exactly the shape a symmetric Normal distribution structurally can't represent.
2. model trial log — how v4 actually got here
›§1.3.1Trial 1 — Normal baseline2.4235 NLL/obs
NBS's own reported forecast spread isn't well-calibrated — it's persistently biased and more volatile than it admits, especially at longer lead times.
Correct the reported spread with two fitted parameters: a lead-time-dependent inflation and a fixed bias shift, keeping the residual distribution Normal.
›§1.3.2Trial 2 — skew-t2.3783 NLL/obs (−0.045 vs. Trial 1)
Forecast residuals aren't just mis-scaled Normal noise — they show real skew and fatter tails than a symmetric bell curve can represent.
MLE fit (Nelder-Mead) of a skew-t distribution — location, scale, skew, and tail-heaviness each a function of lead time. 8 parameters.
›§1.3.3Trial 3 — skew-t + features2.2953 NLL/obs (−0.044 vs. Trial 2)
Lead time alone doesn't capture everything relevant to forecast error — recent forecast drift, recent realized bias, seasonal skew, and the day's climatological variance should matter too.
Same skew-t MLE framework, extended to 11 parameters across forecast-drift, recent-bias, climatological-variance, and day-of-year terms.
›§1.3.4v4 — production (live)LR stat 81 (need 3.84)
A partial-dependence diagnostic on the Trial-3 fit showed clim_var's relationship to predicted outcome wasn't flat — real signal was still being left on the table.
Added a quadratic clim_var term, refit via MLE, confirmed with a likelihood-ratio test against Trial 3 on the same window, then re-confirmed on an annual walk-forward CV (not just one split). This is what fv_live.py trades on today.
› further exploration — not all of it adopted
- v5 (doy_cos² added to the skew term): the same PDP process flagged a sharp, non-linear turn in doy_cos right around New Year. The quadratic extension gave a marginal, inconclusive OOS gain on walk-forward CV — not adopted.
- v6 (forecast_drift_std² added to scale): written to address a leftover flat/rise/dip/rise pattern the existing drift term couldn't express. Not yet fit or walk-forward tested — an open thread, not a result.
- Parameter audit, not just parameter addition: periodically re-tested whether an old parameter (df, tail-heaviness, live since v2) was still pulling its weight. LR stat 1.19 (not significant) — dropping it would cost ~0 NLL, confirmed across 3 full annual folds. Verified, but not yet adopted — would need a full re-fit and downstream re-run before it's safe to call done.
3. partial-dependence diagnostics — what justified each step
recent_bias
clim_var
doy_cos
forecast_drift_std
spread_f
Partial dependence of a one-shot diagnostic classifier’s predicted P(payout) on each variable, others held at their real values. Flat = the FV model already explains it; sloped = signal was still on the table. This is the actual diagnostic that flagged clim_var and doy_cos (Trial 4 / v5 above).
4. model vs. market, one real day
real candidate set, September 26, 2026
5. what happens next — turning that into a trade
›Kill switch
flag-file halt, checked before every order — manual or auto-tripped on repeated data-refresh failure
›Live FV
fv_live.py computes fresh fair value per bucket from live prices + the production v4 model
›Best candidate
argmax positive-edge bucket for the day — one trade, not several correlated bets on the same underlying
›Kelly size
0.10× Kelly fraction, capped at 2% of live balance — the backtested sizing cell
›Idempotency check
DB unique constraint on client_order_id — an exact duplicate submission is blocked, a genuinely different decision isn't
›Submit
marketable IOC limit order — fills now at the checked price or not at all, never rests unmonitored
›Reconcile
on a network-ambiguous failure (timeout, no response), searches Kalshi's own order history by client_order_id instead of guessing