PORTFOLIO UPDATED — SEPTEMBER 23, 2026

§1.2 — Trading

Kalshi Weather Model

Skew-t fair-value model for NYC temperature contracts, driving the live Kalshi execution stack.

Kalshi's KXHIGHNY markets pay $1 if NYC's daily high temperature lands in a given 1°F bucket (or above/below a threshold) — the model below estimates the true probability of each bucket to find the mispriced ones.

v1 → v4: −5.3% NLL end-to-end, each step LR-tested and walk-forward validated

see it trading → PnL Dashboard

the academic version of this question → Senior Honors Thesis

TradingResearch

1. why skew-t, not Normal

35°F forecast
Normal (Trial 1) skew-t (v4, live)

Real fitted densities for one actual day (2022-12-27) — the skew-t curve is visibly shifted and asymmetric with a fatter left tail, exactly the shape a symmetric Normal distribution structurally can't represent.

2. model trial log — how v4 actually got here

›§1.3.1Trial 1 — Normal baseline2.4235 NLL/obs

NBS's own reported forecast spread isn't well-calibrated — it's persistently biased and more volatile than it admits, especially at longer lead times.

Correct the reported spread with two fitted parameters: a lead-time-dependent inflation and a fixed bias shift, keeping the residual distribution Normal.

›§1.3.2Trial 2 — skew-t2.3783 NLL/obs (−0.045 vs. Trial 1)

Forecast residuals aren't just mis-scaled Normal noise — they show real skew and fatter tails than a symmetric bell curve can represent.

MLE fit (Nelder-Mead) of a skew-t distribution — location, scale, skew, and tail-heaviness each a function of lead time. 8 parameters.

›§1.3.3Trial 3 — skew-t + features2.2953 NLL/obs (−0.044 vs. Trial 2)

Lead time alone doesn't capture everything relevant to forecast error — recent forecast drift, recent realized bias, seasonal skew, and the day's climatological variance should matter too.

Same skew-t MLE framework, extended to 11 parameters across forecast-drift, recent-bias, climatological-variance, and day-of-year terms.

›§1.3.4v4 — production (live)LR stat 81 (need 3.84)

A partial-dependence diagnostic on the Trial-3 fit showed clim_var's relationship to predicted outcome wasn't flat — real signal was still being left on the table.

Added a quadratic clim_var term, refit via MLE, confirmed with a likelihood-ratio test against Trial 3 on the same window, then re-confirmed on an annual walk-forward CV (not just one split). This is what fv_live.py trades on today.

› further exploration — not all of it adopted
  • v5 (doy_cos² added to the skew term): the same PDP process flagged a sharp, non-linear turn in doy_cos right around New Year. The quadratic extension gave a marginal, inconclusive OOS gain on walk-forward CV — not adopted.
  • v6 (forecast_drift_std² added to scale): written to address a leftover flat/rise/dip/rise pattern the existing drift term couldn't express. Not yet fit or walk-forward tested — an open thread, not a result.
  • Parameter audit, not just parameter addition: periodically re-tested whether an old parameter (df, tail-heaviness, live since v2) was still pulling its weight. LR stat 1.19 (not significant) — dropping it would cost ~0 NLL, confirmed across 3 full annual folds. Verified, but not yet adopted — would need a full re-fit and downstream re-run before it's safe to call done.

3. partial-dependence diagnostics — what justified each step

0.1970.190-2.03.0

recent_bias

0.1960.19238.6105.6

clim_var

0.1990.191-1.01.0

doy_cos

0.1950.1900.01.5

forecast_drift_std

0.2110.1911.07.0

spread_f

Partial dependence of a one-shot diagnostic classifier’s predicted P(payout) on each variable, others held at their real values. Flat = the FV model already explains it; sloped = signal was still on the table. This is the actual diagnostic that flagged clim_var and doy_cos (Trial 4 / v5 above).

4. model vs. market, one real day

real candidate set, September 26, 2026

T69 *+20.1¢
B68.5+7.9¢
B66.5+2.9¢
B64.5-9.5¢
B62.5-19.6¢
T62-13.8¢
ask (market) fv (model)* traded — 122 contracts @ 1¢

5. what happens next — turning that into a trade

1
›Kill switch

flag-file halt, checked before every order — manual or auto-tripped on repeated data-refresh failure

2
›Live FV

fv_live.py computes fresh fair value per bucket from live prices + the production v4 model

3
›Best candidate

argmax positive-edge bucket for the day — one trade, not several correlated bets on the same underlying

4
›Kelly size

0.10× Kelly fraction, capped at 2% of live balance — the backtested sizing cell

5
›Idempotency check

DB unique constraint on client_order_id — an exact duplicate submission is blocked, a genuinely different decision isn't

6
›Submit

marketable IOC limit order — fills now at the checked price or not at all, never rests unmonitored

7
›Reconcile

on a network-ambiguous failure (timeout, no response), searches Kalshi's own order history by client_order_id instead of guessing

›data provenance
Caught a real settlement bug rather than trusting a data source blindly: GHCN's archived high for July 15, 2026 loaded as 81°F, but the NWS's own primary CLI report — what Kalshi actually settles on — already had the correct 95°F. Confirmed independently before changing anything. The live system now pulls the CLI report directly for the last ~10 days of settlement; the full 2021–2025 training window still runs on GHCN, tracked as an open audit item rather than assumed fine.
← back to Trading