Calibration

When we say 60%, does it happen 60% of the time?

That is the only question this page answers — and the answer includes the places where we’re wrong. Calibration is the one property a probability model owes you. It is not an edge, and nothing on this page claims one.

Calibration error by bucket

Each bar is actual − predicted for one probability bucket in one season. Zero is perfect. Below zero means the model said it would happen more often than it did — it was over-confident.

expected calibration error

mean of the four seasons, as the engine grades. Lower is closer to meaning what it says.

mean ECE 0.0583

one dot ≈ 12 settled outcomes · click a bucket to isolate it

2022 ECE 0.04922023 ECE 0.05922024 ECE 0.05072025 ECE 0.0743

This is the engine as it grades. Every bucket sits below the zero line: the raw model says an outcome happens more often than it does, worst in 60–70%. Switch the correction on to see what it fixes.

Where we’re miscalibrated

  • 60–70%The engine overshoots by ~10pp in every season measured (−9.3, −10.2, −10.9, −10.9). This is structural, not noise. When the raw model says 65%, closer to 55% happens.
  • 2025Over-confident across the whole range, including buckets that were near-calibrated in prior years (20–50% all ~10pp over). We think a large rookie / new-team cohort is the likely cause; we haven’t proven it.
  • 50–60%Near-calibrated in all four seasons (worst −4.5pp). This is the band we’d trust most.

The curves in the research browser are the engine’s raw probabilities. Read them with the bias above in mind — especially in the 60–70% band.

Can it be corrected? Yes — and we tested that too

Under purged cross-validation (never fit and evaluate the same season, never train a future season to judge a past one), standard recalibration fixes the miscalibration cleanly:

methodheld-out runECE before → afterBrier before → after
Plattfit 2023 → 20240.10100.02460.25800.2461
Plattfit 2023 → 20250.12530.02490.26180.2438
Plattfit 23+24 → 20250.12530.02330.26180.2440
Isotonicfit 2023 → 20240.10100.01860.25800.2473
Isotonicfit 2023 → 20250.12530.01470.26180.2427
Isotonicfit 23+24 → 20250.12530.01860.26180.2426

Expected calibration error drops roughly (0.10–0.13 → 0.015–0.025) out-of-sample, with Brier improving too — so the fix generalizes rather than memorizing.

And here is the honest part. Recalibrating before grading destroyed the one subset that had looked promising in our market testing — which told us that subset was an artifact of over-confidence, not a real signal. Better probabilities, no better outcomes. So the engine still grades and selects on its raw numbers, and every backtest figure we publish is the uncorrected engine.

What we do ship is the correction as a display figure: in the research browser every curve carries a calibrated companion beside the raw one, so you can read both. It is the same calibrator you just toggled above. It changes what you are shown; it changes nothing about what the engine grades.

The walk-forward verdict

Four seasons of walk-forward backtesting (2022–2025) plus a three-season real-market test against actual sportsbook closing lines. The summary:

  • signalReal. Brier resolution above random in every season measured — the probabilities carry information.
  • edgeNone that survives contact with the closing line. Every apparent edge resolved into favorite price-drift on inspection, and every ROI confidence interval spanned zero.
  • verdictMaintain as research. No commercial picks product. We published that conclusion about our own model rather than burying it.

The full case study — every phase, including the tests we failed — is public in the engine repository. See also the methodology and charter.

this page updates as the 2026 season settles. the 2026 season is a pre-registered forward test: the picks logic and selection rules were fixed in advance and are not re-cut as results arrive. whatever it shows — including a worse calibration record — gets published here.