Record
How right has it been
Two records, kept apart on purpose: the live one was published before the games it scores, and everything after it is a reconstruction.
The live record
Published in advanceNothing yet. Not a single game has been played since forecasts started being stamped, so the live record is empty — which is the honest state, not a missing feature. It will start at one game and be reported at whatever number it reaches.
This is the only measurement on this page that will be able to claim the numbers existed before the results did. It is kept separate from the backtest below permanently, not until it looks respectable.
Reconstructed
Everything below is a backtest
Twenty-three seasons scored under a walk-forward that never lets the model see the game it predicts — the right way to measure, and still not a published record. Why a backtest is never a live record
Against the closing line
| Forecaster | Brier | Log loss | Accuracy | ECE | Gap |
|---|---|---|---|---|---|
| Market (closing line) | 0.2070 | 0.6007 | 67.63% | 0.0066 | — |
| This model | 0.2141 | 0.6166 | 65.75% | 0.0095 | +0.0071 |
| Elo only | 0.2165 | 0.6222 | 65.41% | 0.0377 | +0.0095 |
| Constant base rate | 0.2444 | 0.6819 | 57.51% | 0.0057 | +0.0374 |
14,600 games carry a price and are scored paired; 11,149 do not and are excluded rather than compared against nothing. De-vig: shin.
Paired bootstrap: 0.00712, 95% CI [0.00573, 0.00849] — the market is better and the interval excludes zero. That is the expected and wanted result: this model carries no market features, so beating the close would mean a bug in the harness, not an edge.
Does it mean what it says
View as a table
| Said | Happened | Games | Gap |
|---|---|---|---|
| 8.1% | 13.3% | 15 | 5.3% |
| 16.7% | 14.1% | 297 | -2.6% |
| 25.8% | 23.5% | 1,231 | -2.3% |
| 35.4% | 33.4% | 2,623 | -2.1% |
| 45.3% | 45.0% | 4,017 | -0.3% |
| 55.2% | 55.2% | 5,461 | 0.0% |
| 64.9% | 65.4% | 5,548 | 0.5% |
| 74.6% | 77.4% | 4,189 | 2.8% |
| 84.1% | 86.0% | 2,076 | 2.0% |
| 92.0% | 93.2% | 292 | 1.2% |
A forecaster that says 70% and is right 70% of the time is telling the truth whatever the schedule looks like. Measured error here is 0.0114.
On every game, priced or not
| Forecaster | Brier | Accuracy | ECE |
|---|---|---|---|
| This model | 0.2106 | 66.45% | 0.0114 |
| Elo only | 0.2137 | 66.08% | 0.0436 |
| Constant base rate | 0.2435 | 58.08% | 0.0000 |
25,749 games. Home teams won 58.1% of them, which is the number the constant baseline predicts every time.
The numbers beside the probability
| Forecast | MAE | RMSE | Bias | Games |
|---|---|---|---|---|
| Margin, this model | 10.04 | 12.82 | 0.43 | 25,749 |
| Margin, the spread | 9.88 | 12.68 | 0.14 | 14,600 |
| Total, this model | 14.86 | 18.87 | -1.53 | 25,749 |
| Total, the posted line | 14.45 | 18.31 | -0.34 | 8,218 |
Points, on the paired subset where a line was published. The market is ahead on margin by 0.306 and on total by 0.760, which is the same story the Brier table tells and for the same reason.
Is the spread on those numbers right
| Stated | Realised | Gap |
|---|---|---|
| 50% | 49.0% | -0.010 |
| 80% | 78.1% | -0.019 |
| 95% | 93.0% | -0.020 |
| Stated | Realised | Gap |
|---|---|---|
| 50% | 52.0% | +0.020 |
| 80% | 81.5% | +0.015 |
| 95% | 95.6% | +0.006 |
Every percentage on the site reads off this same fitted normal, so a spread that runs narrow makes all of them overconfident at once. Why one normal drives everything
As measured: the margin intervals run narrow — its 95% band caught 93.0%, and the total intervals run wide — its 50% band caught 52.0%. Both misses are real and directional yet too small to bend the reliability curve above, and the margin tails run exactly where the distribution’s excess kurtosis says they should. Why a normal, and where it bends
Season by season
| 2025–26 | 1,322 | 0.2069 | 0.1991 | +0.0100 |
| 2024–25 | 1,321 | 0.2107 | — | — |
| 2023–24 | 1,319 | 0.2113 | — | — |
| 2022–23 | 1,320 | 0.2271 | 0.2181 | +0.0090 |
| 2021–22 | 1,323 | 0.2202 | 0.2096 | +0.0107 |
| 2020–21 | 1,171 | 0.2255 | 0.2157 | +0.0097 |
| 2019–20 | 1,143 | 0.2201 | 0.2102 | +0.0099 |
| 2018–19 | 1,312 | 0.2130 | 0.2085 | +0.0106 |
| 2017–18 | 1,312 | 0.2136 | 0.2023 | +0.0105 |
| 2016–17 | 1,309 | 0.2174 | 0.2117 | +0.0059 |
| 2015–16 | 1,316 | 0.1995 | 0.1990 | +0.0005 |
| 2014–15 | 1,311 | 0.2060 | 0.1996 | +0.0067 |
| 2013–14 | 1,319 | 0.2094 | 0.2067 | +0.0020 |
| 2012–13 | 1,314 | 0.2053 | 0.2029 | +0.0018 |
| 2011–12 | 1,074 | 0.2075 | — | — |
| 2010–11 | 1,311 | 0.2032 | — | — |
| 2009–10 | 1,312 | 0.2013 | — | — |
| 2008–09 | 1,315 | 0.1980 | — | — |
| 2007–08 | 1,316 | 0.1997 | — | — |
| 2006–07 | 1,309 | 0.2187 | — | — |
A season with no market column had no published lines in the source. It is shown as absent rather than filled in.
Playoff series
| Forecaster | Brier | Accuracy | ECE |
|---|---|---|---|
| Coin flip | 0.2500 | 50.00% | 0.2200 |
| Higher seed advances | 0.2016 | 72.00% | 0.0000 |
| Series model | 0.1990 | 70.33% | 0.0749 |
300 series since 2007; the higher seed advanced 72.0% of the time. Series reconstruction passes its progression check at 100.0% — every winner the resolver names does appear in the next round.
The series model does not significantly beat “the higher seed advances”. Paired bootstrap: -0.00261, 95% CI [-0.01807, 0.01357] — the interval straddles zero, so this layer has not yet earned a claim. It ships only because the bracket simulation consumes its probabilities.
benchmark generated Aug 24, 2026, 1:42 PM UTC