Skip to main content
Hardwood

Record

How right has it been

Two records, kept apart on purpose: the live one was published before the games it scores, and everything after it is a reconstruction.

The live record

Published in advance

Nothing yet. Not a single game has been played since forecasts started being stamped, so the live record is empty — which is the honest state, not a missing feature. It will start at one game and be reported at whatever number it reaches.

This is the only measurement on this page that will be able to claim the numbers existed before the results did. It is kept separate from the backtest below permanently, not until it looks respectable.

Reconstructed

Everything below is a backtest

Twenty-three seasons scored under a walk-forward that never lets the model see the game it predicts — the right way to measure, and still not a published record. Why a backtest is never a live record

Against the closing line

ForecasterBrierLog lossAccuracyECEGap
Market (closing line)0.20700.600767.63%0.0066
This model0.21410.616665.75%0.0095+0.0071
Elo only0.21650.622265.41%0.0377+0.0095
Constant base rate0.24440.681957.51%0.0057+0.0374

14,600 games carry a price and are scored paired; 11,149 do not and are excluded rather than compared against nothing. De-vig: shin.

Paired bootstrap: 0.00712, 95% CI [0.00573, 0.00849] — the market is better and the interval excludes zero. That is the expected and wanted result: this model carries no market features, so beating the close would mean a bug in the harness, not an edge.

Does it mean what it says

perfect calibration0%0%25%25%50%50%75%75%100%100%Said 8.1%, happened 13.3% — 15 gamesSaid 16.7%, happened 14.1% — 297 gamesSaid 25.8%, happened 23.5% — 1,231 gamesSaid 35.4%, happened 33.4% — 2,623 gamesSaid 45.3%, happened 45.0% — 4,017 gamesSaid 55.2%, happened 55.2% — 5,461 gamesSaid 64.9%, happened 65.4% — 5,548 gamesSaid 74.6%, happened 77.4% — 4,189 gamesSaid 84.1%, happened 86.0% — 2,076 gamesSaid 92.0%, happened 93.2% — 292 gameswhat the model saidwhat happened
Dot area is the number of games in the bucket. A dot above the dashed line means the model was too cautious; below it, too confident.
View as a table
SaidHappenedGamesGap
8.1%13.3%155.3%
16.7%14.1%297-2.6%
25.8%23.5%1,231-2.3%
35.4%33.4%2,623-2.1%
45.3%45.0%4,017-0.3%
55.2%55.2%5,4610.0%
64.9%65.4%5,5480.5%
74.6%77.4%4,1892.8%
84.1%86.0%2,0762.0%
92.0%93.2%2921.2%

A forecaster that says 70% and is right 70% of the time is telling the truth whatever the schedule looks like. Measured error here is 0.0114.

On every game, priced or not

ForecasterBrierAccuracyECE
This model0.210666.45%0.0114
Elo only0.213766.08%0.0436
Constant base rate0.243558.08%0.0000

25,749 games. Home teams won 58.1% of them, which is the number the constant baseline predicts every time.

The numbers beside the probability

ForecastMAERMSEBiasGames
Margin, this model10.0412.820.4325,749
Margin, the spread9.8812.680.1414,600
Total, this model14.8618.87-1.5325,749
Total, the posted line14.4518.31-0.348,218

Points, on the paired subset where a line was published. The market is ahead on margin by 0.306 and on total by 0.760, which is the same story the Brier table tells and for the same reason.

Is the spread on those numbers right

Interval coverage, margin
StatedRealisedGap
50%49.0%-0.010
80%78.1%-0.019
95%93.0%-0.020
0.0–0.1: 11.5% observed against 10.0% uniform (2,953 games)0.1–0.2: 10.5% observed against 10.0% uniform (2,694 games)0.2–0.3: 9.9% observed against 10.0% uniform (2,543 games)0.3–0.4: 9.8% observed against 10.0% uniform (2,524 games)0.4–0.5: 9.5% observed against 10.0% uniform (2,437 games)0.5–0.6: 10.0% observed against 10.0% uniform (2,578 games)0.6–0.7: 10.0% observed against 10.0% uniform (2,577 games)0.7–0.8: 9.4% observed against 10.0% uniform (2,425 games)0.8–0.9: 9.1% observed against 10.0% uniform (2,343 games)0.9–1.0: 10.4% observed against 10.0% uniform (2,675 games)uniform0model median1
Where the real margin fell inside the model’s own published distribution, in deciles. Flat is correct. Heavy at both ends would mean the intervals are too narrow — and because the win probability is read off this same distribution, that would make every percentage on this site overconfident.
Interval coverage, total
StatedRealisedGap
50%52.0%+0.020
80%81.5%+0.015
95%95.6%+0.006
0.0–0.1: 7.5% observed against 10.0% uniform (1,931 games)0.1–0.2: 9.5% observed against 10.0% uniform (2,435 games)0.2–0.3: 10.3% observed against 10.0% uniform (2,663 games)0.3–0.4: 10.3% observed against 10.0% uniform (2,645 games)0.4–0.5: 10.5% observed against 10.0% uniform (2,701 games)0.5–0.6: 10.4% observed against 10.0% uniform (2,683 games)0.6–0.7: 10.6% observed against 10.0% uniform (2,719 games)0.7–0.8: 10.0% observed against 10.0% uniform (2,570 games)0.8–0.9: 10.0% observed against 10.0% uniform (2,574 games)0.9–1.0: 11.0% observed against 10.0% uniform (2,828 games)uniform0model median1
Where the real total fell inside the model’s own published distribution, in deciles. Flat is correct. Heavy at both ends would mean the intervals are too narrow — and because the win probability is read off this same distribution, that would make every percentage on this site overconfident.

Every percentage on the site reads off this same fitted normal, so a spread that runs narrow makes all of them overconfident at once. Why one normal drives everything

As measured: the margin intervals run narrow — its 95% band caught 93.0%, and the total intervals run wide — its 50% band caught 52.0%. Both misses are real and directional yet too small to bend the reliability curve above, and the margin tails run exactly where the distribution’s excess kurtosis says they should. Why a normal, and where it bends

Season by season

This modelClosing linelower is better — better sits higher
0.1940.2030.2130.2220.231071013161922252013 closing line 0.20292014 closing line 0.20672015 closing line 0.19962016 closing line 0.19902017 closing line 0.21172018 closing line 0.20232019 closing line 0.20852020 closing line 0.21022021 closing line 0.21572022 closing line 0.20962023 closing line 0.21812026 closing line 0.19912007 model 0.21872008 model 0.19972009 model 0.19802010 model 0.20132011 model 0.20322012 model 0.20752013 model 0.20472014 model 0.20872015 model 0.20642016 model 0.19952017 model 0.21762018 model 0.21272019 model 0.21922020 model 0.22012021 model 0.22552022 model 0.22022023 model 0.22712024 model 0.21132025 model 0.21072026 model 0.2091modelmarket
Seasons with no market line have no blue point — those years carried no published price in the source, and an absent benchmark is shown as absent rather than interpolated.
2025261,3220.20690.1991+0.0100
2024251,3210.2107
2023241,3190.2113
2022231,3200.22710.2181+0.0090
2021221,3230.22020.2096+0.0107
2020211,1710.22550.2157+0.0097
2019201,1430.22010.2102+0.0099
2018191,3120.21300.2085+0.0106
2017181,3120.21360.2023+0.0105
2016171,3090.21740.2117+0.0059
2015161,3160.19950.1990+0.0005
2014151,3110.20600.1996+0.0067
2013141,3190.20940.2067+0.0020
2012131,3140.20530.2029+0.0018
2011121,0740.2075
2010111,3110.2032
2009101,3120.2013
2008091,3150.1980
2007081,3160.1997
2006071,3090.2187

A season with no market column had no published lines in the source. It is shown as absent rather than filled in.

Playoff series

ForecasterBrierAccuracyECE
Coin flip0.250050.00%0.2200
Higher seed advances0.201672.00%0.0000
Series model0.199070.33%0.0749

300 series since 2007; the higher seed advanced 72.0% of the time. Series reconstruction passes its progression check at 100.0% — every winner the resolver names does appear in the next round.

The series model does not significantly beat “the higher seed advances”. Paired bootstrap: -0.00261, 95% CI [-0.01807, 0.01357] — the interval straddles zero, so this layer has not yet earned a claim. It ships only because the bracket simulation consumes its probabilities.

benchmark generated Aug 24, 2026, 1:42 PM UTC

Forecasts are scored against the closing line. Nothing here is betting advice. How it works