Accuracy
A fixed MLB benchmark plus live, per-sport calibration updated after graded slates.
The 2025 MLB backtest reports 55.8% moneyline accuracy and +0.0261 Brier skill against its baseline. In the fixed April to June 2026 MLB study, the model's typical gap from the Pinnacle close was 4.2 points. The live panels show how each model sport grades as its sample grows.
These studies do not establish a broad market-beating edge. A value-play stream built on the models lost money, and we retired it.Signed-in members can track one narrow MLB dog result as a hypothesis with its limits stated.
Each row: what we predicted vs what actually happened. Matched bars = calibrated.
Sports publish a skill number once they cross ~500 graded games. Until then they say "gathering", not a placeholder stat.
Measured against the Pinnacle closing line
A betting model is only as good as its calibration. A calibrated model that says a team has a 55% chance wins about 55% of the time across many games. Picking winners at a high rate is not enough on its own, because a model can do that and still be miscalibrated and lose money. This page measures Fairline's calibration directly and shows the working, so you can judge the model instead of taking a claim on faith.
Pinnacle's closing line provides a sharp market benchmark. We pulled it for 855 MLB games the model had already projected (2026-04-12 to 2026-06-15), removed the bookmaker margin from both sides, and lined up three numbers per game: the model's win probability, Pinnacle's, and the result. The calibration card above renders that dataset; this benchmark is MLB only, the sport the study covered.
The model runs about two points more pessimistic on home teams than Pinnacle does (a -2.1 point signed gap).
Where the model disagrees with the sharp line
The fixed study compares model probabilities, Pinnacle probabilities, and results in one MLB window. It does not establish calibration in future samples. One pocket stood out in that study. Betting every underdog blindly was roughly breakeven against Pinnacle. Dogs the model rated at least 3 points higher than Pinnacle won 50.5%of the time against the 41.7% the closing line implied. That difference was +8.8 points over 291 bets in the 2026-04-12 to 2026-06-15 study window.
| Underdog group | n | Model | Pinnacle | Actual win | Beat close | ROI @ Pinn |
|---|---|---|---|---|---|---|
| All underdogs | 848 | 45.6% | 43.9% | 44.8% | +0.9 | +0.4% |
| Model-favored (edge ≥ +3%) | 291 | 48.5% | 41.7% | 50.5% | +8.8 | +18.3% |
| Model-neutral (|edge| < 3%) | 439 | 45.1% | 44.9% | 41.5% | -3.4 | -9.3% |
| Model-warned (edge ≤ -3%) | 118 | 40.3% | 46.0% | 43.2% | -2.8 | -7.3% |
The fixed study shows positive actual-minus-sharp results for both home dogs (+9.6) and away dogs (+8.5) that the model favors. The home-pessimism above does not explain the full result.
The ROI figures use Pinnacle's own price, which a US bettor cannot take. The study did not replay these games at takeable prices. The comparable study statistic is the +8.8 point actual-minus-sharp result. The 2026-04-12 to 2026-06-15 window prompted the study, so the pocket still has to prove itself forward. Signed-in members can follow that test as a dual record: a hypothesis cohort of every sharp-priced dog and an executable record of flagged plays, both entry-frozen. The live sharp feed grades them as each new game closes.
Live calibration, updated as bets grade
The study above is a fixed MLB benchmark over one window. This panel is the ongoing version, and it spans sports. Across the bets the model flags in each sport, it bins the stated probability against how often those bets win, and the claimed edge against the edge they realize. A claimed edge that shrinks once graded is the winner's curse, and showing it is the point. Pick a sport below; the numbers refresh as new games settle.
overall 6.7% claimed -> -1.8% realized (8.6pp cal err) (4124 graded bets)
PREDICTED VS ACTUAL, BY BUCKET
BRIER VS BASE RATE
Brier is the mean squared error of the probabilities (lower is better). The base rate is the score from always guessing the group's average win rate; a positive skill means the model's probabilities beat that baseline.
CLAIMED VS REALIZED EDGE BY BAND
| CLAIMED BAND | N | CLAIMED | REALIZED | SHRINK |
|---|---|---|---|---|
| 0–5% | 2007 | 3.8% | -1.6% | -0.41 |
| 5–10% | 1430 | 7.0% | -3.4% | -0.49 |
| 10–15% | 480 | 11.8% | -1.5% | -0.12 |
| 15–20% | 119 | 17.3% | 0.7% | 0.04 |
| 20–100% | 88 | 28.7% | 10.7% | 0.37 |
shrink = realized / claimed. A band whose realized edge falls short of its claimed edge at real volume is the winner's curse, and showing it is the point.
How to read model accuracy
Betting-model accuracy includes calibration and winner accuracy. Calibration asks whether groups labeled 55% win about 55% over enough games. Winner accuracy counts calls above or below 50%, which can hide probability errors. The charts here emphasize calibration and show how the estimates behave after grading.
Calibration
A model is calibrated when its stated probabilities match how often things actually happen. If it says 55% across a large group of bets, and that group wins close to 55% of the time, the model is calibrated. The same has to hold at every level: its 60% calls win about 60%, its 30% calls win about 30%, and so on up and down the range.
Each displayed probability describes a group frequency, not a guaranteed outcome. A calibrated 55% group still loses about 45% of its games. Calibration must be checked for the relevant sport, market, probability range, and sample.
One game is mostly variance
Flip a fair coin ten times and the result will often differ from five heads. Betting samples behave the same way. A well-calibrated 60% group still loses about 40% of its games, so four losses in ten can occur without contradicting the stated probability.
A single night or week carries little evidence about calibration. Use the sample counts and bucket gaps to judge whether observed win rates line up with stated probabilities.
Winning and losing weeks can occur under the same probability set. A weekly record cannot identify calibration by itself. Reliability and Brier measures summarize larger samples and preserve the probability information that a win-loss count discards.
For the longer version of why this happens, see Variance: Why a Good Model Still Loses for Weeks.
Reading the two charts on this page
Each chart tests a different property of the model probabilities.
- The reliability curve. The model's predicted probability runs along the bottom (x), and how often that group actually won runs up the side (y). The dashed diagonal is perfect calibration: every bucket landed where the model said it would. Points sitting above the diagonal mean the model was too pessimistic (those bets won more often than it claimed). Points sitting below the diagonal mean it was too confident (they won less often than claimed). The closer the points hug the line, the smaller the measured calibration gap. Read every gap with its sample count.
- Claimed edge versus realized edge (shrinkage). When the model flags a bet, it claims some edge over the price. Once those bets grade, we measure the edge they actually delivered. A claimed edge that shrinks once it is graded is the winner's curse: the selected flags are the observations where the model was most optimistic, and some of that optimism can be noise. The chart reports how much survives grading. A positive, stable realized edge at adequate volume would support the claim. A negative or shrinking result warns against relying on the pre-grade edge.
What the charts establish
- Predicted and actual rates that align support calibration in that measured sample.
- Brier skill shows whether the probability distribution improves on the comparison baseline.
- The model-to-sharp gap quantifies disagreement with Pinnacle in the fixed MLB window.
- It does not promise we beat the market broadly. A calibrated model can match the sharp price and find no edge most nights.
- It does not promise any single bet wins. Calibration is a statement about the long run, not the next slip.
- It does not promise a profit on a short sample. Variance can mask a real edge or fake one for weeks.
Closing line value compares an entry price with a closing reference. It measures price movement and selection timing, while profit and loss need a larger sample. Read CLV together with calibration and the source label on the closing reference.
How to use this
Use the model probability as a labeled independent reference. Check the relevant calibration panel before relying on a probability range. Compare takeable prices with the best available market-derived fair reference. Treat an edge as supported only after it survives grading at adequate volume, and size every position for normal variance.
How this is computed
- The model's probability is the projection frozen at lineup lock, before the close. It never reads book lines.
- The sharp probability is Pinnacle's moneyline at first pitch, de-vigged two-way so the bookmaker margin is removed from both sides.
- The benchmark study covers all 855 MLB games from 2026-04-12 to 2026-06-15, every one matched to a Pinnacle line. Every number on it is reproduced from a fixed dataset of those games, so it does not move.
- The live panel runs the same calibration per sport over the bets the model has flagged in the trailing year, recomputed hourly as games grade.
Limits to keep in mind: the dog-pocket edge is measured in the same window that prompted the study and over a short sample (two months, 291 bets in the favored pocket), so it is meaningful but not seasoned. The model probability is fixed pre-close while the sharp probability is the close, so the two are read at slightly different times. Results are not split by team, park, or starter. Pinnacle is a reference line only and is not a book a US bettor can wager at.