Model v5 · Updated 23 Sep 2026
Our track record
Every forecast checked against the final 90-minute result, by league, competition and season. Nothing is hand-picked: every graded match counts.
The backtest runs model v5 over past seasons from Jun 2008 to Sep 2026. For every match it only uses information that existed before kickoff, exactly as if the forecast had been published that day. The model was tuned on matches from January 2025 to April 2026; every other month is an out-of-sample test.
Summary
- Graded matches
- 1,429,771
- Match result hit rate
- 51.8%
- Brier score, 1X2
- 0.591
- Over/Under 2.5 hit rate
- 59.7%
- Both teams to score hit rate
- 55.5%
- Exact score hit rate
- 11.2%
Model v5
Always picking the home side: 44.4%
Lower is better. Guessing scores 0.667.
1,429,771 matches
1,429,771 matches
Out of 121 possible scores
Every market
Results by market
| Market | Matches | Hit rate | Brier score | Log loss |
|---|---|---|---|---|
| Match result (1X2) | 1,429,771 | 51.8% | 0.5908 | 0.9906 |
| Over/Under 1.5 goals | 1,429,771 | 75.0% | 0.1805 | 0.5433 |
| Over/Under 2.5 goals | 1,429,771 | 59.7% | 0.2355 | 0.6633 |
| Over/Under 3.5 goals | 1,429,771 | 69.4% | 0.2028 | 0.5931 |
| Both teams to score | 1,429,771 | 55.5% | 0.2462 | 0.6856 |
| Exact score | 1,428,262 | 11.2% | - | 2.9836 |
Hit rate counts the outcome we marked as most likely. Brier score and log loss grade the full probabilities, so a confident miss costs more than a cautious one.
Season by season
Match result hit rate over time
Show as a table
| Period | Matches | Hit rate | Brier score | Over/Under 2.5 | Both teams to score |
|---|---|---|---|---|---|
| 2026 | 113,336 | 52.6% | 0.5833 | 61.0% | 56.6% |
| 2025 | 164,645 | 52.4% | 0.5867 | 60.5% | 56.2% |
| 2024 | 169,799 | 52.5% | 0.5856 | 60.7% | 55.9% |
| 2023 | 163,709 | 52.3% | 0.5868 | 60.4% | 55.8% |
| 2022 | 154,166 | 52.4% | 0.5863 | 60.2% | 55.6% |
| 2021 | 135,074 | 51.9% | 0.5893 | 59.5% | 54.9% |
| 2020 | 84,631 | 51.3% | 0.5949 | 58.9% | 54.7% |
| 2019 | 95,593 | 50.9% | 0.5971 | 59.0% | 55.4% |
| 2018 | 69,137 | 51.1% | 0.5961 | 58.5% | 55.1% |
| 2017 | 63,863 | 50.9% | 0.5985 | 59.1% | 55.3% |
| 2016 | 52,988 | 50.4% | 0.6019 | 58.2% | 54.9% |
| 2015 | 38,950 | 51.3% | 0.5950 | 58.6% | 55.0% |
| 2014 | 36,329 | 50.7% | 0.6001 | 58.5% | 55.4% |
| 2013 | 34,695 | 50.5% | 0.5993 | 58.4% | 55.4% |
| 2012 | 34,728 | 50.1% | 0.6020 | 58.1% | 55.0% |
| 2011 | 15,796 | 49.4% | 0.6092 | 56.0% | 52.8% |
| 2010 | 2,287 | 46.4% | 0.6279 | 54.4% | 50.9% |
| 2009 | 15 | 60.0% | 0.6145 | 40.0% | 46.7% |
| 2008 | 30 | 46.7% | 0.6255 | 33.3% | 40.0% |
Paler bars have fewer than 100 matches and move a lot by chance. The earliest seasons have few matches and little history for the model to learn from, so they read lower.
Do the percentages mean what they say?
Calibration
Each dot groups home, draw and away probabilities of similar size. Dots close to the diagonal mean that when we say 60%, it happens about 60% of the time.
Where it works best
By competition group
| Group | Matches | Hit rate | Brier score |
|---|---|---|---|
| Major leagues | 150,938 | 49.5% | 0.607 |
| Other men's leagues | 1,019,020 | 51.0% | 0.597 |
| Cups | 109,572 | 55.1% | 0.566 |
| Women's football | 35,577 | 62.6% | 0.494 |
| Youth & reserves | 52,892 | 53.6% | 0.583 |
| National teams | 17,077 | 60.2% | 0.516 |
| Friendlies | 44,695 | 54.7% | 0.577 |
Groups follow the period and model filters. Mismatched competitions such as cups and women's leagues are easier to call than balanced top divisions.
1064 leagues with at least 30 matches
By league
| League | Matches | Hit rate | Brier score | Over/Under 2.5 | Both teams to score |
|---|---|---|---|---|---|
|
|
39,552 | 55.0% | 0.576 | 62.6% | 56.5% |
|
|
10,696 | 48.2% | 0.619 | 62.4% | 57.6% |
|
|
8,450 | 45.7% | 0.635 | 53.5% | 52.5% |
|
|
8,318 | 44.3% | 0.647 | 53.7% | 51.9% |
|
|
8,285 | 46.9% | 0.631 | 52.5% | 52.5% |
|
|
8,200 | 47.1% | 0.625 | 54.0% | 54.7% |
|
|
7,387 | 52.0% | 0.586 | 55.7% | 51.0% |
|
|
7,368 | 43.6% | 0.642 | 66.6% | 58.8% |
|
|
7,300 | 48.7% | 0.614 | 62.0% | 56.9% |
|
|
6,408 | 44.7% | 0.641 | 55.3% | 52.2% |
|
|
6,304 | 49.6% | 0.614 | 56.2% | 52.1% |
|
|
6,202 | 50.0% | 0.617 | 56.8% | 56.9% |
|
|
6,149 | 52.7% | 0.582 | 57.2% | 52.3% |
|
|
6,131 | 53.4% | 0.584 | 54.6% | 53.8% |
|
|
6,120 | 53.1% | 0.583 | 55.1% | 54.6% |
|
|
5,900 | 43.5% | 0.645 | 56.8% | 52.7% |
|
|
5,813 | 49.7% | 0.606 | 57.1% | 52.8% |
|
|
5,761 | 45.9% | 0.639 | 53.7% | 54.2% |
|
|
5,713 | 56.6% | 0.558 | 69.7% | 61.5% |
|
|
5,687 | 41.7% | 0.652 | 63.7% | 57.2% |
|
|
5,610 | 47.0% | 0.625 | 58.9% | 53.0% |
|
|
5,511 | 48.2% | 0.614 | 58.0% | 53.8% |
|
|
5,470 | 40.1% | 0.655 | 65.6% | 57.6% |
|
|
5,298 | 48.3% | 0.618 | 55.8% | 55.0% |
|
|
5,278 | 43.0% | 0.644 | 59.1% | 53.4% |
|
|
5,159 | 59.2% | 0.537 | 68.1% | 59.9% |
|
|
5,110 | 62.8% | 0.520 | 56.2% | 54.1% |
|
|
5,023 | 53.9% | 0.572 | 60.0% | 57.6% |
|
|
5,001 | 50.8% | 0.602 | 53.9% | 53.0% |
|
|
4,985 | 47.3% | 0.630 | 62.8% | 59.9% |
|
|
4,972 | 47.5% | 0.614 | 53.8% | 51.9% |
|
|
4,948 | 51.4% | 0.597 | 58.7% | 56.7% |
|
|
4,915 | 56.1% | 0.572 | 63.3% | 58.0% |
|
|
4,818 | 51.2% | 0.598 | 55.1% | 54.2% |
|
|
4,793 | 47.0% | 0.630 | 55.5% | 55.5% |
|
|
4,784 | 44.0% | 0.639 | 61.7% | 56.2% |
|
|
4,743 | 44.7% | 0.638 | 60.4% | 53.3% |
|
|
4,705 | 54.5% | 0.560 | 56.0% | 52.8% |
|
|
4,654 | 44.7% | 0.645 | 55.6% | 55.4% |
|
|
4,606 | 41.8% | 0.652 | 56.8% | 51.3% |
|
|
4,601 | 47.4% | 0.624 | 55.0% | 51.3% |
|
|
4,582 | 47.4% | 0.623 | 55.4% | 54.8% |
|
|
4,558 | 47.4% | 0.625 | 54.7% | 52.3% |
|
|
4,454 | 50.2% | 0.621 | 54.0% | 52.8% |
|
|
4,446 | 47.2% | 0.621 | 59.2% | 52.9% |
|
|
4,344 | 48.3% | 0.615 | 59.0% | 53.8% |
|
|
4,340 | 48.4% | 0.607 | 57.2% | 52.2% |
|
|
4,334 | 64.0% | 0.479 | 72.5% | 55.5% |
|
|
4,322 | 46.6% | 0.625 | 55.3% | 52.7% |
|
|
4,236 | 46.3% | 0.614 | 60.3% | 52.8% |
With a few hundred matches, a league's hit rate can move by several points from luck alone. Compare leagues with large samples.
Reading the numbers
How we measure accuracy
- Hit rate
- How often the outcome we marked as most likely happened. Simple, but it ignores how confident the forecast was.
- Brier score
- The squared gap between our probabilities and what happened. 0 is perfect; for match results, spreading a third on each outcome scores 0.667.
- Log loss
- Punishes confident mistakes harder than the Brier score. Lower is better.
- Backtest
- The current model run over past matches with only the information available before each kickoff. It shows how the model behaves across many seasons, but it is a simulation, not a record of published forecasts.
- Published forecasts
- Forecasts we actually showed, stored with a timestamp before kickoff and never edited afterwards.
- Result used
- The score after 90 minutes plus stoppage time. Extra time and penalties do not count.
Forecasts are probabilities, not certainties. Past accuracy does not guarantee future results.