Verified historical performance

ScoreProb Probability Model Performance

See how ScoreProb's published football probabilities performed against real match outcomes. This page measures calibration, probability quality, pick accuracy and confidence using only certified pre-kickoff records.

Latest certified pre-kickoff probability per fixture One fixture counted once 3,741 Settled matches
Settled matches 3,741 5,240 Certified predictions
Sample period 2026-08-24 → 2026-10-05
Settlement coverage 79.2% 4,723 matches
1X2 Top-pick accuracy 51.7%
O/U 2.5 Accuracy 61.7%
BTTS Accuracy 57.8%

Probability quality

A probability model should be compared with simple rules, not judged by accuracy alone.

1X2 51.7%
Brier score 0.597 Lower is better
Log loss 0.999 Lower is better
Calibration error 1.80% Closer to zero is better
Naive baseline improvement 10.5% Brier improvement
O/U 2.5 61.7%
Brier score 0.229 Lower is better
Log loss 0.649 Lower is better
Calibration error 2.60% Closer to zero is better
Naive baseline improvement 8.5% Brier improvement
BTTS 57.8%
Brier score 0.242 Lower is better
Log loss 0.678 Lower is better
Calibration error 1.79% Closer to zero is better
Naive baseline improvement 3.0% Brier improvement

Probability calibration

When ScoreProb says 60%, an ideally calibrated model should see that event occur about 60% of the time over a large sample.

Probability band n Predicted Observed

Market detail

How each published probability behaves as a probability, not only as a winning pick.

Home 52.7%
Sample 2,556 Mean predicted 44.8% Observed frequency 45.6% Brier 0.227 ECE 1.76%
Draw 27.0%
Sample 37 Mean predicted 23.3% Observed frequency 20.4% Brier 0.163 ECE 3.13%
Away 50.3%
Sample 1,148 Mean predicted 31.9% Observed frequency 33.9% Brier 0.207 ECE 2.39%
Over 2.5 61.7%
Sample 3,741 Mean predicted 57.0% Observed frequency 59.4% Brier 0.229 ECE 2.60%
BTTS 57.8%
Sample 3,741 Mean predicted 54.0% Observed frequency 55.2% Brier 0.242 ECE 1.79%

Pick accuracy by outcome

When Home, Draw or Away was the highest 1X2 probability, how often was that top-ranked outcome correct?

Home 52.7%
Top pick n=2,556
Draw 27.0%
Top pick n=37
Away 50.3%
Top pick n=1,148

Probability strength

Do stronger top-ranked 1X2 probabilities actually win more often?

33–39% n=592 Mean confidence 38.0% Actual accuracy 43.6%
40–49% n=1,515 Mean confidence 44.7% Actual accuracy 45.1%
50–59% n=980 Mean confidence 54.4% Actual accuracy 55.8%
60–69% n=457 Mean confidence 64.2% Actual accuracy 62.1%
70–79% n=149 Mean confidence 73.8% Actual accuracy 79.9%
80%+ n=48 Mean confidence 85.2% Actual accuracy 89.6%

Confidence diagnostics

1X2 margin groups the gap between the highest and second-highest probability. Uncertainty measures how evenly the three 1X2 probabilities are distributed.

1X2 margin performance

Very tight n=308
43.2%
Tight n=435
42.1%
Moderate n=462
40.9%
Clear edge n=717
46.4%
Strong edge n=1819
60.3%

1X2 uncertainty performance

Very low n=26
88.5%
Low n=104
81.7%
Moderate n=456
64.5%
High n=1274
54.9%
Very high n=1881
44.3%

Rolling 30-day performance

Each point uses the certified settled fixtures available in the preceding 30 calendar days.

1X2 O/U 2.5 BTTS

Baseline comparison

A probability model should be compared with simple rules, not judged by accuracy alone.

1X2

Model
51.7%
Majority-class baseline
45.6%
Brier score 0.597 Naive probability baseline 0.667 Brier improvement 10.5%

O/U 2.5

Model
61.7%
Majority-class baseline
59.4%
Brier score 0.229 Naive probability baseline 0.250 Brier improvement 8.5%

BTTS

Model
57.8%
Majority-class baseline
55.2%
Brier score 0.242 Naive probability baseline 0.250 Brier improvement 3.0%

Outcome distribution

The underlying result mix gives context to every accuracy number.

1X2
Home 45.6% Draw 20.4% Away 33.9%
O/U 2.5
Over 59.4% Under 40.6%
BTTS
Yes 55.2% No 44.8%

League performance

Only leagues meeting the minimum certified settled sample are shown. Minimum sample: 20.

League n 1X2 Home Draw Away O2.5 BTTS Brier ECE
FA Trophy
57 57.9% 62.1% 0.0% 55.6% 71.9% 57.9% 0.579 10.27%
Taça de Portugal
36 75.0% 76.2% 0.0% 78.6% 52.8% 61.1% 0.461 14.09%
Primera Nacional
32 46.9% 50.0% 0.0% 33.3% 62.5% 62.5% 0.643 5.89%
Non League Premier - Northern
30 43.3% 42.1% — 45.5% 56.7% 50.0% 0.654 5.61%
Emperor Cup
26 84.6% 87.0% — 66.7% 53.8% 61.5% 0.388 18.52%
Liga Profesional Argentina
25 48.0% 42.9% — 75.0% 56.0% 44.0% 0.611 6.15%
League Two
24 41.7% 52.9% — 14.3% 62.5% 41.7% 0.679 8.99%
National League - South
22 54.5% 50.0% — 62.5% 45.5% 40.9% 0.608 10.11%
Non League Premier - Southern South
22 54.5% 58.8% — 40.0% 68.2% 72.7% 0.585 6.75%
Cup
21 61.9% 20.0% — 100.0% 76.2% 33.3% 0.547 15.47%
Derde Divisie
21 47.6% 50.0% — 33.3% 71.4% 71.4% 0.590 8.18%
Non League Premier - Southern Central
21 61.9% 60.0% — 66.7% 57.1% 57.1% 0.558 9.75%
Challenge Cup
20 55.0% 52.9% — 66.7% 70.0% 55.0% 0.517 15.05%
First Amateur Division
20 45.0% 54.5% — 33.3% 55.0% 50.0% 0.625 8.49%
Friendlies
20 45.0% 46.2% — 42.9% 60.0% 45.0% 0.627 6.31%
National League - North
20 55.0% 50.0% — 75.0% 45.0% 65.0% 0.615 9.85%
Non League Premier - Isthmian
20 40.0% 45.5% — 33.3% 55.0% 60.0% 0.639 18.00%

Country performance

No country table yet
Certified country metadata is not yet sufficiently populated for a meaningful country-level table. ScoreProb will publish this section once coverage is adequate.

How the metrics are calculated

Top-pick accuracy

Pick accuracy asks whether the highest-probability 1X2 outcome actually occurred.

Brier score

Brier score measures squared probability error. It rewards probabilities that are both accurate and appropriately confident.

Log loss

Log loss penalizes probabilities that assign very little probability to the outcome that actually occurs.

Calibration error

Calibration error measures the weighted gap between predicted probability and observed frequency across probability bands.

Minimum sample

Only leagues meeting the minimum certified settled sample are shown. n=20.

Probability calibration

When ScoreProb says 60%, an ideally calibrated model should see that event occur about 60% of the time over a large sample.

Integrity & publication rules

Every fixture is counted once using its latest certified prediction generated before kickoff. Results are never used to rewrite historical probabilities.
Verified ROI is not published yet
ScoreProb will not manufacture ROI from its own probabilities. Verified ROI requires immutable bookmaker odds captured before kickoff and bound to the certified prediction.