ScoreProb Probability Model Performance
See how ScoreProb's published football probabilities performed against real match outcomes. This page measures calibration, probability quality, pick accuracy and confidence using only certified pre-kickoff records.
Probability quality
A probability model should be compared with simple rules, not judged by accuracy alone.
Probability calibration
When ScoreProb says 60%, an ideally calibrated model should see that event occur about 60% of the time over a large sample.
Market detail
How each published probability behaves as a probability, not only as a winning pick.
Pick accuracy by outcome
When Home, Draw or Away was the highest 1X2 probability, how often was that top-ranked outcome correct?
Probability strength
Do stronger top-ranked 1X2 probabilities actually win more often?
Confidence diagnostics
1X2 margin groups the gap between the highest and second-highest probability. Uncertainty measures how evenly the three 1X2 probabilities are distributed.
1X2 margin performance
1X2 uncertainty performance
Rolling 30-day performance
Each point uses the certified settled fixtures available in the preceding 30 calendar days.
Baseline comparison
A probability model should be compared with simple rules, not judged by accuracy alone.
1X2
O/U 2.5
BTTS
Outcome distribution
The underlying result mix gives context to every accuracy number.
League performance
Only leagues meeting the minimum certified settled sample are shown. Minimum sample: 20.
Country performance
Certified country metadata is not yet sufficiently populated for a meaningful country-level table. ScoreProb will publish this section once coverage is adequate.
How the metrics are calculated
Pick accuracy asks whether the highest-probability 1X2 outcome actually occurred.
Brier score measures squared probability error. It rewards probabilities that are both accurate and appropriately confident.
Log loss penalizes probabilities that assign very little probability to the outcome that actually occurs.
Calibration error measures the weighted gap between predicted probability and observed frequency across probability bands.
Only leagues meeting the minimum certified settled sample are shown. n=20.
When ScoreProb says 60%, an ideally calibrated model should see that event occur about 60% of the time over a large sample.
Integrity & publication rules
ScoreProb will not manufacture ROI from its own probabilities. Verified ROI requires immutable bookmaker odds captured before kickoff and bound to the certified prediction.