Brier Score

Summary

The Brier score is a strictly proper scoring rule for measuring the accuracy of probabilistic predictions. For binary outcomes, it is the mean squared error between predicted probabilities and actual outcomes. It was proposed by Glenn W. Brier in 1950.

A strictly proper scoring rule means that the score is optimized only when the predicted probability exactly matches the true probability. If a forecaster reports a probability of 60% for an event that occurs 60% of the time, they get the best possible Brier score. Deviating from the true probability always makes the score worse.

The Brier score ranges from 0 (perfect) to 1 (worst). For binary events: BS = (predicted_prob - actual_outcome)² averaged over all predictions.

The Brier score decomposes into calibration + refinement. Calibration measures how close predicted probabilities are to actual frequencies; refinement measures how much the predictions differentiate between outcomes.

Key Concepts

  • Strictly proper: The score is maximized only when the predicted probability equals the true probability. No hedging or hedging away from confidence can improve it.
  • Binary events: Standard Brier score is for binary (0/1) outcomes. For multi-category predictions, the original Brier definition applies.
  • Calibration component: Measures how well-calibrated the predictions are. If you predict 70% for 100 events, about 70 should actually occur.
  • Refinement component: Measures how sharp/distinct the predictions are. A model that predicts 50% for everything has zero refinement.
  • Brier Skill Score (BSS): Normalized version where 0 = no skill vs. baseline (climatology), 1 = perfect. BSS = 1 - BS/BS_baseline.
  • Limitations: Brier score is less discriminating for very rare or very common events — large samples needed for reliable estimates.

Formulas

Brier score (binary):
$$BS = \frac{1}{N} \sum_{i=1}^{N} (p_i - o_i)^2$$

Where p_i = predicted probability for event i, o_i = actual outcome (1 if event occurred, 0 if not).

Decomposition:
$$BS = \underbrace{\frac{1}{N} \sum_{k} n_k (p_k - \bar{o}k)^2}{\text{reliability (calibration)}} - \underbrace{\frac{1}{N} \sum_{k} n_k (\bar{o}k - \bar{o})^2}{\text{resolution}} + \underbrace{\bar{o}(1 - \bar{o})}_{\text{uncertainty}}$$

Where n_k = number of predictions with probability p_k, ō_k = observed frequency for those predictions, ō = overall observed frequency.

Brier Skill Score:
$$BSS = \frac{BS_{\text{model}} - BS_{\text{climatology}}}{0 - BS_{\text{climatology}}} = 1 - \frac{BS_{\text{model}}}{BS_{\text{climatology}}}$$

Python Implementation

import numpy as np

def brier_score(predicted_probs, actual_outcomes):
    """
    Calculate Brier score for binary predictions.

    Args:
        predicted_probs: array of predicted probabilities (0-1)
        actual_outcomes: array of actual outcomes (0 or 1)

    Returns:
        Brier score (lower is better, 0 = perfect)
    """
    return np.mean((predicted_probs - actual_outcomes) ** 2)

def brier_score_decomposition(predicted_probs, actual_outcomes, n_bins=10):
    """
    Decompose Brier score into calibration and refinement.

    Returns:
        dict with reliability, resolution, uncertainty, and BS
    """
    df = np.array([predicted_probs, actual_outcomes]).T
    df = df[df[:, 0].argsort()]

    bin_size = len(df) // n_bins
    reliability_components = []
    resolution_components = []

    overall_mean = np.mean(actual_outcomes)

    for i in range(n_bins):
        bin_data = df[i*bin_size:(i+1)*bin_size] if i < n_bins-1 else df[i*bin_size:]
        if len(bin_data) == 0:
            continue

        p_bar = np.mean(bin_data[:, 0])
        o_bar = np.mean(bin_data[:, 1])
        n = len(bin_data)

        reliability_components.append(n * (p_bar - o_bar) ** 2)
        resolution_components.append(n * (o_bar - overall_mean) ** 2)

    reliability = np.sum(reliability_components) / len(df)
    resolution = np.sum(resolution_components) / len(df)
    uncertainty = overall_mean * (1 - overall_mean)

    bs = reliability - resolution + uncertainty

    return {
        'brier_score': bs,
        'reliability': reliability,  # lower is better
        'resolution': resolution,      # higher is better
        'uncertainty': uncertainty
    }

def brier_skill_score(predicted_probs, actual_outcomes, baseline_probs=None):
    """
    Calculate Brier Skill Score vs. climatology baseline.
    """
    bs_model = brier_score(predicted_probs, actual_outcomes)

    if baseline_probs is None:
        # Use climatology (overall mean as constant prediction)
        baseline_probs = np.full_like(actual_outcomes, np.mean(actual_outcomes), dtype=float)

    bs_baseline = brier_score(baseline_probs, actual_outcomes)
    bss = 1 - (bs_model / bs_baseline) if bs_baseline > 0 else 0

    return bss

# Example usage
predictions = np.array([0.60, 0.70, 0.30, 0.80, 0.50])
outcomes = np.array([1, 1, 0, 1, 0])
bs = brier_score(predictions, outcomes)
decomp = brier_score_decomposition(predictions, outcomes)
bss = brier_skill_score(predictions, outcomes)

print(f"Brier Score: {bs:.4f}")  # (0.6-1)² + (0.7-1)² + (0.3-0)² + (0.8-1)² + (0.5-0)² / 5
print(f"Decomposition: {decomp}")
print(f"Brier Skill Score: {bss:.4f}")

Notes

  • Brier score is the primary calibration metric for sports betting models — lower is better
  • A Brier score of 0.20 for soccer win/draw/lose predictions is reasonable; 0.15 is good; 0.10 is excellent
  • The decomposition is useful for diagnosis: high reliability = poor calibration (model is overconfident); high resolution = good differentiation ability
  • For the World Cup model: Brier score should be computed separately for each tournament (2010, 2014, 2018, 2022) and averaged for overall assessment
  • Brier score is less sensitive to overconfident predictions on high-probability events than log-loss, making it more appropriate for sports betting where many events have 60-80% implied probabilities