Overfitting in Sports Models

Summary

Overfitting in sports prediction models occurs when the model learns the noise and specific details of the training data to such an extent that it negatively impacts performance on new, unseen data. In sports betting, overfitting is particularly dangerous because: (a) historical data has low signal-to-noise ratio (goals are inherently random), (b) sample sizes are small (64 World Cup matches per tournament), and (c) the market is adversarial — any identified pattern will be arbitraged away.

A model that perfectly fits historical World Cup data but fails to generalize to future tournaments is useless. Overfitting is the primary reason most sports betting models don't work in practice.

Common overfitting patterns in sports betting:
- Too many features relative to training samples
- Fitting to specific tournament conditions that don't repeat
- Using look-ahead bias (future information in training)
- Excessive model complexity (deep trees, many hidden layers)
- Hyperparameter tuning on the test set

Key Concepts

  • Signal-to-noise ratio: Sports outcomes, especially football, have high noise (random goals). A model needs to be simple enough to extract signal without fitting noise.
  • Bias-variance tradeoff: Complex models have low bias but high variance (overfit); simple models have high bias but low variance. For sports betting, lean toward simplicity.
  • Regularization: L1/L2 penalties on model parameters reduce overfitting. In tree models, min_samples_leaf, max_depth serve this role.
  • Feature selection: Using too many features with small samples causes overfitting. The ratio of features to samples matters critically.
  • In-sample vs. out-of-sample: In-sample performance (training set) will always look better than out-of-sample. The gap is a measure of overfitting.
  • Minimum viable complexity: The simplest model that captures the signal is almost always better than a complex one.

Overfitting Patterns in Sports Betting Models

  1. Too many ELO/K-factors: Fitting individual K-factors per team from limited data
  2. Dixon-Coles rho overfitting: Estimating the correlation parameter from small samples
  3. xG model with too many features: Using 50+ shot features when 5 would suffice
  4. Rolling window too short: Adapting too quickly to recent form in volatile tournaments
  5. Cross-validation on temporal data: Random k-fold splitting introduces look-ahead bias
  6. Hyperparameter tuning on test set: Selecting the model that performs best on the test period — this is overfitting to the test set

Prevention Strategies

def detect_overfitting(train_results, test_results, n_params, n_train_samples):
    """
    Detect potential overfitting using metrics.

    Returns:
        dict with indicators and recommendations
    """
    train_metric = train_results['brier_score']
    test_metric = test_results['brier_score']

    gap = test_metric - train_metric  # positive = overfitting

    # Simple heuristic: gap > 0.05 is concerning
    # Corrected for small samples: gap > 0.10 / log(n_params) is concerning

    ratio = n_train_samples / n_params
    expected_gap = 2 * n_params / n_train_samples  # rough approximation

    return {
        'train_metric': train_metric,
        'test_metric': test_metric,
        'gap': gap,
        'sample_param_ratio': ratio,
        'expected_gap': expected_gap,
        'is_overfitting': gap > expected_gap,
        'recommendation': 'simplify model' if gap > expected_gap else 'model OK'
    }

def feature_importance_stability(model, X, y, n_bootstrap=100, threshold=0.05):
    """
    Check feature importance stability via bootstrapping.
    Features with highly variable importance across bootstrap samples are likely overfit.
    """
    importances = []
    for _ in range(n_bootstrap):
        idx = np.random.choice(len(X), len(X), replace=True)
        X_boot, y_boot = X[idx], y[idx]
        model.fit(X_boot, y_boot)
        importances.append(model.feature_importances_)

    importances = np.array(importances)
    cv = importances.std(axis=0) / (importances.mean(axis=0) + 1e-10)

    unstable_features = np.where(cv > threshold)[0]

    return {
        'mean_importance': importances.mean(axis=0),
        'importance_cv': cv,
        'unstable_features': unstable_features
    }

def regularized_poisson_regression(X, y, alpha=1.0):
    """
    L1/L2 regularized Poisson regression for expected goals modeling.
    Penalizes large coefficients to prevent overfitting.
    """
    from sklearn.linear_model import Ridge
    # For Poisson, log-link: log(lambda) = X @ beta
    # Approximate by applying log-link to target then using Ridge
    log_y = np.log(y + 0.1)
    model = Ridge(alpha=alpha)
    model.fit(X, log_y)
    return model

Model Complexity Guidelines

Sample Size Max Features (roughly) Recommended Model
< 500 matches 3-5 Simple Poisson, Elo
500-2000 5-15 Regularized regression, shallow trees
2000+ 15-30 Gradient boosting, neural nets

For World Cup modeling: ~2000 international matches available for training. Use 5-10 features maximum.

Notes

  • The client's walk-forward validation across 4 World Cups is specifically designed to detect overfitting — if performance degrades in later tournaments, the model is overfitting to historical patterns
  • The rule of thumb: number of free parameters should be < N/20 where N is training sample size. For World Cup with 64 matches per tournament and 4 tournaments: N ~ 2000 training samples → < 100 parameters. Poisson with attack/defense for 50 teams = ~100 parameters (right at the limit).
  • Simplicity wins in sports betting: a well-calibrated simple Elo model often outperforms a complex ML model in out-of-sample testing