Elo Rating System¶
Summary¶
The Elo rating system is a method for calculating the relative skill levels of players in zero-sum two-player games. It was invented by Hungarian-American chess master and physics professor Arpad Elo in 1959–1960, replacing the Harkness rating system. The system is named after Elo and is a special case of the Bradley–Terry model.
Elo's central assumption is that the performance of each player in each game is a normally distributed random variable. A player's true skill is the mean of that performance variable. After each game, the winner takes points from the loser — the difference determines how many points are transferred. Upset wins (lower-rated player beating higher-rated) result in larger point transfers.
The system is self-correcting: players whose ratings are too low will outperform their rating and gain points until ratings reflect true strength. Elo ratings are comparative only and valid only within the rating pool — there's no absolute measure of skill.
Elo has been adapted for many sports including association football, American football, baseball, basketball, tennis, and esports. FiveThirtyEight famously uses Elo variants for sports predictions.
Key Concepts¶
- Rating difference → win probability: A 100-point difference → 64% expected score for stronger player; 200-point difference → 76%
- K-factor: Controls how much ratings change after each game. Higher K = faster adjustment but more volatility. Chess uses K=32 for new players, K=10 for veterans. Sports adaptations vary.
- Expected score: E = 1 / (1 + 10^(rating_diff/400)) — logistic function
- Rating update: New_Rating = Old_Rating + K × (Actual_Score - Expected_Score)
- Normal vs. logistic distribution: Elo originally used normal distribution, but logistic is mathematically simpler and works equally well in practice
- Home advantage: Most sports adaptations add 60–100 points to home team's rating to account for home advantage
Formulas¶
Expected score for player A vs player B:
$$E_A = \frac{1}{1 + 10^{(R_B - R_A)/400}}$$
Rating update after game:
$$R_A' = R_A + K(S_A - E_A)$$
Where S_A = 1 (win), 0.5 (draw), 0 (loss); K is the development coefficient.
FIDE performance rating:
$$RP = R_D + 400 \times (W - 0.5)$$
Where W = points scored / games played, R_D = opponent rating average.
Python Implementation¶
import math
def expected_score(rating_a, rating_b):
"""Expected score (win probability) for player A vs player B."""
return 1 / (1 + 10 ** ((rating_b - rating_a) / 400))
def update_rating(rating, expected, actual, k=32):
"""Update Elo rating after a game."""
return rating + k * (actual - expected)
def predict_match(rating_home, rating_away, home_advantage=65):
"""Predict match outcome probabilities using Elo."""
eff_home = rating_home + home_advantage
e = expected_score(eff_home, rating_away)
# For soccer, convert to 1X2 via a distribution assumption
# Simple approach: use normal distribution around rating difference
diff = eff_home - rating_away
prob_home = 1 / (1 + 10 ** (-diff / 400))
# Draw probability from a simple model
prob_draw = 1 / (1 + 10 ** (abs(diff) / 400)) * 0.25 # rough approximation
return {"home_win": prob_home * (1 - prob_draw), "draw": prob_draw, "away_win": (1 - prob_home) * (1 - prob_draw)}
Notes¶
- Elo is a special case of Bradley–Terry with a specific scaling (400-point scale, logistic function)
- Key limitation: Elo doesn't capture uncertainty — a player with few games has the same weight as one with hundreds
- Glicko and Glicko-2 address this with rating deviation (RD) and volatility
- For sports betting, ELO's predictive accuracy is competitive but often needs combination with other models
- The K-factor is crucial: too high = noisy ratings; too low = slow to adapt to form changes
- FiveThirtyEight's soccer Elo uses K=20 and includes home advantage of ~55 rating points