🏐 Forecasting and statistical analysis tool for volleyball championships.
volleyball-simulator docs GLICKO2_IMPLEMENTATION.md
13 kB
Markdown
at main

Glicko-2 Implementation Summary #

Overview #

Successfully implemented proper Glicko-2 rating system for volleyball, achieving +1.60% improvement over traditional ELO with adaptive K-factor. This brings total cumulative improvement to +5.23% over the original baseline.

Key Achievement #

Proper implementation matters! After an initial failed attempt using a "Bayesian ELO" hybrid approach (-6.83%), reimplementing clean Glicko-2 on its own scale and terms achieved positive results (+1.60%).

Validation Results #

Progressive validation on 2024 Superlega season (132 matches, 23 rounds):

System Home Adv Log-loss vs Adaptive ELO Brier Score Accuracy
Traditional ELO (baseline) 250 0.5529 - 0.1881 68.94%
Glicko-2 (optimal) 0.5 0.5441 +1.60% ✅ 0.1842 72.73%
Glicko-2 (home=0.3) 0.3 0.5491 +0.68% 0.1821 73.48%
Glicko-2 (no home) 0.0 0.5958 -7.75% 0.1894 70.45%

Phase-by-Phase Analysis (Optimal Configuration) #

Phase Rounds Traditional Glicko-2 Difference
Early 1-7 0.6199 0.5782 +6.73% ✅
Mid 8-15 0.5191 0.5266 -1.45%
Late 16-23 0.5521 0.6056 -9.69% ❌

Critical trade-off: Glicko-2 excels early season (+6.73%) but significantly underperforms late season (-9.69%). Net improvement is +1.60%, but late-season predictions are crucial for playoff forecasting.

Rating Deviation (φ) Evolution #

Uncertainty properly decreases as teams play more matches:

Round Avg φ Interpretation Matches per Team
1 2.015 Very uncertain 0
7 0.987 Fairly certain ~6
15 0.686 Fairly certain ~13
23 0.559 Very certain ~22

This gradual decay (2.015 → 0.559) is appropriate for the season length.

Implementation Approach #

Key Decisions #

1. Proper Glicko-2 Scale (Not ELO Scale)

  • Initial rating: μ = 0.0 (mean)
  • Initial RD: φ = 2.014761 (high uncertainty)
  • Initial volatility: σ = 0.06
  • Home advantage: μ_home = 0.5 (optimal)
  • Tau constraint: τ = 0.5 (doesn't affect results significantly)

2. Match-as-One-Game Approach

Critical insight: Treat entire volleyball match as ONE Glicko-2 "game" with fractional score, not multiple games per set.

  • 3-1 match → score = 0.75 (not 3 wins + 1 loss)
  • 3-0 match → score = 1.0
  • 2-3 match → score = 0.4

This prevents φ from decreasing too rapidly (which caused -6.52% performance when sets were separate games).

3. Standard Glicko-2 Formulas

Used exact Glicko-2 paper formulas:

  • g(φ) = 1/sqrt(1 + 3φ²/π²)
  • E(μ, μⱼ, φⱼ) = 1 / (1 + exp(-g(φⱼ)(μ - μⱼ)))
  • Illinois algorithm for volatility updates
  • Steps 3-8 from Glickman's paper

Code Structure #

Created vbsimulator/glicko2.py (~400 lines):

class Glicko2RatingSystem:
    """Proper Glicko-2 for volleyball.

    Each team has (μ, φ, σ):
    - μ: Rating (higher is better)
    - φ: Rating deviation (uncertainty)
    - σ: Volatility (consistency)
    """

    def expected_score(rating_a, rating_b, phi_b, a_is_home):
        """Predict set win probability using Glicko-2."""
        mu_a_effective = rating_a + (home_advantage if a_is_home else 0)
        return E(mu_a_effective, rating_b, phi_b)

    def update_rating(team_a, team_b, sets_won_a, sets_won_b, ...):
        """Update ratings using Glicko-2 algorithm.

        Treats match as one game with fractional score.
        """
        score_a = sets_won_a / (sets_won_a + sets_won_b)
        # Apply Glicko-2 steps 3-8 with Illinois algorithm

Why Glicko-2 Works Better Than Previous Attempt #

Failed Attempt: "Bayesian ELO" Hybrid (-6.83%) #

❌ Tried to make Glicko-2 compatible with ELO scale ❌ Used ad-hoc formulas (σ/4 as K-factor) ❌ Treated multiple sets as multiple games ❌ Mixed ELO and Glicko-2 concepts

Result: Worse than traditional ELO

Successful Approach: Proper Glicko-2 (+1.60%) #

✅ Clean implementation on Glicko-2's own scale ✅ Used exact Glicko-2 formulas from paper ✅ Treated match as single game with fractional score ✅ No attempt to match ELO conventions

Result: Better than traditional ELO

Lesson: Use established systems as designed, don't create hybrids.

Trade-Offs and Considerations #

Advantages of Glicko-2 #

  1. Better Early Season Predictions (+6.73%)

    • High initial uncertainty (φ=2.015) allows rapid learning
    • First 7 rounds show significant improvement
    • Useful for early season betting/forecasting
  2. Explicit Uncertainty Quantification

    • Provides confidence intervals on ratings
    • Example: "Team A: μ=1.25 ± 1.1 (95% CI)"
    • Valuable for risk-aware analysis
  3. Handles New Teams Well

    • Promoted teams can start with higher φ
    • Learn quickly initially, stabilize over time
    • Better than fixed K in traditional ELO
  4. Theoretically Sound

    • Bayesian foundation
    • Well-tested in chess, Go, online gaming
    • Published, peer-reviewed algorithm

Disadvantages of Glicko-2 #

  1. Worse Late Season Predictions (-9.69%)

    • Critical issue: Late season is most important for playoffs
    • Adaptive K-factor in ELO seems to handle stabilization better
    • φ may decrease too much, making system too conservative
  2. More Complex

    • 3 parameters per team (μ, φ, σ) vs 1 for ELO
    • Illinois algorithm for volatility updates
    • Harder to understand and explain
  3. Parameter Sensitivity

    • Home advantage must be re-tuned (0.5 optimal, not 250)
    • τ doesn't seem to matter (0.3, 0.5, 0.8 all identical)
    • Different scale makes comparisons harder
  4. No Historical Compatibility

    • Can't directly compare Glicko-2 μ with ELO ratings
    • Must retrain from scratch
    • Breaks continuity with past seasons

Comparison to Previous Improvements #

Feature Log-loss Change Cumulative Status
Baseline (no improvements) 0.5741 - -
Point differential weighting -6.2% ❌ 0.6098 Rejected
Home court advantage (+250) +2.49% ✅ 0.5599 Implemented
Adaptive K-factor (27-37) +1.24% ✅ 0.5529 Implemented
"Bayesian ELO" hybrid -6.83% ❌ 0.5906 Rejected
Proper Glicko-2 +1.60% ✅ 0.5441 Optional

Combined improvement with all features: +5.23% over original baseline

Recommendation #

Option A: Keep Traditional ELO (Conservative) #

Pros:

  • Late season performance is better (-9.69% gap is significant)
  • Simpler to understand and maintain
  • Historical continuity
  • Already achieving +3.7% improvement

Cons:

  • Missing +1.60% overall improvement
  • Early season predictions weaker
  • No uncertainty quantification

Best for: Production systems where late-season accuracy is critical (playoff predictions, betting).

Option B: Adopt Glicko-2 (Progressive) #

Pros:

  • +1.60% overall improvement
  • Excellent early season (+6.73%)
  • Uncertainty quantification
  • Better for new/promoted teams

Cons:

  • Late season degradation (-9.69%)
  • More complex
  • Must retrain all historical data

Best for: Research, analysis, early-season forecasting, systems needing confidence intervals.

Option C: Hybrid Approach (Optimal?) #

Use Glicko-2 for rounds 1-15, switch to Traditional ELO for rounds 16-23:

  • Early/mid season (rounds 1-15): Glicko-2 excels (+6.73% and -1.45% = ~+2.6% avg)
  • Late season (rounds 16-23): Traditional ELO excels (avoid -9.69% loss)

Expected improvement: ~+3-4% overall (better than either alone)

Trade-off: Complexity, discontinuity at transition point

Implementation Files #

New files:

  1. vbsimulator/glicko2.py - Clean Glicko-2 implementation (~400 lines)
  2. validate_glicko2.py - Validation script with phase analysis (~350 lines)
  3. GLICKO2_IMPLEMENTATION.md - This documentation

Preserved files:

  • vbsimulator/elorating.py - Traditional ELO (unchanged)

Usage Example #

from vbsimulator.glicko2 import Glicko2RatingSystem
from vbsimulator.data_parser import load_matches_from_file

# Load data
fixture = load_matches_from_file("data/superlega2024_matches.dat")

# Create Glicko-2 system (all teams start at μ=0, φ=2.015)
glicko2 = Glicko2RatingSystem.create_from_fixture(fixture)

# Print ratings with uncertainty
glicko2.print_ratings()
# Output:
# Team                            μ (Rating)    φ (RD)       σ (Vol)
# ----------------------------------------------------------------------
# Perugia                         1.872         0.547        0.0599
# Trento                          1.654         0.539        0.0601

# Get prediction with confidence
mu_a, phi_a, sigma_a = glicko2.get_rating_full(team_a)
mu_b, phi_b, sigma_b = glicko2.get_rating_full(team_b)

p_set_win = glicko2.expected_score(mu_a, mu_b, phi_b, a_is_home=True)
print(f"Team A set win probability: {p_set_win:.2%}")
print(f"Team A uncertainty: φ = {phi_a:.3f}")

# For match probability, use same formula as traditional ELO
p_match = calculate_match_win_probability(p_set_win)

Future Work #

Immediate Improvements #

  1. Late Season Tuning

    • Try increasing φ decay in late season (slower uncertainty reduction)
    • Test higher τ values (though current tests suggest no effect)
    • Experiment with minimum φ floor (prevent too much certainty)
  2. Hybrid System

    • Implement Option C (Glicko-2 early, ELO late)
    • Find optimal transition point (round 15? 18?)
    • Smooth transition to avoid discontinuity
  3. Cross-League Testing

    • Validate on A2 Men and A1 Women data
    • Test with promoted teams (2024 had none)
    • Verify generalization

Advanced Enhancements #

  1. Per-Set Probability Model

    • Currently uses same set win probability for all sets
    • Could model Set 5 differently (psychological pressure)
    • Expected +0.5-1% improvement
  2. Ensemble Approach

    • Weighted average of Glicko-2 and Traditional ELO
    • Weight by phase: more Glicko-2 early, more ELO late
    • Could capture best of both worlds
  3. Glicko-2 Extensions

    • Player-level ratings (aggregate to team)
    • Time-based uncertainty inflation (off-season)
    • Match importance weighting (playoffs)

Lessons Learned #

1. Implementation Details Matter Critically #

Same algorithm (Glicko-2), different implementations:

  • Hybrid "Bayesian ELO": -6.83% (worse)
  • Proper Glicko-2: +1.60% (better)

Difference: 8.43 percentage points from implementation alone!

2. Don't Force Compatibility #

Trying to make Glicko-2 "compatible" with ELO scale made it worse. Using Glicko-2 on its own terms made it better.

3. Match-as-Game vs Sets-as-Games #

Treating 4-5 sets as 4-5 separate games caused φ to decay too fast (-6.52%). Treating match as one game with fractional score was key to success (+1.60%).

4. Trade-Offs Are Real #

Even successful improvements have costs. +6.73% early but -9.69% late is a genuine trade-off requiring careful consideration.

5. Validation Is Essential #

Phase-by-phase analysis revealed the late-season problem. Overall metrics (+1.60%) hide important details. Always validate by season phase.

Conclusion #

Proper Glicko-2 implementation achieves +1.60% improvement, bringing cumulative gains to +5.23% over baseline.

However, late-season performance degradation (-9.69%) is a significant concern for practical applications focused on playoff predictions.

Recommended approach:

  • Production: Keep Traditional ELO + home advantage + adaptive K (+3.7%)
  • Research: Use Glicko-2 for uncertainty quantification and early season analysis
  • Optimal: Implement hybrid system (Glicko-2 early, ELO late) for best of both

The proper Glicko-2 implementation validates that explicit uncertainty modeling can help when done correctly, but adaptive K-factor in traditional ELO already captures much of this benefit more simply and with better late-season performance.

References #

  • Glicko-2 Rating System: http://www.glicko.net/glicko/glicko2.pdf (Mark Glickman, 2012)
  • Home advantage implementation: HOME_ADVANTAGE_IMPLEMENTATION.md (+2.49%)
  • Adaptive K-factor implementation: ADAPTIVE_K_IMPLEMENTATION.md (+1.24%)
  • Failed Bayesian attempt: BAYESIAN_ELO_IMPLEMENTATION.md (-6.83%)
  • Traditional ELO validation: validate_adaptive_k.py
  • Glicko-2 validation: validate_glicko2.py

Bottom Line: Glicko-2 works (+1.60%) but has trade-offs. Implementation approach matters more than algorithm choice. Consider hybrid system for optimal results.