基于Python优化球队评分:对数似然最小化实现求助
Hey there! Let's work through this problem step by step. I see you're trying to optimize your sports rating model by minimizing the sum of log-likelihoods (or equivalently maximizing the negative log-likelihood), but you're stuck on setting up the right loss function. No worries—let's break this down and fix it together.
Key Issues in Your Current Code
- You precomputed columns like
attack_handdefence_ausing your initial parameter guesses, but these values need to be recalculated dynamically for every set of parameters the optimizer tests. - You haven't defined a proper loss function that accepts parameter inputs and returns the value to minimize, which is required for
scipy.optimize.
Step-by-Step Solution
1. Define the Loss Function
This function will take the parameter vector, split it into individual ratings/advantage values, compute the log-likelihood for all games, and return the negative sum of log-likelihoods. We use the negative because scipy's optimizers minimize the loss function, so minimizing the negative log-likelihood is the same as maximizing the log-likelihood (our actual goal).
import numpy as np import pandas as pd import scipy.optimize import scipy.stats def neg_log_likelihood(params, game_data, id_list): # Split the flat parameter vector into individual components num_teams = len(id_list) attackratings = params[:num_teams] defenceratings = params[num_teams:2*num_teams] stdevratings = params[2*num_teams:3*num_teams] homeadv = params[-1] total_loglik = 0.0 # Iterate through each game to calculate log-likelihood for idx, row in game_data.iterrows(): # Find the index of home/away teams in our sorted ID list home_idx = id_list.index(row['HomeID']) away_idx = id_list.index(row['AwayID']) # Pull the current parameter values for this game attack_h = attackratings[home_idx] defence_a = defenceratings[away_idx] st_dev_h = stdevratings[home_idx] attack_a = attackratings[away_idx] defence_h = defenceratings[home_idx] st_dev_a = stdevratings[away_idx] # Calculate expected mean and standard deviation for scores home_mean = attack_h * defence_a * homeadv home_std = st_dev_h * st_dev_a # Adjust this if your model uses a different std formula away_mean = attack_a * defence_h away_std = st_dev_a * st_dev_h # Get actual observed scores home_pts = row['HomePts'] away_pts = row['AwayPts'] # Compute probability densities for the observed scores home_pdf = scipy.stats.norm.pdf(home_pts, loc=home_mean, scale=home_std) away_pdf = scipy.stats.norm.pdf(away_pts, loc=away_mean, scale=away_std) # Add log of joint probability to total (avoid log(0) errors) if home_pdf <= 0 or away_pdf <= 0: total_loglik += np.log(1e-10) # Use a tiny value instead of 0 else: total_loglik += np.log(home_pdf * away_pdf) # Return negative total log-likelihood for minimization return -total_loglik
2. Set Up Parameter Bounds
Some parameters can't be negative (like standard deviations). We'll define bounds to keep parameters valid during optimization:
# Load your game data first (same as your original code) game = pd.read_csv('Games.csv') id_list = sorted(pd.unique(pd.concat([game['HomeID'], game['AwayID']], axis=0))) num_teams = len(id_list) # Define bounds for each parameter bounds = [] # Attack ratings: positive values (adjust range based on your data) for _ in range(num_teams): bounds.append((0.1, 20)) # Defence ratings: same as attack for _ in range(num_teams): bounds.append((0.1, 20)) # Standard deviations: must be positive (0.1 to 10) for _ in range(num_teams): bounds.append((0.1, 10)) # Home advantage: positive (0.5 to 3, adjust as needed) bounds.append((0.5, 3))
3. Run the Optimization
Now use scipy.optimize.minimize to find the optimal parameters:
# Initial parameter guess (same as your original setup) init_params = tuple( [5]*num_teams + # Initial attack ratings [5]*num_teams + # Initial defence ratings [2]*num_teams + # Initial standard deviations [1.1] # Initial home advantage ) # Run the optimization result = scipy.optimize.minimize( fun=neg_log_likelihood, x0=init_params, args=(game, id_list), # Pass extra data to the loss function method='L-BFGS-B', # Good choice for bounded optimization bounds=bounds, options={'maxiter': 1000, 'disp': True} # Show progress updates ) # Check results and extract optimal parameters if result.success: print("\nOptimization converged successfully!") opt_params = result.x # Split optimized parameters into individual components opt_attack = opt_params[:num_teams] opt_defence = opt_params[num_teams:2*num_teams] opt_stdev = opt_params[2*num_teams:3*num_teams] opt_homeadv = opt_params[-1] # Create a dataframe to view team ratings clearly team_ratings = pd.DataFrame({ 'TeamID': id_list, 'AttackRating': opt_attack, 'DefenceRating': opt_defence, 'ScoreStdDev': opt_stdev }) print("\nOptimal Team Ratings:") print(team_ratings) print(f"\nOptimal Home Advantage: {opt_homeadv:.4f}") else: print(f"\nOptimization failed: {result.message}")
Important Notes
- Avoiding Log(0): We added a check for zero PDF values and replaced them with
1e-10to prevent errors fromnp.log(0). - Std Dev Calculation: Double-check if
st_dev_h * st_dev_ais the correct standard deviation formula for your model. If your model assumes independent scores, you might need a different calculation—adjust this if needed. - Performance: For large datasets, looping with
iterrows()can be slow. You could vectorize calculations using NumPy to speed things up, but the loop is easier to read for smaller datasets. - Optimizer Choice:
L-BFGS-Bis ideal here because it handles bounded parameters well. If you don't need bounds, you could useNelder-Meadinstead.
内容的提问来源于stack exchange,提问作者bobman

