如何基于numpy、pandas在poly_fit中计算R²与调整R²?
Hey there! Let's fix up your poly_fit_r2 function to correctly calculate R² (coefficient of determination) and adjusted R² using NumPy and Pandas. I'll walk through the issues in your current code, then share a polished, working version with explanations.
First, let's spot the problems in your modified code:
- You reference
df['poly']but never definedf(you're usingdf_polyfor your DataFrame) - You call
np.polyfit()twice unnecessarily (this wastes computation) - You don't calculate adjusted R², which is critical for comparing models of different orders (like linear vs. quadratic)
- Error handling doesn't properly track metrics when fitting fails
Here's the corrected function:
import numpy as np import pandas as pd # Optional: Use scikit-learn for R², or use the manual calculation below from sklearn.metrics import r2_score def poly_fit_r2(id_col, x, y, z, p, n): """ Fits a polynomial to sample data, generates predictions for full dataset, and calculates R²/adjusted R². Args: id_col: ID values to match with original table (used as index) x: Sample input data for fitting y: Sample target data for fitting z: Full input dataset to generate predictions for p: Fallback values to use if fitting fails n: Order of the polynomial (1 = linear, 2 = quadratic, etc.) Returns: df_poly: DataFrame with ID index and fitted values, plus metrics as columns results: Dictionary with polynomial coefficients, R², and adjusted R² """ # Initialize DataFrame with ID as index (cleaner than your original empty setup) df_poly = pd.DataFrame({'id': id_col}).set_index('id') results = {} try: # Fit polynomial ONCE (avoid redundant calls) coeffs = np.polyfit(x, y, n) poly_func = np.poly1d(coeffs) # Generate fitted values for the full dataset df_poly['new_fit'] = poly_func(z) # Store polynomial coefficients results['polynomial'] = coeffs.tolist() # Calculate R² (two options below - pick one) # Option 1: Use scikit-learn's built-in function (simpler) r2 = r2_score(y, poly_func(x)) # Option 2: Manual calculation (no scikit-learn dependency) # y_pred = poly_func(x) # ss_residual = np.sum((y - y_pred) ** 2) # ss_total = np.sum((y - np.mean(y)) ** 2) # r2 = 1 - (ss_residual / ss_total) # Calculate adjusted R² (penalizes higher-order models for complexity) num_samples = len(y) adjusted_r2 = 1 - (1 - r2) * (num_samples - 1) / (num_samples - n - 1) # Store metrics in results dictionary results['r2'] = round(r2, 4) results['adjusted_r2'] = round(adjusted_r2, 4) # Optional: Add metrics as columns to the DataFrame (all rows get the same value) df_poly['r2'] = results['r2'] df_poly['adjusted_r2'] = results['adjusted_r2'] except Exception as e: # Fallback to original data if fitting fails df_poly['new_fit'] = p # Use your fallback prices, not z (matches your original logic) results['polynomial'] = None results['r2'] = 0.0 results['adjusted_r2'] = 0.0 df_poly['r2'] = 0.0 df_poly['adjusted_r2'] = 0.0 print(f"Fitting failed: {str(e)}") return df_poly, results
Key improvements explained:
- Efficient fitting: Only calls
np.polyfit()once instead of twice, saving computation time. - Adjusted R² calculation: This is essential for comparing linear (n=1) vs. quadratic (n=2) models. Unlike raw R², adjusted R² accounts for the number of parameters in the model, so it doesn't automatically favor higher-order polynomials.
- Cleaner DataFrame setup: Removes the useless
df_poly.emptyline and initializes the DataFrame in one step. - Dual return values: Gives you both the fitted data (for your original table) and a metrics dictionary (for model comparison).
- Flexible R² calculation: Choose between using scikit-learn's convenient function or a manual calculation if you want to avoid external dependencies.
Example usage:
# Test data id_list = [1, 2, 3, 4, 5] x_sample = np.array([1, 2, 3, 4, 5]) y_sample = np.array([2, 4, 5, 4, 2]) # Quadratic pattern z_full = np.array([1, 2, 3, 4, 5, 6, 7]) fallback_prices = np.array([2, 4, 5, 4, 2, 1, 0]) # Test linear regression (n=1) linear_fit, linear_metrics = poly_fit_r2(id_list, x_sample, y_sample, z_full, fallback_prices, 1) print("Linear Regression Metrics:") print(linear_metrics) # Test quadratic regression (n=2) quad_fit, quad_metrics = poly_fit_r2(id_list, x_sample, y_sample, z_full, fallback_prices, 2) print("\nQuadratic Regression Metrics:") print(quad_metrics)
How to choose between linear and quadratic models:
Compare the adjusted R² values. The model with the higher adjusted R² is the better fit (since it balances goodness-of-fit and model complexity). For the test data above, the quadratic model will have a much higher adjusted R², which makes sense given the quadratic pattern in y_sample.
内容的提问来源于stack exchange,提问作者HeadOverFeet
相关产品推荐
相关产品推荐

