ElasticNet模型拟合效果差(负R²),求模型优化建议
问题:ElasticNet模型拟合效果差,出现负R²值排查与优化
模型评估结果
- Best lambda (alpha): 1.0
- Best l1_ratio: 0.2
- R-squared on testing data: -0.00499349856926945
- RMSE on testing data: 0.8576623398551885
代码
import os import pandas as pd from sklearn.linear_model import ElasticNet from sklearn.model_selection import train_test_split, GridSearchCV, cross_validate from sklearn.metrics import mean_squared_error, r2_score import numpy as np # Set the working directory to the desktop desktop_path = os.path.expanduser("~/Desktop") os.chdir(desktop_path) # Load data from CSV file data = pd.read_csv('stan_func_conn.csv') # Split data into input features (X) and target variable (y) X = data.iloc[:, 1:] # Independent variables (excluding the first column) y = data.iloc[:, 0] # Dependent variable (the first column) # Split data into training and testing sets X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # Define the parameter grid for tuning param_grid = { 'alpha': [0.1, 0.5, 1.0], # Adjust these values as per your requirement 'l1_ratio': [0.2, 0.5, 0.8] # Adjust these values as per your requirement } # Create an instance of the ElasticNet model elastic_net = ElasticNet(max_iter=10000) # Perform grid search for hyperparameter tuning and model evaluation grid_search = GridSearchCV(estimator=elastic_net, param_grid=param_grid, cv=5, scoring='neg_mean_squared_error') grid_search.fit(X_train, y_train) # Get the best lambda (alpha) and l1_ratio parameters best_alpha = grid_search.best_params_['alpha'] best_l1_ratio = grid_search.best_params_['l1_ratio'] print("Best lambda (alpha):", best_alpha) print("Best l1_ratio:", best_l1_ratio) # Create an instance of the ElasticNet model with the best lambda (alpha) and l1_ratio parameters elastic_net_tuned = ElasticNet(alpha=best_alpha, l1_ratio=best_l1_ratio, max_iter=10000) # Fit the tuned model to the training data elastic_net_tuned.fit(X_train, y_train) # Make predictions on the testing data y_pred = elastic_net_tuned.predict(X_test) # Compute the R-squared on the testing data r2 = r2_score(y_test, y_pred) print('R-squared on testing data:', r2) # Compute the RMSE on testing data mse = mean_squared_error(y_test, y_pred) rmse = np.sqrt(mse) print('RMSE on testing data:', rmse) # Compare RMSE to the scale of the target variable if rmse < 0.5: print('The model has a very good fit.') elif rmse < 1: print('The model has a good fit.') else: print('The model has a moderate to poor fit.')
排查与优化建议
代码逻辑排查
- 数据划分验证:打印
data.head()确认第一列确实是目标变量,其余列是特征,避免特征与标签搞反。 - 模型收敛性检查:训练后打印
elastic_net_tuned.n_iter_,如果结果等于设置的max_iter=10000,说明模型未收敛,需增大max_iter(如调到100000)或降低tol参数的收敛阈值。 - 网格搜索参数扩展:当前
alpha范围仅到1.0,建议添加0.01, 0.05, 2.0, 5.0;l1_ratio补充0.0, 1.0(对应纯Ridge和纯Lasso),覆盖更多正则化组合。
数据预处理优化
- 特征标准化:ElasticNet对特征尺度敏感,必须做标准化处理,示例代码:
from sklearn.preprocessing import StandardScaler scaler = StandardScaler() X_train_scaled = scaler.fit_transform(X_train) X_test_scaled = scaler.transform(X_test) # 后续用缩放后的特征训练模型 - 数据分布检查:用
sns.histplot(y)查看目标变量分布,若为非线性分布,可对y做对数、Box-Cox等变换后再训练。 - 特征筛选:若特征维度远大于样本量,计算特征与
y的相关性,剔除低相关特征;或用方差筛选去掉方差接近0的冗余特征。
模型评估与调整
- 对比训练集指标:计算训练集的R²和RMSE,若训练集指标也差,说明欠拟合,需减小
alpha降低正则化强度;若训练集好但测试集差,说明过拟合,需增大alpha。 - 分析交叉验证结果:查看
grid_search.cv_results_中不同参数的交叉验证得分,若最优参数在网格边界(如当前alpha为1.0是最大值),需继续扩展参数范围测试。 - 尝试其他模型:若线性模型不适合当前数据,可尝试随机森林、XGBoost等树模型,对比拟合效果。
内容的提问来源于stack exchange,提问作者Toiba
相关产品推荐
相关产品推荐

