You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ElasticNet模型拟合效果差(负R²),求模型优化建议

问题:ElasticNet模型拟合效果差,出现负R²值排查与优化

模型评估结果

  • Best lambda (alpha): 1.0
  • Best l1_ratio: 0.2
  • R-squared on testing data: -0.00499349856926945
  • RMSE on testing data: 0.8576623398551885

代码

import os
import pandas as pd
from sklearn.linear_model import ElasticNet
from sklearn.model_selection import train_test_split, GridSearchCV, cross_validate
from sklearn.metrics import mean_squared_error, r2_score
import numpy as np

# Set the working directory to the desktop
desktop_path = os.path.expanduser("~/Desktop")
os.chdir(desktop_path)

# Load data from CSV file
data = pd.read_csv('stan_func_conn.csv')

# Split data into input features (X) and target variable (y)
X = data.iloc[:, 1:]  # Independent variables (excluding the first column)
y = data.iloc[:, 0]  # Dependent variable (the first column)

# Split data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Define the parameter grid for tuning
param_grid = {
    'alpha': [0.1, 0.5, 1.0],  # Adjust these values as per your requirement
    'l1_ratio': [0.2, 0.5, 0.8]  # Adjust these values as per your requirement
}

# Create an instance of the ElasticNet model
elastic_net = ElasticNet(max_iter=10000)

# Perform grid search for hyperparameter tuning and model evaluation
grid_search = GridSearchCV(estimator=elastic_net, param_grid=param_grid, cv=5, scoring='neg_mean_squared_error')
grid_search.fit(X_train, y_train)

# Get the best lambda (alpha) and l1_ratio parameters
best_alpha = grid_search.best_params_['alpha']
best_l1_ratio = grid_search.best_params_['l1_ratio']
print("Best lambda (alpha):", best_alpha)
print("Best l1_ratio:", best_l1_ratio)

# Create an instance of the ElasticNet model with the best lambda (alpha) and l1_ratio parameters
elastic_net_tuned = ElasticNet(alpha=best_alpha, l1_ratio=best_l1_ratio, max_iter=10000)

# Fit the tuned model to the training data
elastic_net_tuned.fit(X_train, y_train)

# Make predictions on the testing data
y_pred = elastic_net_tuned.predict(X_test)

# Compute the R-squared on the testing data
r2 = r2_score(y_test, y_pred)
print('R-squared on testing data:', r2)

# Compute the RMSE on testing data
mse = mean_squared_error(y_test, y_pred)
rmse = np.sqrt(mse)
print('RMSE on testing data:', rmse)

# Compare RMSE to the scale of the target variable
if rmse < 0.5:
    print('The model has a very good fit.')
elif rmse < 1:
    print('The model has a good fit.')
else:
    print('The model has a moderate to poor fit.')

排查与优化建议

代码逻辑排查

  • 数据划分验证:打印data.head()确认第一列确实是目标变量,其余列是特征,避免特征与标签搞反。
  • 模型收敛性检查:训练后打印elastic_net_tuned.n_iter_,如果结果等于设置的max_iter=10000,说明模型未收敛,需增大max_iter(如调到100000)或降低tol参数的收敛阈值。
  • 网格搜索参数扩展:当前alpha范围仅到1.0,建议添加0.01, 0.05, 2.0, 5.0;l1_ratio补充0.0, 1.0(对应纯Ridge和纯Lasso),覆盖更多正则化组合。

数据预处理优化

  • 特征标准化:ElasticNet对特征尺度敏感,必须做标准化处理,示例代码:
    from sklearn.preprocessing import StandardScaler
    scaler = StandardScaler()
    X_train_scaled = scaler.fit_transform(X_train)
    X_test_scaled = scaler.transform(X_test)
    # 后续用缩放后的特征训练模型
    
  • 数据分布检查:用sns.histplot(y)查看目标变量分布,若为非线性分布,可对y做对数、Box-Cox等变换后再训练。
  • 特征筛选:若特征维度远大于样本量,计算特征与y的相关性,剔除低相关特征;或用方差筛选去掉方差接近0的冗余特征。

模型评估与调整

  • 对比训练集指标:计算训练集的R²和RMSE,若训练集指标也差,说明欠拟合,需减小alpha降低正则化强度;若训练集好但测试集差,说明过拟合,需增大alpha。
  • 分析交叉验证结果:查看grid_search.cv_results_中不同参数的交叉验证得分,若最优参数在网格边界(如当前alpha为1.0是最大值),需继续扩展参数范围测试。
  • 尝试其他模型:若线性模型不适合当前数据,可尝试随机森林、XGBoost等树模型,对比拟合效果。

内容的提问来源于stack exchange,提问作者Toiba

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 10:18:11