Sklearn中random_state选择问题:不同取值下Lasso预测效果差异
如何选择合适的random_state?
我正在学习机器学习中的回归算法,发现将random_state设为42时,Lasso模型的预测效果很差,但设为2时效果相反。请问该如何选择合适的random_state?
以下是我编写的基础代码:
import pandas as pd import matplotlib.pyplot as plt import yfinance as yf data = yf.download('NG=F', '2010-06-06', '2023-06-06', auto_adjust=True) # 天然气期货 print(data.head(20)) print(data.shape) print(data.info()) data.Close.plot(figsize=(10,7)) plt.show() import seaborn as sbn from sklearn.model_selection import train_test_split x = data.drop('Close', axis=1) y = data['Close'] x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=42) from sklearn.linear_model import LinearRegression line_near = LinearRegression() line_near.fit(x_train, y_train) predictions = line_near.predict(x_test) print(f'Actual values: {y_test[0:10]}') print(f'Predictions: {predictions[0:10]}') from sklearn.metrics import mean_squared_error mean_error = mean_squared_error(y_test, predictions) print(mean_error) from sklearn.linear_model import Ridge ridge = Ridge() ridge.fit(x_train, y_train) ridge_predictions = ridge.predict(x_test) print(f'Actual: {ridge_predictions[0:10]}') print(f'Predictions: {y_test[0:10]}') mean_error_ridge = mean_squared_error(y_test, ridge_predictions) print(mean_error_ridge) from sklearn.linear_model import Lasso lasso = Lasso() lasso.fit(x_train, y_train) lasso_predictions = lasso.predict(x_test) print(f'Actual: {y_test[0:10]}') print(f'predictions: {lasso_predictions[0:10]}') mean_error_lasso = mean_squared_error(y_test, lasso_predictions) print(mean_error_lasso)
解决方案:
- 不要依赖单一随机种子:你遇到的是典型的随机划分偶然性问题——不同的
random_state会生成不同的训练/测试集,若测试集中刚好包含Lasso拟合难度高的样本,就会出现效果骤降。单一种子的结果不能代表模型真实性能。 - 用交叉验证评估模型:用K折交叉验证替代单次
train_test_split,比如sklearn.model_selection.cross_val_score,它会多次划分数据集并取性能平均值,能稳定反映模型真实能力,避免随机划分干扰。 - 检查数据集分布一致性:验证不同划分下训练集和测试集的特征、目标变量分布是否一致(比如画直方图、计算统计量)。你的数据是期货价格这类时序数据,随机划分会破坏时间顺序,导致模型学到未来信息,此时应改用时序划分(按时间节点切分训练/测试集)。
- 调优Lasso超参数:Lasso默认的
alpha=1可能不适配你的数据,结合交叉验证用GridSearchCV或RandomizedSearchCV搜索最优alpha值,这能大幅提升模型性能,同时规避单一随机划分的影响。 - 固定随机种子的场景:仅需复现特定实验结果时,才固定
random_state;评估模型性能时,必须多次随机划分取平均或用交叉验证。
内容的提问来源于stack exchange,提问作者nnguyenquy
相关产品推荐
相关产品推荐

