多项式回归折线绘制异常求助:如何修正来回波动的折线?
问题原因及解决方案
核心问题
你绘制出波动折线的根本原因是:绘制回归曲线时使用的x数据未按从小到大排序,plt.plot()会严格按照原始数据的顺序连接预测点,当原始x样本是乱序时,就会出现来回跳跃的折线,而非平滑的多项式曲线。
另外你的代码存在变量名不匹配问题:代码里用了未定义的x、hx,但实际定义的输入变量是x_train_full,预测结果是y_prediction,这会导致未定义变量报错。
修正后的完整代码
import pandas as pd from sklearn.preprocessing import PolynomialFeatures from sklearn.linear_model import LinearRegression import matplotlib.pyplot as plt df_full = pd.read_excel(r'lab_test.xlsx', sheet_name='tests') x_train_full = df_full.loc[:, 'test(mg)'].values y_train_full = df_full.loc[:, 'chance %'].values # 多项式特征转换 poly = PolynomialFeatures(degree=2) x_poly = poly.fit_transform(x_train_full.reshape(-1, 1)) # 训练模型 model = LinearRegression() model.fit(x_poly, y_train_full) # 关键修正:对x排序并生成对应排序后的预测值 sorted_indices = x_train_full.argsort() x_sorted = x_train_full[sorted_indices] # 生成排序后x的多项式特征再预测,保证对应关系正确 x_poly_sorted = poly.transform(x_sorted.reshape(-1, 1)) y_pred_sorted = model.predict(x_poly_sorted) # 绘图 plt.xlabel('test(mg)') plt.ylabel('chance %') plt.scatter(x_train_full, y_train_full, label='original data') plt.plot(x_sorted, y_pred_sorted, 'r', label='regression line') plt.legend(loc='upper left') plt.show()
关键步骤说明
- 用
argsort()获取x数据从小到大排序的索引,通过该索引重新排序x,确保绘图时x轴是递增顺序 - 必须基于排序后的x重新生成多项式特征再预测,不能直接使用原x的预测值,避免因排序导致x与预测值的对应关系错乱
- 用排序后的x和对应预测值绘图,
plt.plot()会按x递增顺序连接点,得到平滑的二次多项式曲线
内容的提问来源于stack exchange,提问作者aye
相关产品推荐
相关产品推荐

