如何仅对指定自变量应用scikit-learn的PolynomialFeatures变换
无需额外安装包的实现方案
不需要依赖sklearn_pandas这类第三方工具,直接基于现有导入的pandas、numpy、sklearn组件,通过「拆分特征分别处理+结果拼接」的逻辑就能实现需求。
核心思路
- 仅将
x1列单独输入PolynomialFeatures做3阶多项式变换,设置include_bias=False关闭默认生成的全1偏置列,避免和线性回归自带的截距项产生共线性 - 保留所有不需要做多项式变换的特征:
x2列、color独热编码生成的所有衍生列 - 将多项式变换后的特征、保留的原始特征按样本维度横向拼接,得到最终的训练特征集,后续直接输入线性回归模型即可
修改后的可运行代码
只需要替换原代码中多项式生成、模型训练部分的逻辑即可:
import pandas as pd import numpy as np from sklearn.linear_model import LinearRegression from sklearn.preprocessing import PolynomialFeatures x1 = [28.0, 29.0, 12.0, 12.0, 42.0, 35.0, 28.0, 30.0, 32.0, 46.0, 18.0, 28.0, 28.0, 64.0, 38.0, 18.0, 49.0, 37.0, 25.0, 24.0, 42.0, 50.0, 12.0, 64.0, 23.0, 35.0, 22.0, 16.0, 44.0, 77.0, 26.0, 44.0, 38.0, 37.0, 45.0, 42.0, 24.0, 42.0, 12.0, 46.0, 12.0, 26.0, 37.0, 15.0, 67.0, 36.0, 43.0, 36.0, 45.0, 82.0, 44.0, 30.0, 33.0, 51.0, 50.0] x2 = [0.36, 0.53, 0.45, 0.48, 0.4, 0.44, 0.44, 0.6, 0.39, 0.39, 0.29, 0.52, 0.46, 0.55, 0.62, 0.53, 0.79, 0.57, 0.49, 0.23, 0.55, 0.54, 0.44, 0.74, 0.36, 0.46, 0.37, 0.38, 0.75, 0.8, 0.43, 0.43, 0.58, 0.38, 0.63, 0.39, 0.14, 0.26, 0.14, 0.62, 0.49, 0.46, 0.49, 0.53, 0.73, 0.48, 0.5, 0.47, 0.49, 0.83, 0.56, 0.22, 0.49, 0.43, 0.46] y = [59.5833333333333, 59.5833333333333, 10.0, 10.0, 47.0833333333333, 51.2499999999999, 34.5833333333333, 88.75, 63.7499999999999, 34.5833333333333, 51.2499999999999, 10.0, 63.7499999999999, 51.0, 59.5833333333333, 47.0833333333333, 49.5625, 43.5624999999999, 63.7499999999999, 10.0, 76.25, 47.0833333333333, 10.0, 51.2499999999999, 47.0833333333333, 10.0, 35.0, 51.2499999999999, 76.25, 100.0, 51.2499999999999, 59.5833333333333, 63.7499999999999, 76.25, 100.0, 51.2499999999999, 10.0, 22.5, 10.0, 88.75, 10.0, 59.5833333333333, 47.0833333333333, 34.5833333333333, 51.2499999999999, 63.7499999999999, 63.7499999999999, 10.0, 76.25, 62.1249999999999, 47.0833333333333, 10.0, 76.25, 47.0833333333333, 88.75] color = ['green','red','blue','purple','black','white','orange','grey ','gold','yellow','white','orange','grey ','green','red','purple','orange','grey ','gold','yellow','white','orange','grey ','green','red','blue','black','white','orange','grey ','gold','yellow','white','orange','grey ','green','red','blue','purple','orange','grey ','gold','green','red','blue','purple','black','white','orange','grey ','gold','yellow','white','orange','grey '] df_final = pd.DataFrame({ 'x1': x1, 'x2' :x2, 'y': y, 'color':color}) # 独热编码处理分类变量 df_final = pd.get_dummies(df_final, columns=['color']) color_cols = [c for c in df_final.columns if c.startswith('color_')] # --- 核心修改部分开始 --- # 仅对x1生成3阶多项式特征,关闭偏置项 poly = PolynomialFeatures(degree=3, include_bias=False) x1_poly_features = poly.fit_transform(df_final[['x1']]) # 提取不需要做多项式变换的特征 x_keep = df_final[['x2'] + color_cols].values # 横向拼接所有特征 X_train = np.hstack([x1_poly_features, x_keep]) # --- 核心修改部分结束 --- # 模型训练 lin2 = LinearRegression() lin2.fit(X_train, y) intercept = lin2.intercept_ # 截距项 coeff = lin2.coef_ # 特征系数,顺序对应:x1、x1²、x1³、x2、各color独热列 r_squared_lin2 = lin2.score(X_train, y) print(r_squared_lin2)
可选优化
如果需要把系数和特征名一一对应,可以不用转numpy数组,直接用pandas做拼接保留列名:
# 给多项式特征命名 poly_col_names = ['x1', 'x1_2', 'x1_3'] x1_poly_df = pd.DataFrame(x1_poly_features, columns=poly_col_names, index=df_final.index) x_keep_df = df_final[['x2'] + color_cols] X_train_named = pd.concat([x1_poly_df, x_keep_df], axis=1)
内容的提问来源于stack exchange,提问作者user032020
相关产品推荐
相关产品推荐

