Pipeline转DataFrame报DataFrame constructor not properly called如何解决
错误根因
你调用gblpq.fit(X_train,y_train)返回的是Pipeline实例对象,不是经过标准化处理后的数值数组,直接将Pipeline对象传入pd.DataFrame()构造函数不符合参数要求,因此抛出该错误。fit()方法仅用于对训练集拟合预处理/模型参数,不会返回转换后的数据集。
修复方案
把fit()替换为fit_transform(),直接返回拟合+转换后的训练集标准化结果,再传入DataFrame即可。如果后续需要对测试集做相同标准化处理,直接调用拟合好的Pipeline的transform(X_test)方法即可,无需重复拟合。
修正后可运行代码
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler from sklearn.pipeline import Pipeline # 假设DATA是你读取后的鸢尾花数据集 X = DATA.drop(['class'], axis = 'columns') y = DATA['class'].values X_train, X_test, y_train, y_test = train_test_split(X,y, test_size=0.20, random_state=42) gbl_pl = [] gbl_pl.append(('standard_scaler_gb', StandardScaler())) gblpq = Pipeline(gbl_pl) # 替换fit为fit_transform,得到标准化后的数组 scaled_arr = gblpq.fit_transform(X_train, y_train) # 直接复用X_train的列名,避免手动写列名出现顺序错误 scaled_df = pd.DataFrame(scaled_arr, columns=X_train.columns) # 绘图部分统一用训练集的数据对比,避免全量数据和训练集分布有偏差 fig, (ax1, ax2) = plt.subplots(ncols=2, figsize=(10, 5)) ax1.set_title('Before Scaling') sns.kdeplot(X_train['petal_length'], ax=ax1) sns.kdeplot(X_train['petal_width'], ax=ax1) sns.kdeplot(X_train['sepal_length'], ax=ax1) sns.kdeplot(X_train['sepal_width'], ax=ax1) ax2.set_title('After Standard Scaler') sns.kdeplot(scaled_df['petal_length'], ax=ax2) sns.kdeplot(scaled_df['petal_width'], ax=ax2) sns.kdeplot(scaled_df['sepal_length'], ax=ax2) sns.kdeplot(scaled_df['sepal_width'], ax=ax2) plt.savefig("output73.png")
补充说明
如果Pipeline里包含多个预处理步骤,fit_transform()会按顺序执行所有步骤的拟合+转换操作,返回最终处理后的数组,直接转DataFrame即可,不需要单独提取每一步的输出。
内容的提问来源于stack exchange,提问作者Ctrl7
相关产品推荐
相关产品推荐

