如何在Scikit-learn Pipeline中删除生成新特征后的原列
问题:Sklearn Pipeline中拆分字符串列后无法移除原列?
我有一份示例数据,其中一列包含字符串值(如"34 12")。我在预处理步骤中创建了两个新列,分别存储该字符串列拆分后的左右整数。最终我希望移除原字符串列,但不知道如何在Pipeline内实现该操作。
尝试在ColumnTransformer中使用("column_dropper", "drop", ["string1"]),但查看x_transformed时发现原字符串列仍存在:
array([[1.0, 6.5, '34 12', 34, 12], [2.0, 6.0, '34 5', 34, 5], [1.5, 7.0, '56 6', 56, 6]], dtype=object)
复现代码如下:
import pandas as pd from sklearn.preprocessing import FunctionTransformer from sklearn.impute import SimpleImputer from sklearn.pipeline import Pipeline from sklearn.compose import ColumnTransformer from sklearn.preprocessing import LabelEncoder from sklearn.base import BaseEstimator, TransformerMixin # 创建示例数据 data= {"string1": ["34 12", "34 5", "56 6"], "age": [1, 2, None], "grade": [None, 6, 7]} x_train = pd.DataFrame(data=data) # 定义拆分函数 def extract_int2(x): num = x.split(" ")[-1] if num.isnumeric(): return int(num) else: return 0 def extract_int1(x): num = x.split(" ")[0] if num.isnumeric(): return int(num) else: return 0 def int_features(df): df["num1"] = df["string1"].apply(extract_int1) df["num2"] = df["string1"].apply(extract_int2) return df columns_to_drop="string1" # 定义Pipeline num_vals = Pipeline([("imputer", SimpleImputer(strategy = "mean"))]) features_vals = Pipeline([("new_features", FunctionTransformer(int_features, validate=False))]) preprocess_pipeline = ColumnTransformer(transformers=[ ("num_preprocess", num_vals, ["age", "grade"]), ("feature_preprocess", features_vals, ["string1"]), ("column_dropper", "drop", ["string1"])]) preprocess_pipeline.fit(x_train) x_transformed = preprocess_pipeline.transform(x_train) x_transformed
我还尝试用自定义删除函数结合FunctionTransformer(),但同样无效:
def drop_column(df): df = df.drop(columns=["string1"]) return df # 定义Pipeline num_vals = Pipeline([("imputer", SimpleImputer(strategy = "mean"))]) features_vals = Pipeline([("new_features", FunctionTransformer(int_features, validate=False))]) dropping= Pipeline([("drop_string", FunctionTransformer(drop_column))]) preprocess_pipeline = ColumnTransformer(transformers=[ ("num_preprocess", num_vals, ["age", "grade"]), ("feature_preprocess", features_vals, ["string1"]), ("drop_preprocess", dropping, ["string1"])] ) preprocess_pipeline.fit(x_train) x_transformed = preprocess_pipeline.transform(x_train) x_transformed
解决方案
问题原因
ColumnTransformer的工作逻辑是对每个transformer单独处理指定的列,然后将所有transformer的输出按顺序拼接。之前的写法中:
feature_preprocess处理["string1"]列,返回的是包含原string1列+新num1/num2列的DataFramedrop或者自定义删除transformer处理["string1"]列,返回空(因为丢弃了该列)- 最终拼接结果是:数值列处理结果 +
feature_preprocess的完整结果(含原列) + 空,所以原列依然存在。
正确实现方式
方法1:修改字符串处理函数,仅返回新特征列
让处理字符串列的步骤只输出需要保留的新列,不包含原字符串列:
import pandas as pd from sklearn.preprocessing import FunctionTransformer from sklearn.impute import SimpleImputer from sklearn.pipeline import Pipeline from sklearn.compose import ColumnTransformer # 创建示例数据 data= {"string1": ["34 12", "34 5", "56 6"], "age": [1, 2, None], "grade": [None, 6, 7]} x_train = pd.DataFrame(data=data) # 重新定义拆分函数 def extract_int2(x): num = x.split(" ")[-1] return int(num) if num.isnumeric() else 0 def extract_int1(x): num = x.split(" ")[0] return int(num) if num.isnumeric() else 0 def create_num_features(df): # 只返回新生成的两列,不保留原string1 return pd.DataFrame({ "num1": df["string1"].apply(extract_int1), "num2": df["string1"].apply(extract_int2) }) # 定义各Pipeline num_vals = Pipeline([("imputer", SimpleImputer(strategy="mean"))]) str_to_num = Pipeline([("new_features", FunctionTransformer(create_num_features, validate=False))]) preprocess_pipeline = ColumnTransformer(transformers=[ ("num_preprocess", num_vals, ["age", "grade"]), ("str_to_num", str_to_num, ["string1"]) ]) preprocess_pipeline.fit(x_train) x_transformed = preprocess_pipeline.transform(x_train) print(x_transformed)
输出结果:
array([[ 1. , 6.5, 34. , 12. ], [ 2. , 6. , 34. , 5. ], [ 1.5, 7. , 56. , 6. ]])
方法2:用自定义Transformer统一处理所有列(含删除原列)
自定义Transformer,一次性完成生成新特征+删除原列的操作,再和数值列预处理结合:
import pandas as pd from sklearn.preprocessing import FunctionTransformer from sklearn.impute import SimpleImputer from sklearn.pipeline import Pipeline from sklearn.compose import ColumnTransformer from sklearn.base import BaseEstimator, TransformerMixin # 创建示例数据 data= {"string1": ["34 12", "34 5", "56 6"], "age": [1, 2, None], "grade": [None, 6, 7]} x_train = pd.DataFrame(data=data) class StringSplitAndDrop(BaseEstimator, TransformerMixin): def fit(self, X, y=None): return self def transform(self, X): X_copy = X.copy() # 生成新特征 X_copy["num1"] = X_copy["string1"].apply(lambda x: int(x.split()[0]) if x.split()[0].isnumeric() else 0) X_copy["num2"] = X_copy["string1"].apply(lambda x: int(x.split()[1]) if x.split()[1].isnumeric() else 0) # 删除原列 X_copy.drop("string1", axis=1, inplace=True) return X_copy # 定义Pipeline full_preprocess = Pipeline([ ("split_and_drop", StringSplitAndDrop()), ("imputer", SimpleImputer(strategy="mean")) ]) full_preprocess.fit(x_train) x_transformed = full_preprocess.transform(x_train) print(x_transformed)
输出结果:
array([[ 1. , 6.5, 34. , 12. ], [ 2. , 6. , 34. , 5. ], [ 1.5, 7. , 56. , 6. ]])
关键说明
- ColumnTransformer的每个transformer独立作用于指定列,若某个transformer返回了原列,后续的drop操作无法移除它,因为drop处理的是单独的列子集,不会影响其他transformer的输出。
- 最佳实践是让处理字符串列的步骤仅输出需要保留的新特征,或用完整的Transformer统一处理列的生成与删除,避免拼接时混入不需要的列。
内容的提问来源于stack exchange,提问作者user22815609
相关产品推荐
相关产品推荐

