You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scikit-learn Pipeline中删除生成新特征后的原列

问题:Sklearn Pipeline中拆分字符串列后无法移除原列?

我有一份示例数据,其中一列包含字符串值(如"34 12")。我在预处理步骤中创建了两个新列,分别存储该字符串列拆分后的左右整数。最终我希望移除原字符串列,但不知道如何在Pipeline内实现该操作。

尝试在ColumnTransformer中使用("column_dropper", "drop", ["string1"]),但查看x_transformed时发现原字符串列仍存在:

array([[1.0, 6.5, '34 12', 34, 12],
       [2.0, 6.0, '34 5', 34, 5],
       [1.5, 7.0, '56 6', 56, 6]], dtype=object)

复现代码如下:

import pandas as pd
from sklearn.preprocessing import FunctionTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import LabelEncoder
from sklearn.base import BaseEstimator, TransformerMixin

# 创建示例数据
data= {"string1": ["34 12", "34 5", "56 6"], "age": [1, 2, None], "grade": [None, 6, 7]}
x_train = pd.DataFrame(data=data)

# 定义拆分函数
def extract_int2(x):
    num = x.split(" ")[-1]
    if num.isnumeric():
        return int(num)
    else:
        return 0
            
def extract_int1(x):
    num = x.split(" ")[0]
    if num.isnumeric():
        return int(num)
    else:
        return 0

def int_features(df):
    df["num1"] = df["string1"].apply(extract_int1)
    df["num2"] = df["string1"].apply(extract_int2)
    return df

columns_to_drop="string1"

# 定义Pipeline
num_vals =  Pipeline([("imputer", SimpleImputer(strategy = "mean"))])
features_vals = Pipeline([("new_features", FunctionTransformer(int_features, validate=False))])

preprocess_pipeline = ColumnTransformer(transformers=[
    ("num_preprocess", num_vals, ["age", "grade"]),
    ("feature_preprocess", features_vals, ["string1"]),
    ("column_dropper", "drop", ["string1"])])

preprocess_pipeline.fit(x_train)

x_transformed = preprocess_pipeline.transform(x_train)
x_transformed

我还尝试用自定义删除函数结合FunctionTransformer(),但同样无效:

def drop_column(df):
    df = df.drop(columns=["string1"])
    return df

# 定义Pipeline
num_vals =  Pipeline([("imputer", SimpleImputer(strategy = "mean"))])
features_vals = Pipeline([("new_features", FunctionTransformer(int_features, validate=False))])
dropping= Pipeline([("drop_string", FunctionTransformer(drop_column))])

preprocess_pipeline = ColumnTransformer(transformers=[
    ("num_preprocess", num_vals, ["age", "grade"]),
    ("feature_preprocess", features_vals, ["string1"]),
    ("drop_preprocess", dropping, ["string1"])]
                              )

preprocess_pipeline.fit(x_train)

x_transformed = preprocess_pipeline.transform(x_train)
x_transformed

解决方案

问题原因

ColumnTransformer的工作逻辑是对每个transformer单独处理指定的列,然后将所有transformer的输出按顺序拼接。之前的写法中:

  • feature_preprocess处理["string1"]列,返回的是包含原string1列+新num1/num2列的DataFrame
  • drop或者自定义删除transformer处理["string1"]列,返回空(因为丢弃了该列)
  • 最终拼接结果是:数值列处理结果 + feature_preprocess的完整结果(含原列) + 空,所以原列依然存在。

正确实现方式

方法1:修改字符串处理函数,仅返回新特征列

让处理字符串列的步骤只输出需要保留的新列,不包含原字符串列:

import pandas as pd
from sklearn.preprocessing import FunctionTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer

# 创建示例数据
data= {"string1": ["34 12", "34 5", "56 6"], "age": [1, 2, None], "grade": [None, 6, 7]}
x_train = pd.DataFrame(data=data)

# 重新定义拆分函数
def extract_int2(x):
    num = x.split(" ")[-1]
    return int(num) if num.isnumeric() else 0
            
def extract_int1(x):
    num = x.split(" ")[0]
    return int(num) if num.isnumeric() else 0

def create_num_features(df):
    # 只返回新生成的两列,不保留原string1
    return pd.DataFrame({
        "num1": df["string1"].apply(extract_int1),
        "num2": df["string1"].apply(extract_int2)
    })

# 定义各Pipeline
num_vals = Pipeline([("imputer", SimpleImputer(strategy="mean"))])
str_to_num = Pipeline([("new_features", FunctionTransformer(create_num_features, validate=False))])

preprocess_pipeline = ColumnTransformer(transformers=[
    ("num_preprocess", num_vals, ["age", "grade"]),
    ("str_to_num", str_to_num, ["string1"])
])

preprocess_pipeline.fit(x_train)
x_transformed = preprocess_pipeline.transform(x_train)
print(x_transformed)

输出结果:

array([[ 1. ,  6.5, 34. , 12. ],
       [ 2. ,  6. , 34. ,  5. ],
       [ 1.5,  7. , 56. ,  6. ]])

方法2:用自定义Transformer统一处理所有列(含删除原列)

自定义Transformer,一次性完成生成新特征+删除原列的操作,再和数值列预处理结合:

import pandas as pd
from sklearn.preprocessing import FunctionTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.base import BaseEstimator, TransformerMixin

# 创建示例数据
data= {"string1": ["34 12", "34 5", "56 6"], "age": [1, 2, None], "grade": [None, 6, 7]}
x_train = pd.DataFrame(data=data)

class StringSplitAndDrop(BaseEstimator, TransformerMixin):
    def fit(self, X, y=None):
        return self
    
    def transform(self, X):
        X_copy = X.copy()
        # 生成新特征
        X_copy["num1"] = X_copy["string1"].apply(lambda x: int(x.split()[0]) if x.split()[0].isnumeric() else 0)
        X_copy["num2"] = X_copy["string1"].apply(lambda x: int(x.split()[1]) if x.split()[1].isnumeric() else 0)
        # 删除原列
        X_copy.drop("string1", axis=1, inplace=True)
        return X_copy

# 定义Pipeline
full_preprocess = Pipeline([
    ("split_and_drop", StringSplitAndDrop()),
    ("imputer", SimpleImputer(strategy="mean"))
])

full_preprocess.fit(x_train)
x_transformed = full_preprocess.transform(x_train)
print(x_transformed)

输出结果:

array([[ 1. ,  6.5, 34. , 12. ],
       [ 2. ,  6. , 34. ,  5. ],
       [ 1.5,  7. , 56. ,  6. ]])

关键说明

  • ColumnTransformer的每个transformer独立作用于指定列,若某个transformer返回了原列,后续的drop操作无法移除它,因为drop处理的是单独的列子集,不会影响其他transformer的输出。
  • 最佳实践是让处理字符串列的步骤仅输出需要保留的新特征,或用完整的Transformer统一处理列的生成与删除,避免拼接时混入不需要的列。

内容的提问来源于stack exchange,提问作者user22815609

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 13:34:57