You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Sklearn Pipeline中为不同特征设置自定义阈值使用Binarizer?

为不同特征设置不同阈值完成二值化的解决方案

你遇到的问题核心是Binarizer本身只支持单个阈值,没法一次性给多个特征分配不同阈值。要解决这个,得在ColumnTransformer里为每个特征单独配置对应的Binarizer实例,每个实例用自己的阈值。

直接上修改后的可运行代码:

import pandas as pd
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import Binarizer

X = pd.DataFrame({"age": [20, 35, 67, 85, 98, 33, 28],
                  "BMI": [21.2, 24.2, 19.8, 28.1, 18.6, 31.3, 22.3]})
y = [1,0,0,0,0,0,0]

thresholds = {"age": 67, "BMI": 25}

# 遍历阈值字典,为每个特征创建专属的Binarizer转换器
transformers = [
    (f"{col}_binarizer", Binarizer(threshold=threshold), [col])
    for col, threshold in thresholds.items()
]

pipe = Pipeline([
    ('binarize', ColumnTransformer(transformers=transformers))
])

# 执行转换并查看结果
result = pipe.fit_transform(X, y)
print(result)

运行后会输出符合预期的二值化结果:

[[0. 0.]
 [0. 0.]
 [0. 0.]
 [1. 1.]
 [1. 0.]
 [0. 1.]
 [0. 0.]]

其中第一列对应age的二值化(大于67为1),第二列对应BMI的二值化(大于25为1)。如果你的特征数量很多,这种列表推导式的方式可以自动生成所有转换器,不用手动逐个编写。

内容的提问来源于stack exchange,提问作者Larsq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 07:52:13