Python中分组数据先Scale再Binarize的预处理实现问题
按组实现Scale+Binarize一站式预处理
问题核心
要在groupby流程内完成两步预处理:先按组标准化(scale),再将均值以上数据设为1、以下设为0(binarize),避免后续单独reshape操作。
解决方案
直接在transform的lambda函数内整合两个步骤,无需拆分处理。这里提供两种实现方式:
方式1:整合sklearn的scale与binarize
利用reshape适配binarize的输入要求,将两步逻辑合并到同一个lambda中:
import pandas as pd from sklearn import preprocessing df = pd.DataFrame({'group': ['A', 'A', 'A', 'B','B', 'B'], 'column_to_scale': [4, 2, 6, 4, 9, 6]}) # 一站式完成按组scale+binarize df['column_to_scale'] = df.groupby("group")['column_to_scale'].transform( lambda x: preprocessing.binarize(preprocessing.scale(x).reshape(-1, 1), threshold=0).flatten() ) print(df)
方式2:利用标准化特性简化逻辑
因为preprocessing.scale会将数据标准化为均值0、方差1,所以直接判断标准化后的值是否大于0即可实现需求,无需调用binarize,更高效:
import pandas as pd from sklearn import preprocessing df = pd.DataFrame({'group': ['A', 'A', 'A', 'B','B', 'B'], 'column_to_scale': [4, 2, 6, 4, 9, 6]}) # 更简洁的实现:scale后直接判断并转int df['column_to_scale'] = df.groupby("group")['column_to_scale'].transform( lambda x: (preprocessing.scale(x) > 0).astype(int) ) print(df)
输出验证
两种方式都会得到预期结果:
group column_to_scale 0 A 0 1 A 0 2 A 1 3 B 0 4 B 1 5 B 0
关键说明
transform会自动保留原数据的索引结构,无需额外对齐操作- 方式2的布尔判断逻辑完全等价于
binarize(threshold=0),且避免了数组形状转换的麻烦
内容的提问来源于stack exchange,提问作者Tim Romero
相关产品推荐
相关产品推荐

