You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中分组数据先Scale再Binarize的预处理实现问题

按组实现Scale+Binarize一站式预处理

问题核心

要在groupby流程内完成两步预处理:先按组标准化(scale),再将均值以上数据设为1、以下设为0(binarize),避免后续单独reshape操作。

解决方案

直接在transform的lambda函数内整合两个步骤,无需拆分处理。这里提供两种实现方式:


方式1:整合sklearn的scale与binarize

利用reshape适配binarize的输入要求,将两步逻辑合并到同一个lambda中:

import pandas as pd
from sklearn import preprocessing

df = pd.DataFrame({'group': ['A', 'A', 'A', 'B','B', 'B'],
                    'column_to_scale': [4, 2, 6, 4, 9, 6]})

# 一站式完成按组scale+binarize
df['column_to_scale'] = df.groupby("group")['column_to_scale'].transform(
    lambda x: preprocessing.binarize(preprocessing.scale(x).reshape(-1, 1), threshold=0).flatten()
)

print(df)

方式2:利用标准化特性简化逻辑

因为preprocessing.scale会将数据标准化为均值0、方差1,所以直接判断标准化后的值是否大于0即可实现需求,无需调用binarize,更高效:

import pandas as pd
from sklearn import preprocessing

df = pd.DataFrame({'group': ['A', 'A', 'A', 'B','B', 'B'],
                    'column_to_scale': [4, 2, 6, 4, 9, 6]})

# 更简洁的实现:scale后直接判断并转int
df['column_to_scale'] = df.groupby("group")['column_to_scale'].transform(
    lambda x: (preprocessing.scale(x) > 0).astype(int)
)

print(df)

输出验证

两种方式都会得到预期结果:

group  column_to_scale
0     A                0
1     A                0
2     A                1
3     B                0
4     B                1
5     B                0

关键说明

  • transform会自动保留原数据的索引结构,无需额外对齐操作
  • 方式2的布尔判断逻辑完全等价于binarize(threshold=0),且避免了数组形状转换的麻烦

内容的提问来源于stack exchange,提问作者Tim Romero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 21:17:43