You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于不同分组的阈值条件,优雅实现Pandas数据打标签?

按分组自定义阈值打标签的优雅实现方案

我需要根据数值为数据打标签,目前已经实现了统一阈值的版本,代码如下:

import pandas as pd
import numpy as np

data = pd.DataFrame({
    'Group': ['group1', 'group1', 'group1', 'group2', 'group2', 'group2'],
    'Value': [30, 40, 10, 40, 60, 70]
})

conditions = [
    (data['Value'] < 50) & (data['Value'] >= 40),
    (data['Value'] < 40) & (data['Value'] >= 30)
]

results = ['Large', 'Small']

data['Label'] = np.select(conditions, results, default='Other')

这段代码运行正常,但我的目标是按分组执行操作,不同分组使用不同的阈值。我现在可以通过单独处理每个分组的方式实现,比如处理group1:

import pandas as pd
import numpy as np

data = pd.DataFrame({
    'Group': ['group1', 'group1', 'group1', 'group2', 'group2', 'group2'],
    'Value': [30, 40, 10, 40, 60, 70]
})

conditions = [
    (data.loc[data['Group']=='group1','Value'] < 50) & (data.loc[data['Group']=='group1','Value'] >= 40),
    (data.loc[data['Group']=='group1','Value'] < 40) & (data.loc[data['Group']=='group1','Value'] >= 30)
]

results = ['Large', 'Small']

data.loc[data['Group']=='group1','Label'] = np.select(conditions, results, default='Other')

再处理group2:

import pandas as pd
import numpy as np

data = pd.DataFrame({
    'Group': ['group1', 'group1', 'group1', 'group2', 'group2', 'group2'],
    'Value': [30, 40, 10, 40, 60, 70]
})

conditions = [
    (data.loc[data['Group']=='group2','Value'] < 60) & (data.loc[data['Group']=='group2','Value'] >= 50),
    (data.loc[data['Group']=='group2','Value'] < 50) & (data.loc[data['Group']=='group2','Value'] >= 40)
]

results = ['Large', 'Small']

data.loc[data['Group']=='group2','Label'] = np.select(conditions, results, default='Other')

但真实数据集包含更多分组与更多条件,这种重复写法效率太低,希望找到更优雅的解决方案。

内容的提问来源于stack exchange,提问作者Derek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 05:55:15