You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:按分组规则对DataFrame列进行条件分类

问题描述

原始DataFrame如下:

group   id
A   009x
A   010x
B   009x
B   002x
C   002x
C   003x

需要按group分组生成新列new,规则如下:

  • 若组内所有id仅包含009x和010x,则分类为g1
  • 若组内同时存在009x/010x以及其他id值,则分类为g2
  • 其余情况(组内无009x/010x),直接显示id值

期望结果:

group   id  new
A   009x    g1
A   010x    g1
B   009x    g2
B   002x    g2
C   002x    002x
C   003x    003x

示例数据生成代码:

import pandas as pd

data = {
    'group': ['A', 'A', 'B', 'B', 'C', 'C'],
    'id': ['009x', '010x', '009x', '002x', '002x', '003x'], 
}  
df = pd.DataFrame(data)
解决方案

可以通过分组计算组内特征,结合numpy.select批量判断赋值:

import pandas as pd
import numpy as np

# 生成示例数据
data = {
    'group': ['A', 'A', 'B', 'B', 'C', 'C'],
    'id': ['009x', '010x', '009x', '002x', '002x', '003x'], 
}  
df = pd.DataFrame(data)

# 定义目标id集合
target_ids = {'009x', '010x'}

# 按group分组,计算组内是否包含目标id、是否全为目标id
group_features = df.groupby('group')['id'].agg(
    has_target=lambda x: any(i in target_ids for i in x),
    all_target=lambda x: all(i in target_ids for i in x)
).reset_index()

# 合并分组特征到原表
df = df.merge(group_features, on='group', how='left')

# 按条件设置new列
conditions = [
    df['all_target'],
    df['has_target'] & ~df['all_target'],
]
choices = ['g1', 'g2']
df['new'] = np.select(conditions, choices, default=df['id'])

# 清理临时列
df = df.drop(['has_target', 'all_target'], axis=1)

print(df)

代码说明

  1. 先定义目标id集合target_ids,简化后续判断逻辑
  2. 分组计算每个组的两个核心特征:是否包含目标id、是否全为目标id
  3. 将分组特征合并到原DataFrame,让每行都能获取所在组的特征
  4. 用np.select按优先级匹配条件,满足对应规则则赋值g1/g2,否则直接使用原id值
  5. 删除临时生成的特征列,得到最终结果

内容的提问来源于stack exchange,提问作者shsh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 20:35:30