You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中按指定规则筛选DataFrame的行?

实现指定规则的DataFrame筛选操作

原始数据

首先定义初始的DataFrame:

import pandas as pd

big = pd.DataFrame({'group': ['A', 'A', 'A','A', 'B','B','C','D','D', 'D'], 
                    'animal': ['ALL other', 'cat','rabbit', 'dog', 'rabbit','ALL other', 'ALL', 'ALL other', 'dog','cat']})

输出内容:

group     animal
0     A  ALL other
1     A        cat
2     A     rabbit
3     A        dog
4     B     rabbit
5     B  ALL other
6     C        ALL
7     D  ALL other
8     D        dog
9     D        cat

筛选规则

  • 若组内包含rabbit,选取该行
  • 若animal字段为ALL,选中该行
  • 若组内无rabbit,选取animal为ALL other的行

解决方案

方法一:向量式条件筛选(高效型)

适合大数据集,利用Pandas的向量操作提升效率:

# 给每行标记所在组是否包含rabbit
has_rabbit = big.groupby('group')['animal'].transform(lambda x: 'rabbit' in x.values)

# 组合筛选条件
filter_condition = (
    (has_rabbit & (big['animal'] == 'rabbit')) |
    (big['animal'] == 'ALL') |
    (~has_rabbit & (big['animal'] == 'ALL other'))
)

# 执行筛选并重置索引
result = big[filter_condition].reset_index(drop=True)

方法二:分组自定义函数(直观型)

逻辑完全贴合规则描述,可读性更强:

def process_single_group(group):
    if 'rabbit' in group['animal'].values:
        return group[group['animal'] == 'rabbit']
    elif 'ALL' in group['animal'].values:
        return group[group['animal'] == 'ALL']
    else:
        return group[group['animal'] == 'ALL other']

# 分组处理后合并结果
result = big.groupby('group', group_keys=False).apply(process_single_group).reset_index(drop=True)

两种方法最终输出结果一致:

group     animal
0     A     rabbit
1     B     rabbit
2     C        ALL
3     D  ALL other

内容的提问来源于stack exchange,提问作者huo shankou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 08:01:02