You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pandas根据companyId及type字段过滤DataFrame行?

Pandas DataFrame 按规则过滤行

需求说明

  • 当同一companyId存在多条记录时,仅保留type为actual的行
  • 当同一companyId仅存在一条记录时,无论type字段取值,均保留该行

实现方法

方法一:分组后条件筛选

通过groupby按companyId分组,对每个组根据记录数执行不同的筛选逻辑:

import pandas as pd

# 构造原始DataFrame
data = {'companyId': {0: 198236, 1: 198236, 2: 900814, 3: 153421, 4: 153421, 5: 337815},
        'region': {0: 'Europe', 1: 'Europe', 2: 'Asia-Pacific', 3: 'North America', 4: 'North America', 5:'Africa'},
        'value': {0: 560, 1: 771, 2: 964, 3: 217, 4: 433, 5: 680},
        'type': {0: 'actual', 1: 'forecast', 2: 'actual', 3: 'forecast', 4: 'actual', 5: 'forecast'}}
df = pd.DataFrame(data)

# 分组筛选逻辑
result = df.groupby('companyId').apply(
    lambda x: x[x['type'] == 'actual'] if len(x) > 1 else x
).reset_index(drop=True)

print(result)

方法二:布尔索引结合transform计数

先计算每个companyId的记录数,再通过布尔条件直接过滤,效率更高:

# 计算每个companyId的记录数
counts = df.groupby('companyId')['companyId'].transform('size')

# 构造过滤条件:要么记录数为1,要么记录数>1且type为actual
mask = (counts == 1) | ((counts > 1) & (df['type'] == 'actual'))
result = df[mask]

print(result)

输出结果

两种方法都会得到以下过滤后的DataFrame:

companyId         region  value     type
0     198236         Europe    560   actual
2     900814    Asia-Pacific    964   actual
4     153421  North America    433   actual
5     337815         Africa    680  forecast

内容的提问来源于stack exchange,提问作者A.N.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 21:39:21