You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中标记跨不同分组的重复行?

实现基于分组的重复行标记需求

初始需求:识别在foo和bar两组中均存在的行

示例数据

import pandas as pd

df = pd.DataFrame({'A': [1, 2, 2, 3, 3, 1],
                   'B': ['foo', 'bar', 'foo', 'bar', 'foo', 'foo']})
df = df.sort_values('B')
print(df)

输出:

A    B
1  2  bar
3  3  bar
0  1  foo
2  2  foo
4  3  foo
5  1  foo

预期结果

A    B  Indicator
1  2  bar  True  # value 2 also present in foo, so returns True
3  3  bar  True  # value 3 also present in foo, so returns True
0  1  foo  False  # value 1 only present in foo, so returns False
2  2  foo  True  # value 2 also present in bar, so returns True
4  3  foo  True  # value 3 also present in bar, so returns True
5  1  foo  False  # value 1 only present in foo, so returns False

实现代码

# 筛选出B列为foo或bar的行(原数据仅含这两类可跳过此步)
filtered_df = df[df['B'].isin(['foo', 'bar'])]

# 统计每个A值对应的不同B分组数量
group_count = filtered_df.groupby('A')['B'].nunique()

# 生成标记列:A值在至少2个目标分组中存在则标记为True
df['Indicator'] = df['A'].map(group_count) >= 2

逻辑说明

  1. 先过滤出目标分组数据,排除其他分组的干扰;
  2. 用nunique()统计每个A值覆盖的不同B分组数;
  3. 通过map将统计结果映射回原表,判断是否满足跨2个分组的条件,生成标记列。

更新需求:标记A列值出现在至少3个不同分组中的行

示例数据

df = pd.DataFrame({'A': [1, 2, 2, 3, 3, 2, 1],  'B': ['foo', 'bar', 'foo', 'bar', 'foo', 'baz', 'baz']})
df = df.sort_values('B')
print(df)

输出:

A    B
1  2  bar
3  3  bar
5  2  baz
6  1  baz
0  1  foo
2  2  foo
4  3  foo

预期结果

A    B  Indicator
1  2  bar  True  # The value 2 occurs in categories baz, bar, and foo, so returns True.
3  3  bar  False  # The value 3 only occurs in categories bar and foo, so returns False.
5  2  baz  True  # The value 2 occurs in categories baz, bar, and foo, so returns True.
6  1  baz  False  # The value 1 only occurs in categories baz and foo, so returns False.
0  1  foo  False  # The value 1 only occurs in categories baz and foo, so returns False.
2  2  foo  True  # The value 2 occurs in categories baz, bar, and foo, so returns True.
4  3  foo  False  # The value 3 only occurs in categories bar and foo, so returns False.

实现代码

# 统计每个A值对应的不同B分组数量
group_count = df.groupby('A')['B'].nunique()

# 生成标记列:A值在至少3个分组中存在则标记为True
df['Indicator'] = df['A'].map(group_count) >= 3

逻辑说明

和初始需求核心逻辑一致,仅将判断条件从跨2个分组调整为跨3个分组。无需指定特定分组,直接统计所有B分组中每个A值的覆盖数,再做条件判断即可。


内容的提问来源于stack exchange,提问作者ah bon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 11:37:08