如何在Pandas DataFrame的3列中去重统计值的出现次数?
多列指定值去重统计实现方案
针对你提出的需求——统计DataFrame中col1、col2、col3三列里orange、apple、a fruit的出现次数,且同一行内的重复值只计一次,完全可以通过「合并目标列+按行去重+扁平化统计」的方式实现,具体步骤如下:
1. 构造示例DataFrame(模拟你的数据)
import pandas as pd import numpy as np data = { 'index id': [4, 2, 5, 3, 1], 'col1': ['apple', 'orange', 'orange', 'apple', np.nan], 'col2': ['orange', 'orange', 'a fruit', 'a fruit', np.nan], 'col3': [np.nan, np.nan, np.nan, 'a fruit', np.nan], 'othercolumns': [''] * 5 } df = pd.DataFrame(data)
2. 核心处理逻辑
# 1. 选择需要统计的目标列 target_cols = ['col1', 'col2', 'col3'] # 2. 按行处理:过滤空值后去重,再将每行的唯一值拆分为单独行(扁平化) flattened_unique = df[target_cols].apply( lambda row: row.dropna().unique(), axis=1 ).explode() # 3. 统计指定值的出现次数 target_values = ['orange', 'apple', 'a fruit'] result = flattened_unique.value_counts().loc[target_values]
3. 输出结果
执行后会得到符合要求的统计结果:
orange 3 apple 2 a fruit 2 dtype: int64
逻辑说明
apply(axis=1):对每一行单独处理,确保同一行的重复值只保留一个dropna().unique():先过滤空值,再提取该行的唯一值,避免空值和重复值干扰统计explode():将每行的唯一值数组拆分为单独的行,把二维结构扁平化,方便后续用value_counts()统计次数value_counts().loc[target_values]:只保留你关心的三个值的统计结果,避免其他无关值干扰
内容的提问来源于stack exchange,提问作者user10735181
相关产品推荐
相关产品推荐

