You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame的3列中去重统计值的出现次数?

多列指定值去重统计实现方案

针对你提出的需求——统计DataFrame中col1、col2、col3三列里orange、apple、a fruit的出现次数,且同一行内的重复值只计一次,完全可以通过「合并目标列+按行去重+扁平化统计」的方式实现,具体步骤如下:

1. 构造示例DataFrame(模拟你的数据)

import pandas as pd
import numpy as np

data = {
    'index id': [4, 2, 5, 3, 1],
    'col1': ['apple', 'orange', 'orange', 'apple', np.nan],
    'col2': ['orange', 'orange', 'a fruit', 'a fruit', np.nan],
    'col3': [np.nan, np.nan, np.nan, 'a fruit', np.nan],
    'othercolumns': [''] * 5
}
df = pd.DataFrame(data)

2. 核心处理逻辑

# 1. 选择需要统计的目标列
target_cols = ['col1', 'col2', 'col3']

# 2. 按行处理:过滤空值后去重,再将每行的唯一值拆分为单独行(扁平化)
flattened_unique = df[target_cols].apply(
    lambda row: row.dropna().unique(), axis=1
).explode()

# 3. 统计指定值的出现次数
target_values = ['orange', 'apple', 'a fruit']
result = flattened_unique.value_counts().loc[target_values]

3. 输出结果

执行后会得到符合要求的统计结果:

orange      3
apple       2
a fruit     2
dtype: int64

逻辑说明

  • apply(axis=1):对每一行单独处理,确保同一行的重复值只保留一个
  • dropna().unique():先过滤空值,再提取该行的唯一值,避免空值和重复值干扰统计
  • explode():将每行的唯一值数组拆分为单独的行,把二维结构扁平化,方便后续用value_counts()统计次数
  • value_counts().loc[target_values]:只保留你关心的三个值的统计结果,避免其他无关值干扰

内容的提问来源于stack exchange,提问作者user10735181

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 09:00:51