You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas逐行统计不同值数与众数 无唯一众数取最大值方法

Pandas逐行统计不同值数量与高频值实现方案

针对集成模型结果逐行统计的需求,直接使用apply方法指定按行遍历(axis=1)即可实现,不需要逐列编写计算逻辑。

核心逻辑说明

  • 不同值计数:单行序列直接调用nunique()方法,即可返回该行不重复值的总数
  • 高频值计算:先统计单行内各值的出现次数,提取出现次数最高的所有值,若存在多个值并列最高频,直接取其中的最大值即可,符合平票时取最大值的规则

完整可运行代码

import pandas as pd

# 测试数据
df = pd.DataFrame({
    "col1":[1,2,3],
    "col2":[2,2,2],
    "col3":[3,2,2]
})

# 逐行计算不同值数量
df["different_values"] = df.apply(lambda row: row.nunique(), axis=1)

# 逐行计算最高频值,平票取最大值
def cal_most_freq(row):
    count_res = row.value_counts()
    max_freq = count_res.max()
    freq_candidates = count_res[count_res == max_freq].index
    return max(freq_candidates)

df["most_frequent"] = df.apply(cal_most_freq, axis=1)

运行结果

执行代码后得到的DataFrame和预期完全一致:

col1  col2  col3  different_values  most_frequent
0     1     2     3                 3              3
1     2     2     2                 1              2
2     3     2     2                 2              2

注意:如果表格中存在不需要参与统计的列,可以先筛选出所有分类结果列再执行apply操作,例如df[["col1", "col2", "col3"]].apply(...),避免无关列干扰计算结果。

内容的提问来源于stack exchange,提问作者Marco_CH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 13:36:21