You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:计算多同维度数据框对应位置元素的众数

生成多同维度分类数据框对应位置的众数数据框

问题描述

我有一个包含多个同维度数据框的列表,所有数据框列名完全一致,且均为带因子水平的分类类型。多数因子水平相同,但部分因子水平可能未在所有数据框中出现。需要生成一个新的数据框,其中每个元素是所有数据框对应位置元素的众数;若众数存在并列情况,任选其一即可。

示例数据

数据框df1、df2、df3、df4存储在列表df <- list(df1,df2,df3,df4)中:

df1
  col1 col2  col3
1    e    6 FALSE
2    b    1 FALSE
3    d    1  TRUE
4    e    2  TRUE
5    d    5  TRUE
df2
  col1 col2  col3
1    b    2 FALSE
2    f    0  TRUE
3    e    5 FALSE
4    e    1  TRUE
5    b    1 FALSE
df3
  col1 col2  col3
1    r    0  TRUE
2    d    1  TRUE
3    d    0 FALSE
4    b    5  TRUE
5    e    2  TRUE
df4
  col1 col2  col3
1    d    5  TRUE
2    e    1  TRUE
3    b    2 FALSE
4    d    0  TRUE
5    e    5  TRUE

期望结果

col1 col2  col3
1    e    6 FALSE
2    b    1  TRUE
3    d    1  FALSE
4    e    2  TRUE
5    e    5  TRUE

可复现代码

df1 = data.frame(col1 = c("e", "b", "d", "e", "d") ,
                 col2 = c(6, 1, 1, 2, 5),
                 col3= c(FALSE, FALSE, TRUE,TRUE, TRUE))
df1 <- data.frame(lapply(df1,factor))

df2 = data.frame(col1 = c("b", "f", "e", "e", "b") ,
                 col2 = c(2, 0, 5, 1, 1),
                 col3= c(FALSE, TRUE, FALSE,TRUE, FALSE))
df2 <- data.frame(lapply(df2,factor))

df3 = data.frame(col1 = c("r", "d", "d", "b", "e") ,
                 col2 = c(0, 1, 0, 5, 2),
                 col3= c(TRUE, TRUE, FALSE,TRUE, TRUE))
df3 <- data.frame(lapply(df3,factor))

df4 = data.frame(col1 = c("d", "e", "b", "d", "e") ,
                 col2 = c(5, 1, 2, 0, 5),
                 col3= c(TRUE, TRUE, FALSE,TRUE, TRUE))
df4 <- data.frame(lapply(df4,factor))

df <- list(df1,df2,df3,df4)   

解决方案

思路

  1. 定义众数函数:处理因子类型时先转为字符统计频率,避免不同数据框因子水平差异导致的统计错误;众数并列时返回第一个出现的结果,最后转换回原因子类型以保留分类属性。
  2. 将数据框列表转换为三维数组,方便按行-列位置提取所有数据框的对应元素。
  3. 遍历每个行-列位置计算众数,最终重构为目标数据框。

代码实现

# 定义众数计算函数
get_mode <- function(x) {
  # 转换为字符统一统计基准,避免因子水平差异影响
  x_char <- as.character(x)
  freq_table <- table(x_char)
  max_freq <- max(freq_table)
  # 取第一个出现的众数
  mode_val <- names(freq_table[freq_table == max_freq])[1]
  # 转换回原因子类型
  factor(mode_val, levels = levels(x))
}

# 将数据框列表转为三维数组(行×列×数据框数量)
df_array <- array(unlist(df), dim = c(nrow(df[[1]]), ncol(df[[1]]), length(df)))

# 对每个行-列位置计算众数
result_matrix <- apply(df_array, c(1, 2), get_mode)

# 转换为数据框并设置列名
result_df <- as.data.frame(result_matrix)
colnames(result_df) <- colnames(df[[1]])

# 输出结果
result_df

运行上述代码后,即可得到符合要求的众数数据框。

内容的提问来源于stack exchange,提问作者Kevin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 05:27:21