You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中提取数据框每行最大值对应的列名

在R中提取每行最大值对应的列名的高效实现方法

输入数据

假设我们有如下R数据框:

df <- structure(list(n1 = c(10L, 9L, 6L, 8L, 3L, 14L, 13L, 10L, 4L, 12L, 11L, 14L, 1L, 4L, 6L, 2L, 10L, 2L),
                     n2 = c(5L, 10L, 7L, 11L, 9L, 2L, 6L, 11L, 8L, 9L, 12L, 7L, 11L, 5L, 9L, 6L, 13L, 10L),
                     n3 = c(13L, 11L, 2L, 13L, 7L, 3L, 5L, 3L, 11L, 2L, 6L, 5L, 13L, 10L, 13L, 11L, 12L, 4L),
                     n4 = c(3L, 1L, 14L, 1L, 8L, 10L, 9L, 5L, 7L, 11L, 7L, 11L, 2L, 8L, 14L, 5L, 9L, 7L),
                     n5 = c(2L, 12L, 11L, 12L, 12L, 6L, 7L, 12L, 10L, 1L, 8L, 8L, 14L, 13L, 12L, 9L, 5L, 13L),
                     n6 = c(8L, 7L, 10L, 3L, 5L, 12L, 8L, 7L, 3L, 14L, 14L, 12L, 6L, 7L, 2L, 13L, 14L, 5L),
                     n7 = c(1L, 2L, 12L, 4L, 10L, 5L, 10L, 14L, 13L, 6L, 10L, 4L, 9L, 11L, 3L, 8L, 11L, 9L),
                     n8 = c(9L, 14L, 5L, 2L, 2L, 13L, 11L, 4L, 14L, 10L, 13L, 3L, 7L, 3L, 7L, 10L, 6L, 8L),
                     n9 = c(4L, 5L, 13L, 10L, 1L, 8L, 2L, 9L, 12L, 8L, 3L, 1L, 4L, 2L, 10L, 3L, 1L, 12L),
                     n10 = c(7L, 8L, 8L, 9L, 6L, 9L, 3L, 8L, 2L, 3L, 4L, 9L, 12L, 14L, 8L, 4L, 3L, 11L),
                     n11 = c(6L, 13L, 1L, 14L, 14L, 7L, 4L, 2L, 5L, 7L, 9L, 6L, 3L, 9L, 1L, 14L, 4L, 14L),
                     n12 = c(12L, 4L, 9L, 7L, 11L, 4L, 1L, 13L, 9L, 5L, 2L, 13L, 5L, 12L, 5L, 7L, 8L, 3L),
                     n13 = c(11L, 3L, 3L, 5L, 4L, 1L, 12L, 1L, 1L, 4L, 1L, 10L, 8L, 6L, 4L, 1L, 7L, 1L),
                     n14 = c(14L, 6L, 4L, 6L, 13L, 11L, 14L, 6L, 6L, 13L, 5L, 2L, 10L, 1L, 11L, 12L, 2L, 6L)),
                class = "data.frame", row.names = c("3557", "3558", "3559", "3560", "3561", "3562", "3563", "3564", "3565", "3566", "3567", "3568", "3569", "3570", "3571", "3572", "3573", "3574"))

需求

提取每行中数值最大的元素对应的列名,得到如下结果:

#         Choice
# 3557    n14
# 3558    n8
# 3559    n4
# 3560    n11
# 3561    n11
# 3562    n1
# 3563    n14
# 3564    n7
# 3565    n8
# 3566    n6
# 3567    n6
# 3568    n1
# 3569    n5
# 3570    n10
# 3571    n4
# 3572    n11
# 3573    n6
# 3574    n11

高效实现方法

1. 基础R:apply 函数

适合中等规模数据,写法直观:

df$Choice <- colnames(df)[apply(df, 1, which.max)]
  • apply(df, 1, which.max) 逐行找出最大值所在的列索引
  • 通过colnames(df)匹配索引对应的列名

2. 向量化操作(性能最优,适合大数据)

避免循环开销,完全向量化处理:

df$Choice <- colnames(df)[max.col(df, ties.method = "first")]
  • max.col(df, ties.method = "first") 返回每行最大值的列索引,ties.method可指定多最大值场景的处理规则(如"last"取最后一个匹配列)
  • 该方法性能远优于apply,是大规模数据的首选

3. tidyverse/dplyr 方法

适配tidyverse生态的写法:

library(dplyr)

df_result <- df %>%
  rowwise() %>%
  mutate(Choice = colnames(df)[which.max(c_across(everything()))]) %>%
  ungroup() %>%
  select(Choice)
  • c_across(everything()) 提取当前行的所有数值
  • rowwise() 确保计算逻辑逐行生效

4. data.table 方法(超大数据场景首选)

处理超大规模数据时,data.table的速度优势明显:

library(data.table)

dt <- as.data.table(df, keep.rownames = TRUE)
dt[, Choice := colnames(df)[max.col(.SD, ties.method = "first")], by = rn]
dt_result <- dt[, .(Choice), keyby = rn]
  • .SD代表当前分组(此处为每行)的所有列
  • data.table的分组计算效率极高,适合处理百万级以上行数的数据

内容的提问来源于stack exchange,提问作者Apook

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 23:19:00