如何在R语言的data.table每行中提取指定数量的最高值与最低值对应的列名(或列索引)
我来给你几个实用的解决方案,既能灵活提取指定数量的最高/最低值对应的列名(或索引),还能完美适配任意列数的data.table,调整提取数量也超方便:
方法1:简洁的向量化实现(推荐)
先写一个自定义函数处理单行数据,返回指定数量的最高/最低值对应的列名(或索引),再用apply按行批量处理,最后合并到原表:
library(data.table) set.seed(1) x <- data.table(a = runif(n = 5, min = -4, max = 4), b = runif(n = 5, min = -4, max = 4), c = runif(n = 5, min = -4, max = 4), d = runif(n = 5, min = -4, max = 4), e = runif(n = 5, min = -4, max = 4), f = runif(n = 5, min = -4, max = 4), g = runif(n = 5, min = -4, max = 4)) # 自定义函数:提取单行的top_n个最高值和bottom_n个最低值的列名 get_top_bottom_cols <- function(row_data, top_n = 2, bottom_n = 2) { # 对行数据降序排序,得到索引 sorted_idx <- order(row_data, decreasing = TRUE) # 提取top_n的列名并命名 top_cols <- names(row_data)[sorted_idx[1:top_n]] names(top_cols) <- paste0("max", 1:top_n) # 提取bottom_n的列名并命名 bottom_cols <- names(row_data)[sorted_idx[(length(sorted_idx)-bottom_n+1):length(sorted_idx)]] names(bottom_cols) <- paste0("min", 1:bottom_n) # 合并结果 c(top_cols, bottom_cols) } # 设置要提取的数量:比如2个最高,2个最低 top_n <- 2 bottom_n <- 2 # 按行应用函数,转成data.table后合并到原表 result_cols <- as.data.table(t(apply(x, 1, get_top_bottom_cols, top_n = top_n, bottom_n = bottom_n))) x_result <- cbind(x, result_cols) # 查看第一行结果 x_result[1]
运行后第一行的输出和你期望的完全一致:
a b c d e f g max1 max2 min1 min2 1: -2.854667 2.707073 0.3029194 -0.1426434 0.8779915 -3.010211 -2.638051 b e f a
如果只需要索引而不是列名,只需要修改函数里的names(row_data)[...]为sorted_idx[...]即可:
get_top_bottom_idx <- function(row_data, top_n = 2, bottom_n = 2) { sorted_idx <- order(row_data, decreasing = TRUE) top_idx <- sorted_idx[1:top_n] names(top_idx) <- paste0("max", 1:top_n) bottom_idx <- sorted_idx[(length(sorted_idx)-bottom_n+1):length(sorted_idx)] names(bottom_idx) <- paste0("min", 1:bottom_n) c(top_idx, bottom_idx) } # 使用方法同上 result_idx <- as.data.table(t(apply(x, 1, get_top_bottom_idx, top_n = top_n, bottom_n = bottom_n))) x_result_idx <- cbind(x, result_idx)
方法2:改进你的循环实现
如果你更习惯用循环,我把你现有的代码优化了一下,支持提取任意数量的最值:
q <- list() top_n <- 2 bottom_n <- 2 for (i in 1:nrow(x)) { row_vals <- unlist(x[i, ]) # 降序排序得到索引 sorted_idx <- order(row_vals, decreasing = TRUE) # 提取top和bottom的索引(要列名的话用names(x)[sorted_idx[...]]) top_idx <- sorted_idx[1:top_n] bottom_idx <- sorted_idx[(length(sorted_idx)-bottom_n+1):length(sorted_idx)] # 构建临时表,包含原数据和新列 temp <- cbind(x[i, ], as.data.table(t(setNames(c(top_idx, bottom_idx), c(paste0("max", 1:top_n), paste0("min", 1:bottom_n)))))) q[[i]] <- temp } # 合并所有临时表 q <- rbindlist(q)
关键优势
- 灵活调整数量:只要修改
top_n和bottom_n,就能提取1个、3个甚至更多的最值,比如top_n=3, bottom_n=3就会生成max1/max2/max3和min1/min2/min3列; - 适配任意列数:不管你的
data.table有7列还是70列,只要top_n + bottom_n不超过总列数,就能正常运行; - 两种输出形式:既可以返回列名,也可以返回列索引,满足不同需求。
内容的提问来源于stack exchange,提问作者Михаил Табаков
相关产品推荐
相关产品推荐

