如何基于指定行名向量统一排序不同规模的R数据框?
基于自定义行名向量对多个R数据框统一排序
问题背景
需要按照自定义的行名顺序,对多个行名不完全重叠、行数不同的数据框进行排序,替代默认的字母顺序排序,确保所有数据框遵循同一指定顺序排列现有行名。
可复现示例数据
# 初始化完整数据框(行名未按目标顺序排列) tabComplete <- data.frame( Group = c("E", "A", "D", "B", "C"), Values = c("order", "This", "correct", "is", "the"), row.names = c("Fifth", "First", "Fourth", "Second", "Third") ) # 提取两个行名不全的子数据框 tabPortion1 <- tabComplete[c(1, 3, 5), ] tabPortion2 <- tabComplete[c(1, 2, 4), ]
自定义排序向量
vecMyOrder <- c("First", "Second", "Third", "Fourth", "Fifth")
解决方案:基础R函数批量处理
编写通用排序函数,自动匹配每个数据框的现有行名并按自定义顺序排列:
# 定义自定义排序函数 sort_by_custom_rowname <- function(df, custom_order) { # 筛选出当前数据框中存在的行名,保留自定义顺序 matched_rows <- intersect(custom_order, rownames(df)) # 按筛选后的顺序索引数据框,drop=FALSE避免单行列转化为向量 df[matched_rows, , drop = FALSE] } # 应用函数到所有数据框 tabComplete_sorted <- sort_by_custom_rowname(tabComplete, vecMyOrder) tabPortion1_sorted <- sort_by_custom_rowname(tabPortion1, vecMyOrder) tabPortion2_sorted <- sort_by_custom_rowname(tabPortion2, vecMyOrder)
排序结果
tabComplete_sorted:
tabComplete_sorted #> Group Values #> First A This #> Second B is #> Third C the #> Fourth D correct #> Fifth E order
tabPortion1_sorted:
tabPortion1_sorted #> Group Values #> Third C the #> Fourth D correct #> Fifth E order
tabPortion2_sorted:
tabPortion2_sorted #> Group Values #> First A This #> Second B is #> Fifth E order
可选方案:tidyverse风格实现
如果习惯使用tidyverse工具链,可通过以下方式实现:
library(dplyr) library(tibble) sort_by_custom_rowname_tidy <- function(df, custom_order) { df %>% rownames_to_column("rowname") %>% # 将行名转为普通列 mutate(rowname = factor(rowname, levels = custom_order)) %>% # 按自定义顺序设置因子水平 arrange(rowname) %>% # 按因子水平排序 filter(!is.na(rowname)) %>% # 过滤掉不存在的行名 column_to_rownames("rowname") # 转回行名格式 } # 应用示例 tabComplete_tidy <- sort_by_custom_rowname_tidy(tabComplete, vecMyOrder)
内容的提问来源于stack exchange,提问作者Yacine Hajji
相关产品推荐
相关产品推荐

