分组内重新定义因子水平与顺序的R语言方案问询
按分组自定义子因子顺序并补全缺失水平的优化实现
需求说明
需要按预设顺序呈现数据汇总结果,核心要求:
- 针对
col1的不同分组,单独指定col2的排序规则 - 保留各
col1分组内未出现在原始数据中的col2因子水平(类似group_by(..., .drop=FALSE)的补全效果) col2的值可跨分组重复,且无全局统一排序逻辑,属于双水平因子
示例输入数据
df <- read.table( header = TRUE, sep=",", text = " col1,col2 Tunnels,Dick Tunnels,Tom Tunnels,Tom Beatles,George Beatles,Paul Beatles,Ringo Beatles,Ringo UK Artists,Gilbert " )
期望输出
col1 col2 n Beatles John 0 Beatles Paul 1 Beatles George 1 Beatles Ringo 2 UK Artists Gilbert 1 UK Artists George 0 Tunnels Tom 2 Tunnels Dick 1 Tunnels Harry 0
无效实现代码
以下代码无法满足需求:全局统一的col2因子水平无法适配不同分组的自定义排序,还会保留跨分组的多余水平,导致结果不符合预期。
col2_tunnels <- c("Tom", "Dick", "Harry") col2_beatles <- c("John", "Paul", "George", "Ringo") col2_artists <- c("Gilbert", "George") col2_order <- unique(c(col2_tunnels, col2_beatles, col2_artists)) # 无法保留重复项 col1_order <- c("Beatles", "UK Artists", "Tunnels") df %>% mutate( col1 = factor(col1, levels = col1_order), col2 = factor(col2, levels = col2_order) ) %>% group_by(col1, col2, .drop = FALSE) %>% summarise(n = n(), )
现有可行方案
通过按col1分组拆分数据,结合命名列表定义每个分组的col2因子顺序,可实现需求:
col2_fctlist <- list( Tunnels = c("Tom", "Dick", "Harry"), Beatles = c("John", "Paul", "George", "Ringo"), 'UK Artists' = c("Gilbert", "George") ) x <- lapply(col1_order, function(col1grp) df %>% filter(col1==col1grp) %>% mutate(col2 = factor(col2, levels = col2_fctlist[[col1grp]])) %>% group_by(col1, col2, .drop = FALSE) %>% summarise(n = n(), ) ) do.call(rbind, x)
更优实现方法
方法1:使用dplyr::group_modify(推荐)
利用group_modify在分组内直接处理因子水平,无需手动拆分合并数据,保持管道操作的连贯性:
col2_fctlist <- list( Beatles = c("John", "Paul", "George", "Ringo"), 'UK Artists' = c("Gilbert", "George"), Tunnels = c("Tom", "Dick", "Harry") ) col1_order <- names(col2_fctlist) # 直接复用列表名作为col1顺序 df %>% mutate(col1 = factor(col1, levels = col1_order)) %>% group_by(col1, .drop = FALSE) %>% group_modify(function(.x, .y) { # 获取当前分组对应的col2因子水平 col2_levels <- col2_fctlist[[.y$col1]] .x %>% mutate(col2 = factor(col2, levels = col2_levels)) %>% count(col2, .drop = FALSE) }) %>% ungroup()
优势:代码连贯易读,减少冗余步骤,逻辑清晰,无需手动处理数据拆分与合并。
方法2:结合tidyr::complete补全水平
适合习惯tidyr工具链的场景,直接补全分组内缺失的col2水平后统计数量:
library(tidyr) col2_fctlist <- list( Beatles = c("John", "Paul", "George", "Ringo"), 'UK Artists' = c("Gilbert", "George"), Tunnels = c("Tom", "Dick", "Harry") ) col1_order <- names(col2_fctlist) df %>% mutate(col1 = factor(col1, levels = col1_order)) %>% group_by(col1) %>% # 补全当前分组的所有col2水平 complete(col2 = col2_fctlist[[cur_group()$col1]]) %>% # 统计有效数量,缺失值对应0 summarise(n = sum(!is.na(col2)), .groups = "drop_last") %>% # 按分组内的col2顺序排序 arrange(factor(col2, levels = col2_fctlist[[cur_group()$col1]])) %>% ungroup()
内容的提问来源于stack exchange,提问作者moreQthanA
相关产品推荐
相关产品推荐

