基于分组权重保留TOP4因子:求更简洁的R实现方案
问题
我有一个tibble数据集,其中x列为字符向量(本质为因子类型),每个因子重复出现且对应数值列w。希望保留按w列总和排序的前4个因子,将其余因子替换为"other"。目前已实现需求,但想寻求更简洁的处理方法(例如使用forcats包)。
原实现代码:
library(tidyverse) df <- tibble(x=rep(letters[1:10], 10), w=seq(100)) # 提取按w总和排序的前4个因子 ff <- df |> group_by(x) |> summarise(w_tot=sum(w)) |> ungroup() |> arrange(desc(w_tot)) |> slice(1:4) |> pull(x) # 重编码数据 df_new <- df |> mutate(x=if_else(x %in% ff, x, "other"))
简洁解决方案(使用forcats包)
利用forcats包中的fct_reorder()和fct_lump_n()函数,可以一步完成需求,无需单独提取目标因子:
library(tidyverse) df <- tibble(x=rep(letters[1:10], 10), w=seq(100)) df_new <- df |> mutate( x = fct_lump_n( fct_reorder(x, w, sum, .desc = TRUE), n = 4, other_level = "other" ) ) # 查看结果 df_new #> # A tibble: 100 × 2 #> x w #> <fct> <int> #> 1 other 1 #> 2 other 2 #> 3 other 3 #> 4 other 4 #> 5 other 5 #> 6 other 6 #> 7 g 7 #> 8 h 8 #> 9 i 9 #> 10 j 10 #> # … with 90 more rows
代码解释
fct_reorder(x, w, sum, .desc = TRUE):以每个x对应的w列总和为依据,对x的因子水平进行降序重排,确保总和最高的因子排在最前面fct_lump_n(..., n = 4, other_level = "other"):保留重排后的前4个因子水平,将其余所有水平统一替换为"other"
这种方法直接在mutate()中完成所有逻辑,省去了单独提取目标因子的中间步骤,代码更紧凑且可读性强。
内容的提问来源于stack exchange,提问作者larry77
相关产品推荐
相关产品推荐

