You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于分组权重保留TOP4因子:求更简洁的R实现方案

问题

我有一个tibble数据集,其中x列为字符向量(本质为因子类型),每个因子重复出现且对应数值列w。希望保留按w列总和排序的前4个因子,将其余因子替换为"other"。目前已实现需求,但想寻求更简洁的处理方法(例如使用forcats包)。

原实现代码:

library(tidyverse)

df <- tibble(x=rep(letters[1:10], 10), w=seq(100))

# 提取按w总和排序的前4个因子
ff <- df |>
    group_by(x) |>
    summarise(w_tot=sum(w)) |>
    ungroup() |>
    arrange(desc(w_tot)) |>
    slice(1:4) |>
    pull(x)

# 重编码数据
df_new <- df |>
    mutate(x=if_else(x %in% ff, x, "other"))
简洁解决方案(使用forcats包)

利用forcats包中的fct_reorder()和fct_lump_n()函数,可以一步完成需求,无需单独提取目标因子:

library(tidyverse)

df <- tibble(x=rep(letters[1:10], 10), w=seq(100))

df_new <- df |>
  mutate(
    x = fct_lump_n(
      fct_reorder(x, w, sum, .desc = TRUE),
      n = 4,
      other_level = "other"
    )
  )

# 查看结果
df_new
#> # A tibble: 100 × 2
#>    x         w
#>    <fct> <int>
#>  1 other     1
#>  2 other     2
#>  3 other     3
#>  4 other     4
#>  5 other     5
#>  6 other     6
#>  7 g         7
#>  8 h         8
#>  9 i         9
#> 10 j        10
#> # … with 90 more rows

代码解释

  • fct_reorder(x, w, sum, .desc = TRUE):以每个x对应的w列总和为依据,对x的因子水平进行降序重排,确保总和最高的因子排在最前面
  • fct_lump_n(..., n = 4, other_level = "other"):保留重排后的前4个因子水平,将其余所有水平统一替换为"other"

这种方法直接在mutate()中完成所有逻辑,省去了单独提取目标因子的中间步骤,代码更紧凑且可读性强。

内容的提问来源于stack exchange,提问作者larry77

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 12:31:23