You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言合并tibble嵌套列表按分组统计区间唯一值问题求解

统计分组数值区间覆盖的唯一值总数

问题场景

需要对包含多组from/to数值区间对的表格,按分组统计所有区间覆盖的不重复数值总个数,测试数据如下:

tmp <- tribble(
  ~group, ~from, ~to,
       1,     1,  10,
       1,     5,   8,
       1,    15,  20,
       2,     1,  10,
       2,     5,  10,
       2,    15,  18
)

原有尝试代码运行结果不符合预期:

tmp %>%
  group_by(group) %>%
  rowwise() %>%
  mutate(nrs = list(c(from:to))) %>%
  summarise(n_uni = length(unique(unlist(list(nrs)))))

预期正确输出为:

# A tibble: 2 × 2
   group n_uni
   <dbl> <int>
1      1    16
2      2    14

解决方案

方案1:修正原有列表列逻辑

原有代码的问题有两处:

  • group_by和rowwise顺序错误,需要先按行生成每个区间的数值序列,再按group字段分组聚合
  • 聚合时多余嵌套了一层list(),导致unlist没有正确打平同组下所有行的序列向量

修正后代码:

library(tidyverse)

tmp %>%
  rowwise() %>%
  mutate(nrs = list(from:to)) %>%
  group_by(group) %>%
  summarise(n_uni = length(unique(unlist(nrs))))

运行后直接得到预期结果。

方案2:区间合并法(适合大数值区间场景)

如果区间数值跨度大,生成全量序列会占用大量内存,可以先合并同组内重叠的区间,再直接计算总长度,无需生成所有数值:

tmp %>%
  arrange(group, from, to) %>%
  group_by(group) %>%
  summarise(
    n_uni = {
      cur_max_to <- to[1]
      total_cnt <- cur_max_to - from[1] + 1
      for (i in 2:n()) {
        if (from[i] <= cur_max_to) {
          if (to[i] > cur_max_to) {
            total_cnt <- total_cnt + to[i] - cur_max_to
            cur_max_to <- to[i]
          }
        } else {
          total_cnt <- total_cnt + to[i] - from[i] + 1
          cur_max_to <- to[i]
        }
      }
      total_cnt
    }
  )

内容的提问来源于stack exchange,提问作者Rumpl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 18:12:58