You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中按分组统计值大于等于指定序列各值的行数的实现方法

解决方案

Tidyverse 实现

你之前编写的循环没有存储每次计算的结果,因此无法得到结构化输出,且多层循环在数据量较大时效率很低,使用tidyverse的向量化操作可以快速实现需求:

library(tidyverse)

# 定义待校验的阈值序列
thresholds <- seq(300, 1000, 100)

# 核心计算逻辑
result <- df %>%
  group_by(Company) %>%
  reframe(
    Greater.or.equal.to = thresholds,
    Count = purrr::map_int(thresholds, ~ sum(Sales >= .x))
  )

执行后得到的result和你给出的预期结果完全一致,可通过all.equal(result, 你定义的预期result)验证。

如果使用的dplyr版本低于1.1.0(不支持reframe函数),可改用如下写法:

result <- df %>%
  group_by(Company) %>%
  summarise(
    Greater.or.equal.to = list(thresholds),
    Count = list(purrr::map_int(thresholds, ~ sum(Sales >= .x)))
  ) %>%
  unnest(c(Greater.or.equal.to, Count))

通用封装函数

如果需要频繁在不同列上执行该类统计,可以封装为通用函数:

count_ge_by_group <- function(data, group_col, value_col, thresholds) {
  group_col <- dplyr::enquo(group_col)
  value_col <- dplyr::enquo(value_col)
  
  data %>%
    dplyr::group_by(!!group_col) %>%
    dplyr::reframe(
      Greater.or.equal.to = thresholds,
      Count = purrr::map_int(thresholds, ~ sum(!!value_col >= .x))
    )
}

# 调用示例
result <- count_ge_by_group(df, Company, Sales, seq(300, 1000, 100))

大数据量优化方案

如果数据集行数多、阈值序列长,可以先对分组内的数值排序,通过findInterval批量计算,避免重复遍历数据:

result <- df %>%
  dplyr::group_by(Company) %>%
  dplyr::summarise(sorted_sales = list(sort(Sales, decreasing = TRUE))) %>%
  tidyr::crossing(Greater.or.equal.to = thresholds) %>%
  dplyr::rowwise() %>%
  dplyr::mutate(Count = findInterval(Greater.or.equal.to - 1e-9, sorted_sales, rightmost.closed = TRUE)) %>%
  dplyr::select(-sorted_sales)

内容的提问来源于stack exchange,提问作者Economist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 21:06:03