You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中用滑动窗口统计位置并过滤(偏好Tidyverse)

问题:基于Tidyverse实现浮动滑动窗口的位置统计

需求:给定包含positions列的dataframe,使用长度为100的浮动滑动窗口(以每个positions值作为窗口起始,窗口范围为[起始值, 起始值+100)),识别出所有区间内包含至少3个位置值的窗口,返回窗口起始值及对应计数,优先使用Tidyverse工具。

可复现示例

set.seed(0)
possible_values <- c("A", "B", "C", "D")
col1 <- sample(possible_values, 17, replace = TRUE)
positions <- c(5,17,57,101,105,123,400,578,698,707,717,735,787,811,832,853,919)
df <- data.frame(col1, positions) 

输出的df:

col1 positions
1     B         5
2     A        17
3     D        57
4     D       101
5     B       105
6     B       123
7     D       400
8     C       578
9     D       698
10    C       707
11    D       717
12    A       735
13    B       787
14    A       811
15    C       832
16    B       853
17    A       919

预期结果

col2 <- c(5,17,57,101,707,717,735,787)
count <- c(5,4,4,3,4,4,4,4)
expected <- data.frame(col2,count)

输出的expected:

col2 count
1    5     5
2   17     4
3   57     4
4  101     3
5  707     4
6  717     4
7  735     4
8  787     4

尝试过的代码(未得到预期结果)

window_size <- 100
df_window_counts <- df %>% arrange(positions) %>%
     mutate(window_start = findInterval(positions, seq(min(positions), max(positions), 
        by = window_size))) %>%
     group_by(window_start) %>%
     summarise(count = sum(!is.na(positions))) %>%
     filter(count >= 3)

输出结果:

# A tibble: 3 × 2
  window_start count
         <int> <int>
1            1     4
2            8     4
3            9     3

解决方案

这里的核心是浮动窗口——每个窗口的起始是positions中的每个值,而非固定步长的区间。可以结合purrr(Tidyverse成员)实现:

library(tidyverse)

window_size <- 100

result <- df %>%
  arrange(positions) %>%
  mutate(
    # 对每个positions值,统计在[positions, positions+window_size)范围内的元素数量
    count = map_int(positions, ~sum(between(positions, .x, .x + window_size)))
  ) %>%
  filter(count >= 3) %>%
  select(col2 = positions, count)

print(result)

输出结果:

col2 count
1    5     5
2   17     4
3   57     4
4  101     3
5  707     4
6  717     4
7  735     4
8  787     4

解释

  1. 先按positions排序,确保统计逻辑的准确性;
  2. 使用map_int遍历每个positions值作为窗口起始点;
  3. 用between函数判断每个positions是否落在当前窗口范围内,求和得到计数;
  4. 筛选计数≥3的行,重命名列后得到预期结果。

如果数据量较大追求性能,可使用slider包的slide_index_dbl函数:

library(slider)

result_slider <- df %>%
  arrange(positions) %>%
  mutate(
    count = slide_index_dbl(
      .x = positions,
      .i = positions,
      .f = ~sum(.x <= current_pos + window_size),
      .before = Inf,
      .complete = FALSE,
      current_pos = positions
    )
  ) %>%
  filter(count >= 3) %>%
  select(col2 = positions, count)

内容的提问来源于stack exchange,提问作者Mata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 20:39:50