You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按id分组筛选type为x时category含连续两个no的组(允许NA间隔)

R语言分组筛选数据框问题

原始数据

首先定义目标数据框:

dt <- data.frame(id=c(1,1,1,1,1,1,2,2,2,2,2,3,3,3,3,3,4,4,4,4,4,5,5,5,5,6,6,6,6,6),
                 type=c('x','x','x', 'x', 'y','y','x','x','x', 'y', 'y','x','x','x','x','x','x','x','x', 'w', 'w', 'x','x','x', 'w','x','x','x', 'y', 'y'),
                 category=c(NA,'no',NA,'no',NA,'yes',NA,NA,'no',NA,'no','no',NA,'no','yes','no','no',NA,'no',NA,'yes','no',NA,NA,'no',NA,'no','no','no','yes'))

需求说明

按id分组,筛选满足以下条件的组:当type等于'x'时,category列中存在两个可被NA间隔的连续'no'值,保留符合条件组内的所有行。

尝试的错误代码

使用dplyr编写的代码未得到正确结果:

library(dplyr)
dt1 <- dt %>% 
  group_by(id) %>%  
  filter(any(category[!is.na(category)] == 'no' & 
               lag(category[!is.na(category)]) == 'no'&type=='x'))

代码问题分析

这段代码的核心问题:

  • category[!is.na(category)]会过滤NA值,但过滤后的向量与原数据type列的位置对应关系断裂,导致type=='x'的条件无法匹配到原数据中type为x的行。
  • lag(category[!is.na(category)])是对过滤后的向量取滞后值,与原分组逻辑脱节,无法准确判断type为x的行中是否存在符合要求的'no'。

正确解决方案

我们需要在每个分组内,先提取type为'x'的行的category值,去掉NA后检查是否存在连续的'no',再以此为条件筛选分组:

library(dplyr)

dt_filtered <- dt %>%
  group_by(id) %>%
  filter(
    # 提取当前组type为x的非NA category向量
    let(
      cat_x = na.omit(category[type == 'x']),
      # 检查向量中是否存在至少两个连续的no
      any(cat_x == 'no' & lag(cat_x) == 'no', na.rm = TRUE)
    )
  ) %>%
  ungroup()

# 查看结果
print(dt_filtered)

如果不想使用let,也可以分步处理:

dt_filtered <- dt %>%
  group_by(id) %>%
  mutate(
    # 存储当前组type为x的非NA category向量
    cat_x_list = list(na.omit(category[type == 'x'])),
    # 判断是否满足条件
    meets_condition = any(cat_x_list[[1]] == 'no' & lag(cat_x_list[[1]]) == 'no', na.rm = TRUE)
  ) %>%
  filter(meets_condition) %>%
  select(-cat_x_list, -meets_condition) %>%
  ungroup()

运行后得到的结果与预期完全一致。

内容的提问来源于stack exchange,提问作者Mahlet Tadesse

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 18:01:10