You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr移除DataFrame中含连续重复值的分组?

解决方案:用dplyr移除包含连续重复值的分组

没问题,用dplyr就能轻松搞定这个需求!咱们一步步来拆解实现思路:

思路拆解

  • 按Variable 1分组,对每个组内的Variable 2做连续重复值检查
  • 标记存在连续重复的组,最终筛选掉这些组的所有行

代码实现(dplyr版本)

首先构造示例数据,然后执行核心处理逻辑:

library(dplyr)

# 构造你的示例数据
df <- tibble(
  `Variable 1` = c(1, 1, 2, 2, 2, 3, 3, 3),
  `Variable 2` = c("a", "b", "a", "a", "b", "a", "c", "a")
)

# 核心处理逻辑
cleaned_df <- df %>%
  group_by(`Variable 1`) %>%
  # 标记当前组是否存在连续重复的Variable2值
  mutate(has_consecutive_duplicates = any(`Variable 2` == lag(`Variable 2`, default = ""))) %>%
  # 筛选出没有连续重复的组
  filter(!has_consecutive_duplicates) %>%
  # 移除辅助标记列
  select(-has_consecutive_duplicates) %>%
  ungroup()

# 查看结果
cleaned_df

代码解释

  • lag(Variable 2, default = ""):获取当前行的前一行Variable 2的值,第一行默认用空字符串(避免第一行被误判为重复)
  • any(...):只要组内有任意一行和前一行的值相同,就标记该组存在连续重复
  • filter(!has_consecutive_duplicates):只保留没有连续重复的组的所有行

备选方案(data.table版本)

如果习惯用data.table,可以这样实现:

library(data.table)

setDT(df)
cleaned_df_dt <- df[, if (!any(`Variable 2` == shift(`Variable 2`, fill = ""))) .SD, by = `Variable 1`]

最终输出结果

执行代码后会得到你想要的DataFrame:

# A tibble: 5 × 2
  `Variable 1` `Variable 2`
         <dbl> <chr>       
1            1 a           
2            1 b           
3            3 a           
4            3 c           
5            3 a           

内容的提问来源于stack exchange,提问作者SSMFTKIB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:05:45