You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言如何使用tidyverse/dplyr对含重叠列值的行进行分组

实现方案

该需求本质是识别col2中数值的连通分量:只要两个数值出现在同一行就视为连通,所有连通的数值归为同一组,最后将分组映射回原数据行即可。你可以用tidyverse搭配igraph快速实现:

1. 加载依赖包

library(tidyverse)
library(igraph)

2. 构造示例数据

df <- tibble(
  col1 = paste0("a", 1:8),
  col2 = c("2;3", "2", "3;4", "4", "2;4", "5", "5;6", "6;7")
)

3. 核心处理代码

# 给原数据加行索引,拆分col2为长格式
df_long <- df %>%
  mutate(row_id = row_number()) %>%
  separate_longer_delim(col2, delim = ";", convert = TRUE)

# 构造连通边:同一行的数值两两相连
edges <- df_long %>%
  group_by(row_id) %>%
  filter(n() >= 2) %>%
  summarise(from = first(col2), to = last(col2))

# 构建无向图,计算连通分量分组
val_group <- graph_from_data_frame(edges, directed = FALSE) %>%
  components() %>%
  `$`(membership) %>%
  enframe(name = "col2", value = "group") %>%
  mutate(col2 = as.integer(col2))

# 将分组映射回原数据
result <- df_long %>%
  left_join(val_group, by = "col2") %>%
  group_by(row_id) %>%
  summarise(group = first(group)) %>%
  left_join(df, by = "row_id") %>%
  select(col1, col2, group)

4. 输出结果

> print(result)
# A tibble: 8 × 3
  col1  col2  group
  <chr> <chr> <dbl>
1 a1    2;3       1
2 a2    2         1
3 a3    3;4       1
4 a4    4         1
5 a5    2;4       1
6 a6    5         2
7 a7    5;6       2
8 a8    6;7       2

内容的提问来源于stack exchange,提问作者Makunata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 02:54:05