在R中分组扩展data.frame,生成列内指定动物配对的方法
解决方案:生成分组内动物的无序两两配对
要实现按Area(及对应Location)分组,生成Animals列中仅存在的无序两两配对(不含反向重复,同动物多次出现时按不同行组合生成对应次数的自身配对),可以用两种简洁的方法:
方法1:dplyr + tidyr行号配对法
通过给分组内的行添加索引,生成索引配对后关联原数据获取动物名称,逻辑直观:
library(dplyr) library(tidyr) # 给原数据添加分组内的行索引 stack_df_with_id <- stack_df %>% group_by(Area, Location) %>% mutate(row_idx = row_number()) %>% ungroup() # 生成目标配对结果 result <- stack_df_with_id %>% group_by(Area, Location) %>% # 生成分组内所有行号的两两组合 expand(row_idx1 = row_idx, row_idx2 = row_idx) %>% # 过滤出无序配对(仅保留row_idx1 < row_idx2,避免反向重复) filter(row_idx1 < row_idx2) %>% # 关联获取第一个配对动物 left_join(stack_df_with_id, by = c("Area", "Location", "row_idx1" = "row_idx")) %>% # 关联获取第二个配对动物 left_join(stack_df_with_id, by = c("Area", "Location", "row_idx2" = "row_idx")) %>% # 整理最终列 select(Area, Location, Animal1 = Animals.x, Animal2 = Animals.y) %>% ungroup()
方法2:dplyr + purrr的combn法
combn函数天生用于生成向量的两两无序组合,结合分组汇总和列表展开,代码更简洁:
library(dplyr) library(purrr) library(tidyr) result <- stack_df %>% group_by(Area, Location) %>% # 对每个分组的Animals列生成所有两两无序组合,存储为列表 summarise(pairs = list(combn(Animals, 2, simplify = FALSE)), .groups = "drop") %>% # 展开列表中的每个配对 unnest(pairs) %>% # 将配对列表拆分为两列 mutate( Animal1 = map_chr(pairs, ~ .x[1]), Animal2 = map_chr(pairs, ~ .x[2]) ) %>% # 移除临时列表列 select(-pairs)
你之前尝试失效的原因
- 使用
expand(nesting(Location, Animals, Animals))时,nesting()的作用是保留原数据中多列之间的现有组合,而非同一列的两两组合,且会自动去重,无法保留行级的重复动物配对 crossing()会生成全笛卡尔积,自然会产生所有可能组合,不符合你的需求
内容的提问来源于stack exchange,提问作者Purrsia
相关产品推荐
相关产品推荐

