You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tidyverse按组聚合列表列:问题原因及解决方案

按组合并列表列的tidyverse解决方案

原始数据与需求

原始数据:

library(tidyverse)
df <- tibble(group = c("a", "b", "b"), val = list(1:3, 4:6, 7:12))
# 输出结构:
## A tibble: 3 × 2
#  group val      
#  <chr> <list>   
#1 a     <int [3]>
#2 b     <int [3]>
#3 b     <int [6]>

需求:按group列合并val列中的列表条目,预期输出:

df_out <- tibble(group = c("a", "b"), val = list(1:3, 4:12))

补充说明:val中的元素不一定是有序序列,比如某组的val为c(1, 10, 2)和c(4, 7, 7),合并后需得到c(1, 10, 2, 4, 7, 7)。

尝试的错误代码及问题

尝试执行以下代码未得到预期结果:

df %>% group_by(group) %>% summarise(val = map(val, c), .groups = "drop")

返回结果仍为3行,且触发警告:

# A tibble: 3 × 2
  group val      
  <chr> <list>   
1 a     <int [3]>
2 b     <int [3]>
3 b     <int [6]>
Warning message:
Returning more (or less) than 1 row per `summarise()` group was deprecated in dplyr 1.1.0.
ℹ Please use `reframe()` instead.
ℹ When switching from `summarise()` to `reframe()`, remember that `reframe()` always returns an ungrouped data frame and adjust
  accordingly.
Call `lifecycle::last_lifecycle_warnings()` to see where this warning was generated. 

疑问:为何summarise()会返回“每组多行”?

原因解释

map(val, c)的作用是对val列中的每个单独列表元素执行c()操作(这里相当于无意义的包装,单个列表元素用c()处理后还是原样),所以每组有多少行,就会返回多少个列表元素,导致summarise()每组输出多行,既没有合并列表,还触发了dplyr的版本警告。

解决方案

使用group_by() + summarise()的两步方案,通过reduce()将组内所有列表合并为一个向量,再包装成列表存入val列:

df %>% 
  group_by(group) %>% 
  summarise(val = list(reduce(val, c)), .groups = "drop")

也可以用解包语法!!!直接合并组内所有列表:

df %>% 
  group_by(group) %>% 
  summarise(val = list(c(!!!val)), .groups = "drop")

验证补充示例:

df <- tibble(group = c("a", "b", "b"), val = list(1:3, c(1, 10, 2), c(4, 7, 7)))
# 执行上述代码后得到:
## A tibble: 2 × 2
#  group val          
#  <chr> <list>       
#1 a     <int [3]>    
#2 b     <dbl [6]>    
#  其中b组的val为c(1, 10, 2, 4, 7, 7),符合预期

内容的提问来源于stack exchange,提问作者FKneip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 22:12:28