使用dplyr拆分列表列并保留ID变量的实现方法
用dplyr拆分列表列并保留对应ID
原始数据集
df <- structure(list(spp = list(c("Other species (please list species name; i.e. Tarpon)", "Other species (please list species name; i.e. Amberjack)"), c("Red Drum (a.k.a. Redfish or Red)", "Other species (please list species name; i.e. Tarpon)", "Other species (please list species name; i.e. Amberjack)" ), c("Red Drum (a.k.a. Redfish or Red)", "Other species (please list species name; i.e. Tarpon)" )), ID = c("1", "2", "3")), row.names = 3:5, class = "data.frame")
需求说明
需要将spp列(每个单元格为物种名称组成的列表)拆分为每行对应单个物种的格式,同时保留对应的ID变量,最终输出格式如下:
> harv spp ID 1 Other species (please list species name; i.e. Tarpon) 1 2 Other species (please list species name; i.e. Amberjack) 1 3 Red Drum (a.k.a. Redfish or Red) 2 4 Other species (please list species name; i.e. Tarpon) 2 5 Other species (please list species name; i.e. Amberjack) 2 6 Red Drum (a.k.a. Redfish or Red) 3 7 Other species (please list species name; i.e. Tarpon) 3
解决方案
使用dplyr配合tidyr的unnest()函数即可实现,代码如下:
# 加载依赖包 library(dplyr) library(tidyr) # 拆分列表列并整理结果 harv <- df %>% unnest(cols = spp) %>% # 展开spp列的列表元素,自动保留对应ID arrange(ID) # 按ID排序,匹配期望输出的顺序(可省略)
说明
unnest(cols = spp)是核心操作:它会把spp列里的每个列表元素拆成单独行,同时自动关联对应的ID值;arrange(ID)用于调整行的排序,和目标输出顺序一致,如果不需要特定顺序可以直接省略这一步。
运行上述代码后,得到的harv数据框即为目标格式。
内容的提问来源于stack exchange,提问作者David Smith
相关产品推荐
相关产品推荐

