在R中使用separate_wider_regex拆分半结构化数据为多行多列
问题描述
我有一个数据框,其中的comment列包含不同研究站点多个区域的物种存在情况的半结构化信息,典型数据格式如下:
| Site | Comment |
|---|---|
| 1 | 1: species A, species B 2a: species A, species C 4: species A |
| 2 | 2a: species B 2b: species A 4: species C |
说明:
- 并非所有5个区域都会被列出,未列出代表该区域未发现对应物种
- Zone 1常缺失,Zone 4几乎都会存在,匹配组长度不均
我希望使用Tidyverse工具拆分comment列,得到包含Site、Zone、Species的多行数据,最终格式如下:
| Site | Zone | Species |
|---|---|---|
| 1 | 1 | species A, species B |
| 1 | 2a | species A, species C |
| 1 | 4 | species A |
| 2 | 2a | species B |
| 2 | 2b | species A |
| 2 | 4 | species C |
后续可以通过pivot_wider将每个区域转为单独列。我尝试过用正则表达式"[1-4][ab]?[:]?"匹配区域,用".+?(?=[1-4][ab]?:)|(.*)"匹配物种,但未成功。以下是我的样本数据:
library(tidyr) sample_df <- structure(list(site = 1:2, comment = c("2a: species A, species B 2b: species A, species C 3: species C, species D 4: species A, species B, species C", "2a: species A, species B, species D 2b: species B 3: species C 4: species C, species D, species E" )), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame" ))
解决方案
可以结合字符串拆分、行展开和正则拆分实现需求,具体代码如下:
library(tidyverse) sample_df %>% # 按区域块拆分comment,正则正向预查匹配区域标识的开头 mutate(zone_block = str_split(comment, "(?=\\d[ab]?:)")) %>% # 将拆分后的区域块展开为单独行 unnest(zone_block) %>% # 拆分每个区域块中的Zone和Species信息 separate_wider_regex(zone_block, patterns = list( Zone = "\\d[ab]?", colon = ": ", Species = ".*" )) %>% # 移除无用的colon列,整理Species列的空格 select(-colon) %>% mutate(Species = str_trim(Species))
代码说明:
str_split(comment, "(?=\\d[ab]?:)"):用正向预查正则,把comment拆分成以数字/字母:开头的独立区域块,比如"2a: species A..."这类片段unnest(zone_block):将每个站点的多个区域块展开为单独行,实现一行转多行separate_wider_regex:精准拆分每个区域块,提取Zone(数字加可选的ab后缀)、冒号分隔符,剩下的内容即为Species- 最后清理无用列和字符串空格,得到目标格式的数据
如果需要将区域转为单独列,在上述代码后追加:
%>% pivot_wider(names_from = Zone, values_from = Species)
内容的提问来源于stack exchange,提问作者Nick
相关产品推荐
相关产品推荐

