You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中使用separate_wider_regex拆分半结构化数据为多行多列

问题描述

我有一个数据框,其中的comment列包含不同研究站点多个区域的物种存在情况的半结构化信息,典型数据格式如下:

SiteComment
11: species A, species B 2a: species A, species C 4: species A
22a: species B 2b: species A 4: species C

说明:

  • 并非所有5个区域都会被列出,未列出代表该区域未发现对应物种
  • Zone 1常缺失,Zone 4几乎都会存在,匹配组长度不均

我希望使用Tidyverse工具拆分comment列,得到包含Site、Zone、Species的多行数据,最终格式如下:

SiteZoneSpecies
11species A, species B
12aspecies A, species C
14species A
22aspecies B
22bspecies A
24species C

后续可以通过pivot_wider将每个区域转为单独列。我尝试过用正则表达式"[1-4][ab]?[:]?"匹配区域,用".+?(?=[1-4][ab]?:)|(.*)"匹配物种,但未成功。以下是我的样本数据:

library(tidyr)

sample_df <- structure(list(site = 1:2, comment = c("2a: species A, species B 2b: species A, species C 3: species C, species D 4: species A, species B, species C", 
"2a: species A, species B, species D 2b: species B 3: species C 4: species C, species D, species E"
)), row.names = c(NA, -2L), class = c("tbl_df", "tbl", "data.frame"
))
解决方案

可以结合字符串拆分、行展开和正则拆分实现需求,具体代码如下:

library(tidyverse)

sample_df %>%
  # 按区域块拆分comment,正则正向预查匹配区域标识的开头
  mutate(zone_block = str_split(comment, "(?=\\d[ab]?:)")) %>%
  # 将拆分后的区域块展开为单独行
  unnest(zone_block) %>%
  # 拆分每个区域块中的Zone和Species信息
  separate_wider_regex(zone_block, patterns = list(
    Zone = "\\d[ab]?",
    colon = ": ",
    Species = ".*"
  )) %>%
  # 移除无用的colon列,整理Species列的空格
  select(-colon) %>%
  mutate(Species = str_trim(Species))

代码说明:

  1. str_split(comment, "(?=\\d[ab]?:)"):用正向预查正则,把comment拆分成以数字/字母:开头的独立区域块,比如"2a: species A..."这类片段
  2. unnest(zone_block):将每个站点的多个区域块展开为单独行,实现一行转多行
  3. separate_wider_regex:精准拆分每个区域块,提取Zone(数字加可选的ab后缀)、冒号分隔符,剩下的内容即为Species
  4. 最后清理无用列和字符串空格,得到目标格式的数据

如果需要将区域转为单独列,在上述代码后追加:

%>% pivot_wider(names_from = Zone, values_from = Species)

内容的提问来源于stack exchange,提问作者Nick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 18:15:58