You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整tidyr的extract函数正则以捕获可选industry字段?

调整正则捕获industry字段的解决方案

修改后的代码如下:

data.frame(dat) %>%
  # 清理特殊字符
  mutate(dat = gsub('["\\]\\[}{]', '', dat, perl = TRUE)) %>%
  # 按company拆分条目
  separate_rows(dat, sep = '(?<!^)(?=company)') %>%
  # 提取目标字段
  extract(dat,
          into = c("company", "user_name", "position", "duration", "industry"),
          regex = 'company:([^,]+).*?url:([^,]+).*?\\btitle:([^,]+).*?duration:([^,]+)(?:.*?industry:(.*?)(?=,|$))?')

问题原因与调整说明

原正则的核心问题是最后部分的.*?(?:industry:(.*))?逻辑缺陷:惰性匹配.*?会优先匹配最少内容,导致即使存在industry字段,可选捕获组也不会被触发,最终返回空值。

调整后的正则通过以下逻辑修复:

  • (?:.*?industry:(.*?)(?=,|$))?:
    1. .*?惰性匹配到industry:前的内容,精准定位字段位置;
    2. (.*?)(?=,|$)捕获industry:后到逗号或字符串结尾的内容,确保只获取industry的有效值;
    3. 整个组通过?标记为可选,适配没有industry的条目。

运行后会得到正确结果:

# A tibble: 3 × 5
  company user_name        position                     duration                              industry               
  <chr>   <chr>            <chr>                        <chr>                                 <chr>                  
1 Orange  https://www.xyz  CEO                          "Dec 2021 - Present 7 months "        Non-profit Organizations
2 Fig     https://www.xyz2 Business Development Manager "Feb 2019 Dec 2021 2 years 11 months" ""                     
3 Papaya  https://www.xyz3 Business Development Manager "Jan 2018 Oct 2018 10 months"         High Tech

内容的提问来源于stack exchange,提问作者Chris Ruehlemann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 00:35:18