如何调整tidyr的extract函数正则以捕获可选industry字段?
调整正则捕获industry字段的解决方案
修改后的代码如下:
data.frame(dat) %>% # 清理特殊字符 mutate(dat = gsub('["\\]\\[}{]', '', dat, perl = TRUE)) %>% # 按company拆分条目 separate_rows(dat, sep = '(?<!^)(?=company)') %>% # 提取目标字段 extract(dat, into = c("company", "user_name", "position", "duration", "industry"), regex = 'company:([^,]+).*?url:([^,]+).*?\\btitle:([^,]+).*?duration:([^,]+)(?:.*?industry:(.*?)(?=,|$))?')
问题原因与调整说明
原正则的核心问题是最后部分的.*?(?:industry:(.*))?逻辑缺陷:惰性匹配.*?会优先匹配最少内容,导致即使存在industry字段,可选捕获组也不会被触发,最终返回空值。
调整后的正则通过以下逻辑修复:
(?:.*?industry:(.*?)(?=,|$))?:.*?惰性匹配到industry:前的内容,精准定位字段位置;(.*?)(?=,|$)捕获industry:后到逗号或字符串结尾的内容,确保只获取industry的有效值;- 整个组通过
?标记为可选,适配没有industry的条目。
运行后会得到正确结果:
# A tibble: 3 × 5 company user_name position duration industry <chr> <chr> <chr> <chr> <chr> 1 Orange https://www.xyz CEO "Dec 2021 - Present 7 months " Non-profit Organizations 2 Fig https://www.xyz2 Business Development Manager "Feb 2019 Dec 2021 2 years 11 months" "" 3 Papaya https://www.xyz3 Business Development Manager "Jan 2018 Oct 2018 10 months" High Tech
内容的提问来源于stack exchange,提问作者Chris Ruehlemann
相关产品推荐
相关产品推荐

