R语言用str_extract正则提取参考文献指定内容失败如何解决
R 语言参考文献标题提取正则修改方案
问题原因
你原有正则中的.*属于贪婪匹配规则,会尽可能匹配最长的符合条件的字符串,因此会一直匹配到整条参考文献末尾的最后一个句号之前,不会停在第一个句号的位置。同时原有后行断言没有匹配年份括号后的点和空格,会导致提取结果开头多出空格。
修正方案
使用排除字符集代替贪婪匹配,保证匹配到第一个句号就停止,同时修正后行断言的匹配范围,直接匹配年份右括号+点+空格的固定格式,避免提取结果开头出现冗余空格。
修正后的正则表达式有两种等价写法,推荐使用第一种性能更稳定:
- 排除字符集写法:
(?<=\\)\\.\\s)[^.]*\\. - 非贪婪匹配写法:
(?<=\\)\\.\\s).*?\\.
测试代码示例
library(stringr) # 构造示例参考文献向量 df1 <- data.frame( References = c( "Acemoglu, D., & Robinson, J. A. (2012). Why nations fail: The origins of power, prosperity, and poverty. Crown Books.", "Adam, S., & Kriesi, H. (2007). The network approach. In Sabatier, P. A. (ed.), Theories of the policy process (2nd Ed.). Cambridge, MA: Westview Press.", "Adams-Webber, J. R. (1969). Cognitive complexity and sociality. British Journal of Social and Clinical Psychology, 8, 211-216." ) ) # 执行提取 result <- str_extract(df1$References, pattern = "(?<=\\)\\.\\s)[^.]*\\.") print(result)
输出结果
[1] "Why nations fail: The origins of power, prosperity, and poverty." [2] "The network approach." [3] "Cognitive complexity and sociality."
完全符合预期提取要求。
内容的提问来源于stack exchange,提问作者Tiangeng Lu
相关产品推荐
相关产品推荐

