如何用Python正则表达式提取论文References部分的特定格式参考文献
提取指定参考文献的Python正则方案
问题描述
给定如下论文文本:
inp = """Something at the beginning References 1. Ryff, C.D. (2014) Psychological Well-Being Revisited: Advances in the Science and Practice of Eudaimonia. 2. Deci, E.L. & Ryan, R.M. (2002) Self-determination research: reflections and future directions. 3. Acedo, F. J., & Casillas, J. C. (2005). Current paradigms in the international management field. Other References 1. Tarelli, E. (2003), “How to transfer responsibilities from expatriates to local nationals”. 2. Riusala, K. and Suutari, V. (2004), “International knowledge transfers through expatriates”. 3. Wallace, J. (2001), “The benefits of mentoring for female lawyers”. Something at the end 12. Wallace, J. (2001), “The benefits of mentoring for female lawyers”. Something else at the end"""
需实现:
- 仅提取**
References标题下**的参考文献条目 - 适配条目被换行/空格拆分的场景(如条目内容跨多行)
- 排除
Other References等其他参考文献区块,以及文本其他位置的类似格式条目 - 最终得到格式规范的条目列表:
['Ryff, C.D. (2014) Psychological Well-Being Revisited: Advances in the Science and Practice of Eudaimonia.', 'Deci, E.L. & Ryan, R.M. (2002) Self-determination research: reflections and future directions.', 'Acedo, F. J., & Casillas, J. C. (2005). Current paradigms in the international management field.']
解决方案代码
import re inp = """Something at the beginning References 1. Ryff, C.D. (2014) Psychological Well-Being Revisited: Advances in the Science and Practice of Eudaimonia. 2. Deci, E.L. & Ryan, R.M. (2002) Self-determination research: reflections and future directions. 3. Acedo, F. J., & Casillas, J. C. (2005). Current paradigms in the international management field. Other References 1. Tarelli, E. (2003), “How to transfer responsibilities from expatriates to local nationals”. 2. Riusala, K. and Suutari, V. (2004), “International knowledge transfers through expatriates”. 3. Wallace, J. (2001), “The benefits of mentoring for female lawyers”. Something at the end 12. Wallace, J. (2001), “The benefits of mentoring for female lawyers”. Something else at the end""" # 正则匹配References标题后的所有有效条目,直到下一个标题或文本结束 pattern = r'References\s*(?:\d+\.\s+(.*?))(?=\s*(?:\d+\.\s+|Other References|$))' # 启用DOTALL模式让.匹配换行符,适配跨条目换行的情况 matches = re.findall(pattern, inp, re.DOTALL) # 清理条目中的多余空白,合并换行和连续空格 cleaned_matches = [' '.join(match.split()) for match in matches] print(cleaned_matches)
代码说明
正则逻辑:
References\s*:定位目标标题,匹配后续所有空白(含换行)(?:\d+\.\s+(.*?)):非捕获组匹配条目开头的数字+点+空白,捕获组提取条目核心内容,非贪婪匹配避免过度捕获(?=\s*(?:\d+\.\s+|Other References|$)):正向预查,遇到下一个条目开头、其他参考文献标题或文本结尾时停止匹配
模式适配:
re.DOTALL让正则中的.可以匹配换行符,解决条目跨多行的问题格式清理:
' '.join(match.split())将条目中的换行、多空格统一为单个空格,输出格式规整的条目内容
内容的提问来源于stack exchange,提问作者Rostyslav Chayka
相关产品推荐
相关产品推荐

