请求编写正则表达式提取科学PDF参考文献的作者与标题
嘿,我之前刚好处理过类似的科学文献参考文献提取需求——这类格式虽然五花八门,但抓准规律用正则还是能搞定的。下面我分常见格式给你拆解方案,还有可直接用的代码示例:
先明确核心规律
科学文献的参考文献不管是APA、MLA还是IEEE格式,基本结构都是:作者列表 → 年份/标点 → 标题 → 出版物信息。我们的正则就是要精准定位这两块内容。
针对不同格式的正则表达式
1. APA格式(最常见)
示例条目:
Smith, J. D., & Johnson, A. B. (2020). The impact of climate change on coral reefs. Marine Biology Quarterly, 45(2), 123-145.
正则表达式:
^([\w\s,.&-]+?)\s*\((\d{4})\)\.\s*([^.]+)\.\s*.*$
- 捕获组1:作者部分(匹配姓名、缩写点、逗号、&符号,非贪婪模式避免过度匹配)
- 捕获组3:标题部分(匹配到下一个句号为止,因为APA格式标题后直接跟句号)
2. MLA格式
示例条目:
Smith, John Doe, and Anna Brown. "The impact of climate change on coral reefs." Marine Biology Quarterly 45.2 (2020): 123-145.
正则表达式:
^([\w\s,.&-]+?)\.\s*["“]([^"”]+)["”]\.\s*.*$
- 捕获组1:作者部分(匹配到第一个句号为止)
- 捕获组2:标题部分(匹配引号内的内容,兼容中文引号“”)
3. IEEE格式
示例条目:
J. D. Smith and A. B. Johnson, "The impact of climate change on coral reefs," Marine Biology Quarterly, vol. 45, no. 2, pp. 123-145, 2020.
正则表达式:
^([\w\s,.&-]+?)\,\s*["“]([^"”]+)["”]\,\s*.*$
- 捕获组1:作者部分(匹配到第一个逗号为止)
- 捕获组2:标题部分(匹配引号内的内容)
关键注意事项
- 处理斜体标题:如果你的PDF解析文本里斜体是用
*包裹的(比如*Marine Biology Quarterly*),可以把标题匹配部分改成\*([^*]+)\*来精准定位。 - 避开交叉引用:一定要先定位到参考文献列表的起始位置(比如找到“References”或“参考文献”标题),只在这个区域内应用正则,避免匹配正文里的交叉引用(比如[1]、[Smith 2020])。
- 非贪婪模式:正则里的
+?是非贪婪匹配,能防止作者/标题部分不小心匹配到后面的出版物信息。
实际代码示例(Python)
我写了个小脚本,能同时匹配上面三种格式,你可以直接套用:
import re # 替换成你的参考文献文本 reference_text = """ Smith, J. D., & Johnson, A. B. (2020). The impact of climate change on coral reefs. Marine Biology Quarterly, 45(2), 123-145. J. D. Smith and A. B. Johnson, "The impact of climate change on coral reefs," Marine Biology Quarterly, vol. 45, no. 2, pp. 123-145, 2020. Smith, John Doe, and Anna Brown. "The impact of climate change on coral reefs." Marine Biology Quarterly 45.2 (2020): 123-145. """ # 组合三种格式的正则 ref_patterns = [ r'^([\w\s,.&-]+?)\s*\(\d{4}\)\.\s*([^.]+)\.\s*.*$', # APA r'^([\w\s,.&-]+?)\.\s*["“]([^"”]+)["”]\.\s*.*$', # MLA r'^([\w\s,.&-]+?)\,\s*["“]([^"”]+)["”]\,\s*.*$' # IEEE ] # 逐行处理参考文献 for line in reference_text.strip().split('\n'): line_clean = line.strip() for pattern in ref_patterns: match_result = re.match(pattern, line_clean) if match_result: authors = match_result.group(1).strip() title = match_result.group(2).strip() print(f"📝 作者: {authors}\n 标题: {title}\n---") break
如果你的参考文献有特殊格式(比如中文文献、非标准排版),可以给我几条示例条目,我再帮你调整正则~
内容的提问来源于stack exchange,提问作者rabbitdoubt

