You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求编写正则表达式提取科学PDF参考文献的作者与标题

从科学PDF参考文献中提取作者与标题的正则方案

嘿,我之前刚好处理过类似的科学文献参考文献提取需求——这类格式虽然五花八门,但抓准规律用正则还是能搞定的。下面我分常见格式给你拆解方案,还有可直接用的代码示例:

先明确核心规律

科学文献的参考文献不管是APA、MLA还是IEEE格式,基本结构都是:作者列表 → 年份/标点 → 标题 → 出版物信息。我们的正则就是要精准定位这两块内容。

针对不同格式的正则表达式

1. APA格式(最常见)

示例条目:

Smith, J. D., & Johnson, A. B. (2020). The impact of climate change on coral reefs. Marine Biology Quarterly, 45(2), 123-145.

正则表达式:

^([\w\s,.&-]+?)\s*\((\d{4})\)\.\s*([^.]+)\.\s*.*$
  • 捕获组1:作者部分(匹配姓名、缩写点、逗号、&符号,非贪婪模式避免过度匹配)
  • 捕获组3:标题部分(匹配到下一个句号为止,因为APA格式标题后直接跟句号)

2. MLA格式

示例条目:

Smith, John Doe, and Anna Brown. "The impact of climate change on coral reefs." Marine Biology Quarterly 45.2 (2020): 123-145.

正则表达式:

^([\w\s,.&-]+?)\.\s*["“]([^"”]+)["”]\.\s*.*$
  • 捕获组1:作者部分(匹配到第一个句号为止)
  • 捕获组2:标题部分(匹配引号内的内容,兼容中文引号“”)

3. IEEE格式

示例条目:

J. D. Smith and A. B. Johnson, "The impact of climate change on coral reefs," Marine Biology Quarterly, vol. 45, no. 2, pp. 123-145, 2020.

正则表达式:

^([\w\s,.&-]+?)\,\s*["“]([^"”]+)["”]\,\s*.*$
  • 捕获组1:作者部分(匹配到第一个逗号为止)
  • 捕获组2:标题部分(匹配引号内的内容)

关键注意事项

  • 处理斜体标题:如果你的PDF解析文本里斜体是用*包裹的(比如*Marine Biology Quarterly*),可以把标题匹配部分改成\*([^*]+)\*来精准定位。
  • 避开交叉引用:一定要先定位到参考文献列表的起始位置(比如找到“References”或“参考文献”标题),只在这个区域内应用正则,避免匹配正文里的交叉引用(比如[1]、[Smith 2020])。
  • 非贪婪模式:正则里的+?是非贪婪匹配,能防止作者/标题部分不小心匹配到后面的出版物信息。

实际代码示例(Python)

我写了个小脚本,能同时匹配上面三种格式,你可以直接套用:

import re

# 替换成你的参考文献文本
reference_text = """
Smith, J. D., & Johnson, A. B. (2020). The impact of climate change on coral reefs. Marine Biology Quarterly, 45(2), 123-145.
J. D. Smith and A. B. Johnson, "The impact of climate change on coral reefs," Marine Biology Quarterly, vol. 45, no. 2, pp. 123-145, 2020.
Smith, John Doe, and Anna Brown. "The impact of climate change on coral reefs." Marine Biology Quarterly 45.2 (2020): 123-145.
"""

# 组合三种格式的正则
ref_patterns = [
    r'^([\w\s,.&-]+?)\s*\(\d{4}\)\.\s*([^.]+)\.\s*.*$',  # APA
    r'^([\w\s,.&-]+?)\.\s*["“]([^"”]+)["”]\.\s*.*$',    # MLA
    r'^([\w\s,.&-]+?)\,\s*["“]([^"”]+)["”]\,\s*.*$'     # IEEE
]

# 逐行处理参考文献
for line in reference_text.strip().split('\n'):
    line_clean = line.strip()
    for pattern in ref_patterns:
        match_result = re.match(pattern, line_clean)
        if match_result:
            authors = match_result.group(1).strip()
            title = match_result.group(2).strip()
            print(f"📝 作者: {authors}\n   标题: {title}\n---")
            break

如果你的参考文献有特殊格式(比如中文文献、非标准排版),可以给我几条示例条目,我再帮你调整正则~

内容的提问来源于stack exchange,提问作者rabbitdoubt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:13:52