You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用正则表达式匹配提取文本中的arXiv DOI链接

你原有正则的问题

  1. 开头的^限制了只能匹配整段文本的起始位置,但你要的链接大概率不在段落开头,直接去掉即可
  2. org*的写法不符合需求,这里的*是正则限定符,代表匹配前面的g任意次,不是匹配org后面的任意内容
  3. 没有覆盖不带https://前缀的arxiv.org开头的链接场景

正确实现方案

用正则模式r'https?://arxiv\.org\S+|arxiv\.org\S+'即可匹配所有符合要求的链接,规则说明:

  • https? 同时兼容http和https两种前缀
  • \. 对.做转义,避免被正则识别为匹配任意字符的通配符
  • \S+ 匹配链接后续所有非空白字符,直到遇到空格、换行、括号等边界为止
  • | 是或规则,同时覆盖带协议头和不带协议头的两种链接格式

匹配完成后可以加一步简单清洗,去掉链接末尾可能带上的括号、句号等多余标点即可。

完整示例代码:

import re

# 输入文本
input_str = """Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged.

Further information can be referred to this [arXiv article]`(https://arxiv.org/abs/2109.05857).`
"""

# 匹配arXiv链接
arxiv_pattern = r'https?://arxiv\.org\S+|arxiv\.org\S+'
raw_results = re.findall(arxiv_pattern, input_str)

# 清洗末尾多余标点
clean_results = [re.sub(r'[^a-zA-Z0-9/\-:.]+$', '', res) for res in raw_results]

print(clean_results)
# 输出:['https://arxiv.org/abs/2109.05857']

内容的提问来源于stack exchange,提问作者Haris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 07:06:04