如何用正则匹配被and分隔的姓名并排除and本身
匹配文献作者字段中的加粗姓名
示例文本
title={AUTOMATIC ROCKING DEVICE},
author={Diaz, Navarro David and Gines, Rodriguez Noe},
year={2006},title={The sitting position in neurosurgery: a retrospective analysis of 488 cases},
author={Standefer, Michael and Bay, Janet W and Trusso, Russell},
journal={Neurosurgery},title={Fuel cells and their applications},
author={Kordesch, Karl and Simader, G{"u}nter and Wiley, John},
volume={117},
你之前使用的正则(?<=author={).+(?=})会匹配author={}内的全部内容,以下是几种精准提取加粗姓名的方案:
方案1:直接匹配加粗标签内的姓名
最直接的方式是定位<strong>标签,捕获其中的姓名内容:
<strong>([^<]+)</strong>
- 逻辑:
<strong>匹配起始标签,([^<]+)捕获所有非<的字符(即标签内的姓名),</strong>匹配结束标签,可直接提取所有加粗姓名。
方案2:限定在author字段内匹配
如果需要严格只提取author={}范围内的加粗姓名,可使用带范围限定的正则:
(?<=author={[^}]*<strong>)([^<]+)(?=</strong>[^}]*})
- 逻辑:
(?<=author={[^}]*<strong>):反向预查,确保当前位置处于author={之后且紧邻<strong>标签([^<]+):捕获姓名内容,直到遇到<(?=</strong>[^}]*}):正向预查,确保后续存在</strong>标签和author块的闭合大括号
方案3:拆分author块内容
若要按and拆分author块内容后提取姓名,可先获取author块内的全部内容,再拆分处理(以Python为例):
import re # 替换为你的目标文本 target_text = """title={AUTOMATIC ROCKING DEVICE},<br> author={<strong>Diaz, Navarro David</strong> and <strong>Gines, Rodriguez Noe</strong>},<br> year={2006}, title={The sitting position in neurosurgery: a retrospective analysis of 488 cases},<br> author={<strong>Standefer, Michael</strong> and <strong>Bay, Janet W</strong> and <strong>Trusso, Russell</strong>},<br> journal={Neurosurgery},<br> title={Fuel cells and their applications},<br> author={<strong>Kordesch, Karl</strong> and <strong>Simader, G{"u}nter</strong> and <strong>Wiley, John</strong>},<br> volume={117},""" # 提取所有author块内容 author_content_list = re.findall(r'(?<=author={).+(?=})', target_text) # 拆分并清理姓名 for content in author_content_list: name_items = [re.sub(r'<strong>|</strong>', '', item.strip()) for item in content.split('and')] print(name_items)
运行后会输出每个author字段对应的姓名列表。
内容的提问来源于stack exchange,提问作者Arpit Shukla
相关产品推荐
相关产品推荐

