You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式匹配多字符并拆分:提取文本三类信息的技术问题

正则文本提取问题解决方案

需求说明

需要从目标文本中提取三类信息并整合为元组:

  • 提取**包裹的姓名文本
  • 提取#x格式中的数字部分
  • 提取当前条目到下一个姓名前的所有剩余内容(含最后一段无姓名的内容)

现有问题

现有正则仅能拆分匹配姓名和编号,返回结果包含大量空元素,无法提取剩余内容,也不能将单条目的三类信息整合为一个元组。

示例文本

note = "**Jane Greiz** `#1`: Should be open here .
**Thomas Fitzpatrick** `#90`: Anim: Can we start the movement.
**Anthony Smith** `#91`: Her left shoulder.
https://google.com"

现有代码

import re
all_pattern = r"\*\*(.+?)\*\*|\`#(.+?)\`:"
result = re.findall(all_pattern, note)

现有输出

[('Jane Greiz', ''), ('', '1'), ('Thomas Fitzpatrick', ''), ('', '90'), ('Anthony Smith', ''), ('', '91')]

修正方案

使用包含三个捕获组的正则表达式,匹配完整条目结构,同时兼容最后一段无姓名的内容:

修正后代码

import re

note = "**Jane Greiz** `#1`: Should be open here .
**Thomas Fitzpatrick** `#90`: Anim: Can we start the movement.
**Anthony Smith** `#91`: Her left shoulder.
https://google.com"

# 匹配带姓名的条目 + 最后一段无姓名的收尾内容
pattern = r"\*\*(.+?)\*\*\s*`#(\d+)`:\s*(.*?)(?=\n\*\*|\Z)"
matches = re.findall(pattern, note, re.DOTALL)

# 输出结果
print(matches)

输出结果

[
    ('Jane Greiz', '1', 'Should be open here .'),
    ('Thomas Fitzpatrick', '90', 'Anim: Can we start the movement.'),
    ('Anthony Smith', '91', 'Her left shoulder.\nhttps://google.com')
]

正则规则解释

  • \*\*(.+?)\*\*: 捕获**包裹的姓名(非贪婪匹配避免跨条目)
  • \s*#(\d+):\s*: 捕获#后的数字,自动忽略前后空白字符
  • (.*?): 捕获当前条目的剩余内容(非贪婪匹配)
  • (?=\n\*\*|\Z): 正向预查终止条件,匹配到下一个姓名开头或文本结尾为止
  • re.DOTALL: 让.匹配换行符,确保剩余内容包含换行的最后一段

内容的提问来源于stack exchange,提问作者Zak44

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 06:10:36