You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用多个正则表达式生成包含匹配信息的元组列表

问题解决方案

不需要编写覆盖三类规则的统一正则,直接复用现有匹配逻辑即可,修改成本更低,也方便后续单独调整某一类内容的匹配规则。

具体修改步骤

  • 把三个re.finditer遍历中的打印逻辑替换为往你已初始化的result列表中追加符合要求的元组
  • 所有匹配完成后,按元组第一个元素(匹配起始索引)对列表做升序排序,保证结果按文本中出现的先后顺序排列
  • 可选优化:你当前的网址正则会把匹配内容末尾的标点(比如示例中的逗号)也识别进去,可将规则中[^\s]{2,}替换为[^\s.,!?]{2,},过滤掉末尾常见标点

修改后完整代码

import re

# 测试文本
string = """should we use regex more often? let me know at 012345678@student.eng or bbx@gmail.com. To further notice, contact Khoi at 0957507468 or accessing
https://web.de or maybe www.google.com, or Mr.Q at 0912299922."""

result = []

# 匹配邮箱
email_pattern = r'[\w\.-]+@[\w\.-]+(?:\.[\w]+)+'
for match in re.finditer(email_pattern, string):
    result.append((match.start(), match.end() - match.start(), match.group()))

# 匹配手机号
phone_pattern = r'\(?\d{3}\)?[-.\s]?\d{3}[-.\s]?\d{4}'
for match in re.finditer(phone_pattern, string):
    result.append((match.start(), match.end() - match.start(), match.group()))

# 匹配网址(已优化末尾标点过滤)
website_pattern = r'(https?:\/\/(?:www\.|(?!www))[a-zA-Z0-9][a-zA-Z0-9-]+[a-zA-Z0-9]\.[^\s.,!?]{2,}|www\.[a-zA-Z0-9][a-zA-Z0-9-]+[a-zA-Z0-9]\.[^\s.,!?]{2,}|https?:\/\/(?:www\.|(?!www))[a-zA-Z0-9]+\.[^\s.,!?]{2,}|www\.[a-zA-Z0-9]+\.[^\s.,!?]{2,})'
for match in re.finditer(website_pattern, string):
    result.append((match.start(), match.end() - match.start(), match.group()))

# 按起始索引排序
result.sort(key=lambda x: x[0])

print(result)

运行输出示例

[(47, 21, '012345678@student.eng'), (72, 13, 'bbx@gmail.com'), (122, 10, '0957507468'), (146, 14, 'https://web.de'), (170, 14, 'www.google.com'), (197, 10, '0912299922')]

内容的提问来源于stack exchange,提问作者ukiaqua

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 19:39:03