You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则提取指定关键词附近的对应编号内容

Python正则实现方案

实现思路

将你提供的匹配模式预处理后拼接为正则的可选匹配前缀,匹配到前缀后直接捕获后续连续的非空白字符即可拿到目标编号。

完整代码示例

import re

# 匹配模式数组
patterns = ["reference number", "Ref:", "Invoice:", "Sales Quote No:"]
# 待匹配句子列表
sentences = [
    "My document reference number 25XPOI9876 which is entered",
    "Sales Quote No: SP21-SQ10452 entered quote number",
    "Ref:9874621kl is attached:"
]

# 构建正则表达式
pattern_re = re.compile(rf'(?:{"|".join(map(re.escape, patterns))})\s*(\S+)')

# 提取结果
result = []
for s in sentences:
    match = pattern_re.search(s)
    if match:
        result.append(match.group(1))

# 打印输出
for idx, item in enumerate(result, 1):
    print(f"{idx}. {item}")

输出结果

1. 25XPOI9876
2. SP21-SQ10452
3. 9874621kl

逻辑说明

  • 用re.escape处理每个匹配模式,避免模式中的特殊符号(如冒号)被正则识别为语法字符
  • 用|拼接所有模式为可选前缀,(?:...)标记为非捕获组,避免占用捕获分组位置
  • \s*匹配前缀和目标编号之间可能存在的0个或多个空白字符,兼容前缀后有空格/无空格的情况
  • (\S+)捕获后续连续的非空白字符,即为需要提取的编号内容

内容的提问来源于stack exchange,提问作者user17040675

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 14:15:03