You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多模式交替正则表达式提取数字异常问题求助

解决文本数字提取问题

你的正则之所以返回错误内容,核心问题是用了贪婪匹配的(.*),它会尽可能匹配更多字符,把后缀和后面的部分都包含进去了。下面是两种可靠的解决方法:

方法一:精准正则匹配数字

直接匹配数字部分,用\d+(匹配一个或多个数字),同时用非捕获组处理后缀(st|nd|rd|th),这样就能精准提取出目标数字:

import re

# 处理单个文本
text = 'was the 3001st most popular'
match_result = re.search(r'was the (\d+)(?:st|nd|rd|th)', text)
if match_result:
    target_num = match_result.group(1)
    print(target_num)  # 输出:3001

# 批量处理多个文本
text_list = [
    'was the 3001st most popular',
    'was the 2733rd most popular',
    'was the 3072nd most popular',
    'was the 4747th most popular'
]
# 预编译正则提升效率
pattern = re.compile(r'was the (\d+)(?:st|nd|rd|th)')
for t in text_list:
    match = pattern.search(t)
    if match:
        print(match.group(1))
  • 解释:(\d+)是捕获组,专门提取数字;(?:st|nd|rd|th)是非捕获组,只用来匹配后缀但不会被捕获,避免干扰结果。

方法二:字符串分割(适合格式固定的场景)

如果文本格式完全固定(都是was the Xxx... most popular结构),可以通过字符串分割提取:

text = 'was the 4747th most popular'
# 按空格分割后取第三个元素(索引为2)
num_with_suffix = text.split()[2]
# 去掉后缀(取除最后两位之外的内容)
target_num = num_with_suffix[:-2]
print(target_num)  # 输出:4747
  • 注意:这种方法依赖固定格式,如果后缀长度变化(比如出现特殊情况),可能失效,正则方法更通用。

内容的提问来源于stack exchange,提问作者Sid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 09:20:30