You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python正则表达式提取多对首尾字符串之间的文本

问题原因

你的正则捕获分组索引取错了。捕获分组按照左括号出现的顺序从1开始编号,你写的规则(oi|biu)(.*?)(a|e)中:

  • 第1组对应匹配到的开头字符串oi/biu
  • 第2组对应首尾之间的目标内容
  • 第3组对应匹配到的结尾字符串a/e

修复方案

最简修改

直接将提取的分组编号从1改为2,加上strip()处理掉前后多余的空格即可:

import re
string1 = 'oi sdfdsf a'
string2 = 'biu serdfd e'
pattern = '(oi|biu)(.*?)(a|e)'
substring1 = re.search(pattern, string1).group(2).strip()
substring2 = re.search(pattern, string2).group(2).strip()

print(substring1) # 输出:sdfdsf
print(substring2) # 输出:serdfd

优化版本(推荐)

如果不需要保留首尾匹配结果,可以将首尾的分组改为非捕获组(加?:前缀),避免多余分组占用索引,直接取第1组即可:

import re
string1 = 'oi sdfdsf a'
string2 = 'biu serdfd e'
pattern = '(?:oi|biu)(.*?)(?:a|e)'
substring1 = re.search(pattern, string1).group(1).strip()
substring2 = re.search(pattern, string2).group(1).strip()

多段匹配场景

如果需要一次性提取同一字符串内多段符合规则的内容,改用re.findall即可:

test_str = 'oi first a biu second e oi third a'
pattern = '(?:oi|biu)(.*?)(?:a|e)'
results = [item.strip() for item in re.findall(pattern, test_str)]
print(results) # 输出:['first', 'second', 'third']

内容的提问来源于stack exchange,提问作者philippe brenner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 17:54:01