You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用re.split按正则分割字符串为何出现多余匹配元素

问题原因
  • re.split() 在正则表达式包含捕获组(即普通括号包裹的分组)时,会把所有捕获组匹配到的内容一并加入返回的结果列表,你写的正则里日期、小时的两个()都是捕获组,所以结果里多了10、17这类时间片段内容。
  • 最开头的空字符串是因为第一个时间戳在文本最开头,分割后第一个元素就是前缀空值。
解决方法

方法1:修改分割正则为非捕获组+过滤空值

把所有分组改成非捕获组(分组开头加?:),分割后再过滤掉空值即可:

import re

message = "Aug 10, 17:04 UTCThis is update 1.Aug 10, 15:56 UTCThis is update 2.Aug 10, 15:55 UTCThis is update 3."
# 所有括号加?:改为非捕获组
split_message = re.split(r'[a-zA-Z]{3} (?:0[1-9]|[1-2][0-9]|3[0-1]), (?:[0-1]?[0-9]|2[0-3]):[0-5][0-9] UTC', message)
# 过滤空字符串同时去掉内容末尾的点
result = [item.strip('.') for item in split_message if item]
print(result)

输出符合预期:

["This is update 1", "This is update 2", "This is update 3"]

方法2:直接用findall匹配目标内容,逻辑更简单

不用分割时间戳,直接写正则匹配你要的更新内容即可:

import re

message = "Aug 10, 17:04 UTCThis is update 1.Aug 10, 15:56 UTCThis is update 2.Aug 10, 15:55 UTCThis is update 3."
result = re.findall(r'UTC(.*?)(?=\.[A-Z][a-z]{2} |\.$)', message)
print(result)

输出和上述结果一致。

内容的提问来源于stack exchange,提问作者PritamYaduvanshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 07:00:01