正则拆分含日期字符串:允许首分组含数字的正则优化
问题描述
我有格式如下的字符串:"Description of stuffI\n1 31 2019\nPlace of business",原代码能正常运行,但当第一个分组包含数字时(比如"Description of 2nd stuffI\n1 31 2019\nPlace of business"或"1)Description of stuffI\n1 31 2019\nPlace of business"),原正则无法正确捕获后续日期。请问如何修改正则,允许首分组含数字且不干扰日期捕获?
原代码:
pattern = r"(\D*)(\d.*\d.*\d)\D(.*)$" example_string = "Description of stuffI\n1 31 2019\nPlace of business" match = re.match(pattern, example_string) if match: print("Group 1:", match.group(1).strip()) print("Group 2:", match.group(2).strip()) print("Group 3:", match.group(3).strip()) else: print("No match found!")
解决方案
原正则的问题在于(\D*)仅匹配非数字字符,首段含数字时会截断匹配;同时(\d.*\d.*\d)的贪婪匹配会错误覆盖内容。可以利用字符串的换行结构(首段\n日期段\n末段)来精准分割:
修改后的正则:
pattern = r"(.*?)\n(\d+ \d+ \d+)\n(.*)$"
规则说明
(.*?):非贪婪匹配任意字符,直到遇到第一个换行符,完整捕获首段(无论是否含数字)(\d+ \d+ \d+):精准匹配空格分隔的日期格式,确保只捕获中间的日期行(.*)$:捕获最后一个换行后的所有内容,即末段信息
测试验证
用带数字的首段字符串测试:
pattern = r"(.*?)\n(\d+ \d+ \d+)\n(.*)$" example_string1 = "Description of 2nd stuffI\n1 31 2019\nPlace of business" example_string2 = "1)Description of stuffI\n1 31 2019\nPlace of business" for s in [example_string1, example_string2]: match = re.match(pattern, s) if match: print("---测试字符串:", s.split('\n')[0]) print("Group 1:", match.group(1).strip()) print("Group 2:", match.group(2).strip()) print("Group 3:", match.group(3).strip()) else: print("No match found!")
输出结果:
---测试字符串: Description of 2nd stuffI Group 1: Description of 2nd stuffI Group 2: 1 31 2019 Group 3: Place of business ---测试字符串: 1)Description of stuffI Group 1: 1)Description of stuffI Group 2: 1 31 2019 Group 3: Place of business
如果日期格式有其他变体(如斜杠分隔),可调整日期匹配部分为(\d+[ /]\d+[ /]\d+)以兼容不同分隔符。
内容的提问来源于stack exchange,提问作者Lars Skaug
相关产品推荐
相关产品推荐

