如何用Python re.search获取子串首次匹配并拆分前后内容?
问题原因
你用的正则[\s\S]*(symbol[\s\S]*)里,[\s\S]*是贪婪匹配,会尽可能多的抓取字符,所以它会直接跳过前面两个"symbol",直到碰到最后一个才停下,导致捕获到的是最后一个"symbol"及其后续内容。
解决方法
要匹配第一个"symbol",得把前面的匹配改成非贪婪模式,或者直接精准匹配到第一个"symbol"的位置,下面是具体实现:
方案1:获取第一个"symbol"及后面的内容
给前面的[\s\S]*加个问号开启非贪婪,这样正则会从字符串开头找,碰到第一个"symbol"就停止匹配:
import re test_string = "I have three symbol but I want the first occurence of symbol instead of the symbol in the middle and end" regex = "[\s\S]*?(symbol[\s\S]*)" match_ = re.search(regex, test_string) if match_: result = match_.group(1) print(result) # 输出:symbol but I want the first occurence of symbol instead of the symbol in the middle and end
方案2:同时捕获第一个"symbol"的前后内容
如果需要分别拿到第一个"symbol"前面和后面的内容,可以用两个捕获组:
import re test_string = "I have three symbol but I want the first occurence of symbol instead of the symbol in the middle and end" regex = "([\s\S]*?)symbol([\s\S]*)" match_ = re.search(regex, test_string) if match_: before = match_.group(1) after = match_.group(2) print("前面的内容:", before.strip()) # 输出:前面的内容:I have three print("后面的内容:", after.strip()) # 输出:后面的内容:but I want the first occurence of symbol instead of the symbol in the middle and end
更简洁的写法
如果你的字符串可能包含换行符,可以用.配合re.DOTALL参数替代[\s\S],写法更清爽:
import re test_string = "I have three symbol but I want the first occurence of symbol instead of the symbol in the middle and end" # 获取第一个symbol及后面内容 regex = ".*?(symbol.*)" match_ = re.search(regex, test_string, re.DOTALL) if match_: print(match_.group(1)) # 获取前后内容 regex = "(.*?)symbol(.*)" match_ = re.search(regex, test_string, re.DOTALL) if match_: print("前面:", match_.group(1).strip()) print("后面:", match_.group(2).strip())
内容的提问来源于stack exchange,提问作者Ames ISU
相关产品推荐
相关产品推荐

