Python正则表达式无法匹配TABLEn至google.co.in内容求助
正则匹配问题排查与解决
问题说明
需要匹配从TABLE加数字(如TABLE1)开始,到google.co.in结束的完整内容,但现有正则代码无法匹配到目标内容。
待匹配的示例文本
The paragraph continous here.................... ................................................ TABLE1.. ...........Text continuous........... ......... Text continuous........... ..........Text continuous........... ........Text continuous........... ........Text continuous........... ........Text continuous...........google.co.inFrancisCo. The paragraph continous here.................... ................................................
用户使用的正则代码
match = re.search(r'TABLE(\d+)((?:(?!google.co.in).)*)google.co.in', extractedtext, re.M) if match: ful = match.group() print(f": {ful}") # Additional processing with 'ful' if needed else: print("Not matched")
问题原因
- 默认情况下,正则中的
.不会匹配换行符,而目标内容跨了多行,导致中间的(?:(?!google.co.in).)*无法覆盖换行部分,匹配直接中断。 - 你添加的
re.M(多行模式)无效,这个模式仅改变^和$的匹配范围,不影响.对换行的处理逻辑。
解决办法
方法一:启用re.DOTALL模式
这个模式会让.匹配包括换行在内的所有字符,直接修改代码即可:
# 简洁写法:用非贪婪匹配替代负向预查 match = re.search(r'TABLE(\d+).*?google.co.in', extractedtext, re.DOTALL) if match: ful = match.group() print(f": {ful}") else: print("Not matched")
如果要保留原有的负向预查结构,同样加上re.DOTALL标志就行:
match = re.search(r'TABLE(\d+)((?:(?!google.co.in).)*)google.co.in', extractedtext, re.DOTALL)
方法二:用[\s\S]替代.
[\s\S]能匹配所有空白和非空白字符(等价于开启DOTALL的.),不需要额外添加标志:
match = re.search(r'TABLE(\d+)((?:(?!google.co.in)[\s\S])*)google.co.in', extractedtext)
两种方法都能正确匹配到从TABLE1到google.co.in的完整内容。
内容的提问来源于stack exchange,提问作者ssr1012
相关产品推荐
相关产品推荐

