You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则表达式无法匹配TABLEn至google.co.in内容求助

正则匹配问题排查与解决

问题说明

需要匹配从TABLE加数字(如TABLE1)开始,到google.co.in结束的完整内容,但现有正则代码无法匹配到目标内容。

待匹配的示例文本

The paragraph continous here....................
................................................
TABLE1..
...........Text continuous...........  
......... Text continuous...........  
..........Text continuous...........  
........Text continuous...........  
........Text continuous...........  
........Text continuous...........google.co.inFrancisCo.  
The paragraph continous here....................  
................................................  

用户使用的正则代码

match = re.search(r'TABLE(\d+)((?:(?!google.co.in).)*)google.co.in', extractedtext, re.M)
if match:
    ful = match.group()
    print(f": {ful}")
    # Additional processing with 'ful' if needed
else:
    print("Not matched")

问题原因

  1. 默认情况下,正则中的.不会匹配换行符,而目标内容跨了多行,导致中间的(?:(?!google.co.in).)*无法覆盖换行部分,匹配直接中断。
  2. 你添加的re.M(多行模式)无效,这个模式仅改变^和$的匹配范围,不影响.对换行的处理逻辑。

解决办法

方法一:启用re.DOTALL模式

这个模式会让.匹配包括换行在内的所有字符,直接修改代码即可:

# 简洁写法:用非贪婪匹配替代负向预查
match = re.search(r'TABLE(\d+).*?google.co.in', extractedtext, re.DOTALL)
if match:
    ful = match.group()
    print(f": {ful}")
else:
    print("Not matched")

如果要保留原有的负向预查结构,同样加上re.DOTALL标志就行:

match = re.search(r'TABLE(\d+)((?:(?!google.co.in).)*)google.co.in', extractedtext, re.DOTALL)

方法二:用[\s\S]替代.

[\s\S]能匹配所有空白和非空白字符(等价于开启DOTALL的.),不需要额外添加标志:

match = re.search(r'TABLE(\d+)((?:(?!google.co.in)[\s\S])*)google.co.in', extractedtext)

两种方法都能正确匹配到从TABLE1到google.co.in的完整内容。

内容的提问来源于stack exchange,提问作者ssr1012

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 07:22:24