You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python re模块语法错误:从文件提取正则捕获组的正确方法

Python正则提取指定内容的正确方法

你遇到的语法错误是因为Python不支持在原始字符串末尾直接加s(这是regex101里的DOTALL模式标识),要实现相同效果,需要用Python re模块的两种方式启用DOTALL模式(让正则中的.匹配换行符),具体如下:

方法一:使用re.DOTALL标志参数

import re

# 读取test.txt内容
with open('test.txt', 'r', encoding='utf-8') as f:
    content = f.read()

# 匹配CAR到AIRPLANE之前的内容,启用DOTALL模式
match_result = re.search(r"(CAR.*)AIRPLANE", content, flags=re.DOTALL)
if match_result:
    target_content = match_result.group(1)
    # 写入output.txt
    with open('output.txt', 'w', encoding='utf-8') as f:
        f.write(target_content)
else:
    print("未找到符合条件的内容")

方法二:正则内嵌(?s)模式修饰符

这种方式不需要额外传递flags参数,直接在正则开头加入(?s)即可启用DOTALL模式:

import re

with open('test.txt', 'r', encoding='utf-8') as f:
    content = f.read()

match_result = re.search(r"(?s)(CAR.*)AIRPLANE", content)
if match_result:
    target_content = match_result.group(1)
    with open('output.txt', 'w', encoding='utf-8') as f:
        f.write(target_content)
else:
    print("未找到符合条件的内容")

注意点

  • 必须启用DOTALL模式,否则正则中的.无法匹配换行符,会导致只匹配CAR所在行的内容,无法覆盖后续多行。
  • 加入if match_result判断是为了避免当文件中没有匹配内容时,调用group(1)抛出AttributeError异常。

内容的提问来源于stack exchange,提问作者student123456

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 21:36:12