You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:修复Python正则表达式跨行捕获BEGIN与PAUSE间内容的问题

Fixing Multiline Regex Capture for BEGIN to PAUSE

Got it, let's sort out your regex issue step by step. The problem with your current code comes from two key mistakes: wrong flag usage and incorrect parameter order in re.findall().

What Was Wrong

  1. Incorrect parameter order: You passed re.MULTILINE as the second argument to re.findall(), but the correct order is re.findall(pattern, string, flags). So your flag wasn't being applied at all!
  2. Wrong flag choice: re.MULTILINE only changes how ^ and $ work (matching start/end of each line instead of the whole string). It doesn't make the . character match newline characters, which is what you need for multiline capture.

The Fixed Solution

Use the re.DOTALL (or its alias re.S) flag instead—this makes . match all characters including newlines. Here's the corrected code:

import re

# Your multiline input example
input_text = """START
BEGIN
PROGRAM STARTED
PROGRAM IN PROGRESS
PAUSE
STOP"""

# Correct regex with proper flags and parameter order
matches = re.findall('(?<=BEGIN)(.*?)(?=PAUSE)', input_text, re.DOTALL)

# Optional: Clean up whitespace around the matched content
clean_matches = [match.strip() for match in matches]
print(clean_matches)
# Output: ['PROGRAM STARTED\nPROGRAM IN PROGRESS']

Breakdown of the Fix

  • re.DOTALL: Ensures the .*? pattern spans across line breaks, capturing everything between BEGIN and PAUSE regardless of newlines.
  • Non-greedy .*?: Makes sure we stop at the first occurrence of PAUSE instead of matching all the way to the last one (which would happen with greedy .*).
  • Parameter order: We now pass the input string as the second argument, flags as the third—so the regex engine actually uses the DOTALL flag.

Optional Adjustment (Trim Whitespace)

If you want to automatically ignore leading/trailing newlines or spaces between BEGIN and PAUSE, tweak the regex to account for whitespace:

matches = re.findall('(?<=BEGIN\s*)(.*?)(?=\s*PAUSE)', input_text, re.DOTALL)

This will skip any spaces or newlines right after BEGIN and right before PAUSE, giving you a cleaner match without needing to strip() later.

内容的提问来源于stack exchange,提问作者Smart Techie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 16:43:13