正则表达式解析文件块:提取BEGIN与EXCEPTION间内容求助
How to Extract Content Between BEGIN and EXCEPTION (Ignoring Blocks Without EXCEPTION)
First, let's break down why your original regex isn't working:
- Your pattern
BEGIN.*^[^BEGIN].*EXCEPTIONuses^[^BEGIN]which in multiline mode matches the start of a line followed by a character that isn't "B"—this doesn't help target the right blocks at all. - The greedy
.*will match from the firstBEGINall the way to the lastEXCEPTION, grouping unrelated blocks together instead of treating them as separate entities. - You aren't accounting for the fact that blocks without
EXCEPTIONend withEND;, so we need to explicitly exclude those from our matches.
The Fix: Target Only Blocks With EXCEPTION
We can build a regex that directly extracts the content you need, no extra replace required. It works by:
- Starting with
BEGIN - Capturing content only if it doesn't hit an
END;first (to skip blocks withoutEXCEPTION) - Stopping at the first
EXCEPTIONit encounters
Here's the working code example:
import re # Sample input data matching your structure data = """BEGIN random content without exception END; BEGIN first target content EXCEPTION error1 END; BEGIN second target content EXCEPTION error2 END;""" # Regex to extract content between BEGIN and EXCEPTION (only valid blocks) matches = re.findall(r'BEGIN\s*((?:(?!END;).)*?)\s*EXCEPTION', data, re.DOTALL) # Output the cleaned results for match in matches: print(match.strip()) # Strip extra whitespace if needed
Regex Breakdown:
BEGIN\s*: Matches the openingBEGINplus any leading whitespace after it.((?:(?!END;).)*?): This is the critical part:(?!END;)is a negative lookahead that ensures we don't encounterEND;(the marker for blocks withoutEXCEPTION).(?:...)is a non-capturing group to wrap the lookahead and character match..*?is a non-greedy match, so it stops at the firstEXCEPTIONinstead of overextending to later blocks.
\s*EXCEPTION: Matches the closingEXCEPTIONplus any leading whitespace before it.
Alternative: Using replace (If You Prefer)
If you'd rather first match the full BEGIN...EXCEPTION block and then clean up the BEGIN/EXCEPTION parts, this approach works too:
# Match full BEGIN...EXCEPTION blocks full_blocks = re.findall(r'BEGIN(.*?)EXCEPTION', data, re.DOTALL) # Clean up whitespace to get the core content cleaned_content = [block.strip() for block in full_blocks] print(cleaned_content)
This gives you the exact same result—pick whichever method feels more readable for your use case!
内容的提问来源于stack exchange,提问作者Sarthak
相关产品推荐
相关产品推荐

