You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式解析文件块:提取BEGIN与EXCEPTION间内容求助

How to Extract Content Between BEGIN and EXCEPTION (Ignoring Blocks Without EXCEPTION)

First, let's break down why your original regex isn't working:

  • Your pattern BEGIN.*^[^BEGIN].*EXCEPTION uses ^[^BEGIN] which in multiline mode matches the start of a line followed by a character that isn't "B"—this doesn't help target the right blocks at all.
  • The greedy .* will match from the first BEGIN all the way to the last EXCEPTION, grouping unrelated blocks together instead of treating them as separate entities.
  • You aren't accounting for the fact that blocks without EXCEPTION end with END;, so we need to explicitly exclude those from our matches.

The Fix: Target Only Blocks With EXCEPTION

We can build a regex that directly extracts the content you need, no extra replace required. It works by:

  1. Starting with BEGIN
  2. Capturing content only if it doesn't hit an END; first (to skip blocks without EXCEPTION)
  3. Stopping at the first EXCEPTION it encounters

Here's the working code example:

import re

# Sample input data matching your structure
data = """BEGIN random content without exception END; 
BEGIN first target content EXCEPTION error1 END; 
BEGIN second target content EXCEPTION error2 END;"""

# Regex to extract content between BEGIN and EXCEPTION (only valid blocks)
matches = re.findall(r'BEGIN\s*((?:(?!END;).)*?)\s*EXCEPTION', data, re.DOTALL)

# Output the cleaned results
for match in matches:
    print(match.strip())  # Strip extra whitespace if needed

Regex Breakdown:

  • BEGIN\s*: Matches the opening BEGIN plus any leading whitespace after it.
  • ((?:(?!END;).)*?): This is the critical part:
    • (?!END;) is a negative lookahead that ensures we don't encounter END; (the marker for blocks without EXCEPTION).
    • (?:...) is a non-capturing group to wrap the lookahead and character match.
    • .*? is a non-greedy match, so it stops at the first EXCEPTION instead of overextending to later blocks.
  • \s*EXCEPTION: Matches the closing EXCEPTION plus any leading whitespace before it.

Alternative: Using replace (If You Prefer)

If you'd rather first match the full BEGIN...EXCEPTION block and then clean up the BEGIN/EXCEPTION parts, this approach works too:

# Match full BEGIN...EXCEPTION blocks
full_blocks = re.findall(r'BEGIN(.*?)EXCEPTION', data, re.DOTALL)

# Clean up whitespace to get the core content
cleaned_content = [block.strip() for block in full_blocks]
print(cleaned_content)

This gives you the exact same result—pick whichever method feels more readable for your use case!

内容的提问来源于stack exchange,提问作者Sarthak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:30:07