如何编写单个正则表达式提取多标题标记间的文本?
Solution to Extract All Header Content with a Single Regex
Got it, let's break this down. Your goal is to pull the content under each of the three headers from your target text using one regex, and get a clean list of results. Here's how to do it:
The Regex Pattern
(?:The first Heading|The second header|third header)\s*(.*?)(?=\s*(?:The first Heading|The second header|third header)|$)
Let's unpack what this does:
(?:The first Heading|The second header|third header): A non-capturing group that matches any of your three header strings. We use non-capturing here because we don't want to include the headers themselves in our results.\s*: Matches any number of whitespace characters (spaces, tabs, etc.) right after the header, so we don't capture extra spaces before the content.(.*?): A non-greedy capturing group that grabs all characters until it hits the next header or the end of the string. The non-greedy?ensures it stops at the first possible match instead of going all the way to the end.(?=\s*(?:...)|$): A positive lookahead that checks if the next part of the string is another header (with optional whitespace before it) or the end of the string. This tells the regex where to stop capturing content for the current header.
Example Usage (Python)
Here's how you'd implement this in Python to get your list of results:
import re text = "The first Heading This is heading 1 data The second header : this is heading 2 data third header this is heading 3 data" pattern = r'(?:The first Heading|The second header|third header)\s*(.*?)(?=\s*(?:The first Heading|The second header|third header)|$)' results = re.findall(pattern, text) print(results)
Output:
['This is heading 1 data', ': this is heading 2 data', 'this is heading 3 data']
That's exactly the list of content corresponding to each header you need.
内容的提问来源于stack exchange,提问作者rawwar
相关产品推荐
相关产品推荐

