正则表达式重复匹配:如何精准筛选章节编号并排除无效格式
Solution to Match Valid Chapter Entries
To solve your problem of matching valid chapter entries like "3. Results" or "3.1. Result" while excluding entries like "test 5.", we can craft a regex that targets two key distinctions:
- Valid chapters have a title (non-digit text) immediately following the chapter number.
- Chapter numbers ending in a dot shouldn't be followed by a newline.
Here's the adjusted code with the correct regex:
import re MyStr = """ test 5. 3. Results 3.1. Result 2 3.3.1 Result test test test test""" # The regex to match valid chapter entries matches = re.findall(r'(?<!\S)\d+(?:\.\d+)*\.?(?!\n)\s+(?!\d)\S.*?(?=\s+\d+(?:\.\d+)*\.?(?!\n)|$)', MyStr, re.DOTALL) print(matches)
Output:
['3. Results', '3.1. Result 2', '3.3.1 Result test test test test']
Regex Breakdown:
Let's break down the regex to understand how it works:
(?<!\S): Negative lookbehind to ensure the chapter number is preceded by whitespace or the start of the string (avoids matching numbers embedded in other words).\d+(?:\.\d+)*: Matches the core chapter number (e.g.,3,3.1,3.3.1).\.?(?!\n): Optional dot at the end of the chapter number, with a negative lookahead to ensure the dot isn't immediately followed by a newline (per your requirement).\s+: Matches one or more whitespace characters after the chapter number.(?!\d): Negative lookahead to ensure the next character isn't a digit (this excludes cases like "5. " which is followed by another chapter number3.).\S.*?: Starts matching the title with a non-whitespace character, then continues non-greedily to avoid over-matching.(?=\s+\d+(?:\.\d+)*\.?(?!\n)|$): Positive lookahead to stop matching when we hit the next chapter number or the end of the string.
Flags:
We use re.DOTALL to allow the . character in the regex to match newlines, which is useful if your chapter titles span multiple lines.
内容的提问来源于stack exchange,提问作者Stijn Van Daele
相关产品推荐
相关产品推荐

