如何用正则表达式检测标题前置编号?正则匹配问题排查
Let's break down what's wrong with your current regex and fix it step by step to match prefixes like II -, IV., a), etc.
What's Wrong with Your Original Pattern
Incorrect character group syntax
You used curly braces{}instead of square brackets[]to define sets of possible characters. Curly braces are for quantifiers (e.g.,{1,3}to match 1-3 instances), not for listing characters like spaces or hyphens. This broke how the regex parses your intended separators.Broken Roman numeral matching
Your pattern fragments likeIX?only matchIorIX, andVI{0,3}?is a messy structure that can't properly match valid Roman numerals likeII,III, orVI. It's why you're only getting a singleIinstead of the fullII.Disjointed separator logic
Mixing numeral patterns with separators in an unstructured way caused the regex to stop at the first matching character instead of capturing the full numeral plus separator.
Corrected Regex
Here's a refined pattern that handles the Roman numeral cases you're targeting, plus common separators:
import re # Matches Roman numerals (I-IX) + optional valid separators at string start pattern = r'^(I{1,3}|IV|V|VI{1,3}|IX)(?:[\s\-.,)]|\-\s)?' # Test with your input m = re.search(pattern, "II - Informations supplementaires") if m: print(m.group(0)) # Output: "II -" (full matched prefix) print(m.group(1)) # Output: "II" (just the numeral)
How This Works
Roman numeral capture:
^(I{1,3}|IV|V|VI{1,3}|IX)
This matches all valid single-digit Roman numerals at the start of the string:I{1,3}: Matches I, II, IIIIV: Matches IVV: Matches VVI{1,3}: Matches VI, VII, VIIIIX: Matches IX
Optional separator:
(?:[\s\-.,)]|\-\s)?
This non-capturing group handles the various separators you listed:[\s\-.,)]: Matches a single space, hyphen, dot, comma, or closing parenthesis|\-\s: Also matches a hyphen followed by a space (like yourII -example)?: Makes the separator optional, so it will match just the numeral if no separator exists
Extending to Letter-Based Prefixes
If you also need to match letter-based prefixes like a), C., or C-, expand the pattern to include letter matches:
pattern = r'^([A-Za-z]|I{1,3}|IV|V|VI{1,3}|IX)(?:[\s\-.,)]|\-\s)?' # Test with a letter prefix m = re.search(pattern, "C- Informations supplementaires") print(m.group(0)) # Output: "C-"
内容的提问来源于stack exchange,提问作者dada

