基于Python re库提取文本年龄的正则表达式构建及代码卡顿求助
Hey Maria, let's work through this age extraction issue you're having with Python's re library—both crafting a regex that fits your two scenarios and figuring out why your code was stalling.
First, let's build a regex that handles both of your age prefix scenarios while avoiding the stalling caused by excessive backtracking.
Your two cases:
- Case 1: Colon followed by 0+ spaces then the age (e.g.,
character: eightyorcharacter: eighty) - Case 2: Colon followed by space, some words, then space + age (e.g.,
character: a lovely eighty year old grandma)
We'll make this regex precise and efficient:
- Match the character label up to the colon:
[^:]+:(this works even if character names have spaces or special characters) - Handle 0+ spaces after the colon:
\s* - Match any leading words (for Case 2) without triggering excessive backtracking: use a non-greedy non-capturing group
(?:\S+\s+)*? - Capture the text-based age: we'll use a list of common text numbers to avoid false matches, but you can expand this list as needed.
Here's the full working code:
import re # Define all text-based ages you need to match (expand this list for more cases) text_age_options = r'(one|two|three|four|five|six|seven|eight|nine|ten|eleven|twelve|thirteen|fourteen|fifteen|sixteen|seventeen|eighteen|nineteen|twenty|thirty|forty|fifty|sixty|seventy|eighty|ninety|twenty-one|thirty-two|forty-three)' # Build the optimized regex pattern pattern = rf'^[^:]+:\s*(?:\S+\s+)*?{text_age_options}' # Test with your example cases test_cases = [ "character: eighty", "character: eighty", "character: a lovely eighty year old grandma", "Alice: thirty-two" ] for case in test_cases: match = re.search(pattern, case) if match: print(f"Extracted age: {match.group(1)}") else: print(f"No age found in: {case}")
This regex works because:
[^:]+:safely matches everything up to the colon (no greedy over-matching)(?:\S+\s+)*?uses non-greedy matching (*?) to stop as soon as it hits a valid text age, avoiding unnecessary backtracking- The explicit
text_age_optionsensures we only capture actual age words, not random adjectives or nouns.
If your original code was stalling, it's almost certainly due to catastrophic backtracking—a common regex problem when greedy quantifiers (*, +) are used in nested or ambiguous patterns. Here's how to diagnose and fix it:
- Swap greedy quantifiers for non-greedy ones: If you used something like
.*:\s*(.*)eighty, the greedy.*would match everything first, then backtrack character by character to findeighty—this is brutal on long text. Replace greedy*with non-greedy*?wherever possible. - Use
re.DEBUGto inspect matching steps: Addre.DEBUGto yourre.compilecall to watch exactly how the regex engine processes your string. For example:
This will show you where the engine is getting stuck on backtracking.compiled_pattern = re.compile(pattern, re.DEBUG) compiled_pattern.search("character: a lovely eighty year old grandma") - Avoid nested repeating groups: Patterns like
(a+)+or(\w+\s+)*(\w+)*create exponential backtracking. Stick to flat, non-nested groups whenever you can. - Test with longer text: If stalling only happens on lengthy sentences, it's a sign your regex isn't optimized for large inputs. The regex we built above minimizes backtracking, which should resolve this.
- If ages can be uppercase (e.g.,
Eighty), addre.IGNORECASEto your regex call. - If you need to handle multi-word ages like
eighty five(hyphenated or not), adjust thetext_age_optionsto include those (e.g.,eighty-five|eighty five).
内容的提问来源于stack exchange,提问作者Maria

