You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python re库提取文本年龄的正则表达式构建及代码卡顿求助

Hey Maria, let's work through this age extraction issue you're having with Python's re library—both crafting a regex that fits your two scenarios and figuring out why your code was stalling.

Step 1: A Targeted Regex for Your Two Scenarios

First, let's build a regex that handles both of your age prefix scenarios while avoiding the stalling caused by excessive backtracking.

Your two cases:

  • Case 1: Colon followed by 0+ spaces then the age (e.g., character: eighty or character: eighty)
  • Case 2: Colon followed by space, some words, then space + age (e.g., character: a lovely eighty year old grandma)

We'll make this regex precise and efficient:

  1. Match the character label up to the colon: [^:]+: (this works even if character names have spaces or special characters)
  2. Handle 0+ spaces after the colon: \s*
  3. Match any leading words (for Case 2) without triggering excessive backtracking: use a non-greedy non-capturing group (?:\S+\s+)*?
  4. Capture the text-based age: we'll use a list of common text numbers to avoid false matches, but you can expand this list as needed.

Here's the full working code:

import re

# Define all text-based ages you need to match (expand this list for more cases)
text_age_options = r'(one|two|three|four|five|six|seven|eight|nine|ten|eleven|twelve|thirteen|fourteen|fifteen|sixteen|seventeen|eighteen|nineteen|twenty|thirty|forty|fifty|sixty|seventy|eighty|ninety|twenty-one|thirty-two|forty-three)'

# Build the optimized regex pattern
pattern = rf'^[^:]+:\s*(?:\S+\s+)*?{text_age_options}'

# Test with your example cases
test_cases = [
    "character: eighty",
    "character:  eighty",
    "character: a lovely eighty year old grandma",
    "Alice: thirty-two"
]

for case in test_cases:
    match = re.search(pattern, case)
    if match:
        print(f"Extracted age: {match.group(1)}")
    else:
        print(f"No age found in: {case}")

This regex works because:

  • [^:]+: safely matches everything up to the colon (no greedy over-matching)
  • (?:\S+\s+)*? uses non-greedy matching (*?) to stop as soon as it hits a valid text age, avoiding unnecessary backtracking
  • The explicit text_age_options ensures we only capture actual age words, not random adjectives or nouns.
Step 2: Troubleshooting the Stalling Issue

If your original code was stalling, it's almost certainly due to catastrophic backtracking—a common regex problem when greedy quantifiers (*, +) are used in nested or ambiguous patterns. Here's how to diagnose and fix it:

  • Swap greedy quantifiers for non-greedy ones: If you used something like .*:\s*(.*)eighty, the greedy .* would match everything first, then backtrack character by character to find eighty—this is brutal on long text. Replace greedy * with non-greedy *? wherever possible.
  • Use re.DEBUG to inspect matching steps: Add re.DEBUG to your re.compile call to watch exactly how the regex engine processes your string. For example:
    compiled_pattern = re.compile(pattern, re.DEBUG)
    compiled_pattern.search("character: a lovely eighty year old grandma")
    
    This will show you where the engine is getting stuck on backtracking.
  • Avoid nested repeating groups: Patterns like (a+)+ or (\w+\s+)*(\w+)* create exponential backtracking. Stick to flat, non-nested groups whenever you can.
  • Test with longer text: If stalling only happens on lengthy sentences, it's a sign your regex isn't optimized for large inputs. The regex we built above minimizes backtracking, which should resolve this.
Bonus: Making It More Robust
  • If ages can be uppercase (e.g., Eighty), add re.IGNORECASE to your regex call.
  • If you need to handle multi-word ages like eighty five (hyphenated or not), adjust the text_age_options to include those (e.g., eighty-five|eighty five).

内容的提问来源于stack exchange,提问作者Maria

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:20:34