如何用Python re模块匹配多种年龄表达并替换为统一token(排除特定词)
Got it, let's break down how to solve this clinical note preprocessing task. The goal is to take all those messy, varied age formats and turn them into a consistent [age] year_old token—while making sure we don't accidentally match non-age terms that look similar.
First, Let's List All the Patterns We Need to Cover
From your examples, here's every age variant we need to catch:
- Space-separated:
30 year old,30 years old - Abbreviated with dots:
30 y.o.,30 y.o(sometimes missing the final dot) - Slash-separated:
30 y/o - No-space concatenations:
24yearold,24yearsold - Shortened without dots:
30yo
Building the Regex Pattern
We'll use Python's re module with a pattern that handles all these cases, plus safeguards to avoid false matches. Here's the breakdown:
- Match the age number:
(\d+(?:\.\d+)?)— this catches integers (like 30) and decimals (like 1.5 for infants) - Handle optional space between number and suffix:
\s*— accounts for both "30yearold" and "30 year old" - Match all suffix variants: A non-capturing group
(?:...)that includes every possible age suffix:year[s]?old→ coversyearoldandyearsoldyo→ the short abbreviationy\.?o\.?→ coversy.o.,y.o,yo.(all common dot variations)y\/o→ the slash-separated formatyear[s]?\s+old→ coversyear oldandyears old
- Avoid false matches: Add negative lookarounds to ensure we don't match terms like
yo-yoorx20yearold:(?<![a-zA-Z])→ ensures the age number isn't preceded by a letter(?![a-zA-Z])→ ensures the suffix isn't followed by a letter
Full Implementation Code
Here's a ready-to-use function with test cases to verify it works:
import re def standardize_clinical_age(text): # Regex pattern with case-insensitive matching and false-match safeguards age_pattern = r'(?<![a-zA-Z])(\d+(?:\.\d+)?)\s*(?:year[s]?old|yo|y\.?o\.?|y\/o|year[s]?\s+old)(?![a-zA-Z])' # Replace matches with the standardized token: [age] year_old return re.sub(age_pattern, r'\1 year_old', text, flags=re.IGNORECASE) # Test with your example cases test_notes = [ "Patient is a 30 year old male", "24yearold female presented with abdominal pain", "The 1.5 y.o. infant was admitted for observation", "50 years old patient with chronic hypertension", "He is 45yo, no prior surgical history", "35 y/o male with a 10-year smoking history", "Edge case: Don't match yo-yo toy or x20yearoldtest" ] for note in test_notes: print(f"Original: {note}") print(f"Processed: {standardize_clinical_age(note)}\n")
Customizing for Edge Cases
If you have specific non-age terms you need to exclude (like yo-yo), you can tweak the pattern to add a negative lookahead for those. For example, to skip yo-yo, modify the yo part to yo(?!-yo):
age_pattern = r'(?<![a-zA-Z])(\d+(?:\.\d+)?)\s*(?:year[s]?old|yo(?!-yo)|y\.?o\.?|y\/o|year[s]?\s+old)(?![a-zA-Z])'
This should cover all your age formatting needs while keeping false matches to a minimum.
内容的提问来源于stack exchange,提问作者shaun

