You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python re模块匹配多种年龄表达并替换为统一token(排除特定词)

Standardizing Age Expressions in Clinical Notes with Python Regex

Got it, let's break down how to solve this clinical note preprocessing task. The goal is to take all those messy, varied age formats and turn them into a consistent [age] year_old token—while making sure we don't accidentally match non-age terms that look similar.

First, Let's List All the Patterns We Need to Cover

From your examples, here's every age variant we need to catch:

  • Space-separated: 30 year old, 30 years old
  • Abbreviated with dots: 30 y.o., 30 y.o (sometimes missing the final dot)
  • Slash-separated: 30 y/o
  • No-space concatenations: 24yearold, 24yearsold
  • Shortened without dots: 30yo

Building the Regex Pattern

We'll use Python's re module with a pattern that handles all these cases, plus safeguards to avoid false matches. Here's the breakdown:

  1. Match the age number: (\d+(?:\.\d+)?) — this catches integers (like 30) and decimals (like 1.5 for infants)
  2. Handle optional space between number and suffix: \s* — accounts for both "30yearold" and "30 year old"
  3. Match all suffix variants: A non-capturing group (?:...) that includes every possible age suffix:
    • year[s]?old → covers yearold and yearsold
    • yo → the short abbreviation
    • y\.?o\.? → covers y.o., y.o, yo. (all common dot variations)
    • y\/o → the slash-separated format
    • year[s]?\s+old → covers year old and years old
  4. Avoid false matches: Add negative lookarounds to ensure we don't match terms like yo-yo or x20yearold:
    • (?<![a-zA-Z]) → ensures the age number isn't preceded by a letter
    • (?![a-zA-Z]) → ensures the suffix isn't followed by a letter

Full Implementation Code

Here's a ready-to-use function with test cases to verify it works:

import re

def standardize_clinical_age(text):
    # Regex pattern with case-insensitive matching and false-match safeguards
    age_pattern = r'(?<![a-zA-Z])(\d+(?:\.\d+)?)\s*(?:year[s]?old|yo|y\.?o\.?|y\/o|year[s]?\s+old)(?![a-zA-Z])'
    
    # Replace matches with the standardized token: [age] year_old
    return re.sub(age_pattern, r'\1 year_old', text, flags=re.IGNORECASE)

# Test with your example cases
test_notes = [
    "Patient is a 30 year old male",
    "24yearold female presented with abdominal pain",
    "The 1.5 y.o. infant was admitted for observation",
    "50 years old patient with chronic hypertension",
    "He is 45yo, no prior surgical history",
    "35 y/o male with a 10-year smoking history",
    "Edge case: Don't match yo-yo toy or x20yearoldtest"
]

for note in test_notes:
    print(f"Original: {note}")
    print(f"Processed: {standardize_clinical_age(note)}\n")

Customizing for Edge Cases

If you have specific non-age terms you need to exclude (like yo-yo), you can tweak the pattern to add a negative lookahead for those. For example, to skip yo-yo, modify the yo part to yo(?!-yo):

age_pattern = r'(?<![a-zA-Z])(\d+(?:\.\d+)?)\s*(?:year[s]?old|yo(?!-yo)|y\.?o\.?|y\/o|year[s]?\s+old)(?![a-zA-Z])'

This should cover all your age formatting needs while keeping false matches to a minimum.

内容的提问来源于stack exchange,提问作者shaun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:13:31