You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式提取邮件有效正文并移除冗余信息

Python解决方案:提取邮件核心内容(保留主题+问候与敬语间正文)

Hey there! Let's tackle this problem of pulling out the core content from your email threads. The goal is to keep the Subject line, extract the main body text between greetings (Dear/Hi/Hello) and sign-offs (Sincerely/Regards/Thanks etc.), and strip out all redundant fluff like names, emails, addresses, and job titles.

Approach

Here's how we'll break this down:

  • Split the raw input text into individual emails using the Subject: marker as the delimiter.
  • For each email block:
    1. Capture the Subject line (and avoid repeating duplicate subjects in the final output).
    2. Extract the main body content that sits between a greeting (case-insensitive) and a sign-off (case-insensitive).
    3. Handle edge cases where emails don't have clear greetings or sign-offs (like the third example where the body comes before the greeting, or the fourth with no greeting at all).
    4. Remove any remaining redundant lines (names, emails, addresses, job titles) using targeted regex patterns.

Python Code Implementation

import re

def extract_core_email_content(raw_emails):
    # Split raw text into individual emails (split before each "Subject:" marker)
    email_blocks = re.split(r'(?=Subject:)', raw_emails)
    core_contents = []
    seen_subjects = set()  # Prevent duplicate subject lines in output

    # Regex patterns for matching key elements
    greeting_pattern = re.compile(r'(?:Dear|Hi|Hello)\s*(?:[A-Za-z]+,?)?\s*', re.IGNORECASE)
    signoff_pattern = re.compile(r'\s*(?:Sincerely|Regards|Thanks|Best Regards|Warm Regards),?\s*', re.IGNORECASE)
    # Pattern to catch redundant lines: job titles, emails, addresses, isolated names/companies
    redundant_pattern = re.compile(
        r'^[A-Za-z\s]+(?:Director|Manager|Mr\.|Ms\.|Mrs\.)$|'
        r'^[a-zA-Z0-9_.+-]+@[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+$|'
        r'^\d+\s+[A-Za-z\s]+(?:st|rd|th|ave|blvd|dr)\.?$|'
        r'^[A-Za-z\s]+$',
        re.MULTILINE | re.IGNORECASE
    )

    for block in email_blocks:
        if not block.strip():
            continue
        
        # Extract and track the subject line
        subject_match = re.search(r'Subject:\s*(.*)$', block, re.MULTILINE)
        if subject_match:
            subject = subject_match.group(1).strip()
            if subject not in seen_subjects:
                core_contents.append(f"Subject: {subject}")
                seen_subjects.add(subject)
        
        # Isolate the body by removing the subject line first
        body_candidate = re.sub(r'Subject:\s*(.*)$', '', block, flags=re.MULTILINE).strip()
        
        # Split off everything after the sign-off
        signoff_split = signoff_pattern.split(body_candidate)
        main_body_part = signoff_split[0].strip()
        
        # Remove the greeting from the start of the body
        main_body_part = greeting_pattern.sub('', main_body_part).strip()
        
        # Strip out all remaining redundant lines
        cleaned_body = redundant_pattern.sub('', main_body_part).strip()
        
        # Add cleaned body content if it's not empty
        if cleaned_body:
            core_contents.append(cleaned_body)
    
    # Combine all core content into the final output string
    return '\n'.join(core_contents)

# Test with your sample input
sample_input = """Subject: [EXTERNAL] RE: QUERY regarding supplement 73
Hi Roger,
Yes, an extension until June 22, 2018 is acceptable.
Regards,
Loren
Subject: [EXTERNAL] RE: QUERY regarding supplement 73
Dear Loren,
We had initial discussion with the ABC team us know if you would be able to extend the response due date to June 22, 2018.
Best Regards,
Mr. Roger
Global Director
roger@abc.com
78 Ford st.
Subject: [EXTERNAL] RE: QUERY regarding supplement 73
responding by June 15, 2018.check email for updates
Hello,
John Doe
Senior Director
john.doe@pqr.com
Subject: [EXTERNAL] RE: QUERY regarding supplement 73
Please refer to your January 12, 2018 data containing labeling supplements to add text regarding this symptom. We are currently reviewing your supplements and have made additional edits to your label. Feel free to contact me with any questions.
Warm Regards,
Mr. Roger
Global Director
roger@abc.com
78 Ford st.
Center for Research
Office of New Discoveries
Food and Drug Administration
Loren@mno.com"""

# Generate and print the result
result = extract_core_email_content(sample_input)
print(result)

Key Code Details

  • Email Splitting: Using (?=Subject:) in the split ensures we keep each email block intact (splits before the Subject marker instead of cutting it out).
  • Case Insensitivity: All greeting and sign-off patterns ignore case, so they handle variations like hi, HELLO, Best regards seamlessly.
  • Redundancy Removal: The redundant_pattern targets common non-content lines, including job titles, email addresses, physical addresses, and isolated names/company names.
  • Duplicate Subject Prevention: A set tracks already seen subjects to avoid repeating the same Subject line multiple times in the output.

Expected Output

Running the code with your sample input will produce exactly the output you requested:

Subject: [EXTERNAL] RE: QUERY regarding supplement 73
Yes, an extension until June 22, 2018 is acceptable.
We had initial discussion with the ABC team us know if you would be able to extend the response due date to June 22, 2018.
responding by June 15, 2018.check email for updates
Please refer to your January 12, 2018 data containing labeling supplements to add text regarding this symptom. We are currently reviewing your supplements and have made additional edits to your label. Feel free to contact me with any questions.

内容的提问来源于stack exchange,提问作者Swordsman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:58:28