如何用Python正则表达式提取邮件有效正文并移除冗余信息
Python解决方案:提取邮件核心内容(保留主题+问候与敬语间正文)
Hey there! Let's tackle this problem of pulling out the core content from your email threads. The goal is to keep the Subject line, extract the main body text between greetings (Dear/Hi/Hello) and sign-offs (Sincerely/Regards/Thanks etc.), and strip out all redundant fluff like names, emails, addresses, and job titles.
Approach
Here's how we'll break this down:
- Split the raw input text into individual emails using the
Subject:marker as the delimiter. - For each email block:
- Capture the Subject line (and avoid repeating duplicate subjects in the final output).
- Extract the main body content that sits between a greeting (case-insensitive) and a sign-off (case-insensitive).
- Handle edge cases where emails don't have clear greetings or sign-offs (like the third example where the body comes before the greeting, or the fourth with no greeting at all).
- Remove any remaining redundant lines (names, emails, addresses, job titles) using targeted regex patterns.
Python Code Implementation
import re def extract_core_email_content(raw_emails): # Split raw text into individual emails (split before each "Subject:" marker) email_blocks = re.split(r'(?=Subject:)', raw_emails) core_contents = [] seen_subjects = set() # Prevent duplicate subject lines in output # Regex patterns for matching key elements greeting_pattern = re.compile(r'(?:Dear|Hi|Hello)\s*(?:[A-Za-z]+,?)?\s*', re.IGNORECASE) signoff_pattern = re.compile(r'\s*(?:Sincerely|Regards|Thanks|Best Regards|Warm Regards),?\s*', re.IGNORECASE) # Pattern to catch redundant lines: job titles, emails, addresses, isolated names/companies redundant_pattern = re.compile( r'^[A-Za-z\s]+(?:Director|Manager|Mr\.|Ms\.|Mrs\.)$|' r'^[a-zA-Z0-9_.+-]+@[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+$|' r'^\d+\s+[A-Za-z\s]+(?:st|rd|th|ave|blvd|dr)\.?$|' r'^[A-Za-z\s]+$', re.MULTILINE | re.IGNORECASE ) for block in email_blocks: if not block.strip(): continue # Extract and track the subject line subject_match = re.search(r'Subject:\s*(.*)$', block, re.MULTILINE) if subject_match: subject = subject_match.group(1).strip() if subject not in seen_subjects: core_contents.append(f"Subject: {subject}") seen_subjects.add(subject) # Isolate the body by removing the subject line first body_candidate = re.sub(r'Subject:\s*(.*)$', '', block, flags=re.MULTILINE).strip() # Split off everything after the sign-off signoff_split = signoff_pattern.split(body_candidate) main_body_part = signoff_split[0].strip() # Remove the greeting from the start of the body main_body_part = greeting_pattern.sub('', main_body_part).strip() # Strip out all remaining redundant lines cleaned_body = redundant_pattern.sub('', main_body_part).strip() # Add cleaned body content if it's not empty if cleaned_body: core_contents.append(cleaned_body) # Combine all core content into the final output string return '\n'.join(core_contents) # Test with your sample input sample_input = """Subject: [EXTERNAL] RE: QUERY regarding supplement 73 Hi Roger, Yes, an extension until June 22, 2018 is acceptable. Regards, Loren Subject: [EXTERNAL] RE: QUERY regarding supplement 73 Dear Loren, We had initial discussion with the ABC team us know if you would be able to extend the response due date to June 22, 2018. Best Regards, Mr. Roger Global Director roger@abc.com 78 Ford st. Subject: [EXTERNAL] RE: QUERY regarding supplement 73 responding by June 15, 2018.check email for updates Hello, John Doe Senior Director john.doe@pqr.com Subject: [EXTERNAL] RE: QUERY regarding supplement 73 Please refer to your January 12, 2018 data containing labeling supplements to add text regarding this symptom. We are currently reviewing your supplements and have made additional edits to your label. Feel free to contact me with any questions. Warm Regards, Mr. Roger Global Director roger@abc.com 78 Ford st. Center for Research Office of New Discoveries Food and Drug Administration Loren@mno.com""" # Generate and print the result result = extract_core_email_content(sample_input) print(result)
Key Code Details
- Email Splitting: Using
(?=Subject:)in the split ensures we keep each email block intact (splits before the Subject marker instead of cutting it out). - Case Insensitivity: All greeting and sign-off patterns ignore case, so they handle variations like
hi,HELLO,Best regardsseamlessly. - Redundancy Removal: The
redundant_patterntargets common non-content lines, including job titles, email addresses, physical addresses, and isolated names/company names. - Duplicate Subject Prevention: A set tracks already seen subjects to avoid repeating the same Subject line multiple times in the output.
Expected Output
Running the code with your sample input will produce exactly the output you requested:
Subject: [EXTERNAL] RE: QUERY regarding supplement 73 Yes, an extension until June 22, 2018 is acceptable. We had initial discussion with the ABC team us know if you would be able to extend the response due date to June 22, 2018. responding by June 15, 2018.check email for updates Please refer to your January 12, 2018 data containing labeling supplements to add text regarding this symptom. We are currently reviewing your supplements and have made additional edits to your label. Feel free to contact me with any questions.
内容的提问来源于stack exchange,提问作者Swordsman
相关产品推荐
相关产品推荐

