从Pandas单列数据框提取多来源邮件签名的方法探究
Hey there, great question—extracting email signatures across different providers and formats is tricky but totally doable with a mix of HTML parsing and targeted regex. Let’s break this down step by step:
1. Handling HTML-Formatted Emails
Gmail Signatures (Your Known Case)
You’ve already got this covered, but let’s formalize the function for consistency:
from bs4 import BeautifulSoup import html def extract_gmail_html_signature(html_content): soup = BeautifulSoup(html.unescape(html_content), "html.parser") signature_div = soup.find("div", class_="email_signature") return signature_div.get_text(strip=True) if signature_div else None
Non-Gmail HTML Emails
For other HTML-based providers (Outlook, Yahoo, etc.), signatures usually live in predictable spots. Here’s a flexible function that checks common signature patterns:
def extract_general_html_signature(html_content): soup = BeautifulSoup(html.unescape(html_content), "html.parser") # 1. Look for obvious signature/footer tags/classes signature_candidates = soup.find_all( ["footer", "div"], attrs={"class": lambda c: c and ("signature" in c or "footer" in c or "contact" in c)} ) if signature_candidates: return signature_candidates[-1].get_text(strip=True) # 2. Check for content after a horizontal rule (common separator) hr_tag = soup.find("hr") if hr_tag and hr_tag.next_sibling: return hr_tag.next_sibling.get_text(strip=True) # 3. Fallback: Grab the last text block with contact info all_text_blocks = [block.get_text(strip=True) for block in soup.find_all(["p", "div"]) if block.get_text(strip=True)] for block in reversed(all_text_blocks): if any(keyword in block.lower() for keyword in ["phone", "email", "company", "www.", "@"]): return block return None
2. Handling Plain-Text Emails (Regex Solutions)
Plain-text signatures rely on pattern matching—here are three targeted regex approaches you can combine:
Pattern 1: Match Content After a Separator
Many users use lines like ---, ===, or -- to split body and signature:
import re def extract_text_signature_with_separator(text_content): separator_pattern = r"([-=*]{3,}|\b--\b).*$" match = re.search(separator_pattern, text_content, re.DOTALL) return match.group(0).strip() if match else None
Pattern 2: Match Contact Info-Based Signatures
If no separator exists, target the end of the email for names + contact details:
def extract_text_signature_contact_based(text_content): signature_pattern = r""" (?:^|\n) ([A-Z][a-z]+(?:\s[A-Z][a-z]+)+) # Full name (capitalized words) (?:\s+\(?(?:[A-Za-z\s]+)\)?)? # Optional job title \s* (?: (?:Phone|Tel|Mobile):?\s*\d+[- ]?\d+[- ]?\d+| # Phone number [A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}| # Email address www\.[A-Za-z0-9.-]+\.[A-Z|a-z]{2,} # Website ) .*$ # Capture all remaining signature content """ match = re.search(signature_pattern, text_content, re.VERBOSE | re.DOTALL | re.IGNORECASE) return match.group(0).strip() if match else None
Pattern 3: Match Disclaimer-Led Signatures
Corporate emails often start signatures with confidentiality disclaimers:
def extract_text_signature_disclaimer(text_content): disclaimer_pattern = r"(This email is confidential|The information in this email|If you are not the intended recipient).*$" match = re.search(disclaimer_pattern, text_content, re.DOTALL | re.IGNORECASE) return match.group(0).strip() if match else None
Combine these patterns in priority order for best results:
def extract_text_signature(text_content): # Try separator first sig = extract_text_signature_with_separator(text_content) if sig: return sig # Then contact info sig = extract_text_signature_contact_based(text_content) if sig: return sig # Then disclaimer sig = extract_text_signature_disclaimer(text_content) if sig: return sig # Final fallback: Grab the last paragraph paragraphs = [p.strip() for p in text_content.split("\n\n") if p.strip()] return paragraphs[-1] if paragraphs else None
3. Batch Processing in Pandas
Wrap everything into a single function that auto-detects email format, then apply it to your DataFrame:
import pandas as pd def extract_email_signature(email_content): # Quick check if content is HTML if re.search(r"<[^>]+>", email_content): # First try Gmail-specific signature gmail_sig = extract_gmail_html_signature(email_content) if gmail_sig: return gmail_sig # Fallback to general HTML extraction return extract_general_html_signature(email_content) else: # Process plain text return extract_text_signature(email_content) # Apply to your DataFrame df["extracted_signature"] = df["email_content"].apply(extract_email_signature)
4. Key Tips for Improvement
- Tweak Regex: Adjust the regex patterns to match your specific dataset’s signature quirks (e.g., unique job titles, regional phone formats).
- Test Edge Cases: Keep an eye on emails with no signature, or signatures with unusual formatting—add custom logic for these if needed.
- Clean Output: Use
str.strip()andstr.replace()to remove extra whitespace or unwanted characters from extracted signatures.
内容的提问来源于stack exchange,提问作者Christopher Costello

