You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从Pandas单列数据框提取多来源邮件签名的方法探究

Hey there, great question—extracting email signatures across different providers and formats is tricky but totally doable with a mix of HTML parsing and targeted regex. Let’s break this down step by step:

Extracting Email Signatures from a Pandas DataFrame

1. Handling HTML-Formatted Emails

Gmail Signatures (Your Known Case)

You’ve already got this covered, but let’s formalize the function for consistency:

from bs4 import BeautifulSoup
import html

def extract_gmail_html_signature(html_content):
    soup = BeautifulSoup(html.unescape(html_content), "html.parser")
    signature_div = soup.find("div", class_="email_signature")
    return signature_div.get_text(strip=True) if signature_div else None

Non-Gmail HTML Emails

For other HTML-based providers (Outlook, Yahoo, etc.), signatures usually live in predictable spots. Here’s a flexible function that checks common signature patterns:

def extract_general_html_signature(html_content):
    soup = BeautifulSoup(html.unescape(html_content), "html.parser")
    
    # 1. Look for obvious signature/footer tags/classes
    signature_candidates = soup.find_all(
        ["footer", "div"], 
        attrs={"class": lambda c: c and ("signature" in c or "footer" in c or "contact" in c)}
    )
    if signature_candidates:
        return signature_candidates[-1].get_text(strip=True)
    
    # 2. Check for content after a horizontal rule (common separator)
    hr_tag = soup.find("hr")
    if hr_tag and hr_tag.next_sibling:
        return hr_tag.next_sibling.get_text(strip=True)
    
    # 3. Fallback: Grab the last text block with contact info
    all_text_blocks = [block.get_text(strip=True) for block in soup.find_all(["p", "div"]) if block.get_text(strip=True)]
    for block in reversed(all_text_blocks):
        if any(keyword in block.lower() for keyword in ["phone", "email", "company", "www.", "@"]):
            return block
    return None

2. Handling Plain-Text Emails (Regex Solutions)

Plain-text signatures rely on pattern matching—here are three targeted regex approaches you can combine:

Pattern 1: Match Content After a Separator

Many users use lines like ---, ===, or -- to split body and signature:

import re

def extract_text_signature_with_separator(text_content):
    separator_pattern = r"([-=*]{3,}|\b--\b).*$"
    match = re.search(separator_pattern, text_content, re.DOTALL)
    return match.group(0).strip() if match else None

Pattern 2: Match Contact Info-Based Signatures

If no separator exists, target the end of the email for names + contact details:

def extract_text_signature_contact_based(text_content):
    signature_pattern = r"""
        (?:^|\n)
        ([A-Z][a-z]+(?:\s[A-Z][a-z]+)+)  # Full name (capitalized words)
        (?:\s+\(?(?:[A-Za-z\s]+)\)?)?    # Optional job title
        \s*
        (?:
            (?:Phone|Tel|Mobile):?\s*\d+[- ]?\d+[- ]?\d+|  # Phone number
            [A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}|  # Email address
            www\.[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}  # Website
        )
        .*$  # Capture all remaining signature content
    """
    match = re.search(signature_pattern, text_content, re.VERBOSE | re.DOTALL | re.IGNORECASE)
    return match.group(0).strip() if match else None

Pattern 3: Match Disclaimer-Led Signatures

Corporate emails often start signatures with confidentiality disclaimers:

def extract_text_signature_disclaimer(text_content):
    disclaimer_pattern = r"(This email is confidential|The information in this email|If you are not the intended recipient).*$"
    match = re.search(disclaimer_pattern, text_content, re.DOTALL | re.IGNORECASE)
    return match.group(0).strip() if match else None

Combine these patterns in priority order for best results:

def extract_text_signature(text_content):
    # Try separator first
    sig = extract_text_signature_with_separator(text_content)
    if sig:
        return sig
    # Then contact info
    sig = extract_text_signature_contact_based(text_content)
    if sig:
        return sig
    # Then disclaimer
    sig = extract_text_signature_disclaimer(text_content)
    if sig:
        return sig
    # Final fallback: Grab the last paragraph
    paragraphs = [p.strip() for p in text_content.split("\n\n") if p.strip()]
    return paragraphs[-1] if paragraphs else None

3. Batch Processing in Pandas

Wrap everything into a single function that auto-detects email format, then apply it to your DataFrame:

import pandas as pd

def extract_email_signature(email_content):
    # Quick check if content is HTML
    if re.search(r"<[^>]+>", email_content):
        # First try Gmail-specific signature
        gmail_sig = extract_gmail_html_signature(email_content)
        if gmail_sig:
            return gmail_sig
        # Fallback to general HTML extraction
        return extract_general_html_signature(email_content)
    else:
        # Process plain text
        return extract_text_signature(email_content)

# Apply to your DataFrame
df["extracted_signature"] = df["email_content"].apply(extract_email_signature)

4. Key Tips for Improvement

  • Tweak Regex: Adjust the regex patterns to match your specific dataset’s signature quirks (e.g., unique job titles, regional phone formats).
  • Test Edge Cases: Keep an eye on emails with no signature, or signatures with unusual formatting—add custom logic for these if needed.
  • Clean Output: Use str.strip() and str.replace() to remove extra whitespace or unwanted characters from extracted signatures.

内容的提问来源于stack exchange,提问作者Christopher Costello

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:48:57