You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame中的邮件文本拆分为多行?

解决方案

思路分析

你的邮件内容里,原始正文位于Subject:之后、On ... wrote:之前,回复正文则在On ... wrote:之后。我们可以通过字符串匹配定位分界点,提取对应内容后,将拆分结果展开为同一MessageID下的多行。

代码实现

import pandas as pd

# 模拟你的DataFrame
data = {
    "MessageID": [1, 2],
    "Actual Email": [
        """From: Sales, Team sales@abc.com<br>Sent: Monday, April 10, 2023 2:36 PM<br>To: customer,1 customer1@client.com<br>Subject: Some Message

Hello Customer1,

Welcome and thank you for opting in to receive emails from us

Sincerely,<br>Some Name

On Tue, Apr 11, 2023 at 8:13 AM Important Customer customer1@client.com<br>wrote:

Thanks Sales!<br>-Customer1""",
        """From: Support Team support@abc.com<br>Sent: Tuesday, April 11, 2023 10:00 AM<br>To: customer2 customer2@client.com<br>Subject: Follow Up

Hi Customer2,

We hope you're enjoying our service. Let us know if you need help.

Best Regards,<br>Support Team

On Wed, Apr 12, 2023 at 9:00 AM customer2 customer2@client.com<br>wrote:

Got it, thanks for checking in!<br>-Customer2"""
    ]
}
df = pd.DataFrame(data)

def split_email_content(email):
    # 替换HTML换行符为普通换行,统一格式
    email_clean = email.replace("<br>", "\n")
    # 定位回复起始的分界点
    reply_start = email_clean.find("\nOn ")
    
    if reply_start == -1:
        # 无回复的情况:提取Subject之后的内容,剔除签名
        subject_end = email_clean.find("\n\n", email_clean.find("Subject:"))
        if subject_end == -1:
            return [email_clean.strip().replace("\n", "")]
        raw_body = email_clean[subject_end+2:]
        # 移除签名段
        for sig_keyword in ["\n\nSincerely,", "\n\nBest Regards,"]:
            sig_pos = raw_body.find(sig_keyword)
            if sig_pos != -1:
                raw_body = raw_body[:sig_pos]
        return [raw_body.strip().replace("\n", "")]
    else:
        # 拆分原始正文
        subject_end = email_clean.find("\n\n", email_clean.find("Subject:"))
        raw_body = email_clean[subject_end+2:reply_start]
        # 移除原始正文的签名
        for sig_keyword in ["\n\nSincerely,", "\n\nBest Regards,"]:
            sig_pos = raw_body.find(sig_keyword)
            if sig_pos != -1:
                raw_body = raw_body[:sig_pos]
        raw_body_clean = raw_body.strip().replace("\n", "")
        
        # 拆分回复正文
        reply_body = email_clean[reply_start+len("\nOn "):]
        wrote_end = reply_body.find("\n\n")
        if wrote_end != -1:
            reply_body = reply_body[wrote_end+2:]
        reply_body_clean = reply_body.strip().replace("\n", "")
        
        return [raw_body_clean, reply_body_clean]

# 应用拆分函数,将结果转为列表
df["Actual Email"] = df["Actual Email"].apply(split_email_content)
# 展开列表为多行
result_df = df.explode("Actual Email").reset_index(drop=True)

print(result_df)

代码说明

  1. 格式统一:把HTML的<br>替换为普通换行符,避免格式干扰。
  2. 分界定位:通过\nOn 匹配回复的起始标识,区分原始正文和回复正文的范围。
  3. 正文清洗:
    • 原始正文:截取Subject后到回复分界点前的内容,剔除Sincerely,/Best Regards,开头的签名段,最后合并换行符为单行。
    • 回复正文:截取回复分界点后的内容,跳过wrote:后的空行,合并换行符为单行。
  4. 多行展开:用explode将每个MessageID对应的正文列表拆分为独立行。

运行结果

MessageIDActual Email
1Hello Customer1,Welcome and thank you for opting in to receive emails from us
1Thanks Sales!-Customer1
2Hi Customer2,We hope you're enjoying our service. Let us know if you need help.
2Got it, thanks for checking in!-Customer2

内容的提问来源于stack exchange,提问作者Tpk43

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 11:18:20