如何将Pandas DataFrame中的邮件文本拆分为多行?
解决方案
思路分析
你的邮件内容里,原始正文位于Subject:之后、On ... wrote:之前,回复正文则在On ... wrote:之后。我们可以通过字符串匹配定位分界点,提取对应内容后,将拆分结果展开为同一MessageID下的多行。
代码实现
import pandas as pd # 模拟你的DataFrame data = { "MessageID": [1, 2], "Actual Email": [ """From: Sales, Team sales@abc.com<br>Sent: Monday, April 10, 2023 2:36 PM<br>To: customer,1 customer1@client.com<br>Subject: Some Message Hello Customer1, Welcome and thank you for opting in to receive emails from us Sincerely,<br>Some Name On Tue, Apr 11, 2023 at 8:13 AM Important Customer customer1@client.com<br>wrote: Thanks Sales!<br>-Customer1""", """From: Support Team support@abc.com<br>Sent: Tuesday, April 11, 2023 10:00 AM<br>To: customer2 customer2@client.com<br>Subject: Follow Up Hi Customer2, We hope you're enjoying our service. Let us know if you need help. Best Regards,<br>Support Team On Wed, Apr 12, 2023 at 9:00 AM customer2 customer2@client.com<br>wrote: Got it, thanks for checking in!<br>-Customer2""" ] } df = pd.DataFrame(data) def split_email_content(email): # 替换HTML换行符为普通换行,统一格式 email_clean = email.replace("<br>", "\n") # 定位回复起始的分界点 reply_start = email_clean.find("\nOn ") if reply_start == -1: # 无回复的情况:提取Subject之后的内容,剔除签名 subject_end = email_clean.find("\n\n", email_clean.find("Subject:")) if subject_end == -1: return [email_clean.strip().replace("\n", "")] raw_body = email_clean[subject_end+2:] # 移除签名段 for sig_keyword in ["\n\nSincerely,", "\n\nBest Regards,"]: sig_pos = raw_body.find(sig_keyword) if sig_pos != -1: raw_body = raw_body[:sig_pos] return [raw_body.strip().replace("\n", "")] else: # 拆分原始正文 subject_end = email_clean.find("\n\n", email_clean.find("Subject:")) raw_body = email_clean[subject_end+2:reply_start] # 移除原始正文的签名 for sig_keyword in ["\n\nSincerely,", "\n\nBest Regards,"]: sig_pos = raw_body.find(sig_keyword) if sig_pos != -1: raw_body = raw_body[:sig_pos] raw_body_clean = raw_body.strip().replace("\n", "") # 拆分回复正文 reply_body = email_clean[reply_start+len("\nOn "):] wrote_end = reply_body.find("\n\n") if wrote_end != -1: reply_body = reply_body[wrote_end+2:] reply_body_clean = reply_body.strip().replace("\n", "") return [raw_body_clean, reply_body_clean] # 应用拆分函数,将结果转为列表 df["Actual Email"] = df["Actual Email"].apply(split_email_content) # 展开列表为多行 result_df = df.explode("Actual Email").reset_index(drop=True) print(result_df)
代码说明
- 格式统一:把HTML的
<br>替换为普通换行符,避免格式干扰。 - 分界定位:通过
\nOn匹配回复的起始标识,区分原始正文和回复正文的范围。 - 正文清洗:
- 原始正文:截取
Subject后到回复分界点前的内容,剔除Sincerely,/Best Regards,开头的签名段,最后合并换行符为单行。 - 回复正文:截取回复分界点后的内容,跳过
wrote:后的空行,合并换行符为单行。
- 原始正文:截取
- 多行展开:用
explode将每个MessageID对应的正文列表拆分为独立行。
运行结果
| MessageID | Actual Email |
|---|---|
| 1 | Hello Customer1,Welcome and thank you for opting in to receive emails from us |
| 1 | Thanks Sales!-Customer1 |
| 2 | Hi Customer2,We hope you're enjoying our service. Let us know if you need help. |
| 2 | Got it, thanks for checking in!-Customer2 |
内容的提问来源于stack exchange,提问作者Tpk43
相关产品推荐
相关产品推荐

