You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中提取email.message.Message对象的邮件正文?

获取mbox邮件正文的实现方法

邮件正文常见纯文本(text/plain)和HTML(text/html)两种格式,且邮件分为单部分、多部分两种结构,需针对性处理:

核心处理逻辑

  • 单部分邮件:直接解码获取内容
  • 多部分邮件:遍历所有邮件部分,提取文本类型的内容

完整代码示例

import mailbox
from email import policy
from email.parser import BytesParser

def get_message_body(message):
    # 处理多部分邮件
    if message.is_multipart():
        for part in message.walk():
            content_type = part.get_content_type()
            # 提取纯文本或HTML格式的正文
            if content_type in ['text/plain', 'text/html']:
                charset = part.get_content_charset() or 'utf-8'
                # 解码并处理编码异常
                return part.get_payload(decode=True).decode(charset, errors='replace')
        return ""
    else:
        # 处理单部分邮件
        charset = message.get_content_charset() or 'utf-8'
        return message.get_payload(decode=True).decode(charset, errors='replace')

# 加载mbox文件,使用policy.default保证兼容性
the_mailbox = mailbox.mbox(mbox_fname, factory=lambda f: BytesParser(policy=policy.default).parse(f))

for message in the_mailbox:
    subject = message["subject"]
    content = get_message_body(message)
    # 示例:打印主题和正文
    print(f"主题: {subject}")
    print(f"正文: {content}\n")

关键细节说明

  • policy.default:让邮件解析结果符合现代邮件规范,减少编码和格式兼容问题
  • walk():遍历邮件的所有嵌套部分,确保不会遗漏深层的文本内容
  • decode=True:自动处理邮件的内容传输编码(如base64、quoted-printable)
  • errors='replace':遇到编码无法识别的字符时用�替代,避免程序崩溃,可根据需求改为ignore

内容的提问来源于stack exchange,提问作者canary_in_the_data_mine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 03:11:50