You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取并保存HTML及纯文本格式邮件?

从邮箱读取并处理HTML/纯文本邮件的实现方案

一、读取邮件并判断格式

用Python的imaplib可以对接Gmail或其他支持IMAP的邮箱,先获取邮件内容并区分HTML/纯文本格式:

import imaplib
import email
from email.header import decode_header

# 连接Gmail IMAP服务器(其他邮箱替换为对应IMAP地址)
imap = imaplib.IMAP4_SSL("imap.gmail.com")
# 注意:Gmail开启2FA需用应用密码,而非登录密码
imap.login("your_email@gmail.com", "your_app_password")
imap.select("INBOX")

# 搜索未读邮件,可修改条件(比如"ALL"搜索所有邮件)
status, messages = imap.search(None, "UNSEEN")
message_ids = messages[0].split()

for msg_id in message_ids:
    # 抓取原始邮件内容
    status, msg_data = imap.fetch(msg_id, "(RFC822)")
    raw_email = msg_data[0][1]
    msg = email.message_from_bytes(raw_email)

    # 提取元数据(发件人、主题)
    subject, encoding = decode_header(msg["Subject"])[0]
    if isinstance(subject, bytes):
        subject = subject.decode(encoding if encoding else "utf-8")
    sender = msg.get("From")

    # 初始化内容变量
    is_html = False
    html_content = ""
    plain_content = ""

    # 解析多部分邮件
    if msg.is_multipart():
        for part in msg.walk():
            content_type = part.get_content_type()
            content_disposition = str(part.get("Content-Disposition"))
            try:
                body = part.get_payload(decode=True).decode()
            except:
                continue
            # 过滤附件,提取纯文本/HTML内容
            if content_type == "text/plain" and "attachment" not in content_disposition:
                plain_content = body
            elif content_type == "text/html" and "attachment" not in content_disposition:
                html_content = body
                is_html = True
    else:
        # 非多部分邮件直接解析
        content_type = msg.get_content_type()
        body = msg.get_payload(decode=True).decode()
        if content_type == "text/plain":
            plain_content = body
        elif content_type == "text/html":
            html_content = body
            is_html = True

二、处理并保存HTML邮件(保留布局、图片)

核心是处理邮件中的内嵌图片(通常用cid:引用附件),否则直接保存的HTML在浏览器里看不到图片,以下是两种可行方案:

方案1:将图片转为Base64内嵌到HTML

适合图片较少的邮件,无需额外保存附件:

from bs4 import BeautifulSoup
import base64

if is_html:
    soup = BeautifulSoup(html_content, "html.parser")
    # 替换邮件内指定术语(示例:把"旧术语"换成"新术语")
    for element in soup.find_all(text=lambda text: text and "旧术语" in text):
        element.replace_with(element.replace("旧术语", "新术语"))

    # 替换cid引用的图片为Base64数据
    for img in soup.find_all("img"):
        src = img.get("src")
        if src.startswith("cid:"):
            cid = src.split("cid:")[1]
            # 找到对应cid的附件
            for part in msg.walk():
                if part.get("Content-ID") == f"<{cid}>":
                    img_data = part.get_payload(decode=True)
                    img_type = part.get_content_type()
                    img_base64 = base64.b64encode(img_data).decode()
                    img["src"] = f"data:{img_type};base64,{img_base64}"
                    break

    # 保存HTML文件
    with open(f"{subject}.html", "w", encoding="utf-8") as f:
        f.write(str(soup))

方案2:保存图片到本地,替换cid为本地路径

适合图片较多的邮件,避免HTML文件过大:

from bs4 import BeautifulSoup
import os

if is_html:
    soup = BeautifulSoup(html_content, "html.parser")
    # 替换指定术语
    for element in soup.find_all(text=lambda text: text and "旧术语" in text):
        element.replace_with(element.replace("旧术语", "新术语"))

    # 创建附件文件夹存放图片
    attachments_dir = f"{subject}_attachments"
    os.makedirs(attachments_dir, exist_ok=True)

    # 处理内嵌图片
    for img in soup.find_all("img"):
        src = img.get("src")
        if src.startswith("cid:"):
            cid = src.split("cid:")[1]
            for part in msg.walk():
                if part.get("Content-ID") == f"<{cid}>":
                    filename = part.get_filename() or f"image_{cid}.png"
                    filepath = os.path.join(attachments_dir, filename)
                    # 保存图片到本地
                    with open(filepath, "wb") as f:
                        f.write(part.get_payload(decode=True))
                    # 替换img的src为相对路径
                    img["src"] = os.path.relpath(filepath, os.getcwd())
                    break

    # 保存HTML文件
    with open(f"{subject}.html", "w", encoding="utf-8") as f:
        f.write(str(soup))

三、处理并保存纯文本邮件

直接提取纯文本内容,替换术语后保存,只保留邮件正文:

if not is_html and plain_content:
    # 替换指定术语
    modified_content = plain_content.replace("旧术语", "新术语")
    # 保存为TXT文件
    with open(f"{subject}.txt", "w", encoding="utf-8") as f:
        f.write(modified_content)

关键注意事项

  • Gmail认证:必须开启IMAP服务,开启2FA的账号要使用应用密码登录,普通密码会被拒绝。
  • 编码兼容:部分邮件可能用非UTF-8编码,解码时可尝试chardet库自动检测编码,避免乱码。
  • 复杂邮件:有些邮件会包含多个HTML部分,建议优先选择内容长度最长的那个作为主体。

内容的提问来源于stack exchange,提问作者Zee Kay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 00:05:48