如何用Python读取并保存HTML及纯文本格式邮件?
从邮箱读取并处理HTML/纯文本邮件的实现方案
一、读取邮件并判断格式
用Python的imaplib可以对接Gmail或其他支持IMAP的邮箱,先获取邮件内容并区分HTML/纯文本格式:
import imaplib import email from email.header import decode_header # 连接Gmail IMAP服务器(其他邮箱替换为对应IMAP地址) imap = imaplib.IMAP4_SSL("imap.gmail.com") # 注意:Gmail开启2FA需用应用密码,而非登录密码 imap.login("your_email@gmail.com", "your_app_password") imap.select("INBOX") # 搜索未读邮件,可修改条件(比如"ALL"搜索所有邮件) status, messages = imap.search(None, "UNSEEN") message_ids = messages[0].split() for msg_id in message_ids: # 抓取原始邮件内容 status, msg_data = imap.fetch(msg_id, "(RFC822)") raw_email = msg_data[0][1] msg = email.message_from_bytes(raw_email) # 提取元数据(发件人、主题) subject, encoding = decode_header(msg["Subject"])[0] if isinstance(subject, bytes): subject = subject.decode(encoding if encoding else "utf-8") sender = msg.get("From") # 初始化内容变量 is_html = False html_content = "" plain_content = "" # 解析多部分邮件 if msg.is_multipart(): for part in msg.walk(): content_type = part.get_content_type() content_disposition = str(part.get("Content-Disposition")) try: body = part.get_payload(decode=True).decode() except: continue # 过滤附件,提取纯文本/HTML内容 if content_type == "text/plain" and "attachment" not in content_disposition: plain_content = body elif content_type == "text/html" and "attachment" not in content_disposition: html_content = body is_html = True else: # 非多部分邮件直接解析 content_type = msg.get_content_type() body = msg.get_payload(decode=True).decode() if content_type == "text/plain": plain_content = body elif content_type == "text/html": html_content = body is_html = True
二、处理并保存HTML邮件(保留布局、图片)
核心是处理邮件中的内嵌图片(通常用cid:引用附件),否则直接保存的HTML在浏览器里看不到图片,以下是两种可行方案:
方案1:将图片转为Base64内嵌到HTML
适合图片较少的邮件,无需额外保存附件:
from bs4 import BeautifulSoup import base64 if is_html: soup = BeautifulSoup(html_content, "html.parser") # 替换邮件内指定术语(示例:把"旧术语"换成"新术语") for element in soup.find_all(text=lambda text: text and "旧术语" in text): element.replace_with(element.replace("旧术语", "新术语")) # 替换cid引用的图片为Base64数据 for img in soup.find_all("img"): src = img.get("src") if src.startswith("cid:"): cid = src.split("cid:")[1] # 找到对应cid的附件 for part in msg.walk(): if part.get("Content-ID") == f"<{cid}>": img_data = part.get_payload(decode=True) img_type = part.get_content_type() img_base64 = base64.b64encode(img_data).decode() img["src"] = f"data:{img_type};base64,{img_base64}" break # 保存HTML文件 with open(f"{subject}.html", "w", encoding="utf-8") as f: f.write(str(soup))
方案2:保存图片到本地,替换cid为本地路径
适合图片较多的邮件,避免HTML文件过大:
from bs4 import BeautifulSoup import os if is_html: soup = BeautifulSoup(html_content, "html.parser") # 替换指定术语 for element in soup.find_all(text=lambda text: text and "旧术语" in text): element.replace_with(element.replace("旧术语", "新术语")) # 创建附件文件夹存放图片 attachments_dir = f"{subject}_attachments" os.makedirs(attachments_dir, exist_ok=True) # 处理内嵌图片 for img in soup.find_all("img"): src = img.get("src") if src.startswith("cid:"): cid = src.split("cid:")[1] for part in msg.walk(): if part.get("Content-ID") == f"<{cid}>": filename = part.get_filename() or f"image_{cid}.png" filepath = os.path.join(attachments_dir, filename) # 保存图片到本地 with open(filepath, "wb") as f: f.write(part.get_payload(decode=True)) # 替换img的src为相对路径 img["src"] = os.path.relpath(filepath, os.getcwd()) break # 保存HTML文件 with open(f"{subject}.html", "w", encoding="utf-8") as f: f.write(str(soup))
三、处理并保存纯文本邮件
直接提取纯文本内容,替换术语后保存,只保留邮件正文:
if not is_html and plain_content: # 替换指定术语 modified_content = plain_content.replace("旧术语", "新术语") # 保存为TXT文件 with open(f"{subject}.txt", "w", encoding="utf-8") as f: f.write(modified_content)
关键注意事项
- Gmail认证:必须开启IMAP服务,开启2FA的账号要使用应用密码登录,普通密码会被拒绝。
- 编码兼容:部分邮件可能用非UTF-8编码,解码时可尝试
chardet库自动检测编码,避免乱码。 - 复杂邮件:有些邮件会包含多个HTML部分,建议优先选择内容长度最长的那个作为主体。
内容的提问来源于stack exchange,提问作者Zee Kay
相关产品推荐
相关产品推荐

