You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Django中正确解码Quoted-Printable格式.eml文件以消除=伪影

解决Django中.eml文件Quoted-Printable解码残留=伪影的问题

核心问题分析

你的代码存在重复解码Quoted-Printable内容的错误:Python的email模块调用get_payload(decode=True)时,会自动根据邮件头的Content-Transfer-Encoding完成Quoted-Printable或Base64解码。额外调用quopri.decodestring()会对已解码内容二次处理,导致正常字符被错误解析,出现孤立的=伪影。

修正后的代码

移除二次解码逻辑,直接使用get_payload(decode=True)返回的解码后字节流进行字符集解码:

import os
from email import message_from_bytes
from django.conf import settings

def save_eml_file(self, eml_content):
    file_name = f"{self.date_received.strftime('%Y%m%d_%H%M%S')}_{self.id}.eml"
    file_path = os.path.join(settings.MEDIA_ROOT, "emails", file_name)
    os.makedirs(os.path.dirname(file_path), exist_ok=True)
    email_message = message_from_bytes(eml_content)
    decoded_content = ""

    if email_message.is_multipart():
        for part in email_message.walk():
            if part.get_content_type() in ["text/plain", "text/html"]:
                charset = part.get_content_charset() or "utf-8"
                # get_payload(decode=True)已自动处理Quoted-Printable解码
                payload = part.get_payload(decode=True) or b""
                try:
                    decoded_content += payload.decode(charset, errors="replace")
                except Exception:
                    decoded_content += str(payload)
    else:
        charset = email_message.get_content_charset() or "utf-8"
        payload = email_message.get_payload(decode=True) or b""
        try:
            decoded_content = payload.decode(charset, errors="replace")
        except Exception:
            decoded_content = str(payload)

    with open(file_path, "w", encoding="utf-8") as file:
        file.write(decoded_content)

    self.eml_file_path = f"emails/{file_name}"
    self.save()

额外兜底处理(针对损坏的QP编码)

如果部分邮件本身的Quoted-Printable编码存在损坏(比如缺失=后应有的两位十六进制字符),可在解码为字符串后添加正则替换清理孤立的=:

import re

# 在得到decoded_content后执行:
# 替换孤立的=(=前后非十六进制字符的情况)
decoded_content = re.sub(r'(?<!=[0-9A-Fa-f])=(?![0-9A-Fa-f])', '', decoded_content)
# 清理QP编码中用于换行的行末=
decoded_content = re.sub(r'=\s*', '', decoded_content)

效果验证

修正后,你示例中的错误内容会被正确还原:

  • =ustomers → customers
  • =ill → will
  • =esponsible → responsible

内容的提问来源于stack exchange,提问作者Tofigh Ramazanniya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 23:11:14