Python解析EML文件时如何解码UTF-8 Hex编码(如=E2=80=93)?
解决email.message_from_file解析EML时的quoted-printable编码问题
嘿,我之前处理EML文件时也踩过这个坑!你看到的=E2=80=93是quoted-printable编码格式,对应短破折号(en dash),其实Python的email模块本身就有完整的解码能力,只是需要用对方法,下面给你两种靠谱的解决方案:
方案一:利用
get_payload的decode参数自动解码
这是最省心的方式,email模块会自动识别编码格式并解码,同时我们要注意获取邮件的字符集(避免乱码):import email # 读取EML文件 with open('your_email.eml', 'rb') as eml_file: msg = email.message_from_file(eml_file) # 处理多部分邮件(大部分邮件都是多部分的) if msg.is_multipart(): for part in msg.walk(): # 只处理纯文本内容,如果你需要HTML可以改成'text/html' if part.get_content_type() == 'text/plain': # decode=True自动解码quoted-printable或base64编码 content_bytes = part.get_payload(decode=True) # 获取邮件指定的字符集,默认用utf-8兜底 charset = part.get_content_charset() or 'utf-8' decoded_content = content_bytes.decode(charset) print(decoded_content) else: # 单部分邮件的处理逻辑 content_bytes = msg.get_payload(decode=True) charset = msg.get_content_charset() or 'utf-8' decoded_content = content_bytes.decode(charset) print(decoded_content)方案二:手动调用
quopri模块解码
如果第一种方案因为某些特殊邮件格式没生效,可以手动用quopri模块解码quoted-printable编码:import email import quopri with open('your_email.eml', 'rb') as eml_file: msg = email.message_from_file(eml_file) # 获取原始编码内容 encoded_content = msg.get_payload() # 解码字节流 decoded_bytes = quopri.decodestring(encoded_content) # 转成字符串,同样注意字符集 decoded_content = decoded_bytes.decode('utf-8') print(decoded_content)
关键注意点
- 不要直接读取
get_payload()的原始内容,一定要加上decode=True,否则只会拿到未解码的=E2=80=93这类字符串; - 务必获取邮件头里指定的字符集(
get_content_charset()),不要硬编码utf-8,避免部分非utf-8编码的邮件出现乱码; - 多部分邮件要通过
walk()遍历所有部分,找到你需要的文本类型(plain或html)。
内容的提问来源于stack exchange,提问作者formicaman
相关产品推荐
相关产品推荐

