IMAP协议中=\r\n换行符如何解码?Python邮件爬取解码问题求解
问题根源
你遇到的=\r\n不是异常换行,是**quoted-printable(QP编码,可打印引号编码)**的标准软换行标识,问题出在你跳过了邮件传输编码的解析步骤,直接把原始载荷当成纯文本处理了。
邮件在传输过程中为了兼容仅支持7位ASCII的旧传输链路,非ASCII内容、HTML内容经常会用QP编码,QP规定单行长度不能超过76字符,超出位置就会插入=\r\n作为软换行标记,该标记在解码时需要被删除,同时拼接换行前后的内容。除了软换行,你拿到的原始内容里大概率还存在=XX形式的十六进制转义字符,比如=E4=BD=A0对应中文「你」,手动清理很容易遗漏。
解决方案
不要手动替换=\r\n,直接用Python标准库的内置能力完成解码即可,两种常用实现方式如下:
方式1:基于email模块全流程解析(推荐)
拉取到原始邮件字节后,先用email模块完成邮件结构和编码的自动解析,不需要自己判断编码类型:
import imaplib from email.parser import BytesParser from email.policy import default # 连接IMAP、搜索邮件的逻辑你已经实现,这里只放核心解码部分 # 假设你用imaplib fetch到的原始邮件字节存在raw_email变量中 msg = BytesParser(policy=default).parsebytes(raw_email) def get_clean_html(msg): if msg.is_multipart(): for part in msg.walk(): if part.get_content_type() == 'text/html': # 自动处理QP、base64等所有标准邮件传输编码 content_bytes = part.get_payload(decode=True) # 自动获取内容编码,兜底用utf-8 charset = part.get_content_charset() or 'utf-8' return content_bytes.decode(charset, errors='replace') else: if msg.get_content_type() == 'text/html': content_bytes = part.get_payload(decode=True) charset = part.get_content_charset() or 'utf-8' return content_bytes.decode(charset, errors='replace') return "" # 调用后得到的html_content已经完全清理了QP编码标识 html_content = get_clean_html(msg)
方式2:手动用quopri模块解码QP内容
如果你已经单独拿到了带QP编码的HTML原始字符串,可以直接用quopri库解码:
import quopri # 假设带编码的原始HTML字符串存在raw_html_str变量中 raw_html_bytes = raw_html_str.encode('utf-8') decoded_bytes = quopri.decodestring(raw_html_bytes) clean_html = decoded_bytes.decode('utf-8', errors='replace')
内容的提问来源于stack exchange,提问作者BartMolk
相关产品推荐
相关产品推荐

