使用Python imaplib下载的PDF邮件附件无法打开的问题求助
问题
使用Python的imaplib库下载指定邮件中的PDF附件,代码执行后文件已按预期下载,但所有PDF均无法正常使用:Zathura、Google Chrome无法打开,PdfReader模块也无法识别,且文件体积大于正常大小。
代码示例
import imaplib import base64 import os import email user = 'my email' passwd = 'my password' directory = 'directory I want my downloaded files to go to' mail = imaplib.IMAP4_SSL('imap.gmail.com') mail.login(user, passwd) mail.select('INBOX') res, data = mail.search(None, 'FROM "email I want to download attachments from" SUBJECT "subject I want"') mail_ids = data[0] id_list = mail_ids.split() for num in data[0].split(): res, data = mail.fetch(num, '(RFC822)') raw_email = data[0][1] raw_email_string = raw_email.decode('utf-8') email_message = email.message_from_string(raw_email_string) for part in email_message.walk(): if part.get_content_maintype() == 'multipart': continue if part.get('Content-Disposition') is None: continue fileName = part.get_filename() if bool(fileName): filePath = os.path.join(directory, fileName) if not os.path.isfile(filePath): fp = open(filePath, 'wb') fp.write(part.get_payload(decode=True)) fp.close()
错误提示
Zathura打开时终端输出:
error: cannot recognize version marker
warning: trying to repair broken xref
warning: repairing PDF document
warning: name is too long
warning: ...repeated 1376 times...
error: no objects found
error: could not open document
Chrome打开时提示:
Error
Failed to load PDF document.
原因与修复方案
问题原因
核心问题是错误地将二进制格式的原始邮件内容解码为UTF-8字符串,再用email.message_from_string()解析。邮件的RFC822内容本身是二进制结构,解码为UTF-8会破坏附件的二进制数据(比如PDF的字节流),导致保存的文件损坏。
修复步骤
- 跳过UTF-8解码步骤,直接用
email.message_from_bytes()解析原始二进制邮件内容。 - 保留二进制模式写入文件的逻辑(代码中已实现,无需修改)。
修改后的代码
import imaplib import os import email user = 'my email' passwd = 'my password' directory = 'directory I want my downloaded files to go to' mail = imaplib.IMAP4_SSL('imap.gmail.com') mail.login(user, passwd) mail.select('INBOX') # 替换转义双引号为普通双引号,提升可读性 res, data = mail.search(None, 'FROM "email I want to download attachments from" SUBJECT "subject I want"') mail_ids = data[0] id_list = mail_ids.split() for num in mail_ids.split(): res, data = mail.fetch(num, '(RFC822)') raw_email = data[0][1] # 直接解析二进制邮件内容,避免编码破坏 email_message = email.message_from_bytes(raw_email) for part in email_message.walk(): if part.get_content_maintype() == 'multipart': continue if part.get('Content-Disposition') is None: continue fileName = part.get_filename() if fileName: filePath = os.path.join(directory, fileName) if not os.path.isfile(filePath): # 使用with语句自动管理文件句柄,避免资源泄漏 with open(filePath, 'wb') as fp: fp.write(part.get_payload(decode=True))
额外优化点
- 用
with语句替代手动调用open()和close(),更安全且代码更简洁。 - 将搜索条件中的
"替换为普通双引号,提升代码可读性。
内容的提问来源于stack exchange,提问作者Nickylp
相关产品推荐
相关产品推荐

