如何用imaplib从@me.com账户邮件中提取并下载PDF附件?
问题
我尝试从xy@me.com的Apple账户邮件中下载PDF附件,此前用imaplib在Gmail账户上操作完全正常,但在@me.com账户使用RFC822执行imaplib.fetch时返回空结果:b'34878 ()'。改用BODY获取时得到了内容结构信息,却无法使用检查multipart邮件、.walk()、.get("Content-Disposition")及.get_payload(decode=True)等方法访问附件。
以下是测试代码及返回结果,请问如何访问其中的application部分以下载0304217.pdf文件?
测试代码
import imaplib import email username = 'username' password = 'password' hostname = 'imap.mail.me.com' port = '993' expeditor= 'expeditor' subj = 'subject of the mail' since="01-Oct-2022" # Connect to the server connection = imaplib.IMAP4_SSL(hostname,port) # Login to my account connection.login(username, password) status, messages = connection.select("INBOX") # total number of emails messages_number = int(messages[0]) print(f'{messages_number=}') # Search for specific emails typ_search, msg_ids_search = connection.search(None, 'SINCE', "{since}",'FROM', f'"{expeditor}"', 'SUBJECT', f'"{subj}"') print(f'{msg_ids_search[0]=}') # Select only the first message to test msg_to_test=emails_ids=[elt.decode() for elt in msg_ids_search[0].split()][0] print(f'{msg_to_test=}') # Fetch the message to test res, msg = connection.fetch(msg_to_test, ("BODY")) print(f'{msg=}')
返回结果
messages_number=35780 msg_ids_search[0]=b'34878 34879' msg_to_test='34878' msg=[b'34878 (BODY ((("text" "html" ("CHARSET" "utf-8") NIL NIL "quoted-printable" 4128 74)("image" "jpeg" ("NAME" "img1") "<img1>" NIL "base64" 129910) "related")("application" "pdf" ("NAME" "0304217.pdf") NIL NIL "base64" 433256) "mixed"))']
解决方法
问题核心是你用BODY参数仅获取了邮件的结构框架,没有拿到实际的邮件内容。要正常解析并提取附件,必须获取完整的邮件原始数据,再通过email模块处理。
修正后的代码示例
import imaplib import email import base64 username = 'username' password = 'password' hostname = 'imap.mail.me.com' port = 993 expeditor= 'expeditor' subj = 'subject of the mail' since="01-Oct-2022" # 连接服务器并登录 connection = imaplib.IMAP4_SSL(hostname, port) connection.login(username, password) connection.select("INBOX") # 搜索指定邮件 typ_search, msg_ids_search = connection.search(None, 'SINCE', since, 'FROM', f'"{expeditor}"', 'SUBJECT', f'"{subj}"') email_ids = [elt.decode() for elt in msg_ids_search[0].split()] for msg_id in email_ids: # 用RFC822获取完整邮件内容 res, msg_data = connection.fetch(msg_id, '(RFC822)') # 解析邮件字节数据为可操作的Email对象 raw_email = msg_data[0][1] msg = email.message_from_bytes(raw_email) # 遍历邮件所有部分查找目标附件 for part in msg.walk(): content_disposition = part.get("Content-Disposition") # 判断是否为附件且文件名匹配 if content_disposition and "attachment" in content_disposition: filename = part.get_filename() if filename == "0304217.pdf": # 解码附件内容并保存 payload = part.get_payload(decode=True) with open(filename, "wb") as f: f.write(payload) print(f"附件 {filename} 已成功保存") # 关闭连接 connection.close() connection.logout()
关键说明
- 获取完整邮件:使用
(RFC822)参数执行fetch,能拿到包含所有内容的原始邮件数据,这是后续解析的基础。之前RFC822返回空可能是搜索条件或邮件状态问题,重新尝试即可。 - 解析邮件对象:通过
email.message_from_bytes()将原始字节转为标准EmailMessage对象,这样就能正常使用.walk()、.get()等方法。 - 定位并保存附件:遍历邮件各部分,通过
Content-Disposition判断附件属性,匹配目标文件名后解码二进制内容并写入文件。
若仍出现RFC822返回空的情况,可尝试用BODY.PEEK[]替代,比如connection.fetch(msg_id, '(BODY.PEEK[TEXT])'),但优先推荐RFC822获取完整内容。
内容的提问来源于stack exchange,提问作者DagnyRoarke
相关产品推荐
相关产品推荐

