Python如何通过IMAP获取UNSEEN未读邮件ID并提取对应邮件内容
可行实现方案
该需求可直接实现,你已编写的邮件解析、附件下载逻辑无需大幅修改,核心仅需将原「按序号范围遍历邮件」的逻辑替换为「遍历UNSEEN搜索返回的邮件ID列表」即可。
核心修改逻辑
你之前执行mail.search(None, '(UNSEEN)')返回的new_mails是字节格式的结果集,其中new_mails[0]是空格分隔的未读邮件ID字节串,参考你给出的输出[b'389 393'],只需先做解码拆分即可得到可直接用于fetch的ID列表:
# 解码拆分未读邮件ID,得到字符串格式的ID列表 unread_mail_ids = new_mails[0].decode().split() # 若需要优先处理最新邮件(ID更大的为新邮件),反转列表即可 unread_mail_ids.reverse()
替换原代码中range(messages, messages-N, -1)生成遍历序列的逻辑,直接遍历上述得到的unread_mail_ids即可拉取所有未读邮件内容。
整合后的完整代码
import os import email from email.header import decode_header # 你原有的邮箱连接、select收件箱、搜索未读邮件逻辑 status, messages = mail.select('Inbox') messages = int(messages[0]) _, new_mails = mail.search(None, '(UNSEEN)') unread_mail_ids = new_mails[0].decode().split() recent_mails = len(unread_mail_ids) print("Total Messages that is New:" , recent_mails) print("Unread mail IDs:", unread_mail_ids) # 原有的clean函数(你代码中用于生成附件文件夹名,需保留你自己的实现) def clean(text): # 此处替换为你自己实现的非法字符清理逻辑即可 return "".join(c for c in text if c.isalnum() or c in (' ', '.', '_')).rstrip() # 遍历所有未读邮件ID,复用你原有的解析逻辑 for mail_id in unread_mail_ids: res, msg_data = mail.fetch(mail_id, "(RFC822)") for response in msg_data: if isinstance(response, tuple): msg = email.message_from_bytes(response[1]) # 解码主题 pre_subject, encoding = decode_header(msg["Subject"])[0] if isinstance(pre_subject, bytes): pre_subject = pre_subject.decode(encoding if encoding else 'utf-8') subject = pre_subject.upper() # 解码发件人 From, encoding = decode_header(msg.get("From"))[0] if isinstance(From, bytes): From = From.decode(encoding if encoding else 'utf-8') print("Subject:", pre_subject) print("From:", From) plain = "" # 处理多段邮件 if msg.is_multipart(): for part in msg.walk(): content_type = part.get_content_type() content_disposition = str(part.get("Content-Disposition")) try: body = part.get_payload(decode=True).decode() except: continue if content_type == "text/plain" and "attachment" not in content_disposition: print(body) plain = body elif "attachment" in content_disposition: filename = part.get_filename() if filename: folder_name = clean(subject) if not os.path.isdir(folder_name): os.mkdir(folder_name) filepath = os.path.join(folder_name, filename) open(filepath, "wb").write(part.get_payload(decode=True)) else: content_type = msg.get_content_type() body = msg.get_payload(decode=True).decode() if content_type == "text/plain": print(body) plain = body print("="*100) # 可选:拉取完成后将当前邮件标记为已读,不需要可注释 # mail.store(mail_id, '+FLAGS', '\\Seen')
注意事项
- 若搜索结果为空(无未读邮件),
unread_mail_ids为空列表,循环不会执行,无额外报错风险 - 原代码中解码主题/发件人时未处理编码为空的场景,上述代码补了默认用
utf-8解码的兼容逻辑,避免部分邮件解码报错 - 默认
fetch操作不会修改邮件的未读状态,若需要拉取后自动标记已读,放开代码末尾的store命令注释即可 - 若需要提取HTML格式的正文,可参照
text/plain的判断逻辑新增text/html分支处理即可
内容的提问来源于stack exchange,提问作者CSAPawn
相关产品推荐
相关产品推荐

