You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何通过IMAP获取UNSEEN未读邮件ID并提取对应邮件内容

可行实现方案

该需求可直接实现,你已编写的邮件解析、附件下载逻辑无需大幅修改,核心仅需将原「按序号范围遍历邮件」的逻辑替换为「遍历UNSEEN搜索返回的邮件ID列表」即可。

核心修改逻辑

你之前执行mail.search(None, '(UNSEEN)')返回的new_mails是字节格式的结果集,其中new_mails[0]是空格分隔的未读邮件ID字节串,参考你给出的输出[b'389 393'],只需先做解码拆分即可得到可直接用于fetch的ID列表:

# 解码拆分未读邮件ID,得到字符串格式的ID列表
unread_mail_ids = new_mails[0].decode().split()
# 若需要优先处理最新邮件(ID更大的为新邮件),反转列表即可
unread_mail_ids.reverse()

替换原代码中range(messages, messages-N, -1)生成遍历序列的逻辑,直接遍历上述得到的unread_mail_ids即可拉取所有未读邮件内容。

整合后的完整代码

import os
import email
from email.header import decode_header

# 你原有的邮箱连接、select收件箱、搜索未读邮件逻辑
status, messages = mail.select('Inbox')
messages = int(messages[0])
_, new_mails = mail.search(None, '(UNSEEN)')
unread_mail_ids = new_mails[0].decode().split()
recent_mails = len(unread_mail_ids)
print("Total Messages that is New:" , recent_mails)
print("Unread mail IDs:", unread_mail_ids)

# 原有的clean函数(你代码中用于生成附件文件夹名,需保留你自己的实现)
def clean(text):
    # 此处替换为你自己实现的非法字符清理逻辑即可
    return "".join(c for c in text if c.isalnum() or c in (' ', '.', '_')).rstrip()

# 遍历所有未读邮件ID,复用你原有的解析逻辑
for mail_id in unread_mail_ids:
    res, msg_data = mail.fetch(mail_id, "(RFC822)")
    for response in msg_data:
        if isinstance(response, tuple):
            msg = email.message_from_bytes(response[1])
            # 解码主题
            pre_subject, encoding = decode_header(msg["Subject"])[0]
            if isinstance(pre_subject, bytes):
                pre_subject = pre_subject.decode(encoding if encoding else 'utf-8')
            subject = pre_subject.upper()
            # 解码发件人
            From, encoding = decode_header(msg.get("From"))[0]
            if isinstance(From, bytes):
                From = From.decode(encoding if encoding else 'utf-8')
            print("Subject:", pre_subject)
            print("From:", From)

            plain = ""
            # 处理多段邮件
            if msg.is_multipart():
                for part in msg.walk():
                    content_type = part.get_content_type()
                    content_disposition = str(part.get("Content-Disposition"))
                    try:
                        body = part.get_payload(decode=True).decode()
                    except:
                        continue
                    if content_type == "text/plain" and "attachment" not in content_disposition:
                        print(body)
                        plain = body
                    elif "attachment" in content_disposition:
                        filename = part.get_filename()
                        if filename:
                            folder_name = clean(subject)
                            if not os.path.isdir(folder_name):
                                os.mkdir(folder_name)
                            filepath = os.path.join(folder_name, filename)
                            open(filepath, "wb").write(part.get_payload(decode=True))
            else:
                content_type = msg.get_content_type()
                body = msg.get_payload(decode=True).decode()
                if content_type == "text/plain":
                    print(body)
                    plain = body
            print("="*100)

    # 可选:拉取完成后将当前邮件标记为已读,不需要可注释
    # mail.store(mail_id, '+FLAGS', '\\Seen')

注意事项

  • 若搜索结果为空(无未读邮件),unread_mail_ids为空列表,循环不会执行,无额外报错风险
  • 原代码中解码主题/发件人时未处理编码为空的场景,上述代码补了默认用utf-8解码的兼容逻辑,避免部分邮件解码报错
  • 默认fetch操作不会修改邮件的未读状态,若需要拉取后自动标记已读,放开代码末尾的store命令注释即可
  • 若需要提取HTML格式的正文,可参照text/plain的判断逻辑新增text/html分支处理即可

内容的提问来源于stack exchange,提问作者CSAPawn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 19:12:18