使用Imap_tools与Mailparser提取Outlook邮件邮箱地址报错排查
解决AttributeError: 'Message' object has no attribute 'read'问题
问题核心:mailparser.parse_from_file_obj() 需要传入可读取的文件对象,但imap_tools返回的msg.obj是email库的Message对象,没有read()方法,导致报错。
两种修复方案:
方案一:将邮件原始内容转为BytesIO对象适配mailparser
from io import BytesIO import re from imap_tools import MailBox, A from mailparser import parse_from_file_obj from bs4 import BeautifulSoup with MailBox('outlook.office365.com').xoauth2('MAILBOX@domain.com', result['access_token'], 'INBOX') as mailbox: for msg in mailbox.fetch(A(seen=True, subject='SUBJECT', from_='EMAIL')): print(msg.date_str, msg.subject) # 把邮件原始字节内容转为可读取的BytesIO对象 email_message = parse_from_file_obj(BytesIO(msg.raw)) soup = BeautifulSoup(email_message.body, "html.parser") text = soup.get_text() # 提取邮箱地址(使用更精准的正则规则) emails = re.findall(r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}', text) print(emails) # 避免空列表索引报错 if emails: target_email = emails[0]
方案二:直接使用imap_tools自带的内容属性(更简洁)
imap_tools的msg对象本身提供了html和text属性,无需额外用mailparser解析:
import re from imap_tools import MailBox, A from bs4 import BeautifulSoup with MailBox('outlook.office365.com').xoauth2('MAILBOX@domain.com', result['access_token'], 'INBOX') as mailbox: for msg in mailbox.fetch(A(seen=True, subject='SUBJECT', from_='EMAIL')): print(msg.date_str, msg.subject) # 优先处理HTML内容,没有则用纯文本 if msg.html: soup = BeautifulSoup(msg.html, "html.parser") text = soup.get_text() else: text = msg.text # 提取邮箱地址 emails = re.findall(r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}', text) print(emails) if emails: target_email = emails[0]
额外提示:
- 替换原有的简单正则为更精准的邮箱匹配规则,减少无效匹配
- 必须判断
emails列表非空再取第一个元素,否则会触发索引越界错误
内容的提问来源于stack exchange,提问作者litenchi
相关产品推荐
相关产品推荐

