如何在Bash中从Thunderbird的Inbox文件提取特定邮件内容?
处理Thunderbird MBOX格式Inbox文件的方法
Thunderbird的标准Inbox文件是MBOX格式,这是一种传统邮件存储格式,所有邮件按顺序拼接,每封邮件以From (注意开头是From加空格)作为分隔符。直接用less查看自然会显得杂乱,下面给你两种实用的处理方式:
一、用命令行工具快速筛选提取
用mboxgrep工具,专门针对MBOX文件设计,支持按发件人、主题、日期等条件筛选,还能提取纯文本正文,同时解决编码问题。
安装(以Debian/Ubuntu为例)
sudo apt install mboxgrep
常用命令示例
- 筛选发件人包含
example@domain.com的邮件,输出纯文本正文:
mboxgrep -i -E 'From:.*example@domain.com' --text-only Inbox
-i:忽略大小写匹配-E:启用正则表达式匹配--text-only:仅提取邮件纯文本正文
- 筛选主题包含「会议」的邮件,指定UTF-8编码输出:
mboxgrep -i -E 'Subject:.*会议' --text-only --charset=utf-8 Inbox
- 按日期范围筛选(比如2024年5月的邮件):
mboxgrep -E 'Date:.*May 2024' --text-only Inbox
二、用Python脚本灵活处理
如果需要自定义复杂逻辑(比如批量导出、多条件组合筛选),用Python自带的mailbox模块即可,无需额外安装依赖。
示例脚本
import mailbox from email import policy from email.parser import BytesParser def extract_plain_text(msg): # 递归提取邮件纯文本正文 if msg.is_multipart(): for part in msg.walk(): if part.get_content_type() == 'text/plain': return part.get_content(decode=True).decode('utf-8', errors='replace') else: if msg.get_content_type() == 'text/plain': return msg.get_content(decode=True).decode('utf-8', errors='replace') return '' # 打开MBOX文件 mbox = mailbox.mbox('Inbox', create=False) # 遍历邮件并筛选 for raw_msg in mbox: # 解析邮件头,自动处理编码问题 parsed_msg = BytesParser(policy=policy.default).parsebytes(raw_msg.as_bytes()) sender = parsed_msg['From'] subject = parsed_msg['Subject'] date = parsed_msg['Date'] body = extract_plain_text(parsed_msg) # 自定义筛选条件:比如发件人包含指定邮箱 if 'example@domain.com' in str(sender): print(f"发件人: {sender}") print(f"主题: {subject}") print(f"日期: {date}") print("正文:") print(body) print("-"*50) mbox.close()
脚本说明
- 用
policy.default解析邮件,自动兼容UTF-8等编码格式 extract_plain_text函数遍历邮件的MIME结构,精准提取纯文本内容- 可根据需求修改筛选逻辑(比如日期范围、关键词组合匹配)
补充:用less正确查看编码
如果只是临时查看Inbox内容,指定UTF-8编码即可:
LESSCHARSET=utf-8 less Inbox
不过这种方式仅能改善编码显示,无法实现结构化筛选。
内容的提问来源于stack exchange,提问作者CloudWatcher
相关产品推荐
相关产品推荐

