如何从multipart/alternative类型邮件中提取text/plain部分内容?
提取multipart/alternative邮件中的text/plain内容
你已经通过以下Python代码获取了multipart/alternative类型邮件的结构:
# main.py import email from email.iterators import _structure import sys msg = email.message_from_string(sys.stdin.read()) _structure(msg)
测试输入的邮件内容如下:
./main.py <<EOF From: Nathaniel Borenstein <nsb@bellcore.com> To: Ned Freed <ned@innosoft.com> Subject: Formatted text mail MIME-Version: 1.0 Content-Type: multipart/alternative; boundary=boundary42 --boundary42 Content-Type: text/plain; charset=us-ascii ...plain text version of message goes here.... --boundary42 Content-Type: text/richtext .... richtext version of same message goes here ... --boundary42 Content-Type: text/x-whatever .... fanciest formatted version of same message goes here ... --boundary42-- EOF
得到的结构输出为:
multipart/alternative text/plain text/richtext text/x-whatever
提取text/plain内容的实现代码
要提取邮件中的text/plain部分内容,可以遍历邮件的子部分,匹配类型后读取内容,修改后的代码如下:
# main.py import email import sys msg = email.message_from_string(sys.stdin.read()) # 遍历邮件所有子部分 for part in msg.walk(): # 定位非嵌套的text/plain类型部分 if part.get_content_type() == 'text/plain' and not part.is_multipart(): # 获取字符集,默认用us-ascii兜底 charset = part.get_content_charset() or 'us-ascii' # 解码并输出内容 plain_content = part.get_payload(decode=True).decode(charset) print(plain_content.strip()) break # 找到目标内容后停止遍历
代码说明
msg.walk():递归遍历邮件的所有MIME部分,包括嵌套的multipart结构part.get_content_type():判断当前部分的MIME类型,精准匹配text/plainpart.get_payload(decode=True):自动解码内容(处理base64等编码格式),再用对应字符集转成字符串strip():去除内容前后的空白换行,得到干净的文本内容
运行上述修改后的代码,输入测试邮件即可输出:...plain text version of message goes here....
内容的提问来源于stack exchange,提问作者user1424739
相关产品推荐
相关产品推荐

