从.msg文件提取邮箱至TXT时遇UnicodeEncodeError问题求助
解决.msg文件处理的Unicode编码错误及邮箱提取方案
一、先修复当前的编码错误
你遇到的UnicodeEncodeError是因为写入relat.txt时用了系统默认的charmap编码(Windows下默认编码),它无法处理某些特殊字符。另外第二段代码里write()方法不能传encoding参数,属于语法错误。
修正后的基础代码:
path = r'C:\Dev\Canedo\a.msg' # 读取时用Latin-1保证不丢失字符 with open(path, encoding='Latin-1') as mail: mail_contents = mail.read() # 写入时指定utf-8编码,支持所有Unicode字符 with open("relat.txt", "w+", encoding='utf-8') as t: t.write(mail_contents)
用with语句还能自动管理文件句柄,避免资源泄漏。
二、正确提取.msg邮件正文里的邮箱地址
直接把.msg当文本文件读会包含大量邮件格式冗余内容,建议用专门的库解析Outlook邮件:
方法1:用win32com.client(Windows系统,需先安装pywin32)
import win32com.client import re def extract_email_from_msg(msg_path): outlook = win32com.client.Dispatch("Outlook.Application").GetNamespace("MAPI") msg = outlook.OpenSharedItem(msg_path) # 获取纯文本邮件正文 body = msg.Body # 正则匹配邮箱地址 email_pattern = re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b') emails = email_pattern.findall(body) # 返回第一个匹配的邮箱(可按需求调整逻辑) return emails[0] if emails else None # 处理单个.msg文件 msg_path = r'C:\Dev\Canedo\a.msg' email = extract_email_from_msg(msg_path) if email: with open("relat.txt", "w", encoding='utf-8') as f: f.write(email)
方法2:用msg-parser库(跨平台,需先安装)
先执行安装命令:pip install msg-parser
from msg_parser import MsOxMessage import re def extract_email_from_msg(msg_path): msg = MsOxMessage(msg_path) body = msg.get_body() email_pattern = re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b') emails = email_pattern.findall(body) return emails[0] if emails else None msg_path = r'C:\Dev\Canedo\a.msg' email = extract_email_from_msg(msg_path) if email: with open("relat.txt", "w", encoding='utf-8') as f: f.write(email)
三、批量处理多个.msg文件
如果要处理大量文件,可以遍历目标文件夹:
import os import re from msg_parser import MsOxMessage def extract_email_from_msg(msg_path): try: msg = MsOxMessage(msg_path) body = msg.get_body() email_pattern = re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b') emails = email_pattern.findall(body) return emails[0] if emails else "未找到邮箱" except Exception as e: return f"处理失败:{str(e)}" # 遍历目标文件夹 folder_path = r'C:\Dev\Canedo' for filename in os.listdir(folder_path): if filename.endswith('.msg'): msg_path = os.path.join(folder_path, filename) email = extract_email_from_msg(msg_path) # 每个.msg对应一个同名TXT文件 txt_filename = os.path.splitext(filename)[0] + '.txt' with open(os.path.join(folder_path, txt_filename), 'w', encoding='utf-8') as f: f.write(email)
内容的提问来源于stack exchange,提问作者JameBernabe
相关产品推荐
相关产品推荐

