Python读取Outlook邮件Zip附件内CSV数据的实现疑问
直接读取Outlook邮件Zip附件中的CSV内容(无需本地保存)
问题
已实现Python连接Outlook、定位指定文件夹及获取特定邮件附件的功能,但不清楚如何直接在Python中查看附件内容(现有资料多为保存附件到本地的方法)。
背景
目标数据是Adobe Analytics导出的报表,该报表为CSV文件并压缩为Zip格式作为邮件附件。计划每周自动遍历所有包含该报表的邮件,将所有数据合并为一个DataFrame,整合历史数据与最新周数据后导出文件。
现有代码
#STEP 1--------------------------------------------- #import all methods needed from pathlib import Path import win32com.client import requests import time import datetime import os import zipfile from zipfile import ZipFile import pandas as pd #STEP 2 -------------------------------------------- #connect to outlook outlook = win32com.client.Dispatch("Outlook.Application").GetNamespace("MAPI") #STEP 3 -------------------------------------------- #connect to inbox inbox = outlook.GetDefaultFolder(6) #STEP 4 -------------------------------------------- #connect to adobe data reports folder within inbox adobe_data_reports_folder = inbox.Folders['Cust Insights'].Folders['Adobe data reports'] #STEP 5 -------------------------------------------- #get all messages from adobe reports folder messages_from_adr_folder = adobe_data_reports_folder.Items #STEP 6 --------------------------------------------- #get attachement for a specific message (this is just for testing in real world I'll do this for all messages) for message in messages_from_adr_folder: if message.SentOn.strftime("%d-%m-%y") == '07-12-22': attachments = message.Attachments else: pass #STEP 7 ---------------------------------------------- #get the content of the attachment ##????????????????????????????
解决方案
可以通过内存字节流处理附件,无需将文件保存到本地。具体实现如下:
1. 补充必要导入
添加io模块用于创建内存字节流:
import io
2. 修改附件处理逻辑
替换原有STEP6和STEP7的代码,实现直接读取Zip附件内的CSV数据并合并:
# 初始化空DataFrame用于存储所有合并数据 all_data = pd.DataFrame() # 遍历目标文件夹中的邮件 for message in messages_from_adr_folder: # 按需求筛选目标邮件(示例为指定日期) if message.SentOn.strftime("%d-%m-%y") == '07-12-22': # 遍历当前邮件的所有附件 for att in message.Attachments: # 筛选出Zip格式的附件 if att.FileName.lower().endswith('.zip'): # 创建内存字节流,替代本地文件存储 zip_stream = io.BytesIO() # 将附件内容写入字节流 att.SaveAsFile(zip_stream) # 重置字节流指针到起始位置 zip_stream.seek(0) # 读取内存中的Zip文件 with ZipFile(zip_stream, 'r') as zip_ref: # 遍历Zip内的文件,筛选CSV格式 for csv_file in zip_ref.namelist(): if csv_file.lower().endswith('.csv'): # 直接读取CSV内容到DataFrame df = pd.read_csv(zip_ref.open(csv_file)) # 合并到总DataFrame all_data = pd.concat([all_data, df], ignore_index=True) # 导出合并后的完整数据 all_data.to_csv('合并后的Adobe报表数据.csv', index=False)
关键说明
io.BytesIO创建内存缓冲区,避免磁盘IO操作,提升处理效率att.SaveAsFile()支持写入字节流对象,无需指定本地文件路径- 通过
zip_ref.open()直接读取Zip内的CSV文件,无需解压到本地 - 循环遍历所有目标邮件和附件,自动合并所有CSV数据到统一DataFrame
内容的提问来源于stack exchange,提问作者Thomas Chamberlain
相关产品推荐
相关产品推荐

