You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取Outlook邮件Zip附件内CSV数据的实现疑问

直接读取Outlook邮件Zip附件中的CSV内容(无需本地保存)

问题

已实现Python连接Outlook、定位指定文件夹及获取特定邮件附件的功能,但不清楚如何直接在Python中查看附件内容(现有资料多为保存附件到本地的方法)。

背景

目标数据是Adobe Analytics导出的报表,该报表为CSV文件并压缩为Zip格式作为邮件附件。计划每周自动遍历所有包含该报表的邮件,将所有数据合并为一个DataFrame,整合历史数据与最新周数据后导出文件。

现有代码

#STEP 1---------------------------------------------
#import all methods needed
from pathlib import Path
import win32com.client
import requests
import time
import datetime
import os
import zipfile
from zipfile import ZipFile
import pandas as pd


#STEP 2 --------------------------------------------
#connect to outlook
outlook = win32com.client.Dispatch("Outlook.Application").GetNamespace("MAPI")


#STEP 3 --------------------------------------------
#connect to inbox
inbox = outlook.GetDefaultFolder(6)


#STEP 4 --------------------------------------------
#connect to adobe data reports folder within inbox
adobe_data_reports_folder = inbox.Folders['Cust Insights'].Folders['Adobe data reports']


#STEP 5 --------------------------------------------
#get all messages from adobe reports folder
messages_from_adr_folder = adobe_data_reports_folder.Items


#STEP 6 ---------------------------------------------
#get attachement for a specific message (this is just for testing in real world I'll do this for all messages)
for message in messages_from_adr_folder:
    if message.SentOn.strftime("%d-%m-%y") == '07-12-22':
        attachments = message.Attachments
    else:
        pass


#STEP 7 ----------------------------------------------
#get the content of the attachment

##????????????????????????????

解决方案

可以通过内存字节流处理附件,无需将文件保存到本地。具体实现如下:

1. 补充必要导入

添加io模块用于创建内存字节流:

import io

2. 修改附件处理逻辑

替换原有STEP6和STEP7的代码,实现直接读取Zip附件内的CSV数据并合并:

# 初始化空DataFrame用于存储所有合并数据
all_data = pd.DataFrame()

# 遍历目标文件夹中的邮件
for message in messages_from_adr_folder:
    # 按需求筛选目标邮件(示例为指定日期)
    if message.SentOn.strftime("%d-%m-%y") == '07-12-22':
        # 遍历当前邮件的所有附件
        for att in message.Attachments:
            # 筛选出Zip格式的附件
            if att.FileName.lower().endswith('.zip'):
                # 创建内存字节流,替代本地文件存储
                zip_stream = io.BytesIO()
                # 将附件内容写入字节流
                att.SaveAsFile(zip_stream)
                # 重置字节流指针到起始位置
                zip_stream.seek(0)
                
                # 读取内存中的Zip文件
                with ZipFile(zip_stream, 'r') as zip_ref:
                    # 遍历Zip内的文件,筛选CSV格式
                    for csv_file in zip_ref.namelist():
                        if csv_file.lower().endswith('.csv'):
                            # 直接读取CSV内容到DataFrame
                            df = pd.read_csv(zip_ref.open(csv_file))
                            # 合并到总DataFrame
                            all_data = pd.concat([all_data, df], ignore_index=True)

# 导出合并后的完整数据
all_data.to_csv('合并后的Adobe报表数据.csv', index=False)

关键说明

  • io.BytesIO创建内存缓冲区,避免磁盘IO操作,提升处理效率
  • att.SaveAsFile()支持写入字节流对象,无需指定本地文件路径
  • 通过zip_ref.open()直接读取Zip内的CSV文件,无需解压到本地
  • 循环遍历所有目标邮件和附件,自动合并所有CSV数据到统一DataFrame

内容的提问来源于stack exchange,提问作者Thomas Chamberlain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 05:35:22