You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python zipfile模块无法读取压缩包中的XLS文件

解决Outlook压缩附件中XLS文件读取失败的问题

我写了一个处理Outlook邮件附件的程序,能把压缩附件里的CSV和多数Excel文件转成pandas DataFrame,但读取压缩包中的XLS文件时一直报错:"File is Not a Zip File."

原代码如下:

import os
from tempfile import NamedTemporaryFile
from zipfile import ZipFile
import pandas as pd

def isexcel(file):
    #evaluates file path and return true if excel file. otherwise returns false
    extension = os.path.splitext(file.filename)[1]
    if extension in ['.xls','.xlsx','.xlsm','.xlsb','.odf','.odt','.ods']:
        return True 
    else: 
        return False

def zipattach_to_dfs(attachment, extract_fn=None):
    #evaluates zip file attachments and returns dictionary with file name as key and dataframes as values
    df_objects = {}
    with NamedTemporaryFile(delete=False) as tmp:
        attachment.SaveAsFile(tmp.name)
        zf = ZipFile(tmp, mode = 'a')
        for file in zf.infolist():
            key = (f'{file.filename} ("-".join(map(str, file.date_time[:3])))')
            if isexcel(file) ==True:
                temp_df = pd.read_excel(zf.open(file.filename), header=None)
                df_objects.update({key:temp_df})
            elif file.filename.endswith(".csv"):
                temp_df = pd.read_csv(zf.open(file.filename), header=None)
                df_objects.update({key:temp_df})
            else:
                raise NotImplementedError('Unexpected filetype: '+str(file.filename))
    return (df_objects)

错误原因分析

  1. 压缩文件打开模式错误:使用mode='a'(追加模式)会尝试修改刚保存的压缩文件,破坏原文件结构,导致程序识别失败。
  2. XLS文件读取方式不当:老版本XLS是纯二进制格式(非ZIP封装),直接传递zf.open()返回的对象给pd.read_excel,可能出现文件解析异常。
  3. 字符串格式化语法错误:原代码中key的f-string拼接逻辑错误,"-".join(...)未正确嵌入字符串,会导致语法问题。

修正后的代码

import os
from tempfile import NamedTemporaryFile
from zipfile import ZipFile
import pandas as pd
from io import BytesIO

def isexcel(file):
    # 判断是否为Excel文件,统一转小写避免大小写误判
    extension = os.path.splitext(file.filename)[1].lower()
    return extension in ['.xls','.xlsx','.xlsm','.xlsb','.odf','.odt','.ods']

def zipattach_to_dfs(attachment, extract_fn=None):
    df_objects = {}
    with NamedTemporaryFile(delete=False) as tmp:
        # 保存附件到临时文件
        attachment.SaveAsFile(tmp.name)
    
    try:
        # 以只读模式打开压缩文件,避免修改原文件结构
        with ZipFile(tmp.name, mode='r') as zf:
            for file in zf.infolist():
                # 修正key的格式化逻辑,正确拼接日期部分
                date_part = "-".join(map(str, file.date_time[:3]))
                key = f'{file.filename} ({date_part})'
                
                if isexcel(file):
                    # 读取XLS文件内容到BytesIO内存对象,适配二进制格式解析
                    with zf.open(file) as f:
                        excel_content = BytesIO(f.read())
                        temp_df = pd.read_excel(excel_content, header=None)
                    df_objects[key] = temp_df
                elif file.filename.lower().endswith(".csv"):
                    with zf.open(file) as f:
                        temp_df = pd.read_csv(f, header=None)
                    df_objects[key] = temp_df
                else:
                    raise NotImplementedError(f'Unexpected filetype: {file.filename}')
    finally:
        # 强制清理临时文件,避免磁盘残留
        os.unlink(tmp.name)
    
    return df_objects

关键修改点说明

  • 将ZipFile打开模式改为'r',只读模式不会破坏原压缩文件结构,从根源解决“不是ZIP文件”的错误。
  • 对XLS文件,先将内容读取到BytesIO内存对象中再传递给pd.read_excel,确保二进制格式的文件被正确解析。
  • 修正key的字符串拼接逻辑,解决潜在语法错误。
  • 添加finally块,确保临时文件无论是否执行成功都会被删除,避免磁盘冗余。
  • 优化文件扩展名判断逻辑,统一转为小写,避免因文件名大小写导致的误判。

内容的提问来源于stack exchange,提问作者Possdawgers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 16:48:24